Full text
Benchmarking the HERWIG Event Generator in High-Throughput Computing Infrastructure A Preliminary Approach to Profiling the Environmental Impact of HEP Computations Department of Physics and Astronomy, University of Manchester Oxford Road, Manchester, M13 9PL, United Kingdom May 2025 Luis Miguel Villar Padruno Student ID: 10937290 Abstract This study presents a methodology to profile the energy consumption and CO2e emissions of High Energy Physics (HEP) software using Intel’s RAPL interface on a High-Throughput Computing (HTC) cluster. Power measurements from RAPL were compared with the plug-based Prometheus and Green Algorithms calculator. The methodology was applied to benchmark HERWIG 7.3for Drell–Yan processes [1] at center-of-mass (CoM) energies →s= 13 TeV and →s= 100 TeV and dijet final states at →s= 14 TeV. A functional relation of the form y(n)=a·nb+cwas found between the energy consumed and the number of events, n, compatible with linearity within 3ωat 100 TeV and up to 4.7ωat 13 TeV. Additionally, it was found that up to 108simulated events, the power consumption of the CPU running HERWIG remained constant to a good approximation, suggesting that the energy consumption of the CPU, and therefore the emissions, up to a multiplication factor, can be fully parametrised by the duration of the event generation, a finding supported by previous works [2, 3]. Considering the total events recorded by the Large Hadron Collider (LHC) during 2024 [4], the emissions for generating 1016 events were estimated at (6.75 ±1.72) ·104tonnes of CO2e for →s= 13 TeV and (2.86 ±2.17) ·106 tonnes of CO2e for →s= 100 TeV. The large uncertainties, particularly for →s= 100 TeV, arise because the extrapolation extends far beyond the collected data range. Furthermore, allowing the exponent bto vary as a free parameter in the model makes the fit especially sensitive to uncertainties, as small changes in bare exponentially amplified at high values of n. Nevertheless, these results are of great value since they provide an order-of-magnitude estimate of the environmental impact of Monte Carlo event generators at LHC scales.
Contents Contents ........................................................ 2 Acknowledgements .................................................. 3 1 Introduction .................................................... 4 1.1 Background................................................. 4 1.2 ScopeandExtension ............................................ 4 2 Theoretical Background ............................................. 4 2.1 WorldwideLHCComputingGrid ..................................... 4 2.2 HERWIG and Monte Carlo Event-Generators . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2.3 Selected Particle Physics Processes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 3 Tools and Methodology .............................................. 8 3.1 TheNoetherCluster............................................. 8 3.2 SoftwareTools ............................................... 9 3.3 HERWIG ExecutiononNoetherCluster................................... 11 4 Results and Discussion .............................................. 13 4.1 Results of Drell-Yan Simulations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 4.2 ResultsofDijetSimulations ........................................ 19 4.3 GeneralDiscussion............................................. 19 5 Conclusions and Future Work .......................................... 20 5.1 Conclusions................................................. 20 5.2 FutureWork................................................. 20 References ....................................................... 21 Appendices ....................................................... 24 A Plots ........................................................ 24 B More on Thermal Effects ............................................. 29 C Geographical Dependence ............................................ 30 Word count: 2
Acknowledgements I want to express my sincere gratitude to all the professors, postdoctoral researchers, and PhD students who generously shared their time and expertise, contributing significantly to the success of this research. In particular, I would like to thank Prof Caterina Doglioni, Dr Nicolas Enrique Labra, Dr Tobias Fitschen, Dr Zachery Marshall, Dr Ben Parkes, Prof Michael Seymour, Mr Michael Sparks, Mr Siddharth Sule, and the Blackett Support team. I had the opportunity to present my work and results on many occasions, including at WLCG workshops [5, 6], and I would like to thank my supervisors not only for encouraging me to do so but also for identifying and creating the right opportunities for me. During this project, the foundations were also laid for a GitHub organisation affiliated with The University of Manchester, called UofM-Green-Compute. In doing so, we join the growing efforts of the scientific community to make science not only more rigorous but also more responsible. 3
1 Introduction 1.1 Background Computational techniques have always been an important factor in High Energy Physics (HEP) research [7, 8]. These techniques range from machine learning and deep learning, Bayesian and frequentist statistical methods, to Monte Carlo (MC) simulations [7]. These kinds of simulations are the primary focus of this research. In light of upcoming shortterm collider projects at the Large Hadron Collider (LHC), such as the High-Luminosity LHC (HL-LHC) upgrade, the computational workload is expected to become a major bottleneck due to the vast volume of data that will be collected [8, 9]. This problem gets even more serious when considering long-term projects like the Future Circular Collider (FCC) [10]. The growing need to process increasing volumes of data, combined with the ever-present threat of climate change, has made the assessment of energy efficiency and sustainability in HEP software more relevant than ever before. To do so, an effective approach is to focus on machines specialised in running highly complex computations and analyses, high-performance computing facilities, and data centres [11]. These machines alone emit almost 100 megatonnes of CO2e annually [12], which is more than the total emissions from electricity and heat in the United Kingdom in 2021 [13]. This fact, combined with the consideration that the carbon footprint of scientists worldwide ranges from 4 to 25 tonnes of CO2e per capita per year [14–16], makes this a relevant starting point to evaluate the environmental impact of scientific endeavours. The European Organisation for Nuclear Research (CERN) is by no means unfamiliar with this challenge. In fact, with the creation of the Worldwide LHC Computing Grid (WLCG) [11], as will be discussed in Section 2.1, it has become a major force in global-scale computing. Given this context, it is natural that the attention of physicists and computer scientists has been particularly focused on the efficiency and sustainability of the software and hardware used in physics research. This is illustrated, for example, by the growing number of studies and investigations in the field of computational efficiency [17], as well as by efforts to develop algorithms optimised for parallel computation on GPUs [9, 18]. This investigation itself joins the broader efforts aimed at addressing the sustainability challenges of HEP software, based on a key premise: MC event generators are extensively used in particle physics and rely heavily on CPU architectures to run, and experiments such as ATLAS devote approximately 70% of their CPU time and 60% of their disk storage solely to Monte Carlo simulations [19]. Consequently, the energy consumption of algorithms employing these techniques must be carefully analysed, evaluated, and optimised. 1.2 Scope and Extension This work expanded on the research done in the first semester, Sep - Dec 2024, by Ilan Sasson and Luis Miguel Villar Padruno [2], which focused on investigating the carbon footprint of the HERWIG 7.3Monte Carlo event-generator [20] in an Alienware M15 R6 laptop. In this work, however, the simulations were conducted on a High Throughput Computing Cluster, specifically the Noether Cluster at the University of Manchester, which will be introduced in Section 3.1, instead of on a personal laptop. Compared to last semester, this change in hardware environment represents a natural progression of the project, as computing clusters rather than personal computers (PCs) are typically employed for large-scale MC simulations. Additionally, a process beyond Drell-Yan [1] was considered– namely, dijet final states–. In addition to these new implementations, additional profiling tools will be used to compare with the previously utilised CodeCarbon, described in Section 3.2.3. The following sections will present the technical details of HERWIG, Monte Carlo event generators, and the computing infrastructure (HPC and HTC); Section 3, introduced later, will detail the additional software tools used in this study. 2 Theoretical Background 2.1 Worldwide LHC Computing Grid The European Organisation for Nuclear Research (CERN) is a vast research organisation comprising 24 member states, 10 associate member states, two observer states, and many more with cooperation agreements [21]. This high number of involved parties resulted in an essentially global collaboration, which brought about significant logistical challenges that needed to be addressed in addition to the intrinsic limitations of conducting computationally demanding research, such as data volume and resources. To address these issues, the Worldwide LHC Computing Grid (WLCG) project was established 4
in 2001 and finally launched in 2008 [22]. The WLCG is responsible for distributing the data collected by the LHC and providing data storage and analysis resources to physicists across more than 170 sites in 42 countries [11]. By 2017, the grid processed over 100 petabytes of data each year and counted around 500,000 CPU cores distributed globally [23], from which about 20% were provided by CERN. This number has likely increased significantly, considering that by 2024, ATLAS was utilising approximately 700,000 cores [24], according to Zach Marshall, ATLAS Computing Co-Coordinator. The WLCG is a hierarchical system comprising three computing centres: Tier-0, Tier-1, and Tier-2. Tier-0 refers to the CERN Data Centre, which accounts for approximately 20% of the total computing capacity of the grid [25]. It is responsible for storing the raw data from the LHC experiments and performing the initial data processing, which transforms the raw signals into meaningful physical objects before distributing the data to Tier-1 facilities. Fourteen Tier-1 computer centres worldwide are responsible for storing and reprocessing data received by the Tier-0 layer. They also distribute relevant datasets to the Tier-2 layer, comprising centres typically associated with universities and research institutes. These facilities are primarily used for physics analysis, serving as the main sites where meaningful results are extracted from experimental data. In addition, they play a key role in MC simulations and data reconstruction tasks. The WLCG also recognises an unofficial Tier-3 layer comprising smaller, locally managed computing resources. These Tier-3 centres operate independently and have no formal obligations to the grid, typically serving individual researchers or research groups for analysis and development tasks. For this research, a Tier-3 computing cluster called the Noether Cluster was utilised, as will be described in Section 3.1. 2.2 HERWIG and Monte Carlo Event-Generators This section outlines the main features of HERWIG and Monte Carlo event-generators [7], with a focus on specific aspects relevant to the continuation of the thesis presented in [2]. For a more detailed explanation of Monte Carlo event-generators and a historical overview of HERWIG, one may refer to the work presented in [2]. HERWIG is a software tool used in HEP and functions as a general-purpose event generator. Its primary role is to simulate all aspects of particle collisions, except for particle interactions with the detector (this part is typically handled by dedicated software such as GEANT4 [26]). The stages involved in the event generation process were already mentioned in the previous stage of this investigation [2] and are outlined as follows [27, 28]: •Hard Process: The initial high-energy interaction between partons, typically resulting in the production of heavy particles such as top quarks (t), electroweak bosons (W±and Z0), or Higgs bosons (H). •Parton Shower: The evolution of outgoing partons via successive gluons (g) and quark(q)-antiquark(q¯) pairs emissions. In this step, the structure of jets, discussed in Section 2.3.2, is simulated. •Underlying Event: The collection of additional, less energetic interactions accompanying the hard scatter. These contribute significantly to the overall final-state activity in hadron collisions. •Hadronisation : The process in which coloured partons transform into colour-singlet hadrons due to confinement [29]. It is modelled using phenomenological approaches such as the cluster model (used in HERWIG) [30] or string fragmentation [31]. These stages are essential for a comprehensive understanding of the physics described by the Standard Model; moreover, the predictions produced by Monte Carlo event generators are regarded as crucial for testing theoretical models against experimental data [7]. To be able to do all this, HERWIG 7.3interfaces with MG5_aMC@NLO [32], a software package used to generate matrix elements at both leading order (LO) and next-to-leading order (NLO). The term LO denotes calculations that rely only on the simplest Feynman diagrams and provide a rough estimate of the cross section of a given process. In contrast, NLO calculations include additional corrections, resulting in greater precision. MG5_aMC@NLO can also simulate parton showers and combine them with the matrix element calculation, ensuring a more accurate and realistic event description. It is worth noting that ‘MG5’ refers to MadGraph5 [33], which is responsible for calculating matrix elements at LO. The ‘aMC@NLO’ part stands for ‘automated MadGraph with NLO capabilities’, which allows the software to perform NLO computations and handle the matching 1to parton showers. MG5_aMC@NLO often relies on OpenLoops [34], a tool designed for the fast 1Matching combines NLO matrix elements with parton showers for a single process without double counting emissions. On the other hand, Merging combines matrix elements for multiple jets to improve the description of events with varying numbers of jets. 5
and automated computation of one-loop scattering amplitudes, to compute the corrections required for NLO calculations efficiently. It is important to note that HERWIG is not the only general-purpose event generator; other tools, such as PYTHIA and SHERPA [35], are widely used by particle physicists around the globe. Nevertheless, HERWIG is a robust event-generator that performs exceptionally well in QCD-related calculations. This, combined with the advantage of having one of its developers, Prof. Michael Seymour, based in the department at the University, made HERWIG a natural choice for benchmarking in this study. 2.2.1 HERWIG Workflow The HERWIG workflow consists, in broad terms, of creating an .in file (e.g. LHC-Matchbox.in), which contains all the necessary instructions for the subsequent event generation. These instructions include the process to be simulated, the CoM energy →s, the choice between LO or NLO calculations, whether to include an analysis using Rivet [36] (a tool commonly used in HEP for event analysis), and other relevant configuration options. The choice between LO and NLO, as well as the value of →s, significantly affects the execution time. After the creation of the .in file, it is read by HERWIG, using the command Herwig read my_file.in, triggering the first phase of the event generation process, which consists of integrating the matrix element to compute crosssections and generate phase-space information. At the end of this step, a .run file is created. Note that the integration step is not usually parallelised and HERWIG does not include an option to do this in a single command. Nevertheless, Section 4.3 references a method to do so. In this research, however, the integration step used only a single thread, unless otherwise stated. Finally, the newly created .run file is executed to generate the desired number of events using the command Herwig run my_file.run -N number -j cores, where ‘number’ is the number of events the user wishes to generate, and ‘cores’ is the number of CPU threads the user wants HERWIG to utilise during generation. Further discussion on CPU usage and threading is provided in Section 3. Having established an understanding of HERWIG’s workflow, it is now appropriate to discuss the particle physics processes chosen for simulation. 2.3 Selected Particle Physics Processes 2.3.1 Drell-Yan process The Drell-Yan process is a well-established mechanism in particle physics, involving the annihilation of qand q¯in proton–proton collisions [1]. The Drell–Yan process is purely electroweak at LO; however, it has historically been improved through NLO corrections from Quantum Chromodynamics (QCD), which is the theoretical framework that describes the strong interaction among colour-charged particles, such as quarks and gluons. These QCD corrections enhance the precision of theoretical predictions. As a result, the Drell–Yan process has become a valuable tool for probing the internal structure of the proton [37], particularly through its sensitivity to parton distribution functions (PDFs), which describe how the momentum of a proton is distributed among its constituent quarks and gluons during high-energy collisions. A diagram of this process is presented in Figure 1. The quarks’ annihilation produces a resonance consisting of a highly energetic photon (ε) or Z0boson, which later creates a lepton-antilepton pair. 6
P1 P2 ω+ ω→ q q¯ ε Fig. 1. Diagram of the Drell-Yan Process. The two protons of momenta P1and P2, interact via a qof momentum p1=P1x1and an q¯ of momentum p2=P2x2, where x1/2 are the momentum fractions of the partons. Note that the resonance introduced by Drell and Yan was originally a (virtual) ω, but the definition of the process was relaxed to include other massive bosons like Z0. Diagram by Luis Miguel Villar Padruno. It is now possible to investigate this process in a little more detail by focusing on the concept of coupling constant (ϑi). In Quantum Field Theory (QFT), the theoretical framework of particle physics [38, 39], the coupling constant quantifies how strongly particles interact through a given force. For example, the Drell-Yan process, in the context of QFT and HERWIG, is of order two in ϑEW at LO, which corresponds to a term of the form ϑ2 EW . In the NLO HERWIG runs, the order in ϑS (which is zero at LO) increases automatically, incorporating contributions from QCD corrections. These corrections can be of two types: ‘real’ or ‘virtual’, depending on their order in ϑS[27]. If the correction is of order one, it is referred to as a real correction; if it is of order two, it is considered a virtual correction. Examples of these corrections are shown in Figure 2. Note that real corrections involve the emission of particles that, in principle, can be detected. In contrast, virtual corrections involve particles that exist only as internal lines in Feynman diagrams and cannot be directly observed. q q¯ ω+ ω→ g ε ϑS ϑEW ϑEW (a) Example of real correction: O(εS),O(ε2 EW ) q q¯ ω+ ω→ ε ϑ2 S ϑEW ϑEW (b) Example of virtual correction: O(ε2 S),O(ε2 EW ). Fig. 2. Feynman diagrams illustrating representative examples of real and virtual QCD corrections to the Drell–Yan process. Diagram (a) depicts a real correction, in which the final state is modified by the emission of a gluon, leading to a jet that is, in principle, experimentally detectable. In contrast, Diagram (b) shows a virtual correction, where no additional particles appear in the final state, as the correction originates from internal quantum fluctuations. Both corrections involve the strong interaction, characterised by the coupling constant εS. Diagrams by Luis Miguel Villar Padruno. 2.3.2 Dijet final states (pp ↑jj) Dijet final states refer to a group of processes characterised by the production of two collimated sprays of particles, or jets, resulting from the hadronisation of outgoing quarks or gluons. This study focuses on a specific subset: processes mediated by a gluon propagator. The representative Feynman diagram, shown in Figure 3, involves the annihilation of a quark–antiquark pair and the creation of two quarks or gluons via gluon exchange, which subsequently hadronise into jets. q q¯ q¯ q g (a) Feynman diagram of a qand q¯final state which will hadronise and lead to the formation of jets. q q¯ g g g (b) Feynman diagram of a process leading to two gluons in the final state. Fig. 3. Feynman diagrams illustrating the two main final states corresponding to the studied pp →jj process at LO. The initial qand q¯have momenta determined by the parton distribution functions (PDFs) inside the proton, similar to the Drell–Yan process. Diagrams by Luis Miguel Villar Padruno. The Standard Model dijet processes are of particular interest, as they constitute the dominant background in searches for particles beyond the Standard Model (BSM) that decay into dijets [40, 41], including potential dark matter (DM) 7
candidates, as seen in Figure 4, which are hypothetical particles proposed to account for the unseen mass in the universe. In addition to their relevance in BSM searches, dijet events, much like those from Drell–Yan production, play a crucial role in testing QCD and enabling precision measurements of PDFs and ϑS[29]. This is especially relevant in the context of perturbative QCD, which describes the interactions of quarks and gluons at high energies using a series expansion in ϑS[29], and helps in understanding the energy range of its validity. q q¯ q q¯ Z↑ Fig. 4. Feynman diagram of a dijet final state produced via a Z→dark matter mediator. This diagram is of order zero in εS. Therefore, the background of processes like this is investigated when studying dijet final states mediated by gluon propagators. Diagram by Luis Miguel Villar Padruno. In dijet processes with a gluon propagator, the strong coupling constant ϑSis squared at leading order, while the electroweak coupling constant ϑEW does not contribute. HERWIG will then increase the power of ϑSto account for the previously discussed virtual and real QCD corrections, as shown in Figure 5. Unlike the Drell–Yan case, the electroweak coupling constant remains zero at all orders, making this a purely strong interaction even at NLO. q q¯ q¯ q g g ϑSϑS ϑS (a) Example of real correction: O(ε3 S) q q¯ g g g ϑ2 S ϑSϑS (b) Example of virtual correction: O(ε4 S) Fig. 5. Feynman diagrams illustrating representative examples of real and virtual QCD corrections to a dijet final state process involving a gluon propagator. Diagram (a) shows an additional gluon emission, which would be detected as a jet, effectively resulting in a pp →jjj final state. In Diagram (b), a virtual vertex correction is shown; this is a vertex correction, which plays an important role in QCD by contributing to the cancellation of infinities in perturbative calculations [42]. Diagrams by Luis Miguel Villar Padruno. 3 Tools and Methodology 3.1 The Noether Cluster The Noether Cluster is a Tier-3 research computation cluster located in the Physics Department at the University of Manchester. This cluster runs a Linux environment, and the system consists of a login node, used to access the cluster and submit jobs (the concept of jobs will be discussed in detail in Section 3.2.1); job-scheduling nodes, which manage and distribute jobs across the cluster; and working nodes, where the actual computations are performed. These nodes are managed and coordinated through the HTCondor scheduler [43, 44], which will be discussed in Section 3.2.1. It is illustrative to think of the nodes as small computers: they consist of Central Processing Units (CPUs), Random Access Memory (RAM), storage, integrated cooling systems (fans), and sometimes Graphics Processing Units (GPUs). These components draw energy from a Power Supply Unit (PSU), with each node equipped with its own dedicated PSU. Each PSU, in turn, is powered by the Power Distribution Unit (PDU), which is responsible for delivering energy to all the nodes. 3.1.1 Noether’s Central Processing Units The CPUs are the beating heart of most machines [45]. They consist of several cores, which are physical processing units capable of executing tasks independently. Each core can support multiple threads, which are virtual units of execution that allow the CPU to manage several tasks more efficiently by performing parts of them simultaneously. Although the 8
parallelisation capability of CPUs is lower than that of GPUs [46], it is still effectively utilised by most modern software, including HERWIG, the subject of this investigation. The Noether Cluster has two types of CPUs, spread across the working nodes: 1. Intel®Xeon®Gold 5220R CPUs @ 2.20 GHz (96 Threads) 2. Intel®Xeon®CPU E5-2620 v4 @ 2.10 GHz (16 Threads) The first is the most important for this study, as it is the architecture chosen to run the simulations. The Noether Cluster includes eight of these 96-thread CPUs and twenty 16-thread CPUs. In addition, the cluster provides 9 GPUs. It is essential to note that the Intel®Xeon®Gold 5220R processor comes with 24 physical cores and supports two threads per core by default, resulting in a total of 48 threads. Nevertheless, it supports 2-socket (2S) scalability, meaning it can be deployed in dual-socket systems where two identical processors work together within the same node. Note that in this context, a socket refers to a physical CPU slot on the motherboard, each capable of housing one processor, so this configuration effectively doubles the available computing resources, enabling up to 48 cores and 96 threads in a single system. This is why the Noether Cluster reports 96 threads. The Intel®Xeon®Gold 5220R processor has a specified Thermal Design Power (TDP) of 150 W; in a dual-socket (2S) configuration, this results in a combined TDP of 300 W. This number represents the maximum amount of heat the processor is expected to emit under sustained workloads. This value is important as it provides an estimate of the processor’s energy requirements and aids in the design of appropriate cooling solutions. However, it is essential to note that TDP ↓=Power Draw in most cases; the processor’s real-time power consumption can vary depending on workload characteristics and system settings. This is the reason why RAPL, discussed in Section 3.2.2, was used to measure the power consumption of the CPU in almost real-time. 3.2 Software Tools 3.2.1 HTCondor and Jobs HTCondor is a software developed by the Centre for High Throughput Computing at the University of Wisconsin-Madison [47]. Its primary utility is job scheduling in clusters, and it is designed explicitly for High-Throughput Computing (HTC) environments, a computing paradigm optimised to maximise the total amount of work completed over time, rather than the speed of individual tasks. This makes the Noether Cluster technically an HTC facility rather than an HPC, since HPC systems specialise in maximising the speed of single and complex jobs. In the context of HTCondor, a job refers to an instruction to execute a specific command, along with its associated resource requirements, input and output files, and execution environment. Each job typically consists of two files: an executable script, such as my_job.sh, and a submit file, such as my_job.sub. The former contains the instructions to be executed on a compute node, while the latter specifies the resource requirements and defines any input or output file handling. These so-called jobs can generally run across multiple cores, requesting this via the .sub file, however, a single job is confined to one machine [47]. 3.2.2 Running Average Power Limit (RAPL) on Clusters RAPL is an Intel-developed interface that monitors energy consumption directly at the hardware level. The readings are exposed through the Linux file system at the following path: /sys/class/powercap/intel-rapl. The directory provides energy readings at the socket (package) level via sub-directories such as intel-rapl:0 and intel-rapl:1. However, by default, RAPL does not expose energy consumption per core. This means that only the power usage per CPU package can be measured. This presents a limitation in shared computing environments, where multiple users may run jobs on different subsets of cores within the same socket. In such cases, it is impossible to isolate the energy usage of a specific user or process, leading to inaccurate measurements of individual workloads. Nevertheless, this limitation can be mitigated by requesting exclusive access to all available cores on a single node. Concerning the uncertainties in the measurements of RAPL, a percentage error of 5% was adopted, as supported by findings in literature [48]. 9
On the other hand, the CO2e emissions predicted for pp ↑e+e→at →s= 13 TeV at n= 1016 are around half than those estimated in previous works [2]. This apparent discrepancy stems from two distinct sources. First, during the initial semester of this study, the event generation phase was not parallelised, and as a result, the durations were significantly longer than they should have been. Second, in that earlier study, a constraint was imposed requiring the functional relationship to be linear, thus fixing the exponent to b=1. In the case of pp ↑e+e→, the fitted value of bwas less than one; this fact shrunk the extrapolation results. The floating value of the exponent of the fitted model also explains why, in this study, the predicted emissions for →s= 100 TeV are higher, since for all processes at this CoM, the model was superlinear. The estimated CO2e emissions were (6.75 ±1.72) ·104tonnes for →s= 13 TeV and (2.86 ±2.17) ·106tonnes for →s= 100 TeV. It should be noted that these large uncertainties arise from extrapolating to a region eight orders of magnitude beyond the range in which the data were collected. Moreover, relaxing the constraint b=1, used in previous studies [2], significantly impacted the model’s behaviour at high values of n. For instance, when strictly enforcing b=1, the predicted emissions for →s= 100 TeV decrease by one order of magnitude, and the uncertainty by three. Notwithstanding, a perfectly linear fit of the data yields a ϱ2 red << 1for →s= 13 TeV and ϱ2 red >> 1for →s= 100 TeV. Fig. 7. A comparison is shown of the emissions associated with running HERWIG for the process pp →e+e↑at both ↑s=13TeV and ↑s=100TeV. These extrapolations were made based on the models fitted to the data, including the uncertainties in their parameters. The shadowed areas represent the 1ϖconfidence regions. As expected, generating Drell–Yan events at higher energies results in greater CO2e emissions; this occurs because the simulation durations are longer, further confirming that the duration of the simulation is the most influential parameter. Plot by Luis Miguel Villar Padruno. Reassessment of Last Semester’s Results As mentioned above, the work conducted last semester on an 11th Gen Intel Core i7-11800H @ 2.30 GHz CPU did not parallelise HERWIG’s event generation phase. Therefore, in this brief section, an updated set of measurements is presented to ensure completeness and enable a fair comparison with the Noether Cluster. The average power consumption during the generation phase, utilising all available cores on the laptop’s CPU, was measured to be Pgen laptop = 54.8±8.6W. In contrast, the power consumption during the integration phase corresponded to the value reported last semester, Pint laptop = 24.7±1.4W. Additionally, new models of the form CO2e=a·nb+cwere fitted to the data, and therefore new predictions for the CO2e emissions were established, shown in Figure 8. The results of the fits will be shown in the Appendix, as only the final results will be used for comparison. 16
Fig. 8. Predicted CO2e emissions associated with the generation of nevents on an 11th Gen Intel Core i7-11800H @ 2.30 GHz CPU. It should be noted that GPU energy consumption was neglected, as this hardware was not utilised by HERWIG in these simulations. Plot by Luis Miguel Villar Padruno. Regarding the large uncertainties shown in Figure 8, it is important to note that for the laptop studies, only up to 105events were simulated, and therefore the fit is less significant. Nevertheless, the observed trend is consistent with that identified on the Noether Cluster. A seemingly unexpected result was that the number of events generated per unit of CPU energy consumed was higher for the laptop than for the Noether Cluster, indicating that event generation on the laptop is more energy efficient. This is illustrated in Figure 9, and the fitted function was f(n)= a·n 1+c·n,(4.1.2) where aand care redefined constants for this model; this relationship captures a constant efficiency for sufficiently large numbers of simulated events. However, this model is an approximation, since thermal throttling (the CPU slows down to prevent overheating, reducing performance) will eventually occur during sufficiently long simulations. Fig. 9. Comparison of the number of events generated per unit of energy consumption for the Noether Cluster and the laptop. As shown in the plot, the laptop exhibits higher energy efficiency. The reduced chi-squared values of both fits are acceptable, supporting the conclusion that behaviour is well described by the proposed functional relationship. At sufficiently large n, the expected efficiency of the laptop approaches approximately 30,000 events/Wh, whereas the Noether Cluster reaches a maximum efficiency of around 24,000 events/Wh. Note that, for the laptop case, this claim assumes the model can be extrapolated for higher n, but thermal throttling is expected to occur. Plot by Luis Miguel Villar Padruno. 17
After reaching this conclusion, an obvious question arises: why not use laptops for event generation? The answer is simple: time. In Figure 10, a plot of nover the durations of generations is shown. For short generations, n<104, the laptop outperforms the cluster; nevertheless, beyond this point, the Noether Cluster becomes much faster than the laptop. Fig. 10. Plot of the number of events generated over time. Beyond n=10 4, the CPU of the Cluster outperforms that of the laptop. The reduced chi-squared values of both fits are satisfactory, with a mild overfitting for the Noether Cluster model. Plot by Luis Miguel Villar Padruno. A quantity ςcan be defined as the number of events generated per unit time per unit energy. The plot of ςversus the number of events generated is shown in Figure 11. The fit that best described the data was g(n)= a·nb 1+c·nd(4.1.3) where all constants, a,b,cand dare constants of the model. It can be seen that, for low values of n,ςis greater for the laptop. However, the cluster becomes more efficient beyond n= 105events. It is important to consider that beyond n= 105events, no data was taken for the laptop. Therefore, making predictions past this point is unreliable due to thermal throttling. Fig. 11. Plot of ϱ. Beyond n=10 5, the CPU of the cluster outperforms that of the laptop, but the behaviour of the laptop after this point is unclear due to the lack of data points and known thermal effects such as thermal throttling. The reduced chi-squared values of both fits are satisfactory, with a mild overfitting observed in the Noether Cluster model. Plot by Luis Miguel Villar Padruno. 18
A preliminary conclusion that can be drawn from these results is that, surprisingly, the laptop appears to be more energyefficient than the cluster. However, this becomes less surprising when considering that laptops are specifically designed to operate efficiently on battery power. On the other hand, the energy reported by CodeCarbon reflects only the consumption of the CPU and RAM. In the case of the laptop, however, there are several additional components drawing power, such as the display and GPU, and none of these can be considered negligible. Additionally, one must consider that thermal throttling will occur when the CPU is engaged for long enough, which will decrease its efficiency significantly. 4.2 Results of Dijet Simulations The integration phase of the dijet final states, similar to the Drell–Yan case, utilises only a single CPU thread, as reported by HTCondor; consequently, a similar power consumption to that observed for the Drell–Yan process is expected for the dijet case. As anticipated, the measured power consumption was found to be 98.77 ±0.95 W, a value highly compatible with that obtained from the Drell–Yan integration phase, 96.11 ±1.21 W, their compatibility is evidenced by their zscore z↔1.73. It should be noted that these simulations took significantly longer than the Drell–Yan integrations, around 15 hours, reflecting the greater complexity of the merging algorithm compared to the matching. Using Prometheus data, the average power consumption of the integration phase was found to be 146.16 ±0.10 W, which is lower than the 144.51 ±0.30 W measured for the Drell–Yan case; this time, the difference is significant (z↔5.2). Nevertheless, one must remember that the uncertainties are underestimated for the Prometheus data, since no measurement error was considered, only the simulation durations and the statistical uncertainty from repeated measurements. It has been stated that the integration phase of this process alone takes approximately 15 hours of computing time. This raises the question: why is thermal throttling not observed in the results, given the extended runtime? Power consumption measurements show an almost perfectly constant value throughout. This happens because, during the integration phase, only one thread is active, so the CPU temperature remains relatively low compared to scenarios where all cores are fully utilised, like the generation phase. Since the temperature stays within the processor’s thermal operating limits, there is no need for the system to reduce its clock speed, and thus, no thermal throttling occurs. 4.3 General Discussion Event Generation vs Integration It is important to note that, for simple processes such as Drell–Yan, the event generation phase is significantly more impactful in terms of emissions compared to the integration phase (and this is always the case if the integration phase is parallelised). Therefore, understanding the functional relationship between the number of events and the associated emissions is a valuable tool for assessing the environmental impact of the data produced by event generators. In fact, this study alone resulted in approximately 50 kg of CO2e emissions. In relation to the model which describes the emissions/energy as a function of n, given that the exponent, b, changes with the final state considered, it can be inferred that either the physics of the process might affect the functional relationship or the cluster might have experienced thermal effects like throttling when simulating more computationally demanding final states. Regardless, further study is advised in this direction. Nevertheless, it should be noted that, so far, the variation in blies within the margin of error of 5ω, and the maximum deviation corresponds to 0.1. Thus, the relationship can be approximated as linear for sufficiently low n. Laptop vs Cluster Comparing the results of Drell-Yan from this semester and last semester [2], it is possible to see how important the parallelisation of the generation phase is. Even when using more energy, parallelisation does make the process quicker and, in most cases, greener. On the other hand, it was previously stated that the laptop produced more events per unit energy, which is correct. However, the runtime of the simulations is significantly longer on the laptop, making it impractical for physicists to rely on such machines, as was evidenced in Section 4.1.2. Additionally, HPCC and HTCC are designed to operate nearly 24 hours a day, every day of the week, which is difficult to achieve with a personal computer. On the other hand, it was also shown that the number of events per unit time per unit energy, ς, was higher for the cluster for sufficiently large n, but, as previously stated, no conclusions can be drawn for laptops ςbeyond n= 105, so more data must be taken for this regime, nonetheless, thermal effects are expected. Dijets vs Drell-Yan 19
Focusing on the dijet and Drell-Yan processes, it is plausible to arrive at the same conclusion as in previous works [2] [3]: the ‘greenness’ of the software can be quantified by its runtime duration to a good approximation. More complex algorithms, such as the one used for dijets, take a significantly longer amount of time to complete. However, the machine’s power consumption, and consequently, the emission rate, remained roughly constant throughout the run, at least when the CPU was not fully engaged. That said, thermal throttling is expected when the entire CPU works at full capacity for a long time. A recommendation for this long process is to parallelise the integration using HERWIG’s default method described on their webpage [30]. Following their recommendations reduces the duration from 15 hours to around 400 seconds, making it approximately 135 times faster. Due to time constraints, it was impossible to assess the generation phase of the dijet final states. However, the results are expected to be similar to those obtained for Drell-Yan, with key differences coming from the run-time and thermal effects alone. 5 Conclusions and Future Work 5.1 Conclusions This study contributes to the growing field of sustainable scientific computing by applying a methodology to quantify the energy consumption and CO2e emissions associated with MC simulations in HEP. This was achieved using the HERWIG event generator, executed on an Intel®Xeon®Gold 5220R CPU @ 2.20 GHz from the Noether Cluster. Two physical processes were analysed: the Drell–Yan process at →s= 13 TeV and →s= 100 TeV, as well as dijet production involving a gluon propagator at →s= 14 TeV. The results of this study confirmed that the CPU power draw remained stable during the integration phase of the HERWIG workflow to a very good approximation, whilst minor deviations were noticed for the generation phase, likely due to thermal effects. Consequently, emissions can be effectively parametrised by the time the program is run on a given hardware platform. Additional profiling tools, such as The Green Algorithms and Prometheus, were employed to validate the results obtained through CodeCarbon. Both tools supported the fundamental assumptions underlying this approach. In addition, a power-law plus constant model was found to best describe the relationship between energy consumption and number of events generated; moreover, a deviation from linearity was observed to almost 5ωfor pp ↑ϖ+ϖ→at →s= 13 Tev. These findings suggest that estimating the energy consumed might be less straightforward than just considering the relation outlined in Equation (3.2.3) and assuming a constant power consumption. More investigation is encouraged in this direction, especially to asses how thermal effects might affect the results and current paradigms on how we estimate CO2e emissions. As mentioned in the discussion above, around 50 kg of CO2e were already emitted in this study alone. This highlights how easily the environmental impact of our work as physicists can be underestimated and how frequently such emissions occur without researchers being fully aware of them. The extrapolation for the pp ↑e+e→process estimated that 67 kt of CO2e would be emitted for n= 1016 simulated events at →s= 13 TeV. This figure only increases for →s= 100 TeV, for which the predicted emissions are on the order of millions of tonnes of CO2e. These estimates account solely for CPU energy consumption; however, additional emissions must be expected due to cooling systems, memory usage, data transfer, and other infrastructure components. The challenge now is recognising the role of physicists and researchers in addressing climate change and proposing solutions and strategies that allow continued scientific progress without stagnation. 5.2 Future Work 5.2.1 Developments of this Work A valuable addition to this research would be to establish a relationship between power draw and the number of cores engaged and assess how parallelisation saves energy. Additionally, comparing HERWIG ’s behaviour across different CPU architectures will provide insights into how these variations impact the software’s runtime and efficiency. In addition, collaboration with other profiling software providers is expected. These could include HEPScore [57] and Scaphandre [58], which will benefit the project by offering additional points of comparison between datasets and methodologies. The next step in this investigation will be collecting additional data for laptop runs for n>105events to study thermal throttling effects more directly and to allow meaningful comparison with the cluster. Furthermore, benchmarking the generation phase of dijet events remains a priority, and the methodologies presented in this work serve as a solid foundation for pursuing that goal. Additionally, this project has focused solely on HERWIG, so a natural extension of this work would 20
be to assess the efficiency of other Monte Carlo event generators such as PYTHIA and SHERPA, also available on the Noether Cluster. This comparison would provide a valuable reference point for evaluating efficiency between different generators. Finally, this investigation would benefit from a dedicated team taking charge of a study to complete the hardware’s life cycle assessment (LCA). Thus far, only their operational energy use has been considered. At the same time, the production and disposal stages, both key phases in the life cycle of any electronic device, have not yet been addressed. However, work in this direction is underway by Nicolas Labra et al. in the Department of Engineering of the University of Manchester. 5.2.2 Developments of the Field The study of software sustainability and efficiency in HEP is a relatively new but rapidly growing area of research. As computing continues to play a central role in particle physics, there is increasing awareness of the need to balance performance with environmental responsibility. This opens up a wide range of opportunities for interdisciplinary contributions from both physicists and computer scientists. One of the most promising directions involves the development of event generators and computational tools that can use GPU’s parallelisation capabilities. Recent studies have demonstrated the feasibility of modifying algorithms to exploit GPU architectures effectively [9, 18, 59, 60]. At the same time, there are significant efforts in optimising existing algorithms [61, 62]. Integrating sustainability metrics into job scheduling systems like HTCondor is a valuable practice for the future –some systems already incorporate such features, for example, PanDA [63], used in ATLAS– since comprehensive, real-time metrics for power consumption and CO2e emissions would allow the scientific community to make more informed decisions regarding the use of computing resources. References [1] SIDNEY D. Drell and Tung-mow Yan. “Massive Lepton-Pair Production in Hadron-Hadron Collisions at High Energies”. In: Physical Review Letters 25 (1970), pp. 316–320. URL: https://api.semanticscholar.org/CorpusID:16827178. [2] Luis Miguel Villar Padruno. “Prototyping Sustainability and Performance in High Energy Physics Hardware and Software. Use case: Running the Herwig 7.3 Monte Carlo event generator on a cutting-edge processing unit.” MA thesis. The University of Manchester, 2024. [3] Nicolas van Kempen et al. It’s Not Easy Being Green: On the Energy Efficiency of Programming Languages. 2025. arXiv: 2410.05460 [cs.PL]. URL: https://arxiv.org/abs/2410.05460. [4] CERN. Record Data from the LHC in 2024. Accessed: 2024-12-11. 2024. URL: https : / / home . cern / news / news / accelerators/record-data-lhc-2024. [5] Luis Miguel Villar Padruno and Tobias Fitschen. Green Software in HEP Benchmarks & Studies on Monte Carlo Generators. Presentation at the WLCG Workshop Paris 2025. WLCG Workshop, IJCLab, Paris, France. May 2025. URL: https: //indico.cern.ch/event/1484669/contributions/6456379/. [6] Luis Miguel Villar Padruno. Green software in HEP: Benchmarks & Studies on Monte Carlo Generators. Presentation at SustHEP - 3rd Edition. May 2025. URL: https://indico.global/event/4745/contributions/125384/. [7] S. Navas et al. “Review of Particle Physics”. In: Phys. Rev. D 110 (3 Aug. 2024), pp. 689–774. DOI: 10.1103/PhysRevD. 110.030001. URL: https://link.aps.org/doi/10.1103/PhysRevD.110.030001. [8] Graeme Andrew Stewart et al. “The Critical Importance of Software for HEP”. en. In: (2025). DOI: 10 . 5281 / ZENODO . 15097159. URL: https://zenodo.org/doi/10.5281/zenodo.15097159. [9] Enrico Bothmann et al. “A portable parton-level event generator for the high-luminosity LHC”. In: SciPost Phys. 17.3 (Sept. 2024), p. 081. DOI: 10.21468/SciPostPhys.17.3.081. [10] A. Abada et al. “FCC Physics Opportunities”. In: The European Physical Journal C 79 (2019), p. 474. DOI: 10.1140/epjc/ s10052-019-6904-3. [11] CERN. The Worldwide LHC Computing Grid (WLCG). Accessed: Apr. 12, 2025. 2025. URL: https://home.cern/science/ computing/grid. [12] Loïc Lannelongue, Jason Grealey, and Michael Inouye. “Green Algorithms: Quantifying the Carbon Footprint of Computation”. In: Adv. Sci. 8.12 (2021), p. 2100707. DOI: 10.1002/advs.202100707. [13] Climate Watch. Population based on various sources – with major processing by Our World in Data. Accessed via Climate Watch (2024). 2024. [14] Odin Marc et al. “Comprehensive carbon footprint of Earth, environmental and space science laboratories: Implications for sustainable scientific practice”. In: PLOS Sustainability and Transformation 3 (Oct. 2024). DOI: 10.1371/journal.pstr. 0000135. [15] A. R. H. Stevens, S. Bellstedt, P. J. Elahi, et al. “The imperative to reduce carbon emissions in astronomy”. In: Nature Astronomy 4 (2020), pp. 843–851. DOI: 10.1038/s41550-020-1169-1. 21
[16] ALLEA – All European Academies. Towards Climate Sustainability of the Academic System in Europe and Beyond.https: //doi.org/10.26356/climate-sust-acad. Accessed: 2025-04-13. 2022. [17] Andy Buckley. “Computational challenges for MC event generation”. In: Journal of Physics: Conference Series 1525.1 (Apr. 2020), p. 012023. DOI: 10 . 1088 / 1742 - 6596 / 1525 / 1 / 012023. URL: https : / / dx . doi . org / 10 . 1088 / 1742 - 6596/1525/1/012023. [18] Michael H. Seymour and Siddharth Sule. “An algorithm to parallelise parton showers on a GPU”. In: SciPost Physics Codebases (Aug. 2024). ISSN: 2949-804X. DOI: 10 . 21468 / scipostphyscodeb . 33. URL: http : / / dx . doi . org / 10 . 21468 / SciPostPhysCodeb.33. [19] ATLAS HL-LHC Computing Conceptual Design Report. Tech. rep. Geneva: CERN, 2020. URL: https://cds.cern.ch/ record/2729668. [20] G. Bewick, S. Ferrario Ravasio, S. Gieseke, et al. “Herwig 7.3 release note”. In: Eur. Phys. J. C 84 (2024), p. 1053. DOI: 10.1140/epjc/s10052-024-13211-9. [21] CERN. CERN Member States. Accessed: Apr. 12, 2025. 2025. URL: https://home.cern/about/who-we-are/ourgovernance/member-states. [22] J. Molina et al. “Operating the Worldwide LHC Computing Grid: Current and Future Challenges”. In: J. Phys.: Conf. Ser. 513 (June 2014), p. 062044. DOI: 10.1088/1742-6596/513/6/062044. [23] The HEP Software Foundation, J. Albrecht, A. A. Alves, et al. “A Roadmap for HEP Software and Computing R&D for the 2020s”. In: Comput. Softw. Big Sci. 3 (2019), p. 7. DOI: 10.1007/s41781-018-0018-8. [24] Zach Marshall et al. Carbon, Power, and Sustainability in ATLAS Computing. Tech. rep. Geneva: CERN, 2024. URL: https: //cds.cern.ch/record/2910027. [25] “The Grid: A system of tiers”. In: (2012). URL: https://cds.cern.ch/record/1997396. [26] Geant4 Collaboration. Geant4 Overview. Accessed: Apr. 12, 2025. 2025. URL: https://geant4.web.cern.ch/about/. [27] R. K. Ellis, W. J. Stirling, and B. R. Webber. QCD and Collider Physics. Cambridge Monographs on Particle Physics, Nuclear Physics and Cosmology. Cambridge University Press, 1996. [28] Andy Buckley et al. “General-purpose event generators for LHC physics”. In: Physics Reports 504.5 (July 2011), pp. 145–233. ISSN: 0370-1573. DOI: 10.1016/j.physrep.2011.03.005. URL: http://dx.doi.org/10.1016/j.physrep.2011. 03.005. [29] S. Bethke, A. H. Hoang, S. Kluth, et al. “Quantum Chromodynamics”. In: Prog. Theor. Exp. Phys. 2020.8 (2020). Review of Particle Physics, Particle Data Group, pp. 1–35. DOI: 10.1093/ptep/ptaa104. URL: https://pdg.lbl.gov/2020/ reviews/rpp2020-rev-qcd.pdf. [30] Herwig Collaboration. Herwig Event Generator.https://herwig.hepforge.org/. 2024. [31] S. Prestel. The Lund Hadronization Model (Hadronization in HEP Event Generators). Presented at the NuSTEC Workshop. Gran Sasso Science Institute, Italy. Oct. 2018. URL: https://pythia.org/download/talks/PrestelNu18.pdf. [32] J. Alwall et al. “The automated computation of tree-level and next-to-leading order differential cross sections, and their matching to parton shower simulations”. In: Journal of High Energy Physics 2014.7 (2014). ISSN: 1029-8479. DOI: 10 . 1007 / jhep07(2014)079. URL: http://dx.doi.org/10.1007/JHEP07(2014)079. [33] R. Frederix et al. “The automation of next-to-leading order electroweak calculations”. In: Journal of High Energy Physics 2018.7 (July 2018). ISSN: 1029-8479. DOI: 10.1007/jhep07(2018)185. URL: http://dx.doi.org/10.1007/JHEP07(2018) 185. [34] Federico Buccioni et al. “OpenLoops 2”. In: The European Physical Journal C 79.10 (Oct. 2019). ISSN: 1434-6052. DOI: 10.1140/epjc/s10052-019-7306-2. URL: http://dx.doi.org/10.1140/epjc/s10052-019-7306-2. [35] J. M. Campbell et al. “Event Generators for High-Energy Physics Experiments”. In: (2024), pp. 8–10. arXiv: 2203.11110 [hep-ph]. URL: https://arxiv.org/abs/2203.11110. [36] Christian Bierlich et al. “Robust Independent Validation of Experiment and Theory: Rivet version 3”. In: SciPost Physics 8.2 (Feb. 2020). ISSN: 2542-4653. DOI: 10 . 21468/scipostphys.8.2. 026. URL: http : / / dx . doi .org/10.21468/ SciPostPhys.8.2.026. [37] V. Barger and R.J.N. Phillips. “Quark-parton model relations in deep inelastic lepton scattering”. In: Nuclear Physics B 73.2 (1974), pp. 269–294. ISSN: 0550-3213. DOI: https://doi.org/10.1016/0550-3213(74)90020-0. [38] J. Cardy. Introduction to Quantum Field Theory. Lecture notes, University of Oxford. Accessed: Apr. 12, 2025. 2010. URL: https://www-thphys.physics.ox.ac.uk/people/JohnCardy/qft/qftcomplete.pdf. [39] Particle Data Group. “Review of particle physics: QCD”. In: Chin. Phys. C 38 (2014), pp. 122–135. DOI: 10.1088/16741137/38/9/090001. [40] L. Bryngemark. “Search for New Phenomena in Dijet Angular Distributions at ↑s=8and 13 TeV”. Accessed: Apr. 12, 2025. PhD thesis. Lund, Sweden: Lund University, 2016. URL: https://portal.research.lu.se/files/5666961/8768173. pdf. [41] Sergei V. Chekanov and Rui Zhang. “Enhancing the hunt for new phenomena in dijet final states using anomaly detection filters at the high-luminosity large Hadron Collider”. In: The European Physical Journal Plus 139.3 (Mar. 2024). ISSN: 2190-5444. DOI: 10.1140/epjp/s13360-024-05018-0. URL: http://dx.doi.org/10.1140/epjp/s13360-024-05018-0. [42] P. Z. Skands. “QCD for collider physics”. In: Proceedings of the 2009 CERN-LHC Physics Workshop, CERN Report (2009). DOI: 10.48550/arXiv.1104.2863. [43] Douglas Thain, Todd Tannenbaum, and Miron Livny. “Distributed computing in practice: the Condor experience”. In: Concurrency and Computation: Practice and Experience 17.2-4 (2005), pp. 323–356. DOI: https://doi.org/10.1002/cpe.938. eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/cpe.938. URL: https://onlinelibrary.wiley. com/doi/abs/10.1002/cpe.938. [44] Center for High Throughput Computing. HTCondor High Throughput Computing System. Accessed: Apr. 12, 2025. 2025. URL: https://htcondor.org. 22
[45] Dennis Parsons and Dan Oja. New Perspectives on Computer Concepts 2018: Comprehensive. Boston, MA: Cengage Learning, 2018, p. 33. URL: https://books.google.co.uk/books?id=LAsbEAAAQBAJ. [46] W. Hwu. GPU Computing Gems. Jade Edition. Applications of GPU Computing Series. Waltham, MA: Morgan Kaufmann, 2012. ISBN: 0123859638. DOI: 10.1016/C2010-0-68654-8. [47] Center for High Throughput Computing. HTCondor Users’ Manual. Accessed: Apr. 12, 2025. University of Wisconsin–Madison. 2025. URL: https://htcondor.readthedocs.io/en/latest/users-manual/. [48] Kashif Nizam Khan et al. “RAPL in Action: Experiences in Using RAPL for Power Measurements”. In: ACM Trans. Model. Perform. Eval. Comput. Syst. 3.2 (Mar. 2018). ISSN: 2376-3639. DOI: 10.1145/3177754. URL: https://doi.org/10. 1145/3177754. [49] Dimitrios D. Thomakos and Thomas A. Alexopoulos. “Carbon intensity as a proxy for environmental performance and the informational content of the EPI”. In: Energy Policy 94 (2016), pp. 179–190. ISSN: 0301-4215. DOI: https://doi.org/10.1016/ j.enpol.2016.03.030. URL: https://www.sciencedirect.com/science/article/pii/S0301421516301318. [50] Ember. Global Electricity Review 2024. Accessed: Apr. 12, 2025. 2024. URL: https://ember-climate.org/insights/ research/global-electricity-review-2024/. [51] Loïc Lannelongue, Jason Grealey, and Michael Inouye. “Green Algorithms: Quantifying the Carbon Footprint of Computation”. In: Advanced Science 8.12 (2021), p. 2100707. DOI: https : / / doi . org / 10 . 1002 / advs . 202100707. eprint: https://advanced.onlinelibrary.wiley.com/doi/pdf/10.1002/advs.202100707. URL: https://advanced. onlinelibrary.wiley.com/doi/abs/10.1002/advs.202100707. [52] Green Algorithms Team. GreenAlgorithms4HPC: Green Algorithms for High Performance Computing.https://github. com/GreenAlgorithms/GreenAlgorithms4HPC. Accessed: Apr. 12, 2025. 2025. [53] SchedMD. Slurm Workload Manager Documentation. Accessed: Apr. 12, 2025. 2025. URL: https://slurm.schedmd.com/ documentation.html. [54] A. Karyakin and K. Salem. “Proceedings of the 13th International Workshop on Data Management on New Hardware (DaMoN)”. In: Proc. 13th Int. Workshop Data Management on New Hardware (DaMoN). Chicago, USA: ACM Press, 2017, pp. 1–9. [55] C. Angelini and I. Wilcox. Intel Core i7-5960X, -5930K And -5820K CPU Review: Haswell-E Rises. Accessed: Apr. 12, 2025. Aug. 2014. URL: https://www.tomshardware.com/uk/reviews/intel-core-i7-5960x-haswell-e-cpu,391813.html. [56] The Prometheus Project. Prometheus: Monitoring System and Time Series Database. Accessed: Apr. 12, 2025. 2025. URL: https://prometheus.io/. [57] Domenico Giordano et al. HEPScore: A new CPU benchmark for the WLCG. 2023. arXiv: 2306. 08118 [hep-ex]. URL: https://arxiv.org/abs/2306.08118. [58] Benoit Petit. scaphandre. Version v1.0. 2023. URL: https://github.com/hubblo-org/scaphandre. [59] Enrico Bothmann et al. Many-gluon tree amplitudes on modern GPUs: A case study for novel event generators. 2022. arXiv: 2106.06507 [hep-ph]. URL: https://arxiv.org/abs/2106.06507. [60] Enrico Bothmann et al. “Efficient phase-space generation for hadron collider event simulation”. In: SciPost Physics 15.4 (Oct. 2023). ISSN: 2542-4653. DOI: 10 . 21468 / scipostphys . 15 . 4 . 169. URL: http : / / dx . doi . org / 10 . 21468 / SciPostPhys.15.4.169. [61] S. Lantz et al. “Speeding up particle track reconstruction using a parallel Kalman filter algorithm”. In: Journal of Instrumentation 15.09 (Sept. 2020), P09030–P09030. ISSN: 1748-0221. DOI: 10.1088/1748-0221/15/09/p09030. URL: http://dx.doi. org/10.1088/1748-0221/15/09/P09030. [62] Venkitesh Ayyar et al. “Optimization of Software on High Performance Computing Platforms for the LUX-ZEPLIN Dark Matter Experiment”. In: EPJ Web of Conferences 245 (2020). Ed. by C. Doglioni et al., p. 05012. ISSN: 2100-014X. DOI: 10.1051/epjconf/202024505012. URL: http://dx.doi.org/10.1051/epjconf/202024505012. [63] CERN. PanDA Workload Management System.https://twiki.cern.ch/twiki/bin/view/PanDA/PanDA. 23
Appendices A Plots A.1 Drell-Yan Energy vs Events at 13TeV Fig. A.1. Cumulative energy draw when running HERWIG 7.3for pp →µ+µ↑at ↑s=13TeV. The calculated reduced chi-squared was ϑ2 red =1.02, suggesting that both the fit and the errors considered are highly compatible. In this case, the exponent of the model, b, equals 1.04, that is 3ϖaway from unity. Plot by Luis Miguel Villar Padruno. 24
Fig. A.2. Cumulative energy draw when running HERWIG 7.3for pp →ς+ς↑at ↑s=13TeV. The calculated reduced chi-squared was ϑ2 red =1.27, this number once more suggests a good fit. An interesting note, nonetheless, is the fact that the exponent bof the model deviates from unity, but remains within 5ϖ, the significance is of 4.67ϖ. This is a high deviation that suggests that the physics of the process might modify the relationship between energy and n. Plot by Luis Miguel Villar Padruno. 25