scieee AI-readable full text Open interactive document viewer

Supervised machine learning techniques in high energy physics

Oliveira, António Carlos Pinto

Abstract

O Modelo Padrão da Física de Partículas (MP) é uma teoria extremamente bem sucedida na confrontação experimental. No entanto, a busca por fenómenos que não caibam no seu quadro explicativo é um campo ativo. Várias famílias de teorias que estendem o MP são construídas e novos fenómenos por elas previstos são objeto de pesquisa. Nomeadamente, são feitas buscas por partículas que essas teorias prevêm que se manifestem nos dados adquiridos no detetor ATLAS no LHC a uma energia do centro de massa de 13 TeV. Um tipo particular de buscas consiste em estabelecer limites para certos parâmetros da teoria. Para levar a cabo essas pesquisas vários métodos são usados. Todos eles se baseiam em otimizar a capacidade para distinguir o que é modelizado como acontecimentos esperados assumindo o MP como explicação para as observações (fundo) e o que é modelizado como acontecimentos que deveriam estar presentes se a extensão ao MP fosse correta (sinal). Têm vindo a ser usadas técnicas de aprendizagem de máquina para esse efeito como alternativa a uma análise em que se delimita o espaço de fase da pesquisa usando regiões retangulares e se usam variáveis discriminantes motivadas pelo conhecimento da física do problema em estudo. As redes neuronais escolhem regiões do espaço de fase com formas mais gerais e constroem uma variável discriminante que é opaca no seu significado físico, mas eficaz. Neste trabalho é feita uma comparação do uso de redes neuronais profundas com a análise mais tradicional para estabelecer limites inferiores da massa dum hipotético bosão Z’ usando dados públicos de ATLAS. É também estudado o efeito do uso da variável motivada fisicamente como componente de uma análise baseada em redes neuronais. Por fim, um estudo adicional é feito sobre a transferibilidade de redes neuronais treinadas para reconhecer um sinal específico para discriminar sinais diferentes.

Full text

António Carlos Pinto Oliveira Supervised machine learning techniques in high energy physics Março 2022 UMinho | 2022 António Oliveira Supervised machine learning techniques in high energy physics Universidade do Minho Escola de Ciências Universidade do Minho Escola de Ciências António Carlos Pinto Oliveira Supervised machine learning techniques in high energy physics Dissertação de Mestrado Mestrado em Física Trabalho efetuado sob a orientação do Professor Doutor Nuno Filipe da Silva Fernandes de Castro e do Doutor Miguel Correia dos Santos Crispim Romão Março 2022 ii COPYRIGHT AND TERMS OF USE OF THIS WORK BY A THIRD PARTY This is academic work that can be used by third parties as long as internationally accepted rules and good practices regarding copyright and related rights are respected. Accordingly, this work may be used under the license provided below. If the user needs permission to make use of the work under conditions not provided for in the indicated licensing, they should contact the author through the RepositóriUM of Universidade do Minho. License granted to the users of this work Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International CC BY-NC-SA 4.0 https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en Acknowledgements For the successful completion of this work I relied on the support of my supervisors, Professor Nuno Castro and Doctor Miguel Romão. I am also grateful to my LIP-Minho colleagues for their help, especially to Maura and Henrique. I would like to thank my family for their love and support. iii iv STATEMENT OF INTEGRITY I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the Universidade do Minho. , (Location) (Date) (António Carlos Pinto Oliveira) Resumo Técnicas de aprendizagem automática supervisionada em física de altas energias O Modelo Padrão da Física de Partículas (MP) é uma teoria extremamente bem sucedida na confrontação experimental. No entanto, a busca por fenómenos que não caibam no seu quadro explicativo é um campo ativo. Várias famílias de teorias que estendem o MP são construídas e novos fenómenos por elas previstos são objeto de pesquisa. Nomeadamente, são feitas buscas por partículas que essas teorias prevêm que se manifestem nos dados adquiridos no detetor ATLAS no LHC a uma energia do centro de massa de 13 TeV. Um tipo particular de buscas consiste em estabelecer limites para certos parâmetros da teoria. Para levar a cabo essas pesquisas vários métodos são usados. Todos eles se baseiam em otimizar a capacidade para distinguir o que é modelizado como acontecimentos esperados assumindo o MP como explicação para as observações (fundo) e o que é modelizado como acontecimentos que deveriam estar presentes se a extensão ao MP fosse correta (sinal). Têm vindo a ser usadas técnicas de aprendizagem de máquina para esse efeito como alternativa a uma análise em que se delimita o espaço de fase da pesquisa usando regiões retangulares e se usam variáveis discriminantes motivadas pelo conhecimento da física do problema em estudo. As redes neuronais escolhem regiões do espaço de fase com formas mais gerais e constroem uma variável discriminante que é opaca no seu significado físico, mas eficaz. Neste trabalho é feita uma comparação do uso de redes neuronais profundas com a análise mais tradicional para estabelecer limites inferiores da massa dum hipotético bosão Z’ usando dados públicos de ATLAS. É também estudado o efeito do uso da variável motivada fisicamente como componente de uma análise baseada em redes neuronais. Por fim, um estudo adicional é feito sobre a transferibilidade de redes neuronais treinadas para reconhecer um sinal específico para discriminar sinais diferentes. Palavras-chave: Além do modelo padrão, Aprendizagem automática, ATLAS. v Abstract Supervised machine learning techniques in high energy physics The Standard Model of particle physics (SM) is an extremely successful theory in the comparison with experimental data. However, the search for phenomena that do not fit into its explanatory framework is an active field. Several families of theories extending the SM are constructed and new phenomena predicted by them are the subject of research. Particularly, searches are conducted to find particles that these theories predict will manifest themselves in data acquired at the ATLAS detector at the LHC at a center-of-mass energy of 13 TeV. One particular type of search consists of setting limits on certain parameters of the theory, namely the mass of said particles. To carry out these searches various methods are used. They are all based on optimizing the ability to distinguish between what is modeled as expected events assuming the MP as an explanation for the observations (background) and what is modeled as events that should be present if the extension to the MP were correct (signal). Machine learning techniques have been used for this purpose as an alternative to an analysis in which one delimits the phase space of the search using rectangular regions and uses discriminant variables motivated by knowledge of the physics of the problem under study. Neural networks choose regions of the phase space with more general shapes and construct a discriminant variable that is opaque in its physical meaning, but effective. In this work a comparison is made of the use of deep neural networks with more traditional analysis to establish lower limits on the mass of a hypothetical Z’ boson using ATLAS open data. The effect of using the physically motivated variable as a component of a neural network-based analysis is also studied. Finally, an additional study is done on the transferability of neural networks trained to recognize a specific signal to discriminate different signals Keywords: ATLAS, Beyond the Standard Model, Machine learning. vi Contents List of Figures ix List of Tables xi 1 Theory Overview 1 1.1 The Standard Model ................................ 1 1.1.1 Quantum Electrodynamics ......................... 2 1.1.2 Quantum Chromodynamics ......................... 4 1.1.3 Electroweak Theory ............................. 5 1.1.4 The Brout-Englert-Higgs Mechanism .................... 6 1.2 Beyond the Standard Model Z’ Boson ........................ 10 2 Experimental Setup 12 2.1 LHC ........................................ 12 2.2 ATLAS ....................................... 14 2.2.1 Inner Detector ............................... 15 2.2.2 Calorimeters ................................ 15 2.2.3 Muon Spectrometer ............................ 16 2.2.4 Trigger and Data Acquisition systems .................... 17 3 The 13 TeV ATLAS Open Dataset 19 3.1 Preselection and particle identification ........................ 19 3.2 Search of the decay of Z’ into top quark pairs .................... 22 4 Deep Neural Networks 27 4.1 Machine Learning .................................. 27 4.1.1 Example of regression: Linear Regression .................. 28 vii CHAPTER 1. THEORY OVERVIEW and 𝜂𝜇𝜈 is the Minkowski metric with signature (+ − − −) and a 4x4 identity matrix is implicit. The Proca Lagrangian for a field 𝐴𝜇, L=−1 4(𝜕𝜇𝐴𝜈−𝜕𝜈𝐴𝜇)(𝜕𝜇𝐴𝜈−𝜕𝜈𝐴𝜇) + 𝑚2 𝐴 2𝐴𝜈𝐴𝜈,(1.4) gives rise through the Euler-Lagrange equations to 𝜕𝜇(𝜕𝜇𝐴𝜈−𝜕𝜈𝐴𝜇) +𝑚2 𝐴𝐴𝜈=0,(1.5) which describe a particle of spin 1. For 𝑚𝐴=0, these are Maxwell equations in free space, 𝜕𝜇𝐹𝜇𝜈 =𝑗𝑣, 𝐹 𝜇𝜈 =𝜕𝜇𝐴𝜈−𝜕𝜈𝐴𝜇.(1.6) Adding a source term with a current that obeys the continuity equation, this Lagrangian describes the electromagnetic field. The free fermion Lagrangian is invariant under global phase-rotation transformations, that is, the transformation 𝜓→𝑒−𝑖𝑞𝜒𝜓 does not alter it. If we impose the condition that invariance must be local (called local gauge invariance 1 ), that is, that 𝜒 is a function of space-time coordinates, L must have an additional term in order to preserve the invariance, because the derivative term introduces a dependency on 𝜒that must be cancelled out. As 𝜓→𝑒−𝑖𝑞𝜒(𝑥)𝜓, then Lmust be given by L=[𝑖¯ 𝜓𝛾𝜇𝜕𝜇𝜓−𝑚¯ 𝜓𝜓 ] − (𝑞¯ 𝜓𝛾𝜇𝜓)𝐴𝜇(1.7) where 𝐴𝜇is a vectorial field that transforms as 𝐴𝜇→𝐴𝜇+𝜕𝜇𝜒(𝑥)(1.8) in order to cancel the unwanted term. So, a term that couples the Dirac field with a vectorial field is introduced . Now, one has to add the free part of the Lagrangian for that field, that is, the Proca Lagrangian (1.4). For the Proca Lagrangian to be invariant under (1.8) the mass term must be zero, that is, is necessary to set 𝑚𝐴=0 . As it turns out, imposing local gauge invariance in the Lagrangian for spin 1/2 particles force us to introduce a field for a spin 1 massless particle, the photon, that describes the electromagnetic interaction. The existence of gauge invariance means that not all components of 𝐴𝜇 correspond to physical degrees of freedom because physical quantities must not depend on arbitrary choices of 𝜒(𝑥) . It is important to note, for the discussion that will follow, that the local gauge transformation under which the Lagrangian had to remain invariant is multiplication by a unitary 1x1 matrix, an element of the 𝑈(1) group. Theories for other interactions were built imposing gauge invariance having particular symmetries related to unitary groups of other degrees. It is also worth noting the concept of covariant derivative, which 1For a short account about the use of the term gauge see Reference [8]. 3 CHAPTER 1. THEORY OVERVIEW summarizes the procedure just described to transform a global phase-rotation invariant theory into a local gauge invariant one and to express the interaction between the gauge boson and the fermion. The partial derivative in the free Lagrangian for the fermion is replaced by a covariant derivative 2 D𝜇=𝜕𝜇+𝑖𝑞𝐴𝜇.(1.9) 1.1.2 Quantum Chromodynamics Quantum Chromodynamics (QCD) is the gauge 3 theory of the strong interaction. It describes the interaction between carriers of color charge, that is, quarks, antiquarks, and gluons. Quarks can have three colors, namely, red (r), green (g), blue (b). Anti-quarks have they anti-colors. Color as a quantum number, label states of quarks, antiquarks, and gluons and allow to account for the absence of, for example, free quarks or (q q) hadrons, that are colorless. For a quark of a given flavor with mass 𝑚the free Lagrangian is L=Õ 𝑐∈{𝑟,𝑔,𝑏}[𝑖¯ 𝜓𝑐𝛾𝜇𝜕𝜇𝜓𝑐−𝑚¯ 𝜓𝑐𝜓𝑐].(1.10) Defining the color triplet: 𝜓=©« 𝜓𝑟 𝜓𝑔 𝜓𝑏ª®®®¬ ,¯ 𝜓=(𝜓𝑟𝜓𝑔𝜓𝑏),(1.11) (1.10) can be rewritten as L=𝑖¯ 𝜓𝛾𝜇𝜕𝜇𝜓−𝑚𝑐2¯ 𝜓𝜓 . (1.12) This Lagrangian is invariant under unitary transformations of the new 𝜓 𝜓→𝑈𝜓 , (1.13) where U is an unitary 3x3 matrix. Unitary 3x3 matrices can be written as 𝑈=𝑒𝑖H,(1.14) where His a Hermitian matrix. It can be decomposed further as 𝑈=𝑒𝑖𝜃𝑒𝑖𝝀 2·𝜽,(1.15) where 𝝀 are the eight Gell-Mann matrices 4 . Ignoring the scalar phase factor (the 𝑈(1) symmetry already studied), which amounts to consider only unitary matrices with determinant 1 5 , we can consider then only the 𝑆𝑈 (3)group that is generated by the elements of the algebra defined by [𝑡𝑎, 𝑡𝑏]=𝑖𝑓𝑎𝑏𝑐𝑡𝑐,(1.16) 2This procedure is called the minimal coupling rule. 3 Finding the form of the interaction imposing a local gauge unitary symmetry is based on the 1954 work by Yang and Mills [9] related to isospin in nuclear physics. 4The 3x3 linearly independent traceless Hermitian matrices. 5Due to the relation 𝑑𝑒𝑡 (𝑒𝐴)=𝑒𝑡𝑟 (𝐴). 4 CHAPTER 1. THEORY OVERVIEW where 𝒕=𝝀 2 and 𝑓𝑎𝑏𝑐 are structure constants, not all zero, differently from the 𝑈(1) situation. Using the procedure described in the last subsection for imposing local gauge invariance, namely, considering the transformation 𝜓→𝑒−𝑖𝑔𝝀·𝝓(𝑥)𝜓(1.17) and using the minimal coupling rule D𝜇=𝜕𝜇+𝑖𝑔𝝀·𝑮𝝁(1.18) for which 8 vector fields 𝑮𝝁 must be introduced, that correspond to 8 different gluons, we work out other proprieties of these fields and the complete Lagrangian. Again, the fields must be massless, and due to the noncommutability of the 𝝀 matrices, manifested in the existence of nonzero structure constants, also the field strength tensors 𝐺𝜇𝜈 must have an additional term to ensure the removing of unwanted terms, 𝐺𝜇𝜈 𝑎=𝜕𝜇𝐺𝜈 𝑎−𝜕𝜈𝐺𝜇 𝑎−2𝑔 8 Õ 𝑏,𝑐=1 𝑓𝑎𝑏𝑐𝐺𝜇 𝑏𝐺𝜈 𝑐,(1.19) which results in gluons coupling with each other. The component of the Lagrangian for the quark flavor with mass 𝑚is then given by L=[𝑖¯ 𝜓𝛾𝜇𝜕𝜇𝜓−𝑚¯ 𝜓𝜓] − 1 4G𝜇𝜈 ·G𝜇𝜈 − (𝑔¯ 𝜓𝛾𝜇𝝀𝜓) ·G𝜇.(1.20) The strong interaction has two distinct properties that contrast to electromagnetic interactions. One is asymptotic freedom, which means that the strength of the interaction diminish at shorter distances and higher energies. Other is confinement, whereby the strength of the interaction increases with distance making that, as two quarks are separated, it becomes more energetically favorable to create new quarkantiquark pairs than to have free particles. Hence the observation of jets, a set of hadrons traveling together in a narrow cone, in High Energy Physics (HEP) experiments, where quarks and gluons tracks are expected as decay products. A third propriety is that physical particles must have electric charges that are integer multiples of the electron charge and that limits the possible combinations of quarks and antiquarks in composite particles. 1.1.3 Electroweak Theory The weak interaction was recognized for the first time in nuclear 𝛽 decays. It was in this context that the electron neutrino was introduced to satisfy energy conservation. It has several unique proprieties, namely, it can change the flavor of quarks, it violates parity (and also charge conjugation-parity symmetry), it is mediated by bosons with mass, actually, very significant masses around 90 GeV, which gives the interaction a very short range. Until the development by Salam [10], Weinberg [11], and Glashow [12] of a theory that unified electromagnetic and weak interactions, it was explained considering charged massive 5 CHAPTER 1. THEORY OVERVIEW vector intermediate bosons 𝑊± as force carriers and considering that only left-handed chiral states of leptons and right-handed chiral states of antileptons participated in the interaction. This theory predicted also a a third intermediate boson without electric charge, the 𝑍0 . The gauge symmetry considered was 𝑆𝑈𝐿(2) ⊗ 𝑈(1)𝑌 , where L refers to the lef-handness of the particles interacting and Y refers to the hypercharge that is given by 𝑌=2𝑄−2𝐼3 where Q is the electric charge, in units of the charge of the proton, and 𝐼3 is the value of the third component of isospin, which components are the generators of SU(2), that is 𝐼𝑖=1/2𝜎𝑖 , where 𝜎𝑖 are the Pauli matrices. The fermions are arranged as left-handed doublets and right-handed singlets : 𝜓𝑖 𝐿= 𝜈𝑖 𝐿 𝑙𝑖 𝐿!, 𝑢𝑖 𝐿 𝑑𝑖 𝐿! 𝜓𝑖 𝑅=𝑙𝑖 𝑅,𝑢𝑖 𝑅,𝑑𝑖 𝑅 ,(1.21) where i runs for the three generations of fermions. To build the gauge theory, the following covariant derivative must be used: D𝜇=𝜕𝜇−𝑖𝑔𝑰·𝑾𝝁−𝑖𝑔0𝑌 2𝐵𝜇,(1.22) where g and 𝑔0 are coupling constants, and were gauge fields were introduced, 𝐵𝜇 for the 𝑈(1)𝑌 group and 𝑾𝝁 for 𝑆𝑈𝐿(2) . Also kinetic terms would be necessary as before, constructed once again with the field strengths tensors: 𝑊𝑖 𝜇𝑣 =𝜕𝜇𝑊𝑖 𝜈−𝜕𝑣𝑊𝑖 𝜇−𝑔𝜀𝑖𝑗𝑘𝑊𝑗 𝜇𝑊𝑘 𝜈 𝐵𝜇𝑣 =𝜕𝜇𝐵𝜈−𝜕𝜈𝐵𝜇 (1.23) where 𝜀𝑖𝑗𝑘 are the components of the totally antisymmetric Levi-Civita tensor,the structure constants of SU(2). Then the Lagrangian will result in: L=𝑖¯ 𝜓𝛾𝜇D𝜇𝜓−1 4𝑊𝑖 𝜇𝜈𝑊𝑖,𝜇𝜈 +𝐵𝜇𝑣𝐵𝜇𝑣 .(1.24) Notice that the gauge bosons must be massless but also the fermions (in order to decouple right-handed and left-handed states). However, massive bosons are needed because that is what is observed. Another ingredient was necessary to obviate this problem. It is the Higgs mechanism based on spontaneous symmetry breaking. 1.1.4 The Brout-Englert-Higgs Mechanism The mechanism of spontaneous symmetry breaking can be understood considering the following steps. Consider the Lagrangian [13] [14]: L=1 2(𝜕𝜇𝜙𝜕𝜇𝜙) + 1 2𝜇2𝜙2−1 4𝜆2𝜙4.(1.25) 6 CHAPTER 1. THEORY OVERVIEW Re(𝜙)Im(𝜙) 𝑉(𝜙) The Mexican hat potential A B Figure 1.2: The Brout-Englert-Higgs potential. Consider it like a sum of a kinetic part and the potential part. However, apparently, the mass term has the wrong sign. The minimum for the potential will occur, not for zero but for 𝜙=±𝜇 𝜆 . The Lagrangian must be rewritten in terms of perturbations around the ground state. If we write 𝜂=𝜙±𝜇 𝜆 , the Lagrangian become, L=1 2(𝜕𝜇𝜂𝜕𝜇𝜂) −𝜇2𝜂2±𝜇𝜆𝜂3−1 4𝜆2𝜂4+1 4 𝜇2 𝜆!2 .(1.26) Then, the mass term is the second, corresponding to a mass √2𝜇 . Also, the symmetry ( 𝜙→ −𝜙 ) that was present in 1.25 no longer holds in terms of the field 𝜂. Now, let us consider two fields, L=1 2𝜕𝜇𝜙1𝜕𝜇𝜙1+1 2𝜕𝜇𝜙2𝜕𝜇𝜙2+1 2𝜇2𝜙2 1+𝜙2 2−1 4𝜆2𝜙2 1+𝜙2 22.(1.27) This Lagrangian is invariant under rotations in the space of the linear combinations of both fields. The minimum for the potential is now given by any point in the circle: 𝜙2 1+𝜙2 2=0.(1.28) Choosing the particular solution: 𝜙1= 𝜇 𝜆, 𝜙2=0(1.29) and defining the fluctuations about that minimum: 𝜂=𝜙1−𝜇 𝜆, 𝜉 =𝜙2,(1.30) we can write the Lagrangian as L=1 2𝜕𝜇𝜋𝜕𝜇𝜋−𝜇2𝜂2+1 2𝜕𝜇𝜉𝜕𝜇𝜉 +"𝜇𝜆 𝜂3+𝜂𝜉2−𝜆2 4𝜂4+𝜉4+2𝜂2𝜉2#+𝜇4/4𝜆2. (1.31) 7 CHAPTER 1. THEORY OVERVIEW We recognize a Klein-Gordon field with mass √2𝜇 and a free Lagrangian for 𝜉 with no mass. Also, the original SO(2) symmetry is no longer to be seen. Rewriting 1.27 using: 𝜙=𝜙1+𝑖𝜙2,(1.32) L=1 2(𝜕𝜇𝜙)∗(𝜕𝜇𝜙) + 1 2𝜇2(𝜙∗𝜙) − 1 4𝜆2(𝜙∗𝜙)2.(1.33) We see that in this guise, the Lagrangian has the symmetry U(1) ( 𝜙→𝑒𝑖𝜃𝜙 ). Imposing that the Lagrangian must be invariant under local gauge transformations, and replacing partial derivatives with covariant derivatives, introducing massless vector fields for that effect, the Lagrangian becomes: L=1 2(𝜕𝜇−𝑖𝑞𝐴𝜇)𝜙∗(𝜕𝜇+𝑖𝑞𝐴𝜇)𝜙 +1 2𝜇2(𝜙∗𝜙) − 1 4𝜆2(𝜙∗𝜙)2−1 4𝐹𝜇𝜈𝐹𝜇𝜈 . (1.34) Defining new fields, 𝜂and 𝜉as before, it becomes: L=1 2𝜕𝜇𝜂𝜕𝜇𝜂−𝜇2𝜂2+1 2𝜕𝜇𝜉𝜕𝜇𝜉 +−1 4𝐹𝜇𝜈𝐹𝜇𝜈 +1 2𝑞𝜇 𝜆2𝐴𝜇𝐴𝜇−2𝑖𝜇 𝜆𝑞𝜕𝜇𝜉𝐴𝜇+interaction terms. (1.35) This procedure brought us a massless scalar boson 𝜉 and the particle 𝜂 with mass √2𝜇 , as before, but the field 𝐴𝜇 acquired mass. Still, some problems remain. Using the freedom given by the gauge invariance we can fix the gauge in such a way that, given 𝜙=𝜂+𝜇 𝜆+𝑖𝜉 , (1.36) we make: 𝜙→𝑒𝑖𝜃𝜙=(cos𝜃+𝑖sin𝜃)(𝜙1+𝑖𝜙2)(1.37) real. This happens if we choose 𝜃=−arctan(𝜙2/𝜙1) . Then, the unwanted terms disappear, including the massless boson. We are left with a massive vectorial particle and a scalar massive boson. This procedure was used to break U(1). In a similar manner, the electroweak symmetry can be broken. Consider the doublet scalar field, called Higgs field: Φ= 𝜙+ 𝜙0!=1 √2 𝜙1+𝑖𝜙2 𝜙3+𝑖𝜙4!(1.38) containing an electrically charged and a neutral fields. A Lagrangian to govern this fields has the form: L=(D𝜇)†(D𝜇) − −𝜇2Φ†Φ+𝜆2(Φ†Φ)2,(1.39) 8 CHAPTER 1. THEORY OVERVIEW where the covariant derivative is given in Equation 1.22, and the signs in the Higgs potential introduced after the kinetic term (and depicted in Figure 1.2) were chosen to have a stable minimum different from |Φ|=0. Actually, with this choice, there is an infinite set of degenerate minima in the ring: Φ†Φ=1 2(𝜙2 1+𝜙2 2+𝜙2 3+𝜙2 4)= 𝑣2 2 = 𝜇2 2𝜆2.(1.40) The fields can be parametrized as: Φ=1 √2 0 𝑣+ℎ(𝑥)!.(1.41) The term where to look to find the masses is: (D𝜇Φ)†(D𝜇Φ)=|(𝜕𝜇−𝑖𝑔𝑰·𝑾𝝁−𝑖𝑔0𝑌 2𝐵𝜇)Φ|2.(1.42) Making use of the identity 𝑰=𝝈/2this results in: (D𝜇Φ)†(D𝜇Φ)=1 2𝜕𝜇ℎ𝜕𝜇ℎ+1 8𝑔2 𝑊𝑊(1) 𝜇+𝑖𝑊 (2) 𝜇𝑊(1)𝜇−𝑖𝑊 (2)𝜇(𝑣+ℎ)2 +1 8𝑔𝑊𝑊(3) 𝜇−𝑔0𝐵𝜇𝑔𝑊𝑊(3)𝜇−𝑔0𝐵𝜇(𝑣+ℎ)2. (1.43) Equating the terms quadratic in the boson fields with: 1 2𝑚2 𝑊𝑊(𝑖) 𝜇𝑊(𝑖)𝜇, 𝑓 𝑜𝑟 𝑖 =1,2(1.44) follows the mass of the W boson: 𝑚𝑊=1 2𝑔𝑣 . (1.45) Using the relation: 𝑍𝜇 𝐴𝜇!= cos𝜃𝑊−sin𝜃𝑊 sin𝜃𝑊cos𝜃𝑊! 𝑊3 𝜇 𝐴𝜇!,(1.46) follows 𝑀𝑍=1 2𝑣q𝑔2+𝑔02, 𝑀𝐴=0.(1.47) The three gauge bosons that intermediate the weak interaction gain mass and the photon rests massless. The Weinberg angle can be obtained by cos𝜃𝑊=𝑀𝑊/𝑀𝑍 . The fermion masses are obtained by terms like: Le=−𝑔e"¯𝑣ee𝐿 𝜙+ 𝜙0!e𝑅+e𝑅𝜙++𝜙0+ 𝑣e e!𝐿#(1.48) Using the expansion of Equation 1.41, it follows: 𝑚e=1 √2𝑔e𝑣 . (1.49) 9 CHAPTER 1. THEORY OVERVIEW From Equation 1.42 comes the mass of the Higgs boson: 𝑚𝐻=√2𝜆𝑣 . (1.50) The values of g and g’ can be obtained by 𝑔sin𝜃𝑊=𝑔0cos𝜃𝑊=𝑒 . From Equation 1.45, the value of v, the vacuum expectation value, follows with 𝑣≈ 246 GeV [15]. Only the value of 𝜆 is left to determine the Higgs boson mass. So, this is a parameter to be determined by experiment. 1.2 Beyond the Standard Model Z’ Boson Figure 1.3: 𝑡¯ 𝑡 decaying semileptonically. u’ and d’ stand for up-type and down-type quarks, respectively, and ℓfor electron or muon. Despite its accuracy and success, the SM is thought to be incomplete. The observation of neutrino oscillations indicate that neutrinos have mass, in contradiction to the Standard Model predictions. Astrophysical observations demand the existence of a different kind of matter not described by the SM, known only by its gravitational effects, called Dark Matter because it doesn’t interact via the electromagnetic field. In cosmology, Dark Energy, assumed to exist to justify the acceleration of the expansion of the Universe, is also not explained by the SM. There is the matter-antimatter asymmetry, the fact that SM doesn’t describe gravity, the large number of free parameters, etc. Several extensions of the SM introduce a heavy, electrically neutral, spin-1 boson called Z’, that decays into 𝑡¯ 𝑡 . In this work a specific model will be used, that corresponds to a leptophobic, topophyllic Z’ corresponding to a specific model of the topcolor-assisted-technicolor [16] [17] family, more concretely the Model IV [18], that couples only to first and third generation quarks. A search [19] performed by ATLAS Collaboration found no significant deviation from the Standard Model predictions but set exclusion limits on the production cross-section times branching ratio on the production of Z’. Namely, upper limits on the production cross section vary between 25 pb to 0.02 pb for masses from 0.4 TeV to 5 TeV. Masses of Z’ lower than 2.6 TeV were excluded (c.f. Figure 1.4). 10 CHAPTER 1. THEORY OVERVIEW Figure 1.4: The observed and expected cross-section 95% CL upper limits on the Z’ signal and the theoretical predictions for the production cross-section times branching ratio of Z’ [19]. The high center-of-mass energy of LHC allows it to be a top factory. The top quark decays almost exclusively to a 𝑊+ boson and a quark bottom. Moreover, having a very short lifetime ( ≈4×10−25 s) it do not form hadrons. The semileptonic channel of the pair top antitop (see Figure 1.3) is chosen for study because it is a good compromise between competing advantages. For one part, the all-hadronic channel, where both W bosons decay into quarks, is the dominant channel (46%), which has the advantage of providing more statistics. But the final state of this decay consists in six jets what makes it hard to distinguish from QCD multijets events. Furthermore, the large multiplicity of jets makes difficult the task of ascribing each one to the correct top quark from which they decay. The dileptonic channel where both W bosons decay to leptons has a small branching ratio (9%) which is a drawback in terms of statistics. Its advantage is that it is more easy to distinguish its final states from QCD multijets. It has the drawback, though, of having two undetected particles. Lastly, the semileptonic channel has a branching ratio comparable to the all-hadronic channel (45%) but the decay 𝑊→𝜏𝜈𝜏 (15%) is usually not considered in analyses because it introduces an additional neutrino (missing energy) when it decays via the weak interaction ( 𝜏→𝑊𝜈𝜏 ). But it decays with more probability into quarks, giving rise to more jets and making hard the task of event reconstruction. In this way, omitting this decay, the semileptonic channel is less statistically advantageous than the all-hadronic but still good (30%) [20]. It is, although, easier to distinguish its final state from background and more suitable for event reconstruction. 11 Chapter 2 Experimental Setup 2.1 LHC The Large Hadron Collider (LHC) is a particle accelerator at the European Laboratory for Particle Physics (CERN), in Geneva, Switzerland. It consists of a ring of superconducting magnets located underground, having a circumference of 26.7 km. [22] It is divided in 8 archs and 8 straight sections between them. At 4 of these sections are located four main detectors: ATLAS [23], ALICE [24], CMS [25] and LHCb [26]. Of these, CMS and ATLAS are general-purpose detectors, LHCb is dedicated to the study of B physics, and ALICE to the study of heavy-ion physics. LHC was designed to achieve a center-of-mass energy of 14 TeV, through proton-proton collisions. Lead-lead and proton-lead collisions are also carried out. The purpose is to firmly establish the validity of the SM (namely, the discovery of the Higgs boson [27] [28] announced on 4th July 2012, is a highlight of the goals accomplished so far) and to find new physics beyond it. Protons are accelerated in opposite beams, each one inside its own pipeline (kept at ultra-high vacuum) until they reach the wanted energy. They are then focused to the point of interaction inside the detectors. Beams consist of bunches of particles, around 25 ns apart 1 , guided in the pipelines by very strong magnetic fields created by the superconducting magnets operating at 1.9 K for which a cooling system based on liquid helium is necessary. The acceleration process is done in several phases as illustrated in Figure 2.1. First, protons are collected from a container where hydrogen molecules are split into electrons and protons after application of an intense electric field. Then, they are accelerated in a linear accelerator, called LINAC 2, until they reach an energy of 50 MeV. After, they are injected into the proton synchrotron booster (PSB) where they reach 1.4 GeV before they enter the Proton Synchrotron (PS), that accelerate them until 26 GeV. After, they are injected into the Super Proton Synchrotron (SPS) and are accelerated to 450 Gev. At last the enter the LHC. In the second run of the LHC, from 2015 to 2018, the protons collide with a center-of-mass energy of √𝑠of 13 TeV. 1This refers to run 2 at which the data analyzed in this work was collected. 12 Chapter 3 The 13 TeV ATLAS Open Dataset The 13 TeV ATLAS Open Dataset [33] is a collection of data, Monte Carlo (MC) simulations and tools provided by the ATLAS experiment to be used, among other things, for educational purposes and Machine Learning challenges. It provides tools that implement examples of some physics analysis. These can be the point of departure for further analysis and development of techniques to perform them. Data is available in the ROOT [34] format. The data is comprised of events from 61 runs that were collected in 2016 by the ATLAS detector from p-p collisions at √𝑠= 13 TeV. It consists of approximately 270 millions events. Only events recorded when all subsystems of ATLAS were working acceptably were included. Also quality criteria related to the beam and data were imposed. In total, the dataset correspond to an integrated luminosity of 10.06 ±0.37fb−1. In addition to the data, the 13 TeV ATLAS Open Dataset includes MC simulations that describe various SM processes used to model expected signal and background. Data and MC simulations are submitted to the same quality and trigger criteria as the data. In the end of a loose preselection, performed to reduce subsequent processing time, they are grouped in collections according to type and multiplicity of reconstructed objects with high transverse momentum. For our analysis, the collection of interest is the one labeled 1largeRjet1lep, where among the final state reconstructed objects are at least one jet with large-R with minimum 𝑝𝑇of 250 Gev and exactly one charged lepton with minimum 𝑝𝑇of 25 GeV. 3.1 Preselection and particle identification Electron candidates are reconstructed matching isolated energy deposits (clusters) in the ECAL to tracks in the ID. It is considered only the precision region of the ATLAS detector, called fiducial, defined by |𝜂cluster|<2.47 . It is necessary, in addition, to exclude the transition region between the barrel and the endcap of the ECAL, 1.37 <|𝜂cluster|<1.52 . The candidate must have 𝑝𝑇 > 7 GeV and pass loose identification criteria [35]. Very loose, loose, medium or tight criteria refer to how tight the match 19 CHAPTER 3. THE 13 TEV ATLAS OPEN DATASET between the track and the cluster is enforced, depending on the particle identification efficiency required. The identification is made using a likelihood-based discriminant and the loose operating point correspond to 93 % efficiency for identifying a prompt electron with 𝐸𝑇= 40 GeV. For muons, the reconstruction is based in matching tracks in the ID and in the Muon Spectrometer. The muon candidate must also have 𝑝𝑇 > 7 GeV and pass loose identification criteria [36]. These criteria aims to maximize the reconstruction efficiency while keeping good-quality muon tracks. Isolation criteria are imposed on muons and electrons to reduce contributions from unwanted sources, such as non-prompt leptons, photon conversions and hadrons. Events containing at least an electron or a muon are selected with single-lepton triggers with 𝑝𝑇 threshold of 26 GeV and isolation requirements, or with a larger threshold of 50-60 Gev, looser identification requirements and no isolation requirement. The reconstruction of photon candidates, like for electron candidates, is based on detecting energy clusters in the ECAL not matched by any track in the ID, and, additionally, searching for a process of photon conversion into 𝑒−𝑒+ at the ID and corresponding clusters in the ECAL. To reduce hadronic background, the photon candidates must obey ”loose”isolation criteria. Other criteria are shown in the Table 3.1. Electron (𝑒)Muon (𝜇)Photon (𝛾) ID & ECAL rec. ID & MS rec. ID & ECAL rec. loose identification loose identification tight identification loose isolation loose isolation loose isolation 𝑝𝑇>7 GeV 𝑝𝑇>7 GeV 𝐸𝑇>25 GeV |𝜂|<2.47 |𝜂|<2.5|𝜂|<2.37 Hadronically decaying 𝜏-leptons (𝜏ℎ)Small-R jets LargeR jets ID & ECAL rec. ECAL & HCAL rec. ECAL & HCAL rec. medium identification anti𝑘𝑡, R =0.4 anti𝑘𝑡, R =1.0 𝑃𝑇>20 GeV 𝑃𝑇>20 GeV 𝑃𝑇>250 GeV |𝜂|<2.5|𝜂|<2.5|𝜂|<2.0 1 or 3 associated tracks 𝑏-tagging (MV2c10) trimming: 𝑅sub =0.2,𝑓cut =0.05 Table 3.1: Preselection requirements. Jet candidates are reconstructed using three-dimensional energy clusters in ECAL and HCAL using a specific algorithm called anti𝑘𝑡1 with radius parameter of 0.4 for the ”small-R” jets. They must fulfil |𝜂|< 2.5 and 𝑝𝑇> 20 GeV. One effect that needs to be minimized is ”pile-up”. It consists of low transverse momentum collisions other than the hard-scatter p-p collisions that are of interest, which would hide the rare events that we want to study. These additional collisions from the same bunch crossing are called 1 The procedure goes as follow. The pairwise distance between objects is computed. The two closest objects are merged and the procedure is repeated until no pair of particles are closer than a distance R. Also, a distance between the object i and the beam, 𝑑𝑖𝐵 is considered. If that distance is larger than 𝑑𝑖 𝑗 , the pairwise distance between two objects, those objects are merged, otherwise the object is considered a jet and no more merging is performed. In the anti𝑘𝑡 algorithm, 𝑑𝑖𝐵 is given by 1 𝑝2 𝑇𝑖 and 𝑑𝑖 𝑗 is given by 𝑚𝑖𝑛(1 𝑝2 𝑇𝑖 ,1 𝑝2 𝑇 𝑗 )𝑅2 𝑖 𝑗 /𝑅2, where 𝑅𝑖 𝑗 =p(𝜂𝑖−𝜂𝑗)2+ (𝜙𝑖−𝜙𝑗)2[37]. 20 CHAPTER 3. THE 13 TEV ATLAS OPEN DATASET in-time pile-up. Also, energy deposits associated with previous or following bunch crossings relative to the triggered event must be dealt with (due to the response time of subdetectors being larger than the interval between successive bunch crossings), and they are called out-of-time pile-up. To reduce this effect, a condition is imposed on the score of the jet vertex tagger (JVT) [38] discriminant, for jets with 𝑝𝑇< 60 GeV and |𝜂|<2.4. Large-R jets are built using also the anti𝑘𝑡 algorithm with R = 1.0. Then, the are trimmed. [39] which reduces the effects of pile-up. This consists of recluster the components of the large-R jet into subjets with a 𝑅sub parameter. Subjets with transverse momentum lesser than a fraction 𝑓cut of the original jet are discarded. The parameters 𝑅sub = 0.2 and 𝑓cut = 0.05 are chosen based on a study of sensitivity to pile-up [40]. Large-R jets must have |𝜂|< 2.0 and 𝑝𝑇> 250 GeV. The visible products of the 𝜏 -lepton decays are also reconstructed. It is very important to identify jets containing B-hadrons (b-tagging) for various physical analysis as this allows a huge rejection of background processes. The present analysis demands to find jets originated from b quarks into which top quarks decay almost exclusively, plus a W boson, by way of the weak interaction. Because B-hadrons decay very close to the point of creation (around 0.5 mm, corresponding to a decay time of 10−12 s), leave therefore a secondary vertex. This vertex can be reconstructed by the convergence of tracks to a vertex very near the primary vertex 2 or if it cannot be resolved, the impact parameters of tracks relative to the primary vertex are used. In the pre-selection that is being presented, a multivariate discriminant, MV2c10, is used, that combines this information. For each jet, the value for the discriminant is calculated. For a required efficiency of b-tagging there is a threshold, the working point (WP), the value of which the discriminant must surpass [41]. After this pre-selection, data quality criteria are applied to guarantee that detectors were correctly functioning and tracks were not reconstructed from deposits that were due to cosmic-ray showers, or hardware problems. Also, events must contain at least one reconstructed vertex with at least two tracks with 𝑝𝑇> 0.4 GeV. The tracks associated to muons and electrons must correspond to the primary vertex of the event. Several SM processes that can mimic the signal were simulated, namely the production of 𝑡¯ 𝑡, singletop, W plus jets, Z plus jets and diboson. In addition, simulations of some BSM processes, namely, the Z’ production are included. They are listed in Table 3.2 with references to the software used to perform the MC simulations. The 13 TeV ATLAS Open Dataset includes several SM and BSM physics analyses. For all, the reconstructed physics objects are subject to additional selection (Table 3.3) that correspond to the requirement of calorimeter (etcone20) and track (ptcone30) isolation for electrons, photons and muons. etcone20 is given by the sum of the energy of the clusters located inside a cone of Δ𝑅= 0.2 around the object considered and ptcone30 is defined as the scalar sum of the 𝑝𝑇 of tracks within a cone of Δ𝑅= 0.3. 2This is the vertex that corresponds to highest sum of squared transverse momentum of the tracks associated with it. 21 CHAPTER 3. THE 13 TEV ATLAS OPEN DATASET Process Generator, hadronisation Additional information Top-quark production 𝑡¯ 𝑡+ jets POWHEG-BOX v2 + PYTHIA 8 only 1 ℓand 2 ℓdecays of 𝑡¯ 𝑡-system single (anti)top t-channel POWHEG-BOX v1 + PYTHIA 6 single (anti)top W t-channel POWHEG-BOX v2 + PYTHIA 6 single (anti)top s-channel POWHEG-BOX v2 + PYTHIA 6 W/Z (+ jets) production Z→𝑒𝑒, 𝜇𝜇, 𝜏𝜏 POWHEG-BOX v2 + PYTHIA 8 LO accuracy up to 𝑁𝑗𝑒𝑡𝑠 = 1 W→𝑒𝜈, 𝜇𝜈, 𝜏𝜈 POWHEG-BOX v2 + PYTHIA 8 LO accuracy up to 𝑁𝑗𝑒𝑡𝑠 = 1 W→𝑒𝜈, 𝜇𝜈, 𝜏𝜈 + jets SHERPA 2.2 LO accuracy up to 3-jets final states Z→𝑒𝑒, 𝜇𝜇, 𝜏𝜏 + jets SHERPA 2.2 LO accuracy up to 3-jets final states Dibosons production WW SHERPA 2.2 𝑞𝑞0ℓ𝜈 final states WW SHERPA 2.2 ℓ𝜈ℓ0𝜈0final states ZZ SHERPA 2.2 𝑞𝑞0ℓ+ℓ−final states ZZ SHERPA 2.2 ℓ+ℓ−ℓ0+𝑙0− final states WZ SHERPA 2.2 𝑞𝑞0ℓ+ℓ−final states WZ SHERPA 2.2 ℓ𝜈𝑞𝑞0final states WZ SHERPA 2.2 ℓ𝜈ℓ+ℓ−final states WZ SHERPA 2.2 ℓ𝜈𝜈𝜈0final states Z’ production Z’ →𝑡¯ 𝑡PYTHIA 8 𝑚𝑍0=1 TeV Table 3.2: MC samples contained in 13 Tev ATLAS Open Dataset used in the analysis. Large R-jets with mass lower than 50 GeV or 𝑝𝑇 larger than 1500 GeV are excluded because they belong to a region of phase-space that is not well-calibrated. Electrons and Muons Small-R jets Photons Large-R jets 𝜏ℎ 𝑝𝑇>25 GeV 𝑝𝑇>25 GeV 𝑝𝑇<1500 GeV 𝑝𝑇>25 GeV lep_ptcone30 < 0.15 JVT > 0.59 photon_ptcone30 < 0.065 mass > 50 GeV lep_etcone20 < 0.15 photon_etcone20 < 0.065 Table 3.3: Additional object selection. 3.2 Search of the decay of Z’ into top quark pairs This work will compare a machine learning based analysis to the search of the decay of Z’ into top quark pairs in events that contain a single charged lepton, large R-jets and missing momentum with the analysis included in the 13 Tev ATLAS Open Dataset that will be described below. The analysis is based in previous searches [19] for decays of heavy particles to top-quark pairs in 𝑝𝑝 collisions at √13 = 13 TeV with ATLAS. It applies a selection known as single-lepton boosted topology to the final products of a top and antitop decay. [42]. This correspond to semileptonic decays of the 𝑡¯ 𝑡 system as exemplified in Figure 22 CHAPTER 3. THE 13 TEV ATLAS OPEN DATASET 1.3, where the decay products of the top quark that decays hadronically quark are enclosed within one large-radius jet. Tops decay into a W boson and a bottom quark. One W boson decays into a an electron or a muon plus a neutrino while the other decays into quarks. An all-hadronic topology is more frequent but is more difficult to separate from non𝑡¯ 𝑡 background. This selection requires a single isolated electron or muon, large missing transverse momentum and hadronic jets, one of which must contain a b-hadron. More concretely, the final event selection criteria are: 1. There must be at least a preselected large-R jet. 2. The missing transverse momentum 𝐸miss 𝑇 , defined as the magnitude of the negative of the vector sum of the transverse momentum of all selected physics objects, must be greater than 20 GeV. 3. Single-electron or single-muon trigger satisfied. 4. There is exactly on good lepton, that is, one for which 𝑝𝑇> 30 GeV, and the track associated with it must must match the candidate that triggered the event. Also, the identification must be tight. 5. In order to have a situation consistent with a leptonic W decay, additionally to point 2, 𝐸miss 𝑇+𝑀𝑊 𝑇 must be larger than 60 GeV, where 𝑀𝑊 𝑇 is the transverse mass of the W boson candidate (the selected lepton ℓplus the 𝐸miss 𝑇), given by 𝑀𝑊 𝑇=q2𝑝ℓ 𝑇𝐸miss 𝑇(1−cos Δ𝜙(ℓ, 𝐸miss 𝑇)) 3. 6. At least, one small-R jet close to the lepton, that is, Δ𝑅(lepton, jet) < 2.0. This would correspond to the leptonic W decay. 7. There must be exactly one large-R jet that pass simplified requirements to be compatible with a hadronically top decay (top-tagged), namely, to have mass larger than 100 GeV and to have N-subjettiness ratio [43] 𝜏32 < 0.75. 𝜏32 =𝜏3 𝜏2 , 𝜏𝑁 expressing how well a jet can have N or fewer subjets. This variable allows the discrimination between jets containing three subjets and jets containing two subjets. Also, in addition to the conditions already stated for a well-calibrated region of phase space, we must have 𝑝𝑇>300 Gev and |𝜂|<2. 8. The large-R jet of item 7must be well apart from the small-R jet ( Δ𝑅> 1.5) and from the lepton (Δ𝜙>1.0). 9. At least one b-tagged jet (with the requirement of WP corresponding to 70% of efficiency of tagging). This jet must be within the top-tagged large-R jet or to be the small-R jet close to the lepton, that is Δ𝑅(large-R jet, b-tagged jet)<1.0 or Δ𝑅(small-R jet, b-tagged jet)<0.01. 3 This quantity is defined in a special manner when we have a particle to decay into two particles, of which one is invisible [20]. In this case, 𝑀2 𝑇=(𝐸𝑇(1) +𝐸𝑇(2))2− (® 𝑝𝑇(1) + ® 𝑝𝑇(2))2 , where 𝐸2 𝑇=𝑚2+® 𝑝2 𝑇 is the transverse energy. For the case where the daughter particles are massless or can be considered as such, as in the case under consideration, the result has the form stated. 23 CHAPTER 3. THE 13 TEV ATLAS OPEN DATASET (a) Missing transverse momentum with application of cuts 1,3and 4. (b) Application of the same cuts as in Figure 3.1a. (c) Number of large-R jets that pass simplified top-tagging requirements.Application of the same cuts as in Figure 3.1a plus cuts 2,5,6and at least one b-tagged jet. (d) Application of all cuts except 8and 9. The variable is the Δ𝑅referred in 8. Figure 3.1: Relevant plots to justify the cuts enumerated above. 24 CHAPTER 3. THE 13 TEV ATLAS OPEN DATASET (e) Application of all cuts except 8and 9. The variable is the Δ𝜙referred in 8. (f) Application of all cuts except 9. This variable tells if the small-R jet described in that item is b-tagged. Figure 3.1: Relevant plots to justify the cuts enumerated above. (continuation) Sample Number of Events Single Top 610.78 Diboson 22.34 𝑡¯ 𝑡13600.48 V + jets 682.68 Z’ 1 TeV 426.62 Table 3.4: Expected number of selected events for a luminosity of 10 fb−1. In Figures 3.1a and 3.1b we can see that it is justified to cut in the missing energy as there are disagreement between data and MC samples at low missing energy because multijet background was not simulated due to its large cross-section, which would require simulating a huge number of events. In Figure 3.1c it is possible to see that demanding exactly one large-R jet that pass simplified top-tagging requirement increases the signal to background ratio. The same reasoning applies for Figures 3.1d and 3.1e that justify the cuts 8and Figure 3.1f to justify cut 9. The number of selected events are shown in Table 3.4 distributed according to the type of background. In Figure 3.2 is shown an observable that approximates 4 the mass of the top-antitop system, adding the four-momenta of the charged lepton, of the b-tagged small-R jet and of the top-tagged large-R jet. It can be seen that the Standard Model prediction is consistent with the data, from which the simulated hypothesized Z’ model deviates considerably. 4 For simplicity, the four-momentum of the neutrino is not included as it would involve the difficult reconstruction of the longitudinal component of the missing energy. 25 CHAPTER 3. THE 13 TEV ATLAS OPEN DATASET Figure 3.2: Approximate mass of the top-antitop system. 26 Chapter 4 Deep Neural Networks 4.1 Machine Learning Artificial Intelligence is the area of computer science concerned with implementing the ability to perform tasks usually associated with intelligent beings. Some problems are amenable to be described by formal rules easily translatable to computer programs. The emblematic example is playing chess. Other tasks are intuitive for humans but hard to formally describe by explicit rules (e.g., driving, recognizing faces.) The way to solve them is by allowing the rules to be learned from the data. This ability is called machine learning. Machine Learning (ML) is the subfield of artificial that occupies itself with this problem. Datasets consist of examples of some phenomenon under study, that is, observations, organized as vectors. Each dimension of such vectors, called a feature, represents a relevant aspect of the phenomenon to the task at hand. The process of selecting those features is called feature engineering. For example, if the task is to tell apart signal from background events in a High-Energy Physics experiment, a good guess would be to choose variables used in traditional, non-ML based, analysis or even more low-level variables (i.e., the ones from which the former were constructed.) Machine learning algorithms can be supervised,unsupervised or semi-supervised. The rules to be machine learned are functions (models) that map observations to predictions. Training a machine learning model means finding its parameters. In supervised learning (the only type used in this work), the examples are labeled. The model parameters are found by minimizing an error function, also called a cost function, that measures the difference between predictions and their true values expressed in the labels. When labels refer to discrete categories, it is a classification task. When they take continuous values, it is a regression task, and they are often called targets. In unsupervised learning, the examples are not labeled. The algorithm must extract patterns from the data (for example, clustering, dimensionality reduction, etc). If the model learns noise (statistical fluctuations present in data) during training (in which case we say to have high variance), its predictive power decreases when applied to different data. In addition to what 27 CHAPTER 4. DEEP NEURAL NETWORKS is similar among the various examples of a given phenomenon, it also learns the peculiarities present in that specific sample. This high predictive power of the model for data on which it was trained but which decreases for different data is called overfitting. We say that the model does not generalize well. To avoid this, we use only part of the data to train the model. We use another part to validate the model, that is, to verify that the model is not overfitting. Finally, we use a third part to test the model (in this work we use that part to perform the analysis). In addition to this, other methods, called regularization, are used. Regularization consists in making the model less complex. Various techniques are used for the effect, some will be described later. If the model is less complex, it has more difficulty to fit so well the training data (we say that it has more bias). In that way, it is less able to fit to noise. 4.1.1 Example of regression: Linear Regression We are given some set of 𝑁 observations (examples) 𝑥𝑖 and corresponding targets 𝑦𝑖 , where 𝑖=1, . . . , 𝑁 , and want to predict 𝑦 for some new observation 𝑥 . Each observation has 𝐷 features. We want to build a model 𝑦(𝒙,𝒘)=𝑤0+𝑤1𝑥1+···+𝑤𝐷𝑥𝐷,(4.1) where 𝒙 is the D-dimensional feature vector and 𝒘 are the model parameters to be learned during training. To train the model (4.1) we need to define a cost function and minimize it to determine the parameters 𝒘. In linear regression, the Mean Squared Error 𝐶(𝒘)=1 𝑁 𝑁 Õ 𝑖=1(𝑦(𝒙𝒊,𝒘) −𝑦𝑖)2(4.2) is used. For linear regression, it is possible to minimize mean squared error in closed-form, but, in general, we need to use numerical methods. The Gradient Descent algorithm is commonly used. When minimizing a function, we need to compute its gradient, ∇𝐶(𝒘) , and find where it vanishes. We begin at some initial point in the parameter space. We define a positive learning rate 𝜒 to control the update of the parameters. We compute the gradient of the cost function C (as a function of the model parameters) at that location and move the parameters ˆ𝒘 in the direction opposite to the gradient (that is, in the direction of the steepest descent), in steps proportional to 𝜒, that is, ˆ𝒘←ˆ𝒘−𝜒∇𝒘𝐶|ˆ𝒘,(4.3) until the gradient is sufficiently small. 𝜒 is not learned but specified in advance by the analyst. Parameters of this kind, e.g., the architecture of artificial neural networks, the optimization algorithm (gradient descent being one of them), the cost function, are called hyperparameters. These are parameters that can not be learned from data. The learning rate should not be too small (learning will take longer) nor too large (the 28 CHAPTER 4. DEEP NEURAL NETWORKS steps that help avoid oscillations, making the convergence quicker. This work was based in the use of ADAM [48], a stochastic gradient descent algorithm that uses both improvements. In neural networks, several methods are used for regularization. Some of them were used in this work and will be briefly described. Batch normalization consists in standardize the output of a unit before using it as input to other units. This has a regularization effect. Dropout consist in randomly remove some units in each step of the optimization. This effect is controlled by a dropout rate parameter that specifies the fraction of units to drop. This is a hyperparameter that must be chosen. In this work,for tuning these and other hyperparameters the tool Optuna [49] was used. Early stopping consists in monitoring the performance of the model after each epoch of training in validation data and stop the training if it deteriorates. 35 Chapter 5 Strategy and Results In Chapter 3, it was described a strategy for selecting a signal-enriched region of the phase space, for an hypothetical Z’ signal coupling exclusively to 𝑡¯ 𝑡 . It was shown in a graphical way that the observed data was better described solely by the Standard Model, excluding the BSM hypothesis, building for the effect a discriminant variable, the approximate mass of the 𝑡¯ 𝑡system. In Section 5.2.1, a statistical procedure will be used to interpret the obtained results. More concretely, it will be assumed that the hypothesis of discovery of new physics was rejected and an exclusion limit for the mass of Z’ will be given, meaning that all hypothesized signals corresponding to masses lesser than that limit are excluded, within a certain confidence level, given the observed data. The main purpose of this work is to test if a better result, that is, a larger lower limit on the mass of Z’, can be found constructing a different discriminant variable using an artificial neural network. In the first approach, rectangular regions of phase space are selected for an enhanced ratio of signal to background ( 𝑠/√𝑏 ) for a distribution of a variable chosen for its physical meaning. The machine learning approach will build a new variable that is a non-linear function of the feature space, which distribution better discriminate signal-rich from signalpoor regions. Two variables of this kind will be built, corresponding to two different neural networks. The difference between the two is that for one of them, in addition to low level features, one of more high level is used, namely, the approximate mass of the 𝑡¯ 𝑡 system for the same events selected in the traditional approach and that takes the value zero otherwise, that will be called tt_m. Moreover, the event selection for the machine learning approach is less strict, having in common with the cut-based approach all the cuts already described, mostly to guarantee the correct topological final state, excepting one. That is, the one that demands that there is a single large radius jet that is top-tagged and well separated from the small radius jet from the leptonic decay. A different part of the work consists in studying how well a neural network trained to discriminate a particular signal will perform when used to separate a different signal from the same background. 36 CHAPTER 5. STRATEGY AND RESULTS 5.1 Neural Network Training The new discriminant variables used as discriminant are single-valued non-linear functions of input features of neural networks that determine their output. The output layer is therefore constituted by a single unit, which activation is the logistic function. The network is trained to recognize a given signal through the use of an appropriate cost function, the binary cross-entropy. The event selection was made using Root but building a dataset in tabular form. For that end it was recorded for each event characteristics of a fixed number of jets (10), and of large-R jets (6). The identified jets and large-R jets for each event are sorted by descending order of their transverse momenta. When the recorded number is larger than the actual number, the proprieties of non existent jets are set to zero. Used as features are the following: • The multiplicity of small-R jets close to the lepton that are classified as good 1. • The multiplicity of good b-tagged jets. • The energy, transverse momentum, charge, type, pseudorapidity and azimuthal angle of the lepton. • The energy, transverse momentum, pseudorapidity and azimuthal angle of the jets, and the value of MV2c10, the discriminant used to perform b-tagging. • The energy, invariant mass, transverse momentum, pseudorapidity and azimuthal angle and two additional features of large-R jets. These two last features are 𝜏32 , the weight from the algorithm for top-quark tagging, and a boolean variable that identifies that particular jet as the single large-R jet that is top-tagged and well separated from the small radius jet, for the events for which this condition is satisfied, and is set always to false otherwise. • In one of neural networks, tt_m is also used as a feature. To each event is associated a physical weight (here named so to distinguish them for other kind of weights, like the parameters of the DNNs or the class weights and MC weights described below). By this, it is taken in account that MC simulated events are generated with different luminosities than data and even with distorted distributions. Also, for different physical processes the efficiency of detection is different. Moreover, there are weights assigned to events by the MC generator 𝑤𝑔 (related to specific ways of implementing integration). In the end, a final weight is calculated as follow: 𝑤=𝐹𝜎𝐿𝑤𝑔/Õ𝑤𝑔,(5.1) where 𝐿 is the integrated luminosity (a third of 10 fb−1 ), 𝜎 is the cross-section for the process, and the fraction 𝑤𝑔/Í𝑤𝑔 gives us the inverse of the number of generated events. The sum is over the events 1They have to pass a pile-up exclusion filter and to have an appropriate location in the calorimeter. 37 CHAPTER 5. STRATEGY AND RESULTS corresponding to the same process 2 . 𝐹 represents a combination of scale factors, including for b-tagging, pile-up, etc. The dataset is equally divided in three parts. One part is used for training the model, other for performing validation assessment, and the last to apply the model and obtain results. During training, when computing loss, each event is given a weight that correspond to his physical weights but normalized in the sense that the sum of weights for signal events is equal to the sum of weights for background events, guaranteeing equal importance is given to both classes. The neural network has 100 or 101 units in the input layer, corresponding to the number of features and one in the output layer. The number of hidden layers and the number of units per layer are considered hyperparameters, determined using Optuna. They can assume values between 1 and 9 and between 40 and 60, respectively. The activation function used in units belonging to hidden layers is the ReLU. The outputs of all units except the single unit of the output layer are subject to batch normalization. Instead, the input for this last unit is subject to dropout. The dropout rate, allowed to take values from 0 to 0.5 and the learning rate, allowed to vary between 10−5 and 10−2 , are, jointly with the number of hidden layers and the number of units per layer, hyperparameters. Optuna will search in the hyperparameter space, in the predefined ranges stated above, the combination that will maximize the area under the ROC curve (AUC) of the neural network (the objective function), conveniently pruning unpromising trials to speedup the process. For efficiently sampling the hyperparameter space the Tree-structured Parzen Estimator algorithm [50] is used. It begins with a random search but the history search is taken in account to suggest new values to search based in a probability model of the objective function. This extra effort put in determining the direction of the search in each step is more than compensated by the reduced number of calls to the objective function (each requiring a training session to evaluate the AUC.) The results are presented in Tables 5.1 and 5.2. Z’ mass (GeV) number of hidden layers number of units in each layer learning rate dropout rate 1000 10 56, 48, 58, 41, 54, 54, 44, 59, 60 0.008169 0.5 1250 3 51, 49 0.007106 0.4 1500 9 42, 46, 60, 52, 47, 60, 49, 56 0.009649 0.1 1750 4 50, 50, 48 0.009540 0.3 2000 7 53, 43, 50, 44, 56, 59 0.009865 0.0 Table 5.1: Hyperparameters for NNs for which tt_m was not used as a feature. The training was done using batches of size 1024 for 100 epochs using the early stopping method monitoring the auc with patience of 10 epochs. The performance assessment of the trained DNN is shown in Figure 5.1. 2This procedure is needed because there is a different weight for each event. 38 CHAPTER 5. STRATEGY AND RESULTS 0.0 0.2 0.4 0.6 0.8 1.0 Falsepositiverate 0.0 0.2 0.4 0.6 0.8 1.0 Truepositiverate Z'1000GeV(AUC:0.79) Z'1250GeV(AUC:0.86) Z'1500GeV(AUC:0.90) Z'1750GeV(AUC:0.92) Z'2000GeV(AUC:0.93) (a) ROC curves evaluating the performance of DNNs that don’t use tt_m as feature trained with different signals. 0.0 0.2 0.4 0.6 0.8 1.0 Falsepositiverate 0.0 0.2 0.4 0.6 0.8 1.0 Truepositiverate Z'1000GeV(AUC:0.80) Z'1250GeV(AUC:0.86) Z'1500GeV(AUC:0.90) Z'1750GeV(AUC:0.92) Z'2000GeV(AUC:0.93) (b) ROC curves evaluating the performance of DNNs that use tt_m as feature trained with different signals. Figure 5.1: Performance assessment of DNNs trained to discriminate signals corresponding to Z’ of different masses from background. 39 CHAPTER 5. STRATEGY AND RESULTS Z’ mass (GeV) number of hidden layers number of units in each layer learning rate dropout rate 1000 5 50, 44, 57, 54, 45 0.009198 0.3 1250 6 46, 58, 54, 51, 49, 59 0.009825 0.3 1500 3 53, 55, 59 0.009621 0.2 1750 7 40, 44, 55, 50, 44, 47, 48 0.008370 0.0 2000 1 52 0.006612 0.2 Table 5.2: Hyperparameters for NNs for which tt_m was used as a feature. 5.2 Exclusion Limits 5.2.1 Exclusion Limits The goal of a search in HEP is either to claim that a signal has been discovered or to exclude the presence of the signal. However, a search that has as result an exclusion convey information, notwithstanding. This information is presented as lower or higher limits. In our case that would be to state that if a boson Z’ exists at all it would have to have a mass higher than a certain limit (a lower limit, then) because all signals with a mass smaller than that limit were excluded. The procedure used, the so-called 𝐶𝐿𝑠 method [51], follows closely a standard statistical analysis based on a hypothesis test. However, it departs from it to deal with the lack of sensitivity that occurs when the expected signal is extremely low. In a standard analysis, there is a hypothesis to be rejected, called the null hypothesis ( 𝐻0 ), and an alternative hypothesis ( 𝐻1 ). A quantity q is defined, the test statistic, that is a function of the data. That quantity has a known distribution 𝑓(𝑞|𝐻0) if 𝐻0 is correct. When an observation is made it is then possible to say how probable it is to get an observation that is at least so extreme, assuming that the 𝐻0 is correct. This is made placing the test statistic observed in its distribution. The use of an alternative hypothesis allow us to define what means more extreme, in the sense that the observation is increasingly more likely to be explained by 𝐻1 instead of 𝐻0 . The probability just described is called the p-value for that observation given 𝐻0 . 𝐻0 is rejected at confidence level of 95% if the p-value is less than 0.05. Let us consider the case where the null hypothesis 𝐻𝑠+𝑏 correspond to assume that we have background plus signal and 𝐻𝑏) is the alternative hypothesis that assumes that only background is present. Using the method just described, 𝐻𝑠+𝑏would be rejected at confidence level 95% if 𝑝𝑠+𝑏=∫∞ 𝑞obs 𝑓(𝑞|𝐻𝑠+𝑏)d𝑞<0.05,(5.2) assuming that smaller values of q are more compatible with 𝐻𝑠+𝑏. Likewise, 𝑝𝑏=∫𝑞obs −∞ 𝑓(𝑞|𝐻𝑏)d𝑞 , (5.3) 40 CHAPTER 5. STRATEGY AND RESULTS is the p-value of the observation given 𝐻𝑏 . It could happen that the distributions 𝑓(𝑞|𝐻𝑠+𝑏) and 𝑓(𝑞|𝐻𝑏) almost overlap completely 3 . In this case, the standard analysis, in some cases would lead us to reject 𝐻𝑠+𝑏 with probability close to 5 % in situations to which one has little sensitivity. In order to mitigate this effect, a new ”p-value”is defined, the 𝐶𝐿𝑠: 𝐶𝐿𝑠= 𝑝𝑠+𝑏 1−𝑝𝑏 .(5.4) It can be noticed that when the two distributions are well separated, the denominator is almost 1, and the standard p-value is recovered. When there is no sensitivity, the denominator makes the new ”p-value”larger, allowing that less observations will reject the hypothesis of signal being present. So, in the 𝐶𝐿𝑠 method, the rejected models are only a part of the models that would be rejected using 𝑝𝑠+𝑏 as p-value, being therefore more conservative. In our problem we will test several hypotheses of the type s + b, meaning, 𝑏+𝜇𝑠 , considering different signal strengths 𝜇 , taking values from 0 to 1. An upper limit for 𝜇upper will be found, meaning that all hypotheses with 𝜇>𝜇upper were rejected. This is made adjusting 𝜇 in such a way that 𝐶𝐿𝑠(𝜇)= 𝑝𝜇𝑠+𝑏 1−𝑝𝑏 =0.05 For different masses of Z’, different values of 𝜇upper are found. Then, a value for the mass of Z’ will be calculated that corresponds to a full rejection of the signal hypothesis, that is, for which 𝜇upper =1. All masses lesser than that limit are rejected. The computation of the limits was made using the Python tool pyhf [52]. The test statistic used was ˜𝑞𝜇 appropriated for calculate upper limits for signal strength as presented in Reference [53]. The results for masses lower limits are in Table 5.3 calculated from results from signal strength upper limits presented in Figure 5.5. 5.2.2 Results The following results will be relative to a luminosity of 3.3 fb−1 . The outputs of the DNNs, the predictions for all inputs reserved in the application set (also called test set), shown in Figures 5.2 and 5.3, can be interpreted as probability distributions of an event being the signal in which the DNN was trained. In these plots and the ones in Figure 5.4, the signal is amplified in a way that its integral has the same value has the integral of the background. The signal events are extremely rare. In Figure 5.4 it is shown the distribution of the approximate mass of the 𝑡¯ 𝑡 system. There, events known to correspond to signal or to one of different backgrounds are represented separately, the backgrounds being stacked. It can be seen in Figures 5.1 that DNNs trained with signals corresponding to higher masses have better performances. This is due to the fact that those signals are more distinct from the background as it is shown in Figures [5.2-5.4]. The use of a higher-level feature has a positive impact on the performance as can be seen comparing Figures 5.1a and 5.1b. This feature by itself already encode information that 3That is, not only in the tail region. 41 CHAPTER 5. STRATEGY AND RESULTS allows the process of discrimination. This discussion will be retaken in Section 5.3. The lower limits for the mass of Z’ computed according to the method described in Subsection 5.2.1 are presented in Table 5.3. The same hierarchy among the methods used is obtained. The lower limit calculated based in a DNN that has a high-level feature is higher than the limit obtained using a DNN without the higher-level feature. Using a more traditional analysis, making rectangular cuts in the phase-space and using a physical observable know to have discriminant power results in a smaller lower limit. This is the result of a loss of discriminant power as we pass from one to other of these methods enumerated above. The conclusion is that a non-linear function of the features learned in an automated way can surpass the observable traditionally used based in the knowledge of basic physics, but that that knowledge can help if added to the machine learning workflow. DNN based without tt_m as feature DNN based with tt_m as feature Based on tt_m distribution 1350.21 1364.68 1322.41 Table 5.3: 95% CL lower limits on the Z’ mass (in GeV) at a luminosity of 3.3 fb−1. 5.3 Study of transferability. In this section is presented a study of the transferability of DNNs trained to discriminate a specific Z’ signal from background to be used as a good discriminant if different signals are present amidst the same background. The study was made for DNNs that use and for those that do not use the high-level feature, each one trained for a different Z’ signal. They were used Z’ signals with masses: 1000, 1250, 1500 and 1750 GeV. Then, all DNNs are employed to separate all signals available from the same background. The results are presented in Figures [5.6 -5.9]. It stands out that neural networks trained with signals of larger mass retain more ability to discriminate different signals from background. This follows from the fact evidenced in Figures [5.2-5.4] that the signal is more different than background when the mass of Z’ is larger. This not only explains the better AUC scores for DNNs trained on larger signal masses as shown in Figures 5.1 but can explain also the resilience of these DNNs. The neural networks learn a boundary in the features space between background and signal. As the signal is more different from background, more sharply outlined is the boundary to isolate the background. The ability of all DNNs to isolate the background explains, albeit in different degrees, their ability to discriminate what is not background. Furthermore, the larger the similarity between the signal used for training and the signal being discriminated, the better is the performance of the network. This follows from the fact that the DNNs also learn how to isolate the signal, in addition to isolate the background. When the signal is barely distinguishable from background, as it happens for the mass 1 TeV, it can occur that the DNNs trained with it loose all ability to discriminate when the signals are very different (c.f. the black squares in heatmaps 5.7 and 5.9). 42 CHAPTER 5. STRATEGY AND RESULTS (a) DNN trained in signal Z’ 1000 GeV (b) DNN trained in signal Z’ 1250 GeV (c) DNN trained in signal Z’ 1500 GeV (d) DNN trained in signal Z’ 1750 GeV (e) DNN trained in signal Z’ 2000 GeV Figure 5.2: Outputs of DNNs that don’t use tt_m as feature. 43 CHAPTER 5. STRATEGY AND RESULTS (a) DNN trained in signal Z’ 1000 GeV (b) DNN trained in signal Z’ 1250 GeV (c) DNN trained in signal Z’ 1500 GeV (d) DNN trained in signal Z’ 1750 GeV (e) DNN trained in signal Z’ 2000 GeV Figure 5.3: Outputs of DNNs that use tt_m as feature. 44 CHAPTER 5. STRATEGY AND RESULTS Figure 5.9: Heatmap of the AUC of the DNNs that use tt_m as feature. In the vertical axis is represented the signal with wich the DNN was trained. In the horizontal axis is represented the signal used for prediction. 51 Chapter 6 Conclusions and future work It is possible to conclude from this work that the use of DNNs in establishing exclusion limits in searches for new physics is an improvement over the use of a more traditional approach. Nonetheless, the incorporation of previous knowledge in the form of physically motivated features in the building of the neural networks increases even more their performance. This, apparently, is not a novel result as this constitutes the field of feature engineering. But it also could have happened that all the information contained in the physically motivated feature had be learned by the DNN making it redundant. It is observed that transferability of DNNs occurs but degrades as signals to discriminate becomes more different from the ones used during training. Also stands out that the better situation when one is preparing a single DNN for use to discriminate different signals is to train it with the signal that is more distinct from background. In the case studied, for larger masses of Z’. It must be noticed that the comparisons were limited to basically the same model of Z’, only varying the mass. Furthermore the comparison with analyses not based in machine learning was restricted to a particular case. These limitations could be tackled in a future work. Additionally, this work could be continued in the future exploring the following points. Further comparisons could be done using different models of the boson Z’. How resilient the DNNs will be if the signal will be more different, when they depart from the background in different regions of the phase-space? Another possible study would be to compare Deep Neural Networks with other multivariate methods, for example, Boosted Decision Trees, a method commonly used in data analysis in High-Energy Physics. In addition to the comparison with other supervised methods, a comparison with unsupervised or semi-supervised methods could be performed. The absence of supervision to tell where signal is induce naturally a degradation of the performance to discriminate but these methods can be good at recognizing the background, and detecting anomalies from it. Is the degradation of the discriminative power inherent in the unsupervising nature of the method larger than the loss of ability of the DNNs studied in this work to discriminate signals different from the specific signal to which they were presented during the the supervising training? It would be also interesting to study the the systematic uncertainties related to the detectors and to the modelling of backgrounds, which 52 CHAPTER 6. CONCLUSIONS AND FUTURE WORK are out of the scope of this thesis. 53 Bibliography [1] Wikipedia, the free encyclopedia. Standard model of elementary particles. [Online; accessed October 25, 2021]. 2019. url: https://commons.wikimedia.org/wiki/File:Standard_Model_ of_Elementary_Particles.svg. [2] S. Tomonaga. “On a Relativistically Invariant Formulation of the Quantum Theory of Wave Fields*.” In: Progress of Theoretical Physics 1.2 (Aug. 1, 1946), pp. 27–42. issn: 0033-068X. doi: 10. 1143/PTP.1.27. [3] J. Schwinger. “On Quantum-Electrodynamics and the Magnetic Moment of the Electron.” In: Physical Review 73.4 (Feb. 15, 1948), pp. 416–417. doi: 10.1103/PhysRev.73.416. [4] J. Schwinger. “Quantum Electrodynamics. I. A Covariant Formulation.” In: Physical Review 74.10 (Nov. 15, 1948), pp. 1439–1461. doi: 10.1103/PhysRev.74.1439. [5] R. P. Feynman. “The Theory of Positrons.” In: Physical Review 76.6 (Sept. 15, 1949), pp. 749–759. doi: 10.1103/PhysRev.76.749. [6] R. P. Feynman. “Space-Time Approach to Quantum Electrodynamics.” In: Physical Review 76.6 (Sept. 15, 1949), pp. 769–789. doi: 10.1103/PhysRev.76.769. [7] R. P. Feynman. “Mathematical Formulation of the Quantum Theory of Electromagnetic Interaction.” In: Physical Review 80.3 (Nov. 1, 1950), pp. 440–457. doi: 10.1103/PhysRev.80.440. [8] G. ’. Hooft. “Gauge Theories of the Forces between Elementary Particles.” In: Scientific American 242.6 (1980), pp. 104–141. issn: 00368733, 19467087. [9] C. N. Yang and R. L. Mills. “Conservation of Isotopic Spin and Isotopic Gauge Invariance.” In: Physical Review 96.1 (Oct. 1, 1954), pp. 191–195. doi: 10.1103/PhysRev.96.191. [10] A. Salam and J. C. Ward. “Electromagnetic and Weak Interactions.” In: Physics Letters 13.2 (Nov. 15, 1964), pp. 168–171. issn: 0031-9163. doi: 10.1016/0031-9163(64)90711-5. [11] S. Weinberg. “A Model of Leptons.” In: Physical Review Letters 19.21 (Nov. 20, 1967), pp. 1264– 1266. doi: 10.1103/PhysRevLett.19.1264. 54 BIBLIOGRAPHY [12] S. L. Glashow. “Partial-Symmetries of Weak Interactions.” In: Nuclear Physics 22.4 (Feb. 1, 1961), pp. 579–588. issn: 0029-5582. doi: 10.1016/0029-5582(61)90469-2. [13] E. Abers and B. Lee. “Gauge Theories.” In: Physics Reports 9.1 (Nov. 1, 1973), pp. 1–2. issn: 0370-1573. doi: 10.1016/0370-1573(73)90027-6. [14] D. Griffiths. Introduction to Elementary Particles. Wiley, Dec. 1987. doi: 10.1002/9783527618460 . [15] K. Olive. “Review of Particle Physics.” In: Chinese Physics C 38.9 (Aug. 2014), p. 090001. doi: 10.1088/1674-1137/38/9/090001. [16] C. T. Hill and S. J. Parke. “Top Quark Production: Sensitivity to New Physics.” In: Physical Review D49.9 (May 1, 1994), pp. 4454–4462. doi: 10.1103/PhysRevD.49.4454. [17] C. T. Hill. “Topcolor Assisted Technicolor.” In: Physics Letters B 345.4 (Feb. 23, 1995), pp. 483– 489. issn: 0370-2693. doi: 10.1016/0370-2693(94)01660-5. [18] R. M. Harris, C. T. Hill, and S. J. Parke. Cross Section for Topcolor Z’ Decaying to Top-Antitop. Nov. 9, 1999. arXiv: hep-ph/9911288. [19] The ATLAS Collaboration. “Search for Heavy Particles Decaying into Top-Quark Pairs Using Leptonplus-Jets Events in Proton-Proton Collisions at √𝑠=13 TeV with the ATLAS Detector.” In: The European Physical Journal C 78.7 (July 2018), p. 565. issn: 1434-6044, 1434-6052. doi: 10. 1140/epjc/s10052-018-5995-6. arXiv: 1804.10823. [20] Particle Data Group. “Review of Particle Physics.” In: Progress of Theoretical and Experimental Physics 2020.8 (Aug. 2020). issn: 2050-3911. doi: 10.1093/ptep/ptaa104. [21] E. Mobs. “The CERN accelerator complex. Complexe des accélérateurs du CERN.” In: (July 2016). General Photo. url: https://cds.cern.ch/record/2197559. [22] L. Evans and P. Bryant. “LHC Machine.” In: 3.08 (Aug. 2008), S08001–S08001. issn: 17480221. doi: 10.1088/1748-0221/3/08/S08001 . url: https://doi.org/10.1088/17480221/3/08/s08001. [23] The ATLAS Collaboration. “The ATLAS Experiment at the CERN Large Hadron Collider.” In: Journal of Instrumentation 3.08 (Aug. 2008), S08003–S08003. issn: 1748-0221. doi: 10.1088/17480221/3/08/S08003. [24] The ALICE Collaboration. “The ALICE Experiment at the CERN LHC.” In: Journal of Instrumentation 3.08 (Aug. 2008), S08002–S08002. issn: 1748-0221. doi: 10.1088/1748-0221/3/08/S08002 . [25] The CMS Collaboration. “The CMS Experiment at the CERN LHC.” In: Journal of Instrumentation 3.08 (Aug. 2008), S08004–S08004. issn: 1748-0221. doi: 10.1088/1748-0221/3/08/S08004 . [26] The LHCb Collaboration. “The LHCb Detector at the LHC.” In: Journal of Instrumentation 3.08 (Aug. 2008), S08005–S08005. issn: 1748-0221. doi: 10.1088/1748-0221/3/08/S08005. 55 BIBLIOGRAPHY [27] The ATLAS Collaboration. “Observation of a New Particle in the Search for the Standard Model Higgs Boson with the ATLAS Detector at the LHC.” In: Physics Letters B 716.1 (Sept. 17, 2012), pp. 1–29. issn: 0370-2693. doi: 10.1016/j.physletb.2012.08.020. [28] The CMS Collaboration. “Observation of a New Boson at a Mass of 125 GeV with the CMS Experiment at the LHC.” In: Physics Letters B 716.1 (Sept. 17, 2012), pp. 30–61. issn: 0370-2693. doi: 10.1016/j.physletb.2012.08.021. [29] The ATLAS Collaboration. “Alignment of the ATLAS Inner Detector in Run 2.” In: The European Physical Journal C 80.12 (Dec. 24, 2020), p. 1194. issn: 1434-6052. doi: 10.1140/epjc/ s10052-020-08700-6. [30] J. Pequenão. “Computer Generated image of the ATLAS calorimeter.” Mar. 2008. url: https: //cds.cern.ch/record/1095927. [31] J. Pequenão. “Computer generated image of the ATLAS Muons subsystem.” Mar. 2008. url: https: //cds.cern.ch/record/1095929. [32] The ATLAS Collaboration. “Performance of the ATLAS Trigger System in 2015.” In: The European Physical Journal C 77.5 (May 18, 2017), p. 317. issn: 1434-6052. doi: 10.1140/epjc/s10052017-4852-3. [33] The ATLAS Collaboration. Review of the 13 TeV ATLAS Open Data Release. ATL-OREACH-PUB-2020001. Jan. 24, 2020. url: https://cds.cern.ch/record/2707171. [34] R. Brun and F. Rademakers. “ROOT — An Object Oriented Data Analysis Framework.” In: Nuclear Instruments and Methods in Physics Research Section A: Accelerators, Spectrometers, Detectors and Associated Equipment. New Computing Techniques in Physics Research V 389.1 (Apr. 11, 1997), pp. 81–86. issn: 0168-9002. doi: 10.1016/S0168-9002(97)00048-X. [35] The ATLAS Collaboration. “Electron Reconstruction and Identification in the ATLAS Experiment Using the 2015 and 2016 LHC Proton-Proton Collision Data at √𝑠=13 TeV.” In: The European Physical Journal C 79.8 (Aug. 3, 2019), p. 639. issn: 1434-6052. doi: 10.1140/epjc/s10052019-7140-6. [36] The ATLAS Collaboration. “Muon reconstruction performance of the ATLAS detector in protonproton collision data at √𝑠=13 TeV.” In: The European Physical Journal C 76.5 (May 2016). doi: 10.1140/epjc/s10052-016-4120-y . url: https://doi.org/10.1140/epjc/s10052016-4120-y. [37] M. Cacciari, G. P. Salam, and G. Soyez. “The anti-ktjet clustering algorithm.” In: Journal of High Energy Physics 2008.04 (Apr. 2008), pp. 063–063. doi: 10.1088/1126-6708/2008/04/063 . url: https://doi.org/10.1088/1126-6708/2008/04/063. 56 BIBLIOGRAPHY [38] The ATLAS Collaboration. Tagging and Suppression of Pileup Jets with the ATLAS Detector. ATLASCONF-2014-018. ATLAS-COM-CONF-2014-025, May 12, 2014. url: https://cds.cern.ch/ record/1700870. [39] D. Krohn, J. Thaler, and L. - T. Wang. “Jet Trimming.” In: Journal of High Energy Physics 2010.2 (Feb. 24, 2010), p. 84. issn: 1029-8479. doi: 10.1007/JHEP02(2010)084. [40] The ATLAS Collaboration. Performance of jet substructure techniques in early √𝑠=13 TeV 𝑝𝑝 collisions with the ATLAS detector. Tech. rep. Geneva: CERN, Aug. 2015. url: https://cds. cern.ch/record/2041462. [41] The ATLAS Collaboration. “Measurements of B-Jet Tagging Efficiency with the ATLAS Detector Using $$ T\overline{t} $$events at $$ \sqrt{s}=13 $$TeV.” In: Journal of High Energy Physics 2018.8 (Aug. 16, 2018), p. 89. issn: 1029-8479. doi: 10.1007/JHEP08(2018)089. [42] The ATLAS Collaboration. “Measurements of Top-Quark Pair Differential Cross-Sections in the Lepton+jets Channel in Pp Collisions at √𝑠=13 TeV Using the ATLAS Detector.” In: Journal of High Energy Physics 2017.11 (Nov. 28, 2017), p. 191. issn: 1029-8479. doi: 10 . 1007 / JHEP11(2017)191. [43] The ATLAS Collaboration. Boosted hadronic top identification at ATLAS for early 13 TeV data. Tech. rep. Geneva: CERN, Dec. 2015. url: https://cds.cern.ch/record/2116351. [44] I. Goodfellow, Y. Bengio, and A. Courville. Deep Learning. http://www.deeplearningbook.org . MIT Press, 2016. [45] C. M. Bishop. Pattern Recognition and Machine Learning. Springer, 2006. [46] M. Minsky and S. A. Papert. Perceptrons; an introduction to computational geometry. MIT Press, 1969. [47] X. Glorot, A. Bordes, and Y. Bengio. “Deep Sparse Rectifier Neural Networks.” In: Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics. Ed. by G. Gordon, D. Dunson, and M. Dudík. Vol. 15. Proceedings of Machine Learning Research. Fort Lauderdale, FL, USA: PMLR, 2011, pp. 315–323. url: https://proceedings.mlr.press/v15/glorot11a. html. [48] D. P. Kingma and J. Ba. Adam: A Method for Stochastic Optimization. 2017. arXiv: 1412.6980 [cs.LG]. [49] T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama. “Optuna: A Next-generation Hyperparameter Optimization Framework.” In: Proceedings of the 25rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2019. 57 BIBLIOGRAPHY [50] J. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl. “Algorithms for Hyper-Parameter Optimization.” In: Advances in Neural Information Processing Systems. Ed. by J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, and K. Q. Weinberger. Vol. 24. Curran Associates, Inc., 2011. [51] A. L. Read. “Presentation of Search Results: The CL s Technique.” In: Journal of Physics G: Nuclear and Particle Physics 28.10 (Oct. 1, 2002), pp. 2693–2704. issn: 0954-3899. doi: 10.1088/09543899/28/10/313. [52] L. Heinrich, M. Feickert, G. Stark, and K. Cranmer. “pyhf: pure-Python implementation of HistFactory statistical models.” In: Journal of Open Source Software 6.58 (2021), p. 2823. doi: 10.21105/ joss.02823. [53] G. Cowan, K. Cranmer, E. Gross, and O. Vitells. “Asymptotic Formulae for Likelihood-Based Tests of New Physics.” In: The European Physical Journal C 71.2 (Feb. 9, 2011), p. 1554. issn: 1434-6052. doi: 10.1140/epjc/s10052-011-1554-0. 58