scieee AI-readable full text Open interactive document viewer

Developing a Digital Twin for Discrete Event Processes Using Event Knowledge Graphs and Process Mining Techniques

Abu Sbeit, Abd Alrhman

Abstract

This thesis describes the development of a discrete events digital twin and the evaluation of the twin on a medical kits sterilization process. The research strategy that is considered for the thesis includes both primary and secondary research methodologies through investigating the available documentation at the company and the literature that is related to the topic of the thesis as well as by processing the required data from the sterilization process and the reports of the company and meeting the involved departments for a clearer understanding of the process. The methodology that is followed to achieve the target of the thesis includes analyzing the current sterilization process of the company and developing a digital twin model according to the process extracted from the event data that was provided by the company, then analyzing each activity in the sterilization process in order to replicate the process. In order to evaluate the discrete events digital twin, a comparison between the results of the digital twin and the original event data provided by the company was done for key performance indicators that include the number of sterilized kits per day, the time required for each activity, and the average time required to complete sterilization. According to the results of the evaluation, the developed digital twin shows a similarity with the original process and can be used as a digital replica for process planning and optimizations.

Full text

Eindhoven University of Technology MASTER Developing a Digital Twin for Discrete Event Processes Using Event Knowledge Graphs and Process Mining Techniques Abu Sbeit, A.A.M.S. Award date: 2024 Link to publication Disclaimer This document contains a student thesis (bachelor's or master's), as authored by a student at Eindhoven University of Technology. Student theses are made available in the TU/e repository upon obtaining the required degree. The grade received is not published on the document as presented in the repository. The required complexity or quality of research of student theses may vary by program, and the required minimum study period may vary in duration. General rights Copyright and moral rights for the publications made accessible in the public portal are retained by the authors and/or other copyright owners and it is a condition of accessing publications that users recognise and abide by the legal requirements associated with these rights. • Users may download and print one copy of any publication from the public portal for the purpose of private study or research. • You may not further distribute the material or use it for any profit-making activity or commercial gain Take down policy If you believe that this document breaches copyright please contact us providing details, and we will remove access to the work immediately and investigate your claim. Download date: 19. May. 2025 Department of Mathematics and Computer Science Process Analytics Developing a Digital Twin for Discrete Event Processes Using Event Knowledge Graphs and Process Mining Techniques Master Thesis A.A.M.S. ABU SBEIT [email protected] Supervisors: D. Fahland [email protected] 25-08-2024 Keywords: Digital Twins; Process Mining; Event Knowledge Graphs; Process Modeling; Simulation; Acknowledgment The entire duration of the thesis was an exciting, challenging, and knowledgeable journey. I have experienced many new learnings and challenges not only in the technical areas but also on a personal level. I would like to take this opportunity to express how grateful I am to all the people who have been very supportive throughout the entire period of my thesis. First of all, I would like to convey my profound gratitude towards my supervisor Prof. Dirk Fahland for giving me the opportunity to work on this topic, guiding, encouraging, and inspiring my efforts over the course of my research. Without his constructive feedback and guidance, my thesis research would never have been emerged. Thank you for your flexibility and for enabling Teams meetings. I am also sincerely grateful to the European Union and the universities participating in this program for giving me the opportunity to be part of it. Last but not least, I would like to express my gratitude to my family and friends for their continued support, encouragement, and motivation as well as their constructive feedback and advice during the course of the research. This accomplishment would never have been possible without all of you. Thank you very much. 2 Abstract This thesis describes the development of a discrete events digital twin and the evaluation of the twin on a medical kits sterilization process. The research strategy that is considered for the thesis includes both primary and secondary research methodologies through investigating the available documentation at the company and the literature that is related to the topic of the thesis as well as by processing the required data from the sterilization process and the reports of the company and meeting the involved departments for a clearer understanding of the process. The methodology that is followed to achieve the target of the thesis includes analyzing the current sterilization process of the company and developing a digital twin model according to the process extracted from the event data that was provided by the company, then analyzing each activity in the sterilization process in order to replicate the process. In order to evaluate the discrete events digital twin, a comparison between the results of the digital twin and the original event data provided by the company was done for key performance indicators that include the number of sterilized kits per day, the time required for each activity, and the average time required to complete sterilization. According to the results of the evaluation, the developed digital twin shows a similarity with the original process and can be used as a digital replica for process planning and optimizations. Contents 1 Introduction 4 1.1 ContextandTopic ........................ 5 1.2 StateoftheArt.......................... 6 1.3 Research Question ........................ 7 1.4 Method .............................. 7 1.5 Findings .............................. 10 2 Background 13 2.1 Context Understanding ...................... 13 2.2 Data Understanding ....................... 16 3 Problem Exposition 19 3.1 Preliminaries ........................... 20 3.2 RelatedWork ........................... 24 3.3 Detailed Research Questions ................... 27 4 Research Design and Methodology 28 4.1 Research strategy ......................... 28 4.2 Methodology ........................... 29 4.2.1 Phase One: Digital Twin Modeling ........... 29 4.2.2 PhaseTwo:DigitalTwinSimulation .......... 31 4.2.3 PhaseThree:Evaluation................. 33 5 Digital Twin Modeling 34 5.1 Raw Data Preprocessing ..................... 34 5.2 EventKnowledgeGraph ..................... 37 5.3 Cases................................ 40 5.4 Process Model ........................... 45 5.5 ContextualEnrichment...................... 46 2 6 Digital Twin Simulation 51 6.1 SimulationSoftwares ....................... 52 6.2 SimulationParameterization................... 54 6.2.1 ArrivalRateAnalysis................... 54 6.2.2 Dispatching Rule ..................... 56 6.2.3 ActivityAnalysis ..................... 57 6.2.4 WorkingandWaitingTimesAnalysis.......... 70 6.2.5 Workforce Behavior and Scheduling ........... 71 6.2.6 Machinery Analysis .................... 73 6.3 Simulation Implementation .................... 74 6.4 Conclusion............................. 77 7 Evaluation 78 7.1 Objective ............................. 78 7.2 Setup................................ 79 7.3 Execution ............................. 82 7.4 Results............................... 82 7.4.1 CongestionComparison ................. 83 7.4.2 Goodness-of-Fit Evaluation ............... 85 7.4.3 TemporalComparison .................. 91 7.5 Discussion............................. 94 8 Conclusion 96 8.1 SWOTAnalysis.......................... 96 8.2 Conclusion,Recommendation,andFutureWork........ 97 APPENDICES 100 A Distribution Types Supported By Scylla Software 105 3 Chapter 1 Introduction The main goal of any kind of business is to increase profit and minimize expenses without interrupting the work. One way of achieving this goal can be by optimizing their operations and enhancing decision-making using information technology. One such technology is digital twins. Grieves and Vickers [13] define digital twins as ”a set of virtual information constructs that fully describes a potential or actual physical manufactured product from the micro atomic level to the macro geometrical level”. In this context, a digital twin represents a digital replica of a physical system, connected throughout its entire lifecycle, enabling real-time monitoring, analysis, and optimization. The implementation of digital twins has been encouraged by advancements in sensors, the Internet of Things (IoT), machine learning, and artificial intelligence (AI), which collectively enable the accurate digital representation of complex systems [18]. However, despite these technological advancements, many small-to-medium-sized enterprises (SMEs) face challenges in adopting digital twins due to limited expertise and resources. A survey conducted by Gartner in 2019 revealed that digital twin technology was entering mainstream use, with 75 percent of IoT organizations either utilizing or planning to utilize digital twins by 2020 [11]. Furthermore, Gartner estimated in 2022 that by 2027, over 40 percent of large organizations worldwide will leverage digital twins in their metaverse-based projects to drive revenue growth [12]. While the adoption of digital twins is expanding, there remains a gap 4 in their implementation, particularly among SMEs. Addressing this gap requires targeted strategies that account for the unique challenges faced by these businesses, ensuring that they too can benefit from the significant advantages that digital twins offer in terms of process optimization, decisionmaking, and business growth. The primary objective of this research is to establish a comprehensive guideline for constructing and implementing a discrete events-based digital twin, derived from historical event data, applicable across various industries. The methodology is illustrated through a concrete example of simulating a sterilization process of medical kits, thereby identifying challenges inherent in developing digital twins. The proposed approach addresses these challenges systematically, demonstrating its effectiveness in producing relevant and actionable insights when applied to the specific case study. This research aims to facilitate the development of digital twins with greater ease and efficacy, enabling businesses to achieve their particular objectives through a well-defined process. 1.1 Context and Topic A digital twin is a virtual representation of a tangible object, process, or system. This digital representation can be used for various purposes, including simulation, analysis, and control [14]. Auto-Twin [2] is a research project aiming to develop digital twins. This research project aims to address the limitations of current systems engineering models by introducing a novel method for automated process-aware discovery, facilitating the generation of autonomous digital twins by leveraging data from sensors and other sources to create accurate, real-time models that reflect the current state and behavior of their original equivalent. Auto-Twin is seeking an implementation of a digital twin for one of its SMEs’ use cases. By implementing a digital twin, the Auto-Twin project aims to utilize the digital twin in order to enhance decision-making and simplify process optimization. The intended use-case is a leading company in supplying and managing surgical instruments, they are seeking an optimization for one of their medical instruments’ sterilization process, involving multiple entities contributing 5 to different stages of the sterilization process. In this research, we will be looking into a method to build a digital twin based on multidimensional behavioral understanding by employing event knowledge graphs and process mining, by leveraging a medical kit sterilization process historical event data, exported by the Auto-Twin use cases in order to be used for simulation and process optimization. 1.2 State of the Art Digital twins have evolved from a conceptual ideal for Product Lifecycle Management (PLM) [13] to a crucial element in the digitalization of smart businesses [18]. Recent research has emphasized the importance of advanced technologies, such as event knowledge graphs and process mining, in implementing accurate digital representations of physical systems. Van der Aalst [25] highlights the critical role of process mining techniques, such as process discovery and conformance checking, in constructing Discrete-Event Simulations (DES) and improving their accuracy. While these techniques significantly contribute to the development of simulations that closely mirror real-world scenarios, they fall short in providing a multidimensional perspective on the processes. To address this limitation, Van der Aalst [26] proposes integrating object-centric process mining (OCPM) into simulation modeling, emphasizing the necessity of OCPM to enhance the realism and depth of simulations. Building on the concept of OCPM, the incorporation of knowledge graphs into OCPM further enhances simulation models by offering a structured, multidimensional representation of the relationships and interactions among various entities within the process. Stavropoulo et al. (2024) [19] previously introduced a solution for designing intelligent manufacturing processes by employing a knowledge graph as a data store accessible by a digital twin. Their research demonstrated that utilizing a knowledge graph not only provides additional context based on the available data but also significantly improves process understanding and the accuracy of the digital twin. While existing solutions for digital twins effectively mirror real-world scenarios, they are typically designed for large enterprises with established event-logging policies and dedicated data-processing departments. This research aims to test the feasibility of adapting similar techniques to create 6 Chapter 2 Background Auto-Twin, a European-funded research project represents an innovative method for creating digital twins, which are virtual replicas of physical systems. This innovative method aims to overcome the limitations of the current way of creating system engineering models by introducing a breakthrough method for automated process-aware discovery towards autonomous digital twins generation [2]. By introducing an automated process aware discovery, Auto-Twin seeks to significantly enhance the efficiency and scalability of their developed digital twins. Through Auto-Twin, a leading supplier and manager of surgical instruments seeking to optimize their medical kit sterilization process. By implementing a digital twin for the considered process, to utilize the digital twin in order to enhance decision making, and simplify the process optimization. 2.1 Context Understanding This section provides an introduction about the company and the parameter on which the research is applied (Medical kits Sterilization process). The historical event data of the sterilization process came from a leading supplier and manager of surgical instruments, which is actively engaged in establishing and overseeing surgical medical kit cleaning units within hospital premises. In this research, we analyzed the sterilization process of one of their centers inside a hospital, where they are responsible for cleaning and sterilizing metal and plastic medical kits, in order to implement a digital 13 twin to be used in optimizing the sterilization process and maximizing the efficiency of their process. The sterilization center, shown in Figure 2.1, contains multiple rooms and areas, each equipped with its own machinery, starting from the Hallway, the physical layout of the sterilization unit shows a scanning station used to register the kits entering the station, following that, the Dirty Area where kits are washed equipped with two scanning station used to register the washing activities done to the kits, six unpacking and washing stations, and four washing machines, then an empty MD Device Storage + Entrance Cleaning Area where kits are put before entering the Clean Area. The physical layout also shows the Clean Area where kits are inspected, sterilized, and collected, equipped with two scanning stations used to register the inspection and sterilization activities done to the kits, four inspection stations, and four sterilization machines. Lastly, Storage where kits are stored until the hospital requests them, the storage area is equipped with a scanning station used to register kits requested by the hospital and leave the sterilization station. Through the Auto-Twin project, we were provided with the metal kits sterilization logs. Therefore, the corresponding sterilization process and activities of the metal kits are described. Figure 2.2 represents the activities in each area and how the process currently works. Starting from the Check-in phase (P1) taking place in the Hallway where contaminated kits are registered enter the sterilization station, then, the kits are loaded into the dirty area. In the dirty area, the kits are unpacked and pre-washed (P2), and then they are loaded into the washing machines (P3) where they are batched in racks. After finishing the washing machine cycle, the kits are moved to the clean area where they are inspected for any malfunctions and repacked (P4). Then, the kits are grouped in racks and sterilized in batches (P8) in the sterilization machines. 14 Figure 2.1: Sterilization Center physical layout After finishing the sterilization machine cycle, the kits are collected (P9). At this point, the sterilization cycle is finished and the kits are stored in the storage area, ready to checkout (P11) and to be reused at the hospital again. According to the company’s information, the sterilization process is handled in sequence based on the arrival time of each medical kit. However, they recognize that this approach may negatively impact the overall system performance. Therefore, they seek to deploy a new way of handling contaminated medical kits using the Auto-Twin project to optimize their sterilization processes. Digital twins are solutions for companies to safely and swiftly come up with a new way to handle their processes and optimize their work without 15 Figure 2.2: Sterilization Center logical layout the need to interrupt any production line or activity [21]. This use case aims to leverage comprehensive historical event logs of a sterilization process collected from a sterilization unit built by the company and used in a hospital. Therefore, the company has shared the event logs detailing the current process of cleaning and sterilizing medical kits. The event logs are utilized to implement a digital twin in order to optimize the way of handling the contaminated. Additionally, the company has expressed a strong interest in improving the efficiency of the sterilization process while reducing the workload on their employees. 2.2 Data Understanding To achieve the objectives outlined above, it is crucial to develop a thorough understanding of the data provided by the company. The data consists of detailed event logs that capture the entire sequence of the sterilization process. By examining these logs, we can identify patterns, bottlenecks, and inefficiencies in the current process. Figure 2.3 illustrates the factory layout and the various phases of the sterilization process indicating the exact moment when the kit gets scanned, 16 Figure 2.3: Factory layout Process activity Name Number of rows P1 Entrada de Material Sucio 13537 P2 Asignacion de Rack Lavadora 19692 P3 Carga de Lavadora 24738 P3 Carga de Lavadora Liberada 23617 P4 Inicio de Montaje 24171 P4 Produccion Montada 37931 P8 Carga de Esterilizador 37813 P9 Carga de Esterilizador Liberada 37199 P11 Comisionado 22958 Table 2.1: Activity representation which serves as the basis for event data collection. The event data corresponding to each of these six sterilization phases were transferred into CSV tables. Subsequently, these tables were partitioned into nine distinct tables, each representing the activity within a specific phase, as illustrated in Figure 2.2. Additionally, Figure 2.1 outlines the relationship between the activity names, the number of rows in each table, and the corresponding phase process steps, as represented in the logical layout illustrated in Figure 2.2. The data contains a total of 241656 records with different numbers of records per activity as shown in Table 2.1, and 15 columns represent record 17 Name Describtion Data type Fecha de seguimiento Date of the activity. DD/MM/YYYY format Date Hora de seguimiento Hours and minute of the activity. HH:mm format Time Usuario Employee initials String Nombre punto de control Name of the activity String Es preliminar Is there a priority for sterilizing this kit Boolean Es f´ısico Is it physical Boolean Es denegado Is it denied Boolean Cant. Quantity Integer Tipo de objeto Type of kit String C´odigo Kit unique code String N/S Kit sequence number String Nombre / Descripci´on Name and description of the kit String Tipo de producci´on Production type String Nombre producci´on Production name String Informaci´on adicional Additional information String Table 2.2: Data Description table specifications as shown in the following Table 2.2 and their description. The table shows an additional information column, this column is not restricted to a single format, and to solve this issue, this column was standardized by splitting its data into multiple new columns, allowing for the extraction and utilization of the embedded information. This process enhanced the clarity and usability of the data, enabling more precise analysis. Additionally, the data doesn’t include a unique identifier for each kit, by concatenating the codigo and N/S fields the unique identifiers were generated. This method ensures that each kit is distinctly recognizable within the dataset. Detailed descriptions of these features, along with their implications and usage, are thoroughly discussed in Chapter 5.3. 18 Chapter 3 Problem Exposition The main goal of any kind of business is to create profits for its owners or stakeholders, one aspect of decreasing the expenses and increasing the profit is optimizing the business operations processes and decreasing the processing time while keeping the quality and the expenses the same. Nowadays, many companies try to optimize their operational processes not only by investing in new equipment and machines but also by fully leveraging their existing resources and testing the potential of what they have, one way to achieve that is by building an artificial representation for the whole operation [12]. One of the most effective principles to realize this goal is to build a digital twin that aims to mimic the process activities, fed by the historical data of the process, and gives back information to help the business in decisionmaking to optimize the operations, reducing the time losses and enhancing the processes life-cycle. A digital twin is a virtual representation of a tangible object, process, or system, that serves as a counterpart for experimental purposes such as simulation, analysis, testing, and control, which aim to improve efficiency by optimizing processes by decreasing processing time, optimizing resources, and possibly decreasing expenses. By implementing digital twin, we are aiming to give control to the business to modify and test over its virtual sterilization process representation which allows it to implement and test several optimization techniques without interrupting the real process. 19 3.1 Preliminaries In this chapter, several concepts related to multidimensional process mining are defined, understanding these concepts is essential to keep following the research and its results, since in this research we are leveraging multidimensional process mining and its related definitions. Definition 1 (Events):Anevente∈εdescribes that a specific discrete observation has been made (by a sensor, a system, a human observer, etc.). The observation itself is described by attribute-value pairs through the partial function π:ε×AN  V al. For each event e∈εand each attribute name a∈AN,π(e, a)=vdefines the value vof attribute a. We write π(e, a)=⊥if attribute ais undefined for e(has no value). We also write πa(e)=vor e.a =vfor π(e, a)=v. For each event e∈ε, we require that: •the attribute time is defined, i.e., πtime(e)=⊥. •ecarries a value πa(e)=⊥for some other attribute a∈AN,a=time [9]. Definition 2 (Event Table) An event table T=(E,Attr,#) is a set Eof events, a set Attr of attribute names with act,time ∈Attr. The partial function # : E×Attr Val assigns to an event e∈Eand an attribute name a∈Attr a value #a(e)=v;# a(e)=⊥if ais undefined for e. Each event e∈Erecords an activity and a timestamp, i.e., #act(e)=⊥ and #time(e)∈Valtime [10]. Definition 3 (Event Table with Entities): An event table with entity types T=(E,Attr,#,ENT) additionally designates one or more attributes ∅ =ENT ⊆Attr as names of entity types [10]. Definition 4 (Entities) Let T=(E,Attr,#,ENT) be an event table with entities. Let ent ∈ENT be an entity type. The set of entities in Tof type ent is Entities(ent, T )={n|∃e∈E:n∈e.ent}. An event e∈Ewhich has a value n=e.ent or n∈e.ent is correlated to entity n[10]. 20 Definition 5 (Correlation):LetT=(E,Attr,#,ENT)beaneventtable with entities. Let n∈Entities(ent, T)beanentityoftypeent ∈ENT. Event eis correlated to entity n, written (e, n)∈corrent,T , if and only if n=e.ent or n∈e.ent. We write corr(n, ent, T)={e∈E|(e, n)∈corrent,T } for the set of events correlated to entity n∈Entities(ent, T ) [10]. Definition 6 (Relation):Let T=(E,Attr,#,ENT)beaneventtable with entities. Let ent1,ent 2∈ENT be two entity types. The relation between ent1and ent2in Tis R(ent1,ent2) = {(e.ent1,ent 2)|e.ent1= ⊥,e.ent 2=⊥} [10]. Definition 7 (Directly-Follows Relationship (per Entity)):Let T=(E,Attr,#,ENT) be an event table with entities. Let n∈Entities(ent, T ) be an entity of type ent ∈ENT. Let e1,e 2∈Ebe two events; e2directly follows e1from the perspective of n, written e1n,T e2[10], if and only if: 1. (e1,n),(e2,n)∈corrent,T (both are correlated to n), 2. e1.time <e 2.time (e1occurred before e2), 3. and there is no other event (e,n)∈corrent,T with e1.time <e .time < e2.time. Definition 8 (Labeled Property Graph): A labeled property graph (LPG) G=(N,R,λ,#) is a graph with nodes N, and relationships Rwith the following properties: 1. Each node n∈Ncarries a label λ(n)∈ΛN. 2. Each relationship r∈Rcarries a label λ(r)∈ΛRand defines a directed edge → r=(nsource,n target)∈N×Nbetween two nodes. 3. Any node nand relationship rcan carry properties as attribute-value pairs via function # : (N∪R)×Attr Val [10]. Definition 9 (Event Knowledge Graph): An event knowledge graph (or just graph) is an LPG G=(N,R,λ,#) with node labels {Event, Entity}⊆ ΛNand relationship labels {df , c o r r }⊆ΛRindicating ”directly-follows” and ”correlation” with the following properties: 1. Every event node e∈NEvent records an activity name e.act =⊥and a timestamp e.time =⊥. 21 2. Every entity node n∈NEntity has an entity type n.type =⊥. 3. Every correlation relationship r∈Rcorr,→ r=(e, n), is defined from an event node to an entity node, e∈NEvent,n∈NEntity; we write n∈corr(e) and e∈corr(n) as shorthand. 4. Any directly-follows relationship df ∈Rdf , → df =(e1,e 2) is defined between event nodes e1,e 2∈NEvent and refers to a specific entity df . ent = n∈ NEntity such that (a) e1and e2are correlated to entity n:(e1,n),(e2,n)∈Rcorr; (b) e1occurs before e2:e1.time <e 2.time; and (c) there is no other event e∈NEvent correlated to n,(e,n)∈Rcorr that occurs in between e1.time <e .time <e 2.time [10]. Definition 10 (df-path):LetG=(N,R,λ,#) be a graph. Apathr=r1,...,r k∈(Rdf )∗of df-relationships is a directly-follows path (df-path) if and only if all relationships are defined for the same entity, i.e., for all 1 ≤i<k,ri.ent = ri+1.ent = n;wealsosayris a df-path for entity n.ris maximal if and only if there is no other df-relationship r∈Rdf so that r, r1,...,r kor r1,...,r k,ris also a df-path [10]. Definition 11 (Kolmogorov-Smirnov (K-S) Test): a non-parametric statistical test used to determine whether a sample follows a specific distribution or to compare the distributions of two independent samples. The test evaluates the maximum difference between the empirical cumulative distribution function (ECDF) of the sample and the cumulative distribution function (CDF) of the reference distribution or between the ECDFs of two samples. •K-S Test Statistic one-Sample Dn=sup x|Fn(x)−F(x)| •p-Value one-Sample p-value = P(Dn≥observed Dn|H0) 22 Primary research on the other hand is divided into three parts; the first part is building an event knowledge graph in order to represent the event data and provide behavioral information about each entity contributing to the sterilization process. the second part is investigating the data that was provided by the company in order to obtain parameters to feed the simulator of the sterilization process. The third part is conducting meetings with the representatives from the departments that are involved in the sterilization process in order to build a better understanding of the process and the hidden entities that play a role in the sterilization process but are not completely visible in the event data, which reflects on building the digital twin. The idea of choosing meetings as a research strategy is to be able to have a clearer image of the process and to understand the process in detail by collecting more reliable data about them, it is also important to fill any gaps of missing information by asking the departments’ representatives about any unclear data or activities within the process. 4.2 Methodology The development of a digital twin for the sterilization process involved a structured approach aimed at creating a high-fidelity representation of the actual processes within the company. This methodology was designed to ensure that the digital twin could not only replicate the current operations but also serve as a tool for identifying inefficiencies and testing potential improvements. Figure 4.1 illustrates the process of extracting the data to implement a digital twin. Our approach of implementing a digital twin was divided into three main phases: Digital Twin Modeling, Digital Twin Simulation, and Evaluation. Each phase comprised several critical steps, which are detailed below. 4.2.1 Phase One: Digital Twin Modeling Building a digital twin from multiple sources and sensors comes with a great challenge and several steps and decisions need to be taken to make it real. 29 Figure 4.1: Process of getting the data 1. Raw Data Preprocessing •Data Cleaning Real-world data often contains errors, inconsistencies, and noise, which can compromise the accuracy of the digital twin. The first step involved cleaning the event data provided by the company to remove duplicates and noise. This ensured that the data used in building the digital twin was of high quality, thereby enhancing the reliability of the model. •Data Unification Data from multiple sources often varies in format and structure, hindering accurate modeling and analysis. To address this, we standardized the data into a consistent format. This unification was essential for integrating data from various processing steps into the digital twin, enabling accurate and effective modeling. •Entity Unique Identifier Assignment Without unique identifiers, it is difficult to distinctly track and analyze different entities within the system. We implemented unique identifiers for each entity, such as medical kits, to maintain clarity and accuracy in the digital twin. This step was crucial for precise tracking and monitoring. 30 2. Data Representation Inappropriate data representation can interfere with the efficiency of data processing and make the digital twin less understandable for stakeholders. We selected Event Knowledge Graphs as the data representation method this selection facilitates the ability of event knowledge graphs to visually represent the data. Furthermore, event knowledge graphs provide a robust framework for multidimensional data representation, enhancing the study and enrichment of data to optimize the process model and resource management. 3. Process Case Identification Understanding recurring patterns and flows within the process is critical for accurate modeling and simulation. We identified process cases to recognize these patterns and cycles. This understanding was a key to developing an accurate digital twin model that could be used for optimization. 4. Process Model Discovery The digital twin model must accurately reflect the actual processes and workflows to be effective. Using process mining techniques, we discovered the current process model. This step involved validating the digital twin against the real-world as-is process to ensure accuracy and effectiveness. 5. Contextual Enrichment Insufficient context in data can limit the value and insights derived from it. We enriched the data by adding contextual information, such as details about employees, machines, and programs. This enrichment enhanced the data’s utility for analysis and decision-making, providing a more comprehensive understanding of the process. 4.2.2 Phase Two: Digital Twin Simulation •Simulation Softwares Analysis Studying the available simulation softwares is essential for conducting an in-depth analysis, as it can lead the way for the analysis. •Simulation Parameterization In-depth analyses to configure the chosen simulation software in order to mimic the intended business process, these analyses include the following: 31 1. Arrival Rate Analysis The arrival rate is a crucial factor in simulating a business process accurately as it can affect the process directly by trafficking the activities. 2. Current Process Handling Understanding the current dispatching rules is crucial for an accurate simulation that mirrors the real-world process. We analyzed the existing dispatching rules, such as the First In, First Out (FIFO) approach, to understand their impact on system performance. This analysis helped identify areas for potential improvement. 3. Activity Analysis Studying the activities and the required time to perform an activity is an essential step in mimicking a business process. Without these analyses, the simulation can’t be tuned to reflect the original process correctly. 4. Working and Waiting Time Analysis Identifying bottlenecks and areas for improvement requires an understanding of the working and waiting times within the process. We calculated these times by analyzing the sequence of activities, allowing us to pinpoint inefficiencies and areas for optimization. 5. Work Behavior and Schedule Discovery Accurate simulation of operations requires an understanding of resource availability. We examined the working schedules of employees to ensure that the simulation reflected actual operational constraints and allowed for effective capacity planning. 6. Machinery Analysis Understanding the behavior of machines is essential for accurate simulation and process optimization. We analyzed the behavior of these entities to gain insights into their impact on the process, which was then incorporated into the simulation model. •Simulation Implementation To accurately simulate the sterilization process, the digital twin needed to be fully parameterized. We transformed the discovered process model into a discrete event simulation (DES) model. This involved setting specific parameters that reflect the real-world process, enabling the simulation of both current (as-is) and potential (to-be) process behaviors. 32 4.2.3 Phase Three: Evaluation Ensuring that the simulation model accurately reflects the real process requires a robust verification mechanism. We implemented a verification process that compared the number of processed kits in the simulation to the actual numbers from the real process. This step was crucial for validating the accuracy of the digital twin and ensuring its reliability as a tool for process optimization. 33 Chapter 5 Digital Twin Modeling This chapter includes an overview of the data as well as choosing kits identifiers that are used to build cases and discover the process model. It also includes the analysis of the medical kits sterilization process. 5.1 Raw Data Preprocessing As we mentioned in 2.1, the company provided us with sterilization process event logs and to make this data useful for building a digital twin, several steps are taken to ensure the usability of the data. In this section, we are discussing data cleaning of Phase 1. Section 4.1 established the necessity of data cleaning as a fundamental step in data preprocessing to enhance the accuracy, reliability, and efficiency of data analysis and modeling. Van der Aalst (2016) [22] emphasizes the negative impact of duplicate records on the efficiency and accuracy of analyses, as they lead to inaccuracies in process models and analyses. Therefore, removing duplicates is crucial for maintaining data quality and ensuring reliable results. Mans et al. (2013) [15] stressed the negativities of having inconsistent data formats, as it can create significant barriers to effective process mining, highlighting that standardizing data formats is a necessary step to enable accurate and efficient analysis. Upon analyzing the event dataset provided for the use case, several data quality issues were identified: •Duplicate Records 34 Activity name Number of duplicates Entrada de Material Sucio 298 Asignacion de Rack Lavadora 3727 Carga de Lavadora 9082 Carga de Lavadora Liberada 8816 Inicio de Montaje 673 Carga de Esterilizador 5 Table 5.1: Activity duplication representation The dataset contained a significant number of duplicated rows, which could affect the accuracy and reliability of subsequent analyses. The initial dataset comprised 241,656 records with different numbers of records per activity as shown in Table 2.1. Each table is formed from 15 columns representing the record specifications, Table 2.2 represents columns’ name, description, and expected data type. Upon thorough data analysis, it was determined that there exist 22,637 duplicated records, constituting roughly 10% of the dataset, as illustrated in Table 5.1. By utilizing Python’s Pandas library, duplicate records were identified and removed from each event log table. This resulted in a refined dataset of 219,019 unique records, ensuring the dataset’s integrity and reliability. Accounting for the fact that the data was collected across multiple activities, we aimed to maximize the data’s utility by maintaining consistency, quality, and integrity. To achieve this, we standardized the data into a single format by splitting the ”Additional Information” column into multiple columns based on the respective activities. •Inconsistent Data Formats The ”Additional Information” column contained varied data formats, each specific to a different activity, making the data inconsistent and challenging to analyze. 35 To address this problem, the column was split into multiple new columns based on the respective activities, since the ”Additional Information” column incorporates information about washing racks, machines and programs, and sterilization machines and programs. Splitting the ”Additional Information” column provides information about the most frequently loaded machines and the most commonly used sterilization programs. •Kit identification Upon cleaning and unifying the data, it was essential to accurately identify each medical kit and uniquely relate it to the corresponding events. This step was crucial to ensure the distinction between kits and to prevent any overlap between different kits. The lack of a unique identifier for each kit posed a significant challenge, as it could lead to confusion and inaccuracies in tracking the kits’ behaviors throughout the sterilization process. To address this problem, a new attribute named KitID was introduced. This KitID attribute was constructed by concatenating the Codigo and N/S numbers. The KitID attribute served as a unique identifier for each kit, enabling precise tracking of each kit throughout the sterilization process. The introduction of the KitID attribute opened the door to more significant discoveries. Throughout several meetings with the company, the company revealed that the sterilization unit sterilizes two types of instruments, medical kits and containers. Also, the company provided insights that allowed us to differentiate between them based on the KitID prefix. Additionally, the company explained that these containers are used to hold the kits during the washing process, and they have a different sterilization cycle compared to the medical kits. This differentiation was crucial for accurately modeling the process and ensuring that each entity was correctly represented within the digital twin. 36 This comprehensive approach not only solved the immediate problem of identifying and tracking individual kits but also improved our overall understanding of the data, leading to more accurate and effective process modeling. With the data preprocessed, the next step involved loading the data into a database. 5.2 Event Knowledge Graph Auto-Twin chose Event Knowledge Graphs for data representation for their ability to provide a multidimensional view of the process as mentioned in Section 3.2, making them easier to interpret for multidimensional process mining. To ensure compatibility with the Auto-Twin framework, the cleaned and unified data was loaded into a Neo4J server. This enabled the analysis and inference of relationships between kits and associated activities, establishing clear connections between each kit and its corresponding sterilization activity. This structured approach to data cleaning and preparation ensures the dataset’s reliability and usability, setting a solid foundation for accurate modeling and insightful analysis in the digital twin. The structure of the event knowledge graph is described by an event knowledge graph schema, provided in Figure 5.1. This schema illustrates the steps describing the sterilization processes, along with the entities involved in each step. The main entities comprising the provided schema along with a description of the considered relationships with other entities are as follows: 1. Event A step from the sterilization process describes a specific part of the process and can either be the beginning (NewKitEvent), intermediate, or final (LastKitEvent) sterilization step. It is followed by other event nodes with one or more of the following relationships: •Directly Follows (DF )relationship describes the next sterilization step based on timestamps where both event nodes are related to the same Kit. 37 •Directly Follows Case (DF CASE )relationship describes the next sterilization step based on timestamps where both event nodes related to the same Kit and the same sterilization case (Run). •Directly Follows Employee (DF EMP)relationship describes the next sterilization step an employee worked on, based on timestamps where both event nodes related to the same Employee. •Directly Follows Cycle (DF CYCLE )relationship connects the last sterilization step in a sterilization case with the first sterilization step of the following case, based on timestamps where both event nodes related to the same Kit and different sterilization cases (Run). Each event node is an observation of a class node connected to it via OBSERVED relationship and it correlates to a sterilization case (Run)via CORR relationship, showing the steps taken in each sterilization case. Depending on whether the sterilization step is related to a machine, it can be STERILIZED IN a sterilization machine if it’s related to a sterilization machine or WASHED IN a washing machine if it’s related to a washing machine and WASHED ON arack. A group of events can be related to a Batch node if and only if all the events were executed by the same employee in the same sterilization step in a time period between a sterilization step and the following one doesn’t exceed three minutes, the choice of three minutes comes after deeply exploring the data, especially the steps related to WashingMachine and SterilizationMachine, and finding out that three minutes creates the most number of batches related to the same machine before the employee start loading of freeing a different machine since we are considering working with a different machine a different batch. 2. NewKitEvent A step from the sterilization process describes the starting activity of the sterilization process. 3. LastKitEvent A step from the sterilization process describes the finishing activity of the sterilization process. 4. OnlyKitEvent A special type of event node describes a sterilization activity where only a single activity is recorded from the sterilization cycle. 38 5.4 Process Model Following the identification of the sterilization cycles, it was necessary to ensure that the digital twin model accurately represents the company’s actual sterilization process flow. Without an accurate process model, the digital twin would not reliably simulate the real-world process, leading to potential discrepancies and inefficiencies in the analysis and decision-making. To address this problem, we utilized the ProM extension for Python to discover the process model. By employing the sterilization cycles identified the in previous step, we identified the common starting activities for the sterilization process, as shown in Figure 5.8, and the common ending activities, as shown in Figure 5.9. This step was crucial for delineating the boundaries of the process and ensuring a clear understanding of its structure. Figure 5.8: Common starting activities for the sterilization process Next, we applied the Heuristics Miner algorithm to construct a directlyfollows graph. This graph represents the actual sterilization process, illustrating the sequence of activities and their relationships. By doing so, we ensured that the digital twin accurately mirrors the real-world process flow. Figure 5.10 illustrates the directly-follows graph for the sterilization pro45 Figure 5.9: Common ending activities for the sterilization process cess, providing a visual representation of the process model. This model serves as the foundation for the digital twin, enabling precise simulation and analysis of the sterilization process. Further details on the implementation and validation of this process model will be discussed in the following sections. 5.5 Contextual Enrichment After constructing the directly-follows graph, it was clear that further enrichment of the event knowledge graph was necessary for a more comprehensive representation of the company’s sterilization process. The initial graph provided a basic structure but lacked detailed contextual information about various process elements, which is essential for accurate modeling and analysis. To enhance the event knowledge graph, we added detailed information about employees, washing machines and their washing programs, and sterilization machines and their sterilization programs. This enrichment step aimed to provide a more granular view of the process, allowing for better understanding and analysis. 46 Figure 5.10: Sterilization process directly-follows graph 47 To enhance the event knowledge graph, we enriched it by adding nodes representing detailed information about employees, washing racks, washing machines, and sterilization machines. These new nodes were systematically connected to their relevant existing nodes, thereby creating a more contextualized and comprehensive representation of the process. This enrichment was intended to provide a more granular view, facilitating deeper insights and a more accurate analysis of the process. Then, by utilizing process mining techniques, 20 employees were identified contributed to the sterilization process, 17 of which participated regularly and 3 worked on special occasions. Additionally, five washing machines were identified and analyzed to determine their specific washing programs. The washing machines were categorized into two categories: Lavadora and Jubiter, with each category having its own washing programs. Furthermore, six sterilization machines were identified and studied to determine their respective sterilization programs. These machines were categorized into three categories: Autoclave, VPro, and Eagle, with each category having its own sterilization programs. This additional information was integrated into the event knowledge graph, ensuring that each entity and its attributes were accurately represented. By incorporating these details, the event knowledge graph now offers a richer, multi-dimensional view of the sterilization process, facilitating more precise analysis and enabling more informed decision-making. After enriching the event knowledge graph, we began modeling the digital twin by creating a business process model using the Business Process Model and Notation (BPMN) tool, Camunda [3], which in this research serves as a software platform used for implementing BPMN models and workflows. By utilizing Camunda software, a model closely mirrors the actual sterilization process visualized using Neo4J was structured, as illustrated in Figure 5.11. This BPMN model includes representations of users, washing machines, and sterilization machines, providing a comprehensive view of the current process. The creation of a BPMN model serves as the foundation for our digital twin. By accurately describing the current process, we can simulate the realworld process behavior and identify potential areas for optimization. 48 Figure 5.11: BPMN representation of the real process 49 In the following Chapter 6, a comprehensive analysis to extract simulation parameters and parameterize the digital twin model is provided. 50 Chapter 6 Digital Twin Simulation With the digital twin model now accurately representing the actual sterilization process model and the various dimensions affecting it, the next critical step is to extract the parameters to tune the simulator of the process. To simulate the sterilization process accurately, an in-depth analysis of the entire operation and the available simulation softwares is essential. This involves: •Simulation Softwares Analysis By studying the features provided by many BPMN simulation softwares, in order to select the best simulation software to mimic the current process. •Simulation Parameters Analysis By studying the process, this section aims to provide the needed parameters in order to mimic the real-life process. This study includes several analyses, as follows: – Arrival Rate Analysis Analyzing the number of kits arriving and the arrival rate to the sterilization station. – Dispatching Rule Analysis By examining the current way of handling the contaminated kits followed by the company to understand task prioritization and assignment. – Activity Analysis Performing a detailed analysis of each activity within the sterilization unit to estimate the time required for each activity to be completed at a single kit level and batch level. 51 – Working and Waiting Times Analysis By analyzing the working and waiting times for the kits inside the sterilization unit, aiming to pinpoint bottlenecks and potential areas for improvement within the sterilization process. – Workforce Behavior and Scheduling Analysis Analyzing the roles and contributions of each employee involved in the process to estimate the working schedules and roles of each employee. – Machinery Analysis Analyzing the roles of each machine involved in the process to estimate the time required to complete both the washing and sterilization programs. •Simulation Implementation By configuring the chosen simulation software with the needed parameters in order to mimic the real-life process. By conducting this comprehensive analysis, we can ensure that our simulation accurately reflects the real-world process. This, in turn, allows us to explore various optimization techniques to enhance the efficiency and effectiveness of the sterilization process. In this chapter, a detailed description of each analysis is provided. 6.1 Simulation Softwares In this section, an overview of the process of choosing simulation software is provided, and a detailed view of the chosen software is provided. For this research, we implemented the digital twin simulation over multiple simulation softwares in order to test the best fit for our use case. Apache Airflow [1], an open-source platform designed for orchestrating and automating complex data workflows. Airflow allows users to define, schedule, and monitor workflows as directed acyclic graphs (DAGs) using Python code. First, we simulated the sterilization process using Apache Airflow, as illustrated in Figure 6.1. The simulation was capable of handling a certain task for a certain period of time and assigning it to a specific resource, coded via Python. However, Apache Airflow lacks the ability to schedule an arrival 52 Figure 6.1: Sterilization process implementation via Apache Airflow rate for cases and perform resource locking. Then, Ciw [4] a discrete event simulation library for open queueing networks was utilized. At first glance, it seemed like a good fit for our use case simulation as its core features include the capability to simulate networks of queues, multiple customer classes, and scheduling. However, it lacks the ability to run tasks in batches, which is essential in our use case. Lastly, Scylla [5] is an open-source simulation software designed for process modeling and analysis, particularly in the context of business process management and operations research. Scylla allows users to create detailed models of processes, simulate various scenarios, and analyze outcomes such as throughput, resource utilization, and bottlenecks. Scylla provides functionalities needed in order to simulate our use case sterilization process, such as specifying the number of resources, the distribution of arrival rate and tasks, and batching. Subsequently, Scylla provides an easy-to-use user interface and helps in fully utilizing its functionalities, as illustrated in Figure 6.2. Scylla was selected for its ability to accurately simulate processes while taking into account resource availability, batching, and other complex operational factors. Its advanced simulation capabilities make it an ideal choice 53 Figure 6.2: Scylla user interface for this study, enabling us to replicate the real-world conditions of the sterilization process with high fidelity. After choosing a proper simulation software to reflect the use case sterilization process, a deeper analysis is done in order to parameterize and tune the simulation. 6.2 Simulation Parameterization After analyzing multiple simulation softwares and selecting one to represent the sterilization process, the next step involves conducting a more focused analysis to accurately simulate the sterilization process and optimize the simulation outcomes. 6.2.1 Arrival Rate Analysis The arrival rate of contaminated medical kits is a crucial factor in simulating the sterilization process accurately. In our model, the arrival rate represents the frequency at which contaminated kits enter the sterilization station. This rate is estimated based on the historical event data provided 54 Figure 6.7: Entrada Material Sucio histogram of the time required to perform the activity (a) Cargado en Carro L+D distribution throughout weekdays (b) Cargado en Carro L+D distribution throughout weekends Figure 6.8: Cargado en Carro L+D distribution throughout the day 61 Figure 6.9: Cargado en Carro L+D histogram of the time required to perform the activity After checking the distribution histogram in Figure 6.13 and comparing it with the distribution histograms supported by Scylla software A, it is shown that the distribution histogram resembles a triangular distribution begins at a lower value of 1.0, peaks at 4.0, and then extends to an upper value of 20.0. These values are subsequently utilized to tune the process simulation. 3. Carga L+D Iniciada Figure 6.10(a), represents the distribution histogram for the Carga L+D Iniciada activity throughout the day. Where Figure 6.10(b) represents the distribution throughout the day on normal weekdays, and Figure 6.10 represents the distribution throughout the day on weekends/holidays. Figure 6.11 illustrates the histogram of the time required to perform the Carga L+D Iniciada activity. This histogram provides a detailed visualization of the distribution of time spent on the task, allowing us to observe the frequency of various time intervals. After checking the distribution histogram in Figure 6.11 and comparing 62 (a) Carga L+D Iniciada distribution throughout weekdays (b) Carga L+D Iniciada distribution throughout weekends Figure 6.10: Carga L+D Iniciada distribution throughout the day Figure 6.11: Carga L+D Iniciada histogram of the time required to perform the activity 63 it with the distribution histograms supported by Scylla software A, it is shown that the distribution histogram resembles a normal distribution beginning with mean equals 5.0 and standard deviation equals 3.0. These values are subsequently utilized to tune the process simulation. 4. Washing programs According to the data provided by the company, we identified two different types of washing machines, named Lavadora and Jupiter, each has its own washing programs. •Lavadora We identified four washing programs related to this washing machine type, and the meetings with the company representative specified that each of these programs requires exactly 60 minutes to be finished. – Instrumental Normal – Delicado – Priones – Instrumetal Nuevo •Jupiter As per this washing machine type, we identified two washing programs related to it. Although the company representative did not specify the exact duration required for each sterilization cycle, we estimated that it would require 40 minutes to finish a washing cycle using this washing machine. – Contenedores – Carros 5. Carga L+D liberada Figure 6.12(a), represents the distribution histogram for the Carga L+D Liberada activity throughout the day. Where Figure 6.12(b) represents the distribution throughout the day on normal weekdays, and Figure 6.12 represents the distribution throughout the day on weekends/holidays. Figure 6.13 illustrates the histogram of the time required to perform the Carga L+D Liberad activity. This histogram provides a detailed visualization of the distribution of time spent on the task, allowing us to observe the frequency of various time intervals. 64 (a) Carga L+D Liberad distribution throughout weekdays (b) Carga L+D Liberad distribution throughout weekends Figure 6.12: Carga L+D Liberad distribution throughout the day Figure 6.13: Carga L+D Liberad histogram of the time required to perform the activity 65 After checking the distribution histogram in Figure 6.13 and comparing it with the distribution histograms supported by Scylla software A, it is shown that the distribution histogram resembles a triangular distribution begins at a lower value of 1.0, peaks at 3.0, and then extends to an upper value of 30.0. These values are subsequently utilized to tune the process simulation. 6. Assemply To ensure the simulation accurately reflects the real-world performance, and after deep analysis of the process, we found out that Montaje and Producci´on Montada referred to the same activity, which is assembly, where Montaje represents starting the activity and Producci´on Montada represents finishing it, for that reason, we unite the activities into one activity maned Assembly. The Assembly activity was tuned using a normal distribution. For this activity, the time required to complete it is modeled with a mean of 5 minutes and a standard deviation of 3 minutes. This approach captures the variability in task duration, accounting for both faster and slower completions within the normal operational range. The type of distribution, the mean, and the standard deviation values come after an in-depth analysis of the historical event data, Figure 6.14(a) and Figure 6.14(c) represent the distribution histogram for Montaje and Producci´on Montada activity throughout weekdays. Whereas Figure 6.14(b) and Figure 6.14(d) represent the histogram of the time required to perform the same activities throughout weekends/holidays. Additionally, the simulation was configured to require exactly one employee to perform this activity and in batches, aligning with the staffing constraints observed in the actual process. By incorporating these parameters, the simulation more precisely mirrors the real-world conditions, allowing for better predictions and optimizations of the sterilization process. 7. Composici´on de cargas Figure 6.15(a), represents the distribution histogram for Composici´on de Cargas activity throughout the day. Where Figure 6.15(b) represents the distribution throughout the day on normal weekdays, and Figure 6.15 represents the distribution throughout the day on weekends/holidays. 66 (a) Montaje distribution throughout weekdays (b) Montaje distribution throughout weekends (c) Producci´on Montada distribution throughout weekdays (d) Producci´on Montada distribution throughout weekends Figure 6.14: Assembly distribution throughout the day (a) Composici´on de Cargas distribution throughout weekdays (b) Composici´on de Cargas distribution throughout weekends Figure 6.15: Composici´on de Cargas distribution throughout the day 67 Figure 6.13 illustrates the histogram of the time required to perform Composici´on de Cargas activity. This histogram provides a detailed visualization of the distribution of time spent on the task, allowing us to observe the frequency of various time intervals. Figure 6.16: Composici´on de Cargas histogram of the time required to perform the activity After checking the distribution histogram in Figure 6.16 and comparing it with the distribution histograms supported by Scylla software A, it is shown that the distribution histogram resembles a normal distribution begins with mean equals to 5.0 minutes and standard deviation equals to 3.0 minutes. These values are subsequently utilized to tune the process simulation. 8. Sterilization programs According to the data provided by the company, we identified three different types of sterilization machines, named Lavadora and Jupiter, each has its own sterilization programs. •Autoclave 68 This sterilization machine functions by applying high temperatures to the contaminated medical kits causing the kits to be sterilized, in this sterilization machine, we identified four sterilization programs, and the meetings with the company representative specified that each of these programs requires between 60 and 75 minutes to be finished, without specifying the exact programs. –INTRUM –GOMAS – CONTENE –PRIONES •VPro This sterilization machine functions by applying UV light and temperature to the contaminated medical kits causing the kits to be sterilized, in this sterilization machine, we identified two sterilization programs. –Lumen –NoLumen Although the company representative did not specify the exact duration required for each sterilization cycle, we estimated that the Lumen program requires 65 minutes to finish a sterilization cycle, and NoLumen requires 40 minutes to finish a cycle using this sterilization machine. •Eagle the Eagle sterilization machine has one sterilization program identified related to it, Temperatura Alta, we estimated that it requires 50 minutes to finish a sterilization cycle. 9. Carga de esterilizador liberada For this activity, the distribution histogram was estimated according to the distribution histogram of the Composici´on de Cargas activity shown in Figure 6.16 with mean equals to 4.0 and standard deviation equals to 2.0 minutes since it’s the last step in the sterilization process and the data provided by the company doesn’t contain kits leaving the process. By integrating these findings into our simulation model, we can create a more accurate digital twin of the sterilization process. This digital twin not only mirrors the current process but also serves as a valuable tool for testing various optimization strategies and predicting the impact of changes on the overall efficiency and effectiveness of the sterilization process. 69 6.2.4 Working and Waiting Times Analysis One of the critical aspects of this analysis is to identify and measure the total working time and waiting time for each medical kit. This step aims to pinpoint bottlenecks and potential areas for improvement within the sterilization process. To begin, we assume that employees scan the medical kits at the completion of specific activities and before initiating subsequent ones. These activities include: •Initiation Task Scans: –”Cargado en carro L+D” –”Carga L+D iniciada” –”Montaje” –”Composici´on de cargas” –”Comisionado” •Scan After Task Completion: –”Entrada Material Sucio” –”Carga L+D liberada” –”Producci´on montada” –”Carga de esterilizador liberada” By calculating the time difference between the scans for specific pairs of activities, we identified the waiting and working times for each kit. Waiting Time was calculated by summing the time between the following pairs of activities, representing the duration a kit waits for the next action to be taken: •”Entrada Material Sucio” to ”Cargado en carro L+D” •”Cargado en carro L+D” to ”Carga L+D iniciada” •”Carga L+D liberada” to ”Montaje” •”Producci´on montada” to ”Composici´on de cargas” •”Carga de esterilizador liberada” to ”Comisionado” 70 6.4 Conclusion This chapter provided simulation software analysis highlighted the features needed to mimic the sterilization process. Subsequently, in-depth parameter analysis is conducted to identify the necessary parameters to configure the selected simulation software. Lastly, implementation and execution of the simulation was conducted, indicated the limitations of the selected software and a solution to overcome these limitations was provided. 77 Chapter 7 Evaluation In this chapter, a comprehensive evaluation of the digital twin and the original sterilization process is presented, and a comparison between the original Scylla simulation software and our modified version for handling repetitive activities. 7.1 Objective Chepela-Campa et at. [7] provided a methodology to evaluate business process simulations through a collection of measures to assess the quality of the business process model, this evaluation intended not only to capture how close a business process simulation model is to the actual process behavior, but it also helps to identify sources of discrepancies. The evaluation objectives for this research are centered on assessing the accuracy and reliability of the digital twin developed to simulate the sterilization process side by side with the evaluation of the modified Scylla version to handle repetitive tasks. The key objectives include: •Congestion Comparison To assess the accuracy of the simulation by comparing the number of processed kits in the simulator against the actual process. This comparison involves calculating the time required to complete a full sterilization cycle. •Goodness-of-Fit Evaluation To assess the accuracy of the simulation model by conducting a goodnessof-fit analysis, comparing the distribution of simulated outcomes with actual process data. This evaluation will determine how well the model 78 replicates the real-world process, ensuring its validity for predictive and analytical purposes. •Temporal Comparison To evaluate the model’s accuracy by comparing key performance metrics, such as the average processing time and the time required to perform an activity against real-world data. This comparison determines if the model can reliably predict the outcomes of the sterilization process. By achieving these objectives, the evaluation will provide a comprehensive assessment of the digital twin’s accuracy and its potential as a reliable tool for simulating and optimizing the sterilization process. 7.2 Setup In this research, we utilized two versions of the Scylla simulation software: the latest original version and a locally modified version. These were employed to simulate the BPMN model, with configurations extracted as detailedinSection6. The experiment ran under the assumption that the hospital operates 24/7, while the sterilization unit runs from 7:00 AM to 10:00 PM. Based on this assumption, kits were set to arrive even before the sterilization unit began its daily operations. During the utilization of Scylla simulation software, the following files have to be loaded: 1. General Configuration File This involves specifying the details of the simulation environment as follows: •Total Number of Employees For the total number of employees involved in the sterilization process, fifteen employees were configured to match the average number of employees in the station, thirteen of which are scheduled to work during normal weekdays split into shifts as mentioned in Section 6.2.5, and two are scheduled to work during weekends and holidays. 79 •Washing Machines and Sterilization Machines As per the number of washing machines and sterilization machines available. Eleven machineries were configured, five of which are washing machines, to match the number of identified washing machines in Section 6.2.6, and six sterilization machines, to match the number of the identified sterilization machines in the same section. •Schedules During the analysis, five shifts for employees working during weekdays and two shifts for employees working during weekends and holidays were identified in Section 6.2.5. Also, a change in the events related to machineries is noticed in the events log, as the employees start loading and running the machines overnights, to match this behavior, the machines in the simulation model are configured to run 24/7. These general configurations ensure that the simulation environment closely mimics the actual resource availability and operational constraints. 2. BPMN Model File For this research, the BPMN model illustrated in Figure 5.11 was loaded. This model accurately represents the original sterilization process, including all the entities that contribute to the process. 3. Simulation Configuration File The simulation configuration file is responsible for setting up the tasks and gateways running conditions, containing all the parameters necessary for the experiment. These parameters are detailed in Table 7.1, ensuring that the simulation is accurately tailored to the sterilization process. As we can see, the tasks ”Montaje” and ”Producci´on Montada” are grouped in an expanded sup-process task named ”Assembly”. this comes after deep analysis of the mentioned tasks, concluded that they represent one task, where ”Montaje” represents starting the task and ”Producci´on Montada” represents finishing the task. The simulation was executed separately for each data category (weekdays and weekends/holidays), running the process flow while taking into account resource availability and batching behavior. With a normal distributed arrival rate of a mean of 55.7 minutes and a standard deviation of 19.0, we are expecting no less than 28 kits to enter the 80 Task Distribution Resources Type Parameters Entrada Material Sucio Triangular Lower 2.0 Peak 14.0 Upper 40.0 Usuario Cargado en Carro L+D Triangular Lower 1.0 Peak 4.0 Upper 20.0 Usuario Carga L+D Iniciada Normal Mean 5.0 Standard Deviation 3.0 Usuario Instrumental Normal Constant Constant Value 60.0 Lavadora Delicado Constant Constant Value 60.0 Lavadora Priones Constant Constant Value 60.0 Lavadora Instrumetal Nuevo Constant Constant Value 60.0 Lavadora Contenedores Constant Constant Value 40.0 Lavadora Carros Constant Constant Value 40.0 Lavadora Carga L+D Liberada Triangular Lower 1.0 Peak 3.0 Upper 30.0 Usuario Assembly Triangular Lower 1.0 Peak 2.0 Upper 20.0 Usuario Composici´on de Cargas Normal Mean 5.0 Standard Deviation 3.0 Usuario INTRUM Uniform Lower 60.0 Upper 75.0 Autoclave GOMAS Uniform Lower 60.0 Upper 75.0 Autoclave CONTENE Uniform Lower 60.0 Upper 75.0 Autoclave PRIONES Uniform Lower 60.0 Upper 75.0 Autoclave Lumen Constant Constant Value 65.0 VPro No Lumen Constant Constant Value 40.0 VPro Temperatura Alta Constant Constant Value 50.0 Eagle Carga de Esterilizador Liberada Normal Mean 4.0 Standard Deviation 2.0 Usuario Table 7.1: Scylla sterilization process simulation configuration 81 sterilization unit on weekends/holidays, matching the real data rate. By following these configurations, this experiment would reproduce the evaluation and verify the accuracy of the discrete event-based digital twin under similar conditions. All scripts, datasets, and configuration files used during the evaluation are stored in the research’s repository1. 7.3 Execution The simulation took place on a MacBook Pro equipped with an Apple M2 chip, 8GB of memory, running MacOS Sonoma 14.0, and utilizing Python version 3.9.18. 7.4 Results With Scylla generated a range of performance metrics, including throughput, waiting times, and resource utilization rates, side by side with the event logs, the evaluation of the results is done over three categories, as follows: •Congestion Comparison The congestion comparison provides a comparison between the digital twin and the real-life process on several aspects, such as the number of processed kits per day and the time required to complete a full sterilization cycle. Subsequently, it provides a comparison between the standard and the modified Scylla versions. This comparison was conducted on both event logs, weekdays and weekends. •Goodness-of-Fit Evaluation In the goodness-of-fit evaluation, the K-S test was utilized to provide a distribution comparison between the event logs coming out of the digital twin and the real-life event logs. This evaluation was conducted on the weekends/holidays event logs. •Temporal Comparison An in-detailed comparison conducted on the weekends/holidays event logs to compare key performance metrics, such as the average processing time and the time required to perform an activity against real-world data. 1https://github.com/abu-sbeit/bdma-thesis 82 Date Data Source Mean STD Number of Cases 26/3/2022 Original 4.59 1.64 31 Simulator 4.07 2.90 39 27/3/2022 Original 3.20 0.91 8 Simulator 4.14 1.87 25 Table 7.2: A comparison between the original event data and simulation data over a weekend By conducting these comparisons between the digital twin event logs and the real-life event logs, the reliability of the digital twin is determined. 7.4.1 Congestion Comparison This section provides an overview of the number of kits processed and the time required to complete a sterilization cycle, including the weighted average mean and weighted standard deviation. A comparative analysis between the standard and repetitive activities versions of Scylla is conducted using relative activity frequency comparison to assess accuracy. Starting from the number of processed kits over the simulated period, the simulator shows a higher average number of sterilized kits, as illustrated in Table 7.2. Despite the increase in the number of processed kits, the simulation maintains a close weighted average mean, with values of 4.231 and a weighted standard deviation of 1.609 for the original event data, compared to a weighted average mean of 4.117 and a weighted standard deviation of 2.326 for the simulated data. This similarity suggests the potential for sterilizing more kits over the weekends. Figure 7.1 provides an overview of the histogram generated for case durations during a weekend period for completing a sterilization cycle. Furthermore, Figure 7.1(a) provides the histogram corresponding to the simulated data, while Figure 7.1(b) provides the histogram for the original data. As per on weekdays, Table 7.3 provides a detailed comparison between the original event data and the simulation data, specifically focusing on weekdays. The comparison includes key statistical metrics such as the mean, standard deviation, and the number of cases. 83 (a) Histogram of cases duration for simulated data (b) Histogram of cases duration for original data Figure 7.1: Histogram of cases duration on weekend period Date Data Source Mean STD Number of Cases 1/3/2022 Original 7.27 3.94 281 Simulator 6.64 5.15 294 2/3/2022 Original 7.80 5.25 295 Simulator 6.29 5.12 292 3/3/2022 Original 6.48 3.60 269 Simulator 6.76 5.28 292 Table 7.3: A comparison between the original event data and simulation data on weekdays 84 The comparison shows the ability of the digital twin to handle the same number of cases per day with a lower mean but a higher standard deviation, indicating a less stable sterilization environment leading to some cases taking longer sterilization time as illustrated in Figure 7.2. Figure 7.2: Sterilization cycle time distribution on weekdays Figure 7.3 illustrates a comparison between the original event logs and those generated by the digital twin using both the standard and repetitive activities versions of Scylla. The figure clearly demonstrates the capability of the repetitive activities version of Scylla to effectively manage and replicate the repetitive activities observed in the original event logs, by comparing ”Producci´on Montada”, ”Composici´on de Cargas”, and ”Carga de Esterilizador Liberada” activities between Figure7.3(a) and Figure 7.3(b). 7.4.2 Goodness-of-Fit Evaluation In this section, a detailed view of the results of comparing the distribution of each sterilization activity is given, starting from the point where contaminated kits enter the process until the sterilized kits leave. The similarities of the activities distribution between the digital representation of the process and the actual process were tested to determine whether our digital representation is a good representation of the process or not by 85 (a) Original Scylla Relative Activity Frequency Comparison (b) Modified Scylla Relative Activity Frequency Comparison Figure 7.3: Relative Activity Frequency Comparison utilizing the Kolmogorov–Smirnov (K-S) test. K-S test can be an indicator of whether to build an optimization methodology based on the simulation results or not. K-S test was utilized to compare the distribution between the processing time of the original data and the simulated data activities as shown in Table 7.4. Starting from the point where contaminated kits enter the process until they are sterilized, our results demonstrated a high degree of similarity in the histogram distribution of the time required to perform a sterilization activity. To represent the simulation results of each activity, a comparison between the original log and the digital twin generated logs was done as follows: •Entrada Material Sucio Figure 7.4 illustrates the histograms obtained for the duration of performing the activity Entrada Material Sucio using the digital twin and the original data. •Cargado en Carro L+D Figure 7.5 illustrates the histograms obtained for the duration of performing the activity Cargado en Carro L+D using the digital twin and the original data. Figure 7.5(a) shows less compatibility between the original data and the digital twin, unlike Figure 7.5(b), where it shows a higher similarity between them. 86 (e) Montaje histogram over the day for weekends/holidays (f) Producci´on montada histogram over the day for weekends/holidays (g) Composici´on de cargas histogram over the day for weekends/holidays (h) Carga de esterilizador liberada histogram over the day for weekends/holidays 93 Figure 7.12: Comparison between the time required to perform concatenated activities but rather the time between performing the activities. The approach adopted in this research involved concatenating events based on the factory layout, as illustrated in Figure 2.3. This method proved effective in aligning the detailed business process model developed during this research, as illustrated in Figure 5.11, with the sterilization unit’s process flow illustrated in Figure 2.2. Figure 7.12 illustrates the comparison between the time required to perform concatenated activities. in this Figure, the digital twin shows a higher mean duration for most activities except assembly, which indicates that the digital twin required more time to perform a certain activity and start the next one. As per the variability, the digital twin generally exhibits less variability across all the activities compared to the original process, which indicates that the digital twin is more consistent for the time required to perform a certain activity and start the next one. 7.5 Discussion The evaluation results confirm that the digital twin, configured and simulated using Scylla, provides a good representation of the sterilization process. The close correlation between the simulation and the historical data high94 lights the model’s reliability. Yet, this model can be tuned more to give a more accurate representation, especially with a clearer business understanding of the process itself and the expected duration of each activity. These results support the use of the digital twin for ongoing process analysis and optimization, ultimately contributing to improved operational efficiency and effectiveness. 95 Chapter 8 Conclusion Having the digital twin running results and KPIs comparison outcomes from the previous chapter, this chapter includes the SWOT analysis for the implementation and the conclusions of the research. 8.1 SWOT Analysis In order to simplify the outcomes of the research and to help in making the decision regarding the implementation of the leveled planning model, a SWOT analysis that includes the analyzed KPIs and observations is done as follows: •Strengths – Enhanced decision making. – Requires less manual planning efforts. – Simple to be used and monitored by all the related departments. •Weaknesses – High initiation costs. – Data quality dependency. •Opportunities – Scalability. – Sustainability. – Simplify process optimization. 96 – Provide an initial view of changing workforces or equipment. •Threats – The instability of the processes. 8.2 Conclusion, Recommendation, and Future Work According to the results of the thesis research, it would be recommended to implement a digital twin for the sterilization process due to the improvements that could be achieved by applying such a technology. The findings suggest the need for the company to implement comprehensive data recording practices for all aspects of the process. Specifically, it is recommended that the washing programs and durations be recorded separately from the loading and unloading of washing machines, and the sterilization programs and durations be documented independently of the loading and unloading of sterilization machines. Reflecting this would enhance the accuracy of the digital twin in mimicking the sterilization process. For future work, incorporating all possible scenarios of the sterilization process into the digital twin and presenting the workflow of the containers and the loaned kits, thereby enhancing its representation of the actual process. Additionally, a more enhanced approach for managing repetitive activities should be developed, as the current method relies on randomly generated values, which generates more repetitions than expected, the new approach should manage the number of repetitions. As per the digital twin, an enhanced approach of automatically extracting simulation parameters by incorporating machine learning models should be developed, this approach would make more accurate digital twins representing the real-world processes. 97 Bibliography [1] Apache airflow - a platform to programmatically author, schedule, and monitor workflows. [2] autotwin. [3] Camunda - the universal process orchestrator. [4] Ciw simulation python library. [5] BPT Lab. Scylla - simulation software for business process management. [6] Jason Brownlee. Statistical Methods for Machine Learning. Machine Learning Mastery, 2020. [7] David Chapela-Campa, Ismail Benchekroun, Opher Baron, Marlon Dumas, Dmitry Krass, and Arik Senderovich. Can i trust my simulation model? measuring the quality of business process simulation models. In Chiara Di Francescomarino, Andrea Burattin, Christian Janiesch, and Shazia Sadiq, editors, Business Process Management, pages 20–37, Cham, 2023. Springer Nature Switzerland. [8] Mehdi Naseriparsa Francesco Osborne Ciyuan Peng, Feng Xia. Knowledge graphs: Opportunities and challenges. 2023. [9] Dirk Fahland. Extracting and pre-processing event logs. 2022. [10] Dirk Fahland. Process mining over multiple behavioral dimensions with event knowledge graphs. 448, 2022. [11] Gartner. Gartner survey reveals digital twins are entering mainstream use. 2019. [12] Gartner. Top strategic technology trends 2023. 2023. [13] Michael Grieves. Origins of the digital twin concept. 08 2016. 98 [14] Michael Grieves and John Vickers. Digital Twin: Manufacturing Excellence through Virtual Factory Replication. Springer, 2017. [15] Ronny Mans, Hajo A. Reijers, and Michel van Genuchten. A data preparation framework for process mining: Smoothing out the quality issues. In Information Systems Evolution, pages 137–153. Springer, 2013. [16] Frank J. Massey. The kolmogorov-smirnov test for goodness of fit. Journal of the American Statistical Association, 46(253):68–78, 1951. [17] Douglas C. Montgomery and George C. Runger. Applied Statistics and Probability for Engineers. John Wiley & Sons, 7th edition, 2017. [18] Matteo Perno, Lars Hvam, and Anders Haug. Implementation of digital twins in the process industry: A systematic literature review of enablers and barriers. Computers in Industry, 134:103558, 2022. [19] Georgia Stavropoulou, Konstantinos Tsitseklis, Lydia Mavraidi, Kuo-I Chang, Anastasios Zafeiropoulos, Vasileios Karyotis, and Symeon Papavassiliou. Digital twin meets knowledge graph for intelligent manufacturing processes. Sensors, 24(8), 2024. [20] Dirk Fahland Stefan Esser. Multi-dimensional event data in graph databases. 2021. [21] Fei Tao, Meng Zhang, Yidong Liu, and Andrew Y. C. Nee. Digital twindriven product design, manufacturing and service with big data. The International Journal of Advanced Manufacturing Technology, 94:3563– 3576, 2018. [22] Wil van der Aalst. Process Mining: Data Science in Action. Springer, 2016. [23] Wil M. P. van der Aalst. Process mining: Discovery, conformance and enhancement of business processes. 2011. [24] Wil M. P. van der Aalst. Process mining: A 360 degree overview. 2022. [25] Wil M.P. van der Aalst. Process mining and simulation: a match made in heaven! In Summer Simulation Multiconference, 2018. [26] Wil M.P. van der Aalst. Toward more realistic simulation models using object-centric process mining. In European Conference on Modelling and Simulation, 2023. 99 [27] Wil M.P. van der Aalst. Twin transitions powered by event data - using object-centric process mining to make processes digital and sustainable. In ATAED/PN4TT@Petri Nets, 2023. [28] Wil M.P. van der Aalst and Alessandro Berti. Discovering object-centric petri nets. Fundam. Informaticae, 175:1–40, 2020. [29] Wikipedia contributors. Binomial distribution — Wikipedia, the free encyclopedia, 2024. [30] Wikipedia contributors. Continuous uniform distribution — Wikipedia, the free encyclopedia, 2024. [31] Wikipedia contributors. Erlang distribution — Wikipedia, the free encyclopedia, 2024. [32] Wikipedia contributors. Exponential distribution — Wikipedia, the free encyclopedia, 2024. [33] Wikipedia contributors. Normal distribution — Wikipedia, the free encyclopedia, 2024. [34] Wikipedia contributors. Poisson distribution — Wikipedia, the free encyclopedia, 2024. [35] Wikipedia contributors. Probability distribution — Wikipedia, the free encyclopedia, 2024. [36] Wikipedia contributors. Triangular distribution — Wikipedia, the free encyclopedia, 2024. 100 List of Figures 1.1 Digital Twin implementation overall methodology ....... 8 1.2 Phase One of Digital Twin Implementation overall methodology 9 1.3 Phase Two of Digital Twin Implementation overall methodology 10 2.1 Sterilization Center physical layout ............... 15 2.2 Sterilization Center logical layout . ............... 16 2.3 Factorylayout........................... 17 3.1 Digitaltwinconceptandmainelements............. 25 3.2 information flow representation for our DT solution . ..... 27 4.1 Process of getting the data .................... 30 5.1 EventKnowledgeGraphschema................. 41 5.2 Medical kit sterilization case from starting activity to ending activity .............................. 42 5.3 Medical kit sterilization case from starting activity to starting activity .............................. 42 5.4 Medical kit sterilization case from ending activity to ending activity .............................. 43 5.5 Container sterilization case from starting activity to ending activity .............................. 43 5.6 Container sterilization case from starting activity to starting activity .............................. 44 5.7 Container sterilization case from ending activity to ending activity................................ 44 5.8 Common starting activities for the sterilization process .... 45 5.9 Common ending activities for the sterilization process ..... 46 5.10 Sterilization process directly-follows graph ........... 47 5.11 BPMN representation of the real process ............ 49 6.1 Sterilization process implementation via Apache Airflow . . . 53 6.2 Scyllauserinterface........................ 54 101 6.3 Averagearrivalrateforkitsperhour .............. 55 6.4 Currentdispatchingruleoverone-monthanalysis ....... 58 6.5 Currentdispatchingruleoverdayslevelanalysis........ 59 6.6 Entrada Material Sucio distribution throughout the day .... 60 6.7 Entrada Material Sucio histogram of the time required to performtheactivity ......................... 61 6.8 Cargado en Carro L+D distribution throughout the day . . . 61 6.9 Cargado en Carro L+D histogram of the time required to performtheactivity ......................... 62 6.10 Carga L+D Iniciada distribution throughout the day ..... 63 6.11 Carga L+D Iniciada histogram of the time required to perform theactivity ............................ 63 6.12 Carga L+D Liberad distribution throughout the day ..... 65 6.13 Carga L+D Liberad histogram of the time required to perform theactivity ............................ 65 6.14 Assembly distribution throughout the day ........... 67 6.15 Composici´on de Cargas distribution throughout the day .... 67 6.16 Composici´on de Cargas histogram of the time required to performtheactivity ......................... 68 6.17 Average waiting and working times in minutes for the whole period ............................... 72 6.18Firstandlasteventscannedbyemployees ........... 73 6.19Scyllaglobalconfigurationfileuserinterface .......... 75 6.20 illustration of a process for a medical kit with single pieces . . 76 6.21 illustration of a process for a medical kit with multiple pieces . 76 7.1 Histogramofcasesdurationonweekendperiod ........ 84 7.2 Sterilization cycle time distribution on weekdays ........ 85 7.3 RelativeActivityFrequencyComparison ............ 86 7.4 Histogram of the duration required to perform Entrada MaterialSucioonweekendperiod................... 87 7.5 Histogram of the duration required to perform Cargado en CarroL+Donweekendperiod.................. 88 7.6 Histogram of the duration required to perform Carga L+D Iniciadaonweekendperiod.................... 88 7.7 Histogram of the duration required to perform Cargo L+D Liberadaonweekendperiod ................... 89 7.8 Histogram of the duration required to perform Montaje on weekendperiod .......................... 89 7.9 Histogram of the duration required to perform Producci´on Montadaonweekendperiod................... 90 102 Figure A.6: Normal Distribution The Normal distribution is extensively used in natural and social sciences due to the Central Limit Theorem, which states that the sum of a large number of independent random variables tends to be normally distributed, regardless of the original distribution. •Poisson Distribution The Poisson distribution is a discrete probability distribution that models the number of events occurring within a fixed interval of time or space, given that these events occur with a constant mean rate λand independently of the time since the last event [34]. The probability mass function is given by: f(x;μ, σ)= 1 σ√2πe−(x−μ)2 2σ2 Figure A.7: Poisson Distribution 109 This distribution is frequently used in fields such as traffic engineering, telecommunications, and insurance to model rare events, such as accidents or phone calls arriving at a call center. •Uniform Distribution The Uniform distribution is a continuous probability distribution where all outcomes are equally likely within a specified range [a,b] [30]. The probability density function is constant, given by: f(x;a, b)= 1 b−a,for a≤x≤b Figure A.8: Uniform Distribution This distribution is commonly used in simulations and random sampling, as it assumes all values within the specified range are equally probable. •Discrete Distribution Discrete distributions describe the probability of outcomes for discrete random variables, which can take on a countable number of distinct values [35]. Examples include the Binomial, Poisson, and Geometric distributions. These distributions are used in scenarios where variables represent counts or specific outcomes, such as the number of defective items in a batch or the number of arrivals at a service point. P(X=xi)=pi,for i=1,2,...,n 110 Figure A.9: Discrete Distribution 111