scieee AI-readable full text Open interactive document viewer

Distributed Machine Learning-based Digital Twins Modelling

Mutegeki, Henry; Bunino, Matteo

Abstract

This project delves into the integration of two distinct digital twin (DT) use cases: the DT for Drought Early Warning in the Alps (EURAC) and the Noise Simulation for Gravitational Waves Detector DT (VIRGO), utilizing the itwinai1 library designed for distributed machine learning modelling within high-performance computing (HPC) environments. A key focus is on measuring the performance overhead of itwinai in distributed machine learning across various resource configurations, facilitated by scaling tests incorporated within these DT use cases. Additionally, the project explores Hyperparameter Optimization (HPO) to fine-tune the machine learning models used in these scenarios.

Full text

Distributed Machine Learning-based Digital Twins Modelling September 2024 AUTHOR(S): Henry Mutegeki [email protected] SUPERVISOR(S): Matteo Bunino CERN openlab Report // 2024 PROJECT SPECIFICATION Project Title: Distributed Machine Learning-based Digital twin modelling. Project overview: This project delves into the integration of two distinct digital twin (DT) use cases: the DT for Drought Early Warning in the Alps (EURAC) and the Noise Simulation for Gravitational Waves Detector DT (VIRGO), utilizing the itwinai1library designed for distributed machine learning modelling within high-performance computing (HPC) environments. A key focus is on measuring the performance overhead of itwinai in distributed machine learning across various resource configurations, facilitated by scaling tests incorporated within these DT use cases. Additionally, the project explores Hyperparameter Optimization (HPO) to fine-tune the machine learning models used in these scenarios. Objective: The primary objective is to integrate these DT use cases into the itwinai framework and rigorously evaluate its impact on enhancing the efficiency and effectiveness of distributed machine learning modelling for digital twins with scaling resources. Results: The DT use cases were successfully integrated to use itwinai. Upon integration of scaling tests, we see that itwinai is efficiently leveraging compute resources in HPC environment to cut down compute time/model training time by half with each doubling of GPU resources. HPO reveals the best hyper parameter combinations and shows the power of running it on large compute infrastructure as it is able to make multiple parallel runs with different parameter combinations intended to get the best performance. 1https://itwinai.readthedocs.io/ 2 CERN openlab Report // 2024 ABSTRACT Digital twin applications hold potential for advancing scientific research, yet their adoption is currently hampered by significant overheads related to engineering tasks such as standardization, development, deployment, and testing. These challenges slow down research progress, as scientists are often diverted from their primary focus. In response to these challenges, CERN, in collaboration with other European research institutes, has embarked on a mission to develop the prototype of a unified digital twin engine (interTwin project2) that supports a wide range of reproducible applications, specifically focusing on distributed machine learning (ML) solutions in high-performance computing (HPC) environments. This project explores the integration of two specific digital twin use cases—EURAC and VIRGO—into itwinai, a novel library designed for distributed ML modelling. By embedding these use cases into itwinai, the project aims to validate the library's effectiveness in enhancing the efficiency and scalability of digital twin applications. In addition to integration, also seek to observe and quantify the impact of scaling up computational resources—particularly GPUs—on the training times of machine learning models within the itwinai framework. Our investigation centers on validating itwinai’s ability to train models more efficiently in environments equipped with substantial computational resources. By methodically increasing the number of GPUs, we aim to demonstrate that itwinai not only significantly reduces model training times but also confirms the framework's scalability. This capability is crucial for deploying machine learning models in large-scale, resource-intensive scenarios, proving itwinai’s effectiveness in leveraging high-performance computing environments to optimize both speed and resource utilization. The results have demonstrated that increasing compute resources significantly boosts performance capabilities, leading to decreased model processing times. Specifically, scaling tests within the digital twin use cases have shown that itwinai effectively harnesses additional GPU resources, consistently achieving a substantial reduction in model training times. This efficiency improvement exemplifies the scaling ability of itwinai in model training with increasing resources. This efficient utilization of HPC resources underscores the potential of itwinai to streamline complex ML workflows in scientific research. Moreover, the application of hyper-parameter optimization (HPO) within these use cases has revealed the most effective hyper-parameter combinations, further showcasing the benefits of running such optimizations on large-scale compute infrastructures. By addressing the existing bottlenecks in digital twin development, this 2https://www.intertwin.eu/ 3 CERN openlab Report // 2024 work contributes to making high-level scientific research more accessible and efficient, ultimately accelerating the pace of discovery and innovation. 4 CERN openlab Report // 2024 TABLE OF CONTENTS INTRODUCTION 01 INITIAL PROJECT STATE 02 CONTRIBUTIONS 03 RESULTS 04 CONCLUSIONS & RECOMMENDATIONS 05 5 CERN openlab Report // 2024 1. INTRODUCTION A Digital Twin (DT) is a virtual representation of a physical object, system, or process, designed to mirror the behavior, performance, and characteristics of its real-world counterpart. Digital twin technology represents a significant advancement in modelling and simulation, providing a virtual representation of physical systems that allows scientists to simulate, predict, and optimize the performance of these systems in real-time. This approach is increasingly valuable in research, where the ability to replicate and test complex systems without physical constraints can accelerate discovery and innovation. However, despite its potential, the adoption of digital twins in scientific research is often limited by significant technical overheads. Researchers are frequently required to navigate complex engineering tasks such as standardization, development, deployment, and testing, which diverts their focus from scientific inquiry and slows the progress of their work. Fig 1: interTwin use cases. The interTwin project is a collaborative effort including CERN3and a number of European research institutes4aimed at building a prototype of an Interdisciplinary 4https://www.intertwin.eu/project-partners 3https://home.cern/ 6 CERN openlab Report // 2024 Digital Engine (DTE). This DTE is characterized by being modular, open source, interdisciplinary and having a unified platform. The major contribution of the interTwin project is the itwinai Python library that is aimed at distributed machine learning model training for DT use cases. The interTwin project has a number of DT use cases which can largely be categorized into two: Climate research and Physics. For the scope of this project, we zeroed down to two main use cases i.e EURAC and VIRGO. Fig 2: VIRGO use case. The VIRGO5,6 use case for the VIRGO Gravitational Wave interferometer involves creating a precise virtual model to simulate and analyze noise disturbances affecting measurement accuracy. This model aids in studying how the detector responds to external disruptions and in developing strategies for noise reduction and subtraction. Enhancing the digital twin's capabilities allows for more reliable gravitational wave detections and early alerting of other observatories. 6https://www.intertwin.eu/intertwin-use-case-virgo 5https://itwinai.readthedocs.io/latest/use-cases/use_cases.html#noise-simulation-for-gravitational-waves -detector-virgo-infn-use-case 7 CERN openlab Report // 2024 Fig 3: Architecture of the EURAC use case. EURAC7is developing a digital twin project for Drought Early Warning in the Alps, merging process-based hydrological models with machine learning to enhance predictions and manage water resources more effectively. This system leverages Earth Observation data and a user-friendly interface developed in openEO8to facilitate real-time drought monitoring and forecasting at the river basin scale. The initiative supports local authorities and researchers in improving water resource management and drought resilience, directly addressing the challenges posed by climate change and increasing water demand in the region. In the context of digital twin applications, High-Performance Computing (HPC) plays a crucial role. HPC enables the processing of large datasets and complex Machine Learning modelling at unprecedented speeds, which is essential for the real-time performance and scalability of digital twins. However, leveraging HPC resources effectively requires sophisticated tools and frameworks that can distribute computing tasks across multiple nodes and GPUs. This is where scaling tests become vital—they assess how well a system performs as computational resources are increased, ensuring that the digital twin applications can scale efficiently and make full use of the available HPC infrastructure. Another critical component of optimizing digital twin applications is Hyper-Parameter Optimization (HPO). In machine learning, hyper-parameters are configuration variables set before training an ML model that directly control the model structure & learning 8https://openeo.cloud/ 7https://www.intertwin.eu/intertwin-use-case-a-digital-twin-for-drought-early-warning-in-the-alps 8 CERN openlab Report // 2024 function during the training process of a model, and finding the optimal combination of these configurations is crucial for maximizing model performance. HPO involves systematically exploring different hyper-parameter combinations to identify the best-performing configuration. In the context of digital twins, effective HPO can lead to more accurate and efficient ML modelling, but it requires substantial computational resources, further underscoring the importance of HPC. In this project we applied Ray Tune9to implement HPO for the DT use cases. We selected Ray Tune because of its robust capabilities in handling large-scale optimizations. Built on the distributed computing framework Ray, Ray Tune offers exceptional scalability, allowing it to efficiently manage extensive computational resources across multiple machines and GPUs. This is particularly advantageous for the computationally intensive tasks involved in digital twin modelling. Moreover, Ray Tune supports a wide variety of optimization algorithms, integrates seamlessly with major machine learning frameworks like TensorFlow and PyTorch, and provides sophisticated tools for experiment tracking and early stopping, making it an ideal choice for rigorous and resource-intensive HPO tasks. This project seeks to address these challenges by integrating digital twin use cases into itwinai10, a library designed for distributed machine learning in HPC environments. By focusing on scaling tests and HPO, the project aims to validate the effectiveness of itwinai in enhancing the performance of digital twin applications. The ultimate goal is to develop a standardized, reproducible framework that simplifies the engineering overheads for researchers, enabling them to leverage the full potential of digital twin technology in their scientific endeavors. 10 https://itwinai.readthedocs.io/ 9https://docs.ray.io/ 9 CERN openlab Report // 2024 replicate models across GPUs on multiple nodes, whereas Horovod (originally developed by Uber) provides a framework-agnostic approach to distribute workloads across numerous GPUs. By employing these strategies, computational resources could be utilized more effectively, reducing training time and improving overall throughput in HPC environments. Similar to the GAN integration, a gather method was implemented to ensure that losses from different devices could be collected onto a main device for consolidated analysis. This data aggregation is critical for combining partial results and simplifying the logging process. Ultimately, the VIRGO use case integration showcased itwinai’s capacity to handle high-dimensional, time-series data derived from gravitational wave observations and to efficiently leverage HPC resources for model training. 3.1.3 EURAC Use Case Integration In contrast to VIRGO, the EURAC integration required handling geospatial, sequential, and sometimes sparse datasets typical in environmental and climate research. To address this, the EURAC integration led to the development of an RNNDistributedTrainer class within the itwinai framework. This specialized trainer incorporated modifications in the training and validation steps to adapt to the complexities of ecological data modelling. Custom samplers were integrated into the create_dataloaders method to accommodate the unique characteristics of geospatial datasets. Before aggregating results, the gather method was introduced and explained, enabling the consolidation of intermediate results (e.g., losses, predictions) from multiple distributed devices onto a designated main device. This ensures comprehensive monitoring and logging of training outcomes, which is crucial when dealing with multi-node or multi-GPU setups. The integration also included advanced logging features through MLFlow to facilitate detailed tracking of model performance and diagnostics. DDP and Deepspeed strategies were similarly supported, optimizing the training process across distributed systems. To assist users in deploying this advanced setup, a comprehensive tutorial14 14 https://github.com/interTwin-eu/itwinai/blob/main/use-cases/eurac/README.md 16 CERN openlab Report // 2024 was created, detailing the steps to run the EURAC use case within the itwinai framework, thereby extending its utility in environmental science research. Fig. 7 represents the raw values of the metrics like loss, accuracy, distribution strategy, batch size and other experimental properties. Fig 7: Results tracked from a run in mlflow for the Distributed EURAC use case. 3.2 HPO Integration A significant part of the contribution was the implementation of hyper-parameter optimization (HPO) techniques within these use cases. This significantly enhanced the EURAC and VIRGO use cases by using Ray tune, executed in a high-performance computing (HPC) environment but outside the distributed framework of itwinai. By systematically exploring various hyper-parameter configurations (e.g., learning rates, optimizer types, and batch sizes), the project ensured that the final models for these use cases met both the scientific accuracy demands and computational efficiency requirements. Leveraging Ray Tune’s robust capabilities allowed large-scale experiments to be carried out simultaneously, taking advantage of HPC parallelism. 17 CERN openlab Report // 2024 Consequently, the models were not only customized to meet the domain-specific needs of EURAC and VIRGO but also underwent performance-focused tuning that set a solid foundation for future integrations of advanced ML techniques within itwinai. This methodical approach to HPO underscores the importance of both domain knowledge and computational resource optimization in scientific research, particularly when high-dimensional parameter spaces must be navigated. 3.3 Scaling Analysis Additionally, scaling tests were conducted to assess and enhance the performance of itwinai across increasing HPC resources. The scaling tests were designed to measure time for computation processes while training DT models in itwinai. The tests measured epoch times and the duration per run across various GPU configurations and distribution strategies. Key data points such as the distribution strategy, number of nodes, epoch identifier, and exact epoch timing were systematically logged for each resource configuration. These tests were critical in demonstrating the library’s ability to leverage additional compute resources(like GPUs) to speed up processing times effectively, thereby validating its scalability and performance efficiency in a controlled, rigorous manner. 18 CERN openlab Report // 2024 4. RESULTS The integration of the EURAC and VIRGO use cases using the itwinai library yielded substantial improvements in the operational efficiency and model performance of digital twins. 4.1 Scaling Analysis Results The scaling tests revealed that itwinai could capitalize on increasing GPU resources to significantly reduce model training times, with times decreasing by nearly half for each doubling of GPU availability close to the ideal scenario. This demonstrated the library’s scalability and resource efficiency, key attributes for its applicability in large-scale scientific computing environments. Figures 8 and 9 delineate the efficacy of GPU resource scaling within high-performance computing (HPC) environments. For the VIRGO use case, the speedup results demonstrated by both DDP-torch and Horovod-torch distribution strategies, as shown in Figure 8, reveal a pronounced decrease in model training times as the number of GPUs increased. Although neither strategy achieved ideal linear speedup, DDP-torch closely approached this ideal, indicating an efficient utilization of the expanded GPU resources. This trend underscores the itwinai library's capability to effectively manage increased computational loads, a critical attribute for scaling operations in complex scientific computations. Fig 8: Scaling test results for increasing GPU resources in a run for VIRGO use case. 19 CERN openlab Report // 2024 Conversely, Figure 9, which focuses on the EURAC use case using the DDP-torch strategy, presents a consistent upward trend in processing speed with the addition of more GPUs. The curve, while less steep than in the VIRGO case, still indicates a solid performance in leveraging additional computational resources to accelerate processing times. This consistent improvement across varied computational demands highlights the itwinai library’s robust scalability and adaptability, confirming its suitability for diverse scientific research environments. We noted that the EURAC use case worked specifically well with the DDP-torch strategy for the scaling analysis task compared to other strategies like Horovod and this was largely attributed to the ease of integration of the gather_method in the ddp_torch strategy compared to Horovod where if was difficult to gather metrics across multiple nodes. Fig 9: Scaling test results for increasing GPU resources in a run for EURAC use case. Overall, the observed results from these scaling tests validate the itwinai library's potential to significantly enhance operational efficiency and model performance in digital twin applications. By facilitating faster and more efficient model training sessions, the library not only meets the computational demands of modern scientific research but also propels forward the capabilities of digital twins within large-scale HPC settings. 20 CERN openlab Report // 2024 4.2 HPO Results The Hyper-Parameter Optimization (HPO) processes implemented provided valuable insights into the digital twin modelling for both the VIRGO and EURAC use cases. The optimization focused on refining the parameters such as batch size and learning rates, crucial for enhancing the predictive accuracy and reliability of the models. The tabulated data from Table 1 summarizes the outcomes of ten HPO trials, each representing the best results achieved under varying configurations. Notably, the trials consistently reported low loss values, indicating that the itwinai library's HPO capabilities effectively pinpointed efficient configurations from a complex hyperparameter space. This optimization highlights the library’s proficiency in adjusting parameters to significantly improve model performance. Furthermore, the inclusion of the processing times and configuration details in the table provides a comprehensive view of each trial's computational demands and the efficacy of the selected parameters. trial_id loss training_ iteration time_this_iter(s) time_total(s) batch size config/lr f4491_00000 0.06200347096 50 1.18600893 63.92927647 32 0.0002533304587 f4491_00001 0.06200347096 50 1.17693138 64.02291083 32 0.0009182808133 f4491_00002 0.06200347096 50 1.17662239 63.97996587 32 0.00007978080127 f4491_00003 0.06200347096 50 1.16645956 64.12123585 32 0.0004832588602 f4491_00004 0.06349471211 10 1.15288687 13.69036794 64 0.00002243542462 f4491_00005 0.06349471211 10 1.15916745 13.72511101 64 0.0005037833977 f4491_00006 0.0407073386 50 1.52991056 81.44102979 16 0.0003349316915 f4491_00007 0.04037203267 50 1.54295039 81.44326663 16 0.0001760499314 f4491_00008 0.03925458342 50 1.5378902 80.62055612 16 0.0000104975537 f4491_00009 0.06276089698 10 1.7507894 19.00940681 16 0.00003727533999 Table 1: The best results tracked from an HPO experiment(of 10 runs) for the VIRGO use case. Loss Progression for VIRGO Use Case: The graph in figure 10 depicting the loss progression across nine runs reveals that certain trials were terminated early due to high loss values—a feature enabled by Ray Tune's efficiency in resource management. 21 CERN openlab Report // 2024 This adaptive stopping mechanism is instrumental in conserving computational resources by discontinuing less promising experiments, thereby streamlining the HPO process. The runs that exhibited continuous progression demonstrated a consistent decrease in loss over time, affirming the effectiveness of the selected hyperparameters and the robustness of the itwinai library in managing the training process. Fig 10: Loss progression from an HPO experiment having 9 runs for VIRGO use case. Loss Progression for EURAC Use Case: Similarly in fig. 11, the loss progression for the EURAC use case was visualized over ten runs, each influenced by distinct hyperparameter sets. This graph served as a critical tool in assessing the direct impact of hyperparameter variations on the model's learning process. Runs that quickly achieved lower loss values underscored the potential of optimal parameter settings to enhance model accuracy and efficiency significantly. Such insights are crucial for iteratively refining the model parameters, aiming for an optimal balance between speed and performance in digital twin modelling experiments. Certain hyper-parameter configurations (for example, low learning rates or less optimal architectures) may lead the model to stagnate, causing nearly flat loss curves at higher loss values. 22 CERN openlab Report // 2024 Fig 11: Loss progression from an HPO experiment having 10 runs for EURAC use case. These results collectively underscore the itwinai library’s advanced capabilities in executing and managing complex HPO tasks within high-performance computing environments. By facilitating more precise and dependable ML modelling, the library empowers researchers to advance scientific applications, leveraging efficient computational strategies to achieve high-quality outcomes in a reduced timeframe. This optimization plays a pivotal role in accelerating the pace of scientific discovery and application, particularly in the context of digital twin modeling. 23 CERN openlab Report // 2024 5. CONCLUSIONS & RECOMMENDATIONS This project has successfully demonstrated that the itwinai library can be effectively tailored and utilized for specific digital twin applications within high-performance computing environments, enhancing both scalability and efficiency. The integration of the GAN, EURAC and VIRGO use cases has not only validated the practical utility of itwinai but has also highlighted its potential for broader application across other scientific disciplines. For future endeavors, it is recommended that further research be directed towards enhancing the HPO processes to run in distributed environments for the optimization of digital twin models. Additionally, expanding the scope of scaling tests to include a wider range of computational resources and use cases could provide deeper insights into the library's adaptability and performance under varying conditions. Finally, ongoing enhancements and improvements to the itwinai library are essential to ensure it stays effective and relevant amidst rapid developments in digital twin technologies and machine learning methods. 24