Full text
CERN openlab Report // 2025 PROJECT SPECIFICATION Digital twins are increasingly used in scientific research, but it is important to ensure their scalability across High-Performance Computing (HPC) environments. This project focuses on performing scaling tests for interTwin AI workflows on HPCs. The goal is to assess the scalability of AI workflows using itwinai library[14]. The scaling tests involve distributed machine learning (ML) training and hyperparameter optimization across multiple nodes, evaluating key metrics such as computation time, communication overhead, and resource utilization. The project requires an understanding of Python, HPCs, distributed ML frameworks such as PyTorch DDP[13], Horovod[6], and Ray Tune[5]. This project offers an opportunity to deepen an understanding of scalable AI workflows, HPC environments, and performance optimization in scientific computing. 2
CERN openlab Report // 2025 ABSTRACT This report evaluates how well the interTwin Drought Early Warning in the Alps workflow scales on HPC systems using the open-source itwinai library[14] developed by CERN. We run multi-GPU, multi-node training, comparing distributed strategies (PyTorch DDP[13], DeepSpeed[7], and Horovod[6]). Using itwinai’s built-in Scalability Report[4], we collect and analyze core signals: average epoch time, relative speedup, GPU utilization, GPU power (Wh), and a compute-vs-other breakdown that surfaces non-compute overheads such as communication and data stalls. The study documents how the scaling is observed across multiple nodes/GPUs, and how efficiency varies across strategies. All experiments are reproducible: configurations and code follow the public hython itwinai plugin and itwinai guides, with metrics logged to MLflow for automated report generation. 3
CERN openlab Report // 2025 TABLE OF CONTENTS Introduction 06 Objectives and approach Scope of this report Overview of Hython and its Integration with itwinai 08 Primer on Distributed Training 09 Purpose Data‑parallel training in this workflow Process‑to‑GPU mapping and launch model Strategies implemented Platforms, runtime & measurement setup 10 HPC systems used Launch configuration and process: GPU mapping Measurement and instrumentation Rationale for cross‑site evaluation Scaling results & embedded diagnostics 13 JUWELS: initial (poorly scaling) runs JUWELS: corrected runs Vega: cross-site check (1, 2, 4 nodes) 4
CERN openlab Report // 2025 Conclusion & Artifacts to reproduce 22 Conclusions Artifacts to reproduce References 22 5
CERN openlab Report // 2025 1. Introduction Scientific digital twins are moving from concept to practice across climate, energy, and physics. The EU-funded interTwin project is building an open, interoperable Digital Twin Engine (DTE) [1] to help research teams assemble and operate domain-specific twins at scale. The DTE and its components are released under open-source licenses and are intended to run across federated research infrastructure, from local systems to modern HPC machines. This report focuses on the interTwin use case “Drought Early Warning in the Alps” [2], developed in collaboration with Eurac Research [3]. The goal is to combine process-based hydrological models with machine learning and Earth observation data to support basin-scale drought prediction and early warning for Alpine regions. Because such workflows must ingest large gridded datasets and train models across long time windows, scalability on HPC is not just optional; it is foundational to timely, actionable forecasts. To evaluate and improve scalability, we use itwinai[14], an open-source Python library developed within interTwin, primarily by CERN. itwinai streamlines distributed training, hyperparameter optimization (HPO), experiment logging, and SLURM job management, so that scientific teams spend less time on infrastructure and more on accelerating science. It supports multiple distributed strategies out of the box, like PyTorch DDP, Horovod[13], and Microsoft DeepSpeed[7], and provides a Scalability Report [4] that aggregates key signals (average epoch time, relative speedup, GPU utilization, GPU energy in Wh, and a compute-vs-other/communication view) to pinpoint bottlenecks as node and GPU counts grow. Our distributed ML stack follows widely used, well-documented approaches. For data-parallel training, we rely on PyTorch DistributedDataParallel (DDP)[13] and Horovod[6], and we also exercised DeepSpeed[7] where it made sense for optimizer/communication efficiency. For hyperparameter search, we use Ray Tune, which runs trials at cluster scale. All runs are tracked in MLflow to capture parameters, metrics, artifacts, and code versions for full reproducibility.[5][6][7][8] Early runs exposed a distribution-strategy misconfiguration and an under-provisioned data loader; we therefore report ‘before’ and ‘after’ results and explain the fix. All experiments in this report ran on JUWELS Booster (JSC)[9] and Vega[10], and we deliberately present both the initial poorly scaling runs and the corrected results to show how the fixes changed the picture. a. Objectives and approach We first describe the Hython surrogate and its itwinai plugin, then run controlled scaling experiments on JUWELS Booster and Vega using identical configurations. With itwinai’s Scalability Report[4] and MLflow, we measure time, speedup/efficiency, GPU utilization/energy, and compute-versus-other costs. We present “before/after” results: prior to and after fixing distribution and dataloader issues, and close with guidance on right-sizing training and hyperparameter search. 6
CERN openlab Report // 2025 b. Scope of this report This report documents the scalability evaluation of the hython–itwinai workflow across High-Performance Computing (HPC) systems. It begins with an overview of the Alpine drought workflow, the Hython surrogate model, and its integration with the itwinai framework. A short primer on distributed training follows, providing the conceptual background necessary to interpret the scaling diagnostics. The next section details the computing platforms, runtime configuration, and measurement setup used in the experiments. The subsequent section presents the scaling results obtained on JUWELS Booster and Vega, starting with the initial poorly scaling runs, the applied corrections, and the verified improvements across systems. The report concludes with a summary of findings and a list of artifacts required to reproduce the experiments. 7
CERN openlab Report // 2025 2. Overview of Hython and its Integration with itwinai Hython is an open-source Python package for building deep-learning surrogates of semi-distributed and distributed hydrological models. It provides a dataset and model abstractions suitable for gridded hydrology and ships reference LSTM/ConvLSTM configurations and a command-line interface for preprocessing, training and evaluation. The code is maintained in the interTwin community and is distributed under an open license. The interTwin itwinai[14] framework is a separate open-source library that streamlines scalable AI workflows on HPC systems. It exposes a declarative CLI and Python API for distributed training (PyTorch DistributedDataParallel[13], Horovod[6], DeepSpeed[7]), hyperparameter optimization, experiment tracking with MLflow, SLURM script generation and job submission, and a built-in Scalability Report[4] that aggregates epoch time, relative speedup/efficiency, GPU utilization and related signals. The hython–itwinai plugin bridges these two layers. It packages the Hython drought-forecasting workflow so that it can be launched via itwinai on HPC clusters, logged to MLflow, and analyzed with the Scalability Report, without changing the Hython training code. This study used the public Hython codebase together with the hython–itwinai plugin and the itwinai runtime to assess scaling of the Hython surrogate workflow across multiple GPUs and nodes. Comparative runs exercised PyTorch DDP[13], Horovod[6] and DeepSpeed[7] back-ends; jobs were submitted through itwinai’s SLURM tooling; metrics were recorded in MLflow and summarized by the Scalability Report to diagnose bottlenecks and quantify efficiency. 8
CERN openlab Report // 2025 3. Primer on Distributed Training a. Purpose This brief primer gives readers background to interpret the scalability plots that follow. It explains the training paradigm used in the study, how processes map to GPUs in the HPC runs, and what the chosen back‑ends change. This primer will help the reader to understand the plots that are presented in section 6. b. Data‑parallel training in this workflow itwinai[14] adopts data parallelism strategy, as depicted in Fig 3.1: the AI model is replicated on each GPU; every replica processes a disjoint shard of the mini‑batch, computes gradients locally, and synchronizes them across ranks before applying the update. This preserves a single logical model while allowing throughput to scale with the number of GPUs. In our configuration, itwinai provides the distributed runtime so that the Hython surrogate can be trained with identical code paths on a workstation or across nodes on HPCs like JUWELS Booster and Vega. Fig. 3.1: Visual representation of data-parallel training algorithm. The training dataset is split into partitions which are fed to individual workers. Each worker is assigned one GPU and trains a copy of the model. In this example, two compute nodes with two GPUs each are used. c. Process‑to‑GPU mapping and launch model Jobs are launched under Slurm with the standard one process per GPU, generally mapping on 4‑GPU nodes; scale‑out is achieved by increasing the number of nodes. The total number of processes are referred to as the world size; each process has a global rank and a per‑node local rank. Correctness and performance both rely on per‑rank data sharding so that each process reads a unique subset of the dataset (e.g., via a distributed 9
CERN openlab Report // 2025 Fig 4.1e JUWELS (before): GPU energy consumption b. JUWELS: corrected runs After the training dataset was properly partitioned across workers, the plots exhibited the expected behaviour: the epoch time decreased consistently as the number of workers increased from 4 to 8, 16, and 32 (Fig. 4.2a), while the speedup followed the theoretical baseline at smaller scales before gradually tapering off (Fig 4.2b). These are precisely the trends that the Scalability Report is designed to reveal. The GPU energy consumption plot (Fig. 4.2e) further corroborates the improved scalability. As node count increases, the total GPU energy usage remains approximately constant, reflecting efficient parallelization and reduced wall-clock duration per epoch. In this corrected configuration of the Hython plugin executed through itwinai, each worker now processes a distinct data shard, minimizing redundant I/O and synchronization stalls. Consequently, although additional GPUs are employed, their cumulative energy draw per training run remains nearly unchanged, as the task completes proportionally faster. This behavior is characteristic of well-scaling distributed training, where increasing computational resources yields time-to-solution gains without inflating total energy expenditure. 16
CERN openlab Report // 2025 Fig. 4.2a JUWELS (after): Average epoch time Fig. 4.2b JUWELS (after): Relative speedup 17
CERN openlab Report // 2025 Fig 4.2c JUWELS (after): compute vs. other split Fig 4.2d JUWELS (after): GPU utilization 18
CERN openlab Report // 2025 Fig 4.2e JUWELS (before): GPU energy consumption c. Vega: cross-site check (1, 2, 4 nodes) We repeated the corrected configuration on Vega’s GPU partition using 1, 2, and 4 nodes (4× A100 40 GB per node; 60 GPU nodes in total). As Vega is a shared Slurm system, jobs were executed as nodes became available; during the testing window, the queue was saturated, causing smaller runs to complete while longer multi-node jobs remained pending. Since Slurm partitions function as job queues with priority and fair-share scheduling, large requests can experience extended waiting times. The scaling curves, observed in Fig 4.3a and Fig 4.3b for one to four nodes on Vega followed the same trend observed on JUWELS after the data-loader correction. Vega’s GPU nodes (4× A100 40 GB, dual HDR100) and JUWELS Booster (936 nodes, 4× A100 per node, HDR-200/DragonFly+) thus provided two distinct A100-based systems with different interconnect fabrics on which to verify the fix. 19
CERN openlab Report // 2025 Fig. 4.3a Vega: Average epoch time Fig. 4.3b Vega: Relative speedup 20
CERN openlab Report // 2025 Fig. 4.3c Vega: GPU utilization Fig. 4.3d Vega: GPU energy consumption 21
CERN openlab Report // 2025 6. Conclusion & Artifacts to reproduce a. Conclusions We used the public interTwin stack (Hython + itwinai) to evaluate how the Alpine drought workflow scales on JUWELS Booster and, for a cross-site validation, on Vega. The initial JUWELS runs did not scale as expected because the dataloader failed to shard the dataset, causing every rank to read the full input and eliminating any benefit from additional GPUs. After correcting the loader, epoch time decreased consistently with increased GPU count, and the same behavior was reproduced on Vega, confirming that the fix generalizes across systems. Importantly, the itwinai Scalability Report[4] played a central role in identifying this issue. Its detailed diagnostics, particularly the breakdown of epoch time, and speedup were instrumental in uncovering the dataset sharding problem, which would have been significantly harder to detect through log inspection or standard metrics alone. This demonstrates the effectiveness of itwinai’s integrated reporting in enabling systematic performance analysis and accelerating debugging in distributed training workflows. b. Artifacts to reproduce 1. Code a. Hython (surrogate & configs): GitHub repo. b. itwinai (runner, profiling, Scalability Report): repo + docs[14]. 2. Systems a. JUWELS Booster specs & partitions (booster, develbooster) b. Vega GPU partition (gpu) and A100 layout. 3. How we ran it a. Create a Python venv per itwinai’s docs; install the plugin from source. b. Submit Slurm jobs with itwinai exec-pipeline c. Render figures with the Scalability Report CLI[4]: itwinai generate-scalability-report … 7. References [1] interTwin Project: Digital Twin Engine (DTE) overview (accessed Sept. 2025). [2] interTwin. “Drought Early Warning in the Alps – use case.” (accessed Sept. 2025). [3] Eurac Research. “About Us.” (accessed Sept. 2025). [4] itwinai Docs. “Scalability Report.” (accessed Sept. 2025). [5] Ray Project Docs. “Ray Tune: hyperparameter tuning at scale.” (accessed Sept. 2025). [6] Horovod Docs. “Distributed training overview.” (accessed Sept. 2025). [7] DeepSpeed Docs. “System overview and features.” (accessed Sept. 2025). 22
CERN openlab Report // 2025 [8] MLflow Docs. “Tracking: parameters, metrics, artifacts, and code.” (accessed Sept. 2025). [9] Jülich Supercomputing Centre: JUWELS Booster (system overview) (accessed Sept. 2025). [10] EuroHPC / IZUM. “Vega: GPU partition and node specs.” (accessed Sept. 2025). [11] PyTorch Docs. “DistributedSampler (multi-GPU data loading).” (accessed Sept. 2025). [12] wflow Docs. “wflow_sbm.” (accessed Sept. 2025). [13] PyTorch Docs. “DistributedDataParallel (DDP).” (accessed Oct. 2025). [14] itwinai Docs. “Documentation site.”(accessed Sept. 2025). 23