scieee AI-readable full text Open interactive document viewer

Optimizing I/O for an Exascale Implicit Kinetic Plasma Simulation using the Rabbit Storage System

Lumsden, Ian; Devarajan, Hariharan; Yildirim, Izzet; Markidis, Stefano; Hu, Andong; Peng, Ivy B.; Pennati, Luca; Yokelson, Dewi; Brink, Stephanie; Pearce, Olga; Scogland, Thomas; de Supinski, Bronis; Delzanno, Gian Luca; Kougkas, Anthony; Sun, Xian-He; T

Full text

LLNL-PRES-2010660 This work was performed under the auspices of the U.S. Department of Energy by Lawrence Livermore National Laboratory under contract DE-AC52-07NA27344. Lawrence Livermore National Security, LLC Optimizing I/O for an Exascale Implicit Kinetic Plasma Simulation using the Rabbit Storage System Ian Lumsden, Hariharan Devarajan, Izzet Yildirim, Stefano Markidis, Andong Hu, Ivy Peng, Luca Pennati, Dewi Yokelson, Stephanie Brink, Olga Pearce, Tom Scogland, Bronis R. de Supinski, Gian Luca Delzanno, Anthony Kougkas, Xian-He Sun, Michela Taufer 2 LLNL-PRES-2010660 Exascale Computing Powers Breakthrough Scientific Discovery Running on El Capitan, iPIC3D can achieve fully kinetic, three-dimensional simulation of small-tomedium planetary magnetospheres (e.g., Mercury) with realistic parameters, which was previously computationally prohibitive Check out our paper on arXiv for more information: 3 LLNL-PRES-2010660 The Exascale Bottleneck: Fast Computing, Slow Data Exascale systems speed up computation so much that I/O becomes an increasingly large bottleneck (e.g., 33% of runtime for iPIC3D with 32 nodes) 4 LLNL-PRES-2010660 Closing the I/O Gap: Rack-Level I/O Acceleration with Rabbit Storage The Rabbit storage system is a rack-level, software-defined I/O accelerator engineered to bridge the gap between compute and storage performance on exascale systems SSD SSD SSD SSD GPFS Lustre SSD SSD SSD SSD Compute Nodes Rabbit Nodes Global Storage Sierra El Capitan/Tuolumne 5 LLNL-PRES-2010660 Closing the I/O Gap: Burst Buffer-Style Storage with Rabbit XFS Rabbit XFS provides highly scalable node-local storage, similar to burst buffers on systems like Sierra and Frontier Compute Nodes Rabbit Nodes Global Storage Lustre SSD SSD SSD SSD Rabbit XFS 6 LLNL-PRES-2010660 Closing the I/O Gap: Fast Shared Storage with Rabbit Lustre Rabbit Lustre provides fast, easy-to-use shared storage supporting both collective and noncollective I/O across processes Compute Nodes Rabbit Nodes Global Storage Lustre SSD SSD SSD SSD Rabbit XFS Lustre SSD SSD SSD SSD Rabbit Lustre Lustre 7 LLNL-PRES-2010660 Closing the I/O Gap: Hardware Locality for Storage Rabbits accelerate I/O by moving both local and shared storage close to compute nodes SSD SSD PCIe Switch PCIe Switch AMD EPYC CPU NICs SSD SSD SSD SSD SSD SSD SSD SSD SSD SSD SSD SSD SSD SSD SSD SSD PCIe PCIe Compute Node AMD MI300A APU PCIe PCIe NICs Compute Node PCIe PCIe AMD MI300A APU NICs Hardware Configuration Rabbit Node 8 LLNL-PRES-2010660 Closing the I/O Gap: Dynamic Software Provisioning of Storage Rabbits accelerate I/O by providing dynamic storage provisioning through a cloud-like software stack PCIe Switch PCIe Switch AMD EPYC CPU NICs SSD SSD SSD SSD SSD SSD SSD SSD SSD SSD SSD SSD SSD SSD SSD SSD PCIe PCIe Compute Node AMD MI300A APU PCIe PCIe NICs Compute Node PCIe PCIe AMD MI300A APU NICs Hardware Configuration Rabbit Node NVMe Devices Rabbit Node Software Stack Switchtec and NVMe Drivers XFS Lustre GFS2 Kernel TOSS Kubernetes (Kublet + CRIO) Config Management NVMe Control Containers Calico Overlay Network Lustre CSI Core Kubernetes Services User Containers Orchestration Services Data Movement Services 9 LLNL-PRES-2010660 Optimizing applications for Rabbit is challenging due to diverse I/O behaviors Our Solution: We map iPIC3D’s I/O phases to Rabbit configurations for phase-specific optimization: ▪Analyze iPIC3D’s I/O to identify dominant phases and quantify performance ▪Identify optimal storage backends for iPIC3D’s I/O phases by converting the characterization of each phase into proxy configurations of IOR ▪Reconfigure iPIC3D using the optimal storage backends to optimize iPIC3D’s I/O Our Contributions 16 LLNL-PRES-2010660 Understanding the I/O Phases of iPIC3D Restart Data Field Data Moment Data Number of Files 128 1 1 Processes per File 1 128 128 Total I/O per File (MB) 115074 128 64 Transfer Size (MB) 498.2 ± 0.2 1.0 ± 0.5 0.5 ± 0.0 Percent I/O Time 93.564% 0.007% 0.006% I/O Bandwidth (MB/s) 2029.4 27.4 26.4 We use Caliper and Perfetto to examine the software layers of iPIC3D and identify 3 phases of I/O writes We use DFTracer and DFAnalyzer to characterize each phase There are 3 I/O phases in iPIC3D: •Restart Data Phase: large, sequential ADIOS2 writes with 1 file per process •Field Data Phase: small, collective MPI-IO writes to 1 shared file •Moment Data Phase: small, collective MPI-IO writes with tiny transfer sizes to 1 shared file All runs of iPIC3D were performed on LLNL’s Tuolumne supercomputer 17 LLNL-PRES-2010660 Understanding the I/O Phases of iPIC3D Restart Data Field Data Moment Data Number of Files 128 1 1 Processes per File 1 128 128 Total I/O per File (MB) 115074 128 64 Transfer Size (MB) 498.2 ± 0.2 1.0 ± 0.5 0.5 ± 0.0 Percent I/O Time 93.564% 0.007% 0.006% I/O Bandwidth (MB/s) 2029.4 27.4 26.4 We use Caliper and Perfetto to examine the software layers of iPIC3D and identify 3 phases of I/O writes We use DFTracer and DFAnalyzer to characterize each phase There are 3 I/O phases in iPIC3D: •Restart Data Phase: large, sequential ADIOS2 writes with 1 file per process •Field Data Phase: small, collective MPI-IO writes to 1 shared file •Moment Data Phase: small, collective MPI-IO writes with tiny transfer sizes to 1 shared file All runs of iPIC3D were performed on LLNL’s Tuolumne supercomputer 18 LLNL-PRES-2010660 Understanding the I/O Phases of iPIC3D Restart Data Field Data Moment Data Number of Files 128 1 1 Processes per File 1 128 128 Total I/O per File (MB) 115074 128 64 Transfer Size (MB) 498.2 ± 0.2 1.0 ± 0.5 0.5 ± 0.0 Percent I/O Time 93.564% 0.007% 0.006% I/O Bandwidth (MB/s) 2029.4 27.4 26.4 We use Caliper and Perfetto to examine the software layers of iPIC3D and identify 3 phases of I/O writes We use DFTracer and DFAnalyzer to characterize each phase There are 3 I/O phases in iPIC3D: •Restart Data Phase: large, sequential ADIOS2 writes with 1 file per process •Field Data Phase: small, collective MPI-IO writes to 1 shared file •Moment Data Phase: small, collective MPI-IO writes with tiny transfer sizes to 1 shared file All runs of iPIC3D were performed on LLNL’s Tuolumne supercomputer 19 LLNL-PRES-2010660 Optimizing applications for Rabbit is challenging due to diverse I/O behaviors Our Solution: We map iPIC3D’s I/O phases to Rabbit configurations for phase-specific optimization: ▪Analyze iPIC3D’s I/O to identify dominant phases and quantify performance ▪Identify optimal storage backends for iPIC3D’s I/O phases by converting the characterization of each phase into proxy configurations of IOR ▪Reconfigure iPIC3D using the optimal storage backends to optimize iPIC3D’s I/O Our Contributions 20 LLNL-PRES-2010660 Converting Per-Phase I/O Characterization into Proxies Restart Data Field Data Moment Data Number of Files 128 1 1 Processes per File 1 128 128 Total I/O per File (MB) 115074 128 64 Transfer Size (MB) 498.2 ± 0.2 1.0 ± 0.5 0.5 ± 0.0 Percent I/O Time 93.564% 0.007% 0.006% I/O Bandwidth (MB/s) 2029.4 27.4 26.4 ior –b=16384m –t=512m –i 10 -F –w -m ior –b=16m –t=1m –i10 –a MPIIO –c –w -m ior –b=4m –t=512k –i10 –a MPIIO –c –w -m IOR Parameters for Restart Data Block Size (MiB) 16384 Transfer Size (MiB) 512 Number of Iterations 10 File-per-Process? Yes Collective I/O? No API POSIX IOR Parameters for Field Data Block Size (MiB) 16 Transfer Size (MiB) 1 Number of Iterations 10 File-per-Process? No Collective I/O? Yes API MPI-IO IOR Parameters for Moment Block Size (MiB) 4 Transfer Size (MiB) 0.5 Number of Iterations 10 File-per-Process? No Collective I/O? Yes API MPI-IO We study each of iPIC3D’s I/O phases independently by converting each phase’s characterization into a proxy configuration of the IOR benchmark 21 LLNL-PRES-2010660 Identifying Optimal Storage Backends for I/O Phases using IOR Restart Data Based on our IOR results, we observe: •Rabbit XFS is best suited for the file-per-process restart data. •Rabbit Lustre delivers optimal performance for small MPI-IO collective writes used in field data. •Rabbit Lustre performs best for MPI-IO collective writes with tiny transfer sizes used in moment data. All runs of iPIC3D were performed on LLNL’s Tuolumne supercomputer 22 LLNL-PRES-2010660 Identifying Optimal Storage Backends for I/O Phases using IOR Restart Data Field Data Based on our IOR results, we observe: •Rabbit XFS is best suited for the file-per-process restart data. •Rabbit Lustre delivers optimal performance for small MPI-IO collective writes used in field data. •Rabbit Lustre performs best for MPI-IO collective writes with tiny transfer sizes used in moment data. All runs of iPIC3D were performed on LLNL’s Tuolumne supercomputer 23 LLNL-PRES-2010660 Identifying Optimal Storage Backends for I/O Phases using IOR Restart Data Based on our IOR results, we observe: •Rabbit XFS is best suited for the file-per-process restart data. •Rabbit Lustre delivers optimal performance for small MPI-IO collective writes used in field data. •Rabbit Lustre performs best for MPI-IO collective writes with tiny transfer sizes used in moment data. All runs of iPIC3D were performed on LLNL’s Tuolumne supercomputer Field Data Moment Data 24 LLNL-PRES-2010660 Optimizing applications for Rabbit is challenging due to diverse I/O behaviors Our Solution: We map iPIC3D’s I/O phases to Rabbit configurations for phase-specific optimization: ▪Analyze iPIC3D’s I/O to identify dominant phases and quantify performance ▪Identify optimal storage backends for iPIC3D’s I/O phases by converting the characterization of each phase into proxy configurations of IOR ▪Reconfigure iPIC3D using the optimal storage backends to optimize iPIC3D’s I/O Our Contributions 25 LLNL-PRES-2010660 Reconfiguring iPIC3D Using Optimal Storage Backends from IOR Restart Data Field Data Moment Data IOR Restart Data Field Data Moment Data Rabbit XFS Rabbit Lustre Rabbit Lustre iPIC3D We reconfigure iPIC3D to use the optimal storage backends identified with IOR