PASC 2025 - Shamrock: Scaling SPH to Exascale for Astrophysics Using SYCL
Full text
Scaling SPH to Exascale for astrophysics using SYCL 1 No. 101053020 (Dust2Planets) Github public repo PASC 2025: SYCL minisymposium Timothée David--Cléris (Univ. Grenoble Alpes, CNRS, IPAG)
H. Barreiro, et. al 2 Computational Fluid Dynamics T. Skřivan, et. al A. Chern, et. al •Small scale range •Small density range “Normal” world “Astrophysical” world •Large scale range •Large density range •Many physical effects 10 orders of magnitude Multi-scale ALMA 2014 Wafflard-Fernandez , Lesur 2023 Planets Multi-physics Johansen, Youdin 2007 Vaytet et al. 2018 DUST MHD Lebreuilly et al. 2024 Self-gravity
3 Why going exascale ? Longer integration More Statistics Larger Environment More Resolution
4 Exascale supercomputers Top 500 list : GPU GPU GPU GPU GPU Supercomputers move to GPUs ⇒ (Efficiency, density, AI, …) (But all vendors involved) Could also happen with NPUs Need Portability / Performance portability ⇒
5 Portability/performance Our choice : •C++17 extension •Programming standard (multiple implementations) •Directly compiled through CUDA, Hip, OpenMP, … •Supports all major CPU & GPUs (AMD, Nvidia, Intel, ARM)
6 Use in SHAMROCK •USM allocations •Use out-of-orders queue Direct MPI GPU-GPU communications ⇒ Easier interop with vendors API (cublass, …) ⇒ Automatic scheduling to multiple GPU streams (concurrency/overlap) ⇒ Astrophysics Large range FP64 ⇒ ⇒
7 CFD methods in astrophysics SPH Finite Volumes Finite elements & Interaction criterion Numerical scheme Wait! Are cells just particles that don’t move? Always has been
8 The goal in Shamrock Numerical method + Exascale scalability Generic modules Optimising SPH Optimising AMR ⇔ Generic modules are coded once for all schemes Criterion γ Scheme F Reality is a bit more nuanced but that’s the goal
9 MPI SYCL Python C++ STL Pybind11 Base Comm Backend Algs Math Models SPH Godunov … Bindings •C++17 (templated) •(Multi-GPU) SYCL + MPI •Multiple Hydro solver Particles & Grid (SPH, FE, FV+AMR) •Open source CECILL 2.1 License •Public since march 2025 150.000 lignes of code ∼ Publications : •David--Cléris et al. 2025 (MNRAS), published •David--Cléris 2025 (IWOCL proceeding), published ? •PHD (defended in 2025)
16 Tree performance (A100) Tree Building Tree Traversal 103104105106107108 N 106 107 108 109 1010 particles per second full tree reduction T. Karras int range morton build morton sort 104105106107 N 106 107 108 particles per second cache performance (1 stage cache) reduction 0 reduction 2 reduction 4 reduction 6 reduction 8 104105106107 N 106 107 108 particles per second cache performance (2 stage cache) reduction 0 reduction 2 reduction 4 reduction 6 reduction 8 104105106107 N 106 107 particles per second timestep performance (1 stage cache) reduction 0 reduction 2 reduction 4 reduction 6 reduction 8 104105106107 N 106 107 particles per second timestep performance (2 stage cache) reduction 0 reduction 2 reduction 4 reduction 6 reduction 8 200M parts/second ( 5% timestep) ≃ 100M parts/second ( 10% timestep) ≃ Can recompute tree whenever needed ⇒
17 SPH 0.0 0.1 0.2 0.3 0.4 r 0.5 1.0 1.5 Ω phantom shamrock 0.0 0.1 0.2 0.3 0.4 r 0 100 200 u 0.0 0.1 0.2 0.3 0.4 r 0 2 4 vr 0.0 0.1 0.2 0.3 0.4 r 0.00 0.25 0.50 0.75 1.00 Æ Sedov-blast wave: Reproduce Phantom solver !
18 But also AMR & Zeus 0.0 0.1 0.2 0.3 0.4 r 0.5 1.0 1.5 Ω phantom shamrock 0.0 0.1 0.2 0.3 0.4 r 0 100 200 u 0.0 0.1 0.2 0.3 0.4 r 0 2 4 vr 0.0 0.1 0.2 0.3 0.4 r 0.00 0.25 0.50 0.75 1.00 Æ SPH (Phantom kind) Finite volume (RAMSES kind) Mass based refinement Finite elements (Zeus kind)
19 Performance !!!! (Measured in part/sec)
20 Single GPU/CPU perf (Nvidia) Sedov-blast wave:Sedov-blast wave: ⇒ Acceleration by a factor 7/8 CPU GPU ⇒ Latest benchmarks ≃15Mpart/sec ⇒ Same power consumption !!!
21 Multi-GPU (AMD) 10 100 1000 GPUs 107 108 109 1010 Particles / seconds 1.38e+09 5.34e+09 6.76e+09 8.73e+09 2e6 parts / GPUs 16e6 parts / GPUs 32e6 parts / GPUs 64e6 parts / GPUs Reference : Same test with Phantom SPH = 2e6 part / seconds x4500 speedup ! → •65G particles ( ) •9G part / seconds •7 sec / iterations •1024 GPU (MI250x) •92% parallel efficiency 40003 Shamrock : 40th (top500), GPU + CPU
22 But it is also efficient! Large CPU ≃5000part/s/W"(or"part/J) (Using Adastra’s board power consumption)
23 But also with Ramses 10 100 1000 GPUs 107 108 109 1010 Objects / seconds Adastra 1e6 parts / GPU (SPH) Adastra 8e6 parts / GPU (SPH) Adastra 16e6 parts / GPU (SPH) Adastra 32e6 parts / GPU (SPH) Lumi-G 1283cells / GPU (AMR) Lumi-G 2563cells / GPU (AMR) 84% parallel efficiency on LUMI before optimisation Next: Optimizing the solver to get at least a factor of 10
24 But this isn’t exascale yet … 24
25 Top 500 list : We have access to this one 120 000 GPUs 1 Exaflop On the road to Exascale
32 Vertical resolution 10M particles 100M particles 1G particles 1G particles : 125 particles / H ⇒ Enough for VSI, SI, parametric instability, …
33 Conclusion SPH solver (production ready) : •Exascale disc simulations •Live MCMC radiative transfert •Hydro ✅ •Discs ✅ •MHD ✅ (Y. Lapeyre) •Black hole + warp ✅ •Self-gravity (WIP) Next : AMR solver (WIP) : •Hydro ✅ •Dust ✅ (L. Sewanou) •Self-gravity (WIP) •Portable (all GPUs & CPUs) using •Scalable •Multi-physics •Multi-methods
34 Shamrock Thanks !