scieee Open visual document viewer

D2.5 - Technical report on performance-portable electrostatics

Amaral Coutinho Bartolomeu, Rodrigo; Sutmann, Godehard

Abstract

Review of the development of performance-portable electrostatic solvers developed as a community library.

Full text

Technical epo on pe o mance-po able elec os a ics Mul iXscale Deli e able 2.5 Deli e able Type: Repo Deli e ed in June, 2025 Mul iXscale Eu oHPC Cen e o Excellence o Mul iscale Modelling Acknowledgemen Funded by he Eu opean Union. This wo k has ecei ed unding om he Eu opean High Pe o mance Compu ing Join Unde aking (JU) unde g an ag eemen No 101093169. Disclaime Funded by he Eu opean Union. Views and opinions exp essed a e howe e hose o he au ho (s) only and do no necessa ily e lec hose o he Eu opean Union o he Eu opean High Pe o mance Compu ing Join Unde aking (JU). Nei he he Eu opean Union no he g an ing au ho i y can be held esponsible o hem. Mul iXscale Deli e able 2.5 Page ii P ojec and Deli e able In o ma ion P ojec Ti le Mul iXscale: Eu oHPC Cen e o Excellence o Mul iscale Modelling P ojec Re . G an Ag eemen 101093169 P ojec Websi e h ps://www.mul ixscale.eu Eu oHPC P ojec O ice Ma eo Mascagni Deli e able ID D2.5 Deli e able Na u e Repo Dissemina ion Le el Public Con ac ual Da e o Deli e y P ojec Mon h 30 (30 h June, 2025) Ac ual Da e o Deli e y 27 h June, 2025 Desc ip ion o Deli e able Re iew o he de elopmen o pe o mance-po able elec os a ic sol e s de el- oped as a communi y lib a y. Documen Con ol In o ma ion Documen Ti le: Technical epo on pe o mance-po able elec os a ics ID: D2.5 Ve sion: As o June, 2025 S a us: Accep ed by S ee ing Commi ee A ailable a : h ps://www.mul ixscale.eu/deli e ables Documen his o y: In e nal P ojec Managemen Link Re iew Re iew S a us: Re iewed Au ho ship W i en by: Rod igo Ba olomeu (FZJ) Con ibu o s: Godeha d Su mann (FZJ) Re iewed by: Alan Ó Cais (UB) App o ed by: Alan Ó Cais (UB) Documen Keywo ds Keywo ds: Mul iXscale, HPC, Pe o mance Po abili y, Elec os a ic, Load Balance, ALL 27 h June, 2025 Disclaime : This deli e able has been p epa ed by he esponsible Wo k Package o he P ojec in acco dance wi h he Conso ium Ag eemen and he G an Ag eemen . I solely e lec s he opinion o he pa ies o such ag eemen s on a collec i e basis in he con ex o he P ojec and o he ex en o eseen in such ag eemen s. Copy igh no ices: This deli e able was co-o dina ed by Rod igo Ba olomeu1(FZJ) on behal o he Mul iXscale conso - ium wi h con ibu ions om Godeha d Su mann (FZJ). This wo k is licensed unde he C ea i e Commons A ibu ion 4.0 In e na ional License. To iew a copy o his license, isi : h p://c ea i ecommons.o g/licenses/by/4.0 cb 1 [email p o ec ed] Mul iXscale Deli e able 2.5 Page iii Con en s Execu i e Summa y 1 1 In oduc ion 2 1.1 Scope o he deli e able ................................................ 2 1.2 Ta ge audience . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.3 Repo ou line ...................................................... 2 1.4 Pa ne con ibu ions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 2 Pe o mance Po abili y 3 3 Elec os a ics Me hods and Implemen a ion Conside a ions 7 3.1 Ewald Summa ion Me hod .............................................. 7 3.2 Pa icle-Pa icle Pa icle-Mesh Me hod . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 3.3 Implemen a ion Conside a ions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 4 Elec os a ic Sol e Lib a y 10 4.1 Lib a y In e ace . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 4.2 MPI Communica ion managemen . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 4.3 Benchma ks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 5 In eg a ion o he ALL load balancing lib a y in o LAMMPS 14 5.1 ALL in e ace wi h LAMMPS . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 5.2 Gene al LAMMPS Benchma ks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 5.3 Benchma k wi h balancing me hods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 6 Conclusion and Pe spec i es 18 Acknowledgemen s 19 Re e ences 19 Lis o Figu es 1 Sequence o ope a ions and p og am low wi hin he N-body p oblem. The main phase o he p og am is epea ed o N i e a ions, each compu ing a disc e e ime s ep o size d . P.I.: pa icle ini ializa ion; F.C.: o ce compu a ion; ∆ : p opaga ion o posi ions in Veloci y Ve le ; ∆ : p opaga ion o hal s ep eloci ies in Veloci y Ve le in eg a o . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 2 In e ac ions pe second s numbe o pa icles Non he in es iga ed a chi ec u es by he di e en im- plemen a ions. No e he di e en scale o he y-axis o GH200 (b). We do no include he esul s o OpenACC on he MI 250 since hey a e abou a ac o o 100 lowe han hose om he o he p og am- ming models and nea ly independen o sys em size. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 3 Subdi ision o he global MPI communica o in o sub-communica o s o be used o he compu a ion o he eal( )- and k-space con ibu ions, showing he necessa y da a ans e be ween global commu- nica o and each subcommunica o . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 4 S ong scaling benchma k o he Ewald summa ion sol e implemen ed in he lib a y using a sys em o 16 million andomly placed pa icles in a pe iodic sys em o size 1003on Juwels Boos e (Fo schungszen- um Jülich). The esou ces a e e enly spli be ween he -space and k-space compu a ion. . . . . . . . . 12 5 S ong scaling benchma k o he Ewald summa ion sol e implemen ed in he lib a y using a sys em o 16 million andomly placed pa icles in a pe iodic sys em o size 1003on Juwels Boos e (Fo schungszen- um Jülich). Th ee o ou GPUs a e employed o he compu a ion o he k-space con ibu ion, he o he GPUs a e used o he -space con ibu ion. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 6 S ong scaling benchma k o he Ewald summa ion sol e implemen ed in he lib a y using a sys em o 32 million andomly placed pa icles in a pe iodic sys em o size 1263on Juwels Boos e (Fo schungszen- um Jülich). The esou ces a e e enly spli be ween he -space and k-space compu a ion. . . . . . . . . 13 7 Single node weak scaling o LAMMPS using JEDI and JUWELS Boos e (Fo schungzen um Jülich) o he LJ and Rhodopsin benchma ks in LAMMPS using Kokkos. Scaled LJ simula ion con aining 55M a oms, and Scaled Rhodopsin simula ion con aining 16M a oms bo h wi h 1 MPI p ocess pe GPU. No esul s a e p esen ed o Juwels Boos e wi h 1 and 2 GPUs o he Rhodopsin benchma k due o memo y limi- a ion o accommoda e he da a. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 8 S ong scaling o LAMMPS using JEDI o he LJ and Rhodopsin benchma ks in LAMMPS using Kokkos. 16 Mul iXscale Deli e able 2.5 Page i 9 Combined s ong and weak scaling o LAMMPS a he p e-exascale supe compu e JEDI (Fo schungszen- um Jülich). Solid ba s use he s agge ed global load balancing me hod om he ALL lib a y in eg a ed in o LAMMPS and ha ched ba s use he ecu si e bisec ion load balancing me hod na i e o LAMMPS. . 17 Lis o Tables 1 O e iew o implemen a ions on in es iga ed a chi ec u es. ✓indica es ha benchma ks we e pe - o med. ✗indica es ha he implemen a ion was no suppo ed on ha a chi ec u e a he ime o he benchma k. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 2 Fas es measu ed pe o mance da a ac oss di e en sys em sizes, measu ed in Giga-in e ac ions pe second (GIn s) on each pla o m o each p og amming model. The as es o e all measu ed pe o mance is used as e e ence alue o he applica ion e iciency (%) and is indica ed in bold. % (cmp. Eq. 5) is he weigh ed mean o % o e all a chi ec u es, assuming % =0, when benchma ks could no be pe o med. The unde lined esul is he bes achie ed alue o %. (s) indica es a ian s using sha ed memo y, while (nd) indica es he SYCL a ian using nd- anges. The las column shows he pe o mance pe o mance po abili y me ic P ¯ P ¯.................................................. 6 Mul iXscale Deli e able 2.5 Page 1 Execu i e Summa y Wi h he as changing High Pe o mance Compu ing (HPC) landscape, he need o pe o mance po able implemen- a ions has g own signi ican ly in he pas yea s. Among he TOP 50 machines on he TOP500 lis , only 10% do no ely on G aphical P ocessing Uni (GPU)s as accele a o s o i s pe o mance. Mo eo e , accele a o s om mul iple endo s hold he op spo s on he lis . The same pic u e is also p esen on Eu opean High Pe o mance Compu ing Join Unde aking (Eu oHPC) and Gauss Supe compu ing Cen e (GCS) machines. Bo h ha e supe compu e s wi h AMD and NVIDIA GPUs o di e en gene a ions, a chi ec u es, and compu e capabili ies. This s iking eali y poses a g owing demand o pe o mance po abili y o scien i ic codes, lib a ies, and applica- ions. Fo many domain applica ions, he e icien use o CPUs alone is no su icien o ex ac he a ailable pe - o mance and ene gy e iciency a ailable on mode n a chi ec u es. As an example, he wo lds mos ene gy e icien supe compu e JEDI (Eu oHPC/FZJ) is based on N idia’s G ace-Hoppe Supe chip, which combines CPU and GPU. In his sys em he GPUs can be p og ammed using Compu e Uni ied De ice A chi ec u e (CUDA), howe e doing so he code would no be able o un on o he endo ’s pla o ms. Achie ing pe o mance po abili y can be accomplished by pai ing de ices and le e aging o loading o compu e ke - nels h ough amewo ks such as OpenMP, OpenACC, Kokkos, o Raja. Ou p elimina y esea ch on a pe o mance po able Ewald summa ion me hod enligh ens he necessi y o u he compa ison be ween di e en so wa e ame- wo ks o de e mine he mos e ec i e app oach. The e o e, a pe o mance po abili y s udy was conduc ed using he N-body p oblem as a es case o e alua e he e icacy o a ious p og amming models ac oss majo GPU endo s. This s udy aimed o e alua e he ela i e s eng hs and weaknesses o each model, p o iding aluable insigh s in o he op imal app oach o achie ing pe o mance po abili y ac oss di e se ha dwa e a chi ec u es. Fo elec os a ic compu a ions based on a ian s o he Ewald summa ion me hod, which a e o en one o he mos compu e in ense and ime consuming s eps o Molecula Dynamics (MD) simula ions, no only po abili y is desi ed bu also modula i y. Task di ision can be achie ed by sepa a ing he sho - ange and long- ange con ibu ions o he elec os a ic po en ial and o ces. Since he eal-space and Fou ie k-space pa s ha e di e en pe o mance cha ac- e is ics, a mul i-domain pa i ioning app oach can be applied in conjunc ion wi h adjus ing ele an Ewald me hod pa ame e s. This allows o a balanced compu a ional load ac oss independen ly unning pa i ions. To achie e op- imal load balancing be ween he eal-space and k-space pa i ions, i is essen ial o op imize ope a ions pe o med on bo h pa s as well as he amoun o compu e esou ces alloca ed o each ask. In his deli e able, we p esen ad ances, implemen a ion de ails, a ionale, and benchma ks ela ed o he a o emen- ioned me hods. The ou comes o he pe o mance po abili y s udy, as well as he p e iously epo ed FFT bench- ma k, se e as ounda ion o he design choices included in he de elopmen o P3M me hod and he po able elec- os a ic sol e lib a y. Mul iXscale Deli e able 2.5 Page 2 1 In oduc ion 1.1 Scope o he deli e able This deli e able aims a p o iding in o ma ion on he implemen a ion o a pe o mance po able elec os a ics sol e and he in eg a ion o he ALL load balancing lib a y in o LAMMPS. This epo con ains elemen s ele an o Task 2.3 - Pe o mance-po able elec os a ics sol e s and load-balancing lead by Fo schungszen um Jülich. 1.2 Ta ge audience This epo is in ended o eade s wi h echnical expe ise in HPC o in he use o mul i scale pa icle based simula ion so wa e. While po en ial use s migh be in e es ed in o he de ails epo ed he e, he main pu pose is o epo on he esul s and implemen a ions ha migh be mo e o in e es o scien i ic so wa e de elope s willing o in e ace hei code o he po able elec os a ic sol e lib a y. 1.3 Repo ou line This deli e able epo s on all ele an opics wi h espec o Task 2.3, led by Fo schungzen um Jülich. The deli e able is di ided in o ou pa s: •Pe o mance Po abili y: in his sec ion we p esen he up- o-da e s a us on he pe o mance po abili y o he mos commonly used C++ po able p og amming models. Du ing he cou se o Task 2.3 a s udy was conduc ed o access he usabili y and pe o mance o such models. As ou come, an upda ed s a us on he po abili y was epo ed. P ac ical expe ience in he usage as well as a deepe unde s anding o he implemen a ion de ails o each model was acqui ed. This sec ion con ibu ed o he choice o he p og amming model used o he implemen a ions o elec os a ic sol e s. •Theo e ical ounda ions o Ewald sum and P3M me hods: in his sec ion we p esen he o mula ion o he ba- sic Ewald sum, which is he o iginal me hod o compu e he elec os a ic ene gy o an a omis ic sys em unde pe iodic bounda y condi ions. The accele a ed e sion, P3M, makes use o he Fas Fou ie ans o m, e alua ed on a egula mesh, ca ying he mapped pa icle cha ges. Ca dinal B-Splines a e applied o map cha ges o he mesh and o e ie e po en ial ene gies and o ces back o he pa icles. While being able o con ol he app ox- ima ion e o , he P3M me hod is educing he compu a ional complexi y om O(N3/2) down o O(Nlog N), whe e Nis he numbe o pa icles in he sys em. •Elec os a ic Sol e Lib a y: in his sec ion we discuss he design and implemen a ion o he pe o mance po able and scalable lib a y app oach o elec os a ic sol e s. In he p esen p ojec we ocus on he Ewald summa ion as a benchma k implemen a ion and he P3M as an op imized implemen a ion o p oduc ion uns on mode n HPC sys ems. Pe o mance po abili y has been conside ed o bo h he eal and ecip ocal space pa o he me hod. P ope pa ame e choice o e o con ol and aspec s o di e en scalabili y o eal and ecip ocal space pa s a e discussed and conside ed in he design o he lib a y. Fi s benchma ks on di e en a chi ec u es a e shown and p ope ies o scalabili y a e discussed in e ms o di e en asks in he me hod. •Load Balancing: wi h he in en o benchma k he load balancing me hods a ailable in he A Load balancing Lib a y (ALL) wi hin using ou key applica ion codes. We ha e in eg a ed ALL o LAMMPS h ough an in e ace ha p o ide use s he capabili y o using ex ended ix commands. In he con ex o his deli e able we p esen benchma ks compa ing LAMMPS pe o mance using A100 and GH200 GPUs. 1.4 Pa ne con ibu ions JSC con ibu ed as planned o he wo k in his deli e able which is in ended o be con ibu e o ealisa ion o o he asks in he p ojec . Mul iXscale Deli e able 2.5 Page 3 2 Pe o mance Po abili y The ongoing de elopmen o new GPUs om a ious endo s equi es con inuous moni o ing o po able ame- wo ks’ pe o mance and e sa ili y. A p ima y conce n is whe he a po able amewo k is a ailable and unc ional on a gi en a chi ec u e. While he concep o po abili y be ween di e en pla o ms is s aigh o wa d, de ining and measu ing pe o mance po abili y is mo e complex. Success ul p og am execu ion is a clea indica o o po abili y, bu e alua ing pe o - mance po abili y equi es a mo e nuanced app oach. A common me hod o assessing he pe o mance and capabili ies o po able p og amming models in ol es using mini-apps o ull applica ions o ga he ele an da a [1, 2]. This can be achie ed by po ing a single applica ion o a ious pe o mance-po able p og amming models [3–6] o by po ing mul iple mini-apps o one pe o mance- po able p og amming model and compa ing he esul s o o he a ailable e sions [7]. In ou s udy, we ocus on he N-body p oblem [8]. The N-body p oblem is a signi ican p oblem class and has been iden i ied as one o he se en o iginal dwa s in Re . [9], each ep esen ing a compelling use case wi h i s own challenges. Many scien i ic ields equi e compu ing o ces be ween pai s o in e ac ing pa icles, such as g a i a ional o ces in as ophysics, elec os a ic o ces in cha ged sys ems, and an de Waals o ces o modeling a omic sys ems. The g a i a ional po en ial be ween pa icles is gi en by: U(di j )=−Gmimj di j (1) whe e G=6.6743·10−11 m3 kg·s2is he g a i a ional cons an , di j =| i j |is he dis ance be ween pa icles iand j, i j = i− j, and miand mja e he masses o he espec i e pa icles. Gi en a po en ial U(di j ), he o ce i j be ween pa icles iand jis gi en by Fi j =−∇U(di j )=−Gmimj d3 i j i j . (2) The o al o ce ac ing on pa icle iis hen gi en by he sum o e all o he pa icles Fi=∑︂ j=i Fi j . (3) To ob ain he ajec o y o each pa icle, New on’s equa ions o mo ion x ˙i= i, ˙i=ai=Fi mi , (4) a e in eg a ed, whe e he do -no a ion ep esen s he ime de i a i e, aiis he accele a ion, and miis he mass o pa icle i. The g a i a ional N-body p oblem ( he 4 h dwa in Re . [9]) is a ep esen a i e o an algo i hmic me hod ha includes long- ange in e ac ions in open-bounda ies and in ol es summa ion o e all pa icle pai s in he sys em o compu e he o ces on each pa icle, esul ing in a na u al compu a ional complexi y o O(N2). To a oid his quad a ic com- plexi y, mo e e icien me hods like he Ba nes-Hu ee me hod o he as mul ipole me hod (FMM) can be used, e- ducing he complexi y o O(Nlog(N)) o O(N). Howe e , since hese me hods ha e a high le el o complexi y and hei pe o mance s ongly depends on he implemen a ion and op imiza ion, we ocus he e on a simple implemen a ion based on he all-pai s compu a ion o ene gies and o ces, which isola es he pe o mance ou come om algo i hmic issues. To e alua e he u iliza ion o he ha dwa e a chi ec u es in ou benchma ks, we compa ed he h oughpu o pa icle- pa icle compu a ions, whe e he numbe o in e ac ions scales as N2. The s udy used he N-body p oblem as a es om compu e in ense algo i hm ha is p esen in many MD so wa e. An illus a ion o he p og am wo k low is p esen ed in Figu e 1. In o de o e alua e he pe o mance po abili y o di e en p og amming models, he discussed algo i hm was im- plemen ed in ele en a ia ions. The C++ amewo ks conside ed we e: Kokkos [10], OpenMP, OpenACC, SYCL, HIP, CUDA, and pSTL. The implemen a ions we e es ed on ou pla o ms. An o e iew o he es ed pla o ms is p e- sen ed in Table 1. Fo de ails abou pla o m speci ica ions, compile s, implemen a ions, and op imiza ions, eade s may e e o [8]. Mul iXscale Deli e able 2.5 Page 4 ini main I/O Δ𝑣, Δ𝑟 S a P.I. F.C. End Ini . F.C. Δ𝑣 𝑁𝑡 ime s eps Figu e 1: Sequence o ope a ions and p og am low wi hin he N-body p oblem. The main phase o he p og am is epea ed o N i e a ions, each compu ing a disc e e ime s ep o size d . P.I.: pa icle ini ializa ion; F.C.: o ce compu a ion; ∆ : p opaga ion o posi ions in Veloci y Ve le ; ∆ : p opaga ion o hal s ep eloci ies in Veloci y Ve le in eg a o . A chi ec u e OpenMP pSTL Kokkos CUDA HIP SYCL OpenACC N idia A100 ✓ ✓ ✓ ✓ ✓ ✓ ✓ N idia GH200 ✓ ✓ ✓ ✓ ✓ ✓ ✓ AMD MI250 ✓ ✓ ✓ ✗ ✓ ✓ ✓ In el GPU Max 1100 ✓ ✓ ✓ ✗ ✗ ✓ ✗ Table 1: O e iew o implemen a ions on in es iga ed a chi ec u es. ✓indica es ha benchma ks we e pe o med. ✗indica es ha he implemen a ion was no suppo ed on ha a chi ec u e a he ime o he benchma k. As shown in Figu e 2, mo e powe ul GPUs equi e la ge sys em sizes o sa u a e he p ocesso , indica ed by he la ening o he lines in he plo s. The esul s show ha a ian s using sha ed memo y gene ally pe o m be e han hose wi hou i , pa icula ly on he GH200. Howe e , he e a e some excep ions ha equi e u he in es iga ion. No ably, CUDA on he GH200 wi h sha ed memo y achie ed he bes esul s, while SYCL ou pe o med HIP on he MI250. In con as , OpenMP and pSTL ou pe o med SYCL on he In el pla o m. Fu he mo e, OpenACC, OpenMP, and pSTL su p isingly achie ed highe pe o mance on he A100 ca d compa ed o CUDA wi h sha ed memo y. These indings align wi h a s udy by [11], which ound ha SYCL’s so wa e s ack imp o emen s led o ou pe o ming na i e p og amming models on compu e-bound ke nels. These esul s a e also p esen ed nume ically in Figu e 2 o discussion: Using he mean applica ion e iciency % (see Equa ion 5), i can be seen ha he models designed o be po able pe - o m be e han hose a ge ed a speci ic pla o ms (HIP, CUDA). Al hough his is o be expec ed, i is no ewo hy ha pSTL and SYCL (using sha ed memo y) show he bes po abili y quali y. Kokkos shows compa a i ely lowe pe o - mance on he AMD and In el pla o ms. Rega ding AMD, ou esul di e s om ecen s udies, whe e Kokkos showed excellen pe o mance on AMD-based sys ems [2,12]. This may be due o he a ian chosen o i s pe o mance on he GH200. %= %A100 +%GH200 2+%MI250 +%Max1100 3. (5) The espec i e esul s o P ¯ P ¯(Table 2) show li le de ia ion om he esul s o % o he po able p og amming model, bu make i clea ha CUDA, HIP, and OpenACC a e no po able since hey a e no applicable (NA) pe de ini ion o p og amming models ha canno be execu ed on one o mo e o he in es iga ed a chi ec u es. The de ia ion in he pe cen ages o he po able p og amming models comes om he di e en ways o a e aging (see. Equa- ion 5). Fo a mo e de ailed analysis o pe o mance po abili y he numbe o aken measu emen s as well as he numbe o a ge ed pla o ms would need o be inc eased. Despi e his, he pe o med analysis on he p esen ed da a al eady indica es ha using an op imized e sion o a code on a gi en a chi ec u e (in his case he GH200) migh no be op imally sui ed o o he a chi ec u es, bu s ill pe o ms sa is ac o y well on he mos mode n sys em om all majo HPC GPU endo s. Ou esul s show ha OpenMP, pSTL, and SYCL a e well sui ed o po able implemen a ions o he N-body p oblem ac oss a ious GPU endo s, while Kokkos showed sligh ly educed pe o mance on he non-N idia pla o ms. Kokkos Mul iXscale Deli e able 2.5 Page 5 CUDA CUDA (s) HIP HIP (s) Kokkos pSTL OpenMP OpenACC SYCL SYCL (s) SYCL- ange 3236439631283 50 100 150 200 𝑁 GIn / s (a) A100 3236439631283 100 200 300 400 𝑁 GIn / s (b) GH200 3236439631283 50 100 150 200 𝑁 GIn / s (c) MI 250 3236439631283 50 100 150 200 𝑁 GIn / s (d) GPU Max 1100 Figu e 2: In e ac ions pe second s numbe o pa icles Non he in es iga ed a chi ec u es by he di e en imple- men a ions. No e he di e en scale o he y-axis o GH200 (b). We do no include he esul s o OpenACC on he MI 250 since hey a e abou a ac o o 100 lowe han hose om he o he p og amming models and nea ly indepen- den o sys em size. migh show be e pe o mance, when using algo i hms mo e dependen on complex memo y access pa e ns, e.g., sho - ange in e ac ions, as in he case o Ewald Sum and he Pa icle-Pa icle Pa icle-Mesh algo i hms. Mul iXscale Deli e able 2.5 Page 12 numbe s o nodes used, while o he la ge numbe o nodes used he -space un ime becomes dominan . O e all a pa allel e iciency o ≈35% was eached on 128 GPUs (32 nodes). As a consequence o he inding a la ge alloca ion o esou ces o he k-space pa migh be ob ious o imp o e he o al un- ime. The e o e, he benchma k p esen ed in Figu e 5used h ee o ou GPUs pe node o he compu a ion o he k-space con ibu ion and one GPU o each node o he -space con ibu ion. Howe e , i can be seen ha he o e all pa allel e iciency decays as e han in he i s benchma k. This is due o he ac ha he -space compu- a ion los pe o mance and domina es he o e all pa allel e iciency. The imp o ed scaling beha io o he k-space compu a ion can no be u ilized, as he -space compu a ion always akes longe o execu e. In a hi d benchma k a la ge sys em was used o in es iga e he impac o sys em size on s ong scalabili y. The esul s a e shown in Figu e 6. In his benchma k he esou ces we e e enly spli be ween bo h pa s, - and k-space con ibu ions. I can be seen ha an e iciency o ≈60% was achie ed. Fo all benchma ks he di e en compu a ion pa s a e execu ed on sub-communica o s which do no sha e p o- cesses. This means ha hey un independen ly in pa allel bu wi h di e en indi idual pa allel e iciencies. As a esul he sum o bo h pa s is la ge han he o al un ime. Fi s analysis indica es ha his app oach o spli ing compu e esou ces can be bene icial o node coun s whe e he o e all pa allel e iciency s a s o deg ade. In o de o be able o be e guide he communi y in ha espec u he in-dep h analysis will be pe o med. Figu e 4: S ong scaling benchma k o he Ewald summa ion sol e implemen ed in he lib a y using a sys em o 16 million andomly placed pa icles in a pe iodic sys em o size 1003on Juwels Boos e (Fo schungszen um Jülich). The esou ces a e e enly spli be ween he -space and k-space compu a ion. Mul iXscale Deli e able 2.5 Page 13 Figu e 5: S ong scaling benchma k o he Ewald summa ion sol e implemen ed in he lib a y using a sys em o 16 million andomly placed pa icles in a pe iodic sys em o size 1003on Juwels Boos e (Fo schungszen um Jülich). Th ee o ou GPUs a e employed o he compu a ion o he k-space con ibu ion, he o he GPUs a e used o he -space con ibu ion. Figu e 6: S ong scaling benchma k o he Ewald summa ion sol e implemen ed in he lib a y using a sys em o 32 million andomly placed pa icles in a pe iodic sys em o size 1263on Juwels Boos e (Fo schungszen um Jülich). The esou ces a e e enly spli be ween he -space and k-space compu a ion. Mul iXscale Deli e able 2.5 Page 14 5 In eg a ion o he ALL load balancing lib a y in o LAMMPS The A Load balancing Lib a y (ALL) is a modula , open-sou ce so wa e de eloped by he Jülich Supe compu ing Cen e o add ess he c i ical challenge o load imbalance in domain-decomposed simula ions. In high-pe o mance compu ing (HPC), pa icula ly in pa icle-based me hods like molecula dynamics o la ice-Bol zmann simula ions, spa ial domain decomposi ion is a common s a egy o dis ibu e compu a ional wo kload. Howe e , as simula ions e ol e, non-uni o m dis ibu ions o pa icles o a ying local compu a ional cos s can lead o signi ican load imbal- ances, educing o e all e iciency. ALL p o ides a sui e o dynamic load-balancing s a egies speci ically designed o edis ibu e wo k ac oss p ocesso s in such simula ions, ensu ing be e u iliza ion o HPC esou ces. By suppo ing mul iple balancing schemes such as his og am and i e a i e g id pa i ioning ALL o e s lexible and scalable solu ions ailo ed o di e en applica ion needs. I s in eg a ion in o widely-used simula ion codes like HemeLB and DL_MESO_DPD highligh s i s e ec i eness in imp o ing pe o mance on la ge-scale sys ems, including GPU ac- cele a ed pla o ms. Wi hin he scope o he cu en Deli e able, we highligh he mos e icien and used me hods, which a e p esen in he lib a y o being al eady implemen ed in LAMMPS: Tenso Decomposi ion (Local) This me hod spli s he simula ion domain along Ca esian axes using a ixed numbe o pa i ions in each di ec ion (e.g., a g id layou ). I is an i e a i e me hod ha , o e e y dimension, educes he wo k o e he o he dimensions and shi s he bo de s be ween wo neighbou s based on hei cumula i e wo k, making i simple bu less adap i e o highly une en load dis ibu ions. •Classic Fo he classic app oach he wo k pe dimension is collec ed using he a e age o e he domains in ha dimension. •Max The max a ian conside s he maximum wo k in each o he slices in ca esian di ec ion and conside s as objec i e unc ion an e en dis ibu ion o he maxima in he slices o each dimension. S agge ed G id The s agge ed g id is a special domain decomposi ion scheme ha ecu si ely cu s in each dimension in a de ined o de . I is a special case o he iled decomposi ion as used o ecu si e bisec ion and is, in p inciple, able o p oduce a pe ec decomposi ion o e en load in he sys em. •Global (Global + His og am) This me hod di ides he sys em ac oss domains in o g id egions and adjus s he bounda ies using global wo kload his og ams. I op imizes he load balance by analyzing he dis ibu ion o compu a ional wo k ac oss he en i e domain, hen shi s domain bounda ies acco dingly. I is a global me hod and uses his og am-based e inemen o edis ibu e wo k mo e e enly. •Local (Local + I e a i e) Unlike he global a ian , his me hod balances wo kloads based on local neighbo - o- neighbo in e ac ions. I i e a i ely adjus s bounda ies by exchanging wo kload da a wi h neighbo ing domains, g adually imp o ing balance h ough a se ies o local upda es. This local me hod uses i e a ion-based e ine- men , making i scalable and esponsi e o localized imbalances. Recu si e Bisec ion (Global + I e a i e) This is a global and hie a chical me hod ha ecu si ely spli s he domain in o wo sub egions a each s ep, aiming o e enly dis ibu e he o al wo kload. The spli ing p ocess can be guided by ei he wo kload his og ams o i e a i e e alua ion o load imbalance. I is e ec i e o complex o non-uni o m load dis ibu ions bu can in oduce communica ion o e head i applied oo equen ly. 5.1 ALL in e ace wi h LAMMPS ALL balancing me hods ha e been in eg a ed in o LAMMPS using he s anda d de ini ions o so-called ixes. Once LAMMPS is build wi h -D PKG_ALLPKG=yes, he ollowing op ions become a ailable: ix ID g oup-ID balance/all keywo d a gs ... equi ed keywo d/a g pai s e e y a g = ne e y ne e y = pe o m dynamic load balancing e e y his many s eps g id s yle a gs = de ine g id s yle = s agge ed o enso s agge ed a gs = none enso a gs = max o classic max = use TENSOR_MAX me hod o ALL classic = use TENSOR me hod o ALL Mul iXscale Deli e able 2.5 Page 15 weigh s yle a gs = use weigh ed pa icle coun s o he balancing s yle = g oup o neigh o ime o a o s o e g oup a gs = Ng oup g oup1 weigh 1 g oup2 weigh 2 ... Ng oup = numbe o g oups wi h assigned weigh s g oup1, g oup2, ... = g oup IDs weigh 1, weigh 2, ... = co esponding weigh ac o s neigh ac o = compu e weigh based on numbe o neighbo s ac o = scaling ac o (> 0) ime ac o = compu e weigh based on ime spend compu ing ac o = scaling ac o (> 0) a name = ake weigh om a om-s yle a iable name = name o he a om-s yle a iable s o e name = s o e weigh in cus om a om p ope y de ined by ix p ope y/a om command name = a om p ope y name (wi hou d_ p e ix) A leas one o he ollowing keywo d/a g pai s is equi ed. Only one (o none) load balance is used in a load balancing s ep. global a gs = igge _ h eshold bin_wid h igge _ h eshold = loa loa = apply load balance i he imbalance o he load is abo e his h eshold bin_wid h = app oxima e wid h o bin (a bi a y posi i e alue) local a gs = igge _ h eshold (use local balancing) igge _ h eshold = loa loa = apply load balance i he imbalance o he load is abo e his h eshold op ional keywo d a g pai s e bose a gs = none ( e bose ou pu o log/sc een) ile a g = ilename ilename = w i e s a s pe ank o ilename In addi ion o he ix balance/all me hod, a specialized communica ion s yle was added in o de o co ec ly and e icien ly manage he MPI communica o s se by he domain decomposi ion. Some examples o he a ailable bal- ancing me hods a e: # his og am me hod (s agge ed global), weigh ed by numbe o a oms, # wi h bin wid h 0.003, when imbalance >= 1.05 comm_s yle s agge ed zyx ix 5 all balance/all e e y 1000 g id s agge ed global 1.05 0.003 # his og am me hod (s agge ed global), weigh ed by ime, wi h bin wid h 0.003, when imbalance >= 1.05 comm_s yle s agge ed zyx ix 5 all balance/all e e y 1000 g id s agge ed weigh ime 1.0 global 1.05 0.003 # s agge ed local me hod, weigh ed by numbe o a oms, when imbalance >= 1.05 comm_s yle s agge ed zyx ix 5 all balance/all e e y 100 g id s agge ed local 1.05 # enso me hod, minimizing a e age wo k, weigh ed by numbe o a oms, when imbalance >= 1.05 comm_s yle b ick ix 5 all balance/all e e y 50 g id enso classic local 1.05 # enso -max me hod, minimizing maximal wo k, weigh ed by numbe o a oms, when imbalance >= 1.05 comm_s yle b ick ix 5 all balance/all e e y 50 g id enso max local 1.05 Mul iXscale Deli e able 2.5 Page 16 5.2 Gene al LAMMPS Benchma ks Wi h he in en ion o es LAMMPS pe o mance on he mos ecen a chi ec u e p esen a he Eu oHPC la es supe - compu e JUPITER, we ha e un se e al benchma ks o LAMMPS wi h s anda d sys ems. The es consis ed in weak and s ong scaling on JEDI and JUWELS Boos e sys ems bo h a ailable a Fo schungszen um Jülich2. Fo he alida ion benchma ks, LAMMPS was compiled using GCC 12.3.0, CUDA 12.2, and OpenMPI 4.1.6, and he chosen FFT backend was cuFFT. In Figu es 7and 8we p esen he esul s om po ing LAMMPS o he newe a chi- ec u es. (a) LJ (b) Rhodopsin Figu e 7: Single node weak scaling o LAMMPS using JEDI and JUWELS Boos e (Fo schungzen um Jülich) o he LJ and Rhodopsin benchma ks in LAMMPS using Kokkos. Scaled LJ simula ion con aining 55M a oms, and Scaled Rhodopsin simula ion con aining 16M a oms bo h wi h 1 MPI p ocess pe GPU. No esul s a e p esen ed o Juwels Boos e wi h 1 and 2 GPUs o he Rhodopsin benchma k due o memo y limi a ion o accommoda e he da a. As i can be obse ed in Figu e 7bo h sys ems and a chi ec u e combina ions show close o ideal scaling wi hin single node con igu a ions. The expec ed speed up in he newe a chi ec u e (2x) is achie ed and he la ge memo y allows o he simula ion o bigge sys em sizes wi h less compu ing esou ces. In Figu e 7b one can obse e ha due o memo y and communica ion cha ac e is ics, i is possible o achie e simila esul s o he Rhodopsin benchma k using one H100 o he GH200 node ins ead o h ee A100 GPUs. (a) LJ (b) Rhodopsin Figu e 8: S ong scaling o LAMMPS using JEDI o he LJ and Rhodopsin benchma ks in LAMMPS using Kokkos. S ong scaling esul s p esen in he g ouped ba s o Figu e 8show ha o he LJ benchma k he minimum sys em size equi ed o p o i om alloca ing mul i-node jobs was 55.4 million pa icles. Once aking in o conside a ion he mo e complex sys em, i.e., he Rhodopsin benchma k, one may expec 16.5 million pa icles o be su icien a he es ed ange, which seems o esul om he mo e complex in e ac ions. In ac , he Rhodopsin benchma k includes all elec os a ic o ces, bonds, dihed als and imp ope con ibu ions which esul s in a much la ge compu a ional wo k compa ed o he simple LJ case. 2The JEDI sys em has now been decommissioned and has been in eg a ed in o he JUPITER sys em. Mul iXscale Deli e able 2.5 Page 17 Figu e 9: Combined s ong and weak scaling o LAMMPS a he p e-exascale supe compu e JEDI (Fo schungszen- um Jülich). Solid ba s use he s agge ed global load balancing me hod om he ALL lib a y in eg a ed in o LAMMPS and ha ched ba s use he ecu si e bisec ion load balancing me hod na i e o LAMMPS. 5.3 Benchma k wi h balancing me hods Fo es ing he pe o mance o he load balancing lib a y, a syn he ic benchma k was gene a ed emo ing a oms om a sphe ical egion a he cen e o he LJ s anda d benchma k. The sys em was scaled in he same manne as he ones p esen ed in he p e ious sec ion. In Figu e 9 esul s show compa isons be ween load balancing me hods, i.e., RCB, which is one o LAMMPS’s na i e benchma k me hods, and he s agge ed global me hods (ALL) using he same in e al o checking he load imbalance and ac i a ing he load balancing p ocedu e wi h a common c i e ion. The balancing commands added o he inpu sc ip we e: comm_s yle s agge ed zyx ix 5 all balance/all e e y 50 g id s agge ed global 0.9 0.003 in he case o using he me hod a ailable h ough he in eg a ion wi h ALL, and : comm_s yle iled ix 5 all balance 50 cb 0.9 o LAMMPS’s ecu si e bisec ion algo i hm (RCB). I can be obse ed ha o all numbe s o nodes he s agge ed/global me hod (solid ba s) esul ed in a pe o mance imp o emen o 30% o 50% compa ed o RCB (ha ched ba s), which shows a p omising capabili y o enhancing he pe o mance o load imbalanced simula ions on exascale supe compu e s. These simula ions we e pe o med using he GPU package in LAMMPS h ough he use o -s gpu command line. Mul iXscale Deli e able 2.5 Page 18 6 Conclusion and Pe spec i es This deli e able p esen s he de elopmen s achie ed h ough he e o s o Task 2.3. Sec ions 2 o 5co e pe o - mance po abili y, elec os a ic me hods and implemen a ion conside a ions, lib a y API and benchma ks, and he in eg a ion o he ALL load balancing lib a y in o LAMMPS. These indi idual opics a e in e connec ed, wi h po abil- i y aspec s in luencing he design choices o elec os a ic sol e s and hei impac on a ge sys ems o op imal load balancing me hods ha in e ace wi h pa icle simula ions. Kokkos has p o en o be a eliable choice o po ing codes o newe a chi ec u es, o e ing signi ican ad an ages in complex da a s uc u es and widesp ead communi y adop ion. Al hough i may no yield he bes o e all P ¯ P ¯ alues, i consis en ly pe o ms well on sys ems wi h a obus backend so wa e s ack. As demons a ed in Sec ion 2, Kokkos ma ches he pe o mance o na i e/ endo - ecommended pla o m-dependen solu ions wi hou equi ing labo ious compile op imiza ions. Howe e , i is essen ial o no e ha pla o m- ela ed compile op imiza ions may posi i ely impac Kokkos’ pe o mance, and sys em main aine s a e encou aged o explo e his opic. Sec ion 3p esen s algo i hmic conside a ions aimed a ex ac ing he desi ed pe o mance om mode n exascale a - chi ec u es. Each independen s ep can u ilize speci ic sub-communica o s and un in pa allel on CPU+GPU combi- na ions o ac oss modula pa i ions. Bo h Ewald summa ion and P3M me hods employ ei he pa icle o mesh-based pa allelism h ough Cabana me hods o Kokkos Mul i-Dimensional Range loops. The en a i e API o he lib a y and p elimina y benchma ks a e p esen ed in Sec ion 4. These benchma ks alida e and highligh he lib a y’s modula capabili y, demons a ing he spli ing o and k-space me hods and hei espec- i e scalabili ies o he Ewald summa ion me hod (Figu es 4 o 6). These benchma ks we e pe o med by using wo independen alloca ions o he - and k-space pa s on Juwels Boos e (Fo schungszen um Jülich). The discussion includes a c i ical assessmen o he equi emen s and capabili ies o a balancing me hod o un - and k-space pa s mos e icien ly. S ong scaling o 16 and 32 million pa icles scaled om 4 o 128 GPUs e eals ha pa allel e i- ciency is a ec ed by he manne in which compu a ional e o o - and k-space pa s is dis ibu ed on o esou ces. Ini ial analysis sugges s ha his app oach o esou ce alloca ion can be bene icial o node coun s whe e o e all pa allel e iciency begins o deg ade. Fu he in-dep h analysis will be conduc ed o p o ide mo e comp ehensi e guidance. The P3M implemen a ion is cu en ly unde going alida ion, and i s ex ac ion o he new elec os a ic lib a y has commenced. We plan o analyze he impac o modula alloca ion wi h mesh me hods using he same me hodology p esen ed o he Ewald Summa ion. Bo h me hods will be op imized o unning on Eu ope’s la ges supe compu e , JUPITER. Las ly, he LAMMPS po ing o he newes a chi ec u e has exhibi ed expec ed weak and s ong scaling beha io . Pe - o mance on GH200 nodes is expec ed o be wice as as as on he p e ious N idia GPU a chi ec u e (A100). Fo e y la ge simula ions, ene gy e iciency imp o es mo e han wo old, as ewe nodes and GPUs a e equi ed o pe o m equally sized simula ions. Following he analysis o LAMMPS scaling beha io on bo h JEDI and Juwels Boos e , p elimina y esul s using he in eg a ion o ALL in o LAMMPS a e p esen ed in Sec ion 5.3. As he pe o mance o load balancing depends on ini ial con igu a ion and compu ing cos s (desc ibed in Sec ion 5.2), a mo e comp ehensi e s udy will be conduc ed o e alua e he di e ences be ween LAMMPS’ implemen ed balance me hods and hose added by he in eg a ion wi h he ALL lib a y. The s anda diza ion o he in e ace is nea o comple ion and will be made a ailable o he communi y by c ea ing a pull eques o LAMMPS’ main eposi o y. Mul iXscale Deli e able 2.5 Page 19 Acknowledgemen s We hank he Jülich Supe compu ing Cen e, Fo schungszen um Jülich o access o he JURECA-DC E alua ion Pla - o m, JUWELS Boos e and JEDI. This wo k has ecei ed unding om he Eu opean High Pe o mance Compu ing Join Unde aking unde he g an ag eemen 101093169 and was suppo ed by he Ge man Fede al Minis y o Edu- ca ion and Resea ch (BMBF) and he Minis y o Cul u e and Science (MKW) o he s a e o No h-Rhine-Wes phalia h ough unding o he Gauss Cen e o Supe compu ing (GCS). Re e ences Ac onyms used HPC High Pe o mance Compu ing FFT as Fou ie ans o m Eu oHPC Eu opean High Pe o mance Compu ing Join Unde aking GPU G aphical P ocessing Uni MD Molecula Dynamics GCS Gauss Supe compu ing Cen e API Applica ion P og am In e ace So wa e men ioned CUDA Compu e Uni ied De ice A chi ec u e heFFTe highly e icien Fas Fou ie T ans o m o exascale ALL A Load balancing Lib a y ScaFaCoS Scalable Fas Coulomb Sol e s URLs e e enced Page ii h ps://www.mul ixscale.eu ... h ps://www.mul ixscale.eu h ps://www.mul ixscale.eu/deli e ables ... h ps://www.mul ixscale.eu/deli e ables In e nal P ojec Managemen Link ... h ps://gi hub.com/mul ixscale/planning/ .ba olomeu@ z-juelich.de ... mail o: .ba olomeu@ z-juelich.de h p://c ea i ecommons.o g/licenses/by/4.0 ... h p://c ea i ecommons.o g/licenses/by/4.0 Page 1 TOP500 ... h ps://www. op500.o g Page 7 ScaFaCoS ... h p://www.sca acos.de/ Page 9 Cabana ... h ps://gi hub.com/ECP-copa/Cabana #799 ... h ps://gi hub.com/ECP-copa/Cabana/pull/799 #800 ... h ps://gi hub.com/ECP-copa/Cabana/pull/800 Page 18 JUPITER ... h ps://www. op500.o g/lis s/ op500/2025/06/ Ci a ions [1] T. Deakin, S. McIn osh-Smi h, J. P ice, A. Poena u, P. A kinson, C. Popa, and J. Salmon, “Pe o mance po abili y ac oss di e se compu e a chi ec u es,” in 2019 IEEE/ACM In e na ional Wo kshop on Pe o mance, Po abili y and P oduc i i y in HPC (P3HPC). IEEE, 2019, pp. 1–13. [2] J. H. Da is, P. Si a aman, I. Minn, K. Pa asy is, H. Menon, G. Geo gakoudis, and A. Bha ele, “Taking gpu p og am- ming models o ask o pe o mance po abili y,” a Xi p ep in a Xi :2402.08950, 2024. [3] C. Phuong, N. Saied, and C. Tanis, “Assessing Kokkos pe o mance on selec ed a chi ec u es,” in La in Ame ican High Pe o mance Compu ing Con e ence. Sp inge , 2019, pp. 170–184. [4] A. S. Du ek, R. Gaya i, N. Meh a, D. Doe le , B. Cook, Y. Ghada , and C. DeTa , “Case s udy o using Kokkos and SYCL as pe o mance-po able amewo ks o Milc-Dslash benchma k on NVIDIA, AMD and In el GPUs,” Mul iXscale Deli e able 2.5 Page 20 in 2021 In e na ional Wo kshop on Pe o mance, Po abili y and P oduc i i y in HPC (P3HPC), No . 2021, pp. 57–67. [5] M. B eye , A. Van C aen, and D. P lüge , “A compa ison o SYCL, OpenCL, CUDA, and OpenMP o massi ely pa - allel suppo ec o machine classi ica ion on mul i- endo ha dwa e,” in P oceedings o he 10 h In e na ional Wo kshop on OpenCL, se . IWOCL ’22. New Yo k, NY, USA: Associa ion o Compu ing Machine y, May 2022, pp. 1–12. [6] Y. Ding, C. Xu, H. Qiu, Q. Wang, W. Dai, Y. Lin, and Y. Che, “E alua ing pe o mance po abili y o SYCL and Kokkos: A case s udy on LBM simula ions,” in 2023 IEEE In l Con on Pa allel & Dis ibu ed P ocessing wi h Applica ions, Big Da a & Cloud Compu ing, Sus ainable Compu ing & Communica ions, Social Compu ing & Ne wo king (ISPA/BDCloud/SocialCom/Sus ainCom), Dec. 2023, pp. 328–335. [7] W.-C. Lin, T. Deakin, and S. McIn osh-Smi h, “E alua ing ISO C++ pa allel algo i hms on he e ogeneous HPC sys ems,” in 2022 IEEE/ACM In e na ional Wo kshop on Pe o mance Modeling, Benchma king and Simula ion o High Pe o mance Compu e Sys ems (PMBS). IEEE, 2022, pp. 36–47. [8] R. A. Ba olomeu, R. Hal e , J. H. Meinke, and G. Su mann, “E ec o implemen a ions o he n-body p oblem on he pe o mance and po abili y ac oss gpu endo s,” Fu u e Gene a ion Compu e Sys ems, ol. 169, p. 107802, 2025. [Online]. A ailable: h ps://www.sciencedi ec .com/science/a icle/pii/S0167739X25000974 [9] K. Asano ic, R. Bodik, B. C. Ca anza o, J. J. Gebis, P. Husbands, K. Keu ze , D. A. Pa e son, W. L. Plishke , J. Shal , S. W. Williams, and K. A. Yelick, “The landscape o pa allel compu ing esea ch: A iew om Be keley,” Elec ical Enginee ing and Compu e Sciences, Uni e si y o Cali o nia a Be keley, Technical Repo No. UCB/EECS-2006- 183, Decembe , ol. 18, no. 2006-183, p. 19, 2006. [10] C. R. T o , D. Leb un-G andié, D. A nd , J. Ciesko, V. Dang, N. Ellingwood, R. Gaya i, E. Ha ey, D. S. Hollman, D. Ibanez, N. Libe , J. Madsen, J. Miles, D. Poliako , A. Powell, S. Rajamanickam, M. Simbe g, D. Sunde land, B. Tu cksin, and J. Wilke, “Kokkos 3: P og amming model ex ensions o he exascale e a,” IEEE T ansac ions on Pa allel and Dis ibu ed Sys ems, ol. 33, no. 4, pp. 805–817, 2022. [11] J. H. Da is, P. Si a aman, I. Minn, K. Pa asy is, H. Menon, G. Geo gakoudis, and A. Bha ele, “An E alua i e Com- pa ison o Pe o mance Po abili y ac oss GPU P og amming Models,” a Xi p ep in a Xi :2402.08950, 2024. [12] W.-C. Lin, S. McIn osh-Smi h, and T. Deakin, “P elimina y epo : Ini ial e alua ion o S dPa implemen a ions on AMD GPUs o HPC,” Jan. 2024. [13] A. A nold, F. Fah enbe ge , C. Holm, O. Lenz, M. Bol en, H. Dachsel, R. Hal e , I. Kabadshow, F. Gähle , F. Hebe , J. Ise inghausen, M. Ho mann, M. Pippig, D. Po s, and G. Su mann, “Compa ison o scalable as me hods o long- ange in e ac ions,” Phys. Re . E, ol. 88, p. 063308, Dec 2013. [Online]. A ailable: h ps://link.aps.o g/doi/10.1103/PhysRe E.88.063308 [14] M. Dese no and C. Holm, “How o mesh up ewald sums. i. a heo e ical and nume ical compa ison o a ious pa icle mesh ou ines,” The Jou nal o Chemical Physics, ol. 109, no. 18, pp. 7678–7693, 11 1998. [Online]. A ailable: h ps://doi.o g/10.1063/1.477414 [15] R. W. Hockney and J. W. Eas wood, Compu e simula ion using pa icles. c c P ess, 2021. [16] J. J. Ce dà, V. Ballenegge , O. Lenz, and C. Holm, “P3m algo i hm o dipola in e ac ions,” The Jou nal o Chemical Physics, ol. 129, no. 23, p. 234104, 12 2008. [Online]. A ailable: h ps://doi.o g/10.1063/1.3000389 [17] S. Sla e y, S. T. Ree e, C. Junghans, D. Leb un-G andié, R. Bi d, G. Chen, S. Foge y, Y. Qiu, S. Schulz, A. Scheinbe g, A. Isne , K. Chong, S. Moo e, T. Ge mann, J. Belak, and S. Mniszewski, “Cabana: A pe o mance po able lib a y o pa icle-based simula ions,” Jou nal o Open Sou ce So wa e, ol. 7, no. 72, p. 4115, 2022. [Online]. A ailable: h ps://doi.o g/10.21105/joss.04115