Technical epo on pe o mance-po able elec os a ics
Mul iXscale Deli e able 2.5
Deli e able Type: Repo
Deli e ed in June, 2025
Mul iXscale
Eu oHPC Cen e o Excellence o
Mul iscale Modelling
Acknowledgemen
Funded by he Eu opean Union. This wo k has ecei ed unding om he Eu opean High Pe o mance Compu ing Join
Unde aking (JU) unde g an ag eemen No 101093169.
Disclaime
Funded by he Eu opean Union. Views and opinions exp essed a e howe e hose o he au ho (s) only and do no necessa ily
e lec hose o he Eu opean Union o he Eu opean High Pe o mance Compu ing Join Unde aking (JU). Nei he he Eu opean
Union no he g an ing au ho i y can be held esponsible o hem.
Mul iXscale Deli e able 2.5 Page ii
P ojec and Deli e able In o ma ion
P ojec Ti le Mul iXscale: Eu oHPC Cen e o Excellence o Mul iscale Modelling
P ojec Re . G an Ag eemen 101093169
P ojec Websi e h ps://www.mul ixscale.eu
Eu oHPC P ojec O ice Ma eo Mascagni
Deli e able ID D2.5
Deli e able Na u e Repo
Dissemina ion Le el Public
Con ac ual Da e o Deli e y P ojec Mon h 30 (30 h June, 2025)
Ac ual Da e o Deli e y 27 h June, 2025
Desc ip ion o Deli e able Re iew o he de elopmen o pe o mance-po able elec os a ic sol e s de el-
oped as a communi y lib a y.
Documen Con ol In o ma ion
Documen
Ti le: Technical epo on pe o mance-po able elec os a ics
ID: D2.5
Ve sion: As o June, 2025
S a us: Accep ed by S ee ing Commi ee
A ailable a : h ps://www.mul ixscale.eu/deli e ables
Documen his o y: In e nal P ojec Managemen Link
Re iew Re iew S a us: Re iewed
Au ho ship
W i en by: Rod igo Ba olomeu (FZJ)
Con ibu o s: Godeha d Su mann (FZJ)
Re iewed by: Alan Ó Cais (UB)
App o ed by: Alan Ó Cais (UB)
Documen Keywo ds
Keywo ds: Mul iXscale, HPC, Pe o mance Po abili y, Elec os a ic, Load Balance, ALL
27 h June, 2025
Disclaime : This deli e able has been p epa ed by he esponsible Wo k Package o he P ojec in acco dance wi h he
Conso ium Ag eemen and he G an Ag eemen . I solely e lec s he opinion o he pa ies o such ag eemen s on a
collec i e basis in he con ex o he P ojec and o he ex en o eseen in such ag eemen s.
Copy igh no ices: This deli e able was co-o dina ed by Rod igo Ba olomeu1(FZJ) on behal o he Mul iXscale conso -
ium wi h con ibu ions om Godeha d Su mann (FZJ). This wo k is licensed unde he C ea i e Commons A ibu ion
4.0 In e na ional License. To iew a copy o his license, isi :
h p://c ea i ecommons.o g/licenses/by/4.0
cb
1 [email p o ec ed]
Mul iXscale Deli e able 2.5 Page iii
Con en s
Execu i e Summa y 1
1 In oduc ion 2
1.1 Scope o he deli e able ................................................ 2
1.2 Ta ge audience . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2
1.3 Repo ou line ...................................................... 2
1.4 Pa ne con ibu ions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2
2 Pe o mance Po abili y 3
3 Elec os a ics Me hods and Implemen a ion Conside a ions 7
3.1 Ewald Summa ion Me hod .............................................. 7
3.2 Pa icle-Pa icle Pa icle-Mesh Me hod . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
3.3 Implemen a ion Conside a ions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
4 Elec os a ic Sol e Lib a y 10
4.1 Lib a y In e ace . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
4.2 MPI Communica ion managemen . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
4.3 Benchma ks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
5 In eg a ion o he ALL load balancing lib a y in o LAMMPS 14
5.1 ALL in e ace wi h LAMMPS . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
5.2 Gene al LAMMPS Benchma ks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
5.3 Benchma k wi h balancing me hods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
6 Conclusion and Pe spec i es 18
Acknowledgemen s 19
Re e ences 19
Lis o Figu es
1 Sequence o ope a ions and p og am low wi hin he N-body p oblem. The main phase o he p og am
is epea ed o N i e a ions, each compu ing a disc e e ime s ep o size d . P.I.: pa icle ini ializa ion;
F.C.: o ce compu a ion; ∆
: p opaga ion o posi ions in Veloci y Ve le ; ∆
: p opaga ion o hal s ep
eloci ies in Veloci y Ve le in eg a o . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
2 In e ac ions pe second s numbe o pa icles Non he in es iga ed a chi ec u es by he di e en im-
plemen a ions. No e he di e en scale o he y-axis o GH200 (b). We do no include he esul s o
OpenACC on he MI 250 since hey a e abou a ac o o 100 lowe han hose om he o he p og am-
ming models and nea ly independen o sys em size. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
3 Subdi ision o he global MPI communica o in o sub-communica o s o be used o he compu a ion
o he eal( )- and k-space con ibu ions, showing he necessa y da a ans e be ween global commu-
nica o and each subcommunica o . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
4 S ong scaling benchma k o he Ewald summa ion sol e implemen ed in he lib a y using a sys em o
16 million andomly placed pa icles in a pe iodic sys em o size 1003on Juwels Boos e (Fo schungszen-
um Jülich). The esou ces a e e enly spli be ween he -space and k-space compu a ion. . . . . . . . . 12
5 S ong scaling benchma k o he Ewald summa ion sol e implemen ed in he lib a y using a sys em o
16 million andomly placed pa icles in a pe iodic sys em o size 1003on Juwels Boos e (Fo schungszen-
um Jülich). Th ee o ou GPUs a e employed o he compu a ion o he k-space con ibu ion, he
o he GPUs a e used o he -space con ibu ion. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
6 S ong scaling benchma k o he Ewald summa ion sol e implemen ed in he lib a y using a sys em o
32 million andomly placed pa icles in a pe iodic sys em o size 1263on Juwels Boos e (Fo schungszen-
um Jülich). The esou ces a e e enly spli be ween he -space and k-space compu a ion. . . . . . . . . 13
7 Single node weak scaling o LAMMPS using JEDI and JUWELS Boos e (Fo schungzen um Jülich) o he
LJ and Rhodopsin benchma ks in LAMMPS using Kokkos. Scaled LJ simula ion con aining 55M a oms,
and Scaled Rhodopsin simula ion con aining 16M a oms bo h wi h 1 MPI p ocess pe GPU. No esul s
a e p esen ed o Juwels Boos e wi h 1 and 2 GPUs o he Rhodopsin benchma k due o memo y limi-
a ion o accommoda e he da a. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
8 S ong scaling o LAMMPS using JEDI o he LJ and Rhodopsin benchma ks in LAMMPS using Kokkos. 16
Mul iXscale Deli e able 2.5 Page i
9 Combined s ong and weak scaling o LAMMPS a he p e-exascale supe compu e JEDI (Fo schungszen-
um Jülich). Solid ba s use he s agge ed global load balancing me hod om he ALL lib a y in eg a ed
in o LAMMPS and ha ched ba s use he ecu si e bisec ion load balancing me hod na i e o LAMMPS. . 17
Lis o Tables
1 O e iew o implemen a ions on in es iga ed a chi ec u es. ✓indica es ha benchma ks we e pe -
o med. ✗indica es ha he implemen a ion was no suppo ed on ha a chi ec u e a he ime o he
benchma k. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
2 Fas es measu ed pe o mance da a ac oss di e en sys em sizes, measu ed in Giga-in e ac ions pe
second (GIn
s) on each pla o m o each p og amming model. The as es o e all measu ed pe o mance
is used as e e ence alue o he applica ion e iciency (%) and is indica ed in bold. % (cmp. Eq. 5) is he
weigh ed mean o % o e all a chi ec u es, assuming % =0, when benchma ks could no be pe o med.
The unde lined esul is he bes achie ed alue o %. (s) indica es a ian s using sha ed memo y, while
(nd) indica es he SYCL a ian using nd- anges. The las column shows he pe o mance pe o mance
po abili y me ic
P
¯
P
¯.................................................. 6
Mul iXscale Deli e able 2.5 Page 1
Execu i e Summa y
Wi h he as changing High Pe o mance Compu ing (HPC) landscape, he need o pe o mance po able implemen-
a ions has g own signi ican ly in he pas yea s. Among he TOP 50 machines on he TOP500 lis , only 10% do no
ely on G aphical P ocessing Uni (GPU)s as accele a o s o i s pe o mance. Mo eo e , accele a o s om mul iple
endo s hold he op spo s on he lis . The same pic u e is also p esen on Eu opean High Pe o mance Compu ing
Join Unde aking (Eu oHPC) and Gauss Supe compu ing Cen e (GCS) machines. Bo h ha e supe compu e s wi h
AMD and NVIDIA GPUs o di e en gene a ions, a chi ec u es, and compu e capabili ies.
This s iking eali y poses a g owing demand o pe o mance po abili y o scien i ic codes, lib a ies, and applica-
ions. Fo many domain applica ions, he e icien use o CPUs alone is no su icien o ex ac he a ailable pe -
o mance and ene gy e iciency a ailable on mode n a chi ec u es. As an example, he wo lds mos ene gy e icien
supe compu e JEDI (Eu oHPC/FZJ) is based on N idia’s G ace-Hoppe Supe chip, which combines CPU and GPU.
In his sys em he GPUs can be p og ammed using Compu e Uni ied De ice A chi ec u e (CUDA), howe e doing so
he code would no be able o un on o he endo ’s pla o ms.
Achie ing pe o mance po abili y can be accomplished by pai ing de ices and le e aging o loading o compu e ke -
nels h ough amewo ks such as OpenMP, OpenACC, Kokkos, o Raja. Ou p elimina y esea ch on a pe o mance
po able Ewald summa ion me hod enligh ens he necessi y o u he compa ison be ween di e en so wa e ame-
wo ks o de e mine he mos e ec i e app oach. The e o e, a pe o mance po abili y s udy was conduc ed using he
N-body p oblem as a es case o e alua e he e icacy o a ious p og amming models ac oss majo GPU endo s.
This s udy aimed o e alua e he ela i e s eng hs and weaknesses o each model, p o iding aluable insigh s in o
he op imal app oach o achie ing pe o mance po abili y ac oss di e se ha dwa e a chi ec u es.
Fo elec os a ic compu a ions based on a ian s o he Ewald summa ion me hod, which a e o en one o he mos
compu e in ense and ime consuming s eps o Molecula Dynamics (MD) simula ions, no only po abili y is desi ed
bu also modula i y. Task di ision can be achie ed by sepa a ing he sho - ange and long- ange con ibu ions o he
elec os a ic po en ial and o ces. Since he eal-space and Fou ie k-space pa s ha e di e en pe o mance cha ac-
e is ics, a mul i-domain pa i ioning app oach can be applied in conjunc ion wi h adjus ing ele an Ewald me hod
pa ame e s. This allows o a balanced compu a ional load ac oss independen ly unning pa i ions. To achie e op-
imal load balancing be ween he eal-space and k-space pa i ions, i is essen ial o op imize ope a ions pe o med
on bo h pa s as well as he amoun o compu e esou ces alloca ed o each ask.
In his deli e able, we p esen ad ances, implemen a ion de ails, a ionale, and benchma ks ela ed o he a o emen-
ioned me hods. The ou comes o he pe o mance po abili y s udy, as well as he p e iously epo ed FFT bench-
ma k, se e as ounda ion o he design choices included in he de elopmen o P3M me hod and he po able elec-
os a ic sol e lib a y.
Mul iXscale Deli e able 2.5 Page 2
1 In oduc ion
1.1 Scope o he deli e able
This deli e able aims a p o iding in o ma ion on he implemen a ion o a pe o mance po able elec os a ics sol e
and he in eg a ion o he ALL load balancing lib a y in o LAMMPS.
This epo con ains elemen s ele an o Task 2.3 - Pe o mance-po able elec os a ics sol e s and load-balancing
lead by Fo schungszen um Jülich.
1.2 Ta ge audience
This epo is in ended o eade s wi h echnical expe ise in HPC o in he use o mul i scale pa icle based simula ion
so wa e. While po en ial use s migh be in e es ed in o he de ails epo ed he e, he main pu pose is o epo on he
esul s and implemen a ions ha migh be mo e o in e es o scien i ic so wa e de elope s willing o in e ace hei
code o he po able elec os a ic sol e lib a y.
1.3 Repo ou line
This deli e able epo s on all ele an opics wi h espec o Task 2.3, led by Fo schungzen um Jülich.
The deli e able is di ided in o ou pa s:
•Pe o mance Po abili y: in his sec ion we p esen he up- o-da e s a us on he pe o mance po abili y o he
mos commonly used C++ po able p og amming models. Du ing he cou se o Task 2.3 a s udy was conduc ed
o access he usabili y and pe o mance o such models. As ou come, an upda ed s a us on he po abili y was
epo ed. P ac ical expe ience in he usage as well as a deepe unde s anding o he implemen a ion de ails
o each model was acqui ed. This sec ion con ibu ed o he choice o he p og amming model used o he
implemen a ions o elec os a ic sol e s.
•Theo e ical ounda ions o Ewald sum and P3M me hods: in his sec ion we p esen he o mula ion o he ba-
sic Ewald sum, which is he o iginal me hod o compu e he elec os a ic ene gy o an a omis ic sys em unde
pe iodic bounda y condi ions. The accele a ed e sion, P3M, makes use o he Fas Fou ie ans o m, e alua ed
on a egula mesh, ca ying he mapped pa icle cha ges. Ca dinal B-Splines a e applied o map cha ges o he
mesh and o e ie e po en ial ene gies and o ces back o he pa icles. While being able o con ol he app ox-
ima ion e o , he P3M me hod is educing he compu a ional complexi y om O(N3/2) down o O(Nlog N),
whe e Nis he numbe o pa icles in he sys em.
•Elec os a ic Sol e Lib a y: in his sec ion we discuss he design and implemen a ion o he pe o mance
po able and scalable lib a y app oach o elec os a ic sol e s. In he p esen p ojec we ocus on he Ewald
summa ion as a benchma k implemen a ion and he P3M as an op imized implemen a ion o p oduc ion uns
on mode n HPC sys ems. Pe o mance po abili y has been conside ed o bo h he eal and ecip ocal space
pa o he me hod. P ope pa ame e choice o e o con ol and aspec s o di e en scalabili y o eal and
ecip ocal space pa s a e discussed and conside ed in he design o he lib a y. Fi s benchma ks on di e en
a chi ec u es a e shown and p ope ies o scalabili y a e discussed in e ms o di e en asks in he me hod.
•Load Balancing: wi h he in en o benchma k he load balancing me hods a ailable in he A Load balancing
Lib a y (ALL) wi hin using ou key applica ion codes. We ha e in eg a ed ALL o LAMMPS h ough an in e ace
ha p o ide use s he capabili y o using ex ended ix commands. In he con ex o his deli e able we p esen
benchma ks compa ing LAMMPS pe o mance using A100 and GH200 GPUs.
1.4 Pa ne con ibu ions
JSC con ibu ed as planned o he wo k in his deli e able which is in ended o be con ibu e o ealisa ion o o he
asks in he p ojec .
Mul iXscale Deli e able 2.5 Page 3
2 Pe o mance Po abili y
The ongoing de elopmen o new GPUs om a ious endo s equi es con inuous moni o ing o po able ame-
wo ks’ pe o mance and e sa ili y. A p ima y conce n is whe he a po able amewo k is a ailable and unc ional on
a gi en a chi ec u e.
While he concep o po abili y be ween di e en pla o ms is s aigh o wa d, de ining and measu ing pe o mance
po abili y is mo e complex. Success ul p og am execu ion is a clea indica o o po abili y, bu e alua ing pe o -
mance po abili y equi es a mo e nuanced app oach.
A common me hod o assessing he pe o mance and capabili ies o po able p og amming models in ol es using
mini-apps o ull applica ions o ga he ele an da a [1, 2]. This can be achie ed by po ing a single applica ion
o a ious pe o mance-po able p og amming models [3–6] o by po ing mul iple mini-apps o one pe o mance-
po able p og amming model and compa ing he esul s o o he a ailable e sions [7]. In ou s udy, we ocus on he
N-body p oblem [8].
The N-body p oblem is a signi ican p oblem class and has been iden i ied as one o he se en o iginal dwa s in
Re . [9], each ep esen ing a compelling use case wi h i s own challenges. Many scien i ic ields equi e compu ing
o ces be ween pai s o in e ac ing pa icles, such as g a i a ional o ces in as ophysics, elec os a ic o ces in cha ged
sys ems, and an de Waals o ces o modeling a omic sys ems.
The g a i a ional po en ial be ween pa icles is gi en by:
U(di j )=−Gmimj
di j
(1)
whe e G=6.6743·10−11 m3
kg·s2is he g a i a ional cons an , di j =| i j |is he dis ance be ween pa icles iand j, i j =
i− j, and miand mja e he masses o he espec i e pa icles.
Gi en a po en ial U(di j ), he o ce i j be ween pa icles iand jis gi en by
Fi j =−∇U(di j )=−Gmimj
d3
i j
i j . (2)
The o al o ce ac ing on pa icle iis hen gi en by he sum o e all o he pa icles
Fi=∑︂
j=i
Fi j . (3)
To ob ain he ajec o y o each pa icle, New on’s equa ions o mo ion
x
˙i= i,
˙i=ai=Fi
mi
, (4)
a e in eg a ed, whe e he do -no a ion ep esen s he ime de i a i e, aiis he accele a ion, and miis he mass o
pa icle i.
The g a i a ional N-body p oblem ( he 4 h dwa in Re . [9]) is a ep esen a i e o an algo i hmic me hod ha includes
long- ange in e ac ions in open-bounda ies and in ol es summa ion o e all pa icle pai s in he sys em o compu e
he o ces on each pa icle, esul ing in a na u al compu a ional complexi y o O(N2). To a oid his quad a ic com-
plexi y, mo e e icien me hods like he Ba nes-Hu ee me hod o he as mul ipole me hod (FMM) can be used, e-
ducing he complexi y o O(Nlog(N)) o O(N). Howe e , since hese me hods ha e a high le el o complexi y and hei
pe o mance s ongly depends on he implemen a ion and op imiza ion, we ocus he e on a simple implemen a ion
based on he all-pai s compu a ion o ene gies and o ces, which isola es he pe o mance ou come om algo i hmic
issues.
To e alua e he u iliza ion o he ha dwa e a chi ec u es in ou benchma ks, we compa ed he h oughpu o pa icle-
pa icle compu a ions, whe e he numbe o in e ac ions scales as N2. The s udy used he N-body p oblem as a es
om compu e in ense algo i hm ha is p esen in many MD so wa e. An illus a ion o he p og am wo k low is
p esen ed in Figu e 1.
In o de o e alua e he pe o mance po abili y o di e en p og amming models, he discussed algo i hm was im-
plemen ed in ele en a ia ions. The C++ amewo ks conside ed we e: Kokkos [10], OpenMP, OpenACC, SYCL, HIP,
CUDA, and pSTL. The implemen a ions we e es ed on ou pla o ms. An o e iew o he es ed pla o ms is p e-
sen ed in Table 1.
Fo de ails abou pla o m speci ica ions, compile s, implemen a ions, and op imiza ions, eade s may e e o [8].
Mul iXscale Deli e able 2.5 Page 4
ini main
I/O Δ𝑣, Δ𝑟
S a P.I. F.C. End
Ini . F.C. Δ𝑣
𝑁𝑡 ime s eps
Figu e 1: Sequence o ope a ions and p og am low wi hin he N-body p oblem. The main phase o he p og am
is epea ed o N i e a ions, each compu ing a disc e e ime s ep o size d . P.I.: pa icle ini ializa ion; F.C.: o ce
compu a ion; ∆
: p opaga ion o posi ions in Veloci y Ve le ; ∆
: p opaga ion o hal s ep eloci ies in Veloci y Ve le
in eg a o .
A chi ec u e OpenMP pSTL Kokkos CUDA HIP SYCL OpenACC
N idia A100 ✓ ✓ ✓ ✓ ✓ ✓ ✓
N idia GH200 ✓ ✓ ✓ ✓ ✓ ✓ ✓
AMD MI250 ✓ ✓ ✓ ✗ ✓ ✓ ✓
In el GPU Max 1100 ✓ ✓ ✓ ✗ ✗ ✓ ✗
Table 1: O e iew o implemen a ions on in es iga ed a chi ec u es. ✓indica es ha benchma ks we e pe o med.
✗indica es ha he implemen a ion was no suppo ed on ha a chi ec u e a he ime o he benchma k.
As shown in Figu e 2, mo e powe ul GPUs equi e la ge sys em sizes o sa u a e he p ocesso , indica ed by he
la ening o he lines in he plo s.
The esul s show ha a ian s using sha ed memo y gene ally pe o m be e han hose wi hou i , pa icula ly on
he GH200. Howe e , he e a e some excep ions ha equi e u he in es iga ion. No ably, CUDA on he GH200 wi h
sha ed memo y achie ed he bes esul s, while SYCL ou pe o med HIP on he MI250. In con as , OpenMP and
pSTL ou pe o med SYCL on he In el pla o m. Fu he mo e, OpenACC, OpenMP, and pSTL su p isingly achie ed
highe pe o mance on he A100 ca d compa ed o CUDA wi h sha ed memo y. These indings align wi h a s udy
by [11], which ound ha SYCL’s so wa e s ack imp o emen s led o ou pe o ming na i e p og amming models on
compu e-bound ke nels. These esul s a e also p esen ed nume ically in Figu e 2 o discussion:
Using he mean applica ion e iciency % (see Equa ion 5), i can be seen ha he models designed o be po able pe -
o m be e han hose a ge ed a speci ic pla o ms (HIP, CUDA). Al hough his is o be expec ed, i is no ewo hy ha
pSTL and SYCL (using sha ed memo y) show he bes po abili y quali y. Kokkos shows compa a i ely lowe pe o -
mance on he AMD and In el pla o ms. Rega ding AMD, ou esul di e s om ecen s udies, whe e Kokkos showed
excellen pe o mance on AMD-based sys ems [2,12]. This may be due o he a ian chosen o i s pe o mance on
he GH200.
%=
%A100 +%GH200
2+%MI250 +%Max1100
3. (5)
The espec i e esul s o
P
¯
P
¯(Table 2) show li le de ia ion om he esul s o % o he po able p og amming model,
bu make i clea ha CUDA, HIP, and OpenACC a e no po able since hey a e no applicable (NA) pe de ini ion
o p og amming models ha canno be execu ed on one o mo e o he in es iga ed a chi ec u es. The de ia ion
in he pe cen ages o he po able p og amming models comes om he di e en ways o a e aging (see. Equa-
ion 5).
Fo a mo e de ailed analysis o pe o mance po abili y he numbe o aken measu emen s as well as he numbe o
a ge ed pla o ms would need o be inc eased. Despi e his, he pe o med analysis on he p esen ed da a al eady
indica es ha using an op imized e sion o a code on a gi en a chi ec u e (in his case he GH200) migh no be
op imally sui ed o o he a chi ec u es, bu s ill pe o ms sa is ac o y well on he mos mode n sys em om all majo
HPC GPU endo s.
Ou esul s show ha OpenMP, pSTL, and SYCL a e well sui ed o po able implemen a ions o he N-body p oblem
ac oss a ious GPU endo s, while Kokkos showed sligh ly educed pe o mance on he non-N idia pla o ms. Kokkos
Mul iXscale Deli e able 2.5 Page 5
CUDA CUDA (s) HIP HIP (s)
Kokkos pSTL OpenMP OpenACC
SYCL SYCL (s) SYCL- ange
3236439631283
50
100
150
200
𝑁
GIn / s
(a) A100
3236439631283
100
200
300
400
𝑁
GIn / s
(b) GH200
3236439631283
50
100
150
200
𝑁
GIn / s
(c) MI 250
3236439631283
50
100
150
200
𝑁
GIn / s
(d) GPU Max 1100
Figu e 2: In e ac ions pe second s numbe o pa icles Non he in es iga ed a chi ec u es by he di e en imple-
men a ions. No e he di e en scale o he y-axis o GH200 (b). We do no include he esul s o OpenACC on he
MI 250 since hey a e abou a ac o o 100 lowe han hose om he o he p og amming models and nea ly indepen-
den o sys em size.
migh show be e pe o mance, when using algo i hms mo e dependen on complex memo y access pa e ns, e.g.,
sho - ange in e ac ions, as in he case o Ewald Sum and he Pa icle-Pa icle Pa icle-Mesh algo i hms.
Mul iXscale Deli e able 2.5 Page 12
numbe s o nodes used, while o he la ge numbe o nodes used he -space un ime becomes dominan . O e all a
pa allel e iciency o ≈35% was eached on 128 GPUs (32 nodes).
As a consequence o he inding a la ge alloca ion o esou ces o he k-space pa migh be ob ious o imp o e he
o al un- ime. The e o e, he benchma k p esen ed in Figu e 5used h ee o ou GPUs pe node o he compu a ion
o he k-space con ibu ion and one GPU o each node o he -space con ibu ion. Howe e , i can be seen ha he
o e all pa allel e iciency decays as e han in he i s benchma k. This is due o he ac ha he -space compu-
a ion los pe o mance and domina es he o e all pa allel e iciency. The imp o ed scaling beha io o he k-space
compu a ion can no be u ilized, as he -space compu a ion always akes longe o execu e.
In a hi d benchma k a la ge sys em was used o in es iga e he impac o sys em size on s ong scalabili y. The
esul s a e shown in Figu e 6. In his benchma k he esou ces we e e enly spli be ween bo h pa s, - and k-space
con ibu ions. I can be seen ha an e iciency o ≈60% was achie ed.
Fo all benchma ks he di e en compu a ion pa s a e execu ed on sub-communica o s which do no sha e p o-
cesses. This means ha hey un independen ly in pa allel bu wi h di e en indi idual pa allel e iciencies. As a
esul he sum o bo h pa s is la ge han he o al un ime. Fi s analysis indica es ha his app oach o spli ing
compu e esou ces can be bene icial o node coun s whe e he o e all pa allel e iciency s a s o deg ade. In o de
o be able o be e guide he communi y in ha espec u he in-dep h analysis will be pe o med.
Figu e 4: S ong scaling benchma k o he Ewald summa ion sol e implemen ed in he lib a y using a sys em o 16
million andomly placed pa icles in a pe iodic sys em o size 1003on Juwels Boos e (Fo schungszen um Jülich).
The esou ces a e e enly spli be ween he -space and k-space compu a ion.
Mul iXscale Deli e able 2.5 Page 13
Figu e 5: S ong scaling benchma k o he Ewald summa ion sol e implemen ed in he lib a y using a sys em o 16
million andomly placed pa icles in a pe iodic sys em o size 1003on Juwels Boos e (Fo schungszen um Jülich).
Th ee o ou GPUs a e employed o he compu a ion o he k-space con ibu ion, he o he GPUs a e used o he
-space con ibu ion.
Figu e 6: S ong scaling benchma k o he Ewald summa ion sol e implemen ed in he lib a y using a sys em o 32
million andomly placed pa icles in a pe iodic sys em o size 1263on Juwels Boos e (Fo schungszen um Jülich).
The esou ces a e e enly spli be ween he -space and k-space compu a ion.
Mul iXscale Deli e able 2.5 Page 14
5 In eg a ion o he ALL load balancing lib a y in o LAMMPS
The A Load balancing Lib a y (ALL) is a modula , open-sou ce so wa e de eloped by he Jülich Supe compu ing
Cen e o add ess he c i ical challenge o load imbalance in domain-decomposed simula ions. In high-pe o mance
compu ing (HPC), pa icula ly in pa icle-based me hods like molecula dynamics o la ice-Bol zmann simula ions,
spa ial domain decomposi ion is a common s a egy o dis ibu e compu a ional wo kload. Howe e , as simula ions
e ol e, non-uni o m dis ibu ions o pa icles o a ying local compu a ional cos s can lead o signi ican load imbal-
ances, educing o e all e iciency. ALL p o ides a sui e o dynamic load-balancing s a egies speci ically designed o
edis ibu e wo k ac oss p ocesso s in such simula ions, ensu ing be e u iliza ion o HPC esou ces.
By suppo ing mul iple balancing schemes such as his og am and i e a i e g id pa i ioning ALL o e s lexible and
scalable solu ions ailo ed o di e en applica ion needs. I s in eg a ion in o widely-used simula ion codes like HemeLB
and DL_MESO_DPD highligh s i s e ec i eness in imp o ing pe o mance on la ge-scale sys ems, including GPU ac-
cele a ed pla o ms.
Wi hin he scope o he cu en Deli e able, we highligh he mos e icien and used me hods, which a e p esen in
he lib a y o being al eady implemen ed in LAMMPS:
Tenso Decomposi ion (Local) This me hod spli s he simula ion domain along Ca esian axes using a ixed numbe
o pa i ions in each di ec ion (e.g., a g id layou ). I is an i e a i e me hod ha , o e e y dimension, educes he wo k
o e he o he dimensions and shi s he bo de s be ween wo neighbou s based on hei cumula i e wo k, making i
simple bu less adap i e o highly une en load dis ibu ions.
•Classic Fo he classic app oach he wo k pe dimension is collec ed using he a e age o e he domains in ha
dimension.
•Max The max a ian conside s he maximum wo k in each o he slices in ca esian di ec ion and conside s as
objec i e unc ion an e en dis ibu ion o he maxima in he slices o each dimension.
S agge ed G id The s agge ed g id is a special domain decomposi ion scheme ha ecu si ely cu s in each dimension
in a de ined o de . I is a special case o he iled decomposi ion as used o ecu si e bisec ion and is, in p inciple,
able o p oduce a pe ec decomposi ion o e en load in he sys em.
•Global (Global + His og am) This me hod di ides he sys em ac oss domains in o g id egions and adjus s he
bounda ies using global wo kload his og ams. I op imizes he load balance by analyzing he dis ibu ion o
compu a ional wo k ac oss he en i e domain, hen shi s domain bounda ies acco dingly. I is a global me hod
and uses his og am-based e inemen o edis ibu e wo k mo e e enly.
•Local (Local + I e a i e) Unlike he global a ian , his me hod balances wo kloads based on local neighbo - o-
neighbo in e ac ions. I i e a i ely adjus s bounda ies by exchanging wo kload da a wi h neighbo ing domains,
g adually imp o ing balance h ough a se ies o local upda es. This local me hod uses i e a ion-based e ine-
men , making i scalable and esponsi e o localized imbalances.
Recu si e Bisec ion (Global + I e a i e) This is a global and hie a chical me hod ha ecu si ely spli s he domain
in o wo sub egions a each s ep, aiming o e enly dis ibu e he o al wo kload. The spli ing p ocess can be guided
by ei he wo kload his og ams o i e a i e e alua ion o load imbalance. I is e ec i e o complex o non-uni o m
load dis ibu ions bu can in oduce communica ion o e head i applied oo equen ly.
5.1 ALL in e ace wi h LAMMPS
ALL balancing me hods ha e been in eg a ed in o LAMMPS using he s anda d de ini ions o so-called ixes. Once
LAMMPS is build wi h -D PKG_ALLPKG=yes, he ollowing op ions become a ailable:
ix ID g oup-ID balance/all keywo d a gs ...
equi ed keywo d/a g pai s
e e y a g = ne e y
ne e y = pe o m dynamic load balancing e e y his many s eps
g id s yle a gs = de ine g id
s yle = s agge ed o enso
s agge ed a gs = none
enso a gs = max o classic
max = use TENSOR_MAX me hod o ALL
classic = use TENSOR me hod o ALL
Mul iXscale Deli e able 2.5 Page 15
weigh s yle a gs = use weigh ed pa icle coun s o he balancing
s yle = g oup o neigh o ime o a o s o e
g oup a gs = Ng oup g oup1 weigh 1 g oup2 weigh 2 ...
Ng oup = numbe o g oups wi h assigned weigh s
g oup1, g oup2, ... = g oup IDs
weigh 1, weigh 2, ... = co esponding weigh ac o s
neigh ac o = compu e weigh based on numbe o neighbo s
ac o = scaling ac o (> 0)
ime ac o = compu e weigh based on ime spend compu ing
ac o = scaling ac o (> 0)
a name = ake weigh om a om-s yle a iable
name = name o he a om-s yle a iable
s o e name = s o e weigh in cus om a om p ope y de ined by ix p ope y/a om
command
name = a om p ope y name (wi hou d_ p e ix)
A leas one o he ollowing keywo d/a g pai s is equi ed.
Only one (o none) load balance is used in a load balancing s ep.
global a gs = igge _ h eshold bin_wid h
igge _ h eshold = loa
loa = apply load balance i he imbalance o he load is abo e his
h eshold
bin_wid h = app oxima e wid h o bin (a bi a y posi i e alue)
local a gs = igge _ h eshold (use local balancing)
igge _ h eshold = loa
loa = apply load balance i he imbalance o he load is abo e his
h eshold
op ional keywo d a g pai s
e bose a gs = none ( e bose ou pu o log/sc een)
ile a g = ilename
ilename = w i e s a s pe ank o ilename
In addi ion o he ix balance/all me hod, a specialized communica ion s yle was added in o de o co ec ly and
e icien ly manage he MPI communica o s se by he domain decomposi ion. Some examples o he a ailable bal-
ancing me hods a e:
# his og am me hod (s agge ed global), weigh ed by numbe o a oms,
# wi h bin wid h 0.003, when imbalance >= 1.05
comm_s yle s agge ed zyx
ix 5 all balance/all e e y 1000 g id s agge ed global 1.05 0.003
# his og am me hod (s agge ed global), weigh ed by ime, wi h bin wid h 0.003, when imbalance >= 1.05
comm_s yle s agge ed zyx
ix 5 all balance/all e e y 1000 g id s agge ed weigh ime 1.0 global 1.05 0.003
# s agge ed local me hod, weigh ed by numbe o a oms, when imbalance >= 1.05
comm_s yle s agge ed zyx
ix 5 all balance/all e e y 100 g id s agge ed local 1.05
# enso me hod, minimizing a e age wo k, weigh ed by numbe o a oms, when imbalance >= 1.05
comm_s yle b ick
ix 5 all balance/all e e y 50 g id enso classic local 1.05
# enso -max me hod, minimizing maximal wo k, weigh ed by numbe o a oms, when imbalance >= 1.05
comm_s yle b ick
ix 5 all balance/all e e y 50 g id enso max local 1.05
Mul iXscale Deli e able 2.5 Page 16
5.2 Gene al LAMMPS Benchma ks
Wi h he in en ion o es LAMMPS pe o mance on he mos ecen a chi ec u e p esen a he Eu oHPC la es supe -
compu e JUPITER, we ha e un se e al benchma ks o LAMMPS wi h s anda d sys ems. The es consis ed in weak
and s ong scaling on JEDI and JUWELS Boos e sys ems bo h a ailable a Fo schungszen um Jülich2.
Fo he alida ion benchma ks, LAMMPS was compiled using GCC 12.3.0, CUDA 12.2, and OpenMPI 4.1.6, and he
chosen FFT backend was cuFFT. In Figu es 7and 8we p esen he esul s om po ing LAMMPS o he newe a chi-
ec u es.
(a) LJ (b) Rhodopsin
Figu e 7: Single node weak scaling o LAMMPS using JEDI and JUWELS Boos e (Fo schungzen um Jülich) o he
LJ and Rhodopsin benchma ks in LAMMPS using Kokkos. Scaled LJ simula ion con aining 55M a oms, and Scaled
Rhodopsin simula ion con aining 16M a oms bo h wi h 1 MPI p ocess pe GPU. No esul s a e p esen ed o Juwels
Boos e wi h 1 and 2 GPUs o he Rhodopsin benchma k due o memo y limi a ion o accommoda e he da a.
As i can be obse ed in Figu e 7bo h sys ems and a chi ec u e combina ions show close o ideal scaling wi hin single
node con igu a ions. The expec ed speed up in he newe a chi ec u e (2x) is achie ed and he la ge memo y allows
o he simula ion o bigge sys em sizes wi h less compu ing esou ces.
In Figu e 7b one can obse e ha due o memo y and communica ion cha ac e is ics, i is possible o achie e simila
esul s o he Rhodopsin benchma k using one H100 o he GH200 node ins ead o h ee A100 GPUs.
(a) LJ (b) Rhodopsin
Figu e 8: S ong scaling o LAMMPS using JEDI o he LJ and Rhodopsin benchma ks in LAMMPS using Kokkos.
S ong scaling esul s p esen in he g ouped ba s o Figu e 8show ha o he LJ benchma k he minimum sys em
size equi ed o p o i om alloca ing mul i-node jobs was 55.4 million pa icles. Once aking in o conside a ion he
mo e complex sys em, i.e., he Rhodopsin benchma k, one may expec 16.5 million pa icles o be su icien a he
es ed ange, which seems o esul om he mo e complex in e ac ions. In ac , he Rhodopsin benchma k includes
all elec os a ic o ces, bonds, dihed als and imp ope con ibu ions which esul s in a much la ge compu a ional
wo k compa ed o he simple LJ case.
2The JEDI sys em has now been decommissioned and has been in eg a ed in o he JUPITER sys em.
Mul iXscale Deli e able 2.5 Page 17
Figu e 9: Combined s ong and weak scaling o LAMMPS a he p e-exascale supe compu e JEDI (Fo schungszen-
um Jülich). Solid ba s use he s agge ed global load balancing me hod om he ALL lib a y in eg a ed in o LAMMPS
and ha ched ba s use he ecu si e bisec ion load balancing me hod na i e o LAMMPS.
5.3 Benchma k wi h balancing me hods
Fo es ing he pe o mance o he load balancing lib a y, a syn he ic benchma k was gene a ed emo ing a oms om
a sphe ical egion a he cen e o he LJ s anda d benchma k. The sys em was scaled in he same manne as he ones
p esen ed in he p e ious sec ion.
In Figu e 9 esul s show compa isons be ween load balancing me hods, i.e., RCB, which is one o LAMMPS’s na i e
benchma k me hods, and he s agge ed global me hods (ALL) using he same in e al o checking he load imbalance
and ac i a ing he load balancing p ocedu e wi h a common c i e ion. The balancing commands added o he inpu
sc ip we e:
comm_s yle s agge ed zyx
ix 5 all balance/all e e y 50 g id s agge ed global 0.9 0.003
in he case o using he me hod a ailable h ough he in eg a ion wi h ALL, and :
comm_s yle iled
ix 5 all balance 50 cb 0.9
o LAMMPS’s ecu si e bisec ion algo i hm (RCB).
I can be obse ed ha o all numbe s o nodes he s agge ed/global me hod (solid ba s) esul ed in a pe o mance
imp o emen o 30% o 50% compa ed o RCB (ha ched ba s), which shows a p omising capabili y o enhancing he
pe o mance o load imbalanced simula ions on exascale supe compu e s.
These simula ions we e pe o med using he GPU package in LAMMPS h ough he use o -s gpu command line.
Mul iXscale Deli e able 2.5 Page 18
6 Conclusion and Pe spec i es
This deli e able p esen s he de elopmen s achie ed h ough he e o s o Task 2.3. Sec ions 2 o 5co e pe o -
mance po abili y, elec os a ic me hods and implemen a ion conside a ions, lib a y API and benchma ks, and he
in eg a ion o he ALL load balancing lib a y in o LAMMPS. These indi idual opics a e in e connec ed, wi h po abil-
i y aspec s in luencing he design choices o elec os a ic sol e s and hei impac on a ge sys ems o op imal load
balancing me hods ha in e ace wi h pa icle simula ions.
Kokkos has p o en o be a eliable choice o po ing codes o newe a chi ec u es, o e ing signi ican ad an ages in
complex da a s uc u es and widesp ead communi y adop ion. Al hough i may no yield he bes o e all
P
¯
P
¯ alues,
i consis en ly pe o ms well on sys ems wi h a obus backend so wa e s ack. As demons a ed in Sec ion 2, Kokkos
ma ches he pe o mance o na i e/ endo - ecommended pla o m-dependen solu ions wi hou equi ing labo ious
compile op imiza ions. Howe e , i is essen ial o no e ha pla o m- ela ed compile op imiza ions may posi i ely
impac Kokkos’ pe o mance, and sys em main aine s a e encou aged o explo e his opic.
Sec ion 3p esen s algo i hmic conside a ions aimed a ex ac ing he desi ed pe o mance om mode n exascale a -
chi ec u es. Each independen s ep can u ilize speci ic sub-communica o s and un in pa allel on CPU+GPU combi-
na ions o ac oss modula pa i ions. Bo h Ewald summa ion and P3M me hods employ ei he pa icle o mesh-based
pa allelism h ough Cabana me hods o Kokkos Mul i-Dimensional Range loops.
The en a i e API o he lib a y and p elimina y benchma ks a e p esen ed in Sec ion 4. These benchma ks alida e
and highligh he lib a y’s modula capabili y, demons a ing he spli ing o and k-space me hods and hei espec-
i e scalabili ies o he Ewald summa ion me hod (Figu es 4 o 6). These benchma ks we e pe o med by using wo
independen alloca ions o he - and k-space pa s on Juwels Boos e (Fo schungszen um Jülich). The discussion
includes a c i ical assessmen o he equi emen s and capabili ies o a balancing me hod o un - and k-space pa s
mos e icien ly. S ong scaling o 16 and 32 million pa icles scaled om 4 o 128 GPUs e eals ha pa allel e i-
ciency is a ec ed by he manne in which compu a ional e o o - and k-space pa s is dis ibu ed on o esou ces.
Ini ial analysis sugges s ha his app oach o esou ce alloca ion can be bene icial o node coun s whe e o e all
pa allel e iciency begins o deg ade. Fu he in-dep h analysis will be conduc ed o p o ide mo e comp ehensi e
guidance.
The P3M implemen a ion is cu en ly unde going alida ion, and i s ex ac ion o he new elec os a ic lib a y has
commenced. We plan o analyze he impac o modula alloca ion wi h mesh me hods using he same me hodology
p esen ed o he Ewald Summa ion. Bo h me hods will be op imized o unning on Eu ope’s la ges supe compu e ,
JUPITER.
Las ly, he LAMMPS po ing o he newes a chi ec u e has exhibi ed expec ed weak and s ong scaling beha io . Pe -
o mance on GH200 nodes is expec ed o be wice as as as on he p e ious N idia GPU a chi ec u e (A100). Fo e y
la ge simula ions, ene gy e iciency imp o es mo e han wo old, as ewe nodes and GPUs a e equi ed o pe o m
equally sized simula ions.
Following he analysis o LAMMPS scaling beha io on bo h JEDI and Juwels Boos e , p elimina y esul s using he
in eg a ion o ALL in o LAMMPS a e p esen ed in Sec ion 5.3. As he pe o mance o load balancing depends on
ini ial con igu a ion and compu ing cos s (desc ibed in Sec ion 5.2), a mo e comp ehensi e s udy will be conduc ed
o e alua e he di e ences be ween LAMMPS’ implemen ed balance me hods and hose added by he in eg a ion
wi h he ALL lib a y. The s anda diza ion o he in e ace is nea o comple ion and will be made a ailable o he
communi y by c ea ing a pull eques o LAMMPS’ main eposi o y.
Mul iXscale Deli e able 2.5 Page 19
Acknowledgemen s
We hank he Jülich Supe compu ing Cen e, Fo schungszen um Jülich o access o he JURECA-DC E alua ion Pla -
o m, JUWELS Boos e and JEDI. This wo k has ecei ed unding om he Eu opean High Pe o mance Compu ing
Join Unde aking unde he g an ag eemen 101093169 and was suppo ed by he Ge man Fede al Minis y o Edu-
ca ion and Resea ch (BMBF) and he Minis y o Cul u e and Science (MKW) o he s a e o No h-Rhine-Wes phalia
h ough unding o he Gauss Cen e o Supe compu ing (GCS).
Re e ences
Ac onyms used
HPC High Pe o mance Compu ing
FFT as Fou ie ans o m
Eu oHPC Eu opean High Pe o mance Compu ing Join Unde aking
GPU G aphical P ocessing Uni
MD Molecula Dynamics
GCS Gauss Supe compu ing Cen e
API Applica ion P og am In e ace
So wa e men ioned
CUDA Compu e Uni ied De ice A chi ec u e
heFFTe highly e icien Fas Fou ie T ans o m o exascale
ALL A Load balancing Lib a y
ScaFaCoS Scalable Fas Coulomb Sol e s
URLs e e enced
Page ii
h ps://www.mul ixscale.eu ... h ps://www.mul ixscale.eu
h ps://www.mul ixscale.eu/deli e ables ... h ps://www.mul ixscale.eu/deli e ables
In e nal P ojec Managemen Link ... h ps://gi hub.com/mul ixscale/planning/
.ba olomeu@ z-juelich.de ... mail o: .ba olomeu@ z-juelich.de
h p://c ea i ecommons.o g/licenses/by/4.0 ... h p://c ea i ecommons.o g/licenses/by/4.0
Page 1
TOP500 ... h ps://www. op500.o g
Page 7
ScaFaCoS ... h p://www.sca acos.de/
Page 9
Cabana ... h ps://gi hub.com/ECP-copa/Cabana
#799 ... h ps://gi hub.com/ECP-copa/Cabana/pull/799
#800 ... h ps://gi hub.com/ECP-copa/Cabana/pull/800
Page 18
JUPITER ... h ps://www. op500.o g/lis s/ op500/2025/06/
Ci a ions
[1] T. Deakin, S. McIn osh-Smi h, J. P ice, A. Poena u, P. A kinson, C. Popa, and J. Salmon, “Pe o mance po abili y
ac oss di e se compu e a chi ec u es,” in 2019 IEEE/ACM In e na ional Wo kshop on Pe o mance, Po abili y
and P oduc i i y in HPC (P3HPC). IEEE, 2019, pp. 1–13.
[2] J. H. Da is, P. Si a aman, I. Minn, K. Pa asy is, H. Menon, G. Geo gakoudis, and A. Bha ele, “Taking gpu p og am-
ming models o ask o pe o mance po abili y,” a Xi p ep in a Xi :2402.08950, 2024.
[3] C. Phuong, N. Saied, and C. Tanis, “Assessing Kokkos pe o mance on selec ed a chi ec u es,” in La in Ame ican
High Pe o mance Compu ing Con e ence. Sp inge , 2019, pp. 170–184.
[4] A. S. Du ek, R. Gaya i, N. Meh a, D. Doe le , B. Cook, Y. Ghada , and C. DeTa , “Case s udy o using Kokkos
and SYCL as pe o mance-po able amewo ks o Milc-Dslash benchma k on NVIDIA, AMD and In el GPUs,”
Mul iXscale Deli e able 2.5 Page 20
in 2021 In e na ional Wo kshop on Pe o mance, Po abili y and P oduc i i y in HPC (P3HPC), No . 2021, pp.
57–67.
[5] M. B eye , A. Van C aen, and D. P lüge , “A compa ison o SYCL, OpenCL, CUDA, and OpenMP o massi ely pa -
allel suppo ec o machine classi ica ion on mul i- endo ha dwa e,” in P oceedings o he 10 h In e na ional
Wo kshop on OpenCL, se . IWOCL ’22. New Yo k, NY, USA: Associa ion o Compu ing Machine y, May 2022,
pp. 1–12.
[6] Y. Ding, C. Xu, H. Qiu, Q. Wang, W. Dai, Y. Lin, and Y. Che, “E alua ing pe o mance po abili y o SYCL and
Kokkos: A case s udy on LBM simula ions,” in 2023 IEEE In l Con on Pa allel & Dis ibu ed P ocessing wi h
Applica ions, Big Da a & Cloud Compu ing, Sus ainable Compu ing & Communica ions, Social Compu ing &
Ne wo king (ISPA/BDCloud/SocialCom/Sus ainCom), Dec. 2023, pp. 328–335.
[7] W.-C. Lin, T. Deakin, and S. McIn osh-Smi h, “E alua ing ISO C++ pa allel algo i hms on he e ogeneous HPC
sys ems,” in 2022 IEEE/ACM In e na ional Wo kshop on Pe o mance Modeling, Benchma king and Simula ion o
High Pe o mance Compu e Sys ems (PMBS). IEEE, 2022, pp. 36–47.
[8] R. A. Ba olomeu, R. Hal e , J. H. Meinke, and G. Su mann, “E ec o implemen a ions o he n-body p oblem on
he pe o mance and po abili y ac oss gpu endo s,” Fu u e Gene a ion Compu e Sys ems, ol. 169, p. 107802,
2025. [Online]. A ailable: h ps://www.sciencedi ec .com/science/a icle/pii/S0167739X25000974
[9] K. Asano ic, R. Bodik, B. C. Ca anza o, J. J. Gebis, P. Husbands, K. Keu ze , D. A. Pa e son, W. L. Plishke , J. Shal ,
S. W. Williams, and K. A. Yelick, “The landscape o pa allel compu ing esea ch: A iew om Be keley,” Elec ical
Enginee ing and Compu e Sciences, Uni e si y o Cali o nia a Be keley, Technical Repo No. UCB/EECS-2006-
183, Decembe , ol. 18, no. 2006-183, p. 19, 2006.
[10] C. R. T o , D. Leb un-G andié, D. A nd , J. Ciesko, V. Dang, N. Ellingwood, R. Gaya i, E. Ha ey, D. S. Hollman,
D. Ibanez, N. Libe , J. Madsen, J. Miles, D. Poliako , A. Powell, S. Rajamanickam, M. Simbe g, D. Sunde land,
B. Tu cksin, and J. Wilke, “Kokkos 3: P og amming model ex ensions o he exascale e a,” IEEE T ansac ions on
Pa allel and Dis ibu ed Sys ems, ol. 33, no. 4, pp. 805–817, 2022.
[11] J. H. Da is, P. Si a aman, I. Minn, K. Pa asy is, H. Menon, G. Geo gakoudis, and A. Bha ele, “An E alua i e Com-
pa ison o Pe o mance Po abili y ac oss GPU P og amming Models,” a Xi p ep in a Xi :2402.08950, 2024.
[12] W.-C. Lin, S. McIn osh-Smi h, and T. Deakin, “P elimina y epo : Ini ial e alua ion o S dPa implemen a ions
on AMD GPUs o HPC,” Jan. 2024.
[13] A. A nold, F. Fah enbe ge , C. Holm, O. Lenz, M. Bol en, H. Dachsel, R. Hal e , I. Kabadshow, F. Gähle ,
F. Hebe , J. Ise inghausen, M. Ho mann, M. Pippig, D. Po s, and G. Su mann, “Compa ison o scalable
as me hods o long- ange in e ac ions,” Phys. Re . E, ol. 88, p. 063308, Dec 2013. [Online]. A ailable:
h ps://link.aps.o g/doi/10.1103/PhysRe E.88.063308
[14] M. Dese no and C. Holm, “How o mesh up ewald sums. i. a heo e ical and nume ical compa ison o a ious
pa icle mesh ou ines,” The Jou nal o Chemical Physics, ol. 109, no. 18, pp. 7678–7693, 11 1998. [Online].
A ailable: h ps://doi.o g/10.1063/1.477414
[15] R. W. Hockney and J. W. Eas wood, Compu e simula ion using pa icles. c c P ess, 2021.
[16] J. J. Ce dà, V. Ballenegge , O. Lenz, and C. Holm, “P3m algo i hm o dipola in e ac ions,” The Jou nal o
Chemical Physics, ol. 129, no. 23, p. 234104, 12 2008. [Online]. A ailable: h ps://doi.o g/10.1063/1.3000389
[17] S. Sla e y, S. T. Ree e, C. Junghans, D. Leb un-G andié, R. Bi d, G. Chen, S. Foge y, Y. Qiu, S. Schulz,
A. Scheinbe g, A. Isne , K. Chong, S. Moo e, T. Ge mann, J. Belak, and S. Mniszewski, “Cabana: A pe o mance
po able lib a y o pa icle-based simula ions,” Jou nal o Open Sou ce So wa e, ol. 7, no. 72, p. 4115, 2022.
[Online]. A ailable: h ps://doi.o g/10.21105/joss.04115