Ci a ion: Vysocky, O.; Holze , M.;
S a elbach, G.; Va ik, R.; Riha, L.
Ene gy-E icien Implemen a ion o
he La ice Bol zmann Me hod.
Ene gies 2024,17, 502. h ps://
doi.o g/10.3390/en17020502
Academic Edi o : O lando Ayala
Recei ed: 29 No embe 2023
Re ised: 8 Janua y 2024
Accep ed: 16 Janua y 2024
Published: 19 Janua y 2024
Copy igh : © 2024 by he au ho s.
Licensee MDPI, Basel, Swi ze land.
This a icle is an open access a icle
dis ibu ed unde he e ms and
condi ions o he C ea i e Commons
A ibu ion (CC BY) license (h ps://
c ea i ecommons.o g/licenses/by/
4.0/).
ene gies
A icle
Ene gy-E icien Implemen a ion o he La ice Bol zmann Me hod
Ond ej Vysocky 1,* , Ma kus Holze 2,3 , Gab iel S a elbach 3, Radim Va ik 1and Lubomi Riha 1
1IT4Inno a ions Na ional Supe compu ing Cen e , VŠB—Technical Uni e si y o Os a a,
708 00 Os a a-Po uba, Czech Republic; [email p o ec ed] (R.V.); lubomi [email p o ec ed] (L.R.)
2
Chai o Sys em Simula ion, F ied ich-Alexande -Uni e si a E langen-Nu nbe g, 91058 E langen, Ge many;
[email p o ec ed]
3CERFACS, 31057 Toulouse Cedex 1, F ance; gab iel.s a [email p o ec ed]
*Co espondence: ond [email p o ec ed]
Abs ac : Ene gy cos s a e now one o he leading c i e ia when p ocu ing new compu ing ha dwa e.
Un il ecen ly, de elope s and use s ocused only on pu e pe o mance in e ms o ime- o-solu ion.
Recen ad ances in ene gy-awa e un ime sys ems ende he op imiza ion o bo h un ime and
ene gy- o-solu ion possible by including ha dwa e uning depending on he applica ion’s wo kload.
This wo k p esen s he impac ha ene gy-sensi i e uning s a egies ha e on a s a e-o - he-a
high-pe o mance compu ing code based on he la ice Bol zmann app oach called WALBERLA. We
e alua e bo h CPU-only and GPU-accele a ed supe compu e s. This pape demons a es ha , wi h
li le use in e en ion, when using he ene gy-e icien un ime sys em called MERIC, i is possible
o sa e a signi ican amoun o ene gy while main aining pe o mance.
Keywo ds: HPC; GPU accele a o s; DVFS; MERIC; ene gy-awa e un ime sys em; dynamic esou ce
managemen
1. In oduc ion
Wi h he inc easing challenges in de eloping as e ha dwa e, he indus y has shi ed
i s ocus om a pu e pe o mance pe spec i e, as dic a ed by Moo e’s law, o he me ic
o pe o mance pe Wa . This pa adigm shi , in oduced by In el in he mid-2000s wi h
hei i s mul i-co e p ocesso s, p io i ized ene gy e iciency and powe consump ion as
pi o al ac o s in new chip design. Despi e he adop ion o pe o mance pe Wa me ics by
a ious endo s, simila o he heo e ical peak pe o mance, hese me ics a e o en based
on undisclosed and nons anda dized benchma ks. Consequen ly, hey do no accu a ely
e lec he ue powe consump ion o an applica ion. Mo eo e , he di e se ways in which
applica ions u ilize s anda dized ha dwa e make i essen ial o cus omize de aul p ocesso
se ings o enhance pe o mance pe Wa on a pe -applica ion basis.
In his s udy, ou ocus lies in op imizing he ene gy consump ion o he la ice Bol z-
mann me hod (LBM)-based massi ely pa allel mul iphysics amewo k WALBERLA [
1
].
WALBERLA s ands as con empo a y open-sou ce C
++
so wa e designed o ha ness he ull
po en ial o la ge-scale supe compu e s o add ess in ica e esea ch ques ions in he a ea
o Compu a ional Fluid Dynamics (CFD). The WALBERLA is one o Eu oHPC Cen e o
Excellence o Exascale CFD (CEEC) [
2
] applica ions. The amewo k de elopmen p ocess
p io i izes pe o mance and e iciency, leading o s a egic choices such as ully dis ibu ed
da a s uc u es on an oc ee o blocks. Each da a block con ains in o ma ion only abou i sel
and i s nea es neighbo s, allowing e icien dis ibu ion ac oss supe compu e s h ough
he Message Passing In e ace (MPI) [1,3,4].
Op imizing ha dwa e e iciency begins a he indi idual chip and co e le el, necessi a -
ing low-le el a chi ec u e-speci ic op imiza ions, like ec o iza ion wi h Single Ins uc ion,
Mul iple Da a (SIMD) ins uc ions. Challenges escala e wi h code po ing o accele a o s,
such as GPUs, demanding compa ibili y adjus men s. WALBERLA add esses his com-
plexi y h ough me a-p og amming echniques wi hin he lbmpy and pys encils Py hon
Ene gies 2024,17, 502. h ps://doi.o g/10.3390/en17020502 h ps://www.mdpi.com/jou nal/ene gies
Ene gies 2024,17, 502 2 o 15
amewo ks [
5
–
8
]. These echniques enable he o mula ion o algo i hms in a symbolic
o m close o a ma hema ical ep esen a ion. Subsequen ly, au oma ed p ocesses handle
disc e iza ion and he gene a ion o low-le el C code, subs an ially ele a ing he le el o
abs ac ion and sepa a ion o conce ns.
LBM-based applica ions a e in e es ing o analyze o hei dynamic beha io —in
gene al, e e y sol e i e a ion consis s o wo phases wi h di e en equi emen s on
ha dwa e esou ces. Calo e e al. analyzed hese ke nels on a ious ha dwa e a chi ec u es
o possible ene gy sa ings [
9
,
10
] bu using a e y simple C code [
11
]. We build on hei
indings, especially he e ec i e usage o an ene gy-e icien un ime sys em [12].
This pape p o ides an o e iew o he WALBERLA amewo k, elucida ing i s heo-
e ical unde pinnings and echnologies. We in eg a e his unde s anding wi h pe o mance
uning, pinpoin ing scena ios conduci e o powe e iciency gains while minimizing he
impac on he ime- o-solu ion o bo h CPU and GPU ha dwa e con igu a ions.
2. La ice Bol zmann Me hod—Theo e ical Backg ound
The la ice Bol zmann me hod is a mesoscopic app oach si ua ed be ween mac oscopic
solu ions o he Na ie –S okes equa ions (NSEs) and mic oscopic me hods. I s o igins can
be aced back o an ex ension o la ice gas au oma a; howe e , mo e mode nly, he heo y
is de i ed by disc e izing he Bol zmann equa ion [
13
,
14
]. F om his, he la ice Bol zmann
equa ion (LBE) eme ges ha can be s a ed as:
i(x+ci∆ , +∆ )= i(x, )+Ωi(x, ).
I desc ibes he e olu ion o a local pa icle dis ibu ion unc ion (PDF)
wi h
q
-en ies
s o ed in each la ice si e. Typically, he g id is a
d
-dimensional Ca esian la ice wi h g id
spacing
∆x∈R+
, gi ing he me hod i s name. The PDF ec o desc ibes he p obabili y o
a i ual luid pa icle in posi ion
x∈Rd
and ime,
∈R+
a eling wi h disc e e la ice
eloci y
ci∈∆x/∆ {−
1,0,1
}d
[
14
]. Thus, ins ead o acking indi idual eal exis ing
pa icles as mic oscopic app oaches do, ensembles o i ual pa icles a e simula ed in he
LBM app oach.
The LBM can be sepa a ed in a s eaming s ep, whe e PDFs a e ad ec ed acco ding o
hei eloci ies, and a collision s ep ha ea anges he popula ion cell locally. Thus, in he
eme ging algo i hm, all nonlinea ope a ions a e cell local, while all nonlocal ope a ions
a e linea . This gi es he me hod i s algo i hmic simplici y and ease o he pa alleliza ion
p ocess. The collision ope a o
Ωi(x, )∈R
, o edis ibu ion o he PDFs, can be s a ed as
Ωi(x, )=T−1(T( ) + S( eq −T ( )))
whe e he PDFs a e ans o med o he collision space wi h a bijec i e mapping
T
[
8
]. In
he collision space, he collision is esol ed by sub ac ing he equilib ium o he PDFs
eq(ρ
,
u)∈Rq
om he PDFs. The e o e, each en y in he eme ging ec o co esponds
o di e en physical p ope ies. Thus, o model dis inc physical p ocesses, di e en
elaxa ion a es a e applied o each quan i y, which a e s o ed in a diagonal elaxa ion
ma ix
S
. Typically, each elaxa ion a e
ωi<
2
/∆
, he in e se o which is e e ed o as
he elaxa ion ime
τi=
1
/ωi
. Fo example, o eco e he co ec kinema ic iscosi y
ν
o a
luid, he elaxa ion ime o he co esponding collision quan i ies can be ob ained h ough
ν=c2
sτ−∆
2.
The basis o mos LBM o mula ions is he Maxwell–Bol zmann dis ibu ion ha
de ines he equilib ium s a e o he pa icles [14]
Ψ(ρ,u,c)=ρ1
2πc2
s3/2
exp −∥c−u∥2
2c2
s!,
Ene gies 2024,17, 502 3 o 15
whe e
ρ≡ρ(x, )
and
u≡u(x, )∈Rd
desc ibe he mac oscopic densi y and eloci y,
espec i ely. Fu he mo e, he speed o sound csis de ined as cs=√1/3 ∆x/∆ .
3. Code Gene a ion o LBM Ke nels
W i ing highly pe o man and lexible so wa e is a se e e challenge in many ame-
wo ks. On one side, he p oblem a ises in desc ibing he equa ions o sol e in a way ha
is close o he ma hema ical desc ip ion, while on he o he side, he code needs o be
specialized o di e en p ocessing uni s, like SIMD, and accele a o s, such as GPUs. In
he massi ely pa allel mul iphysics amewo k, WALBERLA his is sol ed by employing
me a-p og amming echniques. An o e iew o he app oach is depic ed in Figu e 1. A he
highes le el, he Py hon package lbmpy encapsula es he comple e symbolic ep esen a ion
o he la ice Bol zmann me hod. Fo his, he open-sou ce lib a y SymPy is used and
ex ended by [
15
]. This wo k low allows o he sys ema ic dissec ion o he LBM in o
i s cons i uen pa s, subsequen ly modula izing and s eamlining each s ep. Howe e ,
modula iza ion occu s di ec ly on he ma hema ical le el o o m a inal op imized upda e
ule. A de ailed desc ip ion o his p ocess can be ound in [
8
]. Finally, his leads o highly
specialized, p oblem-speci ic LBM compu e ke nels wi h minimal loa ing poin ope a ions
(FLOPs), all while main aining a ema kable deg ee o modula i y wi hin he sou ce code.
Py hon
C++
lbmpy
• LB gene ic and symbolic
• Highes abs ac ion le el
•
De i a ion o disc e ized
equa ions
pys encils
• pys encils IR
• Code ans o ma ions
•
Code gene a ion (CPU,
GPU)
WALBERLA
•
Massi ely pa allel mul i-
physics amewo k
• Domain decomposi ion
• Load balancing
• LB-Ke nel
• Bounda y Condi ions
• MPI Packing ke nels
MERIC
•
Ene gy consump ion mea-
su emen
• HW esou ces uning
• Applica ion acing
Figu e 1. O e iew o he so wa e s ack wi h lbmpy,pys encils,WALBERLA and so wa e uning. Wi h
he high-le el Py hon packages lbmpy and pys encils, he nume ical equa ions a e de i ed, disc e ized,
and, inally, lowe -le el C-Code is gene a ed om his symbolic ep esen a ion. The gene a ed code
can be combined wi h he C
++
amewo k WALBERLA and compiled. The execu able is op imized in
e ms o ene gy consump ion.
F om he symbolic desc ip ion, an Abs ac Syn ax T ee (AST) is cons uc ed wi hin
he pys encils In e mi ed Rep esen a ion (IR). This ee-based ep esen a ion inco po a es
a chi ec u e-speci ic AST nodes and poin e access in subsequen ke nels. Wi hin his ep e-
sen a ion, spa ial access pa icula s a e encapsula ed h ough pys encils ields. Addi ionally,
cons an exp essions o ixed special model pa ame e s can be di ec ly e alua ed o educe
he compu a ional o e head. Gi en ha he LBM compu e ke nel is symbolically de ined,
encompassing all ield da a accesses, he au oma ed de i a ion o compu e ke nels na u-
ally ex ends o encompass bounda y condi ions. This p ocess also in ol es he c ea ion o
ke nels o packing and unpacking. This sui e o ke nels plays a pi o al ole in popula ing
communica ion bu e s o MPI ope a ions.
Finally, he in e media e ep esen a ion o he compu e, bounda y and packing/un-
packing ke nels is p in ed by he C o he CUDA backend o pys encils o a clea ly de ined
in e ace. Each unc ion akes aw poin e s o a ay accesses oge he wi h hei ep esen a-
i e shape and s ide in o ma ion as well as all emaining ee pa ame e s. This simple and
Ene gies 2024,17, 502 4 o 15
consis en in e ace makes i possible o easily in eg a e he ke nels in exis ing C/C
++
so -
wa e s uc u es. Fu he mo e, wi h Py hon C-API, he low-le el ke nels can be mapped o
Py hon unc ions, which enables in e ac i e de elopmen by u ilizing lbmpy/pys encils as
s and-alone packages.
LBM is known o i s high memo y demand, and hus i was o en shown ha highly
op imized compu e ke nels a e only limi ed by he memo y bandwid h o a p ocesso o
accele a o [
6
,
7
]. Thus, na u ally, he ques ion a ises as o whe he i is possible o educe
he ene gy consump ion by educing he equency o he CPU compu e uni s (CPU co es)
while main aining he ull memo y subsys em pe o mance. Fu he mo e, he high le el o
op imiza ion employed o lbmpy leads o especially low FLOP numbe s in he ho spo o
he code [8].
4. Ene gy-Awa e Ha dwa e Tuning—Theo e ical Backg ound
Ene gy e iciency is commonly de ined as he pe o mance achie ed pe uni o powe
consump ion, ypically exp essed as loa ing poin ope a ions pe second pe Wa . How-
e e , when dealing wi h codes based on he la ice Bol zmann me hod (LBM), pe o mance
is be e cha ac e ized by he numbe o La ice Upda es execu ed Pe Second (LUPS). In
his s udy, we quan i y he ene gy e iciency o he WALBERLA applica ion as Millions o
La ice Upda es pe Second pe Wa (MLUPs/W).
To accu a ely measu e he ene gy consump ion o an applica ion, a high- equency
powe moni o ing sys em is impe a i e. This sys em should p o ide eal- ime powe o
ene gy consump ion eadings o he en i e compu a ional node o , a a minimum, i s key
compu a ional componen s.
The o al ene gy consumed (in Joules) can be calcula ed om powe samples (in Wa s)
ob ained a a speci ic sampling equency (in He z) as depic ed in he ollowing equa ion:
Ene gy( ) = Z
0Powe (x),dx≈∑n
i=0Powe Samplei
SamplingF equency.
The e a e wo undamen al app oaches o inc ease ene gy e iciency: (1) op imizing
applica ions o ully exploi compu a ional esou ces, ensu ing ha he wo kload aligns
wi h he uppe limi s de ined by he ha dwa e’s oo line model, o (2) judiciously limi ing
unused esou ces o p e en powe was age.
Mode n high-pe o mance CPUs and GPUs o e a leas one unable pa ame e
con ollable om he use space. Typically, i is he equency o compu a ion uni s (CPU
co es) which di ec ly impac s he peak pe o mance o he chip, and i is c ucial o compu e-
in ensi e compu ing asks. These pa ame e s can be adjus ed ei he s a ically o dynamically.
S a ic uning in ol es con igu ing speci ic ha dwa e se ings a he s a o an appli-
ca ion execu ion and main aining his con igu a ion un il i s comple ion. Howe e , such
s a ic se ups a e a ely op imal o complex applica ions, leading o subop imal ene gy
sa ings. S a ic uning lacks adap abili y o wo kload changes du ing applica ion execu ion,
hinde ing he achie emen o maximum a ailable e iciencies.
In con as , dynamic uning adjus s pa ame e s con inuously du ing applica ion
execu ion. This unc ionali y is achie ed by ene gy-awa e un ime sys ems ha can iden i y
op imal se ings o di e en phases o he applica ion and modi y ha dwa e con igu a ion.
One such sys em is COUNTDOWN [
16
], main ained by CINECA and he Uni e si y
o Bologna. COUNTDOWN dynamically scales CPU co e equency du ing he MPI com-
munica ion and synch oniza ion phases, while ensu ing ha he applica ion’s pe o mance
is p ese ed.
The Ba celona Supe compu ing Cen e de elops EAR [
17
], a lib a y ha i e a i ely
adjus s he CPU co e equency o powe cap based on bina y ins uc ions and pe o mance
coun e s alues.
LLNL Conduc o [
18
] employs a powe limi app oach, iden i ying c i ical communi-
ca ion pa hs and alloca ing mo e powe o slowe p ocesses o educe wai ing imes, hus
enhancing o e all pe o mance. Simila ly, LLNL Unco e Powe Sca enge [
19
] dynami-
Ene gies 2024,17, 502 5 o 15
cally unes he In el CPU con igu a ion by sampling RAPL DRAM powe consump ion
and ins uc ions pe cycle a ia ion. Op imal ene gy sa ings a e achie ed wi h a 200ms
sampling in e al.
Fu he mo e, he Run ime Exploi a ion o Applica ion Dynamism o Ene gy-e icien
eXascale compu ing (READEX) p ojec [
20
] in oduced a dynamic uning me hodology [
21
]
and i s implemen a ion. The ools de eloped in his p ojec p o ide HPC applica ion de el-
ope s wi h ways o exploi he dynamic beha io o hei applica ions. The me hodology is
based on he assump ion ha each egion o an applica ion may equi e speci ic ha dwa e
con igu a ions. The READEX app oach iden i ies hese equi emen s o each egion and
dynamically adjus s he ha dwa e con igu a ion when en e ing a egion. MERIC [
22
], an
implemen a ion o he READEX app oach de eloped a IT4Inno a ions, de ines a min-
imum un ime o he egion as 100ms o ensu e eliable ene gy measu emen s and o
accommoda e la ency when changing ha dwa e con igu a ions.
These ools a e able o b ing signi ican ene gy sa ings wi hou o wi h limi ed pe -
o mance penal y. Howe e , hese ools a e designed o wo k well on non-accele a ed
machines, and hey do no une GPU pa ame e s. Howe e , in mode n accele a ed HPC
clus e s, GPUs consume he majo i y o he compu e node ene gy. The ene gy e iciency o
da a cen e GPUs is signi ican ly highe when compa ed o ha o se e CPUs. This is
con i med by he ac ha all he op- anked supe compu e s in G een500 [
23
] ( he lis o
he mos ene gy-e icien HPC sys ems) a e based on N idia o AMD GPUs. Taking all his
in o accoun , i is s ill possible o imp o e he ene gy e iciency o hese GPUs by a ound
ens o pe cen as p esen ed in [24].
A su ey o GPU ene gy-e iciency analysis and op imiza ion echniques [
25
] e e s
o a ious app oaches o iden i y he op imal con igu a ion o he GPU equency o
ob ain ene gy sa ings, bu in each lis ed case, he pape p esen s a single con igu a ion
o he whole execu ion o he applica ion, which we e e o as s a ic uning. Simila ly,
K aljic e al. [
26
] iden i ied he execu ion phases o an analyzed applica ion by sampling
GPU ene gy consump ion. Ghazan a e al. [
27
] ained a neu al ne wo k model o iden i y
he op imal GPU SM equency. Howe e , in all cases abo e, he au ho s did no y o
change he con igu a ion dynamically du ing he execu ion o an applica ion.
5. Resul s o he waLBe la Ene gy Consump ion Op imiza ion
Fo s udying he ene gy consump ion o WALBERLA in an indus ially ele an es
case, he so-called LAGOON (LAnding-Gea nOise da abase o CAA alida iON), which
is simpli ied Ai bus plane landing gea , was chosen [
28
,
29
]. The es se up consis s o
he LAGOON geome y (see Figu e 2) in a i ual wind unnel wi h a uni o m esolu ion
o 512
×
512
×
512 cells. Fo he in low wall bounce back, bounda y condi ions wi h an
in low eloci y o 0.05 in la ice uni s we e used, while he ou low was modeled wi h
non- e lec ing ou low bounda ies. Fo he pu pose o he ene gy uning, he benchma k
case was simula ed o 100 ime s eps.
Benchma king and pe o mance and ene gy measu emen s we e pe o med a he ol-
lowing machines o he IT4Inno a ions supe compu ing cen e : (1) Ba bo a non-accele a ed
pa i ion, and (2) Ka olina GPU accele a ed pa i ion.
The Ba bo a sys em is equipped wi h wo In el Xeon Gold 6240 CPUs (codename
Cascade lake) pe node. Each CPU has 18 co es ( he hype - h eading is disabled) and is
designed o wo k a 150W TDP. The nominal equency o he CPU is 2.6GHz, bu i can
each up o (i) 3.9GHz o he u bo equency when only wo co es a e ac i e o (ii) 3.3GHz
o he u bo equency when all co es a e ac i e and execu e SSE ins uc ions. The CPU
co e equency can be educed all he way o 1.1GHz by a use o ope a ing sys em. Since
he Nehalem a chi ec u e, In el has been using ‘unco e’ o e e o he equency o he
subsys ems in he physical p ocesso package ha a e sha ed by mul iple p ocesso co es
e.g., las -le el cache, on-chip ing in e connec o in eg a ed memo y con olle s. Unco e
egions o e all occupy app oxima ely 30% o a chip a ea [
30
]. While he co e equency is
c i ical o compu e-bound egions, he memo y-bound egions a e much mo e sensi i e
Ene gies 2024,17, 502 6 o 15
o unco e equency [
31
]. The Ba bo a CPUs can scale he unco e equency be ween 1.2
and 2.4 GHz.
Figu e 2. Geome y o he LAGOON plane landing gea (le ). The geome y was s udied in a i ual
wind unnel ( igh ).
Ba bo a’s compu a ional nodes a e equipped by he on-boa d A os|Bull High De ini-
ion Ene gy E icien Moni o ing (HDEEM) sys em [
32
], which eads powe consump ion
om he mainboa d ha dwa e senso s and s o es he da a o dedica ed memo y. The senso
ha moni o s he consump ion o he whole node p o ides 1000 powe samples pe second,
and he es o he senso s ha moni o he compu e node sub-uni s p o ide 100 samples
pe second. Bo h agg ega ed alues and powe samples can be ead om he use space
using a dedica ed lib a y o command-line u ili y. Since Sandy B idge gene a ion, In el p o-
cesso s ha e in eg a ed a Running A e age Powe Limi (RAPL) ha dwa e powe con olle
ha p o ides a powe measu emen and mechanism o limi he powe consump ion o
se e al domains o he CPU [
33
]. In el RAPL con ols he CPU co e and unco e equencies
o keep he a e age powe consump ion o he CPU package below he TDP. The In el
RAPL in e ace allows a educ ion in his powe limi bu no an inc ease.
Ka olina clus e ( ank 71. in Top500 11/2021, 8. in G een500 11/2021 [
34
]) nodes a e
equipped wi h wo AMD EPYC 7763 CPUs, and eigh N idia A100-SXM4 GPUs. One MPI
p ocess pe GPU is used, ou pe CPU. To imp o e he ene gy e iciency o he applica ion,
we ha e speci ied he equency o he GPU s eaming mul ip ocesso s (SMs), which is
analogous o he CPU co e equency. Fo his pu pose, we use he N idia Managemen
Lib a y (NVML), which p o ides he unc ion n mlDe iceSe Applica ionsClocks() ha se s a
speci ic clock speed o a a ge GPU o bo h (i) memo y and (ii) s eaming mul ip ocesso s.
Howe e , he A100-SXM4 uses HBM2 memo y, whose equency canno be uned as i is
possible in he case o he GDDR memo y. The e o e, on da a-cen e -g ade GPUs, like A100,
only he equency o SMs can be con olled.
The ene gy consump ion o he GPU-accele a ed applica ion execu ions was measu ed
using pe o mance coun e s o he GPU (accessed using he N idia Managemen Lib a y)
and CPU (AMD RAPL wi h a simila powe moni o ing in e ace as In el RAPL wi hou
he suppo o powe capping).
To imp o e he ene gy e iciency o he GPU-accele a ed execu ions o he WALBERLA,
we pe o med he s a ic uning o he GPUs. The MERIC un ime sys em suppo s he
dynamic uning o GPU-accele a ed applica ions based on CPU egions only i he GPU
wo kload is synch onized. This limi a ion comes om he equi emen ha he GPU
equency canno be con olled om a ke nel. I is he CPU ha c ea es he eques o
change he equency h ough he GPU d i e .
By de aul , when unning a wo kload, he A100-SXM4 GPU (400W TDP) uses he
maximum u bo equency o 1.410GHz (i no o ced o educe he equency by he
powe consump ion exceeding he powe limi o by he mal h o ling) and swi ches o
he nominal equency o he GPU (1.095GHz) when copying he da a o/ om he GPU
memo y. The equency can be educed o 210 MHz in 81 s eps.
Du ing he execu ion o he CUDA ke nels on he GPUs, we also e alua ed he impac
o he CPU co e equency uning. The nominal equency o he AMD EPYC 7763 (280W
Ene gies 2024,17, 502 7 o 15
TDP) is 2.45 GHz, while he CPU can un up o 3.525 GHz boos equency. To educe
he numbe o es s, he 100 MHz s ep was used ins ead o he 25 MHz s ep, which is he
highes esolu ion suppo ed.
To con ol he ha dwa e pa ame e s men ioned abo e and o measu e he esou ce
consump ion o he execu ed applica ion, we used he MERIC un ime sys em. In he case
o he non-accele a ed e sion o WALBERLA, bo h s a ic (single ha dwa e con igu a ion o
he en i e applica ion execu ion) and dynamic uning (a speci ic ha dwa e con igu a ion o
each pa o he applica ion) we e used. In he case o he GPU-accele a ed e sion, s a ic
uning was used only since he MERIC does no ha e suppo o iden i y which CUDA
ke nel is unning on a GPU. A un ime sys em wi h suppo o dynamic GPU uning is
s ill a wo k in p og ess.
5.1. S a ic Tuning o he CPU Pa ame e s
The WALBERLA ene gy e iciency analysis s a ed wi h he s a ic uning o i s non-
accele a ed e sion. We pe o med an exhaus i e s a e–space sea ch, es ing all possible
CPU co e and unco e equency con igu a ions, using he 0.2GHz s ep o bo h he CPU
co e and he unco e equency. The lowes equencies we e omi ed since one can expec
high pe o mance penal y in hese con igu a ions.
Table 1(pe o mance penal y), Table 2(HDEEM ene gy sa ings) and Table 3(In el
RAPL ene gy sa ings) show he consump ion o esou ces o he WALBERLA sol e in
a ious con igu a ions, using colo coding o indica e which alues a e be e (g een) o
wo se ( ed).
F om all he e alua ed con igu a ions, he highes ene gy sa ings based on he HDEEM
measu emen s a e 22.8%. These we e eached o he ollowing con igu a ion: CF 1.9 GHz
and UCF 1.8 GHz. Fo he same con igu a ion, he sa ings calcula ed om he RAPL
measu emen s a e 29.6%. Howe e , in his con igu a ion, he pe o mance d ops by abou
13.9%.The majo di e ence in ene gy sa ings be ween HDEEM and RAPL comes om he
se o powe domains moni o ed by hem. RAPL only moni o s he powe consump ion
o he CPU, which is he only componen ha b ings ene gy sa ings due o uning. The
powe consump ion o he emaining node componen s emains unchanged, which has a
majo impac on ene gy sa ings i he un ime is ex ended. Since HDEEM moni o s he
powe consump ion o he en i e node, i s esul s a e mo e ep esen a i e.
S a ic uning usually p o ides a limi ed possibili y o ob ain majo ene gy sa ings wi h-
ou a pe o mance penal y. WALBERLA eached 10.2% ene gy sa ings based on HDEEM
measu emen s and 12.1% ene gy sa ings based on RAPL measu emen s a a cos o 1.6%
pe o mance deg ada ion when he co e equency was educed o 2.8 GHz and he unco e
equency emained wi hou any limi a ions.
The pe o mance impac o he e alua ed CPU equencies con igu a ions on he
WALBERLA sol e is shown in Table 1. The espec i e ene gy sa ings a e in Table 2 o he
HDEEM measu emen s and in Table 3 o he In el RAPL measu emen s. Based on he
measu emen s in hese ables, we iden i ied con igu a ions ha cause up o 2, 5, 10 and
unlimi ed un ime ex ension while b inging he maximum possible ene gy sa ings based
on HDEEM measu emen s. The summa y o he esul s is in Table 4, which p esen s he
bes con igu a ions o a ious pe o mance penal y limi s. I also shows one hand-picked
con igu a ion which p o ides 19.6% ene gy sa ings based on HDEEM measu emen s while
ex ending he un ime only abou 6.9%.
Table 1. Impac o he s a ic uning on he o e all un ime in [%] o WALBERLA when unning he
Lagoon use case.
co e[GHz]
unco e[GHz] 1.9 2.1 2.3 2.5 2.6 2.7 2.8 2.9 3.0 3.1 3.2 3.3
1.6 −17.46 −15.91 −14.6 −13.68 −14.2 −13.06 −13.93 −14.35 −11.71 −11.49 −11.26 −10.82
1.8 −13.9 −11.75 −10.35 −9.97 −9.24 −8.83 −8.45 −8.34 −7.64 −8.01 −7.32 −6.93
2−10.89 −10.03 −7.79 −6.86 −6.37 −5.78 −5.57 −5.36 −4.79 −4.27 −4.44 −4.01
2.2 −8.87 −6.88 −5.85 −4.75 −4−3.84 −3.38 −3−2.94 −2.22 −1.85 −1.89
2.4 −7.3 −5.17 −4.24 −3.1 −2.81 −3.69 −1.59 −1.33 −1.06 −0.96 −0.27 −0.09
Ene gies 2024,17, 502 8 o 15
Table 2. Ene gy sa ings o di e en con igu a ions using HDEEM measu emen s in [%] o WAL-
BERLA when unning he Lagoon use case.
co e[GHz]
unco e[GHz] 1.9 2.1 2.3 2.5 2.6 2.7 2.8 2.9 3.0 3.1 3.2 3.3
1.6 22.61 21.39 19.78 16.63 15 14.29 11.81 9.22 9.53 7.07 4.86 2.45
1.8 22.84 22.28 20.84 17.3 16.63 15.36 14.18 12.18 10.59 7.68 6.02 3.56
2 21.95 20.51 19.62 16.64 15.79 14.75 13.35 11.59 9.92 8.14 5.71 3.21
2.2 20.19 19.61 17.74 15.15 14.47 13.14 11.99 10.49 8.14 6.54 4.72 1.82
2.4 17.69 17.26 15.54 12.89 11.81 9.56 10.17 8.33 6.36 4.18 2.68 −0.27
Table 3. Ene gy sa ings o di e en con igu a ions using In el RAPL measu emen s in [%] o
WALBERLA when unning he Lagoon use case.
co e[GHz]
unco e[GHz] 1.9 2.1 2.3 2.5 2.6 2.7 2.8 2.9 3.0 3.1 3.2 3.3
1.6 31.25 28.8 26.42 22.46 20.46 19.34 16.38 13.61 13.62 10.44 7.81 4.63
1.8 29.63 29.2 26.93 22.23 21.68 19.95 18.42 16.05 13.77 10.5 8.08 5.16
2 28.04 26.1 24.99 21 19.96 18.57 16.67 14.28 12.38 10.06 7.37 4.34
2.2 25.4 24.6 21.84 18.77 17.82 16.08 14.58 12.92 9.88 7.81 5.84 2.36
2.4 22.35 21.28 19.48 15.61 14.42 11.66 12.06 9.88 7.69 4.97 3.33 −0.32
Table 4. Resul s o he s a ic uning o he WALBERLA sol e when unning he Lagoon use case.
The able p esen s ene gy sa ings om bo h he HDEEM and In el RAPL ene gy measu emen s o
a ious pe o mance deg ada ion ade-o s. In addi ion, one hand-picked con igu a ion ha b ings
meaning ul ene gy sa ings wi h modes pe o mance penal y is also p esen ed.
−2% −5% −10% No Selec ed
Limi Limi Limi Limi Con igu a ion
Run ime [%] −1.6 −4.8 −8.9 −13.9 −6.9
HDEEM [%] 10.2 15.2 20.2 22.8 19.6
RAPL [%] 12.1 18.8 25.4 29.6 24.6
CF; UCF [GHz] 2.8; 2.4 2.5; 2.2 1.9; 2.2 1.9; 1.8 2.1, 2.2
Finally, Figu e 3shows he powe consump ion imeline o he en i e compu e node
(blade) and selec ed node componen s. Samples we e collec ed by HDEEM du ing he
execu ion o he Lagoon es case. One can see dynamic changes in powe consump ion,
which indica es ha a s a ic ha dwa e con igu a ion is no op imal o he whole applica ion
un because he ha dwa e equi emen s change o e ime.
5.2. Dynamic Tuning o he CPU Pa ame e s
This sec ion p esen s a WALBERLA analysis using dynamic uning, which se s ha d-
wa e con igu a ions ha bes sui each ins umen ed sec ion o he code. MERIC suppo s
au oma ic bina y ins umen a ion, which gene a es a copy o he applica ion execu able
bina y ile ha includes he MERIC API calls a he beginning and he end o all selec ed
egions. Due o he excep ion o usage in he WALBERLA code, i was no possible o use
ully au oma ic bina y ins umen a ion because he execu ion hen esul ed in a un ime
e o o uncaugh excep ions. We manually ins umen ed he WALBERLA sou ce code wi h
he MERIC unc ion calls, which esul ed in less ine-g ain ins umen a ion han would
be possible wi h ull bina y ins umen a ions. Despi e he ac ha he ins umen a ion
consis s o only eigh egions, which is no op imal, we we e able o co e 99% o he
applica ion un ime.
In he case o he iden i ica ion o an op imal dynamic con igu a ion, we also execu ed
he applica ion in a ious ha dwa e con igu a ions, while o each ins umen ed egion,
we iden i ied i s op imal con igu a ion. The s a e–space sea ch was pe o med wice—one
o ob ain a con igu a ion ha does no cause any pe o mance deg ada ion, and ano he
one o b ing maximum HDEEM ene gy sa ings wi hou any un ime ex ension limi a ion.
Ene gies 2024,17, 502 9 o 15
Figu e 3. Powe imeline based on HDEEM measu emen s o WALBERLA when unning he Lagoon
use case. The da a a e collec ed o en i e compu e node (blade), and i s componen s ( wo CPUs and
wo g oups o memo y channels). The ime windows shows one synch oniza ion phase ollowed by
he six sol e i e a ions.
Table 5compa es ou di e en execu ions o WALBERLA, he de aul ha dwa e con-
igu a ion, he comp omise s a ic con igu a ion wi h 2.1 GHz co e and 1.8 GHz unco e
equency, and wo execu ions using dynamic uning wi h and wi hou pe o mance
penal y. The WALBERLA execu ion Lagoon con igu a ion has changed o show ha hese
pa ame e s ha e a majo impac on s a ic execu ion op imal con igu a ion and wha sa -
ings i b ings. In con as o he p e ious sec ion, he e, we p esen alues o he whole
applica ion un ime because he dynamic uning op imized he whole un ime. In his
case, he sol e akes 3/4 o he un ime. While in he p e ious Sec ion 5.1, we p esen a
un ime ex ension o abou 12.6% in he comp omise s a ic con igu a ion, now he same
con igu a ion ex ends he un ime by jus abou 5.8%. The p oblem ha each execu ion
con igu a ion may esul in a di e en op imal s a ic con igu a ion is sol ed by using
dynamic uning since each egion o he applica ion may ake a di e en ime; howe e ,
he egions’ ha dwa e equi emen s a e he same, and hus he op imal con igu a ion is
he same.
Table 5shows ha dynamic uning can achie e highe ene gy sa ings han s a ic
uning. The dynamically uned execu ion o WALBERLA consumed 7.9% less ene gy
wi hou ex ending he un ime. The highes ene gy sa ings achie ed wi h dynamic uning
is 19.1% a a cos o ex ending he un ime by 16.2%.
Please no e ha in his case, he sol e consis s o a single egion only. In Figu e 3, i
is isible ha he sol e should be spli in o a leas wo di e en egions because hese
egions ha e di e en ha dwa e equi emen s. Howe e , hei un ime is e y sho (up
o 5ms), while MERIC equi es egions o a leas 100ms. In es iga ing possible ene gy
sa ings gained om he dynamic uning o he sol e is ou goal when a new elease o he
MERIC ha b ings be e suppo o ine-g ain uning becomes a ailable.