scieee Science in your language
[en] (orig)

Energy-efficient implementation of the lattice Boltzmann method

Abstract

Energy costs are now one of the leading criteria when procuring new computing hardware. Until recently, developers and users focused only on pure performance in terms of time-to-solution. Recent advances in energy-aware runtime systems render the optimization of both runtime and energy-to-solution possible by including hardware tuning depending on the application’s workload. This work presents the impact that energy-sensitive tuning strategies have on a state-of-the-art high-performance computing code based on the lattice Boltzmann approach called WALBERLA. We evaluate both CPU-only and GPU-accelerated supercomputers. This paper demonstrates that, with little user intervention, when using the energy-efficient runtime system called MERIC, it is possible to save a significant amount of energy while maintaining performance.

Read accessible full text

Energy-efficient implementation of the lattice Boltzmann method

Author: Vysocký, Ondřej
Publisher: MDPI
Year: 2024
DOI: 10.3390/en17020502
Source: https://dspace.vsb.cz/bitstreams/62344931-77c7-477d-897b-13f0e06afc41/download
Ci a ion: Vysocky, O.; Holze , M.;
S a elbach, G.; Va ik, R.; Riha, L.
Ene gy-E icien Implemen a ion o
he La ice Bol zmann Me hod.
Ene gies 2024,17, 502. h ps://
doi.o g/10.3390/en17020502
Academic Edi o : O lando Ayala
Recei ed: 29 No embe 2023
Re ised: 8 Janua y 2024
Accep ed: 16 Janua y 2024
Published: 19 Janua y 2024
Copy igh : © 2024 by he au ho s.
Licensee MDPI, Basel, Swi ze land.
This a icle is an open access a icle
dis ibu ed unde he e ms and
condi ions o he C ea i e Commons
A ibu ion (CC BY) license (h ps://
c ea i ecommons.o g/licenses/by/
4.0/).
ene gies
A icle
Ene gy-E icien Implemen a ion o he La ice Bol zmann Me hod
Ond ej Vysocky 1,* , Ma kus Holze 2,3 , Gab iel S a elbach 3, Radim Va ik 1and Lubomi Riha 1
1IT4Inno a ions Na ional Supe compu ing Cen e , VŠB—Technical Uni e si y o Os a a,
708 00 Os a a-Po uba, Czech Republic; [email p o ec ed] (R.V.); lubomi [email p o ec ed] (L.R.)
2
Chai o Sys em Simula ion, F ied ich-Alexande -Uni e si a E langen-Nu nbe g, 91058 E langen, Ge many;
[email p o ec ed]
3CERFACS, 31057 Toulouse Cedex 1, F ance; gab iel.s a [email p o ec ed]
*Co espondence: ond [email p o ec ed]
Abs ac : Ene gy cos s a e now one o he leading c i e ia when p ocu ing new compu ing ha dwa e.
Un il ecen ly, de elope s and use s ocused only on pu e pe o mance in e ms o ime- o-solu ion.
Recen ad ances in ene gy-awa e un ime sys ems ende he op imiza ion o bo h un ime and
ene gy- o-solu ion possible by including ha dwa e uning depending on he applica ion’s wo kload.
This wo k p esen s he impac ha ene gy-sensi i e uning s a egies ha e on a s a e-o - he-a
high-pe o mance compu ing code based on he la ice Bol zmann app oach called WALBERLA. We
e alua e bo h CPU-only and GPU-accele a ed supe compu e s. This pape demons a es ha , wi h
li le use in e en ion, when using he ene gy-e icien un ime sys em called MERIC, i is possible
o sa e a signi ican amoun o ene gy while main aining pe o mance.
Keywo ds: HPC; GPU accele a o s; DVFS; MERIC; ene gy-awa e un ime sys em; dynamic esou ce
managemen
1. In oduc ion
Wi h he inc easing challenges in de eloping as e ha dwa e, he indus y has shi ed
i s ocus om a pu e pe o mance pe spec i e, as dic a ed by Moo e’s law, o he me ic
o pe o mance pe Wa . This pa adigm shi , in oduced by In el in he mid-2000s wi h
hei i s mul i-co e p ocesso s, p io i ized ene gy e iciency and powe consump ion as
pi o al ac o s in new chip design. Despi e he adop ion o pe o mance pe Wa me ics by
a ious endo s, simila o he heo e ical peak pe o mance, hese me ics a e o en based
on undisclosed and nons anda dized benchma ks. Consequen ly, hey do no accu a ely
e lec he ue powe consump ion o an applica ion. Mo eo e , he di e se ways in which
applica ions u ilize s anda dized ha dwa e make i essen ial o cus omize de aul p ocesso
se ings o enhance pe o mance pe Wa on a pe -applica ion basis.
In his s udy, ou ocus lies in op imizing he ene gy consump ion o he la ice Bol z-
mann me hod (LBM)-based massi ely pa allel mul iphysics amewo k WALBERLA [
1
].
WALBERLA s ands as con empo a y open-sou ce C
++
so wa e designed o ha ness he ull
po en ial o la ge-scale supe compu e s o add ess in ica e esea ch ques ions in he a ea
o Compu a ional Fluid Dynamics (CFD). The WALBERLA is one o Eu oHPC Cen e o
Excellence o Exascale CFD (CEEC) [
2
] applica ions. The amewo k de elopmen p ocess
p io i izes pe o mance and e iciency, leading o s a egic choices such as ully dis ibu ed
da a s uc u es on an oc ee o blocks. Each da a block con ains in o ma ion only abou i sel
and i s nea es neighbo s, allowing e icien dis ibu ion ac oss supe compu e s h ough
he Message Passing In e ace (MPI) [1,3,4].
Op imizing ha dwa e e iciency begins a he indi idual chip and co e le el, necessi a -
ing low-le el a chi ec u e-speci ic op imiza ions, like ec o iza ion wi h Single Ins uc ion,
Mul iple Da a (SIMD) ins uc ions. Challenges escala e wi h code po ing o accele a o s,
such as GPUs, demanding compa ibili y adjus men s. WALBERLA add esses his com-
plexi y h ough me a-p og amming echniques wi hin he lbmpy and pys encils Py hon
Ene gies 2024,17, 502. h ps://doi.o g/10.3390/en17020502 h ps://www.mdpi.com/jou nal/ene gies
Ene gies 2024,17, 502 2 o 15
amewo ks [
5
–
8
]. These echniques enable he o mula ion o algo i hms in a symbolic
o m close o a ma hema ical ep esen a ion. Subsequen ly, au oma ed p ocesses handle
disc e iza ion and he gene a ion o low-le el C code, subs an ially ele a ing he le el o
abs ac ion and sepa a ion o conce ns.
LBM-based applica ions a e in e es ing o analyze o hei dynamic beha io —in
gene al, e e y sol e i e a ion consis s o wo phases wi h di e en equi emen s on
ha dwa e esou ces. Calo e e al. analyzed hese ke nels on a ious ha dwa e a chi ec u es
o possible ene gy sa ings [
9
,
10
] bu using a e y simple C code [
11
]. We build on hei
indings, especially he e ec i e usage o an ene gy-e icien un ime sys em [12].
This pape p o ides an o e iew o he WALBERLA amewo k, elucida ing i s heo-
e ical unde pinnings and echnologies. We in eg a e his unde s anding wi h pe o mance
uning, pinpoin ing scena ios conduci e o powe e iciency gains while minimizing he
impac on he ime- o-solu ion o bo h CPU and GPU ha dwa e con igu a ions.
2. La ice Bol zmann Me hod—Theo e ical Backg ound
The la ice Bol zmann me hod is a mesoscopic app oach si ua ed be ween mac oscopic
solu ions o he Na ie –S okes equa ions (NSEs) and mic oscopic me hods. I s o igins can
be aced back o an ex ension o la ice gas au oma a; howe e , mo e mode nly, he heo y
is de i ed by disc e izing he Bol zmann equa ion [
13
,
14
]. F om his, he la ice Bol zmann
equa ion (LBE) eme ges ha can be s a ed as:
i(x+ci∆ , +∆ )= i(x, )+Ωi(x, ).
I desc ibes he e olu ion o a local pa icle dis ibu ion unc ion (PDF)
wi h
q
-en ies
s o ed in each la ice si e. Typically, he g id is a
d
-dimensional Ca esian la ice wi h g id
spacing
∆x∈R+
, gi ing he me hod i s name. The PDF ec o desc ibes he p obabili y o
a i ual luid pa icle in posi ion
x∈Rd
and ime,
∈R+
a eling wi h disc e e la ice
eloci y
ci∈∆x/∆ {−
1,0,1
}d
[
14
]. Thus, ins ead o acking indi idual eal exis ing
pa icles as mic oscopic app oaches do, ensembles o i ual pa icles a e simula ed in he
LBM app oach.
The LBM can be sepa a ed in a s eaming s ep, whe e PDFs a e ad ec ed acco ding o
hei eloci ies, and a collision s ep ha ea anges he popula ion cell locally. Thus, in he
eme ging algo i hm, all nonlinea ope a ions a e cell local, while all nonlocal ope a ions
a e linea . This gi es he me hod i s algo i hmic simplici y and ease o he pa alleliza ion
p ocess. The collision ope a o
Ωi(x, )∈R
, o edis ibu ion o he PDFs, can be s a ed as
Ωi(x, )=T−1(T( ) + S( eq −T ( )))
whe e he PDFs a e ans o med o he collision space wi h a bijec i e mapping
T
[
8
]. In
he collision space, he collision is esol ed by sub ac ing he equilib ium o he PDFs
eq(ρ
,
u)∈Rq
om he PDFs. The e o e, each en y in he eme ging ec o co esponds
o di e en physical p ope ies. Thus, o model dis inc physical p ocesses, di e en
elaxa ion a es a e applied o each quan i y, which a e s o ed in a diagonal elaxa ion
ma ix
S
. Typically, each elaxa ion a e
ωi<
2
/∆
, he in e se o which is e e ed o as
he elaxa ion ime
τi=
1
/ωi
. Fo example, o eco e he co ec kinema ic iscosi y
ν
o a
luid, he elaxa ion ime o he co esponding collision quan i ies can be ob ained h ough
ν=c2
sτ−∆
2.
The basis o mos LBM o mula ions is he Maxwell–Bol zmann dis ibu ion ha
de ines he equilib ium s a e o he pa icles [14]
Ψ(ρ,u,c)=ρ1
2πc2
s3/2
exp −∥c−u∥2
2c2
s!,
Ene gies 2024,17, 502 3 o 15
whe e
ρ≡ρ(x, )
and
u≡u(x, )∈Rd
desc ibe he mac oscopic densi y and eloci y,
espec i ely. Fu he mo e, he speed o sound csis de ined as cs=√1/3 ∆x/∆ .
3. Code Gene a ion o LBM Ke nels
W i ing highly pe o man and lexible so wa e is a se e e challenge in many ame-
wo ks. On one side, he p oblem a ises in desc ibing he equa ions o sol e in a way ha
is close o he ma hema ical desc ip ion, while on he o he side, he code needs o be
specialized o di e en p ocessing uni s, like SIMD, and accele a o s, such as GPUs. In
he massi ely pa allel mul iphysics amewo k, WALBERLA his is sol ed by employing
me a-p og amming echniques. An o e iew o he app oach is depic ed in Figu e 1. A he
highes le el, he Py hon package lbmpy encapsula es he comple e symbolic ep esen a ion
o he la ice Bol zmann me hod. Fo his, he open-sou ce lib a y SymPy is used and
ex ended by [
15
]. This wo k low allows o he sys ema ic dissec ion o he LBM in o
i s cons i uen pa s, subsequen ly modula izing and s eamlining each s ep. Howe e ,
modula iza ion occu s di ec ly on he ma hema ical le el o o m a inal op imized upda e
ule. A de ailed desc ip ion o his p ocess can be ound in [
8
]. Finally, his leads o highly
specialized, p oblem-speci ic LBM compu e ke nels wi h minimal loa ing poin ope a ions
(FLOPs), all while main aining a ema kable deg ee o modula i y wi hin he sou ce code.
Py hon
C++
lbmpy
• LB gene ic and symbolic
• Highes abs ac ion le el
•
De i a ion o disc e ized
equa ions
pys encils
• pys encils IR
• Code ans o ma ions
•
Code gene a ion (CPU,
GPU)
WALBERLA
•
Massi ely pa allel mul i-
physics amewo k
• Domain decomposi ion
• Load balancing
• LB-Ke nel
• Bounda y Condi ions
• MPI Packing ke nels
MERIC
•
Ene gy consump ion mea-
su emen
• HW esou ces uning
• Applica ion acing
Figu e 1. O e iew o he so wa e s ack wi h lbmpy,pys encils,WALBERLA and so wa e uning. Wi h
he high-le el Py hon packages lbmpy and pys encils, he nume ical equa ions a e de i ed, disc e ized,
and, inally, lowe -le el C-Code is gene a ed om his symbolic ep esen a ion. The gene a ed code
can be combined wi h he C
++
amewo k WALBERLA and compiled. The execu able is op imized in
e ms o ene gy consump ion.
F om he symbolic desc ip ion, an Abs ac Syn ax T ee (AST) is cons uc ed wi hin
he pys encils In e mi ed Rep esen a ion (IR). This ee-based ep esen a ion inco po a es
a chi ec u e-speci ic AST nodes and poin e access in subsequen ke nels. Wi hin his ep e-
sen a ion, spa ial access pa icula s a e encapsula ed h ough pys encils ields. Addi ionally,
cons an exp essions o ixed special model pa ame e s can be di ec ly e alua ed o educe
he compu a ional o e head. Gi en ha he LBM compu e ke nel is symbolically de ined,
encompassing all ield da a accesses, he au oma ed de i a ion o compu e ke nels na u-
ally ex ends o encompass bounda y condi ions. This p ocess also in ol es he c ea ion o
ke nels o packing and unpacking. This sui e o ke nels plays a pi o al ole in popula ing
communica ion bu e s o MPI ope a ions.
Finally, he in e media e ep esen a ion o he compu e, bounda y and packing/un-
packing ke nels is p in ed by he C o he CUDA backend o pys encils o a clea ly de ined
in e ace. Each unc ion akes aw poin e s o a ay accesses oge he wi h hei ep esen a-
i e shape and s ide in o ma ion as well as all emaining ee pa ame e s. This simple and
Ene gies 2024,17, 502 4 o 15
consis en in e ace makes i possible o easily in eg a e he ke nels in exis ing C/C
++
so -
wa e s uc u es. Fu he mo e, wi h Py hon C-API, he low-le el ke nels can be mapped o
Py hon unc ions, which enables in e ac i e de elopmen by u ilizing lbmpy/pys encils as
s and-alone packages.
LBM is known o i s high memo y demand, and hus i was o en shown ha highly
op imized compu e ke nels a e only limi ed by he memo y bandwid h o a p ocesso o
accele a o [
6
,
7
]. Thus, na u ally, he ques ion a ises as o whe he i is possible o educe
he ene gy consump ion by educing he equency o he CPU compu e uni s (CPU co es)
while main aining he ull memo y subsys em pe o mance. Fu he mo e, he high le el o
op imiza ion employed o lbmpy leads o especially low FLOP numbe s in he ho spo o
he code [8].
4. Ene gy-Awa e Ha dwa e Tuning—Theo e ical Backg ound
Ene gy e iciency is commonly de ined as he pe o mance achie ed pe uni o powe
consump ion, ypically exp essed as loa ing poin ope a ions pe second pe Wa . How-
e e , when dealing wi h codes based on he la ice Bol zmann me hod (LBM), pe o mance
is be e cha ac e ized by he numbe o La ice Upda es execu ed Pe Second (LUPS). In
his s udy, we quan i y he ene gy e iciency o he WALBERLA applica ion as Millions o
La ice Upda es pe Second pe Wa (MLUPs/W).
To accu a ely measu e he ene gy consump ion o an applica ion, a high- equency
powe moni o ing sys em is impe a i e. This sys em should p o ide eal- ime powe o
ene gy consump ion eadings o he en i e compu a ional node o , a a minimum, i s key
compu a ional componen s.
The o al ene gy consumed (in Joules) can be calcula ed om powe samples (in Wa s)
ob ained a a speci ic sampling equency (in He z) as depic ed in he ollowing equa ion:
Ene gy( ) = Z
0Powe (x),dx≈∑n
i=0Powe Samplei
SamplingF equency.
The e a e wo undamen al app oaches o inc ease ene gy e iciency: (1) op imizing
applica ions o ully exploi compu a ional esou ces, ensu ing ha he wo kload aligns
wi h he uppe limi s de ined by he ha dwa e’s oo line model, o (2) judiciously limi ing
unused esou ces o p e en powe was age.
Mode n high-pe o mance CPUs and GPUs o e a leas one unable pa ame e
con ollable om he use space. Typically, i is he equency o compu a ion uni s (CPU
co es) which di ec ly impac s he peak pe o mance o he chip, and i is c ucial o compu e-
in ensi e compu ing asks. These pa ame e s can be adjus ed ei he s a ically o dynamically.
S a ic uning in ol es con igu ing speci ic ha dwa e se ings a he s a o an appli-
ca ion execu ion and main aining his con igu a ion un il i s comple ion. Howe e , such
s a ic se ups a e a ely op imal o complex applica ions, leading o subop imal ene gy
sa ings. S a ic uning lacks adap abili y o wo kload changes du ing applica ion execu ion,
hinde ing he achie emen o maximum a ailable e iciencies.
In con as , dynamic uning adjus s pa ame e s con inuously du ing applica ion
execu ion. This unc ionali y is achie ed by ene gy-awa e un ime sys ems ha can iden i y
op imal se ings o di e en phases o he applica ion and modi y ha dwa e con igu a ion.
One such sys em is COUNTDOWN [
16
], main ained by CINECA and he Uni e si y
o Bologna. COUNTDOWN dynamically scales CPU co e equency du ing he MPI com-
munica ion and synch oniza ion phases, while ensu ing ha he applica ion’s pe o mance
is p ese ed.
The Ba celona Supe compu ing Cen e de elops EAR [
17
], a lib a y ha i e a i ely
adjus s he CPU co e equency o powe cap based on bina y ins uc ions and pe o mance
coun e s alues.
LLNL Conduc o [
18
] employs a powe limi app oach, iden i ying c i ical communi-
ca ion pa hs and alloca ing mo e powe o slowe p ocesses o educe wai ing imes, hus
enhancing o e all pe o mance. Simila ly, LLNL Unco e Powe Sca enge [
19
] dynami-
Ene gies 2024,17, 502 5 o 15
cally unes he In el CPU con igu a ion by sampling RAPL DRAM powe consump ion
and ins uc ions pe cycle a ia ion. Op imal ene gy sa ings a e achie ed wi h a 200ms
sampling in e al.
Fu he mo e, he Run ime Exploi a ion o Applica ion Dynamism o Ene gy-e icien
eXascale compu ing (READEX) p ojec [
20
] in oduced a dynamic uning me hodology [
21
]
and i s implemen a ion. The ools de eloped in his p ojec p o ide HPC applica ion de el-
ope s wi h ways o exploi he dynamic beha io o hei applica ions. The me hodology is
based on he assump ion ha each egion o an applica ion may equi e speci ic ha dwa e
con igu a ions. The READEX app oach iden i ies hese equi emen s o each egion and
dynamically adjus s he ha dwa e con igu a ion when en e ing a egion. MERIC [
22
], an
implemen a ion o he READEX app oach de eloped a IT4Inno a ions, de ines a min-
imum un ime o he egion as 100ms o ensu e eliable ene gy measu emen s and o
accommoda e la ency when changing ha dwa e con igu a ions.
These ools a e able o b ing signi ican ene gy sa ings wi hou o wi h limi ed pe -
o mance penal y. Howe e , hese ools a e designed o wo k well on non-accele a ed
machines, and hey do no une GPU pa ame e s. Howe e , in mode n accele a ed HPC
clus e s, GPUs consume he majo i y o he compu e node ene gy. The ene gy e iciency o
da a cen e GPUs is signi ican ly highe when compa ed o ha o se e CPUs. This is
con i med by he ac ha all he op- anked supe compu e s in G een500 [
23
] ( he lis o
he mos ene gy-e icien HPC sys ems) a e based on N idia o AMD GPUs. Taking all his
in o accoun , i is s ill possible o imp o e he ene gy e iciency o hese GPUs by a ound
ens o pe cen as p esen ed in [24].
A su ey o GPU ene gy-e iciency analysis and op imiza ion echniques [
25
] e e s
o a ious app oaches o iden i y he op imal con igu a ion o he GPU equency o
ob ain ene gy sa ings, bu in each lis ed case, he pape p esen s a single con igu a ion
o he whole execu ion o he applica ion, which we e e o as s a ic uning. Simila ly,
K aljic e al. [
26
] iden i ied he execu ion phases o an analyzed applica ion by sampling
GPU ene gy consump ion. Ghazan a e al. [
27
] ained a neu al ne wo k model o iden i y
he op imal GPU SM equency. Howe e , in all cases abo e, he au ho s did no y o
change he con igu a ion dynamically du ing he execu ion o an applica ion.
5. Resul s o he waLBe la Ene gy Consump ion Op imiza ion
Fo s udying he ene gy consump ion o WALBERLA in an indus ially ele an es
case, he so-called LAGOON (LAnding-Gea nOise da abase o CAA alida iON), which
is simpli ied Ai bus plane landing gea , was chosen [
28
,
29
]. The es se up consis s o
he LAGOON geome y (see Figu e 2) in a i ual wind unnel wi h a uni o m esolu ion
o 512
×
512
×
512 cells. Fo he in low wall bounce back, bounda y condi ions wi h an
in low eloci y o 0.05 in la ice uni s we e used, while he ou low was modeled wi h
non- e lec ing ou low bounda ies. Fo he pu pose o he ene gy uning, he benchma k
case was simula ed o 100 ime s eps.
Benchma king and pe o mance and ene gy measu emen s we e pe o med a he ol-
lowing machines o he IT4Inno a ions supe compu ing cen e : (1) Ba bo a non-accele a ed
pa i ion, and (2) Ka olina GPU accele a ed pa i ion.
The Ba bo a sys em is equipped wi h wo In el Xeon Gold 6240 CPUs (codename
Cascade lake) pe node. Each CPU has 18 co es ( he hype - h eading is disabled) and is
designed o wo k a 150W TDP. The nominal equency o he CPU is 2.6GHz, bu i can
each up o (i) 3.9GHz o he u bo equency when only wo co es a e ac i e o (ii) 3.3GHz
o he u bo equency when all co es a e ac i e and execu e SSE ins uc ions. The CPU
co e equency can be educed all he way o 1.1GHz by a use o ope a ing sys em. Since
he Nehalem a chi ec u e, In el has been using ‘unco e’ o e e o he equency o he
subsys ems in he physical p ocesso package ha a e sha ed by mul iple p ocesso co es
e.g., las -le el cache, on-chip ing in e connec o in eg a ed memo y con olle s. Unco e
egions o e all occupy app oxima ely 30% o a chip a ea [
30
]. While he co e equency is
c i ical o compu e-bound egions, he memo y-bound egions a e much mo e sensi i e

Ene gies 2024,17, 502 6 o 15
o unco e equency [
31
]. The Ba bo a CPUs can scale he unco e equency be ween 1.2
and 2.4 GHz.
Figu e 2. Geome y o he LAGOON plane landing gea (le ). The geome y was s udied in a i ual
wind unnel ( igh ).
Ba bo a’s compu a ional nodes a e equipped by he on-boa d A os|Bull High De ini-
ion Ene gy E icien Moni o ing (HDEEM) sys em [
32
], which eads powe consump ion
om he mainboa d ha dwa e senso s and s o es he da a o dedica ed memo y. The senso
ha moni o s he consump ion o he whole node p o ides 1000 powe samples pe second,
and he es o he senso s ha moni o he compu e node sub-uni s p o ide 100 samples
pe second. Bo h agg ega ed alues and powe samples can be ead om he use space
using a dedica ed lib a y o command-line u ili y. Since Sandy B idge gene a ion, In el p o-
cesso s ha e in eg a ed a Running A e age Powe Limi (RAPL) ha dwa e powe con olle
ha p o ides a powe measu emen and mechanism o limi he powe consump ion o
se e al domains o he CPU [
33
]. In el RAPL con ols he CPU co e and unco e equencies
o keep he a e age powe consump ion o he CPU package below he TDP. The In el
RAPL in e ace allows a educ ion in his powe limi bu no an inc ease.
Ka olina clus e ( ank 71. in Top500 11/2021, 8. in G een500 11/2021 [
34
]) nodes a e
equipped wi h wo AMD EPYC 7763 CPUs, and eigh N idia A100-SXM4 GPUs. One MPI
p ocess pe GPU is used, ou pe CPU. To imp o e he ene gy e iciency o he applica ion,
we ha e speci ied he equency o he GPU s eaming mul ip ocesso s (SMs), which is
analogous o he CPU co e equency. Fo his pu pose, we use he N idia Managemen
Lib a y (NVML), which p o ides he unc ion n mlDe iceSe Applica ionsClocks() ha se s a
speci ic clock speed o a a ge GPU o bo h (i) memo y and (ii) s eaming mul ip ocesso s.
Howe e , he A100-SXM4 uses HBM2 memo y, whose equency canno be uned as i is
possible in he case o he GDDR memo y. The e o e, on da a-cen e -g ade GPUs, like A100,
only he equency o SMs can be con olled.
The ene gy consump ion o he GPU-accele a ed applica ion execu ions was measu ed
using pe o mance coun e s o he GPU (accessed using he N idia Managemen Lib a y)
and CPU (AMD RAPL wi h a simila powe moni o ing in e ace as In el RAPL wi hou
he suppo o powe capping).
To imp o e he ene gy e iciency o he GPU-accele a ed execu ions o he WALBERLA,
we pe o med he s a ic uning o he GPUs. The MERIC un ime sys em suppo s he
dynamic uning o GPU-accele a ed applica ions based on CPU egions only i he GPU
wo kload is synch onized. This limi a ion comes om he equi emen ha he GPU
equency canno be con olled om a ke nel. I is he CPU ha c ea es he eques o
change he equency h ough he GPU d i e .
By de aul , when unning a wo kload, he A100-SXM4 GPU (400W TDP) uses he
maximum u bo equency o 1.410GHz (i no o ced o educe he equency by he
powe consump ion exceeding he powe limi o by he mal h o ling) and swi ches o
he nominal equency o he GPU (1.095GHz) when copying he da a o/ om he GPU
memo y. The equency can be educed o 210 MHz in 81 s eps.
Du ing he execu ion o he CUDA ke nels on he GPUs, we also e alua ed he impac
o he CPU co e equency uning. The nominal equency o he AMD EPYC 7763 (280W
Ene gies 2024,17, 502 7 o 15
TDP) is 2.45 GHz, while he CPU can un up o 3.525 GHz boos equency. To educe
he numbe o es s, he 100 MHz s ep was used ins ead o he 25 MHz s ep, which is he
highes esolu ion suppo ed.
To con ol he ha dwa e pa ame e s men ioned abo e and o measu e he esou ce
consump ion o he execu ed applica ion, we used he MERIC un ime sys em. In he case
o he non-accele a ed e sion o WALBERLA, bo h s a ic (single ha dwa e con igu a ion o
he en i e applica ion execu ion) and dynamic uning (a speci ic ha dwa e con igu a ion o
each pa o he applica ion) we e used. In he case o he GPU-accele a ed e sion, s a ic
uning was used only since he MERIC does no ha e suppo o iden i y which CUDA
ke nel is unning on a GPU. A un ime sys em wi h suppo o dynamic GPU uning is
s ill a wo k in p og ess.
5.1. S a ic Tuning o he CPU Pa ame e s
The WALBERLA ene gy e iciency analysis s a ed wi h he s a ic uning o i s non-
accele a ed e sion. We pe o med an exhaus i e s a e–space sea ch, es ing all possible
CPU co e and unco e equency con igu a ions, using he 0.2GHz s ep o bo h he CPU
co e and he unco e equency. The lowes equencies we e omi ed since one can expec
high pe o mance penal y in hese con igu a ions.
Table 1(pe o mance penal y), Table 2(HDEEM ene gy sa ings) and Table 3(In el
RAPL ene gy sa ings) show he consump ion o esou ces o he WALBERLA sol e in
a ious con igu a ions, using colo coding o indica e which alues a e be e (g een) o
wo se ( ed).
F om all he e alua ed con igu a ions, he highes ene gy sa ings based on he HDEEM
measu emen s a e 22.8%. These we e eached o he ollowing con igu a ion: CF 1.9 GHz
and UCF 1.8 GHz. Fo he same con igu a ion, he sa ings calcula ed om he RAPL
measu emen s a e 29.6%. Howe e , in his con igu a ion, he pe o mance d ops by abou
13.9%.The majo di e ence in ene gy sa ings be ween HDEEM and RAPL comes om he
se o powe domains moni o ed by hem. RAPL only moni o s he powe consump ion
o he CPU, which is he only componen ha b ings ene gy sa ings due o uning. The
powe consump ion o he emaining node componen s emains unchanged, which has a
majo impac on ene gy sa ings i he un ime is ex ended. Since HDEEM moni o s he
powe consump ion o he en i e node, i s esul s a e mo e ep esen a i e.
S a ic uning usually p o ides a limi ed possibili y o ob ain majo ene gy sa ings wi h-
ou a pe o mance penal y. WALBERLA eached 10.2% ene gy sa ings based on HDEEM
measu emen s and 12.1% ene gy sa ings based on RAPL measu emen s a a cos o 1.6%
pe o mance deg ada ion when he co e equency was educed o 2.8 GHz and he unco e
equency emained wi hou any limi a ions.
The pe o mance impac o he e alua ed CPU equencies con igu a ions on he
WALBERLA sol e is shown in Table 1. The espec i e ene gy sa ings a e in Table 2 o he
HDEEM measu emen s and in Table 3 o he In el RAPL measu emen s. Based on he
measu emen s in hese ables, we iden i ied con igu a ions ha cause up o 2, 5, 10 and
unlimi ed un ime ex ension while b inging he maximum possible ene gy sa ings based
on HDEEM measu emen s. The summa y o he esul s is in Table 4, which p esen s he
bes con igu a ions o a ious pe o mance penal y limi s. I also shows one hand-picked
con igu a ion which p o ides 19.6% ene gy sa ings based on HDEEM measu emen s while
ex ending he un ime only abou 6.9%.
Table 1. Impac o he s a ic uning on he o e all un ime in [%] o WALBERLA when unning he
Lagoon use case.
co e[GHz]
unco e[GHz] 1.9 2.1 2.3 2.5 2.6 2.7 2.8 2.9 3.0 3.1 3.2 3.3
1.6 −17.46 −15.91 −14.6 −13.68 −14.2 −13.06 −13.93 −14.35 −11.71 −11.49 −11.26 −10.82
1.8 −13.9 −11.75 −10.35 −9.97 −9.24 −8.83 −8.45 −8.34 −7.64 −8.01 −7.32 −6.93
2−10.89 −10.03 −7.79 −6.86 −6.37 −5.78 −5.57 −5.36 −4.79 −4.27 −4.44 −4.01
2.2 −8.87 −6.88 −5.85 −4.75 −4−3.84 −3.38 −3−2.94 −2.22 −1.85 −1.89
2.4 −7.3 −5.17 −4.24 −3.1 −2.81 −3.69 −1.59 −1.33 −1.06 −0.96 −0.27 −0.09
Ene gies 2024,17, 502 8 o 15
Table 2. Ene gy sa ings o di e en con igu a ions using HDEEM measu emen s in [%] o WAL-
BERLA when unning he Lagoon use case.
co e[GHz]
unco e[GHz] 1.9 2.1 2.3 2.5 2.6 2.7 2.8 2.9 3.0 3.1 3.2 3.3
1.6 22.61 21.39 19.78 16.63 15 14.29 11.81 9.22 9.53 7.07 4.86 2.45
1.8 22.84 22.28 20.84 17.3 16.63 15.36 14.18 12.18 10.59 7.68 6.02 3.56
2 21.95 20.51 19.62 16.64 15.79 14.75 13.35 11.59 9.92 8.14 5.71 3.21
2.2 20.19 19.61 17.74 15.15 14.47 13.14 11.99 10.49 8.14 6.54 4.72 1.82
2.4 17.69 17.26 15.54 12.89 11.81 9.56 10.17 8.33 6.36 4.18 2.68 −0.27
Table 3. Ene gy sa ings o di e en con igu a ions using In el RAPL measu emen s in [%] o
WALBERLA when unning he Lagoon use case.
co e[GHz]
unco e[GHz] 1.9 2.1 2.3 2.5 2.6 2.7 2.8 2.9 3.0 3.1 3.2 3.3
1.6 31.25 28.8 26.42 22.46 20.46 19.34 16.38 13.61 13.62 10.44 7.81 4.63
1.8 29.63 29.2 26.93 22.23 21.68 19.95 18.42 16.05 13.77 10.5 8.08 5.16
2 28.04 26.1 24.99 21 19.96 18.57 16.67 14.28 12.38 10.06 7.37 4.34
2.2 25.4 24.6 21.84 18.77 17.82 16.08 14.58 12.92 9.88 7.81 5.84 2.36
2.4 22.35 21.28 19.48 15.61 14.42 11.66 12.06 9.88 7.69 4.97 3.33 −0.32
Table 4. Resul s o he s a ic uning o he WALBERLA sol e when unning he Lagoon use case.
The able p esen s ene gy sa ings om bo h he HDEEM and In el RAPL ene gy measu emen s o
a ious pe o mance deg ada ion ade-o s. In addi ion, one hand-picked con igu a ion ha b ings
meaning ul ene gy sa ings wi h modes pe o mance penal y is also p esen ed.
−2% −5% −10% No Selec ed
Limi Limi Limi Limi Con igu a ion
Run ime [%] −1.6 −4.8 −8.9 −13.9 −6.9
HDEEM [%] 10.2 15.2 20.2 22.8 19.6
RAPL [%] 12.1 18.8 25.4 29.6 24.6
CF; UCF [GHz] 2.8; 2.4 2.5; 2.2 1.9; 2.2 1.9; 1.8 2.1, 2.2
Finally, Figu e 3shows he powe consump ion imeline o he en i e compu e node
(blade) and selec ed node componen s. Samples we e collec ed by HDEEM du ing he
execu ion o he Lagoon es case. One can see dynamic changes in powe consump ion,
which indica es ha a s a ic ha dwa e con igu a ion is no op imal o he whole applica ion
un because he ha dwa e equi emen s change o e ime.
5.2. Dynamic Tuning o he CPU Pa ame e s
This sec ion p esen s a WALBERLA analysis using dynamic uning, which se s ha d-
wa e con igu a ions ha bes sui each ins umen ed sec ion o he code. MERIC suppo s
au oma ic bina y ins umen a ion, which gene a es a copy o he applica ion execu able
bina y ile ha includes he MERIC API calls a he beginning and he end o all selec ed
egions. Due o he excep ion o usage in he WALBERLA code, i was no possible o use
ully au oma ic bina y ins umen a ion because he execu ion hen esul ed in a un ime
e o o uncaugh excep ions. We manually ins umen ed he WALBERLA sou ce code wi h
he MERIC unc ion calls, which esul ed in less ine-g ain ins umen a ion han would
be possible wi h ull bina y ins umen a ions. Despi e he ac ha he ins umen a ion
consis s o only eigh egions, which is no op imal, we we e able o co e 99% o he
applica ion un ime.
In he case o he iden i ica ion o an op imal dynamic con igu a ion, we also execu ed
he applica ion in a ious ha dwa e con igu a ions, while o each ins umen ed egion,
we iden i ied i s op imal con igu a ion. The s a e–space sea ch was pe o med wice—one
o ob ain a con igu a ion ha does no cause any pe o mance deg ada ion, and ano he
one o b ing maximum HDEEM ene gy sa ings wi hou any un ime ex ension limi a ion.
Ene gies 2024,17, 502 9 o 15
Figu e 3. Powe imeline based on HDEEM measu emen s o WALBERLA when unning he Lagoon
use case. The da a a e collec ed o en i e compu e node (blade), and i s componen s ( wo CPUs and
wo g oups o memo y channels). The ime windows shows one synch oniza ion phase ollowed by
he six sol e i e a ions.
Table 5compa es ou di e en execu ions o WALBERLA, he de aul ha dwa e con-
igu a ion, he comp omise s a ic con igu a ion wi h 2.1 GHz co e and 1.8 GHz unco e
equency, and wo execu ions using dynamic uning wi h and wi hou pe o mance
penal y. The WALBERLA execu ion Lagoon con igu a ion has changed o show ha hese
pa ame e s ha e a majo impac on s a ic execu ion op imal con igu a ion and wha sa -
ings i b ings. In con as o he p e ious sec ion, he e, we p esen alues o he whole
applica ion un ime because he dynamic uning op imized he whole un ime. In his
case, he sol e akes 3/4 o he un ime. While in he p e ious Sec ion 5.1, we p esen a
un ime ex ension o abou 12.6% in he comp omise s a ic con igu a ion, now he same
con igu a ion ex ends he un ime by jus abou 5.8%. The p oblem ha each execu ion
con igu a ion may esul in a di e en op imal s a ic con igu a ion is sol ed by using
dynamic uning since each egion o he applica ion may ake a di e en ime; howe e ,
he egions’ ha dwa e equi emen s a e he same, and hus he op imal con igu a ion is
he same.
Table 5shows ha dynamic uning can achie e highe ene gy sa ings han s a ic
uning. The dynamically uned execu ion o WALBERLA consumed 7.9% less ene gy
wi hou ex ending he un ime. The highes ene gy sa ings achie ed wi h dynamic uning
is 19.1% a a cos o ex ending he un ime by 16.2%.
Please no e ha in his case, he sol e consis s o a single egion only. In Figu e 3, i
is isible ha he sol e should be spli in o a leas wo di e en egions because hese
egions ha e di e en ha dwa e equi emen s. Howe e , hei un ime is e y sho (up
o 5ms), while MERIC equi es egions o a leas 100ms. In es iga ing possible ene gy
sa ings gained om he dynamic uning o he sol e is ou goal when a new elease o he
MERIC ha b ings be e suppo o ine-g ain uning becomes a ailable.