scieee Open visual document viewer

Dynamic load balance of chemical source term evaluation in high-fidelity combustion simulations

Ramirez Miranda, Guillem,Mira Martínez, Daniel,Pérez Sánchez, Eduardo Javier,Surapaneni, Anurag,Borrell Pol, Ricard,Houzeaux, Guillaume,Garcia Gasulla, Marta

Abstract

This paper presents a load balancing strategy for reaction rate evaluation and chemistry integration in reacting flow simulations. The large disparity in scales during combustion introduces stiffness in the numerical integration of the PDEs and generates load imbalance during the parallel execution. The strategy is based on the use of the DLB library to redistribute the computing resources at node level, lending additional CPU-cores to higher loaded MPI processes. This approach does not require explicit data transfer and is activated automatically at runtime. Two chemistry descriptions, detailed and reduced, are evaluated on two different configurations: laminar counterflow flame and a turbulent swirl-stabilized flame. For single-node calculations, speedups of 2.3x and 7x are obtained for the detailed and reduced chemistry, respectively. Results on multi-node runs also show that DLB improves the performance of the pure-MPI code similar to single node runs. It is shown DLB can get performance improvements in both detailed and reduced chemistry calculations.

Full text

Dynamic load balance o chemical sou ce e m e alua ion in high- ideli y combus ion simula ions Guillem Rami ez-Mi anda, Daniel Mi a∗, Edua do J. P´ e ez-S´ anchez, Anu ag Su apaneni, Rica d Bo ell, Guillaume Houzeaux, Ma a Ga cia-Gasulla Ba celona Supe compu ing Cen e (BSC), Plaza Eusebi G¨uell 1-3, 08034, Ba celona Spain Abs ac This pape p esen s a load balancing s a egy o eac ion a e e alua ion and chemis y in eg a ion in eac ing low simula ions. The la ge dispa i y in scales du ing combus ion in oduces s i ness in he nume ical in eg a ion o he PDEs and gene a es load imbal- ance du ing he pa allel execu ion. The s a egy is based on he use o he DLB lib a y o edis ibu e he compu ing esou ces a node le el, lending addi ional CPU-co es o highe loaded MPI p ocesses. This app oach does no equi e explici da a ans e and is ac i a ed au oma ically a un ime. Two chemis y desc ip ions, de ailed and educed, a e e alua ed on wo di e en con igu a ions: lamina coun e low lame and a u bulen swi l-s abilized lame. Fo single-node calcula ions, speedups o 2.3x and 7x a e ob ained o he de ailed and educed chemis y, espec i ely. Resul s on mul i-node uns also show ha DLB imp o es he pe o mance o he pu e-MPI code simila o single node uns. I is shown DLB can ge pe o mance imp o emen s in bo h de ailed and educed chemis y calcula ions. Keywo ds: dynamic load balancing, combus ion, High-Pe o mance Compu ing, compu a ional luid dynamics, DLB lib a y 1. In oduc ion Regula ions applied o he powe and anspo a ion sec o s ha e p opelled esea ch on op imiza ion o he mal engines and o he combus ion sys ems, o simul aneously e- duce uel consump ion and pollu an emissions. The ex ensi e use o Compu a ional Fluid Dynamics (CFD) o his pu pose, as a con en ional and indispensable ool, has made ∗Co esponding au ho Email add ess: [email p o ec ed] (Daniel Mi a) P ep in submi ed o Compu e &Fluids Oc obe 17, 2022 a Xi :2210.07364 1 [physics. lu-dyn] 13 Oc 2022 This is he AAM: The inal e sion o his pape can be ound a : h ps://doi.o g/10.1016/j.comp luid.2022.105723. (h ps://www.sciencedi ec .com/science/a icle/pii/S0045793022003164 Copy igh © 2024 Else ie B.V. CC BY-NC-ND 4.0 impe a i e he sea ch o e icien algo i hms ha can sol e he equa ions o chemically eac ing lows a educed compu a ional cos s. The chemical sou ce e ms a e inse ed in he anspo equa ions, oge he wi h he e alua ion o he con ec ion and di usion e ms, ha cha ac e ize he eac ing low. Mo eo e , he choice o he eac ion mechanism is a c i ical aspec when pe o ming high- ideli y combus ion simula ions. The cos o he e alua ion o he chemical sou ce e m depends on he size and s i ness o he eac ion mechanism, and can esul in he mos demanding pa o he calcula ion. While he wo kload o e alua e he anspo e ms can be pa allelized using s anda d domain decomposi ion s a egies [24], i is ha de o de ine a compu a ionally balanced dis ibu ion o he chemis y wo kload and his speci ic ask can end up aking 90% o he compu a ional ime [29, 35]. Such load imbalance can be explained by he na u e o he combus ion p ocess, whe e he chemical eac ions occu in speci ic egions o he domain and usually along hin laye s, esul ing, in consequence, in high dispa i y o compu a ional load be ween p ocesso s, as no all he subdomains equi e he e alua ion o he chemical eac ion a es. Mo eo e , he s i ness o he ODEs (o dina y di e en ial equa ions) sys em, caused by he non-linea i y o chemis y and he wide ange o ime scales o adicals (a ound 10−8s) and majo species (in he o de o en o mic oseconds), inc eases he compu a ional cos o he chemis y in eg a ion and accen ua es e en mo e he load imbalance. Howe e , di e en om he anspo e ms o he low a iables, which depend on he s a e a he icini y a each poin , chemical sou ce e ms only depend on he punc- ual he mochemical s a e o he mix u e and, hence, i is suscep ible o a high deg ee o pa alleliza ion. Based on hese conside a ions, di e en s a egies ha e been used o achie e load bal- ance in combus ion calcula ions. In he ea ly wo ks by The enin e al. [31], he e was a ans e o poin s be ween neighbo ing p ocesso s, so ha he nodes equi ing mo e com- pu ing ime han he a e age send a ew g id poin s o hei neighbo s. A simila s a egy was p oposed by An onelly and D’Amb a (2011) [2], whe e a cell dis ibu ion based on a dynamic load balancing ha p ese es con igui y o he compu a ional g id cells was used. Ano he impo an aspec is he s i ness o sys em, which can be no iceably di e - en be ween cells. Chemical mechanisms may ha e be ween 50 and 1000 species, which include species ea u ing a wide ange o ime scales depending on he local condi ions. Fo ins ance, he low empe a u e egion de e mines he au oigni ion wi h high deg ee o s i ness, while he high empe a u e egion has low s i ness as species wi h la ge mass ac ions can each equilib ium ela i ely as [2]. These obse a ions can be used o e- duce he compu a ional cos . Muela e al. [25] p oposed a measu e o he s i ness o sepa a e be ween explici and implici algo i hms. Koda asal e al. (2016) [24] p oposed a “s i ness-based” algo i hm o load balancing chemical kine ics using in o ma ion om 2 p e ious ime-s eps. A simila s a egy has been used by Teckg¨ ul e al. (2021) [30], which also includes a zonal e e ence cell mapping me hod o a oid he e alua ion o he ki- ne ic a es in ambien egions wi h low eac i i y. In he con ex o High Pe o mance Compu ing (HPC) wi h massi e use o CPUs, his s a egy can be la gely bene i ed by he hyb id use o CPUs and GPUs p o ided ha explici algo i hms, which a e easily pa al- lelized and consume less memo y, a e applied o GPUs [29]. Zi wes e al. (2018) [34] p oposed a con e sion o eac ion mechanisms in o sou ce codes o es uc u e he da a o he kine ic mechanisms o e icien compu a ion enabling compile op imiza ions. De- spi e hese s a egies ha e esul ed in no iceable compu a ional cos educ ions, ensu ing a load balance s a egy o gene al applica ions in p emixed and non-p emixed combus- ion will be o high alue. All he p e ious s a egies a e based on a e-dis ibu ion o he load a sys em le el, whe e he eac ion a e e alua ion and chemis y in eg a ion is e- dis ibu ed acco ding o he cu en load and an es ima ion o he s i ness using Message Passing In e ace (MPI) s anda d. Howe e , one o he undamen al issues associa ed o hose s a egies is he e alua ion o he chemical s i ness, which is di icul o p edic and can esul in load imbalance. Mo eo e , when applying hese me hods i is necessa y o synch onize he di e en p ocesses and exchange da a h ough MPI communica ion. This pape is de o ed o implemen and analyze new load balancing mechanisms o chemis y in eg a ion. In pa icula , we p opose u ilizing he Dynamic Load Balancing (DLB) lib a y 1[15, 16], which allows eusing CPU-co es associa ed wi h idle MPI p o- cesses by o he p ocesses unning on he same node. I is a load balancing mechanism based on ans e ing idle esou ces a he node le el a he han ans e ing wo kload subse s h ough message passing. DLB ac s as an au oma ic un ime mechanism anspa - en o he use and equi es minimum changes in he sou ce code ( wo lines in he p esen s udy). In ac , DLB can be combined wi h he wo kload ans e ing s a egies men ioned abo e. DLB has al eady been success ully applied o inc ease he load balance o he assembly o he igh -hand side e ms in he Na ie -S okes equa ions [17], he pa icle anspo [22], coupled codes [9] and has been also used in di e en a chi ec u es [18]. He e, i is ex ended o op imize he chemis y pa in eac ing low simula ions. The p o- posed solu ion wi h DLB, does no need o add ex a da a mo emen , because e e y hing is done h ough he sha ed memo y o he node. As i is a dynamic mechanism ha eac s o he load imbalance, i does no need o p edic he s i ness no he compu a ion load associa ed. And las bu no leas , i does no equi e a hea y implemen a ion e o in he applica ion. The emaining o he pape is o ganized as ollows. Sec ion 2 desc ibes he compu- a ional amewo k ha is used o conduc ing he nume ical simula ions including he 1h ps://pm.bsc.es/dlb 3 modelling and nume ical desc ip ions used om he code Alya [32]. Sec ion 3 desc ibes he compu a ional en i onmen in which hese simula ions a e conduc ed and Sec ion 4 de ails he assessmen o his dynamic load balance s a egy on ep esen a i e p oblems in combus ion science. Finally, he conclusions and di ec ions o u u e wo k a e gi en in Sec ion 5. 2. Modelling amewo k 2.1. Go e ning anspo equa ions The simula ion o eac ing lows includes go e ning equa ions o chemical species along wi h ene gy, momen um and con inui y. A low Mach numbe app oxima ion o he Na ie -S okes is conside ed in his s udy o which he conse a ion o con inui y and momen um ead: ∂ρ ∂ +∇·(ρu)=0,(1) ∂(ρu) ∂ +∇·(ρuu)=−∇p+∇·(µ∇u),(2) whe e s anda d no a ion is used o all he quan i ies and ρ,u,pand µ ep esen he densi y, eloci y ec o , p essu e and dynamic iscosi y. Rega ding he e olu ion o he mul icomponen gas, i can be exp essed in e ms o he anspo equa ions o he indi idual species Ykgi en by: ∂(ρYk) ∂ +∇·(ρuYk)=∇·(ρD∇Yk)+˙ωYkk=1,...,Ns.(3) In his equa ion, Dis he di usion mass coe icien , o which a uni y Lewis assump ion has been adop ed, while ˙ωYkdeno es he chemical sou ce e m o species Yk.Nsis he numbe o species conside ed in he chemical mechanism. Finally, he o al en halpy h equa ion, in which hea ing due o iscous o ces is neglec ed, eads as: ∂(ρh) ∂ +∇·(ρuh)=∇·(ρD∇h).(4) 2.2. Chemical in eg a ion The chemical in eg a ion is one o he mos compu a ionally demanding pa s in he in eg a ion o he go e ning equa ions due o he high non-linea i y o he A henius- ype eac ion kine ics. I is, he e o e, clea ha he in eg a ion me hod o chemis y may play 4 an impo an ole in he o al ime o he simula ion, especially when de ailed chemis y models a e conside ed. In his wo k, o educe he s i ness o he in eg a ion o he species go e ning equa- ions 3, a spli ing algo i hm is used o sepa a e he anspo om he chemis y [27]. The solu ion o he chemis y p oblem is achie ed by he in eg a ion o he open sou ce Can- e a [19] so wa e as an ex e nal lib a y in he mul iphysics code Alya [32]. A Fo an o C++ w appe was c ea ed o in eg a e Can e a in o Fo an o Alya, so in e nal unc ions om Can e a could be used in un ime. The eac ion a es a e ob ained om in-buil in- e nal unc ions om Can e a, and he chemical in eg a ion is ob ained using he CVODE algo i hm [11]. A lis ing o he in eg a ion loop o he code is gi en in Lis ing 1. CVODE is a package w i en in C o sol e IVPs (Ini ial Value P oblems) de ined by s i and non-s i ODEs in he o m: ˙y= ( ,y),(5) wi h he ini ial condi ions gi en by y( = 0)=y0. In pa icula , he equa ions o chemis y in eg a ion a e simila o hose o sys em 5 bu wi hou he dependence on ime: dY d = (T,Y),(6) whe e Tis empe a u e, Y=(Y1,...,YN)Tis he ec o o mass ac ion o species. The ini ial condi ions co espond o T( 0)=T0and Y( 0)=Y0. CVODE is in u n based on he ODE sol e packages VODE and VODPK [7] and sol es p e ious IVP by he applica ion o an implici empo al scheme based on ei he Adams-Moul on o mula o backwa d-di e encing o mula (BDF) me hods, wi h he sub- sequen esolu ion o he non-linea equa ion by New on’s me hod. Depending on he o m o he Jacobian ma ix, dedica ed unc ions o dense o banded ma ices can be used allowing, mo eo e , he p econdi ioning o he linea sys em. The code used in Alya o chemical in eg a ion loop wi h no pa alleliza ion is ga he ed in Lis ing 1. 1do ipoin=1,npoin 2i ( eac ion) call c ode_in eg a ion 3end do 4... 5call MPI_All educe (...) Lis ing 1: Chemical in eg a ion loop in Alya code. 5 2.3. Compu a ional pla o m: Alya All he compu a ional s a egies p esen ed in his pape ha e been implemen ed in Alya, he high pe o mance compu a ional mul i-physics code de eloped a he Ba celona Supe compu ing Cen e . Alya is de eloped using a modula a chi ec u e ha includes a module coupling o ackle complex mul i-physics p oblems such as combus ion simula- ions. Alya is w i en in Fo an and is designed o massi ely pa allel supe compu e s; pa icula ly i is one o he wel e simula ion codes o he Uni ied Eu opean Applica ions Benchma k Sui e (UEABS) [8], being egula ly es ed on he Eu opean Tie -0 supe com- pu e s. The pa alleliza ion implemen ed in Alya comp ises h ee le els: dis ibu ed memo y o in e - and in a- node pa allelism, sha ed memo y o in a-node pa allelism, and SIMD (Single Ins uc ion Mul iple Da a) and SIMT (Single Ins uc ion Mul iple Th ead) o CPU ec o iza ion and GPU compu ing, espec i ely. The implemen a ion o such s a - egy combines a ious p og amming models: MPI o in e -p ocess message passing and synch oniza ion, di ec i es-based app oaches (OpenMP, OmpSs, and OpenACC) o loops and ask-based pa allelism and GPU o loading, as well as CUDA o low-le el op imized implemen a ion o speci ic ke nels. The p ima y op ion o mesh pa i ioning is an in-house SFC-based pa i ione [4]. On- line edis ibu ion is also pe o med o adjus he pa i ion o un- ime measu emen s. This las op ion has been exploi ed o co-execu ion on he e ogeneous sys ems, whe e he pa i- ion is adjus ed o a balanced execu ion using CPU and GPU de ices simul aneously [5]. Fo ask-based sha ed-memo y pa allelisms, a second-le el decomposi ion is pe o med in which he size o he esul ing subse s is o he o de o 102elemen s. Finally, a da a e- s uc u ing is pe o med o ec o iza ion: subse s o 8 o 32 elemen s a e packed oge he o be execu ed in a SIMD model. The same s a egy is used o GPU compu ing being, in his case, each pack, o he o de o 105elemen s, launched o he GPU whe e he SIMT pa allel model is exploi ed. 3. Desc ip ion o he es cases Two ep esen a i e combus ion p oblems a e used he e o e alua e he pe o mance o DLB o ensu e a load balancing s a egy o he chemical in eg a ion. The i s case co esponds o a lamina coun e low di usion lame a a mosphe ic p essu e and 298K ai and uel empe a u e. The lame in his con igu a ion is cha ac e - ized by a wide egion whe e uel and oxidize a e mixed and a eac ion laye o med a he icini y o he s oichiome ic mix u e ac ion. Mo eo e , he lame is de e mined by he le el o s ain, which is gi en by he eloci ies o he wo s eams and he dis ance be ween he nozzles. A ep esen a ion o he lame is shown in Fig. 1. A wo-dimensional 6 domain wi h wo di e en mesh sizes was used o e alua e he load balancing s a egy using a single compu ing node (coa se mesh) and a mul i-node calcula ion ( ine mesh). In addi ion, o analyse he e ec o load imbalance due o chemical in eg a ion, wo eac ion mechanisms wi h e y di e en sizes we e chosen o assess he pe o mance o DLB when using de ailed and educed chemis y. The i s eac ion mechanism is a de ailed chemical scheme o ke osene comp ising 189 species and 1327 eac ions [1], while he second is a 2-s ep educed mechanism con aining 6 species [14]. Figu e 1: Coun e low ke osene/ai lame using he de ailed eac ion mechanism: empe a u e ( op) and hea elease a e (bo om). The second es case is a mo e ealis ic p oblem and ea u es a u bulen p emixed lame in a swi l-s abilized bu ne also known in he li e a u e as he PRECCINSTA bu ne [20, 6]. The ope a ing poin a equi alence a io φ=0.67 [6] is aken as e e ence bu changing he uel o ke osene o allow o his compa ison. Fo his case, he go e ning equa ions a e il e ed in space and he compu a ional amewo k is adap ed o un la ge-eddy sim- ula ions (LES). De ails o he modelling and nume ical app oach a e gi en in p e ious wo k [6] and a e omi ed he e o b e i y. Fo his es case, a ini e a e model conside ing he il e ed equa ions o LES a e employed wi hou accoun ing o Tu bulence-Chemis y In e ac ions (TCI). This app oach has been chosen due o i s simplici y and wi h he aim o e alua e he load imbalance in LES and DNS applica ions. No wi hs anding, he con- clusions d awn in he ollowing ela ed o he imp o emen s o hyb idiza ion and DLB a e expec ed o be alid o o he u bulen combus ion models ha do chemical in eg a ion in si u such as he Condi ional Momen Closu e (CMC) [23], Eddy Dissipa ion Concep [13] o he T anspo ed P obabili y Densi y Func ions (TPDFs) models [21], g ea ly expanding he po en ial o DLB o ad anced combus ion simula ions. The compu a ional domain includes he plenum, swi le and combus ion chambe , and is composed by a hyb id mesh 7 including p isms, e ahed ons and py amids. This is shown in Fig. 2 along wi h a sample snapsho o he empe a u e ield. The same p e ious eac ion mechanisms a e also es ed in his con igu a ion. A summa y o he compu a ional cases and de ails o he mesh size and esolu ion is gi en in Table 1. Figu e 2: Swi l-s abilized u bulen p emixed lame using he de ailed eac ion mechanism: mesh esolu ion ( op) and empe a u e (bo om). Iden i ie Con igu a ion Mechanism Species Reac ions Mesh (cells) CF1 Coun e low De ailed 189 1327 6k CF2 Coun e low Reduced 6 2 6k CF3 Coun e low De ailed 189 1327 52k CF4 Coun e low Reduced 6 2 52k SB1 Swi l bu ne De ailed 189 1327 24M SB2 Swi l bu ne Reduced 6 2 24M Table 1: Desc ip ion o he compu a ional cases. Fo bo h o he cases, wo me ics we e ob ained, namely, he ime elapsed o he chemical in eg a ion in a Runge-Ku a sub-s ep and he o al ime equi ed o he in eg a- 8 ion o he whole ime s ep. The analysis includes single-node and mul i-node es s on he coun e low con igu a ion aiming o e alua e he e ec o he hyb idiza ion, g ain size and he op imiza ion wi h DLB. A e he co ec iden i ica ion o he op imal DLB pa ame e s o he coun e low lame, he same se ings a e applied o he u bulen swi l bu ne case in o de o e alua e his me hodology on p oduc ion uns. De ails o he analysis a e gi en in he nex subsec ions. 4. Compu a ional backg ound 4.1. Pe o mance analysis We s a he s udy wi h a pe o mance analysis o he execu ion which is ca ied ou ollowing he POP2me hodology [33, 3]. We use he BSC pe o mance analysis ools3: EXTRAE [28] o gene a e execu ion aces and Pa a e [26] o isualize hem. In Figu e 3, we can see wo imelines om Pa a e . Pa a e imelines show in he x axis he ime and each ow co esponds o one MPI p ocess. The MPI p ocesses a e o de ed going om MPI ank 0 in he op o he iew o he highes ank in he bo om o he iew. The colo can iden i y di e en me ics depending on he iew ha is selec ed. On he op ace, we can see in a colo ba he du a ion o he use ul compu a ion. We conside use ul compu a ion when he p ocess o h ead is doing compu a ion and no wai ing in an MPI call. On he bo om ace, we show he MPI call, in his case black colo co esponds o use ul compu a ion, because i is ou side an MPI call. Bo h aces a e depic ed using he same imescale. Fo illus a i e pu poses, in his iew we show wo ime s eps o one o he cases simula ed in his wo k, namely, he de ailed chemis y simula ion o a coun e low lame. In his case he execu ion is done using 192 MPI anks. We ha e ma ked wi h a yellow squa e one o he chemical in eg a ion loops. We can obse e he impo an load imbalance p esen in his sec ion o he execu ion, whe e some MPI anks ha e a highe load han o he ones. This p oduces an impo an ime spen wai ing in he MPI Ba ie call by some MPI p ocesses. In Figu e 4 we see he same analysis o he simula ion o a coun e low lame wi h educed chemis y. We can see he 192 MPI anks in he di e en ows and wo ime s eps in he ime scale (x axis). The op ace shows he du a ion o he use ul compu a ion, while he bo om one shows he MPI calls being execu ed. This use case also p esen s a high load imbalance in he chemical in eg a ion phases being one o he in eg a ion loops again ma ked wi h a yellow squa e. We can obse e 2h ps://pop-coe.eu/ 3h ps:// ools.bsc.es 9 Figu e 7: Zoom in a Pa a e ace o wo s eps o he de ailed chemis y showing he e ec o DLB. 16 pe o mance ne wo k in e connec and unning SuSE Linux En e p ise Se e as ope a - ing sys em. Compu e nodes a e equipped wi h 2 socke s In el Xeon Pla inum 8160 CPU wi h 24 co es each unning a 2.10GHz o a o al o 48 co es pe node and 96 GB o main memo y (1.88 GB/co e). We ha e used he In el compile 2017.4 and IMPI 2017.4 as MPI lib a y. Fo all he expe imen s we use he mas e b anch o Alya in eg a ed wi h he Can e a lib a y e sion 2.1. Fo he op imiza ion using DLB we ha e used DLB almos 3.0 and OmpSs 19.06. All he esul s ga he ed in his sec ion a e a e ages o 5 uns. As in all he cases, he s anda d de ia ion be ween he di e en uns is below 5%, he e o ba s a e no shown in he plo s. 5.2. Single-node es ing The main aim o his sec ion is o de e mine he impac o he code hyb idiza ion wi h OmpSs on he pe o mance as well as o be able o dis inguish in a second s age he op imiza ions p o ided by DLB. Also, we s udy in de ail he impac o he g ain size when pa allelizing wi h OmpSs o unde s and he impo ance o his ac o on he compu a ional cos and o y o ind i s op imal alue o ange o alues. The i s se o es s comp ises he analysis o he ime educ ions by he use o hy- b idiza ion and DLB on he coun e low lame con igu a ions CF1 and CF2. 5.2.1. Hyb idiza ion and g ain size s udy In his sec ion, we analyze he pe o mance o he hyb id code, ou lined in Lis ing 2, e sus he o iginal MPI-only implemen a ion. Figu e 8 p esen s a nume ical s udy o he coun e low lame conside ing he de ailed chemis y CF1 case using a single node o he Ma eNos um IV supe compu e . In he X axis we can see di e en con igu a ions o MPI p ocesses and OmpSs h eads, e.g. 12 ×4 co esponds o 12 MPI p ocesses and 4 OmpSs h eads each. No e ha he p oduc o hese combina ions is always 48 as all he esul s in he same plo a e using he same numbe o compu a ional esou ces (48 co es). In he Y axis we can see he speed up o he in eg a ion loop wi h espec o he MPI only e sion wi h 48 MPI p ocesses ( he pu e MPI only is depic ed as a ba in he plo as a e e ence). The di e en se ies ep esen ed wi h lines co espond o he hyb id e sion wi h di e en alues o he g ain size. We obse e ha he MPI-only execu ion akes he same ime as he 48 ×1 con ig- u a ion; he e o e, he ac i a ion o OmpSs does no gene a e an app eciable o e head. Mo eo e , he end is ha by inc easing he numbe o OmpSs h eads pe node, he chemis y in eg a ion cos is educed. As he e a e no in e -p ocess communica ions, he main eason o his accele a ion is he implici load-balancing ob ained om he sha ed memo y aski ica ion: s i and non-s i asks a e assigned o OmpSs h eads as hey a e 17 0 0,2 0,4 0,6 0,8 1 1,2 1,4 1,6 1,8 Pu e MPI 48x1 24x2 12x4 8x6 6x8 4x12 Speed up In eg a ion Con ig (MPIs x OmpSs h eads) 14816 32 64 128 Figu e 8: Compa ison o ime o de ailed chemis y in eg a ion (case CF1) be ween only MPI and hy- b idiza ion wi h di e en numbe o h eads o di e en g ain size alues. comple ed. Finally, we obse e ha esul s a e almos independen o he g anula i y as he o e lapping o he di e en lines shows. This means ha he e is no ele an o e head added by he hyb idiza ion o OmpSs and also ha he g anula i y o asks is small enough o p o ide pa allelism and wo k o all he h eads. Only he g anula i y o one elemen pe ask shows a e y sligh educ ion o he speedup o he con igu a ion 12 ×4, his is because in his case 12 OmpSs h eads a e asking o wo k o he un ime sys em, as he asks a e e y small we can s a o see some conges ion in he un ime when accessing he queue o wo k. All in all, 6 ×8 is he bes con igu a ion, and he speedup ob ained e sus he pu e MPI code is 1.6×. In Figu e 9 we see he same esul s o he educed chemis y CF2 case o which quali a i e di e ences a ise wi h ega ds o he de ailed chemis y case CF1. In his case, a di ec in eg a ion o he chemis y can be achie ed wi h much ewe i e a ions in he CVODE sol e , ob aining much lowe o e all compu ing imes compa ed o he de ailed chemis y case CF1. On he one hand, he g ain size signi ican ly a ec s he pe o mance. In his case, unlike o he de ailed chemis y, he cos o he in eg a ion o a single elemen is e y low, so a couple o hem need o be ga he ed pe ask o coun e balance he OmpSs o e head. On he o he hand, we obse e ha he speedup ises up o 3.4×and he bes esul is ob ained wi h he con igu a ion 12×4. This highe speedup comes om an ini ially highe imbalance as we ha e seen in he pe o mance analysis sec ion (Sec ion 4.1), whose co ec ion p oduces mo e no iceable e ec s. As he p oblem is en i ely go e ned by chemis y, he e alua ion o he global a es is local and can s ongly bene i om pa alleliza ion. This s udy p oposes he use o hyb id 18 0 0,5 1 1,5 2 2,5 3 3,5 4 Pu e MPI 48x1 24x2 12x4 8x6 6x8 4x12 Speed up In eg a ion Con ig (MPIs x OmpSs h eads) 1 4 8 16 32 64 128 Figu e 9: Compa ison o ime o educed chemis y in eg a ion (case CF2) be ween only MPI and hy- b idiza ion wi h di e en numbe o h eads o di e en g ain size alues. code ha uses he pa allel p og amming model OmpSs o imp o e he load balance and educe he compu a ional cos . Mo eo e , he OmpSs pa alleliza ion does no add a sig- ni ican o e head al hough, in cases wi h a e y low compu a ional load o he chemical in eg a ion, a small g ain size can add o e head when he numbe o h eads is inc eased. In hese cases, he g ain size migh ha e an impo an impac on he pe o mance and a mid o high alue o he g ain size can app eciably mi iga e such o e head. 5.2.2. DLB e alua ion In his sec ion, we s udy he bene i o using DLB as an addi ional load balancing mechanism in a-node and how he g ain size a ec s in hese simula ions. In Figu e 10, we show he speedup ob ained by he hyb id e sion wi h and wi hou DLB wi h espec o he MPI-only e sion (shown as a blue ba in he plo ) when unning in one node o Ma enos um4. In he X axis we can see he di e en con igu a ions o MPI h eads and OmpSs h eads o ill one node o 48 co es. Fo all he hyb id execu ions we use a g ain size o 32, because ha is he minimal size ha does no show any signi ican o e head in his p oblem. Adding DLB o he de ailed chemis y in eg a ion as shown o CF1, Figu e 10, gene a es a speedup o up o 2.3× e sus he MPI-only implemen a ion, and up o 1.5× e sus he bes hyb id con igu a ion. This indica es ha he use o DLB can imp o e he pe o mance and add ess he imbalance u he han only he hyb idiza ion o he code. I is ele an ha wi h he con igu a ion 48 ×1, we ob ain esul s ha a e no a om op imal, a speed up o 2×. I means ha we can ge mos o he pe o mance wi hou 19 0 0,5 1 1,5 2 2,5 Pu e MPI 48x1 24x2 12x4 8x6 6x8 4x12 Speed up In eg a ion Con ig (MPIs x OmpSs h eads) Hyb id DLB Hyb id No LB Figu e 10: Compa ison o ime o de ailed chemis y in eg a ion (case CF1) be ween only MPI and hy- b idiza ion wi h and wi hou DLB o di e en numbe o h eads wi h g ain size 32. equi ing an o e all hyb idiza ion o he code; only one h ead pe MPI p ocess needs o be ac i a ed in he zones whe e DLB is used. I is impo an o no ice ha he line co esponding o he execu ions wi h DLB (o ange) is la e han he one co esponding o he hyb id code. This means ha he use o DLB makes he pe o mance less dependen on he hyb id con igu a ion, hus, i elie es he p essu e om he use o decide he op imal con igu a ion o MPI p ocesses and h eads. Fo he educed chemis y CF2, we can see he esul s ob ained wi h DLB in Figu e 11. As in he p e ious plo , we show he speedup o e he MPI-only e sion (Y axis), when using di e en con igu a ions o MPI p ocesses and OmpSs h eads (X axis). I is obse ed ha he speedup e sus he MPI-only e sion ises up o 7×, which ep esen s an addi ional 2×accele a ion e sus he bes hyb id op ion. The speedup ob ained by DLB in he educed chemis y in eg a ion is highe han in he de ailed one, as he load imbalance is also highe in he educed chemis y case. No e ha he speedup ha DLB can ob ain is ela ed o he exis ing load imbalance o he applica ion. As DLB e-dis ibu es he asks by he idle ime om he p ocesso s, he e is no need o p edic he s i ness om he chemical p oblem as his is handled by DLB. A second impo an aspec o conside is he impac o he g ain size on he DLB pe o mance in hese applica ions. Figu e 12 shows he speedup e sus he MPI-only e sion o he code (y axis) o he de ailed chemis y case. We can see he di e en con igu a ions o MPI p ocesses and OmpSs h eads o ill a node o 48 co es. The di e en se ies ep esen he di e en g ain sizes used by he hyb id code and DLB. To unde s and he esul s, a ade-o be ween wo aspec s needs o be conside ed: i) he imbalance can be 20 0 1 2 3 4 5 6 7 8 Pu e MPI 48x1 24x2 12x4 8x6 6x8 4x12 Speed up In eg a ion Con ig (MPIs x OmpSs h eads) Hyb id DLB Hyb id No LB Figu e 11: Compa ison o ime o educed chemis y in eg a ion (case CF2) be ween only MPI and hy- b idiza ion wi h and wi hou DLB o di e en numbe o h eads wi h g ain size 32. 1 1,2 1,4 1,6 1,8 2 2,2 2,4 2,6 48x1 24x2 12x4 8x6 6x8 4x12 Speed up In eg a ion Con ig (MPIs x OmpSs h eads) 4 8 16 32 64 128 Figu e 12: Compa ison o g ain size impac when using DLB o de ailed chemis y in eg a ion (case CF1) o di e en numbe o h eads. 21 be e educed wi h hinne g anula i y; ii) he o e head o OmpSs is in e sely p opo ional o he ask size. We can see ha he op imal g ain size is always lowe han o equal o 32. In he de ailed case, he imbalance domina es he ade-o because he a e age cos o he chemis y in eg a ion pe elemen is la ge compa ed o he OmpSs o e heads. Conside ing he op imal g ain size o each con igu a ion, he speedup ob ained anges be ween 2×and 2.3×wi h he 48 ×1 con igu a ion, we achie e 86% o he maximum speedup. Mo eo e , a low a iance be ween di e en con igu a ions is obse ed, which ein o ces p e ious obse a ions abou he use o DLB o isola e he pe o mance om he selec ed con igu a ion. 1 2 3 4 5 6 7 8 48x1 24x2 12x4 8x6 6x8 4x12 Speed up In eg a ion Con ig (MPIs x OmpSs h eads) 4 8 16 32 64 128 Figu e 13: Compa ison o g ain size impac when using DLB o educed chemis y in eg a ion (case CF2) o di e en numbe o h eads. The same s udy was done o he educed chemis y case CF2 and i is shown in Fig- u e 13 whe e we obse e mo e a iabili y wi h espec o he g ain size. No e ha , since he a e age chemis y in eg a ion cos pe elemen is much lowe , he o e head becomes signi ican o low g ain sizes. Consequen ly, o all con igu a ions he op imal g anula i y is always equal o o g ea e han 32. Conside ing he op imal g ain size o each con igu- a ion, he speedup ob ained anges be ween 6.4×and 7×achie ing 91% o he maximum speedup o he 48 ×1 con igu a ion. Rega ding he whole ime-s ep, he implemen a ion o Alya is no comple ely hyb id wha makes mo e adequa e he con igu a ion o 48 ×1 o p oduc ion simula ions since o he wise, o he pa s o he ime-s ep would be penalized because some CPU-co es e- se ed o OmpSs h eads would no be used. Fo he de ailed chemis y model CF1, i s in eg a ion ep esen s 59% o he o e all ime s ep, so he 2.4×accele a ion achie ed wi h 22 he 48 ×1 con igu a ion esul s in a 1.6×o e all accele a ion. In he educed chemis y scena io CF2, i s sha e is 26% o he ime-s ep, he e o e i s 7×accele a ion esul s in a 1.4×o e all accele a ion. In his sec ion, we ha e shown ha DLB can imp o e he pe o mance o he chemical in eg a ion loop u he han he hyb idiza ion o he code. Mo eo e , i is con i med ha he speedup achie ed by DLB depends on he imbalance p esen in he o iginal execu ion achie ing a 7×speedup o highly imbalanced uns wi h he same numbe o esou ces. I is seen ha he selec ion o he g ain size has an impac on he pe o mance o DLB. Finally, la ge g ain sizes o e less lexibili y o DLB o load balance, so he bes ade-o be ween lexibili y and o e head is a ound a g ain size o 32 o all he cases. 5.3. Mul i-node es ing As explained in Sec ion 4.2, DLB can only load balance wi hin a compu a ional node wi h sha ed memo y, so in his sec ion we demons a e how DLB can imp o e he pe o - mance o mul i-node execu ions e en hough i is only ac ing locally inside he node. When conside ing mul i-node simula ions i is impo an o no ice ha , by de aul , esou ce manage s spawn he MPI p ocesses con iguously in he di e en compu a ional nodes. Wi h he con inuous binding, subdomains associa ed wi h p ocesses unning in he same node end o be adjacen in he domain. Wi h he Round Robin (RR) binding, his locali y is a oided on pu pose o make each node o ha e di e en pa s o he domain. This s a egy is well sui ed o combus ion simula ions, since chemical eac ions usually occu in speci ic loca ions o he domain and along hin laye s, so he MPI anks con aining he eac ing laye s end o be close o each o he . In o de o dis ibu e he mos loaded p ocesses among he di e en nodes and imp o e he pe o mance o he load balancing mechanism o DLB, i is usually be e o use a Round Robin (RR) dis ibu ion o MPI p ocesses. Figu e 14: Con iguous s. Round Robin dis ibu ion o MPI anks. 23 In Figu e 14, we show an example o a combus ion domain (le hand side) pa i ioned be ween 8 MPI anks. This pa i ion assigns he nodes o he eac ing laye , which is he mos compu a ionally expensi e o MPI ank 6 and 7. When a con iguous dis ibu ion o MPI anks among nodes is done, he si ua ion depic ed in he op igh hand side pa o he igu e is ound. Whe e MPI anks 1, 2, 3 and 4 a e assigned o Node 1 and MPI, anks 5, 6, 7 and 8 a e assigned o Node 2. Wi h his dis ibu ion, he wo mo e loaded p ocess a e assigned o Node 2, p oducing a load imbalance ac oss nodes ha can no be add essed by DLB. Howe e , conside ing a Round Robin dis ibu ion ins ead, bo om igh hand side, Node 1 ge s MPI anks 1, 3, 5 and 7, and Node 2 ge s MPI anks 2, 4, 6 and 8. A good load balance ac oss nodes is ound wi h his dis ibu ion, so he load imbalance inside he nodes can be add essed wi h DLB. This aspec is in es iga ed he e by he analysis o a coun e low di usion lame, co e- sponding o cases CF3 and CF4 wi h de ailed and educed chemis y, espec i ely. These cases di e om he p e ious cases CF1 and CF2 by ha ing a ine mesh wi h abou 4 imes la ge compu a ional load and hen i will be ex ended o u bulen p emixed lames, cases SB1 and SB2. Fo hese cases, we conside he ange om 1 node (48 CPU-co es) up o 16 nodes (768 CPU-co es). No e ha he chemis y in eg a ion does no equi e any MPI communica ion o synch oniza ion ope a ion and, he e o e, he mos limi ing ac o o scalabili y is he imbalance. All he execu ions in his sec ion a e done using a con ig- u a ion o 48 ×1 o MPI p ocesses and OmpSs h eads pe node, as we wan o mimic a p oduc ion un o Alya and a g ain size o 32, because i is he op imum alue de e mined in he p e ious sec ion. 0 2 4 6 8 10 12 14 48 96 192 384 768 Spped up # MPI anks pu e MPI con iguous Hyb id + DLB Con iguous pu e MPI Round Robin Hyb id + DLB Round Robin Figu e 15: Speedup-up o de ailed chemical in eg a ion (case CF3) up o 16 nodes using DLB and a ying dis ibu ion o MPI anks among nodes, wi h a con igu a ion o 48 ×1 and g ain size 32. 24 In Figu e 15, we see he speedup (Y axis) o he chemis y in eg a ion s age no malized by he MPI-only execu ion on 48 co es as unc ion o he numbe o MPI anks used in he simula ion. In hese plo s, solid lines ep esen con iguous binding, o which MPI anks a e placed con iguously in he nodes, while dashed lines a e used o ep esen he Round Robin binding, whe e MPI anks a e spawned in a Round Robin mode among he compu e nodes. MPI only execu ions a e ep esen ed wi h blue lines, while hyb id DLB execu ions a e ep esen ed wi h o ange lines. We can see ha he binding has low impac on he pe o mance o he pu e MPI imple- men a ion (blue lines o e lap). Con a ily, when DLB is used, he RR binding is help ul o b eak he subdomain’s locali y and a oid si ua ions whe e he subdomains associa ed wi h p ocesses o a node co e egions wi h simila condi ioning. In his case, we can see ha o 2 nodes (96 MPI anks) he e is almos no di e ence. Bu o highe numbe o nodes DLB is able o ob ain a be e speedup when using a Round Robin dis ibu ion. Wi h 768 MPI anks, i ob ains a 12×speedup wi h Round Robin e sus a 10×speedup wi h con- iguous dis ibu ion. Addi ionally, DLB imp o es he pe o mance o he simula ion by a ac o o 2×wi h espec o he o iginal pu e MPI un when using 16 nodes. 0 5 10 15 20 25 48 96 192 384 768 Speed up # MPI anks pu e MPI con iguous Hyb id + DLB Con iguous pu e MPI Round Robin Hyb id + DLB Round Robin Figu e 16: Speedup-up o educed chemical in eg a ion (case CF4) up o 16 nodes using DLB and a ying dis ibu ion o MPI anks among nodes, wi h a con igu a ion o 48 ×1 and g ain size 32. In Figu e 16, we see he same speedup, bu o he educed chemis y case CF4. The blue lines co espond o uns o he o iginal MPI-pu e code, while o ange lines co espond o hyb id uns wi h DLB. I is obse ed ha he dis ibu ion o MPI p ocesses among nodes does no ha e an impac on he pe o mance o he o iginal pu e MPI code (blue lines o e lap), as i also occu s wi h he de ailed chemis y case. Howe e , we obse e ha he impac o he ound Robin dis ibu ion in his case is e en highe han o he 25 [31] D. Th´ e enin, F. Beh end , U. Maas, B. P zywa a, and J. Wa na z. De elopmen o a pa allel di ec simula ion code o in es iga e eac i e lows. Compu e s and Fluids, 25(5):485–496, 1996. [32] M. Vazquez, G. Houzeaux, S. Ko ic, A. A igues, J. Aguado-Sie a, R. A is, D. Mi a, H. Calme , F. Cucchie i, H. Owen, A. Taha, J. M. Cela, and M. Vale o. Mul iphysics enginee ing simula ion owa d exascale. J. Compu . Sci., 14:15 – 27, 2016. [33] M. Wagne , S. Moh , J. Gim´ enez, and J. Laba a. A s uc u ed app oach o pe o - mance analysis. In In e na ional Wo kshop on Pa allel Tools o High Pe o mance Compu ing, pages 1–15. Sp inge , 2017. [34] T. Zi wes, F. Zhang, J. A. Dene , P. Habis eu he , and H. Bockho n. Au oma ed code gene a ion o maximizing pe o mance o de ailed chemis y calcula ions in open oam. In Wol gang E. Nagel, Die ma H. K ¨ one , and Michael M. Resch, edi- o s, High Pe o mance Compu ing in Science and Enginee ing ’ 17, pages 189–204, Cham, 2018. Sp inge In e na ional Publishing. [35] T. Zi wes, F. Zhang, J.A. Dene , P. Habis eu he , and H. Bockho n. Au oma ed code gene a ion o maximizing pe o mance o de ailed chemis y calcula ions in open- oam. High Pe o mance Compu ing in Science and Enginee ing’17, pages 189–204, 2018. 32