E icien mul ilaye shallow-wa e simula ion sys em based on GPUs
Miguel Las aa,∗, Manuel J. Cas o D´ıazb, Ca los U e˜naa, Ma c de la Asunci´ona
aDep o. Lenguajes y Sis emas In o m´a icos. E.T.S Ingen´ıe ia In o m´a ica. Pe iodis a Daniel Saucedo A anda s/n.
Uni e sidad de G anada.
bDep o. An´alisis Ma em´a ico. Facul ad de Ciencias.Campus de Tea inos s/n. Uni e sidad de M´alaga.
Abs ac
The compu a ional simula ion o shallow s a i ied luids is a e y ac i e esea ch opic because hese ypes
o sys ems a e e y common in a a ie y o na u al en i onmen s. The simula ion o such sys ems can
be modelled using mul ilaye shallow-wa e equa ions bu do impose impo an compu a ional equisi es
especially when applied o la ge domains.
Gene al Pu pose Compu ing on G aphics P ocessing Uni s (GPGPU) has become a i id esea ch ield
due o he a i al o massi ely pa allel ha dwa e pla o ms (based on g aphics ca ds) and adequa e p o-
g amming amewo ks which ha e allowed impo an speed-up ac o s wi h espec o no only sequen ial
bu also pa allel CPU based simula ion sys ems.
In his wo k we p esen a p oposal o he simula ion o shallow s a i ied luids wi h an a bi a y numbe
o laye s using GPUs. The designed sys em does ully adap o he many-co e a chi ec u e o mode n GPUs
and se e al expe imen s ha e been ca ied ou o illus a e i s scalabili y and beha io on di e en GPU
models. We p opose a new elabo a ed 3D compu a ional scheme o an unde lying 2D ma hema ical model.
This scheme allowed implemen ing a sys em capable o handling an a bi a y numbe o laye s. The sys em
adds no o e head when used o wo-laye scena ios, compa ed o an exis ing 2D sys em speci ically designed
o jus wo laye s.
Ou p oposal is aimed a c ea ing a GPU-based compu a ional scheme sui able o he simula ion o
mul ilaye la ge-scale eal-wo ld scena ios.
Keywo ds: Shallow-wa e , ini e olumes, simula ion, GPU, CUDA
1. In oduc ion
The simula ion o ee su ace o in e nal wa es in shallow s a i ied luids a e commonly modelled by
he mul ilaye shallow-wa e equa ions o mula ed as a conse a ion law wi h non-conse a i e p oduc s and
sou ce e ms. S a i ied luids a e ubiqui ous in na u e: hey appea in a mosphe ic lows, ocean cu en s, es-
ua ine sys ems,... This is he si ua ion, o ins ance, in he S ai o Gib al a , whe e su ace wa e om he
A lan ic in lows o e sal ie wes wa ds- lowing Medi e anean wa e . Simula ing hose phenomena equi es
e y long las ing simula ions in big compu a ional domains which is why e y e icien implemen a ions a e
needed o be able o analyze hose p oblems in a o dable compu a ional imes.
A 2D mul ilaye shallow-wa e sys em can be disc e ized by he na u al ex ension o 2D domains o a
i s o de PVM pa h-conse a i e ype ini e olume scheme in oduced in [8]. PVM schemes ha e been
in oduced in [8] in he amewo k o balance laws and non-conse a i e hype bolic sys ems. They a e
de ined in e ms o iscosi y ma ices compu ed by a sui able polynomial e alua ion o a Roe Ma ix. These
∗Co esponding au ho
Email add esses: [email p o ec ed] (Miguel Las a), [email p o ec ed] (Manuel J. Cas o D´ıaz), [email p o ec ed]
(Ca los U e˜na), [email p o ec ed] (Ma c de la Asunci´on)
1Co esponding au ho . Tel: +34958246144
P ep in
me hods ha e he ad an age ha hey only need some in o ma ion abou he eigen alues o he sys em and
no spec al decomposi ion o he Roe Ma ix is equi ed. As consequence, hey a e much as e han Roe
schemes o sys ems wi h an inc easing numbe o unknowns. In his wo k he PVM-2U me hod has been
chosen among he PVM schemes in oduced in [8] as i p o ides he bes esul s, conce ning compu a ional
ime and accu acy. This me hod can be seen as a he na u al ex ension o he scheme in oduced by Degond
e al. in [7] o non-conse a i e sys ems.
Since he appea ance o compu ing amewo ks like Cuda [14] and openCL [12] he huge compu a ional
capabili ies o mode n G aphics P ocessing Uni s (GPUs) ha e been un eiled o compu a ionally in ensi e
physically based simula ions. Ne e heless, elabo a ed so wa e designs a e equi ed o ake ull ad ance o
he p ocessing powe o e ed by GPUs.
The Cuda compu a ional model, p o ided by N idia [15], equi es he compu a ional asks o be mapped
o a se o h eads ha will be un in pa allel. Each o hese h eads will un he same p og am (ke nel)
whe e he numbe o h eads unning a a poin in ime will depend on he ha dwa e equi emen s o each
h ead. A he lowe le el, h eads a e un in g oups o 32 h eads called wa ps and i is necessa y o a oid
code di e gences inside a wa p o a oid se ializa ion. On he o he hand, coalesced memo y accesses a e
also essen ial o allow se e al h eads access he equi ed da a s o ed in memo y by only issuing one physical
memo y access o a g oup o h eads. When a wa p equi es a memo y access ope a ion, he whole wa p is
s opped un il ha memo y ope a ion has inished and ano he wa p is un. Mode n GPUs p o ide a la ge
amoun o egis e s o a oid cos ly con ex swi ches by keeping a la ge amoun o h eads in a eady o un
s a e. Taking in o accoun GPUs cu en ly p o ide in he o de o housands o compu ing co es, a la ge
amoun o h eads is equi ed o a oid compu ing co es no ecei ing enough wo kload and idling. This
means any simula ion scheme mus be designed o p o ide a high deg ee o h ead le el pa allelism. The
highe he numbe o h eads ha can be kep ac i e, he highe he alue o a me ic called occupancy.
On he o he hand, as in many a chi ec u es, compu ing co es ha e pipelines which should be kep as
ull as possible o maximize he pe o mance. This equi es an adequa e le el o ins uc ion le el pa allelism
which means ha ing ce ain deg ee o independence be ween he ins uc ions un by each h ead. This ac
equi es ha each h ead mus also ecei e an adequa e le el o wo kload and as much independence as
possible om memo y accesses (a i hme ic in ensi y).
The compu ing co es in N idia b anded ha dwa e a e g ouped in o mul ip ocesso s (SMXs) and h eads
a e also g ouped in o h ead blocks. Each h ead block is un on he same SMX and his allows using an
in a-block synch oniza ion mechanism and also in a-block collabo a ion h ough he use o a high-speed
sha ed memo y (sha ed among he h eads o each block).
Di iding he compu a ions ha make up he simula ion p og am in o independen h eads and g ouping
hem in o blocks equi es also o conside he esou ces equi ed by each h ead and se e al ha dwa e le el
limi s like he maximum esiden h eads and blocks pe SMX o keep he occupancy le el high (al hough
a highe occupancy does no au oma ically ansla e in o a highe pe o mance).
The e a e many examples in li e a u e o success ully using GPUs o shallow-wa e simula ion [5, 4, 2,
10, 3], e en be o e amewo ks like Cuda we e a ailable, o widely used, and physical simula ions we e un
using he s anda d g aphics pipeline [11, 13]. In his wo k we concen a e on mul ilaye en i onmen s wi h
an a bi a y numbe o laye s simula ed using Cuda on N idia GPUs.
This pape is o ganized as ollows: in Sec ion 2 he ma hema ical model ha has been used is desc ibed,
Sec ion 3 desc ibes de nume ical model, Sec ion 4 add esses he mul ilaye compu a ional model p oposed
in his wo k and Sec ion 5 p esen s he se o expe imen s ha ha e been pe o med. Finally, Sec ion 6
p esen s he conclusions ha we e eached.
2
2. Ma hema ical model
Le us conside he sys em o equa ions go e ning he 2D low o msupe posed immiscible laye s o
shallow homogeneous luids wi h cons an densi ies in a subdomain D⊂R2:
∂ hl+∂xqx,l +∂yqy,l = 0,
∂ qx,l +∂x q2
x,l
hl
+1
2gh2
l!+∂yqx,lqy,l
hl+ghl∂x X
k>l
hk+X
k<l
ρk
ρl
hk−H!= 0.
∂ qy,l +∂xqx,lqy,l
hl+∂y q2
y,l
hl
+1
2gh2
l!+ghl∂y X
k>l
hk+X
k<l
ρk
ρl
hk−H!= 0
(1)
whe e l= 1,· · · , m being mis he numbe o laye s, hl(x, ) and ql(x, ) = (qx,l(x, ), qy,l(x, )) a e, espec-
i ely, he hickness and he mass- low o he l- h laye a poin xa ime , and hey a e ela ed o he
mean eloci ies ul(x, )=(ux,l(x, ), uy,l(x, )), by he equali ies: ql(x, ) = ul(x, )hl(x, ); H(x) is he
basin dep h measu ed om a ixed e e ence le el, gis he g a i y cons an and ρj he densi ies o each laye
e i ying
0< ρ1<· · · < ρm.
No ice ha h1is he heigh o he laye o luid on he op and hmis he heigh o he laye o luid o e
he bo om (See Figu e 1).
Figu e 1: Mul ilaye s a i ied low
This sys em can be w i en unde he s uc u e o a sys em o balance laws wi h non-conse a i e
p oduc s
∂ w+∂xFx(w) + ∂yFy(w) + Bx(w)∂xw+By(w)∂yw=Gx(w)∂xH+Gy(w)∂yH(2)
whe e
w=
w1
.
.
.
wm
, wl=
hl
qx,l
qy,l
=
hl
ux,lhl
uy,lhl
,
3
Fx(w) =
Fx,1
.
.
.
Fx,m
, Fx,l =
ux,lhl
u2
x,lhl+1
2gh2
l
ux,luy,lhl
,
Fy(w) =
Fy,1
.
.
.
Fy,m
, Fy,l =
uy,lhl
ux,luy,lhl
u2
y,lhl+1
2gh2
l
Gx(w) =
Gx,1
.
.
.
Gx,m
, Gx,l =
0
ghl
0
, Gy(w) =
Gy,m
.
.
.
Gy,m
, Gy,l =
0
0
ghl
.
Bx(w) is he 3m×3mma ix de ined by
Bx(w) =
0B1,2
xB1,3
x· · · B1,m
x
B2,1
x0B2,3
x· · · B2,m
x
.
.
..
.
.0· · · .
.
.
Bm,1
xBm,2
xBm,3
x· · · 0
whe e Bl,k
x,k, l = 1,· · · , m,k=la e 3 ×3 ma ices de ined by
Bl,k
x=
0 0 0
ρk/ρlhl0 0
0 0 0
i k < l
0 0 0
hl0 0
0 0 0
o he wise.
By(w) is he 3m×3mma ix de ined by
By(w) =
0B1,2
yB1,3
y· · · B1,m
y
B2,1
y0B2,3
y· · · B2,m
y
.
.
..
.
.0· · · .
.
.
Bm,1
yBm,2
yBm,3
y· · · 0
whe e Bl,k
y,k, l = 1,· · · , m,k=la e 3 ×3 ma ices de ined by
Bl,k
y=
0 0 0
0 0 0
ρk/ρlhl0 0
i k < l
0 0 0
0 0 0
hl0 0
o he wise.
Le us de ine he ma ices Aα(w) = Jα(w) + Bα(w), α=x, y whe e Jα(w) = ∂Fα
∂w (w) a e he Jacobians
o he luxes Fα, and we assume ha w∈Ω⊂R3m, so ha (1) is s ic ly hype bolic.
4
Le us also ema k ha he sys em (1) e i ies he p ope y o in a iance by o a ions. E ec i ely, gi en
η∈R2wi h ∥η∥= 1, le us de ine
Tη=
R1,1
η0· · · 0
0R2,2
η· · · 0
.
.
..
.
..
.
..
.
.
0 0 · · · Rm,m
η
, Rl,l
η=
1 0 0
0ηxηy
0−ηyηx
, l = 1,· · · , m
and le us deno e Fη(w) = Fx(w)ηx+Fy(w)ηy,B(w)=(Bx(w), By(w)), and G(w) = (Gx(w), Gy(w)) hen
Fη(w) = T−1
ηFx(Tηw), TηB(w)·η=Bx(Tηw), TηG(w)·η=Gx(Tηw).(3)
Mo eo e , i is easy o check ha Tηw e i ies he sys em
∂ (Tηw) + ∂ηFx(Tηw) + Bx(Tηw)∂ηw=Gx(Tηw)∂ηH−Rη⊥,(4)
whe e Rη⊥=Tη∂η⊥Fη⊥(w) + B(w)·η⊥∂η⊥w−G(w)·η⊥∂η⊥H.
Finally, le us ema k ha he e is no explici o mula o he eigen alues and eigen ec o s o ma ix
Aη(w) = Ax(w)ηx+Ay(w)ηy, bu i is well known in oceanog aphy ha he as es and he slowes wa e o
sys em (1) could be app oxima ed by he exp essions
λη
L= ¯uη−
u
u
g
m
X
l=1
hl, λη
R= ¯uη+
u
u
g
m
X
l=1
hl(5)
whe e ¯uηis de ined by
¯uη=Pm
l=1 ul·η hl
Pm
l=1 hl
.(6)
3. Nume ical Scheme
To disc e ize (1) he compu a ional domain Dis decomposed in o subse s wi h a simple geome y, called
cells o ini e olumes: Vi⊂R2. He e, i is assumed ha he cells a e ec angula wi h edges pa allel o he
Ca esian axes. Le us deno e by T he mesh and by NV he numbe o cells.
Gi en a ini e olume Vi,|Vi|will ep esen i s a ea; Ni∈R2i s cen e ; Ni he se o indexes jsuch ha
Vjis a neighbo o Vi;Eij he common edge o wo neighbo ing cells Viand Vj, and |Eij |i s leng h (∆x
( espec i ely ∆y) he leng h o he ho izon al ( espec i ely e ical) edges); ηij = (ηij,x, ηij,y) he no mal
uni ec o a edge Eij poin ing owa ds he cell Vjand wn
i he app oxima ion o he a e age o he solu ion
in he cell Via ime np o ided by he nume ical scheme:
wn
i∼
=1
|Vi|ZVi
w(x, n)dx.
Le us now b ie ly desc ibe he nume ical scheme ha we use o app oxima e he solu ion o he sys em
(1). He e, we use he na u al ex ension o 2D p oblems o he PVM-2U scheme in oduced in [8] :
wn+1
i=wn
i−∆
|Vi|X
j∈Ni
|Eij|F−
ij (7)
whe e F−
ij is de ined aking in o accoun he p ope y o in a iance by o a ions o sys em (1). The ollowing
s eps a e pe o med o de ine F−
ij :
5
1. We de ine
wηij =h1, qηij
1, h2, qηij
2,· · · , hm, qηij
m, T=Tηij (w)[1,2,4,5··· ,3m−2,3m−1],
and
wη⊥
ij =qη⊥
ij
1, qη⊥
ij
2,· · · , qη⊥
ij
mT
=Tηij (w)[3,6,··· ,3m],
whe e w[i1,··· ,is]is he ec o de ined om ec o w, using i s i1- h, . . . , is- h componen s, qηij
l=ql·ηij
and qη⊥
ij
l=ql·η⊥
ij,l= 1,· · · , m.
2. Le Φ−
ηij be he 1D nume ical PVM-2U lux associa ed o he 1D mul ilaye shallow-wa e sys em
de ined using he 1-s , 2-nd, 4- h and 5- h, · · · , (3m−2)- h and (3m−1)- h equa ions o sys em (4)
whe e he e m Rη⊥
ij has been neglec ed:
Φ−
ηij =1
2(F(wηij
j)− F(wηij
i) + Bij(wηij
j−wηij
i)− Gij(Hj−Hi) (8)
−Qij((wηij
j−wηij
i)−(A∗
ij)−1Gij(Hj−Hi)) + FC(wηij
i)
(9)
whe e
F(wηij ) =
uηij
1h1
(uηij
1)2h1+g
2h2
1
uηij
2h2
(uηij
2)2h2+g
2h2
2
.
.
.
uηij
mhm
(uηij
m)2hm+g
2h2
m
,Gij =
0
ghij
1
0
ghij
2
.
.
.
0
ghij
m
,FC(wηij ) =
uηij
1h1
(uηij
1)2h1
uηij
2h2
(uηij
2)2h2
.
.
.
uηij
mhm
(uηij
m)2hm
,
Bij(wηij
j−wηij
i) =
0
ghij
1 m
X
k=2
(hk,j −hk,i)!
0
ghij
2 m
X
k=3
(hk,j −hk,i) +
1
X
k=1
ρk
ρ2
(hk,j −hk,i)!
.
.
.
0
ghij
m m−1
X
k=1
ρk
ρm
(hk,j −hk,i)!
and
hij
l=hl,j +hl,i
2, l = 1,· · · , m.
Ma ix Qij is de ined by
Qij =αij
0Id +αij
1Aij +αij
2A2
ij (10)
6
wi h
αij
0=(SM)2Sm(sign(Sm)−sign(SM))
(Sm−SM)2,
αij
1=SM(|SM|−|Sm|) + Sm(sign(SM)Sm−SMsign(Sm))
(Sm−SM)2,
αij
2=Sm(sign(Sm)−sign(SM))
(Sm−SM)2,
(11)
whe e
SM=
Sij
Li |Sij
L|≥|Sij
R|,
Sij
Ro he wise
(12)
and
Sm=
Sij
Ri |Sij
L|≥|Sij
R|,
Sij
Lo he wise.
(13)
Sij
Rand Sij
La e wo es ima ions o he as es and slowes wa es, espec i ely, ela ed o he Riemann
p oblem associa ed o edge Eij. He e, he ollowing es ima ions a e used:
Sij
R= max ληij
R,j , λij
R, Sij
L= min ληij
L,i , λij
L,
being
ληij
R,j =uηij
j+
u
u
g
m
X
l=1
hl,j, λij
R=uηij
ij +
u
u
g
m
X
l=1
hij
l
ληij
L,i =uηij
i−
u
u
g
m
X
l=1
hl,i, λij
L=uηij
ij −
u
u
g
m
X
l=1
hij
l
whe e
uηij
k=Pm
l=1 hl,k ul,k ·ηij
Pm
l=1 hl,k
, k =i, j
and
uηij
ij =Pm
l=1 uij
lhij
l
Pm
l=1 hij
l
,
wi h
uij
l=phl,juηij
l,j +phl,iuηij
l,i
phl,j +phl,i
, l = 1,· · · , m and uηij
l,k =ul,k ·ηij , k =i, j.
7
Aij is he 2m×2mma ix de ined by
Aij =
0 1 0 0 · · · 0 0
ghij
1−(uij
1)22uij
1ghij
10· · · ghij
10
0 0 0 1 · · · 0 0
ρ1
ρ2
ghij
20ghij
2−(uij
2)22uij
2· · · ghij
20
.
.
..
.
..
.
..
.
.· · · .
.
..
.
.
0 0 0 0 · · · 0 1
ρ1
ρm
ghij
m0ρ2
ρm
ghij
m0· · · ghij
m−(uij
m)22uij
m
(14)
and A∗
ij is de ined om Aij by se ing uij
l= 0, l= 1,· · · , m.
Taking in o accoun he exp ession o Qij (10) and A∗
ij, Φ−
ηij can be w i en as
Φ−
ηij =1
2(Eij −(αij
0e
Iij +αij
1Eij +αij
2AijEij)) + FC(wηij
i) (15)
whe e
Eij =F(wηij
j)− F(wηij
i) + Bij(wηij
j−wηij
i)− Gij(Hj−Hi)
and
Iij =
1
j− 1
i−( 2
j− 2
i)
qηij
1,j −qηij
1,i
2
j− 2
i−( 3
j− 3
i)
qηij
2,j −qηij
2,i
.
.
.
m
j− m
i
qηij
m,j −qηij
m,i
wi h
l
α=
m
X
k=l
hk,α −Hα, α =i, j.
Finally, le us ema k ha Eij can be w i en in an equi alen way as
Eij =FC(wηij
j)− FC(wηij
i) + Pij
8
whe e Pij is a disc e iza ion o he p essu e e ms gi en by
Pij =
0
ghij
1 1
j− 1
i
0
ghij
2ρ1
ρ2( 1
j− 1
i)−( 2
j− 2
i)+ ( 2
j− 2
i)
.
.
.
0
ghij
m m−1
X
k=1
ρk
ρm( k
j− k
i)−( k+1
j− k+1
i)+ ( m
j− m
i)!
.
Equi alen ly he lux co esponding o cell wjcan be de ined as:
Φ+
ηij =1
2(Eij + (αij
0e
Iij +αij
1Eij +αij
2AijEij)) − FC(wηij
j).(16)
3. Le us de ine
Φ−
η⊥
ij
=
(Φ−
ηij )[1] uη⊥
ij ,∗
1
(Φ−
ηij )[3] uη⊥
ij ,∗
2
.
.
.
(Φ−
ηij )[2m−1] uη⊥
ij ,∗
m
,(17)
whe e uη⊥
ij ,∗
lis de ined as ollows
uη⊥
ij ,∗
l=
qη⊥
ij
l,i
hl,i
i (Φ−
ηij )[2l−1] >0
qη⊥
ij
l,j
hl,j
o he wise
l= 1,· · · , m.
Le us ema k ha Φ−
η⊥
ij
is he nume ical lux associa ed o he 1-s , 6- h, · · · , 3m- h equa ions o
sys em (4) whe e, again, he e m Rη⊥
ij has been neglec ed. I s de i a ion has been done ollowing he
main ideas o he HLLC me hod o he shallow-wa e sys em in oduced in [9] as qη⊥
ij
l,l= 1,· · · , m
can be seen as a passi e scala ha is con ec ed by he low. Equi alen ly, he lux co esponding o
cell wjcan be de ined as
Φ+
η⊥
ij
=−Φ−
η⊥
ij
.
9
o an a bi a y numbe o laye s. In he mul ilaye sys em he bigge ke nel is he one associa ed o he
compu a ion o F−
l,ij and λij
max which equi es 42 egis e s. I can he e o e be concluded ha he s a egy
o a ine -g ained design o he mul ilaye sys em o e s a be e pe o mance o his ype o p oblem as
i inc eases he h ead le el pa allelism while achie ing a su icien ins uc ion le el pa allelism o each
ke nel ins ance. This inc eased pe o mance comes a he cos o a la ge memo y oo p in because o he
in e media e da a s o age equi emen s.
5.1.1. Compa ison wi h a CPU wo-laye sys em
The mul ilaye code was also compa ed o a na i e CPU wo-laye implemen a ion. The same 10 seconds
o simula ion o he dam b eak men ioned in he p e ious sec ion wi h 256 ×256 olumes, equi es 1489
seconds on he CPU (In el Co e i7-2600 CPU @ 3.40GHz). Using ou OpenMP h eads, he un ime is 412
seconds. This esul s in a speed-up o o e 250×wi h espec o he single- h eaded CPU e sion and o e
70×wi h espec o he mul i h eaded CPU e sion. This CPU e sion includes no special op imiza ions
apa om he OpenMP pa alleliza ion and he compile op imiza ions p oduced by he -O3 compile lag.
5.2. Mul ilaye expe imen s
Two expe imen s we e pe o med o es he pe o mance o ou wo k. The i s expe imen used a domain
consis ing in a 10×10 squa e. The basin dep h is de ined by he unc ion: H(x)=1−0.5 exp (−10 (x−5)2)
and all laye s ha e a hickness alue o 1, excep o h ee laye s:
•The wo laye s a ec ed by he p esence o he simula ed dam. The dam heigh is 4.5 and he dis ance
o he op neighbo ing laye is 0.5.
•The heigh o laye mis de ined by H(x).
The unc ion ha de ines he hickness o each laye is shown in equa ion 21.
hl(x) =
H(x) i l=m
1 i l= 1,...,(d−2),(d+ 1),...,(m−1)
(5.5DA(x)) + (1 −DA(x)) i l=d
(0.5DA(x)) + 5 (1 −DA(x)) i l= (d−1)
(21)
whe e d=m
2is he laye on which he dam is de ined. In his expe imen he cen e o he ci cula dam is
loca ed a c= (2,5).
Rega ding he ini ial condi ions, ql(x,0) = 0, l = 1,...m. The CFL alue used was 0.9 and ρl= 2 ρl−1.
Wall bounda y condi ions we e imposed.
The es scena io scheme wi h six laye s and basin is shown in Figu e 7. The e a e also ideo cap u es o
di e en simula ion scena ios a ailable on he complemen a y ma e ial websi e: h p://dici s.ug .es/
so wa e/Mul iLaye SW/
In Table 2 he esul s o di e en numbe o laye s and 2D sizes a e shown. The numbe o olumes
(millions o ) p ocessed pe second is also shown conside ing he numbe o simula ion i e a ions pe o med,
and he o al numbe o compu a ional olumes.
In he i s es , he o al hickness o he wa e inc eases as laye s a e added, and so does he densi y
di e ence be ween he i s and he las laye . This esul s in a dec easing ∆ alue and he e o e an
inc easing numbe o i e a ions as he numbe o laye s inc eases.
Fo he second expe imen , he o al hickness o he wa e was se o 10. This means ha as laye s a e
added, he hickness o each laye dec eases. Addi ionally he densi y alues we e se so ha ρ1= 0.1 and
ρm= 1.0 and all he in e media e ρi alues a e in e pola ed lineally. The es o he expe imen pa ame e s
a e he same as o he p e ious one. This es scena io p oduces a ixed ∆ alue a each i e a ion s ep and
a cons an numbe o i e a ions. In his case, as laye s as added, he simula ion scena io emains ixed bu
mo e esolu ion is added o he dimension ep esen ed by he numbe o laye s.
The simula ion imes o he second expe imen a e shown o a 512 ×512 domain oge he wi h he
esul s o he p e ious expe imen in Figu e 8. This plo shows ha when keeping he numbe o i e a ions
16
#Volumes pe laye #Laye s Time #Vols/s (millions)
256 ×256 4 12.7s 110
256 ×256 8 34.1s 100
256 ×256 16 112.7s 79
512 ×512 4 93.7s 121
512 ×512 8 254.5s 109
512 ×512 16 865.6s 84
1024 ×1024 4 732.7s 125
1024 ×1024 8 1981.6s 113
1024 ×1024 16 6848.1s 86
Table 2: GTX680 - Mul ilaye expe imen s. Time exp essed in seconds
Figu e 7: Tes scena io wi h six laye s
cons an , doubling he numbe o laye s does only inc ease he un ime by a ac o be ween 2 and 3. This is
a e y good esul aking in o accoun he in e -laye dependencies become mo e impo an as he numbe
o laye s inc eases.
A compa ison using a Tesla K20 GPU (2495 Cuda co es @ 706Mhz) was also pe o med. This GPU
o e s a g ea e amoun o Cuda co es (60% mo e) han he GTX 680 (1536 Cuda co es @ 1058Mhz) bu
a a lowe clock speed (33% less). Table 3 shows he esul s ob ained. Finally, a single expe imen was
pe o med using a GTX Ti an GPU (2688 Cuda co es @ 876Mhz). This GPU p o ides 75% mo e co es han
he GTX680 a a 15% lowe clock a e and he esul s ob ained a e shown in able 4
#Volumes pe laye #Laye s GTX 680 ime Tesla K20 ime
256 ×256 8 34,1s 29s
512 ×512 16 865.6s 802s
Table 3: GTX 680-Tesla K20 mul ilaye expe imen . Time exp essed in seconds
Taking in o accoun he 1.5×speed-up ac o ob ained using he GTX Ti an and he up o 1.11× ac o
ob ained wi h he Tesla K20 wi h espec o he esul s o he GTX 680 model, we belie e he sys em shows
a good scalabili y wi h espec o he numbe o a ailable Cuda co es. The impo an clock a e di e ence
o he K20 model explains ha he esul s ob ained wi h his GPU a e wo se han he ones ha could be
expec ed by only conside ing i s numbe o Cuda co es.
17
Figu e 8: 512 ×512 simula ion imes
#Volumes pe laye #Laye s GTX 680 ime GTX Ti an ime
256 ×256 6 23s 15,2s
Table 4: GTX 680-GTX Ti an mul ilaye expe imen s. Time exp essed in seconds
5.3. Volume eo de ing
We ha e also analyzed he bene i s o he olume eo de ing scheme wi h espec o he ad-hoc ol-
ume/da a s uc u e mapping using he i s o he wo es scena ios in oduced in sec ion 5.2.
#Laye s Reo de ed olumes Ad-hoc o de
2 6.1s 5.8s
6 25.3s 26.1s
12 75.6s 82.37s
Table 5: Volume eo de ing bene i s. 65 535 (256 ×256) olumes pe laye . Time exp essed in seconds
Table 5 shows ha he eo de ing scheme does p oduce some o e head because i equi es a highe
amoun o ope a ions o compu e olume and edge indices and his ac p oduces a sligh ly highe un ime
o a low numbe o laye s. As he numbe o edges inc eases, which also inc eases he numbe o memo y
accesses, he bene i o he eo de ing is clea .
6. Conclusions
An e icien GPU based mul ilaye simula ion scheme has been p esen ed which allows an a bi a y
numbe o laye s as i is ully pa ame e ized. The ac ha i is based on a 2D ma hema ical model ha
18
has been implemen ed by using a 3D app oach, a highe le el o h ead le el pa allelism is achie ed which
o e s a good scalabili y as he numbe o laye s inc eases.
The e iciency o he p oposed simula ion scheme has been shown in e ms o he numbe o olumes
which a e p ocessed pe second and he speed-up alues wi h espec o p e ious wo-laye sys ems, using
bo h hei GPU and hei single and mul i- h eaded CPU e sion. The p oposed gene al scheme in oduces
no o e head when compa ed o exis ing GPU-based sys ems which we e speci ically designed o wo laye s.
Mo eo e , i achie es a un ime educ ion om one o wo o de s o magni ude when compa ed o mul i
and single- h eaded CPU implemen a ions. Ou simula ion so wa e p ocesses a ound 100 million o 3D
olumes pe second o scena ios wi h be ween 4 and 16 laye s on a single GPU.
The esul s show ha he simula ion scheme ha has been p esen ed can be applied o la ge-scale
p oblems o simula e eal-wo ld phenomena like lake o oceanic cu en s.
Acknowledgmen s
The wo k o M. J. Cas o was suppo ed in pa by he Spanish Go e nmen and FEDER h ough
he Resea ch p ojec MTM2012-38383-C02-01, and by he Andalusian Go e nmen h ough he p ojec s
P11-FQM8179 and P11-RNM7069
Re e ences
[1] M. de la Asunci´on, J.M. Man as, M.J. Cas o, P og amming CUDA-based GPUs o simula e wo-laye shallow wa e
lows, in: P. D’amb a, M. Gua acino, D. Talia (Eds.), Eu o-Pa 2010, olume 6272 o Lec u e No es in Compu e
Science, Sp inge , 2010, pp. 353–364.
[2] M. de la Asunci´on, J.M. Man as, M.J. Cas o, Simula ion o one-laye shallow wa e sys ems on mul ico e and CUDA
a chi ec u es, Jou nal o Supe compu ing 58 (2011) 206–214.
[3] M. de la Asunci´on, J.M. Man as, M.J. Cas o, E.D. Fe n´andez-Nie o, An MPI–CUDA implemen a ion o an imp o ed oe
me hod o wo-laye shallow wa e sys ems, Jou nal o Pa allel and Dis ibu ed Compu ing. Special Issue on Accele a o s
o High–Pe o mance Compu ing 72 (2012) 1065–1072.
[4] A.R. B od ko b, T.R. Hagen, K.A. Lie, J.R. Na ig, Simula ion and isualiza ion o he Sain -Venan sys em using GPUs,
Compu ing and Visualiza ion in Science 13 (2010) 341–353.
[5] A.R. B od ko b, M.L. Sæ a, Shallow wa e simula ions on mul iple GPUs, in: PARA 2010: S a e o he A in Scien i ic
and Pa allel Compu ing, Reykja ik (Iceland), pp. 56–66.
[6] M.J. Cas o, E.D. Fe n´andez-Nie o, A.M. Fe ei o, J.A. Ga c´ıa, C. Pa ´es, High o de ex ensions o Roe schemes o
wo-dimensional nonconse a i e hype bolic sys ems, Jou nal o Scien i ic Compu ing 39 (2009) 67–114.
[7] P. Degond, P.F. Pey a d, G. Russo, P. Villedieu, Polynomial upwind schemes o hype bolic sys ems, Comp es Rendus de
l’Acad´emie des Sciences - Se ies I - Ma hema ics 328 (1999) 479 – 483. doi:h p://dx.doi.o g/10.1016/S0764-4442(99)
80194-3.
[8] M.J.C. D´ıaz, E.D. Fe n´andez-Nie o, A class o compu a ionally as i s o de ini e olume sol e s: PVM me hods.,
SIAM J. Scien i ic Compu ing 34 (2012).
[9] E.D. Fe n´andez-Nie o, D. B esch, J. Monnie , A consis en in e media e wa e speed o a well-balanced HLLC sol e ,
Comp es Rendus Ma hema ique 346 (2008) 795 – 800. doi:h p://dx.doi.o g/10.1016/j.c ma.2008.05.012.
[10] M. Ge ele , D. Ribb ock, S. Mallach, D. G¨oddeke, A simula ion sui e o La ice-Bol zmann based eal- ime CFD applica-
ions exploi ing mul i-le el pa allelism on mode n mul i- and many-co e a chi ec u es, Jou nal o Compu a ional Science
2 (2011) 113–123.
[11] T.R. Hagen, J.M. Hjelme ik, K.A. Lie, J.R. Na ig, M.O. Hen iksen, Visual simula ions o shallow-wa e wa es, Simula-
ion Modelling P ac ice and Theo y 13 (2005) 716–726.
[12] Kh onos OpenCL Wo king G oup, The OpenCL Speci ica ion, Accessed No embe 2012.
[13] M. Las a, J.M. Man as, C. U e˜na, M.J. Cas o, J.A. Ga c´ıa, Simula ion o shallow-wa e sys ems using g aphics p ocessing
uni s, Ma hema ics and Compu e s in Simula ion 80 (2009) 598–618.
[14] NVIDIA, Cuda, h p://www.n idia.com/objec /cuda_home_new.h ml, Accessed 2013.
[15] NVIDIA, N idia, h p://www.n idia.com/, Accessed 2013.
[16] NVIDIA, Keple -The wo ld’s as es , mos e icien HPC a qui ec u e, h p://www.n idia.com/objec /n idia-keple .
h ml, Accessed Augus 2012.
[17] C. Pa ´es, Nume ical me hods o nonconse a i e hype bolic sys ems: A heo e ical amewo k, SIAM Jou nal on Nume -
ical Analysis 44 (2006) pp. 300–321.
[18] C. Pa ´es, M.J. Cas o, On he well-balance p ope y o Roe’s me hod o nonconse a i e hype bolic sys ems. Applica ions
o shallow-wa e sys ems, ESAIM: Ma hema ical Modelling and Nume ical Analysis 38 (2004) 821–852.
19