scieee Science in your language
[en] (orig)

Analysis of OpenACC Performance using Different Block Geometries

Abstract

Producción Científica

Read accessible full text

Analysis of OpenACC Performance using Different Block Geometries

Author: Barba Gutiérrez, Daniel,González Escribano, Arturo,Llanos Ferraris, Diego Rafael
Publisher: Universidad de Salamanca
Year: 2017
Source: https://uvadoc.uva.es/bitstream/10324/29139/1/cmmse2017_bg_17_04_25.pdf
P oceedings o he 17 h In e na ional Con e ence
on Compu a ional and Ma hema ical Me hods
in Science and Enginee ing, CMMSE 2017
4–8 July, 2017.
Analysis o OpenACC Pe o mance
Using Di e en Block Geome ies
Daniel Ba ba1, A u o Gonzalez-Esc ibano1and Diego R. Llanos1
1Depa amen o de In o ma ica, Uni e sidad de Valladolid, Spain
emails: [email p o ec ed],[email p o ec ed],[email p o ec ed]
Abs ac
OpenACC is a pa allel p og amming model o au oma ic pa alleliza ion o sequen-
ial code using compile di ec i es o p agmas. OpenACC is in ended o be used wi h
accele a o s such as GPUs and Xeon Phi. The di e en implemen a ions o he s an-
da d, al hough s ill in ea ly de elopmen , a e p ima ily ocused on GPU execu ion. In
his s udy, we analyze how he di e en OpenACC compile s a ailable unde ce ain
p emises beha e when he clauses a ec ing he unde lying block geome y implemen-
a ion a e modi ied. These clauses a e he Gang numbe , Wo ke numbe , and Vec o
Size de ined by he s anda d.
Key wo ds: OpenACC, GPU, block geome y, h ead geome y
1 In oduc ion
OpenACC is an open s anda d in ended o au oma ically pa allelize sequen ial code and
manage i s execu ion in accele a o s like GPUs o Xeon Phi cop ocesso s. I de ines a
numbe o compile di ec i es, also called p agmas. The main goal o OpenACC is o
educe bo h lea ning and coding ime in a po able way [1]. The e sion o he OpenACC
s anda d a he ime o w i ing is he 2.5 [2].
The OpenACC s anda d was ounded by N idia, CRAY, CAPS and PGI. The numbe
o membe s now is la ge , including bo h academic ins i u ions and companies like he Oak
Ridge Na ional Labo a o y, he Uni e si y o Hous on, AMD, and he Edinbu gh Pa allel
Compu ing Cen e (EPCC), among o he s.
The e a e se e al compile s ha implemen he OpenACC s anda d. The PGI com-
pile , de eloped by he Po land G oup (subsidia y o N idia) is being dis ibu ed as pa o
c
CMMSE ISBN: 978-84-617-8694-7
Analysis o OpenACC Pe o mance Using Di e en Block Geome ies
he N idia OpenACC Toolki unde a ee 90-day license. C ay Inc. has i s own OpenACC
compile , only a ailable o use wi h hei supe compu e s. Pa hscale Inc., a so wa e de-
elope o compile s and mul ico e so wa e, also has an OpenACC implemen a ion, he
ENZO compile .
Among he many academic o open-sou ce al e na i es he e a e he OpenUH compile
[3] by he Uni e si y o Hous on and accULL [4] om Uni e sidad de La Laguna (Spain).
This wo k p esen s a s udy on he impac o di e en alues o he clauses ha a ec
he unde lying block geome y o OpenACC-gene a ed code. GPUs a e e y sensi i e o
he geome y o he h ead-block chosen [5], and OpenACC makes use o he e ms “gang”,
“wo ke ” and “ ec o ” in o de o de ine di e en le els o pa allelism. Acco ding o [6], he
speci ica ion is ambiguous and his unc ionali y depends di ec ly on how each compile is
implemen ed. In his wo k, we measu e he impac o he choice o an app op ia e h ead-
block geome y when unning a ep esen a i e benchma k. By de aul , he geome y is
decided by he compile unless he speci ic clause is used inside he OpenACC di ec i e.
Ou aim is o compa e he esul ing beha iou among di e en compile s and op ions. Fo
h ead-block geome y es ing, we will modi y he mos ep esen a i e o he benchma ks,
es ing se e al combina ions o alues o he clauses o speci y gang, wo ke s and ec o s,
and analyzing he di e ences in execu ion ime o each compile . This will o e some
insigh on he implemen a ion o hese clauses on each compile .
Ou con ibu ion shows ha he decisions made by each compile is no always op imal,
bu manual uning o he di e en alues is no always possible o e e y compile .
The es o his pape is o ganized as ollows. Sec ion 2 desc ibes he selec ed compile s.
Sec ion 3 shows ou selec ed mic obenchma k o es ing he beha iou o he gene a ed code
when modi ying he block geome y. Sec ion 4 con ains he esul o ou analysis abou he
impac on pe o mance when changing he unde lying block geome y. Finally, Sec ion 5
concludes ou pape .
2 A ailable Compile s
We men ioned se e al compile s in he In oduc ion. In his sec ion we desc ibe wi h mo e
de ail he compile s we we e able o use o his s udy.
2.1 PGI Compile
The PGI Compile [7] is being de eloped by The Po land G oup, being owned by N idia.
This compile is equen ly p esen ed in webina s, wo kshops, and con e ences.
A he ime o w i ing his pape , he PGI compile is a ailable o download as pa o
he OpenACC Toolki om N idia. This oolki includes a 90-day ee ial, he possibili y
o acqui ing an academic license o a whole yea , o buying a comme cial license.
c
CMMSE ISBN: 978-84-617-8694-7
Daniel Ba ba, A u o Gonzalez-Esc ibano, Diego R. Llanos
2.2 accULL
The accULL [4] compile de eloped by Uni e sidad o La Laguna (Spain) is an open-sou ce
ini ia i e. accULL consis s on a s uc u e o wo laye s con aining YaCF [8] (Ye ano he
Compile F amewo k) and F angollo [9], a un ime lib a y. YaCF ac s as a sou ce- o-sou ce
ansla o i, while F angollo wo ks as an in e ace ha p o ides he mos common ope a ions
ound in accele a o s.
2.3 OpenUH
The OpenUH [3] compile , de eloped by he Uni e si y o Hous on (USA) is ano he open-
sou ce ini ia i e. I makes use o Open64, a discon inued open-sou ce op imizing compile .
3 Mic obenchma k Desc ip ion
In OpenACC, block size is de ined by gangs, wo ke s and ec o s.Thei choices a ec he
pe o mance on memo y-bound applica ions. To s udy his issue, we a e going o use a
e y simple ma ix addi ion implemen ed bo h in CUDA and OpenACC. Ou decision is
made by he ac ha he p oblem is emba assingly pa allel, memo y acceses a e pe ec ly
coalesced, and he compu a ional load pe global memo y access is low (memo y-bound
applica ion).
The s anda d only de ines he di e en le els o pa allelism, bu i is up o each imple-
men a ion o decide how a e hese le els exploi ed in he ac ual a chi ec u e. In o de o
ob ain compa able esul s, some de ails should be aken in o accoun . The CUDA e sion
needs o be implemen ed using elas ic ke nels, using a ixed numbe o blocks, which equals
he gang numbe in he OpenACC code. Also, since he OpenACC s anda d es ablishes
ha g id dimensions depend on he use o he collapse clause wi h nes ed loops, we ha e
decided o make a one-dimensional g id. We e alua e a numbe o blocks in he g id anging
om one o 2048 (using only powe s o wo). The sequen ial code can be seen in Fig. 1 and
he CUDA ke nel in Fig. 2.
In OpenACC he X dimension o he CUDA block ansla es o ec o leng h, whe eas
he Y dimension equals he wo ke numbe . We ha e also decided o use 512 h eads pe
block, ying each possible combina ion o X and Y dimension alues using powe s o wo.
4 E alua ion
In his sec ion we analyze he impac o di e en choices o he geome y o he unde lying
h ead-blocks in he OpenACC gene a ed code.
c
CMMSE ISBN: 978-84-617-8694-7
Analysis o OpenACC Pe o mance Using Di e en Block Geome ies
#p agma acc da a copyin(p_A[0:Size*Size],p_B[0:Size*Size]), copyou (p_C[0:Size*Size])
{
in j;
#p agma acc ke nels
#p agma acc loop independen gang(GANG), wo ke (WORKER), ec o (VECTOR)
o (j = 0; j < Size*Size; ++j)
{
p_C[j] = ALPHA*p_A[j] + p_B[j];
}
} //End da a egion
Figu e 1: Sequen ial Code o he Ma ix Addi ion
__global__ oid ma ixKe nel( loa * p_A, loa * p_B, loa * p_C)
{
in i e a ions = ((SIZE/blockDim.x)*(SIZE/blockDim.y)) / GANG;
in i e ;
o (i e = 0; i e < i e a ions; ++i e )
{
in x = (blockIdx.x + i e *GANG)*blockDim.x + h eadIdx.x;
in y = h eadIdx.y;
in i = ( x/SIZE)*blockDim.y + y;
in j = ( x%SIZE);
in o se = i*SIZE + j;
i (o se < SIZE*SIZE) p_C[o se ] = ALPHA*p_A[o se ] + p_B[o se ];
}
}
Figu e 2: CUDA Ke nel o he Ma ix Addi ion
c
CMMSE ISBN: 978-84-617-8694-7
Daniel Ba ba, A u o Gonzalez-Esc ibano, Diego R. Llanos
4.1 Expe imen al Se up
We used a N idia GTX Ti an Black o un he expe imen s. This GPU con ains 2 880
CUDA co es wi h a clock a e o 980MHz and 15 SMs. I has 6GB o RAM, and Compu e
Capabili y 3.5. The hos is a Xeon E5-2690 3 wi h 12 co es a a clock a e o 1.9GHz, and
64GB in ou 12GB modules.
The PGI compile is he one con ained in he N idia OpenACC Toolki , e sion 15.7-
0, published in Jul 13, 2015. We used OpenUH e sion 3.1.0 (published in No embe
4, 2015), based on Open64 e sion 5.0 and using GCC 4.2.0, p ebuil , downloaded om
he High Pe o mance Compu ing Tools g oup websi e [10]. accULL is e sion 0.4alpha
(published in No embe 28, 2013), downloaded om Uni e sidad de La Laguna’s esea ch
g oup “Compu aci´on de Al as P es aciones” [11].
4.2 Block Geome y Sensibili y o Gene a ed Code
Obse ing he esul s ob ained using he CUDA code in Fig. 3, we can see ha he bes
esul s a e ob ained when he block numbe is high enough o make he GPU each p ope
le els o occupa ion. The X dimension plays a huge ole, bu we expec ed o see an im-
p o emen in pe o mance when he X dimension was a leas 16. We suspec his is due o
se e al ac o s, being he mos impo an he beha iou o he cache when elas ic ke nels
a e used.
The esul s o he OpenACC code gene a ed by PGI (Fig. 4) show ha he only ac o
a ec ing he pe o mance is he gangs numbe , whe eas he a ia ion o wo ke s numbe
and ec o leng h does no play a signi ican ole excep when a ec o leng h o one is
used. In his case, pe o mance is se e ely a ec ed, which indica es a poo use o he cache.
O e all, PGI’s beha iou is he closes one o CUDA o his example.
When using OpenUH as he compile o he OpenACC code (Fig. 5), all h ee pa-
ame e s a ec he pe o mance o he gene a ed code, which means ha he pa ame e s
a e ac ually being used o map he compu a ion o he a chi ec u e esou ces. The e is an
excep ion when any o he h ee pa ame e s is se o one. In his case, OpenUH seems o
assume di ec con ol and choose wha i conside s an adequa e se o pa ame e s o gangs
numbe , wo ke s numbe and ec o leng h.
Finally, he esul s ob ained by using accULL (Fig. 6) as ou OpenACC compile show
ha none o he pa ame e s ha e any e ec on he pe o mance o he gene a ed code, and
i is he compile i sel who decides he alue o be used.
5 Conclusions
Du ing his wo k, we ha e ealized ha he OpenACC s anda d is e y unspeci ic abou
how he di e en compile s should implemen he h ee le els o pa allelism. This allows
c
CMMSE ISBN: 978-84-617-8694-7

Analysis o OpenACC Pe o mance Using Di e en Block Geome ies
y1x512
y2x256
y4x128
y8x64
y16x32
y32x16
y64x8
y128x4
y256x2
y512x1
1 4 16 64 256 1024
Block.y Block.x
Numbe o blocks
Th ead-Block sensibili y in CUDA
Time (ms)
4
4.5
5
5.5
6
6.5
7
7.5
8
Figu e 3: E ec s o he Gang Numbe , Wo ke and Vec o Leng h in he measu ed execu ion
ime using CUDA. Execu ion ime is in milliseconds (lowe is be e ). Da ke is lowe .
B igh e is highe .
w1 512
w2 256
w4 128
w8 64
w16 32
w32 16
w64 8
w128 4
w256 2
w512 1
1 4 16 64 256 1024
Wo ke s, Vec o Leng h
Gang Numbe
Th ead-Block sensibili y in PGI
Time (ms)
4
4.5
5
5.5
6
6.5
7
7.5
8
Figu e 4: E ec s o he Gang Numbe , Wo ke and Vec o Leng h in he measu ed execu ion
ime using PGI compile . Execu ion ime is in milliseconds (lowe is be e ). Da ke is lowe .
B igh e is highe .
c
CMMSE ISBN: 978-84-617-8694-7
Daniel Ba ba, A u o Gonzalez-Esc ibano, Diego R. Llanos
w1 512
w2 256
w4 128
w8 64
w16 32
w32 16
w64 8
w128 4
w256 2
w512 1
1 4 16 64 256 1024
Wo ke s, Vec o Leng h
Gang Numbe
Th ead-Block sensibili y in OpenUH
Time (ms)
4
4.5
5
5.5
6
6.5
7
7.5
8
Figu e 5: E ec s o he Gang Numbe , Wo ke and Vec o Leng h in he measu ed execu ion
ime using OpenUH compile . Execu ion ime is in milliseconds (lowe is be e ). Da ke
is lowe . B igh e is highe .
w1 512
w2 256
w4 128
w8 64
w16 32
w32 16
w64 8
w128 4
w256 2
w512 1
1 4 16 64 256 1024
Wo ke s, Vec o Leng h
Gang Numbe
Th ead-Block sensibili y in accULL
Time (ms)
4
4.5
5
5.5
6
6.5
7
7.5
8
Figu e 6: E ec s o he Gang Numbe , Wo ke and Vec o Leng h in he measu ed execu ion
ime using accULL compile . Execu ion ime is in milliseconds (lowe is be e ). Da ke is
lowe . B igh e is highe .
c
CMMSE ISBN: 978-84-617-8694-7
Analysis o OpenACC Pe o mance Using Di e en Block Geome ies
e y di e en beha iou s while uni ying he basic concep s o au oma ic pa alleliza ion o
bo h GPUs and Xeon Phi cop ocesso s.
Due o he s anda d lea ing eedom o he implemen a ion o he di e en le els
o pa allelism o he di e en compile s, he ma u i y o he la e di ec ly a ec s he
pe o mance o he OpenACC-gene a ed code. Thus we ind ha he PGI compile , being
he mo e ma u e o he analyzed compile s, gene a es code ha esembles an op imized
CUDA implemen a ion. OpenUH shows an implemen a ion ha akes he de ined le els
o pa allelism in o conside a ion, bu wi h oom o a pe o mance imp o emen . On he
o he hand, accULL seems o a oid hand-made changes o he le els o pa allelism, choosing
always he same con igu a ion.
Al hough compile implemen a ions a e no e y ma u e ye , he simplici y o ou mi-
c obenchma k allows us o see he e ec s o he a ia ions in he di e en clauses: Gang
numbe , Wo ke numbe , Vec o leng h. Ou esul s ema k ha he pe o mance boos
ob ained by he es ed OpenACC compile s in GPUs is dependan on he implemen a ion
o hese clauses. Howe e , since he mission o hese clauses is o uni y di e en concep s
among GPUs and Xeon Phi accele a o s, we a gue his is complex ask o he compile
implemen a ions.
Acknowledgemen s
This esea ch has been pa ially suppo ed by MICINN (Spain) and ERDF p og am o
he Eu opean Union: HomP og-He Sys p ojec (TIN2014-58876-P), CAPAP-H5 ne wo k
(TIN2014-53522-REDT), and COST P og am Ac ion IC1305: Ne wo k o Sus ainable Ul-
ascale Compu ing (NESUS).
Re e ences
[1] OpenACC-s anda d.o g, “Abou OpenACC.”
[2] OpenACC-S anda d.o g, “The OpenACC applica ion p og amming in e ace e sion
2.5,” oc 2015.
[3] X. Tian, R. Xu, Y. Yan, Z. Yun, S. Chand aseka an, and B. Chapman, “Compiling a
high-le el di ec i e-based p og amming model o GPGPUs,” in Languages and Com-
pile s o Pa allel Compu ing, pp. 105–120, Sp inge , 2014.
[4] R. Reyes, I. L´opez-Rod ´ıguez, J. J. Fume o, and F. de Sande, “accULL: an OpenACC
implemen a ion wi h CUDA and OpenCL suppo ,” in Eu o-Pa 2012 Pa allel P o-
cessing, pp. 871–882, Sp inge , 2012.
c
CMMSE ISBN: 978-84-617-8694-7
Daniel Ba ba, A u o Gonzalez-Esc ibano, Diego R. Llanos
[5] H. O ega-A anz, Y. To es, A. Gonzalez-Esc ibano, and D. R. Llanos, “Op imizing
an apsp implemen a ion o n idia gpus using ke nel cha ac e iza ion c i e ia,” The
Jou nal o Supe compu ing, ol. 70, no. 2, pp. 786–798, 2014.
[6] C. Wang, R. Xu, S. Chand aseka an, B. Chapman, and O. He nandez, “A alida ion
es sui e o OpenACC 1.0,” in Pa allel Dis ibu ed P ocessing Symposium Wo kshops
(IPDPSW), 2014 IEEE In e na ional, pp. 1407–1416, May 2014.
[7] PGI, “Pgi accele a o compile s wi h OpenACC di ec i es.” h ps://www.pg oup.
com/ esou ces/accel.h m, no 2015.
[8] U. de La Laguna, “YaCF.” h ps://bi bucke .o g/ uyman/llcomp, no 2015.
[9] U. de La Laguna, “F angollo.” h ps://bi bucke .o g/ uyman/ angollo, no
2015.
[10] U. o Hous on, “Open-sou ce UH compile .” h p://web.cs.uh.edu/~openuh/
download/, no 2015.
[11] U. de La Laguna, “accULL.” h p://cap.pcg.ull.es/es/accULL, no 2015.
c
CMMSE ISBN: 978-84-617-8694-7