scieee Open visual document viewer

Analysis of OpenACC Performance using Different Block Geometries

Barba Gutiérrez, Daniel,González Escribano, Arturo,Llanos Ferraris, Diego Rafael

Abstract

Producción Científica

Full text

P oceedings o he 17 h In e na ional Con e ence on Compu a ional and Ma hema ical Me hods in Science and Enginee ing, CMMSE 2017 4–8 July, 2017. Analysis o OpenACC Pe o mance Using Di e en Block Geome ies Daniel Ba ba1, A u o Gonzalez-Esc ibano1and Diego R. Llanos1 1Depa amen o de In o ma ica, Uni e sidad de Valladolid, Spain emails: [email p o ec ed],[email p o ec ed],[email p o ec ed] Abs ac OpenACC is a pa allel p og amming model o au oma ic pa alleliza ion o sequen- ial code using compile di ec i es o p agmas. OpenACC is in ended o be used wi h accele a o s such as GPUs and Xeon Phi. The di e en implemen a ions o he s an- da d, al hough s ill in ea ly de elopmen , a e p ima ily ocused on GPU execu ion. In his s udy, we analyze how he di e en OpenACC compile s a ailable unde ce ain p emises beha e when he clauses a ec ing he unde lying block geome y implemen- a ion a e modi ied. These clauses a e he Gang numbe , Wo ke numbe , and Vec o Size de ined by he s anda d. Key wo ds: OpenACC, GPU, block geome y, h ead geome y 1 In oduc ion OpenACC is an open s anda d in ended o au oma ically pa allelize sequen ial code and manage i s execu ion in accele a o s like GPUs o Xeon Phi cop ocesso s. I de ines a numbe o compile di ec i es, also called p agmas. The main goal o OpenACC is o educe bo h lea ning and coding ime in a po able way [1]. The e sion o he OpenACC s anda d a he ime o w i ing is he 2.5 [2]. The OpenACC s anda d was ounded by N idia, CRAY, CAPS and PGI. The numbe o membe s now is la ge , including bo h academic ins i u ions and companies like he Oak Ridge Na ional Labo a o y, he Uni e si y o Hous on, AMD, and he Edinbu gh Pa allel Compu ing Cen e (EPCC), among o he s. The e a e se e al compile s ha implemen he OpenACC s anda d. The PGI com- pile , de eloped by he Po land G oup (subsidia y o N idia) is being dis ibu ed as pa o c CMMSE ISBN: 978-84-617-8694-7 Analysis o OpenACC Pe o mance Using Di e en Block Geome ies he N idia OpenACC Toolki unde a ee 90-day license. C ay Inc. has i s own OpenACC compile , only a ailable o use wi h hei supe compu e s. Pa hscale Inc., a so wa e de- elope o compile s and mul ico e so wa e, also has an OpenACC implemen a ion, he ENZO compile . Among he many academic o open-sou ce al e na i es he e a e he OpenUH compile [3] by he Uni e si y o Hous on and accULL [4] om Uni e sidad de La Laguna (Spain). This wo k p esen s a s udy on he impac o di e en alues o he clauses ha a ec he unde lying block geome y o OpenACC-gene a ed code. GPUs a e e y sensi i e o he geome y o he h ead-block chosen [5], and OpenACC makes use o he e ms “gang”, “wo ke ” and “ ec o ” in o de o de ine di e en le els o pa allelism. Acco ding o [6], he speci ica ion is ambiguous and his unc ionali y depends di ec ly on how each compile is implemen ed. In his wo k, we measu e he impac o he choice o an app op ia e h ead- block geome y when unning a ep esen a i e benchma k. By de aul , he geome y is decided by he compile unless he speci ic clause is used inside he OpenACC di ec i e. Ou aim is o compa e he esul ing beha iou among di e en compile s and op ions. Fo h ead-block geome y es ing, we will modi y he mos ep esen a i e o he benchma ks, es ing se e al combina ions o alues o he clauses o speci y gang, wo ke s and ec o s, and analyzing he di e ences in execu ion ime o each compile . This will o e some insigh on he implemen a ion o hese clauses on each compile . Ou con ibu ion shows ha he decisions made by each compile is no always op imal, bu manual uning o he di e en alues is no always possible o e e y compile . The es o his pape is o ganized as ollows. Sec ion 2 desc ibes he selec ed compile s. Sec ion 3 shows ou selec ed mic obenchma k o es ing he beha iou o he gene a ed code when modi ying he block geome y. Sec ion 4 con ains he esul o ou analysis abou he impac on pe o mance when changing he unde lying block geome y. Finally, Sec ion 5 concludes ou pape . 2 A ailable Compile s We men ioned se e al compile s in he In oduc ion. In his sec ion we desc ibe wi h mo e de ail he compile s we we e able o use o his s udy. 2.1 PGI Compile The PGI Compile [7] is being de eloped by The Po land G oup, being owned by N idia. This compile is equen ly p esen ed in webina s, wo kshops, and con e ences. A he ime o w i ing his pape , he PGI compile is a ailable o download as pa o he OpenACC Toolki om N idia. This oolki includes a 90-day ee ial, he possibili y o acqui ing an academic license o a whole yea , o buying a comme cial license. c CMMSE ISBN: 978-84-617-8694-7 Daniel Ba ba, A u o Gonzalez-Esc ibano, Diego R. Llanos 2.2 accULL The accULL [4] compile de eloped by Uni e sidad o La Laguna (Spain) is an open-sou ce ini ia i e. accULL consis s on a s uc u e o wo laye s con aining YaCF [8] (Ye ano he Compile F amewo k) and F angollo [9], a un ime lib a y. YaCF ac s as a sou ce- o-sou ce ansla o i, while F angollo wo ks as an in e ace ha p o ides he mos common ope a ions ound in accele a o s. 2.3 OpenUH The OpenUH [3] compile , de eloped by he Uni e si y o Hous on (USA) is ano he open- sou ce ini ia i e. I makes use o Open64, a discon inued open-sou ce op imizing compile . 3 Mic obenchma k Desc ip ion In OpenACC, block size is de ined by gangs, wo ke s and ec o s.Thei choices a ec he pe o mance on memo y-bound applica ions. To s udy his issue, we a e going o use a e y simple ma ix addi ion implemen ed bo h in CUDA and OpenACC. Ou decision is made by he ac ha he p oblem is emba assingly pa allel, memo y acceses a e pe ec ly coalesced, and he compu a ional load pe global memo y access is low (memo y-bound applica ion). The s anda d only de ines he di e en le els o pa allelism, bu i is up o each imple- men a ion o decide how a e hese le els exploi ed in he ac ual a chi ec u e. In o de o ob ain compa able esul s, some de ails should be aken in o accoun . The CUDA e sion needs o be implemen ed using elas ic ke nels, using a ixed numbe o blocks, which equals he gang numbe in he OpenACC code. Also, since he OpenACC s anda d es ablishes ha g id dimensions depend on he use o he collapse clause wi h nes ed loops, we ha e decided o make a one-dimensional g id. We e alua e a numbe o blocks in he g id anging om one o 2048 (using only powe s o wo). The sequen ial code can be seen in Fig. 1 and he CUDA ke nel in Fig. 2. In OpenACC he X dimension o he CUDA block ansla es o ec o leng h, whe eas he Y dimension equals he wo ke numbe . We ha e also decided o use 512 h eads pe block, ying each possible combina ion o X and Y dimension alues using powe s o wo. 4 E alua ion In his sec ion we analyze he impac o di e en choices o he geome y o he unde lying h ead-blocks in he OpenACC gene a ed code. c CMMSE ISBN: 978-84-617-8694-7 Analysis o OpenACC Pe o mance Using Di e en Block Geome ies #p agma acc da a copyin(p_A[0:Size*Size],p_B[0:Size*Size]), copyou (p_C[0:Size*Size]) { in j; #p agma acc ke nels #p agma acc loop independen gang(GANG), wo ke (WORKER), ec o (VECTOR) o (j = 0; j < Size*Size; ++j) { p_C[j] = ALPHA*p_A[j] + p_B[j]; } } //End da a egion Figu e 1: Sequen ial Code o he Ma ix Addi ion __global__ oid ma ixKe nel( loa * p_A, loa * p_B, loa * p_C) { in i e a ions = ((SIZE/blockDim.x)*(SIZE/blockDim.y)) / GANG; in i e ; o (i e = 0; i e < i e a ions; ++i e ) { in x = (blockIdx.x + i e *GANG)*blockDim.x + h eadIdx.x; in y = h eadIdx.y; in i = ( x/SIZE)*blockDim.y + y; in j = ( x%SIZE); in o se = i*SIZE + j; i (o se < SIZE*SIZE) p_C[o se ] = ALPHA*p_A[o se ] + p_B[o se ]; } } Figu e 2: CUDA Ke nel o he Ma ix Addi ion c CMMSE ISBN: 978-84-617-8694-7 Daniel Ba ba, A u o Gonzalez-Esc ibano, Diego R. Llanos 4.1 Expe imen al Se up We used a N idia GTX Ti an Black o un he expe imen s. This GPU con ains 2 880 CUDA co es wi h a clock a e o 980MHz and 15 SMs. I has 6GB o RAM, and Compu e Capabili y 3.5. The hos is a Xeon E5-2690 3 wi h 12 co es a a clock a e o 1.9GHz, and 64GB in ou 12GB modules. The PGI compile is he one con ained in he N idia OpenACC Toolki , e sion 15.7- 0, published in Jul 13, 2015. We used OpenUH e sion 3.1.0 (published in No embe 4, 2015), based on Open64 e sion 5.0 and using GCC 4.2.0, p ebuil , downloaded om he High Pe o mance Compu ing Tools g oup websi e [10]. accULL is e sion 0.4alpha (published in No embe 28, 2013), downloaded om Uni e sidad de La Laguna’s esea ch g oup “Compu aci´on de Al as P es aciones” [11]. 4.2 Block Geome y Sensibili y o Gene a ed Code Obse ing he esul s ob ained using he CUDA code in Fig. 3, we can see ha he bes esul s a e ob ained when he block numbe is high enough o make he GPU each p ope le els o occupa ion. The X dimension plays a huge ole, bu we expec ed o see an im- p o emen in pe o mance when he X dimension was a leas 16. We suspec his is due o se e al ac o s, being he mos impo an he beha iou o he cache when elas ic ke nels a e used. The esul s o he OpenACC code gene a ed by PGI (Fig. 4) show ha he only ac o a ec ing he pe o mance is he gangs numbe , whe eas he a ia ion o wo ke s numbe and ec o leng h does no play a signi ican ole excep when a ec o leng h o one is used. In his case, pe o mance is se e ely a ec ed, which indica es a poo use o he cache. O e all, PGI’s beha iou is he closes one o CUDA o his example. When using OpenUH as he compile o he OpenACC code (Fig. 5), all h ee pa- ame e s a ec he pe o mance o he gene a ed code, which means ha he pa ame e s a e ac ually being used o map he compu a ion o he a chi ec u e esou ces. The e is an excep ion when any o he h ee pa ame e s is se o one. In his case, OpenUH seems o assume di ec con ol and choose wha i conside s an adequa e se o pa ame e s o gangs numbe , wo ke s numbe and ec o leng h. Finally, he esul s ob ained by using accULL (Fig. 6) as ou OpenACC compile show ha none o he pa ame e s ha e any e ec on he pe o mance o he gene a ed code, and i is he compile i sel who decides he alue o be used. 5 Conclusions Du ing his wo k, we ha e ealized ha he OpenACC s anda d is e y unspeci ic abou how he di e en compile s should implemen he h ee le els o pa allelism. This allows c CMMSE ISBN: 978-84-617-8694-7 Analysis o OpenACC Pe o mance Using Di e en Block Geome ies y1x512 y2x256 y4x128 y8x64 y16x32 y32x16 y64x8 y128x4 y256x2 y512x1 1 4 16 64 256 1024 Block.y Block.x Numbe o blocks Th ead-Block sensibili y in CUDA Time (ms) 4 4.5 5 5.5 6 6.5 7 7.5 8 Figu e 3: E ec s o he Gang Numbe , Wo ke and Vec o Leng h in he measu ed execu ion ime using CUDA. Execu ion ime is in milliseconds (lowe is be e ). Da ke is lowe . B igh e is highe . w1 512 w2 256 w4 128 w8 64 w16 32 w32 16 w64 8 w128 4 w256 2 w512 1 1 4 16 64 256 1024 Wo ke s, Vec o Leng h Gang Numbe Th ead-Block sensibili y in PGI Time (ms) 4 4.5 5 5.5 6 6.5 7 7.5 8 Figu e 4: E ec s o he Gang Numbe , Wo ke and Vec o Leng h in he measu ed execu ion ime using PGI compile . Execu ion ime is in milliseconds (lowe is be e ). Da ke is lowe . B igh e is highe . c CMMSE ISBN: 978-84-617-8694-7 Daniel Ba ba, A u o Gonzalez-Esc ibano, Diego R. Llanos w1 512 w2 256 w4 128 w8 64 w16 32 w32 16 w64 8 w128 4 w256 2 w512 1 1 4 16 64 256 1024 Wo ke s, Vec o Leng h Gang Numbe Th ead-Block sensibili y in OpenUH Time (ms) 4 4.5 5 5.5 6 6.5 7 7.5 8 Figu e 5: E ec s o he Gang Numbe , Wo ke and Vec o Leng h in he measu ed execu ion ime using OpenUH compile . Execu ion ime is in milliseconds (lowe is be e ). Da ke is lowe . B igh e is highe . w1 512 w2 256 w4 128 w8 64 w16 32 w32 16 w64 8 w128 4 w256 2 w512 1 1 4 16 64 256 1024 Wo ke s, Vec o Leng h Gang Numbe Th ead-Block sensibili y in accULL Time (ms) 4 4.5 5 5.5 6 6.5 7 7.5 8 Figu e 6: E ec s o he Gang Numbe , Wo ke and Vec o Leng h in he measu ed execu ion ime using accULL compile . Execu ion ime is in milliseconds (lowe is be e ). Da ke is lowe . B igh e is highe . c CMMSE ISBN: 978-84-617-8694-7 Analysis o OpenACC Pe o mance Using Di e en Block Geome ies e y di e en beha iou s while uni ying he basic concep s o au oma ic pa alleliza ion o bo h GPUs and Xeon Phi cop ocesso s. Due o he s anda d lea ing eedom o he implemen a ion o he di e en le els o pa allelism o he di e en compile s, he ma u i y o he la e di ec ly a ec s he pe o mance o he OpenACC-gene a ed code. Thus we ind ha he PGI compile , being he mo e ma u e o he analyzed compile s, gene a es code ha esembles an op imized CUDA implemen a ion. OpenUH shows an implemen a ion ha akes he de ined le els o pa allelism in o conside a ion, bu wi h oom o a pe o mance imp o emen . On he o he hand, accULL seems o a oid hand-made changes o he le els o pa allelism, choosing always he same con igu a ion. Al hough compile implemen a ions a e no e y ma u e ye , he simplici y o ou mi- c obenchma k allows us o see he e ec s o he a ia ions in he di e en clauses: Gang numbe , Wo ke numbe , Vec o leng h. Ou esul s ema k ha he pe o mance boos ob ained by he es ed OpenACC compile s in GPUs is dependan on he implemen a ion o hese clauses. Howe e , since he mission o hese clauses is o uni y di e en concep s among GPUs and Xeon Phi accele a o s, we a gue his is complex ask o he compile implemen a ions. Acknowledgemen s This esea ch has been pa ially suppo ed by MICINN (Spain) and ERDF p og am o he Eu opean Union: HomP og-He Sys p ojec (TIN2014-58876-P), CAPAP-H5 ne wo k (TIN2014-53522-REDT), and COST P og am Ac ion IC1305: Ne wo k o Sus ainable Ul- ascale Compu ing (NESUS). Re e ences [1] OpenACC-s anda d.o g, “Abou OpenACC.” [2] OpenACC-S anda d.o g, “The OpenACC applica ion p og amming in e ace e sion 2.5,” oc 2015. [3] X. Tian, R. Xu, Y. Yan, Z. Yun, S. Chand aseka an, and B. Chapman, “Compiling a high-le el di ec i e-based p og amming model o GPGPUs,” in Languages and Com- pile s o Pa allel Compu ing, pp. 105–120, Sp inge , 2014. [4] R. Reyes, I. L´opez-Rod ´ıguez, J. J. Fume o, and F. de Sande, “accULL: an OpenACC implemen a ion wi h CUDA and OpenCL suppo ,” in Eu o-Pa 2012 Pa allel P o- cessing, pp. 871–882, Sp inge , 2012. c CMMSE ISBN: 978-84-617-8694-7 Daniel Ba ba, A u o Gonzalez-Esc ibano, Diego R. Llanos [5] H. O ega-A anz, Y. To es, A. Gonzalez-Esc ibano, and D. R. Llanos, “Op imizing an apsp implemen a ion o n idia gpus using ke nel cha ac e iza ion c i e ia,” The Jou nal o Supe compu ing, ol. 70, no. 2, pp. 786–798, 2014. [6] C. Wang, R. Xu, S. Chand aseka an, B. Chapman, and O. He nandez, “A alida ion es sui e o OpenACC 1.0,” in Pa allel Dis ibu ed P ocessing Symposium Wo kshops (IPDPSW), 2014 IEEE In e na ional, pp. 1407–1416, May 2014. [7] PGI, “Pgi accele a o compile s wi h OpenACC di ec i es.” h ps://www.pg oup. com/ esou ces/accel.h m, no 2015. [8] U. de La Laguna, “YaCF.” h ps://bi bucke .o g/ uyman/llcomp, no 2015. [9] U. de La Laguna, “F angollo.” h ps://bi bucke .o g/ uyman/ angollo, no 2015. [10] U. o Hous on, “Open-sou ce UH compile .” h p://web.cs.uh.edu/~openuh/ download/, no 2015. [11] U. de La Laguna, “accULL.” h p://cap.pcg.ull.es/es/accULL, no 2015. c CMMSE ISBN: 978-84-617-8694-7