scieee Open visual document viewer

Comparative Analysis of OpenACC Compilers

Barba Gutiérrez, Daniel,González Escribano, Arturo,Llanos Ferraris, Diego Rafael

Abstract

Producción Científica

Full text

Compa a i e Analysis o OpenACC Compile s Daniel Ba ba, A u o Gonzalez-Esc ibano, and Diego R. Llanos ? Uni e sidad de Valladolid, Depa amen o de In o ma ica, Valladolid, Spain {daniel|a u o|diego}@in o .u a.es Abs ac . OpenACC has been on de elopmen o a ew yea s now. The OpenACC 2.5 speci ica ion was ecen ly made public and he e a e some ini ia i es o de eloping ull implemen a ions o he s anda d o make use o accele a o capabili ies. The e is much o be done ye , bu cu - en ly, OpenACC o GPUs is eaching a good ma u i y le el in a ious implemen a ions o he s anda d, using CUDA and OpenCL as backends. N idia is in es ing in his p ojec and hey ha e eleased an OpenACC Toolki , including he PGI Compile . The e a e, howe e , mo e de elop- men s ou he e. In his wo k, we analyze di e en a ailable OpenACC compile s ha ha e been de eloped by companies o uni e si ies du - ing he las yea s. We check hei pe o mance and ma u i y, keeping in mind ha OpenACC is designed o be used wi hou ex ensi e knowledge abou pa allel p og amming. Ou esul s show ha he compile s a e on hei way o a easonable ma u i y, p esen ing di e en s eng hs and weaknesses. 1 In oduc ion OpenACC is an open s anda d which de ines a collec ion o compile di ec i es o p agmas o execu ion o code blocks on accele a o s like GPUs o Xeon Phi cop ocesso s. OpenACC aims o educe bo h he equi ed lea ning ime and he pa alleliza ion o sequen ial code in a po able way [1]. OpenACC speci ica ion is cu en ly on i s 2.5 e sion [2], which has been eleased ecen ly. OpenACC was ounded by N idia, CRAY, CAPS and PGI, bu now he e is a la ge lis o conso ium membe s, bo h om he indus y and academy, including he Oak Ridge Na ional Labo a o y, he Uni e si y o Hous on, AMD, and he Edinbu gh Pa allel Compu ing Cen e (EPCC), among o he s. The Co po a e O ice s a e, a he ime o w i ing his pape , om N idia, Oak Ridge Na ional Labo a o y, CRAY and AMD. Academic membe ships a e a ailable o in e es ed ins i u ions. The e a e se e al compile s suppo ing OpenACC. The PGI Compile ( om he Po land G oup, which is a subsidia y o N idia o some ime now) is be- ing dis ibu ed as pa o he N idia OpenACC Toolki , unde a ee 90-day ?This esea ch has been pa ially suppo ed by MICINN (Spain) and ERDF p og am o he Eu opean Union: HomP og-He Sys p ojec (TIN2014-58876-P), CAPAP-H5 ne wo k (TIN2014-53522-REDT), and COST P og am Ac ion IC1305: Ne wo k o Sus ainable Ul ascale Compu ing (NESUS). Thanks also o D . Hec o O ega and D . Ja ie F esno o hei hough ul commen s, help, and encou agemen . 2 ial license, a ee annual academic license, o a comme cial license. The PGI Compile uses CUDA o OpenCL as backend. CAPS en e p ise, a p o ide o so wa e and se ices o he High Pe o mance Compu ing communi y, also de- eloped a compile which suppo s CUDA and OpenCL. Howe e , he company is no longe in business and i s de elopmen is no a ailable anymo e. CRAY Inc. has i s own OpenACC compile , a ailable wi h hei compu e s. I is e- po ed o be one o he mos ma u e comme cial compile s. Howe e , hey ha e no ye o e ed an academic o ial license o he pu poses o s udies like his one. Pa hscale Inc., a compile and mul ico e so wa e de elope , also made an OpenACC implemen a ion on hei ENZO compile . Un o una ely, a e p i a e communica ions hey we e eluc an o allow us o use hei compile o his s udy. The e a e also se e al academic a emp s o de eloping an OpenACC com- pile . In pa icula , OpenUH [3] de eloped a he Uni e si y o Hous on, and accULL [4] om Uni e sidad de La Laguna (Spain). Bo h o hem a e a ailable o ee o anyone in e es ed. This wo k p esen s a s udy abou he le el o suppo o OpenACC in he a ailable compile s, examining hei s eng hs and weaknesses, and gi ing in- sigh s on hei pe o mance. We analyze he le el o ma u i y o each compile , in e ms o hei com- ple eness in he suppo o he s anda d and obus ness, using each compile ’s documen a ion o check wha pa s o he speci ica ion ha e been implemen ed. We check he suppo o OpenACC compile di ec i es wi h he help o a bench- ma k sui e, de eloped by he Edinbu gh Pa allel Compu ing Cen e, which will be desc ibed la e . Ano he aspec o es is he ela i e pe o mance o he gene a ed code. Fo his, we would wan o un mo e complex applica ions and measu e he di e ences be ween he pe o mance o he execu able code gene a ed by each compile . The ideal si ua ion would be o es applica ions as close o eal wo ld p oblems as possible, a oiding syn he ic code agmen s. Since he use o OpenACC is no common ye , we ha e o ely on exis ing benchma ks. A his poin , compa ing he esul s om OpenACC code wi h CUDA o OpenCL di ec implemen a ions migh seem app op ia e, bu po ing he di e en bench- ma ks o hese languages makes he esul dependen on human in e e ence, as he de elope ’s abili y o CUDA p og amming impac s on he pe o mance. The expe imen s conduc ed in his wo k we e ca ied ou using se e al bench- ma ks. Fi s , he EPCC benchma k sui e which con ains a g oup o 13 ke nels po ed o OpenACC, called “Le el 1” benchma ks. This benchma k sui e also con ains h ee eal applica ions called “Himeno”, “27s encil”, and “le co e” [5]. We ha e also used he Pa hscale po o he Rodinia benchma k [6]. We also wan ed o use he OpenACC Valida ion Tes sui e [7] de eloped by he Uni e - si y o Hous on, bu a his momen ha ool is only a ailable o OpenACC membe s. 3 Ou conclusion is ha he di e en compile s a e on hei way o a easonable ma u i y. Howe e , he e is a numbe o ea u es no ully implemen ed ye by some o he compile s. The es o his pape is o ganized as ollows. Sec ion 2 desc ibes he selec ed compile s. Sec ion 3 shows he cha ac e is ics o he benchma k sui es chosen, enume a ing some p oblems encoun e ed when compiling hem wi h he compil- e s selec ed. Sec ion 4 con ains he esul o ou analysis, in e ms o comple eness o he OpenACC ea u es suppo ed, obus ness o compile implemen a ions, and ela i e pe o mance o he gene a ed code. Finally, Sec ion 5 concludes ou pape . 2 A ailable Compile s In he in oduc ion we men ioned se e al compile s. In his sec ion we desc ibe wi h mo e de ail he compile s we we e able o ob ain and use o his s udy, and we will discuss hei ins alla ion pa icula i ies on ou Linux based pla o m. 2.1 PGI Compile The PGI Compile [8] is being de eloped by The Po land G oup, being owned by N idia. This compile is widely used in webina s, wo kshops, and con e ences. The PGI Compile is, a he ime o w i ing his pape , a ailable o download as pa o he OpenACC Toolki om N idia. This oolki includes a 90-day ee ial, he possibili y o acqui ing an academic license o a whole yea , o buying a comme cial license. 2.2 accULL The accULL [4] compile de eloped by Uni e sidad o La Laguna (Spain) is an open sou ce ini ia i e. accULL consis s on a s uc u e o wo laye s con aining YaCF [9] (Ye ano he Compile F amewo k) and F angollo [10], a un ime li- b a y. YaCF ac s as a sou ce- o-sou ce ansla o while F angollo wo ks as an in e ace p o iding he mos common ope a ions ound in accele a o s. 2.3 OpenUH The OpenUH [3] compile , de eloped by he Uni e si y o Hous on (USA) is ano he open sou ce ini ia i e. I makes use o Open64, a discon inued open- sou ce op imizing compile . 3 Benchma k Desc ip ion This sec ion desc ibes he di e en benchma ks used in ou wo k, enume a ing he main cha ac e is ics ha make hem in e es ing o his s udy, and any issue de ec ed du ing hei compila ion wi h he h ee compile s s udied. 4 3.1 EPCC OpenACC Benchma ks This benchma k sui e [11] has been de eloped by he Edinbu gh Pa allel Com- pu ing Cen e (EPCC). The benchma ks a e di ided in h ee ca ego ies: “Le el 0”, “Le el 1” and “Applica ions”. The compiled p og am launches all he bench- ma ks in he sui e sequen ially. By de aul , he numbe o epe i ions is en, and he esul is he a e age o each benchma k. Time is measu ed in mic oseconds in double p ecision, using he OpenMP unc ion omp ge w ime(). We desc ibe b ie ly he benchma ks included bellow: Le el 0 Le el 0 includes a collec ion o small benchma ks ha execu e single hos and accele a o ope a ions, such as memo y ans e s. Le el 1 Le el 1 benchma ks [12] consis on a se ies o BLAS- ype ke nels. They a e based on Polybench [13] and Polybench/GPU ke nels. They measu e he pe o mance o execu ing hose codes. These benchma ks a e un on he CPU i s in o de o ha e esul s o compa e hose ob ained on he GPU. A b ie desc ip ion o he di e en issues ound while unning his sui e o hese compile s ollows. OpenUH: The ollowing benchma ks canno be compiled due o unsuppo ed p agmas p esen in hei code: ke nels i : P oblems using #p agma ke nels i (0). pa allel p i a e: P oblems decla ing pa ams as p i a e. pa allel i s p i a e: P oblems decla ing pa ams as i s p i a e. le co e: P oblems wi h non scala poin e s. himeno: P oblems wi h non scala poin e s. accULL: The e is a p oblem ela ed o a unc ion poin e in he hos p og am. The compile , du ing he sou ce o sou ce ansla ion modi ies he syn ax o he unc ion poin e . A double (* es )( oid) is con e ed o a double *( es ( oid)). This was sol ed by manually co ec ing his change in he in e media e C code gene a ed, e-compiling he objec ile and copying i o he main di ec o y o link all he objec iles again. No wa ning om any o he p agmas was de ec ed so we we e able o un all he benchma ks. 3.2 Rodinia OpenACC Rodinia is a benchma k sui e o he e ogeneous compu ing [14,15]. I includes applica ions and ke nels o mul ico e CPU o GPU applica ions. The e is an e o o po exis ing Rodinia benchma ks o OpenACC. Pa h- scale [6] is wo king on his. We ha e es ed hei Rodinia e sion commi ed o Gi Hub on Ap il 25, 2014. Mos o he sui e wo ks wi h PGI, bu OpenUH and accULL ha e many p oblems o compile mos o he es s. We ha e been able o success ully compile he ollowing benchma ks con ained in he sui e wi h wo o mo e compile s: gaussian, nw, lud, c d, ho spo , pa h inde , and s ad2. 5 4 E alua ion In his sec ion we analyze he OpenACC compile s, using bo h documen a ion and expe imen a ion. We use each compile ’s documen a ion o check he com- ple eness o OpenACC ea u es suppo ed. Then we use he EPCC benchma ks o check bo h obus ness and ela i e pe o mance. Finally we check h ead-block size sensibili y, measu ing he impac on pe o mance o di e en geome ies. 4.1 Expe imen al Se up We used a N idia GTX Ti an Black o un he expe imen s. This GPU con ains 2880 CUDA co es wi h a clock a e o 980Mhz and 15 SMs. I has 6GB o RAM, and Compu e Capabili y 3.5. The hos is a Xeon E5-2690 3 wi h 12 co es a a clock a e o 1.9GHz, and 64GB in ou 12GB modules. The PGI compile is he one con ained in he N idia OpenACC Toolki , e - sion 15.7-0, published in Jul 13, 2015. We used OpenUH e sion 3.1.0 (published in No embe 4, 2015), based on Open64 e sion 5.0 and using GCC 4.2.0, p e- buil , downloaded om he High Pe o mance Compu ing Tools g oup websi e [16]. accULL is e sion 0.4alpha (published in No embe 28, 2013), downloaded om Uni e sidad de La Laguna’s esea ch g oup “Compu aci´on de Al as P es a- ciones” [17]. 4.2 Comple eness o OpenACC Fea u es Suppo ed F om each compile documen a ion we ge some insigh on he comple eness o he OpenACC ea u es suppo ed. F om his in o ma ion, we can conclude ha he OpenACC s anda d is no ully implemen ed ye by any o he a ailable compile s. The e is wo k o be done, bu he h ee compile s a e a a espec able ma u i y le el. 4.3 Robus ness and P agma Implemen a ion The EPCC Benchma k sui e con ains se e al benchma ks o es ing OpenACC di ec i es. These benchma ks a e con ained in he “Le el 0” g oup, which has been desc ibed in he p e ious sec ion. Table 1 con ains he esul s ob ained o he h ee compile s. In his sec ion we enume a e he p oblems wi h each benchma k and we explain he esul s ob ained, including he o e head o he di e en p agma implemen a ions. Excep o Upda e hos ,Ke nels In oc., and Pa allel In oc., he ime shown is he di e ence be ween execu ing and no execu ing each p agma. When he o e head is ze o (o he p agma is no implemen ed), he imes a e e y simila , wi h minimal s ochas ic a ia ion. These a ia ions may p oduce a e y small nega i e esul when calcula ing he di e ence. When di e ences in ime a e on he o de o ens o mic oseconds (posi i e o nega i e), i can be assumed ha he e is no di e ence in ime be ween he di e en e sions es ed in ha benchma k. 6 Table 1: EPCC le el 0: di ec i e’s o e head (in µsec), 1 MB da ase EPCC L0 PGI OpenUH accULL Ke nels i -37.50 Fail 4.54 Pa allel i -30.76 -0.48 1237.02 Pa allel p i a e -21.94 Fail 51.09 Pa allel 1s p i Fail Fail -213.83 Ke nels comb. -1.67 -108.43 -127.17 Pa allel comb. -0.05 -2.74 33.38 Upda e hos 478.63 373.22 548.77 Ke nels In oc. Fail 12.76 2398.20 Pa allel In oc. 31.81 13.47 1377.88 Pa allel educ . -14.85 -164.41 -2168.12 Ke nels educ . -8.49 -172.31 -2009.11 PGI The e was a p oblem wi h he “Ke nels In oca ion” benchma k: I e- u ned an inco ec esul . The code was no being pa allelized and he p agmas we e igno ed because i wasn’ speci ically s a ed ha he i e a ions we e inde- penden . This could be sol ed adding he keywo d es ic o he poin e o he clause independen o he p agma. The “ke nels i ” and “pa allel i ” esul s a e e y simila , and in bo h cases he esul s indica e ha he code wi h he p agma is sligh ly as e han he one wi hou i , e en hough bo h a e being un on he hos . In [5] i was s a ed ha his could be because o op imiza ions done by he compile while o a e p ocessing he p agmas. The “pa allel p i a e” benchma k shows ha he c ea ion o p i a e a iables o each h ead unning he loop is sligh ly as e han he alloca ion o de ice memo y. “Ke nels combined” shows a e y small di e ence o ime be ween w i ing wo p agmas ins ead o a combined one, he o me being sligh ly as e han he la e al hough he di e ence is almos negligible. The same occu s o he “pa allel combined” benchma k, he di e ence being smalle in his case. Finally “Pa allel educ ion” and “Ke nels educ ion” show ha PGI has e y li le o e head o he educ ion clause. In [5] i is s a ed ha he PGI compile does he educ ion e en i i is no anno a ed. This could explain he e y small di e ence in bo h benchma ks. OpenUH We go some e o s du ing compila ion o he “Ke nels i ”, “Pa - allel p i a e” and “Pa allel 1s p i a e” benchma ks so hey a e igno ed in his analysis. Howe e , he “Pa allel i ” di ec i e is suppo ed and he di e ence be- ween using he p agma o un code on he hos o unning i di ec ly is almos negligible. “Ke nels combined” shows an o e head o he combined p agma e sus he sepa a ed e sion. Howe e , his is no he case o he “Pa allel combined” benchma k, whe e he di e ence is much smalle . The in oca ion o ke nels and 7 pa allel di ec i es a e e y simila . And o bo h o hem, he educ ion adds a simila o e head. This migh be ela ed o OpenUH assuming loops o be independen inside ke nels egions. accULL No e o s we e shown while compiling o unning he benchma ks wi h accULL. The e is a big di e ence be ween he wo e sions con ained in he “Ke nels i ” and he “Pa allel i ” benchma ks, whe e he ke nels di ec i e e sion has a e y small o e head compa ed o he non-anno a ed code. This o e head is e y la ge in he pa allel di ec i e e sion. This is explained by he accULL de elope s in [5] whe e hey say ha he absence o a loop clause in he pa allel di ec i e is causing he loop o be execu ed sequen ially in each h ead. The e o e, his clause is no co ec ly suppo ed, as we unde s and om he OpenACC Speci ica ion ha he loop should be execu ed only on he hos . Robus ness Summa y The o e all esul s indica e ha some o he clauses a e no implemen ed ye , bu he h ee compile s a e in hei way o a easonable ma u i y le el and, since he mos used di ec i es a e wo king, hey can ac ually be used o code pa alleliza ion using OpenACC. 4.4 Rela i e Pe o mance o Gene a ed Code In his sec ion we analyze he pe o mance o he gene a ed code desc ibing he impac o p agmas o e head in accULL. Pe o mance measu emen is di- ided in o da a mo emen , whe e we analyze he esul s o he da a mo emen benchma ks in Le el 0 o EPCC OpenACC Benchma k Sui e, and execu ion pe - o mance, using Le el 1 and Applica ion Le el o EPCC OpenACC Benchma k Sui e, and selec ed benchma ks om Rodinia. E ec o P agmas O e head in accULL Some esul s om he Le el 0 o he EPCC Benchma k Sui e show a pe o mance impac in oduced by some clauses and di ec i es in he accULL gene a ed code. Ke nels and Pa allel in oca ions in accULL ha e a highe o e head han o he compile s. This is due o he un ime calls and i is specially no iceable in he educ ion clause. These o e heads accumula ion does no ha e a signi ican impac o complex ke nels, o launching he same ke nel o e and o e again. Howe e , his could be a p oblem when unning simple ke nels o many di e en small ke nels. This is he main eason behind he o e all esul s showing a wo se pe o mance o he accULL compile in his analysis. Da a Mo emen Da a mo emen pe o mance can be measu ed in ou bench- ma ks om he Le el 0 o he EPCC OpenACC benchma k sui e. We ha e launched 10 epe i ions o hose benchma ks wi h da asizes o 1 kB, 1 MB, 10 MB, and 1 GB. The esul s can be seen in Tables 2, 3, 4, and 5. 8 Table 2: EPCC da a mo emen esul s (in µsec), 1 kB da ase . Whi e cells highligh he bes esul s. no m. is he no malized esul using PGI as e e ence. Da a M mn PGI OpenUH accULL 10 eps, 1kB ime no m. ime no m. ime no m. Con igH2D 30.827 1.0 322.699 10.47 338.218 10.97 Con igD2H 14.686 1.0 323.319 22.01 343.919 23.42 SlicedH2D 12.087 1.0 310.897 25.72 315.914 26.13 SlicedD2H 14.948 1.0 324.010 21.67 327.714 21.92 GeoMean 18.93 GeoMean 19.58 Table 3: EPCC da a mo emen esul s (in µsec), 1 MB da ase . Whi e cells highligh he bes esul s. no m. is he no malized esul using PGI as e e ence. Da a M mn PGI OpenUH accULL 10 eps, 1MB ime no m. ime no m. ime no m. Con igH2D 484.347 1.0 950.789 1.96 727.839 1.50 Con igD2H 461.936 1.0 632.691 1.37 792.761 1.72 SlicedH2D 17.094 1.0 267.982 15.68 274.462 16.06 SlicedD2H 36.335 1.0 254.702 7.01 285.685 7.86 GeoMean 4.14 GeoMean 4.24 Table 4: EPCC da a mo emen esul s (in µsec), 10 MB da ase . Whi e cells highligh he bes esul s. no m. is he no malized esul using PGI as e e ence. Da a M mn PGI OpenUH accULL 10 eps, 10MB ime no m. ime no m. ime no m. Con igH2D 4141.402 1.0 6887.984 1.66 3354.666 0.81 Con igD2H 5876.088 1.0 2043.747 0.35 4396.052 0.74 SlicedH2D 27.322 1.0 404.214 14.79 427.203 15.64 SlicedD2H 48.017 1.0 269.818 5.62 280.203 5.84 GeoMean 2.64 GeoMean 2.72 Table 5: EPCC da a mo emen esul s (in µsec), 1 GB da ase . Whi e cells highligh he bes esul s. no m. is he no malized esul using PGI as e e ence. Da a M mn PGI OpenUH accULL 10 eps, 1GB ime no m. ime no m. ime no m. Con igH2D 32310.009 1.0 788945.913 24.42 296340.991 9.17 Con igD2H 55179.119 1.0 553282.976 10.03 347280.359 6.29 SlicedH2D 400.066 1.0 535.011 1.34 533.943 1.34 SlicedD2H 158.071 1.0 2818.100 17.83 4294.407 27.17 GeoMean 8.75 GeoMean 6.76 9 In [5] i was s a ed ha PGI used pinned memo y and ha i was causing issues in smalle da ase s. I seems ha PGI has sol ed his issue since hen and, looking a he documen a ion, i is now possible o speci y he ype o memo y access we wan wi h a compila ion lag. When using la ge da ase s, OpenUH and accULL do no show he expec ed esul s acco ding o he e olu ion shown in ables 2, 3, and 4. We guess ha his is ela ed o he usage o pinned memo y by he PGI Compile , allowing i o ob ain be e esul s when da ase s a e big enough. Execu ion Pe o mance, EPCC Benchma ks In his sec ion we will ana- lyze he pe o mance o he code gene a ed by he PGI, OpenUH, and accULL compile s wi h he benchma ks con ained in he EPCC Le el 1 and Applica ion le el. We use h ee di e en da ase s: 1kB, 1MB, and 10MB. This choice is based on he ac ha bigge da ase s esul in an ou o memo y e o due o how he benchma ks y o alloca e memo y on he de ice. We suspec he memo y alloca ion is being done in each h ead inside he gene a ed ke nels, using mo e memo y han expec ed. In summa y, PGI code ob ains be e esul s in almos e e y benchma k. Howe e , he di e ences sho en when using bigge da ase s. Table 6: EPCC execu ion esul s (in µsec), 1 kB da ase . Whi e cells highligh he bes esul s. Exec. ime PGI OpenUH accULL 10 eps, 1kB ime no m. ime no m. ime no m. 2MM 99.087 1.0 522.304 5.27 2799.229 28.25 3MM 80.204 1.0 380.683 4.75 3799.048 47.37 ATAX 58.103 1.0 327.110 5.63 2564.702 44.14 BICG 72.408 1.0 350.380 4.84 2628.499 36.30 MVT 80.037 1.0 354.743 4.43 2665.299 33.30 SYRK 68.426 1.0 289.512 4.23 2394.803 35.00 COV 87.261 1.0 314.617 3.61 3795.372 43.49 COR 104.976 1.0 337.362 3.21 5208.668 49.62 SYR2K 73.290 1.0 317.574 4.33 2469.765 33.70 GESUMMV 65.613 1.0 312.996 4.77 1500.021 22.86 GEMM 49.710 1.0 323.725 6.51 1237.473 24.89 2DCONV 46.444 1.0 286.174 6.16 1207.528 26.00 3DCONV 45.514 1.0 285.792 6.28 1202.494 26.42 27S 335.884 1.0 432.801 1.29 3273.728 9.75 LE2D 6842374 1.0 * * * * HIMENO 547939 1.0 * * * * GeoMean 4.39 GeoMean 24.38 Fo da ase s o 1kB he esul s can be seen in Table 6. Benchma ks ha ail o execu e wi h a speci ic compile a e shown wi h an as e isk in he able. PGI code shows a e y good pe o mance, ollowed by he OpenUH code, which also