scieee Open visual document viewer

TORMENT OpenACC2016: A benchmarking tool for OpenACC compilers

Barba Gutiérrez, Daniel,González Escribano, Arturo,Llanos Ferraris, Diego Rafael

Abstract

Producción Científica

Full text

TORMENT OpenACC2016: A benchma king ool o OpenACC compile s Daniel Ba ba, A u o Gonzalez-Esc ibano, Diego R. Llanos Dp o. de In o m´ a ica Uni e sidad de Valladolid Valladolid, Spain daniel@in o .u a.es, a u o@in o .u a.es, diego@in o .u a.es Abs ac —OpenACC is a pa allel p og amming model o ha dwa e accele a o s, such as GPUs o Xeon Phi, which has been in de elopmen o se e al yea s by now. Du ing his ime, di e en compile s ha e appea ed, bo h comme - cial and open sou ce, which a e s ill on de elopmen s age. Due o he ac ha bo h he OpenACC s anda d and i s implemen a ions a e ela i ely ecen , we p opose a bench- ma k sui e speci ically designed o check he pe o mance o he OpenACC ea u es in he code gene a ed by di e en compile s on di e en a chi ec u es. Ou benchma k sui e is named TORMENT OpenACC2016. Along wi h his ool we ha e de eloped an adequa e me ic o he compa ison o pe o mance among di e en machine-compile pai s which we ha e named TORMENT ACC2016 Sco e. The e sion 1 o TORMENT OpenACC2016 p esen ed in his pape , con ains six benchma ks, and is a ailable online. Keywo ds-OpenACC, compile s, benchma king. I. INTRODUCTION OpenACC is an open s anda d which de ines a numbe o compila ion di ec i es o p agmas o pa allel code ag- men s execu ion using ha dwa e accele a o s such as GPUs and Xeon Phi. I s objec i e is o ease pa alleliza ion o sequen ial code using his kind o accele a o s and educing he equi ed ime bo h o lea ning and coding [1]. The OpenACC speci ica ion is cu en ly on i s 2.5 e sion [2]. A he ime o w i ing his pape , he e a e se e al bench- ma ks o OpenACC. Fo example, he EPCC OpenACC Benchma k Sui e [3] has been de eloped by he Edinbu gh Pa allel Compu ing Cen e, and consis s on bo h a numbe o mic obenchma ks in ended o check he compliance o he s anda d and measu e he o e head o di e en p agmas implemen a ion. I also measu es he gene al pe o mance o OpenACC gene a ed code o each benchma k, bu se e al compile s lack ull suppo o all he ea u es o his ool ye . This implies ha i is no s aigh o wa d o do compa a i e pe o mance analysis o di e en compile s. On he o he hand, he e a e o he benchma k sui es no o igi- nally designed o OpenACC. Fo example, Rodinia [4] o OpenACC is a po o he o iginal benchma ks om he sui e o he same name [5], de eloped by Pa hscale. As i is no speci ically designed o OpenACC, mos o he benchma ks included in his sui e p esen many compila ion p oblems wi h he s a e-o - he-a OpenACC compile s, and hey do no co e all o he OpenACC ea u es in a sys ema ic way. I would be in e es ing o he communi y o ha e a bench- ma k sui e ha co e s he main ea u es o he OpenACC s anda d and easily allows o compa e he imp o emen s in code gene a ion o di e en a chi ec u es in e ms o pe o mance. The con ibu ion o his wo k includes: •A benchma k sui e designed o OpenACC, ha co e s he main ea u es o he s anda d based on p og am- ming pa e ns o di e en kinds o eal applica ions. •A p oposal o a me ic o p o ide a ela i e sco e which allows o he compa ison o he pe o mance ob ained by he implemen a ion o di e en compile s on di e en pla o ms. The design o his me ic allows o classi y he esul s in a global anking. We named i TORMENT ACC2016 Sco e. •A compila ion-execu ion w appe , along wi h a sys em- a ic way o include in he sou ce code condi ional com- pila ion s a emen s o adap he OpenACC syn ax and de ails o he speci ic pa icula i ies o each compile . This w appe allows o easily adap he benchma ks o di e en OpenACC compile s and pla o ms. I also au oma izes he gene a ion o a inal epo including he me ic measu e and o he ele an da a o he compa isons. We p esen TORMENT OpenACC2016. I is a ool ha includes he benchma ks and he w appe ha au oma izes he compila ion and execu ion. The gene a ed epo s a e p esen ed in a s anda d and easy eadable way, including in o ma ion abou he execu ion pla o m, he compile s, s a is ics o execu ion imes, and he inal alues o he me ic o each case. The epo also o e s s a is ics o execu ion imes o sequen ial and na i e CUDA e sions o he benchma ks o u he e e ence. The ool, in i s cu en s a e, has been alida ed o h ee di e en compile s in wo di e en GPU pla o ms. The suppo o he used compile s o he gene a ion o Xeon Phi code is no ye ma u e enough o ex end ou alida ion o hese kind o de ice. The es o he pape is s uc u ed as ollows: Sec ion II desc ibes wi h mo e de ail p e ious benchma king ools. Sec ion III discusses he ma u i y o OpenACC compile s. Sec ion IV desc ibes TORMENT OpenACC2016, including i s goals, s uc u e o he ool, implemen ed benchma ks, he p oposed me ic, and commen s he epo ob ained in he expe imen al alida ion. Finally, sec ion V concludes he pape . II. EXISTING TOOLS As i was s a ed in he in oduc ion, a he ime o w i ing his pape he e a e p e ious pe o mance e alua ion ools o OpenACC. In his sec ion we will desc ibe mo e ho - oughly examples o hese ools and he p oblems de ec ed o use hem o he goals o his wo k. A. EPCC OpenACC Benchma k Sui e This benchma k sui e de eloped by he EPCC has been designed speci ically o OpenACC. I is composed o h ee di e en pa s, named Le el 0,Le el 1 yApplica ion Le el. Le el 0 con ains se e al mic obenchma ks whose goal is measu ing he o e head caused by he implemen a ion o he di e en p agmas. These mic obenchma ks o e a esul calcula ed as he di e ence be ween he execu ion imes o wo di e en e sions o he use o he OpenACC p agmas. Fo #p agma acc da a di ec i e i simply measu es he equi ed ime o da a ans e s. Le el 1 con ains a se o BLAS ype benchma ks based on Polybench and Polybench/GPU [6]. The esul s o e ed by hese benchma ks a e he execu ion ime o he di e en code agmen s. This means hey a e a good pe o mance indica o o he OpenACC gene a ed code. Howe e , in he cu en s a e o de elopmen o he di e en compile s some o he benchma ks use unsuppo ed p agmas. Due o his, hese benchma ks p esen compila ion o un ime p oblems wi h a numbe o compile s. Applica ion Le el con ains h ee benchma ks o a bigge complexi y han hose om he p e ious sec ion. These benchma ks also o e as hei esul s he execu ion ime o he gene a ed code. The EPCC OpenACC Benchma k Sui e is well planned and designed. Howe e , he esul s ob ained om he mi- c obenchma ks a e no adequa e enough o a ela i e pe - o mance analysis. The o he benchma ks om Le el 1 and Applica ion Le el could be adequa e. Howe e , since hey a e no speci ically designed o es OpenACC ea u es suppo ed by all o he a ailable compile s, hey p esen p oblems. In some cases hey couldn’ be compiled wi h some compile s, and some imes hey p oduced un ime e o s. B. Rodinia o OpenACC Pa hscale Inc. has de eloped an OpenACC e sion [4] o he Rodinia benchma k sui e [5], [7]. The cu en e sion a he ime o w i ing his pape has been published on Gi Hub (Ap il, 24 h 2014). The sui e is composed o benchma ks ha e u n as hei esul he execu ion ime. This means ha hey a e a good s a ing poin o a pe o mance analysis. The main issues a e he low ma u i y le el o he OpenACC use in he code, and he lack o compa ibili y wi h he a ailable compile s. These issues make i impossible o es ablish a ai pe o mance compa ison. III. OPENACC COMPILERS The e is a numbe o choices among compile s wi h suppo o OpenACC, bo h comme cial and ee. In his wo k we ocus on hose compile s which a e a ailable o ee o wi h academic o ial licenses. In gene al, he OpenACC compile s ha e cu en ly e y limi ed suppo o any pla o m o he han GPUs. The compile s conside ed a e: A. PGI Compile The PGI compile [8], de eloped by The Po land G oup and N idia, is, a he ime o w i ing, he compile wi h he highes ma u i y le el. In ou s udies, he PGI compile has shown o be he mos s able and he highes ma u i y le el among he used compile s. I has a p ope implemen a ion o mos o he OpenACC speci ica ion. B. OpenUH OpenUH compile , de eloped by he Uni e si y o Hous- on, is an open sou ce p ojec . Acco ding o ou es s and compa ed o he PGI compile , i s ma u i y and obus ness a e lowe and i lacks se e al unc ionali ies. C. accULL The accULL compile [9] has been de eloped by Uni- e sidad de La Laguna, Spain, and i is an open sou ce compile like OpenUH. Ou es s show ha compa ed o he p e ious compile s, his is he one wi h a lowe obus ness, leading o se e al issues du ing sou ce o sou ce ansla ion, no suppo ing se e al C cha ac e is ics, like unc ion poin - e s. Se e al OpenACC unc ionali ies a e also lacking, like se e al ypes o educ ion. IV. TORMENT OPENACC2016 Ou p oposal is implemen ed as a ool called TORMENT OpenACC2016, an ac onym o T asgo pe ORMance and E alua ioN Tool o OpenACC. The p ojec was bo n o gi e esponse o he need o pe o mance analysis and compa ison o OpenACC code gene a ed by he a ious compile s implemen ing he s anda d. A. Goals The main goal o TORMENT OpenACC2016 is o allow a pe o mance analysis o OpenACC code gene a ed by di e en compile s in an easy way, gene a ing a esul summa y ha can be easily analyzed and which o e s he TORMENT ACC2016 Sco e me ics. This me ic is use ul o compa e machine-compile pai s in e ms o pe o mance. Ou p oposal in ends o o e a ool speci ically p epa ed o OpenACC, de eloped aking in o accoun he ea ly de elopmen s age o he a ailable compile s. Thus, we y o ensu e ha compila ion o un ime p oblems a e a oided o he compile s conside ed, unlike in he p e ious benchma k sui es a ailable. TORMENT OpenACC2016 is p epa ed o be compiled and execu ed wi h minimum in e en ion om he use . The sc ip s accompanying he sui e ge in o ma ion om he benchma king p ocess and inally o e a HTML epo wi h he mos ele an da a. Mo eo e , TORMENT OpenACC2016 makes use o he GCC and NVCC compile s in o de o ob ain da a om sequen ial and CUDA-based code execu ions. Wi h his da a, ou p oposal can o e o he use in o ma ion abou he speedup compa ed wi h sequen ial and CUDA code execu ed in he same machine. B. S uc u e o ou ool TORMENT OpenACC2016 is composed o a numbe o shell sc ip s which a e in cha ge o he compila ion, exe- cu ion and esul s ga he ing p ocesses, emo ing he bu den om he use . Each benchma k has been chosen o ep esen a ele an class o p oblems, and has been implemen ed o be as simple as possible, in o de o a oid compila ion issues. The benchma ks a e launched en imes, wi h an ex a in oca ion a he beginning whose esul s a e disca ded. The alues ha co espond o he bes o he execu ions and he a i hme ic mean o all o hem a e eco ded o la e calcula e he p oposed me ic. These alues a e named peak and a e age espec i ely. C. Fea u es Co e age In o de o make a selec ion o benchma ks o ou sui e we ha e analyzed he de elopmen s a us o he di e en compile s. We ha e s udied hei capabili ies and hei comple eness o he OpenACC s anda d. Ou conclussion is ha he exis ing benchma ks, based on Polybench and Rodinia, o en equi e p agmas ha a e no implemen ed ye on one o he compile s a leas . To achie e ull compa ibili y among he di e en compile s we need o use only he se o di ec i es a ailable o all o hem a he ime o de eloping his e sion o TORMENT OpenACC. We lis below a numbe o GPU cha ac e is ics ha we a gue a e impo an o be es ed o he cu en s a us o compile s’ ma u i y: •Use o sha ed memo y •In ensi e and non-balanced compu a ion •Compu a ion wi hou memo y access •Da a ans e s •Non-pe ec ly coalescen memo y access These a e impo an cha ac e is ics ha play a pi o al ole in he esul ing code o manually w i en CUDA ke nels. The code gene a ed by he OpenACC compile s will also be a ec ed by hese cha ac e is ics, hus he pe o mance o hese gene a ed codes could be measu ed and compa ed. As mos o he well-known benchma ks equi e he use o non-suppo ed di ec i es o one compile a leas , we ha e decided o ollow a bo om-up app oach o de eloping a wo king sui e o benchma ks, ying o co e he cha ac e - is ics lis ed abo e. This decision lea es us wi h a se o e y simple benchma ks, use ul o a compa a i e analysis o he pe o mance o OpenACC gene a ed code using di e en compile s, bu no o a ho ough analysis o comple eness, obus ness and pe o mance o a single compile . Ou wo k aims o ob ain a ool capable o gene a ing a compa a i e analysis epo o a wide numbe o compile s. As he suppo o OpenACC di ec i es imp o es, his design decision should also be e iewed and adap ed in o de o use benchma ks o a highe complexi y as he ones p o- ided in sui es like Polybench (used in he EPCC OpenACC benchma k sui e), Rodinia, and he CUDA Toolki . D. Implemen ed Benchma ks Ve sion 1 o TORMENT OpenACC2016, con ains a se o six benchma ks. These benchma ks y o co e he di e en cha ac e is ics enume a ed in he p e ious sec ion, as i is shown in Table I. 1) T Mon eCa loPi: This is an emba assingly pa allel, wi hou dependencies, pe ec ly egula , well balanced wi h a s a ic pa i ion, and compu a ional bound example. This benchma k consis s on an app oxima ion o Pi using he Mon e Ca lo me hod. This me hod is based on he andom gene a ion o coo dina es in a uni squa e. Fo each coo di- na e, i is checked whe he o no ha coo dina e is loca ed inside a qua e ci cle and, i ha condi ion is sa is ied, hen he coo dina e is accumula ed. Finally, he ollowing o mula is applied: π≈4∗P T Whe e Pis he numbe o coo dina es inside he qua e ci cle and Tis he o al o gene a ed coo dina es. T Mon eCa loPi is a benchma k ha con ains no memo y ans e s and a e y simple compu a ion. Howe e , i can be op imized i each h ead compu es se e al coo dina es and i he sha ed memo y in he h ead blocks is used in o de o a oid global memo y access. The unc ion used o andom numbe gene a ion can be p oblema ic. OpenACC code should use unc ions ha can be o loaded o he GPU and a he same ime can be used on he hos machine. Thus, we canno use na i e CUDA unc ions like he ones included in he cu and lib a y. Ou solu ion has been o eplica e he implemen a ion o he s and unc ion in he s anda d C lib a y. 2) T S ingMa ch: This example is coalesced wi h low memo y bandwid h and i egula da a-dependan loads pe h ead. The T S ingMa ch benchma k is a s ing ma ching p og am. This algo i hm is in e es ing because i combines da a ans e wi h he need o e icien use o he memo y, specially he sha ed memo y. I consis s on he sea ch o he i s occu ence o a small s ing in a la ge one, using anai e algo i hm. In his benchma k, he la ge s ing has a Sha ed Memo y In ensi e Compu a ion Compu a ion wi hou memo y access Da a T ans e Non-pe ec ly coalescen Memo y Access T Mon eCa loPi X X T S ingMa ch X X T 3DS encil X T Mandelb o X T Ma ixMul X X T Re e seA ay X X X Table I GPU CHARACTERISTICS TESTED BY THE BENCHMARKS size o 10 million cha ac e s, ha is, app oxima ely 10MB. We sea ch ou small s ings o 1000 cha ac e s. 3) T 3DS encil: This is a egula , well balanced example wi h dependencies ha lead o i e a i e neighbou synch o- niza ion, and low load pe h ead on each i e a ion. The T 3DS encil benchma k consis s on a 6 poin 3D s encil. Each poin o a 3D ma ix is upda ed wi h he a i hme ic mean o i s neighbou s. This benchma k is in e es ing be- cause, adi ionally, s encil p og ams can be e y e icien when execu ed on a GPU. I includes memo y ans e s only a he beginning and end o he compu a ion. 4) T Mandelb o : This example is emba assingly pa - allel wi hou dependencies wi h non-balanced load. The T Mandelb o implemen s he mos amous o he ac al se s. The implemen ed algo i hm is he escape- ime algo- i hm. This is an emba assingly pa allel p oblem and he e a e no dependencies among he compu a ions. The e o e, he e is no need o use he sha ed memo y. Global memo y access is also negligible. On he o he hand, he compu a ion is e y in ensi e, depending on he maximum numbe o i e a ions. This benchma k should scale e y well on GPU. 5) T Ma ixMul : This example has pe ec ly egula loads wi h coalesced memo y access i sha ed memo y is p ope ly used. Ma ix mul iplica ion is widely used in many domains. I has a la ge ma gin o op imiza ions. When using a GPU, he mos impo an op imiza ion is he co ec use o he sha ed memo y, which equi es a co ec h ead block geome y. As his benchma k has h ee nes ed loops, i p esen s ce ain complexi y o he OpenACC compile s. They need o make choices a he di e en le els o pa al- lelism, including also he geome y o he h ead blocks. 6) T Re e seA ay: This las example has a egula load, bu a di ec implemen a ion o non-aligned a ay sizes de- i es in coalesced eads bu non-pe ec ly coalesced w i es. The las o he benchma ks implemen ed in TORMENT OpenACC2016 consis s on a e y simple ope a ion; e e s- ing an in ege a ay. Al hough his seems o be a i ial ope a ion, unning i on a GPU p esen s a ele an p oblem. Re e sing an a ay on a GPU can be e y ine icien because o non-coalescen global memo y access. The co ec solu- ion in ol es using he sha ed memo y as an in e media e bu e whe e a pa ial e e se is done. This allows o coalescen global memo y access. E. P oposed Me ic Fo he selec ion o he me ic named TORMENT ACC2016 Sco e, we use he same me hodology as SPEC [10] (S anda d Pe o mance E alua ion Co po a ion). I is a widely known e e ence in e ms o benchma king and pe o mance analysis. One o i s s eng hs is o acknowledge ha benchma ks ge olde as ime passes and, in consequence, hey mus be upda ed. SPEC uses he ollowing me hodology. Fi s , he p og am e u ns i s execu ion ime. Then he SPEC a io is calcula ed, which consis s on he a io ob ained om di iding he e e - ence execu ion ime (supplied by SPEC) and he execu ion ime ob ained by he benchma k. Finally, he geome ic mean o all he SPEC a ios is calcula ed [11]. Ou p oposal ollows a simila idea bu wi h se e al changes. Fi s , he execu ion o he benchma ks con ained in TORMENT OpenACC2016 e u n h ee alues. Peak Time is he bes execu ion ime ob ained, measu ed in seconds. A e age Time is he a i hme ic mean o all he execu ions o he same benchma k, also in seconds. Finally, he S anda d De ia ion o he se o imes ob ained gi es an idea o he a iabili y among he se e al execu ions o he benchma ks. Wi h he peak and a e age imes, a a io is calcula ed di id- ing he e e ence imes p o ided by he ool, which a e he execu ion imes o he sequen ial e sion o he benchma ks in a e e ence machine. The cha ac e is ics o his e e ence machine a e shown in h ps:// o men .in o .u a.es. Finally, i compu es he ha monic mean o all he a ios, bo h o peak ime as a e age ime. These alues a e he TORMENT ACC Peak Sco e and TORMENT ACC A e age Sco e. The main di e ence be ween TORMENT OpenACC2016 and SPEC me ics, apa om he ime measu emen me hodology, consis s on he use o he ha monic mean ins ead o he geome ic mean. Fo he goals o TORMENT OpenACC2016, he ha monic mean is mo e adequa e han he geome ic mean. Fi s , e en i he geome ic mean always p oduces a consis en o de ing, i ’s no necessa ily he co ec o de ing [11]. This is because his mean is no in e sely p opo ional o he execu ion ime. In con as , he ha monic mean is in e sely p opo ional o he execu ion ime, which makes i an adequa e mean o a ios (see e.g. [12]). TORMENT ACC2016 Sco e is a HB (Highe is Be e ) me ic. An example o usage can be ound on h ps:// o men .in o .u a.es. The TORMENT Repo Example link shows a epo gene a ed o one o ou machines. The i s pa o he epo consis s on sys em in o ma ion. A e ha sec ion, he esul ables o sequen ial and CUDA code a e ound. These esul s a e me ely in o ma i e, and hey a e no used in he TORMENT ACC2016 Sco e. Howe e , hese esul s allow he use o compa e he pe o mance o he OpenACC code wi h sequen ial and CUDA e sions execu ed in he same machine. A e he sequen ial and CUDA esul s, he OpenACC esul s a e shown o each compile used. These esul s consis on he ables, wi h he execu ion ime s a is ics, he TORMENT ACC2016 Sco e, and he speedup alues using as e e ence he sequen ial and CUDA execu ions. In he Repo Example a ailable online, he PGI compile ob ains he bes esul s, ob aining a 64% o he pe o mance achie ed wi h CUDA implemen a ions. OpenUH ge s a second posi ion, wi h a 41% o he pe o mance achie ed wi h he op imized CUDA implemen a ions. Finally, he accULL compile ge he hi d posi ion, wi h 16% o he pe o mance achie ed wi h CUDA. The TORMENT ACC2016 Sco e also allows o a com- pa ison among compile -machine pai s. Fo example, he Hyd a Sys em in he Repo Example a ailable online has Ti an Black GPUs, and we also an he benchma k sui in a lap op which has a N idia GTX 850M. The TORMENT ACC2016 Peak Sco e in Hyd a is 28.69, while in he lap op is 8.45. On he o he hand he accULL compile ob ains 7.26 in Hyd a and 4.98 in he lap op. This indica es he impo an di e ences o pe o mance in oduced by he code gene a ion and op imiza ion echniques used by bo h compile s. V. CONCLUSIONS AND FUTURE WORK TORMENT OpenACC2016 is a pe o mance analysis and compa ison ool o OpenACC gene a ed code, aking in o conside a ion he ma u i y le el o he suppo o OpenACC in di e en compile s. TORMENT OpenACC2016 con ains a sui e o benchma ks speci ically designed o OpenACC and main aining he maximum po abili y. The esul s o e ed by TORMENT OpenACC2016 include he so called TORMENT ACC2016 Sco e, designed o he compa ison o machine-compile pai s. Mo eo e , i o e s a esul epo o he benchma ks, including a able wi h exe- cu ion imes s a is ics (bo h o peak and a e age alues) and he s anda d de ia ion, and also including execu ion esul s om sequen ial and CUDA code gene a ed by GCC and NVCC compile s. This o e s he use a compa a i e analysis o he OpenACC gene a ed code e sus he sequen ial and CUDA e sions o he benchma ks. Fu u e wo k consis s on he inclusion o new benchma ks, specially when he suppo ed compile s become mo e ma- u e, including a iche s uc u e wi h di e en le els (such as BLAS and eal wo ld applica ions) in addi ion o he exis ing syn he ic applica ions. Mo eo e , an impo an pa o ou u u e wo k consis s on main aining compa ibili y among he suppo ed compile s and new addi ions. We a e in he p ocess o adding suppo o GCC, Omni (Uni e si y o Tsukuba), OpenARC (Oak Ridge Na ional Labo a o y), and RoseACC (Uni e si y o Delawa e). ACKNOWLEDGMENTS This esea ch has been pa ially suppo ed by MICINN (Spain) and ERDF p og am o he Eu opean Union: HomP og-He Sys p ojec (TIN2014-58876-P), and COST P og am Ac ion IC1305: Ne wo k o Sus ainable Ul ascale Compu ing (NESUS). REFERENCES [1] OpenACC-s anda d.o g. Abou OpenACC. [Online]. A ail- able: h p://www.openacc.o g/abou -openacc [2] ——. (2015, oc ) OpenACC 2.5 d a o public commen . [Online]. A ailable: h p://www.openacc.o g/con en /openacc- 25-d a -public-commen [3] EPCC. (2013, sep) Epcc OpenACC benchma k sui e. h ps://gi hub.com/EPCCed/epcc-openacc-benchma ks. [4] Pa hscale. (2014, ap ) Rodinia benchma k sui e 2.1 wi h OpenACC po . h ps://gi hub.com/pa hscale/ odinia. [5] S. Che, M. Boye , J. Meng, D. Ta jan, J. W. Shea e , S.-H. Lee, and K. Skad on, “Rodinia: A benchma k sui e o he - e ogeneous compu ing,” in Wo kload Cha ac e iza ion, 2009. (IISWC), 2009 IEEE In e na ional Symposium on. IEEE, 2009, pp. 44–54. [6] L.-N. Pouche , “Polybench: The polyhed al benchma k sui e,” URL: h p://www. cs. ucla. edu/˜ pouche /so wa e/polybench/[ci ed July,], 2012. [7] S. Che, J. W. Shea e , M. Boye , L. G. Sza a yn, L. Wang, and K. Skad on, “A cha ac e iza ion o he Rodinia bench- ma k sui e wi h compa ison o con empo a y CMP wo k- loads,” in Wo kload Cha ac e iza ion (IISWC), 2010 IEEE In e na ional Symposium on. IEEE, 2010, pp. 1–11. [8] PGI. (2015, no ) Pgi accele a o compile s wi h OpenACC di ec i es. h ps://www.pg oup.com/ esou ces/accel.h m. [9] R. Reyes, I. L´ opez-Rod ´ ıguez, J. J. Fume o, and F. de Sande, “accULL: an OpenACC implemen a ion wi h CUDA and OpenCL suppo ,” in Eu o-Pa 2012 Pa allel P ocessing. Sp inge , 2012, pp. 871–882. [10] K. M. Dixi , “The spec benchma ks,” Pa allel compu ing, ol. 17, no. 10, pp. 1195–1209, 1991. [11] D. J. Lilja, Measu ing compu e pe o mance: a p ac i ione ’s guide. Camb idge Uni e si y P ess, 2005. [12] J. R. Mashey, “Wa o he benchma k means: ime o a uce,” ACM SIGARCH Compu e A chi ec u e News, ol. 32, no. 4, pp. 1–14, 2004.