scieee Science in your language
[en] (orig)

Comparative Analysis of OpenACC Compilers

Abstract

Producción Científica

Read accessible full text

Comparative Analysis of OpenACC Compilers

Author: Barba Gutiérrez, Daniel,González Escribano, Arturo,Llanos Ferraris, Diego Rafael
Publisher: Springer
Year: 2016
DOI: 10.1007/978-3-319-49956-7_7
Source: https://uvadoc.uva.es/bitstream/10324/29131/1/comp_analysis.pdf
Compa a i e Analysis o OpenACC Compile s
Daniel Ba ba, A u o Gonzalez-Esc ibano, and Diego R. Llanos ?
Uni e sidad de Valladolid, Depa amen o de In o ma ica, Valladolid, Spain
{daniel|a u o|diego}@in o .u a.es
Abs ac . OpenACC has been on de elopmen o a ew yea s now. The
OpenACC 2.5 speci ica ion was ecen ly made public and he e a e some
ini ia i es o de eloping ull implemen a ions o he s anda d o make
use o accele a o capabili ies. The e is much o be done ye , bu cu -
en ly, OpenACC o GPUs is eaching a good ma u i y le el in a ious
implemen a ions o he s anda d, using CUDA and OpenCL as backends.
N idia is in es ing in his p ojec and hey ha e eleased an OpenACC
Toolki , including he PGI Compile . The e a e, howe e , mo e de elop-
men s ou he e. In his wo k, we analyze di e en a ailable OpenACC
compile s ha ha e been de eloped by companies o uni e si ies du -
ing he las yea s. We check hei pe o mance and ma u i y, keeping in
mind ha OpenACC is designed o be used wi hou ex ensi e knowledge
abou pa allel p og amming. Ou esul s show ha he compile s a e on
hei way o a easonable ma u i y, p esen ing di e en s eng hs and
weaknesses.
1 In oduc ion
OpenACC is an open s anda d which de ines a collec ion o compile di ec i es
o p agmas o execu ion o code blocks on accele a o s like GPUs o Xeon Phi
cop ocesso s. OpenACC aims o educe bo h he equi ed lea ning ime and he
pa alleliza ion o sequen ial code in a po able way [1]. OpenACC speci ica ion
is cu en ly on i s 2.5 e sion [2], which has been eleased ecen ly.
OpenACC was ounded by N idia, CRAY, CAPS and PGI, bu now he e is a
la ge lis o conso ium membe s, bo h om he indus y and academy, including
he Oak Ridge Na ional Labo a o y, he Uni e si y o Hous on, AMD, and he
Edinbu gh Pa allel Compu ing Cen e (EPCC), among o he s. The Co po a e
O ice s a e, a he ime o w i ing his pape , om N idia, Oak Ridge Na ional
Labo a o y, CRAY and AMD. Academic membe ships a e a ailable o in e es ed
ins i u ions.
The e a e se e al compile s suppo ing OpenACC. The PGI Compile ( om
he Po land G oup, which is a subsidia y o N idia o some ime now) is be-
ing dis ibu ed as pa o he N idia OpenACC Toolki , unde a ee 90-day
?This esea ch has been pa ially suppo ed by MICINN (Spain) and ERDF p og am
o he Eu opean Union: HomP og-He Sys p ojec (TIN2014-58876-P), CAPAP-H5
ne wo k (TIN2014-53522-REDT), and COST P og am Ac ion IC1305: Ne wo k o
Sus ainable Ul ascale Compu ing (NESUS). Thanks also o D . Hec o O ega and
D . Ja ie F esno o hei hough ul commen s, help, and encou agemen .
2
ial license, a ee annual academic license, o a comme cial license. The PGI
Compile uses CUDA o OpenCL as backend. CAPS en e p ise, a p o ide o
so wa e and se ices o he High Pe o mance Compu ing communi y, also de-
eloped a compile which suppo s CUDA and OpenCL. Howe e , he company
is no longe in business and i s de elopmen is no a ailable anymo e. CRAY
Inc. has i s own OpenACC compile , a ailable wi h hei compu e s. I is e-
po ed o be one o he mos ma u e comme cial compile s. Howe e , hey ha e
no ye o e ed an academic o ial license o he pu poses o s udies like his
one. Pa hscale Inc., a compile and mul ico e so wa e de elope , also made an
OpenACC implemen a ion on hei ENZO compile . Un o una ely, a e p i a e
communica ions hey we e eluc an o allow us o use hei compile o his
s udy.
The e a e also se e al academic a emp s o de eloping an OpenACC com-
pile . In pa icula , OpenUH [3] de eloped a he Uni e si y o Hous on, and
accULL [4] om Uni e sidad de La Laguna (Spain). Bo h o hem a e a ailable
o ee o anyone in e es ed.
This wo k p esen s a s udy abou he le el o suppo o OpenACC in he
a ailable compile s, examining hei s eng hs and weaknesses, and gi ing in-
sigh s on hei pe o mance.
We analyze he le el o ma u i y o each compile , in e ms o hei com-
ple eness in he suppo o he s anda d and obus ness, using each compile ’s
documen a ion o check wha pa s o he speci ica ion ha e been implemen ed.
We check he suppo o OpenACC compile di ec i es wi h he help o a bench-
ma k sui e, de eloped by he Edinbu gh Pa allel Compu ing Cen e, which will
be desc ibed la e .
Ano he aspec o es is he ela i e pe o mance o he gene a ed code.
Fo his, we would wan o un mo e complex applica ions and measu e he
di e ences be ween he pe o mance o he execu able code gene a ed by each
compile . The ideal si ua ion would be o es applica ions as close o eal
wo ld p oblems as possible, a oiding syn he ic code agmen s. Since he use
o OpenACC is no common ye , we ha e o ely on exis ing benchma ks. A
his poin , compa ing he esul s om OpenACC code wi h CUDA o OpenCL
di ec implemen a ions migh seem app op ia e, bu po ing he di e en bench-
ma ks o hese languages makes he esul dependen on human in e e ence, as
he de elope ’s abili y o CUDA p og amming impac s on he pe o mance.
The expe imen s conduc ed in his wo k we e ca ied ou using se e al bench-
ma ks. Fi s , he EPCC benchma k sui e which con ains a g oup o 13 ke nels
po ed o OpenACC, called “Le el 1” benchma ks. This benchma k sui e also
con ains h ee eal applica ions called “Himeno”, “27s encil”, and “le co e” [5].
We ha e also used he Pa hscale po o he Rodinia benchma k [6]. We also
wan ed o use he OpenACC Valida ion Tes sui e [7] de eloped by he Uni e -
si y o Hous on, bu a his momen ha ool is only a ailable o OpenACC
membe s.
3
Ou conclusion is ha he di e en compile s a e on hei way o a easonable
ma u i y. Howe e , he e is a numbe o ea u es no ully implemen ed ye by
some o he compile s.
The es o his pape is o ganized as ollows. Sec ion 2 desc ibes he selec ed
compile s. Sec ion 3 shows he cha ac e is ics o he benchma k sui es chosen,
enume a ing some p oblems encoun e ed when compiling hem wi h he compil-
e s selec ed. Sec ion 4 con ains he esul o ou analysis, in e ms o comple eness
o he OpenACC ea u es suppo ed, obus ness o compile implemen a ions,
and ela i e pe o mance o he gene a ed code. Finally, Sec ion 5 concludes ou
pape .
2 A ailable Compile s
In he in oduc ion we men ioned se e al compile s. In his sec ion we desc ibe
wi h mo e de ail he compile s we we e able o ob ain and use o his s udy, and
we will discuss hei ins alla ion pa icula i ies on ou Linux based pla o m.
2.1 PGI Compile
The PGI Compile [8] is being de eloped by The Po land G oup, being owned
by N idia. This compile is widely used in webina s, wo kshops, and con e ences.
The PGI Compile is, a he ime o w i ing his pape , a ailable o download
as pa o he OpenACC Toolki om N idia. This oolki includes a 90-day ee
ial, he possibili y o acqui ing an academic license o a whole yea , o buying
a comme cial license.
2.2 accULL
The accULL [4] compile de eloped by Uni e sidad o La Laguna (Spain) is an
open sou ce ini ia i e. accULL consis s on a s uc u e o wo laye s con aining
YaCF [9] (Ye ano he Compile F amewo k) and F angollo [10], a un ime li-
b a y. YaCF ac s as a sou ce- o-sou ce ansla o while F angollo wo ks as an
in e ace p o iding he mos common ope a ions ound in accele a o s.
2.3 OpenUH
The OpenUH [3] compile , de eloped by he Uni e si y o Hous on (USA) is
ano he open sou ce ini ia i e. I makes use o Open64, a discon inued open-
sou ce op imizing compile .
3 Benchma k Desc ip ion
This sec ion desc ibes he di e en benchma ks used in ou wo k, enume a ing
he main cha ac e is ics ha make hem in e es ing o his s udy, and any issue
de ec ed du ing hei compila ion wi h he h ee compile s s udied.
4
3.1 EPCC OpenACC Benchma ks
This benchma k sui e [11] has been de eloped by he Edinbu gh Pa allel Com-
pu ing Cen e (EPCC). The benchma ks a e di ided in h ee ca ego ies: “Le el
0”, “Le el 1” and “Applica ions”. The compiled p og am launches all he bench-
ma ks in he sui e sequen ially. By de aul , he numbe o epe i ions is en, and
he esul is he a e age o each benchma k. Time is measu ed in mic oseconds
in double p ecision, using he OpenMP unc ion omp ge w ime(). We desc ibe
b ie ly he benchma ks included bellow:
Le el 0 Le el 0 includes a collec ion o small benchma ks ha execu e single
hos and accele a o ope a ions, such as memo y ans e s.
Le el 1 Le el 1 benchma ks [12] consis on a se ies o BLAS- ype ke nels. They
a e based on Polybench [13] and Polybench/GPU ke nels. They measu e he
pe o mance o execu ing hose codes. These benchma ks a e un on he CPU
i s in o de o ha e esul s o compa e hose ob ained on he GPU.
A b ie desc ip ion o he di e en issues ound while unning his sui e o
hese compile s ollows.
OpenUH: The ollowing benchma ks canno be compiled due o unsuppo ed
p agmas p esen in hei code:
ke nels i : P oblems using #p agma ke nels i (0).
pa allel p i a e: P oblems decla ing pa ams as p i a e.
pa allel i s p i a e: P oblems decla ing pa ams as i s p i a e.
le co e: P oblems wi h non scala poin e s.
himeno: P oblems wi h non scala poin e s.
accULL: The e is a p oblem ela ed o a unc ion poin e in he hos p og am.
The compile , du ing he sou ce o sou ce ansla ion modi ies he syn ax o he
unc ion poin e . A double (* es )( oid) is con e ed o a double *( es ( oid)).
This was sol ed by manually co ec ing his change in he in e media e C code
gene a ed, e-compiling he objec ile and copying i o he main di ec o y o
link all he objec iles again. No wa ning om any o he p agmas was de ec ed
so we we e able o un all he benchma ks.
3.2 Rodinia OpenACC
Rodinia is a benchma k sui e o he e ogeneous compu ing [14,15]. I includes
applica ions and ke nels o mul ico e CPU o GPU applica ions.
The e is an e o o po exis ing Rodinia benchma ks o OpenACC. Pa h-
scale [6] is wo king on his. We ha e es ed hei Rodinia e sion commi ed o
Gi Hub on Ap il 25, 2014. Mos o he sui e wo ks wi h PGI, bu OpenUH and
accULL ha e many p oblems o compile mos o he es s. We ha e been able o
success ully compile he ollowing benchma ks con ained in he sui e wi h wo
o mo e compile s: gaussian, nw, lud, c d, ho spo , pa h inde , and s ad2.
5
4 E alua ion
In his sec ion we analyze he OpenACC compile s, using bo h documen a ion
and expe imen a ion. We use each compile ’s documen a ion o check he com-
ple eness o OpenACC ea u es suppo ed. Then we use he EPCC benchma ks
o check bo h obus ness and ela i e pe o mance. Finally we check h ead-block
size sensibili y, measu ing he impac on pe o mance o di e en geome ies.
4.1 Expe imen al Se up
We used a N idia GTX Ti an Black o un he expe imen s. This GPU con ains
2880 CUDA co es wi h a clock a e o 980Mhz and 15 SMs. I has 6GB o RAM,
and Compu e Capabili y 3.5. The hos is a Xeon E5-2690 3 wi h 12 co es a a
clock a e o 1.9GHz, and 64GB in ou 12GB modules.
The PGI compile is he one con ained in he N idia OpenACC Toolki , e -
sion 15.7-0, published in Jul 13, 2015. We used OpenUH e sion 3.1.0 (published
in No embe 4, 2015), based on Open64 e sion 5.0 and using GCC 4.2.0, p e-
buil , downloaded om he High Pe o mance Compu ing Tools g oup websi e
[16]. accULL is e sion 0.4alpha (published in No embe 28, 2013), downloaded
om Uni e sidad de La Laguna’s esea ch g oup “Compu aci´on de Al as P es a-
ciones” [17].
4.2 Comple eness o OpenACC Fea u es Suppo ed
F om each compile documen a ion we ge some insigh on he comple eness o
he OpenACC ea u es suppo ed. F om his in o ma ion, we can conclude ha
he OpenACC s anda d is no ully implemen ed ye by any o he a ailable
compile s. The e is wo k o be done, bu he h ee compile s a e a a espec able
ma u i y le el.
4.3 Robus ness and P agma Implemen a ion
The EPCC Benchma k sui e con ains se e al benchma ks o es ing OpenACC
di ec i es. These benchma ks a e con ained in he “Le el 0” g oup, which has
been desc ibed in he p e ious sec ion. Table 1 con ains he esul s ob ained
o he h ee compile s. In his sec ion we enume a e he p oblems wi h each
benchma k and we explain he esul s ob ained, including he o e head o he
di e en p agma implemen a ions.
Excep o Upda e hos ,Ke nels In oc., and Pa allel In oc., he ime shown
is he di e ence be ween execu ing and no execu ing each p agma. When he
o e head is ze o (o he p agma is no implemen ed), he imes a e e y simila ,
wi h minimal s ochas ic a ia ion. These a ia ions may p oduce a e y small
nega i e esul when calcula ing he di e ence. When di e ences in ime a e
on he o de o ens o mic oseconds (posi i e o nega i e), i can be assumed
ha he e is no di e ence in ime be ween he di e en e sions es ed in ha
benchma k.

6
Table 1: EPCC le el 0: di ec i e’s o e head (in µsec), 1 MB da ase
EPCC L0 PGI OpenUH accULL
Ke nels i -37.50 Fail 4.54
Pa allel i -30.76 -0.48 1237.02
Pa allel p i a e -21.94 Fail 51.09
Pa allel 1s p i Fail Fail -213.83
Ke nels comb. -1.67 -108.43 -127.17
Pa allel comb. -0.05 -2.74 33.38
Upda e hos 478.63 373.22 548.77
Ke nels In oc. Fail 12.76 2398.20
Pa allel In oc. 31.81 13.47 1377.88
Pa allel educ . -14.85 -164.41 -2168.12
Ke nels educ . -8.49 -172.31 -2009.11
PGI The e was a p oblem wi h he “Ke nels In oca ion” benchma k: I e-
u ned an inco ec esul . The code was no being pa allelized and he p agmas
we e igno ed because i wasn’ speci ically s a ed ha he i e a ions we e inde-
penden . This could be sol ed adding he keywo d es ic o he poin e o he
clause independen o he p agma.
The “ke nels i ” and “pa allel i ” esul s a e e y simila , and in bo h cases
he esul s indica e ha he code wi h he p agma is sligh ly as e han he
one wi hou i , e en hough bo h a e being un on he hos . In [5] i was s a ed
ha his could be because o op imiza ions done by he compile while o a e
p ocessing he p agmas.
The “pa allel p i a e” benchma k shows ha he c ea ion o p i a e a iables
o each h ead unning he loop is sligh ly as e han he alloca ion o de ice
memo y.
“Ke nels combined” shows a e y small di e ence o ime be ween w i ing
wo p agmas ins ead o a combined one, he o me being sligh ly as e han
he la e al hough he di e ence is almos negligible. The same occu s o he
“pa allel combined” benchma k, he di e ence being smalle in his case.
Finally “Pa allel educ ion” and “Ke nels educ ion” show ha PGI has e y
li le o e head o he educ ion clause. In [5] i is s a ed ha he PGI compile
does he educ ion e en i i is no anno a ed. This could explain he e y small
di e ence in bo h benchma ks.
OpenUH We go some e o s du ing compila ion o he “Ke nels i ”, “Pa -
allel p i a e” and “Pa allel 1s p i a e” benchma ks so hey a e igno ed in his
analysis. Howe e , he “Pa allel i ” di ec i e is suppo ed and he di e ence be-
ween using he p agma o un code on he hos o unning i di ec ly is almos
negligible.
“Ke nels combined” shows an o e head o he combined p agma e sus he
sepa a ed e sion. Howe e , his is no he case o he “Pa allel combined”
benchma k, whe e he di e ence is much smalle . The in oca ion o ke nels and
7
pa allel di ec i es a e e y simila . And o bo h o hem, he educ ion adds
a simila o e head. This migh be ela ed o OpenUH assuming loops o be
independen inside ke nels egions.
accULL No e o s we e shown while compiling o unning he benchma ks
wi h accULL. The e is a big di e ence be ween he wo e sions con ained in
he “Ke nels i ” and he “Pa allel i ” benchma ks, whe e he ke nels di ec i e
e sion has a e y small o e head compa ed o he non-anno a ed code. This
o e head is e y la ge in he pa allel di ec i e e sion. This is explained by he
accULL de elope s in [5] whe e hey say ha he absence o a loop clause in he
pa allel di ec i e is causing he loop o be execu ed sequen ially in each h ead.
The e o e, his clause is no co ec ly suppo ed, as we unde s and om he
OpenACC Speci ica ion ha he loop should be execu ed only on he hos .
Robus ness Summa y The o e all esul s indica e ha some o he clauses
a e no implemen ed ye , bu he h ee compile s a e in hei way o a easonable
ma u i y le el and, since he mos used di ec i es a e wo king, hey can ac ually
be used o code pa alleliza ion using OpenACC.
4.4 Rela i e Pe o mance o Gene a ed Code
In his sec ion we analyze he pe o mance o he gene a ed code desc ibing
he impac o p agmas o e head in accULL. Pe o mance measu emen is di-
ided in o da a mo emen , whe e we analyze he esul s o he da a mo emen
benchma ks in Le el 0 o EPCC OpenACC Benchma k Sui e, and execu ion pe -
o mance, using Le el 1 and Applica ion Le el o EPCC OpenACC Benchma k
Sui e, and selec ed benchma ks om Rodinia.
E ec o P agmas O e head in accULL Some esul s om he Le el 0 o
he EPCC Benchma k Sui e show a pe o mance impac in oduced by some
clauses and di ec i es in he accULL gene a ed code.
Ke nels and Pa allel in oca ions in accULL ha e a highe o e head han
o he compile s. This is due o he un ime calls and i is specially no iceable in
he educ ion clause. These o e heads accumula ion does no ha e a signi ican
impac o complex ke nels, o launching he same ke nel o e and o e again.
Howe e , his could be a p oblem when unning simple ke nels o many di e en
small ke nels. This is he main eason behind he o e all esul s showing a wo se
pe o mance o he accULL compile in his analysis.
Da a Mo emen Da a mo emen pe o mance can be measu ed in ou bench-
ma ks om he Le el 0 o he EPCC OpenACC benchma k sui e. We ha e
launched 10 epe i ions o hose benchma ks wi h da asizes o 1 kB, 1 MB, 10
MB, and 1 GB. The esul s can be seen in Tables 2, 3, 4, and 5.
8
Table 2: EPCC da a mo emen esul s (in µsec), 1 kB da ase . Whi e cells
highligh he bes esul s. no m. is he no malized esul using PGI as e e ence.
Da a M mn PGI OpenUH accULL
10 eps, 1kB ime no m. ime no m. ime no m.
Con igH2D 30.827 1.0 322.699 10.47 338.218 10.97
Con igD2H 14.686 1.0 323.319 22.01 343.919 23.42
SlicedH2D 12.087 1.0 310.897 25.72 315.914 26.13
SlicedD2H 14.948 1.0 324.010 21.67 327.714 21.92
GeoMean 18.93 GeoMean 19.58
Table 3: EPCC da a mo emen esul s (in µsec), 1 MB da ase . Whi e cells
highligh he bes esul s. no m. is he no malized esul using PGI as e e ence.
Da a M mn PGI OpenUH accULL
10 eps, 1MB ime no m. ime no m. ime no m.
Con igH2D 484.347 1.0 950.789 1.96 727.839 1.50
Con igD2H 461.936 1.0 632.691 1.37 792.761 1.72
SlicedH2D 17.094 1.0 267.982 15.68 274.462 16.06
SlicedD2H 36.335 1.0 254.702 7.01 285.685 7.86
GeoMean 4.14 GeoMean 4.24
Table 4: EPCC da a mo emen esul s (in µsec), 10 MB da ase . Whi e cells
highligh he bes esul s. no m. is he no malized esul using PGI as e e ence.
Da a M mn PGI OpenUH accULL
10 eps, 10MB ime no m. ime no m. ime no m.
Con igH2D 4141.402 1.0 6887.984 1.66 3354.666 0.81
Con igD2H 5876.088 1.0 2043.747 0.35 4396.052 0.74
SlicedH2D 27.322 1.0 404.214 14.79 427.203 15.64
SlicedD2H 48.017 1.0 269.818 5.62 280.203 5.84
GeoMean 2.64 GeoMean 2.72
Table 5: EPCC da a mo emen esul s (in µsec), 1 GB da ase . Whi e cells
highligh he bes esul s. no m. is he no malized esul using PGI as e e ence.
Da a M mn PGI OpenUH accULL
10 eps, 1GB ime no m. ime no m. ime no m.
Con igH2D 32310.009 1.0 788945.913 24.42 296340.991 9.17
Con igD2H 55179.119 1.0 553282.976 10.03 347280.359 6.29
SlicedH2D 400.066 1.0 535.011 1.34 533.943 1.34
SlicedD2H 158.071 1.0 2818.100 17.83 4294.407 27.17
GeoMean 8.75 GeoMean 6.76
9
In [5] i was s a ed ha PGI used pinned memo y and ha i was causing
issues in smalle da ase s. I seems ha PGI has sol ed his issue since hen and,
looking a he documen a ion, i is now possible o speci y he ype o memo y
access we wan wi h a compila ion lag. When using la ge da ase s, OpenUH and
accULL do no show he expec ed esul s acco ding o he e olu ion shown in
ables 2, 3, and 4. We guess ha his is ela ed o he usage o pinned memo y
by he PGI Compile , allowing i o ob ain be e esul s when da ase s a e big
enough.
Execu ion Pe o mance, EPCC Benchma ks In his sec ion we will ana-
lyze he pe o mance o he code gene a ed by he PGI, OpenUH, and accULL
compile s wi h he benchma ks con ained in he EPCC Le el 1 and Applica ion
le el. We use h ee di e en da ase s: 1kB, 1MB, and 10MB. This choice is based
on he ac ha bigge da ase s esul in an ou o memo y e o due o how
he benchma ks y o alloca e memo y on he de ice. We suspec he memo y
alloca ion is being done in each h ead inside he gene a ed ke nels, using mo e
memo y han expec ed. In summa y, PGI code ob ains be e esul s in almos
e e y benchma k. Howe e , he di e ences sho en when using bigge da ase s.
Table 6: EPCC execu ion esul s (in µsec), 1 kB da ase . Whi e cells highligh
he bes esul s.
Exec. ime PGI OpenUH accULL
10 eps, 1kB ime no m. ime no m. ime no m.
2MM 99.087 1.0 522.304 5.27 2799.229 28.25
3MM 80.204 1.0 380.683 4.75 3799.048 47.37
ATAX 58.103 1.0 327.110 5.63 2564.702 44.14
BICG 72.408 1.0 350.380 4.84 2628.499 36.30
MVT 80.037 1.0 354.743 4.43 2665.299 33.30
SYRK 68.426 1.0 289.512 4.23 2394.803 35.00
COV 87.261 1.0 314.617 3.61 3795.372 43.49
COR 104.976 1.0 337.362 3.21 5208.668 49.62
SYR2K 73.290 1.0 317.574 4.33 2469.765 33.70
GESUMMV 65.613 1.0 312.996 4.77 1500.021 22.86
GEMM 49.710 1.0 323.725 6.51 1237.473 24.89
2DCONV 46.444 1.0 286.174 6.16 1207.528 26.00
3DCONV 45.514 1.0 285.792 6.28 1202.494 26.42
27S 335.884 1.0 432.801 1.29 3273.728 9.75
LE2D 6842374 1.0 * * * *
HIMENO 547939 1.0 * * * *
GeoMean 4.39 GeoMean 24.38
Fo da ase s o 1kB he esul s can be seen in Table 6. Benchma ks ha ail
o execu e wi h a speci ic compile a e shown wi h an as e isk in he able. PGI
code shows a e y good pe o mance, ollowed by he OpenUH code, which also