Hardware support for scratchpad memory transactions on GPU architectures
Villegas Fernández, Alejandro,Asenjo-Plaza, Rafael,González-Navarro, María Ángeles,Plata-González, Óscar Guillermo,Ubal, Rafael,Kaeli, David
- Published
- 2017-08-02
- Publisher
- Springer
- Language
- en
Abstract
Graphics Processing Units (GPUs) have become the accelerator of choice for data-parallel applications, enabling the execution of thousands of threads in a Single Instruction - Multiple Thread (SIMT) fashion. Using OpenCL terminology, GPUs offer a global memory space shared by all the threads in the GPU, as well as a low-latency local memory space shared by a subset of the threads. The latter is used as a scratchpad to improve the performance of the applications. We propose GPU-LocalTM, a hardware transactional memory (TM), as an alternative to data locking mechanisms in local memory. GPU-LocalTM allocates transactional metadata in the existing memory resources, minimizing the storage requirements for TM support. In addition, it ensures forward progress through an automatic serialization mechanism. In our experiments, GPU-LocalTM provides up to 100X speedup over serialized execution.
Full text
Ha dwa e suppo o sc a chpad memo y
ansac ions on GPU a chi ec u es
Alejand o Villegas1, Ra ael Asenjo1, Angeles Na a o1, Osca Pla a1,
Ra ael Ubal2, and Da id Kaeli2
1Depa men o Compu e A chi ec u e, Uni e si y o M´alaga, Andaluc´ıa Tech,
29071 M´alaga, Spain,
{a illegas, magonzalez, asenjo, opla a}@uma.es,
2Depa men o Elec ical and Compu e Enginee ing, No heas e n Uni e si y,
Bos on, MA, USA,
{ubal, kaeli}@ece.neu.edu,
Abs ac . G aphics P ocessing Uni s (GPUs) ha e become he accel-
e a o o choice o da a-pa allel applica ions, enabling he execu ion o
housands o h eads in a Single Ins uc ion - Mul iple Th ead (SIMT)
ashion. Using OpenCL e minology, GPUs o e a global memo y space
sha ed by all he h eads in he GPU, as well as a low-la ency local
memo y space sha ed by a subse o he h eads. The la e is used as a
sc a chpad o imp o e he pe o mance o he applica ions.
We p opose GPU-LocalTM, a ha dwa e ansac ional memo y (TM),
as an al e na i e o da a locking mechanisms in local memo y. GPU-
LocalTM alloca es ansac ional me ada a in he exis ing memo y e-
sou ces, minimizing he s o age equi emen s o TM suppo . In ad-
di ion, i ensu es o wa d p og ess h ough an au oma ic se ializa ion
mechanism. In ou expe imen s, GPU-LocalTM p o ides up o 100X
speedup o e se ialized execu ion.
Keywo ds: T ansac ional Memo y, Sc a chpad Memo y, GPGPU
Conclusions
In his pape we p esen GPU-LocalTM as a ha dwa e TM o GPU a chi ec-
u es ha ocuses on he use o local memo y. GPU-LocalTM is in ended o
limi he amoun o addi ional GPU ha dwa e needed o suppo TM. We p o-
pose wo al e na i e con lic de ec ion mechanisms a ge ing di e en ypes o
applica ions. Con lic de ec ion is pe o med pe -bank, ensu ing scalabili y o he
solu ion. We ind ha o some applica ions he use o TM is no op imal and
discuss how o imp o e ou implemen a ion o be e pe o mance. Fu he -
mo e, GPU-LocalTM in oduces a se ializa ion mechanism o ensu e o wa d
p og ess.
Acknowledgemen s
This wo k has been suppo ed by p ojec s TIN2013-42253-P and TIN2016-
80920-R, om he Spanish Go e nmen , P11-TIC8144 and P12- TIC1470, om
Jun a de Andaluc´ıa, and Uni e sidad de M´alaga, Campus de Excelencia In e -
nacional, Andaluc´ıa Tech.