scieee Science in your language
[en] (orig)

GenArchBench: A genomics benchmark suite for arm HPC processors

Abstract

Arm usage has substantially grown in the High-Performance Computing (HPC) community. Japanese supercomputer Fugaku, powered by Arm-based A64FX processors, held the top position on the Top500 list between June 2020 and June 2022, currently sitting in the fourth position. The recently released 7th generation of Amazon EC2 instances for compute-intensive workloads (C7 g) is also powered by Arm Graviton3 processors. Projects like European Mont-Blanc and U.S. DOE/NNSA Astra are further examples of Arm irruption in HPC. In parallel, over the last decade, the rapid improvement of genomic sequencing technologies and the exponential growth of sequencing data has placed a significant bottleneck on the computational side. While most genomics applications have been thoroughly tested and optimized for x86 systems, just a few are prepared to perform efficiently on Arm machines. Moreover, these applications do not exploit the newly introduced Scalable Vector Extensions (SVE). This paper presents GenArchBench, the first genome analysis benchmark suite targeting Arm architectures. We have selected computationally demanding kernels from the most widely used tools in genome data analysis and ported them to Arm-based A64FX and Graviton3 processors. Overall, the GenArch benchmark suite comprises 13 multi-core kernels from critical stages of widely-used genome analysis pipelines, including base-calling, read mapping, variant calling, and genome assembly. Our benchmark suite includes different input data sets per kernel (small and large), each with a corresponding regression test to verify the correctness of each execution automatically. Moreover, the porting features the usage of the novel Arm SVE instructions, algorithmic and code optimizations, and the exploitation of Arm-optimized libraries. We present the optimizations implemented in each kernel and a detailed performance evaluation and comparison of their performance on four different HPC machines (i.e., A64FX, Graviton3, Intel Xeon Skylake Platinum, and AMD EPYC Rome). Overall, the experimental evaluation shows that Graviton3 outperforms other machines on average. Moreover, we observed that the performance of the A64FX is significantly constrained by its small memory hierarchy and latencies. Additionally, as proof of concept, we study the performance of a production-ready tool that exploits two of the ported and optimized genomic kernels.

Read accessible full text

GenArchBench: A genomics benchmark suite for arm HPC processors

Author: López Villellas, Lorien,Langarita Benítez, Rubén,Badouh, Asaf,Soria Pardos, Víctor,Aguado Puig, Quim,López Paradís, Guillem,Doblas Font, Max,Setoain, Javier,Kim, Chulho,Ono, Makoto,Armejach Sanosa, Adrià,Marco Sola, Santiago,Alastruey Benedé, Jesús,Ibáñe
Publisher: Elsevier
Year: 2024
DOI: 10.1016/j.future.2024.03.050
Source: https://upcommons.upc.edu/bitstream/2117/407758/1/1-s2.0-S0167739X24001250-main.pdf
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
A ailable online 2 Ap il 2024
0167-739X/© 2024 The Au ho (s). Published by Else ie B.V. This is an open access a icle unde he CC BY-NC license (h p://c ea i ecommons.o g/licenses/by-
nc/4.0/).
Con en s lis s a ailable a ScienceDi ec
Fu u e Gene a ion Compu e Sys ems
jou nal homepage: www.else ie .com/loca e/ gcs
GenA chBench: A genomics benchma k sui e o a m HPC p ocesso s
Lo ién López-Villellasd,∗,1,Rubén Langa i a-Bení eza,Asa Badouha,Víc o So ia-Pa dosa,
Quim Aguado-Puigb,Guillem López-Pa adísa,Max Doblasa,Ja ie Se oaine,Chulho Kim ,
Mako o Onog,Ad ià A mejacha,c,San iago Ma co-Solaa,c,Jesús Alas uey-Benedéd,
Pablo Ibáñezd,Miquel Mo e óa,c
aBa celona Supe compu ing Cen e , Ba celona, Spain
bDepa men d’A qui ec u a de Compu ado s, Uni e si a Au ònoma de Ba celona, Ba celona, Spain
cDepa men d’A qui ec u a de Compu ado s, Uni e si a Poli ècnica de Ca alunya, Ba celona, Spain
dDepa amen o de In o má ica e Ingenie ía de Sis emas/A agón Ins i u e o Enginee ing Resea ch (I3A), Uni e sidad de Za agoza, Za agoza, Spain
eA m Resea ch, Camb idge, Uni ed Kingdom
Leno o Resea ch, Uni ed S a es
gLeno o In as uc u e Solu ions G oup, Uni ed S a es
ARTICLE INFO
Keywo ds:
Genomics
A m
High-pe o mance compu ing
Pa allel compu ing
Vec o compu ing
Pe o mance cha ac e iza ion
ABSTRACT
A m usage has subs an ially g own in he High-Pe o mance Compu ing (HPC) communi y. Japanese supe -
compu e Fugaku, powe ed by A m-based A64FX p ocesso s, held he op posi ion on he Top500 lis be ween
June 2020 and June 2022, cu en ly si ing in he ou h posi ion. The ecen ly eleased 7 h gene a ion o
Amazon EC2 ins ances o compu e-in ensi e wo kloads (C7 g) is also powe ed by A m G a i on3 p ocesso s.
P ojec s like Eu opean Mon -Blanc and U.S. DOE/NNSA As a a e u he examples o A m i up ion in HPC. In
pa allel, o e he las decade, he apid imp o emen o genomic sequencing echnologies and he exponen ial
g ow h o sequencing da a has placed a signi ican bo leneck on he compu a ional side. While mos genomics
applica ions ha e been ho oughly es ed and op imized o x86 sys ems, jus a ew a e p epa ed o pe o m
e icien ly on A m machines. Mo eo e , hese applica ions do no exploi he newly in oduced Scalable Vec o
Ex ensions (SVE).
This pape p esen s GenA chBench, he i s genome analysis benchma k sui e a ge ing A m a chi ec u es.
We ha e selec ed compu a ionally demanding ke nels om he mos widely used ools in genome da a
analysis and po ed hem o A m-based A64FX and G a i on3 p ocesso s. O e all, he GenA ch benchma k
sui e comp ises 13 mul i-co e ke nels om c i ical s ages o widely-used genome analysis pipelines, including
base-calling, ead mapping, a ian calling, and genome assembly. Ou benchma k sui e includes di e en
inpu da a se s pe ke nel (small and la ge), each wi h a co esponding eg ession es o e i y he
co ec ness o each execu ion au oma ically. Mo eo e , he po ing ea u es he usage o he no el A m SVE
ins uc ions, algo i hmic and code op imiza ions, and he exploi a ion o A m-op imized lib a ies. We p esen
he op imiza ions implemen ed in each ke nel and a de ailed pe o mance e alua ion and compa ison o
hei pe o mance on ou di e en HPC machines (i.e., A64FX, G a i on3, In el Xeon Skylake Pla inum, and
AMD EPYC Rome). O e all, he expe imen al e alua ion shows ha G a i on3 ou pe o ms o he machines
on a e age. Mo eo e , we obse ed ha he pe o mance o he A64FX is signi ican ly cons ained by i s
small memo y hie a chy and la encies. Addi ionally, as p oo o concep , we s udy he pe o mance o a
p oduc ion- eady ool ha exploi s wo o he po ed and op imized genomic ke nels.
1. In oduc ion
Fo many yea s, A m p ocesso s ha e domina ed he mobile de ice
segmen . Thei ene gy e iciency and license-based business model ha e
been he pilla s unde pinning his success.
∗Co espondence o: Depa amen o de In o má ica e Ingenie ía de Sis emas/A agón Ins i u e o Enginee ing Resea ch (I3A), Uni e sidad de Za agoza, Spain.
E-mail add ess: [email p o ec ed] (L. López-Villellas).
1The co esponding au ho conduc ed his wo k while a ilia ed wi h he Ba celona Supe compu ing Cen e .
In ecen yea s, A m has bu s on o he high-pe o mance compu ing
ma ke wi h in luen ial companies and conso iums ha ha e become
licensees, such as Fuji su, Amazon, Apple, NVIDIA, Samsung, AMD,
B oadcom, HUAWEI, and Qualcomm. Cu en ly, he A m-based Fuji su
h ps://doi.o g/10.1016/j. u u e.2024.03.050
Recei ed 3 No embe 2023; Recei ed in e ised o m 25 Ma ch 2024; Accep ed 31 Ma ch 2024
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
314
L. López-Villellas e al.
A64FX p ocesso powe s he Japanese supe compu e Fugaku, which
held he op posi ion on he Top500 lis be ween June 2020 and June
2022 and is cu en ly in he ou h posi ion. Mo eo e , Amazon has
been using A m p ocesso s o powe i s cloud compu ing pla o m
(AWS), s a ing in 2018 wi h he G a i on p ocesso . They ollowed
wi h he second gene a ion o G a i on in 2019 and he ecen ly
eleased G a i on3.
In he nea u u e, NVIDIA G ace CPUs and Ampe e se e s will
be leading u he e o s o b eak h ough A m in HPC. As a esul ,
la ge-scale compu ing in as uc u es, usually equipped wi h x86 and
IBM Powe p ocesso s, now ha e an addi ional compe i i e al e na i e.
Howe e , mos o he scien i ic code o HPC is no ully adap ed and
op imized o A m a chi ec u es.
O e he las decade, genome sequencing has become he co ne -
s one o genomics and mode n p ecision medicine. Due o he apid
imp o emen o sequencing echnologies, i is cu en ly possible o
sequence an indi idual’s genome in less han 24 h. This b eak h ough
has enabled e ec i e pe sonalized heal hca e, allowing he diagnosis
and ea men o diseases based on each pe son’s unique genomic
disposi ion [1]. Fu he mo e, genome sequencing has also been p o en
c ucial in cance s udies [2], d ug de elopmen [3], o COVID-19 ou -
b eak con ol [4]. In he pas 20 yea s, genome sequencing cos s ha e
d opped d ama ically and he amoun o sequencing da a p oduced
yea ly has inc eased exponen ially. Mo e no ably, his inc ease in da a
p oduc ion has ou pe o med he pace o Moo e’s law. As a esul , a
signi ican bo leneck in cu en genome sequencing analysis is placed
on he compu a ional side, execu ing compu a ional-in ensi e genomics
ools and pipelines.
Genome analysis pipelines ha e his o ically been designed o un
e icien ly on x86 a chi ec u es. Wi h he i up ion o A m-based HPC
se e s, adap ing and op imizing genomics ools o exploi HPC A m
a chi ec u es e ec i ely has become pa amoun . Fo ha , we ha e
selec ed 13 compu a ionally-demanding CPU ke nels om he mos
widely-used genomics ools, and we ha e included hem in a bench-
ma k sui e called GenA chBench. All he ke nels exploi mul i-co e
pa allelism and implemen common s ages om widely-used genome
analysis pipelines such as base-calling, ead mapping, a ian call-
ing, and de-no o assembly. Addi ionally, GenA chBench includes inpu
da ase s o each ke nel (i.e., a small da ase and a la ge da ase pe
ke nel) and hei co esponding ou pu s o be used as g ound u h. The
small da ase s ha e been sized o equi e single- h ead execu ion imes
no longe han a ew minu es ( o es ing pu poses); meanwhile, la ge
da ase s equi e se e al minu es ( o pe o mance e alua ion pu poses).
Fo con enience, we p o ide au oma ic eg ession es s o all he
ke nels o e i y he co ec ness o he ou pu s.
Fu he mo e, his wo k in oduces code adap a ions and op imiza-
ions o he genomics ke nels a ge ing A m HPC CPUs. GenA chBench
le e ages A m-speci ic HPC lib a ies (ca e ully op imized o A m p o-
cesso s) and p esen s algo i hmic and code op imiza ions o exploi he
a chi ec u e and esou ces o A m HPC machines. No ably, we ha e
op imized some ke nels by u ilizing he la es A m Scalable Vec o
Ex ensions (SVE) o le e age he po en ial o he la es A m HPC
p ocesso s.
In addi ion o he benchma k sui e po ing and op imiza ion, his
wo k p esen s a pe o mance cha ac e iza ion o GenA chBench on
ou HPC machines ( wo A m-based and wo x86-based nodes). The
expe imen al e alua ion compa es he pe o mance o an A64FX p o-
cesso , a G a i on3 p ocesso , an In el Xeon Skylake Pla inum 8160
p ocesso , and an AMD EPYC 7742 Rome p ocesso . This cha ac e i-
za ion includes he ke nels’ ins uc ion b eakdown, single- h ead and
mul i- h ead pe o mance e alua ions, a mic oa chi ec u e bo leneck
analysis, and an ene gy- o-solu ion s udy in he di e en p ocesso s.
Ul ima ely, we e alua e he pe o mance impac o hese op imiza ions
by in eg a ing wo o he accele a ed ke nels in a p oduc ion- eady ool
used in a my iad o genome analysis pipelines.
In summa y, his wo k makes he ollowing con ibu ions:
•We p esen GenA chBench, he i s benchma k sui e a ge ing
A m HPC a chi ec u es o genome analysis pipelines and ools.
The benchma k sui e is publicly a ailable a h ps://gi hub.com/
Lo ienLV/gena chbench/ eleases/ ag/1.0.0.
•We p opose HPC adap a ions and code op imiza ions applied o
GenA chBench’s ke nels o exploi he po en ial o A m HPC p o-
cesso s, le e aging A m-speci ic HPC lib a ies and A m Scalable
Vec o Ex ension (SVE).
•We pe o m a comp ehensi e pe o mance cha ac e iza ion o
GenA chBench in wo HPC A m p ocesso s (i.e., A64FX and
G a i on3). We compa e he pe o mance o A m agains wo
e e ence HPC x86 machines.
2. Backg ound
Genome da a analysis pipelines comp ise mul iple s ages and com-
pu a ional ools, om sequencing biological samples o de i ing mean-
ing ul da a analysis esul s o scien is s and heal hca e p o essionals.
This sec ion in oduces he main sequencing echnologies, pipelines,
and ools used in common genome analysis (Fig. 1 shows a succinc
g aphic summa y).
2.1. Sequencing echnologies
Be o e any compu a ional analysis can be pe o med, biological
DNA samples mus be con e ed o digi al da a. This p ocess is pe -
o med by he sequencing machines (Fig. 1-1), and, despi e he ema k-
able ad ances in he las decades, hese machines a e s ill unable o
ead a comple e DNA molecule om end o end. Ins ead, sequencing
machines allow eading ela i ely small chunks o DNA, called eads
o agmen s, om andom loca ions wi hin he dono ’s DNA genome.
A e wa ds, sequenced eads mus be jigsaw oge he o econs uc o
eassemble he o iginal dono ’s genome.
Sequencing machines a e commonly ca ego ized in o h ee gene a-
ions based on hei echnological ad ancemen s. The i s sequencing
echnologies (Sange e al. [5] and Maxam e al. [6]) we e de eloped
in 1977 and used o sequence he i s d a o he human genome in
2000 [7]. Since hen, sequencing echnologies ha e e ol ed quickly,
simpli ying he sequencing p ocess and inc easing he da a-p oduc ion
h oughpu . In he mid-2000s, second-gene a ion echnologies [8] we e
in oduced and soon eplaced i s -gene a ion echnologies. Second-
gene a ion echnology can gene a e ixed-leng h sequences o 100–300
bps a a h oughpu o ens o gigaby es pe hou and wi h a low eading
e o a e (0.1% o he ead leng h). A p esen , Illumina domina es he
ma ke o second-gene a ion sequencing machines. Recen ly in oduced
hi d-gene a ion echnologies, known as long- ead sequencing, can ead
a iable-leng h sequences o conside able leng h (i.e., ens o kilo base-
pai s) a he expense o lowe p oduc ion h oughpu (less han 10
Gb/hou ) and highe eading e o a e (0.1%–10% o he ead leng h).
Paci ic Biosciences (PacBio) and Ox o d Nanopo e Technologies (ONT)
a e he mos no able manu ac u e s o hi d-gene a ion sequencing
echnologies.
2.2. Genome da a analysis pipelines and ools
Be o e any p ocessing can be pe o med, he sequencing machines’
aw signals mus be ans o med in o sequences o nucleo ides (A,
C, G, T). This p ocess is called basecalling (Fig. 1-2). Typically, a
specialized basecalling ool is used o pe o m his p ocess ailo ed o
each sequencing echnology. Fo ins ance, Boni o [9] and Guppy [10]
a e wo o he mos widely-used ools o basecalling Ox o d Nanopo e’s
aw-signal ou pu .
Once he sequences o nucleo ides a e decoded, sequenced eads
mus be p ocessed and analyzed o de i e meaning ul biological in-
sigh s. Al hough many di e en genome analyses can be pe o med
using sequenced da a, mos analyses begin wi h ei he genome ese-
quencing (1-3.a) o genome assembly (1-3.b). Bo h analyses seek o
econs uc he sample’s genome by pu ing oge he all he sequenced
eads.
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
315
L. López-Villellas e al.
Fig. 1. Wo k low diag am o common genome analysis pipelines. Going om (1) sequencing, h ough (2) basecalling, o (3.a) genome esequencing o (3.b) and genome assembly.
The igu e shows he di e en compu a ional ke nels used wi hin each s age o ool.
2.2.1. Genome esequencing
The mos common app oach o econs uc ing he sample’s genome
is by esequencing and in ol es econs uc ing he sample’s genome
using a p e iously known e e ence genome. Fo ha , each sequenced
ead is loca ed and ma ched o he mos likely o igina ing posi ion
in he e e ence genome, allowing small di e ences (e.g., misma ches,
inse ions, and dele ions). This p ocesses is called ead mapping (Fig. 1-
3.a.1) and i is implemen ed by many ools like BWA-MEM2 [11,12],
Minimap2 [13], Bow ie2 [14,15], and GEM [16]. Read mapping is one
o he mos compu a ionally expensi e s eps in all genome sequence
analyses. Consequen ly, ead mapping has been ex ensi ely s udied and
op imized.
Mos sequence mappe s a e based on he seed-chain-ex end ech-
nique. This echnique implemen s h ee algo i hmic s eps o swi ly
loca e and align a sequence wi h a e e ence genome. Du ing he i s
s ep, known as seeding (Fig. 1-3.a.1.1), he mappe sea ches small
subsequences o he eads (seeds) in he e e ence le e aging an index
s uc u e. The mos widely-used indexes used o seeding a e FM-
Index [17] and hash- ables [18,19]. Seeding educes he po en ial num-
be o loca ions in he e e ence whe e a sequence can ma ch, dec eas-
ing he amoun o wo k pe o med in subsequen s eps. A e wa ds, a
chaining s ep (Fig. 1-3.a.1.2) is pe o med o educe u he he lis o
possible ma ching loca ions in he e e ence. Du ing he chaining s ep,
all he mapped seeds a e p ocessed o ind a colinea chain o seeds
ha can po en ially ma ch he inpu sequence. Finally, du ing he ex-
ensión o alignmen s ep (Fig. 1-3.a.1.3), he inpu sequence is aligned
agains he candida e loca ion in he e e ence genome, disco e ing he
di e ences be ween he dono ’s sequence and he e e ence genome.
Usually, a dynamic p og amming-based algo i hm, such as Needleman–
Wunsch [20] o Smi h–Wa e man–Go oh [21,22], is used o compu e
he alignmen .
A e sequence mapping, once he eads a e loca ed in he e e -
ence genome, a a ian calling algo i hm (Fig. 1-3.a.2) de e mines he
a ian s and mu a ions be ween he dono ’s genome and he e e -
ence genome. These a ia ions p o ide c ucial insigh s in o he ge-
ne ic makeup o he sequenced indi idual, po en ially e ealing ge-
ne ic a ia ions ha may be associa ed wi h diseases and heal h con-
di ions. No able examples o widely-used a ian calle s a e GATK
Haplo ype-Calle [23], Pla ypus [24], Clai [25,26], DeepVa ian [27]
and Medaka [28].
2.2.2. Genome assembly
Despi e he simplici y and e ec i eness o genome esequencing,
he e is s ill a lack o high-quali y e e ence genomes o many species.
In hose si ua ions, genome de-no o assembly (Fig. 1-3.b) is used o
econs uc he dono ’s genome om sc a ch jigsawing he sequenced
eads oge he .
Mos popula de-no o assembly me hods ely on de B uijn g aphs.
Fo a gi en se o sequences, i s co esponding de B uijn g aph con ains
a node pe each sequence’s k-me (i.e., sub-s ing o leng h 𝑘nu-
cleo ides) and an edge ha connec s adjacen and o e lapping k-me s.
Be o e cons uc ing he de B uijn g aph o a se o inpu sequences, he
numbe o unique k-me s in he eads is coun ed (Fig. 1-3.b.1) o p une
he leas equen ones (likely a i ac s o he sequencing p ocess).
A e wa ds, he de B uijn g aph is cons uc ed (Fig. 1-3.b.2). Then,
he consensus sequence is de i ed using mul iple sequence alignmen
(MSA) algo i hms (Fig. 1-3) and he cons uc ed de B uijn g aph.
No able examples o de B uijn g aph based assemble s a e Flye [29],
Canu [30], and Racon [31].
2.2.3. Me agenomics
Beyond genome esequencing, a ian calling, and de-no o assem-
bly, many p e iously desc ibed analysis s eps and ools can be ound
in o he genome analysis pipelines. This is he case o many me age-
nomics analysis pipelines. Me agenomics pipelines seek o analyze
genomic in o ma ion om mixed mic obial communi ies, p o iding
insigh s in o he di e si y, in e ac ions and unc ion o mic oo gan-
isms p esen in an en i onmen al sample. Me agenomics analyses a e
pe o med using ools such as Cen i uge [32], RawMap [33], UN-
CALLED [34], ReadFish [35], K aken2 [36] and Cla k [37]. These
ools employ k-me coun ing (Fig. 1-3.b.1) and seeding echniques
(Fig. 1-3.a.1.1) o hei analysis. Mo eo e , a ian calle s like GATK
Haplo ype Calle [23] and Pla ypus [24] a e used o cons uc De B uijn
g aphs (Fig. 1-3.b.2) and co ec a i ac s p oduced du ing he map-
ping p ocess (Fig. 1-3.a.1). Fu he mo e, he chaining p ocess (Fig. 1-
3.a.1.2) is also u ilized o genome assembly when using al e na i e
app oaches based on de B uijn g aphs [38].
3. GenA ch benchma k sui e
The GenA ch benchma k sui e comp ises 13 mul i h eaded CPU
ke nels de i ed om he mos widely used genomics ools and co e s
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
316
L. López-Villellas e al.
he mos impo an genome sequencing s eps. I includes en ke nels
om he GenomicsBench [39] benchma k sui e and h ee addi ional
ke nels: he Bi -Pa allel Mye s algo i hm [40] (BPM), he Wa e on
Alignmen algo i hm [41] (WFA), and FAST-CHAIN [42]. BPM and
WFA complemen he sequence alignmen ke nels o GenomicsBench o
be e cap u e con empo a y ends. Addi ionally, FAST-CHAIN [42] is
a ecen ec o -enabled eimplemen a ion o he CHAIN ke nel p esen
in GenomicsBench, which allows us o u he explo e he capabili ies
o SVE.
Addi ionally, GenA chBench includes inpu da ase s o each ke -
nel (i.e., a small da ase and a la ge da ase pe ke nel) and hei
co esponding ou pu s o be used as g ound u h. The small da ase s
ha e been sized o equi e single- h ead execu ion imes no longe
han a ew minu es ( o es ing pu poses); meanwhile, la ge da ase s
equi e se e al minu es ( o pe o mance e alua ion pu poses). Fo
con enience, we p o ide au oma ic eg ession es s o all he ke nels
o e i y he co ec ness o he ou pu s.
Al hough some ke nels included in GenA chBench can exploi he
capabili ies o mode n GPUs, his esea ch ocuses on po ing, accel-
e a ing, and e alua ing he pe o mance o genomics ke nels in A m
p ocesso s. Mo eo e , he A m-sys ems e alua ed in his wo k (A64FX
and G a i on3) a e no equipped wi h GPUs.
The ollowing ex p esen s GenA chBench’s ke nels, b ie ly desc ib-
ing i s unc ionali y, which ools use hem, and a desc ip ion o hei
usage and inpu s.
Adap i e Banded Signal o E en Alignmen (ABEA): ABEA is
a dynamic p og amming algo i hm ha compa es aw nanopo e sig-
nals om ONT sequencing machines o a e e ence genome sequence.
ABEA’s implemen a ion is based on he Suzuki–Kasaha a (SK) [43] al-
go i hm. This s ep is pe o med in some ools, such as Nanopolish [44],
o co ec e o s p oduced in he basecalling p ocess (Fig. 1-2). Fo
GenA chBench, we ha e used he CPU implemen a ion o 5c [45],
a e sion o ABEA based on Nanopolish’s, op imized o bo h CPU-
only and hyb id CPU/GPU execu ions. This implemen a ion o ABEA
exploi s coa se-g ain mul i- h eading by di iding he aw signals o
he inpu be ween he a ailable co es. Since he signals a e no o
egula size, 5c implemen s wo k-s ealing o imp o e load balance.
The small and la ge inpu s comp ise 1K and 10K aw FAST5 (ONT)
eads om ch omosome 22 o NA12878 and GRCh38 as he e e ence
genome [46].
Bi -Pa allel Mye s (BPM): BPM [40] is a dynamic p og amming
algo i hm ha inds all loca ions a que y s ing o size 𝑚ma ches a
e e ence s ing o size 𝑛wi h 𝑘o ewe di e ences (Fig. 1-3.a.1.3). I
compu es he app oxima e s ing ma ching o wo s ings in 𝑂(𝑚𝑛∕𝑤)
ime, whe e 𝑤is he wo d size o he machine. BPM is used in ead map-
ping ools, such as GEM-Mappe [16], Edlib [47], G aphAligne [48]
o Hobbes [49]. Fo GenA chBench, we ha e used an in-house imple-
men a ion o he algo i hm ha exploi s mul i- h eading by assigning
di e en pai s o s ings o di e en h eads. The small and la ge
inpu s comp ise 100K and 10M sequence pai s om human sample
SRR7733443 downloaded om he sequence ead a chi e [50].
Banded Smi h–Wa e man (BSW): The Smi h–Wa e man algo i hm
[21] is a dynamic p og amming algo i hm ha compu es he local
sequence alignmen o wo sequences o leng h 𝑚and 𝑛, espec i ely,
in 𝑂(𝑚𝑛) ime and space. A banded e sion o Smi h–Wa e man [51]
is used o align sequences wi h a maximum o 𝑤inse ions/dele ions,
educing he ime and space complexi y o 𝑂(𝑤𝑛)(Fig. 1-3.a.1.3).
BSW is used in a ian disco e y ools such as GATK [23], and in
sequence alignmen so wa e like BWA-MEM [11,12]. Fo GenA ch-
Bench, we ha e used BWA-MEM2’s x86- ec o ized implemen a ion o
BSW. In o de o exploi mul i- h eading, he se o pai s o s ings o
align is dynamically di ided be ween p ocesso s. The small and la ge
inpu s comp ise 100K and 10M sequence pai s om human sample
SRR7733443 [50].
Seed Chaining (CHAIN): Gi en he se o seeds om a DNA se-
quence ( ead) mapped o ano he sequence, such as he e e ence
genome, he chaining s ep (Fig. 1-3.a.1.2) aims o ind a chain o
colinea seeds. This is a ime-consuming s ep pe o med by alignmen
ools, such as Minimap2, and by de-no o assemble s like Flye [29]
o Canu [30]. We ha e used he implemen a ion o CHAIN ound in
GenomicsBench ha ex ends Minimap2’s o exploi in e - ask pa al-
lelism ac oss eads. The small and la ge inpu s comp ise he seeds om
1K, and 10K eads o Pacbio’s Caeno habdi is elegans wo m sequence
da a [52].
SIMD Seed Chaining (FAST-CHAIN): The p e iously p esen ed
implemen a ion o he CHAIN algo i hm u ilizes heu is ics o s op
execu ing when he esul is su icien ly good. This speedups execu ion
a he cos o accu acy, and i hinde s he ec o iza ion o he ke nel.
FAST-CHAIN [42] is an x86- ec o ized e sion o CHAIN ha emo es
he heu is ics o exploi SIMD compu a ion. As a esul , FAST-CHAIN
ou pu s accu a e esul s and p esen s pe o mance gains compa ed o
CHAIN. FAST-CHAIN uses he same inpu s as CHAIN.
De B uijn G aph Cons uc ion (DBG): The De B uijn g aph (DBG)
o an inpu se o eads is used o ep esen he o e laps be ween he
sub-s ings o leng h 𝑘(k-me s) ound in he inpu (Fig. 1-3.b.2). Each
node o he g aph ep esen s a k-me and he edges connec adjacen
k-me s in he inpu se . The cons uc ion o hese g aphs is a ime-
consuming s ep in de-no o assemble s like Flye [29], Canu [30] o
Racon [31], and in a ian calle s such as GATK [23] and Pla ypus [24].
Fo GenA chBench, we ha e used he DBG cons uc ion o Pla ypus,
which exploi s pa allelism by assigning di e en egions o he inpu
o di e en h eads. Bo h inpu s employ ch omosome 22 o BWA-MEM
aligned eco ds om he Pla inum Genomes da ase [53]. The small
inpu uses bases 16M-16.5M, while he la ge inpu uses he en i e
ch omosome.
FM-Index Sea ch (FMI): The FM-index is a comp essed sub-s ing
index based on he Bu ows–Wheele ans o m [54]. Gi en a sub-s ing
𝑠, FM-index can be used o ind he loca ion o 𝑠in he e e ence
genome in 𝑂(|𝑠|) ime, whe e |𝑠|is he leng h o he sub-s ing (Fig. 1-
3.a.1.1). The FM-index da a s uc u e is used in sequence alignmen
ools such as BWA-MEM [11,12] o Bow ie2 [15], and in me agenomic
classi ica ion so wa e like Cen i uge [32]. Fo GenA chBench, we
ha e used he supe -maximal exac ma ch ke nel o BWA-MEM2, which
u ilizes he FM-Index s uc u e. This ke nel exploi s pa allelism by
dynamically assigning ba ches o eads among h eads. The small and
la ge inpu s comp ise 1M and 10M pai s o 151 bases om human
sample SRR7733443 [50].
K-me Coun ing (KMER-CNT): K-me coun ing aims o coun he
numbe o occu ences o each k-me in an inpu sequence (Fig. 1-
3.b.1). This ask is pe o med in de-no o assemble s such as Flye [29]
o Canu [30] and in me agenomics classi ica ion so wa e like Cla k
[37]. Addi ionally, no e ha he unc ionali y o KMER-CNT is e y
simila o accessing la ge lookup ables, as done in s a e-o - he-a map-
pe s like Minimap2 [13]. Fo GenA chBench, we ha e used he k-me
coun ing ke nel o Flye. This implemen a ion di ides he inpu - eads
be ween h eads and elies on he h ead-sa e hash-map implemen a-
ion o Libcuckoo lib a y [55] o concu en ly inc ease he numbe o
indi idual k-me s shown by each h ead. The small and la ge inpu s
comp ise 1K and 50K Esche ichia coli Ox o d Nanopo e eads sequenced
by Loman Labs [56].
Neu al Ne wo k-based Base Calling (NN-BASE): ONT sequencing
machines moni o changes in an elec ical cu en as single s ands o
DNA o RNA pass h ough a p o ein nanopo e. These changes in he
elec ical cu en a e hen con e ed o a sequence o nucleo ide bases
in he basecalling p ocess (Fig. 1-2). The analog signal ine i ably con-
ains ambigui ies due o noise o measu emen e o s. Some basecalle s,
such as Guppy [10] and Boni o [9], ely on neu al ne wo ks o sol e
hese ambigui ies, de e mining he mos likely obse ed nucleo ide
in each pa o he elec ical cu en . Fo GenA chBench, we ha e
used Boni o’s deep-lea ning base-calle (NN-BASE), which depends on
he PyTo ch lib a y [57]. Boni o spli s he inpu signal in o smalle
chunks o egula size and eeds hem o a PyTo ch neu al ne wo k ha
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
317
L. López-Villellas e al.
Table 1
Cha ac e is ics o e iew o he expe imen al se up.
A64FX G a i on3 SKX Rome
Co es 4 ×12 (+ 4 assis an ) 64 2 ×24 64
SMT No No Disabled Disabled
F equency 2.2 GHz (s a ic) 2.6 GHz 1–2.1 GHz (dynamic) 1.5–2.25 GHz (dynamic)
Max. powe 120 W N/A 2 ×150 W 225 W
Mem. capaci y 4 ×8 GB 8 ×16 GB 2 ×6×8 GB 16 ×64 GB
Mem. echnology on-package HBM2 o -package DDR5 4800 MHz o -package DDR4 2667 MHz o -package DDR4 3200 MHz
Peak bandwid h 4 ×256 GB/s 300 GB/s 2 ×120 GB/s 204.8 GB/s
L1i 64 KB (4-way) 64 KB 32 KB (8-way) 32 KB (8-way)
L1d 64 KB (4-way) 64 KB 32 KB (8-way) 32 KB (8-way)
L2 – 1 MB 1 MB (16-way) 512 KB (8-way)
LLC 4 ×8 MB (16-way) 32 MB 2 ×33 MB (11-way) 16 ×16 MB (16-way)
Vec o ex ension NEON/SVE 512 bi s NEON/SVE 256 bi s SSE/AVX2/AVX512 SSE/AVX2
in e nally exploi s mul i- h eading. The small and la ge inpu s comp ise
1 and 10 aw FAST5 eads om ch omosome 20 o NA12878, ob ained
om he Nanopo e WGS Conso ium [46].
Neu al Ne wo k-based Va ian Calling (NN-VARIANT): Va ian
calling is he p ocess o de ec ing he di e ences ( a ian s o mu a ions)
be ween he aligned eads and he e e ence genome (Fig. 1-3.a.2). This
is a cos ly p ocess pe o med by s a is ics-based a ian calle s, such as
GATK Haplo ypeCalle [23] o Pla ypus [24], and deep-lea ning a ian
calle s, such as Clai [25,26], DeepVa ian [27] o Medaka [28]. Fo
GenA chBench, we ha e used he second gene a ion o Clai a ian
calle (Clai 3), based on he Tenso Flow amewo k [58]. Clai 3 ex-
ploi s pa allelism by di iding he inpu in o egula -size chunks, and
each o hese chunks is p ocessed by one h ead using Tenso Flow. Ou
small and la ge inpu s comp ise 100K and 10M e e ence posi ions,
espec i ely, o ch omosome 20 o HG002 om NITS’s Genome in
a Bo le (GIAB) p ojec [59]. We a e using Clai 3’s ONT p e- ained
model 941_p om_hac_g360+g422 [60].
Pileup Coun ing (PILEUP): Gi en he alignmen da a o a se o
aligned eads o a egion o a e e ence genome, usually a SAM o BAM
ile [61], pileup coun ing is he p ocess o summa izing he base-pai in-
o ma ion a each ch omosomal posi ion. This summa y, called pileup,
is cus oma y he inpu o long- ead neu al ne wo k a ian calle s such
as Clai [25,26] o Meda aka [28] (Fig. 1-3.a.2). Fo GenA chBench
we ha e used he pileup coun ing implemen a ion o Medaka, which
exploi s mul i- h ead pa allelism by dis ibu ing 100 kilobase egions
o he e e ence genome be ween h eads. The small inpu comp ises
bases 1-1499707 o he S aphylococcus au eus genome [10], and he
la ge inpu comp ises bases 1-1412827 o ch omosome 20 o sample
HG002 [59].
Pa ial-O de Alignmen (POA): The cons uc ion o an o e lap
g aph om a se o eads leads o an app oxima e ep esen a ion o he
o iginal sample’s genome. To de e mine he consensus genome o he
sample, he alignmen o all he eads agains each o he is pe o med
in a p ocess called mul iple sequence alignmen (MSA) (Fig. 1-3.b.3).
The Pa ial O de ed Alignmen (POA) algo i hm [62] compu es he
MSA o all sequences by inc emen ally cons uc ing a pa ially-o de
g aph aligning new sequences o i using a dynamic p og amming
algo i hm such as Smi h–Wa e man [21] o Needleman–Wunsch [20].
The mul iple alignmen sequence (consensus sequence) is in e ed om
he g aph by using he Hea ies Bundle algo i hm [63]. POA is used
in so wa e packages such as Nanopolish [44] o Racon [31]. Fo
Gena chBench we ha e used he SIMD-op imized e sion o POA o he
SPOA lib a y [64]. SPOA exploi s mul i- h eading by compu ing he
pa ially-o de ed g aph o mul iple se s o sequences in pa allel. The
small and la ge inpu s comp ise 1K and 6K se s o mul iple sequences
aligned o a e e ence genome, each con aining be ween 5 and 115
sequences. This da a comes om Minimap2’s polishing s ep o he
Flye-assembled S aphylococcus Au eus genome [10].
Wa e on Alignmen (WFA): The wa e on alignmen algo i hm
(WFA) [41] is a pai wise alignmen algo i hm (Fig. 1-3.a.1.3) ha akes
ad an age o homologous egions be ween he sequences o accele a e
he alignmen p ocess. As opposed o adi ional dynamic p og amming
algo i hms ha un in quad a ic ime, WFA ime complexi y is 𝑂(𝑛𝑠),
p opo ional o he ead leng h 𝑛and he alignmen sco e 𝑠, using
𝑂(𝑠2)memo y. The wa e on algo i hm is used in ools such as w -
mash [65], Ancho Wa e [66] o Ances alClus [67]. GenA chBench
uses a cus om mul i- h ead implemen a ion o he algo i hm, in which
each h ead wo ks in he alignmen o a pai o s ings. The small and
la ge inpu s comp ise 100K and 1M sequence pai s om human sample
SRR7733443 [50].
4. Expe imen al se up
Ou expe imen al se up consis s o wo A m and wo x86 HPC
sys ems: a compu e node ea u ing an A m-A64FX p ocesso (A64FX),
a c7 g.16xla ge Amazon-EC2 ins ance (G a i on3), a sys em wi h wo
x86-64 In el Xeon Skylake Pla inum 8160 (SKX), and a compu e node
wi h one x86-64 AMD EPYC 7742 Rome p ocesso (Rome). Table 1
p esen s an o e iew o he main cha ac e is ics o he ou sys ems.
In e ms o compu ing co es, he A64FX is based on ou Non-
Uni o m Memo y Access (NUMA) domains wi hin he chip, also e-
e ed o as co e memo y g oups (CMG). Each NUMA domain has 12
co es, plus one assis ance co e no used o gene al compu ing ( unning
daemons, I/O, asynch onous MPI, e c.). In o al, he A64FX implemen s
48 compu ing co es. G a i on3 implemen s 64 co es in a single NUMA
domain. The AMD Rome CPU comp ises 8 co e chiple s, known as co e
cache dies (CCD), and a cen al I/O die ha con ols all he I/O and
memo y unc ions o he chip. A CCD has wo co e complex (CCX)
clus e s, each wi h 4 co es. Any pai o CCDs can communica e h ough
he I/O die. SKX con ains wo NUMA chips, each wi h 24 physical co es.
Rega ding ope a ional equency, G a i on3 p esen s he highes
maximum equency among he sys ems wi h 2.6 GHz. The o he h ee
sys ems’ maximum equency is e y simila , anging be ween 2.1 and
2.25 GHz. Bo h x86 sys ems dynamically adjus hei equency based
on hei load. Addi ionally, SKX educes i s equency when execu ing
AVX/AVX512 ins uc ions. In con as , he A64FX ope a es a a ixed
equency se o 2.2 GHz. The e is no public in o ma ion abou adap i e
equency ope a ion on G a i on3.
Wi h espec o SIMD ex ensions, he A64FX is he i s CPU o
implemen he A m 8.2-A Scalable Vec o Ex ension (SVE) [68]. One
o SVE’s main ea u es is ha i is Vec o Leng h Agnos ic (VLA);
ha is, he same bina y wo ks on a chi ec u es implemen ing ec o
egis e s o di e en leng hs anging om 128 o 2048 bi s. The A64FX
implemen s 512 bi s SVE egis e s. G a i on3 also implemen s SVE,
wi h a ec o leng h o 256 bi s. The A64FX and G a i on3 also suppo
he A m Neon SIMD ex ension, a non-VLA SIMD ISA ha wo ks wi h
128-bi ec o s. Bo h x86 sys ems implemen he SSE and AVX2 SIMD
ex ensions, wi h a ec o leng h o 128 and 256 bi s, espec i ely. The
SKX also suppo s he AVX512 ex ension, wi h a ec o leng h o 512
bi s. None o he x86 SIMD ex ensions a e VLA.
Conce ning main memo y, each A64FX’s NUMA domain has i s own
local on-chip 8 GB HBM2 main memo y and can access he o he h ee
NUMA domains’ local memo ies ia a ing bus. G a i on3 is connec ed
o 8 ×16 GB DDR5 channels, o a o al o 128 GB o memo y. Each

Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
318
L. López-Villellas e al.
Table 2
Load- o-use memo y la encies in nanoseconds o he expe imen al se up.
A64FX G a i on3 SKX Rome
L1 2.3–5 1.5 1.9 1.8
L2 – 4.6 6.7 3.5
LLC 16.8–21.4 33.1 25.1 13.0
Main Mem. Local 118.2–126.4 153.5 86.2 121.5
Main Mem. Remo e 187.7–242.3 – 144.0 –
chip o he SKX is connec ed o 6 ×8 GB DDR4 local channels and can
access he o he chip’s local memo y. The Rome CPU is connec ed o
16 ×64 GB DDR4 channels, o aling 1 TB o memo y.
The cache hie a chy o ganiza ion o he p ocesso s is ela i ely
di e en . Bo h A m machines ha e wo 64 KB p i a e L1 caches pe
co e (ins uc ions and da a), while he x86 CPUs ea u e wo 32 KB
p i a e L1s pe co e. G a i on3 and SKX include one p i a e 1MB
L2 cache pe co e, and Rome has one 512 KB p i a e L2 pe co e.
The A64FX has one 8 MB las -le el cache (LLC) pe NUMA domain,
G a i on3 includes one 32 MB LLC, SKX has wo 33 MB LLCs (one pe
NUMA domain), and Rome includes one 16 MB LLC pe each 4-co e
CCX.
Conce ning memo y bandwid h, he A64FX is designed o achie e
good pe o mance execu ing high memo y bandwid h-demanding ap-
plica ions. The peak bandwid h o his chip (4 ×256 GB/s) is nea ly
3.5 imes highe han he peak bandwid h o G a i on3 (300 GB/s), he
second sys em among he s udied in e ms o memo y h oughpu . I is
ollowed by SKX, eaching up o 120 GB/s pe chip (240 GB/s in o al),
and Rome holds he las posi ion wi h a peak bandwid h o 204.8 GB/s.
Table 2 p esen s he memo y access la encies o each le el o
he memo y hie a chy o all machines. All la encies on Rome and
G a i on3 and la encies o emo e memo ies on he A64FX ha e been
measu ed using he LMbench benchma k [69]. La encies o cache and
local memo y on he A64FX ha e been ex ac ed om he mic o-
a chi ec u e manual o he CPU. La encies on SKX ha e been measu ed
using In el Memo y La ency Checke . The numbe o cycles o access
he A64FX caches depends on he ype o ins uc ion: scala , loa ing-
poin , sho SIMD, and la ge SIMD. The la encies o access he L1 on
he sys ems ange om 1.5 ns (G a i on3) o 5 ns (la ge SIMD access
on he A64FX). E en hough scala accesses on he A64FX a e as e
(2.3 ns), i s ill p esen s he highes L1 access la ency. As p esen ed
p e iously, he A64FX only implemen s wo le els o caches (L1 and
LLC). The L2 access la encies o he o he sys ems ange be ween 3.5 ns
(Rome) o 6.7 ns (SKX). Rome p esen s he as es access o i s LLC
(13 ns), closely ollowed by he A64FX (16.8 ns o scala access and
21.4 ns o la ge SIMD access). The LLC access la ency on he SKX and
G a i on3 is 25.1 and 33.1 ns, espec i ely. SKX p esen s he as es
access la ency o local main memo y (86.2 ns), ollowed by he A64FX
and Rome, wi h simila la encies (∼120 ns). G a i on3 has he highes
local memo y access la ency, as expec ed om cu en DDR5 SDRAMs.
Accessing emo e main memo ies in he A64FX akes be ween 187.7 ns
(nea - emo e memo y) and 242.3 ns ( a - emo e memo y). Accessing
he o he chip’s main memo y on he SKX machine akes 144 ns, 23%
as e han A64FX’s bes case.
The ou -o -o de esou ces o he expe imen al se up a e p esen ed
in Table 3. We assume ha G a i on3 implemen s he same esou ces as
Neo e se V1 o non-publicly a ailable da a (ma ked wi h *). Una ail-
able da a o nei he G a i on3 no Neo e se V1 is ep esen ed as N/A.
The A64FX is igh in ou -o -o de esou ces compa ed wi h he o he
h ee p ocesso s. The SKX and Rome ha e a simila numbe o physical
egis e s, almos doubling he numbe o gene al-pu pose egis e s o
he A64FX (180 s. 96) and implemen ing 30% mo e SIMD/FP egis e s
han he A64FX (160 s. 128). The A64FX can issue up o 7 mic o-
ope a ions (𝜇OP) pe cycle, G a i on3 can issue up o 15, 8 o SKX, and
11 o Rome. The A64FX and SKX a e capable o commi ing 4 mic o-
ope a ions pe cycle. Howe e , SKX can me ge wo mic o-ope a ions
Table 3
Ou -o -o de esou ces o he expe imen al se up.
A64FX G a i on3 SKX Rome
Gene al egis e s 96 N/A 180 180
SIMD/FP egis e s 128 N/A 168 160
Issue wid h 7 (𝜇OP) 15 (𝜇OP) 8 (𝜇OP) 11 (𝜇OP)
Commi wid h 4 (𝜇OP) N/A 4-8 (𝜇OP) 8 (MOP)
ROB (en ies) 128 256* 224 224
LB (en ies) 40 85* 72 44
SB (en ies) 24 90* 56 48
RS (en ies) 2 ×20 +
2×10 + 19
N/A 97 4 ×16 +
28 + 36
*Neo e se V1 CPU de aul s.
in o one used mic o-ope a ion, inc easing i s heo e ical commi a e
o 8 mic o-ope a ions. Rome can commi up o 8 mac o-ope a ions
(MOP) – i.e., ALU, memo y, o me ged ALU/memo y ope a ion – pe
cycle. The eo de bu e (ROB) o G a i on3 (256 en ies) is wice as
big as he A64FX’s (128 en ies). SKX and Rome ha e an iden ical-size
ROB (224 en ies). The sizes o he load bu e s (LB) and s o e bu e s
(SB) o he CPUs a e ela i ely di e en . G a i on3 and SKX implemen
he la ges LB, wi h 85 and 72 en ies, espec i ely. The LB o he
A64FX has 40 en ies, and Rome implemen s a 44-en y LB. Simila ly,
G a i on3 and SKX ha e he la ges SB (90 and 56 en ies, espec i ely).
The A64FX implemen s a 24-en y SB, hal he size o Rome’s. Addi ion-
ally, a s o e ins uc ion on he A64FX occupies one en y in bo h he
load and he s o e bu e . While SKX implemen s a uni ied ese a ion
s a ion (RS) wi h 97 en ies, bo h he A64FX and Rome ha e se e al
smalle RS. The A64FX di ides i s ese a ion s a ion in o 2 ×20 en ies
o 2 in ege , loa ing-poin , and SIMD pipelines, 2 ×10 en ies o 2
add ess calcula ion pipelines, and 19 en ies o he b anch pipeline.
Rome’s ese a ion s a ion has 4 ×16 en ies o 4 in ege pipelines
(scala +SIMD), 28 en ies o 3 add ess calcula ion pipelines, and 36
en ies o 4 loa ing-poin pipelines (scala +SIMD).
5. A m po ing o genomics ke nels
Mos ke nels p esen ed in Sec ion 3 a ge x86 a chi ec u es and
ha e no been ex ensi ely es ed no op imized o A m machines.
Thus, i was expec ed ha some ke nels could un in o ailu es and
e en gene a e inco ec esul s. To e i y he execu ion o he ke nels,
we used he SKX sys em o compu e he co ec ou pu o all ke nels
and inpu s (i.e., g ound u h).
Fo ou expe imen s, we used he GNU compile (GCC) on G a i on3
( 11.2.0), SKX ( 10.1.0), and Rome ( 10.2.0). On he A64FX, we used
GCC ( 10.2.0) and he Fuji su Compile (FCC) ( 4.2.0b). Fo mos
ke nels, FCC-compiled bina ies exhibi ed be e pe o mance. The Fu-
ji su Compile implemen s wo compila ion modes: a adi ional mode
(T ad) based on compile s o ea lie sys ems and a Clang mode based
on Clang/LLVM. In all cases, we ob ained be e execu ion imes when
compiling wi h FCC’s Clang mode. We lacked FCC-compiled e sions
o key-op imized Py hon lib a ies. Fo hese easons, all he esul s
p esen ed in his documen o he A64FX ha e been ob ained using
he Clang mode o FCC, excluding he wo Py hon ke nels (NN-BASE
and NN-VARIANT), whose lib a ies we e compiled using GCC.
We compile all ke nels wi h a leas -O2 op imiza ion le el and
enable CPU-speci ic op imiza ions: -ma ch=a m 8-a+s e on he
A64FX, -mcpu=na i e on G a i on3 and -ma ch=na i e on SKX
and Rome. Enabling CPU-speci ic op imiza ions in ABEA and POA
esul ed in inco ec execu ions, p obably due o p og amming e o s
in he o iginal sou ce code. The e o e, such op imiza ions a e no used
o hese wo ke nels.
A e pe o ming he app op ia e modi ica ions o he ke nels so
all o hem success ully execu e on A m, we applied u he op i-
miza ions o some ke nels o imp o e he pe o mance ob ained in
his a chi ec u e. Such op imiza ions a e desc ibed in he ollowing
subsec ions.
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
319
L. López-Villellas e al.
Fig. 2. Speedup o SIMD ke nels o e hei scala e sion on he expe imen al se up
using he la ge inpu s.
5.1. Exploi ing ec o iza ion
Some ke nels implemen x86- ec o ized e sions o hei mos ime-
consuming pa s. In pa icula , BSW and FAST-CHAIN include AVX2
and AVX512 e sions o hei c i ical unc ions using in insics. Simi-
la ly, POA implemen s SIMD e sions o i s code using AVX2-in insics
and SIMD E e ywhe e (SIMDe). We ha e implemen ed SVE-in insics
e sions o FAST-CHAIN, BSW, and WFA and a Neon-in insics e sion
o BPM. SIMDe does no ully suppo SVE ye , so we could no le e age
POA’s SIMDe e sion. Fig. 2 shows he speedup o ec o ized ke nels
o e hei scala e sion on he expe imen al se up using he la ge inpu
o he ke nels. No e ha he SVE ec o leng h o G a i on3 (256 bi s)
is hal he A64FX’s (512 bi s), and he e o e he pe o mance speedups
o SVE ke nels o e hei scala e sions a e mo e modes in G a i on3.
BPM: The co e idea behind ec o izing BPM is o ans o m he
alignmen ope a ions used o ill he dynamic p og amming able in o
simple machine-wo d ope a ions. These simple ope a ions a e in ege
addi ions, bi shi s, and bi wise ORs and ANDs. This way, a ious
dynamic p og amming cells a e bi -packed wi hin a machine wo d
and i s dependencies a e encoded using bi -wise ope a ions. In packed
SIMD, ec o ope a ions a e pe o med in independen packe s wi h a
maximum wid h equal o he machine’s maximum wo d wid h, a he
han a whole bi ec o (i.e., i is no possible o pe o m a 128-bi
wid h ope a ion in a 64-bi double wo d machine). Fo example, when
pe o ming a le -shi ope a ion, he le mos bi o each wo d is los .
Howe e , in o de o ec o ize BPM we would wan his bi o be
appended o he closes -le wo d, e ec i ely pe o ming a ec o -wid h
le -shi ope a ion. To ci cum en his p oblem, we mus pe o m
addi ional ope a ions o manually ca y ha bi o he co ec posi ion.
The numbe o addi ional ope a ions equi ed by his app oach o wo k
scales wi h he ec o leng h. Thus, we decided o e alua e he po en ial
o he ec o e sion o BPM using he Neon ec o ex ension (128-bi
ec o s). The ec o ized loop execu es 1.7× mo e ins uc ions han he
o iginal bu pe o ms 2× ewe i e a ions.
On he A64FX, SIMD e sions o simple ins uc ions, like in ege
addi ion, we e much mo e expensi e han scala ones. Fo example,
a simple 64-bi addi ion akes one cycle, while a ec o addi ion o
wo 64-bi wo ds akes ou cycles. This di e ence in la encies leads
o a slow-down o 2×. G a i on3 has lowe SIMD la encies. Howe e ,
he inc ease in he numbe o ins uc ions in he loop leads o a 30%
pe o mance loss. Since we did no gain any pe o mance using he
Neon e sion, i was disca ded in a o o he o iginal scala code.
We belie e ha an in e -sequence o coa se-g ain app oach (i.e., pe -
o m he sequence alignmen o se e al sequences simul aneously) will
deli e be e pe o mance since i simpli ies he ec o iza ion.
BSW: The SVE e sion o BSW [70] is a ansla ion o A m SVE-
in insics o he x86- ec o e sion ound in BWA-MEM2, which g oups
he sequence alignmen o mul iple equal-leng h sequences ia SIMD in-
s uc ions (i.e., in e -sequence ec o iza ion). The x86-in insics e sion
o BSW elies on masks and blend ope a ions o selec alid en ies om
he ec o egis e s. The SVE e sion akes ad an age o SVE’s p edica e
ins uc ions o a oid he need o blend ope a ions, e ec i ely educing
he numbe o o al ins uc ions execu ed. BSW uses in ege s o 16 bi s,
allowing o p ocess 32 elemen s pe i e a ion using SVE-512 (A64FX)
and 16 using SVE-256 (G a i on3).
The SVE e sion o BSW pe o ms 3.4× and 1.3× as e han i s scala
e sion on he A64FX and G a i on3, espec i ely.
FAST-CHAIN: Ou SVE implemen a ion o FAST-CHAIN is a ansla-
ion o SVE in insics o he x86 e sion. The o iginal x86 implemen a-
ion o FAST-CHAIN execu es i s main loop scala e sion (i.e., a oids
execu ing he ec o ized loop) when he numbe o i e a ions o pe -
o m is small. Addi ionally, as usual in x86 ec o loops, i implemen s
a loop- ail o p ocess he emaining elemen s. Since SVE is ec o -leng h
agnos ic, we could a oid mos o he logic o he x86 e sion, educing
he numbe o pe o med ins uc ions.
The x86 ec o ized e sion o FAST-CHAIN uses 32 bi s ancho s. In
some cases, 32-bi ancho s a e no su icien , and his ke nel gene a es
inco ec esul s. To sol e his, we ha e implemen ed 64 and 32 bi s
SVE e sions o FAST-CHAIN. The 64 bi s e sion always ou pu s co -
ec esul s, bu we ha e used he 32 bi s implemen a ion o compa e
agains he 32 bi s x86 implemen a ion.
GenA chBench’s SVE e sion o FAST-CHAIN uns 4.5×and 1.8×
as e han i s scala e sion (CHAIN wi hou heu is ics) on he A64FX
and G a i on3, espec i ely. Expe imen al esul s show ha he pe -
o mance o FAST-CHAIN compa ed o egula CHAIN g ea ly depends
on he inpu used— he usage o heu is ics may lead o pe o mance
a ia ions based on he cha ac e is ics o he inpu . Fo ins ance, us-
ing GenA chBench’s la ge inpu , ou SVE e sion o FAST-CHAIN is
2.2× as e han egula CHAIN on he A64FX, bu i p esen s a 1.4×
slowdown on G a i on3.
WFA: The Wa e on Alignmen Algo i hm consis s o wo ope a-
ions: compu e he nex wa e on (nex ope a ion) and ex end all he
a hes - eaching poin s o a wa e on by exac ma ching cha ac e s
om wo s ings (ex end ope a ion). The nex ope a ion can be au-
oma ically ec o ized by he compile due o i s simple compu a ional
pa e n. In con as , he ex end ope a o canno be au oma ically ec-
o ized as each diagonal equi es an i egula amoun o compu a ions.
To his end, we ha e ec o ized he ex end ope a ion using a cus om
implemen a ion elying on SVE in insics. Each ec o lane ex ends a
di e en diagonal, compa ing ou bases pe lane un il a misma ch is
ound. Because each diagonal equi es a di e en numbe o cha ac e
compa isons, some lanes can equi e mo e i e a ions han o he s. We
ackle his p oblem by masking he lanes as hey inish he ex ension
p ocess. This way, se e al diagonals a e ex ended in pa allel.
The SVE e sion o WFA deli e s a 1.6× and 1.25× speedup o e i s
scala e sion on he A64FX and G a i on3, espec i ely.
5.2. Op imized lib a ies
Many HPC ke nels and ools ely on equen ly used lib a ies. I
is common o endo s, such as A m, Fuji su, o In el, o de elop
op imized e sions o widely used unc ions and lib a ies a ge ing hei
sys ems and a chi ec u es. Fo genome da a analysis, some ools exploi
neu al ne wo ks (NNs) o imp o e he quali y o hei analysis and e-
sul s. Fo he GenA chBench, we ha e es ed di e en implemen a ions
o he lib a ies used by he NN-BASE and NN-VARIANT ke nels.
NN-BASE: The NN-BASE ke nel builds upon he PyTo ch lib a y
[57]. On he A64FX we ha e used an op imized e sion o PyTo ch
o his speci ic CPU p o ided by Fuji su. On G a i on3, we ied
wo di e en PyTo ch backends: PyTo ch compiled wi h OpenBLAS
( ecommended by A m) and PyTo ch compiled wi h oneDNN op imized
wi h ACL (labeled as expe imen al). The docke images wi h he wo
backends a e a ailable in [71]. The oneDNN-ACL backend pe o med
6.5×be e han he OpenBLAS one, and he e o e, we used i o ou
expe imen s. On SKX, we used an op imized e sion o PyTo ch ha
exploi s he AVX512 ec o ex ension. On Rome, we used an op imized
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
320
L. López-Villellas e al.
Fig. 3. Single-co e ( op) and mul i-co e (bo om) execu ion ime o GenA chBench’s ke nels on he expe imen al se up. Mul i-co e esul s co espond o execu ions using all a ailable
co es on each machine: 48 h eads on he A64FX and SKX and 64 h eads on G a i on3 and Rome. The esul s a e no malized o he pe o mance on he A64FX using one co e
( op) and 48 co es (bo om). FCHAIN, KCNT, NNB and NNV a e he abb e ia ions o FAST-CHAIN, KMER-CNT, NN-BASE and NN-VARIANT, espec i ely. NN-VARIANT is no
aken in o conside a ion o he a e age in he mul i-co e plo .
PyTo ch e sion ha suppo s he AVX2 ec o ex ension a ailable on
he machine.
NN-VARIANT: The o iginal NN-VARIANT ke nel om Genomics-
Bench is based on Clai [25] a ian calle . In u n, his a ian calle e-
lies on Tenso Flow [58]. Clai uses Tenso Flow 1 while Fuji su p o ides
an op imized e sion o Tenso Flow 2 o he A64FX. Fo ha eason,
we decided o use Clai 3 [26] ins ead, an upda ed e sion o Clai ha
elies on Tenso Flow 2. To execu e using GenA chBench’s inpu s, we
used he Ox o d Nanopo e 941_p om_hac_g360+g422 [60] p e-
ained model om Clai 3. On G a i on3, we es ed h ee di e en
Tenso Flow backends: Tenso Flow compiled wi h oneDNN op imized
wi h ACL, using Tenso Flow’s Eigen h ead-pool o pa allelism ( ecom-
mended by A m); Tenso Flow compiled wi h oneDNN op imized wi h
ACL, using ACL’s schedule ; and Tenso low compiled wi h he Eigen
backend. The docke images wi h he h ee backends a e a ailable
in [72]. The Eigen backend pe o med mo e han 1.6×be e han he
o he and he e o e i was he one used o un ou expe imen s. We used
op imized Tenso Flow e sions on SKX and Rome capable o exploi ing
he AVX512 and AVX2 ec o ex ensions.
5.3. Algo i hmic and code op imiza ions
This sec ion p esen s he algo i hmic and code op imiza ion we
pe o med o imp o e he pe o mance o FMI and KMER-CNT.
FMI: GenA chBench’s FMI e sion implemen s h ee op imiza ions
p oposed by Langa i a e al. [70]. One o he mos called unc ions in
his ke nel is backwa dEx . To educe he o e head o he calls, his
unc ion is always o ced o be in-lined. FMI uses he
buil in_popcoun unc ion. This unc ion coun s he numbe o
bi s se o one in an in ege . None o he es ed compile s ansla es
his unc ion o SVE’s popula ion coun ins uc ion. Ins ead, hey use
bi wise ope a ions and masks. To o ce exploi ing SVE capabili ies,
all calls o buil in_popcoun a e eplaced by SVE in insics. FMI
pe o mance is hea ily a ec ed by memo y access la encies. To hide
hese la encies, he op imized e sion o FMI in e lea es he execu ion
o se e al sequences, e ec i ely pe o ming se e al memo y accesses in
pa allel. By applying he h ee p esen ed op imiza ions, we imp o ed
he ke nel pe o mance on bo h A m machines by oughly 35%.
KMER-CNT: Ou expe imen al e alua ion shows ha he pe o -
mance o his ke nel is hea ily a ec ed by h ead mig a ions. To a oid
h ead mig a ions, we po ed KMER-CNT om he P h eads lib a y o
OpenMP and se OMP_PROC_BIND clause o ue be o e execu ions.
This change led o mo e han 4×speedups on bo h x86 machines
when using all a ailable co es. Howe e , he pe o mance o he A m
machines emained he same.
KMER-CNT elies on wo global da a s uc u es o s o e he numbe
o indi idual k-me s: an a ay o 4-bi coun e s and libcuckoo’s [55]
mul i- h ead hash-map, which s o es 64-bi coun e s. Each en y o
he global a ay is an 8-bi a omic in ege , which is spli in hal
o c ea e wo 4-bi coun e s. The a ay coun e s a e upda ed using
a omic compa e-and-swap ope a ion. Once he 4-bi coun e o a k-me
sa u a es, he ollowing inc emen s a e pe o med in he global hash
map, also elying on a omic compa e-and-swaps o upda e i s coun e s.
E en by a oiding h ead mig a ions, he scalabili y o he o iginal
ke nel was poo on all he machines. I achie ed a maximum o 7×and
5× s. se ial execu ion on he A64FX (48 h eads) and G a i on3 (64
h eads), espec i ely. We decided o implemen wo new app oaches
o y o imp o e pa allel pe o mance.
The applica ion di ides he inpu be ween he a ailable h eads.
Each h ead i e a es h ough he k-me s o i s pa o he inpu and
inc emen s he coun e o he ead k-me s in he global a ay o
hash map. Since he inpu is ead sequen ially and he e is almos no
compu a ion o pe o m, mos o he execu ion ime is spen accessing
he global coun e s in mu ual exclusion. To educe con en ion and
imp o e da a locali y, ou i s app oach assigns pa o he inpu o
each h ead, and all h eads ead he ull inpu bu only coun pa o
he k-me s. This way, ins ead o ha ing a single global a ay and hash
map, each h ead can ha e a smalle ins ance o he da a s uc u es and
access hem wi hou con en ion.
The o iginal e sion o he ke nel uses compa e-and-swap ins ead
o e ch-and-add o upda e he coun e s because each 8-bi en y o he
a ay s o es wo 4-bi coun e s. In o de o limi memo y usage, when
he k-me size is g ea e han 17, he ke nel does no ins an ia e he
global a ay, and all he coun ing akes place on he hash-map. Fo he
maximum allowed k-me size (17), we equi e 217 4-bi en ies in he
global a ay, esul ing in 8 GB o memo y. Since we ha e mo e han
enough memo y in all sys ems, ou second app oach uses 8-bi ins ead
o 4-bi coun e s, doubling he memo y equi emen s. This enables
e ch-and-add usage and educes he numbe o accesses o he hash
map.
The single h ead execu ion ime o he ke nel did no change wi h
any o he new e sions. Ou i s app oach (p i a e s uc u es) equally
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
321
L. López-Villellas e al.
di ides he possible k-me s be ween h eads. Howe e , some k-me s a e
mo e common in he inpu , causing load imbalance be ween h eads
de i ing in e en poo e scalabili y han he o iginal ke nel. The second
app oach ( e ch-and-add) imp o es he ke nel’s scalabili y on all he
machines: i uns 2.5×,1.4×, 3×and 2.3× as e han he o iginal
e sion on he A64FX (48 h eads), G a i on3 (64 h eads), SKX (48
h eads) and Rome (64 h eads), espec i ely. Consequen ly, we used
he e ch-and-add app oach o he es o he expe imen s.
6. Pe o mance cha ac e iza ion
This sec ion p esen s a de ailed pe o mance cha ac e iza ion o he
ke nels in ou expe imen al se up. We use he op imized e sions o
he ke nels desc ibed in Sec ion 5. Fo all o he s udies p esen ed, we
ha e anno a ed he code o he ke nels o de ine hei egion o in e es ,
i.e., we only s udy he pa o he ke nels dedica ed o meaning ul
compu a ion. All he esul s shown in his sec ion ha e been compu ed
using he la ge inpu o each ke nel. We no ed minimal a ia ion
be ween execu ions o he ke nels, wi h a maximum ela i e s anda d
de ia ion o 5% obse ed ac oss 10 epe i ions o he expe imen s.
Consequen ly, we showcase he esul s based on a single execu ion in
he igu es. While execu ing DBG wi h high h ead coun s in G a i on3,
ou lie execu ion imes occu ed app oxima ely 10% o he ime. In he
case o DBG in G a i on3, we selec i ely p esen esul s om an inlie
execu ion.
6.1. Single- h ead pe o mance
The op plo o Fig. 3 shows he single- h ead execu ion ime o each
ke nel on he expe imen al se up. The esul s a e no malized o he
pe o mance on he A64FX (see Table 3 o he supplemen a y ma e ial
o he execu ion imes o he ke nels).
The A64FX ea u es signi ican ly ewe ou -o -o de esou ces, a
smalle memo y hie a chy, and highe memo y la encies han he
es o he sys ems. On a e age, he o me is 2.4×, 1.8×, and 1.7×
slowe han G a i on3, SKX and Rome on single- h eaded execu ions,
espec i ely. Exploi ing he SVE capabili ies o he A64FX helps o
educe his slowdown. SVE ec o ized ke nels (BSW, FAST-CHAIN,
and WFA) p esen be e - han-a e age pe o mance on he A64FX:
BSW pe o mance is simila o he exhibi ed on G a i on3 and only
17% wo se han he pe o mance on he x86 machines, FAST-CHAIN
pe o ms be e han on Rome, and WFA pe o ms be e han on SKX.
No e ha BSW and FAST-CHAIN exploi AVX512 on SKX while hey
le e age AVX2 on Rome, and ha WFA is no ec o ized on he x86
machines. The deep-lea ning ke nels (NN-BASE and NN-VARIANT) a e
he wo s -pe o ming on he A64FX.
G a i on3 pe o ms excep ionally well in single- h ead execu ions.
On a e age, i p esen s 2.44×, 1.33×, and 1.39×pe o mance speedups
wi h espec o he A64FX, SKX and Rome, espec i ely. FAST-CHAIN
pe o mance on G a i on3 is 1.8×be e han on Rome (AVX2) bu
70% wo se han on SKX since i exploi s AVX512 (512 bi s) on ha
machine. WFA uns 2.5×and 1.8× as e on G a i on3 han on SKX and
Rome, espec i ely. In con as o he A64FX, he deep-lea ning ke nels
(NN-BASE and NN-VARIANT) deli e good pe o mance on G a i on3,
showing speedups o be ween 3.1–6.2×compa ed o he A64FX.
6.2. Pa allel pe o mance
We e alua e he pa allel pe o mance o GenA chBench’s ke nels
using di e en h ead coun s: 2, 8, 24, 48, and 64. The A64FX and
SKX implemen 48 co es. Hence, execu ions wi h mo e han 48 h eads
ha e only been pe o med on G a i on3 and Rome. Con olling h ead
a ini y was manda o y in ou expe imen s o achie e good pa allel
pe o mance on he machines, especially on he A64FX. Fo mos
ke nels, all he execu ions we e pe o med by binding h eads o co es.
Fig. 4. Speedup o e se ial execu ion o GenA chBench’s ke nels on he expe imen al
se up. We show he achie ed speedup using di e en h ead coun s: 2, 8, 24, 48, and
64. The A64FX and SKX 64- h eads poin s a e no shown in he igu e, since hose
machines only implemen 48 co es.
ABEA, NN-BASE, and NN-VARIANT do no allow ull h ead a ini y
con ol. The e o e, h ead mig a ions can occu in hese h ee ke nels.
Fig. 4 shows he speedup o e se ial execu ion achie ed by he
ke nels on he expe imen al se up using he p e iously p esen ed h ead
coun s. Addi ionally, he bo om plo o Fig. 3 compa es he pe o -
mance ob ained using all a ailable co es on each machine: 48 h eads
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
328
L. López-Villellas e al.
[89] H. Li, e al., A su ey o sequence alignmen algo i hms o nex -gene a ion
sequencing, B ie . Bioin o m. 11 (5) (2010) 473–483, h p://dx.doi.o g/10.1093/
bib/bbq015.
[90] A. Zielezinski, e al., Alignmen - ee sequence compa ison: bene i s, applica ions,
and ools, Genome Biol. 18 (1) (2017) h p://dx.doi.o g/10.1186/s13059-017-
1319-7.
[91] Y. Tu akhia, e al., Da win, ACM SIGPLAN No . 53 (2) (2018) 199–213, h p:
//dx.doi.o g/10.1145/3296957.3173193.
[92] A. Nag, e al., Gencache: Le e aging in-cache ope a o s o e icien sequence
alignmen , in: P oceedings o he 52nd Annual IEEE/ACM In e na ional Sym-
posium on Mic oa chi ec u e, 2019, pp. 334–346, h p://dx.doi.o g/10.1145/
3352460.3358308.
[93] D. Fujiki, e al., GenAx: A genome sequencing accele a o , in: 2018 ACM/IEEE
45 h Annual In e na ional Symposium on Compu e A chi ec u e, ISCA, 2018,
pp. 69–82, h p://dx.doi.o g/10.1109/ISCA.2018.00017.
[94] H. Sadasi an, e al., Accele a ed dynamic ime wa ping on GPU o selec i e
nanopo e sequencing, J. Bio echnol. Biomed. 07 (01) (2024) h p://dx.doi.o g/
10.26502/jbb.2642-91280134.
[95] T. Dunn, e al., SquiggleFil e : An accele a o o po able i us de ec ion,
in: MICRO-54: 54 h Annual IEEE/ACM In e na ional Symposium on Mi-
c oa chi ec u e, MICRO ’21, ACM, 2021, h p://dx.doi.o g/10.1145/3466752.
3480117.
[96] P.J. Shih, e al., E icien eal- ime selec i e genome sequencing on
esou ce-cons ained de ices, GigaScience 12 (2022) h p://dx.doi.o g/10.1093/
gigascience/giad046.
[97] T. Robinson, e al., Ha dwa e accele a ion o genomics da a analysis: chal-
lenges and oppo uni ies, Bioin o ma ics (2021) 1–11, h p://dx.doi.o g/10.
1093/bioin o ma ics/b ab017.
Lo ién López-Villellas is a Ph.D. s uden a he Uni e si y
o Za agoza. His esea ch ocuses on exploi ing no el and
consolida ed pa allel and ec o a chi ec u es o scien i ic
applica ions, such as molecula dynamics and genomics.
P io o s a ing his Ph.D., he wo ked as a esea ch enginee
a he Ba celona Supe compu ing Cen e o wo yea s. He
holds a BSc in compu e science om he Uni e si y o
Za agoza and a MSc in High-Pe o mance Compu ing om
he Uni e si a Poli ècnica de Ca alunya.
Rubén Langa i a-Bení ez ecei ed his B.S. deg ee in com-
pu e science om Uni e sidad de Za agoza in 2018. He
spen one academic yea as an E asmus s uden a he
Uni e si y College Co k. His inal deg ee p ojec was abou
op imizing molecula dynamics applica ions. He ecei ed
MS deg ee om UPC in Janua y 2021. He is cu en ly wo k-
ing a he BSC as a esea ch s uden . His esea ch in e es s
include p ocesso mic oa chi ec u e and HPC applica ions.
Asa Badouh is a esea ch so wa e enginee wi h a pas-
sion o high-pe o mance compu ing and machine lea ning.
O e he las 4.5 yea s, Asa has been wo king a he
Ba celona Supe compu ing Cen e as a Resea ch Enginee ,
ocusing on genomics and heal hca e applica ions. P io o
ha , Asa wo ked o 6 yea s as an in e n and enginee in
he In el compile R&D eam in Is ael, de eloping LLVM-
based compile s o OpenCL and C/C++. Aside om ha ,
Asa holds a bachelo ’s deg ee in compu e science om he
Technion, Is ael; and a mas e ’s deg ee in Da a Science om
Uni e si a Poli ècnica de Ca alunya.
Víc o So ia-Pa dos ecei ed a B.Sc. in compu e science
om Uni e si a de Za agoza, in 2019. He ecei ed an
M.Sc. in compu e science om Uni e si a Poli ècnica de
Ca alunya (UPC), in 2022. He is cu en ly a second-yea
Ph.D. in compu e a chi ec u e wi h he UPC. He also wo ks
as a esea che in he Ba celona Supe compu ing Cen e ,
wi hin he Cen e o Excellence pa ne ship wi h A m.
His esea ch in e es s include high-pe o mance compu ing
a chi ec u es, cache cohe ence, and mul ico e a chi ec u es.
Quim Aguado-Puig ecei ed a B.Sc. deg ee in compu e sci-
ence in 2019 om he Uni e si a Au ònoma de Ba celona
(UAB). He ecei ed an MSc a he Uni e si a Poli ècnica de
Ca alunya (UPC) in 2023. He is cu en ly a i s -yea Ph.D.
s uden a UAB. He has p e iously wo ked as a esea ch
enginee in he p ojec Designing RISC-V-based Accele a o s
o nex -gene a ion Compu e s (DRAC) a UAB in collab-
o a ion wi h he Ba celona Supe compu ing Cen e (BSC).
His esea ch in e es s include high-pe o mance compu ing,
massi ely pa allel a chi ec u es, and GPU p og amming;
wi h applica ions o genomics, compu a ional biology, and
sequence alignmen .
Guillem López-Pa adís ecei ed a B.Sc. and M.Sc. om
Uni e si a Poli ecnica de Ca alunya (UPC) in 2017 and
2020, espec i ely. He is cu en ly a hi d yea Ph.D.
s uden in Compu e A chi ec u e a UPC and Ba celona
Supe compu ing Cen e (BSC). He ac i ely pa icipa es in
di e en Eu opean p ojec s, as well as in in e na ional
collabo a ions wi h academia and indus y. He has al-
eady published some pape s in in e na ional con e ences
and pa icipa ed in di e en apeou s designing powe -
e icien RISC-V p ocesso s. His esea ch in e es s include
high-pe o mance compu ing a chi ec u es, scaling RTL Sim-
ula ions, and domain-speci ic accele a o s, wi h special
emphasis on cohe en in e connec s be ween co es and
ha dwa e accele a o s.
Max Doblas ecei ed a B.Sc. in elec ical enginee ing and
compu e science om Uni e si a Poli ècnica de Ca alunya
(UPC), in 2020. He ecei ed an M.Sc. in compu e sci-
ence om UPC, in 2021. He is cu en ly a second-yea
Ph.D. in compu e a chi ec u e wi h he UPC. He also
wo ks as a Resea ch Enginee in he p ojec Designing
RISC-V-based Accele a o s o nex -gene a ion Compu e s
(DRAC) a he Ba celona Supe compu ing Cen e (BSC),
in which he has designed a powe -e icien p ocesso
wi h se e al ex ensions o domain-speci ic applica ions.
His esea ch in e es s include high-pe o mance compu ing
a chi ec u es, and domain-speci ic accele a o s, wi h appli-
ca ions o genomics, compu a ional biology, and sequence
alignmen .
Ja ie Se oain ecei ed his Ph.D. in compu e science and
enginee ing om he Complu ense Uni e si y o Mad id
(UCM), wo king on wo kload op imiza ion o GPUs. A e
wo king as a pos -doc a he Spanish Na ional Cen e
o Bio echnology (CNB), and a esea ch enginee a A m
Resea ch, he is cu en ly a senio membe o echnical s a
in AMD Resea ch and Ad anced De elopmen , wo king on
compile esea ch o AI accele a o s. His esea ch in e es s
cen e a ound compu e accele a o s and specialized a chi-
ec u es, and his cu en esea ch is ocused on au oma ed
ML wo kload op imiza ions o spa ial a chi ec u es.
Chulho Kim is P incipal Consul an wi h he Leno o
In as uc u e Solu ions G oup Se ices in Uni ed S a es.
Chulho has a Bachelo o Science deg ee in Ma hema -
ics o Compu a ion om UCLA in 1989 and joined IBM
Kings on. He joined Leno o US in 2014. He has wo ked
in High Pe o mance Compu ing (HPC) since 1993. He
likes o apply his skills o debug ex emely complex issues,
anging om applica ion pe o mance o sys em and ne -
wo king pe o mance issues. He is esponsible o unning
Top500/G een500 on cus ome clus e s o Leno o (#1
G een500 en y since No embe 2022).
Mako o Ono ecei ed he ME in biophysics om he Osaka
Uni e si y in 1985 and joined IBM Tokyo Resea ch lab. He
wo ked on compu e g aphics esea ch hen s a ed sys em
a chi ec u e de elopmen in IBM Sys em x de elopmen .
He is cu en ly a Dis inguished Enginee a Leno o In as-
uc u e Solu ions G oup and is a lead a chi ec o edge
compu ing. His in e es s include edge compu ing, edge AI,
and he e ogeneous and al e na i e a chi ec u e including
non adi ional CPU / GPU a chi ec u e.

Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
329
L. López-Villellas e al.
Ad ià A mejach is a Lec u e P o esso in Compu e A -
chi ec u e a Uni e si a Poli ècnica de Ca alunya (UPC),
and associa e esea che a he Ba celona Supe compu ing
Cen e (BSC). He ecei ed his Ph.D. om UPC in 2014 and
hen s a ed his esea ch ca ee a BSC, whe e he lead he
echnical con ibu ions o mul iple FP7 and H2020 p ojec s.
His esea ch in e es include memo y sys ems, he e oge-
neous a chi ec u es, simula ion me hodologies and ec o
a chi ec u es. Cu en ly he leads a g oup ha o e sees all
he a chi ec u al simula ion e o s a BSC, enabling esea ch
on mul iple compu e a chi ec u e opics. He has published
mo e han 30 well- anked in e na ional con e ence and
jou nal pape s.
San iago Ma co-Sola ecei ed he M.Sc. and Ph.D. deg ees
in compu e science om he Uni e si a Poli ècnica de
Ca alunya (UPC) in 2012 and 2017, espec i ely. Du -
ing his Ph.D., he wo ked a he Algo i hm De elopmen
and Bioin o ma ics G oup a Spanish Na ional Cen e o
Genome Analysis (CNAG) and lec u ed a he Uni e si a
Au ònoma de Ba celona (UAB). He is cu en ly a Senio
Resea che a he Ba celona Supe compu ing Cen e (BSC)
and a Lec u e a he UPC. His esea ch in e es s include
high-pe o mance compu ing, he e ogeneous a chi ec u es,
genome-da a analysis, and algo i hms in bioin o ma ics and
compu a ional biology.
Jesús Alas uey-Benedé ecei ed he M.S. deg ee in
Telecommunica ion and he Ph.D. deg ee in Compu e Sci-
ence om Uni e sidad de Za agoza in 1997 and 2009,
espec i ely. He is an associa e p o esso in he Compu e
Science and Sys ems Enginee ing Depa men (DIIS), Uni-
e sidad de Za agoza, Spain. His esea ch in e es s include
p ocesso mic oa chi ec u e, memo y hie a chy, and HPC
applica ions.
Pablo Ibáñez ecei ed he M.S. deg ee in compu e science
om he Uni e si a Poli écnica de Ca alunya, Spain, in
1989, and he Ph.D. deg ee in compu e science om he
Uni e sidad de Za agoza, Spain, in 1998. He is an associa e
p o esso wi h he Compu e Science and Sys ems Engi-
nee ing Depa men , Uni e si y o Za agoza. His esea ch
in e es s include p ocesso mic oa chi ec u e, memo y hie -
a chy, pa allel compu e a chi ec u e, and high pe o mance
compu ing applica ions.
Miquel Mo e ó ecei ed he B.Sc., M.Sc., and Ph.D. de-
g ees om Uni e si a Poli ècnica de Ca alunya (UPC),
Spain. Cu en ly, he is a Ramón y Cajal Fellow a UPC
Ba celona. P io o joining UPC, he spen 5 yea s as a
Senio Resea che wi h he Ba celona Supe compu ing Cen-
e (BSC), Spain, and 15 mon hs as a pos -doc o al ellow
wi h he In e na ional Compu e Science Ins i u e (ICSI),
Be keley. His esea ch in e es s include high pe o mance
compu e a chi ec u es, domain-speci ic accele a o s and
ha dwa e-so wa e co-design o u u e massi ely pa allel
sys ems.