Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
A ailable online 2 Ap il 2024
0167-739X/© 2024 The Au ho (s). Published by Else ie B.V. This is an open access a icle unde he CC BY-NC license (h p://c ea i ecommons.o g/licenses/by-
nc/4.0/).
Con en s lis s a ailable a ScienceDi ec
Fu u e Gene a ion Compu e Sys ems
jou nal homepage: www.else ie .com/loca e/ gcs
GenA chBench: A genomics benchma k sui e o a m HPC p ocesso s
Lo ién López-Villellasd,∗,1,Rubén Langa i a-Bení eza,Asa Badouha,Víc o So ia-Pa dosa,
Quim Aguado-Puigb,Guillem López-Pa adísa,Max Doblasa,Ja ie Se oaine,Chulho Kim ,
Mako o Onog,Ad ià A mejacha,c,San iago Ma co-Solaa,c,Jesús Alas uey-Benedéd,
Pablo Ibáñezd,Miquel Mo e óa,c
aBa celona Supe compu ing Cen e , Ba celona, Spain
bDepa men d’A qui ec u a de Compu ado s, Uni e si a Au ònoma de Ba celona, Ba celona, Spain
cDepa men d’A qui ec u a de Compu ado s, Uni e si a Poli ècnica de Ca alunya, Ba celona, Spain
dDepa amen o de In o má ica e Ingenie ía de Sis emas/A agón Ins i u e o Enginee ing Resea ch (I3A), Uni e sidad de Za agoza, Za agoza, Spain
eA m Resea ch, Camb idge, Uni ed Kingdom
Leno o Resea ch, Uni ed S a es
gLeno o In as uc u e Solu ions G oup, Uni ed S a es
ARTICLE INFO
Keywo ds:
Genomics
A m
High-pe o mance compu ing
Pa allel compu ing
Vec o compu ing
Pe o mance cha ac e iza ion
ABSTRACT
A m usage has subs an ially g own in he High-Pe o mance Compu ing (HPC) communi y. Japanese supe -
compu e Fugaku, powe ed by A m-based A64FX p ocesso s, held he op posi ion on he Top500 lis be ween
June 2020 and June 2022, cu en ly si ing in he ou h posi ion. The ecen ly eleased 7 h gene a ion o
Amazon EC2 ins ances o compu e-in ensi e wo kloads (C7 g) is also powe ed by A m G a i on3 p ocesso s.
P ojec s like Eu opean Mon -Blanc and U.S. DOE/NNSA As a a e u he examples o A m i up ion in HPC. In
pa allel, o e he las decade, he apid imp o emen o genomic sequencing echnologies and he exponen ial
g ow h o sequencing da a has placed a signi ican bo leneck on he compu a ional side. While mos genomics
applica ions ha e been ho oughly es ed and op imized o x86 sys ems, jus a ew a e p epa ed o pe o m
e icien ly on A m machines. Mo eo e , hese applica ions do no exploi he newly in oduced Scalable Vec o
Ex ensions (SVE).
This pape p esen s GenA chBench, he i s genome analysis benchma k sui e a ge ing A m a chi ec u es.
We ha e selec ed compu a ionally demanding ke nels om he mos widely used ools in genome da a
analysis and po ed hem o A m-based A64FX and G a i on3 p ocesso s. O e all, he GenA ch benchma k
sui e comp ises 13 mul i-co e ke nels om c i ical s ages o widely-used genome analysis pipelines, including
base-calling, ead mapping, a ian calling, and genome assembly. Ou benchma k sui e includes di e en
inpu da a se s pe ke nel (small and la ge), each wi h a co esponding eg ession es o e i y he
co ec ness o each execu ion au oma ically. Mo eo e , he po ing ea u es he usage o he no el A m SVE
ins uc ions, algo i hmic and code op imiza ions, and he exploi a ion o A m-op imized lib a ies. We p esen
he op imiza ions implemen ed in each ke nel and a de ailed pe o mance e alua ion and compa ison o
hei pe o mance on ou di e en HPC machines (i.e., A64FX, G a i on3, In el Xeon Skylake Pla inum, and
AMD EPYC Rome). O e all, he expe imen al e alua ion shows ha G a i on3 ou pe o ms o he machines
on a e age. Mo eo e , we obse ed ha he pe o mance o he A64FX is signi ican ly cons ained by i s
small memo y hie a chy and la encies. Addi ionally, as p oo o concep , we s udy he pe o mance o a
p oduc ion- eady ool ha exploi s wo o he po ed and op imized genomic ke nels.
1. In oduc ion
Fo many yea s, A m p ocesso s ha e domina ed he mobile de ice
segmen . Thei ene gy e iciency and license-based business model ha e
been he pilla s unde pinning his success.
∗Co espondence o: Depa amen o de In o má ica e Ingenie ía de Sis emas/A agón Ins i u e o Enginee ing Resea ch (I3A), Uni e sidad de Za agoza, Spain.
E-mail add ess: [email p o ec ed] (L. López-Villellas).
1The co esponding au ho conduc ed his wo k while a ilia ed wi h he Ba celona Supe compu ing Cen e .
In ecen yea s, A m has bu s on o he high-pe o mance compu ing
ma ke wi h in luen ial companies and conso iums ha ha e become
licensees, such as Fuji su, Amazon, Apple, NVIDIA, Samsung, AMD,
B oadcom, HUAWEI, and Qualcomm. Cu en ly, he A m-based Fuji su
h ps://doi.o g/10.1016/j. u u e.2024.03.050
Recei ed 3 No embe 2023; Recei ed in e ised o m 25 Ma ch 2024; Accep ed 31 Ma ch 2024
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
314
L. López-Villellas e al.
A64FX p ocesso powe s he Japanese supe compu e Fugaku, which
held he op posi ion on he Top500 lis be ween June 2020 and June
2022 and is cu en ly in he ou h posi ion. Mo eo e , Amazon has
been using A m p ocesso s o powe i s cloud compu ing pla o m
(AWS), s a ing in 2018 wi h he G a i on p ocesso . They ollowed
wi h he second gene a ion o G a i on in 2019 and he ecen ly
eleased G a i on3.
In he nea u u e, NVIDIA G ace CPUs and Ampe e se e s will
be leading u he e o s o b eak h ough A m in HPC. As a esul ,
la ge-scale compu ing in as uc u es, usually equipped wi h x86 and
IBM Powe p ocesso s, now ha e an addi ional compe i i e al e na i e.
Howe e , mos o he scien i ic code o HPC is no ully adap ed and
op imized o A m a chi ec u es.
O e he las decade, genome sequencing has become he co ne -
s one o genomics and mode n p ecision medicine. Due o he apid
imp o emen o sequencing echnologies, i is cu en ly possible o
sequence an indi idual’s genome in less han 24 h. This b eak h ough
has enabled e ec i e pe sonalized heal hca e, allowing he diagnosis
and ea men o diseases based on each pe son’s unique genomic
disposi ion [1]. Fu he mo e, genome sequencing has also been p o en
c ucial in cance s udies [2], d ug de elopmen [3], o COVID-19 ou -
b eak con ol [4]. In he pas 20 yea s, genome sequencing cos s ha e
d opped d ama ically and he amoun o sequencing da a p oduced
yea ly has inc eased exponen ially. Mo e no ably, his inc ease in da a
p oduc ion has ou pe o med he pace o Moo e’s law. As a esul , a
signi ican bo leneck in cu en genome sequencing analysis is placed
on he compu a ional side, execu ing compu a ional-in ensi e genomics
ools and pipelines.
Genome analysis pipelines ha e his o ically been designed o un
e icien ly on x86 a chi ec u es. Wi h he i up ion o A m-based HPC
se e s, adap ing and op imizing genomics ools o exploi HPC A m
a chi ec u es e ec i ely has become pa amoun . Fo ha , we ha e
selec ed 13 compu a ionally-demanding CPU ke nels om he mos
widely-used genomics ools, and we ha e included hem in a bench-
ma k sui e called GenA chBench. All he ke nels exploi mul i-co e
pa allelism and implemen common s ages om widely-used genome
analysis pipelines such as base-calling, ead mapping, a ian call-
ing, and de-no o assembly. Addi ionally, GenA chBench includes inpu
da ase s o each ke nel (i.e., a small da ase and a la ge da ase pe
ke nel) and hei co esponding ou pu s o be used as g ound u h. The
small da ase s ha e been sized o equi e single- h ead execu ion imes
no longe han a ew minu es ( o es ing pu poses); meanwhile, la ge
da ase s equi e se e al minu es ( o pe o mance e alua ion pu poses).
Fo con enience, we p o ide au oma ic eg ession es s o all he
ke nels o e i y he co ec ness o he ou pu s.
Fu he mo e, his wo k in oduces code adap a ions and op imiza-
ions o he genomics ke nels a ge ing A m HPC CPUs. GenA chBench
le e ages A m-speci ic HPC lib a ies (ca e ully op imized o A m p o-
cesso s) and p esen s algo i hmic and code op imiza ions o exploi he
a chi ec u e and esou ces o A m HPC machines. No ably, we ha e
op imized some ke nels by u ilizing he la es A m Scalable Vec o
Ex ensions (SVE) o le e age he po en ial o he la es A m HPC
p ocesso s.
In addi ion o he benchma k sui e po ing and op imiza ion, his
wo k p esen s a pe o mance cha ac e iza ion o GenA chBench on
ou HPC machines ( wo A m-based and wo x86-based nodes). The
expe imen al e alua ion compa es he pe o mance o an A64FX p o-
cesso , a G a i on3 p ocesso , an In el Xeon Skylake Pla inum 8160
p ocesso , and an AMD EPYC 7742 Rome p ocesso . This cha ac e i-
za ion includes he ke nels’ ins uc ion b eakdown, single- h ead and
mul i- h ead pe o mance e alua ions, a mic oa chi ec u e bo leneck
analysis, and an ene gy- o-solu ion s udy in he di e en p ocesso s.
Ul ima ely, we e alua e he pe o mance impac o hese op imiza ions
by in eg a ing wo o he accele a ed ke nels in a p oduc ion- eady ool
used in a my iad o genome analysis pipelines.
In summa y, his wo k makes he ollowing con ibu ions:
•We p esen GenA chBench, he i s benchma k sui e a ge ing
A m HPC a chi ec u es o genome analysis pipelines and ools.
The benchma k sui e is publicly a ailable a h ps://gi hub.com/
Lo ienLV/gena chbench/ eleases/ ag/1.0.0.
•We p opose HPC adap a ions and code op imiza ions applied o
GenA chBench’s ke nels o exploi he po en ial o A m HPC p o-
cesso s, le e aging A m-speci ic HPC lib a ies and A m Scalable
Vec o Ex ension (SVE).
•We pe o m a comp ehensi e pe o mance cha ac e iza ion o
GenA chBench in wo HPC A m p ocesso s (i.e., A64FX and
G a i on3). We compa e he pe o mance o A m agains wo
e e ence HPC x86 machines.
2. Backg ound
Genome da a analysis pipelines comp ise mul iple s ages and com-
pu a ional ools, om sequencing biological samples o de i ing mean-
ing ul da a analysis esul s o scien is s and heal hca e p o essionals.
This sec ion in oduces he main sequencing echnologies, pipelines,
and ools used in common genome analysis (Fig. 1 shows a succinc
g aphic summa y).
2.1. Sequencing echnologies
Be o e any compu a ional analysis can be pe o med, biological
DNA samples mus be con e ed o digi al da a. This p ocess is pe -
o med by he sequencing machines (Fig. 1-1), and, despi e he ema k-
able ad ances in he las decades, hese machines a e s ill unable o
ead a comple e DNA molecule om end o end. Ins ead, sequencing
machines allow eading ela i ely small chunks o DNA, called eads
o agmen s, om andom loca ions wi hin he dono ’s DNA genome.
A e wa ds, sequenced eads mus be jigsaw oge he o econs uc o
eassemble he o iginal dono ’s genome.
Sequencing machines a e commonly ca ego ized in o h ee gene a-
ions based on hei echnological ad ancemen s. The i s sequencing
echnologies (Sange e al. [5] and Maxam e al. [6]) we e de eloped
in 1977 and used o sequence he i s d a o he human genome in
2000 [7]. Since hen, sequencing echnologies ha e e ol ed quickly,
simpli ying he sequencing p ocess and inc easing he da a-p oduc ion
h oughpu . In he mid-2000s, second-gene a ion echnologies [8] we e
in oduced and soon eplaced i s -gene a ion echnologies. Second-
gene a ion echnology can gene a e ixed-leng h sequences o 100–300
bps a a h oughpu o ens o gigaby es pe hou and wi h a low eading
e o a e (0.1% o he ead leng h). A p esen , Illumina domina es he
ma ke o second-gene a ion sequencing machines. Recen ly in oduced
hi d-gene a ion echnologies, known as long- ead sequencing, can ead
a iable-leng h sequences o conside able leng h (i.e., ens o kilo base-
pai s) a he expense o lowe p oduc ion h oughpu (less han 10
Gb/hou ) and highe eading e o a e (0.1%–10% o he ead leng h).
Paci ic Biosciences (PacBio) and Ox o d Nanopo e Technologies (ONT)
a e he mos no able manu ac u e s o hi d-gene a ion sequencing
echnologies.
2.2. Genome da a analysis pipelines and ools
Be o e any p ocessing can be pe o med, he sequencing machines’
aw signals mus be ans o med in o sequences o nucleo ides (A,
C, G, T). This p ocess is called basecalling (Fig. 1-2). Typically, a
specialized basecalling ool is used o pe o m his p ocess ailo ed o
each sequencing echnology. Fo ins ance, Boni o [9] and Guppy [10]
a e wo o he mos widely-used ools o basecalling Ox o d Nanopo e’s
aw-signal ou pu .
Once he sequences o nucleo ides a e decoded, sequenced eads
mus be p ocessed and analyzed o de i e meaning ul biological in-
sigh s. Al hough many di e en genome analyses can be pe o med
using sequenced da a, mos analyses begin wi h ei he genome ese-
quencing (1-3.a) o genome assembly (1-3.b). Bo h analyses seek o
econs uc he sample’s genome by pu ing oge he all he sequenced
eads.
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
315
L. López-Villellas e al.
Fig. 1. Wo k low diag am o common genome analysis pipelines. Going om (1) sequencing, h ough (2) basecalling, o (3.a) genome esequencing o (3.b) and genome assembly.
The igu e shows he di e en compu a ional ke nels used wi hin each s age o ool.
2.2.1. Genome esequencing
The mos common app oach o econs uc ing he sample’s genome
is by esequencing and in ol es econs uc ing he sample’s genome
using a p e iously known e e ence genome. Fo ha , each sequenced
ead is loca ed and ma ched o he mos likely o igina ing posi ion
in he e e ence genome, allowing small di e ences (e.g., misma ches,
inse ions, and dele ions). This p ocesses is called ead mapping (Fig. 1-
3.a.1) and i is implemen ed by many ools like BWA-MEM2 [11,12],
Minimap2 [13], Bow ie2 [14,15], and GEM [16]. Read mapping is one
o he mos compu a ionally expensi e s eps in all genome sequence
analyses. Consequen ly, ead mapping has been ex ensi ely s udied and
op imized.
Mos sequence mappe s a e based on he seed-chain-ex end ech-
nique. This echnique implemen s h ee algo i hmic s eps o swi ly
loca e and align a sequence wi h a e e ence genome. Du ing he i s
s ep, known as seeding (Fig. 1-3.a.1.1), he mappe sea ches small
subsequences o he eads (seeds) in he e e ence le e aging an index
s uc u e. The mos widely-used indexes used o seeding a e FM-
Index [17] and hash- ables [18,19]. Seeding educes he po en ial num-
be o loca ions in he e e ence whe e a sequence can ma ch, dec eas-
ing he amoun o wo k pe o med in subsequen s eps. A e wa ds, a
chaining s ep (Fig. 1-3.a.1.2) is pe o med o educe u he he lis o
possible ma ching loca ions in he e e ence. Du ing he chaining s ep,
all he mapped seeds a e p ocessed o ind a colinea chain o seeds
ha can po en ially ma ch he inpu sequence. Finally, du ing he ex-
ensión o alignmen s ep (Fig. 1-3.a.1.3), he inpu sequence is aligned
agains he candida e loca ion in he e e ence genome, disco e ing he
di e ences be ween he dono ’s sequence and he e e ence genome.
Usually, a dynamic p og amming-based algo i hm, such as Needleman–
Wunsch [20] o Smi h–Wa e man–Go oh [21,22], is used o compu e
he alignmen .
A e sequence mapping, once he eads a e loca ed in he e e -
ence genome, a a ian calling algo i hm (Fig. 1-3.a.2) de e mines he
a ian s and mu a ions be ween he dono ’s genome and he e e -
ence genome. These a ia ions p o ide c ucial insigh s in o he ge-
ne ic makeup o he sequenced indi idual, po en ially e ealing ge-
ne ic a ia ions ha may be associa ed wi h diseases and heal h con-
di ions. No able examples o widely-used a ian calle s a e GATK
Haplo ype-Calle [23], Pla ypus [24], Clai [25,26], DeepVa ian [27]
and Medaka [28].
2.2.2. Genome assembly
Despi e he simplici y and e ec i eness o genome esequencing,
he e is s ill a lack o high-quali y e e ence genomes o many species.
In hose si ua ions, genome de-no o assembly (Fig. 1-3.b) is used o
econs uc he dono ’s genome om sc a ch jigsawing he sequenced
eads oge he .
Mos popula de-no o assembly me hods ely on de B uijn g aphs.
Fo a gi en se o sequences, i s co esponding de B uijn g aph con ains
a node pe each sequence’s k-me (i.e., sub-s ing o leng h 𝑘nu-
cleo ides) and an edge ha connec s adjacen and o e lapping k-me s.
Be o e cons uc ing he de B uijn g aph o a se o inpu sequences, he
numbe o unique k-me s in he eads is coun ed (Fig. 1-3.b.1) o p une
he leas equen ones (likely a i ac s o he sequencing p ocess).
A e wa ds, he de B uijn g aph is cons uc ed (Fig. 1-3.b.2). Then,
he consensus sequence is de i ed using mul iple sequence alignmen
(MSA) algo i hms (Fig. 1-3) and he cons uc ed de B uijn g aph.
No able examples o de B uijn g aph based assemble s a e Flye [29],
Canu [30], and Racon [31].
2.2.3. Me agenomics
Beyond genome esequencing, a ian calling, and de-no o assem-
bly, many p e iously desc ibed analysis s eps and ools can be ound
in o he genome analysis pipelines. This is he case o many me age-
nomics analysis pipelines. Me agenomics pipelines seek o analyze
genomic in o ma ion om mixed mic obial communi ies, p o iding
insigh s in o he di e si y, in e ac ions and unc ion o mic oo gan-
isms p esen in an en i onmen al sample. Me agenomics analyses a e
pe o med using ools such as Cen i uge [32], RawMap [33], UN-
CALLED [34], ReadFish [35], K aken2 [36] and Cla k [37]. These
ools employ k-me coun ing (Fig. 1-3.b.1) and seeding echniques
(Fig. 1-3.a.1.1) o hei analysis. Mo eo e , a ian calle s like GATK
Haplo ype Calle [23] and Pla ypus [24] a e used o cons uc De B uijn
g aphs (Fig. 1-3.b.2) and co ec a i ac s p oduced du ing he map-
ping p ocess (Fig. 1-3.a.1). Fu he mo e, he chaining p ocess (Fig. 1-
3.a.1.2) is also u ilized o genome assembly when using al e na i e
app oaches based on de B uijn g aphs [38].
3. GenA ch benchma k sui e
The GenA ch benchma k sui e comp ises 13 mul i h eaded CPU
ke nels de i ed om he mos widely used genomics ools and co e s
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
316
L. López-Villellas e al.
he mos impo an genome sequencing s eps. I includes en ke nels
om he GenomicsBench [39] benchma k sui e and h ee addi ional
ke nels: he Bi -Pa allel Mye s algo i hm [40] (BPM), he Wa e on
Alignmen algo i hm [41] (WFA), and FAST-CHAIN [42]. BPM and
WFA complemen he sequence alignmen ke nels o GenomicsBench o
be e cap u e con empo a y ends. Addi ionally, FAST-CHAIN [42] is
a ecen ec o -enabled eimplemen a ion o he CHAIN ke nel p esen
in GenomicsBench, which allows us o u he explo e he capabili ies
o SVE.
Addi ionally, GenA chBench includes inpu da ase s o each ke -
nel (i.e., a small da ase and a la ge da ase pe ke nel) and hei
co esponding ou pu s o be used as g ound u h. The small da ase s
ha e been sized o equi e single- h ead execu ion imes no longe
han a ew minu es ( o es ing pu poses); meanwhile, la ge da ase s
equi e se e al minu es ( o pe o mance e alua ion pu poses). Fo
con enience, we p o ide au oma ic eg ession es s o all he ke nels
o e i y he co ec ness o he ou pu s.
Al hough some ke nels included in GenA chBench can exploi he
capabili ies o mode n GPUs, his esea ch ocuses on po ing, accel-
e a ing, and e alua ing he pe o mance o genomics ke nels in A m
p ocesso s. Mo eo e , he A m-sys ems e alua ed in his wo k (A64FX
and G a i on3) a e no equipped wi h GPUs.
The ollowing ex p esen s GenA chBench’s ke nels, b ie ly desc ib-
ing i s unc ionali y, which ools use hem, and a desc ip ion o hei
usage and inpu s.
Adap i e Banded Signal o E en Alignmen (ABEA): ABEA is
a dynamic p og amming algo i hm ha compa es aw nanopo e sig-
nals om ONT sequencing machines o a e e ence genome sequence.
ABEA’s implemen a ion is based on he Suzuki–Kasaha a (SK) [43] al-
go i hm. This s ep is pe o med in some ools, such as Nanopolish [44],
o co ec e o s p oduced in he basecalling p ocess (Fig. 1-2). Fo
GenA chBench, we ha e used he CPU implemen a ion o 5c [45],
a e sion o ABEA based on Nanopolish’s, op imized o bo h CPU-
only and hyb id CPU/GPU execu ions. This implemen a ion o ABEA
exploi s coa se-g ain mul i- h eading by di iding he aw signals o
he inpu be ween he a ailable co es. Since he signals a e no o
egula size, 5c implemen s wo k-s ealing o imp o e load balance.
The small and la ge inpu s comp ise 1K and 10K aw FAST5 (ONT)
eads om ch omosome 22 o NA12878 and GRCh38 as he e e ence
genome [46].
Bi -Pa allel Mye s (BPM): BPM [40] is a dynamic p og amming
algo i hm ha inds all loca ions a que y s ing o size 𝑚ma ches a
e e ence s ing o size 𝑛wi h 𝑘o ewe di e ences (Fig. 1-3.a.1.3). I
compu es he app oxima e s ing ma ching o wo s ings in 𝑂(𝑚𝑛∕𝑤)
ime, whe e 𝑤is he wo d size o he machine. BPM is used in ead map-
ping ools, such as GEM-Mappe [16], Edlib [47], G aphAligne [48]
o Hobbes [49]. Fo GenA chBench, we ha e used an in-house imple-
men a ion o he algo i hm ha exploi s mul i- h eading by assigning
di e en pai s o s ings o di e en h eads. The small and la ge
inpu s comp ise 100K and 10M sequence pai s om human sample
SRR7733443 downloaded om he sequence ead a chi e [50].
Banded Smi h–Wa e man (BSW): The Smi h–Wa e man algo i hm
[21] is a dynamic p og amming algo i hm ha compu es he local
sequence alignmen o wo sequences o leng h 𝑚and 𝑛, espec i ely,
in 𝑂(𝑚𝑛) ime and space. A banded e sion o Smi h–Wa e man [51]
is used o align sequences wi h a maximum o 𝑤inse ions/dele ions,
educing he ime and space complexi y o 𝑂(𝑤𝑛)(Fig. 1-3.a.1.3).
BSW is used in a ian disco e y ools such as GATK [23], and in
sequence alignmen so wa e like BWA-MEM [11,12]. Fo GenA ch-
Bench, we ha e used BWA-MEM2’s x86- ec o ized implemen a ion o
BSW. In o de o exploi mul i- h eading, he se o pai s o s ings o
align is dynamically di ided be ween p ocesso s. The small and la ge
inpu s comp ise 100K and 10M sequence pai s om human sample
SRR7733443 [50].
Seed Chaining (CHAIN): Gi en he se o seeds om a DNA se-
quence ( ead) mapped o ano he sequence, such as he e e ence
genome, he chaining s ep (Fig. 1-3.a.1.2) aims o ind a chain o
colinea seeds. This is a ime-consuming s ep pe o med by alignmen
ools, such as Minimap2, and by de-no o assemble s like Flye [29]
o Canu [30]. We ha e used he implemen a ion o CHAIN ound in
GenomicsBench ha ex ends Minimap2’s o exploi in e - ask pa al-
lelism ac oss eads. The small and la ge inpu s comp ise he seeds om
1K, and 10K eads o Pacbio’s Caeno habdi is elegans wo m sequence
da a [52].
SIMD Seed Chaining (FAST-CHAIN): The p e iously p esen ed
implemen a ion o he CHAIN algo i hm u ilizes heu is ics o s op
execu ing when he esul is su icien ly good. This speedups execu ion
a he cos o accu acy, and i hinde s he ec o iza ion o he ke nel.
FAST-CHAIN [42] is an x86- ec o ized e sion o CHAIN ha emo es
he heu is ics o exploi SIMD compu a ion. As a esul , FAST-CHAIN
ou pu s accu a e esul s and p esen s pe o mance gains compa ed o
CHAIN. FAST-CHAIN uses he same inpu s as CHAIN.
De B uijn G aph Cons uc ion (DBG): The De B uijn g aph (DBG)
o an inpu se o eads is used o ep esen he o e laps be ween he
sub-s ings o leng h 𝑘(k-me s) ound in he inpu (Fig. 1-3.b.2). Each
node o he g aph ep esen s a k-me and he edges connec adjacen
k-me s in he inpu se . The cons uc ion o hese g aphs is a ime-
consuming s ep in de-no o assemble s like Flye [29], Canu [30] o
Racon [31], and in a ian calle s such as GATK [23] and Pla ypus [24].
Fo GenA chBench, we ha e used he DBG cons uc ion o Pla ypus,
which exploi s pa allelism by assigning di e en egions o he inpu
o di e en h eads. Bo h inpu s employ ch omosome 22 o BWA-MEM
aligned eco ds om he Pla inum Genomes da ase [53]. The small
inpu uses bases 16M-16.5M, while he la ge inpu uses he en i e
ch omosome.
FM-Index Sea ch (FMI): The FM-index is a comp essed sub-s ing
index based on he Bu ows–Wheele ans o m [54]. Gi en a sub-s ing
𝑠, FM-index can be used o ind he loca ion o 𝑠in he e e ence
genome in 𝑂(|𝑠|) ime, whe e |𝑠|is he leng h o he sub-s ing (Fig. 1-
3.a.1.1). The FM-index da a s uc u e is used in sequence alignmen
ools such as BWA-MEM [11,12] o Bow ie2 [15], and in me agenomic
classi ica ion so wa e like Cen i uge [32]. Fo GenA chBench, we
ha e used he supe -maximal exac ma ch ke nel o BWA-MEM2, which
u ilizes he FM-Index s uc u e. This ke nel exploi s pa allelism by
dynamically assigning ba ches o eads among h eads. The small and
la ge inpu s comp ise 1M and 10M pai s o 151 bases om human
sample SRR7733443 [50].
K-me Coun ing (KMER-CNT): K-me coun ing aims o coun he
numbe o occu ences o each k-me in an inpu sequence (Fig. 1-
3.b.1). This ask is pe o med in de-no o assemble s such as Flye [29]
o Canu [30] and in me agenomics classi ica ion so wa e like Cla k
[37]. Addi ionally, no e ha he unc ionali y o KMER-CNT is e y
simila o accessing la ge lookup ables, as done in s a e-o - he-a map-
pe s like Minimap2 [13]. Fo GenA chBench, we ha e used he k-me
coun ing ke nel o Flye. This implemen a ion di ides he inpu - eads
be ween h eads and elies on he h ead-sa e hash-map implemen a-
ion o Libcuckoo lib a y [55] o concu en ly inc ease he numbe o
indi idual k-me s shown by each h ead. The small and la ge inpu s
comp ise 1K and 50K Esche ichia coli Ox o d Nanopo e eads sequenced
by Loman Labs [56].
Neu al Ne wo k-based Base Calling (NN-BASE): ONT sequencing
machines moni o changes in an elec ical cu en as single s ands o
DNA o RNA pass h ough a p o ein nanopo e. These changes in he
elec ical cu en a e hen con e ed o a sequence o nucleo ide bases
in he basecalling p ocess (Fig. 1-2). The analog signal ine i ably con-
ains ambigui ies due o noise o measu emen e o s. Some basecalle s,
such as Guppy [10] and Boni o [9], ely on neu al ne wo ks o sol e
hese ambigui ies, de e mining he mos likely obse ed nucleo ide
in each pa o he elec ical cu en . Fo GenA chBench, we ha e
used Boni o’s deep-lea ning base-calle (NN-BASE), which depends on
he PyTo ch lib a y [57]. Boni o spli s he inpu signal in o smalle
chunks o egula size and eeds hem o a PyTo ch neu al ne wo k ha
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
317
L. López-Villellas e al.
Table 1
Cha ac e is ics o e iew o he expe imen al se up.
A64FX G a i on3 SKX Rome
Co es 4 ×12 (+ 4 assis an ) 64 2 ×24 64
SMT No No Disabled Disabled
F equency 2.2 GHz (s a ic) 2.6 GHz 1–2.1 GHz (dynamic) 1.5–2.25 GHz (dynamic)
Max. powe 120 W N/A 2 ×150 W 225 W
Mem. capaci y 4 ×8 GB 8 ×16 GB 2 ×6×8 GB 16 ×64 GB
Mem. echnology on-package HBM2 o -package DDR5 4800 MHz o -package DDR4 2667 MHz o -package DDR4 3200 MHz
Peak bandwid h 4 ×256 GB/s 300 GB/s 2 ×120 GB/s 204.8 GB/s
L1i 64 KB (4-way) 64 KB 32 KB (8-way) 32 KB (8-way)
L1d 64 KB (4-way) 64 KB 32 KB (8-way) 32 KB (8-way)
L2 – 1 MB 1 MB (16-way) 512 KB (8-way)
LLC 4 ×8 MB (16-way) 32 MB 2 ×33 MB (11-way) 16 ×16 MB (16-way)
Vec o ex ension NEON/SVE 512 bi s NEON/SVE 256 bi s SSE/AVX2/AVX512 SSE/AVX2
in e nally exploi s mul i- h eading. The small and la ge inpu s comp ise
1 and 10 aw FAST5 eads om ch omosome 20 o NA12878, ob ained
om he Nanopo e WGS Conso ium [46].
Neu al Ne wo k-based Va ian Calling (NN-VARIANT): Va ian
calling is he p ocess o de ec ing he di e ences ( a ian s o mu a ions)
be ween he aligned eads and he e e ence genome (Fig. 1-3.a.2). This
is a cos ly p ocess pe o med by s a is ics-based a ian calle s, such as
GATK Haplo ypeCalle [23] o Pla ypus [24], and deep-lea ning a ian
calle s, such as Clai [25,26], DeepVa ian [27] o Medaka [28]. Fo
GenA chBench, we ha e used he second gene a ion o Clai a ian
calle (Clai 3), based on he Tenso Flow amewo k [58]. Clai 3 ex-
ploi s pa allelism by di iding he inpu in o egula -size chunks, and
each o hese chunks is p ocessed by one h ead using Tenso Flow. Ou
small and la ge inpu s comp ise 100K and 10M e e ence posi ions,
espec i ely, o ch omosome 20 o HG002 om NITS’s Genome in
a Bo le (GIAB) p ojec [59]. We a e using Clai 3’s ONT p e- ained
model 941_p om_hac_g360+g422 [60].
Pileup Coun ing (PILEUP): Gi en he alignmen da a o a se o
aligned eads o a egion o a e e ence genome, usually a SAM o BAM
ile [61], pileup coun ing is he p ocess o summa izing he base-pai in-
o ma ion a each ch omosomal posi ion. This summa y, called pileup,
is cus oma y he inpu o long- ead neu al ne wo k a ian calle s such
as Clai [25,26] o Meda aka [28] (Fig. 1-3.a.2). Fo GenA chBench
we ha e used he pileup coun ing implemen a ion o Medaka, which
exploi s mul i- h ead pa allelism by dis ibu ing 100 kilobase egions
o he e e ence genome be ween h eads. The small inpu comp ises
bases 1-1499707 o he S aphylococcus au eus genome [10], and he
la ge inpu comp ises bases 1-1412827 o ch omosome 20 o sample
HG002 [59].
Pa ial-O de Alignmen (POA): The cons uc ion o an o e lap
g aph om a se o eads leads o an app oxima e ep esen a ion o he
o iginal sample’s genome. To de e mine he consensus genome o he
sample, he alignmen o all he eads agains each o he is pe o med
in a p ocess called mul iple sequence alignmen (MSA) (Fig. 1-3.b.3).
The Pa ial O de ed Alignmen (POA) algo i hm [62] compu es he
MSA o all sequences by inc emen ally cons uc ing a pa ially-o de
g aph aligning new sequences o i using a dynamic p og amming
algo i hm such as Smi h–Wa e man [21] o Needleman–Wunsch [20].
The mul iple alignmen sequence (consensus sequence) is in e ed om
he g aph by using he Hea ies Bundle algo i hm [63]. POA is used
in so wa e packages such as Nanopolish [44] o Racon [31]. Fo
Gena chBench we ha e used he SIMD-op imized e sion o POA o he
SPOA lib a y [64]. SPOA exploi s mul i- h eading by compu ing he
pa ially-o de ed g aph o mul iple se s o sequences in pa allel. The
small and la ge inpu s comp ise 1K and 6K se s o mul iple sequences
aligned o a e e ence genome, each con aining be ween 5 and 115
sequences. This da a comes om Minimap2’s polishing s ep o he
Flye-assembled S aphylococcus Au eus genome [10].
Wa e on Alignmen (WFA): The wa e on alignmen algo i hm
(WFA) [41] is a pai wise alignmen algo i hm (Fig. 1-3.a.1.3) ha akes
ad an age o homologous egions be ween he sequences o accele a e
he alignmen p ocess. As opposed o adi ional dynamic p og amming
algo i hms ha un in quad a ic ime, WFA ime complexi y is 𝑂(𝑛𝑠),
p opo ional o he ead leng h 𝑛and he alignmen sco e 𝑠, using
𝑂(𝑠2)memo y. The wa e on algo i hm is used in ools such as w -
mash [65], Ancho Wa e [66] o Ances alClus [67]. GenA chBench
uses a cus om mul i- h ead implemen a ion o he algo i hm, in which
each h ead wo ks in he alignmen o a pai o s ings. The small and
la ge inpu s comp ise 100K and 1M sequence pai s om human sample
SRR7733443 [50].
4. Expe imen al se up
Ou expe imen al se up consis s o wo A m and wo x86 HPC
sys ems: a compu e node ea u ing an A m-A64FX p ocesso (A64FX),
a c7 g.16xla ge Amazon-EC2 ins ance (G a i on3), a sys em wi h wo
x86-64 In el Xeon Skylake Pla inum 8160 (SKX), and a compu e node
wi h one x86-64 AMD EPYC 7742 Rome p ocesso (Rome). Table 1
p esen s an o e iew o he main cha ac e is ics o he ou sys ems.
In e ms o compu ing co es, he A64FX is based on ou Non-
Uni o m Memo y Access (NUMA) domains wi hin he chip, also e-
e ed o as co e memo y g oups (CMG). Each NUMA domain has 12
co es, plus one assis ance co e no used o gene al compu ing ( unning
daemons, I/O, asynch onous MPI, e c.). In o al, he A64FX implemen s
48 compu ing co es. G a i on3 implemen s 64 co es in a single NUMA
domain. The AMD Rome CPU comp ises 8 co e chiple s, known as co e
cache dies (CCD), and a cen al I/O die ha con ols all he I/O and
memo y unc ions o he chip. A CCD has wo co e complex (CCX)
clus e s, each wi h 4 co es. Any pai o CCDs can communica e h ough
he I/O die. SKX con ains wo NUMA chips, each wi h 24 physical co es.
Rega ding ope a ional equency, G a i on3 p esen s he highes
maximum equency among he sys ems wi h 2.6 GHz. The o he h ee
sys ems’ maximum equency is e y simila , anging be ween 2.1 and
2.25 GHz. Bo h x86 sys ems dynamically adjus hei equency based
on hei load. Addi ionally, SKX educes i s equency when execu ing
AVX/AVX512 ins uc ions. In con as , he A64FX ope a es a a ixed
equency se o 2.2 GHz. The e is no public in o ma ion abou adap i e
equency ope a ion on G a i on3.
Wi h espec o SIMD ex ensions, he A64FX is he i s CPU o
implemen he A m 8.2-A Scalable Vec o Ex ension (SVE) [68]. One
o SVE’s main ea u es is ha i is Vec o Leng h Agnos ic (VLA);
ha is, he same bina y wo ks on a chi ec u es implemen ing ec o
egis e s o di e en leng hs anging om 128 o 2048 bi s. The A64FX
implemen s 512 bi s SVE egis e s. G a i on3 also implemen s SVE,
wi h a ec o leng h o 256 bi s. The A64FX and G a i on3 also suppo
he A m Neon SIMD ex ension, a non-VLA SIMD ISA ha wo ks wi h
128-bi ec o s. Bo h x86 sys ems implemen he SSE and AVX2 SIMD
ex ensions, wi h a ec o leng h o 128 and 256 bi s, espec i ely. The
SKX also suppo s he AVX512 ex ension, wi h a ec o leng h o 512
bi s. None o he x86 SIMD ex ensions a e VLA.
Conce ning main memo y, each A64FX’s NUMA domain has i s own
local on-chip 8 GB HBM2 main memo y and can access he o he h ee
NUMA domains’ local memo ies ia a ing bus. G a i on3 is connec ed
o 8 ×16 GB DDR5 channels, o a o al o 128 GB o memo y. Each
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
318
L. López-Villellas e al.
Table 2
Load- o-use memo y la encies in nanoseconds o he expe imen al se up.
A64FX G a i on3 SKX Rome
L1 2.3–5 1.5 1.9 1.8
L2 – 4.6 6.7 3.5
LLC 16.8–21.4 33.1 25.1 13.0
Main Mem. Local 118.2–126.4 153.5 86.2 121.5
Main Mem. Remo e 187.7–242.3 – 144.0 –
chip o he SKX is connec ed o 6 ×8 GB DDR4 local channels and can
access he o he chip’s local memo y. The Rome CPU is connec ed o
16 ×64 GB DDR4 channels, o aling 1 TB o memo y.
The cache hie a chy o ganiza ion o he p ocesso s is ela i ely
di e en . Bo h A m machines ha e wo 64 KB p i a e L1 caches pe
co e (ins uc ions and da a), while he x86 CPUs ea u e wo 32 KB
p i a e L1s pe co e. G a i on3 and SKX include one p i a e 1MB
L2 cache pe co e, and Rome has one 512 KB p i a e L2 pe co e.
The A64FX has one 8 MB las -le el cache (LLC) pe NUMA domain,
G a i on3 includes one 32 MB LLC, SKX has wo 33 MB LLCs (one pe
NUMA domain), and Rome includes one 16 MB LLC pe each 4-co e
CCX.
Conce ning memo y bandwid h, he A64FX is designed o achie e
good pe o mance execu ing high memo y bandwid h-demanding ap-
plica ions. The peak bandwid h o his chip (4 ×256 GB/s) is nea ly
3.5 imes highe han he peak bandwid h o G a i on3 (300 GB/s), he
second sys em among he s udied in e ms o memo y h oughpu . I is
ollowed by SKX, eaching up o 120 GB/s pe chip (240 GB/s in o al),
and Rome holds he las posi ion wi h a peak bandwid h o 204.8 GB/s.
Table 2 p esen s he memo y access la encies o each le el o
he memo y hie a chy o all machines. All la encies on Rome and
G a i on3 and la encies o emo e memo ies on he A64FX ha e been
measu ed using he LMbench benchma k [69]. La encies o cache and
local memo y on he A64FX ha e been ex ac ed om he mic o-
a chi ec u e manual o he CPU. La encies on SKX ha e been measu ed
using In el Memo y La ency Checke . The numbe o cycles o access
he A64FX caches depends on he ype o ins uc ion: scala , loa ing-
poin , sho SIMD, and la ge SIMD. The la encies o access he L1 on
he sys ems ange om 1.5 ns (G a i on3) o 5 ns (la ge SIMD access
on he A64FX). E en hough scala accesses on he A64FX a e as e
(2.3 ns), i s ill p esen s he highes L1 access la ency. As p esen ed
p e iously, he A64FX only implemen s wo le els o caches (L1 and
LLC). The L2 access la encies o he o he sys ems ange be ween 3.5 ns
(Rome) o 6.7 ns (SKX). Rome p esen s he as es access o i s LLC
(13 ns), closely ollowed by he A64FX (16.8 ns o scala access and
21.4 ns o la ge SIMD access). The LLC access la ency on he SKX and
G a i on3 is 25.1 and 33.1 ns, espec i ely. SKX p esen s he as es
access la ency o local main memo y (86.2 ns), ollowed by he A64FX
and Rome, wi h simila la encies (∼120 ns). G a i on3 has he highes
local memo y access la ency, as expec ed om cu en DDR5 SDRAMs.
Accessing emo e main memo ies in he A64FX akes be ween 187.7 ns
(nea - emo e memo y) and 242.3 ns ( a - emo e memo y). Accessing
he o he chip’s main memo y on he SKX machine akes 144 ns, 23%
as e han A64FX’s bes case.
The ou -o -o de esou ces o he expe imen al se up a e p esen ed
in Table 3. We assume ha G a i on3 implemen s he same esou ces as
Neo e se V1 o non-publicly a ailable da a (ma ked wi h *). Una ail-
able da a o nei he G a i on3 no Neo e se V1 is ep esen ed as N/A.
The A64FX is igh in ou -o -o de esou ces compa ed wi h he o he
h ee p ocesso s. The SKX and Rome ha e a simila numbe o physical
egis e s, almos doubling he numbe o gene al-pu pose egis e s o
he A64FX (180 s. 96) and implemen ing 30% mo e SIMD/FP egis e s
han he A64FX (160 s. 128). The A64FX can issue up o 7 mic o-
ope a ions (𝜇OP) pe cycle, G a i on3 can issue up o 15, 8 o SKX, and
11 o Rome. The A64FX and SKX a e capable o commi ing 4 mic o-
ope a ions pe cycle. Howe e , SKX can me ge wo mic o-ope a ions
Table 3
Ou -o -o de esou ces o he expe imen al se up.
A64FX G a i on3 SKX Rome
Gene al egis e s 96 N/A 180 180
SIMD/FP egis e s 128 N/A 168 160
Issue wid h 7 (𝜇OP) 15 (𝜇OP) 8 (𝜇OP) 11 (𝜇OP)
Commi wid h 4 (𝜇OP) N/A 4-8 (𝜇OP) 8 (MOP)
ROB (en ies) 128 256* 224 224
LB (en ies) 40 85* 72 44
SB (en ies) 24 90* 56 48
RS (en ies) 2 ×20 +
2×10 + 19
N/A 97 4 ×16 +
28 + 36
*Neo e se V1 CPU de aul s.
in o one used mic o-ope a ion, inc easing i s heo e ical commi a e
o 8 mic o-ope a ions. Rome can commi up o 8 mac o-ope a ions
(MOP) – i.e., ALU, memo y, o me ged ALU/memo y ope a ion – pe
cycle. The eo de bu e (ROB) o G a i on3 (256 en ies) is wice as
big as he A64FX’s (128 en ies). SKX and Rome ha e an iden ical-size
ROB (224 en ies). The sizes o he load bu e s (LB) and s o e bu e s
(SB) o he CPUs a e ela i ely di e en . G a i on3 and SKX implemen
he la ges LB, wi h 85 and 72 en ies, espec i ely. The LB o he
A64FX has 40 en ies, and Rome implemen s a 44-en y LB. Simila ly,
G a i on3 and SKX ha e he la ges SB (90 and 56 en ies, espec i ely).
The A64FX implemen s a 24-en y SB, hal he size o Rome’s. Addi ion-
ally, a s o e ins uc ion on he A64FX occupies one en y in bo h he
load and he s o e bu e . While SKX implemen s a uni ied ese a ion
s a ion (RS) wi h 97 en ies, bo h he A64FX and Rome ha e se e al
smalle RS. The A64FX di ides i s ese a ion s a ion in o 2 ×20 en ies
o 2 in ege , loa ing-poin , and SIMD pipelines, 2 ×10 en ies o 2
add ess calcula ion pipelines, and 19 en ies o he b anch pipeline.
Rome’s ese a ion s a ion has 4 ×16 en ies o 4 in ege pipelines
(scala +SIMD), 28 en ies o 3 add ess calcula ion pipelines, and 36
en ies o 4 loa ing-poin pipelines (scala +SIMD).
5. A m po ing o genomics ke nels
Mos ke nels p esen ed in Sec ion 3 a ge x86 a chi ec u es and
ha e no been ex ensi ely es ed no op imized o A m machines.
Thus, i was expec ed ha some ke nels could un in o ailu es and
e en gene a e inco ec esul s. To e i y he execu ion o he ke nels,
we used he SKX sys em o compu e he co ec ou pu o all ke nels
and inpu s (i.e., g ound u h).
Fo ou expe imen s, we used he GNU compile (GCC) on G a i on3
( 11.2.0), SKX ( 10.1.0), and Rome ( 10.2.0). On he A64FX, we used
GCC ( 10.2.0) and he Fuji su Compile (FCC) ( 4.2.0b). Fo mos
ke nels, FCC-compiled bina ies exhibi ed be e pe o mance. The Fu-
ji su Compile implemen s wo compila ion modes: a adi ional mode
(T ad) based on compile s o ea lie sys ems and a Clang mode based
on Clang/LLVM. In all cases, we ob ained be e execu ion imes when
compiling wi h FCC’s Clang mode. We lacked FCC-compiled e sions
o key-op imized Py hon lib a ies. Fo hese easons, all he esul s
p esen ed in his documen o he A64FX ha e been ob ained using
he Clang mode o FCC, excluding he wo Py hon ke nels (NN-BASE
and NN-VARIANT), whose lib a ies we e compiled using GCC.
We compile all ke nels wi h a leas -O2 op imiza ion le el and
enable CPU-speci ic op imiza ions: -ma ch=a m 8-a+s e on he
A64FX, -mcpu=na i e on G a i on3 and -ma ch=na i e on SKX
and Rome. Enabling CPU-speci ic op imiza ions in ABEA and POA
esul ed in inco ec execu ions, p obably due o p og amming e o s
in he o iginal sou ce code. The e o e, such op imiza ions a e no used
o hese wo ke nels.
A e pe o ming he app op ia e modi ica ions o he ke nels so
all o hem success ully execu e on A m, we applied u he op i-
miza ions o some ke nels o imp o e he pe o mance ob ained in
his a chi ec u e. Such op imiza ions a e desc ibed in he ollowing
subsec ions.
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
319
L. López-Villellas e al.
Fig. 2. Speedup o SIMD ke nels o e hei scala e sion on he expe imen al se up
using he la ge inpu s.
5.1. Exploi ing ec o iza ion
Some ke nels implemen x86- ec o ized e sions o hei mos ime-
consuming pa s. In pa icula , BSW and FAST-CHAIN include AVX2
and AVX512 e sions o hei c i ical unc ions using in insics. Simi-
la ly, POA implemen s SIMD e sions o i s code using AVX2-in insics
and SIMD E e ywhe e (SIMDe). We ha e implemen ed SVE-in insics
e sions o FAST-CHAIN, BSW, and WFA and a Neon-in insics e sion
o BPM. SIMDe does no ully suppo SVE ye , so we could no le e age
POA’s SIMDe e sion. Fig. 2 shows he speedup o ec o ized ke nels
o e hei scala e sion on he expe imen al se up using he la ge inpu
o he ke nels. No e ha he SVE ec o leng h o G a i on3 (256 bi s)
is hal he A64FX’s (512 bi s), and he e o e he pe o mance speedups
o SVE ke nels o e hei scala e sions a e mo e modes in G a i on3.
BPM: The co e idea behind ec o izing BPM is o ans o m he
alignmen ope a ions used o ill he dynamic p og amming able in o
simple machine-wo d ope a ions. These simple ope a ions a e in ege
addi ions, bi shi s, and bi wise ORs and ANDs. This way, a ious
dynamic p og amming cells a e bi -packed wi hin a machine wo d
and i s dependencies a e encoded using bi -wise ope a ions. In packed
SIMD, ec o ope a ions a e pe o med in independen packe s wi h a
maximum wid h equal o he machine’s maximum wo d wid h, a he
han a whole bi ec o (i.e., i is no possible o pe o m a 128-bi
wid h ope a ion in a 64-bi double wo d machine). Fo example, when
pe o ming a le -shi ope a ion, he le mos bi o each wo d is los .
Howe e , in o de o ec o ize BPM we would wan his bi o be
appended o he closes -le wo d, e ec i ely pe o ming a ec o -wid h
le -shi ope a ion. To ci cum en his p oblem, we mus pe o m
addi ional ope a ions o manually ca y ha bi o he co ec posi ion.
The numbe o addi ional ope a ions equi ed by his app oach o wo k
scales wi h he ec o leng h. Thus, we decided o e alua e he po en ial
o he ec o e sion o BPM using he Neon ec o ex ension (128-bi
ec o s). The ec o ized loop execu es 1.7× mo e ins uc ions han he
o iginal bu pe o ms 2× ewe i e a ions.
On he A64FX, SIMD e sions o simple ins uc ions, like in ege
addi ion, we e much mo e expensi e han scala ones. Fo example,
a simple 64-bi addi ion akes one cycle, while a ec o addi ion o
wo 64-bi wo ds akes ou cycles. This di e ence in la encies leads
o a slow-down o 2×. G a i on3 has lowe SIMD la encies. Howe e ,
he inc ease in he numbe o ins uc ions in he loop leads o a 30%
pe o mance loss. Since we did no gain any pe o mance using he
Neon e sion, i was disca ded in a o o he o iginal scala code.
We belie e ha an in e -sequence o coa se-g ain app oach (i.e., pe -
o m he sequence alignmen o se e al sequences simul aneously) will
deli e be e pe o mance since i simpli ies he ec o iza ion.
BSW: The SVE e sion o BSW [70] is a ansla ion o A m SVE-
in insics o he x86- ec o e sion ound in BWA-MEM2, which g oups
he sequence alignmen o mul iple equal-leng h sequences ia SIMD in-
s uc ions (i.e., in e -sequence ec o iza ion). The x86-in insics e sion
o BSW elies on masks and blend ope a ions o selec alid en ies om
he ec o egis e s. The SVE e sion akes ad an age o SVE’s p edica e
ins uc ions o a oid he need o blend ope a ions, e ec i ely educing
he numbe o o al ins uc ions execu ed. BSW uses in ege s o 16 bi s,
allowing o p ocess 32 elemen s pe i e a ion using SVE-512 (A64FX)
and 16 using SVE-256 (G a i on3).
The SVE e sion o BSW pe o ms 3.4× and 1.3× as e han i s scala
e sion on he A64FX and G a i on3, espec i ely.
FAST-CHAIN: Ou SVE implemen a ion o FAST-CHAIN is a ansla-
ion o SVE in insics o he x86 e sion. The o iginal x86 implemen a-
ion o FAST-CHAIN execu es i s main loop scala e sion (i.e., a oids
execu ing he ec o ized loop) when he numbe o i e a ions o pe -
o m is small. Addi ionally, as usual in x86 ec o loops, i implemen s
a loop- ail o p ocess he emaining elemen s. Since SVE is ec o -leng h
agnos ic, we could a oid mos o he logic o he x86 e sion, educing
he numbe o pe o med ins uc ions.
The x86 ec o ized e sion o FAST-CHAIN uses 32 bi s ancho s. In
some cases, 32-bi ancho s a e no su icien , and his ke nel gene a es
inco ec esul s. To sol e his, we ha e implemen ed 64 and 32 bi s
SVE e sions o FAST-CHAIN. The 64 bi s e sion always ou pu s co -
ec esul s, bu we ha e used he 32 bi s implemen a ion o compa e
agains he 32 bi s x86 implemen a ion.
GenA chBench’s SVE e sion o FAST-CHAIN uns 4.5×and 1.8×
as e han i s scala e sion (CHAIN wi hou heu is ics) on he A64FX
and G a i on3, espec i ely. Expe imen al esul s show ha he pe -
o mance o FAST-CHAIN compa ed o egula CHAIN g ea ly depends
on he inpu used— he usage o heu is ics may lead o pe o mance
a ia ions based on he cha ac e is ics o he inpu . Fo ins ance, us-
ing GenA chBench’s la ge inpu , ou SVE e sion o FAST-CHAIN is
2.2× as e han egula CHAIN on he A64FX, bu i p esen s a 1.4×
slowdown on G a i on3.
WFA: The Wa e on Alignmen Algo i hm consis s o wo ope a-
ions: compu e he nex wa e on (nex ope a ion) and ex end all he
a hes - eaching poin s o a wa e on by exac ma ching cha ac e s
om wo s ings (ex end ope a ion). The nex ope a ion can be au-
oma ically ec o ized by he compile due o i s simple compu a ional
pa e n. In con as , he ex end ope a o canno be au oma ically ec-
o ized as each diagonal equi es an i egula amoun o compu a ions.
To his end, we ha e ec o ized he ex end ope a ion using a cus om
implemen a ion elying on SVE in insics. Each ec o lane ex ends a
di e en diagonal, compa ing ou bases pe lane un il a misma ch is
ound. Because each diagonal equi es a di e en numbe o cha ac e
compa isons, some lanes can equi e mo e i e a ions han o he s. We
ackle his p oblem by masking he lanes as hey inish he ex ension
p ocess. This way, se e al diagonals a e ex ended in pa allel.
The SVE e sion o WFA deli e s a 1.6× and 1.25× speedup o e i s
scala e sion on he A64FX and G a i on3, espec i ely.
5.2. Op imized lib a ies
Many HPC ke nels and ools ely on equen ly used lib a ies. I
is common o endo s, such as A m, Fuji su, o In el, o de elop
op imized e sions o widely used unc ions and lib a ies a ge ing hei
sys ems and a chi ec u es. Fo genome da a analysis, some ools exploi
neu al ne wo ks (NNs) o imp o e he quali y o hei analysis and e-
sul s. Fo he GenA chBench, we ha e es ed di e en implemen a ions
o he lib a ies used by he NN-BASE and NN-VARIANT ke nels.
NN-BASE: The NN-BASE ke nel builds upon he PyTo ch lib a y
[57]. On he A64FX we ha e used an op imized e sion o PyTo ch
o his speci ic CPU p o ided by Fuji su. On G a i on3, we ied
wo di e en PyTo ch backends: PyTo ch compiled wi h OpenBLAS
( ecommended by A m) and PyTo ch compiled wi h oneDNN op imized
wi h ACL (labeled as expe imen al). The docke images wi h he wo
backends a e a ailable in [71]. The oneDNN-ACL backend pe o med
6.5×be e han he OpenBLAS one, and he e o e, we used i o ou
expe imen s. On SKX, we used an op imized e sion o PyTo ch ha
exploi s he AVX512 ec o ex ension. On Rome, we used an op imized
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
320
L. López-Villellas e al.
Fig. 3. Single-co e ( op) and mul i-co e (bo om) execu ion ime o GenA chBench’s ke nels on he expe imen al se up. Mul i-co e esul s co espond o execu ions using all a ailable
co es on each machine: 48 h eads on he A64FX and SKX and 64 h eads on G a i on3 and Rome. The esul s a e no malized o he pe o mance on he A64FX using one co e
( op) and 48 co es (bo om). FCHAIN, KCNT, NNB and NNV a e he abb e ia ions o FAST-CHAIN, KMER-CNT, NN-BASE and NN-VARIANT, espec i ely. NN-VARIANT is no
aken in o conside a ion o he a e age in he mul i-co e plo .
PyTo ch e sion ha suppo s he AVX2 ec o ex ension a ailable on
he machine.
NN-VARIANT: The o iginal NN-VARIANT ke nel om Genomics-
Bench is based on Clai [25] a ian calle . In u n, his a ian calle e-
lies on Tenso Flow [58]. Clai uses Tenso Flow 1 while Fuji su p o ides
an op imized e sion o Tenso Flow 2 o he A64FX. Fo ha eason,
we decided o use Clai 3 [26] ins ead, an upda ed e sion o Clai ha
elies on Tenso Flow 2. To execu e using GenA chBench’s inpu s, we
used he Ox o d Nanopo e 941_p om_hac_g360+g422 [60] p e-
ained model om Clai 3. On G a i on3, we es ed h ee di e en
Tenso Flow backends: Tenso Flow compiled wi h oneDNN op imized
wi h ACL, using Tenso Flow’s Eigen h ead-pool o pa allelism ( ecom-
mended by A m); Tenso Flow compiled wi h oneDNN op imized wi h
ACL, using ACL’s schedule ; and Tenso low compiled wi h he Eigen
backend. The docke images wi h he h ee backends a e a ailable
in [72]. The Eigen backend pe o med mo e han 1.6×be e han he
o he and he e o e i was he one used o un ou expe imen s. We used
op imized Tenso Flow e sions on SKX and Rome capable o exploi ing
he AVX512 and AVX2 ec o ex ensions.
5.3. Algo i hmic and code op imiza ions
This sec ion p esen s he algo i hmic and code op imiza ion we
pe o med o imp o e he pe o mance o FMI and KMER-CNT.
FMI: GenA chBench’s FMI e sion implemen s h ee op imiza ions
p oposed by Langa i a e al. [70]. One o he mos called unc ions in
his ke nel is backwa dEx . To educe he o e head o he calls, his
unc ion is always o ced o be in-lined. FMI uses he
buil in_popcoun unc ion. This unc ion coun s he numbe o
bi s se o one in an in ege . None o he es ed compile s ansla es
his unc ion o SVE’s popula ion coun ins uc ion. Ins ead, hey use
bi wise ope a ions and masks. To o ce exploi ing SVE capabili ies,
all calls o buil in_popcoun a e eplaced by SVE in insics. FMI
pe o mance is hea ily a ec ed by memo y access la encies. To hide
hese la encies, he op imized e sion o FMI in e lea es he execu ion
o se e al sequences, e ec i ely pe o ming se e al memo y accesses in
pa allel. By applying he h ee p esen ed op imiza ions, we imp o ed
he ke nel pe o mance on bo h A m machines by oughly 35%.
KMER-CNT: Ou expe imen al e alua ion shows ha he pe o -
mance o his ke nel is hea ily a ec ed by h ead mig a ions. To a oid
h ead mig a ions, we po ed KMER-CNT om he P h eads lib a y o
OpenMP and se OMP_PROC_BIND clause o ue be o e execu ions.
This change led o mo e han 4×speedups on bo h x86 machines
when using all a ailable co es. Howe e , he pe o mance o he A m
machines emained he same.
KMER-CNT elies on wo global da a s uc u es o s o e he numbe
o indi idual k-me s: an a ay o 4-bi coun e s and libcuckoo’s [55]
mul i- h ead hash-map, which s o es 64-bi coun e s. Each en y o
he global a ay is an 8-bi a omic in ege , which is spli in hal
o c ea e wo 4-bi coun e s. The a ay coun e s a e upda ed using
a omic compa e-and-swap ope a ion. Once he 4-bi coun e o a k-me
sa u a es, he ollowing inc emen s a e pe o med in he global hash
map, also elying on a omic compa e-and-swaps o upda e i s coun e s.
E en by a oiding h ead mig a ions, he scalabili y o he o iginal
ke nel was poo on all he machines. I achie ed a maximum o 7×and
5× s. se ial execu ion on he A64FX (48 h eads) and G a i on3 (64
h eads), espec i ely. We decided o implemen wo new app oaches
o y o imp o e pa allel pe o mance.
The applica ion di ides he inpu be ween he a ailable h eads.
Each h ead i e a es h ough he k-me s o i s pa o he inpu and
inc emen s he coun e o he ead k-me s in he global a ay o
hash map. Since he inpu is ead sequen ially and he e is almos no
compu a ion o pe o m, mos o he execu ion ime is spen accessing
he global coun e s in mu ual exclusion. To educe con en ion and
imp o e da a locali y, ou i s app oach assigns pa o he inpu o
each h ead, and all h eads ead he ull inpu bu only coun pa o
he k-me s. This way, ins ead o ha ing a single global a ay and hash
map, each h ead can ha e a smalle ins ance o he da a s uc u es and
access hem wi hou con en ion.
The o iginal e sion o he ke nel uses compa e-and-swap ins ead
o e ch-and-add o upda e he coun e s because each 8-bi en y o he
a ay s o es wo 4-bi coun e s. In o de o limi memo y usage, when
he k-me size is g ea e han 17, he ke nel does no ins an ia e he
global a ay, and all he coun ing akes place on he hash-map. Fo he
maximum allowed k-me size (17), we equi e 217 4-bi en ies in he
global a ay, esul ing in 8 GB o memo y. Since we ha e mo e han
enough memo y in all sys ems, ou second app oach uses 8-bi ins ead
o 4-bi coun e s, doubling he memo y equi emen s. This enables
e ch-and-add usage and educes he numbe o accesses o he hash
map.
The single h ead execu ion ime o he ke nel did no change wi h
any o he new e sions. Ou i s app oach (p i a e s uc u es) equally
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
321
L. López-Villellas e al.
di ides he possible k-me s be ween h eads. Howe e , some k-me s a e
mo e common in he inpu , causing load imbalance be ween h eads
de i ing in e en poo e scalabili y han he o iginal ke nel. The second
app oach ( e ch-and-add) imp o es he ke nel’s scalabili y on all he
machines: i uns 2.5×,1.4×, 3×and 2.3× as e han he o iginal
e sion on he A64FX (48 h eads), G a i on3 (64 h eads), SKX (48
h eads) and Rome (64 h eads), espec i ely. Consequen ly, we used
he e ch-and-add app oach o he es o he expe imen s.
6. Pe o mance cha ac e iza ion
This sec ion p esen s a de ailed pe o mance cha ac e iza ion o he
ke nels in ou expe imen al se up. We use he op imized e sions o
he ke nels desc ibed in Sec ion 5. Fo all o he s udies p esen ed, we
ha e anno a ed he code o he ke nels o de ine hei egion o in e es ,
i.e., we only s udy he pa o he ke nels dedica ed o meaning ul
compu a ion. All he esul s shown in his sec ion ha e been compu ed
using he la ge inpu o each ke nel. We no ed minimal a ia ion
be ween execu ions o he ke nels, wi h a maximum ela i e s anda d
de ia ion o 5% obse ed ac oss 10 epe i ions o he expe imen s.
Consequen ly, we showcase he esul s based on a single execu ion in
he igu es. While execu ing DBG wi h high h ead coun s in G a i on3,
ou lie execu ion imes occu ed app oxima ely 10% o he ime. In he
case o DBG in G a i on3, we selec i ely p esen esul s om an inlie
execu ion.
6.1. Single- h ead pe o mance
The op plo o Fig. 3 shows he single- h ead execu ion ime o each
ke nel on he expe imen al se up. The esul s a e no malized o he
pe o mance on he A64FX (see Table 3 o he supplemen a y ma e ial
o he execu ion imes o he ke nels).
The A64FX ea u es signi ican ly ewe ou -o -o de esou ces, a
smalle memo y hie a chy, and highe memo y la encies han he
es o he sys ems. On a e age, he o me is 2.4×, 1.8×, and 1.7×
slowe han G a i on3, SKX and Rome on single- h eaded execu ions,
espec i ely. Exploi ing he SVE capabili ies o he A64FX helps o
educe his slowdown. SVE ec o ized ke nels (BSW, FAST-CHAIN,
and WFA) p esen be e - han-a e age pe o mance on he A64FX:
BSW pe o mance is simila o he exhibi ed on G a i on3 and only
17% wo se han he pe o mance on he x86 machines, FAST-CHAIN
pe o ms be e han on Rome, and WFA pe o ms be e han on SKX.
No e ha BSW and FAST-CHAIN exploi AVX512 on SKX while hey
le e age AVX2 on Rome, and ha WFA is no ec o ized on he x86
machines. The deep-lea ning ke nels (NN-BASE and NN-VARIANT) a e
he wo s -pe o ming on he A64FX.
G a i on3 pe o ms excep ionally well in single- h ead execu ions.
On a e age, i p esen s 2.44×, 1.33×, and 1.39×pe o mance speedups
wi h espec o he A64FX, SKX and Rome, espec i ely. FAST-CHAIN
pe o mance on G a i on3 is 1.8×be e han on Rome (AVX2) bu
70% wo se han on SKX since i exploi s AVX512 (512 bi s) on ha
machine. WFA uns 2.5×and 1.8× as e on G a i on3 han on SKX and
Rome, espec i ely. In con as o he A64FX, he deep-lea ning ke nels
(NN-BASE and NN-VARIANT) deli e good pe o mance on G a i on3,
showing speedups o be ween 3.1–6.2×compa ed o he A64FX.
6.2. Pa allel pe o mance
We e alua e he pa allel pe o mance o GenA chBench’s ke nels
using di e en h ead coun s: 2, 8, 24, 48, and 64. The A64FX and
SKX implemen 48 co es. Hence, execu ions wi h mo e han 48 h eads
ha e only been pe o med on G a i on3 and Rome. Con olling h ead
a ini y was manda o y in ou expe imen s o achie e good pa allel
pe o mance on he machines, especially on he A64FX. Fo mos
ke nels, all he execu ions we e pe o med by binding h eads o co es.
Fig. 4. Speedup o e se ial execu ion o GenA chBench’s ke nels on he expe imen al
se up. We show he achie ed speedup using di e en h ead coun s: 2, 8, 24, 48, and
64. The A64FX and SKX 64- h eads poin s a e no shown in he igu e, since hose
machines only implemen 48 co es.
ABEA, NN-BASE, and NN-VARIANT do no allow ull h ead a ini y
con ol. The e o e, h ead mig a ions can occu in hese h ee ke nels.
Fig. 4 shows he speedup o e se ial execu ion achie ed by he
ke nels on he expe imen al se up using he p e iously p esen ed h ead
coun s. Addi ionally, he bo om plo o Fig. 3 compa es he pe o -
mance ob ained using all a ailable co es on each machine: 48 h eads
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
328
L. López-Villellas e al.
[89] H. Li, e al., A su ey o sequence alignmen algo i hms o nex -gene a ion
sequencing, B ie . Bioin o m. 11 (5) (2010) 473–483, h p://dx.doi.o g/10.1093/
bib/bbq015.
[90] A. Zielezinski, e al., Alignmen - ee sequence compa ison: bene i s, applica ions,
and ools, Genome Biol. 18 (1) (2017) h p://dx.doi.o g/10.1186/s13059-017-
1319-7.
[91] Y. Tu akhia, e al., Da win, ACM SIGPLAN No . 53 (2) (2018) 199–213, h p:
//dx.doi.o g/10.1145/3296957.3173193.
[92] A. Nag, e al., Gencache: Le e aging in-cache ope a o s o e icien sequence
alignmen , in: P oceedings o he 52nd Annual IEEE/ACM In e na ional Sym-
posium on Mic oa chi ec u e, 2019, pp. 334–346, h p://dx.doi.o g/10.1145/
3352460.3358308.
[93] D. Fujiki, e al., GenAx: A genome sequencing accele a o , in: 2018 ACM/IEEE
45 h Annual In e na ional Symposium on Compu e A chi ec u e, ISCA, 2018,
pp. 69–82, h p://dx.doi.o g/10.1109/ISCA.2018.00017.
[94] H. Sadasi an, e al., Accele a ed dynamic ime wa ping on GPU o selec i e
nanopo e sequencing, J. Bio echnol. Biomed. 07 (01) (2024) h p://dx.doi.o g/
10.26502/jbb.2642-91280134.
[95] T. Dunn, e al., SquiggleFil e : An accele a o o po able i us de ec ion,
in: MICRO-54: 54 h Annual IEEE/ACM In e na ional Symposium on Mi-
c oa chi ec u e, MICRO ’21, ACM, 2021, h p://dx.doi.o g/10.1145/3466752.
3480117.
[96] P.J. Shih, e al., E icien eal- ime selec i e genome sequencing on
esou ce-cons ained de ices, GigaScience 12 (2022) h p://dx.doi.o g/10.1093/
gigascience/giad046.
[97] T. Robinson, e al., Ha dwa e accele a ion o genomics da a analysis: chal-
lenges and oppo uni ies, Bioin o ma ics (2021) 1–11, h p://dx.doi.o g/10.
1093/bioin o ma ics/b ab017.
Lo ién López-Villellas is a Ph.D. s uden a he Uni e si y
o Za agoza. His esea ch ocuses on exploi ing no el and
consolida ed pa allel and ec o a chi ec u es o scien i ic
applica ions, such as molecula dynamics and genomics.
P io o s a ing his Ph.D., he wo ked as a esea ch enginee
a he Ba celona Supe compu ing Cen e o wo yea s. He
holds a BSc in compu e science om he Uni e si y o
Za agoza and a MSc in High-Pe o mance Compu ing om
he Uni e si a Poli ècnica de Ca alunya.
Rubén Langa i a-Bení ez ecei ed his B.S. deg ee in com-
pu e science om Uni e sidad de Za agoza in 2018. He
spen one academic yea as an E asmus s uden a he
Uni e si y College Co k. His inal deg ee p ojec was abou
op imizing molecula dynamics applica ions. He ecei ed
MS deg ee om UPC in Janua y 2021. He is cu en ly wo k-
ing a he BSC as a esea ch s uden . His esea ch in e es s
include p ocesso mic oa chi ec u e and HPC applica ions.
Asa Badouh is a esea ch so wa e enginee wi h a pas-
sion o high-pe o mance compu ing and machine lea ning.
O e he las 4.5 yea s, Asa has been wo king a he
Ba celona Supe compu ing Cen e as a Resea ch Enginee ,
ocusing on genomics and heal hca e applica ions. P io o
ha , Asa wo ked o 6 yea s as an in e n and enginee in
he In el compile R&D eam in Is ael, de eloping LLVM-
based compile s o OpenCL and C/C++. Aside om ha ,
Asa holds a bachelo ’s deg ee in compu e science om he
Technion, Is ael; and a mas e ’s deg ee in Da a Science om
Uni e si a Poli ècnica de Ca alunya.
Víc o So ia-Pa dos ecei ed a B.Sc. in compu e science
om Uni e si a de Za agoza, in 2019. He ecei ed an
M.Sc. in compu e science om Uni e si a Poli ècnica de
Ca alunya (UPC), in 2022. He is cu en ly a second-yea
Ph.D. in compu e a chi ec u e wi h he UPC. He also wo ks
as a esea che in he Ba celona Supe compu ing Cen e ,
wi hin he Cen e o Excellence pa ne ship wi h A m.
His esea ch in e es s include high-pe o mance compu ing
a chi ec u es, cache cohe ence, and mul ico e a chi ec u es.
Quim Aguado-Puig ecei ed a B.Sc. deg ee in compu e sci-
ence in 2019 om he Uni e si a Au ònoma de Ba celona
(UAB). He ecei ed an MSc a he Uni e si a Poli ècnica de
Ca alunya (UPC) in 2023. He is cu en ly a i s -yea Ph.D.
s uden a UAB. He has p e iously wo ked as a esea ch
enginee in he p ojec Designing RISC-V-based Accele a o s
o nex -gene a ion Compu e s (DRAC) a UAB in collab-
o a ion wi h he Ba celona Supe compu ing Cen e (BSC).
His esea ch in e es s include high-pe o mance compu ing,
massi ely pa allel a chi ec u es, and GPU p og amming;
wi h applica ions o genomics, compu a ional biology, and
sequence alignmen .
Guillem López-Pa adís ecei ed a B.Sc. and M.Sc. om
Uni e si a Poli ecnica de Ca alunya (UPC) in 2017 and
2020, espec i ely. He is cu en ly a hi d yea Ph.D.
s uden in Compu e A chi ec u e a UPC and Ba celona
Supe compu ing Cen e (BSC). He ac i ely pa icipa es in
di e en Eu opean p ojec s, as well as in in e na ional
collabo a ions wi h academia and indus y. He has al-
eady published some pape s in in e na ional con e ences
and pa icipa ed in di e en apeou s designing powe -
e icien RISC-V p ocesso s. His esea ch in e es s include
high-pe o mance compu ing a chi ec u es, scaling RTL Sim-
ula ions, and domain-speci ic accele a o s, wi h special
emphasis on cohe en in e connec s be ween co es and
ha dwa e accele a o s.
Max Doblas ecei ed a B.Sc. in elec ical enginee ing and
compu e science om Uni e si a Poli ècnica de Ca alunya
(UPC), in 2020. He ecei ed an M.Sc. in compu e sci-
ence om UPC, in 2021. He is cu en ly a second-yea
Ph.D. in compu e a chi ec u e wi h he UPC. He also
wo ks as a Resea ch Enginee in he p ojec Designing
RISC-V-based Accele a o s o nex -gene a ion Compu e s
(DRAC) a he Ba celona Supe compu ing Cen e (BSC),
in which he has designed a powe -e icien p ocesso
wi h se e al ex ensions o domain-speci ic applica ions.
His esea ch in e es s include high-pe o mance compu ing
a chi ec u es, and domain-speci ic accele a o s, wi h appli-
ca ions o genomics, compu a ional biology, and sequence
alignmen .
Ja ie Se oain ecei ed his Ph.D. in compu e science and
enginee ing om he Complu ense Uni e si y o Mad id
(UCM), wo king on wo kload op imiza ion o GPUs. A e
wo king as a pos -doc a he Spanish Na ional Cen e
o Bio echnology (CNB), and a esea ch enginee a A m
Resea ch, he is cu en ly a senio membe o echnical s a
in AMD Resea ch and Ad anced De elopmen , wo king on
compile esea ch o AI accele a o s. His esea ch in e es s
cen e a ound compu e accele a o s and specialized a chi-
ec u es, and his cu en esea ch is ocused on au oma ed
ML wo kload op imiza ions o spa ial a chi ec u es.
Chulho Kim is P incipal Consul an wi h he Leno o
In as uc u e Solu ions G oup Se ices in Uni ed S a es.
Chulho has a Bachelo o Science deg ee in Ma hema -
ics o Compu a ion om UCLA in 1989 and joined IBM
Kings on. He joined Leno o US in 2014. He has wo ked
in High Pe o mance Compu ing (HPC) since 1993. He
likes o apply his skills o debug ex emely complex issues,
anging om applica ion pe o mance o sys em and ne -
wo king pe o mance issues. He is esponsible o unning
Top500/G een500 on cus ome clus e s o Leno o (#1
G een500 en y since No embe 2022).
Mako o Ono ecei ed he ME in biophysics om he Osaka
Uni e si y in 1985 and joined IBM Tokyo Resea ch lab. He
wo ked on compu e g aphics esea ch hen s a ed sys em
a chi ec u e de elopmen in IBM Sys em x de elopmen .
He is cu en ly a Dis inguished Enginee a Leno o In as-
uc u e Solu ions G oup and is a lead a chi ec o edge
compu ing. His in e es s include edge compu ing, edge AI,
and he e ogeneous and al e na i e a chi ec u e including
non adi ional CPU / GPU a chi ec u e.
Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329
329
L. López-Villellas e al.
Ad ià A mejach is a Lec u e P o esso in Compu e A -
chi ec u e a Uni e si a Poli ècnica de Ca alunya (UPC),
and associa e esea che a he Ba celona Supe compu ing
Cen e (BSC). He ecei ed his Ph.D. om UPC in 2014 and
hen s a ed his esea ch ca ee a BSC, whe e he lead he
echnical con ibu ions o mul iple FP7 and H2020 p ojec s.
His esea ch in e es include memo y sys ems, he e oge-
neous a chi ec u es, simula ion me hodologies and ec o
a chi ec u es. Cu en ly he leads a g oup ha o e sees all
he a chi ec u al simula ion e o s a BSC, enabling esea ch
on mul iple compu e a chi ec u e opics. He has published
mo e han 30 well- anked in e na ional con e ence and
jou nal pape s.
San iago Ma co-Sola ecei ed he M.Sc. and Ph.D. deg ees
in compu e science om he Uni e si a Poli ècnica de
Ca alunya (UPC) in 2012 and 2017, espec i ely. Du -
ing his Ph.D., he wo ked a he Algo i hm De elopmen
and Bioin o ma ics G oup a Spanish Na ional Cen e o
Genome Analysis (CNAG) and lec u ed a he Uni e si a
Au ònoma de Ba celona (UAB). He is cu en ly a Senio
Resea che a he Ba celona Supe compu ing Cen e (BSC)
and a Lec u e a he UPC. His esea ch in e es s include
high-pe o mance compu ing, he e ogeneous a chi ec u es,
genome-da a analysis, and algo i hms in bioin o ma ics and
compu a ional biology.
Jesús Alas uey-Benedé ecei ed he M.S. deg ee in
Telecommunica ion and he Ph.D. deg ee in Compu e Sci-
ence om Uni e sidad de Za agoza in 1997 and 2009,
espec i ely. He is an associa e p o esso in he Compu e
Science and Sys ems Enginee ing Depa men (DIIS), Uni-
e sidad de Za agoza, Spain. His esea ch in e es s include
p ocesso mic oa chi ec u e, memo y hie a chy, and HPC
applica ions.
Pablo Ibáñez ecei ed he M.S. deg ee in compu e science
om he Uni e si a Poli écnica de Ca alunya, Spain, in
1989, and he Ph.D. deg ee in compu e science om he
Uni e sidad de Za agoza, Spain, in 1998. He is an associa e
p o esso wi h he Compu e Science and Sys ems Engi-
nee ing Depa men , Uni e si y o Za agoza. His esea ch
in e es s include p ocesso mic oa chi ec u e, memo y hie -
a chy, pa allel compu e a chi ec u e, and high pe o mance
compu ing applica ions.
Miquel Mo e ó ecei ed he B.Sc., M.Sc., and Ph.D. de-
g ees om Uni e si a Poli ècnica de Ca alunya (UPC),
Spain. Cu en ly, he is a Ramón y Cajal Fellow a UPC
Ba celona. P io o joining UPC, he spen 5 yea s as a
Senio Resea che wi h he Ba celona Supe compu ing Cen-
e (BSC), Spain, and 15 mon hs as a pos -doc o al ellow
wi h he In e na ional Compu e Science Ins i u e (ICSI),
Be keley. His esea ch in e es s include high pe o mance
compu e a chi ec u es, domain-speci ic accele a o s and
ha dwa e-so wa e co-design o u u e massi ely pa allel
sys ems.