scieee Open visual document viewer

GenArchBench: A genomics benchmark suite for arm HPC processors

López Villellas, Lorien,Langarita Benítez, Rubén,Badouh, Asaf,Soria Pardos, Víctor,Aguado Puig, Quim,López Paradís, Guillem,Doblas Font, Max,Setoain, Javier,Kim, Chulho,Ono, Makoto,Armejach Sanosa, Adrià,Marco Sola, Santiago,Alastruey Benedé, Jesús,Ibáñe

Abstract

Arm usage has substantially grown in the High-Performance Computing (HPC) community. Japanese supercomputer Fugaku, powered by Arm-based A64FX processors, held the top position on the Top500 list between June 2020 and June 2022, currently sitting in the fourth position. The recently released 7th generation of Amazon EC2 instances for compute-intensive workloads (C7 g) is also powered by Arm Graviton3 processors. Projects like European Mont-Blanc and U.S. DOE/NNSA Astra are further examples of Arm irruption in HPC. In parallel, over the last decade, the rapid improvement of genomic sequencing technologies and the exponential growth of sequencing data has placed a significant bottleneck on the computational side. While most genomics applications have been thoroughly tested and optimized for x86 systems, just a few are prepared to perform efficiently on Arm machines. Moreover, these applications do not exploit the newly introduced Scalable Vector Extensions (SVE). This paper presents GenArchBench, the first genome analysis benchmark suite targeting Arm architectures. We have selected computationally demanding kernels from the most widely used tools in genome data analysis and ported them to Arm-based A64FX and Graviton3 processors. Overall, the GenArch benchmark suite comprises 13 multi-core kernels from critical stages of widely-used genome analysis pipelines, including base-calling, read mapping, variant calling, and genome assembly. Our benchmark suite includes different input data sets per kernel (small and large), each with a corresponding regression test to verify the correctness of each execution automatically. Moreover, the porting features the usage of the novel Arm SVE instructions, algorithmic and code optimizations, and the exploitation of Arm-optimized libraries. We present the optimizations implemented in each kernel and a detailed performance evaluation and comparison of their performance on four different HPC machines (i.e., A64FX, Graviton3, Intel Xeon Skylake Platinum, and AMD EPYC Rome). Overall, the experimental evaluation shows that Graviton3 outperforms other machines on average. Moreover, we observed that the performance of the A64FX is significantly constrained by its small memory hierarchy and latencies. Additionally, as proof of concept, we study the performance of a production-ready tool that exploits two of the ported and optimized genomic kernels.

Full text

Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329 A ailable online 2 Ap il 2024 0167-739X/© 2024 The Au ho (s). Published by Else ie B.V. This is an open access a icle unde he CC BY-NC license (h p://c ea i ecommons.o g/licenses/by- nc/4.0/). Con en s lis s a ailable a ScienceDi ec Fu u e Gene a ion Compu e Sys ems jou nal homepage: www.else ie .com/loca e/ gcs GenA chBench: A genomics benchma k sui e o a m HPC p ocesso s Lo ién López-Villellasd,∗,1,Rubén Langa i a-Bení eza,Asa Badouha,Víc o So ia-Pa dosa, Quim Aguado-Puigb,Guillem López-Pa adísa,Max Doblasa,Ja ie Se oaine,Chulho Kim , Mako o Onog,Ad ià A mejacha,c,San iago Ma co-Solaa,c,Jesús Alas uey-Benedéd, Pablo Ibáñezd,Miquel Mo e óa,c aBa celona Supe compu ing Cen e , Ba celona, Spain bDepa men d’A qui ec u a de Compu ado s, Uni e si a Au ònoma de Ba celona, Ba celona, Spain cDepa men d’A qui ec u a de Compu ado s, Uni e si a Poli ècnica de Ca alunya, Ba celona, Spain dDepa amen o de In o má ica e Ingenie ía de Sis emas/A agón Ins i u e o Enginee ing Resea ch (I3A), Uni e sidad de Za agoza, Za agoza, Spain eA m Resea ch, Camb idge, Uni ed Kingdom Leno o Resea ch, Uni ed S a es gLeno o In as uc u e Solu ions G oup, Uni ed S a es ARTICLE INFO Keywo ds: Genomics A m High-pe o mance compu ing Pa allel compu ing Vec o compu ing Pe o mance cha ac e iza ion ABSTRACT A m usage has subs an ially g own in he High-Pe o mance Compu ing (HPC) communi y. Japanese supe - compu e Fugaku, powe ed by A m-based A64FX p ocesso s, held he op posi ion on he Top500 lis be ween June 2020 and June 2022, cu en ly si ing in he ou h posi ion. The ecen ly eleased 7 h gene a ion o Amazon EC2 ins ances o compu e-in ensi e wo kloads (C7 g) is also powe ed by A m G a i on3 p ocesso s. P ojec s like Eu opean Mon -Blanc and U.S. DOE/NNSA As a a e u he examples o A m i up ion in HPC. In pa allel, o e he las decade, he apid imp o emen o genomic sequencing echnologies and he exponen ial g ow h o sequencing da a has placed a signi ican bo leneck on he compu a ional side. While mos genomics applica ions ha e been ho oughly es ed and op imized o x86 sys ems, jus a ew a e p epa ed o pe o m e icien ly on A m machines. Mo eo e , hese applica ions do no exploi he newly in oduced Scalable Vec o Ex ensions (SVE). This pape p esen s GenA chBench, he i s genome analysis benchma k sui e a ge ing A m a chi ec u es. We ha e selec ed compu a ionally demanding ke nels om he mos widely used ools in genome da a analysis and po ed hem o A m-based A64FX and G a i on3 p ocesso s. O e all, he GenA ch benchma k sui e comp ises 13 mul i-co e ke nels om c i ical s ages o widely-used genome analysis pipelines, including base-calling, ead mapping, a ian calling, and genome assembly. Ou benchma k sui e includes di e en inpu da a se s pe ke nel (small and la ge), each wi h a co esponding eg ession es o e i y he co ec ness o each execu ion au oma ically. Mo eo e , he po ing ea u es he usage o he no el A m SVE ins uc ions, algo i hmic and code op imiza ions, and he exploi a ion o A m-op imized lib a ies. We p esen he op imiza ions implemen ed in each ke nel and a de ailed pe o mance e alua ion and compa ison o hei pe o mance on ou di e en HPC machines (i.e., A64FX, G a i on3, In el Xeon Skylake Pla inum, and AMD EPYC Rome). O e all, he expe imen al e alua ion shows ha G a i on3 ou pe o ms o he machines on a e age. Mo eo e , we obse ed ha he pe o mance o he A64FX is signi ican ly cons ained by i s small memo y hie a chy and la encies. Addi ionally, as p oo o concep , we s udy he pe o mance o a p oduc ion- eady ool ha exploi s wo o he po ed and op imized genomic ke nels. 1. In oduc ion Fo many yea s, A m p ocesso s ha e domina ed he mobile de ice segmen . Thei ene gy e iciency and license-based business model ha e been he pilla s unde pinning his success. ∗Co espondence o: Depa amen o de In o má ica e Ingenie ía de Sis emas/A agón Ins i u e o Enginee ing Resea ch (I3A), Uni e sidad de Za agoza, Spain. E-mail add ess: [email p o ec ed] (L. López-Villellas). 1The co esponding au ho conduc ed his wo k while a ilia ed wi h he Ba celona Supe compu ing Cen e . In ecen yea s, A m has bu s on o he high-pe o mance compu ing ma ke wi h in luen ial companies and conso iums ha ha e become licensees, such as Fuji su, Amazon, Apple, NVIDIA, Samsung, AMD, B oadcom, HUAWEI, and Qualcomm. Cu en ly, he A m-based Fuji su h ps://doi.o g/10.1016/j. u u e.2024.03.050 Recei ed 3 No embe 2023; Recei ed in e ised o m 25 Ma ch 2024; Accep ed 31 Ma ch 2024 Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329 314 L. López-Villellas e al. A64FX p ocesso powe s he Japanese supe compu e Fugaku, which held he op posi ion on he Top500 lis be ween June 2020 and June 2022 and is cu en ly in he ou h posi ion. Mo eo e , Amazon has been using A m p ocesso s o powe i s cloud compu ing pla o m (AWS), s a ing in 2018 wi h he G a i on p ocesso . They ollowed wi h he second gene a ion o G a i on in 2019 and he ecen ly eleased G a i on3. In he nea u u e, NVIDIA G ace CPUs and Ampe e se e s will be leading u he e o s o b eak h ough A m in HPC. As a esul , la ge-scale compu ing in as uc u es, usually equipped wi h x86 and IBM Powe p ocesso s, now ha e an addi ional compe i i e al e na i e. Howe e , mos o he scien i ic code o HPC is no ully adap ed and op imized o A m a chi ec u es. O e he las decade, genome sequencing has become he co ne - s one o genomics and mode n p ecision medicine. Due o he apid imp o emen o sequencing echnologies, i is cu en ly possible o sequence an indi idual’s genome in less han 24 h. This b eak h ough has enabled e ec i e pe sonalized heal hca e, allowing he diagnosis and ea men o diseases based on each pe son’s unique genomic disposi ion [1]. Fu he mo e, genome sequencing has also been p o en c ucial in cance s udies [2], d ug de elopmen [3], o COVID-19 ou - b eak con ol [4]. In he pas 20 yea s, genome sequencing cos s ha e d opped d ama ically and he amoun o sequencing da a p oduced yea ly has inc eased exponen ially. Mo e no ably, his inc ease in da a p oduc ion has ou pe o med he pace o Moo e’s law. As a esul , a signi ican bo leneck in cu en genome sequencing analysis is placed on he compu a ional side, execu ing compu a ional-in ensi e genomics ools and pipelines. Genome analysis pipelines ha e his o ically been designed o un e icien ly on x86 a chi ec u es. Wi h he i up ion o A m-based HPC se e s, adap ing and op imizing genomics ools o exploi HPC A m a chi ec u es e ec i ely has become pa amoun . Fo ha , we ha e selec ed 13 compu a ionally-demanding CPU ke nels om he mos widely-used genomics ools, and we ha e included hem in a bench- ma k sui e called GenA chBench. All he ke nels exploi mul i-co e pa allelism and implemen common s ages om widely-used genome analysis pipelines such as base-calling, ead mapping, a ian call- ing, and de-no o assembly. Addi ionally, GenA chBench includes inpu da ase s o each ke nel (i.e., a small da ase and a la ge da ase pe ke nel) and hei co esponding ou pu s o be used as g ound u h. The small da ase s ha e been sized o equi e single- h ead execu ion imes no longe han a ew minu es ( o es ing pu poses); meanwhile, la ge da ase s equi e se e al minu es ( o pe o mance e alua ion pu poses). Fo con enience, we p o ide au oma ic eg ession es s o all he ke nels o e i y he co ec ness o he ou pu s. Fu he mo e, his wo k in oduces code adap a ions and op imiza- ions o he genomics ke nels a ge ing A m HPC CPUs. GenA chBench le e ages A m-speci ic HPC lib a ies (ca e ully op imized o A m p o- cesso s) and p esen s algo i hmic and code op imiza ions o exploi he a chi ec u e and esou ces o A m HPC machines. No ably, we ha e op imized some ke nels by u ilizing he la es A m Scalable Vec o Ex ensions (SVE) o le e age he po en ial o he la es A m HPC p ocesso s. In addi ion o he benchma k sui e po ing and op imiza ion, his wo k p esen s a pe o mance cha ac e iza ion o GenA chBench on ou HPC machines ( wo A m-based and wo x86-based nodes). The expe imen al e alua ion compa es he pe o mance o an A64FX p o- cesso , a G a i on3 p ocesso , an In el Xeon Skylake Pla inum 8160 p ocesso , and an AMD EPYC 7742 Rome p ocesso . This cha ac e i- za ion includes he ke nels’ ins uc ion b eakdown, single- h ead and mul i- h ead pe o mance e alua ions, a mic oa chi ec u e bo leneck analysis, and an ene gy- o-solu ion s udy in he di e en p ocesso s. Ul ima ely, we e alua e he pe o mance impac o hese op imiza ions by in eg a ing wo o he accele a ed ke nels in a p oduc ion- eady ool used in a my iad o genome analysis pipelines. In summa y, his wo k makes he ollowing con ibu ions: •We p esen GenA chBench, he i s benchma k sui e a ge ing A m HPC a chi ec u es o genome analysis pipelines and ools. The benchma k sui e is publicly a ailable a h ps://gi hub.com/ Lo ienLV/gena chbench/ eleases/ ag/1.0.0. •We p opose HPC adap a ions and code op imiza ions applied o GenA chBench’s ke nels o exploi he po en ial o A m HPC p o- cesso s, le e aging A m-speci ic HPC lib a ies and A m Scalable Vec o Ex ension (SVE). •We pe o m a comp ehensi e pe o mance cha ac e iza ion o GenA chBench in wo HPC A m p ocesso s (i.e., A64FX and G a i on3). We compa e he pe o mance o A m agains wo e e ence HPC x86 machines. 2. Backg ound Genome da a analysis pipelines comp ise mul iple s ages and com- pu a ional ools, om sequencing biological samples o de i ing mean- ing ul da a analysis esul s o scien is s and heal hca e p o essionals. This sec ion in oduces he main sequencing echnologies, pipelines, and ools used in common genome analysis (Fig. 1 shows a succinc g aphic summa y). 2.1. Sequencing echnologies Be o e any compu a ional analysis can be pe o med, biological DNA samples mus be con e ed o digi al da a. This p ocess is pe - o med by he sequencing machines (Fig. 1-1), and, despi e he ema k- able ad ances in he las decades, hese machines a e s ill unable o ead a comple e DNA molecule om end o end. Ins ead, sequencing machines allow eading ela i ely small chunks o DNA, called eads o agmen s, om andom loca ions wi hin he dono ’s DNA genome. A e wa ds, sequenced eads mus be jigsaw oge he o econs uc o eassemble he o iginal dono ’s genome. Sequencing machines a e commonly ca ego ized in o h ee gene a- ions based on hei echnological ad ancemen s. The i s sequencing echnologies (Sange e al. [5] and Maxam e al. [6]) we e de eloped in 1977 and used o sequence he i s d a o he human genome in 2000 [7]. Since hen, sequencing echnologies ha e e ol ed quickly, simpli ying he sequencing p ocess and inc easing he da a-p oduc ion h oughpu . In he mid-2000s, second-gene a ion echnologies [8] we e in oduced and soon eplaced i s -gene a ion echnologies. Second- gene a ion echnology can gene a e ixed-leng h sequences o 100–300 bps a a h oughpu o ens o gigaby es pe hou and wi h a low eading e o a e (0.1% o he ead leng h). A p esen , Illumina domina es he ma ke o second-gene a ion sequencing machines. Recen ly in oduced hi d-gene a ion echnologies, known as long- ead sequencing, can ead a iable-leng h sequences o conside able leng h (i.e., ens o kilo base- pai s) a he expense o lowe p oduc ion h oughpu (less han 10 Gb/hou ) and highe eading e o a e (0.1%–10% o he ead leng h). Paci ic Biosciences (PacBio) and Ox o d Nanopo e Technologies (ONT) a e he mos no able manu ac u e s o hi d-gene a ion sequencing echnologies. 2.2. Genome da a analysis pipelines and ools Be o e any p ocessing can be pe o med, he sequencing machines’ aw signals mus be ans o med in o sequences o nucleo ides (A, C, G, T). This p ocess is called basecalling (Fig. 1-2). Typically, a specialized basecalling ool is used o pe o m his p ocess ailo ed o each sequencing echnology. Fo ins ance, Boni o [9] and Guppy [10] a e wo o he mos widely-used ools o basecalling Ox o d Nanopo e’s aw-signal ou pu . Once he sequences o nucleo ides a e decoded, sequenced eads mus be p ocessed and analyzed o de i e meaning ul biological in- sigh s. Al hough many di e en genome analyses can be pe o med using sequenced da a, mos analyses begin wi h ei he genome ese- quencing (1-3.a) o genome assembly (1-3.b). Bo h analyses seek o econs uc he sample’s genome by pu ing oge he all he sequenced eads. Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329 315 L. López-Villellas e al. Fig. 1. Wo k low diag am o common genome analysis pipelines. Going om (1) sequencing, h ough (2) basecalling, o (3.a) genome esequencing o (3.b) and genome assembly. The igu e shows he di e en compu a ional ke nels used wi hin each s age o ool. 2.2.1. Genome esequencing The mos common app oach o econs uc ing he sample’s genome is by esequencing and in ol es econs uc ing he sample’s genome using a p e iously known e e ence genome. Fo ha , each sequenced ead is loca ed and ma ched o he mos likely o igina ing posi ion in he e e ence genome, allowing small di e ences (e.g., misma ches, inse ions, and dele ions). This p ocesses is called ead mapping (Fig. 1- 3.a.1) and i is implemen ed by many ools like BWA-MEM2 [11,12], Minimap2 [13], Bow ie2 [14,15], and GEM [16]. Read mapping is one o he mos compu a ionally expensi e s eps in all genome sequence analyses. Consequen ly, ead mapping has been ex ensi ely s udied and op imized. Mos sequence mappe s a e based on he seed-chain-ex end ech- nique. This echnique implemen s h ee algo i hmic s eps o swi ly loca e and align a sequence wi h a e e ence genome. Du ing he i s s ep, known as seeding (Fig. 1-3.a.1.1), he mappe sea ches small subsequences o he eads (seeds) in he e e ence le e aging an index s uc u e. The mos widely-used indexes used o seeding a e FM- Index [17] and hash- ables [18,19]. Seeding educes he po en ial num- be o loca ions in he e e ence whe e a sequence can ma ch, dec eas- ing he amoun o wo k pe o med in subsequen s eps. A e wa ds, a chaining s ep (Fig. 1-3.a.1.2) is pe o med o educe u he he lis o possible ma ching loca ions in he e e ence. Du ing he chaining s ep, all he mapped seeds a e p ocessed o ind a colinea chain o seeds ha can po en ially ma ch he inpu sequence. Finally, du ing he ex- ensión o alignmen s ep (Fig. 1-3.a.1.3), he inpu sequence is aligned agains he candida e loca ion in he e e ence genome, disco e ing he di e ences be ween he dono ’s sequence and he e e ence genome. Usually, a dynamic p og amming-based algo i hm, such as Needleman– Wunsch [20] o Smi h–Wa e man–Go oh [21,22], is used o compu e he alignmen . A e sequence mapping, once he eads a e loca ed in he e e - ence genome, a a ian calling algo i hm (Fig. 1-3.a.2) de e mines he a ian s and mu a ions be ween he dono ’s genome and he e e - ence genome. These a ia ions p o ide c ucial insigh s in o he ge- ne ic makeup o he sequenced indi idual, po en ially e ealing ge- ne ic a ia ions ha may be associa ed wi h diseases and heal h con- di ions. No able examples o widely-used a ian calle s a e GATK Haplo ype-Calle [23], Pla ypus [24], Clai [25,26], DeepVa ian [27] and Medaka [28]. 2.2.2. Genome assembly Despi e he simplici y and e ec i eness o genome esequencing, he e is s ill a lack o high-quali y e e ence genomes o many species. In hose si ua ions, genome de-no o assembly (Fig. 1-3.b) is used o econs uc he dono ’s genome om sc a ch jigsawing he sequenced eads oge he . Mos popula de-no o assembly me hods ely on de B uijn g aphs. Fo a gi en se o sequences, i s co esponding de B uijn g aph con ains a node pe each sequence’s k-me (i.e., sub-s ing o leng h 𝑘nu- cleo ides) and an edge ha connec s adjacen and o e lapping k-me s. Be o e cons uc ing he de B uijn g aph o a se o inpu sequences, he numbe o unique k-me s in he eads is coun ed (Fig. 1-3.b.1) o p une he leas equen ones (likely a i ac s o he sequencing p ocess). A e wa ds, he de B uijn g aph is cons uc ed (Fig. 1-3.b.2). Then, he consensus sequence is de i ed using mul iple sequence alignmen (MSA) algo i hms (Fig. 1-3) and he cons uc ed de B uijn g aph. No able examples o de B uijn g aph based assemble s a e Flye [29], Canu [30], and Racon [31]. 2.2.3. Me agenomics Beyond genome esequencing, a ian calling, and de-no o assem- bly, many p e iously desc ibed analysis s eps and ools can be ound in o he genome analysis pipelines. This is he case o many me age- nomics analysis pipelines. Me agenomics pipelines seek o analyze genomic in o ma ion om mixed mic obial communi ies, p o iding insigh s in o he di e si y, in e ac ions and unc ion o mic oo gan- isms p esen in an en i onmen al sample. Me agenomics analyses a e pe o med using ools such as Cen i uge [32], RawMap [33], UN- CALLED [34], ReadFish [35], K aken2 [36] and Cla k [37]. These ools employ k-me coun ing (Fig. 1-3.b.1) and seeding echniques (Fig. 1-3.a.1.1) o hei analysis. Mo eo e , a ian calle s like GATK Haplo ype Calle [23] and Pla ypus [24] a e used o cons uc De B uijn g aphs (Fig. 1-3.b.2) and co ec a i ac s p oduced du ing he map- ping p ocess (Fig. 1-3.a.1). Fu he mo e, he chaining p ocess (Fig. 1- 3.a.1.2) is also u ilized o genome assembly when using al e na i e app oaches based on de B uijn g aphs [38]. 3. GenA ch benchma k sui e The GenA ch benchma k sui e comp ises 13 mul i h eaded CPU ke nels de i ed om he mos widely used genomics ools and co e s Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329 316 L. López-Villellas e al. he mos impo an genome sequencing s eps. I includes en ke nels om he GenomicsBench [39] benchma k sui e and h ee addi ional ke nels: he Bi -Pa allel Mye s algo i hm [40] (BPM), he Wa e on Alignmen algo i hm [41] (WFA), and FAST-CHAIN [42]. BPM and WFA complemen he sequence alignmen ke nels o GenomicsBench o be e cap u e con empo a y ends. Addi ionally, FAST-CHAIN [42] is a ecen ec o -enabled eimplemen a ion o he CHAIN ke nel p esen in GenomicsBench, which allows us o u he explo e he capabili ies o SVE. Addi ionally, GenA chBench includes inpu da ase s o each ke - nel (i.e., a small da ase and a la ge da ase pe ke nel) and hei co esponding ou pu s o be used as g ound u h. The small da ase s ha e been sized o equi e single- h ead execu ion imes no longe han a ew minu es ( o es ing pu poses); meanwhile, la ge da ase s equi e se e al minu es ( o pe o mance e alua ion pu poses). Fo con enience, we p o ide au oma ic eg ession es s o all he ke nels o e i y he co ec ness o he ou pu s. Al hough some ke nels included in GenA chBench can exploi he capabili ies o mode n GPUs, his esea ch ocuses on po ing, accel- e a ing, and e alua ing he pe o mance o genomics ke nels in A m p ocesso s. Mo eo e , he A m-sys ems e alua ed in his wo k (A64FX and G a i on3) a e no equipped wi h GPUs. The ollowing ex p esen s GenA chBench’s ke nels, b ie ly desc ib- ing i s unc ionali y, which ools use hem, and a desc ip ion o hei usage and inpu s. Adap i e Banded Signal o E en Alignmen (ABEA): ABEA is a dynamic p og amming algo i hm ha compa es aw nanopo e sig- nals om ONT sequencing machines o a e e ence genome sequence. ABEA’s implemen a ion is based on he Suzuki–Kasaha a (SK) [43] al- go i hm. This s ep is pe o med in some ools, such as Nanopolish [44], o co ec e o s p oduced in he basecalling p ocess (Fig. 1-2). Fo GenA chBench, we ha e used he CPU implemen a ion o 5c [45], a e sion o ABEA based on Nanopolish’s, op imized o bo h CPU- only and hyb id CPU/GPU execu ions. This implemen a ion o ABEA exploi s coa se-g ain mul i- h eading by di iding he aw signals o he inpu be ween he a ailable co es. Since he signals a e no o egula size, 5c implemen s wo k-s ealing o imp o e load balance. The small and la ge inpu s comp ise 1K and 10K aw FAST5 (ONT) eads om ch omosome 22 o NA12878 and GRCh38 as he e e ence genome [46]. Bi -Pa allel Mye s (BPM): BPM [40] is a dynamic p og amming algo i hm ha inds all loca ions a que y s ing o size 𝑚ma ches a e e ence s ing o size 𝑛wi h 𝑘o ewe di e ences (Fig. 1-3.a.1.3). I compu es he app oxima e s ing ma ching o wo s ings in 𝑂(𝑚𝑛∕𝑤) ime, whe e 𝑤is he wo d size o he machine. BPM is used in ead map- ping ools, such as GEM-Mappe [16], Edlib [47], G aphAligne [48] o Hobbes [49]. Fo GenA chBench, we ha e used an in-house imple- men a ion o he algo i hm ha exploi s mul i- h eading by assigning di e en pai s o s ings o di e en h eads. The small and la ge inpu s comp ise 100K and 10M sequence pai s om human sample SRR7733443 downloaded om he sequence ead a chi e [50]. Banded Smi h–Wa e man (BSW): The Smi h–Wa e man algo i hm [21] is a dynamic p og amming algo i hm ha compu es he local sequence alignmen o wo sequences o leng h 𝑚and 𝑛, espec i ely, in 𝑂(𝑚𝑛) ime and space. A banded e sion o Smi h–Wa e man [51] is used o align sequences wi h a maximum o 𝑤inse ions/dele ions, educing he ime and space complexi y o 𝑂(𝑤𝑛)(Fig. 1-3.a.1.3). BSW is used in a ian disco e y ools such as GATK [23], and in sequence alignmen so wa e like BWA-MEM [11,12]. Fo GenA ch- Bench, we ha e used BWA-MEM2’s x86- ec o ized implemen a ion o BSW. In o de o exploi mul i- h eading, he se o pai s o s ings o align is dynamically di ided be ween p ocesso s. The small and la ge inpu s comp ise 100K and 10M sequence pai s om human sample SRR7733443 [50]. Seed Chaining (CHAIN): Gi en he se o seeds om a DNA se- quence ( ead) mapped o ano he sequence, such as he e e ence genome, he chaining s ep (Fig. 1-3.a.1.2) aims o ind a chain o colinea seeds. This is a ime-consuming s ep pe o med by alignmen ools, such as Minimap2, and by de-no o assemble s like Flye [29] o Canu [30]. We ha e used he implemen a ion o CHAIN ound in GenomicsBench ha ex ends Minimap2’s o exploi in e - ask pa al- lelism ac oss eads. The small and la ge inpu s comp ise he seeds om 1K, and 10K eads o Pacbio’s Caeno habdi is elegans wo m sequence da a [52]. SIMD Seed Chaining (FAST-CHAIN): The p e iously p esen ed implemen a ion o he CHAIN algo i hm u ilizes heu is ics o s op execu ing when he esul is su icien ly good. This speedups execu ion a he cos o accu acy, and i hinde s he ec o iza ion o he ke nel. FAST-CHAIN [42] is an x86- ec o ized e sion o CHAIN ha emo es he heu is ics o exploi SIMD compu a ion. As a esul , FAST-CHAIN ou pu s accu a e esul s and p esen s pe o mance gains compa ed o CHAIN. FAST-CHAIN uses he same inpu s as CHAIN. De B uijn G aph Cons uc ion (DBG): The De B uijn g aph (DBG) o an inpu se o eads is used o ep esen he o e laps be ween he sub-s ings o leng h 𝑘(k-me s) ound in he inpu (Fig. 1-3.b.2). Each node o he g aph ep esen s a k-me and he edges connec adjacen k-me s in he inpu se . The cons uc ion o hese g aphs is a ime- consuming s ep in de-no o assemble s like Flye [29], Canu [30] o Racon [31], and in a ian calle s such as GATK [23] and Pla ypus [24]. Fo GenA chBench, we ha e used he DBG cons uc ion o Pla ypus, which exploi s pa allelism by assigning di e en egions o he inpu o di e en h eads. Bo h inpu s employ ch omosome 22 o BWA-MEM aligned eco ds om he Pla inum Genomes da ase [53]. The small inpu uses bases 16M-16.5M, while he la ge inpu uses he en i e ch omosome. FM-Index Sea ch (FMI): The FM-index is a comp essed sub-s ing index based on he Bu ows–Wheele ans o m [54]. Gi en a sub-s ing 𝑠, FM-index can be used o ind he loca ion o 𝑠in he e e ence genome in 𝑂(|𝑠|) ime, whe e |𝑠|is he leng h o he sub-s ing (Fig. 1- 3.a.1.1). The FM-index da a s uc u e is used in sequence alignmen ools such as BWA-MEM [11,12] o Bow ie2 [15], and in me agenomic classi ica ion so wa e like Cen i uge [32]. Fo GenA chBench, we ha e used he supe -maximal exac ma ch ke nel o BWA-MEM2, which u ilizes he FM-Index s uc u e. This ke nel exploi s pa allelism by dynamically assigning ba ches o eads among h eads. The small and la ge inpu s comp ise 1M and 10M pai s o 151 bases om human sample SRR7733443 [50]. K-me Coun ing (KMER-CNT): K-me coun ing aims o coun he numbe o occu ences o each k-me in an inpu sequence (Fig. 1- 3.b.1). This ask is pe o med in de-no o assemble s such as Flye [29] o Canu [30] and in me agenomics classi ica ion so wa e like Cla k [37]. Addi ionally, no e ha he unc ionali y o KMER-CNT is e y simila o accessing la ge lookup ables, as done in s a e-o - he-a map- pe s like Minimap2 [13]. Fo GenA chBench, we ha e used he k-me coun ing ke nel o Flye. This implemen a ion di ides he inpu - eads be ween h eads and elies on he h ead-sa e hash-map implemen a- ion o Libcuckoo lib a y [55] o concu en ly inc ease he numbe o indi idual k-me s shown by each h ead. The small and la ge inpu s comp ise 1K and 50K Esche ichia coli Ox o d Nanopo e eads sequenced by Loman Labs [56]. Neu al Ne wo k-based Base Calling (NN-BASE): ONT sequencing machines moni o changes in an elec ical cu en as single s ands o DNA o RNA pass h ough a p o ein nanopo e. These changes in he elec ical cu en a e hen con e ed o a sequence o nucleo ide bases in he basecalling p ocess (Fig. 1-2). The analog signal ine i ably con- ains ambigui ies due o noise o measu emen e o s. Some basecalle s, such as Guppy [10] and Boni o [9], ely on neu al ne wo ks o sol e hese ambigui ies, de e mining he mos likely obse ed nucleo ide in each pa o he elec ical cu en . Fo GenA chBench, we ha e used Boni o’s deep-lea ning base-calle (NN-BASE), which depends on he PyTo ch lib a y [57]. Boni o spli s he inpu signal in o smalle chunks o egula size and eeds hem o a PyTo ch neu al ne wo k ha Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329 317 L. López-Villellas e al. Table 1 Cha ac e is ics o e iew o he expe imen al se up. A64FX G a i on3 SKX Rome Co es 4 ×12 (+ 4 assis an ) 64 2 ×24 64 SMT No No Disabled Disabled F equency 2.2 GHz (s a ic) 2.6 GHz 1–2.1 GHz (dynamic) 1.5–2.25 GHz (dynamic) Max. powe 120 W N/A 2 ×150 W 225 W Mem. capaci y 4 ×8 GB 8 ×16 GB 2 ×6×8 GB 16 ×64 GB Mem. echnology on-package HBM2 o -package DDR5 4800 MHz o -package DDR4 2667 MHz o -package DDR4 3200 MHz Peak bandwid h 4 ×256 GB/s 300 GB/s 2 ×120 GB/s 204.8 GB/s L1i 64 KB (4-way) 64 KB 32 KB (8-way) 32 KB (8-way) L1d 64 KB (4-way) 64 KB 32 KB (8-way) 32 KB (8-way) L2 – 1 MB 1 MB (16-way) 512 KB (8-way) LLC 4 ×8 MB (16-way) 32 MB 2 ×33 MB (11-way) 16 ×16 MB (16-way) Vec o ex ension NEON/SVE 512 bi s NEON/SVE 256 bi s SSE/AVX2/AVX512 SSE/AVX2 in e nally exploi s mul i- h eading. The small and la ge inpu s comp ise 1 and 10 aw FAST5 eads om ch omosome 20 o NA12878, ob ained om he Nanopo e WGS Conso ium [46]. Neu al Ne wo k-based Va ian Calling (NN-VARIANT): Va ian calling is he p ocess o de ec ing he di e ences ( a ian s o mu a ions) be ween he aligned eads and he e e ence genome (Fig. 1-3.a.2). This is a cos ly p ocess pe o med by s a is ics-based a ian calle s, such as GATK Haplo ypeCalle [23] o Pla ypus [24], and deep-lea ning a ian calle s, such as Clai [25,26], DeepVa ian [27] o Medaka [28]. Fo GenA chBench, we ha e used he second gene a ion o Clai a ian calle (Clai 3), based on he Tenso Flow amewo k [58]. Clai 3 ex- ploi s pa allelism by di iding he inpu in o egula -size chunks, and each o hese chunks is p ocessed by one h ead using Tenso Flow. Ou small and la ge inpu s comp ise 100K and 10M e e ence posi ions, espec i ely, o ch omosome 20 o HG002 om NITS’s Genome in a Bo le (GIAB) p ojec [59]. We a e using Clai 3’s ONT p e- ained model 941_p om_hac_g360+g422 [60]. Pileup Coun ing (PILEUP): Gi en he alignmen da a o a se o aligned eads o a egion o a e e ence genome, usually a SAM o BAM ile [61], pileup coun ing is he p ocess o summa izing he base-pai in- o ma ion a each ch omosomal posi ion. This summa y, called pileup, is cus oma y he inpu o long- ead neu al ne wo k a ian calle s such as Clai [25,26] o Meda aka [28] (Fig. 1-3.a.2). Fo GenA chBench we ha e used he pileup coun ing implemen a ion o Medaka, which exploi s mul i- h ead pa allelism by dis ibu ing 100 kilobase egions o he e e ence genome be ween h eads. The small inpu comp ises bases 1-1499707 o he S aphylococcus au eus genome [10], and he la ge inpu comp ises bases 1-1412827 o ch omosome 20 o sample HG002 [59]. Pa ial-O de Alignmen (POA): The cons uc ion o an o e lap g aph om a se o eads leads o an app oxima e ep esen a ion o he o iginal sample’s genome. To de e mine he consensus genome o he sample, he alignmen o all he eads agains each o he is pe o med in a p ocess called mul iple sequence alignmen (MSA) (Fig. 1-3.b.3). The Pa ial O de ed Alignmen (POA) algo i hm [62] compu es he MSA o all sequences by inc emen ally cons uc ing a pa ially-o de g aph aligning new sequences o i using a dynamic p og amming algo i hm such as Smi h–Wa e man [21] o Needleman–Wunsch [20]. The mul iple alignmen sequence (consensus sequence) is in e ed om he g aph by using he Hea ies Bundle algo i hm [63]. POA is used in so wa e packages such as Nanopolish [44] o Racon [31]. Fo Gena chBench we ha e used he SIMD-op imized e sion o POA o he SPOA lib a y [64]. SPOA exploi s mul i- h eading by compu ing he pa ially-o de ed g aph o mul iple se s o sequences in pa allel. The small and la ge inpu s comp ise 1K and 6K se s o mul iple sequences aligned o a e e ence genome, each con aining be ween 5 and 115 sequences. This da a comes om Minimap2’s polishing s ep o he Flye-assembled S aphylococcus Au eus genome [10]. Wa e on Alignmen (WFA): The wa e on alignmen algo i hm (WFA) [41] is a pai wise alignmen algo i hm (Fig. 1-3.a.1.3) ha akes ad an age o homologous egions be ween he sequences o accele a e he alignmen p ocess. As opposed o adi ional dynamic p og amming algo i hms ha un in quad a ic ime, WFA ime complexi y is 𝑂(𝑛𝑠), p opo ional o he ead leng h 𝑛and he alignmen sco e 𝑠, using 𝑂(𝑠2)memo y. The wa e on algo i hm is used in ools such as w - mash [65], Ancho Wa e [66] o Ances alClus [67]. GenA chBench uses a cus om mul i- h ead implemen a ion o he algo i hm, in which each h ead wo ks in he alignmen o a pai o s ings. The small and la ge inpu s comp ise 100K and 1M sequence pai s om human sample SRR7733443 [50]. 4. Expe imen al se up Ou expe imen al se up consis s o wo A m and wo x86 HPC sys ems: a compu e node ea u ing an A m-A64FX p ocesso (A64FX), a c7 g.16xla ge Amazon-EC2 ins ance (G a i on3), a sys em wi h wo x86-64 In el Xeon Skylake Pla inum 8160 (SKX), and a compu e node wi h one x86-64 AMD EPYC 7742 Rome p ocesso (Rome). Table 1 p esen s an o e iew o he main cha ac e is ics o he ou sys ems. In e ms o compu ing co es, he A64FX is based on ou Non- Uni o m Memo y Access (NUMA) domains wi hin he chip, also e- e ed o as co e memo y g oups (CMG). Each NUMA domain has 12 co es, plus one assis ance co e no used o gene al compu ing ( unning daemons, I/O, asynch onous MPI, e c.). In o al, he A64FX implemen s 48 compu ing co es. G a i on3 implemen s 64 co es in a single NUMA domain. The AMD Rome CPU comp ises 8 co e chiple s, known as co e cache dies (CCD), and a cen al I/O die ha con ols all he I/O and memo y unc ions o he chip. A CCD has wo co e complex (CCX) clus e s, each wi h 4 co es. Any pai o CCDs can communica e h ough he I/O die. SKX con ains wo NUMA chips, each wi h 24 physical co es. Rega ding ope a ional equency, G a i on3 p esen s he highes maximum equency among he sys ems wi h 2.6 GHz. The o he h ee sys ems’ maximum equency is e y simila , anging be ween 2.1 and 2.25 GHz. Bo h x86 sys ems dynamically adjus hei equency based on hei load. Addi ionally, SKX educes i s equency when execu ing AVX/AVX512 ins uc ions. In con as , he A64FX ope a es a a ixed equency se o 2.2 GHz. The e is no public in o ma ion abou adap i e equency ope a ion on G a i on3. Wi h espec o SIMD ex ensions, he A64FX is he i s CPU o implemen he A m 8.2-A Scalable Vec o Ex ension (SVE) [68]. One o SVE’s main ea u es is ha i is Vec o Leng h Agnos ic (VLA); ha is, he same bina y wo ks on a chi ec u es implemen ing ec o egis e s o di e en leng hs anging om 128 o 2048 bi s. The A64FX implemen s 512 bi s SVE egis e s. G a i on3 also implemen s SVE, wi h a ec o leng h o 256 bi s. The A64FX and G a i on3 also suppo he A m Neon SIMD ex ension, a non-VLA SIMD ISA ha wo ks wi h 128-bi ec o s. Bo h x86 sys ems implemen he SSE and AVX2 SIMD ex ensions, wi h a ec o leng h o 128 and 256 bi s, espec i ely. The SKX also suppo s he AVX512 ex ension, wi h a ec o leng h o 512 bi s. None o he x86 SIMD ex ensions a e VLA. Conce ning main memo y, each A64FX’s NUMA domain has i s own local on-chip 8 GB HBM2 main memo y and can access he o he h ee NUMA domains’ local memo ies ia a ing bus. G a i on3 is connec ed o 8 ×16 GB DDR5 channels, o a o al o 128 GB o memo y. Each Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329 318 L. López-Villellas e al. Table 2 Load- o-use memo y la encies in nanoseconds o he expe imen al se up. A64FX G a i on3 SKX Rome L1 2.3–5 1.5 1.9 1.8 L2 – 4.6 6.7 3.5 LLC 16.8–21.4 33.1 25.1 13.0 Main Mem. Local 118.2–126.4 153.5 86.2 121.5 Main Mem. Remo e 187.7–242.3 – 144.0 – chip o he SKX is connec ed o 6 ×8 GB DDR4 local channels and can access he o he chip’s local memo y. The Rome CPU is connec ed o 16 ×64 GB DDR4 channels, o aling 1 TB o memo y. The cache hie a chy o ganiza ion o he p ocesso s is ela i ely di e en . Bo h A m machines ha e wo 64 KB p i a e L1 caches pe co e (ins uc ions and da a), while he x86 CPUs ea u e wo 32 KB p i a e L1s pe co e. G a i on3 and SKX include one p i a e 1MB L2 cache pe co e, and Rome has one 512 KB p i a e L2 pe co e. The A64FX has one 8 MB las -le el cache (LLC) pe NUMA domain, G a i on3 includes one 32 MB LLC, SKX has wo 33 MB LLCs (one pe NUMA domain), and Rome includes one 16 MB LLC pe each 4-co e CCX. Conce ning memo y bandwid h, he A64FX is designed o achie e good pe o mance execu ing high memo y bandwid h-demanding ap- plica ions. The peak bandwid h o his chip (4 ×256 GB/s) is nea ly 3.5 imes highe han he peak bandwid h o G a i on3 (300 GB/s), he second sys em among he s udied in e ms o memo y h oughpu . I is ollowed by SKX, eaching up o 120 GB/s pe chip (240 GB/s in o al), and Rome holds he las posi ion wi h a peak bandwid h o 204.8 GB/s. Table 2 p esen s he memo y access la encies o each le el o he memo y hie a chy o all machines. All la encies on Rome and G a i on3 and la encies o emo e memo ies on he A64FX ha e been measu ed using he LMbench benchma k [69]. La encies o cache and local memo y on he A64FX ha e been ex ac ed om he mic o- a chi ec u e manual o he CPU. La encies on SKX ha e been measu ed using In el Memo y La ency Checke . The numbe o cycles o access he A64FX caches depends on he ype o ins uc ion: scala , loa ing- poin , sho SIMD, and la ge SIMD. The la encies o access he L1 on he sys ems ange om 1.5 ns (G a i on3) o 5 ns (la ge SIMD access on he A64FX). E en hough scala accesses on he A64FX a e as e (2.3 ns), i s ill p esen s he highes L1 access la ency. As p esen ed p e iously, he A64FX only implemen s wo le els o caches (L1 and LLC). The L2 access la encies o he o he sys ems ange be ween 3.5 ns (Rome) o 6.7 ns (SKX). Rome p esen s he as es access o i s LLC (13 ns), closely ollowed by he A64FX (16.8 ns o scala access and 21.4 ns o la ge SIMD access). The LLC access la ency on he SKX and G a i on3 is 25.1 and 33.1 ns, espec i ely. SKX p esen s he as es access la ency o local main memo y (86.2 ns), ollowed by he A64FX and Rome, wi h simila la encies (∼120 ns). G a i on3 has he highes local memo y access la ency, as expec ed om cu en DDR5 SDRAMs. Accessing emo e main memo ies in he A64FX akes be ween 187.7 ns (nea - emo e memo y) and 242.3 ns ( a - emo e memo y). Accessing he o he chip’s main memo y on he SKX machine akes 144 ns, 23% as e han A64FX’s bes case. The ou -o -o de esou ces o he expe imen al se up a e p esen ed in Table 3. We assume ha G a i on3 implemen s he same esou ces as Neo e se V1 o non-publicly a ailable da a (ma ked wi h *). Una ail- able da a o nei he G a i on3 no Neo e se V1 is ep esen ed as N/A. The A64FX is igh in ou -o -o de esou ces compa ed wi h he o he h ee p ocesso s. The SKX and Rome ha e a simila numbe o physical egis e s, almos doubling he numbe o gene al-pu pose egis e s o he A64FX (180 s. 96) and implemen ing 30% mo e SIMD/FP egis e s han he A64FX (160 s. 128). The A64FX can issue up o 7 mic o- ope a ions (𝜇OP) pe cycle, G a i on3 can issue up o 15, 8 o SKX, and 11 o Rome. The A64FX and SKX a e capable o commi ing 4 mic o- ope a ions pe cycle. Howe e , SKX can me ge wo mic o-ope a ions Table 3 Ou -o -o de esou ces o he expe imen al se up. A64FX G a i on3 SKX Rome Gene al egis e s 96 N/A 180 180 SIMD/FP egis e s 128 N/A 168 160 Issue wid h 7 (𝜇OP) 15 (𝜇OP) 8 (𝜇OP) 11 (𝜇OP) Commi wid h 4 (𝜇OP) N/A 4-8 (𝜇OP) 8 (MOP) ROB (en ies) 128 256* 224 224 LB (en ies) 40 85* 72 44 SB (en ies) 24 90* 56 48 RS (en ies) 2 ×20 + 2×10 + 19 N/A 97 4 ×16 + 28 + 36 *Neo e se V1 CPU de aul s. in o one used mic o-ope a ion, inc easing i s heo e ical commi a e o 8 mic o-ope a ions. Rome can commi up o 8 mac o-ope a ions (MOP) – i.e., ALU, memo y, o me ged ALU/memo y ope a ion – pe cycle. The eo de bu e (ROB) o G a i on3 (256 en ies) is wice as big as he A64FX’s (128 en ies). SKX and Rome ha e an iden ical-size ROB (224 en ies). The sizes o he load bu e s (LB) and s o e bu e s (SB) o he CPUs a e ela i ely di e en . G a i on3 and SKX implemen he la ges LB, wi h 85 and 72 en ies, espec i ely. The LB o he A64FX has 40 en ies, and Rome implemen s a 44-en y LB. Simila ly, G a i on3 and SKX ha e he la ges SB (90 and 56 en ies, espec i ely). The A64FX implemen s a 24-en y SB, hal he size o Rome’s. Addi ion- ally, a s o e ins uc ion on he A64FX occupies one en y in bo h he load and he s o e bu e . While SKX implemen s a uni ied ese a ion s a ion (RS) wi h 97 en ies, bo h he A64FX and Rome ha e se e al smalle RS. The A64FX di ides i s ese a ion s a ion in o 2 ×20 en ies o 2 in ege , loa ing-poin , and SIMD pipelines, 2 ×10 en ies o 2 add ess calcula ion pipelines, and 19 en ies o he b anch pipeline. Rome’s ese a ion s a ion has 4 ×16 en ies o 4 in ege pipelines (scala +SIMD), 28 en ies o 3 add ess calcula ion pipelines, and 36 en ies o 4 loa ing-poin pipelines (scala +SIMD). 5. A m po ing o genomics ke nels Mos ke nels p esen ed in Sec ion 3 a ge x86 a chi ec u es and ha e no been ex ensi ely es ed no op imized o A m machines. Thus, i was expec ed ha some ke nels could un in o ailu es and e en gene a e inco ec esul s. To e i y he execu ion o he ke nels, we used he SKX sys em o compu e he co ec ou pu o all ke nels and inpu s (i.e., g ound u h). Fo ou expe imen s, we used he GNU compile (GCC) on G a i on3 ( 11.2.0), SKX ( 10.1.0), and Rome ( 10.2.0). On he A64FX, we used GCC ( 10.2.0) and he Fuji su Compile (FCC) ( 4.2.0b). Fo mos ke nels, FCC-compiled bina ies exhibi ed be e pe o mance. The Fu- ji su Compile implemen s wo compila ion modes: a adi ional mode (T ad) based on compile s o ea lie sys ems and a Clang mode based on Clang/LLVM. In all cases, we ob ained be e execu ion imes when compiling wi h FCC’s Clang mode. We lacked FCC-compiled e sions o key-op imized Py hon lib a ies. Fo hese easons, all he esul s p esen ed in his documen o he A64FX ha e been ob ained using he Clang mode o FCC, excluding he wo Py hon ke nels (NN-BASE and NN-VARIANT), whose lib a ies we e compiled using GCC. We compile all ke nels wi h a leas -O2 op imiza ion le el and enable CPU-speci ic op imiza ions: -ma ch=a m 8-a+s e on he A64FX, -mcpu=na i e on G a i on3 and -ma ch=na i e on SKX and Rome. Enabling CPU-speci ic op imiza ions in ABEA and POA esul ed in inco ec execu ions, p obably due o p og amming e o s in he o iginal sou ce code. The e o e, such op imiza ions a e no used o hese wo ke nels. A e pe o ming he app op ia e modi ica ions o he ke nels so all o hem success ully execu e on A m, we applied u he op i- miza ions o some ke nels o imp o e he pe o mance ob ained in his a chi ec u e. Such op imiza ions a e desc ibed in he ollowing subsec ions. Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329 319 L. López-Villellas e al. Fig. 2. Speedup o SIMD ke nels o e hei scala e sion on he expe imen al se up using he la ge inpu s. 5.1. Exploi ing ec o iza ion Some ke nels implemen x86- ec o ized e sions o hei mos ime- consuming pa s. In pa icula , BSW and FAST-CHAIN include AVX2 and AVX512 e sions o hei c i ical unc ions using in insics. Simi- la ly, POA implemen s SIMD e sions o i s code using AVX2-in insics and SIMD E e ywhe e (SIMDe). We ha e implemen ed SVE-in insics e sions o FAST-CHAIN, BSW, and WFA and a Neon-in insics e sion o BPM. SIMDe does no ully suppo SVE ye , so we could no le e age POA’s SIMDe e sion. Fig. 2 shows he speedup o ec o ized ke nels o e hei scala e sion on he expe imen al se up using he la ge inpu o he ke nels. No e ha he SVE ec o leng h o G a i on3 (256 bi s) is hal he A64FX’s (512 bi s), and he e o e he pe o mance speedups o SVE ke nels o e hei scala e sions a e mo e modes in G a i on3. BPM: The co e idea behind ec o izing BPM is o ans o m he alignmen ope a ions used o ill he dynamic p og amming able in o simple machine-wo d ope a ions. These simple ope a ions a e in ege addi ions, bi shi s, and bi wise ORs and ANDs. This way, a ious dynamic p og amming cells a e bi -packed wi hin a machine wo d and i s dependencies a e encoded using bi -wise ope a ions. In packed SIMD, ec o ope a ions a e pe o med in independen packe s wi h a maximum wid h equal o he machine’s maximum wo d wid h, a he han a whole bi ec o (i.e., i is no possible o pe o m a 128-bi wid h ope a ion in a 64-bi double wo d machine). Fo example, when pe o ming a le -shi ope a ion, he le mos bi o each wo d is los . Howe e , in o de o ec o ize BPM we would wan his bi o be appended o he closes -le wo d, e ec i ely pe o ming a ec o -wid h le -shi ope a ion. To ci cum en his p oblem, we mus pe o m addi ional ope a ions o manually ca y ha bi o he co ec posi ion. The numbe o addi ional ope a ions equi ed by his app oach o wo k scales wi h he ec o leng h. Thus, we decided o e alua e he po en ial o he ec o e sion o BPM using he Neon ec o ex ension (128-bi ec o s). The ec o ized loop execu es 1.7× mo e ins uc ions han he o iginal bu pe o ms 2× ewe i e a ions. On he A64FX, SIMD e sions o simple ins uc ions, like in ege addi ion, we e much mo e expensi e han scala ones. Fo example, a simple 64-bi addi ion akes one cycle, while a ec o addi ion o wo 64-bi wo ds akes ou cycles. This di e ence in la encies leads o a slow-down o 2×. G a i on3 has lowe SIMD la encies. Howe e , he inc ease in he numbe o ins uc ions in he loop leads o a 30% pe o mance loss. Since we did no gain any pe o mance using he Neon e sion, i was disca ded in a o o he o iginal scala code. We belie e ha an in e -sequence o coa se-g ain app oach (i.e., pe - o m he sequence alignmen o se e al sequences simul aneously) will deli e be e pe o mance since i simpli ies he ec o iza ion. BSW: The SVE e sion o BSW [70] is a ansla ion o A m SVE- in insics o he x86- ec o e sion ound in BWA-MEM2, which g oups he sequence alignmen o mul iple equal-leng h sequences ia SIMD in- s uc ions (i.e., in e -sequence ec o iza ion). The x86-in insics e sion o BSW elies on masks and blend ope a ions o selec alid en ies om he ec o egis e s. The SVE e sion akes ad an age o SVE’s p edica e ins uc ions o a oid he need o blend ope a ions, e ec i ely educing he numbe o o al ins uc ions execu ed. BSW uses in ege s o 16 bi s, allowing o p ocess 32 elemen s pe i e a ion using SVE-512 (A64FX) and 16 using SVE-256 (G a i on3). The SVE e sion o BSW pe o ms 3.4× and 1.3× as e han i s scala e sion on he A64FX and G a i on3, espec i ely. FAST-CHAIN: Ou SVE implemen a ion o FAST-CHAIN is a ansla- ion o SVE in insics o he x86 e sion. The o iginal x86 implemen a- ion o FAST-CHAIN execu es i s main loop scala e sion (i.e., a oids execu ing he ec o ized loop) when he numbe o i e a ions o pe - o m is small. Addi ionally, as usual in x86 ec o loops, i implemen s a loop- ail o p ocess he emaining elemen s. Since SVE is ec o -leng h agnos ic, we could a oid mos o he logic o he x86 e sion, educing he numbe o pe o med ins uc ions. The x86 ec o ized e sion o FAST-CHAIN uses 32 bi s ancho s. In some cases, 32-bi ancho s a e no su icien , and his ke nel gene a es inco ec esul s. To sol e his, we ha e implemen ed 64 and 32 bi s SVE e sions o FAST-CHAIN. The 64 bi s e sion always ou pu s co - ec esul s, bu we ha e used he 32 bi s implemen a ion o compa e agains he 32 bi s x86 implemen a ion. GenA chBench’s SVE e sion o FAST-CHAIN uns 4.5×and 1.8× as e han i s scala e sion (CHAIN wi hou heu is ics) on he A64FX and G a i on3, espec i ely. Expe imen al esul s show ha he pe - o mance o FAST-CHAIN compa ed o egula CHAIN g ea ly depends on he inpu used— he usage o heu is ics may lead o pe o mance a ia ions based on he cha ac e is ics o he inpu . Fo ins ance, us- ing GenA chBench’s la ge inpu , ou SVE e sion o FAST-CHAIN is 2.2× as e han egula CHAIN on he A64FX, bu i p esen s a 1.4× slowdown on G a i on3. WFA: The Wa e on Alignmen Algo i hm consis s o wo ope a- ions: compu e he nex wa e on (nex ope a ion) and ex end all he a hes - eaching poin s o a wa e on by exac ma ching cha ac e s om wo s ings (ex end ope a ion). The nex ope a ion can be au- oma ically ec o ized by he compile due o i s simple compu a ional pa e n. In con as , he ex end ope a o canno be au oma ically ec- o ized as each diagonal equi es an i egula amoun o compu a ions. To his end, we ha e ec o ized he ex end ope a ion using a cus om implemen a ion elying on SVE in insics. Each ec o lane ex ends a di e en diagonal, compa ing ou bases pe lane un il a misma ch is ound. Because each diagonal equi es a di e en numbe o cha ac e compa isons, some lanes can equi e mo e i e a ions han o he s. We ackle his p oblem by masking he lanes as hey inish he ex ension p ocess. This way, se e al diagonals a e ex ended in pa allel. The SVE e sion o WFA deli e s a 1.6× and 1.25× speedup o e i s scala e sion on he A64FX and G a i on3, espec i ely. 5.2. Op imized lib a ies Many HPC ke nels and ools ely on equen ly used lib a ies. I is common o endo s, such as A m, Fuji su, o In el, o de elop op imized e sions o widely used unc ions and lib a ies a ge ing hei sys ems and a chi ec u es. Fo genome da a analysis, some ools exploi neu al ne wo ks (NNs) o imp o e he quali y o hei analysis and e- sul s. Fo he GenA chBench, we ha e es ed di e en implemen a ions o he lib a ies used by he NN-BASE and NN-VARIANT ke nels. NN-BASE: The NN-BASE ke nel builds upon he PyTo ch lib a y [57]. On he A64FX we ha e used an op imized e sion o PyTo ch o his speci ic CPU p o ided by Fuji su. On G a i on3, we ied wo di e en PyTo ch backends: PyTo ch compiled wi h OpenBLAS ( ecommended by A m) and PyTo ch compiled wi h oneDNN op imized wi h ACL (labeled as expe imen al). The docke images wi h he wo backends a e a ailable in [71]. The oneDNN-ACL backend pe o med 6.5×be e han he OpenBLAS one, and he e o e, we used i o ou expe imen s. On SKX, we used an op imized e sion o PyTo ch ha exploi s he AVX512 ec o ex ension. On Rome, we used an op imized Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329 320 L. López-Villellas e al. Fig. 3. Single-co e ( op) and mul i-co e (bo om) execu ion ime o GenA chBench’s ke nels on he expe imen al se up. Mul i-co e esul s co espond o execu ions using all a ailable co es on each machine: 48 h eads on he A64FX and SKX and 64 h eads on G a i on3 and Rome. The esul s a e no malized o he pe o mance on he A64FX using one co e ( op) and 48 co es (bo om). FCHAIN, KCNT, NNB and NNV a e he abb e ia ions o FAST-CHAIN, KMER-CNT, NN-BASE and NN-VARIANT, espec i ely. NN-VARIANT is no aken in o conside a ion o he a e age in he mul i-co e plo . PyTo ch e sion ha suppo s he AVX2 ec o ex ension a ailable on he machine. NN-VARIANT: The o iginal NN-VARIANT ke nel om Genomics- Bench is based on Clai [25] a ian calle . In u n, his a ian calle e- lies on Tenso Flow [58]. Clai uses Tenso Flow 1 while Fuji su p o ides an op imized e sion o Tenso Flow 2 o he A64FX. Fo ha eason, we decided o use Clai 3 [26] ins ead, an upda ed e sion o Clai ha elies on Tenso Flow 2. To execu e using GenA chBench’s inpu s, we used he Ox o d Nanopo e 941_p om_hac_g360+g422 [60] p e- ained model om Clai 3. On G a i on3, we es ed h ee di e en Tenso Flow backends: Tenso Flow compiled wi h oneDNN op imized wi h ACL, using Tenso Flow’s Eigen h ead-pool o pa allelism ( ecom- mended by A m); Tenso Flow compiled wi h oneDNN op imized wi h ACL, using ACL’s schedule ; and Tenso low compiled wi h he Eigen backend. The docke images wi h he h ee backends a e a ailable in [72]. The Eigen backend pe o med mo e han 1.6×be e han he o he and he e o e i was he one used o un ou expe imen s. We used op imized Tenso Flow e sions on SKX and Rome capable o exploi ing he AVX512 and AVX2 ec o ex ensions. 5.3. Algo i hmic and code op imiza ions This sec ion p esen s he algo i hmic and code op imiza ion we pe o med o imp o e he pe o mance o FMI and KMER-CNT. FMI: GenA chBench’s FMI e sion implemen s h ee op imiza ions p oposed by Langa i a e al. [70]. One o he mos called unc ions in his ke nel is backwa dEx . To educe he o e head o he calls, his unc ion is always o ced o be in-lined. FMI uses he buil in_popcoun unc ion. This unc ion coun s he numbe o bi s se o one in an in ege . None o he es ed compile s ansla es his unc ion o SVE’s popula ion coun ins uc ion. Ins ead, hey use bi wise ope a ions and masks. To o ce exploi ing SVE capabili ies, all calls o buil in_popcoun a e eplaced by SVE in insics. FMI pe o mance is hea ily a ec ed by memo y access la encies. To hide hese la encies, he op imized e sion o FMI in e lea es he execu ion o se e al sequences, e ec i ely pe o ming se e al memo y accesses in pa allel. By applying he h ee p esen ed op imiza ions, we imp o ed he ke nel pe o mance on bo h A m machines by oughly 35%. KMER-CNT: Ou expe imen al e alua ion shows ha he pe o - mance o his ke nel is hea ily a ec ed by h ead mig a ions. To a oid h ead mig a ions, we po ed KMER-CNT om he P h eads lib a y o OpenMP and se OMP_PROC_BIND clause o ue be o e execu ions. This change led o mo e han 4×speedups on bo h x86 machines when using all a ailable co es. Howe e , he pe o mance o he A m machines emained he same. KMER-CNT elies on wo global da a s uc u es o s o e he numbe o indi idual k-me s: an a ay o 4-bi coun e s and libcuckoo’s [55] mul i- h ead hash-map, which s o es 64-bi coun e s. Each en y o he global a ay is an 8-bi a omic in ege , which is spli in hal o c ea e wo 4-bi coun e s. The a ay coun e s a e upda ed using a omic compa e-and-swap ope a ion. Once he 4-bi coun e o a k-me sa u a es, he ollowing inc emen s a e pe o med in he global hash map, also elying on a omic compa e-and-swaps o upda e i s coun e s. E en by a oiding h ead mig a ions, he scalabili y o he o iginal ke nel was poo on all he machines. I achie ed a maximum o 7×and 5× s. se ial execu ion on he A64FX (48 h eads) and G a i on3 (64 h eads), espec i ely. We decided o implemen wo new app oaches o y o imp o e pa allel pe o mance. The applica ion di ides he inpu be ween he a ailable h eads. Each h ead i e a es h ough he k-me s o i s pa o he inpu and inc emen s he coun e o he ead k-me s in he global a ay o hash map. Since he inpu is ead sequen ially and he e is almos no compu a ion o pe o m, mos o he execu ion ime is spen accessing he global coun e s in mu ual exclusion. To educe con en ion and imp o e da a locali y, ou i s app oach assigns pa o he inpu o each h ead, and all h eads ead he ull inpu bu only coun pa o he k-me s. This way, ins ead o ha ing a single global a ay and hash map, each h ead can ha e a smalle ins ance o he da a s uc u es and access hem wi hou con en ion. The o iginal e sion o he ke nel uses compa e-and-swap ins ead o e ch-and-add o upda e he coun e s because each 8-bi en y o he a ay s o es wo 4-bi coun e s. In o de o limi memo y usage, when he k-me size is g ea e han 17, he ke nel does no ins an ia e he global a ay, and all he coun ing akes place on he hash-map. Fo he maximum allowed k-me size (17), we equi e 217 4-bi en ies in he global a ay, esul ing in 8 GB o memo y. Since we ha e mo e han enough memo y in all sys ems, ou second app oach uses 8-bi ins ead o 4-bi coun e s, doubling he memo y equi emen s. This enables e ch-and-add usage and educes he numbe o accesses o he hash map. The single h ead execu ion ime o he ke nel did no change wi h any o he new e sions. Ou i s app oach (p i a e s uc u es) equally Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329 321 L. López-Villellas e al. di ides he possible k-me s be ween h eads. Howe e , some k-me s a e mo e common in he inpu , causing load imbalance be ween h eads de i ing in e en poo e scalabili y han he o iginal ke nel. The second app oach ( e ch-and-add) imp o es he ke nel’s scalabili y on all he machines: i uns 2.5×,1.4×, 3×and 2.3× as e han he o iginal e sion on he A64FX (48 h eads), G a i on3 (64 h eads), SKX (48 h eads) and Rome (64 h eads), espec i ely. Consequen ly, we used he e ch-and-add app oach o he es o he expe imen s. 6. Pe o mance cha ac e iza ion This sec ion p esen s a de ailed pe o mance cha ac e iza ion o he ke nels in ou expe imen al se up. We use he op imized e sions o he ke nels desc ibed in Sec ion 5. Fo all o he s udies p esen ed, we ha e anno a ed he code o he ke nels o de ine hei egion o in e es , i.e., we only s udy he pa o he ke nels dedica ed o meaning ul compu a ion. All he esul s shown in his sec ion ha e been compu ed using he la ge inpu o each ke nel. We no ed minimal a ia ion be ween execu ions o he ke nels, wi h a maximum ela i e s anda d de ia ion o 5% obse ed ac oss 10 epe i ions o he expe imen s. Consequen ly, we showcase he esul s based on a single execu ion in he igu es. While execu ing DBG wi h high h ead coun s in G a i on3, ou lie execu ion imes occu ed app oxima ely 10% o he ime. In he case o DBG in G a i on3, we selec i ely p esen esul s om an inlie execu ion. 6.1. Single- h ead pe o mance The op plo o Fig. 3 shows he single- h ead execu ion ime o each ke nel on he expe imen al se up. The esul s a e no malized o he pe o mance on he A64FX (see Table 3 o he supplemen a y ma e ial o he execu ion imes o he ke nels). The A64FX ea u es signi ican ly ewe ou -o -o de esou ces, a smalle memo y hie a chy, and highe memo y la encies han he es o he sys ems. On a e age, he o me is 2.4×, 1.8×, and 1.7× slowe han G a i on3, SKX and Rome on single- h eaded execu ions, espec i ely. Exploi ing he SVE capabili ies o he A64FX helps o educe his slowdown. SVE ec o ized ke nels (BSW, FAST-CHAIN, and WFA) p esen be e - han-a e age pe o mance on he A64FX: BSW pe o mance is simila o he exhibi ed on G a i on3 and only 17% wo se han he pe o mance on he x86 machines, FAST-CHAIN pe o ms be e han on Rome, and WFA pe o ms be e han on SKX. No e ha BSW and FAST-CHAIN exploi AVX512 on SKX while hey le e age AVX2 on Rome, and ha WFA is no ec o ized on he x86 machines. The deep-lea ning ke nels (NN-BASE and NN-VARIANT) a e he wo s -pe o ming on he A64FX. G a i on3 pe o ms excep ionally well in single- h ead execu ions. On a e age, i p esen s 2.44×, 1.33×, and 1.39×pe o mance speedups wi h espec o he A64FX, SKX and Rome, espec i ely. FAST-CHAIN pe o mance on G a i on3 is 1.8×be e han on Rome (AVX2) bu 70% wo se han on SKX since i exploi s AVX512 (512 bi s) on ha machine. WFA uns 2.5×and 1.8× as e on G a i on3 han on SKX and Rome, espec i ely. In con as o he A64FX, he deep-lea ning ke nels (NN-BASE and NN-VARIANT) deli e good pe o mance on G a i on3, showing speedups o be ween 3.1–6.2×compa ed o he A64FX. 6.2. Pa allel pe o mance We e alua e he pa allel pe o mance o GenA chBench’s ke nels using di e en h ead coun s: 2, 8, 24, 48, and 64. The A64FX and SKX implemen 48 co es. Hence, execu ions wi h mo e han 48 h eads ha e only been pe o med on G a i on3 and Rome. Con olling h ead a ini y was manda o y in ou expe imen s o achie e good pa allel pe o mance on he machines, especially on he A64FX. Fo mos ke nels, all he execu ions we e pe o med by binding h eads o co es. Fig. 4. Speedup o e se ial execu ion o GenA chBench’s ke nels on he expe imen al se up. We show he achie ed speedup using di e en h ead coun s: 2, 8, 24, 48, and 64. The A64FX and SKX 64- h eads poin s a e no shown in he igu e, since hose machines only implemen 48 co es. ABEA, NN-BASE, and NN-VARIANT do no allow ull h ead a ini y con ol. The e o e, h ead mig a ions can occu in hese h ee ke nels. Fig. 4 shows he speedup o e se ial execu ion achie ed by he ke nels on he expe imen al se up using he p e iously p esen ed h ead coun s. Addi ionally, he bo om plo o Fig. 3 compa es he pe o - mance ob ained using all a ailable co es on each machine: 48 h eads Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329 328 L. López-Villellas e al. [89] H. Li, e al., A su ey o sequence alignmen algo i hms o nex -gene a ion sequencing, B ie . Bioin o m. 11 (5) (2010) 473–483, h p://dx.doi.o g/10.1093/ bib/bbq015. [90] A. Zielezinski, e al., Alignmen - ee sequence compa ison: bene i s, applica ions, and ools, Genome Biol. 18 (1) (2017) h p://dx.doi.o g/10.1186/s13059-017- 1319-7. [91] Y. Tu akhia, e al., Da win, ACM SIGPLAN No . 53 (2) (2018) 199–213, h p: //dx.doi.o g/10.1145/3296957.3173193. [92] A. Nag, e al., Gencache: Le e aging in-cache ope a o s o e icien sequence alignmen , in: P oceedings o he 52nd Annual IEEE/ACM In e na ional Sym- posium on Mic oa chi ec u e, 2019, pp. 334–346, h p://dx.doi.o g/10.1145/ 3352460.3358308. [93] D. Fujiki, e al., GenAx: A genome sequencing accele a o , in: 2018 ACM/IEEE 45 h Annual In e na ional Symposium on Compu e A chi ec u e, ISCA, 2018, pp. 69–82, h p://dx.doi.o g/10.1109/ISCA.2018.00017. [94] H. Sadasi an, e al., Accele a ed dynamic ime wa ping on GPU o selec i e nanopo e sequencing, J. Bio echnol. Biomed. 07 (01) (2024) h p://dx.doi.o g/ 10.26502/jbb.2642-91280134. [95] T. Dunn, e al., SquiggleFil e : An accele a o o po able i us de ec ion, in: MICRO-54: 54 h Annual IEEE/ACM In e na ional Symposium on Mi- c oa chi ec u e, MICRO ’21, ACM, 2021, h p://dx.doi.o g/10.1145/3466752. 3480117. [96] P.J. Shih, e al., E icien eal- ime selec i e genome sequencing on esou ce-cons ained de ices, GigaScience 12 (2022) h p://dx.doi.o g/10.1093/ gigascience/giad046. [97] T. Robinson, e al., Ha dwa e accele a ion o genomics da a analysis: chal- lenges and oppo uni ies, Bioin o ma ics (2021) 1–11, h p://dx.doi.o g/10. 1093/bioin o ma ics/b ab017. Lo ién López-Villellas is a Ph.D. s uden a he Uni e si y o Za agoza. His esea ch ocuses on exploi ing no el and consolida ed pa allel and ec o a chi ec u es o scien i ic applica ions, such as molecula dynamics and genomics. P io o s a ing his Ph.D., he wo ked as a esea ch enginee a he Ba celona Supe compu ing Cen e o wo yea s. He holds a BSc in compu e science om he Uni e si y o Za agoza and a MSc in High-Pe o mance Compu ing om he Uni e si a Poli ècnica de Ca alunya. Rubén Langa i a-Bení ez ecei ed his B.S. deg ee in com- pu e science om Uni e sidad de Za agoza in 2018. He spen one academic yea as an E asmus s uden a he Uni e si y College Co k. His inal deg ee p ojec was abou op imizing molecula dynamics applica ions. He ecei ed MS deg ee om UPC in Janua y 2021. He is cu en ly wo k- ing a he BSC as a esea ch s uden . His esea ch in e es s include p ocesso mic oa chi ec u e and HPC applica ions. Asa Badouh is a esea ch so wa e enginee wi h a pas- sion o high-pe o mance compu ing and machine lea ning. O e he las 4.5 yea s, Asa has been wo king a he Ba celona Supe compu ing Cen e as a Resea ch Enginee , ocusing on genomics and heal hca e applica ions. P io o ha , Asa wo ked o 6 yea s as an in e n and enginee in he In el compile R&D eam in Is ael, de eloping LLVM- based compile s o OpenCL and C/C++. Aside om ha , Asa holds a bachelo ’s deg ee in compu e science om he Technion, Is ael; and a mas e ’s deg ee in Da a Science om Uni e si a Poli ècnica de Ca alunya. Víc o So ia-Pa dos ecei ed a B.Sc. in compu e science om Uni e si a de Za agoza, in 2019. He ecei ed an M.Sc. in compu e science om Uni e si a Poli ècnica de Ca alunya (UPC), in 2022. He is cu en ly a second-yea Ph.D. in compu e a chi ec u e wi h he UPC. He also wo ks as a esea che in he Ba celona Supe compu ing Cen e , wi hin he Cen e o Excellence pa ne ship wi h A m. His esea ch in e es s include high-pe o mance compu ing a chi ec u es, cache cohe ence, and mul ico e a chi ec u es. Quim Aguado-Puig ecei ed a B.Sc. deg ee in compu e sci- ence in 2019 om he Uni e si a Au ònoma de Ba celona (UAB). He ecei ed an MSc a he Uni e si a Poli ècnica de Ca alunya (UPC) in 2023. He is cu en ly a i s -yea Ph.D. s uden a UAB. He has p e iously wo ked as a esea ch enginee in he p ojec Designing RISC-V-based Accele a o s o nex -gene a ion Compu e s (DRAC) a UAB in collab- o a ion wi h he Ba celona Supe compu ing Cen e (BSC). His esea ch in e es s include high-pe o mance compu ing, massi ely pa allel a chi ec u es, and GPU p og amming; wi h applica ions o genomics, compu a ional biology, and sequence alignmen . Guillem López-Pa adís ecei ed a B.Sc. and M.Sc. om Uni e si a Poli ecnica de Ca alunya (UPC) in 2017 and 2020, espec i ely. He is cu en ly a hi d yea Ph.D. s uden in Compu e A chi ec u e a UPC and Ba celona Supe compu ing Cen e (BSC). He ac i ely pa icipa es in di e en Eu opean p ojec s, as well as in in e na ional collabo a ions wi h academia and indus y. He has al- eady published some pape s in in e na ional con e ences and pa icipa ed in di e en apeou s designing powe - e icien RISC-V p ocesso s. His esea ch in e es s include high-pe o mance compu ing a chi ec u es, scaling RTL Sim- ula ions, and domain-speci ic accele a o s, wi h special emphasis on cohe en in e connec s be ween co es and ha dwa e accele a o s. Max Doblas ecei ed a B.Sc. in elec ical enginee ing and compu e science om Uni e si a Poli ècnica de Ca alunya (UPC), in 2020. He ecei ed an M.Sc. in compu e sci- ence om UPC, in 2021. He is cu en ly a second-yea Ph.D. in compu e a chi ec u e wi h he UPC. He also wo ks as a Resea ch Enginee in he p ojec Designing RISC-V-based Accele a o s o nex -gene a ion Compu e s (DRAC) a he Ba celona Supe compu ing Cen e (BSC), in which he has designed a powe -e icien p ocesso wi h se e al ex ensions o domain-speci ic applica ions. His esea ch in e es s include high-pe o mance compu ing a chi ec u es, and domain-speci ic accele a o s, wi h appli- ca ions o genomics, compu a ional biology, and sequence alignmen . Ja ie Se oain ecei ed his Ph.D. in compu e science and enginee ing om he Complu ense Uni e si y o Mad id (UCM), wo king on wo kload op imiza ion o GPUs. A e wo king as a pos -doc a he Spanish Na ional Cen e o Bio echnology (CNB), and a esea ch enginee a A m Resea ch, he is cu en ly a senio membe o echnical s a in AMD Resea ch and Ad anced De elopmen , wo king on compile esea ch o AI accele a o s. His esea ch in e es s cen e a ound compu e accele a o s and specialized a chi- ec u es, and his cu en esea ch is ocused on au oma ed ML wo kload op imiza ions o spa ial a chi ec u es. Chulho Kim is P incipal Consul an wi h he Leno o In as uc u e Solu ions G oup Se ices in Uni ed S a es. Chulho has a Bachelo o Science deg ee in Ma hema - ics o Compu a ion om UCLA in 1989 and joined IBM Kings on. He joined Leno o US in 2014. He has wo ked in High Pe o mance Compu ing (HPC) since 1993. He likes o apply his skills o debug ex emely complex issues, anging om applica ion pe o mance o sys em and ne - wo king pe o mance issues. He is esponsible o unning Top500/G een500 on cus ome clus e s o Leno o (#1 G een500 en y since No embe 2022). Mako o Ono ecei ed he ME in biophysics om he Osaka Uni e si y in 1985 and joined IBM Tokyo Resea ch lab. He wo ked on compu e g aphics esea ch hen s a ed sys em a chi ec u e de elopmen in IBM Sys em x de elopmen . He is cu en ly a Dis inguished Enginee a Leno o In as- uc u e Solu ions G oup and is a lead a chi ec o edge compu ing. His in e es s include edge compu ing, edge AI, and he e ogeneous and al e na i e a chi ec u e including non adi ional CPU / GPU a chi ec u e. Fu u e Gene a ion Compu e Sys ems 157 (2024) 313–329 329 L. López-Villellas e al. Ad ià A mejach is a Lec u e P o esso in Compu e A - chi ec u e a Uni e si a Poli ècnica de Ca alunya (UPC), and associa e esea che a he Ba celona Supe compu ing Cen e (BSC). He ecei ed his Ph.D. om UPC in 2014 and hen s a ed his esea ch ca ee a BSC, whe e he lead he echnical con ibu ions o mul iple FP7 and H2020 p ojec s. His esea ch in e es include memo y sys ems, he e oge- neous a chi ec u es, simula ion me hodologies and ec o a chi ec u es. Cu en ly he leads a g oup ha o e sees all he a chi ec u al simula ion e o s a BSC, enabling esea ch on mul iple compu e a chi ec u e opics. He has published mo e han 30 well- anked in e na ional con e ence and jou nal pape s. San iago Ma co-Sola ecei ed he M.Sc. and Ph.D. deg ees in compu e science om he Uni e si a Poli ècnica de Ca alunya (UPC) in 2012 and 2017, espec i ely. Du - ing his Ph.D., he wo ked a he Algo i hm De elopmen and Bioin o ma ics G oup a Spanish Na ional Cen e o Genome Analysis (CNAG) and lec u ed a he Uni e si a Au ònoma de Ba celona (UAB). He is cu en ly a Senio Resea che a he Ba celona Supe compu ing Cen e (BSC) and a Lec u e a he UPC. His esea ch in e es s include high-pe o mance compu ing, he e ogeneous a chi ec u es, genome-da a analysis, and algo i hms in bioin o ma ics and compu a ional biology. Jesús Alas uey-Benedé ecei ed he M.S. deg ee in Telecommunica ion and he Ph.D. deg ee in Compu e Sci- ence om Uni e sidad de Za agoza in 1997 and 2009, espec i ely. He is an associa e p o esso in he Compu e Science and Sys ems Enginee ing Depa men (DIIS), Uni- e sidad de Za agoza, Spain. His esea ch in e es s include p ocesso mic oa chi ec u e, memo y hie a chy, and HPC applica ions. Pablo Ibáñez ecei ed he M.S. deg ee in compu e science om he Uni e si a Poli écnica de Ca alunya, Spain, in 1989, and he Ph.D. deg ee in compu e science om he Uni e sidad de Za agoza, Spain, in 1998. He is an associa e p o esso wi h he Compu e Science and Sys ems Engi- nee ing Depa men , Uni e si y o Za agoza. His esea ch in e es s include p ocesso mic oa chi ec u e, memo y hie - a chy, pa allel compu e a chi ec u e, and high pe o mance compu ing applica ions. Miquel Mo e ó ecei ed he B.Sc., M.Sc., and Ph.D. de- g ees om Uni e si a Poli ècnica de Ca alunya (UPC), Spain. Cu en ly, he is a Ramón y Cajal Fellow a UPC Ba celona. P io o joining UPC, he spen 5 yea s as a Senio Resea che wi h he Ba celona Supe compu ing Cen- e (BSC), Spain, and 15 mon hs as a pos -doc o al ellow wi h he In e na ional Compu e Science Ins i u e (ICSI), Be keley. His esea ch in e es s include high pe o mance compu e a chi ec u es, domain-speci ic accele a o s and ha dwa e-so wa e co-design o u u e massi ely pa allel sys ems.