scieee AI-readable full text Open interactive document viewer

2Pipe: It starts with a question. Matching you with the correct pipeline for MAG reconstruction (Pipeline description)

Yepes García, Jeferyd; Falquet, Laurent

Abstract

Descriptive overview of the main workflow for each pipeline or platform, where important technical considerations such as the type of input (short reads, long reads or both), key tools employed at each step, advantages, limitations and/or special features they depict are documented.

Full text

2Pipe: It starts with a question. Matching you with the correct pipeline for MAG reconstruction Jeferyd Yepes-Garcíaa,b, Laurent Falqueta,b,* aDepartment of Biology, University of Fribourg, Fribourg, Canton of Fribourg, 1700, Switzerland bSwiss Institute of Bioinformatics, Lausanne, Vaud, 1015, Switzerland *Correspondence should be addressed to L.F ([email protected]) __________________________________________________________________________________ Descriptive pipeline overview Below we present a descriptive overview of the main workflow for each pipeline or platform, where important technical considerations such as the type of input (short reads, long reads or both), key tools employed at each step, advantages, limitations and/or special features they depict are documented. 1. Short-read centered pipelines 1.1 Anvi’o1 Anvi’o is a comprehensive modular platform for the analysis and visualization of microbial omics including, but not restricted to, metagenomics, metatranscriptomics and metapangenomics. Anvi’o is developed to be highly customizable through exchangeable programs (tools) that perform specific tasks, empowering the user with a wide range of tools to explore. Being so, a metagenomics workflow is proposed by the developers of the platforms that begins with short-read quality cleaning, proceeds to read assembly to be used for read recruitment (mapping), and finalizes contig annotation (functions, Hidden Markov Models, and taxonomy). Optionally, the user can achieve read taxonomic profiling with KrakenUniq2, and more recently binning tools have been made available such as MetaBAT23, CONCOCT4, MaxBin25 and BinSanity6, as well as DASTool7 as a refinement alternative. Nonetheless, the user must run the analysis manually, requiring them to account with some experience regarding software installation, execution and debugging. Moreover, although Anvi’o is in principle a command line tool, it incorporates a user-friendly graphical interface for data inspection and visualization that is commonly used for contig visualization. 1.2 BugBuster8 BugBuster is an automatic, modular, and reproducible Nextflow9 (DSL2) workflow with specialized modules for taxonomic profiling and resistome characterization. Its workflow encompasses the following steps: initial reads processing for quality filtering and host contamination removal (Bowtie210); taxonomic profiling at the read level using tools like Kraken211/Bracken12 or Sourmash13; and antibiotic resistance gene (ARG) prediction from reads using KARGA14 and KARGVA15. The assembly is carried out with MEGAHIT16, followed by taxonomic and functional annotation of contigs using BLAST17, BlobTools18, DeepARG19, and MetaCerberus20. Afterwards, the contigs are binned with tools such as MetaBAT23, SemiBin221 and COMEBin22, and refined them with a MetaWRAP23-native module; the quality is assessed with CheckM224, and the MAGs are taxonomically affiliated with GTDB-Tk225. BugBuster is fully containerized (Docker) aiming at ensuring ease of installation, high reproducibility, and deployment across various computational environments. Moreover, BugBuster stands out given its inclusion of specific tools to characterize and quantify genes associated with antibiotic resistance. 1.3 DATMA26 DATMA (Distributed AuTomatic Metagenomic Assembly and annotation framework) is a pipeline focused on speed and automation, leveraging distributed computing for efficiency. As a starting point, DATMA applies a quality filter with RAPPIFILT (customized tool developed for this pipeline), Trimmomatic27 and FastQC28, and if the input sequences are paired-end, it merges them using FLASH229 and ForceMerge. Following this procedure, this pipeline identifies and removes 16S rDNA sequences based on RFAM30 (RNA sequence families), NCBI31, Ribosomal Database Project (RDP)32 and SILVA33 to cluster the remaining sequences with CLAME34. The clusters (or bins in definition of the traditional workflow) generated then are assembled in batches by metaSPAdes35, Velvet36, and MEGAHIT16 for a subsequent taxonomic annotation relying on BLAST37 and Kaiju38, as well as ORF prediction with Prodigal39 and GeneMark40. To conclude with the analysis a detailed HTML report is generated with interactive Krona41 plots for taxonomic visualization; this report integrates the 16S rDNA annotation (RDP Classified) along with the annotated bins. As inferred from the described workflow, DATMA performs an inverted approach to generate bins by first grouping the reads using CLAME and attempting to assemble only these groups individually afterwards. Further, this pipeline is wrapped by COMP Superscalar which facilitates the development and execution of parallel applications for distributed infrastructures such as clusters, cloud services and containerized platforms. 1.4 EasyMetagenome42 EasyMetagenome integrates a classical workflow starting with short reads to provide a dereplicated (dRep43) set of bins and pangenome analysis that relies on an Anvi’o module. The assembly is performed with MEGAHIT16, a MetaWRAP23 module is in charge of the binning task, CheckM224 controls the quality of the bins, and GTDB-Tk225 finalizes the execution by taxonomically annotating them. Notably, this pipeline performs functional annotation (GhostKOALA44, eggNOG45, dbCAN346) and taxonomy assignment on the contigs after a pre-filtering step that generates a non-redundant gene set. EasyMetagenome uses Conda environments to assure reproducibility, the user can input multi-sample data, although it is not orchestrated by any workflow manager. As special remarks, it carries out a taxonomic profiling (MetaPhlAn47, HUMAnN348, Kraken211) of the post-filtered (KneadData) reads, and the functional annotation of the gene set is expanded to identify virulence factors (VFDB49) and antibiotic resistant genes (CARD50). 1.5 EURYALE (MEDUSA)51,52 EURYALE is a Nextflow-based reimplementation of the MEDUSA pipeline. It provides a modular and containerized workflow using Nextflow DSL2, with software execution through Docker, Conda or Singularity, which ensures portability, reproducibility, and scalability. The workflow of this pipeline starts with read quality control with FastQC28, trimming and merging using fastp53, and optional host decontamination with Bowtie210; MultiQC54 provides a full report containing visualizations regarding sequence preprocessing. Optionally, clean sequences can be assembled using MEGAHIT16 with a posterior taxonomic classification carried out by Kaiju38 or Kraken211, while functional annotation relies on a DIAMOND55-based alignment to reference databases (NCBi nr by default). It is worthy to mention the flexibility EURYALE offers given its customizable database selection for both taxonomic and functional annotation. 1.6 JAMS56 JAMS (Just a Microbiology System) is an integrated framework originally designed to perform the analysis on the NIH’s Biowulf system. JAMs is divided into two main modules: JAMSα, which performs single sample analyses, and JAMSβ, which focuses on cross-sample comparisons. JAMSα (the pipeline) integrates tools such as Bowtie210 for host removal, MEGAHIT16 or SPAdes57 for read assembly, Kraken211 for taxonomic classification, and Prokka58 and InterProScan59 for gene and protein domain prediction, respectively; JAMSβ uses R-based packages for visualization and statistical analysis. This workflow is executed within Conda environments, and its main advantage relies on the ease to establish comparisons across samples. However, this pipeline does not support binning tools nor genome-quality, and currently, it exhibits restricted deployment flexibility due to optimization for the NIH’s Biowulf system, although JAMS is open source and can be installed on any UNIX-based machine. 1.7 MAGNETO60 MAGNETO is an automated, modularized and scalable pipeline wrapped with Snakemake61 and executed with Conda. It is focused on allowing the user the selection of different assembly and/or binning strategies, involving several steps from read pre-processing until MAG annotation and gene catalog generation. The Pre-processing module leverages fastp53, Bowtie210 and FastQ Screen62, whilst the Assembly mode uses Simka63 and hierarchical agglomerative clustering to cluster the samples if the users pre-defines a co-assembly strategy; the reads are assembled using MEGAHIT16. Furthermore, contig abundances are computed by alignment against the raw reads to be bin by MetaBAT23 afterwards. Quality estimation and dereplication are carried out with CheckM64 v1.0 and dRep43, respectively. To end the workflow, a gene catalog is produced for both the contigs and the MAGs by running Prodigal39, Linclust65 and CD-HIT66, and the MAGs are annotated with GTDB-Tk225 and eggNOG-mapper67. As a special feature, MAGNETO can provide a read-based taxonomy abundance with mOTU68 profiler. MAGNETO exhibits all the advantages Snakemake wrapping, and executed with Conda, represents such as multi-sample handling, scalability across different computing infrastructures and checkpoint control for workflow restarting. 1.8 MAGO69 MAGO is an end-to-end pipeline designed to run over a single execution from a container image (Singularity or Docker); a third option is available as a Virtual Machine (VM). This configuration allows MAGO to offer a streamlined implementation of the entire metagenomics pipeline, including error checking, and computational resource distribution. The tool workflow follows the traditional design with read quality control (fastp53, FastQC28), followed by the assembly step with MEGAHIT16, metaSPAdes35 and/or IBDA-UD70. MAGO performs binning through multiple algorithms (MetaBAT71, MaxBin25, CONCOCT4 and BinSanity with multiple configurations). MAG completeness and contamination of MAGs are estimated with CheckM64. To conclude the execution, MAGO annotates the MAGs with Prokka58, and performs taxonomic classification and phylogenetic placement using GTDB-Tk72. Moreover, to expand its capabilities, the developers included the possibility of generating phylogenetic trees through ezTree73, analyzing the pangenome with Roary and measuring ANI with FastANI74 as an approximation to de-replicate the MAG set. 1.9 metaGEM75 metaGEM represents a traditional end-to-end pipeline designed to reconstruct MAGs from metagenomics raw reads; however, its main feature relies on an integrated module that provides genome scale metabolic models (GEMS). The workflow starts with the read quality cleaning using fastp53 for a subsequent assembly with MEGAHIT16 and a contig coverage estimation with BWA76. The bins are then obtained via three different tools (MetaBAT23, MaxBin25 and CONCOCT4) along a posterior refining by the metaWRAP23 refinement module. As a result, the bins or MAGs are used as input for CarveMe77 (Genome Scale Metabolic Models), and SMETANA78 is called for metabolic interaction predictions and MEMOTE79 is in charge of generating quality reports. The resulting GEMs can then be used for various downstream analyses, such as predicting metabolic interactions within the community, simulating growth under different conditions, and identifying key metabolic pathways. The pipeline ends with MAG characterization through Prokka58 and Roary80 (functional annotation and pangenome analysis), GRiD81 (growth rate estimation), GTDB-Tk225 (taxonomic annotation) and BWA76 (genome abundance). As additional features, metaGEM identifies eukaryotic MAGs via EukRep82 and evaluates contamination with EukCC83. Also, this pipeline produces taxonomic abundance profiles from the filtered reads using mOTUS284. Naturally, this pipeline exhibits the benefits Snakemake61 orchestration provides, as mentioned previously. 1.10 MetaGenePipe85 MetaGenePipe is a pipeline developed with Workflow Definition Language (WDL), selfexecuted within a Singularity container, whose primary goal is performing a contig-based functional and taxonomic analysis from short read sequences. It is composed of 4 subworkflows, where the operation starts with the quality control workflow, the subsequent one assembles the reads with MEGAHIT16 to map them back against the short reads within the third subworkflow. Meanwhile, the last subworkflow is in charge of gene prediction and functional annotation based on two main strategies: alignment with the Swiss-Prot database and Hidden Markov Models search in KOfam database86. Although MetaGenePipe does not include binning software to provide MAGs as main output, its versatility that allows an analysis adapted for eukaryotic and viral analyses with minimal modifications, and its uncommon workflow manager within the pipelines considered in this review, makes MetaGenePipe an interesting alternative for users with advanced computational infrastructures. Additionally, MetaGenePipe is designed to handle a co-assembly strategy in case the user requires this feature. 1.12 Metagenome-Atlas87 Metagenome-Atlas is an end-to-end, Snakemake61-based and Conda-executed pipeline supporting Illumina short reads and providing a modular workflow. It is divided into four modules, namely Quality Control, Assembly, Genomic Binning and Annotation. The initial module removes host, common contaminants and PCR duplicates, and if necessary, trims low-quality sequences according to user pre-specified parameters. The Assembly module corrects sequence errors based on k-mer coverage, merges paired-end sequences, assembles them using MEGAHIT16 and/or metaSPAdes35 along with a contig-length filtering. The following module uses MetaBAT23, MaxBin25, and optionally VAMB88 and SemiBin221 to bin the contigs; CheckM224, BUSCO89 and GUNC90 are run to measure the bin quality, as well as DASTool7 and dRep43 for bin refinement and MAG dereplication, respectively. For the last module, Metagenome-Atlas taxonomically and functionally annotates the MAGs using GTDB-Tk225 and DRAM91, respectively, and it finally produces a gene catalog through mapping the predicted coding sequences using eggNOG-mapper67. Among the main advantages of MetagenomeAtlas, it is possible to describe the possibility of running individual modules and its energetic supporting community and developers. Moreover, the Snakemake wrapper allows for flexibility, multisample handling, and adaptability to medium to large projects running on local servers or HighPerformance Cluster (HPC) environments. 1.13 Metaphor92 Metaphor is a classic metagenomics pipeline aiming at MAG reconstruction and annotation wrapped by Snakemake61 and leveraging Conda as package manager. The pipeline is triggered by the user with a .csv file pointing to the sequence directories and a .yaml file with the pipeline configuration. A quality control will be carried out then with FastQC28 and fastp53, with a posterior assembly with MEGAHIT16, contig evaluation with MetaQUAST93 and mapping against the input sequences using Minimap294 and Samtools; the contigs are binned (VAMB88, MetaBAT23, CONCOCT4) and refined (DASTool7). Metaphor execution finalizes with bin annotation through Prodigal, Diamond, and the NCBI COG database. Complementary to Snakemake orchestration capabilities, Metaphor provides a series of plots depicting runtime and memory with the goal of identifying computational bottlenecks during the analyses. 1.14 MetaWRAP23 MetaWRAP is a popular and customizable pipeline built primarily as a command-line framework with a focus on flexibility and user control. MetaWRAP consists of individual modules that can be run independently or combined into custom workflows. Its core functionalities encompasses read QC and cleaning (FastQC28, Trim Galore and BMTagger), assembly (MEGAHIT16, metaSPAdes35, BWA76 and MetaQUAST93), and a binning suite that incorporates MetaBAT23, MaxBin25, and CONCOCT4. MetaWRAP also includes a native refinement module that produces hybrid bin sets to explore over the different variants of each bin (original and hybridized bin sets) to determine the “best bin” according to the user pre-specified quality values based on completeness and contamination (CheckM64 v1.0). This module is frequently executed in independent metagenomics analysis, and even some pipelines described in this review incorporate it within their workflows. If decided by the user, MetaWRAP offers the possibility of bin re-assembling guided by their previous versions, improving the overall bin quality. For MAG taxonomic and functional analysis, MetaWRAP relies on Prokka58 and Taxator-tk95 (combined with NCBI31 databases), and it provides visualization modules for summarizing results. Analogous to MAGNETO60, MetaWRAP can produce read-based taxonomic profiles in parallel. Although MetaWRAP does not integrate full pipeline automation, its high modularity and straightforward design have promoted a wide supporting community. Nonetheless, at the moment of writing this report, MetaWRAP is not maintained by the developers, with the subsequent lack of tool updates. Nonetheless, given the popularity of MetaWRAP, a Snakemake61 wrapper was developed to automate the metagenomics analysis known as SnakeWRAP96. Therefore, SnakeWRAP can carry out the MetaWRAP end-to-end read processing to generate MAGs in a single run, retaining the flexibility of MetaWRAP while reducing the burden of manual execution and dependency handling. Additionally, SnakeWRAP’s integrated environment management via Conda and support for HPC environments enables seamless execution of multiple MetaWRAP modules and samples in parallel, being particularly useful for multi-sample execution. 1.15 MOSHPIT97 According to its documentation, MOSHPIT (MOdular SHotgun metagenome Pipelines with Integrated provenance Tracking) is a toolkit of plugins for whole metagenome assembly, annotation, and analysis built on the microbiome multi-omics data science framework QIIME 298. MOSHPIT enables flexible, modular, fully reproducible workflows for read-based or assembly-based analysis of metagenome data. The core components of MOSHPIT include q2-assembly, which provides functionalities for genome assembly and quality control, and q2-annotate, which supports contig binning, taxonomic classification, and functional annotation. Additional plugins, such as q2-viromics and q2-amrfinderplus, extend capabilities to viral sequence detection and antimicrobial resistance gene annotation, respectively. In technical terms, MOSHPIT must be run locally or on an HPC environment with the possibility to execute the processes in parallel by the explicit declaration of partitions, a native QIIME2 functionality. Further, the entire QIIME2 ecosystem relies on Conda, and hence this a sine-qua-non requisite to perform MAG reconstruction with MOSHPIT. 1.16 nIMP399 nIMP3 is a Nextflow-based reimplementation of the IMP (Integrated Meta-omic Pipeline) workflow that assembles metagenomics (MG) and metatranscriptomics (MT) datasets together. nIMP3 handles preprocessed and contaminant-free MT and MG reads (FastQC28, SortMeRNA100, BBTools101), and jointly assembles them in a hybrid and iterative process using MEGAHIT16. Additionally, nIMP3 performs taxonomic profiling with mOTUs102 and Kraken211, as well as functional profiling with gffquant103. Unlike the original IMP pipeline, nMP3 does not include a binning module, and thus it cannot recover MAGs. Nonetheless, nIMP3 offers a lighter, reproducible, and integrative pipeline for multi-omics metagenome/metatranscriptome processing. 1.17 SnakeMAGs104 SnakeMAGs is a simple yet useful pipeline that as its name indicates is controlled by a Snakemake61 wrapper with Conda as software administrator. It integrates basic modules starting with quality control with Illumina-utils105 and Trimmomatic27, and if required, host removal with Bowtie210. Afterwards, the reads are assembled through MEGAHIT16, the contigs are binned by MetaBAT23, a quality assessment is carried out with CheckM64 v1.1 and GUNC90, MAG abundances are obtained using CoverM106, and finally the taxonomic classification is performed using GTDB-Tk225. Similar to the previous pipelines governed by Snakemake, SnakeMAGs eases automation, reproducibility, scalability and workflow management. 1.18 SPIRE107 The SPIRE project employs a Nextflow-based pipeline that has been used to process and annotate more than 100,000 metagenomes belonging to more than 700 studies. The workflow incorporates tools such as NGLess108 for read trimming and decontamination, MEGAHIT16 for assembly, Prodigal39 for gene prediction and barrnap109 for ribosomal RNA detection. Moreover, contig binning is carried out with MetaBAT23 with a complementary genome quality assessment using CheckM224 and GUNC90, and the workflow ends with taxonomic classification (GTDB-Tk225) and functional annotation (eggNOG-mapper67, abricate110, RGI50 and Macrel111). Among the advantages SPIRE offers, the possibility to perform antimicrobial resistance gene prediction and the annotation of virulence factors stand out, as well as its scalability, reproducibility across high-performance and cloud environments, and standardized processing, enabling consistent comparisons across global datasets. Nonetheless, at the moment of writing this report, this pipeline is aiming to be executed at online platforms like CloWM112 as it is lacking defined environments or container images, and the input data should be already hosted at the sequencing archives such as ENA, DDBJ or SRA. 1.19 Sunbeam113 Sunbeam is a modular pipeline orchestrated by Snakemake61 with Conda as dependency manager; this configuration makes Sunbeam analysis reliable, reproducible and scalable. The main feature Sunbeam depicts is its modularized and extensible design that allows users to build off the core functionality. The execution backbone of Sunbeam is represented by an initial quality control that encloses adapter trimming, host read removal and low-complexity filtering (Trimmomatic27, FastQC28, BWA76 and Komplexity), followed the assembly of reads into contigs with MEGAHIT16 along with their corresponding annotation with Prodigal39, BLAST37 and Diamond55 (with nucleotide or protein databases). As complementary procedures, Sunbeam maps the reads to reference genomes (user pre-specified) and delivers a taxonomic assignment of the clean reads using Kraken114 v1.0. As previously stated, its modularization and ready-to-use templates to create new modules have enabled the development of additional extensions for assigning metagenomic reads to a full bacterial phylogeny, single genome assembly, among others. 2. Long-read focused pipelines 2.1 EasyNanoMeta115 EasyNanoMeta is a specialized pipeline designed to process ONT long reads either solely or in combination with short reads (hybrid assembly). This pipeline relies on a dual approach that uses both assembly-based and assembly-free strategies. Particularly, EasyNanoMeta incorporates four assemblers (metaFlye116, OPERA-MS117, metaSPAdes35, MetaPlatanus118), five binners (SemiBin221, MetaBAT23, MaxBin25, CONCOCT4, VAMB88) and a polishing tool (NextPolish119) to assure the best possible outcome. Additionally, once the bins are obtained, it performs the common tasks such as functional annotation with Prokka58, quality control with CheckM224, phylogeny inference with PhyloPhlan120 and taxonomic classification with GTDB-Tk225. For the assembly-free methodology, EasyNanoMeta provides a full report containing composition, diversity and correlation among the identified species with Kraken211 and Centrifuge121. Regarding operational characteristics, this pipeline can be run automatically on a Singularity/Apptainer image that streamlines the setup process and minimizes dependency issues or experienced users can execute individual modules through shell scripts that rely on Conda environments. 2.2 Hi-Fi-MAG-Pipeline122 Hi-Fi-MAG is a simple, yet time-saving pipeline developed and maintained by Pacific Biosciences specially designed to build MAGs from Hi-Fi reads (long PacBio reads). It encompasses different binning tools (MetaBAT23 and SemiBin221) along with DASTool7 as refinement software; CheckM224 serves a quality control tool, where contigs above 500 kb are kept as single bins if they show a completeness above 93%, otherwise they are sent back to the binning module. This approach enhances the recovery of high-quality and single-contig MAGs, outperforming traditional binning methods. After MAG de-replication, taxonomic annotation is achieved with GTDB-Tk225, and a complete graphical report is compiled automatically. One important caveat about this workflow is represented by its lack of assembly step, and hence the user must prepare the assembly of the PacBio sequences beforehand using tools such as hifiasm123 in its meta version, metaFlye116, OPERA-MS117, among others. Hi-Fi-MAG-Pipeline requires Conda as software manager, and it is orchestrated by Snakemake61. 2.3 Mapler124 Mapler is a pipeline specifically designed to handle PacBio HiFi long reads. Mapler workflow is orchestrated by Snakemake along with Conda for package management, enabling scalable execution on local or cluster systems. Regarding the specific tools encompassed by Mapler, state-ofthe-art assemblers such as metaMDBG125, hifiasm-meta123, metaFlye116 and OPERA-MS117 are available, with MetaBAT23 as the binning tool. Later on the workflow, each bin is classified taxonomically via GTDB-Tk225 or Kraken211, and genome quality is evaluated using CheckM224 standards. Mapler aligns reads back to contigs with Minimap294 to compute novel metrics including the aligned read percentage and aligned base percentage, stratified across quality categories. It is important to mention that Mapler accepts assemblies and bins as input to skip part of the process, and it includes a parallel analysis, where assembled versus unassembled reads are contrasted by evaluating k-mer distributions (KAT126), read quality (FastQC28), and taxonomic composition (Kraken2 + Krona41). As a result, by combining classic bin-based metrics with read-to-contig alignment statistics, Mapler assists in estimating how much of the sequence diversity remains uncaptured. 2.4 NanoPhase127 NanoPhase is a pipeline that enables building high-quality MAGs from ONT long reads, optionally enhanced with short read-based MAG polishing. The backbone of the pipeline is represented by an assembly with metaFlye116 followed by contig binning with MetaBAT23 and MaxBin25, and bin refinement with a MetaWRAP23 module. To estimate abundance and coverage, the contigs are mapped against the reads, and several polishing rounds with Racon128 and medaka, complete the workflow to generate high-accuracy final bins; If the user decides to include short reads in the analysis, these are used for polishing with Pilon129. Complementary, MetaQuast93 and CheckM64 v1.0 are in charge of MAG quality control, IDEEL130 evaluates the fraction of predicted full-length proteins in each MAG, full-length proteins are detected via alignment with UniProtKB131, and Prokka58 serves as functional annotation software. Remarkably, NanoPhase allows prophage and active prophage identification within the reconstructed MAGs with VIBRANT132 and PropagAtE133. Among pipeline technical specifications, this pipeline requires Conda as package manager and it offers parallelized execution with GNU Parallel to speed up the analysis. 3. Dual pipelines 3.1 GEN-ERA134 GEN-ERA suite is a collection of Nextflow9 pipelines aiming at supporting MAG reconstruction and annotation with as many methodologies as possible starting from either short or long reads. Specifically, this toolbox counts with more that 10 workflows specifically designed for tasks ranging from assembly and binning, quality assessment and decontamination, orthologous inference and maximum likelihood phylogenomic analyses, SSU rRNA phylogeny (constrained by ribosomal phylogenomic), Average Nucleotide Identity (ANI) clustering, taxonomic identification and metabolic modelling. Moreover, GEN-ERA incorporates specific tools designed to handle eukaryotic assembly annotation such as BRAKER2135 and AMAW136. Thus, GEN-ERA suits almost all requirements any user might demand given the variety of goals that can be achieved within a single software suite. From a technical point of view, operational GEN-ERA features, Nextflow-managed and Singularity-executed, ensures portability and reproducibility across environments. 3.2 Metagenomics-Toolkit137 Metagenomics-Toolkit is a workflow designed to increase scalability of task execution, enabling optimal resource allocation from its machine learning-optimized assembly step. This optimized assembly tailors the peak RAM value requested by a metagenome assembler to match actual requirements, thereby minimizing the dependency on dedicated high-memory hardware. Metagenomics-Toolkit is wrapped by Nextflow9 and powered with Docker containerization technology, and it can take either short or Oxford Nanopore (ONT) long reads as input. As a result, this pipeline is highly scalable and adaptable across computational infrastructures with a backbone workflow that relies on the traditional MAG-aimed steps such as quality control, assembly, binning, and annotation, plus an aggregation module that captures the output from each sample to “polish” the final MAGs. Regarding special features offered by Metagenomics-Toolkit, it offers plasmid identification based on various tools, the recovery of unassembled microbial community members, and the discovery of microbial interdependencies through a combination of dereplication, cooccurrence, and genome-scale metabolic modeling. 3.3 metaWGS138 metaWGS is one of the most recently released pipelines whose main differential is related with the possibility to assemble either short reads or long sequences (PacBio). This Nextflow9 pipeline is built off Singularity with consequent benefits this kind of setup brings as discussed previously. It incorporates a wide variety of tools as it must ensure a proper workflow for both types of sequencing technologies in a traditional end-to-end framework divided into 8 steps. The first step aims at cleaning and performing quality control with proper tools according to the input, while the second step allows the assembly of the sequences using either metaSPAdes35/MEGAHIT16 for short sequences and hifiasm123/metaFlye116 for PacBio reads. Following with the process, this pipeline filters the contigs and performs structural annotation during steps 3 and 4, respectively; step 5 is designed to estimate contig abundance by mapping them against the reads. Afterwards, a complete subworkflow for functional annotation is undergone with eggNOG-mapper67 at its core (step 6), and contig taxonomic affiliation is achieved through home-made scripts (step 7) to conclude with step 8, where the contigs are binned with MaxBin25, MetaBAT23 and CONCOCT4. Furthermore, metaWGS utilizes Binette139, a state-of-the-art binning refinement tool designed to construct high-quality MAGs from the output of multiple binning tools. As a special remark, metaWGS performs read taxonomic profiling via Kaiju, as well as contig annotation that includes an in-house algorithm and mapping against the reads. 3.4 MG-TK140 MG-TK (Metagenomic Toolkit) performs read assembly (SPAdes57, MEGAHIT16, Flye141, metaMDBG125) and binning (MetaBAT23, SemiBin221, MetaDecoder142), gene prediction, and clustering into nonredundant gene catalogs, followed by abundance estimation and functional annotation. It is structured around three main phases: processing raw sequences, building a gene catalog, and reconstructing species from MAGs with downstream phylogenetic analyses. It produces a wide range of outputs, including assemblies, MAGs, gene predictions, SNP calls and mapping outputs. A special remark MG-TK exhibits is its ability to generate detailed abundance matrices for both taxonomic and functional features, with hierarchical summaries available at multiple levels. The taxonomic profiles are reported using GTDB143 lineages, while functional annotations are provided for major databases such as KEGG144, SEED145, CAZy146, eggNOG45, and TCDB147. MG-TK also estimates completeness of functional modules, such as KEGG pathways, and links genes to multiple annotations for deeper exploration by the user. Beyond gene catalogs, MG-TK integrates MAG/MGS (Metagenomics Species) information, associating MAGs with their metagenomic species and providing detailed gene content, including representative MAGs for each species. Additionally, MGTK can provide assembly-independent profiles via a wide variety of tools including riboFinder148, MetaPhlAn47 and mOTUs102. 3.5 VEBA149 VEBA (Viral Eukaryotic Bacterial Archaeal) is a Conda-executed pipeline designed that enables the recovery and classification of genomes from all domains of life including archaeas, prokaryotes, microeukaryotes, and viruses. It starts with a common short read-preprocessing and assembly from which the process is bifurcated for prokaryotic and viral binning; unbinned contigs from the viral module are reincorporated into the prokaryotic contig set. Residual contigs from the prokaryotic module are then considered for eukaryotic MAG generation to proceed with the annotation and classification covering the genomes obtained in each module. Hence, several databases are considered at this step such as UniRef50/90150, MIBiG151, VFDB49, CAZy146, KOfamKOALA86, Pfam152, NCBIfam-AMR153 and AntiFam154. Also, a joint phylogeny is obtained based on MAG-gene models and lineage marker detection. An interesting approach VEBA follows is represented by the module coverage.py that collects all the unbinned contigs, from viral, eukaryotic and prokaryotic steps, to pursue a pseudo-coassembly, where iteratively the reference fasta (built from the contigs) and the sorted BAM files used as a final pass through prokaryotic and eukaryotic binning modules. This pseudo-coassembly approach is optional, being easily enabled during the workflow execution; the pipeline documentation widely discusses when this type of assembly should be used in specific cases. Notably, VEBA automates the detection of candidate phyla radiation (CPR) bacteria and integrates a consensus microeukaryotic database to optimize gene modeling and taxonomic classification. 4. Hybrid pipelines 4.1 Aviary155 Aviary is a modular, Snakemake61-based pipeline, with Conda as package manager, designed for single or hybrid metagenomic assembly and MAG recovery, supporting both short and long-read input sequences. The workflow is distributed in 8 modules following a traditional workflow starting with quality and diversity assessment of the reads, followed by a discriminated assembly according to the type of input, MEGAHIT16 or metaSPAdes35 for short reads only or metaFlye116 in case of long reads solely. For hybrid assembly the process is divided into four stages: polishing with Racon128 and Pilon129, metrics-based filtering, assembly and discard of low-quality bins and re-assembly with Unicycler156. The pipeline proceeds with a subsequent assembly evaluation in terms of fragmentation, misassembly detection and diversity quantification, and a complementary module moves forward with a read mapping of the assembly and abundance statistics calculation. To continue with the workflow, the contigs are binned using up to 6 tools (MetaBAT23, Rosella157, MetaBAT171, VAMB88, MaxBin25 and CONCOCT4) and refined afterwards with 5-time loop that includes CheckM224, Rosella Refine and DASTool7. The pipeline ends with MAG recovery assessment via CoverM106, CheckM2 and SingleM to proceed with MAG annotation through GTDB-Tk225, Prodigal39 and eggNOG45. Variant calling, ANI analysis and genotype recovery with Lorikeet158 are interesting attributes offered by Aviary as a complement to the traditional genomic feature detection. Aviary’s design presents a series of advantages that include the possibility of running modules, multi-sample handling and scalability across different computational infrastructures. 4.2 MUFFIN159 MUFFIN is a reproducible pipeline built with Nextflow9 designed for hybrid assembly by integrating short-read (Illumina) and long-read (nanopore) sequencing data. MUFFIN begins its workflow with a quality control of the reads (fastp53 and Filtlong) to progress through hybrid assembly (metaSPAdes35 or metaFlye116 with polishing) and differential binning (CONCOCT4, MetaBAT23, and MaxBin25). After bin refining with the MetaWRAP23 refinement module, a hybrid reassembly is pursued with Unicycler156. The pipeline ends with bin classification through CheckM64 v1.1 and sourmash13 (combined with GTDB143), and with bin annotation with eggNOG45 and a KEGG144 parser, providing high-quality, annotated MAGs and insights into the metabolic potential of the microbial community. Optionally, the user can provide metatranscriptomics data to perform a de novo transcript assembly (Trinity160), quantification (Salmon161) and annotation (eggNOG). Additionally, given its modularity design, the workflow can start as well with user-provided bins, differential reads or only RNA-seq data. MUFFIN can be executed with either Conda or Docker, and its native Nextflow features confer to it the possibility to restart the pipeline in case of failing, run on different computing infrastructures, multisample handling, among others. 4.3 nf-core/mag162 nf-core/mag is a Nextflow9 pipeline developed following the nf-core guidelines that ensures robustness and reproducibility. It supports both short-read and long-read sequences, as well as hybrid datasets, and it leverages a modular design, containerization (Docker, Singularity, among others) and 95. Dröge, J., Gregor, I. & McHardy, A. C. Taxator-tk: precise taxonomic assignment of metagenomes by fast approximation of evolutionary neighborhoods. Bioinformatics 31, 817–824 (2015). 96. Krapohl, J. & Pickett, B. E. SnakeWRAP: a Snakemake workflow to facilitate automated processing of metagenomic data through the metaWRAP pipeline. F1000Research 11, (2022). 97. Ziemski, M. et al. MOSHPIT: accessible, reproducible metagenome data science on the QIIME 2 framework. Preprint at https://doi.org/10.1101/2025.01.27.635007 (2025). 98. Bolyen, E. et al. Reproducible, interactive, scalable and extensible microbiome data science using QIIME 2. Nat. Biotechnol. 37, 852–857 (2019). 99. Narayanasamy, S. et al. IMP: a pipeline for reproducible reference-independent integrated metagenomic and metatranscriptomic analyses. Genome Biol. 17, 260 (2016). 100. Kopylova, E., Noé, L. & Touzet, H. SortMeRNA: fast and accurate filtering of ribosomal RNAs in metatranscriptomic data. Bioinformatics 28, 3211– 3217 (2012). 101. Bushnell, B. BBMap: A Fast, Accurate, Splice-Aware Aligner. LBL Publications, (2014). 102. Sunagawa, S. et al. Metagenomic species profiling using universal phylogenetic marker genes. Nat. Methods 10, 1196–1199 (2013). 103. Schudoma, C. Source code for: gff_quantifier. https://github.com/cschu/gff_quantifier (2023). 104. Tadrent, N. et al. SnakeMAGs: a simple, efficient, flexible and scalable workflow to reconstruct prokaryotic genomes from metagenomes. F1000Research 11, 1522 (2023). 105. Eren, A. M., Vineis, J. H., Morrison, H. G. & Sogin, M. L. A Filtering Method to Generate High Quality Short Reads Using Illumina Paired-End Technology. PLOS ONE 8, e66643 (2013). 106. Aroney, S. T. N. et al. CoverM: read alignment statistics for metagenomics. Bioinformatics 41, btaf147 (2025). 107. Schmidt, T. S. B. et al. SPIRE: a Searchable, Planetary-scale mIcrobiome REsource. Nucleic Acids Res. 52, D777–D783 (2024). 108. Coelho, L. P. et al. NG-meta-profiler: fast processing of metagenomes using NGLess, a domain-specific language. Microbiome 7, 84 (2019). 109. Seemann, T. Source code for: Barrnap-Bacterial ribosomal RNA predictor. https://github.com/tseemann/shovill (2018). 110. Seemann, T. Source code for: ABRicate-Mass screening of contigs for antimicrobial and virulence genes. https://github.com/tseemann/abricate (2020). 111. Santos-Júnior, C. D., Pan, S., Zhao, X.-M. & Coelho, L. P. Macrel: antimicrobial peptide screening in genomes and metagenomes. PeerJ 8, e10555 (2020). 112. Göbel, D., Stoye, J., Sczyrba, A., & Beckstette, M. The Cloud-based Workflow Manager (CloWM) - An integrated platform for highly scalable workflow execution. German Conference on Bioinformatics 2024 (GCB), Bielefeld. Zenodo. https://doi.org/10.5281/zenodo.14039069 (2024). 113. Clarke, E. L. et al. Sunbeam: An extensible pipeline for analyzing metagenomic sequencing experiments. Microbiome 7, 1–13 (2019). 114. Wood, D. E. & Salzberg, S. L. Kraken: ultrafast metagenomic sequence classification using exact alignments. Genome Biol. 15, R46 (2014). 115. Peng, K. et al. Benchmarking of analysis tools and pipeline development for nanopore long-read metagenomics. Sci. Bull. 70, 1591–1595 (2025). 116. Kolmogorov, M. et al. metaFlye: scalable long-read metagenome assembly using repeat graphs. Nat. Methods 17, 1103–1110 (2020). 117. Bertrand, D. et al. Hybrid metagenomic assembly enables high-resolution analysis of resistance determinants and mobile elements in human microbiomes. Nat. Biotechnol. 37, 937–944 (2019). 118. Kajitani, R. et al. MetaPlatanus: a metagenome assembler that combines long-range sequence links and species-specific features. Nucleic Acids Res. 49, e130 (2021). 119. Hu, J., Fan, J., Sun, Z. & Liu, S. NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics 36, 2253–2255 (2020). 120. Asnicar, F. et al. Precise phylogenetic analysis of microbial isolates and genomes from metagenomes using PhyloPhlAn 3.0. Nat. Commun. 11, 1–10 (2020). 121. Kim, D., Song, L., Breitwieser, F. P. & Salzberg, S. L. Centrifuge: rapid and sensitive classification of metagenomic sequences. Genome Res. 26, 1721– 1729 (2016). 122. Portik, D. M. et al. Highly accurate metagenomeassembled genomes from human gut microbiota using long-read assembly, binning, and consolidation methods. Preprint at https://doi.org/10.1101/2024.05.10.593587 (2024). 123. Cheng, H., Concepcion, G. T., Feng, X., Zhang, H. & Li, H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat. Methods 18, 170–175 (2021). 124. Maurice, N., Lemaitre, C., Vicedomini, R. & Frioux, C. Mapler: a pipeline for assessing assembly quality in taxonomically rich metagenomes sequenced with HiFi reads. Bioinformatics 41, btaf334 (2025). 125. Benoit, G. et al. High-quality metagenome assembly from long accurate reads with metaMDBG. Nat. Biotechnol. 42, 1378–1383 (2024). 126. Mapleson, D., Garcia Accinelli, G., Kettleborough, G., Wright, J. & Clavijo, B. J. KAT: a K-mer analysis toolkit to quality control NGS datasets and genome assemblies. Bioinformatics 33, 574–576 (2017). 127. Liu, L., Yang, Y., Deng, Y. & Zhang, T. Nanopore longread-only metagenomics enables complete and highquality genome reconstruction from mock and complex metagenomes. Microbiome 10, 209 (2022). 128. Vaser, R., Sovic, I., Nagarajan, N. & Sikic, M. Fast and accurate de novo genome assembly from long uncorrected reads. Genome Res. 27, 737-746 (2017). 129. Walker, B. J. et al. Pilon: An Integrated Tool for Comprehensive Microbial Variant Detection and Genome Assembly Improvement. PLOS ONE 9, e112963 (2014). 130. Stewart, R. D. et al. Compendium of 4,941 rumen metagenome-assembled genomes for rumen microbiome biology and enzyme discovery. Nat. Biotechnol. 37, 953–961 (2019). 131. The UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2025. Nucleic Acids Res. 53, D609–D617 (2025). 132. Kieft, K., Zhou, Z. & Anantharaman, K. VIBRANT: automated recovery, annotation and curation of microbial viruses, and evaluation of viral community function from genomic sequences. Microbiome 8, 90 (2020). 133. Kieft, K. & Anantharaman, K. Deciphering Active Prophages from Metagenomes. mSystems 7, e00084-22 (2022). 134. Cornet, L. et al. The GEN-ERA toolbox: unified and reproducible workflows for research in microbial genomics. GigaScience 12, 1–10 (2022). 135. Brůna, T., Hoff, K. J., Lomsadze, A., Stanke, M. & Borodovsky, M. BRAKER2: automatic eukaryotic genome annotation with GeneMark-EP+ and AUGUSTUS supported by a protein database. NAR Genomics Bioinforma. 3, lqaa108 (2021). 136. Meunier, L., Baurain, D. & Cornet, L. AMAW: automated gene annotation for non-model eukaryotic genomes. F1000Research 12, 186 (2023). 137. Belmann, P. et al. Metagenomics-Toolkit: the flexible and efficient cloud-based metagenomics workflow featuring machine learning-enabled resource allocation. NAR Genomics Bioinforma. 7, lqaf093 (2025). 138. Mainguy, J. et al. metagWGS, a comprehensive workflow to analyze metagenomic data using Illumina or PacBio HiFi reads. Preprint at https://doi.org/10.1101/2024.09.13.612854 (2024). 139. Mainguy, J. & Hoede, C. Binette: a fast and accurate bin refinement tool to construct high quality Metagenome Assembled Genomes. J. Open Source Softw. 9, 6782 (2024). 140. Hildebrand, F. et al. Dispersal strategies shape persistence and evolution of human gut bacteria. Cell Host Microbe 29, 1167-1176.e9 (2021). 141. Kolmogorov, M., Yuan, J., Lin, Y. & Pevzner, P. A. Assembly of long, error-prone reads using repeat graphs. Nat. Biotechnol. 37, 540–546 (2019). 142. Liu, C.-C. et al. MetaDecoder: a novel method for clustering metagenomic contigs. Microbiome 10, 46 (2022). 143. Parks, D. H. et al. GTDB: an ongoing census of bacterial and archaeal diversity through a phylogenetically consistent, rank normalized and complete genome-based taxonomy. Nucleic Acids Res. 50, D785–D794 (2022). 144. Kanehisa, M., Sato, Y., Kawashima, M., Furumichi, M. & Tanabe, M. KEGG as a reference resource for gene and protein annotation. Nucleic Acids Res. 44, D457– D462 (2016). 145. Overbeek, R. et al. The SEED and the Rapid Annotation of microbial genomes using Subsystems Technology (RAST). Nucleic Acids Res. 42, D206 (2014). 146. Drula, E. et al. The carbohydrate-active enzyme database: functions and literature. Nucleic Acids Res. 50, D571–D577 (2022). 147. Saier, M. H., Jr et al. The Transporter Classification Database (TCDB): 2021 update. Nucleic Acids Res. 49, D461–D467 (2021). 148. Cokelaer, T., Desvillechabrol, D., Legendre, R. & Cardon, M. ‘Sequana’: a Set of Snakemake NGS pipelines. J. Open Source Softw. 2, 352 (2017). 149. Espinoza, J. L. et al. Unveiling the microbial realm with VEBA 2.0: a modular bioinformatics suite for end-to-end genome-resolved prokaryotic, (micro)eukaryotic and viral multi-omics from either shortor long-read sequencing. Nucleic Acids Res. 52, e63 (2024). 150. Suzek, B. E. et al. UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics 31, 926–932 (2015). 151. Zdouc, M. M. et al. MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration. Nucleic Acids Res. 53, D678–D690 (2025). 152. Mistry, J. et al. Pfam: The protein families database in 2021. Nucleic Acids Res. 49, D412–D419 (2021). 153. Feldgarden, M. et al. AMRFinderPlus and the Reference Gene Catalog facilitate examination of the genomic links among antimicrobial resistance, stress response, and virulence. Sci. Rep. 11, 12728 (2021). 154. Eberhardt, R. Y. et al. AntiFam: a tool to help identify spurious ORFs in protein annotation. Database 2012, bas003 (2012). 155. Newell, R. J. P., Aroney, S. T. N., Zaugg, J., Sternes, P., Tyson, G. W., & Woodcroft, B. J. Aviary: Hybrid assembly and genome recovery from metagenomes with Aviary (v0.12.0). Zenodo. https://doi.org/10.5281/zenodo.15208119 (2025). 156. Wick, R. R., Judd, L. M., Gorrie, C. L. & Holt, K. E. Unicycler: Resolving bacterial genome assemblies from short and long sequencing reads. PLOS Comput. Biol. 13, e1005595 (2017). 157. Newell, R. J. P., Tyson, G. W., & Woodcroft, B. J. . Rosella: Metagenomic binning using UMAP and HDBSCAN (v0.5.3). Zenodo. https://doi.org/10.5281/zenodo.10460259 (2024). 158. Newell, R. J. P., McMaster, E. S., Craig, P., Boden, M., Tyson, G. W., & Woodcroft, B. J. Lorikeet: strainresolved metagenome analysis using local reassembly (v0.8.2). Zenodo. https://doi.org/10.5281/zenodo.10275469 (2023). 159. Damme, R. van et al. Metagenomics workflow for hybrid assembly, differential coverage binning, metatranscriptomics and pathway analysis (MUFFIN). PLOS Comput. Biol. 17, 1–13 (2021). 160. Grabherr, M. G. et al. Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat. Biotechnol. 29, 644–652 (2011). 161. Patro, R., Duggal, G., Love, M. I., Irizarry, R. A. & Kingsford, C. Salmon provides fast and bias-aware quantification of transcript expression. Nat. Methods 14, 417–419 (2017). 162. Krakau, S., Straub, D., Gourlé, H., Gabernet, G. & Nahnsen, S. nf-core/mag: a best-practice pipeline for metagenome hybrid assembly and binning. NAR Genomics Bioinforma. 4, (2022). 163. Wick, R. R., Judd, L. M., Gorrie, C. L. & Holt, K. E. Completing bacterial genome assemblies with multiplex MinION sequencing. Microb. Genomics 3, e000132 (2017). 164. Haveman, N. J. et al. Evaluating the lettuce metatranscriptome with MinION sequencing for future spaceflight food production applications. Npj Microgravity 7, 22 (2021). 165. De Coster, W. & Rademakers, R. NanoPack2: population-scale evaluation of long-read sequencing data. Bioinformatics 39, btad311 (2023). 166. Schubert, M., Lindgreen, S. & Orlando, L. AdapterRemoval v2: rapid adapter trimming, identification, and read merging. BMC Res. Notes 9, 88 (2016). 167. Antipov, D., Korobeynikov, A., McLean, J. S. & Pevzner, P. A. hybridSPAdes: an algorithm for hybrid assembly of short and long reads. Bioinformatics 32, 1009–1015 (2016). 168. von Meijenfeldt, F. A. B., Arkhipova, K., Cambuy, D. D., Coutinho, F. H. & Dutilh, B. E. Robust taxonomic classification of uncharted microbial sequences and bins with CAT and BAT. Genome Biol. 20, 217 (2019). 169. Levy Karin, E., Mirdita, M. & Söding, J. MetaEuk— sensitive, high-throughput gene discovery, and annotation for large-scale eukaryotic metagenomics. Microbiome 8, 48 (2020). 170. Borry, M., Hübner, A., Rohrlach, A. B. & Warinner, C. PyDamage: automated ancient damage identification and estimation for contigs in ancient DNA de novo assembly. PeerJ 9, e11845 (2021). 171. Karlicki, M., Antonowicz, S. & Karnkowska, A. Tiara: deep learning-based classification system for eukaryotic sequences. Bioinformatics 38, 344–350 (2022). 172. Camargo, A. P. et al. Identification of mobile genetic elements with geNomad. Nat. Biotechnol. 42, 1303– 1312 (2024). 173. Almeida, F. M. de, Campos, T. A. de & Pappas, G. J. Scalable and versatile container-based pipelines for de novo genome assembly and bacterial annotation. F1000Research 12, 1205 (2023). 174. Koren, S. et al. Canu: scalable and accurate longread assembly via adaptive k-mer weighting and repeat separation. Genome Res. 27, 722–736 (2017). 175. Schwengers, O. et al. Bakta: Rapid and standardized annotation of bacterial genomes via alignment-free sequence identification. Microb. Genomics 7, 000685 (2021). 176. Jolley, K. A. & Maiden, M. C. BIGSdb: Scalable analysis of bacterial genome variation at the population level. BMC Bioinformatics 11, 595 (2010). 177. Graham, E. D., Heidelberg, J. F. & Tully, B. J. Potential for primary productivity in a globallydistributed bacterial phototroph. ISME J. 12, 1861– 1866 (2018). 178. Blin, K. et al. antiSMASH 7.0: new and improved predictions for detection, regulation, chemical structures and visualisation. Nucleic Acids Res. 51, W46–W50 (2023). 179. Hu, K., Huang, N., Zou, Y., Liao, X. & Wang, J. MultiNanopolish: refined grouping method for reducing redundant calculations in Nanopolish. Bioinformatics 37, 2757–2760 (2021). 180. Tamames, J. & Puente-Sánchez, F. SqueezeMeta, a highly portable, fully automatic metagenomic analysis pipeline. Front. Microbiol. 10, 3349 (2019). 181. Bushmanova, E., Antipov, D., Lapidus, A. & Prjibelski, A. D. rnaSPAdes: a de novo transcriptome assembler and its application to RNA-Seq data. GigaScience 8, giz100 (2019). 182. Caspi, R. et al. The MetaCyc database of metabolic pathways and enzymes - a 2019 update. Nucleic Acids Res. 48, D445–D453 (2020). 183. Olson, R. D. et al. Introducing the Bacterial and Viral Bioinformatics Resource Center (BV-BRC): a resource combining PATRIC, IRD and ViPR. Nucleic Acids Res. 51, D678–D689 (2023). 184. Gillespie, J. J. et al. PATRIC: the Comprehensive Bacterial Bioinformatics Resource with a Focus on Human Pathogenic Species. Infect. Immun. 79, 4286–4298 (2011). 185. Brettin, T. et al. RASTtk: A modular and extensible implementation of the RAST algorithm for building custom annotation pipelines and annotating batches of genomes. Sci. Rep. 5, 8365 (2015). 186. Wang, S., Sundaram, J. P. & Spiro, D. VIGOR, an annotation program for small viral genomes. BMC Bioinformatics 11, 451 (2010). 187. The Galaxy Community et al. The Galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2022 update. Nucleic Acids Res. 50, W345–W351 (2022). 188. Kalantar, K. L. et al. IDseq—An open source cloudbased pipeline and analysis service for metagenomic pathogen detection and monitoring. GigaScience 9, giaa111 (2020). 189. Chen, I.-M. A. et al. The IMG/M data management and analysis system v.7: content updates and new features. Nucleic Acids Res. 51, D723–D732 (2023). 190. Kanehisa, M., Furumichi, M., Sato, Y., IshiguroWatanabe, M. & Tanabe, M. KEGG: integrating viruses and cellular organisms. Nucleic Acids Res. 49, D545–D551 (2021). 191. Galperin, M. Y. et al. COG database update 2024. Nucleic Acids Res. 53, D356–D363 (2025). 192. Haft, D. H. et al. TIGRFAMs and Genome Properties in 2013. Nucleic Acids Res. 41, D387–D395 (2013). 193. Arkin, A. P. et al. KBase: The United States Department of Energy Systems Biology Knowledgebase. Nat. Biotechnol. 36, 566–569 (2018). 194. Seaver, S. M. D. et al. The ModelSEED Biochemistry Database for the integration of metabolic annotations and the reconstruction, comparison and analysis of metabolic models for plants, fungi and microbes. Nucleic Acids Res. 49, D575–D588 (2021). 195. Richardson, L. et al. MGnify: the microbiome sequence data analysis resource in 2023. Nucleic Acids Res. 51, D753–D759 (2023). 196. Finn, R. D., Clements, J. & Eddy, S. R. HMMER web server: interactive sequence similarity searching. Nucleic Acids Res. 39, W29–W37 (2011). 197. Weber, N. et al. Nephele: a cloud platform for simplified, standardized and reproducible microbiome data analysis. Bioinformatics 34, 1411– 1413 (2018). 198. Standeven, F. J., Dahlquist-Axe, G., Speller, C. F., Meehan, C. J. & Tedder, A. An efficient pipeline for creating metagenomic-assembled genomes from ancient oral microbiomes. Preprint at https://doi.org/10.1101/2024.09.18.613623 (2024). 199. Jónsson, H., Ginolhac, A., Schubert, M., Johnson, P. L. F. & Orlando, L. mapDamage2.0: fast approximate Bayesian estimates of ancient DNA damage parameters. Bioinformatics 29, 1682–1684 (2013). 200. Zhao, D. et al. Eukfinder: a pipeline to retrieve microbial eukaryote genome sequences from metagenomic data. mBio 16, e00699-25 (2025). 201. Van Nguyen, H. & Lavenier, D. PLAST: parallel local alignment search tool for database comparison. BMC Bioinformatics 10, 329 (2009). 202. Lin, H.-H. & Liao, Y.-C. Accurate binning of metagenomic contigs via automated clustering sequences using information of genomic signatures and marker genes. Sci. Rep. 6, 24175 (2016).