The use of web resources for metabolomics in horticultural crops
Full text
Karakasetal. Horticulture Advances (2025) 3:18 https://doi.org/10.1007/s44281-025-00073-8 REVIEW Open Access © The Author(s) 2025. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http:// creat iveco mmons. org/ licen ses/ by/4. 0/. The use ofweb resources formetabolomics inhorticultural crops Esra Karakas1, Mustafa Bulut1 and Alisdair R. Fernie1* Abstract Metabolomics, a rapidly evolving field, has revolutionized horticultural crop research by enabling comprehensive analysis of metabolites that influence plant yield, growth, quality and nutritional value. The integration of web-based resources, including databases, computational tools and analytical platforms has significantly enhanced metabolomics studies by facilitating data processing, metabolite identification and pathway analysis. Moreover, the application of machine learning algorithms to these web resources has further optimized data interpretation, enabling more accurate prediction of metabolic profiles. Publicly available reference libraries and bioinformatic tools support precision of breeding, postharvest quality assessment and ultimately improving crop yield and sustainability. In this mini-review, we explore the current status of the diverse range of plant metabolomics databases in horticultural crops, highlighting the synergy between machine learning and traditional bioinformatics methods, their applications, challenges and future prospects in advancing plant science and agricultural innovation. Keywords Metabolite databases, Bioinformatic tools, Plant metabolomics and horticultural crops Introduction Metabolomics emerged at the end of the last century, with the term coined by analogy with genomics and transcriptomics to refer to the metabolite complement of the cell in a review of Stephen Oliver (Oliver etal. 1998). However, arguably the paper by Fiehn and co-workers published in (2000) wherein gas chromatography (GC) coupled to mass spectrometry (MS) defined the chemical composition of a morphological and metabolic mutant of the model plant Arabidopsis thaliana, and describing the changes in the level of 326 analytes, is a more memorable landmark. This work thus greatly extended the early metabolite profiling study of Sauter etal (Sauter etal. 1991), which presented the technology as a means of putative classification of the mode-of-action of pesticides. Thus the advent of metabolomics in plants arguably preceded that in microbes and mammals although the approach was also rapidly adopted in these communities (Aharoni etal. 2023; Giera etal. 2022). During the next two decades, metabolomics had the considerable advantage over profiling technologies such as transcriptomics and proteomics in that it is not directly reliant on the genome sequence. In this period the species scope of metabolomics rapidly expanded, such that it was no longer merely a tool for identifying biomarkers of cellular circumstance but additionally a cornerstone of systems biology and an approach that could provide mechanistic insight into metabolic regulation (Fernie and Schauer 2009; Meyer etal. 2007). This advantage has subsequently disappeared following the widespread adoption of nextgeneration sequencing, and the lack of linear relationship between the genome and the metabolome now represents part of the problem in identification of unknown analytes (Fernie and Stitt 2012). However, it does mean that metabolomics resources are available for a very wide range of species (~ 500) including a large number of horticultural crops (such as tomatoes, cucumber, strawberry and pepper). Previous reviews have addressed resources available for metabolomics in general (Perez de Souza *Correspondence: Alisdair R. Fernie [email protected] 1 Max-Planck-Institute of Molecular Plant Physiology, 14476 Potsdam-Golm, Germany
Page 2 of 14 Karakasetal. Horticulture Advances (2025) 3:18 and Fernie 2024) as well as plant metabolomics more specifically (Perez de Souza etal. 2017; Tohge and Fernie 2009). However, given the sheer number of publications that have been published over the last quarter of a century, we feel that the time is ripe for a focused review only on metabolomics of horticultural crops. Such resources are particularly important in plants given that whilst the size of the metabolome for prokaryotes has been estimated at a couple of thousand, that of the plant kingdom dwarves these numbers, with estimates ranging between 200 000 and 1 million metabolites (Shen et al. 2023). Over the last two decades, metabolomics has been utilized to address a wide range of important questions in plant biology, including pathway structure (Shen etal. 2023), the influence of metabolism on growth (Meyer etal. 2007; Sulpice etal. 2009), plant ecology (Davey et al. 2008), various aspects of plant genetics including evolution and the domestication syndrome (Alseekh etal. 2021; Beleggia etal. 2016; Kliebenstein 2009) and detailed characterizations of the metabolic response to biotic and abiotic stressors. In the following review, we will discuss several issues the first being the availability of tools for chromatography since we last reviewed this in Perez de Souza and Fernie (2024), the number of resources has massively increased. Thus, we will update our overview of the most useful tools. Secondly, we will review the current status of the board variety of plant metabolomics databases. In this respect, we list sources of archived data and their respective volumes of data. We also briefly discuss recent meta-analyses, there is great potential for cross-study comparisons on metabolite responses in determining common responses between either genetic or environmental perturbations of metabolism. Finally, we will provide an outlook as to the horticulturally important questions that we will ultimately be able to address via the assembly of FAIR (Findable, Accessible, Interoperable, Reusable)-minded database assemblies. Overview ofmetabolomics andkey web‑based databases forhorticultural plant research The main goal of metabolomics is to unravel the intricate metabolic pathways and networks that coordinate diverse biological functions within organisms. Despite nearly two decades of untargeted metabolomics studies on various plant parts and growth conditions, a complete understanding of plant metabolomes remains elusive. Liquid chromatography-mass spectrometry (LC–MS) is widely used due to its high sensitivity and broad metabolite coverage (Theodoridis etal. 2012). However, a major challenge in metabolomics is the shortage of chemical standards, which limits the accurate identification of many detected metabolites. In horticultural plants, this challenge is particularly significant, as precise metabolite identification is crucial for understanding traits related to flavor, nutrition, stress resistance and post-harvest quality. Databases provide researchers with valuable resources to analyze and interpret metabolomics data fostering collaboration and accelerating discoveries in plant science. A wide range of metabolomics databases exist, each providing unique formation, from mass spectrometry spectra to metabolic pathways. These databases enable researchers to identify unknown metabolites, compare metabolic profiles across species and integrate data with omics approaches for a comprehensive understanding of plant metabolism. As metabolomics research has advanced, both databases and analytical tools have been continuously updated to incorporate the latest discoveries. To fully employ these resources, metabolomics software should be capable of (i) processing raw spectral data, (ii) statistical analysis of significantly expressed metabolites, (iii) integration with metabolite databases for accurate identification, (iv) multi-omics data integration and analysis and (v) bioinformatics analysis with advanced visualization of molecular interaction networks (Fig.1). Several metabolomics databases have been developed to support horticultural plants, providing essential resources for studying metabolism in fruits, vegetables and ornamental species. Databases such as the Kyoto Encyclopedia of Genes and Genomes (KEGG) (Kanehisa and Goto 2000) and KNApSAcK (Afendi etal. 2012) contain information on plant metabolites, including those from horticultural crops. The Plant Metabolic Network (PMN) (Schläpfer etal. 2017) hosts species-specific pathway databases, such as SolCyc for tomato and OryzaCyc for rice, aiding in the exploration of metabolic pathways relevant to crop quality and stress responses. By using these databases, researchers gain deeper insights into the metabolic diversity of horticultural plants. Building on these existing resources, we will next introduce some of the novel horticultural metabolomics databases that further enhance research by providing comprehensive metabolic profiles, species-specific datasets and advanced analytical tools tailored for horticultural crops. To enhance clarity and utility for researchers, we categorized over 30 metabolomics databases relevant to horticultural crops into four main groups based on analytical methodology, species focus, metabolite type and functional application. These categories include: (i) Methodology-Based Databases, (ii) Species-Specific and Crop-Focused Databases, (iii) Databases Emphasizing Secondary Metabolites and Natural Products and (iv) Meta-Analysis and Data Integration Tools. A summary of these databases, along with their respective links, is provided in Table1 for easy reference.
Page 3 of 14 Karakasetal. Horticulture Advances (2025) 3:18 Fig. 1 A comprehensive workflow of metabolomics sample preparation to data processing, illustrating key steps from sample preparation and analytical techniques (e.g., LC–MS, GC–MS) to data processing, statistical analysis, and metabolite identification using online databases. The figure highlights the integration of computational tools for peak detection, annotation, and pathway analysis, providing a systematic approach to metabolomics research
Page 4 of 14 Karakasetal. Horticulture Advances (2025) 3:18 Table 1 The recent metabolomics databases Metabolomics databases References Link Methodology Species Scope LC–MS/GC–MS Centric Databases METASPACE-ML (Wadie et al. 2024)https:// apps. embl. de/ metas pacec ontext/ LC–MS Multi-species MetFrag (Ruttkies et al. 2016)https:// msbi. ipbhalle. de/ MetFr ag/ LC–MS/MS Multi-species LIPID MAPS (Conroy et al. 2024)www. lipid maps. org LC–MS, GC–MS Multi-species LipidSig 2.0 (Liu et al. 2024)https:// lipid sig. bioin fomics. org/ LC–MS Multi-species LipidSuite (Mohamed and Hill 2021)http:// suite. lipidr. org LC–MS Multi-species COCONUT (Sorokina et al. 2021)https:// cocon ut. natur alpro ducts. net LC–MS Multi-species PathBank 2.0 (Wishart et al. 2024)https:// pathb ank. org/ Pathway integration Multi-species SCIPDb (Priya et al. 2023)http:// 223. 31. 159.3/ plant_ compl ete/ index_ orang esuns et. php LC–MS Multi-species Multi-Platform Databases (LC–MS, NMR) MetaboLights (Yurekten et al. 2024)https:// www. ebi. ac. uk/ metab oligh ts/ LC–MS, GC–MS, NMR Multi-species PaintOmics 4 (Liu et al. 2022)https:// paint omics. org/ Multi-omics Multi-species RefMetaPlant (Shi et al. 2024a)https:// www. biosi no. org/ RefMe taDB/ LC–MS, NMR Multi-species OmicsSuite (Miao et al. 2023)https:// omics suite. github. io/#/ Multi-omics Multi-species ModelSEED (Seaver et al. 2021)https:// model seed. org/ bioch em Genome-scale modelling Multi-species Horticultural Crop-Specific Databases TOMATOMET (Ara et al. 2021)https:// metab olites. in/ tomatofruits/ LC–MS, GC–MS Tomato PMhub 1.0 (Tian et al. 2024)https:// pmhub. org. cn/#/ LC–MS Multi-species MMHub (Li et al. 2020)https:// biodb. swu. edu. cn/ mmdb/ LC–MS Mullberry ArecaceaeMDB (Yang et al. 2023a)http:// areca ceaegdb. com/#/ LC–MS Arecaceae (Palms) LettuceGDB (Guo et al. 2023)https:// www. lettu cegdb. com/ LC–MS Lettuce BoGDB (Wang et al. 2022)http:// www. bogdb. com/ LC–MS Brassica oleracea BnIR (Yang et al. 2023b)http:// yangl ab. hzau. edu. cn/ BnIR LC–MS Brassica napus PlantMetSuite (Liu et al. 2023)https:// plant metsu ite. veryg enome. com/ LC–MS Multi-species Corriander Genomics Database (Song et al. 2020)http:// cgdb. bio2db. com/ LC–MS Corriander and carrot Metabolite Database for RTB (Price et al. 2020) (Supplementary files of the original article) LC–MS, GC–MS Banana, cassava, potato, sweet potato, yam Broad-Spectrum Plant Metabolomics Databases Plant Reactome Knowledgebase (Gupta et al. 2024)https:// plant react ome. grame ne. org Pathway-based Multi-species PMN 16 (Hawkins et al. 2025)https:// plant cyc. org/ Pathway-based Multi-species MetaCyc (Caspi et al. 2020)https:// metac yc. org/ Pathway-based Multi-species CropMetabolome (Shi et al. 2024b)http:// www. cropm etabo lome. com/ LC–MS, GC–MS Multiple Crops Metadb (Gao et al. 2024)http:// medme tadb. ynau. edu. cn LC–MS Multi-species HypoRiPPAtlas (Lee et al. 2023)https:// hypor ippat las. npana lysis. org/ LC–MS/MS Plants and microbes Databases Emphasizing Secondary Metabolites and Natural Products MPOD (He et al. 2022)http:// medic inalp lants. ynau. edu. cn/ LC–MS Medicinal and food plants CMAUP (Zeng et al. 2019)http:// bidd2. nus. edu. sg/ CMAUP/ LC–MS Medicinal plants Meta-Analysis and Data Integration Tools MODMS (Fang et al. 2023)https:// modms. lzu. edu. cn/ LC–MS, RNA-seq Medicago sativa PCMD (Hu et al. 2024)https:// yangl ab. hzau. edu. cn/ PCMD LC–MS, GC–MS Multi-species The Thing Metabolome Repository family (XMRs) (Sakurai et al. 2023)https:// metab olites. in/ plants/ LC–MS Multi-species WikiPathways (Slenter et al. 2018) wikipathways.org Pathway mapping Multi-species MDSi (Li et al. 2023)http:// sky. sxau. edu. cn/ MDSi. htm LC–MS/MS Setaria italica PEO (Koh et al. 2024)https:// expre ssion. plant. tools/ LC–MS Multi-species
Page 5 of 14 Karakasetal. Horticulture Advances (2025) 3:18 Methodology‑based databases LC–MS/GC–MS centric databases These databases primarily support high-throughput mass spectrometry workflows (e.g., LC–MS, GC–MS), offering spectral libraries, feature annotation tools and pathway analysis capabilities. METASPACE‑ML METASPACE-ML (Wadie et al. 2024) is a machine learning-driven approach that addresses this challenge by incorporating novel scores and optimizing False Discovery Rate (FDR) estimation. It surpasses its rule-based predecessor with greater precision, higher throughput and improved detection of low-intensity biologically relevant metabolites. It employs a FDR-controlled method, reporting metabolite ions at a specified confidence level by ranking them against computationally generated decoy ions. MetFrag MetFrag (Ruttkies etal. 2016) is an open-source tool for identifying small organic compounds through combinatorial fragmentation. It begins by obtaining candidate structures from databases like PubChem, ChemSpider, or KEGG, or by accepting user-uploaded structure data files. The tool then fragments these candidates using a bond dissociation approach and compares the resulting fragments with product ions in the measured mass spectrum to identify the best-matching compounds. LIPID MAPS LIPID Metabolites and Pathways Strategy (LIPID MAPS) (Conroy et al. 2024) is a comprehensive platform for organizing lipid structural and biochemical data. Established 20 years ago, its nomenclature and classification system is widely recognized as the community standard. The platform offers databases for lipid identification, software tools, and educational resources. In 2020, it became an ELIXIR-UK data resource. LIPID MAPS incorporates enhanced metadata, including literature, taxonomic information and improved interoperability, to support FAIR compliance. LipidSig 2.0 LipidSig (Lin etal. 2021) is the first web-based platform designed for comprehensive lipidomic data analysis. The upgraded LipidSig 2.0 (Liu et al. 2024) enhances efficiency by automating lipid identification, assigning 29 characteristics, and supporting 24 data processing methods. It streamlines key analytical tasks such as data preprocessing, lipid annotation, differential expression, enrichment, and network analysis, enabling researchers to explore lipid properties and their biological relevance with ease. LipidSuite LipidSuite (Mohamed and Hill 2021) is an end-to-end platform for differential lipidomics data analysis, offering a structured workflow for preprocessing, exploration, differential analysis, and enrichment analysis. It supports mwTab (Metabolomics Workbench), Skyline CSV Export, and numerical matrix formats. Users can interactively analyze data, filter subsets by sample type or lipid class, and apply univariate, multivariate, and unsupervised analysis methods. COCONUT COlleCtion of Open NatUral producTs (COCONUT) (Sorokina etal. 2021) is a web-based platform for browsing, searching, and downloading natural products. It offers simple and advanced searches, including structurebased queries and molecular feature searches, with no login required. Users can download datasets in various formats or access data programmatically via a REST API. The platform is Docker-based, allowing easy deployment for local installations and integration into workflows. PathBank 2.0 Since its first release in 2020, PathBank has become widely popular, with its pathway data integrated into several major omics databases, including Human Metabolome Database (HMDB) (Wishart etal. 2007), DrugBank (Wishart et al. 2008), MarkerDB (Wishart etal. 2021) and MetaboAnalyst (Xia etal. 2009). The recent update, PathBank 2.0 (Wishart etal. 2024), introduces several key improvements, including significant canonical pathways, improved pathway diagrams and descriptions, a focus on drug metabolism and mechanisms, better compatibility for presentations and publications, new tools for pathway filtering and taxonomy and added pathway analysis for enrichment visualization and calculation. SCIPDb The Stress Combinations and Their Interactions in Plants Database (SCIPDb) (Priya etal. 2023) provides information on plant responses to various stress combinations, covering morpho-physio-biochemical (phenome) and molecular (transcriptome and metabolome) aspects. SCIPDb serves as a comprehensive informatics hub for plant stress research, facilitating data mining on the phenome, transcriptome, trait-gene ontology, and datadriven studies to advance the understanding of combined stress biology. An analysis of 939 studies found that abiotic–abiotic stress combinations significantly affect crop
Page 6 of 14 Karakasetal. Horticulture Advances (2025) 3:18 yield. SCIPDb integrates metabolome data from six studies across four plant species, with drought and heat stress being the most studied. The database provides a detailed catalog of metabolites with their fold change values. Multi‑Platform Databases (LC–MS, NMR) MetaboLights MetaboLights (Yurekten et al. 2024) is a database for metabolomics studies, providing access to raw data and associated metadata for processing and analysis. Established in 2012, MetaboLights (Steinbeck et al. 2012) has undergone recent developments. The latest version features a two-stage process, where data uploaded via FTP or Aspera servers is then synchronized to the study directory. PaintOmics 4 PaintOmics (García-Alcalde etal. 2011) is a web server designed for integrating and visualizing multi-omics data through biological pathway maps. The latest version, PaintOmics 4, (Liu etal. 2022) introduces key enhancements, including support for three pathway databases— KEGG, Reactome, and MapMan—offering broader pathway insights for both animals and plants. Additionally, new metabolite analysis methods address gaps in traditional enrichment approaches. The metabolite hub analysis identifies key compounds influenced by gene expression changes, while the metabolite class activity analysis assesses whether a metabolic class contains more significant elements than expected, indicating experimental regulation. RefMetaPant RefMetaPlant (Shi etal. 2024a) is developed by collecting tissue samples from over 150 plant species representing the five major phyla: Bryophyta, Lycopodiopsida, Pteridophyta, Gymnospermae and Angiospermae. It integrates an extensive collection of metabolomics data, including a vast number of experimental mass spectra from various plant species, a comprehensive library of biologically relevant compounds sourced from public databases and a large dataset of mass spectra that combines both experimental and in silico data. Additionally, it provides a reference metabolome for a diverse range of plant species in ‘.msr’ and ‘.mgf’ formats, serving as a valuable resource for plant metabolomics research. OmicsSuite OmicsSuite (Miao et al. 2023) offers comprehensive, step-by-step workflows tailored for multi-omics applications, making it ideal for horticultural plant breeding and molecular mechanism studies. This tool empowers researchers to deeply explore the molecular data embedded within complex multi-omics datasets. OmicsSuite can process multi-omics raw data in a variety of formats, including FastA, FastQ, Mutation Annotation Format, mzML, Matrix, and HDF5. The platform prioritizes efficient data transfer and robust pipeline analysis functions. It enables users to create pre-publication-quality images and tables, allowing them to focus on deriving meaningful biological insights. ModelSEED ModelSEED (Seaver et al. 2021) is a key resource for draft genome-scale metabolic model construction based on microbial and plant genomes. Its newly released biochemistry database serves as the foundation for ModelSEED and KBase, distinguishing itself through compartmentalization, transport reactions, charge balancing, and proton balancing. It is community-extensible via GitHub and functions as a biochemical ‘Rosetta Stone’, integrating annotations from multiple tools and databases. Built from diverse chemical datasets, it applies standard transformations, redundancy filtering, and thermodynamic calculations for accuracy and consistency. Species‑specific andcrop‑focused databases Horticultural crop‑specific databases TOMATOMET TOMATOMET (Ara et al. 2021) was developed after analysing the compounds in the ripe fruit of 25 different tomato cultivars using liquid chromatography-Orbitrap MS to measure their exact masses. The data was then carefully selected and refined to create the TOMATONET metabolome database, which compiles a vast collection of chemical peaks. Among these, a significant portion of ion peaks was organized into recognized chemical families. The database allows searching peaks by their exact mass values, along with related information such as compound annotations, category classifications, distribution across cultivars, cross-species specificity, retention times, MS/MS spectra, and adduct ion types. PMhub 1.0 The Plant Metabolome Hub (PMhub 1.0) (Tian et al. 2024) provides comprehensive data on plant metabolites, including reference spectra, genetic foundations, chemical reactions, metabolic pathways and biological functions. It contains extensive information on a vast number of metabolites and high-resolution tandem mass spectrometry spectra, both standard and in silico. In addition to literature-based data, PMhub incorporates a large collection of detected features from multiple plant species, with a significant portion successfully annotated using high-resolution tandem mass spectrometry data. This source is further enriched by reference metabolites
Page 7 of 14 Karakasetal. Horticulture Advances (2025) 3:18 features, making it valuable tool for plant metabolomics research. MMHub MMHub (version 1.0) (Li etal. 2020) is the first open public database of mass spectra for small chemical compounds in mulberry leaves. It includes electrospray ionization tandem mass spectrometry (ESI–MS/MS) data from various mulberry resources, collected under independent experimental conditions. The database identifies and annotates numerous metabolites, providing details on their chemical structures. Additional information, such as PubChem data, molecular formulas, and metabolite classifications, is also available. MMHub is a valuable resource for researchers studying mulberry metabolomics, supporting the screening of quality resources and specific metabolites. ArecaceaeMDB The Arecaceae Multi-Omics Database (ArecaceaeMDB) (Yang etal. 2023a) provides comprehensive genomic and multi-omics data for the Arecaceae family. It includes genome sequences from seven accessions across six species, resequencing data from 1631 accessions, and 866 spatiotemporal transcriptome datasets. Additionally, 138 spatiotemporal metabolome datasets have been generated using LC–MS-based metabolic profiling. To enhance usability, ArecaceaeMDB offers 12 commonly used online tools, including Sequence Fetch, BLAST, SynVisio, and Enrichment Analysis. LettuceGDB LettuceGDB (Guo etal. 2023) is a comprehensive omics data hub for lettuce research, integrating genomic, transcriptomic, metabolomic, and phenotypic data alongside a collected research archive. It offers five core functions and eight built-in tools for genome browsing, gene analysis, and data exploration. With its extensive resources and analytical capabilities, LettuceGDB is a valuable platform for functional genomics and lettuce breeding. BoGDB Brassica oleracea Genome Database (BoGDB) (Wang etal. 2022) is the first cross-omics platform for B. oleracea, integrating genome, transcriptome, and metabolome data to support genetic research and molecular breeding. It features multiple functional modules, including gene search, heatmaps, genome browsing, and metabolic analysis, within a user-friendly interface. The database will continue expanding with new genomic data to enhance its resources. BnIR BnIR (Yang etal. 2023b) is a comprehensive multi-omics database for rapeseed (Brassica napus), integrating six omics datasets: genomics, transcriptomics, variomics, epigenetics, phenomics, and metabolomics. It provides numerous “variation–gene expression–phenotype” associations using multiple statistical methods. The platform includes advanced search and analysis tools to enhance data exploration. Case studies demonstrate its effectiveness in identifying candidate genes linked to specific traits and uncovering their regulatory mechanisms. PlantMetSuite PlantMetSuite (Liu etal. 2023) is a web-based platform for plant metabolomics analysis, offering interactive bioinformatics tools and databases for seamless multi-omics research. It requires no installation or programming skills, is freely accessible, and will be regularly updated with new libraries and methods. Designed for ease of use, PlantMetSuite empowers researchers to explore and interpret plant metabolomics data effectively. Coriander genomics database It is a combined genomic, transcriptomic and metabolic database for coriander (Song etal. 2020) which alongside carrot is a model for studying the evolution of the Apiaceae family. Whilst this is mainly focused on genomic and transcriptomic analyses it does contain metabolomic data which can be searched and compared with the more extensive nucleic acid-based data. Metabolite database forroot, tuber andbanana crops The database by Price etal. (2020) is an exemplary compound database and concentration range for metabolites detected in the major RTB (root, tuber and banana) crops: banana (Musa spp.), cassava (Manihot esculenta), potato (Solanum tuberosum), sweet potato (Ipomoea batatas), and yam (Dioscorea spp.), following metabolomics-based diversity screening of global collections held within the CGIAR institutes. The dataset including 711 chemical features provides a valuable resource regarding the comparative biochemical composition of each crop highlight diversity available for incorporation into crop improvement programs. Particularly, the tropical crops cassava, sweet potato and banana displayed more complex compositional metabolite profiles with representations of up to 22 chemical classes (unknowns excluded) than that of potato (which was itself originally a tropical species), for which only metabolites from only ten chemical classes were detected.
Page 8 of 14 Karakasetal. Horticulture Advances (2025) 3:18 Broad‑spectrum plant metabolomics databases Plant reactome knowledgebase Plant Reactome (Gupta etal. 2024) is a free, comprehensive plant pathway knowledge base. It features curated rice reference pathways and gene-orthology-based projections for 129 plant species, covering diverse taxa from single-cell photoautotrophs to higher plants. With 339 reference pathways, it encompasses metabolism, transport, hormone signaling, developmental regulation, and plant responses to environmental stimuli. Beyond a repository, Plant Reactome functions as an interactive platform for analyzing and visualizing omics data, including gene expression, interactions, proteomics, and metabolomics, within a rich biological context. PMN 16 The Plant Metabolic Network (PMN) (Hawkins et al. 2025) is a free online database on plant metabolism. The latest release includes a vast collection of metabolic pathways, enzymes, metabolites, biochemical reactions, and citations from a wide range of plant and green algal genomes. It also features PlantCyc, a comprehensive pan-plant reference database. This update expands on previous versions by incorporating additional genomes, including species from the African Orphan Crop Consortium and nonflowering plants. MetaCyc MetaCyc (Caspi etal. 2020) is a comprehensive, evidencebased database of metabolic pathways and enzymes across all life domains. It includes around 3000 pathways compiled from over 60,000 publications, making it the largest collection of its kind. MetaCyc serves as a key reference for metabolism and is used to generate organismspecific Pathway/Genome Database (PGDBs) (Karp etal. 2019) available on BioCyc.org. It supports various fields, including genome annotation, biochemistry, enzymology, metabolomics and metabolic engineering. CropMetabolome The CropMetabolome (Shi etal. 2024b) database is dedicated to constructing crop reference metabolomes and serving as a comprehensive repository for crop metabolomic data. It facilitates the storage, sharing and analysis of metabolic data while providing advanced analytical tools. Covering over 50 crops across eight categories, the database integrates both self-generated and publicly available mass spectral data (using LC–MS/MS platform). Notably, it establishes a reference metabolome for 59 crop species, similar to the reference genome in genomics, enhancing metabolomic research and applications. The annotated metabolites were classified into 12 groups, including alkaloids, flavonoids, lipids and terpenoids. Metadata is provided in the reference metabolome file (‘.msr’ format). This resource serves as a vital reference for metabolic and physiological studies in crops. Metadb MetaDb (Gao etal. 2024) is a comprehensive database providing publicly available data on medicinal plant genes, transcription factors, metabolic pathways, and metabolites. It serves as a valuable resource for studying natural product synthesis and metabolic regulation in plants. Additionally, MetaDb offers access to bioinformatics tools such as ChemDoole 2D, BLAST, and SWISS-MODEL, aiding researchers in deeper data analysis and interpretation. HypoRiPPAtlas HypoRiPPAtlas (Lee etal. 2023) is a ready-to-use atlas of hypothetical natural product structures for in silico tandem mass spectra searches. It is built using seq2ripp, a machine-learning tool that predicts ribosomally synthesized and post-translationally modified peptides (RiPPs) from genomic data. The atlas identifies RiPPs in microbes and plants and can be expanded to other natural product classes by incorporating additional biosynthetic logic. This resource enables large-scale exploration of biosynthetic pathways and chemical structures of RiPPs. Databases emphasizing secondary metabolites andnatural products MPOD MPOD (He etal. 2022) is a database compiling genomes and transcriptomes of medicinal plants along with newly sequenced data from six genomes, 28 transcriptomes, and five metabolomes. It enables orthologous gene queries, homology comparisons, and correlation analyses between metabolite distribution and gene expression. MPOD provides detailed insights into flavonoid, alkaloid, and terpenoid metabolic pathways. Additionally, it includes a “biosynthetic tools” module with bioinformatics tools like SynVisio, heatmap, and enrichment analysis to support synthetic biology research. CMAUP The Collective Molecular Activities of Useful Plants (CMAUP) (Zeng et al. 2019) provides a comprehensive molecular overview of a vast range of plant species, including medicinal, food, edible, agricultural, and garden plants traditionally used across numerous countries and regions. It integrates various target classes and activity levels, visually represented in a two-dimensional target-ingredient heatmap. Additionally, CMAUP illustrates
Page 9 of 14 Karakasetal. Horticulture Advances (2025) 3:18 the regulatory effects of these plants on gene ontologies, biological pathways, and diseases. Meta‑analysis anddata integration tools MODMS The multi-omics database of Medicago sativa (alfalfa) (MODMS) (Fang etal. 2023) database is a platform dedicated to alfalfa, designed to include seven key components: genomics, transcriptomics, variations, proteomics, metabolomics and guide RNA (gRNA) tools. The metabolomics module contains 13 metabolomics datasets and researchers could identify differential metabolites under various conditions by inputting treatment details and the names of alfalfa varieties. PCMD Plant Comparative Metabolome Database (PCMD) (Hu etal. 2024) is a comprehensive platform for comparing metabolic profiles across 530 plant species. It enables users to analyze species, metabolites, pathways and taxonomy through various online tools, including species comparison, metabolite enrichment and ID conversion. As the most comprehensive plant metabolomics database in terms of species coverage, it offers valuable insights into plant metabolic diversity. The thing metabolome repository family (XMRs) The Metabolome Repository Family Database (Sakurai etal. 2023) encompasses food, plant, and other metabolome repositories, allowing researchers to examine sample-specific localization of unknown compounds detected via LC–MS across diverse samples. This resource aids in identifying and prioritizing unknown metabolites. A set of application programming interfaces for XMRs streamlines access to metabolome data, supporting large-scale analysis and data mining. Additionally, various applications, such as integrated metabolome and genome analyses, are showcased. WikiPathways WikiPathways (Slenter etal. 2018) is a reliable and comprehensive pathway database. Initially focused on genes and proteins, it previously had limited metabolite annotations. Recent efforts have enhanced metabolic pathway data by mapping metabolites to database identifiers and enriching interaction details, doubling the number of annotated metabolite nodes. Additionally, OpenAPI documentation for web services and FAIR-compliant resource annotations have been introduced to improve compatibility with experimental omics data. MDSi Multi-omics database for Setaria italica (MDSi) (Li etal. 2023) is a database containing the xiaomi genome, including protein-coding genes and their expression data across various tissues from xiaomi and JG21 samples, visualized through xEFP in-situ. It also provides wholegenome resequencing (WGS) data for foxtail millets and green foxtails, along with corresponding metabolic data. Users can explore SNPs and Indels interactively and access bioinformatics tools such as BLAST, GBrowse, JBrowse, and map viewer for analysis. PEO Plant Expression Omnibus (PEO) (Koh etal. 2024) is a web application that provides gene expression insights across a wide range of plant species. It allows users to explore gene expression patterns in different organs, identify organ-specific genes, and discover co-expressed genes. Functional annotations are included to help uncover genetic modules and pathways, supporting comparative kingdom-wide gene expression analysis as well as the discovery of metabolic and developmental pathways. Case studies demonstrate its effectiveness in identifying genes involved in pollen coat biosynthesis in Arabidopsis and capsaicin biosynthesis in Capsicum annuum. The database is freely accessible. Bioinformatics tools formetabolomics analysis The plant metabolomics data sets are analyzed and interpreted with the involvement of a series of computational and bioinformatic steps. As the first step the acquired mass spectra are subsequently compared with the spectral reference databases or libraries, such as METLIN (Montenegro-Burke et al. 2020), KEGG (Kanehisa and Goto 2000) and CropMetabolome (Shi et al. 2024b), to identify the related metabolites. Accurate mass and fragmentation measurements, combined with informative data, are obtained using various software tools for metabolite annotation, including R packages such as CAMERA (Kuhl etal. 2012), RAMclust (Broeckling etal. 2014), xMSannotator (Uppal etal. 2017) and MetAssign (Daly etal. 2014). These tools use parameters like retention time, m/z values, isotopic patterns and adducts to enhance annotation accuracy. Metabolites are subsequently quantified by comparing the intensity of specific, medium or low molecular weight metabolites across various treatment groups or between different sample types. Individual metabolites are integrated into metabolic pathways and networks in plant populations in order to reveal potential metabolic relations and regulatory interactions. Various statistical analysis methods are employed to identify hidden features, detect trends and