scieee AI-readable full text Open interactive document viewer

Artificial intelligence supported data provenance of DNA High-troughput sequencing data

Martin Bole; Jonas Coelho Kasmanas; João Saraiva; Ulisses Nunes da Rocha

Abstract

Poster was prepared for the 1st Conference on Research Data Infrastructure Abstract: The NCBI-SRA (Sequence Read Archive) serves as a vital online repository, housing an extensive collection of genetic sequences and associated metadata contributed by researchers worldwide. The availability of correctly annotated data within this repository fosters the reuse of sequencing information for novel analyses and meta-studies. Utilizing such data is particularly crucial for conducting large-scale genomic, metagenomic, and taxonomic studies, as it grants unique insights into the diverse mechanisms governing our world. However, to ensure the reliability and accuracy of results, researchers must often validate the annotated data when reusing deposited sequences, as mislabeled data can lead to false or unreliable outcomes. Addressing this challenge, we present a study that uses the power of Machine Learning (ML) in research data management (RDM) to enhance the identification and prevention of mislabeled metagenomic data within the SRA database. Specifically, we used a trained random forest model to classify metagenomic sequences using the SRA metadata database (last updated on 2022-08-23) while excluding sequences already classified in the previous PARTIE update (last updated on 2020-09-25). PARTIE, a widely available tool, extracts relevant features from submitted sequences using a sub-sampling approach and employs a trained random forest model to classify them into three categories: Whole Genome Sequencing (WGS), amplicon sequencing (Amplicon), or other data types (Other). Our investigation, conducted on 2023-01-13, encompassed 8,206,324 samples from the SRAmetadb metadata database. Among these samples, 844,339 were labeled as WGS, 518,600 as Transcriptome Analysis, and 334,047 as Metagenomic. Notably, 3,787,444 samples lacked labels, 2,479,362 were classified as Other, and the remaining 242,532 samples had different labels (e.g. Population Genomics, Cancer Genomics...). To identify potential mislabeled metagenomic data, we cross-checked the classified run accessions from the PARTIE-provided file with the run accessions in the SRAmetadb database. Our analysis revealed 35,564 samples labeled as Metagenomic that had not undergone PARTIE classification. Leveraging the random forest-trained model, we classified these samples, resulting in 3,311 being labeled as WGS, 13,755 as Amplicon, and 18,498 as Other. Additionally, an in-depth exploration of the SRAmetadb uncovered 4,748,560 samples with run accessions yet to be classified by PARTIE, encompassing study types such as Other, Whole Genome Sequencing, or lacking a label. We plan to subject these run accessions to classification using our local cluster. Our study exemplifies the remarkable potential of AI in enabling FAIR (Findable, Accessible, Interoperable, and Reusable) data usage. It demonstrates how software components can be employed to track metadata and provenance for reused data. The classification of these sequences provides a validated foundation for constructing large, standardized, and manually curated metagenomic metadata databases. Such databases improve the quality and reliability of future metagenomic studies across diverse environments and promise to uncover the underlying mechanisms governing their operation. Integrating Machine Learning techniques with research data management paves the way for enhanced data annotation, quality control, and subsequent advancement in metagenomics. Using AI approaches, researchers can confidently access and utilize vast sequencing data, ensuring accurate results and accelerating scientific discoveries in genomics, metagenomics, and taxonomy.

Full text

Artificial intelligence supported data provenance of DNA High-troughput sequencing data Martin Bole1,2, Jonas Coelho Kasmanas1,2,3, João Saraiva1, Ulisses Nunes da Rocha1*, 1 Department of Environmental Microbiology, Helmholtz Centre for Environmental Research – UFZ GmbH, Leipzig, Saxony, Germany, 2 Department of Computer Science and Interdisciplinary Centre of Bioinformatics, University of Leipzig, Leipzig, Saxony, Germany 3 Institute of Mathematics and Computer Sciences, University of São Paulo, São Carlos, Brazil Contacts: [email protected],[email protected],[email protected],[email protected], @ulisses_rocha REFERENCES To date (Aug 2023) over 2.8 billion entries exist in sequence data repositories(3) The availability of correctly annotated (meta)data within public sequence repositories fosters the reuse of millions of samples (Meta)Data is particularly crucial for conducting large-scale genomic, metagenomic and taxonomic studies with thousands of samples(7,8) Manually validating the annotated data from deposited sequences can take months of work for large-scale studies(2) Study exemplifies the potential of AI in enabling FAIR data usage by using textmining ML for unifying label identifiers of sequences Classified sequences provide a valid foundation for constructing manually curated metagenomic metadata DB’s with thousands of samples Integration of ML techniques with RDM paves the way for enhanced data annotation and quality control, transforming genome sequences into text prior analysis AI approaches enable researchers to access and utilize vast amounts of sequencing data minimizing the time needed to gather validated data 1. For Biotechnology Information (U.S.), N. C. SRA hand-book. National Center for Biotechnology Information, Bethesda. 2. Lohr S. 2014. For big-data scientists, “janitor work” is key hurdle to insights, The New York Times 3. National Center for Biotechnology Information. (n.d.). GenBank statistics. NCBI. https://www.ncbi.nlm.nih.gov/genbank/statistics/ 4. Torres, P. J., Edwards, R. A., and McNair, K. A. (2017). PARTIE: a partition engine to separate metagenomic and amplicon projects in the sequence read archive. Bioinformatics, 33(15):2389–2391. 5. This work was supported by the BMBF-funded de.NBI Cloud within the German Network for Bioinformatics Infrastructure (de.NBI) (031A532B, 031A533A, 031A533B, 031A534A, 031A535A, 031A537A, 031A537B, 031A537C, 031A537D, 031A538A). 6. Schnicke, T., Langenberg, B., Schramm, G., Krause, C., Harzendorf, T., & Strempel, T. (2019). EVE - High-Performance Computing Cluster. Helmholtz-Zentrum für Umweltforschung GmbH - UFZ, Permoserstr. 15, 04318 Leipzig. https://wiki.ufz.de/eve/ 7. Parks, D.H., Rinke, C., Chuvochina, M. et al. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nat Microbiol 2, 1533–1542 (2017). https://doi.org/10.1038/s41564-017-0012-7 8. Pasolli, E., De Filippis, F., Mauriello, I.E. et al. Large-scale genome-wide analysis links lactic acid bacteria from food with the gut microbiome. Nat Commun 11, 2610 (2020). https://doi.org/10.1038/s41467-020-16438-8 9. Zhu, Y.J., Stephens, R.M., Meltzer, P.S., & Davis, S. (2013). SRAdb: query and use public next-generation sequencing data from within R. BMC Bioinformatics, 14. Classification of almost 5 million samples using two tools (PARTIE & BMDSRA) with estimated run time 3.671 h (5 months) per tool Analysis of accuracy and probability of correct labeling for every sample, using outlier analysis and confusion matrices Construction of large Metagenome Metadata DB’s, to improve the quality, reliability and re-usage of data DISCUSSION NEXT STEPS AND FUTURE PROSPECTS METHODOLOGY AND RESULTS EVE - UFZ High-Performance computing cluster(6) True Label Predicted Label Confusion Matrix of BMDSRA after 5-fold cross validation BACKGRUOND AIMS Here, we present a study that uses Machine Learning (ML) in to enhance identification and prevent mislabeled metagenomic data within the Sequence Read Archive (SRA) database (DB). Our study supports Research Data Management (RDM) efforts by utilizing textbased mining for sequence classification and, thus, fostering efforts to improve interoperability by unifying label identifiers. Our approach can be adapted to other types of metadata within the NFDI initiatives SRAmetaDB(1,9) Last updated 2022-08-26 Contains metadata from 8.206.324 sequences PARTIE(4) tool provided file Last updated 2020-09-25 Contains 2.294.941 classified sequences After check for duplicates 4.748.384 samples not classified by PARTIE(4) Estimated run time on cluster: 3.671 h Classifying with PARTIE(4) and in-house developed ML “Boosting Model for Differentiating Sequence Read Archive” sequences (BMDSRA) A validated foundation for constructing large, standardized, and manually curated Metagenomic Metadata DB’s. Additional 3.311 Metagenome samples to be included in the upcoming Metegenomic Metadata DB, combined with the 284.078 from the PARTIE(4) tool classification # Superscript numbers indicate references Other NA WGS Metagenomics Transcriptome Analysis Amplicon Sequencing Metagenome Other NA Amplicon Sequencing Other Metagenome (5) Amplicon Sequencing Isolated Genome Single-cell Amplified Genome Metagenome Amplicon Sequencing Isolated Genome Metagenome Single-cell Amplified Genome Inconsistency issue with labeling conventions in different DB’s regarding Whole Genome Sequencing (WGS), Metagenome/ic and Amplicon sequencing labels Available online Cross-check and check for duplicates 35.564 metagenomic samples Not classified by PARTIE(4) Run time in local EVE cluster(6): 27h h 30 min