Artificial intelligence supported data provenance of DNA High-troughput sequencing data
Abstract
Poster was prepared for the 1st Conference on Research Data Infrastructure Abstract: The NCBI-SRA (Sequence Read Archive) serves as a vital online repository, housing an extensive collection of genetic sequences and associated metadata contributed by researchers worldwide. The availability of correctly annotated data within this repository fosters the reuse of sequencing information for novel analyses and meta-studies. Utilizing such data is particularly crucial for conducting large-scale genomic, metagenomic, and taxonomic studies, as it grants unique insights into the diverse mechanisms governing our world. However, to ensure the reliability and accuracy of results, researchers must often validate the annotated data when reusing deposited sequences, as mislabeled data can lead to false or unreliable outcomes. Addressing this challenge, we present a study that uses the power of Machine Learning (ML) in research data management (RDM) to enhance the identification and prevention of mislabeled metagenomic data within the SRA database. Specifically, we used a trained random forest model to classify metagenomic sequences using the SRA metadata database (last updated on 2022-08-23) while excluding sequences already classified in the previous PARTIE update (last updated on 2020-09-25). PARTIE, a widely available tool, extracts relevant features from submitted sequences using a sub-sampling approach and employs a trained random forest model to classify them into three categories: Whole Genome Sequencing (WGS), amplicon sequencing (Amplicon), or other data types (Other). Our investigation, conducted on 2023-01-13, encompassed 8,206,324 samples from the SRAmetadb metadata database. Among these samples, 844,339 were labeled as WGS, 518,600 as Transcriptome Analysis, and 334,047 as Metagenomic. Notably, 3,787,444 samples lacked labels, 2,479,362 were classified as Other, and the remaining 242,532 samples had different labels (e.g. Population Genomics, Cancer Genomics...). To identify potential mislabeled metagenomic data, we cross-checked the classified run accessions from the PARTIE-provided file with the run accessions in the SRAmetadb database. Our analysis revealed 35,564 samples labeled as Metagenomic that had not undergone PARTIE classification. Leveraging the random forest-trained model, we classified these samples, resulting in 3,311 being labeled as WGS, 13,755 as Amplicon, and 18,498 as Other. Additionally, an in-depth exploration of the SRAmetadb uncovered 4,748,560 samples with run accessions yet to be classified by PARTIE, encompassing study types such as Other, Whole Genome Sequencing, or lacking a label. We plan to subject these run accessions to classification using our local cluster. Our study exemplifies the remarkable potential of AI in enabling FAIR (Findable, Accessible, Interoperable, and Reusable) data usage. It demonstrates how software components can be employed to track metadata and provenance for reused data. The classification of these sequences provides a validated foundation for constructing large, standardized, and manually curated metagenomic metadata databases. Such databases improve the quality and reliability of future metagenomic studies across diverse environments and promise to uncover the underlying mechanisms governing their operation. Integrating Machine Learning techniques with research data management paves the way for enhanced data annotation, quality control, and subsequent advancement in metagenomics. Using AI approaches, researchers can confidently access and utilize vast sequencing data, ensuring accurate results and accelerating scientific discoveries in genomics, metagenomics, and taxonomy.