scieee AI-readable full text Open interactive document viewer

AI-assisted research data annotation in biomedical consortia

Engel, Felix; Watter, Manuel; Benadi, Gita; Giuliani, Claudia; Kalantari Sarcheshmeh, Aref; Binder, Harald; Kaier, Klaus

Abstract

Annotation of research data is a key element of Open Science and has gained additional value as training input for artificial intelligence. However, developing metadata schemas poses a series of challenges, including optimisation and securing both complete coverage and constant completeness and quality. We employ large language models (LLMs) to address some of these challenges while keeping researchers in the loop to ensure reliability of annotations.Our research data management group currently supports seven biomedical research consortia. We develop customised metadata schemas together with consortium members, drawing on established controlled vocabularies (Engel et al. 2025). Schemas are implemented on the fredato research data platform developed at the IMBI (Watter et al. 2023). Schemas are documented and published as knowledge graphs adhering to the Resource Description Framework (RDF), relating metadata to research processes as modelled by commonly used ontologies.LLMs are employed to develop initial schema drafts from related research literature and to predict dataset annotations from scientific papers (Giuliani et al. 2025). The models have proved to perform well with these tasks, supporting researchers with improving metadata coverage in their consortia.

Full text

AI-assisted research data annotation in biomedical consortia Felix Engel, Manuel Watter, Gita Benadi, Claudia Giuliani, Aref Kalantari, Harald Binder, Klaus Kaier Institute of Medical Biometry and Statistics, Faculty of Medicine and Medical Center – University of Freiburg Research Data Management Institute of Medical Biometry and Statistics Assessment of Requirements Schema Development Schema Testing & Tuning Data Annotation Data Presentation Data Reuse The Institute of Medical Biometry and Statistics (IMBI) supports biomedical collaborative research centers (CRCs) with research data management. Together with consortium researchers we develop metadata schemas for annotating the CRC research data output. Researchers use IMBI’s own research data platform fredato to enrich their datasets with metadata. 1234 Annotation Prediction In a four-step process pipeline, LLMs analyse a paper published in the CRC to suggest metadata relevant to the research data that were used in the research documented in the publication. Freely identified keywords are grounded by alignment with the PubTator³ software and the CRC’s metadata schema. This results in suggestions for both data annotation and schema refinement. Evaluation by research data creators has demonstrated a high precision for the annotation prediction. Still, researchers are always kept in the loop to eliminate incorrect suggestions and to ensure sufficient metadata coverage. LLM-based annotation prediction is currently being implemented as a feature in fredato. Schema Inference Large Language Models (LLMs) are presented with current literature from the CRC’s area of research. Output is a list of essential keywords that is discussed with CRC researchers and serves as a basis for schema development and refinement. Automated Dataset Detection and Annotation To take annotation prediction to a new level, we have tasked LLMs to search published journal articles for references to published datasets. The models were asked to visit the repository landing pages of these sources and to suggest annotation keywords from the information found there. These suggestions are further enriched with information from the article itself. This approach provides a more complex and richer annotation template for scientists to complete and approve. Complete and machine-intelligible metadata support further processing with tools powered by artifical intelligence (AI). AI-Readiness Engel, F., Benadi, G., Giuliani, C., Werner, J., Watter, M., Zeiser, R., Köttgen, A., Binder, H., & Kaier, K. (2025). Development of Metadata Schemas For Collaborative Research Centers. FreiData. https://doi.org/10.60493/K1XE3-NPC10 Watter, M., Kahle, L., Brunswiek, B., Fichtner, U., Pfaffenlehner, M., Werner, F., Gebele, D., Binder, H., & Knaus, J. (2023). Standardized metadata collection in a research data management tool to strengthen collaboration in Collaborative Research Centers. EScience-Tage, Heidelberg. https://doi.org/10.11588/HEIDOK.00033131 Giuliani, C., Benadi, G., Engel, F., Werner, J., Watter, M., Schwarzer, G., Groß, O., Zeiser, R., Binder, H., & Kaier, K. (2025). Identifying biomedical entities for datasets in scientific articles – A 4-step cache-augmented generation approach using GPT-4o and PubTator 3.0. medRxiv. https://doi.org/10.1101/2025.03.04.25323310 Watter, M, Giuliani, C., Benadi, G., Engel, F., Binder, H., Kaier, K. (2025) Automated Identification of Contextually Relevant Biomedical Entities with Grounded LLMs. medRxiv. https://doi.org/10.1101/2025.07.07.25331004 This work is licensed under https://creativecommons.org/licenses/by/4.0/