Full text
AI-assisted research data annotation in biomedical consortia Felix Engel, Manuel Watter, Gita Benadi, Claudia Giuliani, Aref Kalantari, Harald Binder, Klaus Kaier Institute of Medical Biometry and Statistics, Faculty of Medicine and Medical Center – University of Freiburg Research Data Management Institute of Medical Biometry and Statistics Assessment of Requirements Schema Development Schema Testing & Tuning Data Annotation Data Presentation Data Reuse The Institute of Medical Biometry and Statistics (IMBI) supports biomedical collaborative research centers (CRCs) with research data management. Together with consortium researchers we develop metadata schemas for annotating the CRC research data output. Researchers use IMBI’s own research data platform fredato to enrich their datasets with metadata. 1234 Annotation Prediction In a four-step process pipeline, LLMs analyse a paper published in the CRC to suggest metadata relevant to the research data that were used in the research documented in the publication. Freely identified keywords are grounded by alignment with the PubTator³ software and the CRC’s metadata schema. This results in suggestions for both data annotation and schema refinement. Evaluation by research data creators has demonstrated a high precision for the annotation prediction. Still, researchers are always kept in the loop to eliminate incorrect suggestions and to ensure sufficient metadata coverage. LLM-based annotation prediction is currently being implemented as a feature in fredato. Schema Inference Large Language Models (LLMs) are presented with current literature from the CRC’s area of research. Output is a list of essential keywords that is discussed with CRC researchers and serves as a basis for schema development and refinement. Automated Dataset Detection and Annotation To take annotation prediction to a new level, we have tasked LLMs to search published journal articles for references to published datasets. The models were asked to visit the repository landing pages of these sources and to suggest annotation keywords from the information found there. These suggestions are further enriched with information from the article itself. This approach provides a more complex and richer annotation template for scientists to complete and approve. Complete and machine-intelligible metadata support further processing with tools powered by artifical intelligence (AI). AI-Readiness Engel, F., Benadi, G., Giuliani, C., Werner, J., Watter, M., Zeiser, R., Köttgen, A., Binder, H., & Kaier, K. (2025). Development of Metadata Schemas For Collaborative Research Centers. FreiData. https://doi.org/10.60493/K1XE3-NPC10 Watter, M., Kahle, L., Brunswiek, B., Fichtner, U., Pfaffenlehner, M., Werner, F., Gebele, D., Binder, H., & Knaus, J. (2023). Standardized metadata collection in a research data management tool to strengthen collaboration in Collaborative Research Centers. EScience-Tage, Heidelberg. https://doi.org/10.11588/HEIDOK.00033131 Giuliani, C., Benadi, G., Engel, F., Werner, J., Watter, M., Schwarzer, G., Groß, O., Zeiser, R., Binder, H., & Kaier, K. (2025). Identifying biomedical entities for datasets in scientific articles – A 4-step cache-augmented generation approach using GPT-4o and PubTator 3.0. medRxiv. https://doi.org/10.1101/2025.03.04.25323310 Watter, M, Giuliani, C., Benadi, G., Engel, F., Binder, H., Kaier, K. (2025) Automated Identification of Contextually Relevant Biomedical Entities with Grounded LLMs. medRxiv. https://doi.org/10.1101/2025.07.07.25331004 This work is licensed under https://creativecommons.org/licenses/by/4.0/