LLM-Based Metadata and Context Extraction for Agricultural Long-Term Experiments
Abstract
Long-term agricultural field experiments (LTEs) provide essential data for crop modeling and understanding agricultural processes under global change. Their reuse is limited by poor FAIR data practices and limited focus on specific experimental factors. Initiatives like the LTE-Overview Map have made a leap in metadata collection. However, manual processes remain inefficient.Our approach uses Large Language Models (LLMs) to automatically extract standardized, interoperable metadata from scientific literature, enhancing data discovery and reusability.
Full text
Susanne Lachmuth, Carsten Hoffmann, Gunnar Lischeid (Leibniz Center for Agricultural Landscape Research (ZALF) on behalf of the FAIRagro consortium. In collaboration with Cenk Dönmez, Masahiro Ryo (both ZALF). Contact: [email protected] LLM-Based Metadata and Context Extraction for Agricultural Long-Term Experiments Long-term agricultural field experiments (LTEs) provide essential data for crop modeling and understanding agricultural processes under global change. Their reuse is limited by poor FAIR data practices and limited focus on specific experimental factors. Initiatives like the LTE-Overview Map have made a leap in metadata collection. However, manual processes remain inefficient. Our approach uses Large Language Models (LLMs) to automatically extract standardized, interoperable metadata from scientific literature, enhancing data discovery and reusability. LTE Overview Map •680 LTE globally, focus on Europe •Comprehensive metadata enabling: •Reviews of LTE practices •Analysis of soil and agricultural indicators •Advanced spatial modeling •Remaining challenges: incomplete metadata and labor-intensive manual collection processes LTE metadata extractor: Literature LTE V140 at ZALF Müncheberg tools.bonares.de/ltfe Icons: Esha chaudary, WR Graphic Garage, Lourenchyus, Hendy Perkasa, pro symbols, icon sea, moch rizki eko waluyo, iconcheese, Yopi Kwb, IconPai via thenounproject.org Standardized output Prompts LLMs Pydantic Pydantic models Structured and validated metadata ? Data–applicationmatrix Measure 3.3 FAIRagro infrastructure Define clear structures that LLMs must adhere to Pydantic enforces type checking / validation Seamless serialization Enable application of domain-specific rules Produce schemas that can be included in prompts Convert types and apply limitations (min/max, regex, etc.) to ensure data quality 1) This work uses the Scientific Compute Cluster at GWDG, the joint data center of Max Planck Society forthe Advancement of Science (MPG) and University of Göttingen. In part funded bythe DFG–405797229 Past LTE data use ? Poster 15