scieee AI-readable full text Open interactive document viewer

D4.2 - Initial list of common terms and metadata standards

hunninck, Marijke; Willekens, Axel; Vangeyte, Jürgen; Wenzlaff, Marie

Abstract

This deliverable is an important step in establishing a standardised framework within AgrifoodTEF. It gathers essential terms and metadata standards, in line with the goal of Task 4.1 to create a common language and metadata standards fundamental for the interoperability of the diverse services offered within AgrifoodTEF. This includes standardised data labels, ontologies for efficient communication, and quality labels for data and AI models.

Full text

D4.2 INITIAL LIST OF COMMON TERMS AND METADATA STANDARDS DECEMBER 2023 Ref. Ares(2024)597007 - 26/01/2024 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 2 Project cofunded by the European Commission within the Digital Europe Programme Dissemination Level PU Public x CO Confidential, only for members of the consortium (including the Commission Services) □ CL Classified, as referred to in Commission decision 2001/844/EC □ This document supports Deliverable Number D4.2 Lead Beneficiary JR Deliverable Name Initial list of common terms and metadata standards Deliverable Description Initial list of common terms and meta-data standards used in the services Type R — Document, Report Dissemination Level PU - Public Due Date (month) 12 Work Package No WP4 Document Revision History Date Issue Author/Editor/Contributor Summary of main change 04/12/2023 V1.0 Hunninck – EV-ILVO First document version 18/12/2023 V1.1 Willekens; Vangeyte – EV-ILVO 3.1 reviewed and improved 19/12/2023 V1.2 Wenzlaff - JR Overall comments on V1.1 23/12/2023 V2.0 Willekens – EV-ILVO Improved 4.0 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 3 Table of Contents Executive summary ........................................................................................................................................................... 5 1. Introduction .................................................................................................................................................................. 7 2. Background ................................................................................................................................................................... 8 2.1. International standardization organizations .......................................................................................................... 9 2.2. ISO standards ....................................................................................................................................................... 11 ISO standards for agriculture .................................................................................................................................. 11 ISO/IEC JTC 1/SC 42: Standardization in the area of Artificial Intelligence ............................................................. 11 ISO/TC 211: Geographic information/Geomatics ................................................................................................... 12 3. Common terms ........................................................................................................................................................... 13 3.1. Service Data Management Plan Survey .............................................................................................................. 13 3.2. Common terms in data collection ........................................................................................................................ 13 Data Types and Categories: .................................................................................................................................... 13 Data formats: .......................................................................................................................................................... 14 Data collection locations: ........................................................................................................................................ 14 Frequency and Context: .......................................................................................................................................... 14 3.3. Common terms in data quality assessment ......................................................................................................... 14 Data quality assessment: ........................................................................................................................................ 14 AI model quality assessment: ................................................................................................................................. 15 Standardized Quality Labels for Data:..................................................................................................................... 15 4. Metadata Analysis ....................................................................................................................................................... 16 5. Ontology for metadata standards ............................................................................................................................... 20 5.1. Established ontology for the metadata: .............................................................................................................. 20 5.2. Maintaining ontology ........................................................................................................................................... 20 6. Data sharing and protection ....................................................................................................................................... 22 6.1. Data storage and data backup: ............................................................................................................................ 22 6.2. Methods used for data sharing: ........................................................................................................................... 22 6.3. Responsible for data management: ..................................................................................................................... 22 7. Working groups ....................................................................................................................................................... 23 8. Discussion, Conclusion and Next Steps ....................................................................................................................... 25 Annex 1 ........................................................................................................................................................................... 27 Annex 2 ........................................................................................................................................................................... 29 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 4 Abtract AgrifoodTEF, a consortium of European testing and validation infrastructures, is committed to supporting agri-food companies in developing AI and robotics solutions through field development within established facilities. The initiative aims to bridge the gap between research efforts in these areas and the tangible creation of products that support efficient and sustainable agriculture aligned with the stringent usability and economic requirements of end users. By leveraging existing experimental farms and AI/Robotics facilities in representative European regions, AgrifoodTEF consolidates resources to strengthen the European Union's central role in ensuring global food security. This consolidation focuses on testing and validating facilities involving the best experts in AI and robotics technology and relevant stakeholders. The strategic framework includes the creation of standards for dataset development, data sovereignty, algorithm interoperability and benchmarking. GAIA-X's agro-domain ambassadors across Europe play an integral role in this endeavor. Although each TEF node emphasizes independent operations tailored to specific specialties and regional needs, a coherent set of guidelines and standards is shared by all nodes. Mutual service support is a feature of the various regions represented by nodes and satellites. The overarching goal is to improve Europe's position by stimulating the widespread adoption of advanced Agri-Food solutions. Task 4.1 within this mission focuses on establishing common terms and metadata standards within AgrifoodTEF. This includes standardized data labeling formats, an ontology of terms and the development of data quality labels and AI models. Existing benchmarks such as ISO standards for AI, results of Horizon 2020 projects (e.g. Atlas and Demeter) and AEF standards form the basis for this effort. This task also initiates specialized working groups (e.g., robotics, computer vision, remote sensing) to facilitate implementation and identify gaps in existing standards. This task is closely related to Task 4.2, where standardization needs emerge and are coordinated for the development of new services. Through collaborative efforts, AgrifoodTEF strives to create a harmonized language and robust standards that strengthen the network's capacity for innovation and coherence in Agrifood technology. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 5 Executive summary This deliverable is an important step in establishing a standardized framework within AgrifoodTEF. It gathers essential terms and metadata standards, in line with the goal of Task 4.1 to create a common language and metadata standards fundamental for the interoperability of the diverse services offered within AgrifoodTEF. This includes standardized data labels, ontologies for efficient communication, and quality labels for data and AI models. After some background information, providing an overview of standards and organizations that form the backbone of uniform practices and protocols within the agri-food sector, this deliverable introduces the template for the Service Data Management Plan (sDMP). The sDMP template was jointly prepared and then discussed and finalized during a workshop of the project meeting in Osnabrück. To efficiently collect the required information on all services and general partners, nodes and satellites were tasked with completing the sDMP for their respective services. The sDMP survey revealed common terminologies used in the AgrifoodTEF portfolio of services, and categorized them across diverse dimensions, highlighting critical aspects such as data collection, quality assessment, and metadata management. This summary provides an overarching insight into the core aspects outlined in the sDMPs, serving as a foundational guide for subsequent operational steps within AgrifoodTEF: o Data Collection: Diverse data types are collected through varied methods and tools, highlighting the breadth and contextuality of data collection practices. o Data Quality and Use: Rigorous quality assessments, including AI model testing, standardized for reliability and suitability across services. The comprehensive explanation of data types, collection practices and quality assessment methods lays the foundation for robust data management and evaluation frameworks within the project. These efforts strengthen uniformity, consistency and integrity in handling data across facets of service delivery, paving the way for future joint developments. o Metadata: Comprehensive metadata management encompassing diverse information, tools, and obligations based on context. o Metadata Standards: Establishment through domain-specific vocabularies, aiming for clarity, understandability, and compliance with FAIR Data principles. A Metadata analysis is performed to examine the types of metadata that are most prevalent, the tools and platforms used for managing this metadata, and the standards applied to ensure consistency and quality across various services. While revealing insights into organizational practices, it also highlighted gaps due to low service readiness. This emphasizes the need for refined metadata strategies as services mature, ensuring robust data organization and management aligned with evolving needs. Another key observation is the predominant use of traditional folder structures for organizing metadata, indicating a preference for user-friendly approaches but possibly lacking in more sophisticated data management systems. Time stamps and location metadata emerge as findamental elements, highlighting chronological and geographical contexts in agrifood experiments. However, it may also indicate a lack of knowledge of data management systems that are much more capable of maintaining consistency and integrity of the data. o Previous European Projects: Varied engagement levels with prior projects, indicating differing approaches in leveraging their outcomes. Engagement in digital agrifood ecosystems through projects like Demeter and IOF2020 hasn't fully translated into strong integration of terminology linked to FAIR Data principles. This signifies a gap between engagement in projects and the adoption of standardized practices within AgrifoodTEF. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 6 o Data Sharing and Protection: Emphasis on secure and collaborative data management, ensuring compliance with ethical guidelines and regulations. Addressing the challenge of secure data sharing and protection is identified as essential. Establishing reliable data management practices to safeguard agricultural data and support collaborative efforts within AgrifoodTEF services remains a critical focus area. The sDMP survey also highlights the need for a comprehensive ontology to improve data consistency and interoperability. iIn this early stage, developing and integrating ontologies present both challenges and opportunities for the AgrifoodTEF project. As the project evolves, continuous refinement and adaptation of these ontologies will be indispensable to stay in line with service and customer needs. While there is a lot of work in implementing consistent, interoperable data management practices, it is essential to balance these efforts. Excessive emphasis on standardization may hinder the final project goal of deploying qualitative services. Data management is essential, yet it must always serve as a means to an end, facilitating improved services for our clients. This is where our working groups play an important role. Their task is to exchange information on existing data ontologies and management tools across various technical domains of AgrifoodTEF, such as robotics, sensor testing, and AI pipelines, to ensure service quality. These groups are key in ensuring that each step towards standardization and a common language is focused on enhancing service delivery. Overall, this pursuit of a common language and standardised framework within AgrifoodTEF is an ongoing journey, with this deliverable being only the starting point of a robust and dynamic process. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 7 1. Introduction This deliverable is an essential step towards a coherent and standardized framework within the AgrifoodTEF ecosystem. It focuses on compiling an initial list of common terms and metadata standards, which is fundamental for the various services offered within the AgrifoodTEF landscape, especially considering the project is still in its early stages This preliminary effort aligns with the primary objective of Task 4.1 to implement a universal language and structured framework, including the development of standardized formats for data labels and a consistent ontology of terms for efficient communication across all services. Emphasizing the early stage of the project, the deliverable also includes the formulation and implementation of quality labels for both data and AI models within the TEF are an integral part of this effort. Drawing on insights and inspiration from established standards such as ISO standards for AI, findings from important Horizon 2020 and HE projects such as Atlas and Demeter, and AEF standards, this work aims to merge and integrate relevant existing and emerging standards into the AgrifoodTEF framework. To ensure the effective adoption and use of these standards across the network, specialized working groups have been formed to focus on different topic areas (e.g., robotics, computer vision, remote sensing, etc.). These groups play a key role in evaluating existing standards and identifying potential gaps to refine and improve the standards landscape, closely linked to the standardization needs identified in Task 4.2. Furthermore, this deliverable functions as a central coordination point for aligning related standards with the development of new services within AgrifoodTEF. By synchronizing emerging services with relevant standards, it aims to ensure coherence, compatibility, and uniformity in the evolving landscape of agrifood innovation. In essence, this deliverable aims to start the groundwork for a standardized language and framework within the AgrifoodTEF, fostering collaboration, interoperability and innovation across services and initiatives in these early stages of the project. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 8 2. Background Common terms and metadata standards increase efficient and effective data sharing in the agrifood sector. Common standards can make different organizations confident that they are using the same data definitions and formats. This helps to avoid errors and create new opportunities for data analysis. In addition, the common standards enable the integration of different data sources into a single system. This can be important for organizations that create added value from the combination of various sources, such as remote sensors, registration or process monitoring tools. Overall, common terms and metadata standards improve the efficiency, accuracy, and interoperability of data in the agrifood sector. This can lead to better decision-making, improved food security or may assist in developing new datadriven technologies. Metadata serves as the foundation of digital curation and plays a major role in the accessibility and usability of digital resources. It encompasses descriptive and contextual information linked to another object or resource, typically organized into structured elements. Metadata not only describes the information resource but also aids users in locating and retrieving it while supporting content and access management. Without metadata, a digital resource may become practically impossible to retrieve, identify, or utilize. Metadata consists of various elements organized according to the specific functions they serve. A metadata standard typically supports multiple defined functions and outlines the elements required to fulfill these functions. These functions may include the following: • Descriptive Metadata: This type of metadata aids users in identifying, locating, and retrieving information resources. It often incorporates controlled vocabularies for classification and indexing and links to related resources. • Technical Metadata: Technical metadata details the processes used to create or the requirements for using a digital object. • Administrative Metadata: Administrative metadata manages various administrative aspects of a digital object, including intellectual property rights, acquisition information, and documentation related to the creation, alteration, and version control of the metadata itself (sometimes referred to as "meta-data"). • Use Metadata: Using metadata facilitates user access, tracks user interactions, and manages information related to multiple versions of a resource. • Preservation Metadata: Preservation metadata documents actions taken to preserve a digital resource, such as migration processes and checksum calculations. Metadata standards typically originate as schemas developed by specific user communities to accurately describe their resource types. This development process involves community consensus and formal processes for submitting, approving, and publishing new elements. Adopted metadata definitions are controlled and guided by the user community. Published specifications are often centralized by consortia, organizations or companies and are made accessible through websites or Metadata Registries. These specifications include semantic definitions of elements and standardized representations in digital formats. Semantic definitions encompass both Metadata Structure Standards, ensuring consistent structure for data sharing, and Metadata Content Standards, ensuring effective machine searches through consistent data entry and using controlled vocabularies like authority files, thesauri, or encoding schemes for access points. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 9 Metadata schemas evolve in response to community needs and can gain widespread acceptance during development. In some scenarios, there is maintenance by nationally or internationally recognized centers, such as the Library of Congress, or support from professional bodies to enhance visibility and adoption within the community. In addition to metadata standards, the agri-food industry relies on technical and data standards. Technical standards regulate the production, processing, and distribution of agri-food products, ensuring safety, quality, and environmental protection. Data standards define data structure and content, enabling data sharing, analysis, and transparency. The development and implementation of these standards contribute to the efficiency, effectiveness, and transparency of the agri-food sector. Furthermore, ontologies play a major role in structuring and organizing information, making it understandable and usable by humans and computers. They are integral to the semantic web, revolutionizing internet interactions by structuring information in a machine-readable manner. Ontologies find application in information modeling, data integration, improving search results, and training artificial intelligence systems to understand the world better. They are a powerful tool for enhancing the organization and accessibility of information in various contexts. 2.1. International standardization organizations Many organizations work with standards in the agriculture sector. Here are a few of the most prominent ones: • International Organization for Standardization (ISO). ISO is the world's largest developer of standards. It has developed standards that are relevant to the agriculture sector, including ISO 22000 on food safety, ISO 14001 on environmental management, and ISO 50001 on energy management. • International Electrotechnical Commission (IEC). IEC is the world's leading organization for developing and publishing international standards for all electrical, electronic, and related technologies. It has developed standards that are relevant to the agriculture sector, including IEC 62061 on functional safety of electrical/electronic/programmable electronic safety-related systems. • International Plant Protection Convention (IPPC). IPPC is an international organization that protects plants from pests and diseases. It has developed several standards that are relevant to the agriculture sector, including ISPM 15 on phytosanitary measures for wood packaging materials. • Codex Alimentarius Commission (CAC). CAC is a joint FAO/WHO food standards program that develops international standards, codes of practice, guidelines, and other recommendations for food safety. It has developed standards that are relevant to the agriculture sector, including CAC/RCP 1-2020 on general principles of food hygiene. • Global Good Agricultural Practices (GLOBALGAP). GLOBALGAP is an international certification program that promotes good agricultural practices. It has developed standards that are relevant to the agriculture sector, including GLOBALGAP for fresh produce. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 16 4. Metadata Analysis The primary goal of the metadata analysis in Section 4 is to comprehensively understand how metadata is currently being utilized within the AgrifoodTEF ecosystem. This involves examining the types of metadata that are most prevalent, the tools and platforms used for managing this metadata, and the standards applied to ensure consistency and quality across various services. By analyzing the responses from the Service Data Management Plan (sDMP) surveys, we aim to identify common practices, pinpoint areas for improvement, and understand the overall landscape of metadata usage within the project. This analysis is indispensabel in guiding the development of more efficient and standardized metadata practices, ultimately enhancing the effectiveness and interoperability of services within AgrifoodTEF." To conduct the metadata analysis, a graph database and a Large Language Model (LLM) were employed to process six key questions from the Service Data Management Plan (sDMP) responses of metadata. The responses yielded common terms related to metadata, which were then categorized under node labels such as “Metadata Type,” “Tools and Platforms,” “Metadata Standard,” and “Succession on previous European projects.” These terms, along with the frequency of their occurrence, are detailed in Annex 2. The collected common terms were organized in a graph database (Neo4j). Each sDMP was linked to its associated common terms, enabling the graph database to manage relationships and connections between data entities within the survey responses. This systematic organization is depicted in Figure 1. A total of 50 sDMPs were analyzed. To further analyze the data, the Louvain algorithm was utilized. This algorithm identifies clusters within the data, based on criteria such as metadata type, tool platforms, and metadata standards. This clustering provided deeper insights into the patterns and relationships inherent in the metadata usage across the AgrifoodTEF project. The Louvain algorithm determined the following clusters in the graph database. 1. The largest cluster identified encompasses responses where no answer was given to the question, a significant observation in the survey. The prevalence of unanswered questions points to a clear trend. One possible reason for this could be that certain questions were not applicable to the specific service being addressed. However, the primary reason appears to be a lack of information or early stage of development of the services. 2. The second largest cluster relates to time stamps as a metadata type and the use of a folder structure to organize the data and metadata. This suggests that timestamps are an important metadata type and folder structures are most often used to store metadata that comprehends timestamps. 3. Other clusters relate to using specific tools with data or metadata types. For example, when the data comprehends the image format, a folder structure is often used to organize the data. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 17 Figure 1: Graph data modelling - Blue Nodes: “sDMP”, Green Nodes: “Metadata Type”, Orange Nodes: “Tools and Platforms”, Purple Nodes: “Metadata Standard”. The LLM (GPT3) was grounded on the composed graph database and was used to create a Neo4j query to question the database. The following questions were asked: 1. Which metadata types were used most, and how many respondents use them? The five Metadata Types used most frequently are time stamp, not answered, location, weather, and system/sensor description. - Time Stamp: As a crucial chronological reference for data points or images. - Location: In image datasets, to denote the geographic origin of the data. - Weather Conditions: In agricultural or environmental studies, to provide context regarding weather during data collection. - Sensor Information: Associated with image datasets, detailing sensor characteristics and configurations. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 18 2. Which tool platforms were used most and how many respondents use them? The five Tool Platforms used most by the respondents are folder structure, not answered, FMS, PostgreSQL, and InfluxDB. - PostgreSQL and InfluxDB: Used as a database management system to store and manage metadata. Postgres, with its PostGIS plugin, for its spatial data capabilities. InfluxDB is used for time-series data storage. - Folder Structure: Often utilized as a basic method for organizing and managing data and metadata. Shared folder structures are a common approach for storing data and associated metadata. - Farm Management Software (FMS): Mentioned as a platform for storing crop management data, likely used for organizing agricultural-related metadata. 3. Which metadata standards were used most, and how many respondents use them? The 5 standards metadata that was used most are not answered, coco forma, Pascal VOC, Yes, undefined, and Json. - COCO Format: Standard dataset format for object detection tasks in computer vision. - Pascal VOC: Dataset format for object detection and classification tasks. - JSON: Lightweight data interchange format using key-value pairs. - YOLO Format: A data structure for YOLO's real-time object detection. - Labelbox: Platform for annotating data for machine learning tasks. - RDF: Standard model for describing web resources in a graph format. 4. Which tools platforms were used most in combination with the metadata type ‘time stamp’? The tool platform used most in combination with the metadata type ‘time stamp’ is folder structure. 5. Which tool platforms and metadata types were mostly combined, give the top 5? 1. ‘Folder structure’ and ‘time stamp’ (18 occurrences) 2. ‘Folder structure’ and ‘location’ (10 occurrences) 3. ‘Folder structure’ and ‘weather’ (9 occurrences) 4. ‘Folder structure’ and ‘acquisition software’ (7 occurrences) 5. No information was provided for the fifth combination 6. What is the strongest association between a metadata type and a tool platform? The folder structure is the strongest association between a metadata type and a tool platform. It was used in 18 instances, making it the most commonly combined combination. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 19 The Graph Database and the Language Model synergy provided a more holistic view of the survey dataset. The most important metadata for the services provided appears to be the timestamp, location, and weather conditions. In question 5 of the graph database, most of this metadata appears to be stored in a folder structure. The answer to question 6 also states a strong association between the metadata and folder structure. 35% of the sDMP’s mentioned ‘Folder structure’ as a tool or platform to maintain the data and metadata organization, and 23% were unanswered. Maintaining a folder structure for data storage is approachable and only needs a few skills. Folder structures are a viable choice for the storage of large data objects. However, they cannot enforce data consistency and integrity to maintain metadata and increase the risks of errors or inconsistent duplicates in the collected metadata. Database technologies (relational, time-series, document-based, graph, …) or higher lever Farm Management Systems (FMS) overcome these issues. Also, other tools mentioned (Excell, R/python) can incorporate data sanity checks, though this has to be carefully implemented. Metadata standards mentioned for AI-related metadata are COCO, Yolo, and Pascal VOC. It is important to note that ISO standards or standards from international standardization organizations were only mentioned once during the data management of the services. Many of the nodes participated in the European Demeter and IOF2020 projects. Diverse Engagement with Prior Projects: Services exhibit various approaches in leveraging the outcomes or findings of previous European projects. Some actively engage, while others consider or apply relevant findings aligning with their domain or needs. Data Sharing and Data Protection (data storage, backup, accessibility, and ethical guidelines) are often performed in collaboration between the researchers and the IT department. Terminology, like the FAIR Data principles, is only mentioned once in all the surveys. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 20 5. Ontology for metadata standards In this section, we look into the development and significance of an ontology for metadata standards within the AgrifoodTEF project. An ontology is a structured and systematic framework designed to categorize and interrelate various metadata terms and standards. This framework is keyl for creating a coherent and unified approach to handling metadata across the diverse range of services and initiatives within AgrifoodTEF. The development of this ontology is not just a technical exercise, but it represents a critical step towards enhancing data interoperability, facilitating clearer communication, and streamlining data management processes within our project. Through this ontology, we aim to establish a common language and a shared understanding of metadata practices, essential for the success and advancement of the AgrifoodTEF project. 5.1. Established ontology for the metadata: The following ontologies were described in the sDMPs and are already implemented in some of the services: - NGSI for Storing Metadata (https://www.nsgi.nl/) - Agrovoc (https://agrovoc.fao.org/browse/agrovoc/en/) - Zenodo Metadata Reference Standard (https://zenodo.org/) - EDI Standards (https://www.ibm.com/docs/en/b2bis?topic=basics-edi-standards) - Semantics and Working Groups: Refers to semantic models like drmAgro, often detailed in research papers or defined within working groups for metadata organization and standardization. - Research Domain-Specific Ontologies: Utilizes terminology specific to the research domain for metadata structuring and categorization. 5.2. Maintaining ontology Adopting a common ontology improves agricultural data's understandability, usability, and interoperability. By providing a structured and standardized representation of the agricultural domain, ontologies facilitate data sharing with the clients and among the AgrifoodTEF nodes and satellites. Therefore, the applied ontology should be easy to consult by all users. This section highlights the various methodologies and practices employed within the AgrifoodTEF project to keep the ontology aligned with the latest agricultural terms, research findings, and data management trends. The applied ontology should continuously reflect the current state of knowledge and practice in the field related to the service. Below, we summarise key practices in the sDMPs that explain used strategies for documenting ontology. This comprehensive set of practices serves as a basis for improved data understanding and uniformity contributing significantly to the effectiveness of data management within the AgrifoodTEF project. - Attribute Naming and Common Agricultural Terms: Uses attribute naming based on common agricultural terms to facilitate understanding. - Ontology Alignment with Research Papers: Aligns ontology structure according to research papers, ensuring consistency with the scientific literature. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 21 - Documentation: Provides comprehensive explanations, protocols, variable descriptions, and definitions in documentation to facilitate understanding. - Client-Provided Descriptions: Descriptions and hypotheses ensure a clear understanding of the metadata used. - Explanatory Vocabulary List and Goal/Scope: An explanatory vocabulary list and a table detailing goal and scope, aiding understanding of metadata related to Life Cycle Assessment. - GitLab Documentation: Documents datasets on a GitLab page, with descriptions for improved comprehension. - Unique Identifiers and Standardized Descriptions: Uses unique identifiers assigned by Zenodo and standardized data descriptions to ensure clarity and understanding. - Semantic Model Definitions: Definitions within semantic models are provided, contributing to a clearer understanding of the ontology. Our Service Data Management Plan (sDMP) survey has highlighted a need for a comprehensive ontology within the AgrifoodTEF project and the knowledge exchange of specialized data management tools. This marks a crucial step towards implementing effective ontologies, aimed at enhancing data coherence and interoperability across our services. However, it's important to acknowledge that we are at the early stages of this initiative. Many of our services are still in the process of being fully developed, which presents both challenges and opportunities in integrating and evolving our ontological frameworks. As AgrifoodTEF continues to grow and mature, our commitment to refining and adapting these ontologies will be indispensabel in ensuring they meet the evolving needs of our project and stakeholders. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 22 6. Data sharing and protection Establishing robust data storage and sharing protocols and implementing responsible data management practices to safeguard valuable data is crucial. Analysis of sDMPs reveals many practices that contribute significantly to effective data protection. These practices, described below, provide a comprehensive framework for securing data and ensuring its integrity throughout the research lifecycle. 6.1. Data storage and data backup: - Server, desktop, and backup: Indicates storage on servers and local storage on desktops, involving regular backups and copies for data safety. - Open Data Repositories: On platforms like GitHub, Gitlab and Zenodo, code and data are stored and versioncontrolled on password-protected servers - MS Teams/SharePoint with Daily Backups: Relies on Microsoft Teams and SharePoint platforms for data storage and conducts daily backups for data security. 6.2. Methods used for data sharing: - Sharing through Website Interface and local Servers: including requesting data access via the specific website interfaces. And servers with authentication, internally hosted. - Data Shared on Network Drives and Repositories: Results, point spectra, and object-level spectra are shared on network drives, SharePoint, Gitlab, GitHub and Zenodo. 6.3. Responsible for data management: A dedicated team of professionals oversees data management. The IT department plays a central role in data infrastructure and security, while researchers are responsible for data creation and processing. The Data Compliance team/officer ensures compliance with data privacy and protection regulations, while the AI Testing department maintains the integrity of data used in AI models. This collaborative approach ensures comprehensive data management and compliance. In conclusion, the fundamental importance of robust data sharing and protection within the AgrifoodTEF project cannot be overstated. The insights gained from our Service Data Management Plans (sDMPs) represent only an initial survey and a preliminary attempt to collect comprehensive information on these practices. Despite being in the early stages, our findings reveal a range of effective methods contributing significantly to data security and integrity. From thorough data storage and backup protocols to secure data sharing methods and responsible data management, these practices form a foundational framework for further deployment of data security approach in our project. This comprehensive approach not only ensures the safeguarding of valuable research data but also promotes accessibility and compliance with regulatory standards. As the AgrifoodTEF project progresses, we anticipate further refinement and enhancement of these data sharing and protection strategies, continuing to support the integrity and success of our testing and experimentation initiative in the agrifood technology sector. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 23 7. Working groups In this section the formation and function of specialized working groups within the AgrifoodTEF project, established in November 2023 is explained within Work Package 4 (WP4). The Working groups serve as a collaborative tool, fostering discussions and information exchange among partners operating at a service level. These groups are formed specifically when multiple partners offer analogous services, aiming to collectively identify and define the fundamental building blocks of these services along with the correlated standards. The working groups are using the preparatory work on common terms, ontologies and meta data standards as a steppingstone towards the common language that needs to be developed in the agrifoodTEF project. These working groups will play a crucial role in coordinating related standards relevant to the development of new services and ensure alignment with standardization needs arising from Task 4.2. The formation of these various working groups (WGs) is a necessary step in fostering collaboration and knowledge exchange among the service level partners of the AgrifoodTEF project and serve as instrumental forums for maintaining industry standards and best practices. Their core responsibilities include identifying, assessing and integrating relevant industry standards and common metadata into new services, while simultaneously studying existing standards to identify areas in need of improvement or potential gaps and promote the development of ontologies. This holistic approach ensures consistent services, interoperability and compliance with quality standards. Simultaneously, the focus of these working groups will also be on the identification of shared and distinctive testing and experimentation methods (essential components of services) which fits seamlessly with the close relationship between common terms and metadata standards. Moreover, the delineation of working groups aligns with different facets within the AI and robotics ecosystem in the agri-food sector, ensuring comprehensive coverage that is absolutely necessary for the AgrifoodTEF project. The formation and focus of each working group were carefully determined through collaborative discussions with partners involved in WP4. To guide this process, these partners developed a strategic matrix that intersected technical competencies, such as AI and standards, with domain expertise, including areas like precision livestock farming. This matrix served as a tool to align the needs of the project with the available expertise, ensuring that each working group was both relevant and well supported. Subsequently, all partners in the AgrifoodTEF project were invited to assign their experts to various domains within this matrix. The initial list of working groups was then created based on the domains that attracted the most expertise. This approach ensured that the working groups were not only aligned with the project's technical and domain-specific needs but also adequately staffed with professionals who could drive significant progress in each area. As a result, the working groups represent a well-balanced blend of technical and domain-specific insights, fitted to address the complex challenges and opportunities within the AgrifoodTEF project. The working groups have only just started their monthly meeting but the results that will be generated by these working groups will potentially offer several service-level benefits: • Improved interoperability: The adoption of uniform terms and standards promotes seamless data exchange between different systems and applications, potentially improving efficiency and productivity. • Improved data quality: Standardized terminologies and standards contribute to the accuracy and consistency of data, thereby increasing the quality of analyses and decision-making processes. • Simplifying collaboration: Common terms and standards facilitate project collaboration and data sharing among stakeholders, which can promote innovation and effective problem solving. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 24 WG Name Topics Lead 1 Robotics Robotics (incl. UAV & implements) (FR) & Marijke H. (BE) 2 Sensor testing Sensor testing (vision, hyperspectral, audio, other) Peter R.-N. (AT) 3 Safety tests Safety tests (obstacle detection, follow-me function, ARPA tests, geofencing) Julien P. (FR) 4 AI pipeline and model testing AI pipeline and model testing (performance, robustness, explain ability) Raphael F. (IT) 5 ELSA, LCA, BM and Usability ELSA (ethical, legal and social aspects), conformity to policy and LCA, Business modeling and Usability & UX testing (UAT) Mireille v. H. (NL) 6 Data data infrastructure and handling (provision of data, interoperability), data analytics and visualization, meta-data standards, synthetic data Raul P. (PL) 7 Simulation and testing in a digital environment Simulation and testing in a digital environment Fredrik B. (SE) 8 Livestock Livestock Magdalena W. (AT) Table 1: Initial Set of Working Groups for agrifoodTEF project The establishment and strategic formation of these working groups within the AgrifoodTEF project represent an indispensable next step in transforming the outcomes of our Service Data Management Plans (sDMPs) into practical measures. This is for achieving a common language in the AgrifoodTEF project. 101100622/AgrifoodTEF AGRIFOOD TEF D5.12 Training catalogue development Y1 December 23th, 2023 25 8. Discussion, Conclusion and Next Steps Establishing a common language within the AgrifoodTEF project, including standardized formats for data labels and a standardized ontology of terms, is an extremely challenging task, particularly when the project is in its infancy and nearly all services are under development. This task not only requires strong coordination and mutual understanding between the partners but also a very dynamic approach to adapt to evolving needs and insights during the early stages of the AgrifoodTEF project. This task set out to make an initial list of common terms and metadata standards and this is only the starting point of a long endeavor towards this AgrifoodTEF common. The introduction of the Service Data Management Plans (sDMPs) played a crucial role in the AgrifoodTEF project, marking the beginning of a collective effort towards a unified language and framework. By engaging all partners in completing this extensive survey, we initiated a process of active reflection and analysis across the consortium. This exercise compelled each partner to thoroughly examine their current practices and to think more critically about the design and structure of their services. It was a start point that challenged our partners to conceptualize and articulate their data management approaches more coherently. This introspective analysis layed the basis for developing a common language and set of standards within the AgrifoodTEF ecosystem, fostering a shared understanding and vision, essential for the collaborative success of AgrifoodTEF. Via the sDMPs we got a view on common terms used within the project and the results show the necessity of a unified terminology framework across various services. The explanation of various aspects of data collection practices, including data types, collection locations, frequency and context, along with the detailed mapping of data quality assessment methods and standardised data quality labels, is a first step for all partners towards a robust data management for their services. As we move forward, these findings set the stage for more efforts in harmonizing terminology within the AgrifoodTEF project. Implementing these common terms is now a first and necessary step. The inclusion of client-specific standards and alignment with internationally recognized frameworks will be prioritized to ensure comprehensive and ethical data processing practices. The observation in the survey, where numerous questions remained unanswered, warrants attention. The lack of responses could be attributed to various factors, but the most prevalent reason appears to be the limited knowledge of the responders, or the early developmental stage of the services. The readiness level of the services plays a significant role in the ability to provide comprehensive metadata information. This finding highlights an area within the AgrifoodTEF project that may require further development and understanding to enhance the overall effectiveness of our metadata strategies as we strive towards establishing a common language within the AgrifoodTEF ecosystem. The metadata analysis reveals significant insights into the current state of data management within the AgrifoodTEF services. A key observation is the prevalent use of folder structures for organizing metadata, indicating a preference for traditional, hierarchical data storage methods. This widespread use of folder structures, particularly in conjunction with metadata types like timestamps, suggests a foundational approach to data organization that prioritizes ease of access and user-friendliness. However, it may also indicate a lack of knowledge of data management systems that are much more capable of maintaining consistency and integrity of the data. The analysis highlights the importance of timestamps as a metadata type underscores their crucial role in providing chronological context to data, an essential aspect in the agricultural sector. Similarly, the importance of geographical context is indicated by the frequent citation of location metadata highlights.