D8.5 Data Management Plan
Abstract
The LiveSeeding Data Management Plan (DMP) has been developed to illustrate how the project will comply with the grant rules on FAIR Data and Open Science, and to provide project partners with guidance on how to operationalise it. This is the 1 st version of the LiveSeeding DMP, which will be updated and expanded by month 36 and month 46 as the project is rolled out and data management needs expand.
Full text
LiveSeeding Project Deliverable 8.5 Funded by the European Union (grant no. 101059872), the Swiss State Secretariat for Education, Research and Innovation (Contract no. 22.0412) and UK Research and Innovation (UKRI). Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or REA, nor SERI or UKRI. Data Management Plan Authors: Mariano Iossa (FiBL Europe) Deliverable Number D8.5 Work Package WP8 Deliverable type Report Dissemination level PU - Public Deliverable Lead partner FiBL Europe Due date 31 March 2023 Submission date 31 March 2023 Version V1.3 Reviewers Monika Messmer (FiBL CH), Ana Marija Špicnagel Ćurko (IPS) Contact [email protected]
D8.5 Data Management Plan 1 History of changes Version Date Author Comments V1.0 6 March 2023 Mariano Iossa Draft version V1.1 21 March 2023 Ana Marija Špicnagel Ćurko Revision of v1.0 version V1.2 30 March 2023 Monika Messmer Revision of v1.1 version V1.3 31 March 2023 Mariano Iossa Final version of 1st DMP LiveSeeding - Organic seed and plant breeding to accelerate sustainable and diverse food systems in Europe is a 4-year Innovation Action funded by the European Union, the Swiss State Secretariat for Education, Research and Innovation (SERI) and UK Research and Innovation (UKRI). The project started in October 2022 and brings together 37 organisations operating in 16 European countries. LiveSeeding provides science-based evidence and best practice solutions to help achieve 100 % organic seed. LiveSeeding contributes to the transition towards environmentally-friendly, climateneutral, healthy and fair food systems through a PUSH-PULL-ENABLE strategy to enhance the availability and adequacy of organic seeds of cultivars appropriate to organic farming (PUSH), increase and stabilise the market demand for organic seeds of cultivars appropriate to organic farming (PULL), foster an enabling policy and regulatory environment where both demand and supply can harmoniously and productively negotiate without irrelevant constraints due to legal restrictions and/or regulatory fragmentation (ENABLE). LiveSeeding addresses the topics in a holistic multi-actor, multi-stakeholder, participatory approach involving stakeholders along the value chain in 17 local Living Labs (LLs) and 3 established networks of organic breeders (ECO-PB), seed savers (ECLLD) and Milan Urban Food Policy Pact (MUFPP). 15 European countries cover the different pedoclimatic zones and socio-economic contexts, including countries with a low level of development in organic seed and breeding in East and South Europe.
2 Table of Contents SUMMARY .............................................................................................................. 5 1. DATA SUMMARY ............................................................................................. 7 DATA TYPE AND FORMATS ................................................................................................ 7 DATA PURPOSE, SIZE AND SOURCES ............................................................................... 13 POTENTIAL USE OF DATA OUTSIDE OF THE PROJECT ...................................................... 15 2. FAIR DATA..................................................................................................... 16 MAKING DATA FINDABLE ................................................................................................ 16 MAKING DATA ACCESSIBLE ............................................................................................. 18 2.2.1 Repositories ........................................................................................................ 18 2.2.2 Data accessibility ............................................................................................... 20 2.2.3 Metadata: ............................................................................................................ 21 MAKING DATA INTEROPERABLE ...................................................................................... 24 INCREASE DATA RE-USE .................................................................................................. 24 2.4.1 Publication and Creative Commons licenses ............................................... 24 2.4.2 Data Quality assurance processes ................................................................. 25 2. OTHER RESEARCH OUTPUTS ......................................................................... 26 3. DATA MANAGEMENT AND ALLOCATION OF RESOURCES............................ 26 4. DATA SECURITY ............................................................................................ 27 5. ETHICS ........................................................................................................... 28 6. OTHER ISSUES ............................................................................................... 28 ANNEX 1 OPEN SCIENCE POLICY UNDER HORIZON EUROPE .............................. 29 ANNEX 2 CREATIVE COMMON LICENSES ............................................................ 30 ANNEX 3 (META)DATA FILE NAMING GUIDELINES (NOMENCLATURE) .............. 31 ANNEX 4 GENERAL TEMPLATE FOR METADATA OF DATA AND REPORTS OBTAINED IN LIVESEEDING ................................................................................. 32 ANNEX 5 TEMPLATE FOR METADATA OF DATA OBTAINED FROM FIELD TRIALS ............................................................................................................................. 34 ANNEX 6 EXAMPLES OF PARAMETERS FOR FIELD TRIALS’ TEMPLATES ............. 36
3 ANNEX 7 METADATA PARAMETERS FOR SURVEYS - EXAMPLES FOR TEMPLATES ............................................................................................................................. 38 ANNEX 8 EXPECTED DATA GENERATED FROM LIVESEEDING WPS ....................... 2 ANNEX 9 PROTOCOL/GUIDELINES TO UPLOAD PROJECT PUBLICATIONS AND DOCUMENTS IN REPOSITORIES ............................................................................. 2 ANNEX 10 ONTOLOGY FOR INTEROPERABILITY ................................................... 5
4 List of abbreviations APC Article Processing Charges API Application Programming Interface BrAPI Breeding Application Programming Interface CA Consortium Agreement CC Creative Commons CPVO Community Plant Variety Office DMP Data Management Plan DoA Description of the Action ECO-PB European Consortium of Organic Plan Breeding ERD European Router Database EU European Union ExCom Executive Committee FAIR Findable, Accessible, Interoperable, and Reusable H2020 Horizon 2020 IP Intellectual Property IPRB Intellectual Property Right Board OA Open Access PC Project Coordinator REST Representational State Transfer SCT Scientific Coordination Team YY.MM Year in 2 digits, Month in 2 digits WP Work Package
5 Summary The LiveSeeding Data Management Plan (DMP) has been developed to illustrate how the project will comply with the grant rules on FAIR Data and Open Science, and to provide project partners with guidance on how to operationalise it. This is the 1st version of the LiveSeeding DMP, which will be updated and expanded by month 36 and month 46 as the project is rolled out and data management needs expand. The objective of this deliverable on Data Management Plan (DMP) is: To ensure good data management throughout the lifecycle of the project but also future sharing and reusing of research data beyond the project and consortium; To present the procedures and repositories used for the project in line with FAIR principles and new requirements of Horizon Europe programme; To provide clear guidance and practical tools to project partners for the management, storage and sharing of data produced in the scope of the project. To that aim, this deliverable includes 6 chapters and 10 annexes that cover: A presentation of the type of (meta)data, size, purpose of its collection, methodology and standards / controls used within LiveSeeding (chapter 1); An explanation on how the project will comply with FAIR (Findable, Accessible, Interoperable, and Reusable) principles (chapter 2); An introduction to procedures and tools used in the project to ensure data security, ethics, privacy, and resources deployed (chapters 3-6) and reference to other project deliverables focused on those issues; A set of easy-to-use (meta)data management tools (templates, info sheets, etc) for partners (Annexes 1-10). The structure of the deliverable is based on the new template developed by the European Commission to guide grantees. It reflects changes in the rules of Horizon Europe in respect to the previous funding programme (Horizon 2020). Some important changes are highlighted in the info box below: Data management - Main changes of rules in Horizon Europe - Scientific publications: No embargo period, immediate OA - Re-use: Author(s) need to retain rights (no copyright transfer) as for a CCby license to assert ownership; acknowledgement for re-use required (see Annex 2) - Costs of APC: recoverable only if journal is a full open access one (not hybrid) - Metadata: must be FAIR and under CC0 license (see Annex 2); deposit as soon as possible after production/generation or after processing control
6 - Metadata: partners should not forget to add project acronym - Data can be confidential if necessary but grantees need to justify why (default approach being FAIR). For example, commercially valuable data, risk to endanger trade secret (soft IP) or make IP more difficult All data will be treated as open source unless it is defined beforehand as confidential data to follow the principle “as open as possible, as closed as necessary” , i.e. to allow for circulation of scientific knowledge while protecting commercial interests and confidentiality. Specific restrictions to the openness / dissemination of background data and results as indicated by partners in the Consortium Agreement will be respected (the exact list of restriction is not annexed due to confidential nature of this information and the fact that this deliverable is public). Each LiveSeeding partner is responsible for careful data recording, validation of data quality, and clear documentation of the data they have generated , in compliance with data protection legislation as for guidelines provided under deliverable D8.4 on Protection of personal data & recruitment of stakeholders. Task and WP leaders are responsible to provide data templates and monitor that data protocols are followed. The DMP is a living document that will be updated over the course of the project by M36 (version 2) and M46 (version 3), as per Grant Agreement. At this stage, some protocols are still under development and more detailed description of datasets and templates need to be defined by WP and task leaders. The Project Coordinator, Scientific Coordination Team and Executive Committee stir this process.
7 1. Data Summary The DMP ensures good data management throughout the lifecycle of the project guaranteeing all information and data collected and produced respects the FAIR principles, i.e. all research data should be Findable, Accessible, Interoperable and Reusable (FAIR). The DMP describes the data (incl. metadata) and software standards, availability, curation and preservation methods and covers the data documents and solutions used and created in the project. This chapter specifically addresses the questions proposed in the deliverable template proposed by the EC on the data type and format (chapter 1.1); the purpose of its collection, the size and sources, i.e. the origin of the data generated by the project or re-used from other sources/projects (chapter 1.2); the potential use of data outside the project, both during and after the project life span (chapter 1.3). Data type and formats LiveSeeding will generate a diverse set of statistical, agronomic, socio-economic and lab data as well as software codes in order to achieve the overall goal to strengthen the organic seed and plant breeding sector for sustainable and diverse food systems in Europe. Based on the expected results of the individual WPs presented in Table 1, each WP leader has identify the expected data and research outputs including their status in terms of type and/or origin (e.g. reused, collected, generated) listed in Annex 8. These list will be amended and further elaborated during the project. LIVESEEDING also uses existing data and software from previous EU funded projects and other sources (see chapter 1.2) that will be integrated with new data and findings as mentioned in the background data of the consortium agreement as well as existing software that will be adjusted for LiveSeeding purposes. Table 2 lists the ICT tools that will be adapted or newly developed for the use of a broader audience. Personal data collected in the scope of LiveSeeding will be anonymised in compliance with the European Data Protection Regulation (GDPR, Regulation (EU) 2016/679) as well as applicable national legislation on data protection as described in detail in deliverable D8.4. Table 1: Project results organised by Work Package WP Project Results WP1 Sound breeding strategies and novel tools suitable for organic farming New organic varieties (OV), organic heterogeneous material (OHM) and improved landraces for numerous crops WP2 Pilots of viable multi-actor organizational models for decentralised, participatory cultivar evaluation
8 New trial designs and analysis for on-farm cultivar testing WP3 Scientific insight on interaction of seed microbiome and seed vigour Holistic seed health strategy WP4 Stepwise approach in the form of national roadmap (tailored for 10 EU countries) to phase out derogations and increase availability of organic seed Data on supply and demand of organic seed WP5 Business models and plans for organic breeding, cultivar testing and seed production initiatives WP6 Quantification of added value of organic breeding Cities food policies including the aspect of farmed diversity from local production Marketing strategies for food products derived from OV and OHM WP1-6 LiveSeeding will have a strong focus on training related to the innovations tested in each WP and specifically targeted to the user target group of each WP outcomes. Trainings modules including different type of material (hand-outs, videos, practical exercises, etc.) will cover all major topics of the project A short description of the 4 software that will be developed/adapted under LiveSeeding project to collect and process data is provided in Table 2: Table 2: Softwares developed/adapted under LiveSeeding project SHiNeMaS to support participatory breeding will be adapted (WP1): Short description : SHiNeMaS is a digital tool developed by INRAE in former projects and dedicated to seed lots history, phenotyping and cultural practices. SHiNeMaS aims: (1) to study on-farm management of crop diversity, (2) to develop populationvarieties adapted to organic and low-inputs agriculture, (3) to co-develop tools and methods adapted to on-farm experiments. Data type : in SHiNeMaS, project partners will collect: - Seed lot description and their genealogy (Reproduction, crossing, mixtures, selection, diffusion (main data collected); - On farm evaluation data of germplasm (plant height, awn, ear weight, thousand kernel weight etc.); - Images illustrating field trials ; - Stock and storage information of seeds. Adaptations : the main developments about SHiNeMaS under/for LiveSeeding will be the implementation of web services to ensure interoperability with other systems
15 For example, data collected under T4.1 on the changes in supply and demand of organic and farm-saved seeds will be compared with latest data collected under the H2020 funded project “LIVESEED”. Potential use of data outside of the project Project partners will be proactive and take specific measures to ensure that their results are exploited outside of the project by a broad range of stakeholders. To that aim a deliverable including dissemination and exploitation plan (D7.1 Plan for the Exploitation, Dissemination and Communication of Results - PEDR) has been prepared by M06 in order to define plans and procedures to maximise dissemination of project activities and exploitation of results. PEDR describes main project’s dissemination, communication, and exploitation tools, channels and methods that will be used in order to maximize the impact of LiveSeeding across the project´s entire range of targeted audiences. During the project various materials will be produced, such as media releases, videos, newsletters, articles, leaflets, posters, roll-ups, presentations, practice abstracts, policy briefs, training materials, booklets and (scientific) publications. These materials will be distributed among the stakeholders and stored on different repositories depending on the type of data and materials. The project publications and output materials will be deposited on 3 main online open access repositories: Zenodo for peer-reviewed scientific publication and data, Organic Eprints (for all kind of publications, data, reports, videos, presentations) and Organic Farm Knowledge platform, section “Seed” (for practice abstracts). LiveSeeding participants are encouraged to share their results via Open Research Europe (ORE). If no sensitive data involved, LiveSeeding data will be made available under the terms of a Creative Commons Public Domain Dedication (CC0 license). The project communication team will work to create awareness of the newly uploaded LiveSeeding data sets and their availability for further use, e.g. by social media outlets and news on webpage. Also, LiveSeeding project website will be a depository of all project results that have an Open Access status. The website will be timely updated by IPS Konzulting and it will serve as a channel for reaching out to stakeholders, citizens and the wider public by providing information on different events, publications, and project deliverables, as well as an inventory of project activities and fast access to results summarized in the form of booklets, videos, newsletters, and policy briefs and recommendations. For long term storage the LiveSeeding website content will be uploaded on the repository Organic Eprints and moved to the website of ECO-PB before the end of the project.
16 2. FAIR data FAIR management of data means that data of the project will comply with the FAIR data principles. The data will be findable and accessible through the use of open access repositories. To make data interoperable, defined standards, formats and vocabularies for data and metadata will be used. The re-usability of the data will be available through Creative Commons Public Domain Dedication (CC0 license). Chapters 2.1 to 2.4 explain what the project will do to make data FAIR but also guide project partners on how to go about it at practical level. FAIRification Process The FAIR Data Principles apply to metadata, data, and supporting infrastructure (e.g., search engines). Most of the requirements for findability and accessibility can be achieved at the metadata level. Interoperability and reuse require more efforts at the data level. For more guidance on the “fairification” process can also be found on the GoFair website. Making data findable LiveSeeding peer reviewed publications and supplementing metadata and data will be assigned, a unique Digital Object Identifier (DOI) which allows to find and track data when deposited in Zenodo and linked to the LiveSeeding community. All LiveSeeding publications and data as well as other materials, such as booklets, Practice Abstracts, PowerPoint presentations, will be uploaded on the repository in Organic eprints and will be findable with keywords (including “LiveSeeding” as a keyword for easy search). Instruction for uploading in Zenodo and organic eprints are given in Annex 9. LiveSeeding research/experiment data files will include "Metadata" in the excel file containing also the “Data”. All Metadata files will include the Project Acronym. More specifically: For agronomic data, following metadata should be accompanied to allow correct interpretation: General Information of experiment Organisation responsible for experiment Objectives
17 Clearly defined treatments or factors to be tested Statistical design, number of replications and randomisation Field plan Information of location (postal code, GPS, altitude, soil type, organic or conventional management, crop rotation, soil management) Crop species (Latin names), Breeding lines and cultivar names, breeder and origin of seed Year of trial Trial management (weed management, fertilization, crop protection, sowing and harvesting time, sowing density, row spacing, temperature and rainfall, irrigation, etc.) Time and developmental stage of data assessment and sampling Methodology used for sampling and data collection Indication of tools used for assessment Clear definition of each qualitative and quantitative parameter assessed Methods used for data validation and statistical analysis Any specific comments or assumptions made. For microbiome or plant genetic data, different formats will be used based on the respective data. For these data following metadata must be accompanied to allow correct interpretation: Exact procedure of sampling for DNA extraction DNA extraction methods, purification and storage conditions Detailed methodology of DNA analysis Laboratory responsible for analysis Time of analysis and respective lab protocols, positive and negative standards used Number of field and lab replications used Method used for data validation and statistical analysis Any specific comments or assumptions made. For socio-economic data in quantitative and qualitative surveys or interviews, the following metadata must be accompanied to allow correct interpretation: Questionnaire / interview guide Recruitment method for participants Demographic data of participants (e.g. age, gender, affinity to organic production, education, residency, country of origin) Country of survey Information provided before survey Tools used for data recording (e.g. online, phone interview, recording device, written, etc.)
18 Clear definition of each qualitative and quantitative parameter assessed Methods used for data validation and statistical or other analysis Any specific comments or assumptions made. Templates for metadata are provided in Annex 4 and 5. Task leaders together with WP leaders are in charge of developing detailed templates for such experiment descriptions and data entering. The involved partners are responsible to provide these metadata for each dataset. The task leaders will monitor the completeness. Metadata will be used to index searchable keywords to improve findability. Based on the specified keywords in the metadata (e.g. organic breeding, microbiome), the LiveSeeding data can be searched in a similar modus like publications and reports under Organic Eprints and also under the Zenodo’s search engine after publication. For both Zenodo and Organic Eprints “LiveSeeding” community will be implemented to allow to extract all LiveSeeding document from the repositories. In Zenodo repository the metadata is exported in several standard formats such as MARCXML, Dublin Core, and DataCite Metadata Schema (according to the OpenAIRE Guidelines). It is worth noting that, several journals already require the deposit of the full research data and script used for statistical analysis as supplement information for publication. Tip for metadata and data stored in separate files In case the metadata and data are stored in separate files (for example in case you need to use a format other than excel that allows several sheets in one file) it is crucial that the link between these files is obvious. This can be achieved by submitting several supplementing files together with the publication: supplement I contains the raw data and supplement II the metadata. Making use of the unique identifier DOI obtained when depositing data in Zenodo, the metadata can be clearly linked with the main data set. Making data accessible 2.2.1 Repositories Different types of data will be transferred to different repositories (FAIR): Statistical data on crop specific level (e.g., area of organic production) will be integrated in existing database of FiBL-CH on Organic Agriculture. Data are publicly available under “Organic World” and summarized in annual reports
19 “The World of Organic Agriculture” published by FiBL-CH and IFOAM Organics International and reported at biofach Nürnberg the world largest organic faire; Practice Abstract will be transferred to EiP-AGRI website, the Organic Farm Knowledge Platform 3 , Organic Eprints 4 and LiveSeeding website (www.liveseeding.eu); All Scientific publications (both peer-reviewed and non-reviewed) will be uploaded on Organic Eprints to reach the organic community; Organic Eprints has been established to promote open access to research results, is running on Linux and other free, open source software and has had almost 7.5 million downloads since May 2018. More statistics about Organic Eprints traffic are available here; Peer-reviewed publications and open research data will also be deposited in Zenodo 5 ; Tools, training and demonstration material will be made available on the Organic Farm knowledge Platform, Section “Seed” the LiveSeeding website as well as on ECO-PB website 6 ; Genomic data, where appropriate, shall be integrated into existing general platforms such as the Plant Genomics and Phenomics Research Data Repository (GPD), the Genetic and Genomic Information System (GnpIS), or 3 The Organic Farm Knowledge platform (www.organic-farmknowledge.org) provides access to a wide range of tools and resources about organic farming that can help improve production. It also aims to serve as a virtual meeting place for cross-border learning. The platform promotes the exchange of knowledge among farmers, farm advisers, and scientists, with the aim of increasing the productivity, quality and sustainability of organic farming across Europe. The core of the Organic Farm Knowledge platform is the toolbox ranging from practice abstracts, leaflets, videos, website, etc. The toolbox offers the possibility for an arbitrary text search as well as displaying tools by theme, language, organisation, project, keyword, type, country of origin and year. 4 Organic Eprints (is an international open access, multilingual online repository of electronic documents related to research in organic food and farming. The archive contains full-text papers together with bibliographic information, abstracts and other metadata. It also offers information on organisations, projects and facilities in the context of organic farming research.To upload papers to Organic Eprints a free of charge user registration is required. A User guide for organic e-print is available in different languages and in Annex 9. 5 Zenodo (www.zenodo.org) is a free of charge, general-purpose open repository developed under the European OpenAIRE program and operated by CERN. Zenodo makes the research results citable via Digital Object Identifier (DOI) and integrates them into existing reporting lines to funding agencies. A LiveSeeding community has been created and guidelines for uploading are described in Annex 9. 6 ECO-PB - European Consortium of Organic Plan Breeding (www.eco-pb.org) is a non-profit network of organisations fostering organic plant breeding since 2001. It brings together the majority of organic plant breeding initiatives across Europe. The open access webpage is managed by LiveSeeding partner FiBL-DE.
20 COpenPlantOmics (COPO), as well as crop-specific databases such as Gramene for cereals or Cool Season Food Legume Genome Database for legumes; Field trials will be stored as a combined file across years and locations using the same parameters. Data will be prepared in such a way that they can be used in meta-analysis within LiveSeeding project and beyond (see chapter 2.5.1 for long term curation). Non-confidential data (see chapter 2.2.2 for Background and Results in order to ensure that data is “as closed as necessary”) will be deposited together with the publication as supplementing data on Zenodo and Organic Eprints. Agronomic data, when appropriate for long term storage, shall be integrated in crop specific repositories like the FAO’s AgroPortal following common reference ontologies (see chapter 2.3 for more information on ontologies and interoperability) Githup, USDA or others; Socioeconomic data including those gathered through surveys after anonymization (considering GDPR compliance) can be archived in domainspecific repositories like Council of European Social Science Data Archives European Research Infrastructure (CESSDA), Eurobarometer, GESIS Data Catalogue, or GESIS Research Data Centers - International Survey Programs (ISSP),). It is worth noting that, several journals already require the deposit of the full research data and script used for statistical analysis as supplement information for publication. 2.2.2 Data accessibility All raw data collected by partners and uploaded on the LiveSeeding SharePoint will be treated as confidential. Preliminary data will only be accessible to the project partners. Access to the LiveSeeding SharePoint is regulated by a personal login and password system administered by FiBL Europe and accessed/used by project partners as for the guidelines of project deliverable D8.1 Internal platform for communication and exchange manual. Greatest care will be given to protect personal data obtained during surveys or interviews and confidential information in line with the project deliverable D8.4 Protection of personal data & recruitment of stakeholders as well as D8.3 Ethics Issues and Non-European compliance. All such deliverables are available in LiveSeeding SharePoint. Without clear consent of each participant to be named, all personal data will by anonymized, before any data will be uploaded on the LiveSeeding SharePoint and made publicly available. As for the terms of the LiveSeeding Consortium Agreement signed by each partner, all information declared as confidential background information will be kept confidential. Access rights to results and Background needed for the performance of the own work of a Party under the LiveSeeding project shall be granted on a royalty-
21 free basis, unless otherwise agreed for Background in the Consortium and Grant Agreement. Also foreground information which is disclosed by one Party (the “Disclosing Party”) to any other Party (the “Recipient”) in connection with LiveSeeding implementation and which has been explicitly marked as “confidential” at the time of disclosure (or when disclosed orally has been identified as confidential at the time of disclosure and has been confirmed and designated in writing within 15 calendar days from oral disclosure as confidential information by the Disclosing Party) is treated as “Confidential Information”. Finalized data will be made available to LiveSeeding partners and linked parties that need these data as input for their tasks. However, certain data and information might be excluded from the open data strategy: if there are commercial interests (e.g., one or several partners claim to protect their IP rights for exploitation of results); if data are not yet published in peer-reviewed papers. For example, in case data is collected also as part of a PhD research project, making it public before the publication as a doctorate research publication would make it unusable for the Doctorate project as no longer considered novel data. The ExCom, on the basis of advice from the IPRB -where relevantwill monitor results and exploitation plan and make sure that commercial interests of involved partners are guaranteed. Data that might be considered confidential will not be disclosed to the broader public before approval by ExCom. The ExCom will act as Data Access Committee (e.g. to evaluate/approve access requests to personal/sensitive data outside of the indications already provided in the present deliverable). If there is no request for confidentiality, compiled data and results will be made available to all LiveSeeding partners and beyond. Thus, data or research results that are not protected by IP rights and are no longer confidential (i.e. after publication) will be published on the LiveSeeding website and added to repositories or databases as for indications under chapter 2.5.1. 2.2.3 Metadata: Metadata can be defined as information that is needed to be able to use and interpret statistics. Metadata describe data by giving definitions of populations, objects, variables, the methodology and quality 7 . 7 https://ec.europa.eu/eurostat/statistics-explained/index.php?title=Glossary:Metadata
22 Another, simpler, way to define Metadata is "data that provides information about other data", 8 but not the content of the data, such as the text of a message or the image itself. Metadata allows users to access resources by "allowing resources to be found by relevant criteria, identifying resources, bringing similar resources together, distinguishing dissimilar resources, and giving location information" 9 While nowadays metadata typically takes a digital form, to exemplify this, we can refer to traditional catalogue cards used in paper archives. Such cards include the name of the author, title, subject, date of publication, kind of literature, edition etc) to help to find relevant information and sources. 10 Figure 1: Catalogue cards, as an example of metadata source (Wikipedia) For a visual and easy to understand explanation of what metadata is with several examples of metadata for different data formats (a photo, a Word document, a spreadsheet, an email, a website, etc) it is also possible to consult Dataedo website. Metadata will be made openly available and licenced under a public domain dedication CC0 (see Annex 2 on Creative commons), as per the Grant Agreement. As a minimum requirement under Horizon Europe grants, Metadata of deposited publications will provide information at least about the following: Publication (author(s), title, date of publication, publication venue); Horizon Europe funding, grant project name, acronym and number; Licensing terms; 8 https://en.wikipedia.org/wiki/Metadata 9 http://www.niso.org/publications/press/UnderstandingMetadata.pdf 10 https://en.wikipedia.org/wiki/Library_catalog
23 Persistent identifiers for the publication, the authors involved in the action and, if possible, for their organisations and the grant; Where applicable, the metadata must include persistent identifiers for any research output or any other tools and instruments needed to validate the conclusions of the publication. Rich metadata allow a computer to automatically accomplish routine and tedious sorting and prioritising tasks that currently demand a lot of attention from researchers. The rationale behind this principle is that someone should be able to find data based on the information provided by their metadata, even without the data’s identifier. 11 Figure 2: Example of the difference between basic metadata and rich metadata 12 Since rich metadata are essential to ensure that data and other research outputs are reusable, LiveSeeding partners are required to provide more extensive metadata as for guidance/templates under Annex 4 and 5. Metadata will remain available and findable, even when the data is no longer available by depositing the data with the attached metadata in repositories, such as Zenodo and Organic Eprints, as for indications provided under chapter 2.5.1. Softwares used in the framework of the LiveSeeding project to collect data have been described under chapter 1.1. Once data is collected and extracted, they will be available as .xls, .cvs and other formats for which the software itself will no longer be needed and therefore data will continue to be accessible and re-usable. 11 https://www.go-fair.org/fair-principles/f2-data-described-rich-metadata/ 12 https://recosenselabs.com/blog/rich-metadata-game-changer-for-media-organization
24 Making data interoperable In order to improve the interoperability of the data produced under the LiveSeeding project, controlled vocabulary will be used to allow the combination of data. In particular, data and metadata should be annotated using vocabularies/ontologies/thesauri that are commonly applied in the field. LiveSeeding project partners are evaluating the use of Breeding API 13 an interoperability standard for plant breeding databases 14 . BrAPI is a standardized RESTful web service API specification for communicating plant breeding data. This community driven standard is free to be used by anyone interested in plant breeding data management. CropOnthology will be used to manage the ontology of the phenotyping traits. Crop and variety will be managed using the codes and standards from CPVO - Community Plant Variety Office. For identifying the seed lots and germplasm it is planned to test the Digital Object Identifier managed by the International Treaty on Plant Genetic Resources for Food and Agriculture. The vocabularies/ontologies/thesauri that are used to describe the data must also be well documented and easily findable, accessible, and interoperable. A dedicated task will ensure that all the tools developed and used in the project respect the FAIR principles. The main relevant data will be accessible using a set of APIS allowing authorized user to search and access data. If the data set builds on already existing data or on complementary information stored in a different data set, this needs to be specified. In particular, the scientific link between the data sets needs to be described. Furthermore, all data sets need to be properly cited (i.e. including their persistent identifiers). If applicable, a Uniform Resource Locator (URL) giving reference to a web resource of such additional data (e.g., reference web pages (http), file transfer (ftp), email (mailto), database access (JDBC)) will be provided. Increase data re-use 2.4.1 Publication and Creative Commons licenses All peer-reviewed scientific publications produced under LiveSeeding will be published as open access immediately after being peer-reviewed and validated (see Annex 1). In fact, under the new Horizon Europe rules no embargo period is allowed for scientific publications, and immediate OA is needed. Authors of publications produced under the LiveSeeding project will retain enough copyrights to be able to allow others reuse the content of their work, by using Creative Commons licenses, CC-BY or CC0 (see Annex 2). 14 https://brapi.org/
31 Annex 3 (Meta)Data file naming guidelines (Nomenclature) The research DATA generated by partners during the project will be stored on the LiveSeeding SharePoint in the respective WP and Tasks folder. Partners are requested to always name files using the following criteria (nomenclature): A. The file name must be unique and informative. The filename should contain involved task using the numbering of the Description of Action, partner organisation (like in GA of LIVESEED), topic, species, year of experiment, B. No spaces should be made in filename, but underscores can be used C. Version number and status indicated by following endings: • _RAW for preliminary data • _FIN for finalized data (validated) • _PRO for protected data • _OA for open access data D. Language: language shall be indicated with 2 letters (EN, ES, FR, DE, IT, in capital letters in the file name Final versions of documents from all WPs, used for the dissemination of LiveSeeding project will be stored in the project SharePoint as follows: Scientific Publications : - under Scientific Publications Other Publications (including practice abstracts, brochures, posters, leaflets) - Under the Publications (PA, brochures, leaflets, factsheets, etc) section of the WP7 Communication, Dissemination and Exploitation folder Photos and Videos: under WP7Communication, Dissemination, and Exploitation. - • Under the Image database LiveSeeding SharePoint folders and sections contain also some guidance on what and how to store files. LiveSeeding_WPn_Tn.n_DATA_filename_Version_Status_LANGUAGE Exemple: LiveSeeding_WP3_Task3.3_DATA_FIBL-CH-pottrial-pea-2019_v2_RAW_DE
32 Annex 4 General template for metadata of data and reports obtained in LiveSeeding Please always add the project acronym (LiveSeeding) in the (meta)data and use “LiveSeeding” as a key word when uploading materials on organic eprints or other repositories Funding Programme Horizon Europe Project Acronym LiveSeeding, Project Number 101059872 File name add Filename Licensing terms CC0 Task no. add Task number related to the file Task Description Title of Task or Subtask of LIVESEED Title/Name Title or short description of dataset Involved Organisation(s) Organisation responsible, who else was involved in data collection (add short name for organisation) Author(s) add responsible persons involved in collecting the dataset Type of data indicate type of data (agronomic, socioeconomic, molecular data) of qualitative or quantitative nature Status of data indicate status and level of protection of research dataset Date add year of data collection Objective of data set add objective why the experiment/research was performed Use of existing data indicate if existing data have been used, if yes, what kind of data were used Original source(s): Download URL: If applicable: The URL of the dataset in its original source Original source(s):Licenses: If applicable: The license under which the dataset is applicable from its original source New collected data indicate if new data have been collected and if yes that kind of data were collected Size of data Estimate or give exact size of dataset in kB
33 References Publications or reports related to the dataset Analysis Indicate software and program of analysis of data Experimental design list factors tested, experimental design, number of replications, survey design for socio-economic data Documentation of experiment list environmental parameters that are important for interpretation of data (e.g. plot size, soil type, farming management, climatic data) Standardisation indicate which standard or template has been used for data collection and if this is in line with international ontology (structures vocabulary) Parameters used list of common parameters with clear descriptions Data repository indicate suited repository for storage and broader re-use of dataset Value of dataset Information about the potential value for exploiting and reusing the dataset. Including references to datasets that can be integrated for added value. Comments Special comments for the interpretation of the dataset Annex a Template for data collection Annex b Report, Presentation, Publications linked to data Annex c Demonstration material linked to data
34 Annex 5 Template for metadata of data obtained from field trials Please always add the project acronym (LiveSeeding) in the (meta)data and use “LiveSeeding” as a key word when uploading materials on organic eprints or other repositories Documentation of Field Trial File name Add Filename Task/Subtask no. Add Task or Subtask related to the file Task Description Title of Task or Subtask of LIVESEED Title/Name Title or short description of dataset involved Organisation(s) which organisation is responsible, who else was involved in data collection (add short name for organisation) Name of Site Manager of trial Farmer involved Location (physical address) Location (GPS data) Year of assessment Soil Type Soil Analysis Average Rainfall per year Average Temperature per year Year of conversion to organic Crop rotation Previous Crop Previous Intercrop CROP Species and Cultivars (species Latin name, cultivar used, seed source, germination rate)
35 Experimental field design copy of field plan and experimental design Time and way of soil preparation Time, quantity and type of fertilisation (incl. nutrient concentration) Time of sowing and type of sowing machine Plot size & Sowing density Time/frequency and quantity of irrigation Time/frequency, quantity type of plant protection (incl. concentration of active compound) Time/frequency and quantity of irrigation Time/frequency and type of weed control Time of harvest and harvesting equipment Parameters assessed: description of measurement, code of parameter, assessment according to Reference or clear description, unit, at which growth stage, when (date), by whom Average plant height at flowering (Code PH_fl_cm) average plant height on plot level estimated based on 3 measures per plot from soil till top of plant in cm in 5 cm intervals at growth stage beginning flowering assessed by person xyz on the 21st May 2018 Leaf rust at flowering (Code LR_fl_5) average severity of leaf rust estimated on plot level rated from 1 = no symptoms till 9= leaves are completely infected = dead tissue according to [give Reference] assessed 5 days after flowering by person xyz on the 26th May 2018
36 Annex 6 Examples of parameters for field trials’ templates (Source: https://github.com/craig-willis/icasa/blob/master/icasa-mgmt-info.csv ) AgMIP Category ICASA Variable_Name ICASA Code_ Displa y Description Unit_or_type Silver harvest_yield_harvest_ dw HWAH Harvested yield at harvest (dry wt) kg/ha Silver harv_yield_harv_f_wt HFAH Harvested yield at harvest (fresh wt) Mg/ha Silver harvest_moisture HMAH Moisture content of fresh yield at harvest kg[water]/kg[harv est] Silver anthesis_dap ADAP Growth stage as days after planting, anthesis dap Silver anthesis_date ADAT Growth stage date, anthesis date Gold tops_dry_weight_matu rity CWAM Tops dry weight at maturity kg/ha Gold byprod_removed_at_h arv BWAH By-product removed at harvest as dry wt kg/ha Gold harvest_maturity_dap R8AP Growth stage days>planting, harv mature dap Gold physiologic_maturity_ dap MDAP Growth stage of physiol. maturity, as days after planting dap Gold physiologic_maturity_ date MDAT Growth stage of physiol. maturity, date date Gold evapotrans_cumul_ma t ETCM Evapotranspiration, cumulative to end of season mm Gold rainfall_in_season PRCM Rainfall, season mm Gold canopy_height_max CHTX Canopy height, max value m Platinum grain_unit_dry_wt_mat ur GWGM Grain unit dry wt at maturity mg Platinum harvest_no_at_maturit y H#AM Harvest number per area at maturity number/m2 Platinum grain_protein_conc_m atur GP%M Grain protein concentration at maturity % Platinum grain_N_at_maturity GNAM Grain N at maturity kg/ha Gold date_of_measurement DATE Date of measurement date Gold ear_number E#AD Ear number number/m2
37 Gold grain_dry_weight GWAD Grain dry weight kg/ha Gold tops_dry_weight CWAD Tops dry weight kg/ha Gold leaf_area_index LAID Leaf area index on a given day m2/m2 Gold tops_(ie_canopy)_N TNAD Tops (ie canopy) N kg[N]/ha Platinum soil_layer_top_depth SLDUB Soil layer depth to upper boundary of sample (top) cm Platinum soil_layer_bot_depth SLDLB Soil layer depth to lower boundary of sample (bottom) cm Platinum soil_water_by_layer SWLD Soil water measured at a specified depth cm3/cm3
38 Annex 7 Metadata parameters for surveys - Examples for templates (Source: Guide to Social Science Data Preparation and Archiving, 6th edition, Interuniversity Consortium for Political and Social Research -ICPSR) XML ICPSR recommends using XML to create structured documentation compliant with the Data Documentation Initiative (DDI) metadata specification, an international standard for the content and exchange of documentation. XML stands for eXtensible Markup Language and was developed by the W3C, the governing body for all Web standards. Structured, XML-based metadata are ideal for documenting research data because the structure provides machine-actionability and the potential for metadata reuse. XML defines structured rules for tagging text in a way that allows the author to express semantic meaning in the markup. Thus, question text — for example, <question>Do you own your own home?</question> — can be tagged separately from the answer categories. This type of tagging embeds “intelligence” in the metadata and permits flexibility in rendering the information for display on the Web. Data Documentation Intiative (DDI) At the outset of a project, we encourage data producers to generate documentation that is tagged according to the Data Documentation Initiative (DDI) metadata specification, an emerging international standard for the content, presentation, transport, and preservation of documentation (Blank and Rasmussen, 2004). The DDI specification is written in XML, which permits the markup, or tagging, of technical documentation content for retrieval and repurposing across the data life cycle. (See “Getting Started with the DDI.”) The Data Documentation Initiative (DDI) provides a set of XML rules specifically for describing social, behavioral, and economic data. DDI is designed to encourage the use of a comprehensive set of elements to describe social science datasets, thereby providing the potential data analyst with broader knowledge about a given collection. In addition, DDI supports a life cycle orientation to data that is crucial for thorough understanding of a dataset. DDI enables the documentation of a project from its earliest stages through questionnaire development, data collection, archiving and dissemination, and beyond, with no metadata loss. DDI authoring options Several XML authoring tools are available to facilitate the creation of DDI metadata. With a generic XML editor, the user imports the DDI rules (i.e., the DDI XML Schema) into the software and is then able to enter text for specific DDI elements and attributes. The resulting document is a valid DDI instance or file. There are also DDI-specific tools, such as Nesstar Publisher and Colectica, which produce DDI-compliant XML markup automatically. For more information on DDI and a list of tools and other XML resources, please consult the DDI website. Depositing DDI metadata ICPSR encourages the deposit of DDI metadata with deposits of research data. There are currently two main versions of the DDI specification — DDI Codebook (Version
39 2.*) and DDI Lifecycle (Version 3.1). Most archives will prefer or at least readily accept documentation submitted in either of the DDI versions. To be in full compliance, a document should have question text integrated into each variable. It may not be possible for a project to produce documentation that is DDIconformant. In those situations, using a uniform, structured format with integrated question text is the best alternative, as it will enable the archive to convert the files to XML format easily. Important metadata elements Since most standard computer programs will produce frequency distributions that show counts and percents for each value of numeric variables, it may seem logical to use that information as the basis for documentation, but there are several reasons why this is not recommended. First, the output typically does not show the exact form of the question or item. Second, it does not contain other important information such as skip patterns, derivations of constructed variables, etc. Key Social Science metadata items A list of the most important items to include in social science metadata is presented below. Note that many of the high-level elements have counterparts in the Dublin Core Metadata Initiative (DCMI) element set. The DCMI is a standard aimed at making it easier to describe and to find resources using the Internet. For more information on the DCMI, please view its website. Principal investigator(s) [Dublin Core -- Creator]. Principal investigator name(s), and affiliation(s) at time of data collection. Title [Dublin Core -- Title]. Official title of the data collection. Funding sources. Names of funders, including grant numbers and related acknowledgments. Data collector/producer. Persons or organizations responsible for data collection, and the date and location of data production. Project description [Dublin Core -- Description]. A description of the project and its intellectual goals and indicates how the data articulate with related datasets. Publications providing essential information about the project should be cited. A brief project history detailing any major difficulties faced or decisions made in the course of the project is useful. Sample and sampling procedures. This section should describe the target population investigated and the methods used to sample it (assuming the entire population is not studied). The discussion of the sampling procedure should indicate whether standard errors based on simple random sampling are appropriate, or if more complex methods are required. If weights were created, they should be described. If available, a copy of the original sampling plan should be included as an appendix. A clear indication of the response rate should be provided, indicating the proportion of those sampled who actually participated in the study. For longitudinal studies, the retention rate across studies should also be noted.
40 Weighting. If weights are required, information on weight variables, how they were constructed, and how they should be used. Substantive, temporal, and geographic coverage of the data collection [Dublin Core -- Coverage]. Descriptions of topics covered, time period, and location Data source(s) [Dublin Core -- Source]. If a dataset draws on resources other than surveys, citations to the original sources or documents from which data were obtained. Unit(s) of analysis/observation. A description of who or what is being studied. Variables. For each variable, the following information should be provided: The exact question wording or the exact meaning of the datum. Sources should be cited for questions drawn from previous surveys or published work. The text of the question integrated into the variable text. If this is not possible, it is useful to have the item or questionnaire number (e.g., Question 3a), so that the archive can make the necessary linkages. Universe information, i.e., who was actually asked the question. Documentation should indicate exactly who was asked and was not asked the question. If a filter or skip pattern indicates that data on the variable were not obtained for all respondents, that information should appear together with other documentation for that variable. Exact meaning of codes. The documentation should show the interpretation of the codes assigned to each variable. For some variables such as occupation or industry, this information might appear in an appendix. Missing data codes. Codes assigned to represent data that are missing. Such codes typically fall outside of the range of valid values. Different types of missing data should have distinct codes. Unweighted frequency distribution or summary statistics. These distributions should show both valid and missing cases. Imputation and editing information. Documentation should identify data that have been estimated or extensively edited. Details on constructed and weight variables. Datasets often include variables constructed using other variables. Documentation should include “audit trails” for such variables, indicating exactly how they were constructed, what decisions were made about imputations, and the like. Ideally, documentation would include the exact programming statements used to construct such variables. Detailed information on the construction of weights should also be provided. Location in the data file. For raw data files, documentation should provide the field or column location and the record number (if there is more than one record per case). If a dataset is in a software-specific system format, location is not important, but the order of the variables is. Ordinarily, the order of variables in the documentation will
3 7. Fill in the details 8. Click on “Next” or on “Subjects” 9. Select at least one main subject category (under Production systems you can find “Breeding, genetics and propagation”) 10. Click on “Next” or on “Affiliation” 11. Click on “European Union” and select “LiveSeeding” by clicking on the respective “Add” 12. Select additional names or codes that where involved in creating the deposit 13. Click on “Next” or on “Upload” 14. Upload File 15. Click on “Show options” and choose under “Visible to” the level of access restriction of this document 16. If you have data to upload, click on “Upload” and add a new file 17. Deposit the item 18. Click on “Manage deposits” to make changes or check the Item Status II. Zenodo (compulsory for peer-reviewed publications) 1. Go to https://zenodo.org 2. Create an account (you need to register in order to be able to deposit items in the repository) 3. Click on “Upload” 4. Click on the green “New upload” button 5. Import the file by dragging it into the appropriate field or by clicking “choose files” 6. Click “start upload” 7. In the “community” box, add the LIVESEEDING community (there are also communities for FiBL and Eucarpia) 8. In the “License” box, enter the creative common code: CC BY-NC-ND 4.0 and choose the appearing: Creative Commons Attribution Non Commercial No Derivatives 4.0 International 9. In the “License” box, you can choose the access right 10. In the “Funding” box, find European Commission (EU) and enter the grant number: 727230. 11. In the “Basic information” box, add the DOI as Digital Object Identifier if your publisher already assigned one 12. Fill in all the other relevant information 13. Click “Publish” 14. The curator will have to approve the file in order for it to be added to the community Note: once a file has been published, it cannot be removed or modified, but the information about the publication can be modified by going to the file and clicking the orange “edit” button. A new version of the file can be added by going to the file and clicking on “new version”. Via ZENODO the publications are automatically stored on OpenAIRE, the European Union’s electronic gateway for peer-reviewed articles and other important publications (www.openaire.eu).
4 By storing and publishing (after the embargo period) the scientific publications and other output on ZENODO, each partner of LIVESEEDING ensures that bibliographic metadata are included such as the funding body, the name of the action, acronym and grant number (already predefined in ZENODO), the publication date, and a persistent identifier. III. EU-Portal (compulsory for peer-reviewed publications) 1. Go to here 2. Sign up now 3. In the row of the LIVESEEDING project click on “Actions” and choose “Manage Project” 4. Click on “Continuous Reporting” 5. In the new window, click on “Publications” 6. Click on “Manually add publications” 7. Fill in the details 8. In the “Repository Link” box, you should provide the direct link to the publication 9. After you have filled all information, click on “Add publication” If you have any questions, please contact Mariano Iossa ([email protected]) - LiveSeeding Project Coordinator, AND Monika Messmer ([email protected]) LiveSeeding Scientific Project Coordinator in CC
5 Annex 10 Ontology for interoperability