WUR DMP Template and Guidance v09-01 Authors: Irene Verhagen https://orcid.org/0000-0001-5588-1333 Sydney Wizenberg https://orcid.org/0009-0007-4989-2424 Date: 20251216 Organization: Wageningen University & Research https://ror.org/04qw24q55 Licence: https://creativecommons.org/licenses/by/4.0/
1 Instructions on how to use this template: ❖ Use this template to fill in a data management plan (DMP) for your research project. This template pertains to research data and research software level low. Research software level low are scripts and code used to validate own research results and is not intended for reuse as-is in other research projects (source: Martinez-Ortiz et al. 2023, DOI: 10.5281/zenodo.7589725). ❖ For PhD candidates: The DMP can be attached to your project proposal as an appendix. ❖ This template is based on the DMP core requirements as defined by Science Europe: ‘Practical Guide to the International Alignment of Research Data’. ❖ To date, there are no known funders that make their own templates mandatory. We strongly advise to use this WUR DMP template as it much better aligns to practices at WUR. Please check the appendix for specific guidance for NWO and ZonMW funders. ❖ You are free to add topics to this template to better align with your project, however, the original topics must be retained. ❖ To get to additional information in the appendix for each section, hold keyboard key CTRL + left-click [info], or right-click [info] and select ‘open hyperlink’. ❖ For a filled in DMP example have a look at DOI: 10.5281/zenodo.7096699. ❖ You can also find this DMP template in DMPonline (https://dmp.wur.nl/). Have a look at DOI: 10.5281/zenodo.7073740 on how to get started in DMPonline. ❖ Any software with level medium and / or high (as outlined by the eScience Center in DOI: 10.5281/zenodo.7589725, page 22-28) for which the intent is to provide a reusable product that others can incorporate in their own research / project / workflow or which is used in decision making at institutional, societal, or governmental level can be described in a software management plan (SMP). Find the WUR SMP template via DOI: 10.5281/zenodo.10473645. ❖ For a review of your DMP, contact
[email protected] through email or through DMPonline. ❖ Questions? Contact d[email protected] or visit the WUR research data management website at https://www.wur.eu/rdm for more information. ❖ Please note that questions can have multiple answers checked where applicable.
2 A. Describe the research project [info] 1. Name researcher: 2. What is the name of your department(s)? ☐ Agrotechnology and Food Sciences ☐ Animal Sciences ☐ Environmental Sciences ☐ Plant Sciences ☐ Social Sciences ☐ Wageningen Food Safety Research ☐ Wageningen Food & Biobased Research ☐ Wageningen Economic Research ☐ Wageningen Plant Research ☐ Wageningen Environmental Research ☐ Wageningen Marine Research ☐ Wageningen Bioveterinary Research ☐ Wageningen Livestock Research
3 3. What is the name of your chair group(s) or business unit(s)? Please copy-paste the English name and abbreviation for: - chair groups from this page. - business units from this page (expand to Wageningen Research and keep expanding to find your specific division / group). Example: Bioprocess Engineering (BPE) or Green Economy and Landuse (GEL). 4. Describe the organisational context of your research project. DMP version (or date last modified) Supervisor / (co-)promotors Graduate School (WU only) Start date of project End date of project Project number Funding body 5. Give a short description of your research project. Title Summary
4 6. List the individuals and / or parties responsible for the following data management tasks. Data collection Data quality Storage and backup Data archiving / publishing Data stewardship / support Any other role [......] 7. I have requested a review of this data management plan from: ☐ WUR Library – Research Data Management Support (
[email protected]) ☐ The (coordinating) data steward of my chair group / business unit ☐ No review requested 8. Name of data management support staff and / or data steward consulted during the preparation of this plan and date of consultation.
5 B. Describe the data to be collected, software used, file formats and data size [info] 9. Will you use existing data for this project? ☐ Yes. Please specify below which data (e.g. DOI, URL, or storage location) and the terms of use (e.g. licence). ☐ No. Please describe below any constraints to reusing existing data. 10. Will new data be produced? ☐ Yes. ☐ No. 11. Please describe the data you expect to generate and / or use in the table below. Include reused existing data as well (as these are files that you manage and store). File contents Data type Software (Open) file format Estimated size of each file (range) Estimated number of files (range) (e.g. lab analysis, gene sequence, interviews, lesion scores, etc.) (e.g. numerical) (e.g. Excel) (e.g. .csv) (e.g. 20 – 50 Mb) (e.g. 50 – 100)
6 12. Estimate how much data storage you require in total (e.g. by using the information in the table at question 11). ☐ 0-10 GB. ☐ 10-100 GB. ☐ 100-1000 GB. ☐ >1000 GB.
7 C. Storage of data and data documentation / metadata during research [info] 13. Where will the data, code and accompanying documentation / metadata be stored and backed up during the research project (see the WUR Data Storage Finder)? Include platforms you use to share data, collect data on, or send data to for processing or analysis. ☐ W:drive Enterprise File Storage (WUR network drive). ☐ W:drive Massive File Storage Disaster Recovery (WUR network drive). ☐ W:drive Massive File Storage without Disaster Recovery (WUR network drive). ☐ M:drive – only when an up-to-date version of the research data is also safely stored on the W:drive or Yoda@WUR. ☐ Yoda@WUR (Yoda hosted by SURF for WUR). ☐ WUR OneDrive for Business - only when an up to date version of the research data is also safely stored on the W:drive or Yoda@WUR. ☐ WUR SharePoint / Teams - only when an up to date version of the research data is also safely stored on the W:drive or Yoda@WUR. ☐ Git@WUR (GitLab locally hosted at WUR). ☐ Other, please specify below the storage medium / system and describe back-up frequency, access management, and geographic location (e.g. within or outside the EU).
8 D. Structuring research data and information [info] 14. Give a (visual) representation of the folder structure you intend to use. 15. Describe the file naming conventions you intend to use. Please give one or multiple example(s). 16. How will you distinguish between versions of files (multiple answers possible)? ☐ Dates within file names are updated when files are modified. ☐ The designation ‘vRAW’ is added to file names that contain raw unaltered data (before any processing and cleaning). Any alteration of RAW data is done on a copy of the RAW data and appended with a version number which increases with each file modification (e.g. v01, v02, v03, etc.). ☐ We will use Git versioning for code / scripts. ☐ Other, please specify below.
15 G. Data archiving and publishing [info] 25. Are there reasons to restrict access to the data or limit which data will be made publicly available? ☐ No. ☐ Privacy / GDPR. ☐ Ethics. ☐ Contractual agreement. ☐ Commercial interests. ☐ Public security. ☐ IP rights. ☐ Other, please specify below. 26. Describe what data from question 11 will be archived (e.g. WUR network drive / Yoda@WUR) and not published, for a minimum of 10 years? Include the exact name of the storage medium chosen (see WUR Data Storage Finder). ☐ Not applicable as data will be published. ☐ Due to sensitivity we will need to archive (part of the) data underlying publications or reports. Please specify below which data and the chosen storage medium. ☐ Other, please specify below.
16 27. What data will be published and made available for reuse via a data repository? ☐ Data underlying publications or reports. Please specify below which data listed in question 11. ☐ Only the metadata is published in a data repository as the data are too sensitive to openly share. ☐ Data not underlying an article or report will also be published. Please specify below which data listed in question 11. ☐ Other, please specify below. 28. When will the data be available for reuse, and for how long will the data be available? ☐ Data will be available for at least 10 years as soon as the article or report is published and not required for any other article publication. ☐ Data will be available for at least 10 years upon completion of the project. ☐ Published data becomes accessible for at least 10 years after expiring of the embargo. Please specify below the reason for an embargo. ☐ Publication of data not underlying an article or report will be considered at the end of the project. ☐ Other, please specify below.
17 29. Which data repository do you intend to use to make the (meta)data findable and accessible (see the WUR Repository Finder)? ☐ DANS Data Stations. ☐ 4TU.ResearchData. ☐ Zenodo. ☐ Other, please specify below. 30. Which metadata standard will be used to describe the data during internal archiving and / or depositing in a data repository? ☐ Yoda metadata (DataCite metadata standard). ☐ Metadata standard from DANS Data Stations, 4TU.ResearchData and / or Zenodo (which often are the DublinCore or DataCite standard). ☐ A discipline-specific metadata standard. Please specify below. ☐ Other, please specify below.
18 31. Which licence/terms of use will be applied to the data? ☐ Open access (Creative Commons Attribution licence (CC BY); anyone can access and reuse with attribution). ☐ Restricted access (custom licence text or data sharing agreement is required, dictating restrictions of access and reuse). When a data sharing agreement is required, the Privacy Officer or Information Security Officer is consulted. ☐ Closed access (only metadata published, data is not allowed to be requested or reused). ☐ Other, please specify and/or provide link to the licence below.
19 H. Data management costs [info] 32. What resources (in time and / or money) will be dedicated to data management, data archiving or publication, and ensuring that data is reusable? Indicate as well how these costs will be covered. ☐ All costs for (long-term) data storage, data publication fees, and long-term access management to the data are covered by the research group / project. ☐ The PhD candidate and supervisor will spend at least 10% of their time on research data management to approach the FAIR principles as much as possible. ☐ Other, please specify below.
20 Guidance: additional information (may be deleted after completion) A guidance. Describe the research project. 1. Name researcher Please add your full name. 2. Name department Please choose your department. If you are working from multiple departments, you can choose multiple answers. 3. Name chair group or business unit Please add the English name of your chair group or business unit exactly as specified on the chair group or business unit webpage. Please carefully check the spelling. Preferably copy and paste directly from aforementioned websites. 4. Organisational context of your research project A data management plan should always include contextual information about the researcher and the project. Then, it is clear to whom / which project the filled in data management plan belong and to which data the described data management practices apply. 5. Description of your research project Giving a short description of your research helps the reader to understand your work and puts the description of management of research data into context. 6. Data management responsibilities Identifying persons who play a role in your daily data management practices helps clarify the data collection process. Identifying these roles are also important in the event you leave WUR or on completion of your PhD. For data to be accessible for at least 10 years, responsibility for the archived data should lie with more than one person. Feel free to add roles. 7. Requested review You can request a review of the DMP from
[email protected] and / or from your local data steward. Provide the contact details of the reviewer in the next question. 8. Data management support staff You can consult with Research Data Management (RDM) support staff or (coordinating) data steward to review your data management plan. For RDM support staff to review your data management plan, contact
[email protected] or request feedback via the ‘Request feedback’ tab in DMPonline.
21 Is this project funded by NWO? If yes, NWO requires that researchers consult with RDM support at an early stage. Plans that have not been consulted with / reviewed by RDM support staff will not be considered by NWO. Please do so by contacting
[email protected] before submitting or via the ‘Request feedback’ tab in DMPonline. B guidance. Describe the data to be collected, software used, file formats and data size 9. Use existing data When others publish data, this gives you the opportunity to search for and find data to reuse within your research. However, make sure that you are aware of the terms of use (e.g. the licence) of the data that you want to reuse. The terms of use provide you with what you are allowed to do with the data. Want to find existing data? Have a look here for places to start. In the case of not reusing data: think of the following potential reasons why the use of existing data was considered but not implemented. Is the data you want to reuse not publicly available or otherwise restricted in access? Would the costs be too high (e.g. acquiring specific software)? Is the data too big in size (e.g. TBs)? 10. New data Indicate whether you will collect / generate new data in the research project. 11. Description newly generated data. Data type refers to whether the data can be classified as for example textual, numeric, audio, film etc. Note that data includes any processing or analysis scripts, protocols, or any other output resulting from research (figures, statistical output, processed data etc.). When filling in the table, be as distinctive as possible for origins of the data. For example, if you do laboratory work and use ELISA, FACS, and manual cell counts, those assays and techniques will likely have their own outputs, it is then not sufficient to only add 1 row 'laboratory output' (you have at least 3 rows of different laboratory data file origins). If you are collecting audio from interviews, focus groups, and other resources, it is not sufficient to only have 1 row specifying 'audio files' (you have at least 3 rows of different audio files origins). Try to use software that let you also save files in an open format, so that others can open the files (now and in the future) even if they don’t have the software.
22 Providing the estimated number of files and file size helps to indicate the storage requirements and costs. For example, in Windows, the size of a file and its size on disk are two different measures. A disk is comprised of ‘disk allocation units’ or clusters. Files occupy as many clusters as required and a cluster cannot be shared between files. As such, a file can leave a cluster incompletely filled (leading to a bigger file size on disk). Cluster size can vary, which influences (also depending on the file sizes) the storage capacity needed. With a cluster size of 1024 KB, many files of 500 KB take up a lot of storage capacity, because 524 KB is ‘wasted’ per cluster. 12. Storage requirement According to the table in the previous question, estimate the amount of storage space you expect to need. C guidance. Data and documentation / metadata storage during research 13. Where will data be stored Ensure that you store data and accompanying documentation / metadata safely in compliance with the WUR research data policy and the WUR Guidelines on Value Creation with Software and Data. Safe storage solutions would be for example the solutions indicated in the WUR Data Storage Finder, possibly in combination with ITapproved cloud storage. Note that if you store data and accompanying documentation / metadata with a WUR cloud service or on the WUR network (except for W:drive – Massive File Storage), they are backed up automatically. Contact your Information Security Officer for help in determining appropriateness of other storage or platform solutions for research data. D guidance. Structuring research data and information 14. Folder structure Designing a logical folder structure (for tips, see here) ensures that you and fellow researchers can easily locate data now and in the future. Provide how you plan to organise your (sub)folders. If you already have a folder structure in place, an easy way to document the structure is via the following PowerShell workflow (note: the following steps do not work with zip files): - Go to your file explorer. - Open the parent folder. - In the address bar, type and press enter afterwards: powershell - A blue screen pops up. Continue in that screen. - Type and press enter afterwards: Get-ChildItem | tree > foldertree.txt
23 - There is now a file in your folder [foldertree.txt], from which you can copy the contents in the answer field. If you’re using Git Bash on Windows, go to your parent directory (or main local Git directory) and use: - cmd //c tree //a > foldertree.txt - There is now a file in your folder [foldertree.txt], from which you can copy the contents in the answer field. If you are using Linux: - Type and press enter afterwards: sudo apt install tree - Go to your project parent directory using the 'cd' command. - Type and press enter afterwards: tree -d > foldertree.txt - There is now a file in your folder [foldertree.txt], from which you can copy the contents in the answer field. 15. File naming conventions Applying a consistent and descriptive file naming convention (i.e. a systematic file naming method) helps to: - identify the content of a (data)file without opening it. - easily and quickly locate, retrieve and filter (data)files, even if they have changed folders. - easily sort and browse through your (data)files. - identify missing (data)files. A basic example of a file name: [project]_[filesubject]_[subsubject]_[date]_[version].[extension] Use the international standard for date (YYYYMMDD) and apply a leading zero to the version number (e.g. v01) which aids in ordering files. Find tips for naming files here. 16. File versioning Make sure that you have a system in place to keep track of file versions. A simple and effective system is to incorporate version numbers (e.g. v01, v02, etc.) in your file names. Additionally, depending on the type of files you work with, you could use systems that keep track of versions of individual files (e.g. Sharepoint) or multiple files (e.g. Git). Note: for files in Git, dates and version numbers should not be added to the filenames. Git is version control software, keeping track of versions for you and without having to apply file versioning yourself. When applying versioning by changing file names, Git would not recognise a new version as such, but as a completely different file.
24 E guidance. Data documentation and metadata It is essential to systematically document research data during your research and when preserving data, whether done by archiving at WUR or depositing data into a data repository. Documentation is information added to data to ensure that the data is understandable and reusable to yourself and to others, both during and after your research. Metadata is also documentation, but in a structured form, describing the data in a way which facilitates cataloguing and discovery of the data. 17. Data documentation and metadata The most common required form of documenting research data is by adding a readme file, which is further supplemented with metadata. Independently of where data is preserved (e.g. at WUR, in a data repository) it should be accompanied by: - a readme file The readme file contains information about e.g. the steps that have been undertaken in processing and analysing data. In short: all information necessary to understand the data, reproduce research and verify results. WUR Library – RDM support advises to use the WUR readme file template (DOI: 10.5281/zenodo.7701727) as the minimum required documentation to add to the data. Feel free to add more documentation where appropriate and required. - metadata Metadata is machine-readable information about the data, according to fixed terms, which makes the data findable and searchable. WUR Library – RDM support advises to use the Yoda metadata terms as the minimum required metadata to add to the data. You can fill in these terms in Yoda@WUR, when applicable as a storage solution, or use the public Yoda metadata editor and download the metadata as a .json file. - a codebook WUR Library – RDM support advises to use the WUR codebook template (DOI: 10.5281/zenodo.7701727) to explain variables, abbreviations, etc, because this template is in .csv format, which makes it easy to import into software such as R, Python, etc. You can find the templates for the readme file and codebook at DOI: 10.5281/zenodo.7701727. Via this link, a filled in example of a readme file, codebook, and metadata file can be found as well.