RECHARGE: Data Management Plan
Abstract
This document outlines the DMP prepared for the RECHARGE project. It informs about the intentions of the consortium partners concerning data collection, its documentation, sharing and short- and long-term storage. It represents the most up-to-date knowledge and awareness of research data management, as well as data to be collected in the scope of the RECHARGE project. The strategy aims to enable effective, safe and sustainable data use, which will be actively evaluated throughout the project. The DMP will evolve during the lifespan of the project and will be regularly updated.
Full text
Resilient European Cultural Heritage As Resource for Growth and Engagement 6.1 Data Management Plan Acknowledgement Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Executive Agency (REA). Neither the European Union nor the granting authority can be held responsible for them. 1
Project description Acronym RECHARGE Title Resilient European Cultural Heritage As Resource for Growth and Engagement Coordinator Erasmus University Rotterdam (EUR) Project number 101061233 Type of action HORIZON Research and Innovation Actions Topic New ways of participatory management and sustainable financing of museums and other cultural institutions Project end date 30 September 2025 Project duration 3 years Website https://recharge-culture.eu/ E-Mail [email protected] Deliverable description Number 6.1 Title Data Management Plan Lead beneficiary EUR Work package 6 Dissemination level Public Type Report Due date 31/03/2023 Submission date 29/03/2023 Resubmission date: Authors Joanna Mania, Ilaria Rosetti Contributors Reviewers Una Hussey, Marta Franceschini, Simon Thompson, Trilce Navarrete Resubmission edits Acronyms and definitions Acronym Meaning DMP Data Management Plan EC European Commission FAIR Findable, Accessible, Interoperable, Re-usable GDPR General Data Protection Regulation WP(s) Work Package(s) 2
Abstract This document outlines the DMP prepared for the RECHARGE project. It informs about the intentions of the consortium partners concerning data collection, its documentation, sharing and shortand long-term storage. It represents the most up-to-date knowledge and awareness of research data management, as well as data to be collected in the scope of the RECHARGE project. The strategy aims to enable effective, safe and sustainable data use, which will be actively evaluated throughout the project. The DMP will evolve during the lifespan of the project and will be regularly updated. About the RECHARGE Project RECHARGE stands for Resilient European Cultural Heritage As Resource for Growth and Engagement, which are keywords that together synthesize the aim of the project itself: to reinvigorate community participation as added economic value for cultural heritage institutions across Europe. Funded within the European Union's HORIZON Europe programme, the key funding programme for research and innovation within the EU, RECHARGE supports cultural heritage institutions in diversifying their funding through a replicable and sustainable participatory business model, to acquire the necessary tools for its future developments, both in the digital realm and onsite. RECHARGE builds on new and existing communities, networks and relationships related to cultural heritage institutions, to engage them in participatory management through cultural heritage Living Labs. These Living Labs, developed collaboratively and open to professionals and the public, aim at testing and devising innovative ways to harness resources, to ensure the development of sustainable future business models, focused on the creation and integration of value within each institution, and in the sector at large. Consortium Partner organisation name Acronym Country Erasmus University Rotterdam EUR NL Centrum Cyfrowe FCC PL Fundación Goteo GOTEO ES Stichting Nederlands Instituut voor Beeld en Geluid NISV NL European Fashion Heritage Association EFHA IT Creativity Lab CLAB EE University of Valladolid UVA ES Estonian Maritime Museum EMM EE Textile Museum Prato TMP IT The Hunt Museum HUNT IE 3
Contents 1. Purpose of the Data Management Plan 5 Table 1: Revisions 5 2. Data summary 5 Table 2: Data Summary 5 3. FAIR Data 6 3.1 Making data findable, including provisions for metadata 7 3.1.1 Persistent Identifier: DOI 7 3.1.2 Metadata schema: Dublin Core Metadata Initiative (DCMI) 7 Table 3: (list of metadata) 7 3.2.1 Repository: Zenodo 7 3.2.2 Accessibility of datasets 8 3.2.3 Accessibility of metadata 8 3.3 Making data interoperable 8 3.4 Increase data re-use 8 3.4.1 License 9 3.4.2 Project-level documentation 9 3.4.3 Naming and versioning 9 3.4.4 Data quality assurance 10 4. Other research outputs 10 5. Allocation of resources 11 5.1 FAIR data preparation 11 5.2 Infrastructure 11 5.3 Expertise 11 6. Data security 11 6.1 Secure storage infrastructure 11 Table 6: Specifications of data storage solution 11 7. Ethics 12 8. GDPR compliance 12 9. Other issues 14 4
1. Purpose of the Data Management Plan The DMP aims to provide a clear outline of the key principles and provisions regarding data within the scope of the project. In short, it specifies what datasets RECHARGE will generate or re-use, how these will be exploited and made available for re-use, and how the datasets will be curated and archived. The RECHARGE DMP also describes how the research data within the project will meet FAIR guiding principles to become findable, accessible, interoperable and reusable. This document draws on the template provided by the European commission. The first section of the DMP summarizes the relevant datasets within the RECHARGE project, and it is followed by FAIR data provisions and allocation of resources. In attrition, the document covers aspects of data security, ethics, and GDPR compliance. It should be noted that, for this first version, the document concerns mostly predicted or expected datasets, since empirical data collection has not fully started yet. Table 1: Revisions Version Submission date Comments Author v0.1 15/03/2023 Internal review version Joanna Mania, Ilaria Rosetti v0.2 15/03/2023 - 23/03/2023 Revision Una Hussey, Marta Franceschini, Simon Thompson, Trilce Navarrete v0.3 28/03/2023 Revision Gai Umali, Chiara Stenico V1.0 31/03/2023 Submission version Joanna Mania, Ilaria Rosetti 2. Data summary A general overview of the datasets collected in the RECHARGE is presented in Table 2: Data Summary. It covers their types, formats, the purpose of collection, expected size, origin, and related considerations. The overview is non exhaustive and will be updated throughout the project. Table 2: Data Summary WP Type Formats Considerations Size Orig in Data Access 1 T1.1 Analysis of the Current Participatory Management and Participatory Business Models in Cultural Heritage Sector and Beyond Desk research (articles, books, reports) .docx, .txt, .xlsx, .csv Keywords search, data citation <1GB G O T1.2 Motivation for Participation Focus groups .mp3, .mp4, .docx, .txt, .xlsx, .csv Informed consent, anonymization, content curation, digitalization of notes, data sharing <10GB G O, R T1.3 Value of Participation Survey .xlsx, .csv, .docx, .txt Informed consent, anonymization, codebooks, data sharing <1GB G O T1.4 Forging Impactful and Durable Partnerships and Collaborations Interviews .mp3, .docx, .docx, .txt, .xlsx, .csv Informed consent, content curation, data sharing <10GB G O, R 5
2 T2.2 Co-creating the RECHARGE models Co-creation workshops .mp3, .mp4, .jpeg, .docx, .txt, , .xlsx, .csv Informed consent, pseudonymization, digitalization of workshop outputs, content curation, data sharing <10GB G O, R T2.4 Running the CH Living Labs TBU TBU TBU TBU G TBU T2.5 Analysis of the Prototypes Observations, workshop .docx, .txt, .mp3, .mp4, .xlsx, .csv Informed consent, pseudonymization, digitalization of workshop outputs, content curation, data sharing <10GB G O, R 3 T3.1 Effectiveness of Participatory Operational Strategies Survey .xlsx, .csv, .docx, .txt Informed consent, anonymization, codebooks, data sharing <1GB G O, R T3.2 Measuring Resilience Interviews .mp3, .docx, .docx, .txt, .xlsx, .csv Informed consent, pseudonymization, content curation, data sharing <10GB G O, R T3.3 Viability and Value Creation Capacity Interviews, survey, secondary datasets .mp3,.docx, .txt, .xlsx, .csv Informed consent, pseudonymization, anonymization, codebooks, data sharing, license, data citation <10GB G, R O, R T3.4 Indicators of Participation Interviews, secondary datasets .mp3,.docx, .txt, .xlsx, .csv Informed consent, pseudonymization, data sharing, data citation <10GB G, R O, R Origin: R – reuse, G – generate, Data Access: O – Open, R – Restricted, TBU – to be updated, NA – not applicable, A/A – as above 3. FAIR Data FAIR stands for findable, accessible, interoperable and reusable. It is a framework for thinking about, preparing and sharing data in a way to maximize its use (and reuse). Finable refers to discovery of datasets, which is achieved through registering them in secure data repositories together with metadata (description of the data) and persistent identifiers (a unique, long-tern link). Accessible means that both humans and computers should gain access to the (meta)data, and when appropriate, under specified restrictions and/or defined access protocols. Research data should be set to be easily combined with other datasets, meaning interoperable. Lastly, preparing data for re-use in a future requires making it self-evident and replicable by including transparent documentation and re-use license. 6
3.1 Making data findable, including provisions for metadata 3.1.1 Persistent Identifier: DOI Datasets, research materials and research outputs are assigned DOI via Zenodo. It is an obligatory element of every record. DOIs of the datasets and related publications are included in the accompanying metadata. 3.1.2 Metadata schema: Dublin Core Metadata Initiative (DCMI) Zenodo follows the DCMI metadata standard. On the day of DMP submission, next to indexing published records in its own search engine, Zenodo sent and indexes metadata in the DataCite servers, links data to OpenAIRE publications and to grant in the EC Participant Portal. A set of keywords will be selected for consistency and to optimize the discoverability and re-usage of the (meta)data. Table 3: (list of metadata) Name of metadata standard Description 1. Title (M) A name given to the resource. 2. Creator (M) An entity responsible for making the resource. 3. Subject (MA) A topic of the resource. 4. Description (MA) Description may include but is not limited to: an abstract, a table of contents, a graphical representation, or a free-text account of the resource. 5. Publisher (MA) An entity responsible for making the resource available. 6. Contributor (R) An entity responsible for making contributions to the resource. 7. Date (M) Recommended practice is to express the date, date/time, or period of time according to ISO 8601-1 (YYYY-MM-YY) 8. Type (M)/(R) The nature or genre of the resource. 9. Format (R) The file format, physical medium, or dimensions of the resource. 10. Identifier (M) An unambiguous reference to the resource within a given context, for example DOI. 11. Source (R) A related resource from which the described resource is derived. 12. Language (R) A language of the resource. 13. Relation (R) A related resource. 14. Coverage (O) Spatial topic and spatial applicability may be a named place or a location specified by its geographic coordinates. 15. Rights (R) Access Rights may include information regarding access or restrictions based on privacy, security, or other policies. M – Mandatory, MA – Mandatory when applicable, R – Recommended, O – Optional 3.2 Making data accessible 3.2.1 Repository: Zenodo RECHARGE chooses to deposit metadata and data in Zenodo, an open, general-purpose repository created under the OpenAIRE programme. All records will be attached to the ‘RECHARGE community’. Metadata and record identifiers can be harvested via OAI-PMH and REST API. For the purpose of the RECHARGE project, the consortium does not foresee additional agreements with Zenodo, except agreeing to the Terms of Use. 7
3.2.2 Accessibility of datasets RECHARGE consortium partners will make datasets and other project outputs accessible for third parties for re-use. The openness of data will depend greatly on the type of data classification. Non-personal data and materials will be published open access. For data containing personal information the rule is as follows. If the balance between pseudonymization/anonymization of qualitative data and data quality can be achieved, the consortium partners will consider making them open access for re-use. If not, the qualitative data will be published under restricted access, and metadata will include information about how and under what condition to access data. Special considerations will be given to community participants, if they are of a vulnerable nature or classified as minor. Embargo options might be considered by some of the project partners. The need of establishing a data access committee will be discussed at a later stage of the project. 3.2.3 Accessibility of metadata The access to the data is public, and no authorization is required to retrieve it. Metadata of the datasets published under restricted access will include information on how to access the dataset. According to the Zenodo FAIR Principles, metadata is stored in high-availability database servers at CERN, which is separate from the data itself. Data and metadata will be retained for the lifetime of the host laboratory CERN, which currently has an experimental program defined for the next 20 years at least. When necessary, information about software (and their versions) needed to open, process or analyze data will be included in the dataset documentation. 3.3 Making data interoperable The research datasets, metadata and documentation will be as much as possible compliant with best standards for interoperability and re-use. Description of the datasets and metadata are in English to increase their discoverability and accessibility. Data formats listed in Table 2 will be revised and converted to non-proprietary formats before publication. The relation between resources can be made through the “Relation” element in the metadata. While uploading data to the repository, consistent keywords will be used. In terms of interoperable metadata, Zenodo uses JSON Schema as an internal representation of metadata and offers export to other popular formats such as Dublin Core or MARCXML. 3.4 Increase data re-use 8
3.4.1 License RECHARGE data, materials, playbooks, workflows, protocols, etc are of interest to cultural heritage institutions and researchers in the field of heritage studies, cultural economics, entrepreneurship and innovation, among others. Striving for a greater impact of the RECHARGE project, better reproducibility, transparency and possibility for rigorous peer-review process, research data (non-personal) and other research outputs, will be labeled with Creative Commons CC0 license. This license places the copyright in the public domain, so that others may freely build upon, enhance, and re-use the works for any purposes without restriction under copyright or database law. 3.4.2 Project-level documentation The primary project-level documentation will consist of descriptive README files. The example README file is included in the SURF Research Drive folder for each WP (file: README-template.txt). The will contain, among others, the following information: - title of the project/dataset - motivation for data collection - research question - a brief description of the contents of each file or group of files - contract information - methodology - tools used for data collection - dates of data collection - software used for data processing and/ or analysis - the provenance of the data and any prior processing (if re-using data) - license under which the data can be reused, of the access level placed on the data 3.4.3 Naming and versioning Each WP is free to create their own system following the best practices listed below. The same list is included in the SURF Research Drive folder for each WP (file: naming-and-versioning.txt). ➔Avoid spaces, punctuation, case sensitivity and special characters ➔Deliberate use of delimiters. For example, a hyphen (-) to mean “different words that are part of the same chunk”, and underscore (_) to separate different chunks of metadata ➔Choose keywords and file names that are sufficiently descriptive, ➔Use YYYY-MM-DD date format (ISO 8601 standard) ➔To order files put date or number first ➔Include the version of the file ➔Create a version control table that lists the number of changes, date, and their purpose ➔Add sequential numbers at the end of the file name (for example: _v1, _v2) ➔Record the date in the file naming Picture 1: Example of folder organization 9