scieee AI-readable full text Open interactive document viewer

Plasma-PEPSC D7.3 Updated Data Management Plan

Williams, Jeremy J.; Hegde, Pratibha Raghupati; Juckeland, Guido

Abstract

Data management is a critical component of research and innovation projects. It involves handling data, preserving it, and making it accessible to other researchers. This updated Data Management Plan (DMP) highlights the datasets produced as part of the project activities that are relevant for inclusion in the DMP. It outlines the types of data generated or gathered during the project, the standards applied, the methods for exploiting and sharing data for verification or reuse, and the strategies for data preservation. The DMP is a living document that evolves throughout the project’s lifecycle, particularly when significant changes occur, such as dataset updates or modifications to Consortium policies. This updated version of the DMP builds on the foundational document delivered in Month 6 of the project. It provides an overview of the datasets produced by the project and outlines the specific conditions attached to them. While this version continues to cover a wide range of aspects related to Plasma-PEPSC data management, it expands on previous iterations by addressing key areas such as data interoperability and the practical data management procedures implemented by the Plasma-PEPSC project consortium. This document has been developed in accordance with established guidelines and serves as a consolidated plan for Plasma-PEPSC partners, ensuring adherence to the project’s DMP policy.

Full text

HORIZON JU Research and Innovation Actions HORIZON-EUROHPC-JU-2021-COE-01-01 European High-Performance Computing Joint Undertaking Plasma Exascale-Performance Simulations CoE 101093261 D7.3 Updated Data Management Plan WP7: Management Date of preparation (latest version): 31/12/2024 Copyright©2023 – 2027 The Plasma-PEPSC Consortium D7.3: Updated Data Management Plan 2 DOCUMENT INFORMATION Deliverable Number D7.3 Deliverable Name Updated Data Management Plan Due Date 31/12/2024 Deliverable lead KTH Authors Jeremy Williams (KTH), Pratibha Hegde (KTH) and Guido Juckeland (HZDR) Responsible Author Jeremy Williams (KTH) E-mail: [email protected] Keywords Information Governance, Lifecycle Management, Privacy and Confidentiality, Validation and Quality Assurance, Access Control Policies, Information Security Measures, and Resource Compliance WP/Task WP7/Task D7.3 Nature DMP Dissemination Level PU Final Version Date 31/12/2024 Reviewed by Erwin Laure (MPG) & David Tskhakaya (IPP CAS) D7.3: Updated Data Management Plan 3 DOCUMENT HISTORY Partner Date Comment Version KTH 03/03/2024 Skeleton version of the deliverable 0.1 KTH 04/11/2024 Draft Sections and Text added 0.2 KTH 27/11/2024 Draft Sections and Initial Text Updated 0.3 KTH 02/12/2024 First draft 0.4 KTH 03/12/2024 Final draft updated for internal review 0.5 KTH 16/12/2024 Revised draft after internal review 0.6 KTH 23/12/2024 Final cleanup for submission 1.0 D7.3: Updated Data Management Plan 4 Executive Summary Data management is a critical component of research and innovation projects. It involves handling data, preserving it, and making it accessible to other researchers. This updated Data Management Plan (DMP) highlights the datasets produced as part of the project activities that are relevant for inclusion in the DMP. It outlines the types of data generated or gathered during the project, the standards applied, the methods for exploiting and sharing data for verification or reuse, and the strategies for data preservation. The DMP is a living document that evolves throughout the project’s lifecycle, particularly when significant changes occur, such as dataset updates or modifications to Consortium policies. This updated version of the DMP builds on the foundational document delivered in Month 6 of the project. It provides an overview of the datasets produced by the project and outlines the specific conditions attached to them. While this version continues to cover a wide range of aspects related to Plasma-PEPSC data management, it expands on previous iterations by addressing key areas such as data interoperability and the practical data management procedures implemented by the Plasma-PEPSC project consortium. This document has been developed in accordance with established guidelines and serves as a consolidated plan for Plasma-PEPSC partners, ensuring adherence to the project’s DMP policy. D7.3: Updated Data Management Plan 5 Contents 1 Introduction 6 2 General Principles Update 7 2.1 FAIR Research Data Management . . . . . . . . . . . . . . . . . . . . . . 7 2.2 Datasources.................................. 7 2.3 Primary research data generated by the project . . . . . . . . . . . . . . 8 2.4 Confidentiality and data protection concerns . . . . . . . . . . . . . . . . 8 2.5 Preferred data formats . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 2.6 Tools for validating results . . . . . . . . . . . . . . . . . . . . . . . . . . 9 2.7 Sourcecode .................................. 9 2.8 Datasecurity ................................. 9 2.9 Dataquality.................................. 10 2.10 Allocation of resources . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 2.11Ethics ..................................... 11 3 Data Repositories and Management of Intellectual Properties Update 11 3.1 Publicrepositories .............................. 11 3.2 Consortium private repository . . . . . . . . . . . . . . . . . . . . . . . . 12 3.3 Management of Intellectual Properties . . . . . . . . . . . . . . . . . . . 12 4 Research Data Outputs Update 13 5 Research Data Outcomes and Impact Update 13 6 Conclusion 14 D7.3: Updated Data Management Plan 6 1 Introduction The purpose of the Data Management Plan (DMP) deliverable is to provide relevant information concerning the data that will be collected and used by the partners of the Plasma-PEPSC project. The DMP describes the data management life cycle for all datasets to be collected, processed, and/or generated by the research project. It covers: •How data should be handled during and after the project. •What types and formats of data will be generated/collected. •Which methodologies and standards will be applied. •Whether the data be shared or made open-access, and how. •How data will be curated and preserved during the project as well as after its conclusion. As with the initial version, this updated Data Management Plan (DMP) provides preliminary information about the data generated by the project, including whether and how it will be exploited or made accessible for verification and reuse, as well as how it will be curated and preserved. The purpose of this DMP is to analyze the main elements of the data management policy that the consortium will follow for all datasets generated during the project. In this updated version, several significant changes have been made compared to the previous version. These include updates to the potential data formats, revisions to the Plasma-PEPSC Code Licenses and Repositories, now featuring mirrored releases to Castiel 2, and the addition of a new section on Research Data Outcomes and Impact. The document has been thoroughly reviewed and expanded to provide more detailed information, ensuring a comprehensive approach to data management. The DMP is a living document that will evolve throughout the project’s lifespan. It will be updated as needed to reflect significant changes, such as the introduction of new datasets, modifications to consortium policies, or changes in the consortium’s composition. The remaining part of the document is organized as follows: •Section 2 describes the updated general principles used to organize project data. •Section 3 describes the updated data repositories used throughout the project and the management of intellectual properties. •Section 4 describes the updated research data output plan for the project. •Section 5 describes the updated research data outcomes and impact. •Section 6 concludes the deliverable. D7.3: Updated Data Management Plan 7 2 General Principles Update This section describes the main aspects to be considered in Plasma-PEPSC’s data management that all partners of the project must follow. In general terms, Plasma-PEPSC’s research data should be “FAIR”, which is findable, accessible, interoperable, and reusable. 2.1 FAIR Research Data Management Plasma-PEPSC will produce two types of data generated by the BIT, GENE, PIConGPU, and Vlasiator plasma codes: •Raw data produced by the simulations as well as performance data from the simulation runs. The former consists of plasma quantities defined on a computational grid (densities, fluid velocities, pressure, temperature, . . . ) and information about particles or phase-space density (positions, velocities). This raw data will only be used for correctness checking as it can easily be reproduced by rerunning the simulations. •Performance data will consist of execution logs (trace files) as well as profile data, such as hardware performance counter samples. This performance data is the basis for analysis that will be based on a variety of tools and methods. The data generated with the help of these tools will be homogenized in the form of written performance reports that will be made public. These reports will always include artifacts that enable the reproducibility of the results, such as the software build setup, the input parameters, the execution parameters of the simulation, the result correctness check data, and the recorded performance data. The artifact will also point to the exact version (release or git-hash). In very few cases, it is expected to also store a reduced set of the simulation raw data as primary data, e.g., for post-mortem visualizations. Plasma-PEPSC will investigate the establishment of different systems for data handling, as well as for accessing performance data, analysis, and white papers generated by the project. As already mentioned before, the data will respect the FAIR principles (findable, accessible, interoperable, and reusable). In particular, the project will investigate the use of openPMD1, a FAIR particle and mesh data format for plasma simulations, developed initially at HZDR, one of its core maintainers to date, to store the direct simulation output. 2.2 Data sources There are three major sources of data in the frame of the project, as shown below. •Simulation Data (SD): SD including input, output, and performance data serves as one of the primary sources of data for the project. •Source Code for Simulations (SCS): SCS for all codes acts as a crucial data source in the project. 1https://github.com/openPMD D7.3: Updated Data Management Plan 8 •Documented Output Data (DOD): DOD such as deliverables, publications, manuals, reports, and more also contribute valuable data to the project. 2.3 Primary research data generated by the project The project’s primary research data will be coming from performance figures, figures supporting software engineering metrics, and other results used to support the research publications produced during the project lifespan. The research is expected to yield results in terms of, e.g., parallel performance (execution time and speedup), reduction in errors, improved time to solution, and improved maintainability. Full consideration will be given to allowing proper statistical analysis of the results using means, standard deviations, regression tests, and other relevant metrics. Performance data and software engineering results will be derived from parallel source programs. When legally and contractually possible, these sources will also be made public. Raw data resulting from research is data that has not been coded, grouped, refined, or modified in any way. Even if raw data has more potential use than modified data, as a general rule, the Plasma-PEPSC project will not provide raw data in open repositories due to data quality checking and Intellectual Property Rights (IPR) controls. Instead, each partner should keep their sets of raw data, which should be maintained untouched, if possible. 2.4 Confidentiality and data protection concerns Most of the data described above will be non-confidential, not refer to human subjects, and not introduce any security concerns to increase dissemination and data re-use. However, all research data will be collected and stored in line with European legislation on data protection, as relevant. For software and data, open-source licenses are promoted in the project. While some partners could retain part of the results with other licenses, this will not be the norm. This rule is explained in the project’s Consortium Agreement. Thus, data will be licensed to permit the most extensive re-use possible as soon as possible in the project timeline. Embargoes are not foreseen for software and data. All data will remain available after the end of the project in open data portals assuming that no budget is needed to keep this data open and reusable. 2.5 Preferred data formats Data will generally be stored as plain text to simplify processing and avoid possible problems with transcribing data between evolving data formats. For research data exchange, CSV (Comma-separated Values) format is preferred. However, each dataset stored in the repository will include a description of the data. Templates are provided for documents and other dissemination data of the Plasma-PEPSC project. Those templates are mandatory to enhance compatibility and inspection inside the project. All data elements must incorporate attribution of their original source, date, and authors. If several contributions are made over time, a change log is required. D7.3: Updated Data Management Plan 9 Table 1 shows the preferred formats and those accepted in the Plasma-PEPSC data catalog. Preferred formats have been chosen because they are suggested to have the highest probability of maintaining accessibility and readability in the future. Accepted formats are commonly used formats that have good prospects of remaining readable in the long term. Document type Preferred format Accepted format Text documents plain (.txt) MS Word PDF Rich Text Format TeX Markup language JSON XML XML Spreadsheets CSV MS Excel (.xls) OpenDocument Spreadsheet Images PNG TIFF SVG JPEG, PDF Databases CSV SQL JSON OpenDocument Base Simulation Outputs ADIOS2, HDF5, JSON ASCII VLSV NetCDF Table 1: Potential data formats in Plasma-PEPSC 2.6 Tools for validating results Complete information will be provided in the research publications about the tools that have been used to produce the results, including details of operating system versions, libraries, and specific software tools, as relevant. Most of the software tools that will be used in the project will be free, open-source software, either produced by third parties or produced by the project and disseminated under open licenses. Full care will, however, be taken to avoid releasing proprietary software and to achieve the best possible commercial exploitation for tools that are developed and/or modified during the project. Where tools are restricted, it will generally be possible to obtain them under some sufficiently liberal license agreement to allow full reproduction of the research experiments. 2.7 Source code Plasma-PEPSC’s source code(s) will be developed using the coding standard defined in the project. This coding standard is not mandatory for already-established projects that are already using their own coding standard that could keep on using the existing coding standard. 2.8 Data security All research data underpinning publications will be made available for verification and re-use unless there are justified reasons for keeping specific datasets confidential. The main elements when considering the confidentiality of datasets are: