scieee AI-readable full text Open interactive document viewer

ClearClimate Data Management Plan

Krejtz, Krzysztof

Full text

D1.1 Data Management Plan i Project: ClearClimate Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Research Executive Agency (REA). Neither the European Union nor the granting authority can be held responsible for them. Grant agreement No. 101059546 “Engaging approaches and services for meaningful climate actions” Call: HORIZON-MSCA-2022-SE-01 Topic: HORIZON-MSCA-2022-SE-01-01 Type of action: HORIZON-TMA-MSCA-SE Duration: 01/11/2023 - 31/10/2027 (48 months) Name of Deliverable Work package: 1 Task: 1.4 Issued by: SWPS Issue date: 29.04.2024 Due date: 30.04.2024 Work Package Leader: UNSPMF Document history (Revisions -Amendments) Version and date Changes v.1, Dissemination level PU Public x PP Restricted to other programme participants (including the EC Services) RE Restricted to a group specified by the consortium (including the EC Services) CO Confidential, only for members of the consortium (including the EC) Ref. Ares(2024)3186684 - 30/04/2024 D1.1 Data Management Plan ii LEGAL NOTICE Neither the Research Executive Agency/European Commission nor any person acting on behalf of the Research Executive Agency/Commission is responsible for the use, which might be made, of the following information. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Executive Agency (REA). Neither the European Union nor the granting authority can be held responsible for them. © ClearClimate Consortium, 2023 Reproduction is authorised provided the source is acknowledged D1.1 Data Management Plan iii Table of contents Summary ................................................................................................................................................. 1 1. Data Summary ............................................................................................................................... 2 1.1 Purpose ......................................................................................................................................... 2 1.2 Data types and formats ................................................................................................................. 2 Data source ......................................................................................................................................... 3 Raw data format ................................................................................................................................. 3 1.3 Re-use of existing data .................................................................................................................. 3 1.4 Origin of the existing data ............................................................................................................. 4 1.5 Generation of new data ................................................................................................................ 4 1.6 Expected size of datasets .............................................................................................................. 4 1.7 Data Usefulness ............................................................................................................................ 4 2. FAIR Principles ............................................................................................................................... 6 2.1. Making data Findable, including provisions for metadata ...................................................... 6 2.2. Making data Accessible ........................................................................................................... 7 2.3. Making data Interoperable ...................................................................................................... 8 2.4. Increase data Re-use ............................................................................................................... 8 3. Allocation of resources ................................................................................................................ 11 4. Data security ................................................................................................................................ 12 5. Ethics ............................................................................................................................................ 13 D1.1 Data Management Plan iv List of Figures D1.1 Data Management Plan 1 SUMMARY This Data Management Plan (DMP) describes the practice of providing online access to scientific information that is free of charge to the reader, while in the context of Research and Development, DMP focuses on access to 'scientific information' or 'research results'. ClearClimate DMP deals with the data collected, generated, and reused during the project lifespan. This DMP also identifies publications and metadata repositories used by the ClearClimate consortium partners. The ClearClimate consortium is committed to ensuring that the outcomes derived from research endeavours undertaken within the ClearClimate project are readily accessible to both the scientific community and the broader public. It is our aim to facilitate the replication of findings by scientific peers and enable the reuse and reanalysis of data compiled by the consortium's partners throughout the ClearClimate studies. To realize this objective, comprehensive datasets along with pertinent metadata will be rendered accessible and easily searchable. Furthermore, the analytical tools and scripts utilized will be made accessible for utilization. The present DMP describes the data management life cycle for all data collected and processed during the studies conducted within the ClearClimate consortium. The DMP covers the following topics: ● Description of types and formats of data ● Datasets administration: database(s) planning, analysis, design, implementation, and maintenance during and after the project lifetime. ● Data curation: organization, integration, and annotation method. ● Data accessibility and sharing policy and protections. ClearClimate adopts the statistical definition of data as the plural of datum, which is the “fact (numerical or otherwise) from which a conclusion may be drawn.” 1 Typically such data consists of values or character strings indicating e.g. answers from study participants, with aligned meta-data containing variables and values, labels and annotations, types of variables, measurement level, etc. The metadata usually contains information about the author and data collection process, including the time and place of data collection or production. 1 Kotz, S., Balakrishnan, N., Read, C. B., Vidakovic, B., & Johnson, N. L. (2005). Encyclopedia of Statistical Sciences, Volume 1. John Wiley & Sons, page 1531. D1.1 Data Management Plan 2 This DMP document is the first version that will be updated in the 24th and 40th month of the project lifespan to reflect the status of data collection and analysis and to update rules for data sharing within the ClearClimate consortium. D1.1 Data Management Plan 2 1. DATA SUMMARY 1.1 Purpose As our planet grapples with the escalating impacts of climate change, the need for effective Climate Information Services (CIS) has never been more urgent. From the complexities of climate extremes to the challenges of data visualisation and decision sciences, navigating the climate landscape demands innovative solutions. The ClearClimate project, assembling an international, interdisciplinary and intersectoral network of scholars with expertise in usercentred design, behavioural science, and cutting-edge computational techniques, aims to revolutionise CIS development by collecting new (e.g., WP2 and WP5) and reusing existing data (WP3, WP4, and WP5). 1.2 Data types and formats This DMP recommends that all the datasets collected/produced in the ClearClimate project will utilize the Hierarchical Data Format (HDF5) for storing the ClearClimate data in the repository. HDF5 stands as a high-performance data software library and file format designed to manage, process, and store diverse scientific data efficiently. Offering support for n-dimensional datasets, the HDF5 format imposes no restrictions on the number or size of datasets within a collection, thus providing significant flexibility conducive to interdisciplinary research, such as that envisaged within the ClearClimate project. Notably, the HDF5 format enables the incorporation of metadata and annotations alongside datasets within a single file, fostering transparent and straightforward data utilization and reuse for all stakeholders, including researchers beyond the ClearClimate consortium. Additionally, the format's versatility is underscored by its compatibility with all major operating systems and provision of highlevel APIs for widely-used computational languages in data science, such as C, C++, Python, Java, and R (a statistical computing language). Consequently, all data amassed within the ClearClimate consortium will be formatted into HDF5 prior to storage in the data repository, regardless of their original format upon collection. The original data collected during studies conducted within ClearClimate project shall follow the following recommendations for formats (Table 1): D1.1 Data Management Plan 3 Table 1: Data Types Data source Raw data format Geospatial data CFand CMIP6-compliant NetCDF Questionnaires quantitative data csv, tsv, hdf5 Neuropsychological measurements, e.g. EEG, eye tracking csv, tsv, hdf5 Behavioural (time reaction) measures csv, tsv, hdf5 Textual data odt, rtf, csv Image data png, jpg, svg Audio data FLAC, mp3 Video data MPEG-4 Storyline Definitions JSON Documentation ODT, PDF Papers & Articles LaTeX, ODT, PDF The aforementioned recommendations stem from the list of official recommendations for datasets as set by the UK data service 2 and recommended for EU projects. The enumeration of permissible data types is deliberately concise, as the inclusion of additional types would necessitate a substantial allocation of resources towards software development to ensure compatibility, preferably prioritizing conversion to the HDF5 format. Newly generated datasets stemming from the ClearClimate project will be appended to the ClearClimate Zenodo community (https://zenodo.org/communities/clearclimatemsca/) if their licensing permits. Whenever feasible, these newly published datasets will be assigned a Digital Object Identifier (DOI) to enhance their discoverability. In instances where adherence to these recommendations proves unfeasible, the partner responsible for dataset addition will provide a rational explanation. Such justifications may encompass constraints related to the inability to alter data from its original format due to licensing restrictions or necessitated usage of software incompatible with the recommended format. However, it's important to note that the necessity of additional effort to format alteration does not constitute a valid justification for straying from these recommendations. 1.3 Re-use of existing data Prior to initiating new data collection efforts, each consortium partner will thoroughly assess the availability of pre-existing datasets that could be repurposed to fulfill the objectives 2 https://www.ukdataservice.ac.uk/manage-data/format/recommended-formats D1.1 Data Management Plan 4 outlined within the ClearClimate research and associated activities. This proactive approach aligns with the overarching principles of data reduction and economy, aiming to minimize redundancy and maximize resource utilization within the consortium's endeavors (principles of data reduction and data economy). 1.4 Origin of the existing data WP2 and WP5 will utilise the existing datasets which originate from the existing ultimately international and national surveys that collect information on attitudes and behaviours related to climate change perception, e.g., International Social Survey Programme (https://issp.org/). In WP3 and WP4 might utilise the datasets collected from the national weather services and climate modelling centres, which provide publicly their datasets to the climate researchers. For example, in WP4, the re-use of existing datasets will help to identify the case studies based on local knowledge, identified users, historical data and high-resolution climate data sets from reliable and maintained climate data servers. 1.5 Generation of new data ClearClimate consortium will obtain new data from internal surveys (interviews, questionnaires, experimental neuropsychological studies – Tasks in WP2, WP4, WP5 and WP7). For example, It is expected that new data will need to be collected to test specific hypotheses that have not been studied before, e.g., eye-tracking data for testing climate change communication (WP5). 1.6 Expected size of datasets The size of data sets will vary from tens-hundreds MB for textual data (interviews) to several GB (climate-related variables, or neuropsychological data, e.g., eye tracking or EEG), even 1TB if uncompressed data is used. 1.7 Data Usefulness The ClearClimate data (existing, new, and metadata) are useful for other researchers by offering: 1) reviews on existing forecasting technologies to identify their limitations and opportunities for improvement (WP3 and WP4) and existing usage of AI technologies in the context of climate change prediction (WP3) 2) new data testing new forecasting technologies that address the limitations. The usercentred approach proposed by the ClearClimate consortium (e.g., WP2 and WP5) can D1.1 Data Management Plan 11 3. ALLOCATION OF RESOURCES ClearClimate Data Management Committee (DMC), composed by two representatives from academia (Paparrizos - WU-DES, Krejtz - SWPS) and two representatives from the private sector (Jelena Mazaj - CESIE and Daniel San Martn Segura - Predictia) prepared the DMP (D1.1), submitted DMP to the Steering Board (STB) for approval, and DMC will monitor DMP adherence. DMP will protect results that are critical to partners' future economic potential. Financial and human labour costs for ensuring datasets and other research outputs align with FAIR standards are estimated to be relatively low. All costs will be minimized by relying on established infrastructure for data storage, archiving, re-use security, etc. at each consortium partner’s institution. The costs will be also minimized by relying on the expertise of involved researchers and staff from the consortium partners. For any further expenses, ClearClimate funds can be used by the partners to cover costs associated with Open Access publications. Consortium partners will utilize Open Access in peer-reviewed scientific publications through free options: deposition of the accepted manuscript in institutional public repositories also referred to as Green Open Access. ClearClimate project will opt for Green Open Access when available leading to timely and wide dissemination benefiting the community. The responsibilities for implementing the FAIR principles within ClearClimate refer to: a) Metadata production and datasets curation: partners conducting study (data collection) b) Ethical clearance: partners conducting particular study or data collection Ethical Board or Institutional Research Board (if available) with the ClearClimate project Ethics Committee supporting this process c) Deposition in appropriate repositories: partners conducting study (data collection) d) Enabling Open Access: partners conducting study or data collection The compliance with defined data management principles is monitored by the project Data Management Committee and reported to the ClearClimate Steering Board. D1.1 Data Management Plan 12 4. DATA SECURITY Zenodo will be used for the collection and sharing of all final datasets and related documents in the ClearClimate project. Initial raw data from studies conducted within ClearClimate project will be stored and managed on servers housed at the ClearClimate project partners’ institutions and overseen by the Data Protection Officer (DPOs) of each institution (if available). These DPOs will uphold FAIR principles and ensure compliance with relevant data protection regulations, including Regulation (EU) 2016/679 (General Data Protection Regulation). Data gathered from studies involving any psychophysiological measures, participants’ attitudes. and opinions will pseudo anonymized (see Section 2.2). The secure storage of data aligns with the data protection policies of each partner institution. Informed consent for data sharing and long-term preservation will be obtained from all studies’ participants before data collection, encompassing various types of data, including surveys, psychophysiological and behavioural data. D1.1 Data Management Plan 13 5. ETHICS ClearClimate will fully adhere to Ethics by Design and Ethics of Use Approaches for Artificial Intelligence. The Project Ethical Committee (PEB) was established. PEB is part of the ClearClimate management structure. PEB will be acting as advisors and will issue ethical recommendations for partners along the life of the project. PEB will monitor the ClearClimate compliances relevant, official EU documents as part of WP1, Task 1.7 and described in D1.4. Every study within ClearClimate project involving human participants will follow the principles outlined in the Declaration of Helsinki and the steps described in D1.4. If available, each consortium partner is responsible for obtaining its institutional Ethical Board or Research Board clearance for research actions within ClearClimate project. Then Ethical Boards/ Research Boards of consortium partners (e.g., Ethical Board of SWPS University) will assess the methodology of proposed studies in the ClearClimate project, including the study protocol, consent, and withdrawal procedures set up for study participants. In case of any further ethical issues the appointed Project Ethical Committee will be consulted . Concerning the participants’ rights including anonymity of the gathered data, in the initial collection phase, in the analysis and technical development phase, and finally in the subsequent testing phase of new solutions, compliance with the GDPR and ethical principles will always be guaranteed. Before participation in each study, the study information and consent forms will be produced in English, and then translated into the other languages of the project if a study will be replicated in the consortium or other countries.