Data Management Plan for LangInLife project
Abstract
Data management plan created with the use of Data Stewardship Wizard for the project A lifetime with language: the nature and ontogeny of linguistic communication (LangInLife). Changes to this version: Detailed information on potential data reuse, resources allocation, and data security added to the corresponding sections.
Full text
Data Management Plan created using Data Stewardship Wizard. Corresponds with Horizon Europe DMP template. A lifetime with language: the nature and ontogeny of linguistic communication (LangInLife) Data Management Plan VERSION 1.1 02 Oct 2025
OP JAK LangInLife 2 / 15 History of changes Version Publication date Changes Version 1.0 02 Jul 2025 - Version 1.1 02 Oct 2025 Detailed information on potential data reuse, resources allocation, and data security added to the corresponding sections.
OP JAK LangInLife 3 / 15 Contributors The following contributors are related to the project of this DMP: • doc. Mgr. Pavel Caha, Ph.D. [email protected], ORCID: 0000-0003-1428-8052 Roles: Data Collector, Data Curator, Project Leader, Researcher Affiliation: Masaryk University (MU) • Mgr. et Mgr. Ľubomíra Nováková, Ph.D. [email protected], ORCID: 0000-0002-8352-439X Roles: Data Collector, Data Curator, Researcher Affiliation: Central European Institute of Technology – Masaryk University (CEITEC MU) • doc. Mgr. Radek Šimík, Ph.D. [email protected], ORCID: 0000-0002-4736-195X Roles: Data Collector, Data Curator, Researcher Affiliation: Charles University • prof. PhDr. Filip Smolík, Ph.D., DSc. [email protected], ORCID: 0000-0003-4160-6124 Roles: Data Collector, Data Curator, Researcher Affiliation: Czech Academy of Sciences (AV ČR) • Mgr. Pavla Martinková [email protected], ORCID: 0000-0002-1456-4224 Roles: Contact Person, Creator of DMP, Data Steward Affiliation: Masaryk University (MU)
OP JAK LangInLife 4 / 15 Project We will be working on the following project for which the data and work are described in this DMP. A lifetime with language: the nature and ontogeny of linguistic communication, CZ.02.01.01/00/23_025/0008726 Acronym LangInLife Start date 2025-01-01 End date 2028-12-31 Funding • https://msmt.gov.cz: CZ.02.01.01/00/23_025/0008726 (granted) The project focuses on the language of the individual, which is an irreplaceable means of communication with the surroundings of every person. It explores how language is acquired in childhood, how it functions in adulthood and what challenges individuals face as they age. The aim of the project is to conduct interdisciplinary research based on a combination of linguistics, psychology and neuroscience, which, by producing cutting-edge results, will enable an adequate and effective response to the global language-related challenges of contemporary society. The project has three main research objectives, carried out by research teams from participating institutes. Research objectives are referred to as RO1, RO2 and RO3 in this DMP. • Research objective 1: Language in childhood: acquisition and early language development (RO1) • Research objective 2: Language in adulthood: nature, communication, language learning and the consequences of migration (RO2) • Research objective 3: Language in old age: loss of language abilities and how to slow it down (RO3)
OP JAK LangInLife 5 / 15 1. Data Summary Different types of data will be created, acquired, and re-used during the project, as detailed in this section. Data related to scientific publications will be made available, in accordance with Open Science principles and following legal requirements, as described in Section 2. Individual datasets will be specified as updates to this DMP. The project will generate a variety of data surrounding nature and ontogeny of linguistic communication. The generated data will be both quantitative (e.g., reaction times) and qualitative (e.g., annotated spoken and written corpora). All data will relate to research questions for each of the research objectives as outlined in the project proposal. The data will encompass different methodological approaches. The expected size of the data differs for each of the research objectives; the expected increase in a year is hundreds of GiB for RO1, under 10 GiB for RO2, and 550 GiB for RO3. The data obtained will be useful not only for the project's research team, but also for the wider community of experts in the fields of linguistics, psychology, neuroscience, pedagogy, and clinical practice. The data can serve as reference material for further research into language development, bilingualism, the impact of aging on language abilities, and the impact of migration on language acquisition. • Equipment data Data will be collected by project members, with our own equipment, which is well described and known. The following types of equipment will be used: • Eye-tracking • Electroencephalogram (EEG) • Magnetic resonance imaging (MRI) • Functional near-infrared spectroscopy (fNIRS) Data will be handled in the following formats: BVCDF, CSV, EDF, FNIRS, NIFTI, and other formats depending on the used equipment and software. • Experimental stimuli We will be using audio, video and image data as stimuli in experiments. We will be working with the following formats: DOCX, JPG, MP3, PNG, TXT, WAV, and other formats depending on the software used.
OP JAK LangInLife 6 / 15 Re-used data We will be re-using the following data: • Openly available linguistic corpora (e.g. Czech national corpus, Clearpond, English Vocabulary Profile) The corpora will be used to study linguistic phenomena related to research objectives. The data can be used in the format provided without any conversion needed. We will work with them using online analytical tools or download them for local analysis. Potential changes in the data will not influence the reproducibility of our results. For some corpora, only parts of them will be used; any filtering or selection will be documented. • Copyrighted text corpora (e.g. Brepolis database, printed books) The corpora will be used to study linguistic phenomena related to research objectives. Texts needed for analysis are available via subscription to the database or purchase of the books (not part of the project costs). There will not be any changes in the data influencing our results. Only parts of the texts will be used, and the selection will be well documented. • Datasets previously created by project members Datasets will be used to study linguistic phenomena related to research objectives. Project members are the owners of the datasets. The datasets can be used in the format provided without any conversion needed. We already have a copy of the data. Those are fixed datasets, changes will not influence the reproducibility of our results. For some, only parts of the dataset will be used; any filtering or selection will be well documented. 1.1 Data formats Information about each of data formats we currently plan to be working with. • BrainVision Core Data Format (BVCDF) o An open proprietary format, compliant with BIDS (Brain Image Data Structure) standard. • Comma-separated Values (CSV) o An open file format, suitable for long-term preservation of tabular data. • Office Open XML Document (DOCX) o A widely used format, following the Office Open XML standard. • European Data Format (EDF) o An open format, suitable for exchange and storage of multichannel biological and physical signals. • JPEG, JPG o A format suitable for long-term preservation of raster images. • JavaScript Object Notation (JSON) o An open file and data interchange format.
OP JAK LangInLife 7 / 15 • Markdown (MD) o An open file format. • MP3 o An open audio file format. • Neuroimaging Informatics Technology Initiative (NIfTI) o An open file format, used to store brain imaging data obtained using MRI methods. • Portable Network Graphics (PNG) o A format suitable for long-term preservation of raster images. • R (R) o A format for programming files used in open-source software for statistical computing and graphics, R. • R Markdown (RMD) o Format of files used in open-source software for statistical computing and graphics, R, using Markdown markup language. • Share Near InfraRed File (SNIRF) o An open file format, with specifications compliant with BIDS (Brain Image Data Structure) standard. • Text File (TXT) o A format suitable for long-term preservation of textual data. • Waveform Audio Format (WAV) o A proprietary audio file format. • Office Open XML Workbook (XLSX) o A widely used format, following the Office Open XML standard.
OP JAK LangInLife 8 / 15 2. FAIR Data 2.1. Making data findable, including provisions for metadata We will share data that can become open in publicly accessible trusted repositories using descriptive metadata as required by the repository. Keywords will be added for better findability. Metadata will be available in a form that can be harvested and indexed. All published data will be assigned a persistent identifier which will be included in their metadata. This will be managed by the repositories. The following metadata standards are relevant for various data we will work with during the project: • Brain Imaging Data Structure (BIDS) • Component Metadata Specification (CMDI) • Investigation Description Format (IDF) • Linguistic Annotation Format (LAF) • Minimum Information about an fMRI Study (MIfMRI) • Open Language Archives Community Metadata (OLAC Metadata) 2.2. Making data accessible We will be working with philosophy as open as possible for our data. Data will be deposited by the publication date of the related scientific publication at the latest; embargo period is currently not planned to be used for any of the datasets. The following repositories are considered for sharing the data: • Zenodo o A trusted general open repository, operated by CERN. • Figshare o A trusted general repository, operated by Digital Science company. • Dataverse o A trusted general repository, operated by Harvard. Other trusted general or disciplinary repositories will be specified as updates to this DMP. Our data cannot become completely open. We will collect data connected to a person, i.e. personal data, as detailed in Section 6. We can use pseudonymization, anonymization and data aggregation to make the data more openly available. There are IP reasons why our data cannot be open, which applies to re-used copyrighted text corpora. For data with limited access, there will be clear instructions on how to get access in the metadata.
OP JAK LangInLife 9 / 15 Metadata will be made openly available under a public domain dedication CC0. This will be managed by the repositories. 2.3. Making data interoperable We will be sharing data in the following formats: • CSV, JSON, TXT More formats will be decided upon closer to the publication date and will be specified as updates to this DMP. For sharing data, we will use open formats as opposed to proprietary ones whenever possible. 2.4. Increase data re-use Rich metadata and documentation will be provided for each dataset; Readme file will be included. Files and folders will be versioned and structured using a file naming convention. All the metadata in the file names will also be available in the proper metadata. For experiments, we will use paper and electronic lab notebooks to make sure that there is good provenance of the data analysis. The following open-source research software will be used for experiments: • PCIbex, an open-source software for internet-based experiments • L-Rex, an open-source software for linguistic experiments • PsychoPY, an open-source behavioral research software It is clear who owns data and documents created during the project and can license them for re-use. We will be employing the following quality processes for instrument data: • Calibrating measurements • Repeat measurements • Standardized data capture and recording • Data Entry validation To validate the integrity of the results, the following will be done: • We will run a subset of our jobs several times across the different computing infrastructures. • We will be instrumenting the tools into pipelines and workflows using automated tools. • We will use independently developed duplicate tools or workflows for critical steps to reduce or eliminate human errors. • We will run part of the data set repeatedly to catch unexpected changes in results.