Research Data Management Strategy and Policy at the Institute for Combustion Technology
Abstract
This institute-wide research data management policy establishes binding baseline standards for the responsible handling of research data across all projects at the Institute for Combustion Technology. It provides staff with clarity and guidance throughout the data lifecycle and supports compliance with the FAIR principles. Where external or project‑specific requirements specify different or stricter measures, those requirements take precedence.
Full text
Research Data Management Strategy and Policy at the Institute for Combustion Technology Raymond Langer, Fabian Fr¨ode, Marianna Cafiero, Florence Cameron, Roman Glaznev, Dominik Golc, Maximillian Hellmuth, Christian Schwenzer, Raik Hesse, Terence Lehmann, Gandolfo Scialabba, Shreyans Sakhare, Hongchao Chu, Michael Gauding, Joachim Beeckmann, and Heinz Pitsch Institute for Combustion Technology, RWTH Aachen University Version: 1.0.0 November 28, 2025 Abstract This institute-wide Research Data Management (RDM) policy establishes binding baseline standards for the responsible handling of research data across all projects at the Institute for Combustion Technology. It provides staff with clarity and guidance throughout the data lifecycle and supports compliance with the FAIR principles. Where external or project-specific requirements specify different or stricter measures, those requirements take precedence. 1 Motivation The availability of published and unpublished data is essential for reproducing scientific results and opens up important opportunities for further research. Despite the crucial role that data availability plays in scientific work, it is not always ensured, even when data are used in peer-reviewed publications. Requesting data from 516 studies published between two and 22 years earlier, Vines et al. [18] found that only one in five datasets could be retrieved. Several other studies found that authors are often unable or unwilling to share their data [14, 21, 16, 19]. Long-term availability is a particular challenge here, as the availability of published research data can decline rapidly over time after publication [18]. In addition to the mere availability of research data, numerous other requirements must also be met. Wilkinson et al. [22] proposed guiding principles aiming to improve the infrastructure supporting the reuse of scholarly data. According For details on how to cite, see doi:10.5281/zenodo.17748592. 1
to these principles, research data should be Findable, Accessible, Interoperable, and Reusable (FAIR) [22]. The present policy aims to consolidate and implement the requirements of the German Research Foundation (DFG) [2, 3], RWTH Aachen University [15], and ITV’s internal objectives while considering the FAIR principles [22]. The policy is structured around a core guideline and three additional guidelines addressing research data management planning, software development, and storage and documentation. The centralized planning, administration, documentation, storage, and archiving of research data ensures long-term usability. The standardized yet flexible approach facilitates successful research data management while enabling evolving best practices. Group leaders and data stewards working across projects enforce and coordinate the implementation and further development of the strategy. The following sections discuss the last three individual guidelines and the most important goals pursued by using them. Examples of important RDMrelated events are illustrated as a timeline in Figure 1. Table 1 summarizes the roles and their current representatives within the ITV data management policy. Core Guideline Research data management (RDM) concerns everyone at ITV. All personnel are expected to familiarize themselves with this policy and to implement the procedures explained according to their role. We encourage active participation in the periodic review and enhancement of this policy to ensure its relevance and effectiveness. Guideline 1 A data management plan maintained in the Research Data Management Organizer (RDMO) is required for each project. Guideline 2 Software developed at ITV and other text-based data that comply with the license conditions and project-dependent confidentiality requirements must be managed in a git repository in the RWTH GitLab instance. Guideline 3 All research data, including L A T EX source code and figures of publications, must be stored or referenced and documented in the Collaborative Scientific Integration Environment (Coscine). 2
Table 1: Roles and their current representative within the ITV data management strategy. Role Current representative Employee All scientific staff Group leaders Joachim Beeckmann, Hongchao Chu, Fabian Fr¨ode, Michael Gauding, Raymond Langer Director Heinz Pitsch Start DMP in RDMO Start COSCINE employee project Perform measurements (Store raw data on an ITVserver) Upload measurement data on COSCINE Develop a chemical kinetic model (use RWTHGitLab) Start kinetic model validation database (use RWTHGitLab) Biannual status meeting (RDMO and COSCINE must be updated) Submit a peer-reviewed paper (store submission with LaTeX source code and figures on COSCINE) Publish peerreviewed paper (store final version and rebuttal with LaTeX source code and figures on COSCINE) Figure 1: Timeline with examples of RDM events for scientific staff at ITV. 2 Research Data Management Plans A data management plan (DMP) describes the life cycle of research data from collection to archiving, including all measures to ensure that the data remains available, usable, and understandable. Many funding agencies, especially governmental sources, require a DMP. Therefore, all scientific staff must update their existing DMPs at least every six months and add a DMP for each new project within the first three months. At the ITV, the RDMO tool is used to create DMPs. The group leader and research staff involved in the project are added as owners of a project. For already existing DMPs, the group leader is responsible for adding relevant new employees. The status of each DMP is tracked within the employees’ status sheet and discussed during the biannual staff strategy meetings. Choose the ITV catalog and the parent project “ITV general” in RDMO unless you are working in a large-scale project that also requires you to maintain a DMP in RDMO. In that case, the project-specific requirements replace the ITV guidelines and the project-specific catalog can be 3
used. 3 Software Development and GitLab usage at ITV Simulations of reactive flows and the evaluation of experiments are part of ITV’s daily work and require complex software. Software developed at ITV and other text-based data that meets the license conditions and (project-dependent) confidentiality requirements must be managed in a git repository in the RWTH GitLab instance. The following GitLab groups have been established to keep project data organized and maintain access control. When creating a GitLab repository, please use one of these groups where applicable: RWTH GitLab/ ITV Code ITV Teaching ITV Kinetics ITV Code: Use this group for software development projects that are broadly relevant to ITV research. This includes any kind of code or scripts used in ongoing research activities at the institute. ITV Teaching: Reserve this group for data, scripts, or code used solely in courses. Avoid using it for student theses or other individual research projects. ITV Kinetics: Place all chemical kinetic models (regardless of detail level) in this group, along with relevant validation databases. If none of the above groups applies, please use your personal GitLab space. The FlameMaster repository demonstrates best practices for collaborative development, continuous integration, and version control in ITV’s GitLab environment. Depending on the size and complexity of your project, these best practices can be adapted to suit your needs. FlameMaster is a publicly available C++ program package for zero-dimensional combustion and one-dimensional laminar flame calculations, which has been maintained using GitLab since 2015. Currently, the repository includes 86 members associated with the ITV (primarily undergraduate and graduate students below referred to as ITV developers) who use the repository jointly to perform various research tasks, such as the analysis of experimental measurements and the development of detailed chemical kinetic models. An essential building block for the long-term success of source code development is the implementation of Continuous Integration (CI), i.e., the practice of frequently building and testing a software system during its development. The implementation for FlameMaster relies on the RWTH GitLab instance focusing on the master branch of the repository, where the primary goals are accessibility and interoperability of the source code and reproducibility of simulation results for all FlameMaster developers and users, including numerous exter4
nal researchers and non-academic users. Each commit that an ITV developer pushes to the remote repository hosted in the RWTH GitLab instance triggers a six-stage CI pipeline. Starting from docker images of recent, minimal Linux operating systems, a GitLab runner performs the following steps automatically: 1. Installation of dependencies 2. Configuration and compilation of the software package 3. Execution of unit tests 4. Execution of preprocessing steps for simulations 5. Execution of simulations 6. Public release of the source code on the ITV website. The success of each step is automatically checked, and a subsequent step is only initiated if the previous one was successful. The ITV developer who pushed the last commit to the repository receives a notification if a step fails. The combined use of CI and GitLab in a large group of developers has several advantages. Executing the CI pipeline after every commit ensures that issues and bugs are detected and fixed promptly. Performing all tests in docker images of recent, minimal Linux operating systems improves the reproducibility of simulation results, a crucial aspect in the field of scientific computing. The automatic release of thoroughly tested software enables external users to access improved software easily, thereby promoting the widespread use of the most recent version of FlameMaster. Moreover, the joint use of an internal git repository by the ITV developers fosters collaboration and long-term improvements of the software, including significant advancements of the source code that the researchers cannot share before peer review in a journal publication. 4 Finding, Storing, and Documenting Research Data in Coscine Coscine is a central tool for implementing the ITV data management policy. It is ITV’s default application for documenting and storing scientific data, such as journal articles, collected validation datasets, experimental measurement data, simulation inputs and results, and chemical kinetic models. Decisions that deviate from this best practice must be discussed with the group leaders. The Coscine team also maintains a news website and a mailing list, which you can subscribe to for the latest updates. 4.1 Fundamentals about Coscine Coscine uses a hierarchical (tree-like) model for all data. Every tree comprises two types of nodes: “Project” nodes and “resource” nodes. Project nodes are 5
used within a chosen tree to explain the structure of data. Resource nodes are leaves of the tree, i.e., they have no children and are used for storing data. Hence, a resource node corresponds to uploaded data, while (sub-)project nodes are used for documentation purposes. Note that entire folder structures can be uploaded into a resource. Different resource types are available. A list of all resource types available and their specification can be found on the corresponding help page of Coscine. Coscine manages permissions using three access models, referred to as owner, member, and guest. The owner of a project node has read and write access and can add new project and resource nodes. In addition, the owner can manage access permissions and project node settings. Guests have only read access and can view and download data. Note that access permissions are not inherited from the parent project node and can/need to be individually specified for each project node. A detailed description of managing access permission of projects can be found on the corresponding help page of Coscine The hierarchical, tree-like structure of Coscine at ITV is illustrated in Figure 2. The baseline structure consists of a central ITV project node at the top of the hierarchy. A distinction between employees and test benches is introduced on the second level, providing sub-project nodes for each employee and test benches separately. While research data management is generally employeebased at ITV, this distinction is introduced to maintain consistency with the well-established server structures for experimental data. It is important to note that only experimental raw data are managed in the resource node of the test bench project nodes. Processed experimental data used, for example, in a publication, are managed within the corresponding employee project node. The employee sub-project node will be created by the group leader during the first days at ITV, and the node will remain available in Coscine after the employee leaves the ITV. Each employee is responsible for managing the subproject node and establishing a meaningful substructure for their data. Keep in mind that managing the nodes and moving project and resource nodes is not straightforward in Coscine and requires the creation of new nodes and reuploading the data. Moreover, it follows from that limitation that the baseline structure is fixed and will not be updated, resulting in a constantly growing tree. The accessibility of data is ensured by granting the corresponding group leader and director ownership rights for all project nodes. This procedure is detailed in Section 4.2. Note that the tree structure shown for Employee 1 and the Jet Burner test bench are examples. The examples in Section 4.3 illustrate how such substructures can be developed depending on the field. 4.2 Workflow RDM with Coscine generally comprises the following steps: 1. Data generation, structuring, and documenting 2. Creation of a project node 6
Figure 2: Schematic representation of the Coscine structure for ITV. 3. Application for resources 4. Creation of resource node(s) 5. Data upload 6. Archiving Data generation, structuring, and documenting: The workflow starts before the actual usage of Coscine during data generation by establishing a meaningful data structure and providing documentation. Note that data generation includes raw data generation (e.g., measured signals, simulation output, etc.) and processed data generation, and a distinction between both is highly recommended. A meaningful structure improves data usability and simplifies the subsequent upload process to Coscine. The examples in Section 4.3 illustrate how data can be structured and documented before uploading to Coscine. All uploaded data must be complete and reproducible by the provided documentation. Providing good README files can be a powerful tool to document data reproduction, particularly of processed data. README files can be used to document the overall project or subordinate data generation steps, such as creating a graph in a publication. A “good” README file is provided in an open file format and is processable by any text editor, e.g., MD (preferred])1or TXT. It is structured, including 1. a short description of the goal and main features to identify what problem the project solves, 1The Markdown format provides rendering on GitLab and light formatting, e.g., headlines, tables, links, etc. 7
2. step-by-step instructions on setting up the environment (system requirements, installation, configuration, testing the environment), 3. basic usage commands and examples representing typical usage scenarios, 4. explanation of folder and file structure with a function description. An exemplary README file is provided in the FlameMaster project or the C3MechV4.0 repository. Open file formats, such as CSV or YAML/JSON text files or HDF5, are preferred over proprietary formats. For example, DOCX and XLSX, are not open and should be converted to an open format before uploading to Coscine. Creation of a project node: Once the data has been generated, structured, and documented, a project node must be created on Coscine if not yet available. Coscine will ask for some metadata. The recommended visibility is “Project Members” for ITV project nodes. As mentioned above, access rights are not inherited from the parent project node. To ensure that the corresponding group leader and director have access to the data, they have to be added as owners. Details about how to add users can be found on the corresponding help page of Coscine. Application for resources: Storage requirements are known once the data has been generated, structured, and documented. Coscine features different resource types, which correspond to different storage options. A detailed description of the available resource types and a decision flow chart is provided on the corresponding help page of Coscine. An application is required for S3 resources and web storage of more than 100 GB (initially this is 25 GB, but the limit can be increased in the project settings). The links to the application forms can be found here. These links open up to a JARDS platform where proposals for storage space can be submitted (more information can be found in Appendix A). A Principal Investigator (PI) must be specified for each application. Joachim Beeckmann and Michael Gauding must always be the PI for experimental and simulation resources, respectively. Creation of resource nodes(s): After the project node creation and resource approval, resource nodes need to be created in Coscine according to the local structure of the data. Step-by-step instructions on creating resource nodes is provided on the corresponding help page of Coscine. The key steps are the resource type selection, Metadata profile selection, and resource metadata. The selection of the resource type follows the guidance provided on the corresponding help page of Coscine. Metadata profiles define a collection of metadata used to describe the data and will be separately presented in Section 4.2.1. In the final step, predefined metadata must be provided, such as the resource name or keywords. In addition, the visibility and license of the resource are asked about. The recommended visibility is “Project Members” and the Creative Commons license CC BY 4.0 is recommended for published data. 8
Data upload: How the data can be uploaded depends on the resource type. Often ITV’s resources are S3 resources, requiring clients to upload data. Different clients and their usage are described on the corresponding help page of Coscine. In addition, it is worth mentioning that Coscine features an Application Programming Interface for automating workflows. Archiving Once the data reaches the end of its lifecycle, it should be archived. Technically, archiving means that the access permissions are set to read-only. In addition, archiving ensures that the data are stored for at least ten years in compliance with good scientific practices. An instruction on how to archive resources can be found on the corresponding help page of Coscine. The following control mechanisms are implemented at ITV to ensure a seamless workflow from data generation to archiving, as outlined in this section. The control mechanisms are based on publications because of their exceptional importance in science. After a publication is accepted (or the conference has taken place), a colleague reviews the uploaded data for completeness. All data required to reproduce figures in the paper must be included in an open file format. A typical strategy is to select a random graph in the publication and try to reproduce it with the provided documentation (e.g., README files) and codes (either provided within the publication or accessible via GitLab). If this check is passed, the corresponding publication will be marked as complete in the employee’s status sheet. The status of each publication is reviewed at the biannual staff strategy meeting. To cover also non-published data, the group leader reviews project nodes for completeness before the employee leaves the ITV. This is part of the standard protocol and is noted in the checklist. 4.2.1 Metadata profiles Metadata is additional data stored with the actual data to improve findability and usability. A Metadata profile defines a collection of metadata for a particular application and must be selected for every resource node added to a Coscine project node. At ITV, it is expected that all uploaded data are complete and reproducible without the metadata provided by the Metadata profile. Hence, the metadata provided by the Metadata profile is mainly intended at ITV to improve findability through Coscine’s search engine. Metadata profiles specific to the requirements of data from ITV were developed. The Metadata profile ITVGeneral defines the baseline entries for all profiles and can be used if none of the application-specific profiles apply. Table 2 shows the required metadata and a description. In addition, Metadata profiles for different research fields at ITV exist: •ITVCFD (Table 3) defines mandatory metadata for CFD (RANS, LES, DNS) simulation data. 9
Table 7: Metadata profile: ITVKineticModel. Field Name Description Git clone URL kinetic model Provide the Git clone URL for the kinetic model repository. For example, [email protected]hen.de:ITV/mech.git Subdirectory Provide the subdirectory in the git repository that contains the Chemkin files and species dictionary, e.g., SootMechanisms/ITV 2015/ Langer2021/kritika liming/soot langer Git hash kinetic model Provide the git hash of the kinetic model. Git clone URL validation database Provide the Git clone URL for the validation database repository. For example, [email protected]hen.de:ITV/mech.git. Git hash validation database Provide the git hash of the commit used for the validation of the kinetic model. Fuels State their common name(s) and the name used in the model. Use a separate field for each fuel (Plus-Button). Validation targets Provide a summary of the validation targets considered. Use a separate field for each validation target (Plus-Button). Possible values are ignition delay time data from shock tubes (IDT(ST)), IDTs from rapid compression machines (IDT(RCM)), laminar burning velocities (LBV), extinction strain rates (ESR), species data from counterflow flames (CF), species data from jet-stirred reactors (JSR), or species data from flow reactors (FR). 16
Table 8: Metadata profile: ITVKineticModelValid. Field Name Description Git clone URL validation database Provide the Git clone URL for the validation data repository. The repository can be either a standalone repository for the kinetic model or a combined repository containing both the kinetic model and the validation data. An example of a combined repository is [email protected]hen.de:ITV/mech.git Subdirectory Provide the subdirectory in the git repository that contains the validation database, e.g, GasolineSootUpdate/validation cases/ Git hash validation database The git hash or tag of the version that is documented in Coscine. For example, this can be the git hash of the validation database as used for the release/publication of a kinetic model. Another status worth documenting would be a major change in the format of the data or the addition of new data. Fuels State their common name(s) and the name used in the model. Use a separate field for each fuel (Plus-Button). Validation targets Provide a summary of the validation targets considered. Use a separate field for each validation target (Plus-Button). Possible values are ignition delay time data from shock tubes (IDT(ST)), IDTs from rapid compression machines (IDT(RCM)), laminar burning velocities (LBV), extinction strain rates (ESR), species data from counterflow flames (CF), species data from jet-stirred reactors (JSR), or species data from flow reactors (FR). 17
Table 9: Metadata profile: ITVLaserDiag. Field Name Description Experimental setup Specify the used burner type. For example, counterflow burner, round jet burner, slot burner, industrial burner, etc. If several burners were used, use a separate field for each burner (Plus-Button). Conditions Provide the relevant conditions of the flame and the measurement location. For example, mixing length, equivalence ratio, oxidizer composition, fuel composition, Re/velocity range, pilot/co-flow conditions, height above nozzle, etc. Measurement technique Provide the measurement techniques used. For example, LIF, LIPF, LIBS, Raman Scattering, Rayleigh Scattering, LII, PIV, etc. Use a separate field for each technique (Plus-Button). Acquisition settings Provide the settings of the measurement system. For example, frequency, gate timing, online processing (binning, accumulation), background correction, etc. 18
4.3 Examples This section presents six best-practice examples, one for each field. The examples are categorized as a data set or publication, demonstrating the separation of raw data generation from processed data. The data set examples demonstrate how to store and document simulation or experimental data, while the publication examples guide the preparation of a publication. 4.3.1 Example 1: Large Eddy Simulation of Nanoparticle Formation in the SpraySyn Flame (Data set) The following example covers the Large Eddy Simulation (LES) data used for the publication of Fr¨ode et al. [4]. Four simulations were performed, covering one simulation without precursor (undoped) and three with precursor (doped). Within the doped cases, three different precursor consumption models were tested assuming instantaneous inception (InstInc), finite rate inception (FiniteInc), and finite rate inception plus adsorption (FiniteIncAds). The data are structured in the following directories: / chemtable generation BSF for pilot createChemtable flamelets mechanisms C2H5OH C2H5OHIronWlokas les grid input runs undoped doped InstInc FiniteInc FiniteIncAds The data set splits up on the top level into the chemtable generation and the actual LES data. The chemtable generation covers three steps: the BurnerStabilized Flame (BSF) computation to estimate the pilot flame, the generation of flamelet solutions, and the creation of the final chemtable. A directory exists for each of the steps. In addition, the kinetic models used are provided in the directory flamelets. A README file is provided in every directory, describing the usage, meaning, and important references (e.g., the reference to the kinetic model) of the files contained in it. For example, the directory createChemtable contains the following files: 19
createChemtable convertTable.input createChemtable.input getNames.sh README.txt The directory les covers the grid generation (grid), required input files (input), and the simulation output (runs). The simulation results are organized according to the structure used in the paper [4]. Again, every directory contains a README file providing documentation. For example, the directory InstInc contains the following file: InstInc data.start data.out part.start part.out spraySyn flame.input links.sh RunInfo.txt README.txt The data set is uploaded into one resource node using the application profile ITVCFD shown in Table 10. The resource node could be placed in the tree under Data from ITV →Employees →Fabian Fr¨ode →Simulation Data → SpraySyn Burner 20
Table 10: Example application of ITVCFD within Example 1 Field Name Response Title Large eddy simulation of nanoparticle formation in the SpraySyn flame: Impact of gas-phase chemistry Contact person Fabian Fr¨ode Contact person ORCID 0009-0009-2215-2762 Contact person email f.fro[email protected]hen.de Producer / Author Fabian Fr¨ode Keywords / Tags LES, nanoparticle formation, Method of Moments, FPVA, HMOM, precursor chemistry, adsorption, SpaySyn, spray flame synthesis Legal information none Date creation 04-2023 Funding SPP1980 Status publication published Publication title Large eddy simulation of iron oxide formation in a laboratory spray flame Publication journal/conference Applications in Energy and Combustion Science Publication DOI 10.1016/j.jaecs.2023.100191 Comments / Explanation Relevance of precursor chemistry and adsorption on the nanoparticle formation in the SpraySyn flame. Software CIAO Git clone URL https://git.rwth-aachen.de/fabian.froede/SPP 1980 Version number / Git hash fca834051cab10ea3fe2f46d2ec291df6520e3b5 Simulation configuration Piloted spray flame Simulation type 3D LES Quantities of interest Temperature, mixture fraction, droplet statistics, particle moments, particle diameters Solver specification CIAO Low-Mach Combustion model Three-stream flamelet progress variable approach (FPVA) Multiphysics Lagrangian spray, moment-based nanoparticle model HPC system CLAIX-2018 21
4.3.2 Example 2: The Role of C3 and C4 Species in Forming Naphthalene in Counterflow Diffusion Flames (Data set) The following example covers the gas chromatograph (GC) speciation data used for the publication of Hellmuth et al. [8]. Two flames were probed multiple times, and numerous calibration measurements were performed. The data are structured in the following directories: / 01 Measurement data YYYYMMDD <name of measurement>.D acqmeth.txt runstart.txt cnorm.ini DATA.MS FID1A.ch PRE POST.INI TCD2B.ch ... 02 GC methods 00 Sequences YYYYMMDD <name of sequence>.S ... YYYYMMDD <name of method> acqmeth.txt Audit.txt method.txt ... ... 03 GMS methods 01 Configuration file Default.mcf 02 Methods Online analysis Online measurement flame.mmf ... Sampling Offline Sampling Flame.mmf ... ... The directory contains three sub-directories: 01 Measurement data has a folder for each measurement, i.e., the analysis of one sample. A flame set comprises 16 measurements. A calibration set contains four or more measurements. Each measurement folder contains the recorded GC detector signals, e.g., DATA.MS for the mass-selective detector, and information on the method (acqmeth.txt) for generating the data. More on the acquisition methods of the 22
GC and the gas measurement system (GMS) is stored in 02 GC methods and 03 GMS methods, respectively. The data set could be placed in the tree under Data from ITV →Test Benches →Gas chromatograph →Data and directly linked to the publication. 23
Table 11: Example application of ITVSpeciation. Field Name Response Title The role of C3and C4species in forming naphthalene in counterflow diffusion flames Contact person Maximilian Hellmuth Contact person ORCID 0000-0003-2325-9798 Contact person email m.hellm[email protected]hen.de Producer / Author Maximilian Hellmuth Keywords / Tags Counterflow diffusion flame; Speciation measurement; Allene; 1-Buten-3-yne (C4H4, vinylacetylene); Naphthalene Legal information none Date of creation 08-2023 - 11-2023 Funding FSC Status publication published Publication title The role of C3and C4species in forming naphthalene in counterflow diffusion flames Publication journal / conference Proceedings of the Combustion Institute Publication DOI 10.1016/j.proci.2024.105620 Comment / Explanation Provide fresh insights into the role of C4chemistry regarding aromatics formation by introducing 1-buten-3-yne (C4H4) as fuel for the first time in counterflow diffusion flames. Experimental setup Oxyflame counterflow burner Measurement technique GC-MS and ToF-MS Fuels C3, C3+C4 Conditions Strain rate of 60 s−1 (momentum balance); C3 flame: fuel flow: 5 % allene, oxidizer flow: 20 % oxygen, diluent: argon; C3+C4 flame: fuel flow: 3.5 % allene & 1.2 % 1-buten-3-yne/vinylacetylene, oxidizer flow: 19.9 % oxygen, diluent: argon. 24
4.3.3 Example 3: LII measurements of soot formation characteristics in ethylene and acetylene counterflow diffusion flames blended with dimethyl carbonate and methyl formate (Data set) The following example covers the raw data from LII measurements in ethylene and acetylene-based counterflow diffusion flames blended with DMC and MeFo as used for the publication of Cameron et al. [1]. The data are uploaded in one resource node exemplary placed under Data from ITV →Test Benches→Jet Burner. The data are structured in the following directories: LII 01 Flame properties 02 Raw data YYYYMMDD <name of measurement> Settings.txt Data.spe Data.mat A README file containing information on the structure and contents of the directory is placed on the first level. Further documentation of the operating conditions is given in the directory 01 Flame properties. The directory 02 Raw data contains for every measurement day the recorded raw data. Besides the raw data files (Data.spe and Data.mat), each measurement directory contains a txt file with a detailed specification of the LII setup. The respective application profile response is given in Table 12. 25
4.3.6 A detailed chemical kinetic model for aromatics formation (Data set) Chemical kinetic models used in various ways to perform reacting flow simulations are among the most commonly shared data at ITV. Here, a detailed chemical kinetic model for aromatics formation is considered as an example to explain the application profile in Table 7. Langer et al. [13] developed this detailed model to analyze small hydrocarbon and gasoline surrogate fuel combustion, but the introduced application profile can also be applied to other kinetic models, including models for different fuels or reduced and optimized models. The de facto standard for providing chemical kinetic models is the CHEMKIN format [12, 10, 11]. Due to the wide availability of CHEMKIN format parsers, this format must always be used as the basis for kinetic model files, and additionally required data, such as data on covalent bonding, multiplicities, excitation states, or species lumping, must be provided in a backward compatible, machine-readable manner. For the detailed model of Langer et al. [13], the following files implement these requirements: soot langer/ ITV PAH.mech ITV PAH.thermo ITV PAH.trans species list default.json make species dict.py style files spec dict hyphenat.sty placeins.sty url.sty xurl.sty The first three files are the CHEMKIN files, which provide gas-phase kinetics, thermochemistry, and transport data. The reader is referred to the literature for a discussion of this format [12, 10, 11]. Langer et al. [13] amended the thermochemistry files by machine-readable comments that provide additional information on the molecular structure of species through International Chemical Identifiers (InChIs) [7, 6] and Simplified Molecular Input Line Entry System (SMILES) formulas [20]. Recently, Pratali Maffei et al. [5] extended this approach to a rigorously defined species model by introducing species definitions based on species multiplicities and excitation states in addition to molecular structure data. This species model is adopted for more recent versions of the Langer et al. [13]. Like Pratali Maffei et al. [5], Langer et al. [13] implemented a Python script, which can be executed to check the numerous constraints on the species data [5] or to generate documentation. The following commands can be executed to perform all checks and generate a documentation: $pip3 install rdkit CairoSVG gitpython --user $./make_species_dict.py 32
$cd SpeciesDict $pdflatex species_dict . tex The metadata for the kinetic model of Langer et al. [13] are summarized in Table 15. The resource node is be placed in the tree under Data from ITV → Employees →Raymond Langer →Kinetic models. Similar to a complex source code, chemical kinetic models are expected to be developed using a Git repository. Therefore, the resource type of the node is “GitLab.” The “Repository kinetic model” field provides the repository location, which can be used with the git clone command, and the field “Git hash kinetic model” enables access to the exact model version that was documented in Coscine. Hence, continued use of the repository is possible and even encouraged, and the metadata in Coscine is considered the documentation of a snapshot of a relevant model version. There are several reasons for creating such a snapshot. The submission and acceptance of a peer-reviewed paper are typical, and Table 15 provides metadata documenting the model developed by Langer et al. [13] after acceptance of their manuscript. However, a new resource node can also be added when a researcher presents results obtained with a work-in-progress version of a model in a meeting or a seminar. A new resource node must be created when a model developer provides a model to a colleague. The application profile summarizes the validation range by specifying considered fuels and a list of validation targets. However, further validation details are provided in a separate resource node, an instance of the validation database “ITVKineticModelValid,” which is discussed in the following subsection. The validation database used to validate a kinetic model should be documented using the fields “Repository validation database” and “Git hash validation database.” A kinetic model and a validation database may be maintained in the same repository, as is the case for the ITV PAH2023 kinetic model described in Table 15. 4.3.7 A validation database a detailed chemical kinetic model (Data set) Validation databases support various tasks related to chemical kinetic modeling, such as developing detailed chemical kinetic models (as discussed in the previous subsection), comparing and benchmarking different kinetic models, and performing model reduction or optimization. Since the data used in these tasks are not specific to a single model, they must be managed independently. Like kinetic models, the relevant data for validation databases, typically collected from the literature, must be managed in a git repository. The metadata fields relevant to validation databases are outlined in Table 8. As the fields overlap those in Table 7 for chemical kinetic models, the reader is referred to the previous subsection for a discussion. The example validation database discussed in this section is the one employed for the work of Langer et al. [13]. Table 16 shows the metadata for this database, which has the following directory structure: validation cases/ ITV Gasoline Validation Cases Cai2015+2019/ 33
BurningVelocityClean/ CounterflowFlames/ FR/ IgnitionDelay/ JSR/ README.md submit.sh submit all.sh The top-level README.md file provides detailed instructions on how to use the validation cases, which are organized by quantity of interest or, in the case of speciation data, configuration type. Within these directories, the cases are organized by fuel composition (not shown), where each subdirectory provides experimental data provided in text files (preferably YAML or CSV format) and scripts for plotting the data (preferably Python). Inlet, initial, and boundary conditions are stored as FlameMaster input files or text files (preferably YAML format) in conjunction with scripts that generate the FlameMaster input files (preferably Python). A script called “run.sh” must be added for each separate validation target that can be used to perform the FlameMaster simulations and generate the figures comparing the measured and predicted results. The toplevel shell scripts can be executed to submit jobs that perform simulations of all validation cases on the HPC system. 34
Table 15: Example application of ITVKineticModel. Field Name Response Title A detailed kinetic model for aromatics formation from small hydrocarbon and gasoline surrogate fuel combustion Contact person Raymond Langer Contact person ORCID 0000-0001-7309-4544 Contact person email [email protected]hen.de Producer / Author Qian Mao, Heinz Pitsch Keywords / Tags PAH; chemical kinetic modeling; flux analysis Legal information none Date creation 23-06-2022 Funding The CRG project under project number URF/1/4688-01-01 funded by KAUST Status publication published Publication title A detailed kinetic model for aromatics formation from small hydrocarbon and gasoline surrogate fuel combustion Publication journal / conference Combustion and Flame Publication DOI doi.org/10.1016/j.combustflame.2022.112574 Comments / Explanation The ITV model is constantly updated; ask members of the kinetics group for recommendations on which kinetic model should be used for a target application. Git clone URL kinetic model [email protected]hen.de:ITV/mech.git Subdirectory SootMechanisms/ITV 2015/ Langer2021/kritika liming/soot langer Git hash kinetic model 9351351f41a05561b4983cac1d2eba9746dc2564 Git clone URL validation database [email protected]hen.de:ITV/mech.git Git hash validation database 9351351f41a05561b4983cac1d2eba9746dc2564 Fuels n-heptane (n-C7H16), iso-octane (i-C8H18), toluene (A1CH3), ... Validation targets IDT(ST), IDT(RCM), LBV, CF, JSR, FR 35
Table 16: Example application of ITVKineticModelValid. Field Name Response Title ITV kinetic model validation database Contact person Raymond Langer Contact person ORCID 0000-0001-7309-4544 Contact person email [email protected]hen.de Producer / Author Qian Mao, Francesca Loffredo, Sanket Girhe, Heinz Pitsch Keywords / Tags chemical kinetic model validation; IDT; LBV; speciation; PAH Legal information none Date creation 23-06-2022 Funding This work received financial support from various projects. Status publication not published Publication title Publication journal / conference Publication DOI Comments / Explanation This is a central repository for collecting kinetic model validation data used at ITV. Git clone URL validation database [email protected]hen.de:ITV/mech.git Subdirectory GasolineSootUpdate/validation cases Git hash validation database 9351351f41a05561b4983cac1d2eba9746dc2564 Fuels Methane (CH4), n-heptane (n-C7H16), toluene (A1CH3), ... Validation targets IDT(ST), IDT(RCM), LBV, CF, JSR, FR 36
A Applying for S3 storage via JARDS Each project automatically gets 25 GB of storage, which can be increased up to 100 GB via the Quota Settings page. To apply for more than 100 GB of storage resources, it is required to submit an application. Among the different available resource types (available here), S3 resources are the preferred choice for the storage of large data. This resource type is only available after successfully applying for storage. The applications must be submitted via the JARDS platform (link). The process is similar to what is required for computing time applications in CLAIX or other HPC systems. A Principal Investigator (PI) is required for each application. Joachim Beeckmann and Michael Gauding must always be the PI for experimental and simulation resources, respectively. The application must include the following information: •Abstract •Grand IDs (e.g., DFG grant number, Computing time grants) •Amount of requested storage •Type of data stored, such as HDF5 (widely used format for large structured data, easy to handle and with easily-accessible data structure), binary data, or other formats •Description of the data and size distribution (e.g., DNS data from a nonpremixed turbulent jet) •Metadata profile (Always use the custom ITV profiles defined in this guide, e.g., ITVCFD) •How the data will be transferred from/to Coscine (e.g., MinIO for direct transfer from HLRS to Coscine, see this link for more info on S3 clients) •Additional info on metadata, data structure and findability (e.g., Coscine structure for ITV, directory structure, READMEs for data usage, unit conversion) S3 storage applications must be saved in Coscine under the project Coscine S3 storage, JARDS, as done for the HPC computing time proposals. An example of a 125 TB storage application for DNS data of a non-premixed turbulent temporal jet flame [17] can be found in Coscine under the same project (JARDS S3 Proposal link). A clear description of the data structure and the metadata profiles used is crucial for such applications. Acknowledgments This work is funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany’s Excellence Strategy – Cluster of Excellence 2186 “The Fuel Science Center” – ID: 390919832. SPP 2419 HyCAM 37
Coordination Funds provided by the DFG under project number 523792626 are gratefully acknowledged. References [1] Florence Cameron, Yihua Ren, Sanket Girhe, Maximilian Hellmuth, Albrecht Kreischer, Qian Mao, and Heinz Pitsch. In-situ laser diagnostic and numerical investigations of soot formation characteristicsin ethylene and acetylene counterflow diffusion flames blended with dimethyl carbonate and methyl formate. Proceedings of the Combustion Institute, 2022. [2] DFG. Leitlinien zum Umgang mit Forschungsdaten. https://www.dfg. de/download/pdf/foerderung/antragstellung/forschungsdaten/ richtlinien_forschungsdaten.pdf, 2015. Accessed: 2024-10-03. [3] Deutsche Forschungsgemeinschaft. Guidelines for safeguarding good research practice. code of conduct, April 2022. [4] Fabian Fr¨ode, Temistocle Grenga, Sophie Dupont, Reinhold Kneer, Ricardo Tischendorf, Orlando Massopo, Hans-Joachim Schmid, and Heinz Pitsch. Large eddy simulation of iron oxide formation in a laboratory spray flame. Applications in Energy and Combustion Science, 16:100191, December 2023. [5] Fabian Fr¨ode, Temistocle Grenga, Sophie Dupont, Reinhold Kneer, Ricardo Tischendorf, Orlando Massopo, Hans-Joachim Schmid, and Heinz Pitsch. Modeling combustion chemistry using c3mechv4.0: an extension to mixtures of hydrogen, ammonia, alkanes, and cycloalkanes. Applications in Energy and Combustion Science, 16, 2025. [6] Jonathan M. Goodman, Igor Pletnev, Paul Thiessen, Evan Bolton, and Stephen R. Heller. InChI version 1.06: now more than 99.99% reliable. Journal of Cheminformatics, 13:40, 2021. [7] Stephen R Heller, Alan McNaught, Igor Pletnev, Stephen Stein, and Dmitrii Tchekhovskoi. InChI, the IUPAC international chemical identifier. Journal of Cheminformatics, 7:1–34, 2015. [8] Maximilian Hellmuth, Raymond Langer, Anita Meraviglia, Joachim Beeckmann, and Heinz Pitsch. The role of C3and C4species in forming naphthalene in counterflow diffusion flames. Proceedings of the Combustion Institute, 40(1):105620, 2024. [9] R. Hesse, R. Glaznev, C. Schwenzer, V. I. Babushok, G. T. Linteris, H. Pitsch, and J. Beeckmann. A microgravity flame speed study on refrigerant mixtures of 2,3,3,3-tetrafluoropropene (R1234yf) and difluoromethane (R32). In European Combustion Meeting, 2023. 38
[10] R. J. Kee, F. M. Rupley, and J. A. Miller. CHEMKIN-II: A FORTRAN chemical kinetics package for the analysis of gas-phase chemical kinetics. Sandia Report SAND-89-8009, Sandia National Laboratories, Livermore, CA, 1989. [11] Robert J Kee, Graham Dixon-Lewis, J¨urgen Warnatz, Michael E Coltrin, and James A Miller. A FORTRAN computer code package for the evaluation of gas-phase multicomponent transport properties. Sandia Report SAND-86-8246, Sandia National Laboratories, Livermore, CA, 1986. [12] Robert J Kee, James A Miller, and Thomas H Jefferson. CHEMKIN: A general-purpose, problem-independent, transportable, FORTRAN chemical kinetics code package. Sandia Report SAND-80-8003, Sandia National Laboratories, Livermore, CA, 1980. [13] Raymond Langer, Qian Mao, and Heinz Pitsch. A detailed kinetic model for aromatics formation from small hydrocarbon and gasoline surrogate fuel combustion. Combustion and Flame, 258:112574, December 2023. [14] Paul L. Leberg and Joseph E. Neigel. Enhancing the retrievability of population genetic survey data? An assessment of animal mitochondrial DNA studies. Evolution, 53(6):1961–1965, 1999. [15] RWTH Aachen University. Leitlinie zum Forschungsdatenmanagement an der RWTH Aachen. https://www.rwth-aachen. de/cms/root/forschung/Forschungsdatenmanagement/~ncfw/ Leitlinie-zum-Forschungsdatenmanagement/, 2024. Accessed: 202410-03. [16] Caroline J. Savage and Andrew J. Vickers. Empirical study of data sharing by authors publishing in PLoS journals. PLOS ONE, 4(9):e7078, September 2009. Publisher: Public Library of Science. [17] Gandolfo Scialabba, Marco Davidovic, Antonio Attili, and Heinz Pitsch. Direct numerical simulation of soot break-through in turbulent non-premixed flames. Combustion and Flame, 275:114093, 2025. [18] Timothy H. Vines, Arianne Y. K. Albert, Rose L. Andrew, Florence D´ebarre, Dan G. Bock, Michelle T. Franklin, Kimberly J. Gilbert, JeanS´ebastien Moore, S´ebastien Renaut, and Diana J. Rennison. The availability of research data declines rapidly with article age. Current Biology, 24(1):94–97, January 2014. [19] Timothy H. Vines, Rose L. Andrew, Dan G. Bock, Michelle T. Franklin, Kimberly J. Gilbert, Nolan C. Kane, Jean-S´ebastien Moore, Brook T. Moyers, S´ebastien Renaut, Diana J. Rennison, Thor Veen, and Sam Yeaman. Mandated data archiving greatly improves access to research data. The FASEB Journal, 27(4):1304–1308, 2013. 39
[20] David Weininger. SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. Journal of Chemical Information & Computer Sciences, 28:31–36, 1988. [21] Jelte M. Wicherts, Denny Borsboom, Judith Kats, and Dylan Molenaar. The poor availability of psychological research data for reanalysis. American Psychologist, 61(7):726–728, 2006. Place: US Publisher: American Psychological Association. [22] Mark D. Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Appleton, Myles Axton, Arie Baak, Niklas Blomberg, Jan-Willem Boiten, Luiz Bonino da Silva Santos, Philip E. Bourne, Jildau Bouwman, Anthony J. Brookes, Tim Clark, Merc`e Crosas, Ingrid Dillo, Olivier Dumon, Scott Edmunds, Chris T. Evelo, Richard Finkers, Alejandra GonzalezBeltran, Alasdair J. G. Gray, Paul Groth, Carole Goble, Jeffrey S. Grethe, Jaap Heringa, Peter A. C. ’t Hoen, Rob Hooft, Tobias Kuhn, Ruben Kok, Joost Kok, Scott J. Lusher, Maryann E. Martone, Albert Mons, Abel L. Packer, Bengt Persson, Philippe Rocca-Serra, Marco Roos, Rene van Schaik, Susanna-Assunta Sansone, Erik Schultes, Thierry Sengstag, Ted Slater, George Strawn, Morris A. Swertz, Mark Thompson, Johan van der Lei, Erik van Mulligen, Jan Velterop, Andra Waagmeester, Peter Wittenburg, Katherine Wolstencroft, Jun Zhao, and Barend Mons. The FAIR guiding principles for scientific data management and stewardship. Scientific Data, 3(1):160018, March 2016. Publisher: Nature Publishing Group. 40