Enhancing Data FAIRness with Persistent Identifiers for Variables - Conceptualisation, Requirements, Challenges, and Implementation
Abstract
This report presents an overview of the benefits of implementing variable-level PIDs with regard to the FAIRness of data and offers a practical checklist outlining the requirements for service providers and the challenges faced by infrastructures aiming to adopt variable-level PIDs. In addition, it provides insights into the technical implementation of the PID test environment.
Full text
Dr. Andreas Daniel Enhancing Data FAIRness with Persistent Identifiers for Variables Conceptualisation, Requirements, Challenges, and Implementation Project Report
This work is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 Germany Licence (CC-BY-NC-SA 4.0 https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en ) Project Lead Dr. Andreas Daniel German Centre for Higher Education Research and Science Studies E-Mail: [email protected] Dr. Janete Saldanha Bach GESIS Leibniz Institute for the Social Sciences E-Mail: [email protected] Peter Mutschke GESIS Leibniz Institute for the Social Sciences E-Mail: [email protected] Project Contributors Tilo Villwock Codematix GmbH E-Mail: [email protected] Theresa Möller Codematix GmbH E-Mail: [email protected] Imprint Published by German Centre for Higher Education and Centre for Higher Education Research and Science Studies GmbH (DZHW) Lange Laube 12 | 30159 Hanover | www.dzhw.eu/en P.O. Box 2920 | 30029 Hanover Tel.: +49 511 450670-0 | Fax: +49 511 450670-960 Management Dr. Marcus Beiner Regina Oelfke Chairman of the Supervisory Board Ministerialdirigent Peter Greisler Registration Court Amtsgericht Hannover | B 210251 VAT No. DE291239300 | TAX ID No. 25/206/21502 October 2025
Tabellen- /Abbildungsverzeichnis I Table of Contents Tabellen-/Abbildungsverzeichnis I 1 Introduction 2 2 Persistent Identifiers for Variables: Why They Matter and the Challenges to Consider 3 3 FDZ-DZHW Metadata System, Registration Service, MetadataMapping and Technical Implementation 9 4 Conclusion and Future Developments 16 5 Literature 17 Tabellen-/Abbildungsverzeichnis Figure 1 - Project cockpit (visible only to registered data providers) of the Nacaps 2018 data package with fully implemented metadata. ....................................................................................................... 10 Figure 2 - Example of a variable in the FDZ-DZHW-MDM .................................................................... 11 Table 1 - Mapping PID-variable schema – FDZ-DZHW metadata schema ............................................ 13 Figure 3 - Release dialog in the current release process ...................................................................... 14 Figure 4 - Release dialog with the option for the registration ............................................................. 15 Figure 5 - Project cockpit with registered variable PIDs ....................................................................... 15
Introduction 2 1 Introduction This report presents an overview of the benefits of implementing variable-level PIDs with regard to the FAIRness of data and offers a practical checklist outlining the requirements for service providers and the challenges faced by infrastructures aiming to adopt variable-level PIDs. In addition, it provides insights into the technical implementation of the PID test environment. The project presented in this report was carried out as a KonsortSWD / NFDI4Society short project at the German Centre for Higher Education Research and Science Studies (DZHW) as part of the National Research Data Infrastructure (NFDI). 1 The project aimed to enhance the infrastructure of the Research Data Center for Higher Education Research and Science Studies at the DZHW (FDZ-DZHW), with a focus on improving the FAIRness of data with fine-grained metadata at the variable level using persistent identifier. Dr. Andreas Daniel (DZHW) was responsible for the conceptual design and oversight of the technical implementation. Tilo Villwock and Theresa Möller (Codematix GmbH) were responsible for the technical implementation within the DZHW metadata system. The infrastructure for registering the persistent identifiers, including the comprehensive metadata schema, was provided by the KonsortSWD Measure TA.5-M.1 “Enhancing PID services as a base for a FAIR data infrastructure”, supervised by Peter Mutschke and Janete Saldanha Bach (Gesis) who also lead this project on the side of Gesis. The author would like to thank Dr. Janete Saldanha Bach and Erdal Bara for their support and valuable consulting during the implementation. Special thanks to Dr. Anne Weber and Ute Hoffstätter (DZHW) for reviewing this report and providing valuable suggestions for improvement.2 1 In Saldanha Bach et al. (2023; 2023b), a different DZHW use case is described. Subsequent to the two publications by Saldanha Bach et al., the DZHW use case transitioned to the one outlined in this report. 2 AI tools were used for the linguistic revision and correction of the English text in this report. The content of the revisions has been reviewed and either originates from the authors or is based on properly attributed references.
Persistent Identifiers for Variables: Why They Matter and the Challenges to Consider 3 2 Persistent Identifiers for Variables: Why They Matter and the Challenges to Consider Persistent identifiers (PIDs) identify objects (both digital and physical) uniquely and permanently. PIDs thus enable the sustainable identification, citation and retrievability of the referenced objects. The metadata also stores long-term information about the referenced object and (in the optimal case) references to related objects. While the practice of uniquely identifying datasets using persistent identifiers has become established in recent years, this practice is not well established regarding the variables they contain. This represents a significant shortcoming, as variables (storing the information from measurements) constituting the core units in quantitative social and economic research for addressing research questions. If quantitative data are to be provided in accordance with the FAIR criteria3, it therefore makes sense to also take persistent identifier for variables into account. The following section outlines the benefits of persistent identifiers for variables across the different FAIR dimensions (2.1) and the practical requirements and challenges (2.2) associated with their assignment.4 2.1 Impact on the FAIRness of Variables 2.1.1 Findability It is almost intuitive that the findability of variables that have been assigned unique identifiers is improved simply because they can be clearly distinguished, even if the storage location or variable name has changed (e.g., due to the creation of a new dataset version). This allows scientists to precisely and sustainably cite the specific variables they used in their research, rather than referencing the entire dataset, which typically contains many more variables than were utilized (thus the findability of the used variables is increased). Additionally, linking metadata to variables enables their registration in research databases, further enhancing their findability.5 2.1.2 Accessibility The stable and unambiguous referencing of variables not only improves findability, but also accessibility, as the variable to be accessed is clearly identified even if organisational conditions and/or storage locations change. The associated metadata (ideally) provides information on access rights or at least the status of accessibility. Thus, even if variables are not / no longer accessible or not directly accessible via direct download, this information is clearly assigned to the variable, which eliminates any doubt about the status of accessibility. This aspect is particularly 3 The FAIR-Criteria are a guideline for implementing a data management that leads to open and sustainable data access considering the dimension Findability, Accessibility, Interoperability, Reusability (Wilkinson et al., 2016). 4 See also Klas et al. (2022, pp. 9ff) for practical use cases that also emphasize the relevance of PIDs at variable level. 5 In this context, it seems reasonable to standardise the metadata or at least to map it to existing metadata standards such as DDI, DublinCore and/or Schema.org. See also the section 2.1.3 on interoperability.
Enhancing Data FAIRness with Persistent Identifiers for Variables 4 important for data from the social sciences, as such data is almost always not directly accessible for download due to data protection measures. 2.1.3 Interoperability The improvement of interoperability through persistent identifiers for variables depends largely on the associated metadata. Although the use of metadata standards is not an absolute requirement for persistent identifiers, they significantly enhance their effectiveness and facilitate their widespread adoption when integrated into the registration process. In combination with the unique identifiability of variables, the use of linked standardized metadata in machine-readable formats such as XML or JSON enhances interoperability and also facilitates the automation of data analysis processes (like using the PIDs directly in the analysis code).6 The persistent identifiers of variables (especially when used with standardised metadata) enable the permanent linkage to related resources with PIDs such as datasets, instruments, questions, publications (Handle7, ePIC8, DOI9), researchers (ORCID10), data centers, repositories (re3data.org11), institutions (ROR12) etc.. This facilitates the implementation of linked (open) data (see Daniel et al., 2024), knowledge graphs or even complex ontologies through standardised models like the Resource Description Framework (RDF) or the Web Ontology Language (OWL). Furthermore, the use of persistent identifiers for variables enables the integration of variables into projects such as the Open Research Knowledge Graph (ORKG)13. While metadata is important for persistent identifiers, it should not be too extensive (more on this topic in section 2.5.1 when it comes to the service provider's metadata schema). Variable metadata should be implemented using existing metadata standards (see Wenzig et al., 2025). 2.1.4 Reusability The ability to permanently and unambiguously cite variables simplifies the reproduction of research results by enabling researchers to clearly identify which variables were used in analyses and to obtain precise information about their accessibility (see above). This significantly reduces the risk of errors caused by variable misidentification. Moreover, the integration of standardized metadata during the PID registration process enhances the reuse potential of variables. By adhering to established metadata standards, variable descriptions remain clear and comprehensible over time, ensuring long-term reusability and interoperability (as mentioned before). 2.2 Requirements for Service Provider and Challenges in the Implementation of PIDs As outlined above, implementing persistent identifiers (PIDs) for variables provides numerous benefits; however, several challenges must also be addressed. Frist the implementation requires conceptual and technical expertise, financial resources, and a reliable service provider for registration. Additionally, several key decisions and (self-)assessments regarding the choice of service provider and the organization’s own infrastructure need to be made. The following section will 6 The accessibility of the resource (metadata and the actual data) (e.g., via APIs) also plays a role here for automated processing. 7 https://www.handle.net/ (retrieved on 23.10.2025) 8 https://www.pidconsortium.net/ (retrieved on 23.10.2025) 9 https://www.doi.org/ (retrieved on 23.10.2025) 10 https://orcid.org/ (retrieved on 23.10.2025) 11 https://www.re3data.org/ (retrieved on 23.10.2025) 12 https://ror.org/ (retrieved on 23.10.2025) 13 https://orkg.org/ (retrieved on 23.10.2025)
Persistent Identifiers for Variables: Why They Matter and the Challenges to Consider 5 briefly discuss the requirements for a service provider and the challenges for a Research Data Center (hereinafter RDC14) seeking to register PIDs, based on the experiences from the outlined project. 2.2.1 Requirements for a Service Provider Is the service provider permanently funded? ▪ If not, is there a plan for a potential succession if the temporary funding is withdrawn? Does the service provider offer a transparent pricing model? ▪ This includes basic fees and costs per registration. Additional costs may arise if a separate test system and first-level support are provided. Is the service provider associated with larger PID consortia, such as PID4NFDI? ▪ This could indicate a broader network and a more standardized approach, helping to avoid being stuck in an isolated solution. Does the metadata schema used by the service provider for PID registration rely on an established metadata standard (at least partially)? ▪ Standardization enhances the interoperability of the metadata associated with the PID (as mentioned above) and makes the PID schema compatible with a variety of registration services. This broadens the options available when switching to a new provider and/or the PID system. Is the metadata schema used by the service provider for PID registration kept simple? ▪ Extensive metadata is often seen as beneficial, and for metadata schemas in general, this holds true. However, when it comes to metadata for persistent identifiers, our focus is on ensuring high compatibility and a streamlined registration process. Both become more challenging if the schema is overly complex. In terms of interoperability, a smaller metadata schema offers advantages, particularly for interdisciplinary use, as it can be mapped to a wider range of domains rather than being limited to discipline-specific metadata. Potential disadvantages (especially the lack of information on provenance and discipline-specific content) can be mitigated by embedding the objects (e.g., variables) within a metadata system that includes rich, domain-specific metadata and by linking related objects. Does the service provider provide a documented API for the registration process? ▪ If not, is there an alternative (partly) automated registration process (e.g., XML import via an interface)? Does the service provider offer a test system where the registration process can be tested? ▪ This is crucial for a smooth implementation, as it helps identify pitfalls that should not be tested in a production environment. Given that the registration process for variables can take some time, does the service provider allow for asynchronous processing to prevent blocking other release processes? ▪ As scientific datasets usually hold many variables. the registration of variable PIDs can take a long time (compared to other PID registration processes). Variables are often registered within the context of a data package (or study) containing other objects 14 For simplicity, we here refer to research data centers, but all the following aspects are also applicable to infrastructures like a RDC.
Enhancing Data FAIRness with Persistent Identifiers for Variables 6 (instruments), which themselves are assigned PIDs. If the variable PID registration process is synchronous with other registration processes, it could lead to confusion, as the entire registration will take as long as it takes to register the last variable. This can be avoided by using asynchronous registration for variables. In this case, the registration process for all objects is completed before the variables are registered, and the landing pages of the respective objects become accessible before the variable landing pages are accessible. It is important to provide a warning message that asynchronous registration for the variables is being applied during the release process, to avoid confusion about the variable landing pages not being available promptly. 2.2.2 Challenges for a RDC Seeking to Register PIDs Does the team in the RDC possess the necessary domain knowledge to establish variable metadata registration both conceptually and technically? ▪ Conceptually, having a solid understanding of common metadata standards, metadata schema mapping, metadata systems, and registration services for PID, as well as the ability to define requirements for IT systems and effectively communicate them to technical staff, is essential. ▪ If the required conceptual expertise is not available in the RDC and cannot be obtained through recruitment or training, the assignment of PIDs at the variable level should be reconsidered. ▪ Technically, expertise is required in the (automated) processing of metadata and creation and querying of databases, API interactions, specific data formats (e.g., JSON, XML), and potentially developing user interfaces for metadata registration. ▪ If the necessary technical expertise is not available within the RDC, and cannot be obtained through recruitment or training, evaluate whether it can be outsourced to a service provider. Is the necessary metadata available for all variables? ▪ Investigate whether all mandatory metadata for the registration process is available in the RDC’s metadata system and if it can be easily mapped to the service provider's metadata schema. If not, can it be readily obtained for all variables that need to be registered, both now and in the future? If this is not possible, the assignment of PIDs at the variable level should be reconsidered. Do all variables have unique landing pages to which the PID can refer? ▪ Landing pages are typically subpages within a publicly accessible metadata system, presenting the metadata of a specific object (e.g., a single variable). The URL of the landing page is referenced in the PID metadata, usually establishing a redirect from the PID URL to the landing page. If the landing page changes, the PID metadata must be updated accordingly. Therefore, while the URL of the landing page may change, the PID URL itself remains stable. If unique landing pages are not available, consideration should be given to setting them up. If this is not possible, the assignment of PIDs at the variable level should be reconsidered.
Persistent Identifiers for Variables: Why They Matter and the Challenges to Consider 7 Should the PID of a variable follow a specific structure? ▪ Typically, the service provider assigns a fixed prefix for the PID, while the suffix can be defined by the RDC (see example in chapter 3.1.2). If other PIDs (e.g., DOIs) are already in use, it makes sense to design the prefix in alignment with the existing PID logic. How should different versions of the same variable be handled? ▪ If versioning of the variable metadata is implemented in the metadata system, it should also be accounted for in the registration process of the variable PIDs. Furthermore, it may be considered to incorporate the version directly into the PID itself. If the PID metadata schema allows it, references to previous versions should be included. ▪ In the optimal case each variable version has its own landing page (see above). ▪ If a version-agnostic approach15 is adopted, a clear plan should be established for how updates will be managed and documented. This is essential in any case, but even more so when new versions overwrite or remove previous ones. In general, version-agnostic approaches limit the FAIRness of the variable metadata, as references to the PID may point to a version that is no longer available, which particularly impacts accessibility and reproducibility. If a version-agnostic approach is applied, consideration should be given to introducing versioning. Which other objects should be linked within the PID metadata (if possible)? As described above, the potential of linked open data is fully realized by including references to related objects in the metadata. If the PID metadata allows for it, consideration should be given to including such references (e.g., dataset, study, instrument, questions, previous versions). This requires a metadata system and a metadata schema capable of storing these references. At what stage of the release process should PID registration occur, and who should be authorized to initiate the registration? ▪ The PID registration is a critical process, as once a PID is registered, it cannot be deleted or modified. The number of individuals with the appropriate rights should be limited and granted only to experienced personnel from the RDC Are sufficient technical, human, and financial resources available in the RDC to sustain PIDs? ▪ PIDs require stable landing pages provided by the data provider, necessitating a permanently (publicly) available metadata system for the landing pages. Ongoing costs for maintaining a metadata management system must be considered. ▪ Ongoing costs for utilizing a registration service must also be considered. Additionally, potential scalability factors should be taken into account, especially if it is anticipated that that more variables than currently will need to be registered in the future. ▪ Variable metadata must be aligned with the standards of the service provider. Given that variable metadata is typically extensive, this can require significant effort if done manually. The effort involved in preparing and registering variables largely depends on the level of automation in generating variable metadata within the respective RDC. 15 A version-agnostic approach essentially means that only the latest version of a variable is referenced, and previous versions are not tracked or maintained. This approach focuses on the most current data, disregarding version history.
Enhancing Data FAIRness with Persistent Identifiers for Variables 14 Figure 1 - Release dialog in the current release process The newly implemented release process now includes an option for registering PIDs for variables via the API of the service provider, which is also hosted by da|ra (see Figure 4).23 If the checkbox for registering the PIDs is checked, the following validations are run (besides the validations that were already in place): - Are PIDs already registered for this version of variables? - Is the PID-Service available? If no variables are registered and the service is available, the registration process begins. Since the registration of the data packages only takes a few minutes, the registration of the variable PIDs is performed asynchronously in the background. The user receives a notification stating that the registration of the variable PIDs may take up to an hour, while the data package DOI will be available within a few minutes. 23 As mentioned earlier, since the service provider only provided a test system until the end of the project, the test implementation is presented here. There are plans to implement a productive release process for variable PIDs in the future.
FDZ-DZHW Metadata System, Registration Service, Metadata-Mapping and Technical Implementation 15 If the variable PID registration process is interrupted due to an unforeseen error, any PIDs that have already been registered will be retained. When the registration process is restarted, it will resume from the point of interruption. If the PID-registration process ended successfully it will be displayed in the project cockpit that the variables have been registered (see Figure 5). Figure 2 - Release dialog with the option for the registration Figure 3 - Project cockpit with registered variable PIDs
Enhancing Data FAIRness with Persistent Identifiers for Variables 16 4 Conclusion and Future Developments The approach of assigning persistent identifiers to objects below the study/data package level offers numerous advantages for delivering FAIR research data (as outlined in Section 2). Although the outlined implementation of PIDs for variables is a test case, it has demonstrated that the metadata schema of the KonsortSWD Measure TA.5-M.1 is highly functional, the mapping is straightforward, the technical implementation is effective, and the overall processes align well with the operational framework of the FDZ-DZHW. With the completion of this project, the foundation is laid for the implementation of a productive system. To move forward, the FDZ-DZHW will engage with Measure TA.5-M.1 and the PID4NFDI24 basic service to plan the next steps. One possible approach is to implement a productive system within the ePIC-PID consortium25, which provides persistent identifiers (PIDs) based on the Handle system. As part of the productive implementation, expanding metadata mapping for related objects should be considered. As mentioned above, the metadata schema of Measure TA.5-M.1 allows for the incorporation of various identifiers, including the type of relationship (e.g., IsNewVersionOf, IsDocumentedBy, IsDerivedFrom) (Saldanha Bach et al., 2023, p. 10). Initially, previous and subsequent versions of variables should be linked, along with DOIs of related publications, ORCIDs of authors, the re3data.org ID of the RDC, the ROR ID of DZHW and funding organizations, as well as related questions and tags. Consequentially the PIDs should be included in the Metadata of Variables. In a collaboration between DZHW, DIW and LIfBi, Wenzig et al. (2025) outlined a concept for the publication of standardized metadata on the variable level using DDI 2.5 . PIDs on the variable level are a keystone for the publication of such standardized metadata. If the metadata is enriched with these linkages, the next step could be the integration of related objects into a triple store using RDF, followed by their incorporation into other platforms such as the Open Research Knowledge Graph, Wikidata, or the Competence Network for Bibliometrics. For certain objects that currently lack a registered PIDs, the outlined infrastructure could also provide a solution. Measure TA.5-M.1 extends its scope to include additional data types such as questions, measurement scales, audio files, and video fragments (Saldanha Bach et al., 2023, p. 12; Saldanha Bach et al., 2023b, Slide 9). The implemented processes could therefore also be used to assign PIDs to objects beyond variables. 24 https://base4nfdi.de/projects/pid4nfdi (retrieved on 23.10.2025) 25 http://www.pidconsortium.net/ (retrieved on 23.10.2025)
Literature 17 5 Literature Daniel, A., Goebel, J., Kern, D., Klein, D., May, A., Momeni, F., Nebelin, J., Saalbach, C., Siegers, P., Wenzig, K., & Zapilko, B. (2024). A pilot study for "Linked Open Research Data" (LORDpilot): a LODbased Concept Registry for social science research data. Zenodo. ttps://doi.org/10.5281/zenodo.11047523 Klas, C.-P., Zloch, M., Saldanha Bach, J., Baran, E., & Mutschke, P. (2022). KonsortSWD Measure 5.1: PID Service for variables report (Version 1.0). Zenodo. https://doi.org/10.5281/zenodo.6397367 Saldanha Bach, J., Klas, C.-P., & Mutschke, P. (2023). KonsortSWD Measure 5.1: metadata schema extended report. Zenodo. https://zenodo.org/records/7588902 Saldanha Bach, J., Klas, C.-P., & Mutschke, P. (2023a). KonsortSWD Measure 5.1: use cases description extended report (Version 1.0). Zenodo. https://doi.org/10.5281/zenodo.7588944 Saldanha Bach, J., Klas, C.-P., & Mutschke, P. (2023b). Persistent identifiers for social science survey variables: An infrastructure developed to foster open science. Zenodo. https://doi.org/10.5281/zenodo.8085183 Wenzig, K. & Hansen, D. (2024). Persistent Identifier (PIDs) für Variablen in SOEP-Core. SOEP Survey Papers 1308: Series G – General Issues and teaching Materials. Berlin: DIW Berlin/SOEP Wenzig, K., Daniel, A., Hansen, D., Koberg, T., & Tudose, M. (2025). Publishing Fine-Grained Standardized Metadata – Lessons Learned from Three Research Data Centers. Zenodo. https://doi.org/10.5281/zenodo.15464650 Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.-W., da Silva Santos, L. B., Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C. T., Finkers, R., ... Mons, B. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3(1), Article 160018 https://doi.org/10.1038/sdata.2016.18