scieee AI-readable full text Open interactive document viewer

Privacy-Preserving Linkage of Distributed Pseudonymised Datasets in a Virtual European Rare Disease Platform

Hayn, Dieter,Sandner, Emanuel,Vengadeswaran, Abishaa,Tãtaru, Elena Alexandra,Wilkinson, Mark D.,Hanauer, Marc,Kreiner, Karl,Schreier, Guenter

Abstract

5 Pág.

Full text

Privacy-Preserving Linkage of Distributed Pseudonymised Datasets in a Virtual European Rare Disease Platform Dieter HAYNa,1 , Emanuel SANDNERa, Abishaa VENGADESWARANb, Elena-Alexandra TÃTARUc, Mark WILKINSONd, Marc HANAUERc, Karl KREINERa and Guenter SCHREIERa a AIT Austrian Institute of Technology GmbH, Graz, Austria b Goethe University Frankfurt, University Hospital, Institute of Medical Informatics (IMI), Frankfurt am Main, Germany cFrench National Institute of Health and Medical Research (INSERM), Paris, France dDepartamento de Biotecnología-Biología Vegetal, Escuela Técnica Superior de Ingeniería Agronómica, Alimentaria y de Biosistemas, Centro de Biotecnología y Genómica de Plantas UPM-INIA, Universidad Politécnica de Madrid (UPM), Instituto Nacional de Investigación y Tecnología Agraria y Alimentaria (INIA/CSIC), Madrid, Spain ORCiD ID: Dieter Hayn https://orcid.org/0000-0003-1822-9033, Marc Hanauer https://orcid.org/0000-0002-6758-2506, Karl Kreiner https://orcid.org/0000-00016066-9708, Guenter Schreier https://orcid.org/0000-0003-3724-4255, Elena-Alexandra Tãtaru https://orcid.org/0009-0007-7339-7175 Abstract. Secondary use of data for research purposes is especially important in rare diseases (RD), since, per definition, data are sparse. The European Joint Programme on Rare Diseases (EJP RD) aims at developing an RD infrastructure which supports the secondary use of data. Significant amounts of RD data are a) distributed and b) available only in pseudonymised format. Privacy-Preserving Record Linkage (PPRL) concerns the linking of such distributed datasets without disclosing the participant’s identities. We present a concept for linking a PPRL Service to the EJP RD Virtual Platform (VP). Level 1 (resource discovery) connection is provided by running an FDP within the PPRL Service. On Level 2 (data discoverability), the PPRL Service can represent both, an individual and a catalog endpoint. Our solution can count patients in PPRL-supporting resources, count duplicates only once, and count only patients registered to multiple resources. Currently, we are preparing the deployment within the EJP RD VP. Keywords. Rare diseases, privacy-preserving record linkage (PPRL), research infrastructure, secondary use, European health data space (EHDS) 1. Introduction Secondary use of data for research purposes is especially important in rare diseases (RD), since, per definition, data are sparse. Support of secondary use of health data is one of 1 Corresponding Author: Dieter Hayn, Reininghausstr. 13/1, Graz, Austria; E-mail: [email protected]. Digital Health and Informatics Innovations for Sustainable Health Care Systems J. Mantas et al. (Eds.) © 2024 The Authors. This article is published online with Open Access by IOS Press and distributed under the terms of the Creative Commons Attribution Non-Commercial License 4.0 (CC BY-NC 4.0). doi:10.3233/SHTI240683 1442 the main aims of the European Health Data Space (EHDS,[1]). For RD, the European Joint Programme on RD (EJP RD) and its successor, the European RD Research Alliance (ERDERA) aim at developing a European RD infrastructure, which links to the EHDS, based on the Findable, Accessible, Interoperable and Reusable (FAIR,[2]) principles. In the absence of clinical guidelines, patients are often treated according to the most recent trial protocol (see e.g.[3]). In addition to electronic trial data capture systems and registries, RD patients’ biological samples (blood, tumour tissue, urine, bone marrow, etc.) and genomic profiles are typically stored in biobanks. Therefore, significant amounts of quality-controlled RD data are a) distributed and b) available only in pseudonymised format, with different pseudonyms for different contexts, as demanded in the General Data Protection Regulation (GDPR) [4]. Privacy-Preserving Data Linkage (PPRL) concerns the linking of different datasets without disclosing the participant’s identity information. PPRL can be applied on personalized data (e.g., between different hospitals) or between pseudonymised datasets (e.g., between different registries). PPRL concerns a huge variety of scenarios. This paper focusses on a small sub-set of three dedicated use cases related to the EJP RD: 1. Count patients in PPRL-supporting resources. As a RD researcher, I would like to count patients registered in European RD resources supporting PPRL so that I can quantify the current state of PPRL support in Europe. 2. Count duplicates only once: As a RD researcher, I would like to count patients registered in European RD resources supporting PPRL and count patients registered to more than one resource only once so that I can avoid biased results due to duplicates. 3. Count patients in multiple resources: As a RD researcher, I would like to count patients registered in a minimum number of European RD resources so that I can analyse the overlap of patients between resources. Currently, these use cases are not supported by any European RD infrastructure. The present paper describes a concept, how these use cases could be supported in the EJP RD. 2. Materials and Methods We have developed a concept for linking PPRL Services to the EJP RD infrastructure based on pre-existing components and specifications as described in the following. 2.1. European Joint Programme on Rare Diseases Virtual Platform (VP) 2.1.1. Overall architecture The EJP RD Virtual Platform (VP) is a service-oriented ecosystem of inter-linked webservices that providing researchers with a unified way to access resources such as registries, biobanks, data repositories and catalogues. The VP is distributed and federated by design, meaning that most services are provided from multiple remote European locations rather than a central location. Researchers can enter the VP either via the VP’s interfaces (see below), or by specific VP portals, which provide graphical user interfaces to specify requests in a user-friendly way. 2.1.2. Meta data model Within the VP, a standardized metadata model is being used to express metadata provided by resources. The ‘Data Catalog Vocabulary’ (DCAT), a widely used D. Hayn et al. / Privacy-Preserving Linkage of Distributed Pseudonymised Datasets 1443 vocabulary, recommended by the World Wide Web Consortium (W3C) for describing resources, is used as the base model for the EJP RD metadata. The current version 1.0 of EJP RD metadata model (Figure 1) provides schemas to describe the following types of resources [5]: Patient Registry, Biobank, Dataset, Data Service, Guideline. Figure 1. European Joint Programme on Rare Diseases DCAT based metadata model [5] 2.1.3. Level 1 – Resource discovery Any resource linked to the VP must provide its metadata automatically (rather than manually) by a standard mechanism. Therefore, the FAIR Data Point (FDP) specification is applied [6]. Depending on the resource type (Patient registry, biobank, or guideline resource), different metadata elements are mandatory, recommended, and optional. Any EJP RD resource’s FDP is linked to the EJP RD central FDP index, which is used to discover RD resources by using the portal of the VP (“Level 1”). 2.1.4. Level 2 – Content discovery Level 1 connected resources are encouraged to also support interrogation of anonymous / aggregated data by supporting queries within the data records (“Level 2”). A level 2 query request will typically return information as either yes/no, or counts. Level 2 queries allow researchers to determine if the resource is likely to contain relevant information to investigate their research question according to their respective study design and methodology. While single resources can be queried via the “individual” endpoint, the “catalog” endpoint can be used to query bundles of resources that are linked to a certain catalog provider in a single request. At level 2, resource discovery queries are carried out programmatically using the GA4GH Beacon 2 framework standard [7], which has slightly been adapted to the requirements of the VP. The Beacon 2 framework provides an API and a JSON exchange format for query and results. 2.2. EUPID Services The EUPID Services are a PPRL Service which was developed within the EU FP7 project European Network for Cancer research in Children and Adolescents (ENCCA) [8]. The EUPID Services support generation of different pseudonyms for one and the same patient in different EUPID Contexts (e.g., clinical trials, registries, biobanks), while keeping the possibility to link the pseudonymised datasets in a privacy-preserving way, without the need to involve all the primary sites who hold the patients’ identity data. Therefore, cryptographic algorithms and Argon-2 Hashes are applied. While the EUPID Services stores encrypted / hashed identity data and pseudonyms of registered patients, all clinical / sensitive data is stored separately in the respective resources. D. Hayn et al. / Privacy-Preserving Linkage of Distributed Pseudonymised Datasets1444 3. Results 3.1. Overall architecture The proposed overall architecture for linking PPRL Services such as the EUPID Services to the EJP RD VP is shown in Figure 2. Figure 2. Overall architecture for connecting Privacy-Preserving Record Linkage (PPRL)-Services such as the EUPID Services to the European Joint Programme on Rare Diseases Virtual Platform 3.2. Level 1 – Resource discovery Level 1 (resource discovery) connection between the VP and the EUPID Services is provided by running an FDP within the EUPID Services infrastructure. Within the FDP, basic metadata concerning the service are provided and links to two Level 2 endpoints as described below are provided. 3.3. Level 2 – Content discovery On Level 2, the EUPID Services can represent both, individual and catalog endpoint. As an individual endpoint, the EUPID Services provide information concerning the overall numbers of patients registered within their database. As a catalog, these data are provided a) for each EUPID Context separately (e.g., number of patients in Context A) AND b) for all identified combinations of EUPID Contexts (e.g., number of patients in Context A & B). The subsequent requirements apply for both endpoints. For Level 2 querying, we specified filter criteria for Beacon-2 queries, which need to be added to the current specification of EJP RD Beacon endpoints (as specified in [9]):  F1 “Count duplicates only once” (boolean, default: FALSE)  F2 “Ignore resources not supporting duplicate detection” (boolean, default: FALSE)  N1 Minimum number of resources a patient must be registered to (integer, default: 1) Portals to the VP need to update their GUI and their Beacon-2-requests to support F1, F2, and N1. Additionally, they either need to send queries only to suitable resources or filter responses based on PPRL-support prior visualization. Finally, they are recommended to visualize the number of resources that do not support the selected flags in addition to the count. Resources supporting duplicate detection should count D. Hayn et al. / Privacy-Preserving Linkage of Distributed Pseudonymised Datasets 1445 accordingly when replying to requests with F1, F2 or N1, and indicate in their reply that duplicate detection was applied. Resources not supporting duplicate detection, which are already connected to the VP, are not required to apply any adaptations on their interfaces. 4. Discussion and Conclusions We have developed a concept for linking PPRL supporting resources to the EJP RD VP based on the EUPID Services. Our solution is capable of addressing the three use cases presented in chapter 1 – i.e., counting patients in PPRL-supporting resources, counting duplicates only once, and counting patients registered to multiple resources. By linking the PPRL Service to the VP, all resources connected to the respective PPRL service are connected, without the need to adapt the interfaces of these resources or of other resources already connected to the VP. Based on our approach, querying is only possible based on those data that are available within the PPRL Service. In the case of the EUPID Services, this means that counting can only be done based on patients and based on metadata available on a resource level. No filtering based on patient-level (e.g., age, sex, etc.) is supported. So far, our concept only supports resource discovery and content discovery. Future versions of the VP will also support federated analyses of data from multiple RD resources in Europe (Level 3). However, our concept focusses on Level 1 and 2, only. Future work includes the deployment of our concept within the VP a) based on test data and b) with real-world data. In addition, PPRL will be integrated into further use cases, partly in the course of the European Rare Diseases Research Alliance (ERDERA) project, which will start in September 2024. Acknowledgements This work was supported by the EJP RD (Grant-number 825575). References [1] European Commission. Proposal for a REGULATION OF THE EUROPEAN PARLIAMENT AND OF THE COUNCIL on the European Health Data Space Strasbourg2022. [2] Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. [3] De Wilde B, Barry E, Fox E, Karres D, Kieran M, Manlay J, et al. The Critical Role of Academic Clinical Trials in Pediatric Cancer Drug Approvals: Design, Conduct, and Fit for Purpose Data for Positive Regulatory Decisions. J Clin Oncol. 2022;40(29):3456. [4] THE EUROPEAN PARLIAMENT AND OF THE COUNCIL. REGULATION (EU) 2016/679 OF THE EUROPEAN PARLIAMENT AND OF THE COUNCIL of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (General Data Protection Regulation). 2016. [5] European Joint Programme on Rare Diseases. Metadata for EJP rare disease patient registries, biobanks and catalogs Metadata for EJP rare disease patient registries, biobanks and catalogs [6] Olavo Bonino L, Burger K, Kaliyaperumal R. FAIR Data Point https://specs.fairdatapoint.org/fdp-specsv1.2.html2023 [7] Unified repository for Beacon v2 Code & Documentation https://github.com/ga4gh-beacon/beacon-v2/ [8] Nitzlnader M, Schreier G. Patient identity management for secondary use of biomedical research data in a distributed computing environment. Stud Health Technol Inform. 2014;198:211-8. [9] European Joint Programme on Rare Diseases consortium. REST API specification for querying RD resources (Virtual Platform Level 2) https://github.com/ejp-rd-vp/vp-api-specs D. Hayn et al. / Privacy-Preserving Linkage of Distributed Pseudonymised Datasets1446