scieee AI-readable full text Open interactive document viewer

D8.4 Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data

UNICOM Consortium

Abstract

In this deliverable D8.4, pertaining to Task T8.2 (IDMP and Big Data), the following items are developed :● a landscape analysis of the Big Data Pharmaco-epidemiology networks, presented in an abstract of the D8.3 that covered this issue● a gap analysis of the remaining issues that infringe the implementation of IDMP in the clinical realm and the realm of science in Drug Utilisation Research and Pharmaco-epidemiology● a forecast analysis of the future of IDMP implementation and the facilitators for its successIn addition, we report on the efforts made to conduct a pharmaco-epidemiological comparative study of exposure to ibuprofen sodium versus ibuprofen potassium on various cardiovascular outcomes.An analysis is made of the common data models in big data networks, comparing the variables for medicinal products in OMOP and in CONCEPTION databases, and the implications for IDMP.The results of a survey on national regulations for substitution and INN prescribing are presented, together with a discussion of their consequences for IDMP implementation.An ontology of dose form is presented, and a hierarchy of substance is proposed.Finally, we report on the construction of a minimal data set of full samples of medicinal product packs, standardised to IDMP, pertaining to 4 substances in 10 countries, to allow first experiments in UNICOM pilots and insights for IDMP implementation in national drug databases and pharmaco-epidemiological Big Data Networks.

Full text

This project has received funding from the European Union’s Horizon 2020 research and innovation programme under the Grant Agreement No. 875299 Project Acronym: UNICOM Project full title: Up-scaling the global univocal identification of medicines in the context of Digital Single Market strategy Call identifier: H2020-SC1-DTH-2019 D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Version: 1.0 Status: Final Dissemination Level1: PU Due date of deliverable: 30.03.2024 Actual submission date: 04.04.2024 Work Package: WP8: Clinical Care, Patients, Pharmacies, Research and Pharmacovigilance Lead partner for this deliverable: I-HD Partner(s) contributing: DWIZ, FOUND, HL7, BIDMC Deliverable type2: R Main author(s): Robert Vander Stichele I-HD Carlos Durán I-HD Dipak Kalra I-HD 1 Dissemination level: PU: Public; CO: Confidential, only for members of the consortium (including the Commission Services); EU-RES: Classified Information: RESTREINT UE (Commission Decision 2005/444/EC); EU-CON: Classified Information: CONFIDENTIEL UE (Commission Decision 2005/444/EC); EU-SEC Classified Information: SECRET UE (Commission Decision 2005/444/EC) 2 Type of the deliverable: R: Document, report; DEM: Demonstrator, pilot, prototype; DEC: Websites, patent fillings, videos, etc.; OTHER; ETHICS: Ethics requirement; ORDP: Open Research Data Pilot UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 2 of 53 Other author(s): Miriam Sturkenboom I-HD Catherine Chronaki HL7 Maurizio Taglialatela FOUND Yuri Quintana BIDMC Geert Thienpont I-HD Geert Byttebier I-HD Christophe Maes I-HD Jens de Clercq I-HD UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 3 of 53 a. Revision history Version Date Changes made Author(s) 1.O 28.12.2023 Initial concept RVS 2.0 28.01.2024 First Draft RVS, CD, GB 3.0 29.02.2024 Version for internal review RVS, CD 3.1 10.3.2024 Internal review Ursula Tschorn UT 4.0 27/3/2994 Submission for external review RVS: ALL Statement of originality This deliverable contains original unpublished work except where clearly indicated otherwise. Acknowledgement of previously published material and of the work of others has been made through appropriate citation, quotation, or both. UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 4 of 53 b. Deliverable abstract In this deliverable D8.4, pertaining to Task T8.2 (IDMP and Big Data), the following items are developed : ● a landscape analysis of the Big Data Pharmaco-epidemiology networks, presented in an abstract of the D8.3 that covered this issue ● a gap analysis of the remaining issues that infringe the implementation of IDMP in the clinical realm and the realm of science in Drug Utilisation Research and Pharmaco-epidemiology ● a forecast analysis of the future of IDMP implementation and the facilitators for its success In addition, we report on the efforts made to conduct a pharmaco-epidemiological comparative study of exposure to ibuprofen sodium versus ibuprofen potassium on various cardiovascular outcomes. An analysis is made of the common data models in big data networks, comparing the variables for medicinal products in OMOP and in CONCEPTION databases, and the implications for IDMP. The results of a survey on national regulations for substitution and INN prescribing are presented, together with a discussion of their consequences for IDMP implementation. An ontology of dose form is presented, and a hierarchy of substance is proposed. Finally, we report on the construction of a minimal data set of full samples of medicinal product packs, standardised to IDMP, pertaining to 4 substances in 10 countries, to allow first experiments in UNICOM pilots and insights for IDMP implementation in national drug databases and pharmacoepidemiological Big Data Networks. Keywords: IDMP, Big Data, Pharmacoepidemiology, Drug Utilisation Research, terminology, substitution rules, INN prescribing, ontology of dose form, RxNorm, OMOP, Conception, ePI. This document contains material, which is the copyright of the members of the UNICOM consortium listed above, and may not be reproduced or copied without their permission. The commercial use of any information contained in this document may require a license from the owner of that information. This document reflects only the views of the authors, and the European Commission is not liable for any use that may be made of its contents. The information in this document is provided “as is”, without warranty of any kind, and accepts no liability for loss or damage suffered by any person using this information. © 2019-2023. The participants of the UNICOM project. UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 5 of 53 TABLE OF CONTENTS Revision history ....................................................................................................................................... 3 Deliverable abstract ................................................................................................................................. 4 Deliverable review ................................................................................................................................... 8 List of abbreviations ............................................................................................................................... 10 1 Executive summary ........................................................................................................................ 13 2 Content of the deliverable .............................................................................................................. 14 2.1 Projected Contents of the deliverable ................................................................................... 14 2.2 Authorship and responsibilities .............................................................................................. 14 3 Introduction ..................................................................................................................................... 16 4 Landscape analysis ........................................................................................................................ 17 4.1 The existing big data Networks in Pharmaco-epidemiology ................................................. 17 4.1.1 Europe ............................................................................................................................... 17 4.1.2 USA ................................................................................................................................... 17 4.1.3 Canada .............................................................................................................................. 17 4.2 Common Data Models in Vaccine Safety Network and in Sentinel ....................................... 17 4.2.1 Sentinel .............................................................................................................................. 18 4.3 A comparison of the Common Data Models of OMOP and Conception ............................... 19 4.3.1 Medicinal product representation within ConcePTION and OMOP CDM ......................... 19 4.3.2 Comparison and conclusions with regard to the implementation in Common Data Models 23 4.4 A concrete illustration of the limitations of current datamodels : the story of a failed attempt to conduct a comparison between two salts of diclofenac study........................................................... 24 4.4.1 Problem statement for the diclofenac study ...................................................................... 24 4.4.2 The protocol for the diclofenac study (excerpt of D8.3) .................................................... 25 4.4.3 Initial steps to conduct the study ....................................................................................... 25 4.4.4 Lessons learned and recommendations ........................................................................... 26 5 Gap analysis ................................................................................................................................... 27 5.1.1 The need for complete implementation of IDMP in the National Medicinal Product Dictionaries ..................................................................................................................................... 27 5.2 The need for improvements in data collection on drug prescriptions at the source .............. 27 5.2.1 Making assessment of drug exposure more precise ......................................................... 27 5.2.2 Precise registration of drug regimen.................................................................................. 28 5.2.3 Measurement of actual drug intake ................................................................................... 28 UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 6 of 53 5.3 Alignment of information models for medicinal products ....................................................... 29 5.3.1 Virtual models of SNOMED, RxNorm, IDMP ........................................................................ 29 5.3.2 Models for actual products on the national market ............................................................... 31 5.4 The need for a hierarchy of Substance ................................................................................. 32 5.5 Need for controlled and robust aggregation of dose forms ................................................... 34 5.5.1 Overview of granularity of dose forms in the different terminologies ................................ 34 5.5.2 The superiority of EDQM as a dose form terminology ...................................................... 34 5.5.3 The creation of UNICOM simple Ontology of Dose Form ................................................. 35 5.5.4 The Validation of the ontology ........................................................................................... 38 5.5.5 Alignment of RxNorm and Snomed-CT dose forms to EDQM .......................................... 38 5.6 Normalisation of strength ....................................................................................................... 38 5.7 Remove the uncertainty about the PhPID algorithms ................................................................. 39 5.8 Implementation of IDMP in Electronic Health Records, Patient Summaries, European Health Data Space and pharmaco-epidemiological Big Data networks. ............................................................... 39 6 Forecast assessment ..................................................................................................................... 40 6.1 The delay in IDMP implementation ....................................................................................... 40 6.2 Prospects on implementation of IDMP .................................................................................. 40 6.2.1 DADI project ...................................................................................................................... 40 6.2.2 The role of SPOR Services ............................................................................................... 40 6.2.3 ePI and SPOR ................................................................................................................... 41 6.2.4 The role of GIDWG ............................................................................................................ 41 6.2.5 Speed up of legacy conversion ......................................................................................... 42 6.2.6 Help of the industry through variation procedures ............................................................. 43 6.3 Spreading the news of IDMP in the scientific community ..................................................... 43 6.3.1 What has been achieved ? ................................................................................................ 43 6.3.2 What can be done in the future ? ...................................................................................... 43 6.4 Possible impact of completing IDMP in the regulatory realm on clinical care and big data .. 43 7 Remedial actions within UNICOM to provide meaningful samples of IDMP data ......................... 44 7.1 The minimal data set of “Data AS IS” in Excel. ..................................................................... 44 7.2 Data collection reports by country ......................................................................................... 46 7.2.1 Greece ............................................................................................................................... 46 7.2.2 Italy .................................................................................................................................... 46 7.2.3 USA ................................................................................................................................... 46 7.2.4 Belgium .............................................................................................................................. 46 7.2.5 Norway ............................................................................................................................... 46 7.2.6 Tunisia ............................................................................................................................... 46 7.2.7 Ecuador ............................................................................................................................. 47 7.2.8 Finland ............................................................................................................................... 47 7.2.9 Spain .................................................................................................................................. 47 UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 7 of 53 7.2.10 France ............................................................................................................................ 47 7.3 Quality assurance of the data ................................................................................................ 47 7.4 The move to an SQL database.............................................................................................. 48 7.5 Export from the SQL database to the UNICOM FHIR Server ............................................... 48 7.6 Potential applications ............................................................................................................. 48 7.6.1 Patient facing apps ............................................................................................................ 48 7.6.2 Substitution and INN prescribing ....................................................................................... 49 7.6.3 Cooperation with other Innovative Medicines Initiative (IMI) projects. .............................. 49 7.6.4 Precision studies in Pharmaco-epidemiology .................................................................... 49 7.6.5 Analysis of the therapeutic arsenals of 4 substances in 10 countries ............................... 49 8 Conclusions .................................................................................................................................... 51 9 List of publications related to this work .......................................................................................... 52 10 Annexes ..................................................................................................................................... 53 10.1 Annex 1. ................................................................................................................................. 53 10.2 Annex 2. ................................................................................................................................. 53 10.3 Annex_3 ................................................................................................................................. 53 10.4 Annex_4 ................................................................................................................................. 53 10.1 Annex_5 ................................................................................................................................. 53 LIST OF FIGURES Figure 1. Hierarchy and dimensions of data quality assessment in the INSIGHT tools, in alignment to the EMA data quality framework------------------------------------------------------------------------------------------------------------ 21 Figure 2. Possible scenarios to construct IDMP concepts according to the registration of medicinal products in European Healthcare databases (Taken from UNICOM deliverable 8.3).------------------------------------------- 25 Figure 3. Virtual Drug Model of SNOMED-CT -------------------------------------------------------------------------------------- 30 Figure 4. Virtual drug model of RxNorm --------------------------------------------------------------------------------------------- 31 Figure 5. Actual and Virtual Drug Model of Dm+d ------------------------------------------------------------------------------- 32 Figure 6. Hierarchy of substance ------------------------------------------------------------------------------------------------------- 33 Figure 7. Granularity of value sets in dose form terminologies -------------------------------------------------------------- 34 Figure 8. Value Set of EDQM Intended Site Charactaristic --------------------------------------------------------------------- 35 Figure 9. A simple ontology of dose form ------------------------------------------------------------------------------------------- 37 Figure 10. Results of initial analyses of the Minimal Data Set “Data AS IS”---------------------------------------------- 50 LIST OF TABLES Table 1: information in vaccines table of VSD CDM------------------------------------------------------------------------------ 18 Table 2: Sentinel CDM dispensing table --------------------------------------------------------------------------------------------- 19 UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 8 of 53 c. Deliverable review Internal reviewer:: Ursula Tschorn External reviewer: ......................................... Answer Comments Type* Answer Comments Type* Is the deliverable in accordance with the Description of Action? ☐x Yes ☐ No ☐ M ☐ m ☐ a ☐ Yes ☐ No ☐ M ☐ m ☐ a the international State of the Art? ☐x Yes ☐ No ☐ M ☐ m ☐ a ☐ Yes ☐ No ☐ M ☐ m ☐ a Is the quality of the deliverable in a status that allows it to be sent to European Commission? ☐x Yes ☐ No With minor additions ☐ M x☐ m ☐ a ☐ Yes ☐ No ☐ M ☐ m ☐ a that needs improvement of the writing by the originator of the deliverable? ☐ Yes ☐ No ☐ M ☐ m ☐ a ☐ Yes ☐ No ☐ M ☐ m ☐ a that needs further work by the Partners responsible for the deliverable? ☐ Yes ☐ No ☐ M ☐ m ☐ a ☐ Yes ☐ No ☐ M ☐ m ☐ a Is the structure and contents of the deliverable structured, logical and easy to understand? ☐x Yes ☐ No ☐ M ☐ m ☐ a ☐ Yes ☐ No ☐ M ☐ m ☐ a suitable to meet its intended scope? ☐ xYes ☐ No ☐ M ☐ m ☐ a ☐ Yes ☐ No ☐ M ☐ m ☐ a Is in conformance with UNICOM deliverable template? ☐x Yes ☐ No ☐ M ☐ m ☐ a ☐ Yes ☐ No ☐ M ☐ m ☐ a UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 9 of 53 * Type of comments: M = Major comment; m = minor comment; a = advice UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 16 of 53 3 Introduction UNICOM Project had the ambition to reflect on the implementation of IDMP not only in regulatory affairs and pharmacovigilance, but also in clinical care, supply chain management, and last but not least, in drug utilisation research and pharmaco-epidemiology. With the advent of the European Health Data Space, also population health management was in the focus. The possibility of extensive networks of electronic Health Records, with clinical data and administrative data, opens the perspective of big data and a learning health care system. In this deliverable we will summarize the landscape analysis from D8.3, and supplement it with an analysis of the common data models of OMOP and Conception Networks. We will provide a gap analysis for the implementation of IDMP into Big data networks and then provide a forecast analysis the next years. We will illustrate our analyses with the story of the protocol of a pharmaco-epidemiological study of two diclofenac salts (diclofenac sodium and potassium), to be compared in number and nature of cardiovascular complications at the population level. In addition, the approach to collecting IDMP-compliant minimal data sets for full collections of medicinal product packs for 4 substances from 10 countries will be described, with an comparison on the number of virtual and actual concepts per country. UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 17 of 53 4 Landscape analysis 4.1 The existing big data Networks in Pharmaco-epidemiology Excerpt from D8.3 Pharmaco-epi has been the driving force for the first examples of clinical big data networks. This development is now boosted in Europe by the IMI projects In many high-income countries regulatory agencies are realizing the benefits of accessing and using real world data. In the previous deliverable D8.3 examples of such systems were described from different continents: 4.1.1 Europe • EU-ADR project (Exploring and Understanding Adverse Drug Reactions by Integrating Mining of Clinical Records and Biomedical Knowledge • VAESCO project (Vaccine Adverse Event Surveillance and Communication). • European Commission funded studies: SOS, ARITMO, SafeGUARD • Infrastructural projects such as EMIF and ADVANCE • DARWIN (‘Data Analysis and Real World Interrogation Network’ It shows that EMA is committed to implement a multi-country/database approach to inform regulatory decision making. 4.1.2 USA • Vaccine Safety Datalink (VSD). • US FDA Sentinel • OMOP (Observational Medical Outcomes Project). • FDA Active Postmarket Risk Identification and Analysis (ARIA). Also in the USA, there is a strong commitment to build proactively research networks that will be instrumental for conducting timely and extensive studies on drug safety and effectiveness 4.1.3 Canada • Drug Safety and Effectiveness Network (DSEN) • Canadian Network for Observational Drug Effect Studies (CNODES). 4.2 Common Data Models in Vaccine Safety Network and in Sentinel Excerpts of D8.3 Since the format of data tables that are created for support of health care data exchanges between organizations and across countries there is a need to harmonize both the structure (syntactic) as well as the meaning (syntactic harmonization) of the variables. This is also referred to as common data models. While there are countless study-specific common data models designed for one-time use, common data models are designed for reuse within a network or community of researchers. Examples of such common data models were listed : UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 18 of 53 • Vaccine Safety Datalink The common data model employed by the Vaccine Safety Datalink (VSD) is an example of a syntactically (structurally) harmonized common data model with limited scope, in which only a limited set of variables relevant to vaccine safety are extracted, transformed, and loaded (ETL) to the CDM. The CDM comprises the following tables: Patient (Demographics and enrolment), Vaccination History (Vaccination dates, types, and manufacturers), Medical Visits (Healthcare encounters and diagnoses), Mortality (Death data), and Birth and Pregnancy (Pregnancy and birth data on mother and child). While high in derivation, it has proven utility to address vaccine safety concerns rapidly3. The vaccines table in the VSD CDM comprises the following information Table 1: information in vaccines table of VSD CDM Variable Explanation StudyID: study site CDCsite: HMO site VACDATE: date of vaccine administration VAC: Vaccine administered (numbered specific to number and type of antigens), no product or brand names FACILITY: facility in which vaccine was administered SITE: Body site of administration MFR: Manufacturer LOT: Lot number VACSOURC location of administration 4.2.1 Sentinel The Sentinel CDM is an example of a CDM which is high in reusability, low in derivation, and somewhat broad in scope. The Sentinel Common Data Model is a product of the United States Food and Drug Administration Sentinel Initiative (https://www.sentinelinitiative.org/) and comprises the following tables: Enrolment (periods of health plan enrolment), Demographic (demographic characteristics), Dispensing (outpatient pharmacy dispensing), Encounter (healthcare encounters), Diagnosis (in and outpatient diagnoses), Procedure (in and outpatient procedures), Death (Death records), Cause of Death (Causes of death related to a death record), Laboratory Result (Results of laboratory tests), Vital Signs (Results of measurements), Inpatient Pharmacy (Inpatient drug administrations), Inpatient Transfusion (Inpatient transfusion administration), and Mother-Infant Linkage (Linkage between mothers and liveborn infants).Data in the Sentinel CDM is developed for the United States and is primarily administrative and claims data from health insurers, collected for reimbursement purposes. Source data is harmonized to a common vocabulary for a subset of variables but for the most part the Sentinel CDM retains source data in its original format. The Dispensing data in version 7.1.0 comprise the following information4 3 https://www.cdc.gov/vaccinesafety/pdf/vsd-data.pdf 4 https://dev.sentinelsystem.org/projects/SCDM/repos/sentinel_common_data_model/browse/files/file0012_admin_dispensing .md?at=refs%2Fheads%2FSCDM7.1.0 UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 19 of 53 Table 2: Sentinel CDM dispensing table Variable Name Variable Type and Length (Bytes) Values Definition / Comments / Guideline Example PatID3 Char (Site specific length) Unique member identifier Arbitrary person-level identifier. Used to link across tables. 123456789012345 RxDate Numeric (4) SAS date Dispensing date (as close as possible to date the person received the dispensing). 11/29/2005 NDC Char (11) National Drug Code Please expunge any place holders (e.g., '-' or extra digit). 00006007431 RxSup2 Numeric (4) Days supply Number of days that the medication supports based on the number of doses as reported by the pharmacist. This amount is typically found on the dispensings record. It should not be necessary to calculate this variable for use in the SCDM. Positive integer values are expected. 30 RxAmt2 Numeric (4) Amount dispensed Number of units (pills, tablets, vials) dispensed. Net amount per NDC per dispensing. This amount is typically found on the dispensings record. It should not be necessary to calculate this variable for use in the SCDM. Positive values are expected. 60 4.3 A comparison of the Common Data Models of OMOP and Conception 4.3.1 Medicinal product representation within ConcePTION and OMOP CDM Description of mapping of medicinal product data into the ConcePTION Common Data Model and data quality assurance. The ConcePTION CDM was developed within the framework of the ConcePTION project.5 It has been designed to preserve the granularity resulting from data heterogeneity of European data sources, while conducting distributed analysis efficiently. A key technical characteristic of the ConcePTION CDM is the absence of semantic harmonization prior to the ETL process, making it faster and more flexible across data sources because there is no need for translation of specific meanings from each data source.6 The full set of ConcePTION CDM tables is divided into 4 sections: i) Routine Healthcare Data, ii) Surveillance, iii) Curated tables, and iv) Metadata.7 Medicinal products are ETL´ed in tailored tables within sections i) and iv). 5 https://www.imi-conception.eu 6 Thurin NH, Pajouheshnia R, Roberto G, Dodd C, Hyeraci G, Bartolini C, et al. From Inception to ConcePTION: Genesis of a Network to Support Better Monitoring and Communication of Medication Safety During Pregnancy and Breastfeeding. Clin Pharmacol Ther. 2022 Jan;111(1):321-331. doi: 10.1002/cpt.2476. 7 https://docs.google.com/spreadsheets/d/1hc-TBOfEzRBthGP78ZWIa13C0RdhU7bK/edit#gid=439480870 UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 20 of 53 Medicines and Vaccines tables corresponding to the Routine Healthcare data section aim to collect data on drug and vaccination prescriptions, dispensing or administrations that occurred during routine care. The mandatory fields in the medicines table are: • person_id • medicinal_product_atc_code • date_dispensing (only if date_prescription is not populated) • date_prescription (only if date_dispensing is not populated) • meaning_of_drug_record • origin_of_drug_record Moreover, this table contains non-mandatory fields that could become mandatory depending on the research question. For instance, the variable called medicinal_product_id allow the registration of a unique identifier of a specific medicinal product. Other non-mandatory fields are: • disp_number_medicinal_product (number of dispensed units of the medicinal_product_id) • presc_quantity_per_day (prescribed quantity of medicinal product to be taken daily) • presc_quantity_unit (unit of measure of the prescribed daily quantity. • presc_duration_days (number of days of medication as prescribed) • product_lot_number (an identifier assigned to a particular quantity or lot of medicinal product from the manufacturer) • indication_code (single indication of a condition/indication for which the medicinal product was prescribed/dispensed) • indication_code_vocabulary (coding system referreing to the indication code) • prescriber_speciality (profile of the healthcare professional who has prescribed the medicinal product) • prescriber_speciality_vocabulary (coding system of the speciality) • vist_ocurrence_id (identifier of the prescription. It is a key linking this record to the VISIT_OCCURRENCE table, indicating the visit where the drug was prescribed or dispensed) Within the Metadata section, product table collects information associated to each marketed product that may have been prescribed, dispensed or administered to a patient. It contains one row per product. As in the Medicines table, there are mandatory and non-mandatory fields to be fulfilled. Mandatory fields are: • medicinal_product_name (example: DALSY 20mg/ml ORAL SOLUTION 1 FLASC of 150ml) • medicinal_product_atc_code • In addition, this table contains several non-mandatory fields aiming to collect information about the medicinal product unique identification, unit of presentation, administrable dose form and route, fixedcombination products, concentration, concentration unit, and manufacturer name. Data converted to the ConcePTION CDM is quality assured by the INSIGHT data quality-R tool. The INSIGHT tool allows a detailed characterization of the data source instance (subset of a data source extracted for the purpose of conducting one or more studies),8 including an overview of the availability of events and exposures (medicinal products) of interest. All INSIGHT scripts are publicly available on GitHub at https://github.com/UMC-Utrecht-RWE. A summary of INSIGHT quality check levels is presented in figure 1. 8 Thurin NH, Pajouheshnia R, Roberto G, et al. From Inception to ConcePTION: Genesis of a Network to Support Better Monitoring and Communication of Medication Safety During Pregnancy and Breastfeeding. Clin Pharmacol Ther. 2022 Jan;111(1):321-331. doi: 10.1002/cpt.2476. UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 21 of 53 Description of mapping of medication data in the OMOP Common Data Model and data quality assurance An important part of ETL to the Observational Medical Outcomes Partnership (OMOP) CDM is mapping the source concepts to the standard concepts in the OMOP vocabulary. This can be as straight forward as mapping ICD-10 codes to SNOMED or as difficult as mapping a national drug vocabulary to RxNorm. Drug mapping is challenging because, in contrast to other mappings, it has multiple components that have to be mapped. These components are ingredient, dose form and strength. The mapping uses the RxNorm hierarchy and consists of four steps: 1. Drugs are mapped to RxNorm Ingredient via the 5th level ATC code. The OMOP relationship ‘ATC - RxNorm’ is used for this purpose. 2. Dose form is added to the ingredient level, to map to Clinical Drug Form level. 3. The information on drug strength (including unit) is added to map to Clinical Drug Component. The strength is rounded to two decimals. 4. The above three mappings are combined to map to a Clinical Drug concept. Manual mappings are added for a small number of frequently prescribed drugs. A major limitation of the mapping is the incomplete RxNorm vocabulary. Multiple drugs do not have a counterpart in RxNorm. To be able to map these drugs, the OHDSI community has proposed to use an extension on RxNorm, called Pseudo-RxNorm. Work is in progress to add as many drugs as possible to Pseudo-RxNorm for a complete mapping Many unmappable drugs are drugs consisting of multiple ingredients and cannot be automatically mapped to one RxNorm ingredient. The ATC concept is often too general. The automatic mapping from ATC to RxNorm ingredient should be revisited to accommodate for mapping of these drugs. Other challenges include: ● Synonymous dose forms (e.g. ‘Cream’ and ‘Topical Cream’) ● Numerator and denominator unit (e.g. ‘GL’ to ‘gram’ and ‘liter’) ● Strength derivation (e.g. 8 gram to 8000 milligram) ● Duplicate mappings (e.g. one drug to multiple Drug Forms). Figure 1 . Hierarchy and dimensions of data quality assessment in the INSIGHT tools, in alignment to the EMA data quality framework UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 22 of 53 Figure 2. Graphical representation of the OMOP Common Data Model. The quality checks within OMOP CDM are organized according to the Kahn Framework9 which uses a system of categories and contexts that represent strategies for assessing data quality. Using this framework, the Data Quality Dashboard takes a systematic-based approach to running data quality checks. Instead of writing thousands of individual checks, the framework uses “data quality check types”. These “check types” are more general, parameterized data quality checks into which OMOP tables, fields, and concepts can be substituted to represent a singular data quality idea. This would be considered an atemporal plausibility verification check because the quality control is looking for implausibly low values in some field based on internal knowledge. Researchers can use this check type to substitute in values for cdmFieldName, cdmTableName, and plausibleValueLow to create a unique data quality check. And, since it is parameterized, it can similarly be applied to DRUG_EXPOSURE.days_supply, e.g. the number and percent of records with a value in the days_supply field of the DRUG_EXPOSURE table is different from 0. Version 1 of the tool includes 24 different check types organized into Kahn contexts and categories. Additionally, each data quality check type is considered either a table check, field check, or conceptlevel check. Table-level checks are those evaluating the table at a high-level without reference to individual fields, or those that span multiple event tables. These include checks making sure required tables are present or that at least some of the people in the PERSON table have records in the event tables. Field-level checks are those related to specific fields in a table. The majority of the check types in version 1 are field-level checks. These include checks evaluating primary key relationship and those investigating if the concepts in a field conform to the specified domain. Concept-level checks are related to individual concepts. These include checks looking for gender-specific concepts in persons of the wrong gender and plausible values for measurement-unit pairs. After systematically applying the 24 check types to an OMOP CDM version approximately 4,000 individual data quality checks are resolved, run against the database, and evaluated based on a pre-specified threshold. The R package then creates a json object that is read into an RShiny application to view the results. 9 Kahn MG, Callahan TJ, Barnard J, et al. A Harmonized Data Quality Assessment Terminology and Framework for the Secondary Use of Electronic Health Record Data. EGEMS (Wash DC). 2016 Sep 11;4(1):1244. doi: 10.13063/23279214.1244. UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 23 of 53 4.3.2 Comparison and conclusions with regard to the implementation in Common Data Models Postmarketing studies could be conducted in two or more population-based databases. Multi-database (MDB) studies are ideal to increase statistical power, for instance, when studying rare exposures or outcomes, or there is a need to inform results from different settings or jurisdictions. The application of a common study protocol to conduct MDB studies is a fundamental difference between them and the classic meta-analysis. By applying a common study design, harmonized definitions of exposures, outcomes, and covariates, and by performing a standardized analysis plan, heterogeneity is reduced among databases, and therefore, the external validity increases. The European ENCePP Methodological Guide has described four approaches to conduct MDB studies: local analyses, sharing of raw data, use of CDM with study-specific data, and use of general CDM.10 The ConcePTION and OMOP CDMs differ on their approaches as the former uses a CDM with studyspecific data and the latter the general CDM approach.11 The ConcePTION CDM is a protocoldependent CDM where a common protocol is agreed by study partners and data is locally extracted and loaded into the CDM. Then, data is processed locally using a single analysis programme. The output of the analysis is transferred to a coordination team to be post-processed. On the contrary, OMOP CDM is a protocol-independent CDM where local data is loaded into the CDM prior to and independent of the study protocol. When a new a study is decided, a protocol is agreed by study partners and data in the CDM (already) is processed locally using a central analysis programme. As in the previous approach, the outputs are shared centrally for post-processing steps. Regarding the ETL processes of medicinal products, in ConcePTION CDM medicines and vaccines are extracted and loaded following the protocol requirements and the ETL specifications, a study specific document as well. ATC codes are mandatory; however, and depending on the research question, a detailed characterization of the medicinal products could be obtained, including a national identification number allowing for the construction of PhPIDs and MPID in different databases. In OMOP CDM, medicinal products are mapped following the RxNorm and Pseudo-RxNorm hierarchy, which does not allow for the ETL´ing of national-based identifiers. To construct PhPIDs, based on substance with the role of Precise Active Ingredient (hence, with modifier if any) and granular administrable dose form will not be possible, even if the strength would be normalised. In conclusion, the generalised introduction of IDMP in national medicinal product dictionaries and data collection systems would greatly advance the precision and ease of representing medicinal products in the Common Data Models of the big pharmaco-epidemiological networks and the resources of the European Health Data Space. 10 https://encepp.europa.eu/encepp-toolkit/methodological-guide_en 11 Gini R, Sturkenboom MCJ, Sultana J, Cave A, Landi A, Pacurariu A, Roberto G, Schink T, Candore G, Slattery J, Trifirò G; Working Group 3 of ENCePP (Inventory of EU data sources and methodological approaches for multisource studies). Different Strategies to Execute Multi-Database Studies for Medicines Surveillance in Real-World Setting: A Reflection on the European Model. Clin Pharmacol Ther. 2020 Aug;108(2):228-235. doi: 10.1002/cpt.1833. UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 24 of 53 4.4 A concrete illustration of the limitations of current datamodels : the story of a failed attempt to conduct a comparison between two salts of diclofenac study 4.4.1 Problem statement for the diclofenac study Diclofenac, a phenylacetic acid derivate, is a Nonsteroidal Anti-inflammatory Drug (NSAID). During decades it has been widely used for the symptomatic treatment of chronic musculoskeletal pain and inflammatory conditions such as osteoarthritis, rheumatoid arthritis and ankylosing spondylitis; periarticular disorders, sprains and strains, and some acute painful conditions, mainly postoperative pain, gout, renal colic, migraine and dysmenorrhea.12 Diclofenac has proven to be effective to control pain in chronic inflammatory conditions as well as for acute pain relief.13 However, during the last 20 years, diclofenac´s cardiovascular safety profile has raised concerns. Several pharmacoepidemiologic safety studies have demonstrated the association between non-selective NSAIDs, such as diclofenac, and cardiovascular adverse outcomes, i.e. acute myocardial infarction, heart failure and stroke14,15,16,17,18. Until today, it has been proved that diclofenac intake increases the risk of acute myocardial infarction, heart failure and death in those who have previously suffered a major cardiovascular event19,20, and do so by 50% in people without history of cardiovascular disease when compared with non-users, even after short periods of use.21 Although most diclofenac products marketed in Europe are bioequivalent potassium and sodium salts for systemic use (oral and parenteral), there has not been performed pharmacoepidemiologic studies aimed to address specific research questions regarding the cardiovascular safety profile derived from the use of the available oral diclofenac potassium or sodium alternatives. Diclofenac pharmaceutical alternatives have been tested for the control of acute postoperative pain. Eighteen studies were evaluated in a Cochrane´ systematic review showing good rates of pain control after diclofenac potassium administration and limited efficacy of diclofenac sodium in this indication.22 Moreover, there were no differences on the rates of adverse events after a single dose administration. Detailed identification of diclofenac oral salts, i.e. through IDMP in large electronic healthcare databases will allow for post-authorization safety studies of the cardiovascular safety outcomes of specific diclofenac pharmaceutical alternatives. 12 Diclofenac. Martindale: The Complete Drug Reference. Royal Pharmaceutical Society; 2020 [cited 2020 Oct 28]. Available from: https://www.medicinescomplete.com/#/content/martindale/13409-w?hspl=diclofenac 13 Derry S, Wiffen PJ, Moore RA. Single dose oral diclofenac for acute postoperative pain in adults. Cochrane Database of Systematic Reviews 2015;2017 14 Masclee GMC, Straatman H, Arfè A, Castellsague J, Garbe E, Herings R, et al. Risk of acute myocardial infarction during use of individual NSAIDs: A nested case-control study from the SOS project. PLoS ONE. 2018 Nov 1;13(11). 15 Arfè A, Scotti L, Varas-Lorenzo C, Nicotra F, Zambon A, Kollhorst B, et al. Non-steroidal anti-inflammatory drugs and risk of heart failure in four European countries: nested case-control study. BMJ. 2016 Sep 28;354:i4857. 16 Schjerning Olsen AM, Fosbøl EL, Lindhardsen J, Andersson C, Folke F, Nielsen MB, et al. Cause-Specific Cardiovascular Risk Associated with Nonsteroidal Anti-Inflammatory Drugs among Myocardial Infarction Patients - A Nationwide Study. PLoS ONE. 2013;8(1). 17 Bally M, Dendukuri N, Rich B, Nadeau L, Helin-Salmivaara A, Garbe E, et al. Risk of acute myocardial infarction with NSAIDs in real world use: Bayesian meta-analysis of individual patient data. BMJ. 2017;357. 18 Baigent C, Bhala N, Emberson J, Merhi A, Abramson S, Arber N, et al. Vascular and upper gastrointestinal effects of nonsteroidal anti-inflammatory drugs: Meta-analyses of individual participant data from randomised trials. The Lancet. 2013;382(9894):769–79. 19 Gislason GH, Jacobsen S, Rasmussen JN, Rasmussen S, Buch P, Friberg J, et al. Risk of death or reinfarction associated with the use of selective cyclooxygenase-2 inhibitors and nonselective nonsteroidal antiinflammatory drugs after acute myocardial infarction. Circulation. 2006;113(25):2906–13. 20 Schjerning Olsen AM, Fosbøl EL, Lindhardsen J, Folke F, Charlot M, Selmer C, et al. Duration of treatment with nonsteroidal anti-inflammatory drugs and impact on risk of death and recurrent myocardial infarction in patients with prior myocardial infarction: A nationwide cohort study. Circulation. 2011 May 24;123(20):2226–35. 21 Schmidt M, Sørensen HT, Pedersen L. Diclofenac use and cardiovascular risks: Series of nationwide cohort studies. BMJ. 2018;36 22 Derry S, Wiffen PJ, Moore RA. Single dose oral diclofenac for acute postoperative pain in adults. Cochrane Database of Systematic Reviews 2015;2017. UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 25 of 53 4.4.2 The protocol for the diclofenac study (excerpt of D8.3) A research protocol was elaborated and entitled “Demonstrating the value of uniquely identifying a medicinal product: the use of diclofenac salts and the risk of cardiovascular events”. This protocol is (included in UNICOM deliverable D8.3. The protocol detailed the rationale and methods to answer the question whether there is a difference in the incidence rate of acute myocardial infarction, heart failure and stroke after the oral intake of diclofenac sodium versus diclofenac potassium in adults ≥ 18 years old. The study was planned as a population-based dynamic cohort study over a period of 10-years and to be conducted using European electronic healthcare secondary data that has been ETL´ed to the ConcePTION Common Data Model (CDM). Two cohorts were planned, one for diclofenac sodium initiators and other for diclofenac potassium initiators, matched by propensity scores methodology. PhPID had to be constructed to properly identify the exposures. The identification of PhPIDs in electronic healthcare databases was described as a main challenge of this project since the identification of the modified substance name (moiety + salt) is a fundamental requisite to construct the study cohorts, but seldomly required in pharmacoepidemiologic studies. Different potential scenarios regarding the availability to construct PhPIDs in European healthcare databases were described in the protocol and presented in the figure 1 of this deliverable. Figure 2. Possible scenarios to construct IDMP concepts according to the registration of medicinal products in European Healthcare databases (Taken from UNICOM deliverable 8.3) 4.4.3 Initial steps to conduct the study In early 2023, several healthcare data providers were contacted to explore their feasibility to conduct the “diclofenac study” in Europe. The primary criteria to contact data providers was their previous experience conducting studies using the ConcePTION CDM. Then, the first question imposed was the availability of the data provider to detect both diclofenac alternatives on the corresponding database. The unavailability to detect one of them, mostly diclofenac potassium was a reason to stop the contact. Diclofenac modifiers not be detected either because the moiety has not received a marketing UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 32 of 53 Figure 5. Actual and Virtual Drug Model of Dm+d The Virtual Medicinal Product is the closest concept to the Pharmaceutical Product and its identifier in iDMP (PhPID), provided the substance is specified and the dose form expressed as granular administrable dose form and strength is normalised. Although IDMP has no real model for aggregation of concepts at a higher level of the Pharmaceutical product, its precision in both substance identification, and dose form identification may provide the robust foundation for the creation of what is called in the Belgian Model the „Virtual Medicinal Product Group“ (VMPGroup), a formal concept (where the substance is represented by the active moiety and the dose form by the intended Site Characteristics of the granular administrable dose form, and the normalised strength) to support a policy of precise regulations for INN prescribing. 5.4 The need for a hierarchy of Substance During the UNICOM project, the European Substance Reference system (EU-SRS), in cooperating with the FDA, and Work Package 2 in UNICOM witnessed intense progress in cleaning the databases UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 33 of 53 governing the identification of substances (chemical and otherwise), following a rigorous protocol.32 Prospects of Global Substance Identifier GSID) became realistic.33 The results of this progress were exported to the SMS of EMA. SMS provides a central dictionary of substance data in multiple languages. SMS supports the continuous exchange of data between information systems across the European medicines regulatory network and across the pharmaceutical industry.34 It is part of the SPOR infrastructure for substances, products, organisations, and referentials (Dose forms, routes of administration, units of presentation, units of measurement). The SMS system will manage one of the four domains of substance, product, organisation and referentials (SPOR) master data in pharmaceutical regulatory processes. During the UNICOM project, a public export of SMS data was discussed and eventually released. This contains information on substances that is not considered confidential. The export is a flat file of many ten thousands of chemical substances and other types of substances (e.g. proteins, nucleic acids, polymers, and others). While within the EU-SRS database relations between substances and their modifiers are present, there is no practical access for third parties to a hierarchy of substances, that would clearly indicate whether or not a given substance has one or more modifiers that are actually used in medicinal products, and whether or not a moiety can exist in a medicinal product without a modifier. Figure 6 provides the strings to describe the values in the different value sets for each concept (moiety, modifier, moiety+modifier, substance with the role of Precise Active Ingredient (PAI), and the grouper of substances with the same moiety). The problem is that the string for values in the value sets of different concepts may be the same. Only two of these concepts have accepted coding systems (moiety, and moiety+modifier). In fact for these two concepts there are several coding systems (WHODRUG, EUSMS, UNII, CAS, SNOMED-CT). Figure 6. Hierarchy of substance 32 https://unicom-project.eu/wp-content/uploads/2021/08/DataCleansingManual_v1.1.pdf 33 https://who-umc.org/idmp/ 34 https://spor.ema.europa.eu/smswi/#/ UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 34 of 53 To decide for each substance whether a search for modifiers needs to be set up, and if yes, whether the modifier must be specified for a particular medicinal product and if yes, which modifier needs to be chosen for a specific medicinal product, is at the crux of the legacy conversion process. Often it means that a historical search in old development and registration documentation needs to be undertaken. This activity has been labelled as “pharmaco-archaeology”. This is an issue that would need cross-checking between the work of different agencies for medicinal products marketed by the same pharmaceutical company, maybe with the assistance of the company. 5.5 Need for controlled and robust aggregation of dose forms 5.5.1 Overview of granularity of dose forms in the different terminologies The terminology for dose forms from EDQM is a precise but extensive collection of values for terms describing dose forms. The system is complex, and a steep learning curve is needed to master it. Hence, standardisation of the local descriptors of dose form can be a challenge.35 Standardisation of local descriptors can also mean a loss of information sometimes, e.g. when the information about split ability of tablets (not considered by EDQM) gets lost. Figure 7. Granularity of value sets in dose form terminologies 5.5.2 The superiority of EDQM as a dose form terminology There are several reasons why the precise terminology of EDQM dose forms is preferable to use, despite the problems in applying it correctly: 1. The description of dose forms is more precise 2. There is a whole apparatus of definitions, descriptors and characteristics 3. The relationship between pharmaceutical dose form (a general term), manufactured dose form and administrable dose form has been made explicit. 4. A robust aggregation can be built on it Use cases for constructing higher levels of aggregation of dose form are : ● Construction of concepts for operationalising INN prescribing36 35 Sass J, Becker K, Ludmann D, Pantazoglou E, Dewenter H, Thun S. Intercoder Reliability of Mapping Between Pharmaceutical Dose Forms in the German Medication Plan and EDQM Standard Terms. Stud Health Technol Inform. 2018;247:845-849. 36 Van Bever E, Wirtz VJ, Azermai M, De Loof G, Christiaens T, Nicolas L, Van Bortel L, Vander Stichele R. Operational rules for the implementation of INN prescribing. Int J Med Inform. 2014 Jan;83(1):47-56. UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 35 of 53 ● Aligning with RxNorm dose forms, SNOMED-CT dose forms37 ● Creating substitution rules for cross-border services38 ● Speeding up legacy conversion processes of national products to EDQM 5.5.3 The creation of UNICOM simple Ontology of Dose Form During UNICOM the value set of administrable dose forms (N=303) was complemented with the definitions and the characteristics. (see Annex 1 The current version of all types of EDQM contains Pharmaceutical Dose forms, Complex Dose Forms and Combined terms. For each of these dose forms, we collaborated with EDQM to create an explicit link between each specific transformable dose form to the administrable dose form (see Annex 2). These two annexes were constructed with the help of Chris Jarvis, terminology expert of EDQM. Methods to create higher levels of aggregation of granular administrable dose forms are : ● Using the Value set of the “Intended Site” characteristics (see figure 8). ● Using the Value set of the ontology of dose form, developed in WP8 in UNICOM. Figure 8. Value Set of EDQM Intended Site Charactaristic The ontology of dose form was created after an analysis of the unique combinations of 4 characteristics of the EDQM dose form terms: ● Basic dose form of the administrable dose form ● Method of Administration ● Release Characteristic ● Intended Site The resulting groupings were analysed for clinical relevance and if needed, split further or concatenated to more clinically relevant groups.39 37 Karapetian N, Vander Stichele R, Quintana Y. Alignment of two standard terminologies for dosage form: RxNorm from the National Library of Medicine for the United States and EDQM from the European Directorate for the Quality in Medicines and Healthcare for Europe. Int J Med Inform. 2022 Sep;165:104826. 38 see UNICOM deliverables of T6.2 39 Vander Stichele RH, Roumier J, van Nimwegen D. How Granular Can a Dose Form Be Described? Considering EDQM Standard Terms for a Global Terminology. Applied Sciences. 2022; 12(9):4337. https://doi.org/10.3390/app12094337 Value Set of EDQM Intended Site Characteristic Auricular Nasal Buccal Ocular Cutaneous Oculonasal Dental Oral Endocervical Oromucosal Environmental Parenteral Extracorporeal Pulmonary Gastric Rectal Gastroenteral Sublingual Intestinal Transdermal Intramammary Unknown/Miscellaneous Intraperitoneal Urethral Intrauterine Vaginal Intravesical UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 36 of 53 In Figure 9, the resulting ontology is presented as a simple taxonomy, able to include all EDQM granular dose forms. This ontology was implemented in WebProtégé, a web-based groupware tool to create and maintain ontologies, also used to construct the ICD-11 of the WHO. UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 37 of 53 Figure 9. A simple ontology of dose form UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 38 of 53 5.5.4 The Validation of the ontology The ontology constructed from the analysis of unique combinations for characteristics of EDQM dose forms was submitted to a validation process. First, the alignment of RxNorm dose forms was explored.40 It was obvious that the granularity of RxNorm was considerably less than EDQM. Many of the concepts of the above simple ontology could be linked to many EDQM dose forms but to none or only a few RxNorm values. RxNorm has a concept of dose form groups, but this is very rudimentary. In the absence of IDMP implementation it is easier to link medicinal products in the different countries involved in the Pharmaco-epi Big Data Networks to the low granularity RxNorm than to the EDQM granular dose forms. The common data model of OMOP is based on this approach. Within UNICOM, the ontology of drug ontology was circulated to a selected group of experts, with access to the WebProtégé website, requesting to explore the taxonomy and comment. Comments were taken into consideration for the construction of a second version of the ontology, which will be reported in the literature. 5.5.5 Alignment of RxNorm and Snomed-CT dose forms to EDQM Alignment to RxNorm dose forms was explored in a study by WP8 during UNICOM. Definitions of dose forms supplied by RxNorm and EDQM technical documentation were used to align the 120 RxNorm dose forms to the 305 EDQM-based administrable dosage form description terms, making of the above mentioned simple ontology (see Annex 3). During the UNICOM project an intense effort was made to map SNOMED-CT to EDQM.41 The May 2023 EDQM Dose Forms to SNOMED CT Map package contained 310 map rows, with correlations between EDQM map sources and SNOMED CT map targets defined as "Equivalent", "Narrower than", or "Broader than". The mapping was based on the March 2023 version of EDQM and the January 2023 International Edition of SNOMED CT“ However, additional work should be done to align EDQM to SNOMED -CT in the opposite direction. The differences between the two terminologies can be considered as trivial. From the SNOMED-CT perspective a limited number of pharmaceutical dose forms, present in EDQM were not withheld in SNOMED -CT, because of lack of clinical relevance. 5.6 Normalisation of strength Getting the expression of strength right on a global scale is a tricky issue. Each country has his own implicit rules, companies have implicit rules, and so do the medicinal product dictionaries, SNOMED - CT and RxNorm. With the implementation of IDMP and the creation of a global pharmaceutical Product and Pharmaceutical Product Identifier (PhPID), a new attempt is made to harmonise the business rules for normalisation of strength. This effort implies the consideration of many issues : ● the relationship between dose form (pattern) and strength expression ● the determination of the basis of strength (usually the molecular weight of the moiety, sometimes the molecular weight of the moiety+substance) and the availability of that information in the public export of the hierarchy of substance. 40 Karapetian N, Vander Stichele R, Quintana Y. Alignment of two standard terminologies for dosage form: RxNorm from the National Library of Medicine for the United States and EDQM from the European Directorate for the Quality in Medicines and Healthcar ANDe for Europe. Int J Med Inform. 2022 Sep;165:104826. 41 https://confluence.ihtsdotools.org/display/RMT/EDQM+Dose+Forms+to+SNOMED+CT+Map+package +PRODUCTION+Release+Notes+-+May+2023 UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 39 of 53 ● the choice between volume and weight ● the choice between expressing concentration as percentage or as weight over volume. ● whether the dose form is a presentation dose form or a concentration dose form ● uni-dose or multidose for injectables ● Strength calculated based on the Administered Dose Form The effort within the GIDWG (Global IDMP Working Group) bringing together WHO UMC, FDA, EMA and other stakeholders has been tremendous, with a lot of support of the Norwegian Agency NOMA. An informal draft of Business Rules for the normalisation of Strength expression has been produced, but still awaits official publication. 5.7 Remove the uncertainty about the PhPID algorithms The debate on the algorithms to produce the global identifiers of medicinal products is still lingering, with regard to coding systems, level of aggregation of substance and dose form, governance (global by the WHO UMC, national by NCAs, regional by EMA, FDA), and use cases. The hesitation of the US and other regions of the world to implement granular EDQM dose form terms is strong, although the International Council for Harmonisation has made a clear choice in this regard.42 The implementation of IDMP has until now mainly been seen with the regulatory marketing authorisation process and pharmacovigilance in mind. It is surprising that for these use cases the level of granularity for substance and dose form is still under debate. When considering other use cases, such as cross-border services, clinical care, data collection for pharmaco-epidemiology and population health management, precision is also of value. A precise and granular PhPID (specified substance, granular administrable dose form, normalized strength) is the rock-solid, robust foundation for credible aggregation of concepts, useful in particular use cases, such as INN prescribing and substitution rules, within countries and cross-border. One can go from precise identification to aggregated concepts and back. Once aggregated, with loss of original precision, going back is more difficult. 5.8 Implementation of IDMP in Electronic Health Records, Patient Summaries, European Health Data Space and pharmaco-epidemiological Big Data networks. Pharmacotherapy is an essential part of healthcare. The adagium “Tell me what you take and I will tell you what you have” is true to some extent. Deep assessment of the benefits and risk of actual prescribing requires good medical documentation of the relevant antecedents and current problem list, as well as a precise identification of the medicinal products on the medication profile of the patient. The implementation of IDMP in the electronic health records (supported by the medical dictionaries of physician and pharmacist vendors) and in patient summaries (used for integrated communication between the echelons of health care and for cross-border services) may be of great benefit for the primary healthcare data, used in everyday practice, but also for the secondary use. Europe is gearing up for the European Health Data Space (EHDS). Making the Electronic Health Record and the Patient Summary semantically interoperable will have a tremendous impact on the quality of secondary data in the EHDS, for population health management and for pharmaco-epidemiological assessment of risk and benefit of the use of medicinal products. The implementation of IDMP will be a cornerstone of this endeavour. The Unicom Project has set the ground for IDMP implementation in 11 of the 27 NCAs in the region. NCAs not involved in UNICOM are beginning to get awareness of the necessity to participate. Continued support from a consortium of stakeholders, as organised in UNICOM, will also be needed in the future, preferably supported by a new Action Programme. 42 https://www.ema.europa.eu/en/documents/other/mandatory-use-iso-icsrich-e2br3-and-edqmterminology-dosage-forms-df-and-routes-administration-roa_en.pdf UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 40 of 53 6 Forecast assessment 6.1 The delay in IDMP implementation It is clear that the implementation of IDMP in the regulatory realm, in clinical care, and in science does not follow the anticipated pace. Tomorrow will not be the eve of the revolution. The outbreak of the Covid-19 pandemic coincided with the start of the UNICOM project, and that caused considerable stress on the regulatory authorities involved in the project. Reaching global consensus takes time, and the longer it lasts, the greater the chance is that national or regional ad-hoc solutions add new layers of complexity to interoperability. Nevertheless, there has been very substantial progress. The UNICOM project has made possible unprecedented cooperation between Standard Development Organisations, NCAs, EMA, FDA, and WHO Collaboration Centres. Many European Agencies have initiated an ongoing process of IDMP implementation by reorganising their internal data and process management systems. The leaders of the European Health Data Space, the eHealth Network, and the European Digital Service infrastructure eHDSI are fully aware of the importance of IDMP implementation. With sufficient support, in the next lustrum, delay could be turned into state-of-the art advance. 6.2 Prospects on implementation of IDMP 6.2.1 DADI project In the UNICOM project, the work foreseen in Work Package 3, on the implementation of IDMP in the marketing authorisation process for new central medicinal products within EMA has been transferred to EMA itself, given the urgency. The DADI project (Digital Application Dataset Integration) is progressing and will also assure the flow of information towards the Product Management System for the new centrally authorised products.43 The interest of the innovative pharmaceutical industry for this process is intense, and good cooperation is expected. So in the long run, the future of IDMP implementation is assured, as the flow of new medicinal products continues and gradually gains in importance. 6.2.2 The role of SPOR Services SPOR will not solve or replace the fundamental work of product assessors who must decide on the substance identification and the dose form identification. In the legacy conversion process for each medicinal product, decisions need to be made whether the substance needs to be specified and what the correct EDQM granular dose form is. The Substance Management System (SMS) will be a great asset for the NCAs. Hopefully the question of hierarchy of substances will be dealt with, as it might considerably help NCAs in organising the legacy conversion process. The Organisation Management System (OMS) may provide very useful services, by clarifying roles of companies (patent holder, marketing authorisation holder, distributor, …) and by standardizing the company names. The Referentials Management System (RMS) provides practical assistance in terminology for dose forms, units of presentation, routes of administration and units of measurement. For technical reasons the codes of the value sets of EDQM and UCUM will be replaced by SPOR-Codes, but mapping tables between the coding systems should minimise the impact of this double coding approach. The Product Management System (PMS) will be fed by a constant stream of new centrally authorized products, and at some point, by medicinal products stemming from the older medicinal products in the 43 https://www.ema.europa.eu/en/events/digital-application-dataset-integration-dadi-and-product-management-service-pmswebinar-variations-form-human-medicinal-products-what-will-happen-go-live UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 41 of 53 national legacy conversion processes. It is a chicken and egg question whether this service will help the legacy conversion process or is dependent on it, especially in the beginning. This service may bring the dream of a European drug database alive again, but it seems wise, given past experiences, to count on national updating power for maintenance of such an endeavour. 6.2.3 ePI and SPOR The use of electronic Product Information (ePI) was tested in a one-year pilot project by EMA and a group of EU national competent authorities starting in July 2023. During the pilot, companies were responsible for creating and managing ePIs throughout regulatory procedures, utilizing an ePI authoring tool on the Product Lifecycle Management Portal. Following approval and publication by regulators, 16 ePIs became publicly available on the portal and through an application programming interface (API).44 The integration of the Substance, Product, Organisation, and Referentials (SPOR) data into electronic Product Information (ePI) is crucial for ensuring the accuracy and reliability of product information. However, up to now, there are no designated fields provided for the inclusion of SPOR data in ePIs. SPOR data plays a crucial role in identifying and tracking medicinal products throughout their lifecycle, including information on substances, manufacturers, and marketing authorization holders. By integrating SPOR data into ePIs, regulators and healthcare professionals can access comprehensive and up-todate information on products, leading to more informed decision-making and improved patient safety. Therefore, the incorporation of SPOR data into ePIs is essential for enhancing the quality, consistency, and completeness of product information, ultimately benefiting both regulatory authorities and healthcare providers. Efforts should be made to develop and implement standardized fields for SPOR data within ePIs to ensure seamless integration and accessibility of this vital information. 6.2.4 The role of GIDWG Already before the start of UNICOM project, a Global IDMP Working Group (GIDWG) was creatd, involving EMA, FDA, The Uppsala Monitoring Centre (UMC), and other stakeholders. The WHO Collaboration Center for Pharmacovigilance, the Uppsala Monitoring Centre (UMC) has played a catalyst role in the search for a global consensus around the algorithms for the construction of a global identifier for the pharmaceutical product (PhPID). This work was done in conjunction with EMA, FDA, and other stakeholders, and strongly supported by the Norwegian NCA NOMA, and the standard Organisations in Work Package WPI of UNICOM. The GIDWG has worked on : - the discussions about aggregation of substance and dose form - the business rules for normalisation of strength expression - the algorithm to use for PhPID production in function of use cases - the creation and maintenance of repository of PhPIDs This repository would provide a global view on global identifiers. Advanced NCAs could link to this repository and hence create illustrations for other countries. Although not generally accepted, the idea that global governance is absolutely needed in this matter, for most, if not all, NCAs. 44 https://www.ema.europa.eu/en/human-regulatory-overview/marketing-authorisation/product-informationrequirements/electronic-product-information-epi UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 48 of 53 For strength, a verification of the basis of strength (in principle the moiety) was made. In the rather rare cases of substances with a modifier and moiety+modifier as basis of strength, the strength of the moiety can be calculated automatically based on molecular weight of the two instances in the EU-SRS database. In the more frequent cases where the moiety is the basis of strength, the strength of the moiety+modifier can similarly be calculated. This was only done for amlodipine products. For ibuprofen products, the validation of the presence of a modifier, if any, was awaited. Normalisation of strength was checked based on the draft business rules of GIDWG. The correctness of the numerical pack size was controlled, as it is a crucial concept for international comparison of therapeutic arsenals. Finally, we controlled the absence of duplication for the primary key of the medicinal product packs (with a concatenation of country code and NDC). In case of medicinal products with the same numerical pack size but different NDCs, the reason for the variability was checked (pack type, pill characteristics, company name variants), as was the case in the US data. In the EU data, medicinal product packs mostly had only one pack type and no variants of pill characteristics. 7.4 The move to an SQL database Once the quality checks were completed, data were transferred to an online SQL database, with a production version publicly available. This database is used to perform comparative analyses of the characteristics of the therapeutic arsenal (see further). 7.5 Export from the SQL database to the UNICOM FHIR Server In the first phase of this project, an input program was written by DataWizard (keeper of the UNICOM FHIR Server) with a partial extraction of the data, directly from the Excel files. In the second phase, a new input program needs to be written, operating on the SQL database (with HL7 APIs). It is possible that the UNICOM FHIR server needs to be adapted with new variables to be able to operate the variables for the substitution rules. The export to the UNICOM FHIR Server allows additional checks of technical specifications of FHIR resources and IHE profiles. It makes the data suitable for the pilots in UNICOM, such as the patient-facing apps and the crossborder experiments with ePrescriptions, eDispensations, and International Patient Summary. 7.6 Potential applications Having complete samples of all medicinal product packs for 4 substances in 10 countries opens possibilities for preliminary pilots, not depending on the progress of IDMP implementation in the regulatory agencies. 7.6.1 Patient facing apps The minimal data set was essential to make the experimental pilot and demonstrations of exchange of medicinal product information between 3 patient facing apps : ● UnicomPharmaWizard from DataWizard in Italy ● HealthPass from Gnomon in Greece ● InfoSage from Beth Israel Deaconess Medical Center in the USA Demonstrations were made at UNICOM consortium meetings and in Connectathons, and a pilot with 25 residents traveling between the USA, Greece and Italy were performed (see D8.5). UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 49 of 53 7.6.2 Substitution and INN prescribing During the UNICOM project, a survey was sent to the NCAs of the countries represented in UNICOM. The results of this survey were used to explore the possibility of country-specific rules for substitution and INN prescribing, respecting the subsidiarity of member state regulation, but allowing cross-border migration of ePrescriptions, eDispensations and International Patient Summaries (see Annex 4). 7.6.3 Cooperation with other Innovative Medicines Initiative (IMI) projects. For the Minimal Data Set the link to public official labelling of all medicinal products was collected in the form of internet Unique Resource Locators (URLs). This provides access for official labeling information in many languages for similar medicinal products. With the IMI projects Gravitate-Health (on Electronic Product Information -ePI)49 and Conception (pregnancy and lactation information)50 a declaration of intent for cooperation was agreed upon during the UNICOM Consortium Meeting in Ghent, 2023 (see Annex 5). Experiments with multilingualism, cross-linking of pharmaco-therapeutic classes, and structuring of labelling information at class level, substance level, and manufactured item level could be performed. 7.6.4 Precision studies in Pharmaco-epidemiology In the Big Data Networks for pharmaco-epidemiology, data on exposure to medicines are often collected at the level of the ATC substance or RxNorm Clinical Drug. Implementation of IDMP would bring an additional level in medicinal product identification, including information on modifiers of active substances, more granular description of dose form, and normalisation of strength expression. The latter would allow more intricate integration with posology information, so that a more accurate calculation of exposure to medicines can be calculated at the individual and population level. For this task in UNICOM a protocol for a study involving a comparison of exposure to two different salts of ibuprofen (sodium and potassium was designed to study differential cardiovascular outcomes. Ibuprofen was selected for that reason as one of the 4 substances under consideration in the Minimal Data Set. From the 10 countries served by the Minimal Data Set, we found 3 countries with Conception databases, able to track the exposure data back to the medicinal product pack level. However, because of the lack of implementation of IDMP, the additional work to go to this level of precision was cumbersome, leading to prohibitive prices for the study. Implementation of IDMP in Big Data Networks will make such studies easier and cheaper. The relationship between prescribed regimen and actual intake can be studied by analysing temporal spread of dispensed packs in function of the posology, contributing to a better understanding of the medication possession rate and non-adherence. 7.6.5 Analysis of the therapeutic arsenals of 4 substances in 10 countries Another advantage of the Minimal Data Set for 4 substances in 10 countries is the possibility to explore the methodology for cross-national comparison of the therapeutic arsenal in Drug Pricing studies51, affordability studies52, and availability studies.53 49 https://www.gravitatehealth.eu/ 50 https://www.imi-conception.eu/ 51 Machado M, O'Brodovich R, Krahn M, Einarson TR. International drug price comparisons: quality assessment. Rev Panam Salud Publica. 2011 Jan;29(1):46-51. 52 Moye-Holz D, Vogler S. Comparison of Prices and Affordability of Cancer Medicines in 16 Countries in Europe and Latin America. Appl Health Econ Health Policy. 2022 Jan;20(1):67-77. 53 Durán C. Regulation, consumption and expenditures of new cancer drugs in Ecuador and Latin America. [Ghent, Belgium]: Ghent University. Faculty of Medicine and Health Sciences; 2021. UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 50 of 53 In Figure 10, the number of instances for the concepts Medicincal Product (MP) and Medicinal Product Package (MPP) is given for the full Minimal Data Set of 5 countries (pending Finland, Ecuador, Tunesia, France and Spain. Figure 10. Results of initial analyses of the Minimal Data Set “Data AS IS” Minimal Data Set for 4substances and 5countries Source MPs MPPs Status ======================================================================================== Belgium SAM 81 141 Complete Greece IDIKA 197 238 Complete Norway NOMA 43 89 In progress Italy ARIA 165 220 Complete US NLM 1003 5167In progress Finland KELA Ini�ated UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 51 of 53 8 Conclusions Substantial progress has been made for the implementation of IDMP in the regulatory realm, with regard to the adaptation of internal process management of marketing authorisation and the construction of regulatory databases. For new medicinal products, the DADI project governs the electronic IDMPcompliant eApplication Forms.The spread of IDMP implementation over pharmaceutical classes is still limited. Flow of information towards Medicinal Product Dictionaries is in its infancy. Big Data pharmacoepidemiological networks are not ready yet. Cross-border Services start to integrate IDMP in their pilots. The European Union is preparing the European Health Data Space, and envisions the use of IDMP in ePrescriptions, eDispensations, the Patient Summary, and the Electronic Health Record. The integration of Substance, Product, Organisation, and Referentials (SPOR) data into electronic Product Information (ePI) is crucial for enhancing the quality and accessibility of product information, benefiting regulatory authorities and healthcare providers. Sustained support by a new Action Programme will be crucial for the ultimate success of this endeavour. First views on the wide range of applications of IDMP implementation for science were explored in this deliverable, through the pragmatic creation of a Minimal Data Set of IDMP-compliant data for all medicinal product packs of 4 substances (amlodipine, carbamazepine, ibuprofen, simvastatin) for 10 countries (Belgium, Greece, Italy, USA, Norway, Ecuador, Tunisia, Finland, Spain, France). UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 52 of 53 9 List of publications related to this work Vander Stichele RH, Hay C, Fladvad M, Sturkenboom MCJM, Chen RT. How to ensure we can track and trace global use of COVID-19 vaccines? Vaccine. 2021 Jan 8;39(2):176-179. doi: 10.1016/j.vaccine.2020.11.055 Vander Stichele R, Kalra D. Aggregations of Substance in Virtual Drug Models Based on ISO/CEN Standards for Identification of Medicinal Products (IDMP). Stud Health Technol Inform. 2022 May 25;294:377-381. doi: 10.3233/SHTI220478. Vander Stichele R.H, Roumier J, van Nimwegen D. How Granular Can a Dose Form Be Described? Considering EDQM Standard Terms for a Global Terminology. Appl. Sci. 2022, 12, 4337. https://doi.org/10.3390/ app12094337 Karapetian N, Vander Stichele R, Quintana Y. Alignment of two standard terminologies for dosage form: RxNorm from the National Library of Medicine for the United States and EDQM from the European Directorate for the Quality in Medicines and Healthcare for Europe. Int J Med Inform. 2022 Sep;165:104826. doi: 10.1016/j.ijmedinf.2022.104826. UNICOM – D8.4: Landscape Analysis, Gap Analysis and Forecast Assessment for Use of IDMP in Big Health Data Projects Page 53 of 53 10 Annexes 10.1 Annex 1. Annex_1_EDQM_admi nistrableDoseForms_d 10.2 Annex 2. Annex_2_EDQM_Phar maceutical Dose Form 10.3 Annex_3 Annex_3_RXNORM_ali gnement_EDQM.csv 10.4 Annex_4 Annex_4_Minutes workshop patient Safe 10.1 Annex_5 Annex_5_ResultsUnico mSurveyWP4.docx