scieee AI-readable full text Open interactive document viewer

GDI Pillar III User Journeys

Hernandez-Ferrer, Carles; Portell-Silva, Laura; Capella-Gutierrez, Salva; Benetti, Elisa; Bucci, Gabriele; González Jiménez, Eugenio; Genova, Francesca; Gandolfi, Giorgia; Kerry, Giselle; Belien, Jeroen; Fernandez, Jose Maria; Rambla, Jordi; Codó, Laia;

Abstract

This document aims to give a high-level description of GDI’s most common usages and, through them, provide a broad overview of expected functionalities and basic requirements. By reading and dissecting each user journey and story, the reader can envision the bricklaying work being done towards a pan-European health data network infrastructure, with real examples of daily research work (ranging from variant search over large populations to supporting cancer multimodal initiatives) where this network will be not only helpful but foundational. The GDI user journey section describes the specific steps a user will take when interacting with the GDI. It gives a high-level description of GDI’s most common usages and, through them, provides a broad overview of expected functionalities and basic requirements. The user journeys map out the user’s experience, from the initial point of contact (i.e. the authentication into GDI Federation) to the final goal (i.e. data access and analysis), providing context and a deeper understanding of the considerations and challenges the user faces along the way. These user journeys are designed to be technology-agnostic, focusing on user needs rather than specific implementations. Subsequently, user stories have been derived from these user journeys, providing practical examples.

Full text

BSC @ GDI Spanish Node 1/37 27/Aug/2024 - v3.0 GDI user journeys Index Aim and scope ............................................................................................................................ 2 Versioning ............................................................................................................................... 3 Glossary .................................................................................................................................. 4 Workflow ........................................................................................................................................ 5 GDI researcher user journeys ................................................................................................... 7 Authenticate into GDI Federation .................................................................................................. 7 Data Discovery using User Portal’s Data Catalogue for Authenticated Users ........................... 8 Data Discovery using GDI’s Beacon Network for Authenticated Users ...................................... 9 Data Access Request .................................................................................................................. 10 Positive application response .................................................................................................... 11 Amend an application ................................................................................................................. 12 Data Applications Overview (for granted access data) ............................................................ 13 Federated Data Computation ..................................................................................................... 14 Download analytical (and mostly aggregated) results ............................................................. 16 Submission of an analytical workflow ....................................................................................... 17 Acceptance of an analytical workflow ....................................................................................... 18 Denial of an analytical workflow ................................................................................................ 19 GDI user stories born from user journeys ............................................................................... 21 Enabling contact with treating healthcare professional ........................................................... 21 Patients diagnosed with metastatic melanoma ....................................................................... 22 Look up variant summary data in and across GDI datasets .................................................... 24 Lookup of individual genetic variants ........................................................................................ 27 Cancer as a Rare Disease ........................................................................................................... 28 Genetic variant risk assessment within healthcare .................................................................. 30 Metastatic Colorectal Cancer .................................................................................................... 31 GWAS analysis ............................................................................................................................ 33 GWAS analyses across federated nodes .................................................................................. 34 Recalibration of polygenic risk scores ...................................................................................... 35 Ancestry-specific imputation ..................................................................................................... 36 Annex .................................................................................................................................... 37 Documentation ............................................................................................................................ 37 Proposed GDI Access Workflow ................................................................................................ 38 BSC @ GDI Spanish Node 2/37 27/Aug/2024 - v3.0 Aim and scope This document aims to give a high-level description of GDI’s most common usages and, through them, provide a broad overview of expected functionalities and basic requirements. By reading and dissecting each user journey and story, the reader can envision the bricklaying work being done towards a pan-European health data network infrastructure, with real examples of daily research work (ranging from variant search over large populations to supporting cancer multimodal initiatives) where this network will be not only helpful but foundational. The GDI user journey section describes the specific steps a user will take when interacting with the GDI. It gives a high-level description of GDI’s most common usages and, through them, provides a broad overview of expected functionalities and basic requirements. The user journeys map out the user’s experience, from the initial point of contact (i.e. the authentication into GDI Federation) to the final goal (i.e. data access and analysis), providing context and a deeper understanding of the considerations and challenges the user faces along the way. These user journeys are designed to be technology-agnostic, focusing on user needs rather than specific implementations. Subsequently, user stories have been derived from these user journeys, providing practical examples. BSC @ GDI Spanish Node 3/37 27/Aug/2024 - v3.0 Versioning Date Mvm Who Description 21/06/2024 v1 EGJ (BSC) First draft, internal discussion 20/07/2024 v2 CHF (BSC) Second draft after discussion with WP7 - Marco Morelli’s team 26/08/2024 v3 CHF (BSC) Version after discussion within full Pillar III 30/08/2024 v4 LP (BSC) Include comments by Pillar I + other feedback 20/09/2024 v5 LP (BSC) Updated numbering of stories, stories names and updated stories content 09/12/2024 v6 CHF (BSC) Added two new user journeys from D7.5 user stories 20/01/2025 v7 LP, CHF, EGJ (BSC) Final version of user journeys to be frozen BSC @ GDI Spanish Node 4/37 27/Aug/2024 - v3.0 Glossary - (Data) User. An identified, authenticated real person belonging to a research or health organisation that uses GDI infrastructure. A user has a series of connotations according to their background according to the HealthyCloud user profiles and sub-profiles1, the 1+MG Governance and ELSI2, and the 1+MG Glossary3: - Researchers. Researchers are those individuals that will interact with GDI's infrastructure with the goal to obtain, process, and/or analyse research data and its potentially associated outcomes. - Healthcare professional. A healthcare professional is a person that has an active role in processing, analysing and/or understanding health-related data. Their main goal while using GDI's infrastructure is to obtain knowledge that could be transferred to clinical practice. - Policy maker. It is a person, or group of people, responsible for or involved in formulating policies4, especially in politics and related to healthcare. In order to fulfil their duties, they might need to access high level statistics of GDI, usage of GDI’s data, or GDI’s data in its fraction or completeness. A user also includes consultancies who develop policies for a policy maker. - Data provider. According to the definition in HealthyCloud1, a data provider is a person or entity that generates and/or makes available health-related data into the 1+MG, either as a product of their activity or as mandated by another individual or organisation. The data provider may not be the one who has collected the data. The data provider does not act vis-a-vis the user but transfers the (legal) responsibility for the disclosure to the user to the Genome EDIC. Then, the 1+MG Governance and ELSI2, and the 1+MG Glossary3 requires for extra detail: - Data controller (can be a joint controller). Data controllers are those that, as a natural or legal person, as well as a public authority, alone or jointly with others, determines the purposes and means of processing GDI’s data or a subset of it (for instance the data in a national node). A data controller does not require to have access to the data nor to host it to exercise the role. In addition, a data controller is a qualification under the GDPR and changes along the data life cycle. - Data provider are controllers for bringing data into the Genome EDIC - The Genome EDIC is controller for the disclosure to the user - The user is controller for the data use such as a research project 1 FAIR Health Data Portal expected users' interactions: https://zenodo.org/records/10225882 2 1+MG Governance and ELSI: https://framework.onemilliongenomes.eu/governance-elsi 3 1+MG Glossary - version for 1+MG Framework: https://zenodo.org/records/8279620 4 Health Researchers and Policy Makers: A Need to Strengthen Relationship: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3191660/ BSC @ GDI Spanish Node 5/37 27/Aug/2024 - v3.0 - When we have subject-level data discovery in advance for an access request, the user and EDIC are joint controllers for those operations. - Data host. A natural or legal person, a public authority, or a third party entity (such as Amazon, Microsoft or Google) that stores any subset of GDI's data, usually on behalf of another entity (usually a data controller). - IT infrastructure provider. It is a legal person, as well as public authority, that contributes to GDI by providing physical and/or computational infrastructure to achieve the project’s goals in terms of data storage capabilities, data processing capabilities, and/or logistic operative support capabilities to enable meeting the project’s long-term needs. - Data processor. According to the 1+MG Glossary3, a Data processor is a natural or legal person, as well as public authority, which processes personal and/or health data on behalf of the Data controller. - (Data) Processing. Data processing covers any sort of process (including holding data) covering the non exhausting aspects of training AI models, analysis of clinical and/or genetic data for research purposes, identification of shared pathologies across datasets (both using clinical and/or genetic data to do so) or high level or low level statistics on data collections, data sets or individual records. - Data. Any digital information susceptible to being used in health research and regulated by previously mentioned laws and stored within GDI’s premises. - (Individual) Record. Data related to individual data subjects. - Data set. Individual records that are grouped in a certain context (for example: project that generated the "data", clinical condition definition, etc.). they can be characterised using non-personal metadata. - Data collection. Data sets that are grouped together in a certain context and coming from the same data controller(s). They can be characterised with nonpersonal metadata. A data collection, for instance, refers to data of a cohort or public registry. - Clinical data. Personal data related to the physical or mental health of an individual data subject independent of its origin (e.g. from healthcare, research, clinical trials contexts - among others). In 1+MG’s Glossary this is referred to as "health data", with other granularity as "healthcare data" and "health research data"; in other projects “phenotypic data” is also used to describe this information and in Beacon v2 it is called “phenoclinic data”. - Genetic data. Personal data related to the inherited or acquired genetic characteristics of an individual which give unique information about their physiology or health, and which results from an analysis of a biological sample from the individual question. BSC @ GDI Spanish Node 6/37 27/Aug/2024 - v3.0 Workflow BSC @ GDI Spanish Node 7/37 27/Aug/2024 - v3.0 GDI researcher user journeys Authenticate into GDI Federation #: ujGdi01 Story name: Triggered by User Role: Researcher and/or health professional Authors and Contributors: EGJ Preconditions: - User belongs to a recognized research or health organisation - User’s national node uses a compatible and configured IAM Description: User aims to authenticate to any of GDI central services. The user goes to the selected central service and is able to, leveraging their national node’s IAM solution, log in into central services properly authenticated with their national digital identity. BSC @ GDI Spanish Node 8/37 27/Aug/2024 - v3.0 Data Discovery using User Portal’s Data Catalogue for Authenticated Users #: ujGdi02 Story name: Discover a dataset of interest for a scientific purpose Triggered by User Role: Researcher and/or health professional Authors and Contributors: ND, LP, EGJ, GC, TP, LW Preconditions: - Data catalogue is defined and included in the GDI User Portal, displaying the metadata for each dataset discoverable. The Data Catalogue will provide basic and advanced querying and filtering options to the user. Currently the data catalogue is open access as it does not include any personal data and can be explored without logging in. - The metadata is granular enough to help the user identify the right datasets - The user understands what metadata from GDI’s datasets is relevant for their purpose. Description: A user aims to discover datasets of interest. The user explores the Data Catalogue and uses the search and filter functionalities to identify the datasets that align with their objectives and requirements. This search is done based only on metadata associated with each dataset, that metadata can be downloaded if the user wishes to do so. Postconditions: - None BSC @ GDI Spanish Node 9/37 27/Aug/2024 - v3.0 Data Discovery using GDI’s Beacon Network for Authenticated Users #: ujGdi03 Story name: Discover a dataset of interest for a scientific purpose Triggered by User Role: Researcher Authors and Contributors: ND, LP, CHF, EGJ Preconditions: - Following ujGdi01 Description: A user aims to discover datasets of interest. They access the GDI’s User Portal and authenticate following ujGdi01. Once logged in, the user decides to explore the data by performing a Beacon’ clinical and/or genomic query through a search box. This query relies on the GDI capabilities of searching over dataset according to the actual phenoclinic and genomic data in the datasets. The response obtained from GDI’s Infrastructure includes the number of individuals in each dataset that meet the properties defined at the query among other information. Postconditions: - None BSC @ GDI Spanish Node 16/37 27/Aug/2024 - v3.0 Download analytical (and mostly aggregated) results #: ujGdi09 Story name: Download analytical results Triggered by User Role: Researcher and/or health professional Authors and Contributors: EGJ, CHF Preconditions: - The user has executed an analysis following ujGdi08. - The user has been granted access to the results of the analysis executed (all nodes had vetoed the results). - All nodes can connect the other nodes with the purpose of accessing the results of that user executions they have stored. The analytical workflow the user selected is composed by two sections: 1) a series of tasks that must be run in multiple national nodes to extract key information from the datasets of interest by the researcher, to generate intermediate results (ujGdi08); 2) an extra step that will aggregate the results from the previous step in a single “final result” (ujGdi08). There will be scenarios where the second step is not necessary and what is called “partial final results” are, in fact, the “final results” (one in each of the involved national nodes). To the user, it is key to be able to download both the “partial final results” and the “final result” of the selected analytical workflow in order to pursue their research. The national node the researcher belongs to will be the one in charge to collect the “partial final results”. If the second step (aggregation) is necessary, it will also be responsible to make the aggregation (for example, the meta analysis done from multiple GWAS analysis). Then, both the “partial final results” and the “final result” should be accessible from the user’s node. Postconditions: - Both “partial final results” and “final results” are accessible by the user at the end of ujGdi08 + ujGdi09 (this user journey). - Other nodes involved in ujGdi08 are allowed to delete temporary results after some time of ujGdi08 BSC @ GDI Spanish Node 17/37 27/Aug/2024 - v3.0 Submission of an analytical workflow #: ujGdi10 Story name: Submission of an analytical workflow Triggered by User Role: Authors and Contributors: EGJ, CHF Preconditions - The user authenticates themselves in GDI following ujGdi01 - The user understands GDI’s Federated Computation Dashboard features and how to define an analytical workflow. Description The user has an idea of an analytical workflow in order to answer their scientific questions, for which they want to use the GDI infrastructure and data resources. In order to do so, the user inspects the Marketplace, where all the applications (software and models) reside. Picking from the list and concatenating the different applications, the user is able to create a new analytical workflow, either using a graphical user interface application such as the Federated Computation Dashboard or using a workflow description language. Once the user has the analytical workflow defined, they can submit it to acceptance into GDI using the same Federated Computation Dashboard. GDI will answer the user in a reasonable amount of time. During this time GDI will: A) check that the analytical workflow does not leak any data nor performs any undesired operation, and B) inform the national nodes about the new workflow, which will have some time (within the original reasonable amount of time) to test the workflow and raise any technical issues in which case they can request changes be made to the workflow and the application to begin anew. If both A and B are successful, ujGdi11. Otherwise ujGdi12. Postconditions - The analytical workflow is submitted into GDI’s system and accessible by GDI’s staff and national nodes for its validation. BSC @ GDI Spanish Node 18/37 27/Aug/2024 - v3.0 Acceptance of an analytical workflow #: ujGdi11 Story name: Acceptance of an analytical workflow Triggered by User Role: Technical personal from national SPE Authors and Contributors: EGJ, CHF Preconditions - The user authenticates themselves in GDI following ujGdi01. - The user has submitted an analytical workflow for validation. Description The GDI validation committee receives a notification from GDI’s Federation Commutation Dashboard indicating that an analytical workflow has been submitted for validations. The committee inspects the multiple tasks of the workflow to ensure no data is leaked during its execution and that all the operations performed are permeated in the project. Seeing that all the tasks are correct, the committee notifies the national nodes for local testing. The national nodes will be given a reasonable amount of time to locally test the workflow if desired. After that amount of time expires or if they submit a positive result, the analytical workflow will be flagged as acceptable. The validation committee will collect all the feedback from the national nodes and, finding them all (actively or passively) positive, they flag the analytical workflow as accepted. Doing so, the GDI’s Federated Computational Dashboard will now include the analytical workflow as usable by GDI’s user and will notify the owner of the positive validation results. Postconditions - The analytical workflow is now listed as usable within GDI federated computation infrastructure. - The user that submitted the workflow is notified of the positive results. BSC @ GDI Spanish Node 19/37 27/Aug/2024 - v3.0 Denial of an analytical workflow #: ujGdi12 Story name: Deny of an analytical workflow Triggered by User Role: Technical personal from national SPE Authors and Contributors: EGJ, CHF Preconditions - The user authenticates themselves in GDI following ujGdi01. - The user has submitted an analytical workflow for validation. Description The GDI validation committee receives a notification from GDI’s Federation Commutation Dashboard indicating that an analytical workflow has been submitted for validations. The committee inspects the multiple tasks of the workflow to ensure no data is leaked during its execution and that all the operations performed are permeated in the project. If the committee identifies a possible data leak or that some of the operations performed by the analytical workflows are unacceptable within GDI’s context, they flag the analytical workflow as rejected (indicating which operations are leaking data and/or which operations are not accepted). If the committee sees that all the tasks are correct, the committee then notifies the national nodes for local testing. The national nodes will be given a reasonable amount of time to locally test the workflow if desired. During that time the national nodes can test the analytical workflow. If they identify technical or legal issues that forbids them to run the workflow locally, they will document the issue and tag the workflow as unacceptable. Either because the validation committee flagged the analytical workflow as unacceptable or because some national node reported issues, the workflow is reported not to be accepted in GDI’s system. In doing so, the GDI’s Federated Computational Dashboard will notify the owner of the negative validation, providing useful feedback that can be used to fix the reported issues. The workflow then could be submitted again to GDI’s network. Postconditions - The analytical workflow was not included in GDI’s system. - The user that submitted the workflow is notified of the negative results and useful feedback is provided in order to create a new analytical workflow with the same functionality but that has higher possibilities to be accepted in GDI. BSC @ GDI Spanish Node 20/37 27/Aug/2024 - v3.0 GDI user stories born from user journeys Enabling contact with treating healthcare professional #: usGdi01 Story name: The user wants to set up a communication channel with the relevant clinical contact related to patients with the same genetic and/or clinical condition as theirs to discuss possible outcomes Triggered by User Role: Health professional Authors and Contributors: MM, EB, GG, FC, LP, CHF, GC, TB Precondition: - The user must be authenticated, authorised, and valid for GDI. - Doing a query constitutes a “data processing” activity; therefore, the organisation where the user belongs needs to have a data-use agreement with Genome-EDIC. Description: A user has an individual case with a particular genetic and/or clinical description. In order to proceed with caution, they want to contact other healthcare professionals that dealt with similar cases. The user will follow the process described in “Lookup of individual genetic variants” (usGdi02.1) to identify datasets containing similar individuals to the one of their interest. The user goes back to the relevant communication channel in the User Portal and describes their case-in-question (providing the ids of datasets of interest) and submits it, expecting a notification when the communication channel between them is available in a specified time frame. Postconditions: - The user ends up having the relevant contact of their specific case mediated through the Use Portal (aka, relevant communication channel). - Although included in the data-use agreement, the user will not, under any setting, contact any patient discovered using this user story. BSC @ GDI Spanish Node 21/37 27/Aug/2024 - v3.0 Patients diagnosed with metastatic melanoma #: usGdi01.1 Story name: Enabling contact with treating healthcare professional for a patient diagnosed with metastatic melanoma Triggered by User Role: Clinician/oncologist Authors and Contributors: From D7.5 - Cancer Research use case User story description: The use case is a refinement of usGdi01 (Enabling contact with treating healthcare professionals to discuss possible outcomes related to patients with similar genetic or clinical conditions) to a real case scenario. A clinician needs to find a second therapeutic option for his/her patient diagnosed with metastatic melanoma, justified by the identification of possible targetable mutation(s). A secondary objective is to find the molecular alterations explaining the emergence of the resistance of the primary cancer treatment. The goal is to find a similar case in the GDI databases, by querying the network, i.e. find melanoma patients with a BRAF point mutation, treated with BRAF inhibitor Vemurafenib, who have a mutation in the PTEN gene (known variant responsible for therapy resistance). The goal is to search for existing information on the PTEN gene, the BRAF mutation and the impact of the mutation on the Kinase pathway that can explain drug resistance in melanoma cancer patients. Preconditions: The clinician must be authenticated and authorised to GDI. Description: ● The clinician logs into the GDI platform (ujGdi01) and browses available datasets (ujGdi02 or ujGdi03). ● The clinician inputs a Beacon query (ujGdi03) using genomic coordinates or gene fusion terms, to find patients with a PTEN deletion in the genomic data. Data are filtered to focus on the most relevant cases: patients with melanoma and with details concerning the characteristics of the patient, the responses to the treatment, and survival rates. ● The GDI returns results, which may include a dataset relative to another patient with a similar disease and therapy history, found on one node of the network. The dataset is BSC @ GDI Spanish Node 22/37 27/Aug/2024 - v3.0 #: usGdi01.1 Story name: Enabling contact with treating healthcare professional for a patient diagnosed with metastatic melanoma Triggered by User Role: Clinician/oncologist Authors and Contributors: From D7.5 - Cancer Research use case composed of genetic data (NGS sequencing of a melanoma treated with Vemurafenib, after developing resistance to the therapy) and the corresponding clinical data. Both raw genetic data (.Fastq file) and analysed data (.vcf file containing somatic mutations and BED file containing structural alterations) are present. Aggregated data could be visualised via a visualisation tool like in cBioportal. ● If needed, a communication channel with the relevant clinical contact related to the identified patient(s) can be set up. ● The clinician applies for data access to 1+MG data (ujGdi04) and the Data Access Committee grants the access (ujGdi05). ● The clinician downloads the genetic sequence data with the corresponding clinical data of the identified patient(s) (ujGdi09). Postconditions: - BSC @ GDI Spanish Node 23/37 27/Aug/2024 - v3.0 Look up variant summary data in and across GDI datasets #: usGdi02 Story name: Look up variant summary data in and across GDI datasets Triggered by User Role: Researcher and/or health professional Authors and Contributors: LW, GC, TP, CHF User story description: “As a [clinician/researcher/policy maker], I would like to look up the frequency of a given allele in a given [disease, disease type, population, ancestry...] in order to [interrogate pathogenicity of the allele/infer the efficacy of a known drug using pharmacogenomics knowledge/...]” Preconditions: - A database containing the following precomputed information on each allele present: - Beacon - Dataset - Population - Genetic Ancestry Group - Genome Build - Allele Count - Allele Number - Homozygous Count - Hemizygous Count - Female Count - Male Count - A user-friendly central web application to perform the query - Web application should state “The information on this website is not intended for direct diagnostic use or medical decision-making without review by a genetics professional. Individuals should not change their health behaviour solely on the basis of information contained on this website. If you have questions about the information contained on this website, please see a healthcare professional.” - from https://clinicalgenome.org/about/ or similar Description: - User visits a web application that looks something like the following: BSC @ GDI Spanish Node 24/37 27/Aug/2024 - v3.0 #: usGdi02 Story name: Look up variant summary data in and across GDI datasets Triggered by User Role: Researcher and/or health professional Authors and Contributors: LW, GC, TP, CHF - Demo exists here: https://af-browser-demo.ega-archive.org/ - The user performs a genomic query searching for a variant using the format chr-posref-alt through GDI’s federated querying system, with the ability to specify the genome build and cohort/dataset (or all cohorts/datasets) Postconditions: - The output should look something like: BSC @ GDI Spanish Node 25/37 27/Aug/2024 - v3.0 #: usGdi02 Story name: Look up variant summary data in and across GDI datasets Triggered by User Role: Researcher and/or health professional Authors and Contributors: LW, GC, TP, CHF - This type of data could be used for: - Clinical and research variant interpretation (overseen by a genetics professional) - Check if your patient’s variant is present in similar patient cohorts (e.g. https://epi25.broadinstitute.org/) Needs discussion and clarification: - One possible function of GDI would be to create a gnomAD-like portal but for the European population and for disease cohorts. With GoE data, this could be achieved but it might not be the goal of the project. Moreover, the remaining datasets and data collections within GDI could be included into this “resource/service” but to do so, the data access policy and its granularity (beyond public/private) needs to be defined and discussed. BSC @ GDI Spanish Node 32/37 27/Aug/2024 - v3.0 GWAS analysis #: usGdi05 Story name: GWAS analysis Triggered by User Role: Researcher Authors and Contributors: From D7.5 - Infectious diseases use case User story description: A researcher needs to query the GDI infrastructure and perform a rapid Genome-Wide Association Study (GWAS) in order to validate risk variants determining the severity of COVID19. As GWAS requires large sample sizes, multiple comparable datasets are often needed. The goal is to find human genomic datasets that include patients with severe COVID-19. Performing a GWAS in a federated analysis scheme. As a researcher, I want to perform a GWAS across multiple comparable population datasets. Preconditions: The researcher must be authenticated and authorised to GDI Description: - The researcher logs into the GDI platform (ujGdi01). - The researcher needs to query and filter the Central Data Catalogue (ujGdi02) for both human genomic data (i.e. SNPs, variants) and phenotypic data associated with COVID19 severity (i.e. hospitalisation, ICU admission, mortality). - The GDI returns results, which could include human genomic datasets annotated with Covid-19 severity indicators and potentially other phenotypic information, found on one or multiple nodes of the network. - The researcher applies for data access to 1+MG data (ujGdi04) and the Data Access Committee grants the access (ujGdi05). - The researcher accesses the annotated genetic sequence data for further analysis. The researcher executes the analysis by distributing an interoperable analysis pipeline across federated locations (ujGdi08). The researcher can add a workflow to the secure processing environment (SPE) (ujGdi10). - The researcher can download the results from his research results account for further processing (ujGdi09). BSC @ GDI Spanish Node 33/37 27/Aug/2024 - v3.0 GWAS analyses across federated nodes #: usGdi05.1 Story name: GWAS analyses across federated nodes Triggered by User Role: Researcher Authors and Contributors: From D7.5 - 1+MG/B1MG use case User story description: As a researcher, I want to perform a GWAS across multiple comparable population datasets. This is an extension of usGdi05. The goal is to identify genes associated with a particular disease or trait using provided whole genomes. The primary output of a GDI GWAS analysis is a list of P values, effect sizes and their directions generated from the association tests of all tested genetic variants with a phenotype of interest. Preconditions: - The researcher must be authenticated and authorised to GDI - Relevant disease or trait and WGS data must be available for cohort building (ujGdi02) Description: - The researcher logs into the GDI platform (ujGdi01) - The researcher needs to define the study cohort using diseases or trait conditions (ujGdi03) - The GDI returns results, which include the list of human genomic datasets that have specified disease or traits. - The researcher applies for data access to 1+MG data (ujGdi04) and the Data Access Committee grants the access (ujGdi05). - If a researcher gets the approval, then the researcher can access genetic sequence data indirectly by using predefined GWAS pipelines. - The researcher executes the GWAS analysis by distributing GWAS pipelines to federated processing locations (ujGdi08 and ujGdi10). Individual participants' data stay within the country. - The researcher gets the relevant GWAS result data (like list of P values, Manhattan plots etc) to his research results account. - The researcher can download the results from his research results account for further processing (ujGdi09). BSC @ GDI Spanish Node 34/37 27/Aug/2024 - v3.0 Recalibration of polygenic risk scores #: usGdi06 Story name: Recalibration of polygenic risk scores Triggered by User Role: Researcher Authors and Contributors: From D7.5 - Genome of Europe use case User story description: As a researcher, I want to develop PRSs so that they accurately predict disease risk for individuals in a specific ancestry population. A researcher wants to improve the accuracy of polygenic risk scores (PRS) for predicting disease risk in a specific ancestry population. The original PRS was developed using data from another ancestry group, but the researcher needs to adjust the PRS to be effective for individuals of the specific ancestry in their study. The goal is to find human genomic data from multiple populations. This data includes both genetic information and health records, such as the presence or absence of the disease of interest. Subsequently, the PRS can be recalculated in the identified dataset. Preconditions: The researcher must be authenticated and authorised to GDI Description: - The researcher logs into the GDI platform (ujGdi01). - The researcher will construct a Beacon query (ujGdi03), i.e. the researcher inputs the details of the variants used in the original PRS into GDI’s Beacon Network, asking whether the variant is present in multiple population datasets. - The GDI returns results, which could include the list of human genomic datasets together with a summary on group level, i.e. the frequency of the variant in the population. Only datasets containing at least a predefined percentage (cutoff) of the required variants are retained. - The researcher applies for data access to 1+MG data (ujGdi04) and the Data Access Committee grants the access (ujGdi05). - The researcher accesses genetic sequence data. The researcher executes the analysis (PRS) by distributing an interoperable analysis pipeline across federated locations (ujGdi08). The researcher can add a workflow to the secure processing environment (SPE) (ujGdi10). Individual participants' data stay within the country. - The researcher can download the results from his research results account for further processing (ujGdi09). BSC @ GDI Spanish Node 35/37 27/Aug/2024 - v3.0 Ancestry-specific imputation #: usGdi07 Story name: Ancestry-specific imputation Triggered by User Role: Researcher Authors and Contributors: From D7.5 - Genome of Europe use case User story description: A researcher wants to impute a genetic dataset using ancestry-specific reference panels The goal is to impute missing genetic data using a multi-ethnic reference panel (Genome of Europe-GoE), which can capture ancestry-specific genetic variants more effectively than general panels. This requires learning correlation patterns between SNPs in the GoE data (reference panel), then using these to infer (impute) the missing genotypes in a new sample. Typically, ±1M variants are measured, and then imputed towards >40M variants using the correlation data. Preconditions: The researcher must be authenticated and authorised to GDI Description: - The researcher logs into the GDI platform (ujGdi01). - The researcher needs to query and filter the Central Data Catalogue (ujGdi02) (filter on ancestry) for multi-ethnic reference panels. - The GDI returns results, which could include aggregated reference panels. - The researcher applies for data access to 1+MG data (ujGdi04) and the Data Access Committee grants the access (ujGdi05). - The researcher accesses the reference panels (aggregated data). The researcher executes the analysis (imputation of genetic datasets). - The researcher can download the results from his research results account for further processing (ujGdi09). BSC @ GDI Spanish Node 36/37 27/Aug/2024 - v3.0 Annex Authors and Contributors CHF: Carles Hernandez-Ferrer EB: Elisa Benetti GB: Gabriele Bucci EGJ: Eugenio González Jiménez FG: Francesca Genova GG: Giorgia Gandolfi GK: Giselle Kerry JAMB: Jeroen Belien JMF: Jose Maria Fernandez JR: Jordi Rambla LC: Laia Codó LP: Laura Portell LW: Laura Whelan MM: Marco Morelli ND: Nicolás Díaz SCG: Salvador Capella-Guitiérrez GC: TP: Documentation - Data request flow proposal: Milestone 7 Coordination - V1.1 BSC @ GDI Spanish Node 37/37 27/Aug/2024 - v3.0 Proposed GDI Access Workflow Figure 1: Proposed GDI / 1+MG Data Access workflow