Full text
AFR AMR EAS EUR SAS HLA-A*33:03 HLA-B*53:01 HLA-C*08:01 HLA-C*03:02 HLA-B*40:06 HLA-A*02:11 HLA-C*07:18 HLA-A*74:01 HLA-C*07:06 HLA-A*02:02 HLA-C*15:05 HLA-A*34:02 HLA-A*36:01 HLA-A*33:01 HLA-B*38:02 Present in 2 Multiple Allele cell lines Not present in any part of the training data Present in 2 Multiple Allele cell lines Not present in any part of the training data Not present in any part of the training data Not present in any part of the training data Not present in any part of the training data Not present in any part of the training data Present in 1 Multiple Allele cell line Not present in any part of the training data Present in 1 Multiple Allele cell line Not present in any part of the training data Not present in any part of the training data Present in 3 Multiple Allele cell lines Not present in any part of the training data HLA-A*31:01 HLA-B*35:01 HLA-C*08:02 HLA-C*03:03 HLA-B*40:02 HLA-A*02:01 HLA-C*07:02 HLA-A*31:01 HLA-C*07:02 HLA-A*02:01 HLA-C*15:02 HLA-A*03:01 HLA-A*01:01 HLA-A*31:01 HLA-B*08:01 2 3 2 2 2 2 2 2 2 2 1 3 2 3 10 Q62R, E63N S77N, N80I, L81A E152T, R156L I95L, Y116S L95W, S97T T73I, H74D K66N, S99Y T9F, I73T K66N, S99Y V95L, L156W L116F F9Y, Q62R, E63N R163T, G167W Q62R, E63N, Y171H D9Y, F67C, D74Y, S77N, N80T, L81A, S97R, Y116F, D156L, A158T Superpopulations Allele Notes Polymorphic distance / polymorphisms Closest allele Represetation for closest allele Visualising the considerable bias within various immunological datasets suggests that targeted data gathering is required to redress the balance and improve the generalisability of methods trained on them. Whilst the majority of human alleles are within 3 amino acid changes of the NetMHCPan pseudosequence, many non-human alleles are at greater distances, and a significant number of HLA-A and HLA-B alleles are at higher distances. One strategy to prioritise data gathering to extend knowledge, would be to view it through the lens of polymorphic distance rather than the addition of alleles, which may be more significant in terms of frequency in underrepresented ancestries and may not add data relevant to enhancing the generalisation of the algorithms. Conclusions HLA-A Year 1 Individuals on the trial Year 2 HLA-C HLA-B Peptide count 0 50k 100k 150k 200k 250k HLA-C HLA-B HLA-A AFR AMR EAS EUR SAS HLA-B*40:10 HLA-C*04:03 HLA-C*18:01 HLA-B*07:05 HLA-B*57:04 Present in 5 Multiple Allele cell lines Not present in any part of the training data Not present in any part of the training data Not present in any part of the training data Not present in any part of the training data HLA-B*40:01 HLA-C*04:01 HLA-C*04:01 HLA-B*07:02 HLA-B*57:01 2 1 2 1 2 H9Y, T24A S9Y S9D, A24S D114N S116D, L156R Figure 3a. Alleles present in Legacy trialists, but absent from the NetMHCPan 4.1 SA EL training dataset. Figure 3b. NetMHCPan 4.1 SA LE representation for the individuals in the trial. Only one individual in Year 1 had all six alleles represented in the training data. Three individuals had only two alleles represented. Figure 3c. Loadable tetramer availability for alleles in the trial. Peptide count 0 50k 100k 150k 200k 250k Understanding ancestry biases in HLA-related datasets. Implications for machine learning predictors and healthcare equity. Christopher Thorpe and Ellen McDonagh 1. European Bioinformatics Institute, Hinxton. 1 1 AFR AMR EAS EUR SAS 0 50 100 150 200 250 300 0.5k 1k 2.5k 5k 10k 25k 50k 75k 100k 125k 150k 175k 200k 225k 250k Superpopulations Allele Count of peptides in the NetMHCPan Single Allele Elution Training Dataset Acknowledgments We would like to acknowledge the LEGACY Network for providing tissue typing data (Figure 3b) and the use case for this analysis. We also acknowledge Tim Elliott, Malcolm Sim, Benny Chain, Andreas Tiffeau-Mayer, Hashem Koohy and his group and the Qimmuno London community for valuable feedback and support on earlier versions of the work. Introduction Machine learning algorithms are mirrors of the data they are trained on. They need abundant, well-labelled, highly diverse training data, ideally with both positive and negative data. Even many state-of-the-art machine learning models generalise poorly away from their training data. Whilst some datasets, such as those for the peptide-binding predictor NetMHCPan, are collected expressly for training the algorithm, many others consist of “exhaust fume” data collected over time in scientific studies. For a variety of reasons, none of the fault of the data collectors and curators, the training data is often significantly biased towards alleles present in White western populations or those related to specific disease contexts, e.g., Abacavir sensitivity in HLAB*57:01 carriers. The bias in immunological training data sets can affect the equity of therapies based on them. Whilst predictive methods cover a wide range of alleles and specificities due to their similarity to alleles in the training data, many alleles with high polymorphic distances in the residues comprising the peptide-binding site remain unstudied. Methods Openly available data on peptides bound to HLA Class I molecules was downloaded from the websites of NetMHCPan, MHCMotifAtlas and Immune Epitope Database. Structual information was obtained from Protein Database in Europe and processed using existing pipelines. HLA Class I sequences were retrieved from the Imuno Polymorphism Database (IPD) and the HLA typing data retrieved from the One Thousand Genomes project. Information on tetramer availability was gathered from the websites of the suppliers. For each dataset the number of items relating to a particular allele were counted using a Python script. The frequency of alleles in different superpopulations was estimated using the HLA typing data from the 1000 Genomes Project. All graphs were created using MatplotLib and Seaborn. Results All studied datasets exhibit a classical “long-tail” distribution relating to the number of peptides per allele (Figure 1). The MHC Motif Atlas dataset has a shorter tail and less alleles as it is a distinct dataset comprised of individual immunopeptidomics experiments. The NetMHCPan dataset contains most of the IEDB data (before the NetMHCPan dataset generation cutoff). All of the datasets have HLAA*02:01 as the predominant allele, apart from the MHC Motif Atlas data, which has HLA-A*03:01. The structure dataset has the smallest number of records, due to the complexity of data generation, and also the largest disparity/bias, with HLA-A*02:01 making up over 40% of the data. This has significant implications for structural prediction methods. For a deeper analysis, we have focused on NetMHCPan, not to criticise that algorithm but to highlight bias in the training data, given its prevalence and deep integration into the fabric of modern-day immunological research. Out of 130 alleles in the NetMHCPan 4.1 Single Allele Elution training dataset, over half have fewer than 10k peptides. Fifty have fewer than 1000 peptides, and nearly thirty have fewer than 250 peptides. In contrast, HLA-A*02:01 has 265,252. This ARISE project has received funding from the European Union's Horizon 2020 research and innovation programme under the Marie Sklodowska-Curie grant agreement number 945405. Case study: The LEGACY Network. Flu vaccination of individuals of diverse ancestry. As part of the development of this analysis, we have been fortunate to collaborate with the LEGACY Network, which has tested a pentavalent flu vaccine with individuals of diverse ancestry. This experience has illustrated some potential shortcomings of current immunological workflows when understanding immunological responses in these individuals. The lack of previously determined epitopes in the IEDB, less confidence in NetMHCPan predictions due to low representation (Figure 3b), where non-conservative or many substitutions occur, and less readily available reagents, such as loadable tetramers (Figure 3c). Alleles present in these individuals are indicated with a small dot at the side of Figures 2a, 2d and Figure 3a. Figure 2a. The 130 alleles present in NetMHCPan 4.1 Single Allele Elution (SA EL) training dataset in the context of allele frequency information per superpopulation in the 1000 Genomes Project HLA typing. A white dot within the tile of the heatmap indicates that the allele is the predominant one in the superpopulation for that locus. The grey box indicates the top 14 alleles in the training set. Figure 1. Representation of peptides bound to a specific allele across four datasets. Figure 2d. Alleles common in 1000 Genomes, missing in NetMHCPan 4.1 SA EL training dataset. Of particular note is HLA-A*33:03, which is present at a high frequencies in East Indian (25.7%), Indonesian (29.7%) and Malaysia Patani (36%) - a similar level to HLA-A*02:01 in European and North America. Figure 2c. Number of IPD alleles at specific polymorphic distances (Hamming distance in the NetMHCPan pseudosequence) from a closest allele in the NetMHCPan 4.1 SA EL training dataset. Figure 2b. The crossover between cumulative top-N alleles and the remainder of the training dataset Top N-alleles (cumulative) Remaining alleles (cumulative) Crossover point (14) Number of Top Alleles Cumulative Peptide Count Number of Providers Peptide count 0 50k 100k 150k 200k 250k [email protected]Poster PDF Online