scieee AI-readable full text Open interactive document viewer

Experiments in generating Multiple Sequence Alignments (MSAs) for AlphaFold, which include specific co-evolutionary information for peptide:MHC complexes

Thorpe, Christopher John

Abstract

The combinatorial complexity of peptide recognition by T cells of antigenic peptides bound to MHC molecules is sufficient to make predictive methods the only realistic path to generating structural information at a repertoire scale. Since the arrival and improvement of AlphaFold Multimer, there have been repeated efforts to improve the performance of the models for immunological tasks, for which the models underperform (see What’s an MSA? panel). Whilst AlphaFold3 performs well on peptide:MHC tasks, its use is somewhat hampered by a non-commercial-use license that limits the open distribution of predictions and forbids their use as synthetic training data for developing task-specific models. AlphaFold3, unlike AlphaFold2, is also trained on bound peptide structures, making benchmarking more difficult. To improve the performance of the fully open source AlphaFold we have built a panel of custom MSAs and tested them on a reference panel of 290 complexes, making over 50,000 predictions.

Full text

Where we started… Results Best Performing Custom MSA MHC + Homologues Sequences from Uniprot Paired MHC:b2m. No peptides b2m Sequences from Uniprot Other Immunoglobulin Sequences vs. AlphaFold MSAs 36,000 lines 8MB of data 4000 lines 800KB of data 13 b2m Sequences from Uniprot ~1000 MHC Sequences from IPD * * * ~ 550 Paired pMHC:b2m ~400 highest binding affinity peptides from NetMHCPan …where we are now. Acknowledgements Thank you to Veridien for the generous donation of dedicated NVIDIA H100 GPU time, which enabled the exhaustive search phase of this project. In particular, thanks go to Sigita Barakauskaite, Natalia Jimenez Lozano, Laurent Gatineau, Ludovic Sauge and Sharon Howe from Veridien for their coordination, time and expertise, and to David Hulcoop from the EBI for facilitating the support. Introduction The combinatorial complexity of peptide recognition by T cells of antigenic peptides bound to MHC molecules is sufficient to make predictive methods the only realistic path to generating structural information at a repertoire scale. Since the arrival and improvement of AlphaFold Multimer, there have been repeated efforts to improve the performance of the models for immunological tasks, for which the models underperform (see What’s an MSA? panel). Whilst AlphaFold3 performs well on peptide:MHC tasks, its use is somewhat hampered by a non-commercialuse license that limits the open distribution of predictions and forbids their use as synthetic training data for developing task-specific models. AlphaFold3, unlike AlphaFold2, is also trained on bound peptide structures, making benchmarking more difficult. To improve the performance of the fully open source AlphaFold we have built a panel of custom MSAs and tested them on a reference panel of 290 complexes, making over 50,000 predictions. a. b. c. d. e. Peptide Information Set Peptide Information Set Peptide Information Set Peptide Information Set Peptide Information Set Interface pLDDT Peptide pLDDT Peptide side chain RMSD Interface side chain RMSD Peptide backbone RMSD Figure 3. a. Comparison of the distribution of per-component RMSDs for all predictions in an experiment, for the best-performing MSA vs the default MSA. b. Comparison of best prediction from experiment 42 vs the ground truth and equivalent ColabFold prediction. c. Comparison of the worst ColabFold prediction vs the ground truth and the equivalent ranked prediction from experiment 42. d-e. Comparison of experiment 42 predictions vs the ground truth and AlphaFold3 3BXN HLA-B*14:02 IRAAPPPLF 5VVP HLA-B*57:03 LSSPVTKSW 7LG0 HLA-B*07:02 SPRWYFYYL Key: Ground truth Experiment 42 ColabFold AlphaFold3 Key: 800 depth 1k depth 6k depth 12k depth ColabFold Experiment 42 backbone Experiment 42 sidechain Experiment 42 overall ColabFold backbone Colabfold sidechain Colabfold overall Figure 1. Illustration of the composition of the custom and default AlphaFold MSAs. The example shown is for experiment 42 which was the best perfoming MSA. Figure 3. Molecular visualisations of AlphaFold 2.3 predictions using a. a minimal test MSA comprising only the antigen binding domain and peptide sequences for HLA alleles for which we have crystal structures and b-c. the best performing MSA. d. and e. illustrate the respective pLDDT plots for these two predictions, showing the improved confidence achieved by blending information about the interplay between polymorphisms and antigens, along with evolutionary-scale coevolutionary information. In all images the ground truth experimental structure is white and the prediction is purple. Method Sequences for MHC molecules of various mammalian species were obtained from the Immuno Polymorphism Database (IPD). Only species and loci without insertions or deletions were considered for this preliminary analysis. Peptide sequences were obtained from the NetMHCPan 4.1 Binding Affinity dataset. Sequences for Beta-2 microglobulin for 13 matching species were obtained from Uniprot. Ground truth experimental structures of the complexes were obtained from the Protein Data Bank in Europe. MSAs were constructed from IPD sequences at different depths using homology-based measures. Peptide data was titrated into the MSA at different depths and using different binding affinity cutoffs - termed Peptide Information Sets in Figure 2. The localcolabfold version of AlphaFold Multimer 2.3 was used to perform the predictions. The code, including an interactive notebook for generating the MSAs, running predictions, and the dataset, will be available upon publication. What’s an MSA? Multiple Sequence Alignments (MSAs) are key inputs to the AlphaFold2 model. They provide evolutionary-scale coevolutionary information on complementary substitutions that are deemed to interact with or be close to each other in three-dimensional space. In immunological molecules, however, there is additional coevolution information to represent. In the case of MHC molecules, this is the interplay between polymorphisms in the antigen-binding site and the peptide repertoire that complements changes in the specificity pockets. [email protected]Poster PDF Online Experiments in generating Multiple Sequence Alignments (MSAs) for AlphaFold, which include specific co-evolutionary information for peptide:MHC complexes Christopher Thorpe 1. European Bioinformatics Institute, Hinxton. 1 This ARISE project has received funding from the European Union's Horizon 2020 research and innovation programme under the Marie Sklodowska-Curie grant agreement number 945405. AlphaFold schematic from Jumper et al Nature volume 596, pages583–589 (2021) Figure 2. a-c. RMSD plots (prediction accuracy) for the peptide side chains, peptide backbone and MHC peptide binding interface side chains, respectively, at different MSA depths. d-e. pLDDT plots for the peptide and the interface side chains. For each dataset, the atom-level RMSDs and residue-level pLDDTs for the five ranked predictions were averaged across 290 complexes. The average measure for an unmodified Colabfold using the latest AlphaFold Multimer 2.3 weights and defaults for MSAs is shown as a baseline. a. a. b. c. d. e. d. b. c. e.