scieee AI-readable full text Open interactive document viewer

Precision proteogenomics reveals pan-cancer impact of germline variants

Martins Rodrigues, Fernanda,Terekhanova, Nadezhda V.,Imbach, Kathleen J.,Clauser, Karl R.,Esai Selvan, Myvizhi,Porta Pardo, Eduard

Abstract

We investigate the impact of germline variants on cancer patients’ proteomes, encompassing 1,064 individuals across 10 cancer types. We introduced an approach, “precision peptidomics,” mapping 337,469 coding germline variants onto peptides from patients’ mass spectrometry data, revealing their potential impact on post-translational modifications, protein stability, allele-specific expression, and protein structure by leveraging the relevant protein databases. We identified rare pathogenic and common germline variants in cancer genes potentially affecting proteomic features, including variants altering protein abundance and structure and variants in kinases (ERBB2 and MAP2K2) impacting phosphorylation. Precision peptidome analysis predicted destabilizing events in signal-regulatory protein alpha (SIRPA) and glial fibrillary acid protein (GFAP), relevant to immunomodulation and glioblastoma diagnostics, respectively. Genome-wide association studies identified quantitative trait loci for gene expression and protein levels, spanning millions of SNPs and thousands of proteins. Polygenic risk scores correlated with distal effects from risk variants. Our findings emphasize the contribution of germline genetics to cancer heterogeneity and high-throughput precision peptidomics.

Full text

Article Precision proteogenomics reveals pan-cancer impact of germline variants Graphical abstract Highlights dPrecision peptidomics reveals germline variants’ impact on cancer patient’s proteomes dGermline variants impact PTMs and protein stability and have allele-specific effects dQTL analyses reveal genes and proteins under germline genetic control in cancer patients dPRSs reveal cumulative impact of common germline variants on cancer proteomes Authors Fernanda Martins Rodrigues, Nadezhda V. Terekhanova, Kathleen J. Imbach, ..., Eduard Porta-Pardo, Li Ding, Clinical Proteomic Tumor Analysis Consortium Correspondence [email protected] (Z.H.G.), matthew.baile[email protected] (M.H.B.), [email protected] (G.G.), [email protected] (E.P.-P.), [email protected] (L.D.) In brief Precision proteogenomics analysis of 1,064 cancer patients across ten cancer types unveils the ways in which rare and common germline variants shape the cancer proteome; the findings highlight the contribution of germline genetics in tumor heterogeneity and oncogenesis. Martins Rodrigues et al., 2025, Cell 188, 2312–2335 May 1, 2025 ª2025 The Author(s). Published by Elsevier Inc. https://doi.org/10.1016/j.cell.2025.03.026 ll Article Precision proteogenomics reveals pan-cancer impact of germline variants Fernanda Martins Rodrigues, 1,2,3,31 Nadezhda V. Terekhanova, 1,2,3,31 Kathleen J. Imbach, 4,5,31 Karl R. Clauser, 6,31 Myvizhi Esai Selvan, 7,8,9,31 Isabel Mendizabal, 10,11,12,32 Yifat Geffen, 6,13,32 Yo Akiyama, 6,32 Myranda Maynard, 6 Tomer M. Yaron, 14 Yize Li, 1,2,3 Song Cao, 1,2,3 Erik P. Storrs, 1,2,3 Olivia S. Gonda, 15 Adrian Gaite-Reguero, 10 Akshay Govindan, 1,2,3 Emily A. Kawaler, 16 Matthew A. Wyczalkowski, 1,2,3 Robert J. Klein, 7 Berk Turhan, 7 (Author list continued on next page) SUMMARY We investigate theimpact of germline variants on cancer patients’ proteomes, encompassing1,064 individuals across 10 cancer types. We introduced an approach, ‘‘precision peptidomics,’’ mapping 337,469 coding germline variants onto peptides from patients’ mass spectrometry data, revealing their potential impact on posttranslational modifications, protein stability, allele-specific expression, and protein structure by leveraging the relevant protein databases. We identified rare pathogenic and common germline variants in cancer genes potentially affecting proteomic features, including variants altering protein abundance and structure and variants in kinases (ERBB2 and MAP2K2) impacting phosphorylation. Precision peptidome analysis predicted destabilizing events in signal-regulatory protein alpha (SIRPA) and glial fibrillary acid protein (GFAP), relevant to immunomodulation and glioblastoma diagnostics, respectively. Genome-wide association studies identified quantitative trait loci for gene expression and protein levels, spanning millions of SNPs and thousands of proteins. Polygenic risk scores correlated with distal effects from risk variants. Our findings emphasize the contribution of germline genetics to cancer heterogeneity and high-throughput precision peptidomics. INTRODUCTION The germline genome of each individual person has a unique combination of millions of genetic variants that influence virtually all biological processes throughout life, including cancer evolution. Many studies have demonstrated the critical importance of germline genomics, from cancer risk assessment to the development of tailored treatments. 1 The earliest germline genomics 1 Department of Medicine, Washington University in St. Louis, Saint Louis, MO, USA 2 McDonnell Genome Institute, Washington University in St. Louis, Saint Louis, MO, USA 3 Department of Genetics, Washington University in St. Louis, St. Louis, MO 63110, USA 4 Josep Carreras Leukaemia Research Institute (IJC), Badalona, Barcelona, Spain 5 Universitat Autonoma de Barcelona, Barcelona, Spain 6 Broad Institute of MIT and Harvard, Cambridge, MA, USA 7 Department of Genetics and Genomic Sciences, Icahn School of Medicine at Mount Sinai, New York, NY, USA 8 Center for Thoracic Oncology, Tisch Cancer Institute, Icahn School of Medicine at Mount Sinai, New York, NY, USA 9 Precision Immunology Institute, Icahn School of Medicine at Mount Sinai, New York, NY, USA 10 Center for Cooperative Research in Biosciences (CIC bioGUNE), Basque Research and Technology Alliance (BRTA), Bizkaia Technology Park, Derio, Spain 11 Ikerbasque, Basque Foundation for Science, Bilbao, Spain 12 Translational Prostate Cancer Research Lab, CIC bioGUNE-Basurto, Biocruces Bizkaia Health Research Institute, Derio, Spain 13 Cancer Center and Department of Pathology, Massachusetts General Hospital, Boston, MA, USA 14 Meyer Cancer Center, Department of Medicine, Department of Physiology and Biophysics, Weill Cornell Medicine, New York, NY, USA 15 Department of Biology, Brigham Young University, Salt Lake City, UT, USA 16 Applied Bioinformatics Laboratories, New York University Langone Health, New York City, NY, USA 17 Department of Pathology, University of Michigan, Ann Arbor, MI, USA 18 Department of Computational Medicine and Bioinformatics, University of Michigan, Ann Arbor, MI, USA 19 Institute for Systems Genetics, NYU Grossman School of Medicine, New York, NY, USA 20 Department of Public Health Sciences, University of Miami Miller School of Medicine, Miami, FL, USA (Affiliations continued on next page) ll OPEN ACCESS 2312 Cell 188, 2312–2335, May 1, 2025 ª2025 The Author(s). Published by Elsevier Inc. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). studies of cancer-prone families identified highly penetrant risk genes. 2–6 These targeted-gene and linkage studies were followed by array-based genome-wide association studies (GWASs), which identified many common variants (minor allele frequency [MAF] R1%) associated with tissue-specific 7,8 or pan-cancer risk. 9 While common risk variants typically have small effect sizes when seen individually, they discriminate individuals at high risk when combined as polygenic risk scores (PRSs). 10,11 Furthermore, many common germline variants regulate proximal and distal expression of genes in specific tissues and tumors with potentially additive effects. 12,13 With advances in sequencing technologies, it has become feasible to identify rare and low-frequency variants (MAF < 1%) with moderate to high penetrance, associated with tissue-specific 14,15 or overall cancer risk, 16–18 mechanisms and pathways in tumor development, 19 tumor immune microenvironment, 20,21 mutational burden, 22,23 mutational signatures, 24,25 loss of heterozygosity (LOH), 19 and clinical variables such as age of cancer onset 16,22 and survival. 26 However, the impact of germline variants on the cancer proteome and post-translational modification (PTM) landscapes is poorly understood, specifically on oncogenic signaling pathways and their impact on cancer formation and evolution. We analyzed the pan-cancer Clinical Proteomic Tumor Analysis Consortium (CPTAC) datasets from genomic, transcriptomic, proteomic, acetylomic, and phosphoproteomic analytes to generate precision proteogenomic profiles. These datasets provide a unique resource to study the impact of germline genomics on molecular oncogenic processes. Integrative multi-omic analyses revealed new putative pathogenic (P) rare germline variants in cancer predisposition genes (CPGs). Furthermore, common variants in these genes were associated with reduced levels of tumor suppressors in both primary tumors and normal adjacent tissues (NATs). Additionally, common germline variants at specific protein phosphorylation and acetylation sites influenced phosphorylation and acetylation levels or resulted in the emergence of new PTM sites. Our precision peptidomics data also identified allele-specific protein and PTM (ASP) effects and germline indels associated with protein stabilization, destabilization, or alternative products. Finally, whole-genome sequencing (WGS) and quantitative trait loci (QTL) analyses identified common variants affecting protein expression levels in normal and tumor tissues, impacting cancer-associated pathways. Our results highlight the power of integrative multi-omic approaches to illuminate the impact of germline variants across cancer phenotypes, revealing important biological insights into the role of germline genomics. These findings suggest that precision proteogenomics could inform patient risk stratification and prevention and interception approaches. RESULTS Precision peptidomic and PTM analysis of coding germline variants CPTAC provides a proteogenomic dataset that includes common and rare germline variants across 10 cancer types. We processed and analyzed proteogenomic, clinical, and demographic data from 1,064 prospectively collected tumor and matching blood samples, including: whole-exome sequencing (WES), RNA sequencing (RNA-seq), proteome, and phosphoproteome data from all 10 cancer types; WGS from seven cancer types; and, acetylome data from six (Figure 1A). CPTAC also includes proteogenomic data from paired NATs from eight cancer types (n= 548/1,064 cases). All 1,064 matching blood samples passed WES quality control criteria and were used for germline variant calling, with average coverage ranging between 1053and 3573across target regions, and overall coverage of 253–2803across a prioritized list of 160 CPGs (STAR Methods; Figures S1A and S1B; Table S1). Karsten Krug, 6 D.R. Mani, 6 Felipe da Veiga Leprevost, 17 Alexey I. Nesvizhskii, 17,18 Steven A. Carr, 6 David Fenyo ¨, 19 Michael A. Gillette, 6 Antonio Colaprico, 20,21 Antonio Iavarone, 21,22 Ana I. Robles, 23 Kuan-lin Huang, 7,24 Chandan Kumar-Sinha, 17,25 Franc¸ ois Aguet, 6 Alexander J. Lazar, 26 Lewis C. Cantley, 27 Urko M. Marigorta, 10,11 Zeynep H. Gu¨mu¨sx, 7,8,9, *Matthew H. Bailey, 15, *Gad Getz, 6,13,28, *Eduard Porta-Pardo, 4,29, *Li Ding, 1,2,3,30,33, *and Clinical Proteomic Tumor Analysis Consortium 21 Sylvester Comprehensive Cancer Center, University of Miami Miller School of Medicine, Miami, FL, USA 22 Department of Neurological Surgery, Department of Biochemistry and Molecular Biology, University of Miami, Miller School of Medicine, Miami, FL, USA 23 Office of Cancer Clinical Proteomics Research, National Cancer Institute, Rockville, MD, USA 24 Center for Transformative Disease Modeling, Tisch Cancer Institute, Icahn Institute for Data Science and Genomic Technology, Icahn School of Medicine at Mount Sinai, New York, NY, USA 25 Michigan Center for Translational Pathology, University of Michigan, Ann Arbor, MI, USA 26 Departments of Pathology and Genomic Medicine, The University of Texas MD Anderson Cancer Center, Houston, TX, USA 27 Dana Farber Cancer Institute, Boston, MA, USA 28 Harvard Medical School, Boston, MA, USA 29 Barcelona Supercomputing Center (BSC), Barcelona, Spain 30 Siteman Cancer Center, Washington University in St. Louis, Saint Louis, MO, USA 31 These authors contributed equally 32 These authors contributed equally 33 Lead contact *Correspondence: zeynep.g[email protected] (Z.H.G.), matthew.bai[email protected] (M.H.B.), [email protected] (G.G.), eporta@ carrerasresearch.org (E.P.-P.), [email protected] (L.D.) https://doi.org/10.1016/j.cell.2025.03.026 ll OPEN ACCESS Cell 188, 2312–2335, May 1, 2025 2313 Article A total of 185,724,997 germline variants were called from WES data (STAR Methods). Variants were filtered and annotated, resulting in 27,104,152 germline variant calls (563,036 unique variants) in exonic regions (25,474 variants per sample). Individuals of African genetic ancestry (AFR) showed the highest average number of exonic germline variants per individual (30,510), with the lowest being for those of European (EUR) ancestry (25,205; Figure S1C). The germline exomes exhibited an average transition-transversion (TiTv) ratio of 2.74% and >99% concordance with dbSNP (Figure S1D). We derived ancestral (ANC) status information for 27,104,152 exonic germline variants (STAR Methods). Throughout this manuscript, we refer to individual variants in terms of ANC or derived (DER) alleles instead of major and minor alleles, respectively, according to their ANC status. We also characterized the impact of non-coding variants on gene expression and protein abundance in 779 CPTAC samples from seven cancer types for which WGS data were available (STAR Methods). Given the low-pass nature of our WGS dataset, we phased and imputed the genotypes using GLIMPSE 27 using a set of high-quality variants from 2,504 unrelated samples from Phase 3 of the 1,000 Genomes Project, which were resequenced to high coverage by the New York Genome Center (NYGC). 28 For quality control, variants called from WGS and WES were compared for the seven cancer types (STAR Methods). Overall, 94.6% of variants overlapping between WES and WGS had the same genotypes in the same samples (Figure S1E; Table S1D). WGS data was also used to confirm genetic ancestry predictions obtained from WES. Ancestry was first predicted from WES using a random forest classifier for all individuals (STAR Methods), while WGS was used to refine ancestry for 9 individuals of Slavic origin in the glioblastoma (GBM), head and neck squamous cell carcinoma (HNSCC), lung squamous cell carcinoma (LSCC), pancreatic ductal adenocarcinoma (PDAC), and A BC Figure 1. CPTAC dataset overview and precision peptidomics workflow (A) The CPTAC cohort of 1,064 individuals of different genetic ancestries across 10 cancer types and available data types. Colors in top distribution represent genetic ancestry: African (AFR); admixed American (AMR); East Asian (EAS); European (EUR); South Asian (SAS). (B) Our precision peptidomics workflow, representing the implementation of the Spectrum Mill workflow on the LC-MS/MS datasets to yield peptide spectrum matches (PSMs) that detected 18,599 germline variants in the proteome, phosphoproteome, and acetylome datasets. (C) Overview of phosphorylation (upper) and acetylation (lower) sites affected by germline variants across cancer types based on the precision peptide data. Variants occur nearby or directly at the site, with 78% of phosphosites and 84% of acetylsites having germline variants located at 10 or fewer amino acids from the PTM site. See also Figure S1. ll OPEN ACCESS 2314 Cell 188, 2312–2335, May 1, 2025 Article A C EF D B Figure 2. Impact of rare pathogenic and common germline variants on gene and protein expression (A) Schematic of filtering and classification of germline variants. Purple boxes describe the prioritization procedure for rare variants; yellow boxes show processing for common variants. (B) (Left) Distribution of rare pathogenic/likely pathogenic (P/LP) variants across 10 cancer types. (Right) Distribution of variants previously reported in any of the TCGA, gnomAD, and UKBB datasets (light blue) or novel to this study (dark blue). (C) Gene expression (x axis) and protein abundance (y axis) quantiles of proteins from P/LP variant carriers. Yellow and pink shading indicates variants with impact on protein levels and gene expression, respectively; gray denotes variants with effect in both. (legend continued on next page) ll OPEN ACCESS Cell 188, 2312–2335, May 1, 2025 2315 Article uterine corpus endometrial carcinoma (UCEC) cohorts which were misclassified in the WES-based predictions, but correctly classified as EUR using WGS (Figure S1F; STAR Methods). Next, we combined proteomics and genomics datasets to create protein sequence databases for each individual using the proteogenomic integration tool Quantitative Integrated Library of Translated SNPs/Splicing (QUILTS) 29 (STAR Methods). From the total of 185,724,997 germline variants from WES, we incorporated 337,469 unique patient-specific coding variants that mapped to Gencode v34 reference protein sequences (Figure 1B). Using these databases for each cancer cohort, the proteomics liquid chromatography-tandem mass spectrometry (LC-MS/MS) datasets were searched with the Spectrum Mill workflow (STAR Methods) to yield peptide spectrum matches (PSMs) that matched peptides from the reference proteome, germline variants, or somatic mutations. We detected peptides for 18,599 unique coding germline variants in the proteome, phosphoproteome, and acetylome, with the majority having low frequencies at the cohort level. Among the variants detected, 1,828 were in more than one dataset, while the majority were in only one: 12,330 in the proteome, 4,081 in the phosphoproteome, and 360 in the acetylome (Figure 1B). Scrutiny of the location of these variants revealed 8,046 PTM sites (7,353 phosphosites and 693 acetylation sites) affected by germline variants, 150 of which were detected across all cancers and 5,459 detected in a single cancer (Figure 1C). The pattern of high cancer-type specificity of PTM sites is consistent with our previous study 30 and suggests a role for PTMs in tissue/cell-specific regulation and signaling. Looking at the peptide-length distribution for reference and alternative alleles (Figure S1G), there is a tendency for peptides carrying the alternative allele to be longer than all reference proteome-derived peptides, regardless of the specific proteomics dataset (protein, phosphorylation, or acetylation). While the higher minimum-score thresholds employed in the subset specific false discovery rate (ssFDR) filtering of the proteome dataset to maintain suitable false discovery rate (FDR) levels will bias against shorter peptides, a random variant in a protein is more likely to occur in a peptide that spans a longer proportion of that protein. Proteogenomic modeling of rare pathogenic and common germline variants Germline variants associated with cancer likely have different P mechanisms depending on allele frequency (AF): rare P variants are oftentimes more damaging to protein function than common variants. 31,32 Here, we investigated the landscape of rare P and common germline variants in the CPTAC cohort, leveraging multi-omics information from tumor and matching NAT samples. From 27,104,152 total exonic germline variants, a minority of them were rare (1,528,083 variants; gnomAD AF %0.05%), followed by low frequency (993,176; 0.05% < gnomAD AF < 1%), and common variants (24,582,893, gnomAD AF R1%; Figure 2A). These proportions are similar to other large-scale databases of population genomics, such as UK Biobank (UKBB). Considering that rare P germline variants play important roles in cancer susceptibility, 16,18 we aimed to identify such events using CharGer 33 (STAR Methods;Figures 2A and S2). We identified 119 P and likely pathogenic (LP) variants across CPGs (Table S1) affecting 115 individuals (10.8% of the cohort; Figure 2B). The majority of P/LP variants likely represent loss-of-function events (i.e., nonsense, frameshift, start-loss, and splice-site variants; 75%, n= 89), with the remaining being missense variants predicted to be deleterious (n= 30; Table S2). These variants were also observed in other cohorts (The Cancer Genome Atlas [TCGA], 16 gnomAD, and UKBB) at extremely low frequencies (mean gnomAD AF = 0.0001, and mean UKBB AF = 0.0002). Furthermore, 34 variants (29%) were private to the CPTAC cohort (Figure 2B). We also observed that carriers were younger at diagnosis compared with non-carriers for the breast cancer (BRCA), colorectal adenocarcinoma (COAD), and clear cell renal cell carcinoma (ccRCC) cohorts (Figure S2A). To evaluate the impact of germline variants in a somatic context, we investigated LOH events using allele fractions from tumor-normal data to identify variants positively selected in the tumor based on the two-hit hypothesis 17,34,35 (STAR Methods). From 119 P/LP variants, we observed 21 (17.6%) and 11 (9.2%) variants undergoing significant (FDR %0.05) and suggestive (0.05 < FDR %0.15) LOH in the tumor, respectively (Figure S2C). For 15 of 21 (71%) significant LOH, we observed deletion of the respective gene detected by the tool Genomic Identification of Significant Targets in Cancer, version 2 (GISTIC2). 36 Also, 6 (5%) of 119 P/LP variants co-occurred with non-silent somatic mutations in the same gene. Next, we explored the molecular consequences of these 119 P/LP variants using protein and RNA expression data, focusing on 65 P/LP variants for which both RNA and protein levels were available. Consistent with a loss-of-function phenotype, P/LP variant carriers displayed lower RNA expression and protein levels (within-cancer-type quantile means of 0.36 and 0.29, respectively, compared with 0.5 for the entire cohort; Figures 2C and S2D). This was observed for variants affecting members of the mismatch repair (MMR) pathway (PMS2, MSH2, and MSH6) associated with low RNA expression and protein levels (expression quantiles < 0.25). We observed that 4 of 5 carriers of P/LP variants in those genes (MSH2 L277*, MSH6 E744fs, MSH2 Q518*, and PMS2 I611fs) were also identified as microsatellite instability (MSI)-high samples (Table S1 from Li et al. 37 ), consistent with the fact that carriers of P/LP variants (D) Effects of common germline variants in cancer genes in their protein abundance (y axis) and RNA expression (x axis). Effect is calculated as the slope in the regression model. Dot size reflects the log 10 of the FDR adjusted pvalues from the regression model. (E) Protein levels (y axis) in NAT or tumor samples for ATM, SDHA, and ERCC2 according to genotype (x axis). pvalues from pairwise Wilcoxon tests between genotype groups are provided, and data are represented as median and interquartile range. (F) Mapping of the ERCC2 K751 position on the PDB: 6RO4. Blue represents the residue, gray denotes ERCC2, pink represents ERCC3, orange highlights the DNA molecule. See also Figure S2. ll OPEN ACCESS 2316 Cell 188, 2312–2335, May 1, 2025 Article A B EF D C Figure 3. Impact of missense germline variants on PTM sites based on linear distances (A) Depiction of how missense variants may impact PTM sites based on linear distances: direct hit (colocalizes with PTM site); proximal (within 5 amino acids); or distal (located beyond 5 amino acids). Created in Biorender. (B) Direct hits are classified based on their consequence: loss, change, and gain. Created in Biorender. (C) Distribution of direct, proximal, and distal events detected in CPTAC, beside a bar plot summarizing the distribution of direct hits across the top 30 cancerrelated genes. (D) Significance (y axis) and effect (x axis) of direct-hit events on global protein levels in NATs (left) and tumor samples (right) from linear model results. Points are colored by variant consequence and shaped according to PTM type. Effect (x axis) is calculated as the slope in the regression model, and y axis reflects the log 10 of the FDR adjusted pvalues from model. (E) Significance (y axis) and effect (x axis) of proximal or distal variants in cancer-related genes on phosphorylation and acetylation levels of their corresponding PTM sites in NATs (left) and tumors (right). (Top and bottom) Results from rare/low frequency (gnomAD AF < 1%) and common variants (gnomAD AF R1%), respectively. Colors represent variant distance to the PTM sites; shapes represent PTM type; and sizes represent the frequency of the event in the CPTAC cohort (pan-cancer level). Only events for which protein abundance differences were not observed are labeled. Events in HLA-A and HLA-B were removed from common (legend continued on next page) ll OPEN ACCESS Cell 188, 2312–2335, May 1, 2025 2317 Article in core MMR pathway genes tend to develop an MSI cancer phenotype. 38 Most variants had comparable quantiles of both gene and protein expression (rho = 0.49, p= 3.08 310 5 ;Figure 2C). Interestingly, we also observed outliers, including TP53 M1I, ERCC2 A717G, and ATM L1283fs, which were associated with high RNA expression but low protein abundance of the respective genes, highlighting the importance of proteomics to assess the functional impact of variants. Next, we explored the potential effects of common germline variants (gnomAD AF R1%) in our list of 160 CPGs and 299 cancer driver genes 37,39 (Figure 2D). We observed variants in ATM, SDHA, and ERCC2 with no detectable effect on RNA expression, but lower protein levels in carriers of the DER alleles in tumor and matched NAT samples (Figure 2E). ERCC2 K751Q has been associated with lower DNA-repair activity in vitro and better outcomes in patients treated with chemotherapy, 40,41 consistent with the DER allele lowering DNA-repair efficiency. A structural alignment of the AlphaFold2 model for ERCC2 (Protein Data Bank [PDB]: 6RO4) suggests that K751 could sit at the binding interface between ERCC2 and ERCC3 (Figure 2F). This, together with previous in vitro and clinical data, and lower protein levels, suggests that the DER allele may damage the stability of the complex. Further experiments are needed to validate this hypothesis. In conclusion, the overall lower protein levels for core proteins of the DNA-repair machinery suggest that, even if these are common variants and with no detectable effects at the RNA level, they could potentially have important clinical impacts. Direct, proximal, and distal effects of germline variants on PTM sites Germline variants may mediate cancer risk through dysregulation of signaling pathways. 42,43 For example, variants might change a PTM site to abrogate its ability to become phosphorylated or acetylated 44–48 or alter the motifs that make it recognizable by enzymes, making it more or less likely to become modified. We explored the impact of rare/low frequency (gnomAD AF < 1%) and common (gnomAD AF R1%) missense variants co-localizing, proximal (within 5 amino acids), or distal (beyond 5 amino acids) to PTM sites at the linear distance (Figure 3A). For germline variants directly overlapping a PTM site, three scenarios were assessed: (1) loss of a PTM site; (2) creation of a new site; or, in the case of phosphorylations, (3) change of the substrate, e.g., serine to threonine (Figure 3B). To focus on protein-coding variants, we evaluated missense germline variants from WES to identify reference peptides in the (phopsho/ acetyl)proteomics datasets with the matching amino acids for both alleles across the entire cohort (STAR Methods). We observed 532,142 proximal, distal, and direct-hit events involving single phosphorylation sites and 42,014 events involving acetylation sites. Of these, 1,706 variants directly overlapped a site, 4,660 were proximal, and 567,790 were distal to a site on the same protein (Figure 3C; Table S3). Most PTM-related genetic variants (92.6%) were associated with phosphorylation rather than acetylation sites, reflecting the higher abundance of phosphorylation PTMs in our dataset (Figures 1A and 3C). Regarding variants overlapping a PTM site, PTM losses were the most frequent events: 1,578 losses detected across all proteins, compared to 120 gain and 8 changes (Figure 3C). Of these, we observe 115 loss and 5 gain events across the lists of 160 CPGs (Table S1), 299 cancer driver genes, 37,39 and 624 other cancer genes 17 including ATRX,BRCA1,TP53BP1, and PARP4 (Figure 3C; Table S3). Samples with variants affecting PTMs in these proteins displayed differences in protein abundance compared with those with reference alleles (STAR Methods). Specifically, 16 proteins with variants located at a PTM site exhibited significant dysregulation in NATs, of which 14 were also observed in tumors (generalized linear model [GLM] FDR %0.05; Figure 3D; Table S3). For example, we noted a small but statistically significant increase in the level of DEP containing MTOR interacting protein (DEPTOR) in the presence of the S389N phosphosite loss allele. DEPTOR is associated with suppression of the mechanistic target of rapamycyin kinase (mTOR) complexes 1/2 (mTORC1/2), 49 and the S389N variant (rs4871827, gnomAD AF = 0.33) is at the interface between DEPTOR and mTOR. 50 To understand whether this variant has broader mTOR pathway effects, we tested for changes in proteins or phosphoproteins in pathway members between variant carriers and non-carriers (STAR Methods), as even modest changes in protein abundance may elicit downstream effects (Table S3). We found a slight decrease in MAP2K2 T25 phosphorylation levels in HNSCC (GLM FDR = 0.0163; Wilcoxon FDR = 0.00097 between non-carriers and heterozygous (HET) individuals; Figure S3A). In PDAC, EIF4EBP1 showed decreased phosphorylation at S83/S101 and T36/T37 (GLM FDR = 0.02 and 0.027, respectively). Moreover, patients homozygous for the DER allele of DEPTOR S389N showed the lowest phosphorylation levels at both EIF4EBP1 sites (Wilcoxon FDR = 0.018 and 0.036, respectively; Figures S3B and S3C; Table S3). The T37 site in EIF4EBP1 is involved in hyperphosphorylation-dependent disruption of eIF4E binding. 51 Beyond DEPTOR, several other PTM-overlapping variants showed associations with phosphorylation levels of pathway members, including ERBB2 P1170A on PAK1 S220s/T225t phosphorylation in the pan-cancer cohort (GLM FDR = 0.043; Figure S3D), HLA-B V69A on HSP90AA1 S763s phosphorylation in BRCA (GLM FDR = 0.005; Figure S3E), and CASP8 D344H on SEPTIN4 S605s phosphorylation in GBM (GLM FDR = 0.048; Figure S3F). Next, we quantified the association of proximal or distal variants with phosphorylation/acetylation abundance differences on reference peptides at the pan-cancer level (STAR Methods). For rare/low-frequency variants, to increase statistical power, we collapsed all individuals harboring a proximal or distal variant into a single variable at the gene level (STAR Methods). We identified 46 variants associated with phosphorylation and variant results (bottom) (see Table S3D for complete list of tested events and Table S3E for protein abundance differences results). Effect (x axis) is calculated as the slope in the regression model, and y axis reflects the log 10 of the FDR adjusted pvalues from the model. (F) PTM levels according to patient genotype status for variants proximal (top) and distal (bottom) to the sites. FDR adjusted pvalues from pairwise Wilcoxon tests between genotype groups are provided, and data are represented as median and interquartile range. See also Figure S3. ll OPEN ACCESS 2318 Cell 188, 2312–2335, May 1, 2025 Article A C E D B Figure 4. Spatially interacting missense germline variants, somatic mutations, and PTM sites (A) Depiction of how missense germline variants and somatic mutations may interact with PTM sites based on spatial distances, showing an overview of HotPho analyses, which map input mutations and PTM sites onto protein structures. Created in Biorender. (B) HotPho pipeline. Created in Biorender. (C) (Left) Number of intramolecular hybrid clusters in cancer-related proteins detected in AFDB and PDB (inner). (Right) Number of germline variants and somatic mutations in each hybrid cluster that are directly overlapping a PTM site in the same cluster at a linear distance (direct), within 5 amino acids (proximal), or beyond 5 amino acids (distal). (D) Protein level differences in samples involved in hybrid clusters detected in AlphaFoldDB structures vs. not. Dots represent a cluster, where color depicts its type based on involved events. AFDB cluster ID is shown beside protein names. Effect (x axis) is calculated as the slope in the regression model, and y axis reflects the log 10 of the FDR adjusted pvalues from the model. (legend continued on next page) ll OPEN ACCESS Cell 188, 2312–2335, May 1, 2025 2319 Article OAS1, CARD8, CASP7, and ITIH1 (Figure S6). We highlight a novel association between a common variant in the 30UTR and protein expression of the glial fibrillary acid protein (GFAP) in GBM tumors (Figure 6E, top x axis). No association is detectable between indel carriers and non-carriers at RNA expression level, but there is a drastic shift in protein abundance (Figure 6E, right y axis, Welch’s t test pvalue = 1.323 310 8 ). This association found in the UTR of GFAP has not been previously reported, probably because of its lack of influence on RNA levels and scarce proteomics data in this disease. GFAP is a critical GBM biomarker, 145 and a ‘‘promising therapeutic target’’ 146 . A breakdown of carriers and non-carriers at the peptide level indicated an increase in the abundance and stability of the entire protein (Figure 6F). Furthermore, miRWalk, 147 an miRNA binding prediction tool, suggests strong binding of miR-137 at this indel site. Collectively, these analyses underscore the utility of multi-omic integration in linking genomic, expression, and proteomic changes to cancer mechanisms. Omics-wide association of common germline variants and ANC variants with proteomics impacts Most germline variations occur in non-coding regions of the genome, which regulate cellular processes. To characterize their regulatory impact on gene expression and protein abundance, we performed QTL analyses. Germline variant calling was performed on blood-derived WGS samples followed by imputation using the NYGC 1,000 Genomes Project genome. 28 QTLs affecting transcript (eQTL) and protein (pQTL) abundance were mapped in NAT and tumors across ccRCC, HNSCC, LSCC, LUAD, and PDAC (Figures 7A and 7B; Table S7A). We observed that the expression levels of 5% and 10% of the total tested genes (eGenes) and the abundance levels of 4% and 5% total tested proteins (pProteins) were associated with WGS germline variants in tumor and NAT, respectively (Table S7). Furthermore, 12% and 15% of pQTLs were also eQTLs in tumor and NAT, respectively. Pan-cancer analysis identified 237 eGenes and 47 pProteins that were shared across all NATs and cancers we studied, suggesting cross-tissue QTLs (Figure 7C). Interestingly, ERAP2, HLA-DQB1, and PPIL3 were under germline genetic control across all NATs and cancers at both gene expression and protein levels. To determine whether the causal genetic variant was the same for transcript expression and protein abundance for each of these three genes, we conducted a Bayesian test for colocalization of all eQTL-pQTL cis-pairs. We discovered that for ERAP2 (Figure 7D), the same variant drives both eQTLs and pQTLs (Figure 7E). We also show the effect of the lowest pvalue cis-SNP (rs2927608) on ERAP2 in Figure 7D. Similarly, in HLADQB1 and PPIL3, we observed that eQTLs and pQTLs shared the same causal variants in most of the NAT and tumor tissues (Table S7F). As a positive control, we compared the cis-eQTLs of NAT and tumor tissues of LSCC and LUAD with normal lung eQTL data from the Genotype-Tissue Expression (GTEx) Consortium (Figure 7F; Table S7G), showing that 60% of eQTLs (65% eGenes) and 50% of eQTLs (60% eGenes) in NAT and tumors, respectively, were also identified in GTEx at 1% FDR. Furthermore, >95% of common cis-eQTLs had the same allelic effects (beta direction) in both lung GTEx and our lung NAT. Given their prevalence across tissues and -omics datasets, and their role in disease risk in other immune related diseases, we tested whether the expressions of ERAP2,HLA-DQB1, and PPIL3 correlated with patient survival. Indeed, the expression of ERAP2 and HLA-DQB1 was positively associated with overall survival in HNSCC. Note that 109 of 110 HNSCC individuals were HPV-negative. Furthermore, we observed the same trend in the TCGA HNSCC cohort (Figure S7A). We calculated PRSs using variants discovered through prior GWAS to evaluate the global impact of personal risk in CPTAC participants (Table S7H). For GBM, LSCC, and PDAC, PRSs were associated with cancer diagnosis as compared with other cancer types in CPTAC and healthy controls from UKBB (Figures 7G and S7B). PRSs also stratified patients by disease aggressiveness, as indicated by disease recurrence and overall survival rates in PDAC (Figure 7H; same patterns observed for LSCC). Considering the potential of PRSs, and that most risk variants from GWAS are non-coding, we characterized their regulatory impact on the tumor proteome. We modeled the effect of PRSs on protein abundance while controlling for clinical, demographic, and molecular covariates. We observed few proteins Figure 7. eQTL, pQTL, and polygenic risk assessment of samples with tumor and normal WGS (A) Shared number of eGenes (genes with significant eQTLs). (B) pProteins (proteins with significant pQTLs) across NAT and tumor tissues of different cancer types indicated by an UpSetR plot 148 (top 40). Asterisks indicate that eQTL analysis was not performed for normals due to the limited number of samples. (C) Intersection of eGenes and pProteins from a Pan-CPTAC (CCRCC, HNSCC, LSCC, and LUAD) comparison across NAT and tumors. (D) All pvalues of cis-eQTLs and -pQTLs associated with ERAP2 in LUAD tumor samples. Plot insets highlight the effect of rs2927608 alleles on ERAP2 RNA expression (top) and protein abundance (bottom). (E) Colocalization results of eQTLs and pQTLs in ERAP2 across NAT and tumors of different cancer types (PP: posterior probabilities supporting each hypothesis; H0: no causal variant; H1: causal variant for RNA expression only; H2: causal variant for protein abundance only; H3: distinct causal variants; H4: common causal variants for eQTL and pQTL). (F) Comparison of beta coefficients of common cis-eQTLs at 1% FDR between GTEx lung and CPTAC LSCC NAT (top) and LUAD NAT (bottom). (G) Literature-based polygenic risk scores (PRSs) calculated on CPTAC PDAC samples and compared with orthogonal datasets. Data are represented as median and interquartile range. pvalues for statistical significance for the comparison against ‘‘Cancer CPTAC’’ and ‘‘Controls UK Biobank’’ are provided (t test). (H) Kaplan-Myer plots estimating recurrence free survival (top) and overall survival (bottom) for samples with high and low PRSs. (I) Protein abundance changes that correlate with PRS, highlighting that proximal genes (magenta) change less than distal genes. (J) GSEA shows that high PRS samples are enriched for genes in the adaptive immune system and the RAF/MAPK cascades. (K) The gnomAD ANC allele frequency (AF) for our top findings, separated by the section in which they are described. Top annotations show overall gnomAD AF separated by rare and common germline variants (left: gnomAD AF %0.05%, right: gnomAD AF > 0.05%). y axis displays the ANC population. See also Figure S7. ll OPEN ACCESS 2326 Cell 188, 2312–2335, May 1, 2025 Article associated with PRS (Figure 7I), implying limited impact at the single-protein level in CPTAC. However, a pathway-based approximation of these results with gene-set enrichment analyses (GSEAs) showed significant overrepresentation of several biological processes (Figure 7J; Table S7I), suggesting that genetic risk has a cumulative impact that converges in certain biological processes rather than large alterations in specific proteins. Antigen presentation was among the top pathways associated with common risk for PDAC, consistent with its high heritability estimated by pan-cancer immunity studies, 20 in addition to platelet function 149 and L1 cell adhesion molecule (L1CAM) related neural microenvironment remodeling. 150 Common variants also impacted protein levels of the RAS/MAPK pathway, which is mutated in 96% of pancreatic ductal tumors. 151 We also examined whether variants in this study vary in prevalence across genetic ancestries. While our analyses accounted for ancestry as a covariate (STAR Methods), we recognize that some variants may differ in frequency among individuals from different genetic backgrounds. To explore this, we selected 150 statistically significant variants from our analyses and compared their ancestry-specific AF using gnomAD for the groups relevant to CPTAC: admixed-American (AMR), East Asian (EAS), non-Finnish European (NFE), and South Asian (SAS). We observed some variants with varying AFs among the five ancestry groups, while others showed consistent AF across all groups (Figure 7K). For instance, the truncating SIRPA indel is more common in EAS individuals, while the CHD4 E139D variant exhibiting strong ASE is more frequent in AFR individuals. In contrast, variants like the top SNP from QTL analysis for HLADQB1 (rs9273472) and CASP8 D344H, which influenced a distal phosphorylation site, showed similar AFs across all ancestries in gnomAD. DISCUSSION While most cancer genomics studies have focused on the role of somatic mutations, the number of germline variants greatly exceed that of somatic mutations in a cancer cell. The composition of these variants is unique, and their effects in oncogenic processes and cancer evolution remain poorly understood. We have leveraged the CPTAC cohort with multiple cancer types to explore the impact of germline variations on cancer-relevant genes through multiple-omics layers: from DNA to RNA, protein abundance, and PTM. To assess the effects of coding variants and their association with cognate proteins (and PTMs), we used precision peptidomics, i.e., the quantification of peptides carrying genetic variants from individual patients. Integrating bulk proteomic and transcriptomic data with germline variants, we derived mechanistic inferences on the effects of coding variants. Point mutations at or near phosphosites altering downstream biological processes were noted in both tumor and NAT samples. Similar regulatory mechanisms are seen for mutations far from phosphosites in linear distance. We have highlighted examples where a distal linear effect is likely caused by the genetic variants and the PTM sites being close in 3D, benefiting from predicted 3D models by AlphaFold2. We are mindful that those models are imperfect, particularly regarding the relative spatial arrangement of different domains within the same protein. 152 Finally, we also show that germline indels can shape peptide and protein abundance through effects that cannot be discerned at the RNA level. We explored the impact of non-coding variants on both gene expression and protein abundance (QTL analyses), reporting genes and proteins under germline genetic control across different NATs and tumors (https://immuneregulation.mssm. edu). Comparison of our lung NAT eQTLs with lung eQTLs from GTEx showed an extensive overlap, validating our approach. Beyond highlighted genes from colocalization and survival analyses, there are additional tissue-specific or multicancer eGenes and pProteins that merit further investigation. In recent years, large consortia like GTEx have generated genome-wide catalogs of regulatory effects that were critical in understanding the molecular consequences of germline loci identified by GWAS. 153 Here, we provide a pan-tissue catalog of matched gene expression and protein abundance in tumors and NATs that expands such efforts. We also observed that the collective effect of known GWAS risk variants in PDAC, measured as PRS, correlated better with protein levels within oncogenic pathways that are distal to the loci that are part of the PRS. These results suggest that, on top of their local impact in cis, GWAS loci can collectively alter global proteomic regulation in trans. Despite the case-control design of the cancer discovery GWASs performed to date, our results confirm that a PRS can stratify patients according to disease aggressiveness and overall survival rates. 10 These findings underscore the value of proteogenomics in interpreting germline variant effects on cancer phenotypes and clinical outcomes. Finally, genetic ancestry might influence the effects of germline variants. 154,155 While diverse, spanning five key genetic ancestries—EUR (n=786),AFR(n=40),EAS(n=194),SAS(n=5), and AMR (n= 39)—the CPTAC cohort remains underpowered for discovery of novel contributors to cancer phenotypes for specific genetic ancestries other than EUR. Also, our cohort is relatively small compared with larger genomic studies. 39,156–158 Despite this limitation, we uncovered ancestry-independent associations of proteomic, phospho-proteomic, and transcriptomic variations by accounting for genetic ancestry in our analyses. In conclusion, the germline genome is the fundamental arena where the drama of cancer unfolds and is depicted. Amid mutational chaos, the germline plays a critical role that can enable or constrain the evolution of cancer, dictating the odds of many clinically relevant phenomena: from cancer driver mutations 11,159 to immune responses against cancer cells. 20 A deeper understanding, afforded by proteomics, illuminates this complexity, unveiling altered protein function as pivotal in carcinogenesis. Limitations of the study While our dataset is one of the largest multi-omic resources available, we remain underpowered due to sample size. Our cohort included patients predominantly of EUR genetic ancestry, with smaller subsets of other ancestries. Future proteogenomic studies need to include more diverse populations. All -omics datasets were from bulk analytes, limiting our ability to resolve impacts of germline variants on specific cell types. We only used the common variants imputed from the 1,000 Genomes ll OPEN ACCESS Cell 188, 2312–2335, May 1, 2025 2327 Article Project, 28 as we did not have high-coverage WGS data. Current proteomic pipelines rely on generic peptide references to quantify peptide abundance. We addressed this limitation by identifying personalized peptides, but single peptides reflect diverse allele frequencies from populations and our cancer-specific cohort. While protein and gene-level quantification are mitigated by aggregating many peptides, we remain conservative when addressing the impact of single peptides. AlphaFoldDB expanded our structural analysis to all human proteins, but its models are not experimentally validated. Finally, validation of our findings is challenging due to the limited availability of comparable comprehensive datasets, so some of our results will likely evolve as more samples are analyzed. RESOURCE AVAILABILITY Lead contact Further information and requests for resources and reagents should be directed to and will be fulfilled by Dr. Li Ding ([email protected]). Materials availability This study did not generate new unique reagents. Data and code availability Raw and processed proteomics as well as open-access genomic data, can be obtained via Proteomic Data Commons (PDC) at https://pdc.cancer.gov/pdc/ cptac-pancancer. Raw genomic and transcriptomic data files can be accessed via the Genomic Data Commons (GDC) Data Portal at https://portal. gdc.cancer.gov with dbGaP Study Accession: phs001287.v17.p6. Complete CPTAC Pan-Cancer controlled and processed data, including the precision proteogenomics data generated in this manuscript, can be accessed via the Cancer Data Service (CDS). The CPTAC Pan-Cancer data hosted in CDS is controlled data and can be accessed through the NCI DAC approved, dbGaP compiled whitelists. Users can access the data for analysis through the Seven Bridges Cancer Genomics Cloud (SB-CGC) which is one of the NCIfunded Cloud Resource/platform for compute intensive analysis. Instructions to access data are as follows: (1) create an account on CGC, Seven Bridges (https://cgc-accounts.sbgenomics.com/auth/register; (2) get approval from dbGaP to access the controlled study (https://www.ncbi.nlm.nih.gov/projects/ gap/cgi-bin/study.cgi?study_id=phs001287.v17.p6); (3) log into CGC to access Cancer Data Service (CDS) File Explore; (4) copy data into your own space and start analysis and exploration; (5) visit the CDS page to see what studies are available and instructions and guides to use the resources (https:// dataservice.datacommons.cancer.gov/#/data). Data used in this publication were generated by CPTAC, accessible through dbGaP accession numbers phs000892.v6.p1 (‘‘CPTAC Proteogenomic Confirmatory Study’’) and phs001287.v17.p6 (‘‘CPTAC Proteogenomic Study’’). We focused on the CPTAC samples with both genomic and proteomic data available to investigate the Pan-Cancer proteogenomic impacts of oncogenic drivers. DOIs are listed in the key resources table. Any additional information and code required to reanalyze the data reported in this paper is available from the lead contact upon request. CONSORTIA The members of the National Cancer Institute Clinical Proteomic Tumor Analysis Consortium are Eunkyung An, Meenakshi Anurag, Jasmin Bavarva, Chet Birger, Michael J. Birrer, Anna P. Calinawan, Michele Ceccarelli, Daniel W. Chan, Arul M. Chinnaiyan, Hanbyul Cho, Shrabanti Chowdhury, Marcin P. Cieslik, Daniel Cui Zhou, Corbin Day, Marcin J. Domagalski, Yongchao Dou, Brian J. Druker, Nathan Edwards, Matthew J. Ellis, Steven M. Foltz, Alicia Francis, Tania J. Gonzalez Robles, Sara J.C. Gosline, Runyu Hong, Galen Hostetter, Yingwei Hu, Tara Hiltke, Chen Huang, Emily Huntsman, Eric J. Jaehnig, Scott D. Jewell, Jiayi Ji, Wen Jiang, Lizabeth Katsnelson, Karen A. Ketchum, Iga Kolodziejczak, Jonathan T. Lei, Yuxing Liao, Caleb M. Lindgren, Tao Liu, Weiping Ma, Wilson McKerrow, Chelsea J. Newton, Robert Oldroyd, Gilbert S. Omenn, Amanda G. Paulovich, Francesca Petralia, Boris Reva, Karin D. Rodland, Henry Rodriguez, Kelly V. Ruggles, Dmitry Rykunov, Sara R. Savage, Eric E. Schadt, Michael Schnaubelt, Tobias Schraink, Zhiao Shi, Richard D. Smith, Xiaoyu Song, Yizhe Song, Jimin Tan, Ratna R. Thangudu, Nicole Tignor, Joshua M. Wang, Pei Wang, Ying Wang, Bo Wen, Maciej Wiznerowicz, Xinpei Yi, Bing Zhang, Hui Zhang, Xu Zhang, Zhen Zhang, David I. Heiman, Jared L. Johnson, Liang-Bo Wang, Lijun Yao, Mathangi Thiagarajan, Mehdi Mesri, O ¨zgu ¨n Babur, Pietro Pugliese, Qing Zhang, Samuel H. Payne, Saravana M. Dhanasekaran, Shankara Anand, Shankha Satpathy, Stephan Schu ¨rer, Vasileios Stathias, Wen-Wei Liang, Wenke Liu, and Yige Wu. ACKNOWLEDGMENTS We would like to thank the participants and investigators from the National Cancer Institute (NCI) Clinical Proteomic Tumor Analysis Consortium (CPTAC). This work was supported by NCI-CPTAC under award numbers U24CA210955, U24CA210985, U24CA210986, U24CA210954, U24CA2 10967, U24CA210972, U24CA210979, U24CA210993, U01CA214114, U01CA214116, and U01CA214125 as well as U24CA210972 (D.F., and L.D.), U24CA210979 (G.G.), U24CA270823 (M.A.G.), and contract number GR0012005 (L.D.). This work was also supported by NCI U24CA211006 and R01HG009711 to L.D. The Spanish Ministry of Science supports E.P.-P. and K.J.I. (RYC2019-026415-I and PID2019-107043RA-I00) and U.M.M. (RYC2020-030632-I and PID2019-108244RA-I00). I.M. is supported by Fundacio ´n Cris Contra el Ca ´ncer (PR_TPD_2020-19). This research was conducted using the UK Biobank Resource under application numbers 54343 and 74382 (to E.P.-P. and U.M.M., respectively). This project is funded in part with federal funds from the NCI, National Institutes of Health, under contract no. HHSN261201500003I, Task Order no. HHSN26100064. The content of this publication does not necessarily reflect the views or policies of the Department of Health and Human Services, nor does mention of trade names, commercial products, or organizations imply endorsement by the US Government. AUTHOR CONTRIBUTIONS Study conception and design, Z.H.G., E.P.-P., L.D., M.H.B., and G.G.; performed experiments or data collection, F.M.R., N.V.T., Y.L., Y.A., A.I.R., Y.G., F.d.V.L., and A.I.N.; multi-omic & statistical analyses, F.M.R., N.V.T., K.J.I., K.R.C., M.M., K.K., M.E.S., I.M., Y.G., Y.A., T.M.Y., S.C., E.P.S., Y.L., O.S.G., A.G., E.A.K., U.M.M., Z.H.G., M.H.B., E.P.-P., B.T., and R.J.K.; data interpretation & biological analysis, F.M.R., N.V.T., K.R.C., A.C., K.-l.H., C.K.-S., F.A., A.J.L., L.C.C., U.M.M., Z.H.G., M.H.B., G.G., E.P.-P., and L.D.; writing, F.M.R., N.V.T., K.J.I., K.R.C., M.E.S., I.M., Y.G., Y.A., C.K.-S., A.J.L., U.M.M., Z.H.G., D.F., M.A.W., M.H.B, G.G., E.P.-P., and L.D.; supervision, D.R.M., M.A.G., D.F., S.A.C., Z.H.G., M.H.B., G.G., E.P.-P., and L.D.; administration, G.G., A.I.R., and L.D. DECLARATION OF INTERESTS The authors declare no competing interests. STAR+METHODS Detailed methods are provided in the online version of this paper and include the following: dKEY RESOURCES TABLE dEXPERIMENTAL MODEL AND STUDY PARTICIPANT DETAILS BHuman subjects BClinical data annotation dMETHOD DETAILS BHarmonized genome alignment BGermline variant calling and filtering from WES BSomatic mutation and copy number variant calling from WES BGermline variant calling and filtering from WGS ll OPEN ACCESS 2328 Cell 188, 2312–2335, May 1, 2025 Article BComparison of WES and WGS variant calls BAncestry prediction BGene list curation for pathogenic variant classification BInference of the ancestral state of germline variants dQUANTIFICATION AND STATISTICAL ANALYSIS BPathogenicity assessment of rare germline variants BBurden testing analyses of rare P/LP germline variants BLOH analysis of rare P/LP germline variants BProteomics LC-MS/MS data interpretation BGermline Variants Co-localizing with or Around PTM sites BAnalyses of Direct, Proximal, and Distal Impact of Germline Variants on Protein and PTM Levels BHotSpot3D / HotPho analyses BAllele specific expression analysis using RNA-seq data BIndel variant analysis BIdentification of expression and protein quantitative trait loci (eQTLs and pQTLs) BPolygenic Risk Scores and associations with protein abundance dADDITIONAL RESOURCES SUPPLEMENTAL INFORMATION Supplemental information can be found online at https://doi.org/10.1016/j.cell. 2025.03.026. Received: October 9, 2023 Revised: April 29, 2024 Accepted: March 13, 2025 Published: April 14, 2025 REFERENCES 1. Robson, M., Im, S.-A., Senkus, E., Xu, B., Domchek, S.M., Masuda, N., Delaloge, S., Li, W., Tung, N., Armstrong, A., et al. (2017). Olaparib for Metastatic Breast Cancer in Patients with a Germline BRCA Mutation. N. Engl. J. Med. 377, 523–533. https://doi.org/10.1056/ NEJMoa1706450. 2. Villani, A., Tabori, U., Schiffman, J., Shlien, A., Beyene, J., Druker, H., Novokmet, A., Finlay, J., and Malkin, D. (2011). Biochemical and imaging surveillance in germline TP53 mutation carriers with Li-Fraumeni syndrome: a prospective observational study. Lancet Oncol. 12, 559–567. https://doi.org/10.1016/S1470-2045(11)70119-X. 3. Robson, M., and Offit, K. (2007). Clinical practice. Management of an inherited predisposition to breast cancer. N. Engl. J. Med. 357, 154–162. https://doi.org/10.1056/NEJMcp071286. 4. Galiatsatos, P., and Foulkes, W.D. (2006). Familial adenomatous polyposis. Am. J. Gastroenterol. 101, 385–398. https://doi.org/10.1111/j. 1572-0241.2006.00375.x. 5. Rahman, N. (2014). Realizing the promise of cancer predisposition genes. Nature 505, 302–308. https://doi.org/10.1038/nature12981. 6. Lindor, N.M., Petersen, G.M., Hadley, D.W., Kinney, A.Y., Miesfeldt, S., Lu, K.H., Lynch, P., Burke, W., and Press, N. (2006). Recommendations for the care of individuals with an inherited predisposition to Lynch syndrome: a systematic review. JAMA 296, 1507–1517. https://doi.org/10. 1001/jama.296.12.1507. 7. Hung, R.J., McKay, J.D., Gaborieau, V., Boffetta, P., Hashibe, M., Zaridze, D., Mukeria, A., Szeszenia-Dabrowska, N., Lissowska, J., Rudnai, P., et al. (2008). A susceptibility locus for lung cancer maps to nicotinic acetylcholine receptor subunit genes on 15q25. Nature 452, 633–637. https://doi.org/10.1038/nature06885. 8. Easton, D.F., Pooley, K.A., Dunning, A.M., Pharoah, P.D.P., Thompson, D., Ballinger, D.G., Struewing, J.P., Morrison, J., Field, H., Luben, R., et al. (2007). Genome-wide association study identifies novel breast cancer susceptibility loci. Nature 447, 1087–1093. https://doi.org/10.1038/ nature05887. 9. Rashkin, S.R., Graff, R.E., Kachuri, L., Thai, K.K., Alexeeff, S.E., Blatchins, M.A., Cavazos, T.B., Corley, D.A., Emami, N.C., Hoffman, J.D., et al. (2020). Pan-cancer study detects genetic risk variants and shared genetic basis in two large cohorts. Nat. Commun. 11, 4423. https://doi. org/10.1038/s41467-020-18246-6. 10. Klein, R.J., and Gu ¨mu ¨sx, Z.H. (2022). Are polygenic risk scores ready for the cancer clinic?-a perspective. Transl. Lung Cancer Res. 11, 910–919. https://doi.org/10.21037/tlcr-21-698. 11. Porta-Pardo, E., Sayaman, R., Ziv, E., and Valencia, A. (2020). The landscape of interactions between cancer polygenic risk scores and somatic alterations in cancer cells. Preprint at bioRxiv. https://doi.org/10.1101/ 2020.09.28.316851. 12. Bicak, M., Wang, X., Gao, X., Xu, X., Va ¨a ¨na ¨nen, R.-M., Taimen, P., Lilja, H., Pettersson, K., and Klein, R.J. (2020). Prostate cancer risk SNP rs10993994 is a trans-eQTL for SNHG11 mediated through MSMB. Hum. Mol. Genet. 29, 1581–1591. https://doi.org/10.1093/hmg/ ddaa026. 13. Li, Q., Stram, A., Chen, C., Kar, S., Gayther, S., Pharoah, P., Haiman, C., Stranger, B., Kraft, P., and Freedman, M.L. (2014). Expression QTLbased analyses reveal candidate causal genes and loci across five tumor types. Hum. Mol. Genet. 23, 5294–5302. https://doi.org/10.1093/hmg/ ddu228. 14. Esai Selvan, M., Zauderer, M.G., Rudin, C.M., Jones, S., Mukherjee, S., Offit, K., Onel, K., Rennert, G., Velculescu, V.E., Lipkin, S.M., et al. (2020). Inherited Rare, Deleterious Variants in ATM Increase Lung Adenocarcinoma Risk. J. Thorac. Oncol. 15, 1871–1879. https://doi.org/10.1016/j. jtho.2020.08.017. 15. Esai Selvan, M., Klein, R.J., and Gu ¨mu ¨sx, Z.H. (2019). Rare, Pathogenic Germline Variants in Fanconi Anemia Genes Increase Risk for Squamous Lung Cancer. Clin. Cancer Res. 25, 1517–1525. https://doi.org/10.1158/ 1078-0432.CCR-18-2660. 16. Huang, K.-L., Mashl, R.J., Wu, Y., Ritter, D.I., Wang, J., Oh, C., Paczkowska, M., Reynolds, S., Wyczalkowski, M.A., Oak, N., et al. (2018). Pathogenic Germline Variants in 10,389 Adult Cancers. Cell 173, 355–370.e14. https://doi.org/10.1016/j.cell.2018.03.039. 17. Lu, C., Xie, M., Wendl, M.C., Wang, J., McLellan, M.D., Leiserson, M.D.M., Huang, K.-L., Wyczalkowski, M.A., Jayasinghe, R., Banerjee, T., et al. (2015). Patterns and functional implications of rare germline variants across 12 cancer types. Nat. Commun. 6, 10086. https://doi.org/10. 1038/ncomms10086. 18. Esai Selvan, M., Onel, K., Gnjatic, S., Klein, R.J., and Gu ¨mu ¨sx, Z.H. (2023). Germline rare deleterious variant load alters cancer risk, age of onset and tumor characteristics. NPJ Precis. Oncol. 7, 13. https://doi.org/10.1038/ s41698-023-00354-3. 19. Kanchi, K.L., Johnson, K.J., Lu, C., McLellan, M.D., Leiserson, M.D.M., Wendl, M.C., Zhang, Q., Koboldt, D.C., Xie, M., Kandoth, C., et al. (2014). Integrated analysis of germline and somatic variants in ovarian cancer. Nat. Commun. 5, 3156. https://doi.org/10.1038/ncomms4156. 20. Sayaman, R.W., Saad, M., Thorsson, V., Hu, D., Hendrickx, W., Roelands, J., Porta-Pardo, E., Mokrab, Y., Farshidfar, F., Kirchhoff, T., et al. (2021). Germline genetic contribution to the immune landscape of cancer. Immunity 54, 367–386.e8. https://doi.org/10.1016/j.immuni. 2021.01.011. 21. Shahamatdar, S., He, M.X., Reyna, M.A., Gusev, A., AlDubayan, S.H., Van Allen, E.M., and Ramachandran, S. (2020). Germline Features Associated with Immune Infiltration in Solid Tumors. Cell Rep. 30, 2900– 2908.e4. https://doi.org/10.1016/j.celrep.2020.02.039. 22. Qing, T., Mohsen, H., Marczyk, M., Ye, Y., O’Meara, T., Zhao, H., Townsend, J.P., Gerstein, M., Hatzis, C., Kluger, Y., et al. (2020). Germline variant burden in cancer genes correlates with age at diagnosis and somatic mutation burden. Nat. Commun. 11, 2438. https://doi.org/10.1038/ s41467-020-16293-7. 23. Sun, X., Xue, A., Qi, T., Chen, D., Shi, D., Wu, Y., Zheng, Z., Zeng, J., and Yang, J. (2021). Tumor Mutational Burden Is Polygenic and Genetically ll OPEN ACCESS Cell 188, 2312–2335, May 1, 2025 2329 Article Associated with Complex Traits and Diseases. Cancer Res. 81, 1230– 1239. https://doi.org/10.1158/0008-5472.CAN-20-3459. 24. Liu, Y., Gusev, A., Heng, Y.J., Alexandrov, L.B., and Kraft, P. (2022). Somatic mutational profiles and germline polygenic risk scores in human cancer. Genome Med. 14, 14. https://doi.org/10.1186/s13073-02201016-y. 25. Wang, S., Pitt, J.J., Zheng, Y., Yoshimatsu, T.F., Gao, G., Sanni, A., Oluwasola, O., Ajani, M., Fitzgerald, D., Odetunde, A., et al. (2019). Germline variants and somatic mutation signatures of breast cancer across populations of African and European ancestry in the US and Nigeria. Int. J. Cancer 145, 3321–3333. https://doi.org/10.1002/ijc.32498. 26. Musa, J., Cidre-Aranaz, F., Aynaud, M.-M., Orth, M.F., Knott, M.M.L., Mirabeau, O., Mazor, G., Varon, M., Ho ¨lting, T.L.B., Grossete ˆte, S., et al. (2019). Cooperation of cancer drivers with regulatory germline variants shapes clinical outcomes. Nat. Commun. 10, 4128. https://doi.org/ 10.1038/s41467-019-12071-2. 27. Rubinacci, S., Ribeiro, D.M., Hofmeister, R.J., and Delaneau, O. (2021). Efficient phasing and imputation of low-coverage sequencing data using large reference panels. Nat. Genet. 53, 120–126. https://doi.org/10. 1038/s41588-020-00756-0. 28. Byrska-Bishop, M., Evani, U.S., Zhao, X., Basile, A.O., Abel, H.J., Regier, A.A., Corvelo, A., Clarke, W.E., Musunuri, R., Nagulapalli, K., et al. (2022). High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell 185, 3426–3440.e19. https://doi.org/10.1016/j.cell.2022.08.004. 29. Ruggles, K.V., Tang, Z., Wang, X., Grover, H., Askenazi, M., Teubl, J., Cao, S., McLellan, M.D., Clauser, K.R., Tabb, D.L., et al. (2016). An Analysis of the Sensitivity of Proteogenomic Mapping of Somatic Mutations and Novel Splicing Events in Cancer. Mol. Cell. Proteomics 15, 1060– 1071. https://doi.org/10.1074/mcp.M115.056226. 30. Geffen, Y., Anand, S., Akiyama, Y., Yaron, T.M., Song, Y., Johnson, J.L., Govindan, A., Babur, O ¨., Li, Y., Huntsman, E., et al. (2023). Pan-cancer analysis of post-translational modifications reveals shared patterns of protein regulation. Cell 186, 3945–3967.e26. https://doi.org/10.1016/j. cell.2023.07.013. 31. Momozawa, Y., and Mizukami, K. (2021). Unique roles of rare variants in the genetics of complex diseases in humans. J. Hum. Genet. 66, 11–23. https://doi.org/10.1038/s10038-020-00845-2. 32. Park, J.-H., Gail, M.H., Weinberg, C.R., Carroll, R.J., Chung, C.C., Wang, Z., Chanock, S.J., Fraumeni, J.F., and Chatterjee, N. (2011). Distribution of allele frequencies and effect sizes and their interrelationships for common genetic susceptibility variants. Proc. Natl. Acad. Sci. USA 108, 18026–18031. https://doi.org/10.1073/pnas.1114759108. 33. Scott, A.D., Huang, K.-L., Weerasinghe, A., Mashl, R.J., Gao, Q., Martins Rodrigues, F., Wyczalkowski, M.A., and Ding, L. (2019). CharGer: clinical Characterization of germline variants. Bioinformatics 35, 865–867. https://doi.org/10.1093/bioinformatics/bty649. 34. Knudson, A.G. (1971). Mutation and cancer: statistical study of retinoblastoma. Proc. Natl. Acad. Sci. USA 68, 820–823. https://doi.org/10. 1073/pnas.68.4.820. 35. Knudson, A.G. (2001). Two genetic hits (more or less) to cancer. Nat. Rev. Cancer 1, 157–162. https://doi.org/10.1038/35101031. 36. Mermel, C.H., Schumacher, S.E., Hill, B., Meyerson, M.L., Beroukhim, R., and Getz, G. (2011). GISTIC2.0 facilitates sensitive and confident localization of the targets of focal somatic copy-number alteration in human cancers. Genome Biol. 12, R41. https://doi.org/10.1186/gb-2011-124-r41. 37. Li, Y., Porta-Pardo, E., Tokheim, C., Bailey, M.H., Yaron, T.M., Stathias, V., Geffen, Y., Imbach, K.J., Cao, S., Anand, S., et al. (2023). Pan-cancer proteogenomics connects oncogenic drivers to functional states. Cell 186, 3921–3944.e25. https://doi.org/10.1016/j.cell.2023.07.014. 38. Peltoma ¨ki, P., Nystro ¨m, M., Mecklin, J.-P., and Seppa ¨la ¨, T.T. (2023). Lynch Syndrome Genetics and Clinical Implications. Gastroenterology 164, 783–799. https://doi.org/10.1053/j.gastro.2022.08.058. 39. Bailey, M.H., Tokheim, C., Porta-Pardo, E., SenGupta, S., Bertrand, D., Weerasinghe, A., Colaprico, A., Wendl, M.C., Kim, J., Reardon, B., et al. (2018). Comprehensive Characterization of Cancer Driver Genes and Mutations. Cell 173, 371–385.e18. https://doi.org/10.1016/j.cell. 2018.02.060. 40. Xiao, S., Cui, S., Lu, X., Guan, Y., Li, D., Liu, Q., Cai, Y., Jin, C., Yang, J., Wu, S., et al. (2016). The ERCC2/XPD Lys751Gln polymorphism affects DNA repair of benzo[a]pyrene induced damage, tested in an in vitro model. Toxicol. In Vitro 34, 300–308. https://doi.org/10.1016/j.tiv.2016. 04.015. 41. Zhang, G., Guan, Y., Zhao, Y., van der Straaten, T., Xiao, S., Xue, P., Zhu, G., Liu, Q., Cai, Y., Jin, C., et al. (2017). ERCC2/XPD Lys751Gln alter DNA repair efficiency of platinum-induced DNA damage through P53 pathway. Chem. Biol. Interact. 263, 55–65. https://doi.org/10.1016/j. cbi.2016.12.015. 42. Hanahan, D., and Weinberg, R.A. (2011). Hallmarks of Cancer: The Next Generation. Cell 144, 646–674. https://doi.org/10.1016/j.cell.2011. 02.013. 43. Hanahan, D. (2022). Hallmarks of Cancer: New Dimensions. Cancer Discov. 12, 31–46. https://doi.org/10.1158/2159-8290.CD-21-1059. 44. Krassowski, M., Paczkowska, M., Cullion, K., Huang, T., Dzneladze, I., Ouellette, B.F.F., Yamada, J.T., Fradet-Turcotte, A., and Reimand, J. (2018). ActiveDriverDB: human disease mutations and genome variation in post-translational modification sites of proteins. Nucleic Acids Res. 46, D901–D910. https://doi.org/10.1093/nar/gkx973. 45. Kim, Y., Kang, C., Min, B., and Yi, G.-S. (2015). Detection and analysis of disease-associated single nucleotide polymorphism influencing posttranslational modification. BMC Med. Genomics 8(Suppl 2 ), S7. https://doi.org/10.1186/1755-8794-8-S2-S7. 46. Yang, Y., Peng, X., Ying, P., Tian, J., Li, J., Ke, J., Zhu, Y., Gong, Y., Zou, D., Yang, N., et al. (2019). Awesome: a database of SNPs that affect protein post-translational modifications. Nucleic Acids Res. 47, D874–D880. https://doi.org/10.1093/nar/gky821. 47. Patrick, R., Kobe, B., Le ˆCao, K.-A., and Bode ´n, M. (2017). PhosphoPICK-SNP: quantifying the effect of amino acid variants on protein phosphorylation. Bioinformatics 33, 1773–1781. https://doi.org/10.1093/bioinformatics/btx072. 48. Huang, K.-Y., Lee, T.-Y., Kao, H.-J., Ma, C.-T., Lee, C.-C., Lin, T.-H., Chang, W.-C., and Huang, H.-D. (2019). dbPTM in 2019: exploring disease association and cross-talk of post-translational modifications. Nucleic Acids Res. 47, D298–D308. https://doi.org/10.1093/nar/gky1074. 49. Wang, Z., Zhong, J., Inuzuka, H., Gao, D., Shaik, S., Sarkar, F.H., and Wei, W. (2012). An evolving role for DEPTOR in tumor development and progression. Neoplasia 14, 368–375. https://doi.org/10.1593/ neo.12542. 50. Klen, J., Gori car, K., Horvat, S., Stojan, J., and Dol zan, V. (2019). DEPTOR polymorphisms influence late complications in type 2 diabetes patients. Pharmacogenomics 20, 879–890. https://doi.org/10.2217/pgs2019-0058. 51. Gingras, A.C., Raught, B., Gygi, S.P., Niedzwiecka, A., Miron, M., Burley, S.K., Polakiewicz, R.D., Wyslouch-Cieszynska, A., Aebersold, R., and Sonenberg, N. (2001). Hierarchical phosphorylation of the translation inhibitor 4E-BP1. Genes Dev. 15, 2852–2864. https://doi.org/10.1101/gad. 912401. 52. Saadeh, F.S., Morsi, R.Z., El-Kurdi, A., Nemer, G., Mahfouz, R., Charafeddine, M., Khoury, J., Najjar, M.W., Khoueiry, P., and Assi, H.I. (2020). Correlation of genetic alterations by whole-exome sequencing with clinical outcomes of glioblastoma patients from the Lebanese population. PLoS One 15, e0242793. https://doi.org/10.1371/journal.pone. 0242793. ll OPEN ACCESS 2330 Cell 188, 2312–2335, May 1, 2025 Article 53. von der Heyde, S., Wagner, S., Czerny, A., Nietert, M., Ludewig, F., Salinas-Riester, G., Arlt, D., and Beißbarth, T. (2015). mRNA Profiling Reveals Determinants of Trastuzumab Efficiency in HER2-Positive Breast Cancer. PLoS One 10, e0117818. https://doi.org/10.1371/journal.pone. 0117818. 54. Tong, S.Y., Ha, S.Y., Ki, K.D., Lee, J.M., Lee, S.K., Lee, K.B., Kim, M.K., Cho, C.H., and Kwon, S.Y. (2009). The effects of obesity and HER-2 polymorphisms as risk factors for endometrial cancer in Korean women. BJOG 116, 1046–1052. https://doi.org/10.1111/j.1471-0528.2009. 02186.x. 55. Hunter, D.J., Kraft, P., Jacobs, K.B., Cox, D.G., Yeager, M., Hankinson, S.E., Wacholder, S., Wang, Z., Welch, R., Hutchinson, A., et al. (2007). A genome-wide association study identifies alleles in FGFR2 associated with risk of sporadic postmenopausal breast cancer. Nat. Genet. 39, 870–874. https://doi.org/10.1038/ng2075. 56. Breyer, J.P., Sanders, M.E., Airey, D.C., Cai, Q., Yaspan, B.L., Schuyler, P.A., Dai, Q., Boulos, F., Olivares, M.G., Gao, Y.-T., et al. (2009). Heritable Variation of ERBB2 and Breast Cancer Risk. Cancer Epidemiol. Biomarkers Prev. 18, 1252–1258. https://doi.org/10.1158/1055-9965.EPI08-1202. 57. Benusiglio, P.R., Lesueur, F., Luccarini, C., Conroy, D.M., Shah, M., Easton, D.F., Day, N.E., Dunning, A.M., Pharoah, P.D., and Ponder, B.A. (2005). Common ERBB2 polymorphisms and risk of breast cancer in a white British population: a case–control study. Breast Cancer Res. 7, R204–R209. https://doi.org/10.1186/bcr982. 58. Poole, E.M., Curtin, K., Hsu, L., Kulmacz, R.J., Duggan, D.J., Makar, K.W., Xiao, L., Carlson, C.S., Slattery, M.L., Caan, B.J., et al. (2011). Genetic variability in EGFR, Src and HER2 and risk of colorectal adenoma and cancer. Int. J. Mol. Epidemiol. Genet. 2, 300–315. 59. Su, Y., Jiang, Y., Sun, S., Yin, H., Shan, M., Tao, W., Ge, X., and Pang, D. (2015). Effects of HER2 genetic polymorphisms on its protein expression in breast cancer. Cancer Epidemiol. 39, 1123–1127. https://doi.org/10. 1016/j.canep.2015.08.011. 60. Jo, U.H., Han, S.G.L., Seo, J.H., Park, K.H., Lee, J.W., Lee, H.J., Ryu, J.S., and Kim, Y.H. (2008). The genetic polymorphisms of HER-2 and the risk of lung cancer in a Korean population. BMC Cancer 8, 359. https://doi.org/10.1186/1471-2407-8-359. 61. Va ´zquez-Ibarra, K.C., Bustos-Carpinteyro, A.R., Garcı ´a-Ruvalcaba, A., Magaa ˜a-Torres, M.T., Gutie ´rrez-Aguilar, R., Marı ´n-Contreras, M.E., Santiago-Luna, E., and Sa ´nchez-Lo ´pez, J.Y. (2019). The ERBB2 gene polymorphisms rs2643194, rs2934971, and rs1058808 are associated with increased risk of gastric cancer. Braz. J. Med. Biol. Res. 52, e8379. https://doi.org/10.1590/1414-431X20198379. 62. Alanazi, I.O., Shaik, J.P., Parine, N.R., Azzam, N.A., Alharbi, O., Hawsawi, Y.M., Oyouni, A.A.A., Al-Amer, O.M., Alzahrani, F., Almadi, M.A., et al. (2021). Association of HER1 and HER2 Gene Variants in the Predisposition of Colorectal Cancer. J. Oncol. 2021, 6180337. https://doi.org/10. 1155/2021/6180337. 63. Chen, H., Zhai, Z., Xie, Q., Lai, Y., and Chen, G. (2021). Correlation between SNPs of PIK3CA, ERBB2 30UTR, and their interactions with environmental factors and the risk of epithelial ovarian cancer. J. Assist. Reprod. Genet. 38, 2631–2639. https://doi.org/10.1007/ s10815-021-02177-2. 64. Gao, Y., Tang, X., Cao, J., Rong, R., Yu, Z., Liu, Y., Lu, Y., Liu, X., Han, L., Liu, J., et al. (2019). The Effect of HER2 Single Nucleotide Polymorphisms on Cervical Cancer Susceptibility and Survival in a Chinese Population. J. Cancer 10, 378–387. https://doi.org/10.7150/jca.27976. 65. Paska, A.V., and Hudler, P. (2015). Aberrant methylation patterns in cancer: a clinical view. Biochem. Med. 25, 161–176. https://doi.org/10. 11613/BM.2015.017. 66. Appelqvist, F., Yhr, M., Erlandson, A., Martinsson, T., and Enerba ¨ck, C. (2014). Deletion of the MGMT gene in familial melanoma. Genes Chromosomes Cancer 53, 703–711. https://doi.org/10.1002/gcc.22180. 67. De Summa, S.D., Guida, M., Tommasi, S., Strippoli, S., Pellegrini, C., Fargnoli, M.C., Pilato, B., Natalicchio, I., Guida, G., and Pinto, R. (2017). Genetic profiling of a rare condition: co-occurrence of albinism and multiple primary melanoma in a Caucasian family. Oncotarget 8, 29751–29759. https://doi.org/10.18632/oncotarget.12777. 68. Lubahn, J., Berndt, S.I., Jin, C.H., Klim, A., Luly, J., Wu, W.S., Isaacs, S., Wiley, K., Isaacs, W.B., Suarez, B.K., et al. (2010). Association of CASP8 D302H polymorphism with reduced risk of aggressive prostate carcinoma. Prostate 70, 646–653. https://doi.org/10.1002/pros.21098. 69. Vahednia, E., Shandiz, F.H., Bagherabad, M.B., Moezzi, A., Afzaljavan, F., Tajbakhsh, A., Kooshyar, M.M., and Pasdar, A. (2019). The Impact of CASP8 rs10931936 and rs1045485 Polymorphisms as well as the Haplotypes on Breast Cancer Risk: A Case-Control Study. Clin. Breast Cancer 19, e563–e577. https://doi.org/10.1016/j.clbc.2019.02.011. 70. Sanchez-Vega, F., Mina, M., Armenia, J., Chatila, W.K., Luna, A., La, K.C., Dimitriadoy, S., Liu, D.L., Kantheti, H.S., Saghafinia, S., et al. (2018). Oncogenic Signaling Pathways in The Cancer Genome Atlas. Cell 173, 321–337.e10. https://doi.org/10.1016/j.cell.2018.03.035. 71. Porta-Pardo, E., Garcia-Alonso, L., Hrabe, T., Dopazo, J., and Godzik, A. (2015). A Pan-Cancer Catalogue of Cancer Driver Protein Interaction Interfaces. PLoS Comput. Biol. 11, e1004518. https://doi.org/10.1371/ journal.pcbi.1004518. 72. Niu, B., Scott, A.D., SenGupta, S., Bailey, M.H., Batra, P., Ning, J., Wyczalkowski, M.A., Liang, W.-W., Zhang, Q., McLellan, M.D., et al. (2016). Protein-structure-guided discovery of functional mutations across 19 cancer types. Nat. Genet. 48, 827–837. https://doi.org/10.1038/ng.3586. 73. Huang, K.-L., Scott, A.D., Zhou, D.C., Wang, L.-B., Weerasinghe, A., Elmas, A., Liu, R., Wu, Y., Wendl, M.C., Wyczalkowski, M.A., et al. (2021). Spatially interacting phosphorylation sites and mutations in cancer. Nat. Commun. 12, 2313. https://doi.org/10.1038/s41467-021-22481-w. 74. Kamburov, A., Lawrence, M.S., Polak, P., Leshchiner, I., Lage, K., Golub, T.R., Lander, E.S., and Getz, G. (2015). Comprehensive assessment of cancer missense mutation clustering in protein structures. Proc. Natl. Acad. Sci. USA 112, E5486–E5495. https://doi.org/10.1073/pnas. 1516373112. 75. Varadi, M., Anyango, S., Deshpande, M., Nair, S., Natassia, C., Yordanova, G., Yuan, D., Stroe, O., Wood, G., Laydon, A., et al. (2022). AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Res. 50, D439–D444. https://doi.org/10.1093/nar/gkab1061. 76. Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R.,  Zı ´dek, A., Potapenko, A., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589. https://doi.org/10.1038/s41586-021-03819-2. 77. Berman, H.M., Westbrook, J., Feng, Z., Gilliland, G., Bhat, T.N., Weissig, H., Shindyalov, I.N., and Bourne, P.E. (2000). The Protein Data Bank. Nucleic Acids Res. 28, 235–242. https://doi.org/10.1093/nar/28.1.235. 78. Berman, H., Henrick, K., and Nakamura, H. (2003). Announcing the worldwide Protein Data Bank. Nat. Struct. Biol. 10, 980. https://doi.org/ 10.1038/nsb1203-980. 79. The UniProt Consortium (2023). UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res. 51, D523–D531. https://doi.org/ 10.1093/nar/gkac1052. 80. Sidney, J., Peters, B., Frahm, N., Brander, C., and Sette, A. (2008). HLA class I supertypes: a revised and updated classification. BMC Immunol. 9,1.https://doi.org/10.1186/1471-2172-9-1. 81. Al Naqbi, H., Mawart, A., Alshamsi, J., Al Safar, H., and Tay, G.K. (2021). Major histocompatibility complex (MHC) associations with diseases in ethnic groups of the Arabian Peninsula. Immunogenetics 73, 131–152. https://doi.org/10.1007/s00251-021-01204-x. 82. Medhasi, S., and Chantratita, N. (2022). Human Leukocyte Antigen (HLA) System: Genetics and Association with Bacterial and Viral Infections. J. Immunol. Res. 2022, 9710376. https://doi.org/10.1155/2022/9710376. ll OPEN ACCESS Cell 188, 2312–2335, May 1, 2025 2331 Article 83. Janeway, C.A., Travers, P., Walport, M., and Shlomchik, M. (2001). Immunobiology: the Immune System in Health and Disease, Fifth Edition (Garland Publishing). 84. Dou, Y., Kawaler, E.A., Cui Zhou, D., Gritsenko, M.A., Huang, C., Blumenberg, L., Karpova, A., Petyuk, V.A., Savage, S.R., Satpathy, S., et al. (2020). Proteogenomic Characterization of Endometrial Carcinoma. Cell 180, 729–748.e26. https://doi.org/10.1016/j.cell.2020.01.026. 85. Dong, P., Tada, M., Hamada, J.-I., Nakamura, A., Moriuchi, T., and Sakuragi, N. (2007). p53 dominant-negative mutant R273H promotes invasion and migration of human endometrial cancer HHUA cells. Clin. Exp. Metastasis 24, 471–483. https://doi.org/10.1007/s10585-007-9084-8. 86. Li, J., Yang, L., Gaur, S., Zhang, K., Wu, X., Yuan, Y.-C., Li, H., Hu, S., Weng, Y., and Yen, Y. (2014). Mutants TP53 p.R273H and p.R273C but not p.R273G Enhance Cancer Cell Malignancy. Hum. Mutat. 35, 575–584. https://doi.org/10.1002/humu.22528. 87. Xiao, G., Lundine, D., Annor, G.K., Canar, J., Ellison, V., Polotskaia, A., Donabedian, P.L., Reiner, T., Khramtsova, G.F., Olopade, O.I., et al. (2020). Gain-of-Function Mutant p53 R273H Interacts with Replicating DNA and PARP1 in Breast Cancer. Cancer Res. 80, 394–405. https:// doi.org/10.1158/0008-5472.CAN-19-1036. 88. Tan, B.S., Tiong, K.H., Choo, H.L., Chung, F.F.-L., Hii, L.W., Tan, S.H., Yap, I.K., Pani, S., Khor, N.T., Wong, S.F., et al. (2015). Mutant p53R273H mediates cancer cell survival and anoikis resistance through AKT-dependent suppression of BCL2-modifying factor (BMF). Cell Death Dis. 6, e1826. https://doi.org/10.1038/cddis.2015.191. 89. Lang, G.A., Iwakuma, T., Suh, Y.-A., Liu, G., Rao, V.A., Parant, J.M., Valentin-Vega, Y.A., Terzian, T., Caldwell, L.C., Strong, L.C., et al. (2004). Gain of function of a p53 hot spot mutation in a mouse model of LiFraumeni syndrome. Cell 119, 861–872. https://doi.org/10.1016/j.cell. 2004.11.006. 90. Goh, A.M., Coffill, C.R., and Lane, D.P. (2011). The role of mutant p53 in human cancer. J. Pathol. 223, 116–126. https://doi.org/10.1002/ path.2784. 91. Solomon, H., Madar, S., and Rotter, V. (2011). Mutant p53 gain of function is interwoven into the hallmarks of cancer. J. Pathol. 225, 475–478. https://doi.org/10.1002/path.2988. 92. Wang, H., Guo, M., Wei, H., and Chen, Y. (2023). Targeting p53 pathways: mechanisms, structures, and advances in therapy. Signal Transduct. Target. Ther. 8, 1–35. https://doi.org/10.1038/s41392-02301347-1. 93. Belinsky, M., and Jaiswal, A.K. (1993). NAD(P)H:Quinone oxidoreductase1 (DT-diaphorase) expression in normal and tumor tissues. Cancer Metastasis Rev. 12, 103–117. https://doi.org/10.1007/BF00689804. 94. Fagerholm, R., Hofstetter, B., Tommiska, J., Aaltonen, K., Vrtel, R., Syrja ¨koski, K., Kallioniemi, A., Kilpivaara, O., Mannermaa, A., Kosma, V.-M., et al. (2008). NAD(P)H:quinone oxidoreductase 1 NQO1*2 genotype (P187S) is a strong prognostic and predictive factor in breast cancer. Nat. Genet. 40, 844–853. https://doi.org/10.1038/ng.155. 95. Tossetta, G., Fantone, S., Goteri, G., Giannubilo, S.R., Ciavattini, A., and Marzioni, D. (2023). The Role of NQO1 in Ovarian Cancer. Int. J. Mol. Sci. 24, 7839. https://doi.org/10.3390/ijms24097839. 96. Prawan, A., Buranrat, B., Kukongviriyapan, U., Sripa, B., and Kukongviriyapan, V. (2009). Inflammatory cytokines suppress NAD(P)H:quinone oxidoreductase-1 and induce oxidative stress in cholangiocarcinoma cells. J. Cancer Res. Clin. Oncol. 135, 515–522. https://doi.org/10. 1007/s00432-008-0483-2. 97. Kolesar, J.M., Pritchard, S.C., Kerr, K.M., Kim, K., Nicolson, M.C., and McLeod, H. (2002). Evaluation of NQO1 gene expression and variant allele in human NSCLC tumors and matched normal lung tissue. Int. J. Oncol. 21, 1119–1124. https://doi.org/10.3892/ijo.21.5.1119. 98. Cresteil, T., and Jaiswal, A.K. (1991). High levels of expression of the NAD(P)H:quinone oxidoreductase (NQO1) gene in tumor cells compared to normal cells of the same origin. Biochem. Pharmacol. 42, 1021–1027. https://doi.org/10.1016/0006-2952(91)90284-c. 99. Tossetta, G., and Marzioni, D. (2023). Targeting the NRF2/KEAP1 pathway in cervical and endometrial cancers. Eur. J. Pharmacol. 941, 175503. https://doi.org/10.1016/j.ejphar.2023.175503. 100. Siegel, D., Franklin, W.A., and Ross, D. (1998). Immunohistochemical detection of NAD(P)H:quinone oxidoreductase in human lung and lung tumors. Clin. Cancer Res. 4, 2065–2070. 101. Siegel, D., and Ross, D. (2000). Immunodetection of NAD(P)H:quinone oxidoreductase 1 (NQO1) in human tissues. Free Radic. Biol. Med. 29, 246–253. https://doi.org/10.1016/s0891-5849(00)00310-5. 102. Awadallah, N.S., Dehn, D., Shah, R.J., Russell Nash, S., Chen, Y.K., Ross, D., Bentz, J.S., and Shroyer, K.R. (2008). NQO1 expression in pancreatic cancer and its potential use as a biomarker. Appl. Immunohistochem. Mol. Morphol. 16, 24–31. https://doi.org/10.1097/PAI. 0b013e31802e91d0. 103. Tossetta, G., Fantone, S., Montanari, E., Marzioni, D., and Goteri, G. (2022). Role of NRF2 in Ovarian Cancer. Antioxidants (Basel) 11, 663. https://doi.org/10.3390/antiox11040663. 104. Beaver, S.K., Mesa-Torres, N., Pey, A.L., and Timson, D.J. (2019). NQO1: A target for the treatment of cancer and neurological diseases, and a model to understand loss of function disease mechanisms. Biochim. Biophys. Acta Proteins Proteom. 1867, 663–676. https://doi.org/10.1016/j. bbapap.2019.05.002. 105. Lajin, B., and Alachkar, A. (2013). The NQO1 polymorphism C609T (Pro187Ser) and cancer susceptibility: a comprehensive meta-analysis. Br. J. Cancer 109, 1325–1337. https://doi.org/10.1038/bjc.2013.357. 106. Siegel, D., Anwar, A., Winski, S.L., Kepa, J.K., Zolman, K.L., and Ross, D. (2001). Rapid polyubiquitination and proteasomal degradation of a mutant form of NAD(P)H:quinone oxidoreductase 1. Mol. Pharmacol. 59, 263–268. https://doi.org/10.1124/mol.59.2.263. 107. Medina-Carmona, E., Neira, J.L., Salido, E., Fuchs, J.E., Palomino-Morales, R., Timson, D.J., and Pey, A.L. (2017). Site-to-site interdomain communication may mediate different loss-of-function mechanisms in a cancer-associated NQO1 polymorphism. Sci. Rep. 7, 44532. https:// doi.org/10.1038/srep44532. 108. Medina-Carmona, E., Betancor-Ferna ´ndez, I., Santos, J., Mesa-Torres, N., Grottelli, S., Batlle, C., Naganathan, A.N., Oppici, E., Cellini, B., Ventura, S., et al. (2019). Insight into the specificity and severity of pathogenic mechanisms associated with missense mutations through experimental and structural perturbation analyses. Hum. Mol. Genet. 28, 1–15. https:// doi.org/10.1093/hmg/ddy323. 109. Martı ´nez-Limo ´n, A., Alriquet, M., Lang, W.-H., Calloni, G., Wittig, I., and Vabulas, R.M. (2016). Recognition of enzymes lacking bound cofactor by protein quality control. Proc. Natl. Acad. Sci. USA 113, 12156– 12161. https://doi.org/10.1073/pnas.1611994113. 110. Pacheco-Garcia, J.L., Anoz-Carbonell, E., Vankova, P., Kannan, A., Palomino-Morales, R., Mesa-Torres, N., Salido, E., Man, P., Medina, M., Naganathan, A.N., et al. (2021). Structural basis of the pleiotropic and specific phenotypic consequences of missense mutations in the multifunctional NAD(P)H:quinone oxidoreductase 1 and their pharmacological rescue. Redox Biol. 46, 102112. https://doi.org/10.1016/j.redox.2021. 102112. 111. Mun ˜oz, I.G., Morel, B., Medina-Carmona, E., and Pey, A.L. (2017). A mechanism for cancer-associated inactivation of NQO1 due to P187S and its reactivation by the consensus mutation H80R. FEBS Lett. 591, 2826–2835. https://doi.org/10.1002/1873-3468.12772. 112. Medina-Carmona, E., Rizzuti, B., Martı ´n-Escolano, R., Pacheco-Garcı ´a, J.L., Mesa-Torres, N., Neira, J.L., Guzzi, R., and Pey, A.L. (2019). Phosphorylation compromises FAD binding and intracellular stability of wild-type and cancer-associated NQO1: Insights into flavo-proteome stability. Int. J. Biol. Macromol. 125, 1275–1288. https://doi.org/10. 1016/j.ijbiomac.2018.09.108. ll OPEN ACCESS 2332 Cell 188, 2312–2335, May 1, 2025 Article 113. Medina-Carmona, E., Fuchs, J.E., Gavira, J.A., Mesa-Torres, N., Neira, J.L., Salido, E., Palomino-Morales, R., Burgos, M., Timson, D.J., and Pey, A.L. (2017). Enhanced vulnerability of human proteins towards disease-associated inactivation through divergent evolution. Hum. Mol. Genet. 26, 3531–3544. https://doi.org/10.1093/hmg/ddx238. 114. Pacheco-Garcia, J.L., Loginov, D.S., Anoz-Carbonell, E., Vankova, P., Palomino-Morales, R., Salido, E., Man, P., Medina, M., Naganathan, A.N., and Pey, A.L. (2022). Allosteric Communication in the Multifunctional and Redox NQO1 Protein Studied by Cavity-Making Mutations. Antioxidants (Basel) 11, 1110. https://doi.org/10.3390/antiox11061110. 115. Dwight, T., Benn, D.E., Clarkson, A., Vilain, R., Lipton, L., Robinson, B.G., Clifton-Bligh, R.J., and Gill, A.J. (2013). Loss of SDHA Expression Identifies SDHA Mutations in Succinate Dehydrogenase–deficient Gastrointestinal Stromal Tumors. Am. J. Surg. Pathol. 37, 226–233. https://doi. org/10.1097/PAS.0b013e3182671155. 116. Burnichon, N., Brie `re, J.-J., Libe ´, R., Vescovo, L., Rivie `re, J., Tissier, F., Jouanno, E., Jeunemaitre, X., Be ´nit, P., Tzagoloff, A., et al. (2010). SDHA is a tumor suppressor gene causing paraganglioma. Hum. Mol. Genet. 19, 3011–3020. https://doi.org/10.1093/hmg/ddq206. 117. Kamai, T., Higashi, S., Murakami, S., Arai, K., Namatame, T., Kijima, T., Abe, H., Jamiyan, T., Ishida, K., Shirataki, H., et al. (2021). Single nucleotide variants of succinate dehydrogenase A gene in renal cell carcinoma. Cancer Sci. 112, 3375–3387. https://doi.org/10.1111/cas.14977. 118. Knight, J.C. (2004). Allele-specific gene expression uncovered. Trends Genet. 20, 113–116. https://doi.org/10.1016/j.tig.2004.01.001. 119. Castel, S.E., Levy-Moonshine, A., Mohammadi, P., Banks, E., and Lappalainen, T. (2015). Tools and best practices for data processing in allelic expression analysis. Genome Biol. 16, 195. https://doi.org/10.1186/ s13059-015-0762-6. 120. van Beek, D., Verdonschot, J., Derks, K., Brunner, H., de Kok, T.M., Arts, I.C.W., Heymans, S., Kutmon, M., and Adriaens, M. (2023). Allele-specific expression analysis for complex genetic phenotypes applied to a unique dilated cardiomyopathy cohort. Sci. Rep. 13, 564. https://doi.org/10. 1038/s41598-023-27591-7. 121. Robles-Espinoza, C.D., Mohammadi, P., Bonilla, X., and Gutierrez-Arcelus, M. (2021). Allele-specific expression: applications in cancer and technical considerations. Curr. Opin. Genet. Dev. 66, 10–19. https:// doi.org/10.1016/j.gde.2020.10.007. 122. Wang, Z., Fan, X., Shen, Y., Pagadala, M.S., Signer, R., Cygan, K.J., Fairbrother, W.G., Carter, H., Chung, W.K., and Huang, K.-L. (2021). Noncancer-related pathogenic germline variants and expression consequences in ten-thousand cancer genomes. Genome Med. 13, 147. https://doi.org/10.1186/s13073-021-00964-1. 123. Wingo, T.S., Duong, D.M., Zhou, M., Dammer, E.B., Wu, H., Cutler, D.J., Lah, J.J., Levey, A.I., and Seyfried, N.T. (2017). Integrating NextGeneration Genomic Sequencing and Mass Spectrometry to Estimate Allele-Specific Protein Abundance in Human Brain. J. Proteome Res. 16, 3336–3347. https://doi.org/10.1021/acs.jproteome.7b00324. 124. Johansson, H.J., Socciarelli, F., Vacanti, N.M., Haugen, M.H., Zhu, Y., Siavelis, I., Fernandez-Woodbridge, A., Aure, M.R., Sennblad, B., Vesterlund, M., et al. (2019). Breast cancer quantitative proteome and proteogenomic landscape. Nat. Commun. 10, 1600. https://doi.org/10.1038/ s41467-019-09018-y. 125. Wu, L., and Snyder, M. (2015). Impact of allele-specific peptides in proteome quantification. Proteomics Clin. Appl. 9, 432–436. https://doi.org/ 10.1002/prca.201400126. 126. Shi, J., Wang, X., Zhu, H., Jiang, H., Wang, D., Nesvizhskii, A., and Zhu, H.-J. (2018). Determining Allele-Specific Protein Expression (ASPE) Using a Novel Quantitative Concatamer-Based Proteomics Method. J. Proteome Res. 17, 3606–3612. https://doi.org/10.1021/acs.jproteome.8b00620. 127. Liu, Z., Dong, X., and Li, Y. (2018). A Genome-Wide Study of AlleleSpecific Expression in Colorectal Cancer. Front. Genet. 9, 570. https:// doi.org/10.3389/fgene.2018.00570. 128. Bell, C.G., and Beck, S. (2009). Advances in the identification and analysis of allele-specific expression. Genome Med. 1, 56. https://doi.org/ 10.1186/gm56. 129. Polak, P., Kim, J., Braunstein, L.Z., Karlic, R., Haradhavala, N.J., Tiao, G., Rosebrock, D., Livitz, D., Ku ¨bler, K., Mouw, K.W., et al. (2017). A mutational signature reveals alterations underlying deficient homologous recombination repair in breast cancer. Nat. Genet. 49, 1476–1486. https://doi.org/10.1038/ng.3934. 130. Larsen, D.H., Poinsignon, C., Gudjonsson, T., Dinant, C., Payne, M.R., Hari, F.J., Rendtlew Danielsen, J.M., Menard, P., Sand, J.C., Stucki, M., et al. (2010). The chromatin-remodeling factor CHD4 coordinates signaling and repair after DNA damage. J. Cell Biol. 190, 731–740. https://doi.org/10.1083/jcb.200912135. 131. Silva, A.P.G., Ryan, D.P., Galanty, Y., Low, J.K.K., Vandevenne, M., Jackson, S.P., and Mackay, J.P. (2016). The N-terminal Region of Chromodomain Helicase DNA-binding Protein 4 (CHD4) Is Essential for Activity and Contains a High Mobility Group (HMG) Box-like-domain That Can Bind Poly(ADP-ribose). J. Biol. Chem. 291, 924–938. https://doi.org/10. 1074/jbc.M115.683227. 132. De Souza, C., Madden, J., Koestler, D.C., Minn, D., Montoya, D.J., Minn, K., Raetz, A.G., Zhu, Z., Xiao, W.-W., Tahmassebi, N., et al. (2021). Effect of the p53 P72R Polymorphism on Mutant TP53 Allele Selection in Human Cancer. J. Natl. Cancer Inst. 113, 1246–1257. https://doi.org/10. 1093/jnci/djab019. 133. Bojesen, S.E., and Nordestgaard, B.G. (2008). The common germline Arg72Pro polymorphism of p53 and increased longevity in humans. Cell Cycle 7, 158–163. https://doi.org/10.4161/cc.7.2.5249. 134. Marin, M.C., Jost, C.A., Brooks, L.A., Irwin, M.S., O’Nions, J., Tidy, J.A., James, N., McGregor, J.M., Harwood, C.A., Yulug, I.G., et al. (2000). A common polymorphism acts as an intragenic modifier of mutant p53 behaviour. Nat. Genet. 25, 47–54. https://doi.org/10.1038/75586. 135. Siddique, M., and Sabapathy, K. (2006). Trp53-dependent DNA-repair is affected by the codon 72 polymorphism. Oncogene 25, 3489–3500. https://doi.org/10.1038/sj.onc.1209405. 136. Tuch, B.B., Laborde, R.R., Xu, X., Gu, J., Chung, C.B., Monighetti, C.K., Stanley, S.J., Olsen, K.D., Kasperbauer, J.L., Moore, E.J., et al. (2010). Tumor Transcriptome Sequencing Reveals Allelic Expression Imbalances Associated with Copy Number Alterations. PLoS One 5, e9317. https://doi.org/10.1371/journal.pone.0009317. 137. Fagan, R.J., and Dingwall, A.K. (2019). COMPASS Ascending: Emerging clues regarding the roles of MLL3/KMT2C and MLL2/KMT2D proteins in cancer. Cancer Lett. 458, 56–65. https://doi.org/10.1016/j.canlet.2019. 05.024. 138. Fang, F., Liu, C., Li, Q., Xu, R., Zhang, T., and Shen, X. (2022). The Role of SETBP1 in Gastric Cancer: Friend or Foe. Front. Oncol. 12, 908943. https://doi.org/10.3389/fonc.2022.908943. 139. Wang, H., Gao, Y., Qin, L., Zhang, M., Shi, W., Feng, Z., Guo, L., Zhu, B., and Liao, S. (2023). Identification of a novel de novo mutation of SETBP1 and new findings of SETBP1 in tumorgenesis. Orphanet J. Rare Dis. 18, 107. https://doi.org/10.1186/s13023-023-02705-6. 140. Zhang, J., Sun, X., Qian, Y., LaDuca, J.P., and Maquat, L.E. (1998). At least one intron is required for the nonsense-mediated decay of triosephosphate isomerase mRNA: a possible link between nuclear splicing and cytoplasmic translation. Mol. Cell. Biol. 18, 5272–5283. https://doi. org/10.1128/MCB.18.9.5272. 141. Thermann, R., Neu-Yilik, G., Deters, A., Frede, U., Wehr, K., Hagemeier, C., Hentze, M.W., and Kulozik, A.E. (1998). Binary specification of nonsense codons by splicing and cytoplasmic translation. EMBO J. 17, 3484–3494. https://doi.org/10.1093/emboj/17.12.3484. 142. Lindeboom, R.G.H., Supek, F., and Lehner, B. (2016). The rules and impact of nonsense-mediated mRNA decay in human cancers. Nat. Genet. 48, 1112–1118. https://doi.org/10.1038/ng.3664. ll OPEN ACCESS Cell 188, 2312–2335, May 1, 2025 2333 Article 143. Kapoor, G.S., and O’Rourke, D.M. (2023). Editorial Expression of Concern: SIRPa1 receptors interfere with the EGFRvIII signalosome to inhibit glioblastoma cell transformation and migration. Oncogene 42, 2154. https://doi.org/10.1038/s41388-023-02740-4. 144. Gholamin, S., Mitra, S.S., Feroze, A.H., Liu, J., Kahn, S.A., Zhang, M., Esparza, R., Richard, C., Ramaswamy, V., Remke, M., et al. (2017). Disrupting the CD47-SIRPaanti-phagocytic axis by a humanized anti-CD47 antibody is an efficacious treatment for malignant pediatric brain tumors. Sci. Transl. Med. 9, eaaf2968. https://doi.org/10.1126/scitranslmed. aaf2968. 145. Jung, C.S., Foerch, C., Scha ¨nzer, A., Heck, A., Plate, K.H., Seifert, V., Steinmetz, H., Raabe, A., and Sitzer, M. (2007). Serum GFAP is a diagnostic marker for glioblastoma multiforme. Brain 130, 3336–3341. https://doi.org/10.1093/brain/awm263. 146. Radu, R., Petrescu, G.E.D., Gorgan, R.M., and Brehar, F.M. (2022). GFAPd: A Promising Biomarker and Therapeutic Target in Glioblastoma. Front. Oncol. 12, 859247. https://doi.org/10.3389/fonc.2022.859247. 147. Sticht, C., De La Torre, C., Parveen, A., and Gretz, N. (2018). miRWalk: An online resource for prediction of microRNA binding sites. PLoS One 13, e0206239. https://doi.org/10.1371/journal.pone.0206239. 148. Conway, J.R., Lex, A., and Gehlenborg, N. (2017). UpSetR: an R package for the visualization of intersecting sets and their properties. Bioinformatics 33, 2938–2940. https://doi.org/10.1093/bioinformatics/btx364. 149. Mai, S., and Inkielewicz-Stepniak, I. (2021). Pancreatic Cancer and Platelets Crosstalk: A Potential Biomarker and Target. Front. Cell Dev. Biol. 9, 749689. https://doi.org/10.3389/fcell.2021.749689. 150. Na’ara, S., Amit, M., and Gil, Z. (2019). L1CAM induces perineural invasion of pancreas cancer cells by upregulation of metalloproteinase expression. Oncogene 38, 596–608. https://doi.org/10.1038/s41388018-0458-y. 151. Drosten, M., and Barbacid, M. (2020). Targeting the MAPK Pathway in KRAS-Driven Tumors. Cancer Cell 37, 543–550. https://doi.org/10. 1016/j.ccell.2020.03.013. 152. Akdel, M., Pires, D.E.V., Pardo, E.P., Ja ¨nes, J., Zalevsky, A.O., Me ´sza ´ros, B., Bryant, P., Good, L.L., Laskowski, R.A., Pozzati, G., et al. (2022). A structural biology community assessment of AlphaFold2 applications. Nat. Struct. Mol. Biol. 29, 1056–1067. https://doi.org/10.1038/s41594022-00849-w. 153. GTEx Consortium (2015). Human genomics. The Genotype-Tissue Expression (GTEx) pilot analysis: multitissue gene regulation in humans. Science 348, 648–660. https://doi.org/10.1126/science.1262110. 154. Abdellaoui, A., Yengo, L., Verweij, K.J.H., and Visscher, P.M. (2023). 15 years of GWAS discovery: Realizing the promise. Am. J. Hum. Genet. 110, 179–194. https://doi.org/10.1016/j.ajhg.2022.12.011. 155. Shi, H., Burch, K.S., Johnson, R., Freund, M.K., Kichaev, G., Mancuso, N., Manuel, A.M., Dong, N., and Pasaniuc, B. (2020). Localizing Components of Shared Transethnic Genetic Architecture of Complex Traits from GWAS Summary Data. Am. J. Hum. Genet. 106, 805–817. https://doi. org/10.1016/j.ajhg.2020.04.012. 156. Martı ´nez-Jime ´nez, F., Muin ˜os, F., Sentı ´s, I., Deu-Pons, J., Reyes-Salazar, I., Arnedo-Pac, C., Mularoni, L., Pich, O., Bonet, J., Kranas, H., et al. (2020). A compendium of mutational cancer driver genes. Nat. Rev. Cancer 20, 555–572. https://doi.org/10.1038/s41568-020-0290-x. 157. ICGC/TCGA Pan-Cancer Analysis of Whole Genomes Consortium (2020). Pan-cancer analysis of whole genomes. Nature 578, 82–93. https://doi.org/10.1038/s41586-020-1969-6. 158. Ding, L., Bailey, M.H., Porta-Pardo, E., Thorsson, V., Colaprico, A., Bertrand, D., Gibbs, D.L., Weerasinghe, A., Huang, K.L., Tokheim, C., et al. (2018). Perspective on Oncogenic Processes at the End of the Beginning of Cancer Genomics. Cell 173, 305–320.e10. https://doi.org/10.1016/j. cell.2018.03.033. 159. Carrot-Zhang, J., Soca-Chafre, G., Patterson, N., Thorner, A.R., Nag, A., Watson, J., Genovese, G., Rodriguez, J., Gelbard, M.K., Corrales-Rodriguez, L., et al. (2021). Genetic Ancestry Contributes to Somatic Mutations in Lung Cancers from Admixed Latin American Populations. Cancer Discov. 11, 591–598. https://doi.org/10.1158/2159-8290.CD-20-1165. 160. Li, Y., Dou, Y., Da Veiga Leprevost, F., Geffen, Y., Calinawan, A.P., Aguet, F., Akiyama, Y., Anand, S., Birger, C., Cao, S., et al. (2023). Proteogenomic data and resources for pan-cancer analysis. Cancer Cell 41, 1397–1406. https://doi.org/10.1016/j.ccell.2023.06.009. 161. Li, H. (2013). Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. Preprint at arXiv. 162. Wu, T., Hu, E., Xu, S., Chen, M., Guo, P., Dai, Z., Feng, T., Zhou, L., Tang, W., Zhan, L., et al. (2021). clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. Innovation (Camb). 2, 100141. https://doi. org/10.1016/j.xinn.2021.100141. 163. Freed, D., Aldana, R., Weber, J.A., and Edwards, J.S. (2017). The Sentieon Genomics Tools - A fast and accurate solution to variant calling from next-generation sequence data. Preprint at bioRxiv. https://doi. org/10.1101/115717. 164. McLaren, W., Gil, L., Hunt, S.E., Riat, H.S., Ritchie, G.R.S., Thormann, A., Flicek, P., and Cunningham, F. (2016). The Ensembl Variant Effect Predictor. Genome Biol. 17, 122. https://doi.org/10.1186/s13059-0160974-4. 165. Andrews, S. (2010). FastQC: A Quality Control tool for High Throughput Sequence Data. https://www.bioinformatics.babraham.ac.uk/projects/ fastqc/ 166. McKenna, A., Hanna, M., Banks, E., Sivachenko, A., Cibulskis, K., Kernytsky, A., Garimella, K., Altshuler, D., Gabriel, S., Daly, M., et al. (2010). The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res. 20, 1297–1303. https://doi.org/10.1101/gr.107524.110. 167. Robinson, J.T., Thorvaldsdo ´ttir, H., Winckler, W., Guttman, M., Lander, E.S., Getz, G., and Mesirov, J.P. (2011). Integrative Genomics Viewer. Nat. Biotechnol. 29, 24–26. https://doi.org/10.1038/nbt.1754. 168. Kalayci, S., Selvan, M.E., Ramos, I., Cotsapas, C., Harris, E., Kim, E.-Y., Montgomery, R.R., Poland, G., Pulendran, B., Tsang, J.S., et al. (2019). ImmuneRegulation: a web-based tool for identifying human immune regulatory elements. Nucleic Acids Res. 47, W142–W150. https://doi.org/ 10.1093/nar/gkz450. 169. Pedersen, B.S., and Quinlan, A.R. (2018). Mosdepth: quick coverage calculation for genomes and exomes. Bioinformatics 34, 867–868. https://doi.org/10.1093/bioinformatics/btx699. 170. Cibulskis, K., Lawrence, M.S., Carter, S.L., Sivachenko, A., Jaffe, D., Sougnez, C., Gabriel, S., Meyerson, M., Lander, E.S., and Getz, G. (2013). Sensitive detection of somatic point mutations in impure and heterogeneous cancer samples. Nat. Biotechnol. 31, 213–219. 171. Ye, K., Schulz, M.H., Long, Q., Apweiler, R., and Ning, Z. (2009). Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short reads. Bioinformatics 25, 2865–2871. https://doi.org/10.1093/bioinformatics/btp394. 172. Kim, S., Scheffler, K., Halpern, A.L., Bekritsky, M.A., Noh, E., Ka ¨llberg, M., Chen, X., Kim, Y., Beyter, D., Krusche, P., and Saunders, C.T. (2018). Strelka2: fast and accurate calling of germline and somatic variants. Nat. Methods 15, 591–594. 173. Szklarczyk, D., Gable, A.L., Lyon, D., Junge, A., Wyder, S., Huerta-Cepas, J., Simonovic, M., Doncheva, N.T., Morris, J.H., Bork, P., et al. (2019). STRING v11: protein-protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets. Nucleic Acids Res. 47, D607–D613. https://doi.org/10.1093/ nar/gky1131. 174. Kassambara, A., Kosinski, M., and Biecek, P. (2021). survminer: Drawing Survival Curves using ‘ggplot2’. (R package version 0.4.9). https://rpkgs. datanovia.com/survminer/index.html. 175. Zhou, W., Chen, T., Chong, Z., Rohrdanz, M.A., Melott, J.M., Wakefield, C., Zeng, J., Weinstein, J.N., Meric-Bernstam, F., Mills, G.B., et al. (2015). ll OPEN ACCESS 2334 Cell 188, 2312–2335, May 1, 2025 Article TransVar: a multilevel variant annotator for precision genomics. Nat. Methods 12, 1002–1003. https://doi.org/10.1038/nmeth.3622. 176. Koboldt, D.C., Zhang, Q., Larson, D.E., Shen, D., McLellan, M.D., Lin, L., Miller, C.A., Mardis, E.R., Ding, L., and Wilson, R.K. (2012). VarScan 2: somatic mutation and copy number alteration discovery in cancer by exome sequencing. Genome Res. 22, 568–576. https://doi.org/10. 1101/gr.129684.111. 177. Kandoth, C., Gao, J., qwangmsk, Mattioni, M., Struck, A., Boursin, Y., Penson, A., and Chavan, S. (2018). mskcc/vcf2maf: vcf2maf v1.6.16 (Zenodo) https://doi.org/10.5281/zenodo.1185418. 178. Auton, A., Brooks, L.D., Durbin, R.M., Garrison, E.P., Kang, H.M., Korbel, J.O., Marchini, J.L., McCarthy, S., McVean, G.A., et al.; 1000 Genomes Project Consortium (2015). A global reference for human genetic variation. Nature 526, 68–74. https://doi.org/10.1038/nature15393. 179. Paten, B., Herrero, J., Fitzgerald, S., Beal, K., Flicek, P., Holmes, I., and Birney, E. (2008). Genome-wide nucleotide-level mammalian ancestor reconstruction. Genome Res. 18, 1829–1843. https://doi.org/10.1101/ gr.076521.108. 180. Paten, B., Herrero, J., Beal, K., Fitzgerald, S., and Birney, E. (2008). Enredo and Pecan: genome-wide mammalian consistency-based multiple alignment with paralogs. Genome Res. 18, 1814–1828. https://doi.org/ 10.1101/gr.076554.108. 181. Richards, S., Aziz, N., Bale, S., Bick, D., Das, S., Gastier-Foster, J., Grody, W.W., Hegde, M., Lyon, E., Spector, E., et al. (2015). Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genet. Med. 17, 405–424. https://doi.org/10.1038/gim.2015.30. 182. Karczewski, K.J., Francioli, L.C., Tiao, G., Cummings, B.B., Alfo ¨ldi, J., Wang, Q., Collins, R.L., Laricchia, K.M., Ganna, A., Birnbaum, D.P., et al. (2020). The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581, 434–443. https://doi.org/10.1038/ s41586-020-2308-7. 183. Kumar, P., Henikoff, S., and Ng, P.C. (2009). Predicting the effects of coding non-synonymous variants on protein function using the SIFT algorithm. Nat. Protoc. 4, 1073–1081. https://doi.org/10.1038/nprot. 2009.86. 184. Adzhubei, I., Jordan, D.M., and Sunyaev, S.R. (2013). Predicting functional effect of human missense mutations using PolyPhen-2. Curr. Protoc. Hum. Genet. Chapter 7, Unit7.20. https://doi.org/10.1002/ 0471142905.hg0720s76. 185. Basu, S., and Pan, W. (2011). Comparison of statistical tests for disease association with rare variants. Genet. Epidemiol. 35, 606–619. https:// doi.org/10.1002/gepi.20609. 186. Ouspenskaia, T., Law, T., Clauser, K.R., Klaeger, S., Sarkizova, S., Aguet, F., Li, B., Christian, E., Knisbacher, B.A., Le, P.M., et al. (2022). Unannotated proteins expand the MHC-I-restricted immunopeptidome in cancer. Nat. Biotechnol. 40, 209–217. https://doi.org/10.1038/s41587-02101021-3. 187. Niu, L., Stinson, S.E., Holm, L.A., Lund, M.A.V., Fonvig, C.E., Cobuccio, L., Meisner, J., Juel, H.B., Thiele, M., Krag, A., et al. (2025). Plasma Proteome Variation and its Genetic Determinants in Children and Adolescents. Nat. Genet. 57, 635–646. https://doi.org/10.1038/s41588-02502089-2. 188. Liberzon, A., Birger, C., Thorvaldsdo ´ttir, H., Ghandi, M., Mesirov, J.P., and Tamayo, P. (2015). The Molecular Signatures Database (MSigDB) hallmark gene set collection. Cell Syst. 1, 417–425. https://doi.org/10. 1016/j.cels.2015.12.004. 189. Subramanian, A., Tamayo, P., Mootha, V.K., Mukherjee, S., Ebert, B.L., Gillette, M.A., Paulovich, A., Pomeroy, S.L., Golub, T.R., Lander, E.S., et al. (2005). Gene set enrichment analysis: A knowledge-based approach for interpreting genome-wide expression profiles. Proc. Natl. Acad. Sci. USA 102, 15545–15550. https://doi.org/10.1073/pnas. 0506580102. 190. Slenter, D.N., Kutmon, M., Hanspers, K., Riutta, A., Windsor, J., Nunes, N., Me ´lius, J., Cirillo, E., Coort, S.L., Digles, D., et al. (2018). WikiPathways: a multifaceted pathway database bridging metabolomics to other omics research. Nucleic Acids Res. 46, D661–D667. https://academic. oup.com/nar/article/46/D1/D661/4612963. 191. Cunningham, F., Allen, J.E., Allen, J., Alvarez-Jarreta, J., Amode, M.R., Armean, I.M., Austine-Orimoloye, O., Azov, A.G., Barnes, I., Bennett, R., et al. (2022). Ensembl 2022. Nucleic Acids Res. 50, D988–D995. https://doi.org/10.1093/nar/gkab1049. 192. Shabalin, A.A. (2012). Matrix eQTL: ultra fast eQTL analysis via large matrix operations. Bioinformatics 28, 1353–1358. https://doi.org/10.1093/ bioinformatics/bts163. 193. Stegle, O., Parts, L., Piipari, M., Winn, J., and Durbin, R. (2012). Using probabilistic estimation of expression residuals (PEER) to obtain increased power and interpretability of gene expression analyses. Nat. Protoc. 7, 500–507. https://doi.org/10.1038/nprot.2011.457. 194. Gy} orffy, B. (2023). Discovery and ranking of the most robust prognostic biomarkers in serous ovarian cancer. GeroScience 45, 1889–1898. https://doi.org/10.1007/s11357-023-00742-4. 195. Giambartolomei, C., Vukcevic, D., Schadt, E.E., Franke, L., Hingorani, A.D., Wallace, C., and Plagnol, V. (2014). Bayesian test for colocalisation between pairs of genetic association studies using summary statistics. PLoS Genet. 10, e1004383. https://doi.org/10.1371/journal.pgen. 1004383. 196. Strimmer, K. (2008). fdrtool: a versatile R package for estimating local and tail area-based false discovery rates. Bioinformatics 24, 1461– 1462. https://doi.org/10.1093/bioinformatics/btn209. 197. Yu, G., and He, Q.-Y. (2016). ReactomePA: an R/Bioconductor package for reactome pathway analysis and visualization. Mol. Biosyst. 12, 477–479. https://doi.org/10.1039/C5MB00663E. ll OPEN ACCESS Cell 188, 2312–2335, May 1, 2025 2335 Article Peptide spectrum match (PSM) filtering and false discovery rates (FDR) Using the SM Autovalidation module peptide spectrum matches (PSMs) for individual spectra were confidently assigned by applying target-decoy based FDR estimation to achieve <1.0% FDR at the PSM, peptide, VM site and protein levels. For the whole proteome dataset thresholding was done in 3 steps: at the PSM level, the protein level for each TMT-plex, and the protein level for the cohort of 2 TMT-plexes. For the PTM omes (phosphoproteome and acetylome datasets), thresholding was done in two steps: at the PSM level for each TMT-plex and at the VM site level for the cohort of 2 TMT-plexes. In step 1 for all datasets, PSM level autovalidation was done first and separately for each TMT-plex experiment using an auto-thresholds strategy with a minimum sequence length of 7; automatic variable range precursor mass filtering; with score and delta Rank1 - Rank2 score thresholds optimized to yield a PSM level FDR estimate for precursor charges 2 through 4 of < 0.8% for each precursor charge state in each LC-MS/MS run. To achieve reasonable statistics for precursor charges 5-6, thresholds were optimized to yield a PSM-level FDR estimate of < 0.4% across all runs per TMTplex experiment (instead of per each run), since many fewer spectra are generated for the higher charge states. In step 2 for the PTM omes: phosphoproteome and acetylome datasets VM site polishing autovalidation was applied across both TMT plexes to retain all VM site identifications with either a minimum id score of 8.0 or observation in n TMT plexes (n=4, 3, or 2 if > 20, 7, or 1 plexes/cohort, respectively). The intention of the VM site polishing step is to control FDR by eliminating unreliable VM site level identifications, particularly low scoring VM-sites that are only detected as low scoring peptides that are also infrequently detected across TMT plexes in the study. Using the SM Protein/Peptide Summary module to make VM-site reports the ubiquitylome and acetylome datasets are further filtered to remove peptides ending with the regular expression [^K][^K]k since trypsin and Lys-C cannot cleave at a acetylated lysine. The [^K] means retain if unmodified Lys present in one of the last two positions to allow for a missed cleavage with ambiguous PTM-site localization. C-terminally acetylated lysines are present in the acetylome dataset, but have been shown to arise from artifactual modification during TMT-labeling after trypsin digestion. In step 2 for the whole proteome dataset, protein polishing autovalidation was applied separately to each TMT-plex experiment to further filter the PSMs using a target protein level FDR threshold of zero. The primary goal of this step was to eliminate peptides identified with low scoring PSMs that represent proteins identified by a single peptide, so-called ‘‘one-hit wonders.’’ After assembling protein groups from the autovalidated PSMs, protein polishing determined the maximum protein level score of a protein group that consisted entirely of distinct peptides estimated to be false-positive identifications (PSMs with negative delta forward-reverse scores). PSMs were removed from the set obtained in the initial peptide level autovalidation step if they contributed to protein groups that had protein scores below the maximum false-positive protein score. Step 3 was then applied, consisting of protein polishing autovalidation across all TMT plexes in a cohort together using the protein grouping method ‘‘expand subgroups, top uses shared’’ to retain protein subgroups with either a minimum protein score of 25 or observation in TMT plexes (n=4, 3, or 2 if > 20, 7, or 1 plexes/ cohort, respectively). The primary goal of this step was to eliminate low scoring proteins that were infrequently detected in a cohort. As a consequence of these two proteinpolishing steps, each identified protein reported in the study comprised multiple peptides, unless a single excellent scoring peptide was the sole match and that peptide was observed in multiple TMT-plexes. Subset-specific FDR filtering for germline variant containing peptides in the proteome While peptides in the proteome dataset matched to reference proteome sequences are subject to multi-step, protein-level and cohort level FDR filtering as described above, FDR for subsets of rarely observed (<5% of total) classes of peptides required more stringent score thresholding to reach a suitable subset-specific FDR < 1.0%. To this end, we devised and applied subset-specific filtering approaches. The subset of peptides containing single amino acid variants (SAAVs) and indels observed in the proteome was extracted after step 1 of PSM filtering described above using the SM Protein/Peptide Summary module to create a proteogenomics (PG) site report, with quantitation normalized to nullify the effect of differential protein loading using the aggregate protein-level normalization factors from the fully filtered proteome dataset. Germline variants containing peptides were split up into 4 subsets (SAAVs and indels, with each further split by multiple or single representation in a cohort) and each subset was filtered to <1% FDR. Subsets were thresholded independently in each subset using a 2-step approach. First, PSM scoring metric thresholds were tightened in a fixed manner so that distributions for each metric improved to meet or exceed the aggregate distributions. The fixed thresholds were: minimum score: 7; minimum percent scored peak intensity: 50%; normalized precursor mass error: +/-5 ppm. Second, individual subsets with FDR estimates remaining above 1% were further subject to a grid search to determine the lowest values of backbone cleavage score (sequence coverage metric) and score (fragment ion assignment metric) that improved FDR to < 1% for each subset. Quantitation using TMT ratios Using the SM Protein/Peptide Summary module, a protein comparison report was generated for the proteome dataset using the protein grouping method ‘‘expand subgroups, top uses shared’’ (SGT). For the PTM omes (phosphoproteome and acetylome datasets) Variable Modification site comparison reports limited to either phospho, or acetyl sites, respectively, was generated using the protein grouping method ‘‘unexpand subgroups.’’ Relative abundances of proteins and VM-sites were determined in SM using TMT reporter ion log 2 intensity ratios from each PSM. TMT reporter ion intensities were corrected for isotopic impurities in the SM Protein/Peptide Summary module using the afRICA correction method, which implements determinant calculations according to Cramer’s Rule and correction factors obtained from the reagent manufacturer’s certificate of analysis for each cohort. Each protein-level or PTM sitelevel TMT ratio was calculated as the median of all PSM-level ratios contributing to a protein subgroup or PTM site. PSMs were excluded from the calculation if they lacked a TMT label, had a precursor ion purity < 50% (MS/MS has significant precursor isolation ll OPEN ACCESS e7 Cell 188, 2312–2335.e1–e12, May 1, 2025 Article contamination from co-eluting peptides), or had a negative delta forward-reverse identification score (half of all false-positive identifications). Using the SM Process Report module non-quantifiable proteins and PTM sites (ex: unlabeled peptides containing an acetylated protein N-terminus and ending in arginine rather than lysine) were removed, and median/MAD normalization was performed on each TMT channel in each ome to center and scale the aggregate distribution of protein-level or PTM site-level log-ratios around zero in order to nullify the effect of differential protein loading and/or systematic MS variation. When subsets of an ome (nuORF or SAAVs, etc) the TMT ratios were normalized using the normalization factors for the aggregate distribution of the corresponding ome. It is worth noting that current precision database methods separately quantify different forms of a peptide (reference sequence, variant–containing, phosphorylated, unphosphorylated, etc.) having distinct peptide masses and retention times in TMT labeled ratio-based LC-MS/MS experiments. A TMT labeled experiment is purpose-built to measure ratios of an individual peptide form across samples, which are combined so that each sample in a TMT-plex produces a reporter ion of distinct m/z in each MS/MS spectrum. The TMT reporter ion intensities of the reference sequence and variant-containing forms of a peptide cannot be directly combined to form a single value representing the overall peptide abundance since the MS/MS spectra will have been briefly sampled at different points in their corresponding chromatographic peaks. 187 Proteinor gene-level quantification will mitigate this effect by relying on multiple other wild-type (WT)-only peptides. In contrast, PTM measurements may be more affected since they are usually measured as single peptides. Germline Variants Co-localizing with or Around PTM sites Input data From a total of 27,104,152 germline variants called from WES data, we selected 11,962,341 missense germline variants across our 1,064 samples over 10 cancer types to find germline variants directly co-localizing or nearby a PTM site. As per PTM data, we obtained a total of 141,330 unique phosphorylation sites detected in at least one of the samples in our CPTAC cohort (134,244 on reference peptides and 7,086 on variant peptides affected by germline SAAVs) and 23,756 unique acetylation sites (23,190 on reference peptides and 566 on variant peptides affected by germline SAAVs). Sites detected on the same peptide sequence were considered as separate individual sites yielding a total of 168,423 and 9,018 phosphorylation sites on reference and variant peptides, respectively, and 24,109 and 639 acetylation sites on reference and variant peptides, respectively. Calculation of linear distances Missense variants co-localizing with PTM sites involving serine (S), threonine (T), tyrosine (Y), or lysine (K) codons were cross-referenced in the PTM data for cognate positions. PTM associated germline variants were grouped according to the three types of consequences at the PTM level: (1) an amino acid change caused loss of the PTM site; (2) a variant caused gain of a PTM site not encoded by the reference allele; or, (3) one phosphorylated residue changed to another (such as from a serine to a tyrosine, with phosphorylation detected in both). The ancestral and derived alleles were compiled for all the co-localizing variants. In three specific cases: AHNAK S4516N, FAM83B S729T, and FLG S3174C the reference-associated phosphorylated serine detected in the PTM data was derived from the ancestral annotation (T4516N, P729T, and G3174C, respectively). Therefore, these variants were excluded from the analysis. We also detected variants around a PTM site by calculating the linear distance of missense germline variants relative to PTM sites based on amino acid position as extracted from reference peptides, classifying events using 2 categories: missense variants affecting an amino acid within 5 amino acids of the PTM site were categorized as proximal events; variants affecting amino acids beyond 5 amino acids of the PTM site were categorized as distal. We further confirmed if the amino acid changes predicted from germline variant information matched what was detected in the variant peptide information, when existent. In terms of variants proximal or distal to a site, because most variants distal to a PTM site and a portion of proximal variants fell outside the peptide capture of the PTM site in question, we would not expect to detect a variant-derived peptide for such cases. These direct, proximal, and distal events were used for downstream analyses. Analyses of Direct, Proximal, and Distal Impact of Germline Variants on Protein and PTM Levels We assessed the potential influence of a germline variant direct, proximal, or distal to a PTM site on the overall protein abundance levels using a general linear model approach. We also tested the effects of germline variants on phosphorylation and acetylation levels of reference peptides using the same approach, but only those variants for which the position fell outside the peptide capture of the PTM site in question in order to limit the possibility of bias in the mass spec measures (See limitations of the study and quantitation using TMT ratios STAR Methods section). Therefore, for variants directly overlapping a PTM site, we only tested their impact on the overall protein abundance, not on PTM levels. Common germline variants (gnomAD AF R1%) were tested individually. In the case of low frequency and rare germline variants (gnomAD AF <1%), to increase statistical power, we collapsed all individuals harboring a low frequency/rare variant proximal (within 5 amino acids), or all individuals with a low frequency/rare variant distal (> 5 amino acids) to a PTM site into a single variable, at the gene level. In order to test the pan-cancer differences in protein, phosphorylation, or acetylation levels between carriers and non-carriers of a certain germline variant, we ran the following model to learn the bcoefficients: Y=b0+b1Mv+b2P1+b3P2+b4P3+b5C+e ll OPEN ACCESS Cell 188, 2312–2335.e1–e12, May 1, 2025 e8 Article where Y is a (n x 1) vector representing the protein, phosphorylation, or acetylation abundance of the protein of interest for the site of interest; M is a binary vector indicating the germline variant status for the site of interest (v) for each sample; P 1-3 denote the first three PCs for patient genetic ancestry determination (WES-based); and C is the one-hot encoded cancer type for the samples. The error (e) is assumed to be normally distributed with a constant variance. Tumor samples and matching NAT samples were tested separately. Cancer-type specific analyses were also performed. All resulting p-values were adjusted to FDR using the standard BenjaminiHochberg procedure. The results from these tests are provided in Table S3. Using the same approach as above, we also tested for the effects of highlighted variants from the direct/proximal/distal analyses on their Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway partners’ protein and phosphoprotein abundances. That is, we evaluated ‘‘mTOR signaling’’ for DEPTOR S389N (hsa04150), ‘‘ErbB signaling’’ for ERBB2 P1170A (hsa04012), ‘‘MAPK signaling’’ for MAP2K2 P298L (hsa04010), ‘‘Antigen processing and presentation’’ for HLA-B V69A (hsa04612), ‘‘Apoptosis’’ for CASP8 D344H (hsa04210), and ‘‘Cell cycle’’ for ATRX E929Q (hsa04110). Because MGMT is not a member of any KEGG pathways, it was not tested. We similarly did not test SBDS, which is only in the general "Ribosome biogenesis in eukaryotes" pathway. The analyses were done at both pan-cancer and cancer-specific levels, in which we required at least 5 observations each in variant carriers and non-carriers to test. Resulting p-values were FDR adjusted using the standard Benjamini-Hochberg procedure, and all hits from the general linear model with FDR %0.05 were prioritized for plotting by carrier status. Pairwise Wilcoxon tests between carrier groups were performed for plotting, and FDR adjusted p-values are provided within the boxplots. To determine whether the genes harboring PTM-affecting germline variants exhibited any biological bias, we conducted an overrepresentation analysis of curated pathways from the MiSigDB Hallmark 188,189 set and Wiki Pathways. 190 For genes with variants directly overlapping PTM sites imposing phosphorylation loss and gain, the background gene set was defined as all genes detected in the phosphoproteome data. All acetylated proteins detected in the PTM data were similarly used for background adjustment for genes that experience acetylation site loss or gain. The R package clusterProfiler v4.4.2 was used to conduct these analyses for each PTM type and consequence group separately. Results were constricted to a cutoff of 0.05 FDR adjusted p-value and a q-value cutoff of 0.1. A similar analysis was performed for proximal and distal events. In this case, genes harboring variants proximal or distal to a PTM site were used as the test gene set, testing each group separately. The background gene sets and significance cut-offs were defined as above. HotSpot3D / HotPho analyses Input PTM data Here, we collected information for every PTM site detected on both reference and variant peptides in at least one of the samples in our CPTAC cohort via our analyses of proteomics LC-MS/MS data (See proteomics LC-MS/MS data interpretation STAR Methods section for more details). In total, 8,046 PTM sites (7,353 phosphosites and 693 acetylation sites) were detected on variant peptides affected by germline SNVs or Indels in at least one of our samples. For the purposes of HotSpot3D/HotPho analyses, however, we have excluded PTM sites on variant peptides affected by germline Indels. We obtained 141,330 unique phosphorylation sites detected in at least one of the samples in our CPTAC cohort, from which 134,244 are on reference peptides and 7,086 are on variant peptides affected by germline SAAVs. As per acetylation sites, we obtained 23,756 unique acetylation sites, from which 23,190 are on reference peptides and 566 are on variant peptides affected by germline SAAVs. Further, sites detected on the same peptide sequence were considered as separate individual sites for the purposes of using it as an input for HotSpot3D 72 due to the format required by the tool, yielding a total of 168,423 and 9,018 phosphorylation sites on reference and variant peptides, respectively, and 24,109 and 639 acetylation sites on reference and variant peptides, respectively. Of these, 123,676 phosphorylation sites and 23,646 acetylation sites are unique and were used as inputs for HotSpot3D/ HotPho. To map amino acid residues on different protein isoforms between UniProt Knowledge Base (UniProtKB, version 2023_01) 79 and our dataset, we used Transvar, 175 which allowed us to map them to their unique genomic positions. Input somatic mutation and germline variant data Somatic mutations and germline variants detected from WES, as described above, were filtered for missense single nucleotide events. Therefore, from a total of 345,653 and 27,104,152 exonic somatic mutations and germline variants called from WES data, respectively, we selected 183,503 missense somatic mutations and 11,962,341 missense germline variants across our 1,064 samples over 10 cancer types as inputs for HotSpot3D/HotPho. PDB and AlphaFoldDB structures We used the GRCh38 assembly and Ensembl release 100 (Gencode v34) in order to preprocess residue pair data for all human proteins available in two databases: (1) the RCSB Protein Data Bank (RCSB PDB) 77,78 as of June 24 th , 2021, which contains PDB structures for 7,780 proteins; and (2) the AlphaFold Protein Structure Database (AlphaFoldDB - AFDB) 75,76 v4, as of March 16 th , 2023, which contains predicted protein structures from 19,966 proteins present in Uniprot. For PDB, we filtered out chains or structures due to artifacts, as previously described. 73 For AFDB, HotSpot3D’s algorithm pulls information from the web page version of the database, which provides information for proteins up to 2700 amino acids long. For those proteins which are longer than 2700aa, AFDB provides 1400aa long overlapping fragments, for which only the first 1400aa are available in the webpage version used here. ll OPEN ACCESS e9 Cell 188, 2312–2335.e1–e12, May 1, 2025 Article Quality control As described before, 73 HotSpot3D/HotPho takes as input a file containing all PTM sites of interest containing the following information for each: the HUGO gene symbol, the corresponding Ensembl transcript ID, the protein residue position, and a summarized description of the site (e.g. Phosphoserine, Acetyllysine, etc). This information is then passed through the software, together with the input germline variant and somatic mutation information, to find pairwise relationships between mutations and sites. For the purposes of these analyses, we use the word ‘‘mutations’’ to describe both somatic and germline events. For PDB, because the structures provided by uploaders in the database do not always directly map to the associated Uniprot entries, HotSpot3D/HotPho calculates offsets in residue numbers in PDB structures and transcripts. For AFDB, because we are dealing with computationally predicted structures, the residues at the same position between the database structure and the Uniprot entry may not always perfectly match. Therefore, we have filtered out any sites where the residue provided in the PDB or AFDB structure did not match the residue in the input phosphorylation or acetylation site data, resulting in the following results for each input database: (1) PDB: 41,748 mutation-mutation pairs, 13,072 mutation-site pairs (4,625 excluded), and 11,328 site-site pairs (5,414 excluded); (2) AFDB: 110,255 mutation-mutation pairs; 29,888 mutation-site pairs (3,282 excluded), and 32,946 site-site pairs (4,972 excluded). Cluster discovery and filtering We have implemented HotSpot3D 72 and HotPho 73 to allow for the co-clustering of both missense germline variants and somatic mutations with phosphorylation and acetylation sites on the 3D protein structures (Figure 4A), as previously described. 73 Briefly, we used HotSpot3D to calculate the 3D distances between mutations and PTM sites using structures from PDB, as well as predicted structures from AFDB. During this process, missense variants and PTM sites are considered as nodes and the 3D distances between them as edges on an undirected graph. The clusters are then calculated using the Floyd–Warshall shortest-paths algorithm and using recurrence as the vertex type and clustering distance of 10A ˚, as implemented in HotSpot3D. 72 These analyses yielded a total of 15,132 unfiltered clusters across 4,409 unique proteins using PDB structures (2,084 site-only, 9,558 mutation-only, 3,490 hybrid), and 96,719 unfiltered clusters in 15,655 unique proteins using AFDB structures (14,788 site-only, 62,437 mutation-only, 19,494 hybrid). We further filtered clusters based on the cluster closeness score (Cc), for which a high score indicates a cluster enriched in mutations and PTM sites on the 3D protein structure. Here we use a threshold of top 5% to select high confidence intramolecular clusters for downstream analyses, as described in the previous HotSpot3D and HotPho studies. 72,73 This generated a final set of 210 hybrid, 509 mutation-only, and 111 site-only clusters from PDB and 978 hybrid, 3126 mutation-only, 731 site-only clusters from AFDB. These results are provided in Table S4. Impact on protein abundance analyses We applied a linear model to evaluate the protein abundance level differences between carriers and non-carriers of co-clustered mutations and/or PTM sites within the same intramolecular cluster. We ran the model to learn the bcoefficients as follows: Y=b0+b1Mv+b2P1+b3P2+b4P3+b5C+b5N+e where Y is a (n x 1) vector representing the protein abundance of the protein of interest for the cluster of interest; M is a binary vector indicating the co-clustered status (v) for each sample (i.e. if a sample had any event co-clustered in a particular cluster, it was grouped here); P 1-3 denote the first three PCs for patient genetic ancestry determination (WES-based); C is the one-hot encoded cancer type for the samples, and N is the CNV value for the gene being tested, as determined by GISTIC2. The error (e) is assumed to be normally distributed with a constant variance. Cancer-type specific analyses were also performed in the same way, where we evaluated the effect of germline and somatic variants involved in hybrid clusters on protein abundance levels between carriers and non-carriers to find genetic changes potentially associated with a certain cancer type. Analyses of phosphorylation and acetylation levels were not performed in this case due to the limitations addressed in this manuscript (See limitations of the study). Allele specific expression analysis using RNA-seq data To identify allele specific expression (ASE) events based on RNA-seq, we used 1,057 tumor and 340 NAT samples with available RNA-seq data. For these analyses, we used only SNVs in cancer-related genes (624 cancer related genes 17 ). First, germline variants were filtered to the ones that were detected in either of the three datasets: proteome, phosphoproteome, or acetylome. Next, we calculated read counts for each variant in each sample’s RNA-seq BAM files using bam-readcount (v0.7.4 with parameters -q 10, -b 15, and -i so that reads overlapping with an insertion were not included in the per base counts). We retained only variants with at least 10 read counts covering reference and alternative alleles for this analysis. Then, to identify ASE events, we performed a two-sided binomial test with a null probability of success 0.5 in a Bernoulli experiment. The resulting p-values were adjusted using BH procedure, and ASE events were called significant if they reached FDR<0.05. Indel variant analysis Summary statistics of indel counts were measured according to the germline MAF files (above Methods) and restricted to a large set of cancer related genes as previously described. ll OPEN ACCESS Cell 188, 2312–2335.e1–e12, May 1, 2025 e10 Article Indel positioning was performed by mapping variants to the exons and labeling them according to position (First, Middle, or Last exon). When only 1 or 2 exons made up the composition of a gene, then they were assigned first and last, and no did not receive a middle label. Relative position of the mutation within the gene model was calculated for each gene based on the size of the exon as cataloged by Ensembl v100. 191 The Penultimate region next to the last exon junction (<50bp from the last exon junctions [EJC]) was measured. This was performed for frameshift mutations and predicted inframe mutations as annotated in the germline MAF files (see above STAR Methods -germline variant calling and filtering from WES). Again, using the Ensembl gene annotations, relative positions to the last exon start-position was used to determine whether a mutation was assigned to the penultimate position. Kernel density information was estimated and plotted to identify gene position differences of inframe and frameshift mutations (Figure 6B). We also developed two simple algorithms to discover the impact of these germline variants on protein abundances. The first method seeks to determine the impact of indels by looking at the upstream and downstream peptide-level abundances. Simply stated, we used a t-test as the crux of the first analysis. Second, we sought to find mutations that had an effect on protein abundance with respect to the RNA expression. Below we outline the implementation of a multi-omic LDA (moLDA) analysis to accomplish this objective. We used the following criteria to discover variants that had variable upstream and downstream consequences of indels. First, we restricted our indel variants to those that had a predicted frameshift, splice-region, protein-altering designation according to VEP annotations (see above STAR Methods -germline variant calling and filtering from WES). Next, we restricted our search to variants that were observed in at least 20 samples. We ensured only variants with at least 6 measured peptides, upand down-stream, were included. We then split the data based on whether there was a significant difference between upstream peptide abundances to downstream peptide abundances using a t-test. P-values and 95% confidence intervals for all indels and genes that met these criteria are provided in Table S6. The second strategy we implemented to identify the role of indels on protein variability was to leverage an assumed relationship between RNA expression and protein abundance to find examples where mutations clearly associated with an expected relationship. To achieve this objective we implemented a multi-omic linear discriminant analysis (LDA) to classify indel status based on RNA and protein abundance. Briefly, LDA is a statistical method used for classifying or predicting the group membership of observations based on a set of predictor variables. It aims to find a linear combination of predictors that maximally separates two differentiating groups. Here the groups are defined as indel carriers and non-carriers and the predictors are protein abundance and RNA expression. First, we ensured that more than 30 samples had both RNA and protein abundance measurements for a given gene (in cis). Next, we excluded all mutations that didn’t have at least 6 samples with the mutations and at least 6 samples without the mutations. Following a data integration step to merge RNA expression with protein abundance we used the ‘lda‘ function as part of the MASS R library to find linear combinations of protein and RNA that segregated based on mutation status (Figure 6E). Genes and mutations were prioritized based on their singular value decomposition (SVD) scores which provide higher scores for improved separation between predictors Table S6. Identification of expression and protein quantitative trait loci (eQTLs and pQTLs) We performed quantitative trait loci (QTL) mapping to identify common germline genetic variants that affect gene expression (eQTL) and protein abundance (pQTL) in tumor and normal tissues utilizing the linear regression model in MatrixeQTL. 192 For this purpose, we used WGS germline SNPs with MAF R5% and included gender and ten principal components as covariates to adjust for population stratification. We analyzed the data on each cancer and tissue separately. Specifically, we conducted the eQTL and pQTL analyses for tumor and normal tissues of ccRCC, HNSCC, LSCC LUAD and PDAC for which both gene expression and protein abundance data are available (except for eQTL analysis on normal tissue of PDAC patients due to the limited number of samples with normal data). For the eQTL analysis, we utilized the FPKM normalized gene expression generated from the RNA-Seq data as discussed in the Pan-Cancer Data and Resource and Pan-Cancer Driver manuscripts, 37,160 and further performed TPM conversion, quantile normalization, and inverse normal transformation to remove technical noises and allow cross-sample comparisons. The eQTL analysis included individuals for whom genotype and gene expression data were available and genes with TPM > 0.1 in at least 20% of samples (Table S7A). To eliminate the hidden determinants in the expression data, we additionally selected 15 PEER factors as covariates using PEER software. 193 The pQTL analysis included individuals for whom genotype and protein abundance were available and proteins with data in at least 20% of samples (Table S7A). We deemed QTLs at FDR %1% as significant and considered variants within 1 Mb of a genes’ transcription start site as cis-QTLs. The significant eQTLs can be viewed at https://immuneregulation. mssm.edu/. 168 Furthermore, we performed overall survival analysis based on the expression of interesting genes (ERAP2,HLADQB1 and PPIL3) in the ccRCC, HNSCC, LSCC, and LUAD CPTAC cohorts using the best cutoff with Kaplan-Meier Plotter. 194 We also used Kaplan-Meier Plotter to assess the correlation between the expression and overall survival in the TCGA cohort. Colocalization Analysis of eQTLs and pQTLs We performed colocalization analysis to determine whether the leading variants among the ciseQTLs and pQTLs are the same in certain genes of interest using ‘coloc’ R package’s coloc.Abf function. 195 We applied the default values for the prior probabilities for a SNP being associated with gene expression only, protein abundance only and with both. ll OPEN ACCESS e11 Cell 188, 2312–2335.e1–e12, May 1, 2025 Article Polygenic Risk Scores and associations with protein abundance Summary statistics, including risk allele, protective allele, odds ratio (OR) and annotated gene, were obtained for the largest genomewide association study available for each cancer type. This included ccRCC, PDAC, UCEC, GBM, LUAD, and LSCC for a total of 133 risk variants (Table S7H). The same GWAS study was used for LUAD and LSCC as this discovery study comprised a balanced mixture of cases of both lung cancer subtypes. Polygenic risk scores (PRS) for each cancer type were calculated using the score routine available in PLINK 2.0, weighting the allele dosage at each variant by the effect size. In a first pass, we checked the discriminatory power of the PRS by integrating CPTAC and UKBB datasets. To control for population structure, for these specific analyses we selected individuals of European ancestry in both datasets. For each cancer type, we compared the PRSs of the corresponding subtype with: a) patients of other cancer types in CPTAC; b) individuals with cancer diagnosis in the UKBB; and c) individuals without cancer diagnosis in the UKBB. Three out of five tested cancers, namely PDAC, GBM, and LSCC, showed significantly higher PRSs in the corresponding CPTAC patients compared to controls (Figure S7B). We focused on these three cancers in subsequent analyses. We identified tumor proteins that showed significant association with PRS using linear models. To avoid the effects of hidden variables inducing covariance in the protein abundance matrix, we first carried out a principal component analysis. Considering the relatively large number of potential covariates to the sample size in our study, we performed a supervised selection of covariates to be included in the linear models. We tested the correlation between the PRS as well as the first ten principal components (PCs) of protein abundance with relevant clinical, demographic and molecular variables, including: genetic ancestry (first 10 PCs), age at diagnosis, sex, tumor purity, and smoking in the case of lung cancers. In the linear models we only included as covariates those showing significant correlation with the corresponding PRS and/or proteomic PCs. Given our interest in germline variants (which are present in the different tumor compartments), tumor purity was not included in any of the models despite significant correlation. We excluded proteins with more than 20% of missing data across individuals. The effect of the PRS was estimated for each protein using lm() function in R with the following designs: lmprotein PRS GMB +AncestryPC1 +AncestryPC2 +AncestryPC3 +AncestryPC5 +AncestryPC7 +AncestryPC10 +Age +ProteinPC1 +ProteinPC2 +ProteinPC3 +ProteinPC4 +ProteinPC8 lmprotein PRS LSCC +AncestryPC6 +ProteinPC1 +ProteinPC3 +ProteinPC7 +ProteinPC9 lmprotein PRS PDAC +AncestryPC1 +AncestryPC2 +AncestryPC4 +AncestryPC5 +Age +ProteinPC2 +ProteinPC3 +ProteinPC4 +ProteinPC5 +ProteinPC6 False discovery rates were estimated from the p-values using the fdrtool R package. 196 STRINGdb 173 R package (v11.5; https:// www.string-db.org) was used to infer the protein-protein interaction networks and compute enrichments for the number of interactions among the top proteins associated with PRS. The database contains information for 19,566 proteins and over 2.9 million interactions. Over 93% of queried proteins were present in the STRING dataset. We performed Gene Set Enrichment Analyses (GSEA) analyses for Reactome pathways using the R package ReactomePA 197 with 10,000 permutations and significance threshold of 0.05 with BH FDR adjustment. Disease free survival and overall survival plots were generated using survminer (v0.4.9; https://github.com/ kassambara/survminer) and survival (https://github.com/therneau/survival) R packages, stratifying the patients according to the median of the PRS scores. ADDITIONAL RESOURCES Comprehensive information about the CPTAC program, including program initiatives, investigators, and datasets, are available at the CPTAC program website: https://proteomics.cancer.gov/programs/cptac. For the Pan-Cancer proteogenomics collection papers, along with links to the data and supplementary materials associated with these publications, please visit the Proteomic Data Commons (PDC) at https://pdc.cancer.gov/pdc/cptac-pancancer and the Cancer Research Data Commons at https://dataservice.datacommons.cancer.gov/#/data. ll OPEN ACCESS Cell 188, 2312–2335.e1–e12, May 1, 2025 e12 Article Supplemental figures (legend on next page) ll OPEN ACCESS Article Figure S1. Sample and variant quality control procedures, related to Figure 1 (A) Coverage distribution over WES target regions for the entire CPTAC cohort (n= 1,064). WES samples with coverage R203were used for germline variant calling. (B) Average WES coverage of 160 CPGs for the entire cohort. Data are represented as average coverage ±1 standard deviation (mean ±SD). (C) Number of exonic germline variants detected in normal samples from CPTAC WES data. Individuals are represented by dots, which are colored according to their genetic ancestry as predicted by our pipelines. Average number of variants per cancer type ±one standard deviation is also shown (mean ±SD). (D) Boxplots representing the concordance of whole-exome-based variant calls with dbSnP (release 151), showing >99% concordance. Overall concordance for the entire cohort was 97.43%, with a TiTv ratio of 2.74. (E) Overlap rate between variants called from WES and WGS datasets for seven cancer types. (F) Principal-component analysis (PCA) plots showing WES and WGS-based genetic ancestry predictions for the CPTAC cohort. Ancestry predictions were obtained from WES data using a random forest classifier for all 1,064 individuals (STAR Methods). The 9 individuals of Slavic origin in the GBM, HNSCC, LSCC, PDAC, and UCEC cohorts which were misclassified as AMR in the WES-based predictions, but correctly classified as EUR in the WGS-based predictions are labeled (see STAR Methods). (G) Histograms depicting the peptide-length distribution of both reference (top) and alternative (bottom) peptides detected in the proteome (left), phosphoproteome (middle), and acetylome (right). ll OPEN ACCESS Article Figure S2. Impact of pathogenic rare variants, related to Figure 2 (A) Violin plots showing age at diagnosis distributions in carriers and non-carriers of rare pathogenic (P) and likely pathogenic (LP) germline variants across 10 cancer types. (B) Heatmap showing fraction of samples that are carriers of P/LP rare variants across 10 cancer types in the combined sample set of CPTAC and TCGA cohorts. Cancer gene pairs with significant (FDR %0.05) and suggestive (FDR %0.15) enrichment with P/LP variants are indicated with black and gray outline respectively. (legend continued on next page) ll OPEN ACCESS Article (C) Plot showing comparison of variant allele frequencies (VAFs) for P/LP in tumor and normal samples, highlighting variants undergoing LOH in the tumor. Each dot corresponds to one variant, with the diagonal line indicating equal tumor and normal VAFs (i.e., neutral selection). Green indicates suggestive LOH (FDR %0.15); red is significant LOH (FDR %0.05); and blue indicates events not statistically significant. (D) Plot showing protein expression quantiles in NAT (x axis) and tumor (y axis) of the proteins from germline P/LP variant carriers. ll OPEN ACCESS Article (legend on next page) ll OPEN ACCESS Article Figure S6. EJ models for indels and moLDA summary, related to Figure 6 (A) Similar to the main Figure 6, EJ models are shown for each cancer type highlighting the first, middle, and last exons. Two penultimate regions within 50 bps of the last exon and greater than 50 bps away from the last exon are also displayed. (B) Square-root normalized distribution of moLDA SVD scores for each cancer type. Vertical threshold is set at the PanCancer top 2.5% of scores. (C) Genes that exceed the SVD threshold that fall into multiple cancer types. (D) moLDA result for CPNE1 in LUAD. Log normalized RNA-seq gene expression on the x axis and protein abundance on the y axis. Each point represents a tumor from that cancer type. (E) Same as (D) but for the OAS1 gene in LSCC. (F) Same as (D) but for ITIH1 gene in HNSCC. ll OPEN ACCESS Article (legend on next page) ll OPEN ACCESS Article Figure S7. Survival plots and PRS distributions, related to Figure 7 (A) Kaplan-Meier survival plots based on ERAP2 and HLA-DQB1 expression and overall survival in CPTAC and TCGA HNSCC cohorts. (B) Polygenic risk score (PRS) distributions. For 6 cancer types, we calculated the PRS in the CPTAC individuals using common risk variants discovered by the largest GWAS available in each specific cancer type. These values were compared with the distributions of PRSs in three other groups, namely (1) CPTAC individuals for the remaining cancer types (‘‘CPTAC’’), (2) UKBB individuals diagnosed with any cancer type (‘‘Ukbb_cancer’’, and (3) rest of UKBB individuals (‘‘Ukbb_controls’’). The pvalues for statistical significance for the comparisons against CPTAC and Ukbb_controls, respectively, are provided for each cancer type (t test). ll OPEN ACCESS Article