scieee AI-readable full text Open interactive document viewer

Polygenic Risk Scores for Prediction of Breast Cancer and Breast Cancer Subtypes

Mavaddat, Nasim,Michailidou, Kyriaki,Dennis, Joe,Jukkola, Arja

Full text

ARTICLE Polygenic Risk Scores for Prediction of Breast Cancer and Breast Cancer Subtypes Nasim Mavaddat, 1, *Kyriaki Michailidou, 1,2 Joe Dennis, 1 Michael Lush, 1 Laura Fachal, 3 Andrew Lee, 1 Jonathan P. Tyrer, 3 Ting-Huei Chen, 4 Qin Wang, 1 Manjeet K. Bolla, 1 Xin Yang, 1 Muriel A. Adank, 5 Thomas Ahearn, 6 Kristiina Aittoma ¨ki, 7 Jamie Allen, 1 Irene L. Andrulis, 8,9 Hoda Anton-Culver, 10 Natalia N. Antonenkova, 11 Volker Arndt, 12 Kristan J. Aronson, 13 Paul L. Auer, 14,15 Pa ¨ivi Auvinen, 16,17,18 Myrto Barrdahl, 19 Laura E. Beane Freeman, 6 Matthias W. Beckmann, 20 Sabine Behrens, 19 Javier Benitez, 21,22 Marina Bermisheva, 23 Leslie Bernstein, 24 Carl Blomqvist, 25,26 Natalia V. Bogdanova, 11,27,28 Stig E. Bojesen, 29,30,31 Bernardo Bonanni, 32 Anne-Lise Børresen-Dale, 33,34 Hiltrud Brauch, 35,36,37 Michael Bremer, 27 Hermann Brenner, 12,37,38 Adam Brentnall, 39 Ian W. Brock, 40 Angela Brooks-Wilson, 41,42 Sara Y. Brucker, 43 Thomas Bru ¨ning, 44 Barbara Burwinkel, 45,46 Daniele Campa, 19,47 Brian D. Carter, 48 Jose E. Castelao, 49 Stephen J. Chanock, 6 Rowan Chlebowski, 50 Hans Christiansen, 27 Christine L. Clarke, 51 J. Margriet Colle ´e, 52 Emilie Cordina-Duverger, 53 Sten Cornelissen, 54 Fergus J. Couch, 55 Angela Cox, 40 Simon S. Cross, 56 Kamila Czene, 57 (Author list continued on next page) Stratification of women according to their risk of breast cancer based on polygenic risk scores (PRSs) could improve screening and prevention strategies. Our aim was to develop PRSs, optimized for prediction of estrogen receptor (ER)-specific disease, from the largest available genome-wide association dataset and to empirically validate the PRSs in prospective studies. The development dataset comprised 94,075 case subjects and 75,017 control subjects of European ancestry from 69 studies, divided into training and validation sets. Samples were genotyped using genome-wide arrays, and single-nucleotide polymorphisms (SNPs) were selected by stepwise regression or lasso penalized regression. The best performing PRSs were validated in an independent test set comprising 11,428 case subjects and 18,323 control subjects from 10 prospective studies and 190,040 women from UK Biobank (3,215 incident breast cancers). For the best PRSs (313 SNPs), the odds ratio for overall disease per 1 standard deviation in ten prospective studies was 1.61 (95%CI: 1.57–1.65) with area under receiver-operator curve (AUC) ¼0.630 (95%CI: 0.628–0.651). The lifetime risk of overall breast cancer in the top centile of the PRSs was 32.6%. Compared with women in the middle quintile, those in the highest 1% of risk had 4.37and 2.78-fold risks, and those in the lowest 1% of risk had 0.16and 0.27-fold risks, of developing ER-positive and ER-negative disease, respectively. Goodness-of-fit tests indicated that this PRS was well calibrated and predicts disease risk accurately in the tails of the distribution. This PRS is a powerful and reliable predictor of breast cancer risk that may improve breast cancer prevention programs. Introduction Breast cancer is the most common cancer diagnosed among women in Western countries. While rare mutations in genes such as BRCA1 and BRCA2 confer high risks of developing breast cancer, these account for only a small proportion of breast cancer cases in the general population. Multiple common breast cancer susceptibility variants discovered through genome-wide association studies (GWASs) 1,2 confer small risk individually, but their combined effect, when 1 Centre for Cancer Genetic Epidemiology, Department of Public Health and Primary Care, University of Cambridge, Cambridge CB1 8RN, UK; 2 Department of Electron Microscopy/Molecular Pathology, The Cyprus Institute of Neurology and Genetics, 1683 Nicosia, Cyprus; 3 Centre for Cancer Genetic Epidemiology, Department of Oncology, University of Cambridge, Cambridge CB1 8RN, UK; 4 Department of Mathematics and Statistics, Laval University, Que ´bec City, QC G1V 0A6, Canada; 5 Family Cancer Clinic, the Netherlands Cancer Institute - Antoni van Leeuwenhoek hospital, Amsterdam, 1066 CX, the Netherlands; 6 Division of Cancer Epidemiology and Genetics, National Cancer Institute, National Institutes of Health, Department of Health and Human Services, Bethesda, MD 20850, USA; 7 Department of Clinical Genetics, Helsinki University Hospital, University of Helsinki, Helsinki 00290, Finland; 8 Fred A. Litwin Center for Cancer Genetics, Lunenfeld-Tanenbaum Research Institute of Mount Sinai Hospital, Toronto, ON M5G 1X5, Canada; 9 Department of Molecular Genetics, University of Toronto, Toronto, ON M5S 1A8, Canada; 10 Department of Epidemiology, Genetic Epidemiology Research Institute, University of California Irvine, Irvine, CA 92617, USA; 11 NN Alexandrov Research Institute of Oncology and Medical Radiology, Minsk 223040, Belarus; 12 Division of Clinical Epidemiology and Aging Research, German Cancer Research Center (DKFZ), Heidelberg 69120, Germany; 13 Department of Public Health Sciences, and Cancer Research Institute, Queen’s University, Kingston, ON K7L 3N6, Canada; 14 Cancer Prevention Program, Fred Hutchinson Cancer Research Center, Seattle, WA 98109, USA; 15 Zilber School of Public Health, University of Wisconsin-Milwaukee, Milwaukee, WI 53205, USA; 16 Cancer Center, Kuopio University Hospital, Kuopio 70210, Finland; 17 Institute of Clinical Medicine, Oncology, University of Eastern Finland, Kuopio 70210, Finland; 18 Translational Cancer Research Area, University of Eastern Finland, Kuopio 70210, Finland; 19 Division of Cancer Epidemiology, German Cancer Research Center (DKFZ), Heidelberg 69120, Germany; 20 Department of Gynecology and Obstetrics, University Hospital Erlangen, Friedrich-Alexander University Erlangen-Nuremberg, Comprehensive Cancer Center Erlangen-EMN, Erlangen 91054, Germany; 21 Human Cancer Genetics Programme, Spanish National Cancer Research Centre (CNIO), Madrid 28029, Spain; 22 Biomedical Network on Rare Diseases (CIBERER), Madrid 28029, Spain; 23 Institute of Biochemistry and Genetics, Ufa Scientific Center of Russian Academy of Sciences, Ufa 450054, Russia; 24 Department of Population Sciences, Beckman Research Institute of City of Hope, Duarte, CA 91010, USA; 25 Department of Oncology, Helsinki University Hospital, University of Helsinki, Helsinki 00290, Finland; 26 Department of Oncology, O ¨rebro University Hospital, O ¨rebro 70185, Sweden; 27 Department of Radiation Oncology, Hannover Medical (Affiliations continued on next page) The American Journal of Human Genetics 104, 21–34, January 3, 2019 21 Ó2018 The Authors. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). summarized as a polygenic risk score (PRS), can be substantial. 3–5 Such genomic profiles can be used to stratify women according to their risk of developing breast cancer. 6 This in turn holds the promise of improved breast cancer prevention and survival, by targeting screening or other preventative strategies at those women most likely to benefit. We previously derived a PRS based on 77 established breast cancer susceptibility single-nucleotide polymorphisms (SNPs) and reported levels of risk stratification achieved by this PRS. 7 Based on our findings, several studies have investigated the potential for combining PRSs and other known risk factors for risk stratification and evaluated the impact of risk reduction strategies across risk strata defined by the PRS. 8–10 Preliminary studies investigating the use of the PRS to inform targeted breast cancer screening programs are underway (see CORDIS Mary B. Daly, 58 Peter Devilee, 59,60 Thilo Do ¨rk, 28 Isabel dos-Santos-Silva, 61 Martine Dumont, 62 Lorraine Durcan, 63,64 Miriam Dwek, 65 Diana M. Eccles, 64 Arif B. Ekici, 66 A. Heather Eliassen, 67,68 Carolina Ellberg, 69 Christoph Engel, 70,71 Mikael Eriksson, 57 D. Gareth Evans, 72,73 Peter A. Fasching, 20,74 Jonine Figueroa, 6,75,76 Olivia Fletcher, 77 Henrik Flyger, 78 Asta Fo ¨rsti, 79,80 Lin Fritschi, 81 Marike Gabrielson, 57 Manuela Gago-Dominguez, 82,83 Susan M. Gapstur, 48 Jose ´A. Garcı ´a-Sa ´enz, 84 Mia M. Gaudet, 48 Vassilios Georgoulias, 85 Graham G. Giles, 86,87,88 Irina R. Gilyazova, 23,89 Gord Glendon, 8 Mark S. Goldberg, 90,91 David E. Goldgar, 92 Anna Gonza ´lez-Neira, 21 Grethe I. Grenaker Alnæs, 33 Mervi Grip, 93 Jacek Gronwald, 94 Anne Grundy, 95 Pascal Gue ´nel, 53 Lothar Haeberle, 20 Eric Hahnen, 96,97 Christopher A. Haiman, 98 Niclas Ha ˚kansson, 99 Ute Hamann, 100 Susan E. Hankinson, 101 Elaine F. Harkness, 102,103,104 Steven N. Hart, 105 Wei He, 57 Alexander Hein, 20 Jane Heyworth, 106 Peter Hillemanns, 28 Antoinette Hollestelle, 107 Maartje J. Hooning, 107 Robert N. Hoover, 6 John L. Hopper, 87 Anthony Howell, 108 Guanmengqian Huang, 100 Keith Humphreys, 57 David J. Hunter, 68,109,110 (Author list continued on next page) School, Hannover 30625, Germany; 28 Gynaecology Research Unit, Hannover Medical School, Hannover 30625, Germany; 29 Copenhagen General Population Study, Herlev and Gentofte Hospital, Copenhagen University Hospital, Herlev 2730, Denmark; 30 Department of Clinical Biochemistry, Herlev and Gentofte Hospital, Copenhagen University Hospital, Herlev 2730, Denmark; 31 Faculty of Health and Medical Sciences, University of Copenhagen, Copenhagen 2200, Denmark; 32 Division of Cancer Prevention and Genetics, IEO, European Institute of Oncology IRCCS, Milan 20141, Italy; 33 Department of Cancer Genetics, Institute for Cancer Research, Oslo University Hospital-Radiumhospitalet, Oslo 0379, Norway; 34 Institute of Clinical Medicine, Faculty of Medicine, University of Oslo, Oslo 0450, Norway; 35 Dr Margarete Fischer-Bosch-Institute of Clinical Pharmacology, Stuttgart 70376, Germany; 36 University of Tu ¨bingen, Tu ¨bingen 72074, Germany; 37 German Cancer Consortium (DKTK), German Cancer Research Center (DKFZ), Heidelberg 69120, Germany; 38 Division of Preventive Oncology, German Cancer Research Center (DKFZ) and National Center for Tumor Diseases (NCT), Heidelberg 69120, Germany; 39 Centre for Cancer Prevention, Wolfson Institute of Preventive Medicine, Queen Mary University of London, London EC1M 6BQ, UK; 40 Sheffield Institute for Nucleic Acids (SInFoNiA), Department of Oncology and Metabolism, University of Sheffield, Sheffield S10 2TN, UK; 41 Genome Sciences Centre, BC Cancer Agency, Vancouver, BC V5Z 1L3, Canada; 42 Department of Biomedical Physiology and Kinesiology, Simon Fraser University, Burnaby, BC V5A 1S6, Canada; 43 Department of Gynecology and Obstetrics, University of Tu ¨bingen, Tu ¨bingen 72076, Germany; 44 Institute for Prevention and Occupational Medicine of the German Social Accident Insurance, Institute of the Ruhr University Bochum (IPA), Bochum 44789, Germany; 45 Department of Obstetrics and Gynecology, University of Heidelberg, Heidelberg 69120, Germany; 46 Molecular Epidemiology Group, C080, German Cancer Research Center (DKFZ), Heidelberg 69120, Germany; 47 Department of Biology, University of Pisa, Pisa 56126, Italy; 48 Epidemiology Research Program, American Cancer Society, Atlanta, GA 30303, USA; 49 Oncology and Genetics Unit, Instituto de Investigacion Sanitaria Galicia Sur (IISGS), Xerencia de Xestion Integrada de Vigo-SERGAS, Vigo 36312, Spain; 50 Division of Medical Oncology and Hematology, University of California at Los Angeles, Los Angeles, CA 90024, USA; 51 Westmead Institute for Medical Research, University of Sydney, Sydney, NSW 2145, Australia; 52 Department of Clinical Genetics, Erasmus University Medical Center, Rotterdam 3015 CN, the Netherlands; 53 Cancer & Environment Group, Center for Research in Epidemiology and Population Health (CESP), INSERM, University Paris-Sud, University Paris-Saclay, Villejuif 94805, France; 54 Division of Molecular Pathology, the Netherlands Cancer Institute - Antoni van Leeuwenhoek Hospital, Amsterdam 1066 CX, the Netherlands; 55 Department of Laboratory Medicine and Pathology, Mayo Clinic, Rochester, MN 55905, USA; 56 Academic Unit of Pathology, Department of Neuroscience, University of Sheffield, Sheffield S10 2TN, UK; 57 Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Stockholm 171 65, Sweden; 58 Department of Clinical Genetics, Fox Chase Cancer Center, Philadelphia, PA 19111, USA; 59 Department of Pathology, Leiden University Medical Center, Leiden 2333 ZA, the Netherlands; 60 Department of Human Genetics, Leiden University Medical Center, Leiden 2333 ZA, the Netherlands; 61 Department of Non-Communicable Disease Epidemiology, London School of Hygiene and Tropical Medicine, London WC1E 7HT, UK; 62 Genomics Center, Centre Hospitalier Universitaire de Que ´bec - Universite ´Laval Research Center, Universite ´Laval, Que ´bec City, QC G1V 4G2, Canada; 63 Southampton Clinical Trials Unit, Faculty of Medicine, University of Southampton, Southampton SO17 6YD, UK; 64 Cancer Sciences Academic Unit, Faculty of Medicine, University of Southampton, Southampton SO17 6YD, UK; 65 School of Life Sciences, University of Westminster, London W1B 2HW, UK; 66 Institute of Human Genetics, University Hospital Erlangen, Friedrich-Alexander University Erlangen-Nuremberg, Comprehensive Cancer Center Erlangen-EMN, Erlangen 91054, Germany; 67 Channing Division of Network Medicine, Department of Medicine, Brigham and Women’s Hospital, Harvard Medical School, Boston, MA 02115, USA; 68 Department of Epidemiology, Harvard TH Chan School of Public Health, Boston, MA 02115, USA; 69 Department of Cancer Epidemiology, Clinical Sciences, Lund University, Lund 222 42, Sweden; 70 Institute for Medical Informatics, Statistics and Epidemiology, University of Leipzig, Leipzig 04107, Germany; 71 LIFE - Leipzig Research Centre for Civilization Diseases, University of Leipzig, Leipzig 04103, Germany; 72 Division of Evolution and Genomic Sciences, School of Biological Sciences, Faculty of Biology, Medicine and Health, University of Manchester, Manchester Academic Health Science Centre, Manchester M13 9WL, UK; 73 North West Genomic Laboratory Hub, Manchester Centre for Genomic Medicine, St Mary’s Hospital, Manchester University NHS Foundation Trust, Manchester Academic Health Science Centre, Manchester M13 9WL, UK; 74 David Geffen School of Medicine, Department of Medicine Division of Hematology and Oncology, University of California at Los Angeles, Los Angeles, CA 90095, USA; 75 Usher Institute of Population Health Sciences and Informatics, The University of Edinburgh Medical School, Edinburgh EH16 4TJ, UK; 76 Cancer Research UK Edinburgh Centre, Edinburgh EH4 2XR, UK; 77 The Breast Cancer Now Toby Robins Research Centre, The Institute of Cancer Research, London SW7 3RP, UK; 78 Department of Breast Surgery, Herlev and Gentofte Hospital, Copenhagen University Hospital, Herlev 2730, Denmark; 79 Division of Molecular Genetic Epidemiology, German Cancer Research Center (DKFZ), Heidelberg 69120, Germany; 80 Center for Primary Health Care Research, Clinical Research Center, Lund University, Malmo ¨205 02, Sweden; 81 School of Public Health, Curtin University, Perth, WA 6102, (Affiliations continued on next page) 22 The American Journal of Human Genetics 104, 21–34, January 3, 2019 and GenomeCanada in Web Resources). 11,12 Empirical validation and characterization of the PRS in large-scale epidemiological studies has, however, not been carried out previously. In addition, more informative PRSs would improve the clinical utility of risk prediction. GWASs have now identified 170 breast cancer susceptibility loci. 1,2 Moreover, genome-wide heritability estimates indicate that these loci explain only 40% of the heritability explained by all common variants on genome-wide SNP arrays. This suggests that the discrimination provided by the PRS could be improved by incorporating variants associated at more liberal significance thresholds. In addition, many variants confer risks that differ by breast cancer subtype (estrogen-receptor [ER]-positive or -negative), suggesting that subtype-specific PRSs might allow better prediction of subtype-specific disease, including the more aggressive ER-negative breast cancer, and enable selection of women for preventative medication. Milena Jakimovska, 111 Anna Jakubowska, 94,112 Wolfgang Janni, 113 Esther M. John, 114 Nichola Johnson, 77 Michael E. Jones, 115 Arja Jukkola-Vuorinen, 116 Audrey Jung, 19 Rudolf Kaaks, 19 Katarzyna Kaczmarek, 94 Vesa Kataja, 18,117 Renske Keeman, 54 Michael J. Kerin, 118 Elza Khusnutdinova, 23,89 Johanna I. Kiiski, 119 Julia A. Knight, 120,121 Yon-Dschun Ko, 122 Veli-Matti Kosma, 18,123,124 Stella Koutros, 6 Vessela N. Kristensen, 33,34 Ute Kru ¨ger, 69 Tabea Ku ¨hl, 125 Diether Lambrechts, 126,127 Loic Le Marchand, 128 Eunjung Lee, 98 Flavio Lejbkowicz, 129 Jenna Lilyquist, 105 Annika Lindblom, 130 Sara Lindstro ¨m, 131,132 Jolanta Lissowska, 133 Wing-Yee Lo, 35,36 Sibylle Loibl, 134 Jirong Long, 135 Jan Lubi nski, 94 Michael P. Lux, 20 Robert J. MacInnis, 86,87 Tom Maishman, 63,64 Enes Makalic, 87 Ivana Maleva Kostovska, 111 Arto Mannermaa, 18,123,124 Siranoush Manoukian, 136 Sara Margolin, 137,138 John W.M. Martens, 107 Maria Elena Martinez, 83,139 Dimitrios Mavroudis, 85 Catriona McLean, 140 Alfons Meindl, 141 Usha Menon, 142 Pooja Middha, 19,143 Nicola Miller, 118 Fernando Moreno, 84 Anna Marie Mulligan, 144,145 Claire Mulot, 146 Australia; 82 Genomic Medicine Group, Galician Foundation of Genomic Medicine, Instituto de Investigacio ´n Sanitaria de Santiago de Compostela (IDIS), Complejo Hospitalario Universitario de Santiago, SERGAS, Santiago de Compostela 15706, Spain; 83 Moores Cancer Center, University of California San Diego, La Jolla, CA 92093, USA; 84 Medical Oncology Department, Hospital Clı ´nico San Carlos, Instituto de Investigacio ´n Sanitaria San Carlos (IdISSC), Centro Investigacio ´n Biome ´dica en Red de Ca ´ncer (CIBERONC), Madrid 28040, Spain; 85 Department of Medical Oncology, University Hospital of Heraklion, Heraklion 711 10, Greece; 86 Cancer Epidemiology & Intelligence Division, Cancer Council Victoria, Melbourne, VIC 3004, Australia; 87 Centre for Epidemiology and Biostatistics, Melbourne School of Population and Global Health, The University of Melbourne, Melbourne, VIC 3010, Australia; 88 Department of Epidemiology and Preventive Medicine, Monash University, Melbourne, VIC 3004, Australia; 89 Department of Genetics and Fundamental Medicine, Bashkir State University, Ufa 450076, Russia; 90 Department of Medicine, McGill University, Montre ´al, QC H4A 3J1, Canada; 91 Division of Clinical Epidemiology, Royal Victoria Hospital, McGill University, Montre ´al, QC H4A 3J1, Canada; 92 Department of Dermatology and Huntsman Cancer Institute, University of Utah School of Medicine, Salt Lake City, UT 84112, USA; 93 Department of Surgery, Oulu University Hospital, University of Oulu, Oulu 90220, Finland; 94 Department of Genetics and Pathology, Pomeranian Medical University, Szczecin 71-252, Poland; 95 Centre de Recherche du Centre Hospitalier de Universite ´de Montre ´al (CHUM), Universite ´de Montre ´al, Montre ´al, QC H2X 0A9, Canada; 96 Center for Hereditary Breast and Ovarian Cancer, University Hospital of Cologne, Cologne 50937, Germany; 97 Center for Molecular Medicine Cologne (CMMC), University of Cologne, Cologne 50931, Germany; 98 Department of Preventive Medicine, Keck School of Medicine, University of Southern California, Los Angeles, CA 90033, USA; 99 Institute of Environmental Medicine, Karolinska Institutet, Stockholm 171 77, Sweden; 100 Molecular Genetics of Breast Cancer, German Cancer Research Center (DKFZ), Heidelberg 69120, Germany; 101 Department of Biostatistics & Epidemiology, University of Massachusetts, Amherst, Amherst, MA 1003, USA; 102 Division of Informatics, Imaging and Data Sciences, Faculty of Biology, Medicine and Health, University of Manchester, Manchester Academic Health Science Centre, Manchester M13 9PT, UK; 103 Nightingale Breast Screening Centre, Wythenshawe Hospital, Manchester University NHS Foundation Trust, Manchester M23 9LT, UK; 104 NIHR Manchester Biomedical Research Centre, Manchester University NHS Foundation Trust, Manchester Academic Health Science Centre, Manchester M13 9WL, UK; 105 Department of Health Sciences Research, Mayo Clinic, Rochester, MN 55905, USA; 106 School of Population and Global Health, University of Western Australia, Perth, WA 6009, Australia; 107 Department of Medical Oncology, Family Cancer Clinic, Erasmus MC Cancer Institute, Rotterdam 3015 CN, the Netherlands; 108 Division of Cancer Sciences, University of Manchester, Manchester M13 9PL, UK; 109 Program in Genetic Epidemiology and Statistical Genetics, Harvard TH Chan School of Public Health, Boston, MA 02115, USA; 110 Nuffield Department of Population Health, University of Oxford, Oxford OX3 7LF, UK; 111 Research Centre for Genetic Engineering and Biotechnology ‘‘Georgi D. Efremov,’’ Macedonian Academy of Sciences and Arts, Skopje 1000, Republic of Macedonia; 112 Independent Laboratory of Molecular Biology and Genetic Diagnostics, Pomeranian Medical University, Szczecin 71-252, Poland; 113 Department of Gynecology and Obstetrics, University Hospital Ulm, Ulm 89075, Germany; 114 Department of Medicine, Division of Oncology, Stanford Cancer Institute, Stanford University School of Medicine, Stanford, CA 94304, USA; 115 Division of Genetics and Epidemiology, The Institute of Cancer Research, London SM2 5NG, UK; 116 Department of Oncology, Tampere University Hospital, Tampere, Finland Box 2000, 33521 Tampere, Finland; 117 Central Finland Health Care District, Jyva ¨skyla ¨Central Hospital, Jyva ¨skyla ¨40620, Finland; 118 Surgery, School of Medicine, National University of Ireland, Galway H91TK33, Ireland; 119 Department of Obstetrics and Gynecology, Helsinki University Hospital, University of Helsinki, Helsinki 00290, Finland; 120 Prosserman Centre for Population Health Research, Lunenfeld-Tanenbaum Research Institute, Sinai Health System, Toronto, ON M5T 3L9, Canada; 121 Division of Epidemiology, Dalla Lana School of Public Health, University of Toronto, Toronto, ON M5T 3M7, Canada; 122 Department of Internal Medicine, Evangelische Kliniken Bonn gGmbH, Johanniter Krankenhaus, Bonn 53177, Germany; 123 Institute of Clinical Medicine, Pathology and Forensic Medicine, University of Eastern Finland, Kuopio 70210, Finland; 124 Imaging Center, Department of Clinical Pathology, Kuopio University Hospital, Kuopio 70210, Finland; 125 Cancer Epidemiology Group, University Cancer Center Hamburg (UCCH), University Medical Center Hamburg-Eppendorf, Hamburg 20246, Germany; 126 VIB Center for Cancer Biology, VIB, Leuven 3000, Belgium; 127 Laboratory for Translational Genetics, Department of Human Genetics, University of Leuven, Leuven 3000, Belgium; 128 Epidemiology Program, University of Hawaii Cancer Center, Honolulu, HI 96813, USA; 129 Clalit National Cancer Control Center, Carmel Medical Center and Technion Faculty of Medicine, Haifa 35254, Israel; 130 Department of Molecular Medicine and Surgery, Karolinska Institutet, and Department of Clinical Genetics, Karolinska University Hospital, Stockholm 171 76, Sweden; 131 Department of Epidemiology, University of Washington School of Public Health, Seattle, WA 98195, USA; 132 Public Health Sciences Division, Fred Hutchinson Cancer Research Center, Seattle, WA 98109, USA; 133 Department of Cancer Epidemiology and Prevention, M Sklodowska-Curie Cancer Center - Oncology Institute, Warsaw 02-034, Poland; 134 German Breast Group, GmbH, Neu Isenburg 63263, Germany; 135 Division of Epidemiology, Department (Affiliations continued on next page) (Author list continued on next page) The American Journal of Human Genetics 104, 21–34, January 3, 2019 23 Here, we used data from 79 studies conducted by the Breast Cancer Association Consortium (BCAC) to optimize PRSs for overall and subtype-specific disease, and we validate their performance in independent datasets. 1,13–15 Material and Methods Study Subjects and Genotyping The dataset used for development of the PRSs comprised 94,075 breast cancer-affected case subjects and 75,017 control subjects of European ancestry from 69 studies in the BCAC (Tables S1 and S2). Data collection for individual studies is described previously. 1 Samples were genotyped using one of two arrays: iCOGS 13,14 and OncoArray. 1,15 The dataset was divided into a training and validation set. The validation set was randomly selected (approximately 10% of case and control subjects) from studies that had been genotyped with the OncoArray, after excluding studies of bilateral breast cancer, studies or sub-studies oversampling for family history, and individuals with in situ cancers or case subjects with unknown ER status. The best PRSs were evaluated in an independent test dataset comprising 11,428 invasive breast cancer-affected case subjects of Medicine, Vanderbilt Epidemiology Center, Vanderbilt-Ingram Cancer Center, Vanderbilt University School of Medicine, Nashville, TN 37232, USA; 136 Unit of Medical Genetics, Department of Medical Oncology and Hematology, Fondazione IRCCS Istituto Nazionale dei Tumori (INT), Milan 20133, Italy; 137 Department of Clinical Science and Education, So ¨dersjukhuset, Karolinska Institutet, Stockholm 118 83, Sweden; 138 Department of Oncology, So ¨dersjukhuset, Stockholm 118 83, Sweden; 139 Department of Family Medicine and Public Health, University of California San Diego, La Jolla, CA 92093, USA; 140 Anatomical Pathology, The Alfred Hospital, Melbourne, VIC 3004, Australia; 141 Department of Gynecology and Obstetrics, Ludwig Maximilian University of Munich, Munich 80336, Germany; 142 MRC Clinical Trials Unit at UCL, Institute of Clinical Trials & Methodology, University College London, London WC1V 6LJ, UK; 143 Faculty of Medicine, University of Heidelberg, Heidelberg 69120, Germany; 144 Department of Laboratory Medicine and Pathobiology, University of Toronto, Toronto, ON M5S 1A8, Canada; 145 Laboratory Medicine Program, University Health Network, Toronto, ON M5G 2C4, Canada; 146 Universite ´Paris Sorbonne Cite ´, INSERM UMR-S1147, Paris 75270, France; 147 Radiation Oncology, Hospital Meixoeiro-XXI de Vigo, Vigo 36214, Spain; 148 Leuven Multidisciplinary Breast Center, Department of Oncology, Leuven Cancer Institute, University Hospitals Leuven, Leuven 3000, Belgium; 149 Clinical Genetics Research Lab, Department of Cancer Biology and Genetics, Memorial Sloan Kettering Cancer Center, New York, NY 10065, USA; 150 Clinical Genetics Service, Department of Medicine, Memorial Sloan Kettering Cancer Center, New York, NY 10065, USA; 151 Centre for Cancer Research and Cell Biology, Queen’s University Belfast, Belfast BT7 1NN, UK; 152 University of New Mexico Health Sciences Center, University of New Mexico, Albuquerque, NM 87131, USA; 153 Servicio de Cirugı ´a General y Especialidades, Hospital Monte Naranco, Oviedo 33012, Spain; 154 Medical Oncology Department, Hospital Universitario Puerta de Hierro, Madrid 28222, Spain; 155 Genome Diagnostic Program, IFOM the FIRC (Italian Foundation for Cancer Research) Institute of Molecular Oncology, Milan 20139, Italy; 156 Department of Oncology, Wayne State University School of Medicine, Detroit, MI 48201, USA; 157 Laboratory of Cancer Genetics and Tumor Biology, Cancer and Translational Medicine Research Unit, Biocenter Oulu, University of Oulu, Oulu 90220, Finland; 158 Laboratory of Cancer Genetics and Tumor Biology, Northern Finland Laboratory Centre Oulu, Oulu 90220, Finland; 159 Unit of Molecular Bases of Genetic Risk and Genetic Testing, Department of Research, Fondazione IRCCS Istituto Nazionale dei Tumori (INT), Milan 20133, Italy; 160 Department of Oncology, Mayo Clinic, Rochester, MN 55905, USA; 161 Department of Oncology, University Hospital of Larissa, Larissa 711 10, Greece; 162 Epidemiology Branch, National Institute of Environmental Health Sciences, NIH, Research Triangle Park, NC 27709, USA; 163 Research Oncology, Guy’s Hospital, King’s College London, London SE1 9RT, UK; 164 Faculty of Information Technology, Monash University, Melbourne, VIC 3800, Australia; 165 National Center for Tumor Diseases, University Hospital and German Cancer Research Center, Heidelberg 69120, Germany; 166 Department of Population and Quantitative Health Sciences, Case Western Reserve University, Cleveland, OH 44106, USA; 167 Division of Molecular Medicine, Pathology North, John Hunter Hospital, Newcastle, NSW 2305, Australia; 168 Discipline of Medical Genetics, School of Biomedical Sciences and Pharmacy, Faculty of Health, University of Newcastle, Callaghan, NSW 2308, Australia; 169 Hunter Medical Research Institute, John Hunter Hospital, Newcastle, NSW 2305, Australia; 170 Department of Health Sciences Research, Mayo Clinic College of Medicine, Jacksonville, FL 32224, USA; 171 Precision Medicine, School of Clinical Sciences at Monash Health, Monash University, Clayton, VIC 3168, Australia; 172 Department of Clinical Pathology, The University of Melbourne, Melbourne, VIC 3010, Australia; 173 Population Oncology, BC Cancer, Vancouver, BC V5Z 1G1, Canada; 174 School of Population and Public Health, University of British Columbia, Vancouver, BC, V6T 1Z4, Canada; 175 Saarland Cancer Registry, Saarbru ¨cken 66119, Germany; 176 The Curtin UWA Centre for Genetic Origins of Health and Disease, Curtin University and University of Western Australia, Perth, WA 6000, Australia; 177 Division of Breast Cancer Research, The Institute of Cancer Research, London SW7 3RP, UK; 178 Faculty of Medicine, University of Southampton, Southampton SO17 1BJ, UK; 179 Epigenetic and Stem Cell Biology Laboratory, National Institute of Environmental Health Sciences, NIH, Research Triangle Park, NC 27709, USA; 180 Department of Epidemiology, Mailman School of Public Health, Columbia University, New York, NY 10032, USA; 181 Department of Surgery, Leiden University Medical Center, Leiden 2333 ZA, the Netherlands; 182 Institute of Cancer and Genomic Sciences, University of Birmingham, Birmingham B15 2TT, UK; 183 Wellcome Trust Centre for Human Genetics and Oxford NIHR Biomedical Research Centre, University of Oxford, Oxford OX3 7BN, UK; 184 Department of Pathology, University Hospital of Heraklion, Heraklion 711 10, Greece; 185 Frauenklinik der Stadtklinik Baden-Baden, Baden-Baden 76532, Germany; 186 Department of Gynecology and Obstetrics, Helios Clinics Berlin-Buch, Berlin 13125, Germany; 187 Biostatistics and Computational Biology Branch, National Institute of Environmental Health Sciences, NIH, Research Triangle Park, NC 27709, USA; 188 Department of Health Research and Policy - Epidemiology, Stanford Victor M. Mun ˜oz-Garzon, 147 Susan L. Neuhausen, 24 Heli Nevanlinna, 119 Patrick Neven, 148 William G. Newman, 72,73 Sune F. Nielsen, 29,30 Børge G. Nordestgaard, 29,30,31 Aaron Norman, 105 Kenneth Offit, 149,150 Janet E. Olson, 105 Ha ˚kan Olsson, 69 Nick Orr, 151 V. Shane Pankratz, 152 Tjoung-Won Park-Simon, 28 Jose I.A. Perez, 153 Clara Pe ´rez-Barrios, 154 Paolo Peterlongo, 155 Julian Peto, 61 Mila Pinchev, 129 Dijana Plaseska-Karanfilska, 111 Eric C. Polley, 105 Ross Prentice, 14 Nadege Presneau, 65 Darya Prokofyeva, 89 Kristen Purrington, 156 Katri Pylka ¨s, 157,158 Brigitte Rack, 113 Paolo Radice, 159 Rohini Rau-Murthy, 150 Gad Rennert, 129 Hedy S. Rennert, 129 Valerie Rhenius, 3 Mark Robson, 150 Atocha Romero, 154 Kathryn J. Ruddy, 160 Matthias Ruebner, 20 Emmanouil Saloustros, 161 Dale P. Sandler, 162 Elinor J. Sawyer, 163 Daniel F. Schmidt, 87,164 Rita K. Schmutzler, 96,97 Andreas Schneeweiss, 165 Minouk J. Schoemaker, 115 Fredrick Schumacher, 166 Peter Schu ¨rmann, 28 Lukas Schwentner, 113 Christopher Scott, 105 Rodney J. Scott, 167,168,169 Caroline Seynaeve, 107 Mitul Shah, 3 Mark E. Sherman, 170 (Author list continued on next page) (Affiliations continued on next page) 24 The American Journal of Human Genetics 104, 21–34, January 3, 2019 and 18,323 control subjects from ten studies nested within prospective cohorts, all genotyped using the OncoArray (Tables S3 and S4). The overall breast cancer PRS was also evaluated among 190,040 women of European ancestry from the UK Biobank cohort who had not had any cancer diagnosis or mastectomy prior to recruitment. A total of 3,215 incident registry-confirmed invasive breast cancers developed over 1,381,019 person years of prospective follow-up. Follow-up started 6 months after age of baseline questionnaire. The primary endpoint was invasive breast cancer. Follow-up was censored at the earliest of: risk-reducing mastectomy, diagnosis of any type of cancer, death, or January 15, 2017. Genotype calling, quality control, and imputation for iCOGS and OncoArray were performed as previously described. 1,14 Briefly, imputation was performed for the iCOGS and OncoArray datasets separately using the Phase 3 (October 2014) release of the 1000 Genomes data as reference. 16 We followed a two-stage approach using SHAPEIT for phasing 17 and IMPUTE2 for the imputation. 15 Where samples were genotyped with iCOGS and OncoArray, the OncoArray calling was used. SNPs with MAF >0.01 and imputation r 2 >0.9 for OncoArray and r 2 >0.3 for iCOGS were included in this analysis (7 million SNPs); a higher threshold was imposed for OncoArray to ensure accurate determination of the PRS in the validation and test datasets. UK Biobank samples were genotyped using Affymetrix UK BiLEVE Axiom array and Affymetrix UK Biobank Axiom array and imputed to the combined 1000 Genomes Project v.3 and UK10K reference panels using SHAPEIT3 and IMPUTE3. 18 The lowest imputation info score for the SNPs used in these analyses was 0.86. Samples were included on the basis of female sex (genetic and self-reported) and ethnicity filter (Europeans/White British ancestry subset). Duplicates, individuals with high degree of relatedness (>10 relatives), and one of each related pair of first degree relatives were removed. Samples were also excluded using standard quality control criteria. Participants provided written informed consent, all studies were approved by the relevant ethics committees, and procedures followed were in accordance with the ethical standards of these committees. Statistical Analysis The general aim was to derive a PRS of the form: PRS ¼b1x1þb2x2þ.þbkxk.þbnxn where b k is the per-allele log odds ratio (OR) for breast cancer associated with SNP k,x k is the allele dosage for SNP k, and n is the total number of SNPs included in the PRS. Previous analyses found no evidence for statistically significant interactions between SNPs 19,20 and little evidence for departures from a log-additive model for individual SNPs. Assuming this is true in general, the PRS summarizes efficiently the combined effects of SNPs on disease risk. The main challenge is how to determine which SNPs to include and the weighting parameters b k to assign. Inclusion of only those SNPs reaching a stringent significance threshold (‘‘genome-wide significant,’’ p <5310 8 ) threshold ignores information from larger numbers of SNPs that are likely, but not certain, to be associated with the risk of breast cancer. We used two general Martha J. Shrubsole, 135 Xiao-Ou Shu, 135 Susan Slager, 105 Ann Smeets, 148 Christof Sohn, 165 Penny Soucy, 62 Melissa C. Southey, 171,172 John J. Spinelli, 173,174 Christa Stegmaier, 175 Jennifer Stone, 87,176 Anthony J. Swerdlow, 115,177 Rulla M. Tamimi, 67,68,109 William J. Tapper, 178 Jack A. Taylor, 162,179 Mary Beth Terry, 180 Kathrin Tho ¨ne, 125 Rob A.E.M. Tollenaar, 181 Ian Tomlinson, 182,183 The ´re `se Truong, 53 Maria Tzardi, 184 Hans-Ulrich Ulmer, 185 Michael Untch, 186 Celine M. Vachon, 105 Elke M. van Veen, 72,73 Joseph Vijai, 149,150 Clarice R. Weinberg, 187 Camilla Wendt, 137,138 Alice S. Whittemore, 188,189 Hans Wildiers, 148 Walter Willett, 68,190,191 Robert Winqvist, 157,158 Alicja Wolk, 99,192 Xiaohong R. Yang, 6 Drakoulis Yannoukakos, 193 Yan Zhang, 12 Wei Zheng, 135 Argyrios Ziogas, 10 ABCTB Investigators, 194 kConFab/AOCS Investigators, 195 NBCS Collaborators, 33,34,196,197,198,199,200,201,202,203,204,205 Alison M. Dunning, 3 Deborah J. Thompson, 1 Georgia Chenevix-Trench, 206 Jenny Chang-Claude, 19,125 Marjanka K. Schmidt, 54,207 Per Hall, 57,138 Roger L. Milne, 86,87,171 Paul D.P. Pharoah, 1,3 Antonis C. Antoniou, 1 Nilanjan Chatterjee, 6,208,209 Peter Kraft, 68,109 Montserrat Garcı ´a-Closas, 6 Jacques Simard, 62 and Douglas F. Easton 1,3 University School of Medicine, Stanford, CA 94305, USA; 189 Department of Biomedical Data Science, Stanford University School of Medicine, Stanford, CA 94305, USA; 190 Department of Nutrition, Harvard TH Chan School of Public Health, Boston, MA 02115, USA; 191 Channing Division of Network Medicine, Brigham and Women’s Hospital and Harvard Medical School, Boston, MA 02115, USA; 192 Department of Surgical Sciences, Uppsala University, Uppsala 751 05, Sweden; 193 Molecular Diagnostics Laboratory, INRASTES, National Centre for Scientific Research ‘‘Demokritos,’’ Athens 15310, Greece; 194 Australian Breast Cancer Tissue Bank, Westmead Institute for Medical Research, University of Sydney, Sydney, NSW 2145, Australia; 195 Peter MacCallum Cancer Center, Melbourne, VIC 3000, Australia; 196 Department of Research, Vestre Viken Hospital, Drammen 3019, Norway; 197 Department of Cancer Genetics, Vestre Viken Hospital, Drammen 3019, Norway; 198 Section for Breast and Endocrine Surgery, Department of Cancer, Division of Surgery, Cancer and Transplantation Medicine, Oslo University Hospital-Ulleva ˚l, Oslo 0450, Norway; 199 Department of Radiology and Nuclear Medicine, Oslo University Hospital, Oslo 0379, Norway; 200 Department of Pathology, Akershus University Hospital, Lørenskog 1478, Norway; 201 Department of Tumor Biology, Institute for Cancer Research, Oslo University Hospital, Oslo 0379, Norway; 202 Department of Oncology, Division of Surgery, Cancer and Transplantation Medicine, Oslo University Hospital-Radiumhospitalet, Oslo 0379, Norway; 203 National Advisory Unit on Late Effects after Cancer Treatment, Oslo University Hospital-Radiumhospitalet, Oslo 0379, Norway; 204 Department of Oncology, Akershus University Hospital, Lørenskog 1478, Norway; 205 Breast Cancer Research Consortium, Oslo University Hospital, Oslo 0379, Norway; 206 Department of Genetics and Computational Biology, QIMR Berghofer Medical Research Institute, Brisbane, QLD 4006, Australia; 207 Division of Psychosocial Research and Epidemiology, the Netherlands Cancer Institute - Antoni van Leeuwenhoek hospital, Amsterdam 1066 CX, the Netherlands; 208 Department of Biostatistics, Bloomberg School of Public Health, John Hopkins University, Baltimore, MD 21205, USA; 209 Department of Oncology, School of Medicine, John Hopkins University, Baltimore, MD 21205, USA *Correspondence: [email protected] https://doi.org/10.1016/j.ajhg.2018.11.002. The American Journal of Human Genetics 104, 21–34, January 3, 2019 25 approaches for model selection: ‘‘hard-thresholding,’’ based on a stepwise regression model that retained SNPs significantly associated with overall or subtype-specific disease at a given threshold, and penalized regression using lasso. 21,22 A schema for the analyses is shown in Figure S1. To prioritize SNPs for analysis, single SNP association tests were first conducted in the training set. Per-allele ORs and standard errors were estimated separately in the iCOGS and OncoArray datasets, adjusting for study and nine ancestry informative principal components (PCs) in the iCOGS dataset and by country and ten PCs in the OncoArray dataset, using a purpose-written program. 1 Combined p values were then derived using a fixed-effects meta-analysis with the software METAL. 23 SNPs were sorted by p value and filtered on LD, such that uncorrelated SNPs (correlation r 2 <0.9) with lowest p value for association with overall breast cancer in the training set were retained (more rigorous pruning, for example at r 2 <0.2, would have removed from consideration informative SNPs from regions with multiple correlated signals 24,25 ). In the hard thresholding approach, a series of stepwise forward regression analyses were first carried out in 1 Mb regions centered on SNPs significant at a pre-specified threshold for association with either overall and/or subtype-specific disease in the training set. Only SNPs passing the specified p value thresholds were included in each 1 Mb region. Two analyses were performed in parallel: for overall breast cancer and ER-negative disease. At each stage the SNP with the smallest (conditional) p value for any analysis was added to the model, the threshold for the stepwise regression being the same as that for pre-selection. The process was repeated until no further SNPs could be added at the pre-defined threshold. A second stage of stepwise regressions were then carried out across all regions in each chromosome, to take into account correlated SNPs in different regions. Finally, the effect sizes for the selected SNPs were jointly estimated in a single logistic regression model. For the best-performing PRSs, SNPs associated with ER-positive at p <10 6 but not with overall breast cancer (at p <10 5 ) were added at the end of the final SNP list. A third round of stepwise forward regression was then carried out with p value for selection of p <10 6 for ER-positive disease. For completeness we added to this final PRS two rarer variants (BRCA2 p.Lys3326X and CHEK2 p.Ile157Tyr) which are established to confer a moderate risk of breast cancer and were genotyped on the OncoArray but did not pass the allele frequency threshold in the PRS development phase. For the penalized regression using lasso, we used the program glmnet 21 .SNPs with p <0.001 in overall BC or ER-negative disease in the training set were pre-selected for inclusion in the lasso, and BRCA2 p.Lys3326X and CHEK2 p.Ile157Thr were added. Covariates for 19 PCs (9 for iCOGs and 10 for Oncoarray) and country were included in each model. For overall breast cancer, the penalty parameter (lambda) giving the best overall breast cancer PRS in the validation set was selected. To construct subtype-specific PRSs, we evaluated four different methods: (1) using effect sizes for overall breast cancer (for each of the subtypes), (2) using effect sizes for subtype-specific (ER-positive or ER-negative) disease, (3) using a hybrid method, in which effect sizes were estimated in the relevant subtype for SNPs passing a certain optimal significance threshold in a case-only logistic regression (ER-positive versus ER-negative disease), and otherwise, using effect sizes estimated for overall breast cancer, or (4) by estimating case-only ORs using lasso and combining these with the overall breast cancer ORs to derive subtype-specific estimates, using the formulae: bERpositive ¼boverall þhbcaseonly bERnegative ¼boverall ð1hÞbcaseonly where h¼0.27 was the proportion of ER-negative tumors in the validation set. For the lasso analysis, effect sizes for subtype-specific disease were estimated using method 4 above, combining the estimates from a case-only lasso analysis with the coefficients for overall breast cancer from the lasso analysis. The lambda for the case-only model giving the best subtype-specific PRS in the validation set was selected. To evaluate the performance of each potential PRS, we standardized the PRSs to have unit standard deviation (SD) in the validation set of control subjects. The association of the standardized PRSs was evaluated in the validation and test (prospective studies) datasets, by logistic regression. We used a Cox proportional hazards regression model to assess the association with risk of breast cancer in UK Biobank. Models were also compared in terms of the area under the receiver operator characteristic curves (AUC), adjusted for study, calculated using the Stata command comproc. Meta-analysis of study-specific effects was carried out using the Stata command metan. The goodness of fit of the continuous model (i.e., assuming a linear association between log(OR) and risk) was tested using the Hosmer-Lemeshow (HL) test to compare the observed and predicted risks by quantile and using the tail-based test proposed by Song et al. 26 In addition, we considered specifically the risks in the highest and lowest 1% of the distribution. Effect modification of the PRS by age and family history of breast cancer in first-degree relatives was evaluated by fitting additional interaction terms in the model. The validation and prospective test datasets were combined for this analysis. The absolute risks of developing breast cancer (overall and subtype-specific disease) were calculated taking into account the competing risk of dying from causes other than breast cancer, as described previously, 7 with the PRS modeled as a continuous covariate and including a linear ‘‘age 3PRS’’ interaction term. The absolute risk of developing subtype-specific disease was obtained constraining to the incidence of overall incidence of ER-negative and ER-positive disease in the UK. Women are at risk of developing both ER-negative and ER-positive disease, so the absolute risks were calculated given that the individual has been free of breast cancer of any subtype. Analyses were carried out in R v.3.0.2 and Stata v.14.2. All tests of statistical significance were two-sided. Further details are provided in the Supplemental Material and Methods. Results Development of the PRS We tried several approaches to develop PRSs; here we report results for models giving the highest prediction accuracy. Using stepwise forward selection, the best PRS for prediction of overall breast cancer was obtained at a p value threshold for pre-selection and stepwise regression of p <10 5 (Table 1). The OR per unit standard deviation (SD) for this 305-SNP PRS with overall breast cancer in 26 The American Journal of Human Genetics 104, 21–34, January 3, 2019 the validation set was 1.65 (95%CI: 1.58–1.72), compared with 1.59 (95%CI: 1.52–1.66) using a ‘‘genome-wide’’ (p <5310 8 ) threshold (123 SNPs). Using lasso regression, the best PRS (OR ¼1.71, 95%CI: 1.64–1.79) was more predictive than the best PRS developed using the stepwise regression model. In the best model (l¼0.003), 3,820 SNPs were selected (Table 1). Optimizing the PRS for Prediction of Subtype-Specific Disease For evaluation of subtype-specific models following stepwise regression, SNP effect sizes were estimated, in the first instance, in each disease subtype. The best subtype-specific PRSs using this method were also obtained at a p value threshold of p <10 5 (Table S5). The 305-SNP PRS was supplemented with 6 additional SNPs associated with ER-positive at p value <10 6 and, in addition, by two known rare breast cancer susceptibility variants in the BRCA2 and CHEK2 genes, bringing the total number of SNPs included to 313 (PRS 313 ). The optimum subtype-specific PRS was obtained when a subset of these 313 SNPs (196 SNPs with a case-only p value for association with ER-negative versus ER-positive disease of p <0.025) were given subtype-specific weights, while the remaining SNPs were given overall breast cancer weights. For ER-negative disease, the OR improved from OR ¼1.45 (95%CI: 1.35–1.56) to OR ¼1.47 (95%CI: 1.37–1.58) using the hybrid method compared with using only subtype-specific estimates, while for ER-positive disease the results were similar (OR ¼1.74) (Tables S6 and S7). Subtype-specific prediction using the lasso analysis was optimized using case-only lasso analysis. The OR per 1 SD in the validation set was 1.81 (95%CI: 1.73– 1.89) for ER-positive and 1.48 (95%CI: 1.37–1.59) for ERnegative disease (Tables 2 and S8). Validation of the PRS in the Prospective Test Dataset The final PRSs were evaluated using data from 11,428 invasive breast cancer-affected case subjects and 18,323 control subjects from ten prospective studies. The ORs for both the overall and subtype-specific PRSs were slightly lower in the prospective test set compared to the validation set (Table 2). The difference between validation and test set may reflect some overfitting due to choosing the optimum p value threshold and for the lasso, the optimum lambda, in the validation set, but could also be due to somewhat different characteristics of the prospective studies. The ORs for overall and ER-positive, but not ER-negative, breast cancer were slightly higher for the 3,820-SNP PRS (PRS 3820 ) compared with PRS 313 . The odds ratio (OR) for overall disease per 1 standard deviation (SD) of the PRS 313 in the prospective studies was 1.61 (95%CI: 1.57–1.65) while for the 77-SNP PRS (PRS 77 ) derived previously OR ¼1.46 (95%CI: 1.42–1.49). For ERnegative disease the difference was OR ¼1.45 (95%CI: 1.37–1.53) versus 1.35 (95%CI: 1.27–1.43) (Table 2). The associations between the PRS and overall, ERpositive, and ER-negative breast cancer by percentiles of the PRS 313 are shown in Figure 1 and Table S9. Compared with women in the middle quintile (40th to 60th percentile), those in the highest 1% of risk for the subtype-specific PRS 313 had 4.37 (95%CI: 3.59–5.33)- and 2.78 (95%CI: 1.83–4.24)-fold risks, and those in the lowest 1% had 0.16 (95%CI: 0.09–0.30)- and 0.27 (95%CI: 0.09–0.86)- fold risks of developing ER-positive and ER-negative disease, respectively. The ORs by percentile of the PRS 3820 were similar (Table S10). Goodness of Fit of the PRS The remaining analyses concentrated on PRS 313 .The associations between the PRS and breast cancer risk by Table 1. Comparison of Methods for Deriving the PRS: Results for Overall Breast Cancer in the Validation Set p Value Cutoff a SNPs Entering Model (n) SNPs Selected (n) OR b 95% CI AUC Published PRS 7 77 77 1.49 1.44–1.56 0.612 Hard-Thresholding Stepwise Forward Regression <5310 8 1,817 123 1.59 1.52–1.66 0.626 <10 6 2,603 197 1.62 1.55–1.68 0.634 <10 5 3,818 305 1.65 1.58–1.72 0.637 <10 4 6,743 669 1.62 1.56–1.69 0.631 <10 3 14,760 1,707 1.55 1.49–1.62 0.623 Penalized Regression Lasso 15,032 3,820 1.71 1.64–1.79 0.647 a The p value cut off refers to the SNPs considered based on their marginal associations in the training set; the same p value threshold was used in each case in the stepwise regression. Parameter selection and effect size estimation for derivation of the PRS was carried out in the training set as described in the Material and Methods. b OR per 1 SD for the PRS. OR for association with breast cancer in the validation set was derived using logistic regression adjusting for country and ten PCs. AUCs were adjusted for country. The lasso was carried out after pre-selecting SNPs at p <10 3 based on their marginal association in the training set. For the lasso l¼0.003 gave the optimal PRS in the validation set. The American Journal of Human Genetics 104, 21–34, January 3, 2019 27 percentiles of the risk score were compared with those predicted under a simple polygenic model with the PRS considered as a continuous covariate. The effect sizes did not differ from those predicted, and in particular the estimates for the highest and lowest centile were consistent with the predicted estimates (Table S9). Further tests for goodness of fit and tail-based tests (see Material and Methods) were not statistically significant at p <0.05. There was no evidence of heterogeneity in the effect sizes among studies (Figure 2). All studies showed a significant association with similar effect sizes for overall and ER-positive breast cancer, and all but one study (FHRISK, based on only six case subjects) showed a significant effect for ER-negative breast cancer. In the UK Biobank, the estimated hazard ratio (HR) for overall breast cancer per unit PRS (including 306 of the 313 SNPs) was HR ¼1.59 (95%CI: 1.54–1.64) (Figure 2). By way of comparison, we also evaluated a PRS based on 177 previously published susceptibility loci. 1,2 The effect size for this PRS (OR ¼1.61, 95%CI: 1.57–1.65) in the ten prospective studies was similar to the PRS 313 . However, this estimated effect size is biased because the validation and test datasets used here contributed to the GWAS discovery datasets; in the UK Biobank this PRS (based on 174 of 177 available SNPs) performed worse (HR ¼1.53, 95%CI: 1.48–1.58). PRS Effects by Age A weak decline in the OR with age was observed for ER-positive disease (p ¼0.001, for the combined validation and test set). There was some evidence that the decline in PRS OR was not linear, driven by a lower estimate below age 40 years (Table S11,Figure S2). There was no evidence of a decline in the OR by age for ER-negative disease (p ¼0.39). Combined Effects of PRS and Breast Cancer Family History The association between PRS and disease risk was observed for women with and without a family history (Table 3). However, there was some evidence that for ER-positive disease, the PRS OR was smaller in women with a family history (interaction OR ¼0.91, p ¼0.004). The log OR for family history was attenuated by 21% (1.59 to 1.44) and 12% (1.66 to 1.56) for ER-positive and ER-negative disease, respectively, after adjusting for the PRS (Tables 3 and S12). Absolute Risk of Developing Breast Cancer According to the PRS Estimated lifetime and 10-year absolute risks for UK women in percentiles of the PRS are shown in Figure 3. For ER-positive disease, the estimated lifetime absolute risk by age 80 years ranged from 2% for women in the lowest centile to 31% in the highest centile, while for ER-negative disease, the absolute risks ranged from 0.55% to 4%. The average 10-year absolute risk of breast cancer for a 47-year-old woman (i.e., the age at which women become eligible to enter the UK breast cancer screening program) in the general population is 2.6%. However, the 19% of women with the highest PRSs will attain this level of risk by age 40 years. Table 2. Association between PRS and Breast Cancer Risk in the Validation Set and Prospective Test Datasets Validation Set Prospective Test Set OR a 95% CI AUC OR a 95% CI AUC 77 SNP PRS (PRS 77 ) Overall BC 1.49 1.44–1.56 0.612 1.46 1.42–1.49 0.603 ER-positive 1.56 1.49–1.63 0.623 1.52 1.48–1.56 0.615 ER-negative 1.40 1.30–1.50 0.596 1.35 1.27–1.43 0.584 313 SNP PRS (PRS 313 ) Overall BC 1.65 1.59–1.72 0.639 1.61 1.57–1.65 0.630 ER-positive 1.74 1.66–1.82 0.651 1.68 1.63–1.73 0.641 ER-negative 1.47 1.37–1.58 0.611 1.45 1.37–1.53 0.601 3,820 SNP PRS (PRS 3820 ) Overall BC 1.71 1.64–1.79 0.646 1.66 1.61–1.70 0.636 ER-positive 1.81 1.73–1.89 0.659 1.73 1.68–1.78 0.647 ER-negative 1.48 1.37–1.59 0.611 1.44 1.36–1.53 0.600 Parameter selection and effect size estimation for derivation of the PRS was carried out in the training set as described in the Material and Methods. The optimal subtype-specific PRS was obtained by carrying out case-only logistic regression and estimating effect sizes in the relevant subtype for SNPs passing a p value of 0.025 in case-only ordinary logistic regression (ER-positive versus ER-negative disease). OR for association with breast cancer in the validation set derived using logistic regression adjusting for country and ten PCs. AUCs were adjusted for by country. In the prospective test set, logistic regression models were adjusted for study and 15 PCs. AUCs were adjusted for by study. a OR per 1 SD for the PRS. 28 The American Journal of Human Genetics 104, 21–34, January 3, 2019 Discussion We report development and independent validation of polygenic risk scores for breast cancer, optimized for prediction of subtype-specific disease and based on the largest available GWAS dataset. The best PRS based on a hard thresholding approach included 313 SNPs and was significantly more predictive of risk than the previously reported 77-SNP PRS 7 (OR per 1 SD in the prospective test set: 1.61 versus 1.46; Table 2). The effect sizes were remarkably Figure 2. Prospective Validation for the 313 SNP Polygenic Risk Score Prospective validation for the 313 SNP polygenic risk score (PRS) by study for (A) overall breast cancer, (B) ER-positive disease, and (C) ER-negative disease. Association between the 313 SNP PRS and breast cancer risk in women of European origin. Odds ratios and 95% confidence intervals are shown. I-squared and p value for heterogeneity were calculated using fixed effect meta-analysis. Figure 1. Association between the 313 SNP Polygenic Risk Score and Breast Cancer Risk Association between the 313 SNP polygenic risk score (PRS) and breast cancer risk in women of European origin for (A) overall breast cancers, (B) estrogen receptor (ER)-positive disease, and (C) ER-negative disease, in the validation (dashed line) and test (solid line) sets. Odds ratios are for different quantiles of the PRS relative to the mean PRS. Odds ratios and 95% confidence intervals are shown. The American Journal of Human Genetics 104, 21–34, January 3, 2019 29