scieee Open visual document viewer

Atlas of prostate cancer heritability in European and African-American men pinpoints tissue-specific regulation

Gusev, Alexander,Shi, Huwenbo,Kichaev, Gleb,Schleutker, Johanna,Auvinen, Anssi,Murtola, Teemu,Tammela, Teuvo,Wahlfors, Tiina

Abstract

Although genome-wide association studies have identified over 100 risk loci that explain ~33% of familial risk for prostate cancer (PrCa), their functional effects on risk remain largely unknown. Here we use genotype data from 59,089 men of European and African American ancestries combined with cell-type-specific epigenetic data to build a genomic atlas of single-nucleotide polymorphism (SNP) heritability in PrCa. We find significant differences in heritability between variants in prostate-relevant epigenetic marks defined in normal versus tumour tissue as well as between tissue and cell lines. The majority of SNP heritability lies in regions marked by H3k27 acetylation in prostate adenoc7arcinoma cell line (LNCaP) or by DNaseI hypersensitive sites in cancer cell lines. We find a high degree of similarity between European and African American ancestries suggesting a similar genetic architecture from common variation underlying PrCa risk. Our findings showcase the power of integrating functional annotation with genetic data to understand the genetic basis of PrCa.

Full text

ARTICLE Recei ed 26 Dec 2014 |Accep ed 3 Feb 2016 |Published 7 Ap 2016 A las o p os a e cance he i abili y in Eu opean and A ican-Ame ican men pinpoin s issue-specific egula ion Alexande Guse e al.# Al hough genome-wide associa ion s udies ha e iden ified o e 100 isk loci ha explain B33% o amilial isk o p os a e cance (P Ca), hei unc ional e ec s on isk emain la gely unknown. He e we use geno ype da a om 59,089 men o Eu opean and A ican Ame ican ances ies combined wi h cell- ype-specific epigene ic da a o build a genomic a las o single-nucleo ide polymo phism (SNP) he i abili y in P Ca. We find significan di e ences in he i abili y be ween a ian s in p os a e- ele an epigene ic ma ks defined in no mal e sus umou issue as well as be ween issue and cell lines. The majo i y o SNP he i abili y lies in egions ma ked by H3k27 ace yla ion in p os a e adenoc7a cinoma cell line (LNCaP) o by DNaseI hype sensi i e si es in cance cell lines. We find a high deg ee o simila i y be ween Eu opean and A ican Ame ican ances ies sugges ing a simila gene ic a chi ec u e om common a ia ion unde lying P Ca isk. Ou findings showcase he powe o in eg a ing unc ional anno a ion wi h gene ic da a o unde s and he gene ic basis o P Ca. Co espondence and eques s o ma e ials should be add essed o A.G. (email: [email p o ec ed].edu) o o B.P (email: [email p o ec ed]). #A ull lis o au ho s and hei a filia ions appea s a he end o he pape . DOI: 10.1038/ncomms10979 OPEN NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions 1 Family his o y is a well-es ablished isk ac o o p os a e cance (P Ca), which has an es ima ed he i abili y o 58%—one o he highes ac oss common cance s1. Genome-wide associa ion s udies (GWAS) ha e been pa icula ly success ul in iden i ying o e 100 isk loci ha cap u e B33% o he es ima ed amilial isk2. Al hough mos o he GWAS P Ca a ian s o e lap p os a e-specific egula o y elemen s ( o example, and ogen ecep o -binding si es (ARBS))2–8, a quan ifica ion o he con ibu ion o gene ic a ia ion om a ious ch oma in ma ks o P Ca isk is cu en ly lacking. Recen wo k o m he ENCODE/ROADMAP conso ia9has shown ha a la ge ac ion o he genome plays a ole in a leas one biochemical e en , in a leas one issue. Al hough his unc ional a las o he human genome has g ea ly enhanced ou unde s anding o egula o y elemen s, such unc ional elemen s a e o en issue specific10,11 making hei in e p e abili y in he con ex o P Ca isk challenging. Exis ing s udies ha ha e in eg a ed P Ca GWAS findings wi h issue-specific unc ional anno a ions ha e elied only on he GWAS significan a ian s (B100 in he mos ecen s udy) o single-nucleo ide polymo phisms (SNPs) agging hem2,7, hus igno ing loci ha do no each genome-wide significance. Recen me hodological ad ances ha e shown ha he en i e polygenic a chi ec u e o common ai s can be in e oga ed using a iance componen s ac oss all assayed SNPs ( yped and/o impu ed) o inc ease powe o de ec ing ai -specific unc ional anno a ions12. In addi ion o o e ing supe io pe o mance ela i e o me hods ha e alua e only GWAS SNPs, he a iance componen s me hods also allow o compa ison o es ima es ac oss di e en s udies and sample sizes. This is because a iance componen s yield an unbiased es ima e (unde s anda d assump ions) o SNP he i abili y ðh2 gÞ— he a iance in ai explained by SNPs ha eside wi hin elemen s o a gi en unc ional ca ego y12–15. He e, we use a ge ed and genome-wide SNP a ay da a om 59,089 male P Ca cases and con ols o Eu opean (BPC3 ( e . 16) and iCOGS ( e . 4), espec i ely, see Me hods) and A ican Ame ican (AAPC ( e . 17), see Me hods) ances y o dissec he gene ic isk o P Ca. We es ima e he SNP he i abili y o p e iously implica ed egula o y anno a ions7,18 and pe o m a b oad analysis o 544 epigene ic ma ks om ENCODE/ROADMAP ( e . 9). Ou app oach in e oga es he en i e common polygenic a chi ec u e o P Ca while accoun ing o po en ial co ela ions be ween ela ed unc ional ca ego ies. Fi s , we find ha SNPs nea ARBS assayed in p os a e umou explain significan ly mo e o he he i abili y o P Ca han ARBS SNPs assayed in p os a e no mal issue. Second, we localize mos o he he i abili y o P Ca o egions in he genome ma ked by h ee unc ional ca ego ies: (i) H3K27ac his one modifica ions in p os a e adenoca cinoma cell lines (LNCaP; ypically ma king ac i e enhance s19); (ii) and ogen ecep o s in p os a e issue18; and (iii) DNase I hype sensi i i y si es (DHS) in cance cell lines. We eplica e he LNCaP H3K27ac and DHS esul s ac oss di e en ances ies and show ha isk p edic ion om genome-wide SNP da a is significan ly imp o ed wi h a p edic o ha inco po a es he unc ional a las as p io . O e all, ou esul s sugges a simila gene ic a chi ec u e om common a ia ion o P Ca isk ac oss men o Eu opean and A ican ances y and highligh H3k27ac his one ma k in LNCaP and ARBS in p os a e issue o ollow-up s udies o P Ca isk. Resul s Pa i ioning he gene ic isk o p os a e cance . We analysed mul iple unc ional anno a ions and quan ified he ac ion o a iance in ai explained by SNPs ha a e localized wi hin each unc ional class. Ou app oach models he pheno ype (P Ca) o a se o indi iduals as being d awn om a mul i a ia e no mal dis ibu ion wi h a iance componen s es ima ed based on gene ic da a ( ha is, SNPs) plus an en i onmen al e m (see Me hods)13,14. Fo each unc ional ca ego y i, a gene ic ela ionship ma ix ac oss all indi iduals is compu ed om all he SNPs esiding in he gi en unc ional ca ego y o se e as a a iance componen . Mul iple componen s a e hen join ly fi ed using he es ic ed maximum likelihood (REML) as implemen ed in he GCTA so wa e14 o es ima e a iance pa ame e s s2 i  o each componen . The SNP he i abili y o componen iis hen es ima ed as h2 g;i¼s2 i=Pjs2 j, whe e he sum in he denomina o is ac oss all fi ed componen s including he en i onmen al e m. The e o e, we iew h2 g;ias an es ima e o he a iance in ai ha can be explained by all he SNPs in he co esponding unc ional ca ego y wi h a linea model o he ai ( ha is, SNP he i abili y)12. We expec unc ional ca ego ies ha a e en iched wi h casual a ian s o P Ca o a ain a highe es ima ed SNP he i abili y as compa ed wi h unc ional ca ego ies deple ed o causal a ian s o P Ca. To ocus ou esul s on noncoding a ia ion and accoun o po en ial con ounde s because o linkage disequilib ium (LD), we explici ly included coding and coding-p oximal egula o y a ia ion as ‘backg ound’ compo- nen s whene e we quan ified he e ec o each unc ional anno a ion es ed (see Me hods). The a iance componen model has p e iously been shown o yield obus es ima es unde he assump ion ha causal a ian s a e yped and uni o mly sampled om a gi en componen 13,20,21. He e, we pe o m addi ional simula ions using he UK10K whole-genome sequence da a o confi m he alidi y o his model o ou da a, and o assess how ep esen a i e SNP es ima es a e o ue unde lying biology a common sequenced a ian s. The simula ion amewo k uses eal geno ype da a om he UK10K conso ium o gene a e addi i e, polygenic pheno ypes wi h a gi en he i abili y and hen pe o ms he i abili y es ima ion wi h he a iance componen model (see Me hods). Al hough he UK10K da a con ains a much smalle se o indi iduals as he iCOGS da a (3,047 e sus 42,613 indi iduals, see Me hods), i con ains a ia ion om whole- genome sequencing; his allows us o e alua e model pe o mance by simula ion when es ic ing o SNPs geno yped on he iCOGS pla o m. We ocused on he LNCaP: H3k27ac anno a ion (which was mos significan in ou da a, see below) o e alua e he mul iple componen models. O e housands o simula ions, we confi med ha he a iance componen s app oach co ec ly eco e ed he causal con ibu ion o ai om a gi en unc ional ca ego y when causal a ian s we e yped (Supplemen a y Table 1, see Me hods). Unde bo h null and en iched scena ios he es ima es we e unbiased and s anda d e o s p ope ly calib a ed (Supplemen a y Table 1). Fo common sequenced a ian s no p esen on he iCOGS pla o m, ela i e es ima es o noncoding en ichmen /deple ion we e conse a i e, wi h he agged e ec s dis ibu ed ac oss he yped componen s (Supplemen a y Table 2). De ia ions om he s anda d a iance componen s model assump ions on he dis ibu ion o e ec -sizes and ances y-specific e ec s in A ican Ame icans yielded ei he well calib a ed o conse a i e es ima es o SNP he i abili y in he ocal LNCaP: H3k27ac ca ego y (see Me hods, Supplemen a y Tables 1–3). Ou p ima y unc ional analyses ocus on he densely geno yped iCOGS sample (21,678 cases and 20,935 con ols), whose la ge sample size allowed o highly accu a e es ima es o componen -specific h2 g. Al hough he iCOGS chip is cus om buil o o e sample isk loci, i p o ides a b oad co e age o he common a ia ion genome wide4. To showcase he powe o he a iance componen s app oach, we es ima ed he o al SNP ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979 2NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions he i abili y o P Ca a 0.28 (s.e. 0.01) in he iCOGS da a (no significan ly di e en om he o al SNP he i abili y es ima e o 0.26 (s.e. 0.05) in he BPC3 da a), a significan inc ease om he a iance explained only by he known GWAS a ian s h2 GWAS  o 0.06 (s.e.m. 0.001) (see Me hods; Supplemen a y Table 4). In e es ingly, he o al SNP he i abili y in he A ican Ame ican sample, which was geno yped on a di e en pla o m han iCOGS (see Me hods), was es ima ed a 0.32 (s.e. 0.06) indica ing a simila agg ega e con ibu ion o common a ia ion o P Ca isk ac oss he wo e hnici ies despi e highe o e all isk in A ican Ame icans22 (Supplemen a y Table 4). En ichmen a and ogen ecep o -binding si es in umou s.We fi s ocused on SNPs localized in he ARBS: an epigene ic p ofile causally implica ed in p os a e umo igensis. In con as o ypical assays ha ocus on cell lines, he ARBS we e defined by ch oma in immunop ecipi a ion and high- h oughpu sequencing (ChIP-seq) di ec ly in p ima y human issue (se en no mal and 13 umou specimens)18. We obse ed ha a ian s wi hin 5 kb o umou - specific ARBS explained 17.0% o he genome-wide h2 g(s.e. 1.7%; P¼2.6 1016 by Z- es ), whe eas he a ian s nea no mal-specific ARBS explained 0.0% o he h2 g(s.e. 0.9%; P¼0.11 by Z- es ) (Fig. 1). The di e ence be ween hese wo g oups was highly significan and demons a es he impo ance o assaying unc ional ma ks in bo h no mal and umou issues. We no e ha he 5 kb ex ension may also include o he egula o y a ian s nea he umou /no mal-specific ARBS (bu no he i abili y om coding/un ansla ed egion (UTR)/p omo e a ian s, which we e explici ly modelled, see Me hods). Smalle flanking egions we e also in es iga ed bu did no include enough ma ke s o he a iance componen s model o con e ge. We also quan ified he p opo ion o SNP he i abili y explained di ec ly by all ARBS a ian s (bo h no mal and umou wi hou 5 kb flanks) a 10.7% o h2 g; significan ly di e en om he SNP he i abili y o ARBS a ian s assayed in p os a e adenoca cinoma cance cell line (LNCaP; 3.2% o h2 g)(P¼4.4 107 o di e ence by Z- es ) (Fig. 1). This di e ence is pa ially explained by he e y low numbe o SNPs wi hin cell line ARBS making hei agg ega e con ibu ion small bu no empowe ing us o place a s ong bound on he en ichmen . O e all, hese findings highligh he inc eased complexi y o ARBS in a sample o issues as compa ed wi h he single LNCaP cell line. Iden ifica ion o unc ional ma ks ele an o P Ca isk. Nex , we looked o ma ks ha con ibu e o he he i abili y o P Ca ac oss a b oad spec um o unc ional anno a ions wi hou p io assump ions on ele ance o disease. We in es iga ed 544 epigene ic anno a ions spanning six majo classes (DHS; H3k4me1; H3k4me3; H3k9ac; H3k27ac; and compu a ionally p edic ed unc ional classes o ‘segmen a ions’23,24) a e aging 101 cell ypes pe class (see Me hods). A e accoun ing o mul iple es ing, we iden ified 82 anno a ions ha exhibi ed s a is ically significan de ia ions in SNP he i abili y om wha was expec ed based on he p opo ion o he genome co e ed by ha pa icula anno a ion (see Fig. 2 and Supplemen a y Da a). We fi s ocused on 17 unc ional ma ks measu ed in he p os a e, o which 14 we e s a is ically significan (Supplemen a y Table 5). The single mos significan en ichmen was obse ed o H3k27ac ma ks in LNCaP (P¼11032 by Z- es ), which localized 22% o he o al h2 g o he 2.9% o geno yped SNPs wi hin he anno a ion. This was ollowed by a ian s in DHS ma ks in LNCaP (P¼21018 by Z- es ; 16.7% o h2 glocalized in 3.1% o genome). The DHS anno a ions allowed us o compa e es ima es ac oss h ee majo p os a e cell lines: LNCaP; no mal p os a e epi helial (P EC); and immo alized p os a e epi helial (RWPE1) (o e lapping by 25–50% wi h ARBS, Supplemen a y Fig. 1). We obse ed he i abili y explained by LNCaP DHS o be nominally significan ly highe han P EC (P¼0.01 by Z- es ); and bo h LNCaP and P EC o be significan ly highe han RWPE1 (P¼1.5 109,P¼1.2 105, espec i ely, by Z- es ) (Fig. 3). Mo e b oadly, 10 ou o 16 DHS ma ks measu ed in cance cell lines we e obse ed as significan , wi h colo ec al cance as he nex mos significan cance (P¼6.0 1010 by Z- es ; 9.4% o he i abili y localized in 2.0% o genome; Supplemen a y Da a). H3k27ac in LNCaP emained he mos significan ly en iched ma k ac oss all 544 anno a ions (p esen ed in de ail in he Supplemen a y Da a). The mos deple ed ca ego ies we e ep essed egions compu a ionally p edic ed by Segway-ch omHMM in HepG2 cells (P¼1.3 1019 by Z- es ; 51.9% o h2 g om 74.3% o SNPs; Supplemen a y Da a), wi h simila le els o deple ion in ep essed egions om o he cell ypes. These egions a e ypically associa ed wi h dec eased gene exp ession and ep essi e his one ma ks23–25, u he emphasizing he impo ance o ac i e egula ion. As H3k27ac ypically ma ks ac i e enhance s, we u he e alua ed a ian s wi h espec o hei enhance o ‘supe ’-enhance s a us (la ge clus e s o enhance s ha a e en iched o genes in ol ed in cell iden i y26) (see Me hods). We did no obse e di e ences in a e age he i abili y explained by SNPs wi hin he wo ma ks ac oss 49 cell lines (see Me hods), wi h an a e age o 1.51 (1.47)- old inc ease o e andom SNPs o NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979 ARTICLE 0.00 0.05 0.10 0.15 0.20 ARBS LNCaP cell line ARBS p os a e issue no mal+ umou 0.05 0.10 0.15 0.200.00 ARBS p os a e issue umou -only ±5kb ARBS p os a e issue no mal-only ±5kb %SNP-he i abili y %SNP-he i abili y a b Figu e 1 | Func ional pa i ioning o a ian s wi hin ARBS o P Ca. Ba s g aphs de ailing %SNP he i abili y es ima es om wo models o P Ca ele an unc ional anno a ions. (a) Join compa ison o a ian s wi hin 5 kb o umou -only and no mal-only egions in he ARBS in p os a e issue (P¼2.1 1019 o di e ence by Z- es ). (b) Es ima es om ARBS in p os a e issue (no longe using a 5 kb flank) and ARBS in LNCaP cell lines7(P¼4.4 107 o di e ence). The null ð%h 2 g¼%SNPsÞis labelled by he dashed lines. E o ba s show analy ical s anda d e o o es ima e. NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions 3 enhance s (supe enhance s) (Fig. 4). Su p isingly, we obse ed an indi idually significan di e ence only in LNCaP, wi h 4.9 (1.7)- old en ichmen a enhance s (supe enhance s), in con as o p e ious hypo heses26 (Fig. 4). Genomic unc ional a las o p os a e cance SNP he i abili y. Al hough he esul s abo e showcase he powe o he a iance componen app oach in finding epigene ic ma ks ele an o P Ca, such ma ks o en o e lap making he causal ma k di ficul o iden i y (Supplemen a y Fig. 1). To accoun o he co ela ion among ma ks we g ouped he 82 ma ginally significan anno a ions in o 15 biologically ele an , non-o e - lapping g oups o ganized by ma k and cell line, and pa i ioned h2 gac oss all g oups in a join model (see Me hods, Table 1, Fig. 5 and Supplemen a y Table 6). Fi e componen s we e nominally significan in he join model a Po0.05; ou o he fi e componen s h ee emained significan a e accoun ing o 15 es s: H3k27ac ma ks in LNCaP (P¼2.5 1020 by Z- es ); DHS ma ks in o he cance cell ypes (P¼3.9 105by Z- es ); and ep essed segmen a ions (P¼2.1 1020 by Z- es ). To u he efine ou model, we es ic ed o he significan anno a ions (and he backg ound componen s accoun ing o LD o coding egions) and e-e alua ed hem join ly, e e ed o as he ‘selec ed’ model. This selec ed model localized 51.0% o he h2 g wi hin 12.1% o SNPs (LNCaP: H3K27ac þARBS þDHS cance ), whe eas coding egions only explained 3.3% (s.e. 1.4%) o h2 g wi hin 1.8% o SNPs (Supplemen a y Table 7). The localiza ion was e en s onge wi h impu ed da a, whe e 86% o he h2 gwas localized o 8.6% o SNPs (Table 1 and Supplemen a y Tables 8 and 9). Es ima es om impu ed ma ke s we e mo e ep esen a- i e o unde lying en ichmen in ou simula ions (see Me hods, Supplemen a y Table 2) bu may include he e ec s o nea by ma ke s12 and so we conside hem as an uppe bound. None o he es ima es changed significan ly a e adjus ing o known GWAS associa ions2(79 o which we e yped in his da a), unde sco ing he polygenic na u e o his e ec . Ha ing in e ed he selec ed model, we e-analysed each o he 82 ma ginally significan ca ego ies join ly wi h he selec ed model (see Me hods). Only h ee ma ks emained significan : wo H3k27ac anno a ions in he colon c yp and one H3k27ac anno a ion in panc eas (Supplemen a y Da a). This implies ha he ma ginal en ichmen o he 82 anno a ions was p ima ily d i en by he o e lap wi h unc ional ma ks in he selec ed model. Fo example, he H3K4me1 ma k in penis o eskin ke a inocy es ha was p e iously highly significan (24.6% h2 g, P¼3.0 1016 by Z- es , Fig. 1) was no longe en iched a e condi ioning on he selec ed model (7.1% h2 g,P¼0.29 by Z- es , Supplemen a y Da a). The educ ion o a small numbe o ca ego ies in he selec ed model wi h limi ed loss in signal u he emphasizes he ex en o which he selec ed model has localized he unc ional sou ces o en ichmen . Focusing on he wo mos en iched ca ego ies in he selec ed model, we ound ha SNPs p esen in bo h he p os a e issue ARBS and LNCaP H3k27ac ma ks yielded significan ly highe a e age he i abili y pe SNP han ei he ma k indi idually (Supplemen a y Table 10). ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979 0.00 0.05 0.10 0.15 0.00 0.05 0.10 0.15 0.20 0.25 0.00 0.05 0.10 0.15 0.20 0.25 0.00 0.05 0.10 0.15 0.20 0.25 0.00 0.05 0.10 0.15 0.20 0.25 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.05 0.10 0.15 0.20 0.25 DHS % SNP % SNP-he i abili y LnCAP Mamma y epi helium 0.00 0.05 0.10 0.15 H3K27ac % SNP % SNP-he i abili y LNCaP+DHT LNCaP 0.00 0.05 0.10 0.15 H3K4me1 % SNP % SNP-he i abili y Penis o eskin Rec al mucosa 0.00 0.05 0.10 0.15 H3K4me3 % SNP % SNP-he i abili y HUES6 cell line Penis o eskin 0.00 0.05 0.10 0.15 H3K9ac % SNP % SNP-he i abili y Colonic mucosa Rec al mucosa 0.0 0.2 0.4 0.6 0.8 1.0 O he % SNP % SNP-he i abili y Rep essed (HepG2) ARBS Figu e 2 | Func ional pa i ioning o he i abili y ac oss six main epigene ic classes. Each poin co esponds o an es ima e o % SNP he i abili y (yaxis) om SNPs wi hin a cell- ype-specific unc ional anno a ion e sus anno a ion size (%SNPs, xaxis). O e all, 544 anno a ions we e es ed, and ed poin s indica e significan de ia ions om he null o %h 2 gequal o %SNPs a e accoun ing o all es s. The wo mos significan anno a ions in each class a e shown wi h iangle/c oss, espec i ely, and labelled in bo om igh (see Supplemen a y Da a o all anno a ions). 4NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions In con as , he a ian s specific o ARBS o H3k27ac we e compa able in SNP he i abili y. Replica ion o genomic unc ional a las ac oss ances ies.We e alua ed eplica ion o ou model using wo sepa a e genome-wide SNP da a se s o P Ca, one o Eu opean ances y (BPC3; 6,953 samples) and one o A ican ances y (AAPC; 9,522 samples) o P Ca (see Me hods). To accoun o he smalle sample size, we ocused on he eigh -componen selec ed model, only e aining significan componen s and h ee coding-p oximal classes (coding, UTR, p omo e )12. Because o pla o m di e ences be ween he popula ions, we used pos -QC impu ed a ian s in each da a se , which a e mos eflec i e o unde lying en ichmen in ou simula ions (see Me hods). We eplica ed he significan de ia ion in h2 ga H3k27ac and he ep essed loci ac oss bo h BPC3 and AAPC (Supplemen a y Tables 11 and 12). Howe e , cance DHS was only significan in he BPC3 da a and ARBS no significan in ei he ( hough he es ima es we e no significan ly di e en om he iCOGS es ima e). The en ichmen did no change a e es ic ing o e y high-quali y impu ed ma ke s (Supplemen a y Table 13). Al hough he ela i ely small alida ion sample size did no p o ide enough powe o es di e ences be ween he ances ies, he mean SNP he i abili y o a ian s wi hin each ma k we e ema kably simila ( ¼0.90 be ween AAPC and BPC3 ac oss eigh componen s), sugges ing a simila pa e n o agg ega e con ibu ion o isk coming om common a ian s ma ked by epigene ic classes ac oss Eu opean and A ican Ame ican ances ies ( hough indi idual isk a ian s hemsel es may di e ). H3k27ac ma k in LNCaP is specific o P Ca. As a nega i e con ol, we e alua ed he selec ed model wi h impu ed SNPs ac oss 11 common non-cance diseases om he Wellcome T us Case Con ol Conso ium (WTCCC) (see Me hods, Supplemen a y Table 14) whe e we obse ed wo main di e ences: he LNCaP H3k27ac anno a ion was no longe significan ly en iched (1.1% h2 gwi h 2.6% o SNPs); and he ep essed egions we e much less deple ed om he null (28.1% NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979 ARTICLE PREC PREC RWPE1 LNCAP LNCAP RWPE1 2.0% SNPs 6.0% h 2g 0.8% SNPs 4.1% h 2g 2.3% SNPs 11.0% h 2g 2.1% SNPs 9.6% h 2g 0.7% SNPs 0.6% h 2g 0.3% SNPs 1.1% h 2g 2.7% SNPs 14.1% h 2g 0.3% SNPs 1.2% h 2g (P=0.2) 0.7% SNPs 0.7% h 2g (P=7×10 –3 )(P=6×10 –4 )(P=2×10 –7 ) (P=1×10 –6 )(P=0.9) (P=0.2) (P=4×10 –8 )(P=0.9) a b c Figu e 3 | Pai wise analysis o DHS ma ks in h ee p os a e cell ypes. Join model om all pai s o DHS ma ks shown o : cance cell line (LNCAP); no mal p os a e epi helial (PREC); and immo alized p os a e epi helial (RWPE1). Ci cle size co esponds o % SNPs, wi h % SNP he i abili y and significance labelled. P alue was compu ed o di e ence be ween %h 2 gand %SNP, wi h bold ep esen ing significance a e co ec ing o nine es s. The obse ed end is LNCAP4PREC4RWPE1: (a)%h 2 gin LNCaP DHS was nominally significan ly highe han P EC (P¼0.01); and %h 2 gin LNCaP and P EC was significan ly highe han RWPE1 (b,c;P¼1.5 109,P¼1.2 105, espec i ely). All P alues compu ed by Z- es using %h 2 ges ima e and analy ical s anda d e o . CD4 memo y p ima y B ain angula gy us B ain hippocampus middle B ain in e io empo al lobe K562 B ain cingula e Gy us MM1S S omach smoo h muscle Ad enal gland B ain an e io cauda e U87 B ain hippocampus middle Le en icle Ao a Skele al muscle Fe al hymus Adipose nuclei Fe al muscle CD56 T−cell leukemia CD14 CD34 p ima y Lung ib oblas CD20 Lung CD4 De mal ib oblas s Os eoblas s HUVEC As ocy es CD8 p ima y Duodenum smoo h muscle Ju ka Esophagus HeLa HSMM ube Panc ea ic cance B eas cance Lung cance Gas ic Fe al in es ine la ge Skele al muscle myoblas Mamma y epi helial Fe al in es ine Colon cance Colon c yp 1 Colon c yp 2 Panc eas LnCAP Enhance %SNP-he i abili y / %SNP 0246 Supe -enhance %SNP-he i abili y / %SNP 0246 Figu e 4 | Compa ison o enhance s and supe enhance s ac oss 49 cell ypes. Each ba ep esen s he %SNP he i abili y %h 2 g/ %SNP o enhance s (le ) and supe enhance s ( igh ) om a gi en cell ype es ed ma ginally. Red indica es significan di e ence om 1.0 (no en ichmen ) a e accoun ing o 49 es s. Enhance LNCAP is mos significan , wi h o he cance s also appea ing significan and non-cance issues leas significan . E o ba s show analy ical s.e. o es ima e. NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions 5 h2 gwi h 87.8% o SNPs) compa ed wi h he 0.3% o h2 gobse ed in iCOGS impu ed da a (P¼2.2 104 o di e ence by Z- es ). In e es ingly, al hough ARBS we e significan ly en iched in all 11 ai s, he en ichmen was no longe significan a e excluding au oimmune ai s. O e all, hese di e ences indica e ha he LNCaP H3k27ac ma k is uniquely in o ma i e o P Ca, whe eas a ian s nea he ARBS and DHS cance elemen s (which o e lap o he DHS anno a ions by 56%; Supplemen a y Fig. 2) may play a gene ally impo an ole ac oss o he common diseases12. Genomic unc ional a las imp o es polygenic isk p edic ion. To alida e ou SNP he i abili y genomic a las, we compa ed he accu acy o p edic ing case/con ol s a us om gene ic da a wi h o wi hou he unc ional a las. We e alua ed h ee dis inc p edic ion models in he iCOGS sample: (i) a gene ic isk sco e (GRS) om he genome-wide significan SNPs; (ii) he single bes linea unbiased p edic o (BLUP) using a single a iance componen om all SNPs; and (iii) he weigh ed sum o indi idual BLUPs om each epigene ic ca ego y in he selec ed model (mul i-BLUP; see Me hods). E alua ed by c oss- alida ion, he GRS yielded an R2¼0.029 wi h ue pheno ype, whe eas he single BLUP yielded an R2¼0.065 and he mul i-BLUP had an R2¼0.071 (Supplemen a y Table 15). In a join model wi h all h ee p edic o s, he mul i-BLUP was highly significan (P¼5.3 1031 om mul iple eg ession). When we cons uc ed he GRS om SNPs ecen ly disco e ed in a much la ge P Ca GWAS ( e . 2), he esul ing p edic ion R2inc eased o 0.084. Howe e , including he single BLUP o he mul i-BLUP as an addi ional p edic o s ill inc eased he p edic ion R2 o 0.096 (join P¼6.7 104 om mul iple eg ession) and 0.098 (join P¼1.3 1023 om mul iple eg ession), espec i ely (Supplemen a y Table 15). The consis en s a is ical significance and inc eased p edic ion accu acy confi ms he alidi y o he selec ed model in his da a and in la ge GWAS. Discussion Using la ge-scale geno ype da a om o e 59,089 men o Eu opean and A ican Ame ican ances ies join ly wi h epigene ic anno a ions, we iden ified highly significan di e ences in SNP he i abili y ðh2 gÞo P Ca ac oss a ian s om di e en epigene ic classes, issue ypes and cell lines. Focusing on ma ks measu ed in p os a e, we obse ed significan ly highe h2 ga ound umou -specific ARBS; ARBS measu ed in p ima y issue ela i e o cell line; and DHS measu ed in P Ca cell line ela i e o p os a e epi helial cell line. The en ichmen a umou -specific ARBS was consis en wi h ecen findings showing ha hese si es we e en iched o nea by genes highly exp essed in umou s18. These analyses a e comp ehensi e and co e mos commonly s udied p os a e cell lines excep o e eb al cance o he p os a e, which we e no well ep esen ed in he ENCODE/ ROADMAP. A sea ch ac oss 544 di e se unc ional anno a ions es ic ed mos o he h2 g o a small ac ion o he genome ma ked by p os a e egula o y elemen s. Consis en wi h p e ious findings in common disease, unc ionally ep essed egions we e significan ly deple ed in he i abili y, highligh ing he ole o ac i e egula ion in P Ca suscep ibili y. Subsequen model selec ion localized he en ichmen om 82 indi idually significan anno a ions o six ha emained significan in a join model. In pa icula , he abundance o en ichmen in H3k27ac ma ks (ac i e enhance s) ela i e o H3k4me1/H3k4me3 (poised enhance s/p omo e s) unde sco es hei ole in P Ca, hough u he en ichmen in supe enhance s was no obse ed. The en ichmen wi hin LNCaP: H3K27ac and deple ion a ep essed egions was eplica ed ac oss di e en ances ies and yielded significan imp o emen s in polygenic isk p edic ion. Wi h mos GWAS associa ions alling ou side coding egions, ou analyses o e an impo an esou ce o p io i izing po en ial loci and ocusing u u e s udies on he mos he i able genomic egions27. The ma ginal analyses p o ide a anking o 544 common unc ional assays, while he selec ed model localizes he i abili y o only hose unc ional classes ha a e independen ly en iched. Eme ging unc ional ca ego ies may u he efine his signal o e eal o he ele an epigene ic ma ks, hough li le en ichmen beyond he selec ed model was obse ed in he comp ehensi e sampling o unc ional da a analysed he e. In gene al, he a iance componen model o e s an oppo uni y o e alua e biological hypo heses in silico and wi hou s ic ly elying on indi idually significan SNPs. Howe e , as wi h any analysis o a ay-based da a, he h2 ges ima es will no include he con ibu ion o SNPs ha a e un yped o poo ly agged, such ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979 Table 1 | Pa i ioning o he i abili y ac oss unc ional classes in p os a e cance . Func ional ca ego y %SNPs Full Model Selec ed model iCOGS geno yped iCOGS impu ed BPC3 impu ed AAPC impu ed %h 2 gs.e.m. %h 2 gs.e.m. %h 2 gs.e.m. %h 2 gs.e.m. Coding 1.8 3.0 1.3 0.9 2.9 0.2 10.1 3.3 11.1 UTR 1.9 1.6 1.4 3.0 3.1 21.0 11.3 5.9 11.2 P omo e 3.4 *7.8 1.8 8.9 4.1 0.0 12.7 0.0 14.7 LNCaP: H3k27ac 3.2 **22.3 2.1 **27.0 3.8 *30.3 12.1 *28.9 12.7 ARBS 1.0 *3.3 1.1 *9.1 3.3 1.1 12.1 15.2 12.1 LNCaP: FOXA1 1.5 1.5 1.3 LNCaP: H3k4me1 2.0 1.3 1.4 LNCaP: DHS 2.9 5.4 1.6 DHS p os a e 1.8 2.6 1.4 DHS cance 4.7 **14.1 2.3 **49.6 6.3 *47.4 21.4 46.6 22.4 H3k4me1 (o he ) 16.3 19.6 3.5 H3k27ac (o he ) 7.3 4.1 2.4 DHS (o he ) 1.8 0.2 1.3 ep essed 48.7 **11.0 4.1 **0.3 7.0 **0.0 23.8 **0.0 24.5 all o he 1.7 0.7 1.2 0.2 2.7 0.0 9.2 0.0 7.6 ARBS, and ogen ecep o -binding si es; DHS, DNase I hype sensi i i y si es; SNP, single-nucleo ide polymo phism; UTR, un ansla ed egion. Full model deno es a 15- a iance componen s model while ‘selec ed’ model deno es a model es ic ed o he fi e componen s a aining significance in he ‘ ull’ model (and h ee componen s o backg ound). * (**) deno es significan de ia ion a Po0.05 (Po0.05/15) o ac ion o SNP he i abili y ð%h 2 gÞ om null model o %h 2 g¼%SNPs (by Z- es ; see Supplemen a y Table 6 o P alues). 6NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions as a e a ian s o o he con ibu o s o he missing he i abili y. Fu u e analyses o whole-genome sequencing, addi ional unc ional anno a ions, and la ge sample sizes can yield impo an insigh s in o unc ional mechanisms ha a e s ill no localized. O e all, ou esul s sugges simila pa e ns o unc ional en ichmen ac oss men o Eu opean and A ican Ame ican ances y and highligh ARBS, H3k27ac ma ks in LNCaP cell lines and DHS in cance cell lines o ollow-up s udies o P Ca isk. Me hods Epigene ic anno a ions.Sample collec ion and p ocessing o unc ional anno- a ions was made publically a ailable by he ENCODE/ROADMAP conso ia28. DHS, H3k4me1, H3k4me3, H3k9ac anno a ions and genome segmen a ions20,29, enhance s and supe enhance s26 and P Ca-specific anno a ions7,18 we e assay and p ocessed as de ailed in he o iginal s udies. Tumou -only and no mal-only ARBS we e defined in se en no mal and 13 umou specimens in he o iginal s udy18. All anno a ions cu a ed o his pape (ENCODE/ROADMAP; Pome an z e al.; and Hazele e al.) a e a ailable a h ps://da a.b oadins i u e.o g/alkesg oup/ ANNOTATIONS/PRCA/. The ull lis o indi idual anno a ions wi h web-links o he co esponding bounda y defini ions is p o ided in he Supplemen a y Da a. Some unc ional ma ks a e lis ed mul iple imes due o mul iple independen assays o labo a o y p o ocols. ARBS ChIP-seq in human issue specimens.The ARBS assay was pe o med as desc ibed in REF ( e . 18) and summa ized he e. Fou een subjec s o Eu opean Ame ican ances y we e selec ed o ChIP analysis. Thei ch oma in was incuba ed o e nigh wi h 6 mg an ibody AR (N-20, San a C uz Bio echnology, Dallas, TX) NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979 ARTICLE iCOGS main model Coding P omo e * LnCAP H3K27ac** ARBS* LnCaP FOXA1 LnCaP H3K4me1 LnCaP DHS DHS p os a e DHS cance ** H3K4me1 (o he ) H3K27ac (o he ) DHS (o he ) D Rep essed** O he × 0.5 × 5 × 10 iCOGS selec ed model Coding UTR P omo e * LnCAP H3K27ac** ARBS* DHS cance ** Rep essed** O he × 0.5 × 5 × 10 AAPC selec ed model Coding UTR P omo e LnCAP H3K27ac* ARBS DHS cance Rep essed** O he × 0.5 × 5 × 10 × 25 BPC3 selec ed model UTR LnCAP H3K27ac* DHS cance * Rep essed** O he × 0.5 × 5 × 10 × 25 ab cd Figu e 5 | Pa i ioning o he i abili y ac oss unc ional classes in p os a e cance . Visual ep esen a ion o he i abili y en ichmen in h ee s udies a,b: iCOGS; c: AAPC; d: BPC3 (shown nume ically in Table 1). Each subplo co esponds o an analysis o he lis ed join model, wi h colou ed slices ep esen ing he unc ional anno a ions e alua ed. Volume o each in e io (ligh colou ed) pie-cha slice ep esen s he %SNP o he unc ional anno a ion, which is equal o he expec ed %h 2 gunde he null o no en ichmen . Volume o each shaded pie-cha slice ep esen s he ac ual %h 2 g in e ed by he model. Slices ex ending ou side/inside he middle pie co espond o en ichmen /deple ion in SNP he i abili y, as indica ed by he do ed lines. Colou coding is consis en ac oss all subpanels. * (**) deno es significan de ia ion a Po0.05 (Po0.05/15) o ac ion o SNP he i abili y (%h 2 g om null model o %h 2 g¼%SNPs by Z- es ; see Supplemen a y Table 6 o P alues). NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions 7 bound o p o ein A and p o ein G beads (Li e Technologies, Ca lsbad, CA). A ac ion o he sample was no exposed o an ibody o be used as con ol (inpu ). The samples we e de-c osslinked, ea ed wi h RNase and p o einase K, and DNA was ex ac ed. The samples we e hen e-shea ed o 100–300 base pai s using he Co a is ul a-sonica o , and concen a ions o he ChIP DNA we e quan ified by Qubi Fluo ome e (Li e Technologies). DNA sequencing lib a ies we e p epa ed using he Th uPLEX-FD P ep Ki (Rubicon Genomics, Ann A bo , MI). Lib a ies we e sequenced using 50-base pai eads on he Illumina pla o m (Illumina, San Diego, CA) a Dana-Fa be Cance Ins i u e. AR binding si es we e gene a ed using Model-Based Analysis o ChIP-seq 2 (MACS2), wi h a q alue ( alse disco e y a e, FDR) h eshold o 0.01. The 13 umou s used in his s udy we e and ogen dependen and no exposed o and ogen dep i a ion he apies. All o he umou s we e specimens ob ained om adical p os a ec omies, de i ed om men wi h ea ly s age disease. These samples we e no selec ed based on any specific ea u es; he e o e, we would expec ha he dis ibu ion o isk a ian s would be simila o a andom sampling o P Ca cases. La ge-scale gene ic su eys ha e shown ha soma ically acqui ed al e a ions in p ima y localized p os a e umou s ( he ype o umou e alua ed in his s udy) a e in equen . Based on hese p e ious esul s, we belie e ha soma ically acqui ed gene ic e en s in egions ela ed o and ogen biology a e no common and, he e o e, do no influence ou esul s. Pa ien ma e ial.In o med consen was ob ained om all subjec s and all s udies we e app o ed by local Resea ch E hics Commi ees and/o Ins i u ional Re iew Boa ds. Da a quali y con ol.Quali y con ol is c ucial o accu a e he i abili y es ima ion, whe e many small a i ac s can add up o la ge biases. All da a se s wen h ough a s ingen QC p ocess wi h he ollowing exclusion c i e ia: mino allele equency (MAF)o1%; ac ion o missing/uncalled SNPs45%; Ha dy–Weinbe g equilib ium P alueo0.01; case–con ol missingness P alueo0.05; impu a ion INFO sco e40.30. In addi ion, close ela i es we e p uned such ha no pai o indi iduals had gene ic ela edness (GRM) coe ficien s40.05. The op 10 p incipal componen s and a coded s udy label we e always included as fixed-e ec s. All analysed samples, cases and con ols, we e males. iCOGS da a.The iCOGS conso ium geno yped balanced cases and con ols on a cus om a ge ed a ay4. A e quali y con ol, 42,613 samples and 153,621 geno yped SNPs emained. Impu a ion was pe o med o he 1000 Genomes e e ence panel using HAPI-UR ( e . 30) o phasing and IMPUTE2 ( e . 31) o impu a ion. O e all, 1,910,827 impu ed and geno yped SNPs passed QC. Because o compu a ional es ic ions, he he i abili y es ima ion was ca ied ou in wo equally sized hal es o he ICOGS, wi h o al e ec s compu ed by in e se- a iance me a-analysis. We pa i ioned he geno yped SNP he i abili y by MAF bu obse ed no end and only sligh en ichmen o % h2 ga high- equency a ian s (Supplemen a y Table 16). BPC3 da a.The Na ional Cance Ins i u e B eas & P os a e Cance Coho Conso ium (BPC3) conso ium geno yped indi iduals on he Illumina Human- Hap610 quad a ay32. A e quali y con ol, 6,953 samples and 4,004,229 geno yped and impu ed SNPs emained. Age was a ailable o all samples and addi ionally included as a co a ia e. AAPC da a.The AAPC conso ium geno yped indi iduals o A ican ances y on he Illumina Human1M a ay2,33,34. A e quali y con ol, 9,522 samples and 10,468,389 geno yped and impu ed SNPs emained. WTCCC da a.The Wellcome T us Case Con ol Conso ium Geno yping geno yped cases o 11 ai s as well as sha ed con ols on mul iple Illumina and A yme ix a ays35–37. The pheno ypes analysed he e we e ankylosing spondyli is (AS); bipola diso de (BD); co ona y a e y disease (CAD); C ohn’s disease (CD); hype ension (HT); mul iple scle osis (MS); heuma oid a h i is (RA); schizoph enia (SP); ype 1 diabe es (T1D); ype 2 diabe es (T2D); and ulce a i e coli is (UC). A e quali y con ol, a o al o 47,053 samples and 4–5 million geno yped and impu ed SNPs emained. Repo ed h2 g alues we e es ima ed o each pheno ype sepa a ely and me a-analysed using in e se- a iance weigh ing. UK10K da a.The UK10K whole-genome sequence da a om ALSPAC and TWINSUK (h p://www.uk10k.o g) was used only o simula ion, and so s ingen quali y con ol was no applied. A e ela edness fil e ing, 3,047 samples and 15,691,225 non-single on a ian s we e e ained. He i abili y es ima ion o indi idual anno a ions.We es ima ed he SNP he i abili y ðh2 gÞcap u ed by unc ional ca ego ies in a join a iance componen model using GCTA as desc ibed in REF ( e . 20). B iefly, his model assumes he pheno ype is d awn om a mul i a ia e no mal dis ibu ion wi h a iance-co a iance modelled by componen s compu ed om he SNPs and a no mal esidual. Fo each unc ional ca ego y ( o example, DHS) i¼1..Mwhe e Mis he o al numbe o ca ego ies in he model, a GRM ac oss all pai s o indi iduals is compu ed es ic ing o SNPs wi hin he unc ional ca ego y. Va iance componen s o all GRMs in he model a e hen fi ed using REML as implemen ed in GCTA o es ima e a a iance pa ame e s2 i  used o compu e %h 2 i¼s2 i=PM j¼1s2 j.Theh2 ico esponds o he ac ion o ai a iance ha can be explained by he BLUP es ic ed o SNPs in he co esponding unc ional ca ego y (o anno a ion). Fo a gi en unc ional anno a ion, SNPs we e ca ego ized in o a hie a chy o se en non-o e lapping componen s: (1) coding; (2) UTR; (3) p omo e ( unc ional anno a ion o in e es ); (4) DHS; (5) in on; and (6) in e genic. SNPs belonging o mul iple ca ego ies we e pa i ioned explici ly in o he fi s ca ego y in his lis . The coding and coding-p oximal componen s we e included o ensu e ha he anno a ion he i abili y was no infla ed by SNPs ha we e in high LD wi h coding a ia ion. A gene ic ela edness ma ix was compu ed o each componen by fi s s anda dizing he co esponding SNPs and hen compu ing a SNP co a iance o e all pai s o samples. Componen -specific s2and e o s we e fi ed i e a i ely using he A e age In o ma ion algo i hm38.The analy ical s anda d e o o %h 2 iwas es ima ed by ans o ming he GCTA- in e ed s2 iand e o co a iance ma ix using he del a me hod. As in REF ( e . 20) s a is ical significance was e alua ed by compa ing he %h 2 gexplained by he ca ego y and i ’s s anda d e o o he %SNPs in he ca ego y using a Z- es (compa ing nes ed models using a likelihood a io es yielded simila esul s). To al h2 ges ima es we e compu ed as h2 g¼PM j¼1s2 j=PMþ1 j¼1s2 ja e ans o ming o he liabili y scale assuming a p e alence o 0.14 and using he s udy-specific case/ con ol a io. Hie a chical join models.Fo specific models o in e es , we ex ended he indi idual anno a ion model desc ibed abo e o es in e sec ing and non-in e - sec ing componen s. This allowed us o e alua e p ecisely which sub-anno a ions o o e lapping componen s we e likely o be causal. Fo he umou /no mal model, we expanded each umou /no mal ma k by 5 kb in bo h di ec ions om he cen e o cap u e nea by genes and o he egula o y egions so ha umou (no mal) co e ed 3.3% (1.4%) o he SNPs, espec i ely. We es ima ed h2 g om he join hie a chical model: (1) coding; (2) UTR; (3) p omo e ; (4) no mal-only; (5) umou -only; (6) DHS; (7) in on; and (8) O he . When compa ing ARBS om issue and ARBS LNCaP om cell line, only 59 SNPs (0.03%) o e lapped be ween he wo ca ego ies, and so we es ed wo sepa a e models: (1) coding; (2) UTR; (3) p omo e ; (4) (ARBS issue/ARBS LNCaP); (5) DHS; (6) in on; and (7) o he . Fo compa isons be ween LNCAP, PREC and RWPE1 using DHS we es ed each pai o cell lines using he join model: (1) coding; (2) UTR; (3) p omo e ; (4) DHS pa icula o one cell line; (5) DHS common o bo h cell lines; (6) DHS pa icula o o he cell line; (7) DHS o he cell lines; (8) In on; and (9) O he . Fo com- pa isons be ween enhance s and supe enhance s, we used he 86 cell- ype-specific anno a ions om REF ( e . 26), es ing each enhance o supe enhance sepa a ely in he ollowing join model: (1) coding; (2) UTR; (3) p omo e , (4) (enhance / supe enhance o cell- ype o in e es ); (5) DHS; (6) in on; (7) o he . O hese, 49 cell ypes yielded model con e gence o bo h he enhance and co esponding supe enhance and we e used o es ima e means and co ela ion. The o de and g ouping o ma ginally significan anno a ions in o epigene ic ma k and cell ype ( o example, in Table 1) a e lis ed in he Supplemen a y Da a. Fo each o he 82 indi idually significan anno a ions, we e-e alua ed hem join ly wi h he selec ed model in he ollowing hie a chical join model: (1) coding; (2) UTR; (3) p omo e ; (4) LNCaP:H3k27ac; (5) ARBS; (6) DHS cance ; (7) ( unc ional anno a ion o in e es ); (8) DHS; (9) in on; and (10) o he . Only unc ional anno a ions ha con e ged we e epo ed in he Supplemen a y Da a. Accu acy o h2 ges ima es om yped a ian s in simula ions.The a iance componen model has p e iously been shown o yield obus es ima es unde he assump ion ha causal a ian s a e yped and uni o mly sampled om a gi en unc ional ca ego y13,20,21. He e, we pe o m simula ions using he UK10K whole- genome sequence da a o confi m he alidi y o his model o ou anno a ions, and o assess how ep esen a i e SNP es ima es a e o ue unde lying biology a common sequenced a ian s. O e all, he simula ions in ol e using eal ma ke s o gene a e addi i e, polygenic pheno ypes wi h a gi en he i abili y and hen es ima ing he he i abili y wi h he a iance componen model. We e alua ed he UK10K da a o h ee ypes o SNPs: (i) common sequenced a ian s (7,534,538 SNPs); (ii) UK10K SNPs yped by he iCOGS pla o m (178,509; 95% o iCOGS SNPs); and (iii) UK10K SNPs yped and impu ed by he iCOGS pla o m (1,655,723; 87% o he iCOGS impu ed SNPs). We ocused on he LNCaP:H3k27ac anno a ion (which was mos significan in ou da a) o e alua e he main join model. All pheno ypes we e simula ed by d awing 5,000 causal a ian s andomly om he specified ca ego ies and sampling causal e ec -sizes om a no mal dis ibu ion such ha SNPs ei he explain equal a iance ( he model assump ion) o a iance in p opo ion o hei MAF. The pheno ype was hen gene a ed as he do p oduc o geno ype and e ec -size wi h andom noise added o fix he i abili y a 50%. Pheno ypes we e simula ed housands o imes un il he s anda d e o o e simula ions was low enough o e alua e unbiasedness. ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979 8NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions We confi med ha es ima es o h2 g om a polygenic ai we e accu a e unde he model whe e causal a ian s a e yped (Supplemen a y Table S1). Unde he null, he LNCaP H3k27ac componen is expec ed o explain 3.22% o he SNP he i abili y, and he model es ima ed 3.50% (0.22%) and 3.68% (0.21%) unde a low- equency and high- equency disease a chi ec u e, espec i ely (Supplemen a y Table S1). None o he es ima es we e significan ly di e en om he u h gi en he numbe o componen s es ed. Unde a scena io whe e LNCaP H3k27ac explains 50% o he h2 g, he model es ima ed 51.13% (0.40%) and 46.98% (0.35%) unde a low- equency and high- equency disease a chi ec u e, espec i ely (Supplemen a y Table S1). Al hough he high- equency a chi ec u e (whe e common a ian s explain mo e a iance in ai han a e a ian s) ep esen s a subs an ial model misspecifica ion, ou simula ions show ha his does no in oduce subs an ial bias and is likely o sligh ly unde es ima e he SNP he i abili y a he ocal ch oma in ma k. In all cases, he empi ical s anda d de ia ion o e 500 simula ions was simila o he a e age analy ical s.e.m. compu ed by GCTA (REML algo i hm), hus showing ha ha analy ical s anda d e o is well calib a ed (Supplemen a y Table S1). We no e ha he s anda d e o is in e sely ela ed o he sample size39,40, and is he e o e much highe in hese simula ions han in he iCOGS da a which is 14- old la ge . Las ly, we pe o med he eal da a pa i ioned analysis in subse s o indi iduals o e alua e biasedness and powe o de ec significan en ichmen . We confi med ha no significan di e ences we e obse ed be ween es ima es om he en i e s udy compa ed wi h hose a e aged ac oss subse s o he s udy (Supplemen a y Fig. 3). As such, we can confiden ly epo es ima es and bounds on he en ichmen obse ed in he en i e s udy ha will hold o la ge s udies. Fu he mo e, all bu one o he significan componen s om he main model emained significan in smalle samples (ARBS), making i unlikely ha hey we e a ec ed by winne ’s cu se. Recen wo k has quan ified he heo e ical ela ionship be ween es ima ion e o and e ec i e sample size o indi idual componen s39,40. Causal a ian s no agged on he iCOGS geno yping pla o m.We used he sequenced UK10K common a ian s o e alua e how well he iCOGS geno yped and impu ed SNPs cap u ed unde lying he i abili y by simula ing pheno ypes using causal a ian s om sequencing and es ima ing he i abili y om he iCOGS SNPs ( ha is, hiding a ian s ha we e no geno yped o impu ed, Supplemen a y Table S2). 83% o common UK10K SNPs lie wi hin 100kb o an iCOGS SNP, so some common a - ia ion is likely o be pa ially agged by he chip. I he impu ed and/o geno yped SNPs se ed as a good p oxy o he common sequence a ia ion, hen we would expec hei es ima es o %h 2 g o ma ch he simula ed ac ions. When no unc ional ca ego y was en iched wi h causal a ian s, small bu significan di e ences we e obse ed o geno yped coding a ian s (4.75% h2 ges ima ed as compa ed wi h simula ed 0.67%) and impu ed in e genic a ian s (56.09% h2 gas compa ed wi h 50.52% simula ed) bu no he ocal LNCaP:H3k27ac ca ego y. Simila de ia ions we e obse ed o he disease a chi ec u e whe e common a ian s explain mo e a iance in ai han a e a ian s (Supplemen a y Table S2). When causal a ian s whe e en iched wi hin LNCaP:H3k27ac ca ego y, de ia ions be ween simula ed and es i- ma ed SNP he i abili y we e la ge (Supplemen a y Table S2). Mos o his de ia ion was due o a significan unde es ima e a LNCaP:H3k27ac, which was simula ed o explain 50% o h2 gbu explained only 12.55% (s.e.m.0.92%) and 30.92% (s.e.m. 1.09%) om geno yped and impu ed SNPs, espec i ely. This he i abili y was dis ibu ed ac oss all he emaining componen s, pa icula ly in in e genic SNPs o he geno- yped es ima e and DHS SNPs o he impu ed es ima e, which end o be nea by. O e all, ou simula ions showed ha he model is highly accu a e when all causal a ian s a e yped. When conside ing en ichmen om un yped causal a ian s, he impu ed es ima e was consis en ly close o he u h han he geno yped es ima e. Mos impo an ly, he es ima e om he ocal ca ego y (LNCaP H3k27ac in ou simula ions) was shown o be highly conse a i e bo h in he null and in he en iched scena io and unlikely o be biased due o agging o un yped ma ke s. We no e ha p e ious wo k has shown es ima es om impu ed SNPs (bu no geno yped SNPs) may be con amina ed by ma ke s e y close o an en iched anno a ion12;assuchwe ocused ou esul s on he densely geno yped iCOGS a ian s which a e expec ed o be conse a i e, and p ima ily used impu ed da a o alida ion ac oss da a se s. Es ima es o h2 g om A ican Ame ican samples.To assess po en ial biases in es ima ing h2 g om an admixed popula ion, we pe o med sepa a e simula ions in he AAPC da a whe e causal a ian s we e specifically sampled om a ying F ST bins. This amewo k e alua ed he po en ial bias esul ing om ma ke s ha had d i ed o di e en equencies in he wo popula ions. The F ST was es ima ed ou - o -sample in he HapMap CEU Eu opean and YRI Yo uba popula ions. We es ed he null six-componen model (Coding, UTR, P omo e , DHS, In on, O he ) and obse ed no significan de ia ions om he null unde any class o di e en ia ed SNPs (Supplemen a y Table S3). Howe e , we no e ha o al h2 gwas simula ed a 0.50 bu was in e ed a 0.38–0.66 ac oss inc easing quin iles o causal F ST (Supplemen a y Table 3), indica ing ha e en wi h well-calib a ed es ima es o en ichmen he o al es ima e may be biased upwa ds i he causal SNPs a e highly di e en ia ed (obse ed in his simula ion when mean causal F ST 4¼0.35). Gene ic p edic ion.We sough o alida e he u ili y o ou unc ional a las by applying i o gene ic p edic ion. The aim o gene ic p edic ion is o use aining indi iduals wi h gene ics ( o example, SNPs) and diagnosed pheno ype o accu a ely p edic he pheno ype in o indi iduals wi h only gene ic da a a ailable41,42.He e,we ocus on co ela ion o p edic ed pheno ype wi h ue pheno ype (R2), as i has a na u al ela ionship o SNP he i abili y12,42. In ui i ely, be e localiza ion o he ue e ec -sizes will educe noise in aining he p edic o and inc ease accu acy. I he unc ional a las iden ified egions wi h inc eased he i abili y, his in o ma ion should significan ly imp o e he p edic ion. We e alua ed h ee s anda d models o isk p edic ion: GRS; BLUP ( e . 43); and mul i-componen BLUP ( e . 14). The GRS was compu ed as a sum o e SNPs o he log odds- a ios om he aining sample41.The se o SNPs used was ei he he genome-wide significan ma ke s in he aining se ( es ic ed o one pe 1MB locus) o he genome-wide significan ma ke s iden ified in a ecen la ge GWAS o P Ca2. In con as o he GRS, he BLUP used all ma ke s in he da a o o m he p edic ion. The s anda d BLUP was es ima ed using GCTA o e all SNPs. The mul i-componen BLUP was es ima ed using he componen s in he selec ed model (join ly) o compu e a single sco e equal o he sum o he p edic ions om each componen weigh ed by hei componen -specific h2 g.Thisis analogous o speci ying a di e en p io on he e ec -size a iance in each componen . All p edic ions we e ca ied ou by c oss- alida ion in he ull iCOGS da a, emo ing 1,000 indi iduals in each old. P edic ion R2was hen compu ed om a eg ession o pheno ype on he p edic o sco e wi h 10 PCs included as co a ia es o accoun o ances y, subsequen ly sub ac ing he R2¼0.021 om a model wi h PCs only. P alues we e es ima ed o each o he coe ficien s in he mul iple eg ession o pheno ype BGRS þsingle-BLUP þmul i-BLUP þPCs. To ensu e ha p edic ion ac oss da a se s was independen , we ca e ully emo ed all iCOGS indi iduals wi h a GRM alue o 40.05 o any indi idual in he BPC3 when compu ing BLUP coe ficien s. We sepa a ely analysed he p edic o in 26,000 iCOGS samples ha had age a diagnosis, bu did no obse e significan di e ences be o e/ a e including age as a co a ia e. Re e ences 1. Hjelmbo g, J. B. e al. The he i abili y o p os a e cance in he No dic win s udy o cance . Cance Epidemiol. Bioma ke s P e . 23, 2303–2310 (2014). 2. Al Olama, A. A. e al. A me a-analysis o 87,040 indi iduals iden ifies 23 new suscep ibili y loci o p os a e cance . Na . Gene . 46, 1103–1109 (2014). 3. Cas o, E. e al. Ge mline BRCA mu a ions a e associa ed wi h highe isk o nodal in ol emen , dis an me as asis, and poo su i al ou comes in p os a e cance . J. Clin. Oncol. 31, 1748–1757 (2013). 4. Eeles, R. A. e al. Iden ifica ion o 23 new p os a e cance suscep ibili y loci using he iCOGS cus om geno yping a ay. Na . Gene . 45, 385–391 (2013). 5. Saunde s, E. J. e al. Fine-mapping he HOXB egion de ec s common a ian s agging a a e coding allele: e idence o syn he ic associa ion in p os a e cance . PLoS Gene . 10, e1004129 (2014). 6. Ewing, C. M. e al. Ge mline mu a ions in HOXB13 and p os a e-cance isk. N. Engl. J. Med. 366, 141–149 (2012). 7. Hazele , D. J. e al. Comp ehensi e unc ional anno a ion o 77 p os a e cance isk loci. PLoS Gene . 10, e1004102 (2014). 8. Hazele , D. J., Coe zee, S. G. & Coe zee, G. A. A a e a ian , which des oys a FoxA1 si e a 8q24, is associa ed wi h p os a e cance isk. Cell Cycle 12, 379–380 (2013). 9. ENCODE P ojec Conso ium e al. An in eg a ed encyclopedia o DNA elemen s in he human genome. Na u e 489, 57–74 (2012). 10. S ama oyannopoulos, J. A. Wha does ou genome encode? Genome Res. 22, 1602–1611 (2012). 11. Mau ano, M. T. e al. Sys ema ic localiza ion o common disease-associa ed a ia ion in egula o y DNA. Science 337, 1190–1195 (2012). 12. Guse , A. e al. Pa i ioning he i abili y o egula o y and cell- ype-specific a ian s ac oss 11 common diseases. Am. J. Hum. Gene . 95, 535–552 (2014). 13. Yang, J. e al. Common SNPs explain a la ge p opo ion o he he i abili y o human heigh . Na . Gene . 42, 565–569 (2010). 14. Yang, J., Lee, S. H., Godda d, M. E. & Vissche , P. M. GCTA: a ool o genome-wide complex ai analysis. Am. J. Hum. Gene . 88, 76–82 (2011). 15. Yang, J. e al. Genome pa i ioning o gene ic a ia ion o complex ai s using common SNPs. Na . Gene . 43, 519–525 (2011). 16. Schumache , F. R. e al. Genome-wide associa ion s udy iden ifies new p os a e cance suscep ibili y loci. Hum. Mol. Gene . 20, 3867–3875 (2011). 17. Haiman, C. A. e al. Cha ac e izing gene ic isk a known p os a e cance suscep ibili y loci in A ican Ame icans. PLoS Gene . 7, e1001387 (2011). 18. Pome an z, M. e al.The and ogen ecep o cis ome is ex ensi ely ep og ammed in human p os a e umo igenesis.Na .Gene .47, 1346–1351 (2015). 19. Shlyue a, D., S amp el, G. & S a k, A. T ansc ip ional enhance s: om p ope ies o genome-wide p edic ions. Na . Re . Gene . 15, 272–286 (2014). 20. Guse , A. e al. Pa i ioning he i abili y o egula o y and cell- ype-specific a ian s ac oss 11 common diseases. Am. J. Hum. Gene . 95, 535–552. 21. Lee, S. H. e al. Es ima ion o SNP he i abili y om dense geno ype da a. Am. J. Hum. Gene . 93, 1151–1155 (2013). 22. Cance Fac s & Figu es o A ican Ame icans 2009–2010. Accessed on: Decembe 2015. NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979 ARTICLE NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions 9