Full text
ARTICLE
Recei ed 26 Dec 2014 |Accep ed 3 Feb 2016 |Published 7 Ap 2016
A las o p os a e cance he i abili y in Eu opean
and A ican-Ame ican men pinpoin s
issue-specific egula ion
Alexande Guse e al.#
Al hough genome-wide associa ion s udies ha e iden ified o e 100 isk loci ha explain
B33% o amilial isk o p os a e cance (P Ca), hei unc ional e ec s on isk emain
la gely unknown. He e we use geno ype da a om 59,089 men o Eu opean and A ican
Ame ican ances ies combined wi h cell- ype-specific epigene ic da a o build a genomic
a las o single-nucleo ide polymo phism (SNP) he i abili y in P Ca. We find significan
di e ences in he i abili y be ween a ian s in p os a e- ele an epigene ic ma ks defined in
no mal e sus umou issue as well as be ween issue and cell lines. The majo i y o SNP
he i abili y lies in egions ma ked by H3k27 ace yla ion in p os a e adenoc7a cinoma cell line
(LNCaP) o by DNaseI hype sensi i e si es in cance cell lines. We find a high deg ee o
simila i y be ween Eu opean and A ican Ame ican ances ies sugges ing a simila gene ic
a chi ec u e om common a ia ion unde lying P Ca isk. Ou findings showcase he powe
o in eg a ing unc ional anno a ion wi h gene ic da a o unde s and he gene ic basis o P Ca.
Co espondence and eques s o ma e ials should be add essed o A.G. (email: [email p o ec ed].edu) o o B.P (email: [email p o ec ed]).
#A ull lis o au ho s and hei a filia ions appea s a he end o he pape .
DOI: 10.1038/ncomms10979 OPEN
NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions 1
Family his o y is a well-es ablished isk ac o o p os a e
cance (P Ca), which has an es ima ed he i abili y o
58%—one o he highes ac oss common cance s1.
Genome-wide associa ion s udies (GWAS) ha e been
pa icula ly success ul in iden i ying o e 100 isk loci ha
cap u e B33% o he es ima ed amilial isk2. Al hough mos
o he GWAS P Ca a ian s o e lap p os a e-specific
egula o y elemen s ( o example, and ogen ecep o -binding
si es (ARBS))2–8, a quan ifica ion o he con ibu ion o
gene ic a ia ion om a ious ch oma in ma ks o P Ca isk is
cu en ly lacking.
Recen wo k o m he ENCODE/ROADMAP conso ia9has
shown ha a la ge ac ion o he genome plays a ole in a
leas one biochemical e en , in a leas one issue. Al hough his
unc ional a las o he human genome has g ea ly enhanced
ou unde s anding o egula o y elemen s, such unc ional
elemen s a e o en issue specific10,11 making hei
in e p e abili y in he con ex o P Ca isk challenging. Exis ing
s udies ha ha e in eg a ed P Ca GWAS findings wi h
issue-specific unc ional anno a ions ha e elied only on he
GWAS significan a ian s (B100 in he mos ecen s udy) o
single-nucleo ide polymo phisms (SNPs) agging hem2,7, hus
igno ing loci ha do no each genome-wide significance. Recen
me hodological ad ances ha e shown ha he en i e polygenic
a chi ec u e o common ai s can be in e oga ed using
a iance componen s ac oss all assayed SNPs ( yped and/o
impu ed) o inc ease powe o de ec ing ai -specific unc ional
anno a ions12. In addi ion o o e ing supe io pe o mance
ela i e o me hods ha e alua e only GWAS SNPs, he
a iance componen s me hods also allow o compa ison o
es ima es ac oss di e en s udies and sample sizes. This is
because a iance componen s yield an unbiased es ima e (unde
s anda d assump ions) o SNP he i abili y ðh2
gÞ— he a iance in
ai explained by SNPs ha eside wi hin elemen s o a gi en
unc ional ca ego y12–15.
He e, we use a ge ed and genome-wide SNP a ay da a om
59,089 male P Ca cases and con ols o Eu opean (BPC3 ( e . 16)
and iCOGS ( e . 4), espec i ely, see Me hods) and A ican
Ame ican (AAPC ( e . 17), see Me hods) ances y o dissec he
gene ic isk o P Ca. We es ima e he SNP he i abili y o p e iously
implica ed egula o y anno a ions7,18 and pe o m a b oad analysis
o 544 epigene ic ma ks om ENCODE/ROADMAP ( e . 9). Ou
app oach in e oga es he en i e common polygenic a chi ec u e o
P Ca while accoun ing o po en ial co ela ions be ween ela ed
unc ional ca ego ies. Fi s , we find ha SNPs nea ARBS assayed in
p os a e umou explain significan ly mo e o he he i abili y o P Ca
han ARBS SNPs assayed in p os a e no mal issue. Second, we
localize mos o he he i abili y o P Ca o egions in he genome
ma ked by h ee unc ional ca ego ies: (i) H3K27ac his one
modifica ions in p os a e adenoca cinoma cell lines (LNCaP;
ypically ma king ac i e enhance s19); (ii) and ogen ecep o s in
p os a e issue18; and (iii) DNase I hype sensi i i y si es (DHS) in
cance cell lines. We eplica e he LNCaP H3K27ac and DHS esul s
ac oss di e en ances ies and show ha isk p edic ion om
genome-wide SNP da a is significan ly imp o ed wi h a p edic o
ha inco po a es he unc ional a las as p io . O e all, ou esul s
sugges a simila gene ic a chi ec u e om common a ia ion o
P Ca isk ac oss men o Eu opean and A ican ances y and
highligh H3k27ac his one ma k in LNCaP and ARBS in p os a e
issue o ollow-up s udies o P Ca isk.
Resul s
Pa i ioning he gene ic isk o p os a e cance . We analysed
mul iple unc ional anno a ions and quan ified he ac ion o
a iance in ai explained by SNPs ha a e localized wi hin each
unc ional class. Ou app oach models he pheno ype (P Ca) o a
se o indi iduals as being d awn om a mul i a ia e no mal
dis ibu ion wi h a iance componen s es ima ed based on
gene ic da a ( ha is, SNPs) plus an en i onmen al e m
(see Me hods)13,14. Fo each unc ional ca ego y i, a gene ic
ela ionship ma ix ac oss all indi iduals is compu ed om all he
SNPs esiding in he gi en unc ional ca ego y o se e as a
a iance componen . Mul iple componen s a e hen join ly fi ed
using he es ic ed maximum likelihood (REML) as implemen ed
in he GCTA so wa e14 o es ima e a iance pa ame e s s2
i
o
each componen . The SNP he i abili y o componen iis hen
es ima ed as h2
g;i¼s2
i=Pjs2
j, whe e he sum in he denomina o
is ac oss all fi ed componen s including he en i onmen al e m.
The e o e, we iew h2
g;ias an es ima e o he a iance in ai ha
can be explained by all he SNPs in he co esponding unc ional
ca ego y wi h a linea model o he ai ( ha is, SNP
he i abili y)12. We expec unc ional ca ego ies ha a e en iched
wi h casual a ian s o P Ca o a ain a highe es ima ed SNP
he i abili y as compa ed wi h unc ional ca ego ies deple ed o
causal a ian s o P Ca. To ocus ou esul s on noncoding
a ia ion and accoun o po en ial con ounde s because o
linkage disequilib ium (LD), we explici ly included coding and
coding-p oximal egula o y a ia ion as ‘backg ound’ compo-
nen s whene e we quan ified he e ec o each unc ional
anno a ion es ed (see Me hods).
The a iance componen model has p e iously been shown o
yield obus es ima es unde he assump ion ha causal a ian s
a e yped and uni o mly sampled om a gi en componen 13,20,21.
He e, we pe o m addi ional simula ions using he UK10K
whole-genome sequence da a o confi m he alidi y o his
model o ou da a, and o assess how ep esen a i e SNP
es ima es a e o ue unde lying biology a common sequenced
a ian s. The simula ion amewo k uses eal geno ype da a
om he UK10K conso ium o gene a e addi i e, polygenic
pheno ypes wi h a gi en he i abili y and hen pe o ms
he i abili y es ima ion wi h he a iance componen model
(see Me hods). Al hough he UK10K da a con ains a much
smalle se o indi iduals as he iCOGS da a (3,047 e sus 42,613
indi iduals, see Me hods), i con ains a ia ion om whole-
genome sequencing; his allows us o e alua e model pe o mance
by simula ion when es ic ing o SNPs geno yped on he iCOGS
pla o m. We ocused on he LNCaP: H3k27ac anno a ion (which
was mos significan in ou da a, see below) o e alua e he
mul iple componen models. O e housands o simula ions, we
confi med ha he a iance componen s app oach co ec ly
eco e ed he causal con ibu ion o ai om a gi en unc ional
ca ego y when causal a ian s we e yped (Supplemen a y
Table 1, see Me hods). Unde bo h null and en iched scena ios
he es ima es we e unbiased and s anda d e o s p ope ly
calib a ed (Supplemen a y Table 1). Fo common sequenced
a ian s no p esen on he iCOGS pla o m, ela i e es ima es o
noncoding en ichmen /deple ion we e conse a i e, wi h he
agged e ec s dis ibu ed ac oss he yped componen s
(Supplemen a y Table 2). De ia ions om he s anda d
a iance componen s model assump ions on he dis ibu ion o
e ec -sizes and ances y-specific e ec s in A ican Ame icans
yielded ei he well calib a ed o conse a i e es ima es o SNP
he i abili y in he ocal LNCaP: H3k27ac ca ego y (see Me hods,
Supplemen a y Tables 1–3).
Ou p ima y unc ional analyses ocus on he densely
geno yped iCOGS sample (21,678 cases and 20,935 con ols),
whose la ge sample size allowed o highly accu a e es ima es o
componen -specific h2
g. Al hough he iCOGS chip is cus om buil
o o e sample isk loci, i p o ides a b oad co e age o he
common a ia ion genome wide4. To showcase he powe o
he a iance componen s app oach, we es ima ed he o al SNP
ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979
2NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions
he i abili y o P Ca a 0.28 (s.e. 0.01) in he iCOGS da a (no
significan ly di e en om he o al SNP he i abili y es ima e o
0.26 (s.e. 0.05) in he BPC3 da a), a significan inc ease om he
a iance explained only by he known GWAS a ian s h2
GWAS
o
0.06 (s.e.m. 0.001) (see Me hods; Supplemen a y Table 4).
In e es ingly, he o al SNP he i abili y in he A ican Ame ican
sample, which was geno yped on a di e en pla o m han iCOGS
(see Me hods), was es ima ed a 0.32 (s.e. 0.06) indica ing a
simila agg ega e con ibu ion o common a ia ion o P Ca isk
ac oss he wo e hnici ies despi e highe o e all isk in A ican
Ame icans22 (Supplemen a y Table 4).
En ichmen a and ogen ecep o -binding si es in umou s.We
fi s ocused on SNPs localized in he ARBS: an epigene ic p ofile
causally implica ed in p os a e umo igensis. In con as o ypical
assays ha ocus on cell lines, he ARBS we e defined by ch oma in
immunop ecipi a ion and high- h oughpu sequencing (ChIP-seq)
di ec ly in p ima y human issue (se en no mal and 13 umou
specimens)18. We obse ed ha a ian s wi hin 5 kb o umou -
specific ARBS explained 17.0% o he genome-wide h2
g(s.e. 1.7%;
P¼2.6 1016 by Z- es ), whe eas he a ian s nea
no mal-specific ARBS explained 0.0% o he h2
g(s.e. 0.9%;
P¼0.11 by Z- es ) (Fig. 1). The di e ence be ween hese wo
g oups was highly significan and demons a es he impo ance o
assaying unc ional ma ks in bo h no mal and umou issues. We
no e ha he 5 kb ex ension may also include o he egula o y
a ian s nea he umou /no mal-specific ARBS (bu no
he i abili y om coding/un ansla ed egion (UTR)/p omo e
a ian s, which we e explici ly modelled, see Me hods). Smalle
flanking egions we e also in es iga ed bu did no include enough
ma ke s o he a iance componen s model o con e ge. We also
quan ified he p opo ion o SNP he i abili y explained di ec ly by
all ARBS a ian s (bo h no mal and umou wi hou 5 kb flanks) a
10.7% o h2
g; significan ly di e en om he SNP he i abili y o
ARBS a ian s assayed in p os a e adenoca cinoma cance cell line
(LNCaP; 3.2% o h2
g)(P¼4.4 107 o di e ence by Z- es )
(Fig. 1). This di e ence is pa ially explained by he e y low
numbe o SNPs wi hin cell line ARBS making hei agg ega e
con ibu ion small bu no empowe ing us o place a s ong bound
on he en ichmen . O e all, hese findings highligh he inc eased
complexi y o ARBS in a sample o issues as compa ed wi h he
single LNCaP cell line.
Iden ifica ion o unc ional ma ks ele an o P Ca isk. Nex ,
we looked o ma ks ha con ibu e o he he i abili y o P Ca
ac oss a b oad spec um o unc ional anno a ions wi hou p io
assump ions on ele ance o disease. We in es iga ed 544
epigene ic anno a ions spanning six majo classes (DHS;
H3k4me1; H3k4me3; H3k9ac; H3k27ac; and compu a ionally
p edic ed unc ional classes o ‘segmen a ions’23,24) a e aging 101
cell ypes pe class (see Me hods). A e accoun ing o mul iple
es ing, we iden ified 82 anno a ions ha exhibi ed s a is ically
significan de ia ions in SNP he i abili y om wha was expec ed
based on he p opo ion o he genome co e ed by ha pa icula
anno a ion (see Fig. 2 and Supplemen a y Da a).
We fi s ocused on 17 unc ional ma ks measu ed in he
p os a e, o which 14 we e s a is ically significan (Supplemen a y
Table 5). The single mos significan en ichmen was obse ed o
H3k27ac ma ks in LNCaP (P¼11032 by Z- es ), which
localized 22% o he o al h2
g o he 2.9% o geno yped SNPs
wi hin he anno a ion. This was ollowed by a ian s in DHS
ma ks in LNCaP (P¼21018 by Z- es ; 16.7% o h2
glocalized
in 3.1% o genome). The DHS anno a ions allowed us o compa e
es ima es ac oss h ee majo p os a e cell lines: LNCaP; no mal
p os a e epi helial (P EC); and immo alized p os a e epi helial
(RWPE1) (o e lapping by 25–50% wi h ARBS, Supplemen a y
Fig. 1). We obse ed he i abili y explained by LNCaP DHS o be
nominally significan ly highe han P EC (P¼0.01 by Z- es );
and bo h LNCaP and P EC o be significan ly highe han
RWPE1 (P¼1.5 109,P¼1.2 105, espec i ely, by Z- es )
(Fig. 3). Mo e b oadly, 10 ou o 16 DHS ma ks measu ed in
cance cell lines we e obse ed as significan , wi h colo ec al
cance as he nex mos significan cance (P¼6.0 1010 by
Z- es ; 9.4% o he i abili y localized in 2.0% o genome;
Supplemen a y Da a). H3k27ac in LNCaP emained he mos
significan ly en iched ma k ac oss all 544 anno a ions (p esen ed
in de ail in he Supplemen a y Da a). The mos deple ed
ca ego ies we e ep essed egions compu a ionally p edic ed by
Segway-ch omHMM in HepG2 cells (P¼1.3 1019 by Z- es ;
51.9% o h2
g om 74.3% o SNPs; Supplemen a y Da a), wi h
simila le els o deple ion in ep essed egions om o he cell
ypes. These egions a e ypically associa ed wi h dec eased gene
exp ession and ep essi e his one ma ks23–25, u he
emphasizing he impo ance o ac i e egula ion.
As H3k27ac ypically ma ks ac i e enhance s, we u he
e alua ed a ian s wi h espec o hei enhance o
‘supe ’-enhance s a us (la ge clus e s o enhance s ha a e
en iched o genes in ol ed in cell iden i y26) (see Me hods). We
did no obse e di e ences in a e age he i abili y explained by
SNPs wi hin he wo ma ks ac oss 49 cell lines (see Me hods),
wi h an a e age o 1.51 (1.47)- old inc ease o e andom SNPs o
NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979 ARTICLE
0.00 0.05 0.10 0.15 0.20
ARBS LNCaP cell line
ARBS p os a e issue
no mal+ umou
0.05 0.10 0.15 0.200.00
ARBS p os a e issue
umou -only ±5kb
ARBS p os a e issue
no mal-only ±5kb
%SNP-he i abili y
%SNP-he i abili y
a
b
Figu e 1 | Func ional pa i ioning o a ian s wi hin ARBS o P Ca. Ba s g aphs de ailing %SNP he i abili y es ima es om wo models o P Ca ele an
unc ional anno a ions. (a) Join compa ison o a ian s wi hin 5 kb o umou -only and no mal-only egions in he ARBS in p os a e issue (P¼2.1 1019
o di e ence by Z- es ). (b) Es ima es om ARBS in p os a e issue (no longe using a 5 kb flank) and ARBS in LNCaP cell lines7(P¼4.4 107 o
di e ence). The null ð%h
2
g¼%SNPsÞis labelled by he dashed lines. E o ba s show analy ical s anda d e o o es ima e.
NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions 3
enhance s (supe enhance s) (Fig. 4). Su p isingly, we obse ed
an indi idually significan di e ence only in LNCaP, wi h 4.9
(1.7)- old en ichmen a enhance s (supe enhance s), in con as
o p e ious hypo heses26 (Fig. 4).
Genomic unc ional a las o p os a e cance SNP he i abili y.
Al hough he esul s abo e showcase he powe o he a iance
componen app oach in finding epigene ic ma ks ele an o
P Ca, such ma ks o en o e lap making he causal ma k
di ficul o iden i y (Supplemen a y Fig. 1). To accoun o he
co ela ion among ma ks we g ouped he 82 ma ginally
significan anno a ions in o 15 biologically ele an , non-o e -
lapping g oups o ganized by ma k and cell line, and pa i ioned
h2
gac oss all g oups in a join model (see Me hods, Table 1, Fig. 5
and Supplemen a y Table 6). Fi e componen s we e nominally
significan in he join model a Po0.05; ou o he fi e
componen s h ee emained significan a e accoun ing o 15
es s: H3k27ac ma ks in LNCaP (P¼2.5 1020 by Z- es );
DHS ma ks in o he cance cell ypes (P¼3.9 105by
Z- es ); and ep essed segmen a ions (P¼2.1 1020 by Z- es ).
To u he efine ou model, we es ic ed o he significan
anno a ions (and he backg ound componen s accoun ing o LD
o coding egions) and e-e alua ed hem join ly, e e ed o as
he ‘selec ed’ model. This selec ed model localized 51.0% o he h2
g
wi hin 12.1% o SNPs (LNCaP: H3K27ac þARBS þDHS cance ),
whe eas coding egions only explained 3.3% (s.e. 1.4%) o h2
g
wi hin 1.8% o SNPs (Supplemen a y Table 7). The localiza ion
was e en s onge wi h impu ed da a, whe e 86% o he h2
gwas
localized o 8.6% o SNPs (Table 1 and Supplemen a y Tables 8
and 9). Es ima es om impu ed ma ke s we e mo e ep esen a-
i e o unde lying en ichmen in ou simula ions (see Me hods,
Supplemen a y Table 2) bu may include he e ec s o nea by
ma ke s12 and so we conside hem as an uppe bound. None o
he es ima es changed significan ly a e adjus ing o known
GWAS associa ions2(79 o which we e yped in his da a),
unde sco ing he polygenic na u e o his e ec .
Ha ing in e ed he selec ed model, we e-analysed each o he
82 ma ginally significan ca ego ies join ly wi h he selec ed
model (see Me hods). Only h ee ma ks emained significan : wo
H3k27ac anno a ions in he colon c yp and one H3k27ac
anno a ion in panc eas (Supplemen a y Da a). This implies ha
he ma ginal en ichmen o he 82 anno a ions was p ima ily
d i en by he o e lap wi h unc ional ma ks in he selec ed
model. Fo example, he H3K4me1 ma k in penis o eskin
ke a inocy es ha was p e iously highly significan (24.6% h2
g,
P¼3.0 1016 by Z- es , Fig. 1) was no longe en iched a e
condi ioning on he selec ed model (7.1% h2
g,P¼0.29 by Z- es ,
Supplemen a y Da a). The educ ion o a small numbe o
ca ego ies in he selec ed model wi h limi ed loss in signal u he
emphasizes he ex en o which he selec ed model has localized
he unc ional sou ces o en ichmen . Focusing on he wo mos
en iched ca ego ies in he selec ed model, we ound ha SNPs
p esen in bo h he p os a e issue ARBS and LNCaP H3k27ac
ma ks yielded significan ly highe a e age he i abili y pe SNP
han ei he ma k indi idually (Supplemen a y Table 10).
ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979
0.00 0.05 0.10 0.15
0.00
0.05
0.10
0.15
0.20
0.25
0.00
0.05
0.10
0.15
0.20
0.25
0.00
0.05
0.10
0.15
0.20
0.25
0.00
0.05
0.10
0.15
0.20
0.25
0.0
0.2
0.4
0.6
0.8
1.0
0.00
0.05
0.10
0.15
0.20
0.25
DHS
% SNP
% SNP-he i abili y
LnCAP
Mamma y epi helium
0.00 0.05 0.10 0.15
H3K27ac
% SNP
% SNP-he i abili y
LNCaP+DHT
LNCaP
0.00 0.05 0.10 0.15
H3K4me1
% SNP
% SNP-he i abili y
Penis o eskin
Rec al mucosa
0.00 0.05 0.10 0.15
H3K4me3
% SNP
% SNP-he i abili y
HUES6 cell line
Penis o eskin
0.00 0.05 0.10 0.15
H3K9ac
% SNP
% SNP-he i abili y
Colonic mucosa
Rec al mucosa
0.0 0.2 0.4 0.6 0.8 1.0
O he
% SNP
% SNP-he i abili y
Rep essed (HepG2)
ARBS
Figu e 2 | Func ional pa i ioning o he i abili y ac oss six main epigene ic classes. Each poin co esponds o an es ima e o % SNP he i abili y (yaxis)
om SNPs wi hin a cell- ype-specific unc ional anno a ion e sus anno a ion size (%SNPs, xaxis). O e all, 544 anno a ions we e es ed, and ed poin s
indica e significan de ia ions om he null o %h
2
gequal o %SNPs a e accoun ing o all es s. The wo mos significan anno a ions in each class a e
shown wi h iangle/c oss, espec i ely, and labelled in bo om igh (see Supplemen a y Da a o all anno a ions).
4NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions
In con as , he a ian s specific o ARBS o H3k27ac we e
compa able in SNP he i abili y.
Replica ion o genomic unc ional a las ac oss ances ies.We
e alua ed eplica ion o ou model using wo sepa a e
genome-wide SNP da a se s o P Ca, one o Eu opean ances y
(BPC3; 6,953 samples) and one o A ican ances y (AAPC; 9,522
samples) o P Ca (see Me hods). To accoun o he smalle
sample size, we ocused on he eigh -componen selec ed model,
only e aining significan componen s and h ee coding-p oximal
classes (coding, UTR, p omo e )12. Because o pla o m
di e ences be ween he popula ions, we used pos -QC impu ed
a ian s in each da a se , which a e mos eflec i e o unde lying
en ichmen in ou simula ions (see Me hods). We eplica ed he
significan de ia ion in h2
ga H3k27ac and he ep essed loci
ac oss bo h BPC3 and AAPC (Supplemen a y Tables 11 and 12).
Howe e , cance DHS was only significan in he BPC3 da a and
ARBS no significan in ei he ( hough he es ima es we e no
significan ly di e en om he iCOGS es ima e). The en ichmen
did no change a e es ic ing o e y high-quali y impu ed
ma ke s (Supplemen a y Table 13). Al hough he ela i ely small
alida ion sample size did no p o ide enough powe o es
di e ences be ween he ances ies, he mean SNP he i abili y o
a ian s wi hin each ma k we e ema kably simila ( ¼0.90
be ween AAPC and BPC3 ac oss eigh componen s), sugges ing a
simila pa e n o agg ega e con ibu ion o isk coming om
common a ian s ma ked by epigene ic classes ac oss Eu opean
and A ican Ame ican ances ies ( hough indi idual isk a ian s
hemsel es may di e ).
H3k27ac ma k in LNCaP is specific o P Ca. As a nega i e
con ol, we e alua ed he selec ed model wi h impu ed
SNPs ac oss 11 common non-cance diseases om he Wellcome
T us Case Con ol Conso ium (WTCCC) (see Me hods,
Supplemen a y Table 14) whe e we obse ed wo main
di e ences: he LNCaP H3k27ac anno a ion was no longe
significan ly en iched (1.1% h2
gwi h 2.6% o SNPs); and he
ep essed egions we e much less deple ed om he null (28.1%
NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979 ARTICLE
PREC
PREC
RWPE1
LNCAP
LNCAP
RWPE1
2.0% SNPs
6.0% h
2g
0.8% SNPs
4.1% h
2g
2.3% SNPs
11.0% h
2g
2.1% SNPs
9.6% h
2g
0.7% SNPs
0.6% h
2g
0.3% SNPs
1.1% h
2g
2.7% SNPs
14.1% h
2g
0.3% SNPs
1.2% h
2g
(P=0.2)
0.7% SNPs
0.7% h
2g
(P=7×10
–3
)(P=6×10
–4
)(P=2×10
–7
)
(P=1×10
–6
)(P=0.9) (P=0.2)
(P=4×10
–8
)(P=0.9)
a
b
c
Figu e 3 | Pai wise analysis o DHS ma ks in h ee p os a e cell ypes.
Join model om all pai s o DHS ma ks shown o : cance cell line
(LNCAP); no mal p os a e epi helial (PREC); and immo alized p os a e
epi helial (RWPE1). Ci cle size co esponds o % SNPs, wi h % SNP
he i abili y and significance labelled. P alue was compu ed o di e ence
be ween %h
2
gand %SNP, wi h bold ep esen ing significance a e
co ec ing o nine es s. The obse ed end is LNCAP4PREC4RWPE1:
(a)%h
2
gin LNCaP DHS was nominally significan ly highe han P EC
(P¼0.01); and %h
2
gin LNCaP and P EC was significan ly highe han
RWPE1 (b,c;P¼1.5 109,P¼1.2 105, espec i ely). All P alues
compu ed by Z- es using %h
2
ges ima e and analy ical s anda d e o .
CD4 memo y p ima y
B ain angula gy us
B ain hippocampus middle
B ain in e io empo al lobe
K562
B ain cingula e Gy us
MM1S
S omach smoo h muscle
Ad enal gland
B ain an e io cauda e
U87
B ain hippocampus middle
Le en icle
Ao a
Skele al muscle
Fe al hymus
Adipose nuclei
Fe al muscle
CD56
T−cell leukemia
CD14
CD34 p ima y
Lung ib oblas
CD20
Lung
CD4
De mal ib oblas s
Os eoblas s
HUVEC
As ocy es
CD8 p ima y
Duodenum smoo h muscle
Ju ka
Esophagus
HeLa
HSMM ube
Panc ea ic cance
B eas cance
Lung cance
Gas ic
Fe al in es ine la ge
Skele al muscle myoblas
Mamma y epi helial
Fe al in es ine
Colon cance
Colon c yp 1
Colon c yp 2
Panc eas
LnCAP
Enhance
%SNP-he i abili
y
/ %SNP
0246
Supe -enhance
%SNP-he i abili
y
/ %SNP
0246
Figu e 4 | Compa ison o enhance s and supe enhance s ac oss 49 cell
ypes. Each ba ep esen s he %SNP he i abili y %h
2
g/ %SNP o
enhance s (le ) and supe enhance s ( igh ) om a gi en cell ype es ed
ma ginally. Red indica es significan di e ence om 1.0 (no en ichmen )
a e accoun ing o 49 es s. Enhance LNCAP is mos significan , wi h
o he cance s also appea ing significan and non-cance issues leas
significan . E o ba s show analy ical s.e. o es ima e.
NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions 5
h2
gwi h 87.8% o SNPs) compa ed wi h he 0.3% o h2
gobse ed in
iCOGS impu ed da a (P¼2.2 104 o di e ence by Z- es ).
In e es ingly, al hough ARBS we e significan ly en iched in all 11
ai s, he en ichmen was no longe significan a e excluding
au oimmune ai s. O e all, hese di e ences indica e ha he
LNCaP H3k27ac ma k is uniquely in o ma i e o P Ca, whe eas
a ian s nea he ARBS and DHS cance elemen s (which o e lap
o he DHS anno a ions by 56%; Supplemen a y Fig. 2) may play a
gene ally impo an ole ac oss o he common diseases12.
Genomic unc ional a las imp o es polygenic isk p edic ion.
To alida e ou SNP he i abili y genomic a las, we compa ed he
accu acy o p edic ing case/con ol s a us om gene ic da a wi h
o wi hou he unc ional a las. We e alua ed h ee dis inc
p edic ion models in he iCOGS sample: (i) a gene ic isk sco e
(GRS) om he genome-wide significan SNPs; (ii) he single
bes linea unbiased p edic o (BLUP) using a single a iance
componen om all SNPs; and (iii) he weigh ed sum o
indi idual BLUPs om each epigene ic ca ego y in he selec ed
model (mul i-BLUP; see Me hods). E alua ed by c oss- alida ion,
he GRS yielded an R2¼0.029 wi h ue pheno ype, whe eas he
single BLUP yielded an R2¼0.065 and he mul i-BLUP had an
R2¼0.071 (Supplemen a y Table 15). In a join model wi h all
h ee p edic o s, he mul i-BLUP was highly significan (P¼5.3
1031 om mul iple eg ession). When we cons uc ed he
GRS om SNPs ecen ly disco e ed in a much la ge P Ca
GWAS ( e . 2), he esul ing p edic ion R2inc eased o 0.084.
Howe e , including he single BLUP o he mul i-BLUP as an
addi ional p edic o s ill inc eased he p edic ion R2 o 0.096
(join P¼6.7 104 om mul iple eg ession) and 0.098 (join
P¼1.3 1023 om mul iple eg ession), espec i ely
(Supplemen a y Table 15). The consis en s a is ical significance
and inc eased p edic ion accu acy confi ms he alidi y o he
selec ed model in his da a and in la ge GWAS.
Discussion
Using la ge-scale geno ype da a om o e 59,089 men o
Eu opean and A ican Ame ican ances ies join ly wi h epigene ic
anno a ions, we iden ified highly significan di e ences in SNP
he i abili y ðh2
gÞo P Ca ac oss a ian s om di e en epigene ic
classes, issue ypes and cell lines. Focusing on ma ks measu ed
in p os a e, we obse ed significan ly highe h2
ga ound
umou -specific ARBS; ARBS measu ed in p ima y issue ela i e
o cell line; and DHS measu ed in P Ca cell line ela i e o
p os a e epi helial cell line. The en ichmen a umou -specific
ARBS was consis en wi h ecen findings showing ha hese si es
we e en iched o nea by genes highly exp essed in umou s18.
These analyses a e comp ehensi e and co e mos commonly
s udied p os a e cell lines excep o e eb al cance o
he p os a e, which we e no well ep esen ed in he ENCODE/
ROADMAP. A sea ch ac oss 544 di e se unc ional anno a ions
es ic ed mos o he h2
g o a small ac ion o he genome
ma ked by p os a e egula o y elemen s. Consis en wi h p e ious
findings in common disease, unc ionally ep essed egions we e
significan ly deple ed in he i abili y, highligh ing he ole o ac i e
egula ion in P Ca suscep ibili y. Subsequen model selec ion
localized he en ichmen om 82 indi idually significan
anno a ions o six ha emained significan in a join model. In
pa icula , he abundance o en ichmen in H3k27ac ma ks
(ac i e enhance s) ela i e o H3k4me1/H3k4me3 (poised
enhance s/p omo e s) unde sco es hei ole in P Ca, hough
u he en ichmen in supe enhance s was no obse ed.
The en ichmen wi hin LNCaP: H3K27ac and deple ion a
ep essed egions was eplica ed ac oss di e en ances ies and
yielded significan imp o emen s in polygenic isk p edic ion.
Wi h mos GWAS associa ions alling ou side coding egions,
ou analyses o e an impo an esou ce o p io i izing po en ial
loci and ocusing u u e s udies on he mos he i able genomic
egions27. The ma ginal analyses p o ide a anking o 544
common unc ional assays, while he selec ed model localizes
he i abili y o only hose unc ional classes ha a e independen ly
en iched. Eme ging unc ional ca ego ies may u he efine his
signal o e eal o he ele an epigene ic ma ks, hough li le
en ichmen beyond he selec ed model was obse ed in he
comp ehensi e sampling o unc ional da a analysed he e. In
gene al, he a iance componen model o e s an oppo uni y o
e alua e biological hypo heses in silico and wi hou s ic ly elying
on indi idually significan SNPs. Howe e , as wi h any analysis o
a ay-based da a, he h2
ges ima es will no include he
con ibu ion o SNPs ha a e un yped o poo ly agged, such
ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979
Table 1 | Pa i ioning o he i abili y ac oss unc ional classes in p os a e cance .
Func ional ca ego y %SNPs Full Model Selec ed model
iCOGS geno yped iCOGS impu ed BPC3 impu ed AAPC impu ed
%h
2
gs.e.m. %h
2
gs.e.m. %h
2
gs.e.m. %h
2
gs.e.m.
Coding 1.8 3.0 1.3 0.9 2.9 0.2 10.1 3.3 11.1
UTR 1.9 1.6 1.4 3.0 3.1 21.0 11.3 5.9 11.2
P omo e 3.4 *7.8 1.8 8.9 4.1 0.0 12.7 0.0 14.7
LNCaP: H3k27ac 3.2 **22.3 2.1 **27.0 3.8 *30.3 12.1 *28.9 12.7
ARBS 1.0 *3.3 1.1 *9.1 3.3 1.1 12.1 15.2 12.1
LNCaP: FOXA1 1.5 1.5 1.3
LNCaP: H3k4me1 2.0 1.3 1.4
LNCaP: DHS 2.9 5.4 1.6
DHS p os a e 1.8 2.6 1.4
DHS cance 4.7 **14.1 2.3 **49.6 6.3 *47.4 21.4 46.6 22.4
H3k4me1 (o he ) 16.3 19.6 3.5
H3k27ac (o he ) 7.3 4.1 2.4
DHS (o he ) 1.8 0.2 1.3
ep essed 48.7 **11.0 4.1 **0.3 7.0 **0.0 23.8 **0.0 24.5
all o he 1.7 0.7 1.2 0.2 2.7 0.0 9.2 0.0 7.6
ARBS, and ogen ecep o -binding si es; DHS, DNase I hype sensi i i y si es; SNP, single-nucleo ide polymo phism; UTR, un ansla ed egion.
Full model deno es a 15- a iance componen s model while ‘selec ed’ model deno es a model es ic ed o he fi e componen s a aining significance in he ‘ ull’ model (and h ee componen s o
backg ound). * (**) deno es significan de ia ion a Po0.05 (Po0.05/15) o ac ion o SNP he i abili y ð%h
2
gÞ om null model o %h
2
g¼%SNPs (by Z- es ; see Supplemen a y Table 6 o
P alues).
6NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions
as a e a ian s o o he con ibu o s o he missing he i abili y.
Fu u e analyses o whole-genome sequencing, addi ional
unc ional anno a ions, and la ge sample sizes can yield
impo an insigh s in o unc ional mechanisms ha a e s ill no
localized. O e all, ou esul s sugges simila pa e ns o
unc ional en ichmen ac oss men o Eu opean and A ican
Ame ican ances y and highligh ARBS, H3k27ac ma ks in
LNCaP cell lines and DHS in cance cell lines o ollow-up
s udies o P Ca isk.
Me hods
Epigene ic anno a ions.Sample collec ion and p ocessing o unc ional anno-
a ions was made publically a ailable by he ENCODE/ROADMAP conso ia28.
DHS, H3k4me1, H3k4me3, H3k9ac anno a ions and genome segmen a ions20,29,
enhance s and supe enhance s26 and P Ca-specific anno a ions7,18 we e assay and
p ocessed as de ailed in he o iginal s udies. Tumou -only and no mal-only ARBS
we e defined in se en no mal and 13 umou specimens in he o iginal s udy18.
All anno a ions cu a ed o his pape (ENCODE/ROADMAP; Pome an z e al.;
and Hazele e al.) a e a ailable a h ps://da a.b oadins i u e.o g/alkesg oup/
ANNOTATIONS/PRCA/. The ull lis o indi idual anno a ions wi h web-links o
he co esponding bounda y defini ions is p o ided in he Supplemen a y Da a.
Some unc ional ma ks a e lis ed mul iple imes due o mul iple independen assays
o labo a o y p o ocols.
ARBS ChIP-seq in human issue specimens.The ARBS assay was pe o med as
desc ibed in REF ( e . 18) and summa ized he e. Fou een subjec s o Eu opean
Ame ican ances y we e selec ed o ChIP analysis. Thei ch oma in was incuba ed
o e nigh wi h 6 mg an ibody AR (N-20, San a C uz Bio echnology, Dallas, TX)
NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979 ARTICLE
iCOGS main model
Coding
P omo e *
LnCAP H3K27ac**
ARBS*
LnCaP FOXA1
LnCaP H3K4me1
LnCaP DHS
DHS p os a e
DHS cance **
H3K4me1 (o he )
H3K27ac (o he )
DHS (o he )
D
Rep essed**
O he
× 0.5
× 5
× 10
iCOGS selec ed model
Coding
UTR
P omo e *
LnCAP H3K27ac**
ARBS*
DHS cance **
Rep essed**
O he
× 0.5
× 5
× 10
AAPC selec ed model
Coding
UTR
P omo e
LnCAP H3K27ac*
ARBS
DHS cance
Rep essed**
O he
× 0.5
× 5
× 10
× 25
BPC3 selec ed model
UTR
LnCAP H3K27ac*
DHS cance *
Rep essed**
O he
× 0.5
× 5
× 10
× 25
ab
cd
Figu e 5 | Pa i ioning o he i abili y ac oss unc ional classes in p os a e cance . Visual ep esen a ion o he i abili y en ichmen in h ee s udies
a,b: iCOGS; c: AAPC; d: BPC3 (shown nume ically in Table 1). Each subplo co esponds o an analysis o he lis ed join model, wi h colou ed slices
ep esen ing he unc ional anno a ions e alua ed. Volume o each in e io (ligh colou ed) pie-cha slice ep esen s he %SNP o he unc ional
anno a ion, which is equal o he expec ed %h
2
gunde he null o no en ichmen . Volume o each shaded pie-cha slice ep esen s he ac ual %h
2
g
in e ed by he model. Slices ex ending ou side/inside he middle pie co espond o en ichmen /deple ion in SNP he i abili y, as indica ed by he do ed
lines. Colou coding is consis en ac oss all subpanels. * (**) deno es significan de ia ion a Po0.05 (Po0.05/15) o ac ion o SNP he i abili y (%h
2
g
om null model o %h
2
g¼%SNPs by Z- es ; see Supplemen a y Table 6 o P alues).
NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions 7
bound o p o ein A and p o ein G beads (Li e Technologies, Ca lsbad, CA). A
ac ion o he sample was no exposed o an ibody o be used as con ol (inpu ).
The samples we e de-c osslinked, ea ed wi h RNase and p o einase K, and DNA
was ex ac ed. The samples we e hen e-shea ed o 100–300 base pai s using he
Co a is ul a-sonica o , and concen a ions o he ChIP DNA we e quan ified by
Qubi Fluo ome e (Li e Technologies). DNA sequencing lib a ies we e p epa ed
using he Th uPLEX-FD P ep Ki (Rubicon Genomics, Ann A bo , MI). Lib a ies
we e sequenced using 50-base pai eads on he Illumina pla o m (Illumina, San
Diego, CA) a Dana-Fa be Cance Ins i u e. AR binding si es we e gene a ed using
Model-Based Analysis o ChIP-seq 2 (MACS2), wi h a q alue ( alse disco e y a e,
FDR) h eshold o 0.01.
The 13 umou s used in his s udy we e and ogen dependen and no exposed
o and ogen dep i a ion he apies. All o he umou s we e specimens ob ained
om adical p os a ec omies, de i ed om men wi h ea ly s age disease. These
samples we e no selec ed based on any specific ea u es; he e o e, we would
expec ha he dis ibu ion o isk a ian s would be simila o a andom sampling
o P Ca cases. La ge-scale gene ic su eys ha e shown ha soma ically acqui ed
al e a ions in p ima y localized p os a e umou s ( he ype o umou e alua ed in
his s udy) a e in equen . Based on hese p e ious esul s, we belie e ha
soma ically acqui ed gene ic e en s in egions ela ed o and ogen biology a e no
common and, he e o e, do no influence ou esul s.
Pa ien ma e ial.In o med consen was ob ained om all subjec s and all s udies
we e app o ed by local Resea ch E hics Commi ees and/o Ins i u ional Re iew
Boa ds.
Da a quali y con ol.Quali y con ol is c ucial o accu a e he i abili y es ima ion,
whe e many small a i ac s can add up o la ge biases. All da a se s wen h ough a
s ingen QC p ocess wi h he ollowing exclusion c i e ia: mino allele equency
(MAF)o1%; ac ion o missing/uncalled SNPs45%; Ha dy–Weinbe g
equilib ium P alueo0.01; case–con ol missingness P alueo0.05; impu a ion
INFO sco e40.30. In addi ion, close ela i es we e p uned such ha no pai o
indi iduals had gene ic ela edness (GRM) coe ficien s40.05. The op 10 p incipal
componen s and a coded s udy label we e always included as fixed-e ec s. All
analysed samples, cases and con ols, we e males.
iCOGS da a.The iCOGS conso ium geno yped balanced cases and con ols on a
cus om a ge ed a ay4. A e quali y con ol, 42,613 samples and 153,621
geno yped SNPs emained. Impu a ion was pe o med o he 1000 Genomes
e e ence panel using HAPI-UR ( e . 30) o phasing and IMPUTE2 ( e . 31) o
impu a ion. O e all, 1,910,827 impu ed and geno yped SNPs passed QC. Because
o compu a ional es ic ions, he he i abili y es ima ion was ca ied ou in wo
equally sized hal es o he ICOGS, wi h o al e ec s compu ed by in e se- a iance
me a-analysis. We pa i ioned he geno yped SNP he i abili y by MAF bu
obse ed no end and only sligh en ichmen o % h2
ga high- equency a ian s
(Supplemen a y Table 16).
BPC3 da a.The Na ional Cance Ins i u e B eas & P os a e Cance Coho
Conso ium (BPC3) conso ium geno yped indi iduals on he Illumina Human-
Hap610 quad a ay32. A e quali y con ol, 6,953 samples and 4,004,229 geno yped
and impu ed SNPs emained. Age was a ailable o all samples and addi ionally
included as a co a ia e.
AAPC da a.The AAPC conso ium geno yped indi iduals o A ican ances y on
he Illumina Human1M a ay2,33,34. A e quali y con ol, 9,522 samples and
10,468,389 geno yped and impu ed SNPs emained.
WTCCC da a.The Wellcome T us Case Con ol Conso ium Geno yping
geno yped cases o 11 ai s as well as sha ed con ols on mul iple Illumina and
A yme ix a ays35–37. The pheno ypes analysed he e we e ankylosing spondyli is
(AS); bipola diso de (BD); co ona y a e y disease (CAD); C ohn’s disease (CD);
hype ension (HT); mul iple scle osis (MS); heuma oid a h i is (RA);
schizoph enia (SP); ype 1 diabe es (T1D); ype 2 diabe es (T2D); and ulce a i e
coli is (UC). A e quali y con ol, a o al o 47,053 samples and 4–5 million
geno yped and impu ed SNPs emained. Repo ed h2
g alues we e es ima ed o
each pheno ype sepa a ely and me a-analysed using in e se- a iance weigh ing.
UK10K da a.The UK10K whole-genome sequence da a om ALSPAC and
TWINSUK (h p://www.uk10k.o g) was used only o simula ion, and so s ingen
quali y con ol was no applied. A e ela edness fil e ing, 3,047 samples and
15,691,225 non-single on a ian s we e e ained.
He i abili y es ima ion o indi idual anno a ions.We es ima ed he SNP
he i abili y ðh2
gÞcap u ed by unc ional ca ego ies in a join a iance
componen model using GCTA as desc ibed in REF ( e . 20). B iefly, his model
assumes he pheno ype is d awn om a mul i a ia e no mal dis ibu ion wi h
a iance-co a iance modelled by componen s compu ed om he SNPs and a
no mal esidual. Fo each unc ional ca ego y ( o example, DHS) i¼1..Mwhe e
Mis he o al numbe o ca ego ies in he model, a GRM ac oss all pai s o
indi iduals is compu ed es ic ing o SNPs wi hin he unc ional ca ego y.
Va iance componen s o all GRMs in he model a e hen fi ed using REML as
implemen ed in GCTA o es ima e a a iance pa ame e s2
i
used o compu e
%h
2
i¼s2
i=PM
j¼1s2
j.Theh2
ico esponds o he ac ion o ai a iance ha
can be explained by he BLUP es ic ed o SNPs in he co esponding unc ional
ca ego y (o anno a ion). Fo a gi en unc ional anno a ion, SNPs we e ca ego ized
in o a hie a chy o se en non-o e lapping componen s: (1) coding; (2) UTR; (3)
p omo e ( unc ional anno a ion o in e es ); (4) DHS; (5) in on; and (6)
in e genic. SNPs belonging o mul iple ca ego ies we e pa i ioned explici ly in o
he fi s ca ego y in his lis . The coding and coding-p oximal componen s we e
included o ensu e ha he anno a ion he i abili y was no infla ed by SNPs ha
we e in high LD wi h coding a ia ion. A gene ic ela edness ma ix was compu ed
o each componen by fi s s anda dizing he co esponding SNPs and hen
compu ing a SNP co a iance o e all pai s o samples. Componen -specific s2and
e o s we e fi ed i e a i ely using he A e age In o ma ion algo i hm38.The
analy ical s anda d e o o %h
2
iwas es ima ed by ans o ming he GCTA-
in e ed s2
iand e o co a iance ma ix using he del a me hod. As in REF ( e . 20)
s a is ical significance was e alua ed by compa ing he %h
2
gexplained by he
ca ego y and i ’s s anda d e o o he %SNPs in he ca ego y using a Z- es
(compa ing nes ed models using a likelihood a io es yielded simila esul s).
To al h2
ges ima es we e compu ed as h2
g¼PM
j¼1s2
j=PMþ1
j¼1s2
ja e ans o ming
o he liabili y scale assuming a p e alence o 0.14 and using he s udy-specific case/
con ol a io.
Hie a chical join models.Fo specific models o in e es , we ex ended he
indi idual anno a ion model desc ibed abo e o es in e sec ing and non-in e -
sec ing componen s. This allowed us o e alua e p ecisely which sub-anno a ions o
o e lapping componen s we e likely o be causal. Fo he umou /no mal model,
we expanded each umou /no mal ma k by 5 kb in bo h di ec ions om he cen e
o cap u e nea by genes and o he egula o y egions so ha umou (no mal)
co e ed 3.3% (1.4%) o he SNPs, espec i ely. We es ima ed h2
g om he join
hie a chical model: (1) coding; (2) UTR; (3) p omo e ; (4) no mal-only; (5)
umou -only; (6) DHS; (7) in on; and (8) O he . When compa ing ARBS om
issue and ARBS LNCaP om cell line, only 59 SNPs (0.03%) o e lapped be ween
he wo ca ego ies, and so we es ed wo sepa a e models: (1) coding; (2) UTR; (3)
p omo e ; (4) (ARBS issue/ARBS LNCaP); (5) DHS; (6) in on; and (7) o he . Fo
compa isons be ween LNCAP, PREC and RWPE1 using DHS we es ed each pai
o cell lines using he join model: (1) coding; (2) UTR; (3) p omo e ; (4) DHS
pa icula o one cell line; (5) DHS common o bo h cell lines; (6) DHS pa icula
o o he cell line; (7) DHS o he cell lines; (8) In on; and (9) O he . Fo com-
pa isons be ween enhance s and supe enhance s, we used he 86 cell- ype-specific
anno a ions om REF ( e . 26), es ing each enhance o supe enhance sepa a ely
in he ollowing join model: (1) coding; (2) UTR; (3) p omo e , (4) (enhance /
supe enhance o cell- ype o in e es ); (5) DHS; (6) in on; (7) o he . O hese, 49
cell ypes yielded model con e gence o bo h he enhance and co esponding
supe enhance and we e used o es ima e means and co ela ion. The o de and
g ouping o ma ginally significan anno a ions in o epigene ic ma k and cell ype
( o example, in Table 1) a e lis ed in he Supplemen a y Da a. Fo each o he 82
indi idually significan anno a ions, we e-e alua ed hem join ly wi h he selec ed
model in he ollowing hie a chical join model: (1) coding; (2) UTR; (3) p omo e ;
(4) LNCaP:H3k27ac; (5) ARBS; (6) DHS cance ; (7) ( unc ional anno a ion o
in e es ); (8) DHS; (9) in on; and (10) o he . Only unc ional anno a ions ha
con e ged we e epo ed in he Supplemen a y Da a.
Accu acy o h2
ges ima es om yped a ian s in simula ions.The a iance
componen model has p e iously been shown o yield obus es ima es unde he
assump ion ha causal a ian s a e yped and uni o mly sampled om a gi en
unc ional ca ego y13,20,21. He e, we pe o m simula ions using he UK10K whole-
genome sequence da a o confi m he alidi y o his model o ou anno a ions,
and o assess how ep esen a i e SNP es ima es a e o ue unde lying biology a
common sequenced a ian s. O e all, he simula ions in ol e using eal ma ke s o
gene a e addi i e, polygenic pheno ypes wi h a gi en he i abili y and hen
es ima ing he he i abili y wi h he a iance componen model. We e alua ed he
UK10K da a o h ee ypes o SNPs: (i) common sequenced a ian s (7,534,538
SNPs); (ii) UK10K SNPs yped by he iCOGS pla o m (178,509; 95% o iCOGS
SNPs); and (iii) UK10K SNPs yped and impu ed by he iCOGS pla o m
(1,655,723; 87% o he iCOGS impu ed SNPs). We ocused on he LNCaP:H3k27ac
anno a ion (which was mos significan in ou da a) o e alua e he main join
model. All pheno ypes we e simula ed by d awing 5,000 causal a ian s andomly
om he specified ca ego ies and sampling causal e ec -sizes om a no mal
dis ibu ion such ha SNPs ei he explain equal a iance ( he model assump ion)
o a iance in p opo ion o hei MAF. The pheno ype was hen gene a ed as he
do p oduc o geno ype and e ec -size wi h andom noise added o fix he i abili y
a 50%. Pheno ypes we e simula ed housands o imes un il he s anda d e o
o e simula ions was low enough o e alua e unbiasedness.
ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979
8NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions
We confi med ha es ima es o h2
g om a polygenic ai we e accu a e unde
he model whe e causal a ian s a e yped (Supplemen a y Table S1). Unde he
null, he LNCaP H3k27ac componen is expec ed o explain 3.22% o he
SNP he i abili y, and he model es ima ed 3.50% (0.22%) and 3.68% (0.21%)
unde a low- equency and high- equency disease a chi ec u e, espec i ely
(Supplemen a y Table S1). None o he es ima es we e significan ly di e en om
he u h gi en he numbe o componen s es ed. Unde a scena io whe e LNCaP
H3k27ac explains 50% o he h2
g, he model es ima ed 51.13% (0.40%) and 46.98%
(0.35%) unde a low- equency and high- equency disease a chi ec u e,
espec i ely (Supplemen a y Table S1). Al hough he high- equency a chi ec u e
(whe e common a ian s explain mo e a iance in ai han a e a ian s)
ep esen s a subs an ial model misspecifica ion, ou simula ions show ha his does
no in oduce subs an ial bias and is likely o sligh ly unde es ima e he SNP
he i abili y a he ocal ch oma in ma k. In all cases, he empi ical s anda d
de ia ion o e 500 simula ions was simila o he a e age analy ical s.e.m.
compu ed by GCTA (REML algo i hm), hus showing ha ha analy ical s anda d
e o is well calib a ed (Supplemen a y Table S1). We no e ha he s anda d e o
is in e sely ela ed o he sample size39,40, and is he e o e much highe in hese
simula ions han in he iCOGS da a which is 14- old la ge .
Las ly, we pe o med he eal da a pa i ioned analysis in subse s o indi iduals
o e alua e biasedness and powe o de ec significan en ichmen . We confi med
ha no significan di e ences we e obse ed be ween es ima es om he en i e
s udy compa ed wi h hose a e aged ac oss subse s o he s udy (Supplemen a y
Fig. 3). As such, we can confiden ly epo es ima es and bounds on he en ichmen
obse ed in he en i e s udy ha will hold o la ge s udies. Fu he mo e, all bu
one o he significan componen s om he main model emained significan in
smalle samples (ARBS), making i unlikely ha hey we e a ec ed by winne ’s
cu se. Recen wo k has quan ified he heo e ical ela ionship be ween es ima ion
e o and e ec i e sample size o indi idual componen s39,40.
Causal a ian s no agged on he iCOGS geno yping pla o m.We used he
sequenced UK10K common a ian s o e alua e how well he iCOGS geno yped and
impu ed SNPs cap u ed unde lying he i abili y by simula ing pheno ypes using causal
a ian s om sequencing and es ima ing he i abili y om he iCOGS SNPs ( ha is,
hiding a ian s ha we e no geno yped o impu ed, Supplemen a y Table S2). 83% o
common UK10K SNPs lie wi hin 100kb o an iCOGS SNP, so some common a -
ia ion is likely o be pa ially agged by he chip. I he impu ed and/o geno yped
SNPs se ed as a good p oxy o he common sequence a ia ion, hen we would
expec hei es ima es o %h
2
g o ma ch he simula ed ac ions. When no unc ional
ca ego y was en iched wi h causal a ian s, small bu significan di e ences we e
obse ed o geno yped coding a ian s (4.75% h2
ges ima ed as compa ed wi h
simula ed 0.67%) and impu ed in e genic a ian s (56.09% h2
gas compa ed wi h
50.52% simula ed) bu no he ocal LNCaP:H3k27ac ca ego y. Simila de ia ions we e
obse ed o he disease a chi ec u e whe e common a ian s explain mo e a iance
in ai han a e a ian s (Supplemen a y Table S2). When causal a ian s whe e
en iched wi hin LNCaP:H3k27ac ca ego y, de ia ions be ween simula ed and es i-
ma ed SNP he i abili y we e la ge (Supplemen a y Table S2). Mos o his de ia ion
was due o a significan unde es ima e a LNCaP:H3k27ac, which was simula ed o
explain 50% o h2
gbu explained only 12.55% (s.e.m.0.92%) and 30.92% (s.e.m. 1.09%)
om geno yped and impu ed SNPs, espec i ely. This he i abili y was dis ibu ed
ac oss all he emaining componen s, pa icula ly in in e genic SNPs o he geno-
yped es ima e and DHS SNPs o he impu ed es ima e, which end o be nea by.
O e all, ou simula ions showed ha he model is highly accu a e when all causal
a ian s a e yped. When conside ing en ichmen om un yped causal a ian s, he
impu ed es ima e was consis en ly close o he u h han he geno yped es ima e.
Mos impo an ly, he es ima e om he ocal ca ego y (LNCaP H3k27ac in ou
simula ions) was shown o be highly conse a i e bo h in he null and in he en iched
scena io and unlikely o be biased due o agging o un yped ma ke s. We no e ha
p e ious wo k has shown es ima es om impu ed SNPs (bu no geno yped SNPs)
may be con amina ed by ma ke s e y close o an en iched anno a ion12;assuchwe
ocused ou esul s on he densely geno yped iCOGS a ian s which a e expec ed o be
conse a i e, and p ima ily used impu ed da a o alida ion ac oss da a se s.
Es ima es o h2
g om A ican Ame ican samples.To assess po en ial biases in
es ima ing h2
g om an admixed popula ion, we pe o med sepa a e simula ions in
he AAPC da a whe e causal a ian s we e specifically sampled om a ying F
ST
bins. This amewo k e alua ed he po en ial bias esul ing om ma ke s ha had
d i ed o di e en equencies in he wo popula ions. The F
ST
was es ima ed ou -
o -sample in he HapMap CEU Eu opean and YRI Yo uba popula ions. We es ed
he null six-componen model (Coding, UTR, P omo e , DHS, In on, O he ) and
obse ed no significan de ia ions om he null unde any class o di e en ia ed
SNPs (Supplemen a y Table S3). Howe e , we no e ha o al h2
gwas simula ed a
0.50 bu was in e ed a 0.38–0.66 ac oss inc easing quin iles o causal F
ST
(Supplemen a y Table 3), indica ing ha e en wi h well-calib a ed es ima es o
en ichmen he o al es ima e may be biased upwa ds i he causal SNPs a e highly
di e en ia ed (obse ed in his simula ion when mean causal F
ST
4¼0.35).
Gene ic p edic ion.We sough o alida e he u ili y o ou unc ional a las by
applying i o gene ic p edic ion. The aim o gene ic p edic ion is o use aining
indi iduals wi h gene ics ( o example, SNPs) and diagnosed pheno ype o accu a ely
p edic he pheno ype in o indi iduals wi h only gene ic da a a ailable41,42.He e,we
ocus on co ela ion o p edic ed pheno ype wi h ue pheno ype (R2), as i has a
na u al ela ionship o SNP he i abili y12,42. In ui i ely, be e localiza ion o he ue
e ec -sizes will educe noise in aining he p edic o and inc ease accu acy. I he
unc ional a las iden ified egions wi h inc eased he i abili y, his in o ma ion should
significan ly imp o e he p edic ion. We e alua ed h ee s anda d models o isk
p edic ion: GRS; BLUP ( e . 43); and mul i-componen BLUP ( e . 14). The GRS was
compu ed as a sum o e SNPs o he log odds- a ios om he aining sample41.The
se o SNPs used was ei he he genome-wide significan ma ke s in he aining se
( es ic ed o one pe 1MB locus) o he genome-wide significan ma ke s iden ified
in a ecen la ge GWAS o P Ca2. In con as o he GRS, he BLUP used all ma ke s
in he da a o o m he p edic ion. The s anda d BLUP was es ima ed using GCTA
o e all SNPs. The mul i-componen BLUP was es ima ed using he componen s in
he selec ed model (join ly) o compu e a single sco e equal o he sum o he
p edic ions om each componen weigh ed by hei componen -specific h2
g.Thisis
analogous o speci ying a di e en p io on he e ec -size a iance in each
componen . All p edic ions we e ca ied ou by c oss- alida ion in he ull iCOGS
da a, emo ing 1,000 indi iduals in each old. P edic ion R2was hen compu ed om
a eg ession o pheno ype on he p edic o sco e wi h 10 PCs included as co a ia es
o accoun o ances y, subsequen ly sub ac ing he R2¼0.021 om a model wi h
PCs only. P alues we e es ima ed o each o he coe ficien s in he mul iple
eg ession o pheno ype BGRS þsingle-BLUP þmul i-BLUP þPCs. To ensu e
ha p edic ion ac oss da a se s was independen , we ca e ully emo ed all iCOGS
indi iduals wi h a GRM alue o 40.05 o any indi idual in he BPC3 when
compu ing BLUP coe ficien s. We sepa a ely analysed he p edic o in 26,000 iCOGS
samples ha had age a diagnosis, bu did no obse e significan di e ences be o e/
a e including age as a co a ia e.
Re e ences
1. Hjelmbo g, J. B. e al. The he i abili y o p os a e cance in he No dic win
s udy o cance . Cance Epidemiol. Bioma ke s P e . 23, 2303–2310 (2014).
2. Al Olama, A. A. e al. A me a-analysis o 87,040 indi iduals iden ifies 23 new
suscep ibili y loci o p os a e cance . Na . Gene . 46, 1103–1109 (2014).
3. Cas o, E. e al. Ge mline BRCA mu a ions a e associa ed wi h highe isk o
nodal in ol emen , dis an me as asis, and poo su i al ou comes in p os a e
cance . J. Clin. Oncol. 31, 1748–1757 (2013).
4. Eeles, R. A. e al. Iden ifica ion o 23 new p os a e cance suscep ibili y loci
using he iCOGS cus om geno yping a ay. Na . Gene . 45, 385–391 (2013).
5. Saunde s, E. J. e al. Fine-mapping he HOXB egion de ec s common a ian s
agging a a e coding allele: e idence o syn he ic associa ion in p os a e
cance . PLoS Gene . 10, e1004129 (2014).
6. Ewing, C. M. e al. Ge mline mu a ions in HOXB13 and p os a e-cance isk.
N. Engl. J. Med. 366, 141–149 (2012).
7. Hazele , D. J. e al. Comp ehensi e unc ional anno a ion o 77 p os a e cance
isk loci. PLoS Gene . 10, e1004102 (2014).
8. Hazele , D. J., Coe zee, S. G. & Coe zee, G. A. A a e a ian , which des oys
a FoxA1 si e a 8q24, is associa ed wi h p os a e cance isk. Cell Cycle 12,
379–380 (2013).
9. ENCODE P ojec Conso ium e al. An in eg a ed encyclopedia o DNA
elemen s in he human genome. Na u e 489, 57–74 (2012).
10. S ama oyannopoulos, J. A. Wha does ou genome encode? Genome Res. 22,
1602–1611 (2012).
11. Mau ano, M. T. e al. Sys ema ic localiza ion o common disease-associa ed
a ia ion in egula o y DNA. Science 337, 1190–1195 (2012).
12. Guse , A. e al. Pa i ioning he i abili y o egula o y and cell- ype-specific
a ian s ac oss 11 common diseases. Am. J. Hum. Gene . 95, 535–552 (2014).
13. Yang, J. e al. Common SNPs explain a la ge p opo ion o he he i abili y o
human heigh . Na . Gene . 42, 565–569 (2010).
14. Yang, J., Lee, S. H., Godda d, M. E. & Vissche , P. M. GCTA: a ool o
genome-wide complex ai analysis. Am. J. Hum. Gene . 88, 76–82 (2011).
15. Yang, J. e al. Genome pa i ioning o gene ic a ia ion o complex ai s using
common SNPs. Na . Gene . 43, 519–525 (2011).
16. Schumache , F. R. e al. Genome-wide associa ion s udy iden ifies new p os a e
cance suscep ibili y loci. Hum. Mol. Gene . 20, 3867–3875 (2011).
17. Haiman, C. A. e al. Cha ac e izing gene ic isk a known p os a e
cance suscep ibili y loci in A ican Ame icans. PLoS Gene . 7, e1001387
(2011).
18. Pome an z, M. e al.The and ogen ecep o cis ome is ex ensi ely ep og ammed
in human p os a e umo igenesis.Na .Gene .47, 1346–1351 (2015).
19. Shlyue a, D., S amp el, G. & S a k, A. T ansc ip ional enhance s: om
p ope ies o genome-wide p edic ions. Na . Re . Gene . 15, 272–286 (2014).
20. Guse , A. e al. Pa i ioning he i abili y o egula o y and cell- ype-specific
a ian s ac oss 11 common diseases. Am. J. Hum. Gene . 95, 535–552.
21. Lee, S. H. e al. Es ima ion o SNP he i abili y om dense geno ype da a.
Am. J. Hum. Gene . 93, 1151–1155 (2013).
22. Cance Fac s & Figu es o A ican Ame icans 2009–2010. Accessed on:
Decembe 2015.
NATURE COMMUNICATIONS | DOI: 10.1038/ncomms10979 ARTICLE
NATURE COMMUNICATIONS | 7:10979 | DOI: 10.1038/ncomms10979 | www.na u e.com/na u ecommunica ions 9