Genome-wide association study in 79,366 European-ancestry individuals informs the genetic architecture of 25-hydroxyvitamin D levels
Full text
ARTICLE
Genome-wide associa ion s udy in 79,366
Eu opean-ances y indi iduals in o ms he gene ic
a chi ec u e o 25-hyd oxy i amin D le els
Xia Jiang e al.
#
Vi amin D is a s e oid ho mone p ecu so ha is associa ed wi h a ange o human ai s and
diseases. P e ious GWAS o se um 25-hyd oxy i amin D concen a ions ha e iden ified ou
genome-wide significan loci (GC, NADSYN1/DHCR7, CYP2R1, CYP24A1). In his s udy, we
expand he p e ious SUNLIGHT Conso ium GWAS disco e y sample size om 16,125 o
79,366 (all Eu opean descen ). This la ge GWAS yields wo addi ional loci ha bo ing
genome-wide significan a ian s (P=4.7×10−9a s8018720 in SEC23A, and P=1.9×10−14 a
s10745742 in AMDHD1). The o e all es ima e o he i abili y o 25-hyd oxy i amin D se um
concen a ions a ibu able o GWAS common SNPs is 7.5%, wi h s a is ically significan loci
explaining 38% o his o al. Fu he in es iga ion iden ifies signal en ichmen in immune and
hema opoie ic issues, and clus e ing wi h au oimmune diseases in cell- ype-specific analysis.
La ge s udies a e equi ed o iden i y addi ional common SNPs, and o explo e he ole o
a e o s uc u al a ian s and gene–gene in e ac ions in he he i abili y o ci cula ing 25-
hyd oxy i amin D le els.
DOI: 10.1038/s41467-017-02662-2 OPEN
Co espondence and eques s o ma e ials should be add essed o E.H. (email: [email p o ec ed]) o o P.K. (email: [email p o ec ed])
o o D.P.K. (email: [email p o ec ed])
#A ull lis o au ho s and hei a flia ions appea s a he end o he pape
NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions 1
1234567890():,;
Vi amin D is an essen ial a soluble i amin and s e oid
p o-ho mone ha plays an impo an ole in muscu-
loskele al heal h. Vi amin D deficiency has been linked o
au oimmune1,2and in ec ious disease3, ca dio ascula disease4,
cance 5, and neu odegene a i e condi ions6. Se um 25-
hyd oxy i amin D, a p ima y ci cula ing o m o i amin D
and a measu e ha bes eflec s i amin D s o es, is influenced by
many ac o s including sun exposu e, age, body mass index7,
die a y in ake o ce ain oods such as o ified dai y p oduc s and
oily fish, supplemen s, and gene ic ac o s8. The concen a ion o
25-hyd oxy i amin D has been epo ed o be highly he i able,
wi h he i abili y es ima es o 50–80% om classical win
s udies9,10.
A genome-wide associa ion s udy (GWAS) me a-analysis o
se um 25-hyd oxy i amin D11 in 4501 pa icipan s o Eu opean
ances y and eplica ion in 2221 samples iden ified a ian s in
h ee loci (g oup componen (GC), 7-dehyd ochles e ol educ ase
(NADSYN1/DHCR7), and 25-hyd oxylase (CYP2R1)). A la ge
GWAS conduc ed by he SUNLIGHT conso ium in 16,125
Eu opean ances y indi iduals, wi h a eplica ion sample o
17,871, eplica ed hese h ee loci and disco e ed one addi ional
locus (CYP24A1)8. Howe e , despi e hese loci being in o nea
genes encoding p o eins in ol ed in i amin D syn hesis, he
associa ed a ian s collec i ely explain only a small ac ion o he
a iance in 25-hyd oxy i amin D concen a ions (~5%)8,11,12.
The e o e, o ex end ou p e ious findings and be e unde s and
he gene ic a chi ec u e unde lying se um 25-hyd oxy i amin D,
as well as es o in e ac ions be ween die a y i amin D in ake
and gene ic ac o s, we conduc ed a la ge-scale GWAS me a-
analysis on his impo an i amin.
Ou GWAS wi h a 79,366 disco e y sample and a 40,562
eplica ion sample eplica es ou p e ious loci and iden ifies wo
new gene ic loci o se um le els o 25-hyd o xy i amin D. We
u he find e idence o a sha ed gene ic basis be ween ci cu-
la ing 25-hyd oxy i amin D and au oimmune diseases. Ou
analyses sugges a ela i ely modes SNP-he i abili y a e o 25-
hyd oxy i amin D when conside ing only common a ian s.
La ge s udies a e equi ed o iden i y addi ional common SNPs,
and o explo e he ole o a e o s uc u al a ian s. The gene ic
ins umen s iden ified by ou esul s could be used in u u e
Mendelian Randomiza ion analyses o he associa ion be ween
i amin D and complex ai s.
Resul s
S udy desc ip ion and GWAS. This s udy ep esen s an expan-
sion o ou p e ious SUNLIGHT conso ium GWAS8. He e, we
combine he 5 disco e y coho s and 5 in-silico eplica ion
coho s om ha s udy, and augmen hese wi h 21 addi ional
coho s ha ha e joined he SUNLIGHT conso ium since 2010
(s udy cha ac e is ics a e desc ibed in Supplemen a y Table 1,
Supplemen a y No e 1). In con as o he p e ious me a-analysis
which in ol ed disco e y, in-silico and de-no o geno yping
s ages, we pe o med a fi s s age disco e y me a-analysis on a
o al o up o 79,366 indi iduals and eplica ed no el findings in
wo independen sepa a e in-silico da a se s (40,562 indi iduals
collec ed by EPIC and 2195 indi iduals collec ed by SOCCS). To
assess and con ol o popula ion s a ifica ion, we examined QQ-
plo s and genomic con ol infla ion ac o s o each con ibu ing
coho p io o me a-analysis. We did no obse e e idence o
widesp ead infla ion (median λ
GC
=0.92; only 1/31 samples wi h
λ
GC
>1.01), indica ing ha ou GWAS esul s we e no infla ed
by popula ion s a ifica ion o c yp ic ela edness (Supplemen a y
Fig. 1). Despi e he sligh ly defla ed λ
GC
obse ed in some o he
cons i uen coho s mos p obably due o o e -co ec ion o es
s a is ics, ou λ
GC
o 0.99 in all samples indica ed app op ia e
con ol o popula ion s a ifica ions and con ounde s. We
iden ified six suscep ibili y loci ha bo ing genome-wide sig-
nifican SNPs, confi ming ou p e iously epo ed loci a GC
(P=4.7×10−343 a s3755967), NADSYN1/DHCR7 (P=3.8×10−62
a s12785878), CYP2R1 (P=2.1×10−46 a s10741657), CYP24A1
(P=8.1×10−23 a s17216707), and wo no el loci a AMDHD1
(P=1.9×10−14 a s10745742) and SEC23A (P=4.7×10−9a
s8018720) (Table 1; Manha an plo s and QQ-plo s o o e all
samples a e p esen ed in Fig. 1, and egional associa ion plo s a e
p esen ed in Supplemen a y Fig. 2). The associa ions a bo h
no el loci we e confi med in he wo independen in-silico
eplica ion coho s (EPIC: P=1.21×10−8a s10745742, P=
5.24×10−4a s8018720; SOCCS: P=0.03 a s10745742, P=0.04
a s8018720) wi h consis en di ec ion o e ec bu sligh ly la ge
e ec sizes and wide confidence in e als obse ed in SOCCS
which could be due o he educed sample size. When analyzing
he wo eplica ion da a se s oge he wi h he disco e y da a se ,
he P- alues in pooled samples became mo e significan (P
pooled
=
2.10×10−20 a s10745742, P
pooled
=1.11×10−11 a s8018720)
(Table 1). We also ound mo e han one dis inc signal a ising
om a ian s a he GC,CYP2R1, and AMDHD1 loci h ough
condi ional and join analysis, bu no o he NADSYN1/DHCR7,
CYP24A1, and SEC23A loci whe e only one p ima y associa ed
SNP was iden ified (Supplemen a y Table 2).
SNP by die a y i amin D in ake in e ac ion. In addi ion o
pe o ming he ma ginal e ec me a-analysis using all samples,
we also es ed a model wi h SNPs and die a y i amin D in ake as
main e ec s and a e m o hei in e ac ion in a subse o
samples. Die ques ionnai es, including i amin D in ake, we e
a ailable o a subse o 13 coho s and an addi ional 2 coho s
ha we e no included in he o e all me a-GWAS analysis (a o al
o 15 coho s, N=41,981). We pe o med a GWAS explici ly
allowing o an in e ac ion be ween i amin D in ake and SNP
geno ypes, in which die a y i amin D was coded as a con inuous
a iable. We pe o med wo es s: (i) a 1 deg ee-o - eedom
in e ac ion es be ween each SNP and i amin D in ake, and (ii)
a 2 deg ee-o - eedom join es o main gene ic and in e ac ion
e ec s. Howe e , o compa ison pu poses, we also pe o med,
(iii) a s anda d es o ma ginal gene ic e ec a e adjus ing o
i amin D in ake in he same sub-samples (Supplemen a y Fig. 3,
and Supplemen a y Fig. 4). The ma ginal gene ic e ec analyses
confi med exis ing associa ion signals a GC (lead SNP s2282679,
in comple e linkage disequilib ium wi h he lead SNP s3755967
iden ified om me a-GWAS using all indi iduals), CYP2R1,
NADSYN1/DHCR7 (lead SNP s4944062, in comple e linkage
disequilib ium wi h he lead SNP s12785878 iden ified om
me a-GWAS using all indi iduals) and CYP24A1, as well as he
no el associa ion a AMDHD1 (P=5.7×10−9, Supplemen a y
Table 3). The join analysis also iden ified he fi e genes abo e,
bu wi h less significan P- alues. Fo ins ance, he associa ion a
AMDHD1 achie ed only sugges i e genome-wide significance (P
=1.2×10−7, Table 2). The in e ac ion es did no iden i y any
a ian s a genome-wide significance le el. Among he 5 SNPs
significan in ma ginal e ec es s, he lead SNP in CYP2R1
showed nominal significance o in e ac ion wi h die a y i amin
D in ake ( s10741657, P=0.028), bu no in e ac ions we e
obse ed o he o he SNPs ( s2282679, P=0.45; s4944062, P=
0.74; s10745742, P=0.64; s17216707, P=0.46). Repea ing he
analysis using a e ile coding o i amin D ins ead o a con-
inuous coding did no quali a i ely change he esul s.
SNP-he i abili y o 25-hyd oxy i amin D. We u he e alua ed
he SNP-he i abili y, defined as he he i abili y explained by
GWAS SNPs o 25-hyd oxy i amin D, using LD sco e eg ession
ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2
2NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions
(see Me hods). The o e all obse ed he i abili y o 25-
hyd oxy i amin D es ima ed by using all common SNPs
(2,579,296 a e QC) was 7.54% (s anda d e o (SE): 1.88%).
A e excluding genome-wide significan SNPs (533 SNPs wi h
P≤5×10−8) om six loci and all SNPs wi hin ±500 kb o hose
loci, he he i abili y dec eased o 4.70% (SE: 0.72%). The es ima e
u he dec eased o 1.73% (SE: 0.32%) a e excluding all SNPs
ha eached nominal significance (156,675 SNPs wi h P≤0.05).
These esul s indica e ha common a ian s agged by GWAS
chips explain a modes ac ion o o e all a iabili y in ci cula ing
30
GC NADSYN1/DHCR7
CYP2R1
AMDHD1
SEC23A
CYP24A1
Know
No el
25
20
–log10 (p – alue)
15
10
5
200 = 0.99
01234
Expec ed –log10 co ec ed p- alue
56
Obse ed –log10 co ec ed p- alue
100
50
0
1
2
3
4
5
6
7
8
Ch omosome
9
10
11
12
13
14
15
16
17
18
19
20
21
22
a
b
Fig. 1 Genome-wide associa ion o ci cula ing 25-hyd oxy i amin D g aphed by ch omosome posi ions and −log10 P- alue (Manha an plo ), and quan ile-
quan ile plo o all SNPs om he me a-analysis (QQ-plo ). aManha an plo : The P- alues we e ob ained om he single s age fixed-e ec s in e se
a iance weigh ed me a-analysis. The Yaxis shows −log
10
P- alues, and he Xaxis shows ch omosome posi ions. Ho izon al g ay dash line ep esen s he
h esholds o P=5×10−8 o genome-wide significance. Known loci we e colo ed coded as ed, and no el loci we e colo coded as g een. bQQ-plo : The Y
axis shows obse ed −log
10
P- alues, and he Xaxis shows he expec ed −log
10
P- alues. Each SNP is plo ed as a black do , and he dash line indica es null
hypo hesis o no ue associa ion. De ia ion om he expec ed P- alue dis ibu ion is e iden only in he ail a ea, wi h a lambda o 0.99, sugges ing ha
popula ion s a ifica ion was adequa ely con olled
Table 1 Single nucleo ide polymo phisms iden ified in genome-wide analyses o ci cula ing 25-hyd oxy i amin D concen a ions
Gene SNP Ch omosome: Posi ion E ec / e e ence
allele
Allele equency Me a-GWAS es ima es
E ec (Be a) S anda d E o P- alue
Fi s s age disco e y me a-GWAS (N=79,366)
GC s3755967 4:72828262 T/C 0.28 −0.089 0.0023 4.74E–343
NADSYN1/ DHCR7 s12785878 11:70845097 T/G 0.75 0.036 0.0022 3.80E–62
CYP2R1 s10741657 11:14871454 A/G 0.4 0.031 0.0022 2.05E–46
CYP24A1 s17216707 20:52165769 T/C 0.79 0.026 0.0027 8.14E–23
AMDHD1 s10745742 12:94882660 T/C 0.4 0.017 0.0022 1.88E–14
SEC23A s8018720 14:38625936 C/G 0.82 −0.017 0.0029 4.72E–09
Replica ion da a se 1: samples collec ed by EPIC (N=40,562)
AMDHD1 s10745742 12:94882660 T/C 0.41 0.041 0.0071 1.21E–08
SEC23A s8018720 14:38625936 C/G 0.83 −0.032 0.0093 5.24E–04
Replica ion da a se 2: addi ional con ol samples collec ed by SOCCS (N=2195)
AMDHD1 s10745742 12:94882660 T/C 0.37 0.045 0.021 0.03
SEC23A s8018720 14:38625936 C/G 0.81 −0.051 0.026 0.04
Pooled analysis (disco e y me a-GWAS + eplica ion 1 + eplica ion 2) (N=122,123)
AMDHD1 s10745742 12:94882660 T/C 0.39 0.019 0.002 2.10E–20
SEC23A s8018720 14:38625936 C/G 0.82 −0.019 0.0027 1.11E–11
In he pooled analysis, P
he e ogenei y
=0.003 o AMDHD1 s10745742; P
he e ogenei y =
0.14 o SEC23A s8018720.
NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2 ARTICLE
NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions 3
25-hyd oxy i amin D, and ha an app eciable p opo ion o his
SNP-he i abili y is explained by he six gene ic egions o asso-
cia ed SNPs iden ified h ough GWAS.
Pa i ioning he o al he i abili y o 25-hyd oxy i amin D.We
nex pa i ioned he he i abili y by unc ional elemen s using
baseline model wi h 24 publicly a ailable anno a ions (see
Me hods), and obse ed la ge and significan en ichmen o
se e al unc ional ca ego ies (Fig. 2, Supplemen a y Table 4). Fo
example, we ound he la ges en ichmen in weak enhance s,
wi h 2.1% o SNPs explaining 42.3% o he o e all he i abili y
(20- old en ichmen , P=0.02), ollowed by conse ed egions
(13.8- old en ichmen , P=0.03), open ch oma in (as eflec ed by
DHS, 8.5- old en ichmen , P=0.02), ansc ip ion ac o binding
si es (5.7- old en ichmen , P=0.048), supe -enhance s (1.9- old
en ichmen , P=0.04), and all ou his one ma ks we e en iched
(bo h e sions o H3K27ac (one e sion p ocessed by Hnisz e al.,
and ano he e sion used by he Psychia ic Genomics Con-
so ium (PGC), H3K4me1, H3K4me3 (500 bp), H3K29ac (500
bp)). We also obse ed deple ion o ep essed egions (0.06- old
en ichmen , P=0.006). Howe e , none o hose anno a ions
wi hs ood mul iple- es ing co ec ions (Bon e oni co ec ed P-
h eshold: 0.05/24) excep o he ac i e enhance his one ma k
H3K27ac (PGC) (4.2- old en ichmen , P=8×10−4) and
H3K4me1 (1.8- old en ichmen , P=0.0019).
We subsequen ly pe o med cell- ype-specific analysis by using
10 b oad cell- ype g oups. As shown in Table 3, he op h ee
en ichmen s we e in he immune and hema opoie ic issues (4.3-
old en ichmen , P=2.2×10−5), gas oin es inal issues (4.4- old
en ichmen , P=0.0017), and CNS (3.6- old en ichmen , P=
0.0039). The e was also significan en ichmen o li e , kidney,
and connec i e and bone issues, bu hese esul s did no su i e
mul iple- es ing co ec ions. When u he analyzing 220 cell-
ype-specific anno a ions, we obse ed he mos significan
en ichmen in CD19 cells (app oxima ely 8- old en ichmen ,
P~0.001), ollowed by CD20 cells (6.4- old en ichmen , P=0.003)
and CD3 cells (7.8- old en ichmen , P=0.01) (Supplemen a y
Da a 1).
Gene ic co ela ions be ween 25-hyd oxy i amin D and ai s.
We con inued o assess he gene ic co ela ion be ween 25-
hyd oxy i amin D and each o he 37 ai s wi h publicly a ailable
GWAS summa y s a is ics da a (Supplemen a y Table 5). None o
he gene ic co ela ions emained significan a e Bon e oni
co ec ion (co ec ed P- h eshold: 0.05/37, Fig. 3). Wi hou
mul iple- es ing co ec ion, he e we e some co ela ions wi h
nominal s a is ical significance. Fo example, e e smoking
( g(SE): −0.17 (0.073), P=0.019), p ima y bilia y ci hosis
( g(SE): −0.18 (0.076), P=0.019) and BMI adjus ed wais -hip-
a io ( g(SE): −0.10 (0.050), P=0.042) we e obse ed o be
in e sely co ela ed wi h 25-hyd oxy i amin D; whe eas lung
unc ion ( g(SE): 0.14 (0.046), P=0.0036) showed a posi i e
co ela ion wi h 25-hyd oxy i amin D. Subsequen di ec ional
gene ic co ela ion analysis did no e eal any appa en pu a i e
causal ela ionship o 25-hyd o y i amin D wi h o he ai s,
excep o a po en ial link be ween 25-hyd oxy i amin D and
Table 2 Resul s om he SNP-by-die a y i amin D in ake in e ac ion analysis
Gene SNP Ch omosome:
Posi ion
E ec /
Re e ence
Allele
Allele
F equency
SNP-by-die a y i amin D in ake In e ac ion analysis
Main
Gene ic
E ec
In e ac ion
E ec
P- alue o
in e ac ion
P- alue
o join
es
E ec
(Be a_G)
S anda d
E o
E ec
(Be a_In )
S anda d
E o
Fi s s age disco e y me a-GWAS(N=79,366)
GC s3755967 4:72828262 T/C 0.28 −0.082 0.0042 −2.01E–05 1.67E–05 0.23 2.92E–171
s2282679* 4:72827247 T/G 0.28 0.085 0.004 1.20E–05 1.60E–05 0.45 1.40E–187
NADSYN1/
DHCR7
s12785878 11:70845097 T/G 0.75 0.033 0.0039 6.61E–06 1.62E–05 0.68 3.52E–29
s4944062* 11:70864942 T/G 0.75 0.034 0.004 5.30E–06 1.60E–05 0.74 1.90E–31
CYP2R1 s10741657 11:14871454 A/G 0.4 0.03 0.0035 3.21E–05 1.46E–05 0.028 2.23E–38
CYP24A1 s17216707 20:52165769 T/C 0.79 0.025 0.0048 1.39E–05 1.88E–05 0.46 1.32E–14
AMDHD1 s10745742 12:94882660 T/C 0.4 0.016 0.0036 −7.05E–06 1.49E–05 0.64 1.20E–07
SEC23A s8018720 14:38625936 C/G 0.82 −0.013 0.0051 −2.40E–05 2.06E–05 0.24 1.94E–05
* Top SNPs iden ified in he SNP-by-die a y i amin D in ake in e ac ion analysis, pe o med in a subse o indi iduals. Fo GC and NADSYN1/DHCR7, he op SNPs iden ified h ough he ma ginal e ec
eg ession me a-analysis using all indi iduals we e in high linkage disequilib ium wi h he op SNPs iden ified h ough he SNP-by-die a y i amin D in ake in e ac ion analysis using a subse o indi iduals
( 2 o s3755967 and s2282679: 1.0; 2 o s12785878 and s4944062: 1.0). Be a_G indica es he main e ec o he SNP, Be a_In indica es he in e ac ion e ec o SNP-by-die a y i amin D in ake
20
Type
24anno a ions
24anno a ions_500 bp
En ichmen =P op.H2/P op.SNPs
10
0
Weak enhance
Conse ed
TSS
Fe al DHS
TFBS
H3K27ac (PGC)
H3K27ac
H3K4me1
H3K4me3
Supe enhance
H3K27ac (Hnisz)
Rep essed
Func ional ca ego ies
**
**
Fig. 2 He i abili y en ichmen o he op 12 genomic unc ional elemen s.
We pa i ioned he SNP-he i abili y o se um 25-hyd oxy i amin D
concen a ions in o 24 publicly a ailable genomic unc ional elemen s using
LD-sco e eg ession. We plo ed he en ichmen (Yaxis) o each o he 12
op anno a ions (as shown in Xaxis) in o a ba cha . G ay ba s and blue
ba s ep esen he anno a ions wi h and wi hou he 500 base-pai
windows. The heigh o each ba ep esen s magni ude o en ichmen .
Significan es ima es o en ichmen ha passed Bon e oni co ec ions
(P- alue o en ichmen <0.05/24) a e ma ked wi h double s a s. TSS
ansc ip ion s a si es, DHS DNase I hype sensi i e si es, TFBS
ansc ip ion ac o binding si es, Rep essed ep essed egions
ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2
4NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions
HDL (Supplemen a y Table 6). Howe e , wi h only six
25-hyd oxy i amin D associa ed SNPs included in he analysis,
we conside an o e all null di ec ional co ela ion as ou main
finding, and u he well-designed la ge-scale Mendelian ando-
miza ion analyses a e wa an ed.
Finally, we analyzed he 220 cell- ype-specific anno a ions in
each o he 37 ai s and compa ed he cell- ype-specific
en ichmen s o 25-hyd oxy i amin D o he en ichmen s o
hese ai s. The en ichmen pa e n o 25-hyd oxy i amin D
di e ed no ably om he pa e ns o psychia ic diseases and
me abolic ela ed ai s. Psychia ic diseases showed en ichmen
o his one ma ks specific o CNS cell ypes, and me abolic
diseases showed en ichmen o gas oin es inal cell ypes, while
hese anno a ions we e dep essed in 25-hyd oxy i amin D.
Con e sely, 25-hyd oxy i amin D showed simila pa e ns wi h
au oimmune inflamma o y diseases, whe e mul iple immune cell
ypes we e en iched. We consis en ly obse ed ha 25-
hyd oxy i amin D was clus e ed wi h au oimmune diseases
(Supplemen a y Fig. 5).
Discussion
Vi amin D inadequacy has been linked o many diseases such as
cance , au oimmune diso de and ca dio ascula condi ions in
addi ion o musculoskele al diseases, which has led o subs an ial
in e es in he de e minan s o i amin D s a us, especially i s
gene ic componen s. We ha e pe o med a la ge 25-
hyd oxy i amin D me a-GWAS in ol ing 31 s udies wi h a
o al o 79,366 indi iduals. Ou esul s ecapi ula ed se e al p e-
iously epo ed findings. Fi s o all, we confi med he ole o
common gene ic a ian s in egula ion o ci cula ing 25-
hyd oxy i amin D concen a ions. Ou s udy alida ed h ee
loci, GC,NADSYN1/DHCR7, and CYP2R1, all we e es ablished
25-hyd oxy i amin D isk loci iden ified h ough wo ea lie
GWASs8,11. In addi ion, we we e able o confi m he associa ion
o a locus con aining CYP24A1 wi h 25-hyd oxy i amin D con-
cen a ions using ou la ge sample size, which highligh s he
impo ance o his p o ein in he deg ada ion o i amin D
molecule, by ca alyzing hyd oxyla ion eac ions a he side chain
o 1,25-dihyd oxy i amin D, he physiologically ac i e o m
(ho monal o m) o i amin D. Significan finding a his locus
was only shown in he pooled analyses in ol ing bo h disco e y
and eplica ion samples in an ea lie GWAS8.
We ex ended p e iously epo ed findings by iden i ying wo
addi ional new loci. SEC23A (Sec23 Homolog A, coa p o ein
complex II (COPII) componen ) encodes a membe o
SEC23 sub amily. In euka yo ic cells, sec e ed p o eins a e syn-
hesized in he endoplasmic e iculum (ER), packaged in o
COPII-coa ed esicles, and a fic o he Golgi appa a us. As pa
o COPII complex, SEC23 plays a ole in p omo ing ER-Golgi
p o ein a ficking. SEC23A mu a ions ha e been epo ed o
cause c aniolen iculosu u al dysplasia, a disease cha ac e ized by
c anio acial and skele al mal o ma ion such as delay in closu e o
on anels, su u al ca a ac s and acial dysmo phisms, due o
de ec i e collagen sec ec ion13,14. The second no el locus is
AMDHD1 (amidohyd olase domain con aining 1). This gene
encodes an enzyme in ol ed in he his idine, lysine, phenylala-
nine, y osine, p oline and yp ophan ca abolic pa hway. Mu a-
ions in AMDHD1 a e ound o be associa ed wi h a ypical
lipoma ous umo , a cance o connec i e issues ha esemble a
cells15.
Ou SNP-he i abili y esul s sugges ha 25-hyd oxy i amin D
has a modes o e all he i abili y due o common genome-wide
SNPs o 7.5%, and ha an app eciable p opo ion (2.84% ou o
7.5%, i.e., 38%) o his o al could be explained by known gene ic
egions iden ified h ough GWAS. Ou findings a e in line wi h a
p e ious published epo (by Hi aki e al.12) which es ima ed he
a iance in ci cula ing 25-hyd oxy i amin D explained by SNPs
in a o al o 5575 indi iduals12. Acco ding o ha epo , by
employing a linea mixed model fi ing he addi i e gene ic
ma ix c ea ed om all geno yped and impu ed SNPs, he p o-
po ion o a iance explained was 8.9%; by employing a polygenic
sco e app oach comp ised o he hen GWAS-disco e ed SNPs
(GC, CYP2R1, DHCR7/NADSYN1), he p opo ion o a iance
explained was 5%. Bo h o hese es ima es we e close o ou s. In
Hi aki e al., he known 25-hyd oxy i amin D associa ed en i -
onmen al ac o s such as age, BMI, season o blood d awn,
i amin D die a y in ake, i amin D supplemen in ake, egion o
esidence and e hnici y, explained ~18% o he obse ed a -
iance12. Ou esul s, in ag eemen wi h hese findings, sugges
ha al hough he e appea s o be some polygenic signals ou side
o he iden ified egions, he emaining common e ec s may be
small. The e also may be low equency a ian s wi h la ge
e ec s ha we e no in es iga ed he e. Fo example, while his
pape was unde e iew, a ela ed s udy iden ified low- equency
(MAF =2.5%) synonymous coding a ian s117913124_A a
CYP2R1 con e ing a la ge e ec on 25-hyd oxy im ain D le els,
which was ou imes g ea e in magni ude and independen o a
p e iously desc ibed associa ion o a common a ian
( s10741657) nea CYP2R116.
Resul s o win and amilial s udies ha e e ealed a subs an ial
gene ic basis in he a iabili y o ci cula ing 25-hyd oxy i amin D
le els, wi h es ima es o he i abili y eaching as high as
86%9,10,17–19. These es ima es, howe e , seem o be influenced by
en i onmen al condi ions. Fo example, in a s udy conduc ed by
O on e al. wi h 40 monozygo ic and 59 dizygo ic win pai s,
bloods we e collec ed a he end o win e and a he i abili y o
77% was epo ed10. Simila ly, he s udy conduc ed by Ka ohl
e al. wi h 310 monozygo ic and 200 dizygo ic male wins
obse ed a he i abili y o 70% du ing win e , whe eas in summe ,
se um 25-hyd oxy i amin D concen a ions appea ed o be
Table 3 He i abili y en ichmen o en g ouped cell ypes
Ca ego y P opo ion o SNPs (%) P opo ion o h2
g(%) En ichmen (s anda d e o s) P- alue
Kidney 4.26 27.27 6.4 (2.44) 0.027
Li e 7.22 37.68 5.22 (1.55) 0.01
Gas oin es inal 16.77 72.88 4.35 (0.97) 0.0017
Immune and hema opoie ic 23.34 100.17 4.29 (0.76) 2.20E-05
Cen al ne ous sys em 14.88 54.09 3.64 (0.87) 0.0039
Ca dio ascula 11.11 35.74 3.22 (1.26) 0.078
Connec i e issue/bone 11.5 35.65 3.1 (1.04) 0.037
Ad enal/panc eas 9.36 26.17 2.8 (1.31) 0.18
O he 20.27 56.68 2.8 (0.98) 0.076
Skele al Muscle 10.38 14.29 1.38 (1.25) 0.76
Black bold on indica es significan P- alues a e mul iple co ec ions (P<0.05/10)
NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2 ARTICLE
NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions 5
en i ely de e mined by non-gene ic ac o s (he i abili y: 0%)9.
Compa able es ima es we e also iden ified in a sligh ly la ge
s udy conduc ed by Mills e al. (win e : 90% s. summe : 56%)18.
Consis en wi h season dependency, sex di e ences we e also
obse ed (males: 86% s. emales: 17%)17. While hese es ima es
should be ea ed wi h cau ion due o small samples and ela ed
imp ecision, hey confi m he subs an ial a ia ion in 25-
hyd oxy i amin D le els by season (as shown p e iously20) and
illus a e ha he i abili y es ima es de i ed om a homogenous
sou ce may be highly infla ed. In a ele an ly well-powe ed win
s udy wi h a o al o ~2100 emale wins, he he i abili y o 25-
hyd oxy i amin D was calcula ed o be 40%, indica ing a la ge
p opo ion o a iance explained by non-gene ic ac o s21. He -
i abili y es ima es ob ained using GWAS SNPs ha e ypically
been ound o be app oxima ely hal o hose om classical win
s udies9,10, bu ou es ima e o 7.5%, calcula ed using common
genome-wide SNPs, is a lowe han epo ed he i abili y om
win and amily based s udies. In addi ion o po en ially infla ed
es ima es om win s udies, he di e ence may eflec he p o-
po ion o he i abili y explained by a e SNPs o s uc u al a -
ian s ha we e no included in ou da a, and he po en ial gene-
gene in e ac ions ha emain o be iden ified. The combina ion o
ou samples om all seasons is also likely o dec ease he p ob-
abili y o finding gene ic a ian s, and hence defla e he i abili y
es ima es.
Th ough pa i ioning he SNP-he i abili y o se um 25-
hyd oxy i amin D le els, we obse ed a significan en ichmen
in immune and hema opoie ic issues; likewise, he cell- ype-
specific analysis e ealed clus e ing o 25-hyd oxy i amin D and
au oimmune diseases, indica ing ha hese ai s sha e a majo i y
T ai s Gene ic co ela ion
(95%CI)
Au oimmune/in lamma o y diseases
Age- ela ed macula degene a ion –0.02 (–0.1, 0.06)
p- alue Signi icance
*
*
*
*
0.57
0.89
0.25
0.98
0.21
0.019
0.44
0.59
0.9
0.91
0.71
0.39
0.7
0.64
0.63
0.49
0.61
1
0.25
0.66
0.98
0.21
0.36
0.18
0.26
0.54
0.019
0.13
0.25
0.45
0.63
0.69
0.0036
0.64
0.078
0.8
0.042
–0.02 (–0.17, 0.15)
–0.07 (–0.19, 0.05)
0 (–0.15, 0.14)
–0.17 (–0.45, 0.1)
–0.18 (–0.33, –0.03)
0.05 (–0.08, 0.19)
0.04 (–0.1, 0.18)
–0.01 (–0.11, 0.1)
–0.01 (–0.12, 0.1)
–0.02 (–0.13, 0.09)
–0.08 (–0.28, 0.11)
–0.03 (–0.15, 0.1)
0.03 (–0.11, 0.17)
0.04 (–0.11, 0.18)
0.04 (–0.08, 0.17)
0.02 (–0.07, 0.11)
0 (–0.08, 0.08)
0.14 (–0.1, 0.38)
–0.03 (–0.16, 0.1)
0 (–0.12, 0.12)
0.07 (–0.04, 0.19)
0.06 (–0.07, 0.18)
–0.06 (–0.16, 0.03)
–0.08 (–0.22, 0.06)
0.02 (–0.05, 0.1)
–0.17 (–0.32, –0.03)
–0.07 (–0.02, 0.15)
–0.07 (–0.19, 0.05)
–0.03 (–0.11, 0.05)
0.02 (–0.06, 0.09)
–0.01 (–0.08, 0.06)
0.14 (0.04, 0.23)
–0.02 (–0.09, 0.05)
0.06 (–0.01, 0.12)
–0.01 (–0.08, 0.06)
–0.1 (–0.2, 0)
–0.5 –0.25 00.25 0.5
Celiac Disease
C ohn’s Disease
Lupus
Mul iple Scle osis
P ima y bilia y ci hosis
Rheuma oid A h i is
Ulce a i e Coli is
UKBiobank As hma
UKBiobank Eczema
In lamma o y Bowel Disease
Alzheime ’s
Psychia ic diso de / ai s
Me abolism ela ed ai s
Co ona y A e y Disease
Fas ing Glucose
HDL
LDL
T iglyce ides
Type 2 Diabe es
UKBiobank Hype ension
E e Smoked
UKBiobank Age a Mena che
UKBiobank Age a Menopause
UKBiobank BMI
UKBiobank Dias olic
UKBiobank FEV1FVC
UKBiobank FVC
UKBiobank Heel TSco e
UKBiobank Heigh
UKBiobank Sys olic
UKBiobank Wais -Hip Ra io
O he s
Ano exia
Au ism
Bipola Diso de
Dep essi e Synd ome
Neu o icism
Schizoph enia
Subjec Well Being
Fig. 3 Gene ic co ela ions be ween 25-hyd oxy i amin D and 37 ai s. We collec ed GWAS summa y s a is ics o 37 diseases and ai s spanning a wide
ange o pheno ypes (au oimmune inflamma o y diseases, psychia ic diso de s, me abolic ai s, and an h opome ic index) om publicly a ailable
esou ces, and es ima ed hei sha ed gene ic simila i ies wi h se um 25-hyd oxy i amin D le els. We plo ed he gene ic co ela ion oge he wi h 95%
confidence in e als using a blue squa e and g ay ho izon al lines. Red e ical line indica es no gene ic co ela ion ( g=0). S a is ical significance was
defined as P- alue <0.05. None o he pai wise co ela ions passed Bon e oni co ec ions
ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2
6NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions
o common cell ypes. The link be ween i amin D deficiency and
inc eased isk o au oimmune inflamma o y diseases has long
been ecognized by epidemiological in es iga ions22,23. Al hough
he unde lying mechanisms emain unclea , i is now e iden ha
i amin D is in ol ed in many biological p ocesses ha egula e
bo h inna e and adap i e immune esponses, h ough ligand-
ecep o binding, ac i a ion, in e ac ion wi h esponse elemen s
in he p omo e egions o di e en genes, and e en ually lead o
unc ional changes in a wide a ie y o immune cells including
Th1, Th2, Th17, T egula o y and na u al kille T cells22,23. The
sha ed cell ype en ichmen s be ween i amin D and au oimmune
diseases obse ed in ou s udy, u he sugges ha i amin D no
only a ec s au oimmune diseases h ough i s di ec e ec (as a
ligand), bu also h ough hei sha ed gene ic e iology. Thus,
indi iduals wi h i amin D deficiency may be mo e suscep ible o
hese diso de s, bo h because o en i onmen al and gene ic
influences.
Ou genome-wide in e ac ion analysis be ween gene ic a ian s
and die a y in ake o i amin D did no iden i y new signals. All
significan associa ions obse ed in he join es o main gene ic
and in e ac ion e ec s we e o equal o highe significance le el
(i.e., lowe P- alues) in he GWAS o ma ginal gene ic e ec
pe o med in he same indi iduals, indica ing no majo con-
ibu ion o in e ac ion e ec s a hese loci. Indeed, only one o
he op 5 loci om he o e all ma ginal GWAS showed nominally
significan in e ac ion e ec , and none passed Bon e oni co -
ec ions. While smalle gene-die in e ac ion e ec s emain o be
disco e ed, ou esul s p o ide some e idence agains la ge
in e ac ions be ween common SNPs and die a y i amin D
in ake. S ill, one canno comple ely ule ou he possibili y o
in e ac ion, bu only conclude ha gene ic e ec s appea s able
wi hin i amin D in ake ange in he popula ions s udied. Indeed,
as o any gene-en i onmen in e ac ion es s, s a is ical powe is
highly dependen on he a iance o exposu e in he samples
analyzed24, and in e ac ions would emain unobse ed i he
exposu e is homogeneous among indi iduals. Also, we we e no
able o cap u e i amin D supplemen a ion adequa ely o include
his in he die a y in ake a iable, and we e no able o es ima e
sunligh exposu e as a sou ce o i amin D p oduc ion in he skin.
Se um 25-hyd oxy i amin D concen a ions a e mainly
de e mined by modifiable en i onmen al ac o s, and con a y o
es ima es om p e ious win s udies, ou la ge-scale analyses
sugges a SNP-he i abili y a e ha is ela i ely modes in mag-
ni ude when conside ing common a ian s. Ou s udy also
showed ha common gene ic a ian s a e unlikely o ha e a
s ong modi ying e ec on inc eases in 25-hyd oxy i amin D
ollowing ypical die a y in akes, sugges ing ha conside a ion o
gene ic backg ound is no equi ed when de e mining popula ion
based i amin D in ake ecommenda ions. Howe e , ou esul s
suppo he ole o i amin D in immunological diseases as we
obse ed om cell- ype-specific analysis o clus e ing o i amin
D and au oimmune diseases, and he e idence o signal
en ichmen o immune and hema opoie ic issues. These find-
ings a e in line wi h p e ious Mendelian Randomiza ion s udies
which ound a pu a i e causal associa ion be ween i amin D and
au oimmune diseases such as mul iple scle osis1,2and ype 1
diabe es25. The addi ional gene ic ins umen s iden ified by ou
esul s could also be used in u u e Mendelian Randomiza ion
analyses o he associa ion be ween i amin D and complex ai s.
Me hods
S udy coho s. We expanded ou p e ious SUNLIGHT conso ium GWAS, and
unde ook a la ge, mul icen e , genome-wide associa ion s udy o 31 coho s in
Eu ope, Canada and USA. Ou fi s s age disco e y me a-analysis consis ed o
79,366 samples o Eu opean descen d awn om 31 epidemiological coho s.
Among hose 31 coho s, en we e used as disco e y and in-silico eplica ion
samples in ou p e ious GWAS publica ion ( he 1958 B i ish Bi h Coho
(1958BC), he Ca dio ascula Heal h S udy (CHS), he F amingham Hea S udy
(FHS), he Go henbu g Os eopo osis and Obesi y De e minan s s udy (GOOD),
he Heal h, Aging, and Body Composi ion s udy (Heal h ABC), he Indiana
Women coho , he No h Finland Bi h Coho 1966 (NFBC1966), he Old O de
Amish S udy (OOA), he Ro e dam S udy (RS), and he TwinsUK), and an
addi ional 21 coho s we e included o he cu en analysis ( he Alpha-Toco-
phe ol, Be a-Ca o ene Cance P e en ion S udy (ATBC), he A he oscle osis Risk
in Communi ies S udy (ARIC), he A he oGene egis y, B- i amins o he P e-
en ion O Os eopo o ic F ac u es (B-PROOF), he Epidemiology o Diabe es
In e en ions and Complica ions (EDIC), he Case-Con ol S udy o Me abolic
Synd ome (GenMe s), he Helsinki Bi h Coho S udy (HBCS), he Heal h P o-
essional Follow-up S udy (HPFS, nes ed a co ona y hea disease case-con ol
s udy), he In ecchia e in Chian i S udy (InChian i), he Coope a i e Heal h
Resea ch in he egion Augsbu g (KORA), he Leiden Longe i y S udy (LLS), he
Ludwigsha en Risk and Ca dio ascula Heal h S udy (LURIC), he Mul i-E hnic
S udy o A he oscle osis (MESA), he Nijmegen Biomedische S udie (NBS), he
Nu ses’Heal h S udy (NHS, nes ed a b eas cance case-con ol s udy, and a ype2
diabe es case-con ol s udy), he O kney Complex Disease S udy (ORCADES), he
P os a e, Lung, Colo ec al, and O a ian Cance Sc eening T ial (PLCO), he
PROspec i e S udy o P a as a in in he Elde ly a Risk (PROSPER), he S udy o
Heal h in Pome ania (SHIP), he Sco ish Colo ec al Cance S udy (SOCCS), he
Ca dio ascula Risk in Young Finns S udy (YFS), and mo e samples om he RS
(RSI, RSII, and RSIII)). Full desc ip ions o all pa icipa ing coho s, de ails o
geno yping pla o ms used, numbe o SNPs, and he measu emen s o se um 25-
hyd oxy i amin D concen a ions in each coho a e shown in Supplemen a y
Table 1and Supplemen a y No e 1. W i en in o med consen was ob ained om
all pa icipan s in he included coho s, and he s udy p o ocols we e e iewed and
app o ed by local ins i u ional e iew boa ds.
Powe calcula ion. Ou la ge sample size p o ided good s a is ical powe o
associa ion analysis. A he genome-wide significance h eshold o 5×10−8, wi h a
disco e y sample size o 75,000, ou s udy had 85% powe o de ec a gene ic
a ian (single nucleo ide polymo phism, SNP) accoun ing o 0.06% o he o al
a iance in se um 25-hyd oxy i amin D concen a ions, and 99% powe o de ec a
a ian ha explained 0.1% o he o al a iance. We also had powe o de ec gene-
en i onmen in e ac ion e ec s e en smalle han he obse ed ma ginal e ec s. In
he case whe e a SNP has no ma ginal e ec on ci cula ing 25-Hyd oxy i amin D
concen a ions (and so could no ha e been disco e ed ia he ma ginal GWAS),
we had 80% powe o de ec an in e ac ion ha explained 0.07% o he o al
a iance in 25-hyd oxy i amin D concen a ions.
Associa ion analysis. Genome-wide analyses we e pe o med wi hin each coho
acco ding o a uni o m analysis plan. We fi addi i e gene ic models using linea
eg ession on na u al-log ans o med 25-hyd oxy i amin D, and adjus ed he
models o mon h o sample collec ion (12 ca ego ies), age, sex, and body mass
index, and p incipal componen s cap u ing gene ic ances y. Fu he adjus men s
included coho -specific a iables, such as geog aphical loca ion and assay ba ch,
whe e ele an . Fo pa icipa ing s udies wi h a case-con ol design, we analyzed
cases and con ols sepa a ely. We pe o med a fixed-e ec s in e se a iance
weigh ed me a-analysis ac oss he con ibu ing coho s, as implemen ed in he
so wa e METAL26, wi h con ol o popula ion s uc u e wi hin each coho and
quali y con ol h esholds o mino allele equency (MAF) >0.05, impu a ion in o
sco e >0.8, Ha dy-Weinbe g equilib ium (HWE) >1×10−6, and a minimum o
wo s udies and 10,000 indi iduals con ibu ing o each epo ed SNP-pheno ype
associa ion. We ega ded P- alues <5×10−8as genome-wide significan .
Replica ion s udy. We eplica ed he iden ified no el loci in wo independen da a
se s o which geno ype da a we e a ailable: he Eu opean P ospec i e In es iga ion
in o Cance and Nu i ion (EPIC) s udy wi h 40,562 indi iduals ac oss wo nes ed
case-con ol s udies (EPIC-In e Ac and EPIC-CVD) and he coho -wide EPIC-
No olk s udy (Supplemen a y No e 1); and a coho o 2195 indi iduals (all
con ols) addi ionally collec ed as pa o he SOCCS ha we e no included in ou
disco e y s age. As o he pheno ype, EPIC indi iduals we e assayed o plasma
25-hyd oxy i amin D
3
and SOCCS indi iduals we e assayed o o al 25-
hyd oxy i amin D. We pe o med he associa ion analysis in a simila manne ,
adjus ed o age, sex, ime o sample collec ion, and s udy cen e whe e ele an .
We ega ded P- alue <0.05 in he eplica ion samples, and P- alue <5×10−8in he
pooled analysis as success ul eplica ion.
Condi ional analysis. A e iden i ying he p ima y associa ed a ian a each locus
selec ed acco ding o he s eng h o i s associa ion, we u he es ed whe he he e
we e any o he SNPs significan ly associa ed wi h 25-hyd oxy i amin D a e
accoun ing o he e ec o lead SNP. We hus pe o med a s epwise model
selec ion p ocedu e o hose ch omosomes whe e a significan a ian was p e-
iously iden ified. We s a ed wi h he mos significan ly associa ed SNP, scanning
h ough he whole ch omosome, selec ing addi ional independen ly associa ed
SNPs using a s epwise p ocedu e, one a a ime, based on hei condi ional P-
alues. Finally, we fi all selec ed SNP in o one model o es ima e hei join e ec s.
NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2 ARTICLE
NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions 7
We used GCTA-COJO so wa e o accommoda e ou summa y le el GWAS
da a27, and he Cance Gene ic Ma ke o Suscep ibili y (CGEMS) GWAS wi h
2287 indi iduals o Eu opean descen and 2,543,887 geno yped and impu ed
(HapMap22) SNPs as e e ence panel.
SNP-by-die in e ac ion. We pe o med a genome-wide associa ion sc eening o
ci cula ing 25-hyd oxy i amin D while accoun ing o po en ial in e ac ion e ec
be ween SNP and die a y i amin D in ake. Ou es s inco po a ing gene-die
in e ac ion we e based on he ollowing model:
ln 25 OH
ðÞ
D
ðÞ
¼β0þβ1´Gþβ2´Eþβ3´G´EþβZ´Z
whe e Gis a SNP ha was coded addi i ely, Eis he aw i amin D in ake,
measu ed on a con inuous scale. The pa ame e s β0,β1,β2, and β3a e he
in e cep , he main e ec o SNP, he main e ec o die a y i amin D in ake, and
he in e ac ion e ec be ween Gand E. The model also included he same
co a ia es Zas o he ma ginal e ec sc eening, he e ec s o which we e cap u ed
in he pa ame e βZ. We conside ed bo h a s anda d 1 deg ee-o - eedom es o
in e ac ion e ec (i.e., null hypo hesis o β3¼0), and a join 2 deg ee-o - eedom
es o main gene ic e ec plus gene-by-die in e ac ion (i.e., null hypo hesis o
β1¼0 and β3¼0). Fo compa ison pu poses, we also conside ed a model
adjus ing o i amin D in ake bu no modeling in e ac ion (i.e., no including he
β3´G´E e m) using he same subse o indi iduals.
Vi amin D in ake was a ailable o 15 coho s on a o al o 41,981 indi iduals. I
included bo h he popula ion based coho s (ARIC, 1958BC, B-PROOF, FHS,
Heal h ABC, MESA, NFBC, RS, RS III, and YFS as pa o he o e all Me a-GWAS,
plus wo addi ional coho s, he P ospec i e In es iga ion o he Vascula u e in
Uppsala Senio s (PIVUS), and he Uppsala Longi udinal S udy o Adul Men
(ULSAM), geno yped on cus om a ay ha we e no included in he o e all me a-
GWAS bu we e included in his SNP-by-die in e ac ion analysis), and case-con ol
s udies (HPFS (HPFS_CHD), NHS (NHS_BRCA, NHS_T2D), and SOCCS). Fo he
la e s udies, all analyses we e pe o med sepa a ely in cases and con ols. The
a o emen ioned in e ac ion model was applied o each o he included coho s, and
s udy-specific esul s we e me a-analyzed using in e se- a iance weigh ed sum o
e ec es ima es as implemen ed in METAL26. Fo he 2 deg ee-o - eedom es we
used join amewo k desc ibed in wo p e ious published pape s28,29. Quali y
con ol fil e ing was pe o med on each s udy be o e me a-analysis. Only SNPs wi h
impu a ion in o sco e >0.8, MAF >0.05, HWE >1×10−6, and a minimum o o al
sample size in he me a-analysis >10,000 we e e ained.
Linkage Disequilib ium sco e eg ession. We pe o med linkage disequilib ium
sco e eg ession (LDSC) analysis o es ima e he SNP-he i abili y o se um 25-
hyd oxy i amin D concen a ions30,31. This me hod is based on a alida ed ela-
ionship be ween LD sco e and χ2-s a is ics:
Eχ2
j
hi
Njh2
g
Mljþ1
whe e Eχ2
j
hi
deno es he expec ed χ2-s a is ics o he associa ion be ween
ou come and SNP j,N
j
is he s udy sample size a ailable o SNP j, M is he o al
numbe s o a ian s and ljdeno es he LD sco e o SNP j defined as lj¼P
k
2ðj;kÞ.
LDSC calcula es he i abili y using only summa y-le el da a ins ead o indi idual
geno ypes, and is compu a ionally cos e ec i e a la ge sample sizes. We used he
summa y s a is ics om 25-hyd oxy i amin D me a-GWAS esul s, wi h SNPs
a ailable in a leas 2 s udies and a sample size o a leas 10,000. We fi s analyzed
he SNP-he i abili y by using (1) SNPs ac oss he en i e genome ha passed quali y
con ol; (2) SNPs excluding he op associa ions (SNPs eaching genome-wide
significance, P≤5×10−8), as well as all SNPs wi hin ±500 kb o he op hi s in he
egion; and (3) SNPs excluding he nominally significan associa ions (P≤0.05). We
subsequen ly pa i ioned he he i abili y h ough h ee di e en models: (1) a ull
baseline model including he 24 publicly a ailable main anno a ions ha a e no
specific o any cell ype, he 500-bp windows a ound each anno a ion, as well as
100-bp windows a ound ch oma in immunop ecipi a ion and sequencing peaks
(ChIP-seq) when app op ia e. This esul ed in a o al o 52 o e lapping unc ional
ca ego ies in he ull baseline model; (2) a cell- ype-specific model including 10 cell
ype g oups: ad enal and panc eas, cen al ne ous sys em (CNS), ca dio ascula ,
connec i e and bone, gas oin es inal, immune and hema opoie ic, kidney, li e ,
skele al muscle, and o he ; (3) a cell- ype-specific model including 220 cell- ype-
specific anno a ions o he ou his one ma ks wi h pu a i e enhance o p omo e
unc ions, H3K4me1, H3K4me3, H3K9ac, and H3K27ac.
De ails o he 24 publicly a ailable anno a ions, he 220 cell- ype-specific
anno a ions, as well as he 10 cell ype g oups we e desc ibed by Finucane e al.31.
B iefly, he 24 anno a ions included coding, UTR (3′UTR and 5′UTR), p omo e
and in onic egions, acqui ed om he UCSC Genome B owse 32 and pos -
p ocessed by Guse e al.33; he h ee his one ma ks (mono-me hyla ion
(H3K4me1) o his one H3 a lysine 4, i-me hyla ion (H3K4me3) o his one H3 a
lysine 4, and ace yla ion o his one H3 a lysine 9 (H3K9ac) p ocessed by T ynka
e al.34–36 and wo e sions o ace yla ion o his one H3 a lysine 27 (H3K27ac, one
e sion p ocessed by Hnisz e al.37, ano he e sion used by he Psychia ic
Genomics Conso ium (PGC)38); open ch oma in, as eflec ed by DNase I
hype sensi i i y si es (DHSs and e al DHSs)33, ob ained as a combina ion o
Encyclopedia o DNA Elemen s (ENCODE) and Roadmap Epigenomics da a, and
p ocessed by T ynka e al.36; combined ch omHMM and Segway p edic ions
ob ained om Ho man e al.39, which le e age on many anno a ions o pa i ion
he genome in o se en unde lying ch oma in s a es ( he CCCTC-binding ac o
(CTCF), p omo e -flanking, ansc ibed egion, ansc ip ion s a si e (TSS),
s ong enhance , weak enhance , and he ep essed egion); egions ha a e
conse ed in mammals, p o ided by Lindblad-Toh e al.40 and pos -p ocessed by
Wa d and Kellis41; supe -enhance s, which a e la ge g oups o pu a i e enhance s
wi h high le els o ac i i y, p o ided by Hnisz e al.37; FANTOM5 enhance s
mapped by using cap analysis o gene exp ession in he FANTOM5 panel o
samples, ob ained om Ande sson e al.42; digi al genomic oo p in (DGF) and
ansc ip ion ac o binding si e (TFBS) anno a ions downloaded om ENCODE35
and pos -p ocessed by Guse e al.33. We included 500-bp windows a ound each o
he 24 main anno a ions in he baseline model, and 100-bp windows a ound ChIP-
seq when app op ia e, o p e en upwa d bias o es ima es gene a ed by
en ichmen in he nea by egions.
In addi ion o he baseline model using 24 main anno a ions, we also pe o med
cell- ype-specific analyses using anno a ions o he ou his one ma ks (H3K4me1,
H3K4me3, H3K9ac and H3K27ac). Each cell- ype-specific anno a ion co esponds
o a his one ma k in a single cell ype ( o example, H3K27ac in adipose nuclei
issues), and he e was a o al o 220 such anno a ions. We u he subdi ided hese
220 cell- ype-specific anno a ions in o 10 ca ego ies by agg ega ing he cell- ype-
specific anno a ions wi hin each g oup ( o example, SNPs ela ed wi h any o he
ou his one modifica ions in any hema opoie ic and immune cells we e conside ed
as one big ca ego y). When gene a ing he cell- ype-specific models, we added each
anno a ion indi idually (one a a ime) o he baseline model, c ea ing sepa a e
models o con ol o o e lap wi h he genomic unc ional elemen s in he ull
baseline model bu no o e lap wi h he o he cell ypes.
We addi ionally assembled he summa y s a is ics om GWAS o 37 ai s o
diseases pe o med in indi iduals o Eu opean descen , which a e publicly
a ailable38,43–55 o applied om he UK Biobank. These s udies span a wide ange
o pheno ypes, om an h opome ic indices such as heigh , weigh , BMI, o men al
diso de s ( o example dep essi e synd ome and schizoph enia) o au oimmune
and inflamma o y diseases ( o example heuma oid a h i is and celiac diseases).
We calcula ed he pai wise gene ic co ela ion ( g, c oss ai he i abili y) be ween
25-hyd oxy i amin D and each o he 37 ai s. We u he conduc ed he same
cell- ype-specific analysis o each ai , and plo ed be a-coe ficien z-sco e ma ix,
cons uc ed om he o al 220 anno a ions by 37 ai s, in o ou hea -maps based
on he ou his one ma ks.
Finally, in addi ion o he gene ic co ela ion analysis which eflec s sha ed
gene ic ac o s ac oss di e en ai s bu does no in o m di ec ion, we also
a emp ed o iden i y di ec ions o such co ela ion using an algo i hm p oposed by
Pick ell e al.56. The me hod adop s a simila in ui ion as he Mendelian
Randomiza ion app oach, whe e, i a ai X influences ai Y, hen SNPs
influencing X should also influence Y, and he SNP-specific e ec sizes o he wo
ai s should be co ela ed. Fu he , since Y does no influence X, bu could be
influenced by mechanisms independen o X, gene ic a ian s ha influence Y do
no necessa ily influence X. Based on his in ui ion, he me hod p oposes wo
“causal”models and wo “non-causal”models, and calcula es he ela i e likelihood
a io o he bes non-causal model compa ed o he bes causal model. We
de e mined significan SNPs o each gi en ai by selec ed genome-wide
significan (P<5×10−8) SNPs and p uned he numbe s based on hei LD-pa e n
in he Eu opean popula ions in Phase1 o 1000 Genome P ojec . We scanned
h ough all pai s o 25-hyd oxy i amin D and ai s o iden i y di ec ional
co ela ions. We conside pai s o ai s wi h likelihood a io
non-causal s. causal
<
0.05 as ha ing e idence o di ec ional co ela ions.
Da a a ailabili y. The GWAS summa y s a is ics on se um ci cula ing i amin D
concen a ions is a ailable a dbGap h ps://d i e.google.com/d i e/ olde s/
0BzYD Co_doHJRFRKR0l ZHZWZjQ; all ele an da a a e a ailable om he
au ho s upon eques .
Recei ed: 21 July 2017 Accep ed: 15 Decembe 2017
Re e ences
1. Mok y, L. E. e al. Vi amin D and isk o mul iple scle osis: a mendelian
andomiza ion s udy. PLoS Med. 12, e1001866 (2015).
2. Rhead, B. e al. Mendelian andomiza ion shows a causal e ec o low i amin
D on mul iple scle osis isk. Neu ol. Gene . 2, e97 (2016).
3. Ma ineau, A. R. e al. Vi amin D supplemen a ion o p e en acu e espi a o y
ac in ec ions: sys ema ic e iew and me a-analysis o indi idual pa icipan
da a. B . Med. J. 356, i6583 (2017).
ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2
8NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions
4. Pilz, S., Ve heyen, N., G üble , M. R., Tomaschi z, A. & Mä z, W. Vi amin D
and ca dio ascula disease p e en ion. Na . Re . Ca diol. 13, 404–417 (2016).
5. Ga land, C. F. e al. The ole o i amin D in cance p e en ion. Am. J. Public
Heal h 96, 252–261 (2006).
6. Fe nandes de Ab eu, D. A., Eyles, D. & Fé on, F. Vi amin D, a neu o-
immunomodula o : implica ions o neu odegene a i e and au oimmune
diseases. Psychoneu oendoc inology 34 Suppl 1, S265–277 (2009).
7. Laguno a, Z., Po ojnicu, A. C., Lindbe g, F., Hexebe g, S. & Moan, J. The
dependency o i amin D s a us on body mass index, gende , age and season.
An icance Res. 29, 3713–3720 (2009).
8. Wang, T. J. e al. Common gene ic de e minan s o i amin D insu ficiency: a
genome-wide associa ion s udy. Lance Lond. Engl. 376, 180–188 (2010).
9. Ka ohl, C. e al. He i abili y and seasonal a iabili y o i amin D
concen a ions in male wins. Am. J. Clin. Nu . 92, 1393–1398 (2010).
10. O on, S.-M. e al. E idence o gene ic egula ion o i amin D s a us in wins
wi h mul iple scle osis. Am. J. Clin. Nu . 88, 441–447 (2008).
11. Ahn, J. e al. Genome-wide associa ion s udy o ci cula ing i amin D le els.
Hum. Mol. Gene . 19, 2739–2745 (2010).
12. Hi aki, L. T. e al. Explo ing he gene ic a chi ec u e o ci cula ing 25-
hyd oxy i amin D. Gene . Epidemiol. 37,92–98 (2013).
13. Boyadjie , S. e al. C anio-len iculo-su u al dysplasia associa ed wi h de ec s in
collagen sec e ion. Clin. Gene . 80, 169–176 (2011).
14. Boyadjie , S. A. e al. C anio-len iculo-su u al dysplasia is caused by a SEC23A
mu a ion leading o abno mal endoplasmic- e iculum- o-Golgi a ficking. Na .
Gene . 38, 1192–1197 (2006).
15. Myung, J. K. e al. Well-di e en ia ed liposa coma o he oesophagus:
clinicopa hological, immunohis ochemical and a ay CGH analysis. Pa hol.
Oncol. Res. 17, 415–420 (2011).
16. Manousaki, D. e al. Low- equency synonymous coding a ia ion in CYP2R1
has la ge e ec s on i amin D le els and isk o mul iple scle osis. Am. J. Hum.
Gene . 101, 227–238 (2017).
17. A guelles, L. M. e al. He i abili y and en i onmen al ac o s a ec ing i amin
D s a us in u al Chinese adolescen wins. J. Clin. Endoc inol. Me ab. 94,
3273–3281 (2009).
18. Mills, N. T. e al. He i abili y o ans o ming g ow h ac o -β1 and umo
nec osis ac o - ecep o ype 1 exp ession and i amin D le els in heal hy
adolescen wins. Twin Res. Hum. Gene . O . J. In . Soc. Twin S ud. 18,28–35
(2015).
19. Li shi s, G., Ka asik, D. & Seibel, M. J. S a is ical gene ic analysis o plasma
le els o i amin D: amilial s udy. Ann. Hum. Gene . 63, 429–439 (1999).
20. Yu, H.-J., Kwon, M.-J., Woo, H.-Y. & Pa k, H. Analysis o 25-hyd oxy i amin
D s a us acco ding o age, gende , and seasonal a ia ion. J. Clin. Lab. Anal. 30,
905–911 (2016).
21. Hun e , D. e al. Gene ic con ibu ion o bone me abolism, calcium exc e ion,
and i amin D and pa a hy oid ho mone egula ion. J. Bone Mine . Res. O . J.
Am. Soc. Bone Mine . Res. 16, 371–378 (2001).
22. Agmon-Le in, N., Theodo , E., Segal, R. M. & Shoen eld, Y. Vi amin D in
sys emic and o gan-specific au oimmune diseases. Clin. Re . Alle gy Immunol.
45, 256–266 (2013).
23. Yang, C.-Y., Leung, P. S. C., Adamopoulos, I. E. & Ge shwin, M. E. The
implica ion o i amin D and au oimmuni y: a comp ehensi e e iew. Clin.
Re . Alle gy Immunol. 45, 217–226 (2013).
24. Ascha d, H. A pe spec i e on in e ac ion e ec s in gene ic associa ion s udies.
Gene . Epidemiol. 40, 678–688 (2016).
25. Coope , J. D. e al. Inhe i ed a ia ion in i amin D genes is associa ed wi h
p edisposi ion o au oimmune disease ype 1 diabe es. Diabe es 60, 1624–1631
(2011).
26. Wille , C. J., Li, Y. & Abecasis, G. R. METAL: as and e ficien me a-analysis o
genomewide associa ion scans. Bioin o ma. Ox . Engl. 26, 2190–2191
(2010).
27. Yang, J. e al. Condi ional and join mul iple-SNP analysis o GWAS summa y
s a is ics iden ifies addi ional a ian s influencing complex ai s. Na . Gene .
44, 369–375 (2012).
28. Ascha d, H., Hancock, D. B., London, S. J. & K a , P. Genome-wide me a-
analysis o join es s o gene ic and gene-en i onmen in e ac ion e ec s.
Hum. He ed. 70, 292–300 (2010).
29. Manning, A. K. e al. Me a-analysis o gene-en i onmen in e ac ion: join
es ima ion o SNP and SNP×en i onmen eg ession coe ficien s. Gene .
Epidemiol. 35,11–18 (2011).
30. Bulik-Sulli an, B. K. e al. LD sco e eg ession dis inguishes con ounding om
polygenici y in genome-wide associa ion s udies. Na . Gene . 47, 291–295
(2015).
31. Finucane, H. K. e al. Pa i ioning he i abili y by unc ional anno a ion using
genome-wide associa ion summa y s a is ics. Na . Gene . 47, 1228–1235 (2015).
32. Ken , W. J. e al. The human genome b owse a UCSC. Genome Res. 12,
996–1006 (2002).
33. Guse , A. e al. Pa i ioning he i abili y o egula o y and cell- ype-specific
a ian s ac oss 11 common diseases. Am. J. Hum. Gene . 95, 535–552 (2014).
34. Roadmap Epigenomics Conso ium. e al. In eg a i e analysis o 111 e e ence
human epigenomes. Na u e 518, 317–330 (2015).
35. ENCODE P ojec Conso ium. An in eg a ed encyclopedia o DNA elemen s in
he human genome. Na u e 489,57–74 (2012).
36. T ynka, G. e al. Ch oma in ma ks iden i y c i ical cell ypes o fine mapping
complex ai a ian s. Na . Gene . 45, 124–130 (2013).
37. Hnisz, D. e al. Supe -enhance s in he con ol o cell iden i y and disease. Cell
155, 934–947 (2013).
38. Schizoph enia Wo king G oup o he Psychia ic Genomics Conso ium.
Biological insigh s om 108 schizoph enia-associa ed gene ic loci. Na u e 511,
421–427 (2014).
39. Ho man, M. M. e al. In eg a i e anno a ion o ch oma in elemen s om
ENCODE da a. Nucleic Acids Res. 41, 827–841 (2013).
40. Lindblad-Toh, K. e al. A high- esolu ion map o human e olu iona y
cons ain using 29 mammals. Na u e 478, 476–482 (2011).
41. Wa d, L. D. & Kellis, M. E idence o abundan pu i ying selec ion in humans
o ecen ly acqui ed egula o y unc ions. Science 337, 1675–1678 (2012).
42. Ande sson, R. e al. An a las o ac i e enhance s ac oss human cell ypes and
issues. Na u e 507, 455–461 (2014).
43. Bo aska, V. e al. A genome-wide associa ion s udy o ano exia ne osa. Mol.
Psychia y 19, 1085–1094 (2014).
44. Global Lipids Gene ics Conso ium. e al. Disco e y and efinemen o loci
associa ed wi h lipid le els. Na . Gene . 45, 1274–1283 (2013).
45. Ben ham, J. e al. Gene ic associa ion analyses implica e abe an egula ion o
inna e and adap i e immuni y genes in he pa hogenesis o sys emic lupus
e y hema osus. Na . Gene . 47, 1457–1464 (2015).
46. Okada, Y. e al. Gene ics o heuma oid a h i is con ibu es o biology and
d ug disco e y. Na u e 506, 376–381 (2014).
47. Okbay, A. e al. Gene ic a ian s associa ed wi h subjec i e well-being,
dep essi e symp oms, and neu o icism iden ified h ough genome-wide
analyses. Na . Gene . 48, 624–633 (2016).
48. Jos ins, L. e al. Hos -mic obe in e ac ions ha e shaped he gene ic a chi ec u e
o inflamma o y bowel disease. Na u e 491, 119–124 (2012).
49. C oss-Diso de , G oup o he Psychia ic Genomics Conso ium Iden ifica ion
o isk loci wi h sha ed e ec s on fi e majo psychia ic diso de s: a genome-
wide analysis. Lance Lond. Engl. 381, 1371–1379 (2013).
50. Co dell, H. J. e al. In e na ional genome-wide me a-analysis iden ifies new
p ima y bilia y ci hosis isk loci and a ge able pa hogenic pa hways. Na .
Commun. 6, 8019 (2015).
51. Schunke , H. e al. La ge-scale associa ion analysis iden ifies 13 new
suscep ibili y loci o co ona y a e y disease. Na . Gene . 43, 333–338
(2011).
52. Mo is, A. P. e al. La ge-scale associa ion analysis p o ides insigh s in o he
gene ic a chi ec u e and pa hophysiology o ype 2 diabe es. Na . Gene . 44,
981–990 (2012).
53. Psychia ic GWAS Conso ium Bipola Diso de Wo king G oup. La ge-scale
genome-wide associa ion analysis o bipola diso de iden ifies a new
suscep ibili y locus nea ODZ4. Na . Gene . 43, 977–983 (2011).
54. Dubois, P. C. A. e al. Mul iple common a ian s o celiac disease influencing
immune gene exp ession. Na . Gene . 42, 295–302 (2010).
55. Tho gei sson, T. E. e al. Sequence a ian s a CHRNB3-CHRNA6 and
CYP2A6 a ec smoking beha io . Na . Gene . 42, 448–453 (2010).
56. Pick ell, J. K. e al. De ec ion and in e p e a ion o sha ed gene ic influences on
42 human ai s. Na . Gene . 48, 709–717 (2016).
Acknowledgemen s
A ull lis o acknowledgemen s can be ound in Supplemen a y No e 2.
Au ho con ibu ions
C.P., E.H., M.I.M., E.A.S., P.L.L., E.B., D.A., S.J.W., N.D.F., W.H., L.C.P.G.M.D.G.,
N.M.V.S., N. .V., J.B.R., B.K., T.J.W., D.P.K., R.S.V., C.O., M.L., J.G.E., S.B.K., M.B., E.S.,
I.H.D.B., J.I.R., S.S.R., M.J., P.K., J.F.W., L.L., K.M., L.Z., H.C., E.T., S.M.F., M.G.D., T.L.,
M.K., O.T.R., V.M., M.A.I., J.W.J., N.J.W., C.L., N.G.F., K.K., A.B., and J.D. designed and
managed indi idual s udies. C.P., E.H., M.I.M., E.A.S., P.L.L., D.A., S.J.W., W.H.,
N.M.V.S., A.E., J.B.R., R.S.V., C.O., L.V., M.L., J.G.E., S.B.K., M.B., D.V.H., I.H.D.B.,
J.I.R., S.S.R., J.F.W., L.L., E.I., K.M., H.C., E.T., S.M.F., M.G.D., T.L., M.K., O.T.R., V.M.,
M.A.I., N.S., J.W.J., N.J.W., C.L., N.G.F., K.K., A.B., and J.D. collec ed da a. E.B., A.E.,
C.O., M.L., Y.L., M.B., E.M.K., J.I.R., S.S.R., J.F.W., L.L., E.I., K.M., S.T., S.M.F., M.G.D.,
T.L., L.L., J.W.J., N.J.W., C.L., A.B., and J.D. pe o med he geno yping. C.P., D.B., E.H.,
M.I.M., W.T., E.B., N.M.V.S., A.E., J.D., C.O., L.V., M.L., K.K.L., J.D., E.M.K., J.I.R.,
S.S.R., J.F.W., C.H., E.I., K.M., S.T., A.G.U., F.R., L.Z., S.M.F., M.G.D., A.M.V., T.L., L.L.,
M.A.I., N.S., N.J.W., C.L., A.B., and J.D. p epa ed he geno ype da a. D.B., E.H., M.I.M.,
E.A.S., P.L.L., A.E., J.B.R., S.B., R.S.V., C.O., L.V., M.L., J.G.E., D.K.H., D.V.H., E.M.K.,
I.H.D.B., A.C.W., J.F.W., L.L., K.M., M.C.Z., A.G.U., F.R., L.Z., E.T., S.M.F., M.G.D.,
A.M.V., E.T., L.L., N.J.W., C.L., N.G.F., A.B., and J.D. p epa ed he pheno ype da a. D.B.,
E.H., M.I.M., J.B.R., T.J.W., D.P.K., Y.H., C.L., A.C.W., C.R.C., P.F.O., M.J., X.J., H.A.,
N.J.W., C.L., N.G.F., A.B., and J.D. de eloped he analysis plan. A.Z., D.B., E.H., M.I.M.,
NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2 ARTICLE
NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions 9