scieee Open visual document viewer

Genome-wide association study in 79,366 European-ancestry individuals informs the genetic architecture of 25-hydroxyvitamin D levels

Jiang, Xia,O´Reilly, Paul F,Aschard, Hugues,Lehtimäki, Terho,Kähönen, Mika,Lyytikäinen, Leo-Pekka

Full text

ARTICLE Genome-wide associa ion s udy in 79,366 Eu opean-ances y indi iduals in o ms he gene ic a chi ec u e o 25-hyd oxy i amin D le els Xia Jiang e al. # Vi amin D is a s e oid ho mone p ecu so ha is associa ed wi h a ange o human ai s and diseases. P e ious GWAS o se um 25-hyd oxy i amin D concen a ions ha e iden ified ou genome-wide significan loci (GC, NADSYN1/DHCR7, CYP2R1, CYP24A1). In his s udy, we expand he p e ious SUNLIGHT Conso ium GWAS disco e y sample size om 16,125 o 79,366 (all Eu opean descen ). This la ge GWAS yields wo addi ional loci ha bo ing genome-wide significan a ian s (P=4.7×10−9a s8018720 in SEC23A, and P=1.9×10−14 a s10745742 in AMDHD1). The o e all es ima e o he i abili y o 25-hyd oxy i amin D se um concen a ions a ibu able o GWAS common SNPs is 7.5%, wi h s a is ically significan loci explaining 38% o his o al. Fu he in es iga ion iden ifies signal en ichmen in immune and hema opoie ic issues, and clus e ing wi h au oimmune diseases in cell- ype-specific analysis. La ge s udies a e equi ed o iden i y addi ional common SNPs, and o explo e he ole o a e o s uc u al a ian s and gene–gene in e ac ions in he he i abili y o ci cula ing 25- hyd oxy i amin D le els. DOI: 10.1038/s41467-017-02662-2 OPEN Co espondence and eques s o ma e ials should be add essed o E.H. (email: [email p o ec ed]) o o P.K. (email: [email p o ec ed]) o o D.P.K. (email: [email p o ec ed]) #A ull lis o au ho s and hei a flia ions appea s a he end o he pape NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions 1 1234567890():,; Vi amin D is an essen ial a soluble i amin and s e oid p o-ho mone ha plays an impo an ole in muscu- loskele al heal h. Vi amin D deficiency has been linked o au oimmune1,2and in ec ious disease3, ca dio ascula disease4, cance 5, and neu odegene a i e condi ions6. Se um 25- hyd oxy i amin D, a p ima y ci cula ing o m o i amin D and a measu e ha bes eflec s i amin D s o es, is influenced by many ac o s including sun exposu e, age, body mass index7, die a y in ake o ce ain oods such as o ified dai y p oduc s and oily fish, supplemen s, and gene ic ac o s8. The concen a ion o 25-hyd oxy i amin D has been epo ed o be highly he i able, wi h he i abili y es ima es o 50–80% om classical win s udies9,10. A genome-wide associa ion s udy (GWAS) me a-analysis o se um 25-hyd oxy i amin D11 in 4501 pa icipan s o Eu opean ances y and eplica ion in 2221 samples iden ified a ian s in h ee loci (g oup componen (GC), 7-dehyd ochles e ol educ ase (NADSYN1/DHCR7), and 25-hyd oxylase (CYP2R1)). A la ge GWAS conduc ed by he SUNLIGHT conso ium in 16,125 Eu opean ances y indi iduals, wi h a eplica ion sample o 17,871, eplica ed hese h ee loci and disco e ed one addi ional locus (CYP24A1)8. Howe e , despi e hese loci being in o nea genes encoding p o eins in ol ed in i amin D syn hesis, he associa ed a ian s collec i ely explain only a small ac ion o he a iance in 25-hyd oxy i amin D concen a ions (~5%)8,11,12. The e o e, o ex end ou p e ious findings and be e unde s and he gene ic a chi ec u e unde lying se um 25-hyd oxy i amin D, as well as es o in e ac ions be ween die a y i amin D in ake and gene ic ac o s, we conduc ed a la ge-scale GWAS me a- analysis on his impo an i amin. Ou GWAS wi h a 79,366 disco e y sample and a 40,562 eplica ion sample eplica es ou p e ious loci and iden ifies wo new gene ic loci o se um le els o 25-hyd o xy i amin D. We u he find e idence o a sha ed gene ic basis be ween ci cu- la ing 25-hyd oxy i amin D and au oimmune diseases. Ou analyses sugges a ela i ely modes SNP-he i abili y a e o 25- hyd oxy i amin D when conside ing only common a ian s. La ge s udies a e equi ed o iden i y addi ional common SNPs, and o explo e he ole o a e o s uc u al a ian s. The gene ic ins umen s iden ified by ou esul s could be used in u u e Mendelian Randomiza ion analyses o he associa ion be ween i amin D and complex ai s. Resul s S udy desc ip ion and GWAS. This s udy ep esen s an expan- sion o ou p e ious SUNLIGHT conso ium GWAS8. He e, we combine he 5 disco e y coho s and 5 in-silico eplica ion coho s om ha s udy, and augmen hese wi h 21 addi ional coho s ha ha e joined he SUNLIGHT conso ium since 2010 (s udy cha ac e is ics a e desc ibed in Supplemen a y Table 1, Supplemen a y No e 1). In con as o he p e ious me a-analysis which in ol ed disco e y, in-silico and de-no o geno yping s ages, we pe o med a fi s s age disco e y me a-analysis on a o al o up o 79,366 indi iduals and eplica ed no el findings in wo independen sepa a e in-silico da a se s (40,562 indi iduals collec ed by EPIC and 2195 indi iduals collec ed by SOCCS). To assess and con ol o popula ion s a ifica ion, we examined QQ- plo s and genomic con ol infla ion ac o s o each con ibu ing coho p io o me a-analysis. We did no obse e e idence o widesp ead infla ion (median λ GC =0.92; only 1/31 samples wi h λ GC >1.01), indica ing ha ou GWAS esul s we e no infla ed by popula ion s a ifica ion o c yp ic ela edness (Supplemen a y Fig. 1). Despi e he sligh ly defla ed λ GC obse ed in some o he cons i uen coho s mos p obably due o o e -co ec ion o es s a is ics, ou λ GC o 0.99 in all samples indica ed app op ia e con ol o popula ion s a ifica ions and con ounde s. We iden ified six suscep ibili y loci ha bo ing genome-wide sig- nifican SNPs, confi ming ou p e iously epo ed loci a GC (P=4.7×10−343 a s3755967), NADSYN1/DHCR7 (P=3.8×10−62 a s12785878), CYP2R1 (P=2.1×10−46 a s10741657), CYP24A1 (P=8.1×10−23 a s17216707), and wo no el loci a AMDHD1 (P=1.9×10−14 a s10745742) and SEC23A (P=4.7×10−9a s8018720) (Table 1; Manha an plo s and QQ-plo s o o e all samples a e p esen ed in Fig. 1, and egional associa ion plo s a e p esen ed in Supplemen a y Fig. 2). The associa ions a bo h no el loci we e confi med in he wo independen in-silico eplica ion coho s (EPIC: P=1.21×10−8a s10745742, P= 5.24×10−4a s8018720; SOCCS: P=0.03 a s10745742, P=0.04 a s8018720) wi h consis en di ec ion o e ec bu sligh ly la ge e ec sizes and wide confidence in e als obse ed in SOCCS which could be due o he educed sample size. When analyzing he wo eplica ion da a se s oge he wi h he disco e y da a se , he P- alues in pooled samples became mo e significan (P pooled = 2.10×10−20 a s10745742, P pooled =1.11×10−11 a s8018720) (Table 1). We also ound mo e han one dis inc signal a ising om a ian s a he GC,CYP2R1, and AMDHD1 loci h ough condi ional and join analysis, bu no o he NADSYN1/DHCR7, CYP24A1, and SEC23A loci whe e only one p ima y associa ed SNP was iden ified (Supplemen a y Table 2). SNP by die a y i amin D in ake in e ac ion. In addi ion o pe o ming he ma ginal e ec me a-analysis using all samples, we also es ed a model wi h SNPs and die a y i amin D in ake as main e ec s and a e m o hei in e ac ion in a subse o samples. Die ques ionnai es, including i amin D in ake, we e a ailable o a subse o 13 coho s and an addi ional 2 coho s ha we e no included in he o e all me a-GWAS analysis (a o al o 15 coho s, N=41,981). We pe o med a GWAS explici ly allowing o an in e ac ion be ween i amin D in ake and SNP geno ypes, in which die a y i amin D was coded as a con inuous a iable. We pe o med wo es s: (i) a 1 deg ee-o - eedom in e ac ion es be ween each SNP and i amin D in ake, and (ii) a 2 deg ee-o - eedom join es o main gene ic and in e ac ion e ec s. Howe e , o compa ison pu poses, we also pe o med, (iii) a s anda d es o ma ginal gene ic e ec a e adjus ing o i amin D in ake in he same sub-samples (Supplemen a y Fig. 3, and Supplemen a y Fig. 4). The ma ginal gene ic e ec analyses confi med exis ing associa ion signals a GC (lead SNP s2282679, in comple e linkage disequilib ium wi h he lead SNP s3755967 iden ified om me a-GWAS using all indi iduals), CYP2R1, NADSYN1/DHCR7 (lead SNP s4944062, in comple e linkage disequilib ium wi h he lead SNP s12785878 iden ified om me a-GWAS using all indi iduals) and CYP24A1, as well as he no el associa ion a AMDHD1 (P=5.7×10−9, Supplemen a y Table 3). The join analysis also iden ified he fi e genes abo e, bu wi h less significan P- alues. Fo ins ance, he associa ion a AMDHD1 achie ed only sugges i e genome-wide significance (P =1.2×10−7, Table 2). The in e ac ion es did no iden i y any a ian s a genome-wide significance le el. Among he 5 SNPs significan in ma ginal e ec es s, he lead SNP in CYP2R1 showed nominal significance o in e ac ion wi h die a y i amin D in ake ( s10741657, P=0.028), bu no in e ac ions we e obse ed o he o he SNPs ( s2282679, P=0.45; s4944062, P= 0.74; s10745742, P=0.64; s17216707, P=0.46). Repea ing he analysis using a e ile coding o i amin D ins ead o a con- inuous coding did no quali a i ely change he esul s. SNP-he i abili y o 25-hyd oxy i amin D. We u he e alua ed he SNP-he i abili y, defined as he he i abili y explained by GWAS SNPs o 25-hyd oxy i amin D, using LD sco e eg ession ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2 2NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions (see Me hods). The o e all obse ed he i abili y o 25- hyd oxy i amin D es ima ed by using all common SNPs (2,579,296 a e QC) was 7.54% (s anda d e o (SE): 1.88%). A e excluding genome-wide significan SNPs (533 SNPs wi h P≤5×10−8) om six loci and all SNPs wi hin ±500 kb o hose loci, he he i abili y dec eased o 4.70% (SE: 0.72%). The es ima e u he dec eased o 1.73% (SE: 0.32%) a e excluding all SNPs ha eached nominal significance (156,675 SNPs wi h P≤0.05). These esul s indica e ha common a ian s agged by GWAS chips explain a modes ac ion o o e all a iabili y in ci cula ing 30 GC NADSYN1/DHCR7 CYP2R1 AMDHD1 SEC23A CYP24A1 Know No el 25 20 –log10 (p – alue) 15 10 5 200  = 0.99 01234 Expec ed –log10 co ec ed p- alue 56 Obse ed –log10 co ec ed p- alue 100 50 0 1 2 3 4 5 6 7 8 Ch omosome 9 10 11 12 13 14 15 16 17 18 19 20 21 22 a b Fig. 1 Genome-wide associa ion o ci cula ing 25-hyd oxy i amin D g aphed by ch omosome posi ions and −log10 P- alue (Manha an plo ), and quan ile- quan ile plo o all SNPs om he me a-analysis (QQ-plo ). aManha an plo : The P- alues we e ob ained om he single s age fixed-e ec s in e se a iance weigh ed me a-analysis. The Yaxis shows −log 10 P- alues, and he Xaxis shows ch omosome posi ions. Ho izon al g ay dash line ep esen s he h esholds o P=5×10−8 o genome-wide significance. Known loci we e colo ed coded as ed, and no el loci we e colo coded as g een. bQQ-plo : The Y axis shows obse ed −log 10 P- alues, and he Xaxis shows he expec ed −log 10 P- alues. Each SNP is plo ed as a black do , and he dash line indica es null hypo hesis o no ue associa ion. De ia ion om he expec ed P- alue dis ibu ion is e iden only in he ail a ea, wi h a lambda o 0.99, sugges ing ha popula ion s a ifica ion was adequa ely con olled Table 1 Single nucleo ide polymo phisms iden ified in genome-wide analyses o ci cula ing 25-hyd oxy i amin D concen a ions Gene SNP Ch omosome: Posi ion E ec / e e ence allele Allele equency Me a-GWAS es ima es E ec (Be a) S anda d E o P- alue Fi s s age disco e y me a-GWAS (N=79,366) GC s3755967 4:72828262 T/C 0.28 −0.089 0.0023 4.74E–343 NADSYN1/ DHCR7 s12785878 11:70845097 T/G 0.75 0.036 0.0022 3.80E–62 CYP2R1 s10741657 11:14871454 A/G 0.4 0.031 0.0022 2.05E–46 CYP24A1 s17216707 20:52165769 T/C 0.79 0.026 0.0027 8.14E–23 AMDHD1 s10745742 12:94882660 T/C 0.4 0.017 0.0022 1.88E–14 SEC23A s8018720 14:38625936 C/G 0.82 −0.017 0.0029 4.72E–09 Replica ion da a se 1: samples collec ed by EPIC (N=40,562) AMDHD1 s10745742 12:94882660 T/C 0.41 0.041 0.0071 1.21E–08 SEC23A s8018720 14:38625936 C/G 0.83 −0.032 0.0093 5.24E–04 Replica ion da a se 2: addi ional con ol samples collec ed by SOCCS (N=2195) AMDHD1 s10745742 12:94882660 T/C 0.37 0.045 0.021 0.03 SEC23A s8018720 14:38625936 C/G 0.81 −0.051 0.026 0.04 Pooled analysis (disco e y me a-GWAS + eplica ion 1 + eplica ion 2) (N=122,123) AMDHD1 s10745742 12:94882660 T/C 0.39 0.019 0.002 2.10E–20 SEC23A s8018720 14:38625936 C/G 0.82 −0.019 0.0027 1.11E–11 In he pooled analysis, P he e ogenei y =0.003 o AMDHD1 s10745742; P he e ogenei y = 0.14 o SEC23A s8018720. NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2 ARTICLE NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions 3 25-hyd oxy i amin D, and ha an app eciable p opo ion o his SNP-he i abili y is explained by he six gene ic egions o asso- cia ed SNPs iden ified h ough GWAS. Pa i ioning he o al he i abili y o 25-hyd oxy i amin D.We nex pa i ioned he he i abili y by unc ional elemen s using baseline model wi h 24 publicly a ailable anno a ions (see Me hods), and obse ed la ge and significan en ichmen o se e al unc ional ca ego ies (Fig. 2, Supplemen a y Table 4). Fo example, we ound he la ges en ichmen in weak enhance s, wi h 2.1% o SNPs explaining 42.3% o he o e all he i abili y (20- old en ichmen , P=0.02), ollowed by conse ed egions (13.8- old en ichmen , P=0.03), open ch oma in (as eflec ed by DHS, 8.5- old en ichmen , P=0.02), ansc ip ion ac o binding si es (5.7- old en ichmen , P=0.048), supe -enhance s (1.9- old en ichmen , P=0.04), and all ou his one ma ks we e en iched (bo h e sions o H3K27ac (one e sion p ocessed by Hnisz e al., and ano he e sion used by he Psychia ic Genomics Con- so ium (PGC), H3K4me1, H3K4me3 (500 bp), H3K29ac (500 bp)). We also obse ed deple ion o ep essed egions (0.06- old en ichmen , P=0.006). Howe e , none o hose anno a ions wi hs ood mul iple- es ing co ec ions (Bon e oni co ec ed P- h eshold: 0.05/24) excep o he ac i e enhance his one ma k H3K27ac (PGC) (4.2- old en ichmen , P=8×10−4) and H3K4me1 (1.8- old en ichmen , P=0.0019). We subsequen ly pe o med cell- ype-specific analysis by using 10 b oad cell- ype g oups. As shown in Table 3, he op h ee en ichmen s we e in he immune and hema opoie ic issues (4.3- old en ichmen , P=2.2×10−5), gas oin es inal issues (4.4- old en ichmen , P=0.0017), and CNS (3.6- old en ichmen , P= 0.0039). The e was also significan en ichmen o li e , kidney, and connec i e and bone issues, bu hese esul s did no su i e mul iple- es ing co ec ions. When u he analyzing 220 cell- ype-specific anno a ions, we obse ed he mos significan en ichmen in CD19 cells (app oxima ely 8- old en ichmen , P~0.001), ollowed by CD20 cells (6.4- old en ichmen , P=0.003) and CD3 cells (7.8- old en ichmen , P=0.01) (Supplemen a y Da a 1). Gene ic co ela ions be ween 25-hyd oxy i amin D and ai s. We con inued o assess he gene ic co ela ion be ween 25- hyd oxy i amin D and each o he 37 ai s wi h publicly a ailable GWAS summa y s a is ics da a (Supplemen a y Table 5). None o he gene ic co ela ions emained significan a e Bon e oni co ec ion (co ec ed P- h eshold: 0.05/37, Fig. 3). Wi hou mul iple- es ing co ec ion, he e we e some co ela ions wi h nominal s a is ical significance. Fo example, e e smoking ( g(SE): −0.17 (0.073), P=0.019), p ima y bilia y ci hosis ( g(SE): −0.18 (0.076), P=0.019) and BMI adjus ed wais -hip- a io ( g(SE): −0.10 (0.050), P=0.042) we e obse ed o be in e sely co ela ed wi h 25-hyd oxy i amin D; whe eas lung unc ion ( g(SE): 0.14 (0.046), P=0.0036) showed a posi i e co ela ion wi h 25-hyd oxy i amin D. Subsequen di ec ional gene ic co ela ion analysis did no e eal any appa en pu a i e causal ela ionship o 25-hyd o y i amin D wi h o he ai s, excep o a po en ial link be ween 25-hyd oxy i amin D and Table 2 Resul s om he SNP-by-die a y i amin D in ake in e ac ion analysis Gene SNP Ch omosome: Posi ion E ec / Re e ence Allele Allele F equency SNP-by-die a y i amin D in ake In e ac ion analysis Main Gene ic E ec In e ac ion E ec P- alue o in e ac ion P- alue o join es E ec (Be a_G) S anda d E o E ec (Be a_In ) S anda d E o Fi s s age disco e y me a-GWAS(N=79,366) GC s3755967 4:72828262 T/C 0.28 −0.082 0.0042 −2.01E–05 1.67E–05 0.23 2.92E–171 s2282679* 4:72827247 T/G 0.28 0.085 0.004 1.20E–05 1.60E–05 0.45 1.40E–187 NADSYN1/ DHCR7 s12785878 11:70845097 T/G 0.75 0.033 0.0039 6.61E–06 1.62E–05 0.68 3.52E–29 s4944062* 11:70864942 T/G 0.75 0.034 0.004 5.30E–06 1.60E–05 0.74 1.90E–31 CYP2R1 s10741657 11:14871454 A/G 0.4 0.03 0.0035 3.21E–05 1.46E–05 0.028 2.23E–38 CYP24A1 s17216707 20:52165769 T/C 0.79 0.025 0.0048 1.39E–05 1.88E–05 0.46 1.32E–14 AMDHD1 s10745742 12:94882660 T/C 0.4 0.016 0.0036 −7.05E–06 1.49E–05 0.64 1.20E–07 SEC23A s8018720 14:38625936 C/G 0.82 −0.013 0.0051 −2.40E–05 2.06E–05 0.24 1.94E–05 * Top SNPs iden ified in he SNP-by-die a y i amin D in ake in e ac ion analysis, pe o med in a subse o indi iduals. Fo GC and NADSYN1/DHCR7, he op SNPs iden ified h ough he ma ginal e ec eg ession me a-analysis using all indi iduals we e in high linkage disequilib ium wi h he op SNPs iden ified h ough he SNP-by-die a y i amin D in ake in e ac ion analysis using a subse o indi iduals ( 2 o s3755967 and s2282679: 1.0; 2 o s12785878 and s4944062: 1.0). Be a_G indica es he main e ec o he SNP, Be a_In indica es he in e ac ion e ec o SNP-by-die a y i amin D in ake 20 Type 24anno a ions 24anno a ions_500 bp En ichmen =P op.H2/P op.SNPs 10 0 Weak enhance Conse ed TSS Fe al DHS TFBS H3K27ac (PGC) H3K27ac H3K4me1 H3K4me3 Supe enhance H3K27ac (Hnisz) Rep essed Func ional ca ego ies ** ** Fig. 2 He i abili y en ichmen o he op 12 genomic unc ional elemen s. We pa i ioned he SNP-he i abili y o se um 25-hyd oxy i amin D concen a ions in o 24 publicly a ailable genomic unc ional elemen s using LD-sco e eg ession. We plo ed he en ichmen (Yaxis) o each o he 12 op anno a ions (as shown in Xaxis) in o a ba cha . G ay ba s and blue ba s ep esen he anno a ions wi h and wi hou he 500 base-pai windows. The heigh o each ba ep esen s magni ude o en ichmen . Significan es ima es o en ichmen ha passed Bon e oni co ec ions (P- alue o en ichmen <0.05/24) a e ma ked wi h double s a s. TSS ansc ip ion s a si es, DHS DNase I hype sensi i e si es, TFBS ansc ip ion ac o binding si es, Rep essed ep essed egions ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2 4NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions HDL (Supplemen a y Table 6). Howe e , wi h only six 25-hyd oxy i amin D associa ed SNPs included in he analysis, we conside an o e all null di ec ional co ela ion as ou main finding, and u he well-designed la ge-scale Mendelian ando- miza ion analyses a e wa an ed. Finally, we analyzed he 220 cell- ype-specific anno a ions in each o he 37 ai s and compa ed he cell- ype-specific en ichmen s o 25-hyd oxy i amin D o he en ichmen s o hese ai s. The en ichmen pa e n o 25-hyd oxy i amin D di e ed no ably om he pa e ns o psychia ic diseases and me abolic ela ed ai s. Psychia ic diseases showed en ichmen o his one ma ks specific o CNS cell ypes, and me abolic diseases showed en ichmen o gas oin es inal cell ypes, while hese anno a ions we e dep essed in 25-hyd oxy i amin D. Con e sely, 25-hyd oxy i amin D showed simila pa e ns wi h au oimmune inflamma o y diseases, whe e mul iple immune cell ypes we e en iched. We consis en ly obse ed ha 25- hyd oxy i amin D was clus e ed wi h au oimmune diseases (Supplemen a y Fig. 5). Discussion Vi amin D inadequacy has been linked o many diseases such as cance , au oimmune diso de and ca dio ascula condi ions in addi ion o musculoskele al diseases, which has led o subs an ial in e es in he de e minan s o i amin D s a us, especially i s gene ic componen s. We ha e pe o med a la ge 25- hyd oxy i amin D me a-GWAS in ol ing 31 s udies wi h a o al o 79,366 indi iduals. Ou esul s ecapi ula ed se e al p e- iously epo ed findings. Fi s o all, we confi med he ole o common gene ic a ian s in egula ion o ci cula ing 25- hyd oxy i amin D concen a ions. Ou s udy alida ed h ee loci, GC,NADSYN1/DHCR7, and CYP2R1, all we e es ablished 25-hyd oxy i amin D isk loci iden ified h ough wo ea lie GWASs8,11. In addi ion, we we e able o confi m he associa ion o a locus con aining CYP24A1 wi h 25-hyd oxy i amin D con- cen a ions using ou la ge sample size, which highligh s he impo ance o his p o ein in he deg ada ion o i amin D molecule, by ca alyzing hyd oxyla ion eac ions a he side chain o 1,25-dihyd oxy i amin D, he physiologically ac i e o m (ho monal o m) o i amin D. Significan finding a his locus was only shown in he pooled analyses in ol ing bo h disco e y and eplica ion samples in an ea lie GWAS8. We ex ended p e iously epo ed findings by iden i ying wo addi ional new loci. SEC23A (Sec23 Homolog A, coa p o ein complex II (COPII) componen ) encodes a membe o SEC23 sub amily. In euka yo ic cells, sec e ed p o eins a e syn- hesized in he endoplasmic e iculum (ER), packaged in o COPII-coa ed esicles, and a fic o he Golgi appa a us. As pa o COPII complex, SEC23 plays a ole in p omo ing ER-Golgi p o ein a ficking. SEC23A mu a ions ha e been epo ed o cause c aniolen iculosu u al dysplasia, a disease cha ac e ized by c anio acial and skele al mal o ma ion such as delay in closu e o on anels, su u al ca a ac s and acial dysmo phisms, due o de ec i e collagen sec ec ion13,14. The second no el locus is AMDHD1 (amidohyd olase domain con aining 1). This gene encodes an enzyme in ol ed in he his idine, lysine, phenylala- nine, y osine, p oline and yp ophan ca abolic pa hway. Mu a- ions in AMDHD1 a e ound o be associa ed wi h a ypical lipoma ous umo , a cance o connec i e issues ha esemble a cells15. Ou SNP-he i abili y esul s sugges ha 25-hyd oxy i amin D has a modes o e all he i abili y due o common genome-wide SNPs o 7.5%, and ha an app eciable p opo ion (2.84% ou o 7.5%, i.e., 38%) o his o al could be explained by known gene ic egions iden ified h ough GWAS. Ou findings a e in line wi h a p e ious published epo (by Hi aki e al.12) which es ima ed he a iance in ci cula ing 25-hyd oxy i amin D explained by SNPs in a o al o 5575 indi iduals12. Acco ding o ha epo , by employing a linea mixed model fi ing he addi i e gene ic ma ix c ea ed om all geno yped and impu ed SNPs, he p o- po ion o a iance explained was 8.9%; by employing a polygenic sco e app oach comp ised o he hen GWAS-disco e ed SNPs (GC, CYP2R1, DHCR7/NADSYN1), he p opo ion o a iance explained was 5%. Bo h o hese es ima es we e close o ou s. In Hi aki e al., he known 25-hyd oxy i amin D associa ed en i - onmen al ac o s such as age, BMI, season o blood d awn, i amin D die a y in ake, i amin D supplemen in ake, egion o esidence and e hnici y, explained ~18% o he obse ed a - iance12. Ou esul s, in ag eemen wi h hese findings, sugges ha al hough he e appea s o be some polygenic signals ou side o he iden ified egions, he emaining common e ec s may be small. The e also may be low equency a ian s wi h la ge e ec s ha we e no in es iga ed he e. Fo example, while his pape was unde e iew, a ela ed s udy iden ified low- equency (MAF =2.5%) synonymous coding a ian s117913124_A a CYP2R1 con e ing a la ge e ec on 25-hyd oxy im ain D le els, which was ou imes g ea e in magni ude and independen o a p e iously desc ibed associa ion o a common a ian ( s10741657) nea CYP2R116. Resul s o win and amilial s udies ha e e ealed a subs an ial gene ic basis in he a iabili y o ci cula ing 25-hyd oxy i amin D le els, wi h es ima es o he i abili y eaching as high as 86%9,10,17–19. These es ima es, howe e , seem o be influenced by en i onmen al condi ions. Fo example, in a s udy conduc ed by O on e al. wi h 40 monozygo ic and 59 dizygo ic win pai s, bloods we e collec ed a he end o win e and a he i abili y o 77% was epo ed10. Simila ly, he s udy conduc ed by Ka ohl e al. wi h 310 monozygo ic and 200 dizygo ic male wins obse ed a he i abili y o 70% du ing win e , whe eas in summe , se um 25-hyd oxy i amin D concen a ions appea ed o be Table 3 He i abili y en ichmen o en g ouped cell ypes Ca ego y P opo ion o SNPs (%) P opo ion o h2 g(%) En ichmen (s anda d e o s) P- alue Kidney 4.26 27.27 6.4 (2.44) 0.027 Li e 7.22 37.68 5.22 (1.55) 0.01 Gas oin es inal 16.77 72.88 4.35 (0.97) 0.0017 Immune and hema opoie ic 23.34 100.17 4.29 (0.76) 2.20E-05 Cen al ne ous sys em 14.88 54.09 3.64 (0.87) 0.0039 Ca dio ascula 11.11 35.74 3.22 (1.26) 0.078 Connec i e issue/bone 11.5 35.65 3.1 (1.04) 0.037 Ad enal/panc eas 9.36 26.17 2.8 (1.31) 0.18 O he 20.27 56.68 2.8 (0.98) 0.076 Skele al Muscle 10.38 14.29 1.38 (1.25) 0.76 Black bold on indica es significan P- alues a e mul iple co ec ions (P<0.05/10) NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2 ARTICLE NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions 5 en i ely de e mined by non-gene ic ac o s (he i abili y: 0%)9. Compa able es ima es we e also iden ified in a sligh ly la ge s udy conduc ed by Mills e al. (win e : 90% s. summe : 56%)18. Consis en wi h season dependency, sex di e ences we e also obse ed (males: 86% s. emales: 17%)17. While hese es ima es should be ea ed wi h cau ion due o small samples and ela ed imp ecision, hey confi m he subs an ial a ia ion in 25- hyd oxy i amin D le els by season (as shown p e iously20) and illus a e ha he i abili y es ima es de i ed om a homogenous sou ce may be highly infla ed. In a ele an ly well-powe ed win s udy wi h a o al o ~2100 emale wins, he he i abili y o 25- hyd oxy i amin D was calcula ed o be 40%, indica ing a la ge p opo ion o a iance explained by non-gene ic ac o s21. He - i abili y es ima es ob ained using GWAS SNPs ha e ypically been ound o be app oxima ely hal o hose om classical win s udies9,10, bu ou es ima e o 7.5%, calcula ed using common genome-wide SNPs, is a lowe han epo ed he i abili y om win and amily based s udies. In addi ion o po en ially infla ed es ima es om win s udies, he di e ence may eflec he p o- po ion o he i abili y explained by a e SNPs o s uc u al a - ian s ha we e no included in ou da a, and he po en ial gene- gene in e ac ions ha emain o be iden ified. The combina ion o ou samples om all seasons is also likely o dec ease he p ob- abili y o finding gene ic a ian s, and hence defla e he i abili y es ima es. Th ough pa i ioning he SNP-he i abili y o se um 25- hyd oxy i amin D le els, we obse ed a significan en ichmen in immune and hema opoie ic issues; likewise, he cell- ype- specific analysis e ealed clus e ing o 25-hyd oxy i amin D and au oimmune diseases, indica ing ha hese ai s sha e a majo i y T ai s Gene ic co ela ion (95%CI) Au oimmune/in lamma o y diseases Age- ela ed macula degene a ion –0.02 (–0.1, 0.06) p- alue Signi icance * * * * 0.57 0.89 0.25 0.98 0.21 0.019 0.44 0.59 0.9 0.91 0.71 0.39 0.7 0.64 0.63 0.49 0.61 1 0.25 0.66 0.98 0.21 0.36 0.18 0.26 0.54 0.019 0.13 0.25 0.45 0.63 0.69 0.0036 0.64 0.078 0.8 0.042 –0.02 (–0.17, 0.15) –0.07 (–0.19, 0.05) 0 (–0.15, 0.14) –0.17 (–0.45, 0.1) –0.18 (–0.33, –0.03) 0.05 (–0.08, 0.19) 0.04 (–0.1, 0.18) –0.01 (–0.11, 0.1) –0.01 (–0.12, 0.1) –0.02 (–0.13, 0.09) –0.08 (–0.28, 0.11) –0.03 (–0.15, 0.1) 0.03 (–0.11, 0.17) 0.04 (–0.11, 0.18) 0.04 (–0.08, 0.17) 0.02 (–0.07, 0.11) 0 (–0.08, 0.08) 0.14 (–0.1, 0.38) –0.03 (–0.16, 0.1) 0 (–0.12, 0.12) 0.07 (–0.04, 0.19) 0.06 (–0.07, 0.18) –0.06 (–0.16, 0.03) –0.08 (–0.22, 0.06) 0.02 (–0.05, 0.1) –0.17 (–0.32, –0.03) –0.07 (–0.02, 0.15) –0.07 (–0.19, 0.05) –0.03 (–0.11, 0.05) 0.02 (–0.06, 0.09) –0.01 (–0.08, 0.06) 0.14 (0.04, 0.23) –0.02 (–0.09, 0.05) 0.06 (–0.01, 0.12) –0.01 (–0.08, 0.06) –0.1 (–0.2, 0) –0.5 –0.25 00.25 0.5 Celiac Disease C ohn’s Disease Lupus Mul iple Scle osis P ima y bilia y ci hosis Rheuma oid A h i is Ulce a i e Coli is UKBiobank As hma UKBiobank Eczema In lamma o y Bowel Disease Alzheime ’s Psychia ic diso de / ai s Me abolism ela ed ai s Co ona y A e y Disease Fas ing Glucose HDL LDL T iglyce ides Type 2 Diabe es UKBiobank Hype ension E e Smoked UKBiobank Age a Mena che UKBiobank Age a Menopause UKBiobank BMI UKBiobank Dias olic UKBiobank FEV1FVC UKBiobank FVC UKBiobank Heel TSco e UKBiobank Heigh UKBiobank Sys olic UKBiobank Wais -Hip Ra io O he s Ano exia Au ism Bipola Diso de Dep essi e Synd ome Neu o icism Schizoph enia Subjec Well Being Fig. 3 Gene ic co ela ions be ween 25-hyd oxy i amin D and 37 ai s. We collec ed GWAS summa y s a is ics o 37 diseases and ai s spanning a wide ange o pheno ypes (au oimmune inflamma o y diseases, psychia ic diso de s, me abolic ai s, and an h opome ic index) om publicly a ailable esou ces, and es ima ed hei sha ed gene ic simila i ies wi h se um 25-hyd oxy i amin D le els. We plo ed he gene ic co ela ion oge he wi h 95% confidence in e als using a blue squa e and g ay ho izon al lines. Red e ical line indica es no gene ic co ela ion ( g=0). S a is ical significance was defined as P- alue <0.05. None o he pai wise co ela ions passed Bon e oni co ec ions ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2 6NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions o common cell ypes. The link be ween i amin D deficiency and inc eased isk o au oimmune inflamma o y diseases has long been ecognized by epidemiological in es iga ions22,23. Al hough he unde lying mechanisms emain unclea , i is now e iden ha i amin D is in ol ed in many biological p ocesses ha egula e bo h inna e and adap i e immune esponses, h ough ligand- ecep o binding, ac i a ion, in e ac ion wi h esponse elemen s in he p omo e egions o di e en genes, and e en ually lead o unc ional changes in a wide a ie y o immune cells including Th1, Th2, Th17, T egula o y and na u al kille T cells22,23. The sha ed cell ype en ichmen s be ween i amin D and au oimmune diseases obse ed in ou s udy, u he sugges ha i amin D no only a ec s au oimmune diseases h ough i s di ec e ec (as a ligand), bu also h ough hei sha ed gene ic e iology. Thus, indi iduals wi h i amin D deficiency may be mo e suscep ible o hese diso de s, bo h because o en i onmen al and gene ic influences. Ou genome-wide in e ac ion analysis be ween gene ic a ian s and die a y in ake o i amin D did no iden i y new signals. All significan associa ions obse ed in he join es o main gene ic and in e ac ion e ec s we e o equal o highe significance le el (i.e., lowe P- alues) in he GWAS o ma ginal gene ic e ec pe o med in he same indi iduals, indica ing no majo con- ibu ion o in e ac ion e ec s a hese loci. Indeed, only one o he op 5 loci om he o e all ma ginal GWAS showed nominally significan in e ac ion e ec , and none passed Bon e oni co - ec ions. While smalle gene-die in e ac ion e ec s emain o be disco e ed, ou esul s p o ide some e idence agains la ge in e ac ions be ween common SNPs and die a y i amin D in ake. S ill, one canno comple ely ule ou he possibili y o in e ac ion, bu only conclude ha gene ic e ec s appea s able wi hin i amin D in ake ange in he popula ions s udied. Indeed, as o any gene-en i onmen in e ac ion es s, s a is ical powe is highly dependen on he a iance o exposu e in he samples analyzed24, and in e ac ions would emain unobse ed i he exposu e is homogeneous among indi iduals. Also, we we e no able o cap u e i amin D supplemen a ion adequa ely o include his in he die a y in ake a iable, and we e no able o es ima e sunligh exposu e as a sou ce o i amin D p oduc ion in he skin. Se um 25-hyd oxy i amin D concen a ions a e mainly de e mined by modifiable en i onmen al ac o s, and con a y o es ima es om p e ious win s udies, ou la ge-scale analyses sugges a SNP-he i abili y a e ha is ela i ely modes in mag- ni ude when conside ing common a ian s. Ou s udy also showed ha common gene ic a ian s a e unlikely o ha e a s ong modi ying e ec on inc eases in 25-hyd oxy i amin D ollowing ypical die a y in akes, sugges ing ha conside a ion o gene ic backg ound is no equi ed when de e mining popula ion based i amin D in ake ecommenda ions. Howe e , ou esul s suppo he ole o i amin D in immunological diseases as we obse ed om cell- ype-specific analysis o clus e ing o i amin D and au oimmune diseases, and he e idence o signal en ichmen o immune and hema opoie ic issues. These find- ings a e in line wi h p e ious Mendelian Randomiza ion s udies which ound a pu a i e causal associa ion be ween i amin D and au oimmune diseases such as mul iple scle osis1,2and ype 1 diabe es25. The addi ional gene ic ins umen s iden ified by ou esul s could also be used in u u e Mendelian Randomiza ion analyses o he associa ion be ween i amin D and complex ai s. Me hods S udy coho s. We expanded ou p e ious SUNLIGHT conso ium GWAS, and unde ook a la ge, mul icen e , genome-wide associa ion s udy o 31 coho s in Eu ope, Canada and USA. Ou fi s s age disco e y me a-analysis consis ed o 79,366 samples o Eu opean descen d awn om 31 epidemiological coho s. Among hose 31 coho s, en we e used as disco e y and in-silico eplica ion samples in ou p e ious GWAS publica ion ( he 1958 B i ish Bi h Coho (1958BC), he Ca dio ascula Heal h S udy (CHS), he F amingham Hea S udy (FHS), he Go henbu g Os eopo osis and Obesi y De e minan s s udy (GOOD), he Heal h, Aging, and Body Composi ion s udy (Heal h ABC), he Indiana Women coho , he No h Finland Bi h Coho 1966 (NFBC1966), he Old O de Amish S udy (OOA), he Ro e dam S udy (RS), and he TwinsUK), and an addi ional 21 coho s we e included o he cu en analysis ( he Alpha-Toco- phe ol, Be a-Ca o ene Cance P e en ion S udy (ATBC), he A he oscle osis Risk in Communi ies S udy (ARIC), he A he oGene egis y, B- i amins o he P e- en ion O Os eopo o ic F ac u es (B-PROOF), he Epidemiology o Diabe es In e en ions and Complica ions (EDIC), he Case-Con ol S udy o Me abolic Synd ome (GenMe s), he Helsinki Bi h Coho S udy (HBCS), he Heal h P o- essional Follow-up S udy (HPFS, nes ed a co ona y hea disease case-con ol s udy), he In ecchia e in Chian i S udy (InChian i), he Coope a i e Heal h Resea ch in he egion Augsbu g (KORA), he Leiden Longe i y S udy (LLS), he Ludwigsha en Risk and Ca dio ascula Heal h S udy (LURIC), he Mul i-E hnic S udy o A he oscle osis (MESA), he Nijmegen Biomedische S udie (NBS), he Nu ses’Heal h S udy (NHS, nes ed a b eas cance case-con ol s udy, and a ype2 diabe es case-con ol s udy), he O kney Complex Disease S udy (ORCADES), he P os a e, Lung, Colo ec al, and O a ian Cance Sc eening T ial (PLCO), he PROspec i e S udy o P a as a in in he Elde ly a Risk (PROSPER), he S udy o Heal h in Pome ania (SHIP), he Sco ish Colo ec al Cance S udy (SOCCS), he Ca dio ascula Risk in Young Finns S udy (YFS), and mo e samples om he RS (RSI, RSII, and RSIII)). Full desc ip ions o all pa icipa ing coho s, de ails o geno yping pla o ms used, numbe o SNPs, and he measu emen s o se um 25- hyd oxy i amin D concen a ions in each coho a e shown in Supplemen a y Table 1and Supplemen a y No e 1. W i en in o med consen was ob ained om all pa icipan s in he included coho s, and he s udy p o ocols we e e iewed and app o ed by local ins i u ional e iew boa ds. Powe calcula ion. Ou la ge sample size p o ided good s a is ical powe o associa ion analysis. A he genome-wide significance h eshold o 5×10−8, wi h a disco e y sample size o 75,000, ou s udy had 85% powe o de ec a gene ic a ian (single nucleo ide polymo phism, SNP) accoun ing o 0.06% o he o al a iance in se um 25-hyd oxy i amin D concen a ions, and 99% powe o de ec a a ian ha explained 0.1% o he o al a iance. We also had powe o de ec gene- en i onmen in e ac ion e ec s e en smalle han he obse ed ma ginal e ec s. In he case whe e a SNP has no ma ginal e ec on ci cula ing 25-Hyd oxy i amin D concen a ions (and so could no ha e been disco e ed ia he ma ginal GWAS), we had 80% powe o de ec an in e ac ion ha explained 0.07% o he o al a iance in 25-hyd oxy i amin D concen a ions. Associa ion analysis. Genome-wide analyses we e pe o med wi hin each coho acco ding o a uni o m analysis plan. We fi addi i e gene ic models using linea eg ession on na u al-log ans o med 25-hyd oxy i amin D, and adjus ed he models o mon h o sample collec ion (12 ca ego ies), age, sex, and body mass index, and p incipal componen s cap u ing gene ic ances y. Fu he adjus men s included coho -specific a iables, such as geog aphical loca ion and assay ba ch, whe e ele an . Fo pa icipa ing s udies wi h a case-con ol design, we analyzed cases and con ols sepa a ely. We pe o med a fixed-e ec s in e se a iance weigh ed me a-analysis ac oss he con ibu ing coho s, as implemen ed in he so wa e METAL26, wi h con ol o popula ion s uc u e wi hin each coho and quali y con ol h esholds o mino allele equency (MAF) >0.05, impu a ion in o sco e >0.8, Ha dy-Weinbe g equilib ium (HWE) >1×10−6, and a minimum o wo s udies and 10,000 indi iduals con ibu ing o each epo ed SNP-pheno ype associa ion. We ega ded P- alues <5×10−8as genome-wide significan . Replica ion s udy. We eplica ed he iden ified no el loci in wo independen da a se s o which geno ype da a we e a ailable: he Eu opean P ospec i e In es iga ion in o Cance and Nu i ion (EPIC) s udy wi h 40,562 indi iduals ac oss wo nes ed case-con ol s udies (EPIC-In e Ac and EPIC-CVD) and he coho -wide EPIC- No olk s udy (Supplemen a y No e 1); and a coho o 2195 indi iduals (all con ols) addi ionally collec ed as pa o he SOCCS ha we e no included in ou disco e y s age. As o he pheno ype, EPIC indi iduals we e assayed o plasma 25-hyd oxy i amin D 3 and SOCCS indi iduals we e assayed o o al 25- hyd oxy i amin D. We pe o med he associa ion analysis in a simila manne , adjus ed o age, sex, ime o sample collec ion, and s udy cen e whe e ele an . We ega ded P- alue <0.05 in he eplica ion samples, and P- alue <5×10−8in he pooled analysis as success ul eplica ion. Condi ional analysis. A e iden i ying he p ima y associa ed a ian a each locus selec ed acco ding o he s eng h o i s associa ion, we u he es ed whe he he e we e any o he SNPs significan ly associa ed wi h 25-hyd oxy i amin D a e accoun ing o he e ec o lead SNP. We hus pe o med a s epwise model selec ion p ocedu e o hose ch omosomes whe e a significan a ian was p e- iously iden ified. We s a ed wi h he mos significan ly associa ed SNP, scanning h ough he whole ch omosome, selec ing addi ional independen ly associa ed SNPs using a s epwise p ocedu e, one a a ime, based on hei condi ional P- alues. Finally, we fi all selec ed SNP in o one model o es ima e hei join e ec s. NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2 ARTICLE NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions 7 We used GCTA-COJO so wa e o accommoda e ou summa y le el GWAS da a27, and he Cance Gene ic Ma ke o Suscep ibili y (CGEMS) GWAS wi h 2287 indi iduals o Eu opean descen and 2,543,887 geno yped and impu ed (HapMap22) SNPs as e e ence panel. SNP-by-die in e ac ion. We pe o med a genome-wide associa ion sc eening o ci cula ing 25-hyd oxy i amin D while accoun ing o po en ial in e ac ion e ec be ween SNP and die a y i amin D in ake. Ou es s inco po a ing gene-die in e ac ion we e based on he ollowing model: ln 25 OH ðÞ D ðÞ ¼β0þβ1´Gþβ2´Eþβ3´G´EþβZ´Z whe e Gis a SNP ha was coded addi i ely, Eis he aw i amin D in ake, measu ed on a con inuous scale. The pa ame e s β0,β1,β2, and β3a e he in e cep , he main e ec o SNP, he main e ec o die a y i amin D in ake, and he in e ac ion e ec be ween Gand E. The model also included he same co a ia es Zas o he ma ginal e ec sc eening, he e ec s o which we e cap u ed in he pa ame e βZ. We conside ed bo h a s anda d 1 deg ee-o - eedom es o in e ac ion e ec (i.e., null hypo hesis o β3¼0), and a join 2 deg ee-o - eedom es o main gene ic e ec plus gene-by-die in e ac ion (i.e., null hypo hesis o β1¼0 and β3¼0). Fo compa ison pu poses, we also conside ed a model adjus ing o i amin D in ake bu no modeling in e ac ion (i.e., no including he β3´G´E e m) using he same subse o indi iduals. Vi amin D in ake was a ailable o 15 coho s on a o al o 41,981 indi iduals. I included bo h he popula ion based coho s (ARIC, 1958BC, B-PROOF, FHS, Heal h ABC, MESA, NFBC, RS, RS III, and YFS as pa o he o e all Me a-GWAS, plus wo addi ional coho s, he P ospec i e In es iga ion o he Vascula u e in Uppsala Senio s (PIVUS), and he Uppsala Longi udinal S udy o Adul Men (ULSAM), geno yped on cus om a ay ha we e no included in he o e all me a- GWAS bu we e included in his SNP-by-die in e ac ion analysis), and case-con ol s udies (HPFS (HPFS_CHD), NHS (NHS_BRCA, NHS_T2D), and SOCCS). Fo he la e s udies, all analyses we e pe o med sepa a ely in cases and con ols. The a o emen ioned in e ac ion model was applied o each o he included coho s, and s udy-specific esul s we e me a-analyzed using in e se- a iance weigh ed sum o e ec es ima es as implemen ed in METAL26. Fo he 2 deg ee-o - eedom es we used join amewo k desc ibed in wo p e ious published pape s28,29. Quali y con ol fil e ing was pe o med on each s udy be o e me a-analysis. Only SNPs wi h impu a ion in o sco e >0.8, MAF >0.05, HWE >1×10−6, and a minimum o o al sample size in he me a-analysis >10,000 we e e ained. Linkage Disequilib ium sco e eg ession. We pe o med linkage disequilib ium sco e eg ession (LDSC) analysis o es ima e he SNP-he i abili y o se um 25- hyd oxy i amin D concen a ions30,31. This me hod is based on a alida ed ela- ionship be ween LD sco e and χ2-s a is ics: Eχ2 j hi Njh2 g Mljþ1 whe e Eχ2 j hi deno es he expec ed χ2-s a is ics o he associa ion be ween ou come and SNP j,N j is he s udy sample size a ailable o SNP j, M is he o al numbe s o a ian s and ljdeno es he LD sco e o SNP j defined as lj¼P k 2ðj;kÞ. LDSC calcula es he i abili y using only summa y-le el da a ins ead o indi idual geno ypes, and is compu a ionally cos e ec i e a la ge sample sizes. We used he summa y s a is ics om 25-hyd oxy i amin D me a-GWAS esul s, wi h SNPs a ailable in a leas 2 s udies and a sample size o a leas 10,000. We fi s analyzed he SNP-he i abili y by using (1) SNPs ac oss he en i e genome ha passed quali y con ol; (2) SNPs excluding he op associa ions (SNPs eaching genome-wide significance, P≤5×10−8), as well as all SNPs wi hin ±500 kb o he op hi s in he egion; and (3) SNPs excluding he nominally significan associa ions (P≤0.05). We subsequen ly pa i ioned he he i abili y h ough h ee di e en models: (1) a ull baseline model including he 24 publicly a ailable main anno a ions ha a e no specific o any cell ype, he 500-bp windows a ound each anno a ion, as well as 100-bp windows a ound ch oma in immunop ecipi a ion and sequencing peaks (ChIP-seq) when app op ia e. This esul ed in a o al o 52 o e lapping unc ional ca ego ies in he ull baseline model; (2) a cell- ype-specific model including 10 cell ype g oups: ad enal and panc eas, cen al ne ous sys em (CNS), ca dio ascula , connec i e and bone, gas oin es inal, immune and hema opoie ic, kidney, li e , skele al muscle, and o he ; (3) a cell- ype-specific model including 220 cell- ype- specific anno a ions o he ou his one ma ks wi h pu a i e enhance o p omo e unc ions, H3K4me1, H3K4me3, H3K9ac, and H3K27ac. De ails o he 24 publicly a ailable anno a ions, he 220 cell- ype-specific anno a ions, as well as he 10 cell ype g oups we e desc ibed by Finucane e al.31. B iefly, he 24 anno a ions included coding, UTR (3′UTR and 5′UTR), p omo e and in onic egions, acqui ed om he UCSC Genome B owse 32 and pos - p ocessed by Guse e al.33; he h ee his one ma ks (mono-me hyla ion (H3K4me1) o his one H3 a lysine 4, i-me hyla ion (H3K4me3) o his one H3 a lysine 4, and ace yla ion o his one H3 a lysine 9 (H3K9ac) p ocessed by T ynka e al.34–36 and wo e sions o ace yla ion o his one H3 a lysine 27 (H3K27ac, one e sion p ocessed by Hnisz e al.37, ano he e sion used by he Psychia ic Genomics Conso ium (PGC)38); open ch oma in, as eflec ed by DNase I hype sensi i i y si es (DHSs and e al DHSs)33, ob ained as a combina ion o Encyclopedia o DNA Elemen s (ENCODE) and Roadmap Epigenomics da a, and p ocessed by T ynka e al.36; combined ch omHMM and Segway p edic ions ob ained om Ho man e al.39, which le e age on many anno a ions o pa i ion he genome in o se en unde lying ch oma in s a es ( he CCCTC-binding ac o (CTCF), p omo e -flanking, ansc ibed egion, ansc ip ion s a si e (TSS), s ong enhance , weak enhance , and he ep essed egion); egions ha a e conse ed in mammals, p o ided by Lindblad-Toh e al.40 and pos -p ocessed by Wa d and Kellis41; supe -enhance s, which a e la ge g oups o pu a i e enhance s wi h high le els o ac i i y, p o ided by Hnisz e al.37; FANTOM5 enhance s mapped by using cap analysis o gene exp ession in he FANTOM5 panel o samples, ob ained om Ande sson e al.42; digi al genomic oo p in (DGF) and ansc ip ion ac o binding si e (TFBS) anno a ions downloaded om ENCODE35 and pos -p ocessed by Guse e al.33. We included 500-bp windows a ound each o he 24 main anno a ions in he baseline model, and 100-bp windows a ound ChIP- seq when app op ia e, o p e en upwa d bias o es ima es gene a ed by en ichmen in he nea by egions. In addi ion o he baseline model using 24 main anno a ions, we also pe o med cell- ype-specific analyses using anno a ions o he ou his one ma ks (H3K4me1, H3K4me3, H3K9ac and H3K27ac). Each cell- ype-specific anno a ion co esponds o a his one ma k in a single cell ype ( o example, H3K27ac in adipose nuclei issues), and he e was a o al o 220 such anno a ions. We u he subdi ided hese 220 cell- ype-specific anno a ions in o 10 ca ego ies by agg ega ing he cell- ype- specific anno a ions wi hin each g oup ( o example, SNPs ela ed wi h any o he ou his one modifica ions in any hema opoie ic and immune cells we e conside ed as one big ca ego y). When gene a ing he cell- ype-specific models, we added each anno a ion indi idually (one a a ime) o he baseline model, c ea ing sepa a e models o con ol o o e lap wi h he genomic unc ional elemen s in he ull baseline model bu no o e lap wi h he o he cell ypes. We addi ionally assembled he summa y s a is ics om GWAS o 37 ai s o diseases pe o med in indi iduals o Eu opean descen , which a e publicly a ailable38,43–55 o applied om he UK Biobank. These s udies span a wide ange o pheno ypes, om an h opome ic indices such as heigh , weigh , BMI, o men al diso de s ( o example dep essi e synd ome and schizoph enia) o au oimmune and inflamma o y diseases ( o example heuma oid a h i is and celiac diseases). We calcula ed he pai wise gene ic co ela ion ( g, c oss ai he i abili y) be ween 25-hyd oxy i amin D and each o he 37 ai s. We u he conduc ed he same cell- ype-specific analysis o each ai , and plo ed be a-coe ficien z-sco e ma ix, cons uc ed om he o al 220 anno a ions by 37 ai s, in o ou hea -maps based on he ou his one ma ks. Finally, in addi ion o he gene ic co ela ion analysis which eflec s sha ed gene ic ac o s ac oss di e en ai s bu does no in o m di ec ion, we also a emp ed o iden i y di ec ions o such co ela ion using an algo i hm p oposed by Pick ell e al.56. The me hod adop s a simila in ui ion as he Mendelian Randomiza ion app oach, whe e, i a ai X influences ai Y, hen SNPs influencing X should also influence Y, and he SNP-specific e ec sizes o he wo ai s should be co ela ed. Fu he , since Y does no influence X, bu could be influenced by mechanisms independen o X, gene ic a ian s ha influence Y do no necessa ily influence X. Based on his in ui ion, he me hod p oposes wo “causal”models and wo “non-causal”models, and calcula es he ela i e likelihood a io o he bes non-causal model compa ed o he bes causal model. We de e mined significan SNPs o each gi en ai by selec ed genome-wide significan (P<5×10−8) SNPs and p uned he numbe s based on hei LD-pa e n in he Eu opean popula ions in Phase1 o 1000 Genome P ojec . We scanned h ough all pai s o 25-hyd oxy i amin D and ai s o iden i y di ec ional co ela ions. We conside pai s o ai s wi h likelihood a io non-causal s. causal < 0.05 as ha ing e idence o di ec ional co ela ions. Da a a ailabili y. The GWAS summa y s a is ics on se um ci cula ing i amin D concen a ions is a ailable a dbGap h ps://d i e.google.com/d i e/ olde s/ 0BzYD Co_doHJRFRKR0l ZHZWZjQ; all ele an da a a e a ailable om he au ho s upon eques . Recei ed: 21 July 2017 Accep ed: 15 Decembe 2017 Re e ences 1. Mok y, L. E. e al. Vi amin D and isk o mul iple scle osis: a mendelian andomiza ion s udy. PLoS Med. 12, e1001866 (2015). 2. Rhead, B. e al. Mendelian andomiza ion shows a causal e ec o low i amin D on mul iple scle osis isk. Neu ol. Gene . 2, e97 (2016). 3. Ma ineau, A. R. e al. Vi amin D supplemen a ion o p e en acu e espi a o y ac in ec ions: sys ema ic e iew and me a-analysis o indi idual pa icipan da a. B . Med. J. 356, i6583 (2017). ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2 8NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions 4. Pilz, S., Ve heyen, N., G üble , M. R., Tomaschi z, A. & Mä z, W. Vi amin D and ca dio ascula disease p e en ion. Na . Re . Ca diol. 13, 404–417 (2016). 5. Ga land, C. F. e al. The ole o i amin D in cance p e en ion. Am. J. Public Heal h 96, 252–261 (2006). 6. Fe nandes de Ab eu, D. A., Eyles, D. & Fé on, F. Vi amin D, a neu o- immunomodula o : implica ions o neu odegene a i e and au oimmune diseases. Psychoneu oendoc inology 34 Suppl 1, S265–277 (2009). 7. Laguno a, Z., Po ojnicu, A. C., Lindbe g, F., Hexebe g, S. & Moan, J. The dependency o i amin D s a us on body mass index, gende , age and season. An icance Res. 29, 3713–3720 (2009). 8. Wang, T. J. e al. Common gene ic de e minan s o i amin D insu ficiency: a genome-wide associa ion s udy. Lance Lond. Engl. 376, 180–188 (2010). 9. Ka ohl, C. e al. He i abili y and seasonal a iabili y o i amin D concen a ions in male wins. Am. J. Clin. Nu . 92, 1393–1398 (2010). 10. O on, S.-M. e al. E idence o gene ic egula ion o i amin D s a us in wins wi h mul iple scle osis. Am. J. Clin. Nu . 88, 441–447 (2008). 11. Ahn, J. e al. Genome-wide associa ion s udy o ci cula ing i amin D le els. Hum. Mol. Gene . 19, 2739–2745 (2010). 12. Hi aki, L. T. e al. Explo ing he gene ic a chi ec u e o ci cula ing 25- hyd oxy i amin D. Gene . Epidemiol. 37,92–98 (2013). 13. Boyadjie , S. e al. C anio-len iculo-su u al dysplasia associa ed wi h de ec s in collagen sec e ion. Clin. Gene . 80, 169–176 (2011). 14. Boyadjie , S. A. e al. C anio-len iculo-su u al dysplasia is caused by a SEC23A mu a ion leading o abno mal endoplasmic- e iculum- o-Golgi a ficking. Na . Gene . 38, 1192–1197 (2006). 15. Myung, J. K. e al. Well-di e en ia ed liposa coma o he oesophagus: clinicopa hological, immunohis ochemical and a ay CGH analysis. Pa hol. Oncol. Res. 17, 415–420 (2011). 16. Manousaki, D. e al. Low- equency synonymous coding a ia ion in CYP2R1 has la ge e ec s on i amin D le els and isk o mul iple scle osis. Am. J. Hum. Gene . 101, 227–238 (2017). 17. A guelles, L. M. e al. He i abili y and en i onmen al ac o s a ec ing i amin D s a us in u al Chinese adolescen wins. J. Clin. Endoc inol. Me ab. 94, 3273–3281 (2009). 18. Mills, N. T. e al. He i abili y o ans o ming g ow h ac o -β1 and umo nec osis ac o - ecep o ype 1 exp ession and i amin D le els in heal hy adolescen wins. Twin Res. Hum. Gene . O . J. In . Soc. Twin S ud. 18,28–35 (2015). 19. Li shi s, G., Ka asik, D. & Seibel, M. J. S a is ical gene ic analysis o plasma le els o i amin D: amilial s udy. Ann. Hum. Gene . 63, 429–439 (1999). 20. Yu, H.-J., Kwon, M.-J., Woo, H.-Y. & Pa k, H. Analysis o 25-hyd oxy i amin D s a us acco ding o age, gende , and seasonal a ia ion. J. Clin. Lab. Anal. 30, 905–911 (2016). 21. Hun e , D. e al. Gene ic con ibu ion o bone me abolism, calcium exc e ion, and i amin D and pa a hy oid ho mone egula ion. J. Bone Mine . Res. O . J. Am. Soc. Bone Mine . Res. 16, 371–378 (2001). 22. Agmon-Le in, N., Theodo , E., Segal, R. M. & Shoen eld, Y. Vi amin D in sys emic and o gan-specific au oimmune diseases. Clin. Re . Alle gy Immunol. 45, 256–266 (2013). 23. Yang, C.-Y., Leung, P. S. C., Adamopoulos, I. E. & Ge shwin, M. E. The implica ion o i amin D and au oimmuni y: a comp ehensi e e iew. Clin. Re . Alle gy Immunol. 45, 217–226 (2013). 24. Ascha d, H. A pe spec i e on in e ac ion e ec s in gene ic associa ion s udies. Gene . Epidemiol. 40, 678–688 (2016). 25. Coope , J. D. e al. Inhe i ed a ia ion in i amin D genes is associa ed wi h p edisposi ion o au oimmune disease ype 1 diabe es. Diabe es 60, 1624–1631 (2011). 26. Wille , C. J., Li, Y. & Abecasis, G. R. METAL: as and e ficien me a-analysis o genomewide associa ion scans. Bioin o ma. Ox . Engl. 26, 2190–2191 (2010). 27. Yang, J. e al. Condi ional and join mul iple-SNP analysis o GWAS summa y s a is ics iden ifies addi ional a ian s influencing complex ai s. Na . Gene . 44, 369–375 (2012). 28. Ascha d, H., Hancock, D. B., London, S. J. & K a , P. Genome-wide me a- analysis o join es s o gene ic and gene-en i onmen in e ac ion e ec s. Hum. He ed. 70, 292–300 (2010). 29. Manning, A. K. e al. Me a-analysis o gene-en i onmen in e ac ion: join es ima ion o SNP and SNP×en i onmen eg ession coe ficien s. Gene . Epidemiol. 35,11–18 (2011). 30. Bulik-Sulli an, B. K. e al. LD sco e eg ession dis inguishes con ounding om polygenici y in genome-wide associa ion s udies. Na . Gene . 47, 291–295 (2015). 31. Finucane, H. K. e al. Pa i ioning he i abili y by unc ional anno a ion using genome-wide associa ion summa y s a is ics. Na . Gene . 47, 1228–1235 (2015). 32. Ken , W. J. e al. The human genome b owse a UCSC. Genome Res. 12, 996–1006 (2002). 33. Guse , A. e al. Pa i ioning he i abili y o egula o y and cell- ype-specific a ian s ac oss 11 common diseases. Am. J. Hum. Gene . 95, 535–552 (2014). 34. Roadmap Epigenomics Conso ium. e al. In eg a i e analysis o 111 e e ence human epigenomes. Na u e 518, 317–330 (2015). 35. ENCODE P ojec Conso ium. An in eg a ed encyclopedia o DNA elemen s in he human genome. Na u e 489,57–74 (2012). 36. T ynka, G. e al. Ch oma in ma ks iden i y c i ical cell ypes o fine mapping complex ai a ian s. Na . Gene . 45, 124–130 (2013). 37. Hnisz, D. e al. Supe -enhance s in he con ol o cell iden i y and disease. Cell 155, 934–947 (2013). 38. Schizoph enia Wo king G oup o he Psychia ic Genomics Conso ium. Biological insigh s om 108 schizoph enia-associa ed gene ic loci. Na u e 511, 421–427 (2014). 39. Ho man, M. M. e al. In eg a i e anno a ion o ch oma in elemen s om ENCODE da a. Nucleic Acids Res. 41, 827–841 (2013). 40. Lindblad-Toh, K. e al. A high- esolu ion map o human e olu iona y cons ain using 29 mammals. Na u e 478, 476–482 (2011). 41. Wa d, L. D. & Kellis, M. E idence o abundan pu i ying selec ion in humans o ecen ly acqui ed egula o y unc ions. Science 337, 1675–1678 (2012). 42. Ande sson, R. e al. An a las o ac i e enhance s ac oss human cell ypes and issues. Na u e 507, 455–461 (2014). 43. Bo aska, V. e al. A genome-wide associa ion s udy o ano exia ne osa. Mol. Psychia y 19, 1085–1094 (2014). 44. Global Lipids Gene ics Conso ium. e al. Disco e y and efinemen o loci associa ed wi h lipid le els. Na . Gene . 45, 1274–1283 (2013). 45. Ben ham, J. e al. Gene ic associa ion analyses implica e abe an egula ion o inna e and adap i e immuni y genes in he pa hogenesis o sys emic lupus e y hema osus. Na . Gene . 47, 1457–1464 (2015). 46. Okada, Y. e al. Gene ics o heuma oid a h i is con ibu es o biology and d ug disco e y. Na u e 506, 376–381 (2014). 47. Okbay, A. e al. Gene ic a ian s associa ed wi h subjec i e well-being, dep essi e symp oms, and neu o icism iden ified h ough genome-wide analyses. Na . Gene . 48, 624–633 (2016). 48. Jos ins, L. e al. Hos -mic obe in e ac ions ha e shaped he gene ic a chi ec u e o inflamma o y bowel disease. Na u e 491, 119–124 (2012). 49. C oss-Diso de , G oup o he Psychia ic Genomics Conso ium Iden ifica ion o isk loci wi h sha ed e ec s on fi e majo psychia ic diso de s: a genome- wide analysis. Lance Lond. Engl. 381, 1371–1379 (2013). 50. Co dell, H. J. e al. In e na ional genome-wide me a-analysis iden ifies new p ima y bilia y ci hosis isk loci and a ge able pa hogenic pa hways. Na . Commun. 6, 8019 (2015). 51. Schunke , H. e al. La ge-scale associa ion analysis iden ifies 13 new suscep ibili y loci o co ona y a e y disease. Na . Gene . 43, 333–338 (2011). 52. Mo is, A. P. e al. La ge-scale associa ion analysis p o ides insigh s in o he gene ic a chi ec u e and pa hophysiology o ype 2 diabe es. Na . Gene . 44, 981–990 (2012). 53. Psychia ic GWAS Conso ium Bipola Diso de Wo king G oup. La ge-scale genome-wide associa ion analysis o bipola diso de iden ifies a new suscep ibili y locus nea ODZ4. Na . Gene . 43, 977–983 (2011). 54. Dubois, P. C. A. e al. Mul iple common a ian s o celiac disease influencing immune gene exp ession. Na . Gene . 42, 295–302 (2010). 55. Tho gei sson, T. E. e al. Sequence a ian s a CHRNB3-CHRNA6 and CYP2A6 a ec smoking beha io . Na . Gene . 42, 448–453 (2010). 56. Pick ell, J. K. e al. De ec ion and in e p e a ion o sha ed gene ic influences on 42 human ai s. Na . Gene . 48, 709–717 (2016). Acknowledgemen s A ull lis o acknowledgemen s can be ound in Supplemen a y No e 2. Au ho con ibu ions C.P., E.H., M.I.M., E.A.S., P.L.L., E.B., D.A., S.J.W., N.D.F., W.H., L.C.P.G.M.D.G., N.M.V.S., N. .V., J.B.R., B.K., T.J.W., D.P.K., R.S.V., C.O., M.L., J.G.E., S.B.K., M.B., E.S., I.H.D.B., J.I.R., S.S.R., M.J., P.K., J.F.W., L.L., K.M., L.Z., H.C., E.T., S.M.F., M.G.D., T.L., M.K., O.T.R., V.M., M.A.I., J.W.J., N.J.W., C.L., N.G.F., K.K., A.B., and J.D. designed and managed indi idual s udies. C.P., E.H., M.I.M., E.A.S., P.L.L., D.A., S.J.W., W.H., N.M.V.S., A.E., J.B.R., R.S.V., C.O., L.V., M.L., J.G.E., S.B.K., M.B., D.V.H., I.H.D.B., J.I.R., S.S.R., J.F.W., L.L., E.I., K.M., H.C., E.T., S.M.F., M.G.D., T.L., M.K., O.T.R., V.M., M.A.I., N.S., J.W.J., N.J.W., C.L., N.G.F., K.K., A.B., and J.D. collec ed da a. E.B., A.E., C.O., M.L., Y.L., M.B., E.M.K., J.I.R., S.S.R., J.F.W., L.L., E.I., K.M., S.T., S.M.F., M.G.D., T.L., L.L., J.W.J., N.J.W., C.L., A.B., and J.D. pe o med he geno yping. C.P., D.B., E.H., M.I.M., W.T., E.B., N.M.V.S., A.E., J.D., C.O., L.V., M.L., K.K.L., J.D., E.M.K., J.I.R., S.S.R., J.F.W., C.H., E.I., K.M., S.T., A.G.U., F.R., L.Z., S.M.F., M.G.D., A.M.V., T.L., L.L., M.A.I., N.S., N.J.W., C.L., A.B., and J.D. p epa ed he geno ype da a. D.B., E.H., M.I.M., E.A.S., P.L.L., A.E., J.B.R., S.B., R.S.V., C.O., L.V., M.L., J.G.E., D.K.H., D.V.H., E.M.K., I.H.D.B., A.C.W., J.F.W., L.L., K.M., M.C.Z., A.G.U., F.R., L.Z., E.T., S.M.F., M.G.D., A.M.V., E.T., L.L., N.J.W., C.L., N.G.F., A.B., and J.D. p epa ed he pheno ype da a. D.B., E.H., M.I.M., J.B.R., T.J.W., D.P.K., Y.H., C.L., A.C.W., C.R.C., P.F.O., M.J., X.J., H.A., N.J.W., C.L., N.G.F., A.B., and J.D. de eloped he analysis plan. A.Z., D.B., E.H., M.I.M., NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-02662-2 ARTICLE NATURE COMMUNICATIONS | (2018) 9:260 |DOI: 10.1038/s41467-017-02662-2 |www.na u e.com/na u ecommunica ions 9