scieee Science in your language
[en] (orig)

CNV-association meta-analysis in 191,161 European adults reveals new loci associated with anthropometric traits

Read accessible full text

CNV-association meta-analysis in 191,161 European adults reveals new loci associated with anthropometric traits

Author: Mace, Aurelien,Tuke, Marcus A,Deelen, Patrik,Kähönen, Mika,Lehtimäki, Terho
Year: 2017
Source: https://trepo.tuni.fi/bitstream/10024/102211/1/CNV_association_meta-analysis_2017.pdf
ARTICLE
CNV-associa ion me a-analysis in 191,161 Eu opean
adul s e eals new loci associa ed wi h
an h opome ic ai s
Au élien Macé e al.
#
The e a e ew examples o obus associa ions be ween a e copy numbe a ian s (CNVs)
and complex con inuous human ai s. He e we p esen a la ge-scale CNV associa ion me a-
analysis on an h opome ic ai s in up o 191,161 adul samples om 26 coho s. The s udy
e eals fi e CNV associa ions a 1q21.1, 3q29, 7q11.23, 11p14.2, and 18q21.32 and confi ms wo
known loci a 16p11.2 and 22q11.21, implica ing a leas one an h opome ic ai . The dis-
co e ed CNVs a e ecu en and a e (0.01–0.2%), wi h la ge e ec s on heigh (>2.4 cm),
weigh (>5 kg), and body mass index (BMI) (>3.5 kg/m2). Bu den analysis shows a 0.41 cm
dec ease in heigh , a 0.003 inc ease in wais - o-hip a io and inc ease in BMI by 0.14 kg/m2
o each Mb o o al dele ion bu den (P=2.5 × 10−10, 6.0 × 10−5, and 2.9 × 10−3). Ou s udy
p o ides e idence ha he same genes (e.g., MC4R,FIBIN, and FMO5) ha bo bo h common
and a e a ian s a ec ing body size and ha an h opome ic ai s sha e gene ic loci wi h
de elopmen al and psychia ic diso de s.
DOI: 10.1038/s41467-017-00556-x OPEN
Co espondence and eques s o ma e ials should be add essed o Z.K. (email: [email p o ec ed]).
#A ull lis o au ho s and hei a flia ions appea s a he end o he pape .
NATURE COMMUNICATIONS |8: 744 |DOI: 10.1038/s41467-017-00556-x |www.na u e.com/na u ecommunica ions 1
Many human an h opome ic ai s a e highly he i able.
Twin s udies ha e es ima ed ha gene ic ac o s
con ibu e o 40–80% o he obse ed a iabili y o
body mass index (BMI)1–5and up o 80% o heigh 6,7. Findings
om he la ges genome-wide associa ion s udies (GWAS) on
BMI8and heigh 9, including o e 250,000 samples, e ealed
97 and 697 single nucleo ide polymo phisms (SNPs) explaining
cumula i ely only 2.7 and 20% o he a iance o he espec i e
pheno ypes. Using geno yping a ays en iched o coding
egions (exome-chip) la ge me a-analysis GWAS o heigh and
BMI disco e ed se e al a e coding single nucleo ide a ian s
(SNVs) associa ed wi h hese ai s. S ill, hese SNVs ha e hus
a explained only a e y small a ia ion in hese ai s (e.g., 0.51%
explained heigh a iance10). Ne e heless, andom e ec
models accoun ing o impe ec impu a ion es ima e ha he
o al addi i e e ec o all SNVs explain 56 and 27% o heigh and
BMI a iabili y, espec i ely11. While he e is a g owing
consensus ha p edominan ly SNVs con ibu e o he he i abili y,
he impac o he s uc u al a chi ec u e o he genome (copy
numbe a ian s, complex ea angemen s, e c.) is unde s udied
and no negligible12. I has been shown ha a e and la ge copy
numbe a ian s (CNVs), such as he 600 kb b eakpoin 4–5
(BP4–BP5) 16p11.2 ea angemen 13,14, can exe subs an ial
impac on BMI, bu li le e o has been made owa ds assessing
he genome-wide impac o CNVs on complex ai s. To ou
knowledge, only one genome-wide CNV-associa ion s udy (on
schizoph enia) has been pe o med in la ge adul popula ion
samples15. The aim o ou s udy is o es ablish a genome-wide
ca alog o CNVs and o iden i y CNVs associa ed wi h heigh ,
weigh , wais - o-hip a io (WHR) and BMI. To his end, we apply
he same CNV calling16 and associa ion pipeline o 25 s udies o
he Gene ic In es iga ion o An h opome ic T ai s (GIANT)
Conso ium combined wi h he UK Biobank and pe o m a
genome-wide associa ion me a-analysis s udy in up o 191,161
un ela ed Eu opean adul s. These analyses show ha o e all CNV
bu den is linked o sho e s a u e and highe WHR. The
genome-wide scans e eal a e a ia ions in se e al genomic
egions (1q21.1, 7q11.23, 3q29, 16p11.2, FIBIN/BBOX1, and
MC4R) o be associa ed wi h an h opome ic measu es. Some o
hese loci ha e a iable equencies ac oss coho s and explain o
o e lap p e ious SNP o a e a ian associa ions. These esul s
highligh he impo an con ibu ion o a e CNVs o complex
human ai s.
Resul s
Summa y o he me hods. All he 25 GIANT coho s we e
geno yped on Illumina a ays, whe eas he UK Biobank used he
A yme ix Axiom chip. Only un ela ed adul samples o Eu -
opean o igin we e included. As PennCNV was ini ially designed
o da a gene a ed on Illumina a ays, we ook ex a ca e wi h he
signal no maliza ion and p e-p ocessing o he UK Biobank da a
(see “Me hods”). Each coho applied ou s anda dised CNV
pipeline o call CNVs16 and o es associa ions be ween p ob-
abilis ic CNV dosages (a con inuous alue be ween −1 (dele ion)
and 1 (duplica ion)) a each p obe on he geno yping chip and
ou a ge an h opome ic ai s. In b ie , ou pipeline combines
pennCNV calls, CNV- and sample pa ame e s o yield a mo e
accu a e p obabilis ic CNV call, especially in he case o a e o
low confidence CNVs. The numbe o p obes a ied be ween
~680,000 and 2,500,000 ac oss he 26 coho s. We hen impu ed
he summa y s a is ics o he Illumina 1M Duo V3 p obe se in
o de o ha e a common se o p obes o he me a-analysis.
Based on an in-house coho we conse a i ely es ima ed he
numbe o e ec i e es s17 o be ~29,400, esul ing in a P- alue
h eshold o 1.7 × 10−6 o con ol amily-wise e o a e
(see “Me hods”). We pe o med a CNV bu den and a genome-
wide CNV associa ion me a-analysis o BMI, weigh , heigh , and
wais –hip a io. The genome-wide CNV associa ion scan was
pe o med conside ing a mi o e ec model (assuming opposi e
and equal sized e ec o dele ions and duplica ions a any gi en
locus) and he genome-wide significan signals we e u he es ed
o dele ion-only and duplica ion-only e ec s. As seconda y
analysis, we also es ed U-shaped (assuming he same e ec o
dele ions and duplica ions), dele ion-only, duplica ion-only
models genome-wide. All epo ed CNV e ec sizes (unless
specified o he wise) ep esen he impac o one addi ional copy
ela i e o he popula ion a e age. Fo bu den analysis all ou
abo emen ioned models (mi o , U-shaped, dele ion, duplica-
ion) we e es ed. Depending on he ai , he sample sizes a ied
be ween 161,244 and 191,161.
To al CNV bu den. The inc eased bu den o a e CNVs has
al eady been obse ed o pe sons wi h sho s a u e18, highe
BMI19, and also schizoph enia15. Indi ec ly, inc eased dele ion
bu den is also eflec ed in longe egions wi h loss-o -he e o-
zygosi y, which has shown o associa e wi h s a u e and
Table 1 Lis o he CNVs associa ed wi h one o se e al ai s
Ch S a End F equency (%) BMI Weigh Heigh Wais –hip a io
(Mb) (Mb) Del Dup βP alue βP alue βP alue βP alue
1 145 145.9 0.03 0.049 –– 6.66 1.73E−06 3.46 3.75E−10 ––
3 197.7 197.9 0.004 0.005 –– 22.55 1.20E−06 –– – –
3 198.2 198.4 0.007 0.007 –– – – 13.3 2.32E–08 ––
7 72.61 72.75 0.005 0.005 –– – – –– 0.11 1.49E–06
11 26.97 27.19 0.126 0.011 –– – – 2.43 1.46E−06 ––
16 28.73 28.95 0.028 0.041 −3.07 5.31E−08 −10.35 5.03E−09 –– – –
16 29.5 30.1 0.027 0.031 −3.66 1.39E−12 –– 5.21 1.20E−14 −0.041 2.30E−07
18 55.81 56.05 0.018 0.004 −5.06 2.03E−07 15.9 1.45E−08 –– – –
All posi ions a e hg18 in megabase (Mb). In case o genome-wide significan CNV- ai associa ions (P<1.7 × 10−6), we epo e ec sizes (β) and P alues coming om a mi o e ec model, assuming
opposi e and equal sized e ec o dele ions and duplica ions a any gi en locus. Fu he in o ma ion and esul s om o he models a e a ailable in Supplemen a y Table 2. The e ec s co espond o
change in he ai o each addi ional copy o he egion: posi i e e ec means ha dele ion o he co esponding egion dec eases he ai alue and duplica ions inc ease i . The genes in ol ed in hese
egions a e as ollows: RN7SL261P, RNVU1-8, CHD1L, NBPF13P, GJA8, OR13Z3P, LINC00624, OR13Z2P, OR13Z1P, PDIA3P1, FMO5, RPL7AP15, CCT8P1, PRKAB2, GJA5, GPR89B, BCL9, ACP6, (Ch 1:145–146 Mb);
PIGX, (Ch 3:197.7–197.9 Mb); DLG1, MFI2, MFI2-AS1, (Ch 3:198.2–198.4 Mb); VPS37D, DNAJC30, WBSCR22, MLXIPL, (Ch 7:72.61–72.75 Mb); FIBIN, BBOX1, BBOX1-AS1, (Ch 11:26.97–27.19 Mb); MIR4721,
MIR4517, ATXN2L, SH2B1, CD19, RABEP2, TUFM, ATP2A1, NFATC2IP, ATP2A1-AS1, LAT, SPNS1, (Ch 16:28.7–29 Mb); MIR3680-2, RN7SKP127, C16o 54, PAGR1, CORO1A, MAZ, ALDOA, CDIPT, MVP, ZG16,
SEZ6L2, CDIPT-AS1, PRRT2, YPEL3, TMEM219, DOC2A, GDPD3, INO80E, KCTD13, HIRIP3, ASPHD1, MAPK3, TAOK2, PPP4C, FAM57B, C16o 92, SMG1P2, SLC7A5P1, CA5AP1, QPRT, SPN, TBX6, KIF22,
(Ch 16:29.5–30.1 Mb); RNU4-17P, RNU6-567P, SDCCAG3P1, FAM60CP, RPS3AP49, (Ch 18:55.8–56.1 Mb).
ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-00556-x
2NATURE COMMUNICATIONS |8: 744 |DOI: 10.1038/s41467-017-00556-x |www.na u e.com/na u ecommunica ions
cogni ion20. In his s udy we confi m he link be ween CNV
bu den, measu ed as he o al numbe o copy a ian p obes, and
heigh and BMI and we also ound an e ec on he wais –hip
a io (Supplemen a y Table 1). Indi iduals wi h an addi ional 1
Mb o copy-al e ed in e al end o ha e 0.144 kg/m2highe BMI
(P=2.9 × 10−3). The e ec o CNV bu den was much s onge on
wais –hip a io and heigh , o which inc eased CNV bu den
(be i duplica ion o dele ion) was associa ed wi h a 0.001
highe WHR (P=6.9 × 10−5) and 0.132 cm sho e s a u e (P=
4.5 × 10−7). Fo bo h ai s he impac was domina ed by he
bu den o dele ions a he han duplica ions (0.003 WHR uni
(P=6×10
−5) and 0.41 cm (P=2.5 × 10−10) pe Mb dele ion,
espec i ely). We did no obse e any CNV bu den e ec on
human weigh .
Genome-wide scan. The analyses on hese ou an h opome ic
ai s e ealed se en independen CNV egions associa ed wi h
one o se e al ai s wi h P alue below he genome-wide sig-
nificance h eshold (1.7 × 10−6, see “Me hods”) (Table 1, Fig. 1).
Two o hem co espond o he well known BP2–BP3 and
BP4–BP5 CNVs in he 16p11.2 egion associa ed wi h BMI and
neu ode elopmen . Th ee u he CNVs (1q21.1, 3q29, and
7q11.23) o e lap ecu en synd omic CNV egions, associa ed
wi h a iable neu ode elopmen al ai s, schizoph enia and
de elopmen al delay. One CNV (nea MC4R) o e laps wi h SNPs
associa ed wi h BMI in GWAS and is pa o a la ge dele ion
epo ed o be associa ed wi h obesi y21. And finally one dele ion,
encompassing BBOX1 and FIBIN genes ( he la e ha bo ing a e,
heigh -lowe ing coding a ian s10), seems o be pa icula ly e-
quen in he Finnish popula ion (0.89% s 0.02% in he non-
Finnish coho s). In he ollowing we p o ide a de ailed
desc ip ion o he impac o each o hese CNVs (bo h dele ions
and duplica ions). We es ed U-shaped, dele ion-only,
duplica ion-only models, bu hese did no yield u he
significan associa ions (Supplemen a y Table 2).
New insigh s on he 16p11.2 egion. The 16p11.2 egion is well
known o se e al dis inc ecu en CNVs, wo o hem asso-
cia ed wi h an h opome ic ai s. The 220 kb BP2–BP3 dele ion
was associa ed wi h se e e ea ly-onse obesi y and de elopmen al
delay22. The 600 kb BP4–BP5 ea angemen was fi s known o
i s impac on au ism bu i has also been p o en o ha e e ec s on
BMI and head ci cum e ence13,14. Bo h ha e been ecen ly
epo ed as associa ed wi h lowe IQ and schizoph enia15. Fi s ,
we eplica ed he known e ec s o he 220 kb dele ion (β=+ 3.07
kg/m2,P=5.3 × 10−8) and he mi o e ec o he 600 kb ea -
angemen (mi o e ec : β=−3.66 kg/m2,P=1.4 × 10−12; dele-
ion: β=6.15 kg/m2,P=4.5 × 10−14; duplica ion: β=−1.81
kg/m2,P=1.2 × 10−2) on BMI. In addi ion, we ound ha while
he 220 kb dele ion inc eases BMI h ough inc easing weigh (by
10.35 kg, P=5×10
−9), he 600 kb dele ion does so by bo h
dec easing heigh (by 5.21 cm, P=1.1 × 10−14) and inc easing
weigh (6.57 kg, P=5.3 × 10−5) (Supplemen a y Figs. 1–4). Fu -
he mo e, ou analysis e ealed ha he 600 kb ea angemen
also impac s wais –hip a io (β=−0.04, P=2.3 × 10−7) (Supple-
men a y Figs. 3C–4C). Nei he analyzing dele ions and duplica-
ions sepa a ely, no hei absolu e e ec showed s onge signal
o he 16p11.2 220 kb ea angemen han he mi o e ec
associa ion. On he con a y, he obse ed e ec om he 600 kb
seems almos exclusi ely d i en by he dele ion, which demon-
s a ed a s onge signal han he duplica ions o he pooled
esul s (Supplemen a y Table 2). The lis o he Online Mendelian
Inhe i ance in Man (OMIM) diseases co esponding o he genes
p esen in hese wo CNVs is a ailable in Supplemen a y Table 3.
The op associa ions be ween hese CNVs and 27 es ed ai s in
he UK Biobank a e lis ed in Supplemen a y Tables 4–7.
We could no na ow down he BMI associa ion signal o he
p e iously p oposed SH2B1 (lowes P=7.7 × 10−8) as i co e s
o he genes, including SPNS1 and LAT (lowes P=5.3 × 10−8)
(Fig. 2). Fine-mapping o he signal using a iable b eakpoin s
would be necessa y, which a e ex emely a e due o he egional
12 14
12
10
8
6
4
2
0
10
8
6
–log10(P)
–log10(P)
–log10(P)
–log10(P)
4
2
0
8
7
6
5
4
3
2
1
0
6
4
2
0
12 4 6
Ch omosome
Ch omosome Ch omosome
Weigh
BMI Heigh
Wais –hip a io
Ch omosome
811 14 18 1 2 4 6 8 11 14 18
1246
811 14 18
12 4 6 8 11 1418
Fig. 1 Genome-wide Manha an plo s o ou an h opome ic ai s. Genome-wide associa ion s udy o CNVs associa ed wi h BMI, heigh , weigh , and
wais - o-hip a io in 191,161 Eu opeans
NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-00556-x ARTICLE
NATURE COMMUNICATIONS |8: 744 |DOI: 10.1038/s41467-017-00556-x |www.na u e.com/na u ecommunica ions 3
a chi ec u e shaped by segmen al duplica ions and non-allelic
homologous ecombina ions (Supplemen a y Fig. 2).
Resul s om p e ious GWAS e ealed BMI-associa ed SNPs
nea SH2B1 loca ed in he 220 kb ea angemen and heigh -
associa ed SNPs nea FLJ25404 loca ed in he 600 kb ea ange-
men highligh ing he impo ance o bo h common and a e
a ian s in hese egions (Supplemen a y Table 8). Nex , we es ed
whe he he p e iously published h ee independen BMI-
associa ed SNPs a his locus ( s3888190, s2650492, and
s4787491) could be explained by he CNV associa ions in he
16p11.2 egion using he UK Biobank da a o which bo h SNPs
and CNV calls we e a ailable. Ou analysis showed ha he
o iginal BMI-SNP associa ion P alues inc eased subs an ially
( om 2.28 × 10−8, 6.45 × 10−5, 8.67 × 10−6 o 2.94 × 10−4, 3.80 ×
10−3, 1.56 × 10−2, espec i ely) when he mos BMI-associa ed
CNV p obe was included in he mul i a ia e model. In he
mean ime he BMI-CNV associa ion signals emained
unchanged. Simila ly, he heigh -FLJ25404 ( s11642612) associa-
ion P alue inc eased mo e han 100- old om P=1.5 × 10−5 o
P=4.35 × 10−3when including he 16p11.2 CNV p obe wi h
s onges heigh associa ion, indica ing ha he p e iously
obse ed SNP-heigh associa ion may be (a leas pa ially)
explained by he 16p11.2 CNV-heigh associa ion.
Cis eQTL analysis o heigh - and BMI-associa ed SNPs loca ed
in he 600 kb ea angemen showed a po en ial e ec o hese
SNPs modula ing he exp ession le els (in whole-blood) o
CORO1A ( s11150581, s11642612) and INO80E ( s11150581,
s11642612, s2278557, s6565173, s9925915) genes (Supple-
men a y Table 9).
1q21.1 dis al ea angemen . A CNV egion on ch omosome 1
(145–145.9 Mb, Supplemen a y Figs. 5and 6) was associa ed wi h
bo h heigh (β=3.46 cm, P=3.8 × 10−10) and weigh (β=6.66
kg, P=1.7 × 10−6). This ea angemen co esponds o he dis al
pa o he 1q21.1 ecu en CNV (OMIM dele ion: #612474;
OMIM duplica ion: #612475). As o he 16p11.2 600 kb CNV,
his CNV is known o ha e a mi o e ec on head ci cum e ence
and o be a po en ial cause o au ism and schizoph enia15,23.An
e ec on heigh has been epo ed o he dele ion, wi h 25–50%
o he ca ie s ha ing sho s a u e24. In con as , duplica ion
ca ie s end o be in he uppe pe cen iles o heigh bu he e ec
is less clea . Suppo ing he e ec on heigh , a common a ian in
his egion, nea he FMO5 gene ( s6658763) was associa ed wi h
heigh in p e ious GWAS9(Supplemen a y Table 8). This SNP
was no significan ly associa ed wi h heigh in he UK Biobank
(P=0.14), and hus, no condi ional analysis was pe o med.
Howe e , he SNP seems o be independen o he CNV (Dele-
ion: 2=0, Dʹ=0.005−Duplica ion 2=0, Dʹ=0.022 in he UK
Biobank). As o he 16p11.2 600 kb ea angemen , he obse ed
e ec s seem o be mainly due o he dele ions (Supplemen a y
Table 2). The lis o he OMIM diseases co esponding o he
genes p esen in he CNV is a ailable in Supplemen a y Table 3.
The op associa ions be ween his CNV and 27 es ed ai s in he
UK Biobank a e lis ed in Supplemen a y Tables 10 and 11.
A CNV o e lapping FIBIN and BBOX1. A 220 kb CNV (ch 11:
26.97–27.19 Mb, Supplemen a y Figs. 7and 8) was associa ed
wi h heigh (β=2.43 cm, P=1.5 × 10−6). While he duplica ion
equency is low in all coho s (0.008–0.016%), he dele ion e-
quency is much highe in he Finnish popula ion han in he
o he s (0.89% s 0.016%). This egion has added in e es , because
a case- epo desc ibed an I anian sho -s a u ed gi l wi h
homozygous dele ion o his egion25. Sepa a e analysis o he
dele ions and duplica ions showed a highly significan e ec om
he dele ions (β=2.56 cm, P=8.2 × 10−8). The in ol emen o
he FIBIN gene o heigh was also confi med by he GIANT-
exome s udy on heigh 10 including 381,625 indi iduals. This
s udy e ealed a s ong associa ion be ween he a e (0.3% in
ExAC) missense a ian s138273386, loca ed in he FIBIN gene,
and heigh (P=5.79 × 10−12) (Supplemen a y Table 12). The op
0
1
2
3
4
5
6
7
8
9
10
P alue (−log10)
0
0.05
0.1
F equency
−1.25
−1.00
−0.75
−0.50
−0.25
0.00
E ec size
ATXN2L
TUFM
SH2B1
ATP2A1
RABEP2
CD19
NFATC2IP
SPNS1
LAT
28700000 28750000 28800000 28850000 28900000 28950000
P obes posi ions
Fig. 2 Regional associa ion plo o he 16p11.2 220 kb ea angemen . The blue do s ep esen −log10 BMI-associa ion P alues, he ed do s show he
co esponding e ec sizes. A he bo om he black and g ay lines a e he dele ion and duplica ion equencies. Finally, he do s a he bo om indica e he
p obe posi ions o he GIANT coho s (abo e) and he UK Biobank (below). Posi ions o he p o ein-coding genes a e shown a he op o he plo . The
p obes posi ions co espond o he human genome build 36
ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-00556-x
4NATURE COMMUNICATIONS |8: 744 |DOI: 10.1038/s41467-017-00556-x |www.na u e.com/na u ecommunica ions
associa ions be ween his CNV and 27 es ed ai s in he UK
Biobank a e lis ed in Supplemen a y Tables 13 and 14.
CNV in he MC4R egion. Single nucleo ide coding mu a ions in
he MC4R gene cause se e e obesi y, and common a ian s nea
he gene a e associa ed wi h BMI8,26,27. Ou analysis e ealed a
a e ( equency 0.018% (del), 0.004% (dup)), 300 kb long CNV
(55.81–56.05 Mb, Supplemen a y Figs. 9and 10) associa ed wi h
BMI (β=−5.06 kg/m2,P=2×10
−7) and weigh (β=−15.94 kg,
P=1.4 × 10−8). Follow-up analysis demons a ed ha he
obse ed signal is exclusi ely due o dele ions (Supplemen a y
Table 2). This CNV encompasses he BMI-associa ed lead SNP
( s6567160)8,28 (Fig. 3, Supplemen a y Table 8), bu we obse ed
i ually iden ical BMI-associa ion P alues o he SNP and he
CNV in uni a ia e and mul i a ia e analysis. Hence, he wo
associa ions a e mos p obably independen , u he e idenced by
he low LD be ween hem ( 2=0.0014, D′=0.31 in he UK
Biobank). While a e heigh -inc easing MC4R a ian s27 ha e
been p e iously epo ed, we ound no heigh -e ec o any CNV
p obes in his egions. P e ious e idence o CNVs a ec ing BMI
in he MC4R gene is sca ce— he e is only one case epo o a
9-yea -old obese boy ca ying a la ge (2.6 Mb) dele ion encom-
passing he MC4R gene21. The op associa ions be ween his CNV
and 27 es ed ai s in he UK Biobank a e lis ed in Supplemen-
a y Tables 15 and 16.
7q11.23 ea angemen . The only WHR-specific genome-wide
significan associa ion implica ed a CNV in he 7q11.23 egion
(72.6–73.58 Mb, Supplemen a y Figs. 11 and 12). Duplica ion
ca ie s end o ha e a highe wais -hip a io (β=0.11,
P=1.5 × 10−6). Sepa a e analysis o dele ions and duplica ions
and absolu e e ec associa ion did no show any s onge asso-
cia ion, ne e heless, he e ec o he duplica ion is sligh ly la ge
han ha o he dele ion (Supplemen a y Table 2). This CNV was
ecen ly ound o be associa ed wi h schizoph enia in a la ge
case–con ol s udy15. I also o e laps wi h a 1.55–1.84 Mb long
egion known as he Williams–Beu en (WB) synd ome c i ical
egion (WBSCR)29,30. WB synd ome31,32 is esponsible o
se e al complica ions: ca dio ascula disease, neu ologic
abno mali ies, a en ion defici hype ac i i y diso de , cogni i e
impai men , dis inc i e beha io al, and social ai s. Due o
selec ion bias, ou p e alence es ima ion o he duplica ion
(0.005%) and he dele ion (0.005%) is somewha lowe han wha
is es ima ed o he WBSCR in he li e a u e (0.005–0.013% o
he duplica ion33,34 and 0.008–0.013% o he dele ion35). The
op associa ions be ween his CNV and 27 es ed ai s in he UK
Biobank a e lis ed in Supplemen a y Tables 17 and 18.
3q29 ea angemen . We disco e ed wo CNVs in he 3q29
egion, one 256 kb long (197.6–197.9 Mb), a ec ing weigh (β=
22.55 kg, P=1.6 × 10−6, dele ion equency =0.004%, duplica-
ions equency =0.005%) and one 212 kb long (198.2–198.4 Mb)
a ec ing heigh (β=13.3 cm, P=2.3 × 10−8, dele ion equency
=0.007%, duplica ions equency =0.007%) (Supplemen a y
Figs. 13 and 14). Running he me a-analysis sepa a ely on he
UKBB and he o he coho s, i appea s ha he signal comes
mainly om he UK Biobank, howe e , wi hou e idence o
s ong he e ogenei y (Coch an P>0.05). The p opo ional e ec s
on heigh and weigh a e conco dan wi h he ac ha no asso-
cia ion has been ound wi h BMI o WHR. Child en wi h his
3q29 dele ion su e om eeding p oblems, which may esul in
educed adul weigh 36. A ecu en synd omic CNVs encom-
passing he wo segmen s has ecen ly been epo ed o be asso-
cia ed wi h schizoph enia15. On he an h opome ic aspec , case
epo s om he li e a u e a e in ag eemen wi h ou findings
ega ding he dele ion impac on bo h weigh and heigh 24,37,38.
Conce ning he duplica ion, he pheno ype spec um is wide and
he li e a u e mainly epo s obese/o e weigh cases, which is in
ag eemen wi h ou weigh es ima es. The epo ed e ec on
heigh is less p onounced. Ou (median) dele ion equency
0
1
2
3
4
5
6
7
8
9
10
P alue (−log10)
0
0.1
0.2
0.3
F equency
–2.5
–2.0
–1.5
–1.0
–0.5
0.0
E ec size
RPS3AP49
MC4R
55600000 55800000 56000000 56200000 56400000
P obes posi ions
Fig. 3 Regional associa ion plo o he ea angemen nea MC4R. The blue do s ep esen –log10 BMI-associa ion P alues, he ed do s show he
co esponding e ec sizes. A he bo om he black and g ay lines a e he dele ions and duplica ions equencies. Finally, he do s a he bo om a e he
p obes posi ions o he GIANT coho s abo e and he UK BioBank below. Posi ions o he p o ein-coding genes a e shown a he op o he plo along wi h
he posi ion o he BMI-associa ed GWAS SNP. The p obes posi ions co espond o he human genome build 36
NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-00556-x ARTICLE
NATURE COMMUNICATIONS |8: 744 |DOI: 10.1038/s41467-017-00556-x |www.na u e.com/na u ecommunica ions 5

(0.004 and 0.007%, o he wo segmen s espec i ely) is sligh ly
o e he epo ed alue in a con ol popula ion (0.003%)36.
Finally, upon close inspec ion o he egion (Supplemen a y
Fig. 14) we obse ed ha he cen ome ic pa o he CNV is
implica ed in weigh egula ion, while he elome ic end impac s
heigh (implica ed genes a e lis ed in he legend o Table 1).
Dele ion and duplica ion equencies we e oo low o be able o
eliably es ablish he e ec s o dele ions and duplica ions sepa-
a ely. The op associa ions be ween his CNV and 27 es ed ai s
in he UK Biobank a e lis ed in Supplemen a y Tables 19–22.
CNVs wi h a iable equency ac oss geog aphic loca ions. Ou
me a-analysis e ealed wo popula ion-specific CNVs. The fi s
one is he CNV o e lapping FIBIN and BBOX1, o which Finnish
popula ion coho s ha e much highe dele ion equency. The
second one, nea MC4R, is specific o UK popula ion coho s. In
bo h cases we compa ed po en ial con ounding ac o s, such as
p obe densi ies, call a es, and CNV quali y, bu none o hese
could explain he equency di e ences (Supplemen a y Figs. 8
and 10). No e ha he equency o he MC4R CNV bo h in he
UK Biobank (0.028% (del), 0.005% (dup)) and in o he UK
coho s geno yped on Illumina a ays (0.018% (del), 0.009%
(dup)) is consis en ly highe han he equency in non-UK
samples (0.006% (del), 0.005% (dup)). Thus, he obse ed e-
quency di e ence is, a leas in pa , no due o a ay e ec .
Discussion
Ou genome-wide CNV associa ion me a-analysis on ou
an h opome ic ai s in ~190,000 un ela ed adul s showed a non-
negligible CNV bu den e ec on BMI, heigh , and WHR. Fu -
he mo e, we iden ified se en CNVs significan ly associa ed wi h
a leas one ai and h ee addi ional CNV egions ha e a close o
genome-wide significan e ec on one o he ou ai s
(Supplemen a y Table 23). The analysis also ga e new insigh s
in o he wo 16p11.2 ea angemen s13,14,22.
As a p oo o concep , we looked a CNVs known o be
associa ed wi h BMI o obesi y39,40. Only one CNV (22q11.21)
ou side he 16p11.2 egion was confi med (Supplemen a y
Fig. 15, Supplemen a y Tables 24 and 25). This di e ence migh
in pa be explained by insu ficien powe o by he ac ha ,
con a y o mos p e iously published CNV s udies13,15,39,41,
ou samples come om gene al popula ions. In addi ion, some o
he p e iously epo ed CNVs migh ha e been popula ion-
specific o simply spu ious42.
CNV bu den analysis confi med he al eady obse ed e ec on
BMI19 and heigh 18, and showed an impo an e ec on a dis-
ibu ion (WHR). These obse ed signals a e domina ed by he
dele ions (up o fi e- old la ge e ec s), while he duplica ion
e ec s a e mino , excep o WHR. The o e all CNV bu den has
seemingly opposi e e ec s on heigh and BMI, compa ible wi h
ha ing no significan CNV bu den e ec on weigh .
O e all, he genome-wide P alues showed good adhe ence o
he null dis ibu ion (Supplemen a y Fig. 16). Fo well-powe ed
GWAS s udies on he i able ai s (e.g., heigh 9and mena che43)
high genomic con ol lambda alue a he eflec s ue polygenic
signals han unco ec ed popula ion s a ifica ion44. This was he
case o ou s udy oo: while we obse ed infla ed genomic
lambda coe ficien s (λ=1.16 (heigh ), λ=1.12 (weigh ), λ=1.08
(BMI) and λ=1.05 (WHR)), upon applying LD sco e eg es-
sion44 in he UK Biobank sample he in e cep e ms e ealed no
unaccoun ed popula ion s a ifica ion (λ
LD
=0.971(heigh ),
λ
LD
=1.005(weigh ), λ
LD
=0.993 (BMI), λ
LD
=0.942 (WHR)).
Among he se en significan CNVs, wo migh be ances y-
specific, one Finnish and one B i ish. I is no su p ising, as hese
wo popula ions ha e con ibu ed he mos samples o ou me a-
analysis. These esul s show he need o collec ing la ge popu-
la ion coho s o he same o igin since he equency o many
CNVs may a y ac oss popula ions. The e o e, we belie e ha in
he u u e, collec ing la ge , geno yped popula ion-based coho s
om o he coun ies and e hnici ies could be an e ficien way o
disco e no el ai -associa ed CNVs wi h la ge e ec s.
Al hough we would ha e had highe powe o de ec associa-
ions wi h common CNVs, all o he an h opome ic ai -
associa ed CNVs iden ified in his s udy a e a e (0.01–0.07%).
This may be explained by he massi e shi o CNV equency
spec um compa ed o ha o SNVs: based on CNV calls om
>191,000 samples we obse ed ha mo e han 92.4% o he
CNVs a e p esen in <1 in 1000 samples and 99.4% o hem a e
a e (<1%). We a e unsu e whe he he eason o he e y low
numbe o common CNVs is due o he de ec ion echnology o
whe he i eflec s he na u e o he unde lying genomic e en s.
Gi en he low equency o mos o he disco e ed CNVs and he
neighbo ing genome s uc u e, he majo i y o hem may esul
om de no o and ecen ea angemen s. The o al explained
a iance o all hese ea angemen s is ∼0.09% o BMI, 0.10%
o weigh , 0.14% o heigh , and 0.04% o wais –hip a io.
Ou condi ional analysis showed ha CNV p obes in he
16p11.2 egion explain a subs an ial ac ion o he associa ion
be ween all p e iously published SNPs nea SH2B1 and BMI and,
simila ly, he associa ion be ween he SNP nea FLJ25404 and
heigh . None o he emaining associa ed CNVs showed e idence
o agging common SNPs, no do hey explain known heigh /
BMI-SNP associa ions. No e, howe e , ha ou CNV da a a e
much noisie han SNP calls and hus he measu ed CNVs a e
poo e p oxies o he ue CNV s a us, which biases he condi-
ional analysis owa ds he null (no agging). S ill, mos o he
ob ained esul s a e in line wi h he p oposed heo y ha he
majo i y o he disco e ed disease-associa ed common SNPs a e
no syn he ic associa ions due o a e a ian agging45.
Ou s udy, besides epo ing he associa ion wi h an h opo-
me ic ai s, can se e as an a las o CNV maps based
on a la ge gene al popula ion o Eu opean ances y46,47 (h ps://
cn ca alogue.bbm i.nl/ and unde lying da a in Supplemen a y
Da a 1). Simila ly o la ge compendia o sequenced popula ion
indi iduals (e.g., EXaC48) o whole exome-/genome-sequence
analysis o a e diseases, ou in en o y o CNV equencies could
help es ima ing hei pa hogenici y in he a e disease se ing.
So a , many an h opome ic GWASs ha e ocused on BMI o
heigh , bu less on weigh . In ou analysis we ound ha s udying
he e ec o CNVs on heigh and weigh sepa a ely can ca y
impo an addi ional in o ma ion beyond wha we can lea n om
looking only a BMI. All he CNVs ound o be associa ed wi h
BMI we e also associa ed ei he wi h heigh o weigh , bu he
opposi e does no hold. CNVs a ec ing heigh and weigh in he
same di ec ion (e.g., 1q21.1) ha e less impac on BMI.
Ou s udy has se e al weaknesses, which we ied o mi iga e.
Despi e he ac ha a ple ho a o so wa e has been de eloped o
de ec CNVs om SNP a ay pla o ms, hese geno yping chips
we e no ini ially designed o his pu pose. This d awback
educes s a is ical powe in ou analysis by in oducing alse
posi i e and alse nega i e CNV calls. Impo an ly, he e is no
pa icula eason o belie e ha CNV calling a e ac s appea
specifically o samples en iched o low/high ai alues. Thus,
we belie e ha alse CNV calls do no ansla e o alse posi i e
findings, bu o cou se can subs an ially educe s a is ical powe .
We did no pe o m independen (e.g., qPCR) expe imen s o
confi m hese CNV findings, bu p o ided se e al lines o e i-
dences o suppo ou claims: (i) mos o ou epo ed CNVs ha e
been epo ed be o e wi h simila equency and b eakpoin s; (ii)
many o ou CNVs all in o egions al eady associa ed wi h
obesi y; (iii) QQ-plo s o all ai s show excellen adhe ence o
ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-00556-x
6NATURE COMMUNICATIONS |8: 744 |DOI: 10.1038/s41467-017-00556-x |www.na u e.com/na u ecommunica ions
he null o he bulk o he CNV p obes (Supplemen a y Fig. 16);
(i ) ou op CNVs show li le o no he e ogenei y ac oss s udies;
( ) coho s used 15 di e en geno yping a ays elimina ing a ay-
specific a e ac s. C ucially, hese geno yping pla o ms a e mo e
eliable han low-co e age sequencing o in e CNVs and hus
emain he mos cos -e ficien o pe o m such la ge CNV-
associa ion s udies. Ano he limi a ion is selec ion bias: As shown
in he esul s, many o he an h opome ic ai -associa ed CNVs
we disco e ed a e synd omic and we e al eady obse ed in
pa ien s wi h specific genomic diso de s (including ai s like e.g.,
de elopmen al delay41, schizoph enia15, e c.). In such si ua ions
we canno dis inguish whe he he e ec o hose a e me e con-
sequence o he p ima y synd omes o igge molecula
mechanisms ha ac independen ly on an h opome ic ai s.
No e, howe e , ha his c i icism is alid o any GWAS. The
an h opome ic e ec s o mos o he synd omic CNVs a e o en
poo ly epo ed due o he small numbe o cases. We ound
e idence ha 1q21.1 duplica ion ca ie s all o he ~80 h
popula ion heigh pe cen ile49, equi alen o ~4 cm o heigh
inc ease, close o ou obse ed e ec o 3.46 cm. Ca ie s o he
22q11.2 dele ion ha e on a e age 3 kg/m2highe BMI by he age
o 2050, which is compa able o ou es ima ed e ec o 4 kg/m2.
Mo eo e , ou pheWAS analysis on 27 ai s in he UK Biobank
did no iden i y any non-an h opome ic ai o be s onge
associa ed wi h he disco e ed CNVs. Fu he mo e, we could no
iden i y any end be ween e ec s on an h opome ic ai s and
schizoph enia (Supplemen a y Table 26), indica ing ha he
an h opome ic associa ions we obse ed canno be seconda y o
schizoph enia. These lines o e idence indica e ha mos o he
disco e ed CNVs a ec an h opome ic ai s ei he p ima ily o
in a disease-independen ashion. Impo an ly, he samples we
analyzed come om popula ion-based coho s, heal hie han he
gene al popula ion51. Selec ion bias, hus, emo es many ca ie s
o CNVs wi h la ge e ec s52, implying ha he e ec s seen in ou
s udy a e po en ially smalle han he eal ones. A u he lim-
i a ion is he a iabili y in he equency o such a e CNVs ac oss
popula ions. This phenomenon may ende some o hese dis-
co e ies di ficul o eplica e ac oss popula ions, as no only
simila ly la ge eplica ion s udies would be necessa y, bu also
popula ions in which he CNV equency is high enough o yield
su ficien s a is ical powe . The final weakness o men ion is ha
— o educe coho analys bu den—al hough we adjus ed he
analyzed ai s o gende , bu did no pe o m sex-specific ana-
lysis, which will be he cen al ocus o a u u e s udy.
CNV s udies in gene al popula ion coho s ha e been neglec ed
in he pas due o da a a ailabili y issues. We ha e shown ha
such s udies a e easible h ough a ca e ul e-analysis o exis ing
geno yping da a. Ou s udy has iden ified se e al heigh - and
obesi y-associa ed a e CNVs wi h subs an ial e ec . We hope
ha ou s udy will open new a enues o esea ch o unde s and
he impac o CNVs on human heal h on an unp eceden ed scale.
The pipeline used o his me a-analysis could be applied wi h any
o he ype o quan i a i e ai , and, wi h some modifica ion, o
any bina y ai . Gi en he conside able o e lap be ween CNVs
associa ed wi h an h opome ic ai s, de elopmen al delay and
schizoph enia, in he u u e i would be insigh ul o swi ch poin
o iew and apply a pheWAS app oach in la ge, pheno ype- ich
coho s such as he UK BioBank, allowing deepe in e p e a ion
o candida e CNVs.
Me hods
Coho s. We conduc ed he me a-analysis o BMI (N UKBB =119,873, N GIANT
=71,288), weigh (N UKBB =119,767, N GIANT =55,416), heigh (N UKBB =
116,259, N GIANT =65,706) and wais -hip a io (N UKBB =119,867, N GIANT
=41,377) (Supplemen a y Table 27). All GIANT samples we e geno yped on
Illumina pla o ms and he UK BioBank was geno yped on A yme ix Axiom.
Pa icipan s o each coho ha e signed he in o med consen o m o he
espec i e s udy. In addi ion o he e hical commi ee app o al o each indi idual
s udy, his me a-analysis e o was also app o ed by he s ee ing commi ee o he
GIANT Conso ium and he e hical e iew boa d o he UK Biobank (applica ions
#17085, #16389, #9072). Only un ela ed adul s (gene ic ela edness <0.1) o
Eu opean ances y we e included in he s udy.
CNV calling. Fo he A yme ix Axiom chip, addi ional p ep ocessing was
necessa y: Raw p obese in ensi y alues we e quan ile-no malized. B iefly, in en-
si ies we e so ed nume ically ac oss each ch omosome wi h missing alues being
alloca ed an o e all median alue o acili a e no maliza ion. The mean in ensi y
ac oss each geno yping ba ch was hen calcula ed o each so ed posi ion. Mean
in ensi ies we e hen subs i u ed in place o he equi alen ly anked aw in ensi ies
whils igno ing missing alues. Each ans o med in ensi y alue was hen log
2
ans o med o p ocessing in PennCNV-A y. Mean alues o all in ensi ies ac oss
each ch omosome we e checked o ensu e ha hey we e he same wi hin each
geno yping ba ch o he UK Biobank. PennCNV-A y was used o in e geno ype
clus e s, gene a e Log R Ra io (LRR) and B-Allele F equencies (BAF). All o he
coho s used Illumina a ays, so LRR and BAF alues we e eadily a ailable.
We de ised a pipeline ha akes as inpu no malized BAF and LRR o each
p obe and sample (using he PennCNV so wa e53), assigns p obabilis ic CNV
calls16 and uns associa ion wi h each ai . The pipeline c ea ed a “popula ion B
allele equency”(PFB) file o each coho based on 200 andomly selec ed final
epo s. Adjacen CNVs wi h small gaps (gap sho e han 20% o he o al leng h)
we e me ged using he de aul PennCNV pa ame e s. Finally, samples wi h mo e
han 200 CNVs we e excluded om u he analysis.
Fo each CNV p obe he pipeline compu ed a quali y sco e (QS)16. The QS,
based on he PennCNV quali y me ics, es ima es (Supplemen a y Table 28) he
p obabili y o a pennCNV call o be a ue posi i e CNV call. I is a con inuous
alue be ween −1 and 1, ep esen ing he p oduc o he ela i e copy numbe
(+1 o duplica ion, −1 o dele ion) and he p obabili y o he call being ue, i.e., i
is he expec ed copy numbe dosage ela i e o he copy neu al (2 copy) s a e. The
QS is compu ed o each p obe jand sample i(QS
i,j
) and used as geno ype alue o
CNV- ai associa ion assuming a mi o e ec o dele ions and duplica ions. Since
he QS accoun s o a ious CNV cha ac e is ics (leng h, numbe o p obes, e c.)
we did no apply any fil e ing on hese sco es, which was shown o be he mos
powe ul s a egy o associa ion16. Howe e , p obes wi h low impu a ion quali y
(see below) a e fil e ed ou in each coho .
CNV associa ions wi h an h opome ic ai s. We ocussed on he ollowing
an h opome ic ai s: BMI, weigh , heigh and wais –hip a io. BMI (kg/m2),
weigh (kg), and wais –hip a io we e adjus ed o sex, age, age2and he fi s fi e
p incipal componen s o he geno ype da a when a ailable. Heigh (in me e s) was
adjus ed o sex, age and he fi s fi e p incipal componen s o he geno ype da a
when a ailable. The esul ing ai esiduals we e hen in e se no mal quan ile
ans o med.
As CNV bounda ies a y ac oss indi iduals, all associa ions we e pe o med a
he p obe le el. Fo his, CNV calls and quali y measu es we e ansla ed o p obe
le el. Fo each p obe in each coho , associa ion summa y s a is ics a e compu ed
and collec ed o me a-analysis. The summa y s a is ics o each p obe a e he
mean QS, he sum o squa ed QS s, he sum o he pheno ype–geno ype p oduc s,
he pheno ype means and he sum o he squa ed pheno ype alues. These
quan i ies a e su ficien o compu e eg ession coe ficien s as i we had access o
each indi idual coho da a, which is, o a e a ian associa ions, ad an ageous
compa ed o s anda d in e se- a iance me a-analysis.
Summa y s a is ics impu a ion. As we collec ed di e en SNP a ays wi h a i-
able p obe con en we impu ed summa y s a is ics o he Illumina 1M p obes as
e e ence p obe se . Fo each Illumina 1M p obe no p esen in he summa y
s a is ic p obe lis o a s udy we impu ed i s summa y s a is ics based on he
closes neighbo ing p obes on each side wi hin a 5 kb window. The impu a ion
weigh s a e se o be in e sely p opo ional o he dis ance be ween he a ge p obe
and he neighbo ing p obes.
CNV impu a ion quali y. Analogously o geno ype impu a ion, we used he
MACH ^
2measu e54 o es ima e he quali y o he CNVs es ima ion using he QS.
This measu e is he a io o he a iance o he Be noulli dis ibu ed p obabilis ic
CNV ( aking alue 1 wi h p obabili y |QS
i,j
|, 0 o he wise) a e aged o e he samples
o he empi ical a iance o an expec ed dosage ac oss all samples
^
2
j¼PiQSi;j
QS2
i;j
PiQSi;j
QS;j

2
whe e QS
i,j
ep esen s he QS o indi idual iand p obe j, and QS;jis he a e age
QS
i,j
o p obe jac oss he samples. Fo each coho , only p obes wi h ^
2
j0:1
we e kep o he me a-analysis. Impu a ion quali ies we e also me a-analyzed
using sample-size weigh ing.
NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-00556-x ARTICLE
NATURE COMMUNICATIONS |8: 744 |DOI: 10.1038/s41467-017-00556-x |www.na u e.com/na u ecommunica ions 7
Pipeline. Ou published pipeline16 based on bash, pe l and R has been imple-
men ed o un all p e-me a-analysis s eps. In b ie , his pipeline o ma s he gen-
o ype files o un CNV calls using PennCNV, i cleans and me ges he aw CNV
calls, i compu es a QS o each CNV and i finally calcula es he summa y s a is ics
a he p obe le el. All pa icipa ing coho s an he exac same pipeline and sha ed
summa y s a is ics wi h us. An example configu a ion file can be ound in he
Supplemen a y No e 1.
Me a-analysis. We an fixed e ec s me a-analysis as desc ibed by RAR-
EMETAL55. We di ec ly compu ed he me a β
Me a
and se
Me a
o a gi en CNV
p obe om he summa y s a is ics om he mul iple coho s:
βMe a ¼Pcpgc
Pcg2
cNcgc2
seMe a ¼ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
1
Pcg2
cNcgc2
s
whe e g2
cis he sum o he squa ed CNV dosage o all indi iduals in coho c,pg
c
is
he sum o he p oduc s o pheno ype × CNV dosage alues o all indi iduals in
coho c,gcis he a e age CNV dosage in coho c, and N
c
is he sample size o
coho c. An o e all Zsco e can be es ima ed as: Z
Me a
=β
Me a
/se
Me a
. E en ually
P alues a e compu ed o each p obe as: P
alue
=2∗ϕ(−|Z
Me a
|). All epo ed
esul s in he pape a e based on he ull s udy popula ion, unless s a ed o he wise
(e.g., condi ional analysis).
In o de o dec ease he numbe o es s and o a oid spu ious associa ions we
kep only p obes ha we e CNV in a leas ou coho s and had a equency o a
leas 0.01%. Beside, using a 1 Mb sliding window o e he en i e genome, we
me ged p obes wi h exac ly he same summa y s a is ics ( equency, e ec size, SE),
i.e., highly likely ha ing he same p ofile ac oss all indi iduals, wi hin ha window.
Numbe o e ec i e es s. Based on an in-house coho (HYPERGENES,
N=2930) we es ima ed he numbe o e ec i e es s o he p obes ha we e no
disca ded a he p e ious s ep (N
o al
=399,665). We compu ed he QS o his
coho and kep only he p obes o which he QS was no ze o (N
non-ze o
). Fo each
ch omosome we used a 1 Mb window o calcula e he numbe o e ec i e es s
locally. The numbe o e ec i e es s co esponds o he numbe o p obes
explaining 99.5% o he a iance in he window17. The esul s o each window and
ch omosome we e hen summed o ob ain an o e all N
e
numbe o e ec i e es s
o he non-ze o p obes. This indica ed he s eng h o dependence be ween CNV
p obes, =N
e
/N
non-ze o
. We ob ained a a io o =8.23 and applied his scaling
cons an o he 242,022 p obes es ed in ou me a-analysis s udy, yielding 29,407
independen es s and subsequen ly a 1.7 × 10−6genome-wide significan h eshold.
To ensu e obus ness, we epea ed he same analysis o each o he 33 ba ches o
he UK Biobank samples and ob ained a sligh ly less s ingen a io (median
=13.83, CI
95%
=[8.20, 20.17]), bu we p e e ed o use he mo e conse a i e
h eshold o 1.7 × 10−6.
CNV bu den analysis. Fo each sample we calcula ed he o al numbe o
(impu ed) CNV p obes showing de ia ion om he copy neu al s a e. To accoun
o unce ain y in he calls and o a oid a bi a y h esholding, we used he
absolu e QS o each p obe ( o he U-shaped model) and summed hem up o he
whole genome. Fo he o he models (dele ion-, duplica ion-only) we used he
espec i e modifica ions (minus dele ion QS, duplica ion QS). We hen an a linea
eg ession be ween he o al bu den sco e and he a ious ai s and me a-analyzed
he esul s om he 26 coho s.
Candida e CNV egions. In o de o alida e ou me hodology, we decided o fi s
look a egions al eady associa ed wi h BMI, weigh o heigh . Regions ha e been
defined based on p oximi y o GWAS SNP8,9, CNVs epo 39, genes om OMIM
eposi o y15 and om a e y ecen sys ema ic e iew o known genes implica ed in
gene ic synd omes wi h obesi y (Table 1o Kau e al.40). Rega ding he candida e
CNVs/genes o he OMIM egions, all high quali y ( 2>0.5) p obes alling in o he
egions we e selec ed. Fo each GWAS SNP we selec ed all he p obes wi h asso-
cia ion esul s in a ±500 kb egion a ound he SNP posi ion. The CNV epo
ca aloged 84 BMI and obesi y-associa ed CNVs om esea ch published since 2008
ia PubMed sea ch (see Supplemen a y Table 2o he publica ion by Pe e sen
e al.39). Ou o he 84 CNVs, we had good quali y p obes o 48 o hem ha we
subsequen ly es ed. Ou o he 79 OMIM egions o weigh and BMI, 37 had good
quali y p obes ( 2>0.5). The 96 Kau e al. genes40 ep esen 65 egions, ou o
which 57 a e on au osomes and 32 o hose con ained p obes wi hin 10 kb wi h
good impu a ion quali y ( 2>0.5). In e e y candida e egion we compu ed he
minimal P alue and mul iplied i wi h he numbe o e ec i e es s o ha egion
o ob ain one (co ec ed) P alue pe egion. We hen compu ed quan ile–quan ile
plo o isually inspec po en ial infla ion and compu ed he old-en ichmen o
egions wi h low P alues (P<0.05).
GWAS and eQTL lookup. In o de o u he in e p e ou findings, we checked
whe he heigh -associa ed coding a ian s10 we e loca ed wi hin ou heigh -
associa ed CNVs. Genes whose exp ession is modula ed by bo h ai -associa ed
SNPs and CNVs a e good gene candida es and can help na owing down he
c i ical egion. To iden i y such genes, we asked whe he known heigh /BMI-
associa ed SNPs8,9ac also as cis eQTLs in blood56 o he genes loca ed wi hin
heigh /BMI-associa ed CNVs.
Da a a ailabili y. All associa ion esul s a e a ailable in Supplemen a y Da a 1and
can also be b owsed a h ps://cn ca alogue.bbm i.nl
Recei ed: 28 Feb ua y 2017 Accep ed: 10 July 2017
Re e ences
1. Maes, H. H., Neale, M. C. & Ea es, L. J. Gene ic and en i onmen al ac o s in
ela i e body weigh and human adiposi y. Beha . Gene . 27, 325–351 (1997).
2. Vissche , P. M., B own, M. A., McCa hy, M. I. & Yang, J. Fi e yea s o GWAS
disco e y. Am. J. Hum. Gene . 90,7–24 (2012).
3. Zai len, N. e al. Using ex ended genealogy o es ima e componen s o
he i abili y o 23 quan i a i e and dicho omous ai s. PLoS Gene . 9, e1003520
(2013).
4. Hjelmbo g, J. e al. Gene ic influences on g ow h ai s o BMI: a longi udinal
s udy o adul wins. Obesi y 16, 847–852 (2008).
5. S unka d, A. J., Ha is, J. R., Pede sen, N. L. & McClea n, G. E. The body-mass
index o wins who ha e been ea ed apa . N. Engl. J. Med. 322, 1483–1487
(1990).
6. Vissche , P. M. e al. Assump ion- ee es ima ion o he i abili y om genome-
wide iden i y-by-descen sha ing be ween ull siblings. PLoS Gene . 2, e41
(2006).
7. Sil en oinen, K. e al. He i abili y o adul body heigh : a compa a i e s udy o
win coho s in eigh coun ies. Twin Res. 6, 399–408 (2003).
8. Locke, A. E. e al. Gene ic s udies o body mass index yield new insigh s o
obesi y biology. Na u e 518, 197–206 (2015).
9. Wood, A. R. e al. Defining he ole o common a ia ion in he genomic and
biological a chi ec u e o adul human heigh . Na . Gene . 46, 1173–1186
(2014).
10. Ma ouli, E. e al. Ra e and low- equency coding a ian s al e human adul
heigh . Na u e 542, 186–190 (2017).
11. Yang, J. e al. Gene ic a iance es ima ion wi h impu ed a ian s finds negligible
missing he i abili y o human heigh and body mass index. Na . Gene . 47,
1114–1120 (2015).
12. Gamazon, E. R., Cox, N. J. & Da is, L. K. S uc u al a chi ec u e o SNP e ec s
on complex ai s. Am. J. Hum. Gene . 95, 477–489 (2014).
13. Jacquemon , S. e al. Mi o ex eme BMI pheno ypes associa ed wi h gene
dosage a he ch omosome 16p11.2 locus. Na u e 478,97–9102 (2011).
14. Zu e ey, F. e al. A 600 kb dele ion synd ome a 16p11.2 leads o ene gy
imbalance and neu opsychia ic diso de s. J. Med. Gene . 49, 660–668 (2012).
15. Ma shall, C. R., e al. Con ibu ion o copy numbe a ian s o schizoph enia
om a genome-wide s udy o 41,321 subjec s. Na . Gene .49,27–35 (2017).
16. Mace, A. e al. New quali y measu e o SNP a ay based CNV de ec ion.
Bioin o ma ics 32, 3298–3305 (2016).
17. Gao, X., S a me , J. & Ma in, E. R. A mul iple es ing co ec ion me hod o
gene ic associa ion s udies using co ela ed single nucleo ide polymo phisms.
Gene . Epidemiol. 32, 361–369 (2008).
18. Daube , A. e al. Genome-wide associa ion o copy-numbe a ia ion e eals an
associa ion be ween sho s a u e and he p esence o low- equency genomic
dele ions. Am. J. Hum. Gene . 89, 751–759 (2011).
19. Wheele , E. e al. Genome-wide SNP and CNV analysis iden ifies common and
low- equency a ian s associa ed wi h se e e ea ly-onse obesi y. Na . Gene .
45, 513–517 (2013).
20. Joshi, P. K. e al. Di ec ional dominance on s a u e and cogni ion in di e se
human popula ions. Na u e 523, 459–462 (2015).
21. Tu ne , L., G ego y, A., Twells, L., G ego y, D. & S a opoulos, D. J. Dele ion o
he MC4R gene in a 9-yea -old obese boy. Child Obes. 11, 219–223 (2015).
22. Bachmann-Gagescu, R. e al. Recu en 200-kb dele ions o 16p11.2 ha
include he SH2B1 gene a e associa ed wi h de elopmen al delay and obesi y.
Gene . Med. 12, 641–647 (2010).
23. B une i-Pie i, N. e al. Recu en ecip ocal 1q21.1 dele ions and duplica ions
associa ed wi h mic ocephaly o mac ocephaly and de elopmen al and
beha io al abno mali ies. Na . Gene . 40, 1466–1471 (2008).
24. Zahnlei e , D. e al. Ra e copy numbe a ian s a e a common cause o sho
s a u e. PLoS Gene . 9, e1003365 (2013).
25. Rashidi-Nezhad, A., Talebi, S., Saebnou i, H., Ak ami, S. M. & Reymond, A.
The e ec o homozygous dele ion o he BBOX1 and Fibin genes on ca ni ine
le el and acyl ca ni ine p ofile. BMC Med. Gene . 15, 75 (2014).
ARTICLE NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-00556-x
8NATURE COMMUNICATIONS |8: 744 |DOI: 10.1038/s41467-017-00556-x |www.na u e.com/na u ecommunica ions
26. Loos, R. J. e al. Common a ian s nea MC4R a e associa ed wi h a mass,
weigh and isk o obesi y. Na . Gene . 40, 768–775 (2008).
27. Loos, R. J. The gene ic epidemiology o melanoco in 4 ecep o a ian s. Eu . J.
Pha macol. 660, 156–164 (2011).
28. Spelio es, E. K. e al. Associa ion analyses o 249,796 indi iduals e eal 18 new
loci associa ed wi h body mass index. Na . Gene . 42, 937–948 (2010).
29. Makeye , A. V. e al. GTF2IRD2 is loca ed in he Williams-Beu en synd ome
c i ical egion 7q11.23 and encodes a p o ein wi h wo TFII-I-like helix-loop-
helix epea s. P oc. Na l Acad. Sci. USA 101, 11052–11057 (2004).
30. Me is, C. B., Mo is, C. A., Klein-Tasman, B. P., Velleman, S. L. & Osbo ne, L.
R. 7q11.23 Duplica ion Synd ome. In GeneRe iews (eds Pagon, R. A. e al.)
(Uni . o Washing on P ess, 1993–2017).
31. F ancke, U. Williams-Beu en synd ome: genes and mechanisms. Hum. Mol.
Gene . 8, 1947–1954 (1999).
32. Mo is, C. A. Williams Synd ome. In GeneRe iews (eds Pagon, R. A. e al.)
(Uni . o Washing on P ess, 1993-2017).
33. Van de Aa, N. e al. Fou een new cases con ibu e o he cha ac e iza ion o
he 7q11.23 mic oduplica ion synd ome. Eu . J. Med. Gene . 52,94–100 (2009).
34. Velleman, S. L. & Me is, C. B. Child en wi h 7q11.23 duplica ion synd ome:
speech, language, cogni i e, and beha io al cha ac e is ics and hei
implica ions o in e en ion. Pe spec . Lang. Lea n. Educ. 18, 108–116 (2011).
35. S omme, P., Bjo ns ad, P. G. & Rams ad, K. P e alence es ima ion o Williams
synd ome. J. Child Neu ol. 17, 269–271 (2002).
36. Glass o d, M. R., Rosen eld, J. A., F eedman, A. A., Zwick, M. E. & Mulle, J. G.
No el ea u es o 3q29 dele ion synd ome: esul s om he 3q29 egis y. Am. J.
Med. Gene . A 170A, 999–1006 (2016).
37. Willa , L. e al. 3q29 mic odele ion synd ome: clinical and molecula
cha ac e iza ion o a new synd ome. Am. J. Hum. Gene . 77, 154–160 (2005).
38. Lisi, E. C. e al. 3q29 in e s i ial mic oduplica ion: a new synd ome in a h ee-
gene a ion amily. Am. J. Med. Gene . A 146A, 601–609 (2008).
39. Pe e son, R. E. e al. On he associa ion o common and a e gene ic a ia ion
influencing body mass index: a combined SNP and CNV analysis. BMC
Genomics 15, 368 (2014).
40. Kau , Y., de Souza, R. J., Gibson, W. T. & Mey e, D. A sys ema ic e iew o
gene ic synd omes wi h obesi y Obes. Re .18 603–634 (2017).
41. Coe, B. P. e al. Refining analyses o copy numbe a ia ion iden ifies specific
genes associa ed wi h de elopmen al delay. Na . Gene . 46, 1063–1071
(2014).
42. Wal e s, R. G. e al. Ra e genomic s uc u al a ian s in complex disease:
lessons om he eplica ion o associa ions wi h obesi y. PLoS ONE 8, e58048
(2013).
43. Day, F. R. e al. Genomic analyses iden i y hund eds o a ian s associa ed wi h
age a mena che and suppo a ole o pube y iming in cance isk. Na .
Gene .49, 834–841 (2017).
44. Bulik-Sulli an, B. K. e al. LD Sco e eg ession dis inguishes con ounding om
polygenici y in genome-wide associa ion s udies. Na . Gene . 47, 291–295
(2015).
45. W ay, N. R., Pu cell, S. M. & Vissche , P. M. Syn he ic associa ions c ea ed by
a e a ian s do no explain mos GWAS esul s. PLoS Biol. 9, e1000579
(2011).
46. Swe z, M. A. e al. The MOLGENIS oolki : apid p o o yping o bioso wa e a
he push o a bu on. BMC Bioin o ma ics 11, S12 (2010).
47. Swe z, M. A. & Jansen, R. C. Beyond s anda diza ion: dynamic so wa e
in as uc u es o sys ems biology. Na . Re . Gene . 8, 235–243 (2007).
48. Lek, M. e al. Analysis o p o ein-coding gene ic a ia ion in 60,706 humans.
Na u e 536, 285–291 (2016).
49. Dolce i, A. e al. 1q21.1 Mic oduplica ion exp ession in adul s. Gene . Med. 15,
282–289 (2013).
50. Habel, A., McGinn, M. J. 2nd, Zackai, E. H., Unanue, N. & McDonald-McGinn,
D. M. Synd ome-specific g ow h cha s o 22q11.2 dele ion synd ome in
Caucasian child en. Am. J. Med. Gene . A 158A, 2665–2671 (2012).
51. Coope , A. J., Lamb, M. J., Sha p, S. J., Simmons, R. K. & G i fin, S. J. Bidi ec ional
associa ion be ween physical ac i i y and muscula s eng h in olde adul s: esul s
om he UK Biobank s udy. In . J. Epidemiol.46,141–148 (2016).
52. Mannik, K. e al. Copy numbe a ia ions and cogni i e pheno ypes in
unselec ed popula ions. JAMA 313, 2044–2054 (2015).
53. Wang, K. e al. PennCNV: an in eg a ed hidden Ma ko model designed o
high- esolu ion copy numbe a ia ion de ec ion in whole-genome SNP
geno yping da a. Genome Res. 17, 1665–1674 (2007).
54. Li, Y., Wille , C. J., Ding, J., Schee , P. & Abecasis, G. R. MaCH: using sequence
and geno ype da a o es ima e haplo ypes and unobse ed geno ypes. Gene .
Epidemiol. 34, 816–834 (2010).
55. Feng, S., Liu, D., Zhan, X., Wing, M. K. & Abecasis, G. R. RAREMETAL: as
and powe ul me a-analysis o a e a ian s. Bioin o ma ics 30, 2828–2829
(2014).
56. Wes a, H. J. e al. Sys ema ic iden ifica ion o ans eQTLs as pu a i e
d i e s o known disease associa ions. Na . Gene . 45, 1238–1243 (2013).
Acknowledgemen s
This esea ch has been conduc ed using he UK Biobank Resou ce. This esea ch has
been conduc ed using he Danish Na ional Biobank esou ce. The au ho s a e g a e ul o
he Raine S udy pa icipan s and hei amilies, and o he Raine S udy esea ch s a o
coho co-o dina ion and da a collec ion. QIMR is g a e ul o he wins and hei amilies
o hei gene ous pa icipa ion in hese s udies. We would like o hank s a a he
Queensland Ins i u e o Medical Resea ch: Anjali Hende s, Dixie S a ham, Lisa Bowdle ,
Ann Eld idge, and Ma lene G ace o sample collec ion, p ocessing and geno yping, Sco
Go don, B ian McE oy, Belinda Co nes and Beben Benyamin o da a QC and p e-
pa a ion, and Da id Smy h and Ha y Beeby o IT suppo . HBCS Acknowledgemen s:
We hank all s udy pa icipan s as well as e e ybody in ol ed in he Helsinki Bi h
Coho S udy. Helsinki Bi h Coho S udy has been suppo ed by g an s om he
Academy o Finland, he Finnish Diabe es Resea ch Socie y, Folkhälsan Resea ch
Founda ion, No o No disk Founda ion, Finska Läka esällskape , Juho Vainio Founda-
ion, Signe and Ane Gyllenbe g Founda ion, Uni e si y o Helsinki, Minis y o Edu-
ca ion, Ahokas Founda ion, Emil Aal onen Founda ion. Fin isk s udy is g a e ul o he
THL DNA labo a o y o i s skill ul wo k o p oduce he DNA samples used in his s udy
and hanks he Sange Ins i u e and FIMM geno yping acili ies o geno yping he
samples. We hank he MOLGENIS eam and Genomics Coo dina ion Cen e o he
Uni e si y Medical Cen e G oningen o so wa e de elopmen and da a managemen ,
in pa icula Ma ieke Bijlsma and Edi h Ad iaanse. This wo k was suppo ed by he
Leena ds Founda ion ( o Z.K.), he Swiss Na ional Science Founda ion (31003A_169929
o Z.K., Sine gia g an CRSII33-133044 o AR), Simons Founda ion (SFARI274424 o
AR) and Sys emsX.ch (51RTP0_151019 o Z.K.). A.R.W., H.Y. and T.M.F. a e suppo ed
by he Eu opean Resea ch Council g an : 323195:SZ-245. M.A.T., M.N.W. and An.M. a e
suppo ed by he Wellcome T us Ins i u ional S a egic Suppo Awa d
(WT097835MF). Fo ull unding in o ma ion o all pa icipa ing coho s see Supple-
men a y No e 2.
Au ho con ibu ions
Z.K. and T.M.F. designed he s udy, Au.M. and Z.K. designed he analysis pipeline, Au.
M. analyzed, QC’ed he da a and pe o med he me a-analysis, Au.M. and Z.K. con-
duc ed all ollow-up analyses. Au.M., T.M.F., and Z.K. w o e he pape . A.B.N., A.L., An .
M., A.R.W., A.S.H., C.H., E.P., E.P.B., E.V., F.G., F.R., G.W.M., H.Y., I.M.H., I.M.N., J.C.
C., J.K., K.S.R., M.A.T., M.F.F., M.K.W., M.M., M.N., M.P., N.S., Pan.D., Pa .D., P.L., R.
M.F., S.A., S.E.J., T.M., T.T.P., T.W.W., U.S., V.S., and W.Z. con ibu ed o indi idual
s udy design and managemen ; A.B.N., A.L., An.M., Au.M., A.R.W., A.S.H., C.P., E.P.B.,
E.S., E.V., G.W.M., H.B., H.Y., I.M.H., J.C.C., J.K., J.S., J.S.K., J.T., M.A.P., M.F.F., M.L.,
M.N., M.P., N.G.M., N.S., Pan.D., Pa .D., P.L., R.M., R.M.F., S.A., S.E.J., S.M., T.M., T.M.
F., T.T.P., T.W.W., U.S., W.Z. played a ole in da a collec ion; A.B.N., Au.M., A.P., C.H.,
C.P., F.G., G.W.M., I.M.H., J.C.C., M.A.P., M.M., M.N., M.N.W., M.P., N.S., Pan.D., P.L.,
R.M., R.M.F., S.E.J., W.Z. con ibu ed o he geno yping; A.B.N., Au.M., A.R.W., A.S.H.,
E.S., F.G., G.W.M., I.M.H., J.K., J.T., M.F.F., M.L., M.N., M.N.W., M.P., N.G.M., Pa .D.,
P.K., R.N.B., S.E.J., V.S., W.Z. pa icipa ed in he pheno ype p epa a ion; A.B.N., Au.M.,
B.F., E.S., L.F., M.L., M.N., M.N.W., N.R.W., Pa .D., P.J.V., R.M.F., R.N.B., Y.S. pe o med
s udy specific analysis; A.J.O., Au.M., C.V., D.C., D.P., D.P.S., D.R.N., G.L., H.S., K.B., K.
C., K.C., K.K., K.M., M.A.S., M.A.T., M.F.F., M.K.W., M.M., M.N., M.N.W., N.S., R.J.F.L.,
S.E.M., T.D.S., T.T.P., W.A., X.L., Z.K. con ibu ed o/supe ised s udy specific analysis.
Addi ional in o ma ion
Supplemen a y In o ma ion accompanies his pape a doi:10.1038/s41467-017-00556-x.
Compe ing in e es s: The au ho s decla e no compe ing financial in e es s.
Rep in s and pe mission in o ma ion is a ailable online a h p://npg.na u e.com/
ep in sandpe missions/
Publishe 's no e: Sp inge Na u e emains neu al wi h ega d o ju isdic ional claims in
published maps and ins i u ional a filia ions.
Open Access This a icle is licensed unde a C ea i e Commons
A ibu ion 4.0 In e na ional License, which pe mi s use, sha ing,
adap a ion, dis ibu ion and ep oduc ion in any medium o o ma , as long as you gi e
app op ia e c edi o he o iginal au ho (s) and he sou ce, p o ide a link o he C ea i e
Commons license, and indica e i changes we e made. The images o o he hi d pa y
ma e ial in his a icle a e included in he a icle’s C ea i e Commons license, unless
indica ed o he wise in a c edi line o he ma e ial. I ma e ial is no included in he
a icle’s C ea i e Commons license and you in ended use is no pe mi ed by s a u o y
egula ion o exceeds he pe mi ed use, you will need o ob ain pe mission di ec ly om
he copy igh holde . To iew a copy o his license, isi h p://c ea i ecommons.o g/
licenses/by/4.0/.
© The Au ho (s) 2017
NATURE COMMUNICATIONS | DOI: 10.1038/s41467-017-00556-x ARTICLE
NATURE COMMUNICATIONS |8: 744 |DOI: 10.1038/s41467-017-00556-x |www.na u e.com/na u ecommunica ions 9