Gene ics and popula ion analysis
me aCCA: summa y s a is ics-based
mul i a ia e me a-analysis o genome-wide
associa ion s udies using canonical
co ela ion analysis
Anna Cichonska
1,2,
*, Juho Rousu
2
, Pekka Ma inen
2
, An i J. Kangas
3
,
Pasi Soininen
3,4
, Te ho Leh ima¨ki
5
, Olli T. Rai aka i
6,7
,
Ma jo-Rii a Ja¨ elin
8,9,10,11
, Veikko Salomaa
12
, Mika Ala-Ko pela
3,4,13
,
Samuli Ripa i
1,14,15
and Ma i Pi inen
1,
*
1
Ins i u e o Molecula Medicine Finland FIMM, Uni e si y o Helsinki, Helsinki, Finland,
2
Helsinki Ins i u e o
In o ma ion Technology HIIT, Depa men o Compu e Science, Aal o Uni e si y, Espoo, Finland,
3
Compu a ional
Medicine, Uni e si y o Oulu, Oulu Uni e si y Hospi al and Biocen e Oulu, Oulu, Finland,
4
NMR Me abolomics
Labo a o y, School o Pha macy, Uni e si y o Eas e n Finland, Kuopio, Finland,
5
Depa men o Clinical Chemis y,
Fimlab Labo a o ies, Uni e si y o Tampe e School o Medicine, Tampe e, Finland,
6
Depa men o Clinical
Physiology and Nuclea Medicine, Uni e si y o Tu ku and Tu ku Uni e si y Hospi al, Tu ku, Finland,
7
Resea ch
Cen e o Applied and P e en i e Ca dio ascula Medicine, Uni e si y o Tu ku and Depa men o Clinical
Physiology and Nuclea Medicine, Tu ku Uni e si y Hospi al, Tu ku, Finland,
8
Depa men o Epidemiology and
Bios a is ics, MRC-PHE Cen e o En i onmen & Heal h, School o Public Heal h, Impe ial College London,
London, UK,
9
Cen e o Li e Cou se Epidemiology, Facul y o Medicine, Uni e si y o Oulu, Oulu, Finland,
10
Biocen e Oulu, Uni e si y o Oulu, Oulu, Finland,
11
Uni o P ima y Ca e, Oulu Uni e si y Hospi al, Oulu, Finland,
12
Na ional Ins i u e o Heal h and Wel a e, Helsinki, Finland,
13
Compu a ional Medicine, School o Social and
Communi y Medicine and he Medical Resea ch Council In eg a i e Epidemiology Uni , Uni e si y o B is ol,
B is ol, UK,
14
Public Heal h, Uni e si y o Helsinki, Helsinki, Finland and
15
Wellcome T us Sange Ins i u e,
Wellcome T us Genome Campus, Hinx on, UK
*To whom co espondence should be add essed.
Associa e Edi o : Jane Kelso
Recei ed on 16 July 2015; e ised on 4 Decembe 2015; accep ed on 19 Janua y 2016
Abs ac
Mo i a ion: A dominan app oach o gene ic associa ion s udies is o pe o m uni a ia e es s be-
ween geno ype-pheno ype pai s. Howe e , analyzing ela ed ai s oge he inc eases s a is ical
powe , and ce ain complex associa ions become de ec able only when se e al a ian s a e es ed
join ly. Cu en ly, modes sample sizes o indi idual coho s, and es ic ed a ailabili y o indi id-
ual-le el geno ype-pheno ype da a ac oss he coho s limi conduc ing mul i a ia e es s.
Resul s: We in oduce me aCCA, a compu a ional amewo k o summa y s a is ics-based analysis
o a single o mul iple s udies ha allows mul i a ia e ep esen a ion o bo h geno ype and pheno ype.
I ex ends he s a is ical echnique o canonical co ela ion analysis o he se ing whe e o iginal indi id-
ual-le el eco ds a e no a ailable, and employs a co a iance sh inkage algo i hm o achie e
obus ness.
Mul i a ia e me a-analysis o wo Finnish s udies o nuclea magne ic esonance me abolomics by
me aCCA, using s anda d uni a ia e ou pu om he p og am SNPTEST, shows an excellen
V
CThe Au ho 2016. Published by Ox o d Uni e si y P ess. 1981
This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion Non-Comme cial License (h p://c ea i ecommons.o g/licenses/by-nc/4.0/),
which pe mi s non-comme cial e-use, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly ci ed. Fo comme cial e-use, please con ac
jou nals.pe [email protected]
Bioin o ma ics, 32(13), 2016, 1981–1989
doi: 10.1093/bioin o ma ics/b w052
Ad ance Access Publica ion Da e: 19 Feb ua y 2016
O iginal Pape
ag eemen wi h he pooled indi idual-le el analysis o o iginal da a. Mo i a ed by s ong mul i a i-
a e signals in he lipid genes es ed, we en ision ha mul i a ia e associa ion es ing using
me aCCA has a g ea po en ial o p o ide no el insigh s om al eady published summa y s a is ics
om high- h oughpu pheno yping echnologies.
A ailabili y and implemen a ion: Code is a ailable a h ps://gi hub.com/aal o-ics-kepaco
Con ac s: [email p o ec ed] o [email p o ec ed]
Supplemen a y in o ma ion: Supplemen a y da a a e a ailable a Bioin o ma ics online.
1 In oduc ion
Mos human diseases and ai s ha e a s ong gene ic componen .
Genome-wide associa ion s udies (GWAS) ha e p o en e ec i e in
iden i ying gene ic a ia ion con ibu ing o common complex dis-
o de s, including ype 2 diabe es (Mahajan e al., 2014), ca dio as-
cula disease (Deloukas e al., 2013), schizoph enia (Schizoph enia
Wo king G oup o he Psychia ic Genomics Conso ium, 2014),
and quan i a i e ai s, such as lipid le els (Global Lipids Gene ics
Conso ium, 2013;Su akka e al., 2015) and me abolomics
(Ke unen e al., 2012;Shin e al., 2014).
A dominan app oach o GWAS is o es one single-nucleo ide
polymo phism (SNP) a a ime agains one quan i a i e pheno ype
measu e o a bina y disease indica o . This uni a ia e app oach is un-
likely o be op imal when millions o SNPs and a g owing numbe o
pheno ypes, including se um me abolomic p o iles (Ke unen e al.,
2012;Shin e al.,2014), h ee-dimensional images (Wang e al.,
2013), and gene exp ession da a (A dlie e al.,2015)becomea ailable
simul aneously. Indeed, a ecen compa ison demons a ed ha u iliz-
ing mul i a ia e pheno ype ep esen a ion inc eases s a is ical powe ,
and leads o iche indings in he associa ion es s compa ed o he
uni a ia e analysis (Inouye e al., 2012). Mo eo e , some complex
geno ype-pheno ype co ela ions can be de ec ed only when es ing
se e al gene ic a ian s simul aneously (Ma inen e al.,2014), and
mul i-geno ype es s a e common p ac ice in a e a ian associa ion
s udies, whe e s a is ical powe o de ec any single a ian is e y
small (Feng e al.,2014;Lee e al.,2014).
Un o una ely, es ic ed a ailabili y o comple e mul i a ia e indi-
idual-le el eco ds ac oss he coho s cu en ly limi s mul i a ia e
analyses. O en, only he uni a ia e GWAS summa y s a is ics, i.e. uni-
a ia e eg ession coe icien s wi h hei s anda d e o s, om indi id-
ual coho s a e publicly a ailable. Hence, a majo ques ion is how we
can use hese uni a ia e associa ion esul s o ca y ou a mul i a ia e
me a-analysis o GWAS (E angelou and Ioannidis, 2013), which is c u-
cial o inc ease he powe o iden i y no el gene ic associa ions.
Recen ly, wo kinds o mul i a ia e es ing app oaches ope a ing
on uni a ia e summa y s a is ics ha e been in oduced: (i) one SNP
agains mul iple ai s (S ephens, 2013; an de Sluis e al., 2013;
Vucko ic e al., 2015;Zhu e al., 2015) and (ii) mul iple SNPs
agains one ai (Feng e al., 2014;Yang e al., 2012). We p opose a
new amewo k, me aCCA, ha uni ies bo h o he exis ing
app oaches by allowing canonical co ela ion analysis (CCA) o
mul iple SNPs agains mul iple ai s based on uni a ia e summa y
s a is ics and publicly a ailable da abases.
CCA is a well-es ablished s a is ical echnique o iden i ying lin-
ea ela ionships be ween wo se s o a iables, and has been suc-
cess ully applied o GWAS (Fe ei a and Pu cell, 2009;Inouye e al.,
2012;Ma inen e al., 2013;Tang and Fe ei a, 2012). Ou
me aCCA me hod ex ends CCA o he se ing whe e o iginal indi-
idual-le el measu emen s a e no a ailable. Ins ead, me aCCA
wo ks wi h h ee pieces o he ull da a co a iance ma ix, and
applies a co a iance sh inkage algo i hm o achie e obus ness.
We demons a e he pe o mance o me aCCA using SNP and me-
aboli e da a om h ee Finnish coho s. In summa y, his pape
makes he ollowing con ibu ions.
•To ou knowledge, we p o ide he i s compu a ional ame-
wo k o associa ion es ing be ween mul i a ia e geno ype and
mul i a ia e pheno ype, based on uni a ia e summa y s a is ics
om single o mul iple GWAS. Ou implemen a ion is eely
a ailable.
•We demons a e how o accu a ely es ima e co ela ion s uc-
u es o pheno ypic and geno ypic a iables wi hou an access o
he indi idual-le el da a.
•We a oid alse posi i e associa ions by a co a iance sh inkage al-
go i hm based on s abiliza ion o he leading canonical
co ela ion.
•Ou app oach, me aCCA, is a gene al amewo k o conduc
CCA when ull da a a e no a ailable, and he e o e i is widely
applicable also ou side GWAS.
A de ailed discussion on he ela ionship be ween me aCCA and
p e iously published mul i a ia e associa ion me hods can be ound
in Supplemen a y Da a.
2 Me hods
This sec ion is o ganized as ollows. Fi s , Sec ion 2.1 explains uni-
a ia e GWAS, he esul s o which, in he o m o c oss-co a iance
ma ix, cons i u e an inpu o me aCCA desc ibed in Sec ion 2.2;
Sec ion 2.3 demons a es how a me a-analysis o se e al s udies is
conduc ed in ou amewo k; Sec ion 2.4 ou lines a p ocedu e o
choosing SNPs ep esen a i e o a gi en locus; inally, Sec ion 2.5
in oduces he da a we used o es me aCCA in he me a-analy ic
se ing.
2.1 Uni a ia e GWAS
Le Xand Ydeno e geno ype and pheno ype ma ices o dimensions
NGand NP, espec i ely, s o ing he indi idual-le el da a; N
he numbe o samples; Gand P he numbe o geno ypic and
pheno ypic a iables, espec i ely. The columns o Xand Ya e
s anda dized o ha e mean 0 and s anda d de ia ion 1.
Typically, uni a ia e GWAS analysis o quan i a i e ai s es s
o an associa ion be ween each pai o geno ype xg2RNand
pheno ype yp2RNsepa a ely using a linea model:
yp¼agp þxgbgp þe:(1)
Coe icien b
gp
, co esponding o he slope o he eg ession line,
is he pa ame e o in e es , since i depic s he size o he e ec
o he gene ic a ian xgon he ai yp.Pa ame e a
gp
is an in e cep
on he y-axis, and eindica es a Gaussian e o e m o noise. The
model is i by he me hod o leas squa es ha leads o a closed- o m
1982 A.Cichonska e al.
es ima e o he unknown pa ame e bgp ¼½xT
gyp½xT
gxg1¼
½ðN1Þsxy½ðN1Þsxx1¼sxy,whe es
xy
is a sample co a iance o
xgand yp,ands
xx
¼1 is a sample a iance o xg. Hence, he c oss-co-
a iance ma ix R
XY
be ween all geno ypic and pheno ypic a iables is
made o uni a ia e eg ession coe icien s b
gp
:
RXY ¼XTY
N1¼
b11 b12 b1P
b21 b22 b2P
.
.
..
.
...
..
.
.
bG1bG2 bGP
0
B
B
B
B
B
B
@
1
C
C
C
C
C
C
A
:(2)
An impo an no e is ha i he indi idual-le el da ase s Xand Y
we e no s anda dized be o e applying he linea eg ession, he
s anda diza ion can be achie ed a e wa ds by a ans o ma ion
bSTANDR
gp ¼1
ffiffiffiffiffi
N
pSEgp bgp;(3)
whe e SE
gp
indica es he s anda d e o o b
gp
, as gi en by GWAS
so wa e. (Typically, SEgp p=ðffiffiffiffiffi
N
pffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
2 gð1 gÞ
pÞ, whe e
p
is he
s anda d de ia ion o he ai p, and
g
is he mino allele equency
o SNP g, bu unce ain y in geno ype impu a ion causes de ia ions
om his exp ession.)
2.2 me aCCA
Conduc ing mul i a ia e associa ion es s equi es es ima es o he
dependencies be ween geno ypic and pheno ypic a iables, deno ed
R
XX
and R
YY,
espec i ely. Typically, hey a e calcula ed based on
he indi idual-le el measu emen s Xand Y:
RXX ¼XTX
N1;(4)
RYY ¼YTY
N1:(5)
me aCCA ope a es on he c oss-co a iance ma ix R
XY
(Equa ion
2), and co ela ion s uc u es ^
RXX;^
RYY, es ima ed wi hou an access
o he indi idual-le el da a Xand Y(Fig. 1A, B). To make he esul -
ing ull co a iance ma ix Ra alid co a iance ma ix, me aCCA
applies a sh inkage algo i hm (Fig. 1C).
The es o his sec ion desc ibes he de ails o me aCCA
amewo k.
2.2.1 Es ima ion o geno ypic co ela ion s uc u e
Gene ic a ia ion is o ganized in haplo ype blocks, whose s uc u e
is de e mined by mu a ion and ecombina ion e en s, oge he wi h
demog aphic e ec s, including popula ion g ow h, admix u e and
bo lenecks (Wall and P i cha d, 2003). Hence, co ela ion s uc u e
o gene ic a ian s di e s be ween popula ions, such as, e.g. he
Finns, Icelande s o Cen al Eu opeans. In me aCCA,^
RXX is calcu-
la ed using a e e ence da abase ep esen ing he s udy popula ion,
such as he 1000 Genomes da abase (1000 Genomes P ojec
Conso ium, 2012, www.1000genomes.o g), o o he geno ypic
da a a ailable on he a ge popula ion. In he Sec ion 3, we demon-
s a e ha es ima ing ^
RXX om he a ge popula ion (in ou case,
he Finns) leads o be e esul s han u ilizing he da a comp ising
indi iduals ac oss dis inc popula ions (e.g. he Finns and o he
Eu opeans). Howe e , since e e ence da a on he a ge popula ion
may no always be a hand, we also p esen a obus bu less powe -
ul solu ion o mul i a ia e associa ion es ing by simply using
geno ypes o all indi iduals om a ce ain b oade geog aphical e-
gion (e.g. a con inen ) a ailable unde he 1000 Genomes P ojec .
2.2.2 Es ima ion o pheno ypic co ela ion s uc u e
In ou amewo k, pheno ypic co ela ion s uc u e ^
RYY is compu ed
based on R
XY.
Each en y o ^
RYY co esponds o a Pea son co el-
a ion be ween wo column ec o s o R
XY
- uni a ia e eg ession co-
e icien s o wo pheno ypic a iables sand ac oss Ggene ic
a ian s:
^
RYYðs; Þ¼ X
G
g¼1ðbgs lsÞðbg l Þ
ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
X
G
g¼1ðbgs lsÞ2
u
u
ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
X
G
g¼1ðbg l Þ2
u
u
;(6)
whe e l
s
and l
a e he mean alues ls¼1
GPG
g¼1bgs and
l ¼1
GPG
g¼1bg . (The de ailed jus i ica ion is p o ided in
Supplemen a y Da a.) In Supplemen a y Table S2, we demons a e
ha he highe he numbe o geno ypic a iables G, he lowe he
e o o he es ima e. Thus, ^
RYY should be calcula ed om summa y
Fig. 1. Schema ic pic u e showing an o e iew o me aCCA amewo k o
summa y s a is ics-based mul i a ia e associa ion es ing using canonical
co ela ion analysis. (A)me aCCA ope a es on h ee pieces o he ull co a i-
ance ma ix R:R
XY
o uni a ia e geno ype-pheno ype associa ion esul s, R
XX
o geno ype-geno ype co ela ions, and R
YY
o pheno ype-pheno ype co el-
a ions. (B)^
RXX is es ima ed om a e e ence da abase ma ching he s udy
popula ion, e.g. he 1000 Genomes, and pheno ypic co ela ion s uc u e ^
RYY
is es ima ed om R
XY.
(C) A co a iance sh inkage algo i hm is applied o add
obus ness o he me hod. Numbe s in b acke s e e o subsec ions in
Me hods. Me a-analysis o se e al s udies is pe o med by pooling co a i-
ance ma ices o he same ype, be o e s ep (C), as desc ibed in Sec ion 2.3.
The da a educ ion achie ed by me aCCA can be seen in Supplemen a y
Figu e S1
me aCCA 1983
s a is ics o all a ailable gene ic a ian s, e en i only a subse o
hem is aken o he u he analysis.
2.2.3 Canonical co ela ion analysis
CCA (Ho elling, 1936) is a mul i a ia e echnique o de ec ing lin-
ea ela ionships be ween wo g oups o a iables X2RNGand
Y2RNP, whe e Xand Ycons i u e wo di e en iews o he same
objec . The objec i e is o ind maximally co ela ed linea combin-
a ions o columns o each ma ix. This co esponds o inding ec-
o s a2RGand b2RP ha maximize
¼ðXaÞTðYbÞ
kXakkYbk¼aTRXY b
ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
aTRXXa
pffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
bTRYYb
q:(7)
The maximized co ela ion is called canonical co ela ion be-
ween Xand Y. We p o ide he echnical de ails o he me hod, as
well as i s ex ension o subsequen canonical co ela ions and hei
signi icance es ing in Supplemen a y Da a.
2.2.4 Sh inkage
A his poin , we ha e h ee co a iance ma ices, namely R
XY,
^
RXX,and
^
RYY . Howe e , in many cases, he esul ing ull co a iance ma ix
R¼
^
RXX RXY
RT
XY ^
RYY
!
is no posi i e semide ini e (PSD), and he e o e i s building blocks
canno be jus plugged in o he CCA amewo k (Equa ion 7). To
o e come his p oblem, in me aCCA, we apply sh inkage o ind a
nea es alid R(Ledoi and Wol , 2003). We use an i e a i e p oced-
u e whe e he magni udes o he o -diagonal en ies a e being
sh unk owa ds ze o un il Rbecomes PSD (Algo i hm 1).
Assu ing he PSD p ope y o he ull co a iance ma ix is neces-
sa y, al hough, as we demons a e in he Sec ion 3, no su icien o
ob ain eliable esul s o he associa ion analysis when he es ima e
^
RXX (and/o ^
RYY) is noisy. In o de o add ess his issue, we p opose a
a ian o me aCCA, called me aCCA1, whe e he ull co a iance
ma ix Ris sh unk beyond he le el gua an eeing i s PSD p ope y. A
challenge, howe e , is o ind an op imal sh inkage in ensi y.
Sh inkage applied wi hou any s opping c i e ion would lead o g ad-
ual emo al o all dependencies be ween geno ypic and pheno ypic
a iables. Ledoi and Wol (2003) in oduced an analy ic app oach o
de e mining he op imal sh inkage le el bu i equi es he indi idual-
le el da ase s Xand Y.Inme aCCAþ, we moni o he leading canon-
ical co ela ion alue , and we con inue he sh inkage o he ull co-
a iance ma ix Run il s abilizes. Speci ically, we ack he pe cen
change pc o be ween subsequen sh inkage i e a ions, and we de e -
mine an app op ia e amoun o sh inkage using an elbow heu is ic,
simila o he c i e ion o inding he numbe o clus e s, equen ly
used in he li e a u e (Tibshi ani e al., 2001). The idea is ha he slope
o he g aph should be s eep o he le o he elbow, bu s able o he
igh o i . We ind he elbow, and hus he app op ia e numbe o
sh inkage i e a ions, by aking he poin closes o he o igin o he
plo o pc e sus i e a ion numbe , as schema ically shown in
Supplemen a y Figu e S2.
Building blocks ^
RXY ;^
RXX;^
RYY o he esul ing ull co a iance ma-
ix R, sh unk un il i became PSD o beyond, a e hen plugged in o he
CCA amewo k o ge he inal geno ype-pheno ype associa ion esul .
In p ac ice, in o de o p o ec om alse posi i e signals, he sh inkage
mode o me aCCAþshould be applied whene e ^
RYY is es ima ed
om summa y s a is icso asmallnumbe o gene ic a ian s,and/o
^
RXX is calcula ed using a gene ic e e ence popula ion.
2.2.5 Types o he mul i a ia e associa ion analysis
We conside he ollowing wo ypes o he mul i a ia e analysis.
1. Uni a ia e geno ype – mul i a ia e pheno ype
One gene ic a ian es ed o an associa ion wi h a se o pheno-
ypic a iables (ma ix ^
RXX no needed).
2. Mul i a ia e geno ype – mul i a ia e pheno ype
A se o gene ic a ian s es ed o an associa ion wi h a se o
pheno ypic a iables.
The i s ype co esponds o a s anda d mul i- ai analysis. The
second ype akes in o accoun he e ec s ac oss genomic a ian s
on mul iple ai s, which a e igno ed when analyzing only a single
SNP o a single ai a a ime.
2.3 Me a-analysis
me aCCA allows o conduc summa y s a is ics-based mul i a ia e
analysis o one o mul iple GWAS. In he me a-analy ic se ing, co-
a iance ma ices RðiÞ
XY ;^
RðiÞ
XX, and ^
RðiÞ
YY co esponding o i¼1, ...,M
independen s udies on he same opic a e pooled using a weigh ed
a e age:
RXY ¼ðN11ÞRð1Þ
XY þþðNM1ÞRðMÞ
XY
NM;(8)
whe e N
i
deno es he numbe o samples in he i h coho , and
N¼N1þþNM. This s ep is pe o med be o e applying he
sh inkage o he ull co a iance ma ix. As is ypical o a ixed-
e ec s me a-analysis, he weigh ed a e age is used in o de o
accoun o he a ying p ecision o he es ima es. The o mulas o
^
RXX and ^
RYY a e analogous o (8). Howe e , i all coho s included
in he me a-analysis ha e he same unde lying popula ion, only one
geno ypic co ela ion es ima e is needed.
2.4 Choosing SNPs ep esen ing a locus
When analyzing mul iple gene ic a ian s oge he , we use a p oced-
u e o selec ing om a gi en locus a se o SNPs ha join ly cap u e
a maximal amoun o gene ic a ia ion in he locus, as measu ed by
a linkage disequilib ium (LD) sco e.
In each i e a ion, a SNP g ha maximizes LD-sco e, which we
de ine as Xk^
2
gk 2
k, is selec ed, whe e he sum is o e all SNPs k ha
ha e no ye been chosen; ^
gk deno es a pa ial co ela ion be ween
SNPs gand k; 2
kindica es empi ical a iance o he esiduals o
SNP ka e he e ec s o he selec ed SNPs ha e been eg essed ou .
The esidual a iance 2
kge s smalle , i he SNP has al eady been
well explained by he p e iously chosen ones; hence, highly co e-
la ed SNPs will no be selec ed oge he . In he i s i e a ion, ^
gk is
he Pea son co ela ion coe icien be ween SNPs gand k, and
2
k¼1, meaning ha he s a ing SNP is he one cap u ing he high-
es amoun o gene ic a ia ion in he egion. Fo each locus, we
Algo i hm 1
jwhile Rno PSD
jR¼0:999 R;
diagðRÞ¼1;
1984 A.Cichonska e al.
selec he smalles numbe o SNPs ha explain, a median, o e
95% o he a iance o he emaining SNPs in he locus.
2.5 Da ase s
In o de o es ou app oach, we used geno ypic and pheno ypic
da a om h ee Finnish popula ion coho s: he Ca dio ascula Risk
in Young Finns S udy (YFS, N
1
¼2390; Rai aka i e al., 2008), he
FINRISK s udy su ey o 1997 (N
2
¼3661; Va iainen e al., 2010),
and he No he n Finland Bi h Coho 1966 (NFBC, N
3
¼4702;
Ran akallio, 1969). The de ailed desc ip ion o he coho s can be
ound in Supplemen a y Da a.
Ou pheno ype da a consis o 81 lipid measu es (Supplemen a y
Table S1) om a high- h oughpu nuclea magne ic esonance
(NMR) pla o m (Soininen e al., 2009,2015). As a p e-p ocessing
s ep, wi hin each coho , each ai was quan ile no malized, and he
e ec s o age, sex and en leading p incipal componen s o he gen-
e ic popula ion s uc u e we e eg essed ou using a linea model.
All coho s we e geno yped using Illumina a ays, and impu ed by
IMPUTE2 (Howie e al., 2009) using he 1000 Genomes P ojec e -
e ence panel (1000 Genomes P ojec Conso ium, 2012). In he ana-
lyses, we included 455 521 SNPs on ch omosome 1 and,
addi ionally, he SNPs in he ollowing 5 genes:
•APOE (apolipop o ein E), 259 SNPs on ch 19;
•CETP (choles e yl es e ans e p o ein), 387 SNPs on ch 16;
•GCKR (glucokinase (hexokinase 4) egula o ), 160 SNPs on ch 2;
•PCSK9 (p op o ein con e ase sub ilisin/kexin ype 9), 265 SNPs
on ch 1;
•NOD2 (nucleo ide-binding oligome iza ion domain con aining
2), 145 SNPs on ch 16.
We expec ed ha his se o genes would p o ide a comp ehen-
si e spec um o associa ions wi h ou pheno ypes, since APOE,
CETP,GCKR, and PCSK9 ha e well-known associa ions o lipid
le els, whe eas NOD2 is no known o ha e such an associa ion
(NHGRI GWAS ca alogue, Hindo e al., 2011, www.genome.
go /gwas udies). All SNPs used we e o good quali y: IMPUTE2
in o 0.8 (Ma chini and Howie, 2010), and mino allele equency
0.05.
Fo mul i-SNP models, we compa ed he esul s om Finnish
geno ype da a wi h hose ob ained by es ima ing he geno ypic co -
ela ion s uc u e ^
RXX om he 1000 Genomes P ojec da a on 503
Eu opean indi iduals ( elease 20130502).
Fo each coho , geno ypic and pheno ypic co ela ion s uc u es
compu ed based on XðiÞand YðiÞ, as shown in he Equa ions (4) and
(5), can be ound in Supplemen a y Figu es S3 and S4.
3 Resul s
3.1 Pe o mance assessmen
The pu pose o his sec ion is o alida e ha me aCCA applied o
summa y s a is ics p oduces simila esul s o he s anda d CCA
(MATLAB unc ion canonco ) applied o he indi idual-le el da a.
Fo me aCCA, we always use ^
RYY es ima ed by he me hod
desc ibed in Sec ion 2.2.2 using summa y s a is ics o he en i e
ch omosome 1.
We ocus on he e ec s o (i) he amoun o sh inkage applied o
he ull co a iance ma ix (me aCCA/me aCCAþ) and (ii) es ima -
ing ^
RXX om he popula ion unde lying he analysis (he e, Finnish),
o om a mo e he e ogeneous panel (he e, Eu opean indi iduals
om he 1000 Genomes da abase).
3.1.1 Uni a ia e geno ype – mul i a ia e pheno ype
We conduc ed a me a-analysis o he h ee coho s (YFS, FINRISK
and NFBC) by es ing associa ions be ween each SNP in he i e
genes (as lis ed in Sec ion 2.5; 1 216 SNPs in o al) wi h di e en
numbe s o ai s, anging om 2 o 50. Mul i- ai analyses a e
mos use ul o co ela ed ai s (S ephens, 2013). To e lec his, o
each SNP, we s a ed wi h a andomly selec ed ai , and a each
s ep o he analysis, added he ai mos ly co ela ed wi h he al-
eady chosen ones, excluding co ela ions wi h absolu e alues
abo e 0.95. Fo each SNP, we epea ed he p ocedu e h ee imes
wi h di e en s a ing lipid measu es.
The sca e plo in Figu e 2a shows ha me aCCA applied o he
coho -wise summa y s a is ics p o ides an excellen ag eemen
wi h he s anda d CCA o he pooled indi idual-le el da a. Thus, in
his one-SNP–mul i- ai analysis, due o he eliable ^
RYY es ima e
used, we can base he in e ence on me aCCA, and pu less weigh on
me aCCAþ(Fig. 2b) ha , as expec ed, p oduces conse a i e
P- alues.
The wide ange o he obse ed log 10 P- alues (0–88) shows
ha mul i a ia e associa ion es s can be e y powe ul in ealis ic
se ings, and ha ou example assesses he pe o mance o
me aCCA h oughou he ange ha is impo an in p ac ical ana-
lyses. Supplemen a y Figu e S5 u he e ines he beha iou o
me aCCA wi hin he ange mos encoun e ed in genome-wide asso-
cia ion s udies (0–10).
3.1.2 Mul i a ia e geno ype – mul i a ia e pheno ype
When bo h geno ype and pheno ype a e mul i a ia e, geno ypic co -
ela ion s uc u e ^
RXX needs o be es ima ed in addi ion o ^
RYY.We
conduc ed he me a-analysis o wo s udy coho s (YFS and NFBC),
and compu ed ^
RXX ei he om FINRISK (FIN) o om a mo e gen-
e ic popula ion o he 1000 Genomes Eu opean indi iduals
(1000G). (Supplemen a y Table S3 shows e o s o ^
RXX es ima es.)
We analyzed oge he be ween 2 and 10 highly co ela ed lipid
measu es, chosen sequen ially as in he single-SNP es s in Sec ion
3.1.1. Fo each o he i e genes, we analyzed oge he be ween 2
and 10 SNPs ha we e chosen o be app oxima ely unco ela ed o
co e a la ge p opo ion o gene ic a ia ion wi hin he gene. Each
se o SNPs was es ed o an associa ion wi h each g oup o co e-
la ed lipid measu es. We epea ed he p ocedu e en imes o each
gene, wi h di e en s a ing pheno ypes and SNPs.
The esul s a e summa ized in Figu e 2c– .Figu e 2c shows ha
when geno ypic co ela ion ^
RXX is es ima ed om he a ge popula-
ion, me aCCA p oduces highly consis en esul s wi h he s anda d
CCA based on he indi idual-le el da a. When ^
RXX is es ima ed
om a less well ma ching popula ion (Fig. 2e), he accu acy is
educed, and some log 10 P- alues become clea ly o e es ima ed.
In bo h cases, u he sh inkage by me aCCAþ emo es, almos
comple ely, any o e es ima ion (Fig. 2d, ). This p ope y is ex-
pec ed o be impo an in genome-wide associa ion s udies, whe e
me aCCAþcan p o ec om alse posi i es when geno ypic co el-
a ion s uc u e canno be accu a ely es ima ed. me aCCAþhas less
s a is ical powe han he indi idual-le el CCA, bu i is s ill able o
de ec s ong ue associa ions.
3.2 Applica ion o summa y s a is ics om SNPTEST
In he gene ics communi y, es ablished so wa e packages like
SNPTEST (Ma chini and Howie, 2010) a e used o pe o m uni a i-
a e genome-wide es s. In his sec ion, we conduc a me a-analysis o
uni a ia e esul s om s anda d SNPTEST uns on NFBC and YFS
coho s by me aCCA. These coho s ha e been me a-analyzed
me aCCA 1985
p e iously using s anda d CCA applied o pooled indi idual-le el
geno ypes and he same se um me abolomic p o iles ha we con-
side he e (Inouye e al., 2012). This single-SNP–mul i- ai GWAS
highligh ed candida e genes o a he oscle osis, and demons a ed
he powe o inco po a ing mul iple ela ed ai s in o he analysis.
He e, we show ha by me aCCA we ob ain hose same esul s wi h-
ou he access o he indi idual-le el da a, and, in addi ion o ha ,
we can also analyze mul iple SNPs join ly by using only summa y
s a is ics om he o iginal s udies.
We wan ed o choose a se o co ela ed ai s o he join ana-
lysis, and he e o e we p oceeded as ollows. By an agglome a i e
hie a chical clus e ing (a e age linkage) o R
YY
(81 ai s), we iden i-
ied g oups o ela ed lipid measu es. F om he la ges o 6 dis inc
clus e s, we selec ed a se o ai s in such a way ha no pai ex-
hibi ed co ela ion abo e 0.95. We ended up wi h a g oup o 9 lipid
measu es ela ed o 8 VLDL pa icles o di e en sizes and one HDL
pa icle (highligh ed in blue in Supplemen a y Table S1).
We conduc ed wo ypes o me a-analyses o NFBC and YFS:
1. Uni a ia e geno ype – mul i a ia e pheno ype
Each SNP om ch omosome 1 es ed o an associa ion wi h he
se o 9 co ela ed lipid measu es.
2. Mul i a ia e geno ype – mul i a ia e pheno ype
Fo each o he 5 genes (APOE, CETP, GCKR, PCSK9, NOD2),
he smalles se o SNPs ha explained, a median, o e 95% o
he a iance o he emaining SNPs is chosen (see Sec ion 2.4),
and es ed o an associa ion wi h he se o 9 co ela ed lipid
measu es.
The inpu summa y s a is ics o me aCCA we e ob ained by
pe o ming uni a ia e es s o each SNP- ai pai sepa a ely using
SNPTEST applied o he indi idual-le el da a, and ans o ming he
esul ing eg ession coe icien s using (3). The co ela ion s uc u e
o analyzed ai s, ^
RYY, was es ima ed om summa y s a is ics o
SNPs ac oss he en i e genome. The geno ypic co ela ion s uc u e
o mul i-SNP analyses, ^
RXX, was calcula ed om he FINRISK
coho .
We compa ed he esul s o me aCCA and me aCCAþwi h he
pooled indi idual-le el CCA o o iginal da ase s. Figu e 3 shows
sca e plo s o log 10 P- alues o 455 521 SNPs om ch omo-
some 1. The esul s o me aCCA demons a e an excellen ag ee-
men wi h he o iginal P- alues, alida ing ha me aCCA can
conduc eliable mul i a ia e me a-analysis om s anda d uni a ia e
GWAS so wa e ou pu . As an icipa ed, me aCCAþp oduces
Fig. 2. Sca e plo s o log 10 P- alues be ween he pooled indi idual-le el analysis o o iginal da ase s ( ull da a CCA) and me aCCA ( i s ow),
me aCCAþ(second ow). (a,b)Uni a ia e geno ype – mul i a ia e pheno ype; me a-analysis o NFBC, FINRISK and YFS coho s; (c– )Mul i a ia e geno ype –
mul i a ia e pheno ype; me a-analysis o NFBC and YFS coho s; me aCCA/me aCCAþwas used wi h ^
RXX compu ed om FINRISK (FIN; c, d), o om he 1000
Genomes da abase (1000G, 503 EUR indi iduals; e, ) In all he cases, lipid co ela ion s uc u e ^
RYY was calcula ed om uni a ia e summa y s a is ics o SNPs
om he en i e ch omosome 1. Single poin co esponds o he esul o one ou o (a–b) 178 752, (c– ) 4050 mul i a ia e es s. Numbe s a he op o each plo in-
dica e pe cen ages o a leas 0.5 uni o e es ima ed me aCCA’s/me aCCAþ’s log 10 P- alues in he anges [0, 10] (pu ple) o (10, max(log 10 P- alue)] ( ed).
This h eshold is ep esen ed by pu ple and ed lines. Supplemen a y Figu e S5 shows hese esul s es ic ed o he x-axis ange o [0, 10], and Supplemen a y
Figu e S6 illus a es he impac o he numbe o geno ypic and pheno ypic ea u es included in he analysis on he accu acy o me aCCA/me aCCAþ
1986 A.Cichonska e al.
conse a i e P- alues. He e, me aCCA is indeed he me hod o
choice in p ac ice, due o he high quali y o co a iance es ima e
used. Manha an plo s illus a ing P- alues along he ch omosome
a e shown in Supplemen a y Figu e S7. Genome-wide signi ican as-
socia ions (a he h eshold o P¼5108s anda d in he ield)
a e loca ed wi hin wo egions: USP1/DOCK7 and FCGR2A/3A/
2C/3B, which a e known o be associa ed wi h lipid me abolism
(NHGRI GWAS ca alogue, Hindo e al., 2011). me aCCA iden i-
ied bo h egions, and me aCCAþ ound he s onge ou o he wo
signals (DOCK7/USP1). Fo op-SNP in FCGR2A/3A/2C/3B,
me aCCAþ’s log 10 P- alue is 6.11, compa ed o 7.73 p oduced
by CCA on he indi idual-le el da a.
Figu e 4 summa izes he esul s o he mul i-SNP–mul i- ai
me a-analysis, and shows he pe o mance o me aCCA when di e -
en numbe s o SNPs, om 2 up o 25, ep esen ing a gene, a e
es ed join ly o an associa ion wi h he g oup o 9 ela ed lipid
ai s. Numbe s o SNPs ha a e chosen by ou app oach (Sec ion
2.4) a e ma ked wi h x.Figu e 4 alida es ha by using his p o o-
col, a gene is desc ibed well, since when adding mo e SNPs no clea
powe gain is obse ed. Bo h me aCCA and me aCCAþ(Fig. 4,
Supplemen a y Table S4) p oduced e y accu a e P- alues. Fo he
la ges signals (APOE,CETP), log 10 P- alues a e less han one
uni o e es ima ed by me aCCA, and unde es ima ed by
me aCCAþ. These di e ences would be unlikely o lead o alse in-
e ences when a e e ence signi icance le el in a gene-based analysis
was se o 0:05=20000 ¼2:5106, i.e. 5.61 on log 10 scale,
based on he e being abou 20 000 p o ein-coding genes in he
human genome. A his le el, bo h me aCCA and me aCCAþ ound
an associa ion be ween APOE,CETP,GCKR and he ne wo k o
VLDL and HDL pa icles s udied. Fo APOE and CETP, gene-
based signals a e clea ly highe han he uni a ia e ones, e en be o e
accoun ing o di e en numbe s o es s. Mo eo e , in case o
APOE, he mul i-SNP–mul i- ai signal is nea ly 4.5 uni s highe
han he single-SNP–mul i- ai one. No e ha NOD2 has no
(known) associa ion wi h me abolic ai s, and he e o e i se es as
a nega i e con ol Figu e 4 and Supplemen a y Table S4.
4 Discussion
The ad an age o mul i a ia e es ing o gene ic associa ion is well
epo ed in he li e a u e (Inouye e al., 2012;S ephens, 2013), and
also demons a ed in ou esul s (e.g. CETP in Supplemen a y Table
S4 ha has mul i a ia e P- alue 13 o de s o magni ude smalle
han any o he uni a ia e P- alues). Op imal use o co ela ed ai s
is becoming inc easingly impo an as high- h oughpu pheno yping
echnologies a e being mo e widely applied o indi idual s udy co-
ho s and la ge biobanks (Soininen e al., 2015).
We in oduced me aCCA, a compu a ional app oach o he
mul i a ia e me a-analysis o GWAS by using uni a ia e summa y
s a is ics and a e e ence da abase o gene ic da a. Thus, ou ame-
wo k ci cum en s he need o comple e mul i a ia e indi idual-
le el eco ds, and ackles he p oblem o low sample sizes in indi id-
ual coho s by a buil -in me a-analysis app oach. To ou knowledge,
me aCCA is he i s summa y s a is ics-based amewo k ha
allows mul i a ia e ep esen a ion o bo h geno ypic and pheno ypic
a iables.
In la ge me a-analy ic e o s, he abili y o wo k wi h summa y
s a is ics is bene icial, e en when he e is an access o he indi idual-
le el da a. Fo example, wi h a s udy design o he Global Lipids
Gene ics Conso ium (2013), we es ima e ha he educ ion in he
size o inpu da a be ween me aCCA and s anda d CCA could be
o e 750- old (Supplemen a y Figu e S1).
Fig. 3. Sca e plo s o log 10 P- alues om he pooled indi idual-le el CCA
o NFBC and YFS and (a)me aCCA,(b)me aCCAþ. Each poin co esponds o
one gene ic a ian om he ch omosome 1, es ed o an associa ion wi h
he g oup o 9 co ela ed lipid measu es. In o al, 455 521 SNPs we e ana-
lyzed. Red lines indica e he signi icance le el o 5 108(7.301 on log 10
scale)
2 3 5 10 15 20 25
0
2
4
6
8
10
12
14
16
#SNPs
−log10(p− alue)
APOE
CETP
GCKR
PCSK9
NOD2
15.96
13.03
6.38
3.6
0.01
APOE ( ull da a CCA)
APOE (me aCCA)
APOE ( op uni a ia e)
CETP ( ull da a CCA)
CETP (me aCCA)
CETP ( op uni a ia e)
GCKR ( ull da a CCA)
GCKR (me aCCA)
GCKR ( op uni a ia e)
PCSK9 ( ull da a CCA)
PCSK9 (me aCCA)
PCSK9 ( op uni a ia e)
NOD2 ( ull da a CCA)
NOD2 (me aCCA)
NOD2 ( op uni a ia e)
Fig. 4. Mul i-SNP–mul i- ai analysis: log 10 P- alues o CCA on pooled indi idual-le el da ase s (NFBC þYFS), and he me a-analyses conduc ed using
me aCCA, as a unc ion o he numbe o SNPs ep esen ing a gene. Se s o 2–25 SNPs we e es ed o an associa ion wi h he g oup o 9 ela ed lipid measu es.
In p ac ice, he smalles numbe o SNPs ha explain, a median, o e 95% o he a iance o he emaining SNPs would be chosen o ep esen a gene, and is
ma ked wi h x. The e olu ion o he median a iance explained e sus he numbe o SNPs is shown in Supplemen a y Figu e S8. Fo each gene, he la ges
log 10 P- alue om single-SNP–single- ai es s ( op uni a ia e) is ep esen ed by a dashed line. The la ges single-SNP–mul i- ai log 10 P- alues a e 11.54
o APOE, 23.77 o CETP, 9.64 o GCKR, 6.58 o PCSK9 and 0.97 o NOD2. The alues a e summa ized wi h de ails in Supplemen a y Table S4. The numbe o
es s in each gene is 1 o mul i-SNP, G o single-SNP–mul i- ai , and 9 G o single-SNP–single- ai es s, whe e Gis he numbe o SNPs in ha gene
me aCCA 1987
We p o ided wo a ian s o he algo i hm: me aCCA and
me aCCAþ. Based on ou esul s, me aCCA is he me hod o choice
when he accu acy o es ima ed co ela ion ma ices ^
RXX and ^
RYY is
good, i.e. ^
RXX es ima ed om gene ic da a on he a ge popula ion,
and ^
RYY es ima ed om a leas one ch omosome. In such cases, P-
alues om me aCCA we e e y accu a e, meaning ha alse posi-
i e and alse nega i e a es a e close o hose o s anda d CCA
applied o he indi idual-le el da a. We emphasize ha me aCCA
should no be used when he quali y o ^
RXX and/o ^
RYY es ima es is
educed, i.e. when a gene ic e e ence popula ion and/o summa y
s a is ics o only a small numbe o geno ypes a e a ailable. In such
cases, me aCCAþp o ed use ul o p o ec om an inc ease o alse
posi i e associa ions (Fig. 2 and Supplemen a y Figu e S9). This is
impo an in GWAS con ex , whe e alse posi i es could lead o con-
side able was e o esou ces in subsequen expe imen al and unc-
ional s udies. A opic o u u e wo k would be o u he de elop
ou cu en heu is ic s opping c i e ion o me aCCAþ o dec ease i s
alse nega i e a e wi hou sac i icing i s good alse posi i e a e.
We de i ed he amewo k assuming ha all ai s wi hin each
coho ha e been measu ed on he same numbe o indi iduals (N).
We no e ha he dis ibu ion o he es s a is ic depends on N
(Supplemen a y Da a), as do he e ec size ans o ma ion
(Equa ion 3) and me a-analysis app oach (Sec ion 2.3). While a
small p opo ion o missing da a o each ai could be handled by
s a is ical impu a ion me hods, u he wo k is equi ed o s udy
how me aCCA should be used when he sample sizes be ween he
ai s a y conside ably. Howe e , wi h high- h oughpu pheno yp-
ing echnologies, we belie e ha me aCCA can be applied o many
exis ing and o hcoming s udies.
Fo mul i a ia e pheno ype da a, se e al ypes o associa ion
es s a e possible. Na u al ques ion is which one should we p e e in
p ac ice. I is e iden ha single-SNP–mul i- ai es s can de ec
much s onge signals a some SNPs han any o he uni a ia e es s
sepa a ely (e.g. CETP in Supplemen a y Table S4), and iden i y as-
socia ions no ound by uni a ia e app oach (Inouye e al., 2012).
On he con a y, o some o he SNPs, he highes uni a ia e signal
may be clea ly highe han he mul i- ai one, e en a e accoun ing
o he inc ease in he numbe o es s. Fo example, in GCKR
(Supplemen a y Table S4), he op SNP’s ( s1260326) associa ion
was explained al eady by one o he ai s indi idually
(M.VLDL.FC). Gi en he di e ence in deg ees o eedom o he
es s, his led o a 4.6 uni s highe log 10 P- alue in he uni a ia e
es compa ed o he mul i a ia e one. Thus, o single-SNP analysis,
uni a ia e and mul i a ia e es s complemen each o he and nei he
should be excluded om conside a ion.
When also geno ypes a e mul i a ia e, e en mo e possibili ies o
associa ion es ing eme ge. To illus a e ou mul i-SNP app oach,
we p oposed a p ocedu e o selec ing, o each gene, he smalles
numbe o SNPs ha explained, a median, o e 95% o he a i-
ance o he emaining SNPs in he locus. We demons a ed ha es -
ing mul iple SNPs join ly can be mo e powe ul han single-SNP–
single- ai (APOE,CETP in Fig. 4 and Supplemen a y Table S4)
and single-SNP–mul i- ai es s (APOE in Supplemen a y Table
S4). Mo eo e , me aCCA could equally well inco po a e any o he
way o choosing he SNPs, o example, mo i a ed by unc ional an-
no a ions (ENCODE P ojec Conso ium, 2012), known exp ession
e ec s (A dlie e al., 2015) o p e ious GWAS esul s on o he ai s
(Hindo e al., 2011). A opic o u he esea ch could be o ex-
end he co a iance ma ix-based analyses om CCA o dynamic
app oaches ha lea ned om he da a he se o a ian s and ai s
o be conside ed oge he . This would ci cum en he need o es ic
he subse o a iables be o e he analysis.
We en ision ha mul i a ia e associa ion es ing using
me aCCA has a g ea po en ial o p o ide no el insigh s om al-
eady published summa y s a is ics o la ge GWAS me a-analyses on
mul i a ia e high- h oughpu pheno ypes, such as me abolomics
and ansc ip omics. Finally, we hope ha ou wo k helps ex ending
he applica ion a ea o CCA o summa y s a is ics da a also in o he
da a- ich ields ou side gene ics.
Funding
Funding: This wo k was inancially suppo ed by he Helsinki Doc o al
Educa ion Ne wo k in In o ma ion and Communica ions Technology (HICT)
o A.C., he Academy o Finland [257654 and 288509 o M.P.; 251170 o he
Finnish Cen e o Excellence in Compu a ional In e ence Resea ch COIN;
259272 o P.M.; 251217 and 255847 o S.R.] and he Sig id Juselius
Founda ion o S.R., A.J.K., P.S. and M.A.K. S.R. was suppo ed by EU FP7
p ojec s ENGAGE (201413), BioSHaRE (261433), he Finnish Founda ion
o Ca dio ascula Resea ch and Biocen um Helsinki. A.J.K., P.S. and
M.A.K. we e suppo ed by S a egic Resea ch Funding om he Uni e si y o
Oulu and by No o No disk Founda ion.
Con lic o In e es : A.J.K., P.S. and M.A.K. a e sha eholde s o B ainshake
L d., a company o e ing NMR-based me aboli e p o iling. A.J.K. and P.S. e-
po employmen ela ion o B ainshake L d.
Re e ences
1000 Genomes P ojec Conso ium (2012) An in eg a ed map o gene ic a i-
a ion om 1,092 human genomes. Na u e,491, 56–65.
A dlie,K.G. e al. (2015) The Geno ype-Tissue Exp ession (GTEx) pilo ana-
lysis: mul i issue gene egula ion in humans. Sci., 348, 648–660.
Deloukas,P. e al. (2013) La ge-scale associa ion analysis iden i ies new isk
loci o co ona y a e y disease. Na . Gene ., 45, 25–33.
ENCODE P ojec Conso ium (2012) An in eg a ed encyclopedia o DNA
elemen s in he human genome. Na u e,489, 57–74.
E angelou,E. and Ioannidis,J.P. (2013) Me a-analysis me hods o genome-
wide associa ion s udies and beyond. Na . Re . Gene ., 14, 379–389.
Feng,S. e al. (2014) RAREMETAL: as and powe ul me a-analysis o a e
a ian s. Bioin o ma ics, b u367.
Fe ei a,M.A. and Pu cell,S.M. (2009) A mul i a ia e es o associa ion.
Bioin o ma ics,25, 132–133.
Global Lipids Gene ics Conso ium (2013) Disco e y and e inemen o loci
associa ed wi h lipid le els. Na . Gene ., 45, 1274–1283.
Hindo ,L.A. e al. (2011) A Ca alog o Published Genome-Wide Associa ion
S udies. www.genome.go /gwas udies. (July 2015, da e las accessed).
Ho elling,H. (1936) Rela ions be ween wo se s o a ia es. Biome ika,28,
321–377.
Howie,B.N. e al. (2009) A lexible and accu a e geno ype impu a ion me hod
o he nex gene a ion o genome-wide associa ion s udies. PLos Gene ., 5,
e1000529.
Inouye,M. e al. (2012) No el Loci o me abolic ne wo ks and mul i- issue
exp ession s udies e eal genes o a he oscle osis. PLos Gene ., 8,
e1002907.
Ke unen,J. e al. (2012) Genome-wide associa ion s udy iden i ies mul iple
loci in luencing human se um me aboli e le els. Na . Gene ., 44, 269–276.
Ledoi ,O. and Wol ,M. (2003) Imp o ed es ima ion o he co a iance ma ix
o s ock e u ns wi h an applica ion o po olio selec ion. J. Empi . Finance,
10, 603–621.
Lee,S. e al. (2014) Ra e- a ian associa ion analysis: s udy designs and s a is-
ical es s. Am. J. Hum. Gene ., 95, 5–23.
Mahajan,A. e al. (2014) Genome-wide ans-ances y me a-analysis p o ides
insigh in o he gene ic a chi ec u e o ype 2 diabe es suscep ibili y. Na .
Gene ., 46, 234–244.
Ma chini,J. and Howie,B. (2010) Geno ype impu a ion o genome-wide asso-
cia ion s udies. Na . Re . Gene ., 11, 499–511.
Ma inen,P. e al. (2013) Genome-wide associa ion s udies wi h high-dimen-
sional pheno ypes. S a . Appl. Gene . Mol. Biol., 12, 413–431.
1988 A.Cichonska e al.
Ma inen,P. e al. (2014) Assessing mul i a ia e gene-me abolome associ-
a ions wi h a e a ian s using Bayesian educed ank eg ession.
Bioin o ma ics,30, 2026–2034.
Rai aka i,O.T. e al. (2008) Coho p o ile: he Ca dio ascula Risk in Young
Finns S udy. In . J. Epidemiol., 37, 1220–1226.
Ran akallio,P. (1969) G oups a isk in low bi h weigh in an s and pe ina al
mo ali y. Ac a Paedia ica Scand., 193, 193.
Schizoph enia Wo king G oup o he Psychia ic Genomics Conso ium
(2014) Biological insigh s om 108 schizoph enia-associa ed gene ic loci.
Na u e,511, 421–427.
Shin,S.Y. e al. (2014) An a las o gene ic in luences on human blood me abol-
i es. Na . Gene ., 46, 543–550.
Soininen,P. e al. (2009) High- h oughpu se um NMR me abonomics o
cos -e ec i e holis ic s udies on sys emic me abolism. Analys ,134,
1781–1785.
Soininen,P. e al. (2015) Quan i a i e se um nuclea magne ic esonance
me abolomics in ca dio ascula epidemiology and gene ics. Ci cul.:
Ca dio asc. Gene ., 8, 192–206.
S ephens,M. (2013) A uni ied amewo k o associa ion analysis wi h mul-
iple ela ed pheno ypes. PLoS One,8, e65245.
Su akka,I. e al. (2015) The impac o low- equency and a e a ian s on lipid
le els. Na . Gene ., 47, 589–597.
Tang,C.S. and Fe ei a,M.A. (2012) A gene-based es o associa ion using ca-
nonical co ela ion analysis. Bioin o ma ics,28, 845–850.
Tibshi ani,R. e al. (2001) Es ima ing he numbe o clus e s in a da a se ia
he gap s a is ic. J. R. S a . Soc.: Se . B (S a . Me hodol.),63, 411–423.
an de Sluis,S. e al. (2013) TATES: e icien mul i a ia e geno ype-pheno-
ype analysis o genome-wide associa ion s udies. PLos Gene ., 9,
e1003235.
Va iainen,E. e al. (2010) Thi y- i e-yea ends in ca dio ascula isk ac o s
in Finland. In . J. Epidemiol., 39, 504–518.
Vucko ic,D. e al. (2015) Mul iMe a: an R package o me a-analyzing mul i-
pheno ype genome-wide associa ion s udies. Bioin o ma ics, b 222.
Wall,J.D. and P i cha d,J.K. (2003) Haplo ype blocks and linkage disequilib-
ium in he human genome. Na . Re . Gene ., 4, 587–597.
Wang,Y. e al. (2013) Random o es s on Hadoop o Genome-Wide
Associa ion S udies o mul i a ia e neu oimaging pheno ypes. BMC Bioin .,
14, 1–15.
Yang,J. e al. (2012) Condi ional and join mul iple-SNP analysis o GWAS
summa y s a is ics iden i ies addi ional a ian s in luencing complex ai s.
Na . Gene ., 44, 369–375.
Zhu,X. e al. (2015) Me a-analysis o co ela ed ai s ia summa y s a is ics
om GWASs wi h an applica ion in hype ension. Am. J. Hum. Gene ., 96,
21–36.
me aCCA 1989