scieee Science in your language
[en] (orig)

metaCCA: summary statistics-based multivariate meta-analysis of genome-wide association studies using canonical correlation analysis

Abstract

MOTIVATION: A dominant approach to genetic association studies is to perform univariate tests between genotype-phenotype pairs. However, analyzing related traits together increases statistical power, and certain complex associations become detectable only when several variants are tested jointly. Currently, modest sample sizes of individual cohorts, and restricted availability of individual-level genotype-phenotype data across the cohorts limit conducting multivariate tests. RESULTS: We introduce metaCCA, a computational framework for summary statistics-based analysis of a single or multiple studies that allows multivariate representation of both genotype and phenotype. It extends the statistical technique of canonical correlation analysis to the setting where original individual-level records are not available, and employs a covariance shrinkage algorithm to achieve robustness.Multivariate meta-analysis of two Finnish studies of nuclear magnetic resonance metabolomics by metaCCA, using standard univariate output from the program SNPTEST, shows an excellent agreement with the pooled individual-level analysis of original data. Motivated by strong multivariate signals in the lipid genes tested, we envision that multivariate association testing using metaCCA has a great potential to provide novel insights from already published summary statistics from high-throughput phenotyping technologies.

Read accessible full text

metaCCA: summary statistics-based multivariate meta-analysis of genome-wide association studies using canonical correlation analysis

Author: Cichonska, Anna,Rousu, Juho,Marttinen, Pekka,Kangas, Antti J,Soininen, Pasi,Lehtimäki, Terho,Raitakari, Olli T,Järvelin, Marjo-Riitta,Salomaa, Veikko,Ala-Korpela, Mika,Ripatti, Samuli,Pirinen, Matti
Year: 2016
Source: https://trepo.tuni.fi/bitstream/10024/100384/1/metacca_summary_statistics-based_2016.pdf
Gene ics and popula ion analysis
me aCCA: summa y s a is ics-based
mul i a ia e me a-analysis o genome-wide
associa ion s udies using canonical
co ela ion analysis
Anna Cichonska
1,2,
*, Juho Rousu
2
, Pekka Ma inen
2
, An i J. Kangas
3
,
Pasi Soininen
3,4
, Te ho Leh ima¨ki
5
, Olli T. Rai aka i
6,7
,
Ma jo-Rii a Ja¨ elin
8,9,10,11
, Veikko Salomaa
12
, Mika Ala-Ko pela
3,4,13
,
Samuli Ripa i
1,14,15
and Ma i Pi inen
1,
*
1
Ins i u e o Molecula Medicine Finland FIMM, Uni e si y o Helsinki, Helsinki, Finland,
2
Helsinki Ins i u e o
In o ma ion Technology HIIT, Depa men o Compu e Science, Aal o Uni e si y, Espoo, Finland,
3
Compu a ional
Medicine, Uni e si y o Oulu, Oulu Uni e si y Hospi al and Biocen e Oulu, Oulu, Finland,
4
NMR Me abolomics
Labo a o y, School o Pha macy, Uni e si y o Eas e n Finland, Kuopio, Finland,
5
Depa men o Clinical Chemis y,
Fimlab Labo a o ies, Uni e si y o Tampe e School o Medicine, Tampe e, Finland,
6
Depa men o Clinical
Physiology and Nuclea Medicine, Uni e si y o Tu ku and Tu ku Uni e si y Hospi al, Tu ku, Finland,
7
Resea ch
Cen e o Applied and P e en i e Ca dio ascula Medicine, Uni e si y o Tu ku and Depa men o Clinical
Physiology and Nuclea Medicine, Tu ku Uni e si y Hospi al, Tu ku, Finland,
8
Depa men o Epidemiology and
Bios a is ics, MRC-PHE Cen e o En i onmen & Heal h, School o Public Heal h, Impe ial College London,
London, UK,
9
Cen e o Li e Cou se Epidemiology, Facul y o Medicine, Uni e si y o Oulu, Oulu, Finland,
10
Biocen e Oulu, Uni e si y o Oulu, Oulu, Finland,
11
Uni o P ima y Ca e, Oulu Uni e si y Hospi al, Oulu, Finland,
12
Na ional Ins i u e o Heal h and Wel a e, Helsinki, Finland,
13
Compu a ional Medicine, School o Social and
Communi y Medicine and he Medical Resea ch Council In eg a i e Epidemiology Uni , Uni e si y o B is ol,
B is ol, UK,
14
Public Heal h, Uni e si y o Helsinki, Helsinki, Finland and
15
Wellcome T us Sange Ins i u e,
Wellcome T us Genome Campus, Hinx on, UK
*To whom co espondence should be add essed.
Associa e Edi o : Jane Kelso
Recei ed on 16 July 2015; e ised on 4 Decembe 2015; accep ed on 19 Janua y 2016
Abs ac
Mo i a ion: A dominan app oach o gene ic associa ion s udies is o pe o m uni a ia e es s be-
ween geno ype-pheno ype pai s. Howe e , analyzing ela ed ai s oge he inc eases s a is ical
powe , and ce ain complex associa ions become de ec able only when se e al a ian s a e es ed
join ly. Cu en ly, modes sample sizes o indi idual coho s, and es ic ed a ailabili y o indi id-
ual-le el geno ype-pheno ype da a ac oss he coho s limi conduc ing mul i a ia e es s.
Resul s: We in oduce me aCCA, a compu a ional amewo k o summa y s a is ics-based analysis
o a single o mul iple s udies ha allows mul i a ia e ep esen a ion o bo h geno ype and pheno ype.
I ex ends he s a is ical echnique o canonical co ela ion analysis o he se ing whe e o iginal indi id-
ual-le el eco ds a e no a ailable, and employs a co a iance sh inkage algo i hm o achie e
obus ness.
Mul i a ia e me a-analysis o wo Finnish s udies o nuclea magne ic esonance me abolomics by
me aCCA, using s anda d uni a ia e ou pu om he p og am SNPTEST, shows an excellen
V
CThe Au ho 2016. Published by Ox o d Uni e si y P ess. 1981
This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion Non-Comme cial License (h p://c ea i ecommons.o g/licenses/by-nc/4.0/),
which pe mi s non-comme cial e-use, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly ci ed. Fo comme cial e-use, please con ac
jou nals.pe [email protected]
Bioin o ma ics, 32(13), 2016, 1981–1989
doi: 10.1093/bioin o ma ics/b w052
Ad ance Access Publica ion Da e: 19 Feb ua y 2016
O iginal Pape
ag eemen wi h he pooled indi idual-le el analysis o o iginal da a. Mo i a ed by s ong mul i a i-
a e signals in he lipid genes es ed, we en ision ha mul i a ia e associa ion es ing using
me aCCA has a g ea po en ial o p o ide no el insigh s om al eady published summa y s a is ics
om high- h oughpu pheno yping echnologies.
A ailabili y and implemen a ion: Code is a ailable a h ps://gi hub.com/aal o-ics-kepaco
Con ac s: [email p o ec ed] o [email p o ec ed]
Supplemen a y in o ma ion: Supplemen a y da a a e a ailable a Bioin o ma ics online.
1 In oduc ion
Mos human diseases and ai s ha e a s ong gene ic componen .
Genome-wide associa ion s udies (GWAS) ha e p o en e ec i e in
iden i ying gene ic a ia ion con ibu ing o common complex dis-
o de s, including ype 2 diabe es (Mahajan e al., 2014), ca dio as-
cula disease (Deloukas e al., 2013), schizoph enia (Schizoph enia
Wo king G oup o he Psychia ic Genomics Conso ium, 2014),
and quan i a i e ai s, such as lipid le els (Global Lipids Gene ics
Conso ium, 2013;Su akka e al., 2015) and me abolomics
(Ke unen e al., 2012;Shin e al., 2014).
A dominan app oach o GWAS is o es one single-nucleo ide
polymo phism (SNP) a a ime agains one quan i a i e pheno ype
measu e o a bina y disease indica o . This uni a ia e app oach is un-
likely o be op imal when millions o SNPs and a g owing numbe o
pheno ypes, including se um me abolomic p o iles (Ke unen e al.,
2012;Shin e al.,2014), h ee-dimensional images (Wang e al.,
2013), and gene exp ession da a (A dlie e al.,2015)becomea ailable
simul aneously. Indeed, a ecen compa ison demons a ed ha u iliz-
ing mul i a ia e pheno ype ep esen a ion inc eases s a is ical powe ,
and leads o iche indings in he associa ion es s compa ed o he
uni a ia e analysis (Inouye e al., 2012). Mo eo e , some complex
geno ype-pheno ype co ela ions can be de ec ed only when es ing
se e al gene ic a ian s simul aneously (Ma inen e al.,2014), and
mul i-geno ype es s a e common p ac ice in a e a ian associa ion
s udies, whe e s a is ical powe o de ec any single a ian is e y
small (Feng e al.,2014;Lee e al.,2014).
Un o una ely, es ic ed a ailabili y o comple e mul i a ia e indi-
idual-le el eco ds ac oss he coho s cu en ly limi s mul i a ia e
analyses. O en, only he uni a ia e GWAS summa y s a is ics, i.e. uni-
a ia e eg ession coe icien s wi h hei s anda d e o s, om indi id-
ual coho s a e publicly a ailable. Hence, a majo ques ion is how we
can use hese uni a ia e associa ion esul s o ca y ou a mul i a ia e
me a-analysis o GWAS (E angelou and Ioannidis, 2013), which is c u-
cial o inc ease he powe o iden i y no el gene ic associa ions.
Recen ly, wo kinds o mul i a ia e es ing app oaches ope a ing
on uni a ia e summa y s a is ics ha e been in oduced: (i) one SNP
agains mul iple ai s (S ephens, 2013; an de Sluis e al., 2013;
Vucko ic e al., 2015;Zhu e al., 2015) and (ii) mul iple SNPs
agains one ai (Feng e al., 2014;Yang e al., 2012). We p opose a
new amewo k, me aCCA, ha uni ies bo h o he exis ing
app oaches by allowing canonical co ela ion analysis (CCA) o
mul iple SNPs agains mul iple ai s based on uni a ia e summa y
s a is ics and publicly a ailable da abases.
CCA is a well-es ablished s a is ical echnique o iden i ying lin-
ea ela ionships be ween wo se s o a iables, and has been suc-
cess ully applied o GWAS (Fe ei a and Pu cell, 2009;Inouye e al.,
2012;Ma inen e al., 2013;Tang and Fe ei a, 2012). Ou
me aCCA me hod ex ends CCA o he se ing whe e o iginal indi-
idual-le el measu emen s a e no a ailable. Ins ead, me aCCA
wo ks wi h h ee pieces o he ull da a co a iance ma ix, and
applies a co a iance sh inkage algo i hm o achie e obus ness.
We demons a e he pe o mance o me aCCA using SNP and me-
aboli e da a om h ee Finnish coho s. In summa y, his pape
makes he ollowing con ibu ions.
•To ou knowledge, we p o ide he i s compu a ional ame-
wo k o associa ion es ing be ween mul i a ia e geno ype and
mul i a ia e pheno ype, based on uni a ia e summa y s a is ics
om single o mul iple GWAS. Ou implemen a ion is eely
a ailable.
•We demons a e how o accu a ely es ima e co ela ion s uc-
u es o pheno ypic and geno ypic a iables wi hou an access o
he indi idual-le el da a.
•We a oid alse posi i e associa ions by a co a iance sh inkage al-
go i hm based on s abiliza ion o he leading canonical
co ela ion.
•Ou app oach, me aCCA, is a gene al amewo k o conduc
CCA when ull da a a e no a ailable, and he e o e i is widely
applicable also ou side GWAS.
A de ailed discussion on he ela ionship be ween me aCCA and
p e iously published mul i a ia e associa ion me hods can be ound
in Supplemen a y Da a.
2 Me hods
This sec ion is o ganized as ollows. Fi s , Sec ion 2.1 explains uni-
a ia e GWAS, he esul s o which, in he o m o c oss-co a iance
ma ix, cons i u e an inpu o me aCCA desc ibed in Sec ion 2.2;
Sec ion 2.3 demons a es how a me a-analysis o se e al s udies is
conduc ed in ou amewo k; Sec ion 2.4 ou lines a p ocedu e o
choosing SNPs ep esen a i e o a gi en locus; inally, Sec ion 2.5
in oduces he da a we used o es me aCCA in he me a-analy ic
se ing.
2.1 Uni a ia e GWAS
Le Xand Ydeno e geno ype and pheno ype ma ices o dimensions
NGand NP, espec i ely, s o ing he indi idual-le el da a; N
he numbe o samples; Gand P he numbe o geno ypic and
pheno ypic a iables, espec i ely. The columns o Xand Ya e
s anda dized o ha e mean 0 and s anda d de ia ion 1.
Typically, uni a ia e GWAS analysis o quan i a i e ai s es s
o an associa ion be ween each pai o geno ype xg2RNand
pheno ype yp2RNsepa a ely using a linea model:
yp¼agp þxgbgp þe:(1)
Coe icien b
gp
, co esponding o he slope o he eg ession line,
is he pa ame e o in e es , since i depic s he size o he e ec
o he gene ic a ian xgon he ai yp.Pa ame e a
gp
is an in e cep
on he y-axis, and eindica es a Gaussian e o e m o noise. The
model is i by he me hod o leas squa es ha leads o a closed- o m
1982 A.Cichonska e al.
es ima e o he unknown pa ame e bgp ¼½xT
gyp½xT
gxg1¼
½ðN1Þsxy½ðN1Þsxx1¼sxy,whe es
xy
is a sample co a iance o
xgand yp,ands
xx
¼1 is a sample a iance o xg. Hence, he c oss-co-
a iance ma ix R
XY
be ween all geno ypic and pheno ypic a iables is
made o uni a ia e eg ession coe icien s b
gp
:
RXY ¼XTY
N1¼
b11 b12  b1P
b21 b22  b2P
.
.
..
.
...
..
.
.
bG1bG2 bGP
0
B
B
B
B
B
B
@
1
C
C
C
C
C
C
A
:(2)
An impo an no e is ha i he indi idual-le el da ase s Xand Y
we e no s anda dized be o e applying he linea eg ession, he
s anda diza ion can be achie ed a e wa ds by a ans o ma ion
bSTANDR
gp ¼1
ffiffiffiffiffi
N
pSEgp bgp;(3)
whe e SE
gp
indica es he s anda d e o o b
gp
, as gi en by GWAS
so wa e. (Typically, SEgp  p=ðffiffiffiffiffi
N
pffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
2 gð1 gÞ
pÞ, whe e
p
is he
s anda d de ia ion o he ai p, and
g
is he mino allele equency
o SNP g, bu unce ain y in geno ype impu a ion causes de ia ions
om his exp ession.)
2.2 me aCCA
Conduc ing mul i a ia e associa ion es s equi es es ima es o he
dependencies be ween geno ypic and pheno ypic a iables, deno ed
R
XX
and R
YY,
espec i ely. Typically, hey a e calcula ed based on
he indi idual-le el measu emen s Xand Y:
RXX ¼XTX
N1;(4)
RYY ¼YTY
N1:(5)
me aCCA ope a es on he c oss-co a iance ma ix R
XY
(Equa ion
2), and co ela ion s uc u es ^
RXX;^
RYY, es ima ed wi hou an access
o he indi idual-le el da a Xand Y(Fig. 1A, B). To make he esul -
ing ull co a iance ma ix Ra alid co a iance ma ix, me aCCA
applies a sh inkage algo i hm (Fig. 1C).
The es o his sec ion desc ibes he de ails o me aCCA
amewo k.
2.2.1 Es ima ion o geno ypic co ela ion s uc u e
Gene ic a ia ion is o ganized in haplo ype blocks, whose s uc u e
is de e mined by mu a ion and ecombina ion e en s, oge he wi h
demog aphic e ec s, including popula ion g ow h, admix u e and
bo lenecks (Wall and P i cha d, 2003). Hence, co ela ion s uc u e
o gene ic a ian s di e s be ween popula ions, such as, e.g. he
Finns, Icelande s o Cen al Eu opeans. In me aCCA,^
RXX is calcu-
la ed using a e e ence da abase ep esen ing he s udy popula ion,
such as he 1000 Genomes da abase (1000 Genomes P ojec
Conso ium, 2012, www.1000genomes.o g), o o he geno ypic
da a a ailable on he a ge popula ion. In he Sec ion 3, we demon-
s a e ha es ima ing ^
RXX om he a ge popula ion (in ou case,
he Finns) leads o be e esul s han u ilizing he da a comp ising
indi iduals ac oss dis inc popula ions (e.g. he Finns and o he
Eu opeans). Howe e , since e e ence da a on he a ge popula ion
may no always be a hand, we also p esen a obus bu less powe -
ul solu ion o mul i a ia e associa ion es ing by simply using
geno ypes o all indi iduals om a ce ain b oade geog aphical e-
gion (e.g. a con inen ) a ailable unde he 1000 Genomes P ojec .
2.2.2 Es ima ion o pheno ypic co ela ion s uc u e
In ou amewo k, pheno ypic co ela ion s uc u e ^
RYY is compu ed
based on R
XY.
Each en y o ^
RYY co esponds o a Pea son co el-
a ion be ween wo column ec o s o R
XY
- uni a ia e eg ession co-
e icien s o wo pheno ypic a iables sand ac oss Ggene ic
a ian s:
^
RYYðs; Þ¼ X
G
g¼1ðbgs lsÞðbg l Þ
ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
X
G
g¼1ðbgs lsÞ2
u
u
ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
X
G
g¼1ðbg l Þ2
u
u
;(6)
whe e l
s
and l
a e he mean alues ls¼1
GPG
g¼1bgs and
l ¼1
GPG
g¼1bg . (The de ailed jus i ica ion is p o ided in
Supplemen a y Da a.) In Supplemen a y Table S2, we demons a e
ha he highe he numbe o geno ypic a iables G, he lowe he
e o o he es ima e. Thus, ^
RYY should be calcula ed om summa y
Fig. 1. Schema ic pic u e showing an o e iew o me aCCA amewo k o
summa y s a is ics-based mul i a ia e associa ion es ing using canonical
co ela ion analysis. (A)me aCCA ope a es on h ee pieces o he ull co a i-
ance ma ix R:R
XY
o uni a ia e geno ype-pheno ype associa ion esul s, R
XX
o geno ype-geno ype co ela ions, and R
YY
o pheno ype-pheno ype co el-
a ions. (B)^
RXX is es ima ed om a e e ence da abase ma ching he s udy
popula ion, e.g. he 1000 Genomes, and pheno ypic co ela ion s uc u e ^
RYY
is es ima ed om R
XY.
(C) A co a iance sh inkage algo i hm is applied o add
obus ness o he me hod. Numbe s in b acke s e e o subsec ions in
Me hods. Me a-analysis o se e al s udies is pe o med by pooling co a i-
ance ma ices o he same ype, be o e s ep (C), as desc ibed in Sec ion 2.3.
The da a educ ion achie ed by me aCCA can be seen in Supplemen a y
Figu e S1
me aCCA 1983
s a is ics o all a ailable gene ic a ian s, e en i only a subse o
hem is aken o he u he analysis.
2.2.3 Canonical co ela ion analysis
CCA (Ho elling, 1936) is a mul i a ia e echnique o de ec ing lin-
ea ela ionships be ween wo g oups o a iables X2RNGand
Y2RNP, whe e Xand Ycons i u e wo di e en iews o he same
objec . The objec i e is o ind maximally co ela ed linea combin-
a ions o columns o each ma ix. This co esponds o inding ec-
o s a2RGand b2RP ha maximize
¼ðXaÞTðYbÞ
kXakkYbk¼aTRXY b
ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
aTRXXa
pffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
bTRYYb
q:(7)
The maximized co ela ion is called canonical co ela ion be-
ween Xand Y. We p o ide he echnical de ails o he me hod, as
well as i s ex ension o subsequen canonical co ela ions and hei
signi icance es ing in Supplemen a y Da a.
2.2.4 Sh inkage
A his poin , we ha e h ee co a iance ma ices, namely R
XY,
^
RXX,and
^
RYY . Howe e , in many cases, he esul ing ull co a iance ma ix
R¼
^
RXX RXY
RT
XY ^
RYY
!
is no posi i e semide ini e (PSD), and he e o e i s building blocks
canno be jus plugged in o he CCA amewo k (Equa ion 7). To
o e come his p oblem, in me aCCA, we apply sh inkage o ind a
nea es alid R(Ledoi and Wol , 2003). We use an i e a i e p oced-
u e whe e he magni udes o he o -diagonal en ies a e being
sh unk owa ds ze o un il Rbecomes PSD (Algo i hm 1).
Assu ing he PSD p ope y o he ull co a iance ma ix is neces-
sa y, al hough, as we demons a e in he Sec ion 3, no su icien o
ob ain eliable esul s o he associa ion analysis when he es ima e
^
RXX (and/o ^
RYY) is noisy. In o de o add ess his issue, we p opose a
a ian o me aCCA, called me aCCA1, whe e he ull co a iance
ma ix Ris sh unk beyond he le el gua an eeing i s PSD p ope y. A
challenge, howe e , is o ind an op imal sh inkage in ensi y.
Sh inkage applied wi hou any s opping c i e ion would lead o g ad-
ual emo al o all dependencies be ween geno ypic and pheno ypic
a iables. Ledoi and Wol (2003) in oduced an analy ic app oach o
de e mining he op imal sh inkage le el bu i equi es he indi idual-
le el da ase s Xand Y.Inme aCCAþ, we moni o he leading canon-
ical co ela ion alue , and we con inue he sh inkage o he ull co-
a iance ma ix Run il s abilizes. Speci ically, we ack he pe cen
change pc o be ween subsequen sh inkage i e a ions, and we de e -
mine an app op ia e amoun o sh inkage using an elbow heu is ic,
simila o he c i e ion o inding he numbe o clus e s, equen ly
used in he li e a u e (Tibshi ani e al., 2001). The idea is ha he slope
o he g aph should be s eep o he le o he elbow, bu s able o he
igh o i . We ind he elbow, and hus he app op ia e numbe o
sh inkage i e a ions, by aking he poin closes o he o igin o he
plo o pc e sus i e a ion numbe , as schema ically shown in
Supplemen a y Figu e S2.
Building blocks ^
RXY ;^
RXX;^
RYY o he esul ing ull co a iance ma-
ix R, sh unk un il i became PSD o beyond, a e hen plugged in o he
CCA amewo k o ge he inal geno ype-pheno ype associa ion esul .
In p ac ice, in o de o p o ec om alse posi i e signals, he sh inkage
mode o me aCCAþshould be applied whene e ^
RYY is es ima ed
om summa y s a is icso asmallnumbe o gene ic a ian s,and/o
^
RXX is calcula ed using a gene ic e e ence popula ion.
2.2.5 Types o he mul i a ia e associa ion analysis
We conside he ollowing wo ypes o he mul i a ia e analysis.
1. Uni a ia e geno ype – mul i a ia e pheno ype
One gene ic a ian es ed o an associa ion wi h a se o pheno-
ypic a iables (ma ix ^
RXX no needed).
2. Mul i a ia e geno ype – mul i a ia e pheno ype
A se o gene ic a ian s es ed o an associa ion wi h a se o
pheno ypic a iables.
The i s ype co esponds o a s anda d mul i- ai analysis. The
second ype akes in o accoun he e ec s ac oss genomic a ian s
on mul iple ai s, which a e igno ed when analyzing only a single
SNP o a single ai a a ime.
2.3 Me a-analysis
me aCCA allows o conduc summa y s a is ics-based mul i a ia e
analysis o one o mul iple GWAS. In he me a-analy ic se ing, co-
a iance ma ices RðiÞ
XY ;^
RðiÞ
XX, and ^
RðiÞ
YY co esponding o i¼1, ...,M
independen s udies on he same opic a e pooled using a weigh ed
a e age:
RXY ¼ðN11ÞRð1Þ
XY þþðNM1ÞRðMÞ
XY
NM;(8)
whe e N
i
deno es he numbe o samples in he i h coho , and
N¼N1þþNM. This s ep is pe o med be o e applying he
sh inkage o he ull co a iance ma ix. As is ypical o a ixed-
e ec s me a-analysis, he weigh ed a e age is used in o de o
accoun o he a ying p ecision o he es ima es. The o mulas o
^
RXX and ^
RYY a e analogous o (8). Howe e , i all coho s included
in he me a-analysis ha e he same unde lying popula ion, only one
geno ypic co ela ion es ima e is needed.
2.4 Choosing SNPs ep esen ing a locus
When analyzing mul iple gene ic a ian s oge he , we use a p oced-
u e o selec ing om a gi en locus a se o SNPs ha join ly cap u e
a maximal amoun o gene ic a ia ion in he locus, as measu ed by
a linkage disequilib ium (LD) sco e.
In each i e a ion, a SNP g ha maximizes LD-sco e, which we
de ine as Xk^
2
gk 2
k, is selec ed, whe e he sum is o e all SNPs k ha
ha e no ye been chosen; ^
gk deno es a pa ial co ela ion be ween
SNPs gand k; 2
kindica es empi ical a iance o he esiduals o
SNP ka e he e ec s o he selec ed SNPs ha e been eg essed ou .
The esidual a iance 2
kge s smalle , i he SNP has al eady been
well explained by he p e iously chosen ones; hence, highly co e-
la ed SNPs will no be selec ed oge he . In he i s i e a ion, ^
gk is
he Pea son co ela ion coe icien be ween SNPs gand k, and
2
k¼1, meaning ha he s a ing SNP is he one cap u ing he high-
es amoun o gene ic a ia ion in he egion. Fo each locus, we
Algo i hm 1
jwhile Rno PSD
jR¼0:999 R;
diagðRÞ¼1;
1984 A.Cichonska e al.
selec he smalles numbe o SNPs ha explain, a median, o e
95% o he a iance o he emaining SNPs in he locus.
2.5 Da ase s
In o de o es ou app oach, we used geno ypic and pheno ypic
da a om h ee Finnish popula ion coho s: he Ca dio ascula Risk
in Young Finns S udy (YFS, N
1
¼2390; Rai aka i e al., 2008), he
FINRISK s udy su ey o 1997 (N
2
¼3661; Va iainen e al., 2010),
and he No he n Finland Bi h Coho 1966 (NFBC, N
3
¼4702;
Ran akallio, 1969). The de ailed desc ip ion o he coho s can be
ound in Supplemen a y Da a.
Ou pheno ype da a consis o 81 lipid measu es (Supplemen a y
Table S1) om a high- h oughpu nuclea magne ic esonance
(NMR) pla o m (Soininen e al., 2009,2015). As a p e-p ocessing
s ep, wi hin each coho , each ai was quan ile no malized, and he
e ec s o age, sex and en leading p incipal componen s o he gen-
e ic popula ion s uc u e we e eg essed ou using a linea model.
All coho s we e geno yped using Illumina a ays, and impu ed by
IMPUTE2 (Howie e al., 2009) using he 1000 Genomes P ojec e -
e ence panel (1000 Genomes P ojec Conso ium, 2012). In he ana-
lyses, we included 455 521 SNPs on ch omosome 1 and,
addi ionally, he SNPs in he ollowing 5 genes:
•APOE (apolipop o ein E), 259 SNPs on ch 19;
•CETP (choles e yl es e ans e p o ein), 387 SNPs on ch 16;
•GCKR (glucokinase (hexokinase 4) egula o ), 160 SNPs on ch 2;
•PCSK9 (p op o ein con e ase sub ilisin/kexin ype 9), 265 SNPs
on ch 1;
•NOD2 (nucleo ide-binding oligome iza ion domain con aining
2), 145 SNPs on ch 16.
We expec ed ha his se o genes would p o ide a comp ehen-
si e spec um o associa ions wi h ou pheno ypes, since APOE,
CETP,GCKR, and PCSK9 ha e well-known associa ions o lipid
le els, whe eas NOD2 is no known o ha e such an associa ion
(NHGRI GWAS ca alogue, Hindo e al., 2011, www.genome.
go /gwas udies). All SNPs used we e o good quali y: IMPUTE2
in o 0.8 (Ma chini and Howie, 2010), and mino allele equency
0.05.
Fo mul i-SNP models, we compa ed he esul s om Finnish
geno ype da a wi h hose ob ained by es ima ing he geno ypic co -
ela ion s uc u e ^
RXX om he 1000 Genomes P ojec da a on 503
Eu opean indi iduals ( elease 20130502).
Fo each coho , geno ypic and pheno ypic co ela ion s uc u es
compu ed based on XðiÞand YðiÞ, as shown in he Equa ions (4) and
(5), can be ound in Supplemen a y Figu es S3 and S4.
3 Resul s
3.1 Pe o mance assessmen
The pu pose o his sec ion is o alida e ha me aCCA applied o
summa y s a is ics p oduces simila esul s o he s anda d CCA
(MATLAB unc ion canonco ) applied o he indi idual-le el da a.
Fo me aCCA, we always use ^
RYY es ima ed by he me hod
desc ibed in Sec ion 2.2.2 using summa y s a is ics o he en i e
ch omosome 1.
We ocus on he e ec s o (i) he amoun o sh inkage applied o
he ull co a iance ma ix (me aCCA/me aCCAþ) and (ii) es ima -
ing ^
RXX om he popula ion unde lying he analysis (he e, Finnish),
o om a mo e he e ogeneous panel (he e, Eu opean indi iduals
om he 1000 Genomes da abase).
3.1.1 Uni a ia e geno ype – mul i a ia e pheno ype
We conduc ed a me a-analysis o he h ee coho s (YFS, FINRISK
and NFBC) by es ing associa ions be ween each SNP in he i e
genes (as lis ed in Sec ion 2.5; 1 216 SNPs in o al) wi h di e en
numbe s o ai s, anging om 2 o 50. Mul i- ai analyses a e
mos use ul o co ela ed ai s (S ephens, 2013). To e lec his, o
each SNP, we s a ed wi h a andomly selec ed ai , and a each
s ep o he analysis, added he ai mos ly co ela ed wi h he al-
eady chosen ones, excluding co ela ions wi h absolu e alues
abo e 0.95. Fo each SNP, we epea ed he p ocedu e h ee imes
wi h di e en s a ing lipid measu es.
The sca e plo in Figu e 2a shows ha me aCCA applied o he
coho -wise summa y s a is ics p o ides an excellen ag eemen
wi h he s anda d CCA o he pooled indi idual-le el da a. Thus, in
his one-SNP–mul i- ai analysis, due o he eliable ^
RYY es ima e
used, we can base he in e ence on me aCCA, and pu less weigh on
me aCCAþ(Fig. 2b) ha , as expec ed, p oduces conse a i e
P- alues.
The wide ange o he obse ed log 10 P- alues (0–88) shows
ha mul i a ia e associa ion es s can be e y powe ul in ealis ic
se ings, and ha ou example assesses he pe o mance o
me aCCA h oughou he ange ha is impo an in p ac ical ana-
lyses. Supplemen a y Figu e S5 u he e ines he beha iou o
me aCCA wi hin he ange mos encoun e ed in genome-wide asso-
cia ion s udies (0–10).
3.1.2 Mul i a ia e geno ype – mul i a ia e pheno ype
When bo h geno ype and pheno ype a e mul i a ia e, geno ypic co -
ela ion s uc u e ^
RXX needs o be es ima ed in addi ion o ^
RYY.We
conduc ed he me a-analysis o wo s udy coho s (YFS and NFBC),
and compu ed ^
RXX ei he om FINRISK (FIN) o om a mo e gen-
e ic popula ion o he 1000 Genomes Eu opean indi iduals
(1000G). (Supplemen a y Table S3 shows e o s o ^
RXX es ima es.)
We analyzed oge he be ween 2 and 10 highly co ela ed lipid
measu es, chosen sequen ially as in he single-SNP es s in Sec ion
3.1.1. Fo each o he i e genes, we analyzed oge he be ween 2
and 10 SNPs ha we e chosen o be app oxima ely unco ela ed o
co e a la ge p opo ion o gene ic a ia ion wi hin he gene. Each
se o SNPs was es ed o an associa ion wi h each g oup o co e-
la ed lipid measu es. We epea ed he p ocedu e en imes o each
gene, wi h di e en s a ing pheno ypes and SNPs.
The esul s a e summa ized in Figu e 2c– .Figu e 2c shows ha
when geno ypic co ela ion ^
RXX is es ima ed om he a ge popula-
ion, me aCCA p oduces highly consis en esul s wi h he s anda d
CCA based on he indi idual-le el da a. When ^
RXX is es ima ed
om a less well ma ching popula ion (Fig. 2e), he accu acy is
educed, and some log 10 P- alues become clea ly o e es ima ed.
In bo h cases, u he sh inkage by me aCCAþ emo es, almos
comple ely, any o e es ima ion (Fig. 2d, ). This p ope y is ex-
pec ed o be impo an in genome-wide associa ion s udies, whe e
me aCCAþcan p o ec om alse posi i es when geno ypic co el-
a ion s uc u e canno be accu a ely es ima ed. me aCCAþhas less
s a is ical powe han he indi idual-le el CCA, bu i is s ill able o
de ec s ong ue associa ions.
3.2 Applica ion o summa y s a is ics om SNPTEST
In he gene ics communi y, es ablished so wa e packages like
SNPTEST (Ma chini and Howie, 2010) a e used o pe o m uni a i-
a e genome-wide es s. In his sec ion, we conduc a me a-analysis o
uni a ia e esul s om s anda d SNPTEST uns on NFBC and YFS
coho s by me aCCA. These coho s ha e been me a-analyzed
me aCCA 1985

p e iously using s anda d CCA applied o pooled indi idual-le el
geno ypes and he same se um me abolomic p o iles ha we con-
side he e (Inouye e al., 2012). This single-SNP–mul i- ai GWAS
highligh ed candida e genes o a he oscle osis, and demons a ed
he powe o inco po a ing mul iple ela ed ai s in o he analysis.
He e, we show ha by me aCCA we ob ain hose same esul s wi h-
ou he access o he indi idual-le el da a, and, in addi ion o ha ,
we can also analyze mul iple SNPs join ly by using only summa y
s a is ics om he o iginal s udies.
We wan ed o choose a se o co ela ed ai s o he join ana-
lysis, and he e o e we p oceeded as ollows. By an agglome a i e
hie a chical clus e ing (a e age linkage) o R
YY
(81 ai s), we iden i-
ied g oups o ela ed lipid measu es. F om he la ges o 6 dis inc
clus e s, we selec ed a se o ai s in such a way ha no pai ex-
hibi ed co ela ion abo e 0.95. We ended up wi h a g oup o 9 lipid
measu es ela ed o 8 VLDL pa icles o di e en sizes and one HDL
pa icle (highligh ed in blue in Supplemen a y Table S1).
We conduc ed wo ypes o me a-analyses o NFBC and YFS:
1. Uni a ia e geno ype – mul i a ia e pheno ype
Each SNP om ch omosome 1 es ed o an associa ion wi h he
se o 9 co ela ed lipid measu es.
2. Mul i a ia e geno ype – mul i a ia e pheno ype
Fo each o he 5 genes (APOE, CETP, GCKR, PCSK9, NOD2),
he smalles se o SNPs ha explained, a median, o e 95% o
he a iance o he emaining SNPs is chosen (see Sec ion 2.4),
and es ed o an associa ion wi h he se o 9 co ela ed lipid
measu es.
The inpu summa y s a is ics o me aCCA we e ob ained by
pe o ming uni a ia e es s o each SNP- ai pai sepa a ely using
SNPTEST applied o he indi idual-le el da a, and ans o ming he
esul ing eg ession coe icien s using (3). The co ela ion s uc u e
o analyzed ai s, ^
RYY, was es ima ed om summa y s a is ics o
SNPs ac oss he en i e genome. The geno ypic co ela ion s uc u e
o mul i-SNP analyses, ^
RXX, was calcula ed om he FINRISK
coho .
We compa ed he esul s o me aCCA and me aCCAþwi h he
pooled indi idual-le el CCA o o iginal da ase s. Figu e 3 shows
sca e plo s o log 10 P- alues o 455 521 SNPs om ch omo-
some 1. The esul s o me aCCA demons a e an excellen ag ee-
men wi h he o iginal P- alues, alida ing ha me aCCA can
conduc eliable mul i a ia e me a-analysis om s anda d uni a ia e
GWAS so wa e ou pu . As an icipa ed, me aCCAþp oduces
Fig. 2. Sca e plo s o log 10 P- alues be ween he pooled indi idual-le el analysis o o iginal da ase s ( ull da a CCA) and me aCCA ( i s ow),
me aCCAþ(second ow). (a,b)Uni a ia e geno ype – mul i a ia e pheno ype; me a-analysis o NFBC, FINRISK and YFS coho s; (c– )Mul i a ia e geno ype –
mul i a ia e pheno ype; me a-analysis o NFBC and YFS coho s; me aCCA/me aCCAþwas used wi h ^
RXX compu ed om FINRISK (FIN; c, d), o om he 1000
Genomes da abase (1000G, 503 EUR indi iduals; e, ) In all he cases, lipid co ela ion s uc u e ^
RYY was calcula ed om uni a ia e summa y s a is ics o SNPs
om he en i e ch omosome 1. Single poin co esponds o he esul o one ou o (a–b) 178 752, (c– ) 4050 mul i a ia e es s. Numbe s a he op o each plo in-
dica e pe cen ages o a leas 0.5 uni o e es ima ed me aCCA’s/me aCCAþ’s log 10 P- alues in he anges [0, 10] (pu ple) o (10, max(log 10 P- alue)] ( ed).
This h eshold is ep esen ed by pu ple and ed lines. Supplemen a y Figu e S5 shows hese esul s es ic ed o he x-axis ange o [0, 10], and Supplemen a y
Figu e S6 illus a es he impac o he numbe o geno ypic and pheno ypic ea u es included in he analysis on he accu acy o me aCCA/me aCCAþ
1986 A.Cichonska e al.
conse a i e P- alues. He e, me aCCA is indeed he me hod o
choice in p ac ice, due o he high quali y o co a iance es ima e
used. Manha an plo s illus a ing P- alues along he ch omosome
a e shown in Supplemen a y Figu e S7. Genome-wide signi ican as-
socia ions (a he h eshold o P¼5108s anda d in he ield)
a e loca ed wi hin wo egions: USP1/DOCK7 and FCGR2A/3A/
2C/3B, which a e known o be associa ed wi h lipid me abolism
(NHGRI GWAS ca alogue, Hindo e al., 2011). me aCCA iden i-
ied bo h egions, and me aCCAþ ound he s onge ou o he wo
signals (DOCK7/USP1). Fo op-SNP in FCGR2A/3A/2C/3B,
me aCCAþ’s log 10 P- alue is 6.11, compa ed o 7.73 p oduced
by CCA on he indi idual-le el da a.
Figu e 4 summa izes he esul s o he mul i-SNP–mul i- ai
me a-analysis, and shows he pe o mance o me aCCA when di e -
en numbe s o SNPs, om 2 up o 25, ep esen ing a gene, a e
es ed join ly o an associa ion wi h he g oup o 9 ela ed lipid
ai s. Numbe s o SNPs ha a e chosen by ou app oach (Sec ion
2.4) a e ma ked wi h x.Figu e 4 alida es ha by using his p o o-
col, a gene is desc ibed well, since when adding mo e SNPs no clea
powe gain is obse ed. Bo h me aCCA and me aCCAþ(Fig. 4,
Supplemen a y Table S4) p oduced e y accu a e P- alues. Fo he
la ges signals (APOE,CETP), log 10 P- alues a e less han one
uni o e es ima ed by me aCCA, and unde es ima ed by
me aCCAþ. These di e ences would be unlikely o lead o alse in-
e ences when a e e ence signi icance le el in a gene-based analysis
was se o 0:05=20000 ¼2:5106, i.e. 5.61 on log 10 scale,
based on he e being abou 20 000 p o ein-coding genes in he
human genome. A his le el, bo h me aCCA and me aCCAþ ound
an associa ion be ween APOE,CETP,GCKR and he ne wo k o
VLDL and HDL pa icles s udied. Fo APOE and CETP, gene-
based signals a e clea ly highe han he uni a ia e ones, e en be o e
accoun ing o di e en numbe s o es s. Mo eo e , in case o
APOE, he mul i-SNP–mul i- ai signal is nea ly 4.5 uni s highe
han he single-SNP–mul i- ai one. No e ha NOD2 has no
(known) associa ion wi h me abolic ai s, and he e o e i se es as
a nega i e con ol Figu e 4 and Supplemen a y Table S4.
4 Discussion
The ad an age o mul i a ia e es ing o gene ic associa ion is well
epo ed in he li e a u e (Inouye e al., 2012;S ephens, 2013), and
also demons a ed in ou esul s (e.g. CETP in Supplemen a y Table
S4 ha has mul i a ia e P- alue 13 o de s o magni ude smalle
han any o he uni a ia e P- alues). Op imal use o co ela ed ai s
is becoming inc easingly impo an as high- h oughpu pheno yping
echnologies a e being mo e widely applied o indi idual s udy co-
ho s and la ge biobanks (Soininen e al., 2015).
We in oduced me aCCA, a compu a ional app oach o he
mul i a ia e me a-analysis o GWAS by using uni a ia e summa y
s a is ics and a e e ence da abase o gene ic da a. Thus, ou ame-
wo k ci cum en s he need o comple e mul i a ia e indi idual-
le el eco ds, and ackles he p oblem o low sample sizes in indi id-
ual coho s by a buil -in me a-analysis app oach. To ou knowledge,
me aCCA is he i s summa y s a is ics-based amewo k ha
allows mul i a ia e ep esen a ion o bo h geno ypic and pheno ypic
a iables.
In la ge me a-analy ic e o s, he abili y o wo k wi h summa y
s a is ics is bene icial, e en when he e is an access o he indi idual-
le el da a. Fo example, wi h a s udy design o he Global Lipids
Gene ics Conso ium (2013), we es ima e ha he educ ion in he
size o inpu da a be ween me aCCA and s anda d CCA could be
o e 750- old (Supplemen a y Figu e S1).
Fig. 3. Sca e plo s o log 10 P- alues om he pooled indi idual-le el CCA
o NFBC and YFS and (a)me aCCA,(b)me aCCAþ. Each poin co esponds o
one gene ic a ian om he ch omosome 1, es ed o an associa ion wi h
he g oup o 9 co ela ed lipid measu es. In o al, 455 521 SNPs we e ana-
lyzed. Red lines indica e he signi icance le el o 5 108(7.301 on log 10
scale)
2 3 5 10 15 20 25
0
2
4
6
8
10
12
14
16
#SNPs
−log10(p− alue)
APOE
CETP
GCKR
PCSK9
NOD2
15.96
13.03
6.38
3.6
0.01
APOE ( ull da a CCA)
APOE (me aCCA)
APOE ( op uni a ia e)
CETP ( ull da a CCA)
CETP (me aCCA)
CETP ( op uni a ia e)
GCKR ( ull da a CCA)
GCKR (me aCCA)
GCKR ( op uni a ia e)
PCSK9 ( ull da a CCA)
PCSK9 (me aCCA)
PCSK9 ( op uni a ia e)
NOD2 ( ull da a CCA)
NOD2 (me aCCA)
NOD2 ( op uni a ia e)
Fig. 4. Mul i-SNP–mul i- ai analysis: log 10 P- alues o CCA on pooled indi idual-le el da ase s (NFBC þYFS), and he me a-analyses conduc ed using
me aCCA, as a unc ion o he numbe o SNPs ep esen ing a gene. Se s o 2–25 SNPs we e es ed o an associa ion wi h he g oup o 9 ela ed lipid measu es.
In p ac ice, he smalles numbe o SNPs ha explain, a median, o e 95% o he a iance o he emaining SNPs would be chosen o ep esen a gene, and is
ma ked wi h x. The e olu ion o he median a iance explained e sus he numbe o SNPs is shown in Supplemen a y Figu e S8. Fo each gene, he la ges
log 10 P- alue om single-SNP–single- ai es s ( op uni a ia e) is ep esen ed by a dashed line. The la ges single-SNP–mul i- ai log 10 P- alues a e 11.54
o APOE, 23.77 o CETP, 9.64 o GCKR, 6.58 o PCSK9 and 0.97 o NOD2. The alues a e summa ized wi h de ails in Supplemen a y Table S4. The numbe o
es s in each gene is 1 o mul i-SNP, G o single-SNP–mul i- ai , and 9 G o single-SNP–single- ai es s, whe e Gis he numbe o SNPs in ha gene
me aCCA 1987
We p o ided wo a ian s o he algo i hm: me aCCA and
me aCCAþ. Based on ou esul s, me aCCA is he me hod o choice
when he accu acy o es ima ed co ela ion ma ices ^
RXX and ^
RYY is
good, i.e. ^
RXX es ima ed om gene ic da a on he a ge popula ion,
and ^
RYY es ima ed om a leas one ch omosome. In such cases, P-
alues om me aCCA we e e y accu a e, meaning ha alse posi-
i e and alse nega i e a es a e close o hose o s anda d CCA
applied o he indi idual-le el da a. We emphasize ha me aCCA
should no be used when he quali y o ^
RXX and/o ^
RYY es ima es is
educed, i.e. when a gene ic e e ence popula ion and/o summa y
s a is ics o only a small numbe o geno ypes a e a ailable. In such
cases, me aCCAþp o ed use ul o p o ec om an inc ease o alse
posi i e associa ions (Fig. 2 and Supplemen a y Figu e S9). This is
impo an in GWAS con ex , whe e alse posi i es could lead o con-
side able was e o esou ces in subsequen expe imen al and unc-
ional s udies. A opic o u u e wo k would be o u he de elop
ou cu en heu is ic s opping c i e ion o me aCCAþ o dec ease i s
alse nega i e a e wi hou sac i icing i s good alse posi i e a e.
We de i ed he amewo k assuming ha all ai s wi hin each
coho ha e been measu ed on he same numbe o indi iduals (N).
We no e ha he dis ibu ion o he es s a is ic depends on N
(Supplemen a y Da a), as do he e ec size ans o ma ion
(Equa ion 3) and me a-analysis app oach (Sec ion 2.3). While a
small p opo ion o missing da a o each ai could be handled by
s a is ical impu a ion me hods, u he wo k is equi ed o s udy
how me aCCA should be used when he sample sizes be ween he
ai s a y conside ably. Howe e , wi h high- h oughpu pheno yp-
ing echnologies, we belie e ha me aCCA can be applied o many
exis ing and o hcoming s udies.
Fo mul i a ia e pheno ype da a, se e al ypes o associa ion
es s a e possible. Na u al ques ion is which one should we p e e in
p ac ice. I is e iden ha single-SNP–mul i- ai es s can de ec
much s onge signals a some SNPs han any o he uni a ia e es s
sepa a ely (e.g. CETP in Supplemen a y Table S4), and iden i y as-
socia ions no ound by uni a ia e app oach (Inouye e al., 2012).
On he con a y, o some o he SNPs, he highes uni a ia e signal
may be clea ly highe han he mul i- ai one, e en a e accoun ing
o he inc ease in he numbe o es s. Fo example, in GCKR
(Supplemen a y Table S4), he op SNP’s ( s1260326) associa ion
was explained al eady by one o he ai s indi idually
(M.VLDL.FC). Gi en he di e ence in deg ees o eedom o he
es s, his led o a 4.6 uni s highe log 10 P- alue in he uni a ia e
es compa ed o he mul i a ia e one. Thus, o single-SNP analysis,
uni a ia e and mul i a ia e es s complemen each o he and nei he
should be excluded om conside a ion.
When also geno ypes a e mul i a ia e, e en mo e possibili ies o
associa ion es ing eme ge. To illus a e ou mul i-SNP app oach,
we p oposed a p ocedu e o selec ing, o each gene, he smalles
numbe o SNPs ha explained, a median, o e 95% o he a i-
ance o he emaining SNPs in he locus. We demons a ed ha es -
ing mul iple SNPs join ly can be mo e powe ul han single-SNP–
single- ai (APOE,CETP in Fig. 4 and Supplemen a y Table S4)
and single-SNP–mul i- ai es s (APOE in Supplemen a y Table
S4). Mo eo e , me aCCA could equally well inco po a e any o he
way o choosing he SNPs, o example, mo i a ed by unc ional an-
no a ions (ENCODE P ojec Conso ium, 2012), known exp ession
e ec s (A dlie e al., 2015) o p e ious GWAS esul s on o he ai s
(Hindo e al., 2011). A opic o u he esea ch could be o ex-
end he co a iance ma ix-based analyses om CCA o dynamic
app oaches ha lea ned om he da a he se o a ian s and ai s
o be conside ed oge he . This would ci cum en he need o es ic
he subse o a iables be o e he analysis.
We en ision ha mul i a ia e associa ion es ing using
me aCCA has a g ea po en ial o p o ide no el insigh s om al-
eady published summa y s a is ics o la ge GWAS me a-analyses on
mul i a ia e high- h oughpu pheno ypes, such as me abolomics
and ansc ip omics. Finally, we hope ha ou wo k helps ex ending
he applica ion a ea o CCA o summa y s a is ics da a also in o he
da a- ich ields ou side gene ics.
Funding
Funding: This wo k was inancially suppo ed by he Helsinki Doc o al
Educa ion Ne wo k in In o ma ion and Communica ions Technology (HICT)
o A.C., he Academy o Finland [257654 and 288509 o M.P.; 251170 o he
Finnish Cen e o Excellence in Compu a ional In e ence Resea ch COIN;
259272 o P.M.; 251217 and 255847 o S.R.] and he Sig id Juselius
Founda ion o S.R., A.J.K., P.S. and M.A.K. S.R. was suppo ed by EU FP7
p ojec s ENGAGE (201413), BioSHaRE (261433), he Finnish Founda ion
o Ca dio ascula Resea ch and Biocen um Helsinki. A.J.K., P.S. and
M.A.K. we e suppo ed by S a egic Resea ch Funding om he Uni e si y o
Oulu and by No o No disk Founda ion.
Con lic o In e es : A.J.K., P.S. and M.A.K. a e sha eholde s o B ainshake
L d., a company o e ing NMR-based me aboli e p o iling. A.J.K. and P.S. e-
po employmen ela ion o B ainshake L d.
Re e ences
1000 Genomes P ojec Conso ium (2012) An in eg a ed map o gene ic a i-
a ion om 1,092 human genomes. Na u e,491, 56–65.
A dlie,K.G. e al. (2015) The Geno ype-Tissue Exp ession (GTEx) pilo ana-
lysis: mul i issue gene egula ion in humans. Sci., 348, 648–660.
Deloukas,P. e al. (2013) La ge-scale associa ion analysis iden i ies new isk
loci o co ona y a e y disease. Na . Gene ., 45, 25–33.
ENCODE P ojec Conso ium (2012) An in eg a ed encyclopedia o DNA
elemen s in he human genome. Na u e,489, 57–74.
E angelou,E. and Ioannidis,J.P. (2013) Me a-analysis me hods o genome-
wide associa ion s udies and beyond. Na . Re . Gene ., 14, 379–389.
Feng,S. e al. (2014) RAREMETAL: as and powe ul me a-analysis o a e
a ian s. Bioin o ma ics, b u367.
Fe ei a,M.A. and Pu cell,S.M. (2009) A mul i a ia e es o associa ion.
Bioin o ma ics,25, 132–133.
Global Lipids Gene ics Conso ium (2013) Disco e y and e inemen o loci
associa ed wi h lipid le els. Na . Gene ., 45, 1274–1283.
Hindo ,L.A. e al. (2011) A Ca alog o Published Genome-Wide Associa ion
S udies. www.genome.go /gwas udies. (July 2015, da e las accessed).
Ho elling,H. (1936) Rela ions be ween wo se s o a ia es. Biome ika,28,
321–377.
Howie,B.N. e al. (2009) A lexible and accu a e geno ype impu a ion me hod
o he nex gene a ion o genome-wide associa ion s udies. PLos Gene ., 5,
e1000529.
Inouye,M. e al. (2012) No el Loci o me abolic ne wo ks and mul i- issue
exp ession s udies e eal genes o a he oscle osis. PLos Gene ., 8,
e1002907.
Ke unen,J. e al. (2012) Genome-wide associa ion s udy iden i ies mul iple
loci in luencing human se um me aboli e le els. Na . Gene ., 44, 269–276.
Ledoi ,O. and Wol ,M. (2003) Imp o ed es ima ion o he co a iance ma ix
o s ock e u ns wi h an applica ion o po olio selec ion. J. Empi . Finance,
10, 603–621.
Lee,S. e al. (2014) Ra e- a ian associa ion analysis: s udy designs and s a is-
ical es s. Am. J. Hum. Gene ., 95, 5–23.
Mahajan,A. e al. (2014) Genome-wide ans-ances y me a-analysis p o ides
insigh in o he gene ic a chi ec u e o ype 2 diabe es suscep ibili y. Na .
Gene ., 46, 234–244.
Ma chini,J. and Howie,B. (2010) Geno ype impu a ion o genome-wide asso-
cia ion s udies. Na . Re . Gene ., 11, 499–511.
Ma inen,P. e al. (2013) Genome-wide associa ion s udies wi h high-dimen-
sional pheno ypes. S a . Appl. Gene . Mol. Biol., 12, 413–431.
1988 A.Cichonska e al.
Ma inen,P. e al. (2014) Assessing mul i a ia e gene-me abolome associ-
a ions wi h a e a ian s using Bayesian educed ank eg ession.
Bioin o ma ics,30, 2026–2034.
Rai aka i,O.T. e al. (2008) Coho p o ile: he Ca dio ascula Risk in Young
Finns S udy. In . J. Epidemiol., 37, 1220–1226.
Ran akallio,P. (1969) G oups a isk in low bi h weigh in an s and pe ina al
mo ali y. Ac a Paedia ica Scand., 193, 193.
Schizoph enia Wo king G oup o he Psychia ic Genomics Conso ium
(2014) Biological insigh s om 108 schizoph enia-associa ed gene ic loci.
Na u e,511, 421–427.
Shin,S.Y. e al. (2014) An a las o gene ic in luences on human blood me abol-
i es. Na . Gene ., 46, 543–550.
Soininen,P. e al. (2009) High- h oughpu se um NMR me abonomics o
cos -e ec i e holis ic s udies on sys emic me abolism. Analys ,134,
1781–1785.
Soininen,P. e al. (2015) Quan i a i e se um nuclea magne ic esonance
me abolomics in ca dio ascula epidemiology and gene ics. Ci cul.:
Ca dio asc. Gene ., 8, 192–206.
S ephens,M. (2013) A uni ied amewo k o associa ion analysis wi h mul-
iple ela ed pheno ypes. PLoS One,8, e65245.
Su akka,I. e al. (2015) The impac o low- equency and a e a ian s on lipid
le els. Na . Gene ., 47, 589–597.
Tang,C.S. and Fe ei a,M.A. (2012) A gene-based es o associa ion using ca-
nonical co ela ion analysis. Bioin o ma ics,28, 845–850.
Tibshi ani,R. e al. (2001) Es ima ing he numbe o clus e s in a da a se ia
he gap s a is ic. J. R. S a . Soc.: Se . B (S a . Me hodol.),63, 411–423.
an de Sluis,S. e al. (2013) TATES: e icien mul i a ia e geno ype-pheno-
ype analysis o genome-wide associa ion s udies. PLos Gene ., 9,
e1003235.
Va iainen,E. e al. (2010) Thi y- i e-yea ends in ca dio ascula isk ac o s
in Finland. In . J. Epidemiol., 39, 504–518.
Vucko ic,D. e al. (2015) Mul iMe a: an R package o me a-analyzing mul i-
pheno ype genome-wide associa ion s udies. Bioin o ma ics, b 222.
Wall,J.D. and P i cha d,J.K. (2003) Haplo ype blocks and linkage disequilib-
ium in he human genome. Na . Re . Gene ., 4, 587–597.
Wang,Y. e al. (2013) Random o es s on Hadoop o Genome-Wide
Associa ion S udies o mul i a ia e neu oimaging pheno ypes. BMC Bioin .,
14, 1–15.
Yang,J. e al. (2012) Condi ional and join mul iple-SNP analysis o GWAS
summa y s a is ics iden i ies addi ional a ian s in luencing complex ai s.
Na . Gene ., 44, 369–375.
Zhu,X. e al. (2015) Me a-analysis o co ela ed ai s ia summa y s a is ics
om GWASs wi h an applica ion in hype ension. Am. J. Hum. Gene ., 96,
21–36.
me aCCA 1989