Vol. 30 no. 14 2014, pages 2026–2034
BIOINFORMATICS ORIGINAL PAPER doi:10.1093/bioin o ma ics/b u140
Gene ics and popula ion analysis Ad ance Access publica ion Ma ch 24, 2014
Assessing mul i a ia e gene-me abolome associa ions wi h a e
a ian s using Bayesian educed ank eg ession
Pekka Ma inen
1,2
, Ma i Pi inen
3
, An i-Pekka Sa in
3,4
, Jussi Gillbe g
1
,
Johannes Ke unen
3,4
, Ida Su akka
3,4
, An i J. Kangas
5
, Pasi Soininen
5,6
, Paul O’Reilly
7
,
Ma ika Kaakinen
8,9
,MikaKa
¨ho
¨nen
10
, Te ho Leh ima
¨ki
11
, Mika Ala-Ko pela
5,6,12
,Olli
T. Rai aka i
13,14
,VeikkoSalomaa
15
, Ma jo-Rii a Ja
¨ elin
7,8,9,16,17
, Samuli Ripa i
3,4,18,19,
*and
Samuel Kaski
1,20,
*
1
Depa men o In o ma ion and Compu e Science, Helsinki Ins i u e o In o ma ion Technology HIIT, Aal o Uni e si y,
Esbo, Finland,
2
Cen e o Communicable Disease Dynamics, Ha a d School o Public Heal h, Bos on, MA, USA
3
Ins i u e o Molecula Medicine Finland (FIMM), Uni e si y o Helsinki,
4
Uni o Public Heal h Genomics, Na ional Ins i u e
o Heal h and Wel a e, Helsinki,
5
Compu a ional Medicine, Ins i u e o Heal h Sciences, Uni e si y o Oulu and Oulu
Uni e si y Hospi al, Oulu,
6
NMR Me abolomics Labo a o y, School o Pha macy, Uni e si y o Eas e n Finland, Kuopio,
Finland,
7
Depa men o Epidemiology and Bios a is ics, MRC Heal h P o ec ion, Agency (HPA) Cen e o En i onmen
and Heal h, School o Public Heal h, Impe ial College, London, UK,
8
Ins i u e o Heal h Sciences,
9
Biocen e Oulu,
Uni e si y o Oulu, Oulu,
10
Depa men o Clinical Physiology, Tampe e Uni e si y Hospi al and Uni e si y o Tampe e,
11
Depa men o Clinical Chemis y, Fimlab Labo a o ies, Uni e si y o Tampe e School o Medicine, Tampe e, Finland,
12
Compu a ional Medicine, School o Social and Communi y Medicine and he Medical Resea ch Council In eg a i e
Epidemiology Uni , Uni e si y o B is ol, B is ol, UK,
13
Depa men o Clinical Physiology and Nuclea Medicine,
14
Resea ch Cen e o Applied and P e en i e Ca dio ascula Medicine, Uni e si y o Tu ku and Tu ku Uni e si y Hospi al,
Tu ku,
15
Depa men o Ch onic Disease P e en ion, Na ional Ins i u e o Heal h and Wel a e, Helsinki,
16
Uni o P ima y
Ca e, Oulu Uni e si y Hospi al,
17
Depa men o Child en and Young People and Families, Na ional Ins i u e o Heal h
and Wel a e, Oulu, Finland,
18
Wellcome T us Sange Ins i u e, Hinx on, Camb idge, UK,
19
Hjel Ins i u e and
20
Depa men o Compu e Science, Helsinki Ins i u e o In o ma ion Technology HIIT, Uni e si y o Helsinki, Helsinki,
Finland
Associa e Edi o : Je ey Ba e
ABSTRACT
Mo i a ion: A ypical genome-wide associa ion s udy sea ches o
associa ions be ween single nucleo ide polymo phisms (SNPs) and a
uni a ia e pheno ype. Howe e , he e is a g owing in e es o in es i-
ga e associa ions be ween genomics da a and mul i a ia e pheno-
ypes, o example, in gene exp ession o me abolomics s udies. A
common app oach is o pe o m a uni a ia e es be ween each geno-
ype–pheno ype pai , and hen o apply a s ingen signi icance cu o
o accoun o he la ge numbe o es s pe o med. Howe e , his
app oach has limi ed abili y o unco e dependencies in ol ing mul-
iple a iables. Ano he end in he cu en gene ics is he in es iga ion
o he impac o a e a ian s on he pheno ype, whe e he s anda d
me hods o en ail owing o lack o powe when he mino allele is
p esen in only a limi ed numbe o indi iduals.
Resul s: We p opose a new s a is ical app oach based on Bayesian
educed ank eg ession o assess he impac o mul iple SNPs on a
high-dimensional pheno ype. Because o he me hod’s abili y o com-
bine in o ma ion o e mul iple SNPs and pheno ypes, i is pa icula ly
sui able o de ec ing associa ions in ol ing a e a ian s. We demon-
s a e he po en ial o ou me hod and compa e i wi h al e na i es
using he No he n Finland Bi h Coho wi h 4702 indi iduals, o
whom genome-wide SNP da a along wi h lipop o ein p o iles comp is-
ing 74 ai s a e a ailable. We disco e ed wo genes (XRCC4 and
MTHFD2L) wi hou p e iously epo ed associa ions, which eplica ed
in a combined analysis o wo addi ional coho s: 2390 indi iduals om
he Ca dio ascula Risk in Young Finns s udy and 3659 indi iduals
om he FINRISK s udy.
A ailabili y and implemen a ion: R-code eely a ailable o down-
load a h p://use s.ics.aal o. i/pema i/gene_me abolome/.
Con ac : samuli. ipa i@helsinki. i; samuel.kaski@aal o. i
Supplemen a y in o ma ion: Supplemen a y da a a e a ailable a
Bioin o ma ics online.
Recei ed on No embe 7, 2013; e ised on Feb ua y 27, 2014;
accep ed on Ma ch 4, 2014
1 INTRODUCTION
Concen a ions o human me aboli es a e associa ed wi h isk o
many common diseases; o example, low- and high-densi y lipo-
p o ein choles e ol (LDL and HDL) le els a e associa ed wi h
co ona y a e y disease. Fo his eason, human me abolism has
been unde in ensi e in es iga ion and o e he pas ew yea s
se e al genome-wide associa ion s udies ha e success ully un-
co e ed a pa o i s gene ic basis (Ke unen e al., 2012;
*To whom co espondence should be add essed.
ßThe Au ho 2014. Published by Ox o d Uni e si y P ess.
This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion Non-Comme cial License (h p://c ea i ecommons.o g/licenses/
by-nc/3.0/), which pe mi s non-comme cial e-use, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly ci ed. Fo comme cial
e-use, please con ac jou nals.pe [email protected]
Saba i e al., 2008; Suh e e al., 2011; Teslo ich e al., 2010). Fo
example, a la ge me a-analysis (Teslo ich e al., 2010) iden i ied
95 loci in luencing he le els o o al choles e ol, LDL, HDL and
iglyce ides. Mo e ecen s udies used ine subclassi ica ions o
me aboli es and disco e ed dozens o no el loci (Ke unen e al.,
2012; Suh e e al., 2011). Despi e hese ad ances, he a iance
explained by all epo ed single nucleo ide polymo phisms
(SNPs) alls a below he sugges ed he i abili y o he common
me aboli es, as es ima ed ei he om win s udies (Ke unen
e al., 2012) o om mo e dis an ly ela ed indi iduals
(Va iku i e al., 2012). This mo i a es us o de elop new
app oaches o associa ion es ing ha could be e use all in o -
ma ion a ailable o us.
The p esen -day coho s udies o en come wi h a ich se o
pheno ypic ea u es. Examples in addi ion o me abolomics
(Ke unen e al., 2012; Soininen e al., 2009; Suh e e al., 2011)
include s udies o gene exp ession (Acke mann e al., 2013) and
3D- acial imaging (Hammond and Su ie, 2012). As a conse-
quence, we need s a is ical me hods ha inc ease he powe o
unco e geno ype–pheno ype dependencies by combining in o -
ma ion o e se e al ela ed pheno ypes (Fe ei a and Pu cell,
2009; Inouye e al., 2012; O’Reilly e al., 2012). The unde lying
idea is ha i a gene ic a ian a ec s a ai , hen i is likely o
a ec o he ai s ha a e ela ed o he i s one and, by es ing
o associa ion wi h he wo ai s join ly, powe may be
inc eased. This easoning can be aken a s ep u he by es ing
all ai s in high-dimensional omics da a simul aneously, o ex-
ample, all me aboli es in comp ehensi e me abolomic p o iles. A
compa ison o di e en s a is ical me hods a ailable o join
es ing o comple e me abolomics p o iles was ecen ly con-
duc ed (Ma inen e al., 2013).
Besides es ing se e al pheno ypes simul aneously, he abili y
o de ec ce ain kinds o associa ions may be boos ed by com-
bining s a is ical e idence o e se e al SNPs. Usually his is done
in a supe ised manne , such ha SNPs ela ed by loca ion o
unc ion, o example, a e es ed simul aneously. Combining
in o ma ion o e mul iple SNPs is pa icula ly c ucial wi h a e
a ian s, i.e. SNPs whe e he mino allele is p esen in a small
p opo ion o he popula ion. Tes ing such SNPs indi idually is
unlikely o yield signi ican indings because o limi ed powe .
Mos app oaches o handling a e a ian s a e based on collap-
sing se e al a e SNPs in o a single a iable (Bansal e al., 2010).
Fo example, one can simply collapse se e al a e a ian s in o a
single indica o , elling whe he any o he a e a ian s is p esen
in he indi idual (Mo gen hale and Thilly, 2007) o o coun he
numbe o a e a ian s p esen in he indi idual (Mo is and
Zeggini, 2010). The p oblem wi h he collapsing me hods is he
implici assump ion ha he e ec s a e in he same (o a p ede-
ined) di ec ion. A mo e sophis ica ed a iance componen
me hod a oiding his assump ion is able o in es iga e he
impac o se e al a e a ian s on a uni a ia e ai (Wu e al.,
2011).
E en i he e exis s a la ge numbe o me hods o analyzing
a e a ian s, none o hose has been ailo ed o mul i a ia e
pheno ypes. On he o he hand, s anda d me hods o mul i a i-
a e pheno ypes (Fe ei a and Pu cell, 2009) ha e no been ho -
oughly in es iga ed in he con ex o a e a ian s, and we will
see la e in he ex ha se e e o e i ing may occu .
In summa y, me hods o dealing wi h a e a ian s in he con-
ex o mul i a ia e pheno ypes a e clea ly lacking.
In his a icle, we de i e a no el o mula ion o he Bayesian
educed ank eg ession model (Geweke, 1996) o de ec mul i-
a ia e associa ions be ween p ede ined g oups o SNPs and a
high-dimensional pheno ype. In pa icula , ou app oach is sui -
able o analyzing bo h common and a e a ian s. Ou o mu-
la ion inco po a es p io knowledge abou e ec sizes o inc ease
he powe o de ec associa ions. Fu he mo e, i is capable o
co ec ing o he numbe o SNPs conside ed, which is impo -
an when es ing a la ge numbe o SNP g oups o di e en sizes.
We alida e ou me hod by assessing associa ions be ween SNPs
in all human genes, one gene a a ime, and me abolic p o iles
comp ising ine-scale lipop o ein measu emen s o 4702 indi id-
uals om he No he n Finland Bi h Coho 1966 (Ran akallio,
1969; Saba i e al., 2008). Among he op-sco ing genes wi hou
known associa ions o he ai s s udied, wo genes (XRCC4 and
MTHFD2L) eplica ed in a combined analysis o 2390 indi id-
uals om he Ca dio ascula Risk in Young Finns s udy (YFS;
Rai aka i e al., 2008) and 3659 indi iduals om he FINRISK
s udy (Va iainen e al., 2010). Addi ional analyses o he same
da a con i med ha al e na i e me hods disco e ed only one o
hese associa ions and did no iden i y any u he associa ions
ha we e no known be o e.
2 METHODS
2.1 Model
To build ou model, we assume ha he pheno ypes may be a ec ed by
h ee kinds o a iables, (i) known ac o s, such as age, sex o popula ion
s uc u e, (ii) unknown ac o s, such as expe imen al condi ions and o he
ba ch e ec s, and (iii) SNPs unde conside a ion, as schema ically p e-
sen ed in Figu e 1. In con as o s anda d eg ession, whe e each SNP–
pheno ype pai has a pa ame e ep esen ing he e ec o he SNP on he
pheno ype, he e we assume ha a combina ion o se e al SNPs is in lu-
encing se e al pheno ypes h ough some unknown ac o s. This assump-
ion is compac ly exp essed in e ms o he educed ank eg ession,
whe e he SNPs a e i s p ojec ed on o a low-dimensional subspace,
and he p ojec ions a e hen used as eg esso s when p edic ing he
pheno ypes. The educed ank eg ession o mula ion immedia ely
implies some s uc u al assump ions deemed sensible in he cu en se -
ing: i s , i a SNP has an e ec on a pheno ype, hen he SNP is likely o
ha e an e ec on o he pheno ypes; second, i a pheno ype is a ec ed by a
SNP, hen he pheno ype is mo e likely o be a ec ed by o he ela ed
SNPs as well.
Le Ndeno e he numbe o indi iduals, S he numbe o SNPs, P he
numbe o pheno ypes and C he numbe o o he co a ia es. Fo mally,
we conside he Bayesian educed ank eg ession model
Y¼X þZA þHTþEð1Þ
whe e YNPcon ains he pheno ypes, XNScon ains he SNPs, SK1
and K1P ep esen a low- ank app oxima ion o he eg ession coe i-
cien ma ix ¼,ZNC ep esen s o he co a ia es wi h he co es-
ponding coe icien ma ix ACP,HNK2con ains hidden con ounding
ac o s wi h he co esponding coe icien ma ix PK2and
ENP¼½e1,...,eNT,wi heiNð0, Þ,whe e¼diagð2
1,...,2
PÞ.
No e ha by in eg a ing o e he hidden ac o s H, he model is equi a-
len o
yiNðTxiþATzi,TþÞ,i¼1, ...,Nð2Þ
2027
Gene-me abolome associa ions
whe e i is assumed ha he ac o s ha e independen s anda d no mal
p io dis ibu ions. The e o e, we see ha ha ing he la en a iable pa
HTin he model co esponds o assuming a low- ank app oxima ion o
he ull co a iance ma ix, which is impo an when analyzing high-
dimensional da ase s.
Fo compu a ional easons, we es ic he ank o he eg ession
model, K
1
, o uni y in ou genome-wide analysis (w.l.o.g. in he de ec ion
ask, see below) and show esul s wi h K1¼1, 2, 3 o a ew ep esen a-
i e examples. Howe e , he e we p esen a gene al in ini e-dimensional
amewo k a ailable in ou implemen a ion, which does no necessi a e
he selec ion o a ixed ank. In gene al, o use he Bayesian educed ank
eg ession model, he ank o he model, K
1
, and he ank o he low- ank
app oxima ion o he co a iance ma ix, K
2
, mus be selec ed. A ecen
Bayesian in ini e spa se ac o analysis model (Bha acha ya and
Dunson, 2011) ci cum en s he selec ion o a ixed ank o K
2
by assum-
ing in p inciple an in ini e numbe o columns in he ma ix; howe e ,
he columns sh ink p og essi ely such ha only he i s anks a e in lu-
en ial in p ac ice. We assume his p io o mula ion o ou noise model
HTþE. We exploi he idea u he by allowing also he ank K
1
o be
in ini e in p inciple, and en o cing he low- ank na u e by sh inking he
columns o and he ows o inc easingly as he column/ ow index
g ows. In p ac ice, one needs o speci y uppe bounds o K
1
and K
2
.We
selec he uppe bound o K
2
using he adap i e p ocedu e o
Bha acha ya and Dunson (2011), and we ha e implemen ed an analo-
gous me hod o lea ning he uppe bound o K
1
. In p ac ice, we wan ed
o minimize he model complexi y o maximize he powe o de ec asso-
cia ions and speed up he compu a ions. The e o e, we decided o un he
genome-wide analysis (see Sec ion 3.1) using a ixed ank K1¼1. The
model wi h K1¼1 is su icien o ou pu poses o de ec ing whe he a
gene is un ela ed o he pheno ypes (in which case, ank ze o would
al eady be su icien ), al hough o p edic ion a highe ank migh be
mo e sui able. We expe imen ed wi h a selec ed se o known genes
wi h uppe bounds 2 and 3 (see below) and no iced ha he e ec o
inc easing he uppe bound om uni y had only a small impac on he
amoun o a ia ion ha is explained by he model. F om he biological
pe spec i e, his means ha he in luence o a gene on he pheno ypes can
mos ly be desc ibed in e ms o a single la en ac o media ing he e ec .
A de ailed model desc ip ion is gi en in Supplemen a y Sec ion 1.
The model is ela ed o many published me hods, and a ho ough
compa ison can be ound in Supplemen a y Sec ion 2. Co ec ion o
unknown ac o s has ecen ly been conside ed wi h he s anda d eg es-
sion model, ypically o jus one SNP a a ime (Fusi e al., 2012; S egle
e al., 2010). The di e ence o he o iginal Bayesian educed ank eg es-
sion o mula ion (Geweke, 1996) is ha ou model uses low- ank ap-
p oxima ion o he co a iance ma ix, making i mo e sui able o
high-dimensional pheno ypes, and in o ma i e p io dis ibu ions accom-
moda ing p oblem-speci ic knowledge. Fu he mo e, we use he model in
a new way, as desc ibed in he subsequen sec ions. The Bayesian in ini e
spa se ac o analysis model has been used o high-dimensional da a
(Bha acha ya and Dunson, 2011); he e, we use i o ep esen he mul i-
a ia e noise. Canonical co ela ion analysis (CCA) is a classical ool o
modeling mul i a ia e dependencies ha ha e ecen ly been in oduced in
he associa ion s udy con ex (Fe ei a and Pu cell, 2009; Ho elling,
1936). Spa se CCA (Pa khomenko e al., 2009; Waaijenbo g e al.,
2008; Wi en and Tibshi ani, 2009) is mo e sui able o high-dimensional
da ase s; howe e , in oducing p io dis ibu ions o CCA ha would be
in ui i e in he associa ion s udy con ex does no seem s aigh o wa d.
2.2 P opo ion o o al a ia ion explained
A commonly used measu e o he impac o mul iple SNPs, say
x1,...,xS, on a uni a ia e ai yis he p opo ion o a iance explained
(PVE) by he SNPs:
PVE ¼1Va ðyPS
i¼1^
aixiÞ
Va ðyÞ
¼Va ðPS
i¼1^
aixiÞ
Va ðyÞ:
He e, PS
i¼1^
aixiis a linea p edic ion o he pheno ype y, gi en he SNPs.
Analogously, as a measu e o he o e all impac o mul iple SNPs on a
high-dimensional pheno ype, we p opose o use he p opo ion o o al
a ia ion o he pheno ypes explained (PTVE) by he model, namely
PTVE ¼T ðCo ð^
YÞÞ
T ðCo ðYÞÞ ð3Þ
whe e ^
Yis a p edic ion o a high-dimensional pheno ype Y om he
model and T deno es he ace, i.e. he sum o he diagonal elemen s o
he ma ix. In (3) and in gene al, he o al a ia ion o a mul i a ia e
andom a iable is de ined as he ace o he co a iance ma ix, i.e. he
sum o he a iances o he indi idual a iables. The e o e, PTVE meas-
u es he join impac o he SNPs on se e al pheno ypes and hence
is expec ed o yield high sco es o such dependencies in which many
pheno ypes a e a ec ed by he SNPs, e en i none o he e ec s is la ge
by i sel .
In he Bayesian s a is ical amewo k, he in e ences a e based on pos-
e io p obabili y dis ibu ions o he quan i ies o in e es (Gelman e al.,
2004). Wi h he Bayesian educed ank eg ession model, samples om
he pos e io dis ibu ion o he PTVE can be ob ained om
PTVEðiÞ¼T Co XðiÞðiÞ
T Co YðÞðÞ ð4Þ
whe e ðiÞand ðiÞa e samples om he pos e io dis ibu ion o he
pa ame e s and . I is simila ly s aigh o wa d o es ima e he pos-
e io dis ibu ion o he p opo ion o a ia ion explained by he a e
Y X
ψΓ
ZA H
T
E,
Y
X
Z
A
H
T
E
A
B
Fig. 1. G aphical illus a ion o he model. (A) The a iables and depen-
dencies be ween hem. The pheno ypes Ya e assumed o be a ec ed by
known ac o s, such as age o sex, unknown ac o s, such as ba ch e ec s
caused by a ying expe imen al condi ions, and he SNPs. The in luence
o he SNPs is media ed by unknown combina ions o he o iginal SNPs,
ep esen ed by black squa es. (B) The same model using ma ix no a ion.
Ma ices con aining he obse ed a iables, Y( he pheno ypes), X( he
SNPs) and Z(known ac o s), a e blue. The eg ession coe icien ma i-
ces a e ed. No e ha he coe icien ma ix o he SNP e ec s is w i en
as a p oduc o wo ma ices, and , co esponding o a low- ank
app oxima ion o an uncons ained coe icien ma ix. The b own ma i-
ces comp ise unobse ed a iables, H(unknown ac o s) and E(noise
e ms)
2028
P.Ma inen e al.
a ian s, by di iding he a ia ion o he p edic ion in o wo componen s,
one co esponding o he a e a ian s, he o he o he common a ian s.
The sco e ob ained in his way is e e ed o as PTVE- a e in he sequel
and i s exac de ini ion is gi en in Supplemen a y Sec ion 3. In p ac ice,
we app oxima e he pos e io dis ibu ion by using a mode-based poin
es ima e o one o he pa ame e s, , and es ima ing he join dis ibu-
ion o he o he pa ame e s using an Ma ko chain Mon e Ca lo
(MCMC) algo i hm, as desc ibed ho oughly in Supplemen a y
Sec ions 4 and 5.
2.3 In o ma i e p io
In he Bayesian analysis, ex e nal backg ound knowledge may be inco -
po a ed in he s a is ical analysis h ough p io p obabili y dis ibu ions.
As ou analysis is ocused on es ima ing he pos e io dis ibu ion o
PTVE, a sensible p io is ob ained by making he p io dis ibu ion o
he PTVE ep esen ou ue belie s abou his quan i y. The p io dis i-
bu ion o he PTVE appea s no o be a ailable in a closed o m; how-
e e , in Supplemen a y Sec ion 6, we de i e esul s ha show how he
dis ibu ion o he mean o he PTVE, PTVE, depends on model hype -
pa ame e s. Using hese esul s, we se he p io dis ibu ions o sa is y
he ollowing p ope ies:
MedianðPTVEÞ¼106ð5Þ
and
PðPTVE40:001Þ¼0:01 ð6Þ
Equa ion (5) means ha he p io median o PTVE is close o ze o, as
we expec mos o he genes o be un ela ed o he pheno ypes. Equa ion
(6) says ha wi h a small p obabili y, he e equal o 0.01, he gene may
explain40.1% o he o al a ia ion, a alue deemed esonable based on
he backg ound knowledge. The esul ing p io dis ibu ion o he PTVE
is ob ained by in eg a ing o e he dis ib ion o PTVE,andweused
Mon e Ca lo simula ion o in es iga e he dis ibu ion. Figu e 2 shows
PTVE alues sampled om he p io dis ibu ion. The ollowing desi able
cha ac e is ics can be seen: i s , a peak close o ze o, sh inking he coe -
icien s when no e ec is p esen ; second, a long ail, co esponding o he
genes wi h a non-negligible impac on he pheno ypes, wi hou imposing
s ong belie s abou he ac ual size o hese non-ze o e ec s.
Ano he impo an p ope y ha ollows om using he in o ma i e
p io is ha he numbe o SNPs unde conside a ion can be accoun ed
o . De ailed examina ion o Co olla y 1 in Supplemen a y Sec ion 6
e eals ha wi h ixed hype pa ame e s, he expec ed PTVE is p opo -
ional o PS
i¼1Va ðxiÞ, i.e. he o al a ia ion o he SNPs. Roughly, his
means ha i he numbe o SNPs doubles, he expec ed PTVE doubles as
well, i he hype pa ame e s a e kep ixed. Using he Co olla ies 1 and 2
in Supplemen a y Sec ion 6, i is s aigh o wa d o modi y he hype -
pa ame e dis ibu ions o asse he p io condi ions (5) and (6), implying
in ou se ing ha all genes a e expec ed o explain he same amoun o
he a ia ion o he pheno ypes, ega dless o how many SNPs hey
con ain.
2.4 Da a
As a da ase o de ec ing associa ions, we conside a sample o 4702
indi iduals om he No he n Finland Bi h Coho 1966
(NFBC1966), a bi h coho s udy o child en bo n in 1966 in he wo
no he nmos p o inces o Finland (Ran akallio, 1969). The blood sam-
ples o he DNA ex ac ion and pheno ype da a we e collec ed a a
ollow-up isi when he pa icipan s we e 31 yea s o age. Fo eplica-
ion, we conside wo coho s. The YFS is a popula ion-based p ospec i e
coho s udy (Rai aka i e al., 2008) conduc ed in Finland, he pu pose o
which was o in es iga e he le els o ca dio ascula isk ac o s in chil-
d en and adolescen s in di e en pa s o he coun y. The FINRISK
s udy comp ises c oss-sec ional popula ion su eys ha ha e been ca ied
ou e e y 5 yea s since 1972, o assess he isk ac o s o ch onic diseases,
wi h emphasis on ca dio ascula isk ac o s (Va iainen e al., 2010). The
indi iduals analyzed in his s udy belong o he sample om he yea
1997. The blood samples o he wo eplica ion coho s we e collec ed
when he pa icipan s we e 30–45 and 25–71 yea s o age, espec i ely.
The s udy p o ocols o all da ase s ha e been app o ed by he local e hics
commi ees.
The samples we e geno yped wi h Illumina a ays (Illumina, Inc. San
Diego, CA, USA) and impu ed wi h IMPUTE 2 (Howie e al.,2009,
2011) using a 1000 Genomes P ojec e e ence panel (The 1000
Genomes P ojec Conso ium, 2012). O he esul ing good-quali y au o-
somal SNPs (in o 40.4), we ex ac ed SNPs in 24025 human genes by
adding 50 kb lanking egions on bo h sides o he endpoin s o he genes
gi en in NCBI gene da abase (genome assembly GRCh37.p10, NCBI
anno a ion 104, No embe 2012). As a p ep ocessing s ep, we educed
he geno ype space wi hin each gene o he mos p omising 200 SNPs (a
mos ) ha had he highes canonical co ela ion es sco e (Fe ei a and
Pu cell, 2009) wi h he me aboli es; howe e , o p e en o e i ing, he
p io s we e speci ied as desc ibed abo e using he unp uned SNP se . In
p elimina y expe imen s, dec easing he numbe o SNPs in his way om
800 o 200 had no isible e ec on he esul s. Finally, he SNPs we e
scaled o ha e uni a iance.
Pheno ype da a came om he se um NMR me abolomics pla o m
desc ibed ea lie (Soininen e al., 2009). As a p ep ocessing s ep, he ai s
we e quan ile-no malized o ha e s anda d no mal dis ibu ion.
Indi iduals wi h 20% missing alues we e emo ed and he emaining
missing alues we e impu ed by sampling hem om he mul i a ia e
no mal dis ibu ion. In his wo k, we analyzed a subse o 74 lipop o ein
subclass measu es (Supplemen a y Table S3). The empi ical co ela ion
ma ix o he ai s is shown in Supplemen a y Figu e S1. Linea eg es-
sion was used o co ec he pheno ypes o age, sex and popula ion
s uc u e using 10 p incipal componen s (P ice e al., 2006).
P io dis ibu ion o PTVE
PTVE (x10 )
P obabili y
0246810
0 0.3 0.6 0.9
0.000 0.005 0.010 0.015
0 1000 2000 3000 4000 5000
Pos e io PTVE o LIPC gene
PTVE
Densi y
O iginal
Pe mu ed
Fig. 2. P io and pos e io dis ibu ions o he p opo ion o o al a i-
a ion explained (PTVE) by he model. The panel on he le shows he
p io dis ibu ion imposed on he p opo ion o o al a ia ion o he
pheno ypes explained by he SNPs unde conside a ion (he e, he SNPs
om he LIPC gene). The median o he p io dis ibu ion is loca ed a
4e-6. The cha ac e is ic ea u es o he p io dis ibu ion include he
peak a alues close o ze o, e ec i ely emo ing noise unless he e is
s ong e idence abou a possible associa ion, and he long ail allowing a
small pe cen age o genes o explain la ge p opo ions o he pheno ype
a ia ion. The igh mos bin on he x-axis con ains he o al p obabili y
o alues exceeding he maximum alue on he axis. The panel on he
igh shows he pos e io dis ibu ion o he PTVE o he same SNPs.
Two pos e io densi ies a e shown, one showing he dis ibu ion o he
o iginal da a, he o he showing he dis ibu ion o da a in which he
ows o he pheno ype ma ix ha e been pe mu ed. No ice he di e ing
scales on he x-axes o he wo panels
2029
Gene-me abolome associa ions
3RESULTS
3.1 Genome-wide analysis o NFBC1966 da a
We compu ed he PTVE and PTVE- a e sco es o each o he
24 025 human genes in he NFBC1966 da a, one gene a a ime,
using he Bayesian educed ank eg ession. Table 1(a) shows he
op i e genes wi h he highes PTVE sco es, all o which a e
well-known lipid-associa ed loci. Mo e gene ally, all 43 op-sco -
ing genes we e loca ed wi hin 1Mb om p e iously epo ed
genome-wide signi ican lipid associa ions (Ke unen e al.,
2012; Teslo ich e al., 2010; The Global Lipids Gene ics
Conso ium, 2013), and de ailed lis ing o he op genes is
gi en in Supplemen a y Table S4. Fu he mo e, o he op 100
genes, which co espond o 0.2 alse disco e y a e (FDR) (see
he nex sec ion), 73 genes had known associa ions. Toge he ,
hese indings se e as a alida ion o he me hod and i s
implemen a ion.
To illus a e he e ec o he model ank on he esul s, we
ca ied ou a de ailed analysis o h ee known lipid genes: LIPC,
APOB and PLTP,wi h ankK1¼1, 2, 3. The genes we e se-
lec ed such ha hey we e loca ed in di e en ch omosomes
and had dissimila associa ion p o iles [APOB associa ed mos
s ongly o VLDL, IDL and LDL, PLTP o HDL and LIPC o
VLDL, IDL and HDL (Tukiainen e al., 2012)]. In addi ion o
conside ing he genes sepa a ely, we epea ed a join analysis o
all h ee possible pai wise gene combina ions. The esul s a e
shown in Supplemen a y Figu e S2 and a e summa ized as ol-
lows: (i) he ank o he model had a mino e ec on he esul s
o a single gene, (ii) inc easing he ank om K1¼1 oK1¼2
inc eased he a iance explained in all join pai wise analyses,
and he inc ease was signi ican in wo o he h ee cases, (iii) a
combina ion o wo genes always had a la ge e ec han ei he
gene indi idually; howe e , he sum o he indi idual e ec s was
sligh ly la ge han he e ec o he combina ion. We a ibu e
his di e ence mainly o he s onge sh inkage in he second
han he i s geno ype componen . (i ) The i s componen o
a join model was always s ongly co ela ed wi h he single-gene
model wi h he la ge e ec , he second componen wi h he
single-gene model wi h he smalle e ec . This is as expec ed,
as he p io dis ibu ion was designed o iden i y he s onges
associa ions using he i s geno ype componen .
To check whe he no el associa ions could be de ec ed by any
me hod, we ca ied ou a eplica ion expe imen wi h he YFS
and FINRISK da ase s o he mos p omising genes, a e
excluding genes loca ed wi hin 1 Mb om p e iously epo ed
associa ions. We conside ed om each me hod all genes wi h
FDR 50.4 as p omising. Fo he PTVE- a e sco e, six genes
we e es ed, none o which had p e iously epo ed associa ions
o lipids. Fo he PTVE sco e, 305 genes had FDR 50.4; how-
e e , only 167 we e no loca ed close o known associa ions, and
hese 167 genes we e selec ed o eplica ion.
Fo he pu poses o eplica ion, he mul i a ia e dependency
be ween he mul iple SNPs and pheno ypes was educed in o a
uni a ia e es by using pa ame e s es ima ed in he NFBC1966
da a o cons uc uni a ia e geno ype and pheno ype combin-
a ions (see Supplemen a y Sec ion 7 o u he de ails). Then,
he s anda d linea model was used o es o posi i e co ela ion
be ween he geno ype and pheno ype combina ions. A pooled
linea eg ession coe icien combining YFS and FINRISK
da ase s was o med using a ixed e ec model o e he wo
da ase s (Thompson e al., 2011), and he P- alue was ob ained
by ela ing he pooled es ima e o i s SD. A one- ailed es was
used,as we we e only in e es ed in indings in which he e ec s
we e in he same di ec ion in he eplica ion da ase s as in he
NFBC1966 da a. Bon e oni co ec ion was used o accoun o
he numbe o genes es ed wi h each me hod, such ha co ec ed
Table 1. Summa y o esul s om he genome-wide analysis o he eal da a
Ch Locus PTVE (SD) Ra e P- alue Gene ank
PTVE Pai wise S-CCA CCA-single
(a)
15 LIPC 0.01 (5e-04) 0.015 5e-19 1 1 1 132
19 APOC1 0.0046 (3e-04) 0.017 1.4e-26 2 8 18 275
19 PVRL2 0.0045 (1e-04) 0.0015 7e-35 3 9 14 276
2APOB 0.0044 (3e-04) 0.017 1.1e-17 4 45 41 3244
11 APOA5 0.0043 (2e-04) 0.021 5e-10 5 26 33 2433
(b)
16 SPIRE2
a
0.0015 (9e-05) 0.89 0.00091
b
5 ( a e) 8344 5490 4973
5XRCC4 0.0024 (2e-04) 0.55 0.0016 6 ( a e) 2706 5163 2155
(c)
2DTNB
a
0.0015 (2e-04) 0.15 2.6e-04
c
138 4715 338 7652
4MTHFD2L 0.0015 (2e-04) 0.019 7e-06 163 102 1444 1552
No e: (a) Re e ence esul s o genes wi h i e highes PTVE sco es. (b) Replica ed genes om he PTVE- a e sco e (o six genes es ed o eplica ion). (c) Replica ed genes
om he PTVE sco e (o 167 genes). Fi e o he eplica ed genes (PPBP, CXCL5, CXCL2, PF4 and CXCL3) a e no shown as hey we e loca ed wi hin 1 Mb om
MTHFD2L, which had he s onges e ec . Column PTVE (SD) shows he p opo ion o o al a ia ion explained and i s SD, a e speci ies he p opo ion o he a ia ion
explained by he gene a ibu ed o he a e a ian s, P- alue speci ies he P- alue pooled o e YFS and FINRISK eplica ion da ase s (unless s a ed o he wise) and he las
ou columns speci y he anking o he gene among all genes wi h di e en me hods.
a
Deno es genes ha eplica ed signi ican ly in only one o he wo eplica ion da ase s.
b/c
Replica ion P- alue in FINRISK/YFS.
2030
P.Ma inen e al.
P- alue h eshold co esponding o he nominal 0.05 le el o
signi icance was equal o P50:0083 o he six pu a i e no el
genes om PTVE- a e sco e and P50:00030 o he 167 pu a i e
genes om PTVE sco e.
We conside ed a eplica ion signi ican i he es was nomin-
ally signi ican (P50.05) in bo h eplica ion da ase s, and he
P- alue o he pooled es ima e was signi ican a e co ec ing
o he mul iple es s. In o ma ion abou genes ha eplica ed
signi ican ly is p o ided in Table 1(b) and (c) and in Figu e 3.
One gene wi h PTVE- a e and six genes wi h PTVE we e de-
ec ed, co esponding o wo independen genes: XRCC4 and
MTHFD2L.MTHFD2L is loca ed wi hin 1 Mb om wo
SNPs ( s2168889 and s16850360) associa ed wi h ‘me abolic
ne wo ks’ con aining some lipop o ein ai s om ou da a
(Inouye e al., 2012). Howe e , he op me aboli e in he associ-
a ions was Albumin, and when we epea ed he mul i a ia e es
wi h he lipop o ein ai s conside ed he e, hese associa ions
we e no longe genome-wide signi ican (P¼2.5e-4 and
P¼6.4e-7). In addi ion, wo genes, SPIRE2 and DTNB, epli-
ca ed signi ican ly (a e he mul iple es ing co ec ion) in one,
bu no in he o he eplica ion da ase . Table 1 also p esen s he
ankings o hese genes by al e na i e me hods (see below). We
see ha XRCC4,SPIRE2 and DTNB we e comple ely missed by
he o he me hods. Supplemen a y Table S1 shows he SNPs
con ibu ing o he epo ed associa ions. We see ha wi h
XRCC4,SPIRE2 and DTNB, many SNPs, some o which a e
a e, a e equi ed o ep esen he o e all associa ion. This ex-
plains why hese genes did no ecei e any signal om he s and-
a d es ing wi h he pai wise linea model. Supplemen a y Figu e
S3 shows g aphically how he pheno ypes a e a ec ed by he
signi ican genes and he known LIPC gene. We see ha in nei-
he o he new genes is he e ec ocused on any single ai , bu
a he a small e ec is seen on many lipop o ein measu es. This is
no su p ising, as he PTVE sco e is expec ed o gi e high sco es
o p ecisely his kind o associa ion. Supplemen a y Figu e S4
shows he es ima ed SNP coe icien s o he genes and demon-
s a es he use ulness o analyzing all SNPs in a gene simul an-
eously o educe noise esul ing om he co ela ion be ween he
SNPs. Fu he backg ound in o ma ion on hese genes is p e-
sen ed in Supplemen a y Table S2; howe e , a mo e ho ough
biological in e p e a ion o he genes emains o u u e wo k.
Finally, we epea ed a simila analysis wi h h ee al e na i e
me hods: (i) exhaus i e pai wise sea ch wi h a linea model,
whe e he minus loga i hm o he smalles pai wise P- alue
o e all SNPs in he gene and me aboli es was aken as he es
sco e o he gene, (ii) CCA, applied o a single SNP e sus all
me aboli es a a ime as by Fe ei a and Pu cell (2009) and
Inouye e al. (2012) and (iii) spa se CCA, which was used o
compu e he canonical co ela ion be ween all SNPs in a gene
and all pheno ypes (Pa khomenko e al., 2009). The me hods (ii)
and (iii) we e ound o be he mos powe ul in a ecen com-
pa ison o app oaches o a mul i a ia e me abolomics
Fig. 3. Resul s o genes wi h signi ican eplica ion in bo h es se s: XRCC4 and MTHFD2L; o e e ence, he well-known LIPC lipid locus is also
shown. Each panel shows he iden i ied pheno ype combina ion plo ed agains he geno ype combina ion. The le column shows esul s in he
NFBC1966 da ase , in which he associa ions we e de ec ed. The cen e and igh columns show esul s wi h he YFS and FINRISK da ase s, whe e
coe icien ma ices lea ned wi h he NFBC1966 da a we e used o o m he a iable combina ions. The g een and ed backg ound colo ings ma k he
indi iduals wi h he highes /lowes geno ype combina ion alues, and he pheno ype alues in hese ex eme g oups a e in es iga ed in mo e de ail in
Supplemen a y Figu e S3
2031
Gene-me abolome associa ions
pheno ype (Ma inen e al., 2013); howe e , he e he SNPs we e
no p uned using he mino allele equency be o e applying he
me hods as by Ma inen e al. (2013) because he e he ocus is in
he a e a ian se ing. Simila ly o PTVE and PTVE- a e me h-
ods, we ied o eplica e he op-sco ing genes up o FDR ¼0.4.
The las column in Table 2 shows he numbe s o genes included
in he eplica ion wi h de ailed lis ings gi en in Supplemen a y
Tables S4–S8. The associa ion in MTHFD2L was con i med o
be genome-wide signi ican using he s anda d pai wise linea
eg ession (SNP s185567543, ai S.LDL.P, P¼2.5e-15,
whe e he P- alue is pooled o e all h ee da ase s). All o he
eplica ed associa ions we e loca ed wi hin 1 Mb o his o p e-
iously known associa ions, hus yielding no addi ional no el
de ec ions.
3.2 Powe compa ison
To in es iga e he powe o he in oduced me hod and compa e
i wi h al e na i e me hods, we es ima ed FDR co esponding o
di e en h esholds d o decla ing a gene as de ec ed. FDR es-
ima es we e ob ained by pe mu ing he ows o he pheno ype
ma ix and analyzing he pe mu ed da a in exac ly he same way
as he o iginal da a (Benjamini and Hochbe g, 1995; S o ey and
Tibshi ani, 2003; Xie e al., 2005). In de ail, we compu ed he
a e age numbe o genes in he pe mu ed da ase s wi h sco es
exceeding a gi en h eshold d( alse-posi i e a es, FPðdÞ), he
numbe o genes in he o iginal da a wi h sco es exceeding
hesame h esholdd( o al posi i e a es, TPðdÞ)andconside ed
he a io o he wo
FDRðdÞ¼FPðdÞ
TPðdÞ
Because o compu a ional bu den, only one pe mu a ion was
used wi h he Bayesian educed ank eg ession. Wi h he o he
me hods, h ee pe mu a ions we e used. The e o e, he FDR es-
ima es a e app oxima e, which we conside o be su icien o
ou pu poses, especially because he mos ex eme quan iles a e
no conside ed.
The numbe s o genes decla ed de ec ed wi h di e en FDR
h esholds a e p esen ed in Table 2, and he o e all conco dance
o he sco es om di e en me hods is shown in Supplemen a y
Figu e S5. In summa y, he esul s show ha he me hod wi h
which mos associa ions we e ound was exhaus i e pai wise
sea ch and he second-mos powe ul me hod was he Bayesian
educed ank eg ession wi h PTVE as he es sco e. These wo
me hods also had he highes ag eemen in sco ing genes. CCA
applied o es o associa ion be ween indi idual SNPs and he
mul i a ia e pheno ype pe o med badly. These esul s a e some-
wha di e en om wha has been epo ed be o e by us and
o he s (Inouye e al., 2012; Ma inen e al., 2013). In pa icula ,
he CCA applied o indi idual SNPs pe o med much wo se ha
he pai wise es ing, al hough p e iously i has been epo ed o
ha e clea ly highe powe . The explana ion is ha he e we
applied he me hods o all SNPs, including he a e a ian s.
Speci ically, he single-SNP-CCA s a s o o e i when applied
o a e a ian s. Fo example, when we in es iga ed he esul s
mo e ca e ully, we disco e ed SNPs in which he mino allele was
p esen in ew indi iduals, and such indi iduals could be almos
pe ec ly iden i ied by a seemingly andom combina ion o
pheno ypes, leading o a la ge spu ious canonical co ela ion
es sco e (o , equi alen ly, a highly signi ican P- alue).
Bayesian educed ank eg ession and spa se CCA we e less a -
ec ed by he o e i ing because hey exploi ways o con ol he
complexi y o he model, he o me using he in o ma i e p io s,
he la e using he c oss- alida ion.
Combining in o ma ion o e se e al SNPs using CCA o
ela ed me hods has ecen ly been demons a ed o imp o e
powe o de ec associa ions unde ce ain condi ions
(Ma inen e al., 2013; Tang and Fe ei a, 2012; Zhang e al.,
2011). He e we see ha he spa se CCA and also he Bayesian
educed ank eg ession, which use mul i-SNP in o ma ion, ha e
lowe powe in he genome-wide analysis han he simple pai -
wise es ing. The di e ence om he ea lie expe imen s is ha
he e he me hods a e applied o geno ype da a wi h a much
highe SNP densi y, as ob ained h ough ca e ul impu a ion.
The conclusion is ha i he mul i-SNP in o ma ion has al eady
been used wi hin he impu a ion p o ocol, he powe o de ec
associa ions canno , in gene al, be expec ed o imp o e by using
mul i-SNP models.
Fo a mo e de ailed compa ison be ween he Bayesian educed
ank eg ession and he exhaus i e pai wise es ing,
Supplemen a y Figu e S6 shows a Q-Q plo o he sco es (bo h
PTVE and PTVE- a e) agains he expec ed sco es ha ha e
been ob ained by pe mu a ion. The Supplemen a y Figu e S6
also shows genes de ec able by he simple exhaus i e pai wise
sea ch. Bo h Q-Q plo s, bu PTVE in pa icula , indica e an
excess o la ge es sco es, e lec ing he ac ha a leas some
ue associa ions a e de ec ed by he model. We u he see ha
al hough wi h PTVE- a e sco e ewe genes can be de ec ed han
wi h PTVE, none o he op-sco ing genes om PTVE- a e a e
lagged by he pai wise app oach, making PTVE- a e an a ac -
i e sco e o mining associa ions ha migh be missed by he
s anda d me hod.
4 DISCUSSION
We ha e p esen ed a new s a is ical me hod o in es iga ing
associa ions in Genome-wide associa ion s udy (GWAS) da ase s
wi h mul i a ia e pheno ypes. The me hod can combine in o -
ma ion o e mul iple SNPs, making i pa icula ly sui able o
s udying a e a ian s in he high-dimensional pheno ype se ing.
Fo his se up, no me hods known o he au ho s ha e been
Table 2. Powe compa ison o he di e en me hods
Me hod FDR ¼0FDR¼0.1 FDR ¼0.2 FDR ¼0.4 No el
PTVE 36 55 103 305 167
PTVE- a e33366
Pai wise 103 176 243 651 300
CCA,singleSNP77 71111
Spa se CCA 37 50 66 117 51
No e: The able shows he numbe s o gene-me abolome associa ions ha had alse
disco e y a e below he speci ied h eshold. The las column shows he numbe o
pu a i e no el associa ions wi hin genes wi h FDR ¼0.4 a e emo ing he known
associa ions as desc ibed in he main ex .
2032
P.Ma inen e al.
p esen ed be o e. Ou me hod is based on es ima ing he p opo -
ion o o al a iance o he pheno ypes ha is explained by he
SNPs unde conside a ion. Fo his pu pose, we ha e de i ed a
Bayesian o mula ion o he educed ank eg ession model,
which enables us o inco po a e ou knowledge o he expec ed
e ec s sizes in he analysis.
We used he new me hod o analyze a eal GWAS da ase wi h
a mul i a ia e lipop o ein pheno ype. Two no el loci no p e i-
ously associa ed wi h he pheno ype we e disco e ed and epli-
ca ed in an analysis combining wo addi ional da ase s.
Fu he mo e, wo mo e loci we e ound ha eplica ed signi i-
can ly in one bu no in he o he es da ase . Possible easons
o he lack o success in eplica ing he indings in bo h he
da ase s include he ollowing: (i) he associa ions we e alse posi-
i e in he i s place, (ii) he associa ions in ol ed a e a ian s
ha we e no p esen in su icien numbe s o see he e ec s, (iii)
he impu a ion accu acy o he a e a ian s was no su icien in
all da ase s, (i ) he pheno ype da a, al hough p ep ocessed in
exac ly he same way wi h all he da ase s, ha e no been ully
equi alen . Fo example, he scaling o he pheno ypes du ing
p ep ocessing has been done using ac o s no exac ly equal, and
he pa ame e s lea ned in one da a may hus no ep esen he
e ec s adequa ely in ano he da a. Only u he s udies will help
o dis inguish be ween he al e na i e explana ions.
Fo doing in e ence wi h he model, he cu en implemen a-
ion uses MCMC sampling, he compu a ion ime o which is
app oxima ely hal an hou pe gene on a 2.3 GHz p ocesso .
Thus, analyzing all human genes equi es a clus e compu e o
pa allelize he compu a ions o e he genes. Analy ical app oxi-
ma ions, such as he a ia ional o Laplace app oxima ions,
e.g. Bishop (2006), could be used o speed up he compu a ions
and, based on ou expe imen s wi h he cu en me hod, a e
wo h doing in he u u e. Al e na i e ways o use he model
migh also be conside ed. Fo example, ocusing he analysis
on a ian s wi h a p edic ed unc ion could imp o e powe o
de ec associa ions and lessen he compu a ional bu den. As an-
o he example, we ha e used 0.01 as he h eshold o de ining
he a e a ian s when compu ing hei impac on he pheno-
ypes. Resul s based on di e en h esholds could eadily be ex-
ac ed om he ou pu o a single MCMC un and a e likely o
highligh di e en se s o genes.
Funding: A comple e lis o unding is gi en in he Supplemen a y
Ma e ial.
Con lic o In e es : none decla ed.
REFERENCES
Acke mann,M. e al. (2013) Impac o na u al gene ic a ia ion on gene exp ession
dynamics. PLoS Gene .,9, e1003514.
Bansal,V. e al. (2010) S a is ical analysis s a egies o associa ion s udies in ol ing
a e a ian s. Na . Re . Gene .,11, 773–785.
Benjamini,Y. and Hochbe g,Y. (1995) Con olling he alse disco e y a e: a p ac-
ical and powe ul app oach o mul iple es ing. J. R. S a . Soc. B Me hodol.,57,
289–300.
Bha acha ya,A. and Dunson,D. (2011) Spa se Bayesian in ini e ac o models.
Biome ika,98, 291–306.
Bishop,C.M. (2006) Pa e n Recogni ion and Machine Lea ning.Sp inge ,New
Yo k.
Fe ei a,M.A. and Pu cell,S.M. (2009) A mul i a ia e es o associa ion.
Bioin o ma ics,25, 132–133.
Fusi,N. e al. (2012) Join modelling o con ounding ac o s and p ominen gene ic
egula o s p o ides inc eased accu acy in gene ical genomics s udies. PLoS
Compu . Biol.,8,e1002330.
Gelman,A. e al. (2004) Bayesian Da a Analysis. 2nd edn. Chapman & Hall/CRC,
Boca Ra on, FL.
Geweke,J. (1996) Bayesian educed ank eg ession in econome ics. J. Econom.,75,
121–146.
Hammond,P. and Su ie,M. (2012) La ge-scale objec i e pheno yping o 3D acial
mo phology. Hum. Mu a .,33, 817–825.
Ho elling,H. (1936) Rela ions be ween wo se s o a ia es. Biome ika,28, 321–377.
Howie,B. e al. (2011) Geno ype impu a ion wi h housands o genomes. G3
(Be hesda),1, 457–470.
Howie,B.N. e al. (2009) A lexible and accu a e geno ype impu a ion me hod o
he nex gene a ion o genome-wide associa ion s udies. PLoS Gene .,5,
e1000529.
Inouye,M. e al. (2012) No el loci o me abolic ne wo ks and mul i- issue exp es-
sion s udies e eal genes o a he oscle osis. PLoS Gene .,8, e1002907.
Ke unen,J. e al. (2012) Genome-wide associa ion s udy iden i ies mul iple loci
in luencing human se um me aboli e le els. Na . Gene .,44, 269–276.
Ma inen,P. e al. (2013) Genome-wide associa ion s udies wi h high-dimensional
pheno ypes. S a . Appl. Gene . Mol. Biol.,12, 413–431.
Mo gen hale ,S. and Thilly,W.G. (2007) A s a egy o disco e genes ha ca y
mul i-allelic o mono-allelic isk o common diseases: a coho allelic sums
es (CAST). Mu a . Res.,615, 28–56.
Mo is,A.P. and Zeggini,E. (2010) An e alua ion o s a is ical app oaches o
a e a ian analysis in gene ic associa ion s udies. Gene . Epidemiol.,34,
188–193.
O’Reilly,P.F. e al. (2012) Mul iPhen: join model o mul iple pheno ypes can in-
c ease disco e y in GWAS. PLoS One,7,e34861.
Pa khomenko,E. e al. (2009) Spa se canonical co ela ion analysis wi h applica ion
o genomic da a in eg a ion. S a . Appl. Gene . Mol. Biol.,8, 1–34.
P ice,A.L. e al. (2006) P incipal componen s analysis co ec s o s a i ica ion in
genome-wide associa ion s udies. Na . Gene .,38, 904–909.
Rai aka i,O.T. e al. (2008) Coho p o ile: he Ca dio ascula Risk in Young Finns
S udy. In . J. Epidemiol.,37, 1220–1226.
Ran akallio,P. (1969) G oups a isk in low bi h weigh in an s and pe ina al mo -
ali y. Ac a Paedia . Scand.,193 (Suppl. 193), 1þ.
Saba i,C. e al. (2008) Genome-wide associa ion analysis o me abolic ai s in a
bi h coho om a ounde popula ion. Na . Gene .,41,35–46.
Soininen,P. e al. (2009) High- h oughpu se um NMR me abonomics
o cos -e ec i e holis ic s udies on sys emic me abolism. Analys ,134,
1781–1785.
S egle,O. e al. (2010) A Bayesian amewo k o accoun o complex non-gene ic
ac o s in gene exp ession le els g ea ly inc eases powe in eQTL s udies. PLoS
Compu . Biol.,6,e1000770.
S o ey,J.D. and Tibshi ani,R. (2003) S a is ical signi icance o genomewide s udies.
P oc. Na l Acad. Sci. USA,100, 9440–9445.
Suh e,K. e al. (2011) Human me abolic indi iduali y in biomedical and pha ma-
ceu ical esea ch. Na u e,477, 54–60.
Tang,C.S. and Fe ei a,M.A. (2012) A gene-based es o associa ion using canon-
ical co ela ion analysis. Bioin o ma ics,28, 845–850.
Teslo ich,T.M. e al. (2010) Biological, clinical and popula ion ele ance o 95 loci
o blood lipids. Na u e,466, 707–713.
The Global Lipids Gene ics Conso ium. (2013) Disco e y and e inemen o loci
associa ed wi h lipid le els. Na . Gene .,45, 1274–1283.
The 1000 Genomes P ojec Conso ium. (2012) An in eg a ed map o gene ic a i-
a ion om 1,092 human genomes. Na u e,491, 56–65.
Thompson,J.R. e al. (2011) The me a-analysis o genome-wide associa ion s udies.
B ie . Bioin o m.,12, 259–269.
Tukiainen,T. e al. (2012) De ailed me abolic and gene ic cha ac e iza ion e eals
new associa ions o 30 known lipid loci. Hum. Mol. Gene .,21, 1444–1455.
Va iainen,E. e al. (2010) Thi y- i e-yea ends in ca dio ascula isk ac o s in
Finland. In . J. Epidemiol.,39, 504–518.
Va iku i,S. e al. (2012) He i abili y and gene ic co ela ions explained by common
SNPs o me abolic synd ome ai s. PLoS Gene .,8, e1002637.
Waaijenbo g,S. e al. (2008) Quan i ying he associa ion be ween gene exp essions
and dna-ma ke s by penalized canonical co ela ion analysis. S a . Appl. Gene .
Mol. Biol.,7,1–29.
2033
Gene-me abolome associa ions
Wi en,D.M. and Tibshi ani,R. (2009) Ex ensions o spa se canonical co ela ion
analysis wi h applica ions o genomic da a. S a . Appl. Gene . Mol. Biol.,8,
A icle 28.
Wu,M.C. e al. (2011) Ra e- a ian associa ion es ing o sequencing da a wi h he
sequence ke nel associa ion es . Am. J. Hum. Gene .,89, 82–93.
Xie,Y. e al. (2005) A no e on using pe mu a ion-based alse disco e y a e es ima es
o compa e di e en analysis me hods o mic oa ay da a. Bioin o ma ics,21,
4280–4288.
Zhang,F. e al. (2011) Mul ilocus associa ion es ing o quan i a i e ai s based on
pa ial leas -squa es analysis. PLoS One,6, e16739.
2034
P.Ma inen e al.