scieee Open visual document viewer

Assessing multivariate gene-metabolome associations with rare variants using Bayesian reduced rank regression

Marttinen, Pekka,Pirinen, Matti,Sarin, Antti-Pekka,Gillberg, Jussi,Kettunen, Johannes,Surakka, Ida,Kangas, Antti,Soininen, Pasi,O´Reilly, Paul,Kaakinen, Marika,Kähönen, Mika,Lehtimäki, Terho,Ala-Korpela, Mika,Raitakari, Olli T,Salomaa, Veikko,Järvelin, M

Abstract

MOTIVATION: A typical genome-wide association study searches for associations between single nucleotide polymorphisms (SNPs) and a univariate phenotype. However, there is a growing interest to investigate associations between genomics data and multivariate phenotypes, for example, in gene expression or metabolomics studies. A common approach is to perform a univariate test between each genotype-phenotype pair, and then to apply a stringent significance cutoff to account for the large number of tests performed. However, this approach has limited ability to uncover dependencies involving multiple variables. Another trend in the current genetics is the investigation of the impact of rare variants on the phenotype, where the standard methods often fail owing to lack of power when the minor allele is present in only a limited number of individuals. RESULTS: We propose a new statistical approach based on Bayesian reduced rank regression to assess the impact of multiple SNPs on a high-dimensional phenotype. Because of the method's ability to combine information over multiple SNPs and phenotypes, it is particularly suitable for detecting associations involving rare variants. We demonstrate the potential of our method and compare it with alternatives using the Northern Finland Birth Cohort with 4702 individuals, for whom genome-wide SNP data along with lipoprotein profiles comprising 74 traits are available. We discovered two genes (XRCC4 and MTHFD2L) without previously reported associations, which replicated in a combined analysis of two additional cohorts: 2390 individuals from the Cardiovascular Risk in Young Finns study and 3659 individuals from the FINRISK study.

Full text

Vol. 30 no. 14 2014, pages 2026–2034 BIOINFORMATICS ORIGINAL PAPER doi:10.1093/bioin o ma ics/b u140 Gene ics and popula ion analysis Ad ance Access publica ion Ma ch 24, 2014 Assessing mul i a ia e gene-me abolome associa ions wi h a e a ian s using Bayesian educed ank eg ession Pekka Ma inen 1,2 , Ma i Pi inen 3 , An i-Pekka Sa in 3,4 , Jussi Gillbe g 1 , Johannes Ke unen 3,4 , Ida Su akka 3,4 , An i J. Kangas 5 , Pasi Soininen 5,6 , Paul O’Reilly 7 , Ma ika Kaakinen 8,9 ,MikaKa ¨ho ¨nen 10 , Te ho Leh ima ¨ki 11 , Mika Ala-Ko pela 5,6,12 ,Olli T. Rai aka i 13,14 ,VeikkoSalomaa 15 , Ma jo-Rii a Ja ¨ elin 7,8,9,16,17 , Samuli Ripa i 3,4,18,19, *and Samuel Kaski 1,20, * 1 Depa men o In o ma ion and Compu e Science, Helsinki Ins i u e o In o ma ion Technology HIIT, Aal o Uni e si y, Esbo, Finland, 2 Cen e o Communicable Disease Dynamics, Ha a d School o Public Heal h, Bos on, MA, USA 3 Ins i u e o Molecula Medicine Finland (FIMM), Uni e si y o Helsinki, 4 Uni o Public Heal h Genomics, Na ional Ins i u e o Heal h and Wel a e, Helsinki, 5 Compu a ional Medicine, Ins i u e o Heal h Sciences, Uni e si y o Oulu and Oulu Uni e si y Hospi al, Oulu, 6 NMR Me abolomics Labo a o y, School o Pha macy, Uni e si y o Eas e n Finland, Kuopio, Finland, 7 Depa men o Epidemiology and Bios a is ics, MRC Heal h P o ec ion, Agency (HPA) Cen e o En i onmen and Heal h, School o Public Heal h, Impe ial College, London, UK, 8 Ins i u e o Heal h Sciences, 9 Biocen e Oulu, Uni e si y o Oulu, Oulu, 10 Depa men o Clinical Physiology, Tampe e Uni e si y Hospi al and Uni e si y o Tampe e, 11 Depa men o Clinical Chemis y, Fimlab Labo a o ies, Uni e si y o Tampe e School o Medicine, Tampe e, Finland, 12 Compu a ional Medicine, School o Social and Communi y Medicine and he Medical Resea ch Council In eg a i e Epidemiology Uni , Uni e si y o B is ol, B is ol, UK, 13 Depa men o Clinical Physiology and Nuclea Medicine, 14 Resea ch Cen e o Applied and P e en i e Ca dio ascula Medicine, Uni e si y o Tu ku and Tu ku Uni e si y Hospi al, Tu ku, 15 Depa men o Ch onic Disease P e en ion, Na ional Ins i u e o Heal h and Wel a e, Helsinki, 16 Uni o P ima y Ca e, Oulu Uni e si y Hospi al, 17 Depa men o Child en and Young People and Families, Na ional Ins i u e o Heal h and Wel a e, Oulu, Finland, 18 Wellcome T us Sange Ins i u e, Hinx on, Camb idge, UK, 19 Hjel Ins i u e and 20 Depa men o Compu e Science, Helsinki Ins i u e o In o ma ion Technology HIIT, Uni e si y o Helsinki, Helsinki, Finland Associa e Edi o : Je ey Ba e ABSTRACT Mo i a ion: A ypical genome-wide associa ion s udy sea ches o associa ions be ween single nucleo ide polymo phisms (SNPs) and a uni a ia e pheno ype. Howe e , he e is a g owing in e es o in es i- ga e associa ions be ween genomics da a and mul i a ia e pheno- ypes, o example, in gene exp ession o me abolomics s udies. A common app oach is o pe o m a uni a ia e es be ween each geno- ype–pheno ype pai , and hen o apply a s ingen signi icance cu o o accoun o he la ge numbe o es s pe o med. Howe e , his app oach has limi ed abili y o unco e dependencies in ol ing mul- iple a iables. Ano he end in he cu en gene ics is he in es iga ion o he impac o a e a ian s on he pheno ype, whe e he s anda d me hods o en ail owing o lack o powe when he mino allele is p esen in only a limi ed numbe o indi iduals. Resul s: We p opose a new s a is ical app oach based on Bayesian educed ank eg ession o assess he impac o mul iple SNPs on a high-dimensional pheno ype. Because o he me hod’s abili y o com- bine in o ma ion o e mul iple SNPs and pheno ypes, i is pa icula ly sui able o de ec ing associa ions in ol ing a e a ian s. We demon- s a e he po en ial o ou me hod and compa e i wi h al e na i es using he No he n Finland Bi h Coho wi h 4702 indi iduals, o whom genome-wide SNP da a along wi h lipop o ein p o iles comp is- ing 74 ai s a e a ailable. We disco e ed wo genes (XRCC4 and MTHFD2L) wi hou p e iously epo ed associa ions, which eplica ed in a combined analysis o wo addi ional coho s: 2390 indi iduals om he Ca dio ascula Risk in Young Finns s udy and 3659 indi iduals om he FINRISK s udy. A ailabili y and implemen a ion: R-code eely a ailable o down- load a h p://use s.ics.aal o. i/pema i/gene_me abolome/. Con ac : samuli. ipa i@helsinki. i; samuel.kaski@aal o. i Supplemen a y in o ma ion: Supplemen a y da a a e a ailable a Bioin o ma ics online. Recei ed on No embe 7, 2013; e ised on Feb ua y 27, 2014; accep ed on Ma ch 4, 2014 1 INTRODUCTION Concen a ions o human me aboli es a e associa ed wi h isk o many common diseases; o example, low- and high-densi y lipo- p o ein choles e ol (LDL and HDL) le els a e associa ed wi h co ona y a e y disease. Fo his eason, human me abolism has been unde in ensi e in es iga ion and o e he pas ew yea s se e al genome-wide associa ion s udies ha e success ully un- co e ed a pa o i s gene ic basis (Ke unen e al., 2012; *To whom co espondence should be add essed. ßThe Au ho 2014. Published by Ox o d Uni e si y P ess. This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion Non-Comme cial License (h p://c ea i ecommons.o g/licenses/ by-nc/3.0/), which pe mi s non-comme cial e-use, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly ci ed. Fo comme cial e-use, please con ac jou nals.pe [email protected] Saba i e al., 2008; Suh e e al., 2011; Teslo ich e al., 2010). Fo example, a la ge me a-analysis (Teslo ich e al., 2010) iden i ied 95 loci in luencing he le els o o al choles e ol, LDL, HDL and iglyce ides. Mo e ecen s udies used ine subclassi ica ions o me aboli es and disco e ed dozens o no el loci (Ke unen e al., 2012; Suh e e al., 2011). Despi e hese ad ances, he a iance explained by all epo ed single nucleo ide polymo phisms (SNPs) alls a below he sugges ed he i abili y o he common me aboli es, as es ima ed ei he om win s udies (Ke unen e al., 2012) o om mo e dis an ly ela ed indi iduals (Va iku i e al., 2012). This mo i a es us o de elop new app oaches o associa ion es ing ha could be e use all in o - ma ion a ailable o us. The p esen -day coho s udies o en come wi h a ich se o pheno ypic ea u es. Examples in addi ion o me abolomics (Ke unen e al., 2012; Soininen e al., 2009; Suh e e al., 2011) include s udies o gene exp ession (Acke mann e al., 2013) and 3D- acial imaging (Hammond and Su ie, 2012). As a conse- quence, we need s a is ical me hods ha inc ease he powe o unco e geno ype–pheno ype dependencies by combining in o - ma ion o e se e al ela ed pheno ypes (Fe ei a and Pu cell, 2009; Inouye e al., 2012; O’Reilly e al., 2012). The unde lying idea is ha i a gene ic a ian a ec s a ai , hen i is likely o a ec o he ai s ha a e ela ed o he i s one and, by es ing o associa ion wi h he wo ai s join ly, powe may be inc eased. This easoning can be aken a s ep u he by es ing all ai s in high-dimensional omics da a simul aneously, o ex- ample, all me aboli es in comp ehensi e me abolomic p o iles. A compa ison o di e en s a is ical me hods a ailable o join es ing o comple e me abolomics p o iles was ecen ly con- duc ed (Ma inen e al., 2013). Besides es ing se e al pheno ypes simul aneously, he abili y o de ec ce ain kinds o associa ions may be boos ed by com- bining s a is ical e idence o e se e al SNPs. Usually his is done in a supe ised manne , such ha SNPs ela ed by loca ion o unc ion, o example, a e es ed simul aneously. Combining in o ma ion o e mul iple SNPs is pa icula ly c ucial wi h a e a ian s, i.e. SNPs whe e he mino allele is p esen in a small p opo ion o he popula ion. Tes ing such SNPs indi idually is unlikely o yield signi ican indings because o limi ed powe . Mos app oaches o handling a e a ian s a e based on collap- sing se e al a e SNPs in o a single a iable (Bansal e al., 2010). Fo example, one can simply collapse se e al a e a ian s in o a single indica o , elling whe he any o he a e a ian s is p esen in he indi idual (Mo gen hale and Thilly, 2007) o o coun he numbe o a e a ian s p esen in he indi idual (Mo is and Zeggini, 2010). The p oblem wi h he collapsing me hods is he implici assump ion ha he e ec s a e in he same (o a p ede- ined) di ec ion. A mo e sophis ica ed a iance componen me hod a oiding his assump ion is able o in es iga e he impac o se e al a e a ian s on a uni a ia e ai (Wu e al., 2011). E en i he e exis s a la ge numbe o me hods o analyzing a e a ian s, none o hose has been ailo ed o mul i a ia e pheno ypes. On he o he hand, s anda d me hods o mul i a i- a e pheno ypes (Fe ei a and Pu cell, 2009) ha e no been ho - oughly in es iga ed in he con ex o a e a ian s, and we will see la e in he ex ha se e e o e i ing may occu . In summa y, me hods o dealing wi h a e a ian s in he con- ex o mul i a ia e pheno ypes a e clea ly lacking. In his a icle, we de i e a no el o mula ion o he Bayesian educed ank eg ession model (Geweke, 1996) o de ec mul i- a ia e associa ions be ween p ede ined g oups o SNPs and a high-dimensional pheno ype. In pa icula , ou app oach is sui - able o analyzing bo h common and a e a ian s. Ou o mu- la ion inco po a es p io knowledge abou e ec sizes o inc ease he powe o de ec associa ions. Fu he mo e, i is capable o co ec ing o he numbe o SNPs conside ed, which is impo - an when es ing a la ge numbe o SNP g oups o di e en sizes. We alida e ou me hod by assessing associa ions be ween SNPs in all human genes, one gene a a ime, and me abolic p o iles comp ising ine-scale lipop o ein measu emen s o 4702 indi id- uals om he No he n Finland Bi h Coho 1966 (Ran akallio, 1969; Saba i e al., 2008). Among he op-sco ing genes wi hou known associa ions o he ai s s udied, wo genes (XRCC4 and MTHFD2L) eplica ed in a combined analysis o 2390 indi id- uals om he Ca dio ascula Risk in Young Finns s udy (YFS; Rai aka i e al., 2008) and 3659 indi iduals om he FINRISK s udy (Va iainen e al., 2010). Addi ional analyses o he same da a con i med ha al e na i e me hods disco e ed only one o hese associa ions and did no iden i y any u he associa ions ha we e no known be o e. 2 METHODS 2.1 Model To build ou model, we assume ha he pheno ypes may be a ec ed by h ee kinds o a iables, (i) known ac o s, such as age, sex o popula ion s uc u e, (ii) unknown ac o s, such as expe imen al condi ions and o he ba ch e ec s, and (iii) SNPs unde conside a ion, as schema ically p e- sen ed in Figu e 1. In con as o s anda d eg ession, whe e each SNP– pheno ype pai has a pa ame e ep esen ing he e ec o he SNP on he pheno ype, he e we assume ha a combina ion o se e al SNPs is in lu- encing se e al pheno ypes h ough some unknown ac o s. This assump- ion is compac ly exp essed in e ms o he educed ank eg ession, whe e he SNPs a e i s p ojec ed on o a low-dimensional subspace, and he p ojec ions a e hen used as eg esso s when p edic ing he pheno ypes. The educed ank eg ession o mula ion immedia ely implies some s uc u al assump ions deemed sensible in he cu en se - ing: i s , i a SNP has an e ec on a pheno ype, hen he SNP is likely o ha e an e ec on o he pheno ypes; second, i a pheno ype is a ec ed by a SNP, hen he pheno ype is mo e likely o be a ec ed by o he ela ed SNPs as well. Le Ndeno e he numbe o indi iduals, S he numbe o SNPs, P he numbe o pheno ypes and C he numbe o o he co a ia es. Fo mally, we conside he Bayesian educed ank eg ession model Y¼X þZA þHTþEð1Þ whe e YNPcon ains he pheno ypes, XNScon ains he SNPs, SK1 and K1P ep esen a low- ank app oxima ion o he eg ession coe i- cien ma ix ¼,ZNC ep esen s o he co a ia es wi h he co es- ponding coe icien ma ix ACP,HNK2con ains hidden con ounding ac o s wi h he co esponding coe icien ma ix PK2and ENP¼½e1,...,eNT,wi heiNð0, Þ,whe e¼diagð2 1,...,2 PÞ. No e ha by in eg a ing o e he hidden ac o s H, he model is equi a- len o yiNðTxiþATzi,TþÞ,i¼1, ...,Nð2Þ 2027 Gene-me abolome associa ions whe e i is assumed ha he ac o s ha e independen s anda d no mal p io dis ibu ions. The e o e, we see ha ha ing he la en a iable pa HTin he model co esponds o assuming a low- ank app oxima ion o he ull co a iance ma ix, which is impo an when analyzing high- dimensional da ase s. Fo compu a ional easons, we es ic he ank o he eg ession model, K 1 , o uni y in ou genome-wide analysis (w.l.o.g. in he de ec ion ask, see below) and show esul s wi h K1¼1, 2, 3 o a ew ep esen a- i e examples. Howe e , he e we p esen a gene al in ini e-dimensional amewo k a ailable in ou implemen a ion, which does no necessi a e he selec ion o a ixed ank. In gene al, o use he Bayesian educed ank eg ession model, he ank o he model, K 1 , and he ank o he low- ank app oxima ion o he co a iance ma ix, K 2 , mus be selec ed. A ecen Bayesian in ini e spa se ac o analysis model (Bha acha ya and Dunson, 2011) ci cum en s he selec ion o a ixed ank o K 2 by assum- ing in p inciple an in ini e numbe o columns in he ma ix; howe e , he columns sh ink p og essi ely such ha only he i s anks a e in lu- en ial in p ac ice. We assume his p io o mula ion o ou noise model HTþE. We exploi he idea u he by allowing also he ank K 1 o be in ini e in p inciple, and en o cing he low- ank na u e by sh inking he columns o and he ows o inc easingly as he column/ ow index g ows. In p ac ice, one needs o speci y uppe bounds o K 1 and K 2 .We selec he uppe bound o K 2 using he adap i e p ocedu e o Bha acha ya and Dunson (2011), and we ha e implemen ed an analo- gous me hod o lea ning he uppe bound o K 1 . In p ac ice, we wan ed o minimize he model complexi y o maximize he powe o de ec asso- cia ions and speed up he compu a ions. The e o e, we decided o un he genome-wide analysis (see Sec ion 3.1) using a ixed ank K1¼1. The model wi h K1¼1 is su icien o ou pu poses o de ec ing whe he a gene is un ela ed o he pheno ypes (in which case, ank ze o would al eady be su icien ), al hough o p edic ion a highe ank migh be mo e sui able. We expe imen ed wi h a selec ed se o known genes wi h uppe bounds 2 and 3 (see below) and no iced ha he e ec o inc easing he uppe bound om uni y had only a small impac on he amoun o a ia ion ha is explained by he model. F om he biological pe spec i e, his means ha he in luence o a gene on he pheno ypes can mos ly be desc ibed in e ms o a single la en ac o media ing he e ec . A de ailed model desc ip ion is gi en in Supplemen a y Sec ion 1. The model is ela ed o many published me hods, and a ho ough compa ison can be ound in Supplemen a y Sec ion 2. Co ec ion o unknown ac o s has ecen ly been conside ed wi h he s anda d eg es- sion model, ypically o jus one SNP a a ime (Fusi e al., 2012; S egle e al., 2010). The di e ence o he o iginal Bayesian educed ank eg es- sion o mula ion (Geweke, 1996) is ha ou model uses low- ank ap- p oxima ion o he co a iance ma ix, making i mo e sui able o high-dimensional pheno ypes, and in o ma i e p io dis ibu ions accom- moda ing p oblem-speci ic knowledge. Fu he mo e, we use he model in a new way, as desc ibed in he subsequen sec ions. The Bayesian in ini e spa se ac o analysis model has been used o high-dimensional da a (Bha acha ya and Dunson, 2011); he e, we use i o ep esen he mul i- a ia e noise. Canonical co ela ion analysis (CCA) is a classical ool o modeling mul i a ia e dependencies ha ha e ecen ly been in oduced in he associa ion s udy con ex (Fe ei a and Pu cell, 2009; Ho elling, 1936). Spa se CCA (Pa khomenko e al., 2009; Waaijenbo g e al., 2008; Wi en and Tibshi ani, 2009) is mo e sui able o high-dimensional da ase s; howe e , in oducing p io dis ibu ions o CCA ha would be in ui i e in he associa ion s udy con ex does no seem s aigh o wa d. 2.2 P opo ion o o al a ia ion explained A commonly used measu e o he impac o mul iple SNPs, say x1,...,xS, on a uni a ia e ai yis he p opo ion o a iance explained (PVE) by he SNPs: PVE ¼1Va ðyPS i¼1^ aixiÞ Va ðyÞ ¼Va ðPS i¼1^ aixiÞ Va ðyÞ: He e, PS i¼1^ aixiis a linea p edic ion o he pheno ype y, gi en he SNPs. Analogously, as a measu e o he o e all impac o mul iple SNPs on a high-dimensional pheno ype, we p opose o use he p opo ion o o al a ia ion o he pheno ypes explained (PTVE) by he model, namely PTVE ¼T ðCo ð^ YÞÞ T ðCo ðYÞÞ ð3Þ whe e ^ Yis a p edic ion o a high-dimensional pheno ype Y om he model and T deno es he ace, i.e. he sum o he diagonal elemen s o he ma ix. In (3) and in gene al, he o al a ia ion o a mul i a ia e andom a iable is de ined as he ace o he co a iance ma ix, i.e. he sum o he a iances o he indi idual a iables. The e o e, PTVE meas- u es he join impac o he SNPs on se e al pheno ypes and hence is expec ed o yield high sco es o such dependencies in which many pheno ypes a e a ec ed by he SNPs, e en i none o he e ec s is la ge by i sel . In he Bayesian s a is ical amewo k, he in e ences a e based on pos- e io p obabili y dis ibu ions o he quan i ies o in e es (Gelman e al., 2004). Wi h he Bayesian educed ank eg ession model, samples om he pos e io dis ibu ion o he PTVE can be ob ained om PTVEðiÞ¼T Co XðiÞðiÞ  T Co YðÞðÞ ð4Þ whe e ðiÞand ðiÞa e samples om he pos e io dis ibu ion o he pa ame e s and . I is simila ly s aigh o wa d o es ima e he pos- e io dis ibu ion o he p opo ion o a ia ion explained by he a e Y X ψΓ ZA H T E, Y X Z A H T E A B Fig. 1. G aphical illus a ion o he model. (A) The a iables and depen- dencies be ween hem. The pheno ypes Ya e assumed o be a ec ed by known ac o s, such as age o sex, unknown ac o s, such as ba ch e ec s caused by a ying expe imen al condi ions, and he SNPs. The in luence o he SNPs is media ed by unknown combina ions o he o iginal SNPs, ep esen ed by black squa es. (B) The same model using ma ix no a ion. Ma ices con aining he obse ed a iables, Y( he pheno ypes), X( he SNPs) and Z(known ac o s), a e blue. The eg ession coe icien ma i- ces a e ed. No e ha he coe icien ma ix o he SNP e ec s is w i en as a p oduc o wo ma ices, and , co esponding o a low- ank app oxima ion o an uncons ained coe icien ma ix. The b own ma i- ces comp ise unobse ed a iables, H(unknown ac o s) and E(noise e ms) 2028 P.Ma inen e al. a ian s, by di iding he a ia ion o he p edic ion in o wo componen s, one co esponding o he a e a ian s, he o he o he common a ian s. The sco e ob ained in his way is e e ed o as PTVE- a e in he sequel and i s exac de ini ion is gi en in Supplemen a y Sec ion 3. In p ac ice, we app oxima e he pos e io dis ibu ion by using a mode-based poin es ima e o one o he pa ame e s, , and es ima ing he join dis ibu- ion o he o he pa ame e s using an Ma ko chain Mon e Ca lo (MCMC) algo i hm, as desc ibed ho oughly in Supplemen a y Sec ions 4 and 5. 2.3 In o ma i e p io In he Bayesian analysis, ex e nal backg ound knowledge may be inco - po a ed in he s a is ical analysis h ough p io p obabili y dis ibu ions. As ou analysis is ocused on es ima ing he pos e io dis ibu ion o PTVE, a sensible p io is ob ained by making he p io dis ibu ion o he PTVE ep esen ou ue belie s abou his quan i y. The p io dis i- bu ion o he PTVE appea s no o be a ailable in a closed o m; how- e e , in Supplemen a y Sec ion 6, we de i e esul s ha show how he dis ibu ion o he mean o he PTVE, PTVE, depends on model hype - pa ame e s. Using hese esul s, we se he p io dis ibu ions o sa is y he ollowing p ope ies: MedianðPTVEÞ¼106ð5Þ and PðPTVE40:001Þ¼0:01 ð6Þ Equa ion (5) means ha he p io median o PTVE is close o ze o, as we expec mos o he genes o be un ela ed o he pheno ypes. Equa ion (6) says ha wi h a small p obabili y, he e equal o 0.01, he gene may explain40.1% o he o al a ia ion, a alue deemed esonable based on he backg ound knowledge. The esul ing p io dis ibu ion o he PTVE is ob ained by in eg a ing o e he dis ib ion o PTVE,andweused Mon e Ca lo simula ion o in es iga e he dis ibu ion. Figu e 2 shows PTVE alues sampled om he p io dis ibu ion. The ollowing desi able cha ac e is ics can be seen: i s , a peak close o ze o, sh inking he coe - icien s when no e ec is p esen ; second, a long ail, co esponding o he genes wi h a non-negligible impac on he pheno ypes, wi hou imposing s ong belie s abou he ac ual size o hese non-ze o e ec s. Ano he impo an p ope y ha ollows om using he in o ma i e p io is ha he numbe o SNPs unde conside a ion can be accoun ed o . De ailed examina ion o Co olla y 1 in Supplemen a y Sec ion 6 e eals ha wi h ixed hype pa ame e s, he expec ed PTVE is p opo - ional o PS i¼1Va ðxiÞ, i.e. he o al a ia ion o he SNPs. Roughly, his means ha i he numbe o SNPs doubles, he expec ed PTVE doubles as well, i he hype pa ame e s a e kep ixed. Using he Co olla ies 1 and 2 in Supplemen a y Sec ion 6, i is s aigh o wa d o modi y he hype - pa ame e dis ibu ions o asse he p io condi ions (5) and (6), implying in ou se ing ha all genes a e expec ed o explain he same amoun o he a ia ion o he pheno ypes, ega dless o how many SNPs hey con ain. 2.4 Da a As a da ase o de ec ing associa ions, we conside a sample o 4702 indi iduals om he No he n Finland Bi h Coho 1966 (NFBC1966), a bi h coho s udy o child en bo n in 1966 in he wo no he nmos p o inces o Finland (Ran akallio, 1969). The blood sam- ples o he DNA ex ac ion and pheno ype da a we e collec ed a a ollow-up isi when he pa icipan s we e 31 yea s o age. Fo eplica- ion, we conside wo coho s. The YFS is a popula ion-based p ospec i e coho s udy (Rai aka i e al., 2008) conduc ed in Finland, he pu pose o which was o in es iga e he le els o ca dio ascula isk ac o s in chil- d en and adolescen s in di e en pa s o he coun y. The FINRISK s udy comp ises c oss-sec ional popula ion su eys ha ha e been ca ied ou e e y 5 yea s since 1972, o assess he isk ac o s o ch onic diseases, wi h emphasis on ca dio ascula isk ac o s (Va iainen e al., 2010). The indi iduals analyzed in his s udy belong o he sample om he yea 1997. The blood samples o he wo eplica ion coho s we e collec ed when he pa icipan s we e 30–45 and 25–71 yea s o age, espec i ely. The s udy p o ocols o all da ase s ha e been app o ed by he local e hics commi ees. The samples we e geno yped wi h Illumina a ays (Illumina, Inc. San Diego, CA, USA) and impu ed wi h IMPUTE 2 (Howie e al.,2009, 2011) using a 1000 Genomes P ojec e e ence panel (The 1000 Genomes P ojec Conso ium, 2012). O he esul ing good-quali y au o- somal SNPs (in o 40.4), we ex ac ed SNPs in 24025 human genes by adding 50 kb lanking egions on bo h sides o he endpoin s o he genes gi en in NCBI gene da abase (genome assembly GRCh37.p10, NCBI anno a ion 104, No embe 2012). As a p ep ocessing s ep, we educed he geno ype space wi hin each gene o he mos p omising 200 SNPs (a mos ) ha had he highes canonical co ela ion es sco e (Fe ei a and Pu cell, 2009) wi h he me aboli es; howe e , o p e en o e i ing, he p io s we e speci ied as desc ibed abo e using he unp uned SNP se . In p elimina y expe imen s, dec easing he numbe o SNPs in his way om 800 o 200 had no isible e ec on he esul s. Finally, he SNPs we e scaled o ha e uni a iance. Pheno ype da a came om he se um NMR me abolomics pla o m desc ibed ea lie (Soininen e al., 2009). As a p ep ocessing s ep, he ai s we e quan ile-no malized o ha e s anda d no mal dis ibu ion. Indi iduals wi h 20% missing alues we e emo ed and he emaining missing alues we e impu ed by sampling hem om he mul i a ia e no mal dis ibu ion. In his wo k, we analyzed a subse o 74 lipop o ein subclass measu es (Supplemen a y Table S3). The empi ical co ela ion ma ix o he ai s is shown in Supplemen a y Figu e S1. Linea eg es- sion was used o co ec he pheno ypes o age, sex and popula ion s uc u e using 10 p incipal componen s (P ice e al., 2006). P io dis ibu ion o PTVE PTVE (x10 ) P obabili y 0246810 0 0.3 0.6 0.9 0.000 0.005 0.010 0.015 0 1000 2000 3000 4000 5000 Pos e io PTVE o LIPC gene PTVE Densi y O iginal Pe mu ed Fig. 2. P io and pos e io dis ibu ions o he p opo ion o o al a i- a ion explained (PTVE) by he model. The panel on he le shows he p io dis ibu ion imposed on he p opo ion o o al a ia ion o he pheno ypes explained by he SNPs unde conside a ion (he e, he SNPs om he LIPC gene). The median o he p io dis ibu ion is loca ed a 4e-6. The cha ac e is ic ea u es o he p io dis ibu ion include he peak a alues close o ze o, e ec i ely emo ing noise unless he e is s ong e idence abou a possible associa ion, and he long ail allowing a small pe cen age o genes o explain la ge p opo ions o he pheno ype a ia ion. The igh mos bin on he x-axis con ains he o al p obabili y o alues exceeding he maximum alue on he axis. The panel on he igh shows he pos e io dis ibu ion o he PTVE o he same SNPs. Two pos e io densi ies a e shown, one showing he dis ibu ion o he o iginal da a, he o he showing he dis ibu ion o da a in which he ows o he pheno ype ma ix ha e been pe mu ed. No ice he di e ing scales on he x-axes o he wo panels 2029 Gene-me abolome associa ions 3RESULTS 3.1 Genome-wide analysis o NFBC1966 da a We compu ed he PTVE and PTVE- a e sco es o each o he 24 025 human genes in he NFBC1966 da a, one gene a a ime, using he Bayesian educed ank eg ession. Table 1(a) shows he op i e genes wi h he highes PTVE sco es, all o which a e well-known lipid-associa ed loci. Mo e gene ally, all 43 op-sco - ing genes we e loca ed wi hin 1Mb om p e iously epo ed genome-wide signi ican lipid associa ions (Ke unen e al., 2012; Teslo ich e al., 2010; The Global Lipids Gene ics Conso ium, 2013), and de ailed lis ing o he op genes is gi en in Supplemen a y Table S4. Fu he mo e, o he op 100 genes, which co espond o 0.2 alse disco e y a e (FDR) (see he nex sec ion), 73 genes had known associa ions. Toge he , hese indings se e as a alida ion o he me hod and i s implemen a ion. To illus a e he e ec o he model ank on he esul s, we ca ied ou a de ailed analysis o h ee known lipid genes: LIPC, APOB and PLTP,wi h ankK1¼1, 2, 3. The genes we e se- lec ed such ha hey we e loca ed in di e en ch omosomes and had dissimila associa ion p o iles [APOB associa ed mos s ongly o VLDL, IDL and LDL, PLTP o HDL and LIPC o VLDL, IDL and HDL (Tukiainen e al., 2012)]. In addi ion o conside ing he genes sepa a ely, we epea ed a join analysis o all h ee possible pai wise gene combina ions. The esul s a e shown in Supplemen a y Figu e S2 and a e summa ized as ol- lows: (i) he ank o he model had a mino e ec on he esul s o a single gene, (ii) inc easing he ank om K1¼1 oK1¼2 inc eased he a iance explained in all join pai wise analyses, and he inc ease was signi ican in wo o he h ee cases, (iii) a combina ion o wo genes always had a la ge e ec han ei he gene indi idually; howe e , he sum o he indi idual e ec s was sligh ly la ge han he e ec o he combina ion. We a ibu e his di e ence mainly o he s onge sh inkage in he second han he i s geno ype componen . (i ) The i s componen o a join model was always s ongly co ela ed wi h he single-gene model wi h he la ge e ec , he second componen wi h he single-gene model wi h he smalle e ec . This is as expec ed, as he p io dis ibu ion was designed o iden i y he s onges associa ions using he i s geno ype componen . To check whe he no el associa ions could be de ec ed by any me hod, we ca ied ou a eplica ion expe imen wi h he YFS and FINRISK da ase s o he mos p omising genes, a e excluding genes loca ed wi hin 1 Mb om p e iously epo ed associa ions. We conside ed om each me hod all genes wi h FDR 50.4 as p omising. Fo he PTVE- a e sco e, six genes we e es ed, none o which had p e iously epo ed associa ions o lipids. Fo he PTVE sco e, 305 genes had FDR 50.4; how- e e , only 167 we e no loca ed close o known associa ions, and hese 167 genes we e selec ed o eplica ion. Fo he pu poses o eplica ion, he mul i a ia e dependency be ween he mul iple SNPs and pheno ypes was educed in o a uni a ia e es by using pa ame e s es ima ed in he NFBC1966 da a o cons uc uni a ia e geno ype and pheno ype combin- a ions (see Supplemen a y Sec ion 7 o u he de ails). Then, he s anda d linea model was used o es o posi i e co ela ion be ween he geno ype and pheno ype combina ions. A pooled linea eg ession coe icien combining YFS and FINRISK da ase s was o med using a ixed e ec model o e he wo da ase s (Thompson e al., 2011), and he P- alue was ob ained by ela ing he pooled es ima e o i s SD. A one- ailed es was used,as we we e only in e es ed in indings in which he e ec s we e in he same di ec ion in he eplica ion da ase s as in he NFBC1966 da a. Bon e oni co ec ion was used o accoun o he numbe o genes es ed wi h each me hod, such ha co ec ed Table 1. Summa y o esul s om he genome-wide analysis o he eal da a Ch Locus PTVE (SD) Ra e P- alue Gene ank PTVE Pai wise S-CCA CCA-single (a) 15 LIPC 0.01 (5e-04) 0.015 5e-19 1 1 1 132 19 APOC1 0.0046 (3e-04) 0.017 1.4e-26 2 8 18 275 19 PVRL2 0.0045 (1e-04) 0.0015 7e-35 3 9 14 276 2APOB 0.0044 (3e-04) 0.017 1.1e-17 4 45 41 3244 11 APOA5 0.0043 (2e-04) 0.021 5e-10 5 26 33 2433 (b) 16 SPIRE2 a 0.0015 (9e-05) 0.89 0.00091 b 5 ( a e) 8344 5490 4973 5XRCC4 0.0024 (2e-04) 0.55 0.0016 6 ( a e) 2706 5163 2155 (c) 2DTNB a 0.0015 (2e-04) 0.15 2.6e-04 c 138 4715 338 7652 4MTHFD2L 0.0015 (2e-04) 0.019 7e-06 163 102 1444 1552 No e: (a) Re e ence esul s o genes wi h i e highes PTVE sco es. (b) Replica ed genes om he PTVE- a e sco e (o six genes es ed o eplica ion). (c) Replica ed genes om he PTVE sco e (o 167 genes). Fi e o he eplica ed genes (PPBP, CXCL5, CXCL2, PF4 and CXCL3) a e no shown as hey we e loca ed wi hin 1 Mb om MTHFD2L, which had he s onges e ec . Column PTVE (SD) shows he p opo ion o o al a ia ion explained and i s SD, a e speci ies he p opo ion o he a ia ion explained by he gene a ibu ed o he a e a ian s, P- alue speci ies he P- alue pooled o e YFS and FINRISK eplica ion da ase s (unless s a ed o he wise) and he las ou columns speci y he anking o he gene among all genes wi h di e en me hods. a Deno es genes ha eplica ed signi ican ly in only one o he wo eplica ion da ase s. b/c Replica ion P- alue in FINRISK/YFS. 2030 P.Ma inen e al. P- alue h eshold co esponding o he nominal 0.05 le el o signi icance was equal o P50:0083 o he six pu a i e no el genes om PTVE- a e sco e and P50:00030 o he 167 pu a i e genes om PTVE sco e. We conside ed a eplica ion signi ican i he es was nomin- ally signi ican (P50.05) in bo h eplica ion da ase s, and he P- alue o he pooled es ima e was signi ican a e co ec ing o he mul iple es s. In o ma ion abou genes ha eplica ed signi ican ly is p o ided in Table 1(b) and (c) and in Figu e 3. One gene wi h PTVE- a e and six genes wi h PTVE we e de- ec ed, co esponding o wo independen genes: XRCC4 and MTHFD2L.MTHFD2L is loca ed wi hin 1 Mb om wo SNPs ( s2168889 and s16850360) associa ed wi h ‘me abolic ne wo ks’ con aining some lipop o ein ai s om ou da a (Inouye e al., 2012). Howe e , he op me aboli e in he associ- a ions was Albumin, and when we epea ed he mul i a ia e es wi h he lipop o ein ai s conside ed he e, hese associa ions we e no longe genome-wide signi ican (P¼2.5e-4 and P¼6.4e-7). In addi ion, wo genes, SPIRE2 and DTNB, epli- ca ed signi ican ly (a e he mul iple es ing co ec ion) in one, bu no in he o he eplica ion da ase . Table 1 also p esen s he ankings o hese genes by al e na i e me hods (see below). We see ha XRCC4,SPIRE2 and DTNB we e comple ely missed by he o he me hods. Supplemen a y Table S1 shows he SNPs con ibu ing o he epo ed associa ions. We see ha wi h XRCC4,SPIRE2 and DTNB, many SNPs, some o which a e a e, a e equi ed o ep esen he o e all associa ion. This ex- plains why hese genes did no ecei e any signal om he s and- a d es ing wi h he pai wise linea model. Supplemen a y Figu e S3 shows g aphically how he pheno ypes a e a ec ed by he signi ican genes and he known LIPC gene. We see ha in nei- he o he new genes is he e ec ocused on any single ai , bu a he a small e ec is seen on many lipop o ein measu es. This is no su p ising, as he PTVE sco e is expec ed o gi e high sco es o p ecisely his kind o associa ion. Supplemen a y Figu e S4 shows he es ima ed SNP coe icien s o he genes and demon- s a es he use ulness o analyzing all SNPs in a gene simul an- eously o educe noise esul ing om he co ela ion be ween he SNPs. Fu he backg ound in o ma ion on hese genes is p e- sen ed in Supplemen a y Table S2; howe e , a mo e ho ough biological in e p e a ion o he genes emains o u u e wo k. Finally, we epea ed a simila analysis wi h h ee al e na i e me hods: (i) exhaus i e pai wise sea ch wi h a linea model, whe e he minus loga i hm o he smalles pai wise P- alue o e all SNPs in he gene and me aboli es was aken as he es sco e o he gene, (ii) CCA, applied o a single SNP e sus all me aboli es a a ime as by Fe ei a and Pu cell (2009) and Inouye e al. (2012) and (iii) spa se CCA, which was used o compu e he canonical co ela ion be ween all SNPs in a gene and all pheno ypes (Pa khomenko e al., 2009). The me hods (ii) and (iii) we e ound o be he mos powe ul in a ecen com- pa ison o app oaches o a mul i a ia e me abolomics Fig. 3. Resul s o genes wi h signi ican eplica ion in bo h es se s: XRCC4 and MTHFD2L; o e e ence, he well-known LIPC lipid locus is also shown. Each panel shows he iden i ied pheno ype combina ion plo ed agains he geno ype combina ion. The le column shows esul s in he NFBC1966 da ase , in which he associa ions we e de ec ed. The cen e and igh columns show esul s wi h he YFS and FINRISK da ase s, whe e coe icien ma ices lea ned wi h he NFBC1966 da a we e used o o m he a iable combina ions. The g een and ed backg ound colo ings ma k he indi iduals wi h he highes /lowes geno ype combina ion alues, and he pheno ype alues in hese ex eme g oups a e in es iga ed in mo e de ail in Supplemen a y Figu e S3 2031 Gene-me abolome associa ions pheno ype (Ma inen e al., 2013); howe e , he e he SNPs we e no p uned using he mino allele equency be o e applying he me hods as by Ma inen e al. (2013) because he e he ocus is in he a e a ian se ing. Simila ly o PTVE and PTVE- a e me h- ods, we ied o eplica e he op-sco ing genes up o FDR ¼0.4. The las column in Table 2 shows he numbe s o genes included in he eplica ion wi h de ailed lis ings gi en in Supplemen a y Tables S4–S8. The associa ion in MTHFD2L was con i med o be genome-wide signi ican using he s anda d pai wise linea eg ession (SNP s185567543, ai S.LDL.P, P¼2.5e-15, whe e he P- alue is pooled o e all h ee da ase s). All o he eplica ed associa ions we e loca ed wi hin 1 Mb o his o p e- iously known associa ions, hus yielding no addi ional no el de ec ions. 3.2 Powe compa ison To in es iga e he powe o he in oduced me hod and compa e i wi h al e na i e me hods, we es ima ed FDR co esponding o di e en h esholds d o decla ing a gene as de ec ed. FDR es- ima es we e ob ained by pe mu ing he ows o he pheno ype ma ix and analyzing he pe mu ed da a in exac ly he same way as he o iginal da a (Benjamini and Hochbe g, 1995; S o ey and Tibshi ani, 2003; Xie e al., 2005). In de ail, we compu ed he a e age numbe o genes in he pe mu ed da ase s wi h sco es exceeding a gi en h eshold d( alse-posi i e a es, FPðdÞ), he numbe o genes in he o iginal da a wi h sco es exceeding hesame h esholdd( o al posi i e a es, TPðdÞ)andconside ed he a io o he wo FDRðdÞ¼FPðdÞ TPðdÞ Because o compu a ional bu den, only one pe mu a ion was used wi h he Bayesian educed ank eg ession. Wi h he o he me hods, h ee pe mu a ions we e used. The e o e, he FDR es- ima es a e app oxima e, which we conside o be su icien o ou pu poses, especially because he mos ex eme quan iles a e no conside ed. The numbe s o genes decla ed de ec ed wi h di e en FDR h esholds a e p esen ed in Table 2, and he o e all conco dance o he sco es om di e en me hods is shown in Supplemen a y Figu e S5. In summa y, he esul s show ha he me hod wi h which mos associa ions we e ound was exhaus i e pai wise sea ch and he second-mos powe ul me hod was he Bayesian educed ank eg ession wi h PTVE as he es sco e. These wo me hods also had he highes ag eemen in sco ing genes. CCA applied o es o associa ion be ween indi idual SNPs and he mul i a ia e pheno ype pe o med badly. These esul s a e some- wha di e en om wha has been epo ed be o e by us and o he s (Inouye e al., 2012; Ma inen e al., 2013). In pa icula , he CCA applied o indi idual SNPs pe o med much wo se ha he pai wise es ing, al hough p e iously i has been epo ed o ha e clea ly highe powe . The explana ion is ha he e we applied he me hods o all SNPs, including he a e a ian s. Speci ically, he single-SNP-CCA s a s o o e i when applied o a e a ian s. Fo example, when we in es iga ed he esul s mo e ca e ully, we disco e ed SNPs in which he mino allele was p esen in ew indi iduals, and such indi iduals could be almos pe ec ly iden i ied by a seemingly andom combina ion o pheno ypes, leading o a la ge spu ious canonical co ela ion es sco e (o , equi alen ly, a highly signi ican P- alue). Bayesian educed ank eg ession and spa se CCA we e less a - ec ed by he o e i ing because hey exploi ways o con ol he complexi y o he model, he o me using he in o ma i e p io s, he la e using he c oss- alida ion. Combining in o ma ion o e se e al SNPs using CCA o ela ed me hods has ecen ly been demons a ed o imp o e powe o de ec associa ions unde ce ain condi ions (Ma inen e al., 2013; Tang and Fe ei a, 2012; Zhang e al., 2011). He e we see ha he spa se CCA and also he Bayesian educed ank eg ession, which use mul i-SNP in o ma ion, ha e lowe powe in he genome-wide analysis han he simple pai - wise es ing. The di e ence om he ea lie expe imen s is ha he e he me hods a e applied o geno ype da a wi h a much highe SNP densi y, as ob ained h ough ca e ul impu a ion. The conclusion is ha i he mul i-SNP in o ma ion has al eady been used wi hin he impu a ion p o ocol, he powe o de ec associa ions canno , in gene al, be expec ed o imp o e by using mul i-SNP models. Fo a mo e de ailed compa ison be ween he Bayesian educed ank eg ession and he exhaus i e pai wise es ing, Supplemen a y Figu e S6 shows a Q-Q plo o he sco es (bo h PTVE and PTVE- a e) agains he expec ed sco es ha ha e been ob ained by pe mu a ion. The Supplemen a y Figu e S6 also shows genes de ec able by he simple exhaus i e pai wise sea ch. Bo h Q-Q plo s, bu PTVE in pa icula , indica e an excess o la ge es sco es, e lec ing he ac ha a leas some ue associa ions a e de ec ed by he model. We u he see ha al hough wi h PTVE- a e sco e ewe genes can be de ec ed han wi h PTVE, none o he op-sco ing genes om PTVE- a e a e lagged by he pai wise app oach, making PTVE- a e an a ac - i e sco e o mining associa ions ha migh be missed by he s anda d me hod. 4 DISCUSSION We ha e p esen ed a new s a is ical me hod o in es iga ing associa ions in Genome-wide associa ion s udy (GWAS) da ase s wi h mul i a ia e pheno ypes. The me hod can combine in o - ma ion o e mul iple SNPs, making i pa icula ly sui able o s udying a e a ian s in he high-dimensional pheno ype se ing. Fo his se up, no me hods known o he au ho s ha e been Table 2. Powe compa ison o he di e en me hods Me hod FDR ¼0FDR¼0.1 FDR ¼0.2 FDR ¼0.4 No el PTVE 36 55 103 305 167 PTVE- a e33366 Pai wise 103 176 243 651 300 CCA,singleSNP77 71111 Spa se CCA 37 50 66 117 51 No e: The able shows he numbe s o gene-me abolome associa ions ha had alse disco e y a e below he speci ied h eshold. The las column shows he numbe o pu a i e no el associa ions wi hin genes wi h FDR ¼0.4 a e emo ing he known associa ions as desc ibed in he main ex . 2032 P.Ma inen e al. p esen ed be o e. Ou me hod is based on es ima ing he p opo - ion o o al a iance o he pheno ypes ha is explained by he SNPs unde conside a ion. Fo his pu pose, we ha e de i ed a Bayesian o mula ion o he educed ank eg ession model, which enables us o inco po a e ou knowledge o he expec ed e ec s sizes in he analysis. We used he new me hod o analyze a eal GWAS da ase wi h a mul i a ia e lipop o ein pheno ype. Two no el loci no p e i- ously associa ed wi h he pheno ype we e disco e ed and epli- ca ed in an analysis combining wo addi ional da ase s. Fu he mo e, wo mo e loci we e ound ha eplica ed signi i- can ly in one bu no in he o he es da ase . Possible easons o he lack o success in eplica ing he indings in bo h he da ase s include he ollowing: (i) he associa ions we e alse posi- i e in he i s place, (ii) he associa ions in ol ed a e a ian s ha we e no p esen in su icien numbe s o see he e ec s, (iii) he impu a ion accu acy o he a e a ian s was no su icien in all da ase s, (i ) he pheno ype da a, al hough p ep ocessed in exac ly he same way wi h all he da ase s, ha e no been ully equi alen . Fo example, he scaling o he pheno ypes du ing p ep ocessing has been done using ac o s no exac ly equal, and he pa ame e s lea ned in one da a may hus no ep esen he e ec s adequa ely in ano he da a. Only u he s udies will help o dis inguish be ween he al e na i e explana ions. Fo doing in e ence wi h he model, he cu en implemen a- ion uses MCMC sampling, he compu a ion ime o which is app oxima ely hal an hou pe gene on a 2.3 GHz p ocesso . Thus, analyzing all human genes equi es a clus e compu e o pa allelize he compu a ions o e he genes. Analy ical app oxi- ma ions, such as he a ia ional o Laplace app oxima ions, e.g. Bishop (2006), could be used o speed up he compu a ions and, based on ou expe imen s wi h he cu en me hod, a e wo h doing in he u u e. Al e na i e ways o use he model migh also be conside ed. Fo example, ocusing he analysis on a ian s wi h a p edic ed unc ion could imp o e powe o de ec associa ions and lessen he compu a ional bu den. As an- o he example, we ha e used 0.01 as he h eshold o de ining he a e a ian s when compu ing hei impac on he pheno- ypes. Resul s based on di e en h esholds could eadily be ex- ac ed om he ou pu o a single MCMC un and a e likely o highligh di e en se s o genes. Funding: A comple e lis o unding is gi en in he Supplemen a y Ma e ial. Con lic o In e es : none decla ed. REFERENCES Acke mann,M. e al. (2013) Impac o na u al gene ic a ia ion on gene exp ession dynamics. PLoS Gene .,9, e1003514. Bansal,V. e al. (2010) S a is ical analysis s a egies o associa ion s udies in ol ing a e a ian s. Na . Re . Gene .,11, 773–785. Benjamini,Y. and Hochbe g,Y. (1995) Con olling he alse disco e y a e: a p ac- ical and powe ul app oach o mul iple es ing. J. R. S a . Soc. B Me hodol.,57, 289–300. Bha acha ya,A. and Dunson,D. (2011) Spa se Bayesian in ini e ac o models. Biome ika,98, 291–306. Bishop,C.M. (2006) Pa e n Recogni ion and Machine Lea ning.Sp inge ,New Yo k. Fe ei a,M.A. and Pu cell,S.M. (2009) A mul i a ia e es o associa ion. Bioin o ma ics,25, 132–133. Fusi,N. e al. (2012) Join modelling o con ounding ac o s and p ominen gene ic egula o s p o ides inc eased accu acy in gene ical genomics s udies. PLoS Compu . Biol.,8,e1002330. Gelman,A. e al. (2004) Bayesian Da a Analysis. 2nd edn. Chapman & Hall/CRC, Boca Ra on, FL. Geweke,J. (1996) Bayesian educed ank eg ession in econome ics. J. Econom.,75, 121–146. Hammond,P. and Su ie,M. (2012) La ge-scale objec i e pheno yping o 3D acial mo phology. Hum. Mu a .,33, 817–825. Ho elling,H. (1936) Rela ions be ween wo se s o a ia es. Biome ika,28, 321–377. Howie,B. e al. (2011) Geno ype impu a ion wi h housands o genomes. G3 (Be hesda),1, 457–470. Howie,B.N. e al. (2009) A lexible and accu a e geno ype impu a ion me hod o he nex gene a ion o genome-wide associa ion s udies. PLoS Gene .,5, e1000529. Inouye,M. e al. (2012) No el loci o me abolic ne wo ks and mul i- issue exp es- sion s udies e eal genes o a he oscle osis. PLoS Gene .,8, e1002907. Ke unen,J. e al. (2012) Genome-wide associa ion s udy iden i ies mul iple loci in luencing human se um me aboli e le els. Na . Gene .,44, 269–276. Ma inen,P. e al. (2013) Genome-wide associa ion s udies wi h high-dimensional pheno ypes. S a . Appl. Gene . Mol. Biol.,12, 413–431. Mo gen hale ,S. and Thilly,W.G. (2007) A s a egy o disco e genes ha ca y mul i-allelic o mono-allelic isk o common diseases: a coho allelic sums es (CAST). Mu a . Res.,615, 28–56. Mo is,A.P. and Zeggini,E. (2010) An e alua ion o s a is ical app oaches o a e a ian analysis in gene ic associa ion s udies. Gene . Epidemiol.,34, 188–193. O’Reilly,P.F. e al. (2012) Mul iPhen: join model o mul iple pheno ypes can in- c ease disco e y in GWAS. PLoS One,7,e34861. Pa khomenko,E. e al. (2009) Spa se canonical co ela ion analysis wi h applica ion o genomic da a in eg a ion. S a . Appl. Gene . Mol. Biol.,8, 1–34. P ice,A.L. e al. (2006) P incipal componen s analysis co ec s o s a i ica ion in genome-wide associa ion s udies. Na . Gene .,38, 904–909. Rai aka i,O.T. e al. (2008) Coho p o ile: he Ca dio ascula Risk in Young Finns S udy. In . J. Epidemiol.,37, 1220–1226. Ran akallio,P. (1969) G oups a isk in low bi h weigh in an s and pe ina al mo - ali y. Ac a Paedia . Scand.,193 (Suppl. 193), 1þ. Saba i,C. e al. (2008) Genome-wide associa ion analysis o me abolic ai s in a bi h coho om a ounde popula ion. Na . Gene .,41,35–46. Soininen,P. e al. (2009) High- h oughpu se um NMR me abonomics o cos -e ec i e holis ic s udies on sys emic me abolism. Analys ,134, 1781–1785. S egle,O. e al. (2010) A Bayesian amewo k o accoun o complex non-gene ic ac o s in gene exp ession le els g ea ly inc eases powe in eQTL s udies. PLoS Compu . Biol.,6,e1000770. S o ey,J.D. and Tibshi ani,R. (2003) S a is ical signi icance o genomewide s udies. P oc. Na l Acad. Sci. USA,100, 9440–9445. Suh e,K. e al. (2011) Human me abolic indi iduali y in biomedical and pha ma- ceu ical esea ch. Na u e,477, 54–60. Tang,C.S. and Fe ei a,M.A. (2012) A gene-based es o associa ion using canon- ical co ela ion analysis. Bioin o ma ics,28, 845–850. Teslo ich,T.M. e al. (2010) Biological, clinical and popula ion ele ance o 95 loci o blood lipids. Na u e,466, 707–713. The Global Lipids Gene ics Conso ium. (2013) Disco e y and e inemen o loci associa ed wi h lipid le els. Na . Gene .,45, 1274–1283. The 1000 Genomes P ojec Conso ium. (2012) An in eg a ed map o gene ic a i- a ion om 1,092 human genomes. Na u e,491, 56–65. Thompson,J.R. e al. (2011) The me a-analysis o genome-wide associa ion s udies. B ie . Bioin o m.,12, 259–269. Tukiainen,T. e al. (2012) De ailed me abolic and gene ic cha ac e iza ion e eals new associa ions o 30 known lipid loci. Hum. Mol. Gene .,21, 1444–1455. Va iainen,E. e al. (2010) Thi y- i e-yea ends in ca dio ascula isk ac o s in Finland. In . J. Epidemiol.,39, 504–518. Va iku i,S. e al. (2012) He i abili y and gene ic co ela ions explained by common SNPs o me abolic synd ome ai s. PLoS Gene .,8, e1002637. Waaijenbo g,S. e al. (2008) Quan i ying he associa ion be ween gene exp essions and dna-ma ke s by penalized canonical co ela ion analysis. S a . Appl. Gene . Mol. Biol.,7,1–29. 2033 Gene-me abolome associa ions Wi en,D.M. and Tibshi ani,R. (2009) Ex ensions o spa se canonical co ela ion analysis wi h applica ions o genomic da a. S a . Appl. Gene . Mol. Biol.,8, A icle 28. Wu,M.C. e al. (2011) Ra e- a ian associa ion es ing o sequencing da a wi h he sequence ke nel associa ion es . Am. J. Hum. Gene .,89, 82–93. Xie,Y. e al. (2005) A no e on using pe mu a ion-based alse disco e y a e es ima es o compa e di e en analysis me hods o mic oa ay da a. Bioin o ma ics,21, 4280–4288. Zhang,F. e al. (2011) Mul ilocus associa ion es ing o quan i a i e ai s based on pa ial leas -squa es analysis. PLoS One,6, e16739. 2034 P.Ma inen e al.