scieee Open visual document viewer

Allele coding in genomic evaluation

Strandén, Ismo,Christensen, Ole F

Full text

RESEARCH Open Access Allele coding in genomic e alua ion Ismo S andén 1* and Ole F Ch is ensen 2 Abs ac Backg ound: Genomic da a a e used in animal b eeding o assis gene ic e alua ion. Se e al models o es ima e genomic b eeding alues ha e been s udied. In gene al, wo app oaches ha e been used. One app oach es ima es he ma ke e ec s i s and hen, genomic b eeding alues a e ob ained by summing ma ke e ec s. In he second app oach, genomic b eeding alues a e es ima ed di ec ly using an equi alen model wi h a genomic ela ionship ma ix. Allele coding is he me hod chosen o assign alues o he eg ession coe icien s in he s a is ical model. A common allele coding is ze o o he homozygous geno ype o he i s allele, one o he he e ozygo e, and wo o he homozygous geno ype o he o he allele. Ano he common allele coding changes hese eg ession coe icien s by sub ac ing a alue om each ma ke such ha he mean o eg ession coe icien s is ze o wi hin each ma ke . We call his cen e ed allele coding. This s udy conside ed e ec s o di e en allele coding me hods on in e ence. Bo h ma ke -based and equi alen models we e conside ed, and es ic ed maximum likelihood and Bayesian me hods we e used in in e ence. Resul s: Theo e ical de i a ions showed ha pa ame e es ima es and es ima ed ma ke e ec s in ma ke -based models a e he same i espec i e o he allele coding, p o ided ha he model has a ixed gene al mean. Fo he equi alen models, he same esul s hold, e en hough di e en allele coding me hods lead o di e en genomic ela ionship ma ices. Calcula ed genomic b eeding alues a e independen o allele coding when he es ima e o he gene al mean is included in o he alues. Reliabili ies o es ima ed genomic b eeding alues calcula ed using elemen s o he in e se o he coe icien ma ix depend on he allele coding because di e en allele coding me hods imply di e en models. Finally, allele coding a ec s he mixing o Ma ko chain Mon e Ca lo algo i hms, wi h he cen e ed coding being he bes . Conclusions: Di e en allele coding me hods lead o he same in e ence in he ma ke -based and equi alen models when a ixed gene al mean is included in he model. Howe e , eliabili ies o genomic b eeding alues a e a ec ed by he allele coding me hod used. The cen e ed coding has some nume ical ad an ages when Ma ko chain Mon e Ca lo me hods a e used. Backg ound The e has been g owing in e es in he use o ma ke - based models [1] in ecen yea s. In s udies using hese models, desc ip ions o he e ec o allele coding sys em on in e ence and compu a ions a e o en ague o missing. By allele coding, we e e o he coe icien s in he ma ke ma ix o ma ke -based models. Coe icien s, commonly used o allele coding o a ma ke is 0 when he indi idual is homozygous o he i s allele, 1 when he indi idual is he e ozygous, and 2 when he indi idual is homozygous o he second allele. Depending on which o he alleles has been chosen as he i s allele, he coe icien s a e di e en . Thus, his allele coding me hod does no gi e unique eg ession coe icien s. The e a e o he allele coding me hods such as he one ha use coe icien s -1, 0, and 1 ins ead o 0, 1, and 2, espec i ely. Di e en allele coding me hods a ec coe i- cien s in he s a is ical models bu hey do no seem o change he amoun o in o ma ion o s a is ical in e ence. Hence, one would expec ha he use o di e en allele coding me hods would lead o he same in e ence. How- e e , allele coding can be o i al impo ance in compu a- ions. Fi s , con e gence o i e a i e me hods such as Ma ko chain Mon e Ca lo (McMC) me hods o en used in Bayesian in e ence can be assumed o be a ec ed by he allele coding me hod used because di e en allele coding me hods change he co ela ion s uc u e be ween ma ke * Co espondence: [email p o ec ed]i Full lis o au ho in o ma ion is a ailable a he end o he a icle S andén and Ch is ensen Gene ics Selec ion E olu ion 2011, 43:25 h p://www.gsejou nal.o g/con en /43/1/25 Gene ics Selec ion E olu ion © 2011 S andén and Ch is ensen; licensee BioMed Cen al L d. This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion License (h p://c ea i ecommons.o g/licenses/by/2.0), which pe mi s un es ic ed use, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly ci ed. e ec s. Second, equi alen models ha e become popula in animal b eeding [2,3]. An impo an concep in hese me hods is he genomic ela ionship ma ix. Di e ences in allele coding will yield di e en genomic ela ionship ma ices [4]. Thus, some elemen s o he in e se o he coe icien ma ix can be di e en and, consequen ly, eliabili ies may be di e en . We in es iga ed e ec s o di e en allele coding me h- ods using heo e ical de i a ions and a p ac ical example when es ic ed maximum likelihood (REML) o Baye- sian in e ence is used. E ec s on pa ame e es ima es, eliabili ies, and McMC compu a ions we e s udied. We conside ed ma ke models and hei equi alen b eeding alue models. Me hods Genomic ma ke model Le us conside a uni a ia e linea mixed e ec s model o genomic ma ke e ec es ima ion, e.g. [1], y =1μ+Zg +e , (1) whe e yis n× 1 ec o o obse a ions, μis he gen- e al mean, Zis a n×mma ix con aining a column o each ma ke locus, gis a m× 1 ec o o andom SNP ma ke e ec s, and eis a andom esidual ec o . The e can be o he ixed o andom e ec s in he model bu hei inclusion does no change he ollowing de i a ions. Allele coding me hods The e a e se e al al e na i es o coding he coe icien s in he Zma ix. Fou allele coding sys ems a e consid- e ed. A simple ans o ma ion can be made om one allele coding sys em o ano he . Ou basic allele coding sys em coun s he numbe o copies o one o he alleles. Depending on which o he alleles is coun ed, he ma ix can be di e en . In he allele coding sys em 012, he numbe o copies o he less equen allele is coun ed. Thus, he coe icien is 0 i he indi idual is homozygous o he mo e equen allele, 1 i i is he - e ozygous, o 2 i i is homozygous o he less equen allele. In his case, he Zma ix o he basic allele cod- ing sys em 012 is deno ed by Z 0 . A gene al o m o he allele coding ans o ma ion om he basic allele coding sys em is Z0−1n  m whe e m is a m× 1 ec o . This allows many ypes o allele coding me hods. No e ha he ans o ma ion keeps dis- ances be ween allele codes wi hin a ma ke he same. So, he coding 0,1,2 can be changed o - 1,0,1 o 0.5, 1.5, 2.5 by his ans o ma ion, bu no o -10,0,10. We de ine he cen e ed allele coding sys em as Z c = Z 0 -P c , whe e each column o he ma ix P c con ains he a e age allele coun o he co esponding ma ke column. Thus, summing alues in each column will gi e a ec o o ze os, i.e., 1 n (Z0−Pc ) is a ec o o ze os. Fo he cen e ed allele coding sys em, we ha e m=1 n Z 01 n , i.e., Zc=Z0−1 n 1n1 nZ 0 .No e ha m /2 gi es he allele equencies o he ma ke s in he da a. The allele coding ans o ma ion allows shi s in he allele codes. The 101 allele coding sys em is such ha - 1 is assigned o he geno ype homozygous o he mo e equen allele, 0 o he he e ozygous indi idual, and 1 o he indi idual homozygous o he less equen allele. Fo he 101 allele coding sys em, we ha e m =1 m .The 101 allele coding sys em is equal o he cen e ed allele coding sys em when all allele equencies a e equal o 0.5. In he ollowing, he de i a ions will use he gene al allele coding ans o ma ion. The ma ix Z 0 is unique. Howe e , in gene al he decision on which o he alleles o coun is a bi a y. Le he 210 allele coding sys em be such ha he mo e equen allele is coun ed. Then he Zma ix o he 210 allele coding sys em can be calcu- la ed om he 012 coding ma ix by 21n1 m −Z 0 whe e 1 n is n× 1 ec o o ones. The 210 allele coding sys em is he opposi e o he 012 allele coding sys em bu esul s in his pape apply o he 210 allele coding sys- em as well, wi h some modi ica ions men ioned sepa a ely. Di e en allele coding me hods imply di e en models (1) and di e en models may lead o di e en pa ame e es ima es. Howe e , he in e ence conside ed in his pape is no a ec ed by allele coding, as we demons a e below. In e ence in ma ke -based model Va iance componen es ima ion by es ic ed maximum likelihood (REML), p edic ion o andom e ec s, and Bayesian in e ence a e all based on he likelihood a e ma ginaliza ion o he ixed e ec s, i.e., he ixed e ec s ha e been in eg a ed ou . Bayesian in e ence equi es e en mo e ma ginaliza ion. In o de o show ha in e - ence is no a ec ed by allele coding, i is su icien o show ha he likelihood a e in eg a ing ou he gene al mean is he same i espec i e o allele coding. The ol- lowing de i a ion makes no assump ions on he p io densi ies o ma ke e ec s. Thus, he esul s apply o many models, including BLUP, BayesA and BayesB in [1]. The ma ginal likelihood o he mixed e ec s model is p(y|g,θ)= ∞ ∫ − ∞ p(y|μ,g,θ)dμ , whe e p(y|μ,g,θ) is he condi ional densi y o y, o en a Gaussian densi y, and θcon ains all pa ame e s in he dis ibu ion o e, o en only he esidual a iance S andén and Ch is ensen Gene ics Selec ion E olu ion 2011, 43:25 h p://www.gsejou nal.o g/con en /43/1/25 Page 2 o 11 pa ame e σ 2 e . Using he ans o ma ion esul (7) in Appendix A and a change o in eg a ion a iable μ0=μ−  mg , we can w i e p(y|g,θ)= ∞ ∫ −∞ p0(y|μ−  mg,g,θ)d μ = ∞ ∫ −∞ p0(y|μ0,g,θ)dμ0 =p0 ( y|g,θ ) , whe e p 0 deno es he 012 allele coding sys em. Hence, he ma ginal likelihood does no depend on allele cod- ing, a p ope y used in he ollowing de i a ions. The REML-likelihood is de ined when gand ea e mul i a ia e Gaussian dis ibu ed, and equals L ( θ,η ) =∫p ( y|g,θ ) p ( g|η ) dg , whe e hcon ains all pa ame e s in he dis ibu ion o g, commonly only he gene ic ma ke a iance pa ame e σ 2 g . This likelihood is independen o allele coding, and, hence, REML pa ame e es ima ion is independen o allele coding. No e ha maximum likelihood es ima ion is based on L(μ,θ,h)=∫p(y|μ,g,θ)p(g|h)dgand is a ec ed by allele coding because, in his case, he gene al mean is no in eg a ed ou . BLUP es ima ion o ma ke e ec s gassumes ha he a iance pa ame e s (θ,h) a e known. The condi ional dis ibu ion p(g|y,θ,h)=p(y|g,θ)p(g|h)/L(θ,h)is independen o allele coding. Hence, BLUP ĝand asso- cia ed unce ain ies do no depend on he allele coding. In Bayesian in e ence, he join pos e io a e in e- g a ing ou μis p(g,θ,η|y)=∫p(μ,g,θ,η|y)dμ =p(y|g,θ)p(g|η)p(θ,η)  p(y), whe e p(θ,h) is he join p io o he pa ame e s, and he denomina o p(y) is he in eg al o he nume a- o . All e ms in he nume a o a e independen o allele coding, and by ma ginaliza ion p(y) sa is ies he same. Hence, p(g,θ,h|y) does no depend on allele coding. The gene al in e cep μis, howe e , no independen o allele coding. Fo simplici y o he a gumen , we assume ha pa ame e s (θ,h)a eknown,andomi showing hese alues. Acco ding o he ans o ma ion esul (8) in Appendix A and a change o in eg a ion a iable μ0=μ−  mg , he condi ional expec a ion o he gene al mean is ˆμ=∫μ∫p(μ,g|y)dgdμ =∫∫μp0(μ−  mg,g|y)dμdg =∫∫(μ0+  mg)p0(μ0,g|y)dμ0dg =∫μ0p0(μ0|y)dμ0+∫  mgp0(g|y)d g =ˆμ0+  m ˆ g, (2) whe e p 0 deno es densi y o he 012 allele coding, and ˆ μ0 is he condi ional expec a ion o he gene al mean when using he 012 allele coding. Thus, he gen- e al mean es ima e is di e en by allele coding when  m ˆ g is no ze o. When gand ea e mul i a ia e Gaus- sian dis ibu ed, he condi ional expec a ions ĝand ˆ μ equal he BLUP and BLUE es ima es, espec i ely. Finally, he in e ence is indi e en o he allele being coun ed. This is demons a ed by s udying he cen e ed coding sys em and assuming ha allele in he i s ma - ke is coun ed in he opposi e way, i.e., he i s column in Zis minus he i s column in Z c ,o z 1 =-z c1 .We see ha Z g =Zc ˜ g whe e he en ies in ˜ g a e equal o he en ies in g, excep o he i s en y which equals minus he i s en y in g.Sincegand ˜ g ha e he same dis ibu ion, hese wo models a e equi alen . Genomic b eeding alues Es ima ing b eeding alues In b eeding alue e alua ion, he main in e es is in es i- ma ion o genomic b eeding alues o he geno yped animals. In o he wo ds, es ima ion o ˆ a=Zˆ g whe e ĝ a esolu ions o hema ke e ec sbyama ke -based model like model (1). Because he ma ke e ec solu- ions a e he same o di e en allele coding sys ems, he es ima ed genomic b eeding alues a e di e en due o di e ences in he coe icien ma ix Z. Allele coding does no , howe e , change ela i e di - e ences be ween he es ima ed genomic b eeding alues, because Zˆ g−Z0ˆ g=−1n(  m ˆ g ) shows ha hey a e jus shi ed by a cons an . Le us de ine comple e genomic b eeding alues as ˆ ad=1nˆμ+Zˆ g . Subs i u ing Z=Z0−1n  m and using equa ion (2) we ob ain ˆ ad=1nˆμ+(Z0−1n  m)ˆ g =1n(ˆμ−  mˆ g)+Z0ˆ g =1nˆμ0+Z0ˆ g . Consequen ly, he es ima ed comple e b eeding alues ˆ a d a e he same i espec i e o allele coding. Equi alen model and allele coding Assume ha he ma ke e ec s ha e a Gaussian dis i- bu ion g∼N(0,Imσ2 g) whe e I m is an m×miden i y ma ix. The b eeding alues a=Zg can be calcula ed di ec ly wi hou es ima ing ĝby he model [2,3] y =1nμ+a+e , whe e he b eeding alues ha e p io densi y o a∼N(0,ZZ’σ2 g) .O en heco a iancema ixo he b eeding alues is scaled by a alue such as p=2 m  i =1 pi(1 −pi ) whe e p i is he allele equency o ma ke i. Then, he b eeding alues ha e a p io densi y o a∼N(0,Gσ 2 a) whe e he genomic ela ionship ma ix S andén and Ch is ensen Gene ics Selec ion E olu ion 2011, 43:25 h p://www.gsejou nal.o g/con en /43/1/25 Page 3 o 11 is G=ZZ’ 1 p and gene ic a iance is σ2 a=σ2 g p . Assuming ha he esidual dis ibu ion is e~N(0,R), hen he mixed model equa ions o he equi alen model a e  1 nR− 1 1n1nR− 1 R−11nR−1+G−1/σ2 aˆμ ˆ a  =  1 nR−1y R−1y  . (3) The b eeding alue solu ions â om hese mixed model equa ions (3) a e he same as he genomic b eed- ing alues calcula ed by ˆ a=Zˆ g whe e ˆ g a e ma ke e ec s es ima ed by he ma ke -based model (1). The e o e, he conclusion o he ma ke -based models abou ela i e di e ences be ween genomic b eeding alues being una ec ed by allele coding is also ue o he equi alen model, al hough di e en allele coding me hods lead o di e en genomic ela ionship ma ices. Simila ly, a iance componen es ima ion by REML and Bayesian me hods a e una ec ed by he allele coding due o equi a- lence o models wi h he ma ke -based models. The mixed model equa ions in (3) a e no well-de ined when he genomic ela ionship ma ix Gis singula . Howe e , mixed model equa ions no equi ing an in e ible Gma ix do exis ; see page 48 in [5]. The genomic ela ionship ma ix Gcan be singula o se - e al easons. Fo example, he e can be iden ical wins o clones ha ha e he same geno ypes. In addi ion, o he cen e ed allele coding sys em he genomic ela ion- ship ma ix is Gc=ZcZ c .Thelas owo Zc=(I− 1 n 1n1 n)Z 0 is equal o he sum o all he o he ows. Hence, Z c is no o ull ank, and G c is singula . P edic ion e o a iances and eliabili ies Gaussian models a e o en used in p ac ical genomic e alua ion o animals. In hese models, eliabili ies o es ima ed b eeding alues âa e calcula ed using ele- men s o he in e se o he mixed model equa ions such as (3). Reliabili y o â i is 2 i=1−PEVi σ2 a Gii , (4) whe e PEV i is he p edic ion e o a iance, i.e., Va (a i | y), o animal i, and G ii is he diagonal elemen o animal i in he genomic ela ionship ma ix G; e.g. [6], p. 51 in [7]. The p edic ion e o a iance o animal iis he diagonal elemen o he in e se o he coe icien ma ix o mixed model equa ions (3) o animal i. Al e na i ely, PEV = Va (a|y) =Va (Zg|y) =ZVa (g|y)Z ’ =ZCgZ’ , (5) whe e C g is he genomic ma ke e ec subma ix in he in e se o he coe icien ma ix o he mixed model equa ion o ma ke -based model (1) (see Appendix B). The subma ix C g =Va (g|y) is he same i espec i e o he allele coding me hod used as shown in he chap e on in e ence on ma ke -based models. Because he coe - icien ma ix Zis di e en depending on allele coding, PEV is also di e en depending on allele coding. Conse- quen ly, he eliabili y o âdepends on allele coding. Mo e gene ally, o any o he models conside ed in his pape , a=Zg whe e p(g|y) is independen o allele coding and Zdepends on allele coding. The e o e, he dis ibu ion p(a|y) and, in pa icula , he a iance-co - a iance ma ix Va (a|y) and eliabili ies o âdepend on allele coding. Thecomple eb eeding aluedis ibu ionp(a d |y) does no depend on allele coding, unlike PEV associa ed wi h â. The p oo is based on he demons a ion ha when applying any unc ion , he expec a ions a e inde- penden o he allele coding sys em, E[ (ad)|y] =∫∫ (1nμ+Zg)p(μ,g|y)dμdg =∫∫ (1n(μ−  mg)+Z0g) ×p0(μ−  mg,g|y)dμdg =∫∫ (1nμ0+Z0g)p0(μ0,g|y)dμ0d g =E 0[ ( ad ) | y ], whe e E 0 is he expec a ion when using he basic allele coding me hod. The e o e, he a iance-co a iance ma ix Va (a d |y) and all highe o de momen s o he dis ibu ion a e independen o allele coding. Howe e , he esul does no p o ide ac ual o mulas o he momen s. A closed o m o mula o he a iance-co a iance ma ix is de i ed o a Gaussian model. Assume g∼N(0,Imσ2 g) and e~N(0,R). Fo his model, in Appendix B we ob ain ha Va (ad|y)= 1n1 n +(In− 1n1 n R−1)ZCgZ’(In− R−11n1 n ) , whe e =1/(1 n R− 1 1n ) and C g=[Z’(R−1− R−11n1 nR−1)Z+Im/σ2 g ]− 1 .Asdemon- s a ed ea lie in his sec ion, his a iance-co a iance ma ix is independen o allele coding. When R =Inσ 2 e , he a iance-co a iance ma ix simpli ies o Va (ad|y) =1n1 nσ2 e/n+Zc(Z cZc/σ2 e+Im/σ2 g )−1Z c , (6) whe e Z c is based on he cen e ed allele coding me hod. The diagonal elemen s in (6) a e di e en om S andén and Ch is ensen Gene ics Selec ion E olu ion 2011, 43:25 h p://www.gsejou nal.o g/con en /43/1/25 Page 4 o 11 he PEV i s in (4) because hey con ain unce ain y abou he unknown mean μas well. Fo he comple e b eeding alues a d ,weha eshown abo e ha p edic ion e o a iances a e independen o allelecodingandweha ep o ideda o mula o he Gaussian model. Reliabili ies o â d , howe e , can no be de ined in a meaning ul way. Subs i u ing he diagonal elemen s om (6) o he PEV i s in (4) is no app op ia e since he denomina o in (4) is Va (a i )no Va (a d ) ii . The denomina o in he eliabili y o mula should con- ain he ma ginal (uncondi ional) a iance Va (a d ) ii =∫ Va (μ+a i |μ)dμ, bu his a iance is in ini e. McMC compu a ions Theo e ical con e gence a e The con e gence and mixing o an McMC algo i hm depend on he pa ame iza ion o he model and on he algo i hm used. Theo e ical esul s abou he geome ic a e o con e gence o he s a iona y dis ibu ion o Gibbs sampling algo i hms a e shown in [8], and as men- ioned in [9] his a e also desc ibes he mixing o he algo i hm. Below we show speci ic esul s abou he con- e gence a e o Gibbs sampling algo i hms o simu- la ing om [μ,g|y] in he ma ke -based model (1), whe e gis Gaussian g∼N(0,Imσ2 g) and eis Gaussian e∼N(0,Inσ2 e) . These esul s p o ide some ideas abou mo e gene al models and algo i hms whe e heo e ical esul s canno be ob ained. Sec ion 2.2 in [8] con ains esul s abou a ious Gibbs-sampling schemes o simula ing om a mul i- a ia e dis ibu ion. He e we apply hese esul s (see Appendix C o de ails) o wo ypes o Gibbs upda ing schemes. The i s scheme i e a es be ween upda ing μ and a block o all componen s in g, and will be called he block upda ing scheme he eina e . The second scheme upda es μ,g 1 ,...,g m sequen ially one a a ime, and will be called he single si e upda ing scheme. Fo he block scheme, he con e gence a e is ρ1=n σ2 e ¯zC−1 g¯z , whe e ¯z=Z’1n /n is a m×1 ec o ,and C g=Z’Z/σ2 e+Im/σ 2 g . Fo he single si e scheme, he con e gence a e is ρ2=ρl (L−1 g (D−1¯z¯zn/σ2 e−Ug) , whe e L g is he ma ix con aining he lowe iangle and he diagonal o D -1 C g ,Dis he diagonal o C g ,and U g is he ma ix con aining he uppe iangle o D -1 C g . This single si e Gibbs sampling algo i hm is he s ochas- ic coun e pa o he Gauss-Seidel algo i hm o sol ing he mixed model equa ions, and he con e gence esul s a e simila , see [10]. When he cen e ed allele coding me hod Zc=(In−1n1 n /n)Z0is used, ¯ z is a ec o o ze os and, hence, 1 = 0. The cen e ed allele coding me hod b eaks dependency be ween he gene al mean and gene ic ma ke e ec s, as seen om he a iance-co - a iance ma ix (de i ed in Appendix B o a mo e gen- e al si ua ion) Va (μ,g|y) =σ2 en0 0(Z cZcσ2 e+Im  σ2 g) −1 . Consequen ly, abso p ion o he gene al mean is done wi hou needing o compu e abso p ion explici ly. No e ha , in gene al, his holds only when he esidual a - iance-co a iance ma ix is Iσ 2 e . Fo he block McMC scheme, he con e gence and mixing o he algo i hms a e o he same o de as o non-McMC algo i hms ha simula e di ec ly om he dis ibu ion o in e es . Fo he single si e McMC scheme, he cen e ed allele coding me hod s ill b eaks he dependence be ween he gene al mean and ma ke e ec s, bu as 2 >0 illus a es, he indi idual ma ke e ec s g 1 , ..., g m a e no independen and he McMC samples a e au oco ela ed. Da a and me hods Da a Da a o he XII h QTLMAS wo kshop [11] we e used o illus a e he heo y. The simula ed da a had ou gen- e a ions. In each gene a ion, 15 si es and 150 dams we e selec ed andomly o p oduce he nex gene a ion. Each si e was ma ed o 10 dams and each ma ing p o- duced 10 p ogeny. Thus, he base gene a ion had 165 indi iduals, and he subsequen h ee gene a ions, 1500 indi iduals each. In o al, he analyzed da a had 4665 animals wi h pheno ypes. The simula ed ai had a he - i abili y o 0.30. The da a had 6000 equally spaced SNP ma ke s on six ch omosomes. We dele ed ma ke s ha had a mino allele equency less han 1% among he pheno yped indi iduals and his educed he numbe o ma ke s o 5896. The 012 allele coding me hod was used o make he base da a se . So, he leas equen allele was coun ed. In addi ion, 210, 101, and cen e ed allele coding da a se s we e analyzed. Va iance componen analysis The ma ke -based model (1) wi h common gene ic a - iance was used o analyze he da a: e∼N(0,Iσ2 e) and g∼N(0,Iσ 2 g) .McMCcompu a ionsbyasinglesi e upda ing Gibbs sample we e used o calcula e pos e io mean es ima es o he loca ion (μ,g) and dispe sion (σ2 g ,σ2 e ) pa ame e s. The leng h o he McMC chain was 100000 i e a ions, o which he bu n-in pe iod o 10000 S andén and Ch is ensen Gene ics Selec ion E olu ion 2011, 43:25 h p://www.gsejou nal.o g/con en /43/1/25 Page 5 o 11 was omi ed. E e y en h sample was sa ed, gi ing 9000 sa ed samples. E ec i e sample sizes we e calcula ed o all pa ame e s using he ini ial mono one sequence app oach [12]. The app oach es ima es he numbe o independen samples om he pos bu n-in samples. Theo e ical con e gence a es we e calcula ed o all he allele coding me hods in he single si e and block upda ing schemes. He e, he con e gence a e also desc ibes he mixing o he McMC chain and was measu ed by he co - ela ion be ween successi e McMC samples. Because o he Ma ko p ope y, he k-lag co ela ion is k , whe e is he con e gence a e and kis he lag o dis ance be ween samples. In o de o compa e he heo e ical con e gence a e o he obse ed e ec i e sample size, heo e ical mix- ing was calcula ed ela i e o he 012 allele coding sys em as ollows. Le 012 be he con e gence a e om he 012 allele coding sys em, and he con e gence a es om ano he allele coding sys em. Mixing o he o he allele coding sys em is equal o mixing o he 012 allele coding sys em when e e y k h sample is aken om he McMC samples and 012 = k . Thus, he ela i e mixing is k= log ( 012 )/log( ). Pa ame e es ima ion by REML o bo h he ma ke - based model and he equi alen model and o all ou allele coding me hods was done using so wa e DMU [13]. Bo h AI-REML and EM-REML we e in es iga ed. Fo he equi alen model and he cen e ed allele coding me hod, he singula genomic ela ionship ma ix was modi ied by mul iplying he diagonals by 1.001 in o de o be able in e he ma ix. The e ec o allele coding on con e gence was in es iga ed, and i was checked ha he pa ame e es ima es we e he same. Reliabili ies o genomic b eeding alues Reliabili ies we e calcula ed by di e en allele coding me hods using (4) and elemen s o he in e se o he mixed model equa ions. The a iance componen s we e hose es ima ed by he cen e ed allele coding me hod using he single si e McMC app oach. P edic ion e o a iances we e calcula ed by bo h he ma ke -based model equa ions (5) and equi alen model equa ions (3). Mean, minimum, maximum and s anda d de ia ions o eliabili ies we e calcula ed o all allele coding me hods. Also, co ela ions be ween eliabili ies om all allele coding me hods we e calcula ed. Resul s and Discussion Pos e io mean es ima es o ma ke e ec s and a iance componen s we e almos equal be ween di e en allele coding me hods (Table 1). Co ela ions o es ima es o he ma ke e ec s be ween allele coding me hods we e highe han 99.98%. Only he gene al mean (μ)hada di e en es ima e, as expec ed. The es ima ed a iance componen s ag eed well wi h hose used o simula e he da a. Addi i e gene ic a iance was ˆσ2 a=ˆσ2 g p=1.55 6 , whe e p=2 m  i =1 pi(1 −pi) = 2323.0 0 ,andp i a e he obse ed allele equencies in he e e ence da a. Thus, he he i abili y es ima e was ˆ h2=ˆσ2 a /( ˆσ2 a +ˆσ2 e )=0.3 4 , compa ed o he simula ed alue o 0.30. E ec i e sample sizes di e ed depending on allele coding me hod. The cen e ed allele coding me hod had he bes mixing, and he 210 allele coding me hod had he wo s (Table 2). In pa icula , he inc ease in e ec- i e sample size was la ges o he gene al mean. Fo he cen e ed allele coding me hod, he gene al mean was independen om he ma ke e ec g, which led o excellen mixing o his pa ame e . In gene al, he ma - ke e ec s showed excellen mixing. Wi h all allele cod- ing me hods, e ec i e sample sizes we e a leas 5500 o all ma ke e ec s, and on a e age we e equal o abou 8800. Theo e ical con e gence a es (Table 3) displayed he same esul s as he e ec i e sample sizes discussed abo e. No e ha ou Gibbs sample used single si e upda es o all pa ame e s. Fo he single si e upda ing algo i hm, he 210 allele coding sys em was p edic ed o need 5.64 imes mo e i e a ions han he 012 allele cod- ing sys em. The numbe o e ec i e samples o he gene al mean pa ame e (μ) was 5.11 imes bigge o he 012 han o he 210 allele coding sys em. These ig- u es we e 0.48 and 0.48 o he 101 allele coding sys em, and 0.070 and 0.0051 o he cen e ed allele coding sys- em. Theo e ical con e gence a es o he block Gibbs sample showed he same pa e n as o he single si e upda e (Table 3). Su p isingly, he block Gibbs sample was p edic ed o be wo se han he single si e Gibbs sample o all allele coding sys ems excep o he cen e ed allele coding sys- em. Howe e , i is well known in he li e a u e ha block-upda ing schemes may some imes be wo se han Table 1 Pos e io means o selec ed pa ame e s by allele coding Allele coding Pa ame e 012 210 101 cen e ed μ1.698 0.801 1.083 1.359 σ2 g 6.700 × 10 -4 6.703 × 10 -4 6.691 × 10 -4 6.698 × 10 -4 σ2 e 2.996 2.996 2.996 2.996 Table 2 E ec i e sample sizes in McMC compu a ions by allele coding Allele coding Pa ame e 012 210 101 cen e ed μ46 9 96 8961 σ2 g 723 330 1001 1701 σ2 e 7814 6861 7661 7720 S andén and Ch is ensen Gene ics Selec ion E olu ion 2011, 43:25 h p://www.gsejou nal.o g/con en /43/1/25 Page 6 o 11 single si e upda ing schemes, o examples see [8]. The excellen con e gence a e o he block Gibbs sample wi h cen e ed allele coding was expec ed because, in his case, he Gibbs sample is equal o Mon e Ca lo sampling. Table 4 shows con e gence o REML in pa ame e es i- ma ion. Fo he ma ke -based model, he con e gence was independen o he allele coding sys em, whe eas o he equi alen model, he con e gence was as es o he cen e ed coding sys em and slowes o he 210 coding sys em, al hough he di e ences we e small. The pa a- me e es ima es ob ained (Table 5) we e he same, wi h he excep ion o σ 2 e o he cen e ed coding sys em. The di e ence is due o he need o make he genomic ela- ionship ma ix G c o be ull ank by mul iplica ion o he diagonals by 1.001. In summa y, REML pa ame e es i- ma ion is only sligh ly a ec ed by allele coding. Reliabili ies we e a ec ed depending on he allele coding me hod used. Di e ences we e la ge (Table 6). A e age eliabili ies anged om 0.37 wi h he 210 allele coding me hod o 0.80 wi h he cen e ed allele coding me hod. The cen e ed allele coding me hod ga e highe eliabili ies han achie ed by any o he o he allele coding me hods. Reliabili ies calcula ed by di e en allele coding me hods we e also di e en as judged by he co ela ion o each o he (Table 7). Reliabili ies calcula ed by he ma ke - based model and he equi alen model app oaches we e equal wi hin he nume ical ounding e o . The obse ed la ge di e ences in eliabili ies using di - e en allele coding me hods can be explained by di e - ences in es ima ion unce ain y. Di e en allele coding sys ems ha e di e en Zma ices. Conside i s he 012 and 210 allele coding me hods. The 012 allele coding sys- em has a 0 coe icien when he indi idual is homozygous o he mo e equen allele while he 210 allele coding sys em has a coe icien o 2 ins ead. In he ma ke -based model, unce ain y o he in e se o he coe icien ma ix is he same i espec i e o allele coding me hod. Reliabili y is calcula ed by mul iplying he ma ke unce ain y by he Zma ix (5). Consequen ly, unce ain y is less in he 012 allele coding sys em han in he 210 allele coding sys em because he mo e equen homozygous allele mul iplies he ma ke solu ion and unce ain y by ze o. Thus, homo- zygous geno ypes o he mo e equen allele do no inc ease unce ain y when es ima ing genomic b eeding alues. Thus, he 012 allele coding sys em will yield highe eliabili ies han he 210 allele coding sys em, as was obse ed. This a gumen can be gene alized as ollows. In he genomic model conside ed, unce ain y o a geno ype in es ima ing genomic b eeding alue is alued ela i e o a chosen base geno ype. The u he away an obse ed geno ype is om he base geno ype, he la ge he coe i- cien in absolu e alue in he Zma ix and he highe he unce ain y in genomic b eeding alue. In he 012 allele coding sys em, he base geno ype is homozygous o he mo e equen allele, while in he 210 allele coding, i is homozygous o he less equen allele. In he 101 allele coding sys em he base geno ype is he he e ozygo e. The highe he numbe o he e ozygous indi iduals is in he da a he smalle will he unce ain y be, i.e., he highe he eliabili y will be o he 101 allele coding sys em. Fo he cen e ed allele coding sys em he base geno ype is he a e age geno ype in he da a. Thus, o his allele coding sys em he base popula ion is oughly he popula ion we wo k wi h [14], and i has he smalles a e age dis ance o obse ed geno ypes om he base geno ype. In p ac ice, his can be expec ed o lead o he highes eliabili ies. Di e en allele coding sys ems ha e di e en model design ma ices Z, and, hence, imply di e en models. Thus, eliabili ies om di e en allele coding sys ems a e in ac om di e en s a is ical models. Compa ison o eliabili ies om di e en models is meaningless. How- e e , he di e en allele coding sys ems lead o he same pa ame e es ima es. I he co ec allele coding me hod, i.e., s a is ical model, is known, i should be used. Because he ue model is unknown and compa ison o eliabili ies by allele coding me hod is meaningless, some p inciples mus be used o decide on which allele coding me hod should be used. These p inciples will no gua an ee he use o a co ec model o co ec eliabili ies. One such p inciple should be consis ency o eliabili ies be ween e alua ions. The cen e ed allele coding me hod changes model om one e alua ion o he nex because mo e ma - ke da a accumula e. Hence, acco ding o he consis ency p inciple, i canno be ecommended o compu a e Table 3 P edic ed absolu e and ela i e con e gence a es in McMC compu a ions by allele coding Allele coding 012 210 101 cen e ed single si e 2 0.9974795 0.9995523 0.9947515 0.9647670 ela i e 1.00 5.64 0.48 0.070 block 1 0.9995429 0.9998860 0.9989434 0.00 ela i e 1.00 4.01 0.43 - Table 4 Numbe o i e a ions in REML by allele coding and model Model Allele coding AI-REML EM-REML Ma ke 012 9 45 Ma ke 210 9 45 Ma ke 101 9 45 Ma ke cen e ed 9 45 G 012 7 34 G 210 9 44 G 101 7 31 G cen e ed 7 28 S andén and Ch is ensen Gene ics Selec ion E olu ion 2011, 43:25 h p://www.gsejou nal.o g/con en /43/1/25 Page 7 o 11 eliabili ies. Likewise, he base geno ype in he 012 and 210 allele coding me hods depend also on he obse ed allele equencies, i.e., ma ke da a. The cen e ed allele coding me hod is simila o ha in oduced in [4] whe e he allele equencies we e om an unselec ed base popula ion. I was used in o de o “gi e mo e c edi o a e alleles han o common alleles when calcula ing genomic ela ionships”.Asshown, in e ence is he same i espec i e o he allele coding me hod when a ixed gene al mean is in he model. Howe e , eliabili ies a e a ec ed as shown. The use o base popula ion allele equencies in he cen e ed allele coding me hod would emo e he abo e men ioned p o- blem o inconsis ency be ween e alua ions, bu es ima - ing hese allele equencies is elusi e. Recen ly, [15] p esen ed a me hod o adjus ing he G c ela ionship ma ix o become a ela ionship ma ix ela i e o he base popula ion, he eby a oiding he es ima ion o base popula ion allele equencies. The esul s in his pape a e based on he assump ion ha pheno ypes and geno ypes a e a ailable o all animals in he analysis. This assump ion may o en no be sa is ied. Models based on an ex ension o he genomic ela ionship ma ix o include also non-geno yped animals ha e been p esen ed by [16-18]. The esul s in he p esen pape abou pa ame e es ima es and es ima ed b eeding alues no depending on allele coding do no ca y o e o he models wi h an ex ended genomic ela ionship ma ix. Conclusions We showed ha , in heo y, di e en allele coding me h- ods led o he same in e ence in ma ke -based models when he model has a ixed gene al mean e ec . P ac i- cal analyses led o he same conclusions. Also in heo y, he cen e ed allele coding me hod was expec ed o gi e be e mixing p ope ies when Ma ko chain Mon e Ca lo me hods we e used. This was also obse ed in p ac ice. When an equi alen b eeding alue model was used, di e en allele coding me hods p o ed o lead o he same in e ence as in he ma ke -based model. How- e e , eliabili ies o b eeding alues depend on he cho- sen allele coding sys em because di e en allele coding me hods change he amoun o unce ain y in he es i- ma ed b eeding alues. Appendix A In he ollowing, we conside he e ec o allele coding me hod on he densi ies p(y|μ,g)andp(μ,g,|y). Fo simplici y o p esen a ion, he pa ame e s in he dis i- bu ion o gand ea e omi ed. Le p 0 deno e densi y o 012 allele coding. Because he loca ion pa ame e s μand ma ke e ec s g ela e o he obse a ions yonly h ough 1 n μ+Zg, we i s s udy his e m. By subs i u - ing Z=Z0−1n  m in o he e m, we ha e 1nμ+Zg =1nμ+Z0g−1n  m g =1n(μ−  m g)+Z0g. So, when di e en allele coding sys ems a e used, he densi ies ha e equali y by p(y|μ,g)=p0(y|μ−  m g,g) . (7) By changing he in eg a ion a iable μ0=μ−  mg ,we ob ain ∫p(y|μ,g)dμ=∫p 0 (y|μ 0 ,g)dμ 0 and, hence, p(y)=∫∫p(y|μ,g)p(g)dμdg=p 0 (y). F om hese esul s, we see ha p(μ,g|y)=p(y|μ,g)p(g)/p(y) =p0(y|μ−  mg,g)p0(g)/p0(y ) =p0(μ−  m g,g|y). (8) Table 6 Summa y s a is ics o genomic b eeding alue eliabili ies by allele coding Allele coding min mean max s d 012 0.41 0.49 0.59 0.022 210 0.30 0.37 0.42 0.017 101 0.55 0.62 0.73 0.024 cen e ed 0.72 0.80 0.95 0.026 Table 7 Co ela ions be ween genomic b eeding alue eliabili ies by allele coding Allele coding 210 101 cen e ed 012 -0.22 0.32 0.43 210 0.82 0.59 101 0.91 Table 5 REML es ima es by allele coding Allele coding Model Pa ame e 012 210 101 cen e ed Ma ke σ2 g 6.623 × 10 -4 6.623 × 10 -4 6.623 × 10 -4 6.623 × 10 -4 Ma ke σ2 e 2.993 2.993 2.993 2.993 Gσ2 a 1.540 1.540 1.540 1.540 Gσ 2 e 2.993 2.993 2.993 2.992 S andén and Ch is ensen Gene ics Selec ion E olu ion 2011, 43:25 h p://www.gsejou nal.o g/con en /43/1/25 Page 8 o 11 The esul s (7) and (8) a e ai ly gene al in e ms o dis ibu ional assump ions. The only equi emen s a e ha p(y|μ,g)dependsonμand gonly h ough 1 n μ+ Zg, ha an imp ope uni o m p io is used o μ,and ha p(y) is ini e. The la e equi emen is o assu e he pos e io dis ibu ion becomes a p ope dis ibu ion, and his has o be p o en o a model o be alid when an imp ope p io is used. When e~N(0,R), i is no di icul o show ha ∫p(y|μ,g)dμ∝exp(−1 2 (y−Zg)M(y−Zg) ) wi h M=R−1−R−11n1 n R−1/(1 n R−11n). The e o e, ∫ p(y|μ,g)dμ<c 0 whe e he cons an c 0 is independen o g. Thus, p(y)=∫∫p(y|μ,g)dμp(g)dg<c 0 ∫p(g)dg= c 0 is ini e, i espec i e o dis ibu ion o he ma ke e ec s g. Appendix B We conside a Gaussian dis ibu ion model whe e g∼N(0,Imσ 2 g) and e~N(0,R). Consequen ly, he dis- ibu ion [μ,g|y] is a mul i a ia e Gaussian dis ibu- ion. In he ollowing, we de i e he a iance-co a iance ma ix o his dis ibu ion. The condi ional densi y is p(μ,g|y)∝p(y|g,μ)p(g) ∝exp(−1 2(y−1nμ−Zg)R−1(y−1nμ−Zg) −1 2g’g/σ2 g ) ∝exp  −1 2[μ−ˆμg’ −ˆ g]Q  μ−ˆμ g−ˆ g  , whe e Q=[Va (μ,g|y)]−1=  −1¯z ¯z Cg  Wi h =1/(1 n R− 1 1n ) ,¯ z =Z’R−11 n and C g=Z’R−1Z+Imσ− 2 g . The ma ix Qis he coe icien ma ix in he mixed model equa ions and he in e se o his ma ix is Va (μ,g|y)=⎡ ⎣ + 2z Cgz − z Cg − Cgz Cg⎤ ⎦, whe e C g=(C g − ¯z ¯z )− 1 . No e ha he subma ix C g =Va (g|y) is independen o allele coding, because, asshownin hemain ex ,p(g|y) does no depend on allele coding. Va iance-co a iance ma ix o he comple e b eeding alue is Va (ad|y) =1nZ + 2¯z Cg¯z − ¯z Cg − Cg¯z Cg1 n Z = 1n1 n+ 21n¯z Cg¯z 1 n − ZCg¯z 1 n− 1n¯zCgZ+ZCgZ = 1n1 n +(In− 1n1 n R−1)ZCgZ’(In− R−11n1 n ) . Appendix C The esul s in [8] s a e ha he con e gence a e o a Gibbs sample is equal o he la ges modulus eigen a- lue o a ce ain ma ix Bwhe e he eigen alues can be complex numbe s. As men ioned in [9] his con e gence a e is also a a measu e o co ela ion be ween succes- si e McMC samples, i.e., mixing o he algo i hm. The close he con e gence a e is o ze o he less co ela ed a e he successi e samples. The Bma ix is cons uc ed as ollows. Le Qbe he in e se o he a iance-co a iance ma ix o he a ge mul i a ia e no mal dis ibu ion, in ou case he coe i- cien ma ix in he mixed model equa ions. Assume ha a Gibbs sampling scheme is used whe e he a iables a e g ouped in o sblocks and Qis spli acco dingly in o sblocks. Fi s , de ine A=I−diag(Q−1 11 ,...,Q−1 ss )Q . Le Lbe he block lowe iangula ma ix wi h blocks in he lowe diagonal being hose o A, and le U=A- L.Thus,Uis a s ic ly uppe iangle ma ix wi h ze os in he diagonal. Then he ma ix o in e es is B= ( I−L ) − 1 U . We conside he genomic ma ke model (1) whe e gis Gaussian g∼N(0,Imσ2 g) and eis Gaussian e∼N(0,Inσ2 e) . The condi ional dis ibu ion [μ,g|y]is a mul i a ia e no mal dis ibu ion wi h some mean ec- o and a a iance-co a iance ma ix wi h in e se Q=[Va (μ,g|y)]−1=  n/σ2 e1 nZ/σ2 e Z’1n/σ2 eCg , whe e Cg=Z’Z/σ2 e+Im/σ 2 g ; see Appendix B. We con- side wo McMC upda ing schemes. The i s scheme i e a es wo blocks: μand g. The second scheme upda es successi ely all pa ame e s: μ,g 1 , ..., g m . Fo he i s McMC upda ing scheme we ha e S andén and Ch is ensen Gene ics Selec ion E olu ion 2011, 43:25 h p://www.gsejou nal.o g/con en /43/1/25 Page 9 o 11