scieee Open visual document viewer

Genome-wide association analysis of milk yield traits in Nordic Red Cattle using imputed whole genome sequence variants

Iso-Touru, T.,Sahana, G.,Guldbrandtsen, B.,Lund, M. S.,Vilkki, J.

Full text

RESEARCH ARTICLE Open Access Genome-wide associa ion analysis o milk yield ai s in No dic Red Ca le using impu ed whole genome sequence a ian s T. Iso-Tou u 1* , G. Sahana 2 , B. Guldb and sen 2 , M. S. Lund 2 and J. Vilkki 1 Abs ac Backg ound: The No dic Red Ca le consis ing o h ee di e en popula ions om Finland, Sweden and Denma k a e unde a join b eeding alue es ima ion sys em. The long his o y o eco ding o p oduc ion and heal h ai s o e s a g ea oppo uni y o s udy p oduc ion ai s and iden i y causal a ian s behind hem. In his s udy, we used whole genome sequence le el da a om 4280 p ogeny es ed No dic Red Ca le bulls o scan he genome o loci a ec ing milk, a and p o ein yields. Resul s: Using a genome-wise signi icance h eshold, egions on Bos au us ch omosomes 5, 14, 23, 25 and 26 we e associa ed wi h a yield. Regions on ch omosomes 5, 14, 16, 19, 20 and 25 we e associa ed wi h milk yield and ch omosomes 5, 14 and 25 had egions associa ed wi h p o ein yield. Signi ican ly associa ed a ia ions we e ound in 227 genes o a yield, 72 genes o milk yield and 30 genes o p o ein yield. Ingenui y Pa hway Analysis was used o iden i y ne wo ks connec ing hese genes displaying signi ican hi s. When compa ed o p e iously mapped genomic egions associa ed wi h e ili y, signi ican ly associa ed a ia ions we e ound in 5 genes common o a yield and e ili y, hus linking hese wo ai s ia biological ne wo ks. Conclusion: This is he i s ime when whole genome sequence da a is u ilized o s udy genomic egions a ec ing milk p oduc ion in he No dic Red Ca le popula ion. Sequence le el da a o e s he possibili y o s udy quan i a i e ai s in de ail bu s ill canno unambiguously e eal which o he associa ed a ia ions is causa i e. Linkage disequilib ium c ea es di icul ies o pinpoin he causa i e genes and a ia ions. One solu ion o o e come hese di icul ies is he iden i ica ion o he unc ional gene ne wo ks and pa hways o e eal impo an in e ac ing genes as candida es o he obse ed e ec s. This in o ma ion on a ge genomic egions may be exploi ed o imp o e genomic p edic ion. Keywo ds: Milk ai s, No dic Red Ca le, Whole genome sequence, Associa ion s udy Backg ound The numbe o dai y cows in he No dic coun ies has been dec easing du ing he 21 s cen u y [1]. Howe e , o al milk p oduc ion le els ha e emained s able, as milk yield pe cow has inc eased. Fo example in Finland (including all dai y b eeds) he a e age p oduc ion pe cow pe yea has inc eased om 6786 l (2000) o 8201 l (2014), while a and milk con en s ha e emained ai ly cons an [2]. Global yea ly milk consump ion pe capi a is inc easing, and global demand o animal based oods is expec ed o be doubled by 2050 [3], d i en by bo h popula ion g ow h and inc eased consume p e e ences o mea and milk p oduc s. Ruminan s a e unique in hei capaci y o diges ib e and con e non-edible esou ces in o high quali y human nu i ion, making hem highly ele an o mee ing he inc easing global demand o ood. While animal b eede s ha e achie ed conside able imp o emen s in p oduc ion ai s, cow e ili y has been declining [4–6]. Howe e , du ing he ecen yea s, he dec ease in cow e - ili y in No dic coun ies has been slowing down and e en is e ac ed, due o he weigh ing o e ili y ai s in he b eeding p og am [7]. Many emale e ili y ai s in dai y ca le show an agonis ic gene ic co ela ions wi h milk p oduc ion ai s [8] bu wi h low o mode a e co ela- ions [9]. This implies ha simul aneous gene ic selec ion o inc eased milk yield and ep oduc i e pe o mance is * Co espondence: [email p o ec ed] 1 Animal Genomics, G een Technology, Na u al Resou ces Ins i u e Finland (Luke), Jokioinen, Finland Full lis o au ho in o ma ion is a ailable a he end o he a icle © 2016 Iso-Tou u e al. Open Access This a icle is dis ibu ed unde he e ms o he C ea i e Commons A ibu ion 4.0 In e na ional License (h p://c ea i ecommons.o g/licenses/by/4.0/), which pe mi s un es ic ed use, dis ibu ion, and ep oduc ion in any medium, p o ided you gi e app op ia e c edi o he o iginal au ho (s) and he sou ce, p o ide a link o he C ea i e Commons license, and indica e i changes we e made. The C ea i e Commons Public Domain Dedica ion wai e (h p://c ea i ecommons.o g/publicdomain/ze o/1.0/) applies o he da a made a ailable in his a icle, unless o he wise s a ed. Iso-Tou u e al. BMC Gene ics (2016) 17:55 DOI 10.1186/s12863-016-0363-8 possible [9]. Simul aneous b eeding o bo h p oduc i e and e ile cows would bene i subs an ially om knowing he gene ic and physiological links be ween p oduc ion and heal h o disen angle he e ec s on hese ai s. Recen esul s in Hols ein and Je sey b eeds indica e li le o no o e lap be ween genomic egions associa ed wi h milk yield and e ili y [10, 11]. Genome wide associa ion s udies (GWAS) ha e bene- i ed om he apid de elopmen o single nucleo ide polymo phism (SNP) geno yping echnologies, bu des- pi e o he ela i ely high densi y o he a ailable SNP chips, inding he causa i e mu a ion is no s aigh o - wa d. The high le el o linkage disequilib ium in dai y ca le esul s in long quan i a i e ai loci (QTL) egions wi h se e al possible candida e genes. Using whole genome le el sequence a ian s o associa ion analyses would be an ul ima e choice, because hen he causa i e a ian is mos likely included among he s udied a ian s. Po en ially his helps o pinpoin he causa i e mu a ions hus leading o a be e unde s anding o bio- logical mechanisms behind he QTL [12] and imp o e he e iciency o genomic selec ion [13]. Using sequence le el SNPs will also enable iden i ica ion o SNPs ha explain a small ac ion o he ai a ia ion because ei he he causal SNP and/o SNP(s) wi h high linkage disequilib ium (LD) wi h he causal a ian a e included in he analysis [13]. His o ically sepa a ed h ee dai y b eeds Finnish Ay shi e om Finland, Danish Red om Denma k and Swedish Red om Sweden a e a p esen unde a join b eeding alue es ima ion sys em, known as he No dic Ca le Gene ic E alua ion [14]. P e ious QTL s udies o milk ai s in No dic Red Ca le (NRC) ha e been done wi hin he subpopula ions wi h mic osa elli e ma ke s and ai ly small sample sizes (e.g. [15, 16]). The objec i e o his s udy was o use a ia ions a he genome sequence le el o ca y ou associa ion s udy o milk, a and p o ein yields in NRC; o iden i y po en ial causal a ian s and unde s and he gene ic a chi ec u e o hese ai s. In addi ion, he da a p o ides he possibili y o compa e he esul s o simila s udies o e ili y ai s in he NRC [17], o e eal po en ial QTL wi h an agonis- ic e ec s o milk p oduc ion and e ili y ai s. Me hods No animal expe imen s we e pe o med in his s udy, and, he e o e, app o al om he e hics commi ee was no equi ed. Semen samples we e collec ed o b eeding pu poses by local o ganiza ions wi h app op ia e pe mi s. Milk, a and p o ein yields’ ai de ini ions a e s an- da dized ac oss he No dic coun ies. Pheno ypic eco ds o dai y ca le a e housed in a cen alized da abase [14]. B eeding alues o milk, a and p o ein yield (MY, FY and PY) a e based on p oduc ion igu es exp essed in kilog ams aken om ou ine milk eco ds and hen com- bined in o an index o each ai . Fo de ails on gene ic e alua ion o milk yield ai s in No dic coun ies see [18]. The b eeding alues used o associa ion analysis we e de- eg essed b eeding alues om he ou ine gen- e ic e alua ion by NAV (No dic ca le gene ic e alua ion) and we e a ailable o 4280 p ogeny es ed NRC bulls (2127 om Finland, 1217 om Sweden, 915 om Denma k and 21 om o he coun ies). The eliabili ies o he de eg essed b eeding alues we e in he ange o 0.67 o 0.99 wi h a mean o 0.95 and he i s qua ile a 0.94. SNP a ay geno yping All 4280 NRC bulls wi h de eg essed b eeding alues we e geno yped using Bo ineSNP50 BeadChip SNP a ay e sion 1 o 2 (Illumina Inc., San Diego, CA). DNA was ex ac ed using s anda d p ocedu es om semen samples. Chip ypings we e done by GenoSkan A/S, Tjele, Denma k o labs belonging o Aa hus Uni e si y. The quali y pa ame e s used o selec ion o SNPs we e minimum call a es o 85 % o indi iduals and 95 % o loci. Ma ke loci wi h mino allele equencies below 5 % and de ia ion om Ha dy-Weinbe g p opo ion (P< 0.00001) we e excluded. The minimal accep able GC sco e was 0.60 o indi idual ypings, and indi iduals wi h a e age GC sco es below 0.65 we e excluded. The numbe o SNP emaining a e quali y con ol was 43,415 in he geno ypes ob ained om Bo ineSNP50 BeadChip SNP a ay (50 K da a se ). The genome posi ions o he SNPs we e acco ding o he UMD3.1 Bo ine genome assembly [19]. Impu a ion o whole genome sequences The 50 K geno ypes o hese bulls we e impu ed o whole genome sequence da a using a wo-s ep app oach [20]. Geno ypes om 50 K chip o each bull we e i s impu ed o a high-densi y SNP a ay (HD) using a mul i-b eed e e ence o 3383 animals (1222 Hols ein, 1326 NRC and 835 Danish Je sey indi iduals) which had been geno yped wi h he Illumina Bo ineHD chip (Illumina Inc., San Diego, CA). The numbe o SNPs, a e impu a ion o he Bo ineHD chip, was 648,219. These impu ed HD geno ypes we e subsequen ly impu ed o he whole genome sequence le el using a mul i-b eed e e ence panel o 1228 animals om Run4 o he 1000 Bull Genomes P ojec [13, 21] and addi ional whole genome sequences om Aa hus Uni e si y [22] including 368 Hols ein, 86 RDC, 88 Je sey and es om numbe o ca le b eeds. Da ase s (SNP a ay ypes and whole sequence) we e p e-phased wi h BEAGLE4 1274 [23] and geno ype impu a ion we e done using Minimac2 [24]. The impu a ion accu acy o his da a is epo ed ea lie [25], bu wi h a smalle whole genome sequence e e - ence popula ion. Sequence a ian s ha ing impu a ion accu acy 2 ( a io o empi ically obse ed a iance o Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 2 o 12 he allele dosages o he expec ed binomial a iance a Ha dy-Weinbe g equilib ium and was ob ained om Minimac2 so wa e ou pu ) less han 0.5 we e il e ed away. The mean accu acy o he a ian s wi h 2 >0.5 was 0.94. Associa ion analysis The associa ion analysis o each o he impu ed sequence a ian s (mino allele equency, MAF > 0.005 and de i- a ion om Ha dy-Weinbe g p opo ion > 0.00001) was ca ied ou using a wo-s ep a iance componen s-based app oach o accoun o popula ion s a i ica ion imple- men ed in he EMMAX so wa e ool [26]. In a i s s ep, he polygenic and e o a iances a e es ima ed using ollowing a iance componen model: y¼1μþaþe whe e yis a ec o o de- eg essed b eeding alues, 1is a ec o o ones, μis he in e cep , Gis he kinship ma ix buil based on high-densi y SNP geno ypes using EMMAX so wa e, ais a ec o o b eeding alues assumed o ha e a mul i a ia e no mal dis ibu ion a~N(0,Gσ a 2 ), eis a ec o o andom esiduals assumed o ha e a mul i a ia e no mal dis ibu ion e~N(0,Iσ e 2 ), whe e Iis an iden i y ma ix, σ a 2 is he addi i e gene ic a iance and σ e 2 is he e o a iance. In a second s ep, he SNP e ec is ob ained using a linea eg ession model: y¼1μþxb þη; whe e xis a ec o o impu ed geno ype dosages ( anged be ween 0 and 2), bis he allele subs i u ion e ec and ηis a ec o o andom esidual de ia es wi h (co) a iance s uc u e Gσ a 2 +Iσ e 2 . Sea ch o mul iple QTL in a genomic egion To es i mul iple QTL a e seg ega ing in a genomic egion we included he mos signi ican o he known causal a ian as co ac o in he model and check o addi ional QTL in a genomic egion o a yield on ch omosomes 14, 25 and 26 and o milk yield on ch omosome 14. We i ed he SNPs (Addi ional ile 1) as ixed e ec o a linea mixed model. The s a is ical model is desc ibed by he o mula: y¼1μþqsnp op þxgþZu þe whe e y,1,μa e desc ibed as in he EMMAX model, snp op is he e ec o he SNP i ed as co- ac o in he model, and gis headdi i egene ice ec o he heSNP unde s udy, qand xa e ec o s o SNP geno ype dosages ( anging om 0 o 2), and uis a ec o o andom poly- genic e ec s, which a e no mally dis ibu ed u~ N(0, Aσ u 2 ), whe e Ais he pedig ee-based addi i e ela ionship ma ix, σ u 2 is he polygenic a iance, Zis an incidence ma ix ela ing pheno ypes o he co esponding andom poly- genic e ec s, and eis a ec o o esidual e ec s, which a e no mally dis ibu ed e~N(0,Dσ e 2 ), whe e Dis a diagonal ma ix wi h elemen s d ii =(1− DRP 2 )/ DRP 2 o accoun o he e ogeneous esidual a iances due o di e en eliabil- i ies o DRP ( DRP 2 ), and σ e 2 is he esidual a iance. Analyses we e pe o med using he DMU package [27]. Signi i- cance es ing o SNP e ec s was pe o med using a wo-sided - es . The null hypo hesis was g = 0. A e ha , he Bon e oni co ec ion was applied same as in he EMMAX analysis o con ol o alse posi i e associa ions. The genome-wise signi icance h eshold co esponding o an e o a e o 0.05 was se a 3.16 × 10 −9 a e co ec- ion o mul iple es ing using a Bon e oni co ec ion o 15,679,852, 15,679,853 and 15,679,844 independen es s o a , milk and p o ein yield espec i ely. Only SNPs wi h he p alue less han 3.16 × 10 −9 (−log 10 (p)≥8.50) we e anno a ed wi h he a ian e ec p edic o (VEP) ool using he Ensembl da abase, Release 82 [28]. The p edic ion whe he an amino acid subs i u ion caused by missense a ia ion a ec s p o ein unc ion was es ima ed by SIFT analysis [29] implemen ed in VEP ool [28]. The SIFT p edic ion is based on sequence homology and he physical p ope ies o amino acids. Manha an plo s we e c ea ed wi h he qqman .0.1.2 R package [30]. In addi ion, we compa ed ou indings o esul s ob ained om s udy by Höglund e al. [17] whe e a simila genome-wide associa ion s udy o emale e ili y in No dic Red ca le was conduc ed. Ingenui y pa hway analysis and en ichmen analysis Lis s o genes wi h signi ican hi s om he QTL peak egions associa ed wi h each milk ai we e uploaded in o he Qiagen’s Ingenui y® Pa hway Analysis IPA® [31]. Fo his pu pose, also SNPs signi ican ly associa ed wi h Fe ili y index (FI) in he No dic Red Ca le [17] we e anno a ed wi h he VEP ool [28] and genes ha ing one o mo e signi ican SNP we e analyzed wi h IPA® [31]. Bioma ool [32] embedded in Ensembl da abase [33], was used o sea ching human homologs o cow Ensembl IDs o he genes. In case he e was mo e han one, all epo ed human o hologs we e kep . The human homologue lis s o each ai included 214, 69, 66 and 263 genes o a , milk, p o ein yields and e ili y index, espec i ely. A e unning co e analysis o each ai , ne wo ks based on he in o ma ion o gene connec i i y in Ingenui y Knowledge Da abase wi h highes sco e- alues we e conside ed. Sco e- alue ep esen s he nega i e log o he p- alue o he likeli- hood ha he molecules would be ound oge he by chance. Gene on ology (GO) e m en ichmen analysis wi h genes ound wi hin he op SNPs was pe o med wi h a Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 3 o 12 Singula En ichmen Analysis (SEA, Fishe ’s exac es , FDR < 0.01) p o ided by Ag iGO webpage [34]. Resul s Depending on he ai we could iden i y se e al hou- sands (3594, FY), less han a housand (755, MY) o less han a hund ed (85, PY) signi ican ly associa ed SNPs (Bon e oni co ec ed h eshold o signi icance -log 10 (p)≥8.50; Addi ional iles 2, 3 and 4). No common signi ican ly associa ed SNPs we e ound be ween his s udy and wi h hose ound o emale e ili y ai s [17]. Howe e , signi ican ly associa ed SNPs we e ound in i e common genes be ween e ili y and a yield. These i e genes a e loca ed on ch omosomes 25 (ENSBTAG00000034643) and 26 (GBF1,TMEM180, ACTR1B, and b a-mi -146b). The summa y o he anno a ions o signi ican ly associ- a ed SNPs a e p esen ed in Table 1. Among anno a ed SNPs, in on a ian s a e he mos common ype o each o he ai s. SNPs changing amino acid (missense a ia- ions) a e a e, 21 o FY, 1 o MY, 3 o PY and 18 sha ed be ween FY and MY. Se en missense a ia ions we e p edic ed by SIFT analysis [29] o be dele e ious, i.e. po en ially leading o changes in he unc ion o he p o ein (Addi ional ile 5). Splice egion a ian s a e a e o ming only one pe cen o less om he o al amoun o signi ican ly associa ed SNPs (Addi ional ile 5). Resul s o a yield wi hin he la ge associa ed a eas on ch omosomes 14, 25 and 26 we e u he examined by ixing he e ec o op SNP(s). Only peaks ha emained a e ixing he o he op SNPs we e consid- e ed as po en ial QTL in hese ch omosomes. Se en, eigh and ou sepa a e QTL egions (“peaks”) we e de ined o a , milk and p o ein yield, espec i ely (Fig. 1, Addi ional iles 6, 7, 8, 9, 10, 11, 12 and 13). The peaks we e de ined as con inuous egions con aining SNPs ha ing –log 10 (p)≥8.50. Top SNPs wi h he conse- quences o each de ined QTL egion pe ai a e lis ed in Table 2. The highes peak was obse ed on Bos au us ch omosome 14 (BTA14) (Addi ional ile 7) spanning he egion om 1,448,510 bp o 2,271,832 bp o a yield (509 SNPs ha ing –log 10 (p)≥8.50) and om 1,448,510 bp o 2,271,832 bp o milk yield (455 SNPs ha ing –log 10 (p)≥8.50). The highes –log 10 (p) alues wi hin hese egions we e ob ained o SNPs s136783505 (bp 1,807,140) o a yield and s133033480 (bp 1,743,939) o milk yield. The highes peak o p o ein yield was loca ed on BTA25 (Addi ional ile 11) wi hin he egion 3,306,363–3,516,671 which con ained 40 SNPs, he highes p alue being o SNP s110749311 (bp 3,498,960). Fa yield We iden i ied se en di e en QTL egions on i e di e en ch omosomes a ec ing a yield (Table 2, Fig. 1, Addi ional ile 2). The s onges associa ion ound o a yield was lo- ca ed on BTA14 (Addi ional ile 7) in he DGAT1 (Diacyl- glyce ol O-acyl ans e ase 1) gene egion. In ou da a he s onges associa ion is o he a ia ion s136783505 (bp 1,807,140) wi h no unc ional anno a ion, loca ed 2578 bp downs eam o he DGAT1 gene (Table 2). Howe e , se - e al o he a ian s loca ed nea by, including he p e iously iden i ied causa i e a ian K232A [35] a bp 1,802,266, show simila ly high -log 10 (p) (Addi ional ile 2). To in es i- ga e he signi icance o o he SNPs in he egion, we i ed he a ia ion K232A as ixed e ec . None o he o he SNPs emained signi ican (Addi ional ile 14) a e he ixa ion o he K232A a ia ion. O he SNPs wi h s ong associa ions wi h a yield we e ound on BTA5 (Addi ional ile 6), BTA23 (Addi ional ile 11), BTA25 (Addi ional ile 12) and BTA26 (Addi ional ile 13). The QTL egion on BTA5 is loca ed be ween 92,372,732 bp and 94,425,668 bp and he a ia ion wi h he s onges associa ion ( s209818856, pos. 93,945,694) is loca ed in an in on o he gene MGST1 (Table 2). On BTA23 he associa ion signal o a yield comes om he egion 28,567,796–28,591,530, he op a ia ion being lo- ca ed a bp 28,567,796 in he in on o gene TRIM26 (Table 2). On BTA25 and BTA26, complex pa e ns o associ- a ion we e seen (Addi ional iles 12 and 13). To cla i y he numbe o independen QTL wi hin hese egions we in es iga ed he signi icance o he SNPs by i ing Table 1 The numbe o signi ican SNPs (- log 10 (p)≥8.50) o each ai and how SNPs a e di ided in o di e en consequences. SNPs we e anno a ed wi h he a ian e ec p edic o – ool [28]. One SNP can ha e mo e han one anno a ion T ai Numbe o signi ican SNPs In on a ian In e genic a ian Downs eam gene a ian Ups eam gene a ian Synonymous a ian Missense a ian 3′UTR a ian Splice egion a ian , in on a ian 5′UTR a ian Splice egion a ian , synonymous a ian Fa yield 3594 1641 1307 600 551 94 40 39 9 6 2 Milk yield 755 322 195 301 240 42 20 16 5 1 2 P o . yield 85 50 14 32 12 3 3 1 1 Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 4 o 12 he op SNPs ( i e o BTA25 and eigh o BTA26, Addi ional ile 1) as ixed e ec s alone and in di e en combina ions. On BTA 25, signi ican associa ions emained a bp 9,870,005 (in onic egion o he CLEC16A gene) and a bp 36,226,978 (in e genic egion) (Table 2). Two QTL emained also on BTA26, one in he egion o he NEURL1 gene (posi ion 24,379,571) and he o he in an in e genic egion ( op SNP a bp 44,802,991) (Table 2). Milk yield In all, eigh QTL egions we e ound o milk yield (Table 2, Fig. 1, Addi ional ile 3). They we e loca ed on six di e en ch omosomes (BTA5, BTA14, BTA16, BTA19, BTA20 and BTA25). The s onges associa ion wi h milk yield was ound on BTA14 (1,448,510– 2,271,832) ha ing op a ia ion loca ed a bp 1,743,939. This a ia ion is wi hin wo o e lapping genes, CPSF1 and ADCK5.TheQTL egionis hesameaswas ound associa ed wi h a yield bu wi h di e en op a i- a ion. As o a yield, he signi icance o o he han he p e iously known causa i e DGAT1 a ia ion [35] was es ed by i ing he a ia ion K232A as a ixed e ec . None o he o he SNPs emained signi ican a e ixing he K232A e ec (Addi ional ile 15). The weakes signi ican associa ion was loca ed on BTA5 wi h op a ia ion s383553819 (posi ion 112,343,204) loca edinanin ono hegeneMKL1 (Addi ional ile 6). On BTA19 he associa ed egion con ained no anno a ed genes (61,447,138–61,449,096), and he op a ia ion s210324693 (bp 61,449,096 bp) was loca ed in an in e - genic egion (Table 2, Addi ional ile 9). Two QTL egions we e de ec ed on BTA20 (Addi ional ile 10). The QTL we e loca ed in he egions 30,531,217–32,952,019 and 37,766,226–39,183,141. The known causa i e a ia ion F279Y (bp. 31,909,478) o milk ai s in he gene GHR [36]was he opSNPinou analysis o heQTLlo- ca ed in he i s egion. The op a ia ion wi hin he second QTL was loca ed a bp 38,828,254 in he in e genic egion, bu his QTL egion also includes he PRLR gene p e iously indica ed o be linked o milk p oduc ion (e.g. [37]). BTA25 ha bo s wo QTL, he i s a bp 2,669,704 and he second in he egion om 3,494,706 bp o 3,516,671 bp, he op SNP lo- ca ed a bp 3,498,960 downs eam om he gene PAM16 (Addi ional ile 12). Fig. 1 Genome-wide Manha an plo s o he a yield (FY), milk yield (MY) and p o ein yield (PY). Red line indica es he genome-wide signi icance le el (−log 10 (p) = 8.50) Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 5 o 12 Table 2 QTL egions o each o he ai . Top SNP o each QTL a e shown including posi ion, -log 10 (p)- alues, mino allele equency (MAF), gene in o ma ion, anno a ion o he op SNP, allele subs i u ion e ec (b. alue) and s anda d e o o b. alue (SE) CHR S a (bp) End (bp) Leng h o he QTL egion (bp) Signi ican SNPs in egion No o genes wi h signi ican SNPs wi hin he QTL egion Top SNP Posi ion o he op SNP (bp) -log 10 (p) MAF Gene Anno a ion o he op SNP b. alue SE Fa yield 5 92,372,732 94,425,668 2,052,936 330 3 s209818856 93,945,694 27.49 0.38 MGST1 in on a ian −2.807 0.253 14 1,448,510 2271832 823,322 509 49 s136783505 1,807,140 42.01 0.07 DGAT1/HSF1 downs eam a ian / in on a ian −6.709 0.484 23 28,567,796 28,591,530 23,734 5 1 s381390819 28,567,796 9.36 0.45 TRIM26 in on a ian 1.148 0.184 25 8,222,347 11,507,986 3,285,639 883 16 s379546164 9,870,005 14.40 0.24 CLEC16A in on a ian −1.769 0.224 25 36,226,978 36,227,132 154 2 s109480808 36,226,978 8.72 0.16 in e genic a ian 1.635 0.272 26 22,144,777 24,793,744 2,648,967 595 32 s438420348 24,379,571 14.28 0.16 NEURL1 in on a ian −2.273 0.290 26 44,802,991 44,802,991 0 1 s135624939 44,802,991 9.72 0.28 in e genic a ian −1.474 0.231 Milk yield 5 112,343,204 11,2450,860 107,656 2 1 s383553819 112,343,204 8.70 0.36 MKL1 in on a ian 1.437 0.239 14 1,448,510 2,271,832 823,322 455 48 s133033480 1,743,939 33.01 0.09 CPSF1/ADCK5 downs eam a ian / splice egion a ian , in on a ian 6.266 0.513 16 1,322,611 1,322,611 0 1 1 s108979795 1,322,611 8.58 0.28 LAX1 ups eam a ian −1.451 0.243 19 61,447,138 61,449,096 1958 5 s210324693 61,449,096 8.93 0.32 in e genic a ian 1.490 0.244 20 30,531,217 32,952,019 2,420,802 74 5 s385640152 31,909,478 15.56 0.11 GHR missense a ian −3.877 0.472 20 37,766,226 39,183,141 1,416,915 34 3 NA 38,828,254 9.54 0.16 in e genic a ian −2.250 0.356 25 2,669,704 2,669,704 0 1 s209691835 2,669,704 9.21 0.13 in e genic a ian −2.855 0.460 25 3,494,706 3,516,671 21,965 13 4 s110749311 3,498,960 9.32 0.41 PAM16/GLIS2 downs eam a ian 1.225 0.196 P o ein yield 5 112,450,860 112,450,860 0 1 s109041054 112,450,860 8.88 0.48 in e genic a ian −1.473 0.242 14 1,802,667 1,802,667 0 1 2 NA 1,802,667 8.52 0.06 DGAT1/HSF1 in on a ian / downs eam a ian 3.354 0.564 25 1,094,996 1,257,612 162,616 12 3 s136085792 1,103,856 10.83 0.22 UNKL in on a ian 1.694 0.250 25 3,306,363 3,516,671 210,308 40 8 s110749311 3,498,960 11.70 0.41 PAM16/GLIS2 downs eam a ian 1.427 0.202 Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 6 o 12 P o ein yield In all, p o ein yield did no show as many signi ican ly associa ed SNPs as obse ed o a and milk yield (Table 2, Fig. 1, Addi ional ile 4). Ch omosomes ha ing QTL o p o ein yield we e BTA5, BTA14, and BTA25. Fo bo h BTA5 and BTA14, only one a ian o each eached he signi icance cu -o le el; on BTA5 an in e - genic a ia ion a bp 112,450,860 (Addi ional ile 6) close o he a ian ound o milk yield and on BTA14, he a ia ion (bp 1,802,667) loca ed in an in on o he DGAT1 gene (Addi ional ile 7). The wo o he QTL o p o ein yield we e bo h lo- ca ed on BTA25. The QTL on BTA25 a 3,306,363– 3,516,671 o e lapped wi h he QTL ound o milk yieldanddisplayed hesame opSNP(bp3,498,960). The o he QTL on BTA25 was unique o p o ein yield, op SNP (bp 1,103,856) loca ed in he gene UNKL. Ne wo ks o associa ed genes and en ichmen analysis The ne wo ks wi h highes sco es o each ai a e p esen ed in Addi ional iles 16, 17 and 18. The wo op ne wo ks om genes associa ed wi h a yield (Addi ional ile 16) had sco es o 49, he ne wo k associ- a ed wi h ca bohyd a e me abolism, gene exp ession and lipid me abolism is p esen ed in Addi ional ile 16a. Addi ional ile 16b shows he ne wo k gene a ed om he genes associa ed o e ili y index. The sco e alue o his ne wo k is 41 and i consis s o genes associa ed wi h in lamma o y esponse, cell- o-cell signaling and lymphoid issue s uc u e and de elopmen . E en hough common signi ican ly associa ed SNPs we e no ound be ween his s udy and ha o [17] on emale e ili y ai s, signi ican ly associa ed SNPs we e ound in i e common genes. Two o hem (GBF1 and b a-mi -146b) a e p esen in bo h a and e ili y ne wo ks (Addi ional ile 16a and b). The op ne wo k o milk yield (sco e 43, Addi ional ile 17) was associa ed wi h unc ions molecula anspo , o gan mo phology and o ganismal de elopmen . I includes he known milk candida e genes DGAT1,GHR and PRLR,as well as he op hi s om BTA5 (MLK1) and BTA16 (LAX1). The op ne wo k o p o ein yield (sco e 35, Addi ional ile 18) is connec ed o unc ions cell dea h and su i al, cance and o ganismal inju y. Al oge he 18, 18, 11 and 16 GO e ms we e signi i- can ly ( alse disco e y a e, FDR < 0.01) en iched o FY, MY, PY and FI, espec i ely. A b oad GO e m, mul i- cellula o ganismal p ocess, was he mos signi ican o all ou ai s, o ally eigh e ms we e sha ed be ween hem. Fo example all ai s we e ha ing QTL in he egions con aining signi ican en ichmen o genes ela ed o ep oduc ion and ep oduc i e p o- cesses (Table 3). Discussion A la ge numbe o a ian s we e ound signi ican ly associa ed wi h milk, a and p o ein yield. Ou indings suppo p e ious QTL indings om he No dic Red b eeds, e.g. on BTA 5, 14, 20 [37, 38] and loca e new a ia ions ha a e good candida es o be causa i e a ia ions. This is he i s ime when NRC popula ion is s udied wi h impu ed whole genome sequence a ian s in o de o e ine QTL associa ed wi h milk p oduc ion. Howe e , i is s ill di icul o pinpoin he causa i e a ian among se e al closely linked, almos equally signi ican ly associa ed a ia ions. One way o classi y he a ia ions is o look a he p edic ed unc ional con- sequence o he SNP [29]. The possibili y ha a a ia ion has an impac on he pheno ype is highe i he a ia ion causes an amino acid change (missense a ia ion) which is p edic ed (e.g. wi h SIFT analysis, [29]) o ha e an e ec o p o ein unc ion, is loca ed on splicing si e, o is loca ed downs eam o ups eam o he known gene (possible egula o y egions o he ansc ip ion). On he o he hand, genome anno a ion o ca le is s ill incom- ple e and mos egula o y elemen s emain unknown. In he sea ch o biologically ele an ma ke s he in o ma- ion o in e ac ions be ween genes in known pa hways o ne wo ks can be use ul. In his s udy, we iden i ied some in e es ing gene in e ac ion ne wo ks based on he signi i- can ly associa ed a ian s wi hin genes (e en hough he unc ional e ec s o he a ian s could no be p edic ed). The esul smaybeused oha eaclose looka alsoo he genes in he indica ed ne wo ks o unc ional a ian s. Al hough no common SNPs we e ound associa ed wi h milk p oduc ion ai s and e ili y, he i e common genes be ween a yield and e ili y gi e some indica ion o he ela ionship be ween hose ai s. The genes b a-mi -146b and GFB1 a e associa ed wi h a e ili y gene ne wo k linked wi h in lamma o y esponse and cell- o-cell signal- ing and he a yield ne wo k connec ed wi h lipid and ca bohyd a e me abolism. Fu he suppo was gained om he gene en ichmen analysis, bo h he ai s show signi ican en ichmen o he genes ela ed o example o ep oduc ion and ep oduc ion and ep oduc i e p o- cesses, al oge he ha ing 16 common GO e ms. F om he ou ch omosomes epo ed o ha bo highes numbe o QTL o milk p oduc ion [39], wo we e indi- ca ed by ou da a (BTA14 and BTA20). S ucken e al. [40] summa ized 14 genes om en di e en ch omosomes o be he majo genes in ol ed in milk p oduc ion. Among hose genes a e DGAT1 and GHR. Some commonly ound QTL (e.g. BTA6, [41]) we e no seen in ou da a; ha could be due o ixa ion o he QTL o e y low MAF in he NRC popula ion. One explana ion could be ha he EMMAX me hod chosen o associa ion analysis migh be oo conse a i e. EMMAX uses app oxima ions o con- s uc ing es o he ixed SNP e ec s o in e es in he Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 7 o 12 linea mixed model because i ing a ull linea mixed model o each SNP in u n ac oss he genome is compu a- ionally challenging [42]. This leads o sys ema ic unde - es ima ion o he mos signi ican p alues [43], bu makes EMMAX one o he as es LMM based p og ams [42]; an a gumen ha has o be conside ed when ha ing whole genome sequence le el da a om subs an ially la ge amoun o indi iduals. DGAT1 (BTA14) The s onges signal o associa ion was ound om BTA14 o a and milk yield. Also p o ein yield is signi ican ly associa ed wi h he same egion on BTA14. The op SNP a ied depending on he ai (see Table 2). Dae wyle e al. [13] analyzed associ- a ion o ea ly lac a ion milk a pe cen age wi h whole genome sequence a ia ion da a in Fleck ieh and Hols ein bulls. As in ou s udy, he p e iously epo ed causa i e a ia ion K232A (bp 1,802,266, [35]) o DGAT1 gene was no he a ian wi h he lowes p alue in Hol- s ein and Fleck ieh o milk a p oduc ion al hough K232A was among he op SNPs. Pausch e al. [44] used whole genome sequence da a o impu e Ge man Fleck ieh and Hols ein-F iesian ca le geno ypes om a la ge se o animals candida e egions and we e able o con i m he associa ion o K232A mu a ion, howe e , hei app oach was biased as only known candida e SNPs we e es ed o associa ion. In ou da a, close inspec ion o he associa ed a ia ions (−log 10 (p)≥8.50, a yield) in he DGAT1 egion e ealed ha when he e ec o he K232A mu a ion was ixed, no addi ional s a is ically signi ican SNP e ec s we e le . The causa i e mu a ion (K232A/ s109326954) o DGAT1 was epo ed al eady o e a decade ago [45] and has been unc ionally con i med [35]. The K allele inc eases milk a pe cen age [35], whe eas allele A inc eases milk p oduc ion [46]. The e a e di e en possible explana ions why K232A did no u n ou o be he mos signi ican ly associa ed a ia ion in his s udy. Impe ec impu a ion may a ec associa ion esul s. Accu acy o he impu a ion is conside - ably imp o ed by inc easing he size o e e ence panel, i.e. sequenced animals [47] and impu a ion accu acy seems o be highe when popula ions unde s udy a e combined o he impu a ion p ocesses [13]. Ou e e ence panel consis ed o a mul i-b eed popula ion wi h 1228 indi- iduals om se e al b eeds including bo h dai y and bee ca le. The DGAT1 egion would be an in e es - ing candida e o s udy wi h he in o ma ion om 1000 Bulls Genomes P ojec [13, 21]. I would gi e a chance o s udy he haplo ype s uc u e o he egion in he ca le popula ion wo ldwide and possibly ace back he e olu ion o he QTL e ec . Table 3 GO en ichmen e ms ha ing FDR < 0.01 om he genes ha ing signi ican a ia ions (−log 10 (p)≥8.50) o a yield (FY), milk yield (MY), p o ein yield (PY) and e ili y index (FI) FY MY PY FI GO e m Desc ip ion FDR FDR FDR FDR GO:0022610 Biological adhesion 6.20E-07 0.0013 2.00E-06 GO:0065007 Biological egula ion 9.40E-12 6.80E-05 0.072 2.40E-28 GO:0044085 Cellula componen biogenesis 3.60E-09 0.048 1.00E-22 GO:0016043 Cellula componen o ganiza ion 1.00E-37 2.90E-12 0.00027 1.80E-76 GO:0009987 Cellula p ocess 4.90E-14 4.60E-05 0.0074 3.80E-20 GO:0016265 Dea h 7.80E-21 2.00E-07 4.00E-25 GO:0032502 De elopmen al p ocess 8.80E-102 2.10E-41 1.40E-14 GO:0051234 Es ablishmen o localiza ion 1.80E-14 2.80E-05 0.0057 1.70E-27 GO:0040007 G ow h 5.60E-42 2.10E-17 9.40E-45 GO:0002376 Immune sys em p ocess 0.00022 0.0032 9.80E-10 GO:0051179 Localiza ion 5.90E-20 1.10E-07 0.0059 4.70E-44 GO:0008152 Me abolic p ocess 9.40E-12 0.00035 0.57 7.80E-18 GO:0032501 Mul icellula o ganismal p ocess 8.00E-126 1.10E-44 1.50E-20 1.20E-161 GO:0048519 Nega i e egula ion o biological p ocess 2.80E-50 1.30E-25 1.20E-07 GO:0048518 Posi i e egula ion o biological p ocess 1.70E-32 1.60E-08 GO:0050789 Regula ion o biological p ocess 7.10E-10 0.00011 0.32 6.90E-22 GO:0000003 Rep oduc ion 3.70E-57 3.70E-14 3.50E-09 1.20E-54 GO:0022414 Rep oduc i e p ocess 1.70E-50 7.50E-11 1.90E-07 6.30E-38 GO:0050896 Response o s imulus 6.00E-31 5.50E-06 8.80E-07 8.60E-44 Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 8 o 12 MGST1 and MKL1 (BTA5) Vii ala e al. [16] showed ha Finnish Ay shi e has a milk p oduc ion QTL a he p oximal end o BTA5. Wang e al. [48] epo ed o a QTL o milk a pe cen age in Ge man Hols ein-F iesian a he loca ion o 94,551,792 bp and sugges ed a candida e gene o be EPS8. One o he mos signi ican SNP in he s udy by Aliloo e al. [11] o milk yield (Je sey and Hols ein) was loca ed a bp 94,518,850 on BTA5. Fo a yield, we obse ed an associ- a ion peak a bp 93,945,694 (Table 2). This a ia ion is in in onic egion o he gene MGST1 ha ing a ole in oxida i e s ess eac ion [49] QTL peak o he milk yield on BTA5 is loca ed wi hin an in on o he MKL1gene (bp 112,343,204). MKL1 is ela ed o an- sc ip ion egula ion. P o ein yield associa ion peak is loca ed a bp 112,450,860 on BTA5 and is anno a ed o he non-coding egion. P e iously QTL ela ed o body weigh ha e been mapped nea by [50]. GHR and PRLR (BTA20) BTA20 is among he ch omosomes ha bo ing many QTL ela ed o milk p oduc ion [39]. GHR has been epo ed as one o he majo genes in ol ed in milk p oduc ion [40]. Fi s Blo e al. [36] ound ha a ia ion F279Y (bp 31,909,478) in he GHR gene is associa ed wi h a s ong e ec on milk yield and composi ion and o he s udies ha e con i med i (e.g. [44]). O he a ia ions han F279Y ha e also been ound o a ec milk p oduc ion nea by he GHR gene [38]. In ou s udy, F279Y has he s onges associa ion o milk yield among he NRC popula ion on BTA20. Fu he con i ma ion o he causali y o he F279Y comes om SIFT analysis p edic ing he mu a ion o be dele e ious o p o ein s uc u e, i.e. po en ially al e - ing he p o ein s uc u e hus possible leading o changes in unc ion o he p o ein. The op SNP (bp 31,909,478) in he GHR gene egion is clea ly he mos signi ican one (Addi ional ile 10) in con as o DGAT1 egion whe e se e al a ia ions a e s ongly associa ed o milk yield (Addi ional ile 7). The Y allele is p edic ed o be un a o - able o milk p oduc ion [36], bu i is s ill ai ly common in he NRC popula ion. GHR has been sugges ed o be unde balancing selec ion because o he obse ed high a ia ion in he cy oplasmic egion [51]. Ano he gene on BTA20 o special in e es is PRLR and he a ia ion S18N (posi ions 39,115,344-39,115,345) [37]. Howe e , i has been sugges ed ha S18N is a he linked o he causa i e mu a ion han being causa i e i sel [44]. We ound ha he a ia ion a bp 38,828,254 on BTA20 loca ed in he in e genic egion, was indica ed o be he mos likely candida e esponsible o he QTL e ec seen. This a ia ion is loca ed app oxima ely 245,000 base pai s downs eam o he PRLR gene and addi ional s udies a e equi ed o esol e he mechanisms how i may in luence milk p oduc ion. TRIM26 (BTA23) An in onic a ian in he gene TRIM26 a bp 28,567,796 on BTA23 has an associa ion wi h a yield. The unc ion o he TRIM26 gene, a membe o he ipa i e mo i (TRIM) gene amily, is unknown [52]. I is loca ed close o he majo his ocompa ibili y complex (MHC) class I egion. Feed in ake QTL ha e been mapped close o he associa ion peak obse ed in his s udy [53]. PAM16,UNKL and CLEC16A (BTA25) Al oge he six associa ion peaks (QTL) we e obse ed om BTA25 o di e en ai s. The same a ia ion (bp 3,498,960) in gene PAM16 is associa ed wi h bo h milk and p o ein yield. The gene has a c i ical ole in p o ein ansloca ion ac oss he inne mi ochond ial memb ane [54]. O he QTL on BTA25 we e ound o milk yield in he in e genic egion (bp 2,669,704) and o p o einyieldQTLa bp1,103,856in hegene UNKL ha has a ole on p o ein, zink ion and me al ion binding. Two dis inc QTL we e iden i ied o a yield, peak a ia ions loca ed a posi ions 9,870,005 and 36,226,978. Va ia ion a bp 9,870,005 is loca ed in he CLEC16A gene (Table 2). Va ia ions o CLEC16A in humans a e associa ed wi h inc eased ype I diabe es isk [55]. In addi ion, milk p o ein pe cen age QTL has p e i- ously been ound om he egion 9.3–10.6 Mb [56]. NEURL1 (BTA26) A e he signi icance es by ixing o he op SNPs, wo QTL we e le on BTA26 o a yield a bp 24,379,571 in NEURL1 gene and a bp 44,802,991 (Table 2). NEURL1 gene is associa ed wi h lac a ion (GO e m 0007595) hus making he a ia ion an in e es ing candi- da e o be a causa i e mu a ion. Conclusions Associa ion analyses among No dic Red Ca le using o e 15 million sequence a ia ions ac oss he whole genome impu ed o o e 4000 p ogeny es ed No dic Red Ca le bulls indica ed se e al a ia ions likely o ha e an impac o milk p oduc ion. We show ha impu a ion is obus and cos -e ec i e way o expand he in o ma ion a ailable and o inc ease knowledge o he causa i e mu a ions a ec ing ai s impo an o p o- duc ion animals. The a ailabili y o he whole genome le el sequence da a opens endless possibili ies o s udy quan i a i e ai a chi ec u e mo e closely. S ill inding he quan i a i e ai nucleo ides is challenging, wi h linkage disequilib ium and many small-e ec QTL c ea ing he puzzle ha is no easy o sol e. Fu he mo e, be e anno- a ion o he ca le genome is equi ed o be able o p edic he e ec s o a ia ions on he pheno ypes mo e accu - a ely. The knowledge om gene in e ac ions (al hough human/ oden based) may help o iden i y likely candida e Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 9 o 12