scieee Science in your language
[en] (orig)

Genome-wide association analysis of milk yield traits in Nordic Red Cattle using imputed whole genome sequence variants

Read accessible full text

Genome-wide association analysis of milk yield traits in Nordic Red Cattle using imputed whole genome sequence variants

Author: Iso-Touru, T.,Sahana, G.,Guldbrandtsen, B.,Lund, M. S.,Vilkki, J.
Publisher: BioMed Central,London,gb
Year: 2016
Source: https://jukuri.luke.fi/bitstream/10024/532243/1/IsoTouru.pdf
RESEARCH ARTICLE Open Access
Genome-wide associa ion analysis o milk
yield ai s in No dic Red Ca le using
impu ed whole genome sequence a ian s
T. Iso-Tou u
1*
, G. Sahana
2
, B. Guldb and sen
2
, M. S. Lund
2
and J. Vilkki
1
Abs ac
Backg ound: The No dic Red Ca le consis ing o h ee di e en popula ions om Finland, Sweden and Denma k
a e unde a join b eeding alue es ima ion sys em. The long his o y o eco ding o p oduc ion and heal h ai s
o e s a g ea oppo uni y o s udy p oduc ion ai s and iden i y causal a ian s behind hem. In his s udy, we
used whole genome sequence le el da a om 4280 p ogeny es ed No dic Red Ca le bulls o scan he genome
o loci a ec ing milk, a and p o ein yields.
Resul s: Using a genome-wise signi icance h eshold, egions on Bos au us ch omosomes 5, 14, 23, 25 and 26
we e associa ed wi h a yield. Regions on ch omosomes 5, 14, 16, 19, 20 and 25 we e associa ed wi h milk yield
and ch omosomes 5, 14 and 25 had egions associa ed wi h p o ein yield. Signi ican ly associa ed a ia ions we e
ound in 227 genes o a yield, 72 genes o milk yield and 30 genes o p o ein yield. Ingenui y Pa hway Analysis
was used o iden i y ne wo ks connec ing hese genes displaying signi ican hi s. When compa ed o p e iously
mapped genomic egions associa ed wi h e ili y, signi ican ly associa ed a ia ions we e ound in 5 genes
common o a yield and e ili y, hus linking hese wo ai s ia biological ne wo ks.
Conclusion: This is he i s ime when whole genome sequence da a is u ilized o s udy genomic egions a ec ing
milk p oduc ion in he No dic Red Ca le popula ion. Sequence le el da a o e s he possibili y o s udy quan i a i e
ai s in de ail bu s ill canno unambiguously e eal which o he associa ed a ia ions is causa i e. Linkage disequilib ium
c ea es di icul ies o pinpoin he causa i e genes and a ia ions. One solu ion o o e come hese di icul ies is he
iden i ica ion o he unc ional gene ne wo ks and pa hways o e eal impo an in e ac ing genes as candida es o he
obse ed e ec s. This in o ma ion on a ge genomic egions may be exploi ed o imp o e genomic p edic ion.
Keywo ds: Milk ai s, No dic Red Ca le, Whole genome sequence, Associa ion s udy
Backg ound
The numbe o dai y cows in he No dic coun ies has
been dec easing du ing he 21
s
cen u y [1]. Howe e , o al
milk p oduc ion le els ha e emained s able, as milk yield
pe cow has inc eased. Fo example in Finland (including
all dai y b eeds) he a e age p oduc ion pe cow pe yea
has inc eased om 6786 l (2000) o 8201 l (2014), while
a and milk con en s ha e emained ai ly cons an [2].
Global yea ly milk consump ion pe capi a is inc easing,
and global demand o animal based oods is expec ed o
be doubled by 2050 [3], d i en by bo h popula ion g ow h
and inc eased consume p e e ences o mea and milk
p oduc s. Ruminan s a e unique in hei capaci y o diges
ib e and con e non-edible esou ces in o high quali y
human nu i ion, making hem highly ele an o mee ing
he inc easing global demand o ood. While animal
b eede s ha e achie ed conside able imp o emen s in
p oduc ion ai s, cow e ili y has been declining [4–6].
Howe e , du ing he ecen yea s, he dec ease in cow e -
ili y in No dic coun ies has been slowing down and e en
is e ac ed, due o he weigh ing o e ili y ai s in he
b eeding p og am [7]. Many emale e ili y ai s in dai y
ca le show an agonis ic gene ic co ela ions wi h milk
p oduc ion ai s [8] bu wi h low o mode a e co ela-
ions [9]. This implies ha simul aneous gene ic selec ion
o inc eased milk yield and ep oduc i e pe o mance is
* Co espondence: [email p o ec ed]
1
Animal Genomics, G een Technology, Na u al Resou ces Ins i u e Finland
(Luke), Jokioinen, Finland
Full lis o au ho in o ma ion is a ailable a he end o he a icle
© 2016 Iso-Tou u e al. Open Access This a icle is dis ibu ed unde he e ms o he C ea i e Commons A ibu ion 4.0
In e na ional License (h p://c ea i ecommons.o g/licenses/by/4.0/), which pe mi s un es ic ed use, dis ibu ion, and
ep oduc ion in any medium, p o ided you gi e app op ia e c edi o he o iginal au ho (s) and he sou ce, p o ide a link o
he C ea i e Commons license, and indica e i changes we e made. The C ea i e Commons Public Domain Dedica ion wai e
(h p://c ea i ecommons.o g/publicdomain/ze o/1.0/) applies o he da a made a ailable in his a icle, unless o he wise s a ed.
Iso-Tou u e al. BMC Gene ics (2016) 17:55
DOI 10.1186/s12863-016-0363-8
possible [9]. Simul aneous b eeding o bo h p oduc i e
and e ile cows would bene i subs an ially om knowing
he gene ic and physiological links be ween p oduc ion
and heal h o disen angle he e ec s on hese ai s.
Recen esul s in Hols ein and Je sey b eeds indica e li le
o no o e lap be ween genomic egions associa ed wi h
milk yield and e ili y [10, 11].
Genome wide associa ion s udies (GWAS) ha e bene-
i ed om he apid de elopmen o single nucleo ide
polymo phism (SNP) geno yping echnologies, bu des-
pi e o he ela i ely high densi y o he a ailable SNP
chips, inding he causa i e mu a ion is no s aigh o -
wa d. The high le el o linkage disequilib ium in dai y
ca le esul s in long quan i a i e ai loci (QTL) egions
wi h se e al possible candida e genes. Using whole
genome le el sequence a ian s o associa ion analyses
would be an ul ima e choice, because hen he causa i e
a ian is mos likely included among he s udied
a ian s. Po en ially his helps o pinpoin he causa i e
mu a ions hus leading o a be e unde s anding o bio-
logical mechanisms behind he QTL [12] and imp o e
he e iciency o genomic selec ion [13]. Using sequence
le el SNPs will also enable iden i ica ion o SNPs ha
explain a small ac ion o he ai a ia ion because
ei he he causal SNP and/o SNP(s) wi h high linkage
disequilib ium (LD) wi h he causal a ian a e included
in he analysis [13].
His o ically sepa a ed h ee dai y b eeds Finnish
Ay shi e om Finland, Danish Red om Denma k and
Swedish Red om Sweden a e a p esen unde a join
b eeding alue es ima ion sys em, known as he No dic
Ca le Gene ic E alua ion [14]. P e ious QTL s udies o
milk ai s in No dic Red Ca le (NRC) ha e been done
wi hin he subpopula ions wi h mic osa elli e ma ke s
and ai ly small sample sizes (e.g. [15, 16]). The objec i e
o his s udy was o use a ia ions a he genome
sequence le el o ca y ou associa ion s udy o milk,
a and p o ein yields in NRC; o iden i y po en ial causal
a ian s and unde s and he gene ic a chi ec u e o hese
ai s. In addi ion, he da a p o ides he possibili y o
compa e he esul s o simila s udies o e ili y ai s
in he NRC [17], o e eal po en ial QTL wi h an agonis-
ic e ec s o milk p oduc ion and e ili y ai s.
Me hods
No animal expe imen s we e pe o med in his s udy, and,
he e o e, app o al om he e hics commi ee was no
equi ed. Semen samples we e collec ed o b eeding
pu poses by local o ganiza ions wi h app op ia e pe mi s.
Milk, a and p o ein yields’ ai de ini ions a e s an-
da dized ac oss he No dic coun ies. Pheno ypic eco ds
o dai y ca le a e housed in a cen alized da abase [14].
B eeding alues o milk, a and p o ein yield (MY, FY
and PY) a e based on p oduc ion igu es exp essed in
kilog ams aken om ou ine milk eco ds and hen com-
bined in o an index o each ai . Fo de ails on gene ic
e alua ion o milk yield ai s in No dic coun ies see
[18]. The b eeding alues used o associa ion analysis
we e de- eg essed b eeding alues om he ou ine gen-
e ic e alua ion by NAV (No dic ca le gene ic e alua ion)
and we e a ailable o 4280 p ogeny es ed NRC bulls
(2127 om Finland, 1217 om Sweden, 915 om
Denma k and 21 om o he coun ies). The eliabili ies o
he de eg essed b eeding alues we e in he ange o 0.67
o 0.99 wi h a mean o 0.95 and he i s qua ile a 0.94.
SNP a ay geno yping
All 4280 NRC bulls wi h de eg essed b eeding alues we e
geno yped using Bo ineSNP50 BeadChip SNP a ay
e sion 1 o 2 (Illumina Inc., San Diego, CA). DNA was
ex ac ed using s anda d p ocedu es om semen samples.
Chip ypings we e done by GenoSkan A/S, Tjele, Denma k
o labs belonging o Aa hus Uni e si y. The quali y
pa ame e s used o selec ion o SNPs we e minimum call
a es o 85 % o indi iduals and 95 % o loci. Ma ke loci
wi h mino allele equencies below 5 % and de ia ion
om Ha dy-Weinbe g p opo ion (P< 0.00001) we e
excluded. The minimal accep able GC sco e was 0.60 o
indi idual ypings, and indi iduals wi h a e age GC sco es
below 0.65 we e excluded. The numbe o SNP emaining
a e quali y con ol was 43,415 in he geno ypes ob ained
om Bo ineSNP50 BeadChip SNP a ay (50 K da a se ).
The genome posi ions o he SNPs we e acco ding o he
UMD3.1 Bo ine genome assembly [19].
Impu a ion o whole genome sequences
The 50 K geno ypes o hese bulls we e impu ed o
whole genome sequence da a using a wo-s ep app oach
[20]. Geno ypes om 50 K chip o each bull we e i s
impu ed o a high-densi y SNP a ay (HD) using a
mul i-b eed e e ence o 3383 animals (1222 Hols ein,
1326 NRC and 835 Danish Je sey indi iduals) which had
been geno yped wi h he Illumina Bo ineHD chip
(Illumina Inc., San Diego, CA). The numbe o SNPs,
a e impu a ion o he Bo ineHD chip, was 648,219.
These impu ed HD geno ypes we e subsequen ly impu ed
o he whole genome sequence le el using a mul i-b eed
e e ence panel o 1228 animals om Run4 o he 1000
Bull Genomes P ojec [13, 21] and addi ional whole
genome sequences om Aa hus Uni e si y [22] including
368 Hols ein, 86 RDC, 88 Je sey and es om numbe o
ca le b eeds. Da ase s (SNP a ay ypes and whole
sequence) we e p e-phased wi h BEAGLE4 1274 [23] and
geno ype impu a ion we e done using Minimac2 [24]. The
impu a ion accu acy o his da a is epo ed ea lie
[25], bu wi h a smalle whole genome sequence e e -
ence popula ion. Sequence a ian s ha ing impu a ion
accu acy
2
( a io o empi ically obse ed a iance o
Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 2 o 12
he allele dosages o he expec ed binomial a iance a
Ha dy-Weinbe g equilib ium and was ob ained om
Minimac2 so wa e ou pu ) less han 0.5 we e il e ed
away. The mean accu acy o he a ian s wi h
2
>0.5
was 0.94.
Associa ion analysis
The associa ion analysis o each o he impu ed sequence
a ian s (mino allele equency, MAF > 0.005 and de i-
a ion om Ha dy-Weinbe g p opo ion > 0.00001) was
ca ied ou using a wo-s ep a iance componen s-based
app oach o accoun o popula ion s a i ica ion imple-
men ed in he EMMAX so wa e ool [26]. In a i s s ep,
he polygenic and e o a iances a e es ima ed using
ollowing a iance componen model:
y¼1μþaþe
whe e yis a ec o o de- eg essed b eeding alues, 1is a
ec o o ones, μis he in e cep , Gis he kinship ma ix
buil based on high-densi y SNP geno ypes using EMMAX
so wa e, ais a ec o o b eeding alues assumed o ha e a
mul i a ia e no mal dis ibu ion a~N(0,Gσ
a
2
), eis a ec o
o andom esiduals assumed o ha e a mul i a ia e no mal
dis ibu ion e~N(0,Iσ
e
2
), whe e Iis an iden i y ma ix, σ
a
2
is
he addi i e gene ic a iance and σ
e
2
is he e o a iance.
In a second s ep, he SNP e ec is ob ained using a
linea eg ession model:
y¼1μþxb þη;
whe e xis a ec o o impu ed geno ype dosages
( anged be ween 0 and 2), bis he allele subs i u ion
e ec and ηis a ec o o andom esidual de ia es wi h
(co) a iance s uc u e Gσ
a
2
+Iσ
e
2
.
Sea ch o mul iple QTL in a genomic egion
To es i mul iple QTL a e seg ega ing in a genomic egion
we included he mos signi ican o he known causal
a ian as co ac o in he model and check o addi ional
QTL in a genomic egion o a yield on ch omosomes 14,
25 and 26 and o milk yield on ch omosome 14. We i ed
he SNPs (Addi ional ile 1) as ixed e ec o a linea mixed
model.
The s a is ical model is desc ibed by he o mula:
y¼1μþqsnp op þxgþZu þe
whe e y,1,μa e desc ibed as in he EMMAX model,
snp
op
is he e ec o he SNP i ed as co- ac o in he
model, and gis headdi i egene ice ec o he heSNP
unde s udy, qand xa e ec o s o SNP geno ype dosages
( anging om 0 o 2), and uis a ec o o andom poly-
genic e ec s, which a e no mally dis ibu ed u~ N(0, Aσ
u
2
),
whe e Ais he pedig ee-based addi i e ela ionship ma ix,
σ
u
2
is he polygenic a iance, Zis an incidence ma ix
ela ing pheno ypes o he co esponding andom poly-
genic e ec s, and eis a ec o o esidual e ec s, which a e
no mally dis ibu ed e~N(0,Dσ
e
2
), whe e Dis a diagonal
ma ix wi h elemen s d
ii
=(1−
DRP
2
)/
DRP
2
o accoun o
he e ogeneous esidual a iances due o di e en eliabil-
i ies o DRP (
DRP
2
), and σ
e
2
is he esidual a iance. Analyses
we e pe o med using he DMU package [27]. Signi i-
cance es ing o SNP e ec s was pe o med using a
wo-sided - es . The null hypo hesis was g = 0. A e
ha , he Bon e oni co ec ion was applied same as
in he EMMAX analysis o con ol o alse posi i e
associa ions.
The genome-wise signi icance h eshold co esponding
o an e o a e o 0.05 was se a 3.16 × 10
−9
a e co ec-
ion o mul iple es ing using a Bon e oni co ec ion o
15,679,852, 15,679,853 and 15,679,844 independen es s
o a , milk and p o ein yield espec i ely. Only SNPs
wi h he p alue less han 3.16 × 10
−9
(−log
10
(p)≥8.50)
we e anno a ed wi h he a ian e ec p edic o (VEP)
ool using he Ensembl da abase, Release 82 [28]. The
p edic ion whe he an amino acid subs i u ion caused by
missense a ia ion a ec s p o ein unc ion was es ima ed
by SIFT analysis [29] implemen ed in VEP ool [28]. The
SIFT p edic ion is based on sequence homology and he
physical p ope ies o amino acids. Manha an plo s we e
c ea ed wi h he qqman .0.1.2 R package [30]. In
addi ion, we compa ed ou indings o esul s ob ained
om s udy by Höglund e al. [17] whe e a simila
genome-wide associa ion s udy o emale e ili y in
No dic Red ca le was conduc ed.
Ingenui y pa hway analysis and en ichmen analysis
Lis s o genes wi h signi ican hi s om he QTL peak
egions associa ed wi h each milk ai we e uploaded
in o he Qiagen’s Ingenui y® Pa hway Analysis IPA® [31].
Fo his pu pose, also SNPs signi ican ly associa ed wi h
Fe ili y index (FI) in he No dic Red Ca le [17] we e
anno a ed wi h he VEP ool [28] and genes ha ing one
o mo e signi ican SNP we e analyzed wi h IPA® [31].
Bioma ool [32] embedded in Ensembl da abase [33],
was used o sea ching human homologs o cow
Ensembl IDs o he genes. In case he e was mo e han
one, all epo ed human o hologs we e kep . The
human homologue lis s o each ai included 214, 69,
66 and 263 genes o a , milk, p o ein yields and e ili y
index, espec i ely. A e unning co e analysis o
each ai , ne wo ks based on he in o ma ion o gene
connec i i y in Ingenui y Knowledge Da abase wi h
highes sco e- alues we e conside ed. Sco e- alue
ep esen s he nega i e log o he p- alue o he likeli-
hood ha he molecules would be ound oge he by
chance.
Gene on ology (GO) e m en ichmen analysis wi h
genes ound wi hin he op SNPs was pe o med wi h a
Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 3 o 12
Singula En ichmen Analysis (SEA, Fishe ’s exac es ,
FDR < 0.01) p o ided by Ag iGO webpage [34].
Resul s
Depending on he ai we could iden i y se e al hou-
sands (3594, FY), less han a housand (755, MY) o less
han a hund ed (85, PY) signi ican ly associa ed SNPs
(Bon e oni co ec ed h eshold o signi icance
-log
10
(p)≥8.50; Addi ional iles 2, 3 and 4). No common
signi ican ly associa ed SNPs we e ound be ween his
s udy and wi h hose ound o emale e ili y ai s
[17]. Howe e , signi ican ly associa ed SNPs we e ound
in i e common genes be ween e ili y and a yield.
These i e genes a e loca ed on ch omosomes 25
(ENSBTAG00000034643) and 26 (GBF1,TMEM180,
ACTR1B, and b a-mi -146b).
The summa y o he anno a ions o signi ican ly associ-
a ed SNPs a e p esen ed in Table 1. Among anno a ed
SNPs, in on a ian s a e he mos common ype o each
o he ai s. SNPs changing amino acid (missense a ia-
ions) a e a e, 21 o FY, 1 o MY, 3 o PY and 18 sha ed
be ween FY and MY. Se en missense a ia ions we e
p edic ed by SIFT analysis [29] o be dele e ious, i.e.
po en ially leading o changes in he unc ion o he
p o ein (Addi ional ile 5). Splice egion a ian s a e a e
o ming only one pe cen o less om he o al amoun o
signi ican ly associa ed SNPs (Addi ional ile 5).
Resul s o a yield wi hin he la ge associa ed a eas
on ch omosomes 14, 25 and 26 we e u he examined
by ixing he e ec o op SNP(s). Only peaks ha
emained a e ixing he o he op SNPs we e consid-
e ed as po en ial QTL in hese ch omosomes.
Se en, eigh and ou sepa a e QTL egions (“peaks”)
we e de ined o a , milk and p o ein yield, espec i ely
(Fig. 1, Addi ional iles 6, 7, 8, 9, 10, 11, 12 and 13). The
peaks we e de ined as con inuous egions con aining
SNPs ha ing –log
10
(p)≥8.50. Top SNPs wi h he conse-
quences o each de ined QTL egion pe ai a e lis ed
in Table 2.
The highes peak was obse ed on Bos au us
ch omosome 14 (BTA14) (Addi ional ile 7) spanning
he egion om 1,448,510 bp o 2,271,832 bp o a
yield (509 SNPs ha ing –log
10
(p)≥8.50) and om
1,448,510 bp o 2,271,832 bp o milk yield (455 SNPs
ha ing –log
10
(p)≥8.50). The highes –log
10
(p) alues
wi hin hese egions we e ob ained o SNPs s136783505
(bp 1,807,140) o a yield and s133033480 (bp 1,743,939)
o milk yield. The highes peak o p o ein yield was
loca ed on BTA25 (Addi ional ile 11) wi hin he egion
3,306,363–3,516,671 which con ained 40 SNPs, he highes
p alue being o SNP s110749311 (bp 3,498,960).
Fa yield
We iden i ied se en di e en QTL egions on i e di e en
ch omosomes a ec ing a yield (Table 2, Fig. 1, Addi ional
ile 2). The s onges associa ion ound o a yield was lo-
ca ed on BTA14 (Addi ional ile 7) in he DGAT1 (Diacyl-
glyce ol O-acyl ans e ase 1) gene egion. In ou da a he
s onges associa ion is o he a ia ion s136783505 (bp
1,807,140) wi h no unc ional anno a ion, loca ed 2578 bp
downs eam o he DGAT1 gene (Table 2). Howe e , se -
e al o he a ian s loca ed nea by, including he p e iously
iden i ied causa i e a ian K232A [35] a bp 1,802,266,
show simila ly high -log
10
(p) (Addi ional ile 2). To in es i-
ga e he signi icance o o he SNPs in he egion, we i ed
he a ia ion K232A as ixed e ec . None o he o he SNPs
emained signi ican (Addi ional ile 14) a e he ixa ion
o he K232A a ia ion.
O he SNPs wi h s ong associa ions wi h a yield
we e ound on BTA5 (Addi ional ile 6), BTA23
(Addi ional ile 11), BTA25 (Addi ional ile 12) and
BTA26 (Addi ional ile 13).
The QTL egion on BTA5 is loca ed be ween
92,372,732 bp and 94,425,668 bp and he a ia ion wi h
he s onges associa ion ( s209818856, pos. 93,945,694) is
loca ed in an in on o he gene MGST1 (Table 2). On
BTA23 he associa ion signal o a yield comes om he
egion 28,567,796–28,591,530, he op a ia ion being lo-
ca ed a bp 28,567,796 in he in on o gene TRIM26
(Table 2).
On BTA25 and BTA26, complex pa e ns o associ-
a ion we e seen (Addi ional iles 12 and 13). To cla i y
he numbe o independen QTL wi hin hese egions
we in es iga ed he signi icance o he SNPs by i ing
Table 1 The numbe o signi ican SNPs (- log
10
(p)≥8.50) o each ai and how SNPs a e di ided in o di e en consequences. SNPs
we e anno a ed wi h he a ian e ec p edic o – ool [28]. One SNP can ha e mo e han one anno a ion
T ai Numbe o
signi ican
SNPs
In on
a ian
In e genic
a ian
Downs eam
gene a ian
Ups eam gene
a ian
Synonymous
a ian
Missense
a ian
3′UTR
a ian
Splice egion
a ian , in on
a ian
5′UTR
a ian
Splice egion
a ian ,
synonymous
a ian
Fa yield 3594 1641 1307 600 551 94 40 39 9 6 2
Milk yield 755 322 195 301 240 42 20 16 5 1 2
P o . yield 85 50 14 32 12 3 3 1 1
Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 4 o 12
he op SNPs ( i e o BTA25 and eigh o BTA26,
Addi ional ile 1) as ixed e ec s alone and in di e en
combina ions. On BTA 25, signi ican associa ions
emained a bp 9,870,005 (in onic egion o he
CLEC16A gene) and a bp 36,226,978 (in e genic egion)
(Table 2). Two QTL emained also on BTA26, one in
he egion o he NEURL1 gene (posi ion 24,379,571)
and he o he in an in e genic egion ( op SNP a bp
44,802,991) (Table 2).
Milk yield
In all, eigh QTL egions we e ound o milk yield
(Table 2, Fig. 1, Addi ional ile 3). They we e loca ed on
six di e en ch omosomes (BTA5, BTA14, BTA16,
BTA19, BTA20 and BTA25). The s onges associa ion
wi h milk yield was ound on BTA14 (1,448,510–
2,271,832) ha ing op a ia ion loca ed a bp 1,743,939.
This a ia ion is wi hin wo o e lapping genes, CPSF1
and ADCK5.TheQTL egionis hesameaswas ound
associa ed wi h a yield bu wi h di e en op a i-
a ion. As o a yield, he signi icance o o he han he
p e iously known causa i e DGAT1 a ia ion [35] was
es ed by i ing he a ia ion K232A as a ixed e ec .
None o he o he SNPs emained signi ican a e
ixing he K232A e ec (Addi ional ile 15).
The weakes signi ican associa ion was loca ed on BTA5
wi h op a ia ion s383553819 (posi ion 112,343,204)
loca edinanin ono hegeneMKL1 (Addi ional ile 6).
On BTA19 he associa ed egion con ained no anno a ed
genes (61,447,138–61,449,096), and he op a ia ion
s210324693 (bp 61,449,096 bp) was loca ed in an in e -
genic egion (Table 2, Addi ional ile 9). Two QTL egions
we e de ec ed on BTA20 (Addi ional ile 10). The QTL
we e loca ed in he egions 30,531,217–32,952,019 and
37,766,226–39,183,141. The known causa i e a ia ion
F279Y (bp. 31,909,478) o milk ai s in he gene GHR
[36]was he opSNPinou analysis o heQTLlo-
ca ed in he i s egion. The op a ia ion wi hin he
second QTL was loca ed a bp 38,828,254 in he
in e genic egion, bu his QTL egion also includes
he PRLR gene p e iously indica ed o be linked o
milk p oduc ion (e.g. [37]). BTA25 ha bo s wo QTL,
he i s a bp 2,669,704 and he second in he egion
om 3,494,706 bp o 3,516,671 bp, he op SNP lo-
ca ed a bp 3,498,960 downs eam om he gene
PAM16 (Addi ional ile 12).
Fig. 1 Genome-wide Manha an plo s o he a yield (FY), milk yield (MY) and p o ein yield (PY). Red line indica es he genome-wide signi icance
le el (−log
10
(p) = 8.50)
Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 5 o 12

Table 2 QTL egions o each o he ai . Top SNP o each QTL a e shown including posi ion, -log
10
(p)- alues, mino allele equency (MAF), gene in o ma ion, anno a ion o
he op SNP, allele subs i u ion e ec (b. alue) and s anda d e o o b. alue (SE)
CHR S a (bp) End (bp) Leng h o he
QTL egion (bp)
Signi ican SNPs
in egion
No o genes wi h
signi ican SNPs
wi hin he QTL egion
Top SNP Posi ion o he
op SNP (bp)
-log
10
(p) MAF Gene Anno a ion o
he op SNP
b. alue SE
Fa yield
5 92,372,732 94,425,668 2,052,936 330 3 s209818856 93,945,694 27.49 0.38 MGST1 in on a ian −2.807 0.253
14 1,448,510 2271832 823,322 509 49 s136783505 1,807,140 42.01 0.07 DGAT1/HSF1 downs eam a ian /
in on a ian
−6.709 0.484
23 28,567,796 28,591,530 23,734 5 1 s381390819 28,567,796 9.36 0.45 TRIM26 in on a ian 1.148 0.184
25 8,222,347 11,507,986 3,285,639 883 16 s379546164 9,870,005 14.40 0.24 CLEC16A in on a ian −1.769 0.224
25 36,226,978 36,227,132 154 2 s109480808 36,226,978 8.72 0.16 in e genic a ian 1.635 0.272
26 22,144,777 24,793,744 2,648,967 595 32 s438420348 24,379,571 14.28 0.16 NEURL1 in on a ian −2.273 0.290
26 44,802,991 44,802,991 0 1 s135624939 44,802,991 9.72 0.28 in e genic a ian −1.474 0.231
Milk yield
5 112,343,204 11,2450,860 107,656 2 1 s383553819 112,343,204 8.70 0.36 MKL1 in on a ian 1.437 0.239
14 1,448,510 2,271,832 823,322 455 48 s133033480 1,743,939 33.01 0.09 CPSF1/ADCK5 downs eam a ian /
splice egion a ian ,
in on a ian
6.266 0.513
16 1,322,611 1,322,611 0 1 1 s108979795 1,322,611 8.58 0.28 LAX1 ups eam a ian −1.451 0.243
19 61,447,138 61,449,096 1958 5 s210324693 61,449,096 8.93 0.32 in e genic a ian 1.490 0.244
20 30,531,217 32,952,019 2,420,802 74 5 s385640152 31,909,478 15.56 0.11 GHR missense a ian −3.877 0.472
20 37,766,226 39,183,141 1,416,915 34 3 NA 38,828,254 9.54 0.16 in e genic a ian −2.250 0.356
25 2,669,704 2,669,704 0 1 s209691835 2,669,704 9.21 0.13 in e genic a ian −2.855 0.460
25 3,494,706 3,516,671 21,965 13 4 s110749311 3,498,960 9.32 0.41 PAM16/GLIS2 downs eam a ian 1.225 0.196
P o ein yield
5 112,450,860 112,450,860 0 1 s109041054 112,450,860 8.88 0.48 in e genic a ian −1.473 0.242
14 1,802,667 1,802,667 0 1 2 NA 1,802,667 8.52 0.06 DGAT1/HSF1 in on a ian /
downs eam a ian
3.354 0.564
25 1,094,996 1,257,612 162,616 12 3 s136085792 1,103,856 10.83 0.22 UNKL in on a ian 1.694 0.250
25 3,306,363 3,516,671 210,308 40 8 s110749311 3,498,960 11.70 0.41 PAM16/GLIS2 downs eam a ian 1.427 0.202
Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 6 o 12
P o ein yield
In all, p o ein yield did no show as many signi ican ly
associa ed SNPs as obse ed o a and milk yield
(Table 2, Fig. 1, Addi ional ile 4). Ch omosomes ha ing
QTL o p o ein yield we e BTA5, BTA14, and BTA25.
Fo bo h BTA5 and BTA14, only one a ian o each
eached he signi icance cu -o le el; on BTA5 an in e -
genic a ia ion a bp 112,450,860 (Addi ional ile 6) close
o he a ian ound o milk yield and on BTA14, he
a ia ion (bp 1,802,667) loca ed in an in on o he
DGAT1 gene (Addi ional ile 7).
The wo o he QTL o p o ein yield we e bo h lo-
ca ed on BTA25. The QTL on BTA25 a 3,306,363–
3,516,671 o e lapped wi h he QTL ound o milk
yieldanddisplayed hesame opSNP(bp3,498,960).
The o he QTL on BTA25 was unique o p o ein
yield, op SNP (bp 1,103,856) loca ed in he gene
UNKL.
Ne wo ks o associa ed genes and en ichmen analysis
The ne wo ks wi h highes sco es o each ai a e
p esen ed in Addi ional iles 16, 17 and 18. The wo
op ne wo ks om genes associa ed wi h a yield
(Addi ional ile 16) had sco es o 49, he ne wo k associ-
a ed wi h ca bohyd a e me abolism, gene exp ession and
lipid me abolism is p esen ed in Addi ional ile 16a.
Addi ional ile 16b shows he ne wo k gene a ed om he
genes associa ed o e ili y index. The sco e alue o his
ne wo k is 41 and i consis s o genes associa ed wi h
in lamma o y esponse, cell- o-cell signaling and lymphoid
issue s uc u e and de elopmen . E en hough common
signi ican ly associa ed SNPs we e no ound be ween his
s udy and ha o [17] on emale e ili y ai s, signi ican ly
associa ed SNPs we e ound in i e common genes. Two
o hem (GBF1 and b a-mi -146b) a e p esen in bo h a
and e ili y ne wo ks (Addi ional ile 16a and b). The op
ne wo k o milk yield (sco e 43, Addi ional ile 17) was
associa ed wi h unc ions molecula anspo , o gan
mo phology and o ganismal de elopmen . I includes he
known milk candida e genes DGAT1,GHR and PRLR,as
well as he op hi s om BTA5 (MLK1) and BTA16
(LAX1). The op ne wo k o p o ein yield (sco e 35,
Addi ional ile 18) is connec ed o unc ions cell dea h
and su i al, cance and o ganismal inju y.
Al oge he 18, 18, 11 and 16 GO e ms we e signi i-
can ly ( alse disco e y a e, FDR < 0.01) en iched o FY,
MY, PY and FI, espec i ely. A b oad GO e m, mul i-
cellula o ganismal p ocess, was he mos signi ican
o all ou ai s, o ally eigh e ms we e sha ed
be ween hem. Fo example all ai s we e ha ing QTL
in he egions con aining signi ican en ichmen o
genes ela ed o ep oduc ion and ep oduc i e p o-
cesses (Table 3).
Discussion
A la ge numbe o a ian s we e ound signi ican ly
associa ed wi h milk, a and p o ein yield. Ou indings
suppo p e ious QTL indings om he No dic Red
b eeds, e.g. on BTA 5, 14, 20 [37, 38] and loca e new
a ia ions ha a e good candida es o be causa i e
a ia ions. This is he i s ime when NRC popula ion is
s udied wi h impu ed whole genome sequence a ian s
in o de o e ine QTL associa ed wi h milk p oduc ion.
Howe e , i is s ill di icul o pinpoin he causa i e
a ian among se e al closely linked, almos equally
signi ican ly associa ed a ia ions. One way o classi y
he a ia ions is o look a he p edic ed unc ional con-
sequence o he SNP [29]. The possibili y ha a a ia ion
has an impac on he pheno ype is highe i he a ia ion
causes an amino acid change (missense a ia ion) which
is p edic ed (e.g. wi h SIFT analysis, [29]) o ha e an
e ec o p o ein unc ion, is loca ed on splicing si e, o
is loca ed downs eam o ups eam o he known gene
(possible egula o y egions o he ansc ip ion). On he
o he hand, genome anno a ion o ca le is s ill incom-
ple e and mos egula o y elemen s emain unknown. In
he sea ch o biologically ele an ma ke s he in o ma-
ion o in e ac ions be ween genes in known pa hways
o ne wo ks can be use ul. In his s udy, we iden i ied some
in e es ing gene in e ac ion ne wo ks based on he signi i-
can ly associa ed a ian s wi hin genes (e en hough he
unc ional e ec s o he a ian s could no be p edic ed).
The esul smaybeused oha eaclose looka alsoo he
genes in he indica ed ne wo ks o unc ional a ian s.
Al hough no common SNPs we e ound associa ed wi h
milk p oduc ion ai s and e ili y, he i e common genes
be ween a yield and e ili y gi e some indica ion o he
ela ionship be ween hose ai s. The genes b a-mi -146b
and GFB1 a e associa ed wi h a e ili y gene ne wo k
linked wi h in lamma o y esponse and cell- o-cell signal-
ing and he a yield ne wo k connec ed wi h lipid and
ca bohyd a e me abolism. Fu he suppo was gained
om he gene en ichmen analysis, bo h he ai s show
signi ican en ichmen o he genes ela ed o example o
ep oduc ion and ep oduc ion and ep oduc i e p o-
cesses, al oge he ha ing 16 common GO e ms.
F om he ou ch omosomes epo ed o ha bo highes
numbe o QTL o milk p oduc ion [39], wo we e indi-
ca ed by ou da a (BTA14 and BTA20). S ucken e al. [40]
summa ized 14 genes om en di e en ch omosomes o
be he majo genes in ol ed in milk p oduc ion. Among
hose genes a e DGAT1 and GHR. Some commonly ound
QTL (e.g. BTA6, [41]) we e no seen in ou da a; ha could
be due o ixa ion o he QTL o e y low MAF in he
NRC popula ion. One explana ion could be ha he
EMMAX me hod chosen o associa ion analysis migh be
oo conse a i e. EMMAX uses app oxima ions o con-
s uc ing es o he ixed SNP e ec s o in e es in he
Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 7 o 12
linea mixed model because i ing a ull linea mixed
model o each SNP in u n ac oss he genome is compu a-
ionally challenging [42]. This leads o sys ema ic unde -
es ima ion o he mos signi ican p alues [43], bu makes
EMMAX one o he as es LMM based p og ams [42]; an
a gumen ha has o be conside ed when ha ing whole
genome sequence le el da a om subs an ially la ge
amoun o indi iduals.
DGAT1 (BTA14)
The s onges signal o associa ion was ound om
BTA14 o a and milk yield. Also p o ein yield is
signi ican ly associa ed wi h he same egion on
BTA14. The op SNP a ied depending on he ai
(see Table 2). Dae wyle e al. [13] analyzed associ-
a ion o ea ly lac a ion milk a pe cen age wi h whole
genome sequence a ia ion da a in Fleck ieh and
Hols ein bulls. As in ou s udy, he p e iously epo ed
causa i e a ia ion K232A (bp 1,802,266, [35]) o DGAT1
gene was no he a ian wi h he lowes p alue in Hol-
s ein and Fleck ieh o milk a p oduc ion al hough
K232A was among he op SNPs. Pausch e al. [44] used
whole genome sequence da a o impu e Ge man Fleck ieh
and Hols ein-F iesian ca le geno ypes om a la ge se o
animals candida e egions and we e able o con i m he
associa ion o K232A mu a ion, howe e , hei app oach
was biased as only known candida e SNPs we e es ed o
associa ion.
In ou da a, close inspec ion o he associa ed a ia ions
(−log
10
(p)≥8.50, a yield) in he DGAT1 egion e ealed
ha when he e ec o he K232A mu a ion was ixed, no
addi ional s a is ically signi ican SNP e ec s we e le . The
causa i e mu a ion (K232A/ s109326954) o DGAT1 was
epo ed al eady o e a decade ago [45] and has been
unc ionally con i med [35]. The K allele inc eases milk a
pe cen age [35], whe eas allele A inc eases milk p oduc ion
[46]. The e a e di e en possible explana ions why K232A
did no u n ou o be he mos signi ican ly associa ed
a ia ion in his s udy. Impe ec impu a ion may a ec
associa ion esul s. Accu acy o he impu a ion is conside -
ably imp o ed by inc easing he size o e e ence panel, i.e.
sequenced animals [47] and impu a ion accu acy seems o
be highe when popula ions unde s udy a e combined o
he impu a ion p ocesses [13]. Ou e e ence panel
consis ed o a mul i-b eed popula ion wi h 1228 indi-
iduals om se e al b eeds including bo h dai y and
bee ca le. The DGAT1 egion would be an in e es -
ing candida e o s udy wi h he in o ma ion om
1000 Bulls Genomes P ojec [13, 21]. I would gi e a
chance o s udy he haplo ype s uc u e o he egion
in he ca le popula ion wo ldwide and possibly ace
back he e olu ion o he QTL e ec .
Table 3 GO en ichmen e ms ha ing FDR < 0.01 om he genes ha ing signi ican a ia ions (−log
10
(p)≥8.50) o a yield (FY), milk
yield (MY), p o ein yield (PY) and e ili y index (FI)
FY MY PY FI
GO e m Desc ip ion FDR FDR FDR FDR
GO:0022610 Biological adhesion 6.20E-07 0.0013 2.00E-06
GO:0065007 Biological egula ion 9.40E-12 6.80E-05 0.072 2.40E-28
GO:0044085 Cellula componen biogenesis 3.60E-09 0.048 1.00E-22
GO:0016043 Cellula componen o ganiza ion 1.00E-37 2.90E-12 0.00027 1.80E-76
GO:0009987 Cellula p ocess 4.90E-14 4.60E-05 0.0074 3.80E-20
GO:0016265 Dea h 7.80E-21 2.00E-07 4.00E-25
GO:0032502 De elopmen al p ocess 8.80E-102 2.10E-41 1.40E-14
GO:0051234 Es ablishmen o localiza ion 1.80E-14 2.80E-05 0.0057 1.70E-27
GO:0040007 G ow h 5.60E-42 2.10E-17 9.40E-45
GO:0002376 Immune sys em p ocess 0.00022 0.0032 9.80E-10
GO:0051179 Localiza ion 5.90E-20 1.10E-07 0.0059 4.70E-44
GO:0008152 Me abolic p ocess 9.40E-12 0.00035 0.57 7.80E-18
GO:0032501 Mul icellula o ganismal p ocess 8.00E-126 1.10E-44 1.50E-20 1.20E-161
GO:0048519 Nega i e egula ion o biological p ocess 2.80E-50 1.30E-25 1.20E-07
GO:0048518 Posi i e egula ion o biological p ocess 1.70E-32 1.60E-08
GO:0050789 Regula ion o biological p ocess 7.10E-10 0.00011 0.32 6.90E-22
GO:0000003 Rep oduc ion 3.70E-57 3.70E-14 3.50E-09 1.20E-54
GO:0022414 Rep oduc i e p ocess 1.70E-50 7.50E-11 1.90E-07 6.30E-38
GO:0050896 Response o s imulus 6.00E-31 5.50E-06 8.80E-07 8.60E-44
Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 8 o 12
MGST1 and MKL1 (BTA5)
Vii ala e al. [16] showed ha Finnish Ay shi e has a milk
p oduc ion QTL a he p oximal end o BTA5. Wang
e al. [48] epo ed o a QTL o milk a pe cen age in
Ge man Hols ein-F iesian a he loca ion o 94,551,792 bp
and sugges ed a candida e gene o be EPS8. One o he
mos signi ican SNP in he s udy by Aliloo e al. [11] o
milk yield (Je sey and Hols ein) was loca ed a bp
94,518,850 on BTA5. Fo a yield, we obse ed an associ-
a ion peak a bp 93,945,694 (Table 2). This a ia ion is in
in onic egion o he gene MGST1 ha ing a ole in
oxida i e s ess eac ion [49] QTL peak o he milk
yield on BTA5 is loca ed wi hin an in on o he
MKL1gene (bp 112,343,204). MKL1 is ela ed o an-
sc ip ion egula ion. P o ein yield associa ion peak is
loca ed a bp 112,450,860 on BTA5 and is anno a ed
o he non-coding egion. P e iously QTL ela ed o
body weigh ha e been mapped nea by [50].
GHR and PRLR (BTA20)
BTA20 is among he ch omosomes ha bo ing many QTL
ela ed o milk p oduc ion [39]. GHR has been epo ed as
one o he majo genes in ol ed in milk p oduc ion [40].
Fi s Blo e al. [36] ound ha a ia ion F279Y (bp
31,909,478) in he GHR gene is associa ed wi h a s ong
e ec on milk yield and composi ion and o he s udies
ha e con i med i (e.g. [44]). O he a ia ions han F279Y
ha e also been ound o a ec milk p oduc ion nea by he
GHR gene [38]. In ou s udy, F279Y has he s onges
associa ion o milk yield among he NRC popula ion on
BTA20. Fu he con i ma ion o he causali y o he
F279Y comes om SIFT analysis p edic ing he mu a ion
o be dele e ious o p o ein s uc u e, i.e. po en ially al e -
ing he p o ein s uc u e hus possible leading o changes
in unc ion o he p o ein. The op SNP (bp 31,909,478) in
he GHR gene egion is clea ly he mos signi ican one
(Addi ional ile 10) in con as o DGAT1 egion whe e
se e al a ia ions a e s ongly associa ed o milk yield
(Addi ional ile 7). The Y allele is p edic ed o be un a o -
able o milk p oduc ion [36], bu i is s ill ai ly common
in he NRC popula ion. GHR has been sugges ed o be
unde balancing selec ion because o he obse ed high
a ia ion in he cy oplasmic egion [51].
Ano he gene on BTA20 o special in e es is PRLR and
he a ia ion S18N (posi ions 39,115,344-39,115,345) [37].
Howe e , i has been sugges ed ha S18N is a he linked
o he causa i e mu a ion han being causa i e i sel [44].
We ound ha he a ia ion a bp 38,828,254 on BTA20
loca ed in he in e genic egion, was indica ed o be he
mos likely candida e esponsible o he QTL e ec seen.
This a ia ion is loca ed app oxima ely 245,000 base pai s
downs eam o he PRLR gene and addi ional s udies a e
equi ed o esol e he mechanisms how i may in luence
milk p oduc ion.
TRIM26 (BTA23)
An in onic a ian in he gene TRIM26 a bp 28,567,796
on BTA23 has an associa ion wi h a yield. The unc ion
o he TRIM26 gene, a membe o he ipa i e mo i
(TRIM) gene amily, is unknown [52]. I is loca ed close o
he majo his ocompa ibili y complex (MHC) class I
egion. Feed in ake QTL ha e been mapped close o he
associa ion peak obse ed in his s udy [53].
PAM16,UNKL and CLEC16A (BTA25)
Al oge he six associa ion peaks (QTL) we e obse ed
om BTA25 o di e en ai s. The same a ia ion
(bp 3,498,960) in gene PAM16 is associa ed wi h bo h
milk and p o ein yield. The gene has a c i ical ole in
p o ein ansloca ion ac oss he inne mi ochond ial
memb ane [54]. O he QTL on BTA25 we e ound
o milk yield in he in e genic egion (bp 2,669,704)
and o p o einyieldQTLa bp1,103,856in hegene
UNKL ha has a ole on p o ein, zink ion and me al
ion binding. Two dis inc QTL we e iden i ied o a
yield, peak a ia ions loca ed a posi ions 9,870,005 and
36,226,978. Va ia ion a bp 9,870,005 is loca ed in he
CLEC16A gene (Table 2). Va ia ions o CLEC16A in
humans a e associa ed wi h inc eased ype I diabe es isk
[55]. In addi ion, milk p o ein pe cen age QTL has p e i-
ously been ound om he egion 9.3–10.6 Mb [56].
NEURL1 (BTA26)
A e he signi icance es by ixing o he op SNPs, wo
QTL we e le on BTA26 o a yield a bp 24,379,571
in NEURL1 gene and a bp 44,802,991 (Table 2).
NEURL1 gene is associa ed wi h lac a ion (GO e m
0007595) hus making he a ia ion an in e es ing candi-
da e o be a causa i e mu a ion.
Conclusions
Associa ion analyses among No dic Red Ca le using
o e 15 million sequence a ia ions ac oss he whole
genome impu ed o o e 4000 p ogeny es ed No dic
Red Ca le bulls indica ed se e al a ia ions likely o
ha e an impac o milk p oduc ion. We show ha
impu a ion is obus and cos -e ec i e way o expand
he in o ma ion a ailable and o inc ease knowledge o
he causa i e mu a ions a ec ing ai s impo an o p o-
duc ion animals. The a ailabili y o he whole genome
le el sequence da a opens endless possibili ies o s udy
quan i a i e ai a chi ec u e mo e closely. S ill inding he
quan i a i e ai nucleo ides is challenging, wi h linkage
disequilib ium and many small-e ec QTL c ea ing he
puzzle ha is no easy o sol e. Fu he mo e, be e anno-
a ion o he ca le genome is equi ed o be able o p edic
he e ec s o a ia ions on he pheno ypes mo e accu -
a ely. The knowledge om gene in e ac ions (al hough
human/ oden based) may help o iden i y likely candida e
Iso-Tou u e al. BMC Gene ics (2016) 17:55 Page 9 o 12