E olu iona y Applica ions 2016; 9: 1017–1031 wileyonlinelib a y.com/jou nal/e a
|
1017
© 2016 The Au ho s. E olu iona y Applica ions
published by John Wiley & Sons L d
Recei ed: 20 No embe 2015
|
Accep ed: 7 June 2016
DOI: 10.1111/e a.12407
This is an open access a icle unde he e ms o he C ea i e Commons A ibu ion License, which pe mi s use, dis ibu ion and ep oduc ion in any medium,
p o ided he o iginal wo k is p ope ly ci ed.
Abs ac
Many wild A lan ic salmon (Salmo sala ) popula ions a e h ea ened by in og essi e
hyb idiza ion om domes ica ed ish ha ha e escaped om aquacul u e acili ies. A
de ailed unde s anding o he hyb idiza ion dynamics be ween wild salmon and aqua-
cul u e escapees equi es disc imina ion o di e en hyb id classes; howe e , ma ke s
cu en ly a ailable o disc imina e he wo ypes o pa en al genome ha e limi ed
powe o do his. Using a high- densi y A lan ic salmon single nucleo ide polymo phism
(SNP) a ay, in combina ion wi h pooled- sample allelo yping and an Fs ou lie
app oach, we iden i ied 200 SNPs ha di e en ia ed an impo an A lan ic salmon
s ock om he escapees po en ially hyb idizing wi h i . By simula ing mul iple gene a-
ions o wild–escapee hyb idiza ion, in ol ing wild popula ions in wo majo phyloge-
og aphic lineages and a gene ically di e se se o escapees, we showed ha bo h he
comple e se o SNPs and smalle subse s could eliably assign indi iduals o di e en
hyb id classes up o he hi d hyb id (F3) gene a ion. This se o ma ke s will be a use-
ul ool o in es iga ing he gene ic in e ac ions be ween na i e wild ish and aqua-
cul u e escapees in many A lan ic salmon popula ions.
KEYWORDS
allelo yping, aquacul u e escapee, A lan ic salmon, in og essi e hyb idiza ion, Salmo sala , SNP
a ay
1Depa men o Biology, Uni e si y o Tu ku,
Tu ku, Finland
2Na u al Resou ces Ins i u e Finland (Luke),
U sjoki, Finland
3Cen e o In eg a i e Gene ics
(CIGENE), Depa men o Animal and
Aquacul u al Sciences, No wegian Uni e si y
o Li e Sciences, Aas, No way
Co espondence
Vic o ia L. P i cha d, Depa men o Biology,
Uni e si y o Tu ku, Tu ku FI-20014, Finland.
Emails: ic o ialp i cha [email p o ec ed]; c aig.
[email p o ec ed]
ORIGINAL ARTICLE
Single nucleo ide polymo phisms o disc imina e di e en
classes o hyb id be ween wild A lan ic salmon and aquacul u e
escapees
Vic o ia L. P i cha d1 | Jaakko E kina o2 | Ma hew P. Ken 3 | Ee o Niemelä2 |
Panu O ell2 | Sigbjø n Lien3 | C aig R. P imme 1
1 | INTRODUCTION
In common wi h many o he salmonid ishes (e.g. Ka z, Moyle,
Quiñones, Is ael, & Pu dy, 2013; Me cal e al., 2007; Rand, 2013), wild
A lan ic salmon (Salmo sala ) ha e declined o e he pas wo cen u ies
as a esul o o e ishing and habi a loss. The species has been ex i -
pa ed om hal o he majo Eu opean i e basins and a hi d o he
majo Ame ican i e basins in which i his o ically occu ed, and many
o he emaining popula ions a e conside ed o be unde h ea (ICES
2015; Pa ish, Behnke, Gepha d, McCo mick, & Ree es, 1998). In con-
as , since a i icial cul i a ion began in No way in he la e 1960s, he
cap i e A lan ic salmon popula ion has exploded. In Eu ope o e 1,000
imes mo e salmon is cu en ly p oduced by he aquacul u e indus y
han is caugh in he wild (ICES 2015). Mos o hese domes ic A lan ic
salmon a e ea ed in open ne - pens in he ma ine en i onmen , and
escapes a e equen (Tho s ad e al., 2008). In No way, he wo ld’s
bigges p oduce o a med salmon, ens o housands o aquacul u e
escapees, iden i ied by body mo phology and scale cha ac e is ics,
1018
|
P i cha d e al.
a e caugh annually in he wild. Be ween 1989 and 2006, on a e age,
38% o he salmon ca ch o No wegian coas al ishe ies and 21% o
he indi iduals sampled in spawning a eas we e aquacul u e escapees
(Tho s ad e al., 2008).
Aquacul u e salmon ha e unde gone many gene a ions o selec-
ion in he cap i e en i onmen –bo h inad e en ly, due o ha che y
cul u e condi ions, and delibe a ely. Some No wegian aquacul u e
s ains, o example, ha e been subjec o a i icial selec ion o
mul iple ai s including speed o g ow h, weigh and age a sexual
ma u i y, and esis ance o disease (Bicskei, B on, Glo e , & Tagga ,
2014; Fleming, Agus sson, Fins ad, Johnsson, & Bjö nsson, 2002;
Gjøen, 1997). Co espondingly, a med salmon ha e been shown o
di e om hei wild conspeci ics in mul iple ai s including g ow h
a e and age a ma u a ion (Glo e , O e å, Olsen, & Slinde, 2009;
Debes & Hu chings, 2014), mig a o y beha iou (Jonsson, Jonsson,
& Hansen, 1991), hea a e and swimming endu ance (Johnsson,
Höjesjö, & Fleming, 2011), gene exp ession (Bicskei e al., 2014;
Debes, No mandeau, F ase , Be na chez, & Hu chings, 2012) and
esponse o s ess (Solbe g, Skaala, Nilsen, & Glo e , 2013). These
ai s a e expec ed o make aquacul u e ish less well adap ed o he
na u al en i onmen , and mul iple s udies ha e demons a ed a med
salmon o ha e lowe su i al and ep oduc i e i ness han na i e
conspeci ics in he wild (Fleming, Jonsson, G oss, & Lambe g, 1996;
Jonsson & Jonsson, 2006; Naylo e al., 2005). Ne e heless, ma u e
aquacul u e escapees a e equen ly ound in wild spawning a eas
(E kina o e al., 2010; Fiske, Lund, & Hansen, 2006) and can b eed
wi h wild ish (Cli o d, McGinni y, & Fe guson, 1998). In og essi e
hyb idiza ion – ha is, he in oduc ion o gene ic ma e ial om he
aquacul u e escapees in o he wild popula ion ia successi e gen-
e a ions o in e b eeding – poses a numbe o di e en h ea s o
wild salmon, including he in oduc ion o ai s ha a e no locally
adap ed, a educ ion in o e all gene ic di e si y ac oss popula ions,
and he dis up ion o co- adap ed gene complexes ha ha e become
es ablished wi hin popula ions o e e olu iona y ime (Edmands,
2006). The gene ic con ibu ion o aquacul u e escapees o wild pop-
ula ions can be conside able: s udies in No way (Glo e e al., 2012,
2013), I eland (Cli o d e al., 1998) and No h Ame ica (Bou e ,
O’Reilly, Ca , Be g, & Be na chez, 2011) ha e obse ed ecen gene -
ic changes in many wild A lan ic salmon popula ions ha can be
a ibu ed o he in luence o aquacul u e escapees. Howe e , o h-
e popula ions ha e main ained hei gene ic in eg i y despi e la ge
numbe s o escaped ish being obse ed in hei na al i e s (Glo e
e al., 2012, 2013). I is la gely unknown wha demog aphic o en i-
onmen al ac o s in luence he ulne abili y o wild A lan ic salmon
popula ions o gene ic in asion by aquacul u e escapees, al hough
popula ion densi y may play a ole (Heino, S åsand, Wenne ik, &
Glo e , 2015).
Few se s o ma ke s ha can eliably disc imina e he genomes
o wild A lan ic salmon and aquacul u e s ains ha e been desc ibed
(bu see Ka lsson, Moen, Lien, Glo e , & Hinda , 2011), and his limi s
esea ch in o he dynamics o hyb idiza ion be ween escapees and
wild ish. This ela i e pauci y o ma ke s is pa ly a unc ion o he
gene ic cha ac e is ics o he aquacul u e ish. No wegian aquacul u e
lines, which a e also a med in a numbe o o he coun ies, ha e
mixed o igins. They de i e p ima ily om No wegian popula ions in
he A lan ic e olu iona y lineage o S. sala wi h small con ibu ions
om popula ions in he No h Ba en s/Whi e Sea and Bal ic e o-
lu iona y lineages (Bou e e al., 2013; Gjed em, Gjøen, & Gje de,
1991). Va ious di e en lines ha e been main ained sepa a ely since
hei ini ia ion and a e gene ically e y di e en om one ano he
(Gjøen, 1997; Ka lsson e al., 2011). These lines ha e bo h gi en ise
o u he sublines and been combined o o m new lines (Ka lsson,
Dise ud, Moen, & Hinda , 2014), and a salmon a m may use a a ying
mix o lines (Gjed em e al., 1991; Gjøen, 1997; Tho s ad e al., 2008).
Fu he , he gene ic composi ion and di e si y o hese aquacul u e
lines will ha e changed o e ime as a esul o selec ion and d i and
he addi ion o new ma e ial. Thus, he e may be no consis en gene ic
signa u e whe eby aquacul u e escapees may be disc imina ed om
wild ish.
Using an a ay enabling he simul aneous geno yping o se -
en housand (7K) single nucleo ide polymo phism (SNP) ma ke s in
A lan ic salmon, Ka lsson e al. (2011) iden i ied a sui e o 60 ma k-
e s ha disc imina ed majo No wegian aquacul u e lines om wild
No wegian salmon. Collec i ely, hese SNPs enabled iden i ica ion
o pu e- b ed wild and aquacul u e indi iduals and hei F1 hyb ids.
Howe e , hey ha e no been shown o allow disc imina ion o di -
e en classes o la e gene a ion hyb ids. This is impe a i e o
a ull unde s anding o he gene ic in e ac ions be ween wild and
aquacul u e salmon, bu is a much mo e challenging analy ical ask
(Vähä & P imme , 2006). He e, we use an A lan ic salmon SNP a ay
ha includes 220,000 (220K) mapped SNPs, and combine i wi h a
cos - e ec i e allelo yping app oach, o iden i y a se o SNPs ha
can disc imina e wild ish, a gene ically di e se se o aquacul u e
escapees, and hei i s - , second- and hi d- gene a ion hyb ids. We
ocus on he Teno Ri e o no he n Finland and No way, one o he
wo ld’s la ges and mos di e se wild A lan ic salmon s ocks (Vähä,
E kina o, Niemelä, & P imme , 2007; E kina o e al., 2010; Fig. 1).
Rep oduc i ely ma u e aquacul u e escapees ha e been caugh
h oughou his i e sys em since 1985 (E kina o e al., 2010). We
u he demons a e ha his SNP se pe o ms well o hyb id class
disc imina ion in addi ional popula ions, sugges ing ha i has gene -
al u ili y o examining wild–escapee hyb idiza ion o e a wide geo-
g aphical a ea.
2 | MATERIALS AND METHODS
2.1 | Scale samples & DNA ex ac ion
A chi ed Teno Ri e A lan ic salmon scales, collec ed as pa o a ou -
decade moni o ing p og am (Niemelä, E kina o, Julkunen, & Hassinen,
2005), we e he p ima y sou ce o gene ic ma e ial o his s udy. All
scales had been s o ed d y in en elopes p io o DNA ex ac ion. As
samples o aquacul u e salmon, we used 240 aquacul u e escapees
cap u ed in he Teno Ri e be ween 1987 and 2010 (E kina o e al.,
2010; he ea e e e ed o as ‘Teno Escapees’), and 228 escapees
cap u ed in he adjacen coas al wa e s o Finnma k, No way, in 2008
|
1019
P i cha d e al.
and 2009 (he ea e e e ed o as ‘Finnma k Escapees’). All ish had
been iden i ied as escapees on he basis o mo phological ea u es
and scale g ow h ing pa e ns consis en wi h p e ious cap i e ea -
ing (Fiske, Lund, & Hansen, 2005). We expec hese escapees o ep-
esen mul iple di e en No wegian s ains o aquacul u e salmon
ha a e u ilized in egional ish a ms. As samples o wild Teno salmon
una ec ed by in og ession om aquacul u e escapees, we used
indi iduals caugh in he i e be ween 1982 and 1987. Al hough a
salmon a m was es ablished nea he mou h o he Teno in 1984,
egional le els o salmon aquacul u e we e ela i ely low in he ea ly
1980s, and he hyb id o sp ing o ea ly escapes om his a m a e
no expec ed o e u n o he i e be o e 1989 (E kina o e al., 2010).
We chose no o use p e- 1982 samples due o he dec ease in SNP
geno yping quali y wi h sample age when using simila SNP a ays
(Johns on e al., 2013). Indi iduals we e collec ed om he Teno main-
s em (n = 120, he ea e e e ed o as ‘Old Teno Mains em’), he Teno
headwa e s (Ina ijoki, n = 114) and ou ibu a ies: Ke ojoki (n = 114),
Pulmankijä i (n = 114), Tsa sjoki (n = 114) and U sjoki (n = 120)
(Fig. 1). The salmon spawning in Ke ojoki, Pulmankijä i, Tsa sjoki and
U sjoki a e known o be empo ally s able, gene ically dis inc , popu-
la ions (Vähä, E kina o, Niemelä, & P imme , 2008; Vähä e al., 2007).
The Teno mains em has ecen ly been shown o con ain wo o e -
lapping subpopula ions wi h low gene ic di e gence be ween hem
(Aykana e al., 2015; Johns on e al., 2014); howe e , hese we e no
sepa a ed o ou s udy as he ocus was on wild–aquacul u e hyb id
de ec ion. The Ina ijoki (headwa e ) popula ion is gene ically simila
o he popula ion in he uppe Teno mains em (Vähä e al., 2007). Fo
he Old Teno Mains em sample, we selec ed equal numbe s o mul i-
seawin e (MSW, 3 yea s a sea) and one- seawin e (1SW, 1 yea a
sea) indi iduals (Johns on e al., 2014; Ba son e al., 2015; sea age
de e mined om scale g ow h ing pa e ns) om h oughou he
i e . To ob ain su icien samples, we used scales om indi iduals ha
we e caugh om he las week o July onwa ds (6–8 weeks p io o
spawning, Vähä e al., 2007) – hus, al hough he majo i y a e expec -
ed o belong o he Teno mains em popula ions, some may ha e been
en ou e o o he spawning loca ions. The e a e ewe mul i- seawin e
ish spawning in Ina ijoki and Teno ibu a ies (Vähä e al., 2007), and
we only sampled 1SW ish in hese loca ions. We selec ed an equal
numbe o males and emales wi hin each loca ion o seawin e class.
We ex ac ed DNA om 2 o 4 scales pe indi idual using a QIAmp
DNA mini ki (Qiagen), ollowing he manu ac u e ’s p o ocol and wi h
an ini ial P o einase K diges ion s ep. Olde scale samples a e known
o yield mo e deg aded DNA (Johns on e al., 2013) and he e o e il e
ips we e used wi h p e- 2000 scales o minimize con amina ion isks.
We assessed quali y and concen a ion o all DNA ex ac ions using
a Nanod op ND- 1000 spec opho ome e (The mo Fishe Scien i ic
Inc.). The Nanod op me hod is known o o e es ima e concen a ions
when DNA is pa ly deg aded (Simbolo e al., 2013); he e o e, we also
measu ed DNA concen a ion in a subse o samples (n = 14) ia lu-
o ome ic quan i a ion using a Qubi 2.0 Fluo ome e (The mo Fishe
Scien i ic Inc.).
Pilo s udies showed ha e icacy o A yme ix SNP geno yping
o olde scales declined wi h he DNA concen a ion in he ini ial
ex ac ion: he e o e, only ex ac s wi h >150 ng DNA/μl (as quan i-
ied by Nanod op) we e used u he . Old Teno Mains em and Teno
Escapees we e geno yped indi idually. Indi iduals om Teno headwa-
e and ibu a ies, and Finnma k Escapees, we e pooled by popula ion
o allelo yping, which is a cos - e ec i e app oach o es ima ion o
samplewide allele equencies using SNP chips (Johns on e al., 2013;
Oze o e al., 2013; Sham, Bade , C aig, O’Dono an, & Owen, 2002).
The ollowing numbe s o indi iduals we e included in sample pools:
Ina ijoki 110, Ke ojoki 102, Pulmankijä i 83, Tsa sjoki 107, U sjoki
108, Finnma k Escapees 227. In o ma ion on DNA concen a ion,
es ima ed using Nanod op, was used o equalize he amoun o DNA
con ibu ed by each indi idual o he pool. To accoun o pipe ing
and allelo yping a iabili y, six eplica e pools we e c ea ed ( echni-
cal eplica es) and wo aliquo s (subsamples) o each o hese six
pools we e p o ided o analysis. Fo ex ac s om scales om he
1980s, he e was an app oxima ely linea ela ionship be ween DNA
concen a ion es ima ed by Nanod op and ha es ima ed by Qubi ,
wi h Nanod op es ima es a ound i e imes highe han Qubi es i-
ma es. The e o e, o geno yping on he A yme ix a ay we p o id-
ed Old Teno Mains em and Teno Escapee samples, and Teno pools,
a a Nanod op- es ima ed concen a ion o 70 ng/μl. The Finnma k
Escapee pool, con aining DNA ex ac ed om ecen ly collec ed
scales, was p o ided close o he A yme ix ecommended concen-
a ion (10 ng/μl), a a Nanod op- es ima ed concen a ion o 15 ng/μl.
Fo analysis and quali y con ol pu poses, we used da a om h ee
addi ional se s o scale samples: 530 A lan ic salmon indi iduals col-
lec ed om he Teno mains em be ween 2001 and 2003 as desc ibed
in Johns on e al., 2013 & 2014 (he ea e e e ed o as ‘New Teno
Mains em’); 240 indi iduals collec ed be ween 2006 and 2008 om
he Nää ämö Ri e o Finland and No way, which is adjacen o he
Teno; and 120 indi iduals collec ed be ween 2005 and 2008 om he
FIGURE1 Sampling loca ions. Samples we e collec ed by
ishe man a mul iple loca ions along he Teno mains em, headwa e s
and ibu a ies
To nio
Finnma k
Teno
Nää ämö
Ke ojoki
U sjoki
Tsa sjoki
Ina ijoki
Pulmanki-
ja i
Teno Mains em
50 km
No way
Fjo d wi h salmon
a m 1984-2004
Finland
1020
|
P i cha d e al.
To nio Ri e o Finland and Sweden, which lows in o he Bal ic Sea.
New Teno Mains em and To nio ish we e indi idually geno yped on
he 220K a ay as pa o he Ba son e al. (2015) s udy. New Teno
Mains em indi iduals we e also combined in o ou sepa a e pools and
allelo yped. Pooling o hese samples was as desc ibed in Johns on
e al., 2013: h ee pipe ing eplica es we e gene a ed pe pool and
h ee subsamples allelo yped pe eplica e. Nää ämö samples we e
allelo yped: indi iduals we e combined in o ou di e en pools each
con aining 60 indi iduals; ou pipe ing eplica es we e pe o med
pe pool, and each eplica e allelo yped once.
2.2 | Geno yping and allelo yping
A cus om 220K A yme ix Axiom a ay was used o allelo ype o
geno ype samples on a GeneTi an geno yping pla o m, acco ding o
manu ac u e ’s ins uc ions (A yme ix, USA). The SNPs on his a ay
we e a subse o hose included on he 930K XHD Ssal a ay de eloped
by T. Moen and colleagues (unpublished da a), and had been chosen
o maximum in o ma i eness on he basis o hei SNPolishe pe o -
mance (SNPolishe , V1.4, A yme ix), mino allele equency (MAF) in
aquacul u e samples and physical dis ibu ion. All o hese SNPs ha e
a known loca ion on he NCBI Re Seq A lan ic salmon genome (Lien
e al., 2016, a ailable: h p://www.ncbi.nlm.nih.go /genome/anno a-
ion_euk/Salmo_sala /100/). To ensu e co ec iden i ica ion o geno-
ype clus e s, we applied he A yme ix Bes P ac ices P o ocol o
SNP calling simul aneously o a la ge da a se ha included he Old
Teno Mains em and Teno Escapee samples and all indi iduals geno-
yped o Ba son e al. (2015). Fo allelo yping, pooled samples we e
subjec ed o he s anda d geno yping me hodology, bu no malized
and summa ized Allele A and Allele B p obe in ensi ies we e e u ned
ins ead o geno ype calls.
2.3 | Quali y con ol o indi idually
geno yped samples
One hund ed and one o 112 Old Teno Mains em samples, 199 o 239
Teno Escapee samples, 526 o 530 New Teno Mains em samples and
117 o 120 To nio samples passed quali y con ols on he A yme ix
a ay. Subsequen quali y con ol s eps we e pe o med using Plink
.1.90 (Chang e al., 2015; Pu cell e al., 2007). Fi s , we emo ed
1,112 SNPs no mapped o an assembled S. sala ch omosome, 35
SNPs known o ha e o - a ge a ian s, and 1,208 SNPs de ia ing
om Ha dy–Weinbe g equilib ium a p < .0001 in he combined Old
Teno Mains em and New Teno Mains em samples (indica i e o ech-
nical geno ype calling p oblems). Subsequen ly, we excluded 18,348
SNPs wi h >10% missing da a o a MAF <10% in he combined Old
Teno Mains em and Teno Escapee da a se . Finally, we excluded 17
indi iduals wi h >10% missing da a. Following hese quali y con ol
s eps, 199,297 SNPs, geno yped in 94 Old Teno Mains em, 192 Teno
Escapee, 525 New Teno Mains em and 115 To nio indi iduals, we e
e ained o analysis.
We examined geno yping epea abili y in he Old Teno Mains em
samples by compa ing geno ype calls o ou indi iduals ha had
been geno yped wice, using he me ge unc ion in Plink; his was
compa ed o he epea abili y o i e epea edly geno yped New Teno
Mains em samples.
As an ini ial explo a ion o he gene ic a ia ion in he combined
Old Teno Mains em, New Teno Mains em and Teno Escapee da a
se , we emo ed SNPs wi h >2% missing da a, pe o med a linkage
disequilib ium p uning s ep in Plink (window size = 100, shi = 10,
VIF = 2) and hen used he genome unc ion o calcula e pai wise
iden i y- by- s a e be ween all indi iduals, based on he emaining
48,375 SNPs. The p esence o geno ypic clus e s was in es iga ed
by pe o ming a wo- dimensional mul idimensional scaling analysis
(MDS) on he genomic iden i y- by- s a e (IBS) ma ix in Plink and isu-
alizing he ou pu using ggplo 2 in R 3.1.2 (Wickham, 2009; R Co e
Team 2015, Fig. 2).
2.4 | Es ima ion o allele equency om pooled
indi iduals
We es ima ed allele equency a each SNP in each pool by calcula ing
he ela i e in ensi y o he B- allele p obe signal (=B- allele in ensi y/
(A- allele in ensi y + B- allele in ensi y)), aking he median o e all ep-
lica ed pools, and applying a polynomial- based p obe- speci ic (PPC)
co ec ion o accoun o di e en ial hyb idiza ion e iciency o he
wo p obes (Anan ha aman & Chew, 2009; B ohede, Dunne, McKay,
& Hannan, 2005). To ob ain he PPC co ec ion coe icien s o each
SNP, a second- o de polynomial desc ibing he ela ionship be ween
ela i e B p obe in ensi y and geno ype call (AA, AB o BB) was i o e
he 525 indi idually geno yped New Teno Mains em samples, plus 86
samples om he Teno mains em geno yped o a di e en s udy,
FIGURE2 Mul idimensional scaling analysis plo isualizing
genomewide iden i y- by- s a e amongs Old Teno Mains em, New
Teno Mains em and Teno Escapee samples. Each poin ep esen s an
indi idually geno yped ish. Teno Escapee samples a e colou - coded
by collec ion pe iod
|
1021
P i cha d e al.
using a cus om sc ip in R. The polynomial wi h hese coe icien s was
hen used o co ec he es ima ed allele equency o ha SNP in all
allelo yped pools. SNPs wi h a PPC- co ec ed allele equency >1 o
<0 we e conside ed monomo phic, and allele equencies we e adjus -
ed acco dingly. The accu acy o he PPC- co ec ed allele equency
es ima es o he ou New Teno Mains em pools was in es iga ed
by eg essing hem agains he ue allele equencies ob ained om
indi idual geno yping. Linea eg ession was pe o med using he lm
unc ion in R, speci ying he model as (PPC- co ec ed equency) ~ 0 +
(T ue equency) and using de aul pa ame e s.
2.5 | Selec ion o SNPs o hyb id class
disc imina ion
To selec SNPs po en ially in o ma i e o hyb id class disc imina ion,
we ocused on egions o he genome ha a e unusually di e gen
be ween aquacul u e escapees and wild Teno ish ha we e collec ed
p io o aquacul u e in luence. To iden i y hese egions, we used wo
di e en genome scan app oaches which ake popula ion allele coun s
as inpu and use F s a is ics o iden i y ou lying loci: Fdis 2 (Beaumon
& Nichols, 1996, a ailable: h p://www.ma hs.b is.ac.uk/~mamab/
so wa e/) and Bayescan (Foll & Gaggio i, 2008). The sou ce code
o Fdis 2 was sligh ly modi ied o allow a la ge numbe o ma k-
e s and compiled unde Linux. Genomes scans we e applied o he
en i e da a se o 199,297 SNPs. Al hough ou wild and aquacul u e
popula ions a e expec ed o iola e some o he model assump ions
unde lying hese app oaches (e.g. island mig a ion model; independ-
en di e gence om a common ances o ), as ou p ima y aim was o
iden i y loci ha could be used o disc imina e be ween hem a he
han make in e ences abou di e en ial selec ion we conside ed hei
use jus i ied.
We con e ed allele equency es ima es om pooled samples
in o allele coun s assuming 200 allele copies pe locus pe pool. Fo
Fdis 2 and Bayescan analyses, he ‘All Escapee’ sample was he com-
bined allele coun s o he indi idually geno yped Teno Escapees and
he allelo yped Finnma k Escapee pool. Ini ially, his sample was com-
pa ed o he combined ‘All Wild’ sample (allele coun s om indi idu-
ally geno yped Old Teno Mains em ish plus allele coun s om pooled
Ina ijoki, Ke ojoki, Tsa sjoki, U sjoki and Pulmankijä i). Subsequen ly,
we compa ed he All Escapee sample each o he six Teno subpopula-
ions sepa a ely. We an Bayescan using de aul pa ame e s. To apply
Fdis 2, we used he package unc ion da acal o calcula e obse ed
he e ozygosi y and Fs o each locus. We hen simula ed he expec -
ed null dis ibu ion o he e ozygosi y and Fs , assuming wo demes/
popula ions o 100 indi iduals, o e 10,000,000 loci, using he unc-
ion dis 2. In o de o iden i y he model inpu alue o Expec ed Fs
ha would gene a e a mean simula ed Fs app oxima ing he mean
obse ed Fs be ween each escapee–wild compa ison, we made pilo
uns o 100,000 simula ions i e a i ely changing Expec ed Fs un il
we ob ained he desi ed alues. Inpu alues o Expec ed Fs we e
as ollows: All Wild/All Escapee: 0.143; Ina joki/All Escapee: 0.172;
Ke ojoki/All Escapee: 0.200; Pulmankijä i/All Escapee: 0.207; Old
Teno Mains em/All Escapee: 0.115; Tsa sjoki/All Escapee: 0.361;
U sjoki/All Escapee: 0.228. To ob ain he dis ibu ion o Fs alues
simula ed by dis 2 (wi hin each He bin o 0.04), calcula e empi ical
p obabili ies o ou obse ed Fs alues on he basis o hese dis i-
bu ions, and con e hese in o q alues, we used he R unc ions in
ge P alues.R, w i en by Lo e hos and Whi lock (2014) and a ailable in
he D yad eposi o y (doi: 10.5061/d yad. 8d05).
All ou lying loci iden i ied by Bayescan o Fdis 2 in all compa isons
we e examined o in ach omosome linkage disequilib ium using he
2 unc ion in Plink, applied o he combined Old Teno Mains em and
New Teno Mains em da a se s. F om hese esul s, we iden i ied clus-
e s o physically adjacen linked SNPs: om obse a ion o he da a,
a SNP was a bi a ily conside ed o be wi hin a linked clus e when i s
pai wise 2 wi h any o he SNP wi hin he clus e was >0.2.
We ini ially selec ed a se o 200 SNPs o use in hyb id disc imi-
na ion, as pilo s udies demons a ed no subs an ial inc ease in assign-
men e icacy using addi ional ma ke s (da a no shown), and because
oughly his numbe o SNPs can be con enien ly analysed on wo 96-
well pla es. We i s chose SNPs loca ed in linked clus e s ha we e
iden i ied as ou lie s by bo h Bayescan and Fdis 2 in he All Wild/All
Escapee compa ison. We chose one SNP pe clus e , selec ing he
one ha had he s onges ou lying pa e n o e all he popula ion
compa isons. We supplemen ed hese wi h SNPs in di e en linkage
clus e s ha we e iden i ied as ou lie s by ei he Bayescan o Fdis 2
in a leas ou compa isons be ween di e en Teno subpopula ions
and All Escapees.
2.6 | Popula ion gene ic cha ac e is ics
Fo each popula ion, mean expec ed he e ozygosi y was es ima ed
om allele equencies o he 199,297 SNPs using R. Pai wise unbi-
ased Fs alues be ween all popula ions (weigh ed by he e ozygosi y,
Cocke ham & Wei , 1993; Wei & Cocke ham, 1984) we e calcula ed
om es ima ed allele coun s using da acal. To examine whe he use
o pooled samples biased ou es ima ion o Fs we calcula ed pai -
wise Fs be ween he All Escapee sample and each o he ou New
Teno Mains em pools, i s using allele equencies ob ained om
indi idual geno yping and hen using allele equencies es ima ed
using allelo yping. We also epea ed all pai wise Fs compa isons
using he subse s o 200, 160, 120, 80 and 40 loci selec ed o hyb id
disc imina ion.
2.7 | Simula ion o hyb id popula ions
To es whe he he 200 SNPs could in combina ion be used o dis-
c imina e hyb id classes, we used he Py hon package simuPOP (Peng
& Kimmel, 2005) o simula e popula ions wi h allele equencies a
hese SNPs app oxima ing hose obse ed in ou geno yped and/
o allelo yped samples (All Escapee, Old Teno Mains em, Ina ijoki,
Ke ojoki, Pulmankijä i, Tsa sjoki & U sjoki, plus he combined All
Wild sample; see u he de ails below), and an hem h ough h ee
gene a ions o hyb idiza ion. SimuPOP enables he use o simula e
physical linkage be ween ma ke s on he same ch omosome by p o-
iding linkage dis ances be ween ma ke s. Based on o al physical
1022
|
P i cha d e al.
leng h o he 29 assembled S. sala ch omosomes (≈2200 million bp)
and mean o al linkage map dis ance es ima ed by Gonen e al., 2014
(2190 cM), we used he app oxima ion 1 million bp = 1 cM.
Se e al sou ces o bias could lead us o o e es ima e he e icacy
o ou 200 SNPs o disc imina e he genomes o wild ish and aqua-
cul u e escapees. Fi s , we obse ed ha using allele equencies
ob ained ia allelo yping caused Fs be ween he All Escapee and New
Teno Mains em pools o be o e es ima ed by up o 7% (Table S1, see
below). Second, he es ima e o pai wise Fs om popula ion samples
is in gene al expec ed o be highe han he ue popula ion- le el alue
due o sampling a iance (Ande son, Waples, & Kalinowski, 2008). In
o de o minimize hese biases, we delibe a ely adjus ed he allele e-
quencies in ou simula ed wild and escapee popula ions o be close o
one ano he . Ini ial B- allele equencies o he simula ed popula ions
we e hose es ima ed om geno yping/allelo yping he eal popula-
ion samples. Fi s , based on a compa ison o ac ual allele equencies
in he New Teno Mains em pools o hose es ima ed by allelo yping
(Fig. S1), we adjus ed allele equencies o 0 o 1 o 0.025 o 0.975,
espec i ely. Fo each locus in each popula ion, we hen gene a ed
a new B- allele equency by making andom d aws om he be a-
dis ibu ion wi h pa ame e s ( *200, (1- )*200), whe e is he B- allele
equency ollowing he p e ious adjus men s ep. D aws we e pe -
o med using he R unc ion be a(). Fo simula ed Ina ijoki, Ke ojoki,
Tsa sjoki, Pulmankija i, U sjoki and All Wild popula ions, we chose
he i s alue ha adjus ed he B- allele equency close o he All
Escapee alue; o he simula ed All Escapee popula ion we chose he
i s alue close o he All Wild alue. As Old Teno Mains em allele e-
quencies we e es ima ed om indi idual geno yping only, we did no
adjus hese alues o he simula ed Old Teno Mains em popula ion.
To c ea e Gene a ion 0 in simuPOP, we used he adjus ed allele
equencies o ini ialize an escapee and a wild popula ion each
con aining 500 indi iduals o each sex. We es ima ed pai wise Fs
be ween he simula ed Gene a ion 0 popula ions by con e ing
he adjus ed allele equencies in o allele coun s assuming 1,000
alleles pe popula ion and using hese as inpu in o he Fdis 2 unc-
ion da acal. We hen simula ed h ee gene a ions o hyb idiza ion
in simuPOP as ollows: 100 indi iduals mig a ed om he escapee
popula ion in o he wild popula ion each gene a ion; subsequen ly,
500 o sp ing o each sex we e gene a ed by andom ma ing wi hin
each popula ion, wi h pa en s sampled wi h eplacemen and each
pai p oducing a single o sp ing. Indi idual IDs, pa en al IDs and
geno ypes o all indi iduals in all Gene a ions (0 o 3) we e ou pu
o a single ile in (PED) o ma . We used a cus om R sc ip o calcula e
hyb id class o each indi idual by acing i s lineage om Gene a ion
0. To acili a e compa ison o hyb id class assignmen s amongs
simula ions wi h di e en ances al wild popula ions, assignmen s
we e pe o med on s anda dized subse s o 400 indi iduals o each
simula ion. These subse s we e gene a ed by ha es ing he same
numbe o indi iduals o each hyb id class o each simula ion. The
numbe o indi iduals o be ha es ed was de e mined om he
obse ed equency o ha hyb id class in he ele an gene a ion
o e all simula ions. Fo example, in Gene a ion 2 o e all equen-
cies o Pu e Wild and Wild Backc oss app oxima ed 0.505 and 0.215,
espec i ely; he e o e, he es subse o 400 Gene a ion 2 indi id-
uals o each wild ances al popula ion included 202 Pu e Wild and
86 Wild Backc oss indi iduals. Th ee possible Gene a ion 3 hyb id
classes (Escapee Backc oss X Escapee Backc oss, Escapee Backc oss
X F2, and F2 X F2) we e excluded due o he small p opo ional ep-
esen a ion o hese classes. We also gene a ed an independen se
o Gene a ion 0 indi iduals o each simula ion (50 wild, 50 escapee)
using he same ini ial allele equencies. These indi iduals we e c e-
a ed as wild and escapee e e ence samples and did no con ibu e o
he hyb idizing popula ions.
2.8 | Hyb id class disc imina ion in simula ed
popula ions
We i s used a Bayesian app oach implemen ed in he p og am
NewHyb ids (Ande son & Thompson, 2002) o assign simula ed indi-
iduals o use - de ined pu e o hyb id classes (six possible classes o
Gene a ion 2; 21 possible classes o Gene a ion 3, o which 18 we e
p esen in he da a se and 14 could po en ially be disc imina ed by
NewHyb ids). We p o ided he 100 e e ence indi iduals desc ibed
abo e as e e ence samples: hese indi iduals we e no conside ed
o be pa o he hyb id es popula ion. We addi ionally an he
Gene a ion 2 analysis wi hou a e e ence sample. We used he
command- line e sion o NewHyb ids, compiled unde Linux, and an
he analysis wice o each da a se using di e en andom seeds, wi h
a bu n- in o 10,000 ollowed by 50,000 sweeps and de aul alues
o all o he pa ame e s. Numbe o sweeps was de e mined a p io i
by pe o ming pilo uns in he g aphical e sion o NewHyb ids, wi h
one o he Gene a ion 3 da a se s, using he same de aul pa ame e s,
and isually ollowing he p og ess o he analysis. An indi idual was
conside ed o be co ec ly classi ied when i was assigned o i s own
class wi h p obabili y >0.5.
To examine he e icacy o hyb id class disc imina ion using ewe
SNPs, we epea ed he NewHyb id analyses wi h subse s o he ini-
ial 200 loci. Loci we e anked by allele equency di e ence be ween
he All Wild and All Escapee samples, and hose wi h he smalles
di e ence emo ed i s . Resul s om analyses wi h n = 160, 120,
80 and 40 SNPs we e compa ed o hose om he ull se o 200.
We quan i ied he abili y o di e en numbe s o SNPS o co ec -
ly assign indi iduals o di e en hyb id classes, ollowing Vähä and
P imme (2006), by de ining he ollowing measu es. E iciency is he
p opo ion o indi iduals in a ce ain hyb id class ha we e ac ually
assigned o ha class by NewHyb ids, o example (To al numbe o
simula ed F1 assigned o he F1 class)/(To al numbe o simula ed
F1). Accu acy is he p opo ion o indi iduals assigned o a class by
NewHyb ids ha ac ually belong o ha class, o example (To al
numbe o simula ed F1 assigned o he F1 class)/(To al numbe o
all indi iduals assigned o he F1 class). O e all pe o mance is he
mean o E iciency mul iplied by Accu acy o each hyb id class o e
all popula ions.
As a suppo ing analysis, we also es ima ed he p opo ion o
wild and escapee ances y o each simula ed indi idual using he
command- line e sion o he p og am S uc u e (P i cha d, S ephens,
|
1023
P i cha d e al.
& Donnelly, 2000). Again, we p o ided e e ence samples o 50 wild
ish and 50 escapees and de ined hem using ‘POPDATA’, ‘POPFLAG’
and ‘USEPOPINFO’ wi h MIGRPRIOR = 0.0001. We used k = 2, a
bu n-in o 20,000 ollowed by 200,000 MCMC s eps, eco ded 95%
con idence in e als o es ima ed ances y, and e ained de aul alues
o all o he pa ame e s.
F om S uc u e esul s, i e classes (pu e wild, pu e escapee, F1
o F2 hyb ids, and backc osses in each di ec ion) in Gene a ion 2, and
nine classes in Gene a ion 3, could be disc imina ed on he basis o
expec ed admix u e p opo ions. We assigned indi iduals o class-
es by examining hei p opo ion o ances y om he wild clus e .
Indi iduals we e conside ed o be assigned o a class when he 95%
con idence in e als o hei es ima ed wild ances y did no o e lap
he expec ed mean ances y o adjacen classes (e.g. an F1/F2 indi-
idual has an uppe CL o wild ances y <0.75 and a lowe CL o wild
ances y >0.25; pu e wild o pu e escapee indi iduals had he 95%
CL o hei wild ances y o e lapping 1.0 and 0.0, espec i ely). We
examined he assignmen pe o mance o di e en se s o SNPs as
desc ibed abo e.
2.9 | Hyb id disc imina ion in addi ional wild
popula ions
As he disc imina o y SNPs we e selec ed based on allele equency
di e ences be ween he same popula ions ha we e subsequen ly
used o es hem, ou esul s may o e es ima e he e icacy o SNPs
o disc imina e wild–escapee hyb ids in o he popula ions (‘high g ad-
ing bias’, Ande son, 2010; Waples, 2010). We he e o e es ed he abil-
i y o hese 200 SNPs o disc imina e di e en classes o simula ed
hyb ids using wo addi ional wild popula ions. The Nää ämö A lan ic
salmon popula ion belongs o same e olu iona y lineage (No h
Ba en s/Whi e Sea) as he Teno Ri e , while he To nio popula ion is
wi hin a gene ically di e gen lineage (Bal ic, Bou e e al., 2013, Fig.
S9). Fo Nää ämö, B- allele equency a he 200 SNPs was es ima ed by
aking he median o e all 16 allelo yped pools and applying he PPC
co ec ion; o To nio, i was es ima ed om all 115 indi idually geno-
yped ish. Allele equencies o Nää ämö and To nio we e no u he
adjus ed. Al hough we did no ha e an independen escapee sample,
we made ou analysis mo e conse a i e by es ima ing Gene a ion 0
allele equencies o ou simula ed escapee popula ion om a an-
domly chosen subse o 90 Teno Escapees a he han he en i e All
Escapee sample. The emaining 99 Teno Escapee we e used o es i-
ma e allele equencies o he escapee e e ence sample. As be o e,
h ee gene a ions o hyb idiza ion we e simula ed using simuPOP, and
NewHyb ids used o assign simula ed Gene a ion 2 and Gene a ion 3
indi iduals o di e en hyb id classes as desc ibed abo e.
3 | RESULTS
3.1 | Ini ial explo a ion
O e all, Old Teno Mains em samples, collec ed in he 1980s, ailed
A yme ix quali y con ols mo e equen ly and exhibi ed mo e
missing geno ypes han New Teno Mains em samples, collec ed in
he 2000s. Mean geno ype epea abili y o e he ou eplica e Old
Teno Mains em samples was 98.0%, as compa ed o 99.3% o e he
i e eplica e New Teno Mains em samples.
P elimina y explo a ion o genomewide pa e ns o iden i y- by-
s a e e ealed clea gene ic di e en ia ion be ween he indi idually
geno yped Teno Mains em and Teno Escapee samples (Fig. 2). Fu he ,
he escapee sample comp ised wo, la gely disc e e, gene ic clus-
e s. These clus e s we e no ela ed o he ime pe iod in which he
escapees we e collec ed, and o e all, we obse ed no clea empo-
al pa e ns in he genomic composi ion o he escapee o Teno sam-
ples. Th ee indi iduals classi ied as ‘aquacul u e escapees’ clus e ed
wi h he wild Teno indi iduals and we e conside ed misiden i ied and
emo ed (Fig. S9). No indi iduals classi ied as wild Teno ish clus e ed
wi h he aquacul u e escapees; howe e , se e al New Teno Mains em
indi iduals loca ed be ween he Teno and Escapee clus e s we e con-
side ed po en ial wild–escapee hyb ids (Fig. 2).
3.2 | Accu acy o allele equency es ima ion by
allelo yping
Fo he 199,297 SNPs ha emained a e il e ing, popula ion allele
equencies es ima ed om he ou New Teno Mains em pools wi h
he PPC co ec ion we e closely linea ly ela ed o he ue equen-
cies es ima ed by indi idually geno yping (Fig. S1; each pool: 2 = .996,
esidual SE = 0.037–0.039).
3.3 | Popula ion gene ic pa ame e s
Mean He wi hin he sampled popula ions, calcula ed o e all 199,297
SNPs, was as ollows: Teno Escapees, 0.415; Finnma k Escapees,
0.348; Old Teno Mains em, 0.381; New Teno Mains em 0.376:
Ina ijoki, 0.340; Ke ojoki, 0.377; Pulmankija i, 0.349; Tsa sjoki,
0.276; U sjoki, 0.313; Naa amo, 0.300; To nio, 0.280. Pai wise Fs al-
ues amongs he samples (All Escapee plus eigh wild popula ions) a e
shown in Table 1. Mean pai wise Fs be ween aquacul u e escapees
and he six Teno popula ions was 0.097, compa ed o a mean pai wise
Fs o 0.083 amongs he Teno popula ions. Rela i ely high pai wise
Fs be ween Tsa sjoki and mos o he loca ions e lec ed he educed
gene ic di e si y in his popula ion. Wi h Tsa sjoki emo ed, mean
aquacul u e–wild pai wise Fs and mean amongs - wild pai wise Fs
we e 0.084 and 0.067, espec i ely. The high pai wise Fs be ween
To nio and all he o he popula ions e lec ed he phylogeog aphic
dis inc ness o his popula ion.
As expec ed, pai wise Fs be ween he All Escapee and wild popu-
la ions was much highe when calcula ed only om he disc imina o y
subse o SNPs (Table S1, mean Fs ac oss Teno/All Escapee compa -
isons = 0.464). Pai wise Fs alues es ima ed be ween he New Teno
Mains em pools and he All Escapee sample we e sligh ly highe when
allele equencies had been es ima ed by allelo yping han when hey
had been calcula ed by indi idual geno yping (Table S1, mean Fs om
199,297 loci = 0.067 s. 0.058; mean Fs om 200 loci = 0.324 s.
0.306).
1024
|
P i cha d e al.
3.4 | Ou lie analysis
In he compa ison be ween all wild ish and all aquacul u e escap-
ees (All Wild/All Escapee), we ound 227 ou lying loci (q < .05) wi h
Fdis 2 and 1,112 (q < .05) wi h Bayescan, o which 183 o e lapped
be ween he wo me hodologies (Fig. S2; Table S2). Ou lie loci we e
dis ibu ed ac oss all 29 ch omosomes (Fig. S2). Mo e han 95% o
hese 1,156 loci we e also iden i ied as ou lie s in one o mo e o
he compa isons be ween escapees and ish om di e en Teno
loca ions (Ke ojoki, Ina ijoki, Pulmankijä i, Tsa sjoki, U sjoki o Old
Teno Mains em; 78% o loci we e ou lie s in a leas wo compa i-
sons, Table S2). Ac oss he genome, 67 linked clus e s o SNPs plus
60 single SNPs we e iden i ied by bo h Bayescan (q < .05) and Fdis 2
(q < .05) as ou lie s in he All Wild/Escapee compa ison (Table S2).
The se o 200 SNPs selec ed o use in hyb id disc imina ion we e
dis ibu ed o e 28 o he 29 A lan ic salmon ch omosomes (Fig. S2,
Tables S2, S3). None o hese loci o e lapped wi h hose desc ibed by
Ka lsson e al. (2011). Es ima ed B- allele equencies o hese SNPs
and adjus ed equencies used o simula ions a e p o ided in Table
S3. Pai wise Fs alues be ween simula ed escapee and wild popula-
ions a Gene a ion 0 we e lowe han he unbiased pai wise Fs al-
ues calcula ed om es ima ed allele equencies o he eal samples
(Table S1).
3.5 | Assignmen o simula ed indi iduals o di e en
hyb id classes
Replica e uns o NewHyb ids o he same da a se ga e almos iden-
ical ou comes, and in all cases, we p esen esul s om he i s un.
Fo he second gene a ion o hyb idiza ion be ween wild ish and
aquacul u e ish (Gene a ion 2), he ull se o 200 SNPs exhibi ed a
high o e all pe o mance, assigning all bu 24 indi iduals (1%) o he
co ec hyb id class o e all six simula ed Teno subpopula ions, i e-
spec i e o whe he o no a e e ence sample was p o ided (Fig. 3,
Fig. 5, Tables S4, S5). This pe o mance ba ely changed when num-
be o SNPs was educed o 160 (Fig. 5, Fig. S3, Tables S4, S5: 1.4%
misassignmen ). Fu he educ ions in SNP numbe led o a decline in
o e all pe o mance, pa icula ly when assigning F2 hyb ids and escap-
ee backc osses (Fig. 5, Fig. S3; Tables S4, S5). Howe e , e en using as
ew as 40 SNPs, only 12 hyb id indi iduals (all wild backc oss, 1% o all
hyb ids) we e e oneously iden i ied as pu e wild ish.
Assignmen o indi iduals o he many possible hyb id classes c e-
a ed by h ee gene a ions o wild–escapee hyb idiza ion (Gene a ion
3) is a much mo e di icul analy ical p oblem, and we ocus ou dis-
cussion on he i e mos equen classes gene a ed in ou scena io o
10% escapees pe gene a ion: Pu e Wild, Pu e Escapee, F1 Hyb ids,
Wild Backc oss and Wild Backc oss X Wild. Using he ull se o 200
SNPs, 99.1% o Pu e Wild indi iduals, 96.4% o Pu e Escapees, 94.8%
o F1 Hyb ids, 83.8% o Wild Backc oss X Wild and 73.2% o Wild
Backc oss we e assigned co ec ly (Fig. 4, Fig. 5, Table S6). Howe e ,
a e hyb id classes we e co ec ly assigned wi h much lowe suc-
cess (Fig. 4, Table S6). Misassigned indi iduals we e almos in a iably
assigned o hyb id classes wi h simila p opo ions o wild/escapee
ances y. As expec ed, pe o mance declined wi h dec easing num-
be s o SNPs (Fig. 5, Fig S4, Table S6). Once again, misassignmen o
hyb id indi iduals o he Pu e Wild class was a e and almos en i ely
limi ed o Wild Backc oss X Wild indi iduals; e en wi h 80 SNPs, only
4% o hyb ids we e inco ec ly assigned as Pu e Wild indi iduals.
Replica e uns o S uc u e o he same da a se ga e cong uen
esul s, and we epo esul s om he i s un. In gene al, S uc u e
was a less e ec i e analy ical app oach han NewHyb ids o assign-
ing simula ed indi iduals o hyb id classes, pa icula ly wi h ewe
SNPs and a e h ee gene a ions o hyb idiza ion (Fig. 5, Figs S5, S6,
Tables S7, S8). Ne e heless, a Gene a ion 2, and wi h 120 o mo e
SNPs, S uc u e and NewHyb ids assigned indi iduals o he co ec
class wi h simila high le els o accu acy (Fig. 5, Fig. S5, Table S7).
3.6 | Hyb id disc imina ion in addi ional wild
popula ions
Simula ed hyb ids be ween aquacul u e escapees and ish om
Nää ämö o To nio we e assigned o he co ec hyb id class wi h
TABLE1 Pai wise Fs be ween samples based on all 199,297 SNPs (abo e diagonal) and 200 disc imina o y SNPS (below diagonal)
All escapee Ina ijoki Ke ojoki
Pulmanki
jä i
Teno Old
Mains em
Teno New
Mains em Tsa sjoki U sjoki Nää ämö To nio
All escapee 0.078 0.091 0.094 0.054 0.056 0.163 0.103 0.095 0.170
Ina ijoki 0.470 0.074 0.072 0.018 0.027 0.141 0.066 0.057 0.203
Ke ojoki 0.466 0.087 0.076 0.057 0.078 0.114 0.069 0.131 0.217
Pulmankijä i 0.471 0.096 0.084 0.062 0.074 0.170 0.105 0.111 0.220
Teno Old
Mains em
0.333 0.075 0.103 0.110 0.006 0.125 0.057 0.049 0.172
Teno New
Mains em
0.307 0.110 0.130 0.134 0.010 0.130 0.061 0.039 0.171
Tsa sjoki 0.529 0.141 0.090 0.140 0.161 0.191 0.034 0.162 0.288
U sjoki 0.518 0.101 0.077 0.109 0.137 0.168 0.037 0.080 0.227
Nää ämö 0.401 0.135 0.184 0.178 0.067 0.063 0.252 0.215 0.213
To nio 0.272 0.417 0.410 0.424 0.259 0.232 0.487 0.471 0.319
|
1025
P i cha d e al.
simila accu acy o ha obse ed in he simula ions in ol ing Teno
popula ions (Fig. 6, Fig. S8, Tables S4, S5 & S6), especially using la ge
numbe s o SNPs. Again, misassignmen o simula ed hyb ids as pu e
wild ish was a e. Wi h 200 SNPs, a Gene a ion 2, he e we e no
such misassignmen s o ei he To nio o Nää ämö. A Gene a ion
3, no hyb ids we e misiden i ied as wild ish o Nää ämö and <1.9%
o hyb ids we e misiden i ied o To nio, all o which we e Wild
Backc oss X Wild.
4 | DISCUSSION
He e, we ha e shown ha a sui e o 200 SNPs can collec i ely dis-
c imina e ad anced- gene a ion classes o hyb id be ween wild ish
om a gene ically di e se i e and he aquacul u e escapees ha
may be ep oducing in ha i e . Ou assessmen o he e icacy o
hese ma ke s is expec ed o be conse a i e because we delibe -
a ely simula ed hyb idizing popula ions wi h a lowe le el o gene ic
di e gence be ween hem han was es ima ed om ou eal popula-
ion sample. Fo example, ou simula ed wild popula ions con ained
no monomo phic loci, while many o hese loci may uly be mono-
mo phic in se e al o he Teno subpopula ions. Fo wo gene a ions
o hyb idiza ion, hese ma ke s assign simula ed indi iduals o hei
co ec hyb id class wi h a e y high le el o accu acy. E en a e
h ee gene a ions o hyb idiza ion, many hyb id classes can be eli-
ably iden i ied. Impo an ly, indi iduals wi h hyb id ances y a e
a ely iden i ied as pu e wild ish in ei he o he hyb id gene a ions
examined. This is a much mo e accu a e le el o hyb id class iden i-
ica ion han has p e iously been shown using gene ic ma ke s ha
disc imina e aquacul u e and wild ish (Ka lsson e al., 2011). Fu he ,
we obse e simila ly good disc imina ion o hyb id classes when
simula ing hyb ids om pa en al popula ions ha we e no o iginally
used o iden i y he SNPs, including a popula ion in a highly di e gen
e olu iona y lineage, which sugges s his ma ke se may be use ul
ac oss a ange o popula ions. The a he low numbe o 200 SNPs
can nowadays be assayed ela i ely cheaply h ough geno yping ia
sequencing (Campbell, Ha mon, & Na um, 2014) o simila app oach-
es. Mo eo e , we ha e shown ha a smalle subse o hese SNPs
enables hyb id class disc imina ion o a le el o accu acy ha may be
su icien in many scena ios.
FIGURE3 Resul s o NewHyb ids analysis o 400 indi iduals p oduced by wo gene a ions o simula ed hyb idiza ion be ween aquacul u e
escapees (10% o he popula ion) and wild ish om di e en Teno subpopula ions. Each indi idual is ep esen ed by a e ical ba . Indi iduals
a e a anged along he x- axis by simula ed hyb id class, wi h di e en hyb id classes bounded by black lines. Y- axis indica es he p obabili y,
e u ned by New Hyb ids, ha an indi idual belongs o one o he six possible hyb id classes (‘Assignmen p obabili y’). The di e en possible
hyb id classes a e indica ed by di e en colou s. Fo ‘All Wild’, he wild popula ion was simula ed using he a e age allele equencies o e all
se en subpopula ions
Simula ed
class: