Finding markers that make a difference : DNA pooling and SNP-arrays identify population informative markers for genetic stock identification
Full text
Finding Ma ke s Tha Make a Di e ence: DNA Pooling
and SNP-A ays Iden i y Popula ion In o ma i e Ma ke s
o Gene ic S ock Iden i ica ion
Mikhail Oze o
1
, An i Vasema
¨gi
2,8
, Vida Wenne ik
3
, Rogelio Diaz-Fe nandez
1
, Ma hew Ken
4
,
John Gilbey
5
, Se gey P uso
6
, Ee o Niemela
¨
7
, Juha-Pekka Va
¨ha
¨
1
*
1Ke o Suba c ic Resea ch Ins i u e, Uni e si y o Tu ku, Tu ku, Finland, 2Depa men o Biology, Di ision o Gene ics and Physiology, Uni e si y o Tu ku, Tu ku, Finland,
3Resea ch g oup Popula ion Gene ics and Ecology, Ins i u e o Ma ine Resea ch, Be gen, No way, 4Depa men o Animal and Aquacul u al Sciences, Cen e o
In eg a i e Gene ics (CIGENE), No wegian Uni e si y o Li e Sciences, A
˚s, No way, 5F eshwa e Labo a o y, Ma ine Sco land, Faskally, Pi loch y, Uni ed Kingdom,
6F eshwa e Resou ces Labo a o y, Knipo i ch Pola Resea ch Ins i u e o Ma ine Fishe ies and Oceanog aphy, Mu mansk, Russia, 7Finnish Game and Fishe ies Resea ch
Ins i u e, Oulu, Finland, 8Depa men o Aquacul u e, Ins i u e o Ve e ina y Medicine and Animal Science, Es onian Uni e si y o Li e Sciences, Ta u, Es onia
Abs ac
Gene ic s ock iden i ica ion (GSI) using molecula ma ke s is an impo an ool o managemen o mig a o y species. He e,
we es ed a cos -e ec i e al e na i e o indi idual geno yping, known as allelo yping, o iden i ica ion o highly in o ma i e
SNPs o accu a e gene ic s ock iden i ica ion. We es ima ed allele equencies o 2880 SNPs om DNA pools o 23 A lan ic
salmon popula ions using Illumina SNP-chip. We e alua ed he pe o mance o ou common s a egies (global F
ST
, pai wise
F
ST
,Del a and ou lie app oach) o selec ion o he mos in o ma i e se o SNPs and es ed hei e ec i eness o GSI
compa ed o andom se s o SNP and mic osa elli e ma ke s. Fo he majo i y o cases, SNPs selec ed using he ou lie
app oach pe o med bes ollowed by pai wise F
ST
and Del a me hods. O e all, he selec ion p ocedu e educed he numbe
o SNPs equi ed o accu a e GSI by up o 53% compa ed wi h andomly chosen SNPs. Howe e , GSI accu acy was mo e
a ec ed by popula ions in he asce ainmen g oup a he han he anking me hod i sel . We demons a ed o he i s
ime he compa ibili y o di e en la ge-scale SNP da ase s by compiling he la ges popula ion gene ic da ase o A lan ic
salmon o da e. Finally, we showed an excellen pe o mance o ou op SNPs on an independen se o popula ions
co e ing he main Eu opean dis ibu ion ange o A lan ic salmon. Taken oge he , we demons a e how combina ion o
DNA pooling and SNP a ays can be applied o conse a ion and managemen o salmonids as well as o he species.
Ci a ion: Oze o M, Vasema
¨gi A, Wenne ik V, Diaz-Fe nandez R, Ken M, e al. (2013) Finding Ma ke s Tha Make a Di e ence: DNA Pooling and SNP-A ays
Iden i y Popula ion In o ma i e Ma ke s o Gene ic S ock Iden i ica ion. PLoS ONE 8(12): e82434. doi:10.1371/jou nal.pone.0082434
Edi o : So ia Consueg a, Abe ys wy h Uni e si y, Uni ed Kingdom
Recei ed May 2, 2013; Accep ed Oc obe 24, 2013; Published Decembe 16, 2013
Copy igh : ß2013 Oze o e al. This is an open-access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion License, which pe mi s
un es ic ed use, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal au ho and sou ce a e c edi ed.
Funding: This s udy was unded by he Eu opean Union, Kola c ic ENPI CBC p ojec KO197 ‘‘T ila e al coope a ion in ou common esou ce; he A lan ic salmon
in he Ba en s Region’’ (M.O., R.D.F., E.N. and S.P.), Academy o Finland (J.-P.V.), No wegian Di ec o a e o Na u e Managemen (V.W.), No wegian Resea ch Council
(V.W.), Es onian Minis y o Educa ion and Resea ch (Ins i u ional Resea ch Funding p ojec IUT8-2 o A.V.) and Es onian Science Founda ion (g an numbe s 6802,
8215 o A.V.). The unde s had no ole in s udy design, da a collec ion and analysis, decision o publish, o p epa a ion o he manusc ip .
Compe ing In e es s: The au ho s ha e decla ed ha no compe ing in e es s exis .
* E-mail: juha-pekka. aha@u u. i
In oduc ion
The use o molecula ma ke s o de e mina ion o an
indi idual’s o igin is an impo an ool in he managemen and
conse a ion o domes ic and wild species [1]. Indi idual gene ic
assignmen has also been used in o ensic cases o de ec illegal
ade and ansloca ion o animals [2], illegal ha es ing [1], and
sou ce o o igin o escaped domes ica ed animals [3], [4].
Assigning indi iduals o popula ions o o igin, also known as
gene ic s ock iden i ica ion (GSI), has been pa icula ly impo an
managemen ool in salmonid ishes [5].
Due o hei high a iabili y and a ailabili y, mic osa elli es o
sho andem epea s (STR), ha e been he ma ke s o choice o
GSI o nea ly wo decades [5], [6]. Howe e , he nume ous
ad an ages o single nucleo ide polymo phism (SNP) ma ke s such
as high abundance, p ocessing e iciency, ease o sco ing and
s anda dizing among labo a o ies make SNPs a ac i e o
indi idual gene ic assignmen s udies [5]. On he o he hand,
due o hei bi-allelic na u e, he powe o single SNP loci is
limi ed, equi ing a la ge numbe o independen loci in
compa ison o STRs. To o e come low a e age assignmen powe
o SNPs compa ed o mul i-allelic loci, selec ing a small subse o
highly in o ma i e loci om a la ge numbe o SNPs has been
p oposed [1], [7], [8]. Howe e , ini ial sc eening o highly
in o ma i e ma ke s om among housands o SNPs in mul iple
popula ions is expensi e.
To o e come he high cos o la ge-scale SNP geno yping,
de e mina ion o allele equencies om pooled DNA, i.e.,
‘allelo yping’, has been sugges ed as a cos -e ec i e al e na i e
o ob aining eliable allele equency in o ma ion o housands o
SNPs [9], [10]. These s udies ha e demons a ed high accu acy
and epea abili y o DNA pooling app oach, educing cos s up o
100 old, depending on he numbe o samples [11]. Allelo yping
has also allowed e icien de ec ion o genes associa ed wi h
nume ous ai s and diseases [12], [13]. Since allelo yping allows
de ec ion o ma ke s wi h la ge be ween-g oup allele equency
di e ences [14], his app oach can be applied o iden i y
PLOS ONE | www.plosone.o g 1 Decembe 2013 | Volume 8 | Issue 12 | e82434
popula ion-in o ma i e ma ke s, i.e., a small se o powe ul
ma ke s ha enables accu a e gene ic s ock iden i ica ion.
Fo he pas decade GSI has been an in aluable ool o he
managemen and conse a ion o A lan ic salmon popula ions,
enabling es ima ion o ela i e con ibu ions o a ious popula ions
in mixed s ock ishe ies [15] as well as iden i ying he popula ion o
o igin o indi idual ish [16]. Fo example, hese me hods a e used
o de ec ing popula ion speci ic mig a ion pa e ns [17], es ima -
ing he p opo ion o a m escapees in salmon ishe ies [4], [18]
and iden i ica ion o non-na i e ha che y-b ed indi iduals in wild
popula ions [19].
He e we es ed he easibili y o combining DNA pooling and
SNP a ays o iden i ica ion o highly in o ma i e SNPs o
accu a e GSI in A lan ic salmon, ocusing on indi idual assign-
men . We es ima ed allele equencies o 2880 SNPs om DNA
pools o 23 salmon popula ions using an A lan ic salmon Illumina
SNP-chip. We compa ed he pe o mance o ou common
app oaches (global F
ST
, pai wise F
ST
,Del a and ou lie ) o iden i y
he mos in o ma i e SNPs. We subsequen ly e alua ed he e ec s
o speci ic popula ion da ase and numbe o SNPs on indi idual
assignmen . We compa ed he pe o mance o he op SNPs
agains 31 STR loci. We also es ed he combined powe o
exis ing STR panels wi h he mos in o ma i e SNPs. Finally, we
compiled he la ges popula ion gene ic da ase o A lan ic salmon
o da e, bo h in e ms o geog aphic co e age and he numbe o
samples, by me ging ou da a wi h published da a [9], [20] and
alida ed he pe o mance o ou op SNPs on an independen se
o popula ions co e ing he main Eu opean dis ibu ion ange.
Ma e ials and Me hods
Samples & DNA Pooling
In o al, 1424 indi iduals we e collec ed om 23 A lan ic
salmon popula ions spawning in he i e s along he No wegian
and Russian no h-wes and Bal ic Sea coas s be ween 17uE and
57uE (Fig. 1). Salmon ju eniles we e collec ed by elec o ishing,
sac i iced by decapi a ion and a issue sample o each indi idual
s o ed immedia ely in 70% e hanol. The pe mi s o sample
collec ion we e issued by: 1) Fede al Agency o Fishe ies (Russia),
2) Coun y Go e no o Finnma k and T oms (No way), 3) Cen e
o Economic De elopmen , T anspo and he En i onmen
(Finland), and 4) Minis y o En i onmen (Es onia). As he ish
we e sac i iced immedia ely a e sampling and no expe imen s
wi h li ing ish we e pe o med he app o al o e hics commi ee
was no equi ed (EU di ec i e 2010/63/EU, Russian Fede a ion
go e nmen egula ion 2009/921, No wegian Animal Wel a e Ac
19/06/2009). This da ase included samples om 14 popula ions
s udied by Oze o e al. [9] (Table 1). Simila o ou p e ious
wo k [9], equimola (10 ng/ul) DNA ex ac s om 40 o 70
indi iduals we e pooled o p o ide om h ee o six echnical
eplica es o each popula ion sample (Table 1, Fig. 2). The pooled
DNA samples we e analyzed in he Cen e o In eg a i e
Gene ics (CIGENE, No way) using an Illumina in inium assay
(Illumina, San Diego, CA, USA) and e sion 2 o he A lan ic
salmon SNP-chip [20], [21] ca ying p obes o 5568 SNP
ma ke s. Among ou s udy popula ions, 3928 o hose SNP loci
we e bi-allelic and we e u he analyzed in his s udy. The aw
SNP da a we e analyzed using Geno yping module . 1.9.4
(Genome S udio so wa e . 2011.1, Illumina Inc.). In addi ion,
he same 23 popula ions used o DNA pool cons uc ion we e
indi idually geno yped using 31 commonly applied [17], [22] STR
ma ke s (Table 1, Table S1 in Appendix S1). The STR da a we e
analyzed and he geno ypes we e sco ed wi h Genemappe 4.1
so wa e (Applied Biosys ems).
Allele F equency Es ima ion
Allele equencies o 23 popula ions we e es ima ed om DNA
pools compa ing pool-speci ic alue o he a wi h he e e ence
alues o he a de i ed om 300 A lan ic salmon specimens
geno yped by CIGENE [9]. B ie ly, he aw colo signal da a om
2 al e na e alleles is con e ed in o a he a alue which anges
om 0 o 1. In heo y, an indi idual homozygous o an allele
would ha e a he a alue close o 0 o 1, and a alue o 0.5 would
indica e a he e ozygous geno ype. Howe e , in eali y a SNP’s
he a o geno ype clus e s (AA, AB and BB) a y om heo e ical
alues o 0, 0.5 and 1. The e o e, o es ima ion o allele equency
in a pooled sample, he he a alue o each SNP is compa ed o
he mean he a alues o AA, AB and BB geno ypes calcula ed by
geno yping o indi idual samples applying co ec ion algo i hm,
me hod 2 in [23].
S ingen quali y con ol il e s esul ed in selec ion o 2880 bi-
allelic SNPs showing low e o a es ( a ia ion o he a among
echnical eplica es #0.02) compa ed o in o ma ion con en [9].
Popula ion-speci ic allele equencies o each SNP we e es ima ed
as a mean o e 3–6 echnical eplica es (Appendix S2).
Popula ion Rela ionships and wi hin Popula ion Di e si y
In o de o in e he gene ic ela ionships among popula ions,
pai wise D
A
[24] dis ances and pai wise F
ST
we e calcula ed om
allele equency es ima es de i ed om allelo yping o 2880 SNPs
wi h he Powe Ma ke 3.25 so wa e package [25]. The D
A
gene ic dis ances we e used o cons uc neighbo -joining ees
wi h 1000 boo s ap eplica es. The same app oach was applied
o he 31 STR ma ke s. Consensus dend og ams we e cons uc -
ed sepa a ely o SNPs and STRs by using he p og am
Spli sT ee4 4.11.3 [26]. Simila ly, Powe Ma ke 3.25 [25] was
used o es ima e expec ed he e ozygosi y (H
E
) o popula ions o e
all SNP and STR loci.
Selec ion o he mos In o ma i e SNPs
We e alua ed ou di e en me hods o iden i ica ion o he
mos in o ma i e se o SNPs o GSI: Del a [27], global F
ST
,
pai wise F
ST
[28] and he ou lie app oach [29]. The es ima e o
allele- equency di e en ial, i.e., Del a, is one o he s aigh o wa d
ways o e alua e he in o ma ion con en o a SNP. Fo a bi-allelic
ma ke , like SNP, he Del a alue is es ima ed as |pA
i
-pA
j
|, whe e
pA
i
and pA
j
a e he equencies o allele A in he i
h
and j
h
popula ions, espec i ely. Del a alue o each SNP ma ke was
es ima ed as he mean ac oss all pai -wise compa isons o 23
popula ions. Ano he common c i e ion o selec ing he mos
in o ma i e loci is he popula ion di e en ia ion measu e F
ST
: he
unbiased es ima es o F
ST
we e i s calcula ed o e all popula ions
(global F
ST
) and on a pai wise basis (pai wise F
ST
). SNPs in he
uppe qua ile o dis ibu ion o di e gence alues we e classi ied
as ma ke s ha ing ‘‘high le el’’ o gene ic di e en ia ion (Fig. S1 in
Appendix S1). Fo he i s h ee app oaches (i.e., Del a, global F
ST
and pai wise F
ST
), he op 300 unlinked SNPs (.1 cM dis ance
om each o he ) we e selec ed o subsequen analyses.
Fo iden i ica ion o SNPs de ia ing om he neu al expec a-
ions (ou lie s), a Bayesian likelihood me hod was used, imple-
men ed in Bayescan 2.01 [29]. The me hod p o ides pos e io
odds (PO) as he a io o he pos e io p obabili ies indica ing how
much mo e likely he model wi h selec ion is compa ed o he
neu al model. Pos e io odd alues be ween 10 and 32
(log
10
(PO) = 1–1.5) a e conside ed as s ong e idence o selec ion,
be ween 32 and 100 (log
10
(PO) = 1.5–2) – as e y s ong, and PO
abo e 100 (log
10
(PO) .2) a e iewed as decisi e e idence o
selec ion [1], [29]. Depending on popula ion da ase , 35–111
unlinked ou lie s (.1 cM dis ance om each o he ) po en ially
Finding Ma ke s Tha Make a Di e ence o GSI
PLOS ONE | www.plosone.o g 2 Decembe 2013 | Volume 8 | Issue 12 | e82434
in luenced by di e gen selec ion [30] we e iden i ied and used o
subsequen analyses.
To es ima e how he numbe o SNPs a ec he pe o mance o
GSI, subse s o he op 25, 50, 75, 100, 125, 150, 200, 250 and 300
SNP loci we e selec ed o each o he ou anking app oaches. In
addi ion, we also chose simila ly-sized subse s o andom SNPs (i.e.
25, 50, 75, 100, 125, 150, 200, 250 and 300 SNP loci). To e alua e
he e ec o popula ions on he selec ion o op SNPs we es ed he
o e all assignmen powe o he mos in o ma i e SNPs iden i ied
using h ee di e en popula ion da ase s. The i s da ase included
all 23 popula ions (da ase I); he second se consis ed o 16
popula ions (1–16) excluding he eas e nmos and Bal ic salmon
(da ase II); and inally, he hi d se consis ed o six popula ions (1,
3, 5, 11, 13, 15) e enly dis ibu ed ac oss he No wegian and he
Wes e n Ba en s seas coas s (da ase III; Table 1).
Pe o mance o Top SNPs and STRs o GSI
As he me hods applied o GSI equi e geno ype da a a he
han allele equency es ima es, he mul ilocus geno ypes we e
simula ed om he allele equency es ima es assuming Ha dy-
Weinbe g and linkage equilib ium using bespoke so wa e (see
Appendix S3 o he code). Fo each subse o SNP loci, 100
mul ilocus geno ypes pe popula ion we e simula ed as a baseline
sample. Ano he 500 geno ypes pe popula ion o he mixed s ock
ishe y sample we e simula ed o each SNP subse using ONCOR
[31]. A simila app oach was applied o he STR da a.
The assignmen o indi iduals in a mix u e o baseline
popula ions was pe o med by using ONCOR [31]. This
app oach es ima es a p obabili y ha an indi idual (o unknown
o igin) belongs o a baseline popula ion by assessing es ima es o
he geno ype equencies in each baseline popula ion [32] and an
es ima e o he s ock composi ion o he ishe y [31]. In ONCOR,
he simula ed mixed s ock ishe y sample was es ed agains
baseline da a se along wi h he ‘‘Assign indi iduals o he baseline
popula ion’’ op ion o assign each ish. All indi iduals we e
assigned i espec i e o p ecision. As he so wa e used was
speci ically de eloped o analyzing samples o indi iduals he
in luence o he dis ibu ion o geno ypes in he sample being
es ed was examined using mix u es o ish wi h a ying
composi ions. I was ound ha a he le els o di e en ia ion
obse ed he e hese composi ions had only mino in luence on he
assignmen esul s (see Table S2 in Appendix S1). In addi ion o
SNPs, pe o mance was e alua ed o 31 STRs. We also e alua ed
he pe o mance o a 31-locus STR panel combined wi h di e en
numbe s o SNPs (1, 2, 5, 10, 25, 50 and 100 op- anked loci). We
es ima ed he numbe o STR o SNP alleles equi ed o achie e
80%, 90% and 95% co ec assignmen s o each anking
app oach and popula ion da ase . We made hese es ima es by
i ing a non-linea eg ession model o he cu es o co ec
assignmen pe cen age agains cumula i e numbe o ma ke s. An
Figu e 1. Map indica ing sampling loca ions o he s udied popula ions. See Table 1 o popula ion names. Eu opean A lan ic salmon
samples [20] used o alida ion o op anked SNPs a e indica ed as illed iangles.
doi:10.1371/jou nal.pone.0082434.g001
Finding Ma ke s Tha Make a Di e ence o GSI
PLOS ONE | www.plosone.o g 3 Decembe 2013 | Volume 8 | Issue 12 | e82434
exponen ial eg ession model (y= exp(a+b/x)) was ound o bes i
he da a.
Analysis o he Independen Da ase
To e alua e he use ulness o ou app oach, and he e ec i e-
ness o ou bes SNPs o GSI in di e en se s o popula ions,
independen alida ion was pe o med on he A lan ic salmon
indi iduals geno yped by Bou e e al. [20], DRYAD en y
doi:10.5061/d yad.gm367. Speci ically, we e alua ed he pe o -
mance o ou op- anked loci (da ase II, pai wise F
ST
selec ion
app oach, 25 o 150 SNPs) o GSI on 26 Eu opean anad omous
popula ions anging om Spain (Na cea) o Russia (Se e naja
D ina). Compa ed o ou da a, only h ee popula ions s udied by
Bou e e al. [20] o igina ed om he i e sys ems included in
bo h da ase s (Tana/Teno, Ponoi, Va zuga). Addi ionally, we
es ed he eliabili y o he allelo yping app oach by combining he
allele equency es ima es om DNA pools o 23 popula ions wi h
he 38-popula ion da ase o Bou e e al. [20] which consis ed o
indi idual geno ypes.
Resul s
Gene ic Di e si y and Di e en ia ion: SNPs s. STRs
As expec ed, he gene ic di e si y o SNPs o e all popula ions
was signi ican ly lowe compa ed o STRs (median SNP
He
= 0.36
s. median STR
He
= 0.77, Mann-Whi ney U- es , P,0.001). The
gene ic di e si y (H
E
) o popula ions o e all SNP loci anged om
0.23 o 0.35, whe eas o STR da a H
E
es ima es we e highe ,
anging om 0.64 o 0.74 (Table 1). Howe e , gene ic di e si y
es ima es wi hin popula ions (H
E
) we e signi ican ly co ela ed
be ween he wo ma ke ypes (Pea son’s = 0.93, P,0.0001).
Pai wise popula ion di e en ia ion (F
ST
) es ima es o e 2880 SNPs
a ied om 0.01 (Ti o ka s. U a) o 0.30 (Na a s. Pecho a
Unya), whe eas mean pai wise F
ST
alues o e 31 STRs anged
Table 1. In o ma ion abou popula ions included in he da ase s used o SNP selec ion and hei geog aphic loca ions.
Popula ion Coo dina es
N
SNP
N
STR
H
E
SNPs
H
E
STRs Popula ion da ase
I II III
No wegian Sea
1 Laukhelle* 69u13’N 17u50’E 42 (4) 42 0.35 0.73 x x x
2Ma
˚lsel a 69u13’N 18u29’E 70 (3) 70 0.34 0.72 x x
3 Reisa 69u46’N 21u00’E 70 (3) 70 0.31 0.69 x x x
4 Al a* 69u58’N 23u22’E 70 (3) 65 0.32 0.69 x x
5 Reppa jo del * 70u26’N 24u19’E 69 (4) 67 0.35 0.73 x x x
Wes e n Ba en s Sea
6 Laksel a* 70u04’N 24u55’E 67 (5) 67 0.32 0.68 x x
7 Iesjoki (Teno/Tana) 69u26’N 24u59’E 70 (3) 70 0.32 0.71 x x
8 Ka asjoki (Teno/Tana)* 69u23’N 25u09’E 70 (4) 63 0.33 0.70 x x
9 Ina ijoki (Teno/Tana)* 69u00’N 25u46’E 67 (4) 67 0.33 0.70 x x
10 Yla
¨ko
¨nga
¨s (Teno/Tana) 69u57’N 26u34’E 58 (3) 58 0.32 0.71 x x x
11 Tana B u (Teno/Tana)* 70u12’N 28u11’E 60 (2) 59 0.34 0.70 x x
12 Ves e Jakobsel * 70u06’N 29u19’E 70 (4) 59 0.33 0.72 x x
13 Neiden* 69u42’N 29u31’E 63 (4) 63 0.33 0.72 x x x
14 Ti o ka* 69u30’N 31u58’E 70 (4) 67 0.35 0.74 x x
15 U a* 69u16’N 32u48’E 44 (3) 44 0.34 0.73 x x x
16 Kola* 68u52’N 33u01’E 70 (6) 70 0.33 0.74 x x
Whi e Sea
17 Ponoi* 66u58’N 41u16’E 70 (4) 70 0.33 0.72 x
18 Va zuga* 66u12’N 36u57’E 70 (4) 70 0.31 0.70 x
19 Onega 63u54’N 38u00’E 70 (3) 70 0.25 0.65 x
20 Mezen Pizhma 65u53’N 44u08’E 48 (3) 48 0.28 0.68 x
Eas e n Ba en s Sea
21 Pecho a Pizhma 64u52’N 51u16’E 48 (3) 48 0.26 0.67 x
22 Pecho a Unya 61u47’N 57u52’E 48 (3) 48 0.24 0.64 x
Bal ic Sea
23 Na a 59u23’N 28u12’E 40 (3) 40 0.23 0.64 x
To al pooled (SNPs)
1424
To al indi idual (STRs)
1395
N
SNP
– numbe o indi iduals included in he pools and numbe o echnical eplica es (in b acke s) o SNP-chip analysis, N
STR
– numbe o samples o indi idual
geno yping using 31 STRs, H
E
– o e all expec ed he e ozygosi y o SNPs and STRs.
*Da a om Oze o e al. [9].
doi:10.1371/jou nal.pone.0082434. 001
Finding Ma ke s Tha Make a Di e ence o GSI
PLOS ONE | www.plosone.o g 4 Decembe 2013 | Volume 8 | Issue 12 | e82434
om 0.02 o 0.20 o he same popula ion pai s (Table S3 in
Appendix S1). Simila o gene ic di e si y, gene ic di e gence o
he popula ions (mean pai wise F
ST
) was signi ican ly co ela ed
be ween SNPs and STRs (Man el’s
xy =
0.94, P,0.0001). Thus,
bo h ma ke classes e ealed e y simila popula ion gene ic
s uc u ing as illus a ed by neighbo -joining ees (Fig. 3).
Howe e , he le el o di e en ia ion o SNPs o e all popula ions
was signi ican ly highe han ha o STRs (median SNP
Fs
= 0.077
s. median STR
Fs
= 0.055, Mann-Whi ney U- es , P,0.001).
Iden i ica ion o he mos In o ma i e SNPs
As expec ed, only a small p opo ion o SNPs (ou o 2880)
exhibi ed high le els o gene ic di e en ia ion es ima ed using
h ee measu es (global F
ST
, pai wise F
ST
and Del a, Appendix S4,
Fig. S1 in Appendix S1). The ou lie es o popula ion da ase s I,
II and III iden i ied pu a i e signs o di e gen selec ion (log
10
(PO)
.1, q,0.05) a 141, 120 and 41 SNPs, espec i ely (Fig. 4).
Howe e , as se e al SNPs o med igh ly linked g oups (,1 cM), a
o al o 111, 95 and 35 unlinked ou lie SNPs we e e ained o
subsequen analysis o popula ion da ase s I, II and III,
espec i ely.
The compa ison o allele equency dis ibu ions o 100 op
SNPs anked using Del a o pai wise F
ST
showed ha hese wo
app oaches iden i ied loci wi h a wide ange o allele equencies
among popula ions (Fig. 5). A simila pa e n was e iden o
ou lie s, wi h a sligh ly highe p opo ion o loci (42%) showing
ma ginal allele equency dis ibu ions (close o 0 o 1). In con as ,
he allele equencies o a majo i y o he SNPs (69%) anked using
global F
ST
measu e we e biased owa ds 0 o 1 wi h ela i ely ew
loci exhibi ing in e media e allele equencies.
Despi e he di e ences desc ibed abo e, he e was a subs an ial
amoun o o e lap be ween he op SNPs among all anking
app oaches (Fig. 6). Fo example, 59 o 68 o SNPs ou o he op
100 we e iden i ied using paiwise F
ST
and Del a in all h ee
da ase s. Simila ly, la ge p opo ions o SNPs we e sha ed be ween
global F
ST
and pai wise F
ST
app oach (51%–71%). In addi ion, a
subs an ial p opo ion o ou lie s (up o 58%) was also anked as
op SNPs by he h ee o he app oaches (Fig. 6A).
In con as o SNP anking app oaches, popula ions in he
asce ainmen g oup had a much la ge e ec on anking o he
mos in o ma i e SNPs. Fo example, a ela i ely la ge p opo ion
o SNPs (up o 80%) anked by global F
ST
, pai wise F
ST
and Del a
was iden i ied only in a single da ase (Fig. 6B) and a ela i ely
la ge p opo ion o ou lie s was unique o each da ase (up o
50%).
Pe o mance o Top- anked SNPs o GSI
Compa ed o andomly chosen se s o SNPs, he o e all
assignmen success was conside ably highe o op- anked loci
selec ed using ou di e en app oaches (Fig. 7). O hese, he
ou lie me hod iden i ied he bes pe o ming loci, while loci
selec ed using he global F
ST
app oach esul ed in he lowes
o e all assignmen accu acy. Howe e , when he numbe o SNP
ma ke s eached 100, he o e all assignmen success o loci
iden i ied by global F
ST
, pai wise F
ST
and Del a was a he simila ,
bu s ill lowe han o loci iden i ied by he ou lie me hod (Fig. 7,
Table S4 in Appendix S1). In o de o achie e 80%, 90% and
95% co ec assignmen o 23 A lan ic salmon popula ions, he
ou lie app oach equi ed 15–20% ewe loci han he h ee o he
app oaches (Table 2). Fo example, 95% o e all co ec assign-
men was achie ed using 94 ou lie SNPs iden i ied using
popula ion da ase II, whe eas eaching he same le el o
assignmen powe wi h SNPs anked by he h ee o he
app oaches equi ed 118 o 125 SNPs (Table 2).
Al hough he o e all GSI success was high, he numbe o SNPs
equi ed o p o ide simila accu acy a ied in indi idual
popula ions. Fo example, o Wes e n Ba en s and No wegian
Sea popula ions, app oxima ely 100–150 SNPs we e necessa y o
a ain .90% popula ion assignmen success, whe eas 25–50 op
SNPs we e enough o a ain simila le el o co ec assignmen in
he Eas e n Ba en s, Whi e and Bal ic sea popula ions (Table S5 in
Appendix S1). Howe e , o ma ke s anked using popula ion
da ase s II and III, i.e., when he mos gene ically dis inc
eas e nmos and Bal ic popula ions we e emo ed, he assignmen
success o .90% o Wes e n Ba en s and No wegian Sea salmon
was achie ed using 75–100 op- anked SNPs (Table S5 in
Appendix S1), excep o a g oup o h ee popula ions wi h he
lowes gene ic di e gence (pai wise F
ST
o e 2880 SNPs = 0.016–
0.023; Table S3 in Appendix S1).
When compa ing he assignmen powe o a gi en numbe o
independen alleles be ween wo ma ke classes, he pe o mance
o STRs was lowe han ha o andom and op anked SNPs. Fo
example, an 80% o e all assignmen success was a ained wi h 219
independen alleles o STRs (18 loci) while 96 and 47
independen alleles we e su icien o each simila accu acy o
andom and op- anked SNPs, espec i ely (Table 2, Fig. 7). On
he o he hand, when he assignmen powe was es ima ed o a
gi en numbe o loci, STRs pe o med be e han SNPs as less
mul i-allelic ma ke s we e needed o each gi en le el o
assignmen accu acy compa ed o bi-allelic ma ke s. When
e alua ing he assignmen success o indi idual popula ions, 18
STRs (219 independen alleles) we e su icien o assign salmon
popula ions om he Bal ic and he Whi e and Eas e n Ba en s
seas wi h .99% accu acy, whe eas all 31 STR ma ke s (536
independen alleles) we e equi ed o achie e .90% accu acy o
he Wes e n Ba en s and No wegian Sea popula ions. A combi-
na ion o 31 STR ma ke s and 25 op- anked SNPs inc eased he
o e all assignmen accu acy om 97% o 99% (Fig. S2 in
Appendix S1).
Figu e 2. Wo k low diag am indica ing he main s eps o he
analyses.
doi:10.1371/jou nal.pone.0082434.g002
Finding Ma ke s Tha Make a Di e ence o GSI
PLOS ONE | www.plosone.o g 5 Decembe 2013 | Volume 8 | Issue 12 | e82434
Valida ion o Top- anked SNP on he Independen
Da ase
The se o ou op 100 SNPs iden i ied in da ase II using he
pai wise F
ST
selec ion app oach allowed .98% co ec assign-
men in 13 ou 26 Eu opean anad omous A lan ic salmon
popula ions (Table S6 in Appendix S1). The lowes assignmen
accu acy was obse ed in B i ish, Sco ish and I ish popula ions
(66%–87%) and was in line wi h he lowe le el o gene ic
di e en ia ion among salmon popula ions in his a ea [33]. Simila
o ou No h-Eu opean da ase , he numbe o op SNPs equi ed
o achie e 90% and 95% o e all co ec assignmen o 26
Eu opean popula ions was conside ably lowe compa ed o
andomly chosen SNPs (Table S7 in Appendix S1). Howe e ,
he assignmen accu acy o bo h op- anked and andom ma ke s
eached simila high le els (98%) when o e 150 SNPs we e used
(Table S7 in Appendix S1).
The cons uc ed neighbo -joining ee consis ing o 2763 SNPs
om 61 popula ion da a de i ed by allelo yping and indi idual
geno yping demons a ed he compa ibili y o he wo app oaches
(Fig. 8): he gene ic ela ionships o he popula ions we e consis en
wi h he esul s o Bou e e al. [20]. Howe e , he combined
da ase u he e ealed new insigh s in o gene ic ela ionships
among popula ions such as sepa a ion be ween no he n and
sou he n No wegian popula ions.
Discussion
This is he i s s udy ha combines high- h oughpu SNP
a ays and DNA pooling (i.e., allelo yping) o iden i y he mos
powe ul se s o SNPs o gene ic s ock iden i ica ion. We
demons a e how SNP a ays and DNA pooling enable a as
and cos -e ec i e, ye eliable me hod o iden i ying he mos
in o ma i e ma ke s among housands o SNPs om la ge numbe
o A lan ic salmon popula ions. In line wi h he p e ious s udies o
Figu e 3. Gene ic ela ionships among 23 A lan ic salmon popula ions in no he n Eu ope. Neighbou -joining dend og am is based on
Nei’s D
A
gene ic dis ances es ima ed using (A) 2880 SNPs and (B) 31 STR ma ke s. Dis inc popula ion g oups a e colo ed. The b anches wi h
boo s ap alue suppo ,80% a e d awn as dashed.
doi:10.1371/jou nal.pone.0082434.g003
Finding Ma ke s Tha Make a Di e ence o GSI
PLOS ONE | www.plosone.o g 6 Decembe 2013 | Volume 8 | Issue 12 | e82434
human popula ions [14], he esul s a e mos encou aging o
p ojec s in ol ing high-sample h oughpu and low- o medium-
mul iplex SNP geno yping as a much smalle numbe o SNPs is
needed o accu a e gene ic s ock iden i ica ion compa ed o
andomly chosen SNPs. Mo eo e , we demons a e he applica-
bili y o ou app oach on a da a se compiled om wo sepa a e
SNP geno yping p ojec s and illus a e ans e abili y o he da a
ac oss s udies wi hou he need o labo ious s anda diza ion in
compa ison wi h e.g. mic osa elli es [22].
Reliabili y o DNA Pooling & Allelo yping
Recen ly, we showed ha allelo yping o DNA pools is an
e ec i e me hod o eliable allele equency es ima ion (indi idual
geno yping s. allelo yping, Pea son’s = 0.992) [9]. To u he
alida e he obus ness o allelo yping, we compa ed ou allele
equency es ima es de i ed om allelo yping o allele equencies
ob ained om indi idual geno yping by Bou e e al. [20] o
h ee di e en i e sys ems (Teno/Tana, Ponoi and Va zuga).
Despi e di e en indi idual samples, sampling loca ions and
sampling yea s we obse ed e y high co ela ion be ween he
allele equency es ima es (Pea son’s = 0.952–0.971) indica ing he
eliabili y o allelo yping app oach. Mo eo e , he cos s o analysis
o DNA pools was abou 15 imes lowe compa ed o he analysis
o he same numbe o samples indi idually, i.e. he p ice o
genome-wide analysis o DNA pools is a leas an o de o
magni ude lowe han indi idual geno yping [11–14].
This s udy also allowed a di ec compa ison o a ious
popula ion gene ic pa ame e s de i ed om allelo yping and
geno yping o SNPs and STRs, espec i ely. We obse ed highly
signi ican co ela ion be ween expec ed he e ozygosi y es ima es
o he wo ma ke ypes. Simila ly, he es ima es o pai wise
gene ic di e en ia ion demons a ed highly co ela ed pa e ns o
SNP and STR ma ke s. These esul s a e in ag eemen wi h he
ea lie indings showing high conco dance be ween di e en
ma ke ypes [5], [6], [34], [35]. Likewise, he gene ic ela ionships
among popula ions in e ed by SNPs and STRs we e simila wi h
high boo s ap clus e ing suppo o bo h ma ke classes.
Howe e , in con as o o he s udies [20], [36], we obse ed a
ine sepa a ion o he Wes e n Ba en s salmon in o wo g oups
(Teno and Wes e n Ba en s/No wegian Sea). Fu he mo e, wi h a
combined da ase consis ing o 61 popula ions we we e able o
con i m he gene ic ela ionships among popula ions o e he
whole dis ibu ion ange as well as o e eal no el pa e ns such as
clea sepa a ion be ween no he n and sou he n No wegian
popula ions. Taken oge he , hese esul s no only demons a e
he eliabili y o DNA pooling and allelo yping app oach, bu also
illus a e one o he impo an ad an ages o SNPs – good
ans e abili y o he SNP da a be ween independen s udies. In
con as o STRs, which usually equi e labo ious calib a ion and
s anda diza ion o alleles [22], [37], his acili a es e icien
compila ion o la ge SNP da ase s om independen s udies
c ea ing u he syne gis ic e ec s. Howe e , SNP allele label
swi ching may u n ou p oblema ic when SNP da a se s om
di e en geno yping pla o ms a e compa ed.
De ec ion and Pe o mance o he mos In o ma i e SNPs
The e alua ion o ou SNP selec ion app oaches demons a ed
ha all o hem subs an ially imp o e he accu acy o GSI
compa ed o andomly chosen SNPs. In mos cases, SNPs selec ed
using he ou lie app oach showed he highes assignmen
accu acy, whe eas he global F
ST
app oach usually esul ed in
lowe co ec assignmen a es. On he o he hand, he di e ences
in indi idual assignmen success we e e iden only when he
numbe o SNPs was below 100 while all SNP selec ion
app oaches enabled accu a e GSI when mo e han 100 SNPs
we e used. These esul s a e consis en wi h he ea lie esul s
demons a ing supe io pe o mance o ou lie loci o e neu al
loci o gene ic s ock iden i ica ion [30], [38], [39]. Simila ly, he
ou pe o mance o pai wise SNP selec ion me hods o e he global
F
ST
app oach has been shown in bo h humans [40] and ca le [8].
This is because he global F
ST
app oach ends o selec o ma ke s
ha a e speci ic o he mos dis inc popula ion o g oup o
popula ions while he pai wise app oaches allow selec ion o
ma ke s wi h high he e ozygosi y and mo e e enly dis ibu ed
allele equencies among popula ions, being hus mo e in o ma i e
o indi idual assignmen [8], [40], [41].
The e was also a subs an ial amoun o o e lap be ween he op
loci among all anking app oaches. Fo example, he o e lap o
Figu e 4. Iden i ica ion o ou lie loci using a model-based
genome scan app oach. (A) popula ion da ase I; (B) popula ion
da ase II; (C) popula ion da ase III. Each SNP locus ( illed ci cle) is
ep esen ed by he le el o gene ic di e en ia ion (F
ST
) and log
10
(PO) o
being unde selec ion. Ou lie loci po en ially unde di e gen selec ion
a e inside dashed ec angle.
doi:10.1371/jou nal.pone.0082434.g004
Finding Ma ke s Tha Make a Di e ence o GSI
PLOS ONE | www.plosone.o g 7 Decembe 2013 | Volume 8 | Issue 12 | e82434
op 100 SNPs selec ed using global F
ST
, pai wise F
ST
o Del a
anged om 51 o 71%. Mo eo e , a la ge p opo ion (up o 58%)
o he op 100 SNPs iden i ied using global F
ST
, pai wise F
ST
o
Del a showed signs o di e gen selec ion. This is consis en wi h
he esul s o Lao e al. [7] showing ha he i e mos in o ma i e
SNPs di iding human popula ions om di e en con inen s
Figu e 5. Dis ibu ion o allele equencies o he 100 op- anked SNPs. Allele equencies o he 100 op- anked SNPs in 23 popula ions
iden i ied using ou selec ion app oaches: (A) global F
ST
; (B) pai wise F
ST
; (C) Del a, and (D) ou lie . Ho izon al line, g ey ec angle, whiske s, open
ci cles, and s a s indica e median, 25 h and 75 h qua iles, non-ou lie ange, ou lie s and ex eme ou lie s, espec i ely.
doi:10.1371/jou nal.pone.0082434.g005
Figu e 6. SNP o e lap among di e en anking app oaches and popula ion da ase s. (A) Venn diag ams showing he ex en o o e lap
among ou app oaches (global F
ST
, pai wise F
ST
,Del a and ou lie ) o h ee popula ion da ase s. (B) Venn diag ams showing he ex en o o e lap
among h ee popula ion da ase s o ou anking app oaches. Fo all SNP anking me hods he op 100 SNPs a e p esen ed, excep o he ou lie
app oach whe e 95 and 35 SNPs we e iden i ied as being unde selec ion o da ase II and III, espec i ely.
doi:10.1371/jou nal.pone.0082434.g006
Finding Ma ke s Tha Make a Di e ence o GSI
PLOS ONE | www.plosone.o g 8 Decembe 2013 | Volume 8 | Issue 12 | e82434
exhibi he signs o local posi i e selec ion. Thus, ou esul s
indica e ha despi e he high assignmen powe o non-neu al
ma ke s, mo e simple pai wise me hods a e nea ly as e icien in
anking he mos in o ma i e loci o GSI.
On he o he hand, ou esul s indica e ha he popula ion
da ase migh play a la ge ole o iden i ica ion o he mos
in o ma i e ma ke s han he SNP selec ion app oach [42], [43].
Fo example, he mos in o ma i e se o loci selec ed using one se
o human popula ions ha e been shown o lack powe when
applied o ano he se o popula ions [7], [42]. Indeed, ou analysis
e ealed ha he assignmen accu acy was mo e a ec ed by
popula ions in he asce ainmen g oup used o anking SNPs
a he han anking me hod i sel . While he SNPs anked using
popula ion da ase I allowed quick disc imina ion o he popula-
ions a he la ge geog aphical scale, he ma ke s anked using
popula ion da ase s II and III pe o med be e a egional le el.
This can be explained by highe assignmen success o he
indi iduals om he Wes e n Ba en s Sea and No wegian Sea
popula ions. Indeed, he geog aphically emo e Bal ic popula ion
and salmon o he Whi e and Eas e n Ba en s seas ha e e y
dis inc gene ic p o iles ha allow hei disc imina ion using as ew
as 25–50 andomly chosen SNPs. On he o he hand, gene ic
s uc u e o salmon popula ions o he Wes e n Ba en s and
No wegian seas is mo e sub le [18], [36] and GSI o hese
popula ions equi es a highe numbe o ma ke s. Thus, he
exclusion o he eas e nmos salmon popula ions enabled de ec ion
o loci in o ma i e o he Wes e n Ba en s Sea and No wegian
Sea salmon inc easing he assignmen accu acy o hese popula-
ions. Gi en he esul s o ou and p e ious [7] s udies, pai wise-
based me hods o anking he mos in o ma i e ma ke s o GSI
a e mo e obus o he asce ainmen bias and hus a e mo e
applicable wi h ex ended da ase s.
I mus be ecognized ha ou app oach o pe o m powe
analyses o he op anked SNPs o gene ic s ock iden i ica ion
yields o e ly op imis ic accu acies, see [44]. Ou accu acy le els
a e upwa dly biased o wo easons: i) we simula ed geno ypes
Figu e 7. O e all assignmen success o SNPs and STRs in da ase I. SNPs we e anked using i) global F
ST
(blue), ii) pai wise F
ST
(b own), iii)
Del a (g een) and i ) ou lie app oach ( ed). O e all assignmen success o STRs and andomly chosen SNPs a e shown as g ay and black lines,
espec i ely. The ba s a e ep esen ing s anda d de ia ion o assignmen success among all popula ions o each anking app oach. The s anda d
de ia ion ba s we e a anged o isual pu poses o a oid o e lapping.
doi:10.1371/jou nal.pone.0082434.g007
Table 2. Es ima ed numbe o independen alleles o SNPs and STRs equi ed o achie e 80%, 90%, and 95% o e all co ec
assignmen in 23 A lan ic salmon popula ions o each anking me hod.
Popula ion da ase used o SNP anking
I (23) II (16) III (6)
Global
F
ST
Pai wise
F
ST
Del a
Ou lie
Global
F
ST
Pai wise
F
ST
Del a
Ou lie
Global
F
ST
Pai wise
F
ST
Del a
Ou lie
Random
SNPs STRs*
80% 76 68 59 50 57 50 53 47 62 53 58 n/a 96 219 (,18)
90% 113 104 100 81 90 82 87 71 98 87 95 n/a 149 336 (,24)
95% 167 153 147 114 124 118 125 94 133 123 136 n/a 198 460 (,28)
*Es ima ed numbe o STR loci is indica ed in pa en heses.
doi:10.1371/jou nal.pone.0082434. 002
Finding Ma ke s Tha Make a Di e ence o GSI
PLOS ONE | www.plosone.o g 9 Decembe 2013 | Volume 8 | Issue 12 | e82434