scieee Open visual document viewer

Finding markers that make a difference : DNA pooling and SNP-arrays identify population informative markers for genetic stock identification

Ozerov, Mikhail,Vasemägi, Anti,Wennevik, Vidar,Diaz-Fernandez, Rogelio,Kent, Matthew,Gilbey, John,Prusov, Sergey,Niemelä, Eero,Vähä, Juha-Pekka

Full text

Finding Ma ke s Tha Make a Di e ence: DNA Pooling and SNP-A ays Iden i y Popula ion In o ma i e Ma ke s o Gene ic S ock Iden i ica ion Mikhail Oze o 1 , An i Vasema ¨gi 2,8 , Vida Wenne ik 3 , Rogelio Diaz-Fe nandez 1 , Ma hew Ken 4 , John Gilbey 5 , Se gey P uso 6 , Ee o Niemela ¨ 7 , Juha-Pekka Va ¨ha ¨ 1 * 1Ke o Suba c ic Resea ch Ins i u e, Uni e si y o Tu ku, Tu ku, Finland, 2Depa men o Biology, Di ision o Gene ics and Physiology, Uni e si y o Tu ku, Tu ku, Finland, 3Resea ch g oup Popula ion Gene ics and Ecology, Ins i u e o Ma ine Resea ch, Be gen, No way, 4Depa men o Animal and Aquacul u al Sciences, Cen e o In eg a i e Gene ics (CIGENE), No wegian Uni e si y o Li e Sciences, A ˚s, No way, 5F eshwa e Labo a o y, Ma ine Sco land, Faskally, Pi loch y, Uni ed Kingdom, 6F eshwa e Resou ces Labo a o y, Knipo i ch Pola Resea ch Ins i u e o Ma ine Fishe ies and Oceanog aphy, Mu mansk, Russia, 7Finnish Game and Fishe ies Resea ch Ins i u e, Oulu, Finland, 8Depa men o Aquacul u e, Ins i u e o Ve e ina y Medicine and Animal Science, Es onian Uni e si y o Li e Sciences, Ta u, Es onia Abs ac Gene ic s ock iden i ica ion (GSI) using molecula ma ke s is an impo an ool o managemen o mig a o y species. He e, we es ed a cos -e ec i e al e na i e o indi idual geno yping, known as allelo yping, o iden i ica ion o highly in o ma i e SNPs o accu a e gene ic s ock iden i ica ion. We es ima ed allele equencies o 2880 SNPs om DNA pools o 23 A lan ic salmon popula ions using Illumina SNP-chip. We e alua ed he pe o mance o ou common s a egies (global F ST , pai wise F ST ,Del a and ou lie app oach) o selec ion o he mos in o ma i e se o SNPs and es ed hei e ec i eness o GSI compa ed o andom se s o SNP and mic osa elli e ma ke s. Fo he majo i y o cases, SNPs selec ed using he ou lie app oach pe o med bes ollowed by pai wise F ST and Del a me hods. O e all, he selec ion p ocedu e educed he numbe o SNPs equi ed o accu a e GSI by up o 53% compa ed wi h andomly chosen SNPs. Howe e , GSI accu acy was mo e a ec ed by popula ions in he asce ainmen g oup a he han he anking me hod i sel . We demons a ed o he i s ime he compa ibili y o di e en la ge-scale SNP da ase s by compiling he la ges popula ion gene ic da ase o A lan ic salmon o da e. Finally, we showed an excellen pe o mance o ou op SNPs on an independen se o popula ions co e ing he main Eu opean dis ibu ion ange o A lan ic salmon. Taken oge he , we demons a e how combina ion o DNA pooling and SNP a ays can be applied o conse a ion and managemen o salmonids as well as o he species. Ci a ion: Oze o M, Vasema ¨gi A, Wenne ik V, Diaz-Fe nandez R, Ken M, e al. (2013) Finding Ma ke s Tha Make a Di e ence: DNA Pooling and SNP-A ays Iden i y Popula ion In o ma i e Ma ke s o Gene ic S ock Iden i ica ion. PLoS ONE 8(12): e82434. doi:10.1371/jou nal.pone.0082434 Edi o : So ia Consueg a, Abe ys wy h Uni e si y, Uni ed Kingdom Recei ed May 2, 2013; Accep ed Oc obe 24, 2013; Published Decembe 16, 2013 Copy igh : ß2013 Oze o e al. This is an open-access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion License, which pe mi s un es ic ed use, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal au ho and sou ce a e c edi ed. Funding: This s udy was unded by he Eu opean Union, Kola c ic ENPI CBC p ojec KO197 ‘‘T ila e al coope a ion in ou common esou ce; he A lan ic salmon in he Ba en s Region’’ (M.O., R.D.F., E.N. and S.P.), Academy o Finland (J.-P.V.), No wegian Di ec o a e o Na u e Managemen (V.W.), No wegian Resea ch Council (V.W.), Es onian Minis y o Educa ion and Resea ch (Ins i u ional Resea ch Funding p ojec IUT8-2 o A.V.) and Es onian Science Founda ion (g an numbe s 6802, 8215 o A.V.). The unde s had no ole in s udy design, da a collec ion and analysis, decision o publish, o p epa a ion o he manusc ip . Compe ing In e es s: The au ho s ha e decla ed ha no compe ing in e es s exis . * E-mail: juha-pekka. aha@u u. i In oduc ion The use o molecula ma ke s o de e mina ion o an indi idual’s o igin is an impo an ool in he managemen and conse a ion o domes ic and wild species [1]. Indi idual gene ic assignmen has also been used in o ensic cases o de ec illegal ade and ansloca ion o animals [2], illegal ha es ing [1], and sou ce o o igin o escaped domes ica ed animals [3], [4]. Assigning indi iduals o popula ions o o igin, also known as gene ic s ock iden i ica ion (GSI), has been pa icula ly impo an managemen ool in salmonid ishes [5]. Due o hei high a iabili y and a ailabili y, mic osa elli es o sho andem epea s (STR), ha e been he ma ke s o choice o GSI o nea ly wo decades [5], [6]. Howe e , he nume ous ad an ages o single nucleo ide polymo phism (SNP) ma ke s such as high abundance, p ocessing e iciency, ease o sco ing and s anda dizing among labo a o ies make SNPs a ac i e o indi idual gene ic assignmen s udies [5]. On he o he hand, due o hei bi-allelic na u e, he powe o single SNP loci is limi ed, equi ing a la ge numbe o independen loci in compa ison o STRs. To o e come low a e age assignmen powe o SNPs compa ed o mul i-allelic loci, selec ing a small subse o highly in o ma i e loci om a la ge numbe o SNPs has been p oposed [1], [7], [8]. Howe e , ini ial sc eening o highly in o ma i e ma ke s om among housands o SNPs in mul iple popula ions is expensi e. To o e come he high cos o la ge-scale SNP geno yping, de e mina ion o allele equencies om pooled DNA, i.e., ‘allelo yping’, has been sugges ed as a cos -e ec i e al e na i e o ob aining eliable allele equency in o ma ion o housands o SNPs [9], [10]. These s udies ha e demons a ed high accu acy and epea abili y o DNA pooling app oach, educing cos s up o 100 old, depending on he numbe o samples [11]. Allelo yping has also allowed e icien de ec ion o genes associa ed wi h nume ous ai s and diseases [12], [13]. Since allelo yping allows de ec ion o ma ke s wi h la ge be ween-g oup allele equency di e ences [14], his app oach can be applied o iden i y PLOS ONE | www.plosone.o g 1 Decembe 2013 | Volume 8 | Issue 12 | e82434 popula ion-in o ma i e ma ke s, i.e., a small se o powe ul ma ke s ha enables accu a e gene ic s ock iden i ica ion. Fo he pas decade GSI has been an in aluable ool o he managemen and conse a ion o A lan ic salmon popula ions, enabling es ima ion o ela i e con ibu ions o a ious popula ions in mixed s ock ishe ies [15] as well as iden i ying he popula ion o o igin o indi idual ish [16]. Fo example, hese me hods a e used o de ec ing popula ion speci ic mig a ion pa e ns [17], es ima - ing he p opo ion o a m escapees in salmon ishe ies [4], [18] and iden i ica ion o non-na i e ha che y-b ed indi iduals in wild popula ions [19]. He e we es ed he easibili y o combining DNA pooling and SNP a ays o iden i ica ion o highly in o ma i e SNPs o accu a e GSI in A lan ic salmon, ocusing on indi idual assign- men . We es ima ed allele equencies o 2880 SNPs om DNA pools o 23 salmon popula ions using an A lan ic salmon Illumina SNP-chip. We compa ed he pe o mance o ou common app oaches (global F ST , pai wise F ST ,Del a and ou lie ) o iden i y he mos in o ma i e SNPs. We subsequen ly e alua ed he e ec s o speci ic popula ion da ase and numbe o SNPs on indi idual assignmen . We compa ed he pe o mance o he op SNPs agains 31 STR loci. We also es ed he combined powe o exis ing STR panels wi h he mos in o ma i e SNPs. Finally, we compiled he la ges popula ion gene ic da ase o A lan ic salmon o da e, bo h in e ms o geog aphic co e age and he numbe o samples, by me ging ou da a wi h published da a [9], [20] and alida ed he pe o mance o ou op SNPs on an independen se o popula ions co e ing he main Eu opean dis ibu ion ange. Ma e ials and Me hods Samples & DNA Pooling In o al, 1424 indi iduals we e collec ed om 23 A lan ic salmon popula ions spawning in he i e s along he No wegian and Russian no h-wes and Bal ic Sea coas s be ween 17uE and 57uE (Fig. 1). Salmon ju eniles we e collec ed by elec o ishing, sac i iced by decapi a ion and a issue sample o each indi idual s o ed immedia ely in 70% e hanol. The pe mi s o sample collec ion we e issued by: 1) Fede al Agency o Fishe ies (Russia), 2) Coun y Go e no o Finnma k and T oms (No way), 3) Cen e o Economic De elopmen , T anspo and he En i onmen (Finland), and 4) Minis y o En i onmen (Es onia). As he ish we e sac i iced immedia ely a e sampling and no expe imen s wi h li ing ish we e pe o med he app o al o e hics commi ee was no equi ed (EU di ec i e 2010/63/EU, Russian Fede a ion go e nmen egula ion 2009/921, No wegian Animal Wel a e Ac 19/06/2009). This da ase included samples om 14 popula ions s udied by Oze o e al. [9] (Table 1). Simila o ou p e ious wo k [9], equimola (10 ng/ul) DNA ex ac s om 40 o 70 indi iduals we e pooled o p o ide om h ee o six echnical eplica es o each popula ion sample (Table 1, Fig. 2). The pooled DNA samples we e analyzed in he Cen e o In eg a i e Gene ics (CIGENE, No way) using an Illumina in inium assay (Illumina, San Diego, CA, USA) and e sion 2 o he A lan ic salmon SNP-chip [20], [21] ca ying p obes o 5568 SNP ma ke s. Among ou s udy popula ions, 3928 o hose SNP loci we e bi-allelic and we e u he analyzed in his s udy. The aw SNP da a we e analyzed using Geno yping module . 1.9.4 (Genome S udio so wa e . 2011.1, Illumina Inc.). In addi ion, he same 23 popula ions used o DNA pool cons uc ion we e indi idually geno yped using 31 commonly applied [17], [22] STR ma ke s (Table 1, Table S1 in Appendix S1). The STR da a we e analyzed and he geno ypes we e sco ed wi h Genemappe 4.1 so wa e (Applied Biosys ems). Allele F equency Es ima ion Allele equencies o 23 popula ions we e es ima ed om DNA pools compa ing pool-speci ic alue o he a wi h he e e ence alues o he a de i ed om 300 A lan ic salmon specimens geno yped by CIGENE [9]. B ie ly, he aw colo signal da a om 2 al e na e alleles is con e ed in o a he a alue which anges om 0 o 1. In heo y, an indi idual homozygous o an allele would ha e a he a alue close o 0 o 1, and a alue o 0.5 would indica e a he e ozygous geno ype. Howe e , in eali y a SNP’s he a o geno ype clus e s (AA, AB and BB) a y om heo e ical alues o 0, 0.5 and 1. The e o e, o es ima ion o allele equency in a pooled sample, he he a alue o each SNP is compa ed o he mean he a alues o AA, AB and BB geno ypes calcula ed by geno yping o indi idual samples applying co ec ion algo i hm, me hod 2 in [23]. S ingen quali y con ol il e s esul ed in selec ion o 2880 bi- allelic SNPs showing low e o a es ( a ia ion o he a among echnical eplica es #0.02) compa ed o in o ma ion con en [9]. Popula ion-speci ic allele equencies o each SNP we e es ima ed as a mean o e 3–6 echnical eplica es (Appendix S2). Popula ion Rela ionships and wi hin Popula ion Di e si y In o de o in e he gene ic ela ionships among popula ions, pai wise D A [24] dis ances and pai wise F ST we e calcula ed om allele equency es ima es de i ed om allelo yping o 2880 SNPs wi h he Powe Ma ke 3.25 so wa e package [25]. The D A gene ic dis ances we e used o cons uc neighbo -joining ees wi h 1000 boo s ap eplica es. The same app oach was applied o he 31 STR ma ke s. Consensus dend og ams we e cons uc - ed sepa a ely o SNPs and STRs by using he p og am Spli sT ee4 4.11.3 [26]. Simila ly, Powe Ma ke 3.25 [25] was used o es ima e expec ed he e ozygosi y (H E ) o popula ions o e all SNP and STR loci. Selec ion o he mos In o ma i e SNPs We e alua ed ou di e en me hods o iden i ica ion o he mos in o ma i e se o SNPs o GSI: Del a [27], global F ST , pai wise F ST [28] and he ou lie app oach [29]. The es ima e o allele- equency di e en ial, i.e., Del a, is one o he s aigh o wa d ways o e alua e he in o ma ion con en o a SNP. Fo a bi-allelic ma ke , like SNP, he Del a alue is es ima ed as |pA i -pA j |, whe e pA i and pA j a e he equencies o allele A in he i h and j h popula ions, espec i ely. Del a alue o each SNP ma ke was es ima ed as he mean ac oss all pai -wise compa isons o 23 popula ions. Ano he common c i e ion o selec ing he mos in o ma i e loci is he popula ion di e en ia ion measu e F ST : he unbiased es ima es o F ST we e i s calcula ed o e all popula ions (global F ST ) and on a pai wise basis (pai wise F ST ). SNPs in he uppe qua ile o dis ibu ion o di e gence alues we e classi ied as ma ke s ha ing ‘‘high le el’’ o gene ic di e en ia ion (Fig. S1 in Appendix S1). Fo he i s h ee app oaches (i.e., Del a, global F ST and pai wise F ST ), he op 300 unlinked SNPs (.1 cM dis ance om each o he ) we e selec ed o subsequen analyses. Fo iden i ica ion o SNPs de ia ing om he neu al expec a- ions (ou lie s), a Bayesian likelihood me hod was used, imple- men ed in Bayescan 2.01 [29]. The me hod p o ides pos e io odds (PO) as he a io o he pos e io p obabili ies indica ing how much mo e likely he model wi h selec ion is compa ed o he neu al model. Pos e io odd alues be ween 10 and 32 (log 10 (PO) = 1–1.5) a e conside ed as s ong e idence o selec ion, be ween 32 and 100 (log 10 (PO) = 1.5–2) – as e y s ong, and PO abo e 100 (log 10 (PO) .2) a e iewed as decisi e e idence o selec ion [1], [29]. Depending on popula ion da ase , 35–111 unlinked ou lie s (.1 cM dis ance om each o he ) po en ially Finding Ma ke s Tha Make a Di e ence o GSI PLOS ONE | www.plosone.o g 2 Decembe 2013 | Volume 8 | Issue 12 | e82434 in luenced by di e gen selec ion [30] we e iden i ied and used o subsequen analyses. To es ima e how he numbe o SNPs a ec he pe o mance o GSI, subse s o he op 25, 50, 75, 100, 125, 150, 200, 250 and 300 SNP loci we e selec ed o each o he ou anking app oaches. In addi ion, we also chose simila ly-sized subse s o andom SNPs (i.e. 25, 50, 75, 100, 125, 150, 200, 250 and 300 SNP loci). To e alua e he e ec o popula ions on he selec ion o op SNPs we es ed he o e all assignmen powe o he mos in o ma i e SNPs iden i ied using h ee di e en popula ion da ase s. The i s da ase included all 23 popula ions (da ase I); he second se consis ed o 16 popula ions (1–16) excluding he eas e nmos and Bal ic salmon (da ase II); and inally, he hi d se consis ed o six popula ions (1, 3, 5, 11, 13, 15) e enly dis ibu ed ac oss he No wegian and he Wes e n Ba en s seas coas s (da ase III; Table 1). Pe o mance o Top SNPs and STRs o GSI As he me hods applied o GSI equi e geno ype da a a he han allele equency es ima es, he mul ilocus geno ypes we e simula ed om he allele equency es ima es assuming Ha dy- Weinbe g and linkage equilib ium using bespoke so wa e (see Appendix S3 o he code). Fo each subse o SNP loci, 100 mul ilocus geno ypes pe popula ion we e simula ed as a baseline sample. Ano he 500 geno ypes pe popula ion o he mixed s ock ishe y sample we e simula ed o each SNP subse using ONCOR [31]. A simila app oach was applied o he STR da a. The assignmen o indi iduals in a mix u e o baseline popula ions was pe o med by using ONCOR [31]. This app oach es ima es a p obabili y ha an indi idual (o unknown o igin) belongs o a baseline popula ion by assessing es ima es o he geno ype equencies in each baseline popula ion [32] and an es ima e o he s ock composi ion o he ishe y [31]. In ONCOR, he simula ed mixed s ock ishe y sample was es ed agains baseline da a se along wi h he ‘‘Assign indi iduals o he baseline popula ion’’ op ion o assign each ish. All indi iduals we e assigned i espec i e o p ecision. As he so wa e used was speci ically de eloped o analyzing samples o indi iduals he in luence o he dis ibu ion o geno ypes in he sample being es ed was examined using mix u es o ish wi h a ying composi ions. I was ound ha a he le els o di e en ia ion obse ed he e hese composi ions had only mino in luence on he assignmen esul s (see Table S2 in Appendix S1). In addi ion o SNPs, pe o mance was e alua ed o 31 STRs. We also e alua ed he pe o mance o a 31-locus STR panel combined wi h di e en numbe s o SNPs (1, 2, 5, 10, 25, 50 and 100 op- anked loci). We es ima ed he numbe o STR o SNP alleles equi ed o achie e 80%, 90% and 95% co ec assignmen s o each anking app oach and popula ion da ase . We made hese es ima es by i ing a non-linea eg ession model o he cu es o co ec assignmen pe cen age agains cumula i e numbe o ma ke s. An Figu e 1. Map indica ing sampling loca ions o he s udied popula ions. See Table 1 o popula ion names. Eu opean A lan ic salmon samples [20] used o alida ion o op anked SNPs a e indica ed as illed iangles. doi:10.1371/jou nal.pone.0082434.g001 Finding Ma ke s Tha Make a Di e ence o GSI PLOS ONE | www.plosone.o g 3 Decembe 2013 | Volume 8 | Issue 12 | e82434 exponen ial eg ession model (y= exp(a+b/x)) was ound o bes i he da a. Analysis o he Independen Da ase To e alua e he use ulness o ou app oach, and he e ec i e- ness o ou bes SNPs o GSI in di e en se s o popula ions, independen alida ion was pe o med on he A lan ic salmon indi iduals geno yped by Bou e e al. [20], DRYAD en y doi:10.5061/d yad.gm367. Speci ically, we e alua ed he pe o - mance o ou op- anked loci (da ase II, pai wise F ST selec ion app oach, 25 o 150 SNPs) o GSI on 26 Eu opean anad omous popula ions anging om Spain (Na cea) o Russia (Se e naja D ina). Compa ed o ou da a, only h ee popula ions s udied by Bou e e al. [20] o igina ed om he i e sys ems included in bo h da ase s (Tana/Teno, Ponoi, Va zuga). Addi ionally, we es ed he eliabili y o he allelo yping app oach by combining he allele equency es ima es om DNA pools o 23 popula ions wi h he 38-popula ion da ase o Bou e e al. [20] which consis ed o indi idual geno ypes. Resul s Gene ic Di e si y and Di e en ia ion: SNPs s. STRs As expec ed, he gene ic di e si y o SNPs o e all popula ions was signi ican ly lowe compa ed o STRs (median SNP He = 0.36 s. median STR He = 0.77, Mann-Whi ney U- es , P,0.001). The gene ic di e si y (H E ) o popula ions o e all SNP loci anged om 0.23 o 0.35, whe eas o STR da a H E es ima es we e highe , anging om 0.64 o 0.74 (Table 1). Howe e , gene ic di e si y es ima es wi hin popula ions (H E ) we e signi ican ly co ela ed be ween he wo ma ke ypes (Pea son’s = 0.93, P,0.0001). Pai wise popula ion di e en ia ion (F ST ) es ima es o e 2880 SNPs a ied om 0.01 (Ti o ka s. U a) o 0.30 (Na a s. Pecho a Unya), whe eas mean pai wise F ST alues o e 31 STRs anged Table 1. In o ma ion abou popula ions included in he da ase s used o SNP selec ion and hei geog aphic loca ions. Popula ion Coo dina es N SNP N STR H E SNPs H E STRs Popula ion da ase I II III No wegian Sea 1 Laukhelle* 69u13’N 17u50’E 42 (4) 42 0.35 0.73 x x x 2Ma ˚lsel a 69u13’N 18u29’E 70 (3) 70 0.34 0.72 x x 3 Reisa 69u46’N 21u00’E 70 (3) 70 0.31 0.69 x x x 4 Al a* 69u58’N 23u22’E 70 (3) 65 0.32 0.69 x x 5 Reppa jo del * 70u26’N 24u19’E 69 (4) 67 0.35 0.73 x x x Wes e n Ba en s Sea 6 Laksel a* 70u04’N 24u55’E 67 (5) 67 0.32 0.68 x x 7 Iesjoki (Teno/Tana) 69u26’N 24u59’E 70 (3) 70 0.32 0.71 x x 8 Ka asjoki (Teno/Tana)* 69u23’N 25u09’E 70 (4) 63 0.33 0.70 x x 9 Ina ijoki (Teno/Tana)* 69u00’N 25u46’E 67 (4) 67 0.33 0.70 x x 10 Yla ¨ko ¨nga ¨s (Teno/Tana) 69u57’N 26u34’E 58 (3) 58 0.32 0.71 x x x 11 Tana B u (Teno/Tana)* 70u12’N 28u11’E 60 (2) 59 0.34 0.70 x x 12 Ves e Jakobsel * 70u06’N 29u19’E 70 (4) 59 0.33 0.72 x x 13 Neiden* 69u42’N 29u31’E 63 (4) 63 0.33 0.72 x x x 14 Ti o ka* 69u30’N 31u58’E 70 (4) 67 0.35 0.74 x x 15 U a* 69u16’N 32u48’E 44 (3) 44 0.34 0.73 x x x 16 Kola* 68u52’N 33u01’E 70 (6) 70 0.33 0.74 x x Whi e Sea 17 Ponoi* 66u58’N 41u16’E 70 (4) 70 0.33 0.72 x 18 Va zuga* 66u12’N 36u57’E 70 (4) 70 0.31 0.70 x 19 Onega 63u54’N 38u00’E 70 (3) 70 0.25 0.65 x 20 Mezen Pizhma 65u53’N 44u08’E 48 (3) 48 0.28 0.68 x Eas e n Ba en s Sea 21 Pecho a Pizhma 64u52’N 51u16’E 48 (3) 48 0.26 0.67 x 22 Pecho a Unya 61u47’N 57u52’E 48 (3) 48 0.24 0.64 x Bal ic Sea 23 Na a 59u23’N 28u12’E 40 (3) 40 0.23 0.64 x To al pooled (SNPs) 1424 To al indi idual (STRs) 1395 N SNP – numbe o indi iduals included in he pools and numbe o echnical eplica es (in b acke s) o SNP-chip analysis, N STR – numbe o samples o indi idual geno yping using 31 STRs, H E – o e all expec ed he e ozygosi y o SNPs and STRs. *Da a om Oze o e al. [9]. doi:10.1371/jou nal.pone.0082434. 001 Finding Ma ke s Tha Make a Di e ence o GSI PLOS ONE | www.plosone.o g 4 Decembe 2013 | Volume 8 | Issue 12 | e82434 om 0.02 o 0.20 o he same popula ion pai s (Table S3 in Appendix S1). Simila o gene ic di e si y, gene ic di e gence o he popula ions (mean pai wise F ST ) was signi ican ly co ela ed be ween SNPs and STRs (Man el’s xy = 0.94, P,0.0001). Thus, bo h ma ke classes e ealed e y simila popula ion gene ic s uc u ing as illus a ed by neighbo -joining ees (Fig. 3). Howe e , he le el o di e en ia ion o SNPs o e all popula ions was signi ican ly highe han ha o STRs (median SNP Fs = 0.077 s. median STR Fs = 0.055, Mann-Whi ney U- es , P,0.001). Iden i ica ion o he mos In o ma i e SNPs As expec ed, only a small p opo ion o SNPs (ou o 2880) exhibi ed high le els o gene ic di e en ia ion es ima ed using h ee measu es (global F ST , pai wise F ST and Del a, Appendix S4, Fig. S1 in Appendix S1). The ou lie es o popula ion da ase s I, II and III iden i ied pu a i e signs o di e gen selec ion (log 10 (PO) .1, q,0.05) a 141, 120 and 41 SNPs, espec i ely (Fig. 4). Howe e , as se e al SNPs o med igh ly linked g oups (,1 cM), a o al o 111, 95 and 35 unlinked ou lie SNPs we e e ained o subsequen analysis o popula ion da ase s I, II and III, espec i ely. The compa ison o allele equency dis ibu ions o 100 op SNPs anked using Del a o pai wise F ST showed ha hese wo app oaches iden i ied loci wi h a wide ange o allele equencies among popula ions (Fig. 5). A simila pa e n was e iden o ou lie s, wi h a sligh ly highe p opo ion o loci (42%) showing ma ginal allele equency dis ibu ions (close o 0 o 1). In con as , he allele equencies o a majo i y o he SNPs (69%) anked using global F ST measu e we e biased owa ds 0 o 1 wi h ela i ely ew loci exhibi ing in e media e allele equencies. Despi e he di e ences desc ibed abo e, he e was a subs an ial amoun o o e lap be ween he op SNPs among all anking app oaches (Fig. 6). Fo example, 59 o 68 o SNPs ou o he op 100 we e iden i ied using paiwise F ST and Del a in all h ee da ase s. Simila ly, la ge p opo ions o SNPs we e sha ed be ween global F ST and pai wise F ST app oach (51%–71%). In addi ion, a subs an ial p opo ion o ou lie s (up o 58%) was also anked as op SNPs by he h ee o he app oaches (Fig. 6A). In con as o SNP anking app oaches, popula ions in he asce ainmen g oup had a much la ge e ec on anking o he mos in o ma i e SNPs. Fo example, a ela i ely la ge p opo ion o SNPs (up o 80%) anked by global F ST , pai wise F ST and Del a was iden i ied only in a single da ase (Fig. 6B) and a ela i ely la ge p opo ion o ou lie s was unique o each da ase (up o 50%). Pe o mance o Top- anked SNPs o GSI Compa ed o andomly chosen se s o SNPs, he o e all assignmen success was conside ably highe o op- anked loci selec ed using ou di e en app oaches (Fig. 7). O hese, he ou lie me hod iden i ied he bes pe o ming loci, while loci selec ed using he global F ST app oach esul ed in he lowes o e all assignmen accu acy. Howe e , when he numbe o SNP ma ke s eached 100, he o e all assignmen success o loci iden i ied by global F ST , pai wise F ST and Del a was a he simila , bu s ill lowe han o loci iden i ied by he ou lie me hod (Fig. 7, Table S4 in Appendix S1). In o de o achie e 80%, 90% and 95% co ec assignmen o 23 A lan ic salmon popula ions, he ou lie app oach equi ed 15–20% ewe loci han he h ee o he app oaches (Table 2). Fo example, 95% o e all co ec assign- men was achie ed using 94 ou lie SNPs iden i ied using popula ion da ase II, whe eas eaching he same le el o assignmen powe wi h SNPs anked by he h ee o he app oaches equi ed 118 o 125 SNPs (Table 2). Al hough he o e all GSI success was high, he numbe o SNPs equi ed o p o ide simila accu acy a ied in indi idual popula ions. Fo example, o Wes e n Ba en s and No wegian Sea popula ions, app oxima ely 100–150 SNPs we e necessa y o a ain .90% popula ion assignmen success, whe eas 25–50 op SNPs we e enough o a ain simila le el o co ec assignmen in he Eas e n Ba en s, Whi e and Bal ic sea popula ions (Table S5 in Appendix S1). Howe e , o ma ke s anked using popula ion da ase s II and III, i.e., when he mos gene ically dis inc eas e nmos and Bal ic popula ions we e emo ed, he assignmen success o .90% o Wes e n Ba en s and No wegian Sea salmon was achie ed using 75–100 op- anked SNPs (Table S5 in Appendix S1), excep o a g oup o h ee popula ions wi h he lowes gene ic di e gence (pai wise F ST o e 2880 SNPs = 0.016– 0.023; Table S3 in Appendix S1). When compa ing he assignmen powe o a gi en numbe o independen alleles be ween wo ma ke classes, he pe o mance o STRs was lowe han ha o andom and op anked SNPs. Fo example, an 80% o e all assignmen success was a ained wi h 219 independen alleles o STRs (18 loci) while 96 and 47 independen alleles we e su icien o each simila accu acy o andom and op- anked SNPs, espec i ely (Table 2, Fig. 7). On he o he hand, when he assignmen powe was es ima ed o a gi en numbe o loci, STRs pe o med be e han SNPs as less mul i-allelic ma ke s we e needed o each gi en le el o assignmen accu acy compa ed o bi-allelic ma ke s. When e alua ing he assignmen success o indi idual popula ions, 18 STRs (219 independen alleles) we e su icien o assign salmon popula ions om he Bal ic and he Whi e and Eas e n Ba en s seas wi h .99% accu acy, whe eas all 31 STR ma ke s (536 independen alleles) we e equi ed o achie e .90% accu acy o he Wes e n Ba en s and No wegian Sea popula ions. A combi- na ion o 31 STR ma ke s and 25 op- anked SNPs inc eased he o e all assignmen accu acy om 97% o 99% (Fig. S2 in Appendix S1). Figu e 2. Wo k low diag am indica ing he main s eps o he analyses. doi:10.1371/jou nal.pone.0082434.g002 Finding Ma ke s Tha Make a Di e ence o GSI PLOS ONE | www.plosone.o g 5 Decembe 2013 | Volume 8 | Issue 12 | e82434 Valida ion o Top- anked SNP on he Independen Da ase The se o ou op 100 SNPs iden i ied in da ase II using he pai wise F ST selec ion app oach allowed .98% co ec assign- men in 13 ou 26 Eu opean anad omous A lan ic salmon popula ions (Table S6 in Appendix S1). The lowes assignmen accu acy was obse ed in B i ish, Sco ish and I ish popula ions (66%–87%) and was in line wi h he lowe le el o gene ic di e en ia ion among salmon popula ions in his a ea [33]. Simila o ou No h-Eu opean da ase , he numbe o op SNPs equi ed o achie e 90% and 95% o e all co ec assignmen o 26 Eu opean popula ions was conside ably lowe compa ed o andomly chosen SNPs (Table S7 in Appendix S1). Howe e , he assignmen accu acy o bo h op- anked and andom ma ke s eached simila high le els (98%) when o e 150 SNPs we e used (Table S7 in Appendix S1). The cons uc ed neighbo -joining ee consis ing o 2763 SNPs om 61 popula ion da a de i ed by allelo yping and indi idual geno yping demons a ed he compa ibili y o he wo app oaches (Fig. 8): he gene ic ela ionships o he popula ions we e consis en wi h he esul s o Bou e e al. [20]. Howe e , he combined da ase u he e ealed new insigh s in o gene ic ela ionships among popula ions such as sepa a ion be ween no he n and sou he n No wegian popula ions. Discussion This is he i s s udy ha combines high- h oughpu SNP a ays and DNA pooling (i.e., allelo yping) o iden i y he mos powe ul se s o SNPs o gene ic s ock iden i ica ion. We demons a e how SNP a ays and DNA pooling enable a as and cos -e ec i e, ye eliable me hod o iden i ying he mos in o ma i e ma ke s among housands o SNPs om la ge numbe o A lan ic salmon popula ions. In line wi h he p e ious s udies o Figu e 3. Gene ic ela ionships among 23 A lan ic salmon popula ions in no he n Eu ope. Neighbou -joining dend og am is based on Nei’s D A gene ic dis ances es ima ed using (A) 2880 SNPs and (B) 31 STR ma ke s. Dis inc popula ion g oups a e colo ed. The b anches wi h boo s ap alue suppo ,80% a e d awn as dashed. doi:10.1371/jou nal.pone.0082434.g003 Finding Ma ke s Tha Make a Di e ence o GSI PLOS ONE | www.plosone.o g 6 Decembe 2013 | Volume 8 | Issue 12 | e82434 human popula ions [14], he esul s a e mos encou aging o p ojec s in ol ing high-sample h oughpu and low- o medium- mul iplex SNP geno yping as a much smalle numbe o SNPs is needed o accu a e gene ic s ock iden i ica ion compa ed o andomly chosen SNPs. Mo eo e , we demons a e he applica- bili y o ou app oach on a da a se compiled om wo sepa a e SNP geno yping p ojec s and illus a e ans e abili y o he da a ac oss s udies wi hou he need o labo ious s anda diza ion in compa ison wi h e.g. mic osa elli es [22]. Reliabili y o DNA Pooling & Allelo yping Recen ly, we showed ha allelo yping o DNA pools is an e ec i e me hod o eliable allele equency es ima ion (indi idual geno yping s. allelo yping, Pea son’s = 0.992) [9]. To u he alida e he obus ness o allelo yping, we compa ed ou allele equency es ima es de i ed om allelo yping o allele equencies ob ained om indi idual geno yping by Bou e e al. [20] o h ee di e en i e sys ems (Teno/Tana, Ponoi and Va zuga). Despi e di e en indi idual samples, sampling loca ions and sampling yea s we obse ed e y high co ela ion be ween he allele equency es ima es (Pea son’s = 0.952–0.971) indica ing he eliabili y o allelo yping app oach. Mo eo e , he cos s o analysis o DNA pools was abou 15 imes lowe compa ed o he analysis o he same numbe o samples indi idually, i.e. he p ice o genome-wide analysis o DNA pools is a leas an o de o magni ude lowe han indi idual geno yping [11–14]. This s udy also allowed a di ec compa ison o a ious popula ion gene ic pa ame e s de i ed om allelo yping and geno yping o SNPs and STRs, espec i ely. We obse ed highly signi ican co ela ion be ween expec ed he e ozygosi y es ima es o he wo ma ke ypes. Simila ly, he es ima es o pai wise gene ic di e en ia ion demons a ed highly co ela ed pa e ns o SNP and STR ma ke s. These esul s a e in ag eemen wi h he ea lie indings showing high conco dance be ween di e en ma ke ypes [5], [6], [34], [35]. Likewise, he gene ic ela ionships among popula ions in e ed by SNPs and STRs we e simila wi h high boo s ap clus e ing suppo o bo h ma ke classes. Howe e , in con as o o he s udies [20], [36], we obse ed a ine sepa a ion o he Wes e n Ba en s salmon in o wo g oups (Teno and Wes e n Ba en s/No wegian Sea). Fu he mo e, wi h a combined da ase consis ing o 61 popula ions we we e able o con i m he gene ic ela ionships among popula ions o e he whole dis ibu ion ange as well as o e eal no el pa e ns such as clea sepa a ion be ween no he n and sou he n No wegian popula ions. Taken oge he , hese esul s no only demons a e he eliabili y o DNA pooling and allelo yping app oach, bu also illus a e one o he impo an ad an ages o SNPs – good ans e abili y o he SNP da a be ween independen s udies. In con as o STRs, which usually equi e labo ious calib a ion and s anda diza ion o alleles [22], [37], his acili a es e icien compila ion o la ge SNP da ase s om independen s udies c ea ing u he syne gis ic e ec s. Howe e , SNP allele label swi ching may u n ou p oblema ic when SNP da a se s om di e en geno yping pla o ms a e compa ed. De ec ion and Pe o mance o he mos In o ma i e SNPs The e alua ion o ou SNP selec ion app oaches demons a ed ha all o hem subs an ially imp o e he accu acy o GSI compa ed o andomly chosen SNPs. In mos cases, SNPs selec ed using he ou lie app oach showed he highes assignmen accu acy, whe eas he global F ST app oach usually esul ed in lowe co ec assignmen a es. On he o he hand, he di e ences in indi idual assignmen success we e e iden only when he numbe o SNPs was below 100 while all SNP selec ion app oaches enabled accu a e GSI when mo e han 100 SNPs we e used. These esul s a e consis en wi h he ea lie esul s demons a ing supe io pe o mance o ou lie loci o e neu al loci o gene ic s ock iden i ica ion [30], [38], [39]. Simila ly, he ou pe o mance o pai wise SNP selec ion me hods o e he global F ST app oach has been shown in bo h humans [40] and ca le [8]. This is because he global F ST app oach ends o selec o ma ke s ha a e speci ic o he mos dis inc popula ion o g oup o popula ions while he pai wise app oaches allow selec ion o ma ke s wi h high he e ozygosi y and mo e e enly dis ibu ed allele equencies among popula ions, being hus mo e in o ma i e o indi idual assignmen [8], [40], [41]. The e was also a subs an ial amoun o o e lap be ween he op loci among all anking app oaches. Fo example, he o e lap o Figu e 4. Iden i ica ion o ou lie loci using a model-based genome scan app oach. (A) popula ion da ase I; (B) popula ion da ase II; (C) popula ion da ase III. Each SNP locus ( illed ci cle) is ep esen ed by he le el o gene ic di e en ia ion (F ST ) and log 10 (PO) o being unde selec ion. Ou lie loci po en ially unde di e gen selec ion a e inside dashed ec angle. doi:10.1371/jou nal.pone.0082434.g004 Finding Ma ke s Tha Make a Di e ence o GSI PLOS ONE | www.plosone.o g 7 Decembe 2013 | Volume 8 | Issue 12 | e82434 op 100 SNPs selec ed using global F ST , pai wise F ST o Del a anged om 51 o 71%. Mo eo e , a la ge p opo ion (up o 58%) o he op 100 SNPs iden i ied using global F ST , pai wise F ST o Del a showed signs o di e gen selec ion. This is consis en wi h he esul s o Lao e al. [7] showing ha he i e mos in o ma i e SNPs di iding human popula ions om di e en con inen s Figu e 5. Dis ibu ion o allele equencies o he 100 op- anked SNPs. Allele equencies o he 100 op- anked SNPs in 23 popula ions iden i ied using ou selec ion app oaches: (A) global F ST ; (B) pai wise F ST ; (C) Del a, and (D) ou lie . Ho izon al line, g ey ec angle, whiske s, open ci cles, and s a s indica e median, 25 h and 75 h qua iles, non-ou lie ange, ou lie s and ex eme ou lie s, espec i ely. doi:10.1371/jou nal.pone.0082434.g005 Figu e 6. SNP o e lap among di e en anking app oaches and popula ion da ase s. (A) Venn diag ams showing he ex en o o e lap among ou app oaches (global F ST , pai wise F ST ,Del a and ou lie ) o h ee popula ion da ase s. (B) Venn diag ams showing he ex en o o e lap among h ee popula ion da ase s o ou anking app oaches. Fo all SNP anking me hods he op 100 SNPs a e p esen ed, excep o he ou lie app oach whe e 95 and 35 SNPs we e iden i ied as being unde selec ion o da ase II and III, espec i ely. doi:10.1371/jou nal.pone.0082434.g006 Finding Ma ke s Tha Make a Di e ence o GSI PLOS ONE | www.plosone.o g 8 Decembe 2013 | Volume 8 | Issue 12 | e82434 exhibi he signs o local posi i e selec ion. Thus, ou esul s indica e ha despi e he high assignmen powe o non-neu al ma ke s, mo e simple pai wise me hods a e nea ly as e icien in anking he mos in o ma i e loci o GSI. On he o he hand, ou esul s indica e ha he popula ion da ase migh play a la ge ole o iden i ica ion o he mos in o ma i e ma ke s han he SNP selec ion app oach [42], [43]. Fo example, he mos in o ma i e se o loci selec ed using one se o human popula ions ha e been shown o lack powe when applied o ano he se o popula ions [7], [42]. Indeed, ou analysis e ealed ha he assignmen accu acy was mo e a ec ed by popula ions in he asce ainmen g oup used o anking SNPs a he han anking me hod i sel . While he SNPs anked using popula ion da ase I allowed quick disc imina ion o he popula- ions a he la ge geog aphical scale, he ma ke s anked using popula ion da ase s II and III pe o med be e a egional le el. This can be explained by highe assignmen success o he indi iduals om he Wes e n Ba en s Sea and No wegian Sea popula ions. Indeed, he geog aphically emo e Bal ic popula ion and salmon o he Whi e and Eas e n Ba en s seas ha e e y dis inc gene ic p o iles ha allow hei disc imina ion using as ew as 25–50 andomly chosen SNPs. On he o he hand, gene ic s uc u e o salmon popula ions o he Wes e n Ba en s and No wegian seas is mo e sub le [18], [36] and GSI o hese popula ions equi es a highe numbe o ma ke s. Thus, he exclusion o he eas e nmos salmon popula ions enabled de ec ion o loci in o ma i e o he Wes e n Ba en s Sea and No wegian Sea salmon inc easing he assignmen accu acy o hese popula- ions. Gi en he esul s o ou and p e ious [7] s udies, pai wise- based me hods o anking he mos in o ma i e ma ke s o GSI a e mo e obus o he asce ainmen bias and hus a e mo e applicable wi h ex ended da ase s. I mus be ecognized ha ou app oach o pe o m powe analyses o he op anked SNPs o gene ic s ock iden i ica ion yields o e ly op imis ic accu acies, see [44]. Ou accu acy le els a e upwa dly biased o wo easons: i) we simula ed geno ypes Figu e 7. O e all assignmen success o SNPs and STRs in da ase I. SNPs we e anked using i) global F ST (blue), ii) pai wise F ST (b own), iii) Del a (g een) and i ) ou lie app oach ( ed). O e all assignmen success o STRs and andomly chosen SNPs a e shown as g ay and black lines, espec i ely. The ba s a e ep esen ing s anda d de ia ion o assignmen success among all popula ions o each anking app oach. The s anda d de ia ion ba s we e a anged o isual pu poses o a oid o e lapping. doi:10.1371/jou nal.pone.0082434.g007 Table 2. Es ima ed numbe o independen alleles o SNPs and STRs equi ed o achie e 80%, 90%, and 95% o e all co ec assignmen in 23 A lan ic salmon popula ions o each anking me hod. Popula ion da ase used o SNP anking I (23) II (16) III (6) Global F ST Pai wise F ST Del a Ou lie Global F ST Pai wise F ST Del a Ou lie Global F ST Pai wise F ST Del a Ou lie Random SNPs STRs* 80% 76 68 59 50 57 50 53 47 62 53 58 n/a 96 219 (,18) 90% 113 104 100 81 90 82 87 71 98 87 95 n/a 149 336 (,24) 95% 167 153 147 114 124 118 125 94 133 123 136 n/a 198 460 (,28) *Es ima ed numbe o STR loci is indica ed in pa en heses. doi:10.1371/jou nal.pone.0082434. 002 Finding Ma ke s Tha Make a Di e ence o GSI PLOS ONE | www.plosone.o g 9 Decembe 2013 | Volume 8 | Issue 12 | e82434