METHODOLOGY ARTICLE Open Access
Fish scales and SNP chips: SNP geno yping and
allele equency es ima ion in indi idual and
pooled DNA om his o ical samples o A lan ic
salmon (Salmo sala )
Susan E Johns on
1,5*
, Me i Lindq is
1
, Ee o Niemelä
2
, Panu O ell
2
, Jaakko E kina o
2
, Ma hew P Ken
3
,
Sigbjø n Lien
3
, Juha-Pekka Vähä
1
, An i Vasemägi
1,4
and C aig R P imme
1
Abs ac
Backg ound: DNA ex ac ed om his o ical samples is an impo an esou ce o unde s anding gene ic
consequences o an h opogenic in luences and long- e m en i onmen al change. Howe e , such samples gene ally
yield DNA o a lowe amoun and quali y, and he ex en o which DNA deg ada ion a ec s SNP geno yping
success and allele equency es ima ion is no well unde s ood. We conduc ed high densi y SNP geno yping and
allele equency es ima ion in bo h indi idual DNA samples and pooled DNA samples ex ac ed om d ied A lan ic
salmon (Salmo sala ) scales s o ed a oom empe a u e o up o 35 yea s, and assessed geno yping success,
epea abili y and accu acy o allele equency es ima ion using a high densi y SNP geno yping a ay.
Resul s: In indi idual DNA samples, geno yping success and epea abili y was e y high (> 0.973 and > 0.998,
espec i ely) in samples s o ed o up o 35 yea s; bo h inc eased wi h he p opo ion o DNA o agmen
size > 1000 bp. In pooled DNA samples, allele equency es ima ion was highly epea able (Repea abili y = 0.986)
and highly co ela ed wi h empi ical allele equency measu es (Mean Adjus ed R
2
= 0.991); allele equency could
be accu a ely es ima ed in > 95% o pooled DNA samples wi h a e e ence g oup o a leas 30 indi iduals. SNPs
loca ed in polyploid egions o he genome we e mo e sensi i e o DNA deg ada ion: olde samples had lowe
geno yping success a hese loci, and a la ge e e ence panel o indi iduals was equi ed o accu a ely es ima e
allele equencies.
Conclusions: SNP geno yping was highly success ul in deg aded DNA samples, pa ing he way o he use o
deg aded samples in SNP geno yping p ojec s. DNA pooling p o ides he po en ial o la ge scale popula ion
gene ic s udies wi h ewe assays, p o ided enough e e ence indi iduals a e also geno yped and DNA quali y is
p ope ly assessed be o ehand. We p o ide ecommenda ions o u u e s udies in ending o conduc
high- h oughpu SNP geno yping and allele equency es ima ion in his o ical samples.
Keywo ds: A lan ic salmon, SNP geno yping, Illumina® iSelec SNP-a ay, Deg aded DNA, A chi ed samples,
Fish scales, DNA pooling, Allelo yping, Allele equency, F agmen size
* Co espondence: [email p o ec ed]
1
Depa men o Biology, Uni e si y o Tu ku, Tu ku FIN-20014, Finland
5
Ins i u e o E olu iona y Biology, Uni e si y o Edinbu gh, Edinbu gh EH9
3JT, Uni ed Kingdom
Full lis o au ho in o ma ion is a ailable a he end o he a icle
© 2013 Johns on e al.; licensee BioMed Cen al L d. This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e
Commons A ibu ion License (h p://c ea i ecommons.o g/licenses/by/2.0), which pe mi s un es ic ed use, dis ibu ion, and
ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly ci ed.
Johns on e al. BMC Genomics 2013, 14:439
h p://www.biomedcen al.com/1471-2164/14/439
Backg ound
His o ical a chi ed samples a e an impo an esou ce
o popula ion gene ic moni o ing, as hey allow us o
unde s and he impac o an h opogenic in luences and
en i onmen al change in he wild [1]. Viable DNA ob ained
om his o ical ma e ial such as museum specimens, scales,
ea he s, hai and/o bones [2] has been used o examine
pas popula ion s uc u e [3,4] and i s pe sis ence o e ime
[5,6], popula ion collapses [7] and bo lenecks [8], ounde
e en s [9] and he consequences o s ocking o popula ions
wi h non-na i e indi iduals [10]. Howe e , a c ucial limi a-
ion o his o ical samples is ha hey gene ally yield DNA
o a lowe amoun /quali y han is ecommended o mo-
lecula gene ic s udies. This is due o a numbe o ac o s,
including: DNA deg ada ion o e long- e m s o age in sub-
op imal condi ions; sample age; he sample quali y a ini ial
sampling; he ela i e DNA concen a ion wi hin he sam-
ple; and he educed e iciency o DNA ex ac ion p o ocols
on he sample ype. All o hese ac o s can lead o a isk o
PCR ailu e, geno yping e o s and allelic d op-ou [11,12],
as well as a ailu e o ul il ecommenda ions ega ding he
amoun and quali y o DNA used o geno yping.
Wi h he ad en o nex gene a ion sequencing and
cos -e icien geno yping echnology, single-nucleo ide
polymo phisms (SNPs) a e an inc easingly popula molecu-
la ma ke in gene ic and e olu iona y esea ch. They occu
a highe equencies h oughou he genomes o a wide
ange o species [13,14] and ha e he po en ial o iden i y
unc ionally impo an polymo phisms [15,16]. Compa ed
o using mo e adi ional ma ke s, such as mic osa elli es,
SNP geno yping is as e , mo e cos -e icien and less
e o -p one when conside ing he assessmen o hou-
sands, a he han ens o loci [17,18]. This is because
i can be ca ied ou using low densi y a ays [11,19] and/o
high h oughpu chips [20]. Fu he mo e, i is possible o
es ima e popula ion-wide allele equencies in pooled DNA
samples wi hin a single SNP a ay, meaning ha popula ion
gene ic s udies and ou lie analyses could be conduc ed
using a conside ably educed numbe o assays [21-24].
Consequen ly, SNP ma ke s ha e an eno mous po en ial o
add ess a numbe o ou s anding ques ions in e olu iona y
ecology, conse a ion gene ics and wildli e managemen
[25-27] and indeed, encou aging examples o s udies
u ilising SNPs in his o ical samples a e eme ging in a
numbe o species [4,28-30].
A p esen , he e a e wo p inciple echnologies a ailable
o geno yping housands o SNP loci simul aneously
in cus om a ays (also known as ‘SNP chips’); Illumina
In inium (San Diego, Cali o nia, USA) [31] and A yme ix
Axiom (San a Cla a, Cali o nia, USA) [32]. The ins umen-
a ion and a ay cons uc ion o he wo echnologies a e
e y di e en , bu he basic p inciples o he assay chemis-
y a e simila , and bo h sys ems clus e geno ypes based
on he in ensi y o he signal and he con as be ween he
signals om he wo alleles o he SNP. Bo h pla o ms o e
some ad an ages o geno yping his o ical samples; in
pa icula , he leng h o he DNA agmen size equi ed o
SNP geno yping on bo h pla o ms is small (e.g. 25-100 bp)
compa ed o he leng h o mos mic osa elli e ma ke s
(> 100 bp); indeed SNP geno yping can be mo e success ul
han mic osa elli e geno yping in his o ical samples [33].
Howe e , he e is also po en ial o some bias as a esul
o DNA deg ada ion: p o ocols o bo h a ays equi e
whole-genome ampli ica ion o DNA samples p io o
geno yping [31,34], a p ocess which is sensi i e o DNA
deg ada ion [35]. In addi ion, s udies assessing he use ul-
ness o pooled DNA samples o allele equency es ima ion
on ei he pla o m ha e only used high quali y DNA
[21-24]. The e o e, i emains impe a i e ha de ailed cha -
ac e isa ion o he ex en o DNA deg ada ion and s ingen
pilo es ing o SNP yping on ei he pla o m in his o ical
samples is ca ied ou be o e subsequen scien i ic conclu-
sions and managemen decisions a e made [11].
In his s udy, we es ed he e icacy o SNP geno yping
and allele equency es ima ion in his o ical samples
om A lan ic salmon (Salmo sala ). The economic
impo ance o bo h a med and wild A lan ic salmon
has led o he de elopmen o a high h oughpu cus-
om Illumina® iSelec SNP-a ay, which consis s o 5568
SNPs h oughou i s genome [36]. In addi ion, he e has
been wide-sp ead sampling o A lan ic salmon scales
and/o o holi hs o age de e mina ion and popula ion
moni o ing pu poses [12]. Bo h esou ces p o ide an
excellen ounda ion o empo al gene ic s udies exam-
ining he e olu iona y impac o an h opogenic e ec s
on wild s ocks, bu equi e alida ion ha DNA ex ac ed
om a chi ed ma e ial be o e i can be used eliably wi h
he de eloped a ay. We es ed he eliabili y o SNP geno-
yping in his o ical samples using DNA ex ac ed om
scale samples ha had been collec ed om wild adul
A lan ic salmon o e a hi y yea pe iod. We add ess wo
impo an ques ions in he use o SNP geno yping in
his o ical samples: i s , how call a e and epea abili y o
SNP geno yping a ies wi h sample age and DNA quali y
(using Da ase 1, see below); and second, wha is he
accu acy o allele equency es ima ion in DNA pools
c ea ed using his o ical samples (using Da ase 2). We
hen discuss he applica ion o ou indings and p o ide
ecommenda ions o u u e SNP geno yping s udies in
his o ical samples.
Resul s
Scale samples we e geno yped a 5568 SNP loci using a
modi ied e sion o he cus om-designed Illumina® iSelec
SNP-a ay desc ibed p e iously [36,37] and indi idual
geno ypes we e sco ed using he clus e ing algo i hm
implemen ed in he Illumina® GenomeS udio Geno yping
Analysis Module 2011.1. As a esul o his o ical genome
Johns on e al. BMC Genomics 2013, 14:439 Page 2 o 13
h p://www.biomedcen al.com/1471-2164/14/439
duplica ion in salmonids [38], some SNP ma ke s show
polyploidy and ha e been classi ied as mul i-si e a ian s
[37]. The e o e, we e ained 5317 loci alling wi hin h ee
ca ego ies: ‘SNP’(seg ega es as a no mal diploid SNP,
N = 3928), ‘MSV-3′(whe e a SNP exis s on a single
pa alogue; N = 873) and ‘Mono’(N = 516, whe e loci in
Lien e al. 2011 we e monomo phic).
Da ase 1: Tempo al a ia ion o DNA quali y and
geno yping success in a chi ed scale samples
DNA was ex ac ed om a chi ed scales selec ed om
ish cap u ed in he same i e ibu a y in he yea s
1976, 1987, 1996 and 2006 (4 ish pe yea , wi h wo in-
dependen ex ac ions pe ish and wo independen
geno yping uns pe ex ac ed sample) and no malised
o 50 ng/μl. The ela i e concen a ions o DNA o di -
e en agmen size anges we e hen de e mined o
each ex ac ion.
The e was no ela ionship be ween yea and o al DNA
concen a ion a e no malisa ion (ρ= 0.242, P = 0.182,
Figu e 1A), bu he p opo ion o DNA consis ing o
la ge agmen sizes (> 1000 bp) inc eased wi h yea
(ρ= 0.772, P < 0.001, Figu e 1B). A o al o 4102 loci
(including 341 MSV-3 and 377 Mono loci) passed
isual inspec ion and quali y con ol, accoun ing o 86.2%,
39.1% and 73.1% o SNP, MSV-3 and Mono loci, espec-
i ely; 3238 o hese loci we e polymo phic, wi h a mino
allele equency (MAF) o > 0.05. Ac oss all yea s, geno yp-
ing a e and he p opo ion o epea able geno ypes pe
indi idual we e e y high (> 0.973 and > 0.998 ac oss all
samples, espec i ely), wi h a mean GenCall sco e (a mea-
su e o he eliabili y he geno ype call based on i s posi ion
ela i e o he cen e o he geno ype clus e ) o > 0.666.
The excep ion came om samples o one ish sampled
in 1976 (ID Ss_1976_011), wi h alues o > 0.848, 0.989
and > 0.530, espec i ely. Sample call a e also inc eased wi h
yea (ρ= 0.739, P < 0.001, Figu e 2A) and wi h he p opo -
ion o DNA wi h a agmen size o > 1000 bp (ρ= 0.739,
P < 0.001, Figu e 2B). The ull esul s o all s a is ical com-
pa isons a e gi en in Addi ional ile 1 and he aw da a used
o conduc he analysis is p o ided in Addi ional ile 2.
Da ase 2: Accu acy o allele equency es ima ion
(allelo yping) in pooled DNA samples
DNA was ex ac ed om a chi ed scales om 530 ish
cap u ed in a la ge i e sys em be ween 2001 and 2003
and no malised o 100 ng/μl. Indi iduals we e assigned
o one o ou DNA pools (N = 87–161 pe pool), and
geno yping was hen ca ied ou on all indi idual and
pooled DNA samples. The ‘empi ical’allele equency o
each locus wi h each pool was es ima ed om he indi-
idual geno ypes o i s cons i uen indi iduals, and he
‘es ima ed’allele equency o each locus in each pooled
DNA sample was de e mined by i s allelic in ensi y a io
(The a) ela i e o he mean he a o each geno ype
(de e mined om all indi idually yped samples; (Addi ional
ile 1: Figu e S1).
Compa ison o empi ical and es ima ed allele equencies
A o al o 4642 loci (including 663 MSV-3 and 460
Mono loci) in 514 indi iduals passed isual inspec ion
and quali y con ol, accoun ing o 89.6%, 75.9% and
89.1% o loci ca ego ised as SNP, MSV-3 and Mono,
espec i ely; 3732 o hese loci we e polymo phic
(MAF > 0.05). The co ela ion be ween he empi ical and
es ima ed allele equencies o all alid loci was e y high
ac oss all pools (N = 36; mean adjus ed R
2
=0.991, SE=
2.121 × 10
-4
; Figu e 3) and he mean di e ence be ween
he empi ical and es ima ed allele equencies a e aged o
each locus was 0.0253 (SE = 2.665 × 10
-4
; Figu e 4). Highe
Figu e 1 Tempo al a ia ion o DNA quali y in a chi ed scale samples. We show ela ionships be ween A. To al DNA concen a ion and
Yea , and B. P opo ion o high molecula weigh DNA (> 1000 bp) and yea . Each poin indica es an indi idual DNA ex ac ion. Samples had
been no malised o 50 ng/μl based on NanoD op Spec opho ome e be o e o al DNA concen a ion and agmen sizes we e measu ed using
an Agilen 2100 Bioanalyze .
Johns on e al. BMC Genomics 2013, 14:439 Page 3 o 13
h p://www.biomedcen al.com/1471-2164/14/439
MAF and lowe GenT ain sco es (i.e. a measu e o clus e-
ing e iciency calcula ed in GenomeS udio) we e associa ed
wi h la ge di e ences be ween he empi ical and es ima ed
alues, espec i ely (Gene al linea model, P < 0.001, Table 1)
and loci classi ied as MSV-3 and Mono had signi ican ly
smalle di e ences be ween empi ical and es ima ed e-
quencies compa ed o SNP loci (P < 0.001). Es ima ed allele
equencies we e highly epea able ac oss all he eplica ed
measu es (Analysis-wide epea abili y = 0.986, 95% c edible
in e al = 0.986 - 0.987).
Figu e 2 Tempo al a ia ion o geno yping success in a chi ed scale samples. We show ela ionships be ween A. Sample Call Ra es and
Yea , and B. Sample call a e and p opo ion o DNA > 1000 bp. Each poin indica es an indi idual geno yping un.
Figu e 3 Co ela ion be ween mean es ima ed allele equencies and empi ical allele equencies wi hin each pool. Means we e calcula ed
om nine eplica es wi hin each pool. R2 is he adjus ed R
2
alues om a linea eg ession. N is he numbe o indi iduals included in each pool.
Johns on e al. BMC Genomics 2013, 14:439 Page 4 o 13
h p://www.biomedcen al.com/1471-2164/14/439
Es ima ion o allele equencies om sampled subse s o
indi iduals o e e ence geno ypes
Allele equencies in DNA pools we e e-es ima ed using
mean geno ype clus e posi ions calcula ed om smalle
subse s o cons i uen indi iduals, anging om 10 o
200 indi iduals sampled om he ull da ase . The mean
p opo ion o pool allele equencies ha could be es i-
ma ed inc eased om 0.895 when es ima ed om N = 10
indi iduals, o 0.988 when es ima ed om N = 200 indi-
iduals; a p opo ion o > 0.95 was obse ed when sam-
pling 30 indi iduals o mo e (Figu e 5A). The mean
adjus ed R
2
be ween he empi ical and es ima ed e-
quencies was high o all subse s (> 0.985) and inc eased
wi h he numbe o indi iduals sampled (Figu e 5B).
The mean di e ence be ween he es ima ed and empi ical
allele equencies dec eased as he numbe o sampled in-
di iduals inc eased (Figu e 5C). Fo all h ee es ima es
ca ied ou , he mean alues om each subse size we e
signi ican ly highe om he p e ious ca ego y as he sam-
ple size inc eased ( wo sample - es P < 0.001; Figu e 5).
The mean esul s ob ained o each es ima e a e gi en in
Addi ional ile 1: Table S4.
Re-clus e ing, geno yping and allele equency es ima ion
using subse s o ep esen a i e indi iduals
The en i e analysis o pooled samples in Da ase 2 was
epea ed ( om GenomeS udio clus e ing, geno ype de-
e mina ion and isual examina ion, and allele equency
es ima ion ela i e o geno ype clus e s) using wo single
subse s o 20 and 50 indi iduals as e e ence indi iduals.
Findings we e compa ed o he empi ical alues de e -
mined om he ull Da ase 2 (N = 514). In he N = 20
da ase , 4863 loci passed isual inspec ion and quali y
Figu e 4 His og am o he mean di e ence be ween he
empi ical and es ima ed allele equencies o each locus. The
e ical do ed line indica es he mean o he dis ibu ion (0.0253).
Table 1 Resul s o a gene al linea model es ing ac o s a ec ing o allele equency es ima ion
Pooled indi idual numbe Model e ms Pa ame e es ima e S.E. - alue P- alue
514 Indi iduals Mino Allele F equency 0.0211 0.00180 11.75 < 0.001
(N
LOCI
=4642) GenT ain Sco e −0.0226 0.00196 −11.49 < 0.001
Locus Classi ica ion:
Mono −0.0149 0.00094 −15.74 < 0.001
MSV-3 −0.0045 0.00102 −4.401 < 0.001
20 indi iduals Mino Allele F equency 0.0199 0.00352 5.65 < 0.001
(N
LOCI
=4571) GenT ain Sco e −0.0227 0.00409 −5.55 < 0.001
Locus Classi ica ion:
Mono −0.0150 0.00185 −8.14 < 0.001
MSV-3 0.0063 0.00186 3.37 < 0.001
50 indi iduals Mino Allele F equency 0.0093 0.00346 2.68 0.00735
(N
LOCI
=4681) GenT ain Sco e −0.0302 0.00407 −7.43 < 0.001
Locus Classi ica ion:
Mono −0.0145 0.00182 −7.96 < 0.001
MSV-3 −0.0035 0.00196 −1.79 0.0730
The e ec s o mino allele equency, GenT ain sco e and locus classi ica ion on he accu acy o allele equency es ima es we e es ed based on he mean
di e ence be ween empi ical and es ima ed allele equencies pe locus o all indi iduals. Es ima ed allele equencies we e es ima ed om clus e s o all
indi iduals (N = 514) and andom subse s o 20 and 50 indi iduals. The mean di e ence in allele equency was used as a dependen a iable wi h a Gaussian
e o s uc u e. S.E. is he s anda d e o , - alue is he s a is ic alue and P- alue is he co esponding signi icance o he model e m.
Johns on e al. BMC Genomics 2013, 14:439 Page 5 o 13
h p://www.biomedcen al.com/1471-2164/14/439
Figu e 5 Boxplo demons a ing he accu acy o allele equency es ima ion using subse s o e e ence indi iduals. Pa ame e
es ima ion was ca ied ou 100 imes o each subse . A. The p opo ion o pooled samples o which equency es ima es could be calcula ed.
B. The mean adjus ed R
2
o e all pools and all loci o each simula ion. C. The mean di e ence be ween empi ical and es ima ed allele
equencies o e all pools and loci o each simula ion.
Johns on e al. BMC Genomics 2013, 14:439 Page 6 o 13
h p://www.biomedcen al.com/1471-2164/14/439
con ol, accoun ing o 94.0%, 83.8% and 85.3% o loci
ca ego ised as SNP, MSV-3 and Mono, espec i ely;
3738 loci we e polymo phic (MAF > 0.05). A o al o 111
loci had a leas one geno ype misma ch when compa ed
o he ull da ase (N = 514). In e es ingly, 67 o hese
loci misma ched a > 95% o geno ypes, 64 o which had
been classi ied as MSV-3 and Mono loci, indica ing ha
inco ec clus e posi ioning may be mo e acu e in du-
plica ed egions when using a smalle numbe o e e -
ence indi iduals. The e was a high co ela ion be ween
he empi ical and mean es ima ed allele equencies pe
pool (N
LOCI
= 4571, mean adjus ed R
2
= 0.978 and SE =
1.806 × 10
-4
; Figu e 6A), and he mean di e ence be-
ween he empi ical and es ima ed allele equencies a -
e aged o each locus was 0.0287 (SE = 4.856 × 10
-4
).
In he N = 50 da ase , 4897 loci passed isual inspec-
ion and quali y con ol, accoun ing o 95.9%, 77.0%
and 89.0% o loci ca ego ised as SNP, MSV-3 and Mono,
espec i ely; 3882 loci we e polymo phic (MAF > 0.05).
O hese, 59 loci had a leas one geno ype misma ch
wi h he ull da ase (N = 514), wi h 35 loci misma ching
a >95% o geno ypes (33 o hese classi ied as ‘MSV3′
o ‘Mono’). The inc ease in sample size esul ed in a
highe co ela ion be ween he empi ical and mean es i-
ma ed allele equencies pe pool (N
LOCI
= 4681; mean
adjus ed R
2
= 0.980, SE = 2.198 × 10
-4
; Figu e 6B) and
lowe mean di e ence be ween he empi ical and es i-
ma ed allele equencies a e aged o each locus was
0.0267 (SE = 4.785 × 10
-4
). In bo h he N = 20 and N = 50
da ase s, highe mino allele equencies and lowe
GenT ain sco es we e associa ed wi h la ge di e ences
be ween he empi ical and es ima ed alues, espec i ely
(Gene al linea model, P < 0.001, Table 1); howe e , in
he N = 20 da ase , MSV-3 loci had a signi ican ly la ge
di e ences in be ween he empi ical and es ima ed allele
equencies (P = 7.51 × 10
-4
).
Discussion
DNA agmen p o iling in a chi ed samples (Da ase 1)
A e no malisa ion, he o al concen a ion o DNA
ex ac ed om ai -d ied scale samples did no a y o e
ime, bu he Bioanalyze DNA concen a ions we e con-
sis en ly much lowe han he concen a ion es ima ed
by he NanoD op Spec opho ome e (a mean o
14.76 μl compa ed o he expec ed 50 ng/μl), indica ing
he NanoD op me hod may o e -es ima e he DNA con-
cen a ion in deg aded DNA samples. The e we e
ma ked di e ences in he p opo ion o DNA comp ised
o smalle agmen sizes be ween he wo ea lie sam-
pling pe iods (1976 and 1987) compa ed o he la e
sampling pe iods (1996 onwa ds; Figu e 1).
Geno yping in a chi ed scale samples (Da ase 1)
The geno yping call a e and e o a es in he 1987
samples we e compa able o he 1996 and 2006 samples,
ye he agmen sizes in he 1987 samples we e mo e
simila o he 1976 samples. The sample call a e and
e o a e we e mo e a iable in he 1976 samples, wi h
one sample showing a pa icula ly low call a e and
highe geno ype misma ch a e. In a la ge scale s udy,
samples such as his one could be disca ded du ing qual-
i y con ol, bu i is impo an o no e ha e en in his
sample, 3432 o he 4102 loci assessed ga e sco eable
and epea able geno ypes. Fu he mo e, he emaining
samples om 1976 (call a es > 97.3% and misma ch
a es o less han 0.2%) show ha eliable geno yping is
possible in samples up o 35 yea s old in his case. The e
was a signi ican end o call- a e and geno ype mis-
ma ch o be co ela ed (Addi ional ile 1: Figu e S6), in-
dica ing ha lowe call a es a e a use ul guide o
excluding po en ially un eliable samples. Howe e , in
ou da ase , his end is s ongly in luenced by he sin-
gle, poo ly pe o ming sample, hus i would be
Figu e 6 Co ela ion be ween empi ical and es ima ed allele equencies calcula ed om 20 (le ) and 50 ( igh ) indi iduals. Each poin
ep esen s he mean allele equency calcula ed om nine eplica es wi hin each o he ou pools (= ou poin s pe locus). Empi ical allele
equencies we e de e mined om he ull da ase (N = 514). R2 is he adjus ed R
2
alue om a linea eg ession ( ed line). NB. No e ha some
poin s a e a emo ed om he eg ession line; hese a e cases whe e clus e s in duplica ed egions o he genome ha e been placed inco ec ly
due o he small numbe o e e ence indi iduals.
Johns on e al. BMC Genomics 2013, 14:439 Page 7 o 13
h p://www.biomedcen al.com/1471-2164/14/439
ecommendable o u u e esea ch o ocus on a mo e
de ailed in es iga ion o he obus ness o his end in a
la ge numbe o olde samples. I is no clea why he e
a e di e ences in geno yping e iciency be ween 1976
and 1987 samples when hei agmen size p opo ions
we e highly simila (Addi ional ile 1: Figu e S2). I may
be ha addi ional chemical, physical o biological ac o s
o he han agmen size a ec geno yping success, pos-
sibly h ough inhibi ion o PCR o di ec e ec s on he
s uc u e o he DNA.
Geno yping success in polyploid loci (Da ase s 1 and 2)
The success o geno yping diploid SNP loci was high,
wi h mo e han 86% o loci p o iding sco eable geno-
ypes and allele equency es ima es in bo h da ase s.
Howe e , in Da ase 1, less han 40% o duplica ed
MSV-3 loci passed isual inspec ion, whe e indi idual
clus e s o each geno ype could no be de e mined. As
hese loci ha e a smalle ange o he a (~0.5, compa ed
o 1 in ’SNP’loci), poo ly de ined clus e s a e mo e likely
o o e lap, leading o a la ge numbe unde ined geno-
ypes o indi idual samples. This p oblem is likely o
ha e been mo e acu e in Da ase 1 o se e al easons.
Fi s , al hough 64 samples we e geno yped in Da ase 1,
hese only comp ised o 16 indi iduals. This is likely o
ha e esul ed in a educed chance o sampling a e ge-
no ypes and alleles, and exace ba ed he e ec o lowe
quali y samples on clus e ing. Second, in compa ison o
Da ase 2, samples we e up o 27 yea s olde and he e-
o e mo e likely o ha e a highe deg ee o deg ada ion.
In summa y, ou da a show ha al hough diploid SNP
geno yping is e icien in his o ical samples, geno yping
in duplica ed egions is likely o be mo e sensi i e o
DNA deg ada ion and should be ea ed wi h mo e cau-
ion in quali y con ol and s udy design.
DNA pooling and allele equency es ima ion (Da ase 2)
The high co ela ion and epea abili y be ween he es i-
ma ed and empi ical allele equencies in pooled DNA
samples indica e ha DNA pooling, e en om a chi ed
scale ma e ial, is a cos -e ec i e solu ion o es ima ing
sample and popula ion-wide allele equencies. Accu a e
es ima ion o allele equency may be in luenced by he
indi idual DNA quali y and concen a ion wi hin pools,
especially in cases whe e he amoun o DNA om pa -
icula indi iduals a e unde - o o e - ep esen ed wi hin
he sample [22]. Ou da a indica e ha his e ec is
coun e ac ed by no malisa ion measu es; i is also pos-
sible ha using a la ge numbe o indi iduals pe pool
(>87 indi iduals in ou case) will educe he in luence o
a pa icula indi iduals on allele equency es ima ion
[23]. The e was mo e a ia ion be ween obse ed and
expec ed DNA equencies as he mino allele equency
inc eased in all da ase s, showing ha allele equency
es ima ion may be mo e accu a e a loci wi h lowe
mino allele equencies.
As i is unlikely ha s udies implemen ing a pooled
DNA s a egy o allele equency es ima ion will geno-
ype all cons i uen indi iduals, he e a e se e al impo -
an conside a ions when using subse s o samples o
de e mine clus e posi ions o allele equency es ima-
ion. When sampling smalle subse s o indi iduals om
he ull da ase , each inc ease o 10 indi iduals signi i-
can ly imp o ed he accu acy o allele equency es ima-
ion and educing he amoun o s ochas ic a ia ion
a ec ing he es ima es, al hough he deg ee o change
dec eases wi h each inc ease in sample size. Sampling
geno ypes om jus 30 indi iduals mean ha mo e han
95% o pooled samples could be es ima ed, wi h a mean
di e ence be ween he es ima ed and ue empi ical al-
lele equencies o jus 0.003 highe han ha ob ained
om he ull da ase (Addi ional ile 1: Table S4). The e-
o e, al hough we ecommend ha la ge numbe s o
e e ence indi iduals a e used, u u e s udies can con-
side he ade-o s be ween he numbe o indi iduals
and he numbe o pooled samples ha can be yped,
gi en pa icula inancial limi s o geno yping and he
biological ques ion being conside ed.
Clus e posi ioning using subse s o e e ence indi iduals
(Da ase 2)
When we epea ed he analysis o Da ase 2 using
smalle subse s o e e ence indi iduals o clus e ing
(N = 20 and N = 50), we ound ha using a lowe num-
be o ep esen a i e samples o he ull analysis is likely
o in oduce some deg ee o inaccu acy in bo h clus e
posi ioning and allele equency es ima ion. Fo example,
mo e han 200 loci ha ailed quali y con ol in he ull
da ase bu we e passed in he N = 20 and N = 50
da ase s. This may be because p oblema ic loci (such as
hose which a e polymo phic on bo h pa alogues [37])
a e mo e easily de ec ed in he ull da ase , bu may ap-
pea o seg ega e as no mal biallelic loci when clus e ing
wi h a smalle numbe o indi iduals. Fu he mo e, in
he N = 20 da ase , some clus e posi ions in polyploid
egions (MSV-3 and Mono) we e placed inco ec ly,
leading o la ge disc epancies be ween he empi ical and
es ima ed allele equencies a a hand ul o loci (see de-
sc ip ion o Figu e 6). Clus e ing so wa e, such as
Illumina GenomeS udio which is used in he cu en
s udy, will au oma ically clus e loci as i hey we e dip-
loid SNPs, and so 20 indi iduals may no be enough o
isually de e mine whe he o no a locus is an MSV-3,
pa icula ly i no all geno ypes a a locus a e ep e-
sen ed. This issue disappea ed when he numbe o indi-
iduals used o c ea e he clus e s inc eased o 50,
al hough some allele equencies in e ed be ween he
obse ed and expec ed alues (i.e. an es ima e close o 0
Johns on e al. BMC Genomics 2013, 14:439 Page 8 o 13
h p://www.biomedcen al.com/1471-2164/14/439
becomes close o 1, and ice e sa); his is likely o ha e
a isen in cases whe e a di e en allele o a locus is ixed
(o almos ixed) in he di e en pa alogues. I should
also be no ed ha a e y small p opo ion o ‘SNP’loci
we e misplaced (only 3 ou o 3928 in he N = 20 da ase .
The e o e, in species wi h duplica ed genomes, a la ge
ep esen a i e sample may be equi ed o inc ease bo h
SNP geno yping a e and he accu acy o allele equency
es ima ion.
Conclusions
Al hough his s udy has ocussed on a single species on
a speci ic geno yping pla o m, he me hods ha we ha e
used a e applicable o o he diploid species, and we ha e
also o e ed solu ions o species wi h pa ially dupli-
ca ed genomes. O e all, we ound ha SNP geno yping
was highly success ul in A lan ic salmon scales up o
24 yea s old, wi h high geno yping success and low e o
a es. Fu he mo e, we demons a ed ha allele equen-
cies could be accu a ely es ima ed in pooled DNA
ob ained om 8–10 yea old scales. Ou indings open
up a new ange o oppo uni ies o high h oughpu
SNP analyses using a chi ed ma e ial, and u he in es-
iga ions o e en olde samples may be wo hwhile.
Howe e , we ha e also shown ha i is impe a i e o
ca y ou su icien es ing o bo h DNA condi ion and
SNP geno yping e iciency and o apply s ic quali y
con ol o samples be o e emba king on la ge scale SNP
geno yping s udies, and ha ca e should be aken o se-
lec an app op ia e numbe ep esen a i e indi iduals o
de ec and emo e p oblema ic loci and imp o e he ac-
cu acy o allele equency es ima ion. Building on p e i-
ous ecommenda ions o gene ic analysis in his o ical
ish scales [12], we make he ollowing addi ional ecom-
menda ions o SNP geno yping in his o ical samples:
DNA ex ac ion and quali y checking
1. Pilo es ing should be conduc ed o iden i y sample
con amina ion and o de e mine app op ia e
h esholds o sample inclusion o SNP geno yping
and/o DNA pooling.
2. E ec s o sample age on DNA agmen a ion and
geno yping success should be eliably assessed.
DNA Pooling
3. Indi idual samples included in DNA pools should be
ex ac ed and no malised using he same me hods as
all o he indi iduals wi hin he pool.
4. The po en ial e ec s o inclusion o DNA samples o
a ying quali y (e.g. o di e en ages) in pooled DNA
samples on accu a e allele equency es ima ion
should be conside ed.
5. Pools con aining la ge numbe s o indi iduals a e
ecommended in o de o educe he e ec s o
a ia ion be ween indi iduals on allele equency
es ima ion [23].
Clus e ing o SNP geno yping and/o allele equency
es ima ion
6. Geno yping by clus e ing should be conduc ed
independen ly wi hin each geno yping s udy o
accoun o a ia ion in DNA quan i y and quali y.
Visual inspec ion is ecommended o ensu e ha
clus e ing is accu a e.
7. Re e ence indi iduals should come om he same
popula ion/samples as he pooled indi iduals.
8. Indi idual samples should be geno yped o c ea e
in o ma i e clus e s o accu a e allele equency
es ima ion. In ou s udy, we ound ha geno yping
a leas 30 indi iduals could p o ide su icien
clus e ing accu acy o diploid SNP loci.
9. Species wi h polyploid genomes may equi e a la ge
numbe o e e ence indi iduals o allele equency
es ima ion, o ensu e accu acy in clus e posi ioning.
10.Samples o indi idual and pooled DNA should be
andomised du ing SNP geno yping o minimise any
po en ial biases om pla e posi ion.
11.The es ima ion o allele equencies om DNA pools
should e alua e he e o associa ed wi h
allelo yping using indi idually geno yped da a.
Me hods
Sample collec ion
A lan ic salmon scale samples we e collec ed om wild
adul ish cap u ed wi hin he Teno i e sys em in
No he n Eu ope (No wegian: Tana; 68-70ºN, 25-27ºE)
by local ishe men be ween 1976 and 2006. Scale sam-
ples we e ai -d ied and s o ed in pape en elopes, which
we e a chi ed by he Finnish Game and Fishe ies Ins i u e
a oom empe a u e and humidi y in an ai -condi ioned
acili y as pa o a long- e m ishe ies moni o ing
p ojec [39].
Da ase desc ip ions, DNA Ex ac ion and DNA
quan i ica ion
Da ase 1: Geno yping in a chi ed scale samples
A chi ed scales we e selec ed o 16 ish cap u ed in
Ke ojoki, a ibu a y wi hin he Teno i e sys em, om
he yea s 1976, 1987, 1996 and 2006 (4 ish pe yea ).
Fo each indi idual, wo independen DNA ex ac ions
we e ca ied ou on ou o i e scales using a
NucleoSpin® Tissue ki (Mache ey-Nagel GmbH, Dü en,
Ge many) in he yea 2011. DNA samples we e ini ially
no malised o 50 ng/μl using NanoD op ND-1000 Spec-
opho ome e (The mo Scien i ic). To de e mine he
Johns on e al. BMC Genomics 2013, 14:439 Page 9 o 13
h p://www.biomedcen al.com/1471-2164/14/439