In . J. Mol. Sci. 2013, 14, 5694-5711; doi:10.3390/ijms14035694
In e na ional Jou nal o
Molecula Sciences
ISSN 1422-0067
www.mdpi.com/jou nal/ijms
A icle
De elopmen and Valida ion o Single Nucleo ide Polymo phisms
(SNPs) Ma ke s om Two T ansc ip ome 454-Runs o Tu bo
(Scoph halmus maximus) Using High-Th oughpu Geno yping
Manuel Ve a 1,2,*, Jose-An onio Al a ez-Dios 3, Ca los Fe nandez 2, Ca men Bouza 2,
Roman Vilas 2 and Paulino Ma inez 2
1 Labo a o y o Gene ics Ich hyology, Depa men o Biology, Facul y o Sciences,
Uni e si y o Gi ona, Campus o Mon ili i s/n, Gi ona 17071, Spain
2 Depa men o Gene ics, Facul y o Ve e ina y, Uni e si y o San iago de Compos ela,
Campus o Lugo, Lugo 27002, Spain; E-Mails: ca [email p o ec ed] (C.F.);
[email p o ec ed] (C.B.); om[email p o ec ed] (R.V.); [email p o ec ed] (P.M.)
3 Depa men o Applied Ma hema ics, Facul y o Ma hema ics, Uni e si y o San iago de Compos ela,
San iago de Compos ela 15782, Spain; E-Mail: josean onio.al a [email protected]
* Au ho o whom co espondence should be add essed; E-Mail: m[email p o ec ed];
Tel.: +34-972-418-168; Fax: +34-972-418-277.
Recei ed: 3 Decembe 2012; in e ised o m: 17 Feb ua y 2013 / Accep ed: 22 Feb ua y 2013 /
Published: 12 Ma ch 2013
Abs ac : The u bo (Scoph halmus maximus) is a comme cially aluable la ish and one
o he mos p omising aquacul u e species in Eu ope. Two ansc ip ome 454-py osequencing
uns we e used in o de o de ec Single Nucleo ide Polymo phisms (SNPs) in genes
ela ed o immune esponse and gonad di e en ia ion. A o al o 866 ue SNPs we e
de ec ed in 140 di e en con igs ep esen ing 262,093 bp as a whole. Only one ue SNP
was analyzed in each con ig. One hund ed and hi een SNPs ou o he 140 analyzed
we e easible (geno yped), while Ш we e polymo phic in a wild popula ion.
T ansi ion/ ans e sion a io (1.354) was simila o ha obse ed in o he ish s udies.
Unbiased gene di e si y (He) es ima es anged om 0.060 o 0.510 (mean = 0.351),
minimum allele equency (MAF) om 0.030 o 0.500 (mean = 0.259) and all loci we e in
Ha dy-Weinbe g equilib ium a e Bon e oni co ec ion. A la ge numbe o SNPs (49)
we e loca ed in he coding egion, 33 ep esen ing synonymous and 16 non-synonymous
changes. Mos SNP-con aining genes we e ela ed o immune esponse and gonad
di e en ia ion p ocesses, and could be candida es o unc ional changes leading o
OPEN ACCESS
In . J. Mol. Sci. 2013, 14 5695
pheno ypic changes. These ma ke s will be use ul o popula ion sc eening o look o
adap i e a ia ion in wild and domes ic u bo .
Keywo ds: u bo ; Scoph halmus maximus; SNP alida ion; EST da abase; non-synonymous
subs i u ion; high- h oughpu geno yping
1. In oduc ion
The u bo (Scoph halmus maximus; Scoph halmidae, Pleu onec i o mes) is a comme cially
aluable la ish ha has been in ensi ely cul u ed since he 1980s. I s p oduc ion has s eadily
inc eased up o he p esen igu e o 8549 ons in 2011 (91.2% Eu opean p oduc ion om Spain; [1])
and i appea s o be one o he mos p omising aquacul u e species in Eu ope. In esponse o u bo
indus y demands, gene ic ma ke s ha e been de eloped in his species in o de o e alua e gene ic
esou ces in bo h wild and ha che y popula ions and pe o m pa en age analysis o suppo gene ic
b eeding p og ams [2–4]. These ma ke s ha e also been applied o de elop genomic ools o iden i y
genomic egions associa ed wi h p oduc i e cha ac e s [5–7] and o de ec selec ion oo p in s in wild
popula ions [8]. Inc easing g ow h a e, con olling sex a io ( emales la gely ou g ow males) and
enhancing disease esis ance cu en ly cons i u e he main goals o gene ic b eeding p og ams in
his species.
The necessi y o unde s anding he immune esponse o pa hogens o indus ial ele ance and o
iden i y genes in ol ed in he sex di e en ia ion pa hway led us o inc ease genomic esou ces in
u bo . As a consequence he eo , an Exp essed Sequence Tag (EST) da abase om cDNA lib a ies o
he main immune issues was cons uc ed using Sange sequencing [9]. Recen ly, his da abase has
been ampli ied wi h wo 454 FLX uns [10,11] (454-Li e Sciences, B and o d, CT, USA; o
454- echnique me hodology see [12,13]). Nex Gene a ion Sequencing (NGS) echnologies o e he
abili y o p oduce an eno mous olume o da a wi h a e y low sequencing cos pe base [12]. Thus,
his u bo EST da abase is cu en ly composed o ~70,000 unique sequences (~20,000 con igs and
~50,000 single ons). ESTs a e essen ial o asce ain he gene [14,15], bu also o iden i y polymo phic
gene-associa ed ma ke s, such as mic osa elli es and single nucleo ide polymo phisms (SNPs) ( ype I
ma ke s; [9,16–18]). Type I ma ke s a e e y use ul o cons uc ing gene ic o physical maps, and o
compa a i e mapping [7,19,20].
SNPs ha e se e al ad an ages o e o he ma ke s when i comes o mapping genes o in e ing
popula ion s uc u e [21]. They can be easily e alua ed in silico o public da abases and hei geno ypes
quickly assessed by mini-sequencing eac ions [9,22] o by high- h oughpu echnologies [23,24]. SNP
alleles a e almos exclusi ely iden ical-by-descen (IBD) and hus hey p e en sco ing e o s
associa ed o homoplasy [25]. They a e ex emely s able, due o low mu a ion a es [26], and occu
mo e o en in he genome han o he ma ke s. In he human genome, o ins ance, he e is on a e age
1 SNP pe 300 bp [27], and hei equency in non-model species has been es ima ed a ~1 in 200–500
bases o non-coding DNA and ~1 in 500–1000 bases o coding DNA [28]. In u bo , Ve a e al. [29]
es ima ed 1 ue SNP e e y ~100 bp om he EST da abase composed only o Sange sequences,
sugges ing he exis ence o la ge SNP esou ces in his species. Du ing he las decade, SNP disco e y
In . J. Mol. Sci. 2013, 14 5696
pipelines ha e been de eloped o aquacul u e species including ish [18,30–35], shell ish [36–38] and
c us aceans [39,40]. In u bo , a SNP calling ool was included in he u bo da abase [9] and i has
been e ined in he upda ed e sion [11]. In his s udy, we sc eened genomic esou ces a ailable in an
upda ed e sion o he u bo EST da abase using con igs con aining NGS 454-sequences o iden i y
and cha ac e ize SNPs associa ed o immune- and ep oduc ion- ela ed genes. These ma ke s will be
used o u he s uc u al genomic analysis ocused on quan i a i e ai loci (QTLs) linked o p oduc i e
ai s, as well as o popula ion sc eening o look o adap i e a ia ion in wild and domes ic u bo .
2. Resul s and Discussion
2.1. Da abase Exploi a ion and SNP De ec ion
The main cha ac e is ics o he u bo 454- ansc ip ome sequencing uns ha e been desc ibed in
p e ious s udies [10,11]. The used da abase ( e sion 4.0 Sep embe 2011) was cons i u ed by
71,033 unique sequences, 18,880 con igs and 52,153 single ons including 454-sequences and Sange
sequences [9] wi h a o al leng h o 52,402,177 base pai s (bp, ~52 Mb). Howe e , in o de o a oid
duplica es wi h he p e ious SNPs de eloped om sequences ob ained wi h Sange me hodology [29],
and since we we e mainly in e es ed in alida ing SNPs a new immune- and ep oduc ion- ela ed
genes, only con igs composed exclusi ely o a leas ou 454-sequences we e used o SNP de ec ion.
Thus, 140 con igs om he u bo da abase, which me hese equi emen s, we e aken in o accoun o
he SNP de elopmen . The o al leng h analyzed was 262,093 bp and con ig leng h anged om 728 bp
o 4885 bp, wi h a mean leng h alue o 1872.09 ± 746.69 bp. The o al numbe o ue SNPs de ec ed
using he p og am Quali ySNP ( o ue SNP de ini ion see he expe imen al sec ion) was 866, SNP
numbe pe con ig anged om 1 o 58, wi h a mean alue o 6.18 ± 8.34. Thus, he expec ed
equency o SNP appea ance in he analyzed sequences would be 1 SNP e e y 302 bp. This alue is
lowe han ha p e iously epo ed in S. maximus (1 SNP each ~100 bp; [29]), bu simila o hose
desc ibed in non-model species [28]. The success o any geno yping me hod is e lec ed in wha is
e e ed o as he con e sion a e and he global success a e. The o me only conside s he
polymo phic ma ke s, whe eas he la e conside s all he ma ke s (monomo phic and polymo phic)
ha we e success ully yped wi hin he analyzed samples [41]. O he 140 ue SNPs es ed,
27 (19.3%) could no be geno yped, and hus hey we e conside ed o be geno yping ailu es due o
echnical and/o geno yping p oblems. Only 2 ou o he 113 easible SNPs (see de ini ion in he
expe imen al sec ion) we e monomo phic. The e o e, he global success a e and con e sion a e we e
80.7% and 79.3%, espec i ely. Global success a e was e y simila o ha p e iously desc ibed in
he species (78.4%), bu con e sion a e was much highe han p e iously epo ed using sequences
om cDNA lib a ies (37.7%; see Ve a e al. [29]), likely due o he di e en lib a y cons uc ion me hods
and bioin o ma ic pipeline app oaches ollowed in 454 and Sange con igs (see expe imen al sec ion).
2.2. SNP Pe o mance
A o al o 65 ansi ions (A/G and C/T) and 48 ans e sions (A/C, A/T, C/G and G/T) we e
de ec ed among easible SNPs, A/G being he mos common (35) and A/C he leas common (6)
subs i u ions obse ed (Figu e 1). This ep esen ed a ansi ion/ ans e sion ( s/ ) a io o 1.354. This
In . J. Mol. Sci. 2013, 14 5697
a io was lowe han ha obse ed by Ve a e al. [29] (1.885) and in silico (1.456) by Pa do e al. [9], bu i
was e y simila o ha desc ibed o common ca p (Cyp inus ca pio) (1.310) [42] and gil head
seab eam (Spa us au a a) (1.375) [31]. Also, he mos equen ansi ions and ans e sions di e ed
om p e ious epo s: C/T and G/T, espec i ely [29], and A/G and A/C [9]. These disc epancies
could be due o he opposi e sequencing di ec ions, as all sequences by Ve a e al. [29] and Pa do e al. [9]
we e ob ained om he 3' end using cDNA lib a ies, while hose om he 454- un we e andomly
ob ained by agmen a ion o he whole cDNA acco ding o he cDNA apid lib a y p epa a ion
me hod (Roche Fa ma, S. A. [43]). Mo eo e , he longe coding egion po ion analyzed in 454- uns
ega ding Sange sequencing in ou s udy may de e mine di e ences because o he di e en selec i e
cons ain s o UTR ega ding coding egions. No di e ences we e de ec ed among dis ibu ion o he
a ian s be ween es ed SNPs and easible SNPs (χ2 = 0.3115; p = 0.9974). All polymo phic SNP loci
showed wo alleles and all o hem ag eed wi h hose expec ed om he da abase in o ma ion.
Figu e 1. Dis ibu ion o SNP a ian s analyzed in his s udy (a) using all SNPs es ed;
(b) using only easible SNPs. T ansi ions ( s) and ans e sions ( ) a e indica ed in black
and g ey colou , espec i ely.
(a)
(b)
0
20
40
60
A/G C/T A/C A/T C/G G/T
F equency
T ansi ions/T ans e sions
SNP a ian s
All
0
10
20
30
40
A/G C/T A/C A/T C/G G/T
F equency
T ansi ions/T ans e sions
SNP a ian s
Feasible
In . J. Mol. Sci. 2013, 14 5698
2.3. SNP Di e si y
Only wo loci among he 113 easible SNPs we e monomo phic (SmaSNP_287 and SmaSNP_334).
Among polymo phic SNPs, unbiased gene di e si y (He) es ima es anged om 0.060 a
SmaSNP_237, SmaSNP_245 and SmaSNP_305 o 0.510 a SmaSNP_225 wi h a mean alue o
0.344 ± 0.149. The minimum allele equency (MAF) in he polymo phic ma ke s anged om 0.030
(SmaSNP_237, SmaSNP_245 and SmaSNP_305) o 0.500 in SmaSNP_249 wi h a mean alue o
0.259 ± 0.140. Depa u es om Ha dy-Weinbe g equilib ium (HWE) we e de ec ed in i e ma ke s
(SmaSNP_253, SmaSNP_271, SmaSNP279, SmaSNP_289, SmaSNP_326; Table 1), al hough all
ma ke s we e a equilib ium a e Bon e oni co ec ion (p = 0.0004). The samples om he
Can ab ian u bo popula ion we e globally in acco dance wi h HWE expec a ions when es ed
simul aneously o all loci (p = 0.9999). These polymo phic alues we e in he ange o hose
p e iously desc ibed in he species [29], and hey we e also simila o hose epo ed in o he ish
species [42,44]. No Linkage disequilib ium (LD) was de ec ed among he 6328 loci pai s a e
Bon e oni co ec ion (p = 0.0004).
Table 1. Anno a ion, a ian s and di e si y alues o he 113 echnically easible SNPs in
he Can ab ic u bo popula ion (33 indi iduals) used in his s udy.
SNP Name Anno a ion Va ian s MAF P (HW) He Fis
SmaSNP_211 Cyclin-dependen kinase 2 in e ac ing p o ein A/T A = 0.152 0.1307 0.262 0.307
SmaSNP_212
Zona pellucida spe m-binding p o ein 3
A/G A = 0.152 0.4198 0.265 0.179
SmaSNP_215 Mi o ic speci ic cyclin-B1 C/T T = 0.212 0.2948 0.338 −0.255
SmaSNP_216 P e-mRNA b anch si e p o ein p14 A/T T = 0.348 0.7003 0.460 −0.119
SmaSNP_217 Zona pellucida p o ein C1 A/G G = 0.258 0.1616 0.390 0.301
SmaSNP_218 Mi ochond ial ibosomal p o ein S18A G/T T = 0.333 1.0000 0.452 0.061
SmaSNP_219 U3 small nucleola ibonucleop o ein p o ein IMP3 G/T G = 0.409 1.0000 0.491 −0.050
SmaSNP_220 Coa ome subuni epsilon iso o m 1 C/T T = 0.197 0.5750 0.322 0.153
SmaSNP_222 Signal ecogni ion pa icle 14 kDa p o ein G/T G = 0.203 1.0000 0.329 −0.046
SmaSNP_223 Epi helial cell adhesion p o ein A/T T = 0.333 1.0000 0.452 0.061
SmaSNP_224 T ansc ip ion ini ia ion ac o TFIID subuni D11 C/G G = 0.182 1.0000 0.302 −0.003
SmaSNP_225 Acidic ibosomal p o ein P1 A/G G = 0.480 1.0000 0.510 0.059
SmaSNP_226 Alcohol dehyd ogenase Class-3 C/T T = 0.288 0.6913 0.416 −0.093
SmaSNP_227 Thio edoxin p o ein 4A A/G A = 0.242 1.0000 0.373 0.025
SmaSNP_228 No el p o ein simila o e eb a e THAP domain
con aining 4 (THAP4)
A/G G = 0.212 0.6068 0.340 0.109
SmaSNP_229 Tumo supp esso candida e 2 A/G A = 0.031 1.0000 0.061 −0.016
SmaSNP_230 Op ic a ophy 3 p o ein C/T T = 0.266 0.6477 0.397 0.135
SmaSNP_231 RNA 3'- e minal phospha e cyclase A/C C = 0.266 0.6475 0.397 0.135
SmaSNP_232 RAD1 homolog A/G A = 0.438 0.4921 0.501 0.127
SmaSNP_233 Ubiqui in ca ie p o ein G/T T = 0.409 1.0000 0.491 −0.050
SmaSNP_234 ch oma in accessibili y complex p o ein 1 A/G G = 0.047 0.0504 0.092 0.659
SmaSNP_235 Nucleola p o ein 16 A/G G = 0.258 0.4023 0.389 0.144
SmaSNP_236 Isopen enyl-diphospha e del a-isome ase 1 C/G G = 0.141 0.4763 0.246 0.111
SmaSNP_237 Ran-speci ic GTPase-ac i a ing p o ein G/T T = 0.030 1.0000 0.060 −0.016
In . J. Mol. Sci. 2013, 14 5699
Table 1. Con .
SNP Name Anno a ion Va ian s MAF P (HW) He Fis
SmaSNP_238 Fo khead box H1 A/G A = 0.453 1.0000 0.503 −0.056
SmaSNP_239 S a hmin C/T C = 0.258 0.6436 0.387 −0.174
SmaSNP_240 Ubiquinol-cy och ome c educ ase co e I p o ein C/T C = 0.152 0.5521 0.261 0.072
SmaSNP_241 BolA-like p o ein 3 C/G G = 0.313 0.4371 0.438 0.143
SmaSNP_243 ce ce oid-lipo uscinosis neu onal p o ein 5 G/T G = 0.455 1.0000 0.504 0.038
SmaSNP_244 SSU RNA; Pse a maxima ( u bo ) C/T C = 0.061 1.0000 0.116 −0.049
SmaSNP_245 Ch omobox p o ein homolog 3 G/T T = 0.030 1.0000 0.060 −0.016
SmaSNP_246 T ansmemb ane p o ein 208 A/C A = 0.469 0.7198 0.505 −0.114
SmaSNP_247 Ribosomal p o ein L18a A/C A = 0.234 1.0000 0.365 0.058
SmaSNP_248 P e-mRNA-p ocessing ac o 19 C/T C = 0.318 1.0000 0.440 −0.032
SmaSNP_249 Alpha-L- ucosidase A/G A = 0.500 0.7275 0.509 0.106
SmaSNP_250 P o ein phospha ase 2 (Fo me ly 2A) A/G G = 0.406 0.0598 0.493 0.366
SmaSNP_252 LON pep idase N- e minal domain and RING
inge p o ein 1
G/T T = 0.167 1.0000 0.282 0.034
SmaSNP_253 IK cy okine A/G A = 0.439 0.0348 0.497 −0.402
SmaSNP_256 Ribonuclease UK114 C/T C = 0.232 0.6038 0.364 0.116
SmaSNP_257 Inne cen ome e p o ein A/G G = 0.303 0.4239 0.430 0.154
SmaSNP_259 Be a-galac oside-binding lec in C/T C = 0.379 0.7242 0.477 −0.079
SmaSNP_260 Enoyl-Coenzyme A hyd a ase A/T A = 0.273 0.3819 0.402 −0.208
SmaSNP_261 Sep 2 p o ein A/G G = 0.197 1.0000 0.321 −0.038
SmaSNP_262 DNA-di ec ed RNA polyme ase I subuni RPA34 A/G A = 0.078 1.0000 0.146 −0.069
SmaSNP_263 Epi helial memb ane p o ein 2 A/G G = 0.379 0.1336 0.480 0.306
SmaSNP_264 Re inol dehyd ogenase 3 C/G G = 0.409 1.0000 0.491 −0.050
SmaSNP_265 WD epea -con aining p o ein 54 A/G A = 0.076 1.0000 0.142 −0.067
SmaSNP_266 RNA pseudou idine syn hase 3 C/T C = 0.136 0.4637 0.240 0.115
SmaSNP_267 T ansmemb ane p o ein 167 p ecu so G/T G = 0.258 0.6463 0.387 −0.174
SmaSNP_270 Flo illin-1 C/G G = 0.438 0.1694 0.502 0.253
SmaSNP_271 NAD(P)H dehyd ogenase quinone 1 A/G G = 0.359 0.0488 0.471 0.403
SmaSNP_273 Ubiqui in p o ein ligase E3 componen C/T C = 0.484 0.7353 0.508 0.077
SmaSNP_274 K13213 ma in 3 C/G G = 0.106 0.2983 0.193 0.216
SmaSNP_275 Dolichol-phospha e mannosyl ans e ase A/G A = 0.091 0.2209 0.169 0.281
SmaSNP_276 DNA-di ec ed RNA polyme ases i II and III
subuni pabc1
A/G A = 0.197 0.5728 0.322 0.153
SmaSNP_277 Syndecan 2 A/C A = 0.429 1.0000 0.505 −0.130
SmaSNP_278 Pep ide me hionine sul oxide educ ase C/G C = 0.078 1.0000 0.146 −0.069
SmaSNP_279 Me hyl ans e ase-like p o ein 21D G/T G = 0.470 0.0129 0.509 0.465
SmaSNP_281 Phospha idylinosi ol ans e p o ein be a
iso o m-like iso o m 2
C/T T = 0.318 0.4333 0.439 −0.172
SmaSNP_282 Zona pellucida p o ein C A/T T = 0.121 1.0000 0.216 −0.123
SmaSNP_283 AP-2 complex subuni alpha-2-like G/T G = 0.333 0.2669 0.450 −0.213
SmaSNP_284 Apop osis egula o BAX A/G G = 0.409 0.0780 0.493 0.324
SmaSNP_285 Bo ealin G/T T = 0.032 1.0000 0.063 −0.017
SmaSNP_286 B ain p o ein 44 C/T C = 0.394 0.2691 0.483 −0.255
SmaSNP_287 Exosome componen 8 A/G G =1.000 - 0.000 NA
In . J. Mol. Sci. 2013, 14 5700
Table 1. Con .
SNP Name Anno a ion Va ian s MAF P (HW) He Fis
SmaSNP_288 A ophin-1 domain con aining p o ein G/T T = 0.439 1.0000 0.500 −0.030
SmaSNP_289 simila o connec in/ i in A/T A = 0.303 0.0018 0.433 0.580
SmaSNP_290 Ubiqui in ca boxyl- e minal hyd olase L5 C/T T = 0.469 1.0000 0.506 0.012
SmaSNP_292 His one deace ylase complex subuni SAP18 C/T C = 0.188 0.5568 0.308 −0.216
SmaSNP_293 Replica ion p o ein A 14 kDa subuni C/G G = 0.182 1.0000 0.302 −0.003
SmaSNP_296 Ca bonic anhyd ase G/T G = 0.076 1.0000 0.142 −0.067
SmaSNP_297 UPF0414 ansmemb ane p o ein C/T C = 0.212 0.6080 0.340 0.109
SmaSNP_298 Queuine RNA- ibosyl ans e ase C/T T = 0.061 1.0000 0.116 −0.049
SmaSNP_299 NHP2-like p o ein 1 C/G C = 0.379 0.1358 0.480 0.306
SmaSNP_304 Mic osomal glu a hione S- ans e ase 3 A/G A = 0.091 1.0000 0.168 −0.085
SmaSNP_305 Ac in ela ed p o ein 2/3 complex subuni 4 C/T T = 0.030 1.0000 0.060 −0.016
SmaSNP_306 Cyclophilin B C/G C = 0.061 1.0000 0.116 −0.049
SmaSNP_307 Dynein ligh chain Tc ex- ype 3 C/G C = 0.061 1.0000 0.116 −0.049
SmaSNP_308 Ependymin-1 A/G A = 0.234 0.3135 0.366 0.231
SmaSNP_309 C-4 me hyls e ol oxidase A/G A = 0.297 1.0000 0.424 0.043
SmaSNP_310 Dynein ligh chain LC8- ype G/T T = 0.045 1.0000 0.088 −0.032
SmaSNP_311 Rho- ela ed GTP-binding p o ein RhoF A/T T = 0.394 0.2669 0.483 −0.255
SmaSNP_312 Golgi SNAP ecep o complex membe 1 A/T A = 0.188 0.5587 0.308 −0.216
SmaSNP_314 Ribosomal L1 domain-con aining p o ein 1 A/G A = 0.203 1.0000 0.329 −0.046
SmaSNP_315 N-alpha-ace yl ans e ase 50 A/T A = 0.242 1.0000 0.373 0.025
SmaSNP_316 Oncogene DJ-1 iso o m 1 C/T C = 0.453 1.0000 0.503 −0.056
SmaSNP_317 Wu: j40d12 p o ein n = 7 Tax = Eu eleos omi
RepID = A3KP21_DANRE
A/G A = 0.438 1.0000 0.500 0.000
SmaSNP_318 Mucin mul i-domain p o ein C/G C = 0.167 0.5617 0.281 −0.185
SmaSNP_319 Adenosine kinase A/G A = 0.182 0.5575 0.301 −0.208
SmaSNP_320 No homology ound A/G A = 0.394 0.4901 0.486 0.127
SmaSNP_321 Zymogen g anule memb ane p o ein 16 A/G G = 0.333 1.0000 0.451 −0.076
SmaSNP_322 6-Py u oyl e ahyd obiop e in syn hase C/T C = 0.031 1.0000 0.061 −0.016
SmaSNP_323 P o easome subuni be a C/T T = 0.125 1.0000 0.222 −0.127
SmaSNP_324 RING inge p o ein 4 A/G A = 0.394 0.0652 0.488 0.379
SmaSNP_325 Lipocalin C/G C = 0.136 1.0000 0.239 −0.143
SmaSNP_326 Choline anspo e -like p o ein 2 A/G G = 0.455 0.0311 0.507 0.402
SmaSNP_328 RNA-binding p o eins (RRM domain) C/T C = 0.106 1.0000 0.192 −0.103
SmaSNP_329 Type II ke a in C/G G = 0.061 1.0000 0.116 −0.049
SmaSNP_330 No el p o ein simila o e eb a e hy oid
ho mone ecep o in e ac o 12 (TRIP12)
C/T T = 0.094 1.0000 0.172 −0.088
SmaSNP_332 Ribosomal p o ein S6 kinase A/C A = 0.470 0.7287 0.507 0.103
SmaSNP_333 T ansmemb ane 6 supe amily membe 2 A/T T = 0.288 0.0796 0.419 0.348
SmaSNP_334 PREDICTED: hypo he ical p o ein
LOC100712283 [O eoch omis nilo icus]
C/T T =1.000 - 0.000 NA
SmaSNP_337 1-Alkyl-2-ace ylglyce ophosphocholine es e ase C/T C = 0.234 1.0000 0.365 0.058
SmaSNP_338 CD151 an igen C/T T = 0.266 0.3909 0.395 −0.186
SmaSNP_339 A seni e me hyl ans e ase 1 A/T A = 0.313 1.0000 0.436 −0.002
SmaSNP_340 Recep o exp ession-enhancing p o ein 5 C/T T = 0.234 0.6507 0.364 −0.116
In . J. Mol. Sci. 2013, 14 5701
Table 1. Con .
SNP Name Anno a ion Va ian s MAF P (HW) He Fis
SmaSNP_341 Ca hepsin S C/G G = 0.333 0.1119 0.454 0.332
SmaSNP_342 T ans-1,2-dihyd obenzene-1,2-diol dehyd ogenase A/G A = 0.424 0.2818 0.494 −0.226
SmaSNP_343 High mobili y g oup p o ein 2 G/T G = 0.470 0.2980 0.508 0.224
SmaSNP_346 ATP-binding casse e, sub- amily A (ABC1) C/T T = 0.288 1.0000 0.417 0.055
SmaSNP_347 Myomesin 1a (skelemin) C/T T = 0.091 1.0000 0.168 −0.085
SmaSNP_348 Re inoic acid ecep o esponde p o ein 3 A/G G = 0.439 1.0000 0.500 −0.030
SmaSNP_349 Nucleophosmin 1 A/C A = 0.258 0.6466 0.387 −0.174
2.4. SNP Posi ion wi hin Genes: Synonymous s. Non-Synonymous Subs i u ions
Consensus sequences o con igs con aining polymo phic SNPs we e compa ed using NCBI BLAST
wi h public da abases, namely UniRe 90, NCBI’s n , KEGG, COG, PFAM, LSU and SSU. The
subsequen BLAST ou pu was hen pa sed wi h Au o FACT [45]. All con igs con aining easible
SNPs we e anno a ed (excep SmaSNP_320, Table 1). The in o ma i e s and, eading ame, and s op
codon a each con ig we e eco ded using homology wi h he highes homologous anno a ed sequence
in public da abases. Nine easible SNPs (8.0%) could no be posi ioned, because no consis en eading
ames we e de ec ed (indica ed as “unknown” loca ion on Table 2). Fi y- i e SNPs (48.7%) we e
loca ed in un ansla ed egions (UTR), ei he in he 5' UTR (17, 15.0%) o 3' UTR (38, 33.6%), which
is in acco dance wi h he app oxima ely double leng h o 3' compa ed o 5' UTR [9]. On he o he
hand, 49 SNPs (43.4%) we e localized in he coding egion (Table 2), a pe cen age o SNPs highe
han p e iously epo ed in he species (24.7%, [29]) and in o he aquacul u e ish species (e.g.,
A lan ic salmon 24%, [32]; A lan ic cod 17.4%, [34]). All hese s udies ollowed a 3' UTR Sange
sequencing s a egy, and he e o e he coding egion was less ep esen ed han in he case o he 454
Roche uns a e a cDNA apid lib a y p epa a ion p o ocol, which accoun s o he di e ences
obse ed. This esul shows he u ili y o he NGS me hodologies o SNP de ec ion in he coding
egion. Thi y- h ee (29.2%) o hese 49 SNPs we e synonymous, and 16 (14.2%) we e
non-synonymous. On he o he hand, he ela ionship be ween synonymous s. non-synonymous
changes (2:1) was lowe han in o he species [46,47]. E olu iona y cons ain s should p e e en ially
elimina e non-synonymous a ia ion because i is usually associa ed wi h dele e ious mu a ions [35].
Non-synonymous SNPs in coding egions ep esen al e na i e allelic a ian s o a gene, which can
de e mine unc ional changes in he co esponding p o eins and lead o pheno ypic changes. Among
hese genes he e can be ound a e inol dehyd ogenase (SmaSNP_264), h ee zona pellucida p o eins
(SmaSNP_212, SmaSNP_217, SmaSNP_282) ela ed o ep oduc ion p ocesses, and a lipocalin
(SmaSNP_325) in ol ed in ea sec e ion (Table 2).
In he p esen s udy, we used sequences ob ained om wo ansc ip ome 454-py osequencing uns,
one ela ed o immune sys em [10] and ano he one om he hypo halamic pi ui a y-gonad axis [11].
Thus, GO e ms we e mainly ela ed o immune esponse and ep oduc ion p ocesses (Table 2). The
non-synonymous a ia ion was associa ed wi h genes in ol ing ei he immune esponse o sex
di e en ia ion p ocesses. A la ge numbe o SNP linked o anno a ed genes we e iden i ied and alida ed.
This se o ma ke s a e being used o popula ion genomic s udies and u bo gene ic map en ichmen ,
bo h app oaches p o iding use ul in o ma ion o e olu iona y and u bo indus y applied s udies.
In . J. Mol. Sci. 2013, 14 5702
Table 2. P edic ed posi ion, SNP loca ion wi hin genes and hei co esponden
synonymous s. non-synonymous a ian s o he 113 echnically easible SNPs.
SNP Name SNP loca ion/e ec GO e m
SmaSNP_211 3' UTR phospho yla ion (GO:0016310)
SmaSNP_212 Non synonymous ep oduc ion (GO:0000003)
SmaSNP_215 Synonymous mi o ic cell cycle (GO:0000278)
SmaSNP_216 3' UTR p o ein localiza ion o cell di ision si e (GO:0072741)
SmaSNP_217 Non synonymous binding o spe m o zona pellucida ( GO:0007339)
SmaSNP_218 Non synonymous p o ein impo in o mi ochond ial ma ix (GO:0030150)
SmaSNP_219 3' UTR ibonucleop o ein complex biogenesis (GO:0022613)
SmaSNP_220 Synonymous ibosomal la ge subuni assembly (GO:0000027)
SmaSNP_222 5' UTR egula ion o pep idoglycan ecogni ion p o ein signaling pa hway (GO:0061058)
SmaSNP_223 Synonymous cell adhesion (GO:0007155)
SmaSNP_224 Synonymous DNA-dependen ansc ip ion, ini ia ion (GO:0006352)
SmaSNP_225 3' UTR ibosomal la ge subuni assembly (GO:0000027)
SmaSNP_226 Synonymous cellula alcohol me abolic p ocess (GO:0044107)
SmaSNP_227 Synonymous hio edoxin biosyn he ic p ocess (GO:0042964)
SmaSNP_228 5' UTR egula ion o nucleo ide-binding oligome iza ion domain con aining signaling
pa hway (GO:0070424 )
SmaSNP_229 3' UTR immune esponse o umo cell (GO:0002418)
SmaSNP_230 3' UTR ep oduc ion (GO:0000003)
SmaSNP_231 Non synonymous phospho yla ion o RNA polyme ase II C- e minal domain (GO:0070816)
SmaSNP_232 Synonymous esolu ion o meio ic ecombina ion in e media es (GO:0000712)
SmaSNP_233 Synonymous ubiqui in-dependen p o ein ca abolic p ocess (GO:0006511)
SmaSNP_234 3' UTR egula ion o mac ophage in lamma o y p o ein 1 alpha p oduc ion (GO:0071640)
SmaSNP_235 Synonymous p o ein localiza ion o nucleola DNA epea s (GO:0034503)
SmaSNP_236 Synonymous T-helpe 1 cell ac i a ion (GO:0035711)
SmaSNP_237 5' UTR e mina ion o G-p o ein coupled ecep o signaling pa hway (GO:0038032)
SmaSNP_238 Non synonymous ansc ip ion ini ia ion om RNA polyme ase III ype 2 p omo e (GO:0001023)
SmaSNP_239 5' UTR No ound
SmaSNP_240 Synonymous MHC class I p o ein complex assembly (GO:0002397)
SmaSNP_241 3' UTR ep oduc ion (GO:0000003)
SmaSNP_243 Non synonymous neu onal s em cell main enance (GO:0097150)
SmaSNP_244 Unknown No ound
SmaSNP_245 5' UTR ep oduc ion (GO:0000003)
SmaSNP_246 Synonymous in acellula p o ein ansmemb ane anspo (GO:0065002)
SmaSNP_247 Non synonymous ibosomal p o ein impo in o nucleus (GO:0006610)
SmaSNP_248 3' UTR egula ion o mi o ic ecombina ion (0000019)
SmaSNP_249 3' UTR alpha-L- ucosidase ac i i y (GO:0004560)
SmaSNP_250 3' UTR modula ion by i us o hos p o ein se ine/ h eonine phospha ase ac i i y
(GO:0039517)
SmaSNP_252 3' UTR egula ion o mac ophage in lamma o y p o ein 1 alpha p oduc ion (GO:0071640)
SmaSNP_253 Synonymous egula ion o cy okinesis (GO:0032465)
SmaSNP_256 Synonymous egula ion o ibonuclease ac i i y (GO:0060700)
In . J. Mol. Sci. 2013, 14 5709
25. Bes e , A.E.; Rood -Wilding, R.; Whi ake , H.A. Disco e y and e alua ion o single nucleo ide
polymo phisms (SNPs) o Halio is midae: A a ge ed EST app oach. Anim. Gene ics 2008, 39,
321–324.
26. Sachidanandam, R.; Weissman, D.; Schmid , S.C.; Kakol, J.M.; S ein, L.D.; Ma h, G.; She y, S.;
Mullikin, J.C.; Mo imo e, B.J.; Willey, D.L.; e al. A map o human genome sequence a ia ion
con aining 1.42 million single nucleo ide polymo phisms. Na u e 2001, 409, 928–933.
27. Reich, D.E.; Gab iel, S.B.; Al shule , D. Quali y and comple eness o SNP da abases.
Na . Gene ics 2003, 33, 457–458.
28. B um ield, R.T.; Bee li, P.; Nicke son, D.A.; Edwa ds, S.V. The u ili y o single nucleo ide
polymo phisms in in e ences o popula ion his o y. T ends Ecol. E ol. 2003, 18, 249–256.
29. Ve a, M.; Al a ez-Dios, J.A.; Millan, A.; Pa do, B.G.; Bouza, C.; He mida, M.; Fe nandez, C.;
de la He an, R.; Molina-Luzon, M.J.; Ma inez, P. Valida ion o single nucleo ide polymo phism
(SNP) ma ke s om an immune Exp essed Sequence Tag (EST) u bo ; Scoph halmus maximus;
da abase. Aquacul u e 2011, 313, 31–41.
30. S ickney, H.L.; Schmu z, J.; Woods, I.G.; Hol ze , C.C.; Dickson, M.C.; Kelly, P.D.;
Mye s, R.M.; Talbo , W.S. Rapid mapping o zeb a ish mu a ions wi h SNPs and oligonucleo ide
mic oa ays. Genome Res. 2002, 12, 1929–1934.
31. Cenadelli, S.; Ma an, V.; Bongioni, G.; Fuse i, L.; Pa ma, P.; Aleand i, R. Iden i ica ion o
nuclea SNPs in gil head seab eam. J. Fish. Biol. 2007, 70, 399–405.
32. Hayes, B.; Lae dahl, J.K.; Lien, S.; Moen, T.; Be g, P.; Hinda , K.; Da idson, W.S.; Koop, B.F.;
Adzhubei, A.; Hoyheim, B. An ex ensi e esou ce o single nucleo ide polymo phism ma ke s
associa ed wi h A lan ic salmon (Salmo sala ) exp essed sequences. Aquacul u e 2007, 265, 82–90.
33. Wang, S.L.; Sha, Z.X.; Sons ega d, T.S.; Liu, H.; Xu, P.; Som idhi ej, B.; Pea man, E.;
Kucuk as, H.; Liu, Z.J. Quali y assessmen pa ame e s o EST-de i ed SNPs om ca ish.
BMC Genomics 2008, 9, 450.
34. Hube , S.; Bussey, J.T.; Higgins, B.; Cu is, B.A.; Bowman, S. De elopmen o single nucleo ide
polymo phism ma ke s o A lan ic cod (Gadus mo hua) using exp essed sequences. Aquacul u e
2009, 296, 7–14.
35. Hube , S.; Higgins, B.; Bo za, T.; Bowman, S. De elopmen o a SNP esou ce and a gene ic
linkage map o A lan ic cod (Gadus mo hua). BMC Genomics 2010, 11, 191.
36. Sau age, C.; Bie ne, N.; Lapegue, S.; Boud y, P. Single nucleo ide polymo phisms and hei
ela ionship o codon usage bias in he Paci ic oys e C assos ea gigas. Gene 2007, 406, 13–22.
37. Ve a, M.; Pa do, B.G.; Pino-Que ido, A.; Al a ez-Dios, J.A.; Fuen es, J.; Ma inez, P.
Cha ac e iza ion o single-nucleo ide polymo phism ma ke s in he Medi e anean mussel;
My ilus gallop o incialis. Aquac. Res. 2010, 41, e568–e575.
38. Zhang, L.S.; Guo, X.M. De elopmen and alida ion o single nucleo ide polymo phism ma ke s
in he eas e n oys e C assos ea i ginica Gmelin by mining ESTs and esequencing.
Aquacul u e 2010, 302, 124–129.
39. Du, Z.Q.; Ciobanu, D.C.; On e u, S.K.; Go bach, D.; Mileham, A.J.; Ja amillo, G.; Ro hschild, M.F.
A gene-based SNP linkage map o paci ic whi e sh imp; Li openaeus annamei. Anim. Gene ics
2010, 41, 286–294.
In . J. Mol. Sci. 2013, 14 5710
40. Go bach, D.M.; Hu, Z.L.; Du, Z.Q.; Ro hschild, M.F. Mining ESTs o de e mine he use ulness o
SNPs ac oss sh imp species. Anim. Bio echnol. 2010, 21, 100–103.
41. Lepoi e in, C.; F ige io, J.M.; Ga nie -Ge e, P.; Salin, F.; Ce e a, M.T.; Vo nam, B.;
Ha eng , L.; Plomion, C. In i o s. in silico de ec ed SNPs o he de elopmen o a geno yping
a ay: Wha can we lea n om a non-model species? PLoS One 2010, 5, e11034.
42. Zhu, C.; Cheng, L.; Tong, J.; Yu, X. De elopmen and cha ac e iza ion o new single nucleo ide
polymo phism ma ke s om exp essed sequence ags in common ca p (Cyp inus ca pio). In . J.
Mol. Sci. 2012, 13, 7343–7353.
43. Roche Diagnos ics GmbH. cDNA Rapid Lib a y P epa a ion Me hod Manual; Roche Applied
Science: Manheim, Ge many, 2009.
44. Campbell, N.R.; Amish, S.J.; P i cha d, V.L.; McKel ey, K.S.; Young, M.K.; Schwa z, M.K.;
Ga za, J.C.; Luika , G.; Na um, S.R. De elopmen and e alua ion o 200 no el SNP assays o
popula ion gene ic s udies o wes slope cu h oa ou and gene ic iden i ica ion o ela ed axa.
Mol. Ecol. Resou . 2012, 12, 942–949.
45. Koski, L.B.; G ay, M.W.; Lang, B.F.; Bu ge , G. Au oFACT: An (Au o)unde -ba ma ic
(F)unde -ba unc ional (A)unde -ba nno a ion and (C)unde -ba lassi ica ion (T)unde -ba ool.
BMC Bioin o ma. 2005, 6, 151.
46. Kim, H.; Schmid , C.J.; Decke , K.S.; Ema a, M.G. A double-sc eening me hod o iden i y eliable
candida e non-synonymous SNPs om chicken EST da a. Anim. Gene ics 2003, 34, 249–254.
47. Wondji, C.S.; Hemingway, J.; Ranson, H. Iden i ica ion and analysis o single nucleo ide
polymo phisms (SNPs) in he mosqui o Anopheles unes us; mala ia ec o . BMC Genomics
2007, 8, 5.
48. Che eux, B.; P is e e , T.; D esche , B.; D iesel, A.; Mülle , W.E.G.; We e , T.; Suhai, S. Using
he mi aEST assemble o eliable and au oma ed mRNA ansc ip assembly and d ec ion in
sequenced ESTs. Genome Res. 2004, 14, 1147–1159.
49. Huang, X.; Madan, A. CAP3: A DNA sequence assembly p og am. Genome Res. 1999, 9, 868–877.
50. Ueno, S.; le P o os , G.; Lege , V.; Klopp, C.; Noi o , C.; F ige io, J.-M.; Salin, F.; Salse, J.;
Ab ouk, M.; Mu a , F.; e al. Bioin o ma ic analysis o ESTs collec ed by Sange and
py osequencing me hods o a keys one o es ee species: Oak.BMC Genomics 2010, 11, 650.
51. Tang, J.; Vosman, B.; Voo ips, R.E.; Linden, C.G.; an de Linden, C.G.; Leunissen, J.A.M.
Quali ySNP: A pipeline o de ec ing single nucleo ide polymo phisms and inse ions/dele ions in
EST da a om diploid and polyploid species. BMC Bioin o ma. 2006, 7, 438.
52. MySQL Home Page. A ailable online: h p://www.mysql.com (accessed on 1 July 2009).
53. M-View Home Page. A ailable online: h p://bio-m iew.sou ce o ge.ne (accessed on 1 July 2009).
54. BLAST Home Page. A ailable online: h p://blas .ncbi.nlm.nih.go /Blas .cgi? (accessed on 2
No embe 2012).
55. Samb ook, J.; F i sch, E.F.; Mania is, T. Molecula Cloning: A Labo a o y Manual, 1s ed.;
P ess CSHL: New Yo k, NY, USA, 1989.
In . J. Mol. Sci. 2013, 14 5711
56. Bue ow, K.H.; Edmonson, M.; MacDonald, R.; Cli o d, R.; Yip, P.; Kelley, J.; Li le, D.P.;
S ausbe g, R.; Koes e , H.; Can o , C.R.; e al. High- h oughpu de elopmen and
cha ac e iza ion o a genomewide collec ion o gene-based single nucleo ide polymo phism
ma ke s by chip-based ma ix-assis ed lase deso p ion/ioniza ion ime-o - ligh mass
spec ome y. P oc. Na l. Acad. Sci. USA 2001, 98, 581–584.
57. Oe h, P.; del Mis o, G.; Ma nellos, G.; Shi, T.; an den Boom, D. Quali a i e and quan i a i e
geno yping using single base p ime ex ension coupled wi h ma ix-assis ed lase
deso p ion/ioniza ion ime-o - ligh mass spec ome y (MassARRAY). Me hods Mol. Biol. 2009,
578, 307–343.
58. Goude , J. FSTAT, a p og am o es ima e and es gene di e si ies and ixa ion indices ( e sion 2.9.3).
A ailable online: h p://www.unil.ch/izea/so wa es/ s a .h ml (accessed on 1 Ma ch 2003).
59. Raymond, M.; Rousse , F. GENEPOP (Ve sion 1.2)—Popula ion gene ics so wa e o exac es s
and ecumenicism. J. He ed. 1995, 86, 248–249.
60. Rousse , F. GENEPOP’007: A comple e e-implemen a ion o he GENEPOP so wa e o
Windows and Linux. Mol. Ecol. Resou . 2008, 8, 103–106.
61. Louis, E.J.; Demps e , E.R. An exac es o Ha dy-Weinbe g and mul iple alleles. Biome ics
1987, 43, 805–811.
62. Rice, W.R. Analyzing ables o s a is ical es s. E olu ion 1989, 43, 223–225.
63. ORF Finde Home Page. A ailable online: h p://www.ncbi.nlm.nih.go /go /go .h ml (accessed
on 8 No embe 2012).
64. Thompson, J.D.; Higgins, D.G.; Gibson, T.J. CLUSTAL W imp o ing he sensi i i y o
p og essi e mul iple sequence alignmen h ough sequence weigh ing; posi ion-speci ic gap
penal ies and weigh ma ix choice. Nucl. Acids Res. 1994, 22, 4673–4680.
65. Hall, T.A. BioEdi : A use - iendly biological sequence alignmen edi o and analysis p og am o
Windows 95/98/NT. Nucl. Acids Symp. Se . 1999, 41, 95–98.
66. QuickGO Home Page. A ailable online: h p://www.ebi.ac.uk/QuickGO/ (accessed on 15
Oc obe 2012).
67. AmiGO Home Page. A ailable online: h p://amigo.geneon ology.o g/cgi-bin/amigo/go.cgi
(accessed on 16 Oc obe 2012).
© 2013 by he au ho s; licensee MDPI, Basel, Swi ze land. This a icle is an open access a icle
dis ibu ed unde he e ms and condi ions o he C ea i e Commons A ibu ion license
(h p://c ea i ecommons.o g/licenses/by/3.0/).