Genomic O ganiza ion, Molecula Di e si ica ion, and
E olu ion o An imic obial Pep ide My icin-C Genes in
he Mussel (
My ilus gallop o incialis
)
Manuel Ve a
1
, Paulino Ma ı
´nez
1
*, Lau a Poisa-Bei o
2
, An onio Figue as
2
, Bea iz No oa
2
1Depa amen o de Gene
´ ica, Facul ad de Ve e ina ia. Uni e sidad de San iago de Compos ela, Lugo, Spain, 2Ins i u o de In es igaciones Ma inas, CSIC, Vigo, Spain
Abs ac
My icin-C is a highly a iable an imic obial pep ide associa ed o immune esponse in Medi e anean mussel (My ilus
gallop o incialis). In his s udy, we ied o asce ain he gene ic o ganiza ion and he mechanisms unde lying my icin-C
a ia ion and e olu ion o his gene amily. We ook ad an age o he la ge in on size a ia ion o ind ou he numbe o
my icin-C genes. Using agmen analysis a maximum o ou alleles was de ec ed pe indi idual a bo h in ons in a la ge
mussel sample sugges ing a minimum o wo my icin-C genes. The ansmission pa e n o size a ian s in wo ull-sib
amilies was also used o asce ain he numbe o my icin-C genes unde lying he a iabili y obse ed. Resul s in bo h
amilies we e in acco dance wi h wo my icin-C genes o ganized in andem. A mo e de ailed analysis o my icin-C a ia ion
was ca ied ou by sequencing a la ge sample o complemen a y (cDNA) and genomic DNA (gDNA) in 10 indi iduals. Two
basic sequences we e de ec ed a mos indi iduals and se e al sequences we e cons i u ed by combina ion o wo di e en
basic sequences, s ongly sugges ing soma ic ecombina ion o gene con e sion. Sligh wi hin-basic sequence a ia ion
de ec ed in all indi iduals was a ibu ed o soma ic mu a ion. Such mu a ions we e mo e equen ly a he C- e minal
domain and mos ly de e mined non-synonymous subs i u ions. The ma u e pep ide domain showed he highes a ia ion
bo h in he whole cDNA and in he basic-sequence samples, which is in acco dance wi h he pa hogen ecogni ion unc ion
associa ed o his domain. Al hough mos es s sugges ed neu ali y o my icin-C a ia ion, e idence indica ed posi i e
selec ion in he ma u e pep ide and C- e minal egion. Th ee main highly suppo ed clus e s we e obse ed when
econs uc ing phylogeny on basic sequences, meio ic ecombina ion playing a ele an ole on my icin-C e olu ion. This
s udy demons a es ha mechanisms o gene a e molecula a ia ion simila o ha obse ed in e eb a es a e also
ope a ing in molluscs.
Ci a ion: Ve a M, Ma ı
´nez P, Poisa-Bei o L, Figue as A, No oa B (2011) Genomic O ganiza ion, Molecula Di e si ica ion, and E olu ion o An imic obial Pep ide
My icin-C Genes in he Mussel (My ilus gallop o incialis). PLoS ONE 6(8): e24041. doi:10.1371/jou nal.pone.0024041
Edi o : Bin Tian, UMDNJ-New Je sey Medical School, Uni ed S a es o Ame ica
Recei ed May 6, 2011; Accep ed Augus 2, 2011; Published Augus 31, 2011
Copy igh : ß2011 Ve a e al. This is an open-access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion License, which pe mi s un es ic ed
use, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal au ho and sou ce a e c edi ed.
Funding: This wo k has been unded by he p ojec AGL2008-05111/ACU om he Spanish Minis e io de Ciencia e Inno a ion. The unde s had no ole in s udy
design, da a collec ion and analysis, decision o publish, o p epa a ion o he manusc ip .
Compe ing In e es s: The au ho s ha e decla ed ha no compe ing in e es s exis .
* E-mail: [email p o ec ed]s
In oduc ion
In e eb a es a e a he e ogeneous g oup o animals which
cons i u e he huge majo i y o he su i ing animal phyla.
In e es ingly, only one o abou 35 known animal phyla includes
e eb a es. Usually he en i onmen s whe e in e eb a es dwell
a e abundan on po en ially pa hogenic mic oo ganims. Al hough
in ecen yea s he e ha e been ad ances in he knowledge o
in e eb a e immuni y, a comp ehensi e iew o he immune
mechanisms deployed ac oss he b oad spec um o in e eb a e
phyla [1] is no a ailable. One ecu en ques ion is how hese
animals su i e wi hou an acqui ed immune sys em. In pa icula ,
ma ine in e eb a es, such as bi al es, a e in con ac wi h all so
o po en ial pa hogens such as i uses, bac e ia and pa asi es due
o hei il e ing ac i i ies. Mussels (ex: My ilus gallop o incialis)
p esen a high il e ing ac i i y: one adul mussel can il e oughly
eigh li e s o wa e in one hou [2]–[][4], which implies ha hey
a e in in ima e con ac wi h a wide a ie y o mic oo ganisms.
A key elemen o he immune sys em is he disc imina ion
be ween sel and non-sel . This implies, especially in complex
plu icellula o ganisms, a molecula code o p o ide he singula i y
o each indi idual, whose molecula basis is pa icula ly well
known wi hin e eb a es [5]. On he o he hand, ecogni ion o
non-sel can be achie ed by iden i ying pa hogen-associa ed
molecula pa e ns (PAMPs), mos ly ela ed o inna e immuni y,
o by de ec ing o eign (non-sel ) molecules, cha ac e is ic o he
adap i e immune esponse [6], [7]. The e o e, gene a ion o
molecula di e si y is essen ial o some key elemen s o he
immune sys em and di e en s a egies ha e been de eloped along
e olu ion o molecula di e si ica ion. The p ima y mechanisms
a e ela ed o he exploi a ion o some genome p ope ies such as
ecombina ion, mu a ion, al e na i e splicing and exon shu ling
on speci ic genes which equi e high di e si y o ul ill hei
unc ion. These mechanisms ac mainly in he soma ic cell line,
while ge mline cells main ain hese genes unal e ed h ough
gene a ions and a e only subjec ed o he gene al p ocesses o
genome a ia ion [1], [8], [9]. O he p oposed mechanisms a e
ela ed o in e ac ion o di e en molecules which can p omo e
a iabili y aking ad an age o he high combina o y o di e en
elemen s (syne gism); o changes in he amoun o speci ic
molecules by gene duplica ion; and o a ia ion a egula o y
elemen s (dosage) [9].
PLoS ONE | www.plosone.o g 1 Augus 2011 | Volume 6 | Issue 8 | e24041
Gene a ion o gene ic di e si y o e eb a e immunoglobulin
cons i u es one o he bes s udied p ocesses o he immune sys em,
and i in ol es bo h in agenic ecombina ion and hype mu a ion
[10], [11]. The high allelic a ia ion o he Majo His ocompa -
ibili y Complex (MHC) genes and he main e olu iona y o ces
d i en i , is also well documen ed [12], [13]. Wi hin in e eb a es,
high molecula di e si y has been epo ed a Dscam ( ela ed o
he immunoglobulin supe amily) in D osophila due o al e na i e
splicing leading o mo e han 30.000 di e en iso o ms [14], and
a ib inogen ela ed p o eins (FREPs) in he snail Biomphala ia
glab a a elying on soma ic mu a ion and ecombina ion mecha-
nisms [1]. Howe e , u he s udies a e needed o inc ease
knowledge on in e eb a es immuni y, and pa icula ly, o
unde s anding he mechanisms esponsible o molecula di e si-
ica ion.
An imic obial pep ides (AMPs) a e pep ides o small size which
p omo e e icien binding o s uc u al componen s o mic oo gan-
isms, acili a ing hei elimina ion h ough di e en e ec o
mechanisms in a wide a ie y o o ganisms [15]. The p esence o
AMP iso o ms has been epo ed o be co ela ed wi h an imp o ed
de ense agains pa hogens in well es ablished AMP sys ems such as
de ensins and ca helicidins [16], [17]. AMPs can ac as modi ie s o
inna e and adap i e immune esponse [18]. Fu he , AMPs
syn hesis wi hin insec s has demons a ed o be ac i a ed h ough
Tumou Nec osis Fac o (TNF) ecep o and Toll Like-Recep o
(TLR) molecules, ollowing simila pa hways o mammals [19].
Wi hin molluscs an impo an a ie y o AMPs has been epo ed,
including de ensins, my ilins, my icins and my imycins [20], [21].
The high a ia ion obse ed in some AMP amilies has been
a ibu ed o he high copy numbe a speci ic gene amilies [22].
My icin-C is a highly exp essed AMP du ing Medi e anean mussel
(My ilus gallop o incialis) diseases [23], which shows a ypical AMP
s uc u e including signal pep ide, ma u e pep ide and C- e minal
egion domains (Figu e 1). High sequence a iabili y has been
epo ed a my icin-C, sugges ing ha his wide epe oi e o
sequences may be ela ed o he high disease esis ance obse ed in
Medi e anean mussel (My ilus gallop o incialis) [24]. Rema kably,
his AMP a iabili y was no obse ed in lib a ies om o he
bi al es [25]–[32]. Recen ly, we ha e demons a ed ha my icin-C
p esen s an i i al ac i i y agains wo di e en ish i uses
(en eloped and non-en eloped) and ha is able o modula e he
mussel immune esponse by modi ying he exp ession o mussel
immune- ela ed genes and a ac ing hemocy es [33].
The Medi e anean mussel is a species o g ea ele ance in
aquacul u e wi h a wo ld p oduc ion abo e million Tons [34].
Mo ali ies a e equen in bi al es, bu mussels do no seem o be
suscep ible o he same pa hogens esponsible o massi e dea hs o
o he molluscs. By hei sesile cha ac e and esis ance, mussels a e
used as a model o moni o pollu ion in oceans all o e he wo ld
[35]. Al hough hese animals a e being cul u ed ex ensi ely, we
s ill a om unde s and how hey eac agains pa hogens. In his
wo k, we ha e add essed he s udy o he genomic o ganiza ion o
my icin-C genes aking ad an age o he high a iabili y desc ibed
a hei in ons. Fo his pu pose, in aindi idual and in apopu-
la ion gene ic di e si y o in ons 1 and 2 we e s udied in na u al
popula ions and he pa e n o gene ic ansmission analyzed in
ull-sib amilies. Besides, we in es iga ed he mechanisms ha may
explain he high gene ic di e si y epo ed o my icin-C by
analyzing and compa ing a la ge sample o high quali y
ansc ip omic and genomic sequences om se e al indi iduals.
Using his in o ma ion, we e alua ed he ole o selec ion on he
e olu ion o my icin-C.
Resul s
Gene ic di e si y o my icin-C in ons 1 and 2 in na u al
popula ions
Two mussel samples om NW Spain we e s udied o e alua e
gene ic a iabili y o my icin-C a indi idual and popula ion le els.
My icin-C in ons 1 and 2 we e chosen o his analysis because o
he high leng h a iabili y p e iously epo ed a hese gene
egions [23]. Acco dingly, gDNA agmen analysis was pe o med
o e eal hei a iabili y. We expec ed ha his analysis p o ided
new in o ma ion on my icin-C a ia ion in na u al popula ions
Figu e 1. S uc u e and unc ional domains o my icin-C gene om Medi e anean mussel (
My ilus gallop o incialis
).
doi:10.1371/jou nal.pone.0024041.g001
O ganiza ion and E olu ion o My icin-C Genes
PLoS ONE | www.plosone.o g 2 Augus 2011 | Volume 6 | Issue 8 | e24041
and some insigh s in o i s e olu ion, bu especially i should be
use ul o in e he minimum numbe o genes unde lying he
a ia ion obse ed. Up o ou peaks (alleles) pe indi idual we e
de ec ed o bo h in ons in he whole sample: om 1 o 4 alleles a
in on1 in bo h popula ions and a in on 2 in Co un˜a; and om 1
o 3 alleles a in on 2 in Vigo. This sugges s he exis ence o a
leas wo my icin-C loci. Gene ic a ia ion o my icin-C was highe
a in on 2 han a in on 1 in bo h popula ions, despi e he la ge
allelic ange o in on 1. Th ee main modes we e obse ed in
allelic equency dis ibu ions o in on 1 (184, 221 and 376) and
wo o in on 2 (199/200 and 210/211) in bo h popula ions
(Figu e 2). Gene ic di e si y was e y simila in bo h popula ions:
Vigo (In on 1: A = 12, gene di e si y = 0.779; In on 2: A = 16;
gene di e si y = 0.899); Co un˜a (In on 1: A = 12, gene di e si-
y = 0.761; In on 2: A = 18; gene di e si y = 0.913). No signi ican
di e ences we e obse ed o allele equency dis ibu ions
be ween Co un˜a and Vigo popula ions o bo h in ons
(Wilcoxon-Mann-Whi ney es in on 1: Z = 20.258, P = 0.796;
in on 2: Z = 20.657; P = 0.511).
Mendelian seg ega ion o my icin-C in ons 1 and 2 size
a ian s
As expec ed o in ons o he same gene, ull geno ypic
disequilib ium was obse ed be ween in on 1 and in on 2 in bo h
amilies (Table 1). Thus, in he i s amily, he allele 376 o in on
1 in he a he was ansmi ed always linked o allele 211 o in on
2, while allele 223 was ansmi ed linked o allele 202. O sp ing
geno ypes adjus ed o 1:1:1:1 p opo ions in bo h amilies o bo h
in ons ( amily 1: x
2
= 1.842; P = 0.606; amily 2: x
2
= 1.759;
P = 0.624). These a e he expec ed p opo ions unde a single
locus seg ega ion hypo hesis, when bo h pa en s a e he e ozygous
o di e en alleles. Howe e , one pa en in bo h c osses exhibi ed
mo e han wo alleles in bo h in ons and seg ega ion be ween
hem was no a andom. In ac , alleles 221/223 om he mo he
in he i s c oss and alleles 221/372 om he a he in he second
c oss we e always ansmi ed join ly as a single Mendelian uni a
in on 1. The same occu ed a in on 2, whe e alleles 199/201
om he mo he in he i s c oss and alleles 201/205 om he
a he in he second c oss we e ansmi ed oge he . The exis ence
o wo 201 alleles in he a he o he second c oss was in e ed by
he 200/201/205 o sp ing de ec ed and con i med by he oughly
double heigh o he 201 peak in he a he and in he 201/201/
205 o sp ing. The mos plausible explana ion o hese obse a-
ions is he exis ence o wo my icin-C closely linked genes
a anged in andem (Figu e 3).
T ansc ip omic a ia ion o my icin-C
Nine y h ee high quali y cDNA sequences om 10 M.
gallop o incialis indi iduals (GenBank Accession num-
be s = JF990711–JF990804) we e inally selec ed among he 100
sequences ob ained (10 indi iduals610 cDNA sequences) o s udy
gene ic a ia ion o my icin-C a ansc ip omic le el (Figu e 4).
Be ween wo and, mo e occasionally, h ee highly di e gen cDNA
sequences we e obse ed wi hin each indi idual (ma ked wi h
di e en backg ound colo in Figu e 4). Small di e ences we e
obse ed wi hin each o hese sequences due o single nucleo ide
subs i u ions (single ons), and in some cases, sequences appea ed o
be cons i u ed by combina ion o wo o he a o emen ioned highly
di e gen sequences. As explained below, all da a poin owa d
soma ic mu a ion and ecombina ion o explain he di e ences
obse ed wi hin hese highly di e gen cDNA sequences. Thus, we
de ined basic sequences as hose di e gen cDNA sequences
exis ing in each indi idual excluding single ons and/o ecombi-
na ion e en s.
Two di e en basic sequences we e iden i ied a mos
indi iduals, and only wo mussels showed ei he h ee basic
Figu e 2. Allelic equency dis ibu ion (in base pai s) a my icin-C in ons 1 and 2 in wo popula ions om NW Spain (Co un
˜a and
Vigo).
doi:10.1371/jou nal.pone.0024041.g002
O ganiza ion and E olu ion o My icin-C Genes
PLoS ONE | www.plosone.o g 3 Augus 2011 | Volume 6 | Issue 8 | e24041
cDNA sequences (indi idual 11) o e idences o a hi d one
(indi idual 14). Sligh di e ences we e de ec ed among each one o
hese basic sequences wi hin indi iduals mos ly due o single
nucleo ide subs i u ions: 80.0% sequences showed one; 15% wo;
and 1% h ee. A o al o 30 subs i u ions we e de ec ed in he
whole cDNA sample (all unique), mos ly ep esen ing single ons
(86.7%) and he emaining ou being new nucleo ide a ian s a
ex an a iable si es (13.3%). Th ee o hese ecu en mu a ion
si es we e loca ed in a highly polymo phic egion (be ween
nucleo ides 229 and 246; 35.3% a iable si es) ha could include
mu a ion ho spo si es. These obse a ions sugges ha a ia ion
wi hin each basic sequence is a consequence o soma ic poin
mu a ion mechanisms. Mu a ion a e pe si e would be 1.1*10
23
in he 93 sequences analyzed. Soma ic mu a ions we e non-
andomly dis ibu ed acco ding o he my icin-C polypep ide
s uc u e. Thus, he ma u e pep ide showed he lowes p opo ion
o mu an s (7/120 = 0.058), while he signal pep ide (0.083) and,
especially he C- e minal egion (0.150), displayed a highe
p opo ion. A high pe cen age o hese poin mu a ions (56.7%)
cons i u ed non-synonymous a ian s gi ing ise o aminoacid
subs i u ions, mos ly a ec ing he C- e minal egion (64.7%).
A second ele an ea u e o ansc ip omic analysis o my icin-
C was he e idence o ecombina ion e en s a some indi iduals
e lec ed by he p esence o new sequences a ising as combina ion
o basic cDNA sequences (Figu e 4). Fi e indi iduals ou o en
analyzed showed one o wo ecombinan s a he 9–10 sequences
analyzed pe indi idual. Recombinan s we e he esul o he
combina ion o wo basic sequences, bu in one case appa en ly
h ee basic sequences could be in ol ed (sequence 14_09). The
a e age a e o ecombina ion pe indi idual was 0.06460.023.
A o al o 21 basic sequences (excluding single nucleo ide
a ian s and ecombinan s, as de ined abo e) we e de ec ed
among he 93 cDNA sequences s udied.
Genomic a ia ion o my icin-C
To compa e he ea u es obse ed a ansc ip ome le el and o
ge new insigh s in o he gene ic basis o my icin-C a ia ion, we
analyzed 88 high quali y o wa d and e e se my icin-C gDNA
sequences in eigh indi iduals (be ween 8–15 pe indi idual;
GenBank Accession numbe s = JF990616–JF990710) p e iously
e alua ed o cDNA a ia ion (Fig. S1). The pa e n o my icin-C
gDNA a ia ion was simila o ha obse ed a cDNA. Be ween
one and h ee basic gDNA sequences (de ined as o cDNA
sequences) we e obse ed in each indi idual wi h sligh di e ences
wi hin basic sequences due o one o wo p esumed poin
mu a ions. Basic gDNA and cDNA sequences a each indi idual
we e iden ical, al hough only one basic gDNA sequence was
de ec ed in indi iduals 16 and 17, while hey showed wo di e en
basic cDNAs. Poin mu a ion a e a gDNA (9.1*10
24
) was sligh ly
lowe han ha obse ed in cDNA analysis and only 1 ou o he
21 mu a ions de ec ed a gDNA was also obse ed in hei
co esponden cDNAs. Finally, i was ema kable ha no
ecombinan sequences we e de ec ed a gDNA, when a ound
six should be expec ed acco ding o ecombina ion a e obse ed
a cDNA.
Analysis o gDNA also enabled us a mo e de ailed e alua ion o
a ia ion a in onic egions and i s compa ison wi h agmen
analysis da a. Fi s ly, he h ee and wo allelic modes obse ed,
espec i ely, a in on 1 and in on 2 agmen analysis
dis ibu ions (Figu e 2), we e explained by he p esence o wo
(35 and 155 bp) and one (10 bp) la ge indel/s a in ons 1 and 2,
espec i ely (Fig. S1). Mino and less equen indels (be ween 9–
24 bp) and a iable single mononucleo ide epe i ions (poli A, poli
C, bu especially poli T) ga e accoun o size a ia ion a ound
hese main modes. Second, excluding indels, nucleo ide a ia ion
a bo h in ons was highe han ha obse ed a exons (in on 1:
seg ega ing si es (S) = 35.1%; haplo ype di e si y (Hd) = 0.973, and
Wa e son’s es ima o o nucleo ide di e si y based on he
p opo ion o seg ega ing si es (h
W
) = 0.07759; in on 2:
S = 33.6%; Hd = 0.987, and h
W
= 0.07554), al hough exon 2
showed di e si y igu es e y close o bo h in ons (Tables 2 and 3).
Thi d, simila a ia ion o ha desc ibed o cDNA basic
sequences a ibu ed o soma ic mu a ions was obse ed wi hin
basic sequences a bo h in ons. Fou h, some disco dance was
de ec ed be ween agmen analysis and sequencing a gDNA.
Thus, some leng h a ian s obse ed in he agmen analysis we e
no de ec ed in he gDNA sequencing in se e al indi iduals. I
appea ed like some my icin-C genes showed low o no
Table 1. Inhe i ance o my icin-C in on size a ian s in M. gallop o incialis.
Family 1 In on 1 In on 2 Family 2 In on 1 In on 2
Fa he 223/376 202/211 Fa he 218/221/372 201/201/205
Mo he 221/223/369 198/199/201 Mo he 219/362 201/201
F equency O sp ing geno ypes F equency O sp ing geno ypes
4 221/223 199/201/202 6 218/219 200/201
7 221/223/376 199/201/211 7 219/221/372 200/201/205
3 223/369 198/202 9 218/362 201/201
5 369/376 198/211 7 221/372/362 201/201/205
doi:10.1371/jou nal.pone.0024041. 001
Figu e 3. Hypo hesis on genomic a chi ec u e o my icin-C genes om amilia and popula ion agmen analysis da a o in on 1
and in on 2.
doi:10.1371/jou nal.pone.0024041.g003
O ganiza ion and E olu ion o My icin-C Genes
PLoS ONE | www.plosone.o g 4 Augus 2011 | Volume 6 | Issue 8 | e24041
O ganiza ion and E olu ion o My icin-C Genes
PLoS ONE | www.plosone.o g 5 Augus 2011 | Volume 6 | Issue 8 | e24041
ampli ica ion wi h he p ime s used. To check his possibili y, we
epea ed he agmen analysis in hese indi iduals bu using a wo-
s ep PCR. We i s ampli ied my icin-C genes om o al gDNA
ma e ial, and hen used his DNA in a second s ep using speci ic
p ime s o ampli y in ons 1 and 2. I he e we e di e ences in
ampli ica ion o he wo hypo hesized my icin-C genes, we should
obse e di e ences in hei co esponden in ons 1 and 2
agmen s in compa ison wi h agmen analysis s a ing om
o al DNA. Resul s showed ha , indeed, some agmen s obse ed
in ou p e ious agmen analysis (s a ing om o al DNA) we e
missed when ampli ied my icin-C gDNA was used o pe o m
in on PCRs (Figu e 5). This s ongly sugges s ha one o he wo
hypo hesized my icin-C genes could be unde -ampli ied, p obably
due o misma ches a he p ime empla e egions.
E olu ion o my icin-C
All he 93 cDNA sequences analyzed and he 21 basic sequences
de ec ed in he ansc ip omic analysis we e used o analyze he
e olu iona y pa e n o my icin-C genes. Es ima o s o gene ic
di e si y con i med he high gene ic a ia ion in he whole cDNA
sample (24.0% a iable si es; p= 0.04119; h
W
= 0.04692; Table 2)
and in he basic sequences (16.0%; p= 0.04190 h
W
= 0.04389;
Table 3). Basic cDNA sequences showed a lowe p opo ion o
a iable si es (16%) han whole cDNA sequences (24%) because
soma ic mu a ions we e excluded by de ini ion in hei composi-
ion. Howe e , a e age numbe o nucleo ide di e ences pe si e
(nucleo ide di e si y, p) and nucleo ide di e si y based on he
p opo ion o seg ega ing si es (h
W
) we e e y simila in bo h
cDNA samples p obably because he highe p opo ion o a iable
si es in he whole cDNA sample was coun e balanced by he
epe i ion o basic sequences.
Dissec ion o gene ic di e si y acco ding o he di e en egions
o my icin-C p o ein e ealed ha ma u e pep ide displayed
highe gene ic di e si y han C- e minal egion and signal pep ide.
The di e ence was e en highe in he basic sequences han in he
whole cDNA sample. The highe p opo ion o single ons in he
whole cDNA sample a he C- e minal egion and e en a he
signal pep ide e lec s he highe impac o soma ic poin mu a ion
a hese egions ega ding he ma u e pep ide, as ou lined be o e.
Howe e , a la ge p opo ion o single ons in ma u e pep ide we e
de ec ed in he basic sample indica ing he highe e olu iona y
di e si ica ion o his domain. Mos neu ali y es s we e no
signi ican in he whole pep ide and when applied o he speci ic
domains o my icin-C, hus sugges ing no e ec s o selec ion on he
a ia ion obse ed. Only he Fu and Li es o he whole cDNA
sample esul ed signi ican (D = 23.3483, p,0.02), hus sugges ing
pu i ying selec ion. Howe e , he a io be ween non-synonymous
s synonymous a ia ion was highe han 1 and signi ican unde a
M8-M8a model a ma u e pep ide and C- e minal egion bo h in
he whole cDNA sample and in he basic cDNA sequences. This
sugges s ha , al hough a ia ion is neu al a mos my icin-C
nucleo ide si es, especially a signal pep ide, posi i e selec ion
could be occu ing a some egions o ma u e pep ide and C-
e minal egion, hus ende ing global signi ican es s.
A phylogene ic ee was cons uc ed using a Bayesian me hod
implemen ed in MRBAYES 3.1.2 p og am o analyze phylogene ic
ela ionships o he 21 basic my icin-C cDNA sequences (Figu e 6).
Th ee highly suppo ed clades wi h boo s ap alues close o 100
we e iden i ied. The wo sis e g oups I and II we e much mo e
di e si ied han he III one. Visual inspec ion o basic sequences
sugges ed ha a majo ecombina ion e en could ha e occu ed
in he o igin o hese h ee main g oups (Figu e S2). In ac , i e
ecombina ion e en s we e es ima ed o all sequences using
Dnasp V5.0, e idencing he ole o ecombina ion on he
e olu ion o my icin-C genes.
Discussion
Genomic o ganiza ion o my icin-C
High nucleo ide a iabili y was p e iously epo ed o my icin-
C in ons 1 and 2 by sequence analysis [23] and DGGE
elec opho esis [24]. The a ia ion de ec ed wi h DGGE was so
Figu e 4. Va iable posi ions o my icin-C cDNA in 10 mussels om Co un
˜a na u al popula ion. In whi e, g ay o da k g ay backg ound,
he di e en basic sequences iden i ied a each indi idual. Single nucleo ide a ian s wi hin basic sequences highligh ed in g een (synonymous) and
yellow (non-synonymous). The sequence AM497977 om Genebank was included o e e ence in he analysis.
doi:10.1371/jou nal.pone.0024041.g004
Table 2. Gene ic di e si y pa e n a my icin-C in M. gallop o incialis wi h all cDNA sequences.
Summa y s a is ics Signal pep ide Ma u e pep ide C- e minal egion All egions
N94 949494
Si es 60 120 120 300
S 12 (5, 7) 31 (5, 26) 29 (15, 14) 72 (25,47)
p0.02137 (0.00216) 0.05415 (0.00393) 0.03769 (0.00140) 0.04119 (0.00202)
h
W
0.03910 (0.01448) 0.05009 (0.01495) 0.04724 (0.01422) 0,04692 (0.01258)
D
Tajima
21.2148 (p.0.10)20.1321 (p.0.10)20.9911 (p.0.10)20.7214 (p.0.10)
D
Fu and Li
21.6291 (p.0.10)20.0244 (p.0.10)23.3483 (p,0.02)22.0354 (0.10.p.0.05)
H- es 0.0000 (p = 0.3265)
(Ka/Ks) 0.948 1.027 2.771 1.275
M8-M8a LRT (Selec ion) Non Signi ican p,0.05 p,0.001 p,0.001
The sequence AM497977 om Genebank was included in he analysis. N: numbe o cDNA sequences; S: numbe o seg ega ing si e (in pa en heses single ons and
pa simony in o ma i e si es, espec i ely); p: a e age numbe o nucleo ide di e ences pe si e; h
W
: nucleo ide di e si y based on he p opo ion o seg ega ing si es; :
a io (Ka/Ks) be ween non-synonymous subs i u ions (Ka) and synonymous subs i u ions (Ks). M8-M8a LRT: likelihood a io es among he model M8 and M8a o
e alua e posi i e selec ion.
doi:10.1371/jou nal.pone.0024041. 002
O ganiza ion and E olu ion o My icin-C Genes
PLoS ONE | www.plosone.o g 6 Augus 2011 | Volume 6 | Issue 8 | e24041
high, ha DGGE pa e ns we e unique o each mussel
i espec i e o i s sex o o igin and only ull-sibs sha ed common
bands in elec opho e ic p o iles [24]. In ou s udy, we con i med
his high a iabili y using bo h agmen analysis and gDNA
sequencing and demons a ed ha la ge indels a e on he basis o
he main size a ian s a bo h in ons. Acco dingly, wo la ge
indels a in on 1 would explain he h ee main size a ian s (445,
480 and 615 bp) epo ed a in onic egions [23]. As epo ed by
hese au ho s [23] and sugges ed by he DGGE analysis [24], we
also iden i ied mino indels and a la ge amoun o single nucleo ide
subs i u ions sca e ed along in ons sequences. No di e ences in
allelic size equency we e de ec ed be ween he wo na u al
popula ions analyzed a in on 1 and in on 2, which ag ees wi h
he e y low gene ic s uc u e o Medi e anean mussel obse ed
p e iously wi h mic osa elli es in his a ea (F
ST
= 0.0122) [36].
The numbe o he FREP (Fib inogen- ela ed p o eins) genes, a
amily o highly a iable hemolymph lec ins in ol ed in non-sel
ecogni ion in B. glab a a [37], was in es iga ed by using Sou he n
blo analysis [1]. In ou s udy, we ook ad an age o he high
a ia ion a in ons o asce ain he numbe o genes o my icin-C
Table 3. Gene ic di e si y pa e n a my icin-C in M. gallop o incialis wi h basic cDNA sequences.
Summa y s a is ics Signal pep ide Ma u e pep ide C- e minal egion All egions
n22222222
Si es 60 120 120 300
S 7 (3, 4) 26 (10, 16) 15 (4, 11) 48 (17,31)
p0.02056 (0.00510) 0.05725 (0.00821) 0.03723 (0.00270) 0.04190 (0.00456)
h
W
0.03200 (0.01551) 0.05944 (0.02237) 0.03429 (0.01401) 0,04389 (0.01558)
D
Tajima
21.1440 (p.0.10)20.4034 (p.0.10)20.1536 (p.0.10)20.4692 (p.0.10)
D
Fu and Li
20.6353 (p.0.10)20.6093 (p.0.10)20.0364 (p.0.10)20.4749 (p.0.10)
H- es 0.0000 (p = 0.3246)
(Ka/Ks) 0.858 1.004 3.068 1.302
M8-M8a LRT (Selec ion) Non Signi ican p,0.05 p,0.05 p,0.001
Abb e ia ions co espond o hose indica ed in Table 2.
doi:10.1371/jou nal.pone.0024041. 003
Figu e 5. Compa ison o agmen analysis in he same indi idual o in on 1 and in on 2 s a ing om my icin-C ampli ied genes
and om o al DNA.
doi:10.1371/jou nal.pone.0024041.g005
O ganiza ion and E olu ion o My icin-C Genes
PLoS ONE | www.plosone.o g 7 Augus 2011 | Volume 6 | Issue 8 | e24041
and hei genomic o ganiza ion in Medi e anean mussel.
Palla icini e al. [23] sugges ed ha a single gene could accoun
o my icin-C a ia ion conside ing ha only wo clus e s we e
iden i ied a each indi idual a e analyzing gene ic ela ionships in
a la ge amoun o cDNA clones. Howe e , he de ec ion in ou
s udy o up o ou size a ian s a speci ic indi iduals o bo h
in ons suppo s he exis ence o a leas wo genes. The amily
analysis pe o med s ongly sugges ed he exis ence o wo closely
linked genes ha could be he consequence o a andem
duplica ion. Cos a e al. [24] also analyzed he pa e n o
inhe i ance o DGGE my icin-C a ian s and obse ed sha ed
banding pa e ns be ween ull-sibs, bu hey could no clea ly ace
back he pa e ns obse ed om o sp ing o pa en s. On he o he
hand, he analysis o a la ge sample o cDNA clones in en mussels
in ou s udy showed ha mo e han wo basic a ian s we e
p esen a speci ic indi iduals, con i ming he necessi y o a leas
wo my icin-C genes o explain in aindi idual a ia ion. This
esul is cohe en wi h he iden i ica ion o h ee di e en genomic
clones in a single mussel [24]. All da a indica e he exis ence o wo
genes andemly o ganized in Medi e anean mussel unde lying he
high molecula di e si y obse ed a my icin-C. This esul
con as s wi h he e y high copy numbe de ec ed in o he
AMP amilies in molluscs (de ensines: 48 copies; p olin- ich: 13
copies) [22], and highligh s ha he molecula a iabili y equi ed
o pa hogen ecogni ion and elimina ion by AMPs in molluscs
may ollow di e en di e si ica ion s a egies.
Despi e wo my icin-C genes a e hypo hesized in ou wo k, only
wo cDNA o gDNA a ian s we e de ec ed in mos indi iduals, a
hi d a ian being de ec ed in some indi iduals bu a e y low
equency in he clones analyzed. The compa ison o agmen
analysis wi h gDNA sequencing showed ha some allelic a ian s
de ec ed in he agmen analysis a speci ic indi iduals we e
missed in he gDNA sequencing analysis, despi e ha in some
indi iduals we e sequenced up o 15 clones. The mos likely
explana ion o hese appa en disc epancies is he lack o co ec
ma ching o my icin-C p ime s in one o he wo hypo hesized
genes. Speci ic peaks in he agmen analysis we e missed when
ampli ied my icin-C genes we e used as aw DNA ma e ial o
in on PCR ampli ica ions, hus suppo ing his explana ion.
Acco ding o he exis ence o wo loci and he high size
a iabili y obse ed a bo h in ons (He be ween 0.8 and 0.9), a
high equency o double he e ozygous indi iduals ( ou di e en
alleles) should be expec ed in mussels. This would be pa icula ly
s essed i bo h loci showed simila a iabili y a in ons and
despi e he p obable game ic disequilib ium occu ing a hese loci
since hei close linkage. Howe e , only 6.8% and 3.2%
indi iduals showed ou alleles a in on 1 in Vigo and Co un˜a,
espec i ely, while 60.7% and 57.9% should be expec ed
Figu e 6. Bayesian ee o basic cDNA sequences o my icin-C. Values on b anches indica e Bayesian pos e io p obabili y (only alues .0.75
a e showed). T ee was oo ed using My icin A (AF162334) and my icin B (AF1623354) sequences as ou g oups.
doi:10.1371/jou nal.pone.0024041.g006
O ganiza ion and E olu ion o My icin-C Genes
PLoS ONE | www.plosone.o g 8 Augus 2011 | Volume 6 | Issue 8 | e24041
acco ding o allelic equencies. The same occu ed a in on 2,
whe e 0% and 3.0% double he e ozygo es we e obse ed in he
same popula ions, while 80.8% and 83.3% should be expec ed.
This is an expec able esul i he hypo hesized my icin-C
duplica ion had aken place ecen ly and mu a ion would ha e
no ime enough o inc easing gene ic a iabili y in he new locus.
Al e na i ely, i is possible ha duplica ion is no ixed in mussel
popula ions and a seg ega ing polymo phism exis s, indi iduals
showing be ween wo and ou alleles.
Mechanisms unde lying my icin-C di e si ica ion
Se e al AMP amilies demons a ed high gene ic a iabili y a
coding egions in in e eb a e [9], [22], [38] and e eb a e
species [39]. In he Medi e anean mussel, Palla icini e al. [20]
demons a ed much highe a iabili y in my icin-C han in o he
AMP amilies. These au ho s also epo ed ha o he non
immune- ela ed housekeeping genes, such as b-ac ine, showed
much lowe gene ic di e si y and only wo single nucleo ide
a ian s we e de ec ed among he 33 sequences analyzed, hus
sugges ing di e si ica ion mechanisms speci ically ac ing on
my icin-C genes. The s udy by Palla icini e al. [23] and ha by
Cos a e al. [24] we e conduc ed o desc ibe my icin-C a iabili y,
a he han o ind ou he unde lying mechanisms o such
a ia ion. In ou s udy, we analyzed a la ge sample o cDNA and
gDNA sequences o asce ain he mechanisms esponsible o
ansc ip omic and genomic my icin-C di e si ica ion, ying o
balance he wo sou ces o sampling a iance, indi iduals and
clones wi hin indi iduals. A key poin o add ess his ype o s udies
is o manage high quali y sequences o ensu e he con idence o he
a ian s de ec ed. Thus, all clones in ou s udy we e sequenced
om bo h 59and 39ends and only high quali y sequences we e
conside ed o u he analysis. As p e iously epo ed [23], a ew
basic sequences (2–3) we e iden i ied a each indi idual bo h a
cDNA and gDNA in ou wo k, which ag ees wi h he low numbe
o my icin-C genes hypo hesized. Howe e , equen single
nucleo ide di e ences we e de ec ed among he di e en copies
o each basic sequence, bo h in cDNA and gDNA analysis,
s ongly sugges ing an o igin due o soma ic mu a ion. A simila
mechanism o soma ic a ia ion was epo ed in FREP pep ides in
he snail B. glab a a [1]. In acco dance wi h i s andom mu a ion
o igin, mos o hese a ian s ep esen ed single ons in he whole
sample and only a ew nucleo ide a ian s we e de ec ed a ex an
a iable si es. These appea ed a highly a iable egions ha could
ep esen mu a ion ho spo s, as epo ed in e eb a e immuno-
globulins [40], [41]. Global soma ic mu a ion a e (1.1*10
23
) was
in he uppe ange o ha desc ibed in mouse IgH immunoglob-
ulins ( om 10
23
o 10
25
) [42]. Soma ic mu a ions a ec ed in a
simila ashion o bo h in ons and exons and, despi e some
ho spo s mu a ion si es occu ed, he whole my icin-C gene
appea ed unde hei in luence. Howe e , soma ic mu a ions
seemed o be une enly dis ibu ed among he my icin-C domains
and a highe a e was obse ed especially a he C- e minal
domain, which nea ly doubled ha obse ed a ma u e pep ide.
This obse a ion should be con i med in a la ge sample, and
sugges s a di e en impac o soma ic mu a ion along my icin-C
gene ha may be ela ed o di e en unc ional cons ain s a i s
di e en domains. Rema kably, poin mu a ion si es wi hin each
indi idual showed la ge di e gence be ween gDNA and cDNA
sequences. This could indica e ha he mechanism o soma ic
mu a ion occu s along all he li e o he indi idual, hus
de e mining la ge in e cellula di e ences wi hin indi iduals.
The obse a ion o my icin-C exp ession a di e en adul issues
(man le, diges i e gland and haemolymph) and e en a la al
s ages and o oci es by Cos a e al. [24] suppo s his explana ion.
Also, mRNA edi ing o he exis ence o speci ic molecula
p ocesses a mRNA o mRNA in e media ies could con ibu e
o he p ocess o di e si ica ion explaining he di e ences obse ed
be ween cDNA and gDNA sequences wi hin indi iduals. A mo e
de ailed s udy on my icin-C di e si ica ion along mussel on ogeny
using la ge cDNA and gDNA samples in a ew indi iduals would
be equi ed o disce n be ween hese hypo heses.
Some my icin-C cDNA sequences appea ed o be cons i u ed by
pieces o wo di e en basic sequences, s ongly sugges ing
c ossing-o e o gene con e sion e en s occu ing a soma ic
issues. Recombinan sequences we e also iden i ied by Zhang
e al. [1] in B. glab a a when analyzing he causes o FREPs
di e si ica ion. Thus, appa en ly, soma ic ecombina ion could
also be occu ing in Medi e anean mussel o gene a e molecula
a iabili y a my icin-C genes. The a e age a e o ecombina ion
pe indi idual was no oo high (0.06460.023), bu hal indi iduals
showed a leas one ecombinan in he sample o clones analyzed.
The ole o soma ic ecombina ion in he genesis o a ia ion in
e eb a e immunoglobulin is well known [43], [44]. On he
con a y, no ole o ecombina ion was sugges ed o explain
a ia ion in o he AMPs like in he amphibian Bombina maxima
[45]. In Medi e anean mussel, ema kably, no ecombina ion
signs we e de ec ed in he gDNA sample, when a ce ain
p opo ion should be expec ed acco ding o ecombina ion a e
a cDNA. This may be a esul o he sho cDNA and gDNA
sample size ega ding he low equency o ecombina ion, bu also
some molecula mechanism speci ically ac ing on mRNA canno
be disca ded.
E olu ion o my icin-C
The e olu ion o AMP gene amilies has been add essed in
di e en species [22], [45]. In he Medi e anean mussel, a
phylogene ic analysis was ca ied ou s a ing om la ge samples
o cDNA sequences [23], [46]. Howe e , hese au ho s included
all gene ic a ian s de ec ed in hei cDNA sample. Acco ding o
ou esul s, much o his a ia ion is o soma ic o igin and
he e o e no subjec ed o e olu iona y agen s. In ou s udy, we
spli he analysis using on one hand all cDNA sequences and on
he o he only he 21 basic sequences iden i ied in he 10
indi iduals analyzed. The ma u e pep ide domain showed he
highes nucleo ide di e si y among he h ee my icin-C domains
bo h in he whole cDNA sample and in he basic cDNA sequences,
which highligh s he ele ance o molecula di e si ica ion a his
my icin-C domain, di ec ly ela ed o pa hogen ecogni ion.
Howe e , he C- e minal egion showed he highes impac o
soma ic mu a ion (nea ly wice han ma u e pep ide), sugges ing
he ele ance o andom mu a ion on his likely in acellula
domain o unknown unc ion. The signal pep ide domain was he
leas a iable one in acco dance wi h i s memb ane ecogni ion
unc ion o ans e ence o imma u e my icin-C in o he e iculum
endoplasmic o u he p ocessing. Ou esul s con as wi h ha
ound by o he au ho s [23], who epo ed a highe a iabili y a
he C- e minal egion. This may be explained by he inclusion o
soma ic a ian s in hei s udy ha , as ou lined abo e, showed a
highe impac on C- e minal egion. Padhi and Ve ghese [46]
ound ha some my icin-C a iable codons could be subjec ed o
posi i e selec ion despi e pu i ying selec ion would explain he
pa e n a mos a iable si es. Acco ding o hese au ho s, we
de ec ed e idences o posi i e selec ion in he ma u e pep ide
domain, likely indica ing selec i e p essu es o pa hogen di e si y
de e mining gene ic di e si ica ion a his egion. Conside ing he
high in insic a iabili y obse ed a ma u e pep ide, he signals o
posi i e selec ion a his domain should be mo e p obably ela ed
o balancing selec ion han o di ec ional one. Sound signals o
O ganiza ion and E olu ion o My icin-C Genes
PLoS ONE | www.plosone.o g 9 Augus 2011 | Volume 6 | Issue 8 | e24041