scieee Science in your language
[en] (orig)

Chloroplast genes transferred to the nuclear plant genome have adjusted to nuclear base composition and codon usage

Abstract

During plant evolution, some plastid genes have been moved to the nuclear genome. These transferred genes are now correctly expressed in the nucleus, their products being transported into the chloroplast. We compared the base compositions, the distributions of some dinucleotides and codon usages of transferred, nuclear and chloroplast genes in two dicots and two monocots plant species. Our results indicate that transferred genes have adjusted to nuclear base composition and codon usage, being now more similar to the nuclear genes than to the chloroplast ones in every species analyzed.

Read accessible full text

Chloroplast genes transferred to the nuclear plant genome have adjusted to nuclear base composition and codon usage

Author: Oliver, J.L.; Marín, A.; Martínez Zapater, J.M.
Publisher: Oxford University Press
Year: 1990
DOI: 10.1093/nar/18.1.65
Source: https://idus.us.es/bitstreams/a3e7a856-8702-4755-8f1d-54a6d0e05f2a/download
Nucleic Acids Resea ch, Vol. 18, No. 1
Chlo oplas genes ans e ed o he nuclea plan genome
ha e adjus ed o nuclea base composi ion and codon
usage
J.L.Oli e *, A.Ma n1 and J.M.Ma nez-Zapa e 2
Unidad de Gene ica, Facul ad de Ciencias, Uni e sidad de G anada, E-18071-G anada and
1Depa amen o de Gene ica y Bio ecnia, Facul ad de Biolog a, Uni e sidad de Se illa, Ap do. 1095,
E-41080 Se ille and 2Depa amen o de P o ec ion Vege al, CIT-INIA, Ca e e a
G al.
de La Co una
Km 7, E-28040-Mad id, Spain
Recei ed Oc obe 17, 1989; Re ised and Accep ed No embe 30, 1989
ABSTRACT
Du ing plan e olu ion, some plas id genes ha e been
mo ed o he nuclea genome. These ans e ed genes
a e now co ec ly exp essed in he nucleus, hei
p oduc s being anspo ed in o he chlo oplas . We
compa ed he base composi ions, he dis ibu ions o
some dinucleo ides and codon usages o ans e ed,
nuclea and chlo oplas genes in wo dico s and wo
monoco s plan species. Ou esul s indica e ha
ans e ed genes ha e adjus ed o nuclea base
composi ion and codon usage, being now mo e simila
o he nuclea genes han o he chlo oplas ones in
e e y species analyzed.
INTRODUCTION
The exis ence o adjus men o base composi ion o di e en
genomic
G
+
C
con en s was shown in homologous genes and
noncoding sequences o mic oo ganisms and mi ochond ial
genomes1. Demons a ion o he exis ence o such adjus men
has been hinde ed in highe euka yo es because o he
compa men aliza ion o hei genomes. Only ecen ly, i has been
shown ha he e is a composi ional adjus men o Alu epe i i e
sequences ha a e loca ed in di e en human genome
compa men s2.
Coding sequences ha ha e mo ed be ween di e en genomes
can be good candida es o p obe he exis ence o composi ional
adjus men and o analyze he in ol ed mechanisms. Such gene
mo emen s ha e ecu en ly occu ed along plan e olu ion and
mos o he plas id genes ha e been ans e ed o he nuclea
genome345'6. Since in mos plan species chlo oplas and
nuclea genomes ha e di e en GC con en s7, we hink ha
plas id genes ans e ed o he plan nuclea genome can be an
excellen model sys em o analyze i such composi ional
adjus men exis s and, i so, how i wo ks.
We compa ed nucleo ide composi ion and codon usage o
nuclea , chlo oplas and nuclea genes encoding chlo oplas
p o eins ( ans e ed genes) in wo dico s (pea and obacco) and
wo monoco s (whea and maize) species. The dis ibu ions o
some ele an dinucleo ides we e also s udied. Resul s indica e
ha
—
a he le el o base composi ion, dinucleo ide dis ibu ion
and codon usage
—
ans e ed genes a e mo e simila o nuclea
genes han o chlo oplas ones. We analyzed how hey ha e
adjus ed hei base composi ion and codon usage o ha o he
nuclea en i onmen .
DATA AND METHODS
Gene sequences
Sequences om nuclea and chlo oplas genes we e e ie ed
om he GenBank gene ic sequence da a bank8 ( elease 57), o
di ec ly aken om o iginal publica ions. We selec ed he ou
species wi h highe numbe s o nuclea and chlo oplas genes
sequenced: Nico iana abacum (Solanaceae), Pisum sa i um
(Leguminosae), T i icum aes i um and Zea mays (Poaceae).
Nuclea genes encoding chlo oplas p o eins can be conside ed
as ans e ed chlo oplas genes5-6. Al hough his seems o be a
common si ua ion, wo excep ions o nuclea encoded chlo oplas
p o eins ha p obably e ol ed om nuclea genes ha e al eady
been desc ibed9-10. Table 1 shows he genes we ha e iden i ied
as ans e ed genes. A lis o he emaining nuclea and
chlo oplas genes used in his s udy is shown in he Appendix.
Nucleo ide composi ion da a
Be o e nucleo ide composi ion and codon usage we e analyzed,
in ons o
all
genes and he sequence coding o he signal pep ide
p esen in ans e ed genes we e emo ed. Nucleo ide si es
subjec o silen changes ('silen si es') a e calcula ed acco ding
o e e ence 1 [N = A, C, G, o T(U); R = A o G; Y = C
o
T(U)]:
A, hi d posi ions o
all
codons, plus A in i s posi ions
o AGR codons; C, hi d posi ions o all codons, plus C in i s
posi ions o CTR and CGR codons; G, hi d posi ions o all
codons, minus G in hi d posi ions o ATG and TGG codons;
T, hi d posi ions o all codons, plus T in i s posi ions o TTR
codons.
* To whom co espondence should be add essed
65
66 Nucleic Acids Resea ch
Codon usage da a a ginine, leucine and se ine, each wi h wo codon g oups. We
., ,„ . , ,
„
_ , , . . exclude om die analysis e mina ion codons and single-codon
The ollowing s a egy was used o s udy codon usage in nuclea . ••• i_ TT.- 1 m F
... .B. c- . AC A .u . g oups (me hio une and yp ophan). This lea es 21 codon g oups
and chlo opas genes. Fi s , we de ine as codon g oups he se s •*. * c en * c A . J o
. _,^ • i • . . • J . •_,
wi h
a
o al
o
59 codons. Second,
we
coun codon appea ances
o synonymous codons di e ing only
in he
hi d nucleo ide.
. , , ^ u i »• u A
—,-•,_, ;. •. in
each gene
and
compu e
he
ela i e equency
o
each codon
The e
is a
single codon g oup
o
each ammo acid, excep
o . , .. , ,,. » /-?. . . .• •. A u
6 6 ' in eacn o he
C(Xjon g oups
( he
coun
o
ha codon di ided
by
Table 1. ChJo oplas genes ans e ed o he nucleus o pea (PEA), obacco (TOB), whea (WHT) and maize (MZE)
used in his s udy. Sequences we e e ie ed om GenBank (Release 57) o , when no GenBank LOCUS name is speci ied,
di ec ly om he indica ed sou ce.
GENE GenBank
SYMBOL LOCUS PROTEIN
Chlo ophyll a/b-binding p o ein
Majo ligh ha es ing p o ein AB80
RuBisCo small subuni
Fe edoxin-NADP+ induced p o ein
Ea ly ligh -induced p o ein
Chlo oplas ibosomal p o ein (CL18)
(CL24)
(CL25)
" (CL9)
Chlo oplas GAPDH-A
Chlo oplas GAPDH-B
RuBisCo, small subuni
Ace olac a e syn hase
Majo chlo ophyll a/b-binding p o ein
RuBisCo, small subuni
RuBisCo, small subuni
Glyce aldehide-3-phospha e dehyd ogenase B
Chlo ophyll a/b-binding p o ein
(1) Newman, B.J. and G ay, J.C. (1988) Plan Mol. Biol. 10, 511 -520. (2) Kolanus, W., Scha nho s , C, Kiihne, U.
and He z eld, F. (1987) Mol. Gen. Gene . 209, 234-239. (3) Gan , J.S. (1988) Cu . Gene . 14, 519-528. (4) Mazu ,
B.J., Chui, C.F. and Smi h, J. (1987) Plan Physiol. 85, 1110-1117. (5) Ma suoka, M., Kano-Mu akami, Y., Tanaka,
Y., Ozeki, Y. and Yamamo o, N. (1987) J. Biochem. 102, 673-676. (6) B inkmann, H., Ma inez, P., Quigley, F.,
Ma in, W. and
Ce ,
R. (1987) J. Mol. E ol. 26, 320-328. (7) Ma suoka, M., Kano-Mu akami, Y. and Yamamo o,
N.
(1987) Nucl. Acids Res. 15, 6302.
Table. 2 Nucleo ide composi ion and GC con en in eplacemen (RS) and silen (SS) si es (see ex ) o chlo oplas (CP), ans e ed (TF) and nuclea
(NUC) genes. Species abb e ia ions a e as in Table 1.
PEA:
cab 15
cab80
ubpl5
n
elip
psl8
ps24
ps25
ps9
TOB:
gapA
gapB
bpco
als
WHT:
cab
bca
MZE:
bcs
gapB
cab
PEACAB15
PEACAB80
PEARUBP15
(1)
(2)
(3)
TOBGAPA
TOBGAPB
TOBRBPCO
(4)
WHTCAB
WHTRBCA
(5)
(6)
(7)
GENOME
PEA:
NUC
TF
CP
TOB:
NUC
TF
CP
WHT:
NUC
TF
CP
MZE:
NUC
TF
CP
TOTAL
G+C
.43
.44
.41
.45
.48
.41
.57
.59
.38
.57
.67
.42
A
.31
.30
.25
.27
.28
.27
.28
.25
.31
.26
.26
.30
T
.21
.22
.28
.23
.22
.24
.20
.24
.24
.22
.22
.21
C
.19
.19
.20
.21
.19
.20
.27
.20
.19
.25
.21
.19
RS
G
.29
.29
.27
.29
.31
.29
.25
.31
.26
.27
.31
.30
G+C
.48
.48
.47
.50
.50
.49
.52
.51
.45
.52
.52
.49
A
.29
.29
.27
.24
.21
.30
.20
.07
.32
.14
.02
.33
T
.37
.34
.44
.40
.36
.42
.13
.19
.41
.19
.02
.38
C
.20
.19
.18
.23
.25
.17
.42
.48
.15
.40
.67
.18
SS
G
.14
.18
.11
.13
.18
.11
.25
.26
.12
.27
.29
.11
G+C
.34
.37
.29
.36
.43
.28
.67
.74
.27
.67
.96
.29
Nucleic Acids Resea ch 67
he o al o codons in he pe inen codon g oup); his me hod
d aws a en ion o he speci ic choices made by he o ganism
among di e en op ions ( he synonymous codons) ega dless o
he equencies o
he
di e en amino acids in i s p o eins. Thi d,
we compu e he o e all di e ence in codon usage be ween any
wo genes h ough a dis ance algo i hm which is a e sion o
he 'Manha an me ic' o en used by nume ical axonomis s. The
codon usage dis ance be ween genes A and B is simply he sum
o he absolu e alues o he di e ences in codon equencies:
i=59
D (A,B) =
whe e x(i,A) and x(i,B) a e he equencies o he i h codon in
genes A and B, espec i ely.
RESULTS
Simila i y es ima es be ween non-homologous genes
Measu emen o simila i y be ween nuclea and chlo oplas genes
equi es compa isons o non-homologous sequences.
G an ham" p oposed he combina ion o ou indexes, based on
codon usage and GC con en , o es ima e he simila i y among
non-homologous genes om e y di e en sou ces. When applied
o ou da a, his me hod was unable o di e en ia e among he
nuclea , ans e ed, and chlo oplas gene se s om each species
(da a no shown); only he GC con en o
he
hi d codon posi ion,
aken indi idually, was able o clea ly di e en ia e chlo oplas
om nuclea genes in all ou species analyzed (da a no shown).
Chlo oplas genes ans e ed o he nuclea genome we e always
g ouped among he nuclea genes using his index.
Nucleo ide composi ion analysis
Changes in nucleo ide composi ion o ans e ed genes a e
eloca ion in he nuclea genome we e analyzed by s udying base
Table 3. Dis ibu ions o CpG and TpA double s in nuclea (NUC), ans e ed (TF) and chlo oplas (CP) genes o di e en species. The a io
o he obse ed o he expec ed equencies o hese dinucleo ides a e gi en. De ia ions om expec a ions we e es ed by Chi-squa e. Gene
abb e ia ions a e as in Table 1 and he Appendix.
TOB
GENE
NUC:
gapC
p -la
p -lb
p -lc
ech
pox
hau
TF:
gapA
gapB
bpco
als
CP:
psaA
psaB
psbA
psbC
psbD
a pA
a pB
a pE
a pF
a pH
a pl
ps2
ps4
psl4
psl6
poB
bcl
pe A
pe B
CpG
0.26
0.59
0.55
0.60 '
0.56
0.42 '
***
i.**
***
0.39 ***
0.52 ***
0.48 ***
0.48 *
0.65 ***
0.59 ***
0.66 ***
0.64 *
0.60 **
0.80
0.97
0.99
0.73
1.35
0.85
0.60*
0.94
1.39
1.02
1.35
0.81 *
0.85
0.83
1.10
TpA
0.45 ***
0.84
0.80 gsl
0.86 gs2
0.62 **
0.75 *
0.67 *
0.35 ***
0.52 ***
0.63
0.74 **
0.87
0.85 *
0.96
0.95
0.82
0.93
0.94
0.84
0.66 *
1.01
0.97
0.82
0.90
0.46 **
0.64
0.81 ***
0.88
0.84
1.07
PEA
GENE
abn2
legJ
0.40 ***
0.25 ***
gs3 0.39 ***
lecA 0.52 **
ie 0.45 ***
adh-1
hsp
cab 15
cab80
ubpl5
n
elip
psl8
ps24
ps25
ps9
a pA
cy
psbD
psbC
CpG
0.86
0.59 ***
0.71
0.59
0.55 '
0.72 '
0.57 '
0.39'
C*
|:**
| *
0.43 **
0.43 ***
0.44 ***
0.50 *
0.50 **
0.15 ***
0.43
0.56
0.95
0.38 **
0.96
0.66 *
0.71 *
0.60
TpA
0.65 **
0.48 ***
0.51 ***
0.55 ***
0.71
0.53 **
0.39 **
0.54 ***
0.60**
0.29 ***
0.75
0.36 **
0.59 **
0.95
0.82
0.87
1.03
* P < 0.05; ** P < 0.01; *** P <
0.0001
68 Nucleic Acids Resea ch
Table 3 (con inued)
MZE
GENE
NUC:
ac lG
adhlF
an
eg2R
h3
h4
susysG
zel9A
ze22A
ze22B
zea20M
zea30M
b32
gapA
pepC
gs
ca
cl
TF:
bcs
gapB
cab
CP:
a bB
a bE
a pB
ps4
ubp
psl4
ps8
CpG
0.58 ***
0.69 **
0.46 ***
0.81
1.08
1.15
0.73 ***
0.34 ***
0.46 ***
0.57 **
0.34 ***
0.34 ***
0.88
0.70 **
0.92
1.09
1.11
1.14
0.98
1.17
1.05
0.92
0.70
0.64
1.24
0.93
1.32
0.96
TpA
0.59 ***
0.40 ***
0.62 **
0.50 *
0.21 *
0.76
0.51 ***
0.68 *
0.79
0.75
0.75
0.72
0.47 **
0.46 ***
0.48 ***
0.57
0.68 *
0.41 **
0.91
0.25 ***
0.48 *
0.86
0.79
0.94
0.85
0.89
0.87
0.91
WHT
GENE
glgB
gliABA
glum A
h3
h4
g'
amy
em
cab
bca
a p
a ps
cy
cy b
ps2
xB
CpG
0.37 ***
0.42 ***
0.51 *
1.00
1.19
0.92
0.98
0.96
0.75 *
1.22
0.45
1.19
0.85
1.29
0.96
0.90
TpA
0.31 ***
0.49 ***
1.01
0.44
0.97
0.57 ***
0.59 **
0.37
0.41 ***
0.59
1.03
0.70*
0.74 *
1.04
0.71 *
0.79
* P < 0.05; ** P < 0.01; *** P <
0.0001
composi ion o silen and eplacemen si es. Table 2 shows he
a e age base equencies in silen and eplacemen si es o
nuclea , ans e ed and chlo oplas genes o each species. To al
G
+
C
o he analyzed genes is highe in he nucleus han in he
chlo oplas o all ou species. This di e ence is specially high
in he wo monoco s species. Table 2 also shows ha ans e ed
genes ha e eached simila GC con en han nuclea genes by
inc easing hei GC con en mainly in silen si es. No e ha in
ou sample o genes, GC con en a eplacemen si es is e y
simila among dico and monoco species bu bo h g oups clea ly
di e in hei GC con en a silen si es: GC con en a silen si es
is lowe han a eplacemen si es in dico s bu highe in monoco s.
Dinucleo ide dis ibu ions
Nuclea plan DNA has a high con en o 5-me hylcy osine in
bo h he CG dinucleo ide and also in he C(A/T)G
inucleo ides12 and CpG me hyla ion is no ound in he
chlo oplas genome13. Since chlo oplas genes ans e ed o he
nucleus show an inc ease in GC con en (Table 2), we analyzed
he dis ibu ion o me hyla ion si es in ans e ed genes o ind
ou i he inc ease in GC pa allels an inc ease in he me hyla ion
si es a ailable.
CTG and CAG inucleo ides we e ound a expec ed
equencies in mos genomes (da a no shown). The gene o gene
dis ibu ions o he o he me hyla ion a ge , CpG, a e shown
in Table 3. Chlo oplas genes gene ally show he expec ed
equencies o his dime in he ou species; on he con a y,
mos o he nuclea and ans e ed genes show signi ican
de iciencies. This is specially ue o dico s, while in monoco s
some nuclea genes show he CpG expec ed equencies o e en
an excess o his dinucleo ide.
TpA is o he dinucleo ide o special ele ance, since a gene al
a oidance o his dime has been epo ed in mos
genomes14'15'16'17. Table 3 shows i s dis ibu ion o e e y gene.
Nuclea and ans e ed genes show e y a iable TpA a ios
while chlo oplas genes a e mo e uni o m. Signi ican a oidances
o TpA we e mo e o en ound in nuclea and ans e ed genes
o he ou species han in chlo oplas genes.
Codon usage analysis
Codon usage dis ances. A iangula ma ix con aining all
pai wise compa isons o codon usage dis ances among all genes
was compu ed o each species by means o
he
dis ance algo i hm
desc ibed in he Da a and Me hods sec ion; all pai wise dis ances
we e hen ca ego ised in o six g oups (nuclea s nuclea , nuclea
s ans e ed, e c.) and he a e age dis ance o each g oup
compu ed (Table 4). Codon usage dis ances be ween chlo oplas
and nuclea genes a e small in dico s and highe in monoco
species. In monoco s his is pa alleled by highe dis ances be ween
chlo oplas and ans e ed genes. This me hod gi es a global
idea o codon usage in genes om he h ee di e en gene se s
bu does no conside he a ia ion in codon usage wi hin genes
om he same g oup.
Co espondence analysis. To ge a be e ep esen a ion o he
a iabili y o codon usage wi hin he di e en g oups o genes,
Nucleic Acids Resea ch
69
Table 4. Codon usage dis ances
( S.E.)
among nuclea (NUC), ans e ed (TF)
and chlo oplas
(CP)
genes. Species abb e ia ions
a e as in
Table
1.
NUC
TF
CP
8
1
3
1
2
.
75
.
OS
.
44
*
±
0
0
0
.
28
. 3
1
.
38
1
5
1
5
.
25
.
77
±
0
±
0
.
69
.
49
I
I
1
11.81
z
0
PEA
.
90 |
NUCCP
NUC
TF
CP
1
1
2
1
2
1
3
.
6 1
.
77
.
1
2
x 0
± 0
± 0
.
53
.
68
.
24
1
1
14
.
99
.
52
± 1
± 0
.
45
.
46
I
I10.
46± 0
TOB
.
22 |
NUCCP
NUC
TF
CP
15
.
13
.
23
.
38
59
73
±
±
±
0
0
0
.
93
.
84
.
76
1
|
1
1
1
|
23
.
22
. 8
1
I
£
0.SO | 13.94
WHT
±0.73
|
NUC
NUC
TF
CP
13
13
22
.
06
.
57
.
18
± 0
± 0
I 0
.
32
.
73
.
38
4
29
.
65 ±
.84 ±
1
0
.
1
1
.
5
1
I
I13.
1
5
± 0
MZE
.97 |
NUCCP
we pe o med
a
ac o ial co espondence analysis on a da a ma ix
con aining he n-1 codon equencies
a
each synonymous codon
g oup
o
each gene
(see e . 18 o
me hodology in ol ed
in
co espondence analysis
o
codon usage). Di e ences
in
codon
usage clus e chlo oplas genes sepa a ely om he nuclea ones
in monoco (Fig.
1) bu no in
dico species (da a
no
shown).
In all he species nuclea and ans e ed genes a e always mixed
oge he , being mo e widely sca e ed han he chlo oplas ones.
DISCUSSION
Nucleo ide composi ion
o
ans e ed genes
GC con en is lowe in he chlo oplas genomes han in he nuclea
ones7;
Table
2
shows ha
a
gene le el he di e ences
a e
mo e
ex eme
in
maize and whea han
in
he dico species; his
is
due
o he highe GC con en
o
monoco nuclea genes. Since he e
is
a
co ela ion be ween genomic
GC
con en
and GC
le el
in
he h ee codon posi ions
o
genes119-20,
i
would
be
in e es ing
o in es iga e
i
an adjus men
o
base composi ion occu ed
in
ans e ed genes adap ing hem
o he
high
GC
con en
o
he
nucleus. Table
2
shows ha chlo oplas genes eloca ed in o he
nuclea genome ha e now simila GC con en
o
nuclea genes.
Inc ease
in
GC has been mo e p onounced
a
he silen si es han
a
he
eplacemen si es.
A
silen si es, ans e ed genes show
highe GC con en han nuclea genes, while
a
eplacemen si es
GC con en s
a e
e y simila be ween nuclea
and
ans e ed
glum A
• gliABA
* bca
. glgB
1
•
h3 .em
•
M
•
any
•
cab
.
gi
•
cy b
ps2 • • cy
•
a ps • xB
•
a p
1
B
I a pB
h3
• * cab
bcB
* *
gapB
.
gs
pepC
• •
gopA
eg2R
•
ca
• susyBG
•
a:
•
b32
• odhlF
ac lG
zea20M
zea30M •
2B22B
ze22A
•
ubp
ps4 • • a bZ
•
a bB
psM
• ps8
Figu e
1.
Co espondence analysis
o
codon use
in
di e en genes
o
whea and
maize. The
n-1
codon equencies in each synonymous g oup we e used
o
each
gene.
Gene
and
species abb e ia ions a e
as in
Table 1
and he
Appendix.
The
plo s
o Fl
( e ical) xF2 (ho izon al) ac o s explain he 59%
o
he a iabili y
in codon usage
o
WHT and he 60%
in
MZE.
(•
=nuclea gene;
•
=chlo oplas
gene;
*= ans e ed gene).

70 Nucleic Acids Resea ch
PEATOB
60
55
50
45
40
35
30
o
a
a a
o
a
a
a CD
= 0.21
D
(NS)
0.10.2
0.3
0.4
CG0.50.60.7
WHTMZE
80
70
50
40
30
=
0.76
(P
(
0.95)
S
=
0.19
0.2
0.4 0.6 0.8
CG1.2 1.4
80
70
60
40
30
-
0.93
(P
<
0.99)
S
-
0.27
0.20.4
0.6
CG0.81.2
Figu e 2. Plo s o %(G+C) and CG (Obs/Exp a io) o nuclea and ans e ed genes in each species.
genes.
Simila adjus men s a silen and eplacemen si es
p o oked by GC p essu e ha e been p e iously obse ed when
homologous genes and noncoding sequences a e compa ed in
bac e ial and mi ochond ial genomes wi h di e en GC con en s
(see e e ence
1
and e e ences he ein). The di e ences in GC
con en a silen si es be ween monoco s and dico s (Table 2) a e
p obably indica ing he exis ence o
a
highe GC p essu e in he
wo monoco s s udied.
The ise in GC con en o ans e ed genes o dico s is due
o simila inc emen s in bo h bases G and C (Table 2); howe e ,
in he wo monoco s he ise in GC con en o ans e ed genes
has been due p e e en ially o inc eases in C. This is a simila
si ua ion o wha was ound ea lie in he genomes o wa m-
blooded e eb a es1921.
Dis ibu ions o TpA and CpG dinucleo ides
The composi ional adjus men o ans e ed genes o he nuclea
en i onmen is also e lec ed by a s onge a oidance o he
double TpA. This a oidance is commonly ound in mos
euka yo ic genomes1415, including he nuclea genomes o
plan s16-17.
The highe CpG a oidance in nuclea han in chlo oplas genes
is p obably due o he exis ence o CpG me hyla ion in he nuclea
genome12 bu no in he chlo oplas 13'22. T ans e ed genes also
ollow a simila pa e n in CpG a oidance as nuclea genes, hus
adjus ing hei base composi ion o ha o he nuclea genome.
Nuclea genes o dico species analyzed he e always show
a oidances o TpA and CpG dinucleo ides. Howe e , a di e en
si ua ion was ound in he wo monoco species analyzed, whe e
some genes do no show any a oidance. We hink ha hese
di e ences be ween species a e e lec ing he di e en
composi ional o ganiza ion o hei genomes. The genomes o
obacco and pea a e a mo e homogenous in base composi ion
han he genomes o whea and maize20-23. Ha ing his in mind,
he lack o CpG sho age in ou ou o i e ans e ed genes
o monoco s could be due o he loca ion o hese genes in GC
ich ch omosome egions wi h a dec eased disc imina ion agains
CpG double s. Be na di e al.24 ha e shown ha CpG sho age
dec eases in deg ee when inc easing genomic GC le el in bo h
e eb a es and hei i uses. This could also be happening in
he nucleus o whea and maize, since in hese wo species —
bu no in pea and obacco — we ound ha he CpG double
Nucleic Acids Resea ch
71
le el
is
s ongly co ela ed wi h o e all
GC
con en
o
di e en
genes
(Fig. 2).
Codon usage compa isons
In
all
ou species
he
codon usage
o
ans e ed genes
has
consis en ly become undis inguishable om ha
o
he nucleus
whe e hey
a e
in eg a ed. Table
4
shows ha
he
codon usage
o ans e ed genes
is
mo e dis an ly ela ed
o he
chlo oplas
genome om which hey de i e han
o he
nuclea genome
in
which hey
a e now
loca ed. This
is
pa icula ly clea
in he wo
monoco s species. Since
he GC
con en s
o he
dico nuclea
genes analyzed
a e
mo e simila
o he
chlo oplas ones (Table
2),
codon usage does
no
di e en ia e
so
well genes om
one
o
he
o he compa men .
Mul i a ia e analyses shown
in Fig. 1
allow
o
isualize
he
he e ogenei y p esen
o
codon usage wi hin
he
nuclea
genomes.
The
dispe sion ound
o
nuclea genes con as s wi h
he homogenei y ound
o he
chlo oplas ones. Since cons ain s
on codon usage imposed
by he
aminoacid composi ion o p o eins
can
be
disca ded
due o he
me hod
we
used
o
compu e codon
equencies, his
is
p obably
a
e lec ion
o he
composi ional
o ganiza ion
o he
plan genomes20. T ans e ed genes always
appea ed mixed wi h
he
nuclea genes
by
his analyses (Fig.
1),
which again suppo s hei adjus men
o
he nuclea en i onmen .
A s ong bias
in
codon usage
o he
nuclea encoded
chlo oplas GAPDH
o
maize
has
been epo ed6. These au ho s
hypo hesize ha
he
s ong codon bias ound could
be a
consequence
o a
selec ion
o
highe exp essi i i y. Howe e ,
he same bias
is no
ound
o he
same gene
in
o he species.
A simila pa e n
o
s onge biases
in
codon usage
in
monoco s
han
in
dico species ha e also been epo ed17
o wo
ans e ed genes, bcs (maize)
and
cab (whea ).
Ou
mo e global
esul s indica e ha his bias could
be he
consequence
o he
inc ease
in GC
con en ha ans e ed genes ha e su e ed
o
each
he
le el
o he new
hos genome. Since his inc ease
is
mainly suppo ed
by
changes
in
silen si es (Table 2),
i
p oduces
a s ong bias
in he
codon usage
o
hese
genes.
Bias
is
ex emely
high
in
species like maize whe e di e ences
in GC
con en
be ween chlo oplas
and
nuclea genome
a e
e y high.
As
a
gene al conclusion, chlo oplas genes ans e ed
o he
nucleus seem
o
ha e adjus ed hei base composi ion,
dinucleo ide dis ibu ion
and
codon usage acco ding
o he
cha ac e is ics p e ailing
in
hei
new
hos genomes,
and
hus
hey beha e
as
poli e DNA2526. Table
2
e eals
he
clea end
o ans e ed genes
o
achie e simila GC con en
as he
nuclea
genomes whe e hey
a e
in eg a ed. Consequen ly codon usage
is a ec ed
and
since changes
in
nucleo ide sequence a ec , almos
exclusi ely,
o he
silen si es, p o ein sequences
can
emain
mainly unmodi ied. The e o e,
he
GC inc ease
can
occu wi hou
changing
he
coding capaci y
o
ans e ed genes, which
a he
aminoacid le el
a e
s ill homologous
o
hei p oka yo ic
coun e pa s5. Because
he
mosaic o ganiza ion
o
he euka yo ic
genome19'20'24-27,
a
ce ain le el
o
a iabili y should
be
expec ed
in base composi ion
and
codon usage among genes om
he
nuclea genome
and
his
is in
ac ound
(Fig. 1). As
expec ed
om
he
isocho e o ganiza ion
in
hese ou species, a ia ion
is highe
in
whea
and
maize ha show
a
wide composi ional
he e ogenei y20.
The
dis ibu ions
o
TpA
and CpG
double s
in
ans e ed genes also e lec
he
condi ions
o
e e y nuclea
genome
o
genome compa men .
I seems clea om hese esul s ha he e
is an
e olu ion
owa ds composi ional homogeniza ion wi hin
he
di e en
compa men s
o he
nuclea plan genome.
I
his e olu ion
is
based
on a
selec i e ad an age
due o
imp o ed exp essi i i y
o
o
o he composi ional modi ying mechanisms
no
ela ed wi h
gene exp ession, such
as
di e en mu a ional bias
o DNA
polyme ases
in
ge mline cells28
o
a ia ion
in
mu a ion pa e ns
along
he
eplica ion iming
o
di e en ch omosomal egions
in
he
ge mline29,
is no
known
a
his momen . Expe imen s
es ing
he
exp essi i i y o coding sequences wi h di e en base
composi ion
and
codon usage
a e
equi ed
o
elucida e
he
p esence
o any
selec i e ad an age.
ACKNOWLEDGEMENTS
We
a e
mos g a e ul
o D s. M.
Ruiz Rejdn
and J.
Salinas
by
he c i ical eading
o
he manusc ip . This wo k
was
pa ially
suppo ed
by he
DGICYT (PB87-0881)
and he
INIA (#7556)
o
he
Spanish Go e nmen .
REFERENCES
1.
Jukes,
T.H. and
Bhushan,
V.
(1986)
J. Mol.
E ol. 24,39-44.
2.
Filipski,
J.,
Salinas,
J. and
Rodie ,
F.
(1989)
J. Mol.
Biol. 206,563-566.
3.
Palme ,
J.D.
(1985)
Ann. Re .
Gene . 19,325-354.
4.
Ma in,
W. and Ce , R.
(1986)
Eu . J.
Biochem 159:323-331.
5.
Shih,
M.-C,
Laza ,
G. and
Goodman
H.M.
(1986) Cell 47,73-80.
6. B inkmann,
H.,
Ma inez,
P.,
Quigley,
F.,
Ma in,
W. and Ce , R.
(1987)
J. Mol.
E ol.
26,
320-328.
7.
Boud aa,
M.
(1987) Gene .
Sel.
E ol. 19,143-154.
8. Bilo sky,
H.S.,
Bu ks,
C,
Ficke ,
J.W.,
Goad,
W.B.,
Lewi e ,
F.L.,
Rindone,
W.P.,
Swindell,
CD. and
Tung,
C.S.
(1986) Nucl. Acids
Res.
14,1-4.
9. Tingey,
S.V.,
Tsai, F.-Y., Edwa ds,
J.W.,
Walke ,
E.L. and
Co uzzi,
G.M.
(1988)
J.
Biol. Chem. 263:9551-9657.
10.
Vie ling,
E.,
Nagao,
R.T.,
DeRoche ,
A.E. and
Ha is,
L.M.
(1988) EMBO
J. 7,
575-581.
11.
G an ham,
R.
(1978) FEBS U e s 95(1),1-11.
12.
G uenbaum,
¥.,
Na eh-Many,
T.,
Ceda ,
H. and
Razin,
A.
(1981) Na u e
292,
860-862.
13.
Tewa i,
K.K. and
Wildman,
S.G.
(1966) Science 153,1269-1271.
14.
G an ham,
R.,
G eenland,
T.,
Louail,
S.,
Mouchi oud,
D.,
P a o,
J.L.,
Gouy,
M.
and
Gau ie ,
C.
(1985) Bull. Ins . Pas eu
83,
95-148.
15.
Ohno,
S.
(1988) P oc. Na l. Acad.
Sci. USA
85:9630-9634.
16.
Boud aa,
M. and
Pen-in,
P.
(1987) Nucl. Acids
Res.
15,5729-5743.
17.
Mu ay,
E.E.,
Lo ze ,
J. and
Ebe le,
M.
(1989) Nucl. Acids
Res.
17,477-498.
18.
Holm,
L.
(1986) Nucl. Acids
Res.
14,3075-3087.
19.
Bema di,
G. and
Be na di,
G.
(1986)
J. Mol.
E ol.
24,1-11.
20.
Salinas,
J.,
Ma assi,
G.,
Mon e o,
L.M. and
Be na di,
G.
(1988) Nucl. Acids
Res.
16,4269-4285.
21.
Ma in,
A.,
Be anpe i ,
J.,
Oli e ,
J.L. and
Medina,
J.R.
(1989) Nucl. Acids
Res.
17,6181-6189.
22.
Nge np asi si i,
J.,
Kobayashi,
H. and
Akazawa,
T.
(1988) P oc. Na l. Acad.
Sci.
USA
85:4750-4754.
23.
Ma assi,
G.,
Mon e o,
L.M.,
Salinas,
J. and
Bema di,
G.
(1989) Nucl. Acids
Res.
17,5273-5290.
24.
Bema di,
G.,
Olo sson,
B.,
Filipski,
J.,
Ze ial,
M.,
Salinas,
J.,
Cuny,
G.,
Meunie -Ro i al,
M. and
Rodie ,
F.
(1985) Science 228,953-958.
25.
Zucke kandl,
E.
(1986)
J. Mol.
E ol.
24,
12-27.
26.
Holmquis ,
G.P.
(1989)
J. Mol.
E ol.
28,
469-486.
27.
Ao a,
S. and
Ikemu a,
T.
(1986) Nucl. Acids
Res.
14,6345-6355.
28.
Filipski,
J.
(1987) FEBs Le . 217,184-186.
29.
Wol e,
K.H.,
Sha p,
P.M. and Li, W.-H.
(1989) Na u e 337,283-285.
72 Nucleic Acids Resea ch
APPENDIX
Lis o he emaining nuclea and chlo oplas genes om pea (PEA), obacco (TOB), whea (WHT) and maize (MZE)
used in his s udy. Sequences we e e ie ed om GenBank (Release 57) o , when no GenBank LOCUS name is speci ied,
di ec ly om he indica ed sou ce.
GENE
SYMBOLGenBank
LOCUSPROTEIN
PEA (nucleus):
abn2
legJ
gsl
gs2
gs3
lecA
ie
adh-1
hsp
PEA (chlo oplas ):
a pA
cy
psbD
psbC
TOB (nucleus):
gapC
p -la
p -lb
p -lc
ech
pox
hau
TOB (chlo oplas ):
psaA
psaB
psbA
psbC
psbD
a pA
a pB
a pE
a pF
a pH
a pl
ps2
ps4
psl4
psl6
poB
bel
pe A
pe B
WHT (nucleus):
gliABA
glum A
h3
h4
gi
amy
em
WHT (chlo oplas ):
a p
a ps
cy
cy b
ps2
xB
MZE (nucleus):
ac lG
adhlF
PEAABN2
(1)
PEAGSR1
PEALECA
(2)
(3)
(4)
PEACPATPG
PEACPCYF
PEACPD2
PEACPD2
TOBGAPC
(5)
TOBPR1CR
TOBECH
TOBPXDLF
TOBTHAUR
TOBCPCG
WHTGLGB
WHTGLIABA
WHTGLUMRA
WHTH3
WHTH4
WHTGIR
WHTAMYA
WHTEMR
WHTCPATP
WHTCPATPS
WHTCPCYF
WHTCPCYTB
(6)
(7)
MZEACT1G
MZEADH1F
Albumin 2
'Mino ' legumin polypep ide
Glu amine syn hase 1
2
II
II "1
Seed lec in A
Vicilin
Adh-1
Chlo oplas hsp
ATP syn hase subuni a (aa
1
—247)
Cy och ome p opep ide
PSII D2 p o ein
PSII 44kDa eac ion cen e p o ein
Cy osolic GAPDH-C
Pa hogenesis ela ed p o ein la
Pa hogenesis ela ed p o ein lb
Pa hogenesis ela ed p o ein lc
Endochi inase
Lignin- o ming pe oxidase
TMV induced p o ein homologous o hauma in
PSI P700 apop o ein Al
A2
PSII 32kD p o ein
" 44kD p o ein
" D2 p o ein
ATPase alpha subuni
" be a subuni
" epsilon subuni
" I subuni
" III subuni
" a subuni
Ribosomal p o ein S2
" p o ein S4
" p o ein SI4
" p o ein SI6
RNA polyme ase be a subuni
RuBisCo la ge subuni
Cy och ome
b
Gamma-gliadin B
Alpha-be a-gliadin A-II
High-M- glu en polypep ide
H3 his one
H4 his one
Gibbe ellin esponsi e whea gene
Alpha-amylase
EM p o ein
ATP syn hase p o on- ansloca ing subuni
ATP syn hase CF-0 subuni I p epep ide
Cy och ome
Cy och ome b-559 (aa 1-83)
Ribosomal p o ein S2
xB gene
Ac in 1
Alcohol dehyd ogenase (ADH1-F)
Nucleic Acids Resea ch 73
APPENDIX (con inued)
GENE
SYMBOLGenBank
LOCUSPROTEIN
an
eg2R
h3
h4
susysG
zel9A
ze22A
ze22B
zea20M
zea30M
b32
gapA
pepC
gs
ca
cl
MZE (chlo oplas ):
a bB
a bE
a pB
ps4
ubp
psl4
ps8
REFERENCES
MZEANT
MZEEG2R
MZEH3C2
MZEH4C14
MZESUSYSG
MZEZE19A
MZEZE22A
MZEZE22B
MZEZEA20M
MZEZEA30M
(8)
(9)
MZEPEPCR
MZEGST3A
(10)
(11)
MZECPATBE
MZECPATPB
MZECPRPS4
MZECPRUBP
(12)
(12)
ATP/ADP ansloca o
Endospe m glu elin-2
His one H3
His one H4
Suc ose syn hase
Zein 19 kD p o ein
" 22 kD p o ein
" 26.99 kD p o ein
" (clone a20)
" (clone a30)
b-32 p o ein
Glyce aldehide-3-phospha e dehyd ogenase A
Phosphoenolpy u a e ca boxylase
Glu a ion-S- ans e ase GSTIII
Ca alase
Regula o y cl locus
Coupling ac o complex, be a subuni
" " " , epsilon subuni
ATPase, be a subuni (aa 1-25)
Ribosomal p o ein S4
RuBisCo, la ge subuni
Ribosomal p o ein L14
Ribosomal p o ein S8
1.
Ga ehouse, J.A., Bown, D., Gil oy, J., Le asseu , M., Cas le on, J. and Ellis, T.H.N. (1988) Biochem. J. 250, 15-24.
2.
Wa son, M., Lambe , N., Delauney, A., Ya wood, J.N., C oy, R.R.D., Ga ehouse, J.A., W igh , D.J. and Boul e ,
D.
(1988) Biochem. J. 251, 857-864.
3.
Llewellyn, D.J., Finnegan, E.J., Ellis, J.G., Dennis, E.S. and Peacock, W.J. (1987) J. Mol. Biol. 195, 115-123.
4.
Vie ling, E., Nagao, R.T., DeRoche , A.E. and Ha is, L.M. (1988) EMBO J. 7,
575-581.
5.
Ma suoka, M., Yamamo o, N., Kano-Mu akami, Y., Tanaka, Y., Ozeki, Y., Hi ano, H., Kagawa, H., OsHima,
M. and Ohashi, Y. (1987) Plan Physiol. 85, 942-946.
6. Hoglund, A.S. and G ay, J.C. (1987) Nucl. Acids Res. 15, 10590.
7.
Dunn, P.P.J. and G ay, J.C. (1988) Nucl. Acids Res. 16, 348.
8. Di Fonzo, N., Ha ings, H., B embilla, M., Mo o, M., Soa e, C, Na a o, E., Palau, J., Rhode, W. and Salamini,
F.
(1988) Mol. Gen. Gene . 212, 481-487.
9. B inkmann, H., Ma inez, P., Quigley, F., Ma in, W. and
Ce ,
R. (1987) J. Mol. E ol. 26, 320-328.
10.
Be na ds, L.A., Skadsen, R.W. and Scandalios, J.G. (1987) P oc. Na l. Acad. Sci. USA 84, 6830-6834.
11.
Paz-A es, J., Ghosal, D., Wienand, U., Pe e son, P.A. and Saedle , H. (1987) EMBO J. 6, 3553-3558.
12.
Ma kmann-Mulisch, U. and Sub amanian, R. (1988) Eu . J. Biochem. 170, 507-514.