scieee Open visual document viewer

Chloroplast genes transferred to the nuclear plant genome have adjusted to nuclear base composition and codon usage

Oliver, J.L.; Marín, A.; Martínez Zapater, J.M.

Abstract

During plant evolution, some plastid genes have been moved to the nuclear genome. These transferred genes are now correctly expressed in the nucleus, their products being transported into the chloroplast. We compared the base compositions, the distributions of some dinucleotides and codon usages of transferred, nuclear and chloroplast genes in two dicots and two monocots plant species. Our results indicate that transferred genes have adjusted to nuclear base composition and codon usage, being now more similar to the nuclear genes than to the chloroplast ones in every species analyzed.

Full text

Nucleic Acids Resea ch, Vol. 18, No. 1 Chlo oplas genes ans e ed o he nuclea plan genome ha e adjus ed o nuclea base composi ion and codon usage J.L.Oli e *, A.Ma n1 and J.M.Ma nez-Zapa e 2 Unidad de Gene ica, Facul ad de Ciencias, Uni e sidad de G anada, E-18071-G anada and 1Depa amen o de Gene ica y Bio ecnia, Facul ad de Biolog a, Uni e sidad de Se illa, Ap do. 1095, E-41080 Se ille and 2Depa amen o de P o ec ion Vege al, CIT-INIA, Ca e e a G al. de La Co una Km 7, E-28040-Mad id, Spain Recei ed Oc obe 17, 1989; Re ised and Accep ed No embe 30, 1989 ABSTRACT Du ing plan e olu ion, some plas id genes ha e been mo ed o he nuclea genome. These ans e ed genes a e now co ec ly exp essed in he nucleus, hei p oduc s being anspo ed in o he chlo oplas . We compa ed he base composi ions, he dis ibu ions o some dinucleo ides and codon usages o ans e ed, nuclea and chlo oplas genes in wo dico s and wo monoco s plan species. Ou esul s indica e ha ans e ed genes ha e adjus ed o nuclea base composi ion and codon usage, being now mo e simila o he nuclea genes han o he chlo oplas ones in e e y species analyzed. INTRODUCTION The exis ence o adjus men o base composi ion o di e en genomic G + C con en s was shown in homologous genes and noncoding sequences o mic oo ganisms and mi ochond ial genomes1. Demons a ion o he exis ence o such adjus men has been hinde ed in highe euka yo es because o he compa men aliza ion o hei genomes. Only ecen ly, i has been shown ha he e is a composi ional adjus men o Alu epe i i e sequences ha a e loca ed in di e en human genome compa men s2. Coding sequences ha ha e mo ed be ween di e en genomes can be good candida es o p obe he exis ence o composi ional adjus men and o analyze he in ol ed mechanisms. Such gene mo emen s ha e ecu en ly occu ed along plan e olu ion and mos o he plas id genes ha e been ans e ed o he nuclea genome345'6. Since in mos plan species chlo oplas and nuclea genomes ha e di e en GC con en s7, we hink ha plas id genes ans e ed o he plan nuclea genome can be an excellen model sys em o analyze i such composi ional adjus men exis s and, i so, how i wo ks. We compa ed nucleo ide composi ion and codon usage o nuclea , chlo oplas and nuclea genes encoding chlo oplas p o eins ( ans e ed genes) in wo dico s (pea and obacco) and wo monoco s (whea and maize) species. The dis ibu ions o some ele an dinucleo ides we e also s udied. Resul s indica e ha — a he le el o base composi ion, dinucleo ide dis ibu ion and codon usage — ans e ed genes a e mo e simila o nuclea genes han o chlo oplas ones. We analyzed how hey ha e adjus ed hei base composi ion and codon usage o ha o he nuclea en i onmen . DATA AND METHODS Gene sequences Sequences om nuclea and chlo oplas genes we e e ie ed om he GenBank gene ic sequence da a bank8 ( elease 57), o di ec ly aken om o iginal publica ions. We selec ed he ou species wi h highe numbe s o nuclea and chlo oplas genes sequenced: Nico iana abacum (Solanaceae), Pisum sa i um (Leguminosae), T i icum aes i um and Zea mays (Poaceae). Nuclea genes encoding chlo oplas p o eins can be conside ed as ans e ed chlo oplas genes5-6. Al hough his seems o be a common si ua ion, wo excep ions o nuclea encoded chlo oplas p o eins ha p obably e ol ed om nuclea genes ha e al eady been desc ibed9-10. Table 1 shows he genes we ha e iden i ied as ans e ed genes. A lis o he emaining nuclea and chlo oplas genes used in his s udy is shown in he Appendix. Nucleo ide composi ion da a Be o e nucleo ide composi ion and codon usage we e analyzed, in ons o all genes and he sequence coding o he signal pep ide p esen in ans e ed genes we e emo ed. Nucleo ide si es subjec o silen changes ('silen si es') a e calcula ed acco ding o e e ence 1 [N = A, C, G, o T(U); R = A o G; Y = C o T(U)]: A, hi d posi ions o all codons, plus A in i s posi ions o AGR codons; C, hi d posi ions o all codons, plus C in i s posi ions o CTR and CGR codons; G, hi d posi ions o all codons, minus G in hi d posi ions o ATG and TGG codons; T, hi d posi ions o all codons, plus T in i s posi ions o TTR codons. * To whom co espondence should be add essed 65 66 Nucleic Acids Resea ch Codon usage da a a ginine, leucine and se ine, each wi h wo codon g oups. We ., ,„ . , , „ _ , , . . exclude om die analysis e mina ion codons and single-codon The ollowing s a egy was used o s udy codon usage in nuclea . ••• i_ TT.- 1 m F ... .B. c- . AC A .u . g oups (me hio une and yp ophan). This lea es 21 codon g oups and chlo opas genes. Fi s , we de ine as codon g oups he se s •*. * c en * c A . J o . _,^ • i • . . • J . •_, wi h a o al o 59 codons. Second, we coun codon appea ances o synonymous codons di e ing only in he hi d nucleo ide. . , , ^ u i »• u A —,-•,_, ;. •. in each gene and compu e he ela i e equency o each codon The e is a single codon g oup o each ammo acid, excep o . , .. , ,,. » /-?. . . .• •. A u 6 6 ' in eacn o he C(Xjon g oups ( he coun o ha codon di ided by Table 1. ChJo oplas genes ans e ed o he nucleus o pea (PEA), obacco (TOB), whea (WHT) and maize (MZE) used in his s udy. Sequences we e e ie ed om GenBank (Release 57) o , when no GenBank LOCUS name is speci ied, di ec ly om he indica ed sou ce. GENE GenBank SYMBOL LOCUS PROTEIN Chlo ophyll a/b-binding p o ein Majo ligh ha es ing p o ein AB80 RuBisCo small subuni Fe edoxin-NADP+ induced p o ein Ea ly ligh -induced p o ein Chlo oplas ibosomal p o ein (CL18) (CL24) (CL25) " (CL9) Chlo oplas GAPDH-A Chlo oplas GAPDH-B RuBisCo, small subuni Ace olac a e syn hase Majo chlo ophyll a/b-binding p o ein RuBisCo, small subuni RuBisCo, small subuni Glyce aldehide-3-phospha e dehyd ogenase B Chlo ophyll a/b-binding p o ein (1) Newman, B.J. and G ay, J.C. (1988) Plan Mol. Biol. 10, 511 -520. (2) Kolanus, W., Scha nho s , C, Kiihne, U. and He z eld, F. (1987) Mol. Gen. Gene . 209, 234-239. (3) Gan , J.S. (1988) Cu . Gene . 14, 519-528. (4) Mazu , B.J., Chui, C.F. and Smi h, J. (1987) Plan Physiol. 85, 1110-1117. (5) Ma suoka, M., Kano-Mu akami, Y., Tanaka, Y., Ozeki, Y. and Yamamo o, N. (1987) J. Biochem. 102, 673-676. (6) B inkmann, H., Ma inez, P., Quigley, F., Ma in, W. and Ce , R. (1987) J. Mol. E ol. 26, 320-328. (7) Ma suoka, M., Kano-Mu akami, Y. and Yamamo o, N. (1987) Nucl. Acids Res. 15, 6302. Table. 2 Nucleo ide composi ion and GC con en in eplacemen (RS) and silen (SS) si es (see ex ) o chlo oplas (CP), ans e ed (TF) and nuclea (NUC) genes. Species abb e ia ions a e as in Table 1. PEA: cab 15 cab80 ubpl5 n elip psl8 ps24 ps25 ps9 TOB: gapA gapB bpco als WHT: cab bca MZE: bcs gapB cab PEACAB15 PEACAB80 PEARUBP15 (1) (2) (3) TOBGAPA TOBGAPB TOBRBPCO (4) WHTCAB WHTRBCA (5) (6) (7) GENOME PEA: NUC TF CP TOB: NUC TF CP WHT: NUC TF CP MZE: NUC TF CP TOTAL G+C .43 .44 .41 .45 .48 .41 .57 .59 .38 .57 .67 .42 A .31 .30 .25 .27 .28 .27 .28 .25 .31 .26 .26 .30 T .21 .22 .28 .23 .22 .24 .20 .24 .24 .22 .22 .21 C .19 .19 .20 .21 .19 .20 .27 .20 .19 .25 .21 .19 RS G .29 .29 .27 .29 .31 .29 .25 .31 .26 .27 .31 .30 G+C .48 .48 .47 .50 .50 .49 .52 .51 .45 .52 .52 .49 A .29 .29 .27 .24 .21 .30 .20 .07 .32 .14 .02 .33 T .37 .34 .44 .40 .36 .42 .13 .19 .41 .19 .02 .38 C .20 .19 .18 .23 .25 .17 .42 .48 .15 .40 .67 .18 SS G .14 .18 .11 .13 .18 .11 .25 .26 .12 .27 .29 .11 G+C .34 .37 .29 .36 .43 .28 .67 .74 .27 .67 .96 .29 Nucleic Acids Resea ch 67 he o al o codons in he pe inen codon g oup); his me hod d aws a en ion o he speci ic choices made by he o ganism among di e en op ions ( he synonymous codons) ega dless o he equencies o he di e en amino acids in i s p o eins. Thi d, we compu e he o e all di e ence in codon usage be ween any wo genes h ough a dis ance algo i hm which is a e sion o he 'Manha an me ic' o en used by nume ical axonomis s. The codon usage dis ance be ween genes A and B is simply he sum o he absolu e alues o he di e ences in codon equencies: i=59 D (A,B) = whe e x(i,A) and x(i,B) a e he equencies o he i h codon in genes A and B, espec i ely. RESULTS Simila i y es ima es be ween non-homologous genes Measu emen o simila i y be ween nuclea and chlo oplas genes equi es compa isons o non-homologous sequences. G an ham" p oposed he combina ion o ou indexes, based on codon usage and GC con en , o es ima e he simila i y among non-homologous genes om e y di e en sou ces. When applied o ou da a, his me hod was unable o di e en ia e among he nuclea , ans e ed, and chlo oplas gene se s om each species (da a no shown); only he GC con en o he hi d codon posi ion, aken indi idually, was able o clea ly di e en ia e chlo oplas om nuclea genes in all ou species analyzed (da a no shown). Chlo oplas genes ans e ed o he nuclea genome we e always g ouped among he nuclea genes using his index. Nucleo ide composi ion analysis Changes in nucleo ide composi ion o ans e ed genes a e eloca ion in he nuclea genome we e analyzed by s udying base Table 3. Dis ibu ions o CpG and TpA double s in nuclea (NUC), ans e ed (TF) and chlo oplas (CP) genes o di e en species. The a io o he obse ed o he expec ed equencies o hese dinucleo ides a e gi en. De ia ions om expec a ions we e es ed by Chi-squa e. Gene abb e ia ions a e as in Table 1 and he Appendix. TOB GENE NUC: gapC p -la p -lb p -lc ech pox hau TF: gapA gapB bpco als CP: psaA psaB psbA psbC psbD a pA a pB a pE a pF a pH a pl ps2 ps4 psl4 psl6 poB bcl pe A pe B CpG 0.26 0.59 0.55 0.60 ' 0.56 0.42 ' *** i.** *** 0.39 *** 0.52 *** 0.48 *** 0.48 * 0.65 *** 0.59 *** 0.66 *** 0.64 * 0.60 ** 0.80 0.97 0.99 0.73 1.35 0.85 0.60* 0.94 1.39 1.02 1.35 0.81 * 0.85 0.83 1.10 TpA 0.45 *** 0.84 0.80 gsl 0.86 gs2 0.62 ** 0.75 * 0.67 * 0.35 *** 0.52 *** 0.63 0.74 ** 0.87 0.85 * 0.96 0.95 0.82 0.93 0.94 0.84 0.66 * 1.01 0.97 0.82 0.90 0.46 ** 0.64 0.81 *** 0.88 0.84 1.07 PEA GENE abn2 legJ 0.40 *** 0.25 *** gs3 0.39 *** lecA 0.52 ** ie 0.45 *** adh-1 hsp cab 15 cab80 ubpl5 n elip psl8 ps24 ps25 ps9 a pA cy psbD psbC CpG 0.86 0.59 *** 0.71 0.59 0.55 ' 0.72 ' 0.57 ' 0.39' C* |:** | * 0.43 ** 0.43 *** 0.44 *** 0.50 * 0.50 ** 0.15 *** 0.43 0.56 0.95 0.38 ** 0.96 0.66 * 0.71 * 0.60 TpA 0.65 ** 0.48 *** 0.51 *** 0.55 *** 0.71 0.53 ** 0.39 ** 0.54 *** 0.60** 0.29 *** 0.75 0.36 ** 0.59 ** 0.95 0.82 0.87 1.03 * P < 0.05; ** P < 0.01; *** P < 0.0001 68 Nucleic Acids Resea ch Table 3 (con inued) MZE GENE NUC: ac lG adhlF an eg2R h3 h4 susysG zel9A ze22A ze22B zea20M zea30M b32 gapA pepC gs ca cl TF: bcs gapB cab CP: a bB a bE a pB ps4 ubp psl4 ps8 CpG 0.58 *** 0.69 ** 0.46 *** 0.81 1.08 1.15 0.73 *** 0.34 *** 0.46 *** 0.57 ** 0.34 *** 0.34 *** 0.88 0.70 ** 0.92 1.09 1.11 1.14 0.98 1.17 1.05 0.92 0.70 0.64 1.24 0.93 1.32 0.96 TpA 0.59 *** 0.40 *** 0.62 ** 0.50 * 0.21 * 0.76 0.51 *** 0.68 * 0.79 0.75 0.75 0.72 0.47 ** 0.46 *** 0.48 *** 0.57 0.68 * 0.41 ** 0.91 0.25 *** 0.48 * 0.86 0.79 0.94 0.85 0.89 0.87 0.91 WHT GENE glgB gliABA glum A h3 h4 g' amy em cab bca a p a ps cy cy b ps2 xB CpG 0.37 *** 0.42 *** 0.51 * 1.00 1.19 0.92 0.98 0.96 0.75 * 1.22 0.45 1.19 0.85 1.29 0.96 0.90 TpA 0.31 *** 0.49 *** 1.01 0.44 0.97 0.57 *** 0.59 ** 0.37 0.41 *** 0.59 1.03 0.70* 0.74 * 1.04 0.71 * 0.79 * P < 0.05; ** P < 0.01; *** P < 0.0001 composi ion o silen and eplacemen si es. Table 2 shows he a e age base equencies in silen and eplacemen si es o nuclea , ans e ed and chlo oplas genes o each species. To al G + C o he analyzed genes is highe in he nucleus han in he chlo oplas o all ou species. This di e ence is specially high in he wo monoco s species. Table 2 also shows ha ans e ed genes ha e eached simila GC con en han nuclea genes by inc easing hei GC con en mainly in silen si es. No e ha in ou sample o genes, GC con en a eplacemen si es is e y simila among dico and monoco species bu bo h g oups clea ly di e in hei GC con en a silen si es: GC con en a silen si es is lowe han a eplacemen si es in dico s bu highe in monoco s. Dinucleo ide dis ibu ions Nuclea plan DNA has a high con en o 5-me hylcy osine in bo h he CG dinucleo ide and also in he C(A/T)G inucleo ides12 and CpG me hyla ion is no ound in he chlo oplas genome13. Since chlo oplas genes ans e ed o he nucleus show an inc ease in GC con en (Table 2), we analyzed he dis ibu ion o me hyla ion si es in ans e ed genes o ind ou i he inc ease in GC pa allels an inc ease in he me hyla ion si es a ailable. CTG and CAG inucleo ides we e ound a expec ed equencies in mos genomes (da a no shown). The gene o gene dis ibu ions o he o he me hyla ion a ge , CpG, a e shown in Table 3. Chlo oplas genes gene ally show he expec ed equencies o his dime in he ou species; on he con a y, mos o he nuclea and ans e ed genes show signi ican de iciencies. This is specially ue o dico s, while in monoco s some nuclea genes show he CpG expec ed equencies o e en an excess o his dinucleo ide. TpA is o he dinucleo ide o special ele ance, since a gene al a oidance o his dime has been epo ed in mos genomes14'15'16'17. Table 3 shows i s dis ibu ion o e e y gene. Nuclea and ans e ed genes show e y a iable TpA a ios while chlo oplas genes a e mo e uni o m. Signi ican a oidances o TpA we e mo e o en ound in nuclea and ans e ed genes o he ou species han in chlo oplas genes. Codon usage analysis Codon usage dis ances. A iangula ma ix con aining all pai wise compa isons o codon usage dis ances among all genes was compu ed o each species by means o he dis ance algo i hm desc ibed in he Da a and Me hods sec ion; all pai wise dis ances we e hen ca ego ised in o six g oups (nuclea s nuclea , nuclea s ans e ed, e c.) and he a e age dis ance o each g oup compu ed (Table 4). Codon usage dis ances be ween chlo oplas and nuclea genes a e small in dico s and highe in monoco species. In monoco s his is pa alleled by highe dis ances be ween chlo oplas and ans e ed genes. This me hod gi es a global idea o codon usage in genes om he h ee di e en gene se s bu does no conside he a ia ion in codon usage wi hin genes om he same g oup. Co espondence analysis. To ge a be e ep esen a ion o he a iabili y o codon usage wi hin he di e en g oups o genes, Nucleic Acids Resea ch 69 Table 4. Codon usage dis ances ( S.E.) among nuclea (NUC), ans e ed (TF) and chlo oplas (CP) genes. Species abb e ia ions a e as in Table 1. NUC TF CP 8 1 3 1 2 . 75 . OS . 44 * ± 0 0 0 . 28 . 3 1 . 38 1 5 1 5 . 25 . 77 ± 0 ± 0 . 69 . 49 I I 1 11.81 z 0 PEA . 90 | NUCCP NUC TF CP 1 1 2 1 2 1 3 . 6 1 . 77 . 1 2 x 0 ± 0 ± 0 . 53 . 68 . 24 1 1 14 . 99 . 52 ± 1 ± 0 . 45 . 46 I I10. 46± 0 TOB . 22 | NUCCP NUC TF CP 15 . 13 . 23 . 38 59 73 ± ± ± 0 0 0 . 93 . 84 . 76 1 | 1 1 1 | 23 . 22 . 8 1 I £ 0.SO | 13.94 WHT ±0.73 | NUC NUC TF CP 13 13 22 . 06 . 57 . 18 ± 0 ± 0 I 0 . 32 . 73 . 38 4 29 . 65 ± .84 ± 1 0 . 1 1 . 5 1 I I13. 1 5 ± 0 MZE .97 | NUCCP we pe o med a ac o ial co espondence analysis on a da a ma ix con aining he n-1 codon equencies a each synonymous codon g oup o each gene (see e . 18 o me hodology in ol ed in co espondence analysis o codon usage). Di e ences in codon usage clus e chlo oplas genes sepa a ely om he nuclea ones in monoco (Fig. 1) bu no in dico species (da a no shown). In all he species nuclea and ans e ed genes a e always mixed oge he , being mo e widely sca e ed han he chlo oplas ones. DISCUSSION Nucleo ide composi ion o ans e ed genes GC con en is lowe in he chlo oplas genomes han in he nuclea ones7; Table 2 shows ha a gene le el he di e ences a e mo e ex eme in maize and whea han in he dico species; his is due o he highe GC con en o monoco nuclea genes. Since he e is a co ela ion be ween genomic GC con en and GC le el in he h ee codon posi ions o genes119-20, i would be in e es ing o in es iga e i an adjus men o base composi ion occu ed in ans e ed genes adap ing hem o he high GC con en o he nucleus. Table 2 shows ha chlo oplas genes eloca ed in o he nuclea genome ha e now simila GC con en o nuclea genes. Inc ease in GC has been mo e p onounced a he silen si es han a he eplacemen si es. A silen si es, ans e ed genes show highe GC con en han nuclea genes, while a eplacemen si es GC con en s a e e y simila be ween nuclea and ans e ed glum A • gliABA * bca . glgB 1 • h3 .em • M • any • cab . gi • cy b ps2 • • cy • a ps • xB • a p 1 B I a pB h3 • * cab bcB * * gapB . gs pepC • • gopA eg2R • ca • susyBG • a: • b32 • odhlF ac lG zea20M zea30M • 2B22B ze22A • ubp ps4 • • a bZ • a bB psM • ps8 Figu e 1. Co espondence analysis o codon use in di e en genes o whea and maize. The n-1 codon equencies in each synonymous g oup we e used o each gene. Gene and species abb e ia ions a e as in Table 1 and he Appendix. The plo s o Fl ( e ical) xF2 (ho izon al) ac o s explain he 59% o he a iabili y in codon usage o WHT and he 60% in MZE. (• =nuclea gene; • =chlo oplas gene; *= ans e ed gene). 70 Nucleic Acids Resea ch PEATOB 60 55 50 45 40 35 30 o a a a o a a a CD = 0.21 D (NS) 0.10.2 0.3 0.4 CG0.50.60.7 WHTMZE 80 70 50 40 30 = 0.76 (P ( 0.95) S = 0.19 0.2 0.4 0.6 0.8 CG1.2 1.4 80 70 60 40 30 - 0.93 (P < 0.99) S - 0.27 0.20.4 0.6 CG0.81.2 Figu e 2. Plo s o %(G+C) and CG (Obs/Exp a io) o nuclea and ans e ed genes in each species. genes. Simila adjus men s a silen and eplacemen si es p o oked by GC p essu e ha e been p e iously obse ed when homologous genes and noncoding sequences a e compa ed in bac e ial and mi ochond ial genomes wi h di e en GC con en s (see e e ence 1 and e e ences he ein). The di e ences in GC con en a silen si es be ween monoco s and dico s (Table 2) a e p obably indica ing he exis ence o a highe GC p essu e in he wo monoco s s udied. The ise in GC con en o ans e ed genes o dico s is due o simila inc emen s in bo h bases G and C (Table 2); howe e , in he wo monoco s he ise in GC con en o ans e ed genes has been due p e e en ially o inc eases in C. This is a simila si ua ion o wha was ound ea lie in he genomes o wa m- blooded e eb a es1921. Dis ibu ions o TpA and CpG dinucleo ides The composi ional adjus men o ans e ed genes o he nuclea en i onmen is also e lec ed by a s onge a oidance o he double TpA. This a oidance is commonly ound in mos euka yo ic genomes1415, including he nuclea genomes o plan s16-17. The highe CpG a oidance in nuclea han in chlo oplas genes is p obably due o he exis ence o CpG me hyla ion in he nuclea genome12 bu no in he chlo oplas 13'22. T ans e ed genes also ollow a simila pa e n in CpG a oidance as nuclea genes, hus adjus ing hei base composi ion o ha o he nuclea genome. Nuclea genes o dico species analyzed he e always show a oidances o TpA and CpG dinucleo ides. Howe e , a di e en si ua ion was ound in he wo monoco species analyzed, whe e some genes do no show any a oidance. We hink ha hese di e ences be ween species a e e lec ing he di e en composi ional o ganiza ion o hei genomes. The genomes o obacco and pea a e a mo e homogenous in base composi ion han he genomes o whea and maize20-23. Ha ing his in mind, he lack o CpG sho age in ou ou o i e ans e ed genes o monoco s could be due o he loca ion o hese genes in GC ich ch omosome egions wi h a dec eased disc imina ion agains CpG double s. Be na di e al.24 ha e shown ha CpG sho age dec eases in deg ee when inc easing genomic GC le el in bo h e eb a es and hei i uses. This could also be happening in he nucleus o whea and maize, since in hese wo species — bu no in pea and obacco — we ound ha he CpG double Nucleic Acids Resea ch 71 le el is s ongly co ela ed wi h o e all GC con en o di e en genes (Fig. 2). Codon usage compa isons In all ou species he codon usage o ans e ed genes has consis en ly become undis inguishable om ha o he nucleus whe e hey a e in eg a ed. Table 4 shows ha he codon usage o ans e ed genes is mo e dis an ly ela ed o he chlo oplas genome om which hey de i e han o he nuclea genome in which hey a e now loca ed. This is pa icula ly clea in he wo monoco s species. Since he GC con en s o he dico nuclea genes analyzed a e mo e simila o he chlo oplas ones (Table 2), codon usage does no di e en ia e so well genes om one o he o he compa men . Mul i a ia e analyses shown in Fig. 1 allow o isualize he he e ogenei y p esen o codon usage wi hin he nuclea genomes. The dispe sion ound o nuclea genes con as s wi h he homogenei y ound o he chlo oplas ones. Since cons ain s on codon usage imposed by he aminoacid composi ion o p o eins can be disca ded due o he me hod we used o compu e codon equencies, his is p obably a e lec ion o he composi ional o ganiza ion o he plan genomes20. T ans e ed genes always appea ed mixed wi h he nuclea genes by his analyses (Fig. 1), which again suppo s hei adjus men o he nuclea en i onmen . A s ong bias in codon usage o he nuclea encoded chlo oplas GAPDH o maize has been epo ed6. These au ho s hypo hesize ha he s ong codon bias ound could be a consequence o a selec ion o highe exp essi i i y. Howe e , he same bias is no ound o he same gene in o he species. A simila pa e n o s onge biases in codon usage in monoco s han in dico species ha e also been epo ed17 o wo ans e ed genes, bcs (maize) and cab (whea ). Ou mo e global esul s indica e ha his bias could be he consequence o he inc ease in GC con en ha ans e ed genes ha e su e ed o each he le el o he new hos genome. Since his inc ease is mainly suppo ed by changes in silen si es (Table 2), i p oduces a s ong bias in he codon usage o hese genes. Bias is ex emely high in species like maize whe e di e ences in GC con en be ween chlo oplas and nuclea genome a e e y high. As a gene al conclusion, chlo oplas genes ans e ed o he nucleus seem o ha e adjus ed hei base composi ion, dinucleo ide dis ibu ion and codon usage acco ding o he cha ac e is ics p e ailing in hei new hos genomes, and hus hey beha e as poli e DNA2526. Table 2 e eals he clea end o ans e ed genes o achie e simila GC con en as he nuclea genomes whe e hey a e in eg a ed. Consequen ly codon usage is a ec ed and since changes in nucleo ide sequence a ec , almos exclusi ely, o he silen si es, p o ein sequences can emain mainly unmodi ied. The e o e, he GC inc ease can occu wi hou changing he coding capaci y o ans e ed genes, which a he aminoacid le el a e s ill homologous o hei p oka yo ic coun e pa s5. Because he mosaic o ganiza ion o he euka yo ic genome19'20'24-27, a ce ain le el o a iabili y should be expec ed in base composi ion and codon usage among genes om he nuclea genome and his is in ac ound (Fig. 1). As expec ed om he isocho e o ganiza ion in hese ou species, a ia ion is highe in whea and maize ha show a wide composi ional he e ogenei y20. The dis ibu ions o TpA and CpG double s in ans e ed genes also e lec he condi ions o e e y nuclea genome o genome compa men . I seems clea om hese esul s ha he e is an e olu ion owa ds composi ional homogeniza ion wi hin he di e en compa men s o he nuclea plan genome. I his e olu ion is based on a selec i e ad an age due o imp o ed exp essi i i y o o o he composi ional modi ying mechanisms no ela ed wi h gene exp ession, such as di e en mu a ional bias o DNA polyme ases in ge mline cells28 o a ia ion in mu a ion pa e ns along he eplica ion iming o di e en ch omosomal egions in he ge mline29, is no known a his momen . Expe imen s es ing he exp essi i i y o coding sequences wi h di e en base composi ion and codon usage a e equi ed o elucida e he p esence o any selec i e ad an age. ACKNOWLEDGEMENTS We a e mos g a e ul o D s. M. Ruiz Rejdn and J. Salinas by he c i ical eading o he manusc ip . This wo k was pa ially suppo ed by he DGICYT (PB87-0881) and he INIA (#7556) o he Spanish Go e nmen . REFERENCES 1. Jukes, T.H. and Bhushan, V. (1986) J. Mol. E ol. 24,39-44. 2. Filipski, J., Salinas, J. and Rodie , F. (1989) J. Mol. Biol. 206,563-566. 3. Palme , J.D. (1985) Ann. Re . Gene . 19,325-354. 4. Ma in, W. and Ce , R. (1986) Eu . J. Biochem 159:323-331. 5. Shih, M.-C, Laza , G. and Goodman H.M. (1986) Cell 47,73-80. 6. B inkmann, H., Ma inez, P., Quigley, F., Ma in, W. and Ce , R. (1987) J. Mol. E ol. 26, 320-328. 7. Boud aa, M. (1987) Gene . Sel. E ol. 19,143-154. 8. Bilo sky, H.S., Bu ks, C, Ficke , J.W., Goad, W.B., Lewi e , F.L., Rindone, W.P., Swindell, CD. and Tung, C.S. (1986) Nucl. Acids Res. 14,1-4. 9. Tingey, S.V., Tsai, F.-Y., Edwa ds, J.W., Walke , E.L. and Co uzzi, G.M. (1988) J. Biol. Chem. 263:9551-9657. 10. Vie ling, E., Nagao, R.T., DeRoche , A.E. and Ha is, L.M. (1988) EMBO J. 7, 575-581. 11. G an ham, R. (1978) FEBS U e s 95(1),1-11. 12. G uenbaum, ¥., Na eh-Many, T., Ceda , H. and Razin, A. (1981) Na u e 292, 860-862. 13. Tewa i, K.K. and Wildman, S.G. (1966) Science 153,1269-1271. 14. G an ham, R., G eenland, T., Louail, S., Mouchi oud, D., P a o, J.L., Gouy, M. and Gau ie , C. (1985) Bull. Ins . Pas eu 83, 95-148. 15. Ohno, S. (1988) P oc. Na l. Acad. Sci. USA 85:9630-9634. 16. Boud aa, M. and Pen-in, P. (1987) Nucl. Acids Res. 15,5729-5743. 17. Mu ay, E.E., Lo ze , J. and Ebe le, M. (1989) Nucl. Acids Res. 17,477-498. 18. Holm, L. (1986) Nucl. Acids Res. 14,3075-3087. 19. Bema di, G. and Be na di, G. (1986) J. Mol. E ol. 24,1-11. 20. Salinas, J., Ma assi, G., Mon e o, L.M. and Be na di, G. (1988) Nucl. Acids Res. 16,4269-4285. 21. Ma in, A., Be anpe i , J., Oli e , J.L. and Medina, J.R. (1989) Nucl. Acids Res. 17,6181-6189. 22. Nge np asi si i, J., Kobayashi, H. and Akazawa, T. (1988) P oc. Na l. Acad. Sci. USA 85:4750-4754. 23. Ma assi, G., Mon e o, L.M., Salinas, J. and Bema di, G. (1989) Nucl. Acids Res. 17,5273-5290. 24. Bema di, G., Olo sson, B., Filipski, J., Ze ial, M., Salinas, J., Cuny, G., Meunie -Ro i al, M. and Rodie , F. (1985) Science 228,953-958. 25. Zucke kandl, E. (1986) J. Mol. E ol. 24, 12-27. 26. Holmquis , G.P. (1989) J. Mol. E ol. 28, 469-486. 27. Ao a, S. and Ikemu a, T. (1986) Nucl. Acids Res. 14,6345-6355. 28. Filipski, J. (1987) FEBs Le . 217,184-186. 29. Wol e, K.H., Sha p, P.M. and Li, W.-H. (1989) Na u e 337,283-285. 72 Nucleic Acids Resea ch APPENDIX Lis o he emaining nuclea and chlo oplas genes om pea (PEA), obacco (TOB), whea (WHT) and maize (MZE) used in his s udy. Sequences we e e ie ed om GenBank (Release 57) o , when no GenBank LOCUS name is speci ied, di ec ly om he indica ed sou ce. GENE SYMBOLGenBank LOCUSPROTEIN PEA (nucleus): abn2 legJ gsl gs2 gs3 lecA ie adh-1 hsp PEA (chlo oplas ): a pA cy psbD psbC TOB (nucleus): gapC p -la p -lb p -lc ech pox hau TOB (chlo oplas ): psaA psaB psbA psbC psbD a pA a pB a pE a pF a pH a pl ps2 ps4 psl4 psl6 poB bel pe A pe B WHT (nucleus): gliABA glum A h3 h4 gi amy em WHT (chlo oplas ): a p a ps cy cy b ps2 xB MZE (nucleus): ac lG adhlF PEAABN2 (1) PEAGSR1 PEALECA (2) (3) (4) PEACPATPG PEACPCYF PEACPD2 PEACPD2 TOBGAPC (5) TOBPR1CR TOBECH TOBPXDLF TOBTHAUR TOBCPCG WHTGLGB WHTGLIABA WHTGLUMRA WHTH3 WHTH4 WHTGIR WHTAMYA WHTEMR WHTCPATP WHTCPATPS WHTCPCYF WHTCPCYTB (6) (7) MZEACT1G MZEADH1F Albumin 2 'Mino ' legumin polypep ide Glu amine syn hase 1 2 II II "1 Seed lec in A Vicilin Adh-1 Chlo oplas hsp ATP syn hase subuni a (aa 1 —247) Cy och ome p opep ide PSII D2 p o ein PSII 44kDa eac ion cen e p o ein Cy osolic GAPDH-C Pa hogenesis ela ed p o ein la Pa hogenesis ela ed p o ein lb Pa hogenesis ela ed p o ein lc Endochi inase Lignin- o ming pe oxidase TMV induced p o ein homologous o hauma in PSI P700 apop o ein Al A2 PSII 32kD p o ein " 44kD p o ein " D2 p o ein ATPase alpha subuni " be a subuni " epsilon subuni " I subuni " III subuni " a subuni Ribosomal p o ein S2 " p o ein S4 " p o ein SI4 " p o ein SI6 RNA polyme ase be a subuni RuBisCo la ge subuni Cy och ome b Gamma-gliadin B Alpha-be a-gliadin A-II High-M- glu en polypep ide H3 his one H4 his one Gibbe ellin esponsi e whea gene Alpha-amylase EM p o ein ATP syn hase p o on- ansloca ing subuni ATP syn hase CF-0 subuni I p epep ide Cy och ome Cy och ome b-559 (aa 1-83) Ribosomal p o ein S2 xB gene Ac in 1 Alcohol dehyd ogenase (ADH1-F) Nucleic Acids Resea ch 73 APPENDIX (con inued) GENE SYMBOLGenBank LOCUSPROTEIN an eg2R h3 h4 susysG zel9A ze22A ze22B zea20M zea30M b32 gapA pepC gs ca cl MZE (chlo oplas ): a bB a bE a pB ps4 ubp psl4 ps8 REFERENCES MZEANT MZEEG2R MZEH3C2 MZEH4C14 MZESUSYSG MZEZE19A MZEZE22A MZEZE22B MZEZEA20M MZEZEA30M (8) (9) MZEPEPCR MZEGST3A (10) (11) MZECPATBE MZECPATPB MZECPRPS4 MZECPRUBP (12) (12) ATP/ADP ansloca o Endospe m glu elin-2 His one H3 His one H4 Suc ose syn hase Zein 19 kD p o ein " 22 kD p o ein " 26.99 kD p o ein " (clone a20) " (clone a30) b-32 p o ein Glyce aldehide-3-phospha e dehyd ogenase A Phosphoenolpy u a e ca boxylase Glu a ion-S- ans e ase GSTIII Ca alase Regula o y cl locus Coupling ac o complex, be a subuni " " " , epsilon subuni ATPase, be a subuni (aa 1-25) Ribosomal p o ein S4 RuBisCo, la ge subuni Ribosomal p o ein L14 Ribosomal p o ein S8 1. Ga ehouse, J.A., Bown, D., Gil oy, J., Le asseu , M., Cas le on, J. and Ellis, T.H.N. (1988) Biochem. J. 250, 15-24. 2. Wa son, M., Lambe , N., Delauney, A., Ya wood, J.N., C oy, R.R.D., Ga ehouse, J.A., W igh , D.J. and Boul e , D. (1988) Biochem. J. 251, 857-864. 3. Llewellyn, D.J., Finnegan, E.J., Ellis, J.G., Dennis, E.S. and Peacock, W.J. (1987) J. Mol. Biol. 195, 115-123. 4. Vie ling, E., Nagao, R.T., DeRoche , A.E. and Ha is, L.M. (1988) EMBO J. 7, 575-581. 5. Ma suoka, M., Yamamo o, N., Kano-Mu akami, Y., Tanaka, Y., Ozeki, Y., Hi ano, H., Kagawa, H., OsHima, M. and Ohashi, Y. (1987) Plan Physiol. 85, 942-946. 6. Hoglund, A.S. and G ay, J.C. (1987) Nucl. Acids Res. 15, 10590. 7. Dunn, P.P.J. and G ay, J.C. (1988) Nucl. Acids Res. 16, 348. 8. Di Fonzo, N., Ha ings, H., B embilla, M., Mo o, M., Soa e, C, Na a o, E., Palau, J., Rhode, W. and Salamini, F. (1988) Mol. Gen. Gene . 212, 481-487. 9. B inkmann, H., Ma inez, P., Quigley, F., Ma in, W. and Ce , R. (1987) J. Mol. E ol. 26, 320-328. 10. Be na ds, L.A., Skadsen, R.W. and Scandalios, J.G. (1987) P oc. Na l. Acad. Sci. USA 84, 6830-6834. 11. Paz-A es, J., Ghosal, D., Wienand, U., Pe e son, P.A. and Saedle , H. (1987) EMBO J. 6, 3553-3558. 12. Ma kmann-Mulisch, U. and Sub amanian, R. (1988) Eu . J. Biochem. 170, 507-514.