scieee Science in your language
[en] (orig)

Comparative genomics of citric-acid-producing Aspergillus niger ATCC 1015 versus enzyme-producing CBS 513.88

Abstract

The filamentous fungus Aspergillus niger exhibits great diversity in its phenotype. It is found globally, both as marine and terrestrial strains, produces both organic acids and hydrolytic enzymes in high amounts, and some isolates exhibit pathogenicity. Although the genome of an industrial enzyme-producing A. niger strain (CBS 513.88) has already been sequenced, the versatility and diversity of this species compel additional exploration. We therefore undertook whole- genome sequencing of the acidogenic A. niger wild-type strain (ATCC 1015) and produced a genome sequence of very high quality. Only 15 gaps are present in the sequence, and half the telomeric regions have been elucidated. Moreover, sequence information from ATCC 1015 was used to improve the genome sequence of CBS 513.88. Chromosome-level comparisons uncovered several genome rearrangements, deletions, a clear case of strain-specific horizontal gene transfer, and identi- fication of 0.8 Mb of novel sequence. Single nucleotide polymorphisms per kilobase (SNPs/kb) between the two strains were found to be exceptionally high (average: 7.8, maximum: 160 SNPs/kb). High variation within the species was con- firmed with exo-metabolite profiling and phylogenetics. Detailed lists of alleles were generated, and genotypic differences were observed to accumulate in metabolic pathways essential to acid production and protein synthesis. A transcriptome analysis supported up-regulation of genes associated with biosynthesis of amino acids that are abundant in glucoamylase A, tRNA-synthases, and protein transporters in the protein producing CBS 513.88 strain. Our results and data sets from this integrative systems biology analysis resulted in a snapshot of fungal evolution and will support further optimization of cell factories based on filamentous fungi.

Read accessible full text

Comparative genomics of citric-acid-producing Aspergillus niger ATCC 1015 versus enzyme-producing CBS 513.88

Author: Andersen, Mikael R.; Salazar, Margarita P.; Schaap, Peter J.; Vondervoort, Peter J.I. van de; Culley, David; Thykaer, Jette; Frisvad, Jens C.; Nielsen, Kristian F.; Albang, Richard; Albermann, Kaj; Berka, Randy M.; Braus, Gerhard H.; Braus-Stromeyer, Sus
Publisher: Cold Spring Harbor Laboratory Press
Year: 2011
DOI: 10.1101/gr.112169.110
Source: https://idus.us.es/bitstreams/e87961bd-9c9f-4d68-b766-4b6671e10e29/download
Resea ch
Compa a i e genomics o ci ic-acid-p oducing
Aspe gillus nige ATCC 1015 e sus enzyme-p oducing
CBS 513.88
Mikael R. Ande sen,
1
Ma ga i a P. Salaza ,
1,19
Pe e J. Schaap,
2
Pe e J.I. an de Vonde oo ,
3
Da id Culley,
4
Je e Thykae ,
1
Jens C. F is ad,
1
K is ian F. Nielsen,
1
Richa d Albang,
5
Kaj Albe mann,
5
Randy M. Be ka,
6
Ge ha d H. B aus,
7
Susanna A. B aus-S omeye ,
7
Luis M. Co ochano,
8
Ziyu Dai,
4
Pie W.M. an Dijck,
9
Ge ald Ho mann,
10
Linda L. Lasu e,
4
Jon K. Magnuson,
4
Hildega d Menke,
3
Ma in Meije ,
11
Susan L. Meije ,
1
Jakob B. Nielsen,
1
Michael L. Nielsen,
1
Albe J.J. an Ooyen,
3
He man J. Pel,
3
La s Poulsen,
1
Rob A. Samson,
11
Hein S am,
3
Ad ian Tsang,
12
Johannes M. an den B ink,
13
Alex A kins,
14
And ea Ae s,
14
Ha is Shapi o,
14
Jasmyn Pangilinan,
14
Asa Salamo ,
14
Yigong Lou,
14
E ika Lindquis ,
14
Susan Lucas,
14
Jane G imwood,
15
Igo V. G igo ie ,
14
Ch is ian P. Kubicek,
16
Diego Ma inez,
17,18
Noe¨l N.M.E. an Peij,
3
Johannes A. Roubos,
3
Jens Nielsen,
1,19
and Sco E. Bake
4,20
1–18
[Au ho a ilia ions appea a he end o he pape .]
The ilamen ous ungus Aspe gillus nige exhibi s g ea di e si y in i s pheno ype. I is ound globally, bo h as ma ine and
e es ial s ains, p oduces bo h o ganic acids and hyd oly ic enzymes in high amoun s, and some isola es exhibi
pa hogenici y. Al hough he genome o an indus ial enzyme-p oducing A. nige s ain (CBS 513.88) has al eady been
sequenced, he e sa ili y and di e si y o his species compel addi ional explo a ion. We he e o e unde ook whole-
genome sequencing o he acidogenic A. nige wild- ype s ain (ATCC 1015) and p oduced a genome sequence o e y high
quali y. Only 15 gaps a e p esen in he sequence, and hal he elome ic egions ha e been elucida ed. Mo eo e , sequence
in o ma ion om ATCC 1015 was used o imp o e he genome sequence o CBS 513.88. Ch omosome-le el compa isons
unco e ed se e al genome ea angemen s, dele ions, a clea case o s ain-speci ic ho izon al gene ans e , and iden i-
ica ion o 0.8 Mb o no el sequence. Single nucleo ide polymo phisms pe kilobase (SNPs/kb) be ween he wo s ains
we e ound o be excep ionally high (a e age: 7.8, maximum: 160 SNPs/kb). High a ia ion wi hin he species was con-
i med wi h exo-me aboli e p o iling and phylogene ics. De ailed lis s o alleles we e gene a ed, and geno ypic di e ences
we e obse ed o accumula e in me abolic pa hways essen ial o acid p oduc ion and p o ein syn hesis. A ansc ip ome
analysis suppo ed up- egula ion o genes associa ed wi h biosyn hesis o amino acids ha a e abundan in glucoamylase
A, RNA-syn hases, and p o ein anspo e s in he p o ein p oducing CBS 513.88 s ain. Ou esul s and da a se s om
his in eg a i e sys ems biology analysis esul ed in a snapsho o ungal e olu ion and will suppo u he op imiza ion o
cell ac o ies based on ilamen ous ungi.
[Supplemen al ma e ial is a ailable o his a icle. The A. nige ATCC 1015 whole genome sequence has been submi ed o
GenBank (h p://www.ncbi.nlm.nih.go /genbank/) unde accession no. ACJE00000000. The sequence da a om he
phylogeny s udy ha e been submi ed o GenBank unde accession nos. GU296686–GU296739. The mic oa ay da a
om his s udy ha e been submi ed o he NCBI Gene Exp ession Omnibus (GEO) (h p://www.ncbi.nlm.nih.go /geo/)
unde se ies accession no. GSE10983. The dsmM_ANIGERa_coll511030F lib a y and pla o m in o ma ion ha e been
submi ed o GEO unde accession no. GPL6758.]
The sap o ophic ilamen ous ungus Aspe gillus nige is ound
globally and exhibi s a g ea di e si y in i s pheno ype. A. nige has
become one o he majo wo kho ses in indus ial bio echnology,
being e y e icien in p oducing bo h polysaccha ide-deg ading
enzymes (pa icula ly amylases, pec inases, and xylanases) o o -
ganic acids (mainly ci ic acid) in high amoun s. I also has a long
his o y o sa e use (Schus e e al. 2002; an Dijck e al. 2003; an
19
P esen add ess: Sys ems Biology, Depa men o Chemical and
Biological Enginee ing, Chalme s Uni e si y o Technology, SE-41296
Go
¨ ebo g, Sweden.
20
Co esponding au ho .
E-mail sco .b[email p o ec ed]; ax (509) 372-4732.
A icle published online be o e p in . A icle, supplemen al ma e ial, and pub-
lica ion da e a e a h p://www.genome.o g/cgi/doi/10.1101/g .112169.110.
F eely a ailable online h ough he Genome Resea ch Open Access op ion.
21:885–897 Ó2011 by Cold Sp ing Ha bo Labo a o y P ess; ISSN 1088-9051/11; www.genome.o g Genome Resea ch 885
www.genome.o g
Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om
Dijck 2008). Comme cial impo ance is illus a ed by a wo ld
ma ke o indus ial enzymes o nea ly US$ 5 billion in 2009, o
which ilamen ous ungi accoun o oughly hal o he p o-
duc ion (Lube ozzi and Keasling 2008), and a global ci ic acid
p oduc ion o 9 310
6
me ic ons in 2000 (Ka a a and Kubicek
2003).
In 2007, he genome sequence o A. nige s ain CBS 513.88,
used o indus ial enzyme p oduc ion, was published (Pel e al.
2007). This s ain was de i ed om A. nige NRRL 3122, a s ain
de eloped o glucoamylase A p oduc ion by classical mu agenesis
and sc eening me hods ( an Lanen and Smi h 1968). This wo k
ini ia ed a numbe o new genome-based in es iga ions (Sun e al.
2007; Ande sen e al. 2008ab; Ma ens-Uzuno a and Schaap 2008;
Yuan e al. 2008ab) bu did no e eal di e ences be ween ci ic-
acid-p oducing and enzyme-p oducing A. nige s ains (Cullen
2007).
In his s udy, we p esen he nea ly comple e genome se-
quence o he ci ic-acid-p oducing A. nige wild- ype s ain ATCC
1015 and compa e i o he genome sequence o he enzyme-
p oducing s ain CBS 513.88. The gene ic di e si y o hese wo A.
nige s ains was de e mined by applying sys ems biology ools as
well as new bioin o ma ics me hods o examine mul i-le el di -
e ences ha dis inguish he wild- ype ci ic-acid-p oducing s ain
om he mu agenized glucoamylase A–p oducing s ain.
Resul s
Gene al genome s a is ics
The 34.85-Mb genome sequence o A. nige ATCC 1015 was gen-
e a ed using a sho gun app oach and hen u he imp o ed o
a high-quali y assembly o 24 inished con igs sepa a ed by 15 gaps
(including eigh om cen ome ic egions). Genome s a is ics a e
summa ized in Table 1, wi h de ails in Supplemen al Tex 1. The
ull sequence and anno a ions a e a ailable om he Join Genome
Ins i u e (JGI) Genome Po al (h p://genome.jgi-ps .o g/Aspni5)
and om NCBI (accession numbe ACJE00000000).
The genome sequence o A. nige CBS 513.88 (Pel e al. 2007)
was imp o ed using he ATCC 1015 sequence o close 186 con ig
gaps and modi y he gene models associa ed wi h hese gaps (Table 1;
Supplemen al Table 1). The upda ed A. nige CBS 513.88 genome se-
quence is accessible h ough EMBL (accession numbe s AM269948–
AM270415).
We no e a la ge di e ence (2882) in he numbe o called
genes in he wo s ains (Table 1). A ho ough analysis indica es an
o e p edic ion o genes in CBS 513.88 and an unde p edic ion o
genes in ATCC 1015, and 396/510 unique genes in CBS 513.88/
ATCC 1015 (Supplemen al Tex 2; Supplemen al Fig. 1; o de ails,
see Supplemen al Tables 2–9). Fo u he alida ion o he ab-
sence/p esence o indi idual p o eins, we ha e pe o med gDNA
hyb idiza ions, which can be consul ed o e e ence (Supple-
men al Table 16).
Unique genes in bo h s ains sugges ho izon al gene ans e
o be a cause o he amylase hype -p oduce pheno ype
Fo he genes unique o he amylase-p oducing s ain CBS 513.88
(Supplemen al Table 6), he mos no able genes a e wo alpha-
amylases ha a e iden ical o he Aspe gillus o yzae alpha-amylase,
a possible cause o he amylase hype -p oduce pheno ype. We
discuss his in de ail below (Fig. 2). Fu he mo e, we ind h ee
possible polyke ide syn hases, which sugges s a unique seconda y
me aboli e p o ile o his s ain.
Examining he genes ound only in ATCC 1015 (Supple-
men al Table 6) does no poin o any ob ious cause o ci ic acid
hype -p oduc ion, bu ou possible polyke ide syn hases and
a pu a i e NRPS a e ound o be unique o ATCC 1015, sugges ing
his s ain has unique seconda y me aboli es as well. Mul iple e -
ec s o his ype a e, indeed, seen o bo h s ains in analyses below
(Fig. 1; Supplemen al Fig. 10; Table 2).
Syn eny mapping shows 0.5 Mb o ch omosomal
ea angemen s and a whole-a m in e sion
The genomes o he wo s ains a e la gely syn enic (Fig. 1). Fo
example, he 1429 p o ein-encoding genes o CBS 513.88 supe -
con ig An01 ha ha e p edic ed coun e pa s in ATCC 1015 a e,
wi h one excep ion, all mapped o ch omosome 2b. A simila ex-
ample is seen o all bu one o 1486 p o ein-encoding genes on
supe con ig An02. Bo h excep ions encode pu a i e Tan1-like ans-
posases. No wi hs anding, a numbe o signi ican di e ences in
genome con igu a ion exis (Fig. 1). Mo e han 0.5 Mb o addi ional
genome sequence in ATCC 1015 esides in ou la ge egions on
ch omosomes (Ch ) II, III, and V (Fig. 1). The lanking sequences o
h ee o hese addi ional elemen s in ATCC 1015 a e syn enic o
con inuous sequences in CBS 513.88 and a e hus ue di e ences in
genome con igu a ion (Supplemen al Fig. 2; Supplemen al Table 2;
o de ails on Ch III, see nex sec ion). The ou h egion on he le
a m o Ch V in ATCC 1015 could no be e i ied due o a con ig gap
o CBS 513.88. The con ig gap in he igh a m o Ch V in he ge-
nome sequence o ATCC 1015 is spanned by con inuous sequence
in CBS 513.88, con aining 15 p edic ed genes. An o e iew o he
genes ound in he gaps unique o ATCC 1015 can be ound in
Supplemen al Table 2.
O he s iking di e ences include
a la ge in e sion in Ch VIII and he in-
e sion and ansloca ion o a la ge ag-
men be ween he le a ms o Ch III and
Ch VII (Supplemen al Fig. 3). Bo h we e
con i med wi h PCR spanning he b eak-
poin s (da a no shown).
Finally, he p esence o elome e se-
quences in genome da a con i ms an in-
e sion o he comple e igh a m o Ch VI.
Two ex a alpha-amylase-encoding
genes likely ob ained h ough HGT
An unma ched egion iden i ied o he
le a m o Ch III (Fig. 1) spans 72.5 kb o
Table 1. Gene al genome s a is ics o A. nige ATCC 1015 and A. nige CBS 513.88
A. nige ATCC
1015 ( his s udy)
A. nige CBS 513.88
a
(Pel e al. 2007)
A. nige CBS 513.88
b
( his s udy)
Gene models 11,200 14,165 14,082
Genome size (Mb) 34.85 33.93 34.02
P o ein leng h (amino acids) 484.3 439.9 442.5
Exons pe gene 3.1 3.6 3.6
Exon leng h (bp) 480.8 370.0 371.6
In on leng h (bp) 93.8 97.2 96.9
Excep genome sizes and he numbe o gene models, all alues a e a e ages.
a
The genome assembly published by Pel e al. (2007).
b
Genome assembly o A. nige CBS 513.88 a e gap closu e using sequence in o ma ion om ATCC
1015.
886 Genome Resea ch
www.genome.o g
Ande sen e al.
Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om
unique sequence in ATCC 1015. Rema kably, o CBS 513.88,
a unique 85.3-kb sequence is ound a his loca ion (Fig. 2A). Gene
anno a ion may be ound in Supplemen al Table 10. To con i m
he ac ual absence o he egion in he ch omosome o ATCC 1015,
we pe o med gDNA hyb idiza ions, which did indeed suppo his
inding (Supplemen al Fig. 4).
App oxima ely 12 kb o his egion sha es >99.8% sequence
iden i y a DNA le el wi h genomic DNA om A. o yzae RIB40
(Fig. 2B). The comple e 85-kb agmen including he 12-kb alpha-
amylase egion also appea s o be in ol ed in a ecen duplica ion
ecombina ion e en o Ch VII, (Fig. 2C). Thus, he CBS 513.88
genome ha bo s wo addi ional alpha-amylase-encoding genes
ha a e o hologs o he alpha-amylase-encoding genes
(AO090023000944, AO090120000196) o A. o yzae RIB40. These
indings sugges ha CBS 513.88 and he pa en al s ain NRRL
3122 (da a no shown) acqui ed hese duplica e alpha-amylase
genes (An12g06930, An05g02100) h ough ho izon al gene
ans e (HGT). The occu ence o HGT in Aspe gilli would u he
be suppo ed by he p esence o alpha-amylase-encoding genes in
o he black Aspe gilli ha display >99% nucleo ide iden i y o hose
o A. o yzae RIB40 and he A. nige CBS 513.88 alpha-amylases
(Ko man e al. 1990; Shibuya e al. 1992). The o igin o he
Figu e 1. Syn eny map o he con igs o A. nige ATCC 1015 o he supe con igs o A. nige CBS 513.88. The colo ing o he ch omosomes shows
syn enic egions in A. nige CBS 513.88. A abic nume als show he numbe o he supe con ig in A. nige CBS 513.88. G ay a eas show egions no ound in
he CBS 513.88 genome sequence (Pel e al. 2007). (Filled black ci cles) P oposed loca ions o cen ome ic egions. Sequenced elome es a e ma ked wi h
a T. Ze oes ma k he i s base o he con igs. A black line unde nea h a sec ion o he ch omosomes deno es in e ed sequence. Black his og ams show
SNPs pe kilobase (numbe o single nucleo ide polymo phisms/kilobase) be ween he sequences o he wo s ains (y-axis: 0–160 SNPs/kb). Gaps
be ween con igs and cen ome es a e no o scale. The alignmen demons a es almos comple e syn eny be ween he wo s ains, wi h he excep ion o
a c oss-o e e en be ween he le a ms o ch omosomes III and VII. An o e iew o he genes ound in he gaps unique o ATCC 1015 can be ound in
Supplemen al Table 2.
Aspe gillus nige s ain e olu ion
Genome Resea ch 887
www.genome.o g
Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om
emaining pa o he unma ched egion emains unclea . No sig-
ni ican simila i y was obse ed a he DNA le el, and only a ew o
he encoded p o eins show simila i y wi h o he p o eins p esen
in he NR p o ein da abase.
The e is also e idence ha his HGT may ha e occu ed by
he ac ion o ansposases. The 12-kb HGT egion is lanked by
202-bp in e ed, pe ec , e minal epea s (ITR) (Fig. 2A,B). Fu -
he mo e, his HGT egion ha bo s ano he gene, An12g07000,
ha is iden ical o he A. o yzae npA ansposase gene. In-
e es ingly, in he A. o yzae RIB40genome, hesame12-kb ag-
men has been duplica ed wice and is p esen on mul iple
ch omosomes, bu only one pe ec copy o he ITRs is e ained
(Fig. 2C).
T ansposon p esence is s ain-speci ic
T ansposon-like sequences we e iden i ied in bo h genomes and
quan i ied (Supplemen al Table 11). This compa ison pinpoin s
a ema kable di e ence in he p esence and amoun o ansposon-
ela ed sequences in ATCC 1015 and CBS 513.88 bo h o class I
and o class II ansposons: Only 16 sequences we e iden i ied in
ATCC 1015, bu 55 we e de ec ed in CBS 513.88. Whe eas he
ATCC 1015 genome con ains no class I supe amily copia and
a single copy o class II supe amily Fo 1/pogo, hese sequences a e
much mo e abundan in CBS 513.88, whe e 15 class I copia, and
13 class II Fo 1/pogo a e ound.
SNP analysis e eals high mu a ion a es
and hype a iable egions
We ound 8 616 SNPs/kb (a e age 6s anda d de ia ion) and
a maximum o 163 SNPs/kb single-nucleo ide polymo phisms
(SNPs) be ween A. nige s ains ATCC 1015 and CBS 513.88 (Sup-
plemen al Table 12). This alue is much highe han he maximum
o 9 SNPs/kb ound by Cuomo e al. (2007) in a SNP analysis o
wo Fusa ium g aminea um s ains. A compa ison o he A. nige
ATCC 1015 genome sequence o ha o A. nige ATCC 9029
e ealed ma kedly less a ia ion be ween he wo (2 65 SNPs/kb)
(Supplemen al Fig. 5), indica ing a la ge genomic a ia ion in he
A. nige g oup bu li le be ween ATCC 1015 and 9029. These
polymo phisms a e no uni o mly dis ibu ed bu clus e in hype -
a iable egions (Fig. 1). This is suppo ed by a gDNA hyb idiza ion
s udy, which con i ms he p esence o dis inc hype a iable egions
(Supplemen al Fig. 4).
Gene compa isons e eal indus ially ele an
s ain di e ences
To iden i y s ain-de ining sys emic e ec s o he genome a ia-
ion, esul s om he Imp in gene syn eny analysis (see Supple-
men al Table 5; Me hods) we e used o addi ional genome-scale
in es iga ion.
GO e m o e - ep esen a ion analysis was pe o med on gene
g oups wi h dis inc common p ope ies ( o G oup o e iew, see
Supplemen al Table 5; o de ails, see Supplemen al Table 13). Mos
in e es ing is he obse a ion ha 37 p o eins in ol ed in an-
sc ip ional egula ion a e non unc ional in CBS 513.88 due o
ameshi s o s op codons. This sugges s a less s ingen egula ion
o CBS 513.88 ela i e o ATCC 1015. A compa a i e shake- lask
s udy o he wo s ains also sugges s mu a ions in egula o y ele-
men s as ni ogen sou ce u iliza ion is impai ed in he CBS 513.88
s ain (da a no shown). While he ni ogen ca abolism egula o
A eA was o iginally hough o be mu a ed in CBS 513.88 (Pel e al.
2007), a esequencing o a eA showed he ORFs o be iden ical.
To de ec po en ial di e ences in he me abolism o he wo
s ains, we compa ed all genes encoding p o eins ha display
di e ences a he amino acid le el be ween he wo s ains on he
me abolic ne wo k o A. nige (Ande sen e al. (2008a) and mapped
hem o me abolic pa hways (Supplemen al Fig. 6). Mu a ions
we e ound in he pa hways o biosyn hesis o p oline, aspa a e,
aspa agine, yp ophan, and his idine, which may be ele an o
p o ein p oduc ion. Also, mu a ions we e ound in he plasma
memb ane-bound ATPase, in he enzymes o he GABA shun , o
he TCA cycle, and in componen s o all s eps o he elec on
anspo chain, which could be ele an o he p oduc ion o
ci ic acid.
Phylogene ic analysis con i ms high a iabili y be ween
genome sequences
To place he compa ison o A. nige ATCC 1015 and CBS 513.88
in o he la ge con ex o he Aspe gillus sec ion Nig i, a ia ion in
he wo s ains was analyzed by phylogene ic analysis o a pa o
he be a- ubulin sequence o a numbe o A. nige s ains and o he
black Aspe gilli (Fig. 3A). To u he explo e he genome a iabili y
ac oss he A. nige species, we iden i ied ou 1-kb egions om
Ch II, Ch IV, Ch VI, and Ch VIII ha we e iden ical in he ATCC
1015 and ATCC 9029 s ains bu had ;20 SNPs/kb ela i e o he
genome sequenceo CBS 513.88. These egions we e PCR-sequenced
in se en A. nige s ains including he CBS 513.88 p ogeni o , NRRL
Table 2. Exo-me abolomic p o iling o 11 A. nige s ains based on HPLC-DAD-FLD and HPLC-DAD-HRMS ( his s udy)
Seconda y me aboli es Re e ence
NRRL
3122
CBS
513.88
a
CBS
126.49
ATCC
1015
a
ATCC
9029
a
CBS
554.65
NRRL
328
NRRL
350
NRRL
511
NRRL
1278
NRRL
2270
Au aspe one B Tanaka e al. 1966 • • • • • • •••••
Fumonisin B
2
F is ad e al. 2007b • • • • • •••••
Funalenone Inokoshi e al. 1999 • • • • •••••
Ko anin Bu
¨chi e al. 1971 • • • • •••••
Nig agillin Isogai e al. 1975 ••
Och a oxin A Aba ca e al. 1994 •• •
O landin Cu le e al. 1979 • • • • •••••
O he naph ho-g-py ones • • • • •••••
Py anonig in A Hio e al. 2004 • • • • • • •••••
Tensidol B Fukada e al. 2006 • • • • • • •••••
The e e ence column ela es o he elucida ion o he compound s uc u e.
a
Genome-sequenced s ains.
Ande sen e al.
888 Genome Resea ch
www.genome.o g
Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om
3122, and he A. nige neo ype s ain, CBS 554.65 (Kozakiewicz
e al. 1992) o o m a phylogene ic ee wi h high a ia ion (Fig.
3B). Fo p e iously sequenced s ains, he esequencing esul s
we e iden ical o he p e iously de e mined genomic sequence,
he eby excluding ha he genomic a ia ion is an a i ac o low-
ideli y sequencing.
All A. nige s ains ha e an iden ical sequence o be a- ubulin
and o m a s ongly suppo ed e minal clade oge he wi h As-
pe gillus awamo i (Fig. 3A). The o he axonomically di e en spe-
cies a e phylogene ically dis inc om his g oup, which is in ac-
co dance wi h he se ial classi ica ion o F is ad e al. (2007a). The
la e clus e ing based on he selec ed egions (Fig. 3B) was able o
Figu e 2. Ho izon al gene ans e o alpha-amylase genes om A. o yzae o A. nige CBS 513.88. (A) Unma ched egion iden i ied o he le a m o
ch omosome III spanning 65 kb o ATCC 1015 encoding 30 p edic ed genes and 85 kb o CBS 513.88 encoding 24 p edic ed genes. The unma ched
egion is lanked on one side by a small local in e sion o h ee p edic ed genes (in g een). (B) The 12.4-kb HGT egion is pa o an iden ical 12.7-kb egion
p esen in A. o yzae RIB40 supe con igs SC113 and SC023. In A. nige CBS 513.88, he ans e ed egion is enclosed by a 203-bp in e ed epea ( ed
a ow). An12g07000 (blue) is a pu a i e ansposase and iden ical o he A. o yzae ansposon Ao 1 npA gene. Simila ly, he A. nige CBS 513.88 alpha-
amylase encoding gene An12g06930 (o ange) is iden ical o A. o yzae anno a ed genes AO090120000196 (SC113) and AO090023000944 (SC023). (C)
P oposed duplica ion– ecombina ion e en be ween supe con ig 12 and supe con ig 05 o he 12-kb HGT egion. B eakpoin s a e indica ed wi h do ed
lines. The agmen encoding alpha-amylase An05g02200 is iden ical o An12g06930 (o ange). The b eakpoin s a e lanked wi h addi ional copies o he
203-bp epea egion ( ed a ow). The egion encoding genes An12g06940 o An12g06970 is iden ical o he egion encoding An05g02210 o
An05g02130. The downs eam b eakpoin occu ed in he p edic ed gene coding egion o An12g06970, he eby dele ing he o iginal s a codon and
ups eam egion.
Aspe gillus nige s ain e olu ion
Genome Resea ch 889
www.genome.o g
Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om

sepa a e he A. nige s ains in o h ee g oups. S ain CBS 513.88
and i s p ogeni o NRRL 3122 we e iden ical o he egions, in-
dica ing limi ed SNP in oduc ion due o classical s ain imp o e-
men , and hus con i ms high a ia ion in he A. nige g oup.
Exo-me abolomic p o iling desc ibes h ee dis inc g oups
o A. nige
To u he compa e he wo sequenced s ains on a sys em-wide,
bu nongenomic le el, o each o he and o o he membe s o he
A. nige species g oup, exo-me abolomic p o iles o he wo s ains
we e compa ed o nine o he s ains, including he CBS 513.88
p ogeni o s ain, NRRL 3122; he widely used labo a o y s ain
ATCC 9029; and he A. nige neo ype, CBS 554.65 (Table 2). This
ga e h ee dis inc clades, which ollow he clus e ing o Figu e 3B,
bu does no con o m o he phylogene ic dis ances calcula ed.
S ain CBS 513.88 and i s p ogeni o s ain had e y simila sec-
onda y me aboli e p o iles. These wo s ains di e ed om he
o he s ains analyzed. S ain ATCC 1015 had a p o ile simila o
se en o he s ains, al hough he e we e some quan i a i e di -
e ences. S ain CBS 126.49 had a unique p o ile.
In he case o och a oxin A, i is in e es ing ha in ATCC 1015
and ATCC 9029, a emnan o he PKS gene (An15g07920) o he
pu a i e och a oxin clus e in A. nige CBS 513.88 (Pel e al. 2007)
was iden i ied, which could be ela ed o a 21-kb dele ion in bo h
s ains (Supplemen al Fig. 7). We ha e no ed ha a pu a i e NRPS
is ound in one o he unique ch omosome egions ound in ATCC
1015 (Fig. 1), which may accoun o some o he di e ence
(Supplemen al Table 2).
T ansc ip ome analysis o A. nige ATCC 1015 and CBS 513.88
g owing on glucose
To e alua e he e ec o he di e ences in genome sequence on he
physiology o he wo s ains, ba ch cul i a ions in bio eac o s and
a compa a i e ansc ip ome analysis we e pe o med. S ains
ATCC 1015 and CBS 513.88 we e g own unde he same condi-
ions in ba ch cul u es in a glucose-based minimal medium. The
GlaA-p oducing s ain CBS 513.88 p oduced >1.2 g/L GlaA mo e
han ATCC 1015, while p oducing ;1 g/L biomass less. O he
measu ed cha ac e is ics we e simila o he wo s ains (Table 3).
S a is ical analysis showed 4784 signi ican ly (adj. p<0.05) di -
e en ially exp essed genes, wi h an almos equal numbe o genes
up- egula ed in ei he s ain (2431 in CBS 513.88 s. 2353 o
ATCC 1015).
Examining he di e en ial exp ession in he con ex o me -
abolic pa hways (Supplemen al Fig. 8), only he al e na i e oxi-
da i e pa hway has uni o mly highe exp ession in ATCC 1015,
while a subs an ial subse o me abolism was up- egula ed in CBS
513.88, including glycolysis and he TCA cycle. As he speci ic
g ow h a es o he s ains a e simila , we sugges ha his ex a
ac i i y o cen al me abolism p o ides p ecu so s o he highe
p oduc i i y o glucoamylase. We also ind inc eased exp ession
o genes in ol ed in amino acid me abolism, especially he en-
i e biosyn he ic pa hways o h eonine, se ine, and yp ophan.
Figu e 3. (A) Phylogene ic ela ionship o se e al black Aspe gilli based on pa ial sequencing o be a- ubulin. The ee was oo ed o Aspe gillus aculea us
CBS 172.66. (B) Phylogene ic ela ionship o se en s ains o A. nige based on sequencing o 1-kb a iable egions om ou ch omosomes. The ee was
oo ed o he sequence ob ained om Aspe gillus ca bona ius IMI 388653. Clades based on he exo-me abolomic g oupings o Table 2 a e shown. Fo bo h
ees, boo s ap alues abo e 80% o he 1000 pe o med ei e a ions a e shown.
Table 3. S a is ics o ba ch cul i a ions o A. nige ATCC 1015 and
A. nige CBS 513.88
ATCC 1015 CBS 513.88
mRNA (h) 24.5 61.2 40.2 64.2
Biomass (g/L) 5.0 60.1 4.0 60.5
m
max
(h
1
) 0.17 60.01 0.15 60.01
Glucose (g/L) 10.0 60.6 9.5 60.4
Glyce ol (g/L) 0.09 60.02 0.27 60.03
Y
xs
(Cmol/Cmol)
a
0.67 60.03 0.55 60.03
GlaA (g/L)
b
0.24 60.08 1.57 60.23
Ci ic acid (g/L) 0.10 60.12 0.14 60.03
Fe men a ions we e pe o med in biological iplica es o each s ain.
Values a e p esen ed as a e age 6s anda d de ia ion. m
max
and Y
xs
a e
gene al s a is ics o he e men a ions, while he emaining alues a e
speci ic o he ime o sampling o ansc ip ion analysis (see mRNA ow).
GlaA is glucoamylase A.
a
Biomass was con e ed o Cmol using 24.9 g o biomass/Cmol (Nielsen
e al. 2003).
b
One uni o glucoamylase can be assumed o co espond o 25 mgo
p o ein (PESL p o ein assay; Boeh inge Mannheim).
Ande sen e al.
890 Genome Resea ch
www.genome.o g
Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om
In iguingly, analysis o he amino acid composi ion o he glu-
coamylase A p o ein (Supplemen al Fig. 9) e ealed ha GlaA is
a ypical in ha i has a highe con en o speci ically yp ophan,
h eonine, and se ine han 90% o he p edic ed genes o CBS
513.88. Fo all h ee amino acids, he con en is almos wice as
high as he a e age in he composi ion o o al p o ein (Ch is ias
e al. 1975).
We ini ially hough ha his was due o a ameshi mu a-
ion iden i ied in he gene o he gene al amino acid– egula ing
ansc ip ion ac o (c oss-pa hway con ol p o ein) CpcA (Sup-
plemen al Table 5), bu esequencing o his gene showed his
ameshi o be a sequencing e o . The p esence o wo sense
mu a ions was con i med and is cu en ly being in es iga ed.
To u he suppo ha he inc eased p oduc ion o GlaA is
signi ican enough o a ec amino acid biosyn hesis, we used
a genome-scale me abolic model o A. nige (Ande sen e al. 2008a)
o model he wo s ains unde he g ow h condi ions used (Table
3). The compu ed luxes (Supplemen al Table 14) show ha o all
h ee amino acids, he luxes h ough he biosyn he ic pa hways
mus be a leas wice as high in CBS 513.88 compa ed o ATCC
1015 o suppo he inc eased GlaA p oduc ion, hus co espond-
ing well wi h he ansc ip ome esul s.
Fo a b oade analysis o ends e ealed by he ansc ip ome
p o iles, a GO e m o e - ep esen a ion analysis was conduc ed
on he genes signi ican ly up- egula ed in ei he s ain (Supple-
men al Table 15). The analysis con i med ha ai s ele an o
p o ein p oduc ion—such as amino acid biosyn hesis and RNA
aminoacyla ion—appea o be highly signi ican o CBS 513.88.
Fo ATCC 1015, GO biological unc ions o elec on anspo (adj.
p=2.7 310
5
), ca bohyd a e anspo (adj. p=6.8 310
3
), and
o ganic acid anspo (adj. p=0.035) we e signi ican .
An examina ion o exp ession o indi idual genes showed
ha glaA had signi ican ly inc eased exp ession in CBS 513.88 (adj.
p=84 310
6
), bu he old change (3.2) was lowe han he in-
c ease in enzyme ac i i y (six old) (Table 3). In e es ingly, all iden-
i ied RNA-aminoacyl syn hases we e ound o be wo old o six old
up- egula ed in CBS 513.88.
Mapping o gene exp ession o he genome iden i ies
seconda y me aboli e clus e ac i i ies and unde pins
he whole-a m in e sion in ch omosome VI
To examine di e ences in gene exp ession be ween he wo s ains
ela i e o ch omosome posi ions, he log
2
- a ios o he gene ex-
p ession indices om he ansc ip ome analysis we e mapped o
he syn eny diag ams o Figu e 1 (Supplemen al Fig. 10).
The analysis o his mapping iden i ied a numbe o ea-
u es. Fi s , six egions no ound in CBS 513.88 con ain genes
wi h uni o mly highe exp ession in ATCC 1015, suppo ing he
absence o hese egions in CBS 513.88. Second, wo ac i e sec-
onda y me aboli e clus e s we e iden i ied on Ch VIII (con ains
a polyke ide syn hase [ID: 211885] ha is unique o ATCC 1015)
and Ch I (NRPS clus e , ound in bo h s ains [NRPS ID: 43555/
An09g00520], bu appea s only o be ac i e in CBS 513.88). Thi d,
he in e sion o he en i e igh a m o Ch VI (Fig. 1) is u he
suppo ed. The elome ic posi ion e ec , which de ines de-
c eased exp ession in he icini y o elome es, has been de-
sc ibed o Saccha omyces ce e isiae (Go schling e al. 1990). Thus,
i an a m has been in e ed, educed exp ession should be ound
a opposi e ends in ATCC 1015 and CBS 513.88. This is indeed
e lec ed in he log
2
- a io on he igh a m o Ch VI (Supplemen al
Fig. 10).
Discussion
In his s udy, we p o ide and compa e he genomes o wo s ains
o A. nige . These wo s ains ha e di e en pheno ypes: one, he
p edecesso o e icien enzyme-p oducing s ains ha ing un-
de gone some le el o mu agenesis and selec ion, and he o he
a wild- ype pa en s ain o high ci ic-acid-p oducing s ains. This
makes he compa ison in e es ing bo h in e ms o genomic e-
sea ch and indus ial applica ions. We ha e suppo ed he con-
clusions o ou compa ison wi h u he expe imen s, allowing us
o p opose new hypo heses and conclusions wi hin h ee main
a eas: (1) gene ic di e si y o he A. nige g oup, (2) ho izon al gene
ans e in ungi, and (3) ungal bio echnology (discussed sepa-
a ely below).
The di e si y o he wo s ains was explo ed h ough a mul i-
le el compa ison (DNA, ch omosome, gene, and p o ein) be ween
bo h genome a chi ec u es and was suppo ed by mul iple ypes o
expe imen al wo k. In e es ingly, we obse e a ema kable simi-
la i y a he le el o ch omosome, gene o de , and gene iden i y,
bu also no able di e si y was de ec ed and u he explo ed,
namely, dis inc egions wi h high le els o SNPs, a 0.8-Mb di -
e ence in genome size, se e al la ge inse ion/dele ions o up o
200 kb, di e en ansposon popula ions, a majo ansloca ion/
in e sion, he in e sion o an en i e ch omosomal a m, and a la ge
se o unique genes in bo h species (see poin s 1–5 below):
1. Abou 400–500 unique genes we e iden i ied o each s ain,
mos o which a e e enly dis ibu ed o e he ch omosomes.
This indica es s ain e olu ion by a high equency o loss
and/o up ake o gene agmen s. In e es ingly, he opposi e is
seen in wo s ains o he pa hogenic Aspe gillus umiga us,
whe e 80% o he s ain-speci ic genes (143 and 218, espec-
i ely) a e ound in a ew, la ge isola e-speci ic genomic islands
(Fedo o a e al. 2008). Thus, his in a-species inc eased e-
quency o ans e /loss o gene ic elemen s so a seems o be
speci ic o he A. nige g oup. Mo e genome sequences may
p o e i o be p esen in o he Aspe gilli as well.
2. The whole-a m in e sion on Ch VI is a ema kable e en . O e
he cha ac e ized Aspe gilli sequenced genomes, mos desc ibed
in e sions we e isola ed as mu an s, wi h a b eakpoin in
a gene, leading o a pheno ype. Pel e al. (2007) epo ed high
syn eny be ween he cen ome ic pa s o he a ms o A. nige
CBS 513.88 and Aspe gillus nidulans FGSC A4. E en so, in his
case, i is suppo ed by he genome sequence, he p esence o
elome ic sequences, and a de ec able elome ic posi ioning
e ec . F om Figu e 2 in Galagan e al. 2005, i is also clea ha
cen ome ic in e sions mus ha e occu ed in he genealogy o
he common ances o o A. nidulans,A. o yzae, and A. umiga us.
3. The la ge ansloca ion/in e sion e en o Ch III and Ch VII
explains he disc epancies in ch omosome size epo ed by Pel
e al. (2007) be ween s ain N400 (Ve does e al. 1994) and
calcula ed sizes o Ch III and Ch VII o CBS 513.88. This ac
and ha he b eakpoin s a e wi hin wo o he wise in ac genes
indica e ha he e en occu ed a e he b anching o s ains
N400 and NRRL 3122, possibly in he mu agenesis o he p e-
decesso o CBS 513.88.
4. The p esence and unc ionali y o ansposons a e quali a i ely
and quan i a i ely less complex in he genome o ATCC 1015
han in CBS 513.88 (Supplemen al Table 11). Such a di e en
dis ibu ion o ansposons be ween s ains has been obse ed
be o e (B aumann e al. 2008) and was desc ibed o be inducible
by mu agenesis and o he s ess-induced condi ions (S and and
Aspe gillus nige s ain e olu ion
Genome Resea ch 891
www.genome.o g
Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om
McDonald 1985). Howe e , he numbe and scale o di e ences
make a ecen mu agenesis p og am unlikely o be he basis o
he mul iple HGT e en s. Fu he mo e, he de ec ion o epea
sequences in a numbe o DNA egions was accompanied by
s ain-speci ic di e ences. The e o e, we sugges ha he
ansposon popula ions ha e played a signi ican ole in s ain
e olu ion.
5. Finally, he phylogene ic analysis (Fig. 3B) con i ms he high
gene ic a ia ion desc ibed o wo s ains o be gene al wi hin
he A. nige g oup and no he esul o ei he CBS 513.88 o
ATCC 1015 being a ypical. This is again suppo ed by he exo-
me abolomic p o iles o he 11 examined A. nige s ains (Table
2). While he g ouping o he exo-me aboli es does no accu-
a ely e lec he cladog ams in Figu e 3, we see his as con i -
ma ion o he high gene ic a ia ion in he A. nige g oup, as he
p oduc ion o a gi en exo-me aboli e may be changed by
a single genomic e en . Thus, we belie e ha he di e si y o
he wo genome-sequenced s ains is common o he A. nige
g oup, which seems o ha e highly dynamic genomes.
The p esence/absence o HGT in ungi has caused a lo o discus-
sion (Galagan e al. 2005; Keeling and Palme 2008; Khaldi and
Wol e 2008; Khaldi e al. 2008). In his s udy, we ha e shown HGT
o be a mos likely o igin o wo alpha-amylase genes in CBS 513.88
ha migh ha e been ans e ed o A. nige om ano he species,
possibly A. o yzae. The HGT mus ha e occu ed be o e he sepa-
a ion o CBS 513.88 and ATCC 1015. The amylase genes a e no
ound in he ATCC 1015 s ain (also suppo ed by gDNA hyb id-
iza ion), bu hey a e in CBS 513.88 lanked by a ansposon,
which is ound bo h in A. o yzae RIB40 and in a unca ed o m in
ATCC 1015. Wi h he e idence o HGT in o Aspe gillus cla a us
om a Magnapo he- ela ed dono , p esen ed by Khaldi e al.
(2008), HGT is now iden i ied in wo sepa a e cases wi h dis an ly
ela ed species and may hus be seen as a gene al phenomenon in
ilamen ous ungi.
The ini ial eason o compa ing he wo s ains has been o
gain insigh in he gene ic basis o he wo indus ially ele an
pheno ypes. In ci ic-acid-p oducing ATCC 1015, we a e in igued
o see highe exp ession o elemen s om he elec on anspo ,
ca bohyd a e anspo , and o ganic acid anspo . These ac o s
a e known o be in ol ed in ob aining high ci ic acid yields.
Highe exp ession le els e en hough ci ic acid p oduc ion is no
no ably highe (see discussion below) sugges s ha hese ai s may
ha e been con ibu ing ac o s o he o iginal selec ion o ATCC
1015 o ci ic acid p oduc ion. Fo he GlaA-p oducing s ain (CBS
513.88), we obse e sys emic changes: Mu a ions in egula o y
genes (e.g., cpcA), highe exp ession o glaA i sel , and up- egula ion
o all iden i ied RNA-syn hases may all con ibu e o a mo e e i-
cien enzyme p oduce . We also p esen he hypo hesis ha in-
c eased p oduc ion o he amino acids se ine, h eonine, and yp-
ophan may be equi ed o e icien GlaA p oduc ion, as hese a e
o e - ep esen ed in GlaA.
While we ha e iden i ied mul iple possible con ibu o s o
he e icien amylase p oduc ion o CBS 513.88, we ha e only seen
a ew ac o s associa ed wi h ci ic acid p oduc ion. We p opose
mul iple easons o his: (1) CBS 513.88 was selec ed o inc eased
GlaA p oduc ion, while ATCC 1015 is a wild- ype s ain wi h only
modes ci ic acid p oduc ion (compa ed o cu en yields). (2) The
medium used o he ansc ip ome analysis is no op imized o
ci ic acid p oduc ion, as his would no gi e he dispe sed g ow h
necessa y o ep esen a i e sampling (Ka a a and Kubicek 2003).
(3) P o ein syn hesis is a complex p ocess, wi h mo e ac o s in-
ol ed han in he p oduc ion and sec e ion o a simple o ganic
acid. Since mo e han 6000 genes di e a he amino acid le el
be ween he wo s ains, i would be unlikely no o ind many
di e ences in ol ed in p o ein p oduc ion. Fo his eason, we
ha e pu ou emphasis on genes ound in mul iple ypes o anal-
yses, e ec s in en i e pa hways, and ai s (GO) ha a e ound o be
s a is ically o e - ep esen ed, which allows o obus conclusions.
In conclusion, ou esul s es ablish a i m compa a i e geno-
mics ounda ion on which o build and es hypo heses ega ding
enzyme p oduc ion, o ganic acid p oduc ion, and di e si y wi hin
a ‘‘species g oup.’’
Me hods
ATCC 1015 genome assembly
The sequence eads we e de i ed om ou whole-genome sho -
gun (WGS) lib a ies: one wi h an inse size o 2–3 kb, wo wi h an
inse size o 6–8 kb, and one wi h an inse size o 35–40 kb. The
eads we e sc eened o ec o using c oss_ma ch and hen im-
med o ec o sequence and quali y (J Chapman, N Pu nam, I Ho,
and D Rokhsa , unpubl.). Reads sho e han 100 bases a e im-
ming we e excluded. The inal da a se included: 28,551 o 2–3-kb
eads, con aining 21.5 Mb o sequence; 160,479 o 6–8-kb eads,
con aining 123 Mb o sequence; and 38,651 o 35–40-kb eads,
con aining 23.8 Mb o sequence.
The da a we e assembled using elease 2.7 o Jazz, a WGS as-
semble de eloped a he JGI (Apa icio e al. 2002; J Chapman,
N Pu nam, I Ho, and D Rokhsa , unpubl.). The genome size and se-
quence dep h we e ini ially es ima ed o be 36 Mb and 8.0, e-
spec i ely. A e emo al o sho (<1 kb) and edundan sca olds
(<5 kb wi h 80% o mo e o he leng h ma ching a sca old >5 kb),
he assembly included 43.7 Mb o sca old sequence wi h 8.2 Mb
(18.9%) o gaps in 350 sca olds, wi h hal o he sca old sequence
con ained in he eigh la ges sca olds o 1.81 Mb o longe . The
sequence dep h de i ed om he assembly was 7.88 60.05. To
es ima e he comple eness o he assembly, a se o 50,001 ESTs was
BLAT-aligned o he unassembled immed eads, as well as he
assembly i sel . Fo y- h ee housand h ee hund ed and wen y-
h ee ESTs (86.6%) we e >80% co e ed by he unassembled da a;
43,978 (88.0%) we e >50% co e ed; and 44,206 (88.4%) we e >20%
co e ed. By way o compa ison, 48,798 ESTs (97.6%) showed hi s
o he assembly. O ien a ion o he ch omosome a ms was based
on eigh a ailable elome e sequences, he supe con ig o ien a ion
p oposed by Pel e al. (2007), and esea ch on linkage g oups in
A. nige (Bos e al. 1989; Debe se al.1989, 1990a,b; Swa e al. 1992;
Ve does e al. 1994).
ATCC 1015 genome inishing
To pe o m genome imp o emen on A. nige ATCC 1015, ini ial
ead layou s om he whole-genome sho gun assembly we e
con e ed in o he ph ed/ph ap/consed pipeline (Go don e al.
1998). Following manual inspec ion o he assembled sequences,
inishing was pe o med by esequencing plasmid subclones and
by walking on plasmid subclones o osmids using cus om p ime s.
All inishing eac ions we e pe o med wi h 4:1 BigDye o dGTP
BigDye e mina o chemis y (Applied Biosys ems). Repea s in he
sequence we e esol ed by ansposon-hopping 8-kb plasmid
clones. Fosmid clones we e sho gun sequenced and inished o ill
la ge gaps, esol e la ge epea s, o o esol e ch omosome dupli-
ca ions and ex end in o ch omosome elome e egions.
The esul ing 24 inished sca olds we e o ien a ed whe e
possible in o ch omosome s uc u es using compa isons o Pel e al.
(2007) and elome e posi ions. The imp o ed genome consis s o
Ande sen e al.
892 Genome Resea ch
www.genome.o g
Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om
34,853,277 bp wi h an es ima ed e o a e o less han 1 e o in
100,000 bp.
ATCC 1015 au oma ic anno a ion
Gene models in he genome o A. nige we e p edic ed using Fgenesh
(Salamo and Solo ye 2000), Fgenesh+(Salamo and Solo ye
2000), and Genewise (Bi ney and Du bin 2000) in eg a ed in o he
JGI anno a ion pipeline. Fgenesh was ained on a se o mo e han
2000 pu a i e ull-leng h ansc ip s de i ed om clus e ed A. nige
ESTs and eliable homology-based gene models o show 81% sen-
si i i y and 81% speci ici y o p edic ions on a es se . Homology-
based gene p edic o s we e seeded wi h BLASTX alignmen s o
p o eins om he NCBI non edundan se o p o eins. Thi y-one
housand i e hund ed and se en y-eigh A. nige ESTs we e se-
quenced using Sange echnology and we e ei he di ec ly mapped
o genomic sequence when he ESTs included pu a i e ull-leng h
(FL) genes o used o ex end p edic ed gene models in o FL genes
by adding 59and/o 39UTRs. In addi ion, 386,515 ESTs wi h an
a e age leng h o 104 n we e sequenced wi h 454 Li e Sciences
(Roche) GS20 sequence s and used in alida ion o p edic ed gene
models. Since mul iple gene models we e gene a ed o each locus,
a single ep esen a i e model was chosen based on homology and
EST suppo and used o u he analysis.
All p edic ed gene models we e unc ionally anno a ed by
sequence simila i y o anno a ed genes om he NCBI non-
edundan se and o he specialized da abases (such as KEGG)
(Kanehisa e al. 2002, 2004) using BLAST and ha dwa e accele a ed
double-a ine Smi h-Wa e man alignmen s (h p://www. imelogic.
com). Func ional and s uc u al domains we e p edic ed in p o ein
sequences using he In e P o so wa e (Zdobno and Apweile
2001). All genes we e also anno a ed acco ding o Gene On ology
(Ashbu ne e al. 2000; Ha is e al. 2004), euka yo ic o hologous
g oups (KOGs) (Koonin e al. 2004), and KEGG me abolic pa hways
(Kanehisa e al. 2004).
ATCC 1015 sequence a ailabili y
A. nige assemblies, anno a ions, and analyzes a e a ailable
h ough he in e ac i e JGI Genome Po al a h p://genome.jgi-
ps .o g/Aspni5/Aspni51.home.h ml. Genome assemblies oge he
wi h p edic ed gene models and anno a ions we e also deposi ed a
NCBI unde he p ojec accession numbe ACJE00000000.
S ains
The ollowing A. nige s ains we e used o expe imen s (de-
posi ion numbe s in di e en collec ions a e gi en as well): CBS
513.88 =FGSC A1513 =IBT 29270, ATCC 1015 =IBT 28639 =NCTC
3858a =NRRL 1278 =NRRL 350 =NRRL 511 =NRRL 328 =Thom
167 =CBS 113.46 =ATCC 10582 =IMI 031821 =LSBH Ac4 =Thom
3528.7, NRRL 3 =IBT 28539 =MUCL 30480 =DSM 2466 =CECT
2088 =VTT D-85240 =NRRL 566 =WB3 =ATCC 9029 =N400 =
CBS 120.49 =IMI 041876, NRRL 326 =IBT 27876 =WB 326 =CBS
554.65 =ATCC 16888 =IHEM 3415 =IMI 050566 =Thom 2766 =
JCM 10254 =IFO 33023 (ex annin-gallic acid e men a ion;
A. nige neo ype) (Kozakiewicz e al. 1992) NRRL 328 =IBT 27878 =
NRRL 350 =IBT 27877 =CBS 113.46, NRRL 337 =CBS 126.48 =
ATCC 10254 =DSM 734 =IFO 6428 =IMI 015954 =WB 337, NRRL
363 =IBT 3617 =IBT 5764 =CBS 126.49 =ATCC 10698 =IFO 6648,
NRRL 511 =IBT 27875, NRRL 1278 =IBT 27872, NRRL 2270 =IBT
26391 =ATCC 11414 =VTT D-77050 =IMI 075353 =A60 =S.M.
ma in A-1-233 =Wisconsin 72-4, de i ed om ATCC 1015, and
NRRL 3122 =IBT 23538 =ATCC 22343 =CBS 115989 ( his s ain is
a wild- ype p ogeni o o CBS 513.88). S ain his o ies o ATCC
1015, ATCC 9029, and NRRL 3122 a e e iewed by Bake (2006).
Exo-me aboli e p o iling
A. nige s ains we e inocula ed on Czapek yeas au olysa e (CYA),
yeas ex ac suc ose (YES) aga , and CYA wi h 5% NaCl aga (CYAS
aga ). Fo medium composi ion, see F is ad and Samson (2004). All
s ains we e h ee-poin inocula ed on hese media and incuba ed
a 25°C in da kness o a week, a e which i e plugs (6 mm di-
ame e ) along a diame e o a ungal colony we e cu ou and
ex ac ed (Smedsgaa d 1997). The ex ac s we e analyzed by HPLC-
DAD luo escence (F is ad and Th ane 1987; Smedsgaa d 1997)
and by HPLC-HR-MS (Nielsen and Smedsgaa d 2003; F is ad e al.
2007b). The seconda y me aboli es we e iden i ied by compa ison
wi h au hen ic s anda ds ( umonisin B
2
, och a oxin A, nig agillin,
ko anin, and o landin) and by UV spec a and high- esolu ion
mass spec a using elec osp ay ioniza ion.
CBS 513.88 gap closu e
The nea -comple e ATCC 1015 genome sequence was used o
e i y con ig o de and o ien a ion and o es ima e gap sequence
leng h be ween con igs in CBS 513.88. Gap- lanking PCR p ime s
(20-me s) gi ing ise o PCR p oduc s wi h a minimal o e lap o
100 bp we e au oma ically designed wi h p ime -3 (Rozen and
Skale sky 2000) using cus om sc ip s. PCR p oduc s o expec ed
size we e pu i ied using a QIAquick PCR Pu i ica ion Ki (QIAGEN)
and sequenced by Baseclea . The upda ed sequence o A. nige CBS
513.88 is a ailable om EMBL (h p://www.ebi.ac.uk/embl/) un-
de accession nos. AM269948–AM270415.
Sequencing o a iable egions
Regions we e ampli ied using s anda d PCR echniques wi h he
ou ollowing p ime pai s ( he name desc ibes ch omosome
numbe and di ec ion): (1) Ch 02-Fwd: 59-GGACACTGCTTGATG
TGATG-39, Ch 02-Re : 59-GAGAGACGTACGAAAGGTTG-39; (2)
Ch 04-Fwd: 59-CGATCTGCGACCAGGA-39, Ch 04-Re : 59-CATAA
CGGATTCGTCGCTG-39; (3) Ch 06-Fwd: 59-CTTGAAGGCGTTGA
GGTC-39,Ch 06-Re :59-GCGAGTATGTGGCTAACATC-39; (4) Ch 08-
Fwd: 59-GGTATGTCACATTCCRTCCA-39, Ch 08-Re : 59-GCTTGC
AGTGAGCAAGGA-39. The sequences on Ch II, Ch VI, and Ch VIII
o e lap wi h p edic ed genes. The PCR p oduc s we e pu i ied us-
ing a QIAquick PCR Pu i ica ion Ki (QIAGEN) and sequenced by
Agencou Bioscience Co po a ion (GenBank accession numbe
GU296708–GU296739).
Sequencing o be a- ubulin
Ampli ica ion o pa o he be a- ubulin gene was pe o med using
he p ime s B 2a (GGTAACCAAATCGGTGCTGCTTTC) and B 2b
(ACCCTCAGTGTAGTGACCCTTGGC) (Glass and Donaldson 1995).
Bo h s ands o he PCR agmen s we e sequenced wi h he ABI
P ism Big DyeTM Te mina o .3.0 Ready Reac ion Cycle sequenc-
ing ki . Samples we e analyzed on an ABI PRISM 3700 Gene ic An-
alyze , and con igs we e assembled using he o wa d and e e se
sequences wi h he p og am SeqMan om he Lase Gene package
(GenBank accession numbe s GU296686–GU296707).
Calcula ion o phylogene ic ee
Sequences we e aligned using Clus alX and immed o he i s
common base a bo h ends be o e making he inal alignmen . Fo
he ee o Figu e 3B, ou immed sequences we e joined o
each s ain be o e making he inal alignmen . T ee calcula ions
we e pe o med using T eeCon ( an de Pee and de Wach e
1994). Dis ances we e es ima ed using he Jukes and Can o algo-
i hm aking inse ions and dele ions in o accoun . Topology was
Aspe gillus nige s ain e olu ion
Genome Resea ch 893
www.genome.o g
Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om