scieee Open visual document viewer

Comparative genomics of citric-acid-producing Aspergillus niger ATCC 1015 versus enzyme-producing CBS 513.88

Andersen, Mikael R.; Salazar, Margarita P.; Schaap, Peter J.; Vondervoort, Peter J.I. van de; Culley, David; Thykaer, Jette; Frisvad, Jens C.; Nielsen, Kristian F.; Albang, Richard; Albermann, Kaj; Berka, Randy M.; Braus, Gerhard H.; Braus-Stromeyer, Sus

Abstract

The filamentous fungus Aspergillus niger exhibits great diversity in its phenotype. It is found globally, both as marine and terrestrial strains, produces both organic acids and hydrolytic enzymes in high amounts, and some isolates exhibit pathogenicity. Although the genome of an industrial enzyme-producing A. niger strain (CBS 513.88) has already been sequenced, the versatility and diversity of this species compel additional exploration. We therefore undertook whole- genome sequencing of the acidogenic A. niger wild-type strain (ATCC 1015) and produced a genome sequence of very high quality. Only 15 gaps are present in the sequence, and half the telomeric regions have been elucidated. Moreover, sequence information from ATCC 1015 was used to improve the genome sequence of CBS 513.88. Chromosome-level comparisons uncovered several genome rearrangements, deletions, a clear case of strain-specific horizontal gene transfer, and identi- fication of 0.8 Mb of novel sequence. Single nucleotide polymorphisms per kilobase (SNPs/kb) between the two strains were found to be exceptionally high (average: 7.8, maximum: 160 SNPs/kb). High variation within the species was con- firmed with exo-metabolite profiling and phylogenetics. Detailed lists of alleles were generated, and genotypic differences were observed to accumulate in metabolic pathways essential to acid production and protein synthesis. A transcriptome analysis supported up-regulation of genes associated with biosynthesis of amino acids that are abundant in glucoamylase A, tRNA-synthases, and protein transporters in the protein producing CBS 513.88 strain. Our results and data sets from this integrative systems biology analysis resulted in a snapshot of fungal evolution and will support further optimization of cell factories based on filamentous fungi.

Full text

Resea ch Compa a i e genomics o ci ic-acid-p oducing Aspe gillus nige ATCC 1015 e sus enzyme-p oducing CBS 513.88 Mikael R. Ande sen, 1 Ma ga i a P. Salaza , 1,19 Pe e J. Schaap, 2 Pe e J.I. an de Vonde oo , 3 Da id Culley, 4 Je e Thykae , 1 Jens C. F is ad, 1 K is ian F. Nielsen, 1 Richa d Albang, 5 Kaj Albe mann, 5 Randy M. Be ka, 6 Ge ha d H. B aus, 7 Susanna A. B aus-S omeye , 7 Luis M. Co ochano, 8 Ziyu Dai, 4 Pie W.M. an Dijck, 9 Ge ald Ho mann, 10 Linda L. Lasu e, 4 Jon K. Magnuson, 4 Hildega d Menke, 3 Ma in Meije , 11 Susan L. Meije , 1 Jakob B. Nielsen, 1 Michael L. Nielsen, 1 Albe J.J. an Ooyen, 3 He man J. Pel, 3 La s Poulsen, 1 Rob A. Samson, 11 Hein S am, 3 Ad ian Tsang, 12 Johannes M. an den B ink, 13 Alex A kins, 14 And ea Ae s, 14 Ha is Shapi o, 14 Jasmyn Pangilinan, 14 Asa Salamo , 14 Yigong Lou, 14 E ika Lindquis , 14 Susan Lucas, 14 Jane G imwood, 15 Igo V. G igo ie , 14 Ch is ian P. Kubicek, 16 Diego Ma inez, 17,18 Noe¨l N.M.E. an Peij, 3 Johannes A. Roubos, 3 Jens Nielsen, 1,19 and Sco E. Bake 4,20 1–18 [Au ho a ilia ions appea a he end o he pape .] The ilamen ous ungus Aspe gillus nige exhibi s g ea di e si y in i s pheno ype. I is ound globally, bo h as ma ine and e es ial s ains, p oduces bo h o ganic acids and hyd oly ic enzymes in high amoun s, and some isola es exhibi pa hogenici y. Al hough he genome o an indus ial enzyme-p oducing A. nige s ain (CBS 513.88) has al eady been sequenced, he e sa ili y and di e si y o his species compel addi ional explo a ion. We he e o e unde ook whole- genome sequencing o he acidogenic A. nige wild- ype s ain (ATCC 1015) and p oduced a genome sequence o e y high quali y. Only 15 gaps a e p esen in he sequence, and hal he elome ic egions ha e been elucida ed. Mo eo e , sequence in o ma ion om ATCC 1015 was used o imp o e he genome sequence o CBS 513.88. Ch omosome-le el compa isons unco e ed se e al genome ea angemen s, dele ions, a clea case o s ain-speci ic ho izon al gene ans e , and iden i- ica ion o 0.8 Mb o no el sequence. Single nucleo ide polymo phisms pe kilobase (SNPs/kb) be ween he wo s ains we e ound o be excep ionally high (a e age: 7.8, maximum: 160 SNPs/kb). High a ia ion wi hin he species was con- i med wi h exo-me aboli e p o iling and phylogene ics. De ailed lis s o alleles we e gene a ed, and geno ypic di e ences we e obse ed o accumula e in me abolic pa hways essen ial o acid p oduc ion and p o ein syn hesis. A ansc ip ome analysis suppo ed up- egula ion o genes associa ed wi h biosyn hesis o amino acids ha a e abundan in glucoamylase A, RNA-syn hases, and p o ein anspo e s in he p o ein p oducing CBS 513.88 s ain. Ou esul s and da a se s om his in eg a i e sys ems biology analysis esul ed in a snapsho o ungal e olu ion and will suppo u he op imiza ion o cell ac o ies based on ilamen ous ungi. [Supplemen al ma e ial is a ailable o his a icle. The A. nige ATCC 1015 whole genome sequence has been submi ed o GenBank (h p://www.ncbi.nlm.nih.go /genbank/) unde accession no. ACJE00000000. The sequence da a om he phylogeny s udy ha e been submi ed o GenBank unde accession nos. GU296686–GU296739. The mic oa ay da a om his s udy ha e been submi ed o he NCBI Gene Exp ession Omnibus (GEO) (h p://www.ncbi.nlm.nih.go /geo/) unde se ies accession no. GSE10983. The dsmM_ANIGERa_coll511030F lib a y and pla o m in o ma ion ha e been submi ed o GEO unde accession no. GPL6758.] The sap o ophic ilamen ous ungus Aspe gillus nige is ound globally and exhibi s a g ea di e si y in i s pheno ype. A. nige has become one o he majo wo kho ses in indus ial bio echnology, being e y e icien in p oducing bo h polysaccha ide-deg ading enzymes (pa icula ly amylases, pec inases, and xylanases) o o - ganic acids (mainly ci ic acid) in high amoun s. I also has a long his o y o sa e use (Schus e e al. 2002; an Dijck e al. 2003; an 19 P esen add ess: Sys ems Biology, Depa men o Chemical and Biological Enginee ing, Chalme s Uni e si y o Technology, SE-41296 Go ¨ ebo g, Sweden. 20 Co esponding au ho . E-mail sco .b[email p o ec ed]; ax (509) 372-4732. A icle published online be o e p in . A icle, supplemen al ma e ial, and pub- lica ion da e a e a h p://www.genome.o g/cgi/doi/10.1101/g .112169.110. F eely a ailable online h ough he Genome Resea ch Open Access op ion. 21:885–897 Ó2011 by Cold Sp ing Ha bo Labo a o y P ess; ISSN 1088-9051/11; www.genome.o g Genome Resea ch 885 www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om Dijck 2008). Comme cial impo ance is illus a ed by a wo ld ma ke o indus ial enzymes o nea ly US$ 5 billion in 2009, o which ilamen ous ungi accoun o oughly hal o he p o- duc ion (Lube ozzi and Keasling 2008), and a global ci ic acid p oduc ion o 9 310 6 me ic ons in 2000 (Ka a a and Kubicek 2003). In 2007, he genome sequence o A. nige s ain CBS 513.88, used o indus ial enzyme p oduc ion, was published (Pel e al. 2007). This s ain was de i ed om A. nige NRRL 3122, a s ain de eloped o glucoamylase A p oduc ion by classical mu agenesis and sc eening me hods ( an Lanen and Smi h 1968). This wo k ini ia ed a numbe o new genome-based in es iga ions (Sun e al. 2007; Ande sen e al. 2008ab; Ma ens-Uzuno a and Schaap 2008; Yuan e al. 2008ab) bu did no e eal di e ences be ween ci ic- acid-p oducing and enzyme-p oducing A. nige s ains (Cullen 2007). In his s udy, we p esen he nea ly comple e genome se- quence o he ci ic-acid-p oducing A. nige wild- ype s ain ATCC 1015 and compa e i o he genome sequence o he enzyme- p oducing s ain CBS 513.88. The gene ic di e si y o hese wo A. nige s ains was de e mined by applying sys ems biology ools as well as new bioin o ma ics me hods o examine mul i-le el di - e ences ha dis inguish he wild- ype ci ic-acid-p oducing s ain om he mu agenized glucoamylase A–p oducing s ain. Resul s Gene al genome s a is ics The 34.85-Mb genome sequence o A. nige ATCC 1015 was gen- e a ed using a sho gun app oach and hen u he imp o ed o a high-quali y assembly o 24 inished con igs sepa a ed by 15 gaps (including eigh om cen ome ic egions). Genome s a is ics a e summa ized in Table 1, wi h de ails in Supplemen al Tex 1. The ull sequence and anno a ions a e a ailable om he Join Genome Ins i u e (JGI) Genome Po al (h p://genome.jgi-ps .o g/Aspni5) and om NCBI (accession numbe ACJE00000000). The genome sequence o A. nige CBS 513.88 (Pel e al. 2007) was imp o ed using he ATCC 1015 sequence o close 186 con ig gaps and modi y he gene models associa ed wi h hese gaps (Table 1; Supplemen al Table 1). The upda ed A. nige CBS 513.88 genome se- quence is accessible h ough EMBL (accession numbe s AM269948– AM270415). We no e a la ge di e ence (2882) in he numbe o called genes in he wo s ains (Table 1). A ho ough analysis indica es an o e p edic ion o genes in CBS 513.88 and an unde p edic ion o genes in ATCC 1015, and 396/510 unique genes in CBS 513.88/ ATCC 1015 (Supplemen al Tex 2; Supplemen al Fig. 1; o de ails, see Supplemen al Tables 2–9). Fo u he alida ion o he ab- sence/p esence o indi idual p o eins, we ha e pe o med gDNA hyb idiza ions, which can be consul ed o e e ence (Supple- men al Table 16). Unique genes in bo h s ains sugges ho izon al gene ans e o be a cause o he amylase hype -p oduce pheno ype Fo he genes unique o he amylase-p oducing s ain CBS 513.88 (Supplemen al Table 6), he mos no able genes a e wo alpha- amylases ha a e iden ical o he Aspe gillus o yzae alpha-amylase, a possible cause o he amylase hype -p oduce pheno ype. We discuss his in de ail below (Fig. 2). Fu he mo e, we ind h ee possible polyke ide syn hases, which sugges s a unique seconda y me aboli e p o ile o his s ain. Examining he genes ound only in ATCC 1015 (Supple- men al Table 6) does no poin o any ob ious cause o ci ic acid hype -p oduc ion, bu ou possible polyke ide syn hases and a pu a i e NRPS a e ound o be unique o ATCC 1015, sugges ing his s ain has unique seconda y me aboli es as well. Mul iple e - ec s o his ype a e, indeed, seen o bo h s ains in analyses below (Fig. 1; Supplemen al Fig. 10; Table 2). Syn eny mapping shows 0.5 Mb o ch omosomal ea angemen s and a whole-a m in e sion The genomes o he wo s ains a e la gely syn enic (Fig. 1). Fo example, he 1429 p o ein-encoding genes o CBS 513.88 supe - con ig An01 ha ha e p edic ed coun e pa s in ATCC 1015 a e, wi h one excep ion, all mapped o ch omosome 2b. A simila ex- ample is seen o all bu one o 1486 p o ein-encoding genes on supe con ig An02. Bo h excep ions encode pu a i e Tan1-like ans- posases. No wi hs anding, a numbe o signi ican di e ences in genome con igu a ion exis (Fig. 1). Mo e han 0.5 Mb o addi ional genome sequence in ATCC 1015 esides in ou la ge egions on ch omosomes (Ch ) II, III, and V (Fig. 1). The lanking sequences o h ee o hese addi ional elemen s in ATCC 1015 a e syn enic o con inuous sequences in CBS 513.88 and a e hus ue di e ences in genome con igu a ion (Supplemen al Fig. 2; Supplemen al Table 2; o de ails on Ch III, see nex sec ion). The ou h egion on he le a m o Ch V in ATCC 1015 could no be e i ied due o a con ig gap o CBS 513.88. The con ig gap in he igh a m o Ch V in he ge- nome sequence o ATCC 1015 is spanned by con inuous sequence in CBS 513.88, con aining 15 p edic ed genes. An o e iew o he genes ound in he gaps unique o ATCC 1015 can be ound in Supplemen al Table 2. O he s iking di e ences include a la ge in e sion in Ch VIII and he in- e sion and ansloca ion o a la ge ag- men be ween he le a ms o Ch III and Ch VII (Supplemen al Fig. 3). Bo h we e con i med wi h PCR spanning he b eak- poin s (da a no shown). Finally, he p esence o elome e se- quences in genome da a con i ms an in- e sion o he comple e igh a m o Ch VI. Two ex a alpha-amylase-encoding genes likely ob ained h ough HGT An unma ched egion iden i ied o he le a m o Ch III (Fig. 1) spans 72.5 kb o Table 1. Gene al genome s a is ics o A. nige ATCC 1015 and A. nige CBS 513.88 A. nige ATCC 1015 ( his s udy) A. nige CBS 513.88 a (Pel e al. 2007) A. nige CBS 513.88 b ( his s udy) Gene models 11,200 14,165 14,082 Genome size (Mb) 34.85 33.93 34.02 P o ein leng h (amino acids) 484.3 439.9 442.5 Exons pe gene 3.1 3.6 3.6 Exon leng h (bp) 480.8 370.0 371.6 In on leng h (bp) 93.8 97.2 96.9 Excep genome sizes and he numbe o gene models, all alues a e a e ages. a The genome assembly published by Pel e al. (2007). b Genome assembly o A. nige CBS 513.88 a e gap closu e using sequence in o ma ion om ATCC 1015. 886 Genome Resea ch www.genome.o g Ande sen e al. Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om unique sequence in ATCC 1015. Rema kably, o CBS 513.88, a unique 85.3-kb sequence is ound a his loca ion (Fig. 2A). Gene anno a ion may be ound in Supplemen al Table 10. To con i m he ac ual absence o he egion in he ch omosome o ATCC 1015, we pe o med gDNA hyb idiza ions, which did indeed suppo his inding (Supplemen al Fig. 4). App oxima ely 12 kb o his egion sha es >99.8% sequence iden i y a DNA le el wi h genomic DNA om A. o yzae RIB40 (Fig. 2B). The comple e 85-kb agmen including he 12-kb alpha- amylase egion also appea s o be in ol ed in a ecen duplica ion ecombina ion e en o Ch VII, (Fig. 2C). Thus, he CBS 513.88 genome ha bo s wo addi ional alpha-amylase-encoding genes ha a e o hologs o he alpha-amylase-encoding genes (AO090023000944, AO090120000196) o A. o yzae RIB40. These indings sugges ha CBS 513.88 and he pa en al s ain NRRL 3122 (da a no shown) acqui ed hese duplica e alpha-amylase genes (An12g06930, An05g02100) h ough ho izon al gene ans e (HGT). The occu ence o HGT in Aspe gilli would u he be suppo ed by he p esence o alpha-amylase-encoding genes in o he black Aspe gilli ha display >99% nucleo ide iden i y o hose o A. o yzae RIB40 and he A. nige CBS 513.88 alpha-amylases (Ko man e al. 1990; Shibuya e al. 1992). The o igin o he Figu e 1. Syn eny map o he con igs o A. nige ATCC 1015 o he supe con igs o A. nige CBS 513.88. The colo ing o he ch omosomes shows syn enic egions in A. nige CBS 513.88. A abic nume als show he numbe o he supe con ig in A. nige CBS 513.88. G ay a eas show egions no ound in he CBS 513.88 genome sequence (Pel e al. 2007). (Filled black ci cles) P oposed loca ions o cen ome ic egions. Sequenced elome es a e ma ked wi h a T. Ze oes ma k he i s base o he con igs. A black line unde nea h a sec ion o he ch omosomes deno es in e ed sequence. Black his og ams show SNPs pe kilobase (numbe o single nucleo ide polymo phisms/kilobase) be ween he sequences o he wo s ains (y-axis: 0–160 SNPs/kb). Gaps be ween con igs and cen ome es a e no o scale. The alignmen demons a es almos comple e syn eny be ween he wo s ains, wi h he excep ion o a c oss-o e e en be ween he le a ms o ch omosomes III and VII. An o e iew o he genes ound in he gaps unique o ATCC 1015 can be ound in Supplemen al Table 2. Aspe gillus nige s ain e olu ion Genome Resea ch 887 www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om emaining pa o he unma ched egion emains unclea . No sig- ni ican simila i y was obse ed a he DNA le el, and only a ew o he encoded p o eins show simila i y wi h o he p o eins p esen in he NR p o ein da abase. The e is also e idence ha his HGT may ha e occu ed by he ac ion o ansposases. The 12-kb HGT egion is lanked by 202-bp in e ed, pe ec , e minal epea s (ITR) (Fig. 2A,B). Fu - he mo e, his HGT egion ha bo s ano he gene, An12g07000, ha is iden ical o he A. o yzae npA ansposase gene. In- e es ingly, in he A. o yzae RIB40genome, hesame12-kb ag- men has been duplica ed wice and is p esen on mul iple ch omosomes, bu only one pe ec copy o he ITRs is e ained (Fig. 2C). T ansposon p esence is s ain-speci ic T ansposon-like sequences we e iden i ied in bo h genomes and quan i ied (Supplemen al Table 11). This compa ison pinpoin s a ema kable di e ence in he p esence and amoun o ansposon- ela ed sequences in ATCC 1015 and CBS 513.88 bo h o class I and o class II ansposons: Only 16 sequences we e iden i ied in ATCC 1015, bu 55 we e de ec ed in CBS 513.88. Whe eas he ATCC 1015 genome con ains no class I supe amily copia and a single copy o class II supe amily Fo 1/pogo, hese sequences a e much mo e abundan in CBS 513.88, whe e 15 class I copia, and 13 class II Fo 1/pogo a e ound. SNP analysis e eals high mu a ion a es and hype a iable egions We ound 8 616 SNPs/kb (a e age 6s anda d de ia ion) and a maximum o 163 SNPs/kb single-nucleo ide polymo phisms (SNPs) be ween A. nige s ains ATCC 1015 and CBS 513.88 (Sup- plemen al Table 12). This alue is much highe han he maximum o 9 SNPs/kb ound by Cuomo e al. (2007) in a SNP analysis o wo Fusa ium g aminea um s ains. A compa ison o he A. nige ATCC 1015 genome sequence o ha o A. nige ATCC 9029 e ealed ma kedly less a ia ion be ween he wo (2 65 SNPs/kb) (Supplemen al Fig. 5), indica ing a la ge genomic a ia ion in he A. nige g oup bu li le be ween ATCC 1015 and 9029. These polymo phisms a e no uni o mly dis ibu ed bu clus e in hype - a iable egions (Fig. 1). This is suppo ed by a gDNA hyb idiza ion s udy, which con i ms he p esence o dis inc hype a iable egions (Supplemen al Fig. 4). Gene compa isons e eal indus ially ele an s ain di e ences To iden i y s ain-de ining sys emic e ec s o he genome a ia- ion, esul s om he Imp in gene syn eny analysis (see Supple- men al Table 5; Me hods) we e used o addi ional genome-scale in es iga ion. GO e m o e - ep esen a ion analysis was pe o med on gene g oups wi h dis inc common p ope ies ( o G oup o e iew, see Supplemen al Table 5; o de ails, see Supplemen al Table 13). Mos in e es ing is he obse a ion ha 37 p o eins in ol ed in an- sc ip ional egula ion a e non unc ional in CBS 513.88 due o ameshi s o s op codons. This sugges s a less s ingen egula ion o CBS 513.88 ela i e o ATCC 1015. A compa a i e shake- lask s udy o he wo s ains also sugges s mu a ions in egula o y ele- men s as ni ogen sou ce u iliza ion is impai ed in he CBS 513.88 s ain (da a no shown). While he ni ogen ca abolism egula o A eA was o iginally hough o be mu a ed in CBS 513.88 (Pel e al. 2007), a esequencing o a eA showed he ORFs o be iden ical. To de ec po en ial di e ences in he me abolism o he wo s ains, we compa ed all genes encoding p o eins ha display di e ences a he amino acid le el be ween he wo s ains on he me abolic ne wo k o A. nige (Ande sen e al. (2008a) and mapped hem o me abolic pa hways (Supplemen al Fig. 6). Mu a ions we e ound in he pa hways o biosyn hesis o p oline, aspa a e, aspa agine, yp ophan, and his idine, which may be ele an o p o ein p oduc ion. Also, mu a ions we e ound in he plasma memb ane-bound ATPase, in he enzymes o he GABA shun , o he TCA cycle, and in componen s o all s eps o he elec on anspo chain, which could be ele an o he p oduc ion o ci ic acid. Phylogene ic analysis con i ms high a iabili y be ween genome sequences To place he compa ison o A. nige ATCC 1015 and CBS 513.88 in o he la ge con ex o he Aspe gillus sec ion Nig i, a ia ion in he wo s ains was analyzed by phylogene ic analysis o a pa o he be a- ubulin sequence o a numbe o A. nige s ains and o he black Aspe gilli (Fig. 3A). To u he explo e he genome a iabili y ac oss he A. nige species, we iden i ied ou 1-kb egions om Ch II, Ch IV, Ch VI, and Ch VIII ha we e iden ical in he ATCC 1015 and ATCC 9029 s ains bu had ;20 SNPs/kb ela i e o he genome sequenceo CBS 513.88. These egions we e PCR-sequenced in se en A. nige s ains including he CBS 513.88 p ogeni o , NRRL Table 2. Exo-me abolomic p o iling o 11 A. nige s ains based on HPLC-DAD-FLD and HPLC-DAD-HRMS ( his s udy) Seconda y me aboli es Re e ence NRRL 3122 CBS 513.88 a CBS 126.49 ATCC 1015 a ATCC 9029 a CBS 554.65 NRRL 328 NRRL 350 NRRL 511 NRRL 1278 NRRL 2270 Au aspe one B Tanaka e al. 1966 • • • • • • ••••• Fumonisin B 2 F is ad e al. 2007b • • • • • ••••• Funalenone Inokoshi e al. 1999 • • • • ••••• Ko anin Bu ¨chi e al. 1971 • • • • ••••• Nig agillin Isogai e al. 1975 •• Och a oxin A Aba ca e al. 1994 •• • O landin Cu le e al. 1979 • • • • ••••• O he naph ho-g-py ones • • • • ••••• Py anonig in A Hio e al. 2004 • • • • • • ••••• Tensidol B Fukada e al. 2006 • • • • • • ••••• The e e ence column ela es o he elucida ion o he compound s uc u e. a Genome-sequenced s ains. Ande sen e al. 888 Genome Resea ch www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om 3122, and he A. nige neo ype s ain, CBS 554.65 (Kozakiewicz e al. 1992) o o m a phylogene ic ee wi h high a ia ion (Fig. 3B). Fo p e iously sequenced s ains, he esequencing esul s we e iden ical o he p e iously de e mined genomic sequence, he eby excluding ha he genomic a ia ion is an a i ac o low- ideli y sequencing. All A. nige s ains ha e an iden ical sequence o be a- ubulin and o m a s ongly suppo ed e minal clade oge he wi h As- pe gillus awamo i (Fig. 3A). The o he axonomically di e en spe- cies a e phylogene ically dis inc om his g oup, which is in ac- co dance wi h he se ial classi ica ion o F is ad e al. (2007a). The la e clus e ing based on he selec ed egions (Fig. 3B) was able o Figu e 2. Ho izon al gene ans e o alpha-amylase genes om A. o yzae o A. nige CBS 513.88. (A) Unma ched egion iden i ied o he le a m o ch omosome III spanning 65 kb o ATCC 1015 encoding 30 p edic ed genes and 85 kb o CBS 513.88 encoding 24 p edic ed genes. The unma ched egion is lanked on one side by a small local in e sion o h ee p edic ed genes (in g een). (B) The 12.4-kb HGT egion is pa o an iden ical 12.7-kb egion p esen in A. o yzae RIB40 supe con igs SC113 and SC023. In A. nige CBS 513.88, he ans e ed egion is enclosed by a 203-bp in e ed epea ( ed a ow). An12g07000 (blue) is a pu a i e ansposase and iden ical o he A. o yzae ansposon Ao 1 npA gene. Simila ly, he A. nige CBS 513.88 alpha- amylase encoding gene An12g06930 (o ange) is iden ical o A. o yzae anno a ed genes AO090120000196 (SC113) and AO090023000944 (SC023). (C) P oposed duplica ion– ecombina ion e en be ween supe con ig 12 and supe con ig 05 o he 12-kb HGT egion. B eakpoin s a e indica ed wi h do ed lines. The agmen encoding alpha-amylase An05g02200 is iden ical o An12g06930 (o ange). The b eakpoin s a e lanked wi h addi ional copies o he 203-bp epea egion ( ed a ow). The egion encoding genes An12g06940 o An12g06970 is iden ical o he egion encoding An05g02210 o An05g02130. The downs eam b eakpoin occu ed in he p edic ed gene coding egion o An12g06970, he eby dele ing he o iginal s a codon and ups eam egion. Aspe gillus nige s ain e olu ion Genome Resea ch 889 www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om sepa a e he A. nige s ains in o h ee g oups. S ain CBS 513.88 and i s p ogeni o NRRL 3122 we e iden ical o he egions, in- dica ing limi ed SNP in oduc ion due o classical s ain imp o e- men , and hus con i ms high a ia ion in he A. nige g oup. Exo-me abolomic p o iling desc ibes h ee dis inc g oups o A. nige To u he compa e he wo sequenced s ains on a sys em-wide, bu nongenomic le el, o each o he and o o he membe s o he A. nige species g oup, exo-me abolomic p o iles o he wo s ains we e compa ed o nine o he s ains, including he CBS 513.88 p ogeni o s ain, NRRL 3122; he widely used labo a o y s ain ATCC 9029; and he A. nige neo ype, CBS 554.65 (Table 2). This ga e h ee dis inc clades, which ollow he clus e ing o Figu e 3B, bu does no con o m o he phylogene ic dis ances calcula ed. S ain CBS 513.88 and i s p ogeni o s ain had e y simila sec- onda y me aboli e p o iles. These wo s ains di e ed om he o he s ains analyzed. S ain ATCC 1015 had a p o ile simila o se en o he s ains, al hough he e we e some quan i a i e di - e ences. S ain CBS 126.49 had a unique p o ile. In he case o och a oxin A, i is in e es ing ha in ATCC 1015 and ATCC 9029, a emnan o he PKS gene (An15g07920) o he pu a i e och a oxin clus e in A. nige CBS 513.88 (Pel e al. 2007) was iden i ied, which could be ela ed o a 21-kb dele ion in bo h s ains (Supplemen al Fig. 7). We ha e no ed ha a pu a i e NRPS is ound in one o he unique ch omosome egions ound in ATCC 1015 (Fig. 1), which may accoun o some o he di e ence (Supplemen al Table 2). T ansc ip ome analysis o A. nige ATCC 1015 and CBS 513.88 g owing on glucose To e alua e he e ec o he di e ences in genome sequence on he physiology o he wo s ains, ba ch cul i a ions in bio eac o s and a compa a i e ansc ip ome analysis we e pe o med. S ains ATCC 1015 and CBS 513.88 we e g own unde he same condi- ions in ba ch cul u es in a glucose-based minimal medium. The GlaA-p oducing s ain CBS 513.88 p oduced >1.2 g/L GlaA mo e han ATCC 1015, while p oducing ;1 g/L biomass less. O he measu ed cha ac e is ics we e simila o he wo s ains (Table 3). S a is ical analysis showed 4784 signi ican ly (adj. p<0.05) di - e en ially exp essed genes, wi h an almos equal numbe o genes up- egula ed in ei he s ain (2431 in CBS 513.88 s. 2353 o ATCC 1015). Examining he di e en ial exp ession in he con ex o me - abolic pa hways (Supplemen al Fig. 8), only he al e na i e oxi- da i e pa hway has uni o mly highe exp ession in ATCC 1015, while a subs an ial subse o me abolism was up- egula ed in CBS 513.88, including glycolysis and he TCA cycle. As he speci ic g ow h a es o he s ains a e simila , we sugges ha his ex a ac i i y o cen al me abolism p o ides p ecu so s o he highe p oduc i i y o glucoamylase. We also ind inc eased exp ession o genes in ol ed in amino acid me abolism, especially he en- i e biosyn he ic pa hways o h eonine, se ine, and yp ophan. Figu e 3. (A) Phylogene ic ela ionship o se e al black Aspe gilli based on pa ial sequencing o be a- ubulin. The ee was oo ed o Aspe gillus aculea us CBS 172.66. (B) Phylogene ic ela ionship o se en s ains o A. nige based on sequencing o 1-kb a iable egions om ou ch omosomes. The ee was oo ed o he sequence ob ained om Aspe gillus ca bona ius IMI 388653. Clades based on he exo-me abolomic g oupings o Table 2 a e shown. Fo bo h ees, boo s ap alues abo e 80% o he 1000 pe o med ei e a ions a e shown. Table 3. S a is ics o ba ch cul i a ions o A. nige ATCC 1015 and A. nige CBS 513.88 ATCC 1015 CBS 513.88 mRNA (h) 24.5 61.2 40.2 64.2 Biomass (g/L) 5.0 60.1 4.0 60.5 m max (h 1 ) 0.17 60.01 0.15 60.01 Glucose (g/L) 10.0 60.6 9.5 60.4 Glyce ol (g/L) 0.09 60.02 0.27 60.03 Y xs (Cmol/Cmol) a 0.67 60.03 0.55 60.03 GlaA (g/L) b 0.24 60.08 1.57 60.23 Ci ic acid (g/L) 0.10 60.12 0.14 60.03 Fe men a ions we e pe o med in biological iplica es o each s ain. Values a e p esen ed as a e age 6s anda d de ia ion. m max and Y xs a e gene al s a is ics o he e men a ions, while he emaining alues a e speci ic o he ime o sampling o ansc ip ion analysis (see mRNA ow). GlaA is glucoamylase A. a Biomass was con e ed o Cmol using 24.9 g o biomass/Cmol (Nielsen e al. 2003). b One uni o glucoamylase can be assumed o co espond o 25 mgo p o ein (PESL p o ein assay; Boeh inge Mannheim). Ande sen e al. 890 Genome Resea ch www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om In iguingly, analysis o he amino acid composi ion o he glu- coamylase A p o ein (Supplemen al Fig. 9) e ealed ha GlaA is a ypical in ha i has a highe con en o speci ically yp ophan, h eonine, and se ine han 90% o he p edic ed genes o CBS 513.88. Fo all h ee amino acids, he con en is almos wice as high as he a e age in he composi ion o o al p o ein (Ch is ias e al. 1975). We ini ially hough ha his was due o a ameshi mu a- ion iden i ied in he gene o he gene al amino acid– egula ing ansc ip ion ac o (c oss-pa hway con ol p o ein) CpcA (Sup- plemen al Table 5), bu esequencing o his gene showed his ameshi o be a sequencing e o . The p esence o wo sense mu a ions was con i med and is cu en ly being in es iga ed. To u he suppo ha he inc eased p oduc ion o GlaA is signi ican enough o a ec amino acid biosyn hesis, we used a genome-scale me abolic model o A. nige (Ande sen e al. 2008a) o model he wo s ains unde he g ow h condi ions used (Table 3). The compu ed luxes (Supplemen al Table 14) show ha o all h ee amino acids, he luxes h ough he biosyn he ic pa hways mus be a leas wice as high in CBS 513.88 compa ed o ATCC 1015 o suppo he inc eased GlaA p oduc ion, hus co espond- ing well wi h he ansc ip ome esul s. Fo a b oade analysis o ends e ealed by he ansc ip ome p o iles, a GO e m o e - ep esen a ion analysis was conduc ed on he genes signi ican ly up- egula ed in ei he s ain (Supple- men al Table 15). The analysis con i med ha ai s ele an o p o ein p oduc ion—such as amino acid biosyn hesis and RNA aminoacyla ion—appea o be highly signi ican o CBS 513.88. Fo ATCC 1015, GO biological unc ions o elec on anspo (adj. p=2.7 310 5 ), ca bohyd a e anspo (adj. p=6.8 310 3 ), and o ganic acid anspo (adj. p=0.035) we e signi ican . An examina ion o exp ession o indi idual genes showed ha glaA had signi ican ly inc eased exp ession in CBS 513.88 (adj. p=84 310 6 ), bu he old change (3.2) was lowe han he in- c ease in enzyme ac i i y (six old) (Table 3). In e es ingly, all iden- i ied RNA-aminoacyl syn hases we e ound o be wo old o six old up- egula ed in CBS 513.88. Mapping o gene exp ession o he genome iden i ies seconda y me aboli e clus e ac i i ies and unde pins he whole-a m in e sion in ch omosome VI To examine di e ences in gene exp ession be ween he wo s ains ela i e o ch omosome posi ions, he log 2 - a ios o he gene ex- p ession indices om he ansc ip ome analysis we e mapped o he syn eny diag ams o Figu e 1 (Supplemen al Fig. 10). The analysis o his mapping iden i ied a numbe o ea- u es. Fi s , six egions no ound in CBS 513.88 con ain genes wi h uni o mly highe exp ession in ATCC 1015, suppo ing he absence o hese egions in CBS 513.88. Second, wo ac i e sec- onda y me aboli e clus e s we e iden i ied on Ch VIII (con ains a polyke ide syn hase [ID: 211885] ha is unique o ATCC 1015) and Ch I (NRPS clus e , ound in bo h s ains [NRPS ID: 43555/ An09g00520], bu appea s only o be ac i e in CBS 513.88). Thi d, he in e sion o he en i e igh a m o Ch VI (Fig. 1) is u he suppo ed. The elome ic posi ion e ec , which de ines de- c eased exp ession in he icini y o elome es, has been de- sc ibed o Saccha omyces ce e isiae (Go schling e al. 1990). Thus, i an a m has been in e ed, educed exp ession should be ound a opposi e ends in ATCC 1015 and CBS 513.88. This is indeed e lec ed in he log 2 - a io on he igh a m o Ch VI (Supplemen al Fig. 10). Discussion In his s udy, we p o ide and compa e he genomes o wo s ains o A. nige . These wo s ains ha e di e en pheno ypes: one, he p edecesso o e icien enzyme-p oducing s ains ha ing un- de gone some le el o mu agenesis and selec ion, and he o he a wild- ype pa en s ain o high ci ic-acid-p oducing s ains. This makes he compa ison in e es ing bo h in e ms o genomic e- sea ch and indus ial applica ions. We ha e suppo ed he con- clusions o ou compa ison wi h u he expe imen s, allowing us o p opose new hypo heses and conclusions wi hin h ee main a eas: (1) gene ic di e si y o he A. nige g oup, (2) ho izon al gene ans e in ungi, and (3) ungal bio echnology (discussed sepa- a ely below). The di e si y o he wo s ains was explo ed h ough a mul i- le el compa ison (DNA, ch omosome, gene, and p o ein) be ween bo h genome a chi ec u es and was suppo ed by mul iple ypes o expe imen al wo k. In e es ingly, we obse e a ema kable simi- la i y a he le el o ch omosome, gene o de , and gene iden i y, bu also no able di e si y was de ec ed and u he explo ed, namely, dis inc egions wi h high le els o SNPs, a 0.8-Mb di - e ence in genome size, se e al la ge inse ion/dele ions o up o 200 kb, di e en ansposon popula ions, a majo ansloca ion/ in e sion, he in e sion o an en i e ch omosomal a m, and a la ge se o unique genes in bo h species (see poin s 1–5 below): 1. Abou 400–500 unique genes we e iden i ied o each s ain, mos o which a e e enly dis ibu ed o e he ch omosomes. This indica es s ain e olu ion by a high equency o loss and/o up ake o gene agmen s. In e es ingly, he opposi e is seen in wo s ains o he pa hogenic Aspe gillus umiga us, whe e 80% o he s ain-speci ic genes (143 and 218, espec- i ely) a e ound in a ew, la ge isola e-speci ic genomic islands (Fedo o a e al. 2008). Thus, his in a-species inc eased e- quency o ans e /loss o gene ic elemen s so a seems o be speci ic o he A. nige g oup. Mo e genome sequences may p o e i o be p esen in o he Aspe gilli as well. 2. The whole-a m in e sion on Ch VI is a ema kable e en . O e he cha ac e ized Aspe gilli sequenced genomes, mos desc ibed in e sions we e isola ed as mu an s, wi h a b eakpoin in a gene, leading o a pheno ype. Pel e al. (2007) epo ed high syn eny be ween he cen ome ic pa s o he a ms o A. nige CBS 513.88 and Aspe gillus nidulans FGSC A4. E en so, in his case, i is suppo ed by he genome sequence, he p esence o elome ic sequences, and a de ec able elome ic posi ioning e ec . F om Figu e 2 in Galagan e al. 2005, i is also clea ha cen ome ic in e sions mus ha e occu ed in he genealogy o he common ances o o A. nidulans,A. o yzae, and A. umiga us. 3. The la ge ansloca ion/in e sion e en o Ch III and Ch VII explains he disc epancies in ch omosome size epo ed by Pel e al. (2007) be ween s ain N400 (Ve does e al. 1994) and calcula ed sizes o Ch III and Ch VII o CBS 513.88. This ac and ha he b eakpoin s a e wi hin wo o he wise in ac genes indica e ha he e en occu ed a e he b anching o s ains N400 and NRRL 3122, possibly in he mu agenesis o he p e- decesso o CBS 513.88. 4. The p esence and unc ionali y o ansposons a e quali a i ely and quan i a i ely less complex in he genome o ATCC 1015 han in CBS 513.88 (Supplemen al Table 11). Such a di e en dis ibu ion o ansposons be ween s ains has been obse ed be o e (B aumann e al. 2008) and was desc ibed o be inducible by mu agenesis and o he s ess-induced condi ions (S and and Aspe gillus nige s ain e olu ion Genome Resea ch 891 www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om McDonald 1985). Howe e , he numbe and scale o di e ences make a ecen mu agenesis p og am unlikely o be he basis o he mul iple HGT e en s. Fu he mo e, he de ec ion o epea sequences in a numbe o DNA egions was accompanied by s ain-speci ic di e ences. The e o e, we sugges ha he ansposon popula ions ha e played a signi ican ole in s ain e olu ion. 5. Finally, he phylogene ic analysis (Fig. 3B) con i ms he high gene ic a ia ion desc ibed o wo s ains o be gene al wi hin he A. nige g oup and no he esul o ei he CBS 513.88 o ATCC 1015 being a ypical. This is again suppo ed by he exo- me abolomic p o iles o he 11 examined A. nige s ains (Table 2). While he g ouping o he exo-me aboli es does no accu- a ely e lec he cladog ams in Figu e 3, we see his as con i - ma ion o he high gene ic a ia ion in he A. nige g oup, as he p oduc ion o a gi en exo-me aboli e may be changed by a single genomic e en . Thus, we belie e ha he di e si y o he wo genome-sequenced s ains is common o he A. nige g oup, which seems o ha e highly dynamic genomes. The p esence/absence o HGT in ungi has caused a lo o discus- sion (Galagan e al. 2005; Keeling and Palme 2008; Khaldi and Wol e 2008; Khaldi e al. 2008). In his s udy, we ha e shown HGT o be a mos likely o igin o wo alpha-amylase genes in CBS 513.88 ha migh ha e been ans e ed o A. nige om ano he species, possibly A. o yzae. The HGT mus ha e occu ed be o e he sepa- a ion o CBS 513.88 and ATCC 1015. The amylase genes a e no ound in he ATCC 1015 s ain (also suppo ed by gDNA hyb id- iza ion), bu hey a e in CBS 513.88 lanked by a ansposon, which is ound bo h in A. o yzae RIB40 and in a unca ed o m in ATCC 1015. Wi h he e idence o HGT in o Aspe gillus cla a us om a Magnapo he- ela ed dono , p esen ed by Khaldi e al. (2008), HGT is now iden i ied in wo sepa a e cases wi h dis an ly ela ed species and may hus be seen as a gene al phenomenon in ilamen ous ungi. The ini ial eason o compa ing he wo s ains has been o gain insigh in he gene ic basis o he wo indus ially ele an pheno ypes. In ci ic-acid-p oducing ATCC 1015, we a e in igued o see highe exp ession o elemen s om he elec on anspo , ca bohyd a e anspo , and o ganic acid anspo . These ac o s a e known o be in ol ed in ob aining high ci ic acid yields. Highe exp ession le els e en hough ci ic acid p oduc ion is no no ably highe (see discussion below) sugges s ha hese ai s may ha e been con ibu ing ac o s o he o iginal selec ion o ATCC 1015 o ci ic acid p oduc ion. Fo he GlaA-p oducing s ain (CBS 513.88), we obse e sys emic changes: Mu a ions in egula o y genes (e.g., cpcA), highe exp ession o glaA i sel , and up- egula ion o all iden i ied RNA-syn hases may all con ibu e o a mo e e i- cien enzyme p oduce . We also p esen he hypo hesis ha in- c eased p oduc ion o he amino acids se ine, h eonine, and yp- ophan may be equi ed o e icien GlaA p oduc ion, as hese a e o e - ep esen ed in GlaA. While we ha e iden i ied mul iple possible con ibu o s o he e icien amylase p oduc ion o CBS 513.88, we ha e only seen a ew ac o s associa ed wi h ci ic acid p oduc ion. We p opose mul iple easons o his: (1) CBS 513.88 was selec ed o inc eased GlaA p oduc ion, while ATCC 1015 is a wild- ype s ain wi h only modes ci ic acid p oduc ion (compa ed o cu en yields). (2) The medium used o he ansc ip ome analysis is no op imized o ci ic acid p oduc ion, as his would no gi e he dispe sed g ow h necessa y o ep esen a i e sampling (Ka a a and Kubicek 2003). (3) P o ein syn hesis is a complex p ocess, wi h mo e ac o s in- ol ed han in he p oduc ion and sec e ion o a simple o ganic acid. Since mo e han 6000 genes di e a he amino acid le el be ween he wo s ains, i would be unlikely no o ind many di e ences in ol ed in p o ein p oduc ion. Fo his eason, we ha e pu ou emphasis on genes ound in mul iple ypes o anal- yses, e ec s in en i e pa hways, and ai s (GO) ha a e ound o be s a is ically o e - ep esen ed, which allows o obus conclusions. In conclusion, ou esul s es ablish a i m compa a i e geno- mics ounda ion on which o build and es hypo heses ega ding enzyme p oduc ion, o ganic acid p oduc ion, and di e si y wi hin a ‘‘species g oup.’’ Me hods ATCC 1015 genome assembly The sequence eads we e de i ed om ou whole-genome sho - gun (WGS) lib a ies: one wi h an inse size o 2–3 kb, wo wi h an inse size o 6–8 kb, and one wi h an inse size o 35–40 kb. The eads we e sc eened o ec o using c oss_ma ch and hen im- med o ec o sequence and quali y (J Chapman, N Pu nam, I Ho, and D Rokhsa , unpubl.). Reads sho e han 100 bases a e im- ming we e excluded. The inal da a se included: 28,551 o 2–3-kb eads, con aining 21.5 Mb o sequence; 160,479 o 6–8-kb eads, con aining 123 Mb o sequence; and 38,651 o 35–40-kb eads, con aining 23.8 Mb o sequence. The da a we e assembled using elease 2.7 o Jazz, a WGS as- semble de eloped a he JGI (Apa icio e al. 2002; J Chapman, N Pu nam, I Ho, and D Rokhsa , unpubl.). The genome size and se- quence dep h we e ini ially es ima ed o be 36 Mb and 8.0, e- spec i ely. A e emo al o sho (<1 kb) and edundan sca olds (<5 kb wi h 80% o mo e o he leng h ma ching a sca old >5 kb), he assembly included 43.7 Mb o sca old sequence wi h 8.2 Mb (18.9%) o gaps in 350 sca olds, wi h hal o he sca old sequence con ained in he eigh la ges sca olds o 1.81 Mb o longe . The sequence dep h de i ed om he assembly was 7.88 60.05. To es ima e he comple eness o he assembly, a se o 50,001 ESTs was BLAT-aligned o he unassembled immed eads, as well as he assembly i sel . Fo y- h ee housand h ee hund ed and wen y- h ee ESTs (86.6%) we e >80% co e ed by he unassembled da a; 43,978 (88.0%) we e >50% co e ed; and 44,206 (88.4%) we e >20% co e ed. By way o compa ison, 48,798 ESTs (97.6%) showed hi s o he assembly. O ien a ion o he ch omosome a ms was based on eigh a ailable elome e sequences, he supe con ig o ien a ion p oposed by Pel e al. (2007), and esea ch on linkage g oups in A. nige (Bos e al. 1989; Debe se al.1989, 1990a,b; Swa e al. 1992; Ve does e al. 1994). ATCC 1015 genome inishing To pe o m genome imp o emen on A. nige ATCC 1015, ini ial ead layou s om he whole-genome sho gun assembly we e con e ed in o he ph ed/ph ap/consed pipeline (Go don e al. 1998). Following manual inspec ion o he assembled sequences, inishing was pe o med by esequencing plasmid subclones and by walking on plasmid subclones o osmids using cus om p ime s. All inishing eac ions we e pe o med wi h 4:1 BigDye o dGTP BigDye e mina o chemis y (Applied Biosys ems). Repea s in he sequence we e esol ed by ansposon-hopping 8-kb plasmid clones. Fosmid clones we e sho gun sequenced and inished o ill la ge gaps, esol e la ge epea s, o o esol e ch omosome dupli- ca ions and ex end in o ch omosome elome e egions. The esul ing 24 inished sca olds we e o ien a ed whe e possible in o ch omosome s uc u es using compa isons o Pel e al. (2007) and elome e posi ions. The imp o ed genome consis s o Ande sen e al. 892 Genome Resea ch www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om 34,853,277 bp wi h an es ima ed e o a e o less han 1 e o in 100,000 bp. ATCC 1015 au oma ic anno a ion Gene models in he genome o A. nige we e p edic ed using Fgenesh (Salamo and Solo ye 2000), Fgenesh+(Salamo and Solo ye 2000), and Genewise (Bi ney and Du bin 2000) in eg a ed in o he JGI anno a ion pipeline. Fgenesh was ained on a se o mo e han 2000 pu a i e ull-leng h ansc ip s de i ed om clus e ed A. nige ESTs and eliable homology-based gene models o show 81% sen- si i i y and 81% speci ici y o p edic ions on a es se . Homology- based gene p edic o s we e seeded wi h BLASTX alignmen s o p o eins om he NCBI non edundan se o p o eins. Thi y-one housand i e hund ed and se en y-eigh A. nige ESTs we e se- quenced using Sange echnology and we e ei he di ec ly mapped o genomic sequence when he ESTs included pu a i e ull-leng h (FL) genes o used o ex end p edic ed gene models in o FL genes by adding 59and/o 39UTRs. In addi ion, 386,515 ESTs wi h an a e age leng h o 104 n we e sequenced wi h 454 Li e Sciences (Roche) GS20 sequence s and used in alida ion o p edic ed gene models. Since mul iple gene models we e gene a ed o each locus, a single ep esen a i e model was chosen based on homology and EST suppo and used o u he analysis. All p edic ed gene models we e unc ionally anno a ed by sequence simila i y o anno a ed genes om he NCBI non- edundan se and o he specialized da abases (such as KEGG) (Kanehisa e al. 2002, 2004) using BLAST and ha dwa e accele a ed double-a ine Smi h-Wa e man alignmen s (h p://www. imelogic. com). Func ional and s uc u al domains we e p edic ed in p o ein sequences using he In e P o so wa e (Zdobno and Apweile 2001). All genes we e also anno a ed acco ding o Gene On ology (Ashbu ne e al. 2000; Ha is e al. 2004), euka yo ic o hologous g oups (KOGs) (Koonin e al. 2004), and KEGG me abolic pa hways (Kanehisa e al. 2004). ATCC 1015 sequence a ailabili y A. nige assemblies, anno a ions, and analyzes a e a ailable h ough he in e ac i e JGI Genome Po al a h p://genome.jgi- ps .o g/Aspni5/Aspni51.home.h ml. Genome assemblies oge he wi h p edic ed gene models and anno a ions we e also deposi ed a NCBI unde he p ojec accession numbe ACJE00000000. S ains The ollowing A. nige s ains we e used o expe imen s (de- posi ion numbe s in di e en collec ions a e gi en as well): CBS 513.88 =FGSC A1513 =IBT 29270, ATCC 1015 =IBT 28639 =NCTC 3858a =NRRL 1278 =NRRL 350 =NRRL 511 =NRRL 328 =Thom 167 =CBS 113.46 =ATCC 10582 =IMI 031821 =LSBH Ac4 =Thom 3528.7, NRRL 3 =IBT 28539 =MUCL 30480 =DSM 2466 =CECT 2088 =VTT D-85240 =NRRL 566 =WB3 =ATCC 9029 =N400 = CBS 120.49 =IMI 041876, NRRL 326 =IBT 27876 =WB 326 =CBS 554.65 =ATCC 16888 =IHEM 3415 =IMI 050566 =Thom 2766 = JCM 10254 =IFO 33023 (ex annin-gallic acid e men a ion; A. nige neo ype) (Kozakiewicz e al. 1992) NRRL 328 =IBT 27878 = NRRL 350 =IBT 27877 =CBS 113.46, NRRL 337 =CBS 126.48 = ATCC 10254 =DSM 734 =IFO 6428 =IMI 015954 =WB 337, NRRL 363 =IBT 3617 =IBT 5764 =CBS 126.49 =ATCC 10698 =IFO 6648, NRRL 511 =IBT 27875, NRRL 1278 =IBT 27872, NRRL 2270 =IBT 26391 =ATCC 11414 =VTT D-77050 =IMI 075353 =A60 =S.M. ma in A-1-233 =Wisconsin 72-4, de i ed om ATCC 1015, and NRRL 3122 =IBT 23538 =ATCC 22343 =CBS 115989 ( his s ain is a wild- ype p ogeni o o CBS 513.88). S ain his o ies o ATCC 1015, ATCC 9029, and NRRL 3122 a e e iewed by Bake (2006). Exo-me aboli e p o iling A. nige s ains we e inocula ed on Czapek yeas au olysa e (CYA), yeas ex ac suc ose (YES) aga , and CYA wi h 5% NaCl aga (CYAS aga ). Fo medium composi ion, see F is ad and Samson (2004). All s ains we e h ee-poin inocula ed on hese media and incuba ed a 25°C in da kness o a week, a e which i e plugs (6 mm di- ame e ) along a diame e o a ungal colony we e cu ou and ex ac ed (Smedsgaa d 1997). The ex ac s we e analyzed by HPLC- DAD luo escence (F is ad and Th ane 1987; Smedsgaa d 1997) and by HPLC-HR-MS (Nielsen and Smedsgaa d 2003; F is ad e al. 2007b). The seconda y me aboli es we e iden i ied by compa ison wi h au hen ic s anda ds ( umonisin B 2 , och a oxin A, nig agillin, ko anin, and o landin) and by UV spec a and high- esolu ion mass spec a using elec osp ay ioniza ion. CBS 513.88 gap closu e The nea -comple e ATCC 1015 genome sequence was used o e i y con ig o de and o ien a ion and o es ima e gap sequence leng h be ween con igs in CBS 513.88. Gap- lanking PCR p ime s (20-me s) gi ing ise o PCR p oduc s wi h a minimal o e lap o 100 bp we e au oma ically designed wi h p ime -3 (Rozen and Skale sky 2000) using cus om sc ip s. PCR p oduc s o expec ed size we e pu i ied using a QIAquick PCR Pu i ica ion Ki (QIAGEN) and sequenced by Baseclea . The upda ed sequence o A. nige CBS 513.88 is a ailable om EMBL (h p://www.ebi.ac.uk/embl/) un- de accession nos. AM269948–AM270415. Sequencing o a iable egions Regions we e ampli ied using s anda d PCR echniques wi h he ou ollowing p ime pai s ( he name desc ibes ch omosome numbe and di ec ion): (1) Ch 02-Fwd: 59-GGACACTGCTTGATG TGATG-39, Ch 02-Re : 59-GAGAGACGTACGAAAGGTTG-39; (2) Ch 04-Fwd: 59-CGATCTGCGACCAGGA-39, Ch 04-Re : 59-CATAA CGGATTCGTCGCTG-39; (3) Ch 06-Fwd: 59-CTTGAAGGCGTTGA GGTC-39,Ch 06-Re :59-GCGAGTATGTGGCTAACATC-39; (4) Ch 08- Fwd: 59-GGTATGTCACATTCCRTCCA-39, Ch 08-Re : 59-GCTTGC AGTGAGCAAGGA-39. The sequences on Ch II, Ch VI, and Ch VIII o e lap wi h p edic ed genes. The PCR p oduc s we e pu i ied us- ing a QIAquick PCR Pu i ica ion Ki (QIAGEN) and sequenced by Agencou Bioscience Co po a ion (GenBank accession numbe GU296708–GU296739). Sequencing o be a- ubulin Ampli ica ion o pa o he be a- ubulin gene was pe o med using he p ime s B 2a (GGTAACCAAATCGGTGCTGCTTTC) and B 2b (ACCCTCAGTGTAGTGACCCTTGGC) (Glass and Donaldson 1995). Bo h s ands o he PCR agmen s we e sequenced wi h he ABI P ism Big DyeTM Te mina o .3.0 Ready Reac ion Cycle sequenc- ing ki . Samples we e analyzed on an ABI PRISM 3700 Gene ic An- alyze , and con igs we e assembled using he o wa d and e e se sequences wi h he p og am SeqMan om he Lase Gene package (GenBank accession numbe s GU296686–GU296707). Calcula ion o phylogene ic ee Sequences we e aligned using Clus alX and immed o he i s common base a bo h ends be o e making he inal alignmen . Fo he ee o Figu e 3B, ou immed sequences we e joined o each s ain be o e making he inal alignmen . T ee calcula ions we e pe o med using T eeCon ( an de Pee and de Wach e 1994). Dis ances we e es ima ed using he Jukes and Can o algo- i hm aking inse ions and dele ions in o accoun . Topology was Aspe gillus nige s ain e olu ion Genome Resea ch 893 www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Ap il 26, 2016 - Published by genome.cshlp.o gDownloaded om