scieee Open visual document viewer

Whole-genome sequencing and annotation of the yeast Clavispora santaluciae reveals important insights about its adaptation to the vineyard environment

Franco-Duarte, Ricardo; Čadež, Neža; Rito, Teresa; Drumonde-Neves, João; Dominguez, Yazmid Reyes; Pais, Célia; Sousa, Maria João; Soares, Pedro

Abstract

<i>Clavispora santaluciae</i> was recently described as a novel non-<i>Saccharomyces</i> yeast species, isolated from grapes of Azores vineyards, a Portuguese archipelago with particular environmental conditions, and from Italian grapes infected with <i>Drosophila suzukii</i>. In the present work, the genome of five <i>Clavispora santaluciae</i> strains was sequenced, assembled, and annotated for the first time, using robust pipelines, and a combination of both long- and short-read sequencing platforms. Genome comparisons revealed specific differences between strains of <i>Clavispora santaluciae</i> reflecting their isolation in two separate ecological niches—Azorean and Italian vineyards—as well as mechanisms of adaptation to the intricate and arduous environmental features of the geographical location from which they were isolated. In particular, relevant differences were detected in the number of coding genes (shared and unique) and transposable elements, the amount and diversity of non-coding RNAs, and the enzymatic potential of each strain through the analysis of their CAZyome. A comparative study was also conducted between the <i>Clavispora santaluciae</i> genome and those of the remaining species of the Metschnikowiaceae family. Our phylogenetic and genomic analysis, comprising 126 yeast strains (alignment of 2362 common proteins) allowed the establishment of a robust phylogram of Metschnikowiaceae and detailed incongruencies to be clarified in the future.

Full text

  Ci a ion: F anco-Dua e, R.; ˇ Cadež, N.; Ri o, T.; D umonde-Ne es, J.; Dominguez, Y.R.; Pais, C.; Sousa, M.J.; Soa es, P. Whole-Genome Sequencing and Anno a ion o he Yeas Cla ispo a san aluciae Re eals Impo an Insigh s abou I s Adap a ion o he Vineya d En i onmen . J. Fungi 2022,8, 52. h ps://doi.o g/10.3390/jo 8010052 Academic Edi o : B ian Monk Recei ed: 6 Decembe 2021 Accep ed: 4 Janua y 2022 Published: 5 Janua y 2022 Publishe ’s No e: MDPI s ays neu al wi h ega d o ju isdic ional claims in published maps and ins i u ional a il- ia ions. Copy igh : © 2022 by he au ho s. Licensee MDPI, Basel, Swi ze land. This a icle is an open access a icle dis ibu ed unde he e ms and condi ions o he C ea i e Commons A ibu ion (CC BY) license (h ps:// c ea i ecommons.o g/licenses/by/ 4.0/). Fungi Jou nal o A icle Whole-Genome Sequencing and Anno a ion o he Yeas Cla ispo a san aluciae Re eals Impo an Insigh s abou I s Adap a ion o he Vineya d En i onmen Rica do F anco-Dua e 1,2,* , Neža ˇ Cadež 3, Te esa Ri o 1,2 , João D umonde-Ne es 4, Yazmid Reyes Dominguez 5, Célia Pais 1,2 , Ma ia João Sousa 1,2 and Ped o Soa es 1,2 1CBMA, Cen e o Molecula and En i onmen al Biology, Depa men o Biology, Uni e si y o Minho, 4710-057 B aga, Po ugal; [email p o ec ed] (T.R.); [email p o ec ed] (C.P.); [email p o ec ed] (M.J.S.); ped osoa [email p o ec ed] (P.S.) 2 Ins i u e o Science and Inno a ion o Bio-Sus ainabili y (IB-S), Uni e si y o Minho, 4710-057 B aga, Po ugal 3Depa men o Food Science and Technology, Bio echnical Facul y, Uni e si y o Ljubljana, 101, 1000 Ljubljana, Slo enia; [email p o ec ed] 4IITAA—Ins i u e o Ag icul u al and En i onmen al Resea ch and Technology, Uni e si y o Azo es, 9700-042 Ang a do He oísmo, Po ugal; d [email p o ec ed] 5Laimbu g Resea ch Cen e, Laimbu g 6, 39052 Vadena, I aly; Yazmid.Reyes-Dominguez@laimbu g.i *Co espondence: [email p o ec ed] o [email p o ec ed] Abs ac : Cla ispo a san aluciae was ecen ly desc ibed as a no el non-Saccha omyces yeas species, isola ed om g apes o Azo es ineya ds, a Po uguese a chipelago wi h pa icula en i onmen al condi ions, and om I alian g apes in ec ed wi h D osophila suzukii. In he p esen wo k, he genome o i e Cla ispo a san aluciae s ains was sequenced, assembled, and anno a ed o he i s ime, using obus pipelines, and a combina ion o bo h long- and sho - ead sequencing pla o ms. Genome compa isons e ealed speci ic di e ences be ween s ains o Cla ispo a san aluciae e lec ing hei isola ion in wo sepa a e ecological niches—Azo ean and I alian ineya ds—as well as mechanisms o adap a ion o he in ica e and a duous en i onmen al ea u es o he geog aphical loca ion om which hey we e isola ed. In pa icula , ele an di e ences we e de ec ed in he numbe o coding genes (sha ed and unique) and ansposable elemen s, he amoun and di e si y o non- coding RNAs, and he enzyma ic po en ial o each s ain h ough he analysis o hei CAZyome. A compa a i e s udy was also conduc ed be ween he Cla ispo a san aluciae genome and hose o he emaining species o he Me schnikowiaceae amily. Ou phylogene ic and genomic analysis, comp ising 126 yeas s ains (alignmen o 2362 common p o eins) allowed he es ablishmen o a obus phylog am o Me schnikowiaceae and de ailed incong uencies o be cla i ied in he u u e. Keywo ds: genomics; phylogenomics; unc ional gene analysis; Me schnikowiaceae; Azo es; wine yeas s; bio echnology; adap a ion 1. In oduc ion In ou p e ious su eys on he yeas di e si y o Azo ean ineya ds, in 2009 and 2010 [ 1 – 3 ], we desc ibed a new yeas species Cla ispo a san aluciae [ 4 ], isola ed om g apes. I was cha ac e ized on he basis o he sequences o he in e nal ansc ibed space (ITS) egion (ITS1-5.8S–ITS2), he sequences o he D1/D2 domain o he la ge subuni (LSU) RNA gene, and pa icula physiological cha ac e is ics. Tha s udy also desc ibed his species as being isola ed om g apes in ec ed wi h D osophila suzukii in I aly. This showed iden ical D1/D2 sequences and e y simila ITS egions ( i e nucleo ide subs i u ions) o he Azo ean s ains. The new species was ob ained om pa icula i icul u al en i on- men s, ypical o he Azo es a chipelago, which esul om he in e ac ion be ween speci ic clima ic condi ions, au och honous g ape ine cul i a s, and local i icul u al p ac ices. Pheno ypic cha ac e iza ion o his new species e ealed some in e es ing ea u es ha J. Fungi 2022,8, 52. h ps://doi.o g/10.3390/jo 8010052 h ps://www.mdpi.com/jou nal/jo J. Fungi 2022,8, 52 2 o 18 posi ioned i apa om he closely ela ed ones, such as he inabili y o g ow a empe a- u es abo e 35 ◦C , p oduc ion o ace ic acid, and he capaci y o assimila e s a ch. The ull bio echnological po en ial o his new species emains o be explo ed, as does he unde - s anding o he genomic ea u es associa ed wi h he adap a ion o i s en i onmen . The occu ence o bio echnologically impo an ea u es associa ed wi h species o his clade is no uncommon. (Candida)in e media, a xylose-u ilizing species o he Me schnikowiaceae amily, belonging o he Cla ispo a clade, displays a high-capaci y xylose anspo sys- em [ 5 , 6 ]. Due o hese cha ac e is ics, his species has been ca ego ized as an a ac i e species o p oduce e hanol om lignocellulosic biomass [7]. Ou p e ious s udy was one o he ew o epo he p esence o yeas species belonging o he Cla ispo a clade in ineya ds, e en hough some a e epo s ha e al eady associa ed hese yeas s wi h winemaking, as de ailed below. We highligh ed he a i y o hese occu ences, in a e iew o he associa ion o non-Saccha omyces yeas s wi h i icul u e and winemaking [ 8 ]. In ha s udy, we sys ema ized 80 yea s o he li e a u e desc ibing non- Saccha omyces yeas species isola ed om g apes and/o g ape mus s and compiled a lis o 293 species. Only wo species belonging o he Cla ispo a clade we e iden i ied—Cla ispo a uc us and Cla ispo a lusi aniae. E en hough he e is no s ong associa ion be ween Cla ispo a species and wine, some epo s ha e al eady desc ibed he possible ad an ages o hese yeas s in he p oduc ion o wines wi h al e na i e senso y cha ac e is ics, albei wi h some associa ed disad an ages, such as he p esence o an abno mally high concen a ion o ace aldehyde [ 9 ]. In Azo ean ineya ds, no o he yeas o his genus was ound wi hin he 2910 isola es iden i ied in ou p e ious wo k [ 1 , 2 ], bu a single species o his amily (Me schnikowia pulche ima) was ound in ou di e en islands in wo consecu i e sampling yea s. The use o his species in wine bio echnology was ecen ly e iewed [ 10 ], wi h he au ho s concluding ha i s e sa ili y lies in i s abili y o e men mus in combina ion wi h o he yeas species (mainly o ci cum en i s low e men a i e powe ), as well as modula ing he syn hesis o seconda y e men a ion me aboli es o imp o e and di e si y he senso y p o ile o he wine. The phylog am ob ained in ou p e ious s udy [ 4 ], based on conca ena ed sequences o he D1/D2 domain o he LSU RNA gene, and o he ITS egion, placed Cla ispo a san aluciae s ains nea closely he ela ed species Cla ispo a uc us, (Candida)aspa agi, (Candida) i iphila, (Candida)phyllophila, (Candida)ca ajalis, and Cla ispo a lusi aniae. The desc ip ion o his new species was conside ed as an impo an add-on in unde s anding he phylogene ic ela ionships wi hin he Cla ispo a clade and cla i ying hei biodi e si y and ecology. P e iously, se e al yeas s ha now belong o he Cla ispo a genus, we e placed in he anamo phic genus Candida, due o hei inabili y o o m sexual spo es [ 11 ]. Wi h ecen phylogene ic analysis based on DNA sequences, in combina ion wi h physiological e idence, he ela ionships be ween he species o Candida and he genus Cla ispo a began o be cla i ied, leading o a new classi ica ion o his g oup by Daniel e al. in 2014 [ 11 ], wi h 40 species o Candida assigned o he genus Cla ispo a, based on sequences o he LSU D1/D2 domain, ITS egion and ou coding genes (ACT1,TEF1,MCM7, and RPB2). In 2018, Ku zman e al. [ 12 ] desc ibed ou new species o Me schnikowia and p oposed o ans e se en addi ional Candida species o bo h Me schnikowia and Cla ispo a gene a. Ku zman e al. concluded ha he axonomy o he Cla ispo a clade could only be cla i ied by whole-genome compa isons, which s ill need o be pe o med. Rega ding he Cla - ispo a genus, only wo species ha e been genome sequenced and anno a ed: Cla ispo a lusi aniae ( 11.9–12.1 Mb , 8 ch omosomes, genome accession ASM167369 2 ), and Cla ispo a uc us (11.4 Mb, NCBI genome accession ASM370779 1). In 2016, a phylogeny o he Me schnikowiaceae amily was p esen ed by Lachance e al. [ 13 ], combining d a genomes o 55 s ains and iden i ying 3016 o hologues, 1061 o which exis in all he analyzed s ains. E en hough he Cla ispo a lusi aniae genome was used as a que y o compa e be ween all he genomes, no species belonging o he Cla ispo a genus we e conside ed, wi h he analysis ocused on only he Me schnikowia genus. Mo e ecen ly, Shen e al. [ 14 ] a emp ed o econs uc he phylogeny o 300 budding yeas species, ocusing mainly on J. Fungi 2022,8, 52 3 o 18 he di e si y o Saccha omyco ina. In ha wo k, whole genomes we e used o pa ially desc ibe he phylogeny o Me schnikowiaceae clade conside ing only 22 Me schnikowia species. This analysis used only he ype s ains o each species, which lack in a-species di e si y, and he genome anno a ion was no di ec ed o he analysis o his clade. Thus, a deep, b oad, and ocused analysis using obus ly anno a ed genomes o he phyloge- ne ic ela ions be ween Cla ispo a species (including also he ecen ly assigned Candida species) and he sis e genus Me schnikowia is lacking. Wi h his in mind, he objec i e o he p esen wo k was o sequence, assemble and anno a e whole genomes o he di e en Cla ispo a san aluciae s ains, using a combina ion o bo h long- and sho - ead sequencing pla o ms, in o de o ob ain high-quali y sequences o compa a i e genomics. We used he assembled genomes o u he elucida e he molecula mechanisms unde lying he adap a ion o his species o he pa icula en i onmen al cha ac e is ics om which i was isola ed. In pa icula , he main goal was o un a el he genomic ea u es ha can explain he pheno ypic cha ac e is ics p e iously obse ed wi hin isola es o Cla ispo a san aluciae, as well as o p edic hei bio echnological po en ial. In addi ion, all species om he Me schnikowiaceae amily whose comple e genome was publicly a ailable we e conside ed and combined in phylogene ic and genomic analysis, o help cla i y he phyloge- ne ic placemen o Cla ispo a yeas s wi hin his la ge g oup o impo an non-Saccha omyces yeas s. We plan o use ou esul s o cla i y he posi ioning o (Candida) species in ela ion o he sis e genus Me schnikowia and econs uc Me schnikowiaceae amily phylogeny using comple e genomes. 2. Ma e ials and Me hods 2.1. Cell Cul u e, Sample Collec ion, and DNA Ex ac ion The ype s ain o Cla ispo a san aluciae (A1.18 T = CBS 16465 T ), oge he wi h he h ee s ains isola ed om g apes o Azo ean ineya ds (A1.5, A1.7, and A1.19), and one addi- ional s ain isola ed om g apes in ec ed wi h D osophila suzukii in I aly ( LB-NB-3.3 ) [ 4 ], we e g own in YPD b o h (yeas ex ac , 1% w/ ; pep one, 1% w/ ; glucose 2% w/ ), in 50 mL conical lasks, o 48 h a 28 ◦ C, 220 pm. Genomic DNA was isola ed acco ding o he p o ocol published by Schwa z and She lock [ 15 ], wi h a ew adap a ions o he isola ion o DNA om non-Saccha omyces yeas s. A e washing in 0.9 M so bi ol solu- ion, he cells we e incuba ed by adding 20 µ L o Ly icase (30 mg/mL, Sigma-Ald ich, S . Louis , MO, USA) o he cells. The incuba ion ime was a leas 4 h a 37 ◦ C. Following phenol/chlo o o m (Millipo e, Bu ling on, MA, USA) ex ac ion, he DNA was p ecipi a ed using 40 µ L o 3 M sodium ace a e (pH 5.5) and 1 mL o absolu e e hanol and esuspended in 200 µL o TE bu e . 2.2. Genome Sequencing and Assembly The genomes o all he Cla ispo a san aluciae s ains we e sequenced by using a combi- na ion o he long-and sho - ead sequencing echnologies o PacBio and Illumina, espec- i ely. A e DNA ex ac ion, lib a y p epa a ion and PacBio/Illumina sequencing we e pe o med a No ogene acili ies (No ogene Company LTD, Camb idge, Uni ed Kingdom). Low-quali y eads and adap e s we e emo ed by No ogene, and sequencing quali y was accessed using Fas QC so wa e (h p://www.bioin o ma ics.bbs c.ac.uk/p ojec s/ as qc/; accessed on 1 Augus 2021). The sequencing da a a e a ailable a NCBI BioP ojec ID PRJNA784374. Long- eads ob ained om PacBio sequencing we e de no o assembled using Canu .1.9 [ 16 ] wi h de aul pa ame e s. Illumina pai ed-end eads we e hen used o imp o e assembly quali y, using Masu ca so wa e .4.0.5, in pa icula he Polca package [ 17 , 18 ]. Finally, RagTag so wa e .2.1.0 [ 19 ] was used o assemble all he sca olds in o longe eads, using ch omosome in o ma ion om he closely ela ed species Cla ispo a lusi aniae and (Candida)in e media. Genome assembly quali y me ics, a ailable in Table 1, we e compu ed using QUAST .5.0.2 [20]. J. Fungi 2022,8, 52 4 o 18 Table 1. Genome assembly s a is ics o Cla ispo a san aluciae s ains. A1.18TA1.5 A1.7 A1.19 LB-NB-3.3 Canu assemble Assembly leng h (bp) 11,088,431 11,018,248 10,921,443 10,861,576 11,019,028 Numbe o sca olds 43 13 46 86 30 N50 (bp) 315,943 802,369 355,153 1,329,122 494,470 L50 7 3 6 14 7 Numbe o N’s pe 100 Kb 0 0 0 0 0 Numbe o sca olds > 5000 bp 29 11 28 53 23 To al leng h > 5000 bp 10,780,215 10,974,595 10,557,056 10,118,395 10,856,688 Masu ca assemble Subs i u ion e o s e ised 42 8 141 428 70 Inse ion/Dele ion e o s e ised 1686 536 2825 6144 1607 Assembly leng h (bp) 11,089,145 11,018,616 10,922,446 10,863,639 11,019,715 Numbe o sca olds 43 13 46 86 30 N50 (bp) 532,329 1,048,728 654,799 218,696 650,701 L50 7 3 6 14 7 Numbe o N’s pe 100 Kb 0 0 0 0 0 Numbe o sca olds >5000 bp 29 11 28 53 23 To al leng h >5000 bp 10,780,920 10,974,959 10,558,055 10,120,373 10,857,358 RagTag assemble Assembly leng h (bp) 11,092,545 11,019,016 10,925,846 10,870,339 11,021,815 Numbe o sca olds/ch omosomes 9 9 12 19 9 Numbe o N´s pe 100 Kb 3065 3.63 31.12 61.64 19.05 Numbe o sca olds >5000 bp 4 8 4 3 7 To al leng h >5000 bp 11,025,073 11,000,234 10,766,609 10,606,285 10,966,127 Ploidy haploid haploid haploid haploid haploid GC con en (%) 49.66 49.70 49.73 49.66 49.76 To de e mine ploidy, we used nQui e so wa e [ 21 ] o align sequencing eads o he ype s ain genome assembled a e RagTag and de e mine base equency dis ibu ions be ween equencies 20 and 80. Assessmen o each genomes’ comple eness was pe o med using Benchma king Uni e sal Single-Copy O hologs (BUSCO) so wa e .5.2.2 [ 22 ]. A - e age nucleo ide iden i y (ANI) was calcula ed using he O hoANIu web ool [ 23 ] in pai wise mode, o compa e he nucleo ide con en o genomes. 2.3. Genome Anno a ion Anno a ion o Cla ispo a san aluciae genome assemblies was pe o med using AU- GUSTUS so wa e .3.4.0 [ 24 , 25 ], conside ing 11 di e en p e- ained models, chosen as belonging o he Ascomyco a phyla Saccha omyces ce e isiae S288c, Candida albicans, Meye ozyma (Candida)guillie mondii, Candida opicalis, Deba yomyces hansenii, E emo he- cium gossypii, Kluy e omyces lac is, Lodde omyces elongispo us, Sche e somyces (Pichia) s ipi is, Schizosaccha omyces pombe, and Ya owia lipoly ica. Resul s we e manually e iewed o selec he anno a ion wi h he highe numbe o p edic ed coding genes, which was ob ained using Lodde omyces elongispo us as he p e- ained model, o all he Cla ispo a san aluciae s ains. The po en ial coding egions (nucleo ide sequences) epo ed by AUGUSTUS we e ex ac ed om he comple e genomes o FASTA iles. CMsea ch [ 26 ] and S uc RNA inde [ 27 ] we e used o sc eening he p esence o non-coding RNA (ncRNA). The R am da abase [ 28 ] was employed o ncRNA sea ching, using an e- alue o 0.01. Func ional genomic anno a ion was pe o med wi h eggNOG-mappe .2 [ 29 ] by conside ing p o eins p edic ed by AUGUSTUS and choosing only o hologs ha we e in e ed om he expe imen al e idence. The esul s we e desc ibed conside ing clus e s o o hologous g oups (COGs) wi h hei associa ed unc ional ca ego ies [ 30 ], and also conside ing he Kyo o Encyclopedia o Genes and Genomes (KEGG) pa hways, in pa ic- ula he KEGG O hology (KO) desc ip o s [ 31 , 32 ]. Gene unc ion p edic ions we e also accomplished by assessing he Ca bohyd a e-Ac i e EnZymes (CAZymes) da abase [33]. J. Fungi 2022,8, 52 5 o 18 Final genome anno a ions o all Cla ispo a san aluciae s ains a e a ailable in Supplemen a y Da a S1. 2.4. Homology Analysis, Compa a i e Genomics, and Phylogenomics To compa e he genome o Cla ispo a san aluciae ype s ain A1.18 wi h ha o he emaining s ains, do plo s we e p oduced using he Re-Do -Able ool (h ps://www. bioin o ma ics.bab aham.ac.uk/p ojec s/ edo able/). In e -species di e ences be ween membe s o he amily Me schnikowiaceae we e e alua ed by downloading all he comple e genomes publicly a ailable a NCBI (121 s ains belonging o 48 di e en species). When mo e han one s ain was a ailable o a ce ain species, all s ains we e conside ed. The excep ion was (Candida)au is o which only he ep esen a i e genome was used since he hund eds o s ains wi h genome sequence a ailable would ha e inc eased edundancy. KEGG Mappe was used as a collec ion o KEGG mapping ools o linking genes and p o eins o me abolic pa hways [ 32 , 34 ]. In pa icula , KO gene anno a ions, ob ained om eggnog-mappe , we e used o assess pa hway comple eness using KEGG Mappe – Recons uc web ool (www.genome.jp/kegg/mappe / econs uc .h ml). Resul s we e applied in he cons uc ion o a hea map using Mic oso Excel®. A da abase was p epa ed by conside ing all 126 comple e genomes (121 s ains o Me schnikowiaceae amily plus he 5 Cla ispo a san aluciae isola es). To a oid inconsis- ency, he 121 Me schnikowiaceae genomes we e anno a ed using Augus us wi h he same p e- ained model as was applied o he anno a ion o he Cla ispo a san aluciae genomes. BLASTP analysis was pe o med using he ull p o eome o he Cla ispo a san aluciae ype s ain A1.18 T as a que y agains he o al da abase. An E- alue cu o o 10 −6 was used o exclude alse esul s, and a pipeline adap ed om [ 35 ] was used o pe o m compa a i e genomics be ween all isola es. The BLASTP esul s we e il e ed whe e ep esen a i e p o eins we e de ec ed in he o he 121 isola es. Each se o p obable homologous p o eins (con aining he que y and he espec i e esul s) we e mul iple aligned using he MAFFT algo i hm in FasPa se (h ps://gi hub.com/Sun-Yanbo/FasPa se ) [ 36 ]. All p o eins om a gi en o ganism we e conca ena ed using he alignmen esul s o ob ain he co e con- se ed aligned p o eome con aining mos ly essen ial genes no ela ed o speci ic biological ai s o each species. This alignmen was hen used o phylogene ic econs uc ion by conside ing he maximum likelihood in IQ-TREE (www.iq ee.o g) [ 37 ], wi h he JTT model o amino acid e olu ion and gamma-dis ibu ed a es ( ou a es) wi h 500 boo s ap epli- ca es. Two ou g oups we e conside ed: Lipomyces lipo e and Cybe lindne a jadinii. FigT ee .1.4.4 (h p:// ee.bio.ed.ac.uk/so wa e/ ig ee/) was used o isualize and edi he ee. The second ound o BLASTP analysis, using he p o eomes o he i e Cla ispo a san aluciae s ains as que ies, allowed building Venn diag ams o schema ize he numbe o genes common be ween he i e genomes using he a e age esul s be ween all pai s o be ween g oups o h ee, ou , o in all i e s ains. 3. Resul s and Discussion 3.1. Sequencing, De No o Assembly, and Anno a ion o Cla ispo a San aluciae Genome Genome sequencing o Cla ispo a san aluciae s ains A1.18 T , A1.5, A1.7, A1.19, and LB-NB-3.3 was pe o med using a combina ion o long- and sho - ead sequencing pla - o ms. Be ween 27,903 and 44,977 eads we e ob ained wi h long- ead sequencing, wi h a maximum ead leng h o 110,418 base pai s (bp). Sho - ead sequencing was used o e ine long- ead sequencing esul s. An a e age alue o 3 × 10 6 pai ed-end eads, wi h 250 bp each, was ob ained o each s ain. The i s ound o assembly was pe o med using Canu and Masu ca assemble s (sequencing s a is ics a e p esen ed in Table 1), and hen RagTag so wa e assembled he sca olds in o pu a i e ch omosomes. By using h ee assemble s we we e able o assemble long and sho - ead sequences in o ull ch omosomes o h ee o he s ains, including he ype s ain A1.18 T and s ains A1.5 and LB-NB-3.3. The emain- ing wo s ains, possibly due o lowe sequencing dep h, we e only assembled in o la ge sca olds. The a ained haploid genome size (10.8 Mb o 11.1 Mb) was compa able wi h J. Fungi 2022,8, 52 6 o 18 he p e iously published genomes o Cla ispo a yeas s, in pa icula wi h he 11.9–12.1 Mb o Cla ispo a lusi aniae (8 ch omosomes) [ 38 , 39 ], o wi h he 11.4 Mb o Cla ispo a uc us (NCBI genome accession ASM370779 1). The high-quali y-assembled genomes allowed he p edic ion o be ween 6015 and 6092 p o ein-coding genes o he i e Cla ispo a lusi aniae s ains using AUGUSTUS so wa e (Table 2, Supplemen a y Da a S1). These alues a e among he highes epo ed o yeas s o he Cla ispo a clade, and a e compa able only o he anno a ion o one ((Candida)in e media s ain YCC 4715), o which 6082 coding genes we e p edic ed [ 40 ], bu co esponding o a g ea e genome leng h o 13.08Mb. Table 2. Cla ispo a san aluciae genome anno a ion s a is ics. A1.18TA1.5 A1.7 A1.19 LB-NB-3.3 P o ein coding genes To al numbe 6092 6034 6067 6015 6038 Range o p o ein leng hs (aa) 66–4974 63–4974 57–4974 60–4974 66–5293 A e age p o ein leng h (aa) 557.6 556.6 550.9 543.5 518.3 Non-coding RNAs mic oRNAs (miRNAs) 32 32 33 31 21 small RNAs (sRNA) 20 21 22 20 23 nuclea RNAs (snRNA) 7 7 6 7 7 nucleola RNAs (snoRNA) 93 91 99 94 98 long noncoding RNAs (lncRNA) 8 8 9 8 12 ibosomal RNAs ( RNA) 96 63 42 69 124 ans e RNAs ( RNA) 276 259 279 299 248 O he 29 32 32 35 32 BUSCO O hologs Ascomyco a odb10 da abase Genome Comple eness (%) 93.5 94.4 93.4 90.7 93.6 Comple e BUSCOs 1595 1611 1594 1547 1597 F agmen ed BUSCOs 17 14 18 21 4 Missing BUSCOs 94 81 94 138 94 Saccha omyce es odb10 da abase Genome Comple eness (%) 98.0 99.1 98.0 95.1 98.2 Comple e BUSCOs 2094 2118 2094 2032 2099 F agmen ed BUSCOs 14 11 12 17 13 Missing BUSCOs 29 8 31 88 25 Eggnog-mappe unc ional anno a ion Genes wi h KO assigned 3130 (51.4%) 3129 (51.9%) 3125 (51.6%) 3130 (52.0%) 3101 (51.4%) Genes wi h COG assigned 4180 (68.6%) 4171 (69.1%) 4166 (68.7%) 4119 (68.5%) 4141 (68.6%) CAZymes unc ional anno a ion Numbe o genes anno a ed 120 121 118 112 117 The unusually high numbe o p edic ed p o eins in he genome o Cla ispo a san- aluciae was likely no ela ed, in ou opinion, o any peculia i y o his yeas ´ s genome bu a he o he use o ad anced sequencing echnologies, oge he wi h an imp o ed anno a ion pipeline. The lowes numbe o p edic ed coding sequences was de e mined o s ain A1.19. This could be a ibu ed o lowe sequencing dep h. This was also he sho es genome o he i e, he one wi h he lowes N50 alues (Table 1), and he one wi h lowe BUSCO genome comple eness sco es, bo h in Ascomyco a and Saccha omyce es da abases (Table 2). The highes numbe o p edic ed p o eins was desc ibed in he anno a ion o he genome o he ype s ain A1.18 T , wi h 6092 coding sequences (Table 2). The a e age leng h o he p edic ed p o eins was sligh ly lowe in LB-NB-3.3, al hough he la ges p o ein o 5293 amino acids (aa) was anno a ed in his s ain. This la ge open eading ame encodes he p o ein midasin (Mdn1), an ATPase o 560 kDa ha is essen ial o cell iabili y. I was iden i ied in all Cla ispo a san aluciae s ains and epo ed in o he yeas s, such as in he gene a Saccha omyces and Schizosaccha omyces, as well as in dis an o ganisms as J. Fungi 2022,8, 52 7 o 18 D osophila and A abidopsis [ 41 ]. The lowes coding sequence anno a ed (57 aa) co esponds o a hypo he ical p o ein no ye cha ac e ized in he Me schnikowiaceae (da a no shown) bu iden i ied as a mi ochond ial ATP syn hase ε chain-domain-con aining p o ein in he Te ezia cla e yi myco hizal ungus (NCBI accession KAF8454923.1). The ac ha we ound no p o eins below his size, which could co espond o he anno a ion o alse posi i es, highligh s he high anno a ion quali y ob ained wi h he compu a ional pipeline and he sequencing echnology applied. The o al numbe o non-coding RNAs (ncRNA) p edic ed using s uc RNA inde and he P am da abase was simila among he i e Cla ispo a san aluciae s ains (Table 2, Supplemen a y Da a S2). The e was a high simila i y be ween s ains o he majo i y o he ncRNA anno a ed, wi h he excep ion o ibosomal and ans e RNAs ( RNA and RNA, espec i ely), whose quan i ies showed ele an in e -s ain a ia ion no di ec ly co ela ed wi h he numbe o p edic ed coding sequences o wi h he genome size. Many sequencing p ojec s igno e he compa ison o ncRNA be ween s ains, bu by de ailing hei analysis, i may be possible o unde s and pa icula and in ica e mechanisms o adap a ion o he en i onmen . 3.2. Compa a i e Genomics o Cla ispo a San alucieae S ains To compa e s uc u al a ia ions be ween he genomes o he Cla ispo a san aluciae s ains pai wise, do plo s we e ob ained (Figu e 1A). Resul s showed a s iking pa e n o conse a ion o mos s ains, wi h a high deg ee o mac osyn eny mainly be ween he ype s ain and s ains A1.19 and LB-NB-3.3. On he o he hand, s ains A1.5 and A1.7 showed some di e en ia ion, in pa icula by he p esence o se e al dele ions in pa s o he genome, as ep esen ed by ansloca ions (“jumps” in he do plo ) away om he main diagonal. In pa icula , s ain A1.5 seems o ha e mesosyn eny wi h he ype s ain, since we can gene ally obse e conse a ion o he gene con en . Howe e , in some pa s o he genome, many in e sions (blue lines) and ansloca ions we e de ec ed. This obse a ion is no conco dan wi h he simila i ies obse ed in he ITS and D1/D2 egions [ 4 ], which showed ha s ain A1.5 is mos closely ela ed o he ype s ain. J. Fungi 2022, 8, x FOR PEER REVIEW 8 o 18 o he ou Azo ean s ains (Supplemen a y Da a S1). Acco ding o ou p e ious wo k [42] on he cha ac e iza ion o isogenic isola es o wine S. ce e isiae yeas s, ansposable ele- men s seem o be ela ed o he adap a ion o yeas s o he luc ua ing en i onmen al con- di ions ound in he ha sh en i onmen o he Azo es a chipelago, and hese gene ic ea- u es a e ela ed wi h impo an pheno ypic cha ac e is ics ha de e mine he s ains bi- o echnological po en ial [43,44]. Figu e 1. Compa a i e genomics o Cla ispo a san aluciae genomes: (A) whole-genome do -plo com- pa ison be ween he sequenced s ains in pai wise mode. Homologous egions a e plo ed as do s. Red lines link pa allel homologous pai s, and blue lines link an i-pa allel pai s; (B) Venn diag am indica ing he numbe o sha ed coding genes among Cla ispo a san aluciae s ains. 3.3. Func ional Anno a ion o Cla ispo a San aluciae P o eome Fo his analysis, eggNOG-mappe unc ionally anno a ed he p edic ed open ead- ing ames o Cla ispo a san aluciae, p o iding impo an insigh s in o hei biological sig- ni icance (Table 2, Figu e 2). Be ween 3101 and 3103 genes we e assigned o a KO ca ego y, co esponding o an a e age o 51.7% o all he anno a ed genes. A o al o 4180 genes o Cla ispo a san aluciae ype s ain A1.18T (68.6% o he o al genes) we e clus e ed in o 24 COGs using eggNOG-mappe (Figu e 2A), which we e hen classi ied in o h ee main unc ional ca ego ies (Figu e 2). This analysis e ealed low a ia ion be ween he i e s ains which is in acco dance wi h he emaining anno a ion s a is ics shown be o e. O no e is he ac ha he numbe o unc ionally anno a ed genes ob ained in all s ains a ied be ween 68.5 and 69.1% (Table 2) and is a he low, as indica ed by he high num- be o genes wi h “unknown unc ion” in panel B o Figu e 2 (g ay ba s; be ween 20.7 and 20.9%). Howe e , hese alues a e lowe han hose ob ained o o he species (Figu e 2C, and ca ego y S in panel D), such as Cla ispo a lusi aniae, wi h 24%, (Candida in e media), wi h 25%, and Me schnikowia eukau ii, wi h 24%, o e en o Saccha omyces ce e isiae (22%) o To ulaspo a delb ueckii (22%), as shown in ou p e ious wo k [35]. This low numbe o genes wi h “unknown unc ion” is a consequence o an imp o emen in he sequencing and anno a ion pipelines no mally used o anno a e yeas genomes. Func ional anno a ion o Cla ispo a san aluciae e ealed ha he highes pe cen age o anno a ed genes (Figu e 2B) was ela ed o “me abolism” (be ween 27.3 and 27.7%), ollowed by “cellula p ocesses and signaling” (26.4–26.6%). This esul is in ag eemen wi h ha o o he yeas s o Me schnikowiaceae (Figu e 2, panels C and D), al hough his Figu e 1. Compa a i e genomics o Cla ispo a san aluciae genomes: ( A ) whole-genome do -plo compa ison be ween he sequenced s ains in pai wise mode. Homologous egions a e plo ed as do s. Red lines link pa allel homologous pai s, and blue lines link an i-pa allel pai s; ( B ) Venn diag am indica ing he numbe o sha ed coding genes among Cla ispo a san aluciae s ains. J. Fungi 2022,8, 52 8 o 18 A o al o 5564 coding genes we e ound o be sha ed be ween he i e Cla ispo a san aluciae s ains, co esponding o he pangenome o he species (Figu e 1B). S ain NB- LB-3.3 showed a su p isingly high numbe o unique genes (298), no sha ed by any o he o he s ains, e lec ing i s adap a ion o a di e en ecological niche, as his s ain was isola ed om I alian g apes in ec ed wi h D osophila suzukii. On he o he hand, 283 genes we e sha ed only by he s ains isola ed om Azo ean ineya ds, indica ing adap a ion mechanisms o he in ica e and a duous en i onmen al condi ions o he geog aphical loca ion om which hey we e isola ed. Addi ionally, and o pa icula no e, is he ac ha no ansposable elemen was iden i ied in he genome o s ain LB-NB-3.3, unlike he o he ou Azo ean s ains (Supplemen a y Da a S1). Acco ding o ou p e ious wo k [ 42 ] on he cha ac e iza ion o isogenic isola es o wine S. ce e isiae yeas s, ansposable elemen s seem o be ela ed o he adap a ion o yeas s o he luc ua ing en i onmen al condi ions ound in he ha sh en i onmen o he Azo es a chipelago, and hese gene ic ea u es a e ela ed wi h impo an pheno ypic cha ac e is ics ha de e mine he s ains bio echnological po en ial [43,44]. 3.3. Func ional Anno a ion o Cla ispo a San aluciae P o eome Fo his analysis, eggNOG-mappe unc ionally anno a ed he p edic ed open eading ames o Cla ispo a san aluciae, p o iding impo an insigh s in o hei biological signi i- cance (Table 2, Figu e 2). Be ween 3101 and 3103 genes we e assigned o a KO ca ego y, co esponding o an a e age o 51.7% o all he anno a ed genes. A o al o 4180 genes o Cla ispo a san aluciae ype s ain A1.18 T (68.6% o he o al genes) we e clus e ed in o 24 COGs using eggNOG-mappe (Figu e 2A), which we e hen classi ied in o h ee main unc ional ca ego ies (Figu e 2). This analysis e ealed low a ia ion be ween he i e s ains which is in acco dance wi h he emaining anno a ion s a is ics shown be o e. O no e is he ac ha he numbe o unc ionally anno a ed genes ob ained in all s ains a ied be ween 68.5 and 69.1% (Table 2) and is a he low, as indica ed by he high numbe o genes wi h “unknown unc ion” in panel B o Figu e 2(g ay ba s; be ween 20.7 and 20.9%). Howe e , hese alues a e lowe han hose ob ained o o he species (Figu e 2C, and ca ego y S in panel D), such as Cla ispo a lusi aniae, wi h 24%, (Candida in e media), wi h 25%, and Me schnikowia eukau ii, wi h 24%, o e en o Saccha omyces ce e isiae (22%) o To ulaspo a delb ueckii (22%), as shown in ou p e ious wo k [ 35 ]. This low numbe o genes wi h “unknown unc ion” is a consequence o an imp o emen in he sequencing and anno a ion pipelines no mally used o anno a e yeas genomes. Func ional anno a ion o Cla ispo a san aluciae e ealed ha he highes pe cen age o anno a ed genes (Figu e 2B) was ela ed o “me abolism” (be ween 27.3 and 27.7%), ollowed by “cellula p ocesses and signaling” (26.4–26.6%). This esul is in ag eemen wi h ha o o he yeas s o Me schnikowiaceae (Figu e 2, panels C and D), al hough his no el yeas species has a highe pe cen age o genes ela ed o me abolism, which poin s o a supe io bio echnological po en ial o his species. The impo ance o his alue is e en mo e e iden i we compa e i wi h he unc ional anno a ions o yeas s om o he amilies, o which usually “in o ma ion s o age and p ocessing” is he mos ep esen ed ca ego y, as is he case o T. delb ueckii and S. ce e isiae, as p e iously shown [ 35 ]. The mos abundan COG ca ego y in he genome o Cla ispo a san aluciae A1.18 T (panel A) was “ ansla ion, ibosomal s uc u e, and biogenesis” (333 genes, ep esen ing 8% o he anno a ed genes), ollowed closely by “pos ansla ional modi ica ion, p o ein u no e , chape ones” (328/7.8%). The leas abundan ca ego ies we e “ex acellula s uc u es”, wi h only wo associa ed genes. J. Fungi 2022,8, 52 9 o 18 J. Fungi 2022, 8, x FOR PEER REVIEW 10 o 18 Figu e 2. Func ional anno a ion o Cla ispo a san aluciae genome: (A) p o eome classi ica ion in o 23 unc ional ca ego ies, co esponding o clus e s o o hologous g oups (COGs): A, RNA p ocessing and modi ica ion; B, ch oma in s uc u e and dynamics; C, ene gy p oduc ion and con e sion; D, cell cycle con ol and mi osis; E, amino acid me abolism and anspo ; F, nucleo ide me abolism and anspo ; G, ca bohyd a e me abolism and anspo ; H, coenzyme me abolism; I, lipid me ab- olism; J, ansla ion; K, ansc ip ion; L, eplica ion and epai ; M, cell wall/memb ane/en elop bio- genesis; O, pos ansla ional modi ica ion, p o ein u no e , chape one unc ions; P, ino ganic ion anspo and me abolism; Q, seconda y S uc u e; S, unc ion unknown; T, signal ansduc ion; U, in acellula a icking and sec e ion; Y, nuclea s uc u e; Z, cy oskele on; (B) classi ica ion o he anno a ed genes in o ou la ge unc ional ca ego ies; (C) compa ison be ween he i e Cla ispo a san aluciae s ains and o he ele an yeas species in p opo ions o he la ge unc ional ca ego ies; (D) compa ison be ween ele an yeas species classi ica ion o he anno a ed genes in o 23 COG ca ego ies; (E) pe cen age o CAZymes in he i e sequenced genomes o Cla ispo a san aluciae and o he ele an yeas s, showing he dis ibu ion o p edic ed p o eins in o majo amilies. Func ional anno a ion o Cla ispo a san aluciae was also accomplished using KEGG Mappe —Recons uc Pa hway ool [32,34]. This ool comple ed KO-based mapping Figu e 2. Func ional anno a ion o Cla ispo a san aluciae genome: ( A ) p o eome classi ica ion in o 23 unc ional ca ego ies, co esponding o clus e s o o hologous g oups (COGs): A, RNA p ocessing and modi ica ion; B, ch oma in s uc u e and dynamics; C, ene gy p oduc ion and con e sion; D, cell cycle con ol and mi osis; E, amino acid me abolism and anspo ; F, nucleo ide me abolism and anspo ; G, ca bohyd a e me abolism and anspo ; H, coenzyme me abolism; I, lipid me abolism; J, ansla ion; K, ansc ip ion; L, eplica ion and epai ; M, cell wall/memb ane/en elop biogenesis; O, pos ansla ional modi ica ion, p o ein u no e , chape one unc ions; P, ino ganic ion anspo and me abolism; Q, seconda y S uc u e; S, unc ion unknown; T, signal ansduc ion; U, in acellula a icking and sec e ion; Y, nuclea s uc u e; Z, cy oskele on; ( B ) classi ica ion o he anno a ed genes in o ou la ge unc ional ca ego ies; ( C ) compa ison be ween he i e Cla ispo a san aluciae s ains and o he ele an yeas species in p opo ions o he la ge unc ional ca ego ies; ( D ) compa ison be ween ele an yeas species classi ica ion o he anno a ed genes in o 23 COG ca ego ies; (E) pe cen age o CAZymes in he i e sequenced genomes o Cla ispo a san aluciae and o he ele an yeas s, showing he dis ibu ion o p edic ed p o eins in o majo amilies. J. Fungi 2022,8, 52 16 o 18 Re e ences 1. D umonde-Ne es, J.; F anco-Dua e, R.; Lima, T.; Schulle , D.; Pais, C. Associa ion be ween g ape yeas communi ies and he ineya d ecosys ems. PLoS ONE 2017,12, e0169883. [C ossRe ] [PubMed] 2. D umonde-Ne es, J.; F anco-Dua e, R.; Lima, T.; Schulle , D.; Pais, C. Yeas biodi e si y in ineya d en i onmen s is inc eased by human in e en ion. PLoS ONE 2016,11, e0160579. [C ossRe ] 3. D umonde-Ne es, J.; F anco-Dua e, R.; Viei a, E.; Mendes, I.; Lima, T.; Schulle , D.; Pais, C. Di e en ia ion o Saccha omyces ce e isiae popula ions om ineya ds o he Azo es A chipelago: Geog aphy s. Ecology. Food Mic obiol. 2018 ,74, 151–162. [C ossRe ] [PubMed] 4. D umonde-Ne es, J.; ˇ Cadež, N.; Domínguez, Y.R.; Gallme ze , A.; Schulle #, D.; Lima, T.; Pais, C.; F anco-Dua e, R. Cla ispo a san aluciae .a., sp. no ., a no el ascomyce ous yeas species isola ed om g apes. In . J. Sys . E ol. Mic obiol. 2020 ,70, 6307–6312. [C ossRe ] 5. Gá donyi, M.A.; Ös e be g, M.A.; Rod igues, C.; Spence -Ma ins, I.; Hahn-Häge dal, B. High capaci y xylose anspo in Candida in e media PYCC 4715. FEMS Yeas Res. 2003,3, 45–52. [C ossRe ] 6. Geije , C.; Fa ia-Oli ei a, F.; Mo eno, A.D.; S enbe g, S.; Mazu kewich, S.; Olsson, L. Genomic and ansc ip omic analysis o Candida in e media e eals he gene ic de e minan s o i s xylose-con e ing capaci y. Bio echnol. Bio uels 2020 ,13, 48. [C ossRe ] [PubMed] 7. Mo eno, A.D.; Tomás-Pejó, E.; Olsson, L.; Geije , C. Candida in e media CBS 141442: A no el glucose/xylose co- e men ing isola e o lignocellulosic bioe hanol p oduc ion. Ene gies 2020,13, 5363. [C ossRe ] 8. D umonde-Ne es, J.; Fe nandes, T.; Lima, T.; Pais, C.; F anco-Dua e, R. Lea ning om 80 yea s o s udies: A comp ehensi e ca alogue o non-Saccha omyces yeas s associa ed wi h i icul u e and winemaking. FEMS Yeas Res. 2021 ,21, oab017. [C ossRe ] 9. Mingo ance-Cazo la, L.; Clemen e-Jiménez, J.M.; Ma ínez-Rod íguez, S.; Las He as-Vázquez, F.J.; Rod íguez-Vico, F. Con i- bu ion o di e en na u al yeas s o he a oma o wo alcoholic be e ages. Wo ld J. Mic obiol. Bio echnol. 2003 ,19, 297–304. [C ossRe ] 10. Mo a a, A.; Loi a, I.; Esco , C.; del F esno, J.M.; Bañuelos, M.A.; Suá ez-Lepe, J.A. Applica ions o Me schnikowia pulche ima in wine bio echnology. Fe men a ion 2019,5, 63. [C ossRe ] 11. Daniel, H.M.; Lachance, M.A.; Ku zman, C.P. On he eclassi ica ion o species assigned o Candida and o he anamo phic ascomyce ous yeas gene a based on phylogene ic ci cumsc ip ion. An onie Leeuwenhoek 2014,106, 67–84. [C ossRe ] [PubMed] 12. Ku zman, C.P.; Robne , C.J.; Basehoa , E.; Wa d, T.J. Fou new species o Me schnikowia and he ans e o se en Candida species o Me schnikowia and Cla ispo a as new combina ions. An onie Leeuwenhoek 2018 ,111, 2017–2035. [C ossRe ] [PubMed] 13. Lachance, M.-A.; Hu ado, E.; Hsiang, T. A s able phylogeny o he la ge-spo ed Me schnikowia clade. Yeas 2016 ,33, 261–275. [C ossRe ] [PubMed] 14. Shen, X.X.; Opulen e, D.A.; Kominek, J.; Zhou, X.; S eenwyk, J.L.; Buh, K.V.; Haase, M.A.B.; Wiseca e , J.H.; Wang, M.; Doe ing, D.T .; e al. Tempo and Mode o Genome E olu ion in he Budding Yeas Subphylum. Cell 2018 ,175, 1533–1545.e20. [C ossRe ] 15. Schwa z, K.; She lock, G. P epa a ion o yeas DNA sequencing lib a ies. Cold Sp ing Ha b. P o oc. 2016 ,2016, 871–876. [C ossRe ] 16. Ko en, S.; Walenz, B.P.; Be lin, K.; Mille , J.R.; Be gman, N.H.; Phillippy, A.M. Canu: Scalable and accu a e long- ead assembly ia adap i e κ-me weigh ing and epea sepa a ion. Genome Res. 2017,27, 722–736. [C ossRe ] [PubMed] 17. Zimin, A.V.; Ma çais, G.; Puiu, D.; Robe s, M.; Salzbe g, S.L.; Yo ke, J.A. The MaSuRCA genome assemble . Bioin o ma ics 2013, 29, 2669–2677. [C ossRe ] [PubMed] 18. Zimin, A.V.; Salzbe g, S.L. The genome polishing ool POLCA makes as and accu a e co ec ions in genome assemblies. PLoS Compu . Biol. 2020,16, e1007981. [C ossRe ] [PubMed] 19. Alonge, M.; Soyk, S.; Ramak ishnan, S.; Wang, X.; Goodwin, S.; Sedlazeck, F.J.; Lippman, Z.B.; Scha z, M.C. RaGOO: Fas and accu a e e e ence-guided sca olding o d a genomes. Genome Biol. 2019,20, 224. [C ossRe ] 20. Mikheenko, A.; P jibelski, A.; Sa elie , V.; An ipo , D.; Gu e ich, A. Ve sa ile genome assembly e alua ion wi h QUAST-LG. Bioin o ma ics 2018,34, i142–i150. [C ossRe ] 21. Weib, C.L.; Pais, M.; Cano, L.M.; Kamoun, S.; Bu bano, H.A. nQui e: A s a is ical amewo k o ploidy es ima ion using nex gene a ion sequencing. BMC Bioin o m. 2018,19, 122. [C ossRe ] 22. Seppey, M.; Manni, M.; Zdobno , E. BUSCO: Assessing Genome Assembly and Anno a ion Comple eness. In Me hods in Molecula Biology; Humana: New Yo k, NY, USA, 2019; Volume 1962, pp. 227–245. ISBN 9781493991730. 23. Yoon, S.-H.; Ha, S.-M.; Lim, J.; Kwon, S.; Chun, J. A la ge-scale e alua ion o algo i hms o calcula e a e age nucleo ide iden i y. An onie Leeuwenhoek 2017,110, 1281–1286. [C ossRe ] [PubMed] 24. S anke, M.; Schö mann, O.; Mo gens e n, B.; Waack, S. Gene p edic ion in euka yo es wi h a gene alized hidden Ma ko model ha uses hin s om ex e nal sou ces. BMC Bioin o m. 2006,7, 62. [C ossRe ] [PubMed] 25. S anke, M.; Mo gens e n, B. AUGUSTUS: A web se e o gene p edic ion in euka yo es ha allows use -de ined cons ain s. Nucleic Acids Res. 2005,33, W465–W467. [C ossRe ] 26. Cui, X.; Lu, Z.; Wang, S.; Jing-Yan Wang, J.; Gao, X. CMsea ch: Simul aneous explo a ion o p o ein sequence space and s uc u e space imp o es no only p o ein homology de ec ion bu also p o ein s uc u e p edic ion. Bioin o ma ics 2016 ,32, i332–i340. [C ossRe ] J. Fungi 2022,8, 52 17 o 18 27. A ias-Ca asco, R.; Vásquez-Mo án, Y.; Nakaya, H.I.; Ma acaja-Cou inho, V. S uc RNA inde : An au oma ed pipeline and web se e o RNA amilies p edic ion. BMC Bioin o m. 2018,19, 55. [C ossRe ] [PubMed] 28. G i i hs-Jones, S.; Moxon, S.; Ma shall, M.; Khanna, A.; Eddy, S.R.; Ba eman, A. R am: Anno a ing non-coding RNAs in comple e genomes. Nucleic Acids Res. 2005,33, D121–D124. [C ossRe ] 29. Can alapied a, C.P.; He nández-Plaza, A.; Le unic, I.; Bo k, P.; Hue a-Cepas, J. eggNOG-mappe 2: Func ional Anno a ion, O hology Assignmen s, and Domain P edic ion a he Me agenomic Scale. Mol. Biol. E ol. 2021,38, 5825–5829. [C ossRe ] 30. Ta uso , R.L.; Galpe in, M.Y.; Na ale, D.A.; Koonin, E.V. The COG da abase: A ool o genome-scale analysis o p o ein unc ions and e olu ion. Nucleic Acids Res. 2000,28, 33–36. [C ossRe ] 31. Kanehisa, M. KEGG: Kyo o Encyclopedia o Genes and Genomes. Nucleic Acids Res. 2000,28, 27–30. [C ossRe ] [PubMed] 32. Kanehisa, M.; Sa o, Y.; Kawashima, M. KEGG mapping ools o unco e ing hidden ea u es in biological da a. P o ein Sci. 2021 . [C ossRe ] [PubMed] 33. Can a el, B.I.; Cou inho, P.M.; Rancu el, C.; Be na d, T.; Lomba d, V.; Hen issa , B. The Ca bohyd a e-Ac i e EnZymes da abase (CAZy): An expe esou ce o glycogenomics. Nucleic Acids Res. 2009,37, D233–D238. [C ossRe ] [PubMed] 34. Kanehisa, M.; Sa o, Y. KEGG Mappe o in e ing cellula unc ions om p o ein sequences. P o ein Sci. 2020 ,29, 28–35. [C ossRe ] [PubMed] 35. San iago, C.; Ri o, T.; Viei a, D.; Fe nandes, T.; Pais, C.; Sousa, M.J.; Soa es, P.; F anco-Dua e, R. Imp o emen o o ulaspo a delb ueckii genome anno a ion: Towa ds he exploi a ion o genomic ea u es o a bio echnologically ele an yeas . J. Fungi 2021 , 7, 287. [C ossRe ] 36. Sun, Y.B. FasPa se : A package o manipula ing sequence da a. Zool. Res. 2017,38, 110–112. [C ossRe ] 37. Nguyen, L.T.; Schmid , H.A.; Von Haesele , A.; Minh, B.Q. IQ-TREE: A as and e ec i e s ochas ic algo i hm o es ima ing maximum-likelihood phylogenies. Mol. Biol. E ol. 2015,32, 268–274. [C ossRe ] 38. Du ens, P.; Klopp, C.; Bi eau, N.; Fi on-Ouhabi, V.; Demen hon, K.; Accocebe y, I.; She man, D.J.; Noël, T. Genome Sequence o he Yeas Cla ispo a lusi aniae Type S ain CBS 6936. Genome Announc. 2017,5, 30–31. [C ossRe ] 39. Kannan, A.; Asne , S.A.; T achsel, E.; Kelly, S.; Pa ke , J.; Sangla d, D. Compa a i e Genomics o he Elucida ion o Mul id ug Resis ance in Candida lusi aniae. MBio 2019,10, e02512-19. [C ossRe ] 40. Mo eno, A.D.; Tellg en-Ro h, C.; Sole , L.; Daina , J.; Olsson, L.; Geije , C. Comple e Genome Sequences o he Xylose-Fe men ing Candida in e media S ains CBS 141442 and PYCC 4715. Genome Announc. 2017,5, e00138-17. [C ossRe ] 41. Ga ba ino, J.E.; Gibbons, I.R. Exp ession and genomic analysis o midasin, a no el and highly conse ed AAA p o ein dis an ly ela ed o dynein. BMC Genom. 2002,3, 18. [C ossRe ] 42. F anco-Dua e, R.; Bigey, F.; Ca e o, L.; Mendes, I.; Dequin, S.; San os, M.A.S.; Pais, C.; Schulle , D. In as ain genomic and pheno ypic a iabili y o he comme cial Saccha omyces ce e isiae s ain Zyma lo e VL1 e eals mic oe olu iona y adap a ion o ineya d en i onmen s. FEMS Yeas Res. 2015,15, o 063. [C ossRe ] [PubMed] 43. F anco-Dua e, R.; Umek, L.; Zupan, B.; Schulle , D. Compu a ional app oaches o he gene ic and pheno ypic cha ac e iza ion o a Saccha omyces ce e isiae wine yeas collec ion. Yeas 2009,26, 675–692. [C ossRe ] 44. F anco-Dua e, R.; Umek, L.; Mendes, I.; Cas o, C.C.; Fonseca, N.; Ma ins, R.; Sil a-Fe ei a, A.C.; Sampaio, P.; Pais, C.; Schulle , D . New in eg a i e compu a ional app oaches un eil he Saccha omyces ce e isiae pheno-me abolomic e men a i e p o ile and allow s ain selec ion o winemaking. Food Chem. 2016,211, 509–520. [C ossRe ] 45. Da ies, G.J.; Glos e , T.M.; Hen issa , B. Recen s uc u al insigh s in o he expanding wo ld o ca bohyd a e-ac i e enzymes. Cu . Opin. S uc . Biol. 2005,15, 637–645. [C ossRe ] 46. Piombo, E.; Sela, N.; Wisniewski, M.; Ho mann, M.; Gullino, M.L.; Alla d, M.W.; Le in, E.; Spada o, D.; D oby, S. Genome sequence, assembly and cha ac e iza ion o wo Me schnikowia uc icola s ains used as biocon ol agen s o pos ha es diseases. F on . Mic obiol. 2018,9, 593. [C ossRe ] 47. Fe nandes, T.; Sil a-Sousa, F.; Pe ei a, F.; Ri o, T.; Soa es, P.; F anco-Dua e, R.; Sousa, M.J. Bio echnological Impo ance o To ulaspo a delb ueckii: F om he Obscu i y o he Spo ligh . J. Fungi 2021,7, 712. [C ossRe ] 48. Sahay, S. Wine enzymes: Po en ial and p ac ices. Enzym. Food Bio echnol. P od. Appl. Fu u . P ospec . 2018, 73–92. [C ossRe ] 49. Daenen, L.; Saison, D.; S e ckx, F.; Del aux, F.R.; Ve ach e , H.; De delinckx, G. Sc eening and e alua ion o he glucoside hyd olase ac i i y in Saccha omyces and B e anomyces b ewing yeas s. J. Appl. Mic obiol. 2008 ,104, 478–488. [C ossRe ] [PubMed] 50. S eenwyk, J.L.; Opulen e, D.A.; Kominek, J.; Shen, X.X.; Zhou, X.; Labella, A.L.; B adley, N.P.; Eichman, B.F.; ˇ Cadež, N.; Libkind, D .; e al. Ex ensi e loss o cell-cycle and DNA epai genes in an ancien lineage o bipola budding yeas s. PLoS Biol. 2019,17, e3000255. [C ossRe ] [PubMed] 51. Dujon, B.; She man, D.; Fische , G.; Du ens, P.; Casa egola, S.; La on aine, I.; de Mon igny, J.; Ma ck, C.; Neu église, C.; Talla, E.; e al. Genome e olu ion in yeas s. Na u e 2004,430, 35–44. [C ossRe ] [PubMed] 52. Scannell, D.R.; Bu le , G.; Wol e, K.H. Yeas genome e olu ion— he o igin o he species. Yeas 2007,24, 929–942. [C ossRe ] 53. Gómez, S.; Be dugo, S.; Mena, R. Occu ence o indigenous a buscula myco hizal ungi associa ed wi h he hizosphe e o he naidípalm in Colombia. Cienc. Tecnol. Ag opecu. 2020,21, e1275. 54. Res epo-Co ea, S.; Pineda-Meneses, E.; Rios-Oso io, L. Mechanisms o ac ion o ungi AND bac e ia used as bio e ilize s in ag icul u al soils: A sys ema ic e iew. Co poica. Tecnol. Ag opecu. 2017,18, 335–351. J. Fungi 2022,8, 52 18 o 18 55. Youdkes, D.; Helman, Y.; Bu dman, S.; Ma an, O.; Ju ke i ch, E. Po en ial con ol o po a o so o disease by he obliga e p eda o s bdello ib io and like o ganisms. Appl. En i on. Mic obiol. 2020,86, e02543-19. [C ossRe ] [PubMed] 56. Leon-T acca, B.; A é alo-Ga dini, E.; Bouchon, A.S. Sudden dea h o Theob oma cacao L. caused by Ve icillium dahliae Kleb. In Pe u and i s in i o biocon ol. Cienc. Tecnol. Ag opecu. 2019,20, 133–148. [C ossRe ] 57. López-He nández, F.; Co és, A.J. Las -Gene a ion Genome–En i onmen Associa ions Re eal he Gene ic Basis o Hea Tole ance in Common Bean (Phaseolus ulga is L.). F on . Gene . 2019,10, 954. [C ossRe ] [PubMed] 58. Blai , M.W.; Co és, A.J.; Fa me , A.D.; Huang, W.; Ambachew, D.; Va ma Penme sa, R.; Ca asquilla-Ga cia, N.; Asse a, T.; Cannon, S.B. Une en ecombina ion a e and linkage disequilib ium ac oss a e e ence SNP map o common bean (Phaseolus ulga is L.). PLoS ONE 2018,13, e0189597. [C ossRe ] [PubMed] 59. Co és, A.J.; López-He nández, F.; Oso io-Rod iguez, D. P edic ing The mal Adap a ion by Looking In o Popula ions’ Genomic Pas . F on . Gene . 2020,11, 564515. [C ossRe ] 60. Co és, A.J.; López-He nández, F. Ha nessing c op wild di e si y o clima e change adap a ion. Genes 2021 ,12, 783. [C ossRe ]