MS data linked to manuscript: https://doi.org/10.7554/eLife.109154.1
Full text
Automated genome mining predicts structural diversity and taxonomic distribution of peptide metallophores across bacteria Zachary L. Reitz* 1,2, Bita Pourmohsenin* 3, Melanie Susman4, Emil Thomsen4, Daniel Roth4, Alison Butler4, Nadine Ziemert3, Marnix H. Medema# 1 1. Bioinformatics Group, Wageningen University, 6708 PB Wageningen, the Netherlands 2. Department of Evolution, Ecology, and Marine Biology, University of California, Santa Barbara, CA 93117, USA 3. Interfaculty Institute of Microbiology and Infection Medicine, Institute for Bioinformatics and Medical Informatics, German Centre for Infection Research (DZIF), Tübingen, Germany 4. Department of Chemistry and Biochemistry, University of California, Santa Barbara, CA 93117, USA * Contributed equally # Corresponding authors Supplemental Discussion and Figures Supplemental discussion 1: Details of the detection strategy........................................................2 Figures S1 - S2 Supplemental discussion 2: A BGC from Sporomusa termitida putatively encoding for a menaquinone-related chelating group........................................................................................... 6 Figure S3 Figure S4. Similarity network of complete NRP metallophore BGC regions from RefSeq representative genomes................................................................................................................ 9 Supplemental methods and results for siderophore isolations.................................................... 10 Figures S5-S16 Table S4 Supplemental references.............................................................................................................21 Tables S1-S3 are located in the separate supplemental Excel (.xlsx) file 1
Supplemental discussion 1: Details of the detection strategy Generally, draft pHMMs were built from alignments of known and predicted NRP metallophore biosynthesis genes collected from literature (Fig. S1). Initial bitscore cutoffs were determined by scanning MIBiG 2.014 gene clusters, as well as 38 additional experimentally characterized NRP metallophore BGCs that we added to MIBiG 3.0 (Supplemental Table 1).15 Final cutoffs were determined by scanning 28,688 NRPS cluster regions from the antiSMASH database and manually inspecting hits near the initial cutoff to determine if they are likely true NRP metallophore BGCs based on the presence of genes encoding membrane transport, metal acquisition (ex: ferric reductases16–18), and multiple chelating groups. The pHMMs were iteratively refined where required by adding additional low-scoring putative true positives to the seed alignments until clear bitscore separations appeared between (putative) true and false hits (Fig. S1). Detection of catechols and phenols Like the aromatic amino acids, 2,3-dihydroxybenzoic acid (2,3-DHB) and salicylic acid are derived from chorismate (Figure 1B).19–21 The three-step synthesis of 2,3-DHB is catalyzed by an isochorismate synthase (EntC), an isochorismatase (EntB), and a 2,3-dihydro-2,3dihydroxybenzoate dehydrogenase (EntA).19 The sole presence of EntA matches in a cluster was sufficient to detect 2,3-DHB-containing metallophore BGCs; however, only EntC could be used to eliminate the nearly-identical 3-hydroxyanthranilic acid pathway,22 and thus both EntA and EntC were required to accurately detect 2,3-DHB production. Salicylate biosynthesis was detected by the presence of either an isochorismate pyruvate-lyase (IPL)20 or a bifunctional salicylate synthase (SalSyn).21 Two condensation domain subtypes specific to catecholic and phenolic metallophores were also included as independent detection rules: VibH-like enzymes (VibH), which condense 2,3-DHB to diamines and polyamines,23,24 and tandem heterocyclization domains (Cy_tandem) are sometimes responsible for Ser, Thr, or Cys cyclization.25 Detection of β-hydroxyamino acids An analysis of siderophores containing β-hydroxy-aspartate (β-OHAsp) and β-hydroxy-histidine (β-OHHis) delineated three families of siderophore-specific Fe(II)/α-ketoglutarate-dependent enzymes responsible for β-hydroxylation of Asp (TBH_Asp and IBH_Asp) or His (IBH_His).26 More recently, β-OHAsp-containing siderophores named cyanochelins were isolated from a diverse group of cyanobacteria.27 The BGCs contain two additional classes of β-hydroxylases, 2
tentatively named CyanoBH_Asp1 and CyanoBH_Asp2. Because of the polyphyletic nature of these hydroxylases, well-performing pHMMs could not be made using the standard workflow used for other pathways (Figure S1); even after many iterations, non-metallophore NRPs were captured as false positives, such as the β-OHAsp-containing lipopeptide turnercyclamycin.28 Instead, pHMM construction was guided by a phylogenetic tree of β-hydroxylases from NRPS BGCs (Figure S2). Clades likely to be involved in metallophore biosyntheses were located using known genes from reported BGCs,26,27 and by the presence of nearby metallophore-specific transporter families (vide infra). Several unreported clades of amino acid β-hydroxylases were identified that may also be involved in metallophore biosynthesis based on their co-occurrence with transporter genes (Figure S2). However, these putative β-hydroxylases were not included in the current NRP metallophore rules due to a lack of experimental evidence for metallophore production. A negative constraint (in the form of a competing pHMM for SBH_Asp) was used to exclude the syringomycin family of BGCs, which contain a clade of β-hydroxylases that sit within the IBH_Asp clade.26 Detection of hydroxamic acids Hydroxamates are all produced by the hydroxylation and acylation of a primary amine. In peptidic metallophores, the amine is usually the sidechain of ornithine (Orn) or Lys, which is first hydroxylated by a flavin-dependent monooxygenase (Orn_monoox or Lys_monoox, respectively). Vicibactin is an unusual hydroxamate siderophore29 that could not be accurately captured by either Orn_monoox or Lys_monoox. The Orn hydroxylase VbsO is more similar to those found in NIS pathways, and a VbsO pHMM was non-specific. Instead, the acyl-hydroxyornithine epimerase29 VbsL is used to detect vicibactin. In two non-metallophore pathways,30,31 hydroxyornithine or hydroxylysine are the substrates for N-N bond-forming enzymes. No clear bitscore separation could be achieved to eliminate these false positives, so two negative constraints were added: the presence of KtzT associated with biosynthesis of piperazates,30 and MetRS-like, associated with several hydrazines.31 Unfortunately, false positives may still arise if the ornithine monooxygenase and the negative constraint genes are (distantly) located within the same BGC region (>30 kbp), such as in the himastatin locus (MIBiG BGC0001117). Detection of other chelating groups Although the biosynthesis of the chelating amino acid graminine has not been fully elucidated, gene knockouts and stable isotope studies have revealed two enzymes, GrbD and GrbE, 3
responsible for diazeniumdiolate formation from arginine.32,33 The quinoline chelator Dmaq was first identified in anachelins,34 and the biosynthesis was recently established for fabrubactins.35 Synthesis is initialized by FbnL and FbnM, which form a two-protein heme peroxidase that oxidizes L-tyrosine to L-DOPA.35 GrbDE and FbnLM were previously used as a handle for genome mining,32,35 and our pHMMs gave similar results. We did not include detection of a pathway currently only reported in fabrubactins that produces two α-hydroxycarboxylate chelating moieties (Fig. 1A, bolded atoms).31 Both substructures are proposed to be synthesized by the flavin-dependent monooxygenase FbnE. A profile HMM was built for FbnE (Supplemental dataset) but was not included in antiSMASH, as all BGCs with FbnE hits were either captured independently by FbnL and FbnM (Dmaq biosynthesis), or appeared to be false positives based on a lack of other metallophore-related genes. The diverse pyoverdine family of siderophores is defined by a fluorescent chelating chromophore (Figure 1). Chromophore maturation is driven by the tyrosinase PvdP and the oxidoreductase PvdO.36,37 The two genes are sometimes in separate loci, and including both as independent pHMMs increased the number of pyoverdine BGCs detected. These pHMMs also capture the biosynthesis pathway for azotobactin, which contains a slightly different chromophore (Figure 3).38 4
Figure S1. Workflow for developing an NRP-metallophore-specific profile hidden Markov model (pHMM) and significance score cutoff for an enzyme (sub-)family. (1) Examples are collected from literature, the amino acid sequences are aligned with MUSCLE and (2) a pHMM is constructed with HMMER3. (3) BGCs of known function from MIBiG are scanned for matches to the pHMM to generate a preliminary bitscore cutoff. (4) NRPS BGC regions from the antiSMASH database are scanned for matches to the pHMM and sorted by bitscore. (5) Starting at the bitscore cutoff, BGCs are manually annotated to predict if they encode the biosynthesis of NRP metallophores using features such as genes encoding membrane transport, metal acquisition (ex: ferric reductases), and the biosynthesis of multiple chelating groups. In a properly functioning system, the bitscore cutoff delineates putative true and false positives to accurately detect the NRP-metallophore-related enzymes (bottom right). However, a bitscore cutoff adjustment may be required (bottom middle), or low-scoring true positives may need to be added to the pHMM seed alignment in an iterative process (bottom left). 5
Figure S2. A maximum-likelihood phylogeny of putative β-hydroxylases found in NRPS BGC regions. β-Hydroxylase subtypes found in characterized siderophore BGCs are highlighted and labeled; the red SBH_Asp clade consists of non-metallophore phytotoxins.1 Gray clades indicate possible metallophore β-hydroxylases that currently have no experimentally characterized representative. Amino acid sequences similar to known siderophore β-hydroxylase subtypes were extracted from the antiSMASH database (v3) and dereplicated prior to tree reconstruction. The right-hand bar gives the number of unique siderophore-related transporter families2 present in the BGC region containing the β-hydroxylase. 6
Supplemental discussion 2: A BGC from Sporomusa termitida putatively encoding for a menaquinone-related chelating group One particularly promising novel BGC was found in the genome of Sporomusa termitida DSM 4440, an anaerobic acetogen (Figure S3). The BGC was manually annotated as a putative metallophore due to the presence of a TonB-dependent outer membrane receptor, as well as NRPS domains, methyltransferases, and oxidoreductases homologous to those involved in the biosynthesis of pyochelin and other thiazol(id)ine metallophores. A salicylate synthase gene, present in similar clusters, appears to have been replaced by a partial menaquinone pathway (menFDEB). We thus predicted that the salicylate moiety would be replaced by 1,4-dihydroxy-2-naphthoic acid (Figure S3). The Natural Product Atlas contained one family of bacterial compounds with that substructure, karamomycins, which were isolated from an unsequenced strain of Nonomuraea endophytica (Figure S3).42 Karamomycins also contain thiazol(id)ines, as predicted for the S. termitida DSM 4440 BGC product. Hence, karamomycins are likely produced by a homologous BGC, and both compounds may be involved in trace metal binding and transport. 7
Figure S3 (next page). Analysis of a novel putative NRP metallophore BGC from Sporomusa termitida DSM 4440. (A) Clinker comparison of the S. termitida BGC with homologous loci. The Sporomusa sp. KB1 BGC contains a salicylate synthase gene detectable by the new antiSMASH rules, while S. termitida DSM 4440 instead contains several genes homologous to the menaquinone locus of Desulfitobacterium spp.3 The cluster comparison was generated with clinker v0.0.26.4 (B) A proposed metallophore biosynthesis pathway encoded by the S. termitida DSM 4440 BGC. 1,4-Dihydroxy-2-naphthoic acid is synthesized from chorismic acid by homologs of MenFDHBE, encoded by SPTER_RS05985-06005, and MenC, encoded elsewhere in the genome (SPTER_RS21050). The C4 phenol is likely methylated by O-methyltransferase SPTER_RS06015. The naphthoic acid moiety of Karamomycin C (inset box), characterized from an unsequenced strain of Nonomuraea endophytica,5 is predicted to be synthesized by a similar pathway. Five NRPS genes in S. termitida DSM 4440 encode for the biosynthesis of the core structure by the condensation and cyclization of four Cys residues. The C-methyltransferase of SPTER_RS06025 is predicted to be inactive, as observed in the homologous ulbactin pathway.6 The completed scaffold is released by thioesterase SPTER_RS05970, possibly producing the same tricyclic substructure observed in karamomycin C and ulbactin F (inset box). The final predicted structure accounts for the actions of two thiazoline reductases (SPTER_RS05950 and SPTER_RS05980) and a methyltransferase (SPTER_RS05955), although the regiochemistry and timing of these transformations is unclear. 8
9
Figure S8. UPLC-ESI-MS TIC the peaks with molecular ions consistent with marinobactin A-E are labeled (top). Mass spectra of the labeled peaks in the TIC (bottom). The major peaks present are consistent with m/z value of Marinobactin A-E, 932.4986 m/z, 958.5095 m/z, 960.5315 m/z, 986.5422, and 988.5549 m/z, respectively. 16
Figure S9. ESI-MS/MS fragmentation of marinobactin A [M+H]+, 932 m/z. The b fragments 655.3563 m/z and 742.3990 m/z support the assignment of marinobactin A. Figure S10. ESI-MS/MS fragmentation of marinobactin B [M+H]+, 958 m/z. The b fragments 681.3822 m/z and 768.4156 m/z support the assignment of marinobactin B. 17
Figure S11. ESI-MS/MS fragmentation of marinobactin C [M+H]+, 960 m/z. The b fragments 681.3985 m/z and 768.4298 m/z support the assignment of marinobactin C. Figure S12. ESI-MS/MS of marinobactin D [M+H]+, 986 m/z. The b fragments 709.4139 m/z and 796.4473 m/z support the assignment of marinobactin D. 18
Figure S13. UPLC-MS spectrum of ornicorrugatin siderophore from Pseudomonas brassicacearum DSM 13227. The singly charged mass of m/z 1012.4 Da [M+H]+ and the doubly charged mass, m/z 506.7 Da [M+2H]2+, are consistent with the mass of the charged ornicorrugatin siderophore, m/z 1012.4699 Da [M+H]+. 19
Figure S14. LC-MS tandem MS spectrum of ornicorrugatin siderophore from Pseudomonas brassicacearum DSM 13227. Average collision energies of (top) 60 eV and (bottom) 35 eV. 20
Table S4: Selected fragment ions observed for ornicorrugatin produced by Pseudomonas brassicacearum DSM 13227. All given fragment ions agree with previously reported masses for ornicorrugatin (Matthijs, et al., 2008). Fragment composition Experimental mass (m/z) Calculated mass (m/z) C19H34N5O3+ 380.2681 380.2662 C23H38N7O4+ 476.3011 476.2985 C31H57N11O8+ 711.4194 711.4392 C35H60N11O13+ 842.4422 842.4372 C37H62N11O16+ 916.4429 916.4376 C41H62N13O15+ 976.4544 976.4488 Ornicorrugatin M+H, C41H66N13O17+ 1012.4760 1012.4699 Figure S15. UPLC-MS spectrum of pyoverdine siderophore from Pseudomonas brassicacearum DSM 13227. The singly charged mass of m/z 1134.4 Da [M+H]+ and the doubly charged mass of m/z 567.7 Da [M+2H]2+ are consistent with the m/z values of the pyoverdine, m/z 1134.4339 Da [M+H]+. 21
Figure S16. LC-MS tandem MS spectrum of pyoverdine siderophore, pyoverdine A214, from Pseudomonas brassicacearum DSM 13227. 22
Supplemental references (1) Reitz, Z. L.; Hardy, C. D.; Suk, J.; Bouvet, J.; Butler, A. Genomic Analysis of Siderophore β-Hydroxylases Reveals Divergent Stereocontrol and Expands the Condensation Domain Family. Proc. Natl. Acad. Sci. U. S. A. 2019, 116 (40), 19805–19814. https://doi.org/10.1073/pnas.1903161116. (2) Crits-Christoph, A.; Bhattacharya, N.; Olm, M. R.; Song, Y. S.; Banfield, J. F. Transporter Genes in Biosynthetic Gene Clusters Predict Metabolite Characteristics and Siderophore Activity. Genome Res. 2020, 31, 239–250. https://doi.org/10.1101/gr.268169.120. (3) Kruse, T.; Goris, T.; Maillard, J.; Woyke, T.; Lechner, U.; de Vos, W.; Smidt, H. Comparative Genomics of the Genus Desulfitobacterium. FEMS Microbiol. Ecol. 2017, 93 (12). https://doi.org/10.1093/femsec/fix135. (4) Gilchrist, C. L. M.; Chooi, Y.-H. Clinker & Clustermap.js: Automatic Generation of Gene Cluster Comparison Figures. Bioinformatics 2021. https://doi.org/10.1093/bioinformatics/btab007. (5) Shaaban, K. A.; Shaaban, M.; Rahman, H.; Grün-Wollny, I.; Kämpfer, P.; Kelter, G.; Fiebig, H.-H.; Laatsch, H. Karamomycins A-C: 2-Naphthalen-2-Yl-Thiazoles from Nonomuraea Endophytica. J. Nat. Prod. 2019, 82 (4), 870–877. https://doi.org/10.1021/acs.jnatprod.8b00928. (6) Komaki, H.; Hosoyama, A.; Ichikawa, N.; Igarashi, Y. Draft Genome Sequence of a Sponge-Derived Brevibacillus Sp. TP-B0800, a Producer of Ulbactins with Tumor Cell Migration Inhibitory Activity. Gene Reports 2016, 5, 140–143. https://doi.org/10.1016/j.genrep.2016.10.006. 23