scieee AI-readable full text Open interactive document viewer

Global Spore Sampling Project : A global, standardized dataset of airborne fungal DNA

Ovaskainen, Otso,Abrego, Nerea,Furneaux, Brendan,Hardwick, Bess,Somervuo, Panu,Palorinne, Isabella,Andrew, Nigel R.,Babiy, Ulyana V.,Bao, Tan,Bazzano, Gisela,Bondarchuk, Svetlana N.,Bonebrake, Timothy C.,Brennan, Georgina L.,Bret-Harte, Syndonia,Bässler,

Full text

This is a self-archived version of an original article. This version may differ from the original in pagination and typographic details. Author(s): Title: Year: Version: Copyright: Rights: Rights url: Please cite the original version: CC BY 4.0 https://creativecommons.org/licenses/by/4.0/ Global Spore Sampling Project : A global, standardized dataset of airborne fungal DNA © 2024 the Authors Published version Ovaskainen, Otso; Abrego, Nerea; Furneaux, Brendan; Hardwick, Bess; Somervuo, Panu; Palorinne, Isabella; Andrew, Nigel R.; Babiy, Ulyana V.; Bao, Tan; Bazzano, Gisela; Bondarchuk, Svetlana N.; Bonebrake, Timothy C.; Brennan, Georgina L.; Bret-Harte, Syndonia; Bässler, Claus; Cagnolo, Luciano; Cameron, Erin K.; Chapurlat, Elodie; Creer, Simon; D’Acqui, Luigi P.; de Vere, Natasha; Desprez- Loustau, Marie-Laure; Dongmo, Michel A. K.; Dyrholm Jacobsen, Ida B.; Fisher, Brian L.; Flores de Jesus, Miguel; Gilbert, Gregory S.; Griffith, Gareth W.; Gritsuk, Anna A.; Gross, Andrin; Grudd, Håkan; Halme, Panu; Hanna, Rachid; Hansen, Jannik; Hansen, Lars Holst; Hegbe, Apollon D. M. T.; Hill, Sarah; Hogg, Ian D.; Hultman, Jenni; Hyde, Kevin D.; Hynson, Nicole A.; Ivanova, Natalia; Karisto, Petteri; Kerdraon, Deirdre; Knorre, Anastasia; Krisai-Greilhuber, Irmgard; Kurhinen, Juri; Kuzmina, Masha; Lecomte, Nicolas; Lecomte, Erin; Loaiza, Viviana; Lundin, Erik; Meire, Alexander; Mešić, Armin; Miettinen, Otto; Monkhause, Norman; Mortimer, Peter; Müller, Jörg; Nilsson, R. Henrik; Nonti Puani, Yannick C.; Nordén, Jenni; Nordén, Björn; Paz, Claudia; Pellikka, Petri; Pereira, Danilo; Petch, Geoff; Pitkänen, Juha-Matti; Popa, Flavius; Potter, Caitlin; Purhonen, Jenna; Pätsi, Sanna; Rafiq, Abdullah; Raharinjanahary, Dimby; Rakos, Niklas; Rathnayaka, Achala R.; Raundrup, Katrine; Rebriev, Yury A.; Rikkinen, Jouko; Rogers, Hanna M. K.; Rogovsky, Andrey; Rozhkov, Yuri; Runnel, Kadri; Saarto, Annika; Savchenko, Anton; Schlegel, Markus; Schmidt, Niels Martin; Seibold, Sebastian; Skjøth, Carsten; Stengel, Elisa; Sutyrina, Svetlana V.; Syvänperä, Ilkka; Tedersoo, Leho; Timm, Jebidiah; Tipton, Laura; Toju, Hirokazu; Uscka- Perzanowska, Maria; van der Bank, Michelle; van der Bank, Herman F.; Vandenbrink, Bryan; Ventura, Stefano; Vignisson, Solvi R.; Wang, Xiaoyang; Weisser, Wolfgang W.; Wijesinghe, Subodini N.; Joseph, Wright S.; Yang, Chunyan; Yorou, Nourou S.; Young, Amanda; Yu, Douglas W.; Zakharov, Evgeny V.; Hebert, Paul D. N.; Roslin, Tomas Ovaskainen, O., Abrego, N., Furneaux, B., Hardwick, B., Somervuo, P., Palorinne, I., Andrew, N. R., Babiy, U. V., Bao, T., Bazzano, G., Bondarchuk, S. N., Bonebrake, T. C., Brennan, G. L., Bret- Harte, S., Bässler, C., Cagnolo, L., Cameron, E. K., Chapurlat, E., Creer, S., . . . Roslin, T. (2024). Global Spore Sampling Project : A global, standardized dataset of airborne fungal DNA. Scientific Data, 11, Article 561. https://doi.org/10.1038/s41597-024-03410-0 2024 1 Scientific Data | (2024) 11:561 | https://doi.org/10.1038/s41597-024-03410-0 www.nature.com/scientificdata Global Spore Sampling Project: a global, standardized dataset of airborne fungal DNa Otso Ovaskainen et al.# Novel methods for sampling and characterizing biodiversity hold great promise for reevaluating patterns of life across the planet. the sampling of airborne spores with a cyclone sampler, and the sequencing of their DNA, have been suggested as an efficient and wellcalibrated tool for surveying fungal diversity across various environments. Here we present data originating from the Global Spore Sampling Project, comprising 2,768 samples collected during two years at 47 outdoor locations across the world. Each sample represents fungal DNA extracted from 24 m3 of air. We applied a conservative bioinformatics pipeline that filtered out sequences that did not show strong evidence of representing a fungal species. The pipeline yielded 27,954 species-level operational taxonomic units (OTUs). Each OTU is accompanied by a probabilistic taxonomic classification, validated through comparison with expert evaluations. to examine the potential of the data for ecological analyses, we partitioned the variation in species distributions into spatial and seasonal components, showing a strong effect of the annual mean temperature on community composition. Background & Summary Fungi are one of the most diverse and ecologically important yet unexplored kingdoms of life1. From a practical perspective, fungi are infamously hard to sample2 and characterize3. Recent advancements in DNA-based survey methods have revolutionized studies on fungal diversity, especially its large-scale patterns4–8. Given that fungi occur in nearly every possible environment and substrate, current sampling campaigns and estimates of fungal diversity tend to rely explicitly on substrate-specific sampling9. Sampling of soil has been popular given the relative ease with which the mycobiome of any handful of soil can be characterized through metabarcoding10. Yet, whether biogeographic patterns from those substrates broadly reflect patterns in fungal taxa9 or biodiversity in general11 is unclear. Additionally, there are significant biases in the geographic areas represented in global studies12,13, although there have been recent efforts to expand the coverage of understudied regions10. A recent methodological breakthrough for surveying fungi uses a cyclone sampler to capture fungal spores from the air, followed by DNA sequencing and sequence-based species identification14. Air sampling has revealed high diversity and stronger ecological signals in community composition of fungi than soil sampling15. Air sampling captures any fragments of fungi floating in the air, including the wind-dispersed spores of fungi and fragments of hyphae as well as fungal structures attached to other organisms. Consequently, air sampling detects fungal dispersal at high temporal resolution. In addition to fungal surveys, the sampling of airborne DNA has proved effective in acquiring comprehensive inventories of regional diversity of many other taxa16. Here we present a global-scale database assembled by the Global Spore Sampling Project (GSSP) that was initiated in 2018–201917. The GSSP involves 47 sampling locations distributed across all continents except Antarctica, with each location collecting two 24-hr samples per week, in most cases over a period of one year or more (Fig.1A,B). Sampling is conducted with a cyclone sampler, which orients itself in the direction of the wind. It collects particles >1 μm in size from the air directly into a sampling tube with a single reverse-flow cyclone. For DNA sequencing, we targeted part of the nuclear ribosomal internal transcribed spacer (ITS) region, which is the universal molecular barcode for fungi18. To generate semi-quantitative estimates of DNA content (in units of ng of fungal DNA per m3 of air), we applied a spiking approach17 (Fig.1C). To convert the sequence data into species data, we began by denoising #A full list of authors and their affiliations appears at the end of the paper. DATA DEScriPTOr OPEN 2 Scientific Data | (2024) 11:561 | https://doi.org/10.1038/s41597-024-03410-0 www.nature.com/scientificdata www.nature.com/scientificdata/ the sequence yield into amplicon sequence variants (ASV19). We then applied probabilistic taxonomic placement using Protax-fungi20,21 to assign ASVs to taxa at ranks from phylum to species. Finally, we used a new constrained clustering approach (see Methods) guided by the taxonomic annotations from Protax-fungi to group ASVs into species-level operational taxonomic units (OTUs22). This clustering allowed us to assign OTUs to previously known and unknown taxa (Fig.1D). Using a threshold of >90% probability of correct assignment, this resulted in 27,954 species-level OTUs, of which 1,392 could be reliably assigned to known species. The GSSP data are highly complementary to the Global Soil Mycobiome consortium (GSMc) data10, as among the 10 top ranking orders in the GSSP data, only 5 were found in the 10 top ranking orders of the GSMc data (Table1). Methods Data acquisition. The Global Spore Sampling Project (GSSP) consists of a globally distributed network of 47 sampling sites collecting two 24-hr air samples per week over one to two years (Fig.1). Each sampling site was equipped with a cyclone sampler (Burkard Cyclone Sampler for Field Operation, Burkard Manufacturing Co Ltd; http://burkard.co.uk/product/cyclone-sampler-for-field-operation). The sampling sites represent varying climatic zones and altitudes. Most sampling sites were located in natural environments, with a few in urban settings. Due to logistical reasons, we could not start the global sampling fully synchronously. In some locations, sampling had to stop earlier than expected due to external reasons (e.g., storms breaking the equipment or restrictions caused by COVID-19 lockdown). See Fig.1 for realized sampling periods per site. In October and November 2017, prior to the start of global sampling, a field test was performed in a grassy area at the University of Helsinki Viikki campus (60.2278 N, 25.01653E) to evaluate the quantity of fungal DNA collected over different time frames and in field blanks handled with and without the use of gloves on the part of the human handler. In total we collected seven 24-hour samples, three one-hour samples, and three 10-minute samples, in addition to four field blanks handled with gloves and five field blanks handled without gloves. For field blanks, Eppendorf vials were installed in the cyclone sampler in the field, but the sampler was not activated. The vials were then removed after one minute and sealed. Based on the results of these field tests (see Technical Validation), we decided to use a 24-hr sampling period, and to instruct the participating teams to handle the samples with gloves. The functioning of the cyclone sampler and sample preparation procedure is described in detail in Ovaskainen et al.17. The cyclone samplers were placed at ground level to ensure free airflow through the sampler. The sampler collected particles >1 µm in size from the air directly into a sterile Eppendorf vial. The sampler’s average throughput of air was 16.5 L per minute for a total of 23,800 L (23.8 m3) during each 24-hour sampling period. After sampling, the vial was removed from the cyclone sampler, the lid was closed, and the vials were labelled with the site code and week number. We also recorded the time and duration of the sampling, along with notes on the presence of rainwater or larger objects (e.g., arthropods) in the sampling vial. To avoid contamination, gloves were used while handling the samples and the device. Participants were instructed to clean the cyclone part of the device monthly with water and soap and to rinse it with ethanol, or to sterilize it with dry-heat, chlorine, or UV when such equipment was available. The samples were stored at −20 °C until shipped to the University of Helsinki, Finland. Shipping was done at room temperature. We do not expect much bias across samples due to this approach, as the shipping time was relatively short and most shipments were received with a similar delay. In Helsinki, the samples were separated from visible arthropods. To avoid losing fungal spores attached to arthropod bodies, the surface of any arthropod present in the sample was rinsed by adding sterile water into the sample tube and vortexing. After washing, the arthropods were removed with sterile tweezers. Samples containing any rainwater were dried in a vacuum drier (24 h). Prior to drying, each sample was covered with a porous Parafilm to avoid cross-contamination between samples. After drying, all samples were sent to the University of Guelph, Canada, for DNA extraction and sequencing. DNa extraction, sequencing, and quantifying DNa amount. A detailed description of DNA extraction, primers, and sequencing is given in Ovaskainen et al.17. In brief, the target genetic marker, i.e., the ITS2 region of the rRNA operon, was amplified using the polymerase chain reaction (PCR) for 20 cycles with fusion primers ITS_S2F23, ITS3, and ITS424 tailed with Illumina adapters, and sequenced on Illumina MiSeq with 2 × 300 bp paired end reads. ITS_S2F was included as a second forward primer to specifically amplify plant DNA, in order to include pollen as well as fungal spores in the analysis. However, only a small fraction of reads resulted from the ITS_S2F-ITS4 amplicon, and so these were removed in the early stages of the analysis and not further considered. To quantify the amount of fungal DNA, we applied a spike-in approach17, using nine positive control plasmids prepared from synthetic sequences. These sequences were designed to be generally consistent with fungal ITS sequences, but different from all known natural sequences25. The positive synthetic control (0.01 ng/μl) containing nine plasmids was spiked into the PCR master mix at a ratio of 1:100 for the first 336 samples. For the remaining 2,432 samples, we used a 1:1000 ratio, since the 1:100 ratio produced an unnecessarily high proportion of the sequences representing the spikes. This could have compromised the sequencing depth of the targeted fungal sequences. We converted the ratio of the non-spike vs. spike-sequences into semi-quantitative estimates of DNA amount in units of ng of DNA per m3 of air as described previously17. The resulting estimates of DNA abundance correlated well with a qPCR-based estimate of DNA amount. Each MiSeq run included 84 study samples, one negative control sample introduced in the DNA extraction step, and two negative controls introduced in the PCR step. The only exceptions were two runs (CCDB-35004 and CCDB-35005) which included three extraction negative controls and no PCR negative controls. The same master mix as used for the study samples, including synthetic positive controls, was also used for the negative controls. For the field test samples, DNA was extracted following the same protocol, except that 300 µL of ILB extraction buffer was used instead of 270 µL, and the final DNA extract was eluted into 35 µL of Tris buffer instead of 3 Scientific Data | (2024) 11:561 | https://doi.org/10.1038/s41597-024-03410-0 www.nature.com/scientificdata www.nature.com/scientificdata/ Fig. 1 Study design and data generation pipeline of the Global Spore Sampling Project (GSSP). (A) The sampling design includes 47 sites with a global distribution, with the greatest coverage in Europe (22 sites) and the poorest coverage in the Southern hemisphere (6 sites). The airborne fungal samples were collected by a cyclone sampler, with each sample consisting of fungal spores filtered from 24 m3 of air during the 24-hr sampling period. (B) The study design included weekly samples for a sampling period over one to two years, with some variation among the sites caused mainly by logistical constraints. The sites are ordered according to their mean annual temperature. (C) We employed a metabarcoding approach to sequence the fungal ITS2 marker and quantified the amount of fungal DNA (in units of ng of DNA per m3 of air) using a spiking approach17. (D) We employed a bioinformatics pipeline that utilized denoising to obtain amplicon sequence variants (ASVs). We then combined probabilistic taxonomic placement with a constrained clustering approach to form species-level OTUs, and to place these OTUs in a taxonomic tree to the most resolved taxonomic level possible given the limitations of sequence reference databases. This tree consists of three types of branches: taxa that could be reliably assigned to previously known (black) and novel (red) taxa, and branches that may belong to either known or novel taxa (grey). 4 Scientific Data | (2024) 11:561 | https://doi.org/10.1038/s41597-024-03410-0 www.nature.com/scientificdata www.nature.com/scientificdata/ 45 µL. Two extraction blanks were also included. A fungal DNA standard was extracted from Fleischmann’s Baker’s commercial yeast. Then, approximately one-half package of the commercial yeast was added to 50 mL warm water and proofed with sugar until the formation of active foam. Yeast DNA was extracted using an abbreviated version of the protocol described above, which omitted the initial ILB extraction buffer and homogenization in the TissueLyzer. Instead, six aliquots of 300 µL of yeast suspension were directly transferred to 900 µL each of 5 M GuSCN binding buffer, incubated at 56 °C for 1 hour in an orbital shaker, and then at 65 °C for 1 hour. The six eluates were pooled and quantified using a Qubit fluorometer with the DS DNA high sensitivity kit. The extract, which had a DNA concentration of 2.77 ng/µL, was then diluted to form standards of 1 ng/ µL, 0.1 ng/µL, 0.01 ng/µL, 0.001 ng/µL, and 0.0001 ng/µL. The test samples were quantified by real-time PCR (RT-PCR) on a LightCycler96 (Roche) as described in Ovaskainen et al.17, with two replicates of each of the standards for calibration. Bioinformatic processing. Demultiplexed paired-end reads were first trimmed using Cutadapt version 4.226. Because of low-quality base-calls at the 5′ end of R2 reads, we removed the first 16 bases from all R2 reads. We then trimmed the 3′ end of both reads with a quality threshold of 2 (i.e., remove only N’s), and the 5′ end of R2 with a quality threshold of 10. Reads were then trimmed to the ITS3-ITS4 amplicon, with a minimum 10 bp overlap and error tolerance of 0.2. Primers at the 3′ ends of both reads were optional but read pairs where the 5′ primer was not detected (including reads originating from the ITS_S2F-ITS4 amplicon) were removed. Pairs were discarded after trimming if either read was less than 100 bases or contained ambiguous bases. Reads were then further processed using DADA2 version 1.18.027. First, all pairs where either read matched to the PhiX genome were removed, along with reads where R1 contained more than 3 expected errors or R2 contained more than 5 expected errors. Reads were denoised using separate error profiles fit for each MiSeq run with default parameters, and denoised read pairs were merged to form ASVs with a minimum overlap of 10 bp and a maximum mismatch of 1 bp. An initial de novo chimera check was performed on the merged ASV table using the DADA2 “consensus” method27. A second reference-based chimera check was then performed using the “uchime_ref” option in VSEARCH version 2.22.128 with reference Sanger sequences from the UNITE v9database29, as used by the PlutoF Species Hypothesis matching pipeline30. The synthetic spike sequences were also included as references. Non-chimeric ASVs that were identical except for end gaps were combined, with the most abundant ASV sequence taken as representative. ASVs with a sequence similarity greater than 0.9 to SynMock spike sequences were identified using the “-usearch_global” command in VSEARCH 2.22.128 and labelled as spike sequences. Non-spike sequences were aligned using Infernal 1.1.431 to the covariance model for the combined 5.8 S and 28 S rRNA genes from the FunGene pipeline32 which was truncated to include only the region between the ITS3 and ITS4 primer sites. Sequences that did not match the full length of the model, or which scored less than 50, were discarded. This resulted in a 65,912 ASVs × 2,768 samples matrix, with entries representing read abundance. A taxonomic affiliation was assigned to each non-spike ASV sequence using Protax-fungi21. This procedure gives assignments at each taxonomic rank from phylum to species, along with a calibrated probability that the assignment at each rank is correct. We used the 90% probability threshold for taxonomic assignments. Additionally, because Protax-fungi does not include non-fungi in its reference database, we matched ASVs to the same UNITE Sanger sequences mentioned above using the “usearch_global” command of VSEARCH 2.22.128, with a sequence similarity threshold of 0.8. Sequences whose best match was annotated as belonging to a kingdom other than Fungi, or which had no match at the given threshold, were annotated as potential non-fungi but retained for the next clustering step. Dataset GSSP GSMc Phylum Order %rank %rank Ascomycota Capnodiales 22.0 1 0.8 19 Ascomycota Pleosporales 17.8 2 3.4 9 Basidiomycota Polyporales 10.0 3 0.5 29 Basidiomycota Agaricales 5.9 4 15.8 1 Basidiomycota Tremellales 5.5 5 1.6 16 Ascomycota Helotiales 4.2 6 6.3 3 Basidiomycota Hymenochaetales 3.2 7 0.3 42 Ascomycota Dothideales 2.4 8 0.2 50 Ascomycota Eurotiales 2.0 9 4.3 7 Ascomycota Chaetothyriales 1.8 10 2.8 10 Mortierellomycota Mortierellales 0.06 75 6.3 2 Basidiomycota Russulales 1.0 15 6.0 4 Basidiomycota Thelephorales 0.08 63 5.8 5 Ascomycota Hypocreales 1.0 14 5.0 6 Ascomycota Pezizales 0.09 60 3.4 8 Table 1. The most common orders found in the GSSP data and in the Global Soil Mycobiome consortium (GSMc) data10. The table shows the relative abundance (%) of each order, computed as the mean across samples of the fraction of reads which were assigned to it, as well as the ranking of the order in terms of its abundance. Only orders that rank in the top ten in either of the two datasets are included. 5 Scientific Data | (2024) 11:561 | https://doi.org/10.1038/s41597-024-03410-0 www.nature.com/scientificdata www.nature.com/scientificdata/ Due to frequent intraspecific sequence variants for the ITS region, ITS-based ASVs are not suitable proxies for fungal species33. Consequently, we developed a taxonomically-guided clustering approach using the taxonomic annotations from Protax-fungi to group ASVs into approximately species-level OTUs. Our approach also groups sequences, including those without existing taxonomic annotations, into clusters approximating each taxonomic rank. First, we calculated optimal single-linkage clustering thresholds for each combination of a known taxon at a rank higher than species (henceforth, the “supertaxon”) and a taxonomic rank lower than that taxon (“subrank”) using multi-class F-measure optimization as described for the tool Dnabarcoder34. However, instead of using BLAST to calculate pairwise distances, as in Dnabarcoder, we based our clusters on a sparse pairwise sequence distance matrix generated by the -calc_distmx command in USEARCH 11.0.66735, with an initial kmer dissimilarity threshold of 0.4, maximum global alignment dissimilarity of 0.6, and a gap penalty of 1. For each supertaxon-subrank combination where there were at least five subtaxa represented by a total of at least ten reference sequences, we chose the clustering threshold that generated clusters most closely corresponding to the reference identifications. This match was assessed by the multi-class F-measure. Thus, we generated optimal thresholds for clustering all fungi into ranks from phylum to species; for clustering each phylum into ranks from class to species, and so on. The ASVs were then clustered in three stages for each taxonomic rank from phylum to species, with the species-level clusters forming the final OTUs. In the first step, cluster cores were formed by the ASVs which had been assigned to taxa at that rank by Protax-fungi. These cluster cores were used as a reference for a closed-reference clustering stage, in which unassigned ASVs were matched to the closest cluster core using the optimized sequence similarity threshold for that rank and the nearest enclosing supertaxon. To this aim, we applied the “-usearch_global command” in VSEARCH version 2.22.128. We used the same alignment penalties for closed-reference clustering as for the threshold optimization clustering above to ensure that distance calculations were comparable. Iterations were performed until no new matches were found, generating approximately single-linkage clusters without merging cluster cores. Finally, in the third step, remaining unclustered ASVs at each rank were clustered using de novo single-linkage clustering using distances calculated by USEARCH as above, and again using the optimized sequence similarity threshold for the rank and nearest supertaxon. These de novo clusters, which we refer to as “pseudotaxa”, were assigned placeholder taxonomic names of the form “pseudo{rank}_{number}” (e.g., “pseudogenus_0216” for a cluster at genus rank). At each taxonomic rank after phylum, the three clustering stages were performed within the clusters generated at higher taxonomic ranks. Thus, two ASVs that were assigned to, for instance, different phyla by Protax-fungi, could not be clustered together into the same pseudoclass, even when their sequence similarity was greater than the class-level threshold determined for one or both phyla. Because the current version of Protax-fungi is trained only to identify fungi and not all eukaryotes, the non-fungal sequences were generally unidentified at the phylum level and were grouped into a large number of pseudophyla. We used the kingdom-level results from matching to the UNITE Sanger references (see above) to classify ASVs as “known fungi”, “known non-fungi”, or “unknown kingdom”, and removed pseudotaxa containing more known non-fungal ASVs than known fungal ASVs. At the phylum level, pseudotaxa containing only ASVs of unknown kingdoms were also removed. The final result of this process was a 27,954 species-level OTUs × 2,768 samples read abundance matrix, along with taxonomic annotations at each rank from phylum to species, including pseudotaxon placeholders. The bioinformatics pipeline was implemented using the Targets package version 1.336 in R version 4.2.2. Data records The database has been deposited to Zenodo37 and the sequence data are available at ENA European Nucleotide Archive38. The database is organized in five datasets in a csv format (columns separated by commas): (1) metadata providing the location, date, and time for each sample, along with sequencing depth and other essential information (Table2); (2) species-level OTU tables per sample describing the number of sequences assigned to each species (Table3); (3) taxonomic classification of each species-level OTU (Table4); (4) closest matching sequences and their taxonomy for ASVs in putatively fungal pseudophyla, which are included in (2) and (3)(Table 5); and (5) closest matching sequences and their taxonomy for ASVs in putatively non-fungal pseudophyla, which are not included in the other datasets(Table 6). The first four datasets can be linked to each other using the unique sample codes and the unique identifiers for species-level OTUs. technical Validation Field tests and negative controls. The median DNA amount measured by RT-PCR in the seven 24-hour test samples was 14 fg of DNA. The median DNA content measured in 1-hour samples was 8 fg, and the median for 10-minute samples, as well as for field blanks handled without gloves, were less than 3 fg. The median DNA quantity measured in the field blanks handled with gloves and the extraction blanks were approximately 0.7 fg, and the DNA quantity in the PCR blank was approximately 0.1 fg (Fig.2A). As these values were standardized using genomic DNA extracted from yeast, they cannot be directly translated to other fungi due to varying genome size and ITS copy number. Nonetheless, we note that 24-hour field samples had almost 5 times more ITS copies than blank samples handled without gloves, and twenty times more than blank samples handled with gloves. In the actual study, all samples were handled with gloves. Of the 99 negative controls, 89% of samples (i.e., 88 samples) did not yield any reads of fungal origin at the end of the bioinformatic analysis. For all sequencing runs, at least one negative control sample contained 0 fungal reads, indicating that the reagents were uncontaminated. The 9 negative control samples that did produce fungal reads yielded fewer fungal reads than the study samples (Fig.2B), and, in most cases, these reads belonged to only one or two OTUs. OTUs found in negative control samples were all relatively common in the 6 Scientific Data | (2024) 11:561 | https://doi.org/10.1038/s41597-024-03410-0 www.nature.com/scientificdata www.nature.com/scientificdata/ study. They were no more common in the sequencing runs which contained the negative controls than in other sequencing runs. This suggests that the most likely source of these reads was infrequent cross-contamination from study samples to negative controls. Among the negative controls, sample CCDB-35071NEGPCR2 yielded the highest read count: 2,668 fungal reads. All 18 OTUs detected in this sample were also found in sample COR_41A with abundances 7–60 times as high as in the negative control. Samples CCDB-35071NEGPCR2 and COR_41A were processed in the same sequencing run, indicating that the sample COR_41A was likely the source of cross-contamination. Sufficiency of sequencing depth. The mean sequencing depth among the samples was 86,845, and the median sequencing depth was 79,396. We recommend conducting analyses with samples yielding at least 10,000 sequencing reads, which corresponds to discarding 50 samples and thus 1.8% of the samples (Fig.3A). If rarefying all samples to 10,000 sequence reads, a minor loss of species-level OTU richness is observed for the most diverse samples (Fig.3B). Nonetheless, even the most diverse samples were likely sequenced to an adequate depth, as illustrated by the well-saturating rarefaction curves (Fig.3C). Field name Description sample.id Unique identifier of the sample seqrun The run in which the sample was sequenced site The site of sampling date The year, month, and day of sampling yday The Julian day of sampling, ranging from 1 to 365 duration The duration during which the sample was acquired, in hr water With levels “yes” if the sample contained water and “no.or.NA” if there was no water or the information was missing insect With levels “yes” if the sample contained insect(s) and “no.or.NA” if there were no insect(s) or the information was missing unst.tweezers With levels “yes” if the sample was processed with tweezers sterilized accidentally just by water and “no” if the tweezers were adequately sterilized spike_dilution The dilution level of the spike, either 0.01 or 0.001 numnonspikes The number of sequences assigned to non-spikes numspikes The number of sequences assigned to spikes dna_amount The inferred total amount of fungal DNA in the sample (log10 transformed) lat Latitude of the site (decimal degrees) lon Longitude of the site (decimal degrees) temp.mean Mean annual temperature of the site (°C) Table 2. The fields of the metadata table (metadata.csv). The rows of the metadata correspond to the samples. Field name Description sample.id Unique identifier of the sample Remaining fields Unique OTU identifiers Table 3. The fields of the samples x species-level OTU tables (otu.table.csv). The rows of the OTU tables correspond to the samples. Field name Description OTU Unique OTU identifier nsample The number of samples in which the taxon was found nread The total number of reads assigned to the taxon kingdom Inferred kingdom (always Fungi) phylum Inferred phylum class Inferred class order Inferred order family Inferred family genus Inferred genus species Inferred species sequence The sequence of the taxon Table 4. The fields of the taxonomy tables (taxonomy.csv). Levels of taxonomy that could not be reliably assigned to known taxa are indicated by names that include “pseudo”, numbered to allow identifying species that belong to the same unknown genus/family/order/class/phylum. The rows of the taxonomy tables correspond to the species-level OTUs. 7 Scientific Data | (2024) 11:561 | https://doi.org/10.1038/s41597-024-03410-0 www.nature.com/scientificdata www.nature.com/scientificdata/ Validation of automated taxonomic classifications by manual expert evaluation. Molecular taxonomic identification of fungi from environmental samples is challenging for several reasons21. First, the diversity of fungi is enormous, and most species are still unknown to science. Second, reference sequences are available only for a subset of the scientifically described species. Third, the systematics of fungi remains partially or even largely unresolved and undergoes continuous revisions. Fourth, the reference sequences in standard databases contain errors, and a substantial proportion of the reference sequences are mislabelled. Fifth, unlike the COI region used for molecular identification of animals, the ITS region does not allow for alignment at deep phylogenetic scales (much above the genus level), making sequence comparison more challenging. PROTAX-fungi explicitly accounts for all these sources of uncertainty while performing probabilistic taxonomic classification, and its validity has been tested by cross-validation experiments21. Given the taxonomic breadth of the data and the unexplored nature of airborne fungal diversity, we evaluated the validity of the PROTAX classifications by comparing them to taxonomic classifications carried out by independent experts. To do so, we first clustered the sequences with 97% similarity threshold and selected the most common sequence in each cluster as its representative. We then selected a total of 500 clusters (and their corresponding representatives) as follows: (i) 200 sequences that PROTAX could not reliably (with at least 90% probability) classify to any known phylum, in which case they are unlikely to belong to the fungal kingdom; (ii) 50 sequences that PROTAX reliably classified to a known phylum but an unknown class; (iii) 50 sequences that were reliably classified to a known class but an unknown order; (iv) 50 sequences reliably classified to a known order but an unknown family; (v) 50 sequences reliably classified to a known family but an unknown genus; (vi) 50 sequences reliably classified to a known genus but an unknown species; and (vii) 50 sequences reliably classified to a known species. Within each category, we selected clusters that achieved the highest prevalence Field name Description ASV Unique ASV identifier OTU Unique identifier for the OTU which the ASV belongs to pseudophylum Unique identifier for the phylum-level cluster the ASV belongs to pseudospecies Unique identifier for the species-level cluster the ASV belongs to sh_id Unique identifier for the Unite species hypothesis (SH) of the best match to the ASV dist Sequence dissimilarity between the ASV and the best match. 0.0 = all bases identical, 1.0 = all bases different. kingdom Kingdom of the best matching sequence, as given in Unite phylum Phylum of the best matching sequence, as given in Unite class Class of the best matching sequence, as given in Unite order Order of the best matching sequence, as given in Unite family Family of the best matching sequence, as given in Unite genus Genus of the best matching sequence, as given in Unite species Species of the best matching sequence, as given in Unite Table 5. The fields of the fungal pseudophylum table (pseudophyla_fungi.csv). All amplicon sequence variants (ASVs) that could not be assigned to a fungal phylum, but which belong to a pseudophylum classified as Fungi, are included. These sequences are also represented as OTUs in the main OTU table and taxonomy. For each ASV, the closest matching species hypothesis (SH) in Unite is given, along with the classification of that sequence in Unite. Field name Description ASV Unique ASV identifier pseudophylum Unique identifier for the phylum-level cluster the ASV belongs to pseudospecies Unique identifier for the species-level cluster the ASV belongs to sh_id Unique identifier for the Unite species hypothesis (SH) of the best match to the ASV dist Sequence dissimilarity between the ASV and the best match. 0.0 = all bases identical, 1.0 = all bases different. kingdom Kingdom of the best matching sequence, as given in Unite phylum Phylum of the best matching sequence, as given in Unite class Class of the best matching sequence, as given in Unite order Order of the best matching sequence, as given in Unite family Family of the best matching sequence, as given in Unite genus Genus of the best matching sequence, as given in Unite species Species of the best matching sequence, as given in Unite Table 6. The fields of the nonfungal pseudophylum table (pseudophyla_nonfungi.csv). All amplicon sequence variants (ASVs) that could not be assigned to a fungal phylum, but which belong to a pseudophylum classified as non-Fungi, are included. These sequences are excluded from the main OTU table and taxonomy and so do not have OTU identifiers. For each ASV, the closest matching species hypothesis (SH) in Unite is given, along with the classification of that sequence in Unite. 8 Scientific Data | (2024) 11:561 | https://doi.org/10.1038/s41597-024-03410-0 www.nature.com/scientificdata www.nature.com/scientificdata/ (i.e., that occurred in the highest proportions of the samples) in the GSSP data. Two authors with fungal taxonomic expertise (Otto Miettinen and Anton Savchenko) then manually performed the taxonomic classification of these 500 sequences, up to the taxonomic resolution that they considered possible to reliably achieve. The expert assessment was based on the first 100 BLAST hits between the query sequence and reference sequences in publicly available gene databases, thus incorporating a larger body of information than just a few top hits. In their assessment, the experts accounted for the quality issues in the reference sequences, such as divergent tail regions in poorly trimmed Sanger sequences, or chimeric sequences. Furthermore, naming of the sequences varies wildly, and experts used their judgement on which sequences to trust as the reference, and to what degree. There might be equally good hits under several names, in which case the experts judged which one was most likely correct. The best hit might refer to a name that is a collective, not allowing species-level identification with certainty. An important criterion in judging the reliability of reference sequences was related to the perceived trustworthiness of the sequence authors based on their taxonomic expertise (i.e., their standing in the field). As there is no published, up-to-date taxonomy for all fungal taxa, in many cases the experts had access to more up-to-date information (e.g., unpublished sources) about the classification, and then used this information when deciding on the correct naming at all taxonomic ranks. The taxonomic experts knew the criteria used to select the sequences, whereas the order in which the sequences were provided was randomized, so that the experts did not have a priori information about the Fig. 3 Results illustrating the sufficiency of sequencing depth, i.e., the total number of sequencing reads (including fungal and spike reads) obtained for each sample. Panel A shows the distribution of sequencing depth among the samples, with the dashed vertical line corresponding to the value of 10,000 sequence reads, which we recommend using as a threshold for including a sample for analyses. Panel B shows the decrease in the number of species-level OTUs if rarefying all samples to 10,000 sequence reads. Panel C shows rarefaction curves for all samples that included at least 10,000 sequence reads. Fig. 2 Results from field tests and negative controls. Panel A shows DNA concentration in the field test samples based on either 24-hr sampling, 1-hr sampling, or 10-min sampling, as PCR blanks, extraction blanks, and field blanks handled with and without gloves. Panel B shows the distributions of the number of fungal reads per sample based on either field samples (green bars), field blanks (blue bars), or lab blanks (red bars). Note the logarithmic scale in the x-axis. 15 Scientific Data | (2024) 11:561 | https://doi.org/10.1038/s41597-024-03410-0 www.nature.com/scientificdata www.nature.com/scientificdata/ participated in sample preparation and commented on the manuscript. S. Creer acquired data and commented on the manuscript. L.P. D’Acqui acquired data and commented on the manuscript. N. de Vere acquired data and commented on the manuscript. M. Desprez-Loustau acquired data and commented on the manuscript. M.A. Dongmo acquired data and commented on the manuscript. I.B. Dyrholm Jacobsen acquired data and commented on the manuscript. B.L. Fisher acquired data and commented on the manuscript. M. Flores de Jesus acquired data and commented on the manuscript. G.S. Gilbert acquired data and commented on the manuscript. G.W. Griffith acquired data and commented on the manuscript. A.A. Gritsuk acquired data and commented on the manuscript. A. Gross acquired data and commented on the manuscript. H. Grudd acquired data and commented on the manuscript. P. Halme contributed to the GBIF comparison and commented on the manuscript. R. Hanna acquired data and commented on the manuscript. J. Hansen acquired data and commented on the manuscript. L. Hansen acquired data and commented on the manuscript. A.D. Hegbe acquired data and commented on the manuscript. S. Hill acquired data and commented on the manuscript. I.D. Hogg acquired data and commented on the manuscript. J. Hultman contributed to the development of the bioinformatics pipeline and commented on the manuscript. K.D. Hyde acquired data and commented on the manuscript. N.A. Hynson acquired data and commented on the manuscript. N. Ivanova contributed to the planning and implementation of DNA extraction and sequencing and commented on the manuscript. P. Karisto acquired data and commented on the manuscript. D. Kerdraon participated in project coordination, participated in sample preparation and commented on the manuscript. A. Knorre acquired data and commented on the manuscript. I. Krisai-Greilhuber acquired data and commented on the manuscript. J. Kurhinen facilitated data acquisition and commented on the manuscript. M. Kuzmina contributed to the planning and implementation of DNA extraction and sequencing and commented on the manuscript. N. Lecomte acquired data and commented on the manuscript. E. Lecomte acquired data and commented on the manuscript. V. Loaiza acquired data and commented on the manuscript. E. Lundin acquired data and commented on the manuscript. A. Meire acquired data and commented on the manuscript. A. Mešić acquired data and commented on the manuscript. O. Miettinen performed manual classifications of sequences for technical validation and commented on the manuscript. N. Monkhause contributed to the planning and implementation of DNA extraction and sequencing and commented on the manuscript. P. Mortimer acquired data and commented on the manuscript. J. Müller acquired data and commented on the manuscript. R.H. Nilsson facilitated data acquisition and commented on the manuscript. P.C. Nonti acquired data and commented on the manuscript. J. Nordén acquired data and commented on the manuscript. B. Nordén acquired data and commented on the manuscript. C. Paz acquired data and commented on the manuscript. P. Pellikka acquired data and commented on the manuscript. D. Pereira acquired data and commented on the manuscript. G. Petch acquired data and commented on the manuscript. J. Pitkänen participated in project coordination, participated in sample preparation and commented on the manuscript. F. Popa acquired data and commented on the manuscript. C. Potter acquired data and commented on the manuscript. J. Purhonen contributed to the GBIF comparison and commented on the manuscript. S. Pätsi acquired data and commented on the manuscript. A. Rafiq acquired data and commented on the manuscript. D. Raharinjanahary acquired data and commented on the manuscript. N. Rakos acquired data and commented on the manuscript. A.R. Rathnayaka acquired data and commented on the manuscript. K. Raundrup acquired data and commented on the manuscript. Y.A. Rebriev acquired data and commented on the manuscript. J. Rikkinen acquired data and commented on the manuscript. H.M. Rogers participated in project coordination, participated in sample preparation and commented on the manuscript. A. Rogovsky acquired data and commented on the manuscript. Y. Rozhkov acquired data and commented on the manuscript. K. Runnel acquired data and commented on the manuscript. A. Saarto acquired data and commented on the manuscript. A. Savchenko performed manual classifications of sequences for technical validation and commented on the manuscript. M. Schlegel acquired data and commented on the manuscript. N. Schmidt acquired data and commented on the manuscript. S. Seibold acquired data and commented on the manuscript. C. Skjøth acquired data and commented on the manuscript. E. Stengel acquired data and commented on the manuscript. S.V. Sutyrina acquired data and commented on the manuscript. I. Syvänperä acquired data and commented on the manuscript. L. Tedersoo acquired data and commented on the manuscript. J. Timm acquired data and commented on the manuscript. L. Tipton acquired data and commented on the manuscript. H. Toju acquired data and commented on the manuscript. M. Uscka-Perzanowska participated in sample preparation and commented on the manuscript. M. van der Bank acquired data and commented on the manuscript. F.H. van der Bank acquired data and commented on the manuscript. B. Vandenbrink acquired data and commented on the manuscript. S. Ventura acquired data and commented on the manuscript. S.R. Vignisson acquired data and commented on the manuscript. X. Wang acquired data and commented on the manuscript. W. Weisser acquired data and commented on the manuscript. S.N. Wijesinghe acquired data and commented on the manuscript. S.J. Wright acquired data and commented on the manuscript. C. Yang acquired data and commented on the manuscript. N.S. Yorou acquired data and commented on the manuscript. A. Young acquired data and commented on the manuscript. D.W. Yu acquired data and commented on the manuscript. E. V. Zakharov contributed to the planning and implementation of DNA extraction and sequencing and commented on the manuscript. P.D.N. Hebert contributed to the planning and implementation of DNA extraction and sequencing and commented on the manuscript. T. Roslin conceived the study and contributed to the first draft of the manuscript. competing interests The authors declare no competing interests. additional information Correspondence and requests for materials should be addressed to O.O. 16 Scientific Data | (2024) 11:561 | https://doi.org/10.1038/s41597-024-03410-0 www.nature.com/scientificdata www.nature.com/scientificdata/ Reprints and permissions information is available at www.nature.com/reprints. Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/. © The Author(s) 2024 Otso Ovaskainen 1,2,3 ✉ , Nerea abrego 1,4, Brendan Furneaux 1, Bess Hardwick4, Panu Somervuo2, isabella Palorinne4, Nigel r. Andrew 5,6, Ulyana V. Babiy7, tan Bao8, Gisela Bazzano9, Svetlana N. Bondarchuk 10, Timothy c. Bonebrake11, Georgina L. Brennan12, Syndonia Bret-Harte13, claus Bässler14,15,16, Luciano cagnolo17, Erin K. cameron18, Elodie chapurlat19, Simon creer 20, Luigi P. D’acqui 21,22, Natasha de Vere 23, Marie-Laure Desprez-Loustau24,25, Michel A. K. Dongmo 11,26, ida B. Dyrholm Jacobsen 27, Brian L. Fisher 28,29, Miguel Flores de Jesus30, Gregory S. Gilbert 31, Gareth W. Griffith32, anna a. Gritsuk10, andrin Gross33, Håkan Grudd34, Panu Halme1, rachid Hanna35, Jannik Hansen36, Lars Holst Hansen36, apollon D. M. t. Hegbe 37, Sarah Hill5, ian D. Hogg38,39,40, Jenni Hultman 41,42, Kevin D. Hyde43, Nicole a. Hynson 44, Natalia ivanova45,46, Petteri Karisto 47,48, Deirdre Kerdraon19, Anastasia Knorre49,50, irmgard Krisai-Greilhuber51, Juri Kurhinen2, Masha Kuzmina45, Nicolas Lecomte52, Erin Lecomte 52, Viviana Loaiza53, Erik Lundin34, alexander Meire34, Armin Mešić 54, Otto Miettinen 55, Norman Monkhause45, Peter Mortimer56, Jörg Müller57,58, r. Henrik Nilsson 59, Puani Yannick c. Nonti37, Jenni Nordén60, Björn Nordén 60, claudia Paz61,62, Petri Pellikka63,64,65, Danilo Pereira47,66, Geoff Petch67, Juha-Matti Pitkänen42, Flavius Popa68, caitlin Potter32, Jenna Purhonen 1,69, Sanna Pätsi70, Abdullah rafiq20, Dimby raharinjanahary29, Niklas rakos 34, Achala r. rathnayaka 43,71, Katrine raundrup27, Yury A. rebriev 72, Jouko rikkinen2,55, Hanna M. K. rogers19, Andrey rogovsky49, Yuri rozhkov73, Kadri runnel 74,75, annika Saarto70, anton Savchenko75, Markus Schlegel33, Niels Martin Schmidt 36,76, Sebastian Seibold77,78, carsten Skjøth 67,79, Elisa Stengel57, Svetlana V. Sutyrina10, ilkka Syvänperä80, Leho tedersoo 74,81, Jebidiah Timm13, Laura tipton 82, Hirokazu toju 83,84, Maria Uscka-Perzanowska19, Michelle van der Bank85, F. Herman van der Bank85, Bryan Vandenbrink 38, Stefano Ventura 21,22, Solvi r. Vignisson86, Xiaoyang Wang87, Wolfgang W. Weisser 78, Subodini N. Wijesinghe 43,71, S. Joseph Wright 88, chunyan Yang87, Nourou S. Yorou 37, amanda Young13, Douglas W. Yu 87,89,90, Evgeny V. Zakharov45, Paul D. N. Hebert38,45 & Tomas roslin 2,19 1Department of Biological and Environmental Science, University of Jyväskylä, P.O. Box 35, FI-40014, Jyväskylä, finland. 2Organismal and evolutionary Biology Research Programme, faculty of Biological and environmental Sciences, University of Helsinki, P. O. Box 65, 00014, Helsinki, Finland. 3Department of Biology, centre for Biodiversity Dynamics, Norwegian University of Science and Technology, Trondheim, N-7491, Norway. 4Department of Agricultural Sciences, University of Helsinki, P.O. Box 27, FI-00014, Helsinki, Finland. 5natural History Museum, Zoology, University of New England, Armidale, NSW, 2351, Australia. 6faculty of Science and engineering, Southern cross University, Northern Rivers, NSW, 2480, Australia. 7Wrangel island State nature Reserve, Pevek, Russia. 8Department of Biological Sciences, MacEwan University, 10, 700 – 104 Avenue, Edmonton, AB, T5J 2P2, Canada. 9Universidad nacional de còrdoba, facultad de ciencias exactas físicas y naturales, centro de Zoología Aplicada, córdoba, Argentina. 10Sikhote-Alin State Nature Biosphere Reserve named after K. G. Abramov, 44 Partizanskaya Str., Terney, Primorsky krai, 692150, Russia. 11School of Biological Sciences, The University of Hong Kong, Hong Kong SAR, China. 12CSIC, Institute of Marine Sciences, Passeig Marítim de la Barceloneta, 37-49ES08003, Barcelona, Spain. 13institute of Arctic Biology, University of Alaska, Fairbanks, AK, USA. 14Goethe-University Frankfurt, Faculty of Biological Sciences, Institute for Ecology, Evolution and Diversity, Conservation Biology, D- 60438, Frankfurt am Main, Germany. 15Bavarian Forest National Park, Freyunger Str. 2, D-94481, Grafenau, Germany. 16ecology of fungi, Bayreuth center of Ecology and Environmental Research (BayCEER), University of Bayreuth, Universitätsstraße 30, 95440, Bayreuth, Germany. 17consejo de investigaciones científicas y técnicas (cOnicet), instituto Multidisciplinario de Biología Vegetal, córdoba, Argentina. 18Department of Environmental Science, Saint Mary’s University, 923 Robie St., Halifax, NS, B3H 3C3, Canada. 19Department of ecology, Swedish University of Agricultural Sciences (SLU), Uppsala, Sweden. 17 Scientific Data | (2024) 11:561 | https://doi.org/10.1038/s41597-024-03410-0 www.nature.com/scientificdata www.nature.com/scientificdata/ 20Molecular ecology and evolution at Bangor (MeeB), School of environmental and natural Sciences, Bangor University, Environment Centre Wales, Deiniol Road, Bangor, Gwynedd, Wales, LL57 2UW, UK. 21Research institute on Terrestrial Ecosystems - IRET, National Research Council - CNR, Via Madonna del Piano n° 10, 50019, Sesto Fiorentino, Firenze, Italy. 22national Biodiversity future center, Palermo, italy. 23natural History Museum of Denmark, University of Copenhagen, Gothersgade 130, 1123, København K, Denmark. 24INRAE, BIOGECO, F-33610, Cestas, France. 25University of Bordeaux, BIOGECO, F-33615, Bordeaux, France. 26international institute of tropical Agriculture (iitA), P.O. Box 2008 (Messa), Yaounde, cameroon. 27Greenland Institute of Natural Resources, Kivioq 2, P.O. Box 570, 3900, Nuuk, Greenland. 28Entomology, 55 Music Concourse Drive, California Academy of Sciences, San Francisco, CA, 94118, USA. 29Madagascar Biodiversity Center, Parc Botanique et Zoologique de Tsimbazaza, Antananarivo, 101, Madagascar. 30Legado das Águas, Reserva Votorantin, TPR 188 Km 22, Tapiraí, SP, 18180-000, Brazil. 31environmental Studies Department, University of California, Santa Cruz, 1156 High St., Santa Cruz, CA, 95065, USA. 32Department of Life Sciences, Aberystwyth University, Aberystwyth, Ceredigion, WALES SY23 3DD, UK. 33Research Unit Biodiversity and Conservation Biology, SwissFungi, Swiss Federal Research Institute WSL, Zürcherstrasse 111, CH-8903, Birmensdorf, Switzerland. 34Swedish Polar Research Secretariat, Abisko Scientific Research Station, Vetenskapens väg 38, SE-981 07, Abisko, Sweden. 35center for tropical Research, congo Basin institute, University of california, Los Angeles (UcLA), Los Angeles, CA, 90095, USA. 36Department of Ecoscience, Aarhus University, Dk-4000, Roskilde, Denmark. 37Research Unit in Tropical Mycology and Plant-Soil Fungi Interactions, Faculty of Agronomy, University of Parakou, BP 123, Parakou, Republic of Benin. 38Canadian High Arctic Research Station, Polar Knowledge Canada, PO Box 2150, 1 Uvajuq Road, Cambridge Bay, Nunavut, X0B 0C0, Canada. 39Department of integrative Biology, college of Biological Science, University of Guelph, 50 Stone Road East, Guelph, Ontario, N1G 2W1, Canada. 40School of Science, University of Waikato, Private Bag 3105, Hamilton, 3240, New Zealand. 41Department of Microbiology, University of Helsinki, Viikinkaari 9, FI-00014, Helsinki, Finland. 42Natural Resources Institute Finland, Latokartanonkaari 9, 00790, Helsinki, finland. 43Center of Excellence in Fungal Research, Mae Fah Luang University, Chiang Rai, 57100, Thailand. 44Pacific Biosciences Research center, University of Hawaii at Manoa, Honolulu, Hi, USA. 45Centre for Biodiversity Genomics, University of Guelph, Guelph, ON, N1G 2W1, Canada. 46Nature Metrics North America Ltd., 590 Hanlon Creek Boulevard, Unit 11, Guelph, ON, N1C 0A1, Canada. 47Plant Pathology Group, Institute of Integrative Biology, ETH Zurich, Zurich, Switzerland. 48Plant Health, natural Resources institute finland (Luke), Jokioinen, finland. 49Science Department, National Park Krasnoyarsk Stolby, 26a Kariernaya str., 660006, Krasnoyarsk, Russia. 50institute of Ecology and Geography, Siberian Federal University, 79 Svobodny pr., 660041, Krasnoyarsk, Russia. 51Department of Botany and Biodiversity Research, University of Vienna, Rennweg 14, 1030, Wien, Austria. 52Centre d’études nordiques and Canada Research Chair in Polar and Boreal Ecology, Department of Biology, Pavillon Rémi-Rossignol, 18, Antonine-Maillet, Université de Moncton, Moncton, NB, E1A 3E9, Canada. 53Department of evolutionary Biology and Environmental Sciences, University of Zürich, Zürich, Switzerland. 54Laboratory for Biological Diversity, Rudjer Boskovic Institute, Bijenicka cesta 54, HR-10000, Zagreb, Croatia. 55finnish Museum of natural History, University of Helsinki, P.O. Box 7, 00014, Helsinki, Finland. 56Centre for Mountain Futures, Kunming Institute of Botany, Chinese Academy of Sciences, Kunming, China. 57field Station fabrikschleichach, Department of Animal ecology and tropical Biology (Zoology III), Julius Maximilians University Würzburg, Rauhenebrach, Germany. 58Bavarian forest national Park, Grafenau, Germany. 59Department of Biological and Environmental Sciences, Gothenburg Global Biodiversity Centre, University of Gothenburg, Box 461, 405 30, Göteborg, Sweden. 60norwegian institute for nature Research (NINA), Sognsveien 68, N-0855, Oslo, Norway. 61Department of Biodiversity, institute of Biosciences, São Paulo State University, Av 24A 1515, Rio Claro, SP, 13506-900, Brazil. 62Department of entomology and Acarology, Laboratory of Pathology and Microbial Control, University of São Paulo, CEP 13418-900, Piracicaba, SP, Brazil. 63Department of Geosciences and Geography, Faculty of Science, University of Helsinki, P.O. Box 64, 00014, Helsinki, Finland. 64State Key Laboratory for Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan, 430079, China. 65Wangari Maathai institute for environmental and Peace Studies, University of nairobi, P.O. Box 29053, 00625, Kangemi, Kenya. 66Max Planck Institute for Evolutionary Biology, August-Thienemann-Str. 2, 24306, Plön, Germany. 67School of Science and the Environment, University of Worcester, Henwick Grove, Worcester, WR2 6AJ, UK. 68Department of Ecosystem Monitoring, Research & Conservation, Black Forest National Park, Kniebisstraße 67, 77740, Bad Peterstal-Griesbach, Germany. 69School of Resource Wisdom, University of Jyväskylä, P.O. Box 35, FIN- 40014, Jyväskylä, Finland. 70The Biodiversity Unit of the University of Turku, Henrikinkatu 2, 20500, Turku, Finland. 71School of Science, Mae Fah Luang University, Chiang Rai, 57100, Thailand. 72Southern Scientific center of the Russian Academy of Sciences, 41 Chekhov ave., Rostov-on-Don, 344006, Russia. 73State nature Reserve Olekminsky, Olekminsk, Russian federation, Russia. 74Mycology and Microbiology center, University of tartu, tartu, estonia. 75Institute of Ecology and Earth Sciences, University of Tartu, Liivi 2, 50409, Tartu, Estonia. 76Arctic Research center, Aarhus University, Dk-4000, Roskilde, Denmark. 77tUD Dresden University of technology, forest Zoology, Pienner Str. 7, 01737, Tharandt, Germany. 78Technical University of Munich, Terrestrial Ecology Research Group, Department of Life Science Systems, School of Life Sciences, Hans-Carl-von-Carlowitz-Platz 2, 85354, Freising, Germany. 79Department of Environmental Science, Aarhus University, Frederiksborgvej 399, DK-4000, Roskilde, Denmark. 80the Biodiversity Unit of the University of Turku, Kevontie 470, 99980, Utsjoki, Finland. 81College of Science, King Saud University, Riyadh, Saudi Arabia. 82School of natural Science and Mathematics, chaminade University of Honolulu, Honolulu, Hi, USA. 83Laboratory of Ecosystems and Coevolution, Graduate School of Biostudies, Kyoto University, Kyoto, 606-8501, Japan. 84Center for Living Systems Information Science (CeLiSIS), Graduate School of Biostudies, Kyoto University, Kyoto, 606-8501, Japan. 85African centre for DnA Barcoding (AcDB), University of Johannesburg, PO BOX 524, Auckland Park, 2006, South Africa. 86Sudurnes Science and Learning Center, Garðvegi 1, 245, Sandgerði, iceland. 87State Key Laboratory of Genetic Resources and Evolution, Kunming Institute of Zoology, Chinese Academy of Sciences, Kunming, China. 88Smithsonian Tropical Research Institute, Apartado, 0843–03092, Balboa, Panama. 89School of Biological Sciences, University of East Anglia, Norwich, Norfolk, NR4 7TJ, UK. 90Yunnan Key Laboratory of Biodiversity and Ecological Security of Gaoligong Mountain, Kunming Institute of Zoology, Chinese Academy of Sciences, Kunming, China. ✉e-mail: [email protected]