Assessing Phylogeographic Patterns and Genetic Diversity in Culex quinquefasciatus (Diptera: Culicidae) via mtDNA Sequences from Public Databases
Abstract
García-Escobar, Gian Carlo, González, Juan José Trujillo, Aguirre-Obando, Oscar Alexander (2024): Assessing Phylogeographic Patterns and Genetic Diversity in Culex quinquefasciatus (Diptera: Culicidae) via mtDNA Sequences from Public Databases. Zoological Studies 63 (58): 1-16, DOI: 10.6620/ZS.2024.63-58, URL: http://dx.doi.org/10.5281/zenodo.12831063
Full text
© 2024 Academia Sinica, Taiwan Open Access Assessing Phylogeographic Patterns and Genetic Diversity in Culex quinquefasciatus (Diptera: Culicidae) via mtDNA Sequences from Public Databases Gian Carlo García-Escobar1,2 , Juan José Trujillo González1,3 , and Oscar Alexander AguirreObando1,3,* 1Escuela de investigación en Biomatemática, Universidad del Quindío. Carrera 15, Calle 12 Norte, Armenia, Quindío, Colombia. E-mail: [email protected] (García-Escobar) 2Programa de Licenciatura en Biología y Educación Ambiental, Facultad de Educación, Universidad del Quindío. Carrera 15, Calle 12 Norte, Armenia, Quindío, Colombia 3Programa de Biología, Facultad de Ciencias Básicas y Tecnologías, Universidad del Quindío. Carrera 15, Calle 12 Norte, Armenia, Quindío, Colombia. *Correspondence: E-mail: [email protected] (Aguirre-Obando) E-mail: [email protected] (González) Received 28 May 2024 / Accepted 27 November 2024 / Published 31 December 2024 Communicated by John Wang To identify the worldwide genetic structure, gene flow, and diversity of Culex quinquefasciatus, we conducted phylogeographic and population genetics analyses utilizing publicly available mtDNA sequences. Therefore, the aim of this study was to investigate the genetic structure and diversity of natural populations of C. quinquefasciatus worldwide, using available genetic data reflecting its natural distribution. Our study focused on the cytochrome c oxidase subunit I (COI) gene, mirroring the species’ distribution pattern. We examined COI gene sequences from C. quinquefasciatus populations across Asia (n = 1,698), America (n = 334), Africa (n = 30), Oceania (n = 21), and Europe (n = 1), identifying 69 haplotypes. Genetic links were observed between Asian populations and those from other continents. Global genetic diversity was 0.531, varying from 0.095 in Oceania to 0.648 in South America. Neutrality tests indicated demographic expansions at the continental level in the Americas, North America, and Asia, as well as in some countries within these regions. In contrast, at both global and continental levels (South America, Oceania, and Africa), and in most countries within these continents, neutral populations were observed. AMOVA revealed genetic structuring among and within countries, with no genetic isolation observed (R2 = 0.03144; p > 0.05). Despite lower genetic diversity, Asian populations facilitated gene flow with other continents, suggesting a possible native origin of the species in Asia. The dispersal of this mosquito to new regions, coupled with its ability to transmit various arboviruses, underscores its significance as a potential public health threat. Key words: Epidemiology, Genetic diversity, Public health, Southern house mosquito, Vector capacity Citation: García-Escobar GC, González JJT, Aguirre-Obando OA. 2024. Assessing phylogeographic patterns and genetic diversity in Culex quinquefasciatus (Diptera: Culicidae) via mtDNA sequences from public databases. Zool Stud 63:58. doi:10.6620/ZS.2024.63-58. BACKGROUND The mosquito C. quinquefasciatus (Say 1823), commonly known as the southern house mosquito, is a widely distributed dipteran inhabiting urban, periurban, and rural environments. Its larval stages thrive in stagnant water containing organic debris, often found in artificial containers, subterranean systems, septic tanks, clogged drains, and abandoned wells (Calhoun et al. 2007). Primarily active at night, particularly the females, Zoological Studies 63:58 (2024) doi:10.6620/ZS.2024.63-58 1
© 2024 Academia Sinica, Taiwan adults of this species feed on the blood of both humans and animals. During daylight hours, they seek refuge in shadowed corners, shelters, sewers, and vegetation (Hickner et al. 2011). The taxonomic classification of C. quinquefasciatus has been a topic of debate. Some studies classify it as C. pipiens forma pipiens or C. fatigans, while others include it within a species complex alongside C. pipiens f. pipiens, C. pipiens f. molestus, C. pallens, C. australicus, and C. globocoxitus (Harbach 2012). In this study, C. quinquefasciatus is considered as part of the C. pipiens complex. Hybridization has been observed between populations of C. quinquefasciatus and C. pipiens f. pipiens and f. molestus in Southeast Asia, North America, Argentina, and Madagascar. Although mating between the three taxa is possible under controlled conditions, hybrids often have lower egg fertility and viability, although notable differences between populations persist in South Africa (Aardema et al. 2020). Morphologically, females and interspecies hybrids are nearly identical, necessitating molecular analysis or detailed examination of behavioral and physiological traits such as male genitalia for differentiation (Smith and Fonseca 2004; Harbach 2012). Among the Culex genus, C. quinquefasciatus emerges as the most anthropophilic and endophagic species (Gouge et al. 2019). It serves as a vector for filarial nematodes such as Dirofilaria immitis and Wuchereria bancrofti (Thanchomnang et al. 2013), as well as protozoa like Plasmodium relictum (Ferreira et al. 2022). Moreover, it is implicated in the transmission of a myriad of diseases, including Ross River virus (RRV) (Harley et al. 2001), Venezuelan equine encephalitis (VEE), Eastern equine encephalitis (EEE), Japanese encephalitis (JE), St. Louis encephalitis (SLE), West Nile virus (WNV), Myxoma virus (MV), avian reticuloendotheliosis virus (REV) (Gouge et al. 2019), and Rift Valley fever virus (RVFV) (Sang et al. 2010). Given its pivotal role, comprehending the patterns of genetic structuring and flow within C. quinquefasciatus populations is crucial, as it influences pathogen transmission, vector capacity, and competence (Van Den Eynde et al. 2022). Molecular markers play a pivotal role in unraveling the biology and population dynamics of disease vectors. Mitochondrial DNA (mtDNA) stands out for its small size, rapid evolutionary rate, and exclusive maternal inheritance with minimal genetic recombination. In the phylogeographic context, molecular markers have been extensively employed at the mtDNA level for C. quinquefasciatus in regions like the United States and Australia (Behura et al. 2011). Notably, genes such as dehydrogenase subunit 4 (ND4), cytochrome b (cytb), and cytochrome c oxidase subunit I (COI) have been utilized. ND4 data is available from the United States, Thailand, and South Africa (Rasgon et al. 2006; Chaulk et al. 2016), while cytb data spans regions including Benin, China, the Philippines, the United States, Réunion, South Africa, and Sri Lanka (Ishtiaq et al. 2008; Atyame et al. 2011). Similarly, COI information is accessible from various countries across Asia, Africa, Oceania, the Americas and Europe (Hasan et al. 2009; Shaikevich and Zakharov 2010; Huang et al. 2011; Quintero and Navarro 2012; Pfeiler et al. 2013; Sharma et al. 2013; Low et al. 2014; Ashfaq et al. 2014; Wilke et al. 2014; Daravath et al. 2015; Gunay et al. 2015; Murugan et al. 2016; Shaikevich et al. 2016; Dumas et al. 2016; Talaga et al. 2017; Koosha et al. 2017; Anoopkumar et al. 2019; Cane et al. 2020; Lorenz et al. 2021; Maekawa et al. 2021; Aremu et al. 2022; Panda and Barik 2022; Thankachan et al. 2023). However, a comprehensive analysis of all available genetic data for C. quinquefasciatus is lacking. Thus, this study aimed to investigate the global genetic structure, migration patterns, and diversity of natural populations of C. quinquefasciatus, utilizing existing genetic datasets that adhere to the species’ natural distribution pattern. MATERIALS AND METHODS Sequence retrieval, download, and analysis We conducted an exhaustive search in the NCBI and BoldSystem databases to acquire data on nuclear and mitochondrial genomes, as well as 13 mitochondrial genes. The search query employed the species name (in parentheses) combined with the boolean operator “AND” and the phrases “complete nuclear genome” or “complete mitochondrial genome,” enclosed in quotation marks, to retrieve genomic information. For the retrieval of mitochondrial genes, we used the species name in parentheses, followed by the “AND” operator, the specific gene name enclosed in quotation marks, and the “OR” operator followed by the gene’s abbreviations, also enclosed in quotation marks (Table S1) (Waldbieser et al. 2003). From the results obtained, genetic information was selected for further analysis based on a distribution pattern similar to that of C. quinquefasciatus as provided by GBIF (https://www.gbif.org/). Only genetic sequences demonstrating a similarity percentage above 98% following BLAST analysis (https://blast. ncbi.nlm.nih.gov/Blast.cgi) were considered (Donkor et al. 2014). Given that the species under study is part of a species complex, it was necessary to ensure that the genetic information used was exclusive to C. page 2 of 16Zoological Studies 63:58 (2024)
© 2024 Academia Sinica, Taiwan quinquefasciatus. To achieve this, two phylogenetic clustering analyses were conducted: one employing the maximum likelihood (ML) method (Fig. 4) and the other utilizing Bayesian inference (BI) (Holder and Lewis 2003). Both trees included species from pipiens complex, C. pipiens f. pipiens (GenBank access number: HQ724616), C. pipiens f. molestus (GenBank access number: MN389460), C. pallens (GenBank access number: KT851543), C. australicus (GenBank access number: NC054314), and C. globocoxitus (GenBank access number: KU495003). C. tarsalis (GenBank access number: JX259917), and Aedes aegypti (GenBank access number: MF194022) were used as an outgroup. Exclusive information within the internal group, comprising the genetic data under analysis, was considered valid and was the information used in this study. Based on the previous results, sequences of C. quinquefasciatus were classified according to their geographical location, continent, country, type of coverage, and altitude (Table S2). Sequences without geographical information were excluded from the analyses. The classification of coverage types (urban, peri-urban, and rural) and altitude (in meters) was determined utilizing the geographic coordinates sourced from occurrence data retrieved from the GBIF platform, as well as sequences from GenBank and BoldSystem. This process involved overlaying coverage and altitude layers obtained through DIVA-GIS (https://www.divagis.org/) (Guilherme et al. 2024). Subsequently, a map was generated using QGIS (https://www.qgis.org/ es/site/) to visualize the distribution of geographical information derived from genetic and occurrence data (Fig. 1). The alignment of sequences and construction of a haplotype network We conducted multiple sequence alignment using the MEGA software (Katoh et al. 2018). Haplotypes (H) were identified based on their frequency, with H1 assigned to the most common haplotype and subsequent numbering for others. To identify potential nuclear mitochondrial DNA sequences (NUMTs) among the haplotypes, a stop codon search was performed in the alignment (Leite 2012) (Table S3). Subsequently, to elucidate the genetic relationships between worldwide populations of C. quinquefasciatus, we constructed a haplotype network using NUMT-free haplotypes (Fig. 2). For this purpose, we employed the Population Analysis with Reticulate Trees (PopART) with a parsimony approach (Clement et al. 2000). Additionally, based on the interpretation of the haplotype network, which contains the haplotypes, their connections through lines with their respective mutational steps, and the countries they belong to, we plotted the genetic flow on a world map using arrows to represent the connections suggested by the network, starting from the most frequent haplotype to all other haplotypes. These arrows connect countries that share the same haplotypes and, in turn, those with similar or different haplotypes and/or countries. Population genetic analysis Based on the aligned, trimmed, and countryorganized sequences, we conducted a population genetic analysis that included calculation of haplotype diversity (Hd), nucleotide diversity (π), and neutrality tests (Tajima’s D), at both continental and country levels within each continent (Table 1). In these countrylevel analyses, we considered each country and all the associated sequences as a distinct population due to geographical differences and potential barriers to dispersal that could influence the genetic structure of the mosquito under study. Additionally, an analysis of molecular variance (AMOVA) was performed at a global and continental level to assess the distribution of genetic variation among and within defined populations (countries). These analyses were carried out using the R environment (CoreTeam 2017) with the following packages: ape (Didier 2024), adegenet (Jombart 2023), pegas (Paradis et al. 2016), poppr (Kamvar et al. 2014), usethis (Maintainer 2024), devtools (Hadley et al. 2022), and mmod (Winter et al. 2013). Population genetic structuring was evaluated using the fixation index FST. Pairwise FST values between all populations were calculated as proposed by Weir and Cockerham (1984) implemented in the hierfstat package (Goudet et al. 2022) in R. Gene flow (Nm) was estimated from the FST values using the formula Nm = 1-FST/2.FST, assuming an island model. The FST values obtained were represented using boxplots for each country, showing the distribution of FST values in comparison with all other countries. These plots were generated using the ggplot2 package in R. In these analyses, singleton populations (countries with only one sample) were excluded due to their impact on the precise estimation of FST values and to avoid biases in interpreting genetic differentiation (Hale et al. 2012). Furthermore, we examined the potential relationship between genetic (FST) and geographic (Km) distances using the Mantel test (Fig. 5). The Mantel test was performed using the mantel.randtest function from the ade4 package (Dray and Dufour 2007), with 999 permutations to assess statistical significance. Geographic distances between populations were calculated using the central geographic coordinates of each country. Google Earth page 3 of 16Zoological Studies 63:58 (2024)
© 2024 Academia Sinica, Taiwan Fig. 1. World map for Culex quinquefasciatus presenting information for: A, Distribution of information retrieved from genetic and geographical databases. B, Genotypic distribution of information retrieved from genetic databases associated with three coverages (urban, peri-urban, and rural) and boxplot representing the results of the Kruskal-Wallis analysis for the variables coverage and altitude. C, Phenotypic distribution of information retrieved from the geographical database associated with three coverages (urban, peri-urban, and rural) and boxplot representing the results of the Kruskal-Wallis analysis for the variables coverage and altitude. page 4 of 16Zoological Studies 63:58 (2024)
© 2024 Academia Sinica, Taiwan Fig. 2. Global haplotype network (A) and map representing gene flow for the COI gene (B) in C. quinquefasciatus. In the haplotype network, the size of the circles is proportional to their frequency, and each colored circle belongs to a respectively numbered haplotype; black circles and numbers in parentheses represent mutational steps. In map B, color indicates the location of the sequences used, and arrows represent the genetic relationships between natural populations of Culex quinquefasciatus. page 5 of 16Zoological Studies 63:58 (2024)
© 2024 Academia Sinica, Taiwan was used to calculate the geographic (Km) distances used in the Mantel test. RESULTS Figure 1 shows the world map for C. quinquefasciatus and the distribution of data from genetic and geographical databases (Fig. 1A), including genetic data categorized by urban, peri-urban, and rural coverage (Fig. 1B), with a boxplot of Kruskal-Wallis analysis for coverage and altitude variables, and geographical data categorized similarly with the corresponding boxplot (Fig. 1C). Figure 1A shows the distribution of genetic (GenBank + BoldSystems) and phenotypic information (GBIF) of C. quinquefasciatus, including the COI gene data (n = 1,911) used in this study. When statistically comparing whether countries with phenotypic data Table 1. Global genetic diversity presenting data on haplotypic diversity (Hd), nucleotide diversity (π), and results of the neutrality test (Tajima’s D), calculated for natural populations of the southern house mosquito, by continents and countries within each continent Countries by continent Number of sequences Genetic diversity Number of haplotypes Hd πTajima’s D P-value Global 1,911 69 0.531 0.003 -2.570 0.10 America 167 27 0.383 0.012 -2.239 0.025* South America 42 10 0.648 0.002 -1.322 0.186 Colombia 36 6 0.533 0.002 -1.476 0.140 Argentina 3 2 0.667 0.003 NaN NaN Guayana Francesa 2 1 0 0 NaN NaN Brazil 1 1 NaN NaN NaN NaN North America 125 17 0.268 0.015 -2.080 0.038* United Stated 56 11 0.384 0.022 -1.776 0.076 Puerto Rico 53 1 0 0 NaN NaN Mexico 16 5 0.600 0.034 0.164 0.870 Oceania 21 4 0.095 0.0002 -1.164 0.245 Australia 11 2 0.182 0.0005 -1.129 0.259 New Caledonia 9 1 0 0 NaN NaN New Zealand 1 1 NaN NaN NaN NaN Europe 1 1 NaN NaN NaN NaN Hungary 1 1 NaN NaN NaN NaN Africa 25 5 0.157 0.001 -1.214 0.225 Malawi 14 2 0.143 0.001 -1.155 0.248 Uganda 9 1 0 0 NaN NaN Guinea 1 1 NaN NaN NaN NaN D. R. of the Congo 1 1 NaN NaN NaN NaN Asia 1,697 68 0.480 0.002 -2.521 0.012* Pakistan 1,044 25 0.073 0.0001 -2.345 0.019* China 477 20 0.341 0.002 -2.599 0.009* Malasia 72 2 0.407 0.001 1.196 0.232 India 34 6 0.410 0.001 -1.576 0.115 Thailand 27 2 0.142 0.001 -0.954 0.340 Turkey 15 3 0.362 0.002 -1.451 0.147 Japan 14 2 0.527 0.001 1.434 0.152 Iran 4 2 0.500 0.003 -0.710 0.478 Myanmar 4 3 0.833 0.003 0.592 0.554 Singapore 3 1 0 0 NaN NaN United Arab Emirates 2 1 0 0 NaN NaN Bangladesh 1 1 0 NaN NaN NaN NaN: Not enough sequences to obtain the results, or no segregating sites in the sample. *: indicate statistical significance. page 6 of 16Zoological Studies 63:58 (2024)
© 2024 Academia Sinica, Taiwan also have genetic data, and analyzing the difference between both distributions, the Fischer test indicates a significant difference (p ≤ 0.05), with an odds ratio of 0 (95% CI: 0.0000000–0.1859958), suggesting that the presence of genetic data is not associated with the presence of the species. However, it is observed that several countries where the species has been recorded also have genetic information, suggesting that despite the statistical difference, our results generally reflect a global distribution analysis of C. quinquefasciatus. Additionally, it illustrates the distribution of this information based on genetic data (Fig. 1B) and occurrences (Fig. 1C) categorized by type of coverage. Upon separate analysis of the genetic data and occurrence data using a Kruskal-Wallis test associated with three coverage categories (urban, peri-urban, and rural) concerning altitude, a statistically significant difference is observed in the genetic data (p < 0.0001; Fig. 1B), whereas no such difference is evident in the occurrence data (p > 0.9999; Fig. 1C). This suggests that the available information for genetic data obtained worldwide is not uniform regarding coverage and altitude categories, whereas it is for occurrence data. The species exhibits a global distribution based on genetic and occurrence data, across altitudinal ranges from 1 to 3,540 m. The genetic data suggest that the widest altitudinal distribution is observed in the Americas, with records ranging from 3 to 3,540 m (North America: 4 to 1,545 m, and South America: 3 to 3,540 m). In contrast, in Africa, altitudes range from 22 to 1,209 m, while for Europe, only one record exists at 128 m. Regarding occurrence data, worldwide, the species is recorded altitudinally within a range of 1 to 2,912 m (North America: 1 to 2,638 m, and South America: 1 to 2,912 m), with the lowest altitude in Europe, with records from 168 to 155 m. Distribution by coverage reveals that the majority of records, both in genetic data and occurrence data, are found in urban areas worldwide. Table S2 further elaborates on this information by continent and country, delineating coverage type and altitude for COI gene genetic data and occurrence data sourced from GBIF. Between September 2022 and January 2023, a total of 2,055 genetic sequences were acquired, primarily from the GenBank database (93%) and BoldSystem (7%). These sequences spanned various continents, with distribution as follows: Asia (88.3%), America (8.6%, comprising 6.2% in North America and 2.4% in South America), Africa (1.9%), Oceania (1.1%), and Europe (0.1%). The average sequence length was 698 bp, ranging from 300 to 1,000 bp. Subsequently, following alignment and trimming processes, 1,911 sequences with a consistent length of 388 bp each were retained, covering all aforementioned continents. Asia (88.829%) emerged as the most represented continent, with sequences originating from countries such as Pakistan (54.6%), China (25%), Malaysia (3.8%), India (1.8%), Thailand (1.4%), Turkey (0.8%), Japan (0.7%), Iran (0.21%), Myanmar (0.16%), Singapore (0.2%), United Arab Emirates (0.10%), and Bangladesh (0.05%). North America (6.5%) was represented by sequences from the United States (2.9%), Puerto Rico (2.8%), and Mexico (0.8%). In South America (2.16%), sequences were recorded from Colombia (1.9%), Argentina (0.16%), French Guiana (0.10%), and Brazil (0.05%). Africa (1.33%) was represented by sequences from Malawi (0.73%), Uganda (0.5%), Guinea (0.05%), and the Democratic Republic of the Congo (0.05%). Oceania (1.15%) was represented by sequences from Australia (0.6%), New Caledonia (0.5%), and New Zealand (0.05%). Lastly, Europe (0.05%) was represented by a single sequence from Hungary (0.05%). Table S3 for the COI gene comprises the sequence alignment for each of the 69 identified haplotypes (H), along with their amino acid translations. H1 is the most prevalent, accounting for 61.54% of the sequences, followed by H2 (30.09%), H3 (1.65%), H4 (0.94%), H5 (0.68%), H6 to H7 (each at 0.31%), H8 (0.21%), H9 to H13 (each at 0.16%), H14 to H23 (each at 0.10%), and H24 to H69 (each at 0.05%). While H1 was the most abundant and found in 15 countries, H2 exhibited the broadest geographical distribution, being present in 16 countries. Figure 2 illustrates the geographic origin of the collected genetic material, depicting the distribution of these haplotypes across various continents and countries. Both the haplotype network (Fig. 2A) and the genetic flow map (Fig. 2B) demonstrate genetic connections between Asia and other regions, including America, Africa, Europe, and Oceania, highlighting these connections across multiple populations of C. quinquefasciatus. In Asia, genetic connections are observed among populations from Pakistan and various countries such as China, Myanmar, Malaysia, Thailand, Turkey, India, and Iran. Additionally, China exhibits connections with Pakistan, Japan, and Iran. Asia and America also display genetic relationships through Pakistan, which links with the United States, Mexico, Argentina, Brazil, and French Guiana. Furthermore, China is linked to America, establishing connections with the United States, Puerto Rico, Mexico, and Colombia, while Japan has genetic connections with Argentina and Colombia. Asia is connected to Africa via Pakistan, which links with Uganda and the Democratic Republic of the Congo, and Japan shows genetic connections with Guinea. Moreover, Asia is associated with Europe through Pakistan and China, which connect page 7 of 16Zoological Studies 63:58 (2024)
© 2024 Academia Sinica, Taiwan with Hungary. Lastly, Asia also connects with Oceania, particularly through China, which has genetic ties with Australia, New Zealand, and New Caledonia. These intercontinental genetic connections reflect the diversity and dispersal history of C. quinquefasciatus on a global scale. Additionally, in a global context, we conducted two additional analyses based on the results obtained for the haplotypes. In the first analysis, the objective was to evaluate whether there are significant differences in the altitude at which different haplotypes are found. A non-parametric Kruskal-Wallis test was applied, where the dependent variable was altitude and the independent variable was the haplotypes (mitochondrial lineages H1 to H69). This test was appropriate because altitude does not follow a normal distribution, and there are multiple groups. The results were visualized using boxplots, allowing us to observe the distribution of altitudes associated with each haplotype. In the second analysis, we aimed to determine whether the altitudinal distribution of haplotypes varies according to zones (urban, peri-urban, and rural) and if there is an interaction between the zone and different haplotypes with respect to altitude. A two-way Analysis of Variance (ANOVA) was performed, considering zone and haplotypes as factors and altitude as the dependent variable. Before conducting the ANOVA, we assessed the assumptions of normality and homoscedasticity. The results were presented through a heat map, showing variations in mean altitude among different haplotypes and zones, facilitating the identification of patterns and potential interactions. In both analyses, the altitude data were associated with the results presented in table S2, considering the distribution of haplotypes. All analyses were performed using R (CoreTeam 2017), utilizing the stats package for statistical tests, ggplot2 for data visualization, and reshape2 for data manipulation. Figure 3 presents the results of the KruskalWallis tests (Fig. 3A) and the two-way ANOVA (Fig. 3B). In both cases, the analyses were significant: Kruskal-Wallis (d.f. = 68, p ≤ 0.05) and ANOVA (Zone (Z): d.f. = 2, p ≤ 0.05; Haplotype (H): d.f. = 68, p ≤ 0.05; Zone x Haplotype Interaction (Z x H): d.f. = 14, p ≤ 0.05). Despite these results, the bias in the distribution of genetic information is highlighted again (Tables 1 and S2). Overall, figure 3A suggests that haplotypes are distributed differently across the various altitudes where they are reported, and many of them, due to the limited amount of data, do not show altitudinal variation. It is noteworthy that haplotypes 1 and 2, although the most frequent, do not exhibit the widest range of altitudinal distribution; in contrast, haplotype 6 does show a greater altitudinal range. On the other hand, figure 3B indicates that the distribution of haplotypes based on urban, rural, and peri-urban areas concerning altitude differs. In rural areas, a greater number of haplotypes is observed, which are also more widely distributed altitudinally. In contrast, peri-urban and urban areas have fewer haplotypes compared to rural areas, although the haplotypes present in these zones tend to be concentrated at lower altitudes. Interestingly, haplotype 6 is present in both rural and urban areas. The phylogenetic tree constructed using the maximum likelihood (ML) method indicates that all genetic information for the COI gene of C. quinquefasciatus is exclusive to this species (Fig. 4). However, branch supports were generally ≤ 50 in most cases, indicating limited reliability in the clusters within the tree for certain nodes. However, it is observed that the American populations are different from the Asian ones, and the latter may have been derived from the American populations. Conversely, the phylogenetic tree generated using the Bayesian inference (BI) method did not achieve stabilization of the Markov chains after 1,000 million interactions, and thus it’s results are not presented. Table 1 presents the results for haplotype diversity (Hd), nucleotide diversity (π), and the neutrality test for different continents and countries. The global Hd was 0.531, showing significant variation among continents, ranging from 0.095 in Oceania to 0.648 in South America. At the country level, Hd ranged from 0.0 (observed in places like French Guiana, Puerto Rico, New Caledonia, Uganda, Singapore, and the United Arab Emirates) to 0.833 in Myanmar. Regarding π, the global value was 0.003, with regional disparities ranging from 0.0 (recorded in French Guiana, Puerto Rico, New Caledonia, Uganda, Singapore, and the United Arab Emirates) to 0.034 in Mexico. Our results based on the number of haplotypes distributed by continents suggest the highest diversity in Asia, followed by America, Africa, and Oceania. Tajima’s D test results suggest that, at both global and continental levels (South America, Oceania, and Africa), and in most countries within these continents, negative values (the majority) and positive values (only a few) were observed, though none were statistically significant. In contrast, at the continental level, the Americas, North America, and Asia presented significant negative values. Additionally, in some countries in Asia (Pakistan and China), significant negative Tajima’s D values were also observed. These findings underscore diverse patterns of diversity and natural selection across different geographical regions for C. quinquefasciatus. The results of the analysis of molecular variance (AMOVA) indicated significant genetic structuring both at the level of countries within the same continent (d.f. = 22, ss = 642.23, p < 0.05) and within countries page 8 of 16Zoological Studies 63:58 (2024)
© 2024 Academia Sinica, Taiwan Fig. 3. Boxplot representing the altitudinal distribution of each observed haplotype (A) and heatmap visualizing the interaction between the factors zone and haplotype with respect to altitude (B). Each cell shows the average altitude of a haplotype in a specific zone, with a blue color scale indicating the value: lighter colors represent lower altitude values, while darker colors indicate higher altitude values. 0 1000 2000 3000 H1 H5 H10 H15 H20 H25 H30 H35 H40 H45 H50 H55 H60 H65 H69 Haplotypes Altitude (m) H1 H5 H10 H15 H20 H25 H30 H35 H40 H45 H50 H55 H60 H65 H69 Peri−urban RuralUrban Zone Average altitude 1000 2000 Haplotypes A B page 9 of 16Zoological Studies 63:58 (2024)
© 2024 Academia Sinica, Taiwan Ramírez-Soriano A, Ramos-Onsins SE, Rozas J, Calafell F, Navarro A. 2008. Statistical power analysis of neutrality tests under demographic expansions, contractions and bottlenecks with recombination. Genetics 179:555–567. doi:10.1534/ genetics.107.083006. Rasgon JL, Cornel AJ, Scott TW. 2006. Evolutionary history of a mosquito endosymbiont revealed through mitochondrial hitchhiking. Proc R Soc B Biol Sci 273:1603–1611. doi:10.1098/ rspb.2006.3493. Rose NH, Sylla M, Badolo A, Lutomiah J, Ayala D, Aribodor OB, Ibe N, Akorli J, Otoo S, Mutebi JP, Kriete A, Ewing EG, Sang R, Gloria-Soria A, Powell JR, Baker RE, White BJ, Crawford JE, McBride CS. 2020. Climate and Urbanization Drive Mosquito Preference for Humans. Curr Biol 30:3570–3579.e6. doi:10.1016/j.cub.2020.06.092. Sang R, Kioko E, Lutomiah J, Warigia M, Ochieng C, Guinn MO, Lee JS, Koka H, Godsey M, Hoel D, Hanafi H, Miller B, Schnabel D, Breiman RF, Richardson J. 2010. Rift Valley fever virus epidemic in Kenya, 2006/2007: The entomologic investigations. Am J Trop Med Hyg 83(2_Suppl):28–37. doi:10.4269/ajtmh.2010.09-0319. Shaikevich EV, Vinogradova EB, Bouattour A, de Almeida APG. 2016. Genetic diversity of Culex pipiens mosquitoes in distinct populations from Europe: Contribution of Cx. quinquefasciatus in Mediterranean populations. Parasites and Vectors 9:1–16. doi:10.1186/s13071-016-1333-8. Shaikevich E, Zakharov I. 2010. Polymorphism of mitochondrial COI and nuclear ribosomal ITS2 in the Culex pipiens complex and in Culex torrentium (Diptera: Culicidae). Acta Trop 4:161–174. doi:10.3897/compcytogen.v4i2.45. Sharma AK, Chandel K, Tyagi V, Mendki MJ, Tikar SN, Sukumaran D. 2013. Molecular phylogeney and evolutionary relationship among four mosquito (Diptera: Culicidae) species from India using PCR-RFLP. J Mosq Res 3:58–64. Smith JL, Fonseca DM. 2004. Rapid assays for identification of members of the Culex (Culex) pipiens complex, their hybrids, and other sibling species (Diptera: Culicidae). Am J Trop Med Hyg 70:339–345. doi:10.4269/ajtmh.2004.70.339. Tajima F. 1989. Statistical Method for Testing the Neutral Mutation Hypothesis by DNA Polymorphism. Pharmatherapeutica 3:607– 612. Talaga S, Leroy C, Guidez A, Dusfour I, Girod R, Dejean A, Murienne J. 2017. DNA reference libraries of French Guianese mosquitoes for barcoding and metabarcoding. PLoS ONE 12:e0176993. doi:10.1371/journal.pone.0176993. Tatem AJ, Hay SI, Rogers DJ. 2006. Global traffic and disease vector dispersal. Proc Natl Acad Sci USA 103:6242–6247. doi:10.1073/ pnas.0508391103. Thanchomnang T, Intapan PM, Tantrawatpan C, Lulitanond V, Chungpivat S, Taweethavonsawat P, Kaewkong W, Sanpool O, Janwan P, Choochote W, Maleewong W. 2013. Rapid detection and identification of Wuchereria bancrofti, Brugia malayi, B. pahangi, and Dirofilaria immitis in mosquito vectors and blood samples by high resolution melting real-time PCR. Korean J Parasitol 51:645–650. doi:10.3347/kjp.2013.51.6.645. Thankachan M, Surya P, Sebastian CD. 2023. Molecular Identification and phylogenetic analysis of mosquito vectors from Mananthavady Taluk, Wayanad, Kerala, India. J Vector Borne Dis 60:8. doi:10.4103/0972-9062.361166. Thiemann TC, Lemenager DA, Kluh S, Carrol BD, Lothrop HD, Reisen WK. 2012. Spatial variation in host feeding patterns of Culex tarsalis and the Culex pipiens complex (Diptera: Culicidae) in California. J Med Entomol 49:903–916. doi:10.1603/ME11272. Torres-Gutierrez C, Bergo ES, Emerson KJ, de Oliveira TMP, Greni S, Sallum MAM. 2016. Mitochondrial COI gene as a tool in the taxonomy of mosquitoes Culex subgenus Melanoconion. Acta Trop 164:137–149. doi:10.1016/j.actatropica.2016.09.007. Van Den Eynde C, Sohier C, Matthijs S, De Regge N. 2022. Japanese encephalitis virus interaction with Mosquitoes: a review of vector competence, vector capacity and mosquito immunity. Pathogens 11:317. doi:10.3390/pathogens11030317. Waldbieser GC, Bilodeau AL, Nonneman DJ. 2003. Complete sequence and characterization of the channel catfish mitochondrial genome. DNA Seq - J DNA Seq Mapp 14:265–277. doi:10. 1080/1042517031000149057. Waples RS. 2015. Testing for hardy-weinberg proportions: Have we lost the plot? J Hered 106:1–19. doi:10.1093/jhered/esu062. Weaver SC, Forrester NL, Liu J, Vasilakis N. 2021. Population bottlenecks and founder effects: implications for mosquitoborne arboviral emergence. Nat Rev Microbiol 19:184–195. doi:10.1038/s41579-020-00482-8. Weir B, Cockerham CC. 1984. Estimating F-Statistics for the Analysis of Population Structure. Evolution 38:1358–1370. doi:10.2307/2408641. Wilke A, Vidal P, Suesdek L, Marrelli M. 2014. Population genetics of neotropical Culex quinquefasciatus (Diptera: Culicidae). Parasit Vectors 7:468. doi:10.1186/preaccept-1270950864130552. Winter D, Green P, Kamvar Z, Gosselin T. 2013. Package “mmod” Modern Measures of Population Differentiation. R package version 1:1–7. Zang C, Wang X, Liu Y, Wang H, Sun Q, Cheng P, Liu H. 2024. Wolbachia and mosquitoes: Exploring transmission modes and coevolutionary dynamics in Shandong Province, China. PLoS Negl Trop Dis 18:e0011944. doi:10.1371/journal.pntd.0011944. Supplementary materials Table S1. List of mitochondrial gene names and their abbreviations. Adapted from Waldbieser et al. (2003), for Culex quinquefasciatus. (download) Table S2. Global compilation for Culex quinquefasciatus of altitude and coverage data derived from information retrieved from genetic and geographical databases. (download) Table S3. Global haplotypic diversity from the COI gene for natural populations of Culex quinquefasciatus. The translated amino acid sequence is represented in uppercase blue letters, above the third nucleotide of its corresponding codon. Invariable sites are indicated with dots, whereas alternative nucleotides are represented in uppercase black letters. (download) page 16 of 16Zoological Studies 63:58 (2024)