© The Author(s) 2025. Published by Oxford University Press on behalf of Society for Molecular Biology and Evolution. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited. 1 Title: Whole-genome duplication reshapes adaptation: autotetraploid Arabidopsis arenosa leverages its high 1 genetic variation to compensate for selection constraints. 2 3 Authors: Celestini Sonia1*, Lipánová Veronika2,3, Vlček Jakub1,4, Kolář Filip1* 4 5 Affiliations: 6 1. Department of Botany, Faculty of Science, Charles University, Prague, Czechia. 7 2. Institute of Integrative Biology, ETH Zurich, Zurich, Switzerland. 8 3. Czech Academy of Sciences, Institute of Botany, Průhonice, Czechia. 9 4. Institute of Parasitology, Biology Centre CAS, České Budějovice, Czechia. 10 11 *correspondence:
[email protected],
[email protected] 12 13 Abstract 14 Whole-genome duplication (WGD), a widespread macromutation across eukaryotes, is predicted to affect the 15 tempo and modes of evolutionary processes. By theory, the additional set(s) of chromosomes present in 16 polyploid organisms may reduce the efficiency of selection while, simultaneously, increasing heterozygosity 17 and buffering deleterious mutations. Despite the theoretical significance of WGD, empirical genomic evidence 18 from natural polyploid populations is scarce and direct comparisons of selection footprints between 19 autopolyploids and closely related diploids remains completely unexplored. We therefore combined locally 20 sampled soil data with resequenced genomes of 76 populations of diploid–autotetraploid Arabidopsis arenosa 21 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
2 and tested whether the genomic signatures of adaptation to distinct siliceous and calcareous soils differ 1 between the ploidies. Leveraging multiple independent transitions between these soil types in each ploidy, we 2 identified a set of genes associated with ion transport and homeostasis that were repeatedly selected for 3 across the species’ range. Notably, polyploid populations have consistently retained greater variation at 4 candidate loci compared to diploids, reflecting lower fixation rates. In tetraploids, positive selection 5 predominantly acts on such a large pool of standing genetic variation, rather than targeting de novo mutations. 6 Finally, selection in tetraploids targets genes that are more central within the protein–protein interaction 7 network, potentially impacting a greater number of downstream fitness-related traits. In conclusion, both 8 ploidies thrive across a broad gradient of substrate conditions, but WGD fundamentally alters the ploidies 9 adaptive strategies: tetraploids leverage their greater genetic variation and redundancy to compensate for the 10 predicted constraints on the efficacy of positive selection. 11 12 Introduction 13 Understanding how genetic variation and environmental pressures interact to shape species adaptation over 14 time is a central challenge for research in our rapidly changing world (Rausher and Delph 2015). Among the 15 many genomic processes influencing evolutionary dynamics, whole-genome duplication (WGD, i.e. 16 polyploidization) emerges as a substantial and recurrent macromutation with far-reaching, yet still under17 explored, implications for evolution and ecology (Fox et al. 2020). WGD is widespread in the Eukaryotic 18 kingdom and has significantly affected the evolutionary history of plants in particular. Indeed, all plants are 19 paleo-polyploids, meaning that they all have experienced at least one episode of polyploidization (Wendel 20 2015; Alix et al. 2017), and multiple ancient WGDs have been recognized in evolutionary history of land plants 21 (Garsmeur et al. 2014; Zhang et al. 2020). 22 Earlier ideas saw polyploidy as an evolutionary ‘dead-end’ because of the expected long-term 23 detrimental effects of increased genomic complexity on competition and adaptation (Comai 2005; Wendel 24 2015; Van de Peer et al. 2017). Nevertheless, several ancient polyploidization events seem to coincide with 25 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
3 major lineage diversifications (Van de Peer et al. 2017), adaptive radiations (Zhang et al. 2020), enhanced 1 invasiveness and colonization potential (te Beest et al. 2012; Baduel, Bray, et al. 2018; Baniaga et al. 2020), and 2 plant domestication success (Paterson 2005; Salman-Minkov et al. 2016). The abundance and observed 3 adaptability of polyploids in nature thus reinforces the more recent view of WGD as an advantageous 4 opportunity for evolutionary success (Baduel, Bray, et al. 2018). However, the consequences of WGD on natural 5 selection processes remain controversial (Soltis and Soltis 2000; Comai 2005; Parisod et al. 2010; Carretero - 6 Paulet and Van de Peer 2020). To disentangle this conundrum, the study of autopolyploidy, where genome 7 doubling happens within a single species, helps to isolate the effect of polyploidization from the confounding 8 factor of hybridization present in allopolyploids. Empirical population genomic studies on autopolyploid 9 species are, however, rare (Konečná et al. 2021; Bray et al. 2024; Hämälä et al. 2024; Zhang et al. 2024) and we 10 completely lack a systematic analysis that would compare positive selection footprints in diploid and closely 11 related autopolyploid genomes facing similar selective pressure. 12 Recent theoretical advances in quantitative and population genetics of autopolyploidy have 13 highlighted several interacting factors that shape the tempo and mode of evolution differently in polyploids 14 compared with their diploid relatives (Cuypers and Hogeweg 2014; Baduel, Bray, et al. 2018; Yao et al. 2019; 15 Monnahan and Brandvain 2020; Clo 2022; Ebadi et al. 2023). First, WGD generally increases genetic diversity 16 and polymorphisms within a population, yet this has contrasting consequences. On the one hand, the 17 additional set of chromosomes in autopolyploid organisms masks recessive alleles in a heterozygous state and 18 thereby diminishes the efficiency of purifying and positive selection, increasing the genetic load (Ronfort 1999; 19 Vlček et al. 2025) and slowing fixation (Monnahan and Brandvain 2020). On the other, the increased 20 heterozygosity and genetic redundancy in autopolyploids can buffer the effects of deleterious mutations and 21 thus potentially enhance genetic robustness (Parisod et al. 2010, Yao et al. 2019). Second, from a molecular 22 network perspective, such redundancy provided by duplicated gene networks might relax evolutionary 23 constraints (Yao et al. 2019; Ebadi et al. 2023; Parisod et al. 2024), allowing selection to act on genes/proteins 24 that occupy central positions in molecular networks. Such central genes influence multiple downstream 25 fitness-related traits (Wang et al. 2010; Frachon et al. 2017; Rennison and Peichel 2022) and are thus typically 26 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
4 more constrained in diploids (Olson-Manning et al. 2012; Masalia et al. 2017). Third, weaker genetic linkage 1 and increased effective recombination rates (e.g. Monnahan et al. 2019), which result from the higher number 2 of chromosome combinations at meiosis, theoretically reduce Hill–Robertson interference (e.g. the reduction 3 in the efficacy of selection caused by linkage between loci, Hill and Robertson 1966) and unlink unfavourable 4 allele combinations more efficiently. In simulated polyploid genomes, higher recombination rates and 5 extended fixation times due to masking create a unique genomic footprint, distinct from what is expected in 6 diploids (Monnahan and Brandvain 2020). Hard selective sweeps in polyploids cause a greater absolute 7 reduction in diversity (sweep magnitude) compared to diploids but remain more localized (narrower breadth) 8 and do not reach fixation. However, empirical evidence from natural populations supporting these 9 expectations is missing. 10 Finally, WGD affects evolutionary processes that are known to drive the sources of adaptive variation. 11 At the same census population size, tetraploid populations are expected to experience close to double the 12 occurrences of de novo mutations compared to diploids, due to a doubled number of mutation targets 13 (Selmecki et al. 2015). Whether individual polyploid populations differently leverage such novel genetic 14 variants or instead rely on the high level of standing genetic variation already present in the polyploid lineage 15 can markedly affect their adaptive success. Indeed, the dynamics, speed and outcome of adaptation are 16 affected by the source of variation – standing or de novo – targeted by natural selection (Barrett and Schluter 17 2008; Thompson et al. 2019), and this has been demonstrated by ample genomic evidence in diploids (Reid et 18 al. 2016; Alves et al. 2019; Ji et al. 2020; Bohutínská et al. 2021; Choi et al. 2021; Montejo-Kovacevich et al. 19 2022; Rubin et al. 2022). However, the role of WGD in modulating sources of adaptive variation in natural 20 populations have never been addressed. Konečná et al. (2021) found repeated instances of adaptation to toxic 21 substrates in autotetraploid populations of Arabidopsis arenosa originating predominantly from standing 22 variation. However, the lack of comparable environmental contrasts in diploid populations left the role of ploidy 23 unclear. Taken together, all these genetic processes have influenced the past evolutionary trajectories of 24 polyploids and will affect their evolvability and future survival in a fast-changing world (Wu et al. 2020). 25 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
5 However, the balance between the advantages and disadvantages of WGD in shaping genetic adaptation 1 remains a critical aspect that remains to be addressed empirically. 2 In this study, we address this knowledge gap by investigating how diploid and closely related 3 autotetraploid populations of Arabidopsis arenosa adapt to non-extreme, spatially heterogeneous soil factors 4 commonly experienced by plants, namely those determined by calcareous and siliceous substrates. 5 Arabidopsis arenosa represents a suitable species for this aim given it is an outcrossing species exhibiting very 6 high genetic diversity enabling the efficient detection of hard selective sweeps (Yant and Bomblies 2017). 7 Moreover, it has a well-characterized ploidy distribution, ecology and evolutionary history. Naturally 8 established autotetraploid populations of A. arenosa trace back to a single origin from the Western Carpathian 9 diploid lineage ~ 19–31 k generations ago, after which they have spread across Europe, largely overlapping with 10 the ecological niche of diploids, as the two ploidies followed a similar process of postglacial expansion (Arnold 11 et al. 2015; Monnahan et al. 2019; Padilla-García et al. 2023). Notably, polyploid genomes have retained 12 extensive genetic variation with no detectable effect of any conspicuous bottleneck, possibly reflecting a 13 gradual origin from a large and diverse diploid population (Arnold et al. 2015), and a lack of a transition to selfing 14 – also supported by S-allele diversity study (Vekemans et a. 2025). These findings document a comparable age 15 and demographic histories of populations involved in the colonization of different soils between the ploidies, 16 making them suitable for the purposes of our study. A pervasive effect of ploidy on genetic diversity and 17 footprints of selection has been documented throughout the genome of A. arenosa (Monnahan et al. 2019; 18 Vlček et al. 2025). However, the observed trends primarily reflect genome-wide effects of neutral processes 19 and purifying and linked selection, but the effect of ploidy on local signals of positive selection remains 20 unknown. 21 Siliceous and calcareous soils impose different sets of challenges for plant life, especially in terms of 22 pH and nutrient availability (Ormeño et al. 2008; Bontpart et al. 2024). They are both frequently inhabited by 23 A. arenosa throughout central and southeastern Europe (e.g. Morgan et al. 2020) and Guggisberg et al. (2018) 24 has detected related signatures of adaptation, in a pooled data of seven diploid populations, in genes involved 25 in ion transport, response to nitrate and cellular ion homoeostasis. Here we build on those findings and 26 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
6 expanded the dataset to individually sequenced diploid and autotetraploid genomes of plants from 76 1 populations covering the range-wide diversity of A. arenosa. We complemented each population sample with 2 in situ sampled soil data and applied three methods inferring positive selection to unravel the genetic basis of 3 substrate adaptation in each ploidy level. Then, by comparing the top candidate selection sweeps across 4 ploidies, we sought to answer the following questions: i) Do tetraploids rely on novel private (post-WGD) 5 variation or rather adapt through the same genes selected for also in diploids when facing the same selective 6 pressure? ii) Is there a larger proportion of fixed variants in diploids compared to tetraploids? iii) Do selection 7 footprints differ between ploidies in terms of sweep breadth and magnitude? iv) Do ploidies differ in their ability 8 to repeatedly leverage distinct sources of adaptive variation (standing variation, migration, de novo 9 mutations)? v) Does positive selection in tetraploids target genes that are more evolutionary constrained (i.e. 10 more central in the protein–protein interaction network) than are candidate genes inferred in diploids? 11 12 Results 13 Siliceous and calcareous soils were colonized repeatedly by each ploidy. 14 We first characterised the substrate niche of each population. We collected genomic and soil data for 76 15 populations (35 diploid, 2x, and 41 tetraploid, 4x; 479 individuals in total), covering the majority of the 16 distribution range of Arabidopsis arenosa and different soil types (Fig. 1a). Analyses of locally collected soil 17 parameters (such as scaled values of pH, bioavailable Ca mmol/kg, total content in mmol/kg of Mg, K and Na, 18 and cation exchange capacity – CEC) through PCA identified soil pH, calcium availability and CEC as the major 19 differentiating variables (PC1 explains 44% of the variation, Fig. 1c). As expected, samples of calcareous soils 20 (here defined by PC1 scores smaller than < 0) exhibit high pH, large concentrations of exchangeable calcium, 21 generally high CEC and lower concentrations of potassium, which are typical characteristics of calcareous 22 conditions. In turn, samples of siliceous soils (PC1 score > 0) exhibit the opposite trends. Note that other 23 potentially relevant environmental factors, like temperature and precipitation, do not generally associate with 24 soil conditions (supplementary fig. S1). In addition, neither soil environmental conditions, nor climatic 25 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
7 parameters differ between the ploidies, as both ploidies span the broad range of environmental conditions 1 (Fig. 1c, supplementary fig. S1), in line with previous reports of a lack of ploidy-related niche differentiation 2 (Morgan et al. 2020; Padilla-García et al. 2023). 3 In order to provide a robust evolutionary history background for detection of the selection candidates, 4 we investigated the genetic relationships between the populations and the possible repeated colonization of 5 different soil types, using a set of 1,347,998 putatively neutral fourfold degenerate SNPs called in the full 6 dataset of 479 A. arenosa individuals with an average sequencing depth of 25⨯. The principal component 7 analysis (PCA – supplementary fig. S2) and the tree-based approaches (TreeMix – supplementary fig. S3, Fig. 8 1d, e), as well as the clustering method (Entropy – supplementary fig. S4), corresponded with the previously 9 described internal sub-structuring of A. arenosa into distinct diploid and tetraploid lineages (Kolář et al. 2016; 10 Monnahan et al. 2019; Padilla-García et al. 2023). The overall population grouping by spatial proximity and not 11 by substrate indicates repeated colonization of contrasting soils by multiple lineages within each ploidy 12 cytotype (Fig. 1d, e for the paired dataset, supplementary fig. S2–4 for the full dataset), albeit with an unclear 13 ancestral state. Diploid and tetraploid populations show similar genome-wide Tajima’s D (average D2x = 0.157, 14 D4x = 0.139; Welch two-sample t-test t(51.376) = 0.426, P = 0.672), suggesting ploidy-specific demographic 15 history should have not a marked effect on detecting signals of selection (Fig. 1b). In line with theoretical 16 expectations (Baduel, Bray, et al. 2018), tetraploid populations show greater within-population nucleotide 17 diversity than diploids (average 2x= 0.046, 4x = 0.052; Wilcoxon rank sum test W = 166, P = 0.0005403) (Fig.1b). 18 Information on overall population structure and soil analysis of the full dataset were used to design a 19 replicated pairwise window-based genome-wide scan for directional selection (see further). We selected 20 seven siliceous/calcareous (high-low PC1) population pairs per ploidy, optimizing contrast in soil 21 characteristics, pairwise genome-wide differentiation, geographical distance and sample size. These selected 22 populations well represent the characteristics of our full dataset in terms of lineage diversity, geographical 23 distribution, genetic diversity and soil conditions (Fig. 1). Importantly, neutral pairwise genome-wide 24 differentiation between calcareous and siliceous populations (4dg-Rho) of the paired dataset is very close 25 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
8 between the ploidies (mean 2x: 0.20, SD: 0.05; mean 4x: 0.13, SD: 0.06, supplementary data S2), indicating a 1 comparable neutral divergence between siliceous and calcareous populations across ploidies. 2 3 Fig. 1 Genetic relationships, diversity, distribution and soil conditions of the Arabidopsis arenosa 4 populations under study. a) Geographical distribution in central and southeastern Europe of the 76 genome - 5 sequenced populations studied. Circles indicate diploid populations and triangles indicate tetraploid 6 populations. The subset of populations used for pairwise divergence contrasts (paired dataset) are annotated 7 with their relative pair number (diploid, 2x, 1–7; tetraploid, 4x, 8–14). b) Within-population diversity (pairwise 8 nucleotide diversity, ) and estimate of neutrality (Tajima’s D) calculated from putatively neutral fourfold 9 degenerated SNPs detected in 35 diploid and 41 tetraploid populations. Dots refer to the values for each 10 population used in the paired dataset. c) PCA of the elemental soil composition for each population sampled 11 (dots coloured by ploidy) (left) and variable correlation plot (right). Populations used in the paired dataset are 12 indicated by their corresponding number. Soil variables are coloured by their contribution, expressed in 13 percentage values, to the explained variation along the first two principal component axes. d, e) Relationships 14 between diploid (d) and tetraploid (e) populations of the paired dataset, depicted using an allele frequency 15 covariance graph calculated in Treemix. Each population is labelled by its corresponding pair number and 16 corresponding major genetic lineage, following previous range-wide studies (Monnahan et al. 2019; Padilla17 García et al. 2023). All branches are significantly supported (bootstrap > 96%), the sole exception being marked 18 with a black dot. Throughout the entire figure, the green–violet colour ramp reflects the position of each 19 population along the PC1 axis of elemental soil composition PCA – bar on top in figure c – ranging from 20 calcareous (violet) to siliceous (green). 21 22 23 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
9 Genomic regions associated with adaptation to siliceous and calcareous soil differ between ploidies but 1 share similar functions. 2 We identified candidate genomic regions for adaptation to calcareous and siliceous soils in each ploidy level. 3 To reduce the number of false positives, we used three complementary approaches and their intersection to 4 refine and validate our list of top candidate loci for each cytotype. In addition, we focused on differentiation 5 patterns recurrently present across the A. arenosa range and thus more likely to be result of selection as 6 opposed to local (population pair-specific) signatures of genetic drift. 7 First, we scanned the genome for the top 1% of FST outliers among lineages growing on calcareous 8 versus siliceous soil, calculated in 1-kbp windows including a minimum of 10 SNPs, in a subset of populations 9 representing distinct genetic lineages of each ploidy (paired dataset, Fig. 1). Based on information on overall 10 population structure and soil analyses, we selected seven ‘calcareous/siliceous’ population pairs (contrasts) 11 per ploidy, maximizing the divergence in soil properties while minimizing neutral genome-wide differentiation 12 and geographical distance within the pair (see the Methods section for details). This analysis resulted in an 13 average of 564 and 643 outlier SNP windows in diploid and tetraploid pairs, respectively. After annotating the 14 outlier windows to genic regions, we detected an average number of 406 annotated genes in diploids and 426 15 in tetraploids (supplementary data S3). We then identified a total of 233 genes in diploids and 219 genes in 16 tetraploids that were outliers in at least two population pairs of different lineage within ploidy (Fig. 2c, e). We 17 refer to these genes as parallel candidate genes. Notably, these loci do not cluster in regions characterized by 18 extreme recombination rates, indicating no bias associated with the recombination landscape (the average 19 recombination rates in regions found to be under selection do not differ from the average recombination rate 20 for genes not subject to selection, one-sided permutation test p-value = 0.9426). These genes are significantly 21 enriched (p < 0.05) for an array of biological processes related to substrate adaptation such as xenobiotic 22 transmembrane transport, cellular response to nitrate, potassium ion transmembrane transport, and lateral 23 root development (supplementary data S4). Second, repeated genetic differentiation between populations 24 growing on calcareous versus siliceous soil from the paired dataset was further quantified with PicMin (Booker 25 et al. 2023), a statistical method which uses order statistics to quantify the non-randomness of adaptation 26 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
16 SPINDLY (SPY) – AL3G23090 in quartet 8–9. In tetraploid populations, the increased sharing of adaptive variants 1 through migration (38.3%, n = 18 in 4x; 27.16%, n = 22 in 2x) seems to compensate the lack of selection acting 2 on de novo mutations, although the proportion of loci shared through migration is not significantly greater than 3 in diploids (X²(1, N = 128) = 1.24, p = 0.2659). 4 5 Figure 5. Different evolutionary sources and number of protein–protein interactions of candidate genes 6 associated with substrate adaptation in diploid and tetraploid A. arenosa populations. a) Different 7 representation of the three major evolutionary sources of repeated adaptation quantified by DMC in each 8 ploidy. b) Density distribution of the number of protein–protein interactions among the top candidate genes per 9 each ploidy, compared to the full dataset of A. thaliana genes. The vertical lines show the position of the mean 10 of each group. c) Protein–protein interaction networks of the top candidate genes demonstrating overall greater 11 connectivity (red/violet shading) of genes identified in tetraploids (right). 12 13 Tetraploid adaptive candidates are more central in the protein–protein interaction network. 14 Efficient and faster adaptation might be achieved in tetraploids if selection acts preferentially on genes that 15 exhibit a central position in the protein–protein interaction network and thus likely have a strong pleiotropic 16 effect (Promislow 2004). We thus estimated relevant proxies for pleiotropy using the STRING database v.12 17 (https://string-db.org) (Szklarczyk et al. 2023) of the protein–protein interaction network in Arabidopsis 18 thaliana. In total, we collected data for 25,715 A. thaliana proteins (nodes) annotated to 26,206 genes. The 19 number of edges per connected node ranged from 1 to 6,278, with a mean value of 487.9 edges per node 20 (Fig. 5b). We compared these values with the nodes’ connectivity measures for A. thaliana orthologs of our set 21 of 40/24 top candidate genes in diploid/tetraploid A. arenosa. In diploids, the number of edges per node ranged 22 from 34 to 2,947, with a mean value of 514.6; in tetraploids the range spanned from 107 to 4,587 edges per 23 node, with a mean of 841 (Fig. 5b, supplementary data S6). The average connectivity of diploid top candidate 24 genes does not significantly differ from the average number of edges in the complete A. thaliana dataset (one25 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
17 sided permutation test P = 0.3989), indicating that candidate adaptive genes are neither more peripheral nor 1 central than the average. By contrast, tetraploid top candidate genes show a significantly greater number of 2 edges per node, and thus connectivity, than the average for A. thaliana (one-sided permutation test P = 0.0049). 3 4 Discussion 5 Here, we investigated how diploid and closely related autotetraploid populations of Arabidopsis arenosa adapt 6 to spatially heterogeneous environments and examined the effect of ploidy on genomic footprints of selection. 7 As we used adaptation to calcareous and siliceous substrates as model system, we first discuss the genetic 8 basis of substrate adaptation and its possible functional relevance, highlighting genes involved in ion transport 9 and homeostasis. This provides a biological context for our results and reinforces the robustness of our 10 candidate gene lists, which form the foundation for interpreting subsequent patterns of selection and adaptive 11 variation. Then, we compared selection footprints in these candidate genes between naturally replicated 12 diploid and autotetraploid populations to uncover the role of ploidy in adaptation. We found differences in 13 multiple parameters characterising selective sweeps, availability of adaptive variation and gene interactions 14 and discuss their potential implications for polyploid evolution. 15 16 Repeated colonization of siliceous and calcareous substrates across the range of A. arenosa recruits 17 genes involved in ion transport and homeostasis. 18 Siliceous and calcareous soils are usually scattered throughout the landscape, creating a mosaic of 19 contrasting habitats differing in pH, chemistry and texture. These characteristics are known to affect plant 20 distribution and productivity worldwide (Michalet et al. 2002; Alvarez et al. 2009; Bothe 2015; Nemer et al. 21 2021; Smyčka et al. 2022). In many plant groups, differentiation into populations growing on calcareous or 22 siliceous soils coincided with speciation (e.g. Minuartia, Androsace, Campanula, Gentiana, Phyteuma, 23 Primula and Saxifraga) (Moore and Kadereit 2013; Smyčka et al. 2022; Lipánová et al. 2023). However, lineages 24 and cytotypes of A. arenosa have repeatedly colonized both types of environments without accumulating 25 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
18 significant genome-wide differentiation (Fig. 1). Such pattern rather suggests a scenario of repeated ecotypic 1 differentiation similar to that observed in other plant species with variable substrate preferences (Bastida et 2 al. 2014; Kolář et al. 2014) or in A. arenosa populations adapted to high elevations (Bohutínská et al. 2021) and 3 toxic serpentine soils (Arnold et al. 2016; Konečná et al. 2021). 4 Living on siliceous and calcareous soil requires specific mechanisms for dealing with many chemical 5 constraints, such as high levels of calcium and carbonate, deficiencies of iron, zinc and potassium in 6 calcareous soils, and with aluminium toxicity, nitrogen deficiency and decreased basic cation concentrations 7 in siliceous soils (Clark and Baligar 2000), which may act as selective factors. Indeed, we observed enrichment 8 for genes involved in ion transport and homeostasis (see below for a detailed discussion) among outliers 9 identified in our analysis. This result complements previous findings made in a subset of diploid populations 10 (Guggisberg et al. 2018), altogether suggesting an adaptive genetic response of A. arenosa to contrasting soil 11 conditions. Despite lacking direct experimental evidence of siliceous/calcareous adaptation in A. arenosa, 12 local adaptation to calcareous soil has been experimentally documented in the closely related species 13 A. lyrata (Guggisberg et al. 2018) and A. thaliana (Terés et al. 2019), indicating that calcareous adaptation is a 14 frequent phenomenon in the model genus Arabidopsis (Koch 2018). 15 As also described in Guggisberg et al. (2018), we identified the gene SULTR1;1 as a very strong adaptive 16 candidate in both ploidies (Fig. 2 and 4). This gene is expressed in root epidermis and root hairs in response to 17 sulfur starvation, and its up-regulation positively correlates with the content of sulfate in the roots of 18 Arabidopsis thaliana (Takahashi et al. 2000; Rouached et al. 2009). Notably, sulfur deficiency is common in 19 calcareous soils, as sulfur reacts with the abundant calcium present in the substrate, forming calcium sulfate 20 (CaSO4), which prevents sulfur uptake. Yet, sulfur is a pivotal macronutrient for plants, affecting not only their 21 growth and development, but also their response to oxidative stress, pests, diseases and heavy metals. 22 Indeed, sulfur is required for the biosynthesis of compounds such as the amino acids cysteine (Cys) and 23 methionine (Met), proteins, glutathione (GSH) and phytochelatins, vitamins, and secondary metabolites such 24 as glucosinolates (Shah et al. 2022). One soil-specific SULTR1;1 allele has been found to be shared among 72 25 Arabidopsis accessions, suggesting that selection has recurrently acted on an ancient polymorphism existing 26 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
19 in plants growing on soils of each type (Guggisberg et al. 2018). This is in accordance with our finding that 1 selection concerning the SULTR1;1 gene acts mainly on standing variation across lineages and ploidies of 2 A. arenosa (DMC results, Fig. 4d, supplementary data S8). Similarly to sulfur, potassium has an antagonistic 3 relationship with calcium, which inhibits its uptake by plants (York et al. 1953). Accordingly, we found 4 candidate genes related to potassium deficiency in A. thaliana. It is the GC5 gene, which encodes a protein 5 that functions as a structural component of the Golgi apparatus in diploids (Kang et al. 2004), while in 6 tetraploids selection acted on KAT3 gene, encoding potassium channel subunits, and on the KUP9 potassium 7 ion transmembrane transporter (Hedrich 2012; Šustr et al. 2024). Interestingly, in diploids we have also found 8 candidate genes involved in iron uptake, such as FRO2, FRO4 and ABCG37, which are well characterized in 9 Arabidopsis (Vert et al. 2003; Fourcroy et al. 2016; Anjali et al. 2021). Iron is indeed inaccessible to plants in 10 calcareous conditions, as ferric oxides become more insoluble with every unit increase in pH (Kim and 11 Guerinot 2007). 12 In addition, we detected a strong signal of selection acting on three out of seven genes of the NRT2 13 family of nitrate transporters in Arabidopsis (NRT2.1, NRT2.2 and NRT2.5). These three genes are mostly 14 expressed in the roots and, when induced by nitrogen starvation, mediate high-affinity root NO3− influx (Xu et 15 al. 2024). Although nitrogen soil content does not seem to be particularly linked to any single soil type, 16 microbial activity, responsible for ammonification, nitrification, denitrification and nitrogen fixation, is usually 17 lower at lower pH, negatively affecting how plants absorb nitrogen (Bothe et al. 2006). Many of the above listed 18 genes have been identified as candidates for adaptation to serpentine soil in A. arenosa and A. lyrata, siliceous 19 and calcareous conditions in A. lyrata and A. thaliana, metalliferous soil in A. halleri and salt stress in 20 E. salsugineum and A. pumila, suggesting that they might represent genetic hotspots for adaptation to soil 21 conditions (Turner et al. 2010; Arnold et al. 2016; Guggisberg et al. 2018; Sailer et al. 2018; Yang et al. 2018; 22 Preite et al. 2019; Terés et al. 2019; Konečná et al. 2021; Tran et al. 2021). 23 Finally, one important distinguishing feature of siliceous and calcareous soils is their water and heat 24 holding capacity. Typical calcareous soils are fissured, allowing rainfall to rapidly drain away, exposing plants 25 to water deficiency and higher temperatures. By contrast, siliceous soils have almost no fissures and therefore 26 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
20 retain moisture for longer, also maintaining a lower soil temperature. In agreement with this observation, four 1 of our detected top candidate genes are enriched for drought resistance (RCI3, NIP5;1, DRS1, EXL4) (Llorente 2 et al. 2002; Lee et al. 2010; Yu et al. 2015; Mancini et al. 2022). In summary, we have detected repeated signals 3 of selection acting on a suite of genes affecting ion transport and management as well as drought resistance, 4 all of which correspond with known challenges posed by the calcareous/siliceous soil gradient. 5 6 Targets and strength of positive selection differ between ploidies of A. arenosa. 7 Theory and simulations have raised clear expectations regarding how autopolyploids differently respond to 8 positive selection due to their lower variance in fitness, different manifestation of dominance and propensity 9 to retain genetic variation (Monnahan and Brandvain 2020). We addressed these expectations by comparing 10 candidate genes allele frequency distributions and the extent of gene sharing between diploid and 11 autotetraploid A. arenosa populations repeatedly adapting to similar environmental pressures. 12 First, we found that only six top candidate genes were shared between the cytotypes (14% of diploid 13 top candidates and 21% of tetraploid top candidates, Fig. 2d). The number is greater than what would be 14 expected randomly; however, it is surprisingly low considering the broadly shared pattern of polymorphism in 15 Arabidopsis (Novikova et al. 2018) and the rampant gene-flow between the ploidies (supplementary fig. S2 and 16 S4, Monnahan et al 2019). There are several possible reasons why the two ploidies rely mostly on different 17 adaptive genes in spite of a globally shared pool of variation in the species. The first reason is that selection 18 may act upon a different spectrum of allele dominance in each ploidy cytotype (Marad et al. 2018, Monnahan 19 and Brandvain 2020). Fixation of recessive variants is expected to be less likely in polyploids due to much rarer 20 frequencies of homozygotes (equilibrium q4) compared to diploids (q2). Selection in tetraploids might therefore 21 be more efficient in genes where the adaptive variant is dominant, while diploids may rely on both types of 22 variants (Monnahan and Brandvain 2020). Secondly, genes with different levels of centrality or genetic context 23 could be differently accessible in higher or lower ploidies, which we discuss in more depth in the following 24 section. 25 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
21 In addition, we have found a significant difference between the ploidies in the number of fixed 1 candidate variants subject to selection and the characteristics of their hard selective sweeps. First, we 2 detected an almost complete lack of fixed alleles in tetraploid populations (Fig. 3b), while selective sweeps are 3 greater in magnitude than those of diploids; this means that the genetic differentiation between populations 4 on siliceous and calcareous bedrock at the selected region is elevated higher above the genome-wide 5 background in tetraploids (Fig. 3e, f). In line with the reasoning explained above, we speculate that this might 6 be a result of positive selection acting mainly on dominant variants in tetraploids. Indeed, simulations suggest 7 that (partially) dominant mutations increase in frequency early on but then stabilize at intermediate 8 frequencies and take longer to approach fixation; this is especially true in polyploid populations, where the 9 frequency of the genotypes expressing the dominant phenotype is higher than in diploids (Monnahan and 10 Brandvain 2020). Moreover, in higher ploidies, time to fixation is expected to be longer also because the level 11 of diversity before the onset of selection is usually higher (Otto and Whitton 2000). In turn, our results show 12 that candidate alleles in diploids were more often fixed, which corresponds with the expected greater 13 likelihood of selection acting on recessive alleles in lower ploidies (Marad et al. 2018; Monnahan and Brandvain 14 2020). Incomplete fixation in polyploids might also be explained by an increased plasticity. However, this 15 seems less likely as there is not a consistent trend for increased phenotypic plasticity in autopolyploids (te 16 Beest et al. 2012) and similar levels of plasticity have also been documented for diploid and autotetraploid A. 17 arenosa during alpine adaptation (Wos et al. 2023). 18 A possible additional explanation lies in population-scaled recombination rates, which are expected 19 to be higher in autotetraploids due to the additional chromosomal partners available during recombination 20 (Pecinka et al. 2011). A higher population recombination rate could slow down the fixation of variants, 21 promoting ‘haplotypic homogenization’ (high representation of intermediate allele frequencies at candidate 22 loci). Yant et al. (2013) previously detected a lower-than-double crossover rate in autotetraploids of A. arenosa. 23 This may explain why, in our analysis, sweeps breadth are similar between ploidies (Fig. 3d, e). However, the 24 greater sweep magnitude at similar sweep breadth in polyploids suggests a higher number of recombination 25 events, potentially influencing fixation rates (Fig. 3d, e). Finally, the observed ratio magnitude/breadth could 26 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
22 also indicate stronger selection (s) in tetraploids. Yet, when s is estimated as the amount of decaying within1 population haplotypic similarity around selected loci, tetraploid populations exhibit lower apparent selection 2 intensities during adaptation (as observed in our DMC results, Fig. 3c), which is consistent with an over3 representation of intermediate allele frequencies in candidate loci. In summary, the observed remarkable 4 differences in positive selection footprints between diploid and polyploid genomes likely result from the 5 interplay of multiple factors, with dominance and masking effects playing a major role, consistent with recent 6 findings on purifying selection in A. arenosa (Vlček et al. 2025). 7 8 Sampling pleiotropic genes from shared variation facilitates fast adaptation, especially in tetraploid 9 populations. 10 Distinguishing among sources of adaptive variation (de novo mutations, gene flow or standing variation) 11 provides insight about the evolutionary potential of populations, lineages or cytotypes. Theory suggests that 12 access to sources of adaptation may differ between diploid and autopolyploid populations, but the direction 13 of this difference remains equivocal (Van de Peer et al. 2017; Konečná et al. 2021). In the present study, we 14 addressed this controversy by empirically inferring evolutionary sources of adaptive variation from naturally 15 replicated instances of substrate adaptation using the modelling approach distinguishing among modes of 16 convergent adaptation (DMC). We found that, in both diploids and tetraploids, adaptation primarily draws from 17 the standing genetic variation already present within each cytotype prior to the onset of selection (Fig. 5a). 18 While DMC might bring biases given by its additivity assumption and different expectations on the distribution 19 of dominance between ploidy levels (e.g. Marad et al. 2018), predominant parallel adaptive sourcing from 20 standing variation is in line with expectations for closely related populations (Bohutínská et al. 2021; 21 Bohutínská and Peichel 2024) and has been documented in many diploid systems adapting to diverse 22 challenges, such as vertebrates (Jones et al. 2012; Reid et al. 2016; Terekhanova et al. 2019; Louis et al. 2021; 23 Rubin et al. 2022; Gutiérrez-Guerrero et al. 2024), insects (Ji et al. 2020; Chaturvedi et al. 2022; Montejo24 Kovacevich et al. 2022) and plants (Bohutínská et al. 2021; Choi et al. 2021), but also autotetraploids of A. 25 arenosa adapting to serpentine conditions (Konečná et al. 2021). However, we have found a significant 26 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
23 difference in the relative importance of sources of adaptive variation between ploidies. In particular we found 1 a virtual lack of tetraploid adaptation building on de novo mutations. This might reflect a strong effect of 2 masking on novel mutations in tetraploids, as they are initially present in lower frequencies and rare recessive 3 or additive mutations are less efficiently targeted by selection in polyploids (Otto and Whitton 2000; Monnahan 4 and Brandvain 2020). In fact, a beneficial allele already present as standing variation, rather than arising de 5 novo, is more likely to become fixed because it typically exists at a higher frequency when selective conditions 6 change (Innan and Kim 2004; Barrett and Schluter 2008). Moreover, standing variation is likely to lead to more 7 rapid evolution in novel environments because beneficial alleles are immediately available, in contrast to de 8 novo beneficial mutations, which require time to arise (Ji et al. 2020; Rubin et al. 2022). All else being equal, 9 the fixation time of an allele is affected by the magnitude of its beneficial effect, the effective population size 10 (Ne) and the number of copies of the allele, which is inversely proportional to the likelihood of the allele being 11 lost because of genetic drift (Hermisson and Pennings 2005). Consequently, considering that polyploid 12 populations often display a higher Ne and/or harbour a larger pool of variation (Van de Peer et al. 2017; Baduel, 13 Bray, et al. 2018), polyploids will source from pre-existing variation more often, relatively to the initially rare de 14 novo mutations, than diploids would do. 15 In addition to population genetic parameters, adaptation speed and success can be linked to the 16 propensity of target genes, or their products, to undergo evolutionary changes, and it has been suggested that 17 such a constraint depends on the gene’s position in the molecular network (Orr 2000). However, it remains 18 unclear whether such characteristics differ depending on ploidy. In a protein–protein interaction network 19 (PPIN), proteins occupying a central position have been shown to exert a strong influence over the pathway 20 output, and thus on the final phenotypes and the organism’s fitness (therefore, connectivity correlates 21 positively with pleiotropy – Promislow 2004; Tew et al. 2007; Tyler et al. 2009; Wang et al. 2010). Therefore, 22 positive selection on genes encoding central proteins can drive rapid adaptation; however, here new mutations 23 are more likely to be deleterious and are consequently quickly purged (Olson-Manning et al. 2012). 24 Interestingly, in accord with theory (Wang et al. 2010) and recent empirical studies on sticklebacks (Rennison 25 and Peichel 2022) and Arabidopsis thaliana (Frachon et al. 2017), we found that repeatedly selected candidate 26 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
24 genes of A. arenosa have, on average, an intermediate level of pleiotropy/connectivity (Fig. 5b). Thus, this result 1 adds to the existing body of work that conflicts with earlier theoretical expectations suggesting that higher 2 levels of pleiotropy are constraining and disfavour adaptation (Otto 2004; Kim et al. 2007; Masalia et al. 2017). 3 Notably, large-effect loci have been suggested to possess a selective advantage when a population is 4 colonizing a new environment (Orr 1998; Hämälä et al. 2020), because the cost of complexity imposed by 5 pleiotropy gets balanced out by its positive effect pushing the population faster towards a new fitness optimum 6 (Wang et al. 2010). This process might even be less costly when selection targets adaptive loci standing in the 7 population pool that are likely to have been pre-tested by selection during previous selective pressures (Barrett 8 and Schluter 2008). Here, we suggest that the species A. arenosa as a whole might have been able to quickly 9 adapt to new stress factors associated with siliceous or calcareous bedrock by leveraging intermediate-effect 10 loci that were likely to be advantageous because sourced mainly from pre-tested standing or shared variation. 11 Moreover, increased genetic redundancy of polyploid genomes may translate into lower selective 12 constraints. Theory proposes that expression networks in polyploids are characterized by a greater redundancy 13 given by the multiplication of each node (i.e. gene) and edge (i.e. interaction) (Parisod et al. 2024; Ebadi et al. 14 2023). In line with this expectation, candidate genes of the tetraploid cytotype display a significantly higher 15 average level of connectivity than diploid ones (Fig. 5b, c). Our results suggest that this redundancy, in line with 16 simulations results from (Yao et al. 2019), might promote robustness against new mutations, allowing even 17 more central genes to explore novel protein interactions and outputs while preserving their ancestral functions 18 (Cuypers and Hogeweg 2014; Carretero-Paulet and Van de Peer 2020; Parisod 2024). For the reasons above, 19 WGD has the potential to break genomic constraints imposed by pleiotropy, a mechanism that might 20 compensate for the slower rate of fixation and the dominance-dependent efficiency of selection being 21 experienced, allowing polyploids to ‘jump’ in the fitness landscape quickly reaching new fitness optima – hence 22 the definition of ‘hopeful monsters’ (Cuypers and Hogeweg 2014; Keane et al. 2014; Ebadi et al. 2023). 23 However, further empirical work directly comparing diploid A. arenosa against tetraploid gene regulatory 24 networks would be beneficial in order to corroborate and shed light on the mechanisms of such a process. 25 26 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
25 Conclusions 1 Here, we have identified and compared signatures of positive selection across ploidies, providing insights into 2 how whole-genome duplication influences evolutionary trajectories. Our results support a scenario where, 3 despite the potential constraints imposed by polysomic masking and the delayed fixation of beneficial alleles, 4 tetraploid populations efficiently navigate selective pressures. Their inherent retention of a substantial 5 reservoir of standing genetic variation translates into an evolutionary potential that enables them to persist and 6 diversify in dynamic landscapes, be it in response to gradual ecological shifts or rapid environmental 7 perturbations (Barrett and Schluter 2008; Keane et al. 2014; Fox et al. 2020). 8 9 Materials and Method 10 Genetic and soil dataset collection 11 Siliceous and calcareous soils exhibit a patchy distribution throughout the European landscape and ploidies 12 of Arabidopsis arenosa occur in both these environments throughout their native ranges. Aiming to cover the 13 majority of the species’ distribution along the soil gradients, we assembled whole-genome resequencing data 14 for 76 populations (479 individuals) covering all lineages of A. arenosa (in line with Monnahan et al. (2019) and 15 Padilla-García et al. (2023) except the invasively spreading Ruderal lineage, which exhibits weak substrate 16 preferences, as it occupies human-disturbed non-native habitats (Baduel, Hunter, et al. 2018; Monnahan et al. 17 2019). Specifically, we used previously published data on 35 verified diploid (2x) and 41 verified tetraploid (4x) 18 populations (Hollister et al. 2012; Yant et al. 2013; Arnold et al. 2016; Novikova et al. 2016; Monnahan et al. 19 2019; Bohutínská et al. 2021; Konečná et al. 2021). At each population location, ~ 1 litre of soil was dug from 20 ~ 10–20 cm below the ground in three different spots in the vicinity of A. arenosa plants. Two diploid Baltic 21 populations included in this study lacked soil information because of complicated logistics (occurrence in 22 Ukraine and Lithuania), but they were included in the selection scan (see below), reflecting their sharply 23 contrasting substrate preferences inferred from our field notes on the type of bedrock and accompanying 24 vegetation (a limestone outcrop vs acidic sandy soil). 25 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
32 within-population diversity metrics are used, especially considering the number of samples available. For this 1 reason, we focused on between-population differentiation, calculated as Rho for each population pair 2 designed as above. We calculated allele frequencies and Rho metrics for the paired dataset using the 3 ScanTools pipeline onto a vcf file containing data on all biallelic SNPs (supplementary table S1; the output was 4 used also for the calculations presented in Fig. 3a, b, and d). For each top candidate gene, we measured its 5 selective sweep in terms of magnitude and breadth. ‘Magnitude’ was calculated as the difference between the 6 highest Rho value within the sweep region and the average genome-wide Rho for that population pair. ‘Breadth’ 7 was calculated by extracting Rho values for 100 kbp at each side from the middle of the gene (i.e. the curve 8 region) and fitting a curve using the function loess.as from the R package fANCOVA v.0.6-1, which automatically 9 sets a smoothing parameter using the bias-corrected Akaike information criterion (AICC). Specifically, 10 ‘breadth’ was calculated as the length of the region, in bp, under the sweep curve, where the borders 11 corresponded to the intersection of the fitted curve with the defined threshold of double the value of 12 ‘background’ Rho (average Rho in the 200 kbp flanking the curve region). 13 14 Identifying the evolutionary source of shared adaptive variants 15 After having identified a shortlist of parallel candidate genes (top candidate genes) likely associated with 16 adaptation to siliceous or calcareous soil in A. arenosa, we proceeded to infer their evolutionary source. Using 17 a designated composite likelihood-based method, referred to as the ‘distinguishing among modes of 18 convergent adaptation’ (DMC) (Lee and Coop 2017; Lee and Coop 2019), we modelled patterns of (i) neutral 19 variation, (ii) parallel selection operating on de novo mutations at the same loci, (iii) parallel selection acting 20 on pre-existing standing variation, and (iv) transfer of variants through gene flow. For each scenario, this 21 method compares patterns of allele frequency covariance of deviation from the ancestral state (the coancestry 22 coefficient) at loci near the site subject to selection, within and between (selected and non-selected) 23 populations, as a function of the recombination rate. Because this approach is based on allele frequency 24 covariance, which is a relative measure rather than an absolute one, it remains informative regardless of ploidy 25 level. Whereas autotetraploids exhibit additional complexities due to polysomic masking and potential dosage 26 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
33 effects, the core principles of coancestry-based inference remain valid. By running DMC in a loop 1 (https://github.com/mbohutinska/DMCloop) over all top candidate regions, we estimated the composite log2 likelihoods of each scenario for each parallel gene along each of the 21 population quartets (15 quartets in 3 diploids and 6 in tetraploids), using the same pairs as for the PicMin analysis. We set the model parameters 4 taking into account the populations’ evolutionary history inferred by our previous investigations in A. arenosa 5 (Bohutínská et al. 2021; Konečná et al. 2021), indicating no difference between the ploidies; see 6 supplementary table S2. Composite log-likelihoods were calculated at twelve locations at equal distances 7 along each gene (plus 25 kbp upstream and downstream), as recommended by the author at 8 https://github.com/kristinmlee/dmc and in accordance with the range of LD decay (150–800 bp) previously 9 estimated in A. arenosa (Bohutínská et al. 2021). Not knowing the exact direction of the habitat shift for each 10 quartet, we calculated results for the divergence caused by selection in populations growing on both siliceous 11 and calcareous soil and merged the results. We then identified the best-fitting scenario for each gene and each 12 quartet, only accepting significantly non-neutral cases inferred by comparing models with simulated neutral 13 data in a similar manner as in Bohutínská et al. (2021). 14 15 Construction of protein–protein interaction networks 16 Functional protein-protein interaction networks (PPIN) were constructed using nodes and edges downloaded 17 from the STRING v.12 (https://string-db.org) (Szklarczyk et al. 2023) protein network database (scored links 18 between proteins) for A. thaliana proteins. We did not subset node interactions by applying a threshold on the 19 protein–protein interaction combined score; however, our results do not change when it is applied 20 (supplementary table S3). We associated the ID of each protein sequence (Arabidopsis Genome Initiative, 21 AGI) to a UniProt accession, using the TAIR10 release of Ids conversions 22 (https://www.arabidopsis.org/download/list?dir=Proteins%2FId_conversions) (Berardini et al. 2015). Within 23 the lists of top candidate genes, two and three genes were not included in the construction of the PPIN in 24 diploids and tetraploids, respectively, as they could not be associated to any A. thaliana ortholog. Moreover, 25 for one gene in each ploidy there was no associated protein in the STRING database, leaving us with a final 26 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
34 number of 40 and 24 candidate genes in diploids and tetraploids, respectively. For each ploidy dataset, we 1 tested if the average node number of interactions (connectivity) was significantly greater than the average 2 node connectivity of the full A. thaliana dataset by performing a one-sided permutation test with 10,000 3 permutations in R. We then visualized each ploidy PPIN in R using the packages igraph v.1.3.1 (Csardi and 4 Nepusz 2006) and tcltk v.4.1.2 (Grosjean 2022). 5 6 Acknowledgements 7 We thank Miroslav Poláček and Karolína Havlíková for assistance during fieldwork and Magdalena Bohutínská, 8 Levi Yant, Edita Tylová, Stanislav Kopřiva, Patrick Monnahan and members of the Plant Ecological Genomics 9 group for feedback on an earlier version of this manuscript. 10 This research was funded by the Charles University (GAUK project 243-252135 awarded to SC), the 11 European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation 12 programme (ERC-StG 850852 DOUBLE ADAPT to FK), and the Czech Science Foundation (project 23-07204M 13 to FK). Additional support was provided by the Czech Academy of Sciences (long-term research development 14 project RVO 67985939). Access to computing and storage facilities owned by the parties and projects 15 contributing to the National Grid Infrastructure MetaCentrum was provided under the programme ‘Projects of 16 Large Research, Development, and Innovations Infrastructures’ (CESNET LM2015042) and is greatly 17 appreciated. 18 19 Data Availability 20 Genetic data is accessible under BioProjects PRJNA284572, PRJNA309929, PRJNA357693, PRJNA357372, 21 PRJNA459481, PRJNA493227, PRJEB34247 (ENA), PRJNA506705, PRJNA484107, PRJNA592307, PRJNA667586 22 and PRJNA929698 (see supplementary data S1 for details). Soil measured parameters as pH H2O, pH KCl, 23 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
35 bioavailable Ca, total content of nutrients, such as Ca, Mg, K and Na, and cation exchange capacity (CEC), are 1 available for each population in the supplementary data S1. 2 3 References 4 Alexa A RJ. 2024. topGO: Enrichment Analysis for Gene Ontology. Available from: 5 http://bioconductor.org/packages/topGO/ 6 Alix K, Gérard PR, Schwarzacher T, Heslop-Harrison JSP. 2017. Polyploidy and interspecific hybridization: 7 partners for adaptation, speciation and evolution in plants. Ann. Bot. 120:183–194. 8 Alvarez N, Thiel-Egenter C, Tribsch A, Holderegger R, Manel S, Schönswetter P, Taberlet P, Brodbeck S, 9 Gaudeul M, Gielly L, et al. 2009. History or ecology? Substrate type as a major driver of spatial 10 genetic structure in Alpine plants. Ecol. Lett. 12:632–640. 11 Alves JM, Carneiro M, Cheng JY, Lemos de Matos A, Rahman MM, Loog L, Campos PF, Wales N, Eriksson A, 12 Manica A, et al. 2019. Parallel adaptation of rabbit populations to myxoma virus. Science 363:1319– 13 1326. 14 Anjali, Vanisri, Himabindu, Rathod S, Govindaraj, Nagamani. 2021. Insilico analysis of Arabidopsis ferric 15 reductase oxidases (FRO) proteins associated with iron homeostasis. Pharma Innov. 10:303–310. 16 Arnold B, Kim S-T, Bomblies K. 2015. Single Geographic Origin of a Widespread Autotetraploid Arabidopsis 17 arenosa Lineage Followed by Interploidy Admixture. Mol. Biol. Evol. 32:1382–1395. 18 Arnold BJ, Lahner B, DaCosta JM, Weisman CM, Hollister JD, Salt DE, Bomblies K, Yant L. 2016. Borrowed 19 alleles and convergence in serpentine adaptation. Proc. Natl. Acad. Sci. U. S. A. 113:8320–8325. 20 Baduel P, Bray S, Vallejo-Marin M, Kolář F, Yant L. 2018. The “Polyploid Hop”: Shifting Challenges and 21 Opportunities Over the Evolutionary Lifespan of Genome Duplications. Frontiers in Ecology and 22 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
36 Evolution 6:117. 1 Baduel P, Hunter B, Yeola S, Bomblies K. 2018. Genetic basis and evolution of rapid cycling in railway 2 populations of tetraploid Arabidopsis arenosa. PLoS Genet. 14:e1007510. 3 Baniaga AE, Marx HE, Arrigo N, Barker MS. 2020. Polyploid plants have faster rates of multivariate niche 4 differentiation than their diploid relatives. Ecol. Lett. 23:68–78. 5 Barrett RDH, Schluter D. 2008. Adaptation from standing genetic variation. Trends Ecol. Evol. 23:38–44. 6 Bastida JM, Rey P, Alcántara J. 2014. Plant performance and morpho-functional differentiation in response to 7 edaphic variation in Iberian columbines: cues for range distribution? Journal of Plant Ecology 7:403– 8 412. 9 te Beest M, Le Roux JJ, Richardson DM, Brysting AK, Suda J, Kubesová M, Pysek P. 2012. The more the better? 10 The role of polyploidy in facilitating plant invasions. Ann. Bot. 109:19–45. 11 Berardini TZ, Reiser L, Li D, Mezheritsky Y, Muller R, Strait E, Huala E. 2015. The Arabidopsis information 12 resource: Making and mining the “gold standard” annotated reference plant genome. Genesis 13 53:474–485. 14 Bohutínská M, Peichel CL. 2024. Divergence time shapes gene reuse during repeated adaptation. Trends 15 Ecol. Evol. 39:396–407. 16 Bohutínská M, Vlček J, Yair S, Laenen B, Konečná V, Fracassetti M, Slotte T, Kolář F. 2021. Genomic basis of 17 parallel adaptation varies with divergence in Arabidopsis and its relatives. Proc. Natl. Acad. Sci. U. S. 18 A. 118:e2022713118. 19 Bolger AM, Lohse M, Usadel B. 2014. Trimmomatic: a flexible trimmer for Illumina sequence data. 20 Bioinformatics 30:2114–2120. 21 Bontpart T, Weiss A, Vile D, Gérard F, Lacombe B, Reichheld J-P, Mari S. 2024. Growing on calcareous soils 22 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
37 and facing climate change. Trends Plant Sci. 29:1319–1330. 1 Booker TR, Yeaman S, Whitlock MC. 2023. Using genome scans to identify genes used repeatedly for 2 adaptation. Evolution 77:801–811. 3 Bothe H. 2015. The lime–silicate question. Soil Biol. Biochem. 89:172–183. 4 Bothe H, Ferguson S, Newton WE. 2006. Biology of the Nitrogen Cycle. Elsevier 5 Bray SM, Hämälä T, Zhou M, Busoms S, Fischer S, Desjardins SD, Mandáková T, Moore C, Mathers TC, Cowan 6 L, et al. 2024. Kinetochore and ionomic adaptation to whole-genome duplication in Cochlearia 7 shows evolutionary convergence in three autopolyploids. Cell Rep. 43:114576. 8 Carretero-Paulet L, Van de Peer Y. 2020. The evolutionary conundrum of whole-genome duplication. Am. J. 9 Bot. 107:1101–1105. 10 Caye K, Jumentier B, Lepeule J, François O. 2019. LFMM 2: Fast and Accurate Inference of Gene-Environment 11 Associations in Genome-Wide Studies. Mol. Biol. Evol. 36:852–860. 12 Chaturvedi S, Gompert Z, Feder JL, Osborne OG, Muschick M, Riesch R, Soria-Carrasco V, Nosil P. 2022. 13 Climatic similarity and genomic background shape the extent of parallel adaptation in Timema stick 14 insects. Nat Ecol Evol 6:1952–1964. 15 Choi JY, Dai X, Alam O, Peng JZ, Rughani P, Hickey S, Harrington E, Juul S, Ayroles JF, Purugganan MD, et al. 16 2021. Ancestral polymorphisms shape the adaptive radiation of Metrosideros across the Hawaiian 17 Islands. Proc. Natl. Acad. Sci. U. S. A. 118:e2023801118. 18 Clark R, Baligar V. 2000. Acidic and alkaline soil constraints on plant mineral nutrition. In: Plant-Environment 19 Interactions. CRC Press. p. 133–178. 20 Clo J. 2022. Polyploidization: Consequences of genome doubling on the evolutionary potential of 21 populations. Am. J. Bot. 109:1213–1220. 22 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
38 Comai L. 2005. The advantages and disadvantages of being polyploid. Nat. Rev. Genet. 6:836–846. 1 Csardi G, Nepusz T. 2006. The igraph software package for complex network research. InterJournal Complex 2 Systems:1695. 3 Cuypers TD, Hogeweg P. 2014. A synergism between adaptive effects and evolvability drives whole genome 4 duplication to fixation. PLoS Comput. Biol. 10:e1003547. 5 DePristo MA, Banks E, Poplin R, Garimella KV, Maguire JR, Hartl C, Philippakis AA, del Angel G, Rivas MA, 6 Hanna M, et al. 2011. A framework for variation discovery and genotyping using next-generation DNA 7 sequencing data. Nat. Genet. 43:491–498. 8 Dukić M, Bomblies K. 2022. Male and female recombination landscapes of diploid Arabidopsis arenosa. 9 Genetics 220:iyab236. 10 Durinck S, Spellman PT, Birney E, Huber W. 2009. Mapping identifiers for the integration of genomic datasets 11 with the R/Bioconductor package biomaRt. Nat. Protoc. 4:1184–1191. 12 Ebadi M, Bafort Q, Mizrachi E, Audenaert P, Simoens P, Van Montagu M, Bonte D, Van de Peer Y. 2023. The 13 duplication of genomes and genetic networks and its potential for evolutionary adaptation and 14 survival during environmental turmoil. Proc. Natl. Acad. Sci. 120:e2307289120. 15 Fourcroy P, Tissot N, Gaymard F, Briat J, Dubos C. 2016. Facilitated Fe nutrition by phenolic compounds 16 excreted by the Arabidopsis ABCG37/PDR9 transporter requires the IRT1/FRO2 high-affinity root 17 Fe(2+) transport system. Mol. Plant 9:485–488. 18 Fox DT, Soltis DE, Soltis PS, Ashman T-L, Van de Peer Y. 2020. Polyploidy: A Biological Force From Cells to 19 Ecosystems. Trends Cell Biol. 30:688–694. 20 Frachon L, Libourel C, Villoutreix R, Carrère S, Glorieux C, Huard-Chauveau C, Navascués M, Gay L, Vitalis R, 21 Baron E, et al. 2017. Intermediate degrees of synergistic pleiotropy drive adaptive evolution in 22 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
39 ecological time. Nat. Ecol. Evol. 1:1551–1561. 1 Garsmeur O, Schnable JC, Almeida A, Jourda C, D’Hont A, Freeling M. 2014. Two evolutionarily distinct 2 classes of paleopolyploidy. Mol. Biol. Evol. 31:448–454. 3 Grosjean P. 2022. SciViews-R: A GUI API for R. UMONS Available from: http://www.sciviews.org/SciViews-R 4 Guggisberg A, Liu X, Suter L, Mansion G, Fischer MC, Fior S, Roumet M, Kretzschmar R, Koch MA, Widmer A. 5 2018. The genomic basis of adaptation to calcareous and siliceous soils in Arabidopsis lyrata. Mol. 6 Ecol. 27:5088–5103. 7 Gutiérrez-Guerrero YT, Phifer-Rixey M, Nachman MW. 2024. Across two continents: The genomic basis of 8 environmental adaptation in house mice (Mus musculus domesticus) from the Americas. PLoS 9 Genet. 20:e1011036. 10 Hämälä T, Gorton AJ, Moeller DA, Tiffin P. 2020. Pleiotropy facilitates local adaptation to distant optima in 11 common ragweed (Ambrosia artemisiifolia). PLoS Genet. 16:e1008707. 12 Hämälä T, Moore C, Cowan L, Carlile M, Gopaulchan D, Brandrud MK, Birkeland S, Loose M, Kolář F, Koch MA, 13 et al. 2024. Impact of whole-genome duplications on structural variant evolution in Cochlearia. Nat. 14 Commun. 15:5377. 15 Hedrich R. 2012. Ion channels in plants. Physiol. Rev. 92:1777–1811. 16 Hermisson J, Pennings PS. 2005. Soft sweeps: molecular population genetics of adaptation from standing 17 genetic variation: Molecular population genetics of adaptation from standing genetic variation. 18 Genetics 169:2335–2352. 19 Hill WG, Robertson A. 1966. The effect of linkage on limits to artificial selection. Genetical research 8:269– 20 294. 21 Hollister JD, Arnold BJ, Svedin E, Xue KS, Dilkes BP, Bomblies K. 2012. Genetic adaptation associated with 22 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
40 genome-doubling in autotetraploid Arabidopsis arenosa. PLoS Genet. 8:e1003093. 1 Hu TT, Pattyn P, Bakker EG, Cao J, Cheng J-F, Clark RM, Fahlgren N, Fawcett JA, Grimwood J, Gundlach H, et al. 2 2011. The Arabidopsis lyrata genome sequence and the basis of rapid genome size change. Nat. 3 Genet. 43:476–481. 4 Innan H, Kim Y. 2004. Pattern of polymorphism after strong artificial selection in a domestication event. Proc. 5 Natl. Acad. Sci. U. S. A. 101:10667–10672. 6 Ji Y, Li X, Ji T, Tang J, Qiu L, Hu J, Dong J, Luo S, Liu S, Frandsen PB, et al. 2020. Gene reuse facilitates rapid 7 radiation and independent adaptation to diverse habitats in the Asian honeybee. Sci. Adv. 8 6:eabd3590. 9 Jones FC, Grabherr MG, Chan YF, Russell P, Mauceli E, Johnson J, Swofford R, Pirun M, Zody MC, White S, et 10 al. 2012. The genomic basis of adaptive evolution in threespine sticklebacks. Nature 484:55–61. 11 Kang JG, Pyo YJ, Cho JW, Cho MH. 2004. Comparative proteome analysis of differentially expressed proteins 12 induced by K+ deficiency in Arabidopsis thaliana. Proteomics 4:3549–3559. 13 Keane OM, Toft C, Carretero-Paulet L, Jones GW, Fares MA. 2014. Preservation of genetic and regulatory 14 robustness in ancient gene duplicates of Saccharomyces cerevisiae. Genome Res. 24:1830–1841. 15 Keightley PD, Jackson BC. 2018. Inferring the probability of the derived vs. The ancestral Allelic state at a 16 polymorphic site. Genetics 209:897–906. 17 Kim PM, Korbel JO, Gerstein MB. 2007. Positive selection at the protein network periphery: evaluation in terms 18 of structural constraints and cellular context. Proc. Natl. Acad. Sci. U. S. A. 104:20274–20279. 19 Kim SA, Guerinot ML. 2007. Mining iron: iron uptake and transport in plants. FEBS Lett. 581:2273–2280. 20 Koch M. 2018. The plant model system Arabidopsis set in an evolutionary, systematic, and spatio-temporal 21 context. J. Exp. Bot. 70:55–67. 22 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
41 Kolář F, Dortová M, Lepš J, Pouzar M, Krejčová A, Štech M. 2014. Serpentine ecotypic differentiation in a 1 polyploid plant complex: shared tolerance to Mg and Ni stress among diand tetraploid serpentine 2 populations of Knautia arvensis (Dipsacaceae). Plant Soil 374:435–447. 3 Kolář F, Fuxová G, Záveská E, Nagano AJ, Hyklová L, Lučanová M, Kudoh H, Marhold K. 2016. Northern glacial 4 refugia and altitudinal niche divergence shape genome-wide differentiation in the emerging plant 5 model Arabidopsis arenosa. Mol. Ecol. 25:3929–3949. 6 Konečná V, Bray S, Vlček J, Bohutínská M, Požárová D, Choudhury RR, Bollmann-Giolai A, Flis P, Salt DE, 7 Parisod C, et al. 2021. Parallel adaptation in autopolyploid Arabidopsis arenosa is dominated by 8 repeated recruitment of shared alleles. Nat. Commun. 12:4979. 9 Lee I, Ambaru B, Thakkar P, Marcotte EM, Rhee SY. 2010. Rational association of genes with traits using a 10 genome-scale gene network for Arabidopsis thaliana. Nat. Biotechnol. 28:149–156. 11 Lee KM, Coop G. 2017. Distinguishing Among Modes of Convergent Adaptation Using Population Genomic 12 Data. Genetics 207:1591–1619. 13 Lee KM, Coop G. 2019. Population genomics perspectives on convergent adaptation. Philos. Trans. R. Soc. 14 Lond. B Biol. Sci. 374:20180236. 15 Li H, Durbin R. 2009. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 16 25:1754–1760. 17 Lipánová V, Kabátová KN, Zeisek V, Kolář F, Chrtek J. 2023. Evolution of the Sabulina verna group 18 (Caryophyllaceae) in Europe: A deep split, followed by secondary contacts, multiple 19 allopolyploidization and colonization of challenging substrates. Mol. Phylogenet. Evol. 189:107940. 20 Llorente F, López-Cobollo RM, Catalá R, Martínez-Zapater JM, Salinas J. 2002. A novel cold-inducible gene 21 from Arabidopsis, RCI3, encodes a peroxidase that constitutes a component for stress tolerance. 22 Plant J. 32:13–24. 23 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
48 18:411–424. 1 Van der Auwera GA, Carneiro MO, Hartl C, Poplin R, Del Angel G, Levy-Moonshine A, Jordan T, Shakir K, 2 Roazen D, Thibault J, et al. 2013. From FastQ data to high confidence variant calls: the Genome 3 Analysis Toolkit best practices pipeline. Curr. Protoc. Bioinformatics 43:11.10.1-11.10.33. 4 Venables WN and Ripley BD. 2002. Modern Applied Statistics with S. Fourth Edition. New York, NY: Springer 5 Vert GA, Briat J-F, Curie C. 2003. Dual regulation of the Arabidopsis high-affinity root iron uptake system by 6 local and long-distance signals. Plant Physiol. 132:796–804. 7 Vlček J, Hämälä T, Vives Cobo C, Curran E, Šrámková G, Slotte T, Schmickl R, Yant L, Kolář F. 2025. Whole8 genome duplication increases genetic diversity and load in outcrossing Arabidopsis arenosa. Proc. 9 Natl. Acad. Sci. U. S. A. 122:e2501739122. 10 Wang M, Zhao Y, Zhang B. 2015. Efficient test and visualization of multi-set intersections. Sci. Rep. 5:16923. 11 Wang Z, Liao B-Y, Zhang J. 2010. Genomic patterns of pleiotropy and the evolution of complexity. Proc. Natl. 12 Acad. Sci. U.S.A. 107:18034–18039. 13 Wendel JF. 2015. The wondrous cycles of polyploidy in plants. Am. J. Bot. 102:1753–1756. 14 Wos G, Požárová D, Kolář F. 2023. Role of phenotypic and transcriptomic plasticity in alpine adaptation of 15 Arabidopsis arenosa. Mol. Ecol. 32:5771–5784. 16 Wu S, Han B, Jiao Y. 2020. Genetic Contribution of Paleopolyploidy to Adaptive Evolution in Angiosperms. 17 Mol. Plant 13:59–71. 18 Xu N, Cheng L, Kong Y, Chen G, Zhao L, Liu F. 2024. Functional analyses of the NRT2 family of nitrate 19 transporters in Arabidopsis. Front. Plant Sci. 15:1351998. 20 Yang L, Jin Y, Huang W, Sun Q, Liu F, Huang X. 2018. Full-length transcriptome sequences of ephemeral plant 21 Arabidopsis pumila provides insight into gene expression dynamics during continuous salt stress. 22 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
49 BMC Genomics 19:717. 1 Yant L, Bomblies K. 2017. Genomic studies of adaptive evolution in outcrossing Arabidopsis species. Curr. 2 Opin. Plant Biol. 36:9–14. 3 Yant L, Hollister JD, Wright KM, Arnold BJ, Higgins JD, Franklin FCH, Bomblies K. 2013. Meiotic adaptation to 4 genome duplication in Arabidopsis arenosa. Curr. Biol. 23:2151–2156. 5 Yao Y, Carretero-Paulet L, Van de Peer Y. 2019. Using digital organisms to study the evolutionary 6 consequences of whole genome duplication and polyploidy. PLoS One 14:e0220257. 7 York ET Jr, Bradfield R, Peech M. 1953. Calcium-potassium interactions in soils and plants: II. Reciprocal 8 relationship between calcium and potassium in plants. Soil Sci. 76:481–492. 9 Yu G, Li J, Sun X, Zhang X, Liu J, Pan H. 2015. Overexpression of AcNIP5;1, a novel nodulin-like intrinsic protein 10 from halophyte Atriplex canescens, enhances sensitivity to salinity and improves drought tolerance 11 in Arabidopsis. Plant Mol. Biol. Rep. 33:1864–1875. 12 Zhang F, Long R, Ma Z, Xiao H, Xu X, Liu Z, Wei C, Wang Y, Peng Y, Yang X, et al. 2024. Evolutionary genomics of 13 climatic adaptation and resilience to climate change in alfalfa. Mol. Plant 17:867–883. 14 Zhang L, Wu S, Chang X, Wang X, Zhao Y, Xia Y, Trigiano RN, Jiao Y, Chen F. 2020. The ancient wave of 15 polyploidization events in flowering plants and their facilitated adaptation to environmental stress. 16 Plant Cell Environ. 43:2847–2856. 17 18 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
50 Fig. 1 Genetic relationships, diversity, distribution and soil conditions of the Arabidopsis arenosa populations 1 under study. a) Geographical distribution in central and southeastern Europe of the 76 genome2 sequenced populations studied. Circles indicate diploid populations and triangles indicate 3 tetraploid populations. The subset of populations used for pairwise divergence contrasts (paired 4 dataset) are annotated with their relative pair number (diploid, 2x, 1–7; tetraploid, 4x, 8–14). b) 5 Within-population diversity (pairwise nucleotide diversity, ) and estimate of neutrality (Tajima’s D) 6 calculated from putatively neutral fourfold degenerated SNPs detected in 35 diploid and 41 tetraploid 7 populations. Dots refer to the values for each population used in the paired dataset. c) PCA of the 8 elemental soil composition for each population sampled (dots coloured by ploidy) (left) and variable 9 correlation plot (right). Populations used in the paired dataset are indicated by their corresponding 10 number. Soil variables are coloured by their contribution, expressed in percentage values, to the 11 explained variation along the first two principal component axes. d, e) Relationships between diploid 12 (d) and tetraploid (e) populations of the paired dataset, depicted using an allele frequency 13 covariance graph calculated in Treemix. Each population is labelled by its corresponding pair 14 number and corresponding major genetic lineage, following previous range-wide studies (Monnahan 15 et al. 2019; Padilla-García et al. 2023). All branches are significantly supported (bootstrap > 96%), the 16 sole exception being marked with a black dot. Throughout the entire figure, the green–violet colour 17 ramp reflects the position of each population along the PC1 axis of elemental soil composition PCA – 18 bar on top in figure c – ranging from calcareous (violet) to siliceous (green). 19 Fig. 2 Genomic basis of calcareous/siliceous substrate adaptation in diploid and 20 autotetraploid Arabidopsis arenosa. a, b) Manhattan plots comparing the differentiation 21 between populations on siliceous versus calcareous soils in terms of: repeated 22 differentiation evaluated by means of order statistics in PicMin (upper row), Weir and 23 Cockerham FST per SNP (middle row) and the association of per-SNP allele frequency and 24 soil properties, approximated by scores on the soil-associated PC1 axis by means of latent 25 factor mixed model (LFMM) analysis (lower row), in diploid (a) and tetraploid (b) 26 populations. In PicMin, the dashed line indicates a significance threshold of q = 0.05, 27 inferred by means of order statistics, and nest refers to the estimated number of contrasts 28 (i.e. the couples used in the analyses) exhibiting a pattern of repeated adaptation. Here, the 29 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
51 FST Manhattan plots were generated pooling together the samples within each 1 environment and ploidy to allow a summary visualization. In the LFMM analysis, the black 2 dashed horizontal line refers to the applied threshold calculated as the 0.05 q-value. 3 Selected top gene candidates with known function relevant to substrate adaptation are 4 highlighted in green and annotated with their name. c, e) Venn diagrams for diploid (c), and 5 tetraploid (e) describing the total number of candidate genes overlapping among the three 6 methods. d) Venn diagram depicting the overlap of selected top genes between the 7 ploidies. f, g) Significantly enriched biological processes (BP) gene ontology (GO) terms 8 identified among the top candidate genes in diploids (f) and tetraploids (g). 9 Fig. 3 Different footprints of selection in top candidate genes for substrate adaptation in 10 diploid and tetraploid A. arenosa populations (paired dataset). a) Average increase in 11 genetic differentiation (Rho) in top candidate genes relative to genome-wide values. b) 12 Average number of fixed variants per top candidate gene, calculated across 1-kbp window 13 and population pairs. In both a and b, each population pair was added to the calculation of 14 the mean only if the gene was a significant outlier in that specific pair. c) Proportional 15 representation of bins of selection strengths inferred by the DMC modelling method (see 16 subsection) for each case of top gene parallel adaptation in each ploidy. Strength of 17 selection (maxSel) is calculated as the most likely s at the identified selected gene site for 18 the model with the highest composite log-likelihoods in DMC. We detected a significant 19 association between ploidy and maxSel values of 0.001 and 0.01 using Fisher’s exact test 20 (two-tailed P = 0.009533 and P = 0.01846, respectively). d) Distribution of average allele 21 frequency per SNP (polarised by soil type to aid visualization in the functional context) 22 across ploidies and soil types. Average allele frequency values were calculated per each 23 SNP belonging to a top candidate gene, with a flanking region of 2 kbp. Only candidates 24 from their respective population pairs where they appeared as outliers were included in the 25 calculations. Note that the average allele frequencies in tetraploids are far from fixation 26 and distributed more around intermediate values. e) Sweep magnitude and breadth for 27 each top candidate gene across relevant population pairs. f) Smoothed-profile of the 28 average selective sweep in each ploidy. 29 Figure 4. Different footprints of selection exemplified in one candidate gene found in both 30 diploid and tetraploid A. arenosa (AL6G42430, SULTR1;1). a) Column-clustered heatmap 31 showing allele frequency values for all 114 biallelic SNPs (rows) belonging to the gene 32 (including UTRs, the transcribed region, exons and introns) per each population (column) of 33 the full A. arenosa dataset. Colours in the top bars indicate the population ploidy and score 34 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
52 along the soil-associated PC1 axis. Rows referring to synonymous and nonsynonymous 1 SNPs are coloured dark blue and red, respectively. In the heatmap, full purple cells indicate 2 fixation of the ancestral allele whereas yellow cells indicate absence of the ancestral allele. 3 Tetraploids exhibit a higher incidence of intermediate allele frequencies of individual sites 4 values (0.5) than diploids. b) Local PCA calculated based on variation in the 114 SNPs of the 5 SULTR1;1 gene. Each symbol refers to an individual and colours refer to the population score 6 along the soil-associated PC1 axis. Note that whereas the calcareous/siliceous gradient 7 explains most of the SNP variation (PC1 64.8%), geography, corresponding to scores on PC2, 8 (not shown) has a markedly smaller effect. Tetraploid individuals are spread over a larger 9 area in the plot and exhibit less differentiation between soil types, in line with the overall 10 more intermediate allele frequency at the level of selected genes as compared to diploids. 11 c) Selective sweep profile in diploid population pair 1 (left) and in tetraploid population pair 12 11 (right), indicated as a fitted curve along the distribution of differentiation (Rho) per each 13 SNP (dots) in a genomic window of 200 kbp. The red dotted line represents the local 14 background threshold (twice the average Rho in the 200 kbp flanking regions) and the blue 15 dotted line represents the genome-wide average Rho. The green segment refers to the 16 position of the SULTR1;1 gene model. The vertical black segment is the measured sweep 17 magnitude, and the horizontal black segment is the measured sweep breadth. d) Allele 18 frequency difference (AFD) between populations on calcareous versus siliceous soil for the 19 locus (dots) and the maximum composite log-likelihood (MCL) estimation of the source of 20 the selected alleles, inferred in DMC (lines), one of the diploid quartets on the left and one 21 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
53 of the tetraploid quartets on the right. In both ploidies the standing variation origin scenario 1 exhibits the highest likelihood. 2 Figure 5. Different evolutionary sources and number of protein–protein interactions of 3 candidate genes associated with substrate adaptation in diploid and tetraploid A. arenosa 4 populations. a) Different representation of the three major evolutionary sources of repeated 5 adaptation quantified by DMC in each ploidy. b) Density distribution of the number of 6 protein–protein interactions among the top candidate genes per each ploidy, compared to 7 the full dataset of A. thaliana genes. The vertical lines show the position of the mean of each 8 group. c) Protein–protein interaction networks of the top candidate genes demonstrating 9 overall greater connectivity (red/violet shading) of genes identified in tetraploids (right). 10 Table 1. Description of the six top candidate genes shared between ploidies. 11 A. lyrata ID A. thaliana ID Other name TAIR functional description (Berardini et al. 2015). AL1G18450 AT1G08090 NRT2.1 High-affinity nitrate transporter. Up-regulated by nitrate. Functions as a repressor of lateral root initiation independently of nitrate uptake. AL1G18460 AT1G08100 NRT2.2 Encodes a high-affinity nitrate transporter. AL6G42430 AT4G08620 SULTR1;1 Encodes a high-affinity sulfate transporter. Expressed in roots and guard cells. Upregulated by sulfur deficiency. AL8G13770 AT5G45050 WRKY16 Encodes a member of the WRKY Transcription Factor (Group II-e) family. ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
54 AL8G16980 AT5G43380 TOPP6 Encodes a type I serine/threonine protein phosphatase. AL8G25530 AT5G51270 PUB53 Involved in protein ubiquitination. 1 2 Figure 1 160x210 mm ( x DPI) ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
55 Figure 2 165x209 mm ( x DPI) ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
56 Figure 3 161x110 mm ( x DPI) 1 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025
57 Figure 4 160x170 mm ( x DPI) 1 ACCEPTED MANUSCRIPT Downloaded from https://academic.oup.com/mbe/advance-article/doi/10.1093/molbev/msaf298/8345034 by Department of Plant Physiology, Faculty of Science, Charles University user on 05 December 2025