RESEARCH ARTICLE Pairing statistics and melting of random DNA oligomers: Finding your partner in superdiverse environments Simone Di Leo 1☯ , Stefano MarniID 1☯ , Carlos A. PlataID 2☯ , Tommaso P. FracciaID 3 , Gregory P. SmithID 4 , Amos Maritan 5 , Samir SuweisID 5 , Tommaso BelliniID 1 * 1Dipartimento di Biotecnologie Mediche e Medicina Traslazionale, Universitàdegli Studi di Milano, Milano, Italy, 2Fı ´sica Teo ´rica, Universidad de Sevilla, Sevilla, Spain, 3Institut Pierre-Gilles de Gennes, CBI UMR 8231, ESPCI Paris, Universite ´PSL, CNRS, Paris, France, 4Department of Physics and Soft Materials Research Center, University of Colorado, Boulder, Colorado, United States, 5Dipartimento di Fisica ‘G. Galilei’, INFN, Universitàdi Padova, Padova, Italy ☯These authors contributed equally to this work. *
[email protected] Abstract Understanding of the pairing statistics in solutions populated by a large number of distinct solute species with mutual interactions is a challenging topic, relevant in modeling the complexity of real biological systems. Here we describe, both experimentally and theoretically, the formation of duplexes in a solution of random-sequence DNA (rsDNA) oligomers of length L= 8, 12, 20 nucleotides. rsDNA solutions are formed by 4 L distinct molecular species, leading to a variety of pairing motifs that depend on sequence complementarity and range from strongly bound, fully paired defectless helices to weakly interacting mismatched duplexes. Experiments and theory coherently combine revealing a hybridization statistics characterized by a prevalence of partially defected duplexes, with a distribution of type and number of pairing errors that depends on temperature. We find that despite the enormous multitude of inter-strand interactions, defectless duplexes are formed, involving a fraction up to 15% of the rsDNA chains at the lowest temperatures. Experiments and theory are limited here to equilibrium conditions. Author summary Several biological processes require that specific partner molecules succeed in binding after negotiating their way through a huge number of interactions with other molecules. How such molecular recognition emerges among millions distinct molecular species is an open problem. We have studied, both experimentally and theoretically, such process of “molecular recognition” in pools of highly diverse random DNA oligomers, which binds preferentially, but not exclusively, to its perfect complementary sequence. We find a complex behavior, in which some perfect pairing takes place with a non-trivial temperature dependence that we understand thorough statistical mechanics modelling. The pairing pattern of short random DNA is relevant in the context of the origin of life since the soPLOS COMPUTATIONAL BIOLOGY PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1010051 April 11, 2022 1 / 20 a1111111111 a1111111111 a1111111111 a1111111111 a1111111111 OPEN ACCESS Citation: Di Leo S, Marni S, Plata CA, Fraccia TP, Smith GP, Maritan A, et al. (2022) Pairing statistics and melting of random DNA oligomers: Finding your partner in superdiverse environments. PLoS Comput Biol 18(4): e1010051. https://doi.org/ 10.1371/journal.pcbi.1010051 Editor: Eugene I. Shakhnovich, Harvard University, UNITED STATES Received: January 19, 2022 Accepted: March 22, 2022 Published: April 11, 2022 Peer Review History: PLOS recognizes the benefits of transparency in the peer review process; therefore, we enable the publication of all of the content of peer review and author responses alongside final, published articles. The editorial history of this article is available here: https://doi.org/10.1371/journal.pcbi.1010051 Copyright: ©2022 Di Leo et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Data Availability Statement: All relevant data are within the manuscript and its Supporting information files.
called “RNA World” was most probably based on the mutual recognition of random chains. Introduction One of the defining features of biomolecules is the specificity and selectivity of their mutual interactions. Selectivity is at the heart of virtually all biological processes, including cell signalling, immune response, genetic transmission and regulation of gene expression. These processes are based on the presence of partner molecules that succeed in finding and docking to each other after negotiating their way through a huge number of collisions and interactions with other molecules, some of which exert attraction [1]. Indeed, in all actual cases, specific biomolecular interactions take place in “superdiverse” environments, i.e. in contexts of enormous variety of molecular species. Succeeding in pairing to the target is thus depending not only on the binding strength between partner molecules, but on the whole network of pair interactions between concurring molecular species, possible cooperativity, concentration and degeneracy. While the complexity of interactions in the contexts of biomolecular crowding is universally acknowledged [2,3], the thermodynamics and statistical physics of large pools of distinct interacting molecules have been discussed only in the frame of phase transitions [4,5], and molecular models of such systems have not been presented yet. DNA oligonucleotides mixtures are one of the pre-eminent systems to experimentally recreate the conditions described above in a controlled way, given the natural selectivity of base pairing [6] and the possibility to artificially synthesize large ensembles of distinct DNA sequences with controlled distributions [7]. Herein, we consider systems formed by aqueous solutions of DNA oligonucleotides of length Lin which the four bases (Cytosine, C; Guanine, G; Adenine, A; Thymine, T) are present with equal probability in each position in the sequence as sketched in Fig 1a. Thus, in each of these random-sequence DNA oligomers (rsDNA) solutions, 4 L distinct molecules can be found with approximately the same probability. We consider solutions of rsDNA oligomers with L= 8, 12 and 20 (8N, 12Nand 20N), corresponding to mixtures of �7�10 4 , 2 �10 7 and 10 12 distinct sequences, respectively; in these systems we study interaction selectivity, defect distribution and equilibration properties as a function of the temperature. Because of its relevance, the hybridization of nucleic acids has been a topic of continuous investigation since their identification, which led to well-established tools and models to calculate free energy, melting temperature and secondary structure directly from the DNA (or RNA) sequences involved [8–11]. In particular, the thermodynamics of duplex formation is commonly evaluated in the frame of the so-called “Nearest Neighbor” (NN) model, in which the hybridization free energy is obtained by the summation of elemental contributions. These are extracted from database of melting temperatures of DNA oligomers and effectively take into account both WC and non-canonical pairings [8]. However, available models are accurate only in the case of pools of low numbers of distinct sequences [12], and do not have the capability of predicting the behavior of complex systems such as the rsDNA solutions considered here. In a rsDNA solution, upon colliding, pairs of rsDNA molecules bind to each other with a strength and resultant stability that is mainly determined by the level of complementarity of the base sequences according to the Watson-Crick (WC) pairing rules. This results in an interaction matrix, sketched in Fig 1b, where each position represents the most stable duplex conformation for a given pair. The most probable outcome of a random collision between two PLOS COMPUTATIONAL BIOLOGY Pairing statistics and melting of random DNA oligomers PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1010051 April 11, 2022 2 / 20 Funding: S.D.L., S.M. and T.B. acknowledge support from MIUR-PRIN (Grant No. 2017Z55KCW). T.P.F. acknowledges IPGG Laboratoire d’Excellence, “Investissement d’avenir" program ANR-10-IDEX-0001-02 PSL, ANR-10LABX-31 and ANR-10-EQPX-34. S.S. acknowledges UNIPD grant BIRD209912. C.A.P. acknowledges the support from PGC2018-093998B-I00 funded by FEDER/Ministerio de Ciencia e Innovaciòn-Agencia Estatal de Investigaciòn (Spain) and from the program PAIDI-DOCTOR by Junta de Andalucı `a and European Social Fund. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Competing interests: The authors have declared that no competing interests exist.
rsDNA molecules is a quite unstable pair, with a small number of consequent paired bases (3 in case of L = 12, as in the bottom sketch of Fig 1c). However, even if less probable, more stable structures are formed, ranging from full complementarity, a condition that yields defectless helical duplexes (Fig 1c, top sketch), to duplexes with pairing mismatches of various types, listed in Fig 1c. To each duplex, with any pattern of defects, corresponds a binding free energy that can be computed from the sequences involved with standard tools. At the opposite end of the spectrum, there are rsDNA strands with no complementarity at all, as between an oligo made of only Ts and one made of only Cs. In rsDNA at fixed L, the variety of binding possibilities increases with the number of allowed mismatches, while the binding energy decreases. As shown here, these two factors nearly compensate, leading to non-trivial competition between interaction strength and degeneracy. In a previous study of rsDNA [13], it was observed that, when L>12, rsDNA solutions self-organize into liquid crystal phases. Given the mechanism by which these phases are formed in solutions of DNA oligomers [14,15], this finding suggests that hybridization of rsDNA leads to duplexes with fairly well-paired terminals. This feature of rsDNA solutions remained speculative, with no experimental or statistical support. Besides its value as a platform for exploring hybridization in crowded nucleic acid environments and for describing the network of interactions in superdiverse mixtures, the study of rsDNA is also relevant to the evaluation of scenarios for the origin of life. Indeed, if the RNA world hypothesis is correct, such a state had to be anticipated by a condition in which RNA Fig 1. Description of the system. (a): Solutions of random-sequence DNA (rsDNA) oligomers of length L are mixtures made of 4 L distinct molecules, obtained by all the combinations of the four nucleobases, which are present at any position in the sequence with equal probability. (b): Each rsDNA oligomer can interact with 4 L different rsDNA oligomers, leading to a 4 L ×4 L interaction matrix. Each dot in the matrix represents the most energetically favorable pairing between the two selected rsDNA oligomers, among all the possible mutual shifts. (c) Each position in the interaction matrix corresponds to a specific duplex motif, characterized by pairing errors which are here described by the parameters in α. The most probable duplex in the matrix is highly defected, as the last example in the panel. https://doi.org/10.1371/journal.pcbi.1010051.g001 PLOS COMPUTATIONAL BIOLOGY Pairing statistics and melting of random DNA oligomers PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1010051 April 11, 2022 3 / 20
oligomers were abiotically synthesized with a large degree of randomness, from which ribozyme sequences could have been subsequently selected [16,17]. Whether such molecular mixtures could form WC pairs or whether hybridization was instead prevented by the large variety of species is an information that can shape the RNA world model itself, clarifying the role of complementarity and duplex formation in the prebiotic environment. [18]. In this paper, we study rsDNA solutions by a combination of three complementary strategies: (i) measurement of the overall degree of hybridization by UV absorbance as a function of temperature (T); (ii) measurements of the degree of hybridization and thermal stability of pairs of mutually complementary sequences mixed with rsDNA by fluorescence ContactQuenching (CQ); (iii) development of a theoretical framework, based on a re-parametrization of the NN thermodynamic parameters, enabling quantitative predictions on rsDNA hybridization and pairing error statistics. Materials and methods Description of the system 8N, 12N and 20N rsDNA oligomers were synthesized on solid phase by an A ¨kta Oligopilot. The products were purified via dialysis against a 25mM NaCl solution and lyophilized. The samples were characterized by HPLC (see S2 Text). An analogous previous synthesis was characterized by MALDI-TOF mass spectroscopy [13]. Stock aqueous solutions were prepared at c rsDNA �50g/l, from which the final samples concentrations were obtained by dilution: ranging from c rsDNA = 0.04g/lto c rsDNA = 25g/land with ionic strengths of c NaCl = 0.15M,c NaCl = 0.45M and c NaCl = 1.0M. The statistics of the pairing quality have been explored by mixing rsDNA with a tagged pair of complementary DNA strands. Specifically, we have used the 8-base-long couple 8A�and 8B�and the 12-base-long couple 12A�and 12B�, modified by 5’-Texas Red (A�oligomers) and 3’-6-FAM (Fluorescein) moieties (B�oligomers), respectively, so that upon hybridization the two fluorophores come in contact (see Fig G in S2 Text). Specific sequences are as follows. 8A�: TexasRed-5’-ACAGTCCT-3’. 8B�: 5’-AGGACTGT-3’-FAM. 12A�: TexasRed-5’-ACGA CAGTCCTG-3’. 12B�: 5’-CAGGACTGTCGT-3’-FAM. 8A�, 8B�, 12A�and 12B�were purchased from IDT. Using UV hyperchromicity to detect ensemble rsDNA melting The overall degree of hybridization in rsDNA was evaluated by measuring the absorbance Aat the wavelength λ= 258nm.Ais obtained by averaging over an interval Δλ = 3nm. Experiments were performed with the Evolution 300 UV-Vis spectrophotometer from Thermo Scientific customized with a Quantum Northwest peltier hot/cold stage with hold temperature accuracy of ±0.05˚C. Experiments have been performed with 1 o C/min heating and cooling rate. The cell holder was capable of hosting two different types of cell: standard quartz cuvettes with optical path length ℓ= 1cm, adequate to investigate DNA solutions with c�0.02 −0.04g/l; microfluidic cells with ℓ= 10μm from Starna Scientific Ltd to investigate low volumes (<50μl) of more concentrated DNA solutions (c�20 −35g/l) (see S2 Text). Absorbance data as function of temperature, A(T), were treated according to standard protocols [19] (see S2 Text) to extract the “ensemble” melting curve θ e (T), i.e. the fraction of rsDNA oligomers forming duplexes (of any quality) at temperature T. The melting temperature T m is defined by θ e (T m ) = 1/2. PLOS COMPUTATIONAL BIOLOGY Pairing statistics and melting of random DNA oligomers PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1010051 April 11, 2022 4 / 20
Contact-quenching detects the pairing of specific sequences Fluorescence-based measurements were used to detect how frequently a specific sequence is able to find its exact complementary strand in the midst of the rsDNA solution. This was done by mixing 8A�and 8B�, in equal amount in a 8N solution, and similarly with 12A�and 12B�in 12N. Although the pair of fluorophores Texas Red and FAM were originally chosen to obtain FRET signal, we found that the dominant effect signaling their interaction is the so called “Contact Quenching” (CQ), i.e. the drop in fluorescence quantum yield of both fluorophores when the two fluorophores are in close proximity [20] (see S2 Text). In the case considered here, the quenching is deep (about 80% reduction for Texas Red and 50% for FAM) and can be easily exploited to extract the fraction θ AB of A�oligomers that forms a defectless duplex with its complementary partner B�. To extract θ AB (T) we monitored the quenching of the Texas Red emission, which is known to have a small T dependence [21]. Fluorescence emission vs. T was measured using the Applied Biosystems QuantStudio 5, a Real-Time PCR Instrument by Thermo Fisher Scientific. Calibration and normalization procedures to extract θ AB from raw data are provided in S2 Text. They also include a thermodynamic characterization of the 8A�-8B�and 12A�-12B� duplexes, since their binding free energy ΔG A � B � is slightly modified with respect to their untagged analogs because of the stabilizing effect of the fluorophores at the terminals. CQ experiments were performed by mixing fixed concentrations of both A�and B�,c fluo = 100nM, with rsDNA solutions prepared at 0.04g/l<c rsDNA <25g/l, and ionic strengths c NaCl = 0.15M and c NaCl = 1.0M. Measurements were performed in 15mM TRIS HCl at a pH of 7.4, to minimize fluorescence drift due to pH sensitivity of FAM [21]. In rsDNA solutions, each specific sequence is present at a concentration c rsDNA /4 L . The stoichiometric ratio of each added fluorescent sequence (A�or B�) and the same non-fluorescent sequence already present in the solution of rsDNA is: �¼cfluo crsDNA=4L:ð1Þ When ϕ= 1, i.e. the amount of fluorescently labeled sequence equals that of the same sequence without tag, θ AB offer an approximate evaluation of the degree in which errorless pairing is present within rsDNA solutions. At the same time, the measurement of the ϕdependence of θ AB (ϕ) enables a detailed comparison with theoretical predictions. Experimental results Ensemble melting of rsDNA Hyperchromicity in ultraviolet enables accessing the overall degree of hybridization in rsDNA solutions. Fig 2, green symbols, shows θ e (T) for 12N, c rsDNA = 0.04g/l,c NaCl = 1M, which we compare, as reference, with the melting curves of binary mixtures of complementary DNA 12mers computed, with standard approaches, at two concentrations: c DNA = 0.02g/l(blue dashed line) and c DNA = 0.04/4 12 g/l(red dashed line), both at c NaCl = 1M. As visible, θ e exhibits a behavior intermediate between the two. rsDNA duplexes are way more unstable (of �30˚C) than duplexes of complementary strands at equal total concentration, a clear manifestation of the selectivity of rsDNA pairing. At the same time, rsDNA duplexes appears about 5˚Cmore thermally stable than the 12mer complementary duplexes when solubilized at the same concentration at which they are present in the rsDNA solution, an indication that the formation of defected pairing is a relevant feature of the hybridization of rsDNA, as also suggested by the milder slope of θ e (T). PLOS COMPUTATIONAL BIOLOGY Pairing statistics and melting of random DNA oligomers PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1010051 April 11, 2022 5 / 20
In Fig 3a,θ e (T) measured for 12N and 20N at c rsDNA �0.04g/lare shown for three different salt concentrations c NaCl . Since T m decreases with Lbut increases with c rsDNA , to obtain reliable θ e (T) for 8N, we performed melting experiments at a larger concentration, c rsDNA �25g/l, shown in Fig 3b. We find in all conditions θ e (T) to depend on Tmore mildly than in typical melting curves in binary solutions of complementary strands. Fig 3 also shows that T m of rsDNA grows with c NaCl , with Land with c rsDNA , as it appears by comparing panels (a) and (b), in agreement with DNA melting in less complex systems [8,22]. Experimental data shown here were taken at equilibrium. Attaining this condition is not trivial, since the lifetime of DNA duplexes dramatically depends on the length of the oligomers and on the concentration of the solutions [23,24]. To approach equilibrium of 20N in dilute conditions, we considered only θ e (T) measured upon heating after a long equilibration time at low T. Similar attention had to be given to the behavior of concentrated 12N solutions, as mentioned in the next Section. Further information on equilibrium conditions is provided in the S1 Text. The non-equilibrium behavior of rsDNA will be the topic of a future work. Probability of defectless duplexes The study of θ e (T) hints at hybridization in rsDNA as a combination of selectivity with a certain degree of defects in the duplex formation. However, θ e (T) does not offer much insight on how probable it is to find duplexes with a certain pairing quality.Aiming at this kind of Fig 2. Double strand vs. random sequence DNA melting. Ensemble melting curve as a function of temperature. Green dots: measured ensemble melting of 12N at c rsDNA = 0.04 g/l. Shading marks experimental uncertainty resulting from the average over 8 experimental replicas. Dashed lines: theoretical melting predicted for equimolar solutions of two complementary 12mers at c DNA = 0.02 g/l each (dashed blue line) and c DNA = 0.04/4 12 �2.4 10 −9 g/l each (dashed red line). 12N θ e (T) exhibits a behavior intermediate between the two. Dashed lines are obtained by averaging many melting curves of complementary 12mers. c NaCl = 1M in all curves. https://doi.org/10.1371/journal.pcbi.1010051.g002 PLOS COMPUTATIONAL BIOLOGY Pairing statistics and melting of random DNA oligomers PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1010051 April 11, 2022 6 / 20
Fig 3. Ensemble melting of rsDNA. Measured ensemble melting curves of rsDNA at c rsDNA = 0.04 g/L for 12N (panel a, open circles), 20N (panel a, full diamonds) and 8N at c rsDNA = 25 g/L (panel b, open squares), at various salt concentrations: c NaCl = 0.15M (blue), c NaCl = 0.45M (green) and c NaCl = 1M (red). Shading marks experimental uncertainty resulting from the average over 6–8 experimental replicas for 8N and 12N, whereas, for 20N, just one experiment is shown as described in the text. Dashed lines, with same color code, are the theoretical predictions of Eq (7). https://doi.org/10.1371/journal.pcbi.1010051.g003 PLOS COMPUTATIONAL BIOLOGY Pairing statistics and melting of random DNA oligomers PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1010051 April 11, 2022 7 / 20
information, CQ experiments in solutions of 8N and 12N were performed. We did not perform analogous measurement for the 20N both because of the artifacts due to non-equilibrium pairing and because of the small accessible range of ϕ(4 20 is a large number!). Fig 4 shows the fraction of hybridized 8A�and 8B�in 8N, θ A � B � (T), for various ϕ. In the case of 8N we could reach ϕ= 1, corresponding to c rsDNA = 16g/l. At ϕ= 1 (red dots) the fraction of paired A�B�reaches about 30% at low T. As ϕincreases, the fraction of A�B�duplexes increases, as expected. The dependence of θ A � B � (T) on ϕis a useful tool to test our theoretical model, as discussed below. Similar experiments, shown in S1 Text, have been carried on for 12A�and 12B�in 12N, where however in some conditions (largest c rsDNA ) we cannot reach equilibrium. Theoretical framework Although the combination of the measured θ e (T) and θ A � B � (T) offers important insight on the quality of pairing within the rsDNA system, a deeper understanding of the driving mechanisms in this superdiverse environment requires a statistical model able to take into account the balance between binding energy and degeneracy. Indeed, the strongest binding energy is achieved in defectless duplexes, which are only formed with a fraction 1/4 L of the total number of sequences. On the contrary, defected duplexes are more weakly bound, but they can be assembled with a larger variety of sequences, i.e. the weaker the binding energy, generally, the larger its degeneracy. Fig 4. Probability of defectless duplexes in rsDNA. θ A � B � Fraction of paired 8A�8B�in 8N measured with CQ experiments, at c NaCl = 0.15M. Colors correspond to different values of the stoichiometric ratio ϕbetween the probes A�/B�and rsDNA strands. Black data is the melting curve of a neat A�B�solution (without rsDNA). In all experiments c fluo = 100nM. Each data point is obtained as an average over 5–10 replications of the experiment. The corresponding standard deviation is reported as shaded regions. https://doi.org/10.1371/journal.pcbi.1010051.g004 PLOS COMPUTATIONAL BIOLOGY Pairing statistics and melting of random DNA oligomers PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1010051 April 11, 2022 8 / 20
The hybridization in simple systems formed by a limited number of sequences is well described by the current thermodynamic approaches, such as the NN model. Despite their accuracy, there is yet no theoretical frame to apply this knowledge to systems with high complexity such as the rsDNA, where more than 4 L ×4 L interactions are involved. Explicitly computing the free energy for all pairs involved would be too computationally expensive. We thus develop a mean-field-type theoretical approach that uses averages of the thermodynamic parameters of the NN model and their re-parametrization on a simple counting of defects. With this approach we can compute θ e (T) and θ CQ (T) at equilibrium with no free parameters. Hence, a direct comparison between theory and experiments is achieved. Ensemble melting of rsDNA We consider a rsDNA mixture containing Nchains in a volume V(with a total concentration c=N/V). We assume the mixture to be perfectly balanced (see S2 Text), that is, eachsequence is present through N/4 L copies. We also assume on-off hybridization, with nointermediate state between unbound and paired, which is justified given the limited length (L�20) of the oligomers here considered [11]. Thus, a given couple (iand j) interact with a set of 2L−1 distinct binding free energies DGðasÞ ij , depending on their mutual alignment: we use the shift parameter α s with −(L−1) �α s �L−1 to express such alignment. Specifically, α s = 0 stands for the condition of perfect alignment, i.e. the 3’ terminal base of strand iis aligned with the 5’ terminal base of strand jand vice versa. Positive or negative α s indicate an overhang on the 3’ or 5’ terminal, respectively. We define the Boltzmann factor zðasÞ ij � ½c�expð bDGðasÞ ij Þ;ð2Þ where the brackets around the concentration denotes that it is measured in mol/L ([c] = c/(mol/L)), and β= (k B T) −1 with k B being the Boltzmann constant. In real mixtures, some duplexes are very unlikely, with large ΔGand thus small z. In this setting, we can formally define, using the canonical distribution, the probability of any given hybridization state and the related partition function, which contains all the relevant statistical information of the system (see S3 Text). However, the partition function of the rsDNA system in the thermodynamic limit (N!1) is not in a closed form. Consequently, an exact derivation of the fraction of paired oligomers θ e in such a limit is not simple to obtain. This is thoroughly discussed in the S3 Text, where two opposite limits (high and low temperatures) are carried out allowing the derivation of analytical expressions to bypass this obstacle. Furthermore, an Unification Ansatz that matches the two approximations in their range of validity has been worked out. This theoretical approach allows us to compute yðiÞ eðTÞ, the approximated melting curve for the oligomers with sequence iin the midst of all the 4 L species in the rsDNA solution (see S3 Text for details on the derivation). yðiÞ e¼12 1þffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi 1þ4 4LX4L j¼1XL1 as¼ ðL1ÞzðasÞ ij r:ð3Þ This analytical expression has the same structure of the melting curve as a solution of selfcomplementary sequences (Eq 18a in Ref. [25]), with two relevant differences: the factor 1/4 L normalizing the concentration (the concentration of any specific sequence is c/4 L ), and the double summation of the pairing weight of i: for all possible partner sequence j, and for all the possible shifts. PLOS COMPUTATIONAL BIOLOGY Pairing statistics and melting of random DNA oligomers PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1010051 April 11, 2022 9 / 20
particularly adequate. We would also like to point out that the model validity is limited to equilibrium conditions, as we documented above and in the S1 Text. Also, the model assumes unlimited molecular availability, and does not include the effects of competitive binding which could arise from constrained stoichiometric ratios in limited pools of molecules. The successful comparison with experiments indicates that rsDNA with short enough chain length satisfies all these requirements. In systems with longer molecules, out-of-equilibrium conditions and hybridization states with more than one helical region could become relevant [11]. The agreement with observations also validates the predicted pairing distributions θ α (T) and θ |α| (T) shown in Fig 6, which are worth discussing further. These distributions indicate that, in a superdiverse environment of random sequence oligonucleotides, the selectivity afforded by the free energy of base-pairing is “marginal”. Specifically, the resulting pairing is good, but not perfect, with the majority of sequences being defected, but with less than two pairing errors. Defectless pairing involves at most a fraction of �14% of the rsDNA strands. The pairing statistics of rsDNA largely depends on the compensation between two opposing factors: (i) the degeneracy, g(L,α), which grows in a nearly exponential way with the total number of pairing errors (see Eq (5)), thus favoring the formation of defected duplexes; and (ii) the binding free energy, ΔG, which increases approximately linearly with the number of defects (see S3 Text), yielding a Boltzmann factor with an exponential advantage to duplexes with less defects (see Eq (2)). In DNA duplexes formation, these two factors nearly balance, with a partial dominance of the energetic component, a condition leading to the smooth αand Tdependence of the probability distributions. The result would change if the degeneracy Fig 9. Probability of defectless duplexes in rsDNA: Salt dependence. Fraction of paired 8A�8B�as a function of ϕ, expressing their dilution in 8N, at T= 15˚Cfor c NaCl = 0.15M(blue dots) and c NaCl = 1.0M(red dots). Dashed lines: theoretical predictions, with the shaded regions obtained from the experimental uncertainty on the pairing energy between A�B�, (see S2 Text). https://doi.org/10.1371/journal.pcbi.1010051.g009 PLOS COMPUTATIONAL BIOLOGY Pairing statistics and melting of random DNA oligomers PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1010051 April 11, 2022 16 / 20
grows faster or slower than an exponential with the binding energy. An example of this latter condition is given by the selectivity of PCR primers within the genome, in which the set of competing bindings is limited with respect to the random situation, thus enabling a strong dominance of defectless primer binding. An obvious question is how critical is this marginal condition, and whether modifications in the nucleobase structure, and thus of pairing and stacking energy, could significantly improve selectivity. To answer this question we computed θ α (T) by assuming that all pairing energies equal that of CG (which is roughly double of that of AT). We find a significant, but not dramatic, increment in θ 0 , that at low Treaches 0.3 (see Fig D in S1 Text), indicating that the basic features of θ α (T) are stable within the range of energies involved in natural and artificial nucleobase binding [7]. As a further test of the pairing efficiency in random systems, we computed θ 0 (T) by considering random systems formed by only 2 bases, instead of the four natural ones, using energetic parameters intermediate between AT and CG (see Fig E in S1 Text). Even in this condition, where the degeneracy of the defected duplexes is strongly reduced, the fraction of perfect pairs is larger but still well under 0.4. When, on the contrary, the number of bases are increased (still assuming a WC-type pairing role), θ 0 markedly decreases. These observations strengthen the notion that, in the range of pairing energies of nucleic acids, the presence of randomness appears to be the dominant factor in determining the quality of the pairs. These observations sets a reference for the selectivity in contexts of strong randomness and heterogeneity of sequences such as those that have likely characterized the origin of life and RNA world. Whatever the mechanisms of chain amplification and lengthening, and whatever base pair variants were at the time available, they could not have relied on levels of selectivity much better than those reported here. We previously proposed that one of such mechanisms leading to the formation of long nucleic acid chains could exploit the symmetry breaking and formation of molecular column due to liquid crystal ordering [27,28]. rsDNA can indeed form, in given conditions, columnar liquid crystals [13]. How the distribution of pairing quality here discussed can be compatible with liquid crystal formation appears as a subtle matter that will be the topic of a forthcoming work. Conclusion We introduced rsDNA solutions as a model system of superdiverse mixture, enabling the study of interactions and pair formations in the midst of a huge amount of competing molecular species, a condition offering a conceptual paradigm for the molecular variety and selectivity of biological environments. In the analysis of rsDNA solutions, we could take advantage of the limited polymer heterogeneity given by the four nucleobases, of the rather simple and highly characterized pairing rule, of the availability of solid-state synthesis and of the variety of experimental tools. This combination of factors enabled us to experimentally characterize, and theoretically describe, the selectivity of pairing, and found that the majority of rsDNA duplexes contain pairing errors, but limited to one or two per duplex, a condition that still grants a reliable stability to the structures. Based on the success in the description of rsDNA, we applied our approach to the description of the selective pairing in the context of the PCR technology and miRNA based gene regulation, which is the topic of a forthcoming publication. Extending this statistical approach to biomolecular superdiverse systems closer to cell environments, characterized by a less dramatic diversity but by more complex and less defined interactions, will be the challenging development of this work. PLOS COMPUTATIONAL BIOLOGY Pairing statistics and melting of random DNA oligomers PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1010051 April 11, 2022 17 / 20
Supporting information S1 Text. Further results. Evidence of Out-of-Equilibrium Conditions; Perfect pairing probability with different energies and number of nucleobases types. Fig A: Evidence of out-ofequilibrium behavior for 12N. Tdependence of the fraction of duplexed strands θ e measured while heating and cooling at 1˚C/min.Fig B: Evidence of out-of-equilibrium behavior for 20N. Absorbance vs. Tmeasured upon heating after one month equilibration at 4˚C.Fig C: θ A � B � , fraction of paired 12A�−12B�in 12N, determined from the model and via CQ experiments with different cooling rates. Fig D: yðfCGÞ 0computed with different values of f CG in 8N. Fig E: yðnbÞ 0computed with different values of n b in 12N. (PDF) S2 Text. Materials and methods. Characterization of rsDNA synthesis; Measurement of rsDNA concentration; UV Absorbance: Experimental Setup; Analysis of UV Absorbance Data; Characterization of A�and B�Fluorescence; Contact-Quenching Data Analysis; Free Energy of A�B�Duplex. Fig A: HPLC traces of 12N compared with two different 12mers. HPLC traces of 8N, 12N and 20N. Fig B: Quartz microfluidic. Quantum Northwest Peltier. Fig C: Temperature calibration of Tmeasured by a thermistor in contact with the microfluidic cell vs. the internal control T peltier .Fig D: Steps in the analysis of absorbance data A(T) of 12N used to extract the melting curves θ e .Fig E: Fluorescence Emission Spectra of labeled DNA systems: A�, B�, A�B�, A�B and AB�.Fig F: Absorbance Spectra of labeled DNA systems: A�, B�, A�B�, A�B and AB�.Fig G: Simplified representation of the 12A�B�duplex, showing the relative size of DNA duplex, linker and fluorescent moieties FAM and TexasRed. Fig H: fluorescence intensity of TexasRed vs. Tin a solution of 12A�+ 12B�.Fig I: Normalized fluorescence intensity of TexasRed in a solution of 12A�, 12B�and 12B. Fig J: Average normalized fluorescence of 8A�+8B�. The fit enables determining the linear drift at low temperatures, corresponding to the signal of fully paired A�B�.Fig K: T m vs A�B�concentration measured by CQ for the following systems. (PDF) S3 Text. Comprehensive Description of the theoretical model. Partition Function; Melting Curve Approximations for rsDNA solution; Energetic Parametrization based on α; Pairing Statistics with α i = 0; Effects of the Ionic Strength on the Pairing Statistics.Fig A: Melting Curves of rsDNA, according to the Low T Approximation,High T Approximation and Unification Ansatz. (PDF) Acknowledgments We thank N.A. Clark for useful discussions and E. Toffolo for her precious guidance in using PCR equipment for CQ measurements. Author Contributions Conceptualization: Tommaso P. Fraccia, Amos Maritan, Samir Suweis, Tommaso Bellini. Formal analysis: Stefano Marni, Carlos A. Plata, Amos Maritan, Samir Suweis. Funding acquisition: Tommaso Bellini. Investigation: Simone Di Leo, Stefano Marni, Carlos A. Plata, Tommaso Bellini. PLOS COMPUTATIONAL BIOLOGY Pairing statistics and melting of random DNA oligomers PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1010051 April 11, 2022 18 / 20
Methodology: Simone Di Leo, Stefano Marni, Carlos A. Plata, Samir Suweis, Tommaso Bellini. Project administration: Samir Suweis, Tommaso Bellini. Resources: Gregory P. Smith, Tommaso Bellini. Software: Stefano Marni, Carlos A. Plata. Supervision: Tommaso Bellini. Validation: Simone Di Leo, Stefano Marni, Carlos A. Plata, Tommaso P. Fraccia, Amos Maritan, Samir Suweis, Tommaso Bellini. Writing – original draft: Simone Di Leo, Stefano Marni, Carlos A. Plata, Samir Suweis, Tommaso Bellini. Writing – review & editing: Simone Di Leo, Stefano Marni, Carlos A. Plata, Tommaso P. Fraccia, Gregory P. Smith, Amos Maritan, Samir Suweis, Tommaso Bellini. References 1. Phillips R, Kondev J, Theriot J, Garcia HG. Physical biology of the cell. 2nd ed. New York: Garland Science; 2012. 2. Ellis RJ. Macromolecular crowding: an important but neglected aspect of the intracellular environment. Curr Opin Struct Biol. 2001; 11: 114–119. https://doi.org/10.1016/S0959-440X(00)00239-6 PMID: 11179900 3. Politou A, Temussi PA. Revisiting a dogma: the effect of volume exclusion in molecular crowding. Curr Opin Struct Biol. 2015; 30: 1–6. https://doi.org/10.1016/j.sbi.2014.10.005 PMID: 25464122 4. Jacobs WM, Frenkel D. Phase transitions in biological systems with many components. Biophys J. 2017; 112: 683–691. https://doi.org/10.1016/j.bpj.2016.10.043 PMID: 28256228 5. Harmon TS, Holehouse AS, Pappu RV. To mix, or to demix, that is the question. Biophys J. 2017; 112: 565–567. https://doi.org/10.1016/j.bpj.2016.12.031 PMID: 28256216 6. Calladine CR, Drew H, Luisi B, Travers A. Understanding DNA: The Molecule and it how Works. 2nd ed. San Diego, CA: Elsevier Academic Press; 2004. 7. Hoshika S, Leal NA, Kim MJ, Kim MS, Karalkar NB, Kim HJ, et al. Hachimoji DNA and RNA: A genetic system with eight building blocks. Science. 2019; 363: 884–887. https://doi.org/10.1126/science. aat0971 PMID: 30792304 8. SantaLucia J Jr, Hicks D. The thermodynamics of DNA structural motifs. Annu Rev Biophys Biomol Struct. 2004; 33: 415–440. https://doi.org/10.1146/annurev.biophys.32.110601.141800 PMID: 15139820 9. Zadeh JN, Steenberg CD, Bois JS, Wolfe BR, Pierce MB, Khan AR, et al. NUPACK: analysis and design of nucleic acid systems. J Comput Chem. 2011; 32: 170–173. https://doi.org/10.1002/jcc.21596 PMID: 20645303 10. Owczarzy R, Vallone PM, Gallo FJ, Paner TM, Lane MJ, Benight AS. Predicting sequence-dependent melting stability of short duplex DNA oligomers. Biopolymers. 1997; 44: 217–239. https://doi.org/10. 1002/(SICI)1097-0282(1997)44:3%3C217::AID-BIP3%3E3.0.CO;2-Y PMID: 9591477 11. Vologodskii A, Frank-Kamenetskii MD. DNA melting and energetics of the double helix. Phys Life Rev. 2018; 25: 1–21. https://doi.org/10.1016/j.plrev.2017.11.012 PMID: 29170011 12. Wolfe BR, Porubsky NJ, Zadeh JN,Dirks RM, Pierce NA. Constrained multistate sequence design for nucleic acid reaction pathway engineering. J Am Chem Soc. 2017; 139: 3134–3144. https://doi.org/10. 1021/jacs.6b12693 PMID: 28191938 13. Bellini T, Zanchetta G, Fraccia TP, Cerbino R, Tsai E, Smith GP, et al. Liquid crystal self-assembly of random-sequence DNA oligomers. Proc Natl Acad Sci. 2012; 109: 1110–1115. https://doi.org/10.1073/ pnas.1117463109 PMID: 22233803 14. Nakata M, Zanchetta G, Chapman BD, Jones CD, Cross JO, Pindak R, et al. End-to-end stacking and liquid crystal condensation of 6–to 20–base pair DNA duplexes. Science. 2007; 318: 1276–1279. https://doi.org/10.1126/science.1143826 PMID: 18033877 PLOS COMPUTATIONAL BIOLOGY Pairing statistics and melting of random DNA oligomers PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1010051 April 11, 2022 19 / 20
15. Fraccia TP, Smith GP, Bethge L, Zanchetta G, Nava G, Klussmann S, et al. Liquid crystal ordering and isotropic gelation in solutions of four-base-long DNA oligomers. ACS Nano. 2016; 10: 8508–8516. https://doi.org/10.1021/acsnano.6b03622 PMID: 27571250 16. Bartel DP, Szostak JW. Isolation of new ribozymes from a large pool of random sequences. Science. 1993; 261: 1411–1418. https://doi.org/10.1126/science.7690155 PMID: 7690155 17. Ekland EH, Szostak JW, Bartel DP. Structurally complex and highly active RNA ligases derived from random RNA sequences. Science. 1995; 269: 364–370. https://doi.org/10.1126/science.7618102 PMID: 7618102 18. Kudella PW, Tkachenko AV, Salditt A,Maslov S, Braun D. Structured sequences emerge from random pool when replicated by templated ligation. Proc Natl Acad Sci. 2021; 118: e2018830118. https://doi. org/10.1073/pnas.2018830118 PMID: 33593911 19. Owczarzy R. Melting temperatures of nucleic acids: discrepancies in analysis. Biophys Chem. 2005; 117: 207–215. https://doi.org/10.1016/j.bpc.2005.05.006 PMID: 15963627 20. Marras SAE, Kramer FR, Tyagi S. Efficiencies of fluorescence resonance energy transfer and contact– mediated quenching in oligonucleotide probes. Nucleic Acids Res. 2002; 30: e122. https://doi.org/10. 1093/nar/gnf121 PMID: 12409481 21. You Y, Tataurov AV, Owczarzy R. Measuring thermodynamic details of DNA hybridization using fluorescence. Biopolymers. 2011; 95: 472–486. https://doi.org/10.1002/bip.21615 PMID: 21384337 22. Owczarzy R, You Y, Moreira BG, Manthey JA, Huang L, Behlke MA, et al. Effects of sodium ions on DNA duplex oligomers: improved predictions of melting temperatures. Biochemistry. 2004; 43: 3537– 3554. https://doi.org/10.1021/bi034621r PMID: 15035624 23. Woodside MT, Behnke-Parks WM, Larizadeh K,Travers K, Herschlag D, Block SM. Nanomechanical measurements of the sequence-dependent folding landscapes of single nucleic acid hairpins. Proc Natl Acad Sci. 2006; 103: 6190–6195. https://doi.org/10.1073/pnas.0511048103 PMID: 16606839 24. Dupuis NF, Holmstrom ED, Nesbitt DJ. Single-molecule kinetics reveal cation-promoted DNA duplex formation through ordering of single-stranded helices. Biophys. 2013; 105: 756–766. https://doi.org/10. 1016/j.bpj.2013.05.061 PMID: 23931323 25. Plata CA, Marni S, Maritan A, Bellini T, Suweis S. Statistical physics of DNA hybridization. Phys Rev E. 2021; 103: 042503. https://doi.org/10.1103/PhysRevE.103.042503 PMID: 34005886 26. Ghosh S, Takahashi S, Endoh T, Tateishi-Karimata H, Hazra S, Sugimoto N. Validation of the nearestneighbor model for Watson–Crick self-complementary DNA duplexes in molecular crowding condition. Nucleic Acids Res. 2019; 47: 3284–3294. https://doi.org/10.1093/nar/gkz071 PMID: 30753582 27. Fraccia TP, Smith GP, Zanchetta G, Paraboschi E, Yi Y, Walba DM, et al. Abiotic ligation of DNA oligomers templated by their liquid crystal ordering. Nat Commun. 2015; 6: 6424. https://doi.org/10.1038/ ncomms8463 28. Todisco M, Fraccia TP, Zanchetta G, et al. Nonenzymatic polymerization into long linear RNA templated by liquid crystal self-assembly. ACS nano. 2018; 12: 9750–9762. https://doi.org/10.1021/acsnano. 8b05821 PMID: 30280566 PLOS COMPUTATIONAL BIOLOGY Pairing statistics and melting of random DNA oligomers PLOS Computational Biology | https://doi.org/10.1371/journal.pcbi.1010051 April 11, 2022 20 / 20