Characterization of Oligonucleotide Microarray Hybridization: Microarray Fabrication by Light-Directed in situ Synthesis – Development of an Automated DNA Microarray Synthesizer, Characterization of Single Base Mismatch Discrimination and the Position-Dependent Influence of Point Defects on Oligonucleotide Duplex Binding Affinities
Full text
Characterization of Oligonucleotide Microarray Hybridization Microarray Fabrication by Light-Directed in situ Synthesis – Development of an Automated DNA Microarray Synthesizer, Characterization of Single Base Mismatch Discrimination and the Position-Dependent Influence of Point Defects on Oligonucleotide Duplex Binding Affinities Von der Universit¨at Bayreuth zur Erlangung des Grades eines Doktors der Naturwissenschaften (Dr. rer. nat.) genehmigte Abhandlung von Thomas Naiser geboren in Bayreuth 1. Gutachter: Prof. Dr. Albrecht Ott 2. Gutachter: Prof. Dr. Josef K¨as 3. Gutachter: Prof. Dr. Thomas Fischer Tag der Einreichung: 14.12.2007 Tag des Kolloquiums: 04.07.2008
Abstract The present thesis focuses on nucleic acid hybridization between free-floating target sequences and complementary end-tethered oligonucleotide probes on the surface of DNA microarrays. Hybridization experiments were performed on oligonucleotide microarrays (DNA Chips) which were fabricated with an automated synthesis apparatus (developed in the framework of the present thesis). The working principle of the microarray synthesizer is based on a photochemically controlled in situ synthesis process [Fod91]. By means of the combinatorial approach up to 25000 different (arbitrary) probe sequences can be fabricated in parallel – starting from nucleotide building blocks (NPPOC-phosphoramidites [Has97]) – directly on the surface of the microarray. Great flexibility with regard to the choice of probe sequences is achieved by use of ’virtual photomasks’ [SG99] on the basis of a spatial light modulator (Digital Micromirror Device, DMDTM, Texas Instruments Inc.). A microscope projection photolithographysystem is employed to project the ’virtual masks’ (i.e. the photomask images shown on the DMDTM) onto the surface of the microarray substrate. Spatially controlled photodeprotection of photolabile NPPOC protective groups (followed by coupling of a further nucleotide building block) enables massively parallel synthesis of DNA probe sequences. In the automated synthesis process microarrays are routinely fabricated over night. Comparable in situ synthesis systems are currently operated only at very few institutions around the world. We first report the application of phosphorus dendrimer substrates [LB03] in the in situ synthesis of DNA microarrays. With the phosphorus dendrimer functionalization we obtained superior results in regard to sensitivity, surface homogeneity, signal/backgroundratio and reusability of the microarrays. We performed microarray hybridization experiments to investigate the impact of single base defects (deliberately introduced single base mismatches and single base bulges)on the binding affinity of oligonucleotide duplexes. This is particularly interesting with regard to genotyping microarrays which are increasingly employed as a molecular diagnostics tool for the detection of single nucleotide polymorphisms (SNPs). In a number of experiments we investigatedthe large influence of the single-defect position [Wic06; Poz06; Nai06b] on duplex binding affinity. The origin of this positional dependence – which is apparently not in agreement with the (two-state) nearest-neighbor model – had not been identified so far. We discovered that the influence of the defect position is not restricted to single base mismatches but can also be observed for single base bulge dei
fects. On the basis of the double-ended zipper model [Gib59; Kit69] (assuming fluctuating end-domain-opening of the oligonucleotide duplex) we could reproduce the experimentally observed positional influence. Moreover, our theoretical investigations on the zipper model indicate a significant positional influence in regard to the contributions of the individual Watson-Crick nearest-neighbor pairs to the Gibbs free energy of oligonucleotide duplex formation. The present work provides for the first time a theoretical approach for the positional-dependent nearest-neighbor model (PDNN) of Zhang et al. [Zha03]. In the in situ synthesis process of DNA microarrays random point-mutationsare introduced into the microarray probe sequences. We have shown – experimentally and by means of a numerical model – that synthesis-related defects significantly affect microarray hybridization characteristics. With regard to single base mismatch discrimination, we discovered significant differences between DNA/DNAand RNA/DNA hybridization: experimental results indicate an improved discrimination of purine-purine mismatch base pairs in RNA/DNA-duplexes. For the experimentally observed, unexpectedly high stability of Group II single bulges [Zhu99] we provide an explanatory approach on the basis of the zipper model. The selection of appropriate (specific and sensitive) probe sequences is of crucial importance for successful application of DNA microarray technology. Our experimental results confirm previous results [Lue03] which show that only a small fraction (in piecewise sections about 20-30%) of a long cRNA target sequence is available for hybridization with the complementary microarray probes. Reduced binding affinities are assumed to originate from the influence of target secondary structure. Using software tools for antisense oligonucleotidedesign (accounting for target accessibility) we were able to predict efficient microarray probes. We discovered evidence that mechanically stable secondary structures (e.g. double-helical sections) interfere with the microarray surface (sterical hindrance) and thus result in reduced microarray binding affinities. ii
Kurzzusammenfassung In der vorliegenden Arbeit wurde die Hybridisierung einzelstr¨angiger RNAund DNATarget-Sequenzen mit denf¨ur die einzelnenSequenzen spezifischen Oligonukleotid-ProbeSequenzen auf der Oberfl¨ache von DNA-Microarrays untersucht. Die hierbei verwendeten Oligonukleotid-Microarrays wurden mittels eines im Rahmen dieser Arbeit entwickelten Microarray-Synthese-Systems auf der Basis eines automatisierten, photolithographischkontrollierten Syntheseprozesses [Fod91] hergestellt: Mit Hilfe eines kombinatorischen Verfahrens wurden – ausgehend von chemisch modifizierten NPPOC-Phosphoramidit Basenbausteinen [Has97] – in paralleler Weise bis zu 25000 unterschiedliche (frei w¨ahlbare) Probe-Sequenzen in situ auf dem Microarraysubstrat synthetisiert. Eine hohe Flexibilit¨at hinsichtlich der Auswahl der Probe-Sequenzen wird durch die Verwendung virtueller ”Photomasken” [SG99] – auf der Basis eines Mikrospiegelarrays (DMDTM Digital Micromirror Device, Texas Intruments Inc.) – erreicht. Mittels einer Mikroskop-Projektions-Photolithographie-Konfiguration wird das Bild des Spatial Light Modulators auf die Substratoberfl¨ache abgebildet, um die Entsch¨utzung photolabiler NPPOC-Schutzgruppen – und damit die nachfolgende Ankopplung weiterer Basenbausteine – r¨aumlich kontrolliert zu steuern. Mit den in unseren Experimenten erstmals bei einer in situ Synthese verwendeten Phosphorus-Dendrimer-Substraten [LB03] konnten im Vergleich mit anderen Linker/SpacerMolek¨ulen die besten Resultate in Hinsicht auf Sensitivit¨at, Homogenit¨at, Signal/Untergrund-Verh¨altnis und Wiederverwendbarkeit, erzielt werden. Mit dem Microarray-Synthesizer k¨onnen in einem automatisierten Prozess DNA Microarrays mit Tausenden von beliebig w¨ahlbaren Probe-Sequenzen praktisch ¨uber Nacht hergestellt werden. Vergleichbare Systeme stehen bislang nur wenigen Forschungseinrichtungen zur Verf¨ugung. Anhand von Hybridisierungsexperimenten wurde untersucht, wie sich (gezielt eingebaute) Einzelbasen-Defekte auf die Bindungsaffinit¨at von Oligonukleotid-Duplexen auswirken. Dies ist in Hinsicht auf die Anwendung von SNP-Microarrays interessant, die zur Detektion von Single Nucleotide Polymorphismen – genetisch bedingten Variationen einzelner Basenpaare – in zunehmenden Maße in der molekularen Diagnostik eingesetzt werden. In einer Reihe von Experimenten lag das Augenmerk auf dem starken Einfluss der Defektposition [Wic06; Poz06; Nai06b] auf die Bindungsaffinit¨at. Die Ursache dieser offensichtlich im Widerspruch zum two-state nearest-neighbor-Modell stehenden Positionsabh¨angigkeit konnte bislang nicht erkl¨art werden. Unsere Experimente zeigen erstmals, dass die Positionsabh¨angigkeit nicht nur bei Mismatch-Defekten [Wic06; Poz06; Nai06b], sondern in vergleichbarer St¨arke auch bei single bulge Defekten auftritt. Auf der Basis eines iii
Zipper-Models des Oligonukleotid-Duplexes, bei dem eine fluktuierende partielle Denaturierung der Duplexenden angenommen wird (die auch zur vollst¨andigen Dissoziation f¨uhren kann), konnte der experimentell beobachtete Positionseinfluss reproduziert werden. Dar¨uber hinaus zeigen unsere theoretischen Untersuchungen (auf der Grundlage des Zipper Modells) einen signifikanten Positionseinfluss hinsichtlich der Gewichtung der einzelnen nearest-neighbor-Beitr¨age zur Duplexstabilit¨at auf. Die vorliegende Arbeit liefert damit erstmals einen theoretischen Ansatz f¨ur das positional-dependent nearest-neighbor Modell (PDNN) von Zhang et al. [Zha03]. Verursacht durch Streulicht und andere Einfl¨usse werden im Verlauf der in situ Synthese zuf¨allige Punktmutationen in den Microarray-Probe-Sequenzen generiert. Experimentell und in numerischen Modellen konnte gezeigt werden, dass diese Synthesedefekte maßgeblich die Hybridisierungseigenschaften entsprechender Microarrays beeinflussen. Eine detaillierte Analyse des Einflusses der einzelnen Mismatch-Basenpaare auf die Bindungsaffinit¨at zeigt hinsichtlich der Mismatch-Diskriminierung signifikante Unterschiede zwischen DNA/DNAund RNA/DNA-Hybridisierung auf, die wahrscheinlich auf unterschiedliche Duplexstrukturen zur¨uckzuf¨uhren sind. F¨ur die experimentell beobachtete, vergleichsweise hohe Stabilit¨at von Group II single bulge [Zhu99] Defekten konnte ein Erkl¨arungsansatz auf der Basis des Zipper-Modells gefunden werden. F¨ur die Durchf¨uhrung von Microarrayexperimenten ist die Auswahl geeigneter ProbeSequenzen mit einer hohen Bindungsaffinit¨at hinsichtlich der dazu komplement¨aren Target-Sequenzen von entscheidender Bedeutung. Wir konnten fr¨uhere Resultate [Lue03] best¨atigen, wonach – vermutlich durch den Einfluss der Targetsekund¨arstruktur – nur ein relativ kleiner Teil (abschnittsweise etwa 20 bis 30%) einer mehrere hundert Nukleotide langen cRNA Target-Sequenz f¨ur die Hybridisierung mit den Microarray-Probes zur Verf¨ugung steht. Auf der Grundlage eines Software Tools f¨ur das Design von AntisenseOligonukleotiden (Ber¨ucksichtigung der Targetsekund¨arstruktur) konnten die experimentell bestimmten Hybridisierungseffizienzen der Microarray-Probe-Sequenzen reproduziert werden. Dar¨uber hinaus entdeckten wir Hinweise daf¨ur, dass mechanisch stabile Sekund¨arstrukturen (z.B. doppelhelikale Abschnitte) durch Wechselwirkung mit der MicroarrayOberfl¨ache – aufgrund von sterischer Hinderung der Duplexbildung– die Bindungsaffinit¨at herabsetzen. iv
Contents 1 Introduction 1 2 Fundamentals 7 2.1 Nucleic Acids . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 2.1.1 The Double-Helix Structure . . . . . . . . . . . . . . . . . . . . 8 2.1.2 Stabilizing Interactions . . . . . . . . . . . . . . . . . . . . . . . 9 2.1.3 Differences between DNA and RNA . . . . . . . . . . . . . . . . 14 2.2 Biological Functions of Nucleic Acids . . . . . . . . . . . . . . . . . . . 15 2.2.1 The Central Dogma of Molecular Biology . . . . . . . . . . . . . 15 2.2.2 Genomic DNA . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 2.2.3 Genes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 2.2.4 Gene Expression . . . . . . . . . . . . . . . . . . . . . . . . . . 17 2.2.5 Expression Regulation . . . . . . . . . . . . . . . . . . . . . . . 19 2.2.6 Biological Functions of RNA . . . . . . . . . . . . . . . . . . . 22 2.3 Nucleic Acid Hybridization . . . . . . . . . . . . . . . . . . . . . . . . . 23 2.3.1 Kinetics of Nucleic Acid Hybridization . . . . . . . . . . . . . . 24 2.3.2 The Nearest-Neighbor Model . . . . . . . . . . . . . . . . . . . 27 2.3.3 Zipper-Model of the Oligonucleotide Duplex . . . . . . . . . . . 30 2.3.4 Further Models of the DNA Melting Transition . . . . . . . . . . 34 2.4 Destabilization of Oligonucleotide Duplexes by Point Defects . . . . . . 34 2.4.1 Single Base Mismatches . . . . . . . . . . . . . . . . . . . . . . 35 2.4.2 Single Base Bulges . . . . . . . . . . . . . . . . . . . . . . . . . 36 2.4.3 Influence of the Defect Position . . . . . . . . . . . . . . . . . . 39 2.5 Solid-Phase Synthesis of Nucleic Acids . . . . . . . . . . . . . . . . . . 41 2.5.1 Principles of Solid-Phase Chemical Synthesis . . . . . . . . . . . 41 2.5.2 Nucleic Acid Synthesis by the Phosphoramidite Method . . . . . 42 2.6 DNA Microarrays . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45 2.6.1 Microarray Applications . . . . . . . . . . . . . . . . . . . . . . 47 v
CONTENTS 2.6.2 The Development of DNA Microarray Technologies . . . . . . . 49 2.6.3 Characteristics of Microarray Hybridization . . . . . . . . . . . . 51 2.6.4 Further Reading . . . . . . . . . . . . . . . . . . . . . . . . . . 53 2.7 DNA Chip Fabrication by Light-Directed In Situ Synthesis . . . . . . . . 54 2.7.1 Photolithographic Control of the Combinatorial Synthesis Process 54 2.7.2 ”Maskless” Photolithography and Combinatorial Chemistry . . . 57 3 Development of the DNA Microarray Synthesizer 61 3.1 Motivation and Overview . . . . . . . . . . . . . . . . . . . . . . . . . . 61 3.2 The Maskless Microprojection Photolithography System (MPLS) . . . . . 63 3.2.1 The UV Light Source . . . . . . . . . . . . . . . . . . . . . . . . 63 3.2.2 Digital Mask Projection Using a Digital Micromirror Device . . . 65 3.2.3 The Image Projection Optics . . . . . . . . . . . . . . . . . . . . 66 3.2.4 UV-Sensitive Photochromic Films . . . . . . . . . . . . . . . . . 69 3.2.5 Chromatic Correction of the Projection Optical System . . . . . . 70 3.2.6 UV Light Intensity and Uniformity of Illumination . . . . . . . . 71 3.2.7 Optical System Performance Testing . . . . . . . . . . . . . . . . 73 3.2.8 Outlook - Further Possible Applications . . . . . . . . . . . . . . 78 3.3 The Fluidics System . . . . . . . . . . . . . . . . . . . . . . . . . . . . 80 3.3.1 The Synthesis Cell . . . . . . . . . . . . . . . . . . . . . . . . . 80 3.3.2 Argon Bubble Trapping . . . . . . . . . . . . . . . . . . . . . . 83 3.4 Automated Microarray Synthesis . . . . . . . . . . . . . . . . . . . . . . 84 3.5 Performance of the Microarray Synthesizer . . . . . . . . . . . . . . . . 84 4 Light-directed in situ Synthesis of DNA Microarrays 87 4.1 Light-Directed in situ Synthesis of DNA Microarrays . . . . . . . . . . . 87 4.2 Preparation of Phosphorus Dendrimer Substrates . . . . . . . . . . . . . 91 4.3 Noteworthy Characteristics of the Microarrays . . . . . . . . . . . . . . . 94 4.3.1 Autofluorescence of the Chip Surface . . . . . . . . . . . . . . . 94 4.3.2 Hydrophilicity of DNA Microarray Features . . . . . . . . . . . . 95 4.3.3 Hybridization without Detergent - Unspecific Adsorption . . . . . 96 4.3.4 Irreversible Target Adsorption . . . . . . . . . . . . . . . . . . . 96 4.3.5 Robustness of the Phosphorus Dendrimer Surface Coating . . . . 97 5 DNA Microarray Analysis 99 5.1 Hybridization Signal Acquisition - Experimental Setup . . . . . . . . . . 99 5.1.1 The Hybridization Chamber . . . . . . . . . . . . . . . . . . . . 100 5.1.2 Epifluorescence Microscope . . . . . . . . . . . . . . . . . . . . 102 vi
CONTENTS 5.1.3 Image Acquisition with an EM-CCD Camera . . . . . . . . . . . 103 5.2 Quantitative Analysis of Microarray Hybridization Signals . . . . . . . . 103 5.3 Real-time Monitoring of Microarray Hybridization . . . . . . . . . . . . 105 5.3.1 Hybridization Buffer . . . . . . . . . . . . . . . . . . . . . . . . 106 5.3.2 Microarray Washing Procedures . . . . . . . . . . . . . . . . . . 106 6 Influence of Point-Defects on Oligonucleotide Duplex Binding Affinities 109 6.1 Motivation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 109 6.2 Conception of the Microarray Hybridization Experiments . . . . . . . . . 111 6.3 DNA Microarray Design . . . . . . . . . . . . . . . . . . . . . . . . . . 111 6.3.1 Chip Design - Quantitative Analysis of Hybridization Signals . . 113 6.3.2 Single Base Defect Experiments . . . . . . . . . . . . . . . . . . 114 6.4 Hybridization Assays and Image Analysis . . . . . . . . . . . . . . . . . 115 6.4.1 Oligonucleotide Targets . . . . . . . . . . . . . . . . . . . . . . 115 6.5 Dominant Influence of the Defect Position . . . . . . . . . . . . . . . . . 115 6.6 Mismatch Discrimination in DNA/DNA Duplexes . . . . . . . . . . . . . 121 6.6.1 Experimental Results . . . . . . . . . . . . . . . . . . . . . . . . 121 6.6.2 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 6.7 Influence of Flanking Base Pairs . . . . . . . . . . . . . . . . . . . . . . 128 6.8 Mismatch Discrimination in DNA/DNA and RNA/DNA Duplexes . . . . 132 6.8.1 Outline of the Experiment . . . . . . . . . . . . . . . . . . . . . 132 6.8.2 Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133 6.8.3 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135 6.9 Single Base Bulge Defects . . . . . . . . . . . . . . . . . . . . . . . . . 137 6.9.1 Statistical Analysis . . . . . . . . . . . . . . . . . . . . . . . . . 137 6.9.2 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 140 6.10 Comparison of Single Base Mismatches and Single Base Bulges . . . . . 143 6.11 Binding Affinities of Duplexes Containing Multiple Defects . . . . . . . 146 6.11.1 Results and Discussion . . . . . . . . . . . . . . . . . . . . . . . 147 7 Modeling the Influence of Point Defects on Duplex Stability 151 7.1 The Double-Ended Zipper Model . . . . . . . . . . . . . . . . . . . . . . 151 7.2 Stochastic Simulation of Oligonucleotide Duplex Stability . . . . . . . . 153 7.2.1 Stochastic Simulation with the Gillespie Algorithm . . . . . . . . 153 7.2.2 Simulation Results . . . . . . . . . . . . . . . . . . . . . . . . . 155 7.3 Partition Function Approach of the Double-Ended Zipper Model . . . . . 159 7.3.1 Implementation of the Partition Function Approach (PFA) . . . . 160 vii
Glossary oligonucleotide short nucleic acid strand perfect match duplex consisting of two completely complementary strands; defect-free duplex probe a microarray probe is used to detect/identify one specific nucleic acid target sequence; probes are typically oligonucleotide probes (length<100 nt) or several hundred nt long cDNA sequences; probes are tethered in a regular arrangement (array) – within the microarray features – on the solid support probe sequence motif in the present work this expression is used for the perfect matching probe sequence that is complementary to the oligonucleotide target sequence employed in a single base defect hybridization experiment. Single base defect probes are derived from the ’probe sequence motif’ by substitution, insertion or deletion of a single base. The probe sequence motif may be shorter than the target oligonucleotide used in the experiment. Hybridization signals from the complete set of single base defect probes correspond to the ’defect profile’. single base bulge defect in a nucleic acid duplex which originates from a surplus unpaired base in one of the two strands; the surplus base can adopt a stacked-in conformation or a loopedout conformation and can result in significant reduction of the binding affinity single base mismatch defect in a nucleic acid duplex which originates from a non-Watson-Crick base pair; the reduced binding affinities is employed for detection of SNPs and point-mutations target free nucleic acid sequence whose identity and abundance are to be detected in the microarray assay; for detection target sequences are commonly labeled with fluorescent dyes or with biotin xiv
Chapter 1 Introduction Almost all cells of the human body, regardless of the cell type, contain the same genetic material. However, owing to epigenetic factors (e.g. CpG methylation) the cell types differ in their gene expression – for example, genes which are strongly expressed1in one cell type, may not be expressed in others. Knowledge on gene expression is the key for understanding the individual gene functions and the complex interactions between the about 20,000 to 25,000 genes of the human genome. DNA microarrays are a key technology for massively parallel analysis of gene expression. The working principle of DNA microarrays is based on nucleic acid hybridization: sequential Watson-Crick base pairing between the bases of two complementary nucleic acid strands results in the formation of a relatively stable double-helical duplex. Nucleic acid hybridization is highly specific – already a single mismatched (non-Watson-Crick) base pair can significantly reduce the binding affinity [Nel81; Pat82]. The sequence-specific hybridization between complementary strands is employed for the purpose of molecular recognition (Fig. 1.1): surface-tethered single-stranded probes (of known sequences) are employed as sequence-specific scavengers for complementary target sequences in solution. Hybridized target molecules (bound to the surface) can be detected by means of radioactive or fluorescent dye labels. On DNA microarrays the same detection principle is applied in parallel fashion (Fig. 1.2). Owing to the high specificity of nucleic acid hybridization thousands or even millions of different target sequences can be detected simultaneously. DNA microarrays comprise a regular array of microarray features, small areas, each of which is covered with surfacetethered single-stranded DNA probes of a well-known sequence. Individual microarray features (and thus the corresponding probe and target sequences) can be identified by their position on the microarray. 1Gene expression – the conversion of genetic information into gene products – can be understood as ’gene activity’. 1
Introduction Figure 1.1: Nucleic acid hybridization between surface-tethered probe strands and complementary target strands in solution. Nucleic acid hybridization is based on sequential Watson-Crick base pairing between complementary sequences of nucleotides and results in the formation of a relatively stable double-helical nucleic acid duplex. Nucleic acid hybridization is reversible (dissociation is favored by increased temperatures) because the individual binding interactions (hydrogen bonding and base stacking interactions – no covalent bonds involved) between the base pairs are relatively weak. Targets strands are labeled by covalent linkage of a fluorescent dye, or alternatively, by biotinylation. In a gene expression profiling experiment the messenger RNA (mRNA) sequences (indicators of the individual genes transcriptional activities) are isolated from the biological sample, amplified (if necessary, e.g. by in vitro transcription), and labeled for detection. Subsequently the complex mixture of target sequences to be analyzed is applied (in hybridization buffer solution) onto the surface of the microarray. The target strands can freely diffuse around and interact with the surface-tethered microarray probes, until they are captured by a complementary probe and form a stable duplex. After removal (washing-off) of unhybridized targets, the hybridization signal, which provides information on the quantity of the individual target sequences, is commonly detected by means of fluorescent markers. Comparison of the hybridization signals with the corresponding hybridization signals from a reference sample (by dual-color analysis on the same chip, or by means of two single-channel microarrays) enables identification of genes that have been upor downregulated. Some commercial platforms enable gene expression profiling on a genome-wide scale. Genotyping analysis is a further important microarray application: Single nucleotide polymorphisms (SNPs) – sequence variations in which single nucleotides differ between the members of a species2(or even between the two alleles in diploid cells) – have a strong influence on the phenotype. SNPs are responsible for the majority of genetic variations 2The human genome contains about 3 million SNPs. Thus, about one in a thousand base pairs is subject to this type of inheritable genetic variation. 2
Introduction AB C Figure 1.2: Nucleic acid hybridization on the microarray. Three different surfacetethered probe species a,band care located separate from each other within the corresponding microarray features A, B and C. A complex mixture of different target sequences is applied to the microarray surface. Driven by diffusion (or active mixing) targets move around and interact with the different probes. If a target meets a complementary probe, a stable duplex can arise. Thus, the target gets captured by the complementary probe. After the hybridization the unbound targets can be removed by washing-off. The remaining hybridization signal (fluorescent signal) of the hybridized probes provides information on the quantities of the individual target sequences. In this example we observe hybridization signals only at features A and B. We conclude that the sample mixture contains targets sequences that are complementary to the probes aand b. The sample does not contain targets complementary to probe c. within a single species. They are also associated with a predisposition to a variety of diseases. Moreover, SNPs are associated to individuals’ response to pathogens, chemicals, drugs, vaccines, and other agents. SNP microarrays make use of the specificity of relatively short 12 to 30mer oligonucleotide probes to detect single mismatched base pairs originating from SNPs [Con83]. Genotyping arrays are a valuable tool in genomics research, pharmaceutical research (with a focus on the individual response to pharmaceutical agents) and now increasingly in medical diagnostics. Further applications of DNA microarrays include resequencing assays3and the identification of pathogens. Lab-scale fabrication of DNA microarrays on the basis of standard techniques requires considerable technical and financial efforts.4To provide a flexible and affordable basis for DNA microarray hybridization experiments we developed a DNA microarray synthesizer 3Resequencing arrays are used for the search for mutations with respect to a well-known reference sequence. An important application is the identification of (possibly new) virus strains. 4These include, for example, the acquisition of a microarray spotting robot (to be operated in a clean room environment) and considerable running expenses for presynthesized microarray probes. 3
Introduction based on the work of Singh-Gasson et al. [SG99]. Based on a photochemically controlled in situ synthesis process, the DNA probe sequences are synthesized from nucleotide building blocks, directly on the surface of the microarray. The use of expensive chromium photomasks (and associated mask alignment) is circumvented by means of a spatial light modulator (DMDTM) obtained from a commercial video projector. Comparable in situ synthesis systems are currently operated only at a few institutions around the world. Even though DNA microarrays have become a well-established technology, the underlying physicochemical principles of DNA microarray hybridization are not yet fully understood [Lev05; Poz06]. For example, an unresolved problem in the application of DNA microarrays is the lack of predictability of the hybridization efficiency of DNA microarray probes. Thermal stability of oligonucleotide duplexes (in solution-phase) is well described by the nearest-neighbor model [Cro64; Tin73; Bre86; Fre86], which is accounting for hydrogen bonding and also for base-stacking interactions between adjacent base pairs. Thermodynamic parameters for nearest-neighbor doublets of base pairs were derived from solutionphase hybridizationexperiments [San98]. The nearest-neighbor model is widely employed for the prediction of duplex melting characteristics (melting temperatures, Gibbs free energies of duplex formation) – for example, for the design of PCR primers and for the design of DNA microarray probe sequences. The latter application, however, is questionable: on DNA microarrays, due to various surface-effects and fabrication-related effects, there are significant differences with respect to solution-phase hybridization [Hel03; Lev05; Bin06; Poz06]. Moreover, the secondary structure of long target sequences results in a restricted target accessibility. Thus, the binding affinity of individual microarray probes is also governed by the complex target secondary structure [Lue03; Rat05]. In contrastto solution-phasehybridizationstudies, recent microarraystudies[Wic06; Poz06; Nai06b] report a large influence of the position of single base mismatch defects on the hybridization signal. A position dependent influence of single base defects is not considered by the (two-state) nearest-neighbor model5and hasn’t been explained so far. According to Pozhitkov et al. [Poz06] there is little evidence that microarray hybridization efficiencies can be accurately predicted with software tools on the basis of nearest-neighbor thermodynamic parameters derived from solution-phase experiments. In the experimental part of the present thesis particular interest is on the influence of point defects (single base mismatches and single base bulges) on microarray binding affinities. We systematically investigate the influence of defect type and defect position on probetarget binding affinities. In the same contextwe investigatedifferences between RNA/DNA 5The nearest-neighbor model, on the basis of mismatch base pair nearest-neighbor parameters [All97], is also employed for mismatched duplexes [All97; San04]. The model does not consider a position dependent influence, except for the outermost base positions. 4
Introduction and DNA/DNA hybridization. Our theoretical investigation on the influence of point defects on duplex binding affinities is based on a zipper model [Gib59; Kit69] of the oligonucleotide duplex. Further experiments address a variety of poorly understood influences on DNA microarray hybridization. These include: •random defects in the microarray probes (generated by the in situ synthesis process) affect the hybridization characteristics [Job02; Gar02; Hel03] •the complex secondary structure of long target sequences (widely believed to be a main factor influencing the efficiency of hybridization [Lue03]) •nonspecific cross-hybridization •diffusion limitation – local depletion of the hybridizing target molecules can result in inhomogeneous hybridization signal intensities and slowed-down hybridization kinetics [Pap06; Dan07] 5
Introduction 6
Chapter 2 Fundamentals 2.1 Nucleic Acids For its outstanding role in molecular biology DNA (deoxyribonucleic acid) is often referred to as the ”molecule of life”. Like a blueprint genomic DNA contains the hereditary information, instructions to grow and sustain all forms of life. In the higher eucaryotic organisms the genomic DNA is densely packed on chromosomes inside the cell nucleus. Each chromosomecomprises a single double-helical DNA molecule. The length of the human chromosomes is varying between 50×106and 250 ×106base pairs (corresponding to lengths between 1.7 and 8.5 cm). Overall, stretched end-to-end, the DNA helix contained in a single human diploid cell is about 2 meters long. The information density in the densely packaged nucleus is about 1021 bit/cm3 1 (for comparison: the information storage density on a DVD disc is about 109bit/cm2). The biological function of DNA is the safe storage of genetic information. Genomic DNA is basically a read-only memory and in this way comparable to the CD-ROM drive of a computer. Parts of the genome (the genes) are read in the transcription process to produce RNA transcripts of the DNA sequence. RNA is a rather volatile information carrier in the ongoing processing of genetic sequence information. In the above analogy it is therefore comparable with the working memory (RAM) of a computer. However, RNA is more versatile: its not just an information carrier but rather (in form of functional RNA) a crucial part of the translational machinery and involved in regulatory processes. 1The estimate is based on a cell volume of 8 µm3and a genome size of 3·109base pairs - which is about the size of the human genome. 7
Fundamentals 2.1.1 The Double-Helix Structure of Nucleic Acids The structure of deoxyribonucleic acid (DNA) was discovered by James Watson and Francis Crick in 1953. A few month earlier, Linus Pauling reported a triple-helix model of the DNA-structure, which assumed that the phosphate groups are arranged in the interior of the helix. The model was based on high resolution electron micrographs showing the DNA as cylindric fibrils with a diameter of 1.5 nm. The wrong assumption of a triple helix originated from the incorrect measurement of a too high packaging density. Watson and Crick showed that under physiological conditions DNA has indeed a doublehelical structure (Fig. 2.1). The hydrophobic bases are located in the center, whereas the hydrophilic phosphate groups are located at the outside of the helix. The discovery of Watson and Crick relies on the work of Rosalind Franklin, whose X-Ray structural analysis of DNA fibres proved that DNA has indeed a helical structure. Figure 2.1: Watson-Crick model of the DNA double-helix. Canonical (WatsonCrick) base pairs comprise either adenine (red) and thymine (blue), or guanine (green) and cytosine (yellow) bases. The sugar-phosphate-backbones of the two strands (shown in green and cyan) form a right-handed double helix. The ideal B-DNA structure was generated with the make na web-server which is based on the NAB (Nucleic Acid Builder) by Tom Macke [Mac98]. Image visualization was performed with the UCSF Chimera molecular modeling system [Pet04]. A three-dimensional stereo view of the DNA structure is shown in the appendix, in Fig. B.19. 8
Nucleic Acids Another important hint was provided by Erwin Chargaff. According to Chargaff’s rule the nucleobases A and T, just as the nucleobases C and G always occur with the ratio of about 1:1, independent of the biological origin of the DNA. 2.1.2 Nucleic Acid Duplex Structure - Stabilizing Interactions The DNA double-helix shown in Fig. 2.1 is composed of two complementary singlestranded DNA molecules. It’s well-known that the duplex stability originates from interstrand hydrogen bonding between complementary base pairs A·T and C·G (see Fig. 2.2). However, it is less well-known that a similar degree of stabilization originates from π-π interaction between closely-stacked aromatic bases (π-stacking) [Koo01]. N N N O H H N N N N N O H H HN N N NN H H N N O O CH3 H Adenine Thymine Guanine Cytosine Figure 2.2: Canonical (Watson-Crick) base pairs A·T and C·G, comprise a bicyclic purine base (adenine or guanine) and a monocyclic pyrimidine base (thymine or cytosine). A·T is stabilized by two, C·G by three hydrogen bonds. A DNA molecule is basically a flexible polymer-chain2made up of nucleotide monomers (as shown in Fig. 2.3). A nucleotide consists of a heterocyclic base (i.e. adenine, cytosine, guanine or thymine - in RNA the thymine is replaced by uracil) and a pentose sugar-ring (2-deoxyribose in DNA and ribose in RNA), which in conjunction with a phosphate group constitutes the sugar-phosphate-backbone of the DNA molecule. Apart from the nucleobases listed above, further nucleobases occur naturally in RNA (e.g. pseudouridine in transfer-RNA). Fig. 2.3 shows that subsequent nucleotides are linked via a phosphodiester bond (i.e. over the phosphate group) between the 3’- and 5’-carbons of the deoxyribose sugars. Because 2Here one needs to distinguish between the highly flexible single stranded molecule (persistence length values provided in the literature range from lp≃0.5 nm to 1.3 nm [Koh06]) and the significantly more rigid double-stranded DNA duplex (lp≃45-50 nm [Hag88]). 9
Fundamentals RNA DNA PROTEIN translation transcription DNA replication reverse transcription RNA replication direct translation of DNA Figure 2.7: Francis Cricks Central Dogma of Molecular Biology. There are general transfers of biological sequence information (solid arrows) and specialized transfers (dashed arrows). A general transfer of sequence information is from DNA via the transcription process to messenger RNA. mRNA is translated into a polypeptide chain which folds into a protein. Another general transfer is the replication of DNA during cell division. Specialized transfers are related to virus reproduction (e.g. reverse transcription) or have be performed in vitro (e.g. direct translation of DNA sequences into proteins). 2.2.2 Genomic DNA The genomic DNA contains the hereditary information of an organism. Large parts of the genome are arranged as genes, organizational units that are transcribed into one or several gene products. The function of other noncoding parts of the genome, previously termed ”junk DNA” is less well understood. In the simple procaryotic organisms (e.g. bacteria) the DNA is packaged in ring-shaped chromosomes and plasmids, which are residing in the cytoplasm. In the more complex eucaryotic organisms the DNA is contained in the nucleus, well-separated from the cytoplasm (see Fig. 2.8). Chromosomes contain the DNA in a highly compact, though ordered and accessible form. The double-helical DNA filament contained on a single chromosome can be several centimeters long. Enlarged to a diameter of 2 mm the DNA filament would extend over a length of about 30 km. In conjunction with a complex of histone proteins, acting as spool around which the DNA double-strand is wound up (roughly two superhelical turns of about 80 base pairs around the cylindrical histone octamer), the DNA forms a nucleosome. Countless nucleosomes condense into an ordered superstructure, forming a chromatin fibre with a diameter of about 30 nm. The chromatin fibre (which via certain domains is connected to the nuclear matrix proteins) forms innumerable loops which compose the structure of the chromosome. 16
Biological Functions of Nucleic Acids The degree of chromatin condensation is largely determined by the cell-cycle. The chromatin structure, since it determines the accessibility and readability of genes, has a strong influence on gene activity. Transcriptional active regions correlate with an open chromatin structure (euchromatin). 2.2.3 Genes A gene can be understood as a functional unit of the genetic material, which contains the blue print for a gene product. A gene product can be one or several proteins (or subunits of proteins) or a functional RNA, e.g. microRNA (miRNA), ribosomal RNA (rRNA), or transfer RNA (tRNA).4 2.2.4 Gene Expression Gene expression can be understood as gene activity. It describes how much of a gene product is produced from each particular gene. The gene activity can be regulated at different stages, e.g. at the transcription initiation, or post-transcriptional, at the mRNA level. The various cell types of a higher organism all contain the same genetic information. However, the gene activities are different, depending on the requirements of the particular cell function. Transcription Transcription requires opening of double helix structure first. It is assumed that the reduced duplex stability within the Pribnow-box (comprising the sequence motif TATAAT) supports the opening of the transcription bubble. The transcription bubble extends over about 18 base pairs. Transcription initiationis followedby the elongation process, in which an RNA-copy of the sense-strand (only the sense-strand encodes the sequence information for the gene product) is transcribed until a terminator sequence at the end of the gene is reached. During elongation, the holoenzyme slides along the operon from 5’ to 3’ direction (with respect to sense strand - see Fig. 2.8). The correct nucleotides for the assembly of the mRNA strand are recognized by complementary base pairing with the coding strand. RNA polymerase joins these nucleotides with the growing RNA strand. A proofreading mechanism replaces incorrectly added nucleotides. 4These don’t serve as templates for the synthesis of polypeptide strands but rather constitute a crucial part of the cells molecular machinery or, like miRNA, are involved in the regulation of the expression of other genes. 17
Fundamentals Figure 2.8: Transcription (A) and translation (B). RNA polymerase opens a transcription bubble and produces a RNA copy of the sense strand while sliding in 5’ to 3’ direction (with respect to the sense strand) until transcription termination is encountered. After poly-adenylation (not shown) the mRNA is released from the nucleus through the nuclear pores. Translation of the mRNA sequence into a polypeptide sequence (B) is performed in the cytoplasm. Ribosomes move along the mRNA in 5’ to 3’ direction, thereby translating the genetic code into a polypeptide sequence. Proteins emerge from folding of the polypeptide chains. The transcription ends when a terminator sequence is encountered. Then the transcription complex comes apart, the transcription bubble collapses and the mRNA strand is released. Still in the nucleus the (eucaryotic) mRNA undergoes poly-adenylation (addition of a poly-A-tail at the 3’-end). By binding the poly(A)-binding protein (PABP) the poly-Atail protects the mRNA from degradation and increases the translation of the mRNA. The poly-A sequence is technically employed for the specific extraction of mRNA sequences with poly-T functionalized magnetic beads. Translation Messenger RNA (mRNA) is used as a template for the synthesis of proteins. Single stranded RNAs similar like polypeptide chains can fold and have the capability to form complex tertiary structures, similar as proteins. Ribosomal RNA (rRNA), the most abundant RNA in cells, is not a simple information carrier like DNA, but rather folds itself into a complex ”nanomachine” which is crucial for the synthesis of polypeptide chains. Like tiny robots ribosomes slide along the mRNA strands (downstream from the 5’- to 18
Biological Functions of Nucleic Acids the 3’-end) and translate the nucleic acid sequences via the genetic code into polypeptide sequences (Fig. 2.8). The molecular recognition of the codons (base triplets encoding for amino acids) is performed with transfer RNA (tRNA), another functional RNA structure (Fig. 2.10). The anticodon, an exposed base triplet at the end of the anticodon arm of the tRNA, can specifically bind via base pairing5to a complementary codon sequence on the mRNA strand. Upon binding the corresponding amino acid which was carried by the tRNA to the site of polypeptide synthesis is attached to the growing polypeptide chain. Subsequently the ribosome moves on to the next codon and simultaneously releases the discharged tRNA. Translationof an mRNA strand is performed by manyribosomes simultaneously. While the translation process is going on the mRNA strand is degraded by nucleases in the 5’→3’ direction. 2.2.5 Expression Regulation The functions of a cell (e.g. expression of structural and regulatory proteins, differentiation, control of the life cycle, adaption to environmental influences) are largely controlled by gene regulatory networks. Transcription factor proteins (via specific protein-DNA binding) can activate, amplify or inhibit the translation of the targeted gene(s) and thus control the corresponding gene activity. Post-transcriptional regulatory mechanisms include alternative splicing, RNA silencing, antisense suppression, and the regulation of mRNA stability. DNA microarrays enable simultaneous investigation of the activity of many genes, on a genome-wide scale. Gene expression profiles (which are encoding the complex interactions between genes) are an important tool for the investigation gene regulatory pathways (→functional genomics). Expression profiling has also emerged as a promising diagnostic tool for identification of cancer types or subtypes, thus enabling a well-directed therapeutic response. Expression regulation at the transcription level The most prominent regulation mechanism is transcription initiation. In procaryotes essentially only the holoenzyme RNA-polymerase (composed of several subunits) is directly involved in the transcription process. In eucaryotes a large machinery of proteins (including several holoenzymes) needs to form an initiation complex before the transcription can commence. 5Frequently anticodons contain the relatively unspecifically binding nucleotides inosine or pseudouridine. Unspecific binding accounts for the degeneracy of the genetic code. 19
Fundamentals In the simple procaryotic organisms (e.g. bacteria) the initiation of transcription is regulated by activators and repressors. This shall be explained in the following on the example of the regulation of the lactose genes of the bacterium E. Coli, which has been investigated by Jacob and Monod [Jac61]. E. Coli can digest both food sources - glucose and lactose. To conserve resources the lactose metabolism is only activated if only lactose and no glucose is available. In case both sugars are available E. Coli gives preference to glucose since it is the more efficient source of energy. Only if the glucose is depleted and lactose is still present in the medium, E. Coli begins to express the gene for the protein β-galactosidase, an enzyme which is required for the digestion of lactose. The gene for β-galactosidase lacZ is combined with two further genes, lacY and lacA (auxiliary genes, also required for lactose digestion) in a functional unit called operon (Fig. 2.9A). The operon is typical for procaryotic organisms. Apart from the coding sequences for the protein(s) the operon contains the promoter sequence. This is recognized by RNA-polymerase and enables binding of the RNA-polymerase to the double-stranded DNA. The promoter contains the Pribnow-box with the sequence motif TATAAT (typical for procaryotes), and the so-called operator. In the lac operon the operator is a binding site for a repressor protein. The repressor protein, when bound to the operator site, prevents RNA polymerase from binding to the promoter site (Fig. 2.9D). Another sequence motif, adjacent to the promotor, serves as specific binding site for the activator protein CAP, which supports the binding of RNA-polymerase to the promoter site (Fig. 2.9C). The function of the regulatory proteins (activator and repressor) is controlled by the abundance of glucose and lactose, respectively. The activator CAP (a receptor for cyclic AMP) can only bind to CAP binding site (protein-DNA interaction) upon binding to cyclic AMP, which is abundant in the absence of glucose. The lac repressor protein can only bind to the operator site if lactose is not available, since the binding affinity of the repressor protein to DNA is significantly decreased by a conformational change induced from the presence of allolactose. •Glucose and lactose available: In the presence glucose the activator cannot bind to the CAP site. Since the repressor can neither bind, the expression can occur at a low basal level (Fig. 2.9B). •Lactose available/glucose unavailable: The activator (CAP) can only bind near the promoter site if glucose is not available (Fig. 2.9C). In the presence of lactose only, the activator increases the lacZ expression by a factor of about 40 compared to the basal level [Pta02]. 20
Biological Functions of Nucleic Acids lacY lacA Transcription promoter CAP site operator lacZ promoter CAP site operator lacZ rep promoter CAP site operator lacZ CAP RNA Polymerase A D C B promoter CAP site operator lacZ RNA Polymerase glucose & lactose available glucose unavailable lactose available lactose unavailable basal expression expression activated expression inhibited Figure 2.9: The lac operon (A) and the lac expression (B-D). The lac operon comprises the CAP activator site, the promoter and the genes lacZ,lacY and lacA. The latter are transcribed as a single mRNA. The expression level of the lac genes is controlled by the abundance of glucose and lactose, respectively. Activator and repressor proteins which can bind to specific binding sites (protein-DNA interaction), control the binding RNA polymerase. See text for details. Figures were adapted from [Pta02]. 21
Fundamentals •Lactose unavailable: The lac repressor can only bind to the operator site if lactose is not available. In this case the repressor is bound to the operator site preventing the binding of the RNA-polymerase, no matter if the activator is bound to the CAP site (Fig. 2.9D). The expression of the lac genes is inhibited. 2.2.6 Biological Functions of RNA Figure 2.10: Structure of phenylalanine transfer RNA (visualization of 4TNA.pdb [Hin78] with UCSF Chimera). Transfer-RNA is employed in the translation process as a sequence specific vehicle for amino-acids. The anticodon-arm (near the lower edge of the image) contains a unit of 3 nucleotides corresponding to a codon on the mRNA strand. The amino-acid (not shown here) is attached to the acceptor stem (upper right end) with the characteristic CCA 3’-terminal group. The biological function of RNA is more versatile than that of DNA: •In the process of gene expression messenger RNA (mRNA) is employed as a template for polypeptide synthesis. RNA, unlike DNA, is a volatile information carrier with a rather limited lifetime. •Micro-RNAs (miRNA) have regulatory functions. Via the RNA interference (RNAi) mechanism they can specifically inhibit the expression of the corresponding target genes. •Antisense-RNAs (aRNA) have regulatory functions. An aRNA sequence is produced if the noncoding (antisense)-strand of a gene sequence is also being transcribed. Thus the 22
Nucleic Acid Hybridization aRNA is complementary to the mRNA of the particular gene. By base-pairing between the complementary RNA strands the translation of the corresponding polypeptidesequence is inhibited. In the transgenic Flavr SavrT M tomato antisense RNA is employed to suppress the expression of an enzyme involved in ethylene production. The significant reduction of ethylene delays the ripening and spoiling of the tomato. •RNA sequences, similar as polypeptide chains, can fold into complex secondary and tertiary structures. Ribosomal RNA and transfer RNAs (tRNA) are essential parts of the translation machinery (see section 2.2.4). 2.3 Nucleic Acid Hybridization Two complementary (or partially complementary) nucleic acid strands S1and S2can bind via base pairing and form a stable nucleic acid duplex D. The double-helical duplex structure is stabilized by hydrogen bonding and base stacking interactions. S1+ S2 hybridization −−−−−−−* )−−−−−−− dissociation D (2.1) The formation of nucleic acid duplexes is commonly called hybridization since usually nucleic acid strands from different sources (e.g. DNA probes and RNA targets) are involved. Owing to the non-covalent character of the stabilizing interactions nucleic acid hybridization is reversible: In thermodynamic equilibrium the duplex formation is balanced by duplex dissociation (also called duplex denaturation or melting). Lower temperatures and increased ionic strengths (up to 1 M [Na+]) favor duplex formation. With increasing temperature or reduced ionic strength of the hybridization buffer solution the duplexes are increasingly destabilized. Depending on the particular duplex sequence, nucleic acid duplexes can have a very distinct melting transition. Owing to the cooperative character of the duplex binding, the fraction of melted duplexes can change from close to 0% to 100% within a temperature range of a few Kelvins.6 Only a small fraction of the duplexes is in a partially denatured intermediate state. Therefore, the hybridization/melting transition is frequently described as a two-state transition. An important characteristic of the nucleic acid hybridization is its outstanding sequence specificity. Already a single mismatched base within an oligonucleotide duplex can result in a significantly reduced bindingaffinity. Molecular recognition by nucleic acid hybridization is employed by nature (e.g. in RNA interference) and by various molecular biology applications: 6For oligonucleotide duplexes the width of the melting transition is decreasing with increasing Gibbs free energy of the duplex, thus with increasing duplex length. 23
Fundamentals •DNA microarrays •Fluorescent in situ hybridization(FISH): sequencespecificlabelingofmRNA sequences within cells. •Primer sequences are used as starting points for nucleic acid replication (e.g. in PCR or dideoxy sequencing). For this purpose the primers are hybridized to the template strands. •Molecular beacon probes: this type of hairpin-shaped nucleic acid probe containing a fluorophore-quencher-pair becomes fluorescent upon hybridization with a complementary target sequence. •Antisense RNA sequences (sequence-specific silencing of mRNA transcripts) •RNA interference (sequence-specific silencing of mRNA transcripts) 2.3.1 Kinetics of Nucleic Acid Hybridization The widely used two-state model of nucleic acid hybridization assumes that the single stranded species S1and S2are in equilibrium with the duplexes D. S1+ S2 k+ −* )− k− D (2.2) Equation 2.2 doesn’t describe elementary base pairing processes and is therefore valid only if there are no significantly populated intermediate states. The two-state model is a reasonable approximation, for example, for short linear duplexes. The zipper model of DNA duplex melting transition, which considers individual base pairing and base pair dissociation events, is described in section 2.3.3. In the following, for simplicity’s sake, we assume that are duplexes are not self-complementary and that folding of single stranded species (intrastrand base pairing) can be neglected. Duplex formation is a second order reaction, whereas the denaturation is a first order reaction. d[D] dt=−k−[D] + k+[S1][S2] (2.3) In equilibrium (with d[D]/dt=0) we obtain the equilibrium constant K(as described by the law of mass action). K=k+ k− =[D] [S1][S2]=[D] ([S1]0−[D]) ·([S2]0−[D]) (2.4) 24
Nucleic Acid Hybridization The Gibbs free energy of duplex formation ∆G◦ D(◦is referring to standard conditions) is related to the equilibrium constant Kby ∆G◦ D=−R·T·ln K.(2.5) If the complete temperature dependence of the binding affinity - e.g. from experimentally determined plots of 1/Tmversus ln (CT/4) - is known, the Gibbs free energy ∆G◦ Dcan be determined via the van’t Hoff equation: 1 Tm =R ∆H◦ D ·lnCT 4+∆S◦ D ∆H◦ D (2.6) Tmis the melting temperature of the duplex - the temperature at which per definition (in thermodynamic equilibrium) 50% of the duplexes are dissociated. CTis the total concentration of nucleic acid strands. From the total enthalpy ∆H◦ Dand entropy changes ∆S◦ Dthe Gibbs free energy change ∆G◦ Dof the melting transition can be obtained with ∆G◦ D= ∆H◦ D−T·∆S◦ D.(2.7) Alternatively ∆H◦ Dand ∆S◦ Dcan be predicted from sequence-dependent nearest-neighbor thermodynamic parameters (see section 2.3.2). Fraction of hybridized duplexes The fraction of hybridized oligonucleotides Fb [Koe05] (fraction bound) is a quantity which is directly accessible from experiments (e.g. via the hybridization signal intensity in microarray assays or via the hypochromicity in UV-absorption-based measurements). Fb can be derived from thermodynamic quantities (e.g. via the equilibrium constant K). Fb =[D] min([S1]0,[S2]0)(2.8) [S1]0and [S2]0are the initial concentrations of single-stranded species S1and S2. How Tm,∆G◦ Dand Fb are related and influenced by experimental parameters (duplex length, sequence composition, defects, salt concentration, nucleic acid concentration and temperature) is well discussed in [Koe05]. If the fraction bound F b is compared to microarray hybridization signals one needs to consider that microarray hybridization is affected by many parameters, which are not accounted for in the simple model described above. 25
Fundamentals conformations an open link can adopt [Gib59]. In C. Kittels double-ended zipper model [Kit69] the zipper is consisting of N bonds (corresponding to the base pairs) that can only be opened from the ends. The partition function is determined by summation over the statistical weights of all partially unzipped duplex states. With the partition function the statistical mechanics of the duplex, e.g. the average number of open bonds (corresponding to the degree of partial duplex denaturation) is accessible. Kittel showed that the assumed degeneracy of partially unzipped duplex states (arising from rotational freedom of unpaired nucleotides) - in DNA this degeneracy may be on the order of 104gives rise to a melting transition in the quasi-one-dimensional system.7No phase transition can occur in the non-degenerate case (when the number of rotational degrees of freedom equals 1). Zocchi et al. [Zoc03] reported that a zipper-model based on end-domain opening describes well the temperature dependence of the average number of unzipped base pairs determined in UV absorption experiments. However, they also report that their analysis of transition parameters indicates that, apart from end-domain opening, bubble formation is also important for the denaturation process. Deutsch et al. [Deu04] employed the double-ended zipper model for a statistical mechanics based description of microarray hybridization signals. End-unzipping of the duplexhas also been assumed by Ambj¨ornssonand Metzler [Amb05] for a model to investigate the blinking dynamics of molecular beacons (fluorophore-quencher pair included in a fraying duplex section). Base pairs at the duplex ends are stabilized by stacking interaction with only one neighboring base pair, whereas base pairs in the interior of the duplex are stacked between two neighboring base pairs. The stabilizing stacking interactions from both sides prevent internal denaturation. Therefore unzipping is (largely) restricted to the duplex ends (end fraying) as shown in Fig. 2.14A. Structural constraints arising from the double helix structure may impose further restrictions to internal bubble formation. The influence of the helical structure on duplex stability is, however, not well understood. Denaturation bubbles The above statements, however, do not apply to the denaturation of long duplexes. These denature via theformation ofdenaturationbubblesin the interiorof the duplex(see Fig. 2.14B and C). This is due to several reasons: •due to an exponential decrease of the base pair dissociation probability towards the 7Cuesta and Sanchez [Cue04] discuss why Van Hove’s theorem (simply interpreted: ”No phase transitions occur in 1D particle systems with short-range pair interactions”) doesn’t apply to the melting transition of nucleic acid duplexes. 32
Nucleic Acid Hybridization center of the duplex, end-domain opening is restricted to duplex-ends ⇒thus, long duplexes can only denature via the formation of denaturation bubbles. •occurrence of relatively weakly bound (AT-rich) subsequences in a long duplex •increased melting temperatures of long duplexes ⇒the increased entropy contribution (−T∆S) results in destabilization of the nearest neighbor interactions ∆GNN A C B Figure 2.14: Denaturation of short duplexes (A) occurs mainly via end-domain opening. In long duplexes (B) end-domain opening does’t extend into the middle of the duplex. Rather, denaturation bubbles, forming at weakly bound sections in the interior of the duplexes, propagate and (C) merge with the open end-regions. At increased temperatures denaturation bubble formation leads to dissociation of long duplexes. The relevance of internal denaturation bubble formation depends on duplex length and, in particular, on the individual sequences (i.e. on the distribution of more/less stable NN pairs). To provide a coarse estimate: for duplexes with l < 15 base pairs end-fraying is expected to be the prevailing mode of nucleic acid denaturation, vice versa, for long and intermediate size duplexes with l > 100 base pairs bubble formation is expected to be relevant or more important than end-domain opening [Blo03].8However, Zocchi et al. [Zoc03] reported that denaturation bubbles may be relevant also in the denaturation process of short duplexes. 8Blossey et al. [Blo03]: ”On rather short DNA sequences (∼100 bp’s) the loop entropy contribution is not very important as loops are rare and short and the DNA denatures mainly through unbinding from the edges. A description based on the 1D Ising model with appropriate experimentally determined energy parameters is therefore sufficient [...].” 33
Fundamentals 2.3.4 Further Models of the DNA Melting Transition Further well-established models for the DNA melting transition are the Poland-Scheraga (PS) model [Pol66] and the Peyrard-Bishop-Dauxois (PBD) model [Dau93]. The Poland-Scheraga model describes the helix-coil transition in long polynucleotide duplexes. The duplex comprises alternating double-helical segments and denaturation bubbles. The PS model is essentially a one-dimensional Ising-model. Consideration of the various bubble configurations gives rise to an entropic term. This results in an effective long range interaction, so that in the PS model a phase transition may occur [Blo03]. The PBD model represents a Hamiltonian approach. In the PBD model cooperativity effects - arising from anharmonic nearest neighbor stacking interactions - result in a distinct melting transition. An overview on theoretical models of the nucleic acid melting transition is provided with reference [Zho06]. Further reading on the DNA melting transition: •thermal denaturation of DNA [War85] and DNA oligomers [Zoc03] •DNA breathing dynamics [Amb06] •zipper models [Kit69; Iva04] •end-denaturation [Amb05] •mismatches and bubbles [Zen06] •thermodynamic properties of DNA sequences [Koe05] •further related publications [vE06; Eve07] 2.4 Destabilization of Oligonucleotide Duplexes by Point Defects A high discrimination capability between similar sequences is important in genotyping applications, where single nucleotide polymorphisms(SNPs), variations of single bases, are the subject of interest. SNPs largely determine genetic individuality, but also disposition to genetically caused diseases or response to medicaments, and are therefore of great interest not only for genetic research but also for medical diagnostics and therapy. SNPs can be detected (using DNA microarrays) by hybridization with short oligonucleotide probes. Already a single mismatching (MM) base pair (owing to the SNP) can result in a significant decrease of duplex stability [Nel81; Pat82; Con83]. The impact of a MM base pair on 34
Destabilization of Oligonucleotide Duplexes by Point Defects duplex binding affinity is is determined by the length of the duplex [Koe05], the type of mismatch base pair [All97], the influence of neighboring bases [All97] and by the position of the defect (with respect to the duplex ends) [Wic06; Poz06; Nai06b]. In this study we also investigate single base bulges, another type of point defect, originating from single base insertions and deletions. The insertion of a surplus (unpaired) base into one of the duplex strands results in a small bulge in the regular duplex structure. Similarly a single base deletion creates a bulged base in the opposite strand. Like single base mismatches base bulges can significantly reduce duplex binding affinity. 2.4.1 Single Base Mismatches Figure 2.15: Structure of T·G mismatches in a B-DNA duplex (X-ray diffraction data 113D.pdb [Hun87]). Green arrows indicate the T·G mismatches. Structural investigations (NMR and X-ray studies) have shown that single mismatch base pairs (see Fig. 2.15) mismatches introduce little overall structural distortion on the double helical duplex structure [Hol91; Cog91; Ske93]. Consideration of single base mismatches in the nearest neighbor model The nearest neighbor model has been extended beyond Watson-Crick base pairs to include single base mismatch (MM) defects [All97; San04]. From UV melting experiments Allawi et al. [All97] have established a complete database of MM single base MM thermodynamic parameters for DNA/DNA duplexes. The (mostly) destabilizing MM propagation 35
Fundamentals parameters (a complete table is provided in [San04]) are used for duplex free energy calculations just like the Watson-Crick propagation parameters. Most destabilizing MM nearestneighbor pairs are AC/TC (∆G◦ 37=1.33 kcal/mol), TC/AA (∆G◦ 37=1.33 kcal/mol), TC/AC (∆G◦ 37=1.05 kcal/mol) and GT/CC (∆G◦ 37=0.98 kcal/mol). Least destabilizing are GG/CG (∆G◦ 37=-1.11 kcal/mol) and GT/CG (∆G◦ 37=-0.59 kcal/mol). An order of DNA/DNA base pair stabilities (based on [All97]) is provided in [San04]: G·C>A·T>G·G>G·T≥G·A>T·T≥A·A>T·C≥A·C≥C·C The study of Allawi et al. [All97] also reveals a strong impact of closing base pairs (the base pairs enclosing the MM base pair) - closing C·G base pairs are more stabilizing than A·T base pairs. The two-state nearest neighbor model doesn’t account for the MM position within the duplex sequence. According to SantaLucia [San04] ”[...] with the exception ofthe terminal and penultimate positions, the thermodynamics of a given mismatch in a given context is independent of its position in a duplex, contrary to common opinion”. This, however, is not in agreement with recent observations of a strong influence of defect position on duplex binding affinity [Kie99; Dor03; Wic06; Poz06; Nai06b]. 2.4.2 Single Base Bulges Defects originating from insertion or deletion of a base result in bulged duplexes as shown in Fig. 2.16. Base bulges are a frequent structural motif in RNA structures e.g. in tRNA and rRNA. It is assumed that bulges may play a role in nucleic acid-protein binding [Wu87]. Single bulged bases can adopt looped out (Fig. 2.16) or intrahelically stacked conformations [Yoo01; Bar06]. According to Woodson and Crothers [Woo88] ”[...] evidence from several laboratories suggests that extrahelical purines are generally stacked into the helix, while extrahelical pyrimidines are in equilibrium between stacked and unstacked states [...]” (in this context ”extrahelical base” has the meaning ”bulged base”). The thermodynamics of bulged duplex was first investigated by Fink and Crothers [Fin72]. They reported a destabilizing free energy (25◦C) of 2.8 kcal/mol for a single base bulge. Wartell and coworkers [Ke93; Ke95; Zhu99] investigated the thermodynamics of single base bulges on a larger number of DNA and RNA sequence motifs. The relative stability of bulged RNA duplexes was investigated in temperature gradient gel electrophoresis (TGGE) experiments. For RNA bulges they report an unfavorable free energy (with respect to the bulge-free reference duplex) δ∆G◦ 37 between 2.85 and 4.8 kcal/mol. Wartell and coworkers observed, that the stability of bulged duplexes is increased if the 36
Destabilization of Oligonucleotide Duplexes by Point Defects Figure 2.16: Single base bulge (looped out cytosine base, shown in yellow) in a rRNA helix structure (X-ray diffraction data 1DQF.pdb [Sun00]). bulged base has at least one identical neighboring base. They categorized bulged duplexesdepending on the identity of the bulged base and the duplex sequence - in two groups: •Group I: the bulged base has no identical neighboring bases •Group II: the bulged base has at least one identical neighboring base According to [Zhu99] the local average free energy contribution of a DNA base bulge can be expressed as: ∆G◦ 37,(XNZ)·(X0−Z0)= 2.72 kcal/mol + 0.48∆G◦ 37,(XZ)·(X0Z0)+δg (2.18) For the free energy of RNA single bulges a similar relation was derived [Zhu99]: ∆G◦ 37,(XNZ)·(X0−Z0)= 3.11 kcal/mol + 0.40∆G◦ 37,(XZ)·(X0Z0)+δg (2.19) Notation: The unpaired base N is enclosed by the base pairs X·X’ and Z·Z’. ∆G◦ 37,(XZ)·(X0Z0)is the stacking energy of the base pair doublet (XZ)·(X0Z0). The stabilizing contribution for degenerate Group II bulges δg is -0.4 kcal/mol for DNA and -0.3 kcal/mol for RNA (in both cases δg=0 kcal/mol for Group I bulges). 37
Fundamentals AGGCGTACGTA GTTTCCAGAG TCCGCATGCAT CAAAGGTCTC C AGGCGTACGTA GTTTCCAGAG TCCGCATGCAT CAAAGGTCTC AGGCGTACGT AGTTTCCAGAG TCCGCATGCA TCAAAGGTCTC A A A B Group I base bulge Group II base bulge degenerate conformation Figure 2.17: Positional degeneracy of base bulges. (A) Group I bulge. The nondegenerate bulged base (C, shown in grey) has no identical neighbor bases. (B) Group II bulge. The bulged base A has an identical neighbor, giving rise to positional degeneracy [Ke95] of the bulge conformation. The increased number of possible bulge conformations (here two rather than only one in A) represents an increase in entropy, resulting in a stabilization of the degenerate Goup II bulge with respect to the nondegenerate Group I bulge. The experimentallyobservedfree energy differencebetween Group Iand degenerateGroup II bulges of -0.4 and -0.3 kcal/mol (for DNA and RNA, respectively) is in good agreement with the simplified entropic estimate for a two-position degeneracy of -RT·ln(2)=-0.43 kcal/mol (at 37◦C) [Zhu99]. Znosko et al. [Zno02] report an increased stability of pyrimidine single bulges with respect to purine single bulges (0.4 kcal/mol on average). This study, based optical melting experiments (UV absorption) on RNA duplexes, provided different equations (written here in the notation of [Zhu99]) for the bulge free energies of pyrimidines (eqn. 2.20) and purines (eqn. 2.21). ∆G◦ 37,(XNZ)·(X0−Z0)= 3.9 kcal/mol + 0.10∆G◦ 37,(XZ)·(X0Z0)+δg.(2.20) ∆G◦ 37,(XNZ)·(X0−Z0)= 3.3 kcal/mol −0.3∆G◦ 37,(XZ)·(X0Z0)+δg (2.21) Here, δg is 0 and -0.8 kcal/mol for Group I and Group II bulges, respectively. The reported stabilization of Group II bulges δg=-0.8 kcal/mol is significantly larger than the previously reported stabilization from [Zhu99] (δg=-0.3 to -0.4 kcal/mol), thus raises questions about the mechanisms underlying Group II bulge stabilization. Turner [Tur92] suggested that the stability of a bulged duplex could depend on the proximity of the bulge with respect to the helix end. Znosko et al. [Zno02] didn’t find evidence for an influence of bulge position on duplex stability. 38
Destabilization of Oligonucleotide Duplexes by Point Defects 2.4.3 Influence of the Defect Position Kierzek et al. [Kie99] investigated the effect of the position of a single mismatch within short RNA duplexes (optical melting experiments). They observed that ”[...] moving the position of the mismatch toward the end of the helix enhances the stability for U·U and A·A mismatches by ∼0.5 kcal/mol per each position closer to the helix end [...]”. For A·A mismatches the observed trend is less obvious than for U·U mismatches, G·G mismatches were found to be insensitiveto the position within the helix. Since the study was performed with heptamer duplexes (enabling the comparison of three MM positions) the data base for the observed MM positional influence is rather limited. Dorris et al. [Dor03] observed a similar positional influence for 2-base and 3-base mismatch probes (with respect to cRNA targets) on CodeLinkTM 3D gel arrays. They also report a strong correlation (including the positional influence) between solution-phase melting temperatures and microarray hybridization signals of the MM duplexes. Recent microarray studies [Wic06; Poz06; Nai06b; Nai06a] using extensive sets of probe sequences have shown a very distinct influence of mismatch position and bulge position [Nai06a], respectively. The discrimination between MM and PM is significantly more distinct for defects near the center of the duplex than for defects near the duplex ends. Interestingly, from solution phase hybridization studies (apart from [Kie99] and [Dor03]) an influence of defect position is not been reported. In the nearest neighbor model only terminal and penultimate MM positions are considered to be less destabilizing than MMs in the interior of the duplex [Pey99; San04]. It is not clear whether the positional influence has been overlooked in previous solution-based studies or, if different experimental conditions are the reason, why a distinct positional influence has only been described recently, typically for microarray-based experiments. Typical characteristics of studies not reporting an influence of defect position [Ke95; All97; Pey99; Sug00]: •mostly solution-phase hybridization •presynthesized oligonucleotide probes (thus containing a negligible fraction of synthesis defects) •small probe sets (<100 probes) investigated •the defect is typically restricted to one or few positions (commonly in the center of the duplex), no systematical variation of the defect position •in most studies rather short duplexes ≤10 bp (little margin for variation of defect position) were employed 39
Fundamentals •experimental method: measurement of the melting curves by UV absorbance spectroscopy (for an assumed two-state melting transition the measured fraction of dissociated base pairs is equal to the fraction of dissociated duplexes) •duplex free energies are derived from melting curve analysis According to [Pey99] binding affinity contributions of mismatches more than three base pairs from the end are independent of the position.9[Pey99]: ”Consequently, it can be concluded that the nearest-neighbor model is a good approximation for both Watson-Crick pairs and all single mismatches.” Typical characteristics of studies reporting an influence of defect position [Ura02; Dor03; Wic06; Poz06; Nai06b]: •mostly microarrays studies •microarrays in [Wic06; Poz06; Nai06b] are fabricated by in situ synthesis - probes can therefore contain a considerable amount of synthesis defects •duplex length between 16 and 25 bp •experimental method (typically): measurement of microarray hybridization signals (mostly fluorescence intensity) •measurement of binding affinity variations depending on defect type, defect position and closing base pairs. The PM/MM hybridization signal ratio is a direct measure for the MM discrimination. •microarray studies are favorable for large scale systematic investigations of MM discrimination (improved statistics - many different sequences, ”direct comparison” of binding affinities obtained in the same experiment) The positional influence appears to be most pronounced in [Wic06; Poz06; Nai06b]. However, this may be owing to the fact that the experimental design of these particular studies enables a more systematic and extensive investigation of the position dependence than the other studies. Experimental results in [Kie99] and [Dor03] indicate that a positional influence is not limited to microarray studies but can be observed in solution-phase hybridization studies as well. 9This particular study was performed with a relative small set of 51 relatively short 9-12mer duplexes. 40
Solid-Phase Synthesis of Nucleic Acids Modeling of the positional influence Pozhitkov [Poz02] considered mismatch positional influence empirically in an algorithm for finding specific oligonucleotide probes for species identification. Binder [Bin06] tries to explain the positional influence with a zipper model in which the mismatch affects the base pairing of Watson-Crick base pairs in the duplex section between the MM and the duplex end. Therefore, the impact of a mismatch on duplex stability is getting smaller as its position is closer to the duplex end. However, the assumed base pair opening probability (as shown in Fig. 10 in [Bin06]) doesn’t account for the fact that end fraying under hybridization conditions is largely confined to the two [And06] or three [Lei92] outermost base pairs. Like Binder we use a zipper based model in our analysis, however, we account for the fact that the end fraying is largely restricted to the outermost base pairs and that the base pair opening probability is exponentially decreasing towards the center of the duplex. Partial denaturation of inner base pairs is considered as a rare stochastic event. 2.5 Solid-Phase Synthesis of Nucleic Acids In molecular biosciences synthetic nucleic acid sequences are employed in many of applications. For example, as primers for the amplification of DNA sequences by polymerase chain reaction (PCR), as target-specific probe molecules on DNA microarrays (or in fluorescent in situ hybridization within biological specimens), or as double-stranded RNAs for gene silencing in RNA interference applications. Synthetic nucleic acid sequences are commonly produced in a solid-phase synthesis approach. 2.5.1 Principles of Solid-Phase Chemical Synthesis Solid-phase synthesishas first been employed for the fabrication of polypeptidesequences10 [Mer63]. In solid-phase synthesis the polymer-chains to be synthesized are end-tethered to a solid substrate. This enables efficient separation of uncoupled building blocks (in solution) from the surface-tethered synthesis products, after a synthesis step has been completed. Coupling of monomer building blocks (see Fig. 2.18) is performed via reactive terminal groups. A removable chemical protection group prevents uncontrolled polymerization of 10 For the development of the solid-phase polypeptide synthesis R.B. Merrifield received the Nobel Prize in chemistry in 1984. 41
Fundamentals mRNA redfluorescent cDNA targets mRNA mRNA isolation microarray hybridization reverse transcription labelling cancercells normalcells combine targets greenfluorescent cDNA targets Figure 2.23: Dual color microarray experiment. In this example the expression profile of cancer cells is compared to a reference sample of normal cells. Complex mixtures of messenger RNAs (mRNAs) are isolated from each sample and fluorescently labeled via reverse transcription labeling. The targets from the cancer cell are labeled with a green fluorescent dye, whereas the targets from the reference sample are labeled with a red fluorescent dye. The targets are combined and hybridized on the same microarray. Analysis and comparison of the two color-channels enables identification of upand down-regulated genes. (Adapted from Wikipedia: http://en.wikipedia.org/wiki/DNA microarray) Figure 2.24: Single Nucleotide Polymorphisms (SNPs) are genetic variations of single base pairs between members of the same species, or even between the two copies of a chromosome pair. The DNA strand in 1 differs from the DNA strand in 2 by a single base pair. Genotyping assays enable highthroughput screening for single nucleotide polymorphisms. (Source: Wikipedia, http://en.wikipedia.org/wiki/Single nucleotide polymorphism) 48
DNA Microarrays 2.6.2 The Development of DNA Microarray Technologies An early method (1975) for the analysis of complex nucleic acid mixtures is the Southern blot [Sou75]. Thereby the mixture of unidentified DNA fragments (targets) is separated by gel electrophoresis, transferred and immobilized on a flexible nylon membrane. For identification of the targets radioactively or chemically labeled probes (with well-known sequences) are incubated with the membrane, thus enabling hybridization with the complementary target sequences. The so-called dot blot is a similar technique, in which the (unseparated) target sample is directly applied onto the membrane as ”dots”. After fixation the identification of the target sequences is performed by hybridization with a labeled probe sequence (or a mixture of labeled probes). Miniaturization and parallelization have evolved the dot blot into the high throughput macroarray technique. With the help of automated methods several thousand millimetersized nucleic acid spots can be immobilized on a nylon membrane (typically 10 to 20 cm in size). Here, different from the blotting techniques described above, the known probe sequences (e.g. cDNA or synthetic oligonucleotide probes) are immobilized on the solid substrate, whereas the targets are applied in hybridization solution. Autoradiographic analysis and the large quantity of probe material provide a high sensitivity. However, radioactive labeling with 32Por 33P(requiring precautious handling) and the need for large quantities of probe and target material are serious disadvantages of the macroarray technique. By using rigid substrates rather than flexible nylon membranes, a significant miniaturization was achieved, giving rise to DNA microarray technology. Microarrays are commonly produced on chemically functionalized glass substrates - frequently a microscope slide format is employed. The use of glass substrates, which, unlike nylon-membranes, have low auto-fluorescence, enables highly sensitive detection of fluorescently labeled targets. Different types ofDNA microarrays have been developed in severalindependent approaches: •In 1995 Schena et al. [Sch95] reported the first gene expression assay on a printed microarray. They employed a contact printing technique for deposition of tiny spots (about 0.1-0.2 mm in diam.) of nucleic acids probes (cDNA probes) on a chemically functionalized glass substrate. This now widely-used technique is also known as spotting. The spotting solution with the prefabricated nucleic acid probes is deposited on the surface by a pin. A capillary gap at the tip of the pin releases a small (and reproducible) amount of the spotting solution when the pin is touching the substrate surface. Chemical functionalization of the substrate (e.g. with amino-, epoxyor aldehyde-groups) and the probe molecules (e.g. by attachment of an amino group) 49
Fundamentals enable fixation (immobilization) of the probes. cDNA microarrays are mainly used in gene expression assays. Apart from cDNA and PCR products, presynthesized oligonucleotide probes can be immobilized on microarrays (→oligonucleotide microarray). Microarray robots (arrayers) are commonly employed for a fully automated fabrication process. •Already several years earlier Fodor et al. [Fod91] developed a photolithographically controlledcombinatorialchemistryapproach for thefabrication of high-densityoligonucleotide microarrays.11. Owing to the similarity of the photolithographic fabrication process with semiconductor fabrication techniques, these microarrays are commonly called DNA chips. Unlike the spotting approach the light-directed in situ synthesis approach doesn’t require prefabricated probes for deposition. The probe molecules (DNA oligonucleotides) are fabricated in situ, i.e. nucleotide by nucleotide, on the microarray substrate. The massively parallel synthesis of up to a million different sequences on the same chip is directed by UV light exposure. In the combinatorial synthesis process chrome masks provide a sequence specific exposure scheme (spatially restricted to the particular microarray features) to control the sequence of nucleotide couplings for each probe sequence individually. •Ink-jet techniques (based on piezoelectric deposition) are used for in situ synthesis of microarrays [Bla96] (by deposition of phosphoramidites) and also for spotting of presynthesized DNA [Sch98]. •A rather novel technique is the electrochemical in situ synthesis of DNA microarrays [Mau06]. Thereby nucleic acid coupling is controlled by acid generation on a CMOS addressable electrode array. Depending on the type of probes employed DNA microarrays (not to be confused with other types of microarrays, e.g. protein microarrays) can be categorized into two groups: cDNA microarrays This type of microarray comprises immobilized cDNA probes or PCR products. Owing to the availability of cDNA and PCR products from biological sources, cDNA arrays (microarrays and macroarrays) are frequently prepared by biological labs. Since the long probe sequences (typically one hundred to several hundred nt long) are not suitable for discrimination between similar sequences (e.g. for the identification of single base MMs) the application of cDNA microarrays is restricted to gene expression profiling. 11 ”High density” refers to a high density of microarray features (up to one million per cm−2) 50
DNA Microarrays Oligonucleotide microarrays Oligonucleotidemicroarrays comprise synthetically fabricated probesequences which are typically between 15 and 100 bases long. Unlike cDNA microarrays, oligonucleotide microarrays enable discrimination of very similar genes belonging to the same gene family. Short oligonucleotides (<30 nt) owing to their high discrimination capability are used for genotyping and resequencing applications. Long oligonucleotide probes (∼60 nt) have the advantage of providing a high sensitivity for the detection of low abundance transcripts. Oligonucleotide microarrays are fabricated by immobilization (spotting) of presynthesized oligonucleotides, or by in situ synthesis. A detailed overview of microarray types and fabrication methods is provided in [Gao04]. 2.6.3 Characteristics of Microarray Hybridization Literature reports a large discrepancy between hybridization characteristics in bulk solution and on the microarray surface [Hel03; Bin06; Poz06]. Since NN thermodynamic parameters were determined in solution-phase experiments - under ”ideal hybridization conditions”, the nearest-neighbor model doesn’t necessarily perform satisfactory for the prediction microarray binding affinities. According to Bhanot et al. [Bha03], the loss of translational energy and entropy of microarray-bound probes (with respect to hybridization of free strands in bulk solution), and the constraint that targets can approach the probes only from one half-space, is independent of the sequence. Thus, with respect to bulk-solution, hybridization equilibrium constants, equilibrium constants for microarray hybridization are multiplied by the same sequenceindependent factor. The difference between solution-phase and surface-phase hybridization is of little consequence for specificity and sensitivity when equilibrium is achieved. However, hybridization kinetics (which is different for surfaceand solution-phase hybridization) has a pronounced effect on specificity and sensitivity [Bha03]. Levicky and Horgan [Lev05] reviewed physicochemical aspects of DNA microarray hybridization. In particular they discussed differences with respect to solution-phase hybridization. On DNA microarrays (with respect to solution hybridization)melting temperatures [Hel03] are significantly reduced. Additionally, significantly broadened hybridization isotherms (deviating from Langmuir-type characteristics) [Bin06] are observed. Moreover, on DNA microarrays a strong influence of the position of single base MMs on duplex binding affinities [Wic06; Poz06; Nai06b] is observed. 51
Fundamentals An ”ideal microarray” (in terms of specificity and target quantization) would require the following characteristics: •target-specific hybridization: i.e. probes hybridize only with the complementary target species •there is no intra-strand base pairing leading to formation of probe or target secondary structures. •all probe-target pairs have approximately the same binding affinity •there is a simple (e.g. linear) relation between the hybridization signal and the concentration of the corresponding target sequence Real microarrays deviate from the ”ideal microarray” (above) in several aspects: •complex target mixtures give rise to competitive hybridization processes [Poz06] (unspecific target/target and probe/target cross hybridization) •the use of long relativelylong target sequences (typically between 100 and several hundred nt long) results in target secondary structure formation and increased potential for cross hybridization. Both processes compete with the specific probe-target hybridization. Target secondary structure can prevent probe-target hybridization, thus leading to false negatives. Unspecific cross-hybridization can result in false positives. •surface effects (e.g. electrostatic effects and sterical hindrance) can (in conjunction with the varying length/secondary structure of individual targets) affect the quantitativeness of the measurement (binding affinity is a function of the amount of hybridized targets,target length and target structure) •probes are confined to a small area on the microarray (⇒diffusion-limitation effects) •synthesis defects (originating from in situ synthesis) affect binding affinities •labeling of the target sequences (e.g. with large fluorescent dye molecules like Cy3 attached at random positions) may affect binding affinities Microarray hybridization - a diffusion driven process Microarrays are often fabricated on microscope slides with dimensions of about 75 mm× 25 mm. The hybridization solution (ten to several hundred µl) is inserted into the gap between the microarray and a cover glass, thus forming a thin film with a thickness of 20100 µm. This is better illustrated by the following comparison in which the microarray is assumed to be enlarged to the size of a football field. On this scale, the liquid film corresponds to a puddle between 2 and 10 cm deep. The size of a microarray features may be visualized by a soccer ball. Hybridization in such a configuration is a slow process since diffusion is the dominating transport mechanism for the targets. The hybridization of a target with the corresponding probe is usually limited by the slow diffusion process [Pap06]. With the Einstein52
DNA Microarrays Schmoluchowki relation we find that the average distance a target molecule (with a molecular diffusion coefficient of 10−11 m2/s [Pap06]) is traveling in an overnight hybridization is approximately 1 mm. In microarray assays, by diffusive transport alone, the equilibrium can’t be reached on a reasonable time scale. Novel chaotic micromixingtechniques, e.g. based on surface acoustic waves (SAW) [Toe03], can very efficiently generate microagitation in the capillary gap und thus overcome the diffusion limitation. By scaling down the dimensions of the microarray (⇒increased ratio between the diffusion coefficient and the microarray surface) the hybridization equilibrium can be reached on a realistic time scale [Dan07]. 2.6.4 Further Reading on the Technical and Physical Principles of DNA Microarrays •Sensitivity, specificity, cross hybridization (→detection of false positives) [Bha03; Bin06] •Point defects (mismatches [Dod77; Wal79; Nel81; All97], base bulges [Ke95; Zhu99; Zno02]), influence of defect position [Dor03; Wic06; Poz06], discrimination capability [Ura03; Lee04] •Secondary structure of probes and targets (→detection of false negatives) [Lue03; San04] •Microarray fabrication (immobilization, in situ synthesis) [Sch99; Sch02; Gao04] •Qualityof the probesequences - synthesisdefects [Gar02; Job02; Ric04; Bin06] →heterogeneity of binding affinities •Target preparation [Sch02] (target length, fluorescent labeling, composition of the target mixture, type of nucleic acid target - DNA or RNA) •Competitive effects [Bin06] •Surface density of the probes [Pet02; Wat00; Lev05] (steric hindrance [Hal05], electrostatic repulsion [Vai02; Bin06]) •Attachment of the probes [Sch02], linker/spacer[Bei99], linear and dendrimeric linkers [Cam06] •various hybridization parameters [Sch02] (e.g. ionic strength, temperature, pH, blocking reagents) [Koe05] •Washing characteristics [Poz07] •Microarray size [Dan07], diffusion-limited target transport [Pap06], mixing [Gut05; Toe03] 53
Fundamentals 2.7 DNA Chip Fabrication by Light-Directed In Situ Synthesis Light-directed in situ synthesis of DNA microarrays was developed around 1990 by Fodor and coworkers [Fod91]. Short (typically ≤25mer) oligonucleotide probe sequences are synthesized nucleotide by nucleotide on the surface of the microarray. Spatially addressable photo-deprotection enables a massive parallel synthesis of arbitrary DNA probe sequences on a single microarray. Light-directed in situ synthesis is basically a solid-phase synthesis process (see section 2.5.1), requiring phosphoramidite reagents with photo-labile protection groups. Spatially controlled photo-deprotection is achieved with a photolithographic process and the use of phosphoramidite reagents with photolabile protection groups. Probe sequence information and microarray geometry is encoded in the photomasks. Today, commercial high density oligonucleotide microarrays (fabricated with high resolutionphotomasks)haveup to6.5millionprobes (with a feature size of 5µm). Light-directed in situ synthesis can also be employed for the synthesis of polypeptide sequences [Fod91] (→protein microarrays) or other combinatorial chemistries. 2.7.1 Photolithographic Control of the Combinatorial Synthesis Process For parallel synthesis of different probe sequences spatial control of the phosphoramidite coupling reaction is required. This is achieved by a spatially controlled photo-deprotection of the photolabile 5’-protection group (chemical structure shown in Fig. 2.19). The photocleavage generates a hydroxy-group at the 5’-ends of the exposed sequences and thus determines where on the microarray (i.e. at which microarray features/probe sequences - see Fig. 2.25 h and k) the next phosphoramidite building block (provided in the subsequent coupling step) will elongate the sequence. The fabrication of a microarray comprising arbitrary N-mer sequences requires 4×N deprotection/coupling steps. It is necessary to provide all coupling alternatives (X=A, C, G and T) in each ”nucleotide layer”. Thus, the light-directed combinatorial synthesis comprises a series of 4×N photo-deprotection DXiand associated nucleotide coupling 54
DNA Chip Fabrication by Light-Directed In Situ Synthesis steps CXi: 1.DA1/CA1→DC1/CC1→DG1/CG1→DT1/CT1 2.DA2/CA2→DC2/CC2→DG2/CG2→DT2/CT2 ⇒... N. DAN/CAN→DCN/CCN→DGN/CGN→DTN/CTN Spatially controlled photo-deprotection is shown in more detail in Fig. 2.25 where each probe strand symbolizes an individually addressable microarray feature (whereas in reality each feature comprises millions of identical probes). Assuming a stepwise coupling efficiency fcthe yield Y=fcNof probes which are free of synthesis defects is decreasing with the power of the probe length N. Probes containing defects cannot be repaired or removed as in common solid-phase synthesis (capping, truncation, HPLC separation). Synthesis defects (i.e. single base mismatches, insertions anddeletions) will therefore affect microarray hybridization [Job02]. The length of the microarray-probes is determined by the application. Shorter 15-25mer probes provide a high discrimination capability between PM and MM and are therefore suitable for SNP detection and resequencing assays. Longer probes are less discriminative but rather more sensitive(increased binding affinity), and are therefore favorable for detection of low abundance mRNAs in expression profiling applications. Typically probes on high density oligonucleotide microarrays have a length of ≤25 nt, however, the fabrication/application of arrays with longer 40-60 nt probes has also been reported. The synthesis cycle For light-directedinsitusynthesisphotolabilephosphoramiditereagents, α-methyl-6-nitropiperonyloxycarbonyl (MeNPOC) [Pea94; McG97] or [2-(2-nitrophenyl)-propyloxycarbonyl]-2’-deoxynucleoside (NPPOC) phosphoramidites [Has97] are used. The MeNPOC-chemistry (employed in the fabrication of Affymetrix GeneChips R ) has a stepwise yield of 92 to 94% [McG97]. Significantly better coupling yields have been reported for NPPOC phosphoramidites [Bei99]. Nuwaysir et al. [Nuw02] reported stepwise chemical yields between 96 and 98%.12 Use of NPPOC phosphoramidite reagents (chemical structure shown in Fig. 4.1) has been reported in [Has97; Bei99; Nuw02; Bau03; Wol04; Woe06]. MeNPOC phosphoramidites reagents were used in [Pea94; McG97; SG99; Lue02]. 12 Stepwise synthesis yields of NPPOC phosphoramidites according to [Nuw02]: NPPOC-A(tac) 96%, NPPOC-C(ibu) 99%, NPPOC-G(ipac) 97%, NPPOC-T 98% 55
Fundamentals Figure 2.25: Light-directed in situ synthesis of DNA microarrays [Fod91]. The spatially controlled combinatorial chemistry approach enables parallel synthesis of arbitrary probe sequences. In the (non-optimized) coupling scheme shown here, the probe sequences are synthesized ”layer by layer”: to cover all coupling-alternatives, in each ”nucleotide layer” phosphoramidite-couplings are performed in the order A, C, G, T. In the above series of images each probe strand symbolizes an individually addressable microarray feature (whereas in reality each feature contains millions of probes). Synthesis of the first nucleotide layer (a-f). (a) The substrate is initially functionalized with photo-labile protection groups (depicted as blue balls). Spatially controlled UV exposure (use of photomasks) is restricted to those feature areas where phosphoramidite building blocks are to be attached in the subsequent coupling step. (b) Photo-cleavage of the protection groups created hydroxyl-moieties, which are the binding sites for the subsequent adenosine-phosphoramidite coupling step (c). In the coupling step only one building block can attach to each deprotected strand. Further couplings are prevented by new protection groups (imported with the building blocks). (d) Photo-deprotection of those probes which require cytosine at the first base position. (e) Coupling of cytosine-phosphoramidite. The first nucleotide layer is completed after deprotection and coupling of G nucleotides (not shown) and T nucleotides (f). The second layer is synthesized upon the first layer (g-l). The deprotection/coupling scheme is continued until the final length of the oligonucleotide probes is reached (m). 56
DNA Chip Fabrication by Light-Directed In Situ Synthesis Owing to the the 5’-attachment of the NPPOC protection groups the in situ synthesis is performed in 3’→5’ direction. Therefore the probes are typically 3’-tethered at at the microarray surface. However, 5’-tethered probes can be synthesized (in 5’→3’ direction) with modified phosphoramidite reagents (carrying 3’-NPPOC protection groups) [Alb03]. 5’-tethered microarray probes are, unlike 3’-tethered probes, available for enzymatic modification. In the following we refer to the 5’-NPPOC phosphoramidite chemistry (Fig. 4.1) which has been employed in this work. The synthesis cycle in the light-directed synthesis process (Fig. 4.2) is very similar to the scheme employed for oligonucleotide synthesis on CPG-supports (shown in Fig. 2.20). Exposure with UV light (λ=350-380 nm) induces photo-deprotection and enables coupling of the next phosphoramidite building block. A capping step (as shown in Fig. 2.20), resulting in truncated strands rather than in strands containing single base MMs, is of a rather limited value in microarray synthesis (truncated strands cannot be removed) and is therefore omitted. Coupling and oxidation steps are performed in the same way as in the oligonucleotide synthesis on CPG-supports. Photo-deprotection of NPPOC results in short-lived intermediate states. According to [Wal01] an aci-nitro intermediate is in acid-base equilibrium with its anion. The unstable anion can fragment, thus resulting in the desired deprotection reaction (complete removal of the NPPOC group). However, via a competing reaction pathway, the aci-nitro intermediate can also form a nitroso product, which is not removed from the phosphoramidite residue, thus preventing photo-deprotection. To promote the desired reaction pathwaythe photo-deprotection needs to be performed in a solvent providing sufficient proton acceptors. Therefore the basicity of solvent acetonitrile is increased by addition of a mild base (e.g. piperidine [Bei99]). 2.7.2 Combination of ”Maskless” Digital Photolithography and Combinatorial Chemistry Light-directed in situ synthesis of DNA microarrays with high resolution photomasks has been developed and is employed on an industrial scale by Affymetrix Inc.. High costs for chromium masks, considerable technical effort for the mask alignment, and the lack of flexibility (a new set of photomasks is required for each new microarray design) have so far prevented lab-scale application of the photomask-based fabrication technique. The use of computer-controlled spatial light modulators as ”virtual photomasks” can circumvent the limitations related to the use chromium photomasks and thus provide great flexibility for custom microarray fabrication. 57
The Microarray Synthesizer Figure 3.2: (a) Photograph of the maskless microscope projection photolithography system (top view). Along the optical path (dotted white line): UHP lamp housing, UV cold mirrors, shutter, band pass filters (green and UV), DMD and driver electronics, tube lens, microscope (Zeiss Axiovert 135) and the reaction cell, which is mounted onto the sample holder. (b) Drawing of the lithography system: Ultra High Pressure lamp (UHP) powered by video projector (VP2), plano-concave silica lens (L1), plano-convex lens (L2), UV cold mirror F1, light trap (LT), plano-convex lens (L3), UV cold mirror (F2), shutter (S), bandpass filters for UV (F3) and green (F4) illumination, plano-convex lens (L4), fold mirrors (M1 and M2), DMD and driver electronics of the AstroBeam projector (VP1), tube lens (L5), infinity corrected microscope (ICM), mirror/beamsplitter-assembly (M3), 5×(0.25 NA) Fluar microscope objective (FO), substrate to be patterned (PS). Technical details are provided in Appendix B.3. 64
The Maskless Microprojection Photolithography System (MPLS) 250 W UHP lamp of another video projector (Optoma EP 758). Since UHP lamps require specialized power supplies (integrated in the video projector) the Optoma projector is now employed as a lamp power supply. Due to the requirement for high UV transmission we couldn’t use the highly optimized optics2of the video projector. For the photolithography system a new UV illumination optics had to be designed: the lamp module for the Optoma EP758 projector was built into an air cooled housing and connected via an extension cable to the lamp driver of the Optoma projector (VP2). The arc of the 250 W UHP lamp is located at the inner focal point of the elliptical lamp reflector. To efficiently collimate the strongly divergent beam, a planoconcave diffraction lens (L1) (f=50 mm, 25.4 mm diam., fused silica) is placed between the lamp window and the outer focal point of the reflector. Efficient filtering of the near UV wavelength band required for the photo-deprotection reaction proved to be difficult owing to the high thermal load. Filtering is therefore performed in several steps: A dichroic filter from the Optoma lamp module (F1) (originally designed as a UV protection filter) is employed as a UV cold mirror to cut down the visible light intensity to about 10 percent. UV light below 400 nm is efficiently reflected.3 Infrared radiation is filtered using another UV cold mirror (Oriel) (F2). Finally a band pass interference filter (F3) (bk-370-35-B, Interferenzoptik Elektronik GmbH) is used for selecting the wavelength band in the mercury i-line region (λ= 365 nm) required for the photo-deprotection reaction. Taking into account that the mercury i-line is considerably broadened due the high operation pressure of the lamp, we had to use a relatively wide band pass filter (peak transmission Tmax = 60% at 370 nm, FWHM: 33 nm) to achieve a sufficiently high UV transmission. Use of a broadband filter (color glass UG-5, Schott, transmission between 230 and 430 nm and above 650 nm, Tmax = 90% at 350 nm ) would result in severe chromatic aberration. 3.2.2 Digital Mask Projection Using a Digital Micromirror Device The DMD is a spatial light modulator commonly used for image generation in DLP video projection systems (for technical details on DMD technology see section B.1). In our setup we use a DMD with XGA resolution containing 1024×768=786432 square mirrors (16 µm in size with a pitch of 17 µm) that can be tilted by an angle of +10◦or -10◦relative to the 2Optimized for high light throughput and uniformity of illumination. 3The transmission spectrum of the dichroic filter shows a distinct cutoff at 415 nm - from 420 to 700 nm the transmission is ≥90%. The reflectivity in the i-line range couldn’t be measured with the spectrophotometer available. However, a simple experiment with a 100 mW UV-LED (Nichia NCCU033) shows that the UV reflectivity is (coarsely estimated) between 60 and 80%. 65
The Microarray Synthesizer normal axis of the chip. The two positions are referred to as onand off-state: Mirrors in the on-state reflect the incident light perpendicular to the DMD surface into the projection optical system, whereas mirrors in the off-state reflect light at an angle of 40◦relative to the DMD normal axis into a light trap (Figure 3.3). +10° -10° 40° 20° 20° ON OFF I P T I Figure 3.3: Spatial light modulation with a Digital Micromirror Device. Mirrors in the on-state (blue) reflect the incident light (I) in a direction normal to the DMD surface into the projection optics (P). Mirrors in the off-state (green) reflect the light under an angle of 40◦with respect to the normal axis into a light trap (T). The DMD is oriented perpendicular to the optical axis of the projection system. The micromirrors tilt around their diagonal axis. We have rotated the DMD by 45◦around the optical axis, so that the incident beam and the reflected beam lie both in the horizontal plane of the setup (Figure 3.4). Technical details on the modification of the DLP video projector are provided in Appendix B.2. For better accessibility of the micromirror array the DMD board had to be removed from the projector chassis and reconnected to the driver board via a 148 pin extension cable. Because the driver electronics of the projector remains unchanged, all sorts of video signals can be used to control the image display. Connection to a PC with a dual-head graphics card proved to be useful, as one screen can be used for control purposes (e.g. for running the DNA synthesis control program which automates and coordinates photolithographic pattern display and the fluidics system) while the other one is reserved for pattern display. 3.2.3 The Image Projection Optics To reduce the microarray size to a few mm2we opted for a microscope projection approach. Reduced dimensions of the microarray are beneficial for microarray hybridization due to reduced diffusion times [Dan07] and reduced material requirements (synthesis reagents 66
The Maskless Microprojection Photolithography System (MPLS) P T I a)b)c) Figure 3.4: Rotated DMD arrangement in the maskless microscope projection lithography setup. The DMD is rotated by 45◦around the optical axis, so that the tilting axis of the mirrors is vertical. The incident beam and the reflected beam lie both in the horizontal plane of the setup.(Left image) From the mirrors in on-position (arranged as a ”X”) the incoming light (I) is reflected towards the projection optics (P). Mirrors in the off-position reflect the light into a light trap (T). (Center image) View of the DMD from the projection optics. (Right image) View of the DMD from the light trap. and nucleic acid sample size). By reducing the image area, the illumination intensity is increased by a similar factor: a 250 W UHP lamp does suffice in order to keep the time required for optical deprotection in a reasonable relationship to the total turnover time of the chip synthesis. Use of the microscope also provides superior control of the image focusing and mechanical stability. Image drift occurring from thermal expansion of the optical parts has previously been described as serious problem in the light-directed synthesis process, requiring active control of focusing, e.g. by means of an image-locking technique [Ric04]. An important aspect in the design of the lithography system is image contrast. In lightdirected microarray synthesis stray light is much more critical than for example with photoresist. Photoresist, having a strong nonlinear exposure characteristics, doesn’t respond to small stray light intensities below a threshold value. In microarray synthesis there is no threshold and stray light induced errors can accumulate over many exposure steps. Within the total exposure time of about two hours, stray light causes base insertion errors, affecting most of the synthesized DNA strands. The whole synthesis process involves about 80 exposures with different mask patterns, it extends over about 6.5 hours. Mask alignment requires thermal and mechanical stability. To make use of the maximum pixel resolution of the setup (which is 3.5 µm with a 5×microscope objective) no movements caused by vibrations, tension release, or thermal expansion larger than about 1 µm (in the front focal plane of the objective)can be tolerated. 67
The Microarray Synthesizer Figure 3.5: The projection optics system. Incident light (I) (filtered - either near UV for the exposure or green for focusing); light trap (T); tube lens (TL); beam splitter (BL); microscope objective (MO); synthesis cell (S). The micromirror array (DMD) is placed in the image plane (located outside the microscope frame) of the inverted microscope. With infinity corrected microscope objectives, a tube lens (TL) is necessary to project the image of the DMD to infinity. The adjustment of the distance between DMD and tube lens, which does not exactly equal the nominal focal length of 164.5 mm (as specified by the manufacturer), is crucial for the calibration of the setup, as explained later (in Sec. 3.2.5). A movable half mirror/half beamsplitter optical element (BS), located at the position of the microscope’s fluorescence filter block, is used to reflect the light into the objective back aperture. Using the beamsplitter part, the light reflected from the surface of the microarray substrate can be coupled into the microscope. This is employed for exact focusing and direct observation of the projected image through the eyepiece. For photopatterning, the mirror part is used (exchange is achieved by sliding the plate by hand). In principle for this purpose a dichroic beamsplitter (reflection of UV light and reduced reflection of visible light) could be used. However, the use of a beamsplitter plate (for photo-deprotection) turned out to be problematic since even a small amount of reflection at the backside of the plate can produce ghost images, and thus significantly affect the image contrast. Among several objectives(MO) tested, we found the Zeiss Fluar 5×(0.25 NA) as mostsuitable for DNA chip fabrication, particularly for its superior UV transmittance and its large back aperture allowing for efficient light collection. Over a working distance of 12.5 mm the image of the DMD is projected onto the DNA synthesis substrate - a chemically functionalized glass surface - inside the synthesis cell (S). A 10×(0.30 NA) Plan Neofluar and a 20×(0.5 NA) Plan Neofluar objective (Zeiss) were successfully used to further reduce the image size. Diminished contrast makes these ob68
The Maskless Microprojection Photolithography System (MPLS) jectives less suitable for light directed microarray fabrication. However, patterning of photoresist - having lower requirements on contrast - should be simple with these higher magnification objectives. At a wavelength of 365 nm the diffraction limit of the 5×(0.25 NA) Fluar objective is R=λ/(2 ·NA) = 0.73 µm. However, a significantly larger distance between adjacent features is necessary to achieve a sufficient local contrast for the light-directed fabrication process. Reflective objectives have the advantage of a high UV transmission and are not subject to chromatic aberrations. We therefore tested image projection with a 15×(0.28 NA) Schwarzschild type reflective objective (Ealing). However, a satisfactory image contrast over the whole field couldn’t be achieved. Also, in the given optical system, owing to a narrow back aperture, the light throughput through the reflective objective is very limited. 3.2.4 Fabrication and Application of UV-Sensitive Photochromic Films For evaluation of the imaging quality a fast and simple method for generating patterns upon UV exposure is required. Photographic films and photoresist turned out to be not very useful due to difficult handling and processing efforts. Therefore we have developed a UV-sensitive film based on the photochromic dye spiropyran. Spiropyran undergoes a structural change when exposed to UV-light. This results in a strongly increased light absorption in the visible range. Preparation of photochromic films: We dissolved 10 mg of spiropyran dye (1’,3’-dihydro-1’,3’,3’-trimethyl-6-nitrospiro [2H-1-benzopyran-2,2’-(2H)-indole], Aldrich, Cat.: 27,361-9) in 1 ml of PMMA photoresist (E-beam resist PMMA 200 k; AR-P 641.04, Allresist GmbH, Strausberg, Germany) and spincoated a thin film (thickness about 1 µm) onto a microscope slide. Other resists - we also tried with MicroChem PMMA and MicroChem SU-8 50 - work equally well. The photoresist is used as a carrier material only. After spincoating, and brief heating on a hot plate (1 minute at 100◦C) the slides are ready for use. We found these photochromic films to be a well-suited imaging material. Unlike with photoresist or photographic material no developing or other processing is required. Under UV exposure the film changes from transparent to an almost opaque purple. With the intensities we usually apply (50-100 mW/cm2) this happens within seconds. The process can be reversed by heating or by illumination with bright light (at visible wavelengths). 69
The Microarray Synthesizer Unless the spiropyranhas been bleached with high irradiation doses, the films can be reused several times. For a small exposure dose the optical density increases almost linearly with the dose of UV light. For larger doses D the optical density OD approaches saturation. OD=ODsat(1−exp (−const.·D)) (3.1) Upon very high exposure, photodegradation of the photochromic dye (bleaching) results in reduced OD values. Since the linear exposure characteristics of the spiropyran dye are very similar to that of NPPOC phosphoramidite reagents, spiropyran films are a very useful tool for testing and evaluation of the UV optical system. 3.2.5 Chromatic Correction of the Projection Optical System Since the depth of focus DOF=λ/NA2is only about 6 µm for the 5×(0.25 NA) Fluar objective (at λ=365 nm), it is necessary to perform proper focusing each time a new patterning substrate is mounted on the sample holder. The focus range providing optimum contrast is even smaller than the depth of focus, thus perfect focusing of the pattern onto the surface is crucial. It can be achieved by observing the back reflection of the projected image (from the patterning surface) through the microscope eyepiece. This is easy to perform with visible light, but rather difficult with UV light. If the back-reflected image of the pattern is perfectly focussed in visible (green) light, this usually is not true for UV at the same time. This is owing to chromatic aberration. Longitudinal chromatic aberration causes an axial focus shift usually resulting in a completely blurred image in UV. In the following we describe a method for the correction of this longitudinal chromatic aberration, so that focusing of the near UV image can be performed by observation (through the eyepiece) and focus adjustment under green light illumination. Using photochromic films as a control for the quality of the projected UV pattern, we found that the chromatic aberrations can be compensated by fine-adjustment of the distance dbetween the DMD and the tube lens (see Fig. 3.2). The distance dis roughly the nominal focal length of the tube lens of 164.5 mm. After focusing with green light, the film is exposed with a control pattern in UV and subsequently inspected on a light microscope. The distance dnow can be adjusted iteratively until the patterns imaged on the spiropyran slide indicate perfect focusing. Just a small deviation of a few millimeters from the nominal focal length of the tube lens is necessary for chromatic correction. The tolerance of d, within which a good correction is achieved, is only a few tenths of a millimeter wide. Once the 70
The Maskless Microprojection Photolithography System (MPLS) chromatic correction procedure has been accomplished, focusing can always be performed under illumination with green light. Caution! The above optical adjustment depends on the eye focal length of the experimenter who performed the adjustment. In daily use of the microarray synthesizer, when focusing on the microarray substrate is performed, deviating eye focal lengths (near/far sightedness) of other personnel using the equipment do matter and need to be accounted for. 3.2.6 UV Light Intensity and Uniformity of Illumination For measuring the intensity at the image plane we used a laser power sensor (PS10Q, Coherent Inc.). The thermopile sensor was placed in the focal plane of the microscope objective. To measure the mean intensity, a completely white image was displayed on the DMD. With the measured total power of 7.8 mW we determined the intensity in the image plane as 87 mW/cm2. To study the uniformity of the illumination we projected the image onto a screen. The intensity was measured at different regions of the projected image. An asymmetric large scale deviation with a peak intensity of about 140% of the mean intensity is observed. This is due to the configuration of the illumination system: The UHP lamp’s arc gap is oriented parallel to the optical axis, providing a very inhomogeneous illumination profile. For this reason in a video projection system an integrator element, e.g. an integrator rod (which is a light guide with a rectangular cross section) or a fly-eye lens array is employed to generate a very uniform illumination. Using the integrator rod of the AstroBeam projector turned out to be not feasible as the glass rod absorbs most of the UV light. We decided to flatten the illumination profile by using only a small homogeneous section of the light cone for illuminating the DMD. This way we sacrifice about 80% of the light. Nevertheless, the remaining 20% of light allow photo-deprotection to be performed in a reasonable time. Alternatively, if such parts were available, a quartz integrator rod or an integrator plate (fly-eye lens array [Sun05]) could be used to achieve significantly higher light intensities. To attain a more uniform illumination we employ the DMD for intensity leveling, similar as described by Huebschman et al. [Hue04]. For thispurpose we have created an ”intensity leveling mask”. The black and white images (to be used as a photomasks) can easily be leveled to reduce intensity variations to about ±10% by pixelwise multiplication with this mask. To generate the intensity leveling mask, a fully illuminated image (all mirrors in the on-state) is projected on the screen (as described above, still without using the microscope objective) and photographed with a Nikon Coolpix 4500 digital camera. Deskewing the 71
The Microarray Synthesizer raw image using standard image processing software results in a 1024×768 pixel image, which finally has to be inverted and adjusted in brightness and contrast. The leveling mask is then projected onto the screen and a photometer is used to measure uniformity of illumination. In an iterative way image brightness and contrast are adjusted to achieve a uniform intensity within most of the image area. Contour plots of the light intensity before and after intensity leveling are shown in Fig. 3.6. Only in the outermost corners of the image (comprising about 10% of the total image area) the intensity is reduced to about 50% of the mean intensity. This is due to vignetting: Light reflected from the corners of the DMD, which are located close to the edge of the entrance pupil, is partially blocked by the apertures of the tube lens respectively the microscope objective. Applying intensity leveling we achieved a mean light intensity of 76 mW/cm2. 125 120.6429 116.2857 111.9286 107.5714 103.2143 98.8571 90.1429 81.4286 68.3571 20 40 60 80 100 120 140 160 180 200 20 40 60 80 100 120 140 104 101.1429 98.2857 95.4286 92.5714 89.7143 86.8571 81.1429 66.8571 66.8571 20 40 60 80 100 120 140 160 180 200 20 40 60 80 100 120 140 a) b) Figure 3.6: Uniformity of illumination. (a) Intensity contour map before intensity leveling. (b) After intensity leveling. Using the tube lens, the image of the DMD was projected onto a screen, without the microscope objective in place, and photographed with a digital camera. Vignetting from the microscope objective is neglected here but this effect is small compared to vignetting of the tube lens. The intensity values mentioned above were achieved using an interference filter with a FWHM of 33 nm and a maximum transmission of 60% at a center wave length of 370 nm. Using a narrow i-line filter (FWHM 12 nm at a center wavelength of 365 nm; 35% maximum transmission) provided significantly lower intensities (about one ninth of the intensity achieved with the broad filter). The demand for a wide filter can be explained by the strong line broadening due to the high operation pressure of the UHP mercury arc lamp. 72
The Maskless Microprojection Photolithography System (MPLS) 3.2.7 Optical System Performance Testing with UV-Sensitive Photochromic Films Light-directed synthesis of DNA microarrays requires that the image is projected onto a substrate inside an inert reaction chamber, so that reactions can take place under a moisture free argon atmosphere. The synthesis substrate, a 0.17 mm thickness microscope cover glass, is forming the window of the reaction cell. Hence the image has to be projected onto the inner face of the window. For image focusing (see Sec. 3.2.5) we use the small fraction of green light which is reflected back from the imaging surface into the microscope. Applying a similar approach for contrast measurement is not practicable because the outer face of the cover glass contributes to back-reflection as well. Multiple reflections in the microscope system (e.g. from a beamsplitter) may degrade the image contrast further. Contrast ratios of 1:3000 (as can be found in product specifications of video projection systems) usually refer to the full-on/full-off contrast obtained by comparing the intensities of completely black respectively white images. On our setup (placing a photometer into the focal plane of the microscope objective) we measured a full-on/full-off contrast of about 3400:1. This means that the DMD chip with the mirrors in the off-position reflects only about 0.03% of the exposure intensity onto the imaging substrate. This means that the amount of light scattered by the DMD housing and by the mirrors in the off-position is negligible. Much more relevant for DNA microarray synthesis is the local contrast [Kim04] between neighboring features. The local contrast is diminished by light-scattering and diffraction from mirrors in the on-state, but also by optical aberrations, which cause distortions to the point spread function. It also depends on the feature geometry (i.e. feature size and feature spacing). Reflections within the imaging optics cause flare. This could possibly be improved by using UV anti-reflection coated optical surfaces (DMD window, tube lens). The patterns used for microarray synthesis typically have an array structure with a pitch of 17 µm or less. To obtain an estimate of the stray light induced error rate we have measured the image contrast at high spatial frequencies. We found that the UV-sensitive films we already used for adjustment of the UV optics (see section 3.2.5) are very well suited for testing the performance of the photolithography system. For visual inspection of the patterns we used an optical microscope (Olympus IX81) equipped with an automated X-Y translational stage and with a high resolution CCD camera (C9100 EM-CCD, Hamamatsu Photonics). Patterns of regularly spaced line pairs (a pair comprises a black and a white bar of equal width), were imaged onto photochromic film (Fig. 3.7). The spatial frequency of the pattern was varied between 14 and 70 line pairs per millimeter (lp/mm). Using an exposure 73
The Microarray Synthesizer 3.3 The Fluidics System The modified valve block of a commercial DNA synthesizer (Applied Biosystems ABI 381A)constitutesthemaincomponentofthefluidicssystem(Fig. 3.12). Amicrocontrolleroperated solenoid valve driver (technical details described in Appendix B.8) enables control of the fluidics system via the RS232-interface of the control PC. To ensure waterand oxygen-free conditions the microarray synthesis is performed in an air tight flow cell (Fig. 3.13) under an inert argon atmosphere. Argon gas pressure is employed to drive the reagent transport. A detailed schematic of the fluidics system is shown in Fig. 3.12. 3.3.1 The Synthesis Cell Technical requirements •resistance to the aggressive solvents MeCN, THF and pyridine •use of chemically inert materials (not affecting DNA synthesis) •tight sealing (no seeping of reagents below the gasket) is necessary to enable complete exchange of reagents (e.g. to avoid contamination with water left over from the previous reaction step) •negligible dead volume (required for fast and complete exchange of reagents) •the 0.17 mm microarray substrate must constitute a window of the cell •prevention of gas bubble sticking at the edges of the cell volume •light reflection and scattering must be avoided Implementation The cell volume is formed by a streamlined cutout (shown for example in Fig. 3.15) in an approx. 1 mm thick sheet of polydimethylsiloxane(PDMS) silicone rubber. PDMS is used for its chemical inertness and sealing capability.7The fragile microarray substrate (diam. 22 mm round cover glass) can be reliably sealed with little force, thus without the risk of fracture. The DNA chip substrate is employed as an optical window (Fig. 3.14). Photomasks are projected with a microscope objective onto the inner surface of the substrate, where the DNA probes are synthesized. 7Even though PDMS appears to be chemically inert, we observed (reversible) solvent swelling of the PDMS gasket upon exposure to tetrahydrofurane and pyridine. To prevent excessive deformation of the synthesis volume, oxidation and capping steps should not be longer than necessary. An alternative THF- (and water-) free oxidizer solution (enabling phosphoramidite synthesis on PDMS surfaces) has been described in [Moo05]. 80
The Fluidics System Argon Argon Argon Activator Cap A Cap B Dep. Soln. Argon Oxidizer Argon MeCN 24 23 20 19 18 11 109 8 76 AX T C G Argon Vent 21 22 154 3 2 Argon MeCN Oxidizer Depr. Soln. Cap A Cap B Activator A T Waste X G C 8 7 6 5 1 4 3 2 14 13 12 11 10 9 Waste Argon 17 15 Synthesis Cell Waste 16 T-Piece A B Figure 3.12: Schematic of the fluidics system. The valve block has been adopted from a commercial oligonucleotide synthesizer. The synthesis cell has replaced the synthesis column employed in standard oligonucleotide synthesis. Valve numbers correspond to those used in the synthesis control software. (A) Valve block. (B) Reagent storage bottles. Argon gas pressure (via valves 1, 15, 18-21, 23-24) is employed to drive the reagent transport. 81
The Microarray Synthesizer Figure 3.13: Schematic of the synthesis cell. Syringe needles form the inand outlet of the flow cell. A PDMS gasket with a streamlined cutout forms the synthesis cell volume which is sandwiched between the chip substrate and a glass plate. The assembly is placed on an inverted microscope. UV light from the objective is entering the cell through the chip substrate. Mask patterns are projected onto the inner face of the substrate, where the in situ synthesis takes place. Figure 3.14: The synthesis cell on the microscope. The assembly is mounted on a precision-adjustable aluminium support. The microarray substrate (round cover glass) is located above the microscope objective. Use of transparent materials (polycarbonate, PDMS and glass) simplifies handling and enables visual control. 82
The Fluidics System To prevent the attachment of gas bubbles (argon gas employed to drive the fluidics system tends to form bubbles upon pressure relief) at the edges of the PDMS cell, we put effort in making a cell with very smooth edges. This is achieved using a sharp-edged punching tool (see appendix B.4) fabricated (electrical discharge machining) by the mechanics workshop of the university. Smooth surfaces also improve the reagent exchange between consecutive synthesis steps. For the purpose of chemical inertness the upper side of the synthesis cell consists of a glass microscopy slide which is glued onto a 10 mm thick block of UV absorbing Makrolon R plastics. Back-reflection (and back-scattering) of UV light from the interfaces is reduced (by index matching) with a thin layer of PDMS employed as glue. Inand outlet are formed by syringe needles which are connected to the valveblock via PTFE tubing (Fig. 3.14). More detailed information on the construction of the synthesis cell is provided in appendix B.4. The design of the cell is optimized for DNA in situ synthesis with light-directed photodeprotection. It enables a very small reagent consumption of ca. 40 mg of each NPPOCphosphoramidite for a 25mer synthesis [Nai06b]. 3.3.2 Argon Bubble Trapping The occasional formation of argon bubbles, owing to pressure relief during the reagent transport towards the synthesis cell (the solvent MeCN is saturated with argon) represented a serious problem for the microarray synthesis. Bubbles which have become trapped in the synthesis volume (Fig. 3.15A) do locally increase the stray light intensity during the UV exposure or affect synthesis reactions (coupling etc.) since the substrate surface beneath the bubble is not covered by the reagents. The ”argon bubble problem” has been resolved with a cleverly devised technique: •Large bubbles are captured by a T-piece bubble trap (Fig. 3.15C) which is integrated in the inlet line. •Small bubbles (<2 mm diam.), owing to the increased channel width in the synthesis area, havethe tendencyto get stuck in the synthesis volume (Fig. 3.15A). By employing a short suction pulse the small bubbles are pushed into the inlet region of the synthesis cell (Fig. 3.15B), where they get reliably trapped. This method of bubble catching is highly reliable. In the critical steps of the synthesis process the occurrence of bubbles within the synthesis area is prevented almost completely. 83
The Microarray Synthesizer A C B Figure 3.15: Bubble trapping techniques. Top views of the synthesis cell volume (A) and (B) – inand outlet holes are shown at the right and the left end of the streamlined cell volume. During reagent supply small argon bubbles can get into the synthesis cell (A). They can be removed from the synthesis area (dashed box) by a applying a short suction pulse. Bubbles are moved into the narrow inlet region of the chamber (B), where they are trapped due to a more favorable surface energy. (C) Larger bubbles (too large to get trapped in the inlet region) are captured in a ”T-piece bubble trap” before they reach the synthesis cell. Venting of the accumulated gas is achieved by occasionally opening the valve to the venting line. 3.4 Automated Microarray Synthesis - ControllerHardand Software The original ABI 381A DNA synthesizer control hardware has been substituted by a personal computer based controller. Fully automated light-directed in situ synthesis is performed with the synthesizer control software DNASyn, which is described in detail in appendix B.7. DNASyn integrates fluidics control (via an external microcontroller-based solenoid valve driver - technical details are provided in appendix B.8) with the ”virtual photolithography mask” projection. 3.5 Performance of the Microarray Synthesizer An affordable microarray synthesizer system for lab-scale fabrication of DNA microarrays has been developed from the following widely available components: •Oligonucleotide synthesizer (second-hand) 84
Performance of the Microarray Synthesizer •DLP video projector (second-hand) •Inverted research microscope (second-hand) •Personal computer •Microcontroller-based solenoid valve driver (home-built) •Optical components: Optical table, filters, lenses, mirrors The highly flexible microarray synthesis system enables massively parallel in situ synthesis of almost arbitrary probe sequences. New microarray designs can be developed within hours and automatically synthesized overnight. The synthesis of a 25mer microarray requires about 6.5 hours (plus 1.5 hours for the final deprotection step outside the synthesis apparatus). With our microscope-projection-lithography setup the size of the microarrays has been reduced to a total area of <10 mm2. Owing to the miniaturization, the costs for synthesis reagents (NPPOC-phosphoramidites, MeCN, activator, oxidizer, ethanol) are about 50 Euros per microarray synthesis. Moreover, the small area of the microarray enables hybridization with a very small amount of target solution (in principle less than 10 µl are required). The high stability of our microscope-projection-lithography setup (with respect to image drifting originating from thermal expansion etc.) is beneficial for the quality of the synthesized DNA probes. In principle each micromirror-pixel (in total 1024×768) could be used to synthesize a microarray feature. However, the need for a high local contrast and expected difficulties with the image analysis of the small densely-packed features (image distortions etc.) require the use of composite features consisting of 5×5 DMD pixels (4×4 pixel feature area plus 1 pixel separation gap). In the corners of the synthesis area (i.e. the imaging field defined by the DMD chip) DNA probe quality is suffering from vignetting (→reduced exposure intensity) and uncorrected curvature of field (→reduced local contrast). For quantitative investigations of probe-target binding affinities, a maximum number of about 25000 microarray features is currently achievable. 85
The Microarray Synthesizer 86
Chapter 4 Light-directed in situ Synthesis of DNA Microarrays 4.1 Light-Directed in situ Synthesis of DNA Microarrays Reagents RayDiteTM 3’-phosphoramidites NPPOC-dA(tac), NPPOC-dC(ib), NPPOC-dG(ipac) and NPPOC-dT (see Fig. 4.1) carrying photolabile 5’-nitrophenylpropyloxycarbonyl protective groups were purchased from Sigma-Proligo (Hamburg, Germany). Acetonitrile (ROTISOLV R for DNA synthesis, water<10 ppm, Carl Roth GmbH, Germany); Activator42TM, 0.25 M (Proligo R ); iodine based oxidizer (part no. 401732, Applied Biosystems); Trap-PakTM molecular sieve bags (Applied Biosystems); water-free argon (≤0.5 ppm H2O) Photo-deprotection is carried out in a mildly basic (deprotection) solution of 25 mM piperidine (99%, Aldrich) in waterfree acetonitrile. Alternatively, the use of dimethylsulfoxide (DMSO) has been reported [Woe06]. Final base deprotection is performed (at room temperature for about 90 minutes) in a 1:1 mixture of etylenediamine (analytical grade, Fluka) and ethanol (analytical grade, VWR, Germany). UV glue (Norland optical adhesive 60, Edmund optics) is used to fix the chip onto a stainless steel support. 87
Microarray Synthesis NO2 O CH3 O O NCCH3 CH3 CH3CH3 N P O B O O N N N NH(t ac) N NH O O CH3 N NH N N NH(ipac) O N N NH(ib) O B= Adenine(tac) Thymine Guanine(ipac) Cytosine(ib) ib=isobutyryl tac=tert-butylphenoxyacetyl ipac=isopropylphenoxyacetyl Figure 4.1: 5’-[2-(2-Nitrophenyl)-propyloxycarbonyl]-2’-deoxynucleoside phosphoramidites. Similar as nucleosides, nucleoside phosphoramidites comprise nucleobases and deoxyribose sugar. Additionally, phosphoramidites contain a phosphorus group, which, when chemically activated, can react with the hydroxy group of a growing (deprotected) oligonucleotide strand, This coupling reaction creates the phosphate group in the sugar-phosphate backbone. Various protection groups enable a controlled synthesis of oligonucleotide chains without the risk of unwanted side reactions. The photolabile NPPOC group (blue) substitutes the 5’-hydroxyl of the pentose ring. Its removal (deprotection) enables coupling of another building block. The phosphorus group is protected by a diisopropylamino group (red) (→phosphoramidite) and a 2-cyanoethyl protection group (green). Further protection groups (ib,tac, ipac) are necessary to prevent side reactions of the exocyclic amine groups of the nucleobases during the in situ synthesis process. All protection groups are removed at the end of the synthesis. Preparation of the Microarray Synthesis Light-directed in situ synthesis was performed with NPPOC-phosphoramidites [Has97; Bei99; Nuw02; Nai06b] which differ from the commonly used acid-labile DMT-protected phosphoramiditesby the photo-cleavable5’-nitrophenylpropyloxycarbonylprotectiongroup (NPPOC). Phosphoramidite reagents are highly sensitive to water. To minimize contamination with water, NPPOC-phosphoramidite solutions - 40 mM in water-free MeCN - are prepared only immediately before the start of the synthesis. Deprotection solution, oxidizer1and activator are more stable and can remain on the synthesizer for prolonged times. Contamination with water is particularly critical for the phosphoramidite/MeCN solution contained in the storage bottles. Once degradation due to a small amount of water has started, the phosphoramidites undergo autocatalytic degradation [Kro04]. To minimize water contamination in critical reaction steps, molecular sieve bags (Trap-PakTM) are added to the MeCN storage bottle and to the activator storage bottle. Further hints on phosphoramidite han1The oxidizer solution itself contains a considerable amount of water (several percent) 88
Light-Directed in situ Synthesis of DNA Microarrays dling procedures are provided in appendix B.5.1. The preparation of the automated synthesis should be performed with the synthesis script PrepSyn.prg, which is executed by the controller software DNASyn. The PrepSyn-script includes the preparation of the phosphoramidite solutions, priming of the reagent supply lines and checklist functionality (installation of the synthesis cell, optics, argon pressure, valve function, reagent availability). The automated synthesis cycle An initial photoreactive monolayer is created by coupling of NPPOC-dT-phosphoramidite to thehydroxyl-groupsof the dendrimer functionalizedsubstrate. The synthesiscycle, to be repeated 4×25=100 times for the synthesis of a microarray with 25mer probes2, comprises phosphoramiditecoupling,phosphitetriesterbond oxidation, and photo-deprotection. •Phosphoramidite coupling is carried out for one minute with a 1:1 mixture of 40 mM NPPOC-amidite solution in water-free MeCN and activator solution (Activator42TM, 0.25M) •A iodine based oxidizer solution (ABI) is employed for about 40 s (after every fifth coupling step) to oxidize unstable phosphite triester bonds, thus to form stable phosphotriester linkages •The photo-deprotection step (exposure dose 7 J/cm2at λ=370 nm) is performed in a 25 mM solution of piperidine (Sigma-Aldrich) in MeCN. Piperidine [Bei99] provides the mildly basic conditions necessary for photocleavage of the NPPOC protection group [Wol04]. Between the individual reaction steps extensive washing of the valve block and of the synthesis cell/supply line is performed. It is, for example, absolutely necessary to remove trace amount of water (from previous oxidation steps) from the fluidics system prior to the next coupling reaction. Alternating rinsing with pure MeCN and flushing with argon gas is very efficient to remove remaining reagents from the previous reaction step. However, it is important that solid residues are not allowed to dry on the substrate surface. The final coupling step is followed by complete photo-deprotection of the whole microarray, to remove all remaining NPPOC protection groups, and by a final oxidation step. Capping of unreacted binding sites by acetylation is commonly employed in oligonucleotide synthesis to prevent the synthesis of strands containing point defects. Because of the rather limited benefits of a capping in light-directed microarray fabrication (see section 2.7.1) we do not apply capping in our DNA Chip synthesis scheme. 2In practice, owing to mask optimization, only about 80 cycles are required. 89
Microarray Synthesis Figure 4.6: Wetting characteristics of the microarray surface (microscope image). Hydrophilic features (size about 20 µm) are covered by a closed thin film of water. Regions between feature blocks are covered with tiny droplets. 4.3.3 Hybridization without Detergent - Unspecific Adsorption Omittance of the surfactant (Tween-20TM or SDS) resulted in very strong surface absorption on the entire microarray surface - also in the regions where no probes have been synthesized. Subsequent addition of 0.01% Tween-20 on the same microarray resulted in probe-specific hybridization. Hybridization in this particular experiment was performed with MES hybridization buffer at room temperature. 4.3.4 Irreversible Target Adsorption In MES hybridization buffer at temperatures >55◦C targets tend to bind irreversibly to the substrate surface, making a reuse of the microarrays impossible. The fluorescence intensity is particulary high between the features (see Fig. 4.7). This suggests that targets which have dissociated from the probes are captured by reactivegroups at the substratesurface adjacent to the features. The problem seems to be related to the use of the MES hybridization buffer at high temperatures (>55◦C) . Using 5×SSPE buffer instead, we do not observe this characteristics. However, we found that often (even at high temperatures of 70◦C) the hybridization signals can not be completely removed. This problem, which has also been reported by Hu et al. [Hu05], could be owing to stable duplexes which do not completely dissociate at the temperatures applied. It is also possible that hybridized targets have an increased probability for bonding to unblocked reactive sites at the microarray surface. In 96
Noteworthy Characteristics of the Microarrays Figure 4.7: Fluorescence micrograph of irreversible adsorption. The feature blocks in the center of the image have undergone dissociation in pure MES hybridization buffer. At temperatures of about 60◦C rather than to detach from the surface the fluorescently labeled targets have irreversibly bound to the microarray surface. The brightest signal is visible in the gaps between the features. At the left edge of the image another feature block (with another sequence motif) is shown, which has been hybridized after dissociation conditions have been applied (thus demonstrating that the other probes on the microarray maintained their hybridization capability). either case the targets can be removed completely if RNA targets are used rather than DNA targets. An alkaline stripping procedure (sodium hydroxide) will selectively degrade RNA targets (into nucleotides), whereas DNA probes remain unaffected [Hu05]. 4.3.5 Robustness of the Phosphorus Dendrimer Surface Coating Fig. 4.8 demonstrates that the phosphorus dendrimer functionalization (section 4.2) forms a stable network on the glass surface. Parts of the dendrimer coating (autofluorescence under blue excitation) have come off the surface after harsh treatment with an unsuitable stripping buffer. The robust closed-film structure shown in Fig. 4.8 is rather unexpected since the chemistry of the surface-functionalization would rather suggest a monomolecular layer of unconnected dendrimer molecules. However, dendrimers bound to the aminosilane layer possibly form a densely interwoven network. It is further possible, that the functionalization with the aminosilane APTES results in the formation of a stable multi-layer film. 97
Microarray Synthesis Figure 4.8: Fluorescence micrograph of the phosphorus dendrimer substrate. Use of an unsuitable stripping buffer (10 minutes in boiling 0.1 M Na2CO3solution) revealed the stable network structure of the surface coating. It appears that the phosphorus dendrimer network remained intact, even though the coating is completely detached from the glass surface. 98
Chapter 5 DNA Microarray Analysis 5.1 Hybridization Signal Acquisition - Experimental Setup Microarray hybridization assays were performed in a temperature-controlled hybridization chamber. The design of the flow-through type chamber is similar to that of the synthesis cell (see section 3.3.1). Installation on an epifluorescence microscope setup enables real time monitoring of the hybridization signal. A sensitive electron multiplying CCD-camera (EMCCD) is used for image acquisition. B A Figure 5.1: Microarray analysis setup. (A) Motorized fluorescence microscope with EMCCD-camera (bottom left). (B) Hybridization chamber on the XY-stage of the microscope. 99
Microarray Analysis 5.1.1 The Hybridization Chamber Design considerations: •realtime monitoring (requires a window into the sealed chamber and low background fluorescence) •reagent exchange (e.g. to replace the hybridization buffer by a washing solution) •high mechanical stability to minimize defocusing and xy-drifting of the image upon thermal expansion •temperature control AB Figure 5.2: Hybridization chamber assembly (A). Inlet/outlet tubes enter from the top. The microarray is located at the bottom. Part (B) shows the microarray on its stainless steel support (lying in the front) and the stainless steel top plate (leaning against the brace) with the PDMS gasket. The microarray slide is pressed against the aluminium brace with two fastening screws. The hybridization chamber is made from a 1.5 mm thick PDMS gasket. The chamber volume (about 120 µl) is formed by a 10 mm diam. hole (cut from a sheet of PDMS with a punching tool). Circular cut-offs at the inlet and outlet openings (see Fig. 5.2B) prevent sticking of air bubbles inside the chamber volume. The microarray with its stainless steel support constitutes the bottom side of the hybridization volume. This configuration, using the chip substrate as window, enables observation of the hybridization signal with an inverted microscope. A stainless steel plate forms the upper side of the hybridization volume. Stainless steel is used because it is resistant to the hybridization buffer (no salt corrosion). Also important, since the steel plate is in the background of the microscope field of view: the steel plate isn’t fluorescent and doesn’t adsorb nucleic acid targets. A flexible ThermofoilTM heater (Minco) (with a 15 mm diam. opening in the center - for 100
Hybridization Signal Acquisition - Experimental Setup inlet/outlet tubes) is glued onto the upper side of the steel plate. Temperature is measured with a platinum resistor (Pt-100) which is fixed with thermal adhesive at the edge of the steel plate. Inand outlet tubes (located at opposing ends of the chamber volume) penetrate the steel plate from the top side (see Fig. 5.2A). To avoid dead volume and corrosion and to achieve reliable sealing, the PFA tubes are directly connected to the plate by press-fitting.1 Temperature Control PT100 PC ProfiLabExpert 3.0 PID-Temperature-Controller Meilhaus RedLab USB Measurement Module Toellner TOE 8951 Power Supply USB 2.0 A/D in D/A out Minco Thermofoil heater Pt100 Temperature Sensor Resistance ->Voltage Converter analog Remote Control Hybridization Chamber Figure 5.3: Control of the hybridization temperature. Heating of the hybridization chamber is performed with a Minco ThermofoilTM heater which is in thermal contact with the hybridization solution via a corrosion resistant stainless steel plate. The temperature is measured with a Pt-100 sensor (in thermal contact with the stainless steel plate). The resistance is converted into a voltage signal that is proportional to the temperature. The voltage is read by an A/D input channel of the RedLab USB measurement module. A software based PID-temperature controller (run as a PC application generated with ProfiLabExpert 3.0 - see Fig. B.16) by comparing the actual temperature and the set temperate, determines the control voltage (output via the RedLab D/A output) that is used to operate the remote controlled heater power supply. Temperature ismeasuredwithaPt-100 resistorand convertedintoatemperature-proportional voltage signal. A USB measurement module (ME-Redlab, Meilhaus) is employed for signal acquisition with a personal computer. A software-based PID-controller (see appendix Fig. B.16) designed with ProfiLab-Expert 3.0 (ABACOM Electronics-Software) enables user-defined temperature profiles and temperature-recording. The heating power for the foil heater is provided by a remote-controlled power supply (TOE 8951, Toellner Electronic Instrumente GmbH) which is controlled via the D/A-output of the USB-module. A test with a calibrated Pt-100 resistor - brought in thermal contact with the outside of the microarray substrate - showed that the temperature at the microarray surface is controlled 1The tubes - outer diam. 1.2 mm int. 0.8 mm were drawn through the 1 mm diam. inlet/outlet mounting holes in the steel plate (→stable press-fit-connection). 101
Microarray Analysis with an accuracy of approximately ±1◦C. The temperature can be held constant within a variation of <0.2◦C. 5.1.2 Epifluorescence Microscope The microarray hybridization signal is acquired by epifluorescence microscopy (principle shown in Fig. 5.4). Realtime monitoring of the hybridization signal is performed with an Figure 5.4: Epifluorescence microscopy. The light of a bright mercury arc lamp (A) is filtered by the excitation filter (E) and reflected by the dichroic mirror (D) through the microscope objective (M) onto the fluorescently labeled sample (F). Fluorescent dye molecules absorb the excitation light, and enter an excited electronic state. Due to the Stokes shift the emitted light has a longer wavelength than the excitation light. The Cyanine 3 (Cy3) dye used throughout this study has a peak absorption at 550 nm (green) and shows yellow to orange fluorescence emission (with a peak at 570 nm). A fraction of the fluorescence signal (emitted in all directions) is collected by the microscope objective M and transmitted through the dichroic mirror (D). The barrier filter (B) passes only the fluorescence light to the camera (C). Olympus IX81 inverted research microscope (Fig.5.1). The fluorescence of Cy3 labeled targets is imaged using an UPlanApo 10×0.40 NA microscope objective (Olympus) and the U-MWG 2 filter set (Olympus). 102
Quantitative Analysis of Microarray Hybridization Signals 5.1.3 Image Acquisition with an EM-CCD Camera High resolution image acquisition was performed with a sensitive Hamamatsu EM-CCD C9100-02 electron multiplying camera. Camera specifications: •Peltier cooling: -50◦C •Gain factor: 800 •Read-out noise: <1 electron r.m.s. at high gain mode •Dynamic range: 14 bit •Full resolution: 1000×1000 pixels Camera and microscope (shutter, filter, exposure, focus, XY-stage etc. ) were controlled by the SimplePCI (Compix Inc.) image acquisition software. Shading correction Uneven fluorescence excitation and fluorescence collection, owing to vignetting (larger blockage of off-axis light rays) yield fluorescence micrographs that are brighter at the center and darker at the edges. Intensity gradients due to shading can be a significant source of error for quantitative analysis of hybridization signals. Shading correction (using the SimplePCI setting Ratio shade correction) is therefore performed by dividing the specimen image (microarray) through a fluorescence reference image, which is acquired by imaging a uniformly fluorescent surface. As described by Model et al. [Mod01] spatially uniform fluorescence is obtained from a thin layer of fluorescent dye (e.g. 20 µl hybridization solution with 100 nM of Cy3 labeled targets) sandwiched between a microscopy slide and a cover glass. 5.2 Quantitative Analysis of Microarray Hybridization Signals Fluorescence micrographs of the hybridization signal are saved as 16-bit grayscale TIFF images. Shading correction is performed during image acquisition. Quantization of feature intensities is carried out with the Java program ScanRA (technical details in appendix B.10). The software (which was developed as part of this thesis) enables automatic analysis of microarray feature intensities. To define feature positions a 103
Microarray Analysis Figure 5.5: Raw hybridization signals as imaged with the Hamamatsu EM-CCD camera (original resolution of the image 1000×1000 pixel - size reduced to 500×500). For the image acquisition the 16mer-microarray remained in the hybridization solution (1 nM Cy3-end-labeled RNA oligonucleotide target). The hybridization temperature was 30◦C. readout grid (Fig. 5.6C) is placed on the microarray image. Then the program integrates pixel intensity values over the integration boxes located at the grid points in center of the features. The size of the integration boxes should be chosen to prevent integration over feature boundaries (Fig. 5.6D). The exact placement of the readout grid requires rotation of the image, so that the microarray grid is approx. aligned with the screen axis. Considering small image distortions an orthogonal grid of evenly spaced points is not suitable to determine features positions (Fig. 5.6A). Rather, a quadrilateral grid (defined by the four corner points) is suitable to account for first order distortions of the microarray image. Microarray hybridization signals (16-bit intensity values) are averaged over the integration boxes to provide a 16-bit mean intensity value. The standard deviation of the pixel intensity values provides information about the homogeneity of the individual microarray feature intensities. Large standard deviations can indicate defects (e.g. fluorescent particles or scratches on the microarray surface) or bad alignment of the readout grid. Average brightness, standard deviation of the feature brightness and the position of the individual features are saved in comma-separated value (CSV) format. 104
Real-time Monitoring of Microarray Hybridization A B CD Figure 5.6: Microarray analysis with ScanRA. Readout of the hybridization signal intensities of a feature block. (A) Rotation of the image. (B) An orthogonal readout grid doesn’t match all feature positions exactly if the array is slightly distorted. (C) A quadrilateral readout grid defined by the four corner points is a good first order approximation for small image distortions. (D) Integration boxes (blue) are located at the grid points in the center of the features. Time series of fluorescence micrographs can be analyzed in batch mode – the readout grid needs to be defined only once. To account for drifting of the image owing to thermal expansion of the hybridization chamber (if the temperature has been varied significantly), the position of the first (upper left) corner point has to be provided manually for about five images. The drift offsets of the other images are determined by linear interpolation. 5.3 Real-time Monitoring of Microarray Hybridization Microarray hybridization is usually followed by one or several washing steps to remove unhybridized targets (see Fig. 2.22). Washing is necessary for the detection of small hybridization signals since these are otherwise not visible within the fluorescent background of the hybridization solution. This is typically the case for expression profiling experiments where thousands of different nucleic acid targets comprise the hybridization solution. However, in most of the experiments performed this study only a single target species is contained in the hybridization solution. At a target concentration of 1 nM the concentrated fluorescence of the hybridized targets (surface-bound in the microscope focal plane) can be well-distinguished from the background fluorescence of the hybridization solution. 105
Single Base Defects - Microarray Experiments 0 2 4 6 8 10 12 14 16 0 0.05 0.1 0.15 0.2 T T G A C T T T C G T T T C T G Defect position Hybridization signal (a.u.) 5’-AACTCGCTATAATGACCTGGACTG-Cy3-3’ 3'-TATTACTGGACCTGAC-5’ Target oligonucleotide Probe sequence motif (complementary to a section of the target) Set of point-mutated probe sequences, derived from common probe sequence motif 3'-T TTACTGGACCTGAC-5’ 3'-T TTACTGGACCTGAC-5’ 3'-T TTACTGGACCTGAC-5’ 3'-T TTACTGGACCTGAC-5’ 3'-TA TTACTGGACCTGAC-5’ 3'-TA TTACTGGACCTGAC-5’ 3'-TA TTACTGGACCTGAC-5’ 3'-TA TTACTGGACCTGAC-5’ 3'-T TTACTGGACCTGAC-5’ A C G T - A C G T Defect position 2 3'- ATTACTGGACCTGAC-5’ 3'- ATTACTGGACCTGAC-5’ 3'- ATTACTGGACCTGAC-5’ 3'- ATTACTGGACCTGAC-5’ 3'-T ATTACTGGACCTGAC-5’ 3'-T ATTACTGGACCTGAC-5’ 3'-T ATTACTGGACCTGAC-5’ 3'-T ATTACTGGACCTGAC-5’ 3'- ATTACTGGACCTGAC-5’ A C G - TA C G T Perfect match (PM) probe Single base mismatch (MM) probes Single base deletion probe Single base insertion probes Defect position 1 Defect position 1 Defect position 2 Hybridization signals Defect profile Hybridization with the Target sequence Feature arrangement on the microarray } } Figure 6.1: Design of the experiment: a comprehensive set of point-mutated probes is derived from a common probe sequence motif which is complementary to the target sequence (probe sequences are shown for the first two defect positions only). For each defect position these include 3 single base mismatches (MMs - shown in red), 4 single base insertions (green), one single base deletion (red) and one perfectly matching (PM) control probe (blue). To enhance quantitative analysis, probe sets are arranged on the microarray as a compact feature block. Hybridization signal intensities from hybridization with the target sequence are plotted versus defect position. The defect profile shows relative binding affinities (i.e. the discrimination between the defect hybridization signal and the corresponding PM hybridization signal) as a function of defect type and defect position. 112
DNA Microarray Design •extraction of the defect positional dependence, •comparison of the binding affinities of different defect types, and on the •identification of further influential parameters. The individual chip designs employed differ in selection and spatial arrangement of the probe sequences. 6.3.1 Microarray Design Considerations for Quantitative Analysis of Hybridization Affinities Several factors affect quantitative analysis of microarray hybridization signals: Spatial variations of the photo-deprotection intensity and optical aberrations affecting the imaging contrast can result in gradients of the probe DNA quality (as indicated in Fig. 6.2B). Depending on their position on the microarray, probes contain a varying degree of random synthesis errors. The corners of the rectangular synthesis area are most affected, since here the UV exposure dose, due to vignetting, is significantly smaller than in the center of the synthesis area. Gradients on the fluorescence intensity also arise from optical vignetting in the fluorescence microscope. This is largely compensated by shading correction (see section 5.1.3). To minimize impairments by gradients, probes for which hybridization signals are to be compared directly were arranged in closely spaced feature blocks (as shown in Figs. 6.1 and 6.2). Local target depletion during hybridization (see section 8.5) can likewise result in positiondependent gradients of the hybridization signal intensity. In feature blocks with identical (or very similar) probe sequences, owing to the competition of the probes for the same pool of targets, features in the center of the block (surrounded by 8 competing features) - under unfavorable hybridization conditions [Pap06] - can have smaller hybridization signals than equivalent features at the edges of the feature block. Control features (comprising perfect matching probes) which are evenly distributed over the feature block, are employed to indicate hybridization signal gradients: the variation of the PM signals (e.g. in Fig. 6.4A) shows the magnitude of feature-position dependent bias. Usually the impairment of the hybridization signal by such gradients is relatively small, resulting in variation of the control-probe intensities which is typically smaller than 5-10% of the PM hybridization signal intensity. However, if the hybridization kinetics is very fast - thus incoming targets are preferentially captured by the probes at the edge of the feature block - spatial variations of the hybridization signal of up to 50% of the PM intensity can occur [Pap06]. Unfavorable conditions affecting quantitative measurement are avoided by 113
Single Base Defects - Microarray Experiments using relatively short probes (rather 16mers than 25mers), using sequences with moderate binding affinities, and by application of sufficiently stringent hybridization conditions. 6.3.2 Single Base Defect Experiments A C G T A C G T A C G T A C G T A C G T A C G T A C G T A C G T A C G T A C G T A C G T A C G T A C G T A C G T A C G T A C G T 246 8 10 12 14 16 911 13 15 135 7 AB1 16 Figure 6.2: Microarray feature arrangement (A) for the single base mismatch experiment (compare with Fig. 6.3) and (B) for the direct comparison of various defect types. In Athe feature block comprises 16 MM positions. The substitution base is either A, C, G or T. Depending on the probe sequence motif the substitutions will result in one PM and three MM probes. The design in Bincludes one PM, three MM, four single base insertion and one single base deletion probe four each of the 16 defect positions. The 9 probes belonging to each position are randomly arranged in a 3×3 matrix (depicted by dashed boxes for defect positions 1 and 16). In this arrangement, as shown in (B), the gradient-related variation within the closely spaced 3×3 feature group (belonging to a particular defect position) is significantly smaller than the variation between features (belonging to different defect positions) which are located further apart. Single base mismatches To investigate the positional dependence of single base mismatches and the impact of the mismatch type, we designed microarrays containing comprehensive sets of MM probes derived from a series of twenty-five 16mer probe sequence motifs. As described above, position and type of the mismatch base pair were systematically varied, allowing us later to distinguish between the dominating positional dependence and other influential factors. The features are arranged in groups of four, corresponding to the four possible substituent bases (A, C, G and T) at a particular base position. A group comprises three mismatch probes plus one perfect match probe used for control. Sixteen of these feature groups (one for each base position) are arranged in a square feature block comprising in total 64 features (Figs. 6.3 and 6.2A). 114
Hybridization Assays and Image Analysis Single base bulges Probes containing single base insertions and deletions, owing to an unpaired unpaired nucleotide form bulged duplexes (see Fig. 2.16) with reduced stability. A comprehensive study on the impact of single base insertions was performed. The experiment comprised about 1000 single base insertion probes (insertion base type and position systematically varied) derived from twelve 20 to 25mer probe sequence motifs. The feature arrangement is similar to that in Fig. 6.2A. Direct comparison of single base MMs and single base bulges Probe sets were derived from 16mer probe sequence motifs, complementary to the targets in Table 6.1. For each of the 16 possible defect positions a subset of 9 probes (comprising four single base insertions, one base deletion, three MMs and one PM probe) has been created. To prevent that regular arrangement of the defect types can create a systematic bias on measurement (e.g. due to increased target depletion near the PM probes), the subsets of 9 probes were randomly arranged in 3×3 matrices as shown in Fig. 6.2B. 6.4 Hybridization Assays and Image Analysis 6.4.1 Oligonucleotide Targets DNA and RNA target oligonucleotides (Tab. 6.1) were synthesized by MWG Biotech AG (Ebersberg, Germany) and by IBA Nucleic Acids Synthesis (G¨ottingen, Germany). 5’-Cy3 markers were attached in the final coupling step of the oligonucleotide synthesis via coupling of Cy3-phosphoramidite. The 3’-Cy3 modifications were produced postsynthetically by linkage of amino-reactive NHS-esters. Gibbs free energies ∆G◦ 37 andmelting temperaturesTmofthe PM-duplexes(predicted with the DINAMelt server - two-state hybridization) are provided in Tab. 6.2. Target secondary structure could not be avoided completely - in particular for the longer sequences and for the more stable RNA sequences. Possible target oligonucleotide secondary structure (loop and hairpin formation) was investigated with the DINAMelt Server [Mar05] (see Tab. 6.2). 6.5 Dominant Influence of the Defect Position The ”defect profile” plots (plots of the normalized hybridization signal vs. defect position - e.g. in Figs. 6.4 and 6.15) show that the dominant parameter determining oligonucleotide probe-target-affinity - on the microarray surface - is the position of the defect. 115
Single Base Defects - Microarray Experiments Table 6.1: Fluorescently labeled DNA and RNA target oligonucleotides Name Target sequence (5’→3’) Label Length (nt) URA DNA ACTACAAACTTAGAGTGCAG... 5’-Cy3 38 ...CAGAGGGGAGTGGAATTC NIE DNA ACTCGCAAGCACCACCCTATCA 3’-Cy3 22 LBE DNA GTGATGCTTGTATGGAGGAA... 3’-Cy3 30 ...TACTGCGATT PET DNA ACATCAGTGCCTGTGTACTAGGAC 3’-Cy3 24 BEI DNA ACGGAACTGAAAGCAAAGAC 3’-Cy3 20 COM DNA AACTCGCTATAATGACCTGGACTG 5’-Cy3 24 NCO DNA TAGTGGGAGTTGTTAGTGATGTGA 3’-Cy3 24 PET RNA ACAUCAGUGCCUGUGUACUAGGACA 5’-Cy3 25 LBE RNA GUGAUGCUUGUAUGGAGGAA 5’-Cy3 34 ...UACUGCGAUUCGAU COM RNA AACUCGCUAUAAUGACCUGGACUG 5’-Cy3 24 Table 6.2: Gibbs free energies and melting temperatures of PM duplexes and target secondary structures (DINAMelt server [Mar05]), T=37◦C, [Na+]= M, strand concentration 1 nM. The targets COM (DNA) and NCO don’t form relevant secondary structures. For RNA/DNA duplexes no data on duplex stability is available (NDA). PM duplex Target secondary structure Target Duplex ∆G◦ 37 in Tm∆G◦ 37 in Tm name type kcal/mol in ◦C kcal/mol in ◦C URA DNA/DNA -48.1 77.5 -0.1 40.1 NIE DNA/DNA -29.2 67.1 0.5 27.6 LBE DNA/DNA -36.6 70.7 -1.16 45.3 LBE RNA/DNA NDA NDA -7.1 63.0 PET DNA/DNA -29.6 66.1 -1.23 54.5 PET RNA/DNA NDA NDA -1.23 54.5 BEI DNA/DNA -24.2 59.6 0.08 35.1 COM DNA/DNA -28.7 64.5 - - COM RNA/DNA NDA NDA -0.1 37.4 NCO DNA/DNA -28.7 65.1 - - 116
Dominant Influence of the Defect Position A T G C A T G C 1 2 Figure 6.3: Fluorescence micrograph of two neighboring feature blocks in the 16mer mismatch experiment. The shading-corrected image shows two feature blocks corresponding to two different 16mer probe sequence motifs (3’-TTGAGCGATATTACTG5’ to the left, and 3’-TATTACTGGACCTGAC-5’ to the right) both hybridizing with the fluorescently labeled target sequence COM (5’-Cy3-AACTCGCTATAATGACCTGGACTG-3’). The different hybridization signal intensities of the two feature blocks are owing to different binding affinities of the two probe sequence motifs. The feature size is 21 µm. Each feature block comprises all single base mismatches that can occur in the corresponding probe sequence motif. Groups of four features (as indicated by the marked groups 1 and 2) correspond to each one of the 16 possible mismatch base positions. As indicated by the letters between the feature blocks the uppermost row of features in each group corresponds to an A base at the corresponding base position, followed by probes with C, G and T (see also Fig. 6.2). The brightest feature within each group corresponds to the perfect matching probe. Nonhybridized targets in the hybridization solution contribute to the background intensity between the features. The ”mismatch defect profile” for the probe sequence motif 3’-TATTACTGGACCTGAC-5’ is shown in Fig. 6.4. Defects near the duplexends are distinctly less destabilizing than defects in the center of the duplex. As shown in Fig. 6.4 the hybridization signals of the individual mismatch probes are lined-up along the trough-like ”mean profile” curve (solid black line). A parabolic fit can providea reasonable approximationfor the average position dependence obtained from a large number of different sequence motifs (as shown in [Wic06; Poz06]). The discrimination between PM and MM hybridization signals is largest if the defect is located in the middle of the duplex. For 16mer duplexes (as shown in Fig. 6.4) a single base mismatch (MM) in the center typically yields 0-40% of the perfect match (PM) hybridization signal, whereas at the duplex ends defects have significantly less impact on the hybridization signal. The discrimination between PM and point-mutated probes depends on the stability of the particular probe sequence motif: The more stable 25mer probes (shown in Fig. 6.15) are less discriminative than the shorter 16mer probes (Figs. 6.4A and 6.19). Reduced discrimination is also observed (see Fig. 6.5) for sequences which are stabilized by a high CG-content . 117
Single Base Defects - Microarray Experiments 0 5 10 15 0.05 0.1 0.15 0.2 0.25 0.3 0.35 0.4 Mismatch base position Hybridization signal (a.u.) 0 5 10 15 -0.05 0 0.05 Mismatch base position Deviation from the mean (a.u.) 0 5 10 15 T A T T A C T G G A C C T G A C Mismatch base position A B Chigher lower Figure 6.4: The mismatch defect profile (A) (hybridization signal versus defect base position) was obtained from the analysis of the hybridization signals of the feature block shown in the right part of Figure 6.3. The probe sequence motif 3’- TATTACTGGACCTGAC-5’ is complementary to the target oligonucleotide COM. The different types of base substitutions are highlighted by different markers (A red crosses; C green circles; G blue stars; T cyan triangles). The black line indicates the mean profile (moving average of all mismatch hybridization signals over positions p−2 to p+ 2). PM probes (grey symbols) are used as a control to detect systematic bias (gradient effects) on the hybridization signal. The variation of the PM probe intensities also provides an estimate for the error of the measurement. Bias-related deviations between distant features, owing to gradient effects, are expected to be larger than the errors between the compactly arranged features corresponding to the same defect position. (B) Deviation profile. The strong position dependent component of the hybridization signal was eliminated by subtraction of the mean profile. In the following the hybridization signal deviation from the mean profile is referred to as δImp.(C) Comparison of mean mismatch hybridization signals (average of the three mismatch hybridization signals at a particular defect position) at the sites of C·G base pairs to mean MM hybridization signals at the site of adjacent A·T base pairs. A marker (red star: A·T; blue circle C·G) is set in the upper row if the hybridization signals of the mismatches at the corresponding site is higher than that at the adjacent site; otherwise a marker is set in the lower row. We noticed that mismatched base pairs substituting a C·G base pair usually have systematically lower hybridization signals than mismatches substituting a neighboring A·T base pair. 118
Dominant Influence of the Defect Position The positional influence observed in the mean profiles is largely determined by the defectto-end distance, but is superimposed by a sequence dependent contribution. The variation of the shapes of the mean insertion profiles in Fig. 6.5 indicates that the impact of a defect is affected by the stability of the local sequence environment (i.e. not only by the next nearest neighbor base pairs). We discovered that single base bulge defects, originating from single base insertions (Fig. 6.15) and deletions (Fig. 6.19) - within the individual defect profiles - display the same positional dependence as single base mismatch defects. An attempt to explain the origin of defect positional influence is made in section 7. To investigate other factors influencing oligonucleotide duplex binding affinity (e.g. defect type and defect neighborhood) the dominating positional influence needs to be eliminated. Design (selection and arrangement of probes) and analysis of our experiments enable separation of the different influential factors. 0 5 10 15 20 25 30 35 40 0 0.2 0.4 0.6 0.8 1 T G A T G T T T G A A T C T C A C G T C G T C T C C C C T C A C C T T A A G Insertion base position Hybridization signal (a.u.) 2 134 Figure 6.5: The impact of defects is affected by the local sequence environment. Normalized single base insertion profiles (hybridization signal plotted versus the insertion base position) of four 25mer probe sequence motifs complementary to the same target sequence (URA - shown below). The probe motifs 1 to 4 hybridize to different sections of the target oligonucleotide. Mean profiles (bold lines) were obtained from the moving average of the particular insertion profiles (individual hybridization signals are shown as faint grey symbols - profile 4 is shown in detail in Figure 6.15A). The mean profiles 1 to 3 have a distinct minimum between base positions 15 to 20. The stabilizing CG-rich region between base positions 15 and 33 is the reason for the reduced MM discrimination in profile 4. Discussion We observe a dominating influence of the defect position on duplex binding affinity. Defects located in the center of the oligonucleotide duplexes are significantly more destabilizing than defects at the ends. 119
Single Base Defects - Microarray Experiments Strong influence of MM position has been reported previously mainly by other microarray based studies, but also from hybridization experiments in solution. •From optical melting studies (on 7mer RNA/RNA duplexes in solution) Kierzek et al. [Kie99] report a 0.5 kcal/mol stabilization increment per each base position that the defect is closer to the helix end. A positional influence was observed for U·U and A·A, whereas the G·G mismatch stability was largely unaffected by the position. •Dorris et al. [Dor03] found a similar positional influence for 2-base MM and 3-base MM probes on CodeLink 3D gel arrays. They also report a strong correlation (including the positional influence) between solution-phase melting temperatures and microarray hybridization signals of the MM duplexes. •More recently Wick et al. [Wic06] and Pozhitkov et al. [Poz06] reported a strong influence of the defect position on the binding affinity of single base MM duplexes on DNA microarrays. In accordance with [Poz06] we have identified MM position (relative to the duplex ends) as the strongest influential factor on the hybridization signal, when compared to MM-type (determined by the mismatch base pair X·Y) and nearest neighbors.5 To our knowledge only two studies ([Kie99] and [Dor03]) report a defect positional influence for hybridization in solution. This may partly be due to the unavailability of a large number of of appropriate probes for a systematic study. So far the strong positional influence, mostly observed in microarray experiments, is unexplained. The observation of a strong position dependence is in conflict with the two-state nearestneighbor model of DNA duplex thermal stability, where the thermodynamics of internal mismatches is treated as independent of the MM position [San04]. Also, oligonucleotide duplex stability prediction software (based on a multi-state model) underestimates the MM positional influence when compared to microarray hybridization assays [Wic06]. For single base bulge defects we observed a very similar position dependence as for single base mismatches. Also, the magnitudes of the impacts of the MMs and base bulges on the hybridization signal are very similar (apart from the relative high binding affinity of Group II bulges). This consistency suggests a common origin of the positional influence, expected to be independent of the defect type. Sterical crowding at the surface as discussed by Peterson et al. [Pet02] could possibly introduce a positional dependence on the hybridization signals of defect probes. Reduced accessibility of the probes surface-bound 3’-ends can in principle decrease the impact of 5According to the nearest-neighbor model the flanking base pairs towards both sides of the mismatched base pair X·Y– just like the mismatched base pair itself – determine the base stacking interactions. 120
Mismatch Discrimination in DNA/DNA Duplexes defects located near the 3’-end, and thus result in increased hybridization signals of the corresponding probes. This, however, runs contrary to the largely symmetrical intensity profiles observed (Fig. 6.4) and therefore does not provide a satisfactory explanation for the influence of defect position. Focusing on individual probe sequence motifs, we observe, that the positional influence is not simply a function of the defect-to-end distance: it rather has a sequence-dependent contribution. This indicates that the mismatch discrimination could be affected by the stability of the nearest neighbor pairs between the defect and the proximate duplex end. The observed influence of the duplex sequence and the symmetry of the defect positional influence with respect to both duplex ends suggest that end-domain opening (i.e. sequential unzipping of the double-helix from the duplex ends) is the key mechanism for understanding the influence of defect position on duplex stability. 6.6 Mismatch Discrimination in DNA/DNA Duplexes 6.6.1 Experimental Results For statistical analysis of MM type and nearest-neighbor influences the superimposed positional influence needs to be eliminated. This is achieved by subtraction of the (moving average) mean profile. In the following the hybridization signal deviation from the mean profile is referred to as δImp. The resulting position-independent defect profile (for simplicity we keep using the expression ”defect profile”) comprising defect-type and flanking base pair influences only, is shown in Fig. 6.4B. In the following we use the notation of the mismatch base pair X·Yconsisting of the mismatched base Xin the probe sequence and the base Yin the target sequence. In our experiments the systematic variation was restricted to the bases Xin the microarray probe sequences. Since we had only a limited set of fluorescently labeled target oligonucleotides available (see Tab. 6.1) - the target sequences with the bases Yremained unchanged. To investigate how the particular MM-types X·Yaffect duplex stability we measured probetarget-affinities for 25 different probe sequence motifs (distributed over three different microarrays). The PM hybridization signals of the 16mer probe sequence motifs display a strong variation (up to a factor of 20). The absolute hybridization signals from different probe sets are therefore not directly comparable. However, since the relative intensities (of the various MM probes) within the probe sets are largely unaffected by this variation, we can normalize the ”position-independent defect profiles” by division by their standard 121
Single Base Defects - Microarray Experiments mutations the sequence of the guide strand was preserved) purine-purine MMs resulted in the least silencing of gene activity, whereas U·G, C·U and U·U mismatches resulted in a very efficient gene silencing (see Fig. 6.8c).8”A favored model is that purine-purine mismatches disrupt RISC activity by preventing the formation of a conventional A-form helix between the guide strand and the target mRNA, a structural requirement for RISCmediated cleavage” [RL06]. Interestingly, the reported reduced stability of purine-purine mismatches is in good agreement with the findings of Pozhitkov et al.. However, the inferred RNA/RNA MM stability order in Fig. 6.8c, like that in in [Sug00], is not normalized with respect to the corresponding PM stabilities, but rather reflects the absolute impact of the MM base pairs in a given duplex sequence. Differences between MM discrimination in DNA/DNA hybridization and RNA/DNA hybridization are not surprising since DNA/DNA duplexes (under the experimental conditions employed) occur as B-form helices, whereas RNA/DNA and RNA/RNA duplexes commonly occur as A-form helices (see Fig. 2.5). The apparent discrepancy between the stability orders in the studies discussed above (see Fig. 6.8)motivateda systematiccomparisonofsinglebaseMMdiscriminationinDNA/DNA and RNA/DNA duplexes (see section 6.8). 6.7 Influence of Flanking Base Pairs on Single Base Mismatch Binding Affinities in DNA/DNA Microarray Hybridization Due to stacking interactions the destabilizing impact of a mismatchdefect not only depends on the MM base pair X·Y, but also on the flanking Watson-Crick base pairs A·Aand B·B on both sides of the defect. [Alk82; Sug86]. 50−A Y B −30 30−A XB −50 For a systematic study of the next-nearest-neighbor influence the mismatch hybridization signal data was categorized not only according to the the mismatch type (as discussed in section 6.6), but also according to the flanking base pairs at both sides of the mismatched base pair. 8[Sch06]: ”Mismatches to be well accommodated in an A-form RNA/RNA helix (pyrimidine:pyrimidine, pyrimidine:purine, or purine:pyrimidine) displayed intermediate levels of discrimination, whereas purine:purine mismatches, expected either to destabilize the helix or to promote a stable, but nonhelical, conformation, silenced the reporter least.” 128
Influence of Flanking Base Pairs There are 16 neighborhood classes (combinations of A·Aand B·B) for each of the 12 mismatch types X·Y. -1 -0.5 0 0.5 1 1.5 Mismatch base pair X .Y Deviation from the moving average profile (a.u.) 5'-TYT-3' 3'-AXA-5' 5'-TYG-3' 3'-AXC-5' 5'-TYC-3' 3'-AXG-5' 5'-TYA-3' 3'-AXT-5' 5'-GYT-3' 3'-CXA-5' 5'-GYG-3' 3'-CXC-5' 5'-GYC-3' 3'-CXG-5' 5'-GYA-3' 3'-CXT-5' 5'-CYT-3' 3'-GXA-5' 5'-CYG-3' 3'-GXC-5' 5'-CYC-3' 3'-GXG-5' 5'-CYA-3' 3'-GXT-5' 5'-AYT-3' 3'-TXA-5' 5'-AYG-3' 3'-TXC-5' 5'-AYC-3' 3'-TXG-5' 5'-AYA-3' 3'-TXT-5' AA CA GA TA AC CC GC TC AG CGGGTG AT CT GT TT Figure 6.9: Distribution of the median hybridization signal values (deviation from the moving average profiles) of the various MM neighborhoods classes (see legend) as shown in Figs. A.12 - A.22. Red symbols denote C·G neighbors only, blue symbols denote A·T neighbors only. Green symbols correspond to mixed neighbors. The maximum value of about 1.3 a.u. for A·G MMs (green up-pointing triangle) is probably an outlier (only a single measurement was available for that particular MM class), whereas the value of 0.74 (red star) is based on 10 measurements. The significance of individual data points (which can be affected by lack of experimental data) can be evaluated from the corresponding histograms in Figs. A.12 - A.22. Splitting of the experimental data into 192 subsets (see Figs. A.12 - A.22) results in a relatively small statistical base for the individual MM classes (→large statistical errors and sequence dependent bias). The significance of individualdata points can be evaluated from the corresponding histograms in Figs. A.12 - A.22. The median values of the neighborhood-dependent MM hybridization signal9distributions are shown in Fig. 6.9. To investigate if the experimentally observed influence of flank9normalized hybridization signals, positional influence eliminated 129
Single Base Defects - Microarray Experiments AA CA GA TA AC CC GC TC AG CG GG TG AT CT GT TT -5 -4 -3 -2 -1 0 1 2 3 AA CA GA TA AC CC GC TC AG CG GG TG AT CT GT TT 0 1 2 3 4 5 6 δ ∆G (kcal/mol) Base pair X.Y Base pair X.Y ∆G (kcal/mol) 5'-TYT-3' 3'-AXA-5' 5'-TYG-3' 3'-AXC-5' 5'-TYC-3' 3'-AXG-5' 5'-TYA-3' 3'-AXT-5' 5'-GYT-3' 3'-CXA-5' 5'-GYG-3' 3'-CXC-5' 5'-GYC-3' 3'-CXG-5' 5'-GYA-3' 3'-CXT-5' 5'-CYT-3' 3'-GXA-5' 5'-CYG-3' 3'-GXC-5' 5'-CYC-3' 3'-GXG-5' 5'-CYA-3' 3'-GXT-5' 5'-AYT-3' 3'-TXA-5' 5'-AYG-3' 3'-TXC-5' 5'-AYC-3' 3'-TXG-5' 5'-AYA-3' 3'-TXT-5' A B 37 37 Figure 6.10: Influence of flanking base pairs on MM duplex stability. (A) Gibbs free energies ∆G◦ 37 of mismatched and perfect-matching DNA/DNA trinucleotide duplexes were calculated from MM nearest-neighbor parameters [All97]. C·G flanking base pairs (red markers) are consistently stabilizing, whereas A·T flanking base pairs (blue markers) have a destabilizing influence. (B) Gibbs free energy increments δ∆G◦ 37 between MM and corresponding PM duplexes. In the two-state nearest-neighbor model the discrimination between single base MM and PM duplexes only depends on the identity of the affected trinucleotide sequence (MM base pair and flanking base pairs). δ∆G◦ 37 does not depend on the rest of the duplex sequence or on the position of the defect (unless the defect is located at a terminal position). 130
Influence of Flanking Base Pairs 0123456 -1 -0.5 0 0.5 1 1.5 Hybridization signal (a.u.) δ ∆G37 (kcal/mol) 0123456 -1 -0.5 0 0.5 1 1.5 Hybridization signal (a.u.) δ ∆G37 (kcal/mol) 5'-TYT-3' 3'-AXA-5' 5'-TYG-3' 3'-AXC-5' 5'-TYC-3' 3'-AXG-5' 5'-TYA-3' 3'-AXT-5' 5'-GYT-3' 3'-CXA-5' 5'-GYG-3' 3'-CXC-5' 5'-GYC-3' 3'-CXG-5' 5'-GYA-3' 3'-CXT-5' 5'-CYT-3' 3'-GXA-5' 5'-CYG-3' 3'-GXC-5' 5'-CYC-3' 3'-GXG-5' 5'-CYA-3' 3'-GXT-5' 5'-AYT-3' 3'-TXA-5' 5'-AYG-3' 3'-TXC-5' 5'-AYC-3' 3'-TXG-5' 5'-AYA-3' 3'-TXT-5' AA CA GA AC CC TC AG GG TG CT GT TT Mismatch Base Pair Flanking Base Pairs A B Figure 6.11: Comparison of MM hybridization signals (normalized with respect to PM hybridization signals - thus representing a measure for MM discrimination) with predicted Gibbs free energy increments δ∆G◦ 37. Hybridization signals (as shown in Fig. 6.9) are categorized according to MM base pair type and according to flanking base pairs. Each data point represents the median value of a distribution of hybridization signals (in detail shown in Figs. A.12 to A.22). We observe a significant correlation between the MM hybridization signal and the predicted Gibbs free energy increment δ∆G◦ 37. Part (A) highlights the influence of flanking base pairs on MM discrimination. Flanking A·T base pairs on both sides of the defect (blue symbols) result (on average) in smaller hybridization signals than C·G-only (red symbols) or mixed flanking base pairs (green symbols). However, the influence of flanking base pairs is little consistent compared with the influence of the MM base pair type, which is highlighted in (B): The discrimination of G·G, A·A and T·G mismatches is larger than predicted by MM nearest-neighbor parameters from [All97] and larger than in a similar experiment in [Wic06]. 131
Single Base Defects - Microarray Experiments ing base pairs on binding affinities is in agreement with MM nearest-neighbor parameters [All97], we compared our experimental data (Fig. 6.9) to predicted free energy increments between MM and PM duplexes (Fig. 6.10): The MM nearest-neighbor parameters from [All97] predict a stabilizing influence of C·G flanking base pairs. Fig. 6.10A shows a consistently increased stability of those duplexes with C·G next nearest neighbors only, whereas a systematically decreased stability is seen for duplexes with A·T nearest neighbors only. For the predicted difference δ∆G◦ 37 between PM and MM free energies - which is expected to be reflected in the experimentally determined MM discrimination - this consistency is somewhat reduced (see Fig. 6.10B). The comparison of δ∆G◦ 37 with experimentally determined hybridization signals in Fig. 6.11A confirms a significant influence of flanking base pairs. On average, flanking A·T base pairs result in smaller hybridization signals than C·G or mixed flanking base pairs. However, the influence of the MM-base pairs X·Yon the MM binding affinity (see Fig. 6.11B) is distinctly more consistent than the influence of flanking base pair types. A larger scale investigation of flanking base pair influence (based on a much larger set of oligonucleotide target sequences/probe sequence motifs) would be necessary to increase the statistical significance of the above results. 6.8 Mismatch Discrimination in DNA/DNA and RNA/DNA Duplexes - a Direct Comparison To investigate if the above results from DNA/DNA hybridization also apply to hybridization of RNA/DNA duplexes we performed a direct comparison between DNA/DNA hybridization and RNA/DNA hybridization (employing DNA targets and equivalent RNA target sequences - see Tab. 6.1 ) on the same microarray. 6.8.1 Outline of the Experiment The experiment is basically identical with the experiments described in section 6.6. Hybridization assays are conducted with fluorescently labeled DNA targets and corresponding RNA target sequences (Table 6.1). To avoid fabrication-related variation of the hybridization signals the DNA and RNA hybridization assays were performed on the same chip, first with RNA target oligonucleotides and - after regeneration of the microarray with NaOH (selective degradation of RNA targets) - with the corresponding DNA targets. Three different microarrays were fabricated, each one focussing on one particular target sequence (COM,PET and LBE). The individual microarrays comprise single base MM and insertion probes (→single base bulges) for 6 different probe sequence motifs (probing 132
Mismatch Discrimination in DNA/DNA and RNA/DNA Duplexes different 16 to 20mer subsequences of the target sequence). Two replicates of each feature block provide a test for the reproducibility of the measurement. The subsets of data obtained from the individual microarrays were analyzed independently to check the consistency of the observed results: apart from small sequencerelated biases the three microarrays provided basically the same results. Hybridization was performed with 1 nM target solutions in 5×SSPE (0.01% Tween-20TM). Hybridization temperatures were 30◦Cfor PET and LBE and 40◦Cfor COM (for the target sequence COM the temperature had to be increased to 40◦Csince local depletion led to inhomogeneous hybridization - see section 8.5). 6.8.2 Results The influence of the defect position is very similar for the DNA/DNA and the RNA/DNA binding affinities (see Fig. A.1). However, there are small, though reproducible differences, as the comparison between replicate feature blocks (see Figs. A.2 - A.7) shows. For single base bulges no defect type specific differences between RNA/DNA and DNA/DNA hybridization were found. We observed that under equivalent hybridization conditions the hybridization signal from RNA targets is on average about 1.3 times brighter than that of the corresponding DNA targets. This is anticipated: RNA targets have a slightly larger binding affinity than DNA targets since stacking interactions are stronger in A-form RNA/RNA and RNA/DNA duplexes than in B-DNA duplexes.10 Differences between MM stabilities in DNA/DNA and RNA/DNA duplexes The MM discrimination in RNA/DNA duplexes (Fig. 6.12B) is very similar to that in DNA/DNA duplexes (Fig. 6.12A). However, a closer look reveals systematic differences between DNA/DNA and RNA/DNA hybridization. A statistical analysis (Figs. 6.12 and 6.14) revealed that purine-purine MMs are less stable in RNA/DNA duplexes (Fig. 6.14c) than in DNA/DNA duplexes (Fig. 6.14b). Three independent experiments (performed on different microarrays and with different probe/target sequences) provided the same trends. The decrease of purine-purine MM stabilities becomes obvious in the ranking order of differences between RNA/DNA and DNA/DNA MM stabilities (Fig. 6.14d). The largest differences between RNA/DNA and DNA/DNA MMs are observed for the MM-types G·A and A·G (which are more stable in DNA/DNA duplexes) and, with reversed sign, for the MM-type T·G, which is significantly more stable in RNA/DNA duplexes. 10 Binding affinities: RNA/RNA >RNA/DNA >DNA/DNA 133
Single Base Defects - Microarray Experiments AA AC AG CA CC CT GA GG GT TC TG TT DNA/DNA mismatch hybridization signal -0.5 0 0.5 1 1.5 AA AC AG CA CC CU GA GG GU TC TG TU Hybridization signal (a.u.) RNA/DNA mismatch hybridization signal A B-0.5 0 0.5 1 1.5 Hybridization signal (a.u.) Figure 6.12: Comparison of DNA/DNA and RNA/DNA mismatch hybridization signals - statistical analysis. (A) MM-type related influence in DNA/DNA oligonucleotide duplexes. The positional influence was eliminated by subtraction of the moving average MM profile. Subsequent normalization was performed by division through the mean hybridization signal of the particular MM profile. (B) MM-type related influence in RNA/DNA oligonucleotide duplexes. 134
Mismatch Discrimination in DNA/DNA and RNA/DNA Duplexes -1 -0.5 0 0.5 1 AA AC AG CA CC CT GA GG GT TC TG TT Hybridization signal (a.u.) Differences between MM hybridzation signals of RNA/DNA and DNA/DNA duplexes Figure 6.13: Differences between RNA/DNA and DNA/DNA MM binding affinities. Largest differences between RNA/DNA and DNA/DNA have been found for the MMtypes T·G, G·A and A·G. 6.8.3 Discussion Our investigation on the impact of MM-types in DNA/DNA oligonucleotide duplexes revealed that single base mismatches substituting C·G base pairs are more destabilizing than mismatches substituting A·T base pairs. However, this seemingly plausible result (shown in Fig. 6.6) is not in general agreement with previous work [Sug00; Wic06; Poz06; Sch06] on the influence of the MM type on binding affinities. Our direct comparison (”direct” in the sense of using the same probe sequences on the same microarray) between DNA/DNA and RNA/DNA hybridization on microarrays reveals - for RNA/DNA duplexesan increased destabilization of purine-purine mismatches, with respect to other MM types. However, we did not observe such a distinct impact of purine-purine MMs as reported in [Poz06] and [Sch06]. Rather the MM stability order was very similar to that for DNA/DNA hybridization. From MM stabilityorders in Figs. 6.14c and 6.14b (and Fig. 6.8e) we infer that the stability of MMs in RNA/DNA duplexes is determined by two factors: •In RNA/DNA duplexes purine-purine MMs tend to be more destabilizing (with respect 135
Single Base Defects - Microarray Experiments T U T U U T⋅ ≥ > ⋅ ≈ ⋅ ≈ ⋅ ≈ ⋅ ≈ ⋅ ≈ > ⋅ ≥ ≥ ⋅ >C C C C C CG A G G A A A A A G G G⋅ ⋅ ⋅ ⋅ T C C C T T C T T C T C⋅ > ⋅ ≥ ⋅ ≥ ⋅ > ⋅ ≈ ⋅ > ⋅ ≈ ⋅ ≥ ≈ > >G A G A A A G G A G G A⋅ ⋅ ⋅ ⋅ G A G A G A A A G A G G⋅ ⋅ ⋅ ⋅> ⋅ ≥ ⋅ > ≥ ⋅ ≈ ⋅ > ≈ ⋅ ≈ ⋅ ≈ ⋅ ≥ ⋅ ≥T T T C T C T C C T C C G A A G G A A A A G G G⋅ ⋅ ⋅ ⋅> ⋅ > ≈ ⋅ ≈ ⋅ ≈ > ⋅ > ⋅ ≥ ⋅ ≥ ⋅ ≈ ⋅ >T T C T T C C T T C C C a) DNA/DNA hybridization (large data set) b) DNA/DNA hybridization (small data set for direct comparison with RNA/DNAhybridization) c) RNA/DNA hybridization (small data set - equivalent to the DNA/DNA dataset in b) d) Difference between RNA/DNA and DNA/RNA hybridization signals. Uracil is treated as thymine. (TG to GT positive; AC to GA negative) Figure 6.14: Ranking orders of DNA/DNA MM stabilities in comparison with that of RNA/DNA MMs. (a) For comparison the DNA/DNA MM stability order from an independent experiment (Fig. 6.8) is shown here again. (b) As anticipated the ranking order for DNA/DNA MMs obtained from the smaller data set which is used for the direct comparison between DNA/DNA and RNA/DNA hybridization (Fig. 6.12A) is very similar. The ranking order for RNA/DNA mismatch stabilities (c) (extracted from Fig. 6.12B) reveals significant differences with respect to (b). In part (d) MM-types are ordered according to the hybridization signal differences between RNA/DNA and DNA/DNA MMs (extracted from Fig. 6.12 A and B). Purine bases are highlighted in blue. to other MM-types) than purine-purine MMs in DNA/DNA duplexes. •The influence of the ”affected base pair” - the base pair which has been substituted by the MM base pair - is the other factor that determines the impact of the MM type. In the experiments the PM hybridization signal is used as a reference value for the reduction of the hybridizationsignal due the MM defect. In agreement with [Wic06] we observed that MMs affecting C·G base pairs are more discriminating than MMs affecting A·T base pairs. In the order of RNA/DNA mismatch stabilities (Fig. 6.14c) the latter effect is superimposed by the destabilizing effect of purine-purine MMs, whereas in DNA/DNA duplexes (Fig. 6.14b - our results - in agreement with [Wic06] - see Fig. 6.8d) an increased destabilization of purine-purine MMs is not observed. An explanation for the observed differences between DNA/DNA and RNA/DNA binding affinities is, that purine-purine MMs cause larger steric hindrance in the A-form hybrid duplexes than in the B-form DNA/DNA duplexes. In this study, like in [Poz02], a destabilizing impact of purine-purine MMs was observed in RNA/DNA hybridization. However, we found only a slightly increased destabilization with respect to the corresponding purine-purine MMs in DNA/DNA duplexes, whereas [Poz02] and [Sch06] reported that purine-purine MMs - in absolute terms - are the most 136
Single Base Bulge Defects discriminating MMs with respect to other MM-types.11 Further studies will be necessary to resolve the remaining discrepancy. A more detailed future investigationof MM stabilities should also focus on the influence of the flanking base pairs. This, however, will require a significantly larger database of MM hybridization signals. 6.9 Single Base Bulge Defects Single base insertions and deletions, owing to a surplus unpaired base in one of the two strands, result in bulged duplexes, which like MM duplexeshave a reduced binding affinity. In duplexes with single base insertion probes the bulged base is located on the surfacebound probe strand, whereas in duplexes with single base deletion probes the bulged base is located on the target strand. The positional dependence of the insertion intensity profiles (Figure 6.15A) is very similar to the mismatch intensity profile in Figure 6.4, though the individual insertion profiles (for example the profile of C-insertions - green circles in Figure 6.15) show large deviations from the (moving average) mean profile. Hybridization signals can be significantly increased over two or more consecutive defect positions. In particular, base insertions next to identical bases (Group II bulges [Zhu99]) result in systematically increased binding affinities - in comparison to insertions of nonidentical bases (Group I bulges). In the notation of Zhu et al. [Zhu99] bulged bases without an identical neighboring base (Fig. 2.17A) are defined as Group I bulges, whereas bulges with at least one identical neighboring base (Fig. 2.17B) are referred to as Group II bulges. Increased stability of duplexes with Group II bulges in solution-phase experiments has been described by Ke et al. [Ke95]. Fig. 6.15C demonstrates the systematically increased binding affinity of Group II bulges in DNA microarray hybridization. 6.9.1 Statistical Analysis The observed stabilization of Group II bulges (in comparison to Group I bulges) in our microarray experiments is surprisingly large (see discussion below): Group II bulges located near the center of 16mer probes often show hybridization signals with a similar intensity as the corresponding PM probe, whereas Group I bulges at the same defect position have a significantly smaller binding affinity, with a similar level as single base MMs at the corre11 These studies, however, investigated only DNA/RNA hybridization and RNA/RNA hybrids (RNAi: A-form helix between the guide strand and the target mRNA), respectively. No comparison with DNA/DNA hybridization was made. 137
Summary/Zusammenfassung allein aufgrund der geringf¨ugigen Stabilisierung infolge dieser Entropiezunahme zu erkl¨aren. Unser Erkl¨arungsansatz beruht auf einer durch den bulge-Defekt verursachten Blockade des Zipper-Mechanismus: Die durch den bulge-Defekt hervorgerufene Verschiebung zwischen den Einzelstrang-Sequenzen (frameshift) verhindert ein schnelles Schließen (zipping up) des Duplex. Diese Barriere kann beim Vorliegen eines Group II bulges – aufgrund der Positionsentartung – schneller ¨ubersprungen werden18 als bei Group I bulge-Defekten (bei welchen keine Positionsentartung vorliegt). Die Bindungsaffinit¨at zwischen Probeund Target-Sequenzen wird sehr stark von der Sekund¨arstruktur der Target-Sequenzen beeinflusst [Lue03]. F¨ur ein Experiment zur Untersuchung des Einflusses solcher Sekund¨arstrukturen (Abschnitt 8.6), wurden fluoreszenzmarkierte cRNA-Targets mit einer L¨ange von 300 bzw. 800 Nukleotiden hergestellt. Bei diesen L¨angen sind stabile intramolekulare Sekund¨arstrukturen zu erwarten, die in den dazugeh¨origen Sequenzabschnitten eine Hybridisierung mit komplement¨aren MicroarrayProbes verhindern. Tats¨achlich konnte in dem tiling-array-Experiment19 nur auf etwa 20 bis 30% der L¨ange dieser Target-Sequenzen eine signifikante Hybridisierung erzielt werden. Mit Hilfe von Sfold [Din04], einem Software-Tool welches u. a. zum Auffinden effektiver Antisense Oligonukleotidedient, wurde untersucht, wiesich die infolge der Sekund¨arstruktur verminderte Zug¨anglichkeit von großen Teilen der Targetsequenz auf die Bindungsaffinit¨at der einzelnen Probesequenzen auswirkt. Unsere Ergebnisse zeigen, dass die mit Hilfe von Sfold auf theoretischer Grundlage (unter Ber¨ucksichtung des Boltzmann-Ensembles von Target-Sekund¨arstrukturen) ermittelten Bindungsaffinit¨aten mit unseren experimentell bestimmten Hybridisierungssignalen korreliert sind. Unsere Ergebnisse legen nahe das Sfold auch zum Auffinden effizienter Microarray-Probe-Sequenzen geeignet ist. Weitere Microarray-Hybridisierungsexperimente mit anderen Target-Sequenzen sind erforderlich um die im Rahmen der vorliegenden Arbeit gewonnenen Ergebnisse zu untermauern. Auf der Basis des double-ended Zipper-Modells [Gib59; Kit69] wurde ein thermodynamisches Modell des Oligonukleotid-Duplexes entwickelt (Kapitel 7), um die experimentellen Ergebnisse, inbesondere den starken Einfluss der Defektposition, genauer zu untersuchen. Im Gegensatz zum in der Praxis am h¨aufigsten verwendeten two-state nearest18 Die Stabilisierung von Group II bulges beruht der erh¨ ohten Wahrscheinlichkeit, dass eine der identischen Basen eine g¨ unstige Konformation einnimmt, bei der ein rasches Fortschreiten des ZippingProzesses m¨ oglich ist. 19 Das tiling-array-Experiment beinhaltet einen Satz von 25mer Probe-Sequenzen die entlang der sehr viel l¨ angeren Target-Sequenz relativ zueinander versetzt angeordnet sind. Diese Art von Experiment verfolgt den Zweck, die Bindungsaffinit¨ at der einzelnen Target-Bereiche zu sondieren. 240
Zusammenfassung neighbor Modell werden beim Zipper Modell auch die an den Enden partiell denaturierten Duplexkonformationen ber¨ucksichtigt. Ausgehend von den nearest-neighbor Wechselwirkungen benachbarter Basenpaare werden f¨ur die einzelnen Duplexkonformationen die statistischen Gewichte und daraus schließlich die Zustandssumme berechnet. Die theoretischen Betrachtungen zeigen, dass die Zustandssumme beim Vorliegen von Einzeldefekten umso gr¨oßer ist, je n¨aher der Defekt bei den Duplexenden liegt. Dies best¨atigen die experimentellen Ergebnisse: Oligonukleotid-Duplexe mit endnahen Defekten sind stabiler als entsprechende Duplexe mit in der Mitte liegenden Defekten. Eine numerische Analyse des Defekt-Positionseinflussesauf die Bindungsaffinit¨atzeigt, dass die Oligonukleotidsequenz, in diesem Fall als Abfolge unterschiedlicher starker nearest-neighbor-Wechselwirkungen betrachtet, wie bei auch experimentell beobachtet, einen signifikanten Einfluss auf die Positionsabh¨angigkeit der Bindungsaffinit¨at haben kann. Dies wird vor allem offensichtlich, wenn innerhalb der Duplex-Sequenz st¨arkere und schw¨achere NN-Paare ungleichm¨aßig verteilt sind. Um die experimentell bestimmten Hybridisierungssignale mit den auf theoretischer Basis ermittelten Duplexstabilit¨aten vergleichen zu k¨onnen wurde in einem Microarray-Hybridisierungsexperiment (Abschnitt 7.4) die L¨ange der Probes – und somit die GibbsEnergie ∆Gder DNA-Duplexe – schrittweise variiert. Wir beobachten einen sigmoidalen Zusammenhang θ(∆G)zwischen dem Anteil hybridisierter Probes und der freien Enthalpie der Duplexe ∆G.¨ Uber einen relativ weiten ¨ Ubergangsbereich nimmt das Hybridisierungssignal n¨aherungsweise linear mit der freien Enthalpie der Duplexe zu. Damit weicht das experimentelle Ergebnis deutlich von einem theoretischen Verlauf ab, der durch die Langmuir-Adsorptionsgleichungbeschrieben wird - dieser weist einen vergleichsweise schmalen ¨ Ubergangsbereich auf. Die Diskrepanz konnte anhand einer numerischen Simulation mit dem Einfluss von Synthesedefekten erkl¨art werden: Die in den Experimenten vorliegende breite Verteilung von Bindungsaffinit¨aten, die durch eine variable Anzahl von Defekten in der Probe-Sequenz hervorgerufen wird (die sich zudem an unterschiedlichen Positionen befinden), resultiert in einem stark verbreiterten ¨ Ubergangsbereich in θ(∆G). Die untersuchte Positionsabh¨angigkeit von Defekten kann auch auf die mehr oder weniger starken NN-Wechselwirkungen von Watson-Crick-Basenpaaren ¨ubertragen werden. Unsere Untersuchungen in Abschnitt 7.5 zeigen: Duplexe, die aus identischen NN-Paaren zusammengesetzt, und somit auf der Grundlage des two-state nearest-neighbor Modell thermodynamisch ¨aquivalent sind, weisen im Zipper-Modell die gr¨oßte Stabilit¨at dann auf, wenn die stabilsten NN-Paare in der Mitte des Duplex und die schw¨achsten NN-Paare entsprechend an den Enden des Duplexes angeordnet sind. Bei Raumtemperatur sind die Ergebnisse des Zipper-Modells mit denen des two-state nearest-neighbor Modells praktisch 241
Summary/Zusammenfassung identisch. Erst mit zunehmender Temperatur ist infolge der verst¨arkten Denaturierung an den Duplexenden die beschriebene Positionsabh¨angigkeit zu beobachten. Dieses Ergebnis liefert erstmals eine theoretische Grundlage f¨ur das bislang nur auf empirischer Basis beschriebene positionsabh¨angige nearest-neighbor Modell (PDNN). Im Rahmen der vorliegenden Arbeit wurde auf der Basis von handels¨ublichen Komponenten ein flexibles System zur in situ-Synthese von DNA-Microarrays entwickelt. Aufgrund seiner technischen M¨oglichkeiten(bei vergleichsweise niedrigen Investitionen),aber auch weil es im Gegensatz zu kommerziellen Microarray-Plattformen keine Black-BoxTechnologie darstellt, d¨urfte das hier im Detail beschriebene System eine interessante Ausgangsbasis f¨ur die Entwicklung von Microarray-Synthesizern sein. Eine (evtl. auf einer ”Open Source”-Basis betriebene) Weiterentwicklung des Microarray-Synthesesystems w¨are w¨unschenswert, damit diese vielversprechende und vielseitig einsetzbare Zukunftstechnologie bald breite Anwendung finden kann. In Hinblick auf die zunehmende Bedeutung der DNA-Microarray Technologie ist ein fundiertes Verst¨andnis der zugrunde liegenden physikalisch-chemischen Zusammenh¨ange erforderlich. Vor allem in Hinblick auf die Untersuchungen zur Detektion von Punktmutationen wurde in der vorliegenden Arbeit dazu beigetragen. 242
Summary/Zusammenfassung 244
Bibliography [AB03] G. Altan-Bonnet, A. Libchaber, and O. Krichevsky. Bubble dynamics in doublestranded DNA. Physical Review Letters, 90(13):138101, April 2003. [Alb03] T. J. Albert, J. Norton, M. Ott, T. Richmond, K. Nuwaysir, E. F. Nuwaysir, K. P. Stengele, and R. D. Green. Light-directed 5 ’- 3 ’ synthesis of complex oligonucleotide microarrays. Nucleic Acids Research, 31(7):e35, April 2003. [Alk82] D. Alkema, P. A. Hader, R. A. Bell, and T. Neilson. Effects of flanking GC basepairs on internal watson-crick, GU, and nonbonded base pairs within a short ribonucleic-acid duplex. Biochemistry, 21(9):2109–2117, 1982. [All97] H. T. Allawi and J. SantaLucia. Thermodynamics and NMR of internal GT mismatches in DNA. Biochemistry, 36(34):10581–10594, August 1997. [Amb05] T. Ambjornsson and R. Metzler. Blinking statistics of a molecular beacon triggered by end-denaturation of DNA. Journal of Physics-Condensed Matter, 17(49):S4305–S4316, 2005. [Amb06] T. Ambjornsson, S. K. Banik, O. Krichevsky, and R. Metzler. Sequence sensitivity of breathing dynamics in heteropolymer DNA. Physical Review Letters, 97(12):128105, September 2006. [And06] D. Andreatta, S. Sen, J. L. P. Lustres, S. A. Kovalenko, N. P. Ernsting, C. J. Murphy, R. S. Coleman, and M. A. Berg. Ultrafast dynamics in DNA: ”fraying” at the end of the helix. Journalof the American Chemical Society, 128(21):6885– 6892, May 2006. [App65] J. Applequist and V. Damle. Thermodynamics of helix-coil equilibrium in oligoadenylicacid from hypochromicitystudies. Journal of the American Chemical Society, 87(7):1450–&, 1965. [Bar06] A. Barthel and M. Zacharias. Conformational transitions in rna single uridine and adenosine bulge structures: A molecular dynamics free energy simulation study. Biophysical Journal, 90(7):2450–2462, April 2006. [Bau03] M. Baum, S. Bielau, N. Rittner, K. Schmid, K. Eggelbusch, M. Dahms, A. Schlauersbach, H. Tahedl, M. Beier, R. Guimil, M. Scheffler, C. Hermann, J. M. Funk, A. Wixmerten, H. Rebscher, M. Honig, C. Andreae, D. Buchner, 245
BIBLIOGRAPHY E. Moschel, A. Glathe, E. Jager, M. Thom, A. Greil, F. Bestvater, F. Obermeier, J. Burgmaier, K. Thome, S. Weichert, S. Hein, T. Binnewies, V. Foitzik, M. Muller, C. F. Stahler, and P. F. Stahler. Validation of a novel, fully integrated and flexible microarray benchtop facility for gene expression profiling. Nucleic Acids Research, 31(23):e151, 2003. [Bea81] S. L. Beaucage and M. H. Caruthers. Deoxynucleoside phosphoramidites a new class of key intermediates for deoxypolynucleotide synthesis. Tetrahedron Letters, 22(20):1859–1862, 1981. [Bei99] M. Beier and J. D. Hoheisel. Versatile derivatisation of solid support media for covalent bonding on DNA-microchips. Nucleic Acids Research, 27:1970–1977, 1999. [Ben02] R. Benters, C. M. Niemeyer, D. Drutschmann, D. Blohm, and D. Wohrle. DNA microarrays with PAMAM dendritic linker systems. Nucleic Acids Research, 30(2):e10, January 2002. [Bha03] G. Bhanot, Y. Louzoun, J. H. Zhu, and C. DeLisi. The importance of thermodynamic equilibrium for high throughput gene expression arrays. Biophysical Journal, 84(1):124–135, January 2003. [Bin04] H. Binder, T. Kirsten, M. Loeffler, and P. F. Stadller. Sensitivity of microarray oligonucleotide probes: Variability and effect of base composition. Journal of Physical Chemistry B, 108(46):18003–18014, 2004. [Bin06] H. Binder. Thermodynamics of competitivesurface adsorption on DNA microarrays. Journal of Physics-Condensed Matter, 18(18):S491–S523, 2006. [Bla96] A. P. Blanchard, R. J. Kaiser, and L. E. Hood. High-density oligonucleotide arrays. Biosensors & Bioelectronics, 11(6-7):687–690, 1996. [Blo03] R. Blossey and E. Carlon. Reparametrizing the loop entropy weights: Effect on DNA melting curves. Physical Review E, 68(6):061911, December 2003. [Bre86] K. J. Breslauer, R. Frank, H. Blocker, and L. A. Marky. Predicting DNA duplex stability from the base sequence. Proceedings of the National Academy of Sciences of the United States of America, 83(11):3746–3750, June 1986. [Cam06] A. M. Caminade, C. Padie, R. Laurent, A. Maraval, and J. P. Majoral. Uses of dendrimers for DNA microarrays. Sensors, 6(8):901–914, August 2006. [Car06] E. Carlon and T. Heim. Thermodynamics of RNA/DNA hybridization in highdensity oligonucleotide microarrays. Physica A-Statistical Mechanics and its Applications, 362(2):433–449, April 2006. [Cha05] C. Y. Chan, C. E. Lawrence, and Y. Ding. Structure clustering features on the sfold web server. Bioinformatics, 21(20):3926–3928, October 2005. 246
BIBLIOGRAPHY [Che07] W. W. Chen, S. Kirihara, and Y. Miyamoto. Fabrication of three-dimensional micro photonic crystals of resin-incorporating TiO2 particles and their terahertz wave properties. Journal of the American Ceramic Society, 90(1):92–96, January 2007. [Chi05] P. Y. Chiou, A. T. Ohta, and M. C. Wu. Massively parallel manipulation of single cells and microparticles using optical images. Nature, 436(7049):370– 372, 2005. [Cog91] J. A. H. Cognet, J. Gabarroarpa, M. Lebret, G. A. Vandermarel, J. H. Vanboom, and G. V. Fazakerley. Solution conformation of an oligonucleotide containing a GG mismatch determined by nuclear-magnetic-resonance and molecular mechanics. Nucleic Acids Research, 19(24):6771–6779, December 1991. [Con83] B. J. Conner, A. A. Reyes, C. Morin, K. Itakura, R. L. Teplitz, and R. B. Wallace. Detection of sickle-cell beta-s-globin allele by hybridization with synthetic oligonucleotides. Proceedings of the National Academy of Sciences of the United States of America, 80(1):278–282, 1983. [Cra71] M. E. Craig, D. M. Crothers, and P. Doty. Relaxation kinetics of dimer formation by self complementary oligonucleotides. Journal of Molecular Biology, 62(2):383–&, 1971. [Cri70] F. Crick. Central dogma of molecular biology. Nature, 227(5258):561–&, 1970. [Cro64] D. M. Crothers and B. H. Zimm. Theory of melting transition of synthetic polynucleotides: Evaluation of stacking free energy. Journal of Molecular Biology, 9(1):1–&, 1964. [Cue04] J. A. Cuesta and A. Sanchez. General non-existence theorem for phase transitions in one-dimensional systems with short range interactions, and physical examples of such transitions. Journal of Statistical Physics, 115(3-4):869–893, May 2004. [Dan07] D. S. Dandy, P. Wu, and D. W. Grainger. Array feature size influences nucleic acid surface capture in DNA microarrays. Proceedings of the National Academy of Sciences of the United States of America, 104(4):8223–8228, February 2007. [Dau93] T. Dauxois, M. Peyrard, and A. R. Bishop. Dynamics and thermodynamics of a nonlinear model for DNA denaturation. Physical Review E, 47(1):684–695, January 1993. [Der05] G. Derra, H. Moench, E. Fischer, H. Giese, U. Hechtfischer, G. Hensler, A. Koerber, U. Niemann, F. C. Noertemann, P. Pekarski, J. Pollmann-Retsch, A. Ritz, and U. Weichmann. Uhp lamp systems for projection applications. Journal of Physics D-Applied Physics, 38(17):2995–3010, September 2005. [Deu04] J. M. Deutsch, S. Liang, and O. Narayan. Modelling of microarray data with zippering. Preprint q-bio.BM/0406039 v1, 2004. arXiv:cond-mat/0304567. 247
BIBLIOGRAPHY [Din01] Y. Ding and C. E. Lawrence. Statistical prediction of single-stranded regions in RNA secondary structure and application to predicting effective antisense target sites and beyond. Nucleic Acids Research, 29(5):1034–1046, March 2001. [Din03] Y. Ding and C. E. Lawrence. A statistical sampling algorithm for RNA secondary structure prediction. Nucleic Acids Research, 31(24):7280–7301, December 2003. [Din04] Y. Ding, C. Y. Chan, and C. E. Lawrence. Sfold web server for statistical folding and rational design of nucleic acids. Nucleic Acids Research, 32:W135–W141, July 2004. [Din05] Y. Ding, C. Y. Chan, and C. E. Lawrence. RNA secondary structure prediction by centroids in a boltzmann weighted ensemble. RNA-A Publication of the RNA Society, 11(8):1157–1166, August 2005. [Dod77] J. B. Dodgson and R. D. Wells. Synthesis and thermal melting behaviour of oligomer-polymer complexes containing defined lengths of mismatched da.dg nucleotides. Biochemistry, 16(11):2367–2374, 1977. [Dor03] D. R. Dorris, A. Nguyen, L. Gieser, R. Lockner, A. Lublinsky, M. Patterson, E. Touma, T. J. Sendera, R. Elghanian, and A. Mazumder. Oligodeoxyribonucleotide probe accessibility on a three-dimensional DNA microarray surface and the effect of hybridization time on the accuracy of expression ratios. BMC Biotechnology, 3:6, 2003. [Eve07] R. Everaers, S. Kumar, and C. Simm. Unified description of polyand oligonucleotide DNA melting: Nearest-neighbor, poland-sheraga, and lattice models. Physical Review E, 75:041918, 2007. [Fin72] T. R. Fink and D. M. Crothers. Free-energy of imperfect nucleic-acid helices 1. bulge defect. Journal of Molecular Biology, 66(1):1–&, 1972. [Fod91] S. P. A. Fodor, J. L. Read, M. C. Pirrung, A. T. Stryer, L.and Lu, and D. Solas. Light-directed, spatially addressable parallel chemical synthesis. Science, 251(4995):767–773, 1991. [Fre86] S. M. Freier, R. Kierzek, J. A. Jaeger, N. Sugimoto, M. H. Caruthers, T. Neilson, and D. H. Turner. Improved free-energy parameters for predictions of RNA duplex stability. Proceedings of the National Academy of Sciences of the United States of America, 83(24):9373–9377, December 1986. [Gao01] X. L. Gao, E. LeProust, H. Zhang, O. Srivannavit, E. Gulari, P. L. Yu, C. Nishiguchi, Q. Xiang, and X. C. Zhou. A flexible light-directed DNA chip synthesis gated by deprotection using solution photogenerated acids. Nucleic Acids Research, 29(22):4744–4750, 2001. [Gao04] X. L. Gao, E. Gulari, and X. C. Zhou. In situ synthesis of oligonucleotide microarrays. Biopolymers, 73(5):579–596, April 2004. 248
BIBLIOGRAPHY [Gar02] P. B. Garland and P. J. Serafinowski. Effects of stray light on the fidelity of photodirected oligonucleotide array synthesis. Nucleic Acids Research, 30(19):e99, October 2002. [Gib59] J. H. Gibbs and E. A. Dimarzio. Statistical mechanics of helix-coil transitions in biological macromolecules. Journal of Chemical Physics, 30(1):271–282, 1959. [Gil77] D. T. Gillespie. Exact stochastic simulation of coupled chemical-reactions. Journal Of Physical Chemistry, 81(25):2340–2361, 1977. [Gla06] M. Glazer, J. A. Fidanza, G. H. McGall, M. O. Trulson, J. E. Forman, A. Suseno, and C. W. Frank. Kinetics of oligonucleotide hybridization to photolithographically patterned DNA arrays. Analytical Biochemistry, 358(2):225–238, November 2006. [Got81] O. Gotoh and Y. Tagashira. Stabilities of nearest-neighbor doublets in doublehelical DNA determined by fitting calculated melting profiles to observed profiles. Biopolymers, 20(5):1033–1042, 1981. [Gue87] M. Gueron, M. Kochoyan, and J. L. Leroy. A single-mode of DNA base-pair opening drives imino proton-exchange. Nature, 328(6125):89–92, July 1987. [Gut05] Z. Guttenberg, H. Muller, H. Habermuller, A. Geisbauer, J. Pipper, J. Felbel, M. Kielpinski, J. Scriba, and A. Wixforth. Planar chip device for pcr and hybridization with surface acoustic wave pump. Lab On A Chip, 5(3):308–317, 2005. [Hag88] P. J. Hagerman. Flexibility of DNA. Annual Review of Biophysics and Biophysical Chemistry, 17:265–286, 1988. [Hal04] A. Halperin, A. Buhot, and E. B. Zhulina. Sensitivity, specificity, and the hybridization isotherms of DNA chips. Biophysical Journal, 86(2):718–730, February 2004. [Hal05] A. Halperin, A. Buhot, and E. B. Zhulina. Brush effects on DNA chips: thermodynamics, kinetics, and design guidelines. Biophysical Journal, 89(2):796–811, August 2005. [Has97] A. Hasan, K. P. Stengele, H. Giegrich, P. Cornwell, K. R. Isham, R. A. Sachleben, W. Pfleiderer, and R. S. Foote. Photolabile protecting groups for nucleosides: Synthesis and photodeprotection rates. Tetrahedron, 53(12):4247–4264, 1997. [Hel03] G. A. Held, G. Grinstein, and Y. Tu. Modeling of DNA microarray data by using physical properties of hybridization. Proceedings of the National Academy of Sciences of the United States of America, 100(13):7575–7580, June 2003. [Hel06] G. A. Held, G. Grinstein, and Y. Tu. Relationship between gene expression and observed intensities in DNA microarraysa modeling study. Nucleic Acids Research, 34(9):e70, 2006. 249
BIBLIOGRAPHY [Sin84] N. D. Sinha, J. Biernat, J. Mcmanus, and H. Koster. Polymer support oligonucleotide synthesis .18. use of beta-cyanoethyl-n,n-dialkylamino-/n-morpholino phosphoramidite of deoxynucleosides for the synthesis of DNA fragments simplifying deprotection and isolation of the final product. Nucleic Acids Research, 12(11):4539–4557, 1984. [Ske93] J. V. Skelly, K. J. Edwards, T. C. Jenkins, and S. Neidle. Crystal-structure of an oligonucleotide duplex containing G.G base-pairs influence of mispairing on DNA backbone conformation. Proceedings of the National Academy of Sciences of the United States of America, 90(3):804–808, February 1993. [Sou75] E.M. Southern. Detection of specific sequences among DNA fragments separated by gel electrophoresis. J Mol Biol., 98:503–517, 1975. [Sug86] N. Sugimoto, R. Kierzek, S. M. Freier, and D. H. Turner. Energetics of internal GU mismatches in ribooligonucleotide helixes. Biochemistry, 25(19):5755– 5759, September 1986. [Sug95] N. Sugimoto, S. Nakano, M. Katoh, A. Matsumura, H. Nakamuta, T. Ohmichi, M. Yoneyama, and M. Sasaki. Thermodynamic parameters to predict stability of RNA/DNA hybrid duplexes. Biochemistry, 34(35):11211–11216, September 1995. [Sug00] N. Sugimoto, M. Nakano, and S. Nakano. Thermodynamics-structure relationship of single mismatches in RNA/DNA duplexes. Biochemistry, 39(37):11270– 11281, 2000. [Sun00] M. Sundaralingam and Y. Xiong. Crystal structure of domain II of x-laevis somatic 5s RNA in two conformations. Biophysical Journal, 78(1):311A–311A, January 2000. [Sun05] C. Sun, N. Fang, D. M. Wu, and X. Zhang. Projection micro-stereolithography using digital micro-mirror dynamic mask. Sensors and Actuators A-Physical, 121(1):113–120, 2005. [Tin73] I. Tinoco, P. N. Borer, B. Dengler, M. D. Levine, O. C. Uhlenbeck, D. M. Crothers, and J. Gralla. Improved estimation of secondary structure in ribonucleic-acids. Nature-New Biology, 246(150):40–41, 1973. [Toe03] A. Toegl, R. Kirchner, C. Gauer, and A. Wixforth. Enhancing results of microarray hybridization trough microagitation. Journal of Biomolecular Techniques, 14:197–204, 2003. [Tur92] D. H. Turner. Bulges in nucleic acids. Current Opinion in Structural Biology, 2:334–337, 1992. [Ura02] H. Urakawa, P. A. Noble, S. El Fantroussi, J. J. Kelly, and D. A. Stahl. Singlebase-pair discrimination of terminal mismatches by using oligonucleotide microarrays and neural network analyses. Applied and Environmental Microbiology, 68(1):235–244, January 2002. 256
BIBLIOGRAPHY [Ura03] H. Urakawa, S. El Fantroussi, H. Smidt, J. C. Smoot, E. H. Tribou, J. J. Kelly, P. A. Noble, and D. A. Stahl. Optimization of single-base-pair mismatch discrimination in oligonucleotide microarrays. Applied and environmental microbiology, 69(5):2848–2856, 2003. [Vai02] A. Vainrub and B. M. Pettitt. Coulomb blockage of hybridization in twodimensional DNA arrays. Physical Review E, 66(4):041905, October 2002. [vE06] T. S. van Erp, S. Cuesta-Lopez, and M. Peyrard. Bubbles and denaturation in DNA. European Physical Journal E, 20(4):421–434, August 2006. [Vic00] T. A. Vickers, J. R. Wyatt, and S. M. Freier. Effects of rna secondary structure on cellular antisense activity. Nucleic Acids Research, 28(6):1340–1347, March 2000. [Vij01] R. A. Vijayendran and D. E. Leckband. A quantitative assessment of heterogeneity for surface-immobilized proteins. Analytical Chemistry, 73(3):471–480, February 2001. [Wal79] R. B. Wallace, J. Shaffer, R. F. Murphy, J. Bonner, T. Hirose, and K. Itakura. Hybridization of synthetic oligodeoxyribonucleotidesto phi-chi-174 DNA effect of single base pair mismatch. Nucleic Acids Research, 6(11):3543–3557, 1979. [Wal01] S. Walbert, W. Pfleiderer, and U. E. Steiner. Photolabile protecting groups for nucleosides: Mechanistic studies of the 2-(2-nitrophenyl)ethyl group. Helvetica Chimica Acta, 84(6):1601–1611, 2001. [War85] R. M. Wartell and A. S. Benight. Thermal-denaturation ofDNA-molecules: a comparison of theory with experiment. Physics Reports-Review Section of Physics Letters, 126(2):67–107, 1985. [Wat00] J. H. Watterson, P. A. E. Piunno, C. C. Wust, and U. J. Krull. Effects of oligonucleotide immobilization density on selectivity of quantitative transduction of hybridization of immobilized DNA. Langmuir, 16(11):4984–4992, May 2000. [Wes07] E. M. Westerhout and B. Berkhout. A systematic analysis of the effect of target rna structure an rna interference. Nucleic Acids Research, 35(13):4322–4330, 2007. [Wet68] J. G. Wetmur and N. Davidson. Kinetics of renaturation of DNA. Journal of Molecular Biology, 31(3):349–&, 1968. [Wet91] J. G. Wetmur. DNA probes: Applications of the principles of nucleic-acid hybridization. Critical Reviews in Biochemistry and Molecular Biology, 26(34):227–259, 1991. [Wic06] L. M. Wick, J. M. Rouillard, T. S. Whittam, E. Gulari, J. M. Tiedje, and S. A. Hashsham. On-chip non-equilibrium dissociation curves and dissociation rate constants as methods to assess specificity of oligonucleotide probes. Nucleic Acids Research, 34(3):e26, 2006. 257
BIBLIOGRAPHY [Woe06] D. F. Woell. Neue photolabile Schutzgruppen mit intramolekularer Sensibilisierung - Synthese, photokinetische Charakterisierung und Anwendung f¨ ur die DNA-Chip-Synthese. PhD thesis, Universitaet Konstanz, 2006. [Wol04] D. Woll, S. Walbert, K. P. Stengele, T. J. Albert, T. Richmond, J. Norton, M. Singer, R. D. Green, W. Pfleiderer, and U. E. Steiner. Triplet-sensitized photodeprotection of oligonucleotides in solution and on microarray chips. Helvetica Chimica Acta, 87(1):28–45, 2004. [Won04] C. W. Wong, T. J. Albert, V. B. Vega, J. E. Norton, D. J. Cutler, T. A. Richmond, L. W. Stanton, E. T. Liu, and L. D. Miller. Tracking the evolution of the sars coronavirus using high-throughput, high-density resequencing arrays. Genome Research, 14(3):398–405, March 2004. [Woo88] S. A. Woodson and D. M. Crothers. Structural model for an oligonucleotide containing a bulged guanosine by NMR and energy minimization. Biochemistry, 27(9):3130–3141, May 1988. [Wu87] H. N. Wu and O. C. Uhlenbeck. Role of a bulged-a residue in a specific RNA protein-interaction. Biochemistry, 26(25):8221–8227, December 1987. [Yil04] L. S. Yilmaz and D. R. Noguera. Mechanistic approach to the problem of hybridization efficiency in fluorescent in situ hybridization. Applied And Environmental Microbiology, 70(12):7126–7139, December 2004. [Yoo01] J. S. Yoo, H. K. Cheong, B. J. Lee, Y. B. Kim, and C. Cheong. Solution structure of the SL1 RNA of the m1 double-stranded RNA virus of saccharomyces cerevisiae. Biophysical Journal, 80(4):1957–1966, April 2001. [Zen06] Y. Zeng and G. Zocchi. Mismatches and bubbles in DNA. Biophysical Journal, 90(12):4522–4529, 2006. [Zha03] L. Zhang, M. F. Miles, and K. D. Aldape. A model of molecular interactions on short oligonucleotide microarrays. Nature Biotechnology, 21(7):818–821, July 2003. [Zha07] L.Zhang, C. L.Wu, R. Carta, and H.T.Zhao. Free energy of DNA duplexformation on short oligonucleotide microarrays. Nucleic Acids Research, 35(3):e18, February 2007. [Zho06] H. Zhou, Y. Zhang, and Z. Ou-Yang. Handbook of Theoretical and Computational Nanotechnology, chapter Chapter 9: Theoretical and Computational Treatments of DNA and RNA Molecules, pages 419–487. American Scientific Publishers, 2006. [Zhu99] J. Zhu and R. M. Wartell. The effect of base sequence on the stability of RNA and DNA single base bulges. Biochemistry, 38(48):15986–15993, 1999. [Zim60] B. H. Zimm. Theory of melting of the helical form in double chains of the DNA type. Journal of Chemical Physics, 33(5):1349–1356, 1960. 258
BIBLIOGRAPHY [Zno02] B. M. Znosko, S. B. Silvestri, H. Volkman, B. Boswell, and M. J. Serra. Thermodynamic parameters for an expanded nearest-neighbor model for the formation of RNA duplexes with single nucleotide bulges. Biochemistry, 41(33):10406– 10417, 2002. [Zoc03] G. Zocchi, A. Omerzu, T. Kuriabova, J. Rudnick, and G. Gruner. Duplexsingle strand denaturation transition in DNA oligomers. 2003. arXiv:condmat/0304567. 259
BIBLIOGRAPHY 260
Appendix A Experimental Data 261
Experimental Data A.1 Experimental Data A.1.1 Comparison Between MMs in RNA/DNA and DNA/DNA Duplexes 0 2 4 6 8 10 12 14 16 0 0.2 0.4 0.6 0.8 T A T T A C T G G A C C T G A C 0 2 4 6 8 10 12 14 16 0 0.1 0.2 0.3 0.4 T A T T A C T G G A C C T G A C 0 2 4 6 8 10 12 14 16 0 0.5 1 1.5 T T G A G C G A T A T T A C T G 0 2 4 6 8 10 12 14 16 0 0.5 1 T T G A G C G A T A T T A C T G 0 2 4 6 8 10 12 14 16 0 0.2 0.4 0.6 A G T C A C G G A C A C A T G A 0 2 4 6 8 10 12 14 16 0 0.05 0.1 0.15 0.2 0.25 A G T C A C G G A C A C A T G A 0 2 4 6 8 10 12 14 16 0 0.5 1 1.5 C G A A C A T A C C T C C T T A 0 2 4 6 8 10 12 14 16 0 0.1 0.2 0.3 C G A A C A T A C C T C C T T A D C AB Figure A.1: Direct comparison of DNA/DNA and RNA/DNA mismatch hybridization signals (see section 6.8). Parts A-D compare defect profiles of different sequence motifs (sequences shown at the bottom of the plots). Hybridizations of RNA targets (top image) and equivalent DNA targets (bottom image) were performed subsequently on the same microarrays. The defect positional influence is very similar for DNA/DNA and RNA/DNA hybridization. However, there are systematic differences between the binding affinities of the various MM types in DNA/DNA and RNA/DNA duplexes. The hybridization signal (in a.u.) is plotted versus defect position. Substitution bases A (red cross), C(green circle), G (blue star) and T (cyan triangle) either result in 3 MM duplexes and one PM duplex at every defect position; Hybridization signals of duplexes with single base deletions (yellow line); moving average MM hybridization signal (black line). 262
Experimental Data 0 2 4 6 8 10 12 14 16 -0.2 0 0.2 0.4 0.6 0.8 1 1.2 G A T A T T A C T G G A C C T G 0 2 4 6 8 10 12 14 16 0 0.2 0.4 0.6 0.8 G A T A T T A C T G G A C C T G 0 2 4 6 8 10 12 14 16 0 0.2 0.4 0.6 0.8 1 1.2 G A T A T T A C T G G A C C T G 0 2 4 6 8 10 12 14 16 0 0.1 0.2 0.3 0.4 0.5 G A T A T T A C T G G A C C T G 0 2 4 6 8 10 12 14 16 0 0.5 1 1.5 A G C G A T A T T A C T G G A C 0 2 4 6 8 10 12 14 16 0 0.2 0.4 0.6 0.8 1 A G C G A T A T T A C T G G A C 0 2 4 6 8 10 12 14 16 0 0.5 1 1.5 A G C G A T A T T A C T G G A C 0 2 4 6 8 10 12 14 16 0 0.2 0.4 0.6 0.8 1 1.2 A G C G A T A T T A C T G G A C 0 2 4 6 8 10 12 14 16 0 0.5 1 T T G A G C G A T A T T A C T G 0 2 4 6 8 10 12 14 16 0 0.2 0.4 0.6 0.8 1 T T G A G C G A T A T T A C T G 0 2 4 6 8 10 12 14 16 0 0.5 1 1.5 T T G A G C G A T A T T A C T G 0 2 4 6 8 10 12 14 16 0 0.5 1 T T G A G C G A T A T T A C T G Figure A.2: For details see Fig. A.1. 263
Experimental Data 0 5 10 15 20 0 0.5 1 1.5 2 2.5 G C G A T A T T A C T G G A C C T G A C 0 5 10 15 20 0 0.5 1 1.5 G C G A T A T T A C T G G A C C T G A C 0 5 10 15 20 0 0.5 1 1.5 2 2.5 G C G A T A T T A C T G G A C C T G A C 0 5 10 15 20 0 0.5 1 1.5 2 G C G A T A T T A C T G G A C C T G A C 0 5 10 15 20 0 0.5 1 1.5 2 2.5 3 T T G A G C G A T A T T A C T G G A C C 0 5 10 15 20 0 0.5 1 1.5 2 2.5 T T G A G C G A T A T T A C T G G A C C 0 5 10 15 20 0 0.5 1 1.5 2 2.5 T T G A G C G A T A T T A C T G G A C C 0 5 10 15 20 0 0.5 1 1.5 T T G A G C G A T A T T A C T G G A C C 0 2 4 6 8 10 12 14 16 0 0.2 0.4 0.6 0.8 1 T A T T A C T G G A C C T G A C 0 2 4 6 8 10 12 14 16 0 0.1 0.2 0.3 0.4 0.5 0.6 T A T T A C T G G A C C T G A C 0 2 4 6 8 10 12 14 16 0 0.2 0.4 0.6 0.8 T A T T A C T G G A C C T G A C 0 2 4 6 8 10 12 14 16 0 0.1 0.2 0.3 0.4 T A T T A C T G G A C C T G A C Figure A.3: For details see Fig. A.1. 264
Experimental Data 0 2 4 6 8 10 12 14 16 0 0.1 0.2 0.3 0.4 C A C G G A C A C A T G A T C C 0 2 4 6 8 10 12 14 16 0 0.05 0.1 0.15 C A C G G A C A C A T G A T C C 0 2 4 6 8 10 12 14 16 0 0.1 0.2 0.3 0.4 0.5 C A C G G A C A C A T G A T C C 0 2 4 6 8 10 12 14 16 0 0.05 0.1 0.15 0.2 C A C G G A C A C A T G A T C C 0 2 4 6 8 10 12 14 16 0 0.2 0.4 0.6 A G T C A C G G A C A C A T G A 0 2 4 6 8 10 12 14 16 0 0.05 0.1 0.15 0.2 0.25 A G T C A C G G A C A C A T G A 0 2 4 6 8 10 12 14 16 0 0.2 0.4 0.6 A G T C A C G G A C A C A T G A 0 2 4 6 8 10 12 14 16 0 0.05 0.1 0.15 0.2 0.25 A G T C A C G G A C A C A T G A 0 2 4 6 8 10 12 14 16 0 0.2 0.4 0.6 T G T A G T C A C G G A C A C A 0 2 4 6 8 10 12 14 16 0 0.1 0.2 0.3 T G T A G T C A C G G A C A C A 0 2 4 6 8 10 12 14 16 0 0.1 0.2 0.3 0.4 0.5 0.6 T G T A G T C A C G G A C A C A 0 2 4 6 8 10 12 14 16 0 0.05 0.1 0.15 0.2 0.25 T G T A G T C A C G G A C A C A Figure A.4: For details see Fig. A.1. 265
Experimental Data A.1.3 Single Base Mismatches in DNA/DNA Duplexes - Statistical Analysis to Investigate the Influence of the Flanking Base Pairs −1.5 −1 −0.5 0 0.5 1 1.5 0 5 10 15 20 MM base pair: AA µ= −0.22 −1.5 −1 −0.5 0 0.5 1 1.5 0 5 10 15 20 MM base pair: CA µ= −0.11 −1.5 −1 −0.5 0 0.5 1 1.5 0 5 10 15 20 MM base pair: GA µ= 0.27 −1.5 −1 −0.5 0 0.5 1 1.5 0 5 10 15 20 MM base pair: AC µ= −0.29 −1.5 −1 −0.5 0 0.5 1 1.5 0 5 10 15 20 MM base pair: CC µ= −0.24 −1.5 −1 −0.5 0 0.5 1 1.5 0 5 10 15 20 MM base pair: TC µ= −0.26 −1.5 −1 −0.5 0 0.5 1 1.5 0 5 10 15 20 MM base pair: AG µ= −0.081 −1.5 −1 −0.5 0 0.5 1 1.5 0 5 10 15 20 MM base pair: GG µ= −0.33 −1.5 −1 −0.5 0 0.5 1 1.5 0 5 10 15 20 MM base pair: TG µ= −0.23 −1.5 −1 −0.5 0 0.5 1 1.5 0 5 10 15 20 MM base pair: CT µ= −0.13 −1.5 −1 −0.5 0 0.5 1 1.5 0 5 10 15 20 MM base pair: GT µ= 0.059 −1.5 −1 −0.5 0 0.5 1 1.5 0 5 10 15 20 MM base pair: TT µ= 0.069 Figure A.11: All mismatch base pair types X·Y. Measured hybridization signal distributions (occurrence versus deviation of the particular hybridization signal from the mean profile) as a function of the MM base pair alone, i.e. independent of the flanking base pairs. µdenotes the median value of the distributions. A box-whisker plot of the distributions is shown in Fig. 6.6. On the following pages (Figs. A.12 - A.22) this data is categorized according to the type of flanking base pairs. Owing to the restricted set of target sequences available for this study the sizes of the data sets measured for the individual defect configurations are very different. µdenotes the median values of the distributions. The median values of the nearest neighbor pair dependent subsets are compared in Fig. 6.9. 272
Experimental Data −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAT−3´ 3´−AAA−5´ µ= −0.39 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAG−3´ 3´−AAC−5´ µ= −0.39 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAC−3´ 3´−AAG−5´ µ= 0.0021 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAT−3´ 3´−CAA−5´ µ= −0.0072 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAG−3´ 3´−CAC−5´ µ= −0.24 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAA−3´ 3´−CAT−5´ µ= −0.1 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAT−3´ 3´−GAA−5´ µ= −0.57 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAC−3´ 3´−GAG−5´ µ= −0.087 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAA−3´ 3´−GAT−5´ µ= −0.27 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAG−3´ 3´−TAC−5´ µ= −0.15 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAC−3´ 3´−TAG−5´ µ= −0.22 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAA−3´ 3´−TAT−5´ µ= −0.12 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAA−3´ 3´−AAT−5´ µ= −0.055 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAC−3´ 3´−CAG−5´ µ= −0.33 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAG−3´ 3´−GAC−5´ µ= −0.26 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAT−3´ 3´−TAA−5´ µ= −0.036 Figure A.12: A·A mismatches. Measured hybridization signal distributions categorized according to the flanking base pairs. −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAT−3´ 3´−ACA−5´ µ= −0.24 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAG−3´ 3´−ACC−5´ µ= 0.4 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAC−3´ 3´−ACG−5´ µ= 0.21 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAT−3´ 3´−CCA−5´ µ= 0.064 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAG−3´ 3´−CCC−5´ µ= 0.18 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAA−3´ 3´−CCT−5´ µ= −0.074 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAT−3´ 3´−GCA−5´ µ= −0.47 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAC−3´ 3´−GCG−5´ µ= −0.24 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAA−3´ 3´−GCT−5´ µ= −0.29 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAG−3´ 3´−TCC−5´ µ= −0.09 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAC−3´ 3´−TCG−5´ µ= −0.22 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAA−3´ 3´−TCT−5´ µ= −0.15 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAA−3´ 3´−ACT−5´ µ= −0.29 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAC−3´ 3´−CCG−5´ µ= −0.35 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAG−3´ 3´−GCC−5´ µ= −0.035 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAT−3´ 3´−TCA−5´ µ= −0.26 Figure A.13: C·A mismatches. Measured hybridization signal distributions categorized according to the flanking base pairs. 273
Experimental Data −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAT−3´ 3´−AGA−5´ µ= 0.68 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAG−3´ 3´−AGC−5´ µ= 0.33 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAC−3´ 3´−AGG−5´ µ= 0.017 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAT−3´ 3´−CGA−5´ µ= 0.51 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAG−3´ 3´−CGC−5´ µ= 0.27 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAA−3´ 3´−CGT−5´ µ= −0.28 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAT−3´ 3´−GGA−5´ µ= −0.33 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAC−3´ 3´−GGG−5´ µ= 0.28 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAA−3´ 3´−GGT−5´ µ= 0.18 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAG−3´ 3´−TGC−5´ µ= 0.56 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAC−3´ 3´−TGG−5´ µ= 0.16 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAA−3´ 3´−TGT−5´ µ= −0.00077 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAA−3´ 3´−AGT−5´ µ= −0.17 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAC−3´ 3´−CGG−5´ µ= 0.68 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAG−3´ 3´−GGC−5´ µ= 0.4 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAT−3´ 3´−TGA−5´ µ= 0.35 Figure A.14: G·A mismatches. Measured hybridization signal distributions categorized according to the flanking base pairs. −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TCT−3´ 3´−AAA−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TCG−3´ 3´−AAC−5´ µ= −0.26 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TCC−3´ 3´−AAG−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GCT−3´ 3´−CAA−5´ µ= −0.47 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GCG−3´ 3´−CAC−5´ µ= 0.087 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GCA−3´ 3´−CAT−5´ µ= −0.32 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CCT−3´ 3´−GAA−5´ µ= −0.2 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CCC−3´ 3´−GAG−5´ µ= −0.25 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CCA−3´ 3´−GAT−5´ µ= −0.18 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ACG−3´ 3´−TAC−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ACC−3´ 3´−TAG−5´ µ= −0.16 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ACA−3´ 3´−TAT−5´ µ= −0.48 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TCA−3´ 3´−AAT−5´ µ= −0.28 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GCC−3´ 3´−CAG−5´ µ= −0.11 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CCG−3´ 3´−GAC−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ACT−3´ 3´−TAA−5´ µ= −0.58 Figure A.15: A·C mismatches. Measured hybridization signal distributions categorized according to the flanking base pairs. 274
Experimental Data −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TCT−3´ 3´−ATA−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TCG−3´ 3´−ATC−5´ µ= −0.087 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TCC−3´ 3´−ATG−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GCT−3´ 3´−CTA−5´ µ= −0.32 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GCG−3´ 3´−CTC−5´ µ= −0.63 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GCA−3´ 3´−CTT−5´ µ= −0.31 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CCT−3´ 3´−GTA−5´ µ= −0.094 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CCC−3´ 3´−GTG−5´ µ= −0.042 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CCA−3´ 3´−GTT−5´ µ= 0.094 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ACG−3´ 3´−TTC−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ACC−3´ 3´−TTG−5´ µ= −0.086 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ACA−3´ 3´−TTT−5´ µ= −0.22 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TCA−3´ 3´−ATT−5´ µ= −0.1 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GCC−3´ 3´−CTG−5´ µ= −0.15 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CCG−3´ 3´−GTC−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ACT−3´ 3´−TTA−5´ µ= −0.37 Figure A.16: T·C mismatches. Measured hybridization signal distributions categorized according to the flanking base pairs. −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TGT−3´ 3´−AAA−5´ µ= −0.1 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TGG−3´ 3´−AAC−5´ µ= −0.11 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TGC−3´ 3´−AAG−5´ µ= −0.26 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GGT−3´ 3´−CAA−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GGG−3´ 3´−CAC−5´ µ= 0.74 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GGA−3´ 3´−CAT−5´ µ= 0.1 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CGT−3´ 3´−GAA−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CGC−3´ 3´−GAG−5´ µ= −0.028 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CGA−3´ 3´−GAT−5´ µ= 1.3 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AGG−3´ 3´−TAC−5´ µ= −0.11 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AGC−3´ 3´−TAG−5´ µ= 0.046 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AGA−3´ 3´−TAT−5´ µ= −0.23 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TGA−3´ 3´−AAT−5´ µ= −0.11 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GGC−3´ 3´−CAG−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CGG−3´ 3´−GAC−5´ µ= −0.55 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AGT−3´ 3´−TAA−5´ µ= −0.25 Figure A.17: A·G mismatches. Measured hybridization signal distributions categorized according to the flanking base pairs. 275
Experimental Data −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TGT−3´ 3´−AGA−5´ µ= −0.36 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TGG−3´ 3´−AGC−5´ µ= −0.22 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TGC−3´ 3´−AGG−5´ µ= −0.24 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GGT−3´ 3´−CGA−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GGG−3´ 3´−CGC−5´ µ= −0.17 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GGA−3´ 3´−CGT−5´ µ= −0.15 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CGT−3´ 3´−GGA−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CGC−3´ 3´−GGG−5´ µ= −0.2 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CGA−3´ 3´−GGT−5´ µ= −0.33 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AGG−3´ 3´−TGC−5´ µ= −0.49 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AGC−3´ 3´−TGG−5´ µ= −0.37 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AGA−3´ 3´−TGT−5´ µ= −0.77 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TGA−3´ 3´−AGT−5´ µ= −0.33 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GGC−3´ 3´−CGG−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CGG−3´ 3´−GGC−5´ µ= −0.26 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AGT−3´ 3´−TGA−5´ µ= −0.51 Figure A.18: G·G mismatches. Measured hybridization signal distributions categorized according to the flanking base pairs. −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TGT−3´ 3´−ATA−5´ µ= −0.35 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TGG−3´ 3´−ATC−5´ µ= −0.41 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TGC−3´ 3´−ATG−5´ µ= −0.22 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GGT−3´ 3´−CTA−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GGG−3´ 3´−CTC−5´ µ= 0.056 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GGA−3´ 3´−CTT−5´ µ= −0.1 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CGT−3´ 3´−GTA−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CGC−3´ 3´−GTG−5´ µ= −0.11 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CGA−3´ 3´−GTT−5´ µ= −0.15 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AGG−3´ 3´−TTC−5´ µ= −0.31 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AGC−3´ 3´−TTG−5´ µ= −0.17 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AGA−3´ 3´−TTT−5´ µ= −0.23 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TGA−3´ 3´−ATT−5´ µ= −0.27 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GGC−3´ 3´−CTG−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CGG−3´ 3´−GTC−5´ µ= −0.52 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AGT−3´ 3´−TTA−5´ µ= −0.45 Figure A.19: T·G mismatches. Measured hybridization signal distributions categorized according to the flanking base pairs. 276
Experimental Data −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TTT−3´ 3´−ACA−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TTG−3´ 3´−ACC−5´ µ= 0.053 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TTC−3´ 3´−ACG−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GTT−3´ 3´−CCA−5´ µ= 0.12 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GTG−3´ 3´−CCC−5´ µ= −0.052 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GTA−3´ 3´−CCT−5´ µ= 0.099 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CTT−3´ 3´−GCA−5´ µ= −0.13 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CTC−3´ 3´−GCG−5´ µ= −0.4 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CTA−3´ 3´−GCT−5´ µ= −0.048 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ATG−3´ 3´−TCC−5´ µ= −0.34 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ATC−3´ 3´−TCG−5´ µ= −0.42 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ATA−3´ 3´−TCT−5´ µ= −0.34 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TTA−3´ 3´−ACT−5´ µ= 0.38 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GTC−3´ 3´−CCG−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CTG−3´ 3´−GCC−5´ µ= −0.071 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ATT−3´ 3´−TCA−5´ µ= 0.091 Figure A.20: C·T mismatches. Measured hybridization signal distributions categorized according to the flanking base pairs. −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TTT−3´ 3´−AGA−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TTG−3´ 3´−AGC−5´ µ= 0.24 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TTC−3´ 3´−AGG−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GTT−3´ 3´−CGA−5´ µ= 0.2 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GTG−3´ 3´−CGC−5´ µ= −0.018 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GTA−3´ 3´−CGT−5´ µ= 0.17 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CTT−3´ 3´−GGA−5´ µ= −0.1 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CTC−3´ 3´−GGG−5´ µ= 0.00061 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CTA−3´ 3´−GGT−5´ µ= 0.48 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ATG−3´ 3´−TGC−5´ µ= −0.11 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ATC−3´ 3´−TGG−5´ µ= −0.15 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ATA−3´ 3´−TGT−5´ µ= 0.075 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TTA−3´ 3´−AGT−5´ µ= 0.28 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GTC−3´ 3´−CGG−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CTG−3´ 3´−GGC−5´ µ= 0.14 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ATT−3´ 3´−TGA−5´ µ= 0.27 Figure A.21: G·T mismatches. Measured hybridization signal distributions categorized according to the flanking base pairs. 277
Experimental Data −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TTT−3´ 3´−ATA−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TTG−3´ 3´−ATC−5´ µ= −0.12 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TTC−3´ 3´−ATG−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GTT−3´ 3´−CTA−5´ µ= −0.083 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GTG−3´ 3´−CTC−5´ µ= −0.036 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GTA−3´ 3´−CTT−5´ µ= 0.25 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CTT−3´ 3´−GTA−5´ µ= 0.99 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CTC−3´ 3´−GTG−5´ µ= 0.48 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CTA−3´ 3´−GTT−5´ µ= 0.69 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ATG−3´ 3´−TTC−5´ µ= −0.082 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ATC−3´ 3´−TTG−5´ µ= −0.2 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ATA−3´ 3´−TTT−5´ µ= −0.0012 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TTA−3´ 3´−ATT−5´ µ= −0.017 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GTC−3´ 3´−CTG−5´ no data available −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CTG−3´ 3´−GTC−5´ µ= 0.34 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−ATT−3´ 3´−TTA−5´ µ= 0.52 Figure A.22: T·T mismatches. Measured hybridization signal distributions categorized according to the flanking base pairs. −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAT−3´ 3´−AGA−5´ µ= 0.68 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAG−3´ 3´−AGC−5´ µ= 0.33 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAC−3´ 3´−AGG−5´ µ= 0.017 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAT−3´ 3´−CGA−5´ µ= 0.51 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAG−3´ 3´−CGC−5´ µ= 0.27 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAA−3´ 3´−CGT−5´ µ= −0.28 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAT−3´ 3´−GGA−5´ µ= −0.33 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAC−3´ 3´−GGG−5´ µ= 0.28 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAA−3´ 3´−GGT−5´ µ= 0.18 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAG−3´ 3´−TGC−5´ µ= 0.56 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAC−3´ 3´−TGG−5´ µ= 0.16 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAA−3´ 3´−TGT−5´ µ= −0.00077 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−TAA−3´ 3´−AGT−5´ µ= −0.17 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−GAC−3´ 3´−CGG−5´ µ= 0.68 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−CAG−3´ 3´−GGC−5´ µ= 0.4 −1.5 −1 −0.5 0 0.5 1 1.5 0 2 4 6 5´−AAT−3´ 3´−TGA−5´ µ= 0.35 Figure A.23: G·A mismatches. Measured hybridization signal distributions categorized according to the flanking base pairs. 278
Experimental Data A.1.4 Single Base Insertions - Statistical Analysis −0.4 −0.2 0 0.2 0 5 10 5´−T T−3´ 3´−AAA−5´ µ= 0.023 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−T G−3´ 3´−AAC−5´ µ= 0.048 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−T C−3´ 3´−AAG−5´ µ= 0.0049 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−T A−3´ 3´−AAT−5´ µ= −0.009 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−G T−3´ 3´−CAA−5´ µ= 0.031 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−G G−3´ 3´−CAC−5´ µ= 0 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−G C−3´ 3´−CAG−5´ µ= −0.12 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−G A−3´ 3´−CAT−5´ µ= −0.064 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−C T−3´ 3´−GAA−5´ µ= 0.0063 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−C G−3´ 3´−GAC−5´ µ= 0.015 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−C C−3´ 3´−GAG−5´ µ= −0.012 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−C A−3´ 3´−GAT−5´ µ= −0.098 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−A T−3´ 3´−TAA−5´ µ= 0.017 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−A G−3´ 3´−TAC−5´ µ= 0 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−A C−3´ 3´−TAG−5´ µ= −0.041 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−A A−3´ 3´−TAT−5´ µ= −0.075 Group: I Figure A.24: Insertions of adenine bases - influence of the neighboring base pairs. Distribution of hybridization signal intensities (deviation from the mean profile in a.u.). µdenotes the median value of the distribution. Group II insertions have at least one identical neighbor base, whereas Group I insertions don’t have an identical neighbor. Group II insertion have consistently increased hybridization signals compared to Group I insertions. 279
Experimental Data −0.4 −0.2 0 0.2 0 5 10 5´−T T−3´ 3´−ACA−5´ µ= 0 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−T G−3´ 3´−ACC−5´ µ= 0.045 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−T C−3´ 3´−ACG−5´ µ= −0.043 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−T A−3´ 3´−ACT−5´ µ= 0.0075 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−G T−3´ 3´−CCA−5´ µ= 0.012 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−G G−3´ 3´−CCC−5´ µ= 0.12 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−G C−3´ 3´−CCG−5´ µ= 0.012 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−G A−3´ 3´−CCT−5´ µ= 0.032 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−C T−3´ 3´−GCA−5´ µ= −0.054 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−C G−3´ 3´−GCC−5´ µ= −0.0087 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−C C−3´ 3´−GCG−5´ µ= −0.039 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−C A−3´ 3´−GCT−5´ µ= −0.057 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−A T−3´ 3´−TCA−5´ µ= −0.069 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−A G−3´ 3´−TCC−5´ µ= 0.025 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−A C−3´ 3´−TCG−5´ µ= −0.051 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−A A−3´ 3´−TCT−5´ µ= −0.066 Group: I Figure A.25: Insertions of cytosine bases - influence of the neighboring base pairs. Distribution of hybridization signal intensities (deviation from the mean profile in arbitrary units). −0.4 −0.2 0 0.2 0 5 10 5´−T T−3´ 3´−AGA−5´ µ= 0 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−T G−3´ 3´−AGC−5´ µ= −0.088 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−T C−3´ 3´−AGG−5´ µ= 0.0059 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−T A−3´ 3´−AGT−5´ µ= 0.014 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−G T−3´ 3´−CGA−5´ µ= −0.013 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−G G−3´ 3´−CGC−5´ µ= −0.093 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−G C−3´ 3´−CGG−5´ µ= 0.048 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−G A−3´ 3´−CGT−5´ µ= −0.021 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−C T−3´ 3´−GGA−5´ µ= 0.085 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−C G−3´ 3´−GGC−5´ µ= 0.034 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−C C−3´ 3´−GGG−5´ µ= 0.12 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−C A−3´ 3´−GGT−5´ µ= 0.09 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−A T−3´ 3´−TGA−5´ µ= −0.033 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−A G−3´ 3´−TGC−5´ µ= −0.063 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−A C−3´ 3´−TGG−5´ µ= 0.14 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−A A−3´ 3´−TGT−5´ µ= 0 Group: I Figure A.26: Insertions of guanine bases - influence of the neighboring base pairs. Distribution of hybridization signal intensities (deviation from the mean profile in arbitrary units). 280
Experimental Data −0.4 −0.2 0 0.2 0 5 10 5´−T T−3´ 3´−ATA−5´ µ= 0.032 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−T G−3´ 3´−ATC−5´ µ= −0.081 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−T C−3´ 3´−ATG−5´ µ= −0.035 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−T A−3´ 3´−ATT−5´ µ= 0.0054 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−G T−3´ 3´−CTA−5´ µ= 0.0062 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−G G−3´ 3´−CTC−5´ µ= −0.08 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−G C−3´ 3´−CTG−5´ µ= −0.11 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−G A−3´ 3´−CTT−5´ µ= −0.0068 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−C T−3´ 3´−GTA−5´ µ= 0.028 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−C G−3´ 3´−GTC−5´ µ= −0.029 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−C C−3´ 3´−GTG−5´ µ= 0 Group: I −0.4 −0.2 0 0.2 0 5 10 5´−C A−3´ 3´−GTT−5´ µ= 0 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−A T−3´ 3´−TTA−5´ µ= 0.01 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−A G−3´ 3´−TTC−5´ µ= −0.027 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−A C−3´ 3´−TTG−5´ µ= −0.0097 Group: II −0.4 −0.2 0 0.2 0 5 10 5´−A A−3´ 3´−TTT−5´ µ= 0.043 Group: II Figure A.27: Insertions of thymine bases - influence of the neighboring base pairs. Distribution of hybridization signal intensities (deviation from the mean profile in arbitrary units). 281
Supporting Information 0 0.2 0.4 0.6 0.8 1 0 500 1000 1500 Iin Iout (a.u.) Measured intensity Fit: Iout=Iin2.2 Figure B.5: Gamma function of the AstroBeam projector. The intensity response Iout on the image brightness Iin (normalized on a maximum value of 1) follows a power law with an exponent of 2.2. For an image brightness larger than about 80% of the maximum value a cut-off is observed. The position of the cut-off depends on the contrast and brightness values chosen in the AstroBeams ”Display Settings Menu”. B.3 Optics of the Microscope Projection Photolithography System •UHP: Philips UHP-lamp 250W 1.35 TOP 222 H4 elliptical reflector elliptical reflector geometry: major axis ∼80 mm, minor axis ∼50 mm) •L1: plano-concave lens: f=50 mm, diam. 25 mm (silica), placement between UHP lamp window and the outer focal point of the elliptical reflector •L1-L2: 145 mm •L2: plano-convexlens: f=50 mm, diam. 50 mm •L2-F1: 120 mm •F1: UV cold mirror (UV barrier filter from the Optoma projector lamp module) •F1-L3: 165 mm •L3: plano-convexlens (BK7): f=100 mm, diam. 50 mm •F1-F2: 215 mm •F2: UV cold mirror (Oriel) •F2-F3: 165 mm •F3: UV band pass (bk-370-35-B, Interferenzoptik Elektronik GmbH), diam. 25.4 mm •F2-L4: 250 mm •L4: plano-convexlens (BK7): f=125 mm, diam. 50 mm •L4-M1: 170 mm 288
Optics of the Microscope Projection Photolithography System UHP L1 L2 F1 LT L3 F2 LT S F3 F4 L4 M1 M2 L5 DMD M3 FO PS VP1 VP2 PC ICM Figure B.6: Schematic of the microscope projection photolithography system. •M1: mirror •M1-M2: 380 mm •M2: mirror •M2-DMD: 60 mm •DMD-L5: ca. 164.5 mm, to be fine-adjusted •L5: tube lens, Carl Zeiss, f=164.5 mm •M3: mirror/beam splitter 289
Supporting Information B.4 Fabrication of the Synthesis Cell Figure B.7: Punching tool (top) for the fabrication of the PDMS gasket (center). The tool, producing a diamond-shaped cutout (the cell volume) with clean edges, is essential for smooth operation of synthesis apparatus. Wire-cut EDM (electrical discharge machining) has been employed for producing the sharp-edged structure in hardened steel. Dimensions of diamond-shaped cell volume: length 16 mm; width 5 mm. The outer edge of the gasket was cut with another (smaller) version of the punching tool. Part names are referring to Fig. 3.13. •The top-plate is made from a 10 mm thick plate of transparent Makrolon R plastics (polycarbonate). Produce four tapped holes for fastening screws (not too far away from the center of the plate, to enable proper sealing action). Further, two holes for fastening the cell-assembly on the projection lithography setup are required. •Inlet and outlet tubes are made from syringe needles (0.9×40 mm). By using a drilling machine as a ”lathe” the plastic adapter of the syringe needle is reduced to a cylindric bit as shown in Fig. 3.13. •Produce holes for inlet/outlet needles. (diam. 1 mm on the upper side of the top plate). At the bottom side of the top-plate the needle (blunt end near the coupling) should protrude 1 mm. The needles are fastened with epoxy glue. •To obtain a transparent and chemically inert (solvent resistant) surface, a glass microscopy slide is glued onto the lower side of the top-plate. Before gluing (with transparent PDMS silicone rubber), the slide needs to be cut in 3 pieces to produce gaps for the fastening screws. Moreover, two 1 mm diam. holes for the inlet/outlet tubes have to be drilled into the glass slide by using a diamond tool. By gluing the glass slide onto 290
Fabrication of the Synthesis Cell the top-plate the gaps between the needles and the glass are sealed with PDMS (avoid getting PDMS into the needles!). PDMS (Dow Corning Sylgard R 184) was purchased from World Precision Instruments. •The bottom-platte is made from 5 mm aluminum. The exposure window should not be too large (ideally implemented as a long hole) to achieve proper sealing action by pressing the Chip-substrate/PDMS-gasket against the top-plate. •Fabrication of the PDMS-gasket: PDMS Sylgard 184 (Dow Corning) is mixed thoroughly (ratio between elastomer base and curing agent: 10:1), degassed and poured into a glass petri dish. Curing for 20 minutes at 80◦C. A custom-made punching tool (Fig. B.7) is used to produce the streamlined cutout forming the synthesis volume. •Connectors: PFA (PTFE) tubes (internal diam. 0.8 mm) fit tightly on the 0.9 mm diam. syringe needles. PTFE tube end fittings (UNF 1/4” 28 G) provide a removable connection with the fluidics system. 291
Supporting Information B.5 Technical Notes on Light-directed DNA Chip Synthesis B.5.1 Handling of Phosphoramidite Reagents The coupling efficiency of phosphoramidite reagents is very sensitive to contamination with (even trace amounts of) water. To maintain low moisture conditions the following precautions should be considered: •Storage under moisture free conditions at -20◦C. Use dry argon atmosphere and desiccant. •Open storage bottles only in glove box under dry argon atmosphere. Use silica gel beads to maintain a low moisture content in the glove box. •Use oven-dried glass ware to minimize surface-adsorbed water. •Dissolve phosphoramidites only immediately before synthesis. •Use dry MeCN with <10 ppm of water. •Use molecular sieve bags (in the MeCN storage bottle and in the activator solution) to adsorb water from the solvent. •Phosphoramidite solutions should be used the same day as prepared. Solution stability and degradation pathways of deoxyribonucleoside phosphoramidites in MeCN are discussed in [Kro04]. B.5.2 Additional Notes on the Synthesis Prior to the first phosphoramidite coupling the substrate is soaked in MeCN for about 2 minutes. The initial coupling is performed for 1 minute and then repeated once. According to Richmond et al. [Ric04] an increase of the coupling time (of the first base only) from 20 s to 6 h resulted in an 80% increase in the amount of full-length probes. Coupling and exposure time, washing steps and image quality are the key parameters for high quality synthesis. According to [Ric04] the number of error-free probe sequences could be increased 100-fold by making several technical improvements on their synthesis apparatus. Improvements include the extension of the coupling time from 20 to 60 s and of the exposure time from 50 to 150 s, additional argon drying steps and modifications on the projection optical system (image-locking). Upon prolonged exposure the solvents tetrahydrofurane (THF) and pyridine cause significant swelling of the PDMS gasket. Exposure to these solvents (contained in oxidizer and capping reagents) should therefore be minimized. 292
Technical Notes on Microarray Dendrimer Substrate Preparation B.6 Technical Notes on Microarray Dendrimer Substrate Preparation Figure B.8: (A) Teflon slide holder for up to 12 round cover glasses. The stainless steel pin secures the glasses. For use with dichloroethane the nylon screws should be replaced by stainless steel screws. (B) Substrate functionalization in a 500 ml graduated cylinder requires about 250 ml reagent solution. •For dendrimer functionalization of the microarray substrates a compact slide holder for handling of up to 12 cover glasses was developed. Parts of the teflon (PTFE) slide holder are assembled with stainless steel screws and can thus withstand a bath in dichloroethane solution. The holder enables fast and thorough washing and drying of the slides. Use of the holders resulted in significantly increased quality of the substrates and enabled reduction of the reagent consumption. •To minimize reagent consumption (ethanol analytical grade, dendrimers in dichloroethane) the substrate functionalization is performed in a 500 ml graduated cylinder. Three slide holders (with 36 slides in total) are immersed in about 250 ml of solution. •Drying of the slides under a nitrogen stream should be performed in such a way that the liquid is blown away from the center of the slides. Drying of droplets on the surface has to be avoided because this can produce irremovable stains. 293
Supporting Information B.7 Technical Notes on the Synthesizer Control Software DNASyn The light-directed fabrication of a DNA microarray has been fully automated. The synthesizer control software DNASyn integrates control of the fluidics system with the maskless microphotolithography system (including image display, shutter and filter control). Figure B.9: The graphical user interface of DNASyn. The buttons in the left panel enable manual access to user-defined macro functions. The textbox at the right shows the code of the synthesis script loaded. DNASyn was implemented in JavaTM. It is running with Windows XP Professional (and is also expected to work with Window98). The use with Windows XP Home or Windows Vista is not recommended since these operating systems won’t allow direct access to the hardware ports via the kernel mode driver UserPort. 294
Technical Notes on the Synthesizer Control Software DNASyn B.7.1 Basic Features Manual operation (via GUI) automatic mode Synthesis script (synthesis procedure for the particular DNAchip) Mask display (1024x768 XGA) Standard macros (fluidics parameters etc.) Filter and Shutter control (via parallel port) Mask files (jpeg images) Graphical User Inferface Synthesis script interpreter Fluids control (via serial port) DNASyn Figure B.10: Concept of the DNASyn microarray synthesis control software. DNASyn includes a flexible macro programming language for the automated control of the synthesis process, and a graphical user interface (GUI) for manual control of various synthesizer functions (see Fig. B.9). The macro language comprises only a small number of basic commands. Keywords START Begin of the main program END End of the main program MACRO macroname {...}Macro header PRINT n note DNASyn shows text note in output line n // comment Comment in the source code WAIT n Wait for n seconds VX Y Valve operation X: valve number ; Y: 0=close 1=open DISPLAY imagename.jpg Virtual mask display DISPLAY AGAIN Display the previous image again SHUTTER ON/OFF Shutter control FILTER GREEN/UV Filter changer control •Switching of solenoid valves (fluidics operations) is performed with the V X Y command. •The DISPLAY imagename.jpg command loads the JPG image from the synthesis directory and shows it on the DMD. The keyword AGAIN is used to reload the previous image. •The WAIT ncommand (nduration in seconds) is used for time control of the synthesis processes. 295
Supporting Information •Comments begin with // followed by a space character. Macros Typical routines (e.g. amidite coupling or photo-deprotection) can be combined to macro commands, as shown in the following example. MACRO rinse 20 { //Rinse synthesis cell with MeCN for 20 s - this is a comment V 18 1 V 2 1 V13 1 V17 1 WAIT 20 V 2 0 V13 0 V17 0 V18 0 } Macro commands can be called from the main program and from within other macros. Manual control (via button-click in the control panel) is also based on macro commands. Most control panel buttons are assigned a macro function. Macro codes for these functions are listed (and can be modified if necessary) in the file functions.prg. A synthesis program comprises a list ofmacros (a libraryof standardmacrosandadditional user-defined macros) and the main program. Standard macros describe routine synthesis processes. Basically they are not different from user-defined macros, but since theyinclude critical time parameters (duration of fluidics processes, exposure times etc.) and since they may be called from other macros, modifications in standard macros should be considered cautiously. Upon loading a synthesis program (file extension .prg) the parser of DNASyn initially reads the main program (between the commands START and END). In the next step macro calls are substituted by the corresponding macro codes. To consider nested macros this is repeated until all macros are resolved. A completely resolved synthesis program for a 25mer array synthesis typically comprises about 40000 commands. Frequently used macro functions flush flush synthesis cell with argon flow X flow reagent X through the synthesis cell rinse X fill MeCN into the storage bottle for reagent X 296
Technical Notes on the Synthesizer Control Software DNASyn rinse block rinse valve block with MeCN flush block flush valve block with argon prime X fill the tube between the storage bottle X and the valve block with reagent X reverse flush fast flush of the synthesis cell with argon in reverse direction deprotect photodeprotection couple X coupling of the phoshoramidite X oxidize oxidization of phosphite bonds Number-extensions to the functions name (e.g. flush10) specify the duration of the operation (in seconds). B.7.2 Communications between the Control PC and the Synthesizer Hardware For serial communication with the solenoid valve controller the Java Communications API (Sun Microsystems) is employed. The communications parameters have been set to the requirements of the valve controller (see below). The control of the shutter and filter-changer via the parallel port has been implemented with a Java nativecode. Direct control of the parallel port requires the java packageparport. The library parport.dll needs to be installed in the directory Systems32/drivers. With parport the channels of the parallel port can be set and read in a straightforward way. For direct access on the I/O ports (user mode) the driver UserPort (written by Tomas Franzon) needs to be installed (for this purpose Userport.sys needs to be copied to System32/drivers). Possibly the Windows98 compatibility mode needs to be enabled. With the executable Userport.exethe access to the parallel port (base address $387) is set enabled. B.7.3 Dual Screen Support DNASyn provides dual screen support to display the control panel and the photolithography mask patterns on different devices - TFT monitor and video projector (DMD), respectively. This requires the use of a dualview graphics card and extension of the Windows desktop onto the second display. The control panel is displayed on the primary screen (TFT-monitor with 1280×1024 pixels). Display of the photolithography masks on the secondary display (video projector) is achieved by opening a window at the corresponding desktop coordinates - no further programming tricks are necessary. The Class DisplayFrame, an extension of the Java Class JWindow enables display of the masks without 297
Supporting Information Figure B.15: The readout grid is exactly positioned on the microarray features. Averaging over the readout boxes yields the hybridization signals of the individual microarray features. 304
Temperature Control of the Hybridization Chamber B.11 Temperature Control of the Hybridization Chamber DA0 DA1 DO0 DO1 DO2 DO3 DR0 DR1 DR2 DR3 RES CNT C7 C6 C5 C4 C3 C2 C1 C0 B7 B6 B5 B4 B3 B2 B1 B0 A7 A6 A5 A4 A3 A2 A1 A0 DI3 DI2 DI1 DI0 CH7 CH6 CH5 CH4 CH3 CH2 CH1 CH0 PMD1 PMD-1008 RUN STP E1 E2 E3 E4 YT1 Y( t ) ND1 1.23 E0 E1 A SUB1 - Difference between set temperature and actual temperature Manual temperature setting Actual temperature display E0 E1 A MUL1 * E0 E1 A MUL2 * E0 E1 A MUL3 * E RST A INT1 d(+ ) E RST A DIF1 d(-) E0 E1 E2 A ADD1 + PID-Controller KP KI KD RUN STP E1 E2 E3 E4 YT2 Y( t ) ND2 1.23 ND3 1.23 SR1 ND4 1.23 Control voltage E0 E1 A MUL4 * E0 E1 A MUL5 * FW1 W FW2 W Ain Bin A< B A= B A> B AVG1 Vergl. E0 E1 A MUL6 * Maximum temperature setting SR2 ND5 1.23 EN IN LiH LiLClpL ClpH A LIM1 Limiter FW3 W FW4 W 0...5V SR3 ND6 1.23 LE1 Ain Bin A< B A= B A> B AVG2 Vergl. E0 E1 A MUL7 * ND7 1.23 ND8 1.23 LED1 T1 G1 E A KT1 KT UP DN RST Z CO1 0 0 0 1 T2 ND9 1.23 Temperature program control S1 AND1 & E0 E1 SEL A REL1 S2 manual/automatic selector E0A FRM1 F FW5 W Offset in V PON1 R run program FW6 W FW7 W FW8 W FW9 W E1 E2 Add RST REC1 MWR G2 AND2 &T3 T4 EXOR1 1 = S3 Data Recorder T5 E0 A FRM2 F E0 A FRM3 F E A KT2 KT LED2 E A MW1 MW Signal smoothing Neues File Output: set - and actual temperature Integrator-reset (to prevent strong overshoot) E0 A FRM4 F E0 A FRM5 F Heating current limiter Plotting on/off CK RST U/D ENT ENP RCO Q3 Q2 Q1 Q0 ZBIN1 Zähler (4) T6 a b c d e f g a b c d e fg EN IN S0 S1 S2 A7 A6 A5 A4 A3 A2 A1 A0 ADMX1 ADMX S0 S1 S2 S3 g f e d c b a Seg7-Dekoder FW10 W Program selection display E A KT3 KT E A KT4 KT E A KT5 KT E A KT6 KT E A KT7 KT E A KT8 KT Meilhaus Redlab USB Measurement Module D/A out A/D in program selector Multiplexer program timer Temperature program tables time vs temperature setting To modify temperature programs make corrections here PID parameter output Figure B.16: Software-based PID-temperature controller. Implementation with ProfiLab Expert 3.0 (ABACOM GbR). The RedLab measurement module (Meilhaus) is employed for input/output of analog signals. Temperature can be set manual or in a program mode. Programs are entered as tables (ProfiLab-Function ”Korrekturtabelle”) of time versus temperature (recompilation necessary). Between two successive temperature set-points the temperature is varied linearly. The temperature controller application is run on the ”microscope control PC” in parallel with the image acquisition-software SimplePCI (Compix Inc.). 305
Supporting Information B.12 cRNA Secondary Structures ∆G37 o= -301.8 kcal/mol Minimum free energy secondary structure 5' 3' U A U A A G C A G A GC U G G U U U A G U G A A C C G UCA G A U C C G C U A G C G C U A C CG G U C G C C A C C A U GG U G A G C A A G G G C G A G G A G C U G U U C A C C G G G G U G G U G C C C A U C C U G G U C G A G C U G G A C G G C G AC G U A A ACGGCC A CAA GUUCAGC GUG UCCGG CGAGGGC G A GG G C G A U G C C A C C U ACGGCAA G C U GAC CC U G A A G U U C A U C U G C A C C AC C G G C A A GC U G C C C G U G C C C U G G C C C A C C C U C G U G A C C AC C C U G A C C U A C G G C G U G C A G U G C U U C A G C C G C U A C C CCGACC A C A UGAAGCAG C ACG A C U U C U UC AAGUCC G C C A U G C C C GAA G G CUAC G U C C A G G A G C G C ACC A U C U U C U UC A A G G A C G A C G G C A ACUAC A A G A CC C G C G C C G A G G UGAA G U U C G A G GG C G A C A C C CU G G U GA A C C G C A U C G A G C U G A A G G G CAUC G A C U U C AAG G A G GA C G G C AACA U C C U G G G G C A CAAG C U G G AG U A CAA C U A C A A C A G CC AC A AC G UCU AU AU C A U G G C C GA C A A G C A GAA G A A C G GCA U C A A G G U G A A C U U C AAG A U C C G C C A CAA C A UCGAGGACGGCAGCGUG CAGCUCGCCGA C C ACUACCA G C A G A A C A C C C C C A U C G G C G A C G G C C C C G U G C U G C U G C C C G A C A A C C A C UA C C U G A G C ACCC A G U C C G C C C UG A G CAA A G A C C C C A A C G A G A AGCGC GAUCA C A U G G U C C U G C U G G A G U U C GU G A C C G C C G C C G G G AU C A C U C U C G G C A U G G A C G A G C U G U A CAAG U C C G G A C U C A GAU C U C G A G U G C G U G A G U G C A U C UC C A U C C A CG U U G G C C A G 25 50 75 100 125 150 175 200 225 250 275 300 325 350 375 400 425 450 475 500 525 550 575 600 625 650 675 700 725 750 775 800 825 Figure B.17: The minimum free energy (MFE) secondary structure of the eGFP cRNA target sequence T2 – see section 8.6.2 – was calculated on the Sfold web server [Din04]. Owing to intrastrand base pairing large parts of the sequence are unavailable for hybridization to DNA microarray probes. The base numbering 1 to 825 corresponds to bases 556 to 1380 of the eGFP-Tub plasmid sequence (see section 8.6.2). Compare with the centroid structure in Fig. B.18. Green dots represent base pairs common in the MFE and centroid structures. Blue dots represent base pairs present only in the MFE structure. 306
cRNA Secondary Structures ∆G37 o= -200.66 kcal/mol Ensemble Centroid 5' 3' U A U A A G C A G A G C U G G U U U A G U G A A C C G U C A G A U C C G C U A G C G C U A C C G G U C G C C A C C A UG G U G A GC A A G G G C G AGGA G C U G U U C A C C G G G G U G G U G C C C A U C C U G G U C G A G C U G G A C G G C G A C G U A A ACGGCC A CAA GUUCAGC G U G UCC G G C G A G G G C G A G G G C G A U G C C A C C U ACGGCA A G C U G A C C C U G A A G U U C A U C U G CACCACCGGC A A GCU GCCCGUGCCCUGGCC C A C CCUC G U GACC A C C CU G A C C U A C G G C G U G C A G U G C U U C A G C C G C U A C C CCGACC A C A UGAAGCAG C ACG A C U U C U UC A AGUCC G C C A U G C C C GAA G G CUAC G U C C A G G A G C G C ACC A U C U U C U UC A A G G A C G A C G G C A ACUAC A A G A CC C G C G C C G A G G UGAAGUUC G A GGGCGA CACCCUGGUGAACC GC A U C G A G C U G A A G G G CAUCGACU U C A A G G A G G A CG GCA A C A U C CU G G G GC A C A AG C U G G AG U A CAA C U A C A A C A G C CACAACGUCUAUAUCAUGGCCGACAAGCAGAAGAA C G G C A U C A A G G U G A A C U U C A A G A U C C G C C A C A A C AUCGAGGACGGCAGC GUG CAGC UCGCCG A C CACUACCA G C A G A A C A C C C C C A U C G G C G A C G G C C C C G U G C U G C U G C C C G A CAACCACUACCUGAGCACCCAGUCCGCCCUGAGCAAAGACCCC A A C G A G A A G C G C G A U C A C A U G G U C C U G C U G G A G U U C G U G A C C G C C G C C G G G AU C A C U C U C G G C A U G G A C G A G C U G U A C A A G U C C G G A C U C A G A U C U C G A G U G C G U G A G U G C A U C U C C A U C C A C G U U G G C C A G 25 50 75 100 125 150 175 200 225 250 275 300 325 350 375 400 425 450 475 500 525 550 575 600 625 650 675 700 725 750 775 800 825 Figure B.18: Centroid secondary structure [Din05] of the eGFP cRNA target sequence T2. The centroid structure was calculated on the Sfold web server [Din04; Cha05] from a Boltzmann-weighted structure ensemble. ”The centroid structure can be considered as the single structure that best represents the central tendency of the set” [Cha05]. Compare with the minimum free energy secondary structure in Fig. B.17. Green dots represent base pairs common in the MFE and centroid structures. Red dots represent base pairs present in the centroid structure, not however in the MFE structure. 307
Supporting Information B.13 3-D Visualization of Nucleic Acid Structures Figure B.19: B-DNA structure - stereo view (use cross-eye-technique for 3D effect). Stereo images of the ideal B-DNA structure were created with UCSF Chimera. Figure B.20: A-RNA structure - stereo view (use cross-eye-technique for 3D effect). Stereo images of the ideal A-RNA structure were created with UCSF Chimera. 308
3-D Visualization of Nucleic Acid Structures Figure B.21: Top views of the helix structures - B-DNA (left) and A-RNA (right) - demonstrate significant differences in base stacking 309
Supporting Information 310
List of Publications List of Publications T. Naiser, T. Mai, W. Michel, A. Ott. A versatile maskless microscope projection photolithography system and its application in light-directed fabrication of DNA microarrays. Review of Scientific Instruments, 77(6): 063711, 2006. W. Michel, T. Mai, T. Naiser, A. Ott. Optical study of DNA surface hybridization reveals DNA surface density as a key parameter for microarray hybridization kinetics. Biophysical Journal, 92(3):999-1004, 2007. T. Naiser, O. Ehler, J. Kayser, T. Mai, W. Michel, A. Ott. Impact of point-mutations on the hybridization affinity of surface-bound DNA/DNA and RNA/DNA oligonucleotideduplexes: comparison of single base mismatches and base bulges. BMC Biotechnology 2008, 8:48 Submitted manuscripts T. Naiser, J. Kayser, T. Mai, W. Michel, A. Ott. DNA hybridization to surface bound probes: point defects in experiment and model. Submitted to Phys. Rev. Lett. T. Naiser, J. Kayser, T. Mai, W. Michel A. Ott. Point defects and the stability of surface bound oligonucleotide duplexes - experiments and model. Submitted to Biophysical Journal 311
List of Publications 312
Danksagung Danksagung Als erstes m¨ochte ich mich bei meinem Doktorvater, Herrn Prof. Dr. Albrecht Ott, f¨ur die großartige Betreuung meiner Doktorarbeit bedanken. Vielen Dank Albrecht, f¨ur die ausgesprochen freundschaftliche Zusammenarbeit mit Dir. Danke f¨ur die großen Freir¨aume die Du mir bei der Ausgestaltung der vorliegenden Arbeit gew¨ahrt hast, und auch daf¨ur, dass es praktisch jederzeit m¨oglich war Dich um Rat zu fragen und mit Dir wissenschaftliche Probleme zu er¨ortern. Bei meinen Mit-Doktoranden Timo Mai, Pablo Fernandez, Harish Bokkasam, J´erome Goidin, und ganz besonders bei Wolfgang Michel, mit dem ich ¨uber die vergangen Jahre das B¨uro geteilt habe, m¨ochte ich mich ebenfalls sehr herzlich f¨ur die freundschaftliche Zusammenarbeit und die angenehme und anregende Arbeitsatmosph¨are in unserer Arbeitsgruppe bedanken. Dieser Dank gilt ebenso Pramod Pullarkat und Jordi Soriano-Fradera, die mir beide als Postdocs w¨ahrend der vergangen Jahre stets mit viel Rat und Tat zur Seite standen, und ein wichtiger Quell der Motivationf¨ur mich waren – und sicher immer bleiben werden. War ’ne tolle Zeit mit Euch! Zum guten Arbeitsklima in der Arbeitsgruppe haben auch die Diplomanden Oliver Ehler, Jona Kayser, Benjamin Tr¨ankle und Philipp Baaske ihren Beitrag geleistet. Die gesellige Zeit bei unseren EP1-Kaffeepausen – zusammen mit Michael K¨uken, Ernesto Nicola, Sten R¨udiger, Snigdha Thakur, Cyril Colombo und Roberto Bernal – werd’ ich nie vergessen Freunde! Durch ihre kompetente technische Unterst¨utzung bei molekularbiologischen, softwaretechnischen und elektronischen Problemen haben auch Tobias Mummert, Andrea Hanold, Ralf Pihan und Paul Hurych ganz wesentlich zum Gelingen dieser Arbeit beigetragen. Habt vielen Dank! Auch J¨urgen Gmeiner vom Lehrstuhl EP II m¨ochte ich f¨ur seine Unterst¨utzung in Fragen der organischen Chemie an dieser Stelle meinen Dank aussprechen. F¨ur eine herausragende technische Unterst¨utzung m¨ochte ich mich auch bei Herrn Krejtschi und seinem Mechanikwerkstatt-Team bedanken. Besonders herzlicher Dank gilt der guten Seele des Lehrstuhls Margot Lenich f¨ur eine tolle Unterst¨utzung in administrativer Hinsicht, aber auch f¨ur viele nette Unterhaltungen zwischendurch. Danke Margot! F¨ur ihren Beitrag zur guten Arbeitsatmosph¨are am Lehrstuhl EP I m¨ochte ich mich auch bei Herrn Dr. Uwe Schmelzer, Herrn Prof. Dr. Pascher und seinen Mitarbeitern - meinen Freunden und Kollegen - Jens F¨urst, Wolfgang Kellner, Ralf Lang und Andreas Winter sehr herzlich bedanken. Bei Prof. Dr. Frank J¨ulicherm¨ochte ich mich daf¨ur bedanken, dass er mireinen mehrmonatigen Gastaufenthalt in seiner Arbeitsgruppe am MPIPKS in Dresden, und somit interessante 313