scieee AI-readable full text Open interactive document viewer

Refined characterization of Portuguese mtDNA diversity for forensic purposes

Ana Mafalda Fernandes da Silva Rocha

Full text

FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 2 Refined characterization of Portuguese mtDNA diversity for forensic purposes Ana Mafalda Fernandes da Silva Rocha Forensic Genetics Biology 2012 Supervisor Cíntia Alves, Head of the Genetic Identification and Parentage Testing Unit, Institute of Molecular Pathology and Immunology of the University of Porto (IPATIMUP) Co-supervisor Ana Goios, post-Doc Researcher, IPATIMUP Leonor Gusmão, Researcher, IPATIMUP FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 3 Acknowledgments Em primeiro lugar, gostaria de agradecer à Direção do IPATIMUP, por acreditarem na formação dos seus colaboradores e me terem possibilitado esta oportunidade. À minha orientadora e co-orientadora fisicamente presentes, Cíntia e Ana, muito obrigada pelo vosso apoio e ajuda preciosa. À minha co-orientadora Leonor que me deu o impulso e confiança que precisava. Mesmo distante, senti-te comigo. Obrigada pela amizade. Gosto muito de vocês! Luis (Fernandez), gracias por me ajudares com os problemas estatísticos. À minha Larocas, com quem passo tanto tempo e nunca me farto. Obrigada por estares sempre presente e por me ouvires sempre (sobre isto e sobre tantas outras coisas!). Ao meu Nunocas, por ser um companheiraço e saber bem como me espevitar o espírito. Obrigada pela disponibilidade constante e pela confiança que existe entre nós os três. À malta da hora do almoço, Lara, Inês e Cris, somos tão diferentes e tão iguais. Obrigada pelas gargalhadas, pelo desanuviar diário e pela amizade que tanto aprecio. Iva e Carina, quem me dera ter-vos cá. Obrigada pelos desabafos e encorajamentos. Obrigada também a todas as pessoas que iam passando e paravam para perguntar como estava a correr o trabalho. Aos meus pais e à minha irmã, por acreditarem em mim e no meu trabalho e pelo incentivo “Fazes isso com uma perna às costas!” Obrigada por estarem sempre presentes, nos bons e nos maus momentos. À minha cara-metade, Cirnes, por seres quem és e por tudo o que representas na minha vida. Obrigada pela confiança que depositas em mim, pela paciência e por achares que eu consigo fazer sempre mais que aquilo que eu própria acho. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 4 Ao meu João, que todos os dias me leva a querer ser uma pessoa melhor. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 5 Resumo A análise da diversidade genética tornou-se uma ferramenta útil em áreas como a ciência forense, genética de populações humanas e estudos de evolução molecular. Por todo o mundo, a quantidade de dados acumulados de diferentes populações é colossal. Apesar disso, a população portuguesa está pouco caracterizada no que respeita a região controlo completa do DNA mitocondrial (mtDNA) e carece de uma base de dados comportando sequências de boa qualidade. Com este estudo, proporcionou-se uma melhor caracterização da diversidade do mtDNA na população portuguesa colaborando e contribuindo para o enriquecimento da base de dados EMPOP (www.empop.org). A região controlo completa do DNA mitocondrial de 293 indivíduos não aparentados igualmente distribuídos pelo Norte, Centro e Sul de Portugal foi analisada e cada sequência foi classificada no respetivo haplogrupo. Foi também efetuada a genotipagem de polimorfismos específicos da região codificante (SNPs) em indivíduos portadores do haplogrupo H, que abrange mais de 40% da população da Europa Ocidental. No total, 104 sequências obtidas (35,6%) foram classificadas como haplogrupo H. As restantes sequências foram classificadas nos principais haplogrupos: R - 0,3%, B - 0,7%, J-7.9%, T - 11%, U - 13,7%, K – 6,2%, N - 0,3%, I – 3,8% W – 3,1%, X – 3,8%, R0 - 1,0%, L – 4,5%, M - 0,7% e D - 0,3% (de acordo com nomenclatura em PhyloTree.org – Árvore filogenética do mtDNA versão 13 de 28 dezembro 2011). Estes resultados estão de acordo com o esperado para a Europa Ocidental. Comparando as três regiões, Norte, Centro e Sul, a análise estatística revelou que a amostra portuguesa deve ser considerada como um todo em casos de rotina de genética forense, uma vez que as diferentes regiões não são significativamente distintas. Na amostra total de 293 indivíduos foram observados 239 haplótipos distintos com uma correspondente diversidade haplotípica de 0.9981. A sub-tipagem de sequências do haplogrupo H mostrou que 33% pertenciam ao sub-haplogrupo H1, 13% ao H3 e 9% ao H, corroborando o previamente relatado para a Europa Ocidental. Esta abordagem permitiu um aumento considerável na discriminação entre indivíduos classificados no haplogrupo H, demonstrando-se uma ferramenta valiosa em genética forense. Palavras-chave: genética forense, Portugal, DNA mitocondrial, haplogrupo, EMPOP. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 6 Abstract The analysis of mtDNA polymorphisms has become a useful tool in fields such as forensic science, human population genetics and molecular evolution studies. The amount of data accumulated from different populations all over the world is massive. Still, the Portuguese population is poorly characterized for the mtDNA complete control region and lacks a comprehensive database of good quality sequences. With this study, we have provided a better characterization of the mtDNA diversity in the Portuguese population by collaborating and contributing to the enrichment of the EMPOP database (www.empop.org). We have sequenced the mtDNA complete control region of 293 unrelated individuals equally distributed through North, Center and South of Portugal and undertaken haplogroup classification. Specific coding region single nucleotide polymorphisms (SNPs) were further genotyped by single base extension multiplex reactions in individuals with sequences classified within haplogroup H, which encompasses over 40% Western Europe population. A total of 104 studied mtDNA sequences (35.6%) were classified within haplogroup H. The remaining sequences were classified into the following major haplogroups: R - 0,3%, B - 0,7%, J-7.9%, T - 11%, U - 13,7%, K – 6,2%, N - 0,3%, I – 3,8% W – 3,1%, X – 3,8%, R0 - 1,0%, L – 4,5%, M - 0,7% AND D - 0,3% (nomenclature according to PhyloTree.org - mtDNA tree Build 13, 28 December 2011). These findings are in agreement with what is expected for Western Europe. By comparing the three regions, North, Center and South, statistical analysis revealed that the Portuguese sample should be considered as a whole in routine forensic genetic casework, as they are not significantly distinct. In total sample of 293 individuals, 239 distinct haplotypes were observed with a corresponding haplotype diversity of 0.9981. Sub-typing of haplogroup H sequences showed that 33% belong to sub-haplogroup H1, 13% to H3 and 9% to H, which is also similar to what has been previously reported for Western Europe. This approach allowed for a considerable increase in discrimination between haplogroup H classified individuals, making it a valuable tool in forensic genetics. Keywords: forensic genetics, Portugal, mitochondrial DNA, haplogroup, EMPOP. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 7 Contents Acknowledgments Resumo Abstract Contents List of Tables List of Figures Abbreviations 1. INTRODUCTION -------------------------------------------------------------------------------------- 16 1.1 A Brief Overview on Forensic Genetics --------------------------------------------------- 17 1.2 Forensic DNA Databasing --------------------------------------------------------------------- 18 1.2.1 Genetic basis for DNA databasing ---------------------------------------------------------- 18 1.2.1.1 Short tandem repeats (STRs) ---------------------------------------------------------- 19 1.2.1.2 Single nucleotide polymorphisms (SNPs) ------------------------------------------- 20 1.3 Mitochondrial DNA (mtDNA) ------------------------------------------------------------------ 21 1.3.1 Genetic features of mtDNA ------------------------------------------------------------------- 21 1.3.2 Analysis of mtDNA for forensic purposes -------------------------------------------------- 23 1.3.2.1 EMPOP mtDNA Population Database ------------------------------------------------ 25 1.3.2.2 mtDNA typing concerns ------------------------------------------------------------------ 26 1.3.2.2.1 Technical related ----------------------------------------------------------------------- 26 1.3.2.2.2 Structure related ----------------------------------------------------------------------- 27 1.3.2.3 Accessing population substructure ---------------------------------------------------- 29 1.3.2.4 Using mtDNA for human population history and evolution ---------------------- 29 1.3.2.5 mtDNA – current portrait of Portugal -------------------------------------------------- 32 1.4 Objectives ------------------------------------------------------------------------------------------- 33 FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 8 2. MATERIAL AND METHODS ----------------------------------------------------------------------- 34 2.1 Selection of individuals ------------------------------------------------------------------------- 35 2.2 DNA extraction ------------------------------------------------------------------------------------ 36 2.2.1 Whole blood or FTA® card blood stain ----------------------------------------------------- 36 2.2.2 Buccal smear (in ethanol) --------------------------------------------------------------------- 36 2.3 mtDNA complete control region (CR) ------------------------------------------------------ 37 2.3.1 General considerations ------------------------------------------------------------------------ 37 2.3.2 PCR amplification of mtDNA complete control region (CR) --------------------------- 40 2.3.2.1 PCR amplification verification ---------------------------------------------------------- 41 2.3.2.1.1 Polyacrylamide gel T9C5 and silver nitrate coloration method -------------- 41 2.3.2.1.2 QIAxcel system, automated analysis of DNA fragments --------------------- 41 2.3.2.2 Purification of PCR amplified samples ----------------------------------------------- 42 2.3.3 mtDNA sequencing reaction ------------------------------------------------------------------ 42 2.3.4 Purification of mtDNA sequenced products ----------------------------------------------- 43 2.3.5 Automated sequencing of mtDNA products ----------------------------------------------- 43 2.3.6 Visualization and analysis of mtDNA sequences ---------------------------------------- 44 2.3.7 Classification of haplogroups ----------------------------------------------------------------- 44 2.3.8 Statistical analysis ------------------------------------------------------------------------------- 45 2.4 Determination of sub-haplogroups of haplogroup H ---------------------------------- 46 2.4.1 General considerations ------------------------------------------------------------------------ 46 2.4.2 Amplification by PCR of specific SNPs of haplogroup H ------------------------------- 46 2.4.3 Purification of PCR amplified samples ----------------------------------------------------- 48 2.4.4 SNaPshot reaction ------------------------------------------------------------------------------ 48 2.4.5 Purification of SNaPshot products ----------------------------------------------------------- 50 2.4.6 Automated electrophoretic separation and detection of SNaPshot products ----- 50 2.4.7 Visualization and analysis of SNaPshot products --------------------------------------- 51 2.4.8 Classification of sub-haplogroups of haplogroup H ------------------------------------- 52 FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 9 3. RESULTS AND DISCUSSION -------------------------------------------------------------------- 54 3.1 Short troubleshooting section ---------------------------------------------------------------- 55 3.2 mtDNA control region typing ----------------------------------------------------------------- 57 3.2.1 Point heteroplasmy ----------------------------------------------------------------------------- 60 3.2.2 Length variant polymorphisms --------------------------------------------------------------- 62 3.2.2.1 Indels in homopolymeric stretches ---------------------------------------------------- 63 3.2.2.1.1 Length heteroplasmy in HVS-I ------------------------------------------------------ 64 3.2.2.1.2 Length heteroplasmy in HVS-II and HVS-III ------------------------------------- 66 3.2.2.1 Other indels --------------------------------------------------------------------------------- 68 3.3 mtDNA haplogroup diversity in Portugal ------------------------------------------------- 69 3.4 Statistical analysis results --------------------------------------------------------------------- 71 3.5 Haplogroup H sub-typing ---------------------------------------------------------------------- 74 4. CONCLUSIONS --------------------------------------------------------------------------------------- 77 5. REFERENCES----------------------------------------------------------------------------------------- 79 FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 16 1. Introduction FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 17 1.1 A Brief Overview on Forensic Genetics The word “forensic” comes from the Latin forēnsic, which means “of or before the forum” referring to the public discussion of a criminal charge in the Ancient Roman forum, the center of Roman public life (Lello & Lello, 1980). Nowadays, the term “forensic” is applied whenever a specific subject is related to legal actions and courts. To answer to legal questions, a vast amount of forensic sciences has emerged over time. Of these, forensic genetics and direct analysis of DNA, which began about 30 years ago with the case of Colin Pitchfork in the UK (where DNA evidence was used in a criminal case for the first time), has gained an important role worldwide in what concerns crime investigation and relationship testing. This has brought up the need to standardize and increase the quality of the methods applied (Goodwin et al., 2011). Besides working with material exclusive of each individual (DNA), forensic genetics is different from other forensic sciences because it has the ability to calculate expected values based on theoretical models and, therefore, its results do not present ambiguities (Saks & Koehler, 2005). However, because it is time and money consuming, only small ranges of DNA are analyzed and used for forensic purposes. Consequently, just because two DNA profiles match, that does not mean they belong to the same individual. To help estimate the rarity (frequency) of a DNA profile, it has been necessary to establish and continuously expand population forensic DNA databases, which are collections of DNA profiles obtained from unrelated individuals of a particular ethnic/population group. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 18 1.2 Forensic DNA Databasing Although there are forensic databases composed of offender’s profiles, unidentified human remains or relatives of missing persons, such as UK NDNAD (the first to be established) (Werret, 1997) or CODIS (http://www.fbi.gov/about-us/lab/codis), this work will focus on forensic databases composed of non-offender profiles, randomly sampled, of particular population groups, where donors voluntarily consent the collection of a biological sample for analysis and provide relevant data. These databases, along with providing a better genetic characterization of a specific population (population genetic studies), will also serve as estimators of the frequency of any questioned profile (Goodwin et al., 2011). 1.2.1 Genetic basis for DNA databasing DNA databases in forensic genetics rely on the fact that there are portions of our DNA that differ from one individual to another, generally called polymorphisms. Among other properties, it is important that these polymorphisms vary widely between individuals (high mutation rate). Various markers are combined, with the objective of creating unique profiles. The most common polymorphisms studied in forensic genetics are Short Tandem Repeats (STRs) and Single Nucleotide Polymorphisms (SNPs): FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 19 1.2.1.1 Short tandem repeats (STRs) STR markers are units of 2–6 bp in length repeated in tandem (Figure 1), widespread throughout autosomal and sex chromosomes and account for approximately 3% of the total human genome (Ellegren, 2004). They have a relatively high mutation rate, which is the reason for their high degree of polymorphism. Nevertheless, STR mutation rates are still relatively low (less than 0.1% per generation depending on the loci) which is important when testing, for example, paternity situations where exclusions could be due to mutational events (Butler, 2005). Another great advantage of STRs is that they are easily amplifiable by multiplex PCR. The combination of various autosomal STR loci increases dramatically the power of discrimination (Goodwin et al., 2011). The necessity of standardized methodologies to allow for comparisons between different labs around the world and for sound interchange of genetic information, has resulted over the years in the selection of specific sets of autosomal STR loci. Nowadays, Europe and the USA have adopted their own set of core STR systems for national databasing, composed of 12 and 13 loci respectively, sharing over half of the chosen markers between them. These different adoptions have led commercial companies to develop and validate various STR multiplex systems in order to fulfill the different needs and objectives of the forensic community worldwide. STRs are still the polymorphisms of choice in terms of genetic identification and kinship analysis. However, in some cases its use may be limited, namely in highly degraded DNA analysis. Figure 1: Schematic representation of a STR with 8 repeat units, each with 4bp in length, and one with 10 repeat units. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 20 1.2.1.2 Single nucleotide polymorphisms (SNPs) A SNP is a DNA sequence variation occurring within a single nucleotide — A, T, C or G (Figure 2). SNP mutation rate is very low when compared to STRs [change in the order of once every 108 generations (Brookes, 1999)] and are often found to be populationspecific (Bamshad et al., 2003), which can be helpful when predicting the ethnic origin/geographic ancestry of a profile (Freakish et al., 2003). Another advantage of SNP typing is when samples present a high level of DNA degradation (such as crime scene samples or remains) because the size of the amplified products is significantly lower, making its detection more effective when compared to STR analyses (Dixon et al., 2005; 2006). Due to its lower discrimination power [between 50 and 80 SNPs would be required to achieve the same level of discrimination as the modern STR-based methods (Gill, 2001)], autosomal SNPs are not widely used for forensic identification or establishment of profile databases. Preferably they should be used, whenever possible, to complement previous genetic information (Amorim & Pereira, 2005). The use of SNPs in forensic genetics, however, has become very common in the analysis of mitochondrial DNA (mtDNA), due to its particular characteristics that are described in the next chapter. Figure 2: Schematic representation of a SNP. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 21 1.3 Mitochondrial DNA (mtDNA) 1.3.1 Genetic features of mtDNA Mitochondria are eukaryotic cellular organelles responsible for generating about 90% of the energy the cell needs to function (Mitchell, 1961; Saraste, 1999). They contain their own genome, the mitochondrial DNA (Barrell et al., 1979) (Figure 3), which encompasses 0.25% of total DNA content per cell and has a different genetic code than nuclear DNA (Scheffler, 1999). The human mitochondrial genome is a double-stranded circular molecule composed of approximately 16569bp. It comprehends two parts: the coding region, which encodes for 13 proteins, 22 transfer RNAs and two ribosomal RNAs (12S and 18S) and the 1122bp non-coding or control region, also known as D-loop (displacement loop), with regulatory functions (Anderson et al., 1981; Andrews et al., 1999). Figure 3: A. Location of mtDNA and mitochondria inside an animal eukaryotic cell. B. Representation of the mitochondrial genome. Adapted from http://www.genome.gov/. Illustration by Darryl Leja. B A mtDNA ~16569bp FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 22 Mitochondrial DNA has special features that distinguish it from nuclear DNA. These characteristics have gathered researchers’ attention: High and variable copy number per cell In contrast with the two (maternal and paternal) copies per cell of nuclear DNA, mtDNA has on average 500 copies per cell (Satoh & Kuroiwa, 1991). Its circular nature also makes it less susceptible to exonucleases. This has proven to be extremely important when analyzing specimens where there is insufficient nuclear DNA, such as hair shafts, bones or teeth or degraded samples (Butler, 2005). Maternal inheritance and absence of recombination Maternal inheritance has been accepted as the rule for human mtDNA (Giles et al., 1980; Taylor et al., 2003; Schwartz & Vissing 2004; Bandelt et al., 2005) (Figure 4). As a consequence, mitochondrial DNA does not recombine. These two characteristics enable researchers to trace maternal related lineages back through time, clarifying the history of populations. In the forensic context, it is also important when no direct relatives are available for analysis and other living relatives need to be used as reference samples. It is noteworthy to emphasize here that mtDNA analysis does not discriminate between individuals but rather between mtDNA lineages, since individuals belonging to the same maternal line will share the same mtDNA haplotype (unless a mutation occurs). Figure 4: Schematic representation of mtDNA pattern of inheritance. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 23 High mutation rate The mitochondrial genome presents a significantly higher mutation rate than that observed for nuclear DNA (Brown et al., 1979). This high mutation rate is due, in part, to the exposure of mtDNA to mutagenic reactive oxygen species (ROS) that are produced as byproducts in oxidative phosphorylation (Hashiguchi et al., 2004). Mutation rates vary along the mtGenome. It has been proposed that the coding region evolves at 1.26x10-8 substitutions per site per year (Mishmar et al., 2003); the mutation rate for the control region has been estimated between 5x10-7 and 5x10-6 substitutions per site per year (Parsons et al., 1997). Even within the control region, some blocks are highly conserved but others, which are hypervariable segments I, II and III, contain the highest levels of variation within the mtGenome; furthermore, inside these hypervariable regions, some sites are mutational hotspots (Meyer et al., 1999). The key in forensic investigations is to differentiate individuals and, as mentioned above, the most extensive mtDNA variations occurring in the human populations are found within hypervariable segments I, II and III. Therefore, these are the regions commonly analyzed in forensic genetics. 1.3.2 Analysis of mtDNA for forensic purposes The hypervariable segments I, II and III (HVS-I, HVS-II and HVS-III) of the noncoding region (Figure 5) are normally examined by PCR amplification followed by sequence analysis. The nomenclature used to describe the sequences observed is based on the comparison with a reference sequence. In 1981, Anderson and his team were the first to sequence the mtDNA genome of a European individual (Cambridge Reference Sequence - CRS) (Anderson et al., 1981). The CRS was revised almost 20 years later, correcting some sequencing errors, to what is called today, the rCRS (revised Cambridge Reference Sequence) (Andrews et al., 1999). It is this recent rCRS that is used nowadays as a reference for typing all polymorphisms that are found in the mitochondrial DNA sequences. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 24 When mtDNA was first introduced in population and forensic genetic analyses (more than a decade ago), the source of diversity was thought to be concentrated mainly in HVS-I and HVS-II, and most of the labs would sequence these regions only, on a routine basis. Later, the recognition that HVS-III analysis would introduce a higher discrimination power (Fridman & Gonzalez, 2009), as well as the accompanying evolution of methodologies, has consequently taken the majority of the scientific community to sequence the complete control region (Carracedo et al., 2010). The typing of coding region SNPs also emerged in an attempt to fulfill the need for increasing discrimination power, especially in cases where the control region is not informative enough to discriminate most lineages. Although this methodology is gaining more followers (Quintáns et al., 2004; Coble et al., 2006), it has been recently pointed out that, with new technologies such as next-generation sequencing full mtDNA sequencing would be the best choice (Irwin et al., 2011). Nevertheless, for forensic applications the quantity of information retrieved from mtDNA is not the only factor necessary to take into account. The generated sequences should have excellent quality and undergo quality control validations. According to current guidelines (Carracedo et al., 2000; Bär et al., 2000; Tully et al., 2001; Parson & Dϋr, 2007; Prieto et al., 2011) full double-stranded high quality sequence coverage is the minimum requirement for reporting mtDNA data with forensic value; for particular hard-to-read polymorphisms, new sequence coverage should be retrieved. Nonetheless, because this requires effort and money, several mtDNA sequences available in the literature today do not meet these requirements and, therefore, are not suitable as reference for forensic applications. Figure 5: The three hypervariable (HV) regions of the mtDNA control region. HV1 (HVS-I) spans nucleotide positions 16024–16365 (342 bp), HV2 (HVS-II) spans positions 73–340 (268 bp), and HV3 (HVS-III) spans positions 438–574 (137 bp). Some polymorphic positions for variable regions VR1 and VR2 are also noted. Adapted from Butler, 2005) FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 25 Another effort made to improve the quality of mtDNA sequence data, was the establishment of strict guidelines and the existence of databases exclusively focused on this issue, such as EMPOP. 1.3.2.1 EMPOP mtDNA Population Database EMPOP - http://www.empop.org - was established in 2006 as a result of a collaborative project (Figure 6), under the scientific auspices of the European DNA Profiling Group (EDNAP). Its primary objective is to function as a depository for high quality mtDNA data usable for forensic purposes and educate researchers and technicians on how to produce it (Parson & Dϋr, 2007). EMPOP database relies on powerful tools to monitor and filter all mtDNA data submitted. Since almost 62% of the errors arise during manual transcription of the polymorphisms detected (Parson et al., 2004), EMPOP developed a laboratory information management (LIM) system which monitors every sequence’s journey since its arrival to the database, with a permanent link to raw data information. EMPOP also uses and offers to the general public, by download on its webpage, the NETWORK tool, to be used after data analysis and is an asset in the detection of potential errors. Figure 6: EMPOP logotype FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 32 1.3.2.5 mtDNA – current portrait of Portugal Twelve years ago, an approach to determine diversity of mtDNA lineages in Portugal based on HVS-I and HVS-II sequences (Pereira et al., 2000), revealed that haplogroup H was by far the most frequent haplogroup (about 40%) but haplogroups L (sub-Saharan) and U6 (North African) were present in a higher frequency than expected, with U6 oddly restricted to Northern Portugal – U6 group is characteristic of North Africa (Richards et al., 1998) and it would be expected to show a higher frequency in Southern Portugal; it also concluded that distribution patterns were similar among the geographic regions studied (North, Central and South) and they could be merged into a global Portuguese sample. As for mtDNA coding region SNP typing, Pereira et al. (2006) demonstrated that sub-clades H1 and H3 showed frequency peaks in Iberia and surrounding areas. Nonetheless, the advances in research and technology have improved analysis and have set new guidelines; nowadays, according to current procedures, these results previously published are incomplete and, therefore, need improvement in order to increase its informativeness as a forensic mtDNA database for the Portuguese population. In summary, before this work began, the Portuguese population was poorly characterized for the mtDNA complete control region and lacked a comprehensive database of good quality sequences to be used for forensic purposes. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 33 1.4 Objectives  Provide a characterization of the mtDNA complete control region diversity in the Portuguese population;  Compare the three major geographical regions, North, Center and South, to understand whether the sample should be considered as a whole, when applied to routine forensic genetic casework;  Determine sub-haplogroups of haplogroup H individuals, which encompasses over 40% of the total mtDNA variation in Western Europe, by single base extension multiplex reactions;  Collaborate and contribute to the EMPOP database (www.empop.org) by providing good quality mtDNA complete control region sequences.  Serve as an auxiliary guide, especially for forensic geneticists, by describing and exemplifying technical issues related with mtDNA typing, in what concerns the achievement of high quality sequence data and the problems that one may face with nomenclature issues. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 34 2. Material and Methods A FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 35 2.1 Selection of individuals A total of 293 unrelated individuals, born and living in districts representing three regions of Portugal (North, N=98; Central, N=95; and South, N=100) were selected. These three regions were established according to the country division by the major rivers Douro and Tagus (Figure 8). Blood or buccal smears were collected from voluntary donors, under informed consent. All blood samples were stored in classic FTA ® Whatman ® cards and buccal smears in microtubes with ethanol, before DNA extraction. Figure 8: Representation of Portugal and region division by rivers Douro and Tagus. The map also shows the number of individuals selected from each district. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 36 2.2 DNA extraction DNA extraction was performed by using the resin Chelex 100 method adapted from Lareu et al. (1994). Depending on the initial type of biological material, two different protocols took place: 2.2.1 Whole blood or FTA® card blood stain For whole blood samples, 10µl were taken to a 1.5ml microtube, and for blood stains, a 3x3mm square was cut to a 1.5ml microtube. A volume of 1ml of water was added and incubated for 30 minutes, stirring once in a while. Tubes were then centrifuged at 18620xg for 4 minutes and almost all supernatant was discarded. A volume of 200µl from a 5% Chelex 100 (200-400 mesh) stirring solution was added to the microtubes and these were stirred in a vortex. Samples were then incubated for 8 minutes at 100⁰C. Finally, the tubes were stirred and centrifuged at 18620xg for 4 minutes. Temporary storage was done at 4⁰C while in use; otherwise at -20⁰C for final storage conditions. 2.2.2 Buccal smear (in ethanol) Tubes were centrifuged at 18620xg for 10 minutes and almost all ethanol was discarded by inverting the microtubes. Total evaporation of the ethanol was undertaken using a thermal plate set to 60-80⁰C. A volume of 200µl from a 5% Chelex 100 (200-400 mesh) solution was added to the microtubes and these were stirred in a vortex. Samples were then incubated for 8 minutes at 100⁰C. Finally, the tubes were stirred and centrifuged at 18620xg for 4 minutes. Temporary storage was done at 4⁰C while in use; otherwise at -20⁰C for final storage conditions. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 37 2.3 mtDNA complete control region (CR) 2.3.1 General considerations According to established guidelines for generating high quality data (Carracedo et al., 2000; Bär et al., 2000; Tully et al., 2001; Parson & Dϋr, 2007; Prieto et al., 2011), amplification and sequencing of mtDNA complete control region must be done in a way that all mtDNA sequence between positions 16024-16569 and 1-576 is fully covered, both in forward and reverse directions, at least two times with different primers (see table 1). Table 1: List of primers used for amplification/sequencing of the mtDNA control region. a primers are designated according to the nucleotide number of the 5’-end in the rCRS; b primers designed with Primer3 in http://frodo.wi.mit.edu/primer3/. Primera Primer sequence 5’-3’ Reference L15900 TAA ACT AAT ACA CCA GTC TTG TAA ACC Parson et al., 2007 L15978 CAC CAT TAG CAC CCA AAG CT Alonso et al., 2003 L16268 CAC TAG GAT ACC AAC AAA CC Alonso et al., 2003 L16536 CCC ACA CGT TCC CCT TAA AT New* L00314 CCG CTT CTG GCC ACA GCA CT Parson et al., 2007 H00036 CCC GTG AGT GGT TAA TAG GGT Parson et al., 2007 H00159 AAA TAA TAG GAT GAG GCA GGA ATC Parson et al., 2007 H00599 TTG AGG AGG TAA GCT ACA TA Parson et al., 2007 H00639 GGG TGA TGT GAG CCC GTC TA Parson et al., 2007 FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 38 There are some special situations, such as samples that display length heteroplasmy, sequencing artifacts or sub-optimal reaction and electrophoresis conditions, which need a second amplification with a new pair of primers followed by sequencing reactions also with new initial and intermediate primers, according to the activity diagram shown in Figure 9. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 39 Figure 9: Activity diagram showing the sequential steps for the amplification and sequencing of mtDNA complete control region. The repeat of PCR amplification and/or sequencing reactions due to bad quality sequences only occurred once. If the result was not satisfactory again, the sample was discarded. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 40 2.3.2 PCR amplification of mtDNA complete control region (CR) All samples were amplified using QIAGEN Multiplex PCR Kit, Master Mix 2x, following the guidelines included in QIAGEN Multiplex PCR Handbook (QIAGEN 2010). The primer mix was prepared so that each one is at a final concentration of 2µM in the primer mix. In this case (two primers, forward and reverse) the primer mix solution was prepared by adding 2µl of forward primer plus 2µl of reverse primer taken from the stock solution (100µM) to 96µl of Ultra-Pure Water RNA/DNA free. The reagent mix, for a final volume of 10µl, was formulated by adding 5µl of QIAGEN Multiplex PCR Kit, Master Mix 2x, 1µl of primer mix, 3µl of water and 1µl of DNA (±0.5-4ng/μl). The tubes were placed in a thermal cycler (Figure 10). To avoid contamination issues, all samples were processed in a pre-PCR room (Bär et al., 2000). . Figure 10: Thermal cycler settings for PCR amplification of mtDNA complete control region (CR) (Veriti® 96Well Fast Thermal Cycler, Life Technologies). FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 41 2.3.2.1 PCR amplification verification 2.3.2.1.1 Polyacrylamide gel T9C5 and silver nitrate coloration method To check whether PCR amplification was successful and also to estimate mtDNA quantity, the amplified products were run on a polyacrylamide gel T9C5 (9% acrylamide; 5% N-N-metylene-bis-acrylamide) in 0.375M Tris/HCl buffer (pH=8.8). The gel was set up by mixing 4,4ml of T9C5, 0.35ml of glycerol, 0,25ml of ammonia persulphate 1% and 10µl of TEMED (Tetramethylethylenediamine). The gel polymerized between glass supports on a special film which was later placed on a horizontal refrigerated electrophoresis device. Whatman ® chromatography papers were impregnated with a 0.125M Tris/glycine buffer (pH=8.8) and positioned in each extremity of the gel. Power supply run voltage was set to 180V. The visualization of mtDNA bands was done according to Budowle et al. (1991) by silver nitrate staining method. 2.3.2.1.2 QIAxcel system, automated analysis of DNA fragments Alternatively to the previous procedure, the automated system QIAxcel was used (QIAGEN 2011). Products were run on the instrument according to manufacturer’s specifications (the cartridge used was QIAxcel DNA Screening Kit, alignment marker 15bp-5kb, method chosen was AM320). FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 48 95⁰C 15min 94⁰C 60⁰C 72⁰C 72⁰C 30sec 90sec 1min 15min Initial denaturation x 1 Denaturation Annealing Extension Final extension x30 x 1 To avoid contamination issues, all samples were processed in a pre-PCR room (Bär et al. 2000). 2.4.3 Purification of PCR amplified samples Prior to SNaPshot reaction, it was necessary to remove primers and free nucleotides unincorporated during PCR reaction that, otherwise, would interfere with the SNaPshot reaction. Products were purified by adding 0.5µl of ExoSAP-IT (Exonuclease I + Shrimp Alkaline Phosphatase, GE Healthcare) to 1µl of PCR product. The mixture was incubated first for 15min at 37⁰C and second for 15min at 85⁰C, leading to enzymes inactivation. 2.4.4 SNaPshot reaction The reaction was set to a final volume of 5µl, of which, 1µl SNaPshot Multiplex Kit, 1.5µl of SBE mix, 2µl of water and 0.5µl of purified PCR product. The primer mix was prepared so that each one is at a specific final concentration described in table 3. Figure 12: Thermal cycler settings for amplification by PCR of specific SNPs of haplogroup H (Veriti® 96-Well Fast Thermal Cycler, Life Technologies). FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 49 1438 0.06 2259 0.2 8994 0.06 10166 0.3 10211 Multiplex 3 0.15 Alvarez-Iglesias et al., 2009 11140 0.2 14869 0.2 14872 0.4 15833 0.2 13101C 0.2 Markers SBE Multiplex SBEs conc. (µM) Reference 709 0.2 750 0.2 2581 0.2 3010 0.2 3796 0.1 6253 0.1 6296 0.2 6365 0.2 6776 0.3 7337 Multiplex 1 0.3 Alvarez-Iglesias et al., 2009 10810 0.2 12858 0.3 12957 0.3 13708 0.3 13759 0.1 14365 0.3 14470A 0.4 951 0.2 3915 0.2 3936 0.6 3992 0.3 4310 0.4 4336 0.2 4727 0.2 4745 0.2 4769 0.2 4793 0.4 7028 Multiplex 2 0.2 Alvarez-Iglesias et al., 2009 7645 0.2 8269 0.06 8473 0.2 8592 0.1 8598 0.4 8602 0.4 9066 0.3 9150 0.2 10044 0.6 10394 0.3 13404 0.3 8271T 0.3 961G 0.06 Table 3: Multiplex reactions performed by SNaPshot, for detection of coding region SNPs of haplogroup H mtDNA sequences. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 50 96⁰C 50⁰C 60⁰C 10sec 5sec 30sec Denaturation Annealing Extension x25 The tubes were then placed in a thermal cycler (Figure 13). Figure 13: Thermal cycler settings for SNaPshot reaction (Veriti® 96-Well Fast Thermal Cycler, Life Technologies). 2.4.5 Purification of SNaPshot products Products were purified by adding 1µl of SAP (Shrimp Alkaline Phosphatase, GE Healthcare) to total volume of SNaPshot products. The mixture was incubated first for 1h at 37⁰C and then for 15min at 85⁰C, leading to enzymes inactivation. 2.4.6 Automated electrophoretic separation and detection of SNaPshot products For enabling automated data analysis and achieving high run to run precision in sizing DNA fragments by electrophoresis it is necessary to use an internal lane size standard. For SNaPshot products, with final sizes smaller than 120bp, the choice was GeneScan™ 120 LIZ® Size Standard (Applied Biosystems). The size standard was diluted in Hi-Di Formamide in a proportion of 1:33, which is below manufacturers’ guidelines, but proven to be sufficient. To 10µl of this mixture, 1µl of SNaPshot product was added. All products were run on an Applied Biosystems Genetic Analyzer 3130. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 51 SNP Size (pb) Detection Mutation 750 23 A→G A→G 2581 24 T→C A→G 709 32 G→A G→A 3010 37 C→T G→A 3796 37 A→G A→G 6296 48 C→T C→T 6253 48 A→G T→C 6365 54 A→G T→C 10810 54 T→C T→C 12858 58 C→T C→T 6776 58 A→G T→C 7337 62 G→A G→A 12957 62 T→C T→C 13759 66 G→A G→A 14365 70 G→A C→T 13708 78 G→A G→A 14470A 78 T→A T→A 4336 21 A→G T→C 951 21 C→T G→A 4745 32 A→G A→G 3915 32 C→T G→A 3992 38 C→T C→T 4310 38 A→G A→G 7028 44 C→T C→T 4727 44 A→G A→G 4769 50 A→G A→G 7645 50 T→C T→C 8473 56 T→C T→C 4793 56 A→G A→G 8269 62 G→A G→A 9066 62 T→C A→G 8592 66 G→A G→A 10044 66 T→C A→G 9150 70 A→G A→G 10394 70 C→T C→T 8602 74 T→C T→C 13404 78 T→C T→C 3936 78 G→A C→T 8271T 82 A→T A→T 8598 84 T→C T→C 961G 86 T→G T→G 2259 26 C→T C→T 1438 30 A→G A→G 10211 36 C→T C→T 10166 41 T→C T→C 11140 49 C→T C→T 13101C 54 T→G A→C 8994 64 G→A G→A 14872 68 C→T C→T 15833 74 C→T C→T 14869 77 G→A G→A Multiplex 3 Multiplex 1 Multiplex 2 2.4.7 Visualization and analysis of SNaPshot products For visualization and analysis of SNaPshot products, files were imported to GeneMapper ®4.0 software (Applied Biosystems, 2005). All positions were visually inspected and annotated according to table 4. Table 4: List of all SNPs under analysis, its detection and corresponding mutation. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 52 2.4.8 Classification of sub-haplogroups of haplogroup H The resulting set of SNPs was annotated for each sequence belonging to haplogroup H. Sub-haplogroups were determined according to current guidelines (Álvarez-Iglesias et al., 2009) (Figure 14). Most of the times, there was no need to perform the three multiplexes because the information gathered with one or two multiplexes was sufficient to classify the sequence into H sub-haplogroup. Additionally, when a SNP pattern was unexpected or dubious, a singleplex reaction for that specific SNP was performed. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 53 Figure 14: Phylogeny of haplogroup H adapted from Álvarez-Iglesias et al., (2009). Only positions studied within this work are represented. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 54 3. Results and Discussion FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 55 This chapter will start with a short section about troubleshooting during sequence generation in this work. Then, we will focus on mtDNA control region typing and nomenclature, which includes examples of problematic sequence stretches frequently encountered during mtDNA analysis in the course of this study. Afterwards, there will be a brief description of mtDNA haplogroup diversity found in Portugal. For an easier understanding, the graphical representations focus only on major haplogroups according to PhyloTree´s mtDNA tree Build 13 (28 December 2011). Later, we will concentrate on the results of population comparisons between the three major geographical regions, North, Centre and South, to understand whether the sample should be considered as a whole, when used as a reference in routine forensic genetic casework. Finally, we will refer to sub-haplogroup results of sequences classified as haplogroup H from complete control region data. Complete mtDNA profile data obtained in this work can be consulted in Appendix I and II. 3.1 Short troubleshooting section During the course of this work, we have dealt with some technical difficulties. The most significant problem encountered was when sequencing with reverse primers H00599 and H00159, because the sequences showed, most of the times, a sudden drop of peak intensity with continued low signal or no signal at all. With the aim of increasing the quality of the generated sequences, these issues were solved by changing the sequencing chemistry to dGTP BigDye® Terminator v3.0 Cycle Sequencing Ready Reaction Kit (Applied Biosystems). It is ideal for working with GTand G-rich and other difficult-to-sequence templates enabling the extension through difficult-to-sequence regions, avoiding early signal loss in these samples. Below, some examples are shown for a better elucidation of the problem (Figures 15 and 16). FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 56 EXAMPLE 1 EXAMPLE 2 Figure 15: Electropherograms from sample Z5304 with primer H00599. Top, using regular sequencing kit – BigDyeTerminator v3.1; bottom, using dGTP BigDye Terminator v3.0 sequencing kit. Figure 16: Electropherograms from sample Z6718 with primer H00159. Top, using regular sequencing kit – BigDyeTerminator v3.1; bottom, using dGTP BigDye Terminator v3.0 sequencing kit. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 57 3.2 mtDNA control region typing Each difference to the rCRS was reported (Andrews et al., 1999) according to current guidelines (Carracedo et al., 2000; Bär et al., 2000; Tully et al., 2001; Parson & Dϋr, 2007; Prieto et al., 2011). The number of distinct polymorphic substitutions and indels typed in our total sample are discriminated in table 5 and are in agreement with the fact that transitions occur at a much higher frequency than transversions (Tamura & Nei, 1993). For this calculation, heteroplasmic positions were not included (we only accounted for the dominant variant). Alignment and nomenclature for positions involving transitions and transversions is usually a straightforward task, and this work was no exception. Figures 17 and 18 show the sites where transitions and transversions occurred in HVS-I, HVSII and HVS-III and their frequency. Transitions T↔C Tranversions A↔G Indels North 110 617 Centre 111 414 South 119 725 Table 5: Number of distinct types of polymorphic substitutions and indels typed in this study. Results were computed in Arlequin. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 64 An additional problem that indels in homopolymeric stretches can generate is alignment ambiguities, especially due to multiple polymorphisms. Wilson et al. (2002a; 2002b) were the first to address these issues and have proposed three formal rules based on binary alignment, which distorted phylogenetic considerations. Nowadays, the consensus method to apply in these problematic cases is to take into consideration the phylogenetic alignment, in context with the sample’s complete haplotype (Bandelt & Parson, 2008). The following are some examples found in this study. 3.2.2.1.1 Length heteroplasmy in HVS-I Length heteroplasmy in HVS-I usually occurs along with the 16189C polymorphism. EXAMPLE 1: Sample Z6747 (Figure 23). The haplotype for this region is 16182C, 16183C, 16189C, 16193.1C. Note the G-peak of the dominant population at 16196 – when both G-peaks are equally dominant, it is advised to choose the first dominant G (most parsimony case) (EMPOP collaboration). 1st PCR amplification: L15978+H00639 Sequenced with L15978 2nd PCR amplification: L15900+H00599 Sequenced with L15900 rCRS 1st PCR amplification: L15978+H00639 Sequenced with H00036 dominant G 16182C 16183C 16189C 16193.1C Figure 23: Length heteroplasmy in sample Z6747. Sequence image from SeqScape v2.6. The software converts reverse sequences into forward ones automatically. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 65 EXAMPLE 2: Sample Z6005 (Figure 24). The haplotype for this region is 16189C. This sample shows two different populations of poly-C stretches (the other population has position 16193 deleted) but only the dominant variant is annotated. EXAMPLE 3: Sample Z2290 (Figure 25). The haplotype for this region is 16189C, 16191.1C, 16192T (phylogenetic alignment). According to formal rules, annotation would be 16189C, 16192.1T. 1st PCR amplification: L15978+H00639 Sequenced with L15978 2nd PCR amplification: L15900+H00599 Sequenced with L15900 rCRS 1st PCR amplification: L15978+H00639 Sequenced with H00036 16189C Figure 24: Length heteroplasmy in sample Z6005. Sequence image from SeqScape v2.6. The software converts reverse sequences into forward ones automatically. 16189C 16191.1C 16192T 1st PCR amplification: L15978+H00639 Sequenced with L15978 2nd PCR amplification: L15900+H00599 Sequenced with L15900 rCRS 1st PCR amplification: L15978+H00639 Sequenced with H00036 Figure 25: Length heteroplasmy in sample Z2290. Sequence image from SeqScape v2.6. The software converts reverse sequences into forward ones automatically. dominant G dominant G FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 66 3.2.2.1.2 Length heteroplasmy in HVS-II and HVS-III Insertions resulting in length heteroplasmy in HVS-II and HVS-III are very common around positions 309, 315 and 573. Here, we will present some examples. EXAMPLE 1: Sample Z2616 (Figure 26). The haplotype for this region is 309.1C, 315.1C. There are two different mtDNA populations present. For annotation, the dominant T-peak should be taken into consideration. If there are two equally dominant populations, EMPOP recommends determining the first dominant T (most parsimony case). EXAMPLE 2: Sample Z3585 (Figure 27). The haplotype for this region is 309.1C, 309.2C, 315.1C. There are two different mtDNA populations present, although the dominant population is clearly visible. 1st PCR amplification: L15978+H00639 Sequenced with L16536 1st PCR amplification: L15978+H00639 Sequenced with H00599 rCRS dominant T 309.1C 315.1C Figure 26: Length heteroplasmy in sample Z2616. Sequence image from SeqScape v2.6. The software converts reverse sequences into forward ones automatically. 1st PCR amplification: L15978+H00639 Sequenced with L16536 1st PCR amplification: L15978+H00639 Sequenced with H00599 rCRS 309.1C 309.2C 315.1C Figure 27: Slight length heteroplasmy in sample Z3585. Sequence image from SeqScape v2.6. The software converts reverse sequences into forward ones automatically. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 67 EXAMPLE 3: Sample Z6673 (Figure 28). The haplotype for this region is 573.1C, 573.2C, 573.3C, 573.4C, 573.5C. We took into account the dominant G at position 577 (with primer L00314, as L16536 shows loss of resolution). EXAMPLE 4: Sample Z2306 (Figure 29). The haplotype for this region is 573.1C, 573.2C, 573.3C. We took into account the dominant G at position 577 (with primer L00314, as L16536 shows loss of resolution). Both Gs are equally dominant so we selected the first one (most parsimony case). 1st PCR amplification: L15978+H00639 Sequenced with L16536 1st PCR amplification: L15978+H00639 Sequenced with L00314 rCRS dominant G 573.5C 1st PCR amplification: L15978+H00639 Sequenced with L16536 1st PCR amplification: L15978+H00639 Sequenced with L00314 rCRS dominant G 573.3C Figure 28: Length heteroplasmy in sample Z6673. Sequence image from SeqScape v2.6. The software converts reverse sequences into forward ones automatically. Figure 29: Length heteroplasmy in sample Z2306. Sequence image from SeqScape v2.6. The software converts reverse sequences into forward ones automatically. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 68 3.2.2.1 Other indels During the course of this work we have encountered different types of insertions and deletions unrelated to C-stretch tracts. The most common were reported within positions 523 and 524 (HVS-III). EXAMPLE 1: Sample Z6701 (Figure 30). The haplotype for this region is 524.1A, 524.2C, 524.3A, 524.4C. It was reported the dominant variant, although it is possible to observe length heteroplasmy - a second population without insertions 524.3A and 524.4C is also present. EXAMPLE 2: Sample Z6343 (Figure 31). The haplotype for this region is 523d, 524d. 1st PCR amplification: L15978+H00639 Sequenced with L00314 1st PCR amplification: L15978+H00639 Sequenced with H00639 rCRS 524.1A 524.4C 524.3A 524.2C 1st PCR amplification: L15978+H00639 Sequenced with L15978 1st PCR amplification: L15978+H00639 Sequenced with H00639 rCRS Figure 30: Insertions and length heteroplasmy in sample Z6701. Sequence image from SeqScape v2.6. The software converts reverse sequences into forward ones automatically. Figure 31: Deletions in sample Z6343. Sequence image from SeqScape v2.6. The software converts reverse sequences into forward ones automatically. Position 522 Position 525 FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 69 3.3 mtDNA haplogroup diversity in Portugal The mtDNA haplogroups described in this study are in agreement with the described landscape for Western Europe, with haplogroup H accounting for about 40% of our total sample (Figure 32). Haplogroups T and U are represented in about 25% of the mtDNA sequences. Haplogroups D, G and N, are the least represented by only one sequence for each in the Centre of Portugal (Figures 32, 33B), and are usually less frequent in Europe (Torroni et al., 1996). The haplogroup distribution through the three regions is very similar (Figure 33), except for haplogroup I which is significantly more represented in the southern region. Nonetheless, this haplogroup is also found in Western Europe, although in very low frequencies when compared to the Northern and Eastern regions of this continent (Simoni et al., 2000). Figure 32: Major mtDNA haplogroup distribution in Portugal. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 70 The general results obtained from the control region sequences are in agreement with what was observed in an earlier study (Pereira et al., 2000), except for the South where Pereira et al. (2000) reported 50% of mtDNA sequences belonging to haplogroup H and 1.7% to haplogroup I, contrasting with 33% for the first and 10% for the latter in this study. It is important to refer that the number of samples for the southern region in this study doubled the one analyzed by Pereira et al. (2000), which could have largely influenced the results. As referred by many authors, larger database sizes are more informative and reliable. A. B. C. Figure 33: Major mtDNA haplogroup distribution in A. North, B. Centre and C. South of Portugal FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 71 This previous study (Pereira et al., 2000) also found seven U6 sequences restricted to the North of Portugal, a situation described then as odd - U6 group is characteristic of North Africa (Richards et al., 1998) and it would be expected to show its higher frequency in Southern Portugal. In the present study, that is not the case: we report two U6 sequences in the North, one in the Center and two in the South, showing that a restriction of this haplogroup to the Northern region is not confirmed by this study. The increase in number of samples analyzed in this study (100 more samples in contrast to only 59 of the previous database) is probably responsible for this difference since it has been demonstrated that small databases will not encompass the less common haplotypes within a population (Goodwin et al., 2011). 3.4 Statistical analysis results It is well known that geographic patterns may shape regional population genetic structure (e.g. Cavalli-Sforza et al., 1994; Barbujani, 2000). In the past, the existence of major barriers like Douro and Tagus rivers could have prevented the exchange of genetic material (in this case, mtDNA) between populations of North, Centre and South of Portugal creating a population substructure and, therefore, three genetically differentiated subpopulations (Pereira et al., 2000; Beleza et al., 2006). To determine if this is the case, FST statistic, which is a very convenient and common way of comparing differentiation among two or more subpopulations, was used. In a broad sense, FST results range from 0 (meaning no structure/differentiation) to 1 (completely differentiated). As referred in the Methodology chapter, two approaches were undertaken to calculate these data: 1) mtDNA haplogroup frequencies were directly counted after haplogroup classification according to tables 7 and 8; 2) full CR mtDNA sequence data (haplotype frequencies) was used, between nucleotide positions 16024-576 (table 9). FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 72 Major Haplogroup North Centre South B 2 0 0 D 0 1 0 G 0 1 0 H 34 36 33 HV 7 5 8 I 1 0 10 J 7 9 7 K 7 9 2 L 3 5 7 M 0 1 1 N 0 1 0 R 2 0 0 R0 2 0 1 T 13 11 9 U 14 10 16 W 4 3 1 X 2 4 5 North Centre South North 0.00000 Centre (0.00000) 0.00000 South (0.00031) (0.00486) 0.00000 North Centre South North 0.00000 Centre (0.00000) 0.00000 South 0.00162 0.00259 0.00000 Both approaches showed that there is almost no differentiation (less than 0.5%) among the subpopulations studied. All the values reported for the first approach were non-significant perhaps due to the fact that there is a great unevenness between the haplogroup frequencies (more than 100 individuals in haplogroup H compared to 1 individual in haplogroup N, for example) (Table 8). Table 7: Number of sequences per major haplogroup and geographic area used for population comparison calculations. Table 8: FST genetic distances calculated with Arlequin v.3.5 after direct counting of haplogroup frequencies. Values in brackets are non-significant p>0.05. Table 9: FST genetic distances calculated in Arlequin v.3.5 using full CR mtDNA sequence data (haplotype frequencies), between nucleotide positions 16024-576. Values in brackets are non-significant p>0.05. ( ) ( ) ( ) FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 73 Although the same conclusion was reached, it is advised to use, if available, full sequence data to compute genetic distances, as this is not influenced (except for typing/technical errors) by haplogroup determination misjudgment. Moreover, the very different frequency distribution of the different haplogroups is bound to influence the significance of the results. Since no differentiation was observed between the three analyzed regions, North, Center and South of Portugal, the whole Portuguese data was joined for computation of standard and molecular diversity indices (Table 10). The mtDNA haplotype diversity present in the Portuguese population sample is high (0.9981) and 203 of the 293 haplotypes were observed only once. A maximum of six individuals with the same haplotype occurs only once, corresponding to the most common CR mtDNA haplotype 263G 315.1C 16519C (this data can be consulted through Appendix III). Other shared haplotypes are present in less than five individuals. Population Portugal N293 Number of haplotypes 239 Number of unique haplotypes 203 Segregating sites1210 Haplotype diversity 0.9981 Mean pairwise differences 11.68 ± 5.31 Table 10: Diversity indices for the 293 mtDNA control region sequences from Portugal based on complete control region haplotypes spanning positions 16024-576 (heteroplasmy was disregarded). N – Sample size. 1 Number of polymorphic sites observed in our sample (out of 1122 possible sites). FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 80 Achilli A, Rengo C, Magri C, Battaglia V, Olivieri A, Scozzari R, Cruciani F, Zeviani M, Briem E, Carelli V, Moral P, Dugoujon JM, Roostalu U, Loogväli EL, Kivisild T, Bandelt HJ, Richards M, Villems R, Santachiara-Benerecetti AS, Semino O, Torroni A (2004) The molecular dissection of mtDNA haplogroup H confirms that the Franco-Cantabrian glacial refuge was a major source for the European gene pool. Am J Hum Genet 75:910–918. Alonso A, Albarrán C, Martín P, García P, García O, De La Rúa C, Alzualde A, Fernández de Simón L, Sancho M, Fernández Piqueras J (2003) Multiplex-PCR of short amplicons for mtDNA sequencing from ancient DNA. Progress in Forensic Genetics 9, International Congress Series 1239: 585-588. Alvarez-Iglesias V, Mosquera-Miguel A, Cerezo M, Quintáns B, Zarrabeitia MT, Cuscó I, Lareu MV, García O, Pérez-Jurado L, Carracedo A, Salas A (2009) New population and phylogenetic features of the internal variation within mitochondrial DNA macrohaplogroup R0. PLoS One 4(4):e5112. Amorim A & Pereira L (2005) Pros and cons in the use of SNPs in forensic kinship investigation: a comparative analysis with STRs. Forensic Sci Int 150(1):17-21. Anderson S, Bankier AT, Barrell BG, de Bruijn MH, Coulson AR, Drouin J, Eperon IC, Nierlich DP, Roe BA, Sanger F, Schreier PH, Smith AJ, Staden R, Young IG (1981) Sequence and organization of the human mitochondrial genome. Nature 290:457–465. Andrews RM, Kubacka I, Chinnery PF, Lightowlers RN, Turnbull DM, Howell N (1999) Reanalysis and revision of the Cambridge reference sequence for human mitochondrial DNA. Nat Genet 23:147. Balding, DJ and Nichols, RA (1994) DNA profile match probability calculation: how to allow for population stratification, relatedness, database selection and single bands. Forensic Sci Int 64:125-140. Bamshad MJ, Wooding S, Watkins WS, Ostler CT, Batzer MA, Jorde LB (2003) Human population genetic structure and inference of group membership. Am J Hum Genet 72: 578–589. Bandelt H-J & Parson W (2008) Consistent treatment of length variants in the human mtDNA control region: a reappraisal. Int J Legal Med 122:11–21. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 81 Bandelt HJ, Kong QP, Parson W, Salas A. (2005) More evidence for non-maternal inheritance of mitochondrial DNA? J Med Genet 42(12):957-60. Bandelt HJ, van Oven M, Salas A. (2012) Haplogrouping mitochondrial DNA sequences in Legal Medicine/Forensic Genetics. Int J Legal Med [Epub ahead of print] Bär W, Brinkman B, Budowle B, Carracedo A, Gill P, Holland M, Lincoln P, Mayr W, Morling N, Olaisen B, Schneider P, Tully G, Wilson M (2000) DNA commission of the international society for forensic genetics: guidelines for mitochondrial DNA typing. Int J Legal Med 113: 193-196. Barbujani G (2000) Geographic patterns: how to identify them and why. Hum Biol 72(1):133-53. Barrell BG, Bankier AT, Drouin J (1979) A different genetic code in human mitochondria. Nature 282(5735):189-94. Behar DM, Villems R, Soodyall H, Blue-Smith J, Pereira L, Metspalu E, Scozzari R, Makkan H, Tzur S, Comas D, Bertranpetit J, Quintana-Murci L, Tyler-Smith C, Wells RS, Rosset S; Genographic Consortium (2008) The dawn of human matrilineal diversity. Am J Hum Genet 82(5):1130-40. Beleza S, Gusmão L, Lopes A, Alves C, Gomes I, Giouzeli M, Calafell F, Carracedo A, Amorim A (2006) Micro-phylogeographic and demographic history of Portuguese male lineages. Ann Hum Genet 70(Pt 2):181-94. Brandstätter A, Peterson CT, Irwin JA, Mpoke S, Koech DK, Parson W, Parsons TJ (2004) Mitochondrial DNA control region sequences from Nairobi (Kenya): inferring phylogenetic parameters for the establishment of a forensic database. Int J Legal Med 118(5):294-306. Brookes AJ (1999) The essence of SNPs. Gene 234: 177–186. Brown WM, George M Jr, Wilson AC (1979) Rapid evolution of animal mitochondrial DNA. Proceedings of the National Academy of Sciences of USA. 76(4):1967-71. Budowle B, Chakraborty R, Giusti AM, Eisenberg AJ, Allens RC (1991) Analysis of the VNTR Locus DIS80 by the PCR Followed by High-Resolution PAGE. Am J Hum Genet 48:137–144. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 82 Butler JM (2005) Forensic DNA Typing - Second Edition. Elsevier Academic Press. USA. Carracedo A, Bär W, Lincoln P, Mayr W, Morling N, Olaisen B, Schneider P, Budowle B, Brinkman B, Gill P, Holland M, Tully G, Wilson M (2000) DNA commission of the international society for forensic genetics: guidelines for mitochondrial DNA typing. Forensic Sci Int 110: 79-85. Carracedo A, Butler J, Gusmão L, Parson W, Rower L, Schneider P (2010) Publication of population data for forensic purposes. Forensic Sci. Int Genet. 4: 145-147. Cavalli-Sforza LL, Menozzi P, Piazza A (1994) The history and geography of human genes. Princeton University Press, Princeton, NJ7. Chen YS, Torroni A, Excoffier L, Santachiara-Benerecetti AS, Wallace DC (1995) Analysis of mtDNA variation in African populations reveals the most ancient of all human continent-specific haplogroups. Am J Hum Genet 57(1):133-49. Coble MD, Vallone PM, Just RS, Diegoli TM, Smith BC, Parsons TJ (2006) Effective strategies for forensic analysis in the mitochondrial DNA coding region. Int J Legal Med 120(1):27-32. Dixon LA, Dobbins AE, Pulker HK, Butler JM, Vallone PM, Coble MD, Parson W, Berger B, Grubwieser P, Mogensen HS, Morling N, Nielsen K, Sanchez JJ, Petkovski E, Carracedo A, Sanchez-Diz P, Ramos-Luis E, Briōn M, Irwin JA, Just RS, Loreille O, Parsons TJ, Syndercombe-Court D, Schmitter H, Stradmann-Bellinghausen B, Bender K, Gill P (2006) Analysis of artificially degraded DNA using STRs and SNPs – results of a collaborative European (EDNAP) exercise. Forensic Sci Int 164: 33–44. Dixon LA, Murray CM, Archer EJ, Dobbins AE, Koumi P, Gill P (2005) Validation of a 21-locus autosomal SNP multiplex for forensic identification purposes. Forensic Sci Int 154: 62–77. Ellegren H (2004) Microsatellites: simple sequences with complex evolution. Nat Rev Genet 5, 435-445. Excoffier L, Laval G, Schneider S (2005) Arlequin ver. 3.0: An integrated software package for population genetics data analysis. Evol Bioinform Online 1:47–50 FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 83 Fendt L, Zimmermann B, Daniaux M, Parson W (2009) Sequencing strategy for the whole mitochondrial genome resulting in high quality sequences. BMC Genomics 10:139. Freakish T, Venkateswarlu K, Thomas MJ, Gaskin Z, Ginjupalli S, Gunturi S et al (2003) A classifier for the SNP-based inference of ancestry. J Forensic Sci 48: 771782. Fridman C & Gonzalez RS (2009) HVIII discrimination power to distinguish HVI and HVII common sequences. Forensic Sci Int Genet Supplement Series 2: 320-321. Giles RE, Blanc H, Cann HM, Wallace DC (1980) Maternal inheritance of human mitochondrial DNA. Proceedings of the National Academy of Sciences of USA. 77(11):6715-9. Gill P (2001) An assessment of the utility of single nucleotide polymorphisms (SNPs) for forensic purposes. Int J Legal Med 114: 204–210. Goodwin W, Linacre A, Hadi S (2011) An Introduction to Forensic Genetics – Second Edition. Wiley-Blackwell. United Kingdom. Hashiguchi K, Bohr VA, de Souza-Pinto NC (2004) Oxidative stress and mitochondrial DNA repair: implications for NRTIs induced DNA damage. Mitochondrion. 4(2-3):21522. Holsinger KE, Weir BS (2009) Genetics in geographically structured populations: defining, estimating and interpreting F(ST). Nat Rev Genet 10(9):639-50. Irwin J, Parson W, Coble M, Just R (2011) mtGenome reference population databases and the future of forensic mtDNA analysis. Forensic Sci Int Genet 5: 222-225. Irwin JA, Saunier JL, Niederstätter H, Strouss KM, Sturk KA, Diegoli TM, Brandstätter A, Parson W, Parsons TJ (2009) Investigation of heteroplasmy in the human mitochondrial DNA control region: a synthesis of observations from more than 5000 global population samples. J Mol Evol 68(5):516-27. Kloss-Brandstätter A, Pacher D, Schoenherr S, Weissensteiner H, Binna R, Specht G, Kronenberg F (2010) HaploGrep: a fast and reliable algorithm for automatic FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 84 classification of mitochondrial DNA haplogroups http://www.haplogrep.uibk.ac.at doi: 10.1002/humu.21382. Lareu MV, Phillips CP, Carracedo A, Lincoln AJ, Court DS, Thomson JA (1994) Investigation of the STR locus HUMTH01 using PCR and two electrophoresis formats: UK and Galician Caucasian population surveys and usefulness in paternity investigations, Forensic Sci Int 66: 41–52. Lee HY, Chung U, Yoo JE, Park MJ, Shin KJ (2004) Quantitative and qualitative profiling of mitochondrial DNA length heteroplasmy. Electrophoresis 25(1):28-34. Lee HY, Song I, Ha E, Cho SB, Yang WI, Shin KJ (2008) mtDNAmanager: a Webbased tool for the management and quality analysis of mitochondrial DNA controlregion sequences. BMC Bioinformatics. 17;9:483. Lello J & Lello E (1980) Dicionário Enciclopédico Luso-Brasileiro em 2 Volumes – Volume primeiro (Entrada “Foro”). Lello & Irmão Editores. Porto. Macaulay V, Hill C, Achilli A, Rengo C, Clarke D, Meehan W, Blackburn J, Semino O, Scozzari R, Cruciani F, Taha A, Shaari NK, Raja JM, Ismail P, Zainuddin Z, Goodwin W, Bulbeck D, Bandelt HJ, Oppenheimer S, Torroni A, Richards M (2005) Single, rapid coastal settlement of Asia revealed by analysis of complete mitochondrial genomes. Science. 13: 308(5724):1034-6. Meyer S, Weiss G, von Haeseler A (1999) Pattern of nucleotide substitution and rate heterogeneity in the hypervariable regions I and II of human mtDNA. Genetics. 152(3):1103-10. Mishmar D, Ruiz-Pesini E, Golik P, Macaulay V, Clark AG, Hosseini S, Brandon M, Easley K, Chen E, Brown MD, Sukernik RI, Olckers A, Wallace DC (2003) Natural selection shaped regional mtDNA variation in humans. Proceedings of the National Academy of Sciences of USA. 100(1):171-6. Mitchell P (1961) Coupling of phosphorylation to electron and hydrogen transfer by a chemi-osmotic type of mechanism. Nature 191:144-8. Pakendorf B, Stoneking M (2005) Mitochondrial DNA and human evolution. Annu Rev Genomics Hum Genet 6:165-83. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 85 Parson W & Bandelt HJ (2007) Extended guidelines for mtDNA typing of population data in forensic science. Forensic Sci Int Genet 1: 13-19. Parson W & Dϋr A (2007) EMPOP – A forensic mtDNA database. Forensic Sci Int Genet 1: 88-92. Parson W, Brandstätter A, Alonso A, Brandt N, Brinkmann B, Carracedo A, Corach D, Froment O, Furac I, Grzybowski T, Hedberg K, Keyser-Tracqui C, Kupiec T, LutzBonengel S, Mevag B, Ploski R, Schmitter H, Schneider P, Syndercombe-Court D, Sørensen E, Thew H, Tully G, Scheithauer R (2004) The EDNAP mitochondrial DNA population database (EMPOP) collaborative exercises: organization, results and perspectives. Forensic Sci Int 139(2-3):215-26. Parsons TJ, Muniec DS, Sullivan K, Woodyatt N, Alliston-Greiner R, Wilson MR, Berry DL, Holland KA, Weedn VW, Gill P, Holland MM (1997) A high observed substitution rate in the human mitochondrial DNA control region. Nat Genet 15(4):363-8. Pereira L, Prata MJ, Amorim A (2000) Diversity of mtDNA lineages in Portugal: not a genetic edge of European variation. Ann Hum Genet 64(Pt 6):491-506. Pereira L, Richards M, Goios A, Alonso A, Albarrán C, Garcia O, Behar DM, Gölge M, Hatina J, Al-Gazali L, Bradley DG, Macaulay V, Amorim A (2005) High-resolution mtDNA evidence for the late-glacial resettlement of Europe from an Iberian refugium. Genome Res 15(1):19-24. Pereira L, Richards M, Goios A, Alonso A, Albarrán C, Garcia O, Behar DM, Gölge M, Hatina J, Al-Gazali L, Bradley DG, Macaulay V, Amorim A (2006) Evaluating the forensic informativeness of mtDNA haplogroup H sub-typing on a Eurasian scale. Forensic Sci Int 159(1):43-50. Prieto L, Zimmermann B, Goios A, Rodriguez-Monge A, Paneto GG, Alves C, Alonso A, Fridman C, Cardoso S, Lima G, Anjos MJ, Whittle MR, Montesino M, Cicarelli RM, Rocha AM, Albarrán C, de Pancorbo MM, Pinheiro MF, Carvalho M, Sumita DR, Parson W (2011) The GHEP-EMPOP collaboration on mtDNA population data--A new resource for forensic casework. Forensic Sci Int Genet 5: 146-51. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 86 Quintáns B, Alvarez-Iglesias V, Salas A, Phillips C, Lareu MV, Carracedo A (2004) Typing of mitochondrial DNA coding region SNPs of forensic and anthropological interest using SNaPshot minisequencing. Forensic Sci Int 10;140(2-3):251-7. Rasmussen EM, Sørensen E, Eriksen B, Larsen HJ, Morling N (2002) Sequencing strategy of mitochondrial HV1 and HV2 DNA with length heteroplasmy. Forensic Sci Int 129(3):209-13. Richards M, Macaulay V, Bandelt HJ, Sykes B (1998) Phylogeography of mitochondrial DNA in western Europe. Ann Hum Genet 62:241–246. Richards M, Macaulay V, Hickey E, Vega E, Sykes B, Guida V, Rengo C, Sellitto D, Cruciani F, Kivisild T, Villems R, Thomas M, Rychkov S, Rychkov O, Rychkov Y, Gölge M, Dimitrov D, Hill E, Bradley D, Romano V, Calì F, Vona G, Demaine A, Papiha S, Triantaphyllidis C, Stefanescu G, Hatina J, Belledi M, Di Rienzo A, Novelletto A, Oppenheim A, Nørby S, Al-Zaheri N, Santachiara-Benerecetti S, Scozari R, Torroni A, Bandelt HJ (2000) Tracing European Founder Lineages in the Near Eastern mtDNA Pool. Am J Hum Genet 67:1251–1276. Richards M, Macaulay V, Torroni A, Bandelt HJ (2002) In search of geographical patterns in European mitochondrial DNA. Am J Hum Genet 71:1168–1174. Roostalu U, Kutuev I, Loogväli EL, Metspalu E, Tambets K, Reidla M, Khusnutdinova EK, Usanga E, Kivisild T, Villems R (2007) Origin and expansion of haplogroup H, the dominant human mitochondrial DNA lineage in West Eurasia: the Near Eastern and Caucasian perspective. Mol Biol Evol 24(2):436-48. Saks MJ & Koehler JJ (2005) The Coming Paradigm Shift in Forensic Identification Science. Science 309: 892. Available at SSRN: http://ssrn.com/abstract=962968. Salas A, Comas D, Lareu MV, Bertranpetit J, Carracedo A (1998) mtDNA analysis of the Galician population: a genetic edge of European variation. Eur J Hum Genet 6(4):365-75. Saraste M (1999) Oxidative phosphorylation at the fin de siècle. Science 283(5407):1488-93. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 87 Satoh M & Kuroiwa T (1991) Organization of multiple nucleoids and DNA molecules in mitochondria of a human cell. Exp Cell Res 196: 137–140. Scheffler, IE (1999) Mitochondria. Wiley-Liss. New York. Schwartz M, Vissing J (2004) No evidence for paternal inheritance of mtDNA in patients with sporadic mtDNA mutations. J Neurol Sci 218(1-2):99-101. Shriver MD & Kittles RA (2004) Genetic ancestry and the search for personalized genetic histories. Nat Rev Genet 5: 611-618. Simoni L, Calafell F, Pettener D, Bertranpetit J, Barbujani G (2000) Geographic Patterns of mtDNA Diversity in Europe. Am J Hum Genet 66(1): 262–278. Tamura K & Nei M (1993) Estimation of the number of nucleotide substitutions in the control region of mitochondrial DNA in humans and chimpanzees. Mol Biol Evol 10:512-526. Taylor RW, McDonnell MT, Blakely EL, Chinnery PF, Taylor GA, Howell N, Zeviani M, Briem E, Carrara F, Turnbull DM (2003) Genotypes from patients indicate no paternal mitochondrial DNA contribution. Ann Neurol 54(4):521-4. Torroni A, Achilli A, Macaulay V, Richards M, Bandelt HJ (2006) Harvesting the fruit of the human mtDNA tree. Trends Genet 22(6):339-45. Torroni A, Bandelt H-J, D’Urbano L, Lahermo P, Moral P, Sellitto D, Rengo C, Forster P, Savantaus M-L, Bonné-Tamir B, Scozzari R (1998) mtDNA analysis reveals a major late Paleolithic population expansion from southwestern to northeastern Europe. Am J Hum Genet 62:1137–1152. Torroni A, Huoponen K, Francalacci P, Petrozzi M, Morelli L, Scozzari R, Obinu D, Savontaus ML, Wallace DC (1996) Classification of European mtDNAs From an Analysis of Three European Populations. Genetics 144:1835–1850. Tully G, Bär W, Brinkman B, Carracedo A, Gill P, Morling N, Parson W, Schneider P (2001) Considerations by the European DNA profiling (EDNAP) group on the working practices, nomenclature and interpretation of mitochondrial DNA profiles. Forensic Sci Int 124: 83-91. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 88 Tully G, Barritt SM, Bender K, Brignon E, Capelli C, Dimo-Simonin N, Eichmann C, Ernst CM, Lambert C, Lareu MV, Ludes B, Mevag B, Parson W, Pfeiffer H, Salas A, Schneider PM, Staalstrom E (2004) Results of a collaborative study of the EDNAP group regarding mitochondrial DNA heteroplasmy and segregation in hair shafts. Forensic Sci Int 140(1):1-11. Underhill PA, Kivisild T (2007) Use of Y chromosome and mitochondrial DNA population structure in tracing human migrations. Annu Rev Genet 41:539-64. van Oven M & Kayser M (2009) Updated comprehensive phylogenetic tree of global human mitochondrial DNA variation. Hum Mutat 30(2):E386-E394. http://www.phylotree.org. doi:10.1002/humu.20921. Werret DJ (1997) The national DNA database. Forensic Sci Int 88: 33-42. Wilson MR, Allard MW, Monson KL, Miller KWP, Budowle B (2002a) Recommendations for consistent treatment of length variants in the human mtDNA control region. Forensic Sci Int 129:35–42. Wilson MR, Allard MW, Monson KL, Miller KWP, Budowle B (2002b) Further discussion of the consistent treatment of length variants in the human mitochondrial DNA control region. Forensic Science Community 4:#4. Wright, S (1951) The genetic structure of populations. Annals Eugenics 15:323-354. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 89 User Guides Applied Biosystems (2004a) Sequencing Analysis Software Version 5.2 – Quick Reference Card. Part Number 4359446 Rev. B. Applied Biosystems (2004b) SeqScape Software version 2.5 User Guide. Part Number 4337035 Rev. B Applied Biosystems (2005) GeneMapper® Software Version 4.0 - SNaPshot® Kit Analysis Getting Started Guide. Part Number 4363078 Rev. B Applied Biosystems (2010) BigDye® Terminator v3.1 Cycle Sequencing Kit Protocol. Part Number 4337035 Rev. B. Applied Biosystems (2010) dGTP BigDye™ Terminator v3.0 Ready Reaction Cycle Sequencing Kit Protocol. Part Number 4390038 Rev. D. QIAGEN (2010) QIAGEN® Multiplex PCR Handbook. Part Number 1064922. QIAGEN (2011) QIAxcel® DNA Handbook – Second Edition. Part Number 1066467. FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 96 North Z6675 K1a1a or K1a3a 73G 263G 309.1C 315.1C 497T 16093C 16224C 16290T 16311C 16519C North Z6668 K1a1a or K1a3a 73G 195C 263G 309.1C 315.1C 497T 524.1A 524.2C 16093C 16224C 16311C 16519C North Z6553 K1a4 73G 195C 263G 309.1C 315.1C 497T 524.1A 524.2C 16224C 16519C North Z6342 W73G 119C 189G 195C 204C 263G 315.1C 16223T 16292T 16295T 16519C North Z6589 W73G 189G 195C 204C 207A 263G 315.1C 16223T 16292T 16311C 16519C North Z2577 W73G 189G 195C 198T 204C 207A 263G 309.1C 315.1C 16223T 16292T 16346A 16519C North Z6658 L3 73G 146C 150T 263G 315.1C 398C 523d 524d 16041G 16223T 16519C North Z2113 L3e2b 73G 150T 195C 263G 309.1C 315.1C 16172C 16183C 16189C 16223T 16320T 16519C North Z1544 L3f1b 73G 150T 189G 200G 263G 315.1C 16209C 16218T 16223T 16256T 16292T 16311C 16519C North Z6627 R0a 58C 60.1T 64T 93G 263G 309.1C 315.1C 16093C 16126C 16362C North Z6659 R0a1 60.1T 64T 263G 309.1C 315.1C 451G 16126C 16362C North Z6685 X1b 73G 146C 153G 256T 263G 309.1C 309.2C 315.1C 16182C 16183C 16189C 16223T 16278T 16519C North Z6477 X2b 73G 93G 95C 153G 188G 195C 225A 226C 263G 309.1C 315.1C 16183C 16189C 16223T 16278T 16519C North Z6687 B? 73G 263G 309.1C 315.1C 385G 16183C 16189C 16295Y 16519C North Z6617 B4b 73G 152C 200G 263G 309.1C 309.2C 315.1C 499A 16183C 16189C 16217C 16297C 16519C North Z6679 I3 73G 152C 199C 204C 207A 250C 263G 309.1C 315.1C 573.1C 573.2C 573.3C 573.4C 573.5C 16129A 16223T 16391A 16519C North Z5102 R73G 263G 315.1C 16260T 16519C North Z6753 W73G 189G 195C 204C 207A 228A 263G 315.1C 16223T 16292T 16311C 16519C Center Z6557 HH263G 315.1C 750G 1438G 4769G Center Z2476 HH93G 263G 315.1C 750G 1438G 4769G 16519C Center Z6663 HH263G 315.1C 750G 1438G 4769G 16293G 16519C Center Z2382 HH263G 309.1C 315.1C 750G 1438G 4769G 16192T 16519C Center Z6049 HH1 263G 315.1C 750G 3010A 4769G 16189C Center Z2622 HH1 263G 309.1C 315.1C 750G 3010A 4769G 16519C Center Z6166 HH1 263G 315.1C 523d 524d 750G 3010A 4769G 16519C Center Z2473 HH1 152C 263G 315.1C 329A 750G 3010A 4769G 16519C Center Z2625 HH1 263G 309.1C 315.1C 569T 750G 3010A 4769G 16519C Center Z6570 HH1 150T 263G 309.1C 315.1C 750G 3010A 4769G 16519C Center Z2624 HH1 263G 309.1C 315.1C 750G 3010A 4769G 16129A 16325C 16519C Center Z2628 HH1 263G 309.1C 315.1C 750G 3010A 4769G 16093C 16311C 16519C Center Z2633 HH1 150T 263G 309.1C 315.1C 750G 3010A 4769G 16192T 16519C Center Z2642 HH1 263G 309.1C 315.1C 376G 750G 3010A 4769G 16218T 16519C Center Z5919 HH1 263G 309.1C 315.1C 524.1A 524.2C 750G 3010A 4769G 16519C Center Z6715 HH1 150T 263G 309.1C 309.2C 315.1C 750G 3010A 4769G 16519C Center Z2470 HH1 152C 195C 263G 315.1C 750G 3010A 4769G 16183M 16189C 16193.1C 16519C Center Z2367 HH1-709 152C 263G 309.1C 315.1C 709G 750G 3010A 16519C Center Z2474 H1c H1-709 263G 309.1C 309.2C 315.1C 477C 709G 750G 3010A 4769G 16519C Center Z6639 H1c H1-8473 257G 263G 309.1C 315.1C 477C 750G 3010A 4769G 8473C 16294T 16519C Center Z2464 HH3 263G 315.1C 750G 4769G 6776C 16519C Center Z2319 HH3 263G 309.1C 309.2C 315.1C 750G 4769G 6776C 16519C Center Z2395 HH3 152C 263G 309.1C 315.1C 750G 4769G 6776C 16519C Center Z6236 HH3 10C 263G 309.1C 309.2C 315.1C 750G 4769G 6776C 16129A 16519C Center Z6291 HH3 152C 195C 263G 309.1C 315.1C 750G 4769G 6776C 16126C 16189C 16193.1C 16519C Center Z2465 HH3c 263G 315.1C 573.1C 750G 6776C 12957C 16519C Center Z2797 HH3c 263G 315.1C 573.1C 750G 6776C 12957C 16519C Center Z4718 HH3c 263G 315.1C 573.1C 750G 4769G 6776C 12957C 16185T 16519C Region Sample Via EMPOP Via web tools* H Sub-hg Coding Region Mutations for Haplogroup H mtDNA sequences and Control Region SNPs FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 97 Center Z6656 HH4a1a1 73G 263G 315.1C 523d 524d 750G 3992T 4769G 8269A 10044G 14365T Center Z6549 H5 H5 263G 315.1C 373G 456T 4769G 16304C 16327T Center Z6276 H5 H5a 263G 309.1C 315.1C 456T 1438G 4336C 4769G 16304C Center Z2643 H5 H5a 263G 309.1C 315.1C 456T 1438G 4336C 4769G 16304C 16519C Center Z6544 H5 H5a1 263G 315.1C 456T 523d 524d 1438G 4336C 4769G 15833T 16304C Center Z2475 H5 H5a1 263G 315.1C 456T 523d 524d 1438G 4336C 4769G 15833T 16304C Center Z6688 H1b H7 263G 315.1C 750G 4769G 4793G 16356C 16519C Center Z2369 HH7 263G 315.1C 750G 4769G 4793G 16153A 16519C Center Z4774 T1a 73G 263G 309.1C 315.1C 16126C 16163G 16186T 16189C 16249C 16294T 16311C 16519C Center Z6340 T1a 73G 263G 309.1C 309.2C 315.1C 16126C 16163G 16186T 16189C 16193d 16294T 16519C Center Z6708 T1a 73G 152C 195C 263G 309.1C 315.1C 16126C 16163G 16186T 16189C 16294T 16519C Center Z163 T2b 73G 263G 315.1C 372C 16126C 16294T 16296T 16304C 16519C Center Z6580 T2b 73G 152C 263G 309.1C 315.1C 16126C 16294T 16304C 16519C Center Z2630 T2b 73G 151T 263G 309.1C 315.1C 16126C 16243C 16294T 16296T 16304C Center Z6167 T2b 73G 151T 215G 263G 309.1C 315.1C 16126C 16294T 16304C 16465T 16519C Center Z6664 T2b 73G 151T 263G 315.1C 523d 524d 16126C 16294T 16296T 16304C 16519C Center Z2627 T2b 73G 204C 263G 309.1C 315.1C 16051G 16126C 16294T 16296T 16304C 16519C 16527T Center Z6445 T2c 73G 146C 152C 263G 279C 309.1C 315.1C 16126C 16292T 16294T 16519C Center Z2394 U2e 73G 152C 217C 263G 309.1C 315.1C 340T 508G 523d 524d 16051G 16129C 16189C 16362C 16519C Center Z5987 U2e 73G 152C 217C 263G 309.1C 309.2C 315.1C 340T 508G 524.1A 524.2C 16051G 16129C 16183C 16189C 16193.1C 16327T 16362C 16519C Center Z6734 U4 73G 195C 263G 315.1C 499A 16356C 16519C Center Z2631 U4a1 73G 152C 195C 198T 263G 309.1C 315.1C 499A 16134T 16356C 16519C Center Z6638 U4b2 73G 195C 263G 309.1C 315.1C 499A 524.1A 524.2C 16136C 16356C 16519C Center Z2478 U5b 73G 150T 263G 315.1C 16270T Center Z2621 U5b 73G 150T 263G 315.1C 16270T 16292T 16362C Center Z6458 U5b 73G 150T 263G 309.1C 315.1C 16192T 16270T Center Z2626 U5b1b 73G 146C 150T 152C 263G 315.1C 523d 524d 16145A 16189C 16193.1C 16270T 16465T Center Z2639 U6a 73G 263G 309.1C 315.1C 16145A 16172C 16219G 16235G 16278T 16362Y 16519C Center Z2480 J1b1 73G 242T 263G 295T 309.1C 315.1C 462T 489C 16069T 16126C 16145A 16172C 16222T 16261T Center Z5738 J1c 73G 185A 263G 295T 315.1C 462T 489C 16069T 16126C Center Z2391 J1c 73G 185A 189G 228A 263G 295T 315.1C 462T 489C 523d 524d 16069T 16126C Center Z6132 J1c2 73G 185A 188G 199C 263G 295T 315.1C 462T 489C 16069T 16126C Center Z5939 J1c2 73G 146C 185A 188G 203A 222T 228A 263G 295T 315.1C 462T 489C 16069T 16126C 16519C Center Z2481 J1d 73G 150T 152C 263G 295T 315.1C 489C 16069T 16093C 16126C 16193T Center Z6473 J2b 73G 150T 152C 263G 295T 309.1C 315.1C 489C 523d 524d 16069T 16126C 16193T 16319A 16360T Center Z6568 J2b 73G 150T 152C 263G 295T 309.1C 315.1C 489C 523d 524d 16069T 16126C 16193T 16319A 16360T Center Z6718 J2b 73G 150T 152C 263G 295T 309.1C 315.1C 489C 523d 524d 16069T 16126C 16193T 16319A 16360T Center Z6486 K1a 73G 114T 263G 315.1C 497T 16192T 16224C 16311C 16519C Center Z6202 K1a 73G 263G 309.1C 315.1C 497T 16224C 16234T 16311C 16356C 16519C Center Z6712 K1a 73G 263G 309.1C 315.1C 497T 16168T 16224C 16274A 16311C 16519C Center Z6713 K1a 73G 195C 263G 315.1C 497T 524.1A 524.2C 524.3A 524.4C 16224C 16311C 16519C Center Z6547 K1a 73G 150T 195C 263G 309.1C 309.2C 315.1C 497T 524.1A 524.2C 16093C 16224C 16311C 16519C Center Z2629 K1a1a 73G 263G 315.1C 497T 16093C 16224C 16311C 16519C Center Z6633 K1c 73G 146C 152C 263G 315.1C 498d 16192T 16224C 16311C 16519C Center Z2640 K1c 73G 146C 152C 263G 315.1C 498d 523d 524d 16224C 16311C 16519C Center Z6520 K1c 73G 146C 152C 263G 315.1C 498d 523d 524d 16224C 16311C 16519C Region Sample Via EMPOP Via web tools* H Sub-hg Coding Region Mutations for Haplogroup H mtDNA sequences and Control Region SNPs FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 98 Center Z4477 L1 73G 152C 182T 185C 195C 247A 263G 315.1C 357G 523d 524d 16126C 16145A 16187T 16189C 16223T 16264T 16270T 16278T 16293G 16311C 16519C Center Z6596 L2a1 73G 143A 146C 152C 195C 263G 315.1C 523d 524d 16111T 16223T 16278T 16294T 16309G 16390A Center Z6595 L3b1b 73G 151T 152C 263G 315.1C 523d 524d 16124C 16223T 16278T 16362C 16519C Center Z6733 L3d4 73G 152C 189G 195C 207A 263G 315.1C 523d 524d 16124C 16223T 16519C Center Z6727 HV 57.1C 64T 151T 152C 263G 309.1C 315.1C 16114T 16311C 16362C 16519C 16526A Center Z6526 HV0 72C 263G 309.1C 315.1C 16298C Center Z6269 HV0 72C 195C 263G 309.1C 315.1C 16298C 16355T Center Z6348 HV0 72C 195C 263G 315.1C 524.1A 524.2C 16129A 16298C Center Z2469 HV4a 263G 309.1C 315.1C 16221T 16519C Center Z2453 X2 73G 195C 225A 263G 309.1C 315.1C 16183M 16189C 16193.1C? 16223T 16278T 16519C Center Z2466 X2 73G 195C 225A 263G 309.1C 315.1C 16183M 16189C 16193.1C 16223T 16278T 16519C Center Z6598 X2 73G 143A 153G 189G 195C 198T 225A 226C 263G 309.1C 309.2C 315.1C 475G 16183C 16189C 16193.1C 16220T 16278T 16311C 16519C Center Z2458 X3a 73G 146C 153G 256T 263G 309.1C 309.2C 315.1C 16183C 16189C 16192T 16223T 16278T 16519C Center Z6201 W73G 143A 185A 189G 195C 204C 207A 263G 309.1C 315.1C 16223T 16292T 16311C 16320T 16519C Center Z6515 W1 73G 119C 189G 195C 204C 207A 214G 227G 263G 309.1C 315.1C 16223T 16292T 16519C Center Z2385 W4 73G 143A 189G 194T 195C 204C 207A 263G 315.1C 16223T 16292T 16519C Center Z6750 M1 73G 263G 315.1C 489C 16129A 16183C 16189C 16193.1C 16249C 16311C 16519C Center Z2479 N1b 73G 106d 107d 108d 109d 110d 111d 114T 152C 195C 263G 315.1C 523d 524d 16145A 16176G 16223T 16352C 16390A Center Z6546 G2a? 73G 263G 315.1C 523d 524d 16223T 16278T 16362C 16519C Center Z2455 D4g1 73G 263G 315.1C 523d 524d 16223T 16278T 16362C 16519C South SC60 H2a2b H2a2b1 263G 309.1C 315.1C 16235G 16291T South SC16 H2a2b H2a2b1 93G 263G 309.1C 315.1C 16235G 16291T South SC24 H2a2b H2a2b1 93G 263G 309.1C 315.1C 16235G 16291T South SC59 H2a2b H2a2b1 93G 263G 309.1C 315.1C 16235G 16291T South Z2311 H Missing datab263G 309.1C 315.1C 16362C 16519C South Z2314 HH146C 195C 263G 315.1C 750G 1438G 4769G 16519C South Z6585 H2a3 H263G 309.1C 315.1C 750G 1438G 4769G 16274A 16519C South SC61 HH1 263G 315.1C 750G 3010A 4769G 16519C South SC22 HH1 263G 315.1C 750G 3010A 4769G 16519C South SC47 HH1 263G 315.1C 750G 3010A 4769G 16519C South SC15 H2a2 H1 263G 315.1C 523d 524d 750G 3010A 4769G 16519C South Z3583 HH1 263G 309.1C 309.2C 315.1C 750G 3010A 4769G 16519C South SC66 HH1 263G 315.1C 750G 3010A 4769G 16271C 16309G 16519C South Z2515 HH1 263G 309.1C 315.1C 524.1A 524.2C 750G 3010A 4769G 16519C South SC31 H1a H1a 73G 263G 315.1C 534T 750G 3010A 4769G 16051G 16162G 16519C South SC40 H1a1 H1a1 73G 263G 315.1C 750G 3010A 6365C 16162G 16209C 16519C South SC23 H1a1 H1a1 73G 263G 315.1C 750G 3010A 6365C 16162G 16209C 16519C South SC34 H1a1 H1a1 73G 263G 315.1C 750G 3010A 6365C 16162G 16209C 16519C South Z2303 H1b Missing datab263G 310C 313d 314d 315d 16183C 16189C 16193.1C 16356C 16519C South SC36 H1b H1b 64T 263G 315.1C 750G 3010A 3796G 16189C 16193.1C 16290T 16291T 16356C 16362C 16519C South SC38 HH3 263G 315.1C 750G 4769G 6776C 16519C South SC04 HH3 263G 309.1C 315.1C 750G 4769G 6776C 16519C South Z2613 HH3 152C 188G 263G 315.1C 750G 4769G 6776C 16519C South SC19 HH3 263G 315.1C 524.1A 524.2C 750G 4769G 6776C 16519C South SC54 HH3c 263G 315.1C 573.1C 573.2C 750G 4769G 6776C 12957C 16519C South Z2306 HH3c 263G 315.1C 573.1C 573.2C 573.3C 750G 6776C 12957C 16189Y 16519C Region Sample Via EMPOP Via web tools* H Sub-hg Coding Region Mutations for Haplogroup H mtDNA sequences and Control Region SNPs FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 99 South SC03 H3d H4a1a1 73G 263G 309.1C 315.1C 521d 522d 523d 524d 750G 3992T 4769G 8269G 10044G 14365T South Z5304 H5 Missing datab146C 263G 309.1C 309.2C 315.1C 456T 16304C South SC08 H5 H5a 263G 309.1C 315.1C 456T 1438G 4336C 4769G 16295T 16304C South Z5627 H6 H6a1 239C 263G 309.1C 309.2C 315.1C 3915A 4727G 4769G 16274A 16362C 16482G South Z2316 H1g H7 146C 263G 309.1C 315.1C 750G 4769G 4793G 16124C South Z3973 HH7 263G 309.1C 315.1C 750G 4769G 4793G 16213A 16519C South SC42 H10a H10 263G 309.1C 315.1C 750G 14470A 16114T 16344T 16519C South SC13 U2d 73G 152C 199C 309.1C 309.2C 315.1C 471C 524.1A 524.2C 16051G 16183C 16189C 16193.1C 16234T 16266T 16294T 16525G South Z4823 U2e1 73G 114T 152C 217C 263G 309.1C 309.2C 315.1C 340T 508G 16051G 16129C 16183C 16189C 16193.1C 16362C 16519C South Z6356 U3 73G 150T 263G 315.1C 16343G South Z6241 U5a1 73G 263G 315.1C 16256T 16270T 16399G South Z2318 U5a1 73G 263G 315.1C 16256T 16270T 16399G South SC20 U5a1 73G 263G 315.1C 16192T 16256T 16270T 16362C 16399G South SC55 U5a1 73G 263G 315.1C 16192T 16256T 16270T 16362C 16399G South SC37 U5b 73G 150T 263G 315.1C 533G 16192T 16270T South SC45 U5b 73G 150T 263G 315.1C 533G 16192T 16270T South Z2290 U5b 73G 150T 263G 309.1C 315.1C 16093C 16189C 16191.1C 16192T 16270T 16311C South SC26 U5b1b 73G 150T 263G 315.1C 16189C 16270T South SC32 U5b1b 73G 150T 263G 315.1C 16189C 16270T South Z2507 U5b1b 73G 150T 152C 195C 263G 315.1C 16093C 16183C 16187T 16189C 16192T 16270T South Z2296 U5b1c 73G 150T 263G 309.1C 315.1C 16093C 16189C 16270T 16311C South Z2505 U6a1b 73G 146C 263G 309.1C 315.1C 16172C 16219G 16235G 16278T 16355T 16519C South Z2508 U6a1b 73G 146C 263G 309.1C 315.1C 16172C 16219G 16235G 16278T 16355T 16519C South Z2605 I? 73G 199C 250C 263G 291.1A 309.1C 315.1C 573.1C 573.2C 573.3C 573.4C 573.5C 16140C 16223T 16391A 16519C South Z2300 I73G 93G 150T 199C 204C 250C 263G 315.1C 573.1C 573.2C 573.3C 574.4C 16129A 16223T 16391A 16497G 16519C South Z2609 I1a 73G 199C 203A 204C 250C 263G 315.1C 455.1T 573.1C 573.2C 573.3C 573.4C 16129A 16172C 16223T 16311C 16391A 16519C South SC46 I1a 73G 199C 203A 204C 250C 263G 315.1C 455.1T 573.1C 573.2C 573.3C 573.4C 16129A 16172C 16223T 16311C 16391A 16519C South SC39 I1a 73G 199C 203A 204C 250C 263G 315.1C 455.1T 573.1C 573.2C 573.3C 573.4C 573.5C 16129A 16172C 16223T 16311C 16391A 16519C South SC44 I1a 73G 199C 203A 204C 250C 263G 309.1C 315.1C 455.1T 573.1C 573.2C 573.3C 573.4C 16129A 16172C 16223T 16311C 16391A 16519C South SC05 I1a1 73G 199C 203A 204C 250C 263G 309.1C 315.1C 455.1T 573.1C 573.2C 573.3C 16129A 16172C 16223T 16311C 16390R 16391A 16519C South SC06 I1a1 73G 199C 203A 204C 250C 263G 309.1C 315.1C 455.1T 573.1C 573.2C 573.3C 573.4C 16129A 16172C 16223T 16311C 16391A 16519C South SC21 I1a1 73G 199C 203A 204C 250C 263G 309.1C 315.1C 455.1T 573.1C 573.2C 573.3C 573.4C 16129A 16172C 16223T 16311C 16391A 16519C South SC10 I1a1 73G 199C 203A 204C 250C 263G 309.1C 315.1C 455.1T 573.1C 573.2C 573.3C 573.4C 573.5C 16129A 16172C 16223T 16311C 16391A 16519C South SC62 T1 73G 263G 309.1C 315.1C 16126C 16163G 16186T 16189C 16249C 16294T 16311C 16519C South SC17 T1a 73G 195C 263G 309.1C 309.2C 315.1C 16126C 16163G 16186T 16189C 16294T 16519C South SC27 T1a 73G 263G 309.1C 315.1C 16126C 16153A 16163G 16186T 16189C 16249C 16294T 16311C 16519C South SC29 T1a 73G 152C 195C 263G 309.1C 315.1C 523d 524d 16126C 16163G 16186T 16189C 16294T 16519C South Z4723 T1a 73G 263G 309.1C 309.2C 315.1C 524.1A 524.2C 16126C 16163G 16186T 16189C 16284G 16294T 16519C South Z2292 T1a2 73G 152C 263G 309.1C 315.1C 384G 16126C 16163G 16186T 16189C 16261T 16294T 16519C South Z6242 T1a2 73G 152C 263G 309.1C 315.1C 384G 460Y 16126C 16163G 16186T 16189C 16261T 16294T 16519C South SC48 T2b 73G 263G 309.1C 315.1C 573.1C 16126C 16294T 16296T 16304C 16519C South SC07 T2c 73G 263G 315.1C 16126C 16171G 16292T 16294T 16324C 16519C Region Sample Via EMPOP Via web tools* H Sub-hg Coding Region Mutations for Haplogroup H mtDNA sequences and Control Region SNPs FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 100 South SC35 HV 152C 263G 309.1C 315.1C 16311C South SC02 HV0 72C 195C 263G 315.1C 16298C 16344T South SC57 HV0 72C 195C 263G 315.1C 16298C 16344T South Z2611 HV0 72C 195C 263G 315.1C 16153A 16298C South Z2614 HV0 72C 263G 309.1C 309.2C 315.1C 16298C South Z6099 HV0 72C 195C 263G 315.1C 16153A 16298C South Z2612 HV0 72C 195C 263G 309.1C 315.1C 16298C 16355T South SC41 HV0 72C 195C 198T 263G 309.1C 315.1C 16298C South SC18 J1b 73G 152C 263G 295T 315.1C 462T 489C 16069T 16126C 16145A 16222T 16261T South SC58 J1b1 73G 146C 242T 263G 295T 315.1C 462T 489C 16069T 16093C 16126C 16145A 16172C 16261T South SC49 J1c 73G 185A 228A 263G 295T 309.1C 315.1C 462T 489C 16069T 16126C South SC11 J1c 73G 185A 228A 263G 295T 309.1C 315.1C 462T 489C 16069T 16126C South SC56 J1c 73G 185A 228A 263G 295T 309.1C 315.1C 462T 489C 16069T 16126C South SC51 J1c1 73G 185A 228R 263G 295T 315.1C 462T 482C 489C 16069T 16126C 16153A 16311C South SC30 J1c2 73G 152C 185A 188G 228A 263G 295T 315.1C 462T 489C 523d 524d 16069T 16126C 16519C 64T 93G 95C 152C 189G 207A 236C 247A 263G 309.1C 315.1C 523d 524d 571T 16093C 16148T 16172C 16187T 16188G 16189C 16223T 16230G 16311C 16320T 16519C South SC52 L1b1 73G 152C 182T 185T 195C 198T 247A 263G 315.1C 357G 523d 524d 16126C 16187T 16189C 16223T 16264T 16270T 16278T 16293G 16311C 16400T 16519C South Z2936 L1b1 73G 152C 182T 185T 189G 195C 247A 263G 315.1C 357G 523d 524d 16126C 16187T 16189C 16223T 16261T 16264T 16270T 16278T 16293G 16311C 16519C South Z2500 L1c1? 73G 152C 182T 185T 195C 247A 263G 315.1C 357G 523d 524d 16104T 16187T 16189C 16223T 16270T 16278T 16289G 16293G 16311C 16519C South SC12 L2a1 73G 143A 146C 152C 195C 198T 263G 315.1C 16189C 16192T 16223T 16278T 16294T 16309G 16390A South Z2618 L3f2 73G 191T 263G 315.1C 489C 16129A 16209C 16223T 16519C South Z2317 L3f2? 73G 146C 150T 189G 200G 263G 309.1C 315.1C 16209C 16223T 16311C 16519C South Z3788 X73G 263G 309.1C 309.2C 315.1C 523d 524d 16172C 16183C 16189C 16193.1C 16219G 16239T 16278T South Z2616 X2b 73G 153G 195C 225A 226C 263G 309.1C 315.1C 16189C 16223T 16278T 16519C South SC09 X2b 73G 153G 195C 225A 226C 263G 309.1C 309.2C 315.1C 16182C 16183C 16189C 16223T 16278T 16519C South SC64 X2b 73G 153G 195C 225A 226C 263G 309.1C 309.2C 315.1C 16182C 16183C 16189C 16223T 16278T 16519C South SC53 X2b 73G 153G 195C 225A 226C 263G 309.1C 309.2C 315.1C 16182C 16183C 16189C 16193.1C 16223T 16278T 16519C South SC65 K1a 73G 263G 315.1C 497T 16180G 16224C 16249C 16311C 16519C South SC63 K1c2 73G 146C 152C 169G 183G 263G 309.1C 315.1C 498d 16224C 16311C 16320T 16519C South Z3789 M1b2 73G 195C 263G 309.1C 315.1C 489C 16129A 16183C 16189C 16193.1C 16223T 16249C 16311C 16399G 16519C South Z2514 R0a? 58C 60.1T 64T 263G 315.1C 16126C 16362C South SC28 W73G 189G 194T 195C 204C 207A 263G 309.1C 315.1C 523d 524d 16223T 16292T 16318G 16519C South SC50 L0a2 Region Sample Via EMPOP Via web tools* H Sub-hg Coding Region Mutations for Haplogroup H mtDNA sequences and Control Region SNPs FCUP Refined characterization of Portuguese mtDNA diversity for forensic purposes 101 Hap_1: 1 [Z1544] Hap_61: 1 [Z6650] Hap_121: 1 [Z5738] Hap_181: 3 [SC16 SC24 SC59] Hap_2: 1 [Z2094] Hap_62: 2 [Z6652 Z2642] Hap_122: 2 [Z5919 Z2515] Hap_182: 1 [SC17] Hap_3: 1 [Z2113] Hap_63: 1 [Z6654] Hap_123: 1 [Z5939] Hap_183: 1 [SC18] Hap_4: 1 [Z2121] Hap_64: 1 [Z6658] Hap_124: 1 [Z5987] Hap_184: 1 [SC19] Hap_5: 1 [Z2208] Hap_65: 1 [Z6659] Hap_125: 1 [Z6049] Hap_185: 2 [SC20 SC55] Hap_6: 1 [Z2262] Hap_66: 1 [Z6668] Hap_126: 1 [Z6132] Hap_186: 3 [SC23 SC34 SC40] Hap_7: 1 [Z2442] Hap_67: 1 [Z6673] Hap_127: 1 [Z6167] Hap_187: 2 [SC26 SC32] Hap_8: 1 [Z2577] Hap_68: 1 [Z6675] Hap_128: 1 [Z6201] Hap_188: 1 [SC27] Hap_9: 4 [Z2654 Z6538 Z2622 SC04] Hap_69: 1 [Z6676] Hap_129: 1 [Z6202] Hap_189: 1 [SC28] Hap_10: 1 [Z2675] Hap_70: 1 [Z6677] Hap_130: 1 [Z6236] Hap_190: 1 [SC29] Hap_11: 2 [Z2676 Z6713] Hap_71: 1 [Z6679] Hap_131: 2 [Z6269 Z2612] Hap_191: 1 [SC30] Hap_12: 1 [Z3102] Hap_72: 1 [Z6680] Hap_132: 1 [Z6276] Hap_192: 1 [SC31] Hap_13: 1 [Z4404] Hap_73: 1 [Z6685] Hap_133: 1 [Z6291] Hap_193: 1 [SC35] Hap_14: 1 [Z4446] Hap_74: 1 [Z6687] Hap_134: 1 [Z6340] Hap_194: 1 [SC36] Hap_15: 1 [Z4766] Hap_75: 1 [Z6689] Hap_135: 1 [Z6348] Hap_195: 2 [SC37 SC45] Hap_16: 3 [Z4995 Z2367 Z2395] Hap_76: 1 [Z6701] Hap_136: 1 [Z6445] Hap_196: 2 [SC39 SC46] Hap_17: 1 [Z5102] Hap_77: 1 [Z6710] Hap_137: 1 [Z6458] Hap_197: 1 [SC41] Hap_18: 1 [Z5117] Hap_78: 1 [Z6720] Hap_138: 1 [Z6486] Hap_198: 1 [SC42] Hap_19: 2 [Z5862 Z6526] Hap_79: 2 [Z6721 Z2316] Hap_139: 1 [Z6515] Hap_199: 1 [SC44] Hap_20: 1 [Z6005] Hap_80: 1 [Z6723] Hap_140: 1 [Z6547] Hap_200: 1 [SC48] Hap_21: 1 [Z6055] Hap_81: 1 [Z6724] Hap_141: 1 [Z6557] Hap_201: 1 [SC50] Hap_22: 1 [Z6073] Hap_82: 1 [Z6725] Hap_142: 1 [Z6570] Hap_202: 1 [SC51] Hap_23: 1 [Z6117] Hap_83: 1 [Z6728] Hap_143: 1 [Z6595] Hap_203: 1 [SC52] Hap_24: 4 [Z6237 Z6473 Z6568 Z6718] Hap_84: 1 [Z6731] Hap_144: 1 [Z6596] Hap_204: 1 [SC53] Hap_25: 1 [Z6303] Hap_85: 1 [Z6740] Hap_145: 1 [Z6598] Hap_205: 1 [SC54] Hap_26: 1 [Z6304] Hap_86: 2 [Z6741 Z2311] Hap_146: 1 [Z6633] Hap_206: 1 [SC58] Hap_27: 1 [Z6342] Hap_87: 2 [Z6743 Z6756] Hap_147: 1 [Z6638] Hap_207: 1 [SC63] Hap_28: 4 [Z6343 Z6632 Z6166 SC15] Hap_88: 1 [Z6744] Hap_148: 1 [Z6639] Hap_208: 1 [SC65] Hap_29: 6 [Z6345 Z2464 SC22 SC38 SC47 SC61] Hap_89: 1 [Z6746] Hap_149: 1 [Z6656] Hap_209: 1 [SC66] Hap_30: 4 [Z6351 Z4774 Z6754 SC62] Hap_90: 1 [Z6747] Hap_150: 1 [Z6663] Hap_210: 1 [Z2936] Hap_31: 1 [Z6362] Hap_91: 1 [Z6749] Hap_151: 1 [Z6664] Hap_211: 1 [Z3788] Hap_32: 3 [Z6380 Z2475 Z6544] Hap_92: 1 [Z6751] Hap_152: 1 [Z6688] Hap_212: 1 [Z6585] Hap_33: 1 [Z6382] Hap_93: 1 [Z6753] Hap_153: 1 [Z6708] Hap_213: 1 [Z2290] Hap_34: 1 [Z6420] Hap_94: 1 [Z6759] Hap_154: 1 [Z6712] Hap_214: 2 [Z2292 Z6242] Hap_35: 1 [Z6449] Hap_95: 1 [Z6761] Hap_155: 1 [Z6715] Hap_215: 1 [Z2296] Hap_36: 1 [Z6455] Hap_96: 1 [Z163] Hap_156: 1 [Z6727] Hap_216: 1 [Z2300] Hap_37: 1 [Z6459] Hap_97: 1 [Z2382] Hap_157: 1 [Z6733] Hap_217: 1 [Z2303] Hap_38: 1 [Z6474] Hap_98: 1 [Z2391] Hap_158: 1 [Z6734] Hap_218: 1 [Z2306] Hap_39: 1 [Z6475] Hap_99: 1 [Z2394] Hap_159: 1 [Z6750] Hap_219: 1 [Z2314] Hap_40: 1 [Z6477] Hap_100: 2 [Z2453 Z2466] Hap_160: 1 [Z2369] Hap_220: 1 [Z2317] Hap_41: 2 [Z6478 SC60] Hap_101: 2 [Z2455 Z6546] Hap_161: 1 [Z2385] Hap_221: 2 [Z2318 Z6241] Hap_42: 1 [Z6532] Hap_102: 1 [Z2458] Hap_162: 1 [Z2473] Hap_222: 1 [Z2500] Hap_43: 1 [Z6553] Hap_103: 2 [Z2465 Z2797] Hap_163: 1 [Z2476] Hap_223: 2 [Z2505 Z2508] Hap_44: 1 [Z6565] Hap_104: 1 [Z2469] Hap_164: 1 [Z2478] Hap_224: 1 [Z2507] Hap_45: 1 [Z6589] Hap_105: 1 [Z2470] Hap_165: 1 [Z2479] Hap_225: 1 [Z2514] Hap_46: 1 [Z6615] Hap_106: 1 [Z2474] Hap_166: 1 [Z2481] Hap_226: 1 [Z2605] Hap_47: 1 [Z6617] Hap_107: 1 [Z2480] Hap_167: 1 [Z2629] Hap_227: 1 [Z2609] Hap_48: 1 [Z6618] Hap_108: 1 [Z2621] Hap_168: 1 [Z4718] Hap_228: 2 [Z2611 Z6099] Hap_49: 1 [Z6619] Hap_109: 1 [Z2624] Hap_169: 1 [Z6549] Hap_229: 1 [Z2613] Hap_50: 2 [Z6620 Z6580] Hap_110: 1 [Z2625] Hap_170: 2 [SC02 SC57] Hap_230: 1 [Z2614] Hap_51: 3 [Z6625 Z2319 Z3583] Hap_111: 1 [Z2626] Hap_171: 1 [SC03] Hap_231: 1 [Z2616] Hap_52: 1 [Z6627] Hap_112: 1 [Z2627] Hap_172: 1 [SC05] Hap_232: 1 [Z2618] Hap_53: 1 [Z6629] Hap_113: 1 [Z2628] Hap_173: 2 [SC06 SC21] Hap_233: 1 [Z3789] Hap_54: 1 [Z6635] Hap_114: 1 [Z2630] Hap_174: 1 [SC07] Hap_234: 1 [Z3973] Hap_55: 1 [Z6641] Hap_115: 1 [Z2631] Hap_175: 1 [SC08] Hap_235: 1 [Z4723] Hap_56: 1 [Z6643] Hap_116: 1 [Z2633] Hap_176: 2 [SC09 SC64] Hap_236: 1 [Z4823] Hap_57: 1 [Z6644] Hap_117: 1 [Z2639] Hap_177: 1 [SC10] Hap_237: 1 [Z5304] Hap_58: 1 [Z6646] Hap_118: 2 [Z2640 Z6520] Hap_178: 3 [SC11 SC49 SC56] Hap_238: 1 [Z5627] Hap_59: 1 [Z6647] Hap_119: 1 [Z2643] Hap_179: 1 [SC12] Hap_239: 1 [Z6356] Hap_60: 1 [Z6649] Hap_120: 1 [Z4477] Hap_180: 1 [SC13] Appendix III: Unique and shared Control Region haplotypes in the total Portuguese sample (N=293). Computations made in Arlequin v3.5.