scieee Open visual document viewer

Origins and functional consequences of somatic mitochondrial DNA mutations in human cancer

Ju, Seok Young,Alexandrov, Ludmil B,Gerstung, Moritz,Visakorpi, Tapio,Bova, Steve

Abstract

Recent sequencing studies have extensively explored the somatic alterations present in the nuclear genomes of cancers. Although mitochondria control energy metabolism and apoptosis, the origins and impact of cancer-associated mutations in mtDNA are unclear. In this study, we analyzed somatic alterations in mtDNA from 1675 tumors. We identified 1907 somatic substitutions, which exhibited dramatic replicative strand bias, predominantly C > T and A > G on the mitochondrial heavy strand. This strand-asymmetric signature differs from those found in nuclear cancer genomes but matches the inferred germline process shaping primate mtDNA sequence content. A number of mtDNA mutations showed considerable heterogeneity across tumor types. Missense mutations were selectively neutral and often gradually drifted towards homoplasmy over time. In contrast, mutations resulting in protein truncation undergo negative selection and were almost exclusively heteroplasmic. Our findings indicate that the endogenous mutational mechanism has far greater impact than any other external mutagens in mitochondria and is fundamentally linked to mtDNA replication.

Full text

eli esciences.o g Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 1 o 28 O igins and unc ional consequences o soma ic mi ochond ial DNA mu a ions in human cance Young Seok Ju1, Ludmil B Alexand o 1, Mo i z Ge s ung1, Inigo Ma inco ena1, Se ena Nik-Zainal1, Manasa Ramak ishna1, Helen R Da ies1, Elli Papaemmanuil1, Gunes Gundem1, Adam Shlien1, Niccolo Bolli1, Sam Behja i1, Pa ick S Ta pey1, Jyo i Nangalia1,2,3, Cha les E Massie1,2,3, Adam P Bu le 1, Jon W Teague1, Geo ge S Vassiliou1,2,3, An hony R G een2,3, Ming-Qing Du2, Ashwin Unnik ishnan4, John E Pimanda4, Bin Tean Teh5,6, Nikhil Munshi7, Mel G ea es8, Pa esh Vyas9, Adel K El-Nagga 10, Tom San a ius2, V Pe e Collins2, Richa d G undy11, Jack A Taylo 12, D Neil Hayes13, Da id Malkin14, ICGC B eas Cance G oup1†, ICGC Ch onic Myeloid Diso de s G oup1‡, ICGC P os a e Cance G oup1,8,15§, Ch is ophe S Fos e 16,17, Anne Y Wa en2, Hayley C Whi ake 15, Daniel B ewe 8,18, Rosalind Eeles8, Colin Coope 8,18, Da id Neal15, Tapio Visako pi19, William B Isaacs20, G S e en Bo a19, Ad ienne M Flanagan21,22, P And ew Fu eal1,23, Andy G Lynch15, Pa ick F Chinne y24, Ul an McDe mo 1,2, Michael R S a on1, Pe e J Campbell1,2,3* 1Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, Uni ed Kingdom; 2Camb idge Uni e si y Hospi als NHS Founda ion T us , Camb idge, Uni ed Kingdom; 3Depa men o Haema ology, Uni e si y o Camb idge, Camb idge, Uni ed Kingdom; 4Lowy Cance Resea ch Cen e, Uni e si y o New Sou h Wales, Sydney, Aus alia; 5Labo a o y o Cance Epigenome, Na ional Cance Cen e, Singapo e, Singapo e; 6Duke-NUS G adua e Medical School, Singapo e, Singapo e; 7Depa men o Hema ologic Oncology, Dana-Fa be Cance Ins i u e, Bos on, Uni ed S a es; 8Ins i u e o Cance Resea ch, Su on, London, Uni ed Kingdom; 9Wea he all Ins i u e o Molecula Medicine, Uni e si y o Ox o d, Ox o d, Uni ed Kingdom; 10Depa men o Pa hology, MD Ande son Cance Cen e , Hous on, Uni ed S a es; 11Child en's B ain Tumou Resea ch Cen e, Uni e si y o No ingham, No ingham, Uni ed Kingdom; 12Na ional Ins i u e o En i onmen al Heal h Sciences, Na ional Ins i u e o Heal h, T iangle, No h Ca olina, Uni ed S a es; 13Depa men o In e nal Medicine, Uni e si y o No h Ca olina, Chapel Hill, Uni ed S a es; 14Hospi al o Sick Child en, Uni e si y o To on o, To on o, Canada; 15Cance Resea ch UK Camb idge Ins i u e, Uni e si y o Camb idge, Camb idge, Uni ed Kingdom; 16Depa men o Molecula and Clinical Cance Medicine, Uni e si y o Li e pool, London, Uni ed Kingdom; 17HCA Pa hology Labo a o ies, London, Uni ed Kingdom; 18School o Biological Sciences, Uni e si y o Eas Anglia, No wich, Uni ed Kingdom; 19Ins i u e o Biosciences and Medical Technology - BioMediTech and Fimlab Labo a o ies, Uni e si y o Tampe e and Tampe e Uni e si y Hospi al, Tampe e, Finland; 20Depa men o Oncology, Johns Hopkins Uni e si y, Bal imo e, Uni ed S a es; 21Depa men o His opa hology, Royal Na ional O hopaedic Hospi al, Middlesex, Uni ed Kingdom; 22Uni e si y College London Cance Ins i u e, Uni e si y College London, London, Uni ed Kingdom; 23Depa men o Genomic Medicine, The Uni e si y o Texas, MD Ande son Cance Cen e , Hous on, Texas, Uni ed S a es; 24Wellcome T us Cen e o Mi ochond ial Resea ch, Ins i u e o Gene ic Medicine, Newcas le Uni e si y, Newcas le-upon- yne, Uni ed Kingdom *Fo co espondence: pc8@ sange .ac.uk G oup au ho de ails †ICGC B eas Cance G oup: See page 21 ‡ICGC Ch onic Myeloid Diso de s G oup: See page 22 §ICGC P os a e Cance G oup: See page 23 Compe ing in e es s: The au ho s decla e ha no compe ing in e es s exis . Funding: See page 24 Recei ed: 28 Ma ch 2014 Accep ed: 26 Sep embe 2014 Published: 01 Oc obe 2014 RESEARCH ARTICLE Re iewing edi o : Todd Golub, B oad Ins i u e, Uni ed S a es This is an open-access a icle, ee o all copy igh , and may be eely ep oduced, dis ibu ed, ansmi ed, modi ied, buil upon, o o he wise used by anyone o any law ul pu pose. The wo k is made a ailable unde he C ea i e Commons CC0 public domain dedica ion. Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 2 o 28 Resea ch a icle Abs ac Recen sequencing s udies ha e ex ensi ely explo ed he soma ic al e a ions p esen in he nuclea genomes o cance s. Al hough mi ochond ia con ol ene gy me abolism and apop osis, he o igins and impac o cance -associa ed mu a ions in m DNA a e unclea . In his s udy, we analyzed soma ic al e a ions in m DNA om 1675 umo s. We iden i ied 1907 soma ic subs i u ions, which exhibi ed d ama ic eplica i e s and bias, p edominan ly C > T and A > G on he mi ochond ial hea y s and. This s and-asymme ic signa u e di e s om hose ound in nuclea cance genomes bu ma ches he in e ed ge mline p ocess shaping p ima e m DNA sequence con en . A numbe o m DNA mu a ions showed conside able he e ogenei y ac oss umo ypes. Missense mu a ions we e selec i ely neu al and o en g adually d i ed owa ds homoplasmy o e ime. In con as , mu a ions esul ing in p o ein unca ion unde go nega i e selec ion and we e almos exclusi ely he e oplasmic. Ou indings indica e ha he endogenous mu a ional mechanism has a g ea e impac han any o he ex e nal mu agens in mi ochond ia and is undamen ally linked o m DNA eplica ion. DOI: 10.7554/eLi e.02935.001 In oduc ion All cance s esul om soma ic mu a ions in hei genomes. Beyond he ∼3200 Mb o nuclea genomic DNA, human cells ha e hund eds o housands o mi ochond ia p esen in e e y cell, each ca ying one o a ew copies o he 16,569 bp ci cula mi ochond ial genomes (Smei ink e al., 2001; Leg os e al., 2004; eLi e diges The DNA in a cell's nucleus mus be copied ai h ully, and di ided equally, when a cell di ides o p oduce wo new cells. Mis akes—o mu a ions—a e some imes made du ing he copying p ocess, and mu a ions can also be in oduced by exposing DNA o damaging agen s known as mu agens, such as UV ligh o ciga e e smoke. These mu a ions a e hen main ained in all o he descendan s o he cell. Mos o hese mu a ions ha e no impac on he cell's cha ac e is ics (‘passenge mu a ions’). Howe e , ‘d i e mu a ions’ ha allow cells o di ide uncon ollably and sp ead o o he body si es can lead o cance . Mi ochond ia a e cellula compa men s ha a e esponsible o gene a ing he ene gy a cell needs o su i e and a e also esponsible o ini ia ing p og ammed cell dea h. Mi ochond ia con ain hei own DNA—en i ely sepa a e om ha in he nucleus o he cell— ha encodes he p o eins mos essen ial o ene gy p oduc ion. Mi ochond ial DNA molecules a e equen ly exposed o damaging molecules called eac i e oxygen species ha a e p oduced by he mi ochond ia. The e o e, hese eac i e oxygen species ha e been hough o be one o he mos impo an causes o mi ochond ial DNA mu a ions. In addi ion, because cance cells p oduce ene gy di e en ly o no mal cells, mu a ions in he mi ochond ial DNA ha change he abili y o he mi ochond ia o p oduce ene gy ha e been con en ionally hough o help no mal cells o become cance ous. Howe e , conclusi e e idence o a link be ween cance and mi ochond ial DNA mu a ions is lacking. Ju e al. examined he mi ochond ial DNA sequences aken om 1675 cance biopsies om o e hi y di e en ypes o cance and compa ed hese o no mal issue om he same pa ien s. This e ealed 1907 mu a ions in he mi ochond ial DNA aken om he cance cells. The pa e n o he mu a ions sugges s ha he majo i y o he mu a ions a e no in oduced om eac i e oxygen species, bu om he e o s he mi ochond ia hemsel es make in he p ocess o duplica ing hei DNA when a cell di ides. Unexpec edly, known mu agens, such as ciga e e smoke o UV ligh , had a negligible e ec on mi ochond ial DNA mu a ions. Con a y o con en ional wisdom, Ju e al. ound no e idence ha he mi ochond ial DNA mu a ions help cance o de elop o sp ead. Ins ead, like passenge mu a ions ound in he DNA in he cell nucleus, mos mi ochond ial genome mu a ions ha e no disce nible e ec . Howe e , Ju e al. e ealed ha DNA mu a ions ha damage no mal mi ochond ial ac i i y a e less likely o be main ained in cance cells. P esumably, mi ochond ia con aining hese p o eins p oduce less ene gy, and so a cell con aining oo many o hese mu a ions will ind i ha de o su i e. This shows ha ha ing enough co ec ly unc ioning mi ochond ia is essen ial o e en cance cells o h i e. DOI: 10.7554/eLi e.02935.002 Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 3 o 28 Resea ch a icle Koppenol e al., 2011). In addi ion o hei ole in cellula ene gy balance h ough oxida i e phospho- yla ion, mi ochond ia a e in ol ed in many essen ial cellula unc ions including modula ion o oxida ion– educ ion s a us, con ibu ion o cy osolic biosyn he ic p ecu so s, and ini ia ion o apop osis. Mi ochond ia in euka yo ic cells e ol ed by endosymbiosis om a ee-li ing α-p o eobac e ium (G ay e al., 1999). O e 2 billion yea s o co-e olu ion, many ances al mi ochond ial genes ha e ans e ed o he nucleus (Falkenbe g e al., 2007; Cal o and Moo ha, 2010; Wallace, 2012). Wha emains in he mi ochond ial genome is dis inc i e o he s iking asymme y be ween he wo complemen a y m DNA s ands in e ms o nucleo ide con en and gene dis ibu ion (And ews e al., 1999). The hea y (H) s and is guanine- ich (C/G = 0.4) and is he empla e om which mos mi ochond ial p o eins (12 ou o 13) a e ansc ibed, whe eas only one p o ein-coding gene, MT-ND6, is ansc ibed om he co espondingly cy osine- ich ligh (L) s and. Mu a ions in he mi ochond ial genome cause inhe i ed disease (Chinne y, 1993), wi h a ma e nal inhe i ance pa e n because only eggs con ibu e mi ochond ia o he zygo e. The pene ance o inhe i ed mi ochond ial disease is de e mined s ochas ically by bo h he andom asso men o mu a ed s wild- ype mi ochond ial genomes du ing meiosis and andom d i du ing he ea ly cell di isions a e e iliza ion. In cance , he ole o soma ically acqui ed m DNA mu a ions is con o- e sial. Al hough cance -speci ic mu a ions ha e been p e iously epo ed (Polyak e al., 1998; B andon e al., 2006; Cha e jee e al., 2006; He e al., 2010; La man e al., 2012), he limi ed sample size o poo sensi i i y o capilla y sequencing o he e oplasmic mu a ions has no allowed a comp ehensi e analysis o he mu a ional signa u es o mi ochond ial mu a ions no hei likely unc ional signi icance. I has long been p oposed ha mi ochond ia migh con ibu e o cance de el- opmen gi en hei undamen al impo ance o cellula biology (Wallace, 2012). P e ious epo s sugges ed ha mi ochond ial soma ic mu a ions migh be unde posi i e selec ion and hus con ibu e o cance de elopmen , bu he small numbe o epo ed mu a ions ende s his conclusion unce ain (B andon e al., 2006; Cha e jee e al., 2006; La man e al., 2012; Schon e al., 2012). None heless, he hypo hesis o unc ionally ele an mi ochond ial mu a ions is an appealing one because cance cells ha e g ea ly inc eased ene gy demands o e no mal cells and demons a e a swi ch om ae obic glycolysis in mi ochond ia o lac ic acid e men a ion in he cy osol ( he Wa bu g e ec ) (Hanahan and Weinbe g, 2011; Koppenol e al., 2011). In each cell cycle, he eplica ing genome is a isk o de no o mu a ions, which can p omo e he de elopmen o cance . These mu a ions may be gene a ed by in insic cellula e o s du ing DNA eplica ion o epai o h ough exposu e o mu agens, such as eac i e oxygen species, obacco smoke, and ul a iole ligh (Pleasance e al., 2010a, 2010b). Recen ly, >20 mu a ional signa u es ope a i e in cance s ha e been iden i ied in he nuclea genome (Alexand o e al., 2013). Whe he any o hese mu a ional p ocesses also a ec he mi ochond ial genome has no been s udied. Fu he mo e, whe he he e a e m DNA-speci ic mu a ional p ocesses in soma ic cells emain unclea , al hough he many unique ea u es o m DNA eplica ion and epai , coupled wi h he high concen a ion o eac i e oxygen species gene a ed by he elec on anspo chain, could be associa ed wi h dis inc i e mu a ion signa u es. In his s udy, we compa e 1675 cance and pai ed no mal m DNA sequences ac oss 31 umo ypes using massi ely pa allel DNA sequencing echnologies o ob ain a sys ema ic and unbiased ca alog o soma ic mi ochond ial mu a ions. We ind ha m DNA mu a ions a e almos exclusi ely he p oduc o a mu a ional p ocess ha is speci ic o mi ochond ia and p obably linked o he unique mechanism o genome eplica ion hese o ganelles employ. We ind no e idence o posi i e selec ion o mi ochond ial mu a ions du ing oncogenesis, sugges ing ha hey con e no clonal ad an age on he nascen cance cells. Resul s m DNA sequencing and Mu a ion Calling We ex ac ed he m DNA sequences om 704 whole-genome and 971 whole-exome sequencing da a gene a ed on p ima y cance s and compa ed hem wi h m DNA sequences om hei ma ched no mal samples. Gi en he abundance o m DNA pe cance cell, a s anda d co e age o 30–40× in he nuclea genome p o ides signi ican ly g ea e co e age o he mi ochond ial genome (a e age ead dep h = 7901.0×), enabling accu a e iden i ica ion o soma ic mu a ions including a e he e oplasmic a ian s. We also assessed whe he whole-exome sequencing could be used o iden i y m DNA Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 4 o 28 Resea ch a icle mu a ions om o - a ge eads de i ed om he mi ochond ial genome. We ound an a e age ead dep h o 92.1× ac oss he mi ochond ial genome in exome s udies. F om 139 samples in which we had bo h exome and whole-genome sequencing da a, he o e all ead dep hs co ela ed s ongly (R2 = 0.59, Figu e 1— igu e supplemen 1) as did a ian allele ac ions o m DNA soma ic mu a ions (R2 = 0.97, Figu e 1— igu e supplemen 2). Valida ion expe imen s sugges ed he sensi i i y o whole-exome sequencing o de ec ion o m DNA soma ic mu a ions o be 71.4% compa ed o whole-genome sequencing (Figu e 1— igu e supplemen 3 and ‘Ma e ials and Me hods’, ‘O - a ge m DNA eads in whole-exome sequencing’ and ‘DNA c oss-con amina ion’). To educe po en ial alse-posi i e calls o m DNA soma ic mu a ions, we only epo a ian s called wi h an allele ac ion o >3%. This elimina es he isk o miscalls due o m DNA-de i ed pseudogenes in he nucleus (NuMTs) because m DNA copy numbe s a e 100–1000 imes highe han nuclea genomes in human soma ic cells, and he sequence homology be ween m DNA and NuMTs p esen ed in he human e e ence genome is gene ally <95% (in 96 ou o 101 NuMTs wi h leng h g ea e han 300 bp). Fu he mo e, pai wise compa ison be ween cance and ma ched no mal m DNAs om he same indi idual u he minimizes he con amina ion o NuMTs in he mu a ion calling. The ca alog o m DNA soma ic mu a ions In o al, 1675 umo –no mal pai s ac oss 31 umo ypes we e analyzed (Table 1 and Supplemen a y ile 1). Fo 61 o hese pa ien s, we had sequencing da a a ailable om mul iple si es o he p ima y cance , se e al ime poin s o ma ched p ima y cance s, and me as ases (a o al o 73 such cance samples), allowing us o s udy he iming o m DNA mu a ions in cance e olu ion (Supplemen a y ile 1). We iden i ied 1907 soma ic m DNA subs i u ions (Figu e 1 and Supplemen a y ile 2). In con as o inhe i ed polymo phisms (n = 38,706, a ailable a Supplemen a y ile 2), which we e almos always homoplasmic in bo h he cance and coun e pa no mal, he a ian allele ac ions (VAFs) o hese soma ic subs i u ions we e highly a iable in he cance , anging om ou de ec ion h eshold (3%) o homoplasmy (100%). O hese 1907 soma ic subs i u ions, 1209 (63.4%) we e no egis e ed in he da abases o m DNA common polymo phism (Ingman and Gyllens en, 2006; Le in e al., 2013). In compa ison, when we examined subs i u ions ound in bo h he umo and he no mal samples om a pa ien , only 21 (0.05%) we e no egis e ed in he polymo phism da abases, a signi ican ly di e en ac ion om he umo -only a ian s (p < 10−10; Chi-squa ed es ). We ound 595 (31.2%) ecu en mu a ions ha can be collapsed on o 246 m DNA posi ions, which is a 6.9- old highe le el o ecu ence han expec ed by chance (p < 10−10). This sugges s ha he gene a ion o ixa ion o m DNA mu a ions is no andom, bu in luenced by ac o s such as he unde lying mu a ional p ocess o posi i e selec ion. O he 1675 cance samples, 976 (58.3%) ha bo ed a leas one soma ic subs i u ion and 521 (31.1%) had mul iple subs i u ions, anging om 2 o 7 (Figu e 2A). In hose wi h mul iple subs i u- ions, 72 pai s o mu a ions we e su icien ly close o phase (Nik-Zainal e al., 2012b) such ha we could de e mine whe he hey we e linked on he same m DNA genome o we e on di e en copies. We ound ha 45 (62.5%) pai s o mu a ions we e linked on he same m DNA genome (Supplemen a y ile 3 and Figu e 2— igu e supplemen 1). Fu he mo e, o hese linked mu a ions, 33 showed a clea empo al o de : ha is, one mu a ion was demons ably sub-clonal o he o he . This is a he unexpec ed, since each soma ic cell has 100–1000 copies o he mi ochond ial genome, and we migh an icipa e ha andom mu a ions would, on a e age, a ec di e en copies. Tha many pai s o mu a ions a e phased on he same m DNA genome and ye show a clea sub-clonal ela ionship sugges s ha hey occu su icien ly sepa a ed in ime o allow he mi ochond ial genome ca ying he ea lie mu a ion o d i owa ds a subs an ial ac ion o all genomes in ha cell be o e he second mu a ion occu s, consis en wi h a p e ious epo (De Alwis e al., 2009). The numbe o soma ic m DNA subs i u ions a ied signi ican ly acco ding o umo ype (p = 4.4 × 10−52) a e co ec ing o con ounding a iables such as sequencing co e age: gas ic, hepa ocellula , p os a e, and colo ec al cance s had he highes numbe o m DNA subs i u ions (Figu e 2B). In con as , hema ologic cance s (acu e lymphoblas ic leukemia, myelop oli e a i e disease, and myelo- dysplas ic synd ome) had ewe mu a ions. Se e al possible explana ions could unde pin hese di e - ences ac oss umo ypes. I could be ha he mu a ion a es di e ac oss cell lineages; i could be ha selec ion p essu es shape he numbe o mu a ions; o he numbe o m DNA genome gene a ions could di e ac oss cell lineages. O hese explana ions, we belie e ha he second is unlikely because, as we shall see, posi i e selec ion is no a majo componen o mi ochond ial mu a ions. In e es ingly, Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 5 o 28 Resea ch a icle Table 1. Summa y s a is ics o m DNA sequence da a WGS WXS A e age m RD (WGS) A e age m RD (WXS) To al WGS WXS A e age m RD (WGS) A e age m RD (WXS) To al B eas 284 98 11594.3 52.7 382 Meningioma 0 12 - 42.5 12 Colo ec al 1 75 34916.9 276.6 76 Ependymoma 1 9 10323.7 52.7 10 Lung 60 0 2798.1 - 60 P os a e 80 0 17810.6 - 80 MPD 12 138 1517.0 10.9 150 Hepa ocellula 0 47 - 205.8 47 MDS 3 75 5648.7 44.5 78 Melanoma 13 13 513.9 353.5 26 ALL 64 6 886.6 35.9 70 Gas ic 0 13 - 184.1 13 CLL 6 0 5002.2 - 6 Cholangioca cinoma 0 8 - 143.9 8 AML 1 6 6783.6 27.4 7 Meso helioma 0 6 - 106.3 6 Mul iple myeloma 0 69 - 43.2 69 Bladde 54 0 646.2 - 54 AMKL 0 9 - 24.2 9 Renal 0 23 - 35.4 23 Lymphoma 0 4 - 99.5 4 O a ian 0 38 - 58.9 38 U e ine 27 23 736.0 149.5 50 Os eosa coma 38 90 9525.5 119.2 128 Ce ical 0 52 - 85.2 52 Chond osa coma 0 47 - 99.1 47 Adenoid cys ic ca. 1 60 714.7 75.6 61 Ewing sa coma 0 27 - 69.5 27 Head & Neck 43 3 1369.1 18.8 46 Kaposi sa coma 0 9 - 181.0 9 Cho doma 16 11 1240.0 82.1 27 To al; 31 cance ypes 704 971 1675 WGS, whole-genome sequencing; WXS, whole-exome sequencing; m RD, mi ochond ial ead dep h; MPD, myelop oli e a i e disease; MDS, myelodysplas ic synd ome; ALL, acu e lymphoblas ic leukemia; CLL, ch onic lymphoblas ic leukemia; AML, acu e myeloid leukemia; AMKL, acu e megaka yoblas ic leukemia. DOI: 10.7554/eLi e.02935.003 Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 6 o 28 Resea ch a icle Figu e 1. Mi ochond ial soma ic subs i u ions iden i ied om 1675 Tumo –No mal pai s. m DNA genes and in e genic egions a e shown. The s and o genes is shown based on m DNA s and con aining equi alen sequences o ansc ibed RNA. Subs i u ion ca ego ies (silen , non-silen (missense and nonsense), non-coding ( RNA and RNA), and in e genic) a e shown by he shapes o each subs i u ion. Six classes o subs i u ions a e p esen ed colo -coded. The subs i u ions on he H, and L s and (when six subs i u ional classes we e conside ed) a e shown ou side and inside o m DNA genes, espec i ely. Ve ical axes o H and L s and subs i u ions ep esen he VAF o each a ian . DOI: 10.7554/eLi e.02935.004 The ollowing igu e supplemen s a e a ailable o igu e 1: Figu e supplemen 1. Co ela ion in amoun o m DNA eads be ween whole-genome and whole-exome sequencing. DOI: 10.7554/eLi e.02935.005 Figu e supplemen 2. Co ela ion o he e oplasmy le els be ween whole-genome and whole-exome sequencing. DOI: 10.7554/eLi e.02935.006 Figu e supplemen 3. Valida ion o m DNA soma ic subs i u ions. DOI: 10.7554/eLi e.02935.007 Figu e supplemen 4. Amoun o o - a ge m DNA eads ac oss ou sequencing cen e s. DOI: 10.7554/eLi e.02935.008 Figu e supplemen 5. Fil e ing samples o po en ial DNA con amina ions. DOI: 10.7554/eLi e.02935.009 we ind a posi i e co ela ion be ween he numbe o m DNA soma ic mu a ions and age a diagnosis in b eas cance s (p = 0.0004; Figu e 2C), in keeping wi h he idea ha he numbe o mi ochond ial gene a ions is linked o mu a ion bu den. The mu a ional bu den o an es ablished cance ep esen s Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 7 o 28 Resea ch a icle he accumula ed a ia ion acqui ed in he lineage o cell di isions om e ilized egg o ans o med cell and will include e en s acqui ed in no mal de elopmen and homeos asis as well as hose acqui ed du ing umo igenesis (S a on e al., 2009). In e es ingly, m DNA mu a ions ha e been ound a high Figu e 2. m DNA soma ic subs i u ions o human cance . (A) Numbe o soma ic subs i u ions in a umo sample. (B) A e age numbe o soma ic subs i u ions pe sample ac oss 31 umo ypes. (C) Age o diagnosis and numbe o m DNA soma ic subs i u ions in b eas cance s. DOI: 10.7554/eLi e.02935.010 The ollowing igu e supplemen is a ailable o igu e 2: Figu e supplemen 1. VAFs o phased soma ic m DNA subs i u ions. DOI: 10.7554/eLi e.02935.011 Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 8 o 28 Resea ch a icle a es in no mal colonic c yp cells (Taylo e al., 2003; E icson e al., 2012). Gi en ha we ind high bu dens o mu a ions in colonic umo s as well, he di e ences we see ac oss umo ypes may a ise om p e- o pos - ans o ma ion di e ences in m DNA bu den ac oss issues. Ex ac ing m DNA mu a ional signa u es Wi h espec o signa u es o soma ic subs i u ions, C > T and T > C ansi ions cons i u ed 90.9% o all he 1907 subs i u ions (Figu e 1) among he six classes o possible base subs i u ions. To cha ac- e ize his agg ega ed signa u e o m DNA cance speci ic mu a ions in mo e de ail, we looked o he p esence o m DNA s and bias be ween he complemen a y H and L s ands o m DNA. The wo main subs i u ion classes showed an ex eme le el o m DNA s and bias. 84.1% o he C > T ansi ions we e on he H s and. This le el o s and bias occu ed despi e he ac ha cy osine is 2.4- old less common on he H han he L s and, so he C > T subs i u ion a e is 12.6- old highe on he H s and. By con as , 76.8% o he T > C ansi ions we e on he L s and despi e i s lowe hymine con en (1.3- old less han he H s and). This implies ha he T > C mu a ion a e on he L s and is 4.2- old highe han on he H s and. We hen examined he sequence con ex in which hese mu a ions occu ed by examining he bases immedia ely 5′ and 3′ o he mu a ed bases. This gene a es 96 possible mu a ion classes ( he 6 subs i u ion classes mul iplied by he 16 combina ions o immedia e 5′ and 3′ nucleo ides). Bo h C > T and T > C mu a ions showed highly dis inc i e sequence con ex s. CH > TH subs i u ions (i.e. C > T mu a ions on he H s and) we e en iched o he NpCpG inucleo ide con ex (8- o 15- old mo e equen han expec ed by chance; Figu e 3A). By con as , TL > CL subs i u ions (i.e. T > C mu a ions on he L s and) showed 5- o 8- old en ichmen in NpTpC. This s and-asymme ic mu a ional signa- u e is no simila o any o he 21 cance -associa ed mu a ional signa u es ecen ly iden i ied om he nuclea DNA o 30 di e en cance ypes (Alexand o e al., 2013). O he 18 umo ypes ha p esen ed a leas 25 m DNA soma ic subs i u ions in his s udy, he mu- a ional signa u es we e b oadly consis en ac oss umo ypes (Figu e 3B), wi h he excep ion ha mul iple myeloma had a somewha highe a e o TH > CH changes han o he his ologies (p = 8.1 × 10−6). Thus, in con as o he mu a ional signa u es ound in nuclea genomes, whe e he e is s iking he e o- genei y bo h ac oss umo ypes and ac oss indi iduals wi hin a umo ype (Alexand o e al., 2013), he mu a ional p o ile in he mi ochond ial genome o soma ic cells is ema kably homogeneous. Replica ion-coupled mu a ional p ocess in mi ochond ia The majo known cause o mu a ional s and bias in nuclea DNA is ansc ip ion-coupled nucleo ide excision epai , whe e DNA lesions on he ansc ibed (non-coding) s and a e mo e equen ly epai ed (Alexand o e al., 2013). Howe e , we ind ha he s and bias always a o s CH > TH and TL > CL whe he he gene is ansc ibed om he H s and o om he L s and (Figu e 3— igu e supple- men 1). This is no compa ible wi h ansc ip ion-coupled epai , o which he di ec ion o s and bias is undamen ally dic a ed by which s and is ansc ibed. Ins ead, he m DNA mu a ional s and bias epo ed he e appea s o be d i en by di e ences in eplica ion be ween he wo s ands. m DNA eplica ion ha bo s subs an ial s and asymme y be ween he H and L s ands: m DNA eplica ion ini ia es om an o igin o eplica ion (OH) in he D-loop, wi h he nascen H and he L s and eplica ing as leading and lagging s and, espec i ely (Clay on, 1982; Falkenbe g e al., 2007; Hol and Reyes, 2012). We obse ed ha C > T subs i u- ions we e p e alen in he leading (hea y) s and, whe eas T > C subs i u ions we e ound in he lag- ging (ligh ) s and (Figu e 1). Rema kably, his s and bias was e e sed in he D-loop i sel (Figu es 1 and 3C), u he sugges ing ha he m DNA soma ic mu a ions a e eplica ion-coupled: acco ding o a ecen ly p oposed bidi ec ional model o m DNA eplica ion (Yasukawa e al., 2005, 2006; Hol and Reyes, 2012), m DNA eplica ion is also able o ini ia e om he so-called O i-b si e, ypically loca ed a ound genomic posi ion 16,197 and p oceeds on bo h s ands away om he o igin (Figu e 1). Replica ion o he nascen H s and con inues unimpeded like he adi ional model, bu he nascen L s and e mina es a he so-called OH si e, ypically a ound m DNA posi ion 191 bp. Unde his model, hen, he leading and lagging s and a e e e sed in he ew hund ed base-pai s o he D-loop, which is consis en wi h he e e sed mu a ional signa u e in his egion (Figu es 1 and 3C). Equi alen mu a ional signa u e du ing human m DNA E olu ion I is no en i ely s aigh o wa d o in e he mu a ional signa u es ope a ing on he mi ochond ial genome in he ge mline. De no o mu a ions a e gene ally a e and o en disco e ed because hey Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 9 o 28 Resea ch a icle Figu e 3. Replica i e s and bias o m DNA soma ic subs i u ions. (A) Replica i e s and-speci ic subs i u ion a e (# o obse ed/# o expec ed) by 96 inucleo ide con ex . Subs i u ions in a speci ic m DNA segmen ( om O i-b o OH) a e no included, because hey p esen a di e en subs i u ional signa u e. (B) Mu a ional signa u e ac oss umo ypes. Eigh een umo ypes, which include a leas 25 m DNA mu a ions, we e shown. (C) In e ed subs i u ion signa u e in he O i-b–OH. DOI: 10.7554/eLi e.02935.012 The ollowing igu e supplemen is a ailable o igu e 3: Figu e supplemen 1. Replica i e s and bias obse ed in m DNA subs i u ions. DOI: 10.7554/eLi e.02935.013 cause disease; dis inguishing he ances al base and he de i ed base is challenging o single nucleo- ide polymo phisms; and compa a i e m DNA genomics ac oss species ex ends o e conside able e olu iona y ime. In con as , because ances al and de i ed s a es a e de ined o umo –no mal Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 16 o 28 Resea ch a icle allele equency is ∼50% (45–55%) acco ding o The 1000 Genomes P ojec (1000 Genomes P ojec Conso ium e al., 2010). O he 320 si es, homozygous posi ions in no mal issues (which showed >90% a ian allele ac ion (VAF) wi h bases Q sco e >20) we e compa ed wi h he co esponding geno ypes in he coun e pa cance . Sample pai s we e emo ed i he geno ype misma ch a e was g ea e han 0.1 ( Nhe Nw Nhom Nhe Nw + ++ ; Nhe , numbe o he e ozygo e posi ions; Nhom, numbe o homozygo e posi ions; Nw , numbe o wild- ype posi ions) (Figu e 1— igu e supplemen 5A). We no e 0 is expec ed o he a e when geno yping is pe ec and sample pai s a e om he same indi idual. By con as , 0.5 is expec ed when samples we e om di e en indi iduals. 2. Mino c oss-con amina ion We es ima ed DNA c oss-con amina ion le els wi h he VAF o au osomal homozygous SNPs geno yped om he common (popula ion mino allele equency ∼50%) SNP si es. Theo e ically, i he e is no sequencing (and mapping) e o , all he homozygo e SNP si es in pu e samples should p esen 100% VAFs. Howe e , when samples a e con amina ed, co esponding VAFs a e educed because he con aminan has only an ∼25% o chance o ha ing homozygo e SNPs on he same si e. The e o e, mino con amina ion le els (C) o each cance sequencing da a we e es ima ed as below: ∑ ∑ RCw Ne C RDhom Ne = × ( )– 2, ( )– whe e RDhom is sequencing ead dep h, RCw is ead-coun o wild- ype alleles, and Ne is numbe o sequencing e o s on each au osomal homozygo e SNP si e. Fo high accu acy, we only coun ed base wi h su icien quali y sco e (Q > 20). In o de o es ima e Ne, we assumed a conse a i e a e (sequencing e o a e = 0.001). We conside ed si es co e ed by a leas 10 eads and 90% VAF (Figu e 1— igu e supplemen 5B). 95% con idence in e als o c oss- con amina ion le els we e calcula ed using binomial dis ibu ion. In o de o clea soma ic a ian s, he e we made he e y conse a i e assump ion ha soma ic a ian s p esen in excess o 5- imes o he 95% uppe limi o C le els we e ue soma ic a he han alse-posi i es by low-le el o c oss-con amina ion. 3. Ge mline polymo phisms and back mu a ions We u he checked samples o con amina ion using known m DNA polymo phisms. Because human m DNA is small (16,569 bp) and ex ensi ely explo ed p e iously, mos o ge mline m DNA polymo phisms a e al eady known. Fo example, 97.7% o he 39,036 inhe i ed subs i u- ions we e known polymo phisms in he m DB da abase (Ingman and Gyllens en, 2006). The e o e, when a umo sample is con amina ed by o he samples, many soma ic-like m DNA subs i u ions by con aminan s a e likely o be o e lapped wi h known m DNA polymo phisms. A he same ime, low-le el con amina ion would gene a e excessi e back mu a ions, which appea ed o e e se ge mline common polymo phisms in o wild- ype alleles. Taken oge he , bo h he numbe o soma ic subs i u ions known in m DB and numbe o back mu a ions can be good indica o s o m DNA c oss-con amina ion. The e o e, we il e ed ou umo issues wi h ≥3 known po en ially soma ic mu a ions o wi h ≥2 back mu a ions om he u he analyses (Figu e 1— igu e supplemen 5A and B). Va ian calling We ex ac ed m DNA eads using Sam ools (Li and Du bin, 2009). We used Va Scan2 (Kobold e al., 2012) o ini ial a ian calling wi h a ew op ions (--s and- il e 1 (misma ches should be epo ed by bo h o wa d and e e se eads), --min- a - eq 0.03 (minimum VAF 3%), --min-a g-qual 20 (minimum base quali y 20), --min-co e age 3 and --min- eads2 2). Wi h espec o he --s and- il e , i gene ally emo es a ian when >90% o misma ches a e epo ed om ei he o he H o he L m DNA s and. Howe e , whe e only eads wi h a speci ic o ien a ion a e could be aligned dominan ly (i.e. in bo h ex eme egion o mi ochond ial e e ence genome; only L s and eads could be aligned on he 5′ ex eme o m DNA), we compa ed s and bias be ween ‘pe ec ma ches’ (# pe ec ma ches om L s and Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 17 o 28 Resea ch a icle eads / o al # pe ec ma ches) and misma ches (# misma ches om L s and eads / o al # misma ches). I he di e ence be ween hose wo bias <0.1, he mu a ions we e escued. O he 1907 mu a ions, 54 (2.8%) we e escued acco dingly. Pu a i e soma ic a ian s called by Va Scan2 we e u he il e ed using c i e ia shown below. 1. A leas 4 unique eads suppo ing a ian s and all a ian eads a leas 20 ph ed scale sequencing quali y sco e (Q 20 = 1% sequencing e o a e) and a leas 3% a ian allele ac ions (VAFs). A. Rega dless o in WGS and in WES, he ≥4 misma ches and he ≥3% VAF c i e ia mus be sa is- ied simul aneously. B. Howe e , in WGS, he minimum numbe o eads (n = 4) c i e ion is no essen ial, because he ≥3% VAF c i e ion is much mo e s ingen (3% VAF eques a leas 240 misma ches (>>4) gi en m DNA co e age is ∼8000 o WGS). C. In WES, he ≥3% VAF c i e ion is ela i ely less impo an han in WGS, because he ≥4 mis- ma ches c i e ion is mo e s ingen . Fo example, 4 misma ches in 90x (WXS a e age) co e - age egion (VAF = 4.4%) au oma ically ul ill he ≥3% VAF c i e ion. Fo less co e ed egions (i.e. <40x co e age; n = 285 ou o o al 1907 subs i u ions), he VAF c i e ion becomes less impo an , because 4 misma ches would gene a e ≥10% VAF, much highe han he minimum h eshold (i.e. 3%). As esul s, we a e missing lowe he e oplasmic a ian s (i.e. a ian s wi h 3–10% he e oplasmic le els) om low co e age samples (mos ly by WXS). The lowe sensi i i y o WXS is also con i med in ou alida ion s udy (see “Valida ion o soma ic a ian s” below). 2. The e is no minimum h eshold o o al co e age (# pe ec ma ches + # misma ches). 3. To inc ease sensi i i y o de ec ing mu a ions, we escued mu a ions wi h 3 unique a ian eads (wi h a leas 20 ph ed scale sequencing quali y sco e) when VAFs is ≥ 20%. O 1907 soma ic sub- s i u ions, 32 (1.7%) we e escued acco dingly. 4. All soma ic a ian s p esen ing wi h VAFs lowe han ou e y conse a i e h eshold o mino c oss-con amina ion (5- imes 95% uppe limi o con amina ion le els o each umo sample, see abo e “Mino c oss-con amina ion o DNA samples”) we e emo ed. When we could no es ima e c oss-con amina ion le els because o low sequencing dep h o co e age ( o nuclea genome), a conse a i e c i e ion (10% con amina ion le el h eshold) was explici ly used. 5. Subs i u ions we e u he isually inspec ed using IGV (Tho aldsdo i e al., 2013). Thi een equen alse-posi i e a ian s (shown below) by misalignmen due o ex ensi e le el o homopoly- me s in CRS and due o sequencing e o in he e e ence m DNA genome (3107N, see Mi omap (h p://www.mi omap.o g/bin/ iew.pl/MITOMAP/Camb idgeReanalysis) o mo e in o ma ion) we e explici ly emo ed: 1. Misalignmen due o ACCCCCCCTCCCCC ( CRS 302-315) A302C, C309T, C311T, C312T, C313T, G316C 2. Misalignmen due o GCACACACACACC ( CRS 513-525) C514A, A515G, A523C, C524G 3. Misalignmen due o 3107N in CRS (ACNTT, CRS 3105-3109) C3106A, T3109C, C3110A We compa ed ou a ian calls wi h common inhe i ed m DNA polymo phisms deposi ed in he m DB da abase as o 24 h July 2013 (Ingman and Gyllens en, 2006). Gene anno a ion o soma ic a ian s was done using cus om sc ip based on human m DNA gene in o ma ion (Ruiz-Pesini e al., 2007). Valida ion o soma ic a ian s To alida e he sensi i i y and speci ici y o a ian calling in his s udy, 19 umo and no mal pai s (which we e o iginally whole-genome sequenced) we e whole-exome sequenced and m DNA a ian s we e assessed independen ly. Among he 28 soma ic subs i u ions o iginally de ec ed om he 19 umo – no mal whole-genome sequencing pai s, 20 (71.4%) we e called as soma ic (Figu e 1— igu e supple- men 3). In addi ion, 5 (17.9%) p esen ed e idence o a ian eads in he alida ion se , al hough i was il e ed ou because o i s low ead dep h o co e age in exome sequencings (showed 2–5 a ian eads). Mo eo e , because 3 emaining si es we e no su icien ly co e ed in he alida ion se o call soma ic a ian s, hese could no be e idence o he inaccu acy o whole-genome sequencing da a, he e o e no conside ed in he accu acy alida ion. Taken oge he , all he 25 soma ic subs i u ions by Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 18 o 28 Resea ch a icle whole-genome sequencing we e highly likely o be ue posi i es, he e o e we concluded i p o ided ∼100% accu acy in he m DNA soma ic subs i u ion assessmen . Ac ually, he high accu acy o whole- genome sequencing is e y likely and wha we expec , because i p o ides ex ensi e co e age o m DNA (a e age ead dep h >7,500×), ∼3% he e oplasmic a ian s would p esen >200 a ian eads. By con as , he alida ion se (whole-exome sequencing) is called 21 soma ic subs i u ions. O hese, 20 we e common wi h whole-genome sequencing, and one was inco ec ly called as soma ic hough i was ac ually ge mline subs i u ions in he whole-genome sequencing da a. In addi ion, as men ioned abo e, he alida ion se missed 8 soma ic subs i u ions called by whole-genome sequencing. Six ou o eigh unde calls (75%) we e low he e oplasmic subs i u ions in whole-genome sequencing, anging om 3.36% o 8.68%. Based on hese da a, we sugges 71.4% sensi i i y (20/28) and 95.2% speci ici y (20/21) o exome-sequencing in de ec ing up o 3% he e oplasmic soma ic m DNA subs i u ions in cance . We u he checked he co ela ion o he e oplasmy le el be ween he 20 m DNA soma ic mu a ions called bo h whole-genome and whole-exome sequencing. I showed g ea linea ela ionship (R2 = 0.97, Figu e 1— igu e supplemen 2), u he sugges ing whole-exome sequencing da a a e app op ia e o accu a e de ec ion o m DNA soma ic mu a ions. Subs i u ion phasing We phased 72 soma ic subs i u ion pai s, which a ose in a single cance sample and which loca ed su icien ly close ( om 10 bp o ∼500 bp), he e o e bo h si es could be sequenced by same sequence agmen s (Supplemen a y ile 3 and Figu e 2— igu e supplemen 1). We classi ied hem as ‘di e en s and’, ‘co-clonal’, and ‘sub-clonal’ using c i e ia as ollows: Di e en s and: he wo soma ic subs i u ions a e obliga e on di e en s ands. Reads ha epo wild- ype1(w )-subs i u ion2(subs) and subs1-w 2, bu subs1-subs2, a e obse ed. Co-clonal: eads epo ing w 1-w 2 and subs1-subs2 a e only obse ed. Sub-clonal: One subs i u ion is sub-clonal o he o he , bu he wo a e de ini ely phased. Reads subs1-subs2 and ei he subs1-w 2 o w 1-subs2 a e obse ed. Tumo ype and m DNA soma ic subs i u ions To unde s and he ela ionship be ween umo ypes and numbe o m DNA mu a ions, Poisson e- g ession and ANOVA we e applied o ou da ase using R so wa e (h p://www. -p ojec .o g). sub T N Fi 1 < - glm(N Co + Co , amily = poisson())~ sub T N Fi 2 < glm(N Co + Co + , amily = poisson())-~ ano a(Fi 1, Fi 2, es = Chisq ),“” whe e Nsub is numbe o m DNA subs i u ions o each sample, Co T and Co N a e co e age o umo and no mal m DNA, espec i ely, (i Co is >200, we eplaced i by 200), is umo ypes. Age and m DNA soma ic subs i u ions Poisson eg ession was applied o ou b eas cance da ase . sub T N Fi 1 < glm(N Co + Co + a, amily = poisson()),-~ whe e Nsub is numbe o m DNA subs i u ions o each sample, Co T and Co N a e co e age o umo and no mal m DNA, espec i ely, (i Co is >200, we eplaced i by 200), a is age a diagnosis. p- alue in es ima ion o a was shown in he manusc ip . Mu a ional signa u e and s and bias Di e en mu a ional p ocesses gene a e di e en combina ions o mu a ion ypes, e med ‘signa u es’ (Nik-Zainal e al., 2012a). Fo example, ul a iole (UV) ligh and obacco smoking (polycyclic a oma ic hyd oca bons) equen ly gene a e C > T ansi ions and G > T ans e sions on non- ansc ibed (coding) s ands in melanoma and lung cance s, espec i ely (Pleasance e al., 2010a, 2010b). To unde s and he mu a ional p ocesses in luencing cance m DNA, we co ela ed he 1907 m DNA subs i u ions wi h 21 cance speci ic mu a ional signa u es in he nuclea DNA ecen ly iden i ied Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 19 o 28 Resea ch a icle (Alexand o e al., 2013). Howe e , none o he signa u e could explain he highly unique m DNA subs i u ions. Mu a ional signa u e and s and bias we e assessed as desc ibed in ou p e ious epo s (Alexand o e al., 2013). B ie ly, he immedia e 5′ and 3′ sequence con ex was ex ac ed om CRS. Subs i u ion a e o each inucleo ide con ex was calcula ed wi h he numbe o subs i u ion no malized by he equency o he inucleo ide con ex obse ed in he CRS, in he L and H s and, espec i ely. Fo analyses o subs i u ions alling in he m DNA genes (13 p o ein-coding and 22 RNA genes), ansc ibed/non- ansc ibed s and was also conside ed o compa ison. In o de o p o e he s and bias is no ansc ip ion bu eplica ion-coupled, we checked s and biases o polymo phisms in he 12 L s and p o ein-coding genes, 1 H s and p o ein-coding gene (MT-ND6), and/o 22 RNAs (Figu e 3— igu e supplemen 1). Fo his speci ic pu pose, we did no conside he sequence con ex (immedia e 5′ and 3′ bases) because i o e -classi ies mu a ions (i.e. he numbe o mu a ion classes (n = 96) is la ge han ha o mu a ions). In o he wo ds, 12 classes o subs i u ions (six classes o possible base subs i u ions (C > A, C > G, C > T, T > A, T > C, T > G) × wo s ands (L and H s ands)) we e conside ed. Subs i u ion a es a e a io be ween obse ed and expec ed numbe s (H0 = same mu a ion a e o all subs i u ion classes) o each subs i u ion class. In o de o unde s and which model ( eplica i e o ansc ip ional s and) is app op ia e o explain he s and-bias, Chi-squa e es s we e used be ween he numbe o obse ed mu a ions o each class and expec ed ones unde he backg ound signa u e. m DNA codon usage We coun ed he codon equencies in 13 m DNA p o ein-coding genes. Because 12 L s and p o ein- coding genes and 1 H s and gene (MT-ND6) a e unde opposi e mu a ional p essu e (T > C and G > A o L s and genes; A > G and C > T o MT-ND6), we sepa a ed L and H s and genes o his analysis. T > C skew and G > A skew we e calcula ed as shown below, o unde s and he TL > CL and CH > TH (equi alen o GL > AL) subs i u ions du ing he e olu ion o human m DNA: CT AG skew skew CT AG –– TC = GA = , ++ NN NN and NN NN >> whe e NA, NC, NG, and NT a e numbe o A, C, G, and T base in he 3 d posi ion o iple codons in m DNA genes, espec i ely. Fo he assessmen o m DNA codon usage o o he animal species, we analyzed he m DNA sequence o Caeno habdi is elegans (accession# NC_001328), D osophila melanogas e (accession# NC_001709), D. e io (accession# NC_002333), Xenopus lae is (accession# NC_001573), Mus musculus (accession# EU450583), Gallus domes icus (accession # NC_235570), and Pan oglody es (NC_001643). We conside ed only L s and m DNA genes in he c oss-species analysis. Recu en subs i u ions To compa e he numbe o ecu en subs i u ions be ween silen and missense subs i u ions, we andomly selec ed 100 subs i u ions each om 198 silen subs i u ions in he hi d base o iple codons, 440 missense subs i u ions in he i s base o iple codons, and 405 missense subs i u ions in he second base o iple codons. We coun ed he numbe o ecu en subs i u ions in each g oup. This was i e a ed 300 imes independen ly. ANOVA es ing was applied o de e mine he di e ence be ween he h ee g oups (Figu e 5— igu e supplemen 1). dN/dS a io To es ima e dN/dS alues o missense mu a ions (wmis), we used an adap a ion o he me hod desc ibed p e iously (G eenman e al., 2006). B ie ly, he a e o mu a ions is modeled as a Poisson p ocess, wi h a a e gi en by a p oduc o he mu a ion a e and he impac o selec ion. To ob ain accu a e es ima es o dN/dS, we used wo sepa a e models, one using 12 single-nucleo ide subs i u ion a es and a mo e complex one accoun ing o any con ex dependence e ec by 1-nucleo ide ups eam and downs eam using 192 subs i u ion a es. Fo example in he 12- a e model, he expec ed numbe o A > C mu a ions (λA>C) would be modeled as ollows: syn,A C A C syn,A C = L >> > λ* mis,AC AC mis mis,AC = w L , >> > λ ** Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 20 o 28 Resea ch a icle whe e Lsyn,A>C and Lmis,A>C a e he numbe o si es ha can su e a synonymous and missense A > C mu a ion, espec i ely, which a e calcula ed o any pa icula sequence. The likelihood o obse ing he numbe o missense A > C mu a ions (Nmis,A>C) gi en he expec ed λmis,A>C is hen calcula ed as: mis,A C A C mis Lik = Poisson(N ,w ) >> | and he likelihood o he en i e model is he p oduc o all indi idual likelihoods. Wmis is ixed o be equal in all 12 (o 192) equa ions desc ibing each subs i u ion ype, and a hill-climbing algo i hm is used o ind he maximum likelihood es ima es o all a e and selec ion pa ame e s. Likelihood Ra io Tes s a e hen used o es de ia ions om neu ali y (wmis = 1). The dN/dS a io epo ed in he main ex co esponds o he ull con ex dependen model wi h 192 subs i u ion a es. This me hod allows quan i ying he s eng h o selec ion a oiding he con ounding e ec o gene leng h, sequence com- posi ion, di e en a es o each subs i u ion ype, and con ex -dependen mu agenesis. Sho indels Along wi h he 1907 soma ic m DNA subs i u ions, we iden i ied 109 and 142 soma ic sho inse ions and dele ions, espec i ely, om he 1675 cance m DNA sequences using Va scan2 (Supplemen a y ile 2). E olu iona y dynamics o neu al mi ochond ial mu a ions We model he e olu iona y dynamics o mi ochond ial mu a ions unde andom d i and de i e a simple equa ion o he expec ed numbe o homoplasmic mu a ions. The e exis mul iple le els a which mi ochond ial mu a ions e ol e: wi hin mi ochond ia, in he cy oplasm, and on he cellula le el (Rand, 2011). In his s udy, we ocus on he dynamics in a single cell, which ep esen s he ounde o he las clonal expansion in he umo cell popula ion. The cellula dynamics du ing a clonal expansion is di icul o desc ibe analy ically, bu i is impo an o ealize ha mu a ions o a clonal expansion p ese es he allele equencies o neu al a ian s and ha mu a ions ha occu a e he expansion a e unlikely o con ibu e o measu able allele equencies, as he popula ion becomes la ge. We model he e olu iona y dynamics o mi ochond ial mu a ions in he cy oplasm o a single cell by a W igh –Fishe p ocess (W igh , 1931), in which he numbe o mi ochond ia in a subsequen gene a ion is a binomial sample o he mi ochond ia in he p e ious gene a ion. The numbe o mi ochond ia M is kep ixed. The ma ginal allele equency X o a single si e has wo abso bing bounda ies, X = 0 and X = M (homoplasmy), and he p obabili y o ixa ion o an allele a equency X by neu al d i is ρ = X/M (W igh , 1931). No e ha his p ocess leads, on he popula ion le el, o a dicho omiza ion o he e oplasmic a ian s o ei he go ex inc o become homoplasmic and ixa e in a cell. Mu a ions on any o L (= 16,569 n ) si es in he mi ochond ial genome a e assumed o occu a a uni o m a e μ pe nucleo ide pe cell di ision, which is o o de 10−7, based on a human in e -gene a ional compa ison (Colle e al., 2001). Hence he a e o neu al e olu ion is simply μ LM/M = μ L (Kimu a, 1984). Las ly, he expec ed ime o ixa ion in he W igh –Fishe p ocess is = 2M. Pu ing hese hings oge he , he expec ed numbe o mu an alleles N in a cell ini ially wi hou any mi ochond ial mu a- ions a e T gene a ion is E[ ] = ( – ) N L T 2M µ This equa ion p edic s a linea accumula ion o neu al mu a ions o e ime, wi h a delay imposed by numbe o mi ochond ial copies. A simila beha io has been epo ed using nume ical simula ions (Colle e al., 2001). When also conside ing he e oplasmic mu a ions, he expec ed numbe o al e a ions may be sligh ly highe . To check whe he ou model yields he co ec beha io , we use he ollowing numbe s: he obse ed o de o magni ude o mi ochond ial mu a ions pe pa ien was N = 1. The sequencing co - e age on he mi ochond ial genome indica es ha he e we e o o de M = 100 mi ochond ial genome copies p esen pe cance cell. The expec ed numbe o mu a ions pe cell di ision is μ L = 1.6 × 10−3, i he e o e equi es a ound 1000 cell gene a ions T o accumula e on a e age one homoplasmic mu a ion. This numbe o gene a ions appea s ealis ic o egene a ing issues. As expec ed, epi helial cance s had among he highes obse ed numbe o mi ochond ial mu a ions, while hema opoie ic cance s ypically had lowe numbe s. Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 21 o 28 Resea ch a icle S a is ical es ing S a is ical es ing was pe o med using R so wa e. All p- alues we e calcula ed by wo- ailed es ing. Figu es we e gene a ed using R and Mic oso Excel so wa e. Acknowledgemen s Da a used in his manusc ip a e desc ibed in he supplemen a y ma e ials (Supplemen a y ile 1). We hank Thomas Bleaza d a Facul y o Medical and Human Sciences, Uni e si y o Manches e o dis- cussion and assis ance wi h manusc ip p epa a ion. We would like o hank The Cance Genome A las (TCGA) P ojec Team and hei specimen dono s o p o iding sequencing da a. This wo k was sup- po ed by he Wellcome T us , he B i ish Lung Founda ion, he Heal h Inno a ion Challenge Fund, he Kay Kendall Leukaemia Fund, he Cho doma Founda ion, and he Adenoid Cys ic Ca cinoma Resea ch Founda ion. Y.S.J and I.M. a e suppo ed by EMBO long- e m ellowship (ALTF 1203-2012 and ALTF 1287-2012, espec i ely). PJC. is a Wellcome T us Senio Clinical Fellow. Suppo was p o ided o AMF by he Na ional Ins i u e o Heal h Resea ch (NIHR) UCLH Biomedical Resea ch Cen e. ARG. ecei es suppo om Leukaemia Lymphoma Resea ch, Cance Resea ch UK, and he Leukemia Lymphoma Socie y. Samples om Addenb ooke's Hospi al we e collec ed wi h suppo om he NIHR Camb idge Biomedical Resou ce Cen e. The ICGC B eas Cance Conso ium was suppo ed by a g an om he Eu opean Union (BASIS) and he Wellcome T us . The ICGC P os a e Cance Conso ium was unded by Cance Resea ch UK. We would also like o acknowledge he suppo o he Na ional Cance Resea ch P os a e Cance : Mechanisms o P og ession and T ea men (PROMPT) collabo a i e (g an code G0500966/75466) which has unded issue and u ine collec ions in Camb idge. This esea ch was suppo ed in pa by he In amu al Resea ch P og am o he NIH, Na ional Ins i u e o En i onmen al Heal h Sciences (JAT.). We ob ained in o med consen and consen o publish om pa icipan s en olled. G oup au ho de ails ICGC B eas Cance G oup Elena P o enzano, Camb idge B eas Uni , Addenb ooke’s Hospi al, Camb idge Uni e si y Hospi al NHS Founda ion T us and NIHR Camb idge Biomedical Resea ch Cen e, Camb idge CB2 2QQ, UK; Ma c an de Vij e , Depa men o Pa hology, Academic Medical Cen e , Meibe gd ee 9, 1105 AZ Ams e dam, The Ne he lands; And ea L Richa dson, Depa men o Cance Biology, Dana-Fa be Cance Ins i u e, 450 B ookline A e., Bos on, Massachuse s 02215, USA; Depa men o Pa hology, B igham and Women's Hospi al, Ha a d Medical School, 75 F ancis S ., Bos on, Massachuse s 02115, USA; Colin Pu die, Eas o Sco land B eas Se ice, Ninewells Hospi al, Dundee, Uni ed Kingdom; Sa ah Pinde , Depa men o Resea ch Oncology, Guy’s Hospi al, King’s Heal h Pa ne s AHSC, King’s College London School o Medicine, London SE1 9RT, UK; Gae an MacG ogan, Ins i u Be gonié, 229 cou s de l’A gone, 33076, Bo deaux, F ance; Anne Vincen -Salomon, Ins i u Cu ie, Depa men o Tumo Biology, 26 ue d’Ulm, 75248 Pa is cédex 05, F ance; Ins i u Cu ie, INSERM Uni 830, 26 ue d’Ulm, 75248 Pa is cédex 05, F ance; Denis La simon , Depa men o Pa hology, Jules Bo de Ins i u e, B ussels 1000, Belgium; Do he G abau, Depa men o Pa hology, Skåne Uni e si y Hospi al, Lund Uni e si y, SE-221 85 Lund, Sweden; To ill Saue , Depa men o Pa hology, Oslo Uni e si y Hospi al Ulle al and Uni e si y o Oslo, Facul y o Medicine and Ins i u e o Clinical Medicine, Oslo, No way; Øys ein Ga ed, Depa men o Pa hology, Oslo Uni e si y Hospi al Ulle al and Uni e si y o Oslo, Facul y o Medicine and Ins i u e o Clinical Medicine, Oslo, No way; Anna Ehinge , Depa men o Gynecology & Obs e ics, Depa men o Clinical Sciences, Lund Uni e si y, Skåne Uni e si y Hospi al Lund, SE-221 85 Lund, Sweden; Ge G Van den Eynden, T ansla ional Cance Resea ch Uni , GZA Hospi als S .-Augus inus, An we p, Belgium; C.H.M. an Deu zen, Depa men o Pa hology, E asmus Medical Cen e , Ro e dam, he Ne he lands; Robe o Salgado, B eas Cance T ansla ional Resea ch Labo a o y, Ins i u Jules Bo de , Uni e si é Lib e de B uxelles, B ussels, Belgium; Jane E B ock, Depa men o Pa hology, B igham and Women's Hospi al, Ha a d Medical School, 75 F ancis S ., Bos on, Massachuse s 02115, USA; Sunil R Lakhani, The Uni e si y o Addi ional in o ma ion Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 22 o 28 Resea ch a icle Queensland, School o Medicine, He s on, B isbane, QLD 4006, Aus alia; Pa hology Queensland: The Royal B isbane & Women’s Hospi al, B isbane, QLD 4029, Aus alia; The Uni e si y o Queensland, UQ Cen e o Clinical Resea ch, He s on, B isbane, QLD 4029, Aus alia; Dilip D Gi i, Depa men o Pa hology, Memo ial Sloan-Ke e ing Cance Cen e , New Yo k, NY, USA; Lau en A nould, Cen e Geo ges-F ançois Lecle c, 1 ue du P o esseu Ma ion, 21079, Dijon, F ance; Jocelyne Jacquemie , 0ns i u Paoli Calme es, biopa hology depa men , 232 Bd S e Ma gue i e, 13009, Ma seille, F ance; Isabelle T eilleux, Cen e Léon Bé a d, Lyon, F ance; Uni e si é Claude Be na d Lyon1 - Uni e si é de Lyon, Lyon, F ance; Ca los Caldas, Camb idge B eas Uni , Addenb ooke’s Hospi al, Camb idge Uni e si y Hospi al NHS Founda ion T us and NIHR Camb idge Biomedical Resea ch Cen e, Camb idge CB2 2QQ, UK; Depa men o Oncology, Uni e si y o Camb idge and Cance Resea ch UK Camb idge Resea ch Ins i u e, Li Ka Shin Cen e, Camb idge CB2 0RE; Sue -Feung Chin, Depa men o Oncology, Uni e si y o Camb idge and Cance Resea ch UK Camb idge Resea ch Ins i u e, Li Ka Shin Cen e, Camb idge CB2 0RE; Aquila Fa ima, Depa men o Cance Biology, Dana-Fa be Cance Ins i u e, 450 B ookline A e., Bos on, Massachuse s 02215, USA; Alas ai M Thompson, Dundee Cance Cen e, Ninewells Hospi al, Dundee, UK; Alasdai S enhouse, Dundee Cance Cen e, Ninewells Hospi al, Dundee, UK; John Foekens, E asmus MC Cance Ins i u e, E asmus Uni e si y Medical Cen e , Ro e dam, The Ne he lands; John Ma ens, E asmus MC Cance Ins i u e, E asmus Uni e si y Medical Cen e , Ro e dam, The Ne he lands; Anie a Sieuwe s, E asmus MC Cance Ins i u e, E asmus Uni e si y Medical Cen e , Ro e dam, The Ne he lands; A jen B inkman, Radboud Uni e si y, Depa men o Molecula Biology, Facul y o Science, Nijmegen Cen e o Molecula Li e Sciences, 6500 HB Nijmegen, The Ne he lands; Henk S unnenbe g, Radboud Uni e si y, Depa men o Molecula Biology, Facul y o Science, Nijmegen Cen e o Molecula Li e Sciences, 6500 HB Nijmegen, The Ne he lands; Paul N. Span, Depa men o Radia ion Oncology, Radboud Uni e si y Medical Cen e, Nijmegen, The Ne he lands; F ed Sweep, Depa men o Labo a o y Medicine, Radboud Uni e si y Medical Cen e, Nijmegen, The Ne he lands; Ch is ine Desmed , B eas Cance T ansla ional Resea ch Labo a o y, Ins i u Jules Bo de , Uni e si é Lib e de B uxelles, B ussels, Belgium; Ch is os So i iou, B eas Cance T ansla ional Resea ch Labo a o y, Ins i u Jules Bo de , Uni e si é Lib e de B uxelles, B ussels, Belgium; Gilles Thomas, Uni e si e Lyon1, INCa-Syne gie, Cen e Leon Be a d, 28 ue Laennec Lyon Cedex 08 F ance; Annegein B oeks, Depa men Expe imen al The apy, The Ne he lands Cance Ins i u e, Plesmanlaan 121, 1066 CX Ams e dam, The Ne he lands; Ani a Lange od, Depa men o Gene ics, Ins i u e o Cance Resea ch, The No wegian Radium Hospi al, Oslo Uni e si y Hospi al, O310 Oslo, No way; Samuel Apa icio, Depa men o Molecula Oncology, BC Cance Agency, 675 W10 h A enue, Vancou e V5Z 1L3; Pe e Simpson, The Uni e si y o Queensland, UQ Cen e o Clinical Resea ch, He s on, B isbane, QLD 4029, Aus alia; Lau a an ' Vee , The Ne he lands Cance Ins i u e, Di ision o Molecula Ca cinogenesis, Ams e dam, The Ne he lands; Depa men o Su ge y, Uni e si y o Cali o nia, San F ancisco, San F ancisco, Cali o nia, Uni ed S a es o Ame ica; Jó unn E la Ey jö d, Cance Resea ch Labo a o y, Facul y o Medicine, Uni e si y o Iceland, Reykja ik, Iceland; Holm idu Hilma sdo i , Cance Resea ch Labo a o y, Facul y o Medicine, Uni e si y o Iceland, Reykja ik, Iceland; Jon G Jonasson, Depa men o Pa hology, Uni e si y Hospi al, Reykja ik, Iceland; Icelandic Cance Regis y, Icelandic Cance Socie y, Skoga hlid 8, P.O.Box 5420, 125, Reykja ik, Iceland; Anne-Lise Bø esen- Dale, Depa men o Gene ics, Ins i u e o Cance Resea ch, The No wegian Radium Hospi al, Oslo Uni e si y Hospi al, O310 Oslo, No way; Ins i u e o Clinical Medicine, Facul y o Medicine, Uni e si y o Oslo; Ming Ta Michael Lee, Na ional Geno yping Cen e , Ins i u e o Biomedical Sciences, Academia Sinica, 128 Academia Road, Sec 2, Nankang, Taipei 115, Taiwan, ROC; Be nice Huimin Wong, NCCS- VARI T ansla ional Resea ch Labo a o y, Na ional Cance Cen e Singapo e, 11 Hospi al D i e, 169610, Singapo e; Beni a Kia Tee Tan, Depa men o Gene al Su ge y, Singapo e Gene al Hospi al, Singapo e; Ge i K.J. Hooije , Depa men o Pa hology, Academic Medical Cen e , Meibe gd ee 9, 1105 AZ Ams e dam, The Ne he lands ICGC Ch onic Myeloid Diso de s G oup Luca Malco a i, Fondazione IRCCS Policlinico San Ma eo, Uni e si y o Pa ia, Pa ia, I aly; Sudhi Tau o, Di ision o Medial Sciences, Uni e si y o Dundee, Dundee, UK; Jacqueline Boul wood, Nu ield Depa men o Clinical Labo a o y Sciences, Uni e si y o Ox o d, UK; And ea Pellaga i, Nu ield Depa men o Clinical Labo a o y Sciences, Uni e si y o Ox o d, UK; Michael G o es, Di ision o Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 23 o 28 Resea ch a icle Medial Sciences, Uni e si y o Dundee, Dundee, UK; Alex S e nbe g, Wea he all Ins i u e o Molecula Medicine, Uni e si y o Ox o d, UK; Depa men o Haema ology, G ea Wes e n Hospi al, Swindon, UK; Ca lo Gambaco i-Passe ini, Depa men o Haema ology, Uni e si y o Milan Bicocca, Milan, I aly; Pa esh Vyas, Wea he all Ins i u e o Molecula Medicine, Uni e si y o Ox o d, UK; E a Hells om- Lindbe g, Depa men o Haema ology, Ka olinska Ins i u e, S ockholm, Sweden; Da id Bowen, S James Ins i u e o Oncology, S James Hospi al, Leeds, UK; Nicholas CP C oss, School o Medicine, Uni e si y o Sou hamp on, Sou hamp on, UK; An hony R G een, Depa men o Haema ology, Uni e si y o Camb idge, Camb idge, UK; Ma io Cazzola, Fondazione IRCCS Policlinico San Ma eo, Uni e si y o Pa ia, Pa ia, I aly ICGC P os a e Cance G oup Colin Coope , Di ision o Gene ics and Epidemiology, The Ins i u e O Cance Resea ch, Su on, UK; Depa men o Biological Sciences and School o Medicine, Uni e si y o Eas Anglia, No wich, UK; Senio P incipal In es iga o s o he Cance Resea ch UK unded ICGC P os a e Cance P ojec ; Rosalind Eeles, Di ision o Gene ics and Epidemiology, The Ins i u e O Cance Resea ch, Su on, UK; Royal Ma sden NHS Founda ion T us , London and Su on, UK; Senio P incipal In es iga o s o he Cance Resea ch UK unded ICGC P os a e Cance P ojec ; Da id Wedge, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Pe e Van Loo, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Human Genome Labo a o y, Depa men o Human Gene ics, VIB and KU Leu en, Leu en, Belgium; Gunes Gundem, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Ludmil Alexand o , Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Ba ba a K emeye , Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Adam Bu le , Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; And ew Lynch, S a is ics and Compu a ional Biology Labo a o y, Cance Resea ch UK Camb idge Resea ch Ins i u e, Camb idge, UK; Sand a Edwa ds, Di ision o Gene ics and Epidemiology, The Ins i u e O Cance Resea ch, Su on, UK; Niedzica Camacho, Di ision o Gene ics and Epidemiology, The Ins i u e O Cance Resea ch, Su on, UK; Cha lie Massie, U ological Resea ch Labo a o y, Cance Resea ch UK Camb idge Resea ch Ins i u e, Camb idge, UK; ZSo ia Ko e-Ja ai, Di ision o Gene ics and Epidemiology, The Ins i u e O Cance Resea ch, Su on, UK; Nening Dennis, Royal Ma sden NHS Founda ion T us , London and Su on, UK; Sue Me son, Di ision o Gene ics and Epidemiology, The Ins i u e O Cance Resea ch, Su on, UK; Jo ge Zamo a, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Jona han Kay, U ological Resea ch Labo a o y, Cance Resea ch UK Camb idge Resea ch Ins i u e, Camb idge, UK; Ca hy Co bishley, Depa men o His opa hology, S Geo ges Hospi al, London, UK; Sa ah Thomas, Royal Ma sden NHS Founda ion T us , London and Su on, UK; Se ena Nik-Zainai, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Sa ah O'Mea a, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Lucy Ma hews, Di ision o Gene ics and Epidemiology, The Ins i u e O Cance Resea ch, Su on, UK; Je emy Cla k, Depa men o Biological Sciences and School o Medicine, Uni e si y o Eas Anglia, No wich, UK; Rachel Hu s , Depa men o Biological Sciences and School o Medicine, Uni e si y o Eas Anglia, No wich, UK; Richa d Mi hen, Ins i u e o Food Resea ch, No wich Resea ch Pa k, No wich, UK; Susanna Cooke, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Kei an Raine, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Da id Jones, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; And ew Menzies, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Lucy S ebbings, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Jon Hin on, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Jon Teague, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; S ua McLa en, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Lau a Mudie, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Clai e Ha dy, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Elizabe h Ande son, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Oli ia Joseph, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Vic o ia Goody, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Ben Robinson, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Ma k Maddison, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; S ephen Gamble, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Ch is ophe G eenman, School o Compu ing Sciences, Uni e si y Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 24 o 28 Resea ch a icle o Eas Anglia, No wich, UK; Dan Be ney, Depa men o Molecula Oncology, Ba s Cance Cen e, Ba s and he London School o Medicine and Den is y, London, UK; S e en Hazell, Royal Ma sden NHS Founda ion T us , London and Su on, UK; Naomi Li ni, Royal Ma sden NHS Founda ion T us , London and Su on, UK; Cy il Fishe , Royal Ma sden NHS Founda ion T us , London and Su on, UK; Ch is ophe Ogden, Royal Ma sden NHS Founda ion T us , London and Su on, UK; Pa deep Kuma , Royal Ma sden NHS Founda ion T us , London and Su on, UK; Alan Thompson, Royal Ma sden NHS Founda ion T us , London and Su on, UK; Ch is ophe Woodhouse, Royal Ma sden NHS Founda ion T us , London and Su on, UK; Da id Nicol, Royal Ma sden NHS Founda ion T us , London and Su on, UK; E ik Maye , Royal Ma sden NHS Founda ion T us , London and Su on, UK; Tim Dudde idge, Royal Ma sden NHS Founda ion T us , London and Su on, UK; Nimish Shah, U ological Resea ch Labo a o y, Cance Resea ch UK Camb idge Resea ch Ins i u e, Camb idge, UK; Vincen Gnanap agasam, U ological Resea ch Labo a o y, Cance Resea ch UK Camb idge Resea ch Ins i u e, Camb idge, UK; Pe e Campbell, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; And ew Fu eal, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Senio P incipal In es iga o s o he Cance Resea ch UK unded ICGC P os a e Cance P ojec ; Douglas Eas on, Cen e o Cance Gene ic Epidemiology, Depa men o Oncology, Uni e si y o Camb idge, Camb idge, UK; Senio P incipal In es iga o s o he Cance Resea ch UK unded ICGC P os a e Cance P ojec ; Anne Y Wa en, Depa men o His opa hology, Camb idge Uni e si y Hospi als NHS Founda ion T us , Camb idge, UK; Ch is ophe Fos e , Bos wick Labo a o ies, London, UK; Senio P incipal In es iga o s o he Cance Resea ch UK unded ICGC P os a e Cance P ojec ; Michael S a on, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Senio P incipal In es iga o s o he Cance Resea ch UK unded ICGC P os a e Cance P ojec ; Hayley Whi ake , U ological Resea ch Labo a o y, Cance Resea ch UK Camb idge Resea ch Ins i u e, Camb idge, UK; Ul an McDe mo , Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Senio P incipal In es iga o s o he Cance Resea ch UK unded ICGC P os a e Cance P ojec ; Daniel B ewe , Di ision o Gene ics and Epidemiology, The Ins i u e O Cance Resea ch, Su on, UK; Depa men o Biological Sciences and School o Medicine, Uni e si y o Eas Anglia, No wich, UK; Da id Neal, U ological Resea ch Labo a o y, Cance Resea ch UK Camb idge Resea ch Ins i u e, Camb idge, UK; Depa men o Su gical Oncology, Uni e si y o Camb idge, Addenb ooke's Hospi al, Camb idge, UK; Senio P incipal In es iga o s o he Cance Resea ch UK unded ICGC P os a e Cance P ojec Funding Funde G an e e ence numbe Au ho Wellcome T us Ul an McDe mo , Michael R S a on, Pe e J Campbell Wellcome T us Heal h Inno a ion Challenge Fund (HICF) Pe e J Campbell Kay Kendall Leukaemia Fund An hony R G een, Mel G ea es, Pe e J Campbell Cho doma Founda ion P And ew Fu eal Adenoid Cys ic Ca cinoma Resea ch Founda ion P And ew Fu eal Eu opean Molecula Biology O ganiza ion ALTF 1203_2012 Young Seok Ju Na ional Ins i u e o Heal h Resea ch Biomedical Resea ch Cen e a Uni e si y College London Hospi als Ad ienne M Flanagan Leukaemia and Lymphoma Resea ch An hony R G een Cance Resea ch UK An hony R G een Leukemia and Lymphoma Socie y An hony R G een Genomics and e olu iona y biology Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 25 o 28 Resea ch a icle Funde G an e e ence numbe Au ho Eu opean Union B eas Cance Soma ic Gene ics S udy (BASIS) Michael R S a on Na ional Cance Resea ch Ins i u e PROMPT: G0500966/75466 Rosalind Eeles, Colin Coope , Da id Neal Na ional Ins i u e o En i onmen al Heal h Sciences In amu al Resea ch P og am o he NIH Jack A Taylo Na ional Ins i u e o Heal h Resea ch Camb idge Biomedical Resea ch Cen e An hony R G een, Da id Neal Eu opean Molecula Biology O ganiza ion ALTF 1287-2012 Inigo Ma inco ena Depa men o Heal h Heal h Inno a ion Challenge Fund (HICF) Pe e J Campbell The unde s had no ole in s udy design, da a collec ion and in e p e a ion, o he decision o submi he wo k o publica ion. Au ho con ibu ions YSJ, Concep ion and design, Acquisi ion o da a, Analysis and in e p e a ion o da a, D a ing o e ising he a icle; LBA, Analyzed mu a ional signa u e; MG, IM, Analysis and in e p e a ion o da a, D a ing o e ising he a icle; SN-Z, MR, HRD, EP, GG, AS, NB, SB, PST, JN, CEM, GSV, ARG, M-QD, AU, JEP, BTT, NM, MG, PV, AKE-N, TS, VPC, RG, JAT, DNH, DM, CSF, AYW, HCW, DB, RE, CC, DN, TV, WBI, GSB, AMF, PAF, AGL, PFC, UMD, Con ibu ed samples and scien i ic ad ice; APB, JWT, P o ided bioin o ma ics suppo o sequencing da a acquisi ion; MRS, PJC, Concep ion and design, Analysis and in e p e a ion o da a, D a ing o e ising he a icle E hics Human subjec s: We ob ained in o med consen and consen o publish om pa icipan s en olled in his s udy, E hical app o al e e ences: Genome Analysis o myeloid and lymphoid malignancies (10/H0306/40), Genomic Analysis o Meso helioma (11/EE/0444), Myeloid and lymphoid cance genome analysis (07/S1402/90), The T ea men o Down Synd ome Child en wi h Acu e Myeloid Leukemia and Myelodysplas ic Synd ome(AAML0431), CLL (ch onic lymphocy ic leukaemia) genome analysis (07/Q0104/3), CGP-Exome sequencing o Down synd ome associa ed acu e myeloid leu- kemia samples (IRB 13-010133), Cance Genome P ojec - Global app oaches o cha ac e izing he molecula basis o paedia ic ependymoma (05/MRE04/70), PREDICT-Coho (09/H0801/96), ICGC P os a e (E alua ion o bioma ke s in u ological diseases) (LREC 03/018), ICGC P os a e (779) (P os a e Complex CRUK Sample Coho ) (MREC/0¼/061), ICGC P os a e (Tissue collec ion a ad- ical p os a ec omy) (CRE-2011.373), Soma ic molecula gene ics o human cance s, melanoma and myeloma (Dana Fa be Cance Ins i u e)(08/H0308/303), B eas Cance Genome Analysis o he In e na ional Cance Genome Conso ium Wo king G oup (09/H0306/36), Genome analysis o umou s o he bone (09/H0308/165). Addi ional iles Supplemen a y iles • Supplemen a y ile 1. Sequencing in o ma ion o 1675 umo –no mal pai s. DOI: 10.7554/eLi e.02935.021 • Supplemen a y ile 2. Ca alogs o soma ic mu a ions (subs i u ions and indels) and inhe i ed polymo phisms iden i ied in his s udy. DOI: 10.7554/eLi e.02935.022 • Supplemen a y ile 3. Lis o phased soma ic subs i u ions. DOI: 10.7554/eLi e.02935.023 • Supplemen a y ile 4. dN/dS o 13 p o ein-coding genes in mi ochond ia. DOI: 10.7554/eLi e.02935.024