scieee Science in your language
[en] (orig)

Origins and functional consequences of somatic mitochondrial DNA mutations in human cancer

Abstract

Recent sequencing studies have extensively explored the somatic alterations present in the nuclear genomes of cancers. Although mitochondria control energy metabolism and apoptosis, the origins and impact of cancer-associated mutations in mtDNA are unclear. In this study, we analyzed somatic alterations in mtDNA from 1675 tumors. We identified 1907 somatic substitutions, which exhibited dramatic replicative strand bias, predominantly C > T and A > G on the mitochondrial heavy strand. This strand-asymmetric signature differs from those found in nuclear cancer genomes but matches the inferred germline process shaping primate mtDNA sequence content. A number of mtDNA mutations showed considerable heterogeneity across tumor types. Missense mutations were selectively neutral and often gradually drifted towards homoplasmy over time. In contrast, mutations resulting in protein truncation undergo negative selection and were almost exclusively heteroplasmic. Our findings indicate that the endogenous mutational mechanism has far greater impact than any other external mutagens in mitochondria and is fundamentally linked to mtDNA replication.

Read accessible full text

Origins and functional consequences of somatic mitochondrial DNA mutations in human cancer

Author: Ju, Seok Young,Alexandrov, Ludmil B,Gerstung, Moritz,Visakorpi, Tapio,Bova, Steve
Year: 2014
Source: https://trepo.tuni.fi/bitstream/10024/99817/1/origins_and_functional_2014.pdf
eli esciences.o g
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 1 o 28
O igins and unc ional consequences o
soma ic mi ochond ial DNA mu a ions in
human cance
Young Seok Ju1, Ludmil B Alexand o 1, Mo i z Ge s ung1, Inigo Ma inco ena1,
Se ena Nik-Zainal1, Manasa Ramak ishna1, Helen R Da ies1, Elli Papaemmanuil1,
Gunes Gundem1, Adam Shlien1, Niccolo Bolli1, Sam Behja i1, Pa ick S Ta pey1,
Jyo i Nangalia1,2,3, Cha les E Massie1,2,3, Adam P Bu le 1, Jon W Teague1,
Geo ge S Vassiliou1,2,3, An hony R G een2,3, Ming-Qing Du2, Ashwin Unnik ishnan4,
John E Pimanda4, Bin Tean Teh5,6, Nikhil Munshi7, Mel G ea es8, Pa esh Vyas9,
Adel K El-Nagga 10, Tom San a ius2, V Pe e Collins2, Richa d G undy11,
Jack A Taylo 12, D Neil Hayes13, Da id Malkin14, ICGC B eas Cance G oup1†,
ICGC Ch onic Myeloid Diso de s G oup1‡, ICGC P os a e Cance G oup1,8,15§,
Ch is ophe S Fos e 16,17, Anne Y Wa en2, Hayley C Whi ake 15, Daniel B ewe 8,18,
Rosalind Eeles8, Colin Coope 8,18, Da id Neal15, Tapio Visako pi19, William B Isaacs20,
G S e en Bo a19, Ad ienne M Flanagan21,22, P And ew Fu eal1,23, Andy G Lynch15,
Pa ick F Chinne y24, Ul an McDe mo 1,2, Michael R S a on1, Pe e J Campbell1,2,3*
1Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, Uni ed
Kingdom; 2Camb idge Uni e si y Hospi als NHS Founda ion T us , Camb idge,
Uni ed Kingdom; 3Depa men o Haema ology, Uni e si y o Camb idge,
Camb idge, Uni ed Kingdom; 4Lowy Cance Resea ch Cen e, Uni e si y o New
Sou h Wales, Sydney, Aus alia; 5Labo a o y o Cance Epigenome, Na ional Cance
Cen e, Singapo e, Singapo e; 6Duke-NUS G adua e Medical School, Singapo e,
Singapo e; 7Depa men o Hema ologic Oncology, Dana-Fa be Cance Ins i u e,
Bos on, Uni ed S a es; 8Ins i u e o Cance Resea ch, Su on, London, Uni ed
Kingdom; 9Wea he all Ins i u e o Molecula Medicine, Uni e si y o Ox o d,
Ox o d, Uni ed Kingdom; 10Depa men o Pa hology, MD Ande son Cance Cen e ,
Hous on, Uni ed S a es; 11Child en's B ain Tumou Resea ch Cen e, Uni e si y o
No ingham, No ingham, Uni ed Kingdom; 12Na ional Ins i u e o En i onmen al
Heal h Sciences, Na ional Ins i u e o Heal h, T iangle, No h Ca olina, Uni ed
S a es; 13Depa men o In e nal Medicine, Uni e si y o No h Ca olina, Chapel Hill,
Uni ed S a es; 14Hospi al o Sick Child en, Uni e si y o To on o, To on o, Canada;
15Cance Resea ch UK Camb idge Ins i u e, Uni e si y o Camb idge, Camb idge,
Uni ed Kingdom; 16Depa men o Molecula and Clinical Cance Medicine, Uni e si y
o Li e pool, London, Uni ed Kingdom; 17HCA Pa hology Labo a o ies, London,
Uni ed Kingdom; 18School o Biological Sciences, Uni e si y o Eas Anglia, No wich,
Uni ed Kingdom; 19Ins i u e o Biosciences and Medical Technology - BioMediTech
and Fimlab Labo a o ies, Uni e si y o Tampe e and Tampe e Uni e si y Hospi al,
Tampe e, Finland; 20Depa men o Oncology, Johns Hopkins Uni e si y, Bal imo e,
Uni ed S a es; 21Depa men o His opa hology, Royal Na ional O hopaedic
Hospi al, Middlesex, Uni ed Kingdom; 22Uni e si y College London Cance Ins i u e,
Uni e si y College London, London, Uni ed Kingdom; 23Depa men o Genomic
Medicine, The Uni e si y o Texas, MD Ande son Cance Cen e , Hous on, Texas,
Uni ed S a es; 24Wellcome T us Cen e o Mi ochond ial Resea ch, Ins i u e o
Gene ic Medicine, Newcas le Uni e si y, Newcas le-upon- yne, Uni ed Kingdom
*Fo co espondence: pc8@
sange .ac.uk
G oup au ho de ails
†ICGC B eas Cance G oup:
See page 21
‡ICGC Ch onic Myeloid Diso de s
G oup: See page 22
§ICGC P os a e Cance G oup:
See page 23
Compe ing in e es s: The
au ho s decla e ha no
compe ing in e es s exis .
Funding: See page 24
Recei ed: 28 Ma ch 2014
Accep ed: 26 Sep embe 2014
Published: 01 Oc obe 2014
RESEARCH ARTICLE
Re iewing edi o : Todd Golub,
B oad Ins i u e, Uni ed S a es
This is an open-access a icle,
ee o all copy igh , and may be
eely ep oduced, dis ibu ed,
ansmi ed, modi ied, buil
upon, o o he wise used by
anyone o any law ul pu pose.
The wo k is made a ailable unde
he C ea i e Commons CC0
public domain dedica ion.
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 2 o 28
Resea ch a icle
Abs ac Recen sequencing s udies ha e ex ensi ely explo ed he soma ic al e a ions p esen in
he nuclea genomes o cance s. Al hough mi ochond ia con ol ene gy me abolism and apop osis,
he o igins and impac o cance -associa ed mu a ions in m DNA a e unclea . In his s udy, we
analyzed soma ic al e a ions in m DNA om 1675 umo s. We iden i ied 1907 soma ic subs i u ions,
which exhibi ed d ama ic eplica i e s and bias, p edominan ly C > T and A > G on he mi ochond ial
hea y s and. This s and-asymme ic signa u e di e s om hose ound in nuclea cance genomes
bu ma ches he in e ed ge mline p ocess shaping p ima e m DNA sequence con en . A numbe o
m DNA mu a ions showed conside able he e ogenei y ac oss umo ypes. Missense mu a ions we e
selec i ely neu al and o en g adually d i ed owa ds homoplasmy o e ime. In con as , mu a ions
esul ing in p o ein unca ion unde go nega i e selec ion and we e almos exclusi ely he e oplasmic.
Ou indings indica e ha he endogenous mu a ional mechanism has a g ea e impac han any
o he ex e nal mu agens in mi ochond ia and is undamen ally linked o m DNA eplica ion.
DOI: 10.7554/eLi e.02935.001
In oduc ion
All cance s esul om soma ic mu a ions in hei genomes. Beyond he ∼3200 Mb o nuclea genomic DNA,
human cells ha e hund eds o housands o mi ochond ia p esen in e e y cell, each ca ying one o a ew
copies o he 16,569 bp ci cula mi ochond ial genomes (Smei ink e al., 2001; Leg os e al., 2004;
eLi e diges The DNA in a cell's nucleus mus be copied ai h ully, and di ided equally, when a
cell di ides o p oduce wo new cells. Mis akes—o mu a ions—a e some imes made du ing he
copying p ocess, and mu a ions can also be in oduced by exposing DNA o damaging agen s
known as mu agens, such as UV ligh o ciga e e smoke. These mu a ions a e hen main ained in all
o he descendan s o he cell. Mos o hese mu a ions ha e no impac on he cell's cha ac e is ics
(‘passenge mu a ions’). Howe e , ‘d i e mu a ions’ ha allow cells o di ide uncon ollably and
sp ead o o he body si es can lead o cance .
Mi ochond ia a e cellula compa men s ha a e esponsible o gene a ing he ene gy a cell
needs o su i e and a e also esponsible o ini ia ing p og ammed cell dea h. Mi ochond ia con ain
hei own DNA—en i ely sepa a e om ha in he nucleus o he cell— ha encodes he p o eins
mos essen ial o ene gy p oduc ion. Mi ochond ial DNA molecules a e equen ly exposed o
damaging molecules called eac i e oxygen species ha a e p oduced by he mi ochond ia.
The e o e, hese eac i e oxygen species ha e been hough o be one o he mos impo an causes
o mi ochond ial DNA mu a ions. In addi ion, because cance cells p oduce ene gy di e en ly o
no mal cells, mu a ions in he mi ochond ial DNA ha change he abili y o he mi ochond ia o
p oduce ene gy ha e been con en ionally hough o help no mal cells o become cance ous.
Howe e , conclusi e e idence o a link be ween cance and mi ochond ial DNA mu a ions is lacking.
Ju e al. examined he mi ochond ial DNA sequences aken om 1675 cance biopsies om o e
hi y di e en ypes o cance and compa ed hese o no mal issue om he same pa ien s. This
e ealed 1907 mu a ions in he mi ochond ial DNA aken om he cance cells. The pa e n o he
mu a ions sugges s ha he majo i y o he mu a ions a e no in oduced om eac i e oxygen
species, bu om he e o s he mi ochond ia hemsel es make in he p ocess o duplica ing hei
DNA when a cell di ides. Unexpec edly, known mu agens, such as ciga e e smoke o UV ligh , had
a negligible e ec on mi ochond ial DNA mu a ions.
Con a y o con en ional wisdom, Ju e al. ound no e idence ha he mi ochond ial DNA
mu a ions help cance o de elop o sp ead. Ins ead, like passenge mu a ions ound in he DNA in
he cell nucleus, mos mi ochond ial genome mu a ions ha e no disce nible e ec . Howe e , Ju e
al. e ealed ha DNA mu a ions ha damage no mal mi ochond ial ac i i y a e less likely o be
main ained in cance cells. P esumably, mi ochond ia con aining hese p o eins p oduce less ene gy,
and so a cell con aining oo many o hese mu a ions will ind i ha de o su i e. This shows ha
ha ing enough co ec ly unc ioning mi ochond ia is essen ial o e en cance cells o h i e.
DOI: 10.7554/eLi e.02935.002
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 3 o 28
Resea ch a icle
Koppenol e al., 2011). In addi ion o hei ole in cellula ene gy balance h ough oxida i e phospho-
yla ion, mi ochond ia a e in ol ed in many essen ial cellula unc ions including modula ion o oxida ion–
educ ion s a us, con ibu ion o cy osolic biosyn he ic p ecu so s, and ini ia ion o apop osis.
Mi ochond ia in euka yo ic cells e ol ed by endosymbiosis om a ee-li ing α-p o eobac e ium (G ay
e al., 1999). O e 2 billion yea s o co-e olu ion, many ances al mi ochond ial genes ha e ans e ed
o he nucleus (Falkenbe g e al., 2007; Cal o and Moo ha, 2010; Wallace, 2012). Wha emains in
he mi ochond ial genome is dis inc i e o he s iking asymme y be ween he wo complemen a y
m DNA s ands in e ms o nucleo ide con en and gene dis ibu ion (And ews e al., 1999). The
hea y (H) s and is guanine- ich (C/G = 0.4) and is he empla e om which mos mi ochond ial
p o eins (12 ou o 13) a e ansc ibed, whe eas only one p o ein-coding gene, MT-ND6, is ansc ibed
om he co espondingly cy osine- ich ligh (L) s and.
Mu a ions in he mi ochond ial genome cause inhe i ed disease (Chinne y, 1993), wi h a ma e nal
inhe i ance pa e n because only eggs con ibu e mi ochond ia o he zygo e. The pene ance o
inhe i ed mi ochond ial disease is de e mined s ochas ically by bo h he andom asso men o
mu a ed s wild- ype mi ochond ial genomes du ing meiosis and andom d i du ing he ea ly cell
di isions a e e iliza ion. In cance , he ole o soma ically acqui ed m DNA mu a ions is con o-
e sial. Al hough cance -speci ic mu a ions ha e been p e iously epo ed (Polyak e al., 1998;
B andon e al., 2006; Cha e jee e al., 2006; He e al., 2010; La man e al., 2012), he limi ed
sample size o poo sensi i i y o capilla y sequencing o he e oplasmic mu a ions has no allowed
a comp ehensi e analysis o he mu a ional signa u es o mi ochond ial mu a ions no hei likely
unc ional signi icance. I has long been p oposed ha mi ochond ia migh con ibu e o cance de el-
opmen gi en hei undamen al impo ance o cellula biology (Wallace, 2012). P e ious epo s
sugges ed ha mi ochond ial soma ic mu a ions migh be unde posi i e selec ion and hus con ibu e
o cance de elopmen , bu he small numbe o epo ed mu a ions ende s his conclusion unce ain
(B andon e al., 2006; Cha e jee e al., 2006; La man e al., 2012; Schon e al., 2012). None heless,
he hypo hesis o unc ionally ele an mi ochond ial mu a ions is an appealing one because cance
cells ha e g ea ly inc eased ene gy demands o e no mal cells and demons a e a swi ch om ae obic
glycolysis in mi ochond ia o lac ic acid e men a ion in he cy osol ( he Wa bu g e ec ) (Hanahan and
Weinbe g, 2011; Koppenol e al., 2011).
In each cell cycle, he eplica ing genome is a isk o de no o mu a ions, which can p omo e he
de elopmen o cance . These mu a ions may be gene a ed by in insic cellula e o s du ing DNA
eplica ion o epai o h ough exposu e o mu agens, such as eac i e oxygen species, obacco
smoke, and ul a iole ligh (Pleasance e al., 2010a, 2010b). Recen ly, >20 mu a ional signa u es
ope a i e in cance s ha e been iden i ied in he nuclea genome (Alexand o e al., 2013). Whe he
any o hese mu a ional p ocesses also a ec he mi ochond ial genome has no been s udied.
Fu he mo e, whe he he e a e m DNA-speci ic mu a ional p ocesses in soma ic cells emain unclea ,
al hough he many unique ea u es o m DNA eplica ion and epai , coupled wi h he high concen a ion
o eac i e oxygen species gene a ed by he elec on anspo chain, could be associa ed wi h dis inc i e
mu a ion signa u es.
In his s udy, we compa e 1675 cance and pai ed no mal m DNA sequences ac oss 31 umo ypes
using massi ely pa allel DNA sequencing echnologies o ob ain a sys ema ic and unbiased ca alog o
soma ic mi ochond ial mu a ions. We ind ha m DNA mu a ions a e almos exclusi ely he p oduc o
a mu a ional p ocess ha is speci ic o mi ochond ia and p obably linked o he unique mechanism
o genome eplica ion hese o ganelles employ. We ind no e idence o posi i e selec ion o
mi ochond ial mu a ions du ing oncogenesis, sugges ing ha hey con e no clonal ad an age on he
nascen cance cells.
Resul s
m DNA sequencing and Mu a ion Calling
We ex ac ed he m DNA sequences om 704 whole-genome and 971 whole-exome sequencing da a
gene a ed on p ima y cance s and compa ed hem wi h m DNA sequences om hei ma ched no mal
samples. Gi en he abundance o m DNA pe cance cell, a s anda d co e age o 30–40× in he
nuclea genome p o ides signi ican ly g ea e co e age o he mi ochond ial genome (a e age ead
dep h = 7901.0×), enabling accu a e iden i ica ion o soma ic mu a ions including a e he e oplasmic
a ian s. We also assessed whe he whole-exome sequencing could be used o iden i y m DNA
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 4 o 28
Resea ch a icle
mu a ions om o - a ge eads de i ed om he mi ochond ial genome. We ound an a e age ead
dep h o 92.1× ac oss he mi ochond ial genome in exome s udies. F om 139 samples in which we
had bo h exome and whole-genome sequencing da a, he o e all ead dep hs co ela ed s ongly
(R2 = 0.59, Figu e 1— igu e supplemen 1) as did a ian allele ac ions o m DNA soma ic mu a ions
(R2 = 0.97, Figu e 1— igu e supplemen 2). Valida ion expe imen s sugges ed he sensi i i y o
whole-exome sequencing o de ec ion o m DNA soma ic mu a ions o be 71.4% compa ed o
whole-genome sequencing (Figu e 1— igu e supplemen 3 and ‘Ma e ials and Me hods’, ‘O - a ge
m DNA eads in whole-exome sequencing’ and ‘DNA c oss-con amina ion’).
To educe po en ial alse-posi i e calls o m DNA soma ic mu a ions, we only epo a ian s called
wi h an allele ac ion o >3%. This elimina es he isk o miscalls due o m DNA-de i ed pseudogenes
in he nucleus (NuMTs) because m DNA copy numbe s a e 100–1000 imes highe han nuclea
genomes in human soma ic cells, and he sequence homology be ween m DNA and NuMTs p esen ed
in he human e e ence genome is gene ally <95% (in 96 ou o 101 NuMTs wi h leng h g ea e han
300 bp). Fu he mo e, pai wise compa ison be ween cance and ma ched no mal m DNAs om he
same indi idual u he minimizes he con amina ion o NuMTs in he mu a ion calling.
The ca alog o m DNA soma ic mu a ions
In o al, 1675 umo –no mal pai s ac oss 31 umo ypes we e analyzed (Table 1 and Supplemen a y
ile 1). Fo 61 o hese pa ien s, we had sequencing da a a ailable om mul iple si es o he p ima y
cance , se e al ime poin s o ma ched p ima y cance s, and me as ases (a o al o 73 such cance
samples), allowing us o s udy he iming o m DNA mu a ions in cance e olu ion (Supplemen a y
ile 1). We iden i ied 1907 soma ic m DNA subs i u ions (Figu e 1 and Supplemen a y ile 2). In
con as o inhe i ed polymo phisms (n = 38,706, a ailable a Supplemen a y ile 2), which we e
almos always homoplasmic in bo h he cance and coun e pa no mal, he a ian allele ac ions
(VAFs) o hese soma ic subs i u ions we e highly a iable in he cance , anging om ou de ec ion
h eshold (3%) o homoplasmy (100%). O hese 1907 soma ic subs i u ions, 1209 (63.4%) we e no
egis e ed in he da abases o m DNA common polymo phism (Ingman and Gyllens en, 2006; Le in
e al., 2013). In compa ison, when we examined subs i u ions ound in bo h he umo and he no mal
samples om a pa ien , only 21 (0.05%) we e no egis e ed in he polymo phism da abases, a
signi ican ly di e en ac ion om he umo -only a ian s (p < 10−10; Chi-squa ed es ). We ound
595 (31.2%) ecu en mu a ions ha can be collapsed on o 246 m DNA posi ions, which is a 6.9- old
highe le el o ecu ence han expec ed by chance (p < 10−10). This sugges s ha he gene a ion
o ixa ion o m DNA mu a ions is no andom, bu in luenced by ac o s such as he unde lying
mu a ional p ocess o posi i e selec ion.
O he 1675 cance samples, 976 (58.3%) ha bo ed a leas one soma ic subs i u ion and 521
(31.1%) had mul iple subs i u ions, anging om 2 o 7 (Figu e 2A). In hose wi h mul iple subs i u-
ions, 72 pai s o mu a ions we e su icien ly close o phase (Nik-Zainal e al., 2012b) such ha we
could de e mine whe he hey we e linked on he same m DNA genome o we e on di e en copies.
We ound ha 45 (62.5%) pai s o mu a ions we e linked on he same m DNA genome (Supplemen a y
ile 3 and Figu e 2— igu e supplemen 1). Fu he mo e, o hese linked mu a ions, 33 showed a
clea empo al o de : ha is, one mu a ion was demons ably sub-clonal o he o he . This is a he
unexpec ed, since each soma ic cell has 100–1000 copies o he mi ochond ial genome, and we migh
an icipa e ha andom mu a ions would, on a e age, a ec di e en copies. Tha many pai s o
mu a ions a e phased on he same m DNA genome and ye show a clea sub-clonal ela ionship
sugges s ha hey occu su icien ly sepa a ed in ime o allow he mi ochond ial genome ca ying he
ea lie mu a ion o d i owa ds a subs an ial ac ion o all genomes in ha cell be o e he second
mu a ion occu s, consis en wi h a p e ious epo (De Alwis e al., 2009).
The numbe o soma ic m DNA subs i u ions a ied signi ican ly acco ding o umo ype (p = 4.4 × 10−52)
a e co ec ing o con ounding a iables such as sequencing co e age: gas ic, hepa ocellula ,
p os a e, and colo ec al cance s had he highes numbe o m DNA subs i u ions (Figu e 2B). In
con as , hema ologic cance s (acu e lymphoblas ic leukemia, myelop oli e a i e disease, and myelo-
dysplas ic synd ome) had ewe mu a ions. Se e al possible explana ions could unde pin hese di e -
ences ac oss umo ypes. I could be ha he mu a ion a es di e ac oss cell lineages; i could be ha
selec ion p essu es shape he numbe o mu a ions; o he numbe o m DNA genome gene a ions
could di e ac oss cell lineages. O hese explana ions, we belie e ha he second is unlikely because,
as we shall see, posi i e selec ion is no a majo componen o mi ochond ial mu a ions. In e es ingly,
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 5 o 28
Resea ch a icle
Table 1. Summa y s a is ics o m DNA sequence da a
WGS WXS
A e age m
RD (WGS)
A e age m
RD (WXS) To al WGS WXS
A e age m
RD (WGS)
A e age m
RD (WXS) To al
B eas 284 98 11594.3 52.7 382 Meningioma 0 12 - 42.5 12
Colo ec al 1 75 34916.9 276.6 76 Ependymoma 1 9 10323.7 52.7 10
Lung 60 0 2798.1 - 60
P os a e 80 0 17810.6 - 80 MPD 12 138 1517.0 10.9 150
Hepa ocellula 0 47 - 205.8 47 MDS 3 75 5648.7 44.5 78
Melanoma 13 13 513.9 353.5 26 ALL 64 6 886.6 35.9 70
Gas ic 0 13 - 184.1 13 CLL 6 0 5002.2 - 6
Cholangioca cinoma 0 8 - 143.9 8 AML 1 6 6783.6 27.4 7
Meso helioma 0 6 - 106.3 6 Mul iple myeloma 0 69 - 43.2 69
Bladde 54 0 646.2 - 54 AMKL 0 9 - 24.2 9
Renal 0 23 - 35.4 23 Lymphoma 0 4 - 99.5 4
O a ian 0 38 - 58.9 38
U e ine 27 23 736.0 149.5 50 Os eosa coma 38 90 9525.5 119.2 128
Ce ical 0 52 - 85.2 52 Chond osa coma 0 47 - 99.1 47
Adenoid cys ic ca. 1 60 714.7 75.6 61 Ewing sa coma 0 27 - 69.5 27
Head & Neck 43 3 1369.1 18.8 46 Kaposi sa coma 0 9 - 181.0 9
Cho doma 16 11 1240.0 82.1 27
To al; 31 cance ypes 704 971 1675
WGS, whole-genome sequencing; WXS, whole-exome sequencing; m RD, mi ochond ial ead dep h; MPD, myelop oli e a i e disease; MDS, myelodysplas ic synd ome; ALL, acu e
lymphoblas ic leukemia; CLL, ch onic lymphoblas ic leukemia; AML, acu e myeloid leukemia; AMKL, acu e megaka yoblas ic leukemia.
DOI: 10.7554/eLi e.02935.003

Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 6 o 28
Resea ch a icle
Figu e 1. Mi ochond ial soma ic subs i u ions iden i ied om 1675 Tumo –No mal pai s. m DNA genes and in e genic egions a e shown. The s and o
genes is shown based on m DNA s and con aining equi alen sequences o ansc ibed RNA. Subs i u ion ca ego ies (silen , non-silen (missense and
nonsense), non-coding ( RNA and RNA), and in e genic) a e shown by he shapes o each subs i u ion. Six classes o subs i u ions a e p esen ed
colo -coded. The subs i u ions on he H, and L s and (when six subs i u ional classes we e conside ed) a e shown ou side and inside o m DNA genes,
espec i ely. Ve ical axes o H and L s and subs i u ions ep esen he VAF o each a ian .
DOI: 10.7554/eLi e.02935.004
The ollowing igu e supplemen s a e a ailable o igu e 1:
Figu e supplemen 1. Co ela ion in amoun o m DNA eads be ween whole-genome and whole-exome sequencing.
DOI: 10.7554/eLi e.02935.005
Figu e supplemen 2. Co ela ion o he e oplasmy le els be ween whole-genome and whole-exome sequencing.
DOI: 10.7554/eLi e.02935.006
Figu e supplemen 3. Valida ion o m DNA soma ic subs i u ions.
DOI: 10.7554/eLi e.02935.007
Figu e supplemen 4. Amoun o o - a ge m DNA eads ac oss ou sequencing cen e s.
DOI: 10.7554/eLi e.02935.008
Figu e supplemen 5. Fil e ing samples o po en ial DNA con amina ions.
DOI: 10.7554/eLi e.02935.009
we ind a posi i e co ela ion be ween he numbe o m DNA soma ic mu a ions and age a diagnosis
in b eas cance s (p = 0.0004; Figu e 2C), in keeping wi h he idea ha he numbe o mi ochond ial
gene a ions is linked o mu a ion bu den. The mu a ional bu den o an es ablished cance ep esen s
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 7 o 28
Resea ch a icle
he accumula ed a ia ion acqui ed in he lineage o cell di isions om e ilized egg o ans o med
cell and will include e en s acqui ed in no mal de elopmen and homeos asis as well as hose acqui ed
du ing umo igenesis (S a on e al., 2009). In e es ingly, m DNA mu a ions ha e been ound a high
Figu e 2. m DNA soma ic subs i u ions o human cance . (A) Numbe o soma ic subs i u ions in a umo sample. (B) A e age numbe o soma ic
subs i u ions pe sample ac oss 31 umo ypes. (C) Age o diagnosis and numbe o m DNA soma ic subs i u ions in b eas cance s.
DOI: 10.7554/eLi e.02935.010
The ollowing igu e supplemen is a ailable o igu e 2:
Figu e supplemen 1. VAFs o phased soma ic m DNA subs i u ions.
DOI: 10.7554/eLi e.02935.011
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 8 o 28
Resea ch a icle
a es in no mal colonic c yp cells (Taylo e al., 2003; E icson e al., 2012). Gi en ha we ind high
bu dens o mu a ions in colonic umo s as well, he di e ences we see ac oss umo ypes may a ise
om p e- o pos - ans o ma ion di e ences in m DNA bu den ac oss issues.
Ex ac ing m DNA mu a ional signa u es
Wi h espec o signa u es o soma ic subs i u ions, C > T and T > C ansi ions cons i u ed 90.9% o
all he 1907 subs i u ions (Figu e 1) among he six classes o possible base subs i u ions. To cha ac-
e ize his agg ega ed signa u e o m DNA cance speci ic mu a ions in mo e de ail, we looked o he
p esence o m DNA s and bias be ween he complemen a y H and L s ands o m DNA. The wo main
subs i u ion classes showed an ex eme le el o m DNA s and bias. 84.1% o he C > T ansi ions
we e on he H s and. This le el o s and bias occu ed despi e he ac ha cy osine is 2.4- old less
common on he H han he L s and, so he C > T subs i u ion a e is 12.6- old highe on he H s and.
By con as , 76.8% o he T > C ansi ions we e on he L s and despi e i s lowe hymine con en (1.3-
old less han he H s and). This implies ha he T > C mu a ion a e on he L s and is 4.2- old highe
han on he H s and.
We hen examined he sequence con ex in which hese mu a ions occu ed by examining he
bases immedia ely 5′ and 3′ o he mu a ed bases. This gene a es 96 possible mu a ion classes ( he 6
subs i u ion classes mul iplied by he 16 combina ions o immedia e 5′ and 3′ nucleo ides). Bo h C > T
and T > C mu a ions showed highly dis inc i e sequence con ex s. CH > TH subs i u ions (i.e. C > T
mu a ions on he H s and) we e en iched o he NpCpG inucleo ide con ex (8- o 15- old mo e
equen han expec ed by chance; Figu e 3A). By con as , TL > CL subs i u ions (i.e. T > C mu a ions
on he L s and) showed 5- o 8- old en ichmen in NpTpC. This s and-asymme ic mu a ional signa-
u e is no simila o any o he 21 cance -associa ed mu a ional signa u es ecen ly iden i ied om he
nuclea DNA o 30 di e en cance ypes (Alexand o e al., 2013).
O he 18 umo ypes ha p esen ed a leas 25 m DNA soma ic subs i u ions in his s udy, he mu-
a ional signa u es we e b oadly consis en ac oss umo ypes (Figu e 3B), wi h he excep ion ha
mul iple myeloma had a somewha highe a e o TH > CH changes han o he his ologies (p = 8.1 × 10−6).
Thus, in con as o he mu a ional signa u es ound in nuclea genomes, whe e he e is s iking he e o-
genei y bo h ac oss umo ypes and ac oss indi iduals wi hin a umo ype (Alexand o e al., 2013),
he mu a ional p o ile in he mi ochond ial genome o soma ic cells is ema kably homogeneous.
Replica ion-coupled mu a ional p ocess in mi ochond ia
The majo known cause o mu a ional s and bias in nuclea DNA is ansc ip ion-coupled nucleo ide
excision epai , whe e DNA lesions on he ansc ibed (non-coding) s and a e mo e equen ly epai ed
(Alexand o e al., 2013). Howe e , we ind ha he s and bias always a o s CH > TH and TL > CL
whe he he gene is ansc ibed om he H s and o om he L s and (Figu e 3— igu e supple-
men 1). This is no compa ible wi h ansc ip ion-coupled epai , o which he di ec ion o s and
bias is undamen ally dic a ed by which s and is ansc ibed.
Ins ead, he m DNA mu a ional s and bias epo ed he e appea s o be d i en by di e ences in
eplica ion be ween he wo s ands. m DNA eplica ion ha bo s subs an ial s and asymme y
be ween he H and L s ands: m DNA eplica ion ini ia es om an o igin o eplica ion (OH) in he
D-loop, wi h he nascen H and he L s and eplica ing as leading and lagging s and, espec i ely
(Clay on, 1982; Falkenbe g e al., 2007; Hol and Reyes, 2012). We obse ed ha C > T subs i u-
ions we e p e alen in he leading (hea y) s and, whe eas T > C subs i u ions we e ound in he lag-
ging (ligh ) s and (Figu e 1). Rema kably, his s and bias was e e sed in he D-loop i sel (Figu es 1
and 3C), u he sugges ing ha he m DNA soma ic mu a ions a e eplica ion-coupled: acco ding o
a ecen ly p oposed bidi ec ional model o m DNA eplica ion (Yasukawa e al., 2005, 2006; Hol
and Reyes, 2012), m DNA eplica ion is also able o ini ia e om he so-called O i-b si e, ypically
loca ed a ound genomic posi ion 16,197 and p oceeds on bo h s ands away om he o igin (Figu e 1).
Replica ion o he nascen H s and con inues unimpeded like he adi ional model, bu he nascen L
s and e mina es a he so-called OH si e, ypically a ound m DNA posi ion 191 bp. Unde his model,
hen, he leading and lagging s and a e e e sed in he ew hund ed base-pai s o he D-loop, which
is consis en wi h he e e sed mu a ional signa u e in his egion (Figu es 1 and 3C).
Equi alen mu a ional signa u e du ing human m DNA E olu ion
I is no en i ely s aigh o wa d o in e he mu a ional signa u es ope a ing on he mi ochond ial
genome in he ge mline. De no o mu a ions a e gene ally a e and o en disco e ed because hey
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 9 o 28
Resea ch a icle
Figu e 3. Replica i e s and bias o m DNA soma ic subs i u ions. (A) Replica i e s and-speci ic subs i u ion a e (# o obse ed/# o expec ed) by 96
inucleo ide con ex . Subs i u ions in a speci ic m DNA segmen ( om O i-b o OH) a e no included, because hey p esen a di e en subs i u ional
signa u e. (B) Mu a ional signa u e ac oss umo ypes. Eigh een umo ypes, which include a leas 25 m DNA mu a ions, we e shown. (C) In e ed
subs i u ion signa u e in he O i-b–OH.
DOI: 10.7554/eLi e.02935.012
The ollowing igu e supplemen is a ailable o igu e 3:
Figu e supplemen 1. Replica i e s and bias obse ed in m DNA subs i u ions.
DOI: 10.7554/eLi e.02935.013
cause disease; dis inguishing he ances al base and he de i ed base is challenging o single nucleo-
ide polymo phisms; and compa a i e m DNA genomics ac oss species ex ends o e conside able
e olu iona y ime. In con as , because ances al and de i ed s a es a e de ined o umo –no mal
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 16 o 28
Resea ch a icle
allele equency is ∼50% (45–55%) acco ding o The 1000 Genomes P ojec (1000 Genomes
P ojec Conso ium e al., 2010). O he 320 si es, homozygous posi ions in no mal issues
(which showed >90% a ian allele ac ion (VAF) wi h bases Q sco e >20) we e compa ed
wi h he co esponding geno ypes in he coun e pa cance . Sample pai s we e emo ed i he
geno ype misma ch a e was g ea e han 0.1 (
Nhe Nw
Nhom Nhe Nw
+
++
; Nhe , numbe o he e ozygo e
posi ions; Nhom, numbe o homozygo e posi ions; Nw , numbe o wild- ype posi ions)
(Figu e 1— igu e supplemen 5A). We no e 0 is expec ed o he a e when geno yping is
pe ec and sample pai s a e om he same indi idual. By con as , 0.5 is expec ed when
samples we e om di e en indi iduals.
2. Mino c oss-con amina ion
We es ima ed DNA c oss-con amina ion le els wi h he VAF o au osomal homozygous SNPs
geno yped om he common (popula ion mino allele equency ∼50%) SNP si es. Theo e ically,
i he e is no sequencing (and mapping) e o , all he homozygo e SNP si es in pu e samples
should p esen 100% VAFs. Howe e , when samples a e con amina ed, co esponding VAFs a e
educed because he con aminan has only an ∼25% o chance o ha ing homozygo e SNPs on
he same si e. The e o e, mino con amina ion le els (C) o each cance sequencing da a we e
es ima ed as below:
∑
∑
RCw Ne
C
RDhom Ne
= × ( )–
2,
( )–
whe e RDhom is sequencing ead dep h, RCw is ead-coun o wild- ype alleles, and Ne is numbe
o sequencing e o s on each au osomal homozygo e SNP si e. Fo high accu acy, we only
coun ed base wi h su icien quali y sco e (Q > 20). In o de o es ima e Ne, we assumed a
conse a i e a e (sequencing e o a e = 0.001). We conside ed si es co e ed by a leas
10 eads and 90% VAF (Figu e 1— igu e supplemen 5B). 95% con idence in e als o c oss-
con amina ion le els we e calcula ed using binomial dis ibu ion.
In o de o clea soma ic a ian s, he e we made he e y conse a i e assump ion ha soma ic
a ian s p esen in excess o 5- imes o he 95% uppe limi o C le els we e ue soma ic a he
han alse-posi i es by low-le el o c oss-con amina ion.
3. Ge mline polymo phisms and back mu a ions
We u he checked samples o con amina ion using known m DNA polymo phisms. Because
human m DNA is small (16,569 bp) and ex ensi ely explo ed p e iously, mos o ge mline
m DNA polymo phisms a e al eady known. Fo example, 97.7% o he 39,036 inhe i ed subs i u-
ions we e known polymo phisms in he m DB da abase (Ingman and Gyllens en, 2006).
The e o e, when a umo sample is con amina ed by o he samples, many soma ic-like m DNA
subs i u ions by con aminan s a e likely o be o e lapped wi h known m DNA polymo phisms.
A he same ime, low-le el con amina ion would gene a e excessi e back mu a ions, which
appea ed o e e se ge mline common polymo phisms in o wild- ype alleles. Taken oge he ,
bo h he numbe o soma ic subs i u ions known in m DB and numbe o back mu a ions can be
good indica o s o m DNA c oss-con amina ion. The e o e, we il e ed ou umo issues wi h
≥3 known po en ially soma ic mu a ions o wi h ≥2 back mu a ions om he u he analyses
(Figu e 1— igu e supplemen 5A and B).
Va ian calling
We ex ac ed m DNA eads using Sam ools (Li and Du bin, 2009). We used Va Scan2 (Kobold e al.,
2012) o ini ial a ian calling wi h a ew op ions (--s and- il e 1 (misma ches should be epo ed by bo h
o wa d and e e se eads), --min- a - eq 0.03 (minimum VAF 3%), --min-a g-qual 20 (minimum base
quali y 20), --min-co e age 3 and --min- eads2 2). Wi h espec o he --s and- il e , i gene ally emo es
a ian when >90% o misma ches a e epo ed om ei he o he H o he L m DNA s and. Howe e ,
whe e only eads wi h a speci ic o ien a ion a e could be aligned dominan ly (i.e. in bo h ex eme
egion o mi ochond ial e e ence genome; only L s and eads could be aligned on he 5′ ex eme o
m DNA), we compa ed s and bias be ween ‘pe ec ma ches’ (# pe ec ma ches om L s and

Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 17 o 28
Resea ch a icle
eads / o al # pe ec ma ches) and misma ches (# misma ches om L s and eads / o al # misma ches).
I he di e ence be ween hose wo bias <0.1, he mu a ions we e escued. O he 1907 mu a ions,
54 (2.8%) we e escued acco dingly.
Pu a i e soma ic a ian s called by Va Scan2 we e u he il e ed using c i e ia shown below.
1. A leas 4 unique eads suppo ing a ian s and all a ian eads a leas 20 ph ed scale sequencing
quali y sco e (Q 20 = 1% sequencing e o a e) and a leas 3% a ian allele ac ions (VAFs).
A. Rega dless o in WGS and in WES, he ≥4 misma ches and he ≥3% VAF c i e ia mus be sa is-
ied simul aneously.
B. Howe e , in WGS, he minimum numbe o eads (n = 4) c i e ion is no essen ial, because he
≥3% VAF c i e ion is much mo e s ingen (3% VAF eques a leas 240 misma ches (>>4)
gi en m DNA co e age is ∼8000 o WGS).
C. In WES, he ≥3% VAF c i e ion is ela i ely less impo an han in WGS, because he ≥4 mis-
ma ches c i e ion is mo e s ingen . Fo example, 4 misma ches in 90x (WXS a e age) co e -
age egion (VAF = 4.4%) au oma ically ul ill he ≥3% VAF c i e ion. Fo less co e ed egions
(i.e. <40x co e age; n = 285 ou o o al 1907 subs i u ions), he VAF c i e ion becomes less
impo an , because 4 misma ches would gene a e ≥10% VAF, much highe han he minimum
h eshold (i.e. 3%). As esul s, we a e missing lowe he e oplasmic a ian s (i.e. a ian s wi h
3–10% he e oplasmic le els) om low co e age samples (mos ly by WXS). The lowe sensi i i y
o WXS is also con i med in ou alida ion s udy (see “Valida ion o soma ic a ian s” below).
2. The e is no minimum h eshold o o al co e age (# pe ec ma ches + # misma ches).
3. To inc ease sensi i i y o de ec ing mu a ions, we escued mu a ions wi h 3 unique a ian eads
(wi h a leas 20 ph ed scale sequencing quali y sco e) when VAFs is ≥ 20%. O 1907 soma ic sub-
s i u ions, 32 (1.7%) we e escued acco dingly.
4. All soma ic a ian s p esen ing wi h VAFs lowe han ou e y conse a i e h eshold o mino
c oss-con amina ion (5- imes 95% uppe limi o con amina ion le els o each umo sample, see
abo e “Mino c oss-con amina ion o DNA samples”) we e emo ed. When we could no es ima e
c oss-con amina ion le els because o low sequencing dep h o co e age ( o nuclea genome), a
conse a i e c i e ion (10% con amina ion le el h eshold) was explici ly used.
5. Subs i u ions we e u he isually inspec ed using IGV (Tho aldsdo i e al., 2013). Thi een
equen alse-posi i e a ian s (shown below) by misalignmen due o ex ensi e le el o homopoly-
me s in CRS and due o sequencing e o in he e e ence m DNA genome (3107N, see Mi omap
(h p://www.mi omap.o g/bin/ iew.pl/MITOMAP/Camb idgeReanalysis) o mo e in o ma ion)
we e explici ly emo ed:
1. Misalignmen due o ACCCCCCCTCCCCC ( CRS 302-315)
A302C, C309T, C311T, C312T, C313T, G316C
2. Misalignmen due o GCACACACACACC ( CRS 513-525)
C514A, A515G, A523C, C524G
3. Misalignmen due o 3107N in CRS (ACNTT, CRS 3105-3109)
C3106A, T3109C, C3110A
We compa ed ou a ian calls wi h common inhe i ed m DNA polymo phisms deposi ed in he m DB
da abase as o 24 h July 2013 (Ingman and Gyllens en, 2006). Gene anno a ion o soma ic a ian s
was done using cus om sc ip based on human m DNA gene in o ma ion (Ruiz-Pesini e al., 2007).
Valida ion o soma ic a ian s
To alida e he sensi i i y and speci ici y o a ian calling in his s udy, 19 umo and no mal pai s (which
we e o iginally whole-genome sequenced) we e whole-exome sequenced and m DNA a ian s we e
assessed independen ly. Among he 28 soma ic subs i u ions o iginally de ec ed om he 19 umo –
no mal whole-genome sequencing pai s, 20 (71.4%) we e called as soma ic (Figu e 1— igu e supple-
men 3). In addi ion, 5 (17.9%) p esen ed e idence o a ian eads in he alida ion se , al hough i
was il e ed ou because o i s low ead dep h o co e age in exome sequencings (showed 2–5 a ian
eads). Mo eo e , because 3 emaining si es we e no su icien ly co e ed in he alida ion se o call
soma ic a ian s, hese could no be e idence o he inaccu acy o whole-genome sequencing da a,
he e o e no conside ed in he accu acy alida ion. Taken oge he , all he 25 soma ic subs i u ions by
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 18 o 28
Resea ch a icle
whole-genome sequencing we e highly likely o be ue posi i es, he e o e we concluded i p o ided
∼100% accu acy in he m DNA soma ic subs i u ion assessmen . Ac ually, he high accu acy o whole-
genome sequencing is e y likely and wha we expec , because i p o ides ex ensi e co e age o
m DNA (a e age ead dep h >7,500×), ∼3% he e oplasmic a ian s would p esen >200 a ian eads.
By con as , he alida ion se (whole-exome sequencing) is called 21 soma ic subs i u ions. O
hese, 20 we e common wi h whole-genome sequencing, and one was inco ec ly called as soma ic
hough i was ac ually ge mline subs i u ions in he whole-genome sequencing da a. In addi ion, as
men ioned abo e, he alida ion se missed 8 soma ic subs i u ions called by whole-genome sequencing.
Six ou o eigh unde calls (75%) we e low he e oplasmic subs i u ions in whole-genome sequencing,
anging om 3.36% o 8.68%. Based on hese da a, we sugges 71.4% sensi i i y (20/28) and
95.2% speci ici y (20/21) o exome-sequencing in de ec ing up o 3% he e oplasmic soma ic m DNA
subs i u ions in cance .
We u he checked he co ela ion o he e oplasmy le el be ween he 20 m DNA soma ic mu a ions
called bo h whole-genome and whole-exome sequencing. I showed g ea linea ela ionship (R2 = 0.97,
Figu e 1— igu e supplemen 2), u he sugges ing whole-exome sequencing da a a e app op ia e
o accu a e de ec ion o m DNA soma ic mu a ions.
Subs i u ion phasing
We phased 72 soma ic subs i u ion pai s, which a ose in a single cance sample and which loca ed
su icien ly close ( om 10 bp o ∼500 bp), he e o e bo h si es could be sequenced by same sequence
agmen s (Supplemen a y ile 3 and Figu e 2— igu e supplemen 1). We classi ied hem as
‘di e en s and’, ‘co-clonal’, and ‘sub-clonal’ using c i e ia as ollows:
Di e en s and: he wo soma ic subs i u ions a e obliga e on di e en s ands. Reads ha epo
wild- ype1(w )-subs i u ion2(subs) and subs1-w 2, bu subs1-subs2, a e obse ed.
Co-clonal: eads epo ing w 1-w 2 and subs1-subs2 a e only obse ed.
Sub-clonal: One subs i u ion is sub-clonal o he o he , bu he wo a e de ini ely phased. Reads
subs1-subs2 and ei he subs1-w 2 o w 1-subs2 a e obse ed.
Tumo ype and m DNA soma ic subs i u ions
To unde s and he ela ionship be ween umo ypes and numbe o m DNA mu a ions, Poisson e-
g ession and ANOVA we e applied o ou da ase using R so wa e (h p://www. -p ojec .o g).
sub T N
Fi 1 < - glm(N Co + Co , amily = poisson())~
sub T N
Fi 2 < glm(N Co + Co + , amily = poisson())-~
ano a(Fi 1, Fi 2, es = Chisq ),“”
whe e Nsub is numbe o m DNA subs i u ions o each sample, Co T and Co N a e co e age o umo
and no mal m DNA, espec i ely, (i Co is >200, we eplaced i by 200), is umo ypes.
Age and m DNA soma ic subs i u ions
Poisson eg ession was applied o ou b eas cance da ase .
sub T N
Fi 1 < glm(N Co + Co + a, amily = poisson()),-~
whe e Nsub is numbe o m DNA subs i u ions o each sample, Co T and Co N a e co e age o umo
and no mal m DNA, espec i ely, (i Co is >200, we eplaced i by 200), a is age a diagnosis. p- alue
in es ima ion o a was shown in he manusc ip .
Mu a ional signa u e and s and bias
Di e en mu a ional p ocesses gene a e di e en combina ions o mu a ion ypes, e med ‘signa u es’
(Nik-Zainal e al., 2012a). Fo example, ul a iole (UV) ligh and obacco smoking (polycyclic a oma ic
hyd oca bons) equen ly gene a e C > T ansi ions and G > T ans e sions on non- ansc ibed
(coding) s ands in melanoma and lung cance s, espec i ely (Pleasance e al., 2010a, 2010b).
To unde s and he mu a ional p ocesses in luencing cance m DNA, we co ela ed he 1907 m DNA
subs i u ions wi h 21 cance speci ic mu a ional signa u es in he nuclea DNA ecen ly iden i ied
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 19 o 28
Resea ch a icle
(Alexand o e al., 2013). Howe e , none o he signa u e could explain he highly unique m DNA
subs i u ions.
Mu a ional signa u e and s and bias we e assessed as desc ibed in ou p e ious epo s (Alexand o
e al., 2013). B ie ly, he immedia e 5′ and 3′ sequence con ex was ex ac ed om CRS. Subs i u ion
a e o each inucleo ide con ex was calcula ed wi h he numbe o subs i u ion no malized by
he equency o he inucleo ide con ex obse ed in he CRS, in he L and H s and, espec i ely.
Fo analyses o subs i u ions alling in he m DNA genes (13 p o ein-coding and 22 RNA genes),
ansc ibed/non- ansc ibed s and was also conside ed o compa ison.
In o de o p o e he s and bias is no ansc ip ion bu eplica ion-coupled, we checked s and
biases o polymo phisms in he 12 L s and p o ein-coding genes, 1 H s and p o ein-coding gene
(MT-ND6), and/o 22 RNAs (Figu e 3— igu e supplemen 1). Fo his speci ic pu pose, we did no
conside he sequence con ex (immedia e 5′ and 3′ bases) because i o e -classi ies mu a ions
(i.e. he numbe o mu a ion classes (n = 96) is la ge han ha o mu a ions). In o he wo ds, 12 classes
o subs i u ions (six classes o possible base subs i u ions (C > A, C > G, C > T, T > A, T > C, T > G) ×
wo s ands (L and H s ands)) we e conside ed. Subs i u ion a es a e a io be ween obse ed and
expec ed numbe s (H0 = same mu a ion a e o all subs i u ion classes) o each subs i u ion class.
In o de o unde s and which model ( eplica i e o ansc ip ional s and) is app op ia e o explain he
s and-bias, Chi-squa e es s we e used be ween he numbe o obse ed mu a ions o each class and
expec ed ones unde he backg ound signa u e.
m DNA codon usage
We coun ed he codon equencies in 13 m DNA p o ein-coding genes. Because 12 L s and p o ein-
coding genes and 1 H s and gene (MT-ND6) a e unde opposi e mu a ional p essu e (T > C and G >
A o L s and genes; A > G and C > T o MT-ND6), we sepa a ed L and H s and genes o his
analysis. T > C skew and G > A skew we e calcula ed as shown below, o unde s and he TL > CL and
CH > TH (equi alen o GL > AL) subs i u ions du ing he e olu ion o human m DNA:
CT
AG
skew skew
CT AG
––
TC = GA = ,
++
NN NN
and
NN NN
>>
whe e NA, NC, NG, and NT a e numbe o A, C, G, and T base in he 3 d posi ion o iple codons in
m DNA genes, espec i ely.
Fo he assessmen o m DNA codon usage o o he animal species, we analyzed he m DNA
sequence o Caeno habdi is elegans (accession# NC_001328), D osophila melanogas e (accession#
NC_001709), D. e io (accession# NC_002333), Xenopus lae is (accession# NC_001573), Mus musculus
(accession# EU450583), Gallus domes icus (accession # NC_235570), and Pan oglody es (NC_001643).
We conside ed only L s and m DNA genes in he c oss-species analysis.
Recu en subs i u ions
To compa e he numbe o ecu en subs i u ions be ween silen and missense subs i u ions, we
andomly selec ed 100 subs i u ions each om 198 silen subs i u ions in he hi d base o iple
codons, 440 missense subs i u ions in he i s base o iple codons, and 405 missense subs i u ions
in he second base o iple codons. We coun ed he numbe o ecu en subs i u ions in each g oup.
This was i e a ed 300 imes independen ly. ANOVA es ing was applied o de e mine he di e ence
be ween he h ee g oups (Figu e 5— igu e supplemen 1).
dN/dS a io
To es ima e dN/dS alues o missense mu a ions (wmis), we used an adap a ion o he me hod
desc ibed p e iously (G eenman e al., 2006). B ie ly, he a e o mu a ions is modeled as a Poisson
p ocess, wi h a a e gi en by a p oduc o he mu a ion a e and he impac o selec ion. To ob ain
accu a e es ima es o dN/dS, we used wo sepa a e models, one using 12 single-nucleo ide subs i u ion
a es and a mo e complex one accoun ing o any con ex dependence e ec by 1-nucleo ide ups eam
and downs eam using 192 subs i u ion a es. Fo example in he 12- a e model, he expec ed numbe
o A > C mu a ions (λA>C) would be modeled as ollows:
syn,A C A C syn,A C
= L
>> >
λ*
mis,AC AC mis mis,AC
= w L ,
>> >
λ **
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 20 o 28
Resea ch a icle
whe e Lsyn,A>C and Lmis,A>C a e he numbe o si es ha can su e a synonymous and missense A > C
mu a ion, espec i ely, which a e calcula ed o any pa icula sequence. The likelihood o obse ing
he numbe o missense A > C mu a ions (Nmis,A>C) gi en he expec ed λmis,A>C is hen calcula ed as:
mis,A C A C mis
Lik = Poisson(N ,w )
>>
|
and he likelihood o he en i e model is he p oduc o all indi idual likelihoods. Wmis is ixed o be
equal in all 12 (o 192) equa ions desc ibing each subs i u ion ype, and a hill-climbing algo i hm is
used o ind he maximum likelihood es ima es o all a e and selec ion pa ame e s. Likelihood Ra io
Tes s a e hen used o es de ia ions om neu ali y (wmis = 1). The dN/dS a io epo ed in he main
ex co esponds o he ull con ex dependen model wi h 192 subs i u ion a es. This me hod allows
quan i ying he s eng h o selec ion a oiding he con ounding e ec o gene leng h, sequence com-
posi ion, di e en a es o each subs i u ion ype, and con ex -dependen mu agenesis.
Sho indels
Along wi h he 1907 soma ic m DNA subs i u ions, we iden i ied 109 and 142 soma ic sho inse ions
and dele ions, espec i ely, om he 1675 cance m DNA sequences using Va scan2 (Supplemen a y
ile 2).
E olu iona y dynamics o neu al mi ochond ial mu a ions
We model he e olu iona y dynamics o mi ochond ial mu a ions unde andom d i and de i e a
simple equa ion o he expec ed numbe o homoplasmic mu a ions. The e exis mul iple le els a
which mi ochond ial mu a ions e ol e: wi hin mi ochond ia, in he cy oplasm, and on he cellula
le el (Rand, 2011). In his s udy, we ocus on he dynamics in a single cell, which ep esen s he
ounde o he las clonal expansion in he umo cell popula ion. The cellula dynamics du ing a
clonal expansion is di icul o desc ibe analy ically, bu i is impo an o ealize ha mu a ions o a
clonal expansion p ese es he allele equencies o neu al a ian s and ha mu a ions ha occu
a e he expansion a e unlikely o con ibu e o measu able allele equencies, as he popula ion
becomes la ge.
We model he e olu iona y dynamics o mi ochond ial mu a ions in he cy oplasm o a single cell
by a W igh –Fishe p ocess (W igh , 1931), in which he numbe o mi ochond ia in a subsequen
gene a ion is a binomial sample o he mi ochond ia in he p e ious gene a ion. The numbe o
mi ochond ia M is kep ixed. The ma ginal allele equency X o a single si e has wo abso bing
bounda ies, X = 0 and X = M (homoplasmy), and he p obabili y o ixa ion o an allele a equency
X by neu al d i is
ρ
= X/M (W igh , 1931). No e ha his p ocess leads, on he popula ion le el,
o a dicho omiza ion o he e oplasmic a ian s o ei he go ex inc o become homoplasmic and ixa e
in a cell.
Mu a ions on any o L (= 16,569 n ) si es in he mi ochond ial genome a e assumed o occu a a
uni o m a e
μ
pe nucleo ide pe cell di ision, which is o o de 10−7, based on a human in e -gene a ional
compa ison (Colle e al., 2001). Hence he a e o neu al e olu ion is simply
μ
LM/M =
μ
L (Kimu a,
1984). Las ly, he expec ed ime o ixa ion in he W igh –Fishe p ocess is = 2M. Pu ing hese hings
oge he , he expec ed numbe o mu an alleles N in a cell ini ially wi hou any mi ochond ial mu a-
ions a e T gene a ion is
E[ ] = ( – )
N L T 2M
µ
This equa ion p edic s a linea accumula ion o neu al mu a ions o e ime, wi h a delay imposed
by numbe o mi ochond ial copies. A simila beha io has been epo ed using nume ical simula ions
(Colle e al., 2001). When also conside ing he e oplasmic mu a ions, he expec ed numbe o al e a ions
may be sligh ly highe .
To check whe he ou model yields he co ec beha io , we use he ollowing numbe s: he
obse ed o de o magni ude o mi ochond ial mu a ions pe pa ien was N = 1. The sequencing co -
e age on he mi ochond ial genome indica es ha he e we e o o de M = 100 mi ochond ial genome
copies p esen pe cance cell. The expec ed numbe o mu a ions pe cell di ision is
μ
L = 1.6 × 10−3,
i he e o e equi es a ound 1000 cell gene a ions T o accumula e on a e age one homoplasmic
mu a ion. This numbe o gene a ions appea s ealis ic o egene a ing issues. As expec ed, epi helial
cance s had among he highes obse ed numbe o mi ochond ial mu a ions, while hema opoie ic
cance s ypically had lowe numbe s.
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 21 o 28
Resea ch a icle
S a is ical es ing
S a is ical es ing was pe o med using R so wa e. All p- alues we e calcula ed by wo- ailed es ing.
Figu es we e gene a ed using R and Mic oso Excel so wa e.
Acknowledgemen s
Da a used in his manusc ip a e desc ibed in he supplemen a y ma e ials (Supplemen a y ile 1). We
hank Thomas Bleaza d a Facul y o Medical and Human Sciences, Uni e si y o Manches e o dis-
cussion and assis ance wi h manusc ip p epa a ion. We would like o hank The Cance Genome A las
(TCGA) P ojec Team and hei specimen dono s o p o iding sequencing da a. This wo k was sup-
po ed by he Wellcome T us , he B i ish Lung Founda ion, he Heal h Inno a ion Challenge Fund, he
Kay Kendall Leukaemia Fund, he Cho doma Founda ion, and he Adenoid Cys ic Ca cinoma Resea ch
Founda ion. Y.S.J and I.M. a e suppo ed by EMBO long- e m ellowship (ALTF 1203-2012 and ALTF
1287-2012, espec i ely). PJC. is a Wellcome T us Senio Clinical Fellow. Suppo was p o ided o
AMF by he Na ional Ins i u e o Heal h Resea ch (NIHR) UCLH Biomedical Resea ch Cen e. ARG.
ecei es suppo om Leukaemia Lymphoma Resea ch, Cance Resea ch UK, and he Leukemia
Lymphoma Socie y. Samples om Addenb ooke's Hospi al we e collec ed wi h suppo om he NIHR
Camb idge Biomedical Resou ce Cen e. The ICGC B eas Cance Conso ium was suppo ed by a
g an om he Eu opean Union (BASIS) and he Wellcome T us . The ICGC P os a e Cance Conso ium
was unded by Cance Resea ch UK. We would also like o acknowledge he suppo o he Na ional
Cance Resea ch P os a e Cance : Mechanisms o P og ession and T ea men (PROMPT) collabo a i e
(g an code G0500966/75466) which has unded issue and u ine collec ions in Camb idge. This
esea ch was suppo ed in pa by he In amu al Resea ch P og am o he NIH, Na ional Ins i u e o
En i onmen al Heal h Sciences (JAT.). We ob ained in o med consen and consen o publish om
pa icipan s en olled.
G oup au ho de ails
ICGC B eas Cance G oup
Elena P o enzano, Camb idge B eas Uni , Addenb ooke’s Hospi al, Camb idge Uni e si y Hospi al
NHS Founda ion T us and NIHR Camb idge Biomedical Resea ch Cen e, Camb idge CB2 2QQ, UK;
Ma c an de Vij e , Depa men o Pa hology, Academic Medical Cen e , Meibe gd ee 9, 1105 AZ
Ams e dam, The Ne he lands; And ea L Richa dson, Depa men o Cance Biology, Dana-Fa be
Cance Ins i u e, 450 B ookline A e., Bos on, Massachuse s 02215, USA; Depa men o Pa hology,
B igham and Women's Hospi al, Ha a d Medical School, 75 F ancis S ., Bos on, Massachuse s
02115, USA; Colin Pu die, Eas o Sco land B eas Se ice, Ninewells Hospi al, Dundee, Uni ed
Kingdom; Sa ah Pinde , Depa men o Resea ch Oncology, Guy’s Hospi al, King’s Heal h Pa ne s
AHSC, King’s College London School o Medicine, London SE1 9RT, UK; Gae an MacG ogan, Ins i u
Be gonié, 229 cou s de l’A gone, 33076, Bo deaux, F ance; Anne Vincen -Salomon, Ins i u Cu ie,
Depa men o Tumo Biology, 26 ue d’Ulm, 75248 Pa is cédex 05, F ance; Ins i u Cu ie, INSERM
Uni 830, 26 ue d’Ulm, 75248 Pa is cédex 05, F ance; Denis La simon , Depa men o Pa hology,
Jules Bo de Ins i u e, B ussels 1000, Belgium; Do he G abau, Depa men o Pa hology, Skåne
Uni e si y Hospi al, Lund Uni e si y, SE-221 85 Lund, Sweden; To ill Saue , Depa men o Pa hology,
Oslo Uni e si y Hospi al Ulle al and Uni e si y o Oslo, Facul y o Medicine and Ins i u e o Clinical
Medicine, Oslo, No way; Øys ein Ga ed, Depa men o Pa hology, Oslo Uni e si y Hospi al Ulle al
and Uni e si y o Oslo, Facul y o Medicine and Ins i u e o Clinical Medicine, Oslo, No way; Anna
Ehinge , Depa men o Gynecology & Obs e ics, Depa men o Clinical Sciences, Lund Uni e si y,
Skåne Uni e si y Hospi al Lund, SE-221 85 Lund, Sweden; Ge G Van den Eynden, T ansla ional
Cance Resea ch Uni , GZA Hospi als S .-Augus inus, An we p, Belgium; C.H.M. an Deu zen,
Depa men o Pa hology, E asmus Medical Cen e , Ro e dam, he Ne he lands; Robe o Salgado,
B eas Cance T ansla ional Resea ch Labo a o y, Ins i u Jules Bo de , Uni e si é Lib e de B uxelles,
B ussels, Belgium; Jane E B ock, Depa men o Pa hology, B igham and Women's Hospi al, Ha a d
Medical School, 75 F ancis S ., Bos on, Massachuse s 02115, USA; Sunil R Lakhani, The Uni e si y o
Addi ional in o ma ion

Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 22 o 28
Resea ch a icle
Queensland, School o Medicine, He s on, B isbane, QLD 4006, Aus alia; Pa hology Queensland:
The Royal B isbane & Women’s Hospi al, B isbane, QLD 4029, Aus alia; The Uni e si y o Queensland,
UQ Cen e o Clinical Resea ch, He s on, B isbane, QLD 4029, Aus alia; Dilip D Gi i, Depa men o
Pa hology, Memo ial Sloan-Ke e ing Cance Cen e , New Yo k, NY, USA; Lau en A nould, Cen e
Geo ges-F ançois Lecle c, 1 ue du P o esseu Ma ion, 21079, Dijon, F ance; Jocelyne Jacquemie ,
0ns i u Paoli Calme es, biopa hology depa men , 232 Bd S e Ma gue i e, 13009, Ma seille, F ance;
Isabelle T eilleux, Cen e Léon Bé a d, Lyon, F ance; Uni e si é Claude Be na d Lyon1 - Uni e si é de
Lyon, Lyon, F ance; Ca los Caldas, Camb idge B eas Uni , Addenb ooke’s Hospi al, Camb idge
Uni e si y Hospi al NHS Founda ion T us and NIHR Camb idge Biomedical Resea ch Cen e,
Camb idge CB2 2QQ, UK; Depa men o Oncology, Uni e si y o Camb idge and Cance Resea ch
UK Camb idge Resea ch Ins i u e, Li Ka Shin Cen e, Camb idge CB2 0RE; Sue -Feung Chin,
Depa men o Oncology, Uni e si y o Camb idge and Cance Resea ch UK Camb idge Resea ch
Ins i u e, Li Ka Shin Cen e, Camb idge CB2 0RE; Aquila Fa ima, Depa men o Cance Biology,
Dana-Fa be Cance Ins i u e, 450 B ookline A e., Bos on, Massachuse s 02215, USA; Alas ai M
Thompson, Dundee Cance Cen e, Ninewells Hospi al, Dundee, UK; Alasdai S enhouse, Dundee
Cance Cen e, Ninewells Hospi al, Dundee, UK; John Foekens, E asmus MC Cance Ins i u e,
E asmus Uni e si y Medical Cen e , Ro e dam, The Ne he lands; John Ma ens, E asmus MC Cance
Ins i u e, E asmus Uni e si y Medical Cen e , Ro e dam, The Ne he lands; Anie a Sieuwe s, E asmus
MC Cance Ins i u e, E asmus Uni e si y Medical Cen e , Ro e dam, The Ne he lands; A jen
B inkman, Radboud Uni e si y, Depa men o Molecula Biology, Facul y o Science, Nijmegen
Cen e o Molecula Li e Sciences, 6500 HB Nijmegen, The Ne he lands; Henk S unnenbe g, Radboud
Uni e si y, Depa men o Molecula Biology, Facul y o Science, Nijmegen Cen e o Molecula Li e
Sciences, 6500 HB Nijmegen, The Ne he lands; Paul N. Span, Depa men o Radia ion Oncology,
Radboud Uni e si y Medical Cen e, Nijmegen, The Ne he lands; F ed Sweep, Depa men o
Labo a o y Medicine, Radboud Uni e si y Medical Cen e, Nijmegen, The Ne he lands; Ch is ine
Desmed , B eas Cance T ansla ional Resea ch Labo a o y, Ins i u Jules Bo de , Uni e si é Lib e de
B uxelles, B ussels, Belgium; Ch is os So i iou, B eas Cance T ansla ional Resea ch Labo a o y,
Ins i u Jules Bo de , Uni e si é Lib e de B uxelles, B ussels, Belgium; Gilles Thomas, Uni e si e
Lyon1, INCa-Syne gie, Cen e Leon Be a d, 28 ue Laennec Lyon Cedex 08 F ance; Annegein B oeks,
Depa men Expe imen al The apy, The Ne he lands Cance Ins i u e, Plesmanlaan 121, 1066 CX
Ams e dam, The Ne he lands; Ani a Lange od, Depa men o Gene ics, Ins i u e o Cance Resea ch,
The No wegian Radium Hospi al, Oslo Uni e si y Hospi al, O310 Oslo, No way; Samuel Apa icio,
Depa men o Molecula Oncology, BC Cance Agency, 675 W10 h A enue, Vancou e V5Z 1L3;
Pe e Simpson, The Uni e si y o Queensland, UQ Cen e o Clinical Resea ch, He s on, B isbane,
QLD 4029, Aus alia; Lau a an ' Vee , The Ne he lands Cance Ins i u e, Di ision o Molecula
Ca cinogenesis, Ams e dam, The Ne he lands; Depa men o Su ge y, Uni e si y o Cali o nia, San
F ancisco, San F ancisco, Cali o nia, Uni ed S a es o Ame ica; Jó unn E la Ey jö d, Cance Resea ch
Labo a o y, Facul y o Medicine, Uni e si y o Iceland, Reykja ik, Iceland; Holm idu Hilma sdo i ,
Cance Resea ch Labo a o y, Facul y o Medicine, Uni e si y o Iceland, Reykja ik, Iceland; Jon G
Jonasson, Depa men o Pa hology, Uni e si y Hospi al, Reykja ik, Iceland; Icelandic Cance Regis y,
Icelandic Cance Socie y, Skoga hlid 8, P.O.Box 5420, 125, Reykja ik, Iceland; Anne-Lise Bø esen-
Dale, Depa men o Gene ics, Ins i u e o Cance Resea ch, The No wegian Radium Hospi al, Oslo
Uni e si y Hospi al, O310 Oslo, No way; Ins i u e o Clinical Medicine, Facul y o Medicine, Uni e si y
o Oslo; Ming Ta Michael Lee, Na ional Geno yping Cen e , Ins i u e o Biomedical Sciences, Academia
Sinica, 128 Academia Road, Sec 2, Nankang, Taipei 115, Taiwan, ROC; Be nice Huimin Wong, NCCS-
VARI T ansla ional Resea ch Labo a o y, Na ional Cance Cen e Singapo e, 11 Hospi al D i e,
169610, Singapo e; Beni a Kia Tee Tan, Depa men o Gene al Su ge y, Singapo e Gene al Hospi al,
Singapo e; Ge i K.J. Hooije , Depa men o Pa hology, Academic Medical Cen e , Meibe gd ee 9,
1105 AZ Ams e dam, The Ne he lands
ICGC Ch onic Myeloid Diso de s G oup
Luca Malco a i, Fondazione IRCCS Policlinico San Ma eo, Uni e si y o Pa ia, Pa ia, I aly; Sudhi
Tau o, Di ision o Medial Sciences, Uni e si y o Dundee, Dundee, UK; Jacqueline Boul wood, Nu ield
Depa men o Clinical Labo a o y Sciences, Uni e si y o Ox o d, UK; And ea Pellaga i, Nu ield
Depa men o Clinical Labo a o y Sciences, Uni e si y o Ox o d, UK; Michael G o es, Di ision o
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 23 o 28
Resea ch a icle
Medial Sciences, Uni e si y o Dundee, Dundee, UK; Alex S e nbe g, Wea he all Ins i u e o Molecula
Medicine, Uni e si y o Ox o d, UK; Depa men o Haema ology, G ea Wes e n Hospi al, Swindon,
UK; Ca lo Gambaco i-Passe ini, Depa men o Haema ology, Uni e si y o Milan Bicocca, Milan,
I aly; Pa esh Vyas, Wea he all Ins i u e o Molecula Medicine, Uni e si y o Ox o d, UK; E a Hells om-
Lindbe g, Depa men o Haema ology, Ka olinska Ins i u e, S ockholm, Sweden; Da id Bowen, S
James Ins i u e o Oncology, S James Hospi al, Leeds, UK; Nicholas CP C oss, School o Medicine,
Uni e si y o Sou hamp on, Sou hamp on, UK; An hony R G een, Depa men o Haema ology,
Uni e si y o Camb idge, Camb idge, UK; Ma io Cazzola, Fondazione IRCCS Policlinico San Ma eo,
Uni e si y o Pa ia, Pa ia, I aly
ICGC P os a e Cance G oup
Colin Coope , Di ision o Gene ics and Epidemiology, The Ins i u e O Cance Resea ch, Su on, UK;
Depa men o Biological Sciences and School o Medicine, Uni e si y o Eas Anglia, No wich, UK;
Senio P incipal In es iga o s o he Cance Resea ch UK unded ICGC P os a e Cance P ojec ;
Rosalind Eeles, Di ision o Gene ics and Epidemiology, The Ins i u e O Cance Resea ch, Su on, UK;
Royal Ma sden NHS Founda ion T us , London and Su on, UK; Senio P incipal In es iga o s o he
Cance Resea ch UK unded ICGC P os a e Cance P ojec ; Da id Wedge, Cance Genome P ojec ,
Wellcome T us Sange Ins i u e, Hinx on, UK; Pe e Van Loo, Cance Genome P ojec , Wellcome
T us Sange Ins i u e, Hinx on, UK; Human Genome Labo a o y, Depa men o Human Gene ics, VIB
and KU Leu en, Leu en, Belgium; Gunes Gundem, Cance Genome P ojec , Wellcome T us Sange
Ins i u e, Hinx on, UK; Ludmil Alexand o , Cance Genome P ojec , Wellcome T us Sange Ins i u e,
Hinx on, UK; Ba ba a K emeye , Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on,
UK; Adam Bu le , Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; And ew
Lynch, S a is ics and Compu a ional Biology Labo a o y, Cance Resea ch UK Camb idge Resea ch
Ins i u e, Camb idge, UK; Sand a Edwa ds, Di ision o Gene ics and Epidemiology, The Ins i u e O
Cance Resea ch, Su on, UK; Niedzica Camacho, Di ision o Gene ics and Epidemiology, The Ins i u e
O Cance Resea ch, Su on, UK; Cha lie Massie, U ological Resea ch Labo a o y, Cance Resea ch UK
Camb idge Resea ch Ins i u e, Camb idge, UK; ZSo ia Ko e-Ja ai, Di ision o Gene ics and
Epidemiology, The Ins i u e O Cance Resea ch, Su on, UK; Nening Dennis, Royal Ma sden NHS
Founda ion T us , London and Su on, UK; Sue Me son, Di ision o Gene ics and Epidemiology, The
Ins i u e O Cance Resea ch, Su on, UK; Jo ge Zamo a, Cance Genome P ojec , Wellcome T us
Sange Ins i u e, Hinx on, UK; Jona han Kay, U ological Resea ch Labo a o y, Cance Resea ch UK
Camb idge Resea ch Ins i u e, Camb idge, UK; Ca hy Co bishley, Depa men o His opa hology, S
Geo ges Hospi al, London, UK; Sa ah Thomas, Royal Ma sden NHS Founda ion T us , London and
Su on, UK; Se ena Nik-Zainai, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on,
UK; Sa ah O'Mea a, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Lucy
Ma hews, Di ision o Gene ics and Epidemiology, The Ins i u e O Cance Resea ch, Su on, UK;
Je emy Cla k, Depa men o Biological Sciences and School o Medicine, Uni e si y o Eas Anglia,
No wich, UK; Rachel Hu s , Depa men o Biological Sciences and School o Medicine, Uni e si y o
Eas Anglia, No wich, UK; Richa d Mi hen, Ins i u e o Food Resea ch, No wich Resea ch Pa k,
No wich, UK; Susanna Cooke, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK;
Kei an Raine, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Da id Jones,
Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; And ew Menzies, Cance
Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Lucy S ebbings, Cance Genome
P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Jon Hin on, Cance Genome P ojec , Wellcome
T us Sange Ins i u e, Hinx on, UK; Jon Teague, Cance Genome P ojec , Wellcome T us Sange
Ins i u e, Hinx on, UK; S ua McLa en, Cance Genome P ojec , Wellcome T us Sange Ins i u e,
Hinx on, UK; Lau a Mudie, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK;
Clai e Ha dy, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Elizabe h
Ande son, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Oli ia Joseph,
Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Vic o ia Goody, Cance
Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Ben Robinson, Cance Genome
P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Ma k Maddison, Cance Genome P ojec ,
Wellcome T us Sange Ins i u e, Hinx on, UK; S ephen Gamble, Cance Genome P ojec , Wellcome
T us Sange Ins i u e, Hinx on, UK; Ch is ophe G eenman, School o Compu ing Sciences, Uni e si y
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 24 o 28
Resea ch a icle
o Eas Anglia, No wich, UK; Dan Be ney, Depa men o Molecula Oncology, Ba s Cance Cen e,
Ba s and he London School o Medicine and Den is y, London, UK; S e en Hazell, Royal Ma sden
NHS Founda ion T us , London and Su on, UK; Naomi Li ni, Royal Ma sden NHS Founda ion T us ,
London and Su on, UK; Cy il Fishe , Royal Ma sden NHS Founda ion T us , London and Su on, UK;
Ch is ophe Ogden, Royal Ma sden NHS Founda ion T us , London and Su on, UK; Pa deep Kuma ,
Royal Ma sden NHS Founda ion T us , London and Su on, UK; Alan Thompson, Royal Ma sden NHS
Founda ion T us , London and Su on, UK; Ch is ophe Woodhouse, Royal Ma sden NHS Founda ion
T us , London and Su on, UK; Da id Nicol, Royal Ma sden NHS Founda ion T us , London and Su on,
UK; E ik Maye , Royal Ma sden NHS Founda ion T us , London and Su on, UK; Tim Dudde idge,
Royal Ma sden NHS Founda ion T us , London and Su on, UK; Nimish Shah, U ological Resea ch
Labo a o y, Cance Resea ch UK Camb idge Resea ch Ins i u e, Camb idge, UK; Vincen
Gnanap agasam, U ological Resea ch Labo a o y, Cance Resea ch UK Camb idge Resea ch Ins i u e,
Camb idge, UK; Pe e Campbell, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on,
UK; And ew Fu eal, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Senio
P incipal In es iga o s o he Cance Resea ch UK unded ICGC P os a e Cance P ojec ; Douglas
Eas on, Cen e o Cance Gene ic Epidemiology, Depa men o Oncology, Uni e si y o Camb idge,
Camb idge, UK; Senio P incipal In es iga o s o he Cance Resea ch UK unded ICGC P os a e
Cance P ojec ; Anne Y Wa en, Depa men o His opa hology, Camb idge Uni e si y Hospi als NHS
Founda ion T us , Camb idge, UK; Ch is ophe Fos e , Bos wick Labo a o ies, London, UK; Senio
P incipal In es iga o s o he Cance Resea ch UK unded ICGC P os a e Cance P ojec ; Michael
S a on, Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Senio P incipal
In es iga o s o he Cance Resea ch UK unded ICGC P os a e Cance P ojec ; Hayley Whi ake ,
U ological Resea ch Labo a o y, Cance Resea ch UK Camb idge Resea ch Ins i u e, Camb idge, UK;
Ul an McDe mo , Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, UK; Senio
P incipal In es iga o s o he Cance Resea ch UK unded ICGC P os a e Cance P ojec ; Daniel
B ewe , Di ision o Gene ics and Epidemiology, The Ins i u e O Cance Resea ch, Su on, UK;
Depa men o Biological Sciences and School o Medicine, Uni e si y o Eas Anglia, No wich, UK;
Da id Neal, U ological Resea ch Labo a o y, Cance Resea ch UK Camb idge Resea ch Ins i u e,
Camb idge, UK; Depa men o Su gical Oncology, Uni e si y o Camb idge, Addenb ooke's Hospi al,
Camb idge, UK; Senio P incipal In es iga o s o he Cance Resea ch UK unded ICGC P os a e
Cance P ojec
Funding
Funde
G an e e ence
numbe Au ho
Wellcome T us Ul an McDe mo ,
Michael R
S a on,
Pe e J Campbell
Wellcome T us Heal h Inno a ion
Challenge Fund (HICF)
Pe e J Campbell
Kay Kendall Leukaemia
Fund
An hony R G een,
Mel G ea es,
Pe e J Campbell
Cho doma Founda ion P And ew Fu eal
Adenoid Cys ic Ca cinoma
Resea ch Founda ion
P And ew Fu eal
Eu opean Molecula Biology
O ganiza ion
ALTF 1203_2012 Young Seok Ju
Na ional Ins i u e o Heal h
Resea ch
Biomedical Resea ch
Cen e a Uni e si y
College London Hospi als
Ad ienne M
Flanagan
Leukaemia and Lymphoma
Resea ch
An hony R G een
Cance Resea ch UK An hony R G een
Leukemia and Lymphoma
Socie y
An hony R G een
Genomics and e olu iona y biology
Ju e al. eLi e 2014;3:e02935. DOI: 10.7554/eLi e.02935 25 o 28
Resea ch a icle
Funde
G an e e ence
numbe Au ho
Eu opean Union B eas Cance Soma ic
Gene ics S udy (BASIS)
Michael R
S a on
Na ional Cance Resea ch Ins i u e PROMPT: G0500966/75466 Rosalind Eeles,
Colin Coope ,
Da id Neal
Na ional Ins i u e o En i onmen al
Heal h Sciences
In amu al Resea ch
P og am o he NIH
Jack A Taylo
Na ional Ins i u e o Heal h
Resea ch
Camb idge Biomedical
Resea ch Cen e
An hony R G een,
Da id Neal
Eu opean Molecula Biology
O ganiza ion
ALTF 1287-2012 Inigo
Ma inco ena
Depa men o Heal h Heal h Inno a ion
Challenge Fund (HICF)
Pe e J Campbell
The unde s had no ole in s udy design, da a collec ion and in e p e a ion, o he
decision o submi he wo k o publica ion.
Au ho con ibu ions
YSJ, Concep ion and design, Acquisi ion o da a, Analysis and in e p e a ion o da a, D a ing o
e ising he a icle; LBA, Analyzed mu a ional signa u e; MG, IM, Analysis and in e p e a ion o da a,
D a ing o e ising he a icle; SN-Z, MR, HRD, EP, GG, AS, NB, SB, PST, JN, CEM, GSV, ARG, M-QD,
AU, JEP, BTT, NM, MG, PV, AKE-N, TS, VPC, RG, JAT, DNH, DM, CSF, AYW, HCW, DB, RE, CC, DN,
TV, WBI, GSB, AMF, PAF, AGL, PFC, UMD, Con ibu ed samples and scien i ic ad ice; APB, JWT,
P o ided bioin o ma ics suppo o sequencing da a acquisi ion; MRS, PJC, Concep ion and design,
Analysis and in e p e a ion o da a, D a ing o e ising he a icle
E hics
Human subjec s: We ob ained in o med consen and consen o publish om pa icipan s en olled
in his s udy, E hical app o al e e ences: Genome Analysis o myeloid and lymphoid malignancies
(10/H0306/40), Genomic Analysis o Meso helioma (11/EE/0444), Myeloid and lymphoid cance
genome analysis (07/S1402/90), The T ea men o Down Synd ome Child en wi h Acu e Myeloid
Leukemia and Myelodysplas ic Synd ome(AAML0431), CLL (ch onic lymphocy ic leukaemia) genome
analysis (07/Q0104/3), CGP-Exome sequencing o Down synd ome associa ed acu e myeloid leu-
kemia samples (IRB 13-010133), Cance Genome P ojec - Global app oaches o cha ac e izing he
molecula basis o paedia ic ependymoma (05/MRE04/70), PREDICT-Coho (09/H0801/96), ICGC
P os a e (E alua ion o bioma ke s in u ological diseases) (LREC 03/018), ICGC P os a e (779)
(P os a e Complex CRUK Sample Coho ) (MREC/0¼/061), ICGC P os a e (Tissue collec ion a ad-
ical p os a ec omy) (CRE-2011.373), Soma ic molecula gene ics o human cance s, melanoma and
myeloma (Dana Fa be Cance Ins i u e)(08/H0308/303), B eas Cance Genome Analysis o he
In e na ional Cance Genome Conso ium Wo king G oup (09/H0306/36), Genome analysis o
umou s o he bone (09/H0308/165).
Addi ional iles
Supplemen a y iles
• Supplemen a y ile 1. Sequencing in o ma ion o 1675 umo –no mal pai s.
DOI: 10.7554/eLi e.02935.021
• Supplemen a y ile 2. Ca alogs o soma ic mu a ions (subs i u ions and indels) and inhe i ed
polymo phisms iden i ied in his s udy.
DOI: 10.7554/eLi e.02935.022
• Supplemen a y ile 3. Lis o phased soma ic subs i u ions.
DOI: 10.7554/eLi e.02935.023
• Supplemen a y ile 4. dN/dS o 13 p o ein-coding genes in mi ochond ia.
DOI: 10.7554/eLi e.02935.024