scieee Open visual document viewer

Frequent somatic transfer of mitochondrial DNA into the nuclear genome of human cancer cells

Ju, Seok Young,Tubio, Jose M C,Mifsud, William,Bova, Steven G

Abstract

Mitochondrial genomes are separated from the nuclear genome for most of the cell cycle by the nuclear double membrane, intervening cytoplasm, and the mitochondrial double membrane. Despite these physical barriers, we show that somatically acquired mitochondrial-nuclear genome fusion sequences are present in cancer cells. Most occur in conjunction with intranuclear genomic rearrangements, and the features of the fusion fragments indicate that nonhomologous end joining and/or replication-dependent DNA double-strand break repair are the dominant mechanisms involved. Remarkably, mitochondrial-nuclear genome fusions occur at a similar rate per base pair of DNA as interchromosomal nuclear rearrangements, indicating the presence of a high frequency of contact between mitochondrial and nuclear DNA in some somatic cells. Transmission of mitochondrial DNA to the nuclear genome occurs in neoplastically transformed cells, but we do not exclude the possibility that some mitochondrial-nuclear DNA fusions observed in cancer occurred years earlier in normal somatic cells.

Full text

F equen soma ic ans e o mi ochond ial DNA in o he nuclea genome o human cance cells Young Seok Ju, 1 Jose M.C. Tubio, 1,45 William Mi sud, 1,45 Beiyuan Fu, 2 Helen R. Da ies, 1 Manasa Ramak ishna, 1 Yilong Li, 1 Lucy Ya es, 1 Gunes Gundem, 1 Pa ick S. Ta pey, 1 Sam Behja i, 1 Elli Papaemmanuil, 1 Sancha Ma in, 1 An hony Fullam, 1 Mo i z Ge s ung, 1 ICGC P os a e Cance Wo king G oup, 46 ICGC Bone Cance Wo king G oup, 46 ICGC B eas Cance Wo king G oup, 46 Jyo i Nangalia, 1,3,4 An hony R. G een, 3,4 Ca los Caldas, 3,5 Åke Bo g, 6,7,8 And ew Tu , 9 Ming Ta Michael Lee, 10,11 Lau a J. an’ Vee , 12,13 Beni a K.T. Tan, 14 Samuel Apa icio, 15 Paul N. Span, 16 John W.M. Ma ens, 17 S ian Knappskog, 18,19 Anne Vincen -Salomon, 20 Anne-Lise Bø esen-Dale, 21,22 Jó unn E la Ey jö d, 23 Ola Myklebos , 24 Ad ienne M. Flanagan, 25,26 Ch is ophe Fos e , 27 Da id E. Neal, 28,29 Colin Coope , 30,31 Rosalind Eeles, 32,33 G. S e en Bo a, 34 Sunil R. Lakhani, 35,36,37 Ch is ine Desmed , 38 Gilles Thomas, 39,44 And ea L. Richa dson, 40,41 Colin A. Pu die, 42 Alas ai M. Thompson, 43 Ul an McDe mo , 1 Feng ang Yang, 2 Se ena Nik-Zainal, 1 Pe e J. Campbell, 1 and Michael R. S a on 1 1–43 [Au ho a ilia ions appea a end o pape .] Mi ochond ial genomes a e sepa a ed om he nuclea genome o mos o he cell cycle by he nuclea double memb ane, in e ening cy oplasm, and he mi ochond ial double memb ane. Despi e hese physical ba ie s, we show ha soma ically acqui ed mi ochond ial-nuclea genome usion sequences a e p esen in cance cells. Mos occu in conjunc ion wi h in a- nuclea genomic ea angemen s, and he ea u es o he usion agmen s indica e ha nonhomologous end joining and/o eplica ion-dependen DNA double-s and b eak epai a e he dominan mechanisms in ol ed. Rema kably, mi ochond i- al-nuclea genome usions occu a a simila a e pe base pai o DNA as in e ch omosomal nuclea ea angemen s, in- dica ing he p esence o a high equency o con ac be ween mi ochond ial and nuclea DNA in some soma ic cells. T ansmission o mi ochond ial DNA o he nuclea genome occu s in neoplas ically ans o med cells, bu we do no ex- clude he possibili y ha some mi ochond ial-nuclea DNA usions obse ed in cance occu ed yea s ea lie in no mal soma ic cells. [Supplemen al ma e ial is a ailable o his a icle.] Soma ically acqui ed s uc u al ea angemen s a e common ea- u es o he nuclea genomes o cance cells. These may ange om simple ch omosomal ea angemen s (Campbell e al. 2008) o mo e complex, compound pa e ns, such as ch omo- h ipsis (S ephens e al. 2011) and ch omoplexy (Baca e al. 2013), o mobiliza ion o ansposable elemen s (Lee e al. 2012; Tubio e al. 2014). In ach omosomal ea angemen s a e gene al- ly mo e common han in e ch omosomal ea angemen s, indi- ca ing a highe likelihood o joining a double-s and b eak in a ch omosome o ano he b eak in he same ch omosome despi e he a ailabili yo a much la ge quan i yo nuclea DNA om o h- e ch omosomes (S ephens e al. 2009). In addi ion o he nuclea genome, human cells ha e a ew hund ed o a ew housand mi ochond ia, each ca ying one o a ew copies o he 16,569-bp-long ci cula m DNA (Smei ink e al. 2001; F iedman and Nunna i 2014; Ju e al. 2014). Du ing endo- symbio ic co-e olu ion, mos o he gene ic in o ma ion p esen in he ances al mi ochond ion has ans e ed o he nuclea ge- nome (G ay e al. 1999; Adams and Palme 2003; Timmis e al. 2004). An appa en bu s o m DNA ans e occu ed du ing p i- ma e e olu ion ∼54 million yea s ago (Ghe man e al. 2007) and occasional, p obably mo e ecen , ans e in humans has been obse ed in he ge mline (Tu ne e al. 2003; Goldin e al. 2004; Chen e al. 2005; Milla e al. 2010; Dayama e al. 2014). 44 Deceased. 45 These au ho s con ibu ed equally o his wo k. 46 A ull lis o membe s is p o ided in he Supplemen al Ma e ial. Co esponding au ho : [email p o ec ed] A icle published online be o e p in . A icle, supplemen al ma e ial, and publi- ca ion da e a e a h p://www.genome.o g/cgi/doi/10.1101/g .190470.115. F eely a ailable online h ough he Genome Resea ch Open Access op ion. © 2015 Ju e al. This a icle, published in Genome Resea ch, is a ailable unde a C ea i e Commons License (A ibu ion 4.0 In e na ional), as desc ibed a h p://c ea i ecommons.o g/licenses/by/4.0/. Resea ch 814 Genome Resea ch 25:814–824 Published by Cold Sp ing Ha bo Labo a o y P ess; ISSN 1088-9051/15; www.genome.o g www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Oc obe 16, 2016 - Published by genome.cshlp.o gDownloaded om Cold Sp ing Ha bo Labo a o y P ess on Oc obe 16, 2016 - Published by genome.cshlp.o gDownloaded om Cold Sp ing Ha bo Labo a o y P ess on Oc obe 16, 2016 - Published by genome.cshlp.o gDownloaded om Al hough m DNA nuclea ans e in a HeLa cell line de i a i e, and hus occu ing in i o, has been epo ed (Shay e al. 1991), de no o nuclea ans e o m DNA in animal soma ic issues has no p e iously been comp ehensi ely s udied o ou knowledge. To in es iga e he possibili y o soma ic mi ochond ial-nuclea DNA usion, we analyzed nex -gene a ion pai ed-end DNA whole-genome sequencing da a om 559 p ima y cance s, 28 can- ce cell lines ( e e ed as 587 cance whole genome below) and no - mal DNAs om he same indi iduals (Supplemen al Table 1). Resul s Disco e y o soma ic m DNA ans e s o cance nuclea genomes F om he 587 pai s o cance and no mal whole-genome se- quencing da a, we sea ched o cance -speci ic clus e s o dis- co dan pai ed-end sequence eads in which one membe o he ead-pai mapped o he nuclea genome and he o he o he mi ochond ial genome, and hen cha ac e ized he nuclea - mi ochond ial genome junc ions o nucleo ide esolu ion using indi idual sequence eads ha b idged he junc ion (Fig. 1A). In 12 samples (o e all posi i e a e 2.0%, 12 ou o 587 samples), we obse ed 25 cance -speci ic mi ochond ial-nuclea DNA junc ions (Table 1; Supplemen al Figs. 1–6). Gi en ha he e a e wo junc- ions o a single in eg a ion e en , we conclude ha he e a e mos likely 16 independen m DNA inse ions (Table 1). In addi- ion o soma ic ans e s, we obse ed se e al no el a e ge mline (inhe i ed) e en s ha we esha ed be ween cance and pai ed no - mal samples (Supplemen al Table 2; Supplemen al Ma e ial). B eas cance PD11372a showed a soma ically acqui ed in- eg a ion o almos he en i e human m DNA sequence (16,556 bp) in o a highly ampli ied 2.75-Mb-long egion o Ch omosome 10q22.3. The in eg a ion e en was s ongly sup- po ed by bo h disco dan and spli ead clus e s (Fig. 1B–D) and was con i med by sho - and long- ange PCR ac oss he nu- clea -mi ochond ial genome junc ions (Supplemen al Figs. 7, 8; Supplemen al Table 3). I was no ound in no mal issue (blood) om he same indi idual o om all he o he cases and did no ma ch any known inhe i ed nuclea m DNA-like sequences (known as num s) (Ghe man e al. 2007; Hazkani-Co o e al. 2010). Consis en wi h i s soma ic o igin, he m DNA used o he nuclea genome ha bo ed sequence polymo phisms iden ical o hose p esen in he mi ochond ia o his indi idual (14,905 G > A; 15,028 C > A; 15,043 G > A; 15,326 A > G; 15,452 C > A, and 15,607 A > G). Fluo escence in si u hyb idiza ion (FISH) expe - imen spe o medon o malin- ixedpa a inembedded issuecon- i med ha he used DNA segmen exis s in he nuclei o cance cells (Fig. 1E). In o al, we ound 10 p ima y cance s (1.8%, 10/559) and wo cance cell lines (7.1%, 2/28) wi h soma ic m DNA in eg a ions in o hei nuclea genomes (Table 1; Supplemen al Figs. 1–6). O he 12 cance s, wo (p ima y cance PD13296a and cance cell line NCI-H2087) had mo e han one mi ochond ial-nuclea DNA ansloca ion e en . All in eg a ions we e suppo ed by bo h dis- co dan andspli eadsand u he con i medbyPCRac oss henu- clea -mi ochond ial genome junc ions (Supplemen al Fig. 7; Supplemen alTable3).Allinhe i edm DNAsubs i u ionpolymo - phisms nea hese b eakpoin s we e de ec ed (Table 1). To u he isualize he ans e e en s, we pe o med high- esolu ion FISH on s e ched DNA ibe s ( ibe FISH) om he melanoma cell line, CP66-MEL (Fig. 2A). Soma ic nuclea in eg a ion o m DNA is equen ly combined wi h o he ea angemen s o he nuclea genome The a e o soma ic nuclea ans e o m DNA may a y acco ding o umo ype. T iple-nega i e b eas cance showed a i e old highe equency compa ed o es ogen- ecep o (ER) posi i e b eas cance s (6.2% and 1.2%, espec i ely; Fishe ’s exac es P = 0.002). T iple-nega i e b eas cance genomes ca y a highe numbe o ch omosomal ea angemen s han ER-posi i e b eas (a e age 254 and 94, espec i ely, in ou da a se ). As a esul , he e was a sugges i e posi i e co ela ion be ween he numbe o ch o- mosomal ea angemen s and m DNA ans e s (Mann-Whi ney U es , one-sided P= 0.05) (Fig. 2B). The leng h o m DNA agmen s ans e ed anged om 148 bp o en i e mi ochond ial genomes (16.5 kb) (Table 1). In e es - ingly, b eakpoin s in m DNA we e en iched nea he mi ochond i- al genome hea y s and o igin o eplica ion (χ 2 es , P= 0.0005) (Fig. 2C). This sugges s ha he gene a ion o m DNA segmen s o be in eg a ed in o he nuclea genome is no andom and may occu in a m DNA eplica ion-dependen manne (Lenglez e al. 2010). O he 25 mi ochond ial-nuclea DNA junc ions, a leas 17 (68.0%) we e clea ly associa ed wi h o he nuclea ch omosomal ea angemen s (e.g., in e sions, ansloca ions, and la ge dele- ions) in he icini y (Table 1; Supplemen al Figs. 1–6). Fo example, wi h espec o PD11372a desc ibed ea lie , genomic agmen s om Ch omosomes 10, 11, and m DNA gene a ed com- plex de i a i e ch omosomes (Fig. 3A). In PD6047a, an m DNA agmen was in ol ed in chains o complex genomic ansloca- ions in ol ing Ch omosomes 6, 7, 11, 22, and X (Fig. 3B). In PD10014a, a local in e sion was combined wi h he m DNA in e- g a ion e en (Fig. 3C), and in PD4252a, a 16.5-kb m DNA in eg a- ion was ound in a posi ion on he X Ch omosome om which ∼20 kb o nuclea DNA had been soma ically dele ed (Fig. 3D). Thus, m DNA is o en in eg a ed in o nuclea genomes in he icini y o , o as pa o , complex ea angemen s. Al hough ge m- line num s end o occu nea ansposable elemen s such as LINEs and SINEs (Mishma e al. 2004), we do no obse e his associa ion o soma ic e en s (χ 2 es , wo-sided P= 0.33) (Supplemen al Table 4). The mechanism and iming o soma ic nuclea ans e o m DNA The e was o e lapping sequence mic ohomology ( om 1 o 4 bp) in 20/25 b eakpoin s (80%) (Fig. 4A,B; Table 1; Supplemen al Figs. 1–6), subs an ially mo e han expec ed by chance (χ 2 es , P=5× 10 −26 ). Thus, DNA sequence mic ohomology plays an impo an ole in mi ochond ial-nuclea DNA in eg a ion e en s, al hough blun -end DNA epai was also obse ed. In wo b eakpoin s, we also ound non empla ed sho -nucleo ide inse ions (1 and 4 bp long) (Fig. 4A; Table 1). O e all, hese ea u es a e cha ac e is ic o DNA double-s and b eak epai by nonhomologous end join- ing (NHEJ) (Has ings e al. 2009). Howe e , hey do no ule ou eplica ion-based mechanisms swi ching empla e be ween nucle- a and m DNA, such as mic ohomology-media ed b eak-induced eplica ion (MMBIR) (Liu e al. 2011). We in es iga ed he iming o soma ic m DNA in eg a ion in o he nuclea genome by assessing cases in which a me as a ic sample had been sequenced in addi ion o he p ima y umo . One such case (PD4252a) showed he mi ochond ial-nuclea in e- g a ion e en in he p ima y bu no in he me as asis (Fig. 4C), in- dica ing ha m DNA ans e o he nucleus can occu a e Nuclea in eg a ion o mi ochond ial DNA in cance Genome Resea ch 815 www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Oc obe 16, 2016 - Published by genome.cshlp.o gDownloaded om neoplas ic ans o ma ion and du ing he cou se o subclonal e o- lu ion o he cance . The o he (PD6728b) showed i in bo h he p ima y and me as asis (Fig. 4C), sugges ing ha his e en oc- cu ed in he common ances al cance clone o in no mal soma ic cells p io o neoplas ic change. Nuclea ans e o m DNA is unexpec edly equen in human soma ic cells To ob ain a pe spec i e on he equency o mi ochond ial-nuclea DNA ansloca ion, we compa ed i s a e o ha o in anuclea 1 0 20 40 60 80 100 120 140 160 180 200 220 240 2 0 20 40 60 80 100 120 140 160 180 200 220 240 3 0 20 40 60 80 100 120 140 160 180 4 0 20 40 60 80 100 120 140 160 180 5 0 20 40 60 80 100 120 140 160 180 6 0 20 40 60 80 100 120 140 160 7 0 20 40 60 80 100 120 140 8 0 20 40 60 80 100 120 140 9 0 20 40 60 80 100 120 140 10 0 20 40 60 80 100 120 11 0 20 40 60 80 100 120 12 0 20 40 60 80 100 120 13 0 20 40 60 80 100 14 0 20 40 60 80 100 15 0 20 40 60 80 100 16 0 20 40 60 80 17 0 20 40 60 80 18 0 20 40 60 19 0 20 40 20 0 20 40 60 21 0 20 40 22 0 20 40 x 0 20 40 60 80 100 120 140 umou blood 1 Mb 10 kb 100 bp 1 bp 1 bp 100 bp 10 kb 1 Mb PD11372a CTCCTGGGTG AGAAA CTCCTGGGTG TTGGCCTCAC GATAT TTGGCCTCAC Fusion Ch 10 m DNA TACTGTGGC CC AGACCTCTT ACACT GC AGACCTCTT TACTGTGGC CC CTCAG x 60 x 27 x 25 x 76 ch 10 (+) 81,670,932 MT (-) 15,157 MT (-) 15,171 ch 10 (+) 78,920,385 x 24 x 28 16,556 bp * * * * * * * * * * * * * * * * * * * * * * * ** ch 10 posi ion (Mb) 0 20406080100120 * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * ** * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * ** * * * * * * * * ** * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * ** * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * ** * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * ** * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * 0 2 4 6 8 10 Copy numbe * ** * ** ** * * * * * * * * * * * * * * ** * ** * * ** ** * * * * ******* **** * * * * ** * * ** ** * * * * ** * ******* **** *** * * * * ** * * * * * * * * * * * ** * * **** * * **** *** * * ** ** **** **** *** ** ** * * * * ** * *** **** * * * * * * * * * * * * * * * * * * * * * * ** * ** **** * * * * * * * * * * * * * * * * * * * * * * * ** * * ** **** * * * * * ** * * ** * ** ** **** * * ** * * * ** ** * ** * ** * * * * * * * * * * * * * * * * * * * * * * * * * ** * * * * * ** * * *** * ** * ** * * * * * * * * ** * * * * * * * * * * * * * * * * * * * * * * * * * * * * * ** * * * * * ** * * * * * * * ** ** * * **** **** * * * * * * * * * * * ** ** ** ** ** * * * * ** ** * ** * *** * ** ******** * ** ** **** ** *** *** * * * * * * * * * * * **** * * * ***** * *** * ********** * ** * * * * ** * * * * * * * * ** * * * * m DNA PD11372a dele ion ype andem duplica ion ype in e sion ype (head-head) in e sion ype ( ail- ail) m DNA ch 10: 80Mb ch 10: 78Mb me ged E DC BA mi ochond ial DNA (inse ed segmen ) (2) DRs (1) SRs (Nu) nuclea DNA nuclea DNA Mapping o he e e ence genome (using BWA) cance genome (3) SRs (MT) nuclea (Nu) genome mi ochond ial (MT) genome (2) DRs (1) SRs (Nu) mapped in Nu. genome as ma e-unmmaped (3) SRs (MT) mapped in MT genome as ma e-unmmaped Figu e 1. Disco e y o soma ic nuclea m DNA ans e om PD11372a. (A) The s a egy o de ec ion o nuclea m DNA ans e e en s. See Me hods o a de ailed desc ip ion. (SRs) Spli - eads, (DRs) disco dan eads, (Nu) nucleus, (MT) mi ochond ia. (B) G aphical ep esen a ion o disco dan ead clus- e s in PD11372a and i s pai ed-no mal issue (PD11372b). The ed a ow indica es umo -speci ic disco dan - ead clus e s in Ch 10. Ch omosome ideo- g ams a e shown in he ou e laye . The dis ance be ween each disco dan ead and one p io o i ( he in e - ead dis ance) is plo ed on he e ical axis on a log-scale in he middle ( umo ) and inne laye (blood). Blue do s shown in he middle laye ep esen known num s. (C) m DNA in eg a ion in PD11372a. B eakpoin sequences a e shown.Red ec anglehighligh s mic ohomology. Numbe s o disco dan spli eads a ep esen ed. Inhe i ed m DNA subs i u ion polymo phisms a e shown by ed as e isks. (D) Rea angemen a chi ec u es o Ch omosome 10 o PD11372a. DNA copy numbe s a e shown by black do s. The copy numbe o 2.75-Mb-long egion used wi h m DNA is colo ed in ed. Reads suppo ing ea angemen s (la ge dele ions, andem dupli- ca ions, ail- ail and head-head in e sions) a e shown by a cs and e ical lines. Ch 10-m DNA usions a e shown wi h ed a ows. (E) Nuclea FISH con i ms he mi ochond ial-nuclea DNA usion in he nucleus. (Red) Ch 10 (80 Mb), (blue) Ch 10 (78 Mb), and (g een) m DNA. Ju e al. 816 Genome Resea ch www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Oc obe 16, 2016 - Published by genome.cshlp.o gDownloaded om Table 1. Summa y o soma ic mi ochond ial-nuclea DNA usions iden i ied om 12 cance samples Tissue Sample Le junc ion Righ junc ion F ag. size (bp) Mic o- homology (bp,bp) Va ian s (#D/#P) a Con ex o ea angemen Nuclea MT MT Nuclea P ima y PD11372a 10+:81,670,932] [M−:15,157 M−:15,171] [10+:78,920,385 16,556 (0,1) 6/6 m DNA inse ion wi h complex ea angemen s PD4252a X+:45,631,665] [M+:14,450 M+:14,496] [X+:45,652,120 16,616 (2,1) 2/2 m DNA inse ion wi h la ge ch . dele ion PD6047a X+:14,944,764] [M−:12,735 M−:16,128] [7+:96,923,229 13,177 (1,1) 6/6 Mul iple in e ch omosomal ansloca ions PD10014a 17−:75,618,348] [M−:13,365 M−:9055] [17+:75,688,733 4311 (2,3) 0/0 m DNA inse ion wi h Ch 17 in e sion PD13296a 4+:102,463,870] [M−:14,705 M−:13,235] [4+:102,464,084 1471 (4,0) 0/0 m DNA inse ion wi h la ge ch . dele ion 6+:103,639,248] [M+:14,692 M+:14,972]TA AT [6+:103,690,941 281 (2,0) 2/2 m DNA inse ion wi h la ge ch . dele ion PD6728b 2+:138,664,890] [M−:13,199 M−:13,052] [2−:139,012,040 148 (4,2) 1/1 m DNA inse ion wi h complex ea angemen s PD11397a 19−:12,650,382] [M+:16,233 M+:96] [17+:40,005,738 433 (0,2) 1/1 Mul iple in e ch omosomal ansloca ions PD7404a 1+:44,914,376] [M+:3732 –– >200 (1,–) 0/0 – PD6733b 6−:45,823,498] [M+:16,107 –– >200 (0,–) 1/1 – PD11768a 1−:144,944,326] [M+:16,104 –– >200 (4,–) 1/1 – Cell line CP66-MEL 3+:47,419,506] [M−:7048 M−:16,193] [3+:47,419,447 7425 (1,1) 1/1 m DNA inse ion NCI-H2087 10+:26,775,605] [M+:1690 –– >200 (1,–) 1/1 – 20−:33,836,717] [M−:5666 –– >200 (1,–) 1/1 – 17−:7,481,787] T [M−:3452 –– >200 (1,–) 1/1 b Mul iple in e ch omosomal ansloca ions 17−:31,744,235] [M+:4346 –– >200 (3,–) 1/1 – a Inhe i ed m DNA polymo phisms in he icini y o b eakpoin s. (#D) Numbe o de ec ed, (#P) numbe o p esen . b A soma ically acqui ed he e oplasmic mu a ion in mi ochond ia. Nuclea in eg a ion o mi ochond ial DNA in cance Genome Resea ch 817 www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Oc obe 16, 2016 - Published by genome.cshlp.o gDownloaded om in e ch omosomal ansloca ion, aking in o accoun he sizes and copy numbe s o he mi ochond ial and nuclea genomes. Ou sequencing da a sugges ha each cance cell ca ies ∼500 copies o ci cula m DNA (median alue 495) (Fig. 5A), amoun - ing in agg ega e o ∼8 million base pai s (bp) o m DNA (500 cop- ies × 16.5 kb) enclosed by he mi ochond ial double memb ane in he cy oplasm o each cance cell. The a e age equency in he cance s analyzed o mi ochond ial-nuclea DNA usion was 5.1 × 10 −3 junc ions pe million bp o m DNA, only hal he a - e age a e o in anuclea in e ch omosomal ansloca ion (1.2 × 10 −2 junc ions pe million bp) and simila o ha o Ch omosomes 2, 4, and 13 (Fig. 5B). Gi en he mul iple physical ba ie s o con ac be ween he wo genomes, he esul s indica e ema kably high a es o m DNA escape, con ac , and/o in eg a- ion wi h nuclea DNA in human cance cells. These appea o be conside ably highe han in he ge mline ac oss human e olu- ion bu compa able o hose obse ed in Saccha omyces ce e isiae (Tho sness and Fox 1990) and o chlo oplas DNA mig a ion in o he nucleus in obacco plan s (Me hods; Supplemen al Ma e ial; Huang e al. 2003). Discussion Despi e mul iple physical ba ie s, he e a e plausible mechanisms by which m DNA and nuclea DNA could come in o con ac (Fig. 5C). F ee m DNA can be eleased in o he cy oplasm om de- g adingmi ochond iao a e mi ophagy (Zhang e al. 2008; Eiyama e al. 2013; Higgins and Coughlan 2014). Deg ada- ion o mi ochond ia may be accele a ed in cance cells due o hypoxia and in- c eased ene gy demands (Zhang e al. 2008; Eiyama e al. 2013; Higgins and Coughlan 2014). E en wi hou a bespoke molecula p ocess o anspo a ion, m DNA could hen, in p inciple, mig a e o he nucleus du ing mi o ic me aphase o anaphasewhen henuclea memb ane has b oken down. When hese e en s a e coupled wi h concu en double-s and b eaks (DSBs) and/o eplica ion o k s alling o nuclea ch omosomal DNA, m DNA could be picked up and in eg a - ed in o he nuclea genome as pa o he p ocess o ejoining DSBs (NHEJ) (Has ings e al. 2009) o used as an al- e na i e DNA empla e in eplica ion (MMBIR) (Liu e al. 2011). Mic onuclei in cance cells, which can be gene a ed bye o sinseg ega iono mi o icnuclea ch omosomes, may con ibu e o he e en s.Ch omosomesinmic onuclei e- quen ly unde go de ec i e and delayed DNA eplica ion, esul ing in ex ensi e agmen a ion wi h subsequen jumbled ejoiningcompa ed o hei o iginalo de and o ien a ion (C as a e al. 2012; Fo - men e al. 2012). Thus, m DNA ag- men s inco po a ed in o mic onuclei could end up used o sha e ed nuclea ch omosomes. I is wo hy o no e ha m DNA escaping o he nucleus canbe ac i elyused o DNA epai in Saccha omyces ce e isiae (Ricche i e al. 1999; Yu and Gab iel 1999), pa icula ly when e o - ee DSB DNA epai is no possible. Whe he his applies in mammalian cells is unknown. Some o he soma ic nuclea m DNA in eg a ions we iden i- ied a e di ec ly adjacen o nuclea genes. Fo example, nuclea - m DNA usion in PD11372a occu ed in he i h in on o he KCNMA1 gene, a po assium channel equen ly ampli ied in p os- a e and b eas cance s (Oegge li e al. 2012). Howe e , we do no ind ob ious en ichmen o he nuclea -m DNA usion b eak- poin s nea human nuclea genes. RNA-seq om he NCI-H2087 cell-line indica es ha m DNA agmen s in he nucleus o he cell line a e no exp essed as pa s o mi ochond ial-nuclea usion ansc ip s.Thus, hemajo i yo henuclea m DNA ansloca ion e en s a e likely o be passenge e en s, simila o mu a ions o all o he ypes in mos cance genomes. Howe e , we do no ex- clude hepossibili y ha someo hese e en smayha e unc ional consequences in human cance by gene a ing usion mRNA an- sc ip s (Shay e al. 1991) and/o unca ing cance genes by m DNA inse ion wi hin exons. ch 3 (downs eam)ch 3 (ups eam) m DNA (7.4kb) A CP66-MEL (melanoma cell-line) 250 500 1000 B 78 133 125 Cance issue ypes B eas (ER + e) O he b eas B eas ( iple − e) O he ypes m DNA ans e s -+ cance samples 0 1 2 3 4 0 2 4 6 8 10 12 14 0.5-3kb 3-5.5kb 5.5-8kb 8-11.5kb 11.5-14kb 14-0.5kb Ra io (obs/exp) F equency o e en s Loca ion o b eakpoin s in mi ochond ial genome Expec ed Obse ed Ra io (obs/exp) P = 0.00052 * D-Loop RNAs 16s12s CO1 ND5ND4 ND6 Replica ion o igin ( H s and) Replica ion o igin ( L s and) C CYB Figu e 2. Fea u es o soma ic m DNA nuclea ans e in 12 cance samples. (A) Fibe FISH isualizes he mi ochond ial-nuclea DNA usion om he CP66-MEL cell line. (B) Posi i e co ela ion be ween m DNA ans e and numbe s o nuclea ch omosomal ea angemen s (la ge dele ion, andem duplica- ion, in e sion, and ansloca ion) in cance genomes. Median alues a e shown. (C) m DNA b eak- poin s a e en iched in he 14 kb- o 500-bp egion o he MT genome. (Top) Blue and ed ba s ep esen he expec ed and obse ed numbe s o b eakpoin s in each in e al o MT genome, espec- i ely. G een line shows a io be ween obse ed and expec ed numbe s. A χ 2 es was applied o es en ichmen . (Bo om) Schema ic s uc u al ea u es o he MT genome co esponding o he in e als a e shown. Ju e al. 818 Genome Resea ch www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Oc obe 16, 2016 - Published by genome.cshlp.o gDownloaded om This s udy has shown ha usion o m DNA o nuclea DNA occu s in human soma ic cells a a a e simila o ha o ansloca- ion be ween nuclea ch omosomes. Physical mig a ion o m DNA in o he nucleus may be much mo e equen in s em cells han ones in a e minally di e en ia ed s age (Schneide e al. 2014). Fu he s udies will need o add ess he mechanisms by which he appa en physical ba ie s o con ac be ween mi ochond ial and nuclea DNA a e so e ec i ely o e come. Me hods Samples and sequencing da a We analyzed 559 p ima y umo s and 28 cance cell-lines in his s udy. Pai ed-no mal samples o all he cance s we e also included in his s udy in pa allel. Whole-genome sequencesusedin hiss udywe egene a - ed by Illumina pla o ms (ei he Genome Analyze o HiSeq 2000). Cance ge- nomeswe esequenced oa leas 25×co - e age. Wi h espec o TCGA da a, we downloaded aligned BAM iles h ough UCSC CGHub (h p://cghub.ucsc.edu). Sequencing eads we e aligned on he human e e ence genome build 37 (GRCh37) and human e e ence m DNA sequence ( e ised Camb idge e e ence sequence, CRS) (And ews e al. 1999), mainly by he BWA alignmen ool (Li and Du bin 2009). SAM ools (Li e al. 2009) was used o manipula ing se- quence eads. Calling mi ochond ial-nuclea DNA usion e en s We employed a pipeline o iden i ica- ion o pu a i e m DNA ansloca ion o ch omosomal DNA (Fig. 1A). F om pai ed-end whole-genome sequencing da a o umo s, we ex ac ed disco dan eads (DRs), whe e one end aligned uniquely o m DNA and he o he end o nuclea DNA. In all cases, bo h ends mus ha e a mapping quali y g ea e han ze o. Those disco dan eads a e clus e ed oge he using he ollowing c i e ia: eads sha ing (1) close alignmen posi ions (<500 nucleo ides) o bo h ends on nuclea and m DNA, and (2) he same o ien a ions.In o de o emo e alse posi i es, we emo ed clus e s sup- po ed by less han i e disco dan eads. In o de o emo e po en ial ge m- line calls, se e al il e s a e applied o he umo candida e clus e . The clus e s om umo cells we e emo ed i hey o e lap wi h clus e s iden i ied om ma ched and/o unma ched no mal is- sues by mo e ole able c i e ia (suppo ed by mo e han one disco dan ead) om (1) i s pai ed-no mal issue, and (2) om he o he 586 unma ched no mals. Fil e ed clus e s we e u he e ined wi h known ge mline human num s, a combined se om he human e e ence genome (hg19) de ec ed by BLAT (Ken 2002) (n= 123) and om Simone e al. (2011) (n= 766). Finally, 25 clus e s we e selec ed as soma ic candida es. Nucleo ide- esolu ion b eakpoin s o he ansloca ion junc ions To ob ain nucleo ide- esolu ion b eakpoin s, we sea ched o spli - eads(SRs) wi honeo he endsspanning hejunc iono he ans- loca ion. We ex ac ed “o phan”o “ma e-unmapped” eads (one end o a ead is unmapped by he BWA aligne ) in he icini y (<1000bp)o disco dan - eadclus e sonnuclea andmi ochond i- al genome sequences. Sequences om he unmapped end a e hen e-aligned by BLAT (Ken 2002), which enablesspli - ead mapping. A 10 0 5 5 10 15 20 25 30 35 3 40 45 4 50 55 60 65 6 70 75 80 85 8 90 95 100 105 110 115 1 120 125 130 135 11 0 5 5 10 15 5 20 25 5 30 35 40 45 5 50 55 5 60 65 5 70 75 5 80 85 5 90 95 5 100 105 110 115 5 120 125 130 135 5 MT 0 PD11372a D B ch X(+):45,631,665 ch X(+):45,652,120 m DNA inse ion (16.5kb) nuclea DNA dele ion (20 kb)  ch omosome X 0 20 40 60 80 100 120 140 p11.3 Copy numbe 45.61 45.62 45.63 45.64 45.65 45.66 45.67 024 o oooo o oooooooo o ooooooooooooooo o o o o oo oo o ooo o oooooooooo oooooooooooooo o ooooooooooooooooooooo oo oooooooooooooo ooooooo ooo ooooooooooooooo o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o o 0 8 abe an ead clus e s PD4252a PD6047a 2 0 10 20 30 40 50 60 70 80 90 100 110 120 130 140 150 160 170 180 190 200 210 220 230 240 3 0 10 20 30 40 50 60 70 80 90 100 110 120 130 140 150 160 170 180 190 4 0 10 20 30 40 50 60 70 80 90 100 110 120 130 140 150 160 170 180 190 5 0 10 20 30 40 50 60 70 80 90 100 110 120 130 140 150 160 170 180 6 0 10 20 30 40 50 60 70 80 90 100 110 120 130 140 150 160 170 7 0 10 20 30 40 50 60 70 80 90 100 110 120 130 140 150 8 0 10 20 30 40 50 60 70 80 90 100 110 120 130 140 10 0 10 20 30 40 50 60 70 80 90 100 110 120 130 11 0 10 20 30 40 50 60 70 80 90 100 110 120 130 12 0 10 20 30 40 50 60 70 80 90 100 110 120 130 13 0 10 20 30 40 50 60 70 80 90 100 110 16 0 10 20 30 40 50 60 70 80 90 22 0 10 20 30 40 50 x 0 10 20 30 40 50 60 70 80 90 100 110 120 130 140 150 MT 0 m DNA (-) ch 7(+):96,923,229 ch X(+):14,944,764 13.1 kb m DNA(-): 13,365 - 9,055 (4.3kb) ch 17(+):75,688,733 local in e sion C ch 17(+):75,564,373 ch 17(-):75,655,898 - 75,618,348 (37.6kb) m DNA inse ion PD10014a Figu e 3. Concu ence o soma ic m DNA nuclea ans e s wi h o he s uc u al a ia ions. The com- plex web o ea angemen s in he icini y o mi ochond ial-nuclea DNA usions om ou examples. (A) In PD11372a, m DNA in eg a ion wi h complex ea angemen s be ween Ch 10 and 11. (B)In PD6047a, m DNA in eg a ion wi h complex ea angemen s among Ch 6, 7, 11, 22, and X. (A,B) DNA copy numbe s a e shown by black do s wi h a log scale. Red lines ep esen ansloca ions in ol ing m DNA. (C) In PD10014a, m DNA in eg a ion combined wi h a local in e sion (yellow). (D) In PD4252a, m DNA in eg a ion wi h a local dele ion. DNA copy numbe s a e shown wi h blue do s and lines. Abe an ead clus e s (disco dan and spli eads) a e shown by g een and ed a ows, espec i ely. Nuclea in eg a ion o mi ochond ial DNA in cance Genome Resea ch 819 www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Oc obe 16, 2016 - Published by genome.cshlp.o gDownloaded om Valida ion by PCR A PCR alida ion assay o he soma ic m DNA ans e was pe - o med using genomic DNA om bo h cance and pai ed-no mal issues. P ime s we e designed o ampli y all he b eakpoin s (Supplemen al Table 3). The sho - agmen PCR eac ions we e pe o med as p e iously desc ibed (Tubio e al. 2014). Wi h espec o long- ange PCR, elonga ion ime was inc eased 1 min pe 1 kb. Gene a ion o FISH p obes Human bac e ial a i icial ch omosomes (BAC) and osmid clones used in his s udy we e ob ained om he clone a chi e eam o he Wellcome T us Sange Ins i u e. Plasmid DNA was p epa ed using he PhaseP ep BAC DNA ki (Sigma-Ald ich). Human m DNA was isola ed om lymphoblas- oid cells using a Mi ochond ial DNA Isola ion ki (Abcam). P obes o use in FISH we e made as desc ibed be o e (G ibble e al. 2013). Pu i ied m DNA and plasmid DNA we e i s ampli ied using a GenomePlex Whole Genome Ampli ica ion (WGA) ki (Sigma-Ald ich) ollowing he manu- ac u e ’s p o ocols, hen labeled using a WGA eampli ica ion ki (Sigma- Ald ich) wi h a cus om-made dNTP mix. P obes o in e phase FISH we e labeled di ec ly wi h Aminoallyl-dUTPs - ATTO- 488, -Cy3, -Texas Red, and -Cy5 (Jena Bioscience); p obes o ibe -FISH we e la- beledwi hBio in-16-dUTP,Digoxigenin- 11-dUTP (Roche), and DNP-11-dUTP (Pe kinElme ). Valida ion by ibe -FISH wi h single- molecule DNA ibe s gene a ed by molecula combing Single-molecule DNA ibe s om he cance cell line, CP66-MEL, we e p e- pa ed by molecula combing (Michale e al. 1997) ollowing he manu ac u e ’s ins uc ions (Genomic Vision). B ie ly, he cells we e embedded in a low-mel - poin aga ose plug (1 million cells pe plug), ollowed by p o einase K diges- ion, washing in 1 × TE (10 mM T is, 1 mM EDTA, pH 8.0) and be a-aga ose di- ges ion s eps. The DNA ibe s we e me- chanically s e ched on o saline-coa ed co e slips using a Molecula Combing Sys em (Genomic Vision). Fo ibe -FISH, ∼500 ng o labeled DNA om each p obe and 4 μg o human Co -1 DNA (In i ogen) we e p ecipi- a ed using e hanol, hen esuspended in a mix (1:1) o hyb idiza ion bu e (con aining 2 × SSC, 10% sa kosyl, 2 M NaCl, 10% SDS, and blocking aid [In i- ogen]) and deionized o mamide ( inal concen a ion 50%). Co e slips coa ed wi h combed DNA ibe s we e dehyd a ed h ough a 70%, 90%, and 100% e hanol se ies and aged a 65°C o 30 sec, ollowed by dena u a ion in an alkaline dena u e solu ion (0.5 M NaOH, 1.5 M NaCl) o 1–3 min, h ee washes wi h 1×PBS (In i ogen), and dehyd a ion h ough a 70%, 90%, and 100% e hanol se ies. The p obe mix was dena u ed a 65°C o 10 min be o e being applied on o he co e slips, and he hyb idiza ion was ca ied ou in a 37° C incuba o o e nigh . The pos -hyb idiza ion washes consis ed o wo ounds o washes in 50% o mamide/2 × SSC ( / ), ollowed by wo addi ional washes in 2 × SSC. All pos -hyb idiza ion washes we e done a 25°C, 5 min each ime. Digoxigenin-11-dUTP (Roche) labeled p obes we e de ec ed using a 1:100 dilu ion o monoclonal ch 4: GGCGA AACCC CATT TCTACT Fusion: GGCGA AACCC CATT GGTCGT m DNA: TTTTT CATAT CATT GGTCGT TAAT A ch 4(+)102,463,870 m DNA(-):14,705 PD13296a (2 m DNA nuclea in eg a ion e en s) m DNA(-):13,235 ch 4(+):102,464,084 4bp mic ohomology blun -end DNA joining GCTACGATTT CTTTTGATGT : m DNA GCTACGATTT AAATAACCAC : Fusion GTTATCTTCA AAATAACCAC : ch 4 In eg a ion #1 In eg a ion #2 ch 6(+):103,639,248 m DNA(+) 14,692 m DNA(+) 14,972 ch 6(+):103,690,941 ch 6: TTGTAAGA AC TAATAGAATG Fusion: TTGTAAGA AC AACCACGACC m DNA:CACGGACT AC AACCACGACC GTAAATTATG GCTGAATCAT : m DNA GTAAATTATG TAAT AAAATATTTG : Fusion CTGGGTCCTA AAAATATTTG : ch 6 2bp mic ohomology non- empla e 4bp inse ion PD6728b B ch 2(+):138,664,890 m DNA(-) 13,199 m DNA(-) 13,052 ch 2(-):139,012,040 ch 2: TCATCT TGCT TGCGTTTTGC Fusion: TCATCT TGCT GCGAACAGAG m DNA:GCAGAC TGCT GCGAACAGAG 4bp mic ohomology GGGGTGGGGC CT TCTATGGC : m DNA GGGGTGGGGC CT GACTGCAG : Fusion CTTGGTCTTG CT GACTGCAG : ch 2 2bp mic ohomology C PD4252 Fe ilized egg Blood (PD4252b) MRCA LN me as asis (PD4252c) P ima y locus (PD4252a; subclonal) m DNA ans e PD6728 Fe ilized egg Blood (PD6728a) MRCA LN me as asis (PD6728c; clonal) P ima y locus (PD6728b; clonal) m DNA ans e T ans o ma ion T ans o ma ion Figu e 4. Nucleo ide- esolu ion b eakpoin sequences and he iming o soma ic m DNA nuclea in e- g a ion. (A) B eakpoin sequences o nuclea -m DNA usions in PD13296a. Red ec angles highligh se- quence mic ohomology and non empla e nucleo ides inse ion. (B) B eakpoin sequences o nuclea - m DNA usions in PD6728b. Red ec angles highligh sequence mic ohomology. (C) Phylogene ic ees showing he iming o soma ic m DNA nuclea ans e s in PD4252 and PD6728 samples. (MRCA) Mos ecen common ances o cell. Ju e al. 820 Genome Resea ch www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Oc obe 16, 2016 - Published by genome.cshlp.o gDownloaded om mouse an i-dig an ibody (Sigma-Ald ich) and a 1:100 o Texas Red- X-conjuga ed goa an i-mouse IgG (Molecula P obes/In i ogen); DNP-11-dUTP (Pe kinElme ) labeled p obes we e de ec ed using a 1:100 dilu ion o Alexa 488-conjuga ed abbi an i-DNP IgG and 1:100 Alexa 488-conjuga ed donkey an i- abbi IgG (Molecula P obes/In i ogen); bio in-16-dUTP (Roche) labeled p obes we e de ec ed wi h one laye 1:100 o Cy3-a idin (Sigma-Ald ich). A e de ec ion, slides we e moun ed wi h SlowFade Gold moun ing sol- u ion con aining 4′,6-diamidino-2-phenylindole (Molecula P obes/In i ogen). Images we e isualized on a Zeiss AxioImage D1 mic oscope. Digi al image cap u e and p ocessing we e ca ied ou using he Sma Cap u e so wa e (Digi al Scien i ic UK). Nuclea in e phase FISH Nuclei ex ac ion om pa a in-embed- ded issue o pa ien PD11372a and in e - phase-FISH ollowed Pa e nos e e al. (2002), wi h he excep ion ha 60-μm- hick sec ions we e used in ou s udy. The pos -hyb idiza ion washes consis ed o wo ounds o washes in 50% o mam- ide/2 × SSC ( / ), ollowed by wo ad- di ional washes in 2 × SSC. Slides we e moun ed wi h SlowFade Gold moun - ing solu ion con aining 4′,6-diamid- ino-2-phenylindole (Molecula P obes/ In i ogen). Images we e cap u ed and p ocessed as desc ibed abo e. Co ela ion be ween soma ic m DNA in eg a ion si e and ansposable elemen s We pe o med a s udy simila o he p e- ious epo (Mishma e al. 2004). We calcula ed he dis ance be ween each m DNA-inse ion si e (b eakpoin ) and i s nea es ansposable elemen s (ei- he o SINE, LINE, LTR, simple epea , o DNA ansposon by Repea Maske , downloaded om he UCSC Genome B owse , June 6, 2013). Then, each m DNA-inse ion si e was ca ego ized in o one o ou g oups: (A) b eakpoin wi hin a ansposable elemen ; (B) b eak- poin wi hin 15 bp om a ansposable elemen ; (C) wi hin 15–150 bp; and (D), >150 bp. In o de o unde s and he posi- ional en ichmen o b eakpoin s om ansposable elemen s, we andomly gene a ed in silico b eakpoin posi ions 40 imes as many ( o al n= 1000) as we obse ed om each ch omosome in he eal da a se . In silico b eakpoin s loca ed wi hin gaps o he human e e ence ge- nome we e emo ed and eplaced by newly gene a ed inse ions. Fo hese in silico-gene a ed b eakpoin s, he dis- ances om he nea es ansposable elemen s we e calcula ed and hen ca e- go ized in o one o he ou g oups (A, B,C, and D). Finally, he di e ence in he equency o b eakpoin s in each g oup be ween he obse ed and in silico-gene a ed da a se was compa ed using a χ 2 es . Assessmen o m DNA copy numbe s To unde s and m DNA copy numbe s in a cance cell, we com- pa ed a e age ead dep h o co e age be ween 22 au osomes and m DNA. Wi h espec o he umo sequences by whole-genome sequencing, a e age haploid au osomal co e age (RD au osome ) was ob ained om he ead dep h o 2.685-Gb-long au osomal e- gions (excluding ch omosomal gaps). Likewise, a e age m DNA co e age (RD m DNA ) was ob ained om he ead dep h o he 500 1000 2000 Cance issue ypes B eas Os eosa coma O he ypes P os a e A Es ima ed ci cula m DNA copy numbe s (in cy oplasm) pe cance cell 0.00 0.01 0.02 0.03 ch omosomes T ansloca ion a e (# o e en s pe Mb) ch 2 ch 17 ch 19 m DNA ch 4 ch 13 B mi ophagy mi ochond ial deg ada ion escape o m DNA nuclea memb ane b eaks down beginning o mi osis DNA double-s and b eaks (ch omosome sha e ing) and/o eplica ion o k s alling m DNA mig a ion o he nucleus (mic onucleus) ( a e> 2x10 pcpg) cell memb ane -4 nucleus (mic onucleus) m DNA in eg a ion DSB epai (NHEJ, MMBIR) C Figu e 5. F equency and po en ial mechanisms o soma ic m DNA nuclea ans e in human can- ce . (A) Es ima ed ci cula m DNA copy numbe s (in he cy oplasm) pe cance cell om 587 cance is- sues sequenced. The a io o ead dep hs be ween au osomes and m DNA was used (see Me hods). (B) Simila equency o soma ic nuclea m DNA in eg a ions compa ed o he equency be ween au o- somes (ch omosomal ansloca ion). (C) A model o soma ic m DNA ans e o he nuclea genomes. Nuclea in eg a ion o mi ochond ial DNA in cance Genome Resea ch 821 www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Oc obe 16, 2016 - Published by genome.cshlp.o gDownloaded om 16.5-kb mi ochond ial genome. Finally, m DNA copy numbe in a diploid cell (C m ) is calcula ed as shown below: Cm =2×RDm DNA RDau osome . Assessmen o ansloca ion a e o au osomes and mi ochond ia We iden i ied s uc u al a ia ions among nuclea ch omosomes (la ge dele ions, andem duplica ions, in e sions, and in e ch o- mosomal ansloca ions) using he BRASS II algo i hm (Nik- Zainal e al. 2012), which iden i ies ea angemen s by clus e ing disco dan ead pai s ha poin o he same junc ion and con i ms b eakpoin s by local assembly o unmapped eads. The sensi i i y and speci ici y o he BRASS II algo i hm is equi alen o hose al- ues o he algo i hm used o mi ochond ial-nuclea DNA usions (da a no shown). We ex ac ed in e ch omosomal ansloca ions o calcula e he a e o such e en s. The a e o each haploid au o- some (R ,ch) is calcula ed as shown below: R ,ch(e en s pe megabase)=N ,ch/(2×Lch)/Nsam, whe e N ,ch is he o al numbe o soma ic in e ch omosom- al ansloca ion junc ions in ol ing a speci ic ch omosome, Lch is heleng ho henon edundan egiono hech omosomeinmeg- abases, and Nsam is he o al numbe o samples analyzed. To ob- ain he unique egion leng h (Lch), we excluded edundan (o highly epe i i e) sequence leng hs om he ungapped leng h o each ch omosome. Genomic egions classi ied in one o mo e o he h ee c i e ia shown below we e de ined as edundan , whe e ansloca ion e en s could no be easilyde ec ed due o ambiguous ead alignmen : (1) simple epea s, loca ed by Tandem Repea s Finde (Benson 1999); (2) segmen al duplica ions wi h mode a e o high sequence simila i y (≥95%) (Bailey e al. 2002), o (3) epe - i i e sequences including up o 10 di e en classes o epea s (such as SINE, LINE, LTR, DNA ansposons, andmic osa elli es), loca ed by he Repea Maske p og am (h p://www. epea maske .o g), wi h a low di e gence le el (di e gence < 5%). These non edun- dan sequence egions we e downloaded om he UCSC Genome B owse (h p://genome.ucsc.edu). Simila ly, he a e o mi ochond ial-nuclea DNA ansloca- ions (R ,m ) was calcula ed as below: R ,m (e en s pe megabase)=N ,m /(Cm ×Lm DNA)/Nsam, whe e N ,m is he o al numbe o junc ions o soma ic mi o- chond ial-nuclea DNA usions iden i ied, C m is he median alue o mi ochond ial genome copy numbe s in a diploid cance cell calcula ed abo e (495 copies), and L m DNA is he leng h o he mi- ochond ial genome in megabases (0.016569 Mb). Assessmen o he a es o nuclea m DNA usion and m DNA escape o he nucleus Fusion o m DNA o he nuclea genome equi es a leas wo e en s, each o which could in luence he a e o mi ochond ial- nuclea DNA usion.Theseincludeescapeo m DNA o henucleus and in eg a ion o nuclea DNA. Acco ding o his model, he o e - all numbe o such usion e en s can be calcula ed using he a es o hese p ocesses (ρ escape and ρ in eg a ion , espec i ely): N =Nsam ×Ngen × escape × in eg a ion, whe e N is he numbe o o al soma ic mi ochond ial-nu- clea DNA usion e en s (n= 12), Nsam is he o al numbe o can- ce issues (n= 587), and Ngen is he numbe o a e age cell gene a ion om he e ilized egg. Using a easonable assump ion ha Ngen= 1000, we ob ain he a e o soma ic m DNA usion o he nuclea genome (ρ escape ×ρ in eg a ion ) obe2×10 −5 pe cell pe cell gene a ion (pcpg). Wi h one mo e e y conse a i e assump- ion ha ρ in eg a ion is 0.1, we ob ain ρ escape o be 2 × 10 −4 pcpg, o a leas one escape e en pe 5000 cell gene a ions. We hypo h- esize ha he eal ρ in eg a ion alue is hough o be much lowe han 0.1, which esul s in a highe ρ escape . Fo example, du ing he gene a ion o knockou mice, homologous ecombina ion al- lows one ixa ion e en pe 1000–10,000 mic oinjec ed DNA cop- ies (B ins e e al. 1985). The in eg a ion a e may, howe e , be highe han he a e in cance cells wi h de ec i e homologous e- combina ion-based epai and inc eased a ailabili y o nuclea double-s and b eaks, which can be joined o by NHEJ o MMBIR. The m DNA usion o he nuclea genome in he ge mline ( he a e o num s inse ion) is a ound 5 × 10 −6 pe ge m cell pe indi idual gene a ion in p e ious phylogene ic s udies (Hazkani- Co o e al. 2010). The a e is equi alen o ∼5×10 −8 pcpg, gi en ha he numbe o ge m cell di isions pe human gene a ion is ∼100 (401 in males and 31 in emales [D os and Lee 1995]). Da a access Sequence da a o sample pai s wi h posi i e m DNA nuclea ans- e ha e been submi ed o he Eu opean Genome-phenome A chi e (EGA; h ps://www.ebi.ac.uk/ega/home). The s udy acces- sion numbe is EGAS00001001234. Sample accession numbe s a e a ailable in Supplemen al Table 1. Lis o a ilia ions 1 Cance Genome P ojec , Wellcome T us Sange Ins i u e, Hinx on, Camb idge CB10 1SA, Uni ed Kingdom; 2 Cy ogene ics Facili y, Wellcome T us Sange Ins i u e, Hinx on, Camb idge CB10 1SA, Uni ed Kingdom; 3 Camb idge Uni e si y Hospi als NHS Founda ion T us , Camb idge CB2 0QQ, Uni ed Kingdom; 4 Depa men o Haema ology, Uni e si y o Camb idge, Camb idge CB2 0XY, Uni ed Kingdom; 5 Cance Resea ch UK (CRUK) Camb idge Ins i u e, Uni e si y o Camb idge, Camb idge CB2 0RE, Uni ed Kingdom; 6 BioCa e, S a egic Cance Resea ch P og am, SE-223 81 Lund, Sweden; 7 CREATE Heal h, S a egic Cen e o T ansla ional Cance Resea ch, SE-221 00 Lund, Sweden; 8 Depa men o Oncology and Pa hology, Lund Uni e si y Cance Cen e , SE-221 85 Lund, Sweden; 9 B eak h ough B eas Cance Resea ch Uni , Resea ch Oncology, King’s College London, Guy’s Hospi al, London SE1 9RT, Uni ed Kingdom; 10 Labo a o y o In e na ional Alliance on Genomic Resea ch, RIKEN Cen e o In eg a i e Medical Sciences, 230-0045 Yokohama, Japan; 11 Na ional Cen e o Genome Medicine, Ins i u e o Biomedical Sciences, Academia Sinica, Taipei 115, Taiwan; 12 Depa men o Labo a o y Medicine, Helen Dille Family Comp ehensi e Cance Cen e , Uni e si y o Cali o nia, San F ancisco, Cali o nia 94158, USA; 13 Ne he lands Cance Ins i u e, 1066 CX Ams e dam, Ne he lands; 14 Depa men o Gene al Su ge y, Singapo e Gene al Hospi al, Singapo e 169608; 15 Depa men o Molecula Oncology, B i ish Columbia Cance Agency, Vancou e V5Z 1L3, Canada; 16 Depa men o Radia ion Oncology and Depa men o Labo a o y Medicine, Radboud Uni e si y Medical Cen e , 6525 HP Nijmegen, Ne he lands; 17 Depa men o Medical Oncology, E asmus MC Cance Ins i u e, E asmus Ju e al. 822 Genome Resea ch www.genome.o g Cold Sp ing Ha bo Labo a o y P ess on Oc obe 16, 2016 - Published by genome.cshlp.o gDownloaded om