scieee Open visual document viewer

Segmentum : a tool for copy number analysis of cancer genomes

Afyounian, Ebrahim,Annala, Matti,Nykter, Matti

Abstract

BioMed Central open access

Full text

SOFTWARE Open Access Segmen um: a ool o copy numbe analysis o cance genomes Eb ahim A younian, Ma i Annala and Ma i Nyk e * Abs ac Backg ound: Soma ic al e a ions, including loss o he e ozygosi y, can a ec he exp ession o oncogenes and umo supp esso genes. Whole genome sequencing enables de ailed cha ac e iza ion o such abe a ions. Howe e , due o he limi a ions o cu en high h oughpu sequencing echnologies, his ask emains challenging. Hence, accu a e and eliable de ec ion o such e en s is c ucial o he iden i ica ion o cance - ela ed al e a ions. Resul s: We in oduce a new ool called Segmen um o de e mining soma ic copy numbe s using whole genome sequencing om pai ed umo /no mal samples. In ou app oach, ead dep h and B-allele ac ion signals a e smoo hed, and double sliding windows a e used o de ec b eakpoin s, which makes ou app oach as and s aigh o wa d. Because he b eakpoin de ec ion is pe o med simul aneously a di e en scales, i allows accu a e de ec ion as sugges ed by he e alua ion esul s om simula ed and eal da a. We applied Segmen um o pai ed umo /no mal whole genome sequencing samples om 38 pa ien s wi h low-g ade glioma om he TCGA da ase and we e able o con i m he ecu ence o copy-neu al loss o he e ozygosi y in ch omosome 17p in low-g ade as ocy oma cha ac e ized by IDH1/2 mu a ion and lack o 1p/19q co-dele ion, which was p e iously epo ed using SNP a ay da a. Conclusions: Segmen um is an accu a e, use - iendly ool o soma ic copy numbe analysis o umo samples. We demons a e ha his ool is sui able o he analysis o la ge coho s, such as he TCGA da ase . Keywo ds: Soma ic copy numbe analysis, Loss o he e ozygosi y, Segmen a ion, Whole-genome sequencing, Cance Backg ound Soma iccopynumbe al e a ions(SCNA)a eag oup o genomic abe a ions commonly obse ed in many cance s [1]. Copy numbe is he numbe o copies pe cell o a pa icula gene o DNA sequence. Som- a ically acqui ed ch omosomal ea angemen s such as dele ions and duplica ions may change he copy numbe o a gene. Consequen ly, he exp ession le el o a gene is o en co ela ed wi h i s copy numbe [2] - a phenomenon known as he gene dosage e ec . Loss o he e ozygosi y (LOH) is an e en in which one o he wo alleles a a he e ozygous locus is los due o segmen al aneuploidy, gene con e sion, mi o ic ecombina ion, o mi o ic nondisjunc ion [3]. LOH e en s in ol ing umo supp esso genes such as PTEN,RB1, and TP53 ha e been obse ed in many cance . LOH may al e gene exp ession. Fo example, monoallelic exp ession (MAE), which is he exp es- sion o a gene om only one o wo alleles in a dip- loid o ganism, is associa ed wi h LOH [3]. By analyzing a coho o 23 iple-nega i e b eas cance pa ien s,Hae al.[3]ha eshown ha LOHisa p ominen abe a ion in his ype o cance , and mod- ula es a signi ican po ion o he ansc ip ome in he o m o MAE. Copy-neu al LOH (cnLOH) is a speci ic ype o LOH ha occu s when he los allele is eplaced wi h a duplica ed copy o he su i ing allele, esul ing in he copy numbe emaining unchanged. Suzuki e al. ha e shown ecu ing cnLOH a ch omosome 17p (ha bo ing TP53 gene) in low-g ade as ocy oma [4]. The al e ed exp ession o genes wi h allelic imbalance due o LOH e en s may b ing abou selec i e ad an ages o umo igenesis and umo p og ession. Addi ionally, egions wi h cnLOH may ha bo genes wi h d i e mu a ions [5]. Hence, accu a e and eliable de ec ion and cha ac e iza ion o e en s, such as SCNAs and LOH, * Co espondence: [email p o ec ed] Facul y o Medicine and Li e Sciences and BioMediTech ins i u e, Uni e si y o Tampe e, Tampe e, Finland © The Au ho (s). 2017 Open Access This a icle is dis ibu ed unde he e ms o he C ea i e Commons A ibu ion 4.0 In e na ional License (h p://c ea i ecommons.o g/licenses/by/4.0/), which pe mi s un es ic ed use, dis ibu ion, and ep oduc ion in any medium, p o ided you gi e app op ia e c edi o he o iginal au ho (s) and he sou ce, p o ide a link o he C ea i e Commons license, and indica e i changes we e made. The C ea i e Commons Public Domain Dedica ion wai e (h p://c ea i ecommons.o g/publicdomain/ze o/1.0/) applies o he da a made a ailable in his a icle, unless o he wise s a ed. A younian e al. BMC Bioin o ma ics (2017) 18:215 DOI 10.1186/s12859-017-1626-8 a e c ucial o he iden i ica ion o p ospec i e cance - ela ed genes, such as umo supp esso genes and oncogenes, and e en ually o in o ming new app oaches o ea cance [6]. High h oughpu sequencing (HTS)-based SCNA de ec ion app oaches (including bo h whole exome sequencing (WES) and whole genome sequencing (WGS)) ha e become popula due o hei po en ial o accu a e copy numbe es ima ion and b eakpoin de ec ion wi h single nucleo ide accu acy. Howe e , he sho ead leng h o cu en HTS echnologies makes i di icul o map some eads o unique loca ions in he genome. Fu he - mo e, due o GC-con en bias, GC-con en - ich egions in he genome will ha e inc eased numbe o eads. These ambigui ies make accu a e es ima ion o co e age and consequen ly copy numbe a challenge [7]. Addi ionally, umo ploidy and no mal cell con amina ion in oduce u he challenges in SCNA de ec ion [8]. HTS-based copy numbe analysis is, in mos cases, based on ead dep h (RD) es ima ions a each gen- omic loca ion and u he segmen a ion and quan i i- ca ion o he RD p o iles in o segmen s o consis en copy numbe (Addi ional ile 1: Table S1 o a lis o SCNA ools) [9, 10]. Howe e , such ools a e only capable o de ec ing dele ions and duplica ions. Recen ly, RD-based analysis has been augmen ed o iden i y cnLOH e en s by inco po a ing in o ma ion om an al e na e allele’s ac ion a he e ozygous single nucleo ide poly- mo phism (SNP) posi ions (o B-allele ac ion (BAF)). The BAF o a he e ozygous SNP has an expec ed alue o 0.5 in no mal diploid cells. De ia ion om 0.5 in he he e ozy- gous SNP BAF poin s o an abe a ion. In he case o cnLOH, BAF alues a e expec ed o be ei he 0 o 1 in a pu e umo popula ion. Tools such as Con ol-FREEC [11], Pa chwo k [12], and CLImAT [13] inco po a e BAF da a o ex end SCNA de ec ion. Con ol-FREEC de e - mines he b eakpoin s using a leas absolu e sh inkage es ima o (LASSO) eg ession. Sample ploidy is p o ided by he use o Con ol-FREEC. I also e alua es and co - ec s o no mal cell con amina ion, GC-con en , and map- abili y biases while in e ing he copy numbe p o ile o a umo genome. Pa chwo k pe o ms GC and posi ional no maliza ion and segmen s he genome using a ci cula bina y segmen a ion (CBS) algo i hm. I also es ima es no - mal cell con amina ion and umo ploidy. CLImAT imple- men s co ec ions o GC-con en and mapabili y bias and models he RD and BAF da a wi h a hidden Ma ko model (HMM) o in e he soma ic copy numbe a ia ion, no - mal cell con amina ion and umo ploidy (Addi ional ile 1: O e iew o Tools sec ion o mo e de ails on hese ools). While he abo e ools a e well-sui ed o SCNA de ec ion, hei use has some limi a ions. Con ol-FREEC and Pa ch- wo k u ilize compu a ionally cos ly models, which leads o long analysis imes. The main mo i a ion o ou s udy was o de elop an accu a e and use - iendly ool ha could be used o analyze la ge WGS da ase s, such as he cance genome a las (TCGA) da ase s. In ou app oach, he RD and BAF signals a e smoo hed, and double sliding windows subsequen ly a e used o de ec b eakpoin s, which makes ou app oach as and s aigh o wa d. Because he b eak- poin de ec ion is pe o med simul aneously a di e en scales, i allows accu a e de ec ion. Ou ool, Segmen um, is eely a ailable unde MIT license a : h ps://gi hub.- com/ea younian/Segmen um (Addi ional ile 2 con ains he so wa e code. Fo he la es e sion o he so wa e code please isi he p ojec ’s online eposi o y). Implemen a ion Pipeline Segmen um was de eloped and w i en in he Py hon p og amming language ( e sion 3) and equi es he SciPy lib a y o be ins alled (I he use wishes o use he ‘plo ’sub-command o in o m pa ame e alue selec ion, ma plo lib lib a y is also equi ed). Segmen um employs SAM ools o ex ac RD and he e ozygous SNPs BAF da a om BAM iles con aining WGS da a. These con- s i u e he inpu s equi ed by Segmen um o pe o m copy numbe analysis. Figu e 1 illus a es he Fig. 1 Segmen um pipeline. No mal and umo RDs a e used o calcula e RD log- a ios. RD log- a ios a e hen co ec ed o biases. BAF da a a e simul aneously mi o ed and smoo hed. Using RD log- a ios and BAF, he genome is segmen ed wi h a double sliding window me hod. Segmen a ion esul s a e used o iden i y cnLOH egions in he genome (see he ollowing sec ions o mo e de ails on each s ep) A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 2 o 10 Segmen um pipeline. Each s ep is explained in mo e de- ail in he ollowing sec ions. RD ex ac ion and BAF calcula ion o he e ozygous SNPs To ex ac he RD om he BAM iles, he genome is di ided in o bins o use -de ined leng h (2 kbp by de aul ) and he numbe o eads o e lapping each bin is coun ed o de e mine he RD a each bin. To calcula e he BAF alues, he e ozygous SNPs in he no mal sam- ple a e iden i ied a known SNP si es in he human gen- ome (based on SNP anno a ions such as hose p oduced by he 1000 Genomes p ojec ). Nex , he numbe o e - e ence and al e na i e alleles a each he e ozygous SNP posi ion is ex ac ed om he umo sample and he BAF o he i h he e ozygous SNP is calcula ed using he ollowing equa ion: BAFi¼al i al iþ e i whe e al i and e i e e o he al e na i e and e e ence allele, espec i ely, o he i h he e ozygous SNP. I should be no ed ha by de aul eads wi h mapping quali y sco e 10 a e il e ed ou be o e RD ex ac ion and BAF calcula ion in o de o add ess he challenges aised by eads no mapping o a unique egion in he genome ( he ead il a ion c i e ion based on he map- ping quali y sco e is a pa ame e o Segmen um and can be se by he use ). Log- a io calcula ion The RD log- a io is calcula ed using he ollowing equa ion: log i¼log2 RDi nRDi  whe e log i is he log- a io o he i h genomic window and RD i and nRD i a e RDs ex ac ed om he i h gen- omic window o a speci ic size (de e mined by use ; de aul is 2 kbp) o he umo and no mal samples, espec i ely. Di e ences in he o al numbe o aligned eads in he no mal and umo samples may bias he es ima ion o he RD log- a ios. The co ec ion was pe o med by inding he mode o log- a io alues o each ch omo- some and sub ac ing he median o all o he modes om each log- a io alue. I should be no ed ha median, in he co ec ion s ep, is obus o he changes in one mode. Fo ins ance, one ch omosomal a m ha ing a copy numbe change has no e ec on he co ec ion since i only a ec s one o he ch omosomal modes. Mi o ing and smoo hing o he BAF alues The BAF o a he e ozygous SNP has an expec ed alue o 0.5 in no mal diploid cells. In he p esence o soma ic copy numbe al e a ions, he BAF can di e ge om 0.5 i he ela i e abundance o he wo alleles changes. To make smoo hing and segmen a ion o BAF da a possible, he BAF alues mus be mi o ed abou he 0.5 axis so ha he B allele ac ion always ep esen s he allele ac ion o he dominan allele. Wi hou his mi o ing s ep, he BAF alues will be symme ic abou he BAF = 0.5 axis and smoo hing will unde es ima e he absolu e di e gence om 0.5 [14]. In his s udy, a median il e is used o smoo hing he BAF da a. Simul aneous mi o - ing and smoo hing is implemen ed using he ollowing equa ion: cBAFi¼H0:5−M9BAFi ðÞ jj þ1−HðÞM90:5−BAFi jj ðÞ: whe e BAF i is he BAF alue o he i h he e ozygous SNP, cBAF i is he simul aneously mi o ed and smoo hed BAF i , His a he e ozygosi y measu emen calcula ed wi h he ollowing equa ion: H=1−2∗|0.5 −x|, and M 9 e e s o applying a median il e o 9 SNPs in he icini y o and including he i h SNP. Segmen a ion using a double sliding window app oach To de ec changes in he RD log- a io and BAF signals, wo non-o e lapping, ixed-sized windows (de e mined by he use ) a e slid o e he RD log- a io and BAF alues and a compound sco e (S) is calcula ed o each o he adjacen wo windows. I he compound sco e is g ea e han 1, a change is de ec ed and a b eakpoin is placed a he place whe e he wo windows ouch each o he . The compound sco e is calcula ed using he ollowing equa ion: S¼log wini−log winiþ1 2 τlog þcBAFwini−cBAFwiniþ1 2 τBAF whe e log winiis he mean o he RD log- a io alues in he i h window, cBAFwiniis he mean o he mi o ed and smoo hed BAF alues in he i h window, τ log and τ BAF a e h esholds o he absolu e mean di e ence in he RD log- a ios and he absolu e mean di e ence in he BAF alues in he wo adjacen windows, espec i ely. I is possible ha some b eakpoin s will no be de- ec ed by a single pass o a double sliding window due o a gi en window size. Thus, o inc ease he sensi i i y, Segmen um analyzes he signals o he de ec ion o b eakpoin s mul iple imes wi h di e en window-sizes and h esholds. Each new window is 1.5 imes la ge han he p e ious one. The inc ease in he window size dec eases he de ec ion h esholds. This is due o he ac ha inc easing he window size inc eases he sample A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 3 o 10 size (assuming sampling om no mal dis ibu ion wi h N(μ,σ 2 )) and consequen ly dec eases he s anda d de i- a ion o he mean (mean ha ing p obabili y dis ibu ion o Nμ;σ2.n  ). The new s anda d de ia ion o he mean when window size is inc eased 1.5 imes is 1 ffiffiffiffiffi 1:5 2 p imes he old s anda d de ia ion. Le τ=ασ whe e τis he h eshold and αis a scala and σis he s anda d de- ia ion. I ollows ha : τnew ¼α:σnew ¼1 ffiffiffiffiffiffiffi 1:5 2 p:α:σold ¼1 ffiffiffiffiffiffiffi 1:5 2 p:τold Thus bo h he τ log and τ BAF h esholds a e upda ed using he ollowing equa ion: τnew ¼1 ffiffiffiffiffiffiffi 1:5 2 pτold The p ocess o inc easing he window-size is con in- ued as long as he upda ed h esholds a e g ea e han he h esholds o he me ging wo consecu i e seg- men s (see below). A e de ec ing all he b eakpoin s, a consensus lis o b eakpoin s is c ea ed by accep ing all o he b eakpoin s de ec ed by he i s pass o he double sliding window and adding he b eakpoin s de- ec ed om he la ge windows o he lis only i he b eakpoin is no in he icini y o an exis ing b eakpoin in he lis (i.e., |cp cu en −cp exis ing |>window size, whe e cp is a de ec ed b eakpoin ). Consensus b eakpoin s a e used o c ea e he segmen s. Two consecu i e b eak- poin s cons i u e a segmen . Fo each segmen , he a e - age RD log- a io and a e age mi o ed and smoo hed BAF is calcula ed. Two consecu i e segmen s a e me ged i he ollowing condi ions a e me : log segi −log segiþ1 <τme gelog and jcBAFsegi −cBAFsegiþ1j<τme geBAF whe e log segiis he mean RD log- a io o he i h seg- men , cBAFsegiis he mean mi o ed and smoo hed BAF o he i h segmen , and τme gelog and τme geBAF (de e mined by use ) a e he RD log- a io and BAF me ging h esh- olds, espec i ely. De ec ion o cnLOH e en s wi hin a single sample A segmen is conside ed o be a cnLOH segmen i he ollowing condi ions a e me : log segi <τcnLOHlog and 0:5−cBAFsegi  <τcnLOHBAF whe e log segiis he mean RD log- a io o he i h seg- men , cBAFsegiis he mean mi o ed and smoo hed BAF o he i h segmen , τcnLOHlog and τcnLOHBAF (de e mined by he use ) a e h esholds o calling a cnLOH segmen . De ec ion o ecu en cnLOH egions ac oss mul iple samples To ind genomic egions wi h ecu en cnLOH e en s, all cnLOH egions o indi idual samples a e iden i ied ollowing he p ocedu e desc ibed ea lie . Then, he numbe o occu ences o a cnLOH e en o a speci ic egion ac oss mul iple samples is coun ed using an in e al ee da a s uc u e (Addi ional ile 1: Figu e S1). Simula o To e alua e Segmen um in e ms o segmen a ion ac- cu acy, a simula o capable o simula ing whole- genome RD o bo h no mal and umo samples and BAF based on e en s such as dele ions, ampli ica ions and cnLOH was de eloped. The simula o ecei es a no mal sample RD da a and ou pu s 4 se s o da a in- cluding he simula ed no mal and umo RD, BAF da a and a g ound u h. Fi s , he simula o lea ns he dis ibu ion o he RD da a om he p o ided no mal sample by simply coun ing he numbe o imes wo consecu i e RD alues (e.g., 368 and 299) occu oge he h oughou he genome (Addi ional ile 1: Figu e S2). The lea ned dis ibu ion also ac- coun s o he inhe en noise in he RD da a. Nex , in e se ans o m sampling (Smi no ans o m) is used o gene a e RD alues o each posi ion in he genome based on he lea ned dis ibu ion. Then, noise is emo ed using a median il e . A no mal RD is cons uc ed by adding independen Poisson noise o he simula ed RD da a. To cons uc he umo RD, wo copy numbe acks (because au osomal ch omosomes come in ma e nal and pa e nal pai s) ha bo ing andom SCNAs a e cons uc ed. The umo sample RD is calcula ed using he copy numbe acks, he simula ed no mal sample RD and he no - mal sample con amina ion (i.e., a pa ame e de e - mined by use ). To cons uc he BAF da a, he e ozygous SNPs a e ini ially andomly dis ibu ed ac oss he genome (1 he e ozygous SNP pe 1.5 Kbp). The numbe o B-alleles a a he e ozygous SNP is cal- cula ed using a binomial dis ibu ion wi h he pa am- e e s n( o al numbe o eads a he e ozygous SNP posi ion) and p(p obabili y ha a ead is coming om he B-allele). nis ex ac ed om he simula ed no mal RD a he e ozygous SNP posi ions. pis calcu- la ed using he wo cons uc ed copy numbe acks and he no mal sample con amina ion. Once he numbe o B-alleles is calcula ed, i is used o calcu- la e he BAF alues (Addi ional ile 1: Figu es S3 and S4 o he simula o pipeline and he simula ed da a isualized in he in eg a i e genomics iewe (IGV) [15], espec i ely). A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 4 o 10 Resul s Segmen um segmen a ion accu acy o he simula ed da a Using he simula o (see he ‘Simula o ’sec ion o mo e de ails), RD da a o bo h no mal and umo samples and BAF alues o he e ozygous SNPs om he umo sample as well as a g ound u h we e simula ed wi h di e en pe cen ages o no mal con amina ion (an ex- ample se o simula ed da a is a ailable a Segmen um’s online eposi o y. See he ‘A ailabili y and equi emen s’ sec ion o he link o he eposi o y). The simula ed da a we e analyzed by Segmen um. The segmen a ion esul s we e e alua ed agains he g ound u h. The p ecision, ecall, and he F-measu e alues we e calcu- la ed based on his e alua ion (Fig. 2 and Addi ional ile 1 o he de ini ions o p ecision, ecall, and F-measu e). Segmen um segmen a ion accu acy o eal da a compa ed o o he ools To assess segmen a ion accu acy o Segmen um o eal da a, pai ed umo /no mal whole genome sequencing samples (30x < co e age < 100x) om 10 indi iduals diagnosed wi h low-g ade glioma (LGG) we e down- loaded om he TCGA da ase and used as is. Fu he - mo e, segmen a ion esul s om SNP-a ay da a (le el 3 da a) (comple ed by TCGA using an A yme ix Genome-wide human SNP a ay 6.0) was used as g ound u h (Addi ional ile 1: Table S3). Segmen um’s esul s we e e alua ed agains Con ol-FREEC, Pa ch- wo k, and CLImAT as compe ing ools. To e alua e he segmen a ion accu acy, he genome was b oken in o 100 bp. blocks (excluding all blocks in cen ome es and sex ch omosomes). Using block anno a ions om di e en ools, genome-wide p opo ions o he blocks anno a ed as SCNA by di e en combina ions o ools we e calcula ed and he esul s we e illus a ed by a Venn diag am (Fig. 3). Addi ionally, o measu e he pai wise deg ee o simi- la i y o he segmen a ion esul s be ween wo ools, he Jacca d simila i y index (JSI) was calcula ed o all o he pai s using he ollowing equa ion: JSI ¼∩pai jj ∪pai jj whe e | ∩pai | and | ∪pai | a e he ca dinali ies o in e - sec ion and union, espec i ely. In e sec ion and union alues we e ex ac ed om he Venn diag ams. Figu e 4 ep esen s a hea map o he JSI alues o each pai o ools a e aged o e 10 TCGA LGG samples. Acco ding o he hea map, on a e age, Segmen um p oduces he mos simila esul s o he SNP a ay segmen a ion esul s wi h a JSI sco e o 0.9, ollowed by Pa chwo k wi h a JSI sco e o 0.86. Simila e alua ions using low co e age da a (6x a e - age co e age) a e shown in Addi ional ile 1: Figu es S5 and S6. The low co e age da a is comp ised o he pai ed umo /no mal whole genome sequencing samples o 10 indi iduals diagnosed wi h p os a e adenoca cin- oma (PRAD). Wi h ega d o he low co e age da a, Pa chwo k p oduces he mos simila esul s o he SNP a ay segmen a ion esul s wi h a JSI sco e o 0.93, ollowed by Segmen um wi h a JSI sco e o 0.88. Add- i ional ile 1: Table S4 con ains he names o he 10 TCGA PRAD samples (Addi ional ile 1: Tables S6-S10 ep esen he pa ame e alues used o unning he Fig. 2 Segmen a ion accu acy o Segmen um o simula ed da a wi h di e en deg ees o no mal con amina ion. Es ima ed p ecision, ecall, and F-measu e alues o simula ed da a a di e en no mal con amina ion le els (Addi ional ile 1, De i a ion o he p ecision, ecall, and F-measu e o he simula ed da a) A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 5 o 10 compe ing ools. Addi ional ile 1: Segmen um’s pa am- e e alue selec ion sec ion p o ides guidance on selec - ing pa ame e alues o Segmen um. Addi ional ile 1: Figu e S9 ep esen s an example plo made by Segmen- um’s 'plo ' sub-command ha can be used o guide he pa ame e alue selec ion). Segmen um segmen a ion accu acy o he subsampled eal da a To assess he segmen a ion accu acy o Segmen um o eal da a wi h espec o sample’s co e age, we subsampled one o he LGG samples (i.e. TCGA-CS- 5395) a di e en subsampling ac ions (i.e. 75%, 50%, 25%, 10%, and 5%) using Sam ools ( e sion 1.3.1). We analyzed each subsample by Segmen um and bench- ma ked i agains g ound u h in he same manne as explained ea lie . Figu e 5 ep esen s he JSI sco es o each subsample (Addi ional ile 1: Figu e S7 shows he a e age co e age o he subsamples o no mal and umo pai s). I can be seen ha Segmen um eaches high accu acies e en wi h low co e age da a. Fo in- s ance, he accu acy o he 10%- ac ion subsample was 93.4% (whe e he a e age co e age o umo and no - mal subsamples we e 3 and 4 espec i ely). I should be no ed ha as he co e age dec eases he numbe o iden i ied he e ozygous SNPs dec eases (Addi ional ile 1: Figu e S8). Fo ins ance, o he 10%- ac ion subsample only 1997 he e ozygous SNPs we e iden i ied om he en i e genome (in con as o he o iginal sample whe e he numbe o iden i ied he e ozy- gous SNPs was mo e han 3 million SNPs). E en hough Segmen um is shown o wo k wi h low co e age da a, one should no e he implica ions o low amoun s o de ec ed he e ozygous SNPs on he eliable de ec ion o cnLOH e en s. Time usage e alua ion All o he compu a ions we e comple ed on he same UNIX se e . Table 1 shows he a e age ime equi ed by each ool o pe o m he analysis o 10 TCGA LGG samples (30x < co e age < 100x). Based on he esul s, on a e age, CLImAT appea s o be he as es , ollowed by Segmen um, Pa chwo k, and Con ol-FREEC. I should be no ed ha o assign he allele-speci ic copy numbe o genomic segmen s, Pa chwo k equi es use s o de e - mine some pa ame e alues by in e p e ing plo s p o- duced by he ool, and his in e p e a ion ime is no included he e. Addi ionally, he ime equi ed o c ea e he pileup iles used by Pa chwo k and Con ol-FREEC is di e en due o he use o di e en pa ame e alues in SAM ools. I should be no ed ha ime equi ed o making pileup iles can be dec eased by pa allelizing he p ocess on machines wi h mul iple co es o on com- pu e clus e s (e.g. by assigning one co e o each ch omosome). Simila ly, BAF calcula ion o Segmen um can be pa allelized. Howe e , since his is no a co e ea- u e o he benchma ked ools and no all ools suppo pa alleliza ion, o be ai , only he equi ed linea ime is epo ed he e. A simila ime usage e alua ion, using low co e age da a (a e age co e age 6x), is shown in Addi ional ile 1: Table S2. Wi h ega d o he low co e - age da a, Segmen um comes second a e CLImAT in e ms o analysis ime, which is consis en wi h he esul s om he high co e age da a. Fig. 3 Compa ison o he SCNA esul s wi h di e en ools and he SNP a ay (g ound u h). Venn diag am alues (a e aged o e en TCGA LGG samples) ep esen he pe cen age o o e lap among he SCNA calls Fig. 4 Pai wise JSI sco es a e aged o e en TCGA LGG samples. JSI sco es ange be ween 0 and 1, whe e 0 means no simila i y and 1 ep esen s iden ical esul s be ween wo ools A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 6 o 10 Recu en cnLOH de ec ion case s udy In a s udy o lowe g ade gliomas (LGGs), i.e., g ade II and III gliomas, Suzuki e al. [4] cha ac e ized he mu a- ional landscape o hese glioma ypes by di iding hem in o 3 dis inc sub ypes based on hei dis inc se s o mu a ions and clinical beha io s. These sub ypes a e dis- inguished wi h he ollowing c i e ia: (1) mu a ion in IDH1/2 accompanied by co-dele ion o ch omosomes 1p and 19q (sub ype I), (2) mu a ion in IDH1/2 wi hou co- dele ion o ch omosomes 1p and 19q (sub ype II), and (3) IDH1/2 wild ype (sub ype III). O in e es o ou s udy was he ecu ence o cnLOH e en s in ch omo- some 17p in sub ype II [4]. To show he abili y o Segmen um o de ec such abe a ions om la ge da a- se s, 38 pai ed-end WGS samples om he TCGA da a- se (30x < co e age < 100x)) o pa ien s diagnosed wi h LGG we e downloaded and analyzed by Segmen um. We we e able o dis inguish all h ee sub ypes as cha ac- e ized in [4], including he ecu ence o cnLOH in sub ype II a ch omosome 17p. We also iden i ied a ou h sub ype wi h a mu a ion in IDH1/2 wi hou co- dele ion o ch omosomes 1p and 19q and no cnLOH a 17p (Fig. 6). Discussion By compa ing he simula ed (Fig. 2) and eal da a (Figs. 3, 4 and 5 and Addi ional ile 1: Figu es S5 and S6), we can conclude ha Segmen um can eco e ue copy numbe abe a ions wi h high accu acy e en when he co e age is as low as ~4 eads (Fig. 5, Addi ional ile 1: Figu e S7). On a e age, Segmen um p oduces esul s ha a e he mos conco dan wi h he copy numbe ab- e a ions iden i ied om he SNP a ay da a (i.e. ~90% o conco dance) (Fig. 4). As shown in Table 1, ou ool is mo e han wice as as as he second bes pe o ming ool in e ms o accu acy. Segmen um is also he second as es ool a e CLImAT compa ed o he o he ools e alua ed in his s udy (Table 1). Howe e , CLImAT anks las in e ms o accu acy (Fig. 4). One explana ion o he speed o CLImAT is ha i compu es he BAF alues o a subse o known SNPs (~13.7 million SNPs ha a e e ie ed om he dbSNP da abase [16]). In con as , Segmen um, compu es he BAF alues o he - e ozygous SNPs de e mined om he 1000 Genomes p ojec ’s SNP lis (~85 million SNPs) [17]. The o he eason o he speed o CLImAT migh be ha i does no equi e a no mal sample o analysis. As he no mal con amina ion in he simula ed da a in- c eases, he numbe o alse nega i es inc eases and he ecall a e dec eases (Fig. 2). Howe e , wi hin he anges o ealis ic amoun s o no mal con amina ion (i.e. ~30% o 40%), Segmen um pe o ms consis en ly well. Segmen um is able o epo ecu en cnLOH egions ac oss mul iple cance genome samples; a cha ac e is ic Fig. 5 Pai wise JSI sco es (Segmen um s. SNP a ay as g ound u h) o di e en subsamples. JSI sco es ange be ween 0 and 1, whe e 0 means no simila i y and 1 ep esen s iden ical esul s be ween wo ools Table 1 A e age ool analysis ime o high co e age da a (30x < co e age < 100x) Tool A e age p epa a ion ime A e age analysis ime Segmen um - 10 h 34 min o ex ac ing RD om no mal o umo BAM ile - 4 h 25 min o calcula ing BAF alues - 1 min 45 s Pa chwo k - 29 h 37 min o c ea ing pileups om no mal o umo BAM ile - 3 h 56 min Con ol-FREEC - 33 h 28 min o c ea ing pileups om no mal o umo BAM ile - 7 h 11 min CLImAT - 2 h 12 min o ex ac ing RD - 29 min A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 7 o 10 Fig. 6 SCNA landscape in g ade II and III gliomas. WHO-g ade, his ological class, and molecula sub ype classi ica ion a e shown by colo as indica ed. The hi y-eigh samples a e di ided in o 4 dis inc sub ypes based on he occu ence o a mu a ion in IDH1/2, co-dele ion o ch omosomes 1p and 19q and he p esence o 17p cnLOH. Dele ions and ampli ica ions a e isualized by boxes wi h di e en shades o blue and ed, espec i ely. Whi e egions a e ei he no mal o cnLOH egions. The ba cha s below each box ep esen he mi o ed and smoo hed BAF alues. La ge mi o ed and smoo hed BAF alues (close o 0.5) poin o he e ozygous SNP allelic imbalance. In he second sub ype ( om he op), a ch omosome 17p, ecu ing cnLOH is appa en whe e he ba cha s poin o la ge mi o ed and smoo hed BAF alues, hough no dele ion o ampli ica ion is de ec ed a ha egion (Addi ional ile 1: Table S5 o TCGA LGG sample ba code names) A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 8 o 10 o cance genomes ha has been neglec ed un il ecen ly [4]. By applying Segmen um o TCGA da a, we we e able o eco e ecu en cnLOH e en s om low-g ade glioma samples ha we e epo ed ea lie by SNP a ay- based da a analysis. I is wo h men ioning ha Segmen- um can wo k in wo modes, i.e., wi h o wi hou BAF alue. In he case whe e BAF alues a e no used, Seg- men um canno de ec egions wi h cnLOH. Fu he - mo e, Segmen um is capable o eliably segmen ing he cance genome using bo h high (Figs. 3 and 4) and low (Fig. 5 and Addi ional ile 1: Figu es S5 and S6) sequence co e age da a. Howe e , wi h he low sequence co e age da a, he es ima ed BAF alues o he he e ozygous SNPs will be less eliable. This is also e lec ed in Addi ional ile 1: Figu e S8, whe e i is shown ha he numbe o de ec ed he e ozygous SNPs d op as he a e - age co e age dec eases. The implica ions o low amoun s o de ec ed he e ozygous SNPs on he eliable de ec ion o cnLOH e en s should no be o e looked. E en hough we ha e shown ha Segmen um is highly accu a e a eco e ing he ue copy numbe , o he ools in his s udy do mo e han jus segmen - ing he genome. Fo ins ance, CLImAT and Pa ch- wo k a e capable o es ima ing umo ploidy and umo pu i y and consequen ly, epo ing he in eg al copy numbe s o each segmen . Pa chwo k and Con ol-FREEC a e also capable o epo ing he geno ype o each segmen and CLImAT epo s he geno ype o each SNP wi hin each segmen . This is in con as o Segmen um ha only epo s he mean RD log- a io and BAF alue o each segmen . How- e e , ools such as ABSOLUTE [18] o THe A [19] can be used o es ima e umo impu i y and ploidy om Segmen um’s segmen a ion esul , meaning ha Segmen um can be used as pa o a la ge umo e olu ion analysis pipeline. Finally, a s eng h o ou ool is i s minimum dependence on hi d pa y ools, wi h he excep ion o SAM ools, o calcula ing he RD and BAF. Conclusions We ha e de eloped Segmen um as a ool o he iden i ica ion o SCNAs, including cnLOH in umo samples, using WGS da a. We ha e shown ha Seg- men um is accu a e and as wi h ega ds o o he s a e-o - he-a ools, making i sui able o analyzing coho s wi h a la ge numbe o samples, such as TCGA coho s. A ailabili y and equi emen s P ojec name: Segmen um P ojec homepage: h ps://gi hub.com/ea younian/Segm en um Ope a ing sys em(s): Linux P og amming language: Py hon O he equi emen s: SciPy, Sam ools, and ma plo lib i he ‘plo ’sub-command is used. License: MIT license Any es ic ions o use by non-academics: None Addi ional iles Addi ional ile 1: This ile con ains supplemen a y in o ma ion, ables and igu es suppo ing he manusc ip . Figu e S1. De ec ion o egions ha bo ing ecu en cnLOH ac oss mul iple samples. Figu e S2. Read dep h spa ial co ela ion. Figu e S3. Simula o pipeline. Figu e S4. Simula ed da a isualized in In eg a i e Genomics Viewe (IGV). Figu e S5. Compa ison o SCNA esul s om di e en ools and SNP a ay (g ound u h) o low sequence co e age da a. Figu e S6. Pai wise JSI sco es o low sequence co e age da a (a e aged o 10 TCGA PRAD samples) Figu e S7. Subsample a e age co e ages in he subsampling e alua ion. Figu e S8 De ec ed numbe o he e ozygous SNPs in di e en subsamples Figu e S9. Copy numbe –B-allele ac ion clus e s. Table S1. Lis o SCNA ools using WGS da a. Table S2. A e age ool analysis ime o low sequence co e age da a (a e age co e age 6x). Table S3 TCGA LGG sample ba code names and he es ima ed sample pu i y by ABSOLUTE. Table S4. TCGA PRAD sample ba code names and he es ima ed sample pu i y by ABSOLUTE. Table S5. TCGA LGG sample ba code names ca ego ized based on in e ed sub ype. Table S6. Pa ame e alues o unning DFEx ac . Table S7. Pa ame e alues o unning CLImAT. Table S8. Pa ame e alues o unning Pa chwo k o 10 TCGA LGG samples. Table S9. Pa ame e alues o unning Pa chwo k o 10 TCGA PRAD samples. Table S10. Pa ame e alues o unning Con ol- FREEC. (DOCX 857 kb) Addi ional ile 2: So wa e code. This comp essed ile con ains he so wa e code ( o he la es e sion o he so wa e code please isi he p ojec ’s online eposi o y). (ZIP 32 kb) Abb e ia ions BAF: B-Allele ac ion; BAM: Bina y alignmen map; CBS: ci cula bina y segmen a ion; CGH: A ay compa a i e genomic hyb idiza ion; cnLOH: Copy-neu al loss o he e ozygosi y; CNV: Copy numbe a ia ion; dbSNP: Single nucleo ide polymo phism da abase; DNA: Deoxy ibonucleic acid; FISH: Fluo escen in si u hyb idiza ion; HMM: Hidden Ma ko model; HTS: High h oughpu sequencing; IGV: In eg a i e genomics iewe ; JSI: Jacca d simila i y index; LASSO: Leas absolu e sh inkage eS ima O ; LGG: Low g ade glioma; LOH: Loss o he e ozygosi y; MAE: Monoallelic exp ession; PRAD: PRos a e ADenoca cinoma; RD: Read-dep h; SAM: Sequence alignmen /map; SCNA: Soma ic copy numbe al e a ion; SNP: Single nucleo ide polymo phism; TCGA: The cance genome a las; WES: Whole exome sequencing; WGS: Whole genome sequencing Acknowledgmen s The esul s published he e a e in pa based upon da a gene a ed by The Cance Genome A las p ojec (dbGaP S udy Accession: phs000178. 9.p8) es ablished by he NCI and NHGRI. In o ma ion abou TCGA and he in es iga o s and ins i u ions who cons i u e he TCGA esea ch ne wo k can be ound a h p://cance genome.nih.go . The au ho s would also like o acknowledge CSC —IT Cen e o Science L d. (h ps://www.csc. i/csc) o p o iding he compu a ional esou ces. Funding This wo k was unded by he Academy o Finland (p ojec no. 269474), Cance Socie y o Finland, and Sig id Juselius Founda ion. The unde s had no ole in s udy design, da a collec ion and analysis, decision o publish, o p epa a ion o he manusc ip . A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 9 o 10