scieee Science in your language
[en] (orig)

Segmentum : a tool for copy number analysis of cancer genomes

Abstract

BioMed Central open access

Read accessible full text

Segmentum : a tool for copy number analysis of cancer genomes

Author: Afyounian, Ebrahim,Annala, Matti,Nykter, Matti
Year: 2017
Source: https://trepo.tuni.fi/bitstream/10024/101040/1/segmentum_a_tool_for_2017.pdf
SOFTWARE Open Access
Segmen um: a ool o copy numbe
analysis o cance genomes
Eb ahim A younian, Ma i Annala and Ma i Nyk e
*
Abs ac
Backg ound: Soma ic al e a ions, including loss o he e ozygosi y, can a ec he exp ession o oncogenes and
umo supp esso genes. Whole genome sequencing enables de ailed cha ac e iza ion o such abe a ions.
Howe e , due o he limi a ions o cu en high h oughpu sequencing echnologies, his ask emains challenging.
Hence, accu a e and eliable de ec ion o such e en s is c ucial o he iden i ica ion o cance - ela ed al e a ions.
Resul s: We in oduce a new ool called Segmen um o de e mining soma ic copy numbe s using whole genome
sequencing om pai ed umo /no mal samples. In ou app oach, ead dep h and B-allele ac ion signals
a e smoo hed, and double sliding windows a e used o de ec b eakpoin s, which makes ou app oach
as and s aigh o wa d. Because he b eakpoin de ec ion is pe o med simul aneously a di e en scales,
i allows accu a e de ec ion as sugges ed by he e alua ion esul s om simula ed and eal da a. We applied
Segmen um o pai ed umo /no mal whole genome sequencing samples om 38 pa ien s wi h low-g ade glioma
om he TCGA da ase and we e able o con i m he ecu ence o copy-neu al loss o he e ozygosi y in ch omosome
17p in low-g ade as ocy oma cha ac e ized by IDH1/2 mu a ion and lack o 1p/19q co-dele ion, which was p e iously
epo ed using SNP a ay da a.
Conclusions: Segmen um is an accu a e, use - iendly ool o soma ic copy numbe analysis o umo samples. We
demons a e ha his ool is sui able o he analysis o la ge coho s, such as he TCGA da ase .
Keywo ds: Soma ic copy numbe analysis, Loss o he e ozygosi y, Segmen a ion, Whole-genome sequencing, Cance
Backg ound
Soma iccopynumbe al e a ions(SCNA)a eag oup
o genomic abe a ions commonly obse ed in many
cance s [1]. Copy numbe is he numbe o copies
pe cell o a pa icula gene o DNA sequence. Som-
a ically acqui ed ch omosomal ea angemen s such
as dele ions and duplica ions may change he copy
numbe o a gene. Consequen ly, he exp ession le el
o a gene is o en co ela ed wi h i s copy numbe [2]
- a phenomenon known as he gene dosage e ec .
Loss o he e ozygosi y (LOH) is an e en in which
one o he wo alleles a a he e ozygous locus is los
due o segmen al aneuploidy, gene con e sion, mi o ic
ecombina ion, o mi o ic nondisjunc ion [3]. LOH
e en s in ol ing umo supp esso genes such as
PTEN,RB1, and TP53 ha e been obse ed in many
cance . LOH may al e gene exp ession. Fo example,
monoallelic exp ession (MAE), which is he exp es-
sion o a gene om only one o wo alleles in a dip-
loid o ganism, is associa ed wi h LOH [3]. By
analyzing a coho o 23 iple-nega i e b eas cance
pa ien s,Hae al.[3]ha eshown ha LOHisa
p ominen abe a ion in his ype o cance , and mod-
ula es a signi ican po ion o he ansc ip ome in
he o m o MAE. Copy-neu al LOH (cnLOH) is a
speci ic ype o LOH ha occu s when he los allele
is eplaced wi h a duplica ed copy o he su i ing
allele, esul ing in he copy numbe emaining
unchanged. Suzuki e al. ha e shown ecu ing
cnLOH a ch omosome 17p (ha bo ing TP53 gene) in
low-g ade as ocy oma [4]. The al e ed exp ession o
genes wi h allelic imbalance due o LOH e en s may
b ing abou selec i e ad an ages o umo igenesis
and umo p og ession. Addi ionally, egions wi h
cnLOH may ha bo genes wi h d i e mu a ions [5].
Hence, accu a e and eliable de ec ion and
cha ac e iza ion o e en s, such as SCNAs and LOH,
* Co espondence: [email p o ec ed]
Facul y o Medicine and Li e Sciences and BioMediTech ins i u e, Uni e si y
o Tampe e, Tampe e, Finland
© The Au ho (s). 2017 Open Access This a icle is dis ibu ed unde he e ms o he C ea i e Commons A ibu ion 4.0
In e na ional License (h p://c ea i ecommons.o g/licenses/by/4.0/), which pe mi s un es ic ed use, dis ibu ion, and
ep oduc ion in any medium, p o ided you gi e app op ia e c edi o he o iginal au ho (s) and he sou ce, p o ide a link o
he C ea i e Commons license, and indica e i changes we e made. The C ea i e Commons Public Domain Dedica ion wai e
(h p://c ea i ecommons.o g/publicdomain/ze o/1.0/) applies o he da a made a ailable in his a icle, unless o he wise s a ed.
A younian e al. BMC Bioin o ma ics (2017) 18:215
DOI 10.1186/s12859-017-1626-8
a e c ucial o he iden i ica ion o p ospec i e
cance - ela ed genes, such as umo supp esso genes
and oncogenes, and e en ually o in o ming new
app oaches o ea cance [6].
High h oughpu sequencing (HTS)-based SCNA
de ec ion app oaches (including bo h whole exome
sequencing (WES) and whole genome sequencing (WGS))
ha e become popula due o hei po en ial o accu a e
copy numbe es ima ion and b eakpoin de ec ion wi h
single nucleo ide accu acy. Howe e , he sho ead leng h
o cu en HTS echnologies makes i di icul o map
some eads o unique loca ions in he genome. Fu he -
mo e, due o GC-con en bias, GC-con en - ich egions in
he genome will ha e inc eased numbe o eads. These
ambigui ies make accu a e es ima ion o co e age and
consequen ly copy numbe a challenge [7]. Addi ionally,
umo ploidy and no mal cell con amina ion in oduce
u he challenges in SCNA de ec ion [8].
HTS-based copy numbe analysis is, in mos cases,
based on ead dep h (RD) es ima ions a each gen-
omic loca ion and u he segmen a ion and quan i i-
ca ion o he RD p o iles in o segmen s o consis en
copy numbe (Addi ional ile 1: Table S1 o a lis o
SCNA ools) [9, 10]. Howe e , such ools a e only
capable o de ec ing dele ions and duplica ions. Recen ly,
RD-based analysis has been augmen ed o iden i y cnLOH
e en s by inco po a ing in o ma ion om an al e na e
allele’s ac ion a he e ozygous single nucleo ide poly-
mo phism (SNP) posi ions (o B-allele ac ion (BAF)). The
BAF o a he e ozygous SNP has an expec ed alue o 0.5 in
no mal diploid cells. De ia ion om 0.5 in he he e ozy-
gous SNP BAF poin s o an abe a ion. In he case o
cnLOH, BAF alues a e expec ed o be ei he 0 o 1 in a
pu e umo popula ion. Tools such as Con ol-FREEC
[11], Pa chwo k [12], and CLImAT [13] inco po a e BAF
da a o ex end SCNA de ec ion. Con ol-FREEC de e -
mines he b eakpoin s using a leas absolu e sh inkage
es ima o (LASSO) eg ession. Sample ploidy is p o ided
by he use o Con ol-FREEC. I also e alua es and co -
ec s o no mal cell con amina ion, GC-con en , and map-
abili y biases while in e ing he copy numbe p o ile o a
umo genome. Pa chwo k pe o ms GC and posi ional
no maliza ion and segmen s he genome using a ci cula
bina y segmen a ion (CBS) algo i hm. I also es ima es no -
mal cell con amina ion and umo ploidy. CLImAT imple-
men s co ec ions o GC-con en and mapabili y bias and
models he RD and BAF da a wi h a hidden Ma ko model
(HMM) o in e he soma ic copy numbe a ia ion, no -
mal cell con amina ion and umo ploidy (Addi ional ile 1:
O e iew o Tools sec ion o mo e de ails on hese ools).
While he abo e ools a e well-sui ed o SCNA de ec ion,
hei use has some limi a ions. Con ol-FREEC and Pa ch-
wo k u ilize compu a ionally cos ly models, which leads o
long analysis imes. The main mo i a ion o ou s udy was
o de elop an accu a e and use - iendly ool ha could be
used o analyze la ge WGS da ase s, such as he cance
genome a las (TCGA) da ase s. In ou app oach, he RD
and BAF signals a e smoo hed, and double sliding windows
subsequen ly a e used o de ec b eakpoin s, which makes
ou app oach as and s aigh o wa d. Because he b eak-
poin de ec ion is pe o med simul aneously a di e en
scales, i allows accu a e de ec ion. Ou ool, Segmen um,
is eely a ailable unde MIT license a : h ps://gi hub.-
com/ea younian/Segmen um (Addi ional ile 2 con ains
he so wa e code. Fo he la es e sion o he so wa e
code please isi he p ojec ’s online eposi o y).
Implemen a ion
Pipeline
Segmen um was de eloped and w i en in he Py hon
p og amming language ( e sion 3) and equi es he
SciPy lib a y o be ins alled (I he use wishes o use he
‘plo ’sub-command o in o m pa ame e alue selec ion,
ma plo lib lib a y is also equi ed). Segmen um employs
SAM ools o ex ac RD and he e ozygous SNPs BAF
da a om BAM iles con aining WGS da a. These con-
s i u e he inpu s equi ed by Segmen um o pe o m
copy numbe analysis. Figu e 1 illus a es he
Fig. 1 Segmen um pipeline. No mal and umo RDs a e used o
calcula e RD log- a ios. RD log- a ios a e hen co ec ed o biases.
BAF da a a e simul aneously mi o ed and smoo hed. Using RD
log- a ios and BAF, he genome is segmen ed wi h a double sliding
window me hod. Segmen a ion esul s a e used o iden i y cnLOH
egions in he genome (see he ollowing sec ions o mo e de ails
on each s ep)
A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 2 o 10
Segmen um pipeline. Each s ep is explained in mo e de-
ail in he ollowing sec ions.
RD ex ac ion and BAF calcula ion o he e ozygous SNPs
To ex ac he RD om he BAM iles, he genome is
di ided in o bins o use -de ined leng h (2 kbp by
de aul ) and he numbe o eads o e lapping each bin is
coun ed o de e mine he RD a each bin. To calcula e
he BAF alues, he e ozygous SNPs in he no mal sam-
ple a e iden i ied a known SNP si es in he human gen-
ome (based on SNP anno a ions such as hose p oduced
by he 1000 Genomes p ojec ). Nex , he numbe o e -
e ence and al e na i e alleles a each he e ozygous SNP
posi ion is ex ac ed om he umo sample and he
BAF o he i
h
he e ozygous SNP is calcula ed using he
ollowing equa ion:
BAFi¼al i
al iþ e i
whe e al
i
and e
i
e e o he al e na i e and e e ence
allele, espec i ely, o he i
h
he e ozygous SNP.
I should be no ed ha by de aul eads wi h mapping
quali y sco e 10 a e il e ed ou be o e RD ex ac ion
and BAF calcula ion in o de o add ess he challenges
aised by eads no mapping o a unique egion in he
genome ( he ead il a ion c i e ion based on he map-
ping quali y sco e is a pa ame e o Segmen um and can
be se by he use ).
Log- a io calcula ion
The RD log- a io is calcula ed using he ollowing
equa ion:
log i¼log2
RDi
nRDi

whe e log
i
is he log- a io o he i
h
genomic window
and RD
i
and nRD
i
a e RDs ex ac ed om he i
h
gen-
omic window o a speci ic size (de e mined by use ;
de aul is 2 kbp) o he umo and no mal samples,
espec i ely.
Di e ences in he o al numbe o aligned eads in he
no mal and umo samples may bias he es ima ion o
he RD log- a ios. The co ec ion was pe o med by
inding he mode o log- a io alues o each ch omo-
some and sub ac ing he median o all o he modes
om each log- a io alue. I should be no ed ha
median, in he co ec ion s ep, is obus o he changes
in one mode. Fo ins ance, one ch omosomal a m
ha ing a copy numbe change has no e ec on he
co ec ion since i only a ec s one o he ch omosomal
modes.
Mi o ing and smoo hing o he BAF alues
The BAF o a he e ozygous SNP has an expec ed alue
o 0.5 in no mal diploid cells. In he p esence o soma ic
copy numbe al e a ions, he BAF can di e ge om 0.5
i he ela i e abundance o he wo alleles changes. To
make smoo hing and segmen a ion o BAF da a possible,
he BAF alues mus be mi o ed abou he 0.5 axis so
ha he B allele ac ion always ep esen s he allele
ac ion o he dominan allele. Wi hou his mi o ing
s ep, he BAF alues will be symme ic abou he BAF =
0.5 axis and smoo hing will unde es ima e he absolu e
di e gence om 0.5 [14]. In his s udy, a median il e is
used o smoo hing he BAF da a. Simul aneous mi o -
ing and smoo hing is implemen ed using he ollowing
equa ion:
cBAFi¼H0:5−M9BAFi
ðÞ
jj
þ1−HðÞM90:5−BAFi
jj
ðÞ:
whe e BAF
i
is he BAF alue o he i
h
he e ozygous SNP,
cBAF
i
is he simul aneously mi o ed and smoo hed BAF
i
,
His a he e ozygosi y measu emen calcula ed wi h he
ollowing equa ion: H=1−2∗|0.5 −x|, and M
9
e e s o
applying a median il e o 9 SNPs in he icini y o and
including he i
h
SNP.
Segmen a ion using a double sliding window app oach
To de ec changes in he RD log- a io and BAF signals,
wo non-o e lapping, ixed-sized windows (de e mined
by he use ) a e slid o e he RD log- a io and BAF
alues and a compound sco e (S) is calcula ed o each
o he adjacen wo windows. I he compound sco e is
g ea e han 1, a change is de ec ed and a b eakpoin is
placed a he place whe e he wo windows ouch each
o he . The compound sco e is calcula ed using he
ollowing equa ion:
S¼log wini−log winiþ1
2
τlog þcBAFwini−cBAFwiniþ1
2
τBAF
whe e log winiis he mean o he RD log- a io alues in
he i
h
window, cBAFwiniis he mean o he mi o ed
and smoo hed BAF alues in he i
h
window, τ
log
and
τ
BAF
a e h esholds o he absolu e mean di e ence in
he RD log- a ios and he absolu e mean di e ence in
he BAF alues in he wo adjacen windows,
espec i ely.
I is possible ha some b eakpoin s will no be de-
ec ed by a single pass o a double sliding window due
o a gi en window size. Thus, o inc ease he sensi i i y,
Segmen um analyzes he signals o he de ec ion o
b eakpoin s mul iple imes wi h di e en window-sizes
and h esholds. Each new window is 1.5 imes la ge
han he p e ious one. The inc ease in he window size
dec eases he de ec ion h esholds. This is due o he
ac ha inc easing he window size inc eases he sample
A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 3 o 10
size (assuming sampling om no mal dis ibu ion wi h
N(μ,σ
2
)) and consequen ly dec eases he s anda d de i-
a ion o he mean (mean ha ing p obabili y dis ibu ion
o Nμ;σ2.n

). The new s anda d de ia ion o he
mean when window size is inc eased 1.5 imes is 1
ffiffiffiffiffi
1:5
2
p
imes he old s anda d de ia ion. Le τ=ασ whe e τis
he h eshold and αis a scala and σis he s anda d de-
ia ion. I ollows ha :
τnew ¼α:σnew ¼1
ffiffiffiffiffiffiffi
1:5
2
p:α:σold ¼1
ffiffiffiffiffiffiffi
1:5
2
p:τold
Thus bo h he τ
log
and τ
BAF
h esholds a e upda ed
using he ollowing equa ion:
τnew ¼1
ffiffiffiffiffiffiffi
1:5
2
pτold
The p ocess o inc easing he window-size is con in-
ued as long as he upda ed h esholds a e g ea e han
he h esholds o he me ging wo consecu i e seg-
men s (see below). A e de ec ing all he b eakpoin s, a
consensus lis o b eakpoin s is c ea ed by accep ing all
o he b eakpoin s de ec ed by he i s pass o he
double sliding window and adding he b eakpoin s de-
ec ed om he la ge windows o he lis only i he
b eakpoin is no in he icini y o an exis ing b eakpoin
in he lis (i.e., |cp
cu en
−cp
exis ing
|>window size, whe e
cp is a de ec ed b eakpoin ). Consensus b eakpoin s a e
used o c ea e he segmen s. Two consecu i e b eak-
poin s cons i u e a segmen . Fo each segmen , he a e -
age RD log- a io and a e age mi o ed and smoo hed
BAF is calcula ed. Two consecu i e segmen s a e
me ged i he ollowing condi ions a e me :
log segi
−log segiþ1
<τme gelog
and jcBAFsegi
−cBAFsegiþ1j<τme geBAF
whe e log segiis he mean RD log- a io o he i
h
seg-
men , cBAFsegiis he mean mi o ed and smoo hed BAF
o he i
h
segmen , and τme gelog and τme geBAF (de e mined
by use ) a e he RD log- a io and BAF me ging h esh-
olds, espec i ely.
De ec ion o cnLOH e en s wi hin a single sample
A segmen is conside ed o be a cnLOH segmen i he
ollowing condi ions a e me :
log segi
<τcnLOHlog and 0:5−cBAFsegi

<τcnLOHBAF
whe e log segiis he mean RD log- a io o he i
h
seg-
men , cBAFsegiis he mean mi o ed and smoo hed BAF
o he i
h
segmen , τcnLOHlog and τcnLOHBAF (de e mined
by he use ) a e h esholds o calling a cnLOH segmen .
De ec ion o ecu en cnLOH egions ac oss mul iple
samples
To ind genomic egions wi h ecu en cnLOH
e en s, all cnLOH egions o indi idual samples a e
iden i ied ollowing he p ocedu e desc ibed ea lie .
Then, he numbe o occu ences o a cnLOH e en
o a speci ic egion ac oss mul iple samples is
coun ed using an in e al ee da a s uc u e
(Addi ional ile 1: Figu e S1).
Simula o
To e alua e Segmen um in e ms o segmen a ion ac-
cu acy, a simula o capable o simula ing whole-
genome RD o bo h no mal and umo samples and
BAF based on e en s such as dele ions, ampli ica ions
and cnLOH was de eloped. The simula o ecei es a
no mal sample RD da a and ou pu s 4 se s o da a in-
cluding he simula ed no mal and umo RD, BAF
da a and a g ound u h. Fi s , he simula o lea ns
he dis ibu ion o he RD da a om he p o ided
no mal sample by simply coun ing he numbe o
imes wo consecu i e RD alues (e.g., 368 and 299)
occu oge he h oughou he genome (Addi ional
ile 1: Figu e S2). The lea ned dis ibu ion also ac-
coun s o he inhe en noise in he RD da a. Nex ,
in e se ans o m sampling (Smi no ans o m) is
used o gene a e RD alues o each posi ion in he
genome based on he lea ned dis ibu ion. Then,
noise is emo ed using a median il e . A no mal RD
is cons uc ed by adding independen Poisson noise
o he simula ed RD da a. To cons uc he umo
RD, wo copy numbe acks (because au osomal
ch omosomes come in ma e nal and pa e nal pai s)
ha bo ing andom SCNAs a e cons uc ed. The umo
sample RD is calcula ed using he copy numbe
acks, he simula ed no mal sample RD and he no -
mal sample con amina ion (i.e., a pa ame e de e -
mined by use ). To cons uc he BAF da a,
he e ozygous SNPs a e ini ially andomly dis ibu ed
ac oss he genome (1 he e ozygous SNP pe 1.5 Kbp).
The numbe o B-alleles a a he e ozygous SNP is cal-
cula ed using a binomial dis ibu ion wi h he pa am-
e e s n( o al numbe o eads a he e ozygous SNP
posi ion) and p(p obabili y ha a ead is coming
om he B-allele). nis ex ac ed om he simula ed
no mal RD a he e ozygous SNP posi ions. pis calcu-
la ed using he wo cons uc ed copy numbe acks
and he no mal sample con amina ion. Once he
numbe o B-alleles is calcula ed, i is used o calcu-
la e he BAF alues (Addi ional ile 1: Figu es S3 and
S4 o he simula o pipeline and he simula ed da a
isualized in he in eg a i e genomics iewe (IGV)
[15], espec i ely).
A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 4 o 10
Resul s
Segmen um segmen a ion accu acy o he simula ed
da a
Using he simula o (see he ‘Simula o ’sec ion o mo e
de ails), RD da a o bo h no mal and umo samples
and BAF alues o he e ozygous SNPs om he umo
sample as well as a g ound u h we e simula ed wi h
di e en pe cen ages o no mal con amina ion (an ex-
ample se o simula ed da a is a ailable a Segmen um’s
online eposi o y. See he ‘A ailabili y and equi emen s’
sec ion o he link o he eposi o y). The simula ed
da a we e analyzed by Segmen um. The segmen a ion
esul s we e e alua ed agains he g ound u h. The
p ecision, ecall, and he F-measu e alues we e calcu-
la ed based on his e alua ion (Fig. 2 and Addi ional ile
1 o he de ini ions o p ecision, ecall, and F-measu e).
Segmen um segmen a ion accu acy o eal da a
compa ed o o he ools
To assess segmen a ion accu acy o Segmen um o eal
da a, pai ed umo /no mal whole genome sequencing
samples (30x < co e age < 100x) om 10 indi iduals
diagnosed wi h low-g ade glioma (LGG) we e down-
loaded om he TCGA da ase and used as is. Fu he -
mo e, segmen a ion esul s om SNP-a ay da a (le el 3
da a) (comple ed by TCGA using an A yme ix
Genome-wide human SNP a ay 6.0) was used as
g ound u h (Addi ional ile 1: Table S3). Segmen um’s
esul s we e e alua ed agains Con ol-FREEC, Pa ch-
wo k, and CLImAT as compe ing ools. To e alua e he
segmen a ion accu acy, he genome was b oken in o
100 bp. blocks (excluding all blocks in cen ome es and
sex ch omosomes). Using block anno a ions om
di e en ools, genome-wide p opo ions o he blocks
anno a ed as SCNA by di e en combina ions o ools
we e calcula ed and he esul s we e illus a ed by a
Venn diag am (Fig. 3).
Addi ionally, o measu e he pai wise deg ee o simi-
la i y o he segmen a ion esul s be ween wo ools, he
Jacca d simila i y index (JSI) was calcula ed o all o
he pai s using he ollowing equa ion:
JSI ¼∩pai
jj
∪pai jj
whe e | ∩pai | and | ∪pai | a e he ca dinali ies o in e -
sec ion and union, espec i ely. In e sec ion and union
alues we e ex ac ed om he Venn diag ams. Figu e 4
ep esen s a hea map o he JSI alues o each pai o
ools a e aged o e 10 TCGA LGG samples. Acco ding
o he hea map, on a e age, Segmen um p oduces he
mos simila esul s o he SNP a ay segmen a ion
esul s wi h a JSI sco e o 0.9, ollowed by Pa chwo k
wi h a JSI sco e o 0.86.
Simila e alua ions using low co e age da a (6x a e -
age co e age) a e shown in Addi ional ile 1: Figu es S5
and S6. The low co e age da a is comp ised o he
pai ed umo /no mal whole genome sequencing samples
o 10 indi iduals diagnosed wi h p os a e adenoca cin-
oma (PRAD). Wi h ega d o he low co e age da a,
Pa chwo k p oduces he mos simila esul s o he SNP
a ay segmen a ion esul s wi h a JSI sco e o 0.93,
ollowed by Segmen um wi h a JSI sco e o 0.88. Add-
i ional ile 1: Table S4 con ains he names o he 10
TCGA PRAD samples (Addi ional ile 1: Tables S6-S10
ep esen he pa ame e alues used o unning he
Fig. 2 Segmen a ion accu acy o Segmen um o simula ed da a wi h di e en deg ees o no mal con amina ion. Es ima ed p ecision, ecall, and
F-measu e alues o simula ed da a a di e en no mal con amina ion le els (Addi ional ile 1, De i a ion o he p ecision, ecall, and F-measu e
o he simula ed da a)
A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 5 o 10

compe ing ools. Addi ional ile 1: Segmen um’s pa am-
e e alue selec ion sec ion p o ides guidance on selec -
ing pa ame e alues o Segmen um. Addi ional ile 1:
Figu e S9 ep esen s an example plo made by Segmen-
um’s 'plo ' sub-command ha can be used o guide he
pa ame e alue selec ion).
Segmen um segmen a ion accu acy o he subsampled
eal da a
To assess he segmen a ion accu acy o Segmen um o
eal da a wi h espec o sample’s co e age, we
subsampled one o he LGG samples (i.e. TCGA-CS-
5395) a di e en subsampling ac ions (i.e. 75%, 50%,
25%, 10%, and 5%) using Sam ools ( e sion 1.3.1). We
analyzed each subsample by Segmen um and bench-
ma ked i agains g ound u h in he same manne as
explained ea lie . Figu e 5 ep esen s he JSI sco es o
each subsample (Addi ional ile 1: Figu e S7 shows he
a e age co e age o he subsamples o no mal and
umo pai s). I can be seen ha Segmen um eaches
high accu acies e en wi h low co e age da a. Fo in-
s ance, he accu acy o he 10%- ac ion subsample was
93.4% (whe e he a e age co e age o umo and no -
mal subsamples we e 3 and 4 espec i ely).
I should be no ed ha as he co e age dec eases he
numbe o iden i ied he e ozygous SNPs dec eases
(Addi ional ile 1: Figu e S8). Fo ins ance, o he 10%-
ac ion subsample only 1997 he e ozygous SNPs we e
iden i ied om he en i e genome (in con as o he
o iginal sample whe e he numbe o iden i ied he e ozy-
gous SNPs was mo e han 3 million SNPs). E en hough
Segmen um is shown o wo k wi h low co e age da a,
one should no e he implica ions o low amoun s o
de ec ed he e ozygous SNPs on he eliable de ec ion o
cnLOH e en s.
Time usage e alua ion
All o he compu a ions we e comple ed on he same
UNIX se e . Table 1 shows he a e age ime equi ed
by each ool o pe o m he analysis o 10 TCGA LGG
samples (30x < co e age < 100x). Based on he esul s, on
a e age, CLImAT appea s o be he as es , ollowed by
Segmen um, Pa chwo k, and Con ol-FREEC. I should
be no ed ha o assign he allele-speci ic copy numbe
o genomic segmen s, Pa chwo k equi es use s o de e -
mine some pa ame e alues by in e p e ing plo s p o-
duced by he ool, and his in e p e a ion ime is no
included he e. Addi ionally, he ime equi ed o c ea e
he pileup iles used by Pa chwo k and Con ol-FREEC
is di e en due o he use o di e en pa ame e alues
in SAM ools. I should be no ed ha ime equi ed o
making pileup iles can be dec eased by pa allelizing he
p ocess on machines wi h mul iple co es o on com-
pu e clus e s (e.g. by assigning one co e o each
ch omosome). Simila ly, BAF calcula ion o Segmen um
can be pa allelized. Howe e , since his is no a co e ea-
u e o he benchma ked ools and no all ools suppo
pa alleliza ion, o be ai , only he equi ed linea ime is
epo ed he e. A simila ime usage e alua ion, using
low co e age da a (a e age co e age 6x), is shown in
Addi ional ile 1: Table S2. Wi h ega d o he low co e -
age da a, Segmen um comes second a e CLImAT in
e ms o analysis ime, which is consis en wi h he
esul s om he high co e age da a.
Fig. 3 Compa ison o he SCNA esul s wi h di e en ools and he
SNP a ay (g ound u h). Venn diag am alues (a e aged o e en
TCGA LGG samples) ep esen he pe cen age o o e lap among he
SCNA calls
Fig. 4 Pai wise JSI sco es a e aged o e en TCGA LGG samples.
JSI sco es ange be ween 0 and 1, whe e 0 means no simila i y and
1 ep esen s iden ical esul s be ween wo ools
A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 6 o 10
Recu en cnLOH de ec ion case s udy
In a s udy o lowe g ade gliomas (LGGs), i.e., g ade II
and III gliomas, Suzuki e al. [4] cha ac e ized he mu a-
ional landscape o hese glioma ypes by di iding hem
in o 3 dis inc sub ypes based on hei dis inc se s o
mu a ions and clinical beha io s. These sub ypes a e dis-
inguished wi h he ollowing c i e ia: (1) mu a ion in
IDH1/2 accompanied by co-dele ion o ch omosomes 1p
and 19q (sub ype I), (2) mu a ion in IDH1/2 wi hou co-
dele ion o ch omosomes 1p and 19q (sub ype II), and
(3) IDH1/2 wild ype (sub ype III). O in e es o ou
s udy was he ecu ence o cnLOH e en s in ch omo-
some 17p in sub ype II [4]. To show he abili y o
Segmen um o de ec such abe a ions om la ge da a-
se s, 38 pai ed-end WGS samples om he TCGA da a-
se (30x < co e age < 100x)) o pa ien s diagnosed wi h
LGG we e downloaded and analyzed by Segmen um.
We we e able o dis inguish all h ee sub ypes as cha ac-
e ized in [4], including he ecu ence o cnLOH in
sub ype II a ch omosome 17p. We also iden i ied a
ou h sub ype wi h a mu a ion in IDH1/2 wi hou co-
dele ion o ch omosomes 1p and 19q and no cnLOH a
17p (Fig. 6).
Discussion
By compa ing he simula ed (Fig. 2) and eal da a (Figs. 3,
4 and 5 and Addi ional ile 1: Figu es S5 and S6), we can
conclude ha Segmen um can eco e ue copy
numbe abe a ions wi h high accu acy e en when he
co e age is as low as ~4 eads (Fig. 5, Addi ional ile 1:
Figu e S7). On a e age, Segmen um p oduces esul s
ha a e he mos conco dan wi h he copy numbe ab-
e a ions iden i ied om he SNP a ay da a (i.e. ~90%
o conco dance) (Fig. 4). As shown in Table 1, ou ool
is mo e han wice as as as he second bes pe o ming
ool in e ms o accu acy. Segmen um is also he second
as es ool a e CLImAT compa ed o he o he ools
e alua ed in his s udy (Table 1). Howe e , CLImAT
anks las in e ms o accu acy (Fig. 4). One explana ion
o he speed o CLImAT is ha i compu es he BAF
alues o a subse o known SNPs (~13.7 million SNPs
ha a e e ie ed om he dbSNP da abase [16]). In
con as , Segmen um, compu es he BAF alues o he -
e ozygous SNPs de e mined om he 1000 Genomes
p ojec ’s SNP lis (~85 million SNPs) [17]. The o he
eason o he speed o CLImAT migh be ha i does
no equi e a no mal sample o analysis.
As he no mal con amina ion in he simula ed da a in-
c eases, he numbe o alse nega i es inc eases and he
ecall a e dec eases (Fig. 2). Howe e , wi hin he anges
o ealis ic amoun s o no mal con amina ion (i.e. ~30%
o 40%), Segmen um pe o ms consis en ly well.
Segmen um is able o epo ecu en cnLOH egions
ac oss mul iple cance genome samples; a cha ac e is ic
Fig. 5 Pai wise JSI sco es (Segmen um s. SNP a ay as g ound u h) o di e en subsamples. JSI sco es ange be ween 0 and 1, whe e 0 means
no simila i y and 1 ep esen s iden ical esul s be ween wo ools
Table 1 A e age ool analysis ime o high co e age da a
(30x < co e age < 100x)
Tool A e age p epa a ion ime A e age analysis
ime
Segmen um - 10 h 34 min o ex ac ing RD
om no mal o umo BAM ile
- 4 h 25 min o calcula ing BAF
alues
- 1 min 45 s
Pa chwo k - 29 h 37 min o c ea ing pileups
om no mal o umo BAM ile
- 3 h 56 min
Con ol-FREEC - 33 h 28 min o c ea ing pileups
om no mal o umo BAM ile
- 7 h 11 min
CLImAT - 2 h 12 min o ex ac ing RD - 29 min
A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 7 o 10
Fig. 6 SCNA landscape in g ade II and III gliomas. WHO-g ade, his ological class, and molecula sub ype classi ica ion a e shown by colo
as indica ed. The hi y-eigh samples a e di ided in o 4 dis inc sub ypes based on he occu ence o a mu a ion in IDH1/2, co-dele ion
o ch omosomes 1p and 19q and he p esence o 17p cnLOH. Dele ions and ampli ica ions a e isualized by boxes wi h di e en shades
o blue and ed, espec i ely. Whi e egions a e ei he no mal o cnLOH egions. The ba cha s below each box ep esen he mi o ed
and smoo hed BAF alues. La ge mi o ed and smoo hed BAF alues (close o 0.5) poin o he e ozygous SNP allelic imbalance. In he second sub ype
( om he op), a ch omosome 17p, ecu ing cnLOH is appa en whe e he ba cha s poin o la ge mi o ed and smoo hed BAF alues, hough no
dele ion o ampli ica ion is de ec ed a ha egion (Addi ional ile 1: Table S5 o TCGA LGG sample ba code names)
A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 8 o 10
o cance genomes ha has been neglec ed un il ecen ly
[4]. By applying Segmen um o TCGA da a, we we e
able o eco e ecu en cnLOH e en s om low-g ade
glioma samples ha we e epo ed ea lie by SNP a ay-
based da a analysis. I is wo h men ioning ha Segmen-
um can wo k in wo modes, i.e., wi h o wi hou BAF
alue. In he case whe e BAF alues a e no used, Seg-
men um canno de ec egions wi h cnLOH. Fu he -
mo e, Segmen um is capable o eliably segmen ing he
cance genome using bo h high (Figs. 3 and 4) and low
(Fig. 5 and Addi ional ile 1: Figu es S5 and S6) sequence
co e age da a. Howe e , wi h he low sequence co e age
da a, he es ima ed BAF alues o he he e ozygous
SNPs will be less eliable. This is also e lec ed in
Addi ional ile 1: Figu e S8, whe e i is shown ha he
numbe o de ec ed he e ozygous SNPs d op as he a e -
age co e age dec eases. The implica ions o low
amoun s o de ec ed he e ozygous SNPs on he eliable
de ec ion o cnLOH e en s should no be o e looked.
E en hough we ha e shown ha Segmen um is
highly accu a e a eco e ing he ue copy numbe ,
o he ools in his s udy do mo e han jus segmen -
ing he genome. Fo ins ance, CLImAT and Pa ch-
wo k a e capable o es ima ing umo ploidy and
umo pu i y and consequen ly, epo ing he in eg al
copy numbe s o each segmen . Pa chwo k and
Con ol-FREEC a e also capable o epo ing he
geno ype o each segmen and CLImAT epo s he
geno ype o each SNP wi hin each segmen . This is
in con as o Segmen um ha only epo s he mean
RD log- a io and BAF alue o each segmen . How-
e e , ools such as ABSOLUTE [18] o THe A [19]
can be used o es ima e umo impu i y and ploidy
om Segmen um’s segmen a ion esul , meaning ha
Segmen um can be used as pa o a la ge umo
e olu ion analysis pipeline. Finally, a s eng h o ou
ool is i s minimum dependence on hi d pa y ools,
wi h he excep ion o SAM ools, o calcula ing he
RD and BAF.
Conclusions
We ha e de eloped Segmen um as a ool o he
iden i ica ion o SCNAs, including cnLOH in umo
samples, using WGS da a. We ha e shown ha Seg-
men um is accu a e and as wi h ega ds o o he
s a e-o - he-a ools, making i sui able o analyzing
coho s wi h a la ge numbe o samples, such as
TCGA coho s.
A ailabili y and equi emen s
P ojec name: Segmen um
P ojec homepage: h ps://gi hub.com/ea younian/Segm
en um
Ope a ing sys em(s): Linux
P og amming language: Py hon
O he equi emen s: SciPy, Sam ools, and ma plo lib i
he ‘plo ’sub-command is used.
License: MIT license
Any es ic ions o use by non-academics: None
Addi ional iles
Addi ional ile 1: This ile con ains supplemen a y in o ma ion, ables
and igu es suppo ing he manusc ip . Figu e S1. De ec ion o egions
ha bo ing ecu en cnLOH ac oss mul iple samples. Figu e S2. Read
dep h spa ial co ela ion. Figu e S3. Simula o pipeline. Figu e S4.
Simula ed da a isualized in In eg a i e Genomics Viewe (IGV). Figu e S5.
Compa ison o SCNA esul s om di e en ools and SNP a ay (g ound
u h) o low sequence co e age da a. Figu e S6. Pai wise JSI
sco es o low sequence co e age da a (a e aged o 10 TCGA PRAD
samples) Figu e S7. Subsample a e age co e ages in he subsampling
e alua ion. Figu e S8 De ec ed numbe o he e ozygous SNPs in di e en
subsamples Figu e S9. Copy numbe –B-allele ac ion clus e s. Table S1.
Lis o SCNA ools using WGS da a. Table S2. A e age ool analysis ime
o low sequence co e age da a (a e age co e age 6x). Table S3 TCGA
LGG sample ba code names and he es ima ed sample pu i y by ABSOLUTE.
Table S4. TCGA PRAD sample ba code names and he es ima ed
sample pu i y by ABSOLUTE. Table S5. TCGA LGG sample ba code
names ca ego ized based on in e ed sub ype. Table S6. Pa ame e
alues o unning DFEx ac . Table S7. Pa ame e alues o unning
CLImAT. Table S8. Pa ame e alues o unning Pa chwo k o 10
TCGA LGG samples. Table S9. Pa ame e alues o unning Pa chwo k o
10 TCGA PRAD samples. Table S10. Pa ame e alues o unning Con ol-
FREEC. (DOCX 857 kb)
Addi ional ile 2: So wa e code. This comp essed ile con ains he
so wa e code ( o he la es e sion o he so wa e code please isi
he p ojec ’s online eposi o y). (ZIP 32 kb)
Abb e ia ions
BAF: B-Allele ac ion; BAM: Bina y alignmen map; CBS: ci cula bina y
segmen a ion; CGH: A ay compa a i e genomic hyb idiza ion; cnLOH:
Copy-neu al loss o he e ozygosi y; CNV: Copy numbe a ia ion;
dbSNP: Single nucleo ide polymo phism da abase; DNA: Deoxy ibonucleic
acid; FISH: Fluo escen in si u hyb idiza ion; HMM: Hidden Ma ko model;
HTS: High h oughpu sequencing; IGV: In eg a i e genomics iewe ;
JSI: Jacca d simila i y index; LASSO: Leas absolu e sh inkage eS ima O ;
LGG: Low g ade glioma; LOH: Loss o he e ozygosi y; MAE: Monoallelic
exp ession; PRAD: PRos a e ADenoca cinoma; RD: Read-dep h;
SAM: Sequence alignmen /map; SCNA: Soma ic copy numbe
al e a ion; SNP: Single nucleo ide polymo phism; TCGA: The cance
genome a las; WES: Whole exome sequencing; WGS: Whole genome
sequencing
Acknowledgmen s
The esul s published he e a e in pa based upon da a gene a ed by The
Cance Genome A las p ojec (dbGaP S udy Accession: phs000178. 9.p8)
es ablished by he NCI and NHGRI. In o ma ion abou TCGA and he
in es iga o s and ins i u ions who cons i u e he TCGA esea ch ne wo k
can be ound a h p://cance genome.nih.go . The au ho s would also like
o acknowledge CSC —IT Cen e o Science L d. (h ps://www.csc. i/csc)
o p o iding he compu a ional esou ces.
Funding
This wo k was unded by he Academy o Finland (p ojec no. 269474),
Cance Socie y o Finland, and Sig id Juselius Founda ion. The unde s had
no ole in s udy design, da a collec ion and analysis, decision o publish, o
p epa a ion o he manusc ip .
A younian e al. BMC Bioin o ma ics (2017) 18:215 Page 9 o 10