75CanCe In o ma ICs 2015:14(s5)
In oduc ion
Flow cy ome y enables quan i a i e measu emen o single-
cell p ope ies h ough isible and luo escen ligh in a high-
h oughpu manne . The measu ed signals include luo escence
emission and ligh sca e . Flow cy ome y has been ou inely
used o de ec ing malignancies om blood samples.1 Th ough
echnological ad ances, measu ing he combina ion o luo-
escen signals om se e al di e en channels has enabled he
use o high-dimensional da a o s udies, such as cy ome ic
inge p in ing2 and la ge-scale analysis o cell ypes.3
In his pape , we s udy he analysis o low cy ome y
da a om he ea u e selec ion poin o iew. Mo e speci i-
cally, low cy ome y is able o p oduce la ge quan i ies o
pa ially edundan measu emen da a, and he selec ion o
impo an quan i ies wi hin he la ge body o measu emen s is
o in e es . Mo eo e , a ypical scena io con ains la ge quan-
i ies o measu emen da a bu may be limi ed o only a ew
pa ien s. Thus, an ideal me hod would dis il only he essen ial
pa s o he measu emen s om each pa ien , while p oduc-
ing eliable and well gene alizing esul s when only a small
amoun o indi iduals is a ailable in he aining da a.
We will concen a e ou a en ion on wo pa icula se s
o low cy ome y da a. The i s se o igina es om he acu e
myeloid leukemia (AML) p edic ion challenge o he DREAM
ini ia i e in 2013.4 The compe i ion a ac ed a numbe o
eams, and as se e al esea che s use he da a as pa o hei
wo k, he challenge da a ha e become a s anda d benchma k
wi hin he ield. Fo example, Aghaeepou e al.4 p esen s a
la ge pool o analysis app oaches om he DREAM challenge.
Among classi ica ion me hods p esen ed in he li e a u e,
he e a e se e al sophis ica ed machine lea ning app oaches,
such as lea ning ec o quan iza ion,5 co ela i e ma ix map-
ping, and ela i e en opy di e ences.6 The s eng h o da a-
d i en app oaches elying on supe ised classi ica ion is hei
abili y o handle high-dimensional da a wi hou equi ing
p io knowledge o he biological applica ion.
The DREAM AML da a ep esen a ela i ely la ge-scale
expe imen consis ing o al oge he 179 pa ien s. Al hough i
is a small numbe o adi ional machine lea ning p oblems,
he numbe o pa ien s is unusually la ge o a biological s udy.
To his aim, we use ano he da ase ha ep esen s a mo e
commonly encoun e ed sample size o 16 samples ex ac ed
om a p os a e cance cell line. This da ase , wi h wo di -
e en ea men s and a low numbe o samples, p esen s a
non i ial bu common challenge o p edic ion and ela ed
ea u e selec ion. Mo e in o ma ion on he wo da ase s is
p o ided in Da a sec ion.
In ou ea lie wo k,7 we p esen ed a supe ised classi ica-
ion pipeline based on linea disc iminan analysis (LDA) and
logis ic eg ession (LR) classi ie s. B ie ly, he me hod i s
Flow Cy ome y-Based Classi ica ion in Cance
Resea ch: A View on Fea u e Selec ion
s. saki a Hassan1, Pekka uusu uo i2,3, Leena La onen3 and Heikki Hu unen1
1Depa men o Signal P ocessing, Tampe e Uni e si y o Technology, Tampe e, Finland. 2Po i Depa men , Tampe e Uni e si y o
Technology, Po i, Finland. 3BioMediTech, Uni e si y o Tampe e, Tampe e, Finland.
Supplemen a y Issue: S a is ical Sys ems Theo y in Cance Modeling, Diagnosis, and The apy
Abs Ac : In his pape , we s udy he p oblem o ea u e selec ion in cance - ela ed machine lea ning asks. In pa icula , we s udy he accu acy and
s abili y o di e en ea u e selec ion app oaches wi hin simplis ic machine lea ning pipelines. Ea lie s udies ha e shown ha o ce ain cases, he accu acy
o de ec ion can easily each 100% gi en enough aining da a. He e, howe e , we concen a e on simpli ying he classi ica ion models wi h and seek o
ea u e selec ion app oaches ha a e eliable e en wi h ex emely small sample sizes. We show ha as much as 50% o ea u es can be disca ded wi hou
comp omising he p edic ion accu acy. Mo eo e , we s udy he model selec ion p oblem among he
1 egula iza ion pa h o logis ic eg ession classi ie s.
To his aim, we compa e a mo e adi ional c oss- alida ion app oach wi h a ecen ly p oposed Bayesian e o es ima o .
Keywo ds: AML, leukemia, low cy ome y, logis ic eg ession, e o es ima ion, model selec ion
SUPPLEMENT: s a is ical sys ems heo y in Cance modeling, Diagnosis, and he apy
CITATIoN: Hassan e al. Flow Cy ome y-Based Classi ica ion in Cance Resea ch: A View
on ea u e selec ion. Cance In o ma ics 2015:14(s5) 75–85 doi: 10.4137/CIn.s30795.
TYPE: o iginal esea ch
RECEIVED: no embe 18, 2015. RESUBMITTED: eb ua y 01, 2016. ACCEPTED FoR
PUBLICATIoN: eb ua y 07, 2016.
ACADEMIC EDIToR: J. T. E i d, Edi o in Chie
PEER REVIEw: six pee e iewe s con ibu ed o he pee e iew epo . e iewe s’
epo s o aled 1860 wo ds, excluding any con iden ial commen s o he academic edi o .
FUNDINg: Au ho s disclose no ex e nal unding sou ces.
CoMPETINg INTERESTS: Au ho s disclose no po en ial con lic s o in e es .
CoRRESPoNDENCE: [email p o ec ed]
CoPYRIghT: © he au ho s, publishe and licensee Libe as academica Limi ed. his is
an open-access a icle dis ibu ed unde he e ms o he C ea i e Commons CC-BY-NC
3.0 License.
Pape subjec o independen expe blind pee e iew. all edi o ial decisions made
by independen academic edi o . Upon submission manusc ip was subjec o an i-
plagia ism scanning. P io o publica ion all au ho s ha e gi en signed con i ma ion o
ag eemen o a icle publica ion and compliance wi h all applicable e hical and legal
equi emen s, including he accu acy o au ho and con ibu o in o ma ion, disclosu e o
compe ing in e es s and unding sou ces, compliance wi h e hical equi emen s ela ing
o human and animal s udy pa icipan s, and compliance wi h any copy igh equi emen s
o hi d pa ies. This jou nal is a membe o he Commi ee on Publica ion E hics (COPE).
Published by Libe as academica. Lea n mo e abou his jou nal.
Hassan e al
76 CanCe In o ma ICs 2015:14(s5)
ans o ms he measu emen da a in o highe dimensional
space by gene a ing combined ea u es wi h mul iplica ions and
di isions be ween measu emen s. Following his mapping in o
highe dimensional space, LDA is used o lowe ing he dimen-
sionali y in o a single alue pe measu emen . Then, empi ical
dis ibu ion unc ions (EDFs) a e cons uc ed om LDA esul s
o AML-posi i e and AML-nega i e sample classes and com-
pa ed o aining EDFs o bo h classes. The compa ison esul s
in wo simila i y alues pe g oup o measu emen s, and hese
esul s a e ed o he LR classi ie o a inal AML p edic ion
esul . Ou app oach, oge he wi h al e na i e well-pe o ming
app oaches,4,5 ep esen s a ela i ely complica ed pipeline o
somewha a bi a y compu a ion s eps. Thus, ou in e es is o
simpli y hese pipelines in o a simple collec ion o ob ious ea-
u es, while s ill e aining a good accu acy.
Ou app oach he e is o use LR classi ie applied o sum-
ma ize ea u es, which a e he mean and s anda d de ia ion o
he measu emen s ins ead o he comple e da a. This educes
he numbe o ea u es used in classi ica ion and, subsequen ly,
also he model complexi y. An essen ial pa o classi ie
design is e o es ima ion, which guides model selec ion.8 Ou
s a egy o model selec ion is o apply he ecen ly in oduced
Bayesian e o es ima o (BEE).9 We compa e BEE model
selec ion wi h a adi ional 10- old c oss- alida ion (CV-10)
e o es ima ion, as well as wi h Bayesian in o ma ion c i e-
ion (BIC)-based model selec ion, and conclude ha he p o-
posed app oach enables accu a e p edic ion o low cy ome y
da a wi h ewe measu emen s and a less complex classi ie
model han hose p e iously p esen ed in he li e a u e.
The es o his pape is o ganized as ollows. In Ma e ials
and Me hods sec ion, we desc ibe he da a and me hods used
in his s udy and b ie ly discuss how ea u e selec ion is com-
monly done in machine lea ning. Expe imen al Resul s sec ion
p esen s he esul s o ou expe imen s wi h di e en model
and ea u e selec ion c i e ia o he ma e ials in oduced in
Ma e ials and Me hods sec ion. Finally, in Conclusions sec-
ion, we summa ize he wo k and discuss he conclusions o
he esul s.
Ma e ials and Me hods
In his sec ion, we desc ibe he da ase s used in his pape .
We also gi e a b ie o e iew o modeling me hod o ea u e
selec ion. In addi ion o his, we in oduce he s a e-o - he-a
Bayesian e o es ima o (BEE) o model pa ame e selec-
ion. Finally, we p esen an example whe e he pe o mance o
BEE is benchma ked agains o he model selec ion c i e ia.
da a. In his wo k, we s udy wo da ase s: A la ge se
wi h 179 samples and a smalle se wi h 16 samples. These
wo case s udies ep esen di e en classi ica ion challenges in
e ms o bo h applica ion and sample size.
AML da ase . The low cy ome y da ase o he
AML expe imen has been collec ed om he DREAM6-
FlowCAP2 challenge, which was o ganized by he DREAM
p ojec and he FlowCAP ini ia i e (DREAM challenge
AML da ase can be accessed om Aghaeepou e al).4
We use he aining da ase ha consis s o low cy ome y
measu emen s o 179 pa ien s. Among hem, 23 pa ien s a e
AML posi i e and he emaining 156 pa ien s a e AML
nega i e. The low cy ome y measu emen o each pa ien
co esponds o se en join ly measu ed g oups (he ea e
called ubes) o se en quan i ies wi h a o al o 49 bioma ke
measu emen s pe cell. The bioma ke s a e summa ized in
Table 1 and include Fo wa d Sca e in linea scale (FS Lin),
Sidewa d Sca e in loga i hmic scale (SS Log), and i e luo-
escence in ensi ies (FL1–FL5) in loga i hmic scales. Fo
calib a ion pu poses, FS Lin, SS Log, and CD45-ECD we e
measu ed o all ubes and he o he 28 bioma ke s we e
measu ed only in one ube.
Cance cell line da ase . As ano he case s udy, we use
low cy ome y da a om a small sample se ing. The da a
come om a p os a e cance cell line 22R 1 s ained wi h
p opi dium iodide o cell cycle analysis.10 The cells a e
ans ec ed wi h miRNAs (ei he con ol o miR-193b) and
induced o p oli e a e by o e exp ession o cyclin D. The
da a consis o 16 samples, wi h 8 samples (wi hou cyclin
D o e exp ession) wi h ela i ely consis en cell cycle p o ile
and 8 samples (o e exp essing cyclin D) wi h an al e ed cell
cycle p o ile, ie, induced cell cycle ac i i y wi h an inc ease
in cells in DNA syn hesis phase. The samples o bo h classes
include ou epe i ions o wo ea men s, which a e consid-
e ed he e o ep esen he same class. Each sample con ains
12 measu ed channels, consis ing o wo sca e measu e-
men s and ou luo escence channels, bo h as a ea and
heigh measu emen s.
Table 1. Lis o se en ubes wi h bioma ke s p o ided in DREAM6 AML p edic ion da a.
FL1 Log FL2 Log FL3 Log FL4 Log FL5 Log
ube 1 s Lin ss Log IgG1- I C IgG1-Pe CD45-eCD IgG1-PC5 IgG1-PC7
ube 2 s Lin ss Log Kappa- I Lambda-Pe CD45-eCD CD19-PC5 CD20-PC7
ube 3 s Lin ss Log CD7- I C CD4-Pe CD45-eCD CD8-PC5 CD2-PC7
ube 4 s Lin ss Log CD15- I C CD13-Pe CD45-eCD CD16-PC5 CD56-PC7
ube 5 s Lin ss Log CD14- I C CD11c-Pe CD45-eCD CD64-PC5 CD33-PC7
ube 6 s Lin ss Log HLa-D - I C CD117-Pe CD45-eCD CD34-PC5 CD38-PC7
ube 7 s Lin ss Log CD5- I C CD19-Pe CD45-eCD CD3-PC5 CD10-PC7
Flow cy ome y-based classi ica ion in cance esea ch
77CanCe In o ma ICs 2015:14(s5)
Fea u e ex ac ion. Se e al ea u e ex ac ion me hods
can be used o ob ain meaning ul ea u es om aw low
cy ome y measu emen s. Fo ins ance, among widely
used ea u e ex ac ion echniques a e me hods based on
p incipal componen analysis and his og am compu a-
ion. Biehl e al p oposed s a is ical di e gences o ex ac
ea u es ha include momen s, median, and in e qua ile
ange.5 The leng h o he ea u e ec o was 186 in his case.
Ano he well-pe o med model was based on mul idimen-
sional en opic dis ance-based ea u es.4,5 Manninen e al.7
expanded he cell measu emen s o each ube o a highe
dimensional space. Following his ans o ma ion, LDA is
used o lowe he dimensionali y in o a single alue o each
measu emen . These p e ious s udies a e summa ized in
Table 2. Table 2 also abula es he es accu acy in e ms o he
a ea unde he ecei e ope a ing cha ac e is ics (ROC) cu e
(AUC) measu e o e a single ain/ es spli , which should no
be in e p e ed as a de ini i e measu e o accu acy, as he spli
o he samples is jus one ins ance o all possible spli s.
In his pape , we use one o he simples ea u e ex ac ion
echniques ha include only he mean and he s anda d de ia-
ion o he each measu emen . Fo he i s da ase , he leng h
o his ex ac ed ea u e ec o is 98, comp ising 49 mean
alues and 49 s anda d de ia ions. As seen in he expe imen s
o Expe imen al Resul s sec ion, hese ea u es a e su icien
o sepa a e he classes wi hou comp omising he p edic ion
accu acy. We will conside wo e sions o hese basic ea-
u es: he i s ea u e se con ains only he 49 mean alues o
he measu emen s, while he second ea u e ec o conside s
bo h mean alues and s anda d de ia ions, wi h al oge he 98
ea u es. The same app oach is used wi h he smalle da a-
se , hus p oducing wo di e en expe imen al cases. Be o e
aining he classi ie s, we no malized all ea u es o he in e -
al (0, 1).
L and egula iza ion. LR is a disc imina i e me hod
o modeling he class condi ional p obabili y densi ies
by he logis ic unc ion. Gi en an obse a ion ma ix
wi h N obse a ions, P ea u es, and co espond-
ing class labels y ∈1,…,C, we de ine LR model o he bina y
classi ica ion as,
(1)
and
(2)
He e, x ep esen s a ea u e ec o in he ea u e space
co esponding o class label y ∈ {0,1}, β0 is he in e cep , and β
ep esen s coe icien s o he logis ic model. We can de e mine
he model pa ame e s β0 and β om he aining da ase by
sol ing he
1 penalized LR p oblem,
(3)
whe e λ . 0 is he egula iza ion pa ame e . When he num-
be o aining da a is no la ge compa ed o he numbe o
ea u es, ie, P N, egula iza ion is used o sol e he o e i -
ing p oblem.11 In egula iza ion, an ex a e m, λ, is added,
which con ols he ade-o be ween he loss unc ion and
he size o he coe icien s. Mo e ecen ly, in ea u e selec-
ion,
1- egula ized LR has ecei ed much a en ion, as i
yields a spa se solu ion ha has ela i ely ew nonze o coe i-
cien s.12 This minimiza ion ask is analogous o leas absolu e
sh inkage and selec ion ope a o (Lasso) algo i hm p oposed by
Tibshi ani.13 In addi ion o his, se e al ex ensions o Lasso
ha e also been de eloped, such as g ouped Lasso,14,15 Dan zig
selec o ,16 elas ic ne ,17 and g aphical Lasso.18 In his pape ,
Table 2. S udies based on ea u e ex ac ion s a egies o DREAM
amL challenge da ase .
ACCURACY SIzE oF
FEATURE
VECToR
BRIEF DESCRIPTIoN
Biehl e al.51.00 186 Ex ac ion o
ea u es wi h
momen s, median
and in e qua ile
and lea ning ec o
quan iza ion is used
o p edic ion
Vila e al.41.00 31 Ex ac ion o
ea u es wi h
en opies and
his og am based
classi ie is used o
p edic ion
manninen
e al.7
1.00 (# o
e en s)
x 84
Expand ea u es o
highe dimension
and hen mapping
o 1-D using
linea disc iminan
analysis; logis ic
eg ession is used
o p edic ion
ou solu ion
( his s udy)
0.9989 49 Ex ac ion o ea u e
ec o om means o
measu emen s and
applying egula ized
logis ic eg ession
o p edic ion
ou solu ion
( his s udy)
0.9992 98 Ex ac ion o
ea u e ec o
om means and
s anda d de ia ion o
measu emen s and
applying egula ized
logis ic eg ession
o p edic ion
No e: The accu acy is measu ed in e ms o he AUC o a single ain/ es spli .
Hassan e al
78 CanCe In o ma ICs 2015:14(s5)
we use he GLMNET algo i hm by F iedman e al.19 ha
combines he
2 and
1 penal ies:
(4 )
whe e λ . 0 and α ∈[0,1]. The pa ame e α is a comp omise
be ween he
1 and
2 penal ies, he eby de e mining he ype
o egula iza ion. On he o he hand, he egula iza ion pa am-
e e λ con ols he amoun o egula iza ion. A e y la ge λ will
comple ely sh inks he coe icien s o ze o and may yield a null o
emp y model.
In gene al, he model pa ame e s λ and α a e selec ed using
he CV app oach.20 The da ase is andomly spli in o K mu u-
ally exclusi e subse s o app oxima ely equal sized. In K- old CV,
he p ocess is i e a ed k imes. A he k h i e a ion, he K h old
is e ained as es se and he emaining K − 1 olds a e used as
aining se o ain he model. Each o he K- olds is es ed exac ly
once. The es se assesses he quali y o he ained model. Then,
he K esul s a e combined o a e aged o p oduce a single es i-
ma ion o he model. The mos commonly used alues o K a e 5
and 10. In his expe imen , we se he alue o α = 1 and CV-10
is used o he selec ion o he model pa ame e λ and assessmen
o he model. As he ype o egula iza ion is de e mined by α,
se ing α = 0 p o ides
2 penal y ha is use ul in cases, whe e he
ea u es a e mu ually co ela ed. On he o he hand, α = 1 p o-
ides spa se solu ion wi h ewe coe icien s and, in u n, his is
sui able o implici ea u e selec ion. We ha e also expe imen ed
wi h 5- old CV, bu he esul s do no imp o e signi ican ly.
bayesian e o es ima o . A Bayesian app oach o e o
es ima ion was ecen ly in oduced in he con ex o disc e e
classi ie s21 and linea classi ie s.22 The Bayesian e o es ima-
o (BEE) es ima es he classi ica ion e o di ec ly om he
aining se and has shown o imp o e bo h he accu acy and
speed o he ac ual e o es ima e21,22 compa ed o adi ional
coun ing-based app oaches, such as CV. In ou ea lie pape s,
we ha e shown ha BEE has imp o ed he s abili y and speed
o compu a ion in he model selec ion con ex as well.9,23 We
will nex b ie ly e iew he de ini ion o BEE o a ixed linea
wo-class classi ie speci ied by he pa ame e s β and β0.
The Bayesian e o es ima o o linea classi ica ion
assumes ha he samples om each class a e independen and
iden ically dis ibu ed Gaussian andom a iables. Fo he wo
classes, he pa ame e s (mean and co a iance) o he Gaussian
model a e deno ed as θ0 and θ1 and he co esponding p io s
o he pa a me e s a e deno ed as p0(θ) and p1(θ). Then, he
pos e io p obabili y densi y unc ions (PDFs) o pa ame e s
o class c ∈{0,1} a e gi en by he Bayes’ ule:
(5)
whe e is he Gaussian class condi ional densi y
o c ∈{0,1}.
The Bayesian e o es ima o (BEE) is de ined as he
minimum mean squa ed es ima o by minimizing he expec a-
ion be ween he e o es ima e and he ue e o . This quan-
i y is composed o class-speci ic condi ional expec ed e o s
balanced by he p io s p(c) o he wo classes c ∈{0,1}22:
(6)
wi h he expec ed classi ica ion e o o samples om class c
gi en by
(7)
whe e εc(θ) deno es he ue classi ica ion e o .
The in eg al o Equa ion (7) can be e alua ed by assuming
an in e se Wisha p io o he class condi ional densi y:
(8)
whe e
υ
∈∈∈∈
×
,κ,Sm
PP
, and P
a e he hype pa am-
e e s o he Bayesian model and is he pa ame e
o he Gaussian dis ibu ion. Di e en choices o he alues
o hype pa ame e s lead o di e en e o es ima o s, bu we
will concen a e on a speci ic choice shown o be success ul in
ea lie wo ks9,22: κ = P + 2, ν = 0.5, s = I, and m = 0. Fo he
esul ing simpli ied closed- o m solu ion, e e o Re .9 Ma lab
and Py hon implemen a ions o BEE a e a ailable o down-
load (h ps://si es.google.com/si e/bayesiane o es ima e/).
Model selec ion. Model selec ion is a c i ical aspec in
classi ie design. Mo eo e , mos mode n classi ie s a e uned
by a se o hype pa ame e s, whose selec ion has a subs an ial
e ec on he esul ing accu acy as well. Thus, he selec ion
o an app op ia e model amily and he associa ed hype pa a-
me e s equi es an accu a e measu e o compa ing he accu-
acies o he model candida es. In ou wo k, we a e p ima ily
in e es ed in he selec ion o he egula iza ion pa ame e λ o
an LR classi ie . Howe e , i is o be no ed ha he me hodol-
ogy applies o any linea classi ie .
The p edic ion accu acy and selec ion o he bes model
can be quan i ied by e o es ima o s. CV es ima o is o en
used o selec he bes alue o he model selec ion pa ame e
λ along a egula iza ion pa h. As an example, e o cu es o
di e en alues o λ a e illus a ed in Figu e 1. Fo his pu pose,
we used he low cy ome y aining da a o 49 ea u es and 179
obse a ions. Fo an indi idual ube, each ea u e ep esen s he
a e age o he bioma ke in ensi ies. The e o cu es a e es i-
ma ed o di e en alues o λ anging om 10−9 o 100.
In he example o Figu e 1 (le panel), a 5- old CV p o-
cedu e is epea ed 100 imes and each s ep includes i e ain-
ing i e a ions on pa ial da a. The e o cu es ob ained o
100 i e a ions o 5- old CV illus a e he signi ican de ia ion
Flow cy ome y-based classi ica ion in cance esea ch
79CanCe In o ma ICs 2015:14(s5)
o he egula iza ion pa hs om one i e a ion o ano he . The
de ia ion is due o he andomness in spli ing he aining
da a in o olds, which esul s in an indi idual e o es ima e
o each spli . Mo eo e , o a e y small numbe o samples,
such as 5 o 10, he spli o alida ion and aining se s o he
K- old CV es ima o may no be app op ia e. In ac , in his
expe imen , he K- old CV app oach ails o es ima e he e o s
o smalle λ, as he numbe o samples spli by CV is insu i-
cien . On he o he hand, Figu e 1 ( igh panel) illus a es he
e o es ima e o BEE, which is a single de e minis ic e o
cu e. I is o be no ed ha he cu e ecognizes model o e -
i ing (e o es ima e s a s o inc ease o small egula iza-
ion e ms λ), al hough he e o is es ima ed di ec ly om he
aining se . No spli ing o i e a i e esampling is equi ed,
which in u n accele a es he compu a ion.
expe imen al esul s
In he ollowing sec ion, we p esen he expe imen al esul s.
Fi s , we demons a e di e en model selec ion c i e ia o es i-
ma e he signi ican ea u es. Then, we assess he pe o mances
o hose me hods in he AML classi ica ion case. Finally, we
p esen he esul s o he second, small sample case.
compa ison o model selec ion c i e ia. Typical
app oach o he selec ion o model pa ame e is CV.13 In his
pape , we also conside Bayesian e o es ima o (see Bayesian
E o Es ima o sec ion) and BIC24 as al e na i e app oaches
o es ima e he egula ized pa ame e . In o de o s udy he
beha io o di e en pa ame e selec ion c i e ia, we i s ain
he LR classi ie wi h he aining da a along he dec easing
sequence o egula iza ion pa h wi h log10 (λ)∈{0,–0.05,–0.1,
–0.15,…, –8.90, –8.95, –9.00}. Then, again he whole aining
da a a e used o es ima e he e o a e o each λ. Finally, o
each es ima o , we selec he model wi h λ alue ha achie es
he minimum e o a e. As esampling in CV-10 in oduces
andomness, in his case, we i e a e 200 imes and he esul
is a e aged. The de e minis ic na u e in BEE and BIC will
p oduce he same esul on he aining da a a each i e a ion.
The esul s a e summa ized in Table 3. Fo all me hods,
minimum e o a es, AUC, and he numbe o selec ed ea-
u es a e es ima ed om he whole aining da a. I is o be
no ed ha he epo ed AUC is compu ed om he aining
o emphasize ha all ea u e se s a e enough o pa i ion he
ea u e space in o wo ca ego ies pe ec ly. The es e o is
epo ed la e .
The esul s indica e ha he numbe o ea u es selec ed
by BEE me hod is lowe compa ed o hose o CV and BIC.
Fo he i s ea u e ec o wi h size 49, BEE selec s only 14
ea u es as signi ican , while o he second ea u e ec o wi h
−6 −5 −4 −3 −2 −1 0
0
0.02
0.04
0.06
0.08
0.1
0.12
0.14
0.16
0.18
0.2
Log10 (λ)
5− old c oss alida ion e o
−6 −5 −4 −3 −2 −1 0
0
0.02
0.04
0.06
0.08
0.1
0.12
0.14
0.16
0.18
0.2
Log10 (λ)
BEE
Figu e 1. Le : Examples o egula iza ion pa h e o cu es o 5- old CV o ou low cy ome y da a wi h heal hy and AML posi i e classes. Righ : The
co esponding Bee cu e.
Table 3. Pa ame e selec ion by di e en es ima o s: a e age numbe o selec ed ea u es, λ, aUC, and hei s anda d de ia ions wi h aining
da a.
METhoD FEATURE TYPE NUMBER oF SELECTED
FEATURES
SELECTED
Log10 (λ)
AUC
(TRAININg)
CV-10 mean 19.72 ± 2.41 −2.95 ± 2.30 0.9997 ± 0.0017
CV-10 mean and s d 23.91 ± 0.80 −4.14 ± 3.00 1 ± 0.0000
Bee mean 15 ± 0.00 −2.05 ± 0.00 0.9989 ± 0.0000
Bee mean and s d 13 ± 0.00 −1.80 ± 0.00 0.9992 ± 0.0000
BIC mean 20 ± 0.00 −5.85 ± 0.00 1 ± 0.0000
BIC mean and s d 24 ± 0.00 −5.70 ± 0.00 1 ± 0.0000
Hassan e al
80 CanCe In o ma ICs 2015:14(s5)
leng h 98, BEE selec s only 12 ea u es. Tables 4 and 5 lis
he selec ed ea u es, ie, signi ican bioma ke s along wi h he
co esponding coe icien alues. Due o he andomness in
CV-10, we only p esen he esul s o one i e a ion as an illus-
a ion: he e is a signi ican a ia ion o he selec ed ea u es
depending on he chosen CV spli . Howe e , i is o be no ed
ha he coe icien s o BIC and BEE a e no speci ic o his
pa icula i e a ion, as hey do no include he andom spli .
Pe o mance assessmen o he model selec ion c i e ia.
The pe o mances o he model selec ion me hods a e s udied
in he ollowing sec ion. The classi ica ion e o is conside ed
as he pe o mance c i e ion, and bo h alse posi i es (heal hy
con ol classi ied as AML) and alse nega i es (AML classi ied
as heal hy con ol) a e coun ed wi h equal weigh . The pe o -
mance o he Bayesian e o es ima o is benchma ked agains
hose o CV-10 and BIC o a di e en numbe o sample
sizes. Fo his pu pose, a andomly selec ed p opo ion o 10%,
15%, 20%–90%, and 95% is selec ed o aining he classi ie ,
while he emaining da a a e used o pe o mance assessmen .
Fo each aining sample, he expe imen is execu ed 200 imes
by gene a ing a new aining se each ime.
Classi ica ion e o s. The e o cu es o di e en sample
sizes a e shown in Figu e 2. The p ocedu e is epea ed 200
imes o each aining sample size, and he a e age o clas-
si ica ion e o is compu ed o each model selec ion c i e ion.
Wi h a e y small numbe o aining samples, such as 10%
o 15% o he da ase , BEE p o ides imp o ed accu acy o e
CV-10 and BIC (Fig. 2 le and igh panels). Fo ins ance,
wi h 10% aining samples, he classi ica ion e o s o he
model selec ed by CV-10 a e 7.56% (Fig. 2 le panel) and
7.41% (Fig. 2 igh panel) highe han hose o BEE. In case
o BIC, he classi ica ion e o is 7.81% highe han ha o
BEE (Fig. 2 le panel). As he numbe o aining samples
inc eases, o example, abo e 60%, he pe o mance o BIC
exceeds han ha o BEE (Fig. 2 le panel). Howe e , he
pe o mance o BIC is simila o ha o BEE when mo e ea-
u es a e in ol ed in he expe imen (Fig. 2 igh panel).
Table 4. The nonze o coe icien s o ea u es wi h mean.
TUBE FEATURE 10-FoLD CV BEE BIC
Cons an −13.38 −3.50 −13.54
ube 1 sLin 0.75 00.78
ube 1 ssLog −5.92 −0.73 −6.01
ube 1 L1:IgG1- I C −0.46 −0.19 −0.46
ube 1 L4:IgG1-PC5 −2.08 0−2.13
ube 1 L5:IgG1-PC7 −3.07 −0.19 −3.14
ube 2 sLin 0.001 0 0
ube 2 L5:CD20-PC7 2.76 02.78
ube 3 ssLog 0−0.77 0
ube 3 L4:CD8-PC5 −1.94 −0.10 −1.97
ube 4 sLin 0.97 00.97
ube 4 L1:CD15 - I C −4.77 0−4.82
ube 4 L2:CD13-Pe 3.44 0.21 3.48
ube 4 L4:CD16-PC5 0−0.09 0
ube 4 L5:CD56-PC7 3.29 0.75 3.35
ube 5 sLin 2.35 02.40
ube 5 L2:CD11c-Pe −0.15 0−0.16
ube 5 L3:CD45-eCD −1.84 −0.02 −1.84
ube 5 L4:CD64-PC5 1.66 01.69
ube 5 L5:CD33-PC7 0.75 0.60 0.76
ube 6 L2:CD117-Pe 4.82 0.89 4.88
ube 6 L4:CD34-PC5 6.88 0.72 6.99
ube 6 L5:CD38-PC7 00.41 0
ube 7 L1:CD5 - I C −4.68 −0.19 −4.74
No e: The size o he ea u e se s is 49, o which he CV, BEE, and BIC selec
20, 14, and 19 ea u es, espec i ely.
Table 5. The nonze o coe icien s o ea u es wi h mean and
s anda d de ia ion.
TUBE FEATURE 10-FoLD CV BEE BIC
Cons an −13.47 −3.27 −13.47
ube 1 sLin mean 0.027 00.027
ube 1 ssLog mean −4.49 −0.47 −4.49
ube 1 L1:IgG1- I C s d −3.80 −0.22 −3.80
ube 1 L5:IgG1-PC7 s d −1.60 −0.16 −1.60
ube 2 L5:CD20-PC7 mean 0.50 00.50
ube 3 ssLog mean 0−0.29 0
ube 3 L5:CD2-PC7 mean −0.48 0−0.48
ube 3 L5:CD2-PC7 s d −1.21 0−1.21
ube 4 sLin mean 0.05 00.05
ube 4 L1:CD15 - I C mean −0.14 0−0.14
ube 4 L2:CD13-Pe mean 2.50 02.50
ube 4 L4:CD16-PC5 mean −1.32 0−1.32
ube 4 L4:CD16-PC5 s d 0−0.39 0
ube 4 L5:CD56-PC7 s d 6.07 0.45 6.07
ube 5 L1:CD14- I C s d −0.004 0−0.004
ube 5 L3:CD45-eCD mean −2.51 0−2.51
ube 5 L5:CD33-PC7 mean 1.47 0.32 1.47
ube 5 L5:CD33-PC7 s d −0.70 0−0.70
ube 6 ssLog s d 0−0.23 0
ube 6 L2:CD117-Pe mean 3.19 0.46 3.19
ube 6 L2:CD117-Pe s d 2.19 02.19
ube 6 L4:CD34-PC5 s d 1.88 0.75 1.88
ube 6 L5:CD38-PC7 mean 1.31 0.44 1.31
ube 7 sLin mean 1.39 01.39
ube 7 sLin s d 0−0.16 0
ube 7 L1:CD5 - I C s d −1.16 0−1.16
ube 7 L5:CD10-PC7 s d 0.04 00.04
No e: The size o he ea u e se s is 98, o which he CV, BEE, and BIC selec
23, 12, and 23 ea u es, espec i ely.
Flow cy ome y-based classi ica ion in cance esea ch
81CanCe In o ma ICs 2015:14(s5)
A ea unde he ROC cu e. In his sec ion, we e alua e he
pe o mance in e ms o AUC. Figu e 3 illus a es he a e age o
AUC o di e en aining sample sizes. He e, he BEE me hod
achie es imp o emen o e he o he me hods. Wi h small ain-
ing samples, o ins ance, 10%, he a e age AUC o BEE is 1.11%
(Fig. 3 le panel) and 1.30% (Fig. 3 igh panel) highe han ha
o CV-10. As he numbe o aining samples inc eases, CV-10
and BIC also con e ge owa d he esul s o BEE; howe e , he
BEE selec ed model consis en ly esul s in he highes AUC
sco e. Wi h he la ge ea u e ec o ha includes he measu e-
men s o mean and s anda d de ia ion, he a e age AUC cu es
o BEE and BIC ollow he simila pa e n (Fig. 3 igh panel).
Numbe o selec ed ea u es. We u he assess he pe -
o mances o he es ima o s using ea u e selec ion c i e ia.
A each i e a ion, we de e mine he o al numbe o selec ed
ea u es ha ha e nonze o alues o a di e en numbe o
aining samples. Then, we compu e he a e age and he
a iabili y (ie, s anda d de ia ion) o he selec ed ea u es
o di e en aining samples. The esul s a e illus a ed in
Figu e 4. Fo BEE, he a e age numbe o selec ed ea u es
is lowe in amoun compa ed o hose o CV and BIC (Fig. 4
op panel). Fo ins ance, wi h 95% aining samples, BEE
equi es 36.49% and 33.89% less ea u es han CV and BIC,
espec i ely, o model p edic ion (Fig. 4 op- igh panel).
Mo eo e , he a iabili y in selec ed ea u es using BEE is
also compa able (Fig. 4 bo om panel). The CV-10 has he
wo s pe o mance. Al hough BIC shows ha he de ia-
ion in ea u e selec ion a di e en i e a ions is smalle , he
0 20 40 60 80 100 120 140 160 180
0.02
0.03
0.04
0.05
0.06
0.07
Numbe o samples used o aining
Classi ica ion e o s
CV-10
BEEp
BIC
0 20 40 60 80 100 120 140 160 180
0.02
0.03
0.04
0.05
0.06
0.07
Numbe o samples used o aining
Classi ica ion e o s
CV-10
BEEp
BIC
Figu e 2. The a e age classi ica ion e o cu es o CV-10, BEE wi h p ope p io (BEEp), and BIC.
No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s.
0 20 40 60 80 100 120 140 160 180
0.9
0.91
0.92
0.93
0.94
0.95
0.96
0.97
0.98
0.99
1
Numbe o samples used o aining
AUC
CV-10
BEEp
BIC
0 20 40 60 80 100 120 140 160 180
0.9
0.91
0.92
0.93
0.94
0.95
0.96
0.97
0.98
0.99
1
Numbe o samples used o aining
AUC
CV-10
BEEp
BIC
Figu e 3. The a e age AUCs o CV-10, BEE wi h p ope p io (BEEp), and BIC.
No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s.
Hassan e al
82 CanCe In o ma ICs 2015:14(s5)
a e age numbe o selec ed ea u es is highe han ha o
o he s (Fig. 4 op panel).
Simila i y o he selec ed ea u e se s. Ano he pe o mance
measu emen is he s abili y o selec ing he same ea u e a
di e en i e a ions. Fo his pu pose, Sø ensen–Dice coe -
icien 25 is used, which measu es he deg ee o simila i y
be ween selec ed ea u es o wo di e en i e a ions. The
anges can a y om 0 o 1. The alues closes o 1 indica e a
high-deg ee o simila i y.
Fo di e en aining samples, we i s de e mine
which ea u es a e selec ed a each i e a ion. As he model
selec ion p ocess is epea ed 200 imes, we es ima e he
simila i y as he mean dice coe icien o each o he 200!/
(2! × (200 − 2)!) = 19,900 possible pai s o selec ed ea u e
se s. The esul s a e shown in Figu e 5. In e ms o s abil-
i y, he pe o mance o BEE is subs an ially be e han hose
o he o he me hods, as he selec ed ea u e se s a e mos
simila wi h ha c i e ion – a signi ican issue when ying
o unde s and he biological mechanisms behind he da a.
Fo example, wi h 60% aining samples, he dice coe icien
o BEE is 6.03% highe han ha o CV (Fig. 5 le panel).
On he o he hand, wi h 90% aining samples, he dice coe -
icien o BEE is 5.81% highe han ha o CV and 4.90%
highe han ha o BIC (Fig. 5 igh panel). Indeed, he dice
coe icien is un a o able o CV wi h small aining samples:
The dice coe icien is lowes among he al e na i es, indica -
ing ha he selec ed ea u e se s wi h he CV c i e ion ha e
high a iabili y.
small sample case wi h a cance cell line. Fo u he
con idence on he p esen ed me hod, we analyze da a om a
cance cell line in a small sample se ing. As desc ibed p e i-
ously, we conside ed he classi ica ion accu acy, AUC mea-
su e, and he numbe o selec ed a iables bo h wi h and
wi hou s anda d de ia ion ea u es (Figs. 6–8). In his case,
0 20 40 60 80 100 120 140 160 180
8
10
12
14
16
18
20
22
24
Numbe o samples used o aining
A e age numbe o ea u es
CV-10
BEEp
BIC
0 20 40 60 80 100 120 140 160 180
8
10
12
14
16
18
20
22
24
Numbe o samples used o aining
A e age numbe o ea u es
CV-10
BEEp
BIC
0 20 40 60 80 100 120 140 160 180
1
1.5
2
2.5
3
3.5
4
Numbe o samples used o aining
S anda d de ia ion o numbe o ea u es
CV-10
BEEp
BIC
0 20 40 60 80 100 120 140 160 180
1.5
2
2.5
3
3.5
4
4.5
Numbe o samples used o aining
S anda d de ia ion o numbe o ea u es
CV-10
BEEp
BIC
Figu e 4. Compa isons o he numbe o selec ed ea u es o CV-10, BEE wi h p ope p io (BEEp), and BIC.
No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s. Top: a e age
numbe o selec ed ea u es. Bo om: s anda d de ia ion o numbe o he selec ed ea u es.
Flow cy ome y-based classi ica ion in cance esea ch
83CanCe In o ma ICs 2015:14(s5)
0 20 40 60 80 100 120 140 160 180
0.3
0.4
0.5
0.6
0.7
0.8
0.9
1
Numbe o samples used o aining
Dice index
CV-10
BEEp
BIC
0 20 40 60 80 100 120 140 160 180
0.3
0.4
0.5
0.6
0.7
0.8
0.9
1
Numbe o samples used o aining
Dice index
CV-10
BEEp
BIC
Figu e 5. Compa ison o he s abili y o selec ing ea u es o CV-10, BEEp, and BIC.
No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s.
2 4 6 8 10 12 14
0.2
0.25
0.3
0.35
0.4
0.45
Numbe o samples used o aining
Classi ica ion e o s
LOOCV
BEEp
BIC
2 4 6 8 10 12 14
0.2
0.25
0.3
0.35
0.4
0.45
Numbe o samples used o aining
Classi ica ion e o s
LOOCV
BEEp
BIC
Figu e 6. The a e age classi ica ion e o cu es o LOOCV, BEEp, and BIC.
No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s.
2 4 6 8 10 12 14
0.55
0.6
0.65
0.7
0.75
0.8
0.85
0.9
Numbe o samples used o aining
AUC
LOOCV
BEEp
BIC
2 4 6 8 10 12 14
0.55
0.6
0.65
0.7
0.75
0.8
0.85
0.9
0.95
Numbe o samples used o aining
AUC
LOOCV
BEEp
BIC
Figu e 7. The a e age AUC cu es o LOOCV, BEEp, and BIC.
No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s.