scieee Science in your language
[en] (orig)

Flow Cytometry-Based Classification in Cancer Research: A View on Feature Selection

Abstract

In this paper, we study the problem of feature selection in cancer-related machine learning tasks. In particular, we study the accuracy and stability of different feature selection approaches within simplistic machine learning pipelines. Earlier studies have shown that for certain cases, the accuracy of detection can easily reach 100% given enough training data. Here, however, we concentrate on simplifying the classification models with and seek for feature selection approaches that are reliable even with extremely small sample sizes. We show that as much as 50% of features can be discarded without compromising the prediction accuracy. Moreover, we study the model selection problem among the ℓ1 regularization path of logistic regression classifiers. To this aim, we compare a more traditional cross-validation approach with a recently proposed Bayesian error estimator.

Read accessible full text

Flow Cytometry-Based Classification in Cancer Research: A View on Feature Selection

Author: Hassan, Sakira,Ruusuvuori, Pekka,Latonen, Leena,Huttunen, Heikki
Year: 2016
Source: https://trepo.tuni.fi/bitstream/10024/99726/1/flow-cytometry-based_2016.pdf
75CanCe In o ma ICs 2015:14(s5)
In oduc ion
Flow cy ome y enables quan i a i e measu emen o single-
cell p ope ies h ough isible and luo escen ligh in a high-
h oughpu manne . The measu ed signals include luo escence
emission and ligh sca e . Flow cy ome y has been ou inely
used o de ec ing malignancies om blood samples.1 Th ough
echnological ad ances, measu ing he combina ion o luo-
escen signals om se e al di e en channels has enabled he
use o high-dimensional da a o s udies, such as cy ome ic
inge p in ing2 and la ge-scale analysis o cell ypes.3
In his pape , we s udy he analysis o low cy ome y
da a om he ea u e selec ion poin o iew. Mo e speci i-
cally, low cy ome y is able o p oduce la ge quan i ies o
pa ially edundan measu emen da a, and he selec ion o
impo an quan i ies wi hin he la ge body o measu emen s is
o in e es . Mo eo e , a ypical scena io con ains la ge quan-
i ies o measu emen da a bu may be limi ed o only a ew
pa ien s. Thus, an ideal me hod would dis il only he essen ial
pa s o he measu emen s om each pa ien , while p oduc-
ing eliable and well gene alizing esul s when only a small
amoun o indi iduals is a ailable in he aining da a.
We will concen a e ou a en ion on wo pa icula se s
o low cy ome y da a. The i s se o igina es om he acu e
myeloid leukemia (AML) p edic ion challenge o he DREAM
ini ia i e in 2013.4 The compe i ion a ac ed a numbe o
eams, and as se e al esea che s use he da a as pa o hei
wo k, he challenge da a ha e become a s anda d benchma k
wi hin he ield. Fo example, Aghaeepou e al.4 p esen s a
la ge pool o analysis app oaches om he DREAM challenge.
Among classi ica ion me hods p esen ed in he li e a u e,
he e a e se e al sophis ica ed machine lea ning app oaches,
such as lea ning ec o quan iza ion,5 co ela i e ma ix map-
ping, and ela i e en opy di e ences.6 The s eng h o da a-
d i en app oaches elying on supe ised classi ica ion is hei
abili y o handle high-dimensional da a wi hou equi ing
p io knowledge o he biological applica ion.
The DREAM AML da a ep esen a ela i ely la ge-scale
expe imen consis ing o al oge he 179 pa ien s. Al hough i
is a small numbe o adi ional machine lea ning p oblems,
he numbe o pa ien s is unusually la ge o a biological s udy.
To his aim, we use ano he da ase ha ep esen s a mo e
commonly encoun e ed sample size o 16 samples ex ac ed
om a p os a e cance cell line. This da ase , wi h wo di -
e en ea men s and a low numbe o samples, p esen s a
non i ial bu common challenge o p edic ion and ela ed
ea u e selec ion. Mo e in o ma ion on he wo da ase s is
p o ided in Da a sec ion.
In ou ea lie wo k,7 we p esen ed a supe ised classi ica-
ion pipeline based on linea disc iminan analysis (LDA) and
logis ic eg ession (LR) classi ie s. B ie ly, he me hod i s
Flow Cy ome y-Based Classi ica ion in Cance
Resea ch: A View on Fea u e Selec ion
s. saki a Hassan1, Pekka uusu uo i2,3, Leena La onen3 and Heikki Hu unen1
1Depa men o Signal P ocessing, Tampe e Uni e si y o Technology, Tampe e, Finland. 2Po i Depa men , Tampe e Uni e si y o
Technology, Po i, Finland. 3BioMediTech, Uni e si y o Tampe e, Tampe e, Finland.
Supplemen a y Issue: S a is ical Sys ems Theo y in Cance Modeling, Diagnosis, and The apy
Abs Ac : In his pape , we s udy he p oblem o ea u e selec ion in cance - ela ed machine lea ning asks. In pa icula , we s udy he accu acy and
s abili y o di e en ea u e selec ion app oaches wi hin simplis ic machine lea ning pipelines. Ea lie s udies ha e shown ha o ce ain cases, he accu acy
o de ec ion can easily each 100% gi en enough aining da a. He e, howe e , we concen a e on simpli ying he classi ica ion models wi h and seek o
ea u e selec ion app oaches ha a e eliable e en wi h ex emely small sample sizes. We show ha as much as 50% o ea u es can be disca ded wi hou
comp omising he p edic ion accu acy. Mo eo e , we s udy he model selec ion p oblem among he 
1 egula iza ion pa h o logis ic eg ession classi ie s.
To his aim, we compa e a mo e adi ional c oss- alida ion app oach wi h a ecen ly p oposed Bayesian e o es ima o .
Keywo ds: AML, leukemia, low cy ome y, logis ic eg ession, e o es ima ion, model selec ion
SUPPLEMENT: s a is ical sys ems heo y in Cance modeling, Diagnosis, and he apy
CITATIoN: Hassan e al. Flow Cy ome y-Based Classi ica ion in Cance Resea ch: A View
on ea u e selec ion. Cance In o ma ics 2015:14(s5) 75–85 doi: 10.4137/CIn.s30795.
TYPE: o iginal esea ch
RECEIVED: no embe 18, 2015. RESUBMITTED: eb ua y 01, 2016. ACCEPTED FoR
PUBLICATIoN: eb ua y 07, 2016.
ACADEMIC EDIToR: J. T. E i d, Edi o in Chie
PEER REVIEw: six pee e iewe s con ibu ed o he pee e iew epo . e iewe s’
epo s o aled 1860 wo ds, excluding any con iden ial commen s o he academic edi o .
FUNDINg: Au ho s disclose no ex e nal unding sou ces.
CoMPETINg INTERESTS: Au ho s disclose no po en ial con lic s o in e es .
CoRRESPoNDENCE: [email p o ec ed]
CoPYRIghT: © he au ho s, publishe and licensee Libe as academica Limi ed. his is
an open-access a icle dis ibu ed unde he e ms o he C ea i e Commons CC-BY-NC
3.0 License.
Pape subjec o independen expe blind pee e iew. all edi o ial decisions made
by independen academic edi o . Upon submission manusc ip was subjec o an i-
plagia ism scanning. P io o publica ion all au ho s ha e gi en signed con i ma ion o
ag eemen o a icle publica ion and compliance wi h all applicable e hical and legal
equi emen s, including he accu acy o au ho and con ibu o in o ma ion, disclosu e o
compe ing in e es s and unding sou ces, compliance wi h e hical equi emen s ela ing
o human and animal s udy pa icipan s, and compliance wi h any copy igh equi emen s
o hi d pa ies. This jou nal is a membe o he Commi ee on Publica ion E hics (COPE).
Published by Libe as academica. Lea n mo e abou his jou nal.
Hassan e al
76 CanCe In o ma ICs 2015:14(s5)
ans o ms he measu emen da a in o highe dimensional
space by gene a ing combined ea u es wi h mul iplica ions and
di isions be ween measu emen s. Following his mapping in o
highe dimensional space, LDA is used o lowe ing he dimen-
sionali y in o a single alue pe measu emen . Then, empi ical
dis ibu ion unc ions (EDFs) a e cons uc ed om LDA esul s
o AML-posi i e and AML-nega i e sample classes and com-
pa ed o aining EDFs o bo h classes. The compa ison esul s
in wo simila i y alues pe g oup o measu emen s, and hese
esul s a e ed o he LR classi ie o a inal AML p edic ion
esul . Ou app oach, oge he wi h al e na i e well-pe o ming
app oaches,4,5 ep esen s a ela i ely complica ed pipeline o
somewha a bi a y compu a ion s eps. Thus, ou in e es is o
simpli y hese pipelines in o a simple collec ion o ob ious ea-
u es, while s ill e aining a good accu acy.
Ou app oach he e is o use LR classi ie applied o sum-
ma ize ea u es, which a e he mean and s anda d de ia ion o
he measu emen s ins ead o he comple e da a. This educes
he numbe o ea u es used in classi ica ion and, subsequen ly,
also he model complexi y. An essen ial pa o classi ie
design is e o es ima ion, which guides model selec ion.8 Ou
s a egy o model selec ion is o apply he ecen ly in oduced
Bayesian e o es ima o (BEE).9 We compa e BEE model
selec ion wi h a adi ional 10- old c oss- alida ion (CV-10)
e o es ima ion, as well as wi h Bayesian in o ma ion c i e-
ion (BIC)-based model selec ion, and conclude ha he p o-
posed app oach enables accu a e p edic ion o low cy ome y
da a wi h ewe measu emen s and a less complex classi ie
model han hose p e iously p esen ed in he li e a u e.
The es o his pape is o ganized as ollows. In Ma e ials
and Me hods sec ion, we desc ibe he da a and me hods used
in his s udy and b ie ly discuss how ea u e selec ion is com-
monly done in machine lea ning. Expe imen al Resul s sec ion
p esen s he esul s o ou expe imen s wi h di e en model
and ea u e selec ion c i e ia o he ma e ials in oduced in
Ma e ials and Me hods sec ion. Finally, in Conclusions sec-
ion, we summa ize he wo k and discuss he conclusions o
he esul s.
Ma e ials and Me hods
In his sec ion, we desc ibe he da ase s used in his pape .
We also gi e a b ie o e iew o modeling me hod o ea u e
selec ion. In addi ion o his, we in oduce he s a e-o - he-a
Bayesian e o es ima o (BEE) o model pa ame e selec-
ion. Finally, we p esen an example whe e he pe o mance o
BEE is benchma ked agains o he model selec ion c i e ia.
da a. In his wo k, we s udy wo da ase s: A la ge se
wi h 179 samples and a smalle se wi h 16 samples. These
wo case s udies ep esen di e en classi ica ion challenges in
e ms o bo h applica ion and sample size.
AML da ase . The low cy ome y da ase o he
AML expe imen has been collec ed om he DREAM6-
FlowCAP2 challenge, which was o ganized by he DREAM
p ojec and he FlowCAP ini ia i e (DREAM challenge
AML da ase can be accessed om Aghaeepou e al).4
We use he aining da ase ha consis s o low cy ome y
measu emen s o 179 pa ien s. Among hem, 23 pa ien s a e
AML posi i e and he emaining 156 pa ien s a e AML
nega i e. The low cy ome y measu emen o each pa ien
co esponds o se en join ly measu ed g oups (he ea e
called ubes) o se en quan i ies wi h a o al o 49 bioma ke
measu emen s pe cell. The bioma ke s a e summa ized in
Table 1 and include Fo wa d Sca e in linea scale (FS Lin),
Sidewa d Sca e in loga i hmic scale (SS Log), and i e luo-
escence in ensi ies (FL1–FL5) in loga i hmic scales. Fo
calib a ion pu poses, FS Lin, SS Log, and CD45-ECD we e
measu ed o all ubes and he o he 28 bioma ke s we e
measu ed only in one ube.
Cance cell line da ase . As ano he case s udy, we use
low cy ome y da a om a small sample se ing. The da a
come om a p os a e cance cell line 22R 1 s ained wi h
p opi dium iodide o cell cycle analysis.10 The cells a e
ans ec ed wi h miRNAs (ei he con ol o miR-193b) and
induced o p oli e a e by o e exp ession o cyclin D. The
da a consis o 16 samples, wi h 8 samples (wi hou cyclin
D o e exp ession) wi h ela i ely consis en cell cycle p o ile
and 8 samples (o e exp essing cyclin D) wi h an al e ed cell
cycle p o ile, ie, induced cell cycle ac i i y wi h an inc ease
in cells in DNA syn hesis phase. The samples o bo h classes
include ou epe i ions o wo ea men s, which a e consid-
e ed he e o ep esen he same class. Each sample con ains
12 measu ed channels, consis ing o wo sca e measu e-
men s and ou luo escence channels, bo h as a ea and
heigh measu emen s.
Table 1. Lis o se en ubes wi h bioma ke s p o ided in DREAM6 AML p edic ion da a.
FL1 Log FL2 Log FL3 Log FL4 Log FL5 Log
ube 1 s Lin ss Log IgG1- I C IgG1-Pe CD45-eCD IgG1-PC5 IgG1-PC7
ube 2 s Lin ss Log Kappa- I Lambda-Pe CD45-eCD CD19-PC5 CD20-PC7
ube 3 s Lin ss Log CD7- I C CD4-Pe CD45-eCD CD8-PC5 CD2-PC7
ube 4 s Lin ss Log CD15- I C CD13-Pe CD45-eCD CD16-PC5 CD56-PC7
ube 5 s Lin ss Log CD14- I C CD11c-Pe CD45-eCD CD64-PC5 CD33-PC7
ube 6 s Lin ss Log HLa-D - I C CD117-Pe CD45-eCD CD34-PC5 CD38-PC7
ube 7 s Lin ss Log CD5- I C CD19-Pe CD45-eCD CD3-PC5 CD10-PC7
Flow cy ome y-based classi ica ion in cance esea ch
77CanCe In o ma ICs 2015:14(s5)
Fea u e ex ac ion. Se e al ea u e ex ac ion me hods
can be used o ob ain meaning ul ea u es om aw low
cy ome y measu emen s. Fo ins ance, among widely
used ea u e ex ac ion echniques a e me hods based on
p incipal componen analysis and his og am compu a-
ion. Biehl e al p oposed s a is ical di e gences o ex ac
ea u es ha include momen s, median, and in e qua ile
ange.5 The leng h o he ea u e ec o was 186 in his case.
Ano he well-pe o med model was based on mul idimen-
sional en opic dis ance-based ea u es.4,5 Manninen e al.7
expanded he cell measu emen s o each ube o a highe
dimensional space. Following his ans o ma ion, LDA is
used o lowe he dimensionali y in o a single alue o each
measu emen . These p e ious s udies a e summa ized in
Table 2. Table 2 also abula es he es accu acy in e ms o he
a ea unde he ecei e ope a ing cha ac e is ics (ROC) cu e
(AUC) measu e o e a single ain/ es spli , which should no
be in e p e ed as a de ini i e measu e o accu acy, as he spli
o he samples is jus one ins ance o all possible spli s.
In his pape , we use one o he simples ea u e ex ac ion
echniques ha include only he mean and he s anda d de ia-
ion o he each measu emen . Fo he i s da ase , he leng h
o his ex ac ed ea u e ec o is 98, comp ising 49 mean
alues and 49 s anda d de ia ions. As seen in he expe imen s
o Expe imen al Resul s sec ion, hese ea u es a e su icien
o sepa a e he classes wi hou comp omising he p edic ion
accu acy. We will conside wo e sions o hese basic ea-
u es: he i s ea u e se con ains only he 49 mean alues o
he measu emen s, while he second ea u e ec o conside s
bo h mean alues and s anda d de ia ions, wi h al oge he 98
ea u es. The same app oach is used wi h he smalle da a-
se , hus p oducing wo di e en expe imen al cases. Be o e
aining he classi ie s, we no malized all ea u es o he in e -
al (0, 1).
L and egula iza ion. LR is a disc imina i e me hod
o modeling he class condi ional p obabili y densi ies
by he logis ic unc ion. Gi en an obse a ion ma ix
wi h N obse a ions, P ea u es, and co espond-
ing class labels y ∈1,…,C, we de ine LR model o he bina y
classi ica ion as,
(1)
and
(2)
He e, x ep esen s a ea u e ec o in he ea u e space
co esponding o class label y ∈ {0,1}, β0 is he in e cep , and β
ep esen s coe icien s o he logis ic model. We can de e mine
he model pa ame e s β0 and β om he aining da ase by
sol ing he 
1 penalized LR p oblem,
(3)
whe e λ . 0 is he egula iza ion pa ame e . When he num-
be o aining da a is no la ge compa ed o he numbe o
ea u es, ie, P  N, egula iza ion is used o sol e he o e i -
ing p oblem.11 In egula iza ion, an ex a e m, λ, is added,
which con ols he ade-o be ween he loss unc ion and
he size o he coe icien s. Mo e ecen ly, in ea u e selec-
ion, 
1- egula ized LR has ecei ed much a en ion, as i
yields a spa se solu ion ha has ela i ely ew nonze o coe i-
cien s.12 This minimiza ion ask is analogous o leas absolu e
sh inkage and selec ion ope a o (Lasso) algo i hm p oposed by
Tibshi ani.13 In addi ion o his, se e al ex ensions o Lasso
ha e also been de eloped, such as g ouped Lasso,14,15 Dan zig
selec o ,16 elas ic ne ,17 and g aphical Lasso.18 In his pape ,
Table 2. S udies based on ea u e ex ac ion s a egies o DREAM
amL challenge da ase .
ACCURACY SIzE oF
FEATURE
VECToR
BRIEF DESCRIPTIoN
Biehl e al.51.00 186 Ex ac ion o
ea u es wi h
momen s, median
and in e qua ile
and lea ning ec o
quan iza ion is used
o p edic ion
Vila e al.41.00 31 Ex ac ion o
ea u es wi h
en opies and
his og am based
classi ie is used o
p edic ion
manninen
e al.7
1.00 (# o
e en s)
x 84
Expand ea u es o
highe dimension
and hen mapping
o 1-D using
linea disc iminan
analysis; logis ic
eg ession is used
o p edic ion
ou solu ion
( his s udy)
0.9989 49 Ex ac ion o ea u e
ec o om means o
measu emen s and
applying egula ized
logis ic eg ession
o p edic ion
ou solu ion
( his s udy)
0.9992 98 Ex ac ion o
ea u e ec o
om means and
s anda d de ia ion o
measu emen s and
applying egula ized
logis ic eg ession
o p edic ion
No e: The accu acy is measu ed in e ms o he AUC o a single ain/ es spli .
Hassan e al
78 CanCe In o ma ICs 2015:14(s5)
we use he GLMNET algo i hm by F iedman e al.19 ha
combines he 
2 and 
1 penal ies:
(4 )
whe e λ . 0 and α ∈[0,1]. The pa ame e α is a comp omise
be ween he 
1 and 
2 penal ies, he eby de e mining he ype
o egula iza ion. On he o he hand, he egula iza ion pa am-
e e λ con ols he amoun o egula iza ion. A e y la ge λ will
comple ely sh inks he coe icien s o ze o and may yield a null o
emp y model.
In gene al, he model pa ame e s λ and α a e selec ed using
he CV app oach.20 The da ase is andomly spli in o K mu u-
ally exclusi e subse s o app oxima ely equal sized. In K- old CV,
he p ocess is i e a ed k imes. A he k h i e a ion, he K h old
is e ained as es se and he emaining K − 1 olds a e used as
aining se o ain he model. Each o he K- olds is es ed exac ly
once. The es se assesses he quali y o he ained model. Then,
he K esul s a e combined o a e aged o p oduce a single es i-
ma ion o he model. The mos commonly used alues o K a e 5
and 10. In his expe imen , we se he alue o α = 1 and CV-10
is used o he selec ion o he model pa ame e λ and assessmen
o he model. As he ype o egula iza ion is de e mined by α,
se ing α = 0 p o ides 
2 penal y ha is use ul in cases, whe e he
ea u es a e mu ually co ela ed. On he o he hand, α = 1 p o-
ides spa se solu ion wi h ewe coe icien s and, in u n, his is
sui able o implici ea u e selec ion. We ha e also expe imen ed
wi h 5- old CV, bu he esul s do no imp o e signi ican ly.
bayesian e o es ima o . A Bayesian app oach o e o
es ima ion was ecen ly in oduced in he con ex o disc e e
classi ie s21 and linea classi ie s.22 The Bayesian e o es ima-
o (BEE) es ima es he classi ica ion e o di ec ly om he
aining se and has shown o imp o e bo h he accu acy and
speed o he ac ual e o es ima e21,22 compa ed o adi ional
coun ing-based app oaches, such as CV. In ou ea lie pape s,
we ha e shown ha BEE has imp o ed he s abili y and speed
o compu a ion in he model selec ion con ex as well.9,23 We
will nex b ie ly e iew he de ini ion o BEE o a ixed linea
wo-class classi ie speci ied by he pa ame e s β and β0.
The Bayesian e o es ima o o linea classi ica ion
assumes ha he samples om each class a e independen and
iden ically dis ibu ed Gaussian andom a iables. Fo he wo
classes, he pa ame e s (mean and co a iance) o he Gaussian
model a e deno ed as θ0 and θ1 and he co esponding p io s
o he pa a me e s a e deno ed as p0(θ) and p1(θ). Then, he
pos e io p obabili y densi y unc ions (PDFs) o pa ame e s
o class c ∈{0,1} a e gi en by he Bayes’ ule:
(5)
whe e is he Gaussian class condi ional densi y
o c ∈{0,1}.
The Bayesian e o es ima o (BEE) is de ined as he
minimum mean squa ed es ima o by minimizing he expec a-
ion be ween he e o es ima e and he ue e o . This quan-
i y is composed o class-speci ic condi ional expec ed e o s
balanced by he p io s p(c) o he wo classes c ∈{0,1}22:
(6)
wi h he expec ed classi ica ion e o o samples om class c
gi en by
(7)
whe e εc(θ) deno es he ue classi ica ion e o .
The in eg al o Equa ion (7) can be e alua ed by assuming
an in e se Wisha p io o he class condi ional densi y:
(8)
whe e
υ
∈∈∈∈
×
,κ,Sm
PP
, and P
a e he hype pa am-
e e s o he Bayesian model and is he pa ame e
o he Gaussian dis ibu ion. Di e en choices o he alues
o hype pa ame e s lead o di e en e o es ima o s, bu we
will concen a e on a speci ic choice shown o be success ul in
ea lie wo ks9,22: κ = P + 2, ν = 0.5, s = I, and m = 0. Fo he
esul ing simpli ied closed- o m solu ion, e e o Re .9 Ma lab
and Py hon implemen a ions o BEE a e a ailable o down-
load (h ps://si es.google.com/si e/bayesiane o es ima e/).
Model selec ion. Model selec ion is a c i ical aspec in
classi ie design. Mo eo e , mos mode n classi ie s a e uned
by a se o hype pa ame e s, whose selec ion has a subs an ial
e ec on he esul ing accu acy as well. Thus, he selec ion
o an app op ia e model amily and he associa ed hype pa a-
me e s equi es an accu a e measu e o compa ing he accu-
acies o he model candida es. In ou wo k, we a e p ima ily
in e es ed in he selec ion o he egula iza ion pa ame e λ o
an LR classi ie . Howe e , i is o be no ed ha he me hodol-
ogy applies o any linea classi ie .
The p edic ion accu acy and selec ion o he bes model
can be quan i ied by e o es ima o s. CV es ima o is o en
used o selec he bes alue o he model selec ion pa ame e
λ along a egula iza ion pa h. As an example, e o cu es o
di e en alues o λ a e illus a ed in Figu e 1. Fo his pu pose,
we used he low cy ome y aining da a o 49 ea u es and 179
obse a ions. Fo an indi idual ube, each ea u e ep esen s he
a e age o he bioma ke in ensi ies. The e o cu es a e es i-
ma ed o di e en alues o λ anging om 10−9 o 100.
In he example o Figu e 1 (le panel), a 5- old CV p o-
cedu e is epea ed 100 imes and each s ep includes i e ain-
ing i e a ions on pa ial da a. The e o cu es ob ained o
100 i e a ions o 5- old CV illus a e he signi ican de ia ion
Flow cy ome y-based classi ica ion in cance esea ch
79CanCe In o ma ICs 2015:14(s5)
o he egula iza ion pa hs om one i e a ion o ano he . The
de ia ion is due o he andomness in spli ing he aining
da a in o olds, which esul s in an indi idual e o es ima e
o each spli . Mo eo e , o a e y small numbe o samples,
such as 5 o 10, he spli o alida ion and aining se s o he
K- old CV es ima o may no be app op ia e. In ac , in his
expe imen , he K- old CV app oach ails o es ima e he e o s
o smalle λ, as he numbe o samples spli by CV is insu i-
cien . On he o he hand, Figu e 1 ( igh panel) illus a es he
e o es ima e o BEE, which is a single de e minis ic e o
cu e. I is o be no ed ha he cu e ecognizes model o e -
i ing (e o es ima e s a s o inc ease o small egula iza-
ion e ms λ), al hough he e o is es ima ed di ec ly om he
aining se . No spli ing o i e a i e esampling is equi ed,
which in u n accele a es he compu a ion.
expe imen al esul s
In he ollowing sec ion, we p esen he expe imen al esul s.
Fi s , we demons a e di e en model selec ion c i e ia o es i-
ma e he signi ican ea u es. Then, we assess he pe o mances
o hose me hods in he AML classi ica ion case. Finally, we
p esen he esul s o he second, small sample case.
compa ison o model selec ion c i e ia. Typical
app oach o he selec ion o model pa ame e is CV.13 In his
pape , we also conside Bayesian e o es ima o (see Bayesian
E o Es ima o sec ion) and BIC24 as al e na i e app oaches
o es ima e he egula ized pa ame e . In o de o s udy he
beha io o di e en pa ame e selec ion c i e ia, we i s ain
he LR classi ie wi h he aining da a along he dec easing
sequence o egula iza ion pa h wi h log10 (λ)∈{0,–0.05,–0.1,
–0.15,…, –8.90, –8.95, –9.00}. Then, again he whole aining
da a a e used o es ima e he e o a e o each λ. Finally, o
each es ima o , we selec he model wi h λ alue ha achie es
he minimum e o a e. As esampling in CV-10 in oduces
andomness, in his case, we i e a e 200 imes and he esul
is a e aged. The de e minis ic na u e in BEE and BIC will
p oduce he same esul on he aining da a a each i e a ion.
The esul s a e summa ized in Table 3. Fo all me hods,
minimum e o a es, AUC, and he numbe o selec ed ea-
u es a e es ima ed om he whole aining da a. I is o be
no ed ha he epo ed AUC is compu ed om he aining
o emphasize ha all ea u e se s a e enough o pa i ion he
ea u e space in o wo ca ego ies pe ec ly. The es e o is
epo ed la e .
The esul s indica e ha he numbe o ea u es selec ed
by BEE me hod is lowe compa ed o hose o CV and BIC.
Fo he i s ea u e ec o wi h size 49, BEE selec s only 14
ea u es as signi ican , while o he second ea u e ec o wi h
−6 −5 −4 −3 −2 −1 0
0
0.02
0.04
0.06
0.08
0.1
0.12
0.14
0.16
0.18
0.2
Log10 (λ)
5− old c oss alida ion e o
−6 −5 −4 −3 −2 −1 0
0
0.02
0.04
0.06
0.08
0.1
0.12
0.14
0.16
0.18
0.2
Log10 (λ)
BEE
Figu e 1. Le : Examples o egula iza ion pa h e o cu es o 5- old CV o ou low cy ome y da a wi h heal hy and AML posi i e classes. Righ : The
co esponding Bee cu e.
Table 3. Pa ame e selec ion by di e en es ima o s: a e age numbe o selec ed ea u es, λ, aUC, and hei s anda d de ia ions wi h aining
da a.
METhoD FEATURE TYPE NUMBER oF SELECTED
FEATURES
SELECTED
Log10 (λ)
AUC
(TRAININg)
CV-10 mean 19.72 ± 2.41 −2.95 ± 2.30 0.9997 ± 0.0017
CV-10 mean and s d 23.91 ± 0.80 −4.14 ± 3.00 1 ± 0.0000
Bee mean 15 ± 0.00 −2.05 ± 0.00 0.9989 ± 0.0000
Bee mean and s d 13 ± 0.00 −1.80 ± 0.00 0.9992 ± 0.0000
BIC mean 20 ± 0.00 −5.85 ± 0.00 1 ± 0.0000
BIC mean and s d 24 ± 0.00 −5.70 ± 0.00 1 ± 0.0000

Hassan e al
80 CanCe In o ma ICs 2015:14(s5)
leng h 98, BEE selec s only 12 ea u es. Tables 4 and 5 lis
he selec ed ea u es, ie, signi ican bioma ke s along wi h he
co esponding coe icien alues. Due o he andomness in
CV-10, we only p esen he esul s o one i e a ion as an illus-
a ion: he e is a signi ican a ia ion o he selec ed ea u es
depending on he chosen CV spli . Howe e , i is o be no ed
ha he coe icien s o BIC and BEE a e no speci ic o his
pa icula i e a ion, as hey do no include he andom spli .
Pe o mance assessmen o he model selec ion c i e ia.
The pe o mances o he model selec ion me hods a e s udied
in he ollowing sec ion. The classi ica ion e o is conside ed
as he pe o mance c i e ion, and bo h alse posi i es (heal hy
con ol classi ied as AML) and alse nega i es (AML classi ied
as heal hy con ol) a e coun ed wi h equal weigh . The pe o -
mance o he Bayesian e o es ima o is benchma ked agains
hose o CV-10 and BIC o a di e en numbe o sample
sizes. Fo his pu pose, a andomly selec ed p opo ion o 10%,
15%, 20%–90%, and 95% is selec ed o aining he classi ie ,
while he emaining da a a e used o pe o mance assessmen .
Fo each aining sample, he expe imen is execu ed 200 imes
by gene a ing a new aining se each ime.
Classi ica ion e o s. The e o cu es o di e en sample
sizes a e shown in Figu e 2. The p ocedu e is epea ed 200
imes o each aining sample size, and he a e age o clas-
si ica ion e o is compu ed o each model selec ion c i e ion.
Wi h a e y small numbe o aining samples, such as 10%
o 15% o he da ase , BEE p o ides imp o ed accu acy o e
CV-10 and BIC (Fig. 2 le and igh panels). Fo ins ance,
wi h 10% aining samples, he classi ica ion e o s o he
model selec ed by CV-10 a e 7.56% (Fig. 2 le panel) and
7.41% (Fig. 2 igh panel) highe han hose o BEE. In case
o BIC, he classi ica ion e o is 7.81% highe han ha o
BEE (Fig. 2 le panel). As he numbe o aining samples
inc eases, o example, abo e 60%, he pe o mance o BIC
exceeds han ha o BEE (Fig. 2 le panel). Howe e , he
pe o mance o BIC is simila o ha o BEE when mo e ea-
u es a e in ol ed in he expe imen (Fig. 2 igh panel).
Table 4. The nonze o coe icien s o ea u es wi h mean.
TUBE FEATURE 10-FoLD CV BEE BIC
Cons an −13.38 −3.50 −13.54
ube 1 sLin 0.75 00.78
ube 1 ssLog −5.92 −0.73 −6.01
ube 1 L1:IgG1- I C −0.46 −0.19 −0.46
ube 1 L4:IgG1-PC5 −2.08 0−2.13
ube 1 L5:IgG1-PC7 −3.07 −0.19 −3.14
ube 2 sLin 0.001 0 0
ube 2 L5:CD20-PC7 2.76 02.78
ube 3 ssLog 0−0.77 0
ube 3 L4:CD8-PC5 −1.94 −0.10 −1.97
ube 4 sLin 0.97 00.97
ube 4 L1:CD15 - I C −4.77 0−4.82
ube 4 L2:CD13-Pe 3.44 0.21 3.48
ube 4 L4:CD16-PC5 0−0.09 0
ube 4 L5:CD56-PC7 3.29 0.75 3.35
ube 5 sLin 2.35 02.40
ube 5 L2:CD11c-Pe −0.15 0−0.16
ube 5 L3:CD45-eCD −1.84 −0.02 −1.84
ube 5 L4:CD64-PC5 1.66 01.69
ube 5 L5:CD33-PC7 0.75 0.60 0.76
ube 6 L2:CD117-Pe 4.82 0.89 4.88
ube 6 L4:CD34-PC5 6.88 0.72 6.99
ube 6 L5:CD38-PC7 00.41 0
ube 7 L1:CD5 - I C −4.68 −0.19 −4.74
No e: The size o he ea u e se s is 49, o which he CV, BEE, and BIC selec
20, 14, and 19 ea u es, espec i ely.
Table 5. The nonze o coe icien s o ea u es wi h mean and
s anda d de ia ion.
TUBE FEATURE 10-FoLD CV BEE BIC
Cons an −13.47 −3.27 −13.47
ube 1 sLin mean 0.027 00.027
ube 1 ssLog mean −4.49 −0.47 −4.49
ube 1 L1:IgG1- I C s d −3.80 −0.22 −3.80
ube 1 L5:IgG1-PC7 s d −1.60 −0.16 −1.60
ube 2 L5:CD20-PC7 mean 0.50 00.50
ube 3 ssLog mean 0−0.29 0
ube 3 L5:CD2-PC7 mean −0.48 0−0.48
ube 3 L5:CD2-PC7 s d −1.21 0−1.21
ube 4 sLin mean 0.05 00.05
ube 4 L1:CD15 - I C mean −0.14 0−0.14
ube 4 L2:CD13-Pe mean 2.50 02.50
ube 4 L4:CD16-PC5 mean −1.32 0−1.32
ube 4 L4:CD16-PC5 s d 0−0.39 0
ube 4 L5:CD56-PC7 s d 6.07 0.45 6.07
ube 5 L1:CD14- I C s d −0.004 0−0.004
ube 5 L3:CD45-eCD mean −2.51 0−2.51
ube 5 L5:CD33-PC7 mean 1.47 0.32 1.47
ube 5 L5:CD33-PC7 s d −0.70 0−0.70
ube 6 ssLog s d 0−0.23 0
ube 6 L2:CD117-Pe mean 3.19 0.46 3.19
ube 6 L2:CD117-Pe s d 2.19 02.19
ube 6 L4:CD34-PC5 s d 1.88 0.75 1.88
ube 6 L5:CD38-PC7 mean 1.31 0.44 1.31
ube 7 sLin mean 1.39 01.39
ube 7 sLin s d 0−0.16 0
ube 7 L1:CD5 - I C s d −1.16 0−1.16
ube 7 L5:CD10-PC7 s d 0.04 00.04
No e: The size o he ea u e se s is 98, o which he CV, BEE, and BIC selec
23, 12, and 23 ea u es, espec i ely.
Flow cy ome y-based classi ica ion in cance esea ch
81CanCe In o ma ICs 2015:14(s5)
A ea unde he ROC cu e. In his sec ion, we e alua e he
pe o mance in e ms o AUC. Figu e 3 illus a es he a e age o
AUC o di e en aining sample sizes. He e, he BEE me hod
achie es imp o emen o e he o he me hods. Wi h small ain-
ing samples, o ins ance, 10%, he a e age AUC o BEE is 1.11%
(Fig. 3 le panel) and 1.30% (Fig. 3 igh panel) highe han ha
o CV-10. As he numbe o aining samples inc eases, CV-10
and BIC also con e ge owa d he esul s o BEE; howe e , he
BEE selec ed model consis en ly esul s in he highes AUC
sco e. Wi h he la ge ea u e ec o ha includes he measu e-
men s o mean and s anda d de ia ion, he a e age AUC cu es
o BEE and BIC ollow he simila pa e n (Fig. 3 igh panel).
Numbe o selec ed ea u es. We u he assess he pe -
o mances o he es ima o s using ea u e selec ion c i e ia.
A each i e a ion, we de e mine he o al numbe o selec ed
ea u es ha ha e nonze o alues o a di e en numbe o
aining samples. Then, we compu e he a e age and he
a iabili y (ie, s anda d de ia ion) o he selec ed ea u es
o di e en aining samples. The esul s a e illus a ed in
Figu e 4. Fo BEE, he a e age numbe o selec ed ea u es
is lowe in amoun compa ed o hose o CV and BIC (Fig. 4
op panel). Fo ins ance, wi h 95% aining samples, BEE
equi es 36.49% and 33.89% less ea u es han CV and BIC,
espec i ely, o model p edic ion (Fig. 4 op- igh panel).
Mo eo e , he a iabili y in selec ed ea u es using BEE is
also compa able (Fig. 4 bo om panel). The CV-10 has he
wo s pe o mance. Al hough BIC shows ha he de ia-
ion in ea u e selec ion a di e en i e a ions is smalle , he
0 20 40 60 80 100 120 140 160 180
0.02
0.03
0.04
0.05
0.06
0.07
Numbe o samples used o aining
Classi ica ion e o s
CV-10
BEEp
BIC
0 20 40 60 80 100 120 140 160 180
0.02
0.03
0.04
0.05
0.06
0.07
Numbe o samples used o aining
Classi ica ion e o s
CV-10
BEEp
BIC
Figu e 2. The a e age classi ica ion e o cu es o CV-10, BEE wi h p ope p io (BEEp), and BIC.
No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s.
0 20 40 60 80 100 120 140 160 180
0.9
0.91
0.92
0.93
0.94
0.95
0.96
0.97
0.98
0.99
1
Numbe o samples used o aining
AUC
CV-10
BEEp
BIC
0 20 40 60 80 100 120 140 160 180
0.9
0.91
0.92
0.93
0.94
0.95
0.96
0.97
0.98
0.99
1
Numbe o samples used o aining
AUC
CV-10
BEEp
BIC
Figu e 3. The a e age AUCs o CV-10, BEE wi h p ope p io (BEEp), and BIC.
No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s.
Hassan e al
82 CanCe In o ma ICs 2015:14(s5)
a e age numbe o selec ed ea u es is highe han ha o
o he s (Fig. 4 op panel).
Simila i y o he selec ed ea u e se s. Ano he pe o mance
measu emen is he s abili y o selec ing he same ea u e a
di e en i e a ions. Fo his pu pose, Sø ensen–Dice coe -
icien 25 is used, which measu es he deg ee o simila i y
be ween selec ed ea u es o wo di e en i e a ions. The
anges can a y om 0 o 1. The alues closes o 1 indica e a
high-deg ee o simila i y.
Fo di e en aining samples, we i s de e mine
which ea u es a e selec ed a each i e a ion. As he model
selec ion p ocess is epea ed 200 imes, we es ima e he
simila i y as he mean dice coe icien o each o he 200!/
(2! × (200 − 2)!) = 19,900 possible pai s o selec ed ea u e
se s. The esul s a e shown in Figu e 5. In e ms o s abil-
i y, he pe o mance o BEE is subs an ially be e han hose
o he o he me hods, as he selec ed ea u e se s a e mos
simila wi h ha c i e ion – a signi ican issue when ying
o unde s and he biological mechanisms behind he da a.
Fo example, wi h 60% aining samples, he dice coe icien
o BEE is 6.03% highe han ha o CV (Fig. 5 le panel).
On he o he hand, wi h 90% aining samples, he dice coe -
icien o BEE is 5.81% highe han ha o CV and 4.90%
highe han ha o BIC (Fig. 5 igh panel). Indeed, he dice
coe icien is un a o able o CV wi h small aining samples:
The dice coe icien is lowes among he al e na i es, indica -
ing ha he selec ed ea u e se s wi h he CV c i e ion ha e
high a iabili y.
small sample case wi h a cance cell line. Fo u he
con idence on he p esen ed me hod, we analyze da a om a
cance cell line in a small sample se ing. As desc ibed p e i-
ously, we conside ed he classi ica ion accu acy, AUC mea-
su e, and he numbe o selec ed a iables bo h wi h and
wi hou s anda d de ia ion ea u es (Figs. 6–8). In his case,
0 20 40 60 80 100 120 140 160 180
8
10
12
14
16
18
20
22
24
Numbe o samples used o aining
A e age numbe o ea u es
CV-10
BEEp
BIC
0 20 40 60 80 100 120 140 160 180
8
10
12
14
16
18
20
22
24
Numbe o samples used o aining
A e age numbe o ea u es
CV-10
BEEp
BIC
0 20 40 60 80 100 120 140 160 180
1
1.5
2
2.5
3
3.5
4
Numbe o samples used o aining
S anda d de ia ion o numbe o ea u es
CV-10
BEEp
BIC
0 20 40 60 80 100 120 140 160 180
1.5
2
2.5
3
3.5
4
4.5
Numbe o samples used o aining
S anda d de ia ion o numbe o ea u es
CV-10
BEEp
BIC
Figu e 4. Compa isons o he numbe o selec ed ea u es o CV-10, BEE wi h p ope p io (BEEp), and BIC.
No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s. Top: a e age
numbe o selec ed ea u es. Bo om: s anda d de ia ion o numbe o he selec ed ea u es.
Flow cy ome y-based classi ica ion in cance esea ch
83CanCe In o ma ICs 2015:14(s5)
0 20 40 60 80 100 120 140 160 180
0.3
0.4
0.5
0.6
0.7
0.8
0.9
1
Numbe o samples used o aining
Dice index
CV-10
BEEp
BIC
0 20 40 60 80 100 120 140 160 180
0.3
0.4
0.5
0.6
0.7
0.8
0.9
1
Numbe o samples used o aining
Dice index
CV-10
BEEp
BIC
Figu e 5. Compa ison o he s abili y o selec ing ea u es o CV-10, BEEp, and BIC.
No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s.
2 4 6 8 10 12 14
0.2
0.25
0.3
0.35
0.4
0.45
Numbe o samples used o aining
Classi ica ion e o s
LOOCV
BEEp
BIC
2 4 6 8 10 12 14
0.2
0.25
0.3
0.35
0.4
0.45
Numbe o samples used o aining
Classi ica ion e o s
LOOCV
BEEp
BIC
Figu e 6. The a e age classi ica ion e o cu es o LOOCV, BEEp, and BIC.
No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s.
2 4 6 8 10 12 14
0.55
0.6
0.65
0.7
0.75
0.8
0.85
0.9
Numbe o samples used o aining
AUC
LOOCV
BEEp
BIC
2 4 6 8 10 12 14
0.55
0.6
0.65
0.7
0.75
0.8
0.85
0.9
0.95
Numbe o samples used o aining
AUC
LOOCV
BEEp
BIC
Figu e 7. The a e age AUC cu es o LOOCV, BEEp, and BIC.
No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s.