scieee Open visual document viewer

Flow Cytometry-Based Classification in Cancer Research: A View on Feature Selection

Hassan, Sakira,Ruusuvuori, Pekka,Latonen, Leena,Huttunen, Heikki

Abstract

In this paper, we study the problem of feature selection in cancer-related machine learning tasks. In particular, we study the accuracy and stability of different feature selection approaches within simplistic machine learning pipelines. Earlier studies have shown that for certain cases, the accuracy of detection can easily reach 100% given enough training data. Here, however, we concentrate on simplifying the classification models with and seek for feature selection approaches that are reliable even with extremely small sample sizes. We show that as much as 50% of features can be discarded without compromising the prediction accuracy. Moreover, we study the model selection problem among the ℓ1 regularization path of logistic regression classifiers. To this aim, we compare a more traditional cross-validation approach with a recently proposed Bayesian error estimator.

Full text

75CanCe In o ma ICs 2015:14(s5) In oduc ion Flow cy ome y enables quan i a i e measu emen o single- cell p ope ies h ough isible and luo escen ligh in a high- h oughpu manne . The measu ed signals include luo escence emission and ligh sca e . Flow cy ome y has been ou inely used o de ec ing malignancies om blood samples.1 Th ough echnological ad ances, measu ing he combina ion o luo- escen signals om se e al di e en channels has enabled he use o high-dimensional da a o s udies, such as cy ome ic inge p in ing2 and la ge-scale analysis o cell ypes.3 In his pape , we s udy he analysis o low cy ome y da a om he ea u e selec ion poin o iew. Mo e speci i- cally, low cy ome y is able o p oduce la ge quan i ies o pa ially edundan measu emen da a, and he selec ion o impo an quan i ies wi hin he la ge body o measu emen s is o in e es . Mo eo e , a ypical scena io con ains la ge quan- i ies o measu emen da a bu may be limi ed o only a ew pa ien s. Thus, an ideal me hod would dis il only he essen ial pa s o he measu emen s om each pa ien , while p oduc- ing eliable and well gene alizing esul s when only a small amoun o indi iduals is a ailable in he aining da a. We will concen a e ou a en ion on wo pa icula se s o low cy ome y da a. The i s se o igina es om he acu e myeloid leukemia (AML) p edic ion challenge o he DREAM ini ia i e in 2013.4 The compe i ion a ac ed a numbe o eams, and as se e al esea che s use he da a as pa o hei wo k, he challenge da a ha e become a s anda d benchma k wi hin he ield. Fo example, Aghaeepou e al.4 p esen s a la ge pool o analysis app oaches om he DREAM challenge. Among classi ica ion me hods p esen ed in he li e a u e, he e a e se e al sophis ica ed machine lea ning app oaches, such as lea ning ec o quan iza ion,5 co ela i e ma ix map- ping, and ela i e en opy di e ences.6 The s eng h o da a- d i en app oaches elying on supe ised classi ica ion is hei abili y o handle high-dimensional da a wi hou equi ing p io knowledge o he biological applica ion. The DREAM AML da a ep esen a ela i ely la ge-scale expe imen consis ing o al oge he 179 pa ien s. Al hough i is a small numbe o adi ional machine lea ning p oblems, he numbe o pa ien s is unusually la ge o a biological s udy. To his aim, we use ano he da ase ha ep esen s a mo e commonly encoun e ed sample size o 16 samples ex ac ed om a p os a e cance cell line. This da ase , wi h wo di - e en ea men s and a low numbe o samples, p esen s a non i ial bu common challenge o p edic ion and ela ed ea u e selec ion. Mo e in o ma ion on he wo da ase s is p o ided in Da a sec ion. In ou ea lie wo k,7 we p esen ed a supe ised classi ica- ion pipeline based on linea disc iminan analysis (LDA) and logis ic eg ession (LR) classi ie s. B ie ly, he me hod i s Flow Cy ome y-Based Classi ica ion in Cance Resea ch: A View on Fea u e Selec ion s. saki a Hassan1, Pekka uusu uo i2,3, Leena La onen3 and Heikki Hu unen1 1Depa men o Signal P ocessing, Tampe e Uni e si y o Technology, Tampe e, Finland. 2Po i Depa men , Tampe e Uni e si y o Technology, Po i, Finland. 3BioMediTech, Uni e si y o Tampe e, Tampe e, Finland. Supplemen a y Issue: S a is ical Sys ems Theo y in Cance Modeling, Diagnosis, and The apy Abs Ac : In his pape , we s udy he p oblem o ea u e selec ion in cance - ela ed machine lea ning asks. In pa icula , we s udy he accu acy and s abili y o di e en ea u e selec ion app oaches wi hin simplis ic machine lea ning pipelines. Ea lie s udies ha e shown ha o ce ain cases, he accu acy o de ec ion can easily each 100% gi en enough aining da a. He e, howe e , we concen a e on simpli ying he classi ica ion models wi h and seek o ea u e selec ion app oaches ha a e eliable e en wi h ex emely small sample sizes. We show ha as much as 50% o ea u es can be disca ded wi hou comp omising he p edic ion accu acy. Mo eo e , we s udy he model selec ion p oblem among he  1 egula iza ion pa h o logis ic eg ession classi ie s. To his aim, we compa e a mo e adi ional c oss- alida ion app oach wi h a ecen ly p oposed Bayesian e o es ima o . Keywo ds: AML, leukemia, low cy ome y, logis ic eg ession, e o es ima ion, model selec ion SUPPLEMENT: s a is ical sys ems heo y in Cance modeling, Diagnosis, and he apy CITATIoN: Hassan e al. Flow Cy ome y-Based Classi ica ion in Cance Resea ch: A View on ea u e selec ion. Cance In o ma ics 2015:14(s5) 75–85 doi: 10.4137/CIn.s30795. TYPE: o iginal esea ch RECEIVED: no embe 18, 2015. RESUBMITTED: eb ua y 01, 2016. ACCEPTED FoR PUBLICATIoN: eb ua y 07, 2016. ACADEMIC EDIToR: J. T. E i d, Edi o in Chie PEER REVIEw: six pee e iewe s con ibu ed o he pee e iew epo . e iewe s’ epo s o aled 1860 wo ds, excluding any con iden ial commen s o he academic edi o . FUNDINg: Au ho s disclose no ex e nal unding sou ces. CoMPETINg INTERESTS: Au ho s disclose no po en ial con lic s o in e es . CoRRESPoNDENCE: [email p o ec ed] CoPYRIghT: © he au ho s, publishe and licensee Libe as academica Limi ed. his is an open-access a icle dis ibu ed unde he e ms o he C ea i e Commons CC-BY-NC 3.0 License. Pape subjec o independen expe blind pee e iew. all edi o ial decisions made by independen academic edi o . Upon submission manusc ip was subjec o an i- plagia ism scanning. P io o publica ion all au ho s ha e gi en signed con i ma ion o ag eemen o a icle publica ion and compliance wi h all applicable e hical and legal equi emen s, including he accu acy o au ho and con ibu o in o ma ion, disclosu e o compe ing in e es s and unding sou ces, compliance wi h e hical equi emen s ela ing o human and animal s udy pa icipan s, and compliance wi h any copy igh equi emen s o hi d pa ies. This jou nal is a membe o he Commi ee on Publica ion E hics (COPE). Published by Libe as academica. Lea n mo e abou his jou nal. Hassan e al 76 CanCe In o ma ICs 2015:14(s5) ans o ms he measu emen da a in o highe dimensional space by gene a ing combined ea u es wi h mul iplica ions and di isions be ween measu emen s. Following his mapping in o highe dimensional space, LDA is used o lowe ing he dimen- sionali y in o a single alue pe measu emen . Then, empi ical dis ibu ion unc ions (EDFs) a e cons uc ed om LDA esul s o AML-posi i e and AML-nega i e sample classes and com- pa ed o aining EDFs o bo h classes. The compa ison esul s in wo simila i y alues pe g oup o measu emen s, and hese esul s a e ed o he LR classi ie o a inal AML p edic ion esul . Ou app oach, oge he wi h al e na i e well-pe o ming app oaches,4,5 ep esen s a ela i ely complica ed pipeline o somewha a bi a y compu a ion s eps. Thus, ou in e es is o simpli y hese pipelines in o a simple collec ion o ob ious ea- u es, while s ill e aining a good accu acy. Ou app oach he e is o use LR classi ie applied o sum- ma ize ea u es, which a e he mean and s anda d de ia ion o he measu emen s ins ead o he comple e da a. This educes he numbe o ea u es used in classi ica ion and, subsequen ly, also he model complexi y. An essen ial pa o classi ie design is e o es ima ion, which guides model selec ion.8 Ou s a egy o model selec ion is o apply he ecen ly in oduced Bayesian e o es ima o (BEE).9 We compa e BEE model selec ion wi h a adi ional 10- old c oss- alida ion (CV-10) e o es ima ion, as well as wi h Bayesian in o ma ion c i e- ion (BIC)-based model selec ion, and conclude ha he p o- posed app oach enables accu a e p edic ion o low cy ome y da a wi h ewe measu emen s and a less complex classi ie model han hose p e iously p esen ed in he li e a u e. The es o his pape is o ganized as ollows. In Ma e ials and Me hods sec ion, we desc ibe he da a and me hods used in his s udy and b ie ly discuss how ea u e selec ion is com- monly done in machine lea ning. Expe imen al Resul s sec ion p esen s he esul s o ou expe imen s wi h di e en model and ea u e selec ion c i e ia o he ma e ials in oduced in Ma e ials and Me hods sec ion. Finally, in Conclusions sec- ion, we summa ize he wo k and discuss he conclusions o he esul s. Ma e ials and Me hods In his sec ion, we desc ibe he da ase s used in his pape . We also gi e a b ie o e iew o modeling me hod o ea u e selec ion. In addi ion o his, we in oduce he s a e-o - he-a Bayesian e o es ima o (BEE) o model pa ame e selec- ion. Finally, we p esen an example whe e he pe o mance o BEE is benchma ked agains o he model selec ion c i e ia. da a. In his wo k, we s udy wo da ase s: A la ge se wi h 179 samples and a smalle se wi h 16 samples. These wo case s udies ep esen di e en classi ica ion challenges in e ms o bo h applica ion and sample size. AML da ase . The low cy ome y da ase o he AML expe imen has been collec ed om he DREAM6- FlowCAP2 challenge, which was o ganized by he DREAM p ojec and he FlowCAP ini ia i e (DREAM challenge AML da ase can be accessed om Aghaeepou e al).4 We use he aining da ase ha consis s o low cy ome y measu emen s o 179 pa ien s. Among hem, 23 pa ien s a e AML posi i e and he emaining 156 pa ien s a e AML nega i e. The low cy ome y measu emen o each pa ien co esponds o se en join ly measu ed g oups (he ea e called ubes) o se en quan i ies wi h a o al o 49 bioma ke measu emen s pe cell. The bioma ke s a e summa ized in Table 1 and include Fo wa d Sca e in linea scale (FS Lin), Sidewa d Sca e in loga i hmic scale (SS Log), and i e luo- escence in ensi ies (FL1–FL5) in loga i hmic scales. Fo calib a ion pu poses, FS Lin, SS Log, and CD45-ECD we e measu ed o all ubes and he o he 28 bioma ke s we e measu ed only in one ube. Cance cell line da ase . As ano he case s udy, we use low cy ome y da a om a small sample se ing. The da a come om a p os a e cance cell line 22R 1 s ained wi h p opi dium iodide o cell cycle analysis.10 The cells a e ans ec ed wi h miRNAs (ei he con ol o miR-193b) and induced o p oli e a e by o e exp ession o cyclin D. The da a consis o 16 samples, wi h 8 samples (wi hou cyclin D o e exp ession) wi h ela i ely consis en cell cycle p o ile and 8 samples (o e exp essing cyclin D) wi h an al e ed cell cycle p o ile, ie, induced cell cycle ac i i y wi h an inc ease in cells in DNA syn hesis phase. The samples o bo h classes include ou epe i ions o wo ea men s, which a e consid- e ed he e o ep esen he same class. Each sample con ains 12 measu ed channels, consis ing o wo sca e measu e- men s and ou luo escence channels, bo h as a ea and heigh measu emen s. Table 1. Lis o se en ubes wi h bioma ke s p o ided in DREAM6 AML p edic ion da a. FL1 Log FL2 Log FL3 Log FL4 Log FL5 Log ube 1 s Lin ss Log IgG1- I C IgG1-Pe CD45-eCD IgG1-PC5 IgG1-PC7 ube 2 s Lin ss Log Kappa- I Lambda-Pe CD45-eCD CD19-PC5 CD20-PC7 ube 3 s Lin ss Log CD7- I C CD4-Pe CD45-eCD CD8-PC5 CD2-PC7 ube 4 s Lin ss Log CD15- I C CD13-Pe CD45-eCD CD16-PC5 CD56-PC7 ube 5 s Lin ss Log CD14- I C CD11c-Pe CD45-eCD CD64-PC5 CD33-PC7 ube 6 s Lin ss Log HLa-D - I C CD117-Pe CD45-eCD CD34-PC5 CD38-PC7 ube 7 s Lin ss Log CD5- I C CD19-Pe CD45-eCD CD3-PC5 CD10-PC7 Flow cy ome y-based classi ica ion in cance esea ch 77CanCe In o ma ICs 2015:14(s5) Fea u e ex ac ion. Se e al ea u e ex ac ion me hods can be used o ob ain meaning ul ea u es om aw low cy ome y measu emen s. Fo ins ance, among widely used ea u e ex ac ion echniques a e me hods based on p incipal componen analysis and his og am compu a- ion. Biehl e al p oposed s a is ical di e gences o ex ac ea u es ha include momen s, median, and in e qua ile ange.5 The leng h o he ea u e ec o was 186 in his case. Ano he well-pe o med model was based on mul idimen- sional en opic dis ance-based ea u es.4,5 Manninen e al.7 expanded he cell measu emen s o each ube o a highe dimensional space. Following his ans o ma ion, LDA is used o lowe he dimensionali y in o a single alue o each measu emen . These p e ious s udies a e summa ized in Table 2. Table 2 also abula es he es accu acy in e ms o he a ea unde he ecei e ope a ing cha ac e is ics (ROC) cu e (AUC) measu e o e a single ain/ es spli , which should no be in e p e ed as a de ini i e measu e o accu acy, as he spli o he samples is jus one ins ance o all possible spli s. In his pape , we use one o he simples ea u e ex ac ion echniques ha include only he mean and he s anda d de ia- ion o he each measu emen . Fo he i s da ase , he leng h o his ex ac ed ea u e ec o is 98, comp ising 49 mean alues and 49 s anda d de ia ions. As seen in he expe imen s o Expe imen al Resul s sec ion, hese ea u es a e su icien o sepa a e he classes wi hou comp omising he p edic ion accu acy. We will conside wo e sions o hese basic ea- u es: he i s ea u e se con ains only he 49 mean alues o he measu emen s, while he second ea u e ec o conside s bo h mean alues and s anda d de ia ions, wi h al oge he 98 ea u es. The same app oach is used wi h he smalle da a- se , hus p oducing wo di e en expe imen al cases. Be o e aining he classi ie s, we no malized all ea u es o he in e - al (0, 1). L and egula iza ion. LR is a disc imina i e me hod o modeling he class condi ional p obabili y densi ies by he logis ic unc ion. Gi en an obse a ion ma ix wi h N obse a ions, P ea u es, and co espond- ing class labels y ∈1,…,C, we de ine LR model o he bina y classi ica ion as, (1) and (2) He e, x ep esen s a ea u e ec o in he ea u e space co esponding o class label y ∈ {0,1}, β0 is he in e cep , and β ep esen s coe icien s o he logis ic model. We can de e mine he model pa ame e s β0 and β om he aining da ase by sol ing he  1 penalized LR p oblem, (3) whe e λ . 0 is he egula iza ion pa ame e . When he num- be o aining da a is no la ge compa ed o he numbe o ea u es, ie, P  N, egula iza ion is used o sol e he o e i - ing p oblem.11 In egula iza ion, an ex a e m, λ, is added, which con ols he ade-o be ween he loss unc ion and he size o he coe icien s. Mo e ecen ly, in ea u e selec- ion,  1- egula ized LR has ecei ed much a en ion, as i yields a spa se solu ion ha has ela i ely ew nonze o coe i- cien s.12 This minimiza ion ask is analogous o leas absolu e sh inkage and selec ion ope a o (Lasso) algo i hm p oposed by Tibshi ani.13 In addi ion o his, se e al ex ensions o Lasso ha e also been de eloped, such as g ouped Lasso,14,15 Dan zig selec o ,16 elas ic ne ,17 and g aphical Lasso.18 In his pape , Table 2. S udies based on ea u e ex ac ion s a egies o DREAM amL challenge da ase . ACCURACY SIzE oF FEATURE VECToR BRIEF DESCRIPTIoN Biehl e al.51.00 186 Ex ac ion o ea u es wi h momen s, median and in e qua ile and lea ning ec o quan iza ion is used o p edic ion Vila e al.41.00 31 Ex ac ion o ea u es wi h en opies and his og am based classi ie is used o p edic ion manninen e al.7 1.00 (# o e en s) x 84 Expand ea u es o highe dimension and hen mapping o 1-D using linea disc iminan analysis; logis ic eg ession is used o p edic ion ou solu ion ( his s udy) 0.9989 49 Ex ac ion o ea u e ec o om means o measu emen s and applying egula ized logis ic eg ession o p edic ion ou solu ion ( his s udy) 0.9992 98 Ex ac ion o ea u e ec o om means and s anda d de ia ion o measu emen s and applying egula ized logis ic eg ession o p edic ion No e: The accu acy is measu ed in e ms o he AUC o a single ain/ es spli . Hassan e al 78 CanCe In o ma ICs 2015:14(s5) we use he GLMNET algo i hm by F iedman e al.19 ha combines he  2 and  1 penal ies: (4 ) whe e λ . 0 and α ∈[0,1]. The pa ame e α is a comp omise be ween he  1 and  2 penal ies, he eby de e mining he ype o egula iza ion. On he o he hand, he egula iza ion pa am- e e λ con ols he amoun o egula iza ion. A e y la ge λ will comple ely sh inks he coe icien s o ze o and may yield a null o emp y model. In gene al, he model pa ame e s λ and α a e selec ed using he CV app oach.20 The da ase is andomly spli in o K mu u- ally exclusi e subse s o app oxima ely equal sized. In K- old CV, he p ocess is i e a ed k imes. A he k h i e a ion, he K h old is e ained as es se and he emaining K − 1 olds a e used as aining se o ain he model. Each o he K- olds is es ed exac ly once. The es se assesses he quali y o he ained model. Then, he K esul s a e combined o a e aged o p oduce a single es i- ma ion o he model. The mos commonly used alues o K a e 5 and 10. In his expe imen , we se he alue o α = 1 and CV-10 is used o he selec ion o he model pa ame e λ and assessmen o he model. As he ype o egula iza ion is de e mined by α, se ing α = 0 p o ides  2 penal y ha is use ul in cases, whe e he ea u es a e mu ually co ela ed. On he o he hand, α = 1 p o- ides spa se solu ion wi h ewe coe icien s and, in u n, his is sui able o implici ea u e selec ion. We ha e also expe imen ed wi h 5- old CV, bu he esul s do no imp o e signi ican ly. bayesian e o es ima o . A Bayesian app oach o e o es ima ion was ecen ly in oduced in he con ex o disc e e classi ie s21 and linea classi ie s.22 The Bayesian e o es ima- o (BEE) es ima es he classi ica ion e o di ec ly om he aining se and has shown o imp o e bo h he accu acy and speed o he ac ual e o es ima e21,22 compa ed o adi ional coun ing-based app oaches, such as CV. In ou ea lie pape s, we ha e shown ha BEE has imp o ed he s abili y and speed o compu a ion in he model selec ion con ex as well.9,23 We will nex b ie ly e iew he de ini ion o BEE o a ixed linea wo-class classi ie speci ied by he pa ame e s β and β0. The Bayesian e o es ima o o linea classi ica ion assumes ha he samples om each class a e independen and iden ically dis ibu ed Gaussian andom a iables. Fo he wo classes, he pa ame e s (mean and co a iance) o he Gaussian model a e deno ed as θ0 and θ1 and he co esponding p io s o he pa a me e s a e deno ed as p0(θ) and p1(θ). Then, he pos e io p obabili y densi y unc ions (PDFs) o pa ame e s o class c ∈{0,1} a e gi en by he Bayes’ ule: (5) whe e is he Gaussian class condi ional densi y o c ∈{0,1}. The Bayesian e o es ima o (BEE) is de ined as he minimum mean squa ed es ima o by minimizing he expec a- ion be ween he e o es ima e and he ue e o . This quan- i y is composed o class-speci ic condi ional expec ed e o s balanced by he p io s p(c) o he wo classes c ∈{0,1}22: (6) wi h he expec ed classi ica ion e o o samples om class c gi en by (7) whe e εc(θ) deno es he ue classi ica ion e o . The in eg al o Equa ion (7) can be e alua ed by assuming an in e se Wisha p io o he class condi ional densi y: (8) whe e υ ∈∈∈∈ × ,κ,Sm PP , and P a e he hype pa am- e e s o he Bayesian model and is he pa ame e o he Gaussian dis ibu ion. Di e en choices o he alues o hype pa ame e s lead o di e en e o es ima o s, bu we will concen a e on a speci ic choice shown o be success ul in ea lie wo ks9,22: κ = P + 2, ν = 0.5, s = I, and m = 0. Fo he esul ing simpli ied closed- o m solu ion, e e o Re .9 Ma lab and Py hon implemen a ions o BEE a e a ailable o down- load (h ps://si es.google.com/si e/bayesiane o es ima e/). Model selec ion. Model selec ion is a c i ical aspec in classi ie design. Mo eo e , mos mode n classi ie s a e uned by a se o hype pa ame e s, whose selec ion has a subs an ial e ec on he esul ing accu acy as well. Thus, he selec ion o an app op ia e model amily and he associa ed hype pa a- me e s equi es an accu a e measu e o compa ing he accu- acies o he model candida es. In ou wo k, we a e p ima ily in e es ed in he selec ion o he egula iza ion pa ame e λ o an LR classi ie . Howe e , i is o be no ed ha he me hodol- ogy applies o any linea classi ie . The p edic ion accu acy and selec ion o he bes model can be quan i ied by e o es ima o s. CV es ima o is o en used o selec he bes alue o he model selec ion pa ame e λ along a egula iza ion pa h. As an example, e o cu es o di e en alues o λ a e illus a ed in Figu e 1. Fo his pu pose, we used he low cy ome y aining da a o 49 ea u es and 179 obse a ions. Fo an indi idual ube, each ea u e ep esen s he a e age o he bioma ke in ensi ies. The e o cu es a e es i- ma ed o di e en alues o λ anging om 10−9 o 100. In he example o Figu e 1 (le panel), a 5- old CV p o- cedu e is epea ed 100 imes and each s ep includes i e ain- ing i e a ions on pa ial da a. The e o cu es ob ained o 100 i e a ions o 5- old CV illus a e he signi ican de ia ion Flow cy ome y-based classi ica ion in cance esea ch 79CanCe In o ma ICs 2015:14(s5) o he egula iza ion pa hs om one i e a ion o ano he . The de ia ion is due o he andomness in spli ing he aining da a in o olds, which esul s in an indi idual e o es ima e o each spli . Mo eo e , o a e y small numbe o samples, such as 5 o 10, he spli o alida ion and aining se s o he K- old CV es ima o may no be app op ia e. In ac , in his expe imen , he K- old CV app oach ails o es ima e he e o s o smalle λ, as he numbe o samples spli by CV is insu i- cien . On he o he hand, Figu e 1 ( igh panel) illus a es he e o es ima e o BEE, which is a single de e minis ic e o cu e. I is o be no ed ha he cu e ecognizes model o e - i ing (e o es ima e s a s o inc ease o small egula iza- ion e ms λ), al hough he e o is es ima ed di ec ly om he aining se . No spli ing o i e a i e esampling is equi ed, which in u n accele a es he compu a ion. expe imen al esul s In he ollowing sec ion, we p esen he expe imen al esul s. Fi s , we demons a e di e en model selec ion c i e ia o es i- ma e he signi ican ea u es. Then, we assess he pe o mances o hose me hods in he AML classi ica ion case. Finally, we p esen he esul s o he second, small sample case. compa ison o model selec ion c i e ia. Typical app oach o he selec ion o model pa ame e is CV.13 In his pape , we also conside Bayesian e o es ima o (see Bayesian E o Es ima o sec ion) and BIC24 as al e na i e app oaches o es ima e he egula ized pa ame e . In o de o s udy he beha io o di e en pa ame e selec ion c i e ia, we i s ain he LR classi ie wi h he aining da a along he dec easing sequence o egula iza ion pa h wi h log10 (λ)∈{0,–0.05,–0.1, –0.15,…, –8.90, –8.95, –9.00}. Then, again he whole aining da a a e used o es ima e he e o a e o each λ. Finally, o each es ima o , we selec he model wi h λ alue ha achie es he minimum e o a e. As esampling in CV-10 in oduces andomness, in his case, we i e a e 200 imes and he esul is a e aged. The de e minis ic na u e in BEE and BIC will p oduce he same esul on he aining da a a each i e a ion. The esul s a e summa ized in Table 3. Fo all me hods, minimum e o a es, AUC, and he numbe o selec ed ea- u es a e es ima ed om he whole aining da a. I is o be no ed ha he epo ed AUC is compu ed om he aining o emphasize ha all ea u e se s a e enough o pa i ion he ea u e space in o wo ca ego ies pe ec ly. The es e o is epo ed la e . The esul s indica e ha he numbe o ea u es selec ed by BEE me hod is lowe compa ed o hose o CV and BIC. Fo he i s ea u e ec o wi h size 49, BEE selec s only 14 ea u es as signi ican , while o he second ea u e ec o wi h −6 −5 −4 −3 −2 −1 0 0 0.02 0.04 0.06 0.08 0.1 0.12 0.14 0.16 0.18 0.2 Log10 (λ) 5− old c oss alida ion e o −6 −5 −4 −3 −2 −1 0 0 0.02 0.04 0.06 0.08 0.1 0.12 0.14 0.16 0.18 0.2 Log10 (λ) BEE Figu e 1. Le : Examples o egula iza ion pa h e o cu es o 5- old CV o ou low cy ome y da a wi h heal hy and AML posi i e classes. Righ : The co esponding Bee cu e. Table 3. Pa ame e selec ion by di e en es ima o s: a e age numbe o selec ed ea u es, λ, aUC, and hei s anda d de ia ions wi h aining da a. METhoD FEATURE TYPE NUMBER oF SELECTED FEATURES SELECTED Log10 (λ) AUC (TRAININg) CV-10 mean 19.72 ± 2.41 −2.95 ± 2.30 0.9997 ± 0.0017 CV-10 mean and s d 23.91 ± 0.80 −4.14 ± 3.00 1 ± 0.0000 Bee mean 15 ± 0.00 −2.05 ± 0.00 0.9989 ± 0.0000 Bee mean and s d 13 ± 0.00 −1.80 ± 0.00 0.9992 ± 0.0000 BIC mean 20 ± 0.00 −5.85 ± 0.00 1 ± 0.0000 BIC mean and s d 24 ± 0.00 −5.70 ± 0.00 1 ± 0.0000 Hassan e al 80 CanCe In o ma ICs 2015:14(s5) leng h 98, BEE selec s only 12 ea u es. Tables 4 and 5 lis he selec ed ea u es, ie, signi ican bioma ke s along wi h he co esponding coe icien alues. Due o he andomness in CV-10, we only p esen he esul s o one i e a ion as an illus- a ion: he e is a signi ican a ia ion o he selec ed ea u es depending on he chosen CV spli . Howe e , i is o be no ed ha he coe icien s o BIC and BEE a e no speci ic o his pa icula i e a ion, as hey do no include he andom spli . Pe o mance assessmen o he model selec ion c i e ia. The pe o mances o he model selec ion me hods a e s udied in he ollowing sec ion. The classi ica ion e o is conside ed as he pe o mance c i e ion, and bo h alse posi i es (heal hy con ol classi ied as AML) and alse nega i es (AML classi ied as heal hy con ol) a e coun ed wi h equal weigh . The pe o - mance o he Bayesian e o es ima o is benchma ked agains hose o CV-10 and BIC o a di e en numbe o sample sizes. Fo his pu pose, a andomly selec ed p opo ion o 10%, 15%, 20%–90%, and 95% is selec ed o aining he classi ie , while he emaining da a a e used o pe o mance assessmen . Fo each aining sample, he expe imen is execu ed 200 imes by gene a ing a new aining se each ime. Classi ica ion e o s. The e o cu es o di e en sample sizes a e shown in Figu e 2. The p ocedu e is epea ed 200 imes o each aining sample size, and he a e age o clas- si ica ion e o is compu ed o each model selec ion c i e ion. Wi h a e y small numbe o aining samples, such as 10% o 15% o he da ase , BEE p o ides imp o ed accu acy o e CV-10 and BIC (Fig. 2 le and igh panels). Fo ins ance, wi h 10% aining samples, he classi ica ion e o s o he model selec ed by CV-10 a e 7.56% (Fig. 2 le panel) and 7.41% (Fig. 2 igh panel) highe han hose o BEE. In case o BIC, he classi ica ion e o is 7.81% highe han ha o BEE (Fig. 2 le panel). As he numbe o aining samples inc eases, o example, abo e 60%, he pe o mance o BIC exceeds han ha o BEE (Fig. 2 le panel). Howe e , he pe o mance o BIC is simila o ha o BEE when mo e ea- u es a e in ol ed in he expe imen (Fig. 2 igh panel). Table 4. The nonze o coe icien s o ea u es wi h mean. TUBE FEATURE 10-FoLD CV BEE BIC Cons an −13.38 −3.50 −13.54 ube 1 sLin 0.75 00.78 ube 1 ssLog −5.92 −0.73 −6.01 ube 1 L1:IgG1- I C −0.46 −0.19 −0.46 ube 1 L4:IgG1-PC5 −2.08 0−2.13 ube 1 L5:IgG1-PC7 −3.07 −0.19 −3.14 ube 2 sLin 0.001 0 0 ube 2 L5:CD20-PC7 2.76 02.78 ube 3 ssLog 0−0.77 0 ube 3 L4:CD8-PC5 −1.94 −0.10 −1.97 ube 4 sLin 0.97 00.97 ube 4 L1:CD15 - I C −4.77 0−4.82 ube 4 L2:CD13-Pe 3.44 0.21 3.48 ube 4 L4:CD16-PC5 0−0.09 0 ube 4 L5:CD56-PC7 3.29 0.75 3.35 ube 5 sLin 2.35 02.40 ube 5 L2:CD11c-Pe −0.15 0−0.16 ube 5 L3:CD45-eCD −1.84 −0.02 −1.84 ube 5 L4:CD64-PC5 1.66 01.69 ube 5 L5:CD33-PC7 0.75 0.60 0.76 ube 6 L2:CD117-Pe 4.82 0.89 4.88 ube 6 L4:CD34-PC5 6.88 0.72 6.99 ube 6 L5:CD38-PC7 00.41 0 ube 7 L1:CD5 - I C −4.68 −0.19 −4.74 No e: The size o he ea u e se s is 49, o which he CV, BEE, and BIC selec 20, 14, and 19 ea u es, espec i ely. Table 5. The nonze o coe icien s o ea u es wi h mean and s anda d de ia ion. TUBE FEATURE 10-FoLD CV BEE BIC Cons an −13.47 −3.27 −13.47 ube 1 sLin mean 0.027 00.027 ube 1 ssLog mean −4.49 −0.47 −4.49 ube 1 L1:IgG1- I C s d −3.80 −0.22 −3.80 ube 1 L5:IgG1-PC7 s d −1.60 −0.16 −1.60 ube 2 L5:CD20-PC7 mean 0.50 00.50 ube 3 ssLog mean 0−0.29 0 ube 3 L5:CD2-PC7 mean −0.48 0−0.48 ube 3 L5:CD2-PC7 s d −1.21 0−1.21 ube 4 sLin mean 0.05 00.05 ube 4 L1:CD15 - I C mean −0.14 0−0.14 ube 4 L2:CD13-Pe mean 2.50 02.50 ube 4 L4:CD16-PC5 mean −1.32 0−1.32 ube 4 L4:CD16-PC5 s d 0−0.39 0 ube 4 L5:CD56-PC7 s d 6.07 0.45 6.07 ube 5 L1:CD14- I C s d −0.004 0−0.004 ube 5 L3:CD45-eCD mean −2.51 0−2.51 ube 5 L5:CD33-PC7 mean 1.47 0.32 1.47 ube 5 L5:CD33-PC7 s d −0.70 0−0.70 ube 6 ssLog s d 0−0.23 0 ube 6 L2:CD117-Pe mean 3.19 0.46 3.19 ube 6 L2:CD117-Pe s d 2.19 02.19 ube 6 L4:CD34-PC5 s d 1.88 0.75 1.88 ube 6 L5:CD38-PC7 mean 1.31 0.44 1.31 ube 7 sLin mean 1.39 01.39 ube 7 sLin s d 0−0.16 0 ube 7 L1:CD5 - I C s d −1.16 0−1.16 ube 7 L5:CD10-PC7 s d 0.04 00.04 No e: The size o he ea u e se s is 98, o which he CV, BEE, and BIC selec 23, 12, and 23 ea u es, espec i ely. Flow cy ome y-based classi ica ion in cance esea ch 81CanCe In o ma ICs 2015:14(s5) A ea unde he ROC cu e. In his sec ion, we e alua e he pe o mance in e ms o AUC. Figu e 3 illus a es he a e age o AUC o di e en aining sample sizes. He e, he BEE me hod achie es imp o emen o e he o he me hods. Wi h small ain- ing samples, o ins ance, 10%, he a e age AUC o BEE is 1.11% (Fig. 3 le panel) and 1.30% (Fig. 3 igh panel) highe han ha o CV-10. As he numbe o aining samples inc eases, CV-10 and BIC also con e ge owa d he esul s o BEE; howe e , he BEE selec ed model consis en ly esul s in he highes AUC sco e. Wi h he la ge ea u e ec o ha includes he measu e- men s o mean and s anda d de ia ion, he a e age AUC cu es o BEE and BIC ollow he simila pa e n (Fig. 3 igh panel). Numbe o selec ed ea u es. We u he assess he pe - o mances o he es ima o s using ea u e selec ion c i e ia. A each i e a ion, we de e mine he o al numbe o selec ed ea u es ha ha e nonze o alues o a di e en numbe o aining samples. Then, we compu e he a e age and he a iabili y (ie, s anda d de ia ion) o he selec ed ea u es o di e en aining samples. The esul s a e illus a ed in Figu e 4. Fo BEE, he a e age numbe o selec ed ea u es is lowe in amoun compa ed o hose o CV and BIC (Fig. 4 op panel). Fo ins ance, wi h 95% aining samples, BEE equi es 36.49% and 33.89% less ea u es han CV and BIC, espec i ely, o model p edic ion (Fig. 4 op- igh panel). Mo eo e , he a iabili y in selec ed ea u es using BEE is also compa able (Fig. 4 bo om panel). The CV-10 has he wo s pe o mance. Al hough BIC shows ha he de ia- ion in ea u e selec ion a di e en i e a ions is smalle , he 0 20 40 60 80 100 120 140 160 180 0.02 0.03 0.04 0.05 0.06 0.07 Numbe o samples used o aining Classi ica ion e o s CV-10 BEEp BIC 0 20 40 60 80 100 120 140 160 180 0.02 0.03 0.04 0.05 0.06 0.07 Numbe o samples used o aining Classi ica ion e o s CV-10 BEEp BIC Figu e 2. The a e age classi ica ion e o cu es o CV-10, BEE wi h p ope p io (BEEp), and BIC. No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s. 0 20 40 60 80 100 120 140 160 180 0.9 0.91 0.92 0.93 0.94 0.95 0.96 0.97 0.98 0.99 1 Numbe o samples used o aining AUC CV-10 BEEp BIC 0 20 40 60 80 100 120 140 160 180 0.9 0.91 0.92 0.93 0.94 0.95 0.96 0.97 0.98 0.99 1 Numbe o samples used o aining AUC CV-10 BEEp BIC Figu e 3. The a e age AUCs o CV-10, BEE wi h p ope p io (BEEp), and BIC. No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s. Hassan e al 82 CanCe In o ma ICs 2015:14(s5) a e age numbe o selec ed ea u es is highe han ha o o he s (Fig. 4 op panel). Simila i y o he selec ed ea u e se s. Ano he pe o mance measu emen is he s abili y o selec ing he same ea u e a di e en i e a ions. Fo his pu pose, Sø ensen–Dice coe - icien 25 is used, which measu es he deg ee o simila i y be ween selec ed ea u es o wo di e en i e a ions. The anges can a y om 0 o 1. The alues closes o 1 indica e a high-deg ee o simila i y. Fo di e en aining samples, we i s de e mine which ea u es a e selec ed a each i e a ion. As he model selec ion p ocess is epea ed 200 imes, we es ima e he simila i y as he mean dice coe icien o each o he 200!/ (2! × (200 − 2)!) = 19,900 possible pai s o selec ed ea u e se s. The esul s a e shown in Figu e 5. In e ms o s abil- i y, he pe o mance o BEE is subs an ially be e han hose o he o he me hods, as he selec ed ea u e se s a e mos simila wi h ha c i e ion – a signi ican issue when ying o unde s and he biological mechanisms behind he da a. Fo example, wi h 60% aining samples, he dice coe icien o BEE is 6.03% highe han ha o CV (Fig. 5 le panel). On he o he hand, wi h 90% aining samples, he dice coe - icien o BEE is 5.81% highe han ha o CV and 4.90% highe han ha o BIC (Fig. 5 igh panel). Indeed, he dice coe icien is un a o able o CV wi h small aining samples: The dice coe icien is lowes among he al e na i es, indica - ing ha he selec ed ea u e se s wi h he CV c i e ion ha e high a iabili y. small sample case wi h a cance cell line. Fo u he con idence on he p esen ed me hod, we analyze da a om a cance cell line in a small sample se ing. As desc ibed p e i- ously, we conside ed he classi ica ion accu acy, AUC mea- su e, and he numbe o selec ed a iables bo h wi h and wi hou s anda d de ia ion ea u es (Figs. 6–8). In his case, 0 20 40 60 80 100 120 140 160 180 8 10 12 14 16 18 20 22 24 Numbe o samples used o aining A e age numbe o ea u es CV-10 BEEp BIC 0 20 40 60 80 100 120 140 160 180 8 10 12 14 16 18 20 22 24 Numbe o samples used o aining A e age numbe o ea u es CV-10 BEEp BIC 0 20 40 60 80 100 120 140 160 180 1 1.5 2 2.5 3 3.5 4 Numbe o samples used o aining S anda d de ia ion o numbe o ea u es CV-10 BEEp BIC 0 20 40 60 80 100 120 140 160 180 1.5 2 2.5 3 3.5 4 4.5 Numbe o samples used o aining S anda d de ia ion o numbe o ea u es CV-10 BEEp BIC Figu e 4. Compa isons o he numbe o selec ed ea u es o CV-10, BEE wi h p ope p io (BEEp), and BIC. No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s. Top: a e age numbe o selec ed ea u es. Bo om: s anda d de ia ion o numbe o he selec ed ea u es. Flow cy ome y-based classi ica ion in cance esea ch 83CanCe In o ma ICs 2015:14(s5) 0 20 40 60 80 100 120 140 160 180 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 Numbe o samples used o aining Dice index CV-10 BEEp BIC 0 20 40 60 80 100 120 140 160 180 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 Numbe o samples used o aining Dice index CV-10 BEEp BIC Figu e 5. Compa ison o he s abili y o selec ing ea u es o CV-10, BEEp, and BIC. No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s. 2 4 6 8 10 12 14 0.2 0.25 0.3 0.35 0.4 0.45 Numbe o samples used o aining Classi ica ion e o s LOOCV BEEp BIC 2 4 6 8 10 12 14 0.2 0.25 0.3 0.35 0.4 0.45 Numbe o samples used o aining Classi ica ion e o s LOOCV BEEp BIC Figu e 6. The a e age classi ica ion e o cu es o LOOCV, BEEp, and BIC. No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s. 2 4 6 8 10 12 14 0.55 0.6 0.65 0.7 0.75 0.8 0.85 0.9 Numbe o samples used o aining AUC LOOCV BEEp BIC 2 4 6 8 10 12 14 0.55 0.6 0.65 0.7 0.75 0.8 0.85 0.9 0.95 Numbe o samples used o aining AUC LOOCV BEEp BIC Figu e 7. The a e age AUC cu es o LOOCV, BEEp, and BIC. No es: Le : ea u e ec o o mean alues o measu emen s. Righ : ea u e ec o o mean and s anda d de ia ion o measu emen s.