A Comparison of PSO and GA Approaches for Gene Selection and Classification of Microarray Data
Full text
ACompa ison o PSO and GA App oaches o Gene
Selec ion and Classi ica ion o Mic oa ay Da a
José Ga cía-Nie o,
En ique Alba
Dep . de Leng. y Ciencias de la Compu ación
Uni e si y o Málaga ETSI In o má ica,
Málaga - 29071, Spain
{jnie o,ea }@lcc.uma.es
Lae i ia Jou dan,
El-Ghazali Talbi
LIFL-INRIA Fu u s
B^
a M3, Ci é Scien i ique
59655 Villeneu e d’Ascq, F ance
{jou dan, albi}@li l.
Fea u e selec ion o gene exp ession analysis in cance
p edic ion o en uses w appe classi ica ion me hods o dis-
c imina e a ype o umo , o educe he numbe o genes
o in es iga e in case o a new pa ien . By c ea ing clus-
e s a big educ ion o he numbe o conside ed genes and
an imp o emen o he classi ica ion accu acy can be inally
achie ed. The de ini ion o he ea u e selec ion p oblem is
his: gi en a se o ea u es F={ 1, ..., i, ..., n}, ind a
subse F0⊆F ha maximizes a sco ing unc ion Θ : Γ →G
such ha F0=a gmaxG⊂Γ{Θ(G)},(1)
whe e Γ is he space o all possible ea u e subse s o F
and Ga subse o Γ. The op imal ea u e selec ion p oblem
has been shown o be NP-ha d. The e o e, only heu is ics
app oaches a e able o deal wi h la ge size p oblems.
In his wo k, we a e in e es ed in gene selec ion and classi-
ica ion o DNA Mic oa ay da a in o de o dis inguish u-
mo samples om no mal ones. Fo his pu pose, we p opose
wo hyb id models ha use me aheu is ics and classi ica ion
echniques. The i s one consis s o a Pa icle Swa m Op i-
miza ion (PSO) combined wi h a SVM app oach as w appe
me hod. The second model is based on he popula GA us-
ing a specialized SSOCF [1] c osso e ope a o , ha will
be also combined wi h SVM in ou app oach. A second
impo an con ibu ion consis s in he ac ual disco e y o
new and challenging esul s on six public da ase s iden i-
ying signi ican in he de elopmen o a a ie y o cance s
(leukemia, b eas , colon, o a ian, p os a e, and lung om
he URL h p://sdmc.li .o g.sg/GEDa ase s/Da ase s.
h ml).
Fo ou PSOSV M app oach, a bina y e sion o PSO was
implemen ed in C++ ollowing he skele on a chi ec u e o
he MALLBA lib a y [2]. Fo he GASV M app oach he
GA was implemen ed in C++ using he Pa adisEO F ame-
wo k. Since he posi ion o a pa icle (ch omosome in GA)
x ep esen s a gene subse , he e alua ion is ca ied ou by
means o he SVM classi ie o assess he quali y o he ep-
esen ed gene subse . The i ness o a pa icle/ch omosome
xis calcula ed applying a Lea e One Ou C oss Valida ion
(LOOCV) me hod o calcula e he a e o co ec classi i-
ca ion (accu acy) o a SVM ained wi h his gene subse .
The comple e i ness unc ion is desc ibed in Equa ion 2.
i ness(x) = α·(100/accu acy) + β·# ea u es, (2)
Copy igh is held by he au ho /owne (s).
GECCO ’07, July 7-11, 2007, London, England, Uni ed Kingdom.
ACM 978-1-59593-697-4/07/0007..
whe e αand βa e weigh alues se o 0.75 and 0.25 e-
spec i ely. The objec i e he e consis s o maximizing he
accu acy and minimizing he numbe o genes (# ea u es).
Fo con enience (only minimiza ion o i ness) he i s ac-
o is p esen ed as (100/accu acy).
An special ini ializa ion me hod was adap ed o gene se-
lec ion as ollows. The swa m/popula ion was di ided in o
ou subse s o pa icles/ch omosomes ini ialized in di e en
ways depending on he numbe o ea u es in each pa icle.
Tha is, 10% o pa icles we e ini ialized wi h N(p e ixed
alue) selec ed genes (1s) loca ed andomly. Ano he 20%
o pa icles we e ini ialized wi h 2Ngenes, 30% wi h 3N
genes and inally, he es o pa icles (40%) we e ini ialized
andomly and 50% o he genes we e u ned on.
Table 1: Subse s epo ed wi h 100% es accu acy
Da ase Algo i hm Genes
Leukemia PSOSV M 100(3) K01383 a , U03056 a ,
J04130 s a
B eas PSOSV M 100(4) Con ig49744 RC, Con ig26884 RC
Con ig25936 RC, Con ig13846 RC
Colon PSOSV M 100(3) H64398, H73758
U27699
Lung GASV M 100(3) 33762 a , 34648 a
728 a , 829 s a
O a ian GASV M 100(2) MZ1154.6306
MZ2653.8464
P os a e GASV M 100(3) 35935 a , 39801 a
40069 a
In conclusion, bo h app oaches we e expe imen ally as-
sessed on six well-known cance da ase s disco e ing new
and challenging esul s, and iden i ying speci ic genes ha
ou wo k sugges s as signi ican ones. In his sense, com-
pa isons wi h se e al s a e o a me hods show compe i i e
esul s acco ding o s anda d e alua ion. Resul s o 100%
classi ica ion a e and ew genes pe subse (2, 3 and 4) a e
ob ained in mos o ou execu ions (see Table 1). The use o
an adap ed ini ializa ion me hod has shown a g ea in luence
on he pe o mance o p oposed algo i hms, since i in o-
duces an ea ly se o accep able solu ions in hei e olu ion
p ocess.
1. REFERENCES
[1] L. Jou dan, C. Dhaenens, and E.-G. Talbi. A gene ic
algo i hm o ea u e selec ion in da a-mining o
gene ics. In P oceedings o he 4 h Me aheu is ics
In e na ional Con e encePo o (MIC’2001), pages
29–34, Po o, Po ugal, 2001.
[2] E. Alba and M. g oup. Mallba: A Lib a y o Skele ons
o Combina o ial Op imisa ion. In B. Monien and
R. Feldmann, edi o s, P oceedings o he Eu o-Pa ,
olume LNCS 2400, pages 927–932, 2002.
427