BioTex Re ie e : a ool o Re ie e Rele an
Pape s
Célia Talma Gonçal es
LIACC & Faculdade de Engenha ia da Uni e sidade do Po o &
Ins i u o Supe io de Con abilidade e Adminis ação do Po o & CEISE-STI, Po ugal
Rui Camacho
LIAAD & DEI & Faculdade de Engenha ia da Uni e sidade do Po o, Po ugal
Eugénio Oli ei a
LIACC & DEI & Faculdade de Engenha ia da Uni e sidade do Po o, Po ugal
ABSTRACT
Whene e new sequences o DNA o p o eins ha e been decoded i is almos compulso y o look
a simila sequences and pape s desc ibing hose sequences in o de o bo h collec ele an
in o ma ion conce ning he unc ion and ac i i y o he new sequences and/o know wha is
known al eady abou simila sequences ha migh be use ul in he explana ion o he unc ion o
ac i i y o he newly disco e ed ones.
In cu en web si es and da a bases o sequences he e a e, usually, a se o cu a ed pape
e e ences linked o each sequence. Those links a e e y use ul since he pape s desc ibe use ul
in o ma ion conce ning he sequences. They a e, he e o e, a good s a ing poin o look o
ele an in o ma ion ela ed o a se o sequences. One way is o implemen such app oach is o
do a blas wi h he new decoded sequences, and collec simila sequences. Then one looks a he
pape s linked wi h he simila sequences. Mos o en he numbe o e ie ed pape s is small and
one has o sea ch la ge da a bases o ele an pape s.
In his pape we p opose a p ocess o gene a ing a classi ie based on he ini ially se o ele an
pape s. Fi s we collec simila sequences using an alignmen algo i hm like Blas . We hen use
he enla ges se o pape s o cons uc a classi ie . Finally we use ha classi ie o au oma ically
enla ge he se o ele an pape s by sea ching he MEDLINE using he au oma ically
cons uc ed classi ie . We ha e empi ically e alua ed ou p oposal and epo e y p omising
esul s.
Keywo ds: MEDLINE, Classi ica ion, In o ma ion Re ie al Sys em, Machine Lea ning,
Ensemble Algo i hms, Bioin o ma ics
INTRODUCTION
Molecula Biology and Biomedicine scien i ic publica ions a e a ailable (a leas he abs ac s) in
Medical Li e a u e Analysis and Re ie al Sys em On-line (MEDLINE). MEDLINE is he U.S.
Na ional Lib a y o Medicine (NLM), p emie bibliog aphic da abase: con ains o e 16 million
e e ences o jou nal a icles in li e sciences wi h a concen a ion on Biomedicine. A dis inc i e
ea u e o MEDLINE is ha he eco ds a e indexed wi h NLM’s Medical Subjec Headings
(MeSH e ms). MEDLINE is he majo componen o PubMed (Wheele e al., 2006), a da abase
o ci a ions o he NLM. PubMed comp ises mo e han 19 million ci a ions o biomedical
a icles om MEDLINE and li e science jou nals. The PubMed da abase main ained by he
Na ional Cen e o Bio echnology In o ma ion (NCBI) is a key esou ce o biomedical science,
and is ou i s base o wo k. The NCBIs PubMed sys em is a widely used me hod o accessing
MEDLINE.
The esul o a MEDLINE/PubMed sea ch is a lis o ci a ions (including au ho s, i le, jou nal
name, pape abs ac , keywo ds and MeSH e ms) o jou nal a icles. The esul o such sea ch is,
qui e o en, a huge amoun o documen s, making i e y ha d o esea che s o e icien ly each
he mos ele an documen s. As his is a e y ele an and ac ual opic o in es iga ion we
assess he use o Machine Lea ning-based ex classi ica ion echniques o help in he
iden i ica ion o a easonable amoun o ele an documen s in MEDLINE. The co e o he
epo ed wo k is o s udy he bes way o cons uc he da a se s and he classi ie s om he
s a ing se o sequences.
These expe iences we e done using a se o posi i e examples associa ed o he
sequences/keywo ds gi en by he use and a se o nega i e examples which is he ocuses o his
pape . The nega i e examples we e gene a ed in h ee di e en ways and we in end o show
which is he bes app oach o ou classi ica ion pu pose. In ou expe imen s we ha e used
se e al classi ica ion algo i hms a ailable in he WEKA (Hall e al., 2009) ool including
ensemble algo i hms. We ha e also made some sensi i i y es s o he p uning o a ibu es o
a ibu e educ ion.
The es o he pape is s uc u ed as ollows. The sec ion “An A chi ec u e o an In o ma ion
Re ie al Sys em” p esen s he a chi ec u e o ou in o ma ion e ie al sys em and he
ollowing Sec ion “The Local Da a base” desc ibes he local da a base cons uc ion p ocess and
he p e-p ocessing echniques used. Follows he ela ed wo k and a sec ion dedica ed o he
“Au oma ic Cons uc ion o da a se s” ha desc ibes he di e en al e na i es p oposed o da a
se cons uc ion and he expe iences we ha e done wi h di e en classi ie s. We also include a
sec ion dedica ed o classi ie ensemble whe e we p esen ou expe imen s using classi ie
ensemble me hods and inally we conclude he pape .
AN ARCHITECTURE FOR AN INFORMATION RETRIEVAL SYSTEM
The o e all goal o ou wo k is o implemen a web based sea ch ool ha ecei es a se o
genomic o p o eomic sequences and e u ns an o de ed se o pape s ele an o he s udy o
such sequences. The ini ial se o sequences is supplied by a biologis oge he wi h a se o
ele an keywo ds and an e- alue1. These h ee i ems a e he inpu o BioTex Re ie e
(Gonçal es & Camacho and Oli ei a, 2011) as can be seen in Figu e 1. Figu e 1 p esen s a
summa y o ou app oach ha we will now desc ibe in de ail. In he ollowing desc ip ion we use
NCBI as he sequence Da a Base.
1 A e- alue is a s a is ic o es ima e he signi icance o a ma ch be ween 2 sequences
Figu e 1. Sequence o s eps execu ed by BioTex Re ie e when he use p o ides a se o ini ial
DNA/p o ein sequences.
In S ep 1, he use (a biologis esea che ) p o ides an ini ial o sequences, op ionally a lis o
keywo ds, and an e- alue. Wi h hese h ee i ems (sequences, keywo ds and e- alue) and using
he NCBI BLAST ool we collec a se o simila sequences oge he wi h he pape e e ences
associa ed o hem. We could also use Ensembl wi h he same inpu s because Ensembl may
e u n a di e en se o pape s e e ences. Howe e o he p oposed wo k we ha e only used he
NCBI da abase.
Wi h his lis o pape ci a ions we sea ch o hei abs ac s in a local copy o MEDLINE (LDB
– Local Da abase) (S ep 3). Fo his we ha e p e iously p ep ocessed MEDLINE. S ep 3
sea ches and collec s he ollowing in o ma ion in he p e-p ocessed local copy o MEDLINE:
pmid, jou nal i le, jou nal ISSN, a icle i le, abs ac , lis o au ho s, lis o keywo ds, lis o
MeSH e ms and publica ion da e.
Fo he scope o his pape we a e conside ing only he pape ci a ions ha ha e he abs ac
a ailable in MEDLINE. A e S ep 3 we ha e a da a se o pape s ela ed o he sequences. We
will ake his se o pape s as he posi i e examples o he ull cons uc ion o he da a se (S ep
4) bu we need o ge some nega i e examples. To ob ain he nega i e examples we ha e h ee
possible app oaches.
Thus his s ep is explained in de ail in Sec ion III.
The ollowing s ep, S ep 5, is one o he mos impo an s eps o ou wo k which is o Cons uc
a Classi ie using Machine Lea ning echniques ha is explained in he nex sec ion. As a esul
o his s ep we ha e a ull lis o a icles conside ed ele an by ou classi ie (S ep 6). Howe e ,
we need o p esen hem o he biologis in an o de ed ashion way. So S ep 7 p esen s an o de ed
lis o ele an a icles o
he biologis . He e we will de elop and implemen a anking algo i hm based on ea u es such as
he numbe o ci a ions o he pape and he impac ac o o he jou nal/con e ence whe e i was
published. This pape ocuses on he cons uc ion o he da a se s highligh ing he esea ch om
he di e en app oaches o ob ain he nega i e examples. Acco ding o he igu e and o he
pu pose o his pape we ocus on S ep 4, al hough we ha e made a se o expe iences wi h some
classi ie s o conclude wha was he bes app oach.
THE LOCAL DATA BASE
We ha e downloaded 80GB (617 XML iles) o MEDLINE 2010 om he NCBI websi e. Each
XML ile has in o ma ion cha ac e izing one ci a ion. Among hese cha ac e is ics we ha e
conside ed he ollowing ones: PMID - he PubMed Iden i ie ; he PubMed Da e; he Jou nal
Ti le; he Jou nal ISSN ha co esponds o he ISI Web o Knowledge ISSN; he Ti le; he
Abs ac o he a icle i a ailable; he lis o he Au ho s; he MeSH Headings lis and he
Keywo ds lis . A e download he iles we e p ep ocessed as ollows.
P e-P ocessing MEDLINE XML iles
An independen s ep o ou ool is o main ain a local copy o MEDLINE, ha we will call Local
Da a Base (LDB). The LDB will enable e icien sea ch o he pape and will ha e ha ele an
in o ma ion o each pape in o ma adequa e, he algo i hm ha cons uc a chain as desc ibe
u he in his pape .
The i s s ep is o ead he XML iles and ex ac he ele an in o ma ion o s o e in he LDB.
A icle’s i le and abs ac a e p ep ocessed wi h “ adi ional” ex p e-p ocessing echniques.
Nex sec ion p esen s he p ep ocessing echniques applied.
P e-P ocessing Techniques
We ha e empi ically (Gonçal es & Gonçal es & Camacho & Oli ei a, 2010) e alua e which a e
he bes combina ion o p e-p ocessing echniques o achie e a be e accu acy. Based on his
p e ious s udy and wi h some mo e esea ch in he meanwhile we ha e used he ollowing
p ep ocessing echniques.
Documen Rep esen a ion
Fo each pape wi h he in o ma ion e e ed in he beginning o his sec ion Howe e he ex
ac s o a documen ( i le and abs ac ) a e il e ed using ex p ocessing echniques and
ep esen ed using he ec o space model om In o ma ion Re ie al whe e he alue o a e m
in a documen is gi en by he s anda d e m- equency in e se documen equency
(TFIDF=TF*IDF) unc ion (Zhou & Smalheise & Yu, 2006), o assign weigh s o each e m in
he documen .
TF is he equency o e m in documen
and
1
sin
log
+=
h e mcumen swi numbe o do
collec ioncumen numbe o do
IDF
Named En i y Recogni ion (NER)
NER is he ask o iden i ying e ms ha men ion a known en i y. We ha e used ABNER
(Se les, 2005), which s ands o A Biomedical Named En i y Recogni ion, ha is a so wa e ool
o molecula biology ha iden i ies en i ies in he biology domain: p o eins, RNA, DNA, cell
ype and cell line. Al hough we ha e implemen ed his echnique we ha e concluded ha he
iden i ica ion o NER e ms augmen s signi ican ly he numbe o a ibu es ins ead o educing
hem. We concluded ha he use o NER inc eases s ongly he numbe o e ms which is a
p oblem o he classi ie s. Thus we did no use NER in he p e-p ocessing phase.
Handling Synonyms
We handle synonyms using he Wo dNe (Fellbaum, 1998) o sea ch o simila e ms, in he
case o egula e ms, and used Gene On ology (Ashbu ne , 2000) o ind biological synonyms.
I wo wo ds mean he same hen hey a e synonyms, so hey could be eplaced by one o hem in
he en i e MEDLINE ( i le and abs ac ields) wi hou changing he seman ic meaning o he
e m hus educing he numbe o a ibu es. In his s ep we ha e eplaced all he synonyms
ound by one synonym e m hus educing he numbe o e ms.
Dic iona y Valida ion
A e m is conside ed a alid e m i i appea s in a ailable dic iona ies. We ha e ga he ed se e al
dic iona ies o he common English e ms ( such as Ispell and Wo dNe ) and o he medical and
biological e ms (BioLexicon (Rebholz-Schuhmann e al., 2008), The Hos o d Medical Te ms
Dic iona y (Hos o d, 2004) and Gene On ology (Ashbu ne , 2000). The Hos o d Medical Te ms
Dic iona y consis s o a ile ha con ains a long lis o medical e ms. BioLexicon is a la ge-scale
e minological esou ce de eloped o add ess ex mining equi emen s in he biomedical
domain. The BioLexicon is publicly a ailable bo h as an XML- o ma ed e m eposi o y and as
a ela ional da abase (MySQL) and i adhe es o he LMF ISO s anda ds o lexical esou ces.
We ha e also used he Gene On ology a ailable iles ha a e ela ed o genes, enzymes,
chemical esou ces, species and p o eins. We ha e p ocessed each o hese esou ce iles in o de
o ha e a simple ex ile wi h one e m pe line.
Ou app oach is in he sense ha i a e m appea s in one o hese dic iona ies i is a alid e m,
o he wise, i is no a alid e m, so we emo e i om he collec ion o e ms.
The applica ion o hese echnique is undamen al in a ibu e educ ion once a lo o e ms ha
ha e no biology, medical and no mal signi icance a e disca ded.
S op Wo ds Remo al
S op Wo ds Remo al emo es wo ds ha a e meaningless such as a icles, conjunc ion and
p eposi ions (e.g., a, he, a , e c.). These wo ds a e meaningless o he e alua ion o he
documen con en . We ha e used a se o 659 s op wo ds ile.
Tokeniza ion
Tokeniza ion is he p ocess o b eaking a ex in o okens. A oken is a non emp y sequence o
cha ac e s, excluding spaces and punc ua ion.
Special Cha ac e s Remo al
Special cha ac e emo al emo es all he special cha ac e s (+, -, !, ?, ., ,, ;, :, , g, =, &, #, %, $,
[, ], /, <, >, n, “, ”, j) and digi s.
S emming
S emming is he p ocess o emo ing in lec ional a ixes o wo ds educing he wo ds o hei
s em ( he wo ds compu e , compu ing and compu a ion a e all ans o med in o compu , which
means ha h ee di e en e ms a e ans o med in o only one e m hus educing he numbe o
a ibu es. We implemen ed he Po e ’s S emme Algo i hm (Po e , 1997).
P uning
Using p uning we disca d in he documen s collec ion e ms ha ei he appea oo a ely o oo
equen ly.
RELATED WORK
The e a e some wo k being done on biological and biomedical documen classi ica ion. Some o
hem applied o MEDLINE documen classi ica ion and o he da abases.
The wo k o (Sehgal, 2011) ies o au oma e he p ocess o adding new in o ma ion o TCDB
da abase (T anspo Classi ica ion Da abase) ha is a web ee access da abase
(h p://www. cdb.o g) abou comp ehensi e in o ma ion on anspo p o eins. The au ho s
es ic ed hemsel es o he
documen s in MEDLINE. The main goal is o highligh he use o Machine Lea ning echniques
ou pe o ms ules c ea ed by hand by a human expe . To ain he classi ie hey ha e used a se
o MEDLINE documen s e e ed TCDB as posi i e examples and ha e selec ed andomly also
om MEDLINE a se o nega i e examples.
The au ho s in (Imambi & Sudha, 2011) desc ibe a new model o ex classi ica ion using
es ima ing e m weigh s which imp o es accu acy classi ica ion acco ding o he au ho s
expe iences. Documen s a e ep esen ed as ec o s o e ms wi h hei no malized global
equency. Global weigh s a e unc ions ha coun how many imes a e m appea s in he en i e
collec ion and he no maliza ion p ocess compensa es he disc epancies in he leng hs o he
documen s. They ha e
used 1000 documen s om PubMed; 600 documen s o he aining da a se and 400 o he es
da a se . All hese documen s belong o ou ca ego ies wi h MeSH e ms ela ed o Diabe es
meli us. The au ho s compa e he di e en weigh ing me hods: local-bina y, local-log, local d
and global ele an . They concluded in his s udy ha global ele an weigh ing me hod achie es
a highe p ecision. In ou own wo k we ha e also used all no malized global equency.
BioQSpace (Di oli e al., 2005) is a GUI whe e use s can que y abs ac s om PubMed using an
embedded sea ch acili y. BioQSpace pe o ms pai wise simila i y calcula ions be ween all he
abs ac s based on a se o indi idual a ibu es namely: s uc u e, unc ion, disease and
he apeu ic compounds wo d lis ob ained om MeSH e ms, wo d usage, PubMed ela ed
a icles, publica ion da e among o he s. These a ibu es a e gi en mo e o less impo ance
acco ding o he weigh a ibu ed by use s. A clus e ing algo i hm is used o g oup abs ac s ha
a e e y simila .
(F unza & Inkpen & T an, 2011) desc ibe a me hodology o build an applica ion capable o
iden i ying and dissemina ing heal h ca e in o ma ion using a Machine Lea ning app oach. The
main objec i e o hei wo k is o s udy he bes in o ma ion ep esen a ion model and wha
classi ica ion algo i hms a e sui able o classi ying ele an medical in o ma ion in sho ex s.
They ha e used 6 di e en Machine Lea ning algo i hms. The au ho s concluded ha nai e
Bayes pe o med e y well on sho ex s in he medical domain and ha adaboos had he wo s
esul .
In (Dollah & Seddiqui & Aono, 2010) he au ho s p esen an app oach o classi ying a
collec ion o biomedical abs ac s downloaded om MEDLINE da abase wi h he help o
on ology alignmen . Al hough his wo k classi ies MEDLINE documen s i is based on on ology
alignmen which is ou o ou scope.
Lige Ca (Sa ka e al., 2009) s ands o Li e a u e and Genomic Elec onic Resou ce Ca alogue,
and i is a sys em o explo ing biomedical li e a u e h ough he selec ion o e ms wi hin a
MeSH cloud ha is gene a ed based on an ini ial que y using jou nal, a icle, o gene da a. The
cen al idea o Lige Ca is o c ea e a ag cloud showing an o e iew o impo an concep s and
ends associa ed o he MeSH desc ip o s. Lige Ca agg ega es mul iple a icles in PubMed,
combining
he associa ed MeSH desc ip o s in o a cloud, weigh ed by equency. Lige Ca does no apply
any Machine Lea ning echniques o pape classi ica ion as we p esen in ou s udy.
AUTOMATIC CONSTRUCTION OF DATA SETS
Cons uc ing he Da a Se s
In o de o sol e ou p oblem ha is gi en a se o genomic o p o eomic sequences e u n a se
o ela ed sequences and pape s wi h ele an in o ma ion o he s udy o such sequences, we
need o ob ain i s o all he a icles associa ed wi h he gi en se o sequences and cons uc he
da a se o gi e o he classi ie . Figu e 2 illus a es sequence o s eps included in
BioTex Re i e . The empi ical wo k epea ed in he pape conce ns he cons uc ion o a da a
se (S ep 4).
Figu e 2. Da a Se Cons uc ion
The inpu o ou wo k is a se o sequences gi en in he FASTA o ma . We use he
ne blas -2.2.22 ool ha pe o m a emo e blas sea ch a he NCBI si e. We ha e embedded his
applica ion in o ou code and au oma ically ha e access o bo h, he o iginal sequences and he
se o simila sequences e ie ed by BLAST.
These esul s show us he simila sequences and he e alue associa ed wi h each o he e ie ed
sequence. The e- alue is a s a is ic o es ima e he signi icance o a ma ch be ween wo
sequences. The e- alue is an inpu ha is gi en us by he biologis . We elax his h eshold alue
in o de o ob ain he nega i e examples based on he e- alue as we can see in Figu e 3. The
posi i e examples a e he one’s ha a e lowe han he e- alue p e iously speci ied by he
biologis . We es ablish a “no man’s land“ zone and a e ha zone we collec he nega i e
examples.
Figu e 3. How posi i e and nea -miss (nega i e) examples a e ob ained. e is he e- alue h eshold o
ob ain he posi i e examples. α and β a e pa ame e s o he cu o o he nega i e examples.
The posi i e examples a e he se o pape s associa ed wi h he se o sequences wi h e- alue
below he espec i e h eshold. In his s udy we ha e empi ically e alua ed h ee di e en ways
o ob aining he nega i es examples. We now explain he al e na i es.
Nea -Miss Values (NMV)
To ob ain he Nea -Miss Values (NMV) we collec he pape s associa ed wi h he simila
sequences ha ha e e alue abo e he h eshold bu close o ha . In Figu e 3 he e is a s ip g ay
o be e disc imina e wha a e posi i e examples and nega i e examples. The examples in he
igh mos box con ains nea -miss nega i e examples because hey a e no posi i es bu ha e a
ce ain deg ee o simila i y wi h he sequences. This wo ks on he examples ha ha e a
minimum numbe o nega i e examples. I we do no ha e any nega i e examples wi h his
app oach, o i he nega i e examples a e ew, we can ollow one o he ollowing app oaches: o
use MeSH Random Values o o use Random Values. In ou expe imen s we ha e conside ed
e- alue = 0.001 and we ha e elaxed i o 1, 2 and 5 as we can see in sequences dis ibu ion
ables in he nex Sec ion. We ha e elaxed o hese di e en alues o ob ain mo e nega i e
examples. The a icles associa ed o he simila sequences wi h e- alue less hen 0.001 a e
conside ed posi i e; he a icles associa ed o he simila sequences ha ha e e- alues g ea e
hen 0.001 and e- alues less hen 0.001 plus 10% (_ = 10%) a e conside ed in he g ay s ip so
hey a e no conside ed posi i e o nega i e; he a icles associa ed wi h o he simila sequences
wi h e- alues g ea e hen 0.001 plus 10% a e conside ed nega i e examples (nea -miss alues).
MeSH Random Values (MRV)
This al e na i e o gene a e nega i es is adop ed when we do no ha e su icien numbe o
nega i e examples o he classi ie o lea n. The nega i e examples a e ob ained combining he
nea miss alues, i hey exis , wi h some andom examples gene a ed om he LDB. Bu , hese
MRV examples mus ha e he maximum numbe o MeSH e ms om he posi i e examples. A
he end he numbe o nega i e examples is equal o he numbe o posi i e examples.
Random Values (RV)
The las app oach is o gene a e jus andomly he nega i e examples om ou LDB in a numbe
equal o he numbe o posi i e examples. We gua an ee ha in his se he e is no posi i e
example.
COMPARING THE ALTERNATIVES TO DATA SET CONSTRUCTION
Da a Se Cha ac e iza ion
Fo , his s udy we ha e gene a ed se e al da a se s based on sequences ha belong o six
di e en classes, wi h he ollowing dis ibu ion:
•RNASES: 2 sequences
•Esche ichia Coli: 5 sequences
•Choles e ol: 5 sequences
•Hemoglobin: 5 sequences
•Blood P essu e: 5 sequences
•Alzheime : 5 sequences
We ha e also used h ee di e en elaxa ion alues o he e- alue (1, 2 and 5). I he use en e s
an e- alue o 0.001, hen he posi i e examples a e he ones ha ha e e- alue less o equal o
0.001. And he nega i e examples a e he one’s g ea e hen 0.001. Bu as we can see in Figu e 3
we lea e a g ay s ip o be e sepa a e he posi i e om he nega i e examples. This s ip is also
de ined by he use . Fo hese examples we ha e de ined a s ip o 10% o he numbe o no
simila sequences. So he nega i e nea -miss examples a e he one’s ha a e g ea e hen 0.001
plus 10% o he o he numbe o no simila sequences and lowe han e- alue elaxa ion alue
(1, 2 o 5 in ou examples).
The main idea o his s udy is o s udy he bes way o cons uc he nega i e examples based on
ou expe iences. The dis ibu ions o posi i e and nega i e examples a e show in he Appendix.
Expe imen al Resul s
In ou expe iences we ha e used a se o algo i hms a ailable in he WEKA (Hall e al., 2009)
ools and ha a e lis ed Table 1.
Ac onym Algo i hm Type
Ze oR Majo i y p edic o Rule lea ne
smo Sequen ial Minimal Op imiza ion Suppo Vec o Machines
Random Fo es Ensemble
ibk K-nea es neighbo s Ins ance-based lea ne
BayesNe Bayesan Ne wo k Bayes lea ne
j48 Decision ee (C4.5) Decision ee lea ne
d nb Decision able / naï e bayes hyb id Rule lea ne
AdaBoos Boos ing algo i hm Ensemble lea ne
Bagging Bagging algo i hm Ensemble lea ne
Ensemble Selec ion Combines se e al algo i hms Ensemble lea ne
Table 1. Machine Lea ning Algo i hms used in he s udy.
The da a se s used a e cha ac e ized in Tables 2 and 3. Table 2 cha ac e izes da a se s o which
he nega i e examples a e made only o nea miss examples. Table 3 cha ac e izes he da a se s
o which he e we e no enough nega i e examples and he e o e we ha e used he MRV and
RV s a egies.
Homayouni H. Hashemi S. Hamzeh A., A Lazy Ensemble Lea ning Me hod o
Classi ica ion, IJCSI In e na ional Jou nal o Compu e Science Issues, Vol. 7, Issue 5,
Sep embe 2010.
Hos o d medical e ms dic iona y 3.0, 2004.
Ind a N., Sa ka N., Schenk R., Mille H. and No on C.. Lige Ca : using ”MeSH Clouds” om
jou nal, a icle, o gene ci a ions o acili a e he iden i ica ion o ele an biomedical li e a u e.
AMIA - Annual Symposium p oceedings / AMIA Symposium. AMIA Symposium,
2009:563–567, 2009.
Imambi S. and Sudha T.. Classi ica ion o medline documen s using global ele an weighing
schema. In e na ional Jou nal o Compu e Applica ions, 16(3):45–48, Feb ua y 2011. Published
by Founda ion o Compu e Science.
Ko sian is S. and Pin elas P., Combining Bagging and Boos ing, In e na ional Jou nal o
Compu a ional In elligence, Vol. 1, No. 4 (324-333), 2004.
Opi z D. and Maclin R., Popula Ensemble Me hods: An Empi ical S udy, Jou nal o A i icial
In elligence Resea ch, ol.11, pp. 169-198, 1999.
Po e M. F.. An algo i hm o su ix s ipping. pages 313–316, 1997.
Rebholz-Schuhmann D., Pezik P., Lee V., Kim J-J, Del G a a R., Sasaki Y., McNaugh J.,
Mon emagni S., MonachiniM., Calzola i N. and Ananiadou S. . Biolexicon: Towa ds a e e ence
e minological esou ce in he biomedical domain. In P oceedings o he o he 16 h Annual
In e na ional Con e ence on In elligen Sys ems o Molecula Biology (ISMB- 2008), 2008.
Sehgal A. K., Sanmay D., No o K., Mil on H., Saie J . and Elkan C.. Iden i ying ele an da a
o a biological da abase: Handc a ed ules e sus machine lea ning. IEEE/ACM T ans.
Compu . Biology Bioin o m., 8(3):851–857, 2011.
Se les B.. Abne : an open sou ce ool o au oma ically agging genes, p o eins and o he en i y
names in ex . Bioin o ma ics, 21(14):3191–3192, 2005.
Wheele D. L., Ba e T., Benson D. A, B yan S. H., Canese K., Che e nin V., Chu ch D. M.,
Dicuccio M., Edga R., Fede hen S., Gee L. Y., Helmbe g W., Kapus in Y., Ken on D. L.,
Kho ayko O., Lipman D. J., Madden T. L., Maglo D. R., Os ell j., P ui K. D., Schule G. D.,
Sch iml L. M., Sequei a E., She y S. T., Si o kin K., Sou o o A., S a chenko G., Suzek T. O.,
Ta uso R., Ta uso a T. A., Wagne L., and Yaschenko E.. Da abase esou ces o he na ional
cen e o bio echnology in o ma ion. Nucleic Acids Res, 34(Da abase issue), Janua y 2006.
Zhou W., Smalheise N. R. and Yu C.. A u o ial on in o ma ion e ie al: basic e ms and
concep s. Jou nal o Biomedical Disco e y and Collabo a ion, 1:2, Ma ch 2006.