scieee Science in your language
[en] (orig)

BioTextRetriever: A Tool to Retrieve Relevant Papers.

Abstract

Whenever new sequences of DNA or proteins have been decoded it is almost compulsory to look at similar sequences and papers describing those sequences in order to both collect relevant information concerning the function and activity of the new sequences and/or know what is known already about similar sequences. In current web sites and data bases of sequences there are, usually, a set of curated paper references linked to each sequence. Those links are a good starting point to look for relevant information related to a set of sequences. One way to implement such approach is to do a blast with the new decoded sequences, and collect similar sequences. Then one looks at the papers linked with the similar sequences. Most often the number of retrieved papers is small and one has to search large data bases for relevant papers. This paper proposes a process of generating a classifier based on the initially set of relevant papers. First, the authors collect similar sequences using an alignment algorithm like Blast. Then, the authors use the enlarges set of papers to construct a classifier. Finally a classifier is used to automatically enlarge the set of relevant papers by searching the MEDLINE using the automatically constructed classifier.

Read accessible full text

BioTextRetriever: A Tool to Retrieve Relevant Papers.

Author: Célia Talma Gonçalves,Rui Camacho,Eugénio Oliveira
Year: 2011
DOI: 10.4018/jkdb.2011070102
Source: https://repositorio-aberto.up.pt/bitstream/10216/67120/2/62316.pdf
BioTex Re ie e : a ool o Re ie e Rele an
Pape s
Célia Talma Gonçal es
LIACC & Faculdade de Engenha ia da Uni e sidade do Po o &
Ins i u o Supe io de Con abilidade e Adminis ação do Po o & CEISE-STI, Po ugal
Rui Camacho
LIAAD & DEI & Faculdade de Engenha ia da Uni e sidade do Po o, Po ugal
Eugénio Oli ei a
LIACC & DEI & Faculdade de Engenha ia da Uni e sidade do Po o, Po ugal
ABSTRACT
Whene e new sequences o DNA o p o eins ha e been decoded i is almos compulso y o look
a simila sequences and pape s desc ibing hose sequences in o de o bo h collec ele an
in o ma ion conce ning he unc ion and ac i i y o he new sequences and/o know wha is
known al eady abou simila sequences ha migh be use ul in he explana ion o he unc ion o
ac i i y o he newly disco e ed ones.
In cu en web si es and da a bases o sequences he e a e, usually, a se o cu a ed pape
e e ences linked o each sequence. Those links a e e y use ul since he pape s desc ibe use ul
in o ma ion conce ning he sequences. They a e, he e o e, a good s a ing poin o look o
ele an in o ma ion ela ed o a se o sequences. One way is o implemen such app oach is o
do a blas wi h he new decoded sequences, and collec simila sequences. Then one looks a he
pape s linked wi h he simila sequences. Mos o en he numbe o e ie ed pape s is small and
one has o sea ch la ge da a bases o ele an pape s.
In his pape we p opose a p ocess o gene a ing a classi ie based on he ini ially se o ele an
pape s. Fi s we collec simila sequences using an alignmen algo i hm like Blas . We hen use
he enla ges se o pape s o cons uc a classi ie . Finally we use ha classi ie o au oma ically
enla ge he se o ele an pape s by sea ching he MEDLINE using he au oma ically
cons uc ed classi ie . We ha e empi ically e alua ed ou p oposal and epo e y p omising
esul s.
Keywo ds: MEDLINE, Classi ica ion, In o ma ion Re ie al Sys em, Machine Lea ning,
Ensemble Algo i hms, Bioin o ma ics
INTRODUCTION
Molecula Biology and Biomedicine scien i ic publica ions a e a ailable (a leas he abs ac s) in
Medical Li e a u e Analysis and Re ie al Sys em On-line (MEDLINE). MEDLINE is he U.S.
Na ional Lib a y o Medicine (NLM), p emie bibliog aphic da abase: con ains o e 16 million
e e ences o jou nal a icles in li e sciences wi h a concen a ion on Biomedicine. A dis inc i e
ea u e o MEDLINE is ha he eco ds a e indexed wi h NLM’s Medical Subjec Headings
(MeSH e ms). MEDLINE is he majo componen o PubMed (Wheele e al., 2006), a da abase
o ci a ions o he NLM. PubMed comp ises mo e han 19 million ci a ions o biomedical
a icles om MEDLINE and li e science jou nals. The PubMed da abase main ained by he
Na ional Cen e o Bio echnology In o ma ion (NCBI) is a key esou ce o biomedical science,
and is ou i s base o wo k. The NCBIs PubMed sys em is a widely used me hod o accessing
MEDLINE.
The esul o a MEDLINE/PubMed sea ch is a lis o ci a ions (including au ho s, i le, jou nal
name, pape abs ac , keywo ds and MeSH e ms) o jou nal a icles. The esul o such sea ch is,
qui e o en, a huge amoun o documen s, making i e y ha d o esea che s o e icien ly each
he mos ele an documen s. As his is a e y ele an and ac ual opic o in es iga ion we
assess he use o Machine Lea ning-based ex classi ica ion echniques o help in he
iden i ica ion o a easonable amoun o ele an documen s in MEDLINE. The co e o he
epo ed wo k is o s udy he bes way o cons uc he da a se s and he classi ie s om he
s a ing se o sequences.
These expe iences we e done using a se o posi i e examples associa ed o he
sequences/keywo ds gi en by he use and a se o nega i e examples which is he ocuses o his
pape . The nega i e examples we e gene a ed in h ee di e en ways and we in end o show
which is he bes app oach o ou classi ica ion pu pose. In ou expe imen s we ha e used
se e al classi ica ion algo i hms a ailable in he WEKA (Hall e al., 2009) ool including
ensemble algo i hms. We ha e also made some sensi i i y es s o he p uning o a ibu es o
a ibu e educ ion.
The es o he pape is s uc u ed as ollows. The sec ion “An A chi ec u e o an In o ma ion
Re ie al Sys em” p esen s he a chi ec u e o ou in o ma ion e ie al sys em and he
ollowing Sec ion “The Local Da a base” desc ibes he local da a base cons uc ion p ocess and
he p e-p ocessing echniques used. Follows he ela ed wo k and a sec ion dedica ed o he
“Au oma ic Cons uc ion o da a se s” ha desc ibes he di e en al e na i es p oposed o da a
se cons uc ion and he expe iences we ha e done wi h di e en classi ie s. We also include a
sec ion dedica ed o classi ie ensemble whe e we p esen ou expe imen s using classi ie
ensemble me hods and inally we conclude he pape .
AN ARCHITECTURE FOR AN INFORMATION RETRIEVAL SYSTEM
The o e all goal o ou wo k is o implemen a web based sea ch ool ha ecei es a se o
genomic o p o eomic sequences and e u ns an o de ed se o pape s ele an o he s udy o
such sequences. The ini ial se o sequences is supplied by a biologis oge he wi h a se o
ele an keywo ds and an e- alue1. These h ee i ems a e he inpu o BioTex Re ie e
(Gonçal es & Camacho and Oli ei a, 2011) as can be seen in Figu e 1. Figu e 1 p esen s a
summa y o ou app oach ha we will now desc ibe in de ail. In he ollowing desc ip ion we use
NCBI as he sequence Da a Base.
1 A e- alue is a s a is ic o es ima e he signi icance o a ma ch be ween 2 sequences
Figu e 1. Sequence o s eps execu ed by BioTex Re ie e when he use p o ides a se o ini ial
DNA/p o ein sequences.
In S ep 1, he use (a biologis esea che ) p o ides an ini ial o sequences, op ionally a lis o
keywo ds, and an e- alue. Wi h hese h ee i ems (sequences, keywo ds and e- alue) and using
he NCBI BLAST ool we collec a se o simila sequences oge he wi h he pape e e ences
associa ed o hem. We could also use Ensembl wi h he same inpu s because Ensembl may
e u n a di e en se o pape s e e ences. Howe e o he p oposed wo k we ha e only used he
NCBI da abase.
Wi h his lis o pape ci a ions we sea ch o hei abs ac s in a local copy o MEDLINE (LDB
– Local Da abase) (S ep 3). Fo his we ha e p e iously p ep ocessed MEDLINE. S ep 3
sea ches and collec s he ollowing in o ma ion in he p e-p ocessed local copy o MEDLINE:
pmid, jou nal i le, jou nal ISSN, a icle i le, abs ac , lis o au ho s, lis o keywo ds, lis o
MeSH e ms and publica ion da e.
Fo he scope o his pape we a e conside ing only he pape ci a ions ha ha e he abs ac
a ailable in MEDLINE. A e S ep 3 we ha e a da a se o pape s ela ed o he sequences. We
will ake his se o pape s as he posi i e examples o he ull cons uc ion o he da a se (S ep
4) bu we need o ge some nega i e examples. To ob ain he nega i e examples we ha e h ee
possible app oaches.
Thus his s ep is explained in de ail in Sec ion III.
The ollowing s ep, S ep 5, is one o he mos impo an s eps o ou wo k which is o Cons uc
a Classi ie using Machine Lea ning echniques ha is explained in he nex sec ion. As a esul
o his s ep we ha e a ull lis o a icles conside ed ele an by ou classi ie (S ep 6). Howe e ,
we need o p esen hem o he biologis in an o de ed ashion way. So S ep 7 p esen s an o de ed
lis o ele an a icles o
he biologis . He e we will de elop and implemen a anking algo i hm based on ea u es such as
he numbe o ci a ions o he pape and he impac ac o o he jou nal/con e ence whe e i was
published. This pape ocuses on he cons uc ion o he da a se s highligh ing he esea ch om
he di e en app oaches o ob ain he nega i e examples. Acco ding o he igu e and o he
pu pose o his pape we ocus on S ep 4, al hough we ha e made a se o expe iences wi h some
classi ie s o conclude wha was he bes app oach.
THE LOCAL DATA BASE
We ha e downloaded 80GB (617 XML iles) o MEDLINE 2010 om he NCBI websi e. Each
XML ile has in o ma ion cha ac e izing one ci a ion. Among hese cha ac e is ics we ha e
conside ed he ollowing ones: PMID - he PubMed Iden i ie ; he PubMed Da e; he Jou nal
Ti le; he Jou nal ISSN ha co esponds o he ISI Web o Knowledge ISSN; he Ti le; he
Abs ac o he a icle i a ailable; he lis o he Au ho s; he MeSH Headings lis and he
Keywo ds lis . A e download he iles we e p ep ocessed as ollows.
P e-P ocessing MEDLINE XML iles
An independen s ep o ou ool is o main ain a local copy o MEDLINE, ha we will call Local
Da a Base (LDB). The LDB will enable e icien sea ch o he pape and will ha e ha ele an
in o ma ion o each pape in o ma adequa e, he algo i hm ha cons uc a chain as desc ibe
u he in his pape .
The i s s ep is o ead he XML iles and ex ac he ele an in o ma ion o s o e in he LDB.
A icle’s i le and abs ac a e p ep ocessed wi h “ adi ional” ex p e-p ocessing echniques.
Nex sec ion p esen s he p ep ocessing echniques applied.
P e-P ocessing Techniques
We ha e empi ically (Gonçal es & Gonçal es & Camacho & Oli ei a, 2010) e alua e which a e
he bes combina ion o p e-p ocessing echniques o achie e a be e accu acy. Based on his
p e ious s udy and wi h some mo e esea ch in he meanwhile we ha e used he ollowing
p ep ocessing echniques.
Documen Rep esen a ion
Fo each pape wi h he in o ma ion e e ed in he beginning o his sec ion Howe e he ex
ac s o a documen ( i le and abs ac ) a e il e ed using ex p ocessing echniques and
ep esen ed using he ec o space model om In o ma ion Re ie al whe e he alue o a e m
in a documen is gi en by he s anda d e m- equency in e se documen equency
(TFIDF=TF*IDF) unc ion (Zhou & Smalheise & Yu, 2006), o assign weigh s o each e m in
he documen .
TF is he equency o e m in documen
and
1
sin
log
+=
h e mcumen swi numbe o do
collec ioncumen numbe o do
IDF
Named En i y Recogni ion (NER)
NER is he ask o iden i ying e ms ha men ion a known en i y. We ha e used ABNER
(Se les, 2005), which s ands o A Biomedical Named En i y Recogni ion, ha is a so wa e ool
o molecula biology ha iden i ies en i ies in he biology domain: p o eins, RNA, DNA, cell
ype and cell line. Al hough we ha e implemen ed his echnique we ha e concluded ha he
iden i ica ion o NER e ms augmen s signi ican ly he numbe o a ibu es ins ead o educing
hem. We concluded ha he use o NER inc eases s ongly he numbe o e ms which is a
p oblem o he classi ie s. Thus we did no use NER in he p e-p ocessing phase.
Handling Synonyms
We handle synonyms using he Wo dNe (Fellbaum, 1998) o sea ch o simila e ms, in he
case o egula e ms, and used Gene On ology (Ashbu ne , 2000) o ind biological synonyms.
I wo wo ds mean he same hen hey a e synonyms, so hey could be eplaced by one o hem in
he en i e MEDLINE ( i le and abs ac ields) wi hou changing he seman ic meaning o he
e m hus educing he numbe o a ibu es. In his s ep we ha e eplaced all he synonyms
ound by one synonym e m hus educing he numbe o e ms.
Dic iona y Valida ion
A e m is conside ed a alid e m i i appea s in a ailable dic iona ies. We ha e ga he ed se e al
dic iona ies o he common English e ms ( such as Ispell and Wo dNe ) and o he medical and
biological e ms (BioLexicon (Rebholz-Schuhmann e al., 2008), The Hos o d Medical Te ms
Dic iona y (Hos o d, 2004) and Gene On ology (Ashbu ne , 2000). The Hos o d Medical Te ms
Dic iona y consis s o a ile ha con ains a long lis o medical e ms. BioLexicon is a la ge-scale
e minological esou ce de eloped o add ess ex mining equi emen s in he biomedical
domain. The BioLexicon is publicly a ailable bo h as an XML- o ma ed e m eposi o y and as
a ela ional da abase (MySQL) and i adhe es o he LMF ISO s anda ds o lexical esou ces.
We ha e also used he Gene On ology a ailable iles ha a e ela ed o genes, enzymes,
chemical esou ces, species and p o eins. We ha e p ocessed each o hese esou ce iles in o de
o ha e a simple ex ile wi h one e m pe line.
Ou app oach is in he sense ha i a e m appea s in one o hese dic iona ies i is a alid e m,
o he wise, i is no a alid e m, so we emo e i om he collec ion o e ms.
The applica ion o hese echnique is undamen al in a ibu e educ ion once a lo o e ms ha
ha e no biology, medical and no mal signi icance a e disca ded.
S op Wo ds Remo al
S op Wo ds Remo al emo es wo ds ha a e meaningless such as a icles, conjunc ion and
p eposi ions (e.g., a, he, a , e c.). These wo ds a e meaningless o he e alua ion o he
documen con en . We ha e used a se o 659 s op wo ds ile.
Tokeniza ion
Tokeniza ion is he p ocess o b eaking a ex in o okens. A oken is a non emp y sequence o
cha ac e s, excluding spaces and punc ua ion.
Special Cha ac e s Remo al
Special cha ac e emo al emo es all he special cha ac e s (+, -, !, ?, ., ,, ;, :, , g, =, &, #, %, $,
[, ], /, <, >, n, “, ”, j) and digi s.
S emming
S emming is he p ocess o emo ing in lec ional a ixes o wo ds educing he wo ds o hei
s em ( he wo ds compu e , compu ing and compu a ion a e all ans o med in o compu , which
means ha h ee di e en e ms a e ans o med in o only one e m hus educing he numbe o
a ibu es. We implemen ed he Po e ’s S emme Algo i hm (Po e , 1997).
P uning
Using p uning we disca d in he documen s collec ion e ms ha ei he appea oo a ely o oo
equen ly.

RELATED WORK
The e a e some wo k being done on biological and biomedical documen classi ica ion. Some o
hem applied o MEDLINE documen classi ica ion and o he da abases.
The wo k o (Sehgal, 2011) ies o au oma e he p ocess o adding new in o ma ion o TCDB
da abase (T anspo Classi ica ion Da abase) ha is a web ee access da abase
(h p://www. cdb.o g) abou comp ehensi e in o ma ion on anspo p o eins. The au ho s
es ic ed hemsel es o he
documen s in MEDLINE. The main goal is o highligh he use o Machine Lea ning echniques
ou pe o ms ules c ea ed by hand by a human expe . To ain he classi ie hey ha e used a se
o MEDLINE documen s e e ed TCDB as posi i e examples and ha e selec ed andomly also
om MEDLINE a se o nega i e examples.
The au ho s in (Imambi & Sudha, 2011) desc ibe a new model o ex classi ica ion using
es ima ing e m weigh s which imp o es accu acy classi ica ion acco ding o he au ho s
expe iences. Documen s a e ep esen ed as ec o s o e ms wi h hei no malized global
equency. Global weigh s a e unc ions ha coun how many imes a e m appea s in he en i e
collec ion and he no maliza ion p ocess compensa es he disc epancies in he leng hs o he
documen s. They ha e
used 1000 documen s om PubMed; 600 documen s o he aining da a se and 400 o he es
da a se . All hese documen s belong o ou ca ego ies wi h MeSH e ms ela ed o Diabe es
meli us. The au ho s compa e he di e en weigh ing me hods: local-bina y, local-log, local d
and global ele an . They concluded in his s udy ha global ele an weigh ing me hod achie es
a highe p ecision. In ou own wo k we ha e also used all no malized global equency.
BioQSpace (Di oli e al., 2005) is a GUI whe e use s can que y abs ac s om PubMed using an
embedded sea ch acili y. BioQSpace pe o ms pai wise simila i y calcula ions be ween all he
abs ac s based on a se o indi idual a ibu es namely: s uc u e, unc ion, disease and
he apeu ic compounds wo d lis ob ained om MeSH e ms, wo d usage, PubMed ela ed
a icles, publica ion da e among o he s. These a ibu es a e gi en mo e o less impo ance
acco ding o he weigh a ibu ed by use s. A clus e ing algo i hm is used o g oup abs ac s ha
a e e y simila .
(F unza & Inkpen & T an, 2011) desc ibe a me hodology o build an applica ion capable o
iden i ying and dissemina ing heal h ca e in o ma ion using a Machine Lea ning app oach. The
main objec i e o hei wo k is o s udy he bes in o ma ion ep esen a ion model and wha
classi ica ion algo i hms a e sui able o classi ying ele an medical in o ma ion in sho ex s.
They ha e used 6 di e en Machine Lea ning algo i hms. The au ho s concluded ha nai e
Bayes pe o med e y well on sho ex s in he medical domain and ha adaboos had he wo s
esul .
In (Dollah & Seddiqui & Aono, 2010) he au ho s p esen an app oach o classi ying a
collec ion o biomedical abs ac s downloaded om MEDLINE da abase wi h he help o
on ology alignmen . Al hough his wo k classi ies MEDLINE documen s i is based on on ology
alignmen which is ou o ou scope.
Lige Ca (Sa ka e al., 2009) s ands o Li e a u e and Genomic Elec onic Resou ce Ca alogue,
and i is a sys em o explo ing biomedical li e a u e h ough he selec ion o e ms wi hin a
MeSH cloud ha is gene a ed based on an ini ial que y using jou nal, a icle, o gene da a. The
cen al idea o Lige Ca is o c ea e a ag cloud showing an o e iew o impo an concep s and
ends associa ed o he MeSH desc ip o s. Lige Ca agg ega es mul iple a icles in PubMed,
combining
he associa ed MeSH desc ip o s in o a cloud, weigh ed by equency. Lige Ca does no apply
any Machine Lea ning echniques o pape classi ica ion as we p esen in ou s udy.
AUTOMATIC CONSTRUCTION OF DATA SETS
Cons uc ing he Da a Se s
In o de o sol e ou p oblem ha is gi en a se o genomic o p o eomic sequences e u n a se
o ela ed sequences and pape s wi h ele an in o ma ion o he s udy o such sequences, we
need o ob ain i s o all he a icles associa ed wi h he gi en se o sequences and cons uc he
da a se o gi e o he classi ie . Figu e 2 illus a es sequence o s eps included in
BioTex Re i e . The empi ical wo k epea ed in he pape conce ns he cons uc ion o a da a
se (S ep 4).
Figu e 2. Da a Se Cons uc ion
The inpu o ou wo k is a se o sequences gi en in he FASTA o ma . We use he
ne blas -2.2.22 ool ha pe o m a emo e blas sea ch a he NCBI si e. We ha e embedded his
applica ion in o ou code and au oma ically ha e access o bo h, he o iginal sequences and he
se o simila sequences e ie ed by BLAST.
These esul s show us he simila sequences and he e alue associa ed wi h each o he e ie ed
sequence. The e- alue is a s a is ic o es ima e he signi icance o a ma ch be ween wo
sequences. The e- alue is an inpu ha is gi en us by he biologis . We elax his h eshold alue
in o de o ob ain he nega i e examples based on he e- alue as we can see in Figu e 3. The
posi i e examples a e he one’s ha a e lowe han he e- alue p e iously speci ied by he
biologis . We es ablish a “no man’s land“ zone and a e ha zone we collec he nega i e
examples.
Figu e 3. How posi i e and nea -miss (nega i e) examples a e ob ained. e is he e- alue h eshold o
ob ain he posi i e examples. α and β a e pa ame e s o he cu o o he nega i e examples.
The posi i e examples a e he se o pape s associa ed wi h he se o sequences wi h e- alue
below he espec i e h eshold. In his s udy we ha e empi ically e alua ed h ee di e en ways
o ob aining he nega i es examples. We now explain he al e na i es.
Nea -Miss Values (NMV)
To ob ain he Nea -Miss Values (NMV) we collec he pape s associa ed wi h he simila
sequences ha ha e e alue abo e he h eshold bu close o ha . In Figu e 3 he e is a s ip g ay
o be e disc imina e wha a e posi i e examples and nega i e examples. The examples in he
igh mos box con ains nea -miss nega i e examples because hey a e no posi i es bu ha e a
ce ain deg ee o simila i y wi h he sequences. This wo ks on he examples ha ha e a
minimum numbe o nega i e examples. I we do no ha e any nega i e examples wi h his
app oach, o i he nega i e examples a e ew, we can ollow one o he ollowing app oaches: o
use MeSH Random Values o o use Random Values. In ou expe imen s we ha e conside ed
e- alue = 0.001 and we ha e elaxed i o 1, 2 and 5 as we can see in sequences dis ibu ion
ables in he nex Sec ion. We ha e elaxed o hese di e en alues o ob ain mo e nega i e
examples. The a icles associa ed o he simila sequences wi h e- alue less hen 0.001 a e
conside ed posi i e; he a icles associa ed o he simila sequences ha ha e e- alues g ea e
hen 0.001 and e- alues less hen 0.001 plus 10% (_ = 10%) a e conside ed in he g ay s ip so
hey a e no conside ed posi i e o nega i e; he a icles associa ed wi h o he simila sequences
wi h e- alues g ea e hen 0.001 plus 10% a e conside ed nega i e examples (nea -miss alues).
MeSH Random Values (MRV)
This al e na i e o gene a e nega i es is adop ed when we do no ha e su icien numbe o
nega i e examples o he classi ie o lea n. The nega i e examples a e ob ained combining he
nea miss alues, i hey exis , wi h some andom examples gene a ed om he LDB. Bu , hese
MRV examples mus ha e he maximum numbe o MeSH e ms om he posi i e examples. A
he end he numbe o nega i e examples is equal o he numbe o posi i e examples.
Random Values (RV)
The las app oach is o gene a e jus andomly he nega i e examples om ou LDB in a numbe
equal o he numbe o posi i e examples. We gua an ee ha in his se he e is no posi i e
example.
COMPARING THE ALTERNATIVES TO DATA SET CONSTRUCTION
Da a Se Cha ac e iza ion
Fo , his s udy we ha e gene a ed se e al da a se s based on sequences ha belong o six
di e en classes, wi h he ollowing dis ibu ion:
•RNASES: 2 sequences
•Esche ichia Coli: 5 sequences
•Choles e ol: 5 sequences
•Hemoglobin: 5 sequences
•Blood P essu e: 5 sequences
•Alzheime : 5 sequences
We ha e also used h ee di e en elaxa ion alues o he e- alue (1, 2 and 5). I he use en e s
an e- alue o 0.001, hen he posi i e examples a e he ones ha ha e e- alue less o equal o
0.001. And he nega i e examples a e he one’s g ea e hen 0.001. Bu as we can see in Figu e 3
we lea e a g ay s ip o be e sepa a e he posi i e om he nega i e examples. This s ip is also
de ined by he use . Fo hese examples we ha e de ined a s ip o 10% o he numbe o no
simila sequences. So he nega i e nea -miss examples a e he one’s ha a e g ea e hen 0.001
plus 10% o he o he numbe o no simila sequences and lowe han e- alue elaxa ion alue
(1, 2 o 5 in ou examples).
The main idea o his s udy is o s udy he bes way o cons uc he nega i e examples based on
ou expe iences. The dis ibu ions o posi i e and nega i e examples a e show in he Appendix.
Expe imen al Resul s
In ou expe iences we ha e used a se o algo i hms a ailable in he WEKA (Hall e al., 2009)
ools and ha a e lis ed Table 1.
Ac onym Algo i hm Type
Ze oR Majo i y p edic o Rule lea ne
smo Sequen ial Minimal Op imiza ion Suppo Vec o Machines
Random Fo es Ensemble
ibk K-nea es neighbo s Ins ance-based lea ne
BayesNe Bayesan Ne wo k Bayes lea ne
j48 Decision ee (C4.5) Decision ee lea ne
d nb Decision able / naï e bayes hyb id Rule lea ne
AdaBoos Boos ing algo i hm Ensemble lea ne
Bagging Bagging algo i hm Ensemble lea ne
Ensemble Selec ion Combines se e al algo i hms Ensemble lea ne
Table 1. Machine Lea ning Algo i hms used in he s udy.
The da a se s used a e cha ac e ized in Tables 2 and 3. Table 2 cha ac e izes da a se s o which
he nega i e examples a e made only o nea miss examples. Table 3 cha ac e izes he da a se s
o which he e we e no enough nega i e examples and he e o e we ha e used he MRV and
RV s a egies.
Homayouni H. Hashemi S. Hamzeh A., A Lazy Ensemble Lea ning Me hod o
Classi ica ion, IJCSI In e na ional Jou nal o Compu e Science Issues, Vol. 7, Issue 5,
Sep embe 2010.
Hos o d medical e ms dic iona y 3.0, 2004.
Ind a N., Sa ka N., Schenk R., Mille H. and No on C.. Lige Ca : using ”MeSH Clouds” om
jou nal, a icle, o gene ci a ions o acili a e he iden i ica ion o ele an biomedical li e a u e.
AMIA - Annual Symposium p oceedings / AMIA Symposium. AMIA Symposium,
2009:563–567, 2009.
Imambi S. and Sudha T.. Classi ica ion o medline documen s using global ele an weighing
schema. In e na ional Jou nal o Compu e Applica ions, 16(3):45–48, Feb ua y 2011. Published
by Founda ion o Compu e Science.
Ko sian is S. and Pin elas P., Combining Bagging and Boos ing, In e na ional Jou nal o
Compu a ional In elligence, Vol. 1, No. 4 (324-333), 2004.
Opi z D. and Maclin R., Popula Ensemble Me hods: An Empi ical S udy, Jou nal o A i icial
In elligence Resea ch, ol.11, pp. 169-198, 1999.
Po e M. F.. An algo i hm o su ix s ipping. pages 313–316, 1997.
Rebholz-Schuhmann D., Pezik P., Lee V., Kim J-J, Del G a a R., Sasaki Y., McNaugh J.,
Mon emagni S., MonachiniM., Calzola i N. and Ananiadou S. . Biolexicon: Towa ds a e e ence
e minological esou ce in he biomedical domain. In P oceedings o he o he 16 h Annual
In e na ional Con e ence on In elligen Sys ems o Molecula Biology (ISMB- 2008), 2008.
Sehgal A. K., Sanmay D., No o K., Mil on H., Saie J . and Elkan C.. Iden i ying ele an da a
o a biological da abase: Handc a ed ules e sus machine lea ning. IEEE/ACM T ans.
Compu . Biology Bioin o m., 8(3):851–857, 2011.
Se les B.. Abne : an open sou ce ool o au oma ically agging genes, p o eins and o he en i y
names in ex . Bioin o ma ics, 21(14):3191–3192, 2005.
Wheele D. L., Ba e T., Benson D. A, B yan S. H., Canese K., Che e nin V., Chu ch D. M.,
Dicuccio M., Edga R., Fede hen S., Gee L. Y., Helmbe g W., Kapus in Y., Ken on D. L.,
Kho ayko O., Lipman D. J., Madden T. L., Maglo D. R., Os ell j., P ui K. D., Schule G. D.,
Sch iml L. M., Sequei a E., She y S. T., Si o kin K., Sou o o A., S a chenko G., Suzek T. O.,
Ta uso R., Ta uso a T. A., Wagne L., and Yaschenko E.. Da abase esou ces o he na ional
cen e o bio echnology in o ma ion. Nucleic Acids Res, 34(Da abase issue), Janua y 2006.
Zhou W., Smalheise N. R. and Yu C.. A u o ial on in o ma ion e ie al: basic e ms and
concep s. Jou nal o Biomedical Disco e y and Collabo a ion, 1:2, Ma ch 2006.