scieee Science in your language
[In] (orig)

Software for the Genetic Analysis Domain

Abstract

In this report we overview the state of the art in software for genetic analysis, starting from software tools for genetic analysis, moving on to software tools for genomic analysis, and ending with bioinformatic pipeline development environments.

Read accessible full text

Software for the Genetic Analysis Domain

Author: Villanueva del Pozo, María José,Valverde Giromé, Francisco,Pastor López, Oscar
Publisher: Universitat Politècnica de València
Year: 2015
Source: https://riunet.upv.es/bitstream/10251/57428/1/SoftwareForGeneticAnalysisDomainTR.pdf
In o me Técnico / Technical Repo
Re . #:
P oS-TR-XXXX
Ti le:
So wa e o he Gene ic Analysis Domain
Au ho (s):
Osca Pas o , F ancisco Val e de, and Ma ia Jose Villanue a
Co esponding
au ho (s):
opas o @dsic.up .es
al [email p o ec ed].es
m illanue [email p o ec ed].es
Documen e sion numbe :
Final e sion:
Pages:
Release da e:
Key wo ds:
So wa e o he Gene ic Analysis Domain
Osca Pas o , F ancisco Val e de, and Ma ia Jose Villanue a
1
Con en
So wa e ools ............................................................................................................................... 3
Gene ic Analysis ........................................................................................................................ 3
In oduc ion .......................................................................................................................... 3
Sequenche ............................................................................................................................ 4
SeqScape ............................................................................................................................... 6
Codon Code Aligne ............................................................................................................... 7
Mu a ion Su eyo ................................................................................................................ 8
Polyph ed ............................................................................................................................ 10
InSnp .................................................................................................................................... 12
Compa ison ......................................................................................................................... 13
Conclusion ........................................................................................................................... 15
Re e ences ........................................................................................................................... 15
Genomic Analysis .................................................................................................................... 16
In oduc ion ........................................................................................................................ 16
Re ie ing and anno a ing da a manually ........................................................................... 16
Bioma ................................................................................................................................ 17
VCFTools .............................................................................................................................. 18
Anno a ............................................................................................................................... 19
VEP ...................................................................................................................................... 20
SamTools ............................................................................................................................. 22
SNPE .................................................................................................................................. 22
GATK .................................................................................................................................... 24
Pipeline De elopmen En i onmen s.......................................................................................... 24
Biopy hon, BioPe l, Bioja a, Bio* ............................................................................................ 24
Ta e na .................................................................................................................................... 24
Repo .................................................................................................................................. 26
Wo k low 1: Diagen (Disease diagnosis BREAST Cance om a ia ion de ec ion) ........... 27
Galaxy ...................................................................................................................................... 32
Repo .................................................................................................................................. 33
Wo k low 1: Diagen (Disease diagnosis BREAST Cance om a ia ion de ec ion) ........... 33
EBioFlow .................................................................................................................................. 35
Repo .................................................................................................................................. 35
BSIS .......................................................................................................................................... 36
Repo .................................................................................................................................. 36
2
1 In oduc ion
The main pu pose ha p omo ed he esea ch o he exis en comme cial alignmen ools is o
lea n he die en unc ionali y ha his kind o ools a e oe ing. In o de o accomplish his
a ge will be necessa y o ob ain he die ences be ween hem and o nd he s onges poin s
and deciencies o each one; bu mos ly which unc ionali y may be missing in all o hem.
Ha ing a be e comp ehension abou wha has been al eady de eloped, wha is ac ually
being used and wha a e he needs o he use s o ha ools, could cla i y i he objec i es o
he p esen p ojec o he Genoma g oup could be a eal con ibu ion o he eld.
The s a poin o his p ojec had as a main objec i e o c ea e a so wa e ha mee s
he equi emen s o he biologis s when pe o ming a DNA analysis om a pa ien sample
sea ching o mu a ions ha may cause some disease. As a consequence he so wa e we in end
o de elop will y o co e all he ac i i ies in ol ed on he p ocess in o de o p o ide a
comple e unc ionali y ha his expe s claim all he comme cial ools lack.
Fi s o all i is essen ial o es ablish and delimi all he ac i i ies ha he men ioned p ocess
comp ises. Howe e , inside his collec ion o ac i i ies we will nd ha some o hem canno be
con olled by ou so wa e o canno be au oma ed by any so wa e; ac i i ies like he sequencing
he DNA sample and decision-making he co ec ness o he basecalling espec i ely.
The  s s ep while analyzing a sample o a pa ien is o ex ac a e y small agmen
and pe o m he sequencing. This ac i i y is done by he sequence machine and i is no an
ac i i y ha he so wa e will ake in o accoun . Howe e , he ou pu o his ac i i y, les in
.ABI (Applied Biosys ems Inc.) o ma , will be he inpu o he so wa e. The sequences, he
samples and he e e ence can be exp essed in his o ma bu also in he o ma s .SEQ, .GB
(GeneBank) o FASTA. Be o e ge ing deepe in de ails abou he analysis decomposi ion, i
is impo an o emphasize ha no mally he analysis is es ic ed in only one gene a a ime.
Mo eo e , due biological dicul ies on he sequencing p ocess he DNA sample is sequenced in
pieces abou 800 bp (called con igs). Fo his eason, some imes i is sequenced only he egion
o in e es o he analysis, bu a leas i is sequenced wice (one e e se and o he o wa d
di ec ion) in o de o ensu e he accu acy o he esul . F om bo h sequences i is ob ained a
consensus sequence.
DNA sequence analysis ha he new so wa e will ha e o manage consis s on se e al phases:
1. Assembly each con ig in o i s co ec posi ion: Independen ly o he numbe o con igs
sequenced each one has o be loca ed p ope ly.
2. Clean he esul s om he sequence machine: Using he ool and i s biological knowledge
biologis s pe o m manually a cleaning on he basecalling o he con igs. Expe s ha e
o decide, aided wi h he e e ence sequence, i he basecalling o each sample and he
consensus sequence eec s he e aci y o he o iginal sample, co ec ing hus all he
e o s ha he sequence machine has in oduced. The ool will ha e o pe o m as well a
d opping o he beginnings and ends o sequenced con igs ha no mally a e no use ul o
analysis because o he quali y o he signal. This las cleaning is e e ed as  imming.
3. Compa e consensus wi h he e e ence sequence sea ching o a ia ions: Each die ence
in he consensus sequence espec he e e ence sequence will be conside ed as a a ia ion.
The ool will ha e o sea ch o inse ions, dele ions and indels. He e ozygosis ( wo die -
en signals codi ying wo bases in he same posi ion) is also impo an o he de ec ion o
a ia ions.
4. Sea ch wha does mean each a ia ion: One a ia ion may be p o oking some pheno ype
depending on he base changes and he posi ion o hem. Fo each possible a ia ion-
pheno ype pai i would be c ucial o know he  s publica ion ha suppo s he nding.
3
In his e iew he e has been analyzed he ollowing ools: Sequenche , SeqScape, Mu a ion
Su eyo , CodonCodeAligne , Polyph ed and InSNP. The ollowing 6 sec ions will analyze he
essen ials o each ool conce ning he ins alla ion and use o he ool, he a ia ions de ec ed
in compa ison wi h he concep ual model, he possible connec ion wi h bibliog aphy and some
o he in e es ing ea u es. The las wo sec ions will con ain he compa ison be ween ools and
he conclusion and ecommenda ions o he au ho o his e iew. In he anex i is explained a
 s con ac abou he beha iou o all ools unde he same condi ions.
2 Sequenche
Sequenche is a p op ie a y So wa e o Gene Codes Co po a ion [3] ha oe s he possibili y o
in oduce sequences, pe o m assemblies, he explo a ion and edi ion o sequences in connec ion
wi h i s e e ence and nally he de ec ion o he a ia ions ha die om his e e ence.
Gene Codes Co po a ion oe s a Demo e sion wi h es ic ed unc ionali y and he possi-
bili y o pu chase all he so wa e wi h a license ha once ob ained i ne e expi es.
I is a ailable o MAC and Windows and i is possible o ins all i on a single compu e o
ins all he ne wo k-enabled e sion, wi h a se e and i s clien s.
In o de o help he use s o amilia ize wi h he ool he e is a collec ion o easy and comple e
u o ials on i s webpage (
h p://www.genecodes.com
).
Figu e 1: Die en windows on Sequenche
Sequenche p o ides die en iews o he da a (See Figu e 1): les assembly (le and up
window), p o iding he lis o les con aining he sequences assembled in he same con ig; dis-
ibu ion o con igs a ound he gene ( igh and up window), displaying he e e ence sequence
4
(blue line) and all he con igs si ua ed in i s posi ion (g een lines o o wa d and ed lines o
e e se); disco e ed a ia ions epo ( igh and down window), desc ibing he ound changes
in se e al elds inside a able; bases o each sequence (middle window), si ua ing each sequence
sequen ially and he ch oma og ams o he samples (le and down window). When a a i-
ion is double clicked he bases in ol ed a e highligh ed in he bases ep esen a ion and in he
ch oma og ams.
2.1 Types o de ec ed mu a ions
Inse ions
,
dele ions
and
indels
in homozygosis a e de ec ed and all o hem a e exp essed base
by base, as a ia ions o leng h 1. I , o example, he e is an inse ion o 3 nucleo ides, he
a ia ion is exp essed as 3 inse ions o 1 nucleo ide.
Howe e ,
inse ions
and
dele ions
in he e ozygosis canno be de ec ed. When hey occu ,
he ch oma og ams o he con igs seem a mix o wo signals ha should be almos iden ical bu
now appea shi ed se e al posi ions. This so wa e de ec s
indels
in he e ozygosis because his
shi does no happen.
In o de o iden i y a he e ozygosis base ( wo die en alues in he same posi ion) and die -
en ia e i om a homozygosis one wi h noise, a sensibili y alue can be xed. This sensibili y is
exp essed in pe cen age and when he p esence o wo signican o e lapping uo escence peaks
occu s in a conc e e posi ion, i ep esen s he ela ion be ween he signal o lowe in ensi y
espec he o he .
2.2 Re u n o ma s
All he sequences included he consensus sequence (once all con igs a e assembled and cleaned)
can be impo ed and expo ed in se e al o ma s: in plain ex as ASCII plain o ma o un o -
ma ed (bases only); in specialis and legacy o ma s like AFDIL, Genen ech, IG and S ide ;
in da abases o ma s as GenBank, NBRF and EMBL; in commonly used o ma s like FASTA
(no mal and conca ena ed) and GCG; in phylogenic p og am o ma s like NEXUS/PAUP (in-
e lie ed and sequen ial), Phylip (no mal, 3 and 4) and e en in s anda d ch oma og am o ma s
like SCF (2.0 and 3.0). Fo some o hose o ma s can be indica ed as well he ollowing op ions:
expo wi h uppe /lowe case, selec he o ien a ion o he sequence and lea e o emo e he
gaps.
When he e a e a ia ions in he sample hey a e epo ed in a able. The epo ga he s he
da a abou posi ion and alue o he base on he e e ence, alue o he base on he sample and
numbe o die en samples ha ha e his a ia ion. This a ia ion able can only be expo ed
in TXT and PDF.
2.3 Connec ion wi h bibliog aphy
I is no possible o make any connec ion wi h da a abou a ia ions nei he in he e e ence
sequence no he a ia ions able.
2.4 O he in e es ing ea u es
Du ing he impo and assembly o con igs Sequenche oe s he op ion o classi y he samples
au oma ically only using he name o he le. The samples can be di ided in o wa d and e e se
and a he same ime in die en specimen. I also allows he simul aneous analysis o se e al
pa ien s and all he die ences among hem would be eec ed in he epo s.
5

Figu e 2: SeqScape window
3 SeqScape
SeqScape is a p op ie a y So wa e o Applied Biosys ems [1] wi hou demo a ailable. I s unc-
ionali y goes om impo ing, assembling, edi ing and analyzing samples sea ching o a ia ions
un il compa ison o he segmen s sea ching o pa e ns.
I is a ailable only o Windows and can be congu ed a manage access secu i y con ol wi h
he use o logins and passwo ds allowing h ee die en pe mission p oles. In addi ion o he
secu i y, i can p epa e documen a ion o u u e audi ion, ha is, i is possible o p og am
some e en s o be eco ded.
SeqScape changes he me hodology o ope a ion om ac ions-objec i e, whe e o each ob-
jec i e ha he use wan o achie e has o pe o m se e al conc e e ac ions, o congu a ion-
objec i es, whe e i is necessa y o congu e all he needed op ions be o e pe o ming any ac ion
and once all congu e all o he objec i es a e execu ed oge he . Then i is necessa y o lea n
how o congu e he p ojec s. I has se e al impo an congu a ion phases ha nally will
lead o one bu on ha execu es all he unc ionali y pe o med in one s ep. Then he use will
sea ch o he da a ha wan s o know in each momen (See Figu e 2).
3.1 Types o de ec ed mu a ions
Inse ions
,
dele ions
and
indels
in homozygosis a e de ec ed and SeqScape claims ha i also
de ec s
inse ions
and
dele ions
in he e ozygosis since he e sion 2.5. Howe e he e sion
p o ide by IMEGEN is he 2.0 and his kind o mu a ions can no be loca ed. In e sion 2.0
when hey occu hese con igs a e no included in he assembled ones. S ill
indels
in he e ozygosis
a e de ec ed and i is possible o x a sensibili y alue like was possible in Sequenche .
6
3.2 Re u n o ma s
The consensus sequence can be expo ed as FASTA, SEQ and QUAL o ma . I has he op ion
o eplacing unknown bases wi h a desi ed symbol, in case o gaps o bad quali y bases.
The o ma s o expo ing he epo s abou he a ia ions a e TXT, HTML, PDF o XML.
The expo ed les will con ain some analysis a iables o each specimen (success o analysis,
specimen sco e, mu a ions ound) and in o ma ion abou a ia ions (sample, posi ion, size).
3.3 Connec ion wi h bibliog aphy
I is possible o add known a ia ions o he e e ence sequence in o de o he SeqScape o be
able o iden i y hem as known o unknown. I is also possible o au oma e his in oduc ion o
da a by c ea ing an XLS le ha will con ain all he elds needed o desc ibe each a ia ion.
The a ia ions in oduced can be classi y as inse ions, dele ions, o basechange, indica ing he
ROI, he posi ion, he e e ence base(s), he a ian base(s) and i s desc ip ion.
3.4 In e es ing ea u es
In addi ion o all he secu i y con ol a ound he access i is as well possible o expo da a
signed elec onically.
When he e e ence sequence has been impo ed om GeneBank, SeqScape c ea es egions
o in e es and laye s ha sepa a es he die en exons au oma ically.
In addi ion i exis s he possibili y o c ea e lib a ies sea ching o pa e ns. A lib a y is a
collec ion o se e al segmen s (alleles, geno ypes and haplo ypes) wi h a xed leng h o a egion
o in e es (ROI). The consensus sequence is compa ed wi h he lib a y sea ching o ma ches
on any segmen .
4 CodonCodeAligne
4.1 Types o de ec ed a ia ions
CodonCodeAligne is p opie a y so wa e om CodonCode Co po a ion [2] used o sequence
assembly, con ig edi ing, and mu a ion de ec ion.
CodonCode Co po a ionI makes a ailable a 30 day-demo ha can be yed unde Windows
and Mac.
I s in e ace oe s a isualiza ion (See Figu e 3) o he assembled con igs and he e e -
ence sequence simul aneously wi h he codi ying egions o he gene o (depending o he use
p e e ences) he isualiza ion o he die ences be ween bases.
4.2 Re u n o ma s
CodonCodeAligne can expo he samples wi h he o ma s FASTA and SCF bu he consensus
sequence only in FASTA (No mal bases o haplo ypes). Fo bo h o hem exis s he op ions o
including gaps, append he commen s, eplace p oblem cha ac e s in names and w i e FASTA
quali y les. I he a ge sequences a e he disposi ion inside he whole assembly, i is possible
o sa e i in o an ACE p ojec , a NEXUS/PAUD (in e lea ed and sequen ial) o a Phylip
(in e lea ed and sequen ial) o ma .
The a ia ions ound a e ga he ed in a epo ha can be expo ed in TXT and PDF. This
epo will con ain he ea u e, he sou ce, he ype o sou ce whe e he mu a ion has been
ounded, he pa en con ig, he s a , he end, and he con en .
7
Figu e 3: CodonCode Aligne
4.3 Connec ion wi h bibliog aphy
The e is no op ion o add any in o ma ion abou a ia ions, bibliog aphy o pheno ypes.
4.4 In e es ing Fea u es
G aphical in e ace achie es he comple e na iga ion along sequences, edi ion o sequences al-
lowing all ypes o ope a ions and isualiza ion o all equi ed da a simul aneously.
Fu he mo e he e e ence sequence can be downloaded au oma ically om GeneBank by
only indica ing he accession numbe .
5 Mu a ion Su eyo
Mu a ion Su eyo is a p op ie a y So wa e o So Gene ics [8] ha compa es se e al samples
wi h a e e ence sequence sea ching o a ia ions.
So Gene ics oe s a Demo e sion, he possibili y o ying a ully unc ionali y 30 days ail
and a ee aining o gene ic expe s ha will be he use s o he ool.
The so wa e is a ailable o Windows (NT, 2000, XP, Vis a and 7) and MAC (al hough i
is necessa y a le con e e o PC les o hose les ha will be inpu o he so wa e) and i is
possible o ins all i on a single compu e o ins all he ne wo k-enabled e sion (wi h a se e
and i s clien s).
The use -in e ace has been designed ollowing Mic oso pla o m design guides o be a use -
iendly so wa e (See Figu e 4). Die en iews compose he s uc u e o he in e ace: he ex
iew ( igh ) and he g aphical iew (le ), na iga ing easily along hem.
The e is an addi ional so wa e called Mu a ion Su eyo Au o un ha ha pe mi s he
una ended analysis o mul iple p ojec s and a Log File Edi o o congu e he pa ame e s o
he hole una ended p ocess.
8
Figu e 4: Mu a ion Su eyo iews
5.1 Types o de ec ed mu a ions
Inse ions
,
dele ions
and
indels
homozygosis a e de ec ed and all o hem a e exp essed wi h
i s co ec leng h. A e also de ec ed
indels
in he e ozygosis and a sensibili y alue, now called
d opping ac o , can be xed o nd he he e ozygosis bases.
Rega ding
inse ions
and
dele ions
in he e ozygosis, i has he abili y o iden i y he e ozy-
gous indels down o 5% o he p ima y peak. The sample is decomposed in o wo die en
samples ep esen ing bo h s ands (one om he a he , one o he mo he ). Then one o hose
is shi ed and displayed below acco ding he mu a ion de ec ed in o de o ma ch he e e ence
sequence. Some imes due he na u e o his kind o a ia ions is dicul o he so wa e o
iden i y au oma ically he s a poin . Fo his eason, i allows he expe o e i y i he
ob ained posi ion is he co ec one.
5.2 Re u n o ma s
I is no possible o expo he consensus sequence because his ool does no c ea e any due
he ac i sea ches o a ia ions di ec ly in each o wa d and e e se sample sepa a ely.
Pe con a, he e a e a lo o possibili ies o expo and congu e epo s abou he a ia ions
ound. An s anda d epo could con ain he ex a da a: numbe , sample and e e ence le
names, di ec ion o he GeneBank e e ence sequence, eading ame, s a and end o he
sample, quali y, a mu a ion code o each one ound and se e al mo e in o ma ion. All his
in o ma ion can be expo ed in TXT, XLS, HTML o XML.
The able wi h he mu a ions appea ed on each sample can be expo ed only in a TXT le.
9
Genomic Analysis
In oduc ion
The genomic analysis add esses h ee s eps: 1) i s akes as a base a comple e VCF (con aining
one indi idual o iple s); 2) hen adds all ele an in o ma ion o he VCF ile (anno a ion
p ocess) and inally 3) depending on he analysis he VCF anno a ed is il e ed acco ding some
c i e ia:
The ele an in o ma ion hey wan o anno a e is: 1) S uc u al Ids (no mally om ENSEMBL)
(and hg s no a ion); 2) Snp ids ( s om dbSNP); 3) Popula ion Allele equency ( om 1000G da a
o Exome Sequencing P ojec ); 4) Co e age; 5) Pai E ec p edic ion-T ansc ip (Algo i hm: SIFT,
POLYPHEN o combined, and ansc ip s: all ansc ip s ha a e a ec ed, ansc ip o he issue
whe e i exp esses he mos damaging consequence, mos impo an ansc ip o disease); and
6) combined analysis wi h iple s (Phasing and Combined he e ocigosis o se e al a ia ions
ha a ec a gene le el).
And he di e en il e ing c i e ia a e: 1) S uc u al: Ch , Gene, Type o a ia ion (ins, del, indel,
MNPs, SNPs and CNVs) ype o polymo phism (homozygosis, he e ozygosis, combined
he e ozygosis, de-no o); 2) Posi ion ange; 3) Allele equency; 4) Co e age; 5) T ansc ip ; 6)
Loss o unc ion.
Anno a ions can be classi ied in h ee ca ego ies: 1) s uc u al in o ma ion o he a ia ion, such
as gene, exon, ansc ip s; 2) da abase in o ma ion, such as he s om dbSNP; and 3) e ec s in
di e en ansc ip s, including sco es o SIFT and POLYPHEN and he hg s no a ion.
In o de o e ie e hese da a and anno a e he a ia ion ile, se e al op ions a e a ailable:
• Download da abase ile manually –in GVF, GFF, BED o ma s- o using bioma (in TXT
ile) and anno a e he ile wi h his da a a e wa ds.
• Run he sui able commands o he sui es Anno a , SnpE and/o VEP.
Re ie ing and anno a ing da a manually
A1. The GFF Fo ma
The GFF o ma (Gene ic Fea u e Fo ma Ve sion) speci ies genomic/gene ic ea u es (genes,
exons, CDS, e c.) and hei p ope ies (name, sequence, dbx e , e c.) using some p ede ined
ields and ules and also on ology e ms. I is widely used by he communi y o anno a e
a ia ions wi h s uc u al da a.
Example (cu en e sion 3):
0 ##g - e sion 3
1 ##sequence- egion c g123 1 1497228
2 c g123 . gene 1000 9000 . + . ID=gene00001;Name=EDEN
3 c g123 . TF_binding_si e 1000 1012 . + . ID= bs00001;Pa en =gene00001
4 c g123 . mRNA 1050 9000 . + . ID=mRNA00001;Pa en =gene00001;
5 c g123 . mRNA 1050 9000 . + . ID=mRNA00002;Pa en =gene00001;
6 c g123 . mRNA 1300 9000 . + . ID=mRNA00003;Pa en =gene00001;
7 c g123 . exon 1300 1500 . + . ID=exon00001;Pa en =mRNA00003
8 c g123 . exon 1050 1500 . + . ID=exon00002;Pa en =mRNA00001,mRNA00002
9 c g123 . exon 3000 3902 . + . ID=exon00003;Pa en =mRNA00001,mRNA00003
10 c g123 . exon 5000 5500 . + . ID=exon00004;Pa en =mRNA00001,mRNA00002,mRNA00003
11 c g123 . exon 7000 9000 . + . ID=exon00005;Pa en =mRNA00001,mRNA00002,mRNA00003
16

I can be downloaded om he NCBI p: e _GRCh37.p13_ op_le el.g 3.gz (No e: I choose his
ile obse ing he name and assuming ha i is he one ha con ains all he in o ma ion)
A2. The BED o ma
The BED o ma is de ined by he USCS also o desc ibe anno a ions.
b owse posi ion ch 7:127471196-127495720
b owse hide all
ack name="I emRGBDemo" desc ip ion="I em RGB demons a ion" isibili y=2
i emRgb="On"
ch 7 127471196 127472363 Pos1 0 + 127471196 127472363 255,0,0
ch 7 127472363 127473530 Pos2 0 + 127472363 127473530 255,0,0
ch 7 127473530 127474697 Pos3 0 + 127473530 127474697 255,0,0
ch 7 127474697 127475864 Pos4 0 + 127474697 127475864 255,0,0
ch 7 127475864 127477031 Neg1 0 - 127475864 127477031 0,0,255
ch 7 127477031 127478198 Neg2 0 - 127477031 127478198 0,0,255
ch 7 127478198 127479365 Neg3 0 - 127478198 127479365 0,0,255
ch 7 127479365 127480532 Pos5 0 + 127479365 127480532 255,0,0
ch 7 127480532 127481699 Neg4 0 - 127480532 127481699 0,0,255
A3. The GVF Fo ma
The o ma GVF (Genome Va ia ion Fo ma ) is a specializa ion o he GFF o ma o desc ibe
a ia ions ela i e o a e e ence genome. I is widely used by he communi y o anno a e
a ia ions wi h da abase da a.
GVFLinkGCF
##g - e sion 1.06
##genome-build NCBI B36.3
##sequence- egion ch 16 1 88827254
ch 16 sam ools SNV 49291141 49291141 . + . ID=ID_1;Va ian _seq=A,G;Re e ence_seq=G;
ch 16 sam ools SNV 49291360 49291360 . + . ID=ID_2;Va ian _seq=G;Re e ence_seq=C;
ch 16 sam ools SNV 49302125 49302125 . + . ID=ID_3;Va ian _seq=T,C;Re e ence_seq=C;
ch 16 sam ools SNV 49302365 49302365 . + . ID=ID_4;Va ian _seq=G,C;Re e ence_seq=C;
ch 16 sam ools SNV 49302700 49302700 . + . ID=ID_5;Va ian _seq=T;Re e ence_seq=C;
ch 16 sam ools SNV 49303084 49303084 . + . ID=ID_6;Va ian _seq=G,T;Re e ence_seq=T;
ch 16 sam ools SNV 49303156 49303156 . + . ID=ID_7;Va ian _seq=T,C;Re e ence_seq=C;
ch 16 sam ools SNV 49303427 49303427 . + . ID=ID_8;Va ian _seq=T,C;Re e ence_seq=C;
ch 16 sam ools SNV 49303596 49303596 . + . ID=ID_9;Va ian _seq=T,C;Re e ence_seq=C;
Di e en da ase s con aining his da a can be downloaded om ENSEMBL( p) and NCBI( p).
Acco ding o he README in he ENSEMBL p, we should choose he ile
“homo_sapiens_incl_consequences.g .gz”.
Bioma
Bioma p o ides a use in e ace o que y and e ie e abula iles wi h he da a o hei
da abases.
In o de o e ie e he equi ed da a we can:
1) Choose he da abase o in e es (FROM),
2) Res ic ou que y choosing among he di e en il e ing c i e ia p o ided ( egions, gene
on ology e ms, e c…) (WHERE)
3) Res ic he ields o he esul choosing among he di e en ields (ENSEMBL Ids, HGCN Ids,
e c….) (SELECT)
Gene S a (bp) Gene End (bp) Ensembl Gene ID
17
16573334 16678949 ENSG00000037637
78028101 78149104 ENSG00000036549
29814705 29823405 ENSG00000225011
30117392 30117525 ENSG00000221126
30181698 30182394 ENSG00000228176
VCFTools
Desc ip ion
Pe l sc ip s ha pe o m ope a ion o e c iles
T oubleshoo ing
Requi es ins alling o he linux packages ( ambix, gbzip) and pe l
modules (Tes ::Mos ), and se ing se e al en i onmen a iables
(PATH and PERL5LIB).
Requi es p ep ocessing o c iles: comp essing and index c ea ion.
Few documen a ion abou in e nal de ails o commands (Ex, c -
anno a e has a –d op ion no documen ed
, bu equi ed when
anno a ing a c ile)
Ve sion ins alled
c _ ools_0.1.11
Commands e iewed
c - o- ab, c -que y, c -anno a e
Gene al opinion om
SwEnginee ing
pe spec i e
Ve y in e es ing commands o e c iles, howe e , hose ope a ions
(me g
e, il e ing, in e sec ion, anno a ion), should be pe o med
o e a ia ions ins ead o i s co esponding ex ep esen a ion.
c -que y: Con e s VCF iles in o o ma de ined by he use .
Usage: c -que y [OPTIONS] ile. c .gz
Op ions:
-c, --columns < ile|lis > Lis o comma-sepa a ed column names o
one column name pe line in a ile.
- , -- o ma <s ing> The de aul is '%CHROM:%POS %REF[ %SAMPLE=%GT] n'
-l, --lis -columns Lis columns.
- , -- egion ch : om- o Re ie e he egion. (Runs abix.)
--use-old-me hod Use old e sion o API, slowe bu mo e obus .
Exp essions:
%CHROM The CHROM column (simila ly also o he columns)
%GT T ansla ed geno ype (e.g. C/A)
%GTR Raw geno ype (e.g. 0/1)
%INFO/TAG Any ag in he INFO column
%LINE P in s he whole line
%SAMPLE Sample name
[] The b acke s loop o e all samples
%*<A><B> All o ma ields p in ed as KEY<A>VALUE<B>
Examples:
c -que y ile. c .gz 1:1000-2000 -c NA001,NA002,NA003
c -que y ile. c .gz - 1:1000-2000 -
'%CHROM:%POS %REF %ALT[ %SAMPLE:%*=,] n'
c -que y ile. c .gz - '[%GT ]%LINE n'
c -que y ile. c .gz - '[%GT ]%LINE n'
c -que y ile. c .gz - '%CHROM _%POS %INFO/DP %FILTER n'
18
c - o- ab Con e s he VCF ile in o a ab-delimi ed ex ile lis ing he ac ual a ian s ins ead
o ALT indexes
Usage: c - o- ab [OPTIONS] < in. c > ou . ab
Op ions: -i, --iupac Use one-le e IUPAC codes
c -anno a ion Adds cus om anno a ions o VCF iles.
Usage: c -anno a e [OPTIONS] > ou . c
Op ions:
-a, anno a ion.gz
-d key=INFO,ID=ANN,Numbe =1,Type=In ege ,Desc ip ion='MyAnno a ion'
-c CHROM,FROM,TO,INFO/ANN > ou . c
Vc - il e : Add il e ing c i e ia o il e VCF Files: C omosome, gene, sId (one o a ile), Allele
equency
Usage: c ools – c ile. c –-bed BED ilewi hVa ia ions>
c . il e ed
Anno a
Desc ip ion
Pe l sc ip s ha anno a e c iles wi h s uc u al da a and da abase
in o ma ion.
T oubleshoo ing
Requi es a lo o space in disk. 80M package, 12M Re Gene and 8,5G
dbsnp.
Ve sion ins alled
Dowloaded la es e sion in Oc obe 2013
Commands e iewed
anno a e_ a ia ion, con e 2anno a
Gene al opinion om
SwEnginee ing
pe spec i e
Commands only wo k wi h hei cus om o ma , so use s should
manage hemsel es he ans o ma ion o hei iles o his o ma
(e en o he s anda d o ma c ). I equi es downloading all
da abases locally, which o some use s could be p oblema ic. S ill,
e y in e es ing unc ionali y abou anno a ions, s uc u al da a and
da abase da a and i s sepa a ion in di e en iles. Howe e , I ind
ex ual anno a ions in gene al p oblema ic because ex ual da a is
edious o ead, misses he well-
o med s uc u ed ules o con ol
e o s, and leads o edundancy-which could a ec o access la ency
and space.
Con e 2anno a : Con e s VCF iles in o o ma used by anno a (.a inpu )
Command: con e 2anno a .pl - o ma c 4 ile. c > ile.a inpu
Anno a e_ a ia ion:
• To download a da abase: Conc e ely e Seq, dbsnp and 1000G p ojec (allele equency
da a)
19
Command: anno a e_ a ia ion.pl -downdb -build e hg19 -web om anno a e Gene
humandb
Command:anno a e_ a ia ion.pl -downdb -build e hg19 -web om anno a dbsnp135
humandb
Command:anno a e_ a ia ion.pl -downdb 1000g2012ap humandb –build e hg19
• To sea ch a ia ion p ope ies: Gene, egion.
Command: anno a e_ a ia ion.pl -build e hg19 –geneanno ile.a inpu humandb/
This command c ea es he ile *. a ia ion_ unc ion, which ells whe he he a ian hi a
s uc u al egion and he name o he gene (o neighbou ing genes).
• To sea ch a ia ion in da abases: Conc e ely in dbsnp e ie ing he sId
Command: anno a e_ a ia ion.pl -build e hg19 – il e –db ype dbsnp135
ile.a inpu humandb/
This command c ea es wo iles: *. il e ed ha con ains he SNPs no in dbSNP and
*.d opped ha con ains a ian s ha a e anno a ed in dbSNP oge he wi h hei s iden i ie s
• To sea ch a ia ion p ope ies: S uc u al e ec and hg s no a ion.
Command: anno a e_ a ia ion.pl -build e hg19 -hg s ile.a inpu humandb/
This command c ea es he ile *.exonic_ a ian _ unc ion ha con ains he amino acid
changes as a esul o he exonic a ian (se e al hg sNo a ion wi h i s e SeqIden i ie s)
VEP
Desc ip ion
Pe l Sc ip s ha anno a e a ia ion iles (VCF, TXT and BED) wi h he
e ec s o a a ia ion.
T oubleshoo ing
Requi es o download 5,5 G. Re seq didn’ download co ec ly and I
had o sol e he p oblem manually wi h linux commands (download
and ex ac in he sui able di ec o y: home/. ep/human).
Ve sion ins alled
Dowloaded la es e sion in Oc obe 2013
Commands e iewed
Va ian _e ec _p edic o , il e ep
Gene al opinion om
SwEnginee ing
pe spec i e
Ve y in e es ing unc ionali y abou e ec s. Howe e , he ex ual
o ma has he same p oblem men ioned in anno a and Snpe , he
di e ence is ha VEP exp esses in he
ield “ex a” a se o
p ope ies exp essed using a pai “Key= alue” and sepa a ed wi h
“;”. S ill, o imp o e isualiza ion he analysis c ea es a epo
con aining s a is ics a isual diag ams.
20
Va ian _e ec _p edic o : Anno a es a a ia ion ile (VCF, Pileup, HGVSIden i ie s) wi h s uc u al
da a: Gene (HGNC iden i ie , symbol), CDS posi ions, In on/Exon numbe , and da abase da a: sId
om dbSNP.
Command: pe l a ian _e ec _p edic o –-cache –i ile. c –-symbol –o
ile.o. c
Va ian _e ec _p edic o : Anno a es a a ia ion ile (VCF, Pileup, HGVSIden i ie s) wi h aminoacid
and codon change and allele equency.
Command: pe l a ian _e ec _p edic o –-cache –i ile. c –o ile.o. c
I can also anno a e a ia ions wi h he e ec o SIFT and POLYPHEN, which analyse i he a ian
changes he p o ein unc ion; i calcula es he HGVS no a ion (genomic, coding and some imes
p o ein).
Command: pe l a ian _e ec _p edic o –-cache –i ile. c –-hg s –- e seq –-
si b –-polyphen b –o ile.o. c
Fil e _ ep: Fil e s he ou pu ile o VEP using a ex ual exp ession
Command: pe l il e _ ep –-cache –i ile. c –-hg s –-si b –-polyphen b –o
ile.o. c
21

SamTools
SNPE
Desc ip ion
Ja a p og am ha anno a e a ia ion iles (VCF, TXT and BED) wi h
he e ec s o a a ia ion.
T oubleshoo ing
Requi es o download 1,7G. And ha e a disk Space o 10G.
Ve sion ins alled
Dowloaded la es e sion in Oc obe 2013 ( 3.3)
Commands e iewed
SnpSi , snpE
Gene al opinion om
SwEnginee ing
pe spec i e
I equi es downloading all da abases locally, which o some use s
could be p oblema ic. Ve y in e es ing unc ionali y abou
anno a ions o e ec s and da abase da a. Howe e , he ex ual
o ma has he same p oblem men ioned in anno a . S i
ll, o
imp o e isualiza ion he analysis c ea es a epo con aining
s a is ics a isual diag ams. SnpE p o ides a e y powe ul op ion
o il e a ia ions, as i suppo exp essions in ol ing a big se o
ope a ions and any ield o a VCF ile. Howe e , his ope a ion should
no be pe o med a he ex ual le el.
SnpSi : Anno a es he ID ield o a a ia ion ile (VCF, TXT o BED). This unc ionali y can be used
o anno a e he s o dbSNP. The dbsnp. c can be downloaded om ncbi.
Command: ja a –ja SnpSi .ja anno a e – dbSnp. c ile. c > ile.dbSnp. c
SnpE :
• To download he da abase: F om he e e ence Genome.
Command: ja a –ja SnpE .ja download – GRCh37.69
Command: ja a –ja SnpE .ja download – hg19
• To calcula e s uc u al da a: Gene, e e ence ids. #TODO es his command
Command: ja a –Xmx4g –ja SnpE .ja – GRCh37.69 ile.dbSnp. c > ile.s . c
• To calcula e a ia ion e ec s: Lis o e ec s, hei impac , codon/aminoacid changes,
gene, e e ence ids and he loss o unc ion. Also he hg s no a ion.
Command: ja a –Xmx4g –ja SnpE .ja e –hg s – GRCh37.69 ile.dbSnp. c >
ile.e . c
*Mo e op ions can be used in he analysis: Apply il e s, choose egions o applica ion, p o ide
a lis o ansc ip s, use cus om anno a ions, e c.
22
• To download da abases wi h a ia ion e ec s: The da abase dbNSFP con ains
in o ma ion abou SIFT and POLPHEN as well as MAFs om 1000G among o he s.
Command:wge h p://dbns p.hous onbioin o ma ics.o g/dbNSFPzip/dbNSFP 2.3.zip
Command: unzip dbNSFP2.3.zip
And anno a e a ia ions wi h da a om his da abase
Command: ja a -ja SnpSi .ja dbns p - dbNSFP2.3. x – SIFT_sco e,
Polyphen2_HVAR_p ed ile.anno a ed. c > ile.anno a ed2. c
23
GATK
Pipeline De elopmen En i onmen s
Biopy hon, BioPe l, Bioja a, Bio*
Bio* is a amily o lib a ies o he de elopmen o pe sonalized gene ic analysis ools on op o
well-es ablished p og amming languages. BioPe l (STAJICH2002) in Pe l, BioPy hon (COCK2009)
in Py hon, o BioJa a (HOLLAND2008) in Ja a.
Thei common aim is o p o ide common unc ionali y ega ding DNA sequences manipula ion
and unc ionali y ega ding in eg a ion among so wa e componen s. These lib a ies p o ide a
se o modules, classes and me hods ha implemen algo i hms, da a s uc u es, ad anced
s ing manipula ion ope a ions and so on. Addi ionally, as he domain s ill lacks o a s anda d
nomencla u e o exp ess he ou pu esul s, hey p o ide se e al o ma con e sion ope a ions
o ans o m hese esul s among di e en ools.
Fo example, BioPe l p o ides suppo o : a) Indexa ion, ans o ma ion and anno a ion o
sequences; b) Sequence alignmen ; c) Sequence sea ch; d) Fo ma ans o ma ions; e) Pa e n
ma ching algo i hms o sequence analysis; ) W appe s o da abase e ie al o online se ices
execu ion; and g) 3D ep esen a ion o p o eins.
Figu e 1 BioPe l example
Figu e 6 shows an example w i en wi h BioPe l o access o he EMBL da abase in o de o
e ie e a sequence whose iden i ie is “U14680”. A e e ie al, his sequence is ans o med
o he Genebank o ma .
Ta e na
Ta e na p o ides an en i onmen o design, edi and execu e wo k lows using g aph
ep esen a ions: nodes ha ep esen asks, and a ows ha ep esen links o
communica ions among asks. Gene icis s can choose om a lis o componen s ( ha p o ide
gene ic unc ionali y), he asks hey wan o accomplish, d ag and d op hem on a g aphical
wo kshee and e en ually hey compose a wo k low ha execu es he desi ed gene ic analysis.
Ta e na in eg a es unc ionali y h ough myExpe imen (De ou e e al. 2008), a social ne wo k
o sha e scien i ic wo k lows, and he Bioca alogue (Bhaga e al.), a cu a ed ca alogue o web
se ices o he li e sciences.
Ta e na o e s a use in e ace (Figu e 2) o acili a e he c ea ion o wo k lows. The igh side,
o e s he use a se o abs o wo k wi h he en i onmen ; o example, a ab o disco e se ices
o o o e iew he de ails o a wo k low. The le side o he in e ace shows he g aphical
ep esen a ion o all he asks and da a low among asks ga he ed in he wo k low.
use Bio::DB::EMBL;
use Bio::SeqIO;
my $db= new Bio::DB::EMBL();
my $seq=$db->ge _Seq_by_acc("U14680);
my $seqou =new Bio::SeqIO(- o ma => "genbank");
i (de ined $seq){
$seqou ->w i e_seq($seq);
}
24
Figu e 2 Wo k low Example on Ta e na
Bene i s
A e a wo k low execu ion, i p o ides in e media y esul s om all se ices ha allow he
aceabili y o he esul s.
I p o ides a high abs ac ion o in eg a e command line ools, es se ices, bioma se ices,
and o he s. This abs ac ion makes easie he in eg a ion, bu echnological de ails a e s ill
equi ed.
Disad an ages:
The Ta e na in e ace has a edious wo k low ep esen a ion: Bo h simple and complex
examples con ain o e loading elemen s (a big amoun o se ices).
I is no possible o ep esen he objec s o a da a model.
I does no p o ide clea desc ip ion o unc ionali y and pa ame e s om he in eg a ed
se ices.
Gene al Issues:
- Se ice disco e y: Se ices no o de ed seman ically.
o Same amily se ices a e sca e ed: The a ailable axonomy shows he
echnology and he p o ide in he op. As a second le el some seman ics a e
used, howe e , hey con ain simila and o e lapping ca ego ies such as
{con e sion, con e ing} and {Alignmen , Bioin o ma ics}. In he hi d le el,
seman ics is los again, and se ices a e o de ed by package name.
o Fil e no enough help ul: Se ice sea ch is acili a ed by a key wo d il e .
Howe e , al hough se ice lis is educed disco e y is s ill di icul because o
he ca alog axonomy p e iously desc ibed.
25
4) Re ie e Lis SNPs o GeneBRCA1 om dbSNP, Re ie e Lis mu a ions o Gene om
HGMD, LOVD, BIC.
All his asks can be pe o med in Ta e na, bu in o ma ion is e ie ed om “ENSEMBL
VARIATION 67 (SANGER UK) “
P oblems de ec ed
• Non-consis en da a: Fields e ie ed om Bioma se ices ha e non-
es ic ed con en . Some ields a e no ul illed, o he s do no co espond wi h
he ield meaning and o he s a e exp essed in na u al language. In ou case, i
is no possible o e ie e he ype o a ia ion: inse ion, dele ion, indel.
• The equi ed en i ies o sa ing da a e ie ed a e no a ailable: Fields
e ie ed om any da a sou ce should be sa ed in one h ml/cs /xls ile o
linked o a po .
5) Compa e Lis di e ences s Lis SNPs (To be done)
Galaxy
Galaxy (Gia dine e al. 2005) is an open-sou ce web-based en i onmen o he execu ion o
biological se ices. I s main pu pose is o help gene icis s wi h hei da a in ensi e biological
esea ch h ough he de ini ion o web in e aces o biological da a e ie al and se ices
execu ion. Wi h his pu pose, i p o ides di e en in e aces ha access o some popula
gene ic da abases and oolki s.
Galaxy o e s a use in e ace (Figu e 3) o acili a e he c ea ion o wo k lows made o h ee
panels. The igh panel, shows a lis o ools ha can be execu ed and used o he wo k low.
The cen al panel shows he g aphical ep esen a ion o all he asks and da a low among asks
ga he ed in he wo k low. And inally, he le panel shows he de ails o he ool selec ed in he
cen al panel so ha i can be con igu ed by he use .
32

Figu e 3 Wo k low Example on Galaxy
Repo
Summa y: Galaxy is an en i onmen ha p o ides he use he possibili y o un di e en
biological se ices and c ea e wo k lows combining hose se ices. The en i onmen can be
execu ed locally, using a web in e ace o in he cloud. Galaxy allows he use s o e ie e
local and USCS da abase’ da a se s o be used in hei expe imen s. I allows o combine da a
om independen que ies, o pe o m calcula ions o e hese da a se s (such as il e ing a
da a se , combining se e al da a se s and ans o ming da a using a biological se ice) and,
inally, o isualize he esul s.
PROS: The mos common biological se ices used by gene icis s a e in eg a ed in Galaxy
(such as da abase e ie al, biological algo i hms and isual display uni s). The use s a e
p o ided wi h a use in e ace o each se ice, conc e ely a o m wi h all he pa ame e s o
be illed. Galaxy allows he use o eco d all he s eps o se ices ha a e execu ed, and
a e wa ds, hose selec ed by he use can be composed in a wo k low. All hese asks a e
a ailable o execu e as many imes as equi ed.
CONS: Galaxy ope a es using low le el da a desc ip ions: i s unc ionali y is based in he use
o aw da a se s sa ed in iles. All se ices ecei e da a om a ile and, as a esul , hey ob ain
ano he ile. Da a is o ganized in ows and columns whe e each ield implies a speci ic
meaning. As a consequence, use s con igu e he se ice’s pa ame e s aking in o accoun he
ows and he columns ins ead o he unde lying concep s.
Addi ionally, when composing a wo k low, he use has o wo y abou da a low. The
se ices’ in e aces a e con igu ed o accep iles exp essed in a speci ic o ma , so i can
di e wi hin di e en se ices. I he o ma is di e en a ans o ma ion is equi ed. Hence,
he use has o ind he way o p o ide inside he Galaxy en i onmen a mechanism ha
execu es his ans o ma ion. The au ho s claim ha new unc ionali y can be added, bu
deep knowledge o Galaxy and some p og amming skills a e equi ed.
Wo k low 1: Diagen (Disease diagnosis BREAST Cance om a ia ion de ec ion)
1) Desc ibe Gene, e ie e Pa ien .Sequence, e ie e Gene{BRCA1}.sequence om NCBI.
The h ee asks canno be pe o med in Ta e na.
33
P oblems de ec ed
o En i ies canno be desc ibed: Da a is managed using a “da ase en i y” wi h
some me ada a associa ed (such a name, o ma , e c). Gene and Pa ien a e
ep esen ed as a da ase ha con ains one o se e al sequences espec i ely.
o NCBI da a e ie al canno be pe o med: USCS Genome b owse and
ensemble da abase can be que ied. Addi ionally, gene ic da a can be easily
e ie ed using a Bioma se ice, which p o ides a use in e ace when he
use is guided o c ea e a que y agains a conc e e da a sou ce. This use
in e ace p o ides he a ailable ields o il e he da a sou ce and he en i ies
ha can be e ie ed and hei a ibu es. Using he bioma se ice, NCBI is
no a ailable, bu we used ano he da abase called VEGA (Sange UK), whe e
all he equi ed p ope ies o his wo k low we e a ailable.
o
Desc ip ion: P o eus is a p oblem sol ing en i onmen based on he G id echnology o
composing, compiling and unning bioin o ma ic applica ions. P o eus is based on he use o
on ologies o aid he use s o de ine hei applica ions. The bioin o ma ic on ology used
ep esen s he bioin o ma ics domain by means o he de ini ion o biological da a sou ces,
so wa e componen s and bioin o ma ic p ocesses o asks. The use designs hei
pe sonalized ool b owsing and que ying he on ology o ind o he componen s ha will
be used. Addi ionally, each componen is p o ided wi h me ada a o allow he use o
con igu e i acco ding hei equi emen s.
34
PROS: The on ology is sui able o ind so wa e componen s ha i biologis s’ equi emen s.
Addi ionally, he b owse is easy o use because he on ology is ca ego ized in di e en
axonomies and displayed using showing labeled ela ionships wi h o he i ems o he
on ology.
CONS: The en i onmen does no explain he in e ac ion be ween a domain on ology and
he bioin o ma ics on ology, ei he how he wo k low managemen sys em ans o ms he
g aphical design in o g id sc ip s. As a consequence, i is no clea how da a low among asks
is designed o execu ed.
This wo k was de eloped in 2004-2005, and he en i onmen has no been made a ailable.
Thei au ho s ocused hei e o s in applying hese ideas in mass-spec ome y p o eomics
and c ea ed a speci ic pla o m (MS-Analyze ) ha in eg a es algo i hms and ools o design
ools ega ding his domain.
EBioFlow
eBioFlow (WASSINK2010) is an open-sou ce wo k low managemen sys em o design and execu e
biological wo k lows de eloped in he academic en i onmen as a p oo o concep o a se ies
o Phd disse a ions. I s main pu pose is o imp o e o he wo k low de elopmen en i onmen s
by p o iding a be e usabili y o wo k low design (mul iple pe spec i es o model da a and
con ol low), a be e wo k low enac men wi h suppo o la e binding o se ices, and inally,
he suppo o da a p o enance o imp o ing wo k low sha ing and euse.
Figu e 14 shows he in e ace o eBio low. On he le panel, a se o abs a e a ailable o
na iga e be ween wo k lows, manage da a p o enance, sea ch o se ices, and see p e ious
wo k low uns. On he igh panel, se e al abs p o ide he di e en pe spec i es a ailable o
manage all he di e en conce ns while designing, execu ing and sha ing wo k lows.
Figu e 4 Wo k low Example on eBioFlow
Repo
Desc ip ion: BiosFlow is a wo k low pla o m ha allows he use o disco e web se ices
using on ology cons uc s. I de ines he en i y “node” as a basic uni wi h a se o ea u es
ha ep esen s a se ice. This pla o m p o ides a use en i onmen o disco e , compose
and execu e p ede ined se ice nodes in eg a ed in he pla o m.
35
PROS: This wo k explains he need o simpli y he c ea ion o applica ions by biologis s. Wi h
his pu pose, hey add seman ics o wo k low composi ion, using on ologies, and de elop an
easy use in e ace, based on icons, o d ag and d op se ices and allow hei subsequen
execu ion.
CONS: The e is only a pape ha desc ibes his wo k (2009) and he pla o m is no a ailable.
They claim ha on ologies a e used o sea ch o web se ices, bu i does no explain how
se ices a e ela ed wi h on ologies o allow he se ice disco e y, nei he which on ologies
hey use, and nei he an example o use.
BSIS
Repo
Desc ip ion: BioSe ice In eg a ion Sys em (BSIS) is a amewo k o he de elopmen o
biological wo k lows. BSIS suppo s seman ic disco e y and composi ion o se ices because
p o ides a mechanism o anno a e webse ices wi h on ologies. Conc e ely he on ologies
used a e: 1) Se ice on ologies, which desc ibe p ocesses and ans o ma ions implemen ed
by bioin o ma ic se ices; and 2) Da a on ologies, which desc ibe biological da a.
BSIS p o ides a g aphical wo k low language ( o mally desc ibed) ha allows he use o
c ea e conc e e o abs ac wo k lows ha execu e web se ices. The conc e e wo k lows
a e speci ied by he use s while he abs ac ones a e ins an ia ed by he amewo k
a ending seman ic cons ains added by he use .
PROS: The amewo k p o ides a g aphical use in e ace easy o use o he use . Mo eo e ,
his p oposal akes in o accoun he use o seman ics o he de ini ion o bioin o ma ic
wo k lows. As a consequence, he use is p o ided wi h a high le el o abs ac ion o
implemen a ion de ails. Conc e ely, he use can b owse he biological on ology (se ice o
da a) and use i di ec ly in he wo k low design en i onmen by d ag and d op.
CONS: The phd was de eloped in 2007 and i s main publica ion was on 2010. The con ac
in o ma ion is no alid and he amewo k is no a ailable. As i is no a ailable and he
p ojec seems o be o e , ecen biological on ologies co e age canno be ensu ed. The
“se ice” basic uni is only applicable o web se ice.
36