scieee Open visual document viewer

The SBASE protein domain library, release 6.0: a collection of annoted protein sequence segments

Murvai, János; Vlahovicek, Kristian; Barta, Endre; Szepesvári, Csaba; Acatrinei, Cristina; Pongor, Sándor

Full text

 1999 Ox o d Uni e si y P ess 257–259 Nucleic Acids Resea ch, 1999, Vol. 27, No. 1 The SBASE p o ein domain lib a y, elease 6.0: a collec ion o anno a ed p o ein sequence segmen s János Mu ai1, K is ian Vlaho icek1, End e Ba a2, Csaba Szepes á i3, C is ina Aca inei1 and Sándo Pongo 1,2,* 1In e na ional Cen e o Gene ic Enginee ing and Bio echnology, A ea Science Pa k, 34012 T ies e, I aly, 2ABC Ins i u e o Biochemis y and P o ein Resea ch, 2100 Gödöllö, Hunga y and 3Resea ch G oup on A i icial In elligence, Józse A ila Uni e sis y, 6700 Szeged, Hunga y Recei ed Oc obe 2, 1998; Accep ed Oc obe 7, 1998 ABSTRACT The six h elease o he SBASE p o ein domain lib a y sequences con ains 130 703 anno a ed and c oss e e - enced en ies co esponding o s uc u al, unc ional, ligand-binding and opogenic segmen s o p o eins. The en ies we e g ouped based on s anda d names (2312 g oups) and u he classi ied on he basis o he BLAST simila i y (2463 clus e s). Au oma ed sea ch- ing wi h BLAST and a new sequence-plo ep esen a- ion o local domain simila i ies a e a ailable a he WWW-se e h p://www.icgeb. ies e.i /sbase . A mi o si e is a h p://sbase.abc.hu/sbase . The da abase is eely a ailable by anonymous ‘ p’ ile ans e om p.icgeb. ies e.i INTRODUCTION De ec ion o domains in newly de e mined sequences is usually based on pa e n collec ions ha con ain consensus ep esen a ion domain ypes deduced om mul iple alignmen s. Consensus desc ip ions come in di e en a ie ies such as egula exp ess- ions, sequence p o iles, hidden Ma ko models, e c. De elop- men o such a consesnsus desc ip ion equi es expe ise and ca e ul judgemen hence pa e n collec ions can ha dly keep pace wi h he low o new genome da a. Ano he p oblem is he ine i able s a is ical bias o he consensus. Namely, a ypical domains o which he e a e oo ew known examples, may no i well wi h a consensus pa e n de eloped wi h a nume ous da ase o simila domains. Finally, he e a e domain ypes o which i is no easy o de elop consensus ep esen a ions because o weak simila i y. SBASE is a collec ion o p o ein domain sequences designed o acili a e de ec ion domain homologies wi hou he abo e p oblems (1,2). He e he me hod o domain ecogni ion is da abase sea ch a he han pa e n sea ch, so a ypical and ypical domains a e equally well ecognized. The unde lying da abase, SBASE is p ep ocessed by BLAST simila i y sea ch (3) and he simila i y g oups ( ha can be bes pic u ed as densely connec ed g aphs) o m he basis o domain ecogni ion. Table 1. Inc ease o da a in SBASE 6.0 The cu en elease 6.0 o SBASE con ains o e 100 000 anno a ed p o ein sequence segmen s consis en ly named by s uc u e, unc ion, biased composi ion, binding-speci ici y and/ o simila i y o o he p o eins. The main de elopmen s wi h espec o he p e ious elease can be summa ized as ollows. (i) Release 6.0 con ains 130 703 sequence en ies, 63% mo e han elease 5.0 (Table 1). (ii) All eco ds a e now p o ided wi h s anda d names and an e o was made o use domain names also used by o he squence da abases and pa e n collec ions like P osi e (4) and PFAM (5). (iii) The en ies we e g ouped based on s anda d names (2312 g oups) and hose wi h a leas h ee en ies (1039 g oups) we e u he classi ied on he basis o he BLAST simila i y. A o al o 2463 clus e s wi h a leas h ee membe s a e deposi ed in o a sepa a e da abase, SBASE-CLUSTERS, which is now a ailable h ough anonymous p as well as h ough links on he WWW-se e (a desc ip ion o he clus e ing p ocedu e is gi en a he web-si e). Wi hin each s anda d name g oup he clus e s a e numbe ed, in such a way ha clus e s wi h mo e in e -membe simila i y ha e la ge numbe s. (i ) A new g aphic ou pu acili y is added o he se e whe eby local domain simila i y can be plo ed along he sequence. DESCRIPTION OF THE DATA De ini ion o p o ein domains Domains included in SBASE a e p o ein sequence segmen s wi h known s uc u e and/o unc ion. The main en y classes a e summa ized in Table 2. The bounda ies o he domains a e ei he *To whom co espondence should be add essed a : ICGEB, A ea Science Pa k, 34012 T ies e, I aly. Tel: +39 040 375 7300; Fax: +39 040 226 555; Email: [email p o ec ed] by gues on Oc obe 8, 2015h p://na .ox o djou nals.o g/Downloaded om Nucleic Acids Resea ch, 1999, Vol. 27, No. 1 258 Table 2. Examples o domains in SBASE 6.0 Table 3. C oss- e e ences o o he da abases in SBASE as p e iously de ined in he o iginal publica ions o de e mined by homology o domains wi h known bounda ies. In his elease, he bounda ies used by PFAM (5) we e adop ed o a numbe o domain ypes. Sou ce and o igin o da a SBASE da a o igina e om h ee main sou ces: (i) om he SWISS-PROT p o ein sequence da abank (6); (ii) om he P o ein Sequence Da abase o he PIR In e na ional P o ein sequence da abase (PIR) (7); and (iii) om he li e a u e. F om a o al o 130 703 eco ds in SBASE 6.0, 96 305 (73%), 27 089 (21%) and 6656 (5%) a e o euka yo ic, p oka yo ic and i al o igin, espec i ely. Domain sizes a y in leng h be ween 5 and 1000 amino acids. Redundancy o sequences in SBASE 6.0 is kep a a minimal le el. In some cases, he domain de ini ions o e lap. C oss- e e ences SBASE 6.0 has c oss- e e ences o se e al p o ein and nucleic acid da abanks, as well as o he PROSITE (4), PRINTS (8), PRODOM (9) and BLOCKS (10) da abases (Table 3). In each eco d, he DR-lines con ain he c oss- e e ence da a. Reco d s uc u e The o ma o SBASE 6.0 (Fig. 1) ollows ha o he EMBL and SWISS-PROT da abases and can be di ec ly o ma ed unde he GCG package The ield ypes used a e lis ed in Table 4. The Figu e 1. A sample en y om he SBASE 6.0 p o ein domain lib a y. An annexin epea domain. The unde lined i ems a e linked in he SBASE Wo ld Wide Web se e so ha he co esponding eco ds can be iewed on he sc een by ‘clicking’ on hem. Table 4. Types o commen lines in SBASE 6.0 eco ds clus e s o which a sequence belongs a e de e mined by (i) he s anda d name and (ii) he (op ional) subclass numbe included in he CL ield, e.g. ANNEXINS/8 ( he CE ield o p e ious eleases is now abandoned). DISTRIBUTION AND ACCESS Dis ibu ion SBASE 6.0 (23 Oc obe , 1998) is dis ibu ed by anonymous ‘ p’ ile ans e om p.icgeb. ies e.i . The comple e da abase (including he eco ds and lis o clus e s), is 75 Mb, i s comp essed o m is 8.3 Mb. Access by WWW: eco d e ie al and BLAST sea ch SBASE 6.0 and SBASE-CLUSTERS can be sea ched a he WWW-se e h p://base.icgeb. ies e.i /sbase and a he mi o si e h p://sbase.abc.hu/sbase . Reco d e ie al is wi h he SRS sys em. A p esen , c oss- e e ences o SBASE-CLUSTERS, EMBL, MEDLINE, MIM, PRINTS, PRODOM, PROSITE and SWISS-PROT can be di ec ly accessed h ough he WWW- se e . P edic ion o domain homologies ia BLAST sea ching is possible ei he by (i) unning a sea ch agains SBASE, o (ii) unning a sea ch agains SWISS-PROT and ep ocessing he sea ch ou pu (11,12). In he ou pu o he la e , local domain simila i ies a e also g aphically ep esen ed as a sequence-plo (Fig. 2). by gues on Oc obe 8, 2015h p://na .ox o djou nals.o g/Downloaded om 259 Nucleic Acids Resea ch, 1994, Vol. 22, No. 1 Nucleic Acids Resea ch, 1999, Vol. 27, No. 1 259 Figu e 2. (A) G aphic ou pu o he domain simila i y se e (www. icgeb. ies e.i /sbase ) in esponse o he que y sequence C1S_HUMAN om SWISS-PROT. The known domain s uc u e o his que y is CUB-EGF-CUB- SUSHI-SUSHI-SPR (whe e S = signal, P = p opep ide, SPR = se ine p o ease). The oupu shows he plo o he BLAST simila i ies along wi h he SBASE s anda d names. A ows ha e been added o help iden i ica ion in black and whi e (o iginal is in colo ). (B) Ou pu o he domain homology WWW se e (www.icgeb. ies e.i /sbase ) in esponse o he annexin sequence shown in Figu e 1 (de ail). NSD: numbe o signi ican simila i ies ound in he BLAST ou pu ; GN.: numbe o he gi en domain occu ing in he da abase; Sum.Sco e: cumula i e sum o BLAST sco es belonging o a domain-name in he ou pu ; O e lap Max: maximum simila i y sco e ound (11). The se e ou pu con ains alignmen s p o ided wi h anno a ions and a de ailed explana ion abou e alua ion (no shown). Ci a ion Use s o SBASE and o he WWW/Email se e s a e asked o ci e his a icle in hei publica ions. ACKNOWLEDGEMENTS SBASE was es ablished in 1990 and is main ained collabo a i e- ly by he In e na ional Cen e o Gene ic Enginee ing and Bio echnology, T ies e, I aly and he ABC Ins i u e o Biochem- is y and P o ein Resea ch, Gödöllö, Hunga y. The au ho s wish o hank he suppo o EMBne , he Eu opean Molecula Biology Ne wo k. The P o ein S uc u e and Func ion G oup is suppo ed by EMBne in he amewo k o EU g an ERB- BIO4-CT96-0030. Wo k a ABC was suppo ed by ICGEB collabo a i e esea ch g an no CRP/HUN9603. REFERENCES 1 Pongo ,S., Ske l,V., Cse zo,M., Ha sagi,Z., Simon,G. and Be ilacqua,V. (1993) P o ein Engng., 6, 391–395. 2 Fabian,P., Mu ai,J., Ha sagi,Z., Vlaho icek,K., Hegyi,H. and Pongo ,S. (1997) Nucleic Acids Res., 25, 240–243. 3 Al schul,S.F., Madden,T.L., Scha e ,A.A., Zhang,J., Zhang,Z., Mille ,W. and Lipman,D.J. (1997) Nucleic Acids Res., 25, 3389–3402. 4 Bai och,A., Buche ,P. and Ho mann,K. (1996) Nucleic Acids Res., 24, 189–196. 5 Sonnhamme ,E.L., Eddy,S.R., Bi ney,E., Ba eman,A. and Du bin,R. (1998) Nucleic Acids Res., 26, 320–322. 6 Bai och,A. and Apweile ,R. (1998) Nucleic Acids Res., 26, 38–42. 7 Ba ke ,W.C., Ga a elli,J.S., Ha ,D.H., Hun ,L.T., Ma zec,C.R., O cu ,B.C., S ini asa ao,G.Y., Yeh,L.S.L., Ledley,R.S., Mewes,H.W., P ei e ,F. and Tsugi a,A. (1998) Nucleic Acids Res., 26, 27–32. 8 A wood,T.K., Beck,M.E., Flowe ,D.R., Sco dis,P. and Selley,J.N. (1998) Nucleic Acids Res., 26, 304–308. 9 Co pe ,F., Gouzy,J. and Kahn,D. (1998) Nucleic Acids Res., 26, 323–326. 10 Heniko ,S., Pie oko ski,S. and Heniko ,J.G. (1998) Nucleic Acids Res., 26, 309–312. 11 Mu ai,J., Vlaho icek,K., Ba a,E., PFei e ,F., Hegyi,H. and Pongo ,S. (1998) Bioin o ma ics, in p ess. 12 Hegyi,H. and Pongo ,S. (1993) Compu . Applic. Biosci., 9, 371–372. 13 S oesse ,G., Moseley,M.A., Sleep,J., McGow an,M., Ga cia-Pas o ,M. and S e k,P. (1998) Nucleic Acids Res., 26, 8–15. 14 Be ns ein,F.C., Koe zle,T.F., Williams,G.J., Meye ,E.E.,J , B ice,M.D., Rodge s,J.R., Kenna d,O., Shimanouchi,T. and Tasumi,M. (1977) J. Mol. Biol., 112, 535–542. 15 Pea son,P., F ancomano,C., Fos e ,P., Bocchini,C., Li,P. and McKusick,V. (1994) Nucleic Acids Res., 22, 3470–3473. 16 Flybase Conso ium (1998) Nucleic Acids Res., 26, 85–88. 17 Rudd,K.E., Bou a d,G. and Mille ,G. (1992) In Da ies,K.E. and Tilghman,S.M. (eds), Genome Analysis. Cold Sp ing Ha bo Labo a o y P ess, New Yo k, pp. 1–38. 18 Mye s,F. (1990) Human Re o i us and Aids Da abase. Los Alamos Na ional Labo a o y, Los Alamos, NM, USA. 19 Robe s,R.J. and Macelis,D. (1998) Nucleic Acids Res., 26, 338–350. by gues on Oc obe 8, 2015h p://na .ox o djou nals.o g/Downloaded om