scieee Open visual document viewer

SWAPSC: sliding window analysis procedure to detect selective constraints

Fares, Mario A.

Abstract

Sliding-window analysis procedure to detect selective constraints (SWAPSC) is a software system to dissect the constraints on the evolution of protein-coding genes.The programestimates rates of nucleotide substitutions at specific codon regions in each branch of a phylogenetic tree. The program uses several sets of simulated sequence alignments to estimate the probability of synonymous and nonsynonymous nucleotide substitutions. Thereafter, a statistical analysis is conducted to determine the optimum windowsize to detect selective constraints. Finally, the optimum window size is slid along the real alignment and a test for significance of the estimated number of synonymous and non-synonymous nucleotide substitutions in each sliding step is conducted. A number of friendly useful output files is generated.

Full text

BIOINFORMATICS APPLICATIONS NOTE Vol. 20 no. 16 2004, pages 2867–2868 doi:10.1093/bioin o ma ics/b h303 SWAPSC: sliding window analysis p ocedu e o de ec selec i e cons ain s Ma io A. Fa es Depa men o Biology, Na ional Uni e si y o I eland, Maynoo h, Co. Kilda e, I eland Recei ed on Ap il 15, 2004; e ised on Ap il 21, 2004; accep ed on Ap il 25, 2004 Ad ance Access publica ion May 6, 2004 ABSTRACT Summa y: Sliding-window analysis p ocedu e o de ec selec i e cons ain s (SWAPSC) is a so wa e sys em o dissec he cons ain s on he e olu ion o p o ein-coding genes.Thep og ames ima es a eso nucleo idesubs i u ions a speci ic codon egions in each b anch o a phylogene ic ee. The p og am uses se e al se s o simula ed sequence alignmen s oes ima e hep obabili yo synonymousandnon- synonymous nucleo ide subs i u ions. The ea e , a s a is ical analysisisconduc ed ode e mine heop imumwindowsize o de ec selec i e cons ain s. Finally, he op imum window size is slid along he eal alignmen and a es o signi icance o he es ima ed numbe o synonymous and non-synonymous nucleo ide subs i u ions in each sliding s ep is conduc ed. A numbe o iendly use ul ou pu iles is gene a ed. A ailabili y: SWAPSC is a ailable a h p//www.may.ie/ academic/biology/s a /m molece olandbioin .sh ml. Dis ibu- ion e sions o bo h Linux and Windows ope a ing sys ems a e a ailable, including manual and example iles. Con ac : ma io. a [email protected] Di e en s uc u al and unc ional p o ein domains a e likely o be subjec o dis inc selec i e cons ain s. Co-e olu ion be ween di e en codon si es wi hin hese domains ques ions he use o a single codon as he uni o selec ion, as p e iously demons a edinse e als udies(HughesandNei,1988;Ma ín e al., 2001). Region-speci ic selec i e cons ain s can be de e mined by compa ing he expec ed o he obse ed numbe s o nuc- leo ide subs i u ions. When he phylogeny is p o ided, he same app oach can be applied o speci ic b anches o he ee. SWAPSC enables p ecise analyses o selec i e cons ain s o each egion o he sequence and b anch o he ee. B ie ly, SWAPSC es ima es he a e age numbe o non-synonymous (θN) and synonymous (θS) nucleo ide subs i u ions by he model o Li (1993) using simula ed alignmen s. The p ob- abili ies o θNand θSa e hen calcula ed assuming a binomial dis ibu ion o each a e o subs i u ion. These p obabili ies a eused he ea e oob ain he op imumwindowsize. Unlike o he s udies ha usea andomwindow(Tajima,1991;Hughes and Nei, 1989), SWAPSC pe o ms a s a is ical es o de ine heapp op ia ewindowsize(Fa ese al.,2002). Op imiza ion o he window size is possible sliding a window o a de ined size (e.g. 1–20 codons) along he simula ed sequence align- men and es ing he p obabili ies o synonymous (Ks) and non-synonymous (Ka) subs i u ions unde a Poisson dis ibu- ion. Tha window size showing he 5% lowe P(K a) alue >0.05 is selec ed as he app op ia e one. SWAPSC slides he op imum window along he eal sequence alignmen and es ima es he p obabili y o Kaand Kscompa ing each sequence wi h i s ances o in e ed by maximum pa simony. Due o ha Kaand Ksa e es ed o signi icance, se e al hypo heses ega ding he e olu ion o a speci ic b anch o he ee and egion o he alignmen can be es ed. Hence, SWAPSC can de ec egions ha a e unde adap i e e olu ion, o accele a ed ixa ion a es o non- synonymous subs i u ions, o sa u a ion o synonymous si es o ho spo s. The emphasis in SWAPSC has been ocused in ou main poin s: accu acy o esul s, au oma ic pe o mance, access- ibili y and exhaus i e sc eening o an unlimi ed alignmen size. Files equi edand gene a edby SWAPSC a edepic ed in Figu e 1A. An inpu mul iple-alignmen o coding sequences in PHYLIP o ma , s anda d o many di e en p og ams (e.g. PHYLIP, PAML, e c.), is equi ed. Se e al se s o simula ed sequencealignmen sinPHYLIP o ma and he eeinnewick o ma ha ealso obep o ided.Then heuse has woop ions: (1) ix he window size, which is no ecommended unless biological in o ma ion suppo s ha op ion and (2) in e he app op ia e window size. The p og am gene a es independen ou pu iles in addi ion o he main ou pu ile o help he use o deal wi h he huge amoun o in o ma ion ob ained. These iles a e: (1) an Excel ile con aining he in o ma ion o each egion and b anch; (2) a ile o isualize b anches and loca e cons ain s in an easy way using TREEVIEW p og am; and (3) a ile wi h he amino acid eplacemen s in each b anch o he ee. Figu e1B exempli ies he in o ma ion ha can be ob ained. The pe o mance o he algo i hm has been examined by se e al case s udies (Fa es e al., 2002; Lynn e al., 2004). To ob ain mo e s a is ically obus esul s, he use is ad ised o: (1) use la ge da ase s o sequence alignmen s; (2) use long Bioin o ma ics ol.20 issue 16 © Ox o d Uni e si y P ess 2004;all igh s ese ed. 2867 M.A.Fa es Inpu ile T ee ileSimula ions ile SWAPSC Ou pu ile Excel ile T ee iew ile Amino acid changes ile K. pneumoniae E. ae ogenes E. ca o o o a S. yphimu ium E. coli S. glossinidia A. ac inomy H. in luenzae P. ae uginosa 0.02 Codon si e 0 50 100 150 200 250 300 350 400 450 500 550 Subs i u ions 0.0 0.5 1.0 1.5 2.0 2.5 3.0 Ka Ks W Ka =1.07 Ks = 0.43 = 2.49 PS S Ka =0.21 Ks = 0.18 = 1.18 A B ω ω Fig. 1. In o ma ion gene a ed by SWAPSC. (A) Files equi ed o gene a ed by SWAPSC and (B) selec i e cons ain s ope a ing in one b anch o he ee. Sa u a ion o synonymous si es (S), posi i e selec ion (PS), non-synonymous (Ka) and synonymous (Ks) subs i u ions a e indica ed. sequence alignmen s (200–300 codon si es); and (3) mul iple alignmen s should be eliable. The cu en e sion allows he use o Li’s model, Howe e , I will upg ade my so wa e pe iodically o in oduce mo e Kimu a-based models. ACKNOWLEDGEMENTS I would like o acknowledge D A il Coghland o ca e ul eading o he manusc ip and he be a es e s o SWAPSC o iden i ying bugs. REFERENCES Fa es,M.A., Elena,S.F., O iz,J., Moya,A. and Ba io,E. (2002) A sliding window-based me hod o de ec selec i e cons ain s in p o ein-coding genes and i s applica ion o RNA i uses. J. Mol. E ol.,55, 509–521. Hughes,A.L. and Nei,M. (1988) Pa e n o nucleo ide subs i u ion a majo his ocompa ibili y complex class I loci e eals o e domin- an selec ion. Na u e,335, 167–170. Hughes,A.L. and Nei,M. (1989) Nucleo ide subs i u ion a majo his ocompa ibili y complex class II loci: e idence o o e dominan selec ion. P oc. Na l Acad. Sci., USA,86, 958–962. Li,W.H. (1993) Unbiased es ima ion o he a es o synonymous and nonsynonymous subs i u ion. J. Mol. E ol.,36, 96–99. Lynn,D.J., Lloyd,A.T., Fa es,M.A. and O’Fa elly,C. (2004) E idence o posi i ely selec ed si es in mammalian α-de ensins. Mol. Biol. E ol.,21, 819–827. Ma ín,I.,Fa es,M.A.,González-Candelas,F.,Ba io,E.andMoya,A. (2001) De ec ing changes in he unc ional cons ain s o pa alog- ous genes. J. Mol. E ol.,52, 17–28. Tajima,F. (1991) De e mina ion o window size o analysing DNA sequences. J. Mol. E ol.,33, 470–473. 2868