BIOINFORMATICS APPLICATIONS NOTE Vol. 20 no. 16 2004, pages 2867–2868
doi:10.1093/bioin o ma ics/b h303
SWAPSC: sliding window analysis p ocedu e
o de ec selec i e cons ain s
Ma io A. Fa es
Depa men o Biology, Na ional Uni e si y o I eland, Maynoo h, Co. Kilda e, I eland
Recei ed on Ap il 15, 2004; e ised on Ap il 21, 2004; accep ed on Ap il 25, 2004
Ad ance Access publica ion May 6, 2004
ABSTRACT
Summa y: Sliding-window analysis p ocedu e o de ec
selec i e cons ain s (SWAPSC) is a so wa e sys em o
dissec he cons ain s on he e olu ion o p o ein-coding
genes.Thep og ames ima es a eso nucleo idesubs i u ions
a speci ic codon egions in each b anch o a phylogene ic
ee. The p og am uses se e al se s o simula ed sequence
alignmen s oes ima e hep obabili yo synonymousandnon-
synonymous nucleo ide subs i u ions. The ea e , a s a is ical
analysisisconduc ed ode e mine heop imumwindowsize o
de ec selec i e cons ain s. Finally, he op imum window size
is slid along he eal alignmen and a es o signi icance o
he es ima ed numbe o synonymous and non-synonymous
nucleo ide subs i u ions in each sliding s ep is conduc ed. A
numbe o iendly use ul ou pu iles is gene a ed.
A ailabili y: SWAPSC is a ailable a h p//www.may.ie/
academic/biology/s a /m molece olandbioin .sh ml. Dis ibu-
ion e sions o bo h Linux and Windows ope a ing sys ems
a e a ailable, including manual and example iles.
Con ac : ma io. a [email protected]
Di e en s uc u al and unc ional p o ein domains a e likely
o be subjec o dis inc selec i e cons ain s. Co-e olu ion
be ween di e en codon si es wi hin hese domains ques ions
he use o a single codon as he uni o selec ion, as p e iously
demons a edinse e als udies(HughesandNei,1988;Ma ín
e al., 2001).
Region-speci ic selec i e cons ain s can be de e mined by
compa ing he expec ed o he obse ed numbe s o nuc-
leo ide subs i u ions. When he phylogeny is p o ided, he
same app oach can be applied o speci ic b anches o he ee.
SWAPSC enables p ecise analyses o selec i e cons ain s o
each egion o he sequence and b anch o he ee. B ie ly,
SWAPSC es ima es he a e age numbe o non-synonymous
(θN) and synonymous (θS) nucleo ide subs i u ions by he
model o Li (1993) using simula ed alignmen s. The p ob-
abili ies o θNand θSa e hen calcula ed assuming a binomial
dis ibu ion o each a e o subs i u ion. These p obabili ies
a eused he ea e oob ain he op imumwindowsize. Unlike
o he s udies ha usea andomwindow(Tajima,1991;Hughes
and Nei, 1989), SWAPSC pe o ms a s a is ical es o de ine
heapp op ia ewindowsize(Fa ese al.,2002). Op imiza ion
o he window size is possible sliding a window o a de ined
size (e.g. 1–20 codons) along he simula ed sequence align-
men and es ing he p obabili ies o synonymous (Ks) and
non-synonymous (Ka) subs i u ions unde a Poisson dis ibu-
ion. Tha window size showing he 5% lowe P(K
a) alue
>0.05 is selec ed as he app op ia e one.
SWAPSC slides he op imum window along he eal
sequence alignmen and es ima es he p obabili y o Kaand
Kscompa ing each sequence wi h i s ances o in e ed by
maximum pa simony. Due o ha Kaand Ksa e es ed o
signi icance, se e al hypo heses ega ding he e olu ion o
a speci ic b anch o he ee and egion o he alignmen
can be es ed. Hence, SWAPSC can de ec egions ha a e
unde adap i e e olu ion, o accele a ed ixa ion a es o non-
synonymous subs i u ions, o sa u a ion o synonymous si es
o ho spo s.
The emphasis in SWAPSC has been ocused in ou main
poin s: accu acy o esul s, au oma ic pe o mance, access-
ibili y and exhaus i e sc eening o an unlimi ed alignmen
size. Files equi edand gene a edby SWAPSC a edepic ed in
Figu e 1A. An inpu mul iple-alignmen o coding sequences
in PHYLIP o ma , s anda d o many di e en p og ams (e.g.
PHYLIP, PAML, e c.), is equi ed. Se e al se s o simula ed
sequencealignmen sinPHYLIP o ma and he eeinnewick
o ma ha ealso obep o ided.Then heuse has woop ions:
(1) ix he window size, which is no ecommended unless
biological in o ma ion suppo s ha op ion and (2) in e he
app op ia e window size.
The p og am gene a es independen ou pu iles in addi ion
o he main ou pu ile o help he use o deal wi h he huge
amoun o in o ma ion ob ained. These iles a e: (1) an Excel
ile con aining he in o ma ion o each egion and b anch;
(2) a ile o isualize b anches and loca e cons ain s in an
easy way using TREEVIEW p og am; and (3) a ile wi h he
amino acid eplacemen s in each b anch o he ee. Figu e1B
exempli ies he in o ma ion ha can be ob ained.
The pe o mance o he algo i hm has been examined by
se e al case s udies (Fa es e al., 2002; Lynn e al., 2004). To
ob ain mo e s a is ically obus esul s, he use is ad ised o:
(1) use la ge da ase s o sequence alignmen s; (2) use long
Bioin o ma ics ol.20 issue 16 © Ox o d Uni e si y P ess 2004;all igh s ese ed. 2867
M.A.Fa es
Inpu ile T ee ileSimula ions ile
SWAPSC
Ou pu
ile
Excel
ile
T ee iew
ile
Amino acid
changes ile
K. pneumoniae
E. ae ogenes
E. ca o o o a
S. yphimu ium
E. coli
S. glossinidia
A. ac inomy
H. in luenzae
P. ae uginosa
0.02
Codon si e
0 50 100 150 200 250 300 350 400 450 500 550
Subs i u ions
0.0
0.5
1.0
1.5
2.0
2.5
3.0
Ka
Ks
W
Ka =1.07
Ks = 0.43
= 2.49
PS
S
Ka =0.21
Ks = 0.18
= 1.18
A
B
ω
ω
Fig. 1. In o ma ion gene a ed by SWAPSC. (A) Files equi ed o gene a ed by SWAPSC and (B) selec i e cons ain s ope a ing in one
b anch o he ee. Sa u a ion o synonymous si es (S), posi i e selec ion (PS), non-synonymous (Ka) and synonymous (Ks) subs i u ions a e
indica ed.
sequence alignmen s (200–300 codon si es); and (3) mul iple
alignmen s should be eliable.
The cu en e sion allows he use o Li’s model, Howe e ,
I will upg ade my so wa e pe iodically o in oduce mo e
Kimu a-based models.
ACKNOWLEDGEMENTS
I would like o acknowledge D A il Coghland o ca e ul
eading o he manusc ip and he be a es e s o SWAPSC o
iden i ying bugs.
REFERENCES
Fa es,M.A., Elena,S.F., O iz,J., Moya,A. and Ba io,E. (2002) A
sliding window-based me hod o de ec selec i e cons ain s in
p o ein-coding genes and i s applica ion o RNA i uses. J. Mol.
E ol.,55, 509–521.
Hughes,A.L. and Nei,M. (1988) Pa e n o nucleo ide subs i u ion a
majo his ocompa ibili y complex class I loci e eals o e domin-
an selec ion. Na u e,335, 167–170.
Hughes,A.L. and Nei,M. (1989) Nucleo ide subs i u ion a
majo his ocompa ibili y complex class II loci: e idence o
o e dominan selec ion. P oc. Na l Acad. Sci., USA,86,
958–962.
Li,W.H. (1993) Unbiased es ima ion o he a es o synonymous and
nonsynonymous subs i u ion. J. Mol. E ol.,36, 96–99.
Lynn,D.J., Lloyd,A.T., Fa es,M.A. and O’Fa elly,C. (2004)
E idence o posi i ely selec ed si es in mammalian α-de ensins.
Mol. Biol. E ol.,21, 819–827.
Ma ín,I.,Fa es,M.A.,González-Candelas,F.,Ba io,E.andMoya,A.
(2001) De ec ing changes in he unc ional cons ain s o pa alog-
ous genes. J. Mol. E ol.,52, 17–28.
Tajima,F. (1991) De e mina ion o window size o analysing DNA
sequences. J. Mol. E ol.,33, 470–473.
2868