RSAT a ia ion- ools: An accessible and lexible amewo k o p edic he
impac o egula o y a ian s on ansc ip ion ac o binding
Wal e San ana-Ga cia
a,b
, Ma ia Rocha-Ace edo
b
, Lucia Rami ez-Na a o
b
, Y on Mbouamboua
c,
,
Denis Thie y
a
, Mo gane Thomas-Chollie
a
, B uno Con e as-Mo ei a
d,e
, Jacques an Helden
,g,
⇑
,
Alejand a Medina-Ri e a
b,*
a
Ins i u de Biologie de l’ENS (IBENS), Dépa emen de biologie, École no male supé ieu e, CNRS, INSERM, Uni e si é PSL, 75005 Pa is, F ance
b
Labo a o io In e nacional de In es igación sob e el Genoma Humano, Uni e sidad Nacional Au ónoma de México, Campus Ju iquilla, Bl d Ju iquilla 3001, San iago de Que é a o
76230, Mexico
c
Fonda ion Congolaise pou la Reche che Médicale, B azza ille, People’s Republic o Congo
d
Es ación Expe imen al de Aula Dei-CSIC, Za agoza, Spain
e
Fundación ARAID, Za agoza, Spain
Aix-Ma seille Uni , INSERM UMR S 1090, Theo y and App oaches o Genome Complexi y (TAGC), F-13288 Ma seille, F ance
g
CNRS, Ins i u F ançais de Bioin o ma ique, IFB-co e, UMS 3601, E y, F ance
a icle in o
A icle his o y:
Recei ed 27 Ap il 2019
Recei ed in e ised o m 22 Sep embe
2019
Accep ed 25 Sep embe 2019
A ailable online 7 No embe 2019
Keywo ds:
Regula o y a ian s
T ansc ip ion ac o s
Posi ion speci ic sco ing ma ix
SNPs
Binding mo i s
abs ac
Gene egula o y egions con ain sho and degene a ed DNA binding si es ecognized by ansc ip ion
ac o s (TFBS). When TFBS ha bo SNPs, he DNA binding si e may be a ec ed, he eby al e ing he an-
sc ip ional egula ion o he a ge genes. Such egula o y SNPs ha e been implica ed as causal a ian s in
Genome-Wide Associa ion S udy (GWAS) s udies. In his s udy, we desc ibe imp o ed e sions o he
p og ams Va ia ion- ools designed o p edic egula o y a ian s, and p esen ou case s udies o illus-
a e hei usage and applica ions. In b ie , Va ia ion- ools acili a e i) ob aining a ia ion in o ma ion,
ii) in e con e sion o a ia ion ile o ma s, iii) e ie al o sequences su ounding a ian s, and i ) calcu-
la ing he change on p edic ed ansc ip ion ac o a ini y sco es be ween alleles, using mo i scanning
app oaches. No ably, he ools suppo he analysis o haplo ypes. The ools a e included wi hin he
well-main ained sui e Regula o y Sequence Analysis Tools (RSAT, h p:// sa .eu), and accessible h ough
a web in e ace ha cu en ly enables analysis o i e me azoa and en plan genomes. Va ia ion- ools can
also be used in command-line wi h any locally-ins alled Ensembl genome. Use s can inpu pe sonal
collec ions o a ian s and mo i s, p o iding lexibili y in he analysis.
Ó2019 The Au ho s. Published by Else ie B.V. on behal o Resea ch Ne wo k o Compu a ional and
S uc u al Bio echnology. This is an open access a icle unde he CC BY-NC-ND license (h p://c ea i e-
commons.o g/licenses/by-nc-nd/4.0/).
1. In oduc ion
Genomic DNA sequence ha bo s he gene egula o y in o ma-
ion necessa y spa ial and empo al gene exp ession pa e ns
[38,31]. Gene egula o y egions encompass sho , highly edun-
dan DNA mo i s ecognized by ansc ip ion ac o s (TF) [36].
These egula o y egions may con ain gene ic a ian s, Single
Nucleo ide Polymo phisms (SNPs) o indels, ha al e he DNA
TF binding si e (TFBS), and he eby he binding o TF [20]. Mo e-
o e , i has been epo ed ha 93.7% o a ian s ha ha e been
associa ed wi h human ai s o diseases ha e been ound o be
loca ed in non-coding egions [43,40], and pa icula ly en iched
in open ch oma in egions [57], indica ing ha hese a ian s
h ps://doi.o g/10.1016/j.csbj.2019.09.009
2001-0370/Ó2019 The Au ho s. Published by Else ie B.V. on behal o Resea ch Ne wo k o Compu a ional and S uc u al Bio echnology.
This is an open access a icle unde he CC BY-NC-ND license (h p://c ea i ecommons.o g/licenses/by-nc-nd/4.0/).
Abb e ia ions: RSAT, Regula o y Sequence Analysis Tools; SNP, Single Nucleo-
ide Polymo phism; TF, T ansc ip ion Fac o ; TFBS, T ansc ip ion Fac o Binding
Si e; PSSM, Posi ion Speci ic Sco ing Ma ix; MPRA, Massi ely Pa allel Repo e
Assays: MPRA; LD, Linkage Disequilib ium; sID, Re e ence SNP Iden i ie ; SOIs,
SNPs o In e es ; GWAS, Genome Wide Associa ion S udies; CRM, Cis-Regula o y
Module; eQTL, Exp ession Quan i a i e T ai Loci; ROC, Recei e Ope a ing Cha -
ac e is ic; CEU, No he n Eu opeans om U ah.
⇑
Co esponding au ho s a : Labo a o io In e nacional de In es igación sob e el
Genoma Humano, Uni e sidad Nacional Au ónoma de México, Campus Ju iquilla,
Bl d Ju iquilla 3001, San iago de Que é a o 76230, México (Medina-Ri e a). Aix-
Ma seille Uni , INSERM UMR S 1090, Theo y and App oaches o Genome
Complexi y (TAGC), F-13288 Ma seille, F ance (J. an Helden ).
E-mail add esses: [email p o ec ed] (J. an Helden), amedina@
liigh.unam.mx (A. Medina-Ri e a).
Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428
jou nal homepage: www.else ie .com/loca e/csbj
may a ec ansc ip ional egula o y mechanisms, and he eby
explain he obse ed pheno ypes.
The Regula o y Sequence Analysis Tools (RSAT, h p:// sa .eu)
[47,26] has es ablished i sel in he las 20 yea s as a majo so -
wa e sui e dedica ed o he analysis o egula o y egions, wi h i e
public se e s suppo ing mo e han 500 euka yo e and 9,000
p oka yo e genomes. Wi h a majo ocus on usabili y and accessi-
bili y o use s wi h o wi hou o mal bioin o ma ics aining, RSAT
p o ides ools o e ie e sequences, pe o m mo i s analysis, e al-
ua e TF mo i quali y, compa e and clus e mo i s, con e ile o -
ma s, e c. He e we desc ibe Va ia ion- ools, a subse o ools
included in RSAT ha enable use s o analyse egula o y a ian s
and assess hei pu a i e impac on TF binding si es.
1.1. Cu en app oaches o de ec ing po en ial egula o y a ian s
The ac ha many a ian s a e loca ed in non-coding egions
igge ed he de elopmen o bioin o ma ic ools o iden i y he
egula o y po en ial o hese gene ic a ian s. S a ing om a lis
o SNPs, compu a ional analyses can help o mula ing hypo heses
on which TF may be impac ed by a gene ic a ian . Howe e , he e
a e nume ous challenges o in silico analysis o un a el he impac
o gene ic a ia ions in gene egula o y egions. Se e al ools and
esou ces ha e been published, p o iding al e na i e me hods o
ackle his p oblem (Table 1). Mos o hem a e ei he based on
pa e n-ma ching app oaches o e alua e he impac o alleles on
TF binding, o on machine lea ning models buil using unc ional
anno a ions o he egula o y egions, e.g. epigenomics and an-
sc ip omics da a. S ill, hese esou ces and ools ha e limi a ions
hampe ing hei usage in se e al o ganisms [68,28,35], on new
anno a ed a ian s [5,63,54], and/o on analyses wi h pe sonal col-
lec ions o TF mo i s [68,28,41].
All ools in he Pa e n Ma ching ca ego y, (labeled PM in
Table1) use Posi ion-Speci ic Sco ing Ma ices (PSSMs) o e alua e
he a ini y o a TF o a gi en sequence wi h an allele. Majo di e -
ences be ween hese ools can be ound in (i) hei a ailabili y:
web pages [5], command line [12] o bo h [69]; (ii) lexibili y o
he use o inpu hei own da a [61]; (iii) usabili y: he possibili y
o use se e al a ian o ma s [28]; (i ) esul s ep esen a ion: ig-
u es and/o ables [63]; ( ) a ailable o ganisms: only human [62],
o o he o ganisms [61]; and ( i) he possibili y o calcula e esul s
on- he- ly [41] o access p e-calcula ed ones [5].
Ano he se o ools (labeled ML in Table1) aim o he iden i i-
ca ion o po en ial egula o y a ian s by in eg a ing se e al ypes
o da a, beyond aking in o accoun po en ial dis up ion o TF bind-
ing. Pa icula ly, Lee, e al. [37] in eg a ed DNaseI-seq da a wi h
SVM app oaches o iden i y a ian s ha could po en ially dis up
TF binding. DeepSea [68] in eg a es unc ional genomic da a om
ChIP-seq, DNaseI-seq, RNA-seq and o he unc ional genomic
high- h oughpu da a o assess he po en ial damage o a ian s
ac oss he human genome. P ecalcula ed esul s o anno a ed
a ian s can be accessed on hei websi e.
Bo h ools can be ained on o he o ganisms, p o ided ha
unc ional genomic da a a e a ailable. The main limi a ion o hese
esou ces is he equi ed expe ise in bioin o ma ics and/o com-
pu a ional esou ces o use s o analyse hei own da a se s. O he
ools iden i y po en ial egula o y e ec s o a a ian by compa ing
he measu ed a ini y o a TF o he di e en possible alleles. Ou
ool, named a ia ion-scan, alls wi hin his ca ego y.
1.2. Va ia ion- ools
In his con ex , we ha e de eloped Va ia ion- ools o add ess he
main limi a ions iden i ied in exis ing p og ams (Table 1).
Va ia ion- ools a e composed o ou p og ams ha enable (i)
e ie al o in o ma ion o Ensembl anno a ed a ian s when a ail-
able o a gi en genome in RSAT ( a ia ion-in o), (ii) con e sions
be ween a ian ile o ma s (con e - a ia ions), (iii) e ie al o
he sequences su ounding a ian s ( e ie e- a ia ion-seq), and
(i ) scanning o di e en alleles o a a ian wi h one o se e al
mo i s, compa ing he sco es and p- alues in o de o iden i y
a ec ed TFBS ( a ia ion-scan)(Fig. 1). Ea lie e sions o hese p o-
g ams we e epo ed in 2015 as pa o a RSAT upda e a icle [45],
hese i s e sions we e de eloped in pe l and we e e ac o ed and
imp o ed o he 2018 upda e [47]. In his a icle we p esen he
la es e sions o he ools, wi h op imized memo y usage, and
no el suppo o he inclusion o haplo ype in o ma ion.
In summa y, RSAT Va ia ion- ools p o ide an accessible esou ce
o expe ienced and non-expe use s o analyze egula o y a i-
an s in a web in e ace o i een o ganisms ( i e me azoa
(h p://me azoa. sa .eu) and en plan s (h p://plan s. sa .eu), wi h
lexibili y o upload pe sonal a ian and PSSM collec ions. We
desc ibe he e Va ia ion- ools me hodology, along wi h ou case
s udies demons a ing he lexibili y o he ools, enabling he anal-
ysis o da a se s om di e en o igin (Ensembl a ian s, Genome-
Wide Associa ion S udy (GWAS) da a, ChIP-seq egions, e c.), com-
plexi y, and o ganisms.
2. Me hods
2.1. Va ia ion- ools: om a ian s o iden i ica ion o egula o y
e ec s
Va ia ion- ools consis in a subse o ou ools wi hin RSAT
de o ed o he iden i ica ion o gene ic a ian s pu a i ely a ec -
ing TF binding
1) a ia ion-in o: his ool elies on he Ensembl gene ic a ia-
ion in o ma ion [29] anno a ed and ins alled on he co e-
sponding se e o each pa icula genome (i.e. human
a ian s a e ins alled in he Me azoa se e ). I can ake
wo di e en inpu s: 1) a ian sID o 2) genomic loci in
bed o ma . This ool will e ie e he in o ma ion o he
a ian s ma ching he IDs o he in o ma ion o he a ian s
loca ed in he genomic loci. Va ian s ins alled in RSAT se -
e s ha e been p ocessed o emo e a ian s wi h incom-
ple e anno a ions (no alleles) o ambiguous coo dina es
(non ma ching alleles coo dina es). When use s ha e hei
own a ian s collec ions, hey can skip his ool and use
di ec ly con e - a ia ions.
2) con e - a ia ions: enables he in e con e sion o a ian ile
o ma s such as VCF, GVF and a Bed. a Bed is an in e nal
o ma o RSAT ha acili a es he e ie al o he sequence
su ounding he a ian (Supplemen a y Fig. 1A).
3) e ie e- a ia ion-seq: e ie es he sequence su ounding
he a ian , and p oduces one sequence o each allele (Sup-
plemen a y Fig. 1B). The ool can ake as inpu a a Bed ile
(see con e - a ia ions). Fo o ganisms wi h Ensembl anno-
a ed a ian s, i can ake a lis o IDs o a bed ile lis ing
genomic loci. The ou pu is p o ided in a o ma named a -
Seq, wi h each ow gi ing one allele wi h i s su ounding
sequence. Each a ian has a speci ic in e nal ID o accom-
moda e se e al a ian s wi h a ious alleles in he same ile.
4) a ia ion-scan: pe o ms he scanning o alleles wi h a PSSM
and compa es he sco es and p- alues be ween alleles o
assess he pu a i e e ec on TF binding (see de ails below)
(Supplemen a y Fig. 2). I equi es as inpu a a Seq ile
(see e ie e- a ia ion-seq), a mo i o collec ion o mo i s
(o e wen y suppo ed ile o ma s), and a backg ound
model ( o me hodological de ails on backg ound model,
e e o [59] Box n°3). Di e en backg ound models a e ead-
1416 W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428
Table 1
Tools simila o a ia ion-scan wi h a ailable implemen a ion. PM s ands o Pa e n Ma ching, ML s ands o Machine Lea ning.
Name PMID Sou ce App oach O ganism Inpu Ou pu Ma ix lexibili y Type Las
upda e
del aSVM 26075791 h p://www.
bee lab.o g/
del as m/
Gapped k-me SVM classi ie . Any o ganism DNaseI-seq da a;
pu a i e egula o y
egions as posi i e
aining se and
andomized sequences as
nega i e aining se .
del aSVM, p edic ed impac o a
a ian in ch oma in accessibili y
which is measu ed by adding up
he con ibu ion o all 10-me s in
which he SNP is p esen o
ch oma in accessibili y.
I can only be ained o
one TF a a ime.
ML, non-
s a ic.
Las
upda e
Sep 2015.
DeepSea 26301843 h p://deepsea.
p ince on.edu/
job/analysis/
c ea e/
Deep con olu ional ne wo k. Human SNPs in VCF o ma . Ch oma in ea u e p obabili ies
o e e ence and al e na i e
alleles, ch oma in ea u e
p obabili y log old changes o
each a ian , ch oma in ea u e
p obabili y di e ences o each
a ian s, e- alues o ch oma in
ea u e e ec s, unc ional
signi icance sco e o each a ian .
The e a e 919 ch oma in ea u es
e alua ed.
I con ains 690 TF binding
p o iles o 160 di e en
TFs, bu does no suppo
he addi ion o new
ma ices.
ML, non-
s a ic.
Las
upda e
May 2017.
a SNP 26092860 h ps://
gi hub.com/
keleslab/a SNP
Impo ance sampling algo i hm
o p- alue calcula ion, i s -
o de Ma ko Model o
gene a e andom backg ound
sequences.
Any o ganism
whose
genome is
included in
he
Bioconduc o
BSGenome
package.
SNP lis , mo i ile. p- alue o binding a ini y wi h
al e na i e and e e ence allele, p-
alue o binding a ini y change
based on log-likelihood a io and
log- ank a io. I also p o ides
composi e logo plo s o di ec ly
isualizing he SNP e ec s on
mo i ma ches.
I accep s se e al
ma ices, and se e al
di e en o ma s. I
includes a mo i lib a y o
2,065 PSSMs om
ENCODE and JASPAR, bu
also allows use -de ined
mo i lib a ies.
PM, non-
s a ic.
Las
upda e
No 2018.
BayesPI-BAR 26202972 h p:// olk.uio.no/
junbaiw/BayesPI-
BAR/
Biophysical modeling o
p o ein-DNA in e ac ion,
es ima ion o TF chemical
po en ial ( h ough a bayesian
nonlinea eg ession model)
and di e en ial binding a ini y.
Any o ganism ChIP-seq expe imen o
TFs o be es ed, DNA
sequences o selec ed
SNPs,PSSMs o selec ed
TFs.
Gi en a SNP and a PSSM lis , i
p oduces wo lis s so ed by
signi icance: one composed o
binding mo i s dis up ed by he
SNP, and one by si es wi h an
inc eased a ini y o he TF caused
by he SNP.
Can use se e al PSSMs
simul aneously.
PM,
biophysical
modeling.
Non-s a ic.
No
upda es
lis ed,
so wa e
c ea ed
July 2015.
GWAS4D 29771388 h p://mulinlab.
mu.edu.cn/
gwas4d/gwas4d/
gwas4d/gwas4d_
se e
Va ian p io i iza ion me hod,
ollowed by an in eg a i e
analysis o genome-wide
associa ion.
Human Accep s VCF-like,
coo dina e only, dbSNP
ID and PLINK-like
o ma s.
Regula o y a ian p io i iza ion
able: includes he mos likely
a ec ed mo i by al e na i e
a ian e ec .
The model includes
mo i s o 1,480
ansc ip ional egula o s
om 13 di e en
esou ces. I is no
possible o upload use -
speci ied ma ices.
PM, s a ic Las
upda e
Sep 2018.
sTRAP 20127973 h p:// ap.molgen.
mpg.de/cgi-bin/
home.cgi
P edic ion o local binding
a ini y ollowed by a
no maliza ion o binding
a ini ies o de e mine
di e ence be ween e e ence
allele and SNP.
O ganisms
a ailable in
TRANSFAC.
Accep s only wo
sequences in FASTA
o ma .
Lis o TFs anked acco ding o
changes induced by he SNP.
The e is no op ion o
use -speci ied ma ices,
ma ices om TRANSFAC
e sions can be selec ed.
PM, non-
s a ic
No
upda es
lis ed,
so wa e
c ea ed in
2011.
(con inued on nex page)
W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 1417
Table 1 (con inued)
Name PMID Sou ce App oach O ganism Inpu Ou pu Ma ix lexibili y Type Las
upda e
SNP2TFBS 27899579 h ps://ccg.ep l.
ch//snp2 bs/
Es ima ion based on PSSM
model.
Human. When wo king wi h he
code, he inpu equi ed
is he e e ence genome,
a SNP ca alogue and a
PSSM collec ion.
The web in e ace
accep s SNP IDs and VCF
o ma , as well as a
speci ica ion o a
genomic egion h ough a
bed ile o by speci ying
he s a and end
posi ions.
Lis o a ec ed TFBSs, so ed by
he magni ude o he e ec s.
On he web in e ace,
only ma ices om
JASPAR can be used.
None heless, i is possible
o download he code
used o gene a e he
da abase and use a
di e en inpu .
PM, s a ic. Las
upda e
July 2017.
a SNP
Sea ch
30534948 h p://a snp.
bios a .wisc.edu/
Used a SNP algo i hm wi h
dbSNP build 144 o human
genome assembly 38 agains
JASPAR and Encode mo i s o
c ea e a eposi o y wi h all he
SNP-mo i combina ions
esul ing om he p e ious
esou ces.
Human. I can ecei e a se o
sIDs, a sID and a
window size a ound he
SOI, genomic coo dina es,
a gene symbol and a
window size a ound he
gene o in e es , o a TF
name.
Table including p- alues o mo i
ma ches o bo h e e ence and
al e na e alleles, as well as he
change in he mo i ma ching and
he di ec ion o said change.
Ou pu includes logo plo s,
displaying he sequence logos
aligned o bes mo i ma ches
wi h e e ence and SNP alleles.
Only JASPAR o ENCODE
ma ices can be selec ed,
and i is possible o selec
only one ansc ip ion
ac o a a ime.
PM, s a ic. Las
upda e Jan
2018.
HaploReg 22064851,
26657631
h ps://pubs.
b oadins i u e.
o g/mammals/
haplo eg/haplo eg.
php
I con ains da a om mul iple
genome anno a ion esou ces.
PSSMs a e sco ed agains
e e ence and al e na i e
alleles, and change in log-odds
is calcula ed.
Human Use s can p o ide a lis o
sIDs o ch omosome
egions. Use s can also
selec GWAS s udies om
he NHGRI ca alog.
P o ides da a on allelic
equencies, conse a ion,
ch oma in s a es, and nea genes.
Fo each o he egula o y mo i s
al e ed by he SNP, i p o ides he
change in log-odds and a logo.
HaploReg con ains a
lib a y c ea ed om
li e a u e sou ces,
TRANSFAC, JASPAR and
PBM expe imen s. The e
is no op ion o use -
speci ied ma ices.
PM, s a ic. Las
upda e
No embe
2015.
RegulomeDB 22955989 h p://www.
egulomedb.o g/
RegulomeDB uses in o ma ion
om se e al da ase s, as well as
manual cu a ion and a heu is ic
me hod o dis inguish be ween
unc ional and non- unc ional
a ian s.
Human. Use s can p o ide a lis o
dbSNP IDs, hg19
coo dina es in BED, VCF
o GFF3 o ma , o hg19
ch omosomal egions in
he same o ma s.
Table so ed by likely
unc ionali y, con aining a ian
coo dina es, sco e assigned by he
algo i hm, and e idence o
unc ion including p o ein
binding, mo i s, ch oma in
s uc u e, eQTLs and his one
modi ica ions.
RegulomeDB includes all
PSSMs om TRANSFAC,
JASPAR CORE, and
UniP obe. The e is no
op ion o use -speci ied
ma ices.
PM, s a ic. No
upda es,
lis ed,
so wa e
c ea ed in
Sep 2012.
mo i b eakR 26272984 h ps://
gi hub.com/
Simon-Coe zee/
Mo i B eakR
I has h ee op ions o
algo i hms: he s anda d sum
o log p obabili ies, weigh ed
sum, and an in o ma ion
con en me hod.
O ganisms
included in
BSgenome.
SNPs can be impo ed
om an R package o
p o ided o he algo i hm
in BED o VCF o ma .
PSSMs can be selec ed
om he Mo i Db
package o be use -
speci ied.
Table con aining s a is ics
desc ibing he pe cen o
maximum sco e o a ma ix and
ma ix alues o bo h alleles, as
well as he s and. I also epo s
whe he he TFBS is dis up ed
s ongly o weakly.
PSSMs can be impo ed
om he Mo i Db
package o be use -
speci ied. Mo e han one
ma ix can be used a a
ime.
PM, non-
s a ic.
Las
upda e Jul
2018.
a ia ion-
scan
h p:// sa .eu Es ima ion based on PSSM
model.
web
in e ace:
ins alled
Ensembl
o ganisms.
command-
line: any
locally
ins alled
o ganism.
A collec ion o PSSMs and
a se o a ian s in a Seq
o ma . This o ma can
be ob ained using
e ie e- a ia ion-seq.
A able wi h one line pe pai o
alleles pe mo i (i he e a e mo e
han wo, he e will be one line
pe possible pai ) epo ing he
posi ion, weigh and p- alue o
each allele, weigh di e ence and
p- alue a io.
Use s can selec o he
collec ions a ailable in
RSAT (JASPAR,
HOCOMOCO, CisBP), bu
hey can also use
pe sonal collec ions.
PM-non
s a ic.
Ap il 2019.
1418 W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428
ily a ailable h ough he web in e ace. Howe e , depending
on he biological ques ion and ela ed po en ial biases, we
ecommend he c ea ion o a dedica ed backg ound model,
which can be done using he RSAT ool c ea e-backg ound,
also a ailable ia he RSAT web in e ace.
2.2. Haplo ype p ocessing
Gene ic a ian s can be de ec ed using high- h oughpu ech-
niques. This has enabled he iden i ica ion o millions o a ian s
in he HapMap [30] and 1000 genomes p ojec s [1]. Howe e , he
in o ma ion on he a ian s alone is less use ul han knowing
which g oups o alleles a e co-loca ed on he same ch omosome
(haplo ype). The p ocess o iden i ying he a ian s ha belong
o each ch omosome is known as phasing. Including haplo ype
phasing in o ma ion acili a es he iden i ica ion o ela ions
be ween a ian s [6].
VCF iles can include haplo ype phasing in o ma ion. The ool
con e - a ia ions iden i ies and e ie es he phasing in o ma ion
o he a ian s, while he ool e ie e- a ia ion-seq buil s he co -
esponding haplo ype wi h all he SNPs ha lay wi hin a de ined
window (de aul : 30 bp).
2.3. Compu ing binding speci ici y o a ansc ip ion ac o o a DNA
sequence
a ia ion-scan uses PSSMs o assess he binding speci ici y o a
TF o a DNA sequence wi h di e en alleles in a gi en posi ion.
The i s s ep o a ia ion-scan (i.e., scanning o he sequences wi h
a gi en PSSM) is delega ed o he RSAT ool ma ix-scan. The sco -
ing scheme and p- alue calcula ion a e desc ibed in de ail in [59],
Box n°1 and Box n°2, espec i ely. In b ie :
PSSM a e used o assess he binding speci ici y o a TF. This
a ini y is calcula ed as a weigh sco e (Ws). The Ws o a si e in
a ia ion-scan is calcula ed using [27]:
Ws ¼lnðPðSjMÞ
PðSjBÞ
whe e S is a sequence segmen o he same leng h o M, M is he
PSSM, and B is he backg ound model. Hence, P(S|M) is he p obabil-
i y o he sequence gi en he PSSM and P(S|B) is he p obabili y o
he sequence gi en he backg ound model. Ws has been ela ed o
he a ini y o he TF o he sequence, as i assesses simila i y o a
sequence o a known se o binding si es, p o iding in o ma ion
abou he p obabili y o a sequence o be a new ins ance o a bind-
ing si e [56].
Mo eo e , i is possible o calcula e he p- alue o a gi en sco e
as:
P
alue ¼PðWwjBÞ
whe e he P- alue is calcula ed as he p obabili y o obse ing a
sco e o a leas Wgi en a backg ound model w|B.
When a sequence is longe han he PSSM, he PSSM is shi ed
base by base un il he ull sequence has been sco ed. This scanning
s ep is pe o med on he sequences o all epo ed alleles, so ha
each allele is compa ed wi h all he posi ions o a gi en mo i .
Backg ound models ep esen he nucleo ide composi ion o a
se o sequences (whole genome, all p omo e sequences, e c.).
These models a e used o es ima e he expec ancy o a nucleo ide
being ound. Backg ound models can ep esen dependency
be ween nucleo ides in sequences (e.g. aking in o accoun he e-
Fig. 1. Schema ic ep esen a ion o Va ia ion- ools: This se o ools, included in he Regula o y Sequence Analysis Tools (RSAT), ocuses on assessing he impac o di e en
allelic a ian s on T ansc ip ion ac o binding si es. A) con e - a ia ions allows use s o inpu hei own a ian s and con e hem o o he o ma s (VCF, GVF and a Bed,
he la e is he o ma used in he nex s ep), while a ia ion-in o e ie es he anno a ed in o ma ion o Ensembl a ian s ins alled in RSAT se e s. B) The ool e ie e-
a ia ion-seq e ie es he su ounding sequence o a ian s (including possible haplo ypes) and gene a es a ex ile wi h one line pe allele and pe a ian o haplo ype
( a Seq o ma ). C) Use s can inpu hei a ian s in a Seq o ma and a collec ion o mo i s (di ec inpu by he use o selec ed om RSAT a ailable collec ions) o a ia ion-
scan; he ool hen scans he co esponding sequences wi h all mo i s and pe o m pai wise compa isons be ween he binding sco es o each ansc ip ion ac o on o all
alleles o a a ian o haplo ype.
W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 1419
quencies o dinucleo ides o build a Ma ko model o o de 1 [59]
Box n°3). As backg ound models a e used o calcula e weigh
sco es o a binding si e (P(S|B)), i is impo an o selec an app o-
p ia e model o each analysis. Examples o selec ed backg ound
models a e p esen ed in he di e en s udy cases ma ching each
pa icula biological ques ion.
2.4. Assessmen o allele e ec on ansc ip ion ac o binding
In he second s ep, i.e., e alua ing he impac o SNPs, a ia ion-
scan compa es he ob ained Ws (Ws di e ence = Ws_Allele1 –
Ws_Alelle2) and he P- alue (P- alue a io = P- alue_Allele1/P- al
ue_Allele2) o each o he alleles, posi ion by posi ion h oughou
he scanning window. To e alua e indels, a ia ion-scan compa es
he highes Ws and i s co esponding P- alue o each sequence o
he epo ed alleles. When mo e han wo alleles o a a ian a e
epo ed, all alleles a e compa ed o all alleles in a pai wise
manne .
2.5. a ia ion-scan pe o mance es
2.5.1. Compu ing e iciency
The ools a ia ion-in o and con e - a ia ion a e coded in Pe l,
while e ie e- a ia ion-seq and a ia ion-scan a e coded in C, o
enable he analysis o la ge numbe s o a ian s om euka yo ic
genomes in a easonable ime. To u he imp o e pe o mance,
we educed he da a ans e om he ha d d i e o memo y.
a ia ion-scan pe o mance was assessed by andomly selec ing
a a ian om he 1000 genomes p ojec [1] and a mo i om he
RSAT non- edundan mo i s collec ion [8]. The andomly selec ed
a ian was used o c ea e se s wi h di e en numbe s o epli-
ca es, anging om one housand o nine millions, o es ima e
he ela ion be ween unning ime and he amoun o e alua ed
a ian s. The p ocesses we e un on a Dell Powe Edge C6145 se e
wi h 2 AMD Op e on( m) P ocesso 6386 SE, 16 co es each, P oces-
so speed o 2.8–3.5 Ghz, RAM 256 Gb and wi h an ope a ing sys-
em Cen OS 7 (7.6.1810).
2.5.2. Da ase : expe imen ally-de e mined egula o y a ian s in ed
blood cells
The egula o y ac i i y o 2,756 ed blood cell a ian s has been
sys ema ically measu ed using Massi ely Pa allel Repo e Assays
(MPRA) [60]. MPRA is a high- h oughpu assay in which a lib a y
o pu a i e egula o y elemen s, each ollowed by a unique ba -
code, is inse ed in o a plasmid, hen ans ec ed in o a cell, and
ansc ip s a e hen quan i ied h ough he abundance o ba codes.
These a ian s a e known o be in s ong linkage disequilib ium
(LD) wi h 75 a ian s associa ed wi h common ai s o his cell
ype. Th ee sliding windows pe a ian (le , igh , and cen e )
we e syn hesized, ba coded and used o s udy he e ec o sligh
changes in hei genomic con ex . Following me hods desc ibed
by Uli sch, e al.[60], o each sequence mRNA/DNA a io was com-
pu ed o ob ain a quan i a i e e alua ion o he egula o y e ec o
a sequence a ian .
2.5.3. E alua ion o a ia ion-scan
The a ian da ase was used as inpu o a ia ion-scan; he
a ian s assessed in he ed blood cell assay we e anno a ed wi h
he Ensembl GRCh37 human genome elease, and gi en as inpu
o con e - a ia ions ollowed by e ie e- a ia ion-seq. Since h ee
sliding windows we e used o each a ian in he MPRA, he co -
esponding windows we e me ged be o e compu ing a backg ound
model using he c ea e-backg ound-model ool.
Acco ding o he o iginal s udy [60], binding si es o he ol-
lowing TF we e en iched in he sequences o in e es : GATA1,
KLF1, DHS, TAL1, ETS, FLI1 and AP-1. The e o e, a o al o 48 PSSMs
anno a ed as ela ed o hese TF we e e ie ed om he non-
edundan RSAT mo i collec ion [8], and gi en as inpu o
a ia ion-scan.
A nega i e con ol se o mo i s was c ea ed using he RSAT ool
pe mu e-ma ix [47]; i e pemu ed mo i s we e c ea ed o each o
he 48 mo i s, gene a ing a collec ion o 240 con ol mo i s.
Fo a a ian o be epo ed in a ia ion-scan as posi i e,we
eques ed ha a leas one o he allele sequences was e alua ed
as a binding si es wi h a p- alue o a mos 10
4
(using he pa am-
e e -u h p al 1e-4 in he command line), and ha he p- alue a io
was g ea e o equal o en (a change o one o de o magni ude
be ween he bes and he wo s allele p- alues) (-l h p al_ a io
10).
We compa ed a ia ion-scan o wo o he ools p e iously used
o assess he same se o a ian s by Uli sch, e al. [60]: DeepSea
[50] and [37] del aSVM. In o de o a oid pe sonal biases when cal-
ib a ing ool pa ame e s, we decided o ely on he published ones
[60]. Fo his analysis a ia ion-scan was un wi hou h esholds o
iden i y he impac o he pa ame e s, pa icula ly he h eshold on
p- alue a io.
2.6. Case s udies
2.6.1. Case s udy 1: Iden i ica ion o egula o y a ian s in he
‘‘Pla inum”genomes haplo ypes
The se o high-con idence a ian s om he wo CEU (No he n
Eu opeans om U ah) human Pla inum Genomes NA12877 and
NA12878 [23] we e downloaded h ough he Amazon Web Se ice
(AWS) Command Line In e ace om he Illumina Pla inum Gen-
omes AWS S3 bucke (h ps://gi hub.com/Illumina/Pla -
inumGenomes). The downloaded VCF iles con ained phasing
in o ma ion o each CEU indi idual haplo ype con igu a ion. The
genome e sion used was GRCh37.
We selec ed SNPs in e sec ing wi h he anno a ed DNAseI-seq
clus e ed peaks V3 om he ENCODE p ojec [4]. The VCF ile wi h
he selec ed SNPs was p ocessed using con e - a ia ions wi h he
op ion phased and hen he haplo ype sequences we e econ-
s uc ed wi h e ie e- a ia ion-seq.
Fo a haplo ype SNP se o single posi ion a ian s o be
epo ed in a ia ion-scan, we eques ed ha a leas one o he
sequences was e alua ed as a binding si e wi h a p- alue o a mos
10
4
(-u h p al 1e-4) and ha he p- alue a io be ween he wo
alleles was g ea e o equal o 100 (a change o wo o de s o mag-
ni ude be ween he bes and he wo s alleles p- alues) (-l h
p al_ a io 100). In addi ion, we equi e a change o sign be ween
he bes and wo s sco e as an addi ional il e .
We anno a ed he p edic ed dis up ed TFBS wi h he TF ChIP-
seq non- edundan peak collec ion and wi h he Cis-Regula o y
Modules (CRM) egions om ReMap [10] using bed ools in e sec
e sion 2.27 [49]. We also calcula ed he en ichmen o anno a-
ions in he p o enance sequence segmen s o he p edic ed haplo-
ypes si es.
2.6.2. Case s udy 2: p edic ion o egula o y a ian s associa ed wi h
suscep ibili y o Mycobac e ium ube culosis in ec ion
We collec ed SNPs associa ed wi h he pheno ypic ai ‘‘suscep-
ibili y o Mycobac e ium ube culosis in ec ion measu emen ” (dis-
ease ID EFO_0008407) om he 1.0.2 e sion o he GWAS
ca alog [40] (h ps://www.ebi.ac.uk/gwas/). This que y e u ned
one s udy [58] wi h 67 dis inc a ian s, o which 48 had a alid
e e ence SNP iden i ie ( sID) and could be u he used (deno ed
he ea e as disease-associa ed SNPs, o DA-SNPs). To p edic he
TF binding si es pu a i ely a ec ed by hese selec ed SNPs, we
designed an app oach combining Va ia ion- ools wi h di e en
ex e nal esou ces. We u he collec ed om Ensembl REST in e -
ace (h p:// es .ensembl.o g/) 564 SNPs in linkage disequilib ium
1420 W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428
(LD-SNPs) in he Eu opean popula ion [62], wi h a h eshold on he
eg ession coe icien (
2
0.8) and a maximal dis ance o 200 bp.
Anno a ions (ch omosomal loca ion, ype o genomic egion) o
he esul ing 612 SNPs (48 DA + 564 LD) we e collec ed om
Ensembl BioMa [22,21]. We hen es ic ed he selec ion o SNPs
in non-coding egions, esul ing in a se o 572 SNPs o in e es
(SOIs) o he de ec ion o egula o y a ian s. Using SNPs in LD,
we de e mined LD-Block egions. These we e hen anno a ed based
on o e laps wi h ChIP-seq peaks collec ed om he ReMap da a-
base [10]. We also calcula ed en ichmen o disease anno a ions
using he R XGR package [24].
Finally, we used e ie e- a ia ion-seq o e ie e he sequence
a ian s a ound each SOI, and p edic ed he impac o he a ia ion
on TF binding o each mo i o he JASPAR non- edundan RSAT
mo i collec ion [8] using a ia ion-scan, wi h he h esholds o
1e-4 on he p- alue and 100 on he p- alue a io.
2.6.3. Case s udy 3: Assessmen o he egula o y e ec o GWAS
epo ed a ian s in p omo e s wi h enhance unc ion
The STARR-seq assay [2] is in i s p inciple simila o he MPRA,
and helps iden i y sel - ansc ibing ac i e egula o y egions ha
ha e enhance po en ial. Using his app oach Dao e al. [17], anal-
ysed he enhance po en ial o anno a ed Re Seq p omo e s [48].In
he wo cell lines K562 and HELA, hey iden i ied 632 and 493 p o-
mo e s wi h enhance po en ial (eP omo e s), espec i ely. Mo e-
o e , he au ho s iden i ied en ichmen o eQTL a ian s epo ed
by GTEx [25].
To iden i y eP omo e s a ian s ha could be a ec ing TF bind-
ing, we e ie ed he GWAS ca alog e sion 1.0 (downloaded on
7/01/19) [40]. Using bed ools o e lap e sion 2.26.0 [49], we com-
pu ed he o e lap be ween SNPs and he eP omo e coo dina es
epo ed in [17]. PSSMs ep esen ing TF en iched in eP omo e s
we e also ob ained om [17], co esponding o SMRC1, JUN, FOS,
ATF:MAF:NEF2, YY1, ETS amily, C eb and USF1/2.
Using he selec ed GWAS a ian s ha all wi hin eP omo e s
and he TF mo i s en iched in hese egions, we applied a ia ion-
scan o assess he po en ial egula o y e ec o hese a ian s.
a ia ion-scan was un wi h he pa ame e s – l h w_di 1 – l h
p al_ a io 10, wi h a backg ound model buil using c ea e-
backg ound wi h all Re Seq p omo e sequences. In o de o il a e
a ian s wi h he highes pu a i e egula o y dis up ion, we u -
he selec ed a ian s ha showed a change o sign in he weigh
sco e be ween alleles.
2.6.4. Case s udy 4: iden i ica ion o egula o y a ian s a ec ing VRN1
binding in ba ley
The la es e sion o Ho deum ulga e (ba ley) e e ence gen-
ome [42] and a panel o mapped gene ic a ian s we e impo ed
om Ensembl Genomes elease 42 [34] and ins alled in he RSAT
Plan s se e (h p://plan s. sa .eu). We ob ained expe imen ally
de e mined binding si es (ChIP-seq) o VRN1 om [19]. Since
hese peaks we e o iginally posi ioned wi hin con igs o he 2012
genome assembly [14], hey had o be ma ched o he co espond-
ing egions o he cu en assembly wi h BLAST + 2.9.0 (blas n)
local alignmen s agains he epea -masked genome sequence
(pe ec ma ches) [7]. Using bed ools o e lap e sion 2.26.0 [49],
we selec ed a ian s alling wi hin he VRN1 epo ed binding
peaks. The selec ed a ian s in VCF o ma we e hen p ocessed
using con e - a ia ions and e ie e- a ia ion-seq o ob ain he
sequences wi h he al e na i e alleles.
The VRN1 DNA mo i used o scan he a ian s was ob ained
om he oo p in DB plan collec ion [16] e sion: 2018-06
(h p:// lo es a.eead.csic.es/ oo p in db/index.php?mo i =
AY750993:VRN1:EEADanno ). a ia ion-scan was used wi h a p e-
compu ed backg ound Ma ko model (o de 1) o ba ley o assess
he e ec o a ian s in TF binding, wi h he ollowing pa ame e s:
– l h sco e 1 – l h w_di 1 – l h p al_ a io 10 – u h p al 1e-3.
2.7. A ailabili y
Va ia ion- ools a e a ailable on he web (Me azoa: h p://me a-
zoa. sa .eu/, Plan s: h p://plan s. sa .eu/, Teaching: h p:// each-
ing. sa .eu/). The ools can be also ins alled o command-line
usage wi h he RSAT sui e (h p://download. sa .eu/).
The code and ma e ial o ep oduce he esul s p esen ed in he
a icle can be accessed h ough Gi Hub (h ps://gi hub.com/RSAT-
doc/supp-ma e ial-publica ions.gi ).
3. Resul s
The Va ia ion- ools p o ide complemen a y p og ams enabling
he e ie al o a ian s ( a ia ion-in o) and o hei su ounding
sequences ( e ie e- a ia ion-seq), as well as in e con e sion
be ween ile o ma s (con e - a ia ion). The main p edic i e p o-
g am is a ia ion-scan, which can be used wi h any se o a ian s
p o ided by he use (in VCF o GVF o ma s) o anno a ed in
Ensembl ( om a lis o sIDs o a bed ile o iden i y o e lapping
a ian s in genome coo dina es), wi h any se o mo i s selec ed
om he collec ions a ailable in RSAT, o p o ided by he use .
3.1. a ia ion-scan accu a ely assesses he e ec o expe imen ally
alida ed egula o y a ian s
The o iginal e sion o a ia ion-scan [45] equi ed app oxi-
ma ely i e hou s o assess he allele e ec o nine millions a i-
an s. The no el e sion [47] signi ican ly educes he p ocessing
ime o abou one hou (Supplemen a y Fig. 3).
To e alua e he pe o mance o a ia ion-scan, we used an
expe imen ally alida ed egula o y a ian se ob ained om a
MPRA expe imen [60]. Fo all o he assessed allele pai s, we com-
pa ed he weigh sco e di e ences compu ed wi h a ia ion-scan
wi h he mRNA/DNA a io o he MPRA (see me hods). As shown
in Supplemen a y Fig. 4A, we a e able o eco e only 9.37% o
he expe imen ally alida ed a ian s wi h a ia ion-scan, as we
eques ed a leas one o he alleles o ha e a binding si e o high
con idence (p- alue 10
4
). Focusing on he a ian s epo ed as
posi i e in he MPRA da a se , we obse ed a weak co ela ion
be ween he weigh di e ence and he MPRA mRNA/DNA a io in
posi i e a ian s. Howe e , his co ela ion is no signi ican , as
MPRA alues do no scale wi h he a ia ion-scan weigh di e -
ences. Ne e heless, all a ian s show a p- alue a io indica i e
o allele binding e ec s, showing ha a ia ion-scan gi es accu a e
measu emen s o he impac o egula o y a ian s (Fig. 2A).
Wi h he p oposed h esholds, we can con iden ly ejec 96.35%
o MPRA nega i e sequences, which could be imp o ed using mo e
es ic i e pa ame e , wi h a concomi an educ ion in ue posi-
i es. No ewo hy, as any high- h oughpu assay, MPRA has i s
limi a ions [51] and sequencing biases could inc ease he numbe
o alse nega i es.
We pe o med a nega i e con ol, consis ing o 240 pe mu ed
ma ices ( i e pe mu ed e sions o he 48 mo i s). Wi h his col-
lec ion, i was s ill possible o eco e a g oup o a ian s, bu i
only ep esen ed 31.2% o he MPRA posi i e a ian s (Supplemen-
a y Fig. 4B).
We compa ed he pe o mance o a ia ion-scan o wo o he
ools ha had been p e iously used by Uli sch, e al [60] o assess
he same se o MPRA a ian s: DeepSea [68] and del aSVM [37].
We decided o use he same pa ame e s in o de o a oid pe sonal
biases when calib a ing he ools. The e o e, aining weigh s o
DNAse I hype sensi i i y si es we e used in he del aSVM analysis.
W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 1421
As o DeepSea, he web implemen a ion o he ool was used, wi h
he me ic Func ional Signi icance Sco e.
Tools we e compa ed based on ROC cu es (Fig. 2B), we addi-
ionally an a ia ion-scan using a se o pe mu ed ma ices as neg-
a i e con ol (Fig. 2B, ed line). The h ee ools show e y simila
sensi i i y s speci ici y a he beginning o he cu es, bu only
a ia ion-scan and DeepSea u he emain sepa a ed om he neg-
a i e con ol. As expec ed DeepSea pe o ms sligh ly be e han
a ia ion-scan a he beginning o he cu e, ne e heless his ool
equi es aining using epigene ic da a, while a ia ion-scan
equi es only a mo i and a se o a ian s.
3.2. Va ia ion- ools case s udies
To illus a e he di e se applica ions o Va ia ion- ools o ackle
a ious biological ques ions, we designed ou di e en case
s udies:
1. Impac o egula o y a ian s in he same haplo ype on TF bind-
ing si es.
2. Iden i ica ion o he egula o y po en ial o a ian s epo ed in
GWAS.
3. Assessmen o he egula o y po en ial o GWAS a ian s wi hin
expe imen ally de e mined egula o y egions.
4. De e mina ion o egula o y a ian s wi hin TF binding egions
iden i ied using ChIP-seq [19].
3.2.1. Genome-wide haplo ype a ian in o ma ion can be used o
iden i y se s o egula o y a ian s a ec ing he same TFBS
The lowe ing cos s in sequencing ha e made i possible o
ob ain whole genome sequences o mo e indi iduals, opening he
possibili y o knowing, no only he a ian s o a genome, bu also
he haplo ypes, and de e mining which a ian s a e passed linked
wi hin he same ch omosome. This enables he assessmen o he
egula o y e ec s o se s o a ian s wi hin he same haplo ype
in a gi en TFBS.
Using he high-con idence SNPs om wo ‘‘Pla inum” Genomes
[23], we de e mined haplo ype a ian s ha a e likely o a ec one
TFBS. We selec ed a ian s 30bps apa , loca ed in open ch oma in,
o be analysed wi h a ia ion-scan using he non- edundan mo i
collec ion a RSAT [8]. We de ec ed 7,406 haplo ype si es wi h a
leas wo he e ozygous a ian s and a p obable e ec in binding
o 361 TFs. O e all he numbe o he e ozygous a ian s wi hin a
haplo ype inc eases he measu ed weigh di e ence. This is
expec ed as mo e changes in he binding si es a e mo e likely o
change TF a ini y (Fig. 3A).
To assess he biological ele ance o all he pu a i e dis up ed
TFBS p edic ions, we anno a ed 7,485 p edic ed haplo ypes si es
con aining wo o mo e a ian s wi h a leas one he e ozygous
a ian and 15,396 p edic ed si es con aining a SNP (single ons)
wi h he TF ChIP-seq peaks and he Cis-Regula o y Modules
(CRM) egions om ReMap [10]. We ound ha almos all he p e-
dic ed dis up ed TFBS (~85%) con ain a CRM o peak anno a ion o
bo h (Fig. 3B). In e es ingly, we ound en ichmen o CRM and peak
anno a ions in he p o enance sequence segmen s o he 7,485
p edic ed haplo ypes si es compa ed o he p o enance sequence
segmen s o he single a ian s (Fishe exac es , p- alue < 2.2e-
16).
One o hese anno a ed haplo ypes is composed o he mino
alleles o wo SNPs ( s2732317 and s2732318), whe e we
obse ed a po en ial egula o y e ec likely a ec ing h ee binding
mo i s, o EHF/ELF2, ETV4/ELK1/ETS1/FLI1/ELK4/ETS2/FEV/GABP1,
and ELK3/ELF1/ERG/GABPA (Fig. 3C).
3.2.2. Gene ic a ian s associa ed wi h Mycobac e ium ube culosis
in ec ion show po en ial egula o y e ec s
The second case s udy illus a es a knowledge- ee use o
Va ia ion- ools o iden i y egula o y a ian s om GWAS s udies
o a use -speci ied disease, wi hou p io indica ion abou he
po en ially in ol ed ansc ip ion ac o s o binding mo i s. The
app oach is based on he p edic ion o egula o y a ian s wi h
RSAT Va ia ion- ools, na owed down by selec ing he egula o y
SNPs ha o e lap ChIP-seq peaks in ReMap [10], in o de o iden-
i y con e gen indica ions o a po en ial impac o he a ian s on
he binding o a TF.
A) B)
0.00
0.25
0.50
0.75
1.00
0.000.250.500.751.00 speci ici y
sensi i i y
name
pe mu ed
del aSVM
a ia ion-sca
n
DeepSea
P- alue a io 100
P- alue a io 1000
P- alue a io 10
R=0.12,p=0.5
0
200
400
−9 −6 −3 0
MPRA p− alue o di e en ial ac i i y (log10)
Va ia ion−scan p− alue a io
Fig. 2. Iden i ica ion o expe imen ally alida ed egula o y a ian s using a ia ion-scan. A) Co ela ion o he Massi ely Pa allel Repo e Assays (MPRA) p- alue o he
mRNA/DNA a io o posi i e a ian s and he a ia ion-scan weigh di e ence o he MPRA a ian s wi h signi ican change. B) Recei e Ope a ing Cha ac e is ic (ROC) cu e
compa ing he pe o mance when aiming o classi y MPRA expe imen ally analyzed a ian s using a ia ion-scan ( u quoise), DeepSea (pu ple), del aSVM (g een), and a
nega i e con ol which consis s o pe mu ed mo i s sco ed wi h a ia ion-scan ( ed).
1422 W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428
Fig. 3. Haplo ype analysis in high-quali y human genomes. A) The numbe o he e ozygous a ian s (X-axis) wi hin he same pu a i e binding si e end o ha e a g ea e
impac on he TF binding p obabili y. This is expec ed as he inc ease o weigh di e ence obse ed on he iolin plo co esponds o he expec ed cumula ed impac o
a ia ions a ec ing di e en posi ions o he same binding si e. B) Numbe o p edic ed dis up ed T ansc ip ion Fac o Binding Si es (TFBSs) wi h Cis-Regula o y Modules
(CRMs) and TF ChIP-seq peak anno a ion (blue), wi h only peak anno a ion (yellow), and non-anno a ed p edic ions (g ey). C) Uni e si y o Cali o nia San a C uz (UCSC)
b owse [48] sc een sho , showing a locus encompassing wo SNPs ha compose an he e ozygous haplo ype in one o he No he n Eu opeans om U ah (CEU) indi iduals.
The igu e shows he e e ence genome haplo ype. The a ian s a e loca ed in he FUT10 p omo e ( op). a ia ion-scan p edic s an e ec in h ee mo i s ha ep esen
binding si es o GABPA, ETS1 and ELF2, ac o s ha ha e been p o en o ha e binding si es in his egion by he ENCODE p ojec . The a ian s2732317 has been associa ed
wi h e ec s in gene exp ession by he GTEx p ojec .
W. San ana-Ga cia e al./Compu a ional and S uc u al Bio echnology Jou nal 17 (2019) 1415–1428 1423