New compu a ional me hods
o s uc u al modeling p o ein-p o ein
and p o ein-nucleic acid in e ac ions
Luis Ángel Rod íguez Lumb e as
Aques a esi doc o al es à subjec a a la llicència Reconeixemen - Compa Igual 4.0. Espanya
de C ea i e Commons.
Es a esis doc o al es á suje a a la licencia Reconocimien o - Compa i Igual 4.0. España de
C ea i e Commons.
This doc o al hesis is licensed unde he C ea i e Commons A ibu ion-Sha eAlike 4.0. Spain
License.
FACULTAT DE BIOLOGIA
(Supe iso : Juan Fe nández Recio)
New compu a ional me hods o s uc u al modeling o p o ein-
p o ein and p o ein-nucleic acid in e ac ions
LUIS ÁNGEL RODRÍGUEZ LUMBRERAS
UNIVERSITAT DE BARCELONA
FACULTAT DE BIOLOGIA
DOCTORAT EN BIOMEDICINA
Código HDK05
New compu a ional me hods o s uc u al modeling o
p o ein-p o ein and p o ein-nucleic acid in e ac ions
Memo ia p esen ada po Luis Ángel Rod íguez Lumb e as pa a op a al í ulo de Doc o po
la Uni e sidad de Ba celona
Tesis doc o al ealizada bajo la di ección del Doc o Juan Fe nández Recio, en el Ba celona
Supe compu ing Cen e (BSC) y los dos úl imos años en el Ins i u o de Ciencias de la Vid y
del Vino (ICVV-CSIC).
Tesis adsc i a al Depa amen o de Bioquímica y Biología Molecula de la Facul ad de
Biología (Uni e si a de Ba celona).
Di ec o Tu o Ph.D. candida e
D . Juan Fe nández Recio D . Josep Lluís Gelpí Buchaca Luis Ángel Rod íguez Lumb e as
Ba celona, 2022
i
DECLARACIÓN DE ORIGINALIDAD
Yo, LUIS ANGEL RODRIGUEZ LUMBRERAS, ma iculado en el p og ama de doc o ado de
BIOMEDICINA de la Uni e sidad de Ba celona decla o que la esis i ulada “New
compu a ional me hods o s uc u al modeling o p o ein-p o ein and p o ein-nucleic acid
in e ac ions” es o iginal, que la in es igación que oy a ealiza cumple con los códigos
é icos y las buenas p ác icas y que la esis no incluye plagio. Soy conscien e y acep o po
esc i o que mi esis se some e á al p oceso adecuado pa a p oba la o iginalidad de mis
esul ados.
31 de Oc ub e de 2022
ii
A mis pad es
iii
"Ninguna ciencia, en cuan o a ciencia, engaña; el engaño es á en quien no la sabe."
Miguel de Ce an es
x
6.2.2. P edic i e success o empla e-based docking ............................................................... 70
6.3. Ab ini io docking can imp o e empla e iden i ica ion .......................................................... 71
6.3.1. A new p o ocol o combining ab ini io and empla e-based docking ........................... 71
6.3.2. E alua ion o he combined docking p o ocol. ............................................................... 76
6.4. Docking unc ions can be used o sco e models om empla e-based docking ................... 78
6.4.1. Au oma ic p ocessing o inpu s uc u es ....................................................................... 78
6.4.2. Templa e-base docking ................................................................................................... 79
6.4.3. Ab ini io docking .............................................................................................................. 79
6.4.4. Combining ab ini io and empla e-based me hods......................................................... 80
6.4.5. Inclusion o es ain s om a ailable ex e nal da a ....................................................... 82
6.4.6. Final selec ion o he models. ......................................................................................... 84
6.4.7. Sco ing o p o ein-saccha ide complex models .............................................................. 86
6.5. E alua ion o he p edic i e esul s in CAPRI and CASP ......................................................... 87
6.5.1. 7 h CAPRI ......................................................................................................................... 88
6.5.2. 3 d join CASP-CAPRI expe imen ..................................................................................... 91
6.5.3. 4 h join CASP-CAPRI expe imen ................................................................................... 94
6.6. P o ein docking unc ions o he iden i ica ion o physiological homodime s..................... 96
6.6.1. Applicabili y o pyDock and CCha PPI sco ing unc ions o he iden i ica ion o
physiological homodime s. ....................................................................................................... 98
7. APPLICATION TO A CASE STUDIO: MODELING ELECTRON TRANSFER PROTEIN COMPLEXES. ... 101
7.1. In oduc ion ......................................................................................................................... 103
7.2. Molecula s uc u es and modelling .................................................................................... 103
7.3. P o ein-p o ein docking simula ions .................................................................................... 104
7.3.1. Docking sampling and sco ing ....................................................................................... 104
7.3.2. Minimiza ion o he p o ein-p o ein docking poses ..................................................... 109
7.4. Docking s uc u es o [C :Cc6] and [C :Pc] in ed Algae lineage ........................................... 110
7.6. Conclusions .......................................................................................................................... 115
8. Gene al discussion....................................................................................................................... 117
9. Conclusions.................................................................................................................................. 121
10. Appendices ................................................................................................................................ 123
10.1. Appendix 1. Supplemen a y ma e ial o Chap e 5 .......................................................... 125
10.2. Appendix 2. Supplemen a y ma e ial o Chap e 6 .......................................................... 127
10.3. Appendix 3. Supplemen a y ma e ial o Chap e 7 .......................................................... 140
Bibliog aphy .................................................................................................................................... 142
xi
1
1. INTRODUCTION
2
3
1.1. Biomolecules
1.1.1. F om genes o p o eins
Li ing o ganisms can be seen as da a s o age machines ha igh agains en opy o
conse e and ansmi in o ma ion. As a consequence, he classic unc ions ha de ine a
li ing o ganism, i.e. nu i ion, ela ionship, and ep oduc ion, eme ge om his s uggle.
The e o e, he molecule esponsible o s o ing he in o ma ion mus be s able and
ha e a la ge s o age capaci y. Chemical and biological e olu ion selec ed deoxy ibonucleic
acid (DNA) as such molecule [1]. I is a polyme o med by deoxy ibonucleo ides. The
deoxy ibonucleo ides a e composed o a ni ogenous base, a deoxy ibose (suga ), and a
phospha e g oup. The ni ogenous base can be de i ed om pu ine, such as adenine (A)
and guanine (G), o om py imidine, such as cy osine (C) and hymine (T). Phospha e g oups
o m he backbone o DNA, binding he nucleo ides oge he and gi ing di ec ionali y o he
polyme ic chain. The oxygen in he phospha e is co alen ly a ached o he C5' posi ion o
he suga . The phospha e g oup hen binds o he nex nucleo ide a he suga 's C3' posi ion,
gi ing a 5' o 3' di ec ionali y and o ming a single-s anded polyme .
Howe e , his polyme ic molecule alone would no be s able enough o long ime
in o ma ion s o ing. This s abili y is achie ed by non-co alen binding o wo single-
s anded DNA polyme s, which can adop a s able double helix con o ma ion, acco ding o
a highly speci ic base complemen a i y. The bases A and T mus ace each o he by o ming
wo hyd ogen bonds, while G and C bases ace each o he by o ming h ee hyd ogen bonds.
The epe i ion o hese base pai s builds a double-s anded DNA. The in o ma ion is
encoded in iple s ( h ee consecu i e nucleo ides o he sequence) o codons, which code
up o 20 ypes o amino acids. Amino acids o m a di e en kind o polyme , p o eins
1.1.2. DNA s uc u e
The combina ion o Cha ga 's ules (A+G=T+C) and X- ay s udy o he sodium sal o DNA
o Rosalind F anklin, lead James Wa son and F ancis C ick o disco e he DNA double helix
s uc u e [2]. As can be seen in Figu e 1.1, he double helix is o med by wo chains in an
an ipa allel a angemen . In addi ion, i is possible o obse e he s uc u e o he majo
4
and mino g oo es. The majo g oo e is essen ial in p o eins-DNA in e ac ions (see Chap e
1.3.2).
Figu e 1.1. Le panel shows he double helix s uc u e o DNA. Cen e panel shows he
hyd ogen bonds be ween pai ed bases. Righ panel shows he majo and mino
g oo es, which a e binding si es o DNA binding p o eins du ing ansc ip ion (copying
RNA om DNA) and eplica ion. Image modi ied om "DNA s uc u e and sequencing:"
by OpenS ax College, Biology (CC BY 3.0).
In addi ion o his canonical helical s uc u e (B-DNA), he DNA, can p esen o he
con igu a ions so called A-DNA and Z-DNA.
The helix sense o he A-DNA is igh -handed like he B-DNA. The A-DNA is gene ally
wide and sho e wi h a na ow and deep majo g oo e, as opposed o he mino g oo e,
and like he B-DNA, he N-glycosidic linkage is an i. On he o he hand, he Z-DNA has an
en i ely di e en s uc u e, s a ing wi h a le -handed helix sense. This molecule is
gene ally he na owes and longes o he h ee men ioned con o ma ions, whe e he
majo g oo e is shallow, and he mino g oo e is na ow and deep (see Figu e 1.2).
5
Figu e 1.2. Side and op iew o A-, B-, and Z-DNA con o ma ions. By Mau oesgue o o -
Own wo k, CC BY-SA 4.0,
h ps://commons.wikimedia.o g/w/index.php?cu id=35919357
DNA can ha e al e na i e s uc u es, such as DNA Bubble, slipped loop, c uci o ms,
H-DNA, and G-quad uplex/i-mo i in double-s anded DNA. All o hese s uc u es a e
impo an , as he DNA con o ma ion de ines he ype o p o ein in e ac ions in which i is
in ol ed.
1.1.3. P o ein s uc u e
The e a e ou le els o s uc u al complexi y in p o eins.
P ima y s uc u e e e s o he sequence o amino acids ha o m a polypep idic
chain, de ined om he N- e minal o he C- e minal. This is he o de in which he
ibosomes o m he pep ide bonds. The amino acid sequence depends on he DNA
sequence codons ha o m he speci ic gen.
The seconda y s uc u e o p o eins is o med by local in e ac ions be ween amino
acid esidues ha a e nea ly loca ed in he polypep ide chain. Hyd ogen bonds be ween
he backbone NH and CO g oups s abilise hose local olding e en s. This leads o he wo
mos common ypes o seconda y s uc u e: alpha helices and be a shee s. Such pe iodic
s uc u es we e p edic ed by Linus Pauling and Robe Ca ey in 1951, yea s be o e hey
we e expe imen ally con i med [3].
6
Te ia y s uc u e is o med by he in e ac ion be ween di e en seconda y
s uc u e elemen s in he 3D space, and i is s abilized by hyd ogen bonds, an de Waals,
sal b idges, hyd ophobic packing, and disulphide bonding.
The qua e na y s uc u e o p o eins can be de ined as he in e ac ion be ween
di e en polypep ides ha ha e a e ia y s uc u e and in e ac wi h each o he .
In gene al, he sequence o amino acids o a p o ein de e mines i s 3D s uc u e.
While he majo i y o p o eins ha e e ia y s uc u e, he e a e many p o eins ha a e
pa ially o e en o ally uns uc u ed [4]. In many cases, diso de ed egions can become
s uc u ed when in e ac ing wi h ano he p o ein o biomolecule [5].
1.2. P o ein-p o ein complexes
1.2.1. Impo ance o p o ein in e ac ions
P o ein-p o ein in e ac ions media e mos cellula unc ions. In his ega d, he in e ac ome
is he complex and dynamic ne wo k o med by he en i e se o p o ein in e ac ions in a
li ing o ganism o a cell. Mapping he in e ac ome would p o ide insigh s in o how
biological sys ems wo k, and how hey could be modula ed. Recen s udies ha e made
p og ess owa ds he desc ip ion o his in ica e ne wo k o in e ac ions using p o eome-
wide echniques [6, 7], al hough i is es ima ed ha he eal numbe o p o ein-p o ein
in e ac ions a e much la ge han cu en ly known, be ween 130,000 and 600,000 [8, 9].
All da a om hese expe imen s ha e been collec ed in se e al IPP da abases such as: DIP
[10], MPIDB [11], HPRD [12], BIND [13], MINT [14] , Ma ixDB [15], STRINGS 10[16],
BioGRID [17], Reac ome[18], e c. In addi ion o hese da abases, he e is In e ac ome 3D
and MIn Ac p ojec [19], which y o colle se e al da abases cen alising he in o ma ion
in one place.
Howe e , al hough hese ini ia i es a e sys ema ically ex ending he numbe o
known in e ac ing p o ein pai s, hey ha e di icul y p o iding he s uc u al a chi ec u e o
hese in e ac ions, which is essen ial o un eil he unde lying molecula mechanisms ha
main ain heal h o con ibu e o disease. Such s uc u al knowledge may ul ima ely lead o
7
he a ional design o new d ugs, and he in e p e a ion and design o ele an mu a ions
o biomedical o bio echnological pu poses.
1.2.2. S uc u al cha ac e iza ion o p o ein-p o ein complexes
Biophysical me hods, mainly X- ay c ys allog aphy [20], NMR spec oscopy [21] and c yo-
elec on mic oscopy (c yo-EM) [22, 23], can be used o sol e p o ein–p o ein in e ac ions
a a omic esolu ion. These me hods ha e de e mined de 3D s uc u e o a la ge numbe o
complexes, so many s udies ha e been epo ed o analyze unique cha ac e is ics o
p o ein-p o ein in e ac ions.
Bu i has been di icul o ind gene al s uc u al and ene ge ics ea u es o all
p o ein-p o ein complexes. I has been ound ha in s able complexes, binding a e d i en
by hyd ophobic in e ac ions, as compa ed o mo e ansien complexes [24] ha a e mo e
hyd ophilic. S able complexes show mo e packed in e aces, wi h ewe hyd ogen bonds
be ween he subuni s han he less s able complexes.
Acco ding o se e al s udies, p o ein in e aces a e domina ed by a oma ic (Phe, His,
T y) and alipha ic (Me , Val, Ile, Leu) esidues, and wi h he excep ion o a ginine, hey a e
deple ed in cha ged esidues. In e es ingly, a ginine is one o he esidues ha appea mos
o en a he in e aces. The a ea alues o known p o ein-p o ein in e aces ollow a no mal
dis ibu ion ha peaks in he ange o 600-800 Å2 [25], sugges ing ha complex o ma ion
equi es a minimum numbe o con ac s, in addi ion o he displacemen o wa e
molecules [26].
8
Ano he a ea o esea ch ocused on he ene ge ic con ibu ions o he binding
a ini y. Hyd ophobic [27], elec os a ic [28, 29], and an de Wall [30] o ces play an
essen ial ole in cha ac e ising he binding a ini y. Depending on he complex ype, each
ene ge ic e m can ha e a di e en weigh in he binding a ini y.
1.3. P o ein-DNA complexes
1.3.1. Impo ance o p o ein-DNA in e ac ions
P o ein-DNA in e ac ions egula e many biological p ocesses such as p o ein syn hesis,
signal ansduc ion, DNA s o age, and DNA eplica ion and epai , among o he s. Lea ning
how p o ein and DNA in e ac is undamen al o ully elucida e many cen al biological
p ocesses and disease mechanisms and can also suppo he disco e y o no el he apeu ic
a ge s. The e is a la ge a ie y o p o ein-DNA binding mechanisms. On he one side, DNA-
binding p o eins can be e y speci ic o DNA sequence, such as he es ic ion
endonucleases. Bu on he o he side, he e a e DNA-binding p o eins, such as his one
p o eins and DNA polyme ases, which do no disc imina e DNA sequences when binding.
Figu e 1.3. In e ace size dis ibu ion. In e ace size is calcula ed sepa a ely o each
side o an in e ace. The dis ibu ion has a peak a 600–800 Å2. Abou 25% o he
in e aces ha e a (one-sided) size in he ange o 800 (±200) Å2 . Figu e ep oduced om
Yan, C., e al., Cha ac e iza ion o p o ein-p o ein in e aces. P o ein J, 2008. 27(1): p.
59-70.
15
p o ein-pep ide (T60-64) o p o ein-hepa in (T57) among o he s. Howe e , p o ein-DNA
docking ecei ed limi ed a en ion om he CAPRI communi y and de elope s o
compu a ional me hods. Compa ed o p o ein-p o ein docking, whe e he mos ecen
elease o he s anda d P o ein-P o ein Docking Benchma k 5.5 [87] has 257 en ies, and o
p o ein-RNA docking, whe e he e a e di e en epo ed benchma ks [100-103], o
p o ein-DNA docking he e is only one a ailable benchma k, which con ains 47 complexes
[104]. Using his benchma k, p o ein-DNA docking p o ocols epo mode a e success a es
in unbound condi ions. Fo ins ance, on a subse o 23 cases om his benchma k, HDock
success a e o op 10 models (i.e. a leas one nea -na i e s uc u e wi hin he op 10
models) is less han 10%, while success a e o op 100 is sligh ly o e 30% [95]. NPDock
epo s a maximum success a e (i.e. a leas one nea -na i e con o ma ion ound in he
en i e p edic ion se ) o 7/47 (15%) [94]. P o ein-DNA docking wi h HADDOCK epo ed an
excellen pe o mance [105] when using es ain s based on he eal in e ace. This
ep esen s a e y p omising app oach, bu in a ealis ic scena io, lack o knowledge on he
ac ual complex in e ace migh limi i s applica ion. A mo e ecen coa se- e sion o
HADDOCK p o ein-DNA docking shows simila accu acy wi h ~6- old speed inc ease o e
a omis ic calcula ions [106]. The need o new compu a ional ools o add ess unbound
p o ein-DNA docking is clea .
16
17
2. OBJECTIVES
18
The objec i es o his hesis can be g ouped in hese gene al aims:
i. Imp o e exis ing p o ein-p o ein docking unc ionali ies
ii. De elopmen o new p o ein–DNA docking p o ocols
ii. Explo a ion o new p ocedu es ha in eg a e ab ini io and empla e-based p o ein-
p o ein docking
iii. Implemen a ion o he de eloped ools and p o ocols as web se e s o sha e hem
wi h he communi y
i . E alua ion o he new de elopmen s in blind condi ions (CAPRI-CASP)
. Applica ion o sys ems o biological and bio echnological in e es
19
3. METHODS
20
The p o ocols desc ibed he e o he ins alla ion and execu ion o pyDock ha e been epo ed in
his publica ion: Rosell, M., L.A. Rod íguez-Lumb e as, and J. Fe nández-Recio, Modeling o P o ein
Complexes and Molecula Assemblies wi h pyDock, in P o ein S uc u e P edic ion, D. Kiha a, Edi o .
2020, Sp inge US: New Yo k, NY. p. 175-198.
21
3.1. pyDock
The pyDock me hod is a se o p o ocols o p o ein-p o ein docking and sco ing, p e iously
de eloped [79, 82], alida ed [107] and success ully applied o many cases o biological and
bio echnological in e es [108]. The pyDock so wa e needs he coo dina es o he wo
in e ac ing p o eins, usually as PDB iles. Hyd ogen a oms a e no needed in he PDB iles,
and i p esen , hey will be emo ed and ebuil again by pyDock. In addi ion, all HETATM
coo dina es will be emo ed in he docking calcula ions.
In his hesis, new unc ionali ies ha e been implemen ed o be able o e icien ly
use AMBER coo dina e and opology iles. The newes e sion o he pyDock p og am
op imized du ing his hesis can use AMBER coo dina e iles (wi h ex ensions such as inpc d,
. es , . s7, .c d) and opology iles (wi h ex ensions such as .p m op, .pa m7, . op) c ea ed
by he PARM, LeAP, SANDER, o GIBBS p og ams om AMBER [109]. In his case, he
HETATM coo dina es om he co ac o s and o he compounds will be included in pyDock
calcula ions.
3.1.1. pyDock ins alla ion
The pyDock 3.0 package is a ailable a h ps://li e.bsc.es/pid/pydock/ge _pydock.h ml o
ge pyDock you need o apply o a license by illing in you da a; o academic use, you will
ecei e a link o he pyDock dis ibu ion ile by e-mail; o comme cial use, you will be
con ac ed by he au ho s). Uncomp ess and un a he pyDock dis ibu ion ile o ex ac he
pyDock3 di ec o y.
Nex , we need o change pe missions o he pyDock/da a di ec o y:
> cd pyDock3
> chmod go+ x da a
> chmod u+x pyDock3
The pyDock3 di ec o y can be mo ed o any loca ion o you choice. Fo ins ance,
le us say ha i is mo ed o /us /local/so wa e/ di ec o y; hen, pyDock can be called by:
22
> /us /local/so wa e/pyDock3/pyDock3
Mo eo e , he PYDOCK a iable can be de ined in you .bash c ile, as ollows:
expo PYDOCK=/us /local/so wa e/pyDock3/
so ha he execu able o pyDock can be called in a mo e con enien way:
> $PYDOCK/pyDock3
The pyDock bina y has been compiled o Linux 32-bi o inc ease he compa ibili y wi h
olde S.O.
3.1.2. pyDock ex e nal p og ams
3.1.2.1. SCWRL
In he case o inpu PDB iles wi h incomple e side-chains, pyDock uses SCWRL 3.0
(h p://dunb ack. ccc.edu/) o ebuild hem. We should no e ha his e sion is ou da ed
and canno be di ec ly downloaded om he abo e web, so you need o ob ain he
ins alla ion ile scw l3_lin. a .gz om hei au ho s. Then, unzipping his ile will ex ac i s
con en s o a new di ec o y called scw l3_lin. Inside his di ec o y, un:
> ./se up
This will c ea e a SCWRL 3.0 bina y (scw l3) in ha di ec o y.
3.1.2.2. FTDOCK 2.0
The p og am needs some ex e nal p og ams o gene a e a se o igid-body docking poses.
In his ega d, pyDock is eady o p ocess he ou pu o FTDock 2.0, and we will show he e
how o ins all i . The i s s ep is o ins all he FFTW lib a ies, ha can be downloaded he e:
h ps://www. w.o g/ w-2.1.5. a .gz
23
> sudo ap -ge ins all mpi-de aul -de
> wge h ps://www. w.o g/ w-2.1.5. a .gz
> a -x w-2.1.5. a .gz
Uncomp essing his ile will ex ac i s con en s o a new di ec o y called w-2.1.5.
Wi hin his new di ec o y, compile he lib a ies by:
> ./con igu e --enable- ype-p e ix --enable-mpi --
p e ix='/pa h/ o/lib/ w'
> make ins all
Now, he dock-mpi-mas e .zip ile can be download. This is he cus om FTDOCK
e sion, wi h g id-op imiza ion and eady o un in pa allel [82]. I can be downloaded om
he Gi Hub eposi o y (h ps://gi hub.com/b ianjimenez/ dock-mpi). This zipped ile
should be unpacked, and wi hin he new dock-mpi-mas e di ec o y, he Make ile ile
should be edi ed o se he FFTW_DIR a iable o he ull pa h (--p e ix) used as inpu in he
con igu e command abo e. Now, wi hin he dock-mpi-mas e di ec o y, ype:
> ./make
This will c ea e he p og am bina ies, such as dock. Mo e in o ma ion can be ound
in he README.md ile
3.1.2.3. ZDOCK
Ano he ex e nal ool ha can be used o he gene a ion o igid-body docking poses is
ZDOCK (h ps://zdock.umassmed.edu/). The pyDock pipeline is eady o p ocess he ou pu
o ZDOCK 2.1.
24
3.1.3. Au oma ic use o ex e nal p og ams
Fo au oma ic use o FTDock, ZDOCK, and SCWRL p og ams wi hin pyDock, a e ins alling
hem locally, i is necessa y o indica e he ull pa h o he FTDock and ZDOCK di ec o ies,
and ha o he SCWRL bina y, by modi ying he co esponding lines in he
$PYDOCK/pyDock3/e c/pydock.con ile, as ollows:
(...)
ZDOCK=/<you -ins alla ion-di ec o y>/zdock2.1_linux_64bi /
FTDOCK=/<you -ins alla ion-di ec o y>/ dock-mpi /
SCWRL=/<you -ins alla ion-di ec o y>/scw l3_lin/scw l3
(...)
3.1.4. Running pyDock
pyDock has a highly modula a chi ec u e, wi h a se ies o modules pe o ming he di e en
unc ionali ies o he p og am (Figu e 3.1). The gene al syn ax o unning pyDock is:
> $PYDOCK/pyDock3 DOCKNAME modulename
Thus, he execu able pyDock3 usually needs wo a gumen s: (1) DOCKNAME, which
is he name o he pyDock p ojec and he base o all he iles ha will be c ea ed du ing
he docking pipeline, and (2) modulename, which will call o he speci ic module. The
de ails o he di e en pyDock modules a e desc ibed in he unning ins uc ions below.
31
pep ide chains. And inally, he ATOM eco ds desc ibe he coo dina es o he a oms ha
make up he p o ein. Fo example, he i s ATOM line ep esen s he alpha-N a om o he
i s esidue o pep ide chain A, which is a p oline esidue; he i s h ee loa ing poin
numbe s a e i s x, y and z coo dina es and a e in uni s o Angs oms. The ollowing h ee
columns a e he occupancy, empe a u e ac o , and a om name. The HETATM eco ds
desc ibe he coo dina es o he he e oa oms, i.e., he a oms ha a e no pa o he p o ein
o nucleic acid molecule. They can be co ac o s o me al a oms, such as Heme, Fe, Cu, e c.
(de ails o he o ma in he Figu e 3.2).
3.2.2. PDBx/mmCIF
The cu en e e ence o ma in s uc u al biology is PDBx/mmCIF [112]. I has p ac ically
no limi a ions, om small molecules o la ge mac omolecula complexes. Ac ually,
PDBx/mmCIF is de i ed om he CIF o ma , which was i s used o small molecules. Abou
20 yea s ago, i was upda ed o mmCIF, aiming o be he successo o he PDB o ma . Da a
s o age is done in a simila way o XML o JSON, wi h a dic iona y o schema needed o
know he in e nal s uc u e and o be able o access he in o ma ion.
32
The ac ha PDB o ma is mo e use - iendly han PDBx/mmCIF, and ha he PDB
ile can be easily e ie ed om mmCIF (see Figu e 3.3), makes he PDB o ma o be s ill
widely used, due o i s simplici y.
In he example shown in Figu e 3.3 he ield labels appea i s , and hen
immedia ely below he da a (in PDB-like o ma ). Fo example, _a om_si e.id co esponds
o he a om numbe , and he es o he ields simila ly desc ibe he PDB columns.
3.2.3. Mol2
A T ipos Mol2 (.mol2) ile is a comple e and po able ep esen a ion o a SYBYL molecule
[113]. The mos impo an cha ac e is ics o his ile is ha i explici ly con ains a om ype
and bond in o ma ion. In many cases, i is essen ial o use o con e PDBs o his ile ype
o be able o pe o m he pa ame e iza ion using an echambe . This was he case o he
co ac o s p ocessed in Chap e 7 (see Figu e 3.4).
Figu e 3.3. PDBx/mmCIF unca ed example o X- ay c ys allog aphic s udies o seal
myoglobin (PDB 1MBS)
33
3.3. P og am languages (R, py hon, e c)
In his sec ion, we will b ie ly men ion some o he mos ele an p og amming languages
used in he hesis, wi h ocus on some o hei s ong poin s: py hon 2.7 [114], py hon 3.x.x
[115], pe l [116], [117] and he Linux console bash [118].
Py hon
Guido an Rossum de eloped Py hon in 1991. I is a high-le el, in e p e ed, c oss-pla o m
and objec -o ien ed p og amming language. I is also he mos popula language (as o
2022). I con ains a la ge numbe o ools o he bioin o ma ics communi y, such as
Biopy hon [119], P oDy [120], and pyP oCT [121]. O pa icula no e is i s ex ensi e use in
Machine Lea ning (ML), wi h Sciki -lea n [122], Theano [123] and Tenso Flow [124]. I is also
he main language used in he de elopmen o pyDock 4.0
Figu e 3.4. Mol2 example Benzene.
34
Pe l
La y Wall ini ially de eloped Pe l in 1987. Se e al yea s la e , Andy Doughe y and Tom
Ch is ian, among o he s, joined he p ojec (see h ps://pe ldoc.pe l.o g/pe lhis ). Pe l is
based on a block s yle like AWK and was widely adop ed by genomic bioin o ma icians o
i s ex -p ocessing p owess. Bu in my opinion, i is no an objec -o ien ed language and In
he S uc u al Biology ield, i is no widely used.
R
Ross Ihaka and Robe Gen leman de eloped i in he 1990's. R was bo n as a ool s ic ly
o s a is ical analysis. Bu o e he yea s i has been adap ed and he e a e also mul iple
packages o apply ML. To men ion a ew: da a. able, dply , ggplo 2, ca e , e1071, xgboos ,
andomFo es , e c...
Bash
I is he Swiss a my kni e o con olling Linux-based sys ems and al hough he e a e
excep ions, i is he only way o gi e commands o he la ge compu ing clus e s used du ing
he de elopmen o he hesis. Bash is an in e ac i e command in e p e e and uns in a
e minal, whe e commands a e yped. You can also gene a e a lis o commands in a ile
called a sc ip and execu e hem.
3.4. Molecula Visualiza ion So wa e
The e a e many and a ied p og ams o isualise molecules, bu he mos commonly used
by he bioin o ma ics communi y a e UCSF Chime a [125], UCSF Chime aX [126],
Jmol/JSMol [127], PyMoL , VMD [128], ICM-B owse [129]. Also wo h men ioning NGL
[130] is embedded in he pyDockDNA se e .
In he ollowing sec ions, I will discuss mo e de ails o he h ee molecula isualise s
mainly used du ing he hesis de elopmen .
3.4.1. ICM
ICM-B owse is a ee e sion o he ICM p og am (www.molso .com) wi h many ea u es
o molecula isualiza ion and s uc u al analysis. I can display su aces o ligand binding
pocke s, op imise hyd ogens o a PDB, supe impose (s uc u al align) p o ein s uc u es,
35
measu e dis ances and angles, gene a e and display su aces, among o he s. All he
unc ionali ies can be accessed h ough he command line, bu no all o hem a e easily
ound in he g aphical use in e ace, e.g. h ough he command line one can open a
collec ion o SDF iles o molecules and pe o m di e en measu emen s such as RMSD
calcula ion, bu in he g aphical use in e ace, his op ion is no isible.
3.4.2. UCSF Chime a
The p og am Chime a was used as an al e na i e o ICM, especially o he use o unc ions
ha we e only a ailable in i s comme cial e sion. The mos in e es ing ea u e o Chime a
is he possibili y o using py hon sc ip ing o do speci ic in ensi e asks in addi ion o he
command line (h ps://www. b i.ucs .edu/ ac/chime a/wiki/Sc ip s). You can do isual
ep esen a ions in a simila way as in ICM, as well as Molecula Dynamics (bu only o
eaching pu poses).
3.4.3. Pymol
PyMOL is an open-sou ce bu p op ie a y p og am w i en in he Py hon p og amming
language. I allows he c ea ion o plug-ins ha ex end i s unc ionali y, among which is
he Au odock plugin, which allows he se up o a docking g id (Au oDock Vina [131]) and
iew he docking esul s.
3.5. Benchma ks and e alua ion se s
3.5.1. P o ein-p o ein docking benchma k 4.0
The p o ein-p o ein docking Benchma k 4.0 [132] (BM4) was used as he a ge lib a y o
alida ing he p o ein docking uncionali ies de eloped in his hesis. This p o ein
benchma k p o ides 176 complexes sol ed by x- ay c ys allog aphy (119 dime s and 57
mul ime s) wi h a leas 3.25 Å esolu ion, whe e he bound and unbound s a es a e known.
This benchma k is a non- edundan se o p o ein complexes ha include, among o he s,
enzyme-inhibi o , enzyme-subs a e, and an igen-an ibody complexes. The a ge s a e
classi ied as igid body, medium and challenging in e ms o he expec ed di icul y o ab
36
ini io p o ein–p o ein docking. Fo u he de ails, see
h ps://zlab.umassmed.edu/benchma k/ web si e.
3.5.2. DOCKGROUND
In his s udy, we used he s uc u al empla es included in he DOCKGROUND esou ce 1.0
[133]. The s uc u es o p o ein-p o ein complexes included in his lib a y we e sol ed by x-
ay c ys allog aphy wi h a esolu ion be e han 3.5 Å. Only complexes wi h a mean
accessible su ace a ea bu ied by each chain g ea e han 250 Å2 and con aining a leas 10
in e ace esidues a e included. S uc u al di e si y was ensu ed wi h he MM-align
p og am by using a TM-sco e cu -o o 0.9, which esul ed in a da ase o 7,107 di e se
p o ein–p o ein in e aces. See h p://dockg ound.compbio.ku.edu o a ull da abase
desc ip ion.
3.5.3. P o ein-DNA
In o de o es he new pyDockDNA docking p o ocol de eloped in his hesis, we used a
p e iously epo ed p o ein-DNA docking benchma k ( e sion 1.2) [104]. The benchma k
con ained bound and unbound x- ay c ys allog aphy and NMR s uc u es o 47 p o ein-
DNA complexes in which DNA is in B-DNA con o ma ion. These we e classi ied as "easy",
"in e media e" o "di icul " cases, based on he in e ace RMSD alues be ween he bound
and unbound componen s o he complex (Table 3.2).
37
Table 3.2. P o ein-DNA docking benchma k ( e sion 1.2)
HTH
Zinc-coo d
O he α-helix
β-shee
Β-ha pin
Enzyme
2C5R
1FOK
3CRO
1H9T
1TRO
1RPE
1MNN
1F4K
1K79
1W0T
1Z9C
1DDN
2IRF
1JT0
1ZS4
1O3T
1BY4
1R40
1ZME
1KSY
2FIO
1JJ4
1QRV
1B3T
1HJC
1QNE
1EA4
1AZP
1CMA
1BDT
1PT3
1EMH
1DIZ
1VRR
1KC6
1Z63
1VAS
4KTQ
1G9Z
1A73/1A74
3BAM
1RVA
1DFM
7MHT
2FL3
1EYU
2OAA
Classi ica ion o cases as p e iously desc ibed [134]. Unde lined cases ha e only 1 DNA molecule. G een a e he easy cases
wi h in e ace RMSD
b-u
(be ween bound and unbound o he complex) anging om 0.0 Å o 2.0 Å. Blue a e he medium
cases wi h in e ace RMSDb-u be ween 2.0 Å and 5.0 Å. Red ones a e he ha d cases wi h in e ace RMSDb-u abo e 5.0 Å
An addi ional se o case s udies was compiled ollowing he c i e ia selec ion used
in he abo e-desc ibed p o ein-DNA docking benchma k. This es se is composed o en
p o ein-DNA complexes, whe e bo h bound and unbound s uc u es a e a ailable o each
e e ence complex, and he sequences a e di e en om hose in he i s p o ein-DNA
docking benchma k (Table 3.3). P o ein-DNA complex and unbound s uc u es we e
compiled om he P o ein-DNA In e ace Da abase (PDIdb) [135] and he P o ein Da a Bank
(PDB) [31]. Only complexes ha mee he ollowing condi ions we e conside ed: i) DNA
sequence leng h la ge han eigh base pai s, and ii) p o eins wi hou mu a ions in he co e
o he complex in e ace. To ind he p o ein unbound s uc u es o he selec ed p o ein-
DNA complexes, all he PDB en ies con aining only p o ein s uc u es we e e ie ed,
including s uc u es sol ed by NMR. C ys allog aphic s uc u es wi h a esolu ion wo se
han 3.0 Å we e no conside ed. To a oid edundancy, en ies wi h sequence simila i y ≥
90% we e disca ded. PDBeFOLD [136] was used o ind co espondences be ween bound
and unbound p o ein s uc u es. This ool pe o ms s uc u al alignmen s be ween wo
38
(pai wise alignmen ) o mo e (mul i-alignmen ) molecules using hei 3-dimensional
s uc u es. The alignmen is based on he Seconda y S uc u e Ma ching algo i hm [136].
Alignmen s wi h a Q-sco e highe han 8.0, high P-sco e and sequence simila i y a ound 90-
100% we e accep ed as he co esponding unbound. Then, he bound and unbound
s uc u es o each case, we e pos -p ocessed acco ding o he p o ocol ollowed in a
p e iously de eloped p o ein-DNA docking benchma k, o ins ance by checking
consis ency be ween unbound and bound coo dina es in chain IDs, esidue numbe s and
a om names [104]. The unbound DNA models we e gene a ed by using he so wa e 3DNA
[137, 138], in canonical B-DNA con o ma ion ( ibe model 4).
This addi ional es se (is eely a ailable a he "Help" sec ion o he se e
(h ps://model3dbio.csic.es/pydockdna/in o/ aq_and_help#ex ended_bechma k).
Table 3.3. Lis o he case ex e nal es se .
PDB
complex
P o ein
PDB unbound
p o ein
RMSD unbound-
bound p o ein DNA
RMSD unbound-
bound DNA
5JLT
phage T4 Mo A DNA-
binding domain
1KAF
0.83a
22bp dsDNA
1.89
2X6V
TBX5
2X6V
0.55
11bp DNA
2.03
3POV
SOX
3FHD
1.46
19bp DNA
2.26
4UUV
ETV4 DNA-binding
ETS domain
5ILU
1.24
10bp DNA
2.81
2NTC
s 40 la ge T an igen
2FUF
1.13a
21-n PEN elemen o he SV40
DNA o igin
2.96
2ITL
s 40 la ge T an igen
4NBP
5.37a
24-n PEN elemen o he SV40
DNA o igin
3.84
3MFK
P o ein C-E s1
1GVJ
5.61a
s omelysin-1 p omo e DNA
4.34
2PI0
IRF-3
3QU6
0.76a
PRDIII-I egion o human
in e e on-B p omo e s and 1
4.46
1O3R
ca aboli e gene
ac i a o p o ein
4R8H
0.65
11bp DNA
4.77
3MLO
Eb 1
3LYR
0.71a
22bp DNA
5.11
a In cases wi h mo e han one p o ein-DNA in e ace in he x- ay s uc u e, he a e age alue is p o ided.
39
3.5.4. CAPRI
The C i ical Assessmen o P edic ed In e ac ions (CAPRI) communi y-wide expe imen
s a ed a ound 2001, when he communi y o de elope s o p o ein-p o ein docking
me hods aimed o e alua e he success a e o such algo i hms [98, 139]. CAPRI has been
c ucial in pushing i s communi y membe s o imp o e and add new unc ions o hei
docking p o ocols [140, 141]. Mos o he a ge s p oposed by he CAPRI expe imen
ocused on p o ein-p o ein docking p ocedu es. S ill, ecen ly he e ha e been new
challenges, such as p o ein-pep ide and p o ein-oligosaccha ide docking (see Chap e 6.5.1)
[142].
The expe imen consis s in an open compe i ion in which he p edic i e success
a es o he di e en docking me hods a e compa ed in double-blind condi ions. The
o ganize s choose he a ge s, consis ing o expe imen ally de e mined complex s uc u es
ha a e no ye publicly a ailable. Thus, he a ge s a e blind o he pa icipan s, and he
pa icipan names a e blind o he o ganize s when e alua ing hei p edic ions.
Fo each a ge , he e a e usually wo modes o pa icipa ion in he expe imen :
p edic o s and sco e s. In p edic o s, he g oups a e asked o submi en models om he
sequences o he a ge s uc u es. Gene ally, he 3D s uc u es o easonable empla es o
he in e ac ing molecules a e a ailable, which can be used as a s a ing poin o p o ein-
p o ein docking. In he sco e pa icipa ion, he g oups a e in i ed o e alua e a common
Figu e 3.5. CAPRI-CASP e alua ion p ocess.
40
se o docking models submi ed by he p edic o s g oups. Only he op 5 o he op 10
models by each pa icipan a e conside ed o he assessmen .
A he end o each ound, which may consis o se e al a ge s, he en models
submi ed by each pa icipan (ei he as p edic o s o as sco e s o bo h) a e e alua ed
based on he ligand RMSD, he ac ion o na i e con ac s and he in e ace RMSD wi h
espec o he ac ual complex s uc u e (see Figu e 3.5).
3.5.5. CASP-CAPRI
The CASP and CAPRI communi ies es ablished close ies du ing he CASP 2014 edi ion,
whe e he sec ion o "mul ime ic assemblies" was o ganized join ly wi h he CAPRI
communi y [143]. Since hen, a o al o i e join CASP-CAPRI ounds we e held [107, 144,
145], including his yea (2022) edi ion.
These join CASP-CAPRI ounds ha e encou aged many de elope s o in eg a e
hei ab-ini io docking me hods wi h s uc u e p edic ion me hods. Many o hese
p o ocols a e pe iodically collec ed in books, such as he se en h edi ion o P o ein
S uc u e P edic ion 2020 [146]. Pa icipa ion and e alua ion in hese ounds a e done
independen ly by CASP and CAPRI o ganize s, he la e in a simila way as explained in
he p e ious sec ion o his chap e , see Figu e 3.5.
3.5.6. Physiological/non-physiological homodime s
The g oups o R.L. Dunb ack and E.D. Le y de eloped a benchma k se o physiological/non-
physiological dime s ( e sion 3) in he con ex o he Ac i i y II o he 3DBioIn o ELIXIR
communi y, in which I ha e pa icipa ed du ing his hesis.
The Benchma k con ains a numbe o p o ein homo-dime ic x- ay s uc u es ha
a e classi ied as ei he "physiological" o "non-physiological". Physiological homodime s a e
de ined as hose ha a e likely o occu in he cell. Non-physiological a e homodime ic
in e ac ions ha a e seen in he c ys al s uc u e bu a e unlikely o occu in he cell,
because he known physiological s a e is monome ic o because he ue homodime
in ol es a di e en in e ace.
47
In he ini ial e sions, his ambe module could only be used o p o ein-p o ein
docking wi h modi ied amino acids. This e sion only used he a om cha ges o he AMBER
opology ile. The an de Waals (VDW) pa ame e s and he a omic sol a ion pa ame e s
(ASPs) we e in e nally de ined by pyDock da a iles. The o iginal pyDock e sion used
pa m94 [151], which limi ed he ype o a oms ha can be mapped. Also, he a omic
sol a ion pa ame e s a e unique o pyDock ( hey a e no included in he gene al molecula
mechanics o ce ields) [152].
To make his module wo k on a wide a ie y o molecules, he pyDock 4.0 unc ion
ha eads he opology iles was modi ied o ex ac bo h he cha ge and VDW alues and
w i e hem in he .ambe ile se up pa ame e s ( his ile will be used o calcula e he ene gy
48
sco ing unc ion by dockse module). To use he pyDock sol a ion pa ame e s, we c ea ed
a dic iona y o equi alences be ween he newes ambe a om ypes and he old (pa m94)
a om ypes (see Figu e 4.2)
Finally, he pyDock unc ion ha c ea es he ou pu PDB iles was modi ied. The
esul ing PDB can be used di ec ly wi h FTDOCK, wi hou equi ing u u e modi ica ions. A
use case o his module upg ade can be ound in Chap e 7.3.1.
4.4.2. pyClus e : Clus e ing in pyDock 4.0
The RMSD ma ix equi ed o apply he BSAS algo i hm [153] is compu ed wi h an ad-hoc
ICM sc ip in he in CAPRI and CASP-CAPRI ounds. Bu he use o pa allelisa ion in ICM is
Figu e 4.2. Pa ial snapsho o he equi alence dic iona y be ween pa m94 and he
newes o ce ield
pa ame e s ind in he las ed AMBER e sion.
(h ps://gi hub.com/pyDock/pa allel/blob/mas e /ambe _old_ o_new.map)
49
no i ial. The ini ial op ion was o implemen he clus e ing p o ocol ha we used o es
pyDockDNA, pyP oCT, bu his p og am (as s andalone) is a Py hon 2.7 module, which is
incompa ible wi h he new pyDock e sion. The e o e, we decided o implemen he BSAS
algo i hm as a new module in pyDock 4.0.
The new pyClus e module is able o gene a e he RMSD ma ix e y quickly hanks
o he use o s anda d Py hon 3 mul ip ocessing, jus like he dockse module. And i can be
used bo h o he di ec ou pu o pyDock 4.0 and o collec ions o independen models, as
i is he case o he CAPRI sco e expe imen . See Figu e 4.3A and Figu e 4.3B o examples
o INI con igu a ion iles ha can be used.
A
[clus e ing]
modelslis =
clus e _1AVX_clus e ing.lis
sco ing_ unc ion = PyDock
[ ecep o ]
mol_clus e = A
[ligand]
mol_clus e = B
B
[ ecep o ]
pdb = 1AVX_ _u.pdb
mol = A
newmol = A
[ligand]
pdb = 1AVX_l_u.pdb
mol = B
newmol = B
[ e e ence]
pdb = 1AVX_b.pdb
ecmol = A
ligmol = B
new ecmol = A
newligmol = B
[clus e ing]
RMSD_cu o = 20
Nmodels = 100
Figu e 4.3. PyClus ini ile. (A) The ini ile ha e as inpu s a lis o PDBs. The ligand and he
ecep o chains mus be speci ied. I he RMSD_cu o and Nmodels a e no speci ied,
hey a e se o 4Å and 100 models, espec i ely. (B) Co esponds o an ini ile, using he
di ec ou pu o pyDock 4.0. He e, you can selec he RMSD_cu o and Nmodels
50
51
5. pyDockDNA: A NEW WEB SERVER
FOR ENERGY-BASED PROTEIN-DNA
DOCKING AND SCORING
52
The webse e desc ibed he e ha e been epo ed in his publica ion: Rod íguez-Lumb e as, L.A.,
e al., pyDockDNA: A new web se e o ene gy-based p o ein-DNA docking and sco ing. F on ie s
in Molecula Biosciences, 2022. 9.
53
5.1. De elopmen o pyDockDNA: a new p o ein-DNA docking p ocedu e.
5.1.1. Sampling
In his i s s ep, he inpu iles wi h he coo dina es in PDB o ma o he s uc u es (o
models) o a p o ein and a DNA molecule (which can be B-DNA o any o he con o ma ion)
a e checked o po en ial o ma e o s. Missing side-chains in he p o ein a e ebuil wi h
SCWRL 3.0 [154], and he elec os a ics Ambe 94 o ce ield [151] is loaded, assigning he
cha ges o he a oms. Then, igid-body docking poses be ween he p o ein and he DNA,
ep esen ed as 3D g ids, a e gene a ed wi h a as e and pa allelized e sion o he o iginal
FTDock ( 2.0) so wa e [66] in which he numbe o cells in he g id is op imized o
maximum compu ing e iciency [82]. The molecule (p o ein o DNA) wi h he longes
maximal dis ance be ween any pai o a oms is conside ed he ecep o , ha is, he ixed
molecule, and he o he one is he ligand o mobile molecule. By de aul , he p og am uses
0.7 Å g id cell size, 1.3 Å su ace hickness, 12º o a ion sampling, and keeps he bes 3
poses o each o a ion. Fo each a ge , a o al o 10,000 docking poses a e gene a ed.
5.1.2. Sco ing
Then, he p o ein-DNA docking poses a e anked using a sco ing unc ion composed o
elec os a ics, desol a ion and an de Waals ene gy. This new pyDockDNA sco ing unc ion
is adap ed om he p e iously pyDock sco ing unc ion o p o ein-p o ein docking [82,
155], which now includes a om ypes o nucleo ides om Ambe 94 o ce ield [151] in
o de o calcula e o he modelled p o ein-DNA complexes. The nucleo ide AMBER a om
ypes ha e been mapped o he p e iously de ined a om ypes in pyDock wi hin a new
pa ame e se (nuc.da ).
5.1.3. Clus e ing o p o ein-DNA docking models in benchma king
When es ing his so wa e (see Chap e 5.3) we ha e un se e al docking execu ions in
pa allel, using di e en ini ial andom o a ions o he inpu s uc u es, and he bes -
sco ing 100 esul ing models o each indi idual un we e me ged in o a single pool. To
a oid edundancy in he inal se , all docking o ien a ions we e clus e ed by pyP oCT
54
analysis so wa e [121], which implemen s he GROMOS clus e ing algo i hm [156]. The
dis ance ma ix is buil wi h pyRMSD wi h he op ion "QCP OMP CALCULATOR" o compu e
he ligand oo -mean-squa e de ia ion (L-RMSD) alues o all pai s o docking o ien a ions
a e hei ecep o s we e supe imposed (h ps://gi hub.com/ ic o -gil-
sepul eda/pyRMSD/). A cu -o alue o 4.0 Å was used o L-RMSD o de ine he clus e s.
Fo each de ined clus e o models, he o ien a ion wi h he lowes docking sco e is selec ed
as he clus e ep esen a i e.
5.2. Implemen a ion o pyDockDNA as a web se e
The pyDockDNA p og am has been buil as a module o he new pyDock 4.0 e sion, and i
includes he same hi d-pa y p og ams, modules and ools o he p e ious pyDock e sions,
as well as new unc ionali ies o handle nucleic acid s uc u es in a p ope way (see Chap e
4). The p og am has been implemen ed as a web se e , in a i ual machine hos ed in one
o he da a p ocessing cen es (CPD) o he Spanish Na ional Resea ch Council (CSIC) and is
accessible h ough he ollowing link: h ps://model3dbio.csic.es/pydockdna. The ope a ing
sys em ins alled is Debian10. Fo he co ec managemen o he a ailable esou ces, Slu m
[150] was ins alled, which is a ask managemen sys em o clus e s. Thanks o his sys em,
we can con ol he jobs ecei ed by he se e and keep hem in queue when esou ces a e
limi ed. Fo secu i y easons, he e e se p oxy Nginx was ins alled in conjunc ion wi h he
uWSGI se e , as i has a lowe a e o se ious ulne abili ies compa ed o Apache
(h ps://www.c ede ails.com/).
As a as he web applica ion is conce ned, i is de ined as a backend and a on end.
The backend is essen ially a daemon, i.e. a esiden p og am unning in he backg ound.
This daemon is an adap a ion o he e sion used by pyDockWEB [82] bu upda ed o un
on he newes e sion o py hon so ha i can hos pyDockDNA [157] web applica ion.
Essen ially, i is in cha ge o submi ing jobs o he Slu m queue, moni o ing hei p og ess
and handling execu ion e o s. The job execu es se e al pyDock 4.0 modules in a conce ed
way, acco ding o he op ions selec ed by he use in he on end. This in o ma ion is
epo ed o a da abase ha ac s as a b idge be ween he backend and he on end.
55
The on end is de eloped using web2py, a amewo k o designing websi es using
he py hon language. On he main page, he use can choose he name o he job and an
email add ess o be no i ied a he end o he job. One can use he RCSB code o selec he
s uc u es (ligand and ecep o ) o upload hem om his compu e . In he ollowing s eps
he use can selec he chains o be docked, he ene ge ic sco ing unc ion, and e en include
ex e nal in o ma ion ( om a ailable expe imen al da a o using p edic i e me hods such as
he DBSI se e [158], o ins ance) as esidue-nucleo ide dis ance es ain s o esco e
docking models as p e iously desc ibed o pyDockRST [159]. The ou pu will be a se o
docking models ep esen ed in di e en o ma s: i) he 3D s uc u e o he bes -sco ing 10
docking models in e ms o sco ing can be isualized in he ou pu sc een, ii) he PDB iles
o he bes -sco ing 100 models can be di ec ly downloaded, and iii) he
o a ion/ ansla ion ec o s a e p o ided o gene a e up o a o al o 10,000 docking poses.
A summa y o he docking esul s can be isualized as a plo wi h he dis ibu ion o he
di e en ene gy alues ob ained o all docking poses (Figu e 5.1).
Figu e 5.1. pyDockDNA se e esul s. The op 10 p edic ions based on he use -selec ed
sco ing unc ion a e displayed as a able on he op le and as a 3D ep esen a ion on
he uppe igh . In he lowe igh a ea, a g aph o he ene gy dis ibu ion o he 10,000
models a e ep esen ed.
56
5.3. Pe o mance o pyDockDNA e alua ed on he p o ein-DNA docking
benchma k.
The pyDockDNA web se e has been es ed on he 47 cases o a p e iously epo ed
p o ein-DNA docking benchma k (see Me hods). I is known ha using di e en andomly
o a ed inpu s uc u es can sligh ly a ec docking p edic ions o FFT-based docking
p o ocols as in FTDOCK, because his can modi y he mapping o he a om posi ions on he
3D g ids [70, 160]. To check o con e gence, we applied pyDockDNA o 10 di e en andom
o a ions o he ini ial inpu s uc u es o each benchma k case and compu ed he
p edic i e success a es o he esul s ob ained om each andomly o a ed inpu
s uc u es. The esul s indica e e en mo e di e ences in he p edic i e alues han
p e iously epo ed o p o ein-p o ein docking (Appendix 1: Table 10.1.1). Fo ins ance,
he success a es o he op 10 models anged om 12.8% o 21.3%. The e o e, o a mo e
obus e alua ion, we me ged he esul s o all 10 docking execu ions and clus e ed he
ob ained docking models o emo e simila o ien a ions. Figu e 5.2 shows he p edic i e
success a es o he clus e ep esen a i es esul ing om me ging hese 10 docking uns
(see mo e de ails in Appendix 1: Table 10.1.2). The p edic i e success o he de aul pyDock
sco ing unc ion (including pa ame e s o nucleo ide a oms, see Chap e 5.1.2) a e be e
han hose ob ained o he indi idual docking uns, which means ha inc easing sampling
Figu e 5.2. P edic i e pe o mance o he op N=1, 5, 10, 100 models o pyDockDNA and
di e en combina ions o sco ing e ms on he p o ein-DNA docking benchma k.
63
6. NEW APPROACHES FOR
INTEGRATING AB INITIO AND
TEMPLATE-BASED DOCKING
64
The p o ocols desc ibed and he esul s ha e been epo ed in he ollowing publica ions:
1. Lensink, M.F., e al., Modeling p o ein-p o ein, p o ein-pep ide, and p o ein-
oligosaccha ide complexes: CAPRI 7 h edi ion. P o eins, 2020. 88(8): p. 916-938.
2. Rosell, M., e al., In eg a i e modeling o p o ein-p o ein in e ac ions wi h pyDock o he
new docking challenges. P o eins, 2020. 88(8): p. 999-1008.
3. Lensink, M.F., e al., P edic ion o p o ein assemblies, he nex on ie : The CASP14-CAPRI
expe imen . P o eins, 2021. 89(12): p. 1800-1823.
4. Lensink, M.F., e al., Blind p edic ion o homo- and he e o-p o ein complexes: The CASP13-
CAPRI expe imen . P o eins, 2019. 87(12): p. 1200-1221
65
66
6.1. In oduc ion
The compu a ional p edic ion o 3D p o ein–p o ein complexes ypically includes wo
dis inc s a egies: empla e-based modeling and ab ini io docking. Templa e-based
modeling me hods build a model o a p o ein-p o ein complex based on empla e complex
s uc u es, ha is, o med be ween p o eins ha a e homologous o hose in he modelled
complex, assuming ha binding mode will be conse ed in his si ua ion. Ab ini io docking
app oaches explo e he po en ial binding modes be ween he in e ac ing p o eins h ough
s e ic and physicochemical complemen a i y, in cases whe e no empla e complex s uc u e
is a ailable.
6.1.1. Templa e-based modeling
A wide ange o empla e-based me hods ha e been de eloped, exploi ing he empla e
in o ma ion di e en ly. Homology modeling uses sequence iden i y (S.I.) o empla e
iden i ica ion and model building [163, 164]; h eading me hods' h ead' sequences on o
s uc u al empla es [165, 166]; empla e-based docking usually e e s o global
supe imposi ion o he s uc u es o unbound monome s on o he co esponding subuni s
in a empla e complex s uc u e [52]; and s uc u e in e ace alignmen exploi s local
simila i y and gene a es models by supe imposing he in e ac ing monome s on o he
in e aces o empla es [167, 168]. Howe e , al hough empla e-based me hods emain he
mos eliable [169, 170], hey c i ically depend on he a ailabili y o empla es.
In e es ingly, a s udy has pos ula ed ha he P o ein Da a Bank, [171]
www. csb.o g; [171] al eady con ains s uc u al empla es o model mos cha ac e ized
p o ein in e ac ions [52]. By aligning he in e ac ing s uc u es on o he monome s o
empla es, his s udy shown ha in mos cases i is possible o ind empla es whose
indi idual monome s sha ed TM-sco emin > 0.4 wi h monome s o a ge s uc u es. They
also sugges ed ha alignmen s wi h TM-sco emin alues g ea e han 0.4 ha e he same
mode o binding. Howe e , Neg oni and colleagues [172], ound ha while empla es
indeed exis o model he majo i y o in e ac ions, hey mos ly lead o inco ec complex
s uc u es, especially in cases o emo e homology (e.g cases sha ing sequence iden i y <
30% wi h empla es). To iden i y empla es, hey aligned he monome s o a ge
67
complexes o he in e aces o empla es and obse ed ha he pe o mance signi ican ly
de e io a es when empla es sha e only mode a e s uc u al simila i y wi h he a ge (TM-
sco e ~ 0.4 – 0.6), which is a odds wi h [52] indings.
Ano he limi a ion, in addi ion o he limi ed a ailabili y o empla es, is ha he low
sequence simila i y o emo e homologous makes empla e iden i ica ion di icul when i
is based only on he alignmen o indi idual monome s uc u es. In hese cases, ab ini io
docking can be used o build a pu a i e model o he p o ein-p o ein complex om he
known monome s uc u es (see below).
We ha e s udied in mo e de ail he capabili y o empla e-based modelling unde
di e en condi ions as well as a new app oach in eg a ing ab ini io docking wi h empla e-
based modelling ha can assis in iden i ying ab ini io docking models unde condi ions o
low sequence iden i y.
6.1.2. Ab ini io docking
Cu en ab ini io me hods gene a e a as numbe o con o ma ions using e icien sampling
echniques and disc imina e nea -na i e models om inco ec poses employing
sophis ica ed sco ing unc ions. Fas Fou ie T ans o m (FFT) sampling algo i hms disc e ize
p o eins in o g ids o accele a e he sea ch space p ocess, which a e implemen ed in
p og ams such as GRAMM-X [92], ZDOCK [67] and FTDock [66]. O he app oaches o
gene a ing docking poses use Mon e Ca lo-based sea ching, [129] such as ICM [129], o
Rose aDock [173], molecula dynamics as in HADDOCK [174], and no mal modes as in
ATTRACT [175] o Swa mDock [176]. Va ious sco ing unc ions ha e been de eloped o
selec , among he housands o gene a ed docking poses, hose ones ha a e mos likely o
esemble he na i e s uc u es. These unc ions o en include elec os a ics, desol a ion,
and an de Waals ene gy e ms such as in pyDock [79] and ZRANK [177] o s a is ical
po en ials as in SIPPER [178] o PIE [179]. Howe e , al hough ab ini io docking me hods
ha e p o ed aluable in yielding high-quali y p o ein–p o ein models [180], he limi ed
abili y o sampling me hods o sea ch he con o ma ional space, and he mul iple minima
ha sco ing unc ions gene a e lead o an ex emely high a e o alse posi i es.
68
In any case, docking unc ions can be also used o sco e empla e-based complex models
om emo e empla es.
In addi ion, he e is g owing in e es on epu posing he sco ing unc ions o analyse
ene ge ic aspec s de i ed om c ys allog aphic s uc u es and o in es iga e whe he hese
s uc u es a e biologically meaning ul (see Chap e 6.6 o mo e de ails).
6.2. Limi a ions o empla e-based docking
6.2.1. Templa e-base model gene a ion
We explo ed he e o wha ex en empla e-based docking depends on he quali y o he
a ailable empla es, in e ms o sequence iden i y wi h espec o he a ge .
To alida e he app oach, we used as a ge s uc u es he 176 e e ence complexes
o he p o ein-p o ein benchma k e sion 4.0 (BM4) [132]. We ied o iden i y sui able
empla es o hese complexes om DOCKGROUND ( e sion 1.1), which con ains 7107 non-
edundan PDB s uc u es [181]. Fi s , o ob ain sequence iden i y (SI) alues, sequence
alignmen s we e pe o med be ween he unbound a ge s uc u es o BM4 and he
DOCKGROUND complexes using he SSEARCH p og am o he FASTA package e sion 35.4.7
[182]. The BLOSUM50 ma ix was used as a sco ing ma ix wi h open and gap penal ies o
10 and 0.5, espec i ely. The E- alue equal o 10-5, was conside ed as a h eshold alue o
conside he alignmen s a is ically signi ican . Once hese alignmen s we e ob ained, we
pe o med he ele an SI il e s, by emo ing empla es wi h highe SI han a gi en alue
(see Figu e 6.1).
In addi ion, s uc u al alignmen s we e pe o med be ween he unbound a ge
s uc u es o BM4 and he in e aces o he 7107 complexes ex ac ed a 12 Å om he
DOCKGROUND da abase, as well as on he comple e monome s, by using bo h TM-align,
e sion 20130511 [59] and MM-align[60], e sion 20130815, espec i ely. In he case o
TM-align app oach, we selec ed he bes combina ion o he ou possible alignmen s
be ween he empla es and he a ge s o BM4. F om he wo TM-sco es we calcula ed he
a e aged TM-sco e (TM-sco ea), and he lowes TM-sco e (TM-sco emin).
When MM-align is used, his so wa e au oma ically gene a es he bes
combina ion, bu in e ace monome s o en di e in size, so hei con ibu ion o he global
69
TM-sco e may be di e en . To calcula e such con ibu ion, we compu ed indi idual TM-
sco e o ligands and ecep o s, aligned p e iously wi h MM_align, and he wo alignmen s
gene a ed by TM-align, using he TM-sco e p og am [183], e sion 20130511, and
calcula ed he a e age TM-sco e.
𝑇𝑇𝑇𝑇 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑎𝑎=
𝑇𝑇𝑇𝑇 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠
𝑙𝑙𝑙𝑙𝑙𝑙
+𝑇𝑇𝑇𝑇 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠
𝑟𝑟𝑟𝑟𝑟𝑟
2 (6.1)
Whe e TM-sco e ec e e s o he TM-sco e o he aligned ecep o s, TM-sco elig
ep esen s he TM-sco e o he aligned ligands, and he TM-sco ea is simply he a e age o
TM-sco e ec and TM-sco elig.
Then, he models a e gene a ed by supe imposi ion. When he empla es we e
selec ed by TM-align ( ull monome s), he selec ed empla e is aligned wi h MM-align o
ob ain a o a ion and ansla ion ma ix, hus gene a ing he model by supe posi ion. When
he empla es we e selec ed wi h MM-align (in e aces a 12Å), he o a ion and ansla ion
ma ix is a di ec ou pu , which can be used o gene a e he model in he same way. I
should be no ed ha he esul s p esen ed he e ha e been ob ained by using he BM4
monome s in he unbound con o ma ion.
Finally, models wi h a sol en accessible su ace a ea (SASA) o less han 250 Å we e
disca ded, in line wi h he app oach used by DOCKGROUND o c ea e he lib a y o non-
edundan empla es. The quali y c i e ion ha de ines a empla e as a nea -na i e
s uc u e is ha he Cα-LigRMSD be ween he base model o he empla e and he eal
s uc u e is less han o equal o 10 Å. Finally, we il e ed by sequence iden i y: 100%, 95%,
70%, 30%, and TM-sco e: 0.4, 0.5, 0.6, 0.7 and 0.8.
70
6.2.2. P edic i e success o empla e-based docking
As can be seen in Figu e 6.1, empla e-based docking shows a success a e o e 50% when
SI is 100%. This ela i ely low success a e is mainly due o wo easons. The i s eason is
due o he low edundancy o DOCKGROUND, which means ha he e a e no high SI
empla es o all BM4 cases. The second eason is ha he unbound monome s we e used
o gene a e he inal models, so he con o ma ional changes (be ween unbound and bound
monome s) a e pa ly esponsible o he models no being as good as could be expec ed.
In ac , his app oach is e y close o how empla e-based modelling is done in eal li e, as
in he blind condi ion o CASP o CAPRI expe imen s.
Figu e 6.1 also shows ha i makes i ually no di e ence whe he we use MM-
align o TM-align o selec he empla es. Bu in he so-called " wiligh zone", ha is, in
cases wi h SI below 30%, he use o MM-align p o ides a signi ican ad an age.
In he case o he MM-align esul s, we pe o med a mo e de ailed s udy (Table 6.1), by
analysing he success a es a di e en h eshold alues o sequence iden i y and TM-
sco es. The success a e is usually calcula ed on he o al numbe o he a ge s o BM4, bu
he e we ha e done i on he numbe o cases o which a success ul model can be made.
Thus, o pu he success a e alues in he co ec con ex , he co e age (%) is also
displayed. Consequen ly, he success a es a e appa en ly highe han in Figu e 5.1. Fo
52 49
45 43
30 31
14
19
53 53
46 44
31 28
12 11
0
10
20
30
40
50
60
TM-sco e
a ega e
TM-sco e
min
TM-sco e
a ega e
TM-sco e
min
TM-sco e
a ega e
TM-sco e
min
TM-sco e
a ega e
TM-sco e
min
100% SI 95% SI 70% SI 30% SI
Success Ra e %
Sequence Iden i y
Figu e 6.1. Ba -plo showing he success a es o he TM-align ( illed ba s) and MM-align (pa e ned
ba s) so wa e. Templa es a e il e ed acco ding o hei SI, so ha only empla es wi h lowe o
equal SI a e used o gene a e he models.
71
example, when he TM-sco e is se o 0.4, he TM-sco emin alues a e be e in all SI
condi ions as compa ed o TM-sco ea, bu he co e age dec eases d ama ically be ween
hem. This can also be obse ed o all TM-sco e h esholds.
Table 6.1. Templa e-based docking success a e o he op 10 models and hei co e age
(pe cen age o cases o which a model can be gene a ed o e he 176 o he BM4) o he applied
TM-sco es and sequence iden i y h esholds.
Success Ra e (%) Co e age (%)
Rep esen a i e empla e
-base sco ing
TM-sco e
a
0.4
55.6
47.8
33.6
17.3
93.1
91.4
87.4
76.4
TM-sco emin 0.4 84.2 84.1 74.5 61.5 54.6 47.1 31.6 14.9
TM-sco e
a
0.5
66.9
61.0
46.9
35.6
71.3
67.8
56.3
33.9
TM-sco emin 0.5 85.9 86.1 80.0 75.0 52.9 45.4 28.7 11.5
TM-sco e
a
0.6
83.5
83.3
73.3
64.0
55.7
48.3
34.5
14.4
TM-sco e
min
0.6
86.2
86.1
80.5
76.9
50.0
41.4
23.6
7.5
TM-sco e
a
0.7
84.6
84.4
80.0
71.4
523
44.3
25.9
8.0
TM-sco e
min
0.7
87.2
87.1
81.8
85.7
49.4
40.2
19.0
4.0
TM-sco e
a
0.8
88.2
88.4
81.8
100.0
48.9
39.7
19.0
2.9
TM-sco e
min
0.8
89.0
89.1
80.8
100.0
47.1
36.8
14.9
2.3
100
95
70
30
100
95
70
30
Sequence Iden i y
Bo h TM-sco ea o TM-sco emin ha e ad an ages and disad an ages. In he case o
TM-sco emin i could be di icul o ind a empla e wi h a TM-sco e o e 0.4, bu i ound, i
will likely lead o he gene a ion o good models. The opposi e happens wi h TM-sco ea. I
is easy o make models o almos all a ge s, bu a a cos o a lowe quali y. The e o e, in
o de o ob ain accep able models, a h eshold highe han 0.4 will be needed. Mo e de ails
on he esul ing success a es can be ound in Appendix 2: Table 10.2.1
6.3. Ab ini io docking can imp o e empla e iden i ica ion
6.3.1. A new p o ocol o combining ab ini io and empla e-based docking
We ha e de ised a modeling s a egy ha uses ab ini io docking o imp o e empla e
iden i ica ion, in h ee s ages: i) sampling docking o ien a ions wi h ab ini io docking, ii)
sea ching s uc u al empla es o docking o ien a ions using he MM-align p og am, and
72
iii) sco ing docking o ien a ions using a new unc ion ha combines pyDock docking ene gy
and TM-sco e. Figu e 5.3 illus a es he o e all p ocedu e, and he de ails a e as ollows:
Ab ini io docking
We used he pyDock scheme o gene a e a pool o docking models.
Sampling
The unbound subuni s o he Benchma k 4.0 complexes we e ansla ed and o a ed
andomly o emo e possible bias o ini ial binding condi ions. The docking poses we e
gene a ed using FTDock 2.0 [66], a as Fou ie ans o m algo i hm, which is based on
su ace complemen a i y and elec os a ics, using 0.7 Å g id cell size, su ace hickness o
1.3 Å, a o a ion angle o 12º and 3 ansla ions o each o a ion. Fo each benchma k case,
a o al o 10,000 docking poses we e ob ained.
Sco ing
We used he pyDock sco ing unc ion [184], de eloped in ou g oup, which inco po a es
elec os a ics, sol a ion, and an de Waals ene gy con ibu ions o ank he docking poses.
Fo each a ge case, we selec ed he op 100 anked docking con o ma ions.
Re-sco ing docking con o ma ions using he s uc u al simila i y sco e (TM-sco e)
i. In e ace ex ac ion
In e aces we e ex ac ed om he op 100 selec ed benchma k docking poses and
om he DOCKGROUND empla e lib a y by selec ing only hose esidues wi h Cα
a oms wi hin 12 Å dis ance om he in e ace [185]
79
6.4.2. Templa e-base docking
Templa e-based docking was o en used in bo h join CASP-CAPRI p o ein assembly
p edic ion challenges.
In he 3 d join CASP-CAPRI challenge (CASP13-CAPRI46), complexes we e modelled
based on empla es in almos 50 % o he cases (T137-T144, T152-T154, T158). In he 4 h
challenge (CASP14-CAPRI50) his pe cen age inc eased o 70 % o he cases (T164-T168,
T170, T171, T175-T177, T180, T181). The empla es we e ound hanks o he in o ma ion
p o ided by he models ex ac ed om he CASP-hos ed se e s: ZHANG, ROSETTA, QUARK,
MULTICOM-CONSTRUCT and RAPTORX-DeepModelle , as well as om a BLAST sea ch.
These empla es we e analysed and selec ed based on hei s uc u al simila i y and
biological uni o in e es . Templa es ha did no add ele an in o ma ion we e also
il e ed ou (i.e. edundan empla es we e emo ed). Monome ic models we e
supe imposed on he co esponding subuni s o each non- edundan empla e and
subsequen ly sco ed wi h he pyDock sco ing ene gy unc ion.
6.4.3. Ab ini io docking
In gene al, o he p o ein-p o ein and p o ein-pep ide cases, we gene a ed 10,000 docking
poses wi h FTDock 2.0 [66] (wi h elec os a ics and 0.7 Å g id esolu ion) and 2,000 poses
wi h ZDOCK 2.1 [67], wi h he excep ion o a ew a ge s (T131, T132, T136, T149, T159),
whe e ZDOCK was no used due o he compu a ional cos o e y la ge p o eins. In h ee
cases (T131-T133), we used he s ochas ic docking me hod Ligh Dock [77] o gene a e
addi ional lexible docking poses du ing he sampling p ocess.
In he sco ing phase, he docking poses we e sco ed wi h he pyDock sco ing
unc ion. In he abo e men ioned T131-T133 cases, he pyDockLi e [77] and DFIRE [191]
unc ions we e used. In his de aul p o ocol, co ac o s, wa e molecules and sol en ions
we e no included in ou docking calcula ions. In case o homo-oligome ic a ge s, we kep
only he docking posi ions wi h he expec ed symme y (e.g. C2 o homo-dime s, C3 o
homo- e ame s, e c.).
80
Fo p o ein-saccha ide a ge s (T126-130), we used Dock [192]
(h p:// dock.sou ce o ge.ne /) o gene a e and sco e he models. In addi ion, we
de eloped a new pyDock module specially adap ed o sco ing saccha ide molecules.
6.4.4. Combining ab ini io and empla e-based me hods
In he 7 h CAPRI expe imen and in he 3 d and 4 h join CASP-CAPRI p o ein assembly
p edic ion challenges, some a ge s we e modelled by combining empla e-based and ab
ini io docking.
In he T136 a ge o 7 h CAPRI, he homo-decame in e aces we e modelled based on he
a ailable BLAST empla e, by supe imposing he bina y ab ini io docking models on he
global empla e (PDB 5FKZ).
In he 3 d join CASP-CAPRI expe imen , he e we e a leas h ee a ge s whe e he same
s a egy was used:
• T146 (A2B2): he homodime in e aces we e modelled based on all a ailable
empla es om CASP-hos ed se e s, and he he e ome ic in e aces we e ob ained
om ab ini io docking. We used C yo-EM in o ma ion[193] o localize he ligand-
p o ein and il e he docking esul s.
• T147 (A8): a ailable empla es (PDB codes 2W1V, 2GGL and 5H8I) we e used o
gene a e homodime ic models, hen ab ini io docking was pe o med o build
e ame s, keeping only hose wi h 2- old helical symme y.
• T159 (A6B6C6): he h ee homo-hexame ic ings we e independen ly modelled
based on he empla es ound in he CASP-hos ed se e s (PDB ID: 1Y12, 3EAA, 4HE1
y 3V4H). Then, using he 3J2M empla e, we buil he he e o-dodecame . F om his
mac os uc u e and he 3 d ing, igid body docking was applied o model he inal
homo-oc adecame , keeping only hose models whe e he in e ac ing ings
o e lapped in a consis en way.
In he 4 h join CASP-CAPRI expe imen , he e we e challenging a ge s whe e he
combined sampling s a egy was applied:
81
• T165 (A3H3L3): he homo ime ic glycop o ein was modelled by empla e-based
docking (CASP models o a ailable empla es we e used), while he he e o-dime ic
an ibody was modelled wi h MODELLER 9.19 because no models we e a ailable on
he CASP-hos ed se e s. These models we e docked o o m he inal model.
• T170 (A6B3C12D6): in his case we applied an ad hoc modelling p ocedu e
combining ab ini io docking, empla e-based modelling and manual i ing using a
c yogenic elec on mic oscopy (c yo-EM) map. The a ge consis s o h ee ings
wi h di e en s oichiome y and p o ein composi ion. The i s ing was a homo-
hexame , a anged as a dime o ime s, and was modelled by s uc u al
supe imposi ion i ing he X- ay monome s uc u e (PDB 5NGJ, chain A) in he
a ailable c yo-EM map o he ail o bac e iophage T5 (EMDB ID: 3689). The second
ing is o med by h ee p o ein subuni s o one ype and wel e o a second ype and
was modelled by sequen ially building bina y in e ac ions wi h ab ini io docking and
symme y cons ain s. The hi d ing was modelled in a simila way.
The inal assembly o he modelled ings was pe o med by ab ini io docking,
selec ing only hose models in which he symme y axes o he ings we e aligned
(see Figu e 6.4).
Figu e 6.4. Modeling s a egies o he h ee majo ings o a ge T170
82
• T177 (A20): his complex is o med by wo homodecame ic ings. Each ing was
modelled wi h MODELLER 9.19 om a ailable empla es, and he inal assembly
was buil by applying ab ini io docking o he wo modelled decame s.
6.4.5. Inclusion o es ain s om a ailable ex e nal da a
Besides he abo e men ioned au oma ic modeling p ocedu es, we o en used a ailable da a
o each speci ic case in o de o es ain he docking o ien a ions and help selec ing he
co ec models. Some o he mos impo an es ain s we applied we e based on he
oligome iza ion s a es o he a ge s. This in o ma ion was mainly ob ained om he CAPRI
and CASP desc ip ion o he a ge s, bu we also ex ac ed i om empla es o
oligome iza ion s a e p edic o s [194]. I homo-oligome s we e in ol ed, we assumed
symme ic oligome iza ion, e.g. cyclic symme y C2, C3..., o il e he esul ing docking
models. This s a egy was widely used in he 7 h CAPRI (T125, T124, T136) as well as 3 d
(T137-T141, T143-T144, T147, T148, T152-T154, T158, T159) and 4 h join CASP-CAPRI
expe imen s (T164, T165, T167-T171, T174-T176, T178, T179).
I expe imen al in o ma ion was also a ailable, we included i in he modeling p ocedu e as
dis ance es ain s. We used a a ie y o es ain sou ces (see Figu e 6.5). Fo example, we
es ima ed in e ace esidues o be used as docking and sco ing es ain s om homologous
p o ein s uc u es o conse ed p o ein-p o ein in e ac ions. Mo e speci ically, hese we e
he es ain s applied in each o he expe imen s:
7 h CAPRI:
In he case o he T126-T130 a ge s (p o ein-saccha ide), a mo e speci ic so wa e was used
o gene a e he Dock complexes [195] (h p:// dock.sou ce o ge.ne /). Dock can include
cons ain s o a ge he binding ca i y. The cen e o he ca i y used was de ined as he
cen e o mass o known ligands bound o homologous p o eins (PDB 5F7V o T126-T129;
PDB 3D5Z o T130).
The pyDockRST [159] module was used in mul iple a ge s o e icien ly add dis ance
es ain s. In some cases, we ound homologous empla es om which we deduced he
es ain s o be applied: In a ge s T134-T135 (homologous empla e: PDB 1F95) a dis ance
83
es ain o 5 Å wi h espec o he esidues o he pep ide was used. Simila ly, a dis ance
es ain o 10 Å was applied in a ge s T153 (homologous empla e: PDB 3W36) and T136
(homologous empla e: PDB 5FKZ). In o he a ge s, we di ec ly applied in o ma ion abou
he in e ac ion ha was a ailable in he li e a u e, such as in T122 (T p156 and IL-23A)
[196], T125 (LLT1 Lys169, and NKR-P1 Glu205) [197, 198], and T131-T132 (Ty 35 and Ile92
in he common hCEACAM1 p o ein) [199].
3 d join CASP-CAPRI expe imen :
The a ge T149 was highly challenging as i no only in ol ed he dime iza ion o a
5-domain p o ein, bu i was also necessa y o desc ibe he assembly o he 5 di e en
domains (D) wi hin each monome . The pyDockTET module [141] was used oge he wi h
an ad hoc s a egy o gene a e he models. Each domain was modelled independen ly based
on he ank 1 p edic ion o he QUARK CASP-hos se e (each domain was a CASP13 a ge ).
Nex , he in e molecula o ien a ion be ween he i s domains o each monome (D1-D1')
was modelled based on a empla e (PDB 1DQS). The in e ac ion be ween D1 and D2 o he
same monome was modelled by docking, imposing es ain s de i ed om he in e -
domain bonds wi h he pyDockTET module. Fo each D1-D2 model, a copy o i (D1'-D2')
was supe imposed on D1-D1' o gene a e D2-D2' pai s. This s a egy was i e a i ely applied
o he o he domains (D2-D3 by docking, D2'-D3' by o e lapping, D3-D4 by docking, e c.). A
isual cu a ion o he models was pe o med o a oid la ge clashes be ween domains and
hen he models we e sco ed using he pyDock sco ing unc ion.
Ano he in e es ing a ge s we e T149, T150 and T151, which we e sequen ially eleased
o he same p o ein complex, bu wi h inc easing a ailable expe imen al in o ma ion. In
a ge T150 we e-e alua ed he ab ini io docking o ien a ions gene a ed o a ge T149
wi h he help o SAXS da a, o which we used he pyDockSAXS module. In he case o T151
we also added he newly a ailable c oss-linking in o ma ion in he o m o dis ance
es ain s using pyDockRST.
84
4 h join CASP-CAPRI expe imen :
Ta ge T168 was a ime . We ini ially buil docking ime s wi h ab ini io docking and
symme y cons ain s, which we e compa ed wi h an a ailable empla e (PDB 6FTD), so ha
models wi h Cα-RMSD la ge han 10 Å we e il e ed ou . In he case o T170, we used a
c yo-EM map o selec he inal models. In a ge T181 we also applied es ain s, since he
s uc u e o he sepa a e p o eins (PDB IDs 1N3U and 6XDC, espec i ely) was known. We
also knew ha 6XDC had a ansmemb ane domain ( esidues 44-64, 68-128) o which 1n3u
could no bind. This egion was used as a "nega i e" es ain , elimina ing he ab ini io
docking models ha showed binding in his egion.
6.4.6. Final selec ion o he models.
In gene al, he sco ing o he models, bo h o he p edic o s and sco e s expe imen s, was
pe o med wi h he pyDock bindEy module, which calcula es he docking ene gy o pyDock
CASP-hos ed se e p edic ions
Top5 p edic ions Top5 p edic ions
Deep Modelle
MULTICOM -CONSTRUCT
Docking
A:A
Building
A3
Example T168(A3)
A:A /A:B
3models Up o 25models
CASP Con ibu�on
>> models
>> empla es o CAPRI
-
Fil e ing by a om clashes <250
-
Clus e ing 4Å
-
Minimiza�on (AMBER12)
pyDock
h ps://li e.bsc.es/pid/pydock/
1
2
3
Selec ed empla es Models
Sampling:
Templa e-base
ICM-B owse (Te mpla e -supe posi�on)
•Sequence Iden� y empla es
•Templa es used in CASP-hos ed p edic�ons
Sampling:
P o ein-p o ein Docking
Me ge o he dockings o c ossdocking o m:
FTDOCK(0.7Å g id, elec os a ics)
. dock ile: bes 10,000 con o ma ions
. o ile:FTDOCK o a ion exp essed in Eule angles
ZDOCK 2.1
.zdock ile:bes 2,000 con o ma ions
. o ile: ZDOCK o a ion exp essed in Eule angles
I Fil e ing
Symme y es ain s
Expe imen al da a
[C yo-EM, C oss-linking, SAXS, e c. ]
pyDockRST
Templa e-based
Sequence Iden i y: X - ay /RMN s uc u e a ailable.
pyDock ene gy sco ing
.ene ile: Table wi h all con o ma ion e-
sco ed using pyDock ene gy
Submission
Deep Modelle
MULTICOM -CONSTRUCT
Supe posi ion on
ex ac ed empla es
om S.I. &CASP-
hos ed p edic�ons
empla es.
Figu e 6.5. P o ocol ollowed in he join CAPRI-CASP14 expe imen s using he T168
a ge as an example. The combina ion o empla e-based modelling and ab ini io
docking is shown. I also shows how empla es and ele an in o ma ion can be used o
il e be ween he ab ini io models.
85
[80] ( o de ails, see Chap e 3.1.4.3) o a gi en (expe imen ally de e mined o modelled)
p o ein complex. Fo a ge s whe e possible in e ace esidues could be de ined based on
a ailable expe imen al in o ma ion o homologous complexes, his in o ma ion was usually
included in he inal sco e as dis ance es ain s wi h pyDockRST [159], pyDockSAXS [140]
and/o pyDockTET [141], as desc ibed in p e ious sec ion.
O he sco ing s a egies we e also used o speci ic a ge s, as in T133, an a i icially
designed complex based on an old a ge (T47). Such a edesign inc eased he p edic ion
complexi y wi h espec o T47 and allowed us o explo e di e en s a egies. The e o e,
we changed he adi ional pyDock sco ing me hod success ul applied o T47 by in eg a ing
new me hodologies such as ligh Dock [77], which added mo e lexibili y o he models, and
CCha PPi [78], which p o ided a mo e signi ican numbe o sco ing unc ions, which we e
in eg a ed using he IRaPPA o ing algo i hm [86]. This p o ocol is implemen ed in he
pyDockResco ing (h ps://li e.bsc.es/pid/pydock esco ing/) se e . Fo he p o ein-pep ide
a ge s o he 7 h CAPRI expe imen (T134, T135 and T121), we cons ained he docking
models o adop he an i-pa allel β-chain o ien a ion, which u ned ou o be co ec o
a ge s T134, T135, bu inco ec o a ge T121. A e sco ing, we emo ed edundan
p edic ions using a BSAS algo i hm [153] wi h a dis ance limi o 4.0 Å, as p e iously
desc ibed [200]. Fo models based on empla es o symme y cons ain s, we elimina ed
hose wi h s ong clashes ha would be di icul o sol e wi h minimiza ion.
The numbe o a ailable empla es and hei eliabili y de e mined he pe cen age
o empla e-based complex models included in he inal 5 o 10 models ha we submi ed
o CASP o CAPRI, espec i ely. Models wi h mo e han 250 collisions (i.e., in e molecula
pai s o a oms close han 4 Å) we e also elimina ed. Then he inal en selec ed docking
poses we e minimized using di e en e sions o AMBER (AMBER12 [193] o AMBER17
[194]), wi h implici sol en o imp o e he quali y o he docking models and educe he
numbe o in e a omic clashes, as p e iously desc ibed [196]. We always used he same
a omic pa ame e s: om AMBER 99SB [201] o ce ield o p o eins, and ga o ce ield
o polysaccha ides [202]. The minimiza ion p o ocol consis ed o a 500-cycle s eepes
descen (SD) minimiza ion wi h ha monic cons ain s applied a a o ce cons an o 25
kcal/(mol-Å2) o all backbone a oms in o de o op imize he side chains, ollowed by
86
ano he 500 cycles o uncons ained conjuga e g adien (CG) minimiza ion. In some cases,
due o ime cons ain s o ea lie con e gence, he minimiza ion p o ocol a ied (e.g., in
T134 we used 200-cycle SD and 300-cycle GC; in T135 500-cycle SD and 100-cycle GC; in
T136 some models we e no minimized o we e acuum minimized; he la ge CASP a ge s
T149-151, T159, T165, T170, T177, and T180 we e acuum minimized). In a ge s T131 and
T132, loops p e iously emo ed o docking we e econs uc ed by MODELLER be o e he
inal minimiza ion s ep.
The p o ocol we used o he inal selec ion o models in he sco e s expe imen was
he same as he one we used in he p edic o expe imen s, excep o a ew excep ions, as
ollows: Dis ance es ain s we e no used as sco e s in T121 and T136; IRaPPA was no used
as sco e s in T133. In a ge T174, models wi h a Cα-RMSD<10 we e il e ed agains a
common empla e domain. In a ge T175, no empla e was used in sco e s (while i was
used in p edic o s). In T181, in addi ion o il e ing he models using he 6XDC
ansmemb ane egion (only he A chain was used o ab ini io docking), he in e ace egion
wi h he o he monome , es ima ed om he empla e (amino acids 221-288), was also
used o il e he models.
6.4.7. Sco ing o p o ein-saccha ide complex models
Fo he p o ein-oligosaccha ide a ge s (T126-T130), we used Dock [192]
(h p:// dock.sou ce o ge.ne /) o gene a e and sco e he models. In addi ion, we had o
implemen new unc ionali ies in pyDock (see Chap e 4.4), since he o iginal e sion did
no ha e a omic elec os a ics, sol a ion, and an de Waals pa ame e s o saccha ide
molecules.
A e his implemen a ion, pyDock was able o ead opology and coo dina e iles
om AMBER o all ypes o molecules, and hus calcula e he ene gy-based sco ing
unc ion. The da a ob ained om AMBER iles a e an de Waals ene gies and a omic pa ial
cha ges. As o he a omic sol a ion pa ame e s (ASPs), a dic iona y o equi alences was
c ea ed be ween he new AMBER a om ypes and he pyDock a om ypes o iginally used
o p o eins (h ps://gi hub.com/pyDock/pa allel/blob/mas e /ambe _old_ o_new.map).
Basically, he ASPs o saccha ide C and O a oms we e conside ed as hose o "C alipha ic"
87
and "O hyd oxyls" o o iginal pyDock, espec i ely. To ob ain he AMBER iles men ioned
abo e, we used an echambe wi h he AM1-BCC cha ge model [203], se ing he ne cha ge
o 0, and hen pa mchk2 o ob ain he cha ges, ene gy angle pa ame e s, and a mol2 ile.
We hen used LEaP o load he gene al AMBER o ce ield (GAFF) and ollowed he
p ocedu e o gene a e a lib a y wi h he in o ma ion ob ained om he an echambe . As a
inal s ep, we used LEaP o load each docking pose and ob ain i s coo dina e (.inc d) and
opology (.p m op) iles. These wo iles a e he ones ha can be di ec ly used by pyDock
o calcula e he ene gy wi h he bindEy module. In he sco ing expe imen , we used ano he
cha ge model due o ime cons ain s, he empi ical a omic pa ial cha ges o Gas eige -
Ma sili [203], and he inal sco e was based solely on his new e sion o pyDock adap ed
o glycosidic p o ein in e ac ions ( Dock was no used).
6.5. E alua ion o he p edic i e esul s in CAPRI and CASP
The models submi ed o CAPRI and CASP wi h he de eloped me hodology desc ibed in
he p e ious sec ion we e o icially e alua ed by he o ganiza ion o CAPRI and CASP, which
is a use ul exe cise ha allows us o ha e a mo e objec i e knowledge abou he
applicabili y and limi a ions o ou me hodological app oaches, and a ai compa ison
be ween me hods om o he g oups.
In he 7 h edi ion o CAPRI, we pa icipa ed in all a ge s, as p edic o , sco e s and
se e s, he la e wi h he excep ion o he p o ein-saccha ide cases since ou pyDockWeb
[82] se e was no eady o au oma ic p ocessing o his ype o in e ac ions. The
p edic i e pe o mance o ou g oup as well as ha o o he pa icipan s is desc ibed in ull
de ail in a p e ious publica ion [204].
In he case o he 3 d and 4 h join CASP-CAPRI expe imen s, we pa icipa ed as
p edic o s and sco e s. The p edic i e pe o mance was desc ibed in ull de ail in p e ious
publica ions [142, 144, 145]. Below we summa ize ou esul s o he 51 a ge s p oposed
in he 7 h CAPRI, 3 d and 4 h join CASP-CAPRI expe imen s (conside ing he he e o-me ic
and homo-me ic in e aces a he T125 a ge as wo sepa a e a ge s). The pe o mance
o 7 h CAPRI is summa ized in Table 6.3, and in Appendix 2: Tables 10.2.6, 10.2.7, 10.2.8,
88
o he 3 d join CASP-CAPRI in Table 6.4 and in Appendix 2: Table 10.2.10, and o he 4 h
CASP-CAPRI in Table 6.5 and in Appendix 2: Table 10.2.11.
6.5.1. 7 h CAPRI
The esul s a e consis en wi h p e ious pa icipa ions, as we submmi ed accep able
models o 10 a ge s as p edic o s, ou as se e s and 13 as sco e s (Table 6.3). The o al
numbe o e alua ed in e aces was 19, as he e we e mul iple in e aces ha we e
conside ed independen a ge s. These esul s, conside ing ou 10 bes models, ep esen
a success a e o 53% as p edic o s, 21% as se e s and 68% as sco e s, he la e being he
bes o all pa icipan s. Using he CASP c i e ia in which only he op 5 submi ed models
a e e alua ed, he esul s as p edic o s and se e s did no change, bu he pe o mance as
sco e s was signi ican ly educed. Based on he esul s om he e alua ion o all
pa icipan s, he o ganiza ion conside ed ha he e we e nine di icul a ge s, h ee o
medium di icul y, and i e easy a ge s. The pe o mance summa y is ep esen ed in Table
6.3, which shows he quali y o he 10 bes p edic ions submi ed o each in e ace and/o
a ge in which we pa icipa ed (high***, medium**, and accep able* [204]).
95
Table 6.5 Quali y o submi ed p edic ions o he CASP14-CAPRI50 expe imen p edic ions.
applied an ad-hoc modeling p ocedu e (see Chap e 6.4.4), also combining ab ini io docking
and empla e-based modeling. We ob ained wo accep able models o in e aces #1 (A:B)
and #2 (A:E), whe e we we e he only g oup o submi an accep able model. Howe e , in
he a e aged e alua ion o in e aces #8 (D) and #9 (CiD), we ob ained accep able esul s
as human and se e sco e s (along wi h Venclo as g oup, hese we e he only accep able
ank one submissions om all pa icipan s).
6.5.3.2. Unsuccess ul p edic ions
This 4 h join CASP-CAPRI p o ein assembly p edic ion challenge has p o en o be mo e
complex han he p e ious one, wi h highe p opo ion o di icul cases. Fo a ge s T169,
T165, and T174, no g oup managed o gene a e models o accep able quali y. Fo a ge s
T164, T169, T176 and T174, he p e e ed s a egy in p edic o s and sco e s was ab ini io
Easy Ta ge s S oich. Submission quali y
o P edic o s
Success ul
g oups2
Submission quali y
o Sco e s
(human)
Submission quali y
o Sco e s
(se e s)
Success ul g oups
o Sco e s
T164
A2
-
19/28
*
*
17/23
T166
A1B1
**
17/24
**
**
14/19
T168
A3
**
18/24
**
**
17/20
T177
A20
***/***/-
21/22/14 o 24
**/***/**
***/***/-
17/18/14 o 18
Medium-Di .
Ta ge s S oich. Submission quali y
o P edic o s
Success ul
g oups
Submission quali y
o Sco e s
(human)
Submission quali y
o Sco e s
(se e s)
Success ul g oups
o Sco e s
T169
A2
-
0/27
-
-
0/20
T176
A2
-
3/27
-
-
9/21
T178
A2
*
13/26
-
-
10/19
T179
A2
-
10/25
*
*
17/19
T165
A3H3L3
-
0/27
-
-
0/22
T174
A3
-
0/24
-
-
0/19
T170 A6/B3/C12/D6
*/*/-/-/-/-/-/-/-
13/1/5/4/12/1/1
/7/6 o 23 **/-/-/*/**/-/-/*/*
**/
-/-/*/**/-/-/*/*
17/0/10/9/16/2/2/
13/13 o 17
T180
A240
-/**
1/19 o 25
*/**
*/*
4/17 o 18
1***=high-quali y models; **=medium-quali y models; *=accep able-quali y models o he esul s in he pe o mance [170]. 2In
gene al, docking se e s and CASP14
-CAPRI p edic o s a e included. Fo T167, T175 and T181, he e a e no e alua ed da a p o ided
by
CAPRI. The 3 a ge s o T171-173 a e cancelled by CASP.
96
docking, which p o ed o be unsa is ac o y. This was also he case o a ge T164, whe e
we used a combina ion o empla e-based and ab ini io docking. In e es ingly, we we e
success ul as sco e s in bo h T164 and T179 a ge s, whe e he lack o empla es was a
de e minan ac o o p edic o s bu no o sco e s.
6.6. P o ein docking unc ions o he iden i ica ion o physiological
homodime s
In his sec ion, we ha e e alua ed he applica ion o docking and sco ing unc ions o he
p edic ion o homo-dime assemblies in o he challenging scena ios. As men ioned in he
in oduc ion o his chap e , s uc u al da a on p o ein-p o ein in e ac ions a e c ucial o
elucida e hei mechanism o ac ion a he molecula le el. The o ma ion o hese
mac omolecula assemblies depends on p o ein concen a ions and physicochemical
condi ions such as pH and ionic s eng h [206]. Bu bo h he p o ein concen a ions and he
expe imen al condi ions used du ing he use o such me hods o en di e om he
physiological condi ions unde which he in e ac ions and/o o ma ion o mac omolecula
complexes occu . The e o e, in some cases, when expe imen al pa ame e s di e
su icien ly, he 3D s uc u es de e mined may esul in assemblies ha do no ep esen
physiological eali y. Non-physiological assemblies can also be o med when he s uc u e
unde s udy ep esen s a pa o a la ge complex and has been econs i u ed in he absence
o addi ional componen s.
This p oblem is se e e in he case o X- ay di ac ion c ys allog aphy, by which 86%
o he p o ein s uc u es in PDB a e de e mined (da a om Oc obe 2022). When p o eins
o m a c ys al, spu ious con ac s be ween p o eins can occu only o s abilize he c ys al
la ice, bu migh no be ele an in solu ion. In some p o eins, especially in appa en ly
homo-dime ic c ys als, iden i ying which o hese con ac s a e physiologically ele an is
p oblema ic, which makes i di icul o assign he ue oligome ic s a e, i.e., whe he i is a
monome , a homodime , o a highe o de assembly. The au ho s gene ally p o ide hese
da a du ing he PDB deposi ion p ocess bu emain p one o e o s, as hey equi e
independen biophysical/biochemical cha ac e isa ion, which can gi e ambiguous esul s.
97
Mo eo e , he e is a signi ican ac ion (18%) o p o ein assemblies esol ed by X- ay
c ys allog aphy on he PDB o which no associa ed publica ion exis s.
Se e al compu a ional me hods ha e been de eloped o in e he oligome ic s a e
o p o eins di ec ly om he con ac s made by p o eins in he esol ed c ys al la ice
s uc u e. These me hods make use o he esul s o a la ge numbe o p e ious s udies, in
which he p ope ies o he in e aces o na i e p o ein complexes we e sys ema ically
e alua ed and compa ed wi h hose o he c ys al con ac s, conside ed o ep esen weak
non-speci ic in e ac ions [207]. The mos used me hods a e PISA [194], EPPIC [208] and
PRODIGY-c ys al [209] .
PISA e alua es he chemical and s uc u al p ope ies o he in e aces, while EPPIC
uses geome ic measu emen s and sequence en opy o homolog sequences, which is o en
associa ed wi h egions in ol ed in biological unc ion [210].
To add ess he p oblem, we ha e pa icipa ed in an ac i i y p oposed by he ELIXIR
3D-BioIn o Communi y, which p o ided a benchma k composed o 1677 homo-dime ic
complex s uc u es, o which 841 a e non-physiological and 863 a e physiological (so-called
"dime benchma k e sion 3"; o mo e de ails see Chap e 3.5.6). We ha e e alua ed he
capabili ies o ou pyDock sco ing unc ions, as well as o he desc ip o s in CCha PPI
(Compu a ional Cha ac e isa ion o P o ein-P o ein In e ac ions) se e , ega ding he
iden i ica ion o physiological homo-dime s in his benchma k.
98
6.6.1. Applicabili y o pyDock and CCha PPI sco ing unc ions o he
iden i ica ion o physiological homodime s.
We e alua ed pyDock sco ing [79], as well as each o i s indi idual ene ge ics e ms,
elec os a ic, desol a ion, and an de Waals, oge he wi h 88 desc ip o s in CCha PPI [78]
(h ps://li e.bsc.es/pid/ccha ppi/), by
applying hem o he p oposed cases in
he dime benchma k e sion 3 de eloped
in collabo a ion wi h he ELIXIR 3D-BioIn o
Communi y.
Fo each o hese desc ip o s, we
e alue ed whe he hey we e able o
iden i y he co ec physiological dime s,
and hus calcula ed di e en p edic i e
success me ics, such as sensi i i y o
co e age (TPR), p ecision o posi i e
p edic i e alue (PPV), accu acy (ACC) and Ma hews co ela ion coe icien (MCC).
𝑃𝑃𝑠𝑠𝑠𝑠𝑠𝑠𝑃𝑃𝑠𝑠𝑃𝑃𝑠𝑠𝑃𝑃 𝑠𝑠𝑠𝑠 𝑃𝑃𝑠𝑠𝑠𝑠𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑠𝑠 𝑃𝑃𝑠𝑠𝑠𝑠𝑃𝑃𝑃𝑃𝑠𝑠𝑃𝑃𝑠𝑠𝑃𝑃 𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑠𝑠 (𝑃𝑃𝑃𝑃𝑉𝑉) = 𝑇𝑇𝑃𝑃
𝑇𝑇𝑃𝑃 +𝐹𝐹𝑃𝑃
(6.6)
𝑁𝑁𝑠𝑠𝑁𝑁𝑉𝑉𝑃𝑃𝑃𝑃𝑃𝑃𝑠𝑠 𝑃𝑃𝑠𝑠𝑠𝑠𝑃𝑃𝑃𝑃𝑠𝑠𝑃𝑃𝑠𝑠𝑃𝑃 𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑠𝑠 (𝑁𝑁𝑃𝑃𝑉𝑉)=𝑇𝑇𝑁𝑁
𝑇𝑇𝑁𝑁 +𝐹𝐹𝑁𝑁
(6.7)
𝑆𝑆𝑠𝑠𝑃𝑃𝑠𝑠𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 𝑠𝑠𝑠𝑠 𝑇𝑇𝑠𝑠𝑉𝑉𝑠𝑠 𝑃𝑃𝑠𝑠𝑠𝑠𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑠𝑠 𝑅𝑅𝑉𝑉𝑃𝑃𝑠𝑠 (𝑇𝑇𝑃𝑃𝑅𝑅)=𝑇𝑇𝑃𝑃
𝑇𝑇𝑃𝑃 +𝐹𝐹𝑁𝑁
(6.8)
𝐴𝐴𝑠𝑠𝑠𝑠𝑉𝑉𝑠𝑠𝑉𝑉𝑃𝑃𝑠𝑠𝑃𝑃 (𝐴𝐴𝐴𝐴𝐴𝐴) =
𝑇𝑇𝑃𝑃 +𝑇𝑇𝑁𝑁
𝑇𝑇𝑃𝑃 +𝐹𝐹𝑁𝑁 +𝑇𝑇𝑁𝑁 +𝐹𝐹𝑃𝑃 (6.9)
Ma hews (𝑇𝑇𝐴𝐴𝐴𝐴) =
𝑇𝑇𝑃𝑃 ∗𝑇𝑇𝑁𝑁 +𝐹𝐹𝑃𝑃 ∗𝐹𝐹𝑁𝑁
�(𝑇𝑇𝑃𝑃+𝐹𝐹𝑃𝑃)(𝑇𝑇𝑃𝑃 +𝐹𝐹𝑁𝑁)(𝑇𝑇𝑁𝑁 +𝐹𝐹𝑃𝑃)(𝑇𝑇𝑁𝑁 +𝐹𝐹𝑁𝑁) (6.10)
Figu e 6.6. ROC cu es o he bes desc ip o s om
CCha PII
99
In Figu e 6.6 a ROC cu es compa ainsion be ween he selec ed desc ip o s (aco ding he
AUC (A ea unde he ROC cu e)) o he ou bes CCha PII desc ip o s and he pyDock
ene ge ic e ms,) a e shown. In addi ion, an analysis compa ing sensi i i y and sa e y is
shown in Figu e 6.7. As can be seen, he bes desc ip o s ex ac ed om CCha PII a e
NIPacking desc ibed in [211], change in o a ional en opy upon complexa ion (ROT_S),
change in ansla ional en opy upon complexa ion (TRANS_S) bo h calcula ed as in
CHARMM [212] and he su ace complemen a i y sco e (NSC) desc ibed in [211]. In he case
o he pyDock ene ge ic e ms, he bes desc ip o was desol a ion ene gy e m.
These desc ip o s a e complemen a y wi h o he desc ip o s om di e en g oups
also pa icipa ing in his ac i i y o he 3D-BioIn o communi y. When all desc ip o s a e
agg ega ed using machine lea ning me hods, such as andom o es s, hey yield a consensus
sco ing unc ion ha ob ains a highe disc imina o y powe : 0.94 AUC (A ea Unde he
Cu e).
(AUC 0.78). (AUC 0.77). (AUC 0.77). (AUC 0.77).
(AUC 0.65).
(AUC 0.63).
(AUC 0.73).
(AUC 0.59).
Figu e 6.7. Shows he sensi i i y (dashed line(s)) and accu acy (solid line(s)) o he ou bes
CCha PII desc ip o s, he pyDock ene ge ic e ms and pyDock sco ing unc ion.
100
101
7. APPLICATION TO A CASE STUDIO:
MODELING ELECTRON TRANSFER
PROTEIN COMPLEXES.
102
The esul s desc ibed he e ha e been epo ed in his publica ion: Cas ell, C., e al., New Insigh s
in o he E olu ion o he Elec on T ans e om Cy och ome o Pho osys em I in he G een and
Red B anches o Pho osyn he ic Euka yo es. Plan Cell Physiol, 2021. 62(7): p. 1082-1093
103
7.1. In oduc ion
In his chap e o he hesis, we ha e applied some o he de eloped docking ools o model
p o ein complexes in ol ed in elec on ans e in pho osyn hesis. Pho osyn hesis is
essen ial o cap u ing and s o ing sola ene gy in he biosphe e. Thanks o his p ocess,
plan s abso b millions o onnes o CO2 pe yea on a ne basis. A he molecula le el, he
e iciency o his p ocess lies in a se ies o highly op imized mul i-p o ein complexes o
elec on ans e , such as Pho osys em I, one o he mos e icien pho oelec ic sys ems in
na u e, which can con e sola ene gy in o chemical ene gy a almos 100% e iciency. The
basic mechanisms and componen s o pho osyn hesis ha e been conse ed h oughou
e olu ion om cyanobac e ia o highe plan s, al hough he e a e in e es ing di e ences.
In pho osyn he ic o ganisms, wo p o eins ac as elec on anspo e s om
cy och ome o pho osys em I: plas ocyanin (Pc) (con aining coppe ) and cy och ome c6
(Cc6) (con aining i on). The wo p o eins a e e y di e en in composi ion bu ha e
equi alen s uc u al and unc ional aspec s. By analysing he a ailable h ee-dimensional
s uc u es and applying some o he compu a ional modelling me hods de eloped and/o
op imize in his hesis, i has been possible o iden i y impo an de ails o he molecula
mechanisms o pho osyn hesis in di e en o ganisms.
In cyanobac e ia and many g een algae, he wo p o eins can be used al e na i ely,
depending on he en i onmen al condi ions. Howe e , highe plan s (g een lineage) ha e
only Plas ocyanin, which o ms s ong and e icien complexes o elec on ans e . On he
o he hand, species o he ed lineage, such as ed algae and dia oms, ha e only cy och ome
c6, which o ms weake and less e icien complexes o elec on ans e . In e es ingly,
plas ocyanin genes ha e been ound in oceanic dia oms. In ac , in he case o Thalassiosi a
oceanica, i possesses such genes and weakly exp essing CC6, educing i s dependence on
i on. Bu a he expense o gene a ing less e icien elec on ans e complexes.
7.2. Molecula s uc u es and modelling
Modelling o he p o eins in his s udy was ca ied ou wi h MODELLER e sion 9 23 [205];
h ps://salilab.o g/modelle /, using he de aul se ings [213]. Templa e sea ching and
sequence alignmen s we e ca ied ou wi h BLAST [187] and Clus alW [214].
104
The unca ed C o ms o Phaeodac ylum ico nu um (UniP o KB: A0T0C9 CYF_PHATC) and
Thalassiosi a oceanica (UniP o KB: E7BWE1_THAOC), we e modelled using as empla e a
unca ed C o m ha was ound in Chlamydomonas einha d ii (PDB ID 1CFM) wi h which
hey sha e 58.2% and 44.6% sequence iden i y, espec i ely. The coo dina es o he
co ac o s (heme and Cu ion) we e aken om he empla e s uc u es. The s uc u e o he
small domain dele ion a ian o P. ico nu um C was gene a ed by emo ing esidues 171-
229 om he in ac p o ein s uc u e wi h ICM B owse so wa e
(h p://www.molso .com) [215]. The Thalassiosi a oceanica plas ocyanin model was
ob ained using Chlamydomonas einha d ii plas ocyanin as a empla e (PDB code 2pl ). Fo
each modelled p o ein, 100 models we e buil and he models wi h he bes DOPE (Disc e e
Op imized P o ein Ene gy) sco e we e selec ed [188], as p e iously desc ibed [216]. The 3D
s uc u e o P. ico nu um Cc6 was ob ained om he P o ein Da a Bank, PDB code: 3DMI
[217]. The ep esen a ion o he elec os a ic su ace po en ials, shown in Figu es 7.1 and
7.2, was gene a ed wi h he UCSF Chime a p og am. The Cu a om in Pc was assigned a
cha ge o +2, while he heme a omic cha ges o bo h C and Cc6 (-2) we e dis ibu ed as: Fe
(+2), wo ni ogen a oms o he ing (-1 each), and he wo p opionic acid side chains (-1
each) [218].
7.3. P o ein-p o ein docking simula ions
7.3.1. Docking sampling and sco ing
The p o ein-p o ein docking simula ions we e pe o med by pyDock scheme [79, 82],
adap ed he e o he inclusion o co ac o s du ing docking calcula ions. The se up s ep (see
Chap e 3.1.4) needs he coo dina es o he wo in e ac ion p o eins, usually as PDB iles,
bu i can also ake AMBER coo dina es and opology iles. In he o iginal implemen a ion
o his unc ionali y i was possible o include modi ied amino acids as pa o he in e ac ing
molecules, bu now we needed o modi y he au oma ic p o ocol o inco po a e Cu and he
heme g oup as sepa a ed molecules bu s ill pa o he ecep o o ligand.
Below is a de ailed desc ip ion o he s eps needed o pe o m ab in o docking on
he models o Cy och ome and Plas ocyanin o Thalassiosi a oceanica, using pyDock:
111
The same ene gy-dis ance landscape can also be obse ed o he docking models
o P. ico nu um [C :Cc6] (Figu e 7.5B). Howe e , he o ien a ion o cy och ome c6 is
di e en om ha ob ained in C. einha d ii (Figu e 7.6B, 7.6C). These o ien a ions ha e
smalle in e aces and weake elec os a ic in e ac ions. The hyd ophobic in e ac ions gain
mo e weigh , being simila o wha has been p e iously desc ibed in cyanobac e ia [221,
222]. Compa ing he bes ene gies-dis ance docking models o he [C :Cc6] complex, we
can claim ha he model o C. einha d ii has be e a ini y (-32 a.u. e sus -24.7 a.u.) han
he P. ico nu um
As in p e ious s udies, a e sion o C wi h he small domain dele ion was used o
docking. Su p isingly, o P. ico nu um, he small domain appea s o play only a mino ole
in he in e ac ion wi h Cc6 (Figu e 7.7). This is in con as wi h docking simula ion in C.
Plan C. einha d ii P. ico nu um
B
A
C
Figu e 7.6. Rep esen a i e s uc u es o he [C :Pc] complex o plan s and bes -ene gy
docking models o he [C :Cc6] complexes o C. einha d ii and P. ico nu um. (A)
Rep esen a i e s uc u e o he plan [C :Pc] complex ( u nip C and spinach Pc; PDB
code, 2pc ) (Ubbink e al., 1998). Pc is colou ed in blue and he coppe -bound His87 is
shown. (B, C) Bes -ene gy docking models o e icien ET be ween C (in ligh b own)
and Cc6 (in ed) o he g een alga C. einha d ii ( ank 1, docking ene gy –32.0 a.u.,
dis ance be ween Fe in C o heme in Cc6 o 8.2 Å) and P. ico nu um ( ank 6, docking
ene gy –24.7 a.u., dis ance be ween Fe in C o heme in Cc6 o 8.0 Å, he sho es
dis ance model).
112
einha d ii, whe e dele ion o he small domain o C a ec ed he C binding posi ions o
bo h Pc and Cc6 (Haddadian and G oss, 2006).
Finally, we compa ed models o he [C :Pc] complex om T. oceanica wi h he
p e iously desc ibed [C :Cc6] complex om P. ico n7.5.u um, as well as wi h he
equi alen [C :Pc] complex om C. einha d ii in he g een lineage. The docking be ween C
and Pc om T. oceanica gene a ed a la ge numbe o low-ene gy o ien a ions compa ed o
he [C :Cc6] complex om P. ico nu um. This seems o indica e a highe binding a ini y,
possibly due o he highe elec os a ics o he acqui ed "g een- ype" Plas ocyanin. Bu he
ene gy-dis ance landscape does no con e ge owa ds a single o ien a ion, esul ing in
di e en o ien a ions in a simila ange o dis ances and ene gies. They can coexis and be
biologically unc ional. These possible o ien a ions a e
Figu e 7.7. Supe imposed docking models o P. ico nu um. Supe imposed docking
models o he P. ico nu um [C :Cc6] complex and he model (in blue) co esponding o
a unca ed C wi hou he small domain ( ank 2, docking ene gy –31.0 a.u., dis ance
be ween Fe in C o heme in Cc6 o 8.4 Å).
113
(i) "Head-on" Pc o ien a ion, which is ela i ely simila o ha o some complexes
obse ed in cyanobac e ia. Hyd ophobic in e ac ions a e shown o p edomina e,
and he e is no in e ac ion be ween he elec os a ic zones and he small C domain
(ene gy o -24.2 a.u. and he sho es dis ance be ween Fe and Cu o 11.1 Å) (Figu e
7.9A);
(ii) "Side-on" Pc o ien a ion, simila o ha o he g een lineage complexes ( igu es 7.6A
and 7.9A), which includes he elec os a ic and hyd ophobic pa ches o bo h
p o eins and he small C domain (ene gy o -31.1 a.u.; Fe-Cu dis ance o 12.7 Å)
(iii) "In e media e" Pc a angemen , which includes he hyd ophobic pa ches and some
esidues o he elec os a ic pa ches (ene gy o -30.1 a.u.; Fe-Cu dis ance o 12.2 Å)
(Figu es 7.9C).
Figu e 7.8. Landscape o he compu a ional docking esul s o he [C :Pc] complex o T.
oceanica. The docking o ien a ions showed in he nex Figu e a e highligh ed in ed. The
dis ances be ween he i on in C and he coppe in Pc we e conside ed
114
Bu le 's look o models wi h he lowes ene gy, ega dless o dis ances. We ind a
popula ion o models wi h e en mo e a ou able ene gies (-38 o -36 a.u.), in which s ong
elec os a ic in e ac ions a e es ablished, also in ol ing addi ional posi i e g oups ou side
he usual elec on ans e egion in C . Howe e , he Fe-Cu dis ances a e abou 20 Å
(Figu es 7.9D). Ne e heless, in his o ien a ion, he highly conse ed Y84 esidue in Pc
( ypically named Y83 in cyanobac e ial and euka yo ic Pc) poin s di ec ly owa ds he heme-
binding Y1 o C (Y1-Y84 dis ance o 5.1 Å; Fe-Y84 dis ance o 9.9 Å) (Figu es 7.9D). One
could specula e ha his is an al e na i e binding mode and ha elec on ans e may be
possible.
A
B
C
D
Figu e 7.9 Rep esen a i e s uc u es o he [C :Pc] complex o plan s and bes -ene gy
docking models o he [C :Cc6] complexes o C. einha d ii and P. ico nu um. (A)
Rep esen a i e s uc u e o he plan [C :Pc] complex ( u nip C and spinach Pc; PDB
code, 2PCF) (Ubbink e al., 1998). Pc is colou ed in blue and he coppe -bound His87 is
shown. (B, C) Bes -ene gy docking models o e icien ET be ween C [1](in ligh b own)
and Cc6 (in ed) o he g een alga C. einha d ii ( ank 1, docking ene gy –32.0 a.u.,
dis ance be ween Fe in C o heme in Cc6 o 8.2 Å) and P. ico nu um ( ank 6, docking
ene gy –24.7 a.u., dis ance be ween Fe in C o heme in Cc6 o 8.0 Å, he sho es
dis ance model).
115
7.6. Conclusions
In his chap e , we ha e illus a ed he applicabili y o compu a ional docking using pyDock
on a case s udy ele an o unde s anding pho osyn hesis a he molecula le el. The
s uc u al models p o ide an explana ion o he di e ences in pho osyn he ic e iciency
be ween ed and g een algae. Bu he lowe docking ene gy model ob ained o he [C :Pc]
complex in T. oceanica (Figu e 7.8), despi e some e idence, is s ill highly specula i e. O he
app oaches, such as molecula dynamics (MD) ollowed by expe imen al con i ma ion, will
be necessa y o be able o make a de ini e s a emen
116
117
8. Gene al discussion
118
New de elopmen s o pyDock and pyDockDNA o add ess cu en challenges
This hesis desc ibes he de elopmen o echnical upg ade and new unc ionali ies in he
p o ein-p o ein docking so wa e pyDock, as well as he implemen a ion o a new web
se e o p o ein-DNA docking.
The p og am pyDock has been subs an ially upda ed in o de o ex end i s
applicabili y om a echnical poin o iew and be eady o he new challenges in he ield,
like he use o molecules di e en om p o eins, including co ac o s. Fo he immedia e
u u e, i will be ex ending by in eg a ing pyP oCT in o his new code in o de o inc ease
he ep oducibili y o he clus e ing p o ocol shown in Chap e 5.1.3 ( acili a ing a b oade
applicabili y). Cu en ly pyDock 4.0 is in de elopmen alpha phase, and will be soon upda ed
o he Be a-phase o make i publicly a ailable o he scien i ic communi y, as local
dis ibu ion and also as a web se e .
The new se e pyDockDNA shows easonable p edic i e success a es on he
a ailable benchma ks. Bu he numbe o cases ha a e a ailable o benchma king is s ill
oo low o op imal es ing o new de elopmen s. The numbe o cases in which he
s uc u e o he unbound DNA is a ailable is no likely o inc ease, so we will need o ely
on modeling me hodologies. Fo una ely, he e is a a ie y o modeling s a egies,
especially hose based on deep lea ning, which migh p o ide accu a e models o unbound
p o ein and DNA in a much la ge numbe o cases. These models migh include ensembles
o con o me s, which can also p o ide be e p edic ions when used in docking. In addi ion,
he pyDockDNA se e will be ex ended wi h new unc ionali ies. Fo ins ance, he e a e
plans o apply mo e e icien dis ance es ain s be ween esidues and nucleo ides, o
inc ease he quali y o he gene a ed models.
In eg a ion o ab ini io and empla e-based docking
In his hesis, we ha e explo ed he combina ion o ab ini io docking wi h empla e-based
docking. This s a egy is pa icula ly adequa e o cases in which only low-quali y empla es
a e a ailable, acco ding o ei he hei SI alues (using sequence alignmen s) o hei TM-
sco es (using s uc u al alignmen s). Fo ha , a se o models we e gene a ed by ab ini io
119
sampling o iden i y sui able empla es wi h s uc u al alignmen me hods. Then, he
models we e sco ed by a combined unc ion based on he empla e simila i y (TM-sco e)
and he binding ene gy (pyDock). This s a egy imp o ed he p edic ions wi h espec o
using ab ini io o empla e-based docking alone. The ad an ages o combining ab ini io and
empla e-based docking we e con i med h ough ou pa icipa ion in he CAPRI and CASP-
CAPRI ounds. In pa icula , in he 3 d Join CASP-CAPRI expe imen ou g oup ob ained
excellen esul s, whe e he ene gy-based sco ing unc ion helped o iden i y co ec
models among he di e en al e na i es gene a ed by empla e-based docking app oaches.
In a simila way, he applica ion o ene gy-based sco ing as well as o he unc ions as hose
implemen ed in CCha PPI se e [78, 79] can be ex ended o he disc imina ion o
biologically meaning ul c ys allog aphic homo-dime ic complexes om he a e ac ual
dime ic in e ac ions obse ed in c ys al packing.
Assessmen o p o ein-p o ein docking
In ou pa icipa ion in CASP-CAPRI ounds, ou p o ein-p o ein docking app oaches, ei he
ab ini io o empla e-based, needed he s uc u e o easonable models o he in e ac ing
subuni s. These we e in gene al ob ained om he CASP-hos ed se e s, which p e iously
had au oma ically gene a ed models o hese subuni s, as hey we e also a ge s o CASP
in o he ca ego ies. Du ing he 4 h Join CASP-CAPRI ound, he de elope s o AlphaFold
(AF) pa icipa ed in CASP15 a ge s, ob aining unp eceden ed esul s in he ab ini io
p edic ion o he p o ein s uc u es [64]. Bu since his p og am did no pa icipa e as
se e s, he CAPRI communi y did no use hese models o he mul ime ic assembly ound.
Many o he a ge s p oposed in ha ound (Chap e 6.5.3) did no ha e su icien ly good
empla es o quali y modeling, which a ec ed o he success o he docking p edic ions.
This has changed in he las ound (5 h Join CASP-CAPRI), whe e he AlphaFold models o
he indi idual subuni s and he complexes we e a ailable o he pa icipan s in he
mul ime ic assembly sec ion. In addi ion, he AF-Mul ime e sion can now model p o ein
complexes, wi h epo ed success a es o 51% wi h a alse posi i e a e (FRP) o 1% (pape
in p ep in ) [88]. This Join CASP-CAPRI ound will be a i e es o AF and will p o ide he
120
p edic i e capabili ies o his app oach in blind condi ions. These esul s will be discussed
a he nex mee ing, o be held in An alya (Tu key) om 10 o 13 Decembe 2022.
Applica ion o cases o biological in e es
We ha e illus a ed he applicabili y o compu a ional docking using pyDock on a case s udy
ele an o unde s anding pho osyn hesis a he molecula le el. The s uc u al models
p o ide an explana ion o he di e ences in pho osyn he ic e iciency be ween ed and
g een algae. Bu he lowe docking ene gy model ob ained o he [C :Pc] complex in T.
oceanica (Figu e 7.7), despi e some e idence, is s ill highly specula i e. O he app oaches,
such as molecula dynamics (MD) ollowed by expe imen al con i ma ion, will be necessa y
o be able o make a de ini e s a emen .
127
10.2. Appendix 2. Supplemen a y ma e ial o Chap e 6
Table 10.2.1. Success a e o empla e base modelling by using MM-align.
Success Ra e MM-align 100% SI and 0.4 o TM-
sco e (MM-align)
Success Ra e MM-align 95% SI and 0.4 o TM-
sco e (MM-align)
Success Ra e MM-align 70% SI and 0.4 o TM-
sco e (MM-align)
Success Ra e MM-align 30% SI and 0.4 o TM-
sco e (MM-align)
Top NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
1
83
51.23
78
82.11
71
44.65
67
81.71
46
30.26
39
70.91
18
13.53
14
53.85
5
88
54.32
80
84.21
75
47.17
69
84.15
51
33.55
41
74.55
23
17.29
16
61.54
10
90
55.56
80
84.21
76
47.80
69
84.15
51
33.55
41
74.55
23
17.29
16
61.54
100
91
56.17
80
84.21
77
48.43
69
84.15
54
35.53
41
74.55
27
20.30
16
61.54
Co e age
162
93.10
95
54.60
159
91.38
82
47.13
152
87.36
55
31.61
133
76.44
26
14.94
Success Ra e MM-align 100% SI and 0.5 o TM-
sco e (MM-align)
Success Ra e MM-align 95% SI and 0.5 o TM-
sco e (MM-align)
Success Ra e MM-align 30% SI and 0.5 o TM-
sco e (MM-align)
Success Ra e MM-align 30% SI and 0.5 o TM-
sco e (MM-align)
Top NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
1
81
65.32
78
84.78
69
58.47
67
84.81
43
43.88
39
78.00
16
27.12
14
70.00
5
83
66.94
79
85.87
72
61.02
68
86.08
46
46.94
40
80.00
21
35.59
15
75.00
10
83
66.94
79
85.87
72
61.02
68
86.08
46
46.94
40
80.00
21
35.59
15
75.00
100
83
66.94
79
85.87
72
61.02
68
86.08
47
47.96
40
80.00
21
35.59
15
75.00
Co e age
124
71.26
92
52.87
118
67.82
79
45.40
98
56.32
50
28.74
59
33.91
20
11.49
Success Ra e MM-align 100% SI and 0.6 o TM-
sco e (MM-align)
Success Ra e MM-align 95% SI and 0.6 o TM-
sco e (MM-align)
Success Ra e MM-align 70% SI and 0.6 o TM-
sco e (MM-align)
Success Ra e MM-align 30% SI and 0.6 o TM-
sco e (MM-align)
Top NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
1
80
82.47
75
86.21
69
82.14
62
86.11
43
71.67
33
80.49
15
60.00
10
76.92
5
81
83.51
75
86.21
70
83.33
62
86.11
44
73.33
33
80.49
16
64.00
10
76.92
10
81
83.51
75
86.21
70
83.33
62
86.11
44
73.33
33
80.49
16
64.00
10
76.92
100
81
83.51
75
86.21
70
83.33
62
86.11
44
73.33
33
80.49
16
64.00
10
76.92
Co e age
97
55.75
87
50.00
84
48.28
72
41.38
60
34.48
41
23.56
25
14.37
13
7.47
Success Ra e MM-align 100% SI and 0.7 o TM-
sco e (MM-align)
Success Ra e MM-align 95% SI and 0.7 o TM-
sco e (MM-align)
Success Ra e MM-align 70% SI and 0.7 o TM-
sco e (MM-align)
Success Ra e MM-align 30% SI and 0.7 o TM-
sco e (MM-align)
Top NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
1
77
84.62
75
87.21
65
84.42
61
87.14
36
80.00
27
81.82
10
71.43
6
85.71
5
77
84.62
75
87.21
65
84.42
61
87.14
36
80.00
27
81.82
10
71.43
6
85.71
10
77
84.62
75
87.21
65
84.42
61
87.14
36
80.00
27
81.82
10
71.43
6
85.71
100
77
84.62
75
87.21
65
84.42
61
87.14
36
80.00
27
81.82
10
71.43
6
85.71
Co e age
91
52.30
86
49.43
77
44.25
70
40.23
45
25.86
33
18.97
14
8.05
7
4.02
Success Ra e MM-align 100% SI and 0.8 o TM-
sco e (MM-align)
Success Ra e MM-align 95% SI and 0.8 o TM-
sco e (MM-align)
Success Ra e MM-align 70% SI and 0.8 o TM-
sco e (MM-align)
Success Ra e MM-align 30% SI and 0.8 o TM-
sco e (MM-align)
Top NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
NSC
TM-sco e
a e age
NSC
TM-sco e
min
1
75
88.24
73
89.02
61
88.41
57
89.06
27
81.82
21
80.77
5
100
4
100
5
75
88.24
73
89.02
61
88.41
57
89.06
27
81.82
21
80.77
5
100
4
100
10
75
88.24
73
89.02
61
88.41
57
89.06
27
81.82
21
80.77
5
100
4
100
100
75
88.24
73
89.02
61
88.41
57
89.06
27
81.82
21
80.77
5
100
4
100
Co e age
85
48.85
82
47.13
69
39.66
64
36.78
33
18.97
26
14.94
5
2.87
4
2.30
128
Table 10.2.2. Success a e o empla e base modelling by using TM-align.
Success Ra e TM-align 100% SI and 0.4 o TM-
sco e (TM-align)
Success Ra e TM-align 95% SI and 0.4 o
TM-sco e (TM-align)
Success Ra e TM-align 70% SI and 0.4 o
TM-sco e (TM-align)
Success Ra e TM-align 30% SI and 0.4
o TM-sco e (TM-align)
Top NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC TM-sco e
min
1
90
52.02
87
53.05
75
43.35
71
44.10
48
27.91
44
28.03
17
9.94
13
8.90
5
91
52.60
90
54.88
77
44.51
75
46.58
51
29.65
48
30.57
18
10.53
19
13.01
10
92
53.18
90
54.88
79
45.66
75
46.58
53
30.81
48
30.57
20
11.70
19
13.01
100
92
53.18
90
54.88
79
45.66
75
46.58
53
30.81
48
30.57
21
12.28
20
13.70
Co e age
173
100.00
164
94.80
173
100.00
161
93.06
172
99.42
157
90.75
171
98.84
146
84.39
Success Ra e TM-align 100% SI and 0.5 o TM-
sco e (TM-align)
Success Ra e TM-align 95% SI and 0.5 o
TM-sco e (TM-align)
Success Ra e TM-align 70% SI and 0.5 o
TM-sco e (TM-align)
Success Ra e TM-align 30% SI and 0.5
o TM-sco e (TM-align)
Top NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC TM-sco e
min
1
90
55.21
87
65.41
75
47.47
71
58.20
48
33.10
43
43.00
17
13.82
13
18.84
5
91
55.83
88
66.17
77
48.73
73
59.84
51
35.17
45
45.00
18
14.63
16
23.19
10
91
55.83
88
66.17
78
49.37
73
59.84
52
35.86
45
45.00
19
15.45
16
23.19
100
92
56.44
88
66.17
78
49.37
73
59.84
52
35.86
45
45.00
20
16.26
16
23.19
Co e age
163
94.22
133
76.88
158
91.33
122
70.52
145
83.82
100
57.80
123
71.10
69
39.88
Success Ra e TM-align 100% SI and 0.6 o TM-
sco e (TM-align)
Success Ra e TM-align 95% SI and 0.6 o
TM-sco e (TM-align)
Success Ra e TM-align 70% SI and 0.6 o
TM-sco e (TM-align)
Success Ra e TM-align 30% SI and 0.6
o TM-sco e (TM-align)
Top NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC TM-sco e
min
1
90
66.67
83
76.85
75
59.52
67
72.04
48
46.15
39
60.00
17
24.64
10
29.41
5
91
67.41
84
77.78
77
61.11
69
74.19
51
49.04
41
63.08
18
26.09
13
38.24
10
91
67.41
84
77.78
77
61.11
69
74.19
51
49.04
41
63.08
18
26.09
13
38.24
100
91
67.41
84
77.78
77
61.11
69
74.19
51
49.04
41
63.08
19
27.54
13
38.24
Co e age
135
78.03
108
62.43
126
72.83
93
53.76
104
60.12
65
37.57
69
39.88
34
19.65
Success Ra e TM-align 100% SI and 0.7 o TM-
sco e (TM-align)
Success Ra e TM-align 95% SI and 0.7 o
TM-sco e (TM-align)
Success Ra e TM-align 70% SI and 0.7 o
TM-sco e (TM-align)
Success Ra e TM-align 30% SI and 0.7
o TM-sco e (TM-align)
Top NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC TM-sco e
min
1
85
78.70
82
81.19
69
75.00
65
78.31
41
64.06
33
68.75
11
40.74
7
46.67
5
85
78.70
82
81.19
69
75.00
65
78.31
41
64.06
33
68.75
11
40.74
7
46.67
10
85
78.70
82
81.19
69
75.00
65
78.31
41
64.06
33
68.75
11
40.74
7
46.67
100
85
78.70
82
81.19
69
75.00
65
78.31
41
64.06
33
68.75
11
40.74
7
46.67
Co e age
108
62.43
101
58.38
92
53.18
83
47.98
64
36.99
48
27.75
27
15.61
15
8.67
Success Ra e TM-align 100% SI and 0.8 o TM-
sco e (TM-align)
Success Ra e TM-align 95% SI and 0.8 o
TM-sco e (TM-align)
Success Ra e TM-align 70% SI and 0.8 o
TM-sco e (TM-align)
Success Ra e TM-align 30% SI and 0.8
o TM-sco e (TM-align)
Top NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC
TM-
sco e
min
NSC TM-sco e
a e age NSC TM-sco e
min
1
82
82.83
81
84.38
66
81.48
62
83.78
34
72.34
27
75.00
5
50.00
5
62.50
5
82
82.83
81
84.38
66
81.48
62
83.78
34
72.34
27
75.00
5
50.00
5
62.50
10
82
82.83
81
84.38
66
81.48
62
83.78
34
72.34
27
75.00
5
50.00
5
62.50
100
82
82.83
81
84.38
66
81.48
62
83.78
34
72.34
27
75.00
5
50.00
5
62.50
Co e age
99
57.23
96
55.49
81
46.82
74
42.77
47
27.17
36
20.81
10
5.78
8
4.62
129
Table 10.2.3. Success a e o ab ini io modelling combining he TM-sco e(s) and pyDock sco ing unc ions.
Success Ra e MM-align 100% SI and 0.4 o TM-sco e (MM-align)
Success Ra e MM-align 70% SI and 0.4 o TM-sco e (MM-align)
Top NSC
TM-
sco e
a e age
NSC
TM-
sco e
min
NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min NSC TM-sco e
a e age NSC TM-sco e
min NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min
1 37 24.34 34 31.19 35 23.03 29 26.61
20 13.89 16 17.20 23 15.97 15 16.13
5
41
26.97
37
33.94
41
26.97
35
32.11
25
17.36
20
21.51
25
17.36
20
21.51
10
42
27.63
37
33.94
42
27.63
38
34.86
26
18.06
20
21.51
26
18.06
21
22.58
100 45 29.61 38 34.86 45 29.61 38 34.86
29 20.14 21 22.58 29 20.14 21 22.58
Co e age 152 86.36 109 61.93 152 86.36 109 61.93
144 81.82 93 52.84 144 81.82 93 52.84
Success Ra e MM-align 100% SI and 0.5 o TM-sco e (MM-align)
Success Ra e MM-align 70% SI and 0.5 o TM-sco e (MM-align)
Top NSC
TM-
sco e
a e age
NSC
TM-
sco e
min
NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min NSC TM-sco e
a e age NSC TM-sco e
min NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min
1
37
34.91
26
54.17
35
33.02
26
54.17
20
24.39
12
40.00
21
25.61
13
43.33
5 39 36.79 28 58.33 39 36.79 28 58.33
22 26.83 14 46.67 22 26.83 14 46.67
10
40
37.74
28
58.33
40
37.74
28
58.33
23
28.05
14
46.67
23
28.05
14
46.67
100
40
37.74
28
58.33
40
37.74
28
58.33
23
28.05
14
46.67
23
28.05
14
46.67
Co e age
106
60.23
48
27.27
106
60.23
48
27.27
82
46.59
30
17.05
82
46.59
30
17.05
Success Ra e MM-align 100% SI and 0.6 o TM-sco e (MM-align)
Success Ra e MM-align 70% SI and 0.6 o TM-sco e (MM-align)
Top NSC
TM-
sco e
a e age
NSC
TM-
sco e
min
NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min NSC TM-sco e
a e age NSC TM-sco e
min NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min
1 35 53.03 22 68.75 33 50.00 23 71.88
18 45.00 11 61.11 19 47.50 12 66.67
5
36
54.55
23
71.88
36
54.55
23
71.88
19
47.50
12
66.67
19
47.50
12
66.67
10
36
54.55
23
71.88
36
54.55
23
71.88
19
47.50
12
66.67
19
47.50
12
66.67
100 36 54.55 23 71.88 36 54.55 23 71.88
19 47.50 12 66.67 19 47.50 12 66.67
Co e age 66 37.50 32 18.18 66 37.50 32 18.18
40 22.73 18 10.23 40 22.73 18 10.23
Success Ra e MM-align 100% SI and 0.7 o TM-sco e (MM-align)
Success Ra e MM-align 70% SI and 0.7 o TM-sco e (MM-align)
Top NSC
TM-
sco e
a e age
NSC
TM-
sco e
min
NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min NSC TM-sco e
a e age NSC TM-sco e
min NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min
1
27
72.97
19
70.37
26
70.27
20
74.07
13
72.22
9
69.23
13
72.22
9
69.23
5 28 75.68 20 74.07 28 75.68 20 74.07
13 72.22 10 76.92 13 72.22 10 76.92
10
28
75.68
20
74.07
28
75.68
20
74.07
13
72.22
10
76.92
13
72.22
10
76.92
100
28
75.68
20
74.07
28
75.68
20
74.07
13
72.22
10
76.92
13
72.22
10
76.92
Co e age
37
21.02
27
15.34
37
21.02
27
15.34
18
10.23
13
7.39
18
10.23
13
7.39
Success Ra e MM-align 100% SI
Success Ra e MM-align 70% SI and 0.8 o TM-sco e (MM-align)
Top NSC
TM-
sco e
a e age
NSC
TM-
sco e
min
NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min NSC TM-sco e
a e age NSC TM-sco e
min NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min
1
20
76.92
15
83.33
20
76.92
16
88.89
10
83.33
5
71.43
10
83.33
6
85.71
5
20
76.92
16
88.89
20
76.92
16
88.89
10
83.33
6
85.71
10
83.33
6
85.71
10 20 76.92 16 88.89 20 76.92 16 88.89
10 83.33 6 85.71 10 83.33 6 85.71
100
20
76.92
16
88.89
20
76.92
16
88.89
10
83.33
6
85.71
10
83.33
6
85.71
Co e age 26 14.77 18 10.23 26 14.77 18 10.23
12 6.82 7 3.98 12 6.82 7 3.98
130
Success Ra e MM-align 95% SI and 0.4 o TM-sco e (MM-align)
Success Ra e MM-align 30% SI and 0.4 o TM-sco e (MM-align)
Top NSC
TM-
sco e
a e age
NSC
TM-
sco e
min
NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min NSC TM-sco e
a e age NSC TM-sco e
min NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min
1
30
19.87
27
25.23
30
19.87
22
20.56
7
4.43
2
33.33
8
5.06
2
33.33
5 35 23.18 31 28.97 35 23.18 29 27.10
21 13.29 2 33.33 17 10.76 2 33.33
10
36
23.84
31
28.97
36
23.84
32
29.91
21
13.29
2
33.33
25
15.82
2
33.33
100
39
25.83
32
29.91
39
25.83
32
29.91
35
22.15
2
33.33
35
22.15
2
33.33
Co e age
151
85.80
107
60.80
151
85.80
107
60.80
158
89.77
6
3.41
158
89.77
6
3.41
Success Ra e MM-align 95% SI and 0.5 o TM-sco e (MM-align)
Success Ra e MM-align 30% SI and 0.5 o TM-sco e (MM-align)
Top NSC
TM-
sco e
a e age
NSC
TM-
sco e
min
NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min NSC TM-sco e
a e age NSC TM-sco e
min NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min
1 30 30.30 19 46.34 29 29.29 19 46.34
4 5.88 2 50 5 7.35 2 50
5
32
32.32
21
51.22
32
32.32
21
51.22
12
17.65
2
50
11
16.18
2
50
10
33
33.33
21
51.22
33
33.33
21
51.22
12
17.65
2
50
12
17.65
2
50
100 33 33.33 21 51.22 33 33.33 21 51.22
15 22.06 2 50 15 22.06 2 50
Co e age 99 56.25 41 23.30 99 56.25 41 23.30
68 38.64 4 2.272727 68 38.64 4 2.272727273
Success Ra e MM-align 95% SI and 0.6 o TM-sco e (MM-align)
Success Ra e MM-align 30% SI and 0.6 o TM-sco e (MM-align)
Top NSC
TM-
sco e
a e age
NSC TM-sco e
min NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min NSC TM-sco e
a e age NSC TM-sco e
min NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min
1
28
48.28
15
62.50
27
46.55
16
66.67
2
66.67
2
100
2
66.67
2
100.00
5 29 50.00 16 66.67 29 50.00 16 66.67
2 66.67 2 100 2 66.67 2 100.00
10
29
50.00
16
66.67
29
50.00
16
66.67
2
66.67
2
100
2
66.67
2
100.00
100
29
50.00
16
66.67
29
50.00
16
66.67
2
66.67
2
100
2
66.67
2
100.00
Co e age
58
32.95
24
13.64
58
32.95
24
13.64
3
1.70
2
1.136364
3
1.70
2
1.14
Success Ra e MM-align 95% SI and 0.7 o TM-sco e (MM-align)
Success Ra e MM-align 30% SI and 0.7 o TM-sco e (MM-align)
Top NSC
TM-
sco e
a e age
NSC TM-sco e
min NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min NSC TM-sco e
a e age NSC TM-sco e
min NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min
1
21
72.41
13
72.22
20
68.97
14
77.78
2
100.00
1
100.00
2
100
1
100.00
5
21
72.41
14
77.78
21
72.41
14
77.78
2
100.00
1
100.00
2
100
1
100.00
10 21 72.41 14 77.78 21 72.41 14 77.78
2 100.00 1 100.00 2 100 1 100.00
100
21
72.41
14
77.78
21
72.41
14
77.78
2
100.00
1
100.00
2
100
1
100.00
Co e age 29 16.48 18 10.23 29 16.48 18 10.23
2 1.14 1 0.57 2 1.136363636 1 0.57
Success Ra e MM-align 95% SI and 0.8 o TM-sco e (MM-align)
Success Ra e MM-align 30% SI and 0.8 o TM-sco e (MM-align)
Top NSC
TM-
sco e
a e age
NSC TM-sco e
min NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min NSC TM-sco e
a e age NSC TM-sco e
min NSC pyDock+TM-
sco e a e age NSC pyDock+TM-
sco e min
1
14
77.78
10
76.92
14
77.78
11
84.62
1
100.00
0
0
1
100.00
0
0
5
14
77.78
11
84.62
14
77.78
11
84.62
1
100.00
0
0
1
100.00
0
0
10
14
77.78
11
84.62
14
77.78
11
84.62
1
100.00
0
0
1
100.00
0
0
100
14
77.78
11
84.62
14
77.78
11
84.62
1
100.00
0
0
1
100.00
0
0
Co e age
18
10.23
13
7.39
18
10.23
13
7.39
1
0.57
0
0
1
0.57
0
0
131
Table 10.2.4. S uc u al a ailabili y o he in e ac ing molecules and addi ional in o ma ion o he
p epa a ion o he submi ed models as se e s and p edic o s.
a Unde sco ed: a ge o special di icul y, wi h only 3 o ewe g oups ha submi ed co ec models wi hin hei op 5 submi ed ones.
b PDB code (and chain ID i needed) o he unbound s uc u e (o ha bound o a di e en pa ne in he case o he saccha ides) used o docking. In acke s,
in e ace RMSD wi h espec o he bound s uc u e, calcula ed on all he a oms om in e ace esidues. In e ace esidues a e hose wi h a leas one a om
wi hin 8 Å dis ance om any a om o he pa ne molecule, a e supe imposing he unbound s uc u e on o he co esponding bound s uc u e in he
complex. "N/A": i e e ence complex s uc u e is no ye a ailable. PubChem CID code gi en in case o non-pep idic ligands. cPDB code (and chain ID i
needed) o he empla e used o homology-based modeling o he ecep o o ligand, o empla e-based docking o he complex (o pa o i ), as indica ed.
In some a ge s, I-TASSER was applied, using mul iple empla es, as indica ed. In b acke s, sequence iden i y, and global RMSD o he Cα a oms o he
empla e wi h espec o he bound s uc u e ("N/A": i e e ence complex s uc u e is no ye a ailable). d Use o dis ance es ain s o docking based on
in e ace esidues, ei he om a ailable empla e o he complex, o om li e a u e, as indica ed. e PDB code o he complex e e ence, i i is now a ailable
Ta ge a
RECEPTOR LIGAND COMPLEX
Unbound
S uc u e
(in RMSD)
b
Templa e
(SI, RMSD)C
Unbound
S uc u e
(in RMSD)
b
Templa e
(SI, RMSD)C
Templa e
(SI, RMSD)C
Docking
es ain sd
Ta ge
s uc u ee
P o ein-p o ein
T122
5MXA
(3.4 A)
- - 111R (25%, 10.4 A) - Li e a u e 5MZV
T123
5LMW
(N/A)
5BOZ
(76%, N/A)
-
mul i- empla e l-TASSER
(N/A, N/A)
- - 6EY0
T124
5FWO
(2.2 A)
4GRW:H
(78%, 3.0 A)
-
mul i- empla e l-TASSER
(N/A, N/A)
Dime : 3MTR
(17%, 21.9 A)
- 6EY6
T125
he e o
4QKH
(1.4 A)
- - 3T3A (47%, 10.6 A) - Li e a u e 5MGT
T125
homo
4QKH
(1.4 A)
- 4QKH (1.3 A) -
4QKH (100%,
0.7 Å)
- 5MGT
T131
4WHD
(1.4 Å)
- - 5LP2 (96%,0.9 Å) - Li e a u e 6GBG
132
4WHD
(1.0 Å)
- - 5LP2 (57%,2.9 Å) - Li e a u e 6GBH
T133 -
3U43:B
(87%, 0.8 Å)
- 3U43:A (83%,1.5 Å) - - 6ERE
T136 -
5FKZ:E
(45%, N/A),
mul i- empla e l-
TASSER (N/A)
-
5FKZ:E (45%, N/A), mul i-
empla e l-TASSER (N/A,
N/A)
2VYC (40%,
N/A), 5FKZ
(45%, N/A)
5FKZ
in e ace 6Q6I
P o ein-pep ide
T121 1LR0
(N/A) - -
mul i- empla e l-TASSER
(N/A,N/A), 2HQS (42%,
N/A), 2W8B (38%, N/A)
- 1TOL
in e ace 6S3W
T134 1F3C
(2.0 Å) - -
1F95 (33%,1.2 Å), 1F96
(9%, 2.4 Å), 3P8M (36%,
2.3 Å)
- 1F95
in e ace 6GZL
T135
1F3C
(2.0 Å)
- - 1F95 (33%, 1.2 Å) -
1F95
in e ace
6GZL
P o ein-saccha ide
T126 - 5F7V (30%, N/A)
2C7F.2X8S, 3QEF,
CID 53356682
(N/A)
- - - 6RKH
T127 - 5F7V (30%, N/A)
CID 74539968,
(N/A)
- - - 6RKX
T128 - 5F7V (30%, N/A)
5HOF, CID
74539969 (N/A)
- - - 6RL2
T129 - 5F7V (30%, N/A)
1GYE, CID
74539970 (N/A)
- - - 6RL1
T130
3D5Z
(0.7 Å)
-
5HOF (2.6 Å), CID
74539969 (2.3 Å)
- - - 6F1G
132
Table 10.2.5. In o ma ion on he a ge s o he 7 h CAPRI
CAPRI 7 h expe imen
Ta ge
Le el
S oich.
#In
#Res
PDB
Desc ip ion
P o ein-p o ein complexes
T122 Di icul
A1B1C1
1
198/328/330
5MZV
Human cy okine he e o-dime / ecep o complex
IL23/IL23R
T123
Di icul
A1B1
1
174/121
6EY0
Po M-N /nb(02)
T124
Di icul
A2B1
1
202/141
6EY6
Po M-C /nb(130)
T125
Di icul
A2B4
5
135/146
5MGT
He e o-hexame o LLT1/NKR-P1 (ex a-cellula
domains)
T131
Di icul
A1B1
1
108/404
6BGB
Human CEACAM1/HopQ-Type-I H. pylo i
T132
Medium
A1B1
1
108/418
6BGH
Human CEACAM1/hopQ-Type-II H. pylo i
T133
Easy
A1B1
1
69/95
6ERE
Redesigned Colicin E2 DNase/Im2 complex
T136
Easy
A10
3
751
6Q6I
LdcA P.ae oginosa; EM
P o ein-pep ide complexes
T121 Di icul A1B1 1 115/13
6S3W
P.ae oginosa TolAIII domain/N- e minus
P.ae uginosa TolB
T134
Easy
A2B1
1
88/50
6GZJ
DLC8 dime /MAG 50- esidue agmen
T135
Easy
A2B1
1
88/12
6GZL
DLC8 dime (Ra )/MAG 12- esidue agmen
P o ein-oligosaccha ide complexes
T126 Di icul
A1B1
1 415/6
6RKH
A abino-oligosaccha ide binding p o ein,
G.s ea o he mophilus, wi h AbnE/A6
T127
Di icul
A1B1
1
415/5
6RKX
Idem wi h AbnE/A5
T128
Medium
A1B1
1
415/4
6RL2
Idem wi h AbnE/A4
T129
Medium
A1B1
1
415/3
6RL1
Idem wi h AbnE/A3
T130 Easy
A1B1
1 315/5
6F1G
A abino-oligosaccha ide binding p o ein,
G.s ea o he mophilus, wi h AbnB/A5
No e: The a ge lis is subdi ided in o ca ego ies: p o ein-p o ein (R39:T122-T124; R40:T125; R42:T131, T132; R43:T133; R45:T136), p o ein-pep ide
(R38:T121; T44:T134, T135), and p o ein
-polysaccha ide (R41:T126-
T130). Columns 1 o 3 lis he CAPRI a ge ID, i s di icul y, and he numbe o assessed
in e aces; columns 4 o 6 lis he numbe o esidues, and he PDB ID; column 7 con ains a ex ual desc ip ion o he a ge .
133
Table 10.2.6. O e all 7 h CAPRI pe o mance anking o he p o ein-p o ein a ge s. Ex ac ed
om [204]
op1
op5
op 10
Rank
G oup name
# a ge s
ank
pe o mance1
ank
pe o mance1
ank
pe o mance1
1
Kozako /Vajda
11
2
5/1***/3**
1
6/1***/5**
1
6/1***/5**
2
Venclo as
8
1
5/2***/3**
2
5/2***/3**
2
5/2***/3**
3
Seok
11
6
4/3**
3
5/1***/4**
3
5/1***/4**
3
Pie ce
9
3
4/2***/1**
3
5/2***/2**
3
5/2***/2**
5
And eani/Gue ois
11
6
4/3**
5
5/1***/3**
5
5/1***/3**
6
Zou
11
6
3/1***/2**
6
4/1***/3**
6
4/1***/3**
6
Zacha ias
11
14
4/1***
6
5/1***/2**
6
5/1***/2**
6
Kiha a
11
6
4/1***/1**
6
5/1***/2**
6
5/1***/2**
6
G ay
11
3
5/1***/2**
6
5/1***/2**
6
5/1***/2**
10
Shen
11
14
3/1***/1**
10
4/1***/2**
6
5/1***/2**
10
Moal
8
24
2**
10
4/1***/2**
11
4/1***/2**
10
MDOCKPP
11
5
4/1***/2**
10
4/1***/2**
11
4/1***/2**
10
HADDOCK
11
6
3/1***/2**
10
4/1***/2**
11
4/1***/2**
14
G udinin
11
6
3/1***/2**
14
3/1***/2**
16
3/1***/2**
14
Fe nandez-Recio
11
6
3/1***/2**
14
3/1***/2**
16
3/1***/2**
14
CLUSPRO
11
20
3/2**
14
4/3**
16
4/3**
14
Bon in
11
6
3/1***/2**
14
3/1***/2**
16
3/1***/2**
18
Weng
11
14
3/1***/1**
18
3/1***/1**
20
3/1***/1**
18
SWARMDOCK
11
14
3/1***/1**
18
3/1***/1**
20
3/1***/1**
18
LZERD
11
14
3**
18
3**
20
3**
18
Chang
11
20
2/1***/1**
18
3/1***/1**
11
4/1***/2**
18
Ba es
11
14
3/1***/1**
18
3/1***/1**
11
4/1***/2**
23
Vakse
11
24
2**
23
2/1***/1**
24
2/1***/1**
23
Takeda-Shi aka
8
20
2/1***/1**
23
2/1***/1**
24
2/1***/1**
23
Huang
9
20
2/1***/1**
23
2/1***/1**
20
3/1***/1**
26
Wol son
7
24
2/1***
26
2/1***
27
2/1***
26
PYDOCKWEB
11
28
1***
26
2/1***
24
3/1***
26
HDOCK
8
30
1**
26
2**
27
2**
26
GALAXYPPDOCK
8
24
2**
26
2**
27
2**
30
Laine
3
28
1***
30
1***
30
1***
31
Ri chie
5
30
1**
31
1**
32
1**
31
Iwada e
1
30
1**
31
1**
32
1**
31
INTERPRED
1
30
1**
31
1**
32
1**
31
Del Ca pio
11
30
1**
31
1**
32
1**
31
Ca bone
6
36
0
31
1**
30
2/1**
31
B ini
1
30
1**
31
1**
32
1**
37
ZDOCK
1
36
0
37
0
37
0
37
Wang
0
36
0
37
0
37
0
37
Wallne
4
36
0
37
0
37
0
37
UUcou se
0
36
0
37
0
37
0
37
Tu e y
0
36
0
37
0
37
0
37
Schuele -Fu man
0
36
0
37
0
37
0
37
Schneidman
1
36
0
37
0
37
0
37
Sanne
0
36
0
37
0
37
0
37
Ni
0
36
0
37
0
37
0
37
Negi
4
36
0
37
0
37
0
37
GRAMM-X
3
36
0
37
0
37
0
37
Gong
0
36
0
37
0
37
0
37
Czaplewski
1
36
0
37
0
37
0
37
Ca azo
1
36
0
37
0
37
0
Columns 1 o 3 lis he ank, name g oup and he numbe o assessed in e aces; columns 4 o 6 lis he ank, he quali y o he
models o he op 1, 5 and 10 uploaded models submi ed.
1
This column shows he success o he g oups, he i s numbe
co esponds o he numbe o success ul a ge s, hen sepa a ed by a slash, i necessa y, he numbe and quali y o he models
is displayed. Fo example, 5/1***/3**, means 5 success ul a ge s, o which he e is one a ge wi h high
-
quali y models (***),
3 o medium-quali y models (**) and he emaining one is o accep able quali y.
134
Table 10.2.7. O e all 7 h CAPRI pe o mance anking o he p o ein-pep ide a ge s. Ex ac ed
om [204]
op1
op 5
op 10
Rank
G oup name
# a ge s
ank
pe o mance
ank
pe o mance
ank
pe o mance
1
Zacha ias
3
1
2/1***/1**
1
2***
2
2***
1
Schuele -Fu man
3
1
2/1***/1**
1
2***
1
3/2***/1**
1
And eani/Gue ois
3
1
2/1***/1**
1
2***
2
2***
4
Venclo as
3
11
1
4
2/1***/1**
2
2***
4
Seok
3
7
2/1**
4
3/2**
5
3/2**
4
Moal
3
1
2/1***/1**
4
2/1***/1**
5
2/1***/1**
4
Huang
3
1
2/1***/1**
4
2/1***/1**
5
2/1***/1**
8
HDOCK
2
6
2/1***
8
2/1***
8
2/1***
9
Zou
2
11
1
9
2
12
2
9
Shen
3
8
1**
9
1**
12
1**
9
Kozako /Vajda
2
11
1
9
1**
8
2**
9
GALAXYPPDOCK
3
21
0
9
1**
12
1**
9
Fe nandez-Recio
3
11
1
9
2
12
2
9
CLUSPRO
2
11
1
9
1**
12
1**
9
B ini
2
8
1**
9
1**
12
1**
9
Bon in
3
8
1**
9
1**
8
2**
17
UUcou se
1
11
1
17
1
12
1**
17
SWARMDOCK
3
21
0
17
1
12
2
17
PYDOCKWEB
2
21
0
17
1
23
1
17
Kiha a
3
11
1
17
1
12
1**
17
HADDOCK
3
21
0
17
1
12
2
17
Del Ca pio
2
11
1
17
1
23
1
17
Czaplewski
2
11
1
17
1
12
2
17
Ba es
3
11
1
17
1
11
3
25
Wang
1
21
0
25
0
25
0
25
Wallne
1
21
0
25
0
25
0
25
Vakse
2
21
0
25
0
25
0
25
Tu e y
1
21
0
25
0
25
0
25
Takeda-Shi aka
1
21
0
25
0
25
0
25
Sanne
1
21
0
25
0
25
0
25
Ri chie
1
21
0
25
0
25
0
25
Ni
1
21
0
25
0
25
0
25
Negi
1
21
0
25
0
25
0
25
MDOCKPP
2
21
0
25
0
25
0
25
LZERD
3
21
0
25
0
25
0
25
G udinin
3
21
0
25
0
25
0
25
Gong
1
21
0
25
0
25
0
25
Chang
3
21
0
25
0
25
0
Columns 1 o 3 lis he ank, name g oup and he numbe o assessed in e aces; columns 4 o 6 lis he ank, he quali y o he models
o he op 1, 5 and 10 uploaded models submi ed. 1This column shows he success o he g oups, he i s numbe co esponds
o
he numbe o success ul a ge s, hen sepa a ed by a slash, i necessa y, he numbe and quali y o he models is displayed. Fo
example, 5/1***/3**, means 5 success ul a ge s, o which he e is one a ge wi h high
-quali y models (***), 3 o medium-
quali y
models (**) and he emaining one is o accep able quali y.
135
Table 10.2.8. O e all 7 h CAPRI pe o mance anking o he p o ein-oligosaccha ide a ge s.
Ex ac ed om [204]
op1
op 5
op 10
Rank
G oup name
# a ge s
ank
pe o mance
ank
pe o mance
ank
pe o mance
1
And eani/Gue ois
5
4
4/1**
1
5/1***/2**
1
5/1***/3**
2
Seok
5
1
5/1***/1**
2
5/1***/1**
2
5/1***/2**
2
LZERD
5
4
5
2
5/1***/1**
2
5/1***/2**
2
Chang
5
2
4/1***
2
5/1***/1**
4
5/1***/1**
5
Kozako /Vajda
5
9
3/1**
5
5/2**
8
5/2**
5
Huang
5
9
4
5
5/2**
8
5/2**
5
HDOCK
5
24
1
5
4/3**
4
5/3**
5
CLUSPRO
5
24
1
5
5/2**
8
5/2**
9
Zou
5
9
4
9
5/1**
14
5/1**
9
Zacha ias
5
4
5
9
5/1**
14
5/1**
9
Venclo as
5
4
3/1***
9
4/1***
14
4/1***
9
Takeda-Shi aka
5
9
4
9
5/1**
14
5/1**
9
MDOCKPP
5
14
3
9
5/1**
14
5/1**
9
Kiha a
5
9
4
9
4/2**
4
5/3**
9
G udinin
5
2
4/1***
9
4/1***
8
5/1***
9
G ay
5
24
1
9
4/2**
14
4/2**
9
Fe nandez-Recio
5
19
2
9
5/1**
8
5/2**
9
Bon in
5
14
2/1**
9
3/1***/1**
8
4/1***/1**
19
Shen
5
19
2
19
5
14
5/1**
19
HADDOCK
5
14
2/1**
19
3/1***
4
5/1***/1**
19
Ca bone
4
4
3/1***
19
3/1***
22
3/1***
22
Vakse
4
19
2
22
4
24
4
22
Moal
5
30
0
22
4
22
5
22
GALAXYPPDOCK
5
14
3
22
3/1**
14
4/1***
25
Pie ce
2
14
2/1**
25
2/1**
25
2/1**
25
Ba es
5
19
2
25
3
25
3
27
SWARMDOCK
5
24
1
27
2
28
2
27
Negi
4
19
2
27
2
28
2
29
Ri chie
1
24
1
29
1
28
1**
29
Del Ca pio
5
24
1
29
1
25
3
Columns 1 o 3 lis he ank, name g oup and he numbe o assessed in e aces; columns 4 o 6 lis he ank, he quali y o
he models o he op 1, 5 and 10
uploaded models submi ed. 1This column shows he success o he g oups, he i s numbe
co esponds o he numbe o success ul a ge s, hen sepa a ed by a slash, i necessa y, he numbe and quali y o he model
s
is displayed. Fo example, 5/1***/3**,
means 5 success ul a ge s, o which he e is one a ge wi h high-quali y models
(***), 3 o medium-quali y models (**) and he emaining one is o accep able quali y.
136
Table 10.2.9. In o ma ion on he a ge s o he hi d and ou h CAPRI-CASP expe imen s
3 d Join CAPRI-CASP expe imen
Easy a ge s
CASP ID
S oich.
#In 1
#Res2
PDB3
Desc ip ion
T140
T0973
A2
1
146
6YFN
Bac e iophage ESE058 coa p o ein
T143
T0983
A2
1
245
6UK5
Cals10 p o ein
T144
T0984
A2
1
752
6NQ1
Two-po e calcium channel p o ein; EM
T152
T1003
A2
1
474
6HRH
ALAS2, 50-Aminole ulina e syn hase 2
T153
T1006
A2
1
79
6QEK
Pu a i e memb ane anspo e (C. desul amplus)
T147
T0995
A2/A4/A8
3
330
-
Cyanide dihyd a ase (B. pumilus); EM
T158
T1020
A3
1
577
7WNQ
SLAC1 p o ein
T139
T0961
A4
2
505
6SD8
Acyl-CoA dehyd ogenase om Bdello ib io
bac e io o us
T142
H0974
A1B1
1
70/80
6TRI
Rep esso -an i ep esso complex (lysogeny swi ch)
Di icul a ge s
CASP ID
S oich.
#In 1
#Res2
PDB3
Desc ip ion
T137
T0965
A2
2
326
6D2V
NADP-dependen educ ase
T138
T0966
A2
2
494
5W6L
RasRap1 si e-speci ic endopep idase
T141
T0976
A2
1
252
6MXV
Rhodanese-like amily p o ein, bac e ia
T148
T0997
A2
1
228
-
LD- anspep idase
T149 T0999 A2 5 1589
6HQV
Pen a unc ional AROM polypep ide: i e main
enzymes o he shikima e pa hway
T150
T0999
A2
5
1589
6HQV
Idem; wi h SAXS da a
T151
T0999
A2
5
1589
6HQV
Idem; wi h c osslinking da a
T154
T1009
A2
1
718
6DRU
Alpha-xylosidase
T155
H1015
A1B1
1
89/129
-
CDI_213 p o ein, bac e ia
T156
H1017
A1B1
1
111/129
-
201_INDD4 p o ein, E. coli
T157
H1019
A1B1
1
58/88
-
CDI207 p o ein, E. coli
T146
H0993
A2B2
3
275/112
-
Lipid- anspo , bac e ial ou e memb ane
T159
H1021
A6B6C6
7
148/351/295
6RAP
18-me he e ocomplex; EM
4 h Join CAPRI-CASP expe imen
Easy a ge s
CASP ID
S oich.
#In
1
#Res
2
PDB
3
Desc ip ion
T164
T1032
A2
1
284
6n64
SMCHD1 (human) esidues 1616–1899
T166
H1045
A1B1
1
157/173
6xod
PEX4/PEX22 complex om A abidopsis haliana
T168
T1052
A3
1
832
-
Tail ibe o he Salmonella i us epsilon15
T177 H1081 A20 3
758
7pk6
2 yc
A ginine deca boxylase/bac e ia
Di icul a ge s
CASP ID
S oich.
#In 1
#Res2
PDB3
Desc ip ion
T169 T1054 A2 1
190
6 4
Ou e -memb ane lipop o ein om Acine obac e
baumannii
T176 T1078 A2 1
138
7cwp
Tsp1 om T ichode ma i ens, small sec e ed
cys eine ich p o ein (SSCRP)
T178
T1083
A2
1
98
6nq1
Helical segmen om Ni osococcus oceani
T179
T1087
A2
1
93
-
Helical segmen om Me hylobac e und ipaludum
T165
H1036
A3H3L3 1
931/128/106
6 n1
MC Ab 93 k bound o a icella-zos e i us
glycop o ein gB
T174
T1070
A3
1
335
7 ej
P o ein o a achmen egion o phage ail
T170
H1060
A6/B3/C12/D6
9
464/298/140/142
2 yc
Componen o he T5 phage ail dis al complex
T180
T1099
A240
8 (4)
262
6yhg
Capsid o duck hepa i is B i us