scieee AI-readable full text Open interactive document viewer

New computational methods for structural modeling and flexible docking of protein-protein and protein-nucleic acids interactions

Rodríguez-Lumbreras, Luis A.

Abstract

Trabajo presentado para lograr el titulo de Doctor por la Universidad de Barcelona, Facultad de Biología, Departamento de Bioquímica y Biología Molecular

Full text

 New computational methods for structural modeling protein-protein and protein-nucleic acid interactions Luis Ángel Rodríguez Lumbreras Aquesta tesi doctoral està subjecta a la llicència ReconeixementCompartIgual 4.0. Espanya de Creative Commons. Esta tesis doctoral está sujeta a la licencia Reconocimiento - CompartirIgual 4.0. España de Creative Commons. This doctoral thesis is licensed under the Creative Commons Attribution-ShareAlike 4.0. Spain License. FACULTAT DE BIOLOGIA (Supervisor: Juan Fernández Recio) New computational methods for structural modeling of proteinprotein and protein-nucleic acid interactions LUIS ÁNGEL RODRÍGUEZ LUMBRERAS UNIVERSITAT DE BARCELONA FACULTAT DE BIOLOGIA DOCTORAT EN BIOMEDICINA Código HDK05 New computational methods for structural modeling of protein-protein and protein-nucleic acid interactions Memoria presentada por Luis Ángel Rodríguez Lumbreras para optar al título de Doctor por la Universidad de Barcelona Tesis doctoral realizada bajo la dirección del Doctor Juan Fernández Recio, en el Barcelona Supercomputing Center (BSC) y los dos últimos años en el Instituto de Ciencias de la Vid y del Vino (ICVV-CSIC). Tesis adscrita al Departamento de Bioquímica y Biología Molecular de la Facultad de Biología (Universitat de Barcelona). Director Tutor Ph.D. candidate Dr. Juan Fernández Recio Dr. Josep Lluís Gelpí Buchaca Luis Ángel Rodríguez Lumbreras Barcelona, 2022 i DECLARACIÓN DE ORIGINALIDAD Yo, LUIS ANGEL RODRIGUEZ LUMBRERAS, matriculado en el programa de doctorado de BIOMEDICINA de la Universidad de Barcelona declaro que la tesis titulada “New computational methods for structural modeling of protein-protein and protein-nucleic acid interactions” es original, que la investigación que voy a realizar cumple con los códigos éticos y las buenas prácticas y que la tesis no incluye plagio. Soy consciente y acepto por escrito que mi tesis se someterá al proceso adecuado para probar la originalidad de mis resultados. 31 de Octubre de 2022 ii A mis padres iii "Ninguna ciencia, en cuanto a ciencia, engaña; el engaño está en quien no la sabe." Miguel de Cervantes x 6.2.2. Predictive success of template-based docking ............................................................... 70 6.3. Ab initio docking can improve template identification .......................................................... 71 6.3.1. A new protocol for combining ab initio and template-based docking ........................... 71 6.3.2. Evaluation of the combined docking protocol. ............................................................... 76 6.4. Docking functions can be used to score models from template-based docking ................... 78 6.4.1. Automatic processing of input structures ....................................................................... 78 6.4.2. Template-base docking ................................................................................................... 79 6.4.3. Ab initio docking .............................................................................................................. 79 6.4.4. Combining ab initio and template-based methods......................................................... 80 6.4.5. Inclusion of restraints from available external data ....................................................... 82 6.4.6. Final selection of the models. ......................................................................................... 84 6.4.7. Scoring of protein-saccharide complex models .............................................................. 86 6.5. Evaluation of the predictive results in CAPRI and CASP ......................................................... 87 6.5.1. 7th CAPRI ......................................................................................................................... 88 6.5.2. 3rd joint CASP-CAPRI experiment..................................................................................... 91 6.5.3. 4th joint CASP-CAPRI experiment ................................................................................... 94 6.6. Protein docking functions for the identification of physiological homodimers..................... 96 6.6.1. Applicability of pyDock and CCharPPI scoring functions to the identification of physiological homodimers. ....................................................................................................... 98 7. APPLICATION TO A CASE STUDIO: MODELING ELECTRON TRANSFER PROTEIN COMPLEXES. ... 101 7.1. Introduction ......................................................................................................................... 103 7.2. Molecular structures and modelling .................................................................................... 103 7.3. Protein-protein docking simulations .................................................................................... 104 7.3.1. Docking sampling and scoring ....................................................................................... 104 7.3.2. Minimization of the protein-protein docking poses ..................................................... 109 7.4. Docking structures of [Cf:Cc6] and [Cf:Pc] in red Algae lineage ........................................... 110 7.6. Conclusions .......................................................................................................................... 115 8. General discussion....................................................................................................................... 117 9. Conclusions.................................................................................................................................. 121 10. Appendices ................................................................................................................................ 123 10.1. Appendix 1. Supplementary material for Chapter 5 .......................................................... 125 10.2. Appendix 2. Supplementary material for Chapter 6 .......................................................... 127 10.3. Appendix 3. Supplementary material for Chapter 7 .......................................................... 140 Bibliography .................................................................................................................................... 142 xi 1 1. INTRODUCTION 2 3 1.1. Biomolecules 1.1.1. From genes to proteins Living organisms can be seen as data storage machines that fight against entropy to conserve and transmit information. As a consequence, the classic functions that define a living organism, i.e. nutrition, relationship, and reproduction, emerge from this struggle. Therefore, the molecule responsible for storing the information must be stable and have a large storage capacity. Chemical and biological evolution selected deoxyribonucleic acid (DNA) as such molecule [1]. It is a polymer formed by deoxyribonucleotides. The deoxyribonucleotides are composed of a nitrogenous base, a deoxyribose (sugar), and a phosphate group. The nitrogenous base can be derived from purine, such as adenine (A) and guanine (G), or from pyrimidine, such as cytosine (C) and thymine (T). Phosphate groups form the backbone of DNA, binding the nucleotides together and giving directionality to the polymeric chain. The oxygen in the phosphate is covalently attached to the C5' position of the sugar. The phosphate group then binds to the next nucleotide at the sugar's C3' position, giving a 5' to 3' directionality and forming a single-stranded polymer. However, this polymeric molecule alone would not be stable enough for long time information storing. This stability is achieved by non-covalent binding of two singlestranded DNA polymers, which can adopt a stable double helix conformation, according to a highly specific base complementarity. The bases A and T must face each other by forming two hydrogen bonds, while G and C bases face each other by forming three hydrogen bonds. The repetition of these base pairs builds a double-stranded DNA. The information is encoded in triplets (three consecutive nucleotides of the sequence) or codons, which code up to 20 types of amino acids. Amino acids form a different kind of polymer, proteins 1.1.2. DNA structure The combination of Chargaff's rules (A+G=T+C) and X-ray study of the sodium salt of DNA of Rosalind Franklin, lead James Watson and Francis Crick to discover the DNA double helix structure [2]. As can be seen in Figure 1.1, the double helix is formed by two chains in an antiparallel arrangement. In addition, it is possible to observe the structure of the major 4 and minor grooves. The major groove is essential in proteins-DNA interactions (see Chapter 1.3.2). Figure 1.1. Left panel shows the double helix structure of DNA. Center panel shows the hydrogen bonds between paired bases. Right panel shows the major and minor grooves, which are binding sites for DNA binding proteins during transcription (copying RNA from DNA) and replication. Image modified from "DNA structure and sequencing:" by OpenStax College, Biology (CC BY 3.0). In addition to this canonical helical structure (B-DNA), the DNA, can present other configurations so called A-DNA and Z-DNA. The helix sense of the A-DNA is right-handed like the B-DNA. The A-DNA is generally wider and shorter with a narrow and deep major groove, as opposed to the minor groove, and like the B-DNA, the N-glycosidic linkage is anti. On the other hand, the Z-DNA has an entirely different structure, starting with a left-handed helix sense. This molecule is generally the narrowest and longest of the three mentioned conformations, where the major groove is shallow, and the minor groove is narrow and deep (see Figure 1.2). 5 Figure 1.2. Side and top view of A-, B-, and Z-DNA conformations. By Mauroesguerroto - Own work, CC BY-SA 4.0, https://commons.wikimedia.org/w/index.php?curid=35919357 DNA can have alternative structures, such as DNA Bubble, slipped loop, cruciforms, H-DNA, and G-quadruplex/i-motif in double-stranded DNA. All of these structures are important, as the DNA conformation defines the type of protein interactions in which it is involved. 1.1.3. Protein structure There are four levels of structural complexity in proteins. Primary structure refers to the sequence of amino acids that form a polypeptidic chain, defined from the N-terminal to the C-terminal. This is the order in which the ribosomes form the peptide bonds. The amino acid sequence depends on the DNA sequence codons that form the specific gen. The secondary structure of proteins is formed by local interactions between amino acid residues that are nearly located in the polypeptide chain. Hydrogen bonds between the backbone NH and CO groups stabilise those local folding events. This leads to the two most common types of secondary structure: alpha helices and beta sheets. Such periodic structures were predicted by Linus Pauling and Robert Carey in 1951, years before they were experimentally confirmed [3]. 6 Tertiary structure is formed by the interaction between different secondary structure elements in the 3D space, and it is stabilized by hydrogen bonds, van der Waals, salt bridges, hydrophobic packing, and disulphide bonding. The quaternary structure of proteins can be defined as the interaction between different polypeptides that have a tertiary structure and interact with each other. In general, the sequence of amino acids of a protein determines its 3D structure. While the majority of proteins have tertiary structure, there are many proteins that are partially or even totally unstructured [4]. In many cases, disordered regions can become structured when interacting with another protein or biomolecule [5]. 1.2. Protein-protein complexes 1.2.1. Importance of protein interactions Protein-protein interactions mediate most cellular functions. In this regard, the interactome is the complex and dynamic network formed by the entire set of protein interactions in a living organism or a cell. Mapping the interactome would provide insights into how biological systems work, and how they could be modulated. Recent studies have made progress towards the description of this intricate network of interactions using proteomewide techniques [6, 7], although it is estimated that the real number of protein-protein interactions are much larger than currently known, between 130,000 and 600,000 [8, 9]. All data from these experiments have been collected in several IPP databases such as: DIP [10], MPIDB [11], HPRD [12], BIND [13], MINT [14] , MatrixDB [15], STRINGS v10[16], BioGRID [17], Reactome[18], etc. In addition to these databases, there is Interactome 3D and MIntAct project [19], which try to collet several databases centralising the information in one place. However, although these initiatives are systematically extending the number of known interacting protein pairs, they have difficulty providing the structural architecture of these interactions, which is essential to unveil the underlying molecular mechanisms that maintain health or contribute to disease. Such structural knowledge may ultimately lead to 7 the rational design of new drugs, and the interpretation and design of relevant mutations for biomedical or biotechnological purposes. 1.2.2. Structural characterization of protein-protein complexes Biophysical methods, mainly X-ray crystallography [20], NMR spectroscopy [21] and cryoelectron microscopy (cryo-EM) [22, 23], can be used to solve protein–protein interactions at atomic resolution. These methods have determined de 3D structure of a large number of complexes, so many studies have been reported to analyze unique characteristics of protein-protein interactions. But it has been difficult to find general structural and energetics features for all protein-protein complexes. It has been found that in stable complexes, binding are driven by hydrophobic interactions, as compared to more transient complexes [24] that are more hydrophilic. Stable complexes show more packed interfaces, with fewer hydrogen bonds between the subunits than the less stable complexes. According to several studies, protein interfaces are dominated by aromatic (Phe, His, Try) and aliphatic (Met, Val, Ile, Leu) residues, and with the exception of arginine, they are depleted in charged residues. Interestingly, arginine is one of the residues that appear most often at the interfaces. The area values of known protein-protein interfaces follow a normal distribution that peaks in the range of 600-800 Å2 [25], suggesting that complex formation requires a minimum number of contacts, in addition to the displacement of water molecules [26]. 8 Another area of research focused on the energetic contributions to the binding affinity. Hydrophobic [27], electrostatic [28, 29], and van der Wall [30] forces play an essential role in characterising the binding affinity. Depending on the complex type, each energetic term can have a different weight in the binding affinity. 1.3. Protein-DNA complexes 1.3.1. Importance of protein-DNA interactions Protein-DNA interactions regulate many biological processes such as protein synthesis, signal transduction, DNA storage, and DNA replication and repair, among others. Learning how protein and DNA interact is fundamental to fully elucidate many central biological processes and disease mechanisms and can also support the discovery of novel therapeutic targets. There is a large variety of protein-DNA binding mechanisms. On the one side, DNAbinding proteins can be very specific to DNA sequence, such as the restriction endonucleases. But on the other side, there are DNA-binding proteins, such as histone proteins and DNA polymerases, which do not discriminate DNA sequences when binding. Figure 1.3. Interface size distribution. Interface size is calculated separately for each side of an interface. The distribution has a peak at 600–800 Å2. About 25% of the interfaces have a (one-sided) size in the range of 800 (±200) Å2 . Figure reproduced from Yan, C., et al., Characterization of protein-protein interfaces. Protein J, 2008. 27(1): p. 59-70. 15 protein-peptide (T60-64) or protein-heparin (T57) among others. However, protein-DNA docking received limited attention from the CAPRI community and developers of computational methods. Compared to protein-protein docking, where the most recent release of the standard Protein-Protein Docking Benchmark 5.5 [87] has 257 entries, and to protein-RNA docking, where there are different reported benchmarks [100-103], for protein-DNA docking there is only one available benchmark, which contains 47 complexes [104]. Using this benchmark, protein-DNA docking protocols report moderate success rates in unbound conditions. For instance, on a subset of 23 cases from this benchmark, HDock success rate for top 10 models (i.e. at least one near-native structure within the top 10 models) is less than 10%, while success rate for top 100 is slightly over 30% [95]. NPDock reports a maximum success rate (i.e. at least one near-native conformation found in the entire prediction set) of 7/47 (15%) [94]. Protein-DNA docking with HADDOCK reported an excellent performance [105] when using restraints based on the real interface. This represents a very promising approach, but in a realistic scenario, lack of knowledge on the actual complex interface might limit its application. A more recent coarse-version of HADDOCK protein-DNA docking shows similar accuracy with ~6-fold speed increase over atomistic calculations [106]. The need of new computational tools to address unbound protein-DNA docking is clear. 16 17 2. OBJECTIVES 18 The objectives of this thesis can be grouped in these general aims: i. Improve existing protein-protein docking functionalities ii. Development of new protein–DNA docking protocols ii. Exploration of new procedures that integrate ab initio and template-based proteinprotein docking iii. Implementation of the developed tools and protocols as web servers to share them with the community iv. Evaluation of the new developments in blind conditions (CAPRI-CASP) v. Application to systems of biological and biotechnological interest 19 3. METHODS 20 The protocols described here for the installation and execution of pyDock have been reported in this publication: Rosell, M., L.A. Rodríguez-Lumbreras, and J. Fernández-Recio, Modeling of Protein Complexes and Molecular Assemblies with pyDock, in Protein Structure Prediction, D. Kihara, Editor. 2020, Springer US: New York, NY. p. 175-198. 21 3.1. pyDock The pyDock method is a set of protocols for protein-protein docking and scoring, previously developed [79, 82], validated [107] and successfully applied to many cases of biological and biotechnological interest [108]. The pyDock software needs the coordinates of the two interacting proteins, usually as PDB files. Hydrogen atoms are not needed in the PDB files, and if present, they will be removed and rebuilt again by pyDock. In addition, all HETATM coordinates will be removed in the docking calculations. In this thesis, new functionalities have been implemented to be able to efficiently use AMBER coordinate and topology files. The newest version of the pyDock program optimized during this thesis can use AMBER coordinate files (with extensions such as inpcrd, .restrt, .rs7, .crd) and topology files (with extensions such as .prmtop, .parm7, .top) created by the PARM, LeAP, SANDER, or GIBBS programs from AMBER [109]. In this case, the HETATM coordinates from the cofactors and other compounds will be included in pyDock calculations. 3.1.1. pyDock installation The pyDock 3.0 package is available at https://life.bsc.es/pid/pydock/get_pydock.html to get pyDock you need to apply for a license by filling in your data; for academic use, you will receive a link to the pyDock distribution file by e-mail; for commercial use, you will be contacted by the authors). Uncompress and untar the pyDock distribution file to extract the pyDock3 directory. Next, we need to change permissions of the pyDock/data directory: > cd pyDock3 > chmod go+rx data > chmod u+x pyDock3 The pyDock3 directory can be moved to any location of your choice. For instance, let us say that it is moved to /usr/local/software/ directory; then, pyDock can be called by: 22 > /usr/local/software/pyDock3/pyDock3 Moreover, the PYDOCK variable can be defined in your .bashrc file, as follows: export PYDOCK=/usr/local/software/pyDock3/ so that the executable of pyDock can be called in a more convenient way: > $PYDOCK/pyDock3 The pyDock binary has been compiled for Linux 32-bit to increase the compatibility with older S.O. 3.1.2. pyDock external programs 3.1.2.1. SCWRL In the case of input PDB files with incomplete side-chains, pyDock uses SCWRL 3.0 (http://dunbrack.fccc.edu/) to rebuild them. We should note that this version is outdated and cannot be directly downloaded from the above web, so you need to obtain the installation file scwrl3_lin.tar.gz from their authors. Then, unzipping this file will extract its contents to a new directory called scwrl3_lin. Inside this directory, run: > ./setup This will create a SCWRL 3.0 binary (scwrl3) in that directory. 3.1.2.2. FTDOCK 2.0 The program needs some external programs to generate a set of rigid-body docking poses. In this regard, pyDock is ready to process the output of FTDock 2.0, and we will show here how to install it. The first step is to install the FFTW libraries, that can be downloaded here: https://www.fftw.org/fftw-2.1.5.tar.gz 23 > sudo apt-get install mpi-default-dev > wget https://www.fftw.org/fftw-2.1.5.tar.gz > tar -xvf fftw-2.1.5.tar.gz Uncompressing this file will extract its contents to a new directory called fftw-2.1.5. Within this new directory, compile the libraries by: > ./configure --enable-type-prefix --enable-mpi -- prefix='/path/to/lib/fftw' > make install Now, the ftdock-mpi-master.zip file can be download. This is the custom FTDOCK version, with grid-optimization and ready to run in parallel [82]. It can be downloaded from the GitHub repository (https://github.com/brianjimenez/ftdock-mpi). This zipped file should be unpacked, and within the new ftdock-mpi-master directory, the Makefile file should be edited to set the FFTW_DIR variable to the full path (--prefix) used as input in the configure command above. Now, within the ftdock-mpi-master directory, type: > ./make This will create the program binaries, such as ftdock. More information can be found in the README.md file 3.1.2.3. ZDOCK Another external tool that can be used for the generation of rigid-body docking poses is ZDOCK (https://zdock.umassmed.edu/). The pyDock pipeline is ready to process the output of ZDOCK 2.1. 24 3.1.3. Automatic use of external programs For automatic use of FTDock, ZDOCK, and SCWRL programs within pyDock, after installing them locally, it is necessary to indicate the full path of the FTDock and ZDOCK directories, and that of the SCWRL binary, by modifying the corresponding lines in the $PYDOCK/pyDock3/etc/pydock.conf file, as follows: (...) ZDOCK=/<your-installation-directory>/zdock2.1_linux_64bit/ FTDOCK=/<your-installation-directory>/ftdock-mpi / SCWRL=/<your-installation-directory>/scwrl3_lin/scwrl3 (...) 3.1.4. Running pyDock pyDock has a highly modular architecture, with a series of modules performing the different functionalities of the program (Figure 3.1). The general syntax for running pyDock is: > $PYDOCK/pyDock3 DOCKNAME modulename Thus, the executable pyDock3 usually needs two arguments: (1) DOCKNAME, which is the name of the pyDock project and the base for all the files that will be created during the docking pipeline, and (2) modulename, which will call for the specific module. The details of the different pyDock modules are described in the running instructions below. 31 peptide chains. And finally, the ATOM records describe the coordinates of the atoms that make up the protein. For example, the first ATOM line represents the alpha-N atom of the first residue of peptide chain A, which is a proline residue; the first three floating point numbers are its x, y and z coordinates and are in units of Angstroms. The following three columns are the occupancy, temperature factor, and atom name. The HETATM records describe the coordinates of the heteroatoms, i.e., the atoms that are not part of the protein or nucleic acid molecule. They can be cofactors or metal atoms, such as Heme, Fe, Cu, etc. (details of the format in the Figure 3.2). 3.2.2. PDBx/mmCIF The current reference format in structural biology is PDBx/mmCIF [112]. It has practically no limitations, from small molecules to large macromolecular complexes. Actually, PDBx/mmCIF is derived from the CIF format, which was first used for small molecules. About 20 years ago, it was updated to mmCIF, aiming to be the successor of the PDB format. Data storage is done in a similar way to XML or JSON, with a dictionary or schema needed to know the internal structure and to be able to access the information. 32 The fact that PDB format is more user-friendly than PDBx/mmCIF, and that the PDB file can be easily retrieved from mmCIF (see Figure 3.3), makes the PDB format to be still widely used, due to its simplicity. In the example shown in Figure 3.3 the field labels appear first, and then immediately below the data (in PDB-like format). For example, _atom_site.id corresponds to the atom number, and the rest of the fields similarly describe the PDB columns. 3.2.3. Mol2 A Tripos Mol2 (.mol2) file is a complete and portable representation of a SYBYL molecule [113]. The most important characteristics of this file is that it explicitly contains atom type and bond information. In many cases, it is essential to use or convert PDBs to this file type to be able to perform the parameterization using antechamber. This was the case for the cofactors processed in Chapter 7 (see Figure 3.4). Figure 3.3. PDBx/mmCIF truncated example of X-ray crystallographic studies of seal myoglobin (PDB 1MBS) 33 3.3. Program languages (R, python, etc) In this section, we will briefly mention some of the most relevant programming languages used in the thesis, with focus on some of their strong points: python 2.7 [114], python 3.x.x [115], perl [116], r [117] and the Linux console bash [118]. Python Guido van Rossum developed Python in 1991. It is a high-level, interpreted, cross-platform and object-oriented programming language. It is also the most popular language (as to 2022). It contains a large number of tools for the bioinformatics community, such as Biopython [119], ProDy [120], and pyProCT [121]. Of particular note is its extensive use in Machine Learning (ML), with Scikit-learn [122], Theano [123] and TensorFlow [124]. It is also the main language used in the development of pyDock 4.0 Figure 3.4. Mol2 example Benzene. 34 Perl Larry Wall initially developed Perl in 1987. Several years later, Andy Dougherty and Tom Christian, among others, joined the project (see https://perldoc.perl.org/perlhist ). Perl is based on a block style like AWK and was widely adopted by genomic bioinformaticians for its text-processing prowess. But in my opinion, it is not an object-oriented language and In the Structural Biology field, it is not widely used. R Ross Ihaka and Robert Gentleman developed it in the 1990's. R was born as a tool strictly for statistical analysis. But over the years it has been adapted and there are also multiple packages to apply ML. To mention a few: data.table, dplyr, ggplot2, caret, e1071, xgboost, randomForest, etc... Bash It is the Swiss army knife for controlling Linux-based systems and although there are exceptions, it is the only way to give commands to the large computing clusters used during the development of the thesis. Bash is an interactive command interpreter and runs in a terminal, where commands are typed. You can also generate a list of commands in a file called a script and execute them. 3.4. Molecular Visualization Software There are many and varied programs to visualise molecules, but the most commonly used by the bioinformatics community are UCSF Chimera [125], UCSF ChimeraX [126], Jmol/JSMol [127], PyMoL , VMD [128], ICM-Browser [129]. Also worth mentioning NGL [130] is embedded in the pyDockDNA server. In the following sections, I will discuss more details of the three molecular visualisers mainly used during the thesis development. 3.4.1. ICM ICM-Browser is a free version of the ICM program (www.molsoft.com) with many features for molecular visualization and structural analysis. It can display surfaces of ligand binding pockets, optimise hydrogens to a PDB, superimpose (structural align) protein structures, 35 measure distances and angles, generate and display surfaces, among others. All the functionalities can be accessed through the command line, but not all of them are easily found in the graphical user interface, e.g. through the command line one can open a collection of SDF files of molecules and perform different measurements such as RMSD calculation, but in the graphical user interface, this option is not visible. 3.4.2. UCSF Chimera The program Chimera was used as an alternative to ICM, especially for the use of functions that were only available in its commercial version. The most interesting feature of Chimera is the possibility of using python scripting to do specific intensive tasks in addition to the command line (https://www.rbvi.ucsf.edu/trac/chimera/wiki/Scripts). You can do visual representations in a similar way as in ICM, as well as Molecular Dynamics (but only for teaching purposes). 3.4.3. Pymol PyMOL is an open-source but proprietary program written in the Python programming language. It allows the creation of plug-ins that extend its functionality, among which is the Autodock plugin, which allows the setup of a docking grid (AutoDock Vina [131]) and view the docking results. 3.5. Benchmarks and evaluation sets 3.5.1. Protein-protein docking benchmark 4.0 The protein-protein docking Benchmark 4.0 [132] (BM4) was used as the target library for validating the protein docking funcionalities developed in this thesis. This protein benchmark provides 176 complexes solved by x-ray crystallography (119 dimers and 57 multimers) with at least 3.25 Å resolution, where the bound and unbound states are known. This benchmark is a non-redundant set of protein complexes that include, among others, enzyme-inhibitor, enzyme-substrate, and antigen-antibody complexes. The targets are classified as rigid body, medium and challenging in terms of the expected difficulty for ab 36 initio protein–protein docking. For further details, see https://zlab.umassmed.edu/benchmark/ web site. 3.5.2. DOCKGROUND In this study, we used the structural templates included in the DOCKGROUND resource v1.0 [133]. The structures of protein-protein complexes included in this library were solved by xray crystallography with a resolution better than 3.5 Å. Only complexes with a mean accessible surface area buried by each chain greater than 250 Å2 and containing at least 10 interface residues are included. Structural diversity was ensured with the MM-align program by using a TM-score cut-off of 0.9, which resulted in a dataset of 7,107 diverse protein–protein interfaces. See http://dockground.compbio.ku.edu for a full database description. 3.5.3. Protein-DNA In order to test the new pyDockDNA docking protocol developed in this thesis, we used a previously reported protein-DNA docking benchmark (version 1.2) [104]. The benchmark contained bound and unbound x-ray crystallography and NMR structures for 47 proteinDNA complexes in which DNA is in B-DNA conformation. These were classified as "easy", "intermediate" or "difficult" cases, based on the interface RMSD values between the bound and unbound components of the complex (Table 3.2). 37 Table 3.2. Protein-DNA docking benchmark (version 1.2) HTH Zinc-coord Other α-helix β-sheet Β-harpin Enzyme 2C5R 1FOK 3CRO 1H9T 1TRO 1RPE 1MNN 1F4K 1K79 1W0T 1Z9C 1DDN 2IRF 1JT0 1ZS4 1O3T 1BY4 1R40 1ZME 1KSY 2FIO 1JJ4 1QRV 1B3T 1HJC 1QNE 1EA4 1AZP 1CMA 1BDT 1PT3 1EMH 1DIZ 1VRR 1KC6 1Z63 1VAS 4KTQ 1G9Z 1A73/1A74 3BAM 1RVA 1DFM 7MHT 2FL3 1EYU 2OAA Classification of cases as previously described [134]. Underlined cases have only 1 DNA molecule. Green are the easy cases with interface RMSD b-u (between bound and unbound of the complex) ranging from 0.0 Å to 2.0 Å. Blue are the medium cases with interface RMSDb-u between 2.0 Å and 5.0 Å. Red ones are the hard cases with interface RMSDb-u above 5.0 Å An additional set of case studies was compiled following the criteria selection used in the above-described protein-DNA docking benchmark. This test set is composed of ten protein-DNA complexes, where both bound and unbound structures are available for each reference complex, and the sequences are different from those in the first protein-DNA docking benchmark (Table 3.3). Protein-DNA complex and unbound structures were compiled from the Protein-DNA Interface Database (PDIdb) [135] and the Protein Data Bank (PDB) [31]. Only complexes that meet the following conditions were considered: i) DNA sequence length larger than eight base pairs, and ii) proteins without mutations in the core of the complex interface. To find the protein unbound structures of the selected proteinDNA complexes, all the PDB entries containing only protein structures were retrieved, including structures solved by NMR. Crystallographic structures with a resolution worse than 3.0 Å were not considered. To avoid redundancy, entries with sequence similarity ≥ 90% were discarded. PDBeFOLD [136] was used to find correspondences between bound and unbound protein structures. This tool performs structural alignments between two 38 (pairwise alignment) or more (multi-alignment) molecules using their 3-dimensional structures. The alignment is based on the Secondary Structure Matching algorithm [136]. Alignments with a Q-score higher than 8.0, high P-score and sequence similarity around 90100% were accepted as the corresponding unbound. Then, the bound and unbound structures for each case, were post-processed according to the protocol followed in a previously developed protein-DNA docking benchmark, for instance by checking consistency between unbound and bound coordinates in chain IDs, residue numbers and atom names [104]. The unbound DNA models were generated by using the software 3DNA [137, 138], in canonical B-DNA conformation (fiber model 4). This additional test set (is freely available at the "Help" section of the server (https://model3dbio.csic.es/pydockdna/info/faq_and_help#extended_bechmark). Table 3.3. List of the case external test set. PDB complex Protein PDB unbound protein RMSD unboundbound protein DNA RMSD unboundbound DNA 5JLT phage T4 MotA DNAbinding domain 1KAF 0.83a 22bp dsDNA 1.89 2X6V TBX5 2X6V 0.55 11bp DNA 2.03 3POV SOX 3FHD 1.46 19bp DNA 2.26 4UUV ETV4 DNA-binding ETS domain 5ILU 1.24 10bp DNA 2.81 2NTC sv40 large T antigen 2FUF 1.13a 21-nt PEN element of the SV40 DNA origin 2.96 2ITL sv40 large T antigen 4NBP 5.37a 24-nt PEN element of the SV40 DNA origin 3.84 3MFK Protein C-Ets1 1GVJ 5.61a stromelysin-1 promoter DNA 4.34 2PI0 IRF-3 3QU6 0.76a PRDIII-I region of human interferon-B promoter strand 1 4.46 1O3R catabolite gene activator protein 4R8H 0.65 11bp DNA 4.77 3MLO Ebf1 3LYR 0.71a 22bp DNA 5.11 a In cases with more than one protein-DNA interface in the x-ray structure, the average value is provided. 39 3.5.4. CAPRI The Critical Assessment of Predicted Interactions (CAPRI) community-wide experiment started around 2001, when the community of developers of protein-protein docking methods aimed to evaluate the success rate of such algorithms [98, 139]. CAPRI has been crucial in pushing its community members to improve and add new functions to their docking protocols [140, 141]. Most of the targets proposed by the CAPRI experiment focused on protein-protein docking procedures. Still, recently there have been new challenges, such as protein-peptide and protein-oligosaccharide docking (see Chapter 6.5.1) [142]. The experiment consists in an open competition in which the predictive success rates of the different docking methods are compared in double-blind conditions. The organizers choose the targets, consisting of experimentally determined complex structures that are not yet publicly available. Thus, the targets are blind for the participants, and the participant names are blind for the organizers when evaluating their predictions. For each target, there are usually two modes of participation in the experiment: predictors and scorers. In predictors, the groups are asked to submit ten models from the sequences of the target structures. Generally, the 3D structures or reasonable templates of the interacting molecules are available, which can be used as a starting point for proteinprotein docking. In the scorer participation, the groups are invited to evaluate a common Figure 3.5. CAPRI-CASP evaluation process. 40 set of docking models submitted by the predictors groups. Only the top 5 or the top 10 models by each participant are considered for the assessment. At the end of each round, which may consist of several targets, the ten models submitted by each participant (either as predictors or as scorers or both) are evaluated based on the ligand RMSD, the fraction of native contacts and the interface RMSD with respect to the actual complex structure (see Figure 3.5). 3.5.5. CASP-CAPRI The CASP and CAPRI communities established close ties during the CASP 2014 edition, where the section of "multimeric assemblies" was organized jointly with the CAPRI community [143]. Since then, a total of five joint CASP-CAPRI rounds were held [107, 144, 145], including this year (2022) edition. These joint CASP-CAPRI rounds have encouraged many developers to integrate their ab-initio docking methods with structure prediction methods. Many of these protocols are periodically collected in books, such as the seventh edition of Protein Structure Prediction 2020 [146]. Participation and evaluation in these rounds are done independently by CASP and CAPRI organizers, the latter in a similar way as explained in the previous section of this chapter, see Figure 3.5. 3.5.6. Physiological/non-physiological homodimers The groups of R.L. Dunbrack and E.D. Levy developed a benchmark set of physiological/nonphysiological dimers (version 3) in the context of the Activity II of the 3DBioInfo ELIXIR community, in which I have participated during this thesis. The Benchmark contains a number of protein homo-dimeric x-ray structures that are classified as either "physiological" or "non-physiological". Physiological homodimers are defined as those that are likely to occur in the cell. Non-physiological are homodimeric interactions that are seen in the crystal structure but are unlikely to occur in the cell, because the known physiological state is monomeric or because the true homodimer involves a different interface. 47 In the initial versions, this amber module could only be used for protein-protein docking with modified amino acids. This version only used the atom charges of the AMBER topology file. The van der Waals (VDW) parameters and the atomic solvation parameters (ASPs) were internally defined by pyDock data files. The original pyDock version used parm94 [151], which limited the type of atoms that can be mapped. Also, the atomic solvation parameters are unique to pyDock (they are not included in the general molecular mechanics force fields) [152]. To make this module work on a wider variety of molecules, the pyDock 4.0 function that reads the topology files was modified to extract both the charge and VDW values and write them in the .amber file setup parameters (this file will be used to calculate the energy 48 scoring function by dockser module). To use the pyDock solvation parameters, we created a dictionary of equivalences between the newest amber atom types and the old (parm94) atom types (see Figure 4.2) Finally, the pyDock function that creates the output PDB files was modified. The resulting PDB can be used directly with FTDOCK, without requiring future modifications. A use case for this module upgrade can be found in Chapter 7.3.1. 4.4.2. pyCluster: Clustering in pyDock 4.0 The RMSD matrix required to apply the BSAS algorithm [153] is computed with an ad-hoc ICM script in the in CAPRI and CASP-CAPRI rounds. But the use of parallelisation in ICM is Figure 4.2. Partial snapshot of the equivalence dictionary between parm94 and the newest forcefield parameters find in the lasted AMBER version. (https://github.com/pyDock/parallel/blob/master/amber_old_to_new.map) 49 not trivial. The initial option was to implement the clustering protocol that we used to test pyDockDNA, pyProCT, but this program (as standalone) is a Python 2.7 module, which is incompatible with the new pyDock version. Therefore, we decided to implement the BSAS algorithm as a new module in pyDock 4.0. The new pyCluster module is able to generate the RMSD matrix very quickly thanks to the use of standard Python 3 multiprocessing, just like the dockser module. And it can be used both for the direct output of pyDock 4.0 and for collections of independent models, as it is the case for the CAPRI scorer experiment. See Figure 4.3A and Figure 4.3B for examples of INI configuration files that can be used. A [clustering] modelslist = cluster_1AVX_clustering.list scoring_function = PyDock [receptor] mol_cluster = A [ligand] mol_cluster = B B [receptor] pdb = 1AVX_r_u.pdb mol = A newmol = A [ligand] pdb = 1AVX_l_u.pdb mol = B newmol = B [reference] pdb = 1AVX_b.pdb recmol = A ligmol = B newrecmol = A newligmol = B [clustering] RMSD_cutoff = 20 Nmodels = 100 Figure 4.3. PyClust inifile. (A) The ini file have as inputs a list of PDBs. The ligand and the receptor chains must be specified. If the RMSD_cutoff and Nmodels are not specified, they are set to 4Å and 100 models, respectively. (B) Corresponds to an ini file, using the direct output of pyDock 4.0. Here, you can select the RMSD_cutoff and Nmodels 50 51 5. pyDockDNA: A NEW WEB SERVER FOR ENERGY-BASED PROTEIN-DNA DOCKING AND SCORING 52 The webserver described here have been reported in this publication: Rodríguez-Lumbreras, L.A., et al., pyDockDNA: A new web server for energy-based protein-DNA docking and scoring. Frontiers in Molecular Biosciences, 2022. 9. 53 5.1. Development of pyDockDNA: a new protein-DNA docking procedure. 5.1.1. Sampling In this first step, the input files with the coordinates in PDB format for the structures (or models) of a protein and a DNA molecule (which can be B-DNA or any other conformation) are checked for potential format errors. Missing side-chains in the protein are rebuilt with SCWRL 3.0 [154], and the electrostatics Amber94 force field [151] is loaded, assigning the charges to the atoms. Then, rigid-body docking poses between the protein and the DNA, represented as 3D grids, are generated with a faster and parallelized version of the original FTDock (v2.0) software [66] in which the number of cells in the grid is optimized for maximum computing efficiency [82]. The molecule (protein or DNA) with the longest maximal distance between any pair of atoms is considered the receptor, that is, the fixed molecule, and the other one is the ligand or mobile molecule. By default, the program uses 0.7 Å grid cell size, 1.3 Å surface thickness, 12º rotation sampling, and keeps the best 3 poses for each rotation. For each target, a total of 10,000 docking poses are generated. 5.1.2. Scoring Then, the protein-DNA docking poses are ranked using a scoring function composed of electrostatics, desolvation and van der Waals energy. This new pyDockDNA scoring function is adapted from the previously pyDock scoring function for protein-protein docking [82, 155], which now includes atom types for nucleotides from Amber94 force field [151] in order to calculate for the modelled protein-DNA complexes. The nucleotide AMBER atom types have been mapped to the previously defined atom types in pyDock within a new parameter set (nuc.dat). 5.1.3. Clustering of protein-DNA docking models in benchmarking When testing this software (see Chapter 5.3) we have run several docking executions in parallel, using different initial random rotations for the input structures, and the bestscoring 100 resulting models for each individual run were merged into a single pool. To avoid redundancy in the final set, all docking orientations were clustered by pyProCT 54 analysis software [121], which implements the GROMOS clustering algorithm [156]. The distance matrix is built with pyRMSD with the option "QCP OMP CALCULATOR" to compute the ligand root-mean-square deviation (L-RMSD) values for all pairs of docking orientations after their receptors were superimposed (https://github.com/victor-gilsepulveda/pyRMSD/). A cut-off value of 4.0 Å was used for L-RMSD to define the clusters. For each defined cluster of models, the orientation with the lowest docking score is selected as the cluster representative. 5.2. Implementation of pyDockDNA as a web server The pyDockDNA program has been built as a module of the new pyDock 4.0 version, and it includes the same third-party programs, modules and tools of the previous pyDock versions, as well as new functionalities to handle nucleic acid structures in a proper way (see Chapter 4). The program has been implemented as a web server, in a virtual machine hosted in one of the data processing centres (CPD) of the Spanish National Research Council (CSIC) and is accessible through the following link: https://model3dbio.csic.es/pydockdna. The operating system installed is Debian10. For the correct management of the available resources, Slurm [150] was installed, which is a task management system for clusters. Thanks to this system, we can control the jobs received by the server and keep them in queue when resources are limited. For security reasons, the reverse proxy Nginx was installed in conjunction with the uWSGI server, as it has a lower rate of serious vulnerabilities compared to Apache (https://www.cvedetails.com/). As far as the web application is concerned, it is defined as a backend and a frontend. The backend is essentially a daemon, i.e. a resident program running in the background. This daemon is an adaptation of the version used by pyDockWEB [82] but updated to run on the newest version of python so that it can host pyDockDNA [157] web application. Essentially, it is in charge of submitting jobs to the Slurm queue, monitoring their progress and handling execution errors. The job executes several pyDock 4.0 modules in a concerted way, according to the options selected by the user in the frontend. This information is reported to a database that acts as a bridge between the backend and the frontend. 55 The frontend is developed using web2py, a framework for designing websites using the python language. On the main page, the user can choose the name of the job and an email address to be notified at the end of the job. One can use the RCSB code to select the structures (ligand and receptor) or upload them from his computer. In the following steps the user can select the chains to be docked, the energetic scoring function, and even include external information (from available experimental data or using predictive methods such as the DBSI server [158], for instance) as residue-nucleotide distance restraints to rescore docking models as previously described for pyDockRST [159]. The output will be a set of docking models represented in different formats: i) the 3D structure of the best-scoring 10 docking models in terms of scoring can be visualized in the output screen, ii) the PDB files for the best-scoring 100 models can be directly downloaded, and iii) the rotation/translation vectors are provided to generate up to a total of 10,000 docking poses. A summary of the docking results can be visualized as a plot with the distribution of the different energy values obtained for all docking poses (Figure 5.1). Figure 5.1. pyDockDNA server results. The top 10 predictions based on the user-selected scoring function are displayed as a table on the top left and as a 3D representation on the upper right. In the lower right area, a graph of the energy distribution of the 10,000 models are represented. 56 5.3. Performance of pyDockDNA evaluated on the protein-DNA docking benchmark. The pyDockDNA web server has been tested on the 47 cases of a previously reported protein-DNA docking benchmark (see Methods). It is known that using different randomly rotated input structures can slightly affect docking predictions of FFT-based docking protocols as in FTDOCK, because this can modify the mapping of the atom positions on the 3D grids [70, 160]. To check for convergence, we applied pyDockDNA to 10 different random rotations of the initial input structures for each benchmark case and computed the predictive success rates for the results obtained from each randomly rotated input structures. The results indicate even more differences in the predictive values than previously reported for protein-protein docking (Appendix 1: Table 10.1.1). For instance, the success rates for the top 10 models ranged from 12.8% to 21.3%. Therefore, for a more robust evaluation, we merged the results of all 10 docking executions and clustered the obtained docking models to remove similar orientations. Figure 5.2 shows the predictive success rates of the cluster representatives resulting from merging these 10 docking runs (see more details in Appendix 1: Table 10.1.2). The predictive success for the default pyDock scoring function (including parameters for nucleotide atoms, see Chapter 5.1.2) are better than those obtained for the individual docking runs, which means that increasing sampling Figure 5.2. Predictive performance for the top N=1, 5, 10, 100 models of pyDockDNA and different combinations of scoring terms on the protein-DNA docking benchmark. 63 6. NEW APPROACHES FOR INTEGRATING AB INITIO AND TEMPLATE-BASED DOCKING 64 The protocols described and the results have been reported in the following publications: 1. Lensink, M.F., et al., Modeling protein-protein, protein-peptide, and proteinoligosaccharide complexes: CAPRI 7th edition. Proteins, 2020. 88(8): p. 916-938. 2. Rosell, M., et al., Integrative modeling of protein-protein interactions with pyDock for the new docking challenges. Proteins, 2020. 88(8): p. 999-1008. 3. Lensink, M.F., et al., Prediction of protein assemblies, the next frontier: The CASP14-CAPRI experiment. Proteins, 2021. 89(12): p. 1800-1823. 4. Lensink, M.F., et al., Blind prediction of homoand hetero-protein complexes: The CASP13CAPRI experiment. Proteins, 2019. 87(12): p. 1200-1221 65 66 6.1. Introduction The computational prediction of 3D protein–protein complexes typically includes two distinct strategies: template-based modeling and ab initio docking. Template-based modeling methods build a model of a protein-protein complex based on template complex structures, that is, formed between proteins that are homologous to those in the modelled complex, assuming that binding mode will be conserved in this situation. Ab initio docking approaches explore the potential binding modes between the interacting proteins through steric and physicochemical complementarity, in cases where no template complex structure is available. 6.1.1. Template-based modeling A wide range of template-based methods have been developed, exploiting the template information differently. Homology modeling uses sequence identity (S.I.) for template identification and model building [163, 164]; threading methods' thread' sequences onto structural templates [165, 166]; template-based docking usually refers to global superimposition of the structures of unbound monomers onto the corresponding subunits in a template complex structure [52]; and structure interface alignment exploits local similarity and generates models by superimposing the interacting monomers onto the interfaces of templates [167, 168]. However, although template-based methods remain the most reliable [169, 170], they critically depend on the availability of templates. Interestingly, a study has postulated that the Protein Data Bank, [171] www.rcsb.org; [171] already contains structural templates to model most characterized protein interactions [52]. By aligning the interacting structures onto the monomers of templates, this study shown that in most cases it is possible to find templates whose individual monomers shared TM-scoremin > 0.4 with monomers of target structures. They also suggested that alignments with TM-scoremin values greater than 0.4 have the same mode of binding. However, Negroni and colleagues [172], found that while templates indeed exist to model the majority of interactions, they mostly lead to incorrect complex structures, especially in cases of remote homology (e.g cases sharing sequence identity < 30% with templates). To identify templates, they aligned the monomers of target 67 complexes to the interfaces of templates and observed that the performance significantly deteriorates when templates share only moderate structural similarity with the target (TMscore ~ 0.4 – 0.6), which is at odds with [52] findings. Another limitation, in addition to the limited availability of templates, is that the low sequence similarity of remote homologous makes template identification difficult when it is based only on the alignment of individual monomer structures. In these cases, ab initio docking can be used to build a putative model of the protein-protein complex from the known monomer structures (see below). We have studied in more detail the capability of template-based modelling under different conditions as well as a new approach integrating ab initio docking with templatebased modelling that can assist in identifying ab initio docking models under conditions of low sequence identity. 6.1.2. Ab initio docking Current ab initio methods generate a vast number of conformations using efficient sampling techniques and discriminate near-native models from incorrect poses employing sophisticated scoring functions. Fast Fourier Transform (FFT) sampling algorithms discretize proteins into grids to accelerate the search space process, which are implemented in programs such as GRAMM-X [92], ZDOCK [67] and FTDock [66]. Other approaches for generating docking poses use Monte Carlo-based searching, [129] such as ICM [129], or RosettaDock [173], molecular dynamics as in HADDOCK [174], and normal modes as in ATTRACT [175] or SwarmDock [176]. Various scoring functions have been developed to select, among the thousands of generated docking poses, those ones that are most likely to resemble the native structures. These functions often include electrostatics, desolvation, and van der Waals energy terms such as in pyDock [79] and ZRANK [177] or statistical potentials as in SIPPER [178] or PIE [179]. However, although ab initio docking methods have proved valuable in yielding high-quality protein–protein models [180], the limited ability of sampling methods to search the conformational space, and the multiple minima that scoring functions generate lead to an extremely high rate of false positives. 68 In any case, docking functions can be also used to score template-based complex models from remote templates. In addition, there is growing interest on repurposing the scoring functions to analyse energetic aspects derived from crystallographic structures and to investigate whether these structures are biologically meaningful (see Chapter 6.6 for more details). 6.2. Limitations of template-based docking 6.2.1. Template-base model generation We explored here to what extent template-based docking depends on the quality of the available templates, in terms of sequence identity with respect to the target. To validate the approach, we used as target structures the 176 reference complexes of the protein-protein benchmark version 4.0 (BM4) [132]. We tried to identify suitable templates for these complexes from DOCKGROUND (version 1.1), which contains 7107 nonredundant PDB structures [181]. First, to obtain sequence identity (SI) values, sequence alignments were performed between the unbound target structures of BM4 and the DOCKGROUND complexes using the SSEARCH program of the FASTA package version 35.4.7 [182]. The BLOSUM50 matrix was used as a scoring matrix with open and gap penalties of 10 and 0.5, respectively. The E-value equal to 10-5, was considered as a threshold value to consider the alignment statistically significant. Once these alignments were obtained, we performed the relevant SI filters, by removing templates with higher SI than a given value (see Figure 6.1). In addition, structural alignments were performed between the unbound target structures of BM4 and the interfaces of the 7107 complexes extracted at 12 Å from the DOCKGROUND database, as well as on the complete monomers, by using both TM-align, version 20130511 [59] and MM-align[60], version 20130815, respectively. In the case of TM-align approach, we selected the best combination of the four possible alignments between the templates and the targets of BM4. From the two TM-scores we calculated the averaged TM-score (TM-scorea), and the lowest TM-score (TM-scoremin). When MM-align is used, this software automatically generates the best combination, but interface monomers often differ in size, so their contribution to the global 69 TM-score may be different. To calculate such contribution, we computed individual TMscore of ligands and receptors, aligned previously with MM_align, and the two alignments generated by TM-align, using the TM-score program [183], version 20130511, and calculated the average TM-score. 𝑇𝑇𝑇𝑇 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑎𝑎= 𝑇𝑇𝑇𝑇 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠 𝑙𝑙𝑙𝑙𝑙𝑙 +𝑇𝑇𝑇𝑇 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠 𝑟𝑟𝑟𝑟𝑟𝑟 2 (6.1) Where TM-scorerec refers to the TM-score of the aligned receptors, TM-scorelig represents the TM-score of the aligned ligands, and the TM-scorea is simply the average of TM-scorerec and TM-scorelig. Then, the models are generated by superimposition. When the templates were selected by TM-align (full monomers), the selected template is aligned with MM-align to obtain a rotation and translation matrix, thus generating the model by superposition. When the templates were selected with MM-align (interfaces at 12Å), the rotation and translation matrix is a direct output, which can be used to generate the model in the same way. It should be noted that the results presented here have been obtained by using the BM4 monomers in the unbound conformation. Finally, models with a solvent accessible surface area (SASA) of less than 250 Å were discarded, in line with the approach used by DOCKGROUND to create the library of nonredundant templates. The quality criterion that defines a template as a near-native structure is that the Cα-LigRMSD between the base model of the template and the real structure is less than or equal to 10 Å. Finally, we filtered by sequence identity: 100%, 95%, 70%, 30%, and TM-score: 0.4, 0.5, 0.6, 0.7 and 0.8. 70 6.2.2. Predictive success of template-based docking As can be seen in Figure 6.1, template-based docking shows a success rate over 50% when SI is 100%. This relatively low success rate is mainly due to two reasons. The first reason is due to the low redundancy of DOCKGROUND, which means that there are no high SI templates for all BM4 cases. The second reason is that the unbound monomers were used to generate the final models, so the conformational changes (between unbound and bound monomers) are partly responsible for the models not being as good as could be expected. In fact, this approach is very close to how template-based modelling is done in real life, as in the blind condition of CASP or CAPRI experiments. Figure 6.1 also shows that it makes virtually no difference whether we use MMalign or TM-align to select the templates. But in the so-called "twilight zone", that is, in cases with SI below 30%, the use of MM-align provides a significant advantage. In the case of the MM-align results, we performed a more detailed study (Table 6.1), by analysing the success rates at different threshold values of sequence identity and TMscores. The success rate is usually calculated on the total number of the targets of BM4, but here we have done it on the number of cases for which a successful model can be made. Thus, to put the success rate values in the correct context, the coverage (%) is also displayed. Consequently, the success rates are apparently higher than in Figure 5.1. For 52 49 45 43 30 31 14 19 53 53 46 44 31 28 12 11 0 10 20 30 40 50 60 TM-score avegare TM-score min TM-score avegare TM-score min TM-score avegare TM-score min TM-score avegare TM-score min 100% SI 95% SI 70% SI 30% SI Success Rate % Sequence Identity Figure 6.1. Bar-plot showing the success rates of the TM-align (filled bars) and MM-align (patterned bars) software. Templates are filtered according to their SI, so that only templates with lower or equal SI are used to generate the models. 71 example, when the TM-score is set to 0.4, the TM-scoremin values are better in all SI conditions as compared to TM-scorea, but the coverage decreases dramatically between them. This can also be observed for all TM-score thresholds. Table 6.1. Template-based docking success rate for the top 10 models and their coverage (percentage of cases for which a model can be generated over the 176 of the BM4) for the applied TM-scores and sequence identity thresholds. Success Rate (%) Coverage (%) Representative template -base scoring TM-score a 0.4 55.6 47.8 33.6 17.3 93.1 91.4 87.4 76.4 TM-scoremin 0.4 84.2 84.1 74.5 61.5 54.6 47.1 31.6 14.9 TM-score a 0.5 66.9 61.0 46.9 35.6 71.3 67.8 56.3 33.9 TM-scoremin 0.5 85.9 86.1 80.0 75.0 52.9 45.4 28.7 11.5 TM-score a 0.6 83.5 83.3 73.3 64.0 55.7 48.3 34.5 14.4 TM-score min 0.6 86.2 86.1 80.5 76.9 50.0 41.4 23.6 7.5 TM-score a 0.7 84.6 84.4 80.0 71.4 523 44.3 25.9 8.0 TM-score min 0.7 87.2 87.1 81.8 85.7 49.4 40.2 19.0 4.0 TM-score a 0.8 88.2 88.4 81.8 100.0 48.9 39.7 19.0 2.9 TM-score min 0.8 89.0 89.1 80.8 100.0 47.1 36.8 14.9 2.3 100 95 70 30 100 95 70 30 Sequence Identity Both TM-scorea or TM-scoremin have advantages and disadvantages. In the case of TM-scoremin it could be difficult to find a template with a TM-score over 0.4, but if found, it will likely lead to the generation of good models. The opposite happens with TM-scorea. It is easy to make models for almost all targets, but at a cost of a lower quality. Therefore, in order to obtain acceptable models, a threshold higher than 0.4 will be needed. More details on the resulting success rates can be found in Appendix 2: Table 10.2.1 6.3. Ab initio docking can improve template identification 6.3.1. A new protocol for combining ab initio and template-based docking We have devised a modeling strategy that uses ab initio docking to improve template identification, in three stages: i) sampling docking orientations with ab initio docking, ii) searching structural templates for docking orientations using the MM-align program, and 72 iii) scoring docking orientations using a new function that combines pyDock docking energy and TM-score. Figure 5.3 illustrates the overall procedure, and the details are as follows: Ab initio docking We used the pyDock scheme to generate a pool of docking models. Sampling The unbound subunits of the Benchmark 4.0 complexes were translated and rotated randomly to remove possible bias of initial binding conditions. The docking poses were generated using FTDock 2.0 [66], a fast Fourier transform algorithm, which is based on surface complementarity and electrostatics, using 0.7 Å grid cell size, surface thickness of 1.3 Å, a rotation angle of 12º and 3 translations for each rotation. For each benchmark case, a total of 10,000 docking poses were obtained. Scoring We used the pyDock scoring function [184], developed in our group, which incorporates electrostatics, solvation, and van der Waals energy contributions to rank the docking poses. For each target case, we selected the top 100 ranked docking conformations. Re-scoring docking conformations using the structural similarity score (TM-score) i. Interface extraction Interfaces were extracted from the top 100 selected benchmark docking poses and from the DOCKGROUND template library by selecting only those residues with Cα atoms within 12 Å distance from the interface [185] 79 6.4.2. Template-base docking Template-based docking was often used in both joint CASP-CAPRI protein assembly prediction challenges. In the 3rd joint CASP-CAPRI challenge (CASP13-CAPRI46), complexes were modelled based on templates in almost 50 % of the cases (T137-T144, T152-T154, T158). In the 4th challenge (CASP14-CAPRI50) this percentage increased to 70 % of the cases (T164-T168, T170, T171, T175-T177, T180, T181). The templates were found thanks to the information provided by the models extracted from the CASP-hosted servers: ZHANG, ROSETTA, QUARK, MULTICOM-CONSTRUCT and RAPTORX-DeepModeller, as well as from a BLAST search. These templates were analysed and selected based on their structural similarity and biological unit of interest. Templates that did not add relevant information were also filtered out (i.e. redundant templates were removed). Monomeric models were superimposed on the corresponding subunits of each non-redundant template and subsequently scored with the pyDock scoring energy function. 6.4.3. Ab initio docking In general, for the protein-protein and protein-peptide cases, we generated 10,000 docking poses with FTDock 2.0 [66] (with electrostatics and 0.7 Å grid resolution) and 2,000 poses with ZDOCK 2.1 [67], with the exception of a few targets (T131, T132, T136, T149, T159), where ZDOCK was not used due to the computational cost of very large proteins. In three cases (T131-T133), we used the stochastic docking method LightDock [77] to generate additional flexible docking poses during the sampling process. In the scoring phase, the docking poses were scored with the pyDock scoring function. In the above mentioned T131-T133 cases, the pyDockLite [77] and DFIRE [191] functions were used. In this default protocol, cofactors, water molecules and solvent ions were not included in our docking calculations. In case of homo-oligomeric targets, we kept only the docking positions with the expected symmetry (e.g. C2 for homo-dimers, C3 for homo-tetramers, etc.). 80 For protein-saccharide targets (T126-130), we used rDock [192] (http://rdock.sourceforge.net/) to generate and score the models. In addition, we developed a new pyDock module specially adapted for scoring saccharide molecules. 6.4.4. Combining ab initio and template-based methods In the 7th CAPRI experiment and in the 3rd and 4th joint CASP-CAPRI protein assembly prediction challenges, some targets were modelled by combining template-based and ab initio docking. In the T136 target of 7th CAPRI, the homo-decamer interfaces were modelled based on the available BLAST template, by superimposing the binary ab initio docking models on the global template (PDB 5FKZ). In the 3rd joint CASP-CAPRI experiment, there were at least three targets where the same strategy was used: • T146 (A2B2): the homodimer interfaces were modelled based on all available templates from CASP-hosted servers, and the heteromeric interfaces were obtained from ab initio docking. We used Cryo-EM information[193] to localize the ligandprotein and filter the docking results. • T147 (A8): available templates (PDB codes 2W1V, 2GGL and 5H8I) were used to generate homodimeric models, then ab initio docking was performed to build tetramers, keeping only those with 2-fold helical symmetry. • T159 (A6B6C6): the three homo-hexameric rings were independently modelled based on the templates found in the CASP-hosted servers (PDB ID: 1Y12, 3EAA, 4HE1 y 3V4H). Then, using the 3J2M template, we built the hetero-dodecamer. From this macrostructure and the 3rd ring, rigid body docking was applied to model the final homo-octadecamer, keeping only those models where the interacting rings overlapped in a consistent way. In the 4th joint CASP-CAPRI experiment, there were challenging targets where the combined sampling strategy was applied: 81 • T165 (A3H3L3): the homotrimeric glycoprotein was modelled by template-based docking (CASP models to available templates were used), while the hetero-dimeric antibody was modelled with MODELLERv9.19 because no models were available on the CASP-hosted servers. These models were docked to form the final model. • T170 (A6B3C12D6): in this case we applied an ad hoc modelling procedure combining ab initio docking, template-based modelling and manual fitting using a cryogenic electron microscopy (cryo-EM) map. The target consists of three rings with different stoichiometry and protein composition. The first ring was a homohexamer, arranged as a dimer of trimers, and was modelled by structural superimposition fitting the X-ray monomer structure (PDB 5NGJ, chain A) in the available cryo-EM map of the tail of bacteriophage T5 (EMDB ID: 3689). The second ring is formed by three protein subunits of one type and twelve of a second type and was modelled by sequentially building binary interactions with ab initio docking and symmetry constraints. The third ring was modelled in a similar way. The final assembly of the modelled rings was performed by ab initio docking, selecting only those models in which the symmetry axes of the rings were aligned (see Figure 6.4). Figure 6.4. Modeling strategies for the three major rings of target T170 82 • T177 (A20): this complex is formed by two homodecameric rings. Each ring was modelled with MODELLERv9.19 from available templates, and the final assembly was built by applying ab initio docking to the two modelled decamers. 6.4.5. Inclusion of restraints from available external data Besides the above mentioned automatic modeling procedures, we often used available data for each specific case in order to restrain the docking orientations and help selecting the correct models. Some of the most important restraints we applied were based on the oligomerization states of the targets. This information was mainly obtained from the CAPRI and CASP description of the targets, but we also extracted it from templates or oligomerization state predictors [194]. If homo-oligomers were involved, we assumed symmetric oligomerization, e.g. cyclic symmetry C2, C3..., to filter the resulting docking models. This strategy was widely used in the 7th CAPRI (T125, T124, T136) as well as 3rd (T137-T141, T143-T144, T147, T148, T152-T154, T158, T159) and 4th joint CASP-CAPRI experiments (T164, T165, T167-T171, T174-T176, T178, T179). If experimental information was also available, we included it in the modeling procedure as distance restraints. We used a variety of restraint sources (see Figure 6.5). For example, we estimated interface residues to be used as docking and scoring restraints from homologous protein structures or conserved protein-protein interactions. More specifically, these were the restraints applied in each of the experiments: 7th CAPRI: In the case of the T126-T130 targets (protein-saccharide), a more specific software was used to generate the rDock complexes [195] (http://rdock.sourceforge.net/). rDock can include constraints to target the binding cavity. The center of the cavity used was defined as the center of mass of known ligands bound to homologous proteins (PDB 5F7V for T126-T129; PDB 3D5Z for T130). The pyDockRST [159] module was used in multiple targets to efficiently add distance restraints. In some cases, we found homologous templates from which we deduced the restraints to be applied: In targets T134-T135 (homologous template: PDB 1F95) a distance 83 restraint of 5 Å with respect to the residues of the peptide was used. Similarly, a distance restraint of 10 Å was applied in targets T153 (homologous template: PDB 3W36) and T136 (homologous template: PDB 5FKZ). In other targets, we directly applied information about the interaction that was available in the literature, such as in T122 (Trp156 and IL-23A) [196], T125 (LLT1 Lys169, and NKR-P1 Glu205) [197, 198], and T131-T132 (Tyr35 and Ile92 in the common hCEACAM1 protein) [199]. 3rd joint CASP-CAPRI experiment: The target T149 was highly challenging as it not only involved the dimerization of a 5-domain protein, but it was also necessary to describe the assembly of the 5 different domains (D) within each monomer. The pyDockTET module [141] was used together with an ad hoc strategy to generate the models. Each domain was modelled independently based on the rank 1 prediction of the QUARK CASP-host server (each domain was a CASP13 target). Next, the intermolecular orientation between the first domains of each monomer (D1-D1') was modelled based on a template (PDB 1DQS). The interaction between D1 and D2 of the same monomer was modelled by docking, imposing restraints derived from the interdomain bonds with the pyDockTET module. For each D1-D2 model, a copy of it (D1'-D2') was superimposed on D1-D1' to generate D2-D2' pairs. This strategy was iteratively applied to the other domains (D2-D3 by docking, D2'-D3' by overlapping, D3-D4 by docking, etc.). A visual curation of the models was performed to avoid large clashes between domains and then the models were scored using the pyDock scoring function. Another interesting targets were T149, T150 and T151, which were sequentially released for the same protein complex, but with increasing available experimental information. In target T150 we re-evaluated the ab initio docking orientations generated for target T149 with the help of SAXS data, for which we used the pyDockSAXS module. In the case of T151 we also added the newly available cross-linking information in the form of distance restraints using pyDockRST. 84 4th joint CASP-CAPRI experiment: Target T168 was a trimer. We initially built docking trimers with ab initio docking and symmetry constraints, which were compared with an available template (PDB 6FTD), so that models with Cα-RMSD larger than 10 Å were filtered out. In the case of T170, we used a cryo-EM map to select the final models. In target T181 we also applied restraints, since the structure of the separate proteins (PDB IDs 1N3U and 6XDC, respectively) was known. We also knew that 6XDC had a transmembrane domain (residues 44-64, 68-128) to which 1n3u could not bind. This region was used as a "negative" restraint, eliminating the ab initio docking models that showed binding in this region. 6.4.6. Final selection of the models. In general, the scoring of the models, both for the predictors and scorers experiments, was performed with the pyDock bindEy module, which calculates the docking energy of pyDock CASP-hosted server predictions Top5 predictions Top5 predictions Deep Modeller MULTICOM -CONSTRUCT Docking A:A Building A3 Example T168(A3) A:A /A:B 3models Up to 25models CASP Contribu�on >> models >> templates for CAPRI - Filtering by atom clashes <250 - Clustering 4Å - Minimiza�on (AMBER12) pyDock https://life.bsc.es/pid/pydock/ 1 2 3 Selected templates Models Sampling: Template-base ICM-Browser (Te mplat e -superposi�on) •Sequence Iden�ty templates •Templates used in CASP-hosted predic�ons Sampling: Protein-protein Docking Merge of the dockings or crossdocking form: FTDOCK(0.7Å grid, electrostatics) .ftdock file: best 10,000 conformations .rot file:FTDOCK rotation expressed in Euler angles ZDOCK 2.1 .zdock file:best 2,000 conformations .rot file: ZDOCK rotation expressed in Euler angles If Filtering Symmetry restraints Experimental data [Cryo-EM, Cross-linking, SAXS, etc. ] pyDockRST Template-based Sequence Identity: X -ray /RMN structure available. pyDock energy scoring .ene file: Table with all conformation rescored using pyDock energy Submission Deep Modeller MULTICOM -CONSTRUCT Superposition on extracted templates from S.I. &CASPhosted predic�ons templates. Figure 6.5. Protocol followed in the joint CAPRI-CASP14 experiments using the T168 target as an example. The combination of template-based modelling and ab initio docking is shown. It also shows how templates and relevant information can be used to filter between the ab initio models. 85 [80] (for details, see Chapter 3.1.4.3) for a given (experimentally determined or modelled) protein complex. For targets where possible interface residues could be defined based on available experimental information or homologous complexes, this information was usually included in the final score as distance restraints with pyDockRST [159], pyDockSAXS [140] and/or pyDockTET [141], as described in previous section. Other scoring strategies were also used for specific targets, as in T133, an artificially designed complex based on an old target (T47). Such a redesign increased the prediction complexity with respect to T47 and allowed us to explore different strategies. Therefore, we changed the traditional pyDock scoring method successful applied to T47 by integrating new methodologies such as lightDock [77], which added more flexibility to the models, and CCharPPi [78], which provided a more significant number of scoring functions, which were integrated using the IRaPPA voting algorithm [86]. This protocol is implemented in the pyDockRescoring (https://life.bsc.es/pid/pydockrescoring/) server. For the protein-peptide targets of the 7th CAPRI experiment (T134, T135 and T121), we constrained the docking models to adopt the anti-parallel β-chain orientation, which turned out to be correct for targets T134, T135, but incorrect for target T121. After scoring, we removed redundant predictions using a BSAS algorithm [153] with a distance limit of 4.0 Å, as previously described [200]. For models based on templates or symmetry constraints, we eliminated those with strong clashes that would be difficult to solve with minimization. The number of available templates and their reliability determined the percentage of template-based complex models included in the final 5 or 10 models that we submitted for CASP or CAPRI, respectively. Models with more than 250 collisions (i.e., intermolecular pairs of atoms closer than 4 Å) were also eliminated. Then the final ten selected docking poses were minimized using different versions of AMBER (AMBER12 [193] or AMBER17 [194]), with implicit solvent to improve the quality of the docking models and reduce the number of interatomic clashes, as previously described [196]. We always used the same atomic parameters: from AMBER ff99SB [201] force field for proteins, and gaff force field for polysaccharides [202]. The minimization protocol consisted of a 500-cycle steepest descent (SD) minimization with harmonic constraints applied at a force constant of 25 kcal/(mol-Å2) to all backbone atoms in order to optimize the side chains, followed by 86 another 500 cycles of unconstrained conjugate gradient (CG) minimization. In some cases, due to time constraints or earlier convergence, the minimization protocol varied (e.g., in T134 we used 200-cycle SD and 300-cycle GC; in T135 500-cycle SD and 100-cycle GC; in T136 some models were not minimized or were vacuum minimized; the larger CASP targets T149-151, T159, T165, T170, T177, and T180 were vacuum minimized). In targets T131 and T132, loops previously removed for docking were reconstructed by MODELLER before the final minimization step. The protocol we used for the final selection of models in the scorers experiment was the same as the one we used in the predictor experiments, except for a few exceptions, as follows: Distance restraints were not used as scorers in T121 and T136; IRaPPA was not used as scorers in T133. In target T174, models with a Cα-RMSD<10 were filtered against a common template domain. In target T175, no template was used in scorers (while it was used in predictors). In T181, in addition to filtering the models using the 6XDC transmembrane region (only the A chain was used for ab initio docking), the interface region with the other monomer, estimated from the template (amino acids 221-288), was also used to filter the models. 6.4.7. Scoring of protein-saccharide complex models For the protein-oligosaccharide targets (T126-T130), we used rDock [192] (http://rdock.sourceforge.net/) to generate and score the models. In addition, we had to implement new functionalities in pyDock (see Chapter 4.4), since the original version did not have atomic electrostatics, solvation, and van der Waals parameters for saccharide molecules. After this implementation, pyDock was able to read topology and coordinate files from AMBER for all types of molecules, and thus calculate the energy-based scoring function. The data obtained from AMBER files are van der Waals energies and atomic partial charges. As for the atomic solvation parameters (ASPs), a dictionary of equivalences was created between the new AMBER atom types and the pyDock atom types originally used for proteins (https://github.com/pyDock/parallel/blob/master/amber_old_to_new.map). Basically, the ASPs for saccharide C and O atoms were considered as those for "C aliphatic" 87 and "O hydroxyls" of original pyDock, respectively. To obtain the AMBER files mentioned above, we used antechamber with the AM1-BCC charge model [203], setting the net charge to 0, and then parmchk2 to obtain the charges, energy angle parameters, and a mol2 file. We then used LEaP to load the general AMBER force field (GAFF) and followed the procedure to generate a library with the information obtained from the antechamber. As a final step, we used LEaP to load each docking pose and obtain its coordinate (.incrd) and topology (.prmtop) files. These two files are the ones that can be directly used by pyDock to calculate the energy with the bindEy module. In the scoring experiment, we used another charge model due to time constraints, the empirical atomic partial charges of GasteigerMarsili [203], and the final score was based solely on this new version of pyDock adapted to glycosidic protein interactions (rDock was not used). 6.5. Evaluation of the predictive results in CAPRI and CASP The models submitted to CAPRI and CASP with the developed methodology described in the previous section were officially evaluated by the organization of CAPRI and CASP, which is a useful exercise that allows us to have a more objective knowledge about the applicability and limitations of our methodological approaches, and a fair comparison between methods from other groups. In the 7th edition of CAPRI, we participated in all targets, as predictor, scorers and servers, the latter with the exception of the protein-saccharide cases since our pyDockWeb [82] server was not ready for automatic processing of this type of interactions. The predictive performance of our group as well as that of other participants is described in full detail in a previous publication [204]. In the case of the 3rd and 4th joint CASP-CAPRI experiments, we participated as predictors and scorers. The predictive performance was described in full detail in previous publications [142, 144, 145]. Below we summarize our results for the 51 targets proposed in the 7th CAPRI, 3rd and 4th joint CASP-CAPRI experiments (considering the hetero-meric and homo-meric interfaces at the T125 target as two separate targets). The performance for 7th CAPRI is summarized in Table 6.3, and in Appendix 2: Tables 10.2.6, 10.2.7, 10.2.8, 88 for the 3rd joint CASP-CAPRI in Table 6.4 and in Appendix 2: Table 10.2.10, and for the 4th CASP-CAPRI in Table 6.5 and in Appendix 2: Table 10.2.11. 6.5.1. 7th CAPRI The results are consistent with previous participations, as we submmited acceptable models for 10 targets as predictors, four as servers and 13 as scorers (Table 6.3). The total number of evaluated interfaces was 19, as there were multiple interfaces that were considered independent targets. These results, considering our 10 best models, represent a success rate of 53% as predictors, 21% as servers and 68% as scorers, the latter being the best of all participants. Using the CASP criteria in which only the top 5 submitted models are evaluated, the results as predictors and servers did not change, but the performance as scorers was significantly reduced. Based on the results from the evaluation for all participants, the organization considered that there were nine difficult targets, three of medium difficulty, and five easy targets. The performance summary is represented in Table 6.3, which shows the quality of the 10 best predictions submitted for each interface and/or target in which we participated (high***, medium**, and acceptable* [204]). 95 Table 6.5 Quality of submitted predictions for the CASP14-CAPRI50 experiment predictions. applied an ad-hoc modeling procedure (see Chapter 6.4.4), also combining ab initio docking and template-based modeling. We obtained two acceptable models for interfaces #1 (A:B) and #2 (A:E), where we were the only group to submit an acceptable model. However, in the averaged evaluation for interfaces #8 (D) and #9 (CiD), we obtained acceptable results as human and server scorers (along with Venclovas group, these were the only acceptable rank one submissions from all participants). 6.5.3.2. Unsuccessful predictions This 4th joint CASP-CAPRI protein assembly prediction challenge has proven to be more complex than the previous one, with higher proportion of difficult cases. For targets T169, T165, and T174, no group managed to generate models of acceptable quality. For targets T164, T169, T176 and T174, the preferred strategy in predictors and scorers was ab initio Easy Targets Stoich. Submission quality for Predictors Successful groups2 Submission quality for Scorers (human) Submission quality for Scorers (servers) Successful groups for Scorers T164 A2 - 19/28 * * 17/23 T166 A1B1 ** 17/24 ** ** 14/19 T168 A3 ** 18/24 ** ** 17/20 T177 A20 ***/***/- 21/22/14 of 24 **/***/** ***/***/- 17/18/14 of 18 Medium-Diff. Targets Stoich. Submission quality for Predictors Successful groups Submission quality for Scorers (human) Submission quality for Scorers (servers) Successful groups for Scorers T169 A2 - 0/27 - - 0/20 T176 A2 - 3/27 - - 9/21 T178 A2 * 13/26 - - 10/19 T179 A2 - 10/25 * * 17/19 T165 A3H3L3 - 0/27 - - 0/22 T174 A3 - 0/24 - - 0/19 T170 A6/B3/C12/D6 */*/-/-/-/-/-/-/- 13/1/5/4/12/1/1 /7/6 of 23 **/-/-/*/**/-/-/*/* **/ -/-/*/**/-/-/*/* 17/0/10/9/16/2/2/ 13/13 of 17 T180 A240 -/** 1/19 of 25 */** */* 4/17 of 18 1***=high-quality models; **=medium-quality models; *=acceptable-quality models for the results in the performance [170]. 2In general, docking servers and CASP14 -CAPRI predictors are included. For T167, T175 and T181, there are no evaluated data provided by CAPRI. The 3 targets of T171-173 are cancelled by CASP. 96 docking, which proved to be unsatisfactory. This was also the case for target T164, where we used a combination of template-based and ab initio docking. Interestingly, we were successful as scorers in both T164 and T179 targets, where the lack of templates was a determinant factor for predictors but not for scorers. 6.6. Protein docking functions for the identification of physiological homodimers In this section, we have evaluated the application of docking and scoring functions for the prediction of homo-dimer assemblies in other challenging scenarios. As mentioned in the introduction to this chapter, structural data on protein-protein interactions are crucial to elucidate their mechanism of action at the molecular level. The formation of these macromolecular assemblies depends on protein concentrations and physicochemical conditions such as pH and ionic strength [206]. But both the protein concentrations and the experimental conditions used during the use of such methods often differ from the physiological conditions under which the interactions and/or formation of macromolecular complexes occur. Therefore, in some cases, when experimental parameters differ sufficiently, the 3D structures determined may result in assemblies that do not represent physiological reality. Non-physiological assemblies can also be formed when the structure under study represents a part of a larger complex and has been reconstituted in the absence of additional components. This problem is severe in the case of X-ray diffraction crystallography, by which 86% of the protein structures in PDB are determined (data from October 2022). When proteins form a crystal, spurious contacts between proteins can occur only to stabilize the crystal lattice, but might not be relevant in solution. In some proteins, especially in apparently homo-dimeric crystals, identifying which of these contacts are physiologically relevant is problematic, which makes it difficult to assign the true oligomeric state, i.e., whether it is a monomer, a homodimer, or a higher order assembly. The authors generally provide these data during the PDB deposition process but remain prone to errors, as they require independent biophysical/biochemical characterisation, which can give ambiguous results. 97 Moreover, there is a significant fraction (18%) of protein assemblies resolved by X-ray crystallography on the PDB for which no associated publication exists. Several computational methods have been developed to infer the oligomeric state of proteins directly from the contacts made by proteins in the resolved crystal lattice structure. These methods make use of the results of a large number of previous studies, in which the properties of the interfaces of native protein complexes were systematically evaluated and compared with those of the crystal contacts, considered to represent weak non-specific interactions [207]. The most used methods are PISA [194], EPPIC [208] and PRODIGY-crystal [209] . PISA evaluates the chemical and structural properties of the interfaces, while EPPIC uses geometric measurements and sequence entropy of homolog sequences, which is often associated with regions involved in biological function [210]. To address the problem, we have participated in an activity proposed by the ELIXIR 3D-BioInfo Community, which provided a benchmark composed of 1677 homo-dimeric complex structures, of which 841 are non-physiological and 863 are physiological (so-called "dimer benchmark version 3"; for more details see Chapter 3.5.6). We have evaluated the capabilities of our pyDock scoring functions, as well as other descriptors in CCharPPI (Computational Characterisation of Protein-Protein Interactions) server, regarding the identification of physiological homo-dimers in this benchmark. 98 6.6.1. Applicability of pyDock and CCharPPI scoring functions to the identification of physiological homodimers. We evaluated pyDock scoring [79], as well as each of its individual energetics terms, electrostatic, desolvation, and van der Waals, together with 88 descriptors in CCharPPI [78] (https://life.bsc.es/pid/ccharppi/), by applying them to the proposed cases in the dimer benchmark version 3 developed in collaboration with the ELIXIR 3D-BioInfo Community. For each of these descriptors, we evalueted whether they were able to identify the correct physiological dimers, and thus calculated different predictive success metrics, such as sensitivity or coverage (TPR), precision or positive predictive value (PPV), accuracy (ACC) and Matthews correlation coefficient (MCC). 𝑃𝑃𝑠𝑠𝑠𝑠𝑠𝑠𝑃𝑃𝑠𝑠𝑃𝑃𝑠𝑠𝑃𝑃 𝑠𝑠𝑠𝑠 𝑃𝑃𝑠𝑠𝑠𝑠𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑠𝑠 𝑃𝑃𝑠𝑠𝑠𝑠𝑃𝑃𝑃𝑃𝑠𝑠𝑃𝑃𝑠𝑠𝑃𝑃 𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑠𝑠 (𝑃𝑃𝑃𝑃𝑉𝑉) = 𝑇𝑇𝑃𝑃 𝑇𝑇𝑃𝑃 +𝐹𝐹𝑃𝑃 (6.6) 𝑁𝑁𝑠𝑠𝑁𝑁𝑉𝑉𝑃𝑃𝑃𝑃𝑃𝑃𝑠𝑠 𝑃𝑃𝑠𝑠𝑠𝑠𝑃𝑃𝑃𝑃𝑠𝑠𝑃𝑃𝑠𝑠𝑃𝑃 𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑠𝑠 (𝑁𝑁𝑃𝑃𝑉𝑉)=𝑇𝑇𝑁𝑁 𝑇𝑇𝑁𝑁 +𝐹𝐹𝑁𝑁 (6.7) 𝑆𝑆𝑠𝑠𝑃𝑃𝑠𝑠𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 𝑠𝑠𝑠𝑠 𝑇𝑇𝑠𝑠𝑉𝑉𝑠𝑠 𝑃𝑃𝑠𝑠𝑠𝑠𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑠𝑠 𝑅𝑅𝑉𝑉𝑃𝑃𝑠𝑠 (𝑇𝑇𝑃𝑃𝑅𝑅)=𝑇𝑇𝑃𝑃 𝑇𝑇𝑃𝑃 +𝐹𝐹𝑁𝑁 (6.8) 𝐴𝐴𝑠𝑠𝑠𝑠𝑉𝑉𝑠𝑠𝑉𝑉𝑃𝑃𝑠𝑠𝑃𝑃 (𝐴𝐴𝐴𝐴𝐴𝐴) = 𝑇𝑇𝑃𝑃 +𝑇𝑇𝑁𝑁 𝑇𝑇𝑃𝑃 +𝐹𝐹𝑁𝑁 +𝑇𝑇𝑁𝑁 +𝐹𝐹𝑃𝑃 (6.9) Matthews (𝑇𝑇𝐴𝐴𝐴𝐴) = 𝑇𝑇𝑃𝑃 ∗𝑇𝑇𝑁𝑁 +𝐹𝐹𝑃𝑃 ∗𝐹𝐹𝑁𝑁 �(𝑇𝑇𝑃𝑃+𝐹𝐹𝑃𝑃)(𝑇𝑇𝑃𝑃 +𝐹𝐹𝑁𝑁)(𝑇𝑇𝑁𝑁 +𝐹𝐹𝑃𝑃)(𝑇𝑇𝑁𝑁 +𝐹𝐹𝑁𝑁) (6.10) Figure 6.6. ROC curves of the best descriptors from CCharPII 99 In Figure 6.6 a ROC curves comparainsion between the selected descriptors (acording the AUC (Area under the ROC curve)) of the four best CCharPII descriptors and the pyDock energetic terms,) are shown. In addition, an analysis comparing sensitivity and safety is shown in Figure 6.7. As can be seen, the best descriptors extracted from CCharPII are NIPacking described in [211], change in rotational entropy upon complexation (ROT_S), change in translational entropy upon complexation (TRANS_S) both calculated as in CHARMM [212] and the surface complementarity score (NSC) described in [211]. In the case of the pyDock energetic terms, the best descriptor was desolvation energy term. These descriptors are complementary with other descriptors from different groups also participating in this activity of the 3D-BioInfo community. When all descriptors are aggregated using machine learning methods, such as random forests, they yield a consensus scoring function that obtains a higher discriminatory power: 0.94 AUC (Area Under the Curve). (AUC 0.78). (AUC 0.77). (AUC 0.77). (AUC 0.77). (AUC 0.65). (AUC 0.63). (AUC 0.73). (AUC 0.59). Figure 6.7. Shows the sensitivity (dashed line(s)) and accuracy (solid line(s)) of the four best CCharPII descriptors, the pyDock energetic terms and pyDock scoring function. 100 101 7. APPLICATION TO A CASE STUDIO: MODELING ELECTRON TRANSFER PROTEIN COMPLEXES. 102 The results described here have been reported in this publication: Castell, C., et al., New Insights into the Evolution of the Electron Transfer from Cytochrome f to Photosystem I in the Green and Red Branches of Photosynthetic Eukaryotes. Plant Cell Physiol, 2021. 62(7): p. 1082-1093 103 7.1. Introduction In this chapter of the thesis, we have applied some of the developed docking tools to model protein complexes involved in electron transfer in photosynthesis. Photosynthesis is essential for capturing and storing solar energy in the biosphere. Thanks to this process, plants absorb millions of tonnes of CO2 per year on a net basis. At the molecular level, the efficiency of this process lies in a series of highly optimized multi-protein complexes for electron transfer, such as Photosystem I, one of the most efficient photoelectric systems in nature, which can convert solar energy into chemical energy at almost 100% efficiency. The basic mechanisms and components of photosynthesis have been conserved throughout evolution from cyanobacteria to higher plants, although there are interesting differences. In photosynthetic organisms, two proteins act as electron transporters from cytochrome f to photosystem I: plastocyanin (Pc) (containing copper) and cytochrome c6 (Cc6) (containing iron). The two proteins are very different in composition but have equivalent structural and functional aspects. By analysing the available three-dimensional structures and applying some of the computational modelling methods developed and/or optimize in this thesis, it has been possible to identify important details of the molecular mechanisms of photosynthesis in different organisms. In cyanobacteria and many green algae, the two proteins can be used alternatively, depending on the environmental conditions. However, higher plants (green lineage) have only Plastocyanin, which forms strong and efficient complexes for electron transfer. On the other hand, species of the red lineage, such as red algae and diatoms, have only cytochrome c6, which forms weaker and less efficient complexes for electron transfer. Interestingly, plastocyanin genes have been found in oceanic diatoms. In fact, in the case of Thalassiosira oceanica, it possesses such genes and weakly expressing CC6, reducing its dependence on iron. But at the expense of generating less efficient electron transfer complexes. 7.2. Molecular structures and modelling Modelling of the proteins in this study was carried out with MODELLER version 9v23 [205]; https://salilab.org/modeller/, using the default settings [213]. Template searching and sequence alignments were carried out with BLAST [187] and ClustalW [214]. 104 The truncated Cf forms of Phaeodactylum tricornutum (UniProtKB: A0T0C9 CYF_PHATC) and Thalassiosira oceanica (UniProtKB: E7BWE1_THAOC), were modelled using as template a truncated Cf form that was found in Chlamydomonas reinhardtii (PDB ID 1CFM) with which they share 58.2% and 44.6% sequence identity, respectively. The coordinates of the cofactors (heme and Cu ion) were taken from the template structures. The structure of the small domain deletion variant of P. tricornutum Cf was generated by removing residues 171229 from the intact protein structure with ICM Browser software (http://www.molsoft.com) [215]. The Thalassiosira oceanica plastocyanin model was obtained using Chlamydomonas reinhardtii plastocyanin as a template (PDB code 2plt). For each modelled protein, 100 models were built and the models with the best DOPE (Discrete Optimized Protein Energy) score were selected [188], as previously described [216]. The 3D structure of P. tricornutum Cc6 was obtained from the Protein Data Bank, PDB code: 3DMI [217]. The representation of the electrostatic surface potentials, shown in Figures 7.1 and 7.2, was generated with the UCSF Chimera program. The Cu atom in Pc was assigned a charge of +2, while the heme atomic charges for both Cf and Cc6 (-2) were distributed as: Fe (+2), two nitrogen atoms of the ring (-1 each), and the two propionic acid side chains (-1 each) [218]. 7.3. Protein-protein docking simulations 7.3.1. Docking sampling and scoring The protein-protein docking simulations were performed by pyDock scheme [79, 82], adapted here for the inclusion of cofactors during docking calculations. The setup step (see Chapter 3.1.4) needs the coordinates of the two interaction proteins, usually as PDB files, but it can also take AMBER coordinates and topology files. In the original implementation of this functionality it was possible to include modified amino acids as part of the interacting molecules, but now we needed to modify the automatic protocol to incorporate Cu and the heme group as separated molecules but still part of the receptor or ligand. Below is a detailed description of the steps needed to perform ab into docking on the models of Cytochrome f and Plastocyanin of Thalassiosira oceanica, using pyDock: 111 The same energy-distance landscape can also be observed for the docking models of P. tricornutum [Cf:Cc6] (Figure 7.5B). However, the orientation of cytochrome c6 is different from that obtained in C. reinhardtii (Figure 7.6B, 7.6C). These orientations have smaller interfaces and weaker electrostatic interactions. The hydrophobic interactions gain more weight, being similar to what has been previously described in cyanobacteria [221, 222]. Comparing the best energies-distance docking models for the [Cf:Cc6] complex, we can claim that the model for C. reinhardtii has better affinity (-32 a.u. versus -24.7 a.u.) than the P. tricornutum As in previous studies, a version of Cf with the small domain deletion was used for docking. Surprisingly, for P. tricornutum, the small domain appears to play only a minor role in the interaction with Cc6 (Figure 7.7). This is in contrast with docking simulation in C. Plant C. reinhardtii P. tricornutum B A C Figure 7.6. Representative structures for the [Cf:Pc] complex of plants and best-energy docking models for the [Cf:Cc6] complexes of C. reinhardtii and P. tricornutum. (A) Representative structure for the plant [Cf:Pc] complex (turnip Cf and spinach Pc; PDB code, 2pcf) (Ubbink et al., 1998). Pc is coloured in blue and the copper-bound His87 is shown. (B, C) Best-energy docking models for efficient ET between Cf (in light brown) and Cc6 (in red) of the green alga C. reinhardtii (rank 1, docking energy –32.0 a.u., distance between Fe in Cf to heme in Cc6 of 8.2 Å) and P. tricornutum (rank 6, docking energy –24.7 a.u., distance between Fe in Cf to heme in Cc6 of 8.0 Å, the shortest distance model). 112 reinhardtii, where deletion of the small domain of Cf affected the Cf binding positions of both Pc and Cc6 (Haddadian and Gross, 2006). Finally, we compared models of the [Cf:Pc] complex from T. oceanica with the previously described [Cf:Cc6] complex from P. tricorn7.5.utum, as well as with the equivalent [Cf:Pc] complex from C. reinhardtii in the green lineage. The docking between Cf and Pc from T. oceanica generated a large number of low-energy orientations compared to the [Cf:Cc6] complex from P. tricornutum. This seems to indicate a higher binding affinity, possibly due to the higher electrostatics of the acquired "green-type" Plastocyanin. But the energy-distance landscape does not converge towards a single orientation, resulting in different orientations in a similar range of distances and energies. They can coexist and be biologically functional. These possible orientations are Figure 7.7. Superimposed docking models of P. tricornutum. Superimposed docking models of the P. tricornutum [Cf:Cc6] complex and the model (in blue) corresponding to a truncated Cf without the small domain (rank 2, docking energy –31.0 a.u., distance between Fe in Cf to heme in Cc6 of 8.4 Å). 113 (i) "Head-on" Pc orientation, which is relatively similar to that of some complexes observed in cyanobacteria. Hydrophobic interactions are shown to predominate, and there is no interaction between the electrostatic zones and the small Cf domain (energy of -24.2 a.u. and the shortest distance between Fe and Cu of 11.1 Å) (Figure 7.9A); (ii) "Side-on" Pc orientation, similar to that of the green lineage complexes (figures 7.6A and 7.9A), which includes the electrostatic and hydrophobic patches of both proteins and the small Cf domain (energy of -31.1 a.u.; Fe-Cu distance of 12.7 Å) (iii) "Intermediate" Pc arrangement, which includes the hydrophobic patches and some residues of the electrostatic patches (energy of -30.1 a.u.; Fe-Cu distance of 12.2 Å) (Figures 7.9C). Figure 7.8. Landscape for the computational docking results of the [Cf:Pc] complex of T. oceanica. The docking orientations showed in the next Figure are highlighted in red. The distances between the iron in Cf and the copper in Pc were considered 114 But let's look for models with the lowest energy, regardless of distances. We find a population of models with even more favourable energies (-38 to -36 a.u.), in which strong electrostatic interactions are established, also involving additional positive groups outside the usual electron transfer region in Cf. However, the Fe-Cu distances are about 20 Å (Figures 7.9D). Nevertheless, in this orientation, the highly conserved Y84 residue in Pc (typically named Y83 in cyanobacterial and eukaryotic Pc) points directly towards the hemebinding Y1 of Cf (Y1-Y84 distance of 5.1 Å; Fe-Y84 distance of 9.9 Å) (Figures 7.9D). One could speculate that this is an alternative binding mode and that electron transfer may be possible. A B C D Figure 7.9 Representative structures for the [Cf:Pc] complex of plants and best-energy docking models for the [Cf:Cc6] complexes of C. reinhardtii and P. tricornutum. (A) Representative structure for the plant [Cf:Pc] complex (turnip Cf and spinach Pc; PDB code, 2PCF) (Ubbink et al., 1998). Pc is coloured in blue and the copper-bound His87 is shown. (B, C) Best-energy docking models for efficient ET between Cf [1](in light brown) and Cc6 (in red) of the green alga C. reinhardtii (rank 1, docking energy –32.0 a.u., distance between Fe in Cf to heme in Cc6 of 8.2 Å) and P. tricornutum (rank 6, docking energy –24.7 a.u., distance between Fe in Cf to heme in Cc6 of 8.0 Å, the shortest distance model). 115 7.6. Conclusions In this chapter, we have illustrated the applicability of computational docking using pyDock on a case study relevant to understanding photosynthesis at the molecular level. The structural models provide an explanation for the differences in photosynthetic efficiency between red and green algae. But the lower docking energy model obtained for the [Cf:Pc] complex in T. oceanica (Figure 7.8), despite some evidence, is still highly speculative. Other approaches, such as molecular dynamics (MD) followed by experimental confirmation, will be necessary to be able to make a definite statement 116 117 8. General discussion 118 New developments for pyDock and pyDockDNA to address current challenges This thesis describes the development of technical upgrade and new functionalities in the protein-protein docking software pyDock, as well as the implementation of a new web server for protein-DNA docking. The program pyDock has been substantially updated in order to extend its applicability from a technical point of view and be ready for the new challenges in the field, like the use of molecules different from proteins, including cofactors. For the immediate future, it will be extending by integrating pyProCT into this new code in order to increase the reproducibility of the clustering protocol shown in Chapter 5.1.3 (facilitating a broader applicability). Currently pyDock 4.0 is in development alpha phase, and will be soon updated to the Beta-phase to make it publicly available for the scientific community, as local distribution and also as a web server. The new server pyDockDNA shows reasonable predictive success rates on the available benchmarks. But the number of cases that are available for benchmarking is still too low for optimal testing of new developments. The number of cases in which the structure of the unbound DNA is available is not likely to increase, so we will need to rely on modeling methodologies. Fortunately, there is a variety of modeling strategies, especially those based on deep learning, which might provide accurate models for unbound protein and DNA in a much larger number of cases. These models might include ensembles of conformers, which can also provide better predictions when used in docking. In addition, the pyDockDNA server will be extended with new functionalities. For instance, there are plans to apply more efficient distance restraints between residues and nucleotides, to increase the quality of the generated models. Integration of ab initio and template-based docking In this thesis, we have explored the combination of ab initio docking with template-based docking. This strategy is particularly adequate for cases in which only low-quality templates are available, according to either their SI values (using sequence alignments) or their TMscores (using structural alignments). For that, a set of models were generated by ab initio 119 sampling to identify suitable templates with structural alignment methods. Then, the models were scored by a combined function based on the template similarity (TM-score) and the binding energy (pyDock). This strategy improved the predictions with respect to using ab initio or template-based docking alone. The advantages of combining ab initio and template-based docking were confirmed through our participation in the CAPRI and CASPCAPRI rounds. In particular, in the 3rd Joint CASP-CAPRI experiment our group obtained excellent results, where the energy-based scoring function helped to identify correct models among the different alternatives generated by template-based docking approaches. In a similar way, the application of energy-based scoring as well as other functions as those implemented in CCharPPI server [78, 79] can be extended to the discrimination of biologically meaningful crystallographic homo-dimeric complexes from the artefactual dimeric interactions observed in crystal packing. Assessment of protein-protein docking In our participation in CASP-CAPRI rounds, our protein-protein docking approaches, either ab initio or template-based, needed the structure or reasonable models of the interacting subunits. These were in general obtained from the CASP-hosted servers, which previously had automatically generated models for these subunits, as they were also targets for CASP in other categories. During the 4th Joint CASP-CAPRI round, the developers of AlphaFold (AF) participated in CASP15 targets, obtaining unprecedented results in the ab initio prediction of the protein structures [64]. But since this program did not participate as servers, the CAPRI community did not use these models for the multimeric assembly round. Many of the targets proposed in that round (Chapter 6.5.3) did not have sufficiently good templates for quality modeling, which affected to the success of the docking predictions. This has changed in the last round (5th Joint CASP-CAPRI), where the AlphaFold models for the individual subunits and the complexes were available for the participants in the multimeric assembly section. In addition, the AF-Multimer version can now model protein complexes, with reported success rates of 51% with a false positive rate (FRP) of 1% (paper in preprint) [88]. This Joint CASP-CAPRI round will be a fire test to AF and will provide the 120 predictive capabilities of this approach in blind conditions. These results will be discussed at the next meeting, to be held in Antalya (Turkey) from 10 to 13 December 2022. Application to cases of biological interest We have illustrated the applicability of computational docking using pyDock on a case study relevant to understanding photosynthesis at the molecular level. The structural models provide an explanation for the differences in photosynthetic efficiency between red and green algae. But the lower docking energy model obtained for the [Cf:Pc] complex in T. oceanica (Figure 7.7), despite some evidence, is still highly speculative. Other approaches, such as molecular dynamics (MD) followed by experimental confirmation, will be necessary to be able to make a definite statement. 127 10.2. Appendix 2. Supplementary material for Chapter 6 Table 10.2.1. Success rate of template base modelling by using MM-align. Success Rate MM-align 100% SI and 0.4 of TMscore (MM-align) Success Rate MM-align 95% SI and 0.4 of TMscore (MM-align) Success Rate MM-align 70% SI and 0.4 of TMscore (MM-align) Success Rate MM-align 30% SI and 0.4 of TMscore (MM-align) Top NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min 1 83 51.23 78 82.11 71 44.65 67 81.71 46 30.26 39 70.91 18 13.53 14 53.85 5 88 54.32 80 84.21 75 47.17 69 84.15 51 33.55 41 74.55 23 17.29 16 61.54 10 90 55.56 80 84.21 76 47.80 69 84.15 51 33.55 41 74.55 23 17.29 16 61.54 100 91 56.17 80 84.21 77 48.43 69 84.15 54 35.53 41 74.55 27 20.30 16 61.54 Coverage 162 93.10 95 54.60 159 91.38 82 47.13 152 87.36 55 31.61 133 76.44 26 14.94 Success Rate MM-align 100% SI and 0.5 of TMscore (MM-align) Success Rate MM-align 95% SI and 0.5 of TMscore (MM-align) Success Rate MM-align 30% SI and 0.5 of TMscore (MM-align) Success Rate MM-align 30% SI and 0.5 of TMscore (MM-align) Top NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min 1 81 65.32 78 84.78 69 58.47 67 84.81 43 43.88 39 78.00 16 27.12 14 70.00 5 83 66.94 79 85.87 72 61.02 68 86.08 46 46.94 40 80.00 21 35.59 15 75.00 10 83 66.94 79 85.87 72 61.02 68 86.08 46 46.94 40 80.00 21 35.59 15 75.00 100 83 66.94 79 85.87 72 61.02 68 86.08 47 47.96 40 80.00 21 35.59 15 75.00 Coverage 124 71.26 92 52.87 118 67.82 79 45.40 98 56.32 50 28.74 59 33.91 20 11.49 Success Rate MM-align 100% SI and 0.6 of TMscore (MM-align) Success Rate MM-align 95% SI and 0.6 of TMscore (MM-align) Success Rate MM-align 70% SI and 0.6 of TMscore (MM-align) Success Rate MM-align 30% SI and 0.6 of TMscore (MM-align) Top NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min 1 80 82.47 75 86.21 69 82.14 62 86.11 43 71.67 33 80.49 15 60.00 10 76.92 5 81 83.51 75 86.21 70 83.33 62 86.11 44 73.33 33 80.49 16 64.00 10 76.92 10 81 83.51 75 86.21 70 83.33 62 86.11 44 73.33 33 80.49 16 64.00 10 76.92 100 81 83.51 75 86.21 70 83.33 62 86.11 44 73.33 33 80.49 16 64.00 10 76.92 Coverage 97 55.75 87 50.00 84 48.28 72 41.38 60 34.48 41 23.56 25 14.37 13 7.47 Success Rate MM-align 100% SI and 0.7 of TMscore (MM-align) Success Rate MM-align 95% SI and 0.7 of TMscore (MM-align) Success Rate MM-align 70% SI and 0.7 of TMscore (MM-align) Success Rate MM-align 30% SI and 0.7 of TMscore (MM-align) Top NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min 1 77 84.62 75 87.21 65 84.42 61 87.14 36 80.00 27 81.82 10 71.43 6 85.71 5 77 84.62 75 87.21 65 84.42 61 87.14 36 80.00 27 81.82 10 71.43 6 85.71 10 77 84.62 75 87.21 65 84.42 61 87.14 36 80.00 27 81.82 10 71.43 6 85.71 100 77 84.62 75 87.21 65 84.42 61 87.14 36 80.00 27 81.82 10 71.43 6 85.71 Coverage 91 52.30 86 49.43 77 44.25 70 40.23 45 25.86 33 18.97 14 8.05 7 4.02 Success Rate MM-align 100% SI and 0.8 of TMscore (MM-align) Success Rate MM-align 95% SI and 0.8 of TMscore (MM-align) Success Rate MM-align 70% SI and 0.8 of TMscore (MM-align) Success Rate MM-align 30% SI and 0.8 of TMscore (MM-align) Top NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min NSC TM-score average NSC TM-score min 1 75 88.24 73 89.02 61 88.41 57 89.06 27 81.82 21 80.77 5 100 4 100 5 75 88.24 73 89.02 61 88.41 57 89.06 27 81.82 21 80.77 5 100 4 100 10 75 88.24 73 89.02 61 88.41 57 89.06 27 81.82 21 80.77 5 100 4 100 100 75 88.24 73 89.02 61 88.41 57 89.06 27 81.82 21 80.77 5 100 4 100 Coverage 85 48.85 82 47.13 69 39.66 64 36.78 33 18.97 26 14.94 5 2.87 4 2.30 128 Table 10.2.2. Success rate of template base modelling by using TM-align. Success Rate TM-align 100% SI and 0.4 of TMscore (TM-align) Success Rate TM-align 95% SI and 0.4 of TM-score (TM-align) Success Rate TM-align 70% SI and 0.4 of TM-score (TM-align) Success Rate TM-align 30% SI and 0.4 of TM-score (TM-align) Top NSC TM-score average NSC TMscore min NSC TM-score average NSC TMscore min NSC TM-score average NSC TMscore min NSC TM-score average NSC TM-score min 1 90 52.02 87 53.05 75 43.35 71 44.10 48 27.91 44 28.03 17 9.94 13 8.90 5 91 52.60 90 54.88 77 44.51 75 46.58 51 29.65 48 30.57 18 10.53 19 13.01 10 92 53.18 90 54.88 79 45.66 75 46.58 53 30.81 48 30.57 20 11.70 19 13.01 100 92 53.18 90 54.88 79 45.66 75 46.58 53 30.81 48 30.57 21 12.28 20 13.70 Coverage 173 100.00 164 94.80 173 100.00 161 93.06 172 99.42 157 90.75 171 98.84 146 84.39 Success Rate TM-align 100% SI and 0.5 of TMscore (TM-align) Success Rate TM-align 95% SI and 0.5 of TM-score (TM-align) Success Rate TM-align 70% SI and 0.5 of TM-score (TM-align) Success Rate TM-align 30% SI and 0.5 of TM-score (TM-align) Top NSC TM-score average NSC TMscore min NSC TM-score average NSC TMscore min NSC TM-score average NSC TMscore min NSC TM-score average NSC TM-score min 1 90 55.21 87 65.41 75 47.47 71 58.20 48 33.10 43 43.00 17 13.82 13 18.84 5 91 55.83 88 66.17 77 48.73 73 59.84 51 35.17 45 45.00 18 14.63 16 23.19 10 91 55.83 88 66.17 78 49.37 73 59.84 52 35.86 45 45.00 19 15.45 16 23.19 100 92 56.44 88 66.17 78 49.37 73 59.84 52 35.86 45 45.00 20 16.26 16 23.19 Coverage 163 94.22 133 76.88 158 91.33 122 70.52 145 83.82 100 57.80 123 71.10 69 39.88 Success Rate TM-align 100% SI and 0.6 of TMscore (TM-align) Success Rate TM-align 95% SI and 0.6 of TM-score (TM-align) Success Rate TM-align 70% SI and 0.6 of TM-score (TM-align) Success Rate TM-align 30% SI and 0.6 of TM-score (TM-align) Top NSC TM-score average NSC TMscore min NSC TM-score average NSC TMscore min NSC TM-score average NSC TMscore min NSC TM-score average NSC TM-score min 1 90 66.67 83 76.85 75 59.52 67 72.04 48 46.15 39 60.00 17 24.64 10 29.41 5 91 67.41 84 77.78 77 61.11 69 74.19 51 49.04 41 63.08 18 26.09 13 38.24 10 91 67.41 84 77.78 77 61.11 69 74.19 51 49.04 41 63.08 18 26.09 13 38.24 100 91 67.41 84 77.78 77 61.11 69 74.19 51 49.04 41 63.08 19 27.54 13 38.24 Coverage 135 78.03 108 62.43 126 72.83 93 53.76 104 60.12 65 37.57 69 39.88 34 19.65 Success Rate TM-align 100% SI and 0.7 of TMscore (TM-align) Success Rate TM-align 95% SI and 0.7 of TM-score (TM-align) Success Rate TM-align 70% SI and 0.7 of TM-score (TM-align) Success Rate TM-align 30% SI and 0.7 of TM-score (TM-align) Top NSC TM-score average NSC TMscore min NSC TM-score average NSC TMscore min NSC TM-score average NSC TMscore min NSC TM-score average NSC TM-score min 1 85 78.70 82 81.19 69 75.00 65 78.31 41 64.06 33 68.75 11 40.74 7 46.67 5 85 78.70 82 81.19 69 75.00 65 78.31 41 64.06 33 68.75 11 40.74 7 46.67 10 85 78.70 82 81.19 69 75.00 65 78.31 41 64.06 33 68.75 11 40.74 7 46.67 100 85 78.70 82 81.19 69 75.00 65 78.31 41 64.06 33 68.75 11 40.74 7 46.67 Coverage 108 62.43 101 58.38 92 53.18 83 47.98 64 36.99 48 27.75 27 15.61 15 8.67 Success Rate TM-align 100% SI and 0.8 of TMscore (TM-align) Success Rate TM-align 95% SI and 0.8 of TM-score (TM-align) Success Rate TM-align 70% SI and 0.8 of TM-score (TM-align) Success Rate TM-align 30% SI and 0.8 of TM-score (TM-align) Top NSC TM-score average NSC TMscore min NSC TM-score average NSC TMscore min NSC TM-score average NSC TMscore min NSC TM-score average NSC TM-score min 1 82 82.83 81 84.38 66 81.48 62 83.78 34 72.34 27 75.00 5 50.00 5 62.50 5 82 82.83 81 84.38 66 81.48 62 83.78 34 72.34 27 75.00 5 50.00 5 62.50 10 82 82.83 81 84.38 66 81.48 62 83.78 34 72.34 27 75.00 5 50.00 5 62.50 100 82 82.83 81 84.38 66 81.48 62 83.78 34 72.34 27 75.00 5 50.00 5 62.50 Coverage 99 57.23 96 55.49 81 46.82 74 42.77 47 27.17 36 20.81 10 5.78 8 4.62 129 Table 10.2.3. Success rate of ab initio modelling combining the TM-score(s) and pyDock scoring functions. Success Rate MM-align 100% SI and 0.4 of TM-score (MM-align) Success Rate MM-align 70% SI and 0.4 of TM-score (MM-align) Top NSC TMscore average NSC TMscore min NSC pyDock+TMscore average NSC pyDock+TMscore min NSC TM-score average NSC TM-score min NSC pyDock+TMscore average NSC pyDock+TMscore min 1 37 24.34 34 31.19 35 23.03 29 26.61 20 13.89 16 17.20 23 15.97 15 16.13 5 41 26.97 37 33.94 41 26.97 35 32.11 25 17.36 20 21.51 25 17.36 20 21.51 10 42 27.63 37 33.94 42 27.63 38 34.86 26 18.06 20 21.51 26 18.06 21 22.58 100 45 29.61 38 34.86 45 29.61 38 34.86 29 20.14 21 22.58 29 20.14 21 22.58 Coverage 152 86.36 109 61.93 152 86.36 109 61.93 144 81.82 93 52.84 144 81.82 93 52.84 Success Rate MM-align 100% SI and 0.5 of TM-score (MM-align) Success Rate MM-align 70% SI and 0.5 of TM-score (MM-align) Top NSC TMscore average NSC TMscore min NSC pyDock+TMscore average NSC pyDock+TMscore min NSC TM-score average NSC TM-score min NSC pyDock+TMscore average NSC pyDock+TMscore min 1 37 34.91 26 54.17 35 33.02 26 54.17 20 24.39 12 40.00 21 25.61 13 43.33 5 39 36.79 28 58.33 39 36.79 28 58.33 22 26.83 14 46.67 22 26.83 14 46.67 10 40 37.74 28 58.33 40 37.74 28 58.33 23 28.05 14 46.67 23 28.05 14 46.67 100 40 37.74 28 58.33 40 37.74 28 58.33 23 28.05 14 46.67 23 28.05 14 46.67 Coverage 106 60.23 48 27.27 106 60.23 48 27.27 82 46.59 30 17.05 82 46.59 30 17.05 Success Rate MM-align 100% SI and 0.6 of TM-score (MM-align) Success Rate MM-align 70% SI and 0.6 of TM-score (MM-align) Top NSC TMscore average NSC TMscore min NSC pyDock+TMscore average NSC pyDock+TMscore min NSC TM-score average NSC TM-score min NSC pyDock+TMscore average NSC pyDock+TMscore min 1 35 53.03 22 68.75 33 50.00 23 71.88 18 45.00 11 61.11 19 47.50 12 66.67 5 36 54.55 23 71.88 36 54.55 23 71.88 19 47.50 12 66.67 19 47.50 12 66.67 10 36 54.55 23 71.88 36 54.55 23 71.88 19 47.50 12 66.67 19 47.50 12 66.67 100 36 54.55 23 71.88 36 54.55 23 71.88 19 47.50 12 66.67 19 47.50 12 66.67 Coverage 66 37.50 32 18.18 66 37.50 32 18.18 40 22.73 18 10.23 40 22.73 18 10.23 Success Rate MM-align 100% SI and 0.7 of TM-score (MM-align) Success Rate MM-align 70% SI and 0.7 of TM-score (MM-align) Top NSC TMscore average NSC TMscore min NSC pyDock+TMscore average NSC pyDock+TMscore min NSC TM-score average NSC TM-score min NSC pyDock+TMscore average NSC pyDock+TMscore min 1 27 72.97 19 70.37 26 70.27 20 74.07 13 72.22 9 69.23 13 72.22 9 69.23 5 28 75.68 20 74.07 28 75.68 20 74.07 13 72.22 10 76.92 13 72.22 10 76.92 10 28 75.68 20 74.07 28 75.68 20 74.07 13 72.22 10 76.92 13 72.22 10 76.92 100 28 75.68 20 74.07 28 75.68 20 74.07 13 72.22 10 76.92 13 72.22 10 76.92 Coverage 37 21.02 27 15.34 37 21.02 27 15.34 18 10.23 13 7.39 18 10.23 13 7.39 Success Rate MM-align 100% SI Success Rate MM-align 70% SI and 0.8 of TM-score (MM-align) Top NSC TMscore average NSC TMscore min NSC pyDock+TMscore average NSC pyDock+TMscore min NSC TM-score average NSC TM-score min NSC pyDock+TMscore average NSC pyDock+TMscore min 1 20 76.92 15 83.33 20 76.92 16 88.89 10 83.33 5 71.43 10 83.33 6 85.71 5 20 76.92 16 88.89 20 76.92 16 88.89 10 83.33 6 85.71 10 83.33 6 85.71 10 20 76.92 16 88.89 20 76.92 16 88.89 10 83.33 6 85.71 10 83.33 6 85.71 100 20 76.92 16 88.89 20 76.92 16 88.89 10 83.33 6 85.71 10 83.33 6 85.71 Coverage 26 14.77 18 10.23 26 14.77 18 10.23 12 6.82 7 3.98 12 6.82 7 3.98 130 Success Rate MM-align 95% SI and 0.4 of TM-score (MM-align) Success Rate MM-align 30% SI and 0.4 of TM-score (MM-align) Top NSC TMscore average NSC TMscore min NSC pyDock+TMscore average NSC pyDock+TMscore min NSC TM-score average NSC TM-score min NSC pyDock+TMscore average NSC pyDock+TMscore min 1 30 19.87 27 25.23 30 19.87 22 20.56 7 4.43 2 33.33 8 5.06 2 33.33 5 35 23.18 31 28.97 35 23.18 29 27.10 21 13.29 2 33.33 17 10.76 2 33.33 10 36 23.84 31 28.97 36 23.84 32 29.91 21 13.29 2 33.33 25 15.82 2 33.33 100 39 25.83 32 29.91 39 25.83 32 29.91 35 22.15 2 33.33 35 22.15 2 33.33 Coverage 151 85.80 107 60.80 151 85.80 107 60.80 158 89.77 6 3.41 158 89.77 6 3.41 Success Rate MM-align 95% SI and 0.5 of TM-score (MM-align) Success Rate MM-align 30% SI and 0.5 of TM-score (MM-align) Top NSC TMscore average NSC TMscore min NSC pyDock+TMscore average NSC pyDock+TMscore min NSC TM-score average NSC TM-score min NSC pyDock+TMscore average NSC pyDock+TMscore min 1 30 30.30 19 46.34 29 29.29 19 46.34 4 5.88 2 50 5 7.35 2 50 5 32 32.32 21 51.22 32 32.32 21 51.22 12 17.65 2 50 11 16.18 2 50 10 33 33.33 21 51.22 33 33.33 21 51.22 12 17.65 2 50 12 17.65 2 50 100 33 33.33 21 51.22 33 33.33 21 51.22 15 22.06 2 50 15 22.06 2 50 Coverage 99 56.25 41 23.30 99 56.25 41 23.30 68 38.64 4 2.272727 68 38.64 4 2.272727273 Success Rate MM-align 95% SI and 0.6 of TM-score (MM-align) Success Rate MM-align 30% SI and 0.6 of TM-score (MM-align) Top NSC TMscore average NSC TM-score min NSC pyDock+TMscore average NSC pyDock+TMscore min NSC TM-score average NSC TM-score min NSC pyDock+TMscore average NSC pyDock+TMscore min 1 28 48.28 15 62.50 27 46.55 16 66.67 2 66.67 2 100 2 66.67 2 100.00 5 29 50.00 16 66.67 29 50.00 16 66.67 2 66.67 2 100 2 66.67 2 100.00 10 29 50.00 16 66.67 29 50.00 16 66.67 2 66.67 2 100 2 66.67 2 100.00 100 29 50.00 16 66.67 29 50.00 16 66.67 2 66.67 2 100 2 66.67 2 100.00 Coverage 58 32.95 24 13.64 58 32.95 24 13.64 3 1.70 2 1.136364 3 1.70 2 1.14 Success Rate MM-align 95% SI and 0.7 of TM-score (MM-align) Success Rate MM-align 30% SI and 0.7 of TM-score (MM-align) Top NSC TMscore average NSC TM-score min NSC pyDock+TMscore average NSC pyDock+TMscore min NSC TM-score average NSC TM-score min NSC pyDock+TMscore average NSC pyDock+TMscore min 1 21 72.41 13 72.22 20 68.97 14 77.78 2 100.00 1 100.00 2 100 1 100.00 5 21 72.41 14 77.78 21 72.41 14 77.78 2 100.00 1 100.00 2 100 1 100.00 10 21 72.41 14 77.78 21 72.41 14 77.78 2 100.00 1 100.00 2 100 1 100.00 100 21 72.41 14 77.78 21 72.41 14 77.78 2 100.00 1 100.00 2 100 1 100.00 Coverage 29 16.48 18 10.23 29 16.48 18 10.23 2 1.14 1 0.57 2 1.136363636 1 0.57 Success Rate MM-align 95% SI and 0.8 of TM-score (MM-align) Success Rate MM-align 30% SI and 0.8 of TM-score (MM-align) Top NSC TMscore average NSC TM-score min NSC pyDock+TMscore average NSC pyDock+TMscore min NSC TM-score average NSC TM-score min NSC pyDock+TMscore average NSC pyDock+TMscore min 1 14 77.78 10 76.92 14 77.78 11 84.62 1 100.00 0 0 1 100.00 0 0 5 14 77.78 11 84.62 14 77.78 11 84.62 1 100.00 0 0 1 100.00 0 0 10 14 77.78 11 84.62 14 77.78 11 84.62 1 100.00 0 0 1 100.00 0 0 100 14 77.78 11 84.62 14 77.78 11 84.62 1 100.00 0 0 1 100.00 0 0 Coverage 18 10.23 13 7.39 18 10.23 13 7.39 1 0.57 0 0 1 0.57 0 0 131 Table 10.2.4. Structural availability of the interacting molecules and additional information for the preparation of the submitted models as servers and predictors. a Underscored: target of special difficulty, with only 3 or fewer groups that submitted correct models within their top 5 submitted ones. b PDB code (and chain ID if needed) of the unbound structure (or that bound to a different partner in the case of the saccharides) used for docking. In rackets, interface RMSD with respect to the bound structure, calculated on all the atoms from interface residues. Interface residues are those with at least one atom within 8 Å distance from any atom of the partner molecule, after superimposing the unbound structure onto the corresponding bound structure in the complex. "N/A": if reference complex structure is not yet available. PubChem CID code given in case of non-peptidic ligands. cPDB code (and chain ID if needed) of the template used for homology-based modeling of the receptor or ligand, or template-based docking for the complex (or part of it), as indicated. In some targets, I-TASSER was applied, using multiple templates, as indicated. In brackets, sequence identity, and global RMSD of the Cα atoms of the template with respect to the bound structure ("N/A": if reference complex structure is not yet available). d Use of distance restraints for docking based on interface residues, either from available template for the complex, or from literature, as indicated. e PDB code of the complex reference, if it is now available Targeta RECEPTOR LIGAND COMPLEX Unbound Structure (int RMSD) b Template (SI, RMSD)C Unbound Structure (int RMSD) b Template (SI, RMSD)C Template (SI, RMSD)C Docking restraintsd Target structuree Protein-protein T122 5MXA (3.4 A) - - 111R (25%, 10.4 A) - Literature 5MZV T123 5LMW (N/A) 5BOZ (76%, N/A) - multi-template l-TASSER (N/A, N/A) - - 6EY0 T124 5FWO (2.2 A) 4GRW:H (78%, 3.0 A) - multi-template l-TASSER (N/A, N/A) Dimer: 3MTR (17%, 21.9 A) - 6EY6 T125 hetero 4QKH (1.4 A) - - 3T3A (47%, 10.6 A) - Literature 5MGT T125 homo 4QKH (1.4 A) - 4QKH (1.3 A) - 4QKH (100%, 0.7 Å) - 5MGT T131 4WHD (1.4 Å) - - 5LP2 (96%,0.9 Å) - Literature 6GBG 132 4WHD (1.0 Å) - - 5LP2 (57%,2.9 Å) - Literature 6GBH T133 - 3U43:B (87%, 0.8 Å) - 3U43:A (83%,1.5 Å) - - 6ERE T136 - 5FKZ:E (45%, N/A), multi-template lTASSER (N/A) - 5FKZ:E (45%, N/A), multitemplate l-TASSER (N/A, N/A) 2VYC (40%, N/A), 5FKZ (45%, N/A) 5FKZ interface 6Q6I Protein-peptide T121 1LR0 (N/A) - - multi-template l-TASSER (N/A,N/A), 2HQS (42%, N/A), 2W8B (38%, N/A) - 1TOL interface 6S3W T134 1F3C (2.0 Å) - - 1F95 (33%,1.2 Å), 1F96 (9%, 2.4 Å), 3P8M (36%, 2.3 Å) - 1F95 interface 6GZL T135 1F3C (2.0 Å) - - 1F95 (33%, 1.2 Å) - 1F95 interface 6GZL Protein-saccharide T126 - 5F7V (30%, N/A) 2C7F.2X8S, 3QEF, CID 53356682 (N/A) - - - 6RKH T127 - 5F7V (30%, N/A) CID 74539968, (N/A) - - - 6RKX T128 - 5F7V (30%, N/A) 5HOF, CID 74539969 (N/A) - - - 6RL2 T129 - 5F7V (30%, N/A) 1GYE, CID 74539970 (N/A) - - - 6RL1 T130 3D5Z (0.7 Å) - 5HOF (2.6 Å), CID 74539969 (2.3 Å) - - - 6F1G 132 Table 10.2.5. Information on the targets of the 7th CAPRI CAPRI 7th experiment Target Level Stoich. #Int #Res PDB Description Protein-protein complexes T122 Difficult A1B1C1 1 198/328/330 5MZV Human cytokine hetero-dimer/receptor complex IL23/IL23R T123 Difficult A1B1 1 174/121 6EY0 PorM-Nt/nb(02) T124 Difficult A2B1 1 202/141 6EY6 PorM-Ct/nb(130) T125 Difficult A2B4 5 135/146 5MGT Hetero-hexamer of LLT1/NKR-P1 (extra-cellular domains) T131 Difficult A1B1 1 108/404 6BGB Human CEACAM1/HopQ-Type-I H. pylori T132 Medium A1B1 1 108/418 6BGH Human CEACAM1/hopQ-Type-II H. pylori T133 Easy A1B1 1 69/95 6ERE Redesigned Colicin E2 DNase/Im2 complex T136 Easy A10 3 751 6Q6I LdcA P.aeroginosa; EM Protein-peptide complexes T121 Difficult A1B1 1 115/13 6S3W P.aeroginosa TolAIII domain/N-terminus P.aeruginosa TolB T134 Easy A2B1 1 88/50 6GZJ DLC8 dimer/MAG 50-residue fragment T135 Easy A2B1 1 88/12 6GZL DLC8 dimer (Rat)/MAG 12-residue fragment Protein-oligosaccharide complexes T126 Difficult A1B1 1 415/6 6RKH Arabino-oligosaccharide binding protein, G.stearothermophilus, with AbnE/A6 T127 Difficult A1B1 1 415/5 6RKX Idem with AbnE/A5 T128 Medium A1B1 1 415/4 6RL2 Idem with AbnE/A4 T129 Medium A1B1 1 415/3 6RL1 Idem with AbnE/A3 T130 Easy A1B1 1 315/5 6F1G Arabino-oligosaccharide binding protein, G.stearothermophilus, with AbnB/A5 Note: The target list is subdivided into categories: protein-protein (R39:T122-T124; R40:T125; R42:T131, T132; R43:T133; R45:T136), protein-peptide (R38:T121; T44:T134, T135), and protein -polysaccharide (R41:T126T130). Columns 1 to 3 list the CAPRI target ID, its difficulty, and the number of assessed interfaces; columns 4 to 6 list the number of residues, and the PDB ID; column 7 contains a textual description of the target. 133 Table 10.2.6. Overall 7th CAPRI performance ranking for the protein-protein targets. Extracted from [204] top1 top5 top 10 Rank Group name #targets rank performance1 rank performance1 rank performance1 1 Kozakov/Vajda 11 2 5/1***/3** 1 6/1***/5** 1 6/1***/5** 2 Venclovas 8 1 5/2***/3** 2 5/2***/3** 2 5/2***/3** 3 Seok 11 6 4/3** 3 5/1***/4** 3 5/1***/4** 3 Pierce 9 3 4/2***/1** 3 5/2***/2** 3 5/2***/2** 5 Andreani/Guerois 11 6 4/3** 5 5/1***/3** 5 5/1***/3** 6 Zou 11 6 3/1***/2** 6 4/1***/3** 6 4/1***/3** 6 Zacharias 11 14 4/1*** 6 5/1***/2** 6 5/1***/2** 6 Kihara 11 6 4/1***/1** 6 5/1***/2** 6 5/1***/2** 6 Gray 11 3 5/1***/2** 6 5/1***/2** 6 5/1***/2** 10 Shen 11 14 3/1***/1** 10 4/1***/2** 6 5/1***/2** 10 Moal 8 24 2** 10 4/1***/2** 11 4/1***/2** 10 MDOCKPP 11 5 4/1***/2** 10 4/1***/2** 11 4/1***/2** 10 HADDOCK 11 6 3/1***/2** 10 4/1***/2** 11 4/1***/2** 14 Grudinin 11 6 3/1***/2** 14 3/1***/2** 16 3/1***/2** 14 Fernandez-Recio 11 6 3/1***/2** 14 3/1***/2** 16 3/1***/2** 14 CLUSPRO 11 20 3/2** 14 4/3** 16 4/3** 14 Bonvin 11 6 3/1***/2** 14 3/1***/2** 16 3/1***/2** 18 Weng 11 14 3/1***/1** 18 3/1***/1** 20 3/1***/1** 18 SWARMDOCK 11 14 3/1***/1** 18 3/1***/1** 20 3/1***/1** 18 LZERD 11 14 3** 18 3** 20 3** 18 Chang 11 20 2/1***/1** 18 3/1***/1** 11 4/1***/2** 18 Bates 11 14 3/1***/1** 18 3/1***/1** 11 4/1***/2** 23 Vakser 11 24 2** 23 2/1***/1** 24 2/1***/1** 23 Takeda-Shitaka 8 20 2/1***/1** 23 2/1***/1** 24 2/1***/1** 23 Huang 9 20 2/1***/1** 23 2/1***/1** 20 3/1***/1** 26 Wolfson 7 24 2/1*** 26 2/1*** 27 2/1*** 26 PYDOCKWEB 11 28 1*** 26 2/1*** 24 3/1*** 26 HDOCK 8 30 1** 26 2** 27 2** 26 GALAXYPPDOCK 8 24 2** 26 2** 27 2** 30 Laine 3 28 1*** 30 1*** 30 1*** 31 Ritchie 5 30 1** 31 1** 32 1** 31 Iwadate 1 30 1** 31 1** 32 1** 31 INTERPRED 1 30 1** 31 1** 32 1** 31 Del Carpio 11 30 1** 31 1** 32 1** 31 Carbone 6 36 0 31 1** 30 2/1** 31 Brini 1 30 1** 31 1** 32 1** 37 ZDOCK 1 36 0 37 0 37 0 37 Wang 0 36 0 37 0 37 0 37 Wallner 4 36 0 37 0 37 0 37 UUcourse 0 36 0 37 0 37 0 37 Tuffery 0 36 0 37 0 37 0 37 Schueler-Furman 0 36 0 37 0 37 0 37 Schneidman 1 36 0 37 0 37 0 37 Sanner 0 36 0 37 0 37 0 37 Niv 0 36 0 37 0 37 0 37 Negi 4 36 0 37 0 37 0 37 GRAMM-X 3 36 0 37 0 37 0 37 Gong 0 36 0 37 0 37 0 37 Czaplewski 1 36 0 37 0 37 0 37 Carazo 1 36 0 37 0 37 0 Columns 1 to 3 list the rank, name group and the number of assessed interfaces; columns 4 to 6 list the rank, the quality of the models for the top 1, 5 and 10 uploaded models submitted. 1 This column shows the success of the groups, the first number corresponds to the number of successful targets, then separated by a slash, if necessary, the number and quality of the models is displayed. For example, 5/1***/3**, means 5 successful targets, for which there is one target with high - quality models (***), 3 of medium-quality models (**) and the remaining one is of acceptable quality. 134 Table 10.2.7. Overall 7th CAPRI performance ranking for the protein-peptide targets. Extracted from [204] top1 top 5 top 10 Rank Group name #targets rank performance rank performance rank performance 1 Zacharias 3 1 2/1***/1** 1 2*** 2 2*** 1 Schueler-Furman 3 1 2/1***/1** 1 2*** 1 3/2***/1** 1 Andreani/Guerois 3 1 2/1***/1** 1 2*** 2 2*** 4 Venclovas 3 11 1 4 2/1***/1** 2 2*** 4 Seok 3 7 2/1** 4 3/2** 5 3/2** 4 Moal 3 1 2/1***/1** 4 2/1***/1** 5 2/1***/1** 4 Huang 3 1 2/1***/1** 4 2/1***/1** 5 2/1***/1** 8 HDOCK 2 6 2/1*** 8 2/1*** 8 2/1*** 9 Zou 2 11 1 9 2 12 2 9 Shen 3 8 1** 9 1** 12 1** 9 Kozakov/Vajda 2 11 1 9 1** 8 2** 9 GALAXYPPDOCK 3 21 0 9 1** 12 1** 9 Fernandez-Recio 3 11 1 9 2 12 2 9 CLUSPRO 2 11 1 9 1** 12 1** 9 Brini 2 8 1** 9 1** 12 1** 9 Bonvin 3 8 1** 9 1** 8 2** 17 UUcourse 1 11 1 17 1 12 1** 17 SWARMDOCK 3 21 0 17 1 12 2 17 PYDOCKWEB 2 21 0 17 1 23 1 17 Kihara 3 11 1 17 1 12 1** 17 HADDOCK 3 21 0 17 1 12 2 17 Del Carpio 2 11 1 17 1 23 1 17 Czaplewski 2 11 1 17 1 12 2 17 Bates 3 11 1 17 1 11 3 25 Wang 1 21 0 25 0 25 0 25 Wallner 1 21 0 25 0 25 0 25 Vakser 2 21 0 25 0 25 0 25 Tuffery 1 21 0 25 0 25 0 25 Takeda-Shitaka 1 21 0 25 0 25 0 25 Sanner 1 21 0 25 0 25 0 25 Ritchie 1 21 0 25 0 25 0 25 Niv 1 21 0 25 0 25 0 25 Negi 1 21 0 25 0 25 0 25 MDOCKPP 2 21 0 25 0 25 0 25 LZERD 3 21 0 25 0 25 0 25 Grudinin 3 21 0 25 0 25 0 25 Gong 1 21 0 25 0 25 0 25 Chang 3 21 0 25 0 25 0 Columns 1 to 3 list the rank, name group and the number of assessed interfaces; columns 4 to 6 list the rank, the quality of the models for the top 1, 5 and 10 uploaded models submitted. 1This column shows the success of the groups, the first number corresponds to the number of successful targets, then separated by a slash, if necessary, the number and quality of the models is displayed. For example, 5/1***/3**, means 5 successful targets, for which there is one target with high -quality models (***), 3 of mediumquality models (**) and the remaining one is of acceptable quality. 135 Table 10.2.8. Overall 7th CAPRI performance ranking for the protein-oligosaccharide targets. Extracted from [204] top1 top 5 top 10 Rank Group name #targets rank performance rank performance rank performance 1 Andreani/Guerois 5 4 4/1** 1 5/1***/2** 1 5/1***/3** 2 Seok 5 1 5/1***/1** 2 5/1***/1** 2 5/1***/2** 2 LZERD 5 4 5 2 5/1***/1** 2 5/1***/2** 2 Chang 5 2 4/1*** 2 5/1***/1** 4 5/1***/1** 5 Kozakov/Vajda 5 9 3/1** 5 5/2** 8 5/2** 5 Huang 5 9 4 5 5/2** 8 5/2** 5 HDOCK 5 24 1 5 4/3** 4 5/3** 5 CLUSPRO 5 24 1 5 5/2** 8 5/2** 9 Zou 5 9 4 9 5/1** 14 5/1** 9 Zacharias 5 4 5 9 5/1** 14 5/1** 9 Venclovas 5 4 3/1*** 9 4/1*** 14 4/1*** 9 Takeda-Shitaka 5 9 4 9 5/1** 14 5/1** 9 MDOCKPP 5 14 3 9 5/1** 14 5/1** 9 Kihara 5 9 4 9 4/2** 4 5/3** 9 Grudinin 5 2 4/1*** 9 4/1*** 8 5/1*** 9 Gray 5 24 1 9 4/2** 14 4/2** 9 Fernandez-Recio 5 19 2 9 5/1** 8 5/2** 9 Bonvin 5 14 2/1** 9 3/1***/1** 8 4/1***/1** 19 Shen 5 19 2 19 5 14 5/1** 19 HADDOCK 5 14 2/1** 19 3/1*** 4 5/1***/1** 19 Carbone 4 4 3/1*** 19 3/1*** 22 3/1*** 22 Vakser 4 19 2 22 4 24 4 22 Moal 5 30 0 22 4 22 5 22 GALAXYPPDOCK 5 14 3 22 3/1** 14 4/1*** 25 Pierce 2 14 2/1** 25 2/1** 25 2/1** 25 Bates 5 19 2 25 3 25 3 27 SWARMDOCK 5 24 1 27 2 28 2 27 Negi 4 19 2 27 2 28 2 29 Ritchie 1 24 1 29 1 28 1** 29 Del Carpio 5 24 1 29 1 25 3 Columns 1 to 3 list the rank, name group and the number of assessed interfaces; columns 4 to 6 list the rank, the quality of the models for the top 1, 5 and 10 uploaded models submitted. 1This column shows the success of the groups, the first number corresponds to the number of successful targets, then separated by a slash, if necessary, the number and quality of the model s is displayed. For example, 5/1***/3**, means 5 successful targets, for which there is one target with high-quality models (***), 3 of medium-quality models (**) and the remaining one is of acceptable quality. 136 Table 10.2.9. Information on the targets of the third and fourth CAPRI-CASP experiments 3rd Joint CAPRI-CASP experiment Easy targets CASP ID Stoich. #Int1 #Res2 PDB3 Description T140 T0973 A2 1 146 6YFN Bacteriophage ESE058 coat protein T143 T0983 A2 1 245 6UK5 Cals10 protein T144 T0984 A2 1 752 6NQ1 Two-pore calcium channel protein; EM T152 T1003 A2 1 474 6HRH ALAS2, 50-Aminolevulinate synthase 2 T153 T1006 A2 1 79 6QEK Putative membrane transporter (C. desulfamplus) T147 T0995 A2/A4/A8 3 330 - Cyanide dihydratase (B. pumilus); EM T158 T1020 A3 1 577 7WNQ SLAC1 protein T139 T0961 A4 2 505 6SD8 Acyl-CoA dehydrogenase from Bdellovibrio bacteriovorus T142 H0974 A1B1 1 70/80 6TRI Repressor-antirepressor complex (lysogeny switch) Difficult targets CASP ID Stoich. #Int1 #Res2 PDB3 Description T137 T0965 A2 2 326 6D2V NADP-dependent reductase T138 T0966 A2 2 494 5W6L RasRap1 site-specific endopeptidase T141 T0976 A2 1 252 6MXV Rhodanese-like family protein, bacteria T148 T0997 A2 1 228 - LD-transpeptidase T149 T0999 A2 5 1589 6HQV Pentafunctional AROM polypeptide: five main enzymes of the shikimate pathway T150 T0999 A2 5 1589 6HQV Idem; with SAXS data T151 T0999 A2 5 1589 6HQV Idem; with crosslinking data T154 T1009 A2 1 718 6DRU Alpha-xylosidase T155 H1015 A1B1 1 89/129 - CDI_213 protein, bacteria T156 H1017 A1B1 1 111/129 - 201_INDD4 protein, E. coli T157 H1019 A1B1 1 58/88 - CDI207t protein, E. coli T146 H0993 A2B2 3 275/112 - Lipid-transport, bacterial outer membrane T159 H1021 A6B6C6 7 148/351/295 6RAP 18-mer heterocomplex; EM 4th Joint CAPRI-CASP experiment Easy targets CASP ID Stoich. #Int 1 #Res 2 PDB 3 Description T164 T1032 A2 1 284 6n64 SMCHD1 (human) residues 1616–1899 T166 H1045 A1B1 1 157/173 6xod PEX4/PEX22 complex from Arabidopsis thaliana T168 T1052 A3 1 832 - Tail fiber of the Salmonella virus epsilon15 T177 H1081 A20 3 758 7pk6 2vyc Arginine decarboxylase/bacteria Difficult targets CASP ID Stoich. #Int1 #Res2 PDB3 Description T169 T1054 A2 1 190 6v4v Outer-membrane lipoprotein from Acinetobacter baumannii T176 T1078 A2 1 138 7cwp Tsp1 from Trichoderma virens, small secreted cysteine rich protein (SSCRP) T178 T1083 A2 1 98 6nq1 Helical segment from Nitrosococcus oceani T179 T1087 A2 1 93 - Helical segment from Methylobacter tundripaludum T165 H1036 A3H3L3 1 931/128/106 6vn1 MC Ab 93 k bound to varicella-zoster virus glycoprotein gB T174 T1070 A3 1 335 7rej Protein of attachment region to phage tail T170 H1060 A6/B3/C12/D6 9 464/298/140/142 2vyc Component of the T5 phage tail distal complex T180 T1099 A240 8 (4) 262 6yhg Capsid of duck hepatitis B virus