Departamento de Química Orgánica Facultade de Farmacia Construcción QSAR de Redes Complejas de Compuestos de interés en Química Farmacéutica, Microbiología y Parasitología Memoria presentada por: ISELA GARCÍA PINTOS para optar al grado de Doctor por la Universidade de Santiago de Compostela ISBN 978-84-9887-640-6 (Edición digital PDF)
Construcción QSAR de Redes Complejas de Compuestos de interés en Química Farmacéutica, Microbiología y Parasitología
D. Xerardo García Mera, Prof. Titular y D. Francisco Javier Prado Prado, PDI Doctor Contratado por el Programa Ángeles Albariño, ambos del Departamento de Química Orgánica de la Universidad de Santiago de Compostela (USC). Así como, D. Humberto González Díaz PDI Doctor Contratado por el Programa Isidro Parga Pondal en el Departamento de Microbiología y Parasitología, Área de Parasitología, Facultad de Farmacia, USC. INFORMAN: Que la memoria titulada: “CONSTRUCCIÓN QSAR DE REDES COMPLEJAS DE COMPUESTOS DE INTERÉS EN QUÍMICA FARMACÉUTICA, MICROBIOLOGÍA Y PARASITOLOGÍA”, que para optar al grado de Doctor por la Universidade de Santiago de Compostela presenta Dª ISELA GARCÍA PINTOS, ha sido realizada bajo nuestra dirección, en el Departamento de Química Orgánica de la Facultad de Farmacia de la Universidad de Santiago de Compostela. Y considerando que el trabajo constituye tema de Tesis Doctoral, autorizamos su presentación en la Universidade de Santiago de Compostela. Y para que conste, expedimos el presente certificado en Santiago de Compostela a 20 de marzo de 2011. ___________________________ ____________________________ Fdo. Dr. Francisco Prado Prado Fdo. Prof. Dr. Xerardo García Mera _____________________________ Fdo. Dr. Humberto González Díaz
"I think the next century will be the century of Complexity." Stephen Hawking, January 23, 2000, SAN JOSE MERCURY NEWS
A mis padres y a mis hermanos
índice IV 4. PUBLICACIONES (ANEXOS) ................................................. 47 5. CURRICULUM
ABREVIATURAS
abreviaturas VII SRΠk Matriz estocástica ΘK Distribución de probabilidades ADL (LDA) Linear Discriminant Analysis, término que proviene del inglés: Análisis Discriminante Linear ACP (PCA) Principal Components Analysis, término que proviene del inglés: Análisis de componentes principales Actv Actividad biológica ANN Artificial Neural Network, término que proviene del inglés: Redes Neuronales Artificiales ARN Ácido ribonucleico 3D Tridimensional 4D Cuatro dimensiones D Descriptor Molecular CM Cadenas de Markov CoMFA Comparative Molecular Field Analysis CoMSIA Comparative Molecular Similarity Indices Analysis GSK-3 Enzima glicogen sintasa kinasa-3 HMGR Enzima 3-hidroxi-3-metil-glutaril coenzima A reductasa HMGRIs Inhibidores de la enzima 3-hidroxi-3-metilglutaril coenzima A reductasa
abreviaturas VIII HTS High-Throughput-Screening, término que proviene del inglés: evaluación de alta eficacia LNN Linear Neural Network, término que proviene del inglés: Red Neuronal Lineal m Estado [m]K Magnitudes medias mtMulti-target, término que proviene del inglés: multi-diana M Matriz MARCH-INSIDE Markov Chain Invariants for Network Simulation and Design p Probabilidad QSAR Quantitative-Structure-Activity-Relationship, término que proviene del inglés: RelaciónCuantitativa-Estructura-Actividad QSPR Quantitative-Structure-Property-Relationship, término que proviene del inglés: RelaciónCuantitativa-Estructura-Propiedad QSTR Quantitative-Structure-Toxicity-Relationship, término que proviene del inglés: RelaciónCuantitativa-Estructura-Toxicidad RC Redes Complejas t Tiempo T. cruzi Trypanosoma cruzi TIs Topological Index, término que proviene del inglés: Índices Topológicos
abreviaturas IX v Vector vT Vector transpuesto de v 1b, 2a, 3a… Son ejemplos del sistema usado para identificar los artículos de investigación en este trabajo. En el mismo todos los artículos científicos son identificados con un número que indica su orden de aparición en la Tesis y una letra con formato superíndice que indica el tipo de artículo. Los artículos de tipo (a) son trabajos que usan descriptores moleculares basados en CM y los de tipo (b) otros tipos de descriptores
INTRODUCCIÓN
introducción 3 1.1. Introducción al estudio del QSAR, esquema general de trabajo Actualmente existen más de 15 millones de compuestos que han sido descubiertos o sintetizados en laboratorios químicos. Una gran cantidad de estos compuestos no ha encontrado aún aplicaciones farmacológicas, agroquímicas, industriales o de algún otro tipo. Esto es consecuencia directa de la diferencia existente entre la velocidad con que los nuevos compuestos son preparados y caracterizados, la cantidad de ellos que son sometidos a ensayos experimentales. La situación es más crítica si se tiene en cuenta que un gran número de los compuestos ensayados, de manera masiva por el método clásico de “prueba y error”, da resultados negativos. Este tipo de ensayos experimentales, especialmente los de corte farmacológico y toxicológico, son en general muy caros en términos de recursos materiales, humanos y de tiempo. También es de destacar el aspecto no sólo material, sino de tipo ético que conlleva la investigación con animales y su posterior sacrificio. En todo caso, nuevos paradigmas para el descubrimiento molecular han sido introducidos recientemente, basados en el uso de grandes librerías de compuestos químicos y sistemas robotizados para realizar ensayos biológicos. De tal modo los sistemas HTS (high-throughput screening), permiten la síntesis y ensayo de miles de compuestos cada día.1 En este contexto, la industria farmacéutica ha reorientado las estrategias de búsqueda hacia métodos que permitan una selección o diseño racional de nuevos compuestos. En dicho sentido, los estudios QSAR (quantitative structure-activity-relationships) son usados como 1 Kubinyi, H. Rossiiskii Khimicheskii Zhurnal 2006, 50 (2), 5.
introducción 10 M1 y M2, el número de Harary (H), la invariante de Randic (χ), el índice de conectividad de valencia (χv), el índice de Balaban (J), el índice de topología molecular (MTI) y las auto-correlaciones de Moreau-Boroto (ATSd), por citar sólo algunos, pueden ser expresados como transformaciones v·M·vT.8 También tiene cabida en este grupo los más recientes índices cuadráticos qk(X), lineares fk(X) y estocásticos sk(X), introducidos por Marrero-Ponce et al:12,13 T kk T k T k TTT TvTT TTT XsXfXq ATSMTICJ H MMW wSwuMwwMw wBwuDAvdAd vAvvAvuDu vAvuAvuDu m '' 2 1 '''''' 2 1 2 1 2 1 k21 Como se ha podido ver, todos los símbolos de las matrices y vectores anteriores son de uso común en QSAR. Extensamente explicados en la literatura especializada, no lo serán aquí en detalle, queriéndose hacer hincapié, únicamente, en la idea del carácter unificador de las transformaciones v·M·vT.8,14 Por otra parte, muchos estudios de química computacional hacen uso del concepto de momento espectral. Entre los índices basados en momentos espectrales los más conocidos son los momentos de energía μ(H), los conteos de caminos de auto-retorno srwck, los momentos espectrales de matrices de adyacencia entre enlaces μ(B) y μ(dB), el índice I3 de Estrada, para el grado de plegamiento de proteínas y el número de Kirchhoff (Kf). Todos estos índices pueden 12 Marrero-Ponce, Y. J. Chem. Inf. Comp. Sci. 2004, 44, 2010. 13 Marrero-Ponce, Y.; González-Díaz, H.; Romero-Zaldivar, V.; Torrens, F.; Castro, E. A. Bioorg. Med. Chem. 2004, 12, 5331. 14 Estrada, E; Uriarte, E. Curr. Med. Chem. 2001, 8, 1573.
introducción 11 ser escritos en notación matemática mediante el operador traza de las matrices (Tr), que indica la suma de los valores en la diagonal principal principal de la matriz. Estos descriptores moleculares han sido clasificados tradicionalmente como un grupo aparte de los v·M·vT:815,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38 kk k k k ddkk kk I ,,Tr ! 1 ! 1 TraKfTrμ TrμTrBμTrsrwc 3k kkk ALHH WBBBA A pesar de los esfuerzos realizados con vistas a la unificación de los descriptores moleculares en el esquema v·M·vT, no se han realizado avances en la incorporación de índices como los momentos espectrales, a este prometedor esquema. En el presente trabajo se dará una 15 Gutman, I.; Rosenfield, V.R. Theor. Chim. Acta 1996, 93, 191. 16 Estrada, E. Bioinformatics 2002, 18, 1. 17 Estrada, E. Chem. Phys. Lett. 2000, 319, 713. 18 González, M.P.; Morales, A.H.; Molina R. Polymer 2004, 45, 2773. 19 González, M.P.; Morales, A.H.; González-Díaz H. Polymer 2004, 45, 2073. 20 Morales, A.H.; González, M.P.; Rieumont J. B. Polymer 2004, 45, 2045. 21 Burdett, J.K.; Lee, S. J. Am. Chem. Soc. 1985, 107, 3063. 22 Burdett, J.K.; Lee, S. J. Am.Chem. Soc. 1985, 107, 3050. 23 Lee, S. Acc. Chem. Res. 1991, 24, 249. 24 Gutman, I. Theor. Chim. Acta 1992, 83, 313. 25 Markovic, S.; Gutman, I. J. Mol. Struct. Theochem 1991, 81, 81 26 Jiang, Y.; Tang, A.; Hoffmann, R. Theor. Chim. Acta 1984, 66, 183. 27 Karwowski, J.; Bielinska-Waz, D.; Jurkowski, J. Int. J. Quantum Chem. 1996, 60, 185. 28 Estrada, E.; González-Díaz H. J. Chem. Inf. Comput. Sci. 2003, 43, 75. 29 González, M.P.; Terán, C. Bioorg. Med. Chem. Lett. 2004, 14, 3077. 30 González, M.P.;Terán, C. Bioorg. Med. Chem. 2004, 12, 2985. 31 González, M.P.; Terán, C. Bull. Math. Biol. 2004, 66, 907. 32 González, M.P.; González-Díaz H.; Cabrera-Pérez, M.A.; Molina R. Bioorg. Med. Chem. 2004, 12, 735. 33 González, M.P.; Morales, A.H. J. Comput. Aid. Mol. Des. 2003, 10, 665. 34 González, M.P.; González-Díaz, H.; Molina, R.; Cabrera-Pérez, M.A.; Ramos de Armas, R. J. Chem. Inf. Comput. Sci. 2003, 43, 1192. 35 Cabrera-Pérez, M.A.; Bermejo, M. Bioorg. Med. Chem. 2004, 22, 5833. 36 Cabrera-Pérez, M.A.; García, A.R.; Teruel C.F.; Álvarez, I.G.; Sanz, M.B. Eur. J. Pharm. Biopharm. 2003, 56, 197. 37 Cabrera-Pérez, M.A.; González-Díaz, H.; Fernandez, T.C.; Pla-Delfina, J.M., Bermejo, S.M. Eur. J. Pharm. Biopharm. 2002, 53, 317. 38 Molina, E.; González-Díaz, H.; González, M.P.; Rodríguez, E.; Uriarte, E. J. Chem. Inf. Comput. Sci. 2004, 44, 515.
introducción 12 representación unificada de ambos grupos y esto permitirá simplificar el estudio sistematizado de los descriptores moleculares.
introducción 13 1.3. Cálculo de descriptores moleculares con Cadenas de Markov (CM) Cadenas de Markov es el nombre de una teoría o tipo de modelo matemático definido por Markov.39,40 En nuestro trabajo se utilizará especialmente el método MARCH-INSIDE (del idioma Inglés: Markov Chain Invariants for Network Simulation and Design)41,42,43,44; el cual emplea las CM para calcular descriptores moleculares mediante una aproximación sencilla a fenómenos tales como: a) distribución de electrones de valencia alrededor de los átomos de una molécula. b) propagación de una vibración en una cadena de ARN. c) propagación de interacciones electrostáticas superficiales en una proteína viral o en la estructura plegada 3D de una enzima. d) paso átomo por átomo de un fármaco desde el plasma a un tejido. e) interacción paso por paso de un fármaco con su receptor. Aunque cabe destacar que las CM constituyen unos de los modelos de la teoría de probabilidades más usados, a continuación daremos algunas notas básicas sobre estas: - Las CM estudian fenómenos estocásticos esto es: fenómenos que al ser estudiados mediante mediciones en el tiempo (t) el resultado obtenido no está determinado sino que puede obtenerse cada vez 39 Markov, A.A. Bull. Soc. Phys. Math. Kasan 1906, 15, 155. 40 Bharucha-Reid, A.T. Elements of Theory of Markov Process on the application, McGraw-Hill Series in Probability an Statistic, McGraw-Hill Book Company, New York. 1960, 167. 41 González-Díaz, H.; Prado-Prado, F.;Ubeira, F.M. Curr Top Med Chem. 2008, 8(18), 1676. 42 González-Díaz, H.; Duardo-Sanchez, A.; Ubeira, F.M.; Prado-Prado, F.; Pérez-Montoto, L.G.; Concu, R.; Podda, G.; Shen, B. Curr Drug Metab. 2010, 11(4), 379. 43 González-Díaz, H.; González-Díaz, Y.; Santana, L.; Ubeira, F.M.; Uriarte, E. Proteomics. 2008, 8(4), 750. 44 González-Díaz, H.; Vilar, S.; Santana, L.; Uriarte, E. Curr Top Med Chem. 2007, 7(10), 1015.
i n n troducción co n des t nat u - En C sist e v al o sist e est a de fisi c res p i de n Fig u n cierta p t acar qu e u raleza e C M para e ma ha o o r deter m e ma pas a a dos en t cómo c oquími c p ecto a l o n tificam o u ra 4. Il u p robabil i e todos l e stocásti c referirs e o cupado m inado. P a de un t iempos cada e s c a m j , q o s ejem p o s los si g u stració n i dad un o l os fenó m c os. e a un r e un esta d P or eje m estado i n sucesiv o s tado e s q ue dep e p los ofr e g uientes p n gráfica 14 o de va r m enos a n e sultado d o dado m plo en l n icial al o s. En e s s tá cara c e nde de e cidos a n p ares sis de una c r ios res u n tes me n determi n que cie r l a Figur a t 0 = 0 ( e s te trabaj cterizad o l fenó m n teriorm e tema/es c adena d e u ltados p n cionad o n ado se h r to pará m a 4 se o b e n gris) o habla r o por u m eno en e nte en n tados e n e Marko v p osibles. o s (a-e) s o h abla de m etro t o b serva c o a ocupa r r emos, a d u na m a estudi o n uestro t n la T ab l v . Es de o n por que el o ma un o mo el r otros d emás, a gnitud o . Con t rabajo l a 1.
introducción 15 Tabla 1. Elementos o componentes de una CM para casos concretos estudiados en el presente trabajo. Sistema Estados (mj) Valores del Parámetro Resultados esperados de la CM Capa de electrones de valencia de la molécula. Á tomos (Electronegatividad) ó (electrones compartidos) Presencia o ausencia de los electrones alrededor de cierto átomo j Probs. absolutas con que los electrones se distribuyen alrededor de cada átomo V ibración en una cadena de ARN Nucleótidos (frecuencia de vibración) Propagación permitida o prohibida de la vibración hasta un nucleótido j Probs. absolutas con que un nucleótido participa en la vibración Interacciones superficiales en una proteína viral A minoácidos Superficiales (Carga electrostática) Participación o no de un aminoácido en Inter. Electrost. superficiales Probs. absolutas de Inter. Electrost. superficiales Propagación 3D de interacciones electrostáticas en proteínas A minoácidos (Carga electrostática) Participación o no de un aminoácido en Inter. Electrost. 3D Probs. absolutas de interacción electrostáticas 3D Á tomos de una molécula Á tomo (energía libre estándar) Paso del átomo j a plasma, o tejido Probs. absolutas de partición plasma/tejido
introducción 16 - En la CM se parte de una distribución inicial de probabilidades Ap0(j) que son las probabilidades absolutas iniciales (t0 = 0) con que el sistema ocupa cada estado j. - Dado un sistema con n estados estas probabilidades puede ordenarse en un vector de probabilidades iniciales 0π = [Ap0(1), Ap0(2), Ap0(3), … Ap0(n)]. - Las CM pueden ser representadas por un grafo dirigido, donde los vértices son los estados del sistema y los arcos representan la transición o paso del sistema de un estado a otro. - Lo anterior se complementa con la representación matricial de las CM, convirtiéndolas en una herramienta muy versátil. De tal modo que las relaciones entre los estados del sistema expresadas por el grafo sobre el que se define la CM pueden ser resumidas a través de la matriz estocástica 1П.45 - Los elementos de esta matriz son las probabilidades 1pij de que el sistema pase de un estado i en el tiempo t0 = 0 a otro j en el tiempo t1 = 1 o primer paso, ver Tabla 2: Tabla 2. Representaciones grafo-teórica y matricial de una CM. Grafo Matriz 1П 1 23 4 5 1 2 3 4 5 55 1 53 1 44 1 42 1 35 1 33 1 32 1 23 1 22 1 21 1 12 1 11 1 1 1 000 000 00 00 000 pp pp ppp ppp pp 45 Freund, J.A.; Poschel, T. Eds. Stochastic Processes in Physics, Chemistry, and Biology. In: Lect. Notes Phys. Springer-Verlag, Berlin, Germany 2000.
introducción 17 - Nótese que las transiciones del sistema de un estado i a otro j, que no esté directamente relacionado con él, están prohibidas en el t1 (sistemas con vértices no adyacentes en el grafo). - Las CM cumplen la condición de que las probabilidades de transición kpij tanto para el primer paso (1pij) como para tiempos mayores tk = k > 1 dependeN solamente del estado que el sistema ocupaba en el tiempo inmediatamente anterior tk-1 = k -1 pero no de tiempos anteriores. - Las CM cumplen con las ecuaciones de Chapman-Kolgomorov, por lo que las probabilidades absolutas Apk(j), con que el sistema ocupa determinado estado j en el tiempo k, se determinan como los elementos de los vectores kπ= 0π· kП = 0π·(1П)k. Así, las probabilidades absolutas de evolución del sistema a los tiempos tk = 0, 1, 2, 3, …n quedan determinadas por: nppppnpppp nppppnpppp nppppnpppp nppppnpppp nppppI k A k A k A k A k AAAAkk AAAAAAAA AAAAAAAA AAAAAAAA AAAA n ...3,2,1...3,2,1 . . . ...3,2,1...3,2,1 ...3,2,1...3,2,1 ...3,2,1...3,2,1 ...3,2,1 1 0000 100k 3333 111 0000 3 10303 2222 11 0000 2 10202 1111 1 0000 1 11101 0000 0 0 10000
introducción 18 - Del mismo modo, las probabilidades de transición kpij pueden ser calculadas como los elementos de las matrices 1П = (1П)k. - Las probabilidades kpii para i = j se denominan probabilidades de auto-retorno, el sistema regresa al estado inicial, estas probabilidades se encuentran en la diagonal principal de las matrices estocásticas. - Tanto las probabilidades absolutas iniciales Ap0(j) (elementos de 0π), como las probabilidades 1pij (elementos de 1П) pueden ser determinadas a partir de las magnitudes fisicoquímicas mj que caracterizan al estado: llil jij ij n ll j A m m p m m jp 1 0 - Donde, αij indica la adyacencia entre los dos estados (átomos, aminoácidos, nucleótidos). - De lo anterior se desprende que es posible derivar de la CM ciertos números que la caracterizan, ya que dependen de: i) los estados presentes, ii) de su interconexión caracterizada por αij, y iii) de la tendencia del sistema a ocupar dichos estados caracterizados por su magnitud fisicoquímica mj. - Por tanto de ser aplicada la CM a un sistema molecular como los descritos anteriormente dichos números podrán ser usados como descriptores moleculares (Dh) para encontrar modelos QSAR de una actividad biológica (Actv) dada preferentemente de tipo linear: bDaDaDaDaDaActv hh ... 44332211 - Un ejemplo, en el sistema a) en el cual la CM representa a una molécula donde la posición de los electrones (parámetro) puede
introducción 19 estar alrededor de varios átomos (estados) el aspecto i) está ligado al tipo y cantidad de átomos en la molécula, el aspecto ii) a la presencia de enlaces específicos entre los átomos y el aspecto iii) a la electronegatividad con que cada átomo atrae los electrones. - Entre los números introducidos en este trabajo con ese fin se encuentran: las probabilidades absolutas en sí (que pueden ser usadas como descriptores moleculares locales solamente). los momentos espectrales de la matriz estocástica (SRπk), suma de las probabilidades de autoretorno: jjj k k k k SR pTrTr 1 las entropías de la distribución de probabilidades (Θk): jpjp k A jk A klog y magnitudes medias ([m]k) calculadas como sumas ponderadas de probabilidades absolutas: j jk A kmjpm
RESULTADOS Y DISCUSIÓN
resultados y discusión 29 En este acápite se presentarán todos los resultados obtenidos en forma de artículos de investigación ya publicados por el autor. Los 7 artículos presentados (6 artículos de revista y 1 capítulo de libro) están agrupados de acuerdo al objetivo específico que cumplimentan. Para cada artículo se presenta una breve sección explicativa en español de su importancia y los resultados alcanzados. En el apartado “4. Publicaciones” de esta Tesis se adjuntan las publicaciones correspondientes en el idioma en que fueron publicadas.
resultados y discusión 30 2.1. Estudio mt-QSAR y RC de inhibidores HGMR Medicamentos eficaces como las estatinas o los ácidos mevínicos son HGMRIs, inhibidores de la enzima HMGR; que limita la velocidad de biosíntesis del colesterol. Sin embargo, el elevado número de posibles compuestos candidatos a ensayar crea la necesidad de desarrollar modelos QSAR para guiar la síntesis de los HMGRIs. Los modelos QSAR desarrollados con anterioridad en este sentido (ver trabajo de revisión bibliográfica presentado en la Introducción) presentan dos problemas principales: son aplicables únicamente a series homogéneas de compuestos y no tienen en cuenta la quiralidad de los compuestos. En este trabajo, se propone por primera vez un modelo QSAR para una serie grande y heterogénea de HMGRIs. El modelo se basa en TIs de las estructuras moleculares. Además proponemos la primera red compleja que describe las relaciones de similitud entre HGMRIs usando como entrada las predicciones de este modelo. Ambos, el modelo QSAR y la red, fueron usados para predecir las diferencias en actividad debido a la quiralidad en más de 1600 isómeros quirales de HMGRIs no explorados experimentalmente. También se presentó una versión reducida de esta red (Componente Gigante) que contiene el conjunto más representativo de los compuestos quirales candidatos a ser ensayados como HMGRIs. El trabajo sugiere una nueva aplicación combinado el estudio QSAR y las RC.
resultados y discusión 31 Como conclusión podemos destacar que en este trabajo se propuso por primera vez un modelo QSAR basado en índices topológicos quirales y no quirales de una lista heterogénea de inhibidores de la HMGR. Este modelo se usó para la predicción de la actividad de la HMGR de nuevos inhibidores quirales. Además, la comparación entre isómeros quirales fue realizada mediante redes complejas quirales y no quirales. Estas redes se construyeron mediante previas predicciones QSAR. Una ventaja de este método es la posibilidad que ofrece para buscar nuevos compuestos quirales que no fueron caracterizados experimentalmente. Otra ventaja es que el empleo de este modelo QSAR es una guía importante para los experimentos sintéticos, para la búsqueda de nuevos candidatos a inhibidores de la HMGR y al mismo tiempo, la disminución de costes que eso supone.
resultados y discusión 32 2.2. Estudios QSAR de inhibidores de la GSK-3 α En general, los inhibidores de distintas isoformas de la GSK-3 son candidatos interesantes para el desarrollo de compuestos antiAlzheimer. Los inhibidores GSK-3 también son de interés como compuestos antiparasitarios activos contra Plasmodium falciparum, Trypanosoma brucei y Leishmania donovani, los agentes causantes de la malaria, tripanosomiasis africana humana y la leishmaniosis. Esto ha provocado una búsqueda activa de potentes y selectivos inhibidores de GSK-3. En este sentido, los estudios QSAR podrían desempeñar un papel importante en el descubrimiento de estos inhibidores de GSK-3. Por esta razón, en este trabajo hemos desarrollado modelos QSAR para los inhibidores de la GSK-3α. En el estudio hemos usado ADL y ANN de casi 50.000 casos con más de 700 diferentes inhibidores de GSK-3α. Los compuestos fueron obtenidos desde la base de datos ChEMBL, en total se utilizaron más de 20.000 moléculas diferentes para desarrollar los modelos QSAR. El modelo clasificó correctamente 237 de 275 compuestos activos (86,2%) y 14.870 de 15.970 compuestos inactivos (93,2%) en la serie de entrenamiento. El porcentaje general de buena clasificación fue de 93,0%. La validación del modelo se llevó a cabo mediante una serie de predicción externa. En esta serie el modelo clasifica correctamente 458 de 549 (83,4%) y 29.637 de 31.927 casos control (83,4%). El porcentaje general de buena clasificación fue del 92,7%. En este trabajo, proponemos tres tipos de ANN no lineales y se muestra otro modelo alternativo a los ya existentes en la literatura,
resultados y discusión 33 como ADL. El mejor modelo obtenido fue una ANN lineal (LNN): LNN: 236:236-1-1:1 que tuvo porcentaje de buena clasificación del 96%. Además, hicimos un estudio de los diferentes fragmentos que existen en las moléculas de la base de datos con el fin de ver cuales tenían más influencia en la actividad. Todo esto puede ayudar a diseñar nuevos inhibidores de GSK-3α. Como modelos no lineales calculamos las ANN utilizando los descriptores calculados con el DRAGON viendo que estos modelos eran otros métodos alternativos para estudiar la actividad de distintas familias de moléculas. Otra parte de este trabajo fue el estudio de las contribuciones de los fragmentos usando un modelo QSAR, lo que nos puede ayudar al diseño de los mejores inhibidores de GSK-3α, y poder así luego sintetizarlos en el laboratorio, eliminando la síntesis de moléculas al azar ya que existe la posibilidad de que la mayoría sean inactivos.
resultados y discusión 34 2.3. Uso de Entropía en estudio mt-QSAR de inhibidores de la GSK-3 El desarrollo de modelos QSAR usando índices moleculares simples parece ser una prometedora técnica alternativa o complementaria al Docking fármacos-proteínas. Casi todas las técnicas QSAR se basan en el uso de descriptores moleculares, que son series numéricas que codifican la información química útil y que permiten correlacionar las propiedades estructurales y biológicas. Entropía de Shannon es uno de los parámetros más importantes con el fin de codificar la información estructural sobre los estudios QSAR. En este sentido, nuestro grupo de investigación ha introducido una nueva serie de índices estocásticos que pueden ser calculados con la técnica MARCH-INSIDE. El método MARCHINSIDE se basa en el uso de CM para calcular probabilidades absolutas de la distribución de las diferentes propiedades atómicas dentro de la estructura molecular. Podemos aplicar la fórmula de Shannon a estas probabilidades absolutas para calcular los parámetros de la entropía de la distribución de las propiedades atómicas en la molécula. En este trabajo vamos a explorar el potencial de la técnica MARCH-INSIDE para buscar un modelo mt-QSAR para una serie heterogénea de compuestos inhibidores de la GSK-3. En el primer paso, los descriptores moleculares antes mencionados fueron calculados para una gran serie de compuestos activos/inactivos. Posteriormente se utilizó ADL para ajustar la función de clasificación. El modelo mt-QSAR fue validado después
resultados y discusión 35 con una serie de predicción externa mediante la técnica de resustitución. En conclusión, podemos considerar la técnica MARCH-INSIDE como una buena alternativa para el desarrollo de nuevos inhibidores de la GSK-3 ya que clasifica correctamente a muchos inhibidores con diferente estructura molecular.
CONCLUSIONES
conclusiones 45 Exponemos las conclusiones específicas, en correspondencia con los objetivos trazados, agrupadas en tres grupos, dada la naturaleza de los estudios realizados: 1) estudios QSAR/mt-QSAR de inhibidores de enzimas, 2) estudios de RC, 3) estudio mt-QSAR y de RC de múltiples enzimas: Conclusiones específicas: 1.1. Se pudo desarrollar modelos QSAR para la predicción HGMRIs. 1.2. Desarrollamos modelos mt-QSAR para la predicción de inhibidores de la GSK-3. 1.3. Pudimos desarrollar modelos mt-QSAR para la predicción de compuestos antivirales. 1.4. Según la revisión realizada no existen modelos mt-QSAR en análogos a Vitamina D. Por lo que podemos concluir que el desarrollo de modelos mt-QSAR y RC para estos compuestos es un campo con perspectivas futuras. 2.1. Desarrollamos una metodología QSAR de construcción de RC de compuestos HGMRIs. 2.2. Desarrollamos una metodología de construcción de RC de compuestos anti-virales a partir del modelo mt-QSAR. Conclusión general: Podemos concluir que los nuevos modelos QSAR desarrollados son aplicables a la predicción de la actividad biológica de compuestos contra una única diana o múltiples dianas de interés en química farmacéutica, microbiología y parasitología. Además, podemos concluir que es posible
conclusiones 46 desarrollar nuevas metodologías de construcción de RC de estos compuestos en estudios de Bioinformática a partir de modelos QSAR o mt-QSAR.
PUBLICACIONES
publicaciones 49 A continuación se presenta un ANEXO con las publicaciones que se recogen en la Tesis siguiendo el orden establecido en la misma.
Current Drug Metabolism, 2010, 11, 1389-2002/10 $55.00+.00 © 2010 Bentham Science Publishers Ltd. QSAR & Complex Network Study of the HMGR Inhibitors Structural Diversity Isela García*, Yagamare Fall Diop and Generosa Gómez Department of Organic Chemistry, University of Vigo, Spain Abstract: Efficient drugs such as statins or mevinic acids are inhibitors of the rate-limiting enzyme of cholesterol biosynthesis, 3hydroxy-3-methyl-glutaryl coenzyme A reductase (HMGR), an enzyme responsible for the double reduction of 3-hydroxy-3-methylglutaryl coenzyme A. These compounds promoted the synthesis and evaluation of new inhibitors for HMGR, named HMGRIs. The high number of possible candidates creates the necessity of Quantitative Structure-Activity Relationship models in order to guide the HMGRI (3-hydroxy-3-methyl-glutaryl coenzyme A inhibitor) synthesis. In this work, we revised different computational studies for a very large and heterogeneous series of HMGRIs. First, we revised QSAR studies with conceptual parameters such as flexibility of rotation, probability of availability, etc; we then used the method of regression analysis; and QSAR studies in order to understand the essential structural requirement for binding with receptor. Next, we reviewed 3D QSAR, CoMFA and CoMSIA with different compounds to find out the structural requirements for 3-hydroxy-3-methylglutaryl-CoA reductase (HMGR) inhibitory activity Keywords: QSAR; Complex network, Lipid-lowering agent, Cholesterol level, Atherosclerotic disease, Antiparasite drug, Trypanosoma cruzi, Chagas’ disease, 3-hydroxy-3-methyl-glutaryl coenzyme A reductase, CoMSIA, COMFA, topological indices. 1. INTRODUCTION Hypercholesterolemia (the level of plasma cholesterol, and in particular LDL-associated cholesterol) is well-known as the main risks factor in atherosclerotic, the degenerative disease underlying myocardial infarction and stroke, and coronary heart diseases [1, 2]. In western countries, this is the leading cause of death, being more common than all cancers and leukemias combined. Briefly, atherosclerotic lesions develop as follows: 1. Small mechanical lesions in the vascular endothelium (the innermost layer of blood vessel walls) allow leakage of blood plasma into the muscular layers beneath. Formation of these small leakages is thought to be promoted by high blood pressure. 2. The lipoproteins that leaked into the tissue are degraded. Because of its low solubility, cholesterol released from degraded lipoproteins precipitates. 3. The cholesterol particles trigger invasion of phagocytic cells and in this way contribute to triggering inflammation, which in turn increases the tissue damage and turns the small, potentially reparable defects of the vessel wall into large lesions. Clinical studies with lipid-lowering agents have established that the decrease of high serum cholesterol levels reduces the incidence of cardiovascular mortality. Cholesterol occurs in animals but not in plants or fungi. These have similar sterols for similar purposes, which however cannot be converted into cholesterol. Therefore, it is essential for animals (particularly for plant-feeding ones, such as sheep, goat and vegetarians) to have a pathway for cholesterol biosynthesis. Statins and mevinic acids are two efficient drugs known as inhibitors of the rate-limiting enzyme of cholesterol biosynthesis, 3-hydroxy-3methyl-glutaryl coenzyme A reductase (HMGR), enzyme responsible for the double reduction of 3-hydroxy-3-methyl-glutaryl coenzyme A into mevalonic acid. Synthesis starts with Acetyl-CoA in the mitochondrion, which is used to synthesize HMG-CoA. These reactions also occur in ketogenesis. However, while the entire process of ketogenesis occurs in the mitochondrion, the formation of HMG-CoA in sterol synthesis occurs in the cytosol. All subsequent steps occur in the smooth endoplasmic reticulum. HMG-CoA reductase reduces HMG-CoA to mevalonate, which in turn is con- *Address correspondence to this author at the Department of Organic Chemistry, University of Vigo, Spain; Tel:Fax: E-mail:
[email protected] verted into various isoprene compounds. Several rounds of polymerization lead to the linear hydrocarbon molecule squalene, which is then converted into lanosterol. Subsequent modifications lead to cholesterol. The reactions of the synthetic pathway are shown in Fig. (1). The reaction catalyzed by HMG-CoA reductase is the first committed step, which means that from this point onwards the substrates have no other option than becoming a sterol. Therefore, HMG-CoA reductase is the main target of regulatory mechanisms, which in turn is being exploited in pharmacotherapy. The biosynthesis of cholesterol (the same as for many other metabolites) is regulated both by allosteric control of enzyme activity and by enzyme induction. The committed step is the reduction of HMG-CoA to mevalonate, and it therefore makes sense that HMGCoA reductase is the primary target of regulation. The drug control of this enzyme is efficient in reducing the levels of cholesterol in plasma. [3, 4] The structure of statins and its derivatives is characterized by the desmethylmevalonic acid or by the lactone. The pharmacophore is connected to a lipophilic ring, such as hexahydronaphthalene, indole, pyrrole, pyrimidine or quinine, by a linking element (a two-carbon spacer). The biologically active form of mevinic acids is represented by the open chain hydroxyl-acid, which mimics the HMGR natural substrate [5]. Fig. (1). Pathway. The genetic regulation of cholesterol synthesis has been worked out only in the last decade, and it operates by quite a neat mechanism. Several membrane proteins participate in it (Fig. 2a): The Sterol Response Element Binding Protein (SREBP), the SREBP Cleavage Activating Protein (SCAP), and two SREBP-specific proteases (S1P and S2P). SCAP is the actual cholesterol sensor that changes conformation in response to changes of the cholesterol concentration within the ER membrane. If cholesterol is high,
Current Drug Metabolism, 2010, Vol. 11, No. García et al. [47] Vullo, A.; Frasconi, P.Prediction of protein coarse contact maps. J. Bioinform. Comput. Biol., 2003, 1(2), 411-31. [48] Lange, B. M.; Ghassemian, M.Comprehensive post-genomic data analysis approaches integrating biochemical pathway maps. Phytochemistry, 2005, 66(4), 413-51. [49] Chou, K. C.; Cai, Y. D.Predicting protein-protein interactions from sequences in a hybridization space. J. Proteome Res., 2006, 5(2), 316-22. [50] Yu, X.; Lin, J.; Zack, D. J.; Qian, J.Computational analysis of tissue-specific combinatorial gene regulation: predicting interaction between transcription factors in human tissues. Nucleic Acids Res., 2006, 34(17), 4925-36. [51] Margolin, A. A.; Nemenman, I.; Basso, K.; Wiggins, C.; Stolovitzky, G.; Dalla Favera, R.; Califano, A.ARACNE: an algorithm for the reconstruction of gene regulatory networks in a mammalian cellular context. BMC Bioinformatics, 2006, 7 Suppl 1(S7. [52] González-Díaz, H.; González-Díaz, Y.; Santana, L.; Ubeira, F. M.; Uriarte, E.Proteomics, networks and connectivity indices. Proteomics, 2008, 8(750-778. [53] Estrada, E.Protein bipartivity and essentiality in the yeast proteinprotein interaction network. J. Proteome Res., 2006, 5(9), 2177-84. [54] Newman, M. E.Finding community structure in networks using the eigenvectors of matrices. Phys. Rev. E Stat Nonlin Soft Matter Phys, 2006, 74(3 Pt 2), 036104. [55] Kleinberg, J. M.Navigation in a small world. Nature, 2000, 406(6798), 845. [56] González-Díaz, H.; Prado-Prado, F.Unified QSAR and NetworkBased Computational Chemistry Approach to Antimicrobials, Part 1: Multispecies Activity Models for Antifungals. J. Comput. Chem., 2008, 29(656-657. [57] Prabhakar, Y. S.Analysis of Tetrahedral Carbon in QSAR Studies. J. Chem. Inf. Comput. Sci., 1999, 39(650-653. [58] Mutsukado, M.; Suzuki, M.A New HMG-CoA Reductase Inhibitor, NK-104: QSAR Studies. Kagaku Toronkai, Kozo Kassei Sokan Shinpojiumu Koen Yoshishu, 2000, 23-28(192-193. [59] Saxena, M.; Soni, L. K.; Gupta, A. K.; Wakode, S. R.; Saxena, A. K.; Kaskhedikar, S. G.Development of pharmacophoric model of condensed pyridine and pyrimidine analogs as hydroxymethyl glutaryl coenzyme A reductase inhibitors. Indian J. Biochem & Biophys. (IJBB), 2006, 43(1), 32-36. [60] Web., C. S. o. t. [61] Motoc, I.; Sit, S. Y.; Harte, W. E.; Balasubramanian, N.; Wright, J. J.3-Hydroxy-3-Methylglutaryl-Coenzyme A Reductase: Molecular Modeling, Three-Dimensional Structure-Activity Relationships, Inhibitor Design. QSAR & Combinatorial Sci., 2006, 10(1), 30-35. [62] Garcia, I.; Munteanu, C. R.; Fall, Y.; Gomez, G.; Uriarte, E.; Gonzalez-Diaz, H.QSAR and complex network study of the chiral HMGR inhibitor structural diversity. Bioorg. Med. Chem., 2009, 17(1), 165-75. [63] Istvan; Deisenhofer. Science, 2001, 292(1160-1164. [64] Qing, Y. Z.; Jian, W.; Xin, X.; Guang, F.; Yang, Y. L. R.; Jun, J. L.; Hui, W.; Yu, G.Structure-Based Rational Quest for Potential Novel Inhibitors of Human HMG-CoA Reductase by Combining CoMFA 3D QSAR Modeling and Virtual Screening. J. Comb. Chem., 2007, 9(131-138. [65] Ramasamy Thilagavathi, R.; Kumar, R.; Aparna, V.; Sobhia, M. E.; Gopalakrishnan, B.; Chakraborti, A. K.Three-dimensional quantitative structure (3D QSAR) activity relationship studies on imidazolyl and N-pyrrolyl heptenoates as 3-hydroxy-3-methylglutarylCoA reductase (HMGR) inhibitors by comparative molecular similarity indices analysis (CoMSIA). Bioorg. Med. Chem. Lett., 2005, 15(1027-1032. [66] González-Díaz, H.; Munteanu, C. R. Topological Indices for Medicinal Chemistry, Biology, Parasitology, Neurological and Social Networks, Transworld Research Network: Kerala, India 2010. [67] Kuroda, M.; Endo, A.Biochim. Biophys. Acta, 1977, 486(70. [68] Hill, T.; Lewicki, P. STATISTICS Methods and Applications. A Comprehensive Reference for Science, Industry and Data Mining, StatSoft: Tulsa 2006 [69] Lilien, R. H.; Farid, H.; Donald, B. R.Probabilistic disease classification of expression-dependent proteomic data from mass spectrometry of human serum. J. Comput. Biol., 2003, 10(6), 925-46. [70] James, A.; Hanley, B. J. M.Radiology, 1982, 143(29. [71] Gómez, G.; Rivera, H.; García, I.; Estévez, L.; Fall, Y.The furan approach to carbocyclic systems. Synthesis of cyclohexane derivatives from butenolides through an intramolecular Michael addition Tetrahedron Lett., 2005, 46(35), 5819-5822 [72] Boccaletti, S.; Latora, V.; Moreno, Y.; Chavez, M.; Hwang, D. U.Complex networks: Structure and dynamics. Physics Reports, 2006, 424(175-308. [73] Bianconi, G.; Barabasi, A. L.Bose-Einstein condensation in complex networks. Phys. Rev. Lett., 2001, 86(24), 5632-5. [74] Bornholdt, S.; Schuster, H. G. Handbook of Graphs and Complex Networks: From the Genome to the Internet, WILEY-VCH GmbH & CO. KGa.: Wheinheim 2003. [75] Thilagavathi, R.; Kumar, R.; Aparna, V.; Sobhia, M. E.; Gopalakrishnan, B.; Chakraborti, A. K.Three-dimensional quantitative structure (3-D QSAR) activity relationship studies on imidazolyl and N-pyrrolyl heptenoates as 3-hydroxy-3-methylglutarylCoA reductase (HMGR) inhibitors by comparative molecular similarity indices analysis (CoMSIA). Bioorg. Med. Chem. Lett., 2005, 15(4), 1027-32. Received: )HEUXDU\ Revised: $SULO Accepted: $SULO
2666 Current Pharmaceutical Design, 2010, 16, 2666-2675 1381-6128/10 $55.00+.00 © 2010 Bentham Science Publishers Ltd. QSAR, Docking, and CoMFA Studies of GSK3 Inhibitors Isela García*, Yagamare Fall and Generosa Gómez Department of Organic Chemistry, University of Vigo, Spain Abstract: GSK-3 inhibitors are interesting candidates to develop anti-Alzheimer compounds. GSK-3 are also interesting as antiparasitic compounds active against Plasmodium falciparum, Trypanosoma brucei, and Leishmania donovani; the causative agents for Malaria, African Trypanosomiasis and Leishmaniosis. The high number of possible candidates creates the necessity of Quantitative Structure-Activity Relationship models in order to guide the GSK3 (Glycogen Synthase Kinase 3 inhibitor) synthesis. In this work, we revised different computational studies for a very large and heterogeneous series of GSK-3Is. First, we revised QSAR studies with conceptual parameters such as flexibility of rotation, probability of availability, etc. We then used the method of regression analysis and QSAR studies in order to understand the essential structural requirement for binding with receptor. Next, we reviewed 3D-QSAR, CoMFA and CoMSIA with different compounds to find out the structural requirements for GSK-3 inhibitory activity. Keywords:QSAR, Alzheimer, SAR, parasitic, fungi. INTRODUCTION The enzyme Glycogen Synthase Kinase 3 (GSK-3), so called because of its implication in the phosphorylation of glycogen synthase, is a serine/threoline kinase involved in many cellular processes. Glycogen Synthase Kinase-3 (GSK-3) is a serine-threonine kinase encoded by two isoforms in mammals, termed GSK-3 and GSK-3 [1]. Initially GSK-3 was implicated in muscle energy storage and metabolism, but since its cloning, a more generalized role in cellular regulation has emerged, highlighted by the wide array of substrates controlled by this enzyme that includes cytoplasmic proteins and nuclear transcription factors. GSK-3 targets encompass proteins implicated in Alzheimer´s disease (AD), neurological disorders, in the wnt and insulin signaling pathway, glycogen and protein synthesis, regulation of transcription factors [2], embryonic development, cell proliferation and adhesion, tumorigenesis, apoptosis [3], circadian rhythm, etc. GSK-3 knock-out mice die in utero [4], whereas GSK-3 knock-out mice are viable and display improved glucose tolerance in response to glucose load and elevated hepatic glycogen storage and insulin sensitivity [5, 6]. Alzheimer´s disease [7] is a serious and degenerative disorder that explains the gradual loss of neurons, and in spite of the efforts realized by the big pharmacists of the world, is still not a very clear reason of this pathology, because at present, it is the most recent reason of dementia in main elders. The fundamental characteristic of Alzheimer´s disease is the presence in the brain of two injuries: the Neurofibrillary Tangles (NFTs) (Fig. 1), that are formed by paired helical filaments (PHF) whose main component is Tau Protein kinase (TPK), and insoluble -amyloid (A) plaques (Fig. 2) that are associated with active microglia [5, 8]. NFTs are composed of hyper-phosphorylated forms of the microtubule-associated protein tau, whereas A is derived from the proteolytic cleveage of - amyloid precursor protein (APP). Active GSK-3 appears in neurons with pre-tangle changes [9] and there is an increased GSK3 activity in the frontal cortex in AD as evidenced by immunoblotting for GSK-3 phosphorylated at Tyr216 [10]. GSK3 expression is upregulated in the hippocampus of AD patients and in post-synaptosomal supernatants derived from AD brain, although the latter study reports that there is no increase in GSK-3 enzymatic activity [5]. The functions of GSK-3 and its implication in various human diseases have triggered an active search for potent and selective GSK3 inhibitors [11] in the last years, as shown in Fig. 3. Studies of *Address correspondence to this author at the Department of Organic Chemistry, University of Vigo, Spain; Tel: +34986813679; Fax: +34986812262; E-mail:
[email protected] Fig. (1). Neurofibrillary tangles (NFTs). Fig. (2). -amyloid (A) plaques.
QSAR, Docking, and CoMFA studies of GSK3 inhibitors Current Pharmaceutical Design, 2010, Vol. 16, No. 24 2667 Fig. (3). Search of potent and selective GSK-3 inhibitors. GSK-3 homologues in various organisms have revealed physiological roles for the enzyme in differentiation, cell fate determination, and spatial patterning to establish bilateral embryonic symmetry. Purified GSK-3 and GSK-3 exhibit similar biochemical and substrate properties [12, 13], and is known that in the phosphorylation of TPK/ Tau Protein Kinase takes part actively Glycogen Synthase Kinase 3 (GSK-3), which not only plays a fundamental role in the synthesis of the glycogen (where it was identified by the first time), but it is very important in several processes as cellular signs, metabolic control, embryogenesis, cellular death and oncogenesis [14], and it is related to a wide range of neurodegenerative [15] diseases, bipolar mood disorders [16] and diabetes, hence the inhibition of this enzyme is accepted as a promising therapeutic strategy. In 1988 Ishiguro and col. [17] isolated one enzyme when they were studying an extract of the brain and noticed the presence of paired helical filaments of Tau Protein Kinase, typical injury of Alzheimer´s disease. TPKI and TPKII are the two kinases implied in this process and they found that TPKI has an identical structure to GSK-3. At this moment, there is an increasing interest in the evaluation of kinases from unicellular parasites as targets for potential new anti-parasitic drugs. The evolutionary difference between unicellular kinases and their human homologues might be sufficient to allow the design of parasite-specific inhibitors. The Plasmodium falciparum genome contains 65 genes that encode kinases, including three forms of GSK-3. An initial study showed that P. falciparum exports PfGSK-3 to the cytoplasm of host erythrocytes (which are devoid of GSK-3), where it colocalizes with parasite-generated membrane structures known as Maurer´s clefts. The function of PfGSK-3 is unknown, but the presence of PfCK1, a CK1 homologue, in infected red blood cell supports the hypothesis that both kinases play a role in regulating the strong circadian rhythm of the parasite, which is responsible for the circadian fevers that are the characteristics of this infectious disease [18]. On the other hand, the vector-borne parasitic disease African trypanosomiasis, caused by members of the Trypanosoma brucei complex, is a serious health threat. It is estimated that 300,000 to 500,000 humans in sub-Saharan Africa are infected. If the disease is left inadequately treated, it often has a fatal outcome. Once infection is established, safe and effective therapy is critically important, yet it has been difficult to achieve. Despite the critical need, the available therapies are becoming less satisfactory due to the rising level of resistance to the available drugs, the long period of treatment required to achieve a cure, and the unacceptable and sometimes severe adverse effects associated with current therapies [19]. An urgent priority is to identify and validate new targets for the development of safe, effective, and inexpensive therapeutic alternatives. Compounds that inhibit T. brucei GSK-3 activity and not host GSK-3 might be required for therapy for pregnant women and infants, in that GSK-3 regulates proteins critical in development, such as the wnt gene product. However, optimization of the selectivity of drug candidates for parasite kinases becomes an issue due to the highly conserved amino acids and protein conformation of the catalytic domains [20-23]. Understanding the differences in the substrate binding properties and the three-dimensional structures between mammalian and parasite GSK-3 enzymes is important for the optimization of selected target inhibitors for drug development [24, 25]. In this sense, the development of QSARs using simple molecular indices appears to be a promising alternative or complementary technique to drug-protein docking, high-throughput screening and combinatorial chemistry techniques. Almost all QSAR techniques are based on the use of molecular descriptors, which are numerical series that codify useful chemical information and enable correlations between statistical and biological properties [26-28]. Many recent works, for instance those published by González-Díaz et al. (to cite only one example), illustrated that the parameters used to seek QSAR-like predictive models of low-weight molecules may also be used to predict properties of proteins, RNAs, protein interaction networks (PINs), cerebral cortex, disease spreading, and other more complex systems [29-46]. In any case, a large number of examples have been published in which the use of molecular descriptors has become a rational alternative to massive synthesis and screening of compounds in medicinal chemistry [47, 48]; in the Fig. 4 we can see progress of GSK-3-QSAR studies in the last years. Indeed, experience has shown that the use of models fitted with large data sets of chemicals works as well as the use of models built from a series of homologous compounds and is also a more general method that can be applied in a broad spectrum of cases. The principal deficiency in the use of some molecular indices concerns their lack of physical meaning. In this respect, the introduction of novel molecular indices must obey physicochemical laws in order to ensure a theoretically rigorous interpretation of the results. In the work described here, we review and comment different theoretical studies about GSK-3 inhibitors in the last years. Fig. (4). Progress of GSK-3-QSAR studies in the last years.
2668 Current Pharmaceutical Design, 2010, Vol. 16, No. 24 García et al. 1. 3D-QSAR studies on thiadiazolidinone derivatives as GSK-3 inhibitors. 2. 3D-QSAR and Docking studies of selective GSK-3 inhibitors. 3. QSAR modeling of the inhibition of GSK-3. 4. Docking and 3D-QSAR for pyrimidin-2-amines as GSK-3 inhibitors. 5. 3D-QSAR for GSK-3 inhibition by indirubin analogues. 6. 3D-QSAR and Docking for bisarylmaleimides as GSK-3, CDK-2 and CDK-4 inhibitors. 7. Construction of the pharmacophore model of GSK-3 inhibitors. 8. 3D-QSAR and Docking of pyrazolopyrimidines as GSK-3 inhibitors. 9. 2D image-based approach for prediction of GSK-3 inhibitors. 10. Free-Wilson QSAR Analysis to predict kinase selectivity profiles. 11. CoMFA and Docking for pyrazolo[3,4-b]pyrid[az]ines as GSK-3 inhibitors. 12. 2D and 3D-QSAR models for prediction of indirubin derivatives GSK-3 inhibitors 13. Multi-target QSAR in silico screening for GSK-3 inhibitors. REVIEW AND DISCUSSION 1. 3D-QSAR Studies on Thiadiazolidinone Derivatives as GSK-3 Inhibitors In this article, Ana Martinez [49] reported a biological study following a SAR study about 2,4-disubstituted thiadiazolidinones (TDZD) (see Table 1), compounds that were described as the first ATP-noncompetitive GSK-3 inhibitors, where different structural modifications in the heterocyclic ring aimed to test the influence of each heteroatom. They synthesized various compounds such as hydantoins, dithiazolidindiones, rhodanines, maleimides, and triazoles and were screened as GSK-3 inhibitors. A CoMFA analysis was also performed highlighting the molecular electrostatic field connection in the interaction of TDZDs with GSK-3. CoMFA models were calculated for each molecular field considered alone or in combination, and the best PLS analysis was obtained by combining steric and electrostatic fields, leading to a good correlation between the IC50 values predicted from the principal components extracted from these two fields and the experimental data (r2 = 0.922, q2 = 0.654, and n = 5). Since the relative contributions of steric and electrostatic molecular field were 42% and 58%, respectively. Favorable steric regions are found around the aromatic substituent attached to N4 as well as in the vicinity of N2. An increase of the steric field in the region between N2 and S is detrimental for the inhibitory activity, and a negatively charged region around the ring is favorable for activity. Moreover, first mapping studies indicate two binding modes which in turn might imply relevant differences in the mechanism that undelie the inhibitory activity of TDZDs. 2. 3D-QSAR and Docking Studies of Selective GSK-3 Inhibitors Bureau, R. [50] et al. in this paper carried out a 3D-QSAR (CoMFA) study with several GSK-3inhibitors. The cocrystallographic data of GSK-3vs 3-anilino-4-arylmaleimide were suitable to compare 3D-QSAR results with experimental intermolecular interactions. CoMFA analysis served to start the study of a new compound, a thieno[2,3-b]pyrrolizinone derivative (see Table 1) as GSK-3 inhibitor, because this study was not about the interactions registered in the active site. This comparison based on docking and simulation approaches allowed to confirm one preferential orientation of this ligand inside the active site, explaining the relationship with the reference 3-anilino 4-arylmaleimide derivatives and its biological affinity. 3. QSAR Modeling of the Inhibition of GSK-3 Katritzky, A. R. [51] et al., made a Quantitative StructureActivity Relationship (QSAR) model of the in vitro biological activity (pIC50) of 277 inhibitors of Glycogen Synthase Kinase-3 (GSK-3) (3-anilino-4-arylmaleimides), calculating different molecular descriptors by CODESSA PRO technique, geometrical, topological, quantum mechanical, and electronic descriptors. The linear (multilinear regression) and nonlinear (artificial neural network) models obtained, linked the structures to their reported activity pIC50. Each multilinear model was verified by leave-one-out and internal validation methods confirmed the correct prediction of the inhibitory activity of 3-anilino-4-arylmaleimides (see Table 1), but was found that the Artificial Neural Network (ANN), which was built for all the data points, presented superior prediction over the multilinear models. These studies gave an insight into the dominant role played by the electrostatic, bonding, and steric interactions on the modulation of the inhibitory activity, hence the nature of GSK-3 inhibitor interaction was found to be electrostatic. 4. Docking and 3D-QSAR for Pyrimidin-2-amines as GSK-3 Inhibitors In this paper, Guo [52] et al. carried out a study of Glycogen Synthase Kinase 3 (GSK-3) inhibition. Molecular docking and 3DQSAR approaches were used for the study of the interaction mode of a series of N-phenyl-4-pyrazolo[1,5-b]pyridazin-3-ylpyrimidin2-amine compounds (see Table 1) with human GSK-3. In the 3DQSAR studies, the molecular alignment and conformation determination were of special importance. Flexible docking (AutoDock3.0.5) was used for the determination of ‘active’ conformation and molecular alignment. Comparative molecular field analysis (CoMFA) and comparative molecular similarity indices analysis (CoMSIA) were used to carried out 3D-QSAR models of 80 Nphenyl-4-pyrazolo[1,5-b]pyridazin-3-ylpyrimidin-2-amine compounds. The r2 values of CoMFA were 0.870 and the r2 values of CoMSIA were 0.861. The predictive ability of these models was validated by 10 compounds of the test set. Mapping these models back to the topology of the active site of GSK-3 led to a better understanding of the vital N-phenyl-4-pyrazolo[1,5-b]pyridazin-3ylpyrimidin-2-amines-GSK-3 interactions. The results showed that combination of ligand-based and receptor-based modeling is a powerful approach to develop 3D-QSAR models. 5. 3D-QSAR for GSK-3 Inhibition by Indirubin Analogues Yongjun Jiang [53] et al. studied Indirubin analogues that showed favorable inhibitory activity targeting GSK-3, which is closely related to the property and position of substituents. Two methods were used to build 3D-QSAR models for indirubin derivatives. The conventional 3D-QSAR (ligand-based) studies were performed based on the lower energy conformations employing atom fit alignment rule. The receptor-based 3D-QSAR models were also derived using bioactive conformations obtained by docking the compounds (see Table 1) to the active site of GSK-3. Conclusions of models based on two methods are similar and reliable. The ligand based CoMFA model had cross-validate coefficient q2 of 0.790 with five components and conventional r2 of 0.923, the CoMSIA model with q2of 0.780 for seven components and r2 of 0.916 are obtained. These values indicate that the models were robust. The receptor-based CoMFA model had cross-validation coefficient (q2) of 0.563 and regression coefficient (r2) of 0.897 for four components. q2and r2 for CoMSIA model were 0.766 and 0.908 with five components using the combination of steric + electric + hydrophobic fields, respectively. The results indicated that both ligand-based and receptor-based were feasible tools to build 3D-QSAR models.
QSAR, Docking, and CoMFA studies of GSK3 inhibitors Current Pharmaceutical Design, 2010, Vol. 16, No. 24 2669 Table 1. Summary of QSAR Studies of GSK-3 Inhibitors Author Compounds Structure Study Refs. Martinez, A. TDZD derivatives N NS OO R ´R SAR 3D-QSAR [49] Bureau, R. Pyrrolizinones S N O R ´R 3D-QSAR Docking [50] Katritzky, A. R. Aarylmaleimides NN N R4 R5 R6 N N N R7 XY Z N H N R9 HN R8 O QSAR [51] Guo, Z. Pyrimidines NN NR 6 N N N R CoMSIA CoMFA [52] Jiang, Y. Indirubin derivatives N HNH Y O X W Z R 3D-QSAR [53] Dessalew, N. Maleimides H N R 2 N N O O R 1 3D-QSAR [54] Feng-Chao, J. Maleimides N OO H ´R R CoMFA [55] Bharatam, P. V. Pyrimidines N N N N R3 R1 HN N R2R4 CoMFA CoMSIA [56]
2670 Current Pharmaceutical Design, 2010, Vol. 16, No. 24 García et al. (Table 1) Contd…. Author Compounds Structure Study Refs. Freitas, M. P. Maleimides N OO H ´R R QSAR ANN [57, 58] Sciabola, S. Pyrimidines and Pyrrolopyrazoles N N X 1 X 2 NHHN R 2 R 1 HN N NH R 1 N X 1 X 2 R 2 O X 3 X 4 QSAR [59] Bharatam, P. V. Pyrazoles YX Z N H N N R 2 R 1 H CoMFA Docking [60] Viney, L. Indirubin derivatives N HNH Y O X W Z R 2D-QSAR 3D-QSAR [61] García, I. Maleimides N HNH Y O X W Z R QSAR - 6. 3D-QSAR and Docking for Bisarylmaleimides as GSK-3, CDK-2 and CDK-4 Inhibitors In this paper, Dessalew [54] developed three individual CoMFA models using 36 compounds of bisarylmaleimide series (see Table 1) to correlate with the GSK3, CDK2 and CDK4 inhibitory potencies. Selective Glycogen Synthase Kinase 3 (GSK3) inhibition over Cyclin Dependent Kinases such as Cyclin Dependent Kinase 2 (CDK2) and Cyclin Dependent Kinase 4 (CDK4) was an important requirement for improved therapeutic profile of GSK3 inhibitors. The concepts of selectivity and additivity fields had been employed in developing selective CoMFA models for these related kinases. These models showed a satisfactory statistical significance: CoMFA-GSK3 (r2con, r2cv: 0.931, 0.519), CoMFACDK2 (0.937, 0.563), and CoMFA-CDK4 (0.892, 0.725). Three different selective CoMFA models were then developed using differences in pIC50 values. These three models showed a superior statistical significance: (i) CoMFA-Selective1 (r2con, r2cv: 0.969, 0.768), (ii) CoMFA-Selective 2 (0.974, 0.835) and (iii) CoMFASelective3 (0.963, 0.776). The selective models were found to outperform the individual models in terms of the quality of correlation and were found to be more informative in pinpointing the structural basis for the observed quantitative differences of kinase inhibition. An in-depth comparative investigation was carried out between the individual and selective models to gain an insight into the selectivity criterion. To further validate this approach, a set of new compounds were designed which showed selectivity and were docked into the active site of GSK-3, using FlexX based incremental construction algorithm. The contour maps of the three different selective CoMFA models revealed a way that could be employed to bias the inhibitory activity towards either of the receptors and helped in explaining the structural basis for the experimentally observed differences in activity. The validity of the selective models had been checked using PLS statistical parameters and wes found that it provided a better quality model as compared to the conventional individual models. On the other hand, molecules that showed a good fit in the contours of the selective models were found to exhibit a better selectivity in the inhibitory potencies used to develop the selective model. The final validation criterion designed new molecules using the individual and selective models by making a suitable modification to the basic skeleton in the favorable contour regions.
QSAR, Docking, and CoMFA studies of GSK3 inhibitors Current Pharmaceutical Design, 2010, Vol. 16, No. 24 2671 7. Construction of the Pharmacophore Model of GSK-3 Inhibitors Feng-Chao, J. [55] generated a three dimensional pharmacophore model for glycogen synthase kinase-3 (GSK-3) inhibitors. A dataset consisting of 89 compounds (see Table 1) was selected on the basis of the information content of the structures and activity data as required by the CATALYST system. The better model was selected with a good correlation coefficient, r2 = 0.95. This model was able to predict the activity of other known GSK-3 inhibitors not included in the model generation, and could be used further to identify structurally diverse compounds with desired biological activity by virtual screening. The analysis of the CoMFA study showed that a compound had electrostatic interactions with side chains of K85 and S66 via the NH group in region C around the 4phenyl group. 8. 3D-QSAR and Docking of Pyrazolopyrimidines as GSK-3 Inhibitors Bharatam [56] et al. in this paper studied pyrazolopyrimidines derivatives (see Table 1) as GSK-3 inhibitors. In the present work, they carried out three-dimensional Quantitative Structure Activity Relationship (3D-QSAR) studies on novel class of pyrazolopyrimidine derivatives as GSK-3 inhibitors reported to have improved cellular activity. Docked conformation of the most active molecule in the series, showed desirable interactions in the receptor, was taken as template for alignment of the molecules. CoMFA and CoMSIA models were generated using 49 molecules in training set. By applying leave-one-out (LOO) cross-validation study, r2cv values of 0.53 and 0.48 for CoMFA and CoMSIA, respectively and noncross-validated (r2ncv) values of 0.98 and 0.92 were obtained for CoMFA and CoMSIA models, respectively, which indicates a good internal predictive ability of both models. The predictive ability of CoMFA and CoMSIA models was determined using a test set of 12 molecules, excluded from the model derivation was used, which gave predictive correlation coefficients (r2pred) of 0.47 and 0.48, respectively, indicating good external predictive ability of the model. This field higher r2bs = 0.99 and 0.96 for CoMFA and CoMSIA respectively, further supported the statistical validity of the developed models. Based upon the information derived from CoMFA and CoMSIA contour maps, we have identified some key features that explain the observed variance in the activity and have been used to design new pyrazolopyrimidine derivatives. The designed molecules showed better binding affinity in terms of estimated docking scores with respect to the already reported systems; hence suggesting that newly designed molecules can be more potent and selective towards GSK-3inhibition. 9. 2D Image-Based Approach for Prediction of GSK-3 Inhibitors Freitas, M. P. made different Quantitative Structure-Activity Relationship (QSAR) studies to evaluate GSK-3 inhibitors and their activity. A multivariate image analysis (MIA) [57] applied to QSAR analysis, whose descriptors were derived from pixels of twodimensional (2D) chemical structures that may be built by using any appropriate drawing software, was used to model 17 Glycogen Synthase Kinase 3 (GSK-3) inhibitors. Calibration was carried out using Partial Least Squares (PLS) regression, and an r2 value of 0.93 was obtained for four latent variables. Leave-one-out crossvalidation and a robustness test were performed, and the model proved to be a suitable alternative QSAR approach for modeling this series of compounds. In other paper, Freitas, M. P. [58] et al. calculated different descriptors with Dragon software through three different feature selection methods, namely Genetic Algorithm (GA), Successive Projections Algorithm (SPA), and fuzzy rough set Ant Colony Optimization (fuzzy rough set ACO). Each set of selected descriptors was regressed against the bioactivities of a series of glycogen synthase kinase-3 inhibitors, through linear and nonlinear regression methods, namely Multiple Linear Regression (MLR), Artificial Neural Network (ANN), and Support Vector Machines (SVM). While calibration with MLR yielded QSAR models only reasonably predictable, with r2 ranging from 0.77 to 0.81 and r2 test of 0.67 to 0.76, ANN and specially SVM were capable of estimating and predicting biological activities very accurately. According to PLSbased models, r2 varied from 0.78 to 0.82 for the training set and from 0.70 to 0.80 for the test set. Leave-one-out (LOO) crossvalidation experiments gave q2 ranging from 0.76 to 0.79. These results were only comparable to MLR-based models, independent of the feature selection method used. Overall, Support Vector Machines (SVM) were suggested to be used as a regression method in QSAR studies where linear behavior is not expected or in those studies in which linear approaches do not work well. 10. Free-Wilson QSAR Analysis to Predict Kinase Selectivity Profiles Sciabola, S. [59] et al. wrote this article about kinases that are involved in a variety of diseases such as cancer, diabetes, and arthritis. In recent years, many kinase small molecule inhibitors developed as potential disease treatments. Despite the recent advances, selectivity remains one of the most challenging aspects in kinase inhibitor design. To interrogate kinase selectivity, a panel of 45 kinase assays was developed in-house at Pfizer. The KSS is the Pfizer internal Kinase Selectivity Screening panel consisting of 45 different protein kinase bioassays selected based on bioinformatics and structural data to provide maximal coverage across subfamilies within the kinome. Here, they presented an application of in silico Quantitative Structure Activity relationship (QSAR) models to extract rules from this experimental screening data and make reliable selectivity profile predictions for all compounds enumerated from virtual libraries. They also proposed the construction of Rgroup selectivity profiles by deriving their activity contribution against each kinase using QSAR models. Such selectivity profiles can be used to provide better understanding of subtle structure selectivity relationships during kinase inhibitor design. 11. CoMFA and Docking for Pyrazolo[3,4-b]pyrid[az]ines as GSK-3 Inhibitors Bharatam [60] et al. in this paper carried out a 3D-QSAR study of pyrazolo[3,4-b]pyrid[az]ine derivatives, which includes nonselective and selective GSK-3 inhibitors. The training set of CoMFA models were made from a 59 molecules (see Table 1). A test set containing 14 molecules was used to validate the CoMFA models. The CoMFA model generated by applying leave-one-out (LOO) cross-validation study gave a higher cross-validated correlation coefficient, r2cv = 0.60, which indicated a good internal prediction of the model; and conventional r2conv = 0.97; the predictive correlation coefficient, r2pred = 0.55, indicated good external predictive ability of CoMFA model. The 3D-QSAR study of pyrazolo[3,4-b]pyrid[az]ine derivatives led to the identification of regions of importance for steric and electronic interactions. Validation based on the molecular docking was also made to explain the structural differences between the selective and non-selective molecules in the given series of molecules. 12. 2D and 3D-QSAR Models for Prediction of Indirubin Derivatives GSK-3 Inhibitors The 3D-QSAR study made by Viney, L. et al. [61] of indirubin derivatives (see Table 1) was carried out with molecular descriptors calculated by CODESSA and Molconn-Z and a 3D-QSAR study based on the principle of the alignment of pharmacophoric features by PHASE module of Schrodinger. Statistically significant 2-D (r2 = 0.93) and 3-D (r2 = 0.97) QSAR models were generated using 36 molecules in the training set. The predictive correlation coefficient: r2pred = 0.6 for 2Dand r2pred = 0.91 for 3-D model. These studies
2672 Current Pharmaceutical Design, 2010, Vol. 16, No. 24 García et al. Table 2. Some Rules Used to Classify GSK-3 Inhibitors in Different Assay Conditions Parameter Isoform Specie Conditiona enzyme test (in vitro) % - 0 cKi - >100 IC50 (nM) - >2000 IC50 (nM) - >2000 IC50 (KM) - >2 IC50 (KM) nd - >2 IC50 (KM) / - >2 IC50 E-9 (M) nd - >2 pIC50 - 0 p IC50 nd - 0 parasite EC50 (KM) - T. brucei >2 IC50 (ng/mL) - P. falciparum bNA ED50 (KM) - L. donovani <5 IC50 (ng/mL) - P. falciparum cNA IC50 (Kg/mL) - L. donovani NA IC50 (KM) - P. falciparum >20 IC50 (KM) - L. mexicana >2 IC50 (KM) - P. falciparum >2 IC50 (KM) - P. falciparum D6 NA IC50 (KM) - P. falciparum W2 NA IC90 (Kg/mL) - L. donovani NA bacteria IC50 (Kg/mL) - M. intracellulare NA IC50 (Kg/mL) - MRS NA IC50 (Kg/mL) - S. aureus NA IC50 (KM) - M. intracellulare - IC50 (KM) - MRSA - MIC (Kg/mL) - M. tuberculosi NA cell line IC50 (Kg/mL) - HVC NC IC50 (KM) - Hep2 NA IC50 (KM) - HT29 NA IC50 (KM) - LMM3 NA IC50 (KM) - PTP >2 IC50 (KM) - RD >2 IC50 (KM) - U937 >2 aCondition to be classified as non-active compound (LDA class 0); nd = not determined, NA= not active, NC = not cytotoxicity; bChloroquine resistant W2 clone; c chloroquine sensitive D6 clone. showed that 3D-QSAR model is a better predictive model for the indirubins compared to 2D-QSAR model. 2D-QSAR studies demonstrated that activity of these inhibitors can be increased by changing the polarity, electronegativity of the atoms that might be forming H-bonds with the receptor binding site. This study also correlates with the novel 3D-QSAR study generated volume occluded maps, which demonstrated that the activity can be increased by changing the electronegativity of the N-atoms involved in the Hbond interactions with the binding site of the receptor. Also 3DQSAR results demonstrated that the biological activity of the indirubin-like molecules could be increased by placing the hydrophobic groups at particular positions on the phenyl ring of indirubins. 13. Multi-Target QSAR In Silico Screening for GSK-3 Inhibitors In this paper, we studied the development of a discriminant function [62] that allowed the classification of organic compounds as active or non-active which is the key step in the present approach for the discovery of GSK-3 inhibitors. It was therefore necessary to select a training data set of GSK-3 inhibitors containing wide structural variability. The selection of discriminant techniques instead of regression techniques was determined by the lack of homogeneity in the conditions under which these values were measured. As reported in different sources, numerous IC50 values lie within a range rather than a single value. In other cases, the activity was not scored in terms of IC50 values but was quoted as inhibitory percentages at a given concentration. In Table 2, we summarize different rules used to classify compounds as active or inactive GSK-3 inhibitors in different assay conditions. Once the training series had been designed, forward stepwise Linear Discriminant Analysis (LDA) was carried out in order to derive the QSAR; this equation confirmed our intuitive hypothesis and we could conclude that the deviation of the parameters of one compound from the average values for active compounds tested in a given assay condition (CACq) was very important for the prediction of this compound as active. In particular, for i-GSK-3 small deviations in total polarizability and molecular refractivity indices as well as local deviations in refracitivity and electronegativity for heteroatom-bound hydrogen atoms seem to be especially determinant. Specifically, the introduction of this last parameter in the model coincide with the GSK-3 activity observed in organic acid compounds with [63] this type of model based only on atomic physicochemical parameters and drug connectivity may become an interesting alternative for fast computational pre-screening of large series of compounds in order to rationalize synthetic efforts [52, 60, 64-68] complementary to more elaborated techniques 3D-QSAR, CoMFA, and CoMSIA studies that depend on a detailed knowledge of 3D structure. In any case, the present model is of more general application than the other known methods that apply only to compounds tested in only one CAC and/or belonging to only one homogeneous structural class of compounds. A confirmation of this stamen is that the present classification function has given rise to an efficient separation of all compounds with Accuracy = 89.1% (training series) and Accuracy = 88.3% (validation series). The names, observed classification, predicted classification and subsequent probabilities for all 4508 compounds (see examples in Table 1) in training and average validation are given as supplementary material. This level of total Accuracy, Sensitivity and Specificity is considered as excellent by other researchers that have used LDA for QSAR studies; see for instance the works of Garcia-Domenech, R., Prado-Prado, F. J.; Marrero-Ponce, Y., etc [69-84]. ACKNOWLEDGMENT We are grateful to the Xunta de Galicia (INCITE08PXIB314255PR) for partial financial support.
QSAR, Docking, and CoMFA studies of GSK3 inhibitors Current Pharmaceutical Design, 2010, Vol. 16, No. 24 2673 REFERENCES [1] Woodgett JR. Molecular cloning and expression of glycogen synthase kinase-3/factor A. EMBO J 1990; 9: 2431-8. [2] Troussard AA, Tan C, Yoganathan TN, Dedhar S. Cellextracellular matrix interactions stimulate the AP-1 transcription factor in an integrin-linked kinaseand glycogen synthase kinase 3dependent manner. Mol Cell Biol 1999; 19: 7420-7. [3] Turenne GA, Price BD. Glycogen synthase kinase3 beta phosphorylates serine 33 of p53 and activates p53´s transcriptional activity. BMC Cell Biol 2001; 2: 12-21. [4] Hoeflich KP, Luo J, Rubie EA, Tasao MS, Jin O, Woodgett JR. Requirement for glycogen synthase kinase-3beta in cell survival and NF-kappaB activation. Nature 2000; 406: 86-90. [5] Hooper C, Killick R, Lovestone S. The GSK3 hypothesis of Alzheimer´s disease. J Neurochem 2008; 104: 1433-9. [6] MacAulay K, Doble BW, Patel S, et al. Glycogen synthase kinase 3alpha-specific regulation of murine hepatic glycogen metabolism. Cell Metab 2007; 6: 329-37. [7] Olson RE. Secretase inhibitors as therapeutics for Alzheimer´s disease. Annu Rep Med Chem 2000; 35: 31-40. [8] Vehmas AK, Kawas CH, Stewart WF, Troncoso JC. Immune reactive cells in senile plaques and cognitive decline in Alzheimer´s disease. Neurobiol Aging 2003; 24: 321-31. [9] Pei JJ, Braak H, Grundke-Iqbal I, Iqbal K, Winblad B, Cowburn RF. Distribution of active glycogen synthase kinase 3beta (GSK3beta) in brains staged for Alzheimer disease neurofibrillary changes. J Neuropathol Exp Neurol 1999; 58: 1010-9. [10] Leroy K, Yilmaz Z, Brion JP. Increased level of active GSK-3beta in Alzheimer´ disease and accumulation in argyrophilic grains and in neurones at different stages of neurofibrillary degeneration. Neuropathol Appl Neurobiol 2007; 33: 43-55. [11] Droucheau E. Plasmodium falciparum glycogen synthase kinase-3: molecular model, expression, intracellular localisation and selective inhibitors. Biochim Biophys Acta 2004; 1697: 181-96. [12] Woodgett JR. cDNA cloning and properties of glycogen synthase kinase-3. Methods Enzymol 1991; 200: 564-77. [13] Ali A, Hoeflich KP,Woodgett JR. Glycogen Synthase kinase-3: properties, functions, and regulation. Chem Rev 2001; 101: 252740. [14] Grimes CA, Jope RS. The multifaceted roles of glycogen synthase kinase 3B in cellular signaling. Prog Neurobiol 2001; 65: 391-426. [15] Nadri C, Lipska B, Kozlovsky N, Weinberger DR, Belmaker RH, Agam G. GSK-3 levels and activity in a neurodevelopmental rat model of schizophrenia. Dev Brain Res 2003; 141: 33-7. [16] Gould TD, Zarate CA, Manji HK. Glycogen synthase kinase-3: a target for novel bipolar disorder treatments. J Clin Psychiatry 2004; 65: 10-21. [17] Ishiguro K, Ihara Y, Uchida T, Imahori K. A novel tubulindependent protein kinase forming a paired helical filament epitope on tau. J Bio Chem 1988; 104: 319-21. [18] Meijer L, Flajolet M, Geengard P. Pharmacological inhibitors of glycogen synthase kinase 3. Trends Phamacological Sci 2004; 25: 471-80. [19] Fairlamb AH. Chemotherapy on human African trypanosomiasis:current and future prospects. Trends Parasitol 2003; 19: 488-94. [20] Copeland RA, Pompliano DL, Meek TD. Drug-target residence time and its implications for lead optimization. Nat Rev Drug Discov 2006; 5: 730-2. [21] Liao JJ. Molecular recognition of protein kinase binding pockets for design of potent and selective kinase inhibitors. J Med Chem 2007; 50: 409-24. [22] Pink R, Hudson A, Mouries MA, Bending M. Opportunities and chanllenges in antiparasitic drug discovery. Nat Rev Drug Discov 2005; 4: 727-40. [23] Plyte SE, Hughes K, Nilkolakaki E, Pulverer BJ,Woodgett JR. Glycogen Synthase Kinase-3: functions in oncogenesis and development. Biochim Biophys Acta 1992; 1114: 147-62. [24] Dajani R, Fraser E, Roe SM, et al. Crystal structure of glycogen synthase kinase-3 beta: structural basic for phosphate-primed subtrate specificity and autoinhibition. Cell (Cambridge, Massachusetts) 2001; 105: 721-32. [25] Ojo KK, Gillespie RG, Riechers A, et al. Glycogen Synthase Kinase 3 is a potential drug target for african trypanosomiasis therapy. Antimicrob Agents Chemother 2008; 52: 3710-7. [26] Nunez MB, Maguna FP, Okulik NB,Castro EA. QSAR modeling of the MAO inhibitory activity of xanthones derivatives. Bioorg Med Chem Lett 2004; 14: 5611-7. [27] Todeschini R, Consonni V. Handbook of molecular descriptors. USA: Wiley VCH 2000. [28] Freund JA, Poschel T. Stochastic process in physics, chemistry and biology. In: Lecture Notes on Physics. Berlin, Germany: SpringerVerlag 2000. [29] González-Díaz H, Munteanu CR. Topological indices for medicinal chemistry, biology, parasitology, neurological and social networks. Kerala: Transworld Research Network 2010. [30] González-Díaz H, Duardo-Sanchez A, Ubeira FM, et al. Review of MARCH-INSIDE & complex networks prediction of drugs: ADMET, anti-parasite activity, metabolizing enzymes and cardiotoxicity proteome biomarkers. Curr Drug Metabol 2010; 11: 379-406. [31] Vina D, Uriarte E, Orallo F, Gonzalez-Diaz H. Alignment-free prediction of a drug-target complex network based on parameters of drug connectivity and protein sequence of receptors. Mol Pharm 2009; 6: 825-35. [32] Pérez-Montoto LG, Prado-Prado F, Ubeira FM, González-Díaz H. Study of parasitic infections, cancer, and other diseases with massspectrometry and quantitative proteome-disease relationships. Curr Proteomics 2009; 6: 246-61. [33] González-Díaz H, Prado-Prado F, Pérez-Montoto LG, DuardoSánchez A, López-Díaz A. QSAR models for proteins of parasitic organisms, plants and human guests: theory, applications, legal protection, taxes, and regulatory issues. Curr Proteomics 2009; 6: 214-27. [34] Concu R, Dea-Ayuela MA, Perez-Montoto LG, et al. Prediction of enzyme classes from 3D structure: a general model and examples of experimental-theoretic scoring of peptide mass fingerprints of Leishmania proteins. J Proteome Res 2009; 8: 4372-82. [35] Aguero-Chapin G, Varona-Santos J, de la Riva GA, et al. Alignment-free prediction of polygalacturonases with pseudofolding topological indices: experimental isolation from coffea arabica and prediction of a new sequence. J Proteome Res 2009; 8: 2122-8. [36] Gonzalez-Díaz H, Prado-Prado F,Ubeira FM. Predicting antimicrobial drugs and targets with the MARCH-INSIDE approach. Curr Top Med Chem 2008; 8: 1676-90. [37] González-Díaz H, González-Díaz Y, Santana L, Ubeira FM, Uriarte E. Proteomics, networks and connectivity indices. Proteomics 2008; 8: 750-78. [38] Rodriguez-Soca Y, Munteanu CR, Dorado J, Rabuñal J, Pazos A, González-Díaz. Plasmod-PPI: A web-server predicting complex biopolymer targets in plasmodium with entropy measures of protein-protein interactions. Polymer 2010; 51: 264-73. [39] Munteanu CR, Dorado J, Pazos Sierra A, et al. In information theory analysis of complex networks: statistical methods and applications. Emmert-streib F, Dehmer M, Mehler A, Eds. USA: Springer-Verlag 2010. [40] González-Díaz H, Agüero-Chapin G, Munteanu CR, et al. In advances in genetics Research. Osborne MA, Eds. New York: Nova Sciences 2010; Vol. 1. [41] Rodriguez-Soca Y, Munteanu CR, Prado-Prado FJ, Dorado J, Pazos Sierra A, Gonzalez-Diaz H. Trypano-PPI: A web server for prediction of unique targets in trypanosome proteome by using electrostatic parameters of protein-protein interactions. J Proteome Res 2010; 9(2): 1182-90. [42] Ferino G, Delogu G, Podda G, Uriarte E, González-Díaz H. Clinical chemistry research (ISBN: 978-1-60692-517-1). In: Sharnham BH. Mitchem CHL, Eds. NY: Nova Science Publisher 2009. [43] Concu R, Podda G, Uriarte E, González-Díaz H. Handbook of computational chemistry research. In: Robson CT, Collett CD, Eds. USA: Nova Science Publishers 2009. [44] González-Díaz H, Vilar S, Santana L,Uriarte E. Medicinal chemistry and bioinformatics - current trends in drugs discovery with networks topological indices. Curr Top Med Chem 2007; 7: 1025-39. [45] Gonzalez-Diaz H, Saiz-Urra L, Molina R, Gonzalez-Diaz Y, Sanchez-Gonzalez A. Computational chemistry approach to protein
2674 Current Pharmaceutical Design, 2010, Vol. 16, No. 24 García et al. kinase recognition using 3D stochastic van der Waals spectral moments. J Comput Chem 2007; 28: 1042-8. [46] González-Díaz H, Pérez-Castillo Y, Podda G,Uriarte E. Computational chemistry comparison of stable/nonstable protein mutants classification models based on 3D and topological indices. J Comput Chem 2007; 28: 1990-5. [47] Estrada E, Uriarte E. Recent advances on the role of topological indices in drug discovery research. Curr Med Chem 2001; 8: 157388. [48] Estrada E, Uriarte E, Montero A, Teijeira M, Santana L,De Clercq E. A novel approach for the virtual screening and rational design of anticancer compounds. J Med Chem 2001; 43: 1975-85. [49] Martinez A, Alonso M, Castro A, et al. SAR and 3D-QSAR Studies on thiadiazolidinone derivatives: exploration of structural requirements for glycogen synthase kinase 3 inhibitors. J Med Chem 2005; 48: 7103-12. [50] Lescot E, Bureau R, Sopkova-de Oliveira Santos J, et al. 3D-QSAR and docking studies of selective GSK-3beta inhibitors. Comparison with a thieno[2,3-b]pyrrolizinone derivative, a new potential lead for GSK-3beta ligands. J Chem Inf Model 2005; 45: 708-15. [51] Katritzky AR, Pacureanu LM, Dobchev DA, Fara DC, Duchowicz PR,Karelson M. QSAR modeling of the inhibition of Glycogen Synthase Kinase-3. Bioorg Med Chem 2006; 14: 4987-5002. [52] Xiao J, Guo Z, Guo Y, Chu F, Sun P. Inhibitory mode of N-phenyl4-pyrazolo[1,5-b] pyridazin-3-ylpyrimidin-2-amine series derivatives against GSK-3: molecular docking and 3D-QSAR analyses. Protein Eng Des Sel 2006; 19: 47-54. [53] Jiang Y, Zhang N, Zou J, et al. 3D QSAR for GSK-3 inhibition by indirubin analogues. Eur J Med Chem 2006; 41: 373-8. [54] Dessalew N, Bharatam PV. 3D-QSAR and molecular docking study on bisarylmaleimide series as glycogen synthase kinase 3, cyclin dependent kinase 2 and cyclin dependent kinase 4 inhibitors: An insight into the criteria for selectivity. Eur J Med Chem 2007; 42: 1014-27. [55] Ling L, Li-Na Z, Feng-Chao J. Contruction of the pharmacophore model of glycogen synthase kinase-3 inhibitors. Chinese J Chem 2007; 25: 892-7. [56] Bharatam PV, Dessalew N,Patel DS. 3D-QSAR and molecular docking studies on pyrazolopyrimidine derivatives as glycogen synthase kinase-3B inhibitors. J Mol Graph Model 2007; 25: 88595. [57] Freitas MP. A 2D image-based approach for modelling some glycogen synthase kinase 3 inhibitors. Med Chem Res 2007; 16: 461-7. [58] Freitas MP, Goodarzi M,Jensen R. Feature selection and linear/nonlinear regression methods for the accurate prediction of glycogen synthase kinase-3 inhibitory activities. J Chem Informat Model 2009; 49: 824-32. [59] Sciabola S, Stanton RV, Wittkopp S, et al. Predicting kinase selectivity profiles using free-wilson QSAR analysis. J Chem Inf Model 2008; 48: 1851-67. [60] Patel DS, Bharatam PV. Selectivity criterion for pyrazolo[3,4b]pyrid[az]ine derivatives as GSK-3 inhibitors: CoMFA and molecular docking studies. Eur J Med Chem 2008; 43: 949-57. [61] Viney L, Kristam R, Saini JS, Kristam R, Karthikeyan NA,Balaji VN. QSAR models for prediction of glycogen synthase kinase-3 inhibitory activity of indirubin derivatives. QSAR Comb Sci 2008; 27: 718-28. [62] Van Waterbeemd H. In Chemometric methods in molecular design. Van Waterbeemd H, Ed. New York: Wiley-VCH 1995; Vol. 2, pp. 265-82. [63] Konda VR, Desai A, Darland G, Bland JS,Tripp ML. Rho iso-alpha acids from hops inhibit the GSK-3/NF-kappaB pathway and reduce inflammatory markers associated with bone and cartilage degradation. J Inflamm (Lond) 2009; 6: 26-34. [64] Rochais C, Duc NV, Lescot E, et al. Synthesis of new dipyrroloand furopyrrolopyrazinones related to tripentones and their biological evaluation as potential kinases (CDKs1-5, GSK-3) inhibitors. Eur J Med Chem 2009; 44: 708-16. [65] Simon D, Benitez MJ, Gimenez-Cassina A, et al. Pharmacological inhibition of GSK-3 is not strictly correlated with a decrease in tyrosine phosphorylation of residues 216/279. J Neurosci Res 2008; 86: 668-74. [66] Jacquemard U, Dias N, Lansiaux A, et al. Synthesis of 3,5-bis(2indolyl)pyridine and 3-[(2-indolyl)-5-phenyl]pyridine derivatives as CDK inhibitors and cytotoxic agents. Bioorg Med Chem 2008; 16: 4932-53. [67] Tavares FX, Boucheron JA, Dickerson SH, et al. N-Phenyl-4pyrazolo[1,5-b]pyridazin-3-ylpyrimidin-2-amines as potent and selective inhibitors of glycogen synthase kinase 3 with good cellular efficacy. J Med Chem 2004; 47: 4716-30. [68] Olesen PH, Sorensen AR, Urso B, et al. Synthesis and in vitro characterization of 1-(4-aminofurazan-3-yl)-5-dialkylaminomethyl1H-[1,2,3]triazole-4-carboxyl ic acid derivatives. A new class of selective GSK-3 inhibitors. J Med Chem 2003; 46: 3333-41. [69] Calabuig C, Anton-Fos GM, Galvez J, Garcia-Domenech R. New hypoglycaemic agents selected by molecular topology. Int J Pharm 2004; 278: 111-8. [70] Garcia-Garcia A, Galvez J, de Julian-Ortiz JV, et al. New agents active against Mycobacterium avium complex selected by molecular topology: a virtual screening method. J Antimicrob Chemother 2004; 53: 65-73. [71] Prado-Prado FJ, Ubeira FM, Borges F, Gonzalez-Diaz H. Unified QSAR & network-based computational chemistry approach to antimicrobials. II. Multiple distance and triadic census analysis of antiparasitic drugs complex networks. J Comput Chem 2009; 31: 16473. [72] Prado-Prado FJ, Martinez de la Vega O, Uriarte E, Ubeira FM, Chou KC, Gonzalez-Diaz H. Unified QSAR approach to antimicrobials. 4. Multi-target QSAR modeling and comparative multidistance study of the giant components of antiviral drug-drug complex networks. Bioorg Med Chem 2009; 17: 569-75. [73] Prado-Prado FJ, de la Vega OM, Uriarte E, Ubeira FM, Chou KC, Gonzalez-Diaz H. Unified QSAR approach to antimicrobials. 4. Multi-target QSAR modeling and comparative multi-distance study of the giant components of antiviral drug-drug complex networks. Bioorg Med Chem 2009; 17: 569-75. [74] Prado-Prado FJ, Borges F, Perez-Montoto LG, Gonzalez-Diaz H. Multi-target spectral moment: QSAR for antifungal drugs vs. different fungi species. Eur J Med Chem 2009; 44: 4051-6. [75] Prado-Prado FJ, Gonzalez-Diaz H, de la Vega OM, Ubeira FM,Chou KC. Unified QSAR approach to antimicrobials. Part 3: first multi-tasking QSAR model for input-coded prediction, structural back-projection, and complex networks clustering of antiprotozoal compounds. Bioorg Med Chem 2008; 16: 5871-80. [76] Prado-Prado FJ, Gonzalez-Diaz H, Santana L, Uriarte E. Unified QSAR approach to antimicrobials. Part 2: predicting activity against more than 90 different species in order to halt antibacterial resistance. Bioorg Med Chem 2007; 15: 897-902. [77] Marrero-Ponce Y, Khan MT, Casanola Martin GM, et al. Prediction of Tyrosinase Inhibition Activity Using Atom-Based Bilinear Indices. ChemMedChem 2007; 2: 449-78. [78] Marrero-Ponce Y, Meneses-Marcel A, Castillo-Garit JA, et al. Predicting antitrichomonal activity: a computational screening using atom-based bilinear indices and experimental proofs. Bioorg Med Chem 2006; 14: 6502-24. [79] Meneses-Marcel A, Marrero-Ponce Y, Machado-Tugores Y, et al. A linear discrimination analysis based virtual screening of trichomonacidal lead-like compounds: outcomes of in silico studies supported by experimental results. Bioorg Med Chem Lett 2005; 15: 3838-43. [80] Marrero-Ponce Y, Diaz HG, Zaldivar VR, Torrens F, Castro EA. 3D-chiral quadratic indices of the 'molecular pseudograph's atom adjacency matrix' and their application to central chirality codification: classification of ACE inhibitors and prediction of sigmareceptor antagonist activities. Bioorg Med Chem 2004; 12: 533142. [81] Murcia-Soler M, Perez-Gimenez F, Garcia-March FJ, SalabertSalvador MT, Diaz-Villanueva W, Medina-Casamayor P. Discrimination and selection of new potential antibacterial compounds using simple topological descriptors. J Mol Graph Model 2003; 21: 375-90.
6 Current Bioinformatics, 2011, Vol. 6, No. 1 García et al. 11. STUDY OF RETINOID ACID RECEPTOR In this section, we show all structures of retinoid acid receptor found in RCSB PDB in the last years (Table 2). The structure of 3H0A [63] is the crystal structure of peroxisome proliferator-activated receptor gamma (PPARg) and retinoid acid receptor alpha (RXRa) in complex with 9-cis retinoid acid, co-activator peptide, and a partial agonist. This structure was studied by X-ray diffraction with a resolution of 2.10Å and it is formed by three polymers and two ligands (Fig. 7A). The structure of 2GL8 is the structure of human retinoid acid receptor RXR-gamma ligand-binding domain. This structure was studied by X-ray diffraction with a resolution of 2.40Å and it is formed by one polymer (Fig. 7B). The structure of 1YNW [64] is the crystal structure vitamin D receptor and 9-cis retinoid acid receptor DNA-binding domain bound to a DR3 response element. This structure was studied by X-ray diffraction with a resolution of 3.00Å and it is formed by four polymers and one ligand. The 1XAP [65] is the is the structure of the ligand binding domain of the retinoid acid receptor beta. This structure was studied by Xray diffraction with a resolution of 2.10Å and it is formed by Fig. (6). A. Stereo view showing RH5849 EcR-LBD complex based on retinoid acid. B. Stereo view showing RH5849 EcR-LBD complex based on vitamin D. Table 2. PDB Author Year Classification Experiment Resolution Ref. 3H0A Wang, Z. 2009 transcription x-ray 2.10Å [63] 2GL8 Min, J.R. 2006 Hormone/growth Factor Receptor x-ray 2.40Å - 1YNW Shaffer, P.L. 2005 Transcription/dna x-ray 3.00Å [69] 1XAP Germain, P. 2004 Transcription x-ray 2.10Å [70] 1EXA Klaholz, B.P. 2000 Gene Regulation x-ray 1.59Å [71] 1EXX Klaholz, B.P. 2000 Gene Regulation x-ray 1.67Å [71] 3LBD Klaholz, B.P. 1999 Nuclear Receptor x-ray 2.40Å - 4LBD Klaholz, B.P. 1999 Nuclear Receptor x-ray 2.50Å - 2LBD Renaud, J.-P. 1997 Nuclear Receptor x-ray 2.06Å [72] 1HRA Knegtel, R.M.A. 1994 DNA Binding Receptor NMR - [73]
Trends in Bioinformatics and Chemoinformatics of Vitamin D Current Bioinformatics, 2011, Vol. 6, No. 1 7 one polymer and one ligand (Fig. 7C). The 1EXA is the enantiomer discrimination illustrated by crystal structures of the human retinoid acid receptor hRAR-gamma ligand binding domain; is the complex with the active R-enantiomer BMS270394 [66]. It is formed by one polymer and two ligands studied by X-ray diffraction with a resolution of 1.59Å. The 1EXX is the same that 1EXA but with the inactive S-enantiomer BMS270395 [66] and with a resolution of 1.67Å. The resolution of X-ray diffraction of the 3LBD is 2.40Å and is formed by one polymer and one ligand and 1EXA is the structure of the ligand binding domain of the human retinoid acid receptor gamma bound to 9-cis retinoid acid. The structure of the ligand binding domain of the human retinoid acid receptor gamma bound to the synthetic agonist BMS961 is named as 4LBD and her resolution of the X-ray diffraction is 2.50Å. It is formed by one polymer and one ligand. The 2LBD is the structure of the same receptor but to all-trans retinoid acid. Her resolution of the X-ray diffraction is 2.06Å [67]. The 1HRA is the solution structure of the human retinoid acid receptor-beta DNA-binding domain studied by solution NMR [68] (Fig. 7D). ACKNOWLEDGMENT We are grateful to the Xunta de Galicia (INCITE08PXIB314255PR) for partial financial support. REFERENCES [1] Cernakova M, Kost'alova D, Kettmann V, Plodova M, Toth J, Drimal J. Potential antimutagenic activity of berberine, a constituent of Mahonia aquifolium. BMC Comp Altern Med 2002; 2: 2. [2] Cao JW, Luo HS, Yu BP, Huang XD, Sheng ZX, Yu JP. Effects of berberine on intracellular free calcium in smooth muscle cells of Guinea pig colon. Digestion 2001; 64: 179-83. [3] Jang MJ, Jwa M, Kim JH, Song K. Selective inhibition of MAPKK Wis1 in the stress-activated MAPK cascade of Schizosaccharomyces pombe by novel berberine derivatives. J Biol Chem 2002; 277: 12388-95. [4] Li BX, Yang BF, Zhou J, Xu CQ, Li YR. Inhibitory effects of berberine on IK1, IK, and HERG channels of cardiac myocytes. Acta Pharmacol Sin 2001; 22: 125-31. [5] Choi DS, Kim SJ, Jung MY. Inhibitory activity of berberine on DNA strand cleavage induced by hydrogen peroxide and cytochrome c. Biosci Biotechnol Biochem 2001; 65: 452-5. [6] Iizuka N, Miyamoto K, Okita K, et al. Inhibitory effect of Coptidis Rhizoma and berberine on the proliferation of human esophageal cancer cell lines. Cancer Lett 2000; 148: 19-25. [7] Li H, Miyahara T, Tezuka Y, et al. The effect of kampo formulae on bone resorption in vitro and in vivo. II. Detailed study of berberine. Biol Pharm Bull 1999; 22: 391-6. [8] Chou KC. Structural bioinformatics and its impact to biomedical science. Curr Med Chem 2004; 11: 2105-34. [9] Chou KC, Wei DQ, Du QS, Sirois S, Zhong WZ. Progress in computational approach to drug development against SARS. Curr Med Chem 2006; 13: 3263-70. [10] Todeschini R, Consonni V. Handbook of Molecular Descriptors. Wiley-VCH: 2002. [11] Estrada E., Uriarte E. Recent advances on the role of topological indices in drug discovery research. Curr Med Chem 2001; 8: 157388. [12] González-Díaz H, González-Díaz Y, Santana L, Ubeira FM, Uriarte E. Proteomics, networks and connectivity indices. Proteomics 2008; 8: 750-78. [13] Torrens F, Castellano G. Topological Charge-Transfer Indices: From Small Molecules to Proteins. Curr Proteomics 2009: 204-13. [14] Concu R, Dea-Ayuela MA, Perez-Montoto LG, et al. 3D entropy and moments prediction of enzyme classes and experimentaltheoretic study of peptide fingerprints in Leishmania parasites. Biochim Biophys Acta 2009; 1794: 1784-94. [15] Ivanciuc O. Machine learning Quantitative Structure-Activity Relationships (QSAR) for peptides binding to Human Amphiphysin-1 SH3 domain. Curr Proteomics 2009; 4: 289-302. [16] Vilar S, Gonzalez-Diaz H, Santana L, Uriarte E. A network-QSAR model for prediction of genetic-component biomarkers in human colorectal cancer. J Theor Biol 2009; 261: 449-58. [17] Giuliani A, Di Paola L, Setola R. Proteins as Networks: A Mesoscopic Approach Using Haemoglobin Molecule as Case Study. Curr Proteomics 2009; 6: 235-45. Fig. (7). View of some retinoic acid receptors: A. 3H0A; B. 2GL8; C. 1XAP; D. 1HRA.
8 Current Bioinformatics, 2011, Vol. 6, No. 1 García et al. [18] Chou KC. Pseudo amino acid composition and its applications in bioinformatics, proteomics and system biology. Curr Proteomics 2009; 6: 262-74. [19] Chen J, Shen B. Computational Analysis of Amino Acid Mutation: a Proteome Wide Perspective. Curr Proteomics 2009; 6: 228-34. [20] Gonzalez-Diaz H. Quantitative studies on Structure-Activity and Structure-Property Relationships (QSAR/QSPR). Curr Top Med Chem 2008; 8: 1554. [21] Caballero J, Fernandez M. Artificial neural networks from MATLAB in medicinal chemistry. Bayesian-regularized genetic neural networks (BRGNN): application to the prediction of the antagonistic activity against human platelet thrombin receptor (PAR-1). Curr Top Med Chem 2008; 8: 1580-605. [22] Duardo-Sanchez A, Patlewicz G, Lopez-Diaz A. Current topics on software use in medicinal chemistry: intellectual property, taxes, and regulatory issues. Curr Top Med Chem 2008; 8: 1666-75. [23] Gonzalez MP, Teran C, Saiz-Urra L, Teijeira M. Variable selection methods in QSAR: an overview. Curr Top Med Chem 2008; 8: 1606-27. [24] Gonzalez-Diaz, H, Prado-Prado F, Ubeira FM. Predicting antimicrobial drugs and targets with the MARCH-INSIDE approach. Curr Top Med Chem 2008; 8: 1676-90. [25] Helguera AM, Combes RD, Gonzalez MP, Cordeiro MN. Applications of 2D descriptors in drug design: a DRAGON tale. Curr Top Med Chem 2008; 8: 1628-55. [26] Ivanciuc O. Weka machine learning for predicting the phospholipidosis inducing potential. Curr Top Med Chem 2008; 8: 1691-709. [27] Vilar S, Cozza G, Moro S. Medicinal chemistry and the molecular operating environment (MOE): application of QSAR and molecular docking to drug discovery. Curr Top Med Chem 2008; 8: 1555-72. [28] Wang JF, Wei DQ, Chou KC. Drug candidates from traditional chinese medicines. Curr Top Med Chem 2008; 8: 1656-65. [29] Wang JF, Wei DQ, Chou KC. Pharmacogenomics and personalized use of drugs. Curr Top Med Chem 2008; 8: 1573-9. [30] Lin YY, Qi Y, Lu JY, et al. A comprehensive synthetic genetic interaction network governing yeast histone acetylation and deacetylation. Genes & Dev 2008; 22: 2062-74. [31] Zhong WZ, Zhan J, Kang P, Yamazaki S. Gender specific drug metabolism of PF-02341066 in rats-role of sulfoconjugation. Curr Drug Metab 2010; 11: 296-306. [32] Wang JF, Chou KC. Molecular modeling of cytochrome P450 and drug metabolism. Curr Drug Metab 2010; 11: 342-6. [33] Mrabet Y, Semmar N. Mathematical methods to analysis of topology, functional variability and evolution of metabolic systems based on different decomposition concepts. Curr Drug Metab 2010; 11: 315-41. [34] Martinez-Romero M, Vazquez-Naya JM, Rabunal JR, et al. Artificial intelligence techniques for colorectal cancer drug metabolism: ontology and complex network. Curr Drug Metab 2010; 11: 34768. [35] Khan MT. Predictions of the ADMET properties of candidate drug molecules utilizing different QSAR/QSPR modelling approaches. Curr Drug Metab 2010; 11: 285-95. [36] Gonzalez-Diaz H, Duardo-Sanchez A, Ubeira FM, et al. Review of MARCH-INSIDE & complex networks prediction of drugs: ADMET, anti-parasite activity, metabolizing enzymes and cardiotoxicity proteome biomarkers. Curr Drug Metab 2010; 11: 379-406. [37] Gonzalez-Diaz H. Network topological indices, drug metabolism, and distribution. Curr Drug Metab 2010; 11: 283-4. [38] Garcia I, Diop YF, Gomez G. QSAR & complex network study of the HMGR inhibitors structural diversity. Curr Drug Metab 2010; 11: 307-14. [39] Chou KC. Graphic rule for drug metabolism systems. Curr Drug Metab 2010; 11: 369-78. [40] Concu R, Podda G, Ubeira FM, Gonzalez-Diaz H. Review of QSAR models for enzyme classes of drug targets: Theoretical background and applications in parasites, hosts, and other organisms. Curr Pharm Design 2010; 16: 2710-23. [41] Estrada E, Molina E, Nodarse D, Uriarte E. Structural contributions of substrates to their binding to P-Glycoprotein. A TOPS-MODE approach. Curr Pharm Design 2010; 16: 2676-709. [42] Garcia I, Fall Y, Gomez G. QSAR, docking, and CoMFA studies of GSK3 inhibitors. Curr Pharm Design 2010; 16: 2666-75. [43] Gonzalez-Diaz H. QSAR and complex networks in pharmaceutical design, microbiology, parasitology, toxicology, cancer, and neurosciences. Curr Pharm Design 2010; 16: 2598-600. [44] Gonzalez-Diaz H, Romaris F, Duardo-Sanchez A, et al. Predicting drugs and proteins in parasite infections with topological indices of complex networks: theoretical backgrounds, applications, and legal issues. Curr Pharm Design 2010; 16: 2737-64. [45] Marrero-Ponce Y, Casanola-Martin GM, Khan MT, Torrens F, Rescigno A, Abad C. Ligand-based computer-aided discovery of tyrosinase inhibitors. Applications of the TOMOCOMD-CARDD method to the elucidation of new compounds. Curr Pharm Design 2010; 16: 2601-24. [46] Munteanu CR, Fernandez-Blanco E, Seoane JA, et al. Drug discovery and design for complex diseases through QSAR computational methods. Curr Pharm Design 2010; 16: 2640-55. [47] Roy K, Ghosh G. Exploring QSARs with Extended Topochemical Atom (ETA) indices for modeling chemical and drug toxicity. Curr Pharm Design 2010; 16: 2625-39. [48] Speck-Planche A, Scotti MT, de Paulo-Emerenciano V. Current pharmaceutical design of antituberculosis drugs: future perspectives. Curr Pharm Design 2010; 16: 2656-65. [49] Vazquez-Naya JM, Martinez-Romero M, Porto-Pazos AB, et al. Ontologies of drug discovery and design for neurology, cardiology and oncology. Curr Pharm Design 2010; 16: 2724-36. [50] Tan YL, Goh D, Ong ES. Investigation of differentially expressed proteins due to the inhibitory effects of berberine in human liver cancer cell line HepG2. Mol Biosyst 2006; 2: 250-8. [51] Xu L, Liu Y, He X. Inhibitory effects of berberine on the activation and cell cycle progression of human peripheral lymphocytes. Cell Mol Immunol 2005; 2: 295-300. [52] Gonzalez MP, Gandara Z, Fall Y, Gomez G. Radial Distribution Function descriptors for predicting affinity for vitamin D receptor. Eur J Med Chem 2007. [53] Gonzalez MP, Puente M, Fall Y, Gomez G. In silico studies using Radial Distribution Function approach for predicting affinity of 1 alpha,25-dihydroxyvitamin D(3) analogues for Vitamin D receptor. Steroids 2006; 71: 510-27. [54] Jensen BF, Sorensen MD, Kissmeyer AM, et al. Prediction of in vitro metabolic stability of calcitriol analogs by QSAR. J Comput Aided Mol Des 2003; 17: 849-59. [55] Wang F, Zhou HY, Zhao G, et al. Inhibitory effects of berberine on ion channels of rat hepatocytes. World J Gastroenterol 2004; 10: 2842-5. [56] Ravi M, Hopfinger AJ, Hormann RE, Dinan L. 4D-QSAR analysis of a set of ecdysteroids and a comparison to CoMFA modeling. J Chem Inf Comput Sci 2001; 41: 1587-604. [57] Cernakova M, Kostalova D. Antimicrobial activity of berberine--a constituent of Mahonia aquifolium. Folia Microbiologica 2002; 47: 375-8. [58] Fink H. [On the problem of the minimal inhibitory concentration (MIC) of oxacillin against staphylococci]. Arzneimittel-Forschung 1965; 15: 630-2. [59] Tamura M, Takano S. [Influence of pH of media on the minimal inhibitory concentration of cycloserine to Mycobacterium tuberculosis]. Kekkaku 1965; 40: 213-8. [60] Miura T. Morphological observations of the influence of some chemical drugs on the growth of dermatophytes at concentrations directly below the minimal inhibitory concentration. Tohoku J Exp Med 1963; 80: 103-17. [61] Linser H. [The mechanism of action of growth and inhibitory substances. IV. The concentration-activity curves of various synthetic cell-elongating growth substances in the presence of various quantities of synthetic inhibitors.]. Biochim et Biophys Acta 1954; 15: 25-30. [62] Gero E. [Inhibitory action of vitamin B1 on the oxidation of Lascorbic acid. II. Influence of pH and of the vitamin B1 concentration; role of the various moieties of the vitamin B1 molecule.]. Bull de la Soc de Chimie Biol 1954; 36: 1335-42. [63] Raska, S.B. The Metabolism of the Kidney in Experimental Renal Hypertension: Ii. The Concentration of Cytochrome C and the Activities of the Cytochrome Oxidase and of the Succinic Dehydrogenase Systems in the Kidney of Dogs with Experimental Renal Hypertension. The Inhibitory Effect of Renin and of Kidney Tissue Preparations from Hypertensive Dogs on the Respiratory Enzymes. J Exp Med 1945; 82: 227-40. [64] Shaffer PL, Gewirth DT. Structural analysis of RXR-VDR interactions on DR3 DNA. J Steroid Biochem Mol Biol 2004; 89-90: 2159.
Trends in Bioinformatics and Chemoinformatics of Vitamin D Current Bioinformatics, 2011, Vol. 6, No. 1 9 [65] Germain P, Kammerer S, Peluso-Iltis C, et al. Rational design of RAR-selective ligands revealed by RARbeta crystal structure. Embo Rep 2004; 5: 877-82. [66] Klaholz BP, Mitschler A, Belema M, Zusi C, Moras D. Enantiomer discrimination illustrated by high-resolution crystal structures of the human nuclear receptor hRARgamma. Proc Natl Acad Sci USA 2000; 97: 6322-7. [67] Renaud J.-P, Rochel N, Ruff M, Moras D. Crystal structure of the RAR-gamma ligand-binding domain bound to all-trans retinoic acid. Nature 1995; 378: 681-9. [68] Renaud J.-P, Rochel N, Ruff M, Moras D. The solution structure of the human retinoic acid receptor-beta DNA-binding domain. J Biomol NMR 1995; 3: 1-17. [69] Kong W, Li Z, Xiao X, Zhao Y. Microcalorimetric investigation of the toxic action of berberine on Tetrahymena thermophila BF(5). J Basic Microbiol 2010. [70] Zhang S, Zhang B, Xing K, Zhang X, Tian X, Dai W. Inhibitory effects of golden thread (Coptis chinensis) and berberine on Microcystis aeruginosa. Water Sci Technol 2010; 61: 763-9. [71] Liu L, Yu YL, Yang JS, et al. Berberine suppresses intestinal disaccharidases with beneficial metabolic effects in diabetic states, evidences from in vivo and in vitro study. Naunyn-Schmiedebergs Arch Pharm 2010; 381: 371-81. [72] Yu Y, Liu L, Wang X, Liu X, Xie L, Wang G. Modulation of glucagon-like peptide-1 release by berberine: in vivo and in vitro studies. Biochem Pharmacol 2010; 79: 1000-6. [73] Yang HZ, Zhou MM, Zhao AH, Xing SN, Fan ZQ, Jia W. [Study on effects of baicalin, berberine and Astragalus polysaccharides and their combinative effects on aldose reductase in vitro]. Zhong Yao Cai 2009; 32: 1259-61. Received: 00 02, 2010 Revised: 00 00, 2010 Accepted: 00 00, 2010
Mini-Reviews in Medicinal Chemistry, 2011, X, 00-00 *Corresponding author: Prado-Prado, Francisco:
[email protected]. Review of Theoretical Studies for Prediction neurodegenerative inhibitors. Francisco Prado-Prado1*, Isela García2 1Department of Organic Chemistry, University of Santiago de Compostela, Spain. 2Department of Organic Chemistry, University of Vigo, Spain. Abstract: Alzheimer's disease (AD) is characterize with several pathologies this disease, amyloid plaques, composed of the ȕ-amyloid peptide and Ȗ-amyloid peptide are hallmark neuropathological lesions in Alzheimer's disease brain. Indeed, a wealth of evidence suggests that ȕ-amyloid is central to the pathophysiology of AD and is likely to play an early role in this intractable neurodegenerative disorder. AD is the most prevalent form of dementia, and current indications show that twenty-nine million people live with AD worldwide, a figure expected rise exponentially over the coming decades. Clearly, blocking disease progression or, in the best-case scenario, preventing AD altogether would be of benefit in both social and economic terms. However, current AD therapies are merely palliative and only temporarily slow cognitive decline, and treatments that address the underlying pathologic mechanisms of AD are completely lacking. While familial AD (FAD) is caused by autosomal dominant mutations in either amyloid precursor protein (APP) or the presenilin (PS1, PS2) genes. First, we revised 2D QSAR, 3D QSAR, CoMFA, CoMSIA and Docking of ȕ and Ȗ-secretase inhibitors. Next, we review 2D QSAR, 3D QSAR, CoMFA, CoMSIA and Docking for GSK-3Į and GSK-3ȕ with different compound to find out the structural requirements. Keywords: QSAR; CoMSIA; COMFA; Docking; ȕ-secretase inhibitors; Ȗ-secretase inhibitors; Alzheimer's disease (AD). Introduction Pathologically, Alzheimer's disease (AD) is characterized by the accumulation of amyloid beta peptide (Aȕ), as fibrillar plaques and soluble oligomers in high-order association brain regions. The presence of intracellular neurofibrillary tangles, neuroinflammation, neuronal dysfunction and death further characterizes this disease. Mounting evidence suggests that Aȕ plays a critical early role in AD pathogenesis, and the basic tenant of the amyloid (or Aȕ cascade) hypothesis is that Aȕ aggregates trigger a complex pathological cascade which leads to neurodegeneration [1]. A strong genetic correlation exists between FAD and the 42 amino acid Aȕ form (Aȕ42; reviewed in [2-4]). Aȕ is derived from APP and mutations in APP and PS increase Aȕ42 production and cause FAD with nearly 100% penetrance. Down's syndrome (DS) patients, who have an extra copy of the APP gene on chromosome 21, and FAD families with a duplicated APP gene locus [5], exhibit total Aȕ overproduction and all develop early-onset AD. In FAD, the Aȕ42 increase is present years before AD symptoms arise, suggesting that Aȕ42 is likely to initiate AD pathophysiology. The robust association of Aȕ42 overproduction with FAD argues strongly in favor of a critical role for Aȕ42 in the etiology of AD, including in SAD. Fibrillar and oligomeric forms of Aȕ appear neurotoxic in vitro and in vivo. Importantly, in specific transgenic (Tg) mouse models of AD the lack of Aȕ correlates with the absence of neuronal loss and improved cognitive function [6-8]. Such data provides direct evidence for the amyloid hypothesis in vivo, and also indicates that Aȕ is directly responsible for neuronal death. Consequently, strategies to lower Aȕ42 levels in the brain are anticipated to be of therapeutic benefit in AD. Aȕ peptide is generated following the sequential cleavage of APP by ȕand Ȗ-secretase in the amyloidogenic pathway [9, 10]. Aȕ genesis may be precluded if APP is cleaved by Įsecretase within the Aȕ domain in the nonamyloidogenic pathway Figure 1. Recently, the secretases have been identified and the ȕsecretase is known to be ȕ-site APP cleaving enzyme I (BACE1) [11-14], a novel aspartyl protease. BACE1 cleavage of APP is a prerequisite for Aȕ formation. Aȕ genesis is initiated by BACE1 cleavage of APP at the Asp+1 residue of the Aȕ sequence to form the N-terminus of the peptide. This scission liberates two cleavage fragments: a secreted APP ectodomain, APPsȕ and a membrane-bound carboxyl terminal fragment (CTF). In many instances, an increase in non-amyloidogenic APP metabolism is coupled to a reciprocal decrease in the amyloidogenic processing pathway, and vice-versa, as the Įand ȕ-secretase moieties compete for APP substrate [10, 13]. In the case of Ȗ-secretase is a multisubunit protease complex, itself an integral membrane protein, those cleaves single-pass transmembrane proteins at residues within the
2 transmembrane domain. The most well-known substrate of gamma secretase is amyloid precursor protein, a large integral membrane protein that, when cleaved by both gamma and beta secretase, produces a short 39-42 amino acidpeptide called amyloid beta whose abnormally folded fibrillar form is the primary component of amyloid plaques found in the brains of Alzheimer's disease patients. Gamma secretase is also critical in the related processing of the Notch protein [15, 16]. Figure 1. APP metabolism by the secretase enzymes. Given that both secretase are the initiating enzyme in Aȕ generation, and putatively ratelimiting, it is considered a prime drug target for lowering cerebral Aȕ levels in the treatment and/or prevention of AD. Prior to its identification, numerous studies were undertaken to define the characteristics of ȕsecretase activity. Although the majority of body tissues exhibit ȕ-secretase activity [17], highest activity levels were observed in neural tissue and neuronal cell lines [18]. Indeed, ȕsecretase appeared to predominate in neurons, with the level of ȕ-secretase activity appearing lower in astrocytes [19]. Data showing that ȕsecretase efficiently cleaved only membranebound substrates [20] indicated that the enzyme was likely membrane-bound or closely associated with a [21] prevent the buildup of beta-amyloid and may help slow or stop the disease However, current AD therapies are merely palliative and only temporarily slow cognitive decline, and treatments that address the underlying pathologic mechanisms of AD are completely lacking. In the last years, a number of publications have appeared suggesting GSK-3 as a target for the treatment of AD. Two isoforms of GSK-3 exists, GSK-3Į and GSK-3ȕ, both share a high homology at their catalytic site but the Įform possess an extended N-terminus with respect to the ȕ form [22, 23]. The phosphorylation of proteins by GSK-3 is an important link in neural function [24-26]. Two are the characteristic neuropathological hallmarks of AD, Neurofibrillary Tangles (NFT´s) and increase SURGXFWLRQ RI DP\ORLG EHWD $ȕ SHSWLGHV where NFT´s are composed of highly phosphorylated form of the microtubule associated protein tau [27] and studies have shown that GSK-3 is one of the main in vivo players of phosphorylation of tau protein [28]. It has been reported that Lithium, a GSK-3 inhibitor, block production of $ȕ peptides by interfering with APP cleavage at Ȗ-secretase step, where the target for Lithium is GSK-3Į [21, 22]. Phiel et al. [21] showed that selective reduction in concentration of the Į isoform led WRDGHFUHDVHLQWKHFRQFHQWUDWLRQRI$ȕDQG $ȕSULPDU\FRQVWLWXHQWVRIDP\ORLGSODTXHV in AD. Thus inhibition of GSK-3Į could potentially provide dual therapy against AD, preventing the buildup of amyloid plaques and of neurofibrillary tangles [21, 29, 30]. GSK-3ȕ is a serine/threonine kinase and is thought to be a key factor for aberrant tau phosphorylation [31]. Activated GSK-3ȕ coexists with progression of NFT´s and neurodegeneration in the AD brain [32-34]. A conditional GSK-3ȕ overexpressing transgenic mouse exhibits persistent tau hyperphosphorylation, pretangle-like somatodendritic localization of tau, neuronal death in hippocampus and cognitive deficits [35, 36]. These studies suggest that GSK-3ȕ is associated with AD progression, and GSK-3ȕ inhibition is expected to be a promising therapeutic approach for AD. In this sense, quantitative structure-activity relationships (QSAR) could play an important role in studying these ȕ and Ȗ-secretase inhibitors. QSAR models are necessary in order to guide the ȕ and Ȗ-secretase inhibitors. On the other hand, QSAR models can be used to explore the relationships between the structural spaces of compounds as inhibitors for specific enzymes, such as MAO inhibitors [37], HIV-1 integrase inhibitors [38], and/or protease inhibitors [39] or tyrosinase inhibitors [40-42]. In fact, Almost all QSAR techniques are based on the use of molecular descriptors, which are numerical series that codify useful chemical information and enable correlations between statistical and biological properties [43, 44]. Recently, the field has moved from small molecules to proteins and other systems. For instance, González-Díaz et al. discussed the use of these methods but only from the point of view of proteins [45]. Later, some groups published different papers in one special issue on QSAR but also restricted to the field of protein and proteomics [46-52]. In other recent issue, guestedited by González-Díaz [53] appeared a series of papers devoted to QSAR/QSPR techniques for low-molecular-weight drugs [53-62]. Most recently, Prado-Prado et al. [63] published a mtQSAR for anti-parasitic drugs. This year was
3 published other issue [64] focused on QSAR/QSPR models and graph theory used to approach Drug ADMET processes and Metabolomics [65-72]. Last, one of the most recent issues published is devoted to discuss the applications of QSAR in Pharmaceutical Design [73-82]. In the present work, we firstly revised the state-of-art on the design, synthesis, and biological assay of ȕ and Ȗ-secretase inhibitors. Next, we review previous works based on 2DQSAR, 3D-QSAR, CoMFA, CoMSIA and Docking techniques, which studied different compounds to find out the structural requirements. The topics reviewed, discussed, and/or reported in this paper are: 1. Studies of Ȗ-secretase inhibitors 1.1 Synthesis and Theoretical studies of Ȗ-secretase inhibitors 1.2 Design and synthesis of pyridine derivatives as BACE-1 inhibitors 1.3 Discover non-peptide inhibitors of BACE-1 using VHTS 1.4 Distinct Pharmacological Effects of Inhibitors of Ȗ-Secretase 1.5 3D-QSAR studies of Ȗ-secretase inhibitors 1.6 MD simulaWLRQV RI $ȕ ILEULO interactions 2. Studies of ȕ-secretase inhibitors 2.1 Synthesis, Theoretical studies and Biological Assay of ȕ-secretase inhibitors 2.2 Models of novel pyridinium-based potent ȕ-secretase inhibitory leads 2.3 CoMFA & CoMSIA of hydroxyethylamine derivatives as BACE-1 inhibitors 2.4 Virtual Screening and Protonation States at Asp32 and Asp228 2.5 Docking scoring function based on 2Ddescriptors 2.6 Induced-Fit Docking of Peptidic and Pseudo-peptidic BACE-1 inhibitors 3. Studies of GSK-3Į inhibitors 3.1. 2D-QSAR for 3-anilino-4phenylmaleimides 3.2. 3D-QSAR and docking of 3-anilino-4-phenylmaleimides 3.3. QSAR studies of Some GSK-3Į Inhibitory pyrimidines 4. Studies of GSK-3ȕ inhibitors 4.1. Design, synthesis and SAR of oxadiazole derivatives 4.2. Linear/Nonlinear Regression Methods for Prediction of Glycogen 4.3. Molecular modeling, docking and 3DQSAR studies for maleimides 4.4. Molecuar docking and biological testing of inhibitors of GSK-3ȕ 4.5. 3D-QSAR Modelling of Paullones 4.6. Modeling of Binding Mode of Benzo[e]isoindole-1,3-diones Discusion QSAR and Theoretical studies for neurodegene inhibitors In this section we updated the contents presented in our recent review published in Current Drugs Metabolism [83]. The high number of possible candidates to ȕ-secretase inhibitors creates the necessity of Quantitative StructureActivity Relationship models in order to guide the ȕ-secretase inhibitor synthesis. In this work, we revised different computational studies for a very large and heterogeneous series of ȕ-secretase. First, we revised QSAR studies with conceptual parameters. Next, using method of regression analysis; and QSAR studies in order to understand the essential structural requirement for binding with receptor. Next, we review 3D QSAR, CoMFA and CoMSIA with different compound to find out the structural requirements for ȕsecretase inhibitors. 1. Studies of Ȗ-secretase inhibitors 1.1. Design and synthesis of pyridine derivatives as BACE-1 inhibitors. Soo-Jeong Choi, et al. [84], had designed and synthesized of 1,4-dihydropyridine derivatives as BACE-1 inhibitors using a 1,4dihydropyridine (DHP) scaffold. They had synthesized new inhibitors of BACE-1 (the protein that has been shown to be an attractive therapeutic target in Alzheimer's disease) by modifying the known BACE inhibitor 2 containing a hydroxyethylamine (HEA) motif, see Figure 2. Using structure-based drug design based on computer-aided molecular docking, the isophthalamide ring was replaced with a 1,4dihydropyridine ring as a brain-targeting strategy. After their synthesis, the dihydropyridine derivatives were evaluated their BACE-1-inhibitory activities using a cell-based, reporter gene assay system that measures the cleavage of alkaline phosphatase (AP)-APP fusion protein by BACE-1.
4 Reagent: (a) NH4OAc, Ethanol, 90 ºC, 24 h, 98%; (b) Methane sulfonylchloride, NaH, DMF, 0-60 ºC, 4 h, 35%; (c) AlCl3, Anisole, DCM, -50 ºC to RT, 2 h, 21%; (d) R-methylbenzylamine, PyBOP, DIPA, DCM, 1 h, 70%; (e) AlCl3, Anisole, DCM, 50 ˇ C to RT, 2 h, 30%; (f) Compound 5a, PyBOP, DIPEA, DCM, 15 min, 65%. Figure 2. Synthesis of 1-methylsulfonamide-2,6-dimethyl-1,4-dihydropyridine derivatives. Molecular modeling was performed using CDOCKER, a CHARMm based molecular dynamics docking algorithm Discovery Studio 2.0 (Accelrys). The BACE-1 structure cocrystallized was obtained from the PDBdata bank (PDB code: 2B8L). A protein clean process and a CHARMm-force field were sequentially applied. The area around 2 was chosen as the active site, with the radius set as at 8-A. After removing 2 from the structure of the complex, a binding sphere in the three axis directions was constructed around the active site. Al default parameters were used in the docking process. CHARMm based molecular dynamics (1000 steps) were used to generate random ligand conformations and the position of any ligand was optimized in the binding site using rigid body rotation followed by simulated annealing at 700 K. Final energy minimization was set as the full potential mode. The final binding conformation was determined on the basis of energy (see Figure 3). Figure 3. Overlay of inhibitor 2 (green) and inhibitor 9a in the BACE-1 active site, b. Interaccions of 9a in the active site of BACE1Hydrogen bonds are shown with dotted lines (For interpretation of the references to coloue in this figure legend, the reader is referred to the web version of this article). Based on molecular docking results, we designed 1,4-DHP derivatives as BACE-1 inhibitors using five strategies. These were replacement of the bulky a-methylbenzamide group in 2 with a benzyl ester or smaller acetyl group for binding in the S3 pocket; modification
5 of the sulfonamide group in the aromatic scaffold of 2 with alkyl ester or amide groups, maintaining the important hydrogen bonding with Asn233 in the S2 binding pocket; incorporation of additional hydrophobic interactions into the S1 binding pocket by introduction of alkyl or aryl groups, including methyl, ethyl, propyl, isopropyl, and phenyl groups; alteration of the cyclopropyl group at the R4 position to other aromatic groups, thus changing hydrophobic interactions in the S20 binding pocket by extension toward the primeside of the enzyme; and, alterations at the 2 and 6 positions of the 1,4-DHP scaffold by synthesis of 2-monomethyl and 2,6-unsubstituted analogs. Their results show that most of the 1,4-DHP analogs showed BACE-1-inhibitory activities with IC50 values in the range 8e30 mM, suggesting that the 1,4-DHP skeleton may be utilized to develop brain-targeting BACE-1 inhibitors. 1.2. Discover non-peptide inhibitors of BACE-1 using VHTS A novel series of isatin-based inhibitors of b-secretase (BACE-1) using a virtual highthroughput screening approach have identified by Yi Moka et al. [85]. Structureactivity relationship studies revealed structural features important for inhibition. Docking studies suggest these inhibitors may bind within the BACE-1 active sitethrough H-bonding interactions involving the catalytic aspartate residues. They used AutoDock to separately dock the two proposed favored conformations 19 and 20 of compound 1 into the active site of BACE-1 (PDB code 1M4H). While the docking of conformer 19 did not give any solutions consistent with the observed biological activity, docking of conformer 20 revealed a binding pose which was consistent with the observed activities of compounds 1-8 (see Figure 4). This figure shows the lowest-energy binding pose identified for conformer 20 within BACE1. It is noteworthy that an analogous binding pose was also identified using eHiTS. The acetamide moiety is predicted to occupy the catalytic site, with the acetamide N-H acting as an H-bond donor to the catalytic residue Asp228 (H-bond length = 1.86 ÅA 0). The phenol unit is predicted to make an Hbond contact with the backbone nitrogen of Thr232 (H-bond length = 2.20 ÅA 0) and to partly occupy the P2 substrate pocket. This feature implies the phenol might be involved in both the formation of the intramolecular Hbonding network, and also in intermolecular Hbonding interactions with the enzyme. The nitro group in 1 is predicted to extend into the P4 pocket, possibly participating in weak Hbonding with the side chain of Arg307 (H-bond length = 2.13 ÅA 0), consistent with the slight decrease in the binding affinity exhibited by compound. Figure 4. Binding pose of 1 (corresponding to conformer 20) in the BACE-1 active site generated using AutoDock. In summary, the authors, using the virtual high-throughput screening software eHiTS, have discovered a novel non-peptidic inhibitor of BACE-1 based on an isatin motif. Studies of the biological activity of structural variants in combination with in silico docking suggest the inhibitor adopts a planar conformation, which is stabilized by intramolecular H-bonding from the phenolic moiety. Additionally, binding to BACE-1 appears to involve H-bonding interactions between the p-tolylamide of 1 and the catalytic residue Asp228. A recent report detailing the discovery of a series of potent small molecule BACE-1 inhibitors compares the ligand efficiency (LE) of a range of reported inhibitors of BACE-1. In this study, the authors noted that despite the high potency of the previously reported peptidebased BACE-1 inhibitors such as OM99-2 (Ki = 1.6 nM), the relatively high molecular weights of these systems (e.g., OM99-2 has a molecular weight of 893) often result in them having relatively poor ligand efficiency (e.g., LE = 0.19 for OM99-2). For this study, although still somewhat below the preferred minimum value of LE = 0.3, compound 1 (molecular weight = 461) has LE = 0.22 and therefore, is closer to the preferred value than the potent but considerably larger peptidic inhibitors reported previously. They have demonstrated that eHiTS is a powerful screening tool to identify biologically active compounds quickly and efficiently. 1.3 Distinct Pharmacological Effects of Inhibitors of Ȗ-Secretase Toru Sato, et al. [86], have report that helical peptide inhibitors designed to mimic SPP substrates and interact with the SPP initial substrate-ELQGLQJ VLWH WKH ³GRFNLQJ VLWH´ inhibit both SPP and Ȗ-secretase, but with submicromolar potency for SPP. SPP was labeled by helical peptide and transition-state analogue affinity probes but at distinct sites.
12 4.5. 3D-QSAR Modelling of Paullones D. I. Osolodkin et al. [98] realized 3DQSAR study allows one to suggest ways of modification of the molecule to increase its physiological activity. Comparative molecular field analysis (CoMFA) [7] and comparative molecular similarity indices analysis (CoMSIA) [8] are among the most widely used 3D-QSAR methods. The energy of van der Waals and electrostatic interactions of a probe atom (with the charge +1) with molecules of the training set (CoMFA) or the electrostatic, van der Waals, hydrophobic, and donor/acceptor similarity indices (CoMSIA) were used as descriptors. The equation for activity prediction was derived using the partial least squares (PLS) method. The ability of graphic representation of PLS model coefficients was the advantage of the methods and allowed the user to suggest substitutions affecting activity and/or selectivity of the molecules. They had built a new 3DQSAR model for GSK-3ȕ inhibition by paullones by means of CoMFA method. This model can be used as a guide for design of new paullone GSK-3ȕ inhibitors. 4.6. Modeling of Binding Mode of Benzo[e]isoindole-1,3-diones Zhen Yang et al. [99] synthesized benzo[e]isoindole-1,3-dione derivatives, and the effects on GSK-3ȕ activity and zebrafish embryo growth were evaluated. A series of derivatives showed obvious inhibitory activity against GSK-3ȕ. The most potent inhibitor, 7,8dimethoxy-5-methylbenzo[e]isoindole-1,3dione, showed nanomolar IC50 and obvious phenotype on zebrafish embryo growth associated with the inhibition of GSK-3ȕ at low micromolar concentration. The interaction mode between this compound and GSK-3ȕ was characterized by computational modeling. To rationalize the structure-activity relationships of these compounds, the binding modes of the most potent inhibitors 8a and 8b (see Figure 11) were modeled using docking simulations. Compounds 8a and 8b were docked into the ATP binding site of GSK-3ȕ, and the binding modes of lowest energy were analyzed. Compounds 8a and 8b fit the ATP pocket of GSK-3ȕ well. The maleimide motif of type II formed a pair of hydrogen bonds with the hinge region (Glu133 and Val135) of GSK-3ȕ, similar to the binding mode of other known maleimides GSK-3ȕ inhibitors. The two methoxy oxygen atoms formed another two hydrogen bonds with the positively charged Lys85. The methyl group of the methoxy at C-8 position docked to the small back cleft of GSK-3ȕ. This binding mode explicitly explained the important role of the two methoxy groups at C-7 and C-8 positions. Other result was the 4-ethyl group of 8b docks to the minor hydrophobic pocket formed by Ile62 and Val70 in the front of the ATP binding site of GSK-3ȕ (Figure 12), which contributed to its higher binding affinity compared to 8a. The docking results also provided a template to understand the structure-activity relationships of other compounds. Figure 11. Structure of 8a and 8b Figure 12. Docked binding modes of compounds 8b in the ATP binding site of GSK-3ȕ. CONCLUSIONS Theoretical studies such as QSAR models have become a very useful tool in this context to substantially reduce time and resources consuming experiments. The functions of ȕ and Ȗsecretase and its implication in Alzheimer's disease have triggered an active search for potent and selective ȕ and Ȗ-secretase inhibitors. In this paper we can see that the development of theoretical and QSAR models to study Ȗ-secretase inhibitors are usually not many achieved so far, and most of these works present docking studies. Watching this situation we need to develop QSAR models with Ȗ-secretase inhibitors. In this sense, QSAR could play an important role in studying these Ȗ-secretase inhibitors. QSARs can be used as predictive tools for the development of molecules. In this work we developed a new ANN RBF model using the ModesLab descriptors, based on a large database using about 10,000 different drugs obtained from the ChEMBL server. Acknowledgements Prado-Prado F. thanks sponsorships for research position at the University of Santiago de Compostela from Angeles Alvariño, Xunta de Galicia. All authors acknowledge the Project 07CSA008203PR. References [1] Golde, T. E., Dickson, D. and Hutton, M., Filling the gaps in the abeta cascade hypothesis of Alzheimer's
13 disease. Curr Alzheimer Res,2006, 3, 421-30 [2] Hutton, M., Perez-Tur, J. and Hardy, J., Genetics of Alzheimer's disease. Essays Biochem,1998, 33, 117-31 [3] Younkin, S. G., The role of A beta 42 in Alzheimer's disease. J Physiol Paris, 1998, 92, 289-92 [4] Sisodia, S. S., Alzheimer's disease: perspectives for the new millennium. J Clin Invest,1999, 104, 1169-70 [5] Rovelet-Lecrux, A., Hannequin, D., Raux, G., Le Meur, N., Laquerriere, A., Vital, A., Dumanchin, C., Feuillette, S., Brice, A., Vercelletto, M., Dubas, F., Frebourg, T. and Campion, D., APP locus duplication causes autosomal dominant early-onset Alzheimer disease with cerebral amyloid angiopathy. Nat Genet,2006, 38, 24-6 [6] Ohno, M., Sametsky, E. A., Younkin, L. H., Oakley, H., Younkin, S. G., Citron, M., Vassar, R. and Disterhoft, J. F., BACE1 deficiency rescues memory deficits and cholinergic dysfunction in a mouse model of Alzheimer's disease. Neuron,2004, 41, 27-33 [7] Ohno, M., Cole, S. L., Yasvoina, M., Zhao, J., Citron, M., Berry, R., Disterhoft, J. F. and Vassar, R., BACE1 gene deletion prevents neuron loss and memory deficits in 5XFAD APP/PS1 transgenic mice. Neurobiol Dis,2007, 26, 134-45 [8] Laird, F. M., Cai, H., Savonenko, A. V., Farah, M. H., He, K., Melnikova, T., Wen, H., Chiang, H. C., Xu, G., Koliatsos, V. E., Borchelt, D. R., Price, D. L., Lee, H. K. and Wong, P. C., BACE1, a major determinant of selective vulnerability of the brain to amyloid-beta amyloidogenesis, is essential for cognitive, emotional, and synaptic functions. J Neurosci,2005, 25, 11693-709 [9] Selkoe, D. J., Alzheimer's disease: genes, proteins, and therapy. Physiol Rev,2001, 81, 741-66 [10] Vassar, R., BACE1: the beta-secretase enzyme in Alzheimer's disease. J Mol Neurosci,2004, 23, 105-14 [11] Hussain, I., Powell, D., Howlett, D. R., Tew, D. G., Meek, T. D., Chapman, C., Gloger, I. S., Murphy, K. E., Southan, C. D., Ryan, D. M., Smith, T. S., Simmons, D. L., Walsh, F. S., Dingwall, C. and Christie, G., Identification of a novel aspartic protease (Asp 2) as beta-secretase. Mol Cell Neurosci,1999, 14, 419-27 [12] Sinha, S., Anderson, J. P., Barbour, R., Basi, G. S., Caccavello, R., Davis, D., Doan, M., Dovey, H. F., Frigon, N., Hong, J., Jacobson-Croak, K., Jewett, N., Keim, P., Knops, J., Lieberburg, I., Power, M., Tan, H., Tatsuno, G., Tung, J., Schenk, D., Seubert, P., Suomensaari, S. M., Wang, S., Walker, D., Zhao, J., McConlogue, L. and John, V., Purification and cloning of amyloid precursor protein beta-secretase from human brain. Nature,1999, 402, 53740 [13] Vassar, R., Bennett, B. D., Babu-Khan, S., Kahn, S., Mendiaz, E. A., Denis, P., Teplow, D. B., Ross, S., Amarante, P., Loeloff, R., Luo, Y., Fisher, S., Fuller, J., Edenson, S., Lile, J., Jarosinski, M. A., Biere, A. L., Curran, E., Burgess, T., Louis, J. C., Collins, F., Treanor, J., Rogers, G. and Citron, M., Betasecretase cleavage of Alzheimer's amyloid precursor protein by the transmembrane aspartic protease BACE. Science,1999, 286, 735-41 [14] Yan, R., Bienkowski, M. J., Shuck, M. E., Miao, H., Tory, M. C., Pauley, A. M., Brashier, J. R., Stratman, N. C., Mathews, W. R., Buhl, A. E., Carter, D. B., Tomasselli, A. G., Parodi, L. A., Heinrikson, R. L. and Gurney, M. E., Membrane-anchored aspartyl protease with Alzheimer's disease beta-secretase activity. Nature,1999, 402, 533-7 [15] Goate, A., Chartier-Harlin, M. C., Mullan, M., Brown, J., Crawford, F., Fidani, L., Giuffra, L., Haynes, A., Irving, N., James, L. and et al., Segregation of a missense mutation in the amyloid precursor protein gene with familial Alzheimer's disease. Nature,1991, 349, 704-6 [16] Schellenberg, G. D., Bird, T. D., Wijsman, E. M., Orr, H. T., Anderson, L., Nemens, E., White, J. A., Bonnycastle, L., Weber, J. L., Alonso, M. E. and et al., Genetic linkage evidence for a familial Alzheimer's disease locus on chromosome 14. Science,1992, 258, 668-71 [17] Haass, C., Schlossmacher, M. G., Hung, A. Y., Vigo-Pelfrey, C., Mellon, A., Ostaszewski, B. L., Lieberburg, I., Koo, E. H., Schenk, D., Teplow, D. B. and et al., Amyloid beta-peptide is produced by cultured cells during normal metabolism. Nature,1992, 359, 322-5
14 [18] Seubert, P., Oltersdorf, T., Lee, M. G., Barbour, R., Blomquist, C., Davis, D. L., Bryant, K., Fritz, L. C., Galasko, D., Thal, L. J. and et al., Secretion of beta-amyloid precursor protein cleaved at the amino terminus of the betaamyloid peptide. Nature,1993, 361, 260-3 [19] Zhao, J., Paganini, L., Mucke, L., Gordon, M., Refolo, L., Carman, M., Sinha, S., Oltersdorf, T., Lieberburg, I. and McConlogue, L., Beta-secretase processing of the beta-amyloid precursor protein in transgenic mice is efficient in neurons but inefficient in astrocytes. J Biol Chem,1996, 271, 31407-11 [20] Citron, M., Teplow, D. B. and Selkoe, D. J., Generation of amyloid beta protein from its precursor is sequence specific. Neuron,1995, 14, 661-70 [21] Phiel, C. J., Wilson, C. A., Lee, V. M.- Y. and Klein, P. S., GSK-ĮUHJXODWHV SURGXFWLRQ RI $O]KHLPHU¶V GLVHDVH amyloid-ȕSHSWLGHV Nature,2003, 423, 435-439 [22] Jamloki, A., Karthikeyan, C. and Sharma, S. K., QSAR Studies on Some GSK-Į ,QKLELWRU\ 6-aryl-pyrazolo- (3,4-b)pyrimidines. Asian Journal of Biochemistry,2006, 1, 236-243 [23] Ali, A., Hoeflich, K. P. and Woodgett, J. R., Glycogen Synthase Kinase-3: Properties, Functions, and Regulation. Chem Rev,2001, 101, 2527-2540 [24] Martinez, A., Castro, A., Dorronsoro, I. and Alonso, M., Glycogen synthase kinase 3 (GSK-3) inhibitors as new promising drugs for diabetes, neurodegeneration, cancer, and inflammation. Med Res Rev,2002, 22, 373-84 [25] Martinez, A., Alonso, M., Castro, A., Perez, C. and Moreno, F. J., First NonATP Competitive Glycogen Synthase Kinase 3B (GSK-3B) Inhibitors: Thiadazolidinones (TDZD) as Potential Drugs for the Treatment of Alzheimer´s Disease. J Med Chem, 2002, 45, 1292-1299 [26] Hagit, E.-F., Glycogen synthase kinase 3: an emerging therapeutic target. Trends in Molecular Medicine,2002, 8, 126-132 [27] Lee, V. M., Goedert, M. and Trojanowski, J. Q., Neurodegenerative tauopathies. Ann Rev Neurosci,2001, 24, 1121-1159 [28] Flaherty, D. B., Sorrea, P. J., Tomasienicz, G. H. and Wood, G. J., Phosphorilation of human tau protein by microtubule-associated kinases: GSK3beta and cdk5 are key participants. J Neuro Sci Res,2000, 62, 463-472 [29] Bhat, R. V., Budd Haeberlein, S. L. and Avila, J. J., Glycogen synthase kinase 3: a drug target for CNS therapies. J Neurochem,2004, 89, 1313-1317 [30] Sivaprakasam, P., Xiea, A. and Doerksen, R. J., Probing the physicochemical and structural requirements for glycogen synthase kinase-3a inhibition: 2D-QSAR for 3anilino-4-phenylmaleimides. Bioorganic & Medicinal Chemistry Letters,2006, 14, 8210-8218 [31] Ishiguro, K., Takamatsu, M., Tomizawa, K., Omori, A., Takahashi, M., Arioka, M., Uchida, T. and Imahori, K., Tau protein kinase I converts normal tau protein into A68like component of paired helical filaments. J. Biol. Chem.,1992, 267, 10897-10901 [32] Pei, J. J., Tanaka, T., Tung, Y. C., Braak, E., Iqbal, K. and Grundke-Iqbal, I., Distribution, Levels, and Activity of Glycogen Synthase Kinase-3 in the Alzheimer Disease Brain. J. Neuropathol. Exp. Neurol.,1997, 56, 70-78 [33] Pei, J. J., Braak, H., Grundke-Iqbal, I., Iqbal, K., Winblad, B. and Cowburn, R. F., Distribution of active glycogen synthase kinase 3beta (GSK-3beta) in brains staged for Alzheimer disease neurofibrillary changes. J. Neuropathol. Exp. Neurol.,1999, 58, 1010-1019 [34] Baum, L., Hansen, L., Masliah, E. and Saitoh, T., Glycogen synthase kinase 3 alteration in Alzheimer disease is related to neurofibrillary tangle formation. Mol. Chem. Neuropathol., 1996, 29, 253-261 [35] Lucas, J. J., Hernandez, F., GomezRamos, P., Moran, M. A., Hen, R. and Avila, J., Decreased nuclear -catenin, tau hyperphosphorylation and neurodegeneration in GSK-3 conditional transgenic mice. EMBO J., 2001, 20, 27-39 [36] Hernandez, F., Borrell, J., Guaza, C., Avila, J. and Lucas, J. J., Spatial learning deficit in transgenic mice that conditionally over-express GSK-3beta in the brain but do not form tau
15 filaments. J. Neurochem.,2002, 83, 1529-1533 [37] Santana, L., Uriarte, E., GonzálezDíaz, H., Zagotto, G., Soto-Otero, R. and Mendez-Alvarez, E., A QSAR model for in silico screening of MAOA inhibitors. Prediction, synthesis, and biological assay of novel coumarins. J Med Chem,2006, 49, 1149-56 [38] Marrero-Ponce, Y., Linear indices of the "molecular pseudograph's atom adjacency matrix": definition, significance-interpretation, and application to QSAR analysis of flavone derivatives as HIV-1 integrase inhibitors. J Chem Inf Comput Sci, 2004, 44, 2010-26 [39] Vilar, S., Santana, L. and Uriarte, E., Probabilistic neural network model for the in silico evaluation of anti-HIV activity and mechanism of action. J Med Chem 2006, 49, 1118-1124 [40] Marrero-Ponce, Y., Khan, M. T., Casanola Martin, G. M., Ather, A., Sultankhodzhaev, M. N., Torrens, F. and Rotondo, R., Prediction of tyrosinase inhibition activity using atom-based bilinear indices. ChemMedChem,2007, 2, 449-78 [41] Casanola-Martin, G. M., MarreroPonce, Y., Khan, M. T., Ather, A., Sultan, S., Torrens, F. and Rotondo, R., TOMOCOMD-CARDD descriptorsbased virtual screening of tyrosinase inhibitors: evaluation of different classification model combinations using bond-based linear indices. Bioorg Med Chem,2007, 15, 1483-503 [42] Casanola-Martin, G. M., MarreroPonce, Y., Khan, M. T., Ather, A., Khan, K. M., Torrens, F. and Rotondo, R., Dragon method for finding novel tyrosinase inhibitors: Biosilico identification and experimental in vitro assays. Eur J Med Chem,2007, 42, 1370-81 [43] Nunez, M. B., Maguna, F. P., Okulik, N. B. and Castro, E. A., QSAR modeling of the MAO inhibitory activity of xanthones derivatives. Bioorg Med Chem Lett,2004, 14, 5611-5617 [44] Todeschini, R. and Consonni, V., Handbook of Molecular Descriptors. Wiley VCH,2000, [45] González-Díaz, H., González-Díaz, Y., Santana, L., Ubeira, F. M. and Uriarte, E., Proteomics, networks and connectivity indices. Proteomics,2008, 8, 750-778 [46] Zhao, C. J. and Dai, Q. Y., [Recent advances in study of antinociceptive conotoxins]. Yao Xue Xue Bao,2009, 44, 561-5 [47] Jacob, R. B. and McDougal, O. M., The M-superfamily of conotoxins: a review. Cellular and Molecular Life Sciences,2010, 67, 17-27 [48] Giuliani, A., Di Paola, L. and Setola, R., Proteins as Networks: A Mesoscopic Approach Using Haemoglobin Molecule as Case Study. Curr Proteomics,2009, 6, 235-245 [49] Vilar, S., Gonzalez-Diaz, H., Santana, L. and Uriarte, E., A network-QSAR model for prediction of geneticcomponent biomarkers in human colorectal cancer. Journal of Theoretical Biology,2009, 261, 449-58 [50] Concu, R., Dea-Ayuela, M. A., PerezMontoto, L. G., Prado-Prado, F. J., Uriarte, E., Bolas-Fernandez, F., Podda, G., Pazos, A., Munteanu, C. R., Ubeira, F. M. and Gonzalez-Diaz, H., 3D entropy and moments prediction of enzyme classes and experimentaltheoretic study of peptide fingerprints in Leishmania parasites. Biochimica et Biophysica Acta,2009, 1794, 1784-94 [51] Torrens, F. and Castellano, G., Topological Charge-Transfer Indices: From Small Molecules to Proteins. Curr Proteomics,2009, 204-213 [52] Vázquez, J. M., Aguiar, V., Seoane, J. A., Freire, A., Serantes, J. A., Dorado, J., Pazos, A. and Munteanu, C. R., Star Graphs of Protein Sequences and Proteome Mass Spectra in Cancer Prediction. Curr Proteomics,2009, 6, 275-288 [53] Gonzalez-Diaz, H., Quantitative studies on Structure-Activity and Structure-Property Relationships (QSAR/QSPR). Curr Top Med Chem, 2008, 8, 1554 [54] Ivanciuc, O., Weka machine learning for predicting the phospholipidosis inducing potential. Curr Top Med Chem,2008, 8, 1691-709 [55] Gonzalez-Díaz, H., Prado-Prado, F. and Ubeira, F. M., Predicting antimicrobial drugs and targets with the MARCH-INSIDE approach. Curr Top Med Chem,2008, 8, 1676-90 [56] Duardo-Sanchez, A., Patlewicz, G. and Lopez-Diaz, A., Current topics on software use in medicinal chemistry: intellectual property, taxes, and regulatory issues. Curr Top Med Chem, 2008, 8, 1666-75
16 [57] Wang, J. F., Wei, D. Q. and Chou, K. C., Drug candidates from traditional chinese medicines. Curr Top Med Chem,2008, 8, 1656-65 [58] Helguera, A. M., Combes, R. D., Gonzalez, M. P. and Cordeiro, M. N., Applications of 2D descriptors in drug design: a DRAGON tale. Curr Top Med Chem,2008, 8, 1628-55 [59] Gonzalez, M. P., Teran, C., Saiz-Urra, L. and Teijeira, M., Variable selection methods in QSAR: an overview. Curr Top Med Chem,2008, 8, 1606-27 [60] Caballero, J. and Fernandez, M., Artificial neural networks from MATLAB in medicinal chemistry. Bayesian-regularized genetic neural networks (BRGNN): application to the prediction of the antagonistic activity against human platelet thrombin receptor (PAR-1). Curr Top Med Chem,2008, 8, 1580-605 [61] Wang, J. F., Wei, D. Q. and Chou, K. C., Pharmacogenomics and personalized use of drugs. Curr Top Med Chem,2008, 8, 1573-9 [62] Vilar, S., Cozza, G. and Moro, S., Medicinal chemistry and the molecular operating environment (MOE): application of QSAR and molecular docking to drug discovery. Curr Top Med Chem,2008, 8, 1555-72 [63] Prado-Prado, F. J., Garcia-Mera, X. and Gonzalez-Diaz, H., Multi-target spectral moment QSAR versus ANN for antiparasitic drugs against different parasite species. Bioorg Med Chem, 2010, 18, 2225-31 [64] Gonzalez-Diaz, H., Network topological indices, drug metabolism, and distribution. Curr Drug Metab, 11, 283-4 [65] Khan, M. T., Predictions of the ADMET properties of candidate drug molecules utilizing different QSAR/QSPR modelling approaches. Curr Drug Metab, 11, 285-95 [66] Mrabet, Y. and Semmar, N., Mathematical methods to analysis of topology, functional variability and evolution of metabolic systems based on different decomposition concepts. Curr Drug Metab, 11, 315-41 [67] Martinez-Romero, M., Vazquez-Naya, J. M., Rabunal, J. R., Pita-Fernandez, S., Macenlle, R., Castro-Alvarino, J., Lopez-Roses, L., Ulla, J. L., MartinezCalvo, A. V., Vazquez, S., Pereira, J., Porto-Pazos, A. B., Dorado, J., Pazos, A. and Munteanu, C. R., Artificial intelligence techniques for colorectal cancer drug metabolism: ontology and complex network. Curr Drug Metab, 11, 347-68 [68] Zhong, W. Z., Zhan, J., Kang, P. and Yamazaki, S., Gender specific drug metabolism of PF-02341066 in rats-- role of sulfoconjugation. Curr Drug Metab, 11, 296-306 [69] Wang, J. F. and Chou, K. C., Molecular modeling of cytochrome P450 and drug metabolism. Curr Drug Metab, 11, 342-6 [70] Gonzalez-Diaz, H., Duardo-Sanchez, A., Ubeira, F. M., Prado-Prado, F., Perez-Montoto, L. G., Concu, R., Podda, G. and Shen, B., Review of MARCH-INSIDE & complex networks prediction of drugs: ADMET, anti-parasite activity, metabolizing enzymes and cardiotoxicity proteome biomarkers. Curr Drug Metab, 11, 379-406 [71] Garcia, I., Diop, Y. F. and Gomez, G., QSAR & complex network study of the HMGR inhibitors structural diversity. Curr Drug Metab, 11, 307-14 [72] Chou, K. C., Graphic rule for drug metabolism systems. Curr Drug Metab, 11, 369-78 [73] Concu, R., Podda, G., Ubeira, F. M. and Gonzalez-Diaz, H., Review of QSAR Models for Enzyme Classes of Drug Targets: Theoretical Background and Applications in Parasites, Hosts, and other Organisms. Current Pharmaceutical Design,2010, 16, 2710-23 [74] Estrada, E., Molina, E., Nodarse, D. and Uriarte, E., Structural Contributions of Substrates to their Binding to P-Glycoprotein. A TOPSMODE Approach. Current Pharmaceutical Design,2010, 16, 2676-709 [75] Garcia, I., Fall, Y. and Gomez, G., QSAR, Docking, and CoMFA Studies of GSK3 Inhibitors. Current Pharmaceutical Design,2010, 16, 2666-75 [76] González-Díaz, H., QSAR and Complex Networks in Pharmaceutical Design, Microbiology, Parasitology, Toxicology, Cancer, and Neurosciences. Current Pharmaceutical Design,2010, 16, 2598-600 [77] Gonzalez-Diaz, H., Romaris, F., Duardo-Sanchez, A., Perez-Mototo, L. G., Prado-Prado, F., Patlewicz, G. and
17 Ubeira, F. M., Predicting drugs and proteins in parasite infections with topological indices of complex networks: theoretical backgrounds, aplications, and legal issues. Current Pharmaceutical Design,2010, 16, 2737-64 [78] Marrero-Ponce, Y., Casanola-Martin, G. M., Khan, M. T., Torrens, F., Rescigno, A. and Abad, C., LigandBased Computer-Aided Discovery of Tyrosinase Inhibitors. Applications of the TOMOCOMD-CARDD Method to the Elucidation of New Compounds. Current Pharmaceutical Design,2010, 16, 2601-24 [79] Munteanu, C. R., Fernandez-Blanco, E., Seoane, J. A., Izquierdo-Novo, P., Rodriguez-Fernandez, J. A., PrietoGonzalez, J. M., Rabunal, J. R. and Pazos, A., Drug Discovery and Design for Complex Diseases through QSAR Computational Methods. Current Pharmaceutical Design,2010, 16, 2640-55 [80] Roy, K. and Ghosh, G., Exploring QSARs with Extended Topochemical Atom (ETA) Indices for Modeling Chemical and Drug Toxicity. Current Pharmaceutical Design,2010, 16, 2625-39 [81] Speck-Planche, A., Scotti, M. T. and de Paulo-Emerenciano, V., Current pharmaceutical design of antituberculosis drugs: future perspectives. Current Pharmaceutical Design,2010, 16, 2656-65 [82] Vazquez-Naya, J. M., MartinezRomero, M., Porto-Pazos, A. B., Novoa, F., Valladares-Ayerbes, M., Pereira, J., Munteanu, C. R. and Dorado, J., Ontologies of drug discovery and design for neurology, cardiology and oncology. Current Pharmaceutical Design,2010, 16, 2724-36 [83] Garcia, I., Diop, Y. F. and Gomez, G., QSAR & complex network study of the HMGR inhibitors structural diversity. Curr Drug Metab,2010, 11, 307-14 [84] Choi, S. J., Cho, J. H., Im, I., Lee, S. D., Jang, J. Y., Oh, Y. M., Jung, Y. K., Jeon, E. S. and Kim, Y. C., Design and synthesis of 1,4-dihydropyridine derivatives as BACE-1 inhibitors. Eur J Med Chem,2010, 45, 2578-90 [85] Yi Mok, N., Chadwick, J., Kellett, K. A., Hooper, N. M., Johnson, A. P. and Fishwick, C. W., Discovery of novel non-peptide inhibitors of BACE-1 using virtual high-throughput screening. Bioorg Med Chem Lett, 2009, 19, 6770-4 [86] Sato, T., Ananda, K., Cheng, C. I., Suh, E. J., Narayanan, S. and Wolfe, M. S., Distinct pharmacological effects of inhibitors of signal peptide peptidase and gamma-secretase. J Biol Chem, 2008, 283, 33287-95 [87] Sammi, T., Silakari, O. and Ravikumar, M., Three-dimensional quantitative structure-activity relationship (3D-QSAR) studies of various benzodiazepine analogues of gamma-secretase inhibitors. J Mol Model,2009, 15, 343-8 [88] Al-Nadaf, A., Abu Sheikha, G. and Taha, M. O., Elaborate ligand-based pharmacophore exploration and QSAR analysis guide the synthesis of novel pyridinium-based potent beta-secretase inhibitory leads. Bioorg Med Chem, 2010, 18, 3088-115 [89] Pandey, A., Mungalpara, J. and Mohan, C. G., Comparative molecular field analysis and comparative molecular similarity indices analysis of hydroxyethylamine derivatives as selective human BACE-1 inhibitor. Mol Divers,2010, 14, 39-49 [90] Polgar, T. and Keseru, G. M., Virtual screening for beta-secretase (BACE1) inhibitors reveals the importance of protonation states at Asp32 and Asp228. J Med Chem,2005, 48, 374955 [91] Hetenyi, C., Paragi, G., Maran, U., Timar, Z., Karelson, M. and Penke, B., Combination of a modified scoring function with two-dimensional descriptors for calculation of binding affinities of bulky, flexible ligands to proteins. J Am Chem Soc,2006, 128, 1233-9 [92] Moitessier, N., Therrien, E. and Hanessian, S., A method for inducedfit docking, scoring, and ranking of flexible ligands. Application to peptidic and pseudopeptidic betasecretase (BACE 1) inhibitors. J Med Chem,2006, 49, 5885-94 [93] Sivaprakasam, P., Daga, P. R., Xie, A. and Doerksen, R. J., Glycogen synthase kinase-3 inhibition by 3anilino-4-phenylmaleimides: insights from 3D-QSAR and docking. J Comput Aided Mol Des,2009, 23, 113127 [94] Saitoh, M., Kunitomo, J., Kimura, E., Hayase, Y., Kobayashi, H., Uchiyama,
18 N., Kawamoto, T., Tanaka, T., Mol, C. D., Dougan, D. R., Textor, G. S., Snell, G. P. and Itoh, F., Design, synthesis and structure±activity relationships of 1,3,4-oxadiazole derivatives as novel inhibitors of glycogen synthase kinase3beta. Bioorganic & Medicinal Chemistry,2009, 17, 2017-2029 [95] Freitas, M. P., Goodarzi, M. and Jensen, R., Feature Selection and Linear/Nonlinear Regression Methods for the Accurate Prediction of Glycogen Synthase Kinase-ȕ Inhibitory Activities Journal of Chemical Information and Modeling 2009, 49, 824-832 [96] Kim, K. H., Gaisina, I., Gallier, F., Holzle, D., Blond, S. Y., Mesecar, A. and Kozikowski, A. P., Use of molecular modeling, docking, and 3DQSAR studies for the determination of the binding mode of benzofuran-3-yl- (indol-3-yl)maleimides as GSK-ȕ inhibitors. J Mol Model,2009, 15, 1463-1479 [97] Ryzhova, E. A., Koryakova, A. G., Bulanova, E. A., Mikitas, O. V., Karapetyan, R. N., Lavrovskii, Y. V. and Ivashchenko, A. V., Syntheis, Molecular Docking, and Biological Testing of New Selective Inhibitors of Glycogen Synthase Kinase 3beta. Pharmaceutical Chemistry Journal, 2009, 43, 148-153 [98] Osolodkin, D. I., Shulga, D. A., Tsareva, D. A., Oliferenko, A. A., Palyulin, V. A. and Zefirov, N. S., The Choice of Atomic Charges Calculation Scheme in 3D-QSAR Modelling of GSK-ȕ ,QKLELWLRQ E\ 3DXOORQHV Biochemistry, Biophysics and Molecular Biology,2010, 434, 274-278 [99] Zou, H., Zhou, L., Li, Y., Cui, Y., Zhong, H., Pan, Z., Yang, Z. and Quan, J., Benzo[e]isoindole-1,3-diones as Potential Inhibitors of Glycogen Synthase Kinase-3 (GSK-3). Synthesis, Kinase Inhibitory Activity, Zebrafish Phenotype, and Modeling of Binding Mode. J. Med. Chem.,2010, 53, 9941003
QSAR and complex network study of the chiral HMGR inhibitor structural diversity Isela García a , Cristian Robert Munteanu b,c , Yagamare Fall a , Generosa Gómez a , Eugenio Uriarte d , Humberto González-Díaz c,* a Department of Organic Chemistry, University of Vigo, Spain b Department of Chemistry, REQUIMTE/Faculty of Science, University of Porto, 4169-007 Porto, Portugal c Department of Microbiology and Parasitology, Faculty of Pharmacy, University of Santiago de Compostela, 15782 Santiago de Compostela, Spain d UBICA, Institute of Industrial Pharmacy, Department of Organic Chemistry, Faculty of Pharmacy, University of Santiago de Compostela, 15782 Santiago de Compostela, Spain article info Article history: Received 23 May 2008 Revised 31 October 2008 Accepted 6 November 2008 Available online 9 November 2008 Keywords: QSAR Topological indices Complex network Chiral compound Lipid-lowering agent Cholesterol level Atherosclerotic disease Anti-parasite drug Trypanosoma cruzi Chagas’ disease 3-Hydroxy-3-methyl-glutaryl coenzyme A reductase abstract Efficient drugs such as statins or mevinic acids are inhibitors of the rate-limiting enzyme of cholesterol biosynthesis, 3-hydroxy-3-methyl-glutaryl coenzyme A reductase (HMGR), an enzyme responsible for the double reduction of 3-hydroxy-3-methyl-glutaryl coenzyme A into mevalonic acid. These compounds promoted the synthesis and evaluation of new inhibitors for HMGR, named HMGRIs. The high number of possible candidates creates the necessity of Quantitative Structure–Activity Relationship models in order to guide the HMGRI synthesis. There are two main problems of the reported QSAR models: the homogeneous series of the compounds and the chirality of many candidates. In this work, we propose for the first time a QSAR model for a very large and heterogeneous series of HMGRIs. The model is based on the Topological Indices (TIs) of molecular structures. Using the predictions of this model as input, we construct the first complex network that describes the drug–drug similarity relationships for more than 1600 experimentally non-explored chiral HMGRIs isomers. We also presented a reduced version of this network (Giant Component) that contains the most representative set of chiral HMGRI candidates. The work suggests a new mixed application in the QSAR study of relevant aspects of structural diversity by using chiral/non-chiral TIs, combined with complex networks. Ó2008 Elsevier Ltd. All rights reserved. 1. Introduction Hypercholesterolemia is well-known as the primary risk factor in atherosclerotic and coronary heart diseases. 1,2 Clinical studies with lipid-lowering agents have established that the decrease of high serum cholesterol levels reduces the incidence of cardiovascular mortality. Statins and mevinic acids are two efficient drugs known as inhibitors of the rate-limiting enzyme of cholesterol biosynthesis, 3-hydroxy-3-methyl-glutaryl coenzyme A reductase (HMGR), enzyme responsible for the double reduction of 3-hydroxy-3-methyl-glutaryl coenzyme A into mevalonic acid. The drug control of this enzyme is efficient in reducing the levels of cholesterol in plasma. 3,4 The structure of statins and its derivatives is characterized by the desmethylmevalonic acid or by the lactone. The pharmacophore is connected to a lipophilic ring, such as hexahydronaphthalene, indole, pyrrole, pyrimidine or quinine, by a linking element (a two-carbon spacer). The biologically active form of mevinic acids is represented by the open chain hydroxyl-acid, which mimics the HMGR natural substrate. 5 Another interesting potential use of this type of drugs could be the control of parasite infections. Urbina et al. studied in the protozoan parasite Trypanosoma (Schizotrypanum) cruzi the anti-proliferative effects of mevinolin (lovastatin), a drug from the family of HMGR inhibitors (HMGRIs), and its ability to potentiate the action of specific ergosterol biosynthesis inhibitors, such as ketoconazole and terbinafine (both in vitro and in vivo). These results confirm the synergic action against the proliferative stages of T. cruzi (both in vitro and in vivo) and the ergosterol biosynthesis inhibition (in vivo) by acting in different points of the pathway. In addition, the medical practice suggests that mevinolin, combined with azoles, such as ketoconazole, can be used in the treatment of human Chagas’ disease. 6 The success of HMGRI drugs encourages many researchers to look for new derivatives. The high asymmetry of many interesting compounds (they are commonly present various chiral centres) makes difficult to obtain new drug candidates by low-cost efficient 0968-0896/$ - see front matter Ó2008 Elsevier Ltd. All rights reserved. doi:10.1016/j.bmc.2008.11.007 * Corresponding author. Tel.: +34 981 563100; fax: +34 981 594912. E-mail addresses: [email protected],[email protected] (H. González-Díaz). Bioorganic & Medicinal Chemistry 17 (2009) 165–175 Contents lists available at ScienceDirect Bioorganic & Medicinal Chemistry journal homepage: www.elsevier.com/locate/bmc
asymmetric organic synthesis or synthesis followed by asymmetric separation. 7 Thus, the development of computational methods becomes an important step for the exploration of the large space of potential chiral isomers and for the selection of a higher probability of HMGRI activity in the asymmetric synthesis. In Figure 1 we illustrate the asymmetric structures of some well-known HMGRIs, including those above-mentioned. In order to solve this problem, we use the Quantitative Structure–Activity Relationships (QSAR), a mathematical relationship that links in a quantitative manner the chemical structure and the pharmacological activity of compounds. 8–11 The reported models are efficient and they are based mainly on 3D descriptors (3D-QSAR) and/or based on specific homologous compound series. 12–14 One of the most efficient and fast QSAR techniques (relative to 3D-QSAR) makes use of molecular graph Topological Indices (TIs) 15–18 in order to describe the molecular structure of drugs. The TI-based QSAR (TI-QSAR) models can be used to explore in silico different biological activities in a large series of compounds. In this sense, we can use TI-QSAR to predict general activities such as: anti-fungal, 19 anti-cancer, 20 anti-parasite, 21,22 sedative/hypnotic, anti-bacterial, 23,24 analgesia, 25 hypoglycaemic activity, 26 herbicide action 27 or drug toxicity. 28 Alternatively, we can use TI-QSAR models to explore the relationships between the structural spaces of compounds as inhibitors for specific enzymes, such as MAO inhibitors, 29 HIV-1 integrase inhibitors, 30 and/or protease inhibitors 31 or tyrosinase inhibitors. 32–34 Unfortunately, the classic TIs do not consider important 3D features, such as chirality, and the biological activity of different stereo-isomers of the same molecule can be notably different. Therefore, several authors have recently defined novel Chiral TIs (CTIs), which opens a fast investigation gateway of large spaces of one/many chiral center stereo-isomers. CTI-based QSARs (CTIQSARs) are as fast as TI-QSAR and chiral sensitive because they do not assume the knowledge of 3D-structure. 35–42 However, a TI-QSAR or CTI-QSAR models for a large and heterogeneous series of HMGRIs have not been reported up to date. In general, HMGRIs, such as statins, are molecules with chiral centers. In many cases, only a low percent of the possible stereoisomers of one molecular skeleton was obtained by synthesis, characterized and assayed to know their activity profiles in vitro and in vivo. Using a CTI-QSAR, it is possible to predict the activity profiles of many derivatives of different compounds in a computer simulation. This allows us to select for synthesis the derivatives with a high predicted activity and a similar profile to other successful drugs. Once predicted the most significant derivatives, it may be highly difficult to experimentally access by organic chemical synthesis all selected derivatives (as in the case of many isomers in the entire space of stereo-isomers). Thus, it is highly important to draw some conclusions about the overall degree of similarity in activity profiles of whole spaces of compounds and their derivatives. We can achieve this goal by means of similarity/dissimilarFigure 1. Different types of chiral HMGRIs. 166 I. García et al. / Bioorg. Med. Chem. 17 (2009) 165–175
ity Complex Networks (CNs). 43 These CNs are similar to the graphs, composed by many nodes connected by edges. 44 In molecular sciences, structural units such as atoms, 45–47 aminoacids, 48 nucleotides, 49–51 proteins, 52,53 RNAs, 54 genes 55 or even organisms 56 often play the role of nodes in the CNs. Complementarily, the chemical bonds (in case of classic molecular graph), chemical reactions, nucleotide–nucleotide hydrogen bonds, 57,58 amino acid–amino acid spatial contacts, 59 metabolic pathways steps, 60 protein–protein interactions, 61 RNA–RNA co-expression, 62 gene– gene regulation 63 or any other kind of structural or functional relationships play the role of edges in the CNs. These CNs can be used to characterize overall similarity/dissimilarity relationships 18 as well as properties such as bipartivity, 64 community structure 65 or small world property, 66 in order to unravel complex statistical properties of the phenomena under study. In a previous work, our group used drug–drug CNs to study the overall properties of the activity profiles of several anti-fungals against different species. The CNs were assembled based on the activity predictions made with a TI-QSAR model. 67 There have been reported no CNs for the space of stereo-isomers of potential HMGRIs. In this work, we compiled the biological activity and molecular structure of several well-known HMGRIs. Next, we calculated several TIs and CTIs for these molecules and constructed the first CTI-QSAR models for HMGRIs. Using these models, we predicted the in vitro and in vivo activity profile of all the possible chiral isomers of the studied molecules. With these values we assembled CNs for a large space composed by chiral HMGRIs isomers. We were able to design the CNs for HMGRIs and compared them with the CN resulted when omitted the chirality information. In addition, we proposed a reduced CN that contains the most relevant isomers to be experimentally assayed. The present work may become a useful tool to guide the synthesis of new HMGRI chiral isomers. 2. Results and discussion 2.1. QSAR model In the first step, we performed an LDA in order to obtain a QSAR model that allowed us to classify or not a molecule as a potential HMGRI. The QSAR database is shown in Table SM1 and the correspondent chemical structures and names in Table SM2 (from the Supplementary Materials). The equation of the QSAR model found is the following: HMGRI-score ¼0:169 Cc 0:029 Svd þ0:561 D0:002 W 3:992 Nrc þ12:415 Bn þ12:167 N¼183 Rc ¼0:7k¼0:54 F¼25:15 p-level <0:001 ð1Þ where Nis the number of compounds used for training, Rc is the canonical regression coefficient, kis the Wilks’s statistic parameter, Fis the Fisher ratio and pis the level of error. This model is based on the following parameters presented in Table 1: cluster count (Cc), sum of valence degree (Svd), Wiener index (W), number of R centers (Nrc) and normal chiral/non-chiral balance (Bn). In this work, the high value of Rc = 0.7 (Rc may vary from 0 to 1) indicates a strong correlation and consequently supports the quality of the model. 68 On the other hand, the lower k= 0.54 value found points to a better class or group separation (HMGRIs vs non-HMGRIs). Additionally, the F-test can reject the hypothesis of group overlapping with a level of error (p-level) < 0.001. Thus, we consider that the model performs a good separation of both groups. More direct evidence of the fitting quality are the high values of Sensitivity, Specificity and Accuracy for both training and cv series (see Table 2). 69 All the parameters included in the equation have demonstrated to present statistically a significant contribution to the classification using the Fisher test, see Table 3. The Forward stepwise analysis selects both TIs (Cc, Svd, D, and W) and CTIs (Nrc and Bn), but does not include the In vivo test indicator parameter (Inv). This could indicate that there are no important differences for in vitro versus in vivo results in this database and confirms the use of HMGRI in vitro test results. The results stimulate, in general, the in vivo experiments and, in particular, the search for QSAR models. Cc is related to structure cycles and has a positive contribution of 0.619 to the HMGRI-score. Thus, it could be rationalized, considering that most of these molecules are cycle-containing structures. The diameter of the graph (D) has also a positive contribution of 0.561, possibly due to the fact that drugs need to have certain distance between the separated atoms, in order to guarantee the interaction with the receptor. The negative contribution of 0.002 for Windicates that groups with high level of molecular branching could decrease the HMGRI-score. Considering the relative character of the R-S notation, the general interpretation of the CTI contribution should be avoided and reduced only to specific molecules. In contrast with other QSAR models, with easier structural explanation, 70 we prefer to consider the structural interpretation of the present (TI + CTI)–QSAR models with caution and showing it simply as a prediction instead of an explicative tool. The parametrical assumptions such as the normality, homocedasticity (homogeneity of variances) and non-colinearity have the same importance in the application of multivariate statistic Table 1 Symbols and names of the molecular descriptors used in this work Symbol Name TI Chiral WWiener index Yes No JBalaban index Yes No Cc Cluster count Yes No Svd Sum of valence degree Yes No DDiameter Yes No RRadius Yes No Sha Shape attribute Yes No Shc Shape coefficient Yes No Sod Sum of degree Yes No Mti Molecular topological index Yes No Nrb Number of Rotable Bonds Yes No Tc Total connectivity Yes No Tvc Total valence connectivity Yes No Ntc Number of total centers Yes Yes Ncc Number of chiral centers Yes Yes Nncc Number of non-chiral centers Yes Yes Nsc Number of S centers Yes Yes Nrc Number of R centers Yes Yes Bg = Nncc Ncc General chiral/non-chiral balance Yes Yes Bn = (Ncc Nncc)/Ntc Normal chiral/non-chiral balance Yes Yes Brs = (Nrc Nsc)/(Ncc + 1) Relative RS balance Yes Yes Inv In vivo test indicator No No Psa Polar surface area No No Table 2 QSAR model results Statistics for classes Predicted output Parameter Value (%) Observed input Non-HMGRIs HMGRIs Train Specificity 81.3 Non-HMGRIs 39 9 Sensitivity 89.6 HMGRIs 14 121 Accuracy 87.4 Total 53 130 Cv Specificity 71.4 Non-HMGRIs 10 4 Sensitivity 84.8 HMGRIs 7 39 Accuracy 81.7 Total 17 43 I. García et al. / Bioorg. Med. Chem. 17 (2009) 165–175 167
ncDDD ij ¼X m w m m TI i m TI j ð3aÞ cDDD ij ¼X m w m m TI i m TI j þX n w n ð n CTI i n CTI j Þ ð3bÞ Next, we transformed the distance matrix into a Boolean matrix, assigning a value of 1 to all very similar pairs of drugs and 0 otherwise. The drugs are similar if the DDD ij is lower than a predetermined cut-off value equal to 0.1. We saved the Bollean matrix in a .mat format and used it as input for the Pajek application, (a free-software for CNs analysis). The same software was used to visualize the CN. In the last step, we used the module for analysis of distribution function fit from STATISTICA 6.0 95 in order to investigate the normality of the node degrees (number of similar molecules for each molecule) within the space of chiral derivatives of HMGRIs. Acknowledgements We acknowledge the useful comments and kind attention of the editor Prof. Herbert Waldmann and two unknown reviewers. The authors thank the Portuguese Fundação para a Ciência e a Tecnologia (FCT) (SFRH/BPD/24997/2005). The corresponding author González-Díaz H. acknowledges contract/grant sponsorship for a research position funded by the Program Isidro Parga Pondal of the ‘Xunta de Galicia’. Supplementary data Supplementary data associated with this article can be found, in the online version, at doi:10.1016/j.bmc.2008.11.007. References and notes 1. Shimokata, K.; Yamada, Y.; Kondo, T.; Ichihara, S.; Izawa, H.; Nagata, K.; Murohara, T.; Ohno, M.; Yokota, M. Atherosclerosis 2004,172, 167. 2. Smilde, T. J.; van Wissen, S.; Wollersheim, H.; Kastelein, J. J.; Stalenhoef, A. F. Neth. J. Med. 2001,59, 184. 3. Leitersdorf, E.; Eisenberg, S.; Eliav, O.; Friedlander, Y.; Berkman, N.; Dann, E. J.; Landsberger, D.; Sehayek, E.; Meiner, V.; Wurm, M., et al Circulation 1993,87, III35. 4. Salazar, L. A.; Hirata, M. H.; Quintao, E. C.; Hirata, R. D. J. Clin. Lab. Anal. 2000,14, 125. 5. Suzuki, M.; Iwasaki, H.; Fujikawa, Y.; Sakashita, M.; Kitahara, M.; Sakoda, R. Bioorg. Med. Chem. 2001,9, 2727. 6. Urbina, J. A.; Lazardi, K.; Marchan, E.; Visbal, G.; Aguirre, T.; Piras, M. M.; Piras, R.; Maldonado, R. A.; Payares, G.; de Souza, W. Antimicrob. Agents Chemother. 1993,37, 580. 7. Sit, S. Y.; Parker, R. A.; Motoc, I.; Han, W.; Balasubramanian, N.; Catt, J. D.; Brown, P. J.; Harte, W. E.; Thompson, M. D.; Wright, J. J. J. Med. Chem. 1990,33, 2982. 8. Choulier, L.; Andersson, K.; Hamalainen, M. D.; van Regenmortel, M. H.; Malmqvist, M.; Altschuh, D. Protein Eng. 2002,15, 373. 9. Du, Q. S.; Huang, R. B.; Wei, Y. T.; Du, L. Q.; Chou, K. C. J. Comput. Chem. 2008,29, 211. 10. Li, Y.; Wei, D. Q.; Gao, W. N.; Gao, H.; Liu, B. N.; Huang, C. J.; Xu, W. R.; Liu, D. K.; Chen, H. F.; Chou, K. C. Med. Chem. (Shariqah, United Arab Emirates) 2007,3, 576. 11. Sirois, S.; Tsoukas, C. M.; Chou, K. C.; Wei, D.; Boucher, C.; Hatzakis, G. E. Med. Chem. (Shariqah, United Arab Emirates) 2005,1, 173. 12. Thilagavathi, R.; Kumar, R.; Aparna, V.; Sobhia, M. E.; Gopalakrishnan, B.; Chakraborti, A. K. Bioorg. Med. Chem. Lett. 2005,15, 1027. 13. Prabhakar, Y. S. Drug Des. Discov. 1992,9, 145. 14. Prabhakar, Y. S.; Saxena, A. K.; Doss, M. J. Drug Des. Deliv. 1989,4, 97. 15. González-Díaz, H.; Vilar, S.; Santana, L.; Uriarte, E. Curr. Top. Med. Chem. 2007, 7,1015. 16. Gozalbes, R.; Doucet, J. P.; Derouin, F. Curr. Drug Targets Infect. Disord. 2002,2, 93. 17. Estrada, E.; Uriarte, E. Curr. Med. Chem. 2001,8, 1573. 18. González-Díaz, H.; González-Díaz, Y.; Santana, L.; Ubeira, F. M.; Uriarte, E. Proteomics 2008,8, 750. 19. González-Díaz, H.; Prado-Prado, F. J.; Santana, L.; Uriarte, E. Bioorg. Med. Chem. 2006,14, 5973. 20. Helguera, A. M.; Rodriguez-Borges, J. E.; Garcia-Mera, X.; Fernandez, F.; Cordeiro, M. N. J. Med. Chem. 2007,50, 1537. 21. Marrero-Ponce, Y.; Castillo-Garit, J. A.; Olazabal, E.; Serrano, H. S.; Morales, A.; Castanedo, N.; Ibarra-Velarde, F.; Huesca-Guillen, A.; Sanchez, A. M.; Torrens, F.; Castro, E. A. Bioorg. Med. Chem. 2005,13, 1005. 22. González-Díaz, H.; Olazábal, E.; Santana, L.; Uriarte, E.; Castañedo, N. Bioorg. Med. Chem. 2007,15, 962. 23. Murcia-Soler, M.; Perez-Gimenez, F.; Garcia-March, F. J.; Salabert-Salvador, M. T.; Diaz-Villanueva, W.; Medina-Casamayor, P. J. Mol. Graph. Model. 2003,21, 375. 24. Prado-Prado, F.; González-Díaz, H.; Santana, L.; Uriarte, E. Bioorg. Med. Chem. 2007,15, 897. 25. Mathur, K. C.; Gupta, S.; Khadikar, P. V. Bioorg. Med. Chem. 2003,11, 1915. 26. Calabuig, C.; Anton-Fos, G. M.; Galvez, J.; Garcia-Domenech, R. Int. J. Pharm. 2004,278, 111. 27. González, M. P.; González-Díaz, H.; Molina Ruiz, R.; Cabrera, M. A.; Ramos de Armas, R. J. Chem. Inf. Comput. Sci. 2003,43, 1192. 28. Votano, J. R.; Parham, M.; Hall, L. H.; Kier, L. B.; Oloff, S.; Tropsha, A.; Xie, Q.; Tong, W. Mutagenesis 2004,19, 365. 29. Santana, L.; Uriarte, E.; González-Díaz, H.; Zagotto, G.; Soto-Otero, R.; MendezAlvarez, E. J. Med. Chem. 2006,49, 1149. 30. Marrero-Ponce, Y. J. Chem. Inf. Comput. Sci. 2004,44, 2010. 31. Vilar, S.; Santana, L.; Uriarte, E. J. Med. Chem. 2006,49, 1118. 32. Marrero-Ponce, Y.; Khan, M. T.; Casanola-Martin, G. M.; Ather, A.; Sultankhodzhaev, M. N.; Torrens, F.; Rotondo, R. ChemMedChem 2007,2, 449. 33. Casanola-Martin, G. M.; Marrero-Ponce, Y.; Khan, M. T.; Ather, A.; Sultan, S.; Torrens, F.; Rotondo, R. Bioorg. Med. Chem. 2007,15, 1483. 34. Casanola-Martin, G. M.; Marrero-Ponce, Y.; Khan, M. T.; Ather, A.; Khan, K. M.; Torrens, F.; Rotondo, R. Eur. J. Med. Chem. 2007,42, 1370. 35. González-Díaz, H.; Sanchez, I. H.; Uriarte, E.; Santana, L. Comput. Biol. Chem. 2003,27, 217. 36. Marrero-Ponce, Y.; Castillo-Garit, J. A. J. Comput. Aided Mol. Des. 2005,19, 369. 37. Fabian, W. M.; Stampfer, W.; Mazur, M.; Uray, G. Chirality 2003,15, 271. 38. Golbraikh, A.; Tropsha, A. J. Chem. Inf. Comput. Sci. 2003,43, 144. 39. Golbraikh, A.; Bonchev, D.; Tropsha, A. J. Chem. Inf. Comput. Sci. 2001,41, 147. 40. Castillo-Garit, J. A.; Marrero-Ponce, Y.; Torrens, F. Bioorg. Med. Chem. 2006,14, 2398. 41. de Julián-Ortiz, J. V.; de Gregorio Alapont, C.; Ríos-Santamarina, I.; GarcíaDoménech, R.; Gálvez, J. J. Mol. Graph. Model. 1998,16, 14. 42. Castillo-Garit, J. A.; Marrero-Ponce, Y.; Torrens, F.; Rotondo, R. J. Mol. Graph. Model. 2007,26, 32. 43. Zhang, Z.; Grigorov, M. G. Proteins 2006,62, 470. 44. Bonchev, D.; Buck, G. A. J. Chem. Inf. Model. 2007,47, 909. 45. Ivanciuc, O.; Ivanciuc, T.; Klein, D. J. SAR QSAR Environ. Res. 2001,12,1. Table 5 Behavior of zand Cfor different real versus random CNs CNs a nzC measuredb C randomb Drug–drug activity similarity networks 1 HMGRIs cGC 0.1 479 2 HMGRIs ncGC 3 Anti-fungal 59 22 0.37 0.37 4 Anti-parasites 380 3.34 0.01 0.01 5S. cerevisiae mutants 24 — — — 6 ATC classification 1014 — — — Biological networks 7 Metabolic network 315 28.3 0.59 0.09 8 Neural network 282 14.0 0.28 0.049 9 Food web 134 8.7 0.22 0.065 Social networks 10 Biology collaborations 1 520 251 15.5 0.081 0.00001 11 Mathematics collaborations 253 339 3.9 0.15 0.00002 12 Company directors 7 673 14.4 0.59 0.002 13 Word co-occurrence 460 902 70.1 0.44 0.0002 14 Film-actor collaborations 449 913 113.4 0.2 0.0003 Technological networks 15 WWW (sites) 153 127 35.2 0.11 0.0002 16 Internet 6 374 3.8 0.24 0.0006 17 Power grid 4 941 2.7 0.08 0.0005 a CNs 1 and 2 have been reported here by the first time; CN3—see González-Díaz, H., et al. J. Comput. Chem. 2008,29,656; CN4—Prado-Prado, F. J., et al. Bioorg. Med. Chem. 2008,16(11), 5871. b These are clustering coefficients determined as the ratio C=z/nfor the network measured versus random network with the same number of nodes. 174 I. García et al. / Bioorg. Med. Chem. 17 (2009) 165–175
46. Ivanciuc, O.; Ivanciuc, T.; Klein, D. J.; Seitz, W. A.; Balaban, A. T. J. Chem. Inf. Comput. Sci. 2001,41, 536. 47. Ivanciuc, O. J. Chem. Inf. Comput. Sci. 2000,40, 1412. 48. Gupta, N.; Mangal, N.; Biswas, S. Proteins 2005,59, 196. 49. Gan, H. H.; Pasquali, S.; Schlick, T. Nucleic Acids Res. 2003,31, 2926. 50. Marrero-Ponce, Y.; Nodarse, D.; González-Díaz, H.; Ramos de Armas, R.; Romero-Zaldivar, V.; Torrens, F.; Castro, E. A. Int. J. Mol. Sci. 2004,5, 276. 51. González-Díaz, H.; Agüero-Chapin, G.; Varona, J.; Molina, R.; Delogu, G.; Santana, L.; Uriarte, E.; Gianni, P. J. Comput. Chem. 2007,28, 1049. 52. Rose, A.; Schraegle, S. J.; Stahlberg, E. A.; Meier, I. BMC Evol. Biol. 2005,5, 66. 53. Estrada, E.; Rodriguez-Velazquez, J. A. Phys. Rev. E 2005,71, 056103. 54. Yu, X.; Lin, J.; Masuda, T.; Esumi, N.; Zack, D. J.; Qian, J. Nucleic Acids Res. 2006, 34, 917. 55. Guido, N. J.; Wang, X.; Adalsteinsson, D.; McMillen, D.; Hasty, J.; Cantor, C. R.; Elston, T. C.; Collins, J. J. Nature 2006,439, 856. 56. Williams, R. J.; Berlow, E. L.; Dunne, J. A.; Barabasi, A. L.; Martinez, N. D. Proc. Natl. Acad. Sci. U.S.A. 2002,99, 12913. 57. Marrero-Ponce, Y.; Castillo-Garit, J. A.; Nodarse, D. Bioorg. Med. Chem. 2005,13, 3397. 58. González-Díaz, H.; de Armas, R. R.; Molina, R. Bioinformatics 2003,19, 2079. 59. Vullo, A.; Frasconi, P. J. Bioinform. Comput. Biol. 2003,1, 411. 60. Lange, B. M.; Ghassemian, M. Phytochemistry 2005,66, 413. 61. Chou, K. C.; Cai, Y. D. J. Proteome Res. 2006,5, 316. 62. Yu, X.; Lin, J.; Zack, D. J.; Qian, J. Nucleic Acids Res. 2006,34, 4925. 63. Margolin, A. A.; Nemenman, I.; Basso, K.; Wiggins, C.; Stolovitzky, G.; Dalla Favera, R.; Califano, A. BMC Bioinformatics 2006,7, S7. 64. Estrada, E. J. Proteome Res. 2006,5, 2177. 65. Newman, M. E. Phys. Rev. E: Stat. Nonlin. Soft Matter Phys. 2006,74, 036104. 66. Kleinberg, J. M. Nature 2000,406, 845. 67. González-Díaz, H.; Prado-Prado, F. J. Comput. Chem. 2008,29, 656. 68. Hill, T.; Lewicki, P. STATISTICS Methods and Applications; Tulsa: StatSoft, 2006. 69. Lilien, R. H.; Farid, H.; Donald, B. R. J. Comput. Biol. 2003,10, 925. 70. Vilar, S.; Estrada, E.; Uriarte, E.; Santana, L.; Gutierrez, Y. J. Chem. Inf. Model. 2005,45, 502. 71. Bisquerra Alzina, R. Introducción conceptual al análisis multivariante: Un enfoque informático con los paquetes SPSS-X, BMDP, LISREL y SPAD; Barcelona: PPU, 1989. 72. Stewart, J.; Gill, L. Econometrics; Prentice Hall: London, 1998. 73. Dillon, W. R.; Goldstein, M. Multivariate Analysis: Methods and Applications; Wiley: N.Y., 1984. 74. James, A.; Hanley, B. J. M. Radiology 1982,143, 29. 75. Helguera, A. M.; Rodríguez-Borges, J. E.; García-Mera, X.; Fernández, F.; Cordeiro, M. N. J. med. chem. 2007, 50, 1537. 76. Gómez, G.; Rivera, H.; García, I.; Estévez, L.; Fall, Y. Tetrahedron Lett. 2005,46, 5819. 77. Boccaletti, S.; Latora, V.; Moreno, Y.; Chavez, M.; Hwang, D. U. Phys. Rep. 2006, 424, 175. 78. Bianconi, G.; Barabasi, A. L. Phys. Rev. Lett. 2001,86, 5632. 79. Bornholdt, S.; Schuster, H. G. Handbook of Graphs and Complex Networks; WileyVCH GmbH & Co. KGa: Wheinheim, 2003. 80. Kuroda, M.; Endo, A. Biochim. Biophys. Acta 1977,486,70. 81. Turabi, N.; DiPietro, R. A.; Mantha, S.; Ciosek, C.; Rich, L.; Tu, J. I. Bioorg. Med. Chem. 1995,3, 1479. 82. Watanabe, M.; Koike, H.; Ishiba, T.; Okada, T.; Seo, S.; Hirai, K. Bioorg. Med. Chem. 1997,5, 437. 83. Romo, D.; Harrison, P. H.; Jenkins, S. I.; Riddoch, R. W.; Park, K.; Yang, H. W.; Zhao, C.; Wright, G. D. Bioorg. Med. Chem. 1998,6, 1255. 84. Colle, S.; Taillefumier, C.; Chapleur, Y.; Liebl, R.; Schmidt, A. Bioorg. Med. Chem. 1999,7, 1049. 85. Suzuki, M.; Iwasaki, H.; Fujikawa, Y.; Kitahara, M.; Sakashita, M.; Sakoda, R. Bioorg. Med. Chem. 2001,9, 2727. 86. Taillefumiera, C.; Fornelb, D.; Chapleur, Y. Bioorg. Med. Chem. Lett. 1996,6, 615. 87. Alali, F.; Zeng, L.; Zhang, Y.; Ye, Q.; Hopp, D. C.; Schwedler, J. T.; McLaughlin, J. L. Bioorg. Med. Chem. 1997,5, 549. 88. Schroepfer, G. J., Jr.; Parish, E. J.; Chen, H. W.; Kandutsch, A. A. J. Biol. Chem. 1977,252, 8975. 89. Beck, G.; Kesseler, K.; Baader, E.; Bartmann, W.; Bergmann, A.; Granzer, E.; Jendralla, H.; von Kerekjarto, B.; Krause, R.; Paulus, E., et al J. Med. Chem. 1990, 33, 52. 90. Zhang, Q. Y.; Wan, J.; Xu, X.; Yang, G. F.; Ren, Y. L.; Liu, J. J.; Wang, H.; Guo, Y. J. Comb. Chem. 2007,9, 131. 91. Grabley, S.; Granzer, E.; Hutter, K.; Ludwig, D.; Mayer, M.; Thiericke, R.; Till, G.; Wink, J.; Philipps, S.; Zeeck, A. J. Antibiot. (Tokyo) 1992,45, 56. 92. Gohrt, A.; Zeeck, A.; Hutter, K.; Kirsch, R.; Kluge, H.; Thiericke, R. J. Antibiot. (Tokyo) 1992,45, 66. 93. Grabley, S.; Hammann, P.; Hutter, K.; Kirsch, R.; Kluge, H.; Thiericke, R.; Mayer, M.; Zeeck, A. J. Antibiot. (Tokyo) 1992,45, 1176. 94. Chan, C.; Bailey, E. J.; Hartley, C. D.; Hayman, D. F.; Hutson, J. L.; Inglis, G. G. A.; Jones, P. S.; Keeling, S. E.; Kirk, B. E.; Lamont, R. B.; Lester, M. G.; Pritchard, J. M.; ROSS, B. C.; Scicinski, J. J.; Spooner, S. J.; Smith, G.; Steeples, I. P.; Watson, N. S. J. Med. Chem. 1993,36, 3646. 95. StatSoft, Inc., 2002. 96. Microsoft Corp., Microsoft Excel, 2002. 97. Zhang, W. Environ. Monit. Assess. 2007,124,253. I. García et al. / Bioorg. Med. Chem. 17 (2009) 165–175 175
Mol Divers DOI Theoretical study of GSK-3Į: Neural Networks QSAR studies for the design of new inhibitors using 2D-descriptors Isela García *, Yagamare Fall, Xerardo García-Mera Francisco Prado-Prado Abstract GSK-3 targets encompass proteins implicated in AD, neurological disorders. The functions of GSK-3 and its implication in various human diseases have triggered an active search for potent and selective GSK-3 inhibitors. In this sense, QSAR could play an important role in studying these GSK-3 inhibitors. For this reason we developed QSAR models for GSK3Į, LDA and ANNs from nearly 50000 cases with more than 700 different GSK-3Į inhibitors obtained from ChEMBL database server; in total we used more than 20000 different molecules to develop the QSAR models. The model correctly classified 237 out of 275 active compounds (86.2%) and 14870 out of 15970 non-active compounds (93.2%) in the training series. The overall training performance was 93.0%. Validation of the model was carried out using an external predicting series. In these series the model classified correctly 458 out of 549 (83.4%) compounds and 29637 out of 31927 non-active compounds (83.4%). The overall predictability performance was 92.7%. In this work, we propose three types of non Linear ANN and we show that it is another alternative model to the already existing ones in the literature, such as LDA. The best model obtained was Linear Neural Network: LNN: 236:236-1-1:1 which had an overall training performance of 96%. In addition, we did a study of different fragments that exist in the molecules of the database in order to see which fragments had more influence in the activity. All this can help to design new inhibitors of GSK-3Į. The present work reports the attempts to calculate within a unified framework probabilities of GSK-3Į inhibitors against different molecules found in the literature. Keywords: GSK-3Į, QSAR, Artificial Neural Network, Linear Neural Network, Fragment contribution, Linear Discriminant Analysis. Introduction Alzheimer´s disease (AD) [1] is a serious and degenerative disorder that causes a gradual loss of neurons, and in spite of the efforts realized by the big pharmaceutical companies of the world, the origin of this pathology is still not very clear. Glycogen synthase kinase-3 (GSK-3) is a serine-threonine kinase encoded by two isoforms in mammals, termed GSK-3Į and GSK-3ȕ [2]. GSK-3 targets encompass proteins implicated in AD, neurological I. García Y. Fall G. Gómez Department of Organic Chemistry, University of Vigo, Vigo, Spain e-mail:
[email protected] X. García-Mera F. Prado-Prado Department of Organic Chemistry, Faculty of Pharmacy, USC, 15782 Santiago de Compostela, Spain disorders, in the wnt and insulin signaling pathway, glycogen and protein synthesis, regulation of transcription factors [3], embryonic development, cell proliferation and adhesion, tumorigenesis, apoptosis [4] FLUFDGLDQ UK\WKP« HWF *6.-3ȕ knock-out mice die in utero [5], whereas GSK-3Į knock-out mice are viable and display improved glucose tolerance in response to glucose load and elevated hepatic glycogen storage and insulin sensitivity [6, 7]. The functions of GSK-3 and its implication in various human diseases have triggered an active search for potent and selective GSK-3 inhibitors [8] in the last years. In this sense, quantitative structure-activity relationships (QSAR) could play an important role in studying these GSK-3 inhibitors (GSKI-3Į); QSARs can be used as predictive tools for the development of molecules [9, 10]. Computer-aided drug design techniques based on QSAR could play an important role in drug discovery programs. The QSAR approach involves the development of models that relate the structure of drugs with to their biological activity against different targets [11, 12]. In principle, there are currently more than 1600 molecular descriptors that may be generalized and used to solve the problem outlined above [13, 14]. Many of these indices are known as Dragon Descriptors (DDs). DRAGON provides 1664 molecular descriptors divided into several families: 0D (constitutional descriptors), 1D (e.g., functional group counts), 2D (e.g., topological descriptors and connectivity indices), and 3D (e.g., GETAWAY, WHIM, RDF and 3D-MoRSE descriptors) [15-17]. Numerous different molecular descriptors have been reported to encode chemical structures in QSAR studies. Furthermore, there are multiple chemometric approaches that can, in principle, be selected for this step. Multiple linear regression (MLR), linear discriminant analysis (LDA) [18], partial least squares (PLS) and different kinds of artificial neural networks (ANN) can be used to relate molecular structure (represented by molecular descriptors) with to biological properties. The ANNs are particularly useful in QSAR studies in which the linear models fit poorly due to high data complexity [15, 18], an example was the work of Prado-Prado et. al. in which four types of non Linear Artificial neural networks (ANN) were developed to calculate within an unified framework probabilities of antiparasitic action of drugs against different parasite species [19-21]. There are several different kinds of ANN and these include multilayer perceptron (MLP), radial basis functions (RBF) and PNNs; the latter ANN is a variant of RBF systems. In particular, PNN is a type of neural network that uses a kernel-based approximation to form an estimate of the probability density functions of classes in a classification problem [22]. In this article, we developed QSAR models for GSK-3Į, LDA and ANNs from more than 48000. This database contains 700
different molecules inhibitors of GSK-3Į (active against GSK3Į); in total we used more than 20000 different molecules nonactive or not have interaction with GSK-3Į to develop a different QSAR models. All the tested compounds were compiled from ChEMBL database http://www.ebi.ac.uk/chembldb/index.php/target/browser/classific ation [23, 24]. First we used 237 molecular descriptors calculated with DRAGON software [16]. Next, we developed LDA and ANNs models in order to find new compounds that present inhibitor action against GSK-3Į and hence might be used as new treatments for neurological pathologies such as Parkinson´s disease (PD) or AD. Last, to probe if our QSAR models are consistent, we established a Molecular Fragments study; we used different fragments from the known molecules of the database in order to see which fragments had more influence in the activity, and which fragments interacted more with the GSK-3Į protein.The design of new inhibitors of this enzyme is very important for the study of neurodegenerative diseases [25, 26]. The present is the first work reports the attempts to calculate within an unified framework probabilities of GSK-3Į inhibitors against different molecules using a tool server ChEMBL database. Methods Linear classifier A database from ChEMBL database [23] containing assayed GSK-3Į inhibitors was used (Table SM from the Supplementary Material). The DRAGON software 4.0 [16] was utilized here and provides 1664 descriptors classified as zero- (0D) one- (1D), two- (2D) ane three-dimensional (3D) descriptors depending on whether they are computed from the chemical formula, substructure list representation, molecular graph or geometrical representation of the molecule, respectively [27]. In this work, we calculated the following descriptors: 2D autocorrelations, Burden eigenvalues, topological charge indices, eigenvalue-based indices, functional group counts, atoms-centred fragments, charge descriptors and molecular properties. The QSAR model was constructed with the multivariate regression technique, the LDA, employing the Forward stepwise method for the selection of variables. All statistical analyses and data exploration were carried out in STATISTICA 6.0 [28]. In the actual work, the independent data test is used by splitting the data randomly in a training series used for a model construction and a cross-validation (CV) one. The general formula of the QSAR classification function is the following: 123 0 ¦ WDWGSKI i m mscore D where GSKI-3Įscore is the continuous and dimensionless score value for the GSKI-3Į/non-GSKI3Į classification that gives relatively higher values to molecules with more probability to act as GSKI-3Į, m2Di are the 2Ds of type m, Wm is the coefficient (weights) of these indices in the QSAR model and W0 is the independent term. The reported statistical parameters of the QSAR model DUH WKH IROORZLQJ 1 Ȥ2 and p-level as well as Sensitivity, Specificity, and Accuracy for both training and CV [28]. N is the number of molecules used to WUDLQWKHPRGHOȜLV:LONVVWDWLVWLFSDUDPHWHULV&KLsquare and p-level is the probability of error. Nonlinear classifier We processed our data with different ANNs using the STATISTICA 6.0 software [28] looking for a better model to predict activity against GSK-3Į. Three types of ANNs were used, namely, Radial Basis Function (RBF) [29], Multi Layers Perceptron (MLP) and Linear Neural Network (LNN). The profile of a ANN is: Ni:I-H1-H2-O:No. It means that we have inputs variables (Ni), neurons in the input layer (I), neurons in the first hidden layer (H1), in the second hidden layer (H2), neuron in the output layer (O) and output variable (No). We can used a very simple type of ANN called Linear Neural Network (LNN) to fit this discriminant function. The model deals with the classification of a compound set with or without affinity on different receptors. A dummy variable Affinity Class (AC) was used as input to codify the affinity. This variable indicates either high (AC = 1) or low (AC = 0) affinity of the drug by the receptor. S(DTP)pred or DTP affinity predicted score is the output of the model and it is a continuous dimensionless score that sorts compounds from low to high affinity to the target coinciding DTPs with higher values of S(DTP)pred and nDTPs with lowest values. In Equation 2, b represents the coefficients of the LNN classification function, determined by the ANN module of the STATISTICA 6.0 software package [28]. We used Forward Stepwise algorithm for a variable selection. Let be kȤ* GUXJV PROHFXODU GHVFULSWRUV DQG kȟR) receptor or drug target descriptors for different drugs (d) with different receptor; we can attempt to develop a simple linear classifier of mt-QSAR type with the general formula:
2 5 0 5 0 bRRbGGbDTPS k k k k k k pred ¦¦ [F We assessed the quality of models with different statistical parameters like Specificity (see Equation 2), Sensitivity (see Equation 3), Accuracy (see Equation 4) and ROC curve (Receiver Operating Characteristic curve) which is a graphical plot of the sensitivity RU WUXH SRVLWLYHV YV íspecificity), or false positives, 3 NFPNTN NTN yspecificit 4 NFNNTP NTP ysensitivit 5 TNFPFNNTP NTNNTP accuracy where NTN means number of true negatives, NFP is number of false positives, NTP is number of true positives, NFN is number of false negatives, FN is false negatives, FP is false positives and TN is true negatives. Study of molecular fragments In this work we calculated contributions of different molecular fragments for activity against 15 different molecules inhibitors of GSK-3Į. In so doing we gave the following steps [11, 21]: xFirst, we calculated the specie-dependent atomic descriptors included in the QSAR equation for selected molecular fragments using DRAGON software [16]. xSecond, we calculated the contribution scores of each fragment against 15 species of molecules studied by substituting the atomic descriptors into the QSAR equation using the Microsoft Excell application. xThird, the contributions of each molecular fragment were standardized dividing each value by the sum of all contributions for each molecule. These molecular fragment contributions can indicate the potential relation between molecular fragments with the activity against GSK-3Į, each separated fragment or each fragment inside a molecule. xFourth, contributions for each atom of the drug were scaled into a percentage value. xFifth, scaled atom contributions were grouped into different molecular fragments. xSixth, molecular fragment contributions to the biological activity were back-projected onto the molecular structure for obtaining a colourscaled biological structure-activity. Data set We developed QSAR models for GSK-3Į, LDA and ANNs from more than 48000. This database contains 700 different molecules inhibitors of GSK-3Į (active against GSK-3Į); in total we used more than 20000 different molecules non-active or not have interaction with GSK-3Į to develop a different QSAR models. All the compounds were compiled from ChEMBL database http://www.ebi.ac.uk/chembldb/index.php/target/brow ser/classification [23, 24]. This is a database of bioactive drug-like small molecules, which contains 2D structures, calculated properties (e.g. logP, Molecular Weight, Lipinski Parameters, etc.) and abstracted bioactivities (e.g. binding constants, pharmacology and ADMET data). ChEMBL normalises the bioactivities into a uniform set of endpoints and units where possible, and also tags the links between a molecular target and a published assay with a set of varying confidence levels. The data is abstracted and curated from the primary scientific literature, and covers a significant fraction of the structure activity relationship (SAR) and discovery of modern drugs. The codes and activity for all compounds as well as the references used to collect them are depicted in Table SM of the supplementary material file. Results and Discussion LDA In this paper we obtained a LDA study, Equation 6, and we can observe that eighteen variables entry inside equation: 001.03.3127721,48 )6(7.4024.16342.22731.140.36 30.4188.646.3435.4620.457.333.43 33.731.3910.3769.243.1433.1913.343 2 levelpN VEmJGIGGIBELp BELpBELeBELmBELmeGATSmGATSpMATS eMATSvMATSpATSeATSeATSmATSmATSGSKI score F D The nomenclature used in the descriptors of the equation is the same as establishing the Dragon software, ZKHUH 1 LV WKH QXPEHU RI FDVHV Ȥ2 is the Chi-square and p is the level of error. The model correctly classified 237 out of 275 active compounds (86.2%) and 14870 out of 15970 non-active
compounds (93.2%) in the training series. The overall training performance was 93.0%. Validation of the model was carried out using an external predicting series. In this series the model classified correctly 458 out of 549 (83.4%) antiparasitic compounds and 29637 out of 31927 non-active compounds (92.8%). The overall predictability performance was 92.7% (see Table 1). Table 1 comes about here ANN models The ANN models are non-linear models useful to predict the biological activity of a large datasets of molecules. This technique is an alternative to linear methods such as LDA [30, 31]. Figure 1 depicts the networks maps for some of the ANN models. In general, at least one ANN of every types tested was statically significant. However, one must note that the profiles of each network indicate that these are highly nonlinear and complicated models [32-34]. Figure 1 comes about here In Figure 2, we depict the ROC-curve [35, 36] for LNN tested. Notably, the ROC curve can also be represented equivalently by plotting the fraction of true positives (sensitivity) vs. the fraction of false positives (1-VSHFL¿FLW\ 7KH YLWDOLW\ RI WKLV W\SH RI SURFHGXUHV developing ANN-QSAR models has been demonstrated before [37]; see, for instance, the work of Fernandez and Caballero [38]. The same is true about the ANNs tested in this work, we illustrated a ROC-curve for the best ANN model, in this case was a LNN (236:236-1:1) which results of ROC curve values (AUC) were with an area higher than 0.98. It indicates WKDW WKH SUHVHQW PRGHO JLYHV VWDWLVWLFDOO\ VLJQL¿FDQW results and clearly different from those obtained with a UDQGRP FODVVL¿HU DUHD [39]. To shows how important is this result, we compared the present model with other model used to address the same problem. We processed our data with ANNs looking for a better model [31]. The network found was LNN and it showed training performance higher than 96%. The summary of results is showed in Table 1. After direct inspection of the results reported in Table 1 for ANN methods, we can conclude that a complex ANN method is a good method to predict the activity. We compare different types of networks to obtain a better model; Table 1 shows the classification matrix of the different networks. LNN 236:236-1:1 was taken as the main network because it presented a wider range of variables, 236 inputs in the first layer and 236 neurons in second layer, and two sets of cases (Training and Validation). Another tested networks found were MLP 22:22-27-1:1, RBF 98:98-740-1:1 presented low accuracy and PNN 237:237-16760-2-2:1 had a very low percentage of DTPs leading to possible errors in the model although its accuracy was very good, see Table 1. We depict the ROC-curve for LNN 236:2361:1to show how reliable was the network model developed, see Figure 2. Figure 2 comes about here Study of molecular fragments (F) One application of QSAR is the calculation of the contribution of different molecular fragments to the desired activity [40, 41]. In this sense, one important application of QSAR models is the calculation of molecular fragments contribution to activities or action against different drug targets [11]. Before, to get the different QSAR models we calculated the contribution of different molecular fragments from the better model obtained in this work. The LDA model was better than ANN models obtained, because the LDA model presented 18 variables to obtained one result (active or non-active GSK-3Į), while the best LNN model was presented 236 variables to obtained the same results. In spite of percentage of good classification the LNN model presents the highest classified percentage, but LNN model needs more variables to predict one result than LDA model. This LDA model developed in this work is a single equation that uses few parameters to predict the inhibitory action of a new compound against GSK-3Į. For this reason, to obtain the best and the most reliable results, we calculated fragments of the molecules using the LDA method. As a result, we selected different molecular fragments against 15 molecules of database; we selected these molecules at random. In Figure 3, we observed contribution of all fragments for the 15 molecules selected at random whose results in our model are as follow: green colour indicates major contribution to the activity, yellow colour indicates medium contribution and red colour indicates minor contribution to the biological activity against GSK-3Į. An example is the Dasatinib see Louise N. Johnson [42], where a pair of hydrogen bonds is formed in the hinge region of the ATPbinding site (i.e. between the 3-nitrogen of the aminothiazole ring of dasatinib and the amide nitrogen
of Met318 and between the 2-amino hydrogen of dasatinib with the carbonyl oxygen of Met318). A hydrogen bond is also formed between the side chain hydroxyl oxygen of Thr315 and the amide nitrogen of dasatinib. Figure 3 comes about here Conclusion The functions of GSK-3 and its implication in various human diseases have triggered an active search for potent and selective GSK-3 inhibitors. Nowadays theoretical studies such as QSAR models have become a very useful tool in this context to substantially reduce time and resources consuming experiments. In this work we developed a new LDA model using the Dragon descriptors, using a large data base using about 20000 different drugs obtained from the ChemBL server. We conclude that a large database gives a much more precise model; the use of tools such as ChembL database enables us to develop models with large data bases, and this helps us make the results more reliable. To improve the model we developed non-linear models and compared them to LDA. We proposed non-linear models, and for the first time, we proposed ANN models based on Dragon Descriptors series of GSK-3Į, and we concluded that they are alternative methods to study the activity of different families of molecules compared with other methods found in the literature. The use of tools such as fragment contributions can help us design the best GSK-3Į inhibitors for their later synthesis in the laboratory, eliminating the synthesis of molecules at random which may have few possibilities of being active. Acknowlegments Prado-Prado F. thanks sponsorships for research position at the University of Santiago de Compostela from Angeles Alvariño, Xunta de Galicia. All authors acknowledge the Project 07CSA008203PR. We are grateful to the Xunta de Galicia (INCITE08PXIB314255PR) for partial financial support. References [1] Hübscher U, Maga G, Spadari S. Eukaryotic DNA polymerases. Annu Rev Biochem. 2002;71:133-63. [2] Benek K, Kunkel TA. Functions of DNA polymerases. Adv Protein Chem. 2004;69:137-65. [3] Kornberg A, Baker TA. DNA replication. New York 1992. [4] Takahashi S, Yonezawa Y, Kubota K, Ogawa N, Maeda K, Koshino H, et al. Pyranicin, a non-classical annonaceous acetogenin, is a potent inhibitor of DNA polymerase, topoisomerase and human cancer cell growth. International Jounal of Oncology. 2008;32:451-8. [5] Loeb LA, Agarwal SS, Dube DK, Gopinathan KP, Travaglini EC, Seal G, et al. Inhibitors of mammalian DNA polymerases: possible chemotherapeutic approaches. Pharmacol Ther, Part A:. 1977;2:171-93. [6] Miura S, Izuta S. DNA polymerases as targets of anticancer nucleotides. Curr Drug Targets. 2004;5:191-5. [7] Mizushina Y, Kasai N, Iijima H, Sugawara F, Yoshida H, Sakaguchi TK. Sulfo-quinovosyl-acyl-glycerol (SQAG), a eukaryoticDNA polymerase inhibitor and anticancer agent. Curr Med Chem: Anti-Cancer Agents. 2005;5:613-25. [8] $OEHUWHOOD 05 /DX $ 2¶&RQQRU 0- 7KH overexpression of specialized DNA polymerases in cancer. DNA Repair. 2005;4:583-92. [9] Nakamura R, Takeuchi R, Kuramochi K, Mizushina Y, Ishimaru C, Takakusagi Y, et al. Chemical properties of fatty acid derivatives as inhibitors of DNA polymerases. Organic & Biomolecular Chemistry. 2007;5:3912-21. [10] So AG, Downey KM. Eukaryotic DNA replication. Crit Rev Biochem Mol Biol. 1992; 27:129-55. [11] Mizushina Y, Saito A, Tanaka A, Nakajima N, Kuriyama I, Takemura M, et al. Structural analysis of catechin derivatives as mammalian DNA polymerase inhibitors. Biochem Biophys Res Commun. 2005;333:101-9. [12] Colot V, Rossignol JL. Eukaryotic DNA methylation as an evolutionary device. BioEssays.21:402-11. [13] Jean B, Margot M, Cristina C, Heinrich L. Mammalian DNA Methyltransferases Show Different Subnuclear Distributions. J Cell Biochem. 2001;83:373-9. [14] Sun NJ, Ho Woo S, Cassady JM, Snapka RM. DNA Polymerase and Topoisomerase II Inhibitors from Psoralea corylifolia. J Nat Prod. 1998;61:362-6. [15] Nunez MB, Maguna FP, Okulik NB, Castro EA. QSAR modeling of the MAO inhibitory activity of xanthones derivatives. Bioorg Med Chem Lett. 2004 Nov 15;14(22):5611-7. [16] Todeschini R, Consonni V. Handbook of Molecular Descriptors. 2000. [17] Freund JA, Poschel T. Stochastic processes in physics, chemistry, and biology. Lect Notes Phys. Berlin, Germany: Springer-Verlag 2000. [18] Estrada E, Uriarte E. Recent advances on the role of topological indices in drug discovery research. Curr Med Chem. 2001;8:1573-88. [19] Estrada E, Uriarte E, Montero A, Teijeira M, Santana L, De Clercq E. A Novel Approach for the Virtual Screening and Rational Design of Anticancer Compounds. J Med Chem. 2001;43:1975-85. [20] Verma RP, Kurup A, Hansch C. On the role of polarizability in QSAR. Bioorg Med Chem. 2005 Jan 3;13(1):237-55. [21] WU RS, Wolpert-DeFilippes MK, Quinn FR. Quantitative Structure-Activity Correlations of Rifamycins as Inhibitors of Viral RNA-Directed DNA Polymerase and Mammalian a and p DNA Polymerases. J Med Chem. 1980;23(3):256-61. [22] Quinn FR, Driscoll JS, Hansch C. Structure-Activity Correlations among Rifamycin B Amides and Hydrazides. J Med Chem. 1975;18(4):332-9.
[23] Wright GE, Gambino JJ. Quantitative Structure-Activity Relationships of 6-Anilinouracils as Inhibitors of Bacillus subtilis DNA Polymerase I11. J Med Chem. 1984;27:181-5. [24] Gass KB, Cozzarelli NR. Further Genetic and Enzymological Characterization of the Three Bacillus subtilis Deoxyribonucleic Acid Polymerases. J Biol Chem. 1973;248:7688-700. [25] Clements J, DAmbrosio J, Brown NC. Inhibition of Bacillus subtilis deoxyribonucleic acid polymerase III by phenylhydrazinopyrimidines. Demonstration of a drug-induced deoxyribonucleic acid-enzyme complex. J Biol Chem. 1975;250:522-6. [26] Chakraborty AK, Majumder HK. Mode of action of pentavalent antimonials: specific inhibition of type I DNA topoisomerase of Leishmania donovani. Biochem Biophys Res Commun. 1988;152:605-11. [27] Liu LF. DNA topoisomerase poisons as antitumor drugs. Annu Rev Biomed. 1989;58:351-75. [28] Ray S, Hazra B, Mittra B, Das A, Majumder HK. Diospyrin, a bisnaphthoquinone: a novel inhibitor of type I DNA topoisomerase of Leishmania donovani. Mol Pharmacol. 1998;54:994-9. [29] Sakaguchi K, Sugawara F, Mizushina Y. Inhibitors of eukaryotic DNA polymerase. Seikagaku. 2002;74:244-51. [30] Talete srl, ed. DRAGON for Windows (Software for Molecular Descriptor Calculations). Version 5.3 ed 2005. [31] StatSoft.Inc. STATISTICA (data analysis software system), version 6.0, www.statsoft.com.Statsoft, Inc. 6.0 ed 2002. Received: Revised: Accepted:
Figure 1 Figure 2 Figure 3 Table 1 Model Train Stat. Validation profile Active Non-Active % Par. % Active Non-Active 237 38 86.2 Sn 83.4 458 91 LDA 1100 14870 93.2 Sp 92.8 2290 29637 93.0 Ac 92.7 LNN 258 17 93.8 Sn 93.3 513 37 236:236-1:1 860 15625 94.8 Sp 94.4 1834 31146 96.2 Ac 95.6 MLP 97 178 35.3 Sn 34.6 190 360 22:22-27-1:1 11038 5447 33.0 Sp 33.3 22004 10976 33.1 Ac 33.3 RBF 210 65 76.4 Sn 72.4 398 152 98:98-740-1:1 4713 11772 71.4 Sp 71.4 9441 23539 71.5 Ac 71.4
Mol Divers Fig. 1 Domain of applicability greater than h∗. For the training set this means the chemical is highly influential in determining the model, while for the test set it means the prediction is the result of substantial extrapolation of the model and as such could unreliable [34,35]. To examine this in further detail, a double ordinate Cartesian plot of deleted-residuals (first ordinate), standard residuals (second ordinate), and leverages (abscissa) defined the domain of applicability of the model as a squared area within ±2 band for residuals and a leverage threshold of h=0.0044. As can be noted in Fig. 1, almost all cases used in training and validation lie within this area. Some sequences have leverage higher than the threshold, but show leave-one out (LOO) residuals, deleted-residuals, and standard residuals within the limits. In closing, no apparent outliers were detected and the model can be used with high accuracy in this applicability domain [36–38]. Conclusions In this work, we have shown that the MARCH-INSIDE methodology can be considered as a good alternative for developing GSK-3 inhibitors in a fast and efficient way. This approach is able to correctly classify the GSK-3 inhibitory activity of compounds with different structural patterns. Acknowledgments We are grateful to the Xunta de Galicia (INCITE08PXIB314255PR) for partial financial support. H. GonzálezDíaz acknowledges partial financial support from Program Isidro Parga Pondal, Xunta de Galicia. References 1. Olson RE (2000) Secretase inhibitors as therapeutics for Alzheimer’s disease. Ann Rep Med Chem 35:31–40. doi:10.1016/ S0065-7743(00)35005-9 2. Woodgett JR (1990) Molecular cloning and expression of glycogen synthase kinase-3/factor A. EMBO J 9:2431–2438 3. Woodgett JR (1991) cDNA cloning and properties of glycogen synthase kinase-3 methods. Enzymol 200:564–577. doi:10.1016/ 0076-6879(91)00172-S 4. Ali A, Hoeflich KP, Woodgett JR (2001) Glycogen synthase kinase-3: properties, functions, and regulation. Chem Rev 101: 2527–2540. doi:10.1021/cr000110o 5. Ishiguro K, Ihara Y, Uchida T, Imahori K (1988) A novel tubulindependent protein kinase forming a paired helical filament epitope on tau. J BioChem 104(3):319–321 6. Fairlamb AH (2003) Chemotherapy on human African trypanosomiasis: current and future prospects. Trends Parasitol 19:488–494. doi:10.1016/j.pt.2003.09.002 7. Plyte SE, Hughes K, Nilkolakaki E, Pulverer BJ, Woodgett JR (1992) Glycogen synthase kinase-3: functions in oncogenesis and development. Biochim Biophys Acta 1114:147–162. doi:10. 1016/0304-419X(92)90012-N 8. Ojo KK, Gillespie RG, Riechers A, Napuli AJ, Verlinde CL, Buckner FS et al (2008) Glycogen synthase kinase 3 is a potential drug target for african trypanosomiasis therapy. Antimicrob Agents Chemother 52: 3710–3717 9. Freund JA, Poschel T (2000) Stochastic processes in physics, chemistry, and biology (lecture notes in physics). Springer-Verlag, Berlin 10. Estrada E, Uriarte E (2001) Recent advances on the role of topological indices in drug discovery research. Curr Med Chem 8:1573– 1588. doi:10.2174/0929867013371923 11. Estrada E, Uriarte E, Montero A, Teijeira M, Santana L, De Clercq E (2001) A novel approach for the virtual screening and rational design of anticancer compounds. J Med Chem 43:1975–1985 12. Prado-Prado FJ, Borges F, Perez-Montoto LG, Gonzalez-Diaz H (2009) Multi-target spectral moment: QSAR for antifungal 123
Mol Divers drugs vs. different fungi species. Eur J Med Chem 44:4051–4056. doi:10.1016/j.ejmech.2009.04.040 13. González-Díaz H, Torres-Gomez LA, Guevara Y, Almeida MS, Molina R, Castanedo N et al (2005) Markovian chemicals “in silico” design (MARCH-INSIDE), a promising approach for computer-aided molecular design III: 2.5D indices for the discovery of antibacterials. J Mol Model 11:116–123. doi:10.1007/ s00894-004-0228-3 14. Gonzalez-Díaz H, Prado-Prado F, Ubeira FM (2008) Predicting antimicrobial drugs and targets with the MARCH-INSIDE approach. Curr Top Med Chem 8:1676–1690. doi:10.2174/ 156802608786786543 15. Santana L, Uriarte E, González-Díaz H, Zagotto G, Soto-Otero R, Mendez-Alvarez E (2006) A QSAR model for in silico screening of MAO-A inhibitors. Prediction, synthesis, and biological assay of novel coumarins. J Med Chem 49:1149–1156. doi:10.1021/ jm0509849 16. Concu R, Dea-Ayuela MA, Perez-Montoto LG, Bolas-Fernandez F, Prado-Prado FJ, Podda G et al (2009) Prediction of enzyme classes from 3D structure: a general model and examples of experimental-theoretic scoring of peptide mass fingerprints of Leishmania proteins. J Proteome Res 8:4372–4382. doi:10.1021/pr9003163 17. Kutner MH, Nachtsheim CJ, Neter J, Li W (2005) Standardized multiple regression model. Applied linear statistical models, 5th edn. McGraw Hill, New York, pp. 271–277 18. Hall Ca (1996) The Merck Index, 12th ed. Merck & Co, New Jersey 19. Van Waterbeemd H (1995) Discriminant analysis for activity prediction. In: Van Waterbeemd H (ed) Chemometric methods in molecular design. Wiley-VCH, New York pp 265–282 20. Konda VR, Desai A, Darland G, Bland JS, Tripp ML (2009) Rho iso-alpha acids from hops inhibit the GSK-3/NF-kappaB pathway and reduce inflammatory markers associated with bone and cartilage degradation. J inflamm (Lond) 6:26–34. doi:10.1186/ 1476-9255-6-26 21. Jacquemard U, Dias N, Lansiaux A, Bailly C, Loge C, Robert JM et al (2008) Synthesis of 3,5-bis(2-indolyl)pyridine and 3-[(2indolyl)-5-phenyl]pyridine derivatives as CDK inhibitors and cytotoxic agents. Bioorg Med Chem 16:4932–4953 22. Olesen PH, Sorensen AR, Urso B, Kurtzhals P, Bowler AN, Ehrbar U et al (2003) Synthesis and in vitro characterization of 1-(4-aminofurazan-3-yl)-5-dialkylaminomethyl-1H-[1,2,3]triazole-4-carboxyl ic acid derivatives. A new class of selective GSK-3 inhibitors. J Med Chem 46:3333–3341. doi:10.1021/jm021095d 23. Calabuig C, Anton-Fos GM, Galvez J, Garcia-Domenech R (2004) New hypoglycaemic agents selected by molecular topology. Int J Pharm 278:111–118. doi:10.1016/j.ijpharm.2004.03.012 24. Cercos-del-Pozo RA, Perez-Gimenez F, Salabert-Salvador MT, Garcia-March FJ (2000) Discrimination and molecular design of new theoretical hypolipaemic agents using the molecular connectivity functions. J Chem Inf Comput Sci 40:178–184. doi:10.1021/ ci9900480 25. Murcia-Soler M, Perez-Gimenez F, Garcia-March FJ, Salabert-Salvador MT, Diaz-Villanueva W, Medina-Casamayor P (2003) Discrimination and selection of new potential antibacterial compounds using simple topological descriptors. J Mol Graph Model 21:375–390. doi:10.1016/S1093-3263(02)00184-5 26. Estrada E, Vilar S, Uriarte E, Gutierrez Y (2002) In silico studies toward the discovery of new anti-HIV nucleoside compounds with the use of TOPS-MODE and 2D/3D connectivity indices. 1. Pyrimidyl derivatives. J Chem Inf Comput Sci 42:1194–1203. doi:10. 1021/ci0255331 27. Cronin MT, Aptula AO, Dearden JC, Duffy JC, Netzeva TI, Patel H et al (2002) Structure-based classification of antibacterial activity. J Chem Inf Comput Sci 42:869–878. doi:10.1021/ci025501d 28. Prado-Prado FJ, Ubeira FM, Borges F, Gonzalez-Diaz H (2010) Unified QSAR & network-based computational chemistry approach to antimicrobials. II. Multiple distance and triadic census analysis of antiparasitic drugs complex networks. J Comput Chem 31:164–173. doi:10.1002/jcc.21292 29. Prado-Prado FJ, Martinez de la Vega O, Uriarte E, Ubeira FM, Chou KC, Gonzalez-Diaz H (2009) Unified QSAR approach to antimicrobials. 4. Multi-target QSAR modeling and comparative multi-distance study of the giant components of antiviral drug-drug complex networks. Bioorg Med Chem 17:569–575. doi:10.1016/j. bmc.2008.11.075 30. Oberg T (2004) A QSAR for baseline toxicity: validation, domain of application, and prediction. Chem Res Toxicol 17:1630–1637. doi:10.1021/tx0498253 31. Gramatica P (2007) Principles of QSAR models validation: internal and external. QSAR Comb Sci 26:694–701. doi:10.1002/qsar. 200610151 32. Eriksson L, Jaworska J, Worth AP, Cronin MT, McDowell RM, Gramatica P (2003) Methods for reliability and uncertainty assessment and for applicability evaluations of classificationand regression-based QSARs. Environ Health Perspect 111:1361–1375. doi:10.1289/ehp.5758 33. Tropsha A, Gramatica P, Gombar VK (2003) The importance of being earnest: validation is the absolute essential for successful application and interpretation of QSPR models. QSAR Comb Sci 22:69–77. doi:10.1002/qsar.200390007 34. Melagraki G, Afantitis A, Sarimveis H, Koutentis PA, Kollias G, Igglessi-Markopoulou O (2009) Predictive QSAR workflow for the in silico identification and screening of novel HDAC inhibitors. Mol Divers. 13:301–311. doi:10.1007/s11030-009-9115-2 35. Li J, Gramatica P (2009) The importance of molecular structures, endpoints’ values, and predictivity parameters in QSAR research: QSAR analysis of a series of estrogen receptor binders. Mol Divers. doi:10.1007/s11030-009-9212-2 36. Papa E, Villa F, Gramatica P (2005) Statistically validated QSARs, based on theoretical descriptors, for modeling aquatic toxicity of organic chemicals in Pimephales promelas (fathead minnow). J Chem Inf Model 45:1256–1266. doi:10.1021/ci050212l 37. Liu H, Papa E, Gramatica P (2006) QSAR prediction of estrogen activity for a large set of diverse chemicals under the guidance of OECD principles. Chem Res Toxicol 19:1540–1548. doi:10.1021/ tx0601509 38. Gramatica P, Giani E, Papa E (2006) Statistical external validation and consensus modeling: a QSPR case study for K(oc) prediction. J Mol Graph Model 25:755–766. doi:10.1016/j.jmgm.2006.06.005 123
Transworld Research Network 37/661 (2), Fort P.O. Trivandrum-695 023 Kerala, India Complex Network Entropy: From Molecules to Biology, Parasitology, Technology, Social, Legal, and Neurosciences, 2011: 17-29 ISBN: 978-81-7895-507-0 Editors: Humberto González-Díaz, Francisco J. Prado-Prado and Xerardo García-Mera 2. Entropy Multi-target QSAR model for Anti-Parasitic and Anti-Alzheimer GSK-3 inhibitors Isela García1, Yagamare Fall1, Generosa Gómez1 and Humberto González-Díaz2 1Department of Organic Chemistry, University of Vigo, Spain; 2Department of Microbiology and Parasitology, Faculty of Pharmacy, USC, Santiago de Compostela, 15782, Spain 1. Introduction In this moment, there is an increasing interest in the evaluation of kinases from unicellular parasites as targets for potential new anti-parasitic drugs. The evolutionary difference between unicellular kinases and their human homologues might be sufficient to allow the design of parasite-specific inhibitors. The Plasmodium falciparum genome contains 65 genes that encode kinases, including three forms of Glycogen synthase kinase-3 (GSK-3). An initial study showed that P. falciparum exports PfGSK-3 to the cytoplasm of host erythrocytes (which are devoid of GSK-3), where it colocalizes with parasite-generated membrane structures known as Maurer´s clefts. 3. The function of PfGSK-3 is unknown, but the presence PfCK1, a CK1 homolog, in infected red blood cell supports the hypothesis that both kinases play a role in regulating the strong circadian rhythm of the parasite, which is responsible for the circadian fevers that are characteristic of this infections disease [1]. Correspondence/Reprint request: Dr. Humberto González-Díaz, Department of Microbiology and Parasitology Faculty of Pharmacy, USC, Santiago de Compostela, 15782, Spain. E-mail:
[email protected]
Isela García et al. 18 The vector-borne parasitic disease African trypanosomiasis, caused by members of the Trypanosoma brucei complex, is a serious health threat. It is estimated that 300,000 to 500,000 humans in sub-Saharan African are infected. If the disease is left inadequately treated, it often has a fatal outcome. Once infection is established, safe and effective therapy is critically important, yet it has been difficult to achieve. Despite the critical need, the available therapies are becoming less satisfactory due to the rising level of resistance to the available drugs, the long period of treatment required to achieve a cure, and the unacceptable and sometimes severe adverse effects associated with current therapies [2]. An urgent priority is to identify and validate new targets for the development of safe, effective, and inexpensive therapeutic alternatives. Compounds that inhibit T. brucei GSK-3 activity and not host GSK-3 might be required for therapy for pregnant women and infants, in that GSK-3 regulates proteins critical in development, such as the wnt gene product. However, optimization of the selectivity of drug candidates for parasite kinases becomes an issue due to the highly conserved amino acids and protein conformation of the catalytic domains [3-6]. Understanding the differences in the substrate binding properties and the three-dimensional structures between mammalian and parasite GSK-3 enzymes is important for the optimization of selected target inhibitors for drug development [7, 8]. To report of all these cases, more parasites, fungi, etc. exist that are keep out for compounds that also disable the enzyme GSK-3, and in it consists our aim objective of this work. On the other hand, Alzheimer´s disease (AD) is the most recent reason of dementia in the elders at present [9]. This serious and degenerative disorder explains the gradual loss of neurons, and in spite of the efforts realized by the big pharmacists of the world, still is not very clear the reason of this pathology. The fundamental characteristic of Alzheimer´s disease is the presence in the brain of two injuries: the neurofibrillary tangles that are formed by paired helical filaments (PHF) whose main component is Tau protein kinase (TPK) and the senile plaques formed by the aggregation of the E-amiloide peptide. In addition, GSK-3 is a serine-threonine kinase encoded by two isoforms in mammals, termed GSK-3D and GSK-3E [10]. Initially GSK-3 was implicated in muscle energy storage and metabolism, but since its cloning, a more generalized role in cellular regulation has emerged, highlighted by the wide array of substrates controlled by this enzyme that includes cytoplasmic proteins and nuclear transcription factors. GSK-3 targets encompass proteins implicated in Alzheimer´s disease, neurological disorders, in the wnt and insulin signaling pathway, glycogen and protein synthesis, regulation of transcription factors, embryonic development, cell proliferation and adhesion, tumorigenesis, apoptosis, circadian rhythm,… etc.
Short title 19 The functions of GSK-3 and its implication in various human diseases have stimulated and active search for potent and selective GSK-3 inhibitors [11]. Studies of GSK-3 homologues in various organisms have revealed physiological roles for the enzyme in differentiation, cell fate determination, and spatial patterning to establish bilateral embryonic symmetry [12]. Purified GSK-3Į and GSK-3ȕ exhibit similar biochemical and substrate properties [12, 13], and is known that in the phosphorylation of Tau protein kinase takes part actively glycogen synthase kinase 3ȕ (GSK-3ȕ), which not only plays a fundamental role in the synthesis of the glycogen (where it was identified by the first time), but it is very important in several processes as cellular signs, metabolic control, embryogenesis, cellular death and oncogenesis [14], and it is related to a wide range of neurodegenerative [15] diseases, bipolar mood disorders [16] and diabetes, by this the inhibition of this enzyme is one of the therapeutic aims more promoters of those which are fighting at present. In 1988 Ishiguro and col. [17] isolated one enzyme when they were studying an extract of the brain that there was showing the generation of paired helical filaments of Tau protein kinase, typical injury of Alzheimer´s disease. TPKI and TPKII are the two kinases implied in this process and they found that TPKI has an identical structure to GSK-3ȕ. In parallel, the development of QSARs using simple molecular indices appears to be a promising alternative or complementary technique to drugprotein docking, high-throughput screening and combinatorial chemistry techniques. Almost all QSAR techniques are based on the use of molecular descriptors, which are numerical series that codify useful chemical information and enable correlations between statistical and biological properties [18-20]. Shannon entropy and other entropy-like measures are one of the more prominent parameters in order to codify structural information on QSAR studies [21, 22]. In this direction, our research group has introduced a novel series of stochastic indices in the so called MARCH-INSIDE approach. The method is based on the use of Markov Models (MM) to calculate absolute probabilities of the distribution of different atomic properties within the molecular skeleton with specific bonding patterns. Using these absolute probabilities we can apply the Shannon formula to calculate entropy parameters of the distribution of atomic properties in the molecule [23, 24]. In this work we will explore the potential of MARCHINSIDE to seek a QSAR for GSK-3 inhibitors from a heterogeneous series of compounds. In the first step, the aforementioned molecular descriptors were calculated for a large series of active/non-active compounds. LDA was subsequently used to fit a classification function. The QSAR developed was then validated with an external predicting series by the re-substitution technique.
Isela García et al. 20 2. Materials and methods Computational methods. The MARCH-INSIDE approach [25-27] is based on the calculation of the different physicochemical molecular properties as an average of atomic properties (ap). For instance, it is possible to derive average estimations of molecular descriptors or group indices [28, 29]. Multi-target Linear Discriminant Analysis (LDA) Linear Discriminant Analysis (LDA) was used to construct the classifiers. One of the most important steps in this work was the organization of the spreadsheet containing the raw data used as input for the LDA because this is not a classic classifier. Herein, the schematisation of the paper is peculiar. Our expectation is to use a two-group Discriminant function to classify compounds into two possible groups: compounds that belong to a particular group and compounds that do not belong to this group. To this end, we have to indicate somehow what group we pretend to predict in each case. In this regard, we made the following steps, these steps are essentially the same given by Concu et al. [30, 31] for the QSAR study of six classes of enzymes: 1. We created a raw data representing each compound input as a vector made up of 1 output variable, 108 structural variables (inputs) divided in values (see the first term of the equation 1), averages (see the second term of the equation 1) and differences between values and averages (see the third term of the equation 1); and the Compound Assay Conditions query (CACq) variable. CACq is an auxiliary not used to construct the model. 2. The first element (output) is a dummy variable (Boolean) called Observed Group (OG); OG = 0 if the compound belongs to the class to which we refer in CACq and 1 otherwise (OG = 1). We could repeat each compound more than once in the raw data. In fact, we could repeat each compound 38 times corresponding to 38 CACq Assay Conditions (see Table 1). The first time we used the CACq = CAC number. It means that we used the real CAC class of the compound in CACq. In this case, the LDA model had to give the highest probability to the group OG = 0 because it had to predict the real class of the compound. The remnant 38 times we use an CAC class number different to the real in CACq and then the LDA model had to predict the highest probability for the group OG = 1. This indicated that the compound did not belong to this group.
Short title 21 Table 1. Compound Assay Conditions query (CACq). q Param. Enz. Iso. Target Type Class Cond. Obs. 12 IC50 (ȝg/mL) no no MRS bacterium 0 NA = 13 IC50 (ȝg/mL) no no S. aureus bacterium 0 NA = 23 IC50 (ȝM) no no MRSA bacterium 0 — = 24 IC50 (ȝM) no no Hep2 bacterium 0 — = 36 MIC (ȝg/mL) no no no bacterium 0 NA = 14 IC50 (ȝg/mL) no no HVC CL 0 NA = 26 IC50 (ȝM) no no HVC CL 0 NA = 27 IC50 (ȝM) no no RD CL 0 NC <> 28 IC50 (ȝM) no no U937 CL 0 2 > 25 IC50 (ȝM) no no HT29 CL 0 NA = 1 % GSK-3 ȕ no enzyme 0 0 = 2 cKi GSK-3 Į no enzyme 0 100 = 9 IC50 (nM) GSK-3 Į no enzyme 0 2000 > 10 IC50 (nM) GSK-3 ȕ no enzyme 0 2000 > 11 IC50 (nM) GSK-3 nd M. intracellulare enzyme 0 2000 > 18 IC50 (ȝM) GSK-3 ȕ no enzyme 0 2 > 19 IC50 (ȝM) GSK-3 nd T. brucei enzyme 0 2 > 20 IC50 (ȝM) GSK-3 nd P. falciparum enzyme 0 2 > 22 IC50 (ȝM) GSK-3 Į/ȕM. intracellulare enzyme 0 2 > 34 IC50 E-9 (M) GSK-3 nd L. donovani enzyme 0 2 > 37 pIC50 GSK-3 ȕno enzyme 0 0 = 38 pIC50 GSK-3 n d enzyme 0 0 = 29 IC50 (ȝM ) no n o C. albicans fungi 0 2 > 15 IC50 (ȝg/mL) no n o C. neoformans fungi 0 NC <>
Isela García et al. 22 Table 1. Continued CE = Cell Efficacy glycogen synthase stimulation, HVC = Human Vero cells, W2 = choroquine resistant W2 clone, D6 = choroquine sensitive D6 clone, CL = Cellular Line, M. tuberculosis h = M. tuberculosis (H37Rv), nd = not determined, NA = not active, NC = not cytotoxicity 3. The problem in this type of organization of raw data is that șk(G) values are global or local compound constants that depend only on structure. Consequently, if these latter and LDA are based only on these values, they will necessarily fail when we change OG values. An inconvenient in this regard occurs if we pretend to use the model for a real enzyme, since we have only one unspecific prediction and we need 38 specific probabilities, 1 confirming the real class and 37 giving low probabilities for the other CACq. We can solve this problem introducing variables characteristic of each CAC class referred on the CACq but without giving information in the input about the real CAC class of the protein. To this end, we used the average value of each șk(G)for all enzymes that belonged to the same CAC class. We also calculated the deviation of the șk(G)from the respective group indicated in CACq. Altogether, we have then 36 șk(G) values + 36 șk(G)avg average values for CAC class + 36 șk(G)dev deviation values from CAC class average = 108 input variables. It is of major importance to understand that we never used as input CACq, so the model only includes as input the șk(G)values for the protein entry and the average and deviations of these values from the CACq, which is not necessarily the real CAC class. The general formula for this class of LDA model is shown below, where S(CACq) is not the
Short title 23 probability but a real valued score that predicts the propensity of a compound to act as an inhibitor of a given class: 0 ,,,,,, 0 ,,,,,, 1S(E) aGdGcGb aGGdGcGb dif k DGk k avg k DGk kk DtGk k avg kk DGk k avg k DGk kk DtGk k ¦¦¦ ¦¦¦ TTT TTTT LDA forward stepwise analysis was carried out for variable selection to build up the models [29]. All the variables included in the model were standardized in order to bring them onto the same scale. Subsequently, a standardized linear discriminant equation that allows comparison of their coefficients was obtained [32]. The square of Canonical regression coefficient (Rc) and Wilk’s statistics (U) were examined in order to assess the discriminatory power of the model (U = 0 perfect discrimination, being 0 < U < 1); the separation of the two groups of proteins was statistically verified by the Fisher ratio (F) test with an error level p < 0.05. Data Set. The data set was conformed to a set of marketed and/or reported drugs/receptor pairs where affinity/non-affinity of drugs with the receptors was established taking into consideration the IC50,ki, pki,... values. In consequence, we managed to collect 1012 cases active compounds in different CACq. In addition, we used a negative control series of 2536 cases of non-active results for compounds evaluated at different CACq. The two data sets used were: training series with 249 active + 638 non-active (887 in total) and validation series with 763 + 1898 = 2661 cases in total. The names or codes for all compounds are depicted in the Supporting Information, due to space constraints, as well as the references consulted to compile the data in this table. This series is composed at random by the most representative families of GSK-3 inhibitors taken from the literature (supplementary material). The remaining compounds were a heterogeneous series of inactive compounds including members of the aforementioned families and compounds including in the Merck index [33]. 3. Results and Discussion General QSAR for GSK-3 inhibitors. The development of a discriminant function [34] that allows the classification of organic compounds as active or non-active is the key step in the present approach for the discovery of GSK-3 inhibitors. It was therefore necessary to select a