The Chemical Space Spanned by Manually Curated Datasets of Natural and Synthetic Compounds with Activities against SARS‐CoV‐2
Abstract
This is the accepted manuscript version of the work published in its final form Betow j Y, Turon, G, Metuge, C. S, Akame, S, Shu, V. A, Ebob, O. T, Duran‐frigola, M, & Ntie‐kang, F. (2025). The chemical space spanned by manually curated datasets of natural and synthetic compounds with activities against sars'cov'2. Molecular Informatics, 44(1). https://doi.org/10.1002/minf.202400293 Deposited by shareyourpaper.org and openaccessbutton.org. We've taken reasonable steps to ensure this content doesn't violate copyright. However, if you think it does you can request a takedown by emailing [email protected].
Full text
J. Y. Betow et al. 1 DOI: 10.1002/minf.200((full DOI will be filled in by the editorial staff)) The chemical space spanned by manually curated datasets of natural and synthetic compounds with activities against SARS-CoV-2 Jude Y. Betow,‡[a,b] Gemma Turon,‡[c] Clovis S. Metuge,[a,b] Simeon Akame, [a,d] Vanessa A. Shu,[a,b] Oyere T. Ebob,[e] Miquel Duran-Frigola,*[c] and Fidele Ntie-Kang*[a,b,f] [a] J.Y. Betow, C.S. Metuge, V.A. Shu, S. Akame, F. Ntie-Kang Center for Drug Discovery, Faculty of Science, University of Buea, P. O. Box 63, Buea, Cameroon, [b] J.Y. Betow, C.S. Metuge, V.A. Shu, F. Ntie-Kang Department of Chemistry, Faculty of Science, University of Buea, P. O. Box 63, Buea, Cameroon *e-mail: [email protected], phone:+237 673872475 [c] G. Turon, M. Duran-Frigola Ersilia Open Source Initiative, Barcelona, Spain *email: [email protected], [d] S. Akame Department of Clinical Microbiology, Faculty of Health Sciences, University of Buea, Cameroon [f] O. T. Ebob Department of Chemistry and Forensics, School of Science and Technology (SST), Nottingham Trent University, Clifton Lane, Nottingham, NG11 8NS, UK [g] F. Ntie-Kang Institute of Pharmacy, Martin-Luther University Halle-Wittenberg, Kurt-Mothes Strasse 3, 06120 Halle (Saale), Germany ‡ These authors contributed equally * Co-corresponding authors Abstract: Diseases caused by viruses are challenging to contain, as their outbreak and spread could be very sudden, compounded by rapid mutations, making the development of drugs and vaccines a continued endeavour that requires fast discovery and preparedness. Targeting viral infections with small molecules remains one of the treatment options to reduce transmission and the disease burden. A lesson learned from the recent coronavirus disease (COVID-19) is to collect ready-to-screen small molecule libraries in preparation for the next viral outbreak, and potentially find a clinical candidate before it becomes a pandemic. Public availability of diverse compound libraries, well annotated in terms of chemical structures and scaffolds, modes of action, and bioactivities are, therefore, crucial to ensure the participation of academic laboratories in these screening efforts, especially in resource-limited settings where synthesis, testing and computing capacity are scarce. Here, we demonstrate a lowresource approach to populate the chemical space of naturally occurring and synthetic small molecules that have shown in vitro and/or in vivo activities against the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) and its target proteins. We have manually curated two datasets of small molecules (naturally occurring and synthetically derived) by reading and collecting (handcurating) the published literature. Information from the literature reveals that a majority of the reported SARS-CoV-2 compounds act by inhibiting the main protease, while 25% of the compounds currently have no known target. Scaffold analysis and principal component analysis revealed that the most common scaffolds in the datasets are quite distinct. We then expanded the initially manually curated dataset of over 1200 compounds via an ultra-large scale 2D and 3D similarity search, obtaining an expanded collection of over 150k purchasable compounds. The spanned chemical space significantly extends beyond that of a commercially available coronavirus library of more than 20k small molecules and constitutes a good starting collection for virtual screening campaigns given its manageable size and proximity to handcurated compounds. Keywords: chemical space, data curation, SARS-CoV-2.
Chemical space exploration of SARS-CoV-2 compounds 2 1 Introduction Chemical space is a well-known concept in cheminformatics, often defined simply as the set of molecules in a company’s chemical inventory or vendor catalogue, and other times defined as the totality of molecules that can potentially be constructed using known reactions and building blocks within a certain range of properties.[1] When all possible molecules that abide by a given set of construction principles are characterised, their chemical space refers to the properties spanned by all these compounds.[1,2] With the rise of the popularity of ultra-large-scale chemical libraries, it becomes crucial to develop efficient ways to encircle small, manageable subsets of molecules with desired properties.[2] In their simplest form, chemical spaces are often limited by certain functional groups, chemotypes, or properties that are easy to calculate from a chemical structure alone.[1b] “Drug-like chemical space” is used in the context of drug discovery to reflect the vast number of molecules with physical properties similar to those of existing small-molecule therapeutics. These properties are often encapsulated in “rules of thumb” like Lipinski’s “rule of five”[2a,2e] and other well-known rules and metrics adhered to by most approved drugs, e.g. “Ghose rule”,[2f] “Veber’s rule”,[2g] “Egans’ rule”[2h] and the quantitative estimate of drug-likeness (QED) metric, among others.[2i,2j] Even this relatively straightforward set of rules unveils significant complexity. For example, it has been shown that all currently known drugs only occupy a very minute portion of the available and/or explorable “synthetically accessible chemical space”.[3] This implies that current molecular libraries cover only a small fraction of the total possible drug-like chemical space if one were to enumerate the compounds resulting from an exhaustive combination of feasible chemical reactions and rules.[4] There have been several attempts to estimate the size of the realistic drug-like chemical space,[3b] including compounds that are directly available to be purchased and screened in biological assays.[3e] A widely-cited estimate is that the number of possible Lipinski-compliant (i.e. with MW < 500 Da) molecules surpasses 1060, which is far beyond the chemical space of bioactive compounds reported in literature-curated datasets like ChEMBL (2·106).[3b,4b] Given the intractable size of the drug-like chemical space, it is necessary to further constrain it with targetor disease-specific properties that will make subsequent (virtual) screening campaigns feasible, especially when resources are limited. Different ways of evaluating the chemical space include using molecular assembly trees, scaffold hopping, similarity search techniques, pharmacophore matching, quantum-based machine learning, and chemography.[3d,5] To efficiently narrow down the drug-like space, datasets of compounds with desirable properties or bioactivities are frequently used as starting points. With the advent of artificial intelligence (AI), (deep) generative models are getting a lot of attention, since they have the potential to rapidly suggest new chemical matter in a property-constrained manner, often taking a known molecule as a seed. At its core, the approach is based on the idea that similar compounds tend to bind to similar targets, which is a guiding principle in chemoinformatics.[6] First, we get a set of starting molecules, and then we explore the surroundings of this set[7] to identify a much larger collection of molecules that are still within the space and could capture the relevant chemical features required for exhibiting the desired biological activity. Today, a wide array of computational approaches for exploring chemical spaces exist, with significant improvements to ensure the synthetic feasibility of the compounds[8], leveraged by the incorporation of AI techniques to characterise and learn the plausible structures associated with target properties.[8b] The huge improvement of computational power available to researchers, including cloud computing[8b,9], has made it possible to generate large virtual collections of potentially interesting compounds, the challenge being now how to choose which of them to synthesise and test within a design-make-test cycle.[9f] For small laboratories and drug discovery centres operating under strong resource limitations, as is the case of many computational drug discovery groups. Moreso, in the context of Africa where our facilities are located, rapid synthesis of the compounds is a true limiting factor. In this scenario, a more practical approach is to limit the search to the purchasable chemical space, which nowadays is far beyond the billion-scale. This is particularly relevant in the search for antiviral drugs, since diseases caused by
Chemical space exploration of SARS-CoV-2 compounds 3 viruses spread very quickly and are quite challenging to contain whenever there is a viral outbreak.[10] In many resource-limited countries, vaccine accessibility, and acceptability, have remained challenging, implying that the quick discovery of small molecules that target viral infections remains one of the ways forward, with computational methods playing a crucial role in the pipeline.[11] Thus, it becomes important to develop diverse and focused compound libraries that could be readily screened to keep pace with possible expected viral outbreaks or mutations. This work aims to explore the chemical space of potential antiviral agents, beginning with a manually curated dataset of synthetically derived and naturally occurring compounds with activities against known severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) targets or with cell growth inhibitory properties against the virus. More specifically, we have analysed the properties of small molecules that have shown in vitro or in vivo activities against SARS-CoV-2 and its target proteins, as reported in the literature. We have then explored the chemical space associated with the starting library using a rapid search within the purchasable space of the Enamine REAL[12] and ZINC[13] libraries. The expanded dataset was compared with the Coronavirus Library available from ChemDiv[14] and drug molecules from DrugBank,[15] to verify if the expanded dataset could be used as a reasonable sized, easy-to-test starting library for virtual screening efforts to target the disease, i.e. a library that could be further easily screened in silico via docking and molecular dynamics, followed by in vitro screening. With limited conventional computing capacity, the goal would be to look for a manageable dataset that does not require high performance computing, e.g. 100,000 to 500,000 compounds. 2 Materials and Methods 2.1 Data collection The hand-curated (natural and synthetic) compound libraries were obtained as follows. The electronic databases employed for the assortment of relevant information include Scopus, NISCAIR, SciFinder, PubMed, Springer Link, Science Direct, Google Scholar, Web of Science, and an exhaustive library search for keywords and combinations of keywords related to “COVID-19”, “SARS-CoV-2”, “compounds”, “small molecules”, etc., and a combination of these terms as previously described.[16] Each individual term and a sum of them, e.g. “COVID-19 + compound”, “SARS-CoV-2 + compound”, “small molecule = COVID-19”, etc were used in the search. This was carried out during the period from January to July 2024. The retrieved articles were checked and compounds showing activities against the virus and/or viral targets were selected. The authors then went ahead and double checked the published papers if there were reported bioactive compounds against SARS-CoV-2 in the retrieved literature sources. The compounds were classified into natural products (NPs) and synthetic derivatives (SDs) according to the information available from the literature sources. The chemical structures were downloaded from the PubChem database, when available.[17] Compounds not available in PubChem were drawn using the ChemDraw Ultra software (version 19.1). Additionally, PubChem and ChemSpider databases were used to check the IUPAC names of the compounds, as previously described.[16] Figure 1A provides a workflow of the manual curation procedure. In summary, in checking through the compounds available in the literature found from the various search engines, if a compound had been tested in clinical trials or had been repurposed for the treatment of COVID-19, it was automatically retained. Compounds that had shown activity in viral assays with >50% growth inhibition or had shown activity in a target-based assay were also kept. The retrieved articles were checked and compounds showing activities against the virus or viral targets (e.g. Mpro, PLpro, Spike/ACE2, RdRp, etc., see Table 1) were selected. The selection criteria were based on the phenotypic and/or target assays (IC50, EC50) reported in the literature, with compounds with IC50 or EC50 < 50 μM retained, while those not falling in this cutoff and those not repurposed for COVID-19 treatment were discarded. The mode of action was either determined from the available experimental in vitro assay results against specific viral enzyme targets or through molecular simulations, e.g. by docking and molecular dynamics/binding affinity calculations. Additional information on the modes of action of the compounds were found by searching the COVID-19 HELP[18] and MedChemExpress[19]
Chemical space exploration of SARS-CoV-2 compounds 4 databases. ChemDiv’s Coronavirus Library (containing 21145 small molecule compounds) has been directly retrieved from the ChemDiv website,[14] while the DrugBank dataset used was version 5.1.10.[15] Both were downloaded in July 2024. 2.2 Principal components and scaffold diversity analysis of manually curated synthetic and natural product libraries The molecular descriptors of the compounds in the two datasets were calculated using the Molecular Operating Environment (MOE) software (version 2016.08, 2016).[20] The computed descriptors included 40 well-known physicochemical parameters like molecular weight (MW), the logarithm of the noctanol/water partition coefficient (log P), the number of Lipinski violations (Lip viol), number of atoms (#atom), synthetic accessibility (SA), the energies of the lowest unoccupied molecular orbitals (LUMO) and of the highest occupied molecular orbitals (HOMO), the number of rotatable bonds, the water solubility, the formal charge (Charge), Oprea lead-likeness score (Oprea Lead), the number of chiral centres (#chiral), the number of basic (#basic) and acidic atoms (#acid), the molar refractivity (mr), the total polar surface area (TPSA), the molecular volume (vol), the dipole moments, the polarizabilities, the number of H-bond donors and acceptors, etc. The dimensionality reduction of the computed descriptors was conducted by principal component analysis (PCA) using MOE.[20] Scaffold analysis was preceded by the Retrosynthetic Combinatorial Analysis Procedure (RECAP)[21] implemented in MOE.[20] This consists in fragmenting each molecule by breaking the bonds that are estimated to be those that can be formed when synthesising each molecule from its constituent building blocks by common synthetic reactions. Thus, a unique extended SMILES string and the fragment’s name, which retains the chemical context of the broken bond, was assigned to each resulting fragment, as described by Weininger.[22] This was applied to both the NPs and SDs datasets to determine the most frequent chemical scaffolds and the statistics on the frequency of the individual fragments were generated, while retaining only scaffolds with at least 10 atoms. 2.3 Chemical properties calculation and ADMET prediction To visualize chemical spaces, t-SNE plots were generated using Uni-Mol[23] and WHALES[24] descriptors. Both were calculated using their implementation in the Ersilia Model Hub (https://ersilia.io/model-hub)[25], references eos39co and eos24ur, respectively. The natural product-like score[26] was calculated using the RDKit package via its implementation in the Ersilia Model Hub (reference eos8ioa). The synthetic accessibility has been calculated using the SYBA package[27] (reference eos7pw8). The SARS-CoV-2 predicted activities have been calculated using the Ersilia implementation of REDIAL-2020[28] (reference eos8fth), and ADMET properties have been calculated using the ADMET-AI package[29] (reference eos7d58). We used the openTSNE implementation (Python) with Euclidean distance, perplexity 30 and 500 iterations. 2.4 Exploration of the chemical space of the manually curated dataset by ultra-large library screening We used the freely available CHEESE API (https://cheese.deepmedchem.com) to search against the ZINC15 and Enamine REAL databases. For each query compound, we used four similarity search modes, namely “2D fingerprint”, “3D shape”, “3D electrostatic” and “consensus”; 100 nearest-neighbours using Euclidean distance were retrieved for each search mode with the “high accuracy” option. All molecules were indexed with their InChIKeys and optionally flattened (i.e. stereochemistry removed) to obtain a de-duplicated list. The data aggregation pipeline resulting from this ultra-large-scale search is fully reproducible from the code repository specified in the Code availability section. 2.5 Compound prioritisation To prioritise the compounds obtained from the ultra-large scale similarity search, we developed two criteria. On one hand, we summed the number of occurrences of each retrieved compound across the
Chemical space exploration of SARS-CoV-2 compounds 5 four search methods (namely Morgan, 3D shape, 3D electrostatics, and consensus), multiplied by the Tanimoto coefficient (Tc) with respect to the query compounds. On the other hand, we developed an ensemble of binary classifiers aimed at scoring the probability of a given compound belonging to the antiSARS-CoV-2 chemical space. As a reference chemical space, we used our manually annotated compounds, and as “negative” (null) sets we used DrugBank compounds and three subsamples of the ChEMBL database (v33) with a maximum positive-negative imbalance of 1:10. We also trained a classifier using the ChemDiv Coronavirus Library as positive, and a 100k-scale diversity library from the same vendor as negatives. All classifiers were trained using Ersilia’s LazyQSAR[11b] framework based on Morgan counts fingerprints (radius 3, 2048 dimensions) and the autoML framework FLAML (random forests and LGBM)[30] with a time budget of 60 seconds. Based on five 80:20 stratified train-test splits, all classifiers satisfactorily performed within the range of 0.75-0.85 AUROC. Finally, since the similarity and the classifier ranks are two genuinely different ranking approaches, we merged them into a consensus rank using the rank averages. 3 Results and Discussion 3.1 Literature review provides a comprehensive curated anti-SARS-CoV-2 library The procedure for gathering literature evidence for the biological activities of the SARS-CoV-2 compounds has been summarised in Figure 1A. It was found that the synthetic compounds belong to quite diverse classes like indoles and peptidomimetics, well-known for their antiviral activities[31], as well as antimalarials like chloroquine and its analogues. The naturally occurring compound library was rich in terpenoids, flavonoids, and alkaloids, including the recently discovered hits like salvinorin A and deacetylgedunin which block SARS-CoV-2 viral cell entry by inhibiting the transmembrane protease, serine 2, an enzyme that in humans is encoded by the TMPRSS2 gene[32]. Our analysis rendered a final dataset of 618 unique NPs and 620 unique SDs (Figure 1B). After data collection, we sought to understand the characteristics of our dataset. As expected, NPs present a higher natural product-likeness score and, conversely, a lower synthetic accessibility when compared to SDs (Figure 1C and D). Interestingly, a retrosynthetic analysis of the NP library provided 421 scaffolds, revealing that oxygencontaining rings like sugars and polyphenol moieties are the most abundant chemical building blocks in their biosynthesis (Figure 1E). On the other hand, 793 chemical scaffolds resulted from the retrosynthetic analysis of the SD library, revealing a higher diversity in terms of ring types and constituent atoms, with many halogen-, O-, Nand S-containing chains and rings. A comparison of the top-ten most abundant scaffolds in each dataset and not abundant (freq <3) in DrugBank scaffolds, revealed that the NP fragments contain sugar moieties, polyphenolic rings, and non-oxygenated aliphatics. In contrast, the SD fragments contain heterocyclic rings, aromatic rings, and aliphatic systems, with multiple N-atoms, fewer O-atoms than in the NP fragments, and some S-atoms and halogens (supplementary Figure S1). To visualise the chemical space of our dataset, we chose two different molecular descriptor techniques. On one hand, Uni-Mol[33] (a deep-learning embedding technique pre-trained on over 209 million molecular conformations) has chemical information in 3D space. On the other hand, WHALES descriptors are a small set of physicochemical parameters that capture both molecular 3D shapes and partial charges, making them suited for scaffold hopping exercises. Figure 1F shows how NPs and SDs cluster together much better when represented with WHALES descriptors, indicating that, despite having dissimilar 3D structures, they may retain similar charge patterns, an essential characteristic to bind to the pockets of their targets in SARS-CoV-2. To further inspect the chemical space of our manually-curated compounds, we used MOE descriptors to build a 2D PCA and analysed the top contributors to defining components 1 and 2 (Figure 1G), which highlighted the descriptors corresponding to Tudor Oprea’s test for leadlikeness (Oprea Lead)[34] and synthetic accessibility (SA),[35] along with the number of basic atoms and the number of H-bond donors, which are all empirical rules that generally characterise drugs and lead compounds. The cumulative variances recovered with the two principal components were 46.72% and 58.33%, respectively. The weights of the descriptors used in PCA analysis have been included in the
Chemical space exploration of SARS-CoV-2 compounds 6 updated Supplementary Data (Data S1). This means that, within our literature-curated collection, there is wide variability in terms of drugand lead-likeness. According to the second principal component, the number of acidic and basic atoms, as well as the LUMO features (often associated with chemical reactivities) contribute to the diversity of the dataset. Finally, we aimed to compare our hand-curated dataset with the chemical space of approved drugs available from DrugBank. To that end, we leveraged a recently published AI/ML model, ADMET-AI[36], which has been trained on reference datasets from the Therapeutics Data Commons[37]. In Figure 1H, we show results from the ADMET-AI predictions as percentiles with respect to approved drugs. Thus, a percentile of 50 means that a given value corresponds to the median value of those observed in the drug space, while extremely high (~100) and low (~0) percentiles indicate deviations from the properties observed in approved drugs. When comparing SD and NP compounds, there was no apparent clear distinction between the two datasets for MW, log P, solubility, inhibition of the cytochrome CYP2C9, NRPPAR-𝛄, and SR-ARE. However, for the descriptors BBB, NR-AR-LBD, and skin toxicity, the NP dataset seems to have a higher proportion of compounds above the 50th percentile, whereas this was the contrary for computed descriptors related to drug absorption (e.g. intestinal absorption and bioavailability), distribution, e.g. ability to cross the blood-brain barrier (BBB), metabolism, e.g. the ability to interact with CYP3A4 enzymes and toxicity e.g. drug-induced liver injury (DILI), carcinogenesis, and inhibition of the human ether-a-go-go-related gene (hERG). It must be mentioned that the dysfunction of hERG often causes cardiac arrhythmia and sudden death, implying that compounds that block hERG channels are considered toxic. Collectively, and as expected, this indicates that natural product compounds tend to present more liabilities, which is why they are often considered as starting points that require further optimization from an ADMET perspective. Both NPs and SDs are skewed towards relatively high MW and low solubility with respect to approved drugs, and, as expected in compounds not yet progressed to the clinics, there is an enrichment of potential CYP liabilities and toxicity pathways, reinforcing the notion that this set of compounds should be used as a starting collection to identify a larger set of optimised compounds. Insert Figure 1 here 3.2 Distribution of compounds by drug target based on literature information In addition, we carefully annotated our curated collection with target information, when possible. A summary of the various targets identified in the literature from in vitro assays and putative targets predicted by molecular simulations is given in Table 1. It was observed that the main protease (Mpro) is the most represented target in the two datasets (36% and 48% for NPs and SDs, respectively). Besides, several compounds have more than one target, including dual protease inhibitors like those that inhibit both Mpro and the papain-like protease (PLpro), as well as those that inhibit both Mpro and the RNAdependent-RNA polymerase (RdRp), and those that inhibit both Mpro and the viral spike in complex with the human angiotensin-converting enzyme 2 (spike/ACE2) and other protein targets. In both the NP and SD datasets, a small number of the compounds inhibit more than two targets and are classified as multitarget compounds, while a significant number have no known target. This last category corresponds to 25% of both NPs and SDs (supplementary Figure S2). Insert Table 1 here 3.3 Ultra-large library screening around the anti-SARS-CoV-2 chemical space Having defined and characterised the chemical space of manually-curated compounds, we carried out a systematic similarity search against two of the most widely used compound libraries for virtual screening, namely ZINC15[13] and Enamine REAL[12]. ZINC is a compendium of commercially available molecules, and Enamine REAL offers an enumerated billion-scale library of make-on-demand molecules based on a large collection of building blocks. Even the most basic chemoinformatics operations such as similarity
Chemical space exploration of SARS-CoV-2 compounds 7 search can become prohibitive at such scales, more so in resource-limited settings where computing capacity is low. Thus, we used the online server CHEESE which leverages an embedding-based method to index compounds and speed up the similarity search. The approach capitalises on recent advances in AI embedding techniques initially developed for image and text data, which require fast queries over extremely large databases. In particular, it uses “semantic similarity” search techniques over small molecule embedding vectors, returning the k-neighbors of the seed compound. An advantage of the CHEESE methodology is that it allows performing 3D-based searches, which can be advantageous when the query molecule is IP-protected or difficult to synthesise, as is the case of NP compounds. We successfully carried out a search of 1231 compounds and obtained, in total, a set of unique 225,774 hits, of which 152,901 remained after flattening out stereochemistry information to remove redundancy. The results of the search correspond to four queries (namely, Morgan (2D) similarity, 3D-shape and 3Delectrostatics, and a consensus measure) against both ZINC15 and Enamine REAL. We retrieved 100 nearest neighbours per search request, obtaining a relatively balanced set of structurally similar compounds, with Tanimoto similarity (Tc>0.7) and more distant ones (Figure 2A). The rankings from the classifier and the similarity search were significantly different and, therefore, we argue that they can be combined in a blended measure that captures both magnitudes. Generally, amongst the top-100 list, ZINC compounds were more abundant than those from Enamine REAL, albeit with more redundancy when the stereochemistry was removed. Enamine REAL is a make-on-demand library based on a predefined set of building blocks and, by definition, it enumerates easily synthesizable compounds. Thus, as expected, natural products were generally less similar to Enamine REAL compounds than ZINC compounds from a 2D structure perspective (Figure 2B), and hits from 3D-electrostatics and 3D-shape searches tended to give more distal compounds, which may be helpful for scaffold hopping. Generally, the consensus CHEESE score captures structural similarity while providing a slightly better balance between Enamine REAL and ZINC hits than a mere Morgan fingerprints search (Figure 2B). We then scored the list of compounds based on (a) their similarity to query compounds and (b) their probability of being associated with the SARS-CoV-2 chemical space. These are two simple and indicative measures that can be used to navigate the relatively large collection (>150k) when screening capacity is limited. To assign a score to the latter, we built an ensemble of binary classifiers capable of discriminating between compounds in our manually curated dataset from randomly sampled compounds in the medicinal chemistry space, as well as between compounds from ChemDiv’s Coronavirus Library and a diverse, agnostic collection from the same vendor. As expected, ZINC compounds were ranked higher in the similarity score (Mann-Whitney statistic 2·109, P-value ~ 0), while we could still find 6,066 compounds from the Enamine-REAL database that were dissimilar (Tc < 0.5) to any compound in the query list but still ranked in the top 20% of the classifier list. In Figure 2C, a few examples are shown where starting from a natural product compound with high natural product-likeness (>2), it was possible to find make-on-demand hits from Enamine REAL (some of them only retrievable via a 3D search in the CHEESE embedding space) that appear to have a high probability of being interpolated in the chemical space associated with SARS-CoV-2. All the scores are annotated in an easy-to-navigate table as specified in the Data Availability section. When we inspected the ADMET properties of the expanded collection (Figure 2E), we observed that especially for Enamine REAL compounds, properties like MW, logP, solubility, BBB penetration and bioavailability were quite centred or well distributed with respect to approved drugs, and certainly much better than those of NPs (Figure 1H). While, generally, some ADMET liabilities remained (e.g. CYPs), in some cases such as the toxicity pathway NR-AR-LBD the profile was much improved with respect to hand-curated compounds. Interestingly, when we mapped the abovementioned ChemDiv Coronavirus Library along with DrugBank compounds and our set of >150k molecules, we observed that the ChemDiv set was focused on a relatively well-defined region (Figure 2E) with respect to our set of compounds and the DrugBank
Chemical space exploration of SARS-CoV-2 compounds 8 collection. This suggests that our expanded library can be a good starting point for screening purposes against SARS-CoV-2 generally. Since this set has been generated with a ligand-centred approach using a diverse set of mechanisms of action, and including both natural and synthetic compounds, the library is expected to have broad applicability within this field of research. Insert Figure 2 here 3.4 Proposed libraries retain SARS-CoV-2 predicted activity The goal of this study is not to provide a short list of anti-SARS-CoV-2 molecules with strong confidence. Rather, we wanted to offer a virtual screening library that can be used as a go-to option in this disease area, ensuring that compounds are purchasable and inspired by compounds with reported evidence in the literature. As an exploratory assessment of the potential of our collection across a broad range of virtual screening tasks related to COVID-19, we chose to use REDIAL-2020, a compendium of opensource machine learning (ML) models containing QSAR predictors for in vitro endpoints of viral load reduction. In Figure 3 we can see that, compared to DrugBank compounds (mimicking an unbiased drug repurposing exercise), both our manual collection and the ChemDiv Coronavirus Library tend to perform better in several tasks, most notably in the AlphaLISA screen testing the spike/ACE2 interaction. Differences were observed between NPs and SDs, with NPs having, for example, higher scores in the 3CL and AlphaLISA predictions, and lower in the TruHit counterscreen. Our expanded library was also enriched in high AlphaLISA scores, although, as in the case of the ChemDiv Coronavirus Library and the SD hand-curated compounds, it would be advisable to control for TruHit counterscreen hits. Other enriched predictions are ACE2 blocking and pseudotyped particle entry (PPE) both for SARS-CoV and MERS, suggesting a broad applicability of our collection. We did not obtain particularly high scores in the 3CL predictions, which means that the library is probably not particularly enriched in this class of compounds. However, note that at a classification score above 0.6 (approximately the median for the manually annotated molecules), we still have 26,440 candidates for this activity. Insert Figure 3 here 4 Conclusions In an attempt to understand the chemical space of potential lead compounds for drug discovery against COVID-19, we have characterised the chemical space of naturally occurring and synthetically derived small molecules that inhibit the growth of the SARS-CoV-2 virus. We have compared the two datasets of compounds hand-curated from the literature by descriptor calculation, principal component analysis and scaffold analysis. It was observed that most of the compounds act by inhibiting the main protease, while several compounds could also be dual and multiple inhibitors. We then derived an expanded chemical space of over 150k purchasable compounds with either 2D or 3D relatedness to the manually-curated collection. It is planned that, in follow-up studies, these compounds will be virtually screened through pharmacophore modelling and protein-ligand interactions with the view of identifying a small subset of ligands that could putatively bind to the targets reported in the literature. These will then be screened in vitro to identify novel antivirals which were not originally reported in the literature. With make-on-demand libraries growing at an exponential rate, it is important to devise ways to efficiently exploit these libraries and use them to develop custom and smaller virtual collections[38] like our African Natural Products Database (ANPDB)[39]. Searching across billion-scale libraries is still computationally intensive and becomes prohibitive in resource-limited settings such as laboratories in Africa, as is our case. Here, we have demonstrated how a well-defined methodology for literature curation, coupled with a simple and fast methodology to search ultra-large chemical spaces, can yield a manageable number of molecules to be used in subsequent virtual screening tasks. We have proved the concept for SARS-
Chemical space exploration of SARS-CoV-2 compounds 9 CoV-2, a pathogen for which we have invested efforts in our group, but the approach is disease-agnostic and could be applied to any other area for which some compounds are annotated in the literature. Given the infrastructural limitations of chemistry laboratories in our setting, we chose to use purchasable compounds from either ZINC or Enamine REAL databases. The size of the current library (150k) is amenable for low-resource computing and we expect it to be useful to other researchers pursuing COVID19 treatments based on small molecules and in a cost-effective manner. Indeed, our overarching plan is to apply this pipeline to other disease areas and targets of interest to our team, including neglected tropical diseases that disproportionately affect people living in the global South. Acknowledgments We acknowledge financial support from the Bill & Melinda Gates Foundation through the Calestous Juma Science Leadership Fellowship awarded to FNK (grant award number: INV-036848 through the University of Buea). FNK also acknowledges joint funding from the Bill & Melinda Gates Foundation (award number: INV-055897) and LifeArc (Grant ID: 10646) under the African Drug Discovery Accelerator program. FNK acknowledges further funding from the Alexander von Humboldt Foundation for a Research Group Linkage project (Ref 3.4-1156361-CMR-IP). We acknowledge the technical support of Dr. Conrad V. Simoben. Author contribution statement Conceptualization: Fidele Ntie-Kang and Miquel Duran-Frigola; Data curation: Jude Y. Betow, Clovis S. Metuge, Simeon Akame, Vanessa A. Shu, and Oyere T. Ebob; Formal analysis: Gemma Turon, Fidele Ntie-Kang and Miquel Duran-Frigola; Funding acquisition: Fidele Ntie-Kang; Investigation: Jude Y. Betow, Gemma Turon, Miquel Duran-Frigola and Fidele Ntie-Kang; Methodology: Jude Y. Betow, Gemma Turon, Miquel Duran-Frigola and Fidele Ntie-Kang; Project administration: Jude Y. Betow, Gemma Turon and Fidele Ntie-Kang; Software: Gemma Turon, Miquel Duran-Frigola and Fidele Ntie-Kang; Resources: Gemma Turon, Miquel Duran-Frigola and Fidele Ntie-Kang; Supervision: Fidele Ntie-Kang and Miquel Duran-Frigola; Validation: Fidele Ntie-Kang, Gemma Turon, and Miquel Duran-Frigola; Writing – original draft: Jude Y. Betow, Gemma Turon, Miquel Duran-Frigola and Fidele Ntie-Kang; Writing – review & editing: everyone. Conflict of interest None declared. Data availability All data used in the study is available for download at https://github.com/ersilia-os/sars-cov-2-chemspace. The library of compounds is referenced in the README file of this repository. Code availability All code used in the study is available for download at https://github.com/ersilia-os/sars-cov-2-chemspace. References [1] a) C. W. Coley, Trends Chem. 2020, 3, 133-145.; b) J. Wang, J. Mao, M. Wang, X. Le, Y. Wang, Methods 2023, 210, 52-59. [2] a) C. A. Lipinski, Drug Discov. Today 2004, 1, 337-341 ; b) J. L. Medina-Franco, A. L.Chávez-Hernández, E. López-López, F. I. Saldívar-González, Mol. Inform. 2022, 41, e2200116 ; c) P. S. Gromski, A. B. Henson, J. M. Granda, L. Cronin, Nat. Rev. Chem. 2019, 3, 119-128; d) C. Dobson, Nature 2004, 432, 824-828; C. Lipinski, A. Hopkins, Nature 2004, 432, 855-861; e) C. A. Lipinski, F. Lombardo, B. W. Dominy, P. J. Feeney, Adv. Drug Delivery Rev. 2001, 46, 3−26; f) A. K. Ghose, V. N. Viswanadhan, J. J. A. Wendoloski, J. Comb. Chem. 1999, 1, 55−68; g) D. F. Veber, S. R. Johnson, H.-Y.Cheng, B. R. Smith, K. W. Ward, F. D. Kopple, J. Med. Chem. 2002, 45, 615−2623; h) W. J. Egan, K. M. Merz, J. J. Baldwin, J. Med. Chem. 2000, 43, 3867–3877; i) G. R. Bickerton, G. V. Paolini, J. Besnard, S. Muresan, A. L. Hopkins, Nat. Chem. 2012, 4, 90–98; j) B. Li, Z. Wang, Z. Liu, Y. Tao, C. Sha, M. He, X. Li, Brief. Bioinform. 2024, 25, bbae321