scieee AI-readable full text Open interactive document viewer

Discrete graph generative models for small molecule generation

García Alfocea, Laura

Abstract

This project introduces CatMol, a generative model designed to create small molecules by representing them as graphs, using node and edge matrices with categorical attributes. A discrete diffusion probabilistic model is applied to these representations, leveraging a marginal distribution derived from the dataset, which improves performance compared to a uniform distribution. To ensure that the model effectively learns the graph distribution, an attention mechanism incorporating edge information is employed. The model is trained on multiple datasets and evaluated using various performance metrics. On drug-like datasets, CatMol demonstrates comparable performance to state-of-the-art models, with superior results in certain metrics. Additionally, the model’s performance improves with the size of the dataset, suggesting scalability. Future work will involve training on even larger datasets to further validate its scalability and potential.

Full text

id192425   DISCRETE GRAPH GENERATIVE MODELS FOR SMALL MOLECULE GENERATION LAURA GARCÍA ALFOCEA Thesis supervisor ALEXISMOLINAMARTINEZDELOSREYES(NOSTRUMBIODISCOVERYSL) Tutor:ALFREDOVELLIDOALCACENA(DepartmentofComputerScience) Degree Master'sDegreeinArtificialIntelligence Master's thesis School of Engineering Universitat Rovira i Virgili (URV) Faculty of Mathematics Universitat de Barcelona (UB) Barcelona School of Informatics (FIB) Universitat Politècnica de Catalunya (UPC) - BarcelonaTech  Abstract This project introduces CatMol, a generative model designed to create small molecules by representing them as graphs, using node and edge matrices with categorical attributes. A discrete diffusion probabilistic model is applied to these representations, leveraging a marginal distribution derived from the dataset, which improves performance compared to a uniform distribution. To ensure that the model effectively learns the graph distribution, an attention mechanism incorporating edge information is employed. The model is trained on multiple datasets and evaluated using various performance metrics. On drug-like datasets, CatMol demonstrates comparable performance to state-of-the-art models, with superior results in certain metrics. Additionally, the model’s performance improves with the size of the dataset, suggesting scalability. Future work will involve training on even larger datasets to further validate its scalability and potential. Resumen Este proyecto presenta CatMol, un modelo generativo diseñado para crear pequeñas moléculas representadas como grafos, utilizando matrices de nodos y aristas con atributos categóricos. Se aplica un modelo probabilístico de difusión discreta a estas representaciones, aprovechando una distribución marginal derivada del conjunto de datos, lo que mejora el rendimiento en comparación con una distribución uniforme. Para asegurar que el modelo aprenda correctamente la distribución del grafo, se emplea un mecanismo de atención que incorpora la información de las aristas. El modelo se entrena con varios conjuntos de datos y se evalúa utilizando diferentes métricas de rendimiento. En conjuntos de datos de moléculas similares a fármacos, CatMol demuestra un rendimiento comparable con los modelos más avanzados, con resultados superiores en algunos casos. Además, el rendimiento del modelo mejora con conjuntos de datos más grandes, lo que sugiere su escalabilidad. El trabajo futuro incluirá entrenar con conjuntos de datos aún más grandes para validar aún más su escalabilidad y potencial. 1 Resum Aquest projecte presenta CatMol, un model generatiu dissenyat per crear petites molècules representades com a grafos, utilitzant matrius de nodes i arestes amb atributs categòrics. S’aplica un model probabilístic de difusió discreta a aquestes representacions, aprofitant una distribució marginal derivada del conjunt de dades, la qual millora el rendiment en comparació amb una distribució uniforme. Per assegurar que el model aprengui correctament la distribució del grafo, s’empra un mecanisme d’atenció que incorpora la informació de les arestes. El model s’entrena amb diversos conjunts de dades i es valora utilitzant diferents mètriques de rendiment. En conjunts de dades de molècules similars a fàrmacs, CatMol demostra un rendiment comparable amb els models més avançats, amb resultats superiors en alguns casos. A més, el rendiment del model millora amb conjunts de dades més grans, cosa que suggereix la seva escalabilitat. El treball futur inclourà entrenar amb conjunts de dades encara més grans per validar més la seva escalabilitat i el seu potencial. 2 Acknowledgements I would like to express my deepest gratitude to the entire team at Nostrum Biodiscovery for their support throughout the development of this project. In particular, I would like to offer my sincerest thanks to Alexis Molina and Álvaro Ciudad, whose guidance was instrumental in making this project a reality. Their mentorship has been invaluable, and I have learned so much from them during the course of this work. I would also like to thank Alfredo Vellido for giving me this opportunity and for his continued support throughout the project, as well as my family and colleagues, whose unwavering support has been a constant source of strength and motivation. 3 Contents List of Figures 5 List of Tables 7 1 Introduction and Objectives 8 2 State of the art 10 3 Methods 12 3.1 Graphs...................................... 12 3.2 Diffusionmodels ................................ 14 3.2.1 Denoising Diffusion Probabilistic Models (DDPMs) . . . . . . . . . 15 3.2.2 Discrete Denoising Diffusion Probabilistic Models (D3PM) . . . . . 16 3.2.3 Diffusion for graphs . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 3.3 Graph Neural Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 3.3.1 Graph transformer . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 3.4 Metrics...................................... 22 4 Results 24 4.1 QM9dataset .................................. 24 4.2 Drug-likedatasets................................ 25 4.3 Ablations .................................... 30 5 Analysis of the sustainability and ethical implications 32 6 Conclusions 33 Bibliography 33 Appendix 37 A Datasets..................................... 38 B Resultsofthemetrics.............................. 39 C Configuration .................................. 46 D Moleculargeneration.............................. 47 E Examples of generated molecules . . . . . . . . . . . . . . . . . . . . . . . 48 F Equivariance................................... 49 4 List of Figures 1 Molecular representation of acetamide (CC(=O)N) (representing the categories, not the encoded values). . . . . . . . . . . . . . . . . . . . . . . . . 13 2 Diffusion process for an image. . . . . . . . . . . . . . . . . . . . . . . . . . 14 3 Diffusion process for a molecule. . . . . . . . . . . . . . . . . . . . . . . . . 15 4 Traditional attention mechanism and its modification to add edge feature information.................................... 20 5 Architecture of the Graph Transformer model and architecture of each of itslayers. .................................... 22 6 Visualization of the generation process of a valid molecule using ZINC250k dataset. ..................................... 24 7 Comparation of the distributions of the generated molecules training the model with QM9 dataset along with the distributions of the QM9 dataset. 26 8 Comparation of the distributions of the generated molecules training the model with ZINC250k dataset along with the distributions of the ZINC250k dataset. ..................................... 27 9 Comparation of the distributions of QED and LogP of the generated molecules training the model with ZINC250k dataset along with the distributions of theZINC250kdataset.............................. 28 10 Comparation of the distributions of SA and molecular weights of the generated molecules training the model with ZINC250k dataset along with the distributions of the ZINC250k dataset. . . . . . . . . . . . . . . . . . . . . 29 11 QED and LogP distributions of the generated molecules with ZINC250k dataset using our model and Digress, along with the dataset distributions. 39 12 SA and molecular weight distributions of the generated molecules with ZINC250k dataset using our model and Digress, along with the dataset distributions. .................................. 40 13 Comparation of the distributions of the generated molecules training the model with MOSES dataset along with the distributions of the MOSES dataset. ..................................... 41 14 Comparation of the distributions of QED and LogP of the generated molecules training the model with MOSES dataset and along with the distributions oftheMOSESdataset.............................. 42 15 Comparation of the distributions of SA and molecular weights of the generated molecules training the model with MOSES dataset and along with the distributions of the MOSES dataset. . . . . . . . . . . . . . . . . . . . 43 16 QED and LogP distributions of the generated molecules with a uniform distribution along with the distributions of the ZINC250k dataset. . . . . . 44 5 17 SA and molecular weights distributions of the generated molecules with a uniform distribution along with the distributions of the ZINC250k dataset. 45 18 Examples of generated valid molecules training the model with the QM9 dataset. ..................................... 48 19 Examples of generated valid molecules training the model with the ZINC250k dataset. ..................................... 48 20 Examples of generated valid molecules training the model with the MOSES dataset. ..................................... 48 6 Therefore, a molecule can be represented by these two matrices, Xand E, both containing categorical information: the nodes and bond types. Representing these features using one-hot encoding is particularly advantageous in the generative modeling process. This representation enables the model to directly predict the probabilities of each node belonging to a specific atomic class and each edge corresponding to a specific bond type. By leveraging this probabilistic framework, the generative model can effectively capture the categorical nature of molecular graphs and facilitate the generation of realistic molecular structures. For example, consider the molecule acetamide (C2H5NO), represented by the SMILES string CC(=O)N. In this molecule, there are three atoms: carbon (C), nitrogen (N), and oxygen (O). The node features, X, represent the atomic numbers of each atom in the molecule. The atomic numbers for carbon, nitrogen, and oxygen are 6, 7, and 8, respectively. The adjacency matrix Eis a 4x4 matrix encoding the bonding information. Specifically, there is a single bond between the two carbon atoms (C-C), a single bond between carbon and nitrogen (C-N), and a double bond between carbon and oxygen (C=O). The representation of this molecule, before applying one-hot encoding, can be seen in Figure 1. (a) Acetamide molecule 2D representation using RDKit. X=    6 6 8 7     (b) Node features matrix X. E=    0100 1021 0200 0100     (c) Adjacency matrix E. Figure 1: Molecular representation of acetamide (CC(=O)N) (representing the categories, not the encoded values). However, it is important to note that not every pair of matrices Xand Eactually represents a valid molecule. This is due to inherent chemical constraints that must be satisfied for a molecular structure to be chemically plausible, such as: •Valency: Each atom has a specific valency, determined by its electron configuration, which refers to the number of bonds an atom can form. For instance, carbon can typically form 4 bonds, oxygen 2 and hydrogen just 1. •Bond types: Each type of bond has different properties, such as the number of shared electrons, which must be respected. For instance, in single bonds one pair of electrons is shared while in double bonds two pairs of electrons are shared. •No fragmentation: A graph that represents a molecule cannot have disconnected components unless it’s a set of molecules. These are some of the intrinsic properties that the generative model should learn in order to generate realistic, chemically valid molecules. Another key property of molecular graphs is that the atoms (nodes) do not have a fixed order. This means that the structure of a molecule remains unchanged regardless of how the nodes are permuted or reordered. 13 This permutation invariance is a crucial consideration when designing machine learning models for molecules, as the model must be able to process molecular graphs without relying on any arbitrary node ordering. 3.2 Diffusion models Generative models are a class of machine learning models designed to generate new data that is similar to a given dataset by learning the underlying distribution of the data itself. Once trained, these models can generate new, unseen samples that resemble the data they were trained on, by just sampling from this learned distribution. Many generative models, such as Variational Autoencoders (VAEs) [27] or Generative Adversarial Networks (GANs) [22], work by learning a latent space, i.e., a lower-dimensional representation of the data and then generating data by sampling from this latent space and decoding or transforming it back into the original data space. Another class of generative models, diffusion models [28, 29], operate through a different mechanism. These models are characterized by the progressive introduction of noise into data and progressive removal of that noise. These processes are inspired by principles of thermodynamics, particularly physical diffusion phenomena, such as the spreading of particles in a fluid, where transformations occur incrementally. Denoising Diffusion Probabilistic Models (DDPMs) [28] are a specific subclass of diffusion models that use two Markov chains. The forward process adds noise to the data following a parameterized Markov chain, where the data at each timestep depends only on the previous timestep. Similarly, the reverse process is also modeled as a Markov chain, progressively denoising the data in a stepwise manner. The diffusion process for an image can be seen in Figure 2. Figure 2: Diffusion process for an image. In diffusion models designed for generating new graphs, the forward process begins with a graph that is gradually corrupted over time by introducing noise, leading to a 14 loss of structure and increased randomness. The model then, during the reverse process, learns how to reverse this corruption process, recovering the original graph from its noisy version, effectively learning to generate new graphs that resemble the training data. These two processes can be seen in Figure 3. Figure 3: Diffusion process for a molecule. In this case, since nodes and edges are discrete features representing the type of edge and the type of bond, respectively, it is necessary to use categorical diffusion models. These models extend standard diffusion models to effectively handle discrete or categorical data by progressively introducing noise in a manner that preserves the structure of categorical variables. Discrete Denoising Diffusion Probabilistic Models (D3PM) [30] represent a pioneering framework within this class of categorical diffusion models. It models the diffusion process with transitions between discrete states using a transition probability matrix, which defines the probability of transitioning between discrete categories at each timestep. 3.2.1 Denoising Diffusion Probabilistic Models (DDPMs) In DDPMs, the forward process is modeled as a Gaussian noise-adding process, so, for any data object x, the Markov process for adding noise is the following, denoting by xt its one-hot encoding: q(xt|xt−1) := N(xt;√1−βtxt−1, βtI),(3.1) where βtis the variance schedule. This leads to the following closed-form expression: q(xt|x0) = N(xt;√¯αtx0,(1 −¯αt)I),(3.2) where ¯αt:= ∏t s=1 αs, being αt:= 1 −βt. As it can be seen from equation 3.1, the variance schedule β1, ..., βTdetermines the quantity of noise that is added in each timestep. Therefore, it should be chosen to balance 15 the trade-off between sufficiently destroying the original data to ensure good generative performance. We have implemented it using a cosine beta schedule [31]. After adding noise to the data during the forward process, a neural network is used to learn to remove that noise and recover back the original molecule. This model takes as input the noisy data xtat a given timestep tand the timestep itself and predicts the logits of the clean original data pθ(x|xt). The timestep information is added to the model by using sinusoidal positional embeddings [32], which provide a unique representation for each position in the sequence. The optimization is done by using the variational lower bound on the negative log likelihood, which can be written as the sum of three different terms: L=Lprior +Ldiff −Eq[log pθ(x|x1)] ,(3.3) with Lprior =DKL (q(xT|x)∥p(xT)) ,(3.4) and Ldiff = T ∑ t=1 DKL (q(xt−1|xt,x)∥pθ(xt−1|xt)) ,(3.5) where DKL represents the Kullback-Leibler divergence between two distributions and p(xT=N(xT;0,I)) represents the limit distribution, a Gaussian distribution in the case of DDPMs. The reverse process is also modeled by a Markov chain using the learned Gaussian transitions and starting from the limit distribution. The whole chain can be written as: pθ(x0:T) = p(xT) T ∏ t=1 pθ(xt−1|xt),(3.6) while each step follows: pθ(xt−1|xt) = N(xt−1;µθ(xt, t),Σθ(xt, t)),(3.7) where Σθis usually fixed for simplicity and the expression of µcan be obtained using the prediction of the model. The denoised data is obtained by applying these formula gradually for all the timesteps. 3.2.2 Discrete Denoising Diffusion Probabilistic Models (D3PM) D3PMs are a generalization of DDPMs to categorical data, by going beyond corruption processes to processes with different possible transition matrices [33]. Considering a categorical variable that can take on Kdifferent categories, i.e., xt, xt−1∈ {1, . . . , K}, the forward process is modeled using transition probability matrices Qt, where [Qt]ij =q(xt=j|xt−1=i)specifies the probability of transitioning from category ito category jat timestep t. With these transition matrices, the probability of xtgiven xt−1, representing by xt and xt−1their one-hot encoding, is expressed as: q(xt|xt−1) = Cat(p=xt−1Qt),(3.8) 16 which is a categorical distribution with probabilities given by the row vector p. Therefore, q(xt|x0) = Cat(xt;p=x0Qt),with Qt=Q1Q2. . . Qt.(3.9) It can be seen that, unlike continuous diffusion, categorical diffusion models allow for a variety of transition probability matrices through the design of Qt, which defines the limit distribution. This limit distribution corresponds to the probability distribution of graphs that represent pure noise. Some of the possible transition matrices are the following ones: •Uniform: The transition probability for every state is uniform: Qt= (1 −βt)I+ βt K1K×K. In this case, cumulative products can be computed in close form Qt= αtI+βt K1K×K.The limit distribution is therefore uniform. •Marginal: The marginal transition matrix represents the probabilities of the dataset itself, ensuring that the transitions respect the inherent structure or distribution of the dataset. In this case, the matrix is designed based on the observed frequencies of different categories in the data: Qt= (1−βt)I+βt1KmT, where mrepresents the marginal distribution of the dataset. The cumulative product also has a closed form in this case: Qt=αtI+βt1KmT. The limit distribution in this case corresponds to the distribution of the dataset. It has been tested that considering the marginal distributions of the dataset can speed up both the training and sampling processes. This is because noised graphs are closer to real graphs than in the case of uniform noise, in which noised graphs are really random and therefore very far from representing real molecules. Therefore, we consider this limit distribution, which is different for each of the datasets. In the case of categorical diffusion, the noisy data is sampled from the limit distribution and the sequence of graphs obtained during the denoising process can be obtained using the following formula: pθ(xt−1|xt)∝∑ x0 q(xt−1,xt|x0)pθ(x0|xt),(3.10) where and x0is the predicted clean data obtained from the trained model. The reverse process consists in applying this formula for each of the timesteps until obtaining the final completely denoised data x0. The tractable expression to compute the reverse process is the following: q(xt−1|xt,x0) = q(xt|xt−1,x0)q(xt−1|x0) q(xt|x0)=Cat(xt−1;p=xtQ⊤ t⊙x0Qt−1 x0Qtx⊤ t)(3.11) where ⊙denotes element-wise multiplication. 3.2.3 Diffusion for graphs In the case of considering molecules represented as graphs, noise can be added independently to nodes and edges [3] and, since they both are categorical data, there are two transition matrices QX tand QE tand the categorical distribution of the noisy data is determined by the following probabilities: 17 q(Gt|G) = Cat(XQX t,EGE t),(3.12) with QX t=QX 1QX 2. . . QX tand QE t=QE 1QE 2. . . QE t,(3.13) all of them obtained using the marginal distributions of nodes and edges from the dataset, as explained before. Once this computation is done, the noisy graph at the particular timestep t,Gt= (Xt,Et), is obtained by sampling the nodes and edges from the obtained distributions. Regarding the reverse process, the noisy graph, along with their number of nodes, are sampled from the limit distribution, GT∼qX×qE. This graphs is progressively denoised, obtaining the sequence of graphs GT−1, ..., G1, until obtaining the graph without noise, G0. The sequence of noisy graphs can be obtained using equations 3.10 and 3.11 independently for nodes Xand edges E. With these equations, the graph is denoised step by step for the Ttimesteps until the final graph G0is obtained. After adding noise to the data during the forward process, a neural network is used to learn to remove that noise and recover back the original molecule. This model takes as input the noisy graph Gt= (Xt,Et)at a given timestep tand the timestep embedding and predicts the logits of the clean original data pθ(G|Gt). The optimization of the parameters of the model to predict the original molecule could be done by optimizing the variational lower bound on the negative log likelihood, which can be written as follows: L=Lprior +Ldiff −Eq[log pθ(G|G1)] ,(3.14) with Lprior =DKL (q(GT|G)∥qX×qE),(3.15) and Ldiff = T ∑ t=1 DKL (q(Gt−1|Gt, G)) ∥pθ(Gt−1|Gt).(3.16) The terms qXand qEare the marginal distributions of nodes and edges of the dataset and pθ(G|G1)are the probabilities for the clean graph given the last noisy graph G1. Each of the described terms have to be computed independently for the nodes tensor Xand for the edges tensor E. However, recent diffusion models are optimized using different objectives. We only optimize the last term of the lower bound. The first term (equation 3.14) in the case of categorical diffusion is always zero since both distributions are the limit distributions. The second term (equation 3.15) was very small and has not been considered in our model because of numerical stability (Section 4.3). Therefore, the only term left is the last one, whose expected value represents the cross entropy. Therefore, the model has been trained by optimizing the cross entropy between the probabilities for the predicted node and adjacency matrix tensors and the original ones. 18 3.3 Graph Neural Networks From the representation of molecules as graphs explained in Section 3.1, two key aspects can be highlighted. The first one is the structure. In the case of images, data is represented on a regular grid where each pixel has a fixed number of neighbors, as every point is uniformly connected to its immediate surroundings. In contrast, in graphs, the number of neighbors for each node is determined by the number of edges it possesses, leading to a varying number of neighbors for different nodes. Furthermore, during the molecule generation process, edges are continuously added and removed, resulting in a dynamic and irregular structure. This irregularity introduces challenges for the direct application of traditional convolutional operations. The second aspect concerns the ordering of nodes. A molecule with a reordering of the nodes will have a different nodes tensor and adjacency matrix; however, the molecule it represents remains the same. Therefore, the model’s output should remain consistent for both reordered molecules. To address these challenges, Graph Neural Networks (GNNs) [34] have been developed as specialized architectures capable of effectively processing and learning from the dynamic and irregular nature of graph-structured data, while also preserving node equivariance. Node equivariance ensures that the model’s predictions remain consistent regardless of the node ordering or permutation. This property is crucial for capturing the inherent symmetry in graph data and improving generalization across various graph structures. GNNs are a class of machine learning models specifically designed to work with graphstructured data. At each layer of the GNN, a node aggregates information from its neighbors and updates its own representation. This process enables the network to capture how the surrounding context influences a nodes state, which is specially useful in molecular modeling, where the properties of an atom can be heavily influenced by the atoms it is bonded to. The key idea behind GNNs is message passing [35]. During the forward pass, each node sends information (messages) to its neighbors along the edges of the graph. These messages are aggregated, usually by some form of summation, averaging, or concatenation and used to update the node’s feature representation. This process allows the GNN to learn how the surrounding context influences a node’s state. After the message passing and node updates, a readout function is often used to aggregate information across all nodes in the graph to produce a global representation. 3.3.1 Graph transformer The self-attention mechanism used in Transformer models [32] can be viewed as a flexible and dynamic way to aggregate information in GNNs. This self-attention mechanism allows to process all tokens in a sequence simultaneously, enabling highly efficient parallelization and the ability to capture long-range dependencies. It is obtained by computing a set of attention scores that allow each token to focus on every other token in the sequence using the Query (Q), Key (K) and Value (V) vectors. The self-attention scores are given by: 19 (a) Attention mechanism. (b) Attention mechanism modified using edge features (E). Figure 4: Traditional attention mechanism and its modification to add edge feature information. Attention(Q, K, V ) = softmax(QKT √dk)V. (3.17) We also apply normalization layers to the intermediate representations of the attention scores, due to the fact that attention logits can grow uncontrollably, leading to almost deterministic (one-hot) attention distributions, causing vanishing gradients [36]. This normalization has improved our training stability. The original transformer architecture was designed for Natural Language Processing (NLP) tasks, where it operates on sequences of words that are implicitly represented as fully connected graphs, with connections between all words. However, this approach performs poorly when not every two pairs of nodes are connected. Because of this, the self-attention mechanism should also take into account the information about whether two nodes are connected or not and how they are connected. This is done by multiplying the attention scores by these features using an element-wise product, allowing therefore the attention mechanism to take into account the connectivity of the nodes. The attention mechanism incorporating edge feature information is illustrated in Figure 4. This figure presents a schematic representation of the attention mechanism: one version (Figure 4a) shows the standard formulation using only the matrices Q, K and V derived from the nodes tensor (X), while the other (Figure 4b) includes the additional edge feature information (E) as proposed in [37]. In both cases, the output consists of the attention scores. The graph transformer architecture [37] was designed to address this limitation by adapting the transformer architecture to graph-structured data. This is done by, first of all, incorporating edge features information to the attention mechanism as explained before, enabling the model to capture the relationships between nodes based on the graph structure. Moreover, the positional encodings are modified to better reflect the complex topology of graphs. Given that graphs do not have a natural order (like sequences) and that transformer models process the entire input sequence in parallel, one key challenge is the inability 20 to model sequential order. To address this, positional encodings PE are introduced to provide information about the relative or absolute position of tokens in the sequence. However, as discussed before, there is no single "correct" way to assign unique positions to nodes. Because of this, positional encodings on graphs should be based on their structural properties. One common type of PE used for graphs are Random Walk PE [38]. These are computed using a random walk diffusion process and leverage the concept of random walks, which are paths through a graph that move from one node to another based on the graph’s structure. The idea is to calculate the probability that a random walk starting at node iwill return to node iafter ksteps, which are represented as: pRW P E i= [RWii, RW 2 ii, . . . , RWk ii]∈Rk,(3.18) where RWii denotes the self-return probability for node iin a random walk to return to itself and kis the number of random walk steps considered. It is obtained using the random walk matrix RW =AD−1, where Ais the adjacency matrix and Dthe degree matrix. We use random walks as positional encodings. They are added to the input data during training as learnable parameters, using a dropout factor that randomly turns them to zero, allowing the model to operate without relying on them during inference [38]. The backbone of the model is illustrated in Figure 5a, where the input consists of the node tensor X, augmented with random walks as positional encodings (PE) and timestep embeddings, as well as the edge feature tensor E. These inputs are processed through Feed Forward Networks (FFN), which serve to embed the additional information into the model. A series of layers is then applied, followed by another FFN to produce the final output. The structure of each layer is detailed in Figure 5b. As shown, the process inside each of the layers begins by applying the self-attention mechanism across the heads. The attention scores from each head are then concatenated, followed by the application of a series of linear layers, along with some addition and normalization layers to refine the node and edge representations and generate the final output. 21 (a) Backbone of the Graph Transformer model. (b) Each layer of the Graph Transformer model. Figure 5: Architecture of the Graph Transformer model and architecture of each of its layers. 3.4 Metrics We will explain some of the most relevant metrics in 2D molecule generation, proposed by [4]. It is important to note that while all these metrics assess the quality of the generated set, the relevance of each metric depends on the specific task the model is designed for. Optimizing one metric may come at the expense of another, so the choice of metrics should align with the desired application. •Validity: Measures the proportion of generated molecules that are chemically valid, meaning they satisfy essential chemical constraints. A high validity score indicates the model has learned the dataset’s underlying distribution and intrinsic properties, while a low score suggests the model generates invalid molecules and fails to capture the data distribution properly. •Uniqueness: Reflects the proportion of unique molecules generated, after removing duplicates. A high uniqueness score suggests the model can generate a diverse set of new molecules, whereas a low score indicates model collapse, where the model repeatedly generates similar or identical molecules. •Novelty: Measures the proportion of valid generated molecules not present in the training set. A low novelty score indicates overfitting, where the model replicates the training data instead of generating new, unseen samples. •Fréchet ChemNet Distance (FCD): Compares the distribution of generated molecules with the training dataset using molecular embeddings [39]. These embeddings are derived from a neural network trained to predict biological properties. FCD unifies the validity, chemical and biological relevance of the molecules and diversity. It is important to note that while the Frechet ChemNet Distance (FCD) is a reliable metric for measuring the similarity between the generated distribution and the dataset distribution, it relies on embeddings derived from the activations of a neural network 22 (a) SA distributions. (b) Molecular weights distributions. Figure 10: Comparation of the distributions of SA and molecular weights of the generated molecules training the model with ZINC250k dataset along with the distributions of the ZINC250k dataset. Metric CatMol Digress VAE Validity 73.47% 85.7% 97.7% Uniqueness 99.39% 100.0% 99.8% Novelty 99.26% 95.0% 69.5% FCD 4.29 1.19 0.57 Table 3: Comparison of the evaluation metrics obtained for the MOSES dataset using CatMol, Digress and a VAE. The results have been obtained from [3]. The results obtained on the MOSES dataset surpass those achieved on ZINC250k. This indicates that as the size and diversity of the dataset increase, the model is better able to learn the underlying data distributions and generate molecules that more closely adhere to these patterns. A larger dataset provides more training samples and greater variability, which enhances the model’s capacity to generalize, resulting in higher validity and novelty scores. These findings indicate that the model’s performance is highly dependent on the quality and diversity of the training dataset. To further validate this, the model will be tested on a significantly bigger subset of the ZINC dataset, which encompasses the MOSES and ZINC250k datasets along with many additional samples. This dataset is 29 significantly larger and more diverse, providing a richer foundation for training. This next step aims to prove that increasing the dataset’s size and complexity can further improve the models capacity to generate valid and novel molecules that closely resemble real-world molecular distributions. 4.3 Ablations To demonstrate the importance of including positional encodings as part of the input of the generative model, we conducted additional experiments using the ZINC250k dataset. Specifically, we trained the model without any positional encodings and adding them to the input to all the molecules in the dataset. The results obtained in both cases cases were significantly worse than the results obtained with our standard setup where random walk positional encodings are added to the input with a dropout factor of 0.7. Without positional encodings, the generated atom and bond distributions remained closely aligned with those of the dataset. However, the validity rate dropped to 46.73%, with almost all generated molecules being novel and unique. This suggests that while the model maintains diversity, it struggles to generate a high proportion of valid molecules without positional encodings. Additionally, the FCD metric was measured at 12.19, indicating that the generated molecules were less similar to those in the dataset compared to the results obtained when using positional encodings. On the other hand, we also trained the model with positional encodings always present during training. This was done by setting the dropout factor to zero, ensuring the model consistently had access to these features. This resulted in the generation of molecules composed of many disconnected fragments. From this result we can conclude what, when the model was conditioned on positional encodings during training, it relied heavily on this information to understand the graph’s structure. Because of that, when positional encodings were absent during sampling, the model struggled to generate well-connected graphs. This result emphasizes the importance of positional encodings in capturing the molecular structure connectivity. Similarly, we have also measured the effect of the weight of the loss for the edges. When training with a loss that assigns equal weight to both nodes and edges, the number of valid molecules generated dropped to 46.26%. This suggests that the model generates fewer valid molecules, likely due to incorrect connections between nodes. This empirically demonstrates that edges are the most challenging aspect for the model to generate without compromising molecular validity. By increasing the weight of the edges loss to 10, we achieved the best results with the highest validity score. This outcome aligns with the observation that generating correct bonds is the most difficult task, as invalid molecules often contain incorrect bonds. This issue is particularly evident in the generation of molecular rings, where the model struggles due to the high number of edges involved. We also investigated the effect of the limit distribution on the model’s performance. Specifically, we trained the model using a uniform distribution, where each node and edge was sampled with the same probability, instead of using a limit distribution derived from the dataset. For this experiment, we used the ZINC250k dataset. Although the bond and edge distributions of the generated molecules were again close to those of the dataset, the resulting molecules were less accurate compared to those generated using the marginal distributions as the limit distribution. In particular, the percentage of valid molecules 30 dropped to 42.84%, and the FCD score increased to 11.84, indicating a notable decline in similarity to the dataset. As shown in Table 4 and Figures 16 and 17 in the appendix, the distributions obtained under the uniform limit distribution fail to reproduce the dataset metrics as effectively as those obtained using the marginal distributions (Figures 9 and 10). These findings reinforce that adopting the dataset distribution as the limit distribution significantly enhances the generation process, leading to improved validity and closer resemblance to the original dataset. We also attempted to train the model using the KL divergence term (equation 3.16) from the lower bound as part of the loss function. However, we observed that this term, being very small, led to NaN values during training, making this training impossible. As a result, we opted to train the model solely with the cross-entropy loss, which successfully stabilized the training process. The results of all the ablation experiments can be seen in Table 4. It is important to note that the metrics were not calculated for the model trained adding always the positional encodings. This is because, in this case, the valid molecules generated were limited to very small fragments, often consisting of just one or two atoms. As a result, the metrics in this scenario are not meaningful. Metric Without PE Equal weight Uniform Validity 46.73% 46.26% 42.84% Uniqueness 99.97% 99.98% 99.95% Novelty 100.00% 100.00% 99.99% FCD 12.20 11.98 11.84 Table 4: Evaluation metrics for the ablation experiments. The first experiment involves training the model without any positional encoding, the second experiment uses a loss function with equal weights for both nodes and edges, and the final experiment trains the model using a uniform distribution as the limit distribution. 31 Chapter 5 Analysis of the sustainability and ethical implications The development of the generative model has involved significant environmental, economic, and social considerations. From an environmental standpoint, the training process required multiple experiments with varying model architectures and refinements to ensure optimal performance. This resulted in extensive energy consumption over approximately 1821 hours of training, leading to an estimated emission of 275.34 kg of CO2 1. The economic impact is also considerable, as the training process relied on high-performance A30 with an Ampere architecture, which come with substantial costs both in terms of hardware acquisition and electricity consumption. The execution phase of the project presents a different perspective. Environmentally, the model’s ability to computationally generate new drugs has the potential to reduce the need for laboratory-based testing, which often involves significant plastic consumption and other non-recyclable materials. The model’s efficiency in exploring a vast chemical space computationally minimizes physical waste and experimental errors. Economically, the precise benefits are tied to reduced material costs and faster drug candidate generation. Socially, the use of computational tools reduces the need for experimental molecule verification, a process that would be highly inefficient and labor-intensive given the vastness of chemical space. Regarding the risks and limitations of the project, the economic limitations are clear, as the model could not be trained on larger datasets due to time constraints and the limited availability of computational resources. Socially, while the model is designed for drug discovery in the pharmaceutical industry, there is a risk that such technology could be misused by malicious actors to generate harmful substances, raising ethical concerns about dual-use research [51]. Furthermore, the reliance on advanced computational resources raises questions about the accessibility and equitable distribution of the project, since it could only be effectively used with sufficient technological infrastructure. In alignment with professional ethical standards, the work conducted in this project respects the principles of scientific integrity, transparency, and responsible use of technology. The methodologies and datasets have been chosen and documented with care to avoid potential misuse, and the development process follows guidelines for safe and ethical AI use, emphasizing positive societal impact while mitigating risks. 1https://calculator.linkeddata.es/ 32 Chapter 6 Conclusions To conclude, our experiments demonstrate the effectiveness of the proposed model in generating valid and diverse molecules across multiple datasets, including QM9, ZINC250k, and MOSES. The model’s performance was assessed using various evaluation metrics, including validity, uniqueness, novelty, and FCD, and comparisons with other state-of-theart generative models. The results indicate that our model is able to generate molecules that closely resemble the original dataset’s distribution, particularly in terms of atom and bond distributions, as well as other molecular properties distributions, while also maintaining high novelty and uniqueness scores. The model successfully learned the distribution of the QM9 dataset, and its performance improved when tested on larger datasets, such as ZINC250k and MOSES, which provided more varied and complex molecular structures for training. These results suggest that the models ability to generalize and generate novel molecules improves as the dataset size increases. For future work, we plan to enhance the model by training it on even larger and more diverse datasets. This will allow us to assess whether the model can capture more intricate molecular features and potentially improve its performance in terms of both validity and diversity. Specifically, we aim to train the model on a higher ZINC subset, containing 9.26 million molecules. This will enable us to measure the improvement of the model when increasing the size of the dataset. We also intend to explore downstream applications of the model for generating molecules tailored to specific purposes. A key focus will be on property-conditioned molecule generation, enabling scaffold hopping to create novel compounds with similar properties to existing molecules while avoiding patent restrictions. Another important direction is adapting the model to generate molecules with 3D atomic positions. This extension will allow the model to predict not only the atomic and bond structures but also their spatial arrangements. Such advancements are particularly relevant for generating small molecules capable of binding effectively to target proteins, paving the way for designing compounds with therapeutic potential to modulate protein activity. 33 Bibliography [1] F. Eijkelboom, G. Bartosh, C. A. Naesseth, M. Welling, and J. W. van de Meent. Variational flow matching for graph generation. arXiv preprint arXiv:2406.04843, 2024. [2] H. Tang, C. Li, S. Kamei, Y. Yamanishi, and Y. Morimoto. Molecular generative adversarial network with multi-property optimization. arXiv preprint arXiv:2404.00081, 2024. [3] C. Vignac, I. Krawczuk, A. Siraudin, B. Wang, V. Cevher, and P. Frossard. DiGress: Discrete denoising diffusion for graph generation. In The Eleventh International Conference on Learning Representations, 2023. [4] D. Polykovskiy, A. Zhebrak, B. Sanchez-Lengeling, S. Golovanov, O. Tatanov, S. Belyaev, R. Kurbanov, A. Artamonov, V. Aladinskiy, M. Veselov, A. Kadurin, S. Johansson, H. Chen, S. Nikolenko, A. Aspuru-Guzik, and A. Zhavoronkov. Molecular Sets (MOSES): A Benchmarking Platform for Molecular Generation Models. Frontiers in Pharmacology, 2020. [5] J. P. Hughes, S. Rees, S. B. Kalindjian, and K. L. Philpott. Principles of early drug discovery. British Journal of Pharmacology, 162:1239–1249, 2011. [6] K. Biber, A. Bhattacharya, B. M. Campbell, J. R. Piro, M. Rohe, R. G. W. Staal, R. V. Talanian, and T. Möller. Microglial drug targets in AD: Opportunities and challenges in drug discovery and development. Frontiers in Pharmacology, 10:840, 2019. [7] M. S. Cohen, T. Holtzman, J. Michel, and J. Marshall. Small molecules as tools for modulating protein activity. Nature Reviews Drug Discovery, 3:927–938, 2004. [8] H. O. Rasul, D. D. Ghafour, B. K. Aziz, B. A. Hassan, T. A. Rashid, and A. Kivrak. Decoding drug discovery: Exploring A-to-Z in silico methods for beginners. Applied Biochemistry and Biotechnology, pages 1–51, 2024. [9] S. Chakraborty, P. Kayastha, and R. Ramakrishnan. The chemical space of b, nsubstituted polycyclic aromatic hydrocarbons: Combinatorial enumeration and highthroughput first-principles modeling. The Journal of Chemical Physics, 150(11), 2019. [10] R. O. Fox, P. A. Evans, and C. M. Dobson. Multiple conformations of a protein demonstrated by magnetization transfer NMR spectroscopy. Nature, 320(6058):192– 194, 1986. 34 [11] A. V. Sadybekov and V. Katritch. Computational approaches streamlining drug discovery. Nature, 616(7958):673–685, 2023. [12] Y. Shimizu, M. Ohta, S. Ishida, et al. AI-driven molecular generation of not-patented pharmaceutical compounds using world open patent data. Journal of Cheminformatics, 15(1):120, 2023. [13] P. J. Guns and J. Joossens. Intellectual property management in academic drug discovery: What are the challenges? Pharmaceutical Patent Analyst, 5(2):83–85, 2016. [14] Y. Hu, D. Stumpfe, and J. Bajorath. Recent advances in scaffold hopping: miniperspective. Journal of Medicinal Chemistry, 60(4):1238–1246, 2017. [15] K. K. Mak, Y. H. Wong, and M. R. Pichika. Artificial intelligence in drug discovery and development. In Drug Discovery and Evaluation: Safety and Pharmacokinetic Assays, pages 1–14. Springer, 2023. [16] D. Rigoni, N. Navarin, and A. Sperduti. Conditional constrained graph variational autoencoders for molecule design. In 2020 IEEE Symposium Series on Computational Intelligence (SSCI), pages 729–736. IEEE, 2020. [17] C. Bilodeau, W. Jin, T. Jaakkola, R. Barzilay, and K. F. Jensen. Generative models for molecular discovery: Recent advances and challenges. WIREs Computational Molecular Science, 2022. [18] D. C. Elton, Z. Boukouvalas, M. D. Fuge, and P. W. Chung. Deep learning for molecular designa review of the state of the art. Molecular Systems Design & Engineering, 4(4):828–849, 2019. [19] D. Weininger. SMILES, a chemical language and information system. 1. introduction to methodology and encoding rules. Journal of Chemical Information and Computer Sciences, 28(1):31–36, 1988. [20] W. Jin, R. Barzilay, and T. Jaakkola. Junction tree variational autoencoder for molecular graph generation. In International Conference on Machine Learning, pages 2323–2332. PMLR, 2018. [21] M. J. Kusner, B. Paige, and J. M. Hernández-Lobato. Grammar variational autoencoder. In International Conference on Machine Learning, pages 1945–1954. PMLR, 2017. [22] N. De Cao and T. Kipf. MolGAN: An implicit generative model for small molecular graphs. ICML 2018 workshop on Theoretical Foundations and Applications of Deep Generative Models, 2018. [23] J. N. Wu, T. Wang, Y. Chen, L. J. Tang, H. L. Wu, and R. Q. Yu. t-SMILES: a fragment-based molecular representation framework for de novo ligand design. Nature Communications, 15(1):4993, 2024. [24] M. Krenn, F. Häse, A. Nigam, P. Friederich, and A. Aspuru-Guzik. Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation. Machine Learning: Science and Technology, 1(4):045024, 2020. 35 [25] H. Huang, L. Sun, B. Du, and W. Lv. Conditional diffusion based on discrete graph structures for molecular graph generation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 4302–4311, 2023. [26] H. Chen, C. Xu, L. Zheng, Q. Zhang, and X. Lin. Diffusion-Based Graph Generative Methods. IEEE Transactions on Knowledge & Data Engineering, 36:7954–7972, 2024. [27] Y. Chen, J. Liu, L. Peng, Y. Wu, Y. Xu, and Z. Zhang. Auto-encoding variational bayes. Cambridge Explorations in Arts and Sciences, 2(1), 2024. [28] J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020. [29] J. Song, C. Meng, and S. Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. [30] J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. Van Den Berg. Structured denoising diffusion models in discrete state-spaces. Advances in Neural Information Processing Systems, 34:17981–17993, 2021. [31] A. Q. Nichol and P. Dhariwal. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, pages 8162–8171. PMLR, 2021. [32] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, . Kaiser, and I. Polosukhin. Attention is all you need. Advances in Neural Information Processing Systems, 2017. [33] A. Bansal, E. Borgnia, H-M Chu, J. Li, H. Kazemi, F. Huang, M. Goldblum, J. Geiping, and T. Goldstein. Cold diffusion: Inverting arbitrary image transforms without noise. Advances in Neural Information Processing Systems, 36, 2024. [34] J. Zhou, G. Cui, S. Hu, Z. Zhang, C. Yang, Z. Liu, L. Wang, C. Li, and M. Sun. Graph neural networks: A review of methods and applications. AI Open, 1:57–81, 2020. [35] J. Gilmer, S. S. Schoenholz, P. F. Riley, O. Vinyals, and G. E. Dahl. Neural message passing for quantum chemistry. In International Conference on Machine Learning, pages 1263–1272. PMLR, 2017. [36] M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. P. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, et al. Scaling vision transformers to 22 billion parameters. In International Conference on Machine Learning, pages 7480–7512. PMLR, 2023. [37] V. P. Dwivedi and X. Bresson. A generalization of transformer networks to graphs. AAAI Workshop on Deep Learning on Graphs: Methods and Applications, 2021. [38] V. P. Dwivedi, A. T. Luu, T. Laurent, Y. Bengio, and X. Bresson. Graph neural networks with learnable structural and positional representations. In International Conference on Learning Representations, 2022. 36 [39] K. Preuer, P. Renz, T. Unterthiner, S. Hochreiter, and G. Klambauer. Fréchet ChemNet distance: a metric for generative models for molecules in drug discovery. Journal of Chemical Information and Modeling, 58(9):1736–1741, 2018. [40] J. J. Irwin, T. Sterling, M. M. Mysinger, E. S. Bolstad, and R. G. Coleman. ZINC: A free tool to discover chemistry for biology. Journal of Chemical Information and Modeling, 52(7):1757–1768, 2012. [41] Y. Wang, S. H. Bryant, T. Cheng, J. Wang, A. Gindulyte, B. A. Shoemaker, P. A. Thiessen, S. He, and J. Zhang. PubChem BioAssay: 2017 update. Nucleic Acids Research, 45(D1):D955–D963, 2016. [42] A. P. Bento, A. Gaulton, A. Hersey, L. J. Bellis, J. Chambers, M. Davies, F. A. Krüger, Y. Light, L. Mak, S. McGlinchey, M. Nowotka, G. Papadatos, R. Santos, and J. P. Overington. The ChEMBL bioactivity database: an update. Nucleic Acids Research, 42(D1):D1083–D1090, 2013. [43] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in Neural Information Processing Systems, 30, 2017. [44] G. R. Bickerton, G. V. Paolini, J. Besnard, Sorel Muresan, and A. L. Hopkins. Quantifying the chemical beauty of drugs. Nature Chemistry, 4(2):90–98, 2012. [45] S. A. Wildman and G. M. Crippen. Prediction of physicochemical parameters by atomic contributions. Journal of Chemical Information and Computer Sciences, 39(5):868–873, 1999. [46] P. Ertl and A. Schuffenhauer. Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. Journal of Cheminformatics, 1(1):8, 2009. [47] R. Ramakrishnan, P. O. Dral, M. Rupp, and O. A. von Lilienfeld. Quantum chemistry structures and properties of 134 kilo molecules. Scientific Data, 1, 2014. [48] C. Niu, Y. Song, J. Song, S. Zhao, A. Grover, and S. Ermon. Permutation invariant graph generation via score-based generative modeling. In International Conference on Artificial Intelligence and Statistics, pages 4474–4484. PMLR, 2020. [49] D. Liu, S. Chen, S. Zheng, S. Zhang, and Y. Yang. SE(3) equivalent graph attention network as an energy-based model for protein side chain conformation. In 2023 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pages 120–123. IEEE, 2023. [50] X. Wang, T. Hu, X. Ren, J. Sun, K. Liu, and M. Zhang. TransVae: A novel variational sequence-to-sequence framework for semi-supervised learning and diversity improvement. In 2021 International Joint Conference on Neural Networks (IJCNN), pages 1–8, 2021. [51] W. Feltus and J. Smith. Ethical and security implications of dual-use technologies in drug discovery. Journal of Biosecurity and Ethics, 5(2):123–134, 2017. 37 A Datasets The QM9 dataset [47] is a widely used molecular dataset for training generative models and evaluating property prediction tasks, with 134000 stable small organic molecules. These molecules are composed just of Carbon (C), Hydrogen (H), Oxygen (O), Nitrogen (N), and Fluorine (F). The QM9 dataset focuses on molecules with up to just nine heavy atoms, representing a relatively simple chemical space with limited structural complexity. On the other hand, both MOSES [4] and ZINC250k [40] are more complex datasets, derived from the ZINC database, a free public resource for ligand discovery containing over 230 million commercially available molecule. ZINC250k contains 250000 available molecules with up to 38 heavy atoms, including Hydrogen (H), Carbon (C), Sulfur (S), Oxygen (O), Nitrogen (N), Chlorine (Cl), Fluorine (F), Bromine (Br), Iodine (I) and Phosphorus (P). MOSES is a bigger subset, containing 1.94 million molecules, filtered using different criteria. These molecules contain up to 27 heavy atoms, including Hydrogen (H), Carbon (C), Sulfur (S), Oxygen (O), Nitrogen (N), Chlorine (Cl), Fluorine (F) and Bromine (Br). This dataset is also filtered so that it does not contain cycles larger with more than 8 atoms. Due to the nature of the datasets, metrics such as QED, logP, and others are only measured for the generated samples from the MOSES and ZINC250k datasets, as both contain drug-like molecules. As a result, measuring these metrics is meaningful for assessing the quality and relevance of the generated molecules in these datasets. However, since QM9 consists of simpler molecules that are not drug-like, these metrics are not relevant or useful for evaluating the model’s performance on this dataset. 38 (a) SA distributions. (b) Molecular weights distributions. Figure 17: SA and molecular weights distributions of the generated molecules with a uniform distribution along with the distributions of the ZINC250k dataset. Some other interesting metrics to meassure the performance of molecule generation are the following ones: •Internal Diversity (IntDiv): Assesses the chemical diversity within the generated set, detecting issues like mode collapse where the model fails to explore the full chemical space. A higher score indicates greater diversity in the generated molecules. •Nearest Neighbor Similarity (SNN): Compares the fingerprints, a representation of molecules using binary vectors, of generated molecules with their nearest neighbors in the training dataset using the Jaccard index. This measures how closely the generated molecules resemble the training data. •Fragment Similarity (Frag): Evaluates how similar the distribution of molecular fragments in the generated set is to the training dataset. Fragments are obtained using BRICS fragmentation, which breaks molecules at chemically valid bonds often used in fragment-based drug discovery. •Scaffold Similarity (Scaf): Compares scaffolds, the core structural frameworks of molecules, between the generated set and the training dataset. Scaffolds represent the molecular backbone after removing functional groups and decorations. The fragment and scaffold similarity metrics both assess molecular similarity by breaking molecules into smaller components, but they differ in their decomposition methods. Fragment similarity uses BRICS fragmentation based on chemically valid bond-breaking rules. This is commonly used in fragment-based drug discovery and the that are broken 45 are often those that can be easily re-formed in chemical synthesis. On the other hand, the scaffold similarity uses the scaffolds of the molecules, which are usually referred as the core structural frameworks of molecules. They are obtained after removing side chains and functional groups. We also calculate the Wasserstein-1 distance (Earth Movers Distance) between the molecular properties of the generated set and the training dataset. All these metrics are only computed for MOSES dataset because the reference values for these metrics in some models are derived from the MOSES paper, and the models themselves are trained on MOSES datasets. The results obtained can be seen in Table 5. It can be seen that the results obtained with our model are similar to the ones obtained with one of the state-of-the-art models. However, as explained before, our model is able to obtain a higher number of novel molecules than this model. Metric CatMol VAE SNN 0.47 0.63 Fragment similarity (Frag) 1.00 0.99 Scaffold similarity (Scaff) 1.00 0.94 Internal diversity 0.53 0.86 Wasserstein-1 distance QED 0.04 0.0006 Wasserstein-1 distance weight 0.09 1.8 Wasserstein-1 distance SA 0.80 0.014 Wasserstein-1 distance logP 0.32 0.023 Table 5: More evaluation metrics for the generated samples for the MOSES dataset using CatMol and VAE. The VAE results have been obtained from [4]. C Configuration Regarding the model architecture, all the feedforward networks (FFN) that appear in it are composed of two linear layers with a ReLU activation function in between. The model’s core parameters are summarized in Table 6. Parameter Value Number of layers 10 Number of heads 8 Hidden dim 128 Out dim 128 Dropout 0.1 Table 6: Model parameters used to train the model for the different datasets. The model contains a total of 370189 trainable parameters. In terms of the diffusion process, we consider a total of T=300 timesteps. The random walk positional encodings are added as learnable features with a dropout factor of 0.7. The optimizer used for training is AdamW, a variant of the Adam optimizer with weight decay. The learning rate has been set to 1e-3 with a weight decay of 0.001 for 46 QM9 and ZINC250k dataset while it has been lowered to 1e-4 for MOSES dataset. D Molecular generation When the generation process ends, two matrices are obtained: a nodes matrix of shape (bs, n, dx), where bs is the number of samples generated, n is the number of atoms, and dx is the number of distinct atom types and an edges matrix, and an edge matrix of shape (bs, n, n, de), where de represents the number of possible bond types. These matrices represent the probabilities of assigning a specific atom type to each node and a bond type to each edge. To generate a valid graph, we first sample the atom types and bond types by selecting the most probable ones. Once the atom types and bond types are sampled, we select the largest connected component from the graph, ensuring that only the primary connected fragment is kept. This is important because, in molecular structures, multiple disconnected fragments would not represent a valid single molecule. Afterwards, the molecule undergoes a sanitization process using RDKit, a widely used open-source toolkit for cheminformatics. This process involves a series of checks and corrections designed to ensure that the generated molecule is chemically valid and conforms to standard chemical principles. This sanitization process is the one that determines whether the generated molecule is valid or not, by verifying its structure and making necessary adjustments. 47 E Examples of generated molecules Figure 18: Examples of generated valid molecules training the model with the QM9 dataset. Figure 19: Examples of generated valid molecules training the model with the ZINC250k dataset. Figure 20: Examples of generated valid molecules training the model with the MOSES dataset. 48 F Equivariance A function fis considered equivariant with respect to a transformation group Gif, for an input xand transformation g∈G, the following condition holds f(g·x) = g·f(x).(1) In the context of machine learning, equivariance refers to the property of a model where its output changes predictably under specific transformations applied to the input. Specifically, when a transformation is applied to the input data, the model’s output should transform in a corresponding manner. Our generative model is node equivariant since both its backbone architecture and loss function are designed to preserve this relationship. 49