scieee AI-readable full text Open interactive document viewer

Ontology pre-training improves prediction of molecular properties

Tumescheit, Charlotte; Glauer, Martin; Flügel, Simon; Larralde, Martin; Neuhaus, Fabian; Mossakowski, Till; Hastings, Janna

Abstract

Presentation given at the Ontologies4Chem Workshop 2025. Abstract: We presented how we use ontologies to improve Machine Learning models using a method we call “Ontology Pre-Training” (https://www.biorxiv.org/content/10.1101/2025.09.30.679573v1). We transform the ChEBI hierarchy into a classification problem and add it as an additional training step of a chemical foundation model. We show that this method improves performance for many downstream applications of small molecules, such as toxicity, aqueous solubility, and lipophilicity prediction.

Full text

Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Ontology pre-training improves prediction of molecular properties Ontologies4Chem 2025 Charlotte Tumescheit, Martin Glauer, Simon Flügel, Martin Larralde, Fabian Neuhaus, Till Mossakowski, Janna Hastings 12 Nov 2025 Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Metabolism is at the centre of health 23.07.25Institute for Implementation Science in Health Care However, it is a complex system we do not fully understand yet, and existing models and datasets have many limitations 2 Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Understand biological properties of molecules better 23.07.25Institute for Implementation Science in Health Care Interesting properties ―Aqueous solubility ―Toxicity ―Blood-brain-barrier permeability ― … Interesting applications ―De novo characterisation of molecules ―Pharma (Drug Design) ―Healthy longevity ― … 3 Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Understand biological properties of molecules better 23.07.25Institute for Implementation Science in Health Care Interesting properties ―Aqueous solubility ―Toxicity ―Blood-brain-barrier permeability ― … Interesting applications ―De novo characterisation of molecules ―Pharma (Drug Design) ―Healthy longevity ― … 4 Molecular property prediction Er st el lt du rc h Vor l agen ba uer. ch University of Zurich AI + prior knowledge to understand biological properties of molecules better 23.07.25Institute for Implementation Science in Health Care Machine learning models ―Already used for property prediction ―Can inform about newly discovered/designed molecules 5 Er st el lt du rc h Vor l agen ba uer. ch University of Zurich AI + prior knowledge to understand biological properties of molecules better 23.07.25Institute for Implementation Science in Health Care Machine learning models ―Already used for property prediction ―Can inform about newly discovered/designed molecules Prior knowledge ―Available in the form of e.g. ontologies ―Underutilised in ML models 6 Er st el lt du rc h Vor l agen ba uer. ch University of Zurich AI + prior knowledge to understand biological properties of molecules better 23.07.25Institute for Implementation Science in Health Care Machine learning models ―Already used for property prediction ―Can inform about newly discovered/designed molecules Prior knowledge ―Available in the form of e.g. ontologies ―Underutilised in ML models Incorporating ontologies into ML models ―How to make use of prior knowledge in ML models? 7 Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Different forms of Knowledge Injection 23.07.25Institute for Implementation Science in Health Care 8 Decoding the interactions and functions of non-coding RNA with artificial intelligence, https://doi.org/10.1038/s41580-025-00857-w Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Training a language model 23.07.25Institute for Implementation Science in Health Care 9 Classical Pre-Training - Self-supervised - High amount of general data Fine-Tuning -Supervised - Specialised data Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Ontology-derived knowledge-injection: Ontology Pre-Training 23.07.25Institute for Implementation Science in Health Care ―Goal: enable a model to benefit from the knowledge represented in an ontology by training it to categorise its input according to the class hierarchy from the ontology ―Turn the subsumption hierarchy into classification problem ―Consider classes with a SMILES attached ―Extracted all sub-classes of molecular entity with at least 100 sub-classes ―Result: A multilabel classification problem with 997 classes for over 185,000 SMILES 16 doi.org/10.3233/SW-2331 Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Ontology-derived knowledge-injection: Turn hierarchical structure of ChEBI into a multilabel classification 23.07.25Institute for Implementation Science in Health Care 17 Molecule Purine Hydrocarbon Alkane Pyrimidine Caffeine Heterocyclic compound Molecule (e.g. Caffeine) ChEBI classification Multilabel vector … Molecule Purine Caffeine Alkane Pyrimidine Heterocyclic compound Hydrocarbon Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Excluding rare classes 23.07.25Institute for Implementation Science in Health Care 18 Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Model setup 23.07.25Institute for Implementation Science in Health Care 19 ChEB-AI ―Python library that supports preprocessing of chemical data ―Includes datasets extracted from ChEBI and PubChem ―Provides a training environment for a variety of models, including transformers ―Also includes Chebifier ChEB-AI GitHub https://github.com/ChEB-AI Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Training Pipeline for a Transformer-based Model 23.07.25Institute for Implementation Science in Health Care SMILES - Simplified Molecular Input Line Entry System, BCE –Binary Cross Entropy, MSE –Mean Square Error Standard Pre-Training ―Dataset: 200,000 molecules from PubChem ―Input: SMILES ―Training: Self-supervised (from scratch) ―Loss: Cross entropy ―Aim: model learns the structure of a molecule 20 Fine-Tuning ―Datasets: depending on task ―Input: SMILES ―Output: depending on task ―Training: depending on task ―Loss: Binary Cross Entropy (classification), Mean Square Error (regression) ―Aim: specialising for downstream task Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Training Pipeline for a Transformer-based Model 23.07.25Institute for Implementation Science in Health Care SMILES - Simplified Molecular Input Line Entry System, BCE –Binary Cross Entropy, MSE –Mean Square Error Standard Pre-Training ―Dataset: 200,000 molecules from PubChem ―Input: SMILES ―Training: Self-supervised (from scratch) ―Loss: Cross entropy ―Aim: model learns the structure of a molecule Fine-Tuning ―Datasets: depending on task ―Input: SMILES ―Output: depending on task ―Training: depending on task ―Loss: Binary Cross Entropy (classification), Mean Square Error (regression) ―Aim: specialising for downstream task Ontology Pre-Training ―Dataset: ChEBI ontology ―Input: SMILES ―Output: ChEBI classes ―Training: Multilabel classification ―Loss: Binary Cross Entropy ―Aim: model learns semantics of tokens based on the class hierarchy of ChEBI 21 Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Our experiment workflow 23.07.25Institute for Implementation Science in Health Care 22 Unlabelled molecules (PubChem) Ontology PreTraining Standard PreTraining Molecular Dataset Standard PreTrained Model Standard + Ontology PreTrained Model Ontology (ChEBI) Fine-Tuning Fine-Tuning Molecular Property Prediction without Ontology Molecular Property Prediction with Ontology Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Fine-tuning on various datasets of different types, sizes, … 23.07.25Institute for Implementation Science in Health Care 23 DOI: MoleculeNet - 10.1039/c7sc02664a, Extended Solubility Dataset - 10.1038/s41597-022-01154-3 Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Results 23.07.25Institute for Implementation Science in Health Care 24 Dataset Without OPT With OPT Metric ESOL 0.817 0.863 R 2 Solubility 0.821 0.859 R 2 FreeSolv 0.841 0.880 R 2 Lipophilicity 0.403 0.591 R 2 BACE 0.805 0.783 F1 BBBP 0.939 0.936 F1 SIDER 0.606 0.601 M icro-F1 ClinTox 0.875 0.891 M icro-F1 Tox21 0.370 0.403 M icro-F1 Test set results R2development during training on validation set per epoch for Lipophilicity With Ontology PreTraining Without Ontology PreTraining Er st el lt du rc h Vor l agen ba uer. ch University of Zurich Also improves explainability 23.07.25Institute for Implementation Science in Health Care 25 DOI: 10.1007/978-3-031-42608-7_4 With Ontology PreTraining Without Ontology PreTraining