scieee AI-readable full text Open interactive document viewer

MORE-Q, a dataset for molecular olfactorial receptor engineering by quantum mechanics

Li, Chen; Medrano Sandonas, Leonardo Rafael; Traber, Philipp; Dianat, Arezoo; Tverdokhleb, Nina; Hurevich, Mattan; Yitzchaik, Shlomo; Rafael, Gutierrez; Croy, Alexander; Cuniberti, Gianaurelio (Giovanni)

Abstract

We introduce the MORE-Q dataset, a quantum-mechanical (QM) dataset encompassing the structural and electronic data of non-covalent molecular sensors formed by combining 18 mucin-derived olfactorial receptors with 102 body odor volatilome (BOV) molecules. To have a better understanding of their intra- and inter-molecular interactions, we have performed accurate QM calculations in different stages of the sensor design and, accordingly, MORE-Q splits into three subsets: i) MORE-Q-G1: QM data of 18 receptors and 102 BOV molecules, ii) MORE-Q-G2: QM data of 23,838 BOV-receptor configurations, and iii) MORE-Q-G3: QM data of 1,836 BOV-receptor-graphene systems. Each subset involves geometries optimized using GFN2-xTB with D4 dispersion correction and up to 39 physicochemical properties, including global and local properties as well as binding features, all computed at the tightly converged PBE+D3 level of theory. By addressing BOV-receptor-graphene systems from a QM perspective, MORE-Q can serve as a benchmark dataset for state-of-the-art machine learning methods developed to predict binding features. This, in turn, can provide valuable insights for developing the next-generation mucin-derived olfactory receptor sensing devices.

Full text

1 Scientific Data | (2025) 12:324 | https://doi.org/10.1038/s41597-025-04616-6 www.nature.com/scientificdata MORE-Q, a dataset for molecular olfactorial receptor engineering by quantum mechanics Li Chen1, Leonardo Medrano Sandonas 1 ✉ , Philipp traber2, arezoo Dianat1, Nina tverdokhleb1, Mattan Hurevich3, Shlomo Yitzchaik3, Rafael Gutierrez 1, alexander Croy 2 ✉ & Gianaurelio Cuniberti 1,4 ✉ We introduce the MORE-Q dataset, a quantum-mechanical (QM) dataset encompassing the structural and electronic data of non-covalent molecular sensors formed by combining 18 mucin-derived olfactorial receptors with 102 body odor volatilome (BOV) molecules. To have a better understanding of their intraand inter-molecular interactions, we have performed accurate QM calculations in different stages of the sensor design and, accordingly, MORE-Q splits into three subsets: i) MOREQ-G1: QM data of 18 receptors and 102 BOV molecules, ii) MORE-Q-G2: QM data of 23,838 BOVreceptor configurations, and iii) MORE-Q-G3: QM data of 1,836 BOV-receptor-graphene systems. Each subset involves geometries optimized using GFN2-xTB with D4 dispersion correction and up to 39 physicochemical properties, including global and local properties as well as binding features, all computed at the tightly converged PBE+D3 level of theory. By addressing BOV-receptor-graphene systems from a QM perspective, MORE-Q can serve as a benchmark dataset for state-of-the-art machine learning methods developed to predict binding features. This, in turn, can provide valuable insights for developing the next-generation mucin-derived olfactory receptor sensing devices. Background & Summary Introduction. Nowadays, the increasing progress in artificial intelligence (AI) has boosted the development of AI-based technologies for the recognition of objects, faces, voices, and touch1. However, a significant gap remains in technologies capable of interpreting and predicting the chemical environment around us. In this sense, tailored electronic noses have recently emerged and are already capable of detecting, for example, volatile organic compounds (VOCs)2,3. The detection of VOCs emitted from the human body4–6, also known as body odor volatilomes (BOVs), can be used as characteristic fingerprints and have significant applications in healthcare7. The constituents of BOVs generally indicate the metabolic state of a person and are promising candidates for medical biomarkers in diagnosing a range of diseases5,8–13, e.g. Alzheimer14,15 and Parkinson14–17. A recent report documented that a ’super smeller’ could detect and distinguish the BOVs associated with Parkinson’s disease, emitted from sebum, from those of normal skin18 indicating the powerful body odor perception ability enabled by the olfactory system. Given these facts, the demand for fast and robust sensing materials for detecting BOV molecules remains consistently strong, particularly in medical diagnostics. It is known that odor perception begins in the nose, where tens of thousands of odorants can be detected by the olfactorial receptors, which are composed of a glycoprotein layer (mucin) that covers the epithelium of olfactory and respiratory systems19–21. While there have been considerable efforts in investigating single odor (BOV) molecules22–30 and molecular receptors31–34, less information is available to accurately describe the physical and chemical interactions in BOV-receptor systems. Characterizing these interactions will deepen our understanding of the biomimetic olfactory system, paving the way for the rational design of receptors tailored for sensing applications35. To address this challenge, initial datasets of BOV-receptor systems have been developed36–40 (see Table1). For 1institute for Materials Science and Max Bergmann center for Biomaterials, tUD Dresden University of technology, 01062, Dresden, Germany. 2Institute of Physical Chemistry, Friedrich Schiller University Jena, 07737, Jena, Germany. 3Institute of Chemistry and Center of Nanotechnology, The Hebrew University of Jerusalem, Jerusalem, 91904, israel. 4Dresden center for computational Materials Science (DcMS), tUD Dresden University of technology, 01062, Dresden, Germany. ✉e-mail: [email protected]; alexander[email protected]; gianaurelio. [email protected] Data DESCRiPtOR OPEN 2 Scientific Data | (2025) 12:324 | https://doi.org/10.1038/s41597-025-04616-6 www.nature.com/scientificdata www.nature.com/scientificdata/ instance, Mainland et al.38 provided the in vitro response of 73 odorants against a clone library of 511 human olfactorial receptors. Sharma et al.39 also developed the online platform OlfactionBase, which provides chemoinformatic properties (e.g. drug-likeness, pharmacokinetic profile, molecular weight, Plog ) and odorant-receptor match information for 875 systems. In a more recent study, Lalis et al.40 developed M2OR database of 51,395 odorant-receptor systems for understanding the molecular mechanisms of olfaction. Numerous similar datasets have been published regarding odorant-receptor interactions and match information37,41–43. However, to the best of our knowledge, none of these datasets account for a quantum-mechanical (QM) treatment of physical and chemical interactions in odorant-receptor systems. This is of great relevance because accurately describing intermolecular interactions – such as van der Waals forces, hydrogen bonds, and dipole-dipole interactions – is essential for determining the correct docking configuration and binding features. Indeed, recent efforts have focused on developing QM datasets to understand the structure-property and property-property relationships in both small and large molecular systems44–49. Additionally, a few QM datasets for small molecular dimers have been generated, where the relevant property is solely the interaction energy50–53. Another open challenge in this field is the accurate investigation of the interaction of odorant-receptor systems on a substrate. Carbon-based materials such as graphene are promising sensing materials due to their high charge mobility and favorable surface-to-volume ratio54–58. Hence, it is essential to have a dataset that provides a QM description of both the structural and electronic properties of odorant-receptor systems, as well as their interaction with a sensing material. To address these challenges, we introduce the MORE-Q dataset, which provides an extensive set of QM properties to accurately investigate the BOV-Receptor systems on a graphene surface, see Fig.1. We have modeled 18 mucin-derived receptors and their combinations with 102 relevant skin BOV molecules6—both systems containing heavy atoms C, N, O, and S. The number of atoms of the receptor molecules ranges from 37 to 102 atoms, while for BOV molecules varies from 7 to 53 atoms. We have followed an exhaustive and systematic procedure to generate the final complex systems (i.e. BOV-receptor-graphene system) and compute the QM properties, resulting in the creation of three MORE-Q subsets: i) MORE-Q-G1 contains the QM data of isolated single BOV and receptor molecules, ii) MORE-Q-G2 contains the QM data of diverse configurations of the BOV-receptor systems, and iii) MORE-Q-G3 contains the QM data of dimers deposited on the graphene surface (see Fig.2). The initial BOV and receptor molecules were optimized using the semi-empirical method GFN2-xTB that considers Dataset Receptor source Total dimer configurations QM properties Surface interaction Mainland et al.38 Human-cloned 37,303 No No OlfactionBase39 Human/mouse 875 No No M2OR40 Mammals 51,415 No No OlfactionDB37 Human/mouse ~400 No No MORE-Q-G2 Mucin-derived 23,838 Yes No MORE-Q-G3 Mucin-derived 1,836 Yes Yes Table 1. Main characteristics of recent publicly available olfactory receptor datasets. Note that among these, the MORE-Q dataset uniquely includes quantum-mechanical (QM) properties of the systems under study and is the only dataset that considers the interaction of BOV-receptor systems with a surface. Fig. 1 Graphical representation of the motivation for developing MORE-Q dataset (Molecular Olfactorial Receptor Engineering by Quantum mechanics). Bio-electronic noses (top right panel) are designed as an electronic equivalent to the olfactory system (top left panel), e.g. for sensing body odor volatilomes (BOV). The MORE-Q dataset offers a comprehensive collection of quantum-mechanical properties and structural data that accurately describe intraand intermolecular interactions in molecular sensors, see lower panel. 3 Scientific Data | (2025) 12:324 | https://doi.org/10.1038/s41597-025-04616-6 www.nature.com/scientificdata www.nature.com/scientificdata/ D4 dispersion correction59. Then, the configurations of the 1,836 BOV-receptor systems were screened by molecular docking using the automated Interaction Site Screening (aISS)60 submodule of the xTB code, applying the same level of theory as used in the geometry optimizations. Here, the hierarchical clustering method was used to filter similar geometries for each BOV-receptor combination, resulting in 23,838 dimer configurations. To generate the complex systems, we have considered the most energy-favorable dimer configuration for each unique combination (ranked by the interaction energy Eint) and, then, deposited it on a graphene layer. Finally, each MORE-Q subset includes up to 39 physicochemical properties, encompassing global (molecular) and local (atom-in-a-molecule) properties, as well as binding features. All these properties were computed using the tightly converged PBE+D3 level of theory. As such, the MORE-Q dataset provides a comprehensive set of QM structural and property data for BOV-receptor and BOV-receptor-graphene systems, which can further enhance our understanding and the prediction of the performance of molecular sensors for digital olfaction. Key advancements. The MORE-Q dataset series aims to provide accurate QM data to gain insights into the interaction between BOV molecules and molecular receptors, as well as the effects of substrate deposition on binding features. To achieve this objective, we have ensured that MORE-Q includes the following attributes: Fig. 2 Schematic description of generation procedure of the MORE-Q dataset. MORE-Q is split into three subsets depending on the generation stage: MORE-Q-G1, MORE-Q-G2, MORE-Q-G3. In MORE-Q-G1, We first established a skin body odor set with 102 molecules from the intersection of two body odor resources such as those presented in Drabinska et al.6 and Keller et al.64. The structure of isolated BOV molecules and 18 receptor-graphene systems were optimized using GFN2-xTB with D4 dispersion correction. Physicochemical properties were posteriorly computed at the tightly converged PBE+D3 level of theory. For generating MORE-Q-G2, the conformational search of the BOV-receptor systems from MORE-Q-G1 was carried out using the docking program aISS code which employs xTB-IFF72 force field and a follow-up genetic algorithm, resulting in a total of 83,916 BOV-receptor dimer configurations. Following an RMSD-based hierarchical clustering method, 23,838 non-redundant configurations were selected and their properties were calculated at GFN-xTB+D4 theory level. The most energy-favorable 1,836 configurations were further selected, and their corresponding dimer properties were calculated both at the higher PBE+D3 theory level. In MORE-Q-G3, the selected 1,836 configurations from MORE-Q-G2 were optimized back on the graphene surface, forming the complex BOV-receptor-graphene systems (CPLX). The substrate (SUB) and BOV molecule (OM) systems were constructed by removing the BOV molecules and receptor+graphene, respectively. Next, PBE+D3 level simulations were conducted on the geometries of 1,836 CPLX, SUB, and OM systems, obtaining global and local properties, as well as the binding features of these systems. See “Methods" for more details. 4 Scientific Data | (2025) 12:324 | https://doi.org/10.1038/s41597-025-04616-6 www.nature.com/scientificdata www.nature.com/scientificdata/ • The MORE-Q dataset comprises the QM structural and property data of 18 newly synthesized mucin-derived receptors32,61,62 and 102 skin molecules, carefully selected from an extensive pool of 2,746 BOV molecules, representing a large swath of the chemical space of BOV molecules. • We have exhaustively explored the potential energy surface (PES) of BOV-receptor systems by using the automated interaction site screening method in xTB code60, at the GFN2-xTB level of theory59 with D4 dispersion correction63.This procedure allowed us to determine the 1,836 most energetic-favorable BOV-receptor dimer configurations. • The MORE-Q dataset provides extensive sets of QM global and local properties (up to 39) for single BOV/ receptor molecules (MORE-Q-G1), molecular dimers (MORE-Q-G2), and complex systems (MORE-Q-G3), offering more comprehensive data compared to all other aforementioned BOV-receptor interaction datasets. These properties can be utilized in machine learning (ML) methods (e.g. as QM descriptors) to uncover and elucidate structure-property and property-property relationships between the building blocks, as well as the entire molecular sensor. • The MORE-Q-G3 dataset involves complex electronic structure calculations, such as the determination of the work function (WF) φ, which can be related to the sensing response of molecular devices. This property offers a more realistic validation of the ML models trained on our QM data, intending to design novel molecular sensors. Methods BOV and receptor molecules. The selection of the initial set of BOV molecules dataset was based on two main works6,64. In the first one, Drabińska et al.6 reports a meta-analysis of the available literature on chemical substances that have been documented in the human body. Here, 2,752 molecules are categorized into feces, urine, breath, skin, milk, blood, saliva, and semen. Since a central aspect of digital olfaction is the perception of odor molecules, we have also examined the work done by Keller et al.64, and consider 480 BOV molecules with available perceptions from their dataset. Then, by intersecting both datasets, we ended up with 102 skin-related molecules that include the heavy elements C, S, O, and N (see Fig.2). Their isomeric SMILE strings and the corresponding initial geometries were extracted from the large public molecular database PubChem65. The molecular size of the BOV molecules varies from 7 to 53 atoms. The two-dimensional chemical representation of these molecules has been plotted in Fig.S1 of the Supporting Information (SI). Furthermore, we have modeled newly-synthesised 18 bio-inspired (mucin-derived) receptors32,61,62 that are composed of glycan modified by aromatic decoration for surface adhesion. D-galactose, one of the most common glycans in the extracellular matrix, was used as a scaffold for aromatic decorated monosaccharide receptor’s library, which is obtained by a multistep chemical synthesis62 of monosaccharides. The synthetic approach provides the ability to install specific groups of various natures on the monosaccharide thereby enabling tuning the receptor’s affinity toward odorants. Using this ability we control the rigidity, hydrophobicity, and polarizability of glycan-based receptors. The molecular size of the receptors ranges from 37 to 102 atoms, including the heavy atoms C, N, O, and S. A detailed description and full characterization can be found in Refs. 32,62. We show the two-dimensional chemical structures of the 18 receptors in Fig.S2 of theSI. MORE-Q-G1 dataset generation: properties of monomers. The MORE-Q-G1 dataset contains the quantum-mechanical (QM) properties of the optimized structures of 102 BOV molecules and 18 molecular receptors. The structures of each BOV molecule were first optimized using the semi-empirical method GFN2-xTB that considers D4 dispersion correction as it is implemented in the xTB packages (version 6.6.0)59. A stringent convergence criterion for energies and gradient norms was set to 5×10−8 Eh and Ea510 h 501 ×⋅ −− , respectively. For the molecular receptors, we deposited them directly on a 10×10 graphene layer containing 200 C atoms with the fixed vacuum layer z=50.68 Å in a periodic simulation box. The geometry optimization of the receptor-graphene systems (referred to as the ’SUB’ system in other sections) was conducted using DFTB+ package66,67 and considering the GFN2-xTB Hamiltonian with D4 dispersion correction. The thresholds for SCC convergence and the maximal atomic force were set to 1⋅10−5 and Ea110 h 401 ⋅⋅ −− , respectively. The Fermi smearing in the optimization was set to 300 K and simulation was conducted at Gamma point. During the optimization, lattice vector angles and slab thickness were fixed. For the initial adsorption of each receptor on graphene, configurations were set to have maximal π−π stacking between the pyrene rings of the receptors and the graphene layer. Notice that, instead of optimizing the receptors as an isolated system, we directly optimized them within the periodic graphene system to avoid self π−π stacking in the isolated state, which would impede stable adsorption on the graphene surface. Calculation of physicochemical properties. We have computed 39 QM global and local properties of the optimized structures of BOV and receptor molecules, see Table2. To do this, single-point calculations were conducted employing density-functional theory (DFT) at the PBE level with def2-TZVPP basis set and D3 dispersion correction, as implemented in the ORCA software (version 5.0.3)68. Energy components, orbital energies, C6 dispersion coefficient, atomic forces, dipole moment, and quadrupole moment were extracted from the output files of the self-consistency (SCF) calculations. The molecular isotropic polarizability and the polarizability tensor were analytically calculated through the coupled-perturbed SCF equations (CP-SCF). The radius of gyration was calculated per each structure by R g mr m ii i 2 = Σ⋅ Σ, where mi and ri are the ith atom mass and the corresponding distance to the molecular center of mass, respectively. The atomic charges were computed by performing Mulliken69, Loewdin70 and Mayer71 population analysis. 5 Scientific Data | (2025) 12:324 | https://doi.org/10.1038/s41597-025-04616-6 www.nature.com/scientificdata www.nature.com/scientificdata/ MORE-Q-G2 dataset generation: properties of the molecular dimers. The MORE-Q-G2 dataset contains the QM properties of the optimized structures of the 23,838 BOV-receptor systems (referred to as the dimer system). The configurations of the initial 1,836 BOV-receptor systems (combination of 18 molecular receptors with 102 BOV molecules from MORE-Q-G1) were screened by molecular docking using the automated Interaction Site Screening (aISS)60 submodule of the xTB packages (version 6.6.0), where the GFN2-xTB parameterization and D4 dispersion correction were implemented. The aISS module prescreens potential docking sites (pockets, stack, and angular search) on the receptor, followed by a genetic optimization for stack and angular search, where the intermediate binding energies are evaluated using the xTB-IFF force field72. In the end, the updated dimer structures were optimized again by the GFN2-xTB method. To keep the same structural conformation of the receptor adsorbed on the graphene layer, we have fixed the geometry of the receptor during the docking process. Then, 100 configurations were generated per molecular dimer. We subsequently sent the dimer configurations back to the graphene surface and excluded those configurations in which any atom of the BOV molecule was located between the receptor and the graphene layer, resulting in a subset of 83,916 dimer configurations. It is worth mentioning that the atomic coordinates of the receptor were mapped exactly to the previous #Property Symbol Unit Dimension Type HDF5 keys 1Atomic numbers — — N A ‘atNUM’ 2Atomic positions —Å3N A ‘atXYZ’ 3Total PBE+D3 energy Etot eV 1 M ‘ePBE+D3’ 4Nuclear repulsion energy Enuc eV 1 M ‘eNUC’ 5Electronic repulsion energy Eele eV 1 M ‘eELE’ 6One electron energy E1e eV 1 M ‘e1E’ 7Two electron energy E2e eV 1 M ‘e2E’ 8Virial potential energy Epe eV 1 M ‘ePE’ 9Virial kinetic energy Eke eV 1 M ‘eKE’ 10 Exchange energy ExeV 1 M ‘eX’ 11 Correlation energy EceV 1 M ‘eC’ 12 Exchange-correlation energy Exc eV 1 M ‘eXC’ 13 Total D3 energy ED3 eV 1 M ‘eD3’ 14 Dispersion E6 energy E6eV 1 M ‘eE6’ 15 Dispersion E8 energy E8eV 1 M ‘eE8’ 16 HOMO energy EHOMO eV 1 M ‘eH’ 17 LUMO energy ELUMO eV 1 M ‘eL’ 18 HOMO-LUMO gap EGAP eV 1 M ‘HLgap’ 19 Orbital energies Eoe eV *M ‘eORB’ 20 Isotropic molecular C6 coefficient C6 ⋅Ea h0 6 1 M ‘mC6’ 21 Electronic dipole moment μelc D 3 M ‘vEDIP’ 22 Nuclear dipole moment μnuc D 3 M ‘vNDIP’ 23 Total dipole moment μD 3 M ‘vDIP’ 24 Scalar total dipole moment μsD 1 M ‘DIP’ 25 Rotational spectrum constant BMHz 3 M ‘vRS’ 26 Rotational dipole moment μBd 3 M ‘vRSDIP’ 27 Nuclear quadrupole moment tensor Qnuc ⋅ ea 0 2 6 M ‘NQP’ 28 Electronic quadrupole moment tensor Qele ⋅ea 0 2 6 M ‘EQP’ 29 Total quadrupole moment tensor QBuckingham 6 M ‘TQP’ 30 Isotropic molecular quadrupole QsBuckingham 1 M ‘mQP’ 31 Molecular polarizabillity tensor α a0 3 6 M ‘mTPOL’ 32 Molecular isotropic polarizability αs a0 3 1 M ‘mPOL’ 33 Radius of gyration RgÅ1 M ‘RG’ 34 Inertia moment tensor ITS amu⋅Å26 M ‘IM’ 35 Mulliken atomic charge qmu eN A ‘muCHG’ 36 Loewdin atomic charge qlo eN A ‘loCHG’ 37 Mayer atomic charge qma eN A ‘maCHG’ 38 Atomic forces Fat eV/Å3N A ‘vF’ 39 Atomisation energy Eat eV 1 M ‘eAT’ Table 2. List of physicochemical properties of BOV and molecular receptors contained in MORE-Q-G1 subset. Each property presents a name, symbol, unit, dimension, type, and corresponding key in the HDF5 file. Property types are categorized into atomic (A) and molecular (M). Ehand a0 refer to the atomic unit of Hatree and Bohr radius. * The number of orbital energies varies for each molecule. 6 Scientific Data | (2025) 12:324 | https://doi.org/10.1038/s41597-025-04616-6 www.nature.com/scientificdata www.nature.com/scientificdata/ configuration on graphene from the MORE-Q-G1 dataset. As a final step, the hierarchical clustering method was used to filter similar geometries for each BOV-receptor combination, reducing the number of configurations up to 23,838. This step was carried out by computing the root-mean-square deviation (RMSD) among molecular structures. The detailed clustering process is discussed in Sec.2 of theSI. Calculation of physicochemical properties. We have first computed 24 QM global and local properties of the optimized structures of 23,838 BOV-receptor systems, see Table3. These properties were obtained from the output files of a follow-up single-point calculation using GFN2-xTB, which considers D4 dispersion correction. The SCC convergence for these calculations was set to 1⋅10−6 Eh. To name a few properties, we have energy components, orbital energies, C6 and C8 dispersion coefficients, atomic polarizabilities, dipole moment, quadrupole moment, and binding energy. Moreover, we have selected the most energy-favorable configuration for each dimer (ranked by the binding energy Eint) and computed QM properties at the PBE+D3 level with def2-TZVPP basis set, as was previously done for the generation of MORE-Q-G1. Accordingly, MORE-Q-G2 also contains the 39 QM global and local properties listed in Table2 and Eint for 1,836 dimers at PBE+D3 level. MORE-Q-G3 dataset generation: properties of the complex systems. The MORE-Q-G3 dataset contains the QM properties of the optimized structures of 1,836 BOV-receptor-graphene systems (referred to as the complex (CPLX) system). To generate MORE-Q-G3, we have considered the most energy-favorable dimer configuration for each dimer (ranked by Eint from MORE-Q-G2) and, then, mapped it back to the graphene layer. Next, the CPLX systems underwent geometry optimization using the DFTB+ package, employing the GFN2-xTB Hamiltonian with D4 dispersion correction for the SUB system. We here chose to fix the atomic positions in the graphene layer, as the adsorption of BOV molecules will not significantly affect them. The resulting 1,836 CPLX systems were then split into SUB systems and BOV molecules (OM) to compute binding features. Calculation of physicochemical properties. We have first computed 20 QM global and local properties of the optimized structures of 1,836 CPLX systems, 1,836 SUB systems, and 1,836 OM systems, see the top panel in Table4. In doing so, single-point calculations of these systems were conducted at tightly converged PBE+D3 theory level by Vienna ab initio simulation package73,74 (VASP, version 6.3.1). The energy cutoff for #Property Symbol Unit Dimension Type HDF5 keys 1Atomic number — — N A ‘atNUM’ 2Atomic positions —Å3N A ‘atXYZ’ 3Total GFN2-xTB+D4 energy Etot eV 1 M ‘eXTB+D4’ 4Repulsion energy Erep eV 1 M ‘eREP’ 5SCC total energy Escc eV 1 M ‘eSCC’ 6Isotropic electrostatic energy Eiel eV 1 M ‘eIE’ 7Anisotropic electrostatic energy Eael eV 1 M ‘eAE’ 8Anisotropic exchange-correlation energy Eaxc eV 1 M ‘eAXC’ 9D4 dispersion energy ED4 eV 1 M ‘eD4’ 10 HOMO energy EHOMO eV 1 M ‘eH’ 11 LUMO energy ELUMO eV 1 M ‘eL’ 12 HOMO-LUMO gap EGAP eV 1 M ‘HLgap’ 13 Orbital energies Eoe eV *M ‘eORB’ 14 Atomisation energy Eat eV 1 M ‘eAT’ 15 Atomic coordination number Nac — N A ‘ACN’ 16 Atomic Mulliken charge qmu eN A ‘muCHG’ 17 Atomic C6 dispersion coefficient C6,at ⋅Ea h0 6 N A ‘atC6’ 18 Atomic polarizability αat a0 3 N A ‘atPOL’ 19 Isotropic C6 dispersion coefficient C6 Ea h0 6 ⋅ 1 M ‘mC6’ 20 Isotropic C8 dispersion coefficient C8 ⋅Ea h0 8 1 M ‘mC8’ 21 Isotropic molecular polarizability αs a0 3 1 M ‘mPOL’ 22 Dipole moment μe⋅a03 M ‘vDIP’ 23 Scalar total dipole moment μsD 1 M ‘DIP’ 24 Molecular quadrupole tensor QP ea 0 2 ⋅ 6 M ‘QP’ 25 Binding energy Eint eV 1 M,BD ‘eBIND’ Table 3. List of physicochemical properties of molecular dimers at GFN2-xTB+D4 theory level contained in MORE-Q-G2 subset. Each property presents a name, symbol, unit, dimension, type, and corresponding key in the HDF5 file. Property types are categorized into atomic (A) and molecular (M). BD stands for the binding feature. Eh and a0 refer to the atomic unit of Hatree and Bohr radius. For the most stable dimer conformation, we have also computed the same properties as listed in Table2 at PBE+D3 level of theory, including the binding energy. *The number of orbital energies varies for each molecule. 7 Scientific Data | (2025) 12:324 | https://doi.org/10.1038/s41597-025-04616-6 www.nature.com/scientificdata www.nature.com/scientificdata/ the plane-wave basis set and the SCF convergence threshold were set to 600 and 1⋅10−5 eV, respectively. And all simulations were conducted at Gamma point. The dipole correction along the slab direction (50.68 Å) was switched on to obtain flat electrostatic potential in the slab. Energy components, orbital energies, atomic forces, stress tensor, and electrostatic potential were extracted from the OUTCAR output file. Bader atomic charges q were obtained by postprocessing the information obtained from Bader charge analysis75. The work function of the CPLX and SUB systems were calculated as follows: EE,(1) VF φ=− where EF is the Fermi level and EV is the vacuum energy. And the work function change is defined as the work function difference after and before the BOV adsorption i.e. CPLX and SUB systems: (2) CPLX SUB φφ φΔ= −. EV is obtained by analyzing the flattened region of the electrostatic potential P(z) along the slab direction. P(z) is computed by the following equation: Pz nzdz() () , (3) ∫ = where the planar averaged charge density n(z) is defined as: = nz Anxyzdxdy () 1/ (, ,) (4 ) ∬ and the A denotes the surface area of the cell. #Property Symbol Unit Dimension Type HDF5 keys 1Atomic number — — N A,S ‘atNUM’ 2Atomic coordinates —Å3N A,S ‘atXYZ’ 3Total PBE+D3 energy Etot eV 1 G,S ‘ePBE+D3’ 4Fermi energy EFeV 1 G,S ‘eFE’ 5E6 dispersion energy E6eV 1 G,S ‘eE6’ 6E8 dispersion energy E8eV 1 G,S ‘eE8’ 7Total dispersion energy ED3 eV 1 G,S ‘eD3’ 8Valence band maximum Evbm eV 1 G,S ‘eVBM’ 9Conduction band minimum Ecbm eV 1 G,S ‘eCBM’ 10 VBM-CBM gap Egap eV 1 G,S ‘eGAP’ 11 Band energies Ebe eV *G,S ‘eBE’ 12 Work function φeV 1 G,S ‘WF’ 13 Planar-averaged potential z distance zÅ * G,S ‘zEPOL’ 14 Planar-averaged potential Pavg eV *G,S ‘eEPOL’ 15 Cell parameters lÅ9 G,S ‘CELL’ 16 Cell stress tensor σkB 6 G,S ‘stCELL’ 17 External cell pressure Pcl kB 1 G,S ‘pCELL’ 18 Atomic forces Fat eV/Å3N A,S ‘vF’ 19 Total drift Fdf eV/Å3 A,S ‘vDF’ 20 Bader atomic charge q e N A,S ‘baCHG’ 1Adsorption energy Eads eV 1 G,BD ‘eADS’ 2Graphene Bader charge change ΔQGR e1 G,BD ‘GbaDCHG’ 3Receptor Bader charge change ΔQrec e1 G,BD ‘RbaDCHG’ 4BOV molecule charge change by Bader analysis ΔQom e1 G,BD ‘ObaDCHG’ 5Work function change ΔφeV 1 G,BD ‘DWF’ 6Dispersion energy change ΔED3 eV 1 G,BD ‘DD3’ 7Electronic gap change ΔEgap eV 1 G,BD ‘DGAP’ 8Bader atomic charge change Δqe*A,BD ‘baDCHG’ Table 4. List of physicochemical properties of molecular systems at PBE+D3 theory level contained in MOREQ-G3 subset. Each property presents a name, symbol, unit, dimension, type, and corresponding key in the HDF5 file. Property types are categorized into atomic (A) and global (G). S and BD stand for a single system (e.g. CPLX, SUB, and OM) and for the binding feature, respectively. Eh and a0 refer to the atomic unit of Hatree and Bohr radius. *The dimension of these properties varies for each molecule. 8 Scientific Data | (2025) 12:324 | https://doi.org/10.1038/s41597-025-04616-6 www.nature.com/scientificdata www.nature.com/scientificdata/ To investigate the sensitivity and selectivity of the receptors, we have also calculated 8 binding features for these systems, see the bottom panel in Table4. For example, the adsorption energy Eads, which is defined as the interaction strength between OM and SUB systems, was computed by evaluating: =−−EE EE,(5) adsCPLXSUB OM where ECPLX, ESUB, and EOM are the total energy of each system obtained by VASP. The atomic charge change Δq is obtained by the difference between the Bader atomic charge q of the same atom in CPLX and SUB (OM) systems. By summing up the atomic charge changes of the atoms for individual components, we can obtain the charge change for the receptor (ΔQrec), BOV molecule (ΔQom), and graphene substrate (ΔQGR) upon adsorption of BOV molecules. The other binding features were computed similarly, taking into account the values obtained for the CPLX and SUB systems. Interconnection between the MORE-Q subsets. Here, we summarize the interactions among the three MORE-Q subsets in the following points: • The MORE-Q-G1 subset contains QM property data for 102 BOV molecules and 18 molecular receptors. Among the 39 molecular and atomic properties, we computed the D3 energy, dipole moment, polarizability, and Mulliken charges (see the property list in Table2). • The MORE-Q-G2 subset is built on the geometries from MORE-Q-G1 via the search for molecular docking conformations using BOV molecules and receptors. Accordingly, MORE-Q-G2 contains QM property data for 23,838 dimer conformations at the GFN2-xTB+D4 level and for 1,836 dimers with the lowest binding energies at the PBE+D3 level (see the property list in Tables2 and 3). • The MORE-Q-G3 subset is constructed by depositing 1,836 selected dimers from MORE-Q-G2 onto a graphene surface. Consequently, MORE-Q-G3 includes QM property data at the PBE+D3 level for both the CPLX and SUB systems, as well as binding features that account for property changes in single systems induced by BOV molecule adsorption (see the property list in Table4). Data Records The MORE-Q datasets are available in three HDF5 files in the ZENODO.ORG data repository76. Indeed, one can find there the files MORE-Q-G1.hdf5, MORE-Q-G2.hdf5, and MORE-Q-G3.hdf5 corresponding to the three MORE-Q datasets described in this work. We also provide a README file with technical usage details and examples of how to extract data from the HDF5 files and, then, convert it to Python pandas Dataframes for further analysis (see createDF.py file) HDF5 file format. File sutrcture. Independent of the MORE-Q subset, the information for each molecular structure is stored in a Python dictionary (dict) type containing all relevant properties and recorded in groups in HDF5 file format. The HDF5 file architecture of the MORE-Q subsets is depicted in Fig.3. • For MORE-Q-G1.hdf5 file, containing QM properties of BOV molecules (om) and molecular receptors (rec), nn_om and mm_rec are allocated to the main groups keys, where nn and mm range from 1 to 102 and from 1 to 18, respectively. HDF5 keys to access the atomic numbers, atomic positions (coordinates), and physicochemical properties in each dictionary are provided in Table3. • For MORE-Q-G2.hdf5 file, containing QM properties of the molecular dimers at different levels pf theory, XTB and ORCA become the main groups keys. Under both of them, mm_rec is created as the subgroup. Then, a nn_om subgroup is created per mm_rec subgroup, where nn_om and mm_rec represent the combination of BOV molecules (om) and molecular receptors (rec) that compose a given molecular dimer. Following this, a third-level subgroup, ll_dm, is created to indicate the dimer configurations within each nn_om subgroup, where ll denotes the order of the dimers. Under the ll_dm subgroup, DM (dimer), OM (isolated BOV molecule), and REC (receptor) subgroups were created to store the QM properties listed in Tables3 and 2 for each system. The binding energy values are saved in the DM subgroup. • For MORE-Q-G3.hdf5 file, containing QM properties of the complex systems (i.e. BOV-receptor-graphene), substrate systems (i.e. receptor-graphene), BOV molecules, and the binding features, the main groups keys CPLX, SUB, OM and BD are constructed. Alike the structure of MORE-Q-G2.hdf5 file, we have first created the mm_rec subgroup for each main group. Then, the subgroup, nn_om, is created per nn_om subgroup, representing the BOV molecular and molecular receptor combination. Followed by ll_dm, the dimer configuration order is indicated. QM properties listed in Table4 are then stored per subgroup. Property format. MORE-Q HDF5 files contain various types of data derived from the features of each property. The atomic numbers are stored as a list of strings, with each string representing the atomic number of a corresponding atom in the molecule. The atomic coordinates and forces are saved as a list of 3N vectors, where each vector contains the x, y, and z components. Orbital (band) energies are saved as a vector array, containing 10 energies below HOMO (VBM) level and 10 energies above the LUMO (CBM) level. Vectorial and tensorial properties are also saved as a vector array, with the order maintained as x, y, z for vectorial properties and xx, yy, zz, xy, yz, zx for tensorial properties. For all atomic properties, the order of the vector array is identical to the atomic coordinates vector. For the MORE-Q-G2 and MORE-Q-G3 subsets, the order of the 9 Scientific Data | (2025) 12:324 | https://doi.org/10.1038/s41597-025-04616-6 www.nature.com/scientificdata www.nature.com/scientificdata/ atom type follows receptor→BOV and graphene→receptor→BOV, respectively. This order is also applicable to every atomic property. The other properties are single values that could be directly called by the corresponding HDF5 keys, see Tables2–4. Technical Validation Unlike other datasets in the field of digital olfaction, the MORE-Q datasets contain an extensive set of quantum-mechanical (QM) properties for the building blocks of graphene-based molecular sensors, as well as binding features among the sensor components (see Fig.2). Indeed, structural properties were obtained using the semi-empirical GFN2-xTB method with D4 dispersion correction63, which is known to generate correct geometries at an efficient computational cost. While the energetic, atomic forces and other property calculations were performed at the more accurate level of theory such as PBE with D3 dispersion correction77, as implemented in the VASP code. By doing this, we have guaranteed both sufficiently accurate QM properties and feasible computational time consumption. Before constructing our complex CPLX system (i.e. BOV-receptor-graphene system), we have determined the optimal configuration for the molecular receptor adsorbed on the graphene layer. In this regard, it is known that π−π stacking interaction is the main functionalization mechanism of the mucin-derived receptor. Accordingly, the receptor configuration with the maximal pyrene rings interacting with graphene will be the most stable receptor-graphene system (referred to as the SUB system). Based on this concept, we initially deposited each receptor on graphene in up to four configurations with diverse orientations, which were subsequently optimized using GFN2-xTB methods with D4 correction and tight settings of convergence. Then, the most stable configurations with the lowest adsorption energy were taken as the backbone structures, which were used to further generate the MORE-Q datasets, as shown in Fig.2. Note that a more exhaustive evaluation of the conformational space of the molecular receptors may yield different results. However, the main focus of the current work is to define the pathways for understanding and predicting QM-based structure-property and property-property relationships in potential molecular sensors for digital olfaction. Thus, the selection of these configurations represents the most likely configurations and, therefore, ensures the baseline quality of the dataset. Next, the selected receptor configurations were combined with BOV molecules to form the input geometries for the aISS code60 to determine the docking sites. This method has been successfully used to determine the explicitly solvated structures of peptides and macrocycles from the MPCONF196 dataset78. After performing the docking procedure, hierarchical clustering79 based on root-mean-squared deviation (RMSD) and energetics was carried out to select non-redundant configurations, resulting in a total of 23,838 molecular dimers (see more details for clustering in Fig.S3 of theSI). From this subset, we have selected the most energy-favorable configuration per molecular dimer (ranked by the binding energy Eint) for further examination. Indeed, the principal component analysis (PCA) is conducted on the 23,838 and the selected 1,836 configurations using the global properties stored in MORE-Q-G2 dataset. Fig.4(a) shows the PCA space for both molecular sets, where one can see that the selected 1,836 configurations span the entire region covered by the 23,838 dimers. This indicates that the reduced set is a representative sample of the dimer configurations. As shown in Fig.4(b), even though we sampled only configurations with the lowest Eint, the coverage of the initial Eint values is considerable when considering the reduced molecular set. Here, configurations with Eint smaller than 0.25 eV were filtered out after the energetic selection, as these meta-stable dimers exhibit weak non-covalent interactions and are unlikely to occur during the docking process — another compelling evidence that the reduced set is a representative (a) MORE-Q-G1MORE-Q-G3MORE-Q-G2 (b) (c) 01_om atNUM atXYZ ...... OM ...... REC 01_om atNUM atXYZ ...... ...... 02_om 02_om XTB ORCA 01_om ...... ...... ...... ...... 0_dm 3_dm 01_om ...... 01_rec 0_dm 4_dm 02_rec REC DM OM atXYZ atNUM ...... atXYZ atNUM ...... eBIND ...... ...... 16_om ...... ...... 1_dm 01_om ...... 01_rec 0_dm 4_dm 02_rec REC DM OM atXYZ atNUM ...... atXYZ atNUM ...... eBIND ...... 01_rec ...... atXYZ ...... atNUM CPLX 02_om ...... 01_om 01_rec 02_rec ...... ...... eADS 02_om ...... 01_om 0_dm 0_dm ...... ...... ...... ...... 02_rec 01_rec ...... atXYZ ...... atNUM SUB 02_om ...... 01_om 0_dm ...... ...... 02_rec 01_rec ...... atXYZ ...... atNUM OM 02_om ...... 01_om 0_dm ...... ...... 02_rec BD Fig. 3 Architecture of the HDF5 files corresponding to (a) MORE-Q-G1, (b) MORE-Q-G2, and (c) MOREQ-G3 subsets.