scieee AI-readable full text Open interactive document viewer

Autonomous extraction and building of machine-readable molecular models from publications using large language models

Fleckenstein, Florian; Rodin, Volodymyr; Dieudonne, Isabell; Al Machot, Fadi; Chiacchiera, Silvia; Akagic, Amila; Stephan, Simon

Abstract

Force field models are one of the central elements of molecular simulations. Alarge number of molecular force field models has been developed in the past decades– mostly published in scientific papers. Building machine readable force field inputfiles for simulation engines is a tedious and error-prone task. We developed a methodfor autonomously extracting and building force field files from publications using largelanguage models (LLM). We have tested the new method by extracting 114 forcefield models from 21 scientific publications. The studied force fields comprise 6 - 74parameters. We have compared the performance of different LLMs, namely Gemini2.5 Pro, Claude 4 Sonnet, GPT-4o, Claude 3.7 Sonnet, and Gemini 2.5 Flash. Overall,they yield a similar performance – yet, important differences in individual cases. Theoverall best performance was obtained by the Gemini 2.5 Pro LLM. The force fieldparameters were extracted and identified with an accuracy of 89.1% by the Gemini2.5 Pro LLM. The new autonomous extraction method drastically reduces the timerequired for building force field files – and does not depend on the experience of thesimulator.

Full text

Autonomous extraction and building of machine-readable molecular models from publications using large language models Florian Fleckenstein,†Volodymyr Rodin,‡Isabell Dieudonn´e,†Fadi Al Machot,¶ Silvia Chiacchiera,§Martin Horsch,¶Amila Akagic,‡and Simon Stephan∗,† †Molecular Thermodynamics Group (MTG), RPTU Kaiserslautern, Kaiserslautern 67663, Germany ‡University of Sarajevo, Bosnia and Herzegovina ¶Norwegian University of Life Sciences, Norway §Science and Technology Facilities Council, United Kingdom E-mail: [email protected] Abstract Force field models are one of the central elements of molecular simulations. A large number of molecular force field models has been developed in the past decades – mostly published in scientific papers. Building machine readable force field input files for simulation engines is a tedious and error-prone task. We developed a method for autonomously extracting and building force field files from publications using large language models (LLM). We have tested the new method by extracting 114 force field models from 21 scientific publications. The studied force fields comprise 6 - 74 parameters. We have compared the performance of different LLMs, namely Gemini 2.5 Pro, Claude 4 Sonnet, GPT-4o, Claude 3.7 Sonnet, and Gemini 2.5 Flash. Overall, 1 they yield a similar performance – yet, important differences in individual cases. The overall best performance was obtained by the Gemini 2.5 Pro LLM. The force field parameters were extracted and identified with an accuracy of 89.1% by the Gemini 2.5 Pro LLM. The new autonomous extraction method drastically reduces the time required for building force field files – and does not depend on the experience of the simulator. Introduction Simulations based on molecular models are a powerful tool that enable the prediction of macroscopic thermophysical properties as well as for the modeling of nanoscopic processes. Classical molecular simulations, i.e. molecular dynamics (MD) and Monte Carlo (MC) simulation, have become a standard tool in many scientific disciplines such as computational physics,1–3 physical chemistry,4–7 tribology,8molecular biology,9–13 astronomy,14,15 and chemical engineering.16–21 Matter is modeled in MD and MC simulations using molecular models – so-called force fields – that describe the interactions on the atomistic level. In MD simulations, the force fields are used to calculate the actual forces between interaction sites such that the trajectories of the interaction sites over a given time period can be computed. In MC simulations, the potential energy is used for evaluating the probability for existence of a given atomistic configuration by computing the total potential energy. Thereby, a force field mathematically describes the molecular interactions. Molecular simulations can have impressive predictive capabilities – which facilitates their popularity. Nevertheless, the predictive capabilities depend on the quality of the employed force field.22–29 Moreover, it has been emphasized that the successful reproducibility of predictions from molecular simulations strongly depends on accurate force field models and that even minor mistakes by users can have severe consequences for the results.30–32 Reducing the influence of the simulator as error source upon building force field files and standardizing work flows is desirable.32–35 2 A large number of force fields have been proposed in the past decades, e.g. Refs.36–39 Moreover, various MD/MC simulation engines are in use today, e.g. Refs.40,40–42,42–47 For employing a given force field in a simulation engine, machine readable force field input files are required. Force field models have been published mostly in classical journal articles in the past decades. In most cases, the force field specifications are reported in the main .pdf file of a given publication. Often, the information and data defining a given force field is spread across different entities, i.e. figures, tables, text etc., of a paper. Force field models are often described in text and tables, whereas ready-to-use simulation files are practically never provided in publications (partly due to the lack of standardized force field file formats). Moreover, a publication may report one or multiple force field models for one or more molecules and/ or different molecules, with relevant details either consolidated (e.g. in a table) or scattered throughout the text, sometimes using different potential forms or unit systems. In some cases, relevant information might be included in a graphic such as geometry details. Building force field input files ’by hand’ by simulators is a tedious and error-prone task. While some force field databases exist, e.g. the MolMod database,36,37 that provide input files for simulation engines, these databases comprise only a fraction of the models available in the literature. Presently force field models are not available under FAIR conditions.48 Digitizing, annotating, and depositing force fields in databases is still in its infancy. In this work, we propose a methodology for autonomously digitizing force field models from scientific publications using large language models (LLM). LLMs have gained significant attention across a wide range of natural language processing (NLP) tasks in both academia and industry.49,50 Their performance is being systematically evaluated in diverse applications to better assess their capabilities from multiple perspectives. It is evident that they will continue to play an increasingly important role at the societal level as their potential and associated risks are more thoroughly understood. A key advantage of LLMs, compared to traditional approaches, lies in their versatility to address a broad set of tasks within a single framework. Fundamentally, LLMs are large-scale neural 3 models trained on extensive text corpora to learn patterns of language, semantics, and context.51,52 The core technology underlying most LLMs is the self-attention mechanism within the transformer architecture, which serves as a foundational component for language-related tasks.53 To date, LLMs have been successfully applied to tasks such as information extraction, semantic structuring, classification, summarization of unstructured data and beyond. These applications highlight their potential in the digitization and automated processing of documents and data.54,55 In this work, we propose a novel methodology that introduces LLMs for the autonomous digitization of force field models. This new method is tested using five different LLMs and validated using a variety of force fields extracted from 21 publications from the literature. The results from the digitization are compared with force field files from a benchmark database. This paper is organized as follows: First, a brief introduction of force field models is given – focusing on aspects relevant for their digitization from publications. Then, the methodology for digitizing force field models from publications is presented. Third, the results from the case study based on 21 publications comprising 114 (component-specific rigid classical) force field models are presented and discussed. Finally, conclusions are drawn. Structure of Molecular Force Field Models There is a wide range of force field types available today. Force fields can be classified using different attributes, e.g. atomistic architecture (flexible or rigid), modeling approach (component-specific or transferable), model detail level (all-atom, united-atom, coarse grain), force field chemistry (reactive or non-reactive), underlying parametrization strategy (bottomup or top-down), employed interaction potentials etc. An in-depth discussion is for example given in Ref.56 In this work, classical (i.e. Newtonian) component-specific rigid force fields for molecules and ions are used for testing the new digitization method. In classical force fields, the molecular interactions are modeled by interaction potentials 4 that describe the potential energy as a function of the spatial relations between two (or more particles/sites) u(r). Despite the fact that these interaction potentials provide a relatively crude simplification of the ’true’ molecular interactions, classical force fields have strong predictive capabilities.57–60 In classical molecular force field models, atoms (or groups of atoms) are represented as interaction sites, i.e. points in space. The interactions between these sites are described by the interaction potentials for which different ansatz functions exist, e.g. the Lennard-Jones potential, multipole interactions, etc. In rigid force field models, the interaction sites of a model have fixed relative positions. In flexible force fields, intramolecular degrees of freedom exist, which require additional interaction potentials. In this work, rigid force field models were used for testing the proposed digitization method. An interaction potential (both intermolecular and intramolecular) in a force field is specified by the general mathematical equation and the specific parameters applicable to the interactions under consideration. Moreover, the mass of each site needs to be specified. In summary, a force field consists of (a) the coordinates of the interaction sites and (b) the information on the interaction site type allocated to a given site, which is specified by the mathematical equation and the corresponding parameters. All this information has to be comprised in a well-defined way in force field files to be used by a simulation engine. Intermolecular interactions are usually represented by dispersive-repulsive and electrostatic contributions. Dispersive-repulsive interactions between two particles iand jseparated by a distance rij can for example be modeled by the Mie potential uMie (which is the generalized form of the Lennard-Jones (LJ) potential61) as uMie(rij ) = n n−mn m m n−mεij σij rij n −σij rij m,(1) where nand mdenote the repulsive and dispersive exponents, respectively, εij is the dispersion energy, and σij is the size parameter of the interaction site. For the special case n= 12 and m= 6, the Mie potential reduces to the LJ potential. Electrostatic interactions are commonly described using either point charges, point dipoles, 5 or point quadrupoles. For point charges, the electrostatic interaction is given by Coulomb’s law62 uqq ij (rij, qi, qj) = 1 4πε0 qiqj rij ,(2) where qiand qjdescribe the magnitude of the charge and rij the distance between the two charges. Point dipoles describe the electrostatic field induced by two point charges with equal magnitude, but opposite sign. The interaction of two point dipoles is given by62,63 uDD ij (rij, θi, θj, ϕij , µi, µj) = 1 4πε0 µiµj r3 ij sin θisin θjcos ϕij −2 cos θicos θj,(3) where µiand µjdescribe the magnitude of the point dipoles, rij the distance between the two point dipoles and θi,θj, and ϕij indicate the relative angular orientation of the two point dipoles. Point quadrupoles describe the electrostatic field induced by either two collinear point dipoles with the same moment, but opposite orientation, or by at least three collinear point charges with alternating sign. The interaction potential of two quadrupole sites at a distance rij is given by63,64 uQQ ij (rij, θi, θj, ϕij , Qi, Qj) =1 4πε0 3 4 QiQj r5 ij "1−5cos2θi+ cos2θj−15 cos2θicos2θj + 2sin θisin θjcos ϕij −4 cos θicos θj2#,(4) where Qiand Qjare the magnitude of the quadrupole moments, θi,θj, and ϕij indicate 6 the relative angular orientation of the two point quadrupoles.63 The number of parameters differs for the different interaction potentials, e.g. for a simple LJ potential, the two parameters εij and σij are required – each being a single scalar value (for a Mie potential, additionally mand nare required). For a point dipole, on the other hand, one scalar parameter value is required, namely the dipole moment µ, and one vector specifying its spatial orientation (θand ϕ). An interaction site type is defined by its potential parameters together with the Cartesian coordinates (x, y, z) of the site, and additionally its mass. These interaction site types are not exhaustive, but the vast majority of rigid force field models are built from these interaction potentials. The force field models used for testing the digitization method are built from the above listed interaction potentials. Intramolecular interaction potentials define molecular flexibility and allow for vibrations within molecules. These include bond, angle, torsion, and improper torsion terms. A model in which the molecular geometry is fixed and no vibrations occur is referred to as ’rigid’. In the following, we restrict our approach to rigid models that incorporate both LJ and electrostatic interactions. Figure 1 illustrates two representative molecular force field models, namely for chlorine65 and acetone.66 A force field typically comprises the spatial coordinates of interaction sites (x, y, z), the mass associated with each site (m), and interaction parameters required for the Mie/LJ and electrostatic potential functions (ε, σ, m, n, µ, Q). For implementation in molecular simulation engines, the force field description must usually conform to specific data formats, which may additionally require auxiliary information such as the number of sites assigned to each interaction type and which interaction type is applicable for a given model. 7 Force Field Digitization Method The autonomous digitization method developed in this work extracts, transforms, identifies, and annotates information from scientific publications into simulation-ready force field input files. We use the .pm file format which is already in use within simulation engines as well as data curation infrastructure36,40 – an exemplary .pm file is given in the electronic Supporting Information. The digitization method consists of a sequence of five steps: (1) Optical character recognition (OCR), (2) text cleaning, (3) table detection, (4) data extraction, and (5) file generation. Each step is described in the following in detail, with an implementation provided in the electronic Supporting Information. An overview of the entire pipeline is shown in Figure 2 – in comparison to the classical manual approach. The goal of the workflow is to autonomously extract and identify force field parameters from publications, standardize them into a machine-readable format, and produce reproducible force field files that include both the model and its metadata. Manually creating an input file involves locating all relevant information within a given publication, mapping it to the simulation engine’s data scheme and format, and, if necessary, converting units (cf. Figure 2-right). Potentially, also coordinates have to be transformed to different coordinate systems, e.g. converting from Cartesian to Z-matrix coordinates. Evidently, units, notations, and conventions have to be interpreted correctly for extracting force field information. Ambiguities or missing details frequently require additional effort to resolve in the traditional manual route and transcription into the precise syntax and unit system required by the simulation engine introduces further opportunities for error. Also the validation of the constructed input file is not trivial since simulation benchmark data are often scarce. As a result, preparing even a single ready-to-use input file can consume a considerable amount of time if manually done, e.g. some hours by an experienced simulator. Autonomous extraction using LLM offers an alternative by performing the same mapping and formatting tasks automatically and at scale, greatly reducing the effort required to reuse published force fields. In the autonomous method developed in this work, all these steps are 8 carried out by the LLM. Autonomous Force Field Extraction from Scientific Publications The proposed approach automates the traditionally manual task of transcribing force field data from research papers into force field files. By reducing human involvement, the method circumvents user-induced transcription errors and accelerates data preparation. The workflow of the method, illustrated in Figure 2, progresses through the five steps, which are detailed in the following subsections. Step 1: Optical Character Recognition (OCR) The process begins by selecting a research paper, typically in .pdf format, as input. Many .pdf files, especially older scanned ones, store text as images rather than selectable characters. To make the content usable by our method, the first step is to apply optical character recognition (OCR). For this purpose, we employ the mistral-ocr-2505 model from Mistral AI,67 which acts as a digital reader by analyzing each page image, identifying the shapes of letters and numbers, and converting them into machine-readable text. Conceptually, this step transforms the visual representation of the document into structured textual data that subsequent steps of the pipeline process. Step 2: Text Cleaning The raw text obtained from OCR often contains non-essential elements. Therefore, a text cleaning is applied. In this step, the system processes the raw digital text (as shown in Figure 3-a.) to remove artifacts that are not part of the main textual content or tables relevant for parameter extraction. This includes filtering out footers, bibliography, metadata, and other content that does not contain relevant data. The purpose of cleaning is to ensure that only the core scientific text and data tables are passed on to the next step, preventing interference in subsequent analysis steps. An example of extracted text after text cleaning 9 other LLM. Claude 3.7 Sonnet and Claude 4 Sonnet follow as the next-best performers, delivering consistently good results with only a marginal difference between them. Gemini 2.5 Flash and GPT-4o, lag behind, consistently underperforming the aforementioned LLM across the evaluated papers. Possibly, the superior accuracy of Gemini 2.5 Pro stems from its design as a high-reasoning model, which allocates more computation during inference to produce robust and well-reasoned responses.68 In contrast, Gemini 2.5 Flash is optimized for speed and efficiency,68 making it less suited for tasks that demand strong generalization and domain knowledge, although it may be suitable for applications where processing speed is a critical factor. For all employed LLM, the obtained accuracy differs for different papers, cf. Figure 6 and 7. This scattering is largest for GPT-4o and smallest for Gemini 2.5 Pro. Interestingly, this is in line with the relative accuracy of the studied LLMs. Moreover, the performance of the LLMs for a given paper differs in parts significantly, e.g. Gemini 2.5 Pro yields excellent results for papers #1 and #4 for which Claude 3.7 Sonnet yields relatively poor results. For paper #8, on the other hand, Claude 3.7 Sonnet outperforms Gemini 2.5 Pro. Figure 8 shows per-paper, per-run accuracy results for each model, with the global best across models shown in the background. In some cases, the results from the three independent runs are practically identical, e.g. for the #9 results from Gemini 2.5 Pro. In other cases, significantly different results are obtained from the three independent runs, e.g. for the #8 results from Gemini 2.5 Pro. Gemini 2.5 Pro frequently meets or approaches the global best, reinforcing its leading placement. Claude 3.7 Sonnet and Claude 4 Sonnet follow closely, with occasional drops. The remaining models (GPT-4o and Gemini 2.5 Flash) reveal specific papers where the extraction accuracy is significantly lower. For establishing a better understanding of the nature of the extraction errors, we performed a detailed analysis of 6,049 errors across all employed LLM models, replica runs, force fields, and papers, using the error types defined in Table 4. The distribution, shown in Table 5, reveals that the two most dominant error types are ”Numeric Value Mismatch” 16 (42.0% of all errors) and ”Zero Coordinate Error” (17.0%). Together, these account for 59% of all extraction errors. Other significant error types include ”Count/Index Error” (8.0%), ”Rotation Parameter Error” (7.2%), and ”Placeholder Value Error” (6.4%). Furthermore, specific force fields parameters that were most frequently extracted incorrectly are summarized in Table 6. Basic physical properties such as ‘mass‘ were the most common source of error, followed by atomic coordinates (‘z‘, ‘y‘, ‘x‘). The errors associated with ‘mass‘ were primarily ”Numeric Value Mismatch” and ”Placeholder Value Error”. Coordinates were susceptible to being incorrectly zeroed (”Zero Coordinate Error”) or having their values mismatched (”Coordinate Extraction Error”). Less frequent but still notable errors included ’Sign Error’ for parameters like ‘charge‘ and ’Categorical Misclassification’ for symbolic labels like site types. Conclusions LLM provide a remarkably fast and robust route for digitizing force field models from publications. The autonomous data extraction pipeline proposed in this work transforms a the model data comprised in research paper .pdf files into structured, simulation-ready .pm files. This automated approach offers a significant improvement over manual methods, enhancing efficiency and reducing the potential for human errors in preparing molecular simulation inputs. The accuracy of the autonomous method is roughly 75 to 90% – depending on the employed LLM. Accurate and reproducible molecular simulations demand 100% accuracy, i.e. force field input files without any error. Thus, the method and results from this paper should be seen as a starting point for developing a reliable and autonomous digitization tool. The method proposed in this work is an attractive alternative to tedious and error-prone “by hand” digitization of force fields for molecular simulations. Possible ways to improve the accuracy of the digitization method can be guided from the experiences from the comprehensive evaluation carried out in this work. These results 17 provide a robust benchmark for testing refinements and improvements. A large variety of routes to improve the promising results from this work seem attractive. On the one hand, the prompt may be further refined as to give more context in the translation between Z-matrix and Cartesian coordinates or by providing more physical background/information. The error analysis revealed that numeric mismatches and zero coordinate errors are the most frequent error types, which most frequently occur for the coordinates and the mass. On the other hand, OCR is not 100% error-free and newer, even stronger LLMs might handle this task even better. Additionally, LLMs with greater variability or higher error rates may benefit from ensembling, further prompt adaptation, or supplementary rule-based checks. Moreover, another interesting approach would be to use different LLMs sequentially. The results from this work show that larger, more capable models68,69,71 (Gemini 2.5 Pro, Claude 3.7, and 4 Sonnet) yield an overall more accurate extraction of force field data from scientific publications. The overall best performance was achieved by the Gemini 2.5 Pro LLM with an accuracy rate of 89.1%. Performance variability across papers and occasional outliers further highlight the need to study model stability and generalization in more detail. Furthermore, the comprehensive evaluation of the new method provided insights in challenges in force field publication in the scientific literature: Different unit systems are used in different publications (and communities); different coordinate systems are used (Z-matrix, Cartesian etc.); Information on a given force field are often scattered across a publication with some data given in the bulk text, some in tables, and sometimes even graphics. Sometimes, no spatial coordinates are specified, but have to be determined as the equilibrium configuration from the stated intramolecular interactions. Moreover, there is no universal and standardized force field data format established yet. These insights highlight not only the need for more advanced extraction methods but also for community-wide standards and publication practices to make force fields consistently reusable. 18 Supporting Information Available The Supporting Information is available free of charge at LINK. Acknowledgement The authors gratefully acknowledge funding of the present work by the Federal Ministry of Research, Technology and Space (BMFTR, Germany) 16ME0613 WindHPC and from the European Union’s Horizon Europe research and innovation program under Grant Agreement No. 101137725 (BatCAT). Simulations for validation were carried out on the ELWE supercomputer at Regional University Computing Center Kaiserslautern (RHRZ) under the grant RPTU-MTD and on the HPC machine MOGON NHR at the NHR SW under the grant TUK-MSG (supported by the Federal Ministry of Research, Technology and Space and the state governments). Conflicts of Interest There are no conflicts of interest to declare. Data Availability The python pipeline developed within this work is provided in the Supporting Information. Moreover, the force field .pm files built within this work using the LLMs are provided as well as the reference .pm files that were built ’by hand’. Author contribution statement Florian Fleckenstein: Data Curation, Formal Analysis (equal), Software (equal), Visualization (equal), Writing/Original Draft Preparation (equal); Volodymyr Rodin: Data Curation, 19 Formal Analysis (equal), Software (equal), Visualization (equal), Writing/Original Draft Preparation (equal); Isabell Dieudonne: Data Curation (support), Software; Fadi Al Machot: Writing/Review & Editing (support); Silvia Chiacchiera: Writing/Review & Editing (support); Martin Horsch: Writing/Review & Editing (support); Amila Akagic: Conceptualization (equal), Methodology (equal), Supervision (equal), Writing/Original Draft Preparation (lead); Simon Stephan: Conceptualization (equal), Methodology (equal), Supervision (equal), Writing/Original Draft Preparation (lead), Funding Acquisition (equal). 20 Figures & Tables Table 1: Technical specifications of employed LLMs. Model Modality Max context / tokens Max output / tokens Gemini 2.5 Pro68 Text, Vision, Audio, Video 1,048,576 65,535 Claude 4 Sonnet69 Text, Vision 200,000 64,000 Claude 3.7 Sonnet71 Text, Vision 200,000 64,000 Gemini 2.5 Flash68 Text, Vision, Audio, Video 1,048,576 65,535 GPT-4o70 Text, Vision, Audio 128,000 16,384 Table 2: List of publications used in the validation data set. The columns indicate the publication identifier number (used throughout this work), year and number of force field models (NM) in a given paper. No. Paper Year NMNo. Paper Year NM 1Eckl et al.73 2007 1 12 Abbas & Nordholm 74 1995 3 2Horsch et al.75 2010 1 13 Eckelsbach et al.76 2015 1 3Kohns et al.77 2017 2 14 Eckl et al.78 2008 10 4K¨oster et al.79 2018 1 15 Fischer et al.80 1984 7 5Merker et al.81 2010 1 16 Gao et al.82 1997 8 6Merker et al.83 2012 2 17 Koyama et al.84 2005 1 7Miroshnichenko et al.85 2013 2 18 Romano & Singer 65 1979 2 8Mu˜noz-Mu˜noz et al.86 2017 3 19 Zhang & Siepmann 87 2006 1 9Mu˜noz-Mu˜noz et al.88 2018 1 20 de Santis et al.89 1987 3 10 Vrabec & Fischer 90 1995 1 21 van Leeuwen 91 1994 62 11 Windmann et al.66 2014 1 21 Table 3: Summary of performance metrics results for studied LLMs. Average accuracy is computed by calculating mean accuracy over papers; correct/incorrect/missing counts are aggregate totals of parameters across all papers and runs. Model Avg Acc / % Prec / % Rec / % F1 / % Correct Incorrect Missing Gemini 2.5 Pro 89.1 91.0 99.2 94.9 7382 (90%) 733 (9%) 63 (1%) Claude 4 Sonnet 85.2 92.4 98.9 95.5 7467 (91%) 618 (8%) 81 (1%) Claude 3.7 Sonnet 86.3 89.1 99.3 93.9 7232 (89%) 883 (11%) 51 (1%) Gemini 2.5 Flash 78.0 81.2 99.7 89.5 6618 (81%) 1531 (19%) 17 (0%) GPT-4o 76.2 75.5 98.5 85.5 6094 (75%) 1978 (24%) 94 (1%) Average - - - - - (85%) - (14%) - (1%) Table 4: Classification of extraction errors observed in incorrect results across evaluated datasets. Examples are shown in the format reference →extracted. Error Type Description Example (Reference →Extracted) Numeric Value Mismatch Extracted numeric value differs significantly from reference. mass = 127.91241 →36.46 (Abbas, 1995)74 Zero Coordinate Error Non-zero coordinates incorrectly replaced with 0.0. z = -0.232 →0.0 (de Santis, 1987)89 Count/Index Error Wrong extraction of discrete counts (sites, site types). NSiteTypes = 2 →8 (Munoz-Munoz, 2017)86 Rotation Parameter Error Misextraction of rotational inertia, angles, or axes. theta = 30.0 →45.0 (Gao, 1997)82 Placeholder Value Error Values replaced with placeholder tokens like [MASS VALUE]. mass = 12.011 →[MASS VALUE] (de Santis, 1987)89 Coordinate Extraction Error Incorrect coordinates (non-zero to nonzero). z = -0.232 →1.185 (de Santis, 1987)89 Missing Parameter Parameter completely absent from extraction. mass = 16.042 →(missing) (Koyama, 2005)84 Physical Property Misread Misreading of physical properties (charges, dipoles, etc.). dipole = 3.228 →1.988 (Gao, 1997)82 Categorical Misclassification Site types or labels incorrectly classified. SiteType = Dipole →LJ126 (Munoz-Munoz, 2017)86 Sign Error Correct magnitude with wrong or missing sign. charge = +0.2556 →0.2556 (Horsch, 2010)75 Other Uncategorized errors (mostly unusual formats). mass = 13.019 →#CH (Horsch, 2010)75 22 Table 5: Distribution of extraction error types across 6,049 total errors from 21 scientific papers for all force fields and replica runs. Error Type Count % Num of Papers Affected Numeric Value Mismatch 2,539 42.0% 21 Zero Coordinate Error 1,031 17.0% 13 Count/Index Error 481 8.0% 16 Rotation Parameter Error 438 7.2% 10 Placeholder Value Error 385 6.4% 21 Coordinate Extraction Error 379 6.3% 13 Missing Parameter 306 5.1% 21 Physical Property Misread 297 4.9% 4 Categorical Misclassification 97 1.6% 8 Sign Error 72 1.2% 5 Other 24 0.4% 21 Total 6,049 23 Table 6: Molecular parameters most frequently extracted incorrectly, showing the 15 variables with the highest error counts. Variable Error Count % Primary Error Types mass 2,239 37.0% Numeric Value Mismatch, Placeholder Value Error z 638 10.5% Zero Coordinate Error, Coordinate Extraction Error y 506 8.4% Zero Coordinate Error, Coordinate Extraction Error NSiteTypes 365 6.0% Count/Index Error theta 354 5.9% Rotation Parameter Error x 320 5.3% Zero Coordinate Error, Coordinate Extraction Error, Missing Parameter charge 310 5.1% Missing Parameter, Sign Error epsilon 293 4.8% Missing Parameter, Numeric Value Mismatch dipole 271 4.5% Missing Parameter, Physical Property Misread sigma 235 3.9% Missing Parameter, Numeric Value Mismatch phi 228 3.8% Rotation Parameter Error NSites 121 2.0% Count/Index Error, Missing Parameter SiteType 99 1.6% Categorical Misclassification, Missing Parameter quadrupole 31 0.5% Physical Property Misread, Placeholder Value Error MIE n 15 0.2% Missing Parameter 24 Figure 1: Schematic representation of two exemplary chosen molecular force field models with corresponding geometric, mass, and interaction parameters. Left: Chlorine model of Romano & Singer.65 Right: Acetone model of Windmann et al.66 The total number of parameters is 15 for chlorine (including 12 geometric, mass, and interaction parameters and 3 parameters required for the given file format used within the simulation engine) and 48 for acetone (including 36 geometric, mass, and interaction parameters and 12 parameters required for the given file format used within the simulation engine). 25 References (1) Ruestes, C. J.; Alhafez, I. A.; Urbassek, H. M. Atomistic Studies of Nanoindentation – A Review of Recent Advances. Crystals 2017,7. (2) Ewen, J. P.; Spikes, H. A.; Dini, D. Contributions of molecular dynamics simulations to elastohydrodynamic lubrication. Tribology Letters 2021,69, 24. (3) Urschel, M.; Stephan, S. Determining Brown’s Characteristic Curves Using Molecular Simulation. Journal of Chemical Theory and Computation 2023,5, 1537–1552. (4) Getman, R. B.; Bae, Y.-S.; Wilmer, C. E.; Snurr, R. Q. Review and Analysis of Molecular Simulations of Methane, Hydrogen, and Acetylene Storage in Metal–Organic Frameworks. Chemical Reviews 2012,112, 703–723. (5) Stephan, S.; Hasse, H. Enrichment at Vapour-Liquid Interfaces of Mixtures: Establishing a Link between Nanoscopic and Macroscopic Properties. Int. Rev. Phys. Chem. 2020,39, 319–349. (6) van Gunsteren, W. F.; Berendsen, H. J. C. Computer Simulation of Molecular Dynamics: Methodology, Applications, and Perspectives in Chemistry. Angewandte Chemie International Edition in English 1990,29, 992–1023. (7) Stephan, S.; Thol, M.; Vrabec, J.; Hasse, H. Thermophysical Properties of the LennardJones Fluid: Database and Data Assessment. J. Chem. Inf. Model. 2019,59, 4248– 4265. (8) Stephan, S.; Schmitt, S.; Hasse, H.; Urbassek, H. M. Molecular dynamics simulation of the Stribeck curve: Boundary lubrication, mixedlubrication, and hydrodynamic lubrication on the atomistic level. Friction 2023,11, 2342–2366. (9) Sponer, J.; Bussi, G.; Krepl, M.; Banas, P.; Bottaro, S.; Cunha, R. A.; Gil-Ley, A.; Pinamonti, G.; Poblete, S.; Jurecka, P.; others RNA Structural Dynamics as Captured 32 by Molecular Simulations: A Comprehensive Overview. Chemical Reviews 2018,118, 4177–4338. (10) Salo-Ahen, O. M. H. et al. Molecular Dynamics Simulations in Drug Discovery and Pharmaceutical Development. Processes 2021,9. (11) Levitt, M. The birth of computational structural biology. Nature Structural Biology 2001,8, 392–393. (12) Hollingsworth, S. A.; Dror, R. O. Molecular Dynamics Simulation for All. Neuron 2018, 99, 1129–1143. (13) Fleckenstein, F.; Stephan, S.; Hasse, H. Elucidating the behavior of the SARS-CoV-2 virus surface at vapor–liquid interfaces using molecular dynamics simulation. Proceedings of the National Academy of Sciences 2024,121. (14) Alfaridzi, R.; Stephan, S.; Urbassek, H. M.; Rosandi, Y. Shear flow in an ice layer between two sliding solids. Molecular Simulation 2025, 1–8. (15) Rosandi, Y.; Urbassek, H. M.; Nietiadi, M. L.; Alfaridzi, R. Astrophysical study of dust collision using molecular dynamics method: an overview. J. Phys. Conf. Ser. 2024,2915, 012006. (16) Prausnitz, J. M.; Tavares, F. W. Thermodynamics of fluid-phase equilibria for standard chemical engineering operations. AIChE Journal 2004,50, 739–761. (17) Vrabec, J. et al. SkaSim – Scalable HPC Software for Molecular Simulation in the Chemical Industry. Chemie Ingenieur Technik 2018,90, 295–306. (18) Maginn, E. J.; Elliott, J. R. Historical Perspective and Current Outlook for Molecular Dynamics As a Chemical Engineering Tool. Industrial & Engineering Chemistry Research 2010,49, 3059–3078. 33 (19) R¨oßler, J.; Antolovic, I.; Stephan, S.; Vrabec, J. Assessment of thermodynamic models via Joule-Thomson inversion. Fluid Phase Equilibria 2022,556, 113401. (20) Stephan, S.; Br˚aten, V.; Hasse, H. Mass transfer at vapor-liquid interfaces of H2O + CO2mixtures studied by molecular dynamics simulation. Journal of Non-Equilibrium Thermodynamics 2024,49, 441–461. (21) Schmitt, S.; Hasse, H.; Stephan, S. Entropy scaling for diffusion coefficients in fluid mixtures. Nature Communications 2025,16. (22) Oliveira, M. P.; Gon¸calves, Y. M. H.; Ol Gheta, S. K.; Rieder, S. R.; Horta, B. A. C.; H¨unenberger, P. H. Comparison of the Unitedand All-Atom Representations of (Halo)alkanes Based on Two Condensed-Phase Force Fields Optimized against the Same Experimental Data Set. Journal of Chemical Theory and Computation 2022,18, 6757–6778. (23) Schmitt, S.; Fleckenstein, F.; Hasse, H.; Stephan, S. Comparison of Force Fields for the Prediction of Thermophysical Properties of Long Linear and Branched Alkanes. J. Phys. Chem. B 2023, (24) Ewen, J.; Gattinoni, C.; Thakker, F.; Morgan, N.; Spikes, H.; Dini, D. A Comparison of Classical Force-Fields for Molecular Dynamics Simulations of Lubricants. materials 2016,9, 1–17. (25) Vega, C.; Abascal, J. L. F. Simulating water with rigid non-polarizable models: a general perspective. Phys. Chem. Chem. Phys. 2011,13, 19663–19688. (26) Guvench, O.; MacKerell, A. D. In Molecular Modeling of Proteins; Kukol, A., Ed.; Springer-Humana Press: Totowa, NJ, 2008; pp 63–88. (27) Levitt, M.; Hirshberg, M.; Sharon, R.; Daggett, V. Potential energy function and pa34 rameters for simulations of the molecular dynamics of proteins and nucleic acids in solution. Computer Physics Communications 1995,91, 215–231. (28) Albaugh, A. et al. Advanced Potential Energy Surfaces for Molecular Simulation. The Journal of Physical Chemistry B 2016,120, 9811–9832. (29) Weik, T.; Alt, D.; Stephan, S. Characteristic curves of water: A force field assessment. The Journal of Chemical Physics 2025,163. (30) Schappals, M.; Mecklenfeld, A.; Kr¨oger, L.; Botan, V.; K¨oster, A.; Stephan, S.; Garc´ıa, E. J.; Rutkai, G.; Raabe, G.; Klein, P.; Leonhard, K.; Glass, C. W.; Lenhard, J.; Vrabec, J.; Hasse, H. Round Robin Study: Molecular Simulation of Thermodynamic Properties from Models with Internal Degrees of Freedom. Journal of Chemical Theory and Computation 2017,13, 4270–4280. (31) Craven, N. C. et al. Achieving Reproducibility and Replicability of Molecular Dynamics and Monte Carlo Simulations Using the Molecular Simulation Design Framework (MoSDeF). J. Chem. Eng. Data 2025,70, 2178–2199. (32) Frenkel, D. Simulations: The dark side. The European Physical Journal Plus 2013, 128, 10. (33) Wong-ekkabut, J.; Karttunen, M. The good, the bad and the user in soft matter simulations. Biochimica et Biophysica Acta (BBA) - Biomembranes 2016,1858, 2529–2538. (34) Thompson, M. W.; Gilmer, J. B.; Matsumoto, R. A.; Quach, C. D.; Shamaprasad, P.; Yang, A. H.; Iacovella, C. R.; McCabe, C.; Cummings, P. T. Towards molecular simulations that are transparent, reproducible, usable by others, and extensible (TRUE). Molecular Physics 2020,118, e1742938. (35) Merz, K.; Amaro, R.; Cournia, Z.; Rarey, M.; Soares, T.; Tropsha, A.; Wahab, H.; 35 Wang, R. Editorial: Method and Data Sharing and Reproducibility of Scientific Results. Journal of Chemical Information and Modeling 2020,60, 5868–5869. (36) Stephan, S.; Horsch, M.; Vrabec, J.; Hasse, H. MolMod - an Open Access Database of Force Fields for Molecular Simulations of Fluids. Molecular Simulation 2019,45, 806–814. (37) Schmitt, S.; Kanagalingam, G.; Fleckenstein, F.; Froescher, D.; Hasse, H.; Stephan, S. Extension of the MolMod Database to Transferable Force Fields. Journal of Chemical Information and Modeling 2023,63, 7148–7158. (38) Eggimann, B. L.; Sunnarborg, A. J.; Stern, H. D.; Bliss, A. P.; Siepmann, J. I. An online parameter and property database for the TraPPE force field. Molecular Simulation 2014,40, 101–105. (39) Dodda, L. S.; Cabeza de Vaca, I.; Tirado-Rives, J.; Jorgensen, W. L. LigParGen web server: an automatic OPLS-AA parameter generator for organic ligands. Nucleic Acids Research 2017,45, W331–W336. (40) Nitzke, I.; Guevara-Carrion, G.; Saric, D.; Homes, S.; Stephan, S.; Fingerhut, R.; Bernreuther, M.; Hasse, H.; Vrabec, J. ms2: A molecular simulation tool for thermodynamic properties, release 5.0. Computer Physics Communications 2025,310, 109541. (41) Fingerhut, R.; Guevara-Carrion, G.; Nitzke, I.; Saric, D.; Marx, J.; Langenbach, K.; Prokopev, S.; Celn´y, D.; Bernreuther, M.; Stephan, S.; Kohns, M.; Hasse, H.; Vrabec, J. ms2: A molecular simulation tool for thermodynamic properties, release 4.0. Computer Physics Communications 2021,262, 107860. (42) Thompson, A. P.; Aktulga, H. M.; Berger, R.; Bolintineanu, D. S.; Brown, W. M.; Crozier, P. S.; in ’t Veld, P. J.; Kohlmeyer, A.; Moore, S. G.; Nguyen, T. D.; Shan, R.; Stevens, M. J.; Tranchida, J.; Trott, C.; Plimpton, S. J. LAMMPS - a flexible simulation 36 tool for particle-based materials modeling at the atomic, meso, and continuum scales. Computer Physics Communications 2022,271, 108171. (43) Van Der Spoel, D.; Lindahl, E.; Hess, B.; Groenhof, G.; Mark, A. E.; Berendsen, H. J. C. GROMACS: Fast, flexible, and free. Journal of Computational Chemistry 2005, 26, 1701–1718. (44) Abraham, M. J.; Murtola, T.; Schulz, R.; P´all, S.; Smith, J. C.; Hess, B.; Lindahl, E. GROMACS: High Performance Molecular Simulations Through Multi-level Parallelism from Laptops to Supercomputers. SoftwareX 2015,1-2, 19–25. (45) Smith, W.; Yong, C.; Rodger, P. DL POLY: Application to molecular simulation. Mol. Simulat. 2002,28, 385–471. (46) Niethammer, C.; Becker, S.; Bernreuther, M.; Buchholz, M.; Eckhardt, W.; Heinecke, A.; Werth, S.; Bungartz, H.-J.; Glass, C. W.; Hasse, H.; Vrabec, J.; Horsch, M. ls1 mardyn: The Massively Parallel Molecular Dynamics Code for Large Systems. J. Chem. Theory Comput. 2014,10, 4455–4464. (47) Phillips, J. C.; Braun, R.; Wang, W.; Gumbart, J.; Tajkhorshid, E.; Villa, E.; Chipot, C.; Skeel, R. D.; Kal´e, L.; Schulten, K. Scalable molecular dynamics with NAMD. J. Comput. Chem. 2005,26, 1781–1802. (48) Wilkinson, M. D.; Dumontier, M.; Aalbersberg, I. J.; Appleton, G.; Axton, M.; Baak, A.; Blomberg, N.; Boiten, J.-W.; da Silva Santos, L. B.; Bourne, P. E.; others The FAIR Guiding Principles for scientific data management and stewardship. Scientific data 2016,3, 1–9. (49) Raiaan, M. A. K.; Mukta, M. S. H.; Fatema, K.; Fahad, N. M.; Sakib, S.; Mim, M. M. J.; Ahmad, J.; Ali, M. E.; Azam, S. A review on large language models: Architectures, applications, taxonomies, open issues and challenges. IEEE access 2024,12, 26839– 26874. 37 (50) Chang, Y.; Wang, X.; Wang, J.; Wu, Y.; Yang, L.; Zhu, K.; Chen, H.; Yi, X.; Wang, C.; Wang, Y.; others A survey on evaluation of large language models. ACM transactions on intelligent systems and technology 2024,15, 1–45. (51) Naveed, H.; Khan, A. U.; Qiu, S.; Saqib, M.; Anwar, S.; Usman, M.; Akhtar, N.; Barnes, N.; Mian, A. A comprehensive overview of large language models. ACM Transactions on Intelligent Systems and Technology 2025,16, 1–72. (52) Yang, J.; Jin, H.; Tang, R.; Han, X.; Feng, Q.; Jiang, H.; Zhong, S.; Yin, B.; Hu, X. Harnessing the power of llms in practice: A survey on chatgpt and beyond. ACM Transactions on Knowledge Discovery from Data 2024,18, 1–32. (53) Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; Polosukhin, I. Attention Is All You Need. Advances in Neural Information Processing Systems (NeurIPS). 2017. (54) Polak, M. P.; Morgan, D. Extracting accurate materials data from research papers with conversational language models and prompt engineering. Nature Communications 2024,15, 1569. (55) Schilling-Wilhelmi, M.; R´ıos-Garc´ıa, M.; Shabih, S.; Gil, M. V.; Miret, S.; Koch, C. T.; M´arquez, J. A.; Jablonka, K. M. From text to insight: large language models for chemical data extraction. Chem. Soc. Rev. 2025,54, 1125–1150. (56) Kanagalingam, G.; Schmitt, S.; Fleckenstein, F.; Stephan, S. Data scheme and data format for transferable forcefields for molecular simulation. scientific data 2023,10, 495. (57) Thol, M.; Dubberke, F.; Rutkai, G.; Windmann, T.; K¨oster, A.; Span, R.; Vrabec, J. Fundamental equation of state correlation for hexamethyldisiloxane based on experimental and molecular simulation data. Fluid Phase Equilibria 2016,418, 133–151. 38 (58) Saager, B.; Fischer, J. Predictive Power of Effective Intermolecular Pair Potentials: MD Simulation Results for Methane up to 1000 MPa. Fluid Phase Equilibria 1990,57, 35–46. (59) Fleckenstein, F.; Becker, S.; Hasse, H.; Stephan, S. Vapor–liquid interfacial properties of the system acetone + CO2: Experiments, molecular simulation, density gradient theory, and density functional theory. Fluid Phase Equilibria 2025,596, 114436. (60) Stephan, S.; Fleckenstein, F.; Hasse, H. Vapor-Liquid Interfacial Properties of the Systems (Toluene + CO2) and (Toluene + N2): Experiments, Molecular Simulation, and Density Gradient Theory. Journal of Chemical & Engineering Data 2023, (61) Lenhard, J.; Stephan, S.; Hasse, H. On the History of the Lennard-Jones Potential. Annalen der Physik 2024,536. (62) Allen, M. P.; Tildesley, D. J. Computer Simulation of Liquids, 2nd ed.; Oxford University Press: Oxford, United Kingdom, 2017. (63) C. G. Gray; K. E. Gubbins Theory of Molecular Fluids, Vol. 1: Fundamentals; Clarendon Press, Oxford, 1984. (64) Hirschfelder, J. O.; Curtiss, C. F.; Bird, R. B. The molecular theory of gases and liquids; John Wiley & Sons, 1964. (65) Romano, S.; Singer, K. Calculation of the entropy of liquid chlorine and bromine by computer simulation. Mol. Phys. 1979,37, 1765–1772. (66) Windmann, T.; Linnemann, M.; Vrabec, J. Fluid Phase Behavior of Nitrogen + Acetone and Oxygen + Acetone by Molecular Simulation, Experiment and the Peng–Robinson Equation of State. J. Chem. Eng. Data 2014,59, 28–38. (67) Mistral AI Basic OCR. 2025; https://docs.mistral.ai/capabilities/OCR/basic_ ocr/, Accessed: 2025-08-10. 39 (68) Comanici, G.; Bieber, E.; Schaekermann, M.; Pasupat, I.; Sachdeva, N.; Dhillon, I.; Blistein, M.; Ram, O.; Zhang, D.; Rosen, E.; others Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261 2025, (69) Anthropic System Card: Claude Opus 4 & Claude Sonnet 4. 2025; https://www-cdn. anthropic.com/07b2a3f9902ee19fe39a36ca638e5ae987bc64dd.pdf. (70) Hurst, A.; Lerer, A.; Goucher, A. P.; Perelman, A.; Ramesh, A.; Clark, A.; Ostrow, A.; Welihinda, A.; Hayes, A.; Radford, A.; others Gpt-4o system card. arXiv preprint arXiv:2410.21276 2024, (71) Anthropic Claude 3.7 Sonnet System Card. 2025; https://assets.anthropic.com/ m/785e231869ea8b3b/original/claude-3-7-sonnet-system-card.pdf. (72) Opitz, J. A closer look at classification evaluation metrics and a critical reflection of common evaluation practice. transactions of the association for computational linguistics 2024,12, 820–836. (73) Eckl, B.; Huang, Y.-L.; Vrabec, J.; Hans, H. Vapor pressure of R227ea + ethanol at 343.13 K by molecular simulation. Fluid Phase Equilib. 2007,260, 177–182. (74) Abbas, S.; Nordholm, S. Simple Estimation of Bulk Thermodynamic Properties and Surface Tension of Polar Fluids. J. Colloid Interface Sci. 1995,174, 264–267. (75) Horsch, M.; Heitzig, M.; Merker, T.; Schnabel, T.; Huang, Y.-L.; Hasse, H.; Vrabec, J. Molecular Modeling of Hydrogen Bonding Fluids: Vapor-Liquid Coexistence and Interfacial Properties. High Performance Computing in Science and Engineering ’09. Berlin, Heidelberg, 2010; pp 471–483. (76) Eckelsbach, S.; Janzen, T.; K¨oster, A.; Miroshnichenko, S.; Mu˜noz-Mu˜noz, Y. M.; Vrabec, J. Molecular Models for Cyclic Alkanes and Ethyl Acetate As Well As Surface 40 Tension Data from Molecular Simulation. High Performance Computing in Science and Engineering ‘14. 2015; pp 645–659. (77) Kohns, M.; Werth, S.; Horsch, M.; von Harbou, E.; Hasse, H. Molecular simulation study of the CO2-N2O analogy. Fluid Phase Equilib. 2017,442, 44–52. (78) Eckl, B.; Vrabec, J.; Hasse, H. Set of Molecular Models Based on Quantum Mechanical Ab Initio Calculations and Thermodynamic Data. J. Phys. Chem. B 2008,112, 12710– 12721. (79) K¨oster, A.; Thol, M.; Vrabec, J. Molecular Models for the Hydrogen Age: Hydrogen, Nitrogen, Oxygen, Argon, and Water. J. Chem. Eng. Data 2018,63, 305–320. (80) Fischer, J.; Lustig, R.; Breitenfelder-Manske, H.; Lemming, W. Influence of intermolecular potential parameters on orthobaric properties of fluids consisting of spherical and linear molecules. Mol. Phys. 1984,52, 485–497. (81) Merker, T.; Engin, C.; Vrabec, J.; Hasse, H. Molecular model for carbon dioxide optimized to vapor-liquid equilibria. J. Chem. Phys. 2010,132, 234512. (82) Gao, G.; Wang, W.; Zeng, X. C. Vapor-liquid equilibria for pure HCFCHFC substances by Gibbs ensemble simulation of Stockmayer potential molecules. Fluid Phase Equilib. 1997,137, 87–98. (83) Merker, T.; Vrabec, J.; Hasse, H. Molecular simulation study on the solubility of carbon dioxide in mixtures of cyclohexane + cyclohexanone. Fluid Phase Equilib. 2012,315, 77–83. (84) Koyama, Y.; Tanaka, H.; Koga, K. On the thermodynamic stability and structural transition of clathrate hydrates. J. Chem. Phys. 2005,122, 074503. (85) Miroshnichenko, S.; Gr¨utzner, T.; Staak, D.; Vrabec, J. Molecular simulation of the 41