scieee AI-readable full text Open interactive document viewer

An LLM-based Toolbox for Automated Text Mining on the Uses of Chemicals

Xing, Huadong; Nowack, Bernd; Wang, Zhanyun; Shalin, Anna; Stadelmann, Bianca; Praetorius, Antonia

Abstract

Poster presented at SETAC Europe 35th Annual Meeting

Full text

An LLM-based Toolbox for Automated Text Mining on the Uses of Chemicals Bernd Nowack [email protected] Empa, Technology & Society Laboratory Zhanyun Wang [email protected] Empa, Technology & Society Laboratory Background Preliminary Results Huadong Xing [email protected] Empa, Technology & Society Laboratory Chemical and material technologies underpin countless industrial and consumer applications—from healthcare and agriculture to electronics and manufacturing. Yet the sheer number of chemicals in use—and the gaps in safety data for many of them—raises significant health and environmental concerns. To manage these risks, we need clear, comprehensive information on how and where chemicals are applied and when exposure might occur. Unfortunately, that information is often scattered across diverse scientific and regulatory sources, making it difficult to assemble a complete picture. This study aims to develop an automated and scalable method for extracting structured chemical use data to enhance chemicals assessment and inform regulatory decisions. To do so, we propose to use a two-stage natural language processing (NLP) framework based on large language models (see Figure 1). To better capture semantic relationships, the study avoids traditional BIObased NER models, opting instead for approaches that enable deeper contextual understanding. Objective and Methods of the Study Our NLP pipeline automatically pulls chemical–use information from raw text in two steps. First, a transformer model—fine-tuned on over 40 000 annotations—identifies chemical names. Next, a second model—built on an expert-annotated chemical-use dataset—maps each chemical to its specific application. Finally, post-processing routines standardize terminology and format the output for exposure modeling. Initial tests report high accuracy in both entity detection and chemical–use pairing, with further improvements anticipated as we enlarge our training set. Figure 2 shows the end-to-end project workflow, from raw document ingestion through annotation, model training and inference, post-processing, and finally the integration of structured chemical–use data into database. Acknowledgements We gratefully acknowledge the support of the ETH Domain Open Research Data (ORD) Measure 1 (Project: OpenChemUses), the European Union (project ZeroPM under the Horizon 2020 Research and Innovation Programme, Grant Agreement Number 101036756), and NCCR Catalysis (grant number 180544), a National Centre of Competence in Research funded by the Swiss National Science Foundation. Examples of Preliminary Results Figure 1 Anna Shalin [email protected]onto.ca University of Toronto Bianca Stadelmann [email protected] University of Amsterdam Antonia Praetorius [email protected] University of Amsterdam Prompt: Ethylene dibromide is a known alkyllead thermal stabilizer of considerable effectiveness; but when present with tetramethyllead in a mole ratio as high as 1: 1 (70 percent by weight of the dibromide based on the tetramethyllead), -it does not afford optimum protection against thermal decomposition of the tetramethyllead at elevated temperatures. Alkyllead antiknock compounds must be adequately protected against thermal decomposition during storage and shipment. Generated text (Two stages) : ethylene dibromide: alkyllead thermal stabilizer tetramethyllead: antiknock GPT-4o: ethylene dibromide: alkyllead thermal stabilizer Generated text (Without first stage) : alkyllead: thermal stabilizer Chemical-NER Annotated Model Manual and MachineBased Data Collection Llama-3.1-8B-Instruct Text with compound annotations Patents & Abstracts Text with compound-use pairs Finetuning Chemical-Use RE Annotated Model Finetuning NER - Named entity recognition RE – Relation Extraction Finetuning – LoRA (Low-Rank Adaptation) finetuning Large Language Model Data extraction Figure 2