scieee AI-readable full text Open interactive document viewer

From Idea to Prototype: Using BITS and LLMs to automate the annotation process for SGN Collection Data

Wolodkin, Alexander; Martens, Claudia

Abstract

We will present a workflow at SGN that combines the usage of BITS (https://projects.tib.eu/bits/home) outcome (i.e. ESS collection of the TIB TS) with GPT4all in order to identify gaps in terminologies on the one hand, and provide assistance to scientists, who are working on new collections on the other hand. Based on two major data management challenges facing SGN, Legacy Data Digitisation (historical grown data require systematic transformation into machine-readable formats) and Data Proliferation Management (continuous input of data generated by ongoing collection efforts and research activities), our prototyping process can be divided into several areas: Identifying nominal phrases (NPs) in the collection data and annotating them using BITS TS. Our primary goal was to achieve reliable detection, with a focus on minimising false negatives, while accepting some false positives during annotation. During the prototyping phase, several obstacles were encountered referring to poor NP detection quality in scientific texts and a lack of reliability in conjunction splitting and singularization using common tools. It is also not always possible to determine the correct language of the text, especially with mixed-language content. Revising our requirements had let us choose GPT4all as our preferred solution, specifically the Meta-Llama-3-8B-Instruct.Q4_0.gguf model. This allows us to perform high quality NP detection and transformation, but with very high computational and time requirements. To optimise resource utilisation, GPT4all is employed only for high-level operations. Other operations can be performed by tools with less hardware requirements. Using statistical logging allows us to identify various significant information about the NP detection and usage. This data we can reuse in later development steps. By leveraging the strengths of BITS and GPT4all, SGN is paving the way for more accurate processing of complex scientific data to improve research outcomes.

Full text

●A total of more than 200 terminologies, ●More than 1,097,000 terms ●Collection of 38 terminologies relevant to the ESS From Idea to Prototype: Using BITS and LLMs to automate the annotation process for SGN Collection Data // Current Challenges DFG project number 508107981 Senckenberg Collections: ●More than 40 million physical specimen ●Currently: 1,547,711 digital specimen ○in 124 collections ●Use of mixed languages ●Use of taxonomic structures in object descriptions ●Use of labels and various image data ●Plain text without linked data Example text entry: Minerals: 2. Sulfides and Sulfosalts: 2.C: Metal Sulfides, M: S = 1: 1 (and similar): 2.CD.: 2.CD.10: Galenit The Challenge: Identify annotatable chunks of text, or, in other words, recognize noun groups, subgroups, and relevant single nouns across various languages. // Next Challenges SGN's Collection Data includes not only texts and descriptions, but also a wide range of associated media content, from photographs and scans to complex 3D models. These contents provide us with a wealth of insights that can be made accessible through appropriate semantically annotated textual capture. This process requires the use of suitable multimodal AI models, which are currently being tested and will be extended in the further work. Privacy Policy vs. Resource Limitations // Current Experiences ●LLM Size matters in many cases ○On the one hand, a larger model usually has a better understanding of the tasks, needs and expectations ○On the other hand, significantly higher resource requirements, particularly in RAM and time, necessitate greater investment ○At the same time, some lightweight models below 32B struggle to follow task specifications and produce well-intentioned but useless output ●Important to understand: The structure of noun phrases in a sentence is not strictly deterministic! ○Biggest models may have different ideas in various working sessions ○Simple pre-processing and suggestions, such as pre-selected NP combinations from SpaCy, can enhance accuracy in the following steps ●Too much vs. too little specialisation ○Too much specialisation of a model increases the overall effort. Too little reduces effectiveness ○Lack of training resources for a specific model such as BERT and similar variants ●Find the sweet spot: ○As small as possible, but able to perform tasks without hallucination and with reasonable post-processing effort ○Modern technology (LLM Transformer) with prospects for continuous further development ○MoE as an extension of the Transformer: Replacing Feed-Forward Networks by (a subset of) multiple experts Authors: Alexander Wolodkin https://orcid.org/0000-0003-1556-8750 Alexander.W[email protected] Claudia Martens https://orcid.org/0000-0003-2478-4295 [email protected] Contact BITS: [email protected] Image created by AI Image partially created by AI AI-created image based on the current evaluation process Within BITS, a Terminology Service (TS) will be established for subfields of ESS // Digital Collection Workflow AI-created image based on the current workflow process