scieee AI-readable full text Open interactive document viewer

Knowledge Representation and Ontologies in the Era of Large Language Models

Mendes de Farias, Tarcisio

Abstract

With the advent of large language models (LLMs), we are currently able to build applications with a relatively good accuracy that can process and generate “human”-like text, engage in natural language conversations, and perform several types of tasks based on users’ prompts. Nevertheless, it is well known that LLMs and their applications can have several limitations among them: hallucinations, indecisiveness, a “black-box” approach and lacking domain-specific knowledge. To address these limitations and more, the knowledge representation with ontologies can play a key role by making it possible to use and enhance LLMs for challenging applications that demand a high accuracy and verifiable outcomes. One of these applications is scientific question answering. In this talk, I will present different strategies to address the LLM limitations such as fine- and prompt-tuning, discuss effective relations between ontologies and LLMs, and our latest work on building a question answering system over scientific data in the form of knowledge graphs. In this recent work, we also seek to leverage the ontologies used to build these knowledge graphs.

Full text

Knowledge Representation and Ontologies in the Era of Large Language Models Tarcisio Mendes de Farias Knowledge Representation Unit [email protected] SIB in brief A national network of about 900 scientists A non-profit and independent organization 190 employees in 4 locations The Swiss Node of ELIXIR, the European life science infrastructure Knowledge Representation in AI To begin with, a definition… A brief history of bioinformatics databases… Gene Expression Orthology Protein Interaction FAIR data principles Findable, Accessible, Interoperable, Reusable Importance of identifiers What is a Large Language Model? A (deep learning) model trained to predict the next word in a sentence How is it trained? •Self-supervised learning on HUGE amounts of text •“Fill in the blanks” learning Source: https://amitness.com 7 Generalization ●A (deep learning) model trained to predict the next word in a sentence ○Token = words / amino-acids / genes / …. ●What is a Language? ○A vocabulary + ○Sequences of tokens that represent information in that language token sequence 8 Let’s ask an LLM about TIME FOR COFFEE! How do we extract information from a Knowledge Graph? What are the human genes involved in lung cancer with an ortholog expressed in the mouse? SELECT ?gene ?orthologous_protein2 WHERE { SELECT * { SERVICE <http://sparql.uniprot.org/sparql> { SELECT ?protein1 WHERE { ?protein1 a up:Protein; up:organism/up:scientificName 'Homo sapiens' ; up:annotation ?annotation . ?annotation rdfs:comment ?annotation_text. ?annotation a up:Disease_Annotation . FILTER CONTAINS (?annotation_text, ”lung cancer") } } SERVICE <https://sparql.omabrowser.org/sparql/> { SELECT ?orthologous_protein2 ?protein1 ?gene WHERE { ?protein_OMA a orth:Protein . ?orthologous_protein2 a orth:Protein . ?cluster a orth:OrthologsCluster . ?cluster orth:hasHomologousMember ?node1 . ?cluster orth:hasHomologousMember ?node2 . ?node2 orth:hasHomologousMember* […….] FILTER(?node1 != ?node2) } } SERVICE <https://bgee.org/sparql/> { ?gene genex:isExpressedIn ?anatEntity . ?anatEntity rdfs:label ‘lung' . ?gene orth:organism ?org . ?org obo:RO_0002162 taxon:10090 .} LLMs as interfaces for scientific knowledge graph exploration: a fine-tuning approach Fine-tuning LLMs for SPARQL generation •Joint work with the Data Knowledge Organisation Unit •Dr. Norio Kobayashi, Dr. Julio Rangel Reyes •RIKEN, Japan •Challenge: Very little training data •Databases often propose a few examples of queries online •Some queries are hard to connect back to the question •Solution: Augment and semantically enrich existing dataset of examples automatically! In which taxa is the insulin protein present? SWAT4HCLS 2024, full paper available at https://ceur-ws.org/Vol-3890/paper-4.pdf Fine-tuning an Open LLM for SPARQL generation Evaluation Experiments’ setup Hugging Face SFTTrainer with 2000 steps Nvidia A100 40GB GPUs OpenLLaMA_7b_v2 (7 billion parameters) Temperature parameter equal to zero Knowledge base (KB): Bgee (~7 billion triples) We rely on four different metrics designed for the evaluation of machine translation output : BLEU, SP-BLEU, METEOR and ROUGE. Four main experiments and preliminary results Zero-shot evaluation against: Wikidata and Bgee question-SPARQL query sets All metrics including F1-score: either equal or approximately equal to zero. 3rd experiment 2nd experiment Discussions Systematically augmenting a representative question-to-SPARQL query set over a scientific KG significantly contributes to improving the performance of the OpenLLaMA model for the SPARQL query generation task. Rewriting the SPARQL query to provide more context through comments and meaningful variable names considerably improves OpenLLaMA. Knowledge transfer might deteriorate the LLM performance for the SPARQL query generation over a domain-specific knowledge base. LLMs as interfaces for scientific knowledge graph exploration: a prompt-tuning with RAG approach Problem We need a minimal structured way to describe the SPARQL endpoints’ contents to facilitate LLM-based SPARQL query generation. For large and complex knowledge graphs finding the right context is not trivial. Writing SPARQL queries is hard and time-consuming. LLMs are great at it, but they need context! Generic and reusable methodology that works for most endpoints. •Automatically generate a description of used classes and predicates e.g., Protein isEncodedBy Gene •Provide example queries with their corresponding question in a standard format Make them accessible via the SPARQL endpoint The system will be able to retrieve, index, and use them automatically Properly describing an endpoint with some metadata LLM-based SPARQL Query Generation from Natural Language over Federated Knowledge Graphs, V. Emonet et al, ISWC 2024