Data Mining and Data Warehouses SiKDD 2025
Full text
6. oktober 2025 l 6 October 2025 Ljubljana, Slovenia IS 2025 Odkrivanje znanja in podatkovna skladišča SiKDD Data Mining and Data Warehouses SiKDD Urednika l Editors: Dunja Mladenić, Marko Grobelnik INFORMACIJSKA DRUZBA INFORMATION SOCIETY Zbornik 28. mednarodne multikonference Zvezek C Proceedings of the 28th International Multiconference Volume C ˇ
Zbornik 28. mednarodne multikonference INFORMACIJSKA DRUŽBA – IS 2025 Zvezek C Proceedings of the 28th International Multiconference INFORMATION SOCIETY – IS 2025 Volume C Odkrivanje znanja in podatkovna skladišča - SiKDD Data Mining and Data Warehouses - SiKDD Urednika / Editors Dunja Mladenić, Marko Grobelnik http://is.ijs.si 6. oktober 2025 / 6 October 2025 Ljubljana, Slovenia
Urednika: Dunja Mladenić Department for Artificial Intelligence, Jožef Stefan Institute, Ljubljana Marko Grobelnik Department for Artificial Intelligence, Jožef Stefan Institute, Ljubljana Založnik: Institut »Jožef Stefan«, Ljubljana Priprava zbornika: Mitja Lasič, Vesna Lasič, Lana Zemljak Oblikovanje naslovnice: Vesna Lasič, uporabljena slika iz Pixabay Dostop do e-publikacije: http://library.ijs.si/Stacks/Proceedings/InformationSociety Ljubljana, oktober 2025 Informacijska družba ISSN 2630-371X DOI: https://doi.org/10.70314/is.2025.sikdd Kataložni zapis o publikaciji (CIP) pripravili v Narodni in univerzitetni knjižnici v Ljubljani COBISS.SI-ID 255453699 ISBN 978-961-264-322-5 (PDF)
PREDGOVOR MULTIKONFERENCI INFORMACIJSKA DRUŽBA 2025 28. mednarodna multikonferenca Informacijska družba se odvija v času izjemne rasti umetne inteligence, njenih aplikacij in vplivov na človeštvo. Vsako leto vstopamo v novo dobo, v kateri generativna umetna inteligenca ter drugi inovativni pristopi oblikujejo poti k superinteligenci in singularnosti, ki bosta krojili prihodnost človeške civilizacije. Naša konferenca je tako hkrati tradicionalna znanstvena in akademsko odprta, pa tudi inkubator novih, pogumnih idej in pogledov. Letošnja konferenca poleg umetne inteligence vključuje tudi razprave o perečih temah današnjega časa: ohranjanje okolja, demografski izzivi, zdravstvo in preobrazba družbenih struktur. Razvoj UI ponuja rešitve za številne sodobne izzive, kar poudarja pomen sodelovanja med raziskovalci, strokovnjaki in odločevalci pri oblikovanju trajnostnih strategij. Zavedamo se, da živimo v obdobju velikih sprememb, kjer je ključno, da z inovativnimi pristopi in poglobljenim znanjem ustvarimo informacijsko družbo, ki bo varna, vključujoča in trajnostna. V okviru multikonference smo letos združili dvanajst vsebinsko raznolikih srečanj, ki odražajo širino in globino informacijskih ved: od umetne inteligence v zdravstvu, demografskih in družinskih analiz, digitalne preobrazbe zdravstvene nege ter digitalne vključenosti v informacijski družbi, do raziskav na področju kognitivne znanosti, zdrave dolgoživosti ter vzgoje in izobraževanja v informacijski družbi. Pridružujejo se konference o legendah računalništva in informatike, prenosu tehnologij, mitih in resnicah o varovanju okolja, odkrivanju znanja in podatkovnih skladiščih ter seveda Slovenska konferenca o umetni inteligenci. Poleg referatov bodo okrogle mize in delavnice omogočile poglobljeno izmenjavo mnenj, ki bo pomembno prispevala k oblikovanju prihodnje informacijske družbe. »Legende računalništva in informatike« predstavljajo domači »Hall of Fame« za izjemne posameznike s tega področja. Še naprej bomo spodbujali raziskovanje in razvoj, odličnost in sodelovanje; razširjeni referati bodo objavljeni v reviji Informatica, s podporo dolgoletne tradicije in v sodelovanju z akademskimi institucijami ter strokovnimi združenji, kot so ACM Slovenija, SLAIS, Slovensko društvo Informatika in Inženirska akademija Slovenije. Vsako leto izberemo najbolj izstopajoče dosežke. Letos je nagrado Michie-Turing za izjemen življenjski prispevek k razvoju in promociji informacijske družbe prejel Niko Schlamberger, priznanje za raziskovalni dosežek leta pa Tome Eftimov. »Informacijsko limono« za najmanj primerno informacijsko tematiko je prejela odsotnost obveznega pouka računalništva v osnovnih šolah. »Informacijsko jagodo« za najboljši sistem ali storitev v letih 2024/2025 pa so prejeli Marko Robnik Šikonja, Domen Vreš in Simon Krek s skupino za slovenski veliki jezikovni model GAMS. Iskrene čestitke vsem nagrajencem! Naša vizija ostaja jasna: prepoznati, izkoristiti in oblikovati priložnosti, ki jih prinaša digitalna preobrazba, ter ustvariti informacijsko družbo, ki koristi vsem njenim članom. Vsem sodelujočim se zahvaljujemo za njihov prispevek — veseli nas, da bomo skupaj oblikovali prihodnje dosežke, ki jih bo soustvarjala ta konferenca. Mojca Ciglarič, predsednica programskega odbora Matjaž Gams, predsednik organizacijskega odbora i
FOREWORD TO THE MULTICONFERENCE INFORMATION SOCIETY 2025 The 28th International Multiconference on the Information Society takes place at a time of remarkable growth in artificial intelligence, its applications, and its impact on humanity. Each year we enter a new era in which generative AI and other innovative approaches shape the path toward superintelligence and singularity — phenomena that will shape the future of human civilization. The conference is both a traditional scientific forum and an academically open incubator for new, bold ideas and perspectives. In addition to artificial intelligence, this year’s conference addresses other pressing issues of our time: environmental preservation, demographic challenges, healthcare, and the transformation of social structures. The rapid development of AI offers potential solutions to many of today’s challenges and highlights the importance of collaboration among researchers, experts, and policymakers in designing sustainable strategies. We are acutely aware that we live in an era of profound change, where innovative approaches and deep knowledge are essential to creating an information society that is safe, inclusive, and sustainable. This year’s multiconference brings together twelve thematically diverse meetings reflecting the breadth and depth of the information sciences: from artificial intelligence in healthcare, demographic and family studies, and the digital transformation of nursing and digital inclusion, to research in cognitive science, healthy longevity, and education in the information society. Additional conferences include Legends of Computing and Informatics, Technology Transfer, Myths and Truths of Environmental Protection, Knowledge Discovery and Data Warehouses, and, of course, the Slovenian Conference on Artificial Intelligence. Alongside scientific papers, round tables and workshops will provide opportunities for in-depth exchanges of views, making an important contribution to shaping the future information society. Legends of Computing and Informatics serves as a national »Hall of Fame« honoring outstanding individuals in the field. We will continue to promote research and development, excellence, and collaboration. Extended papers will be published in the journal Informatica, supported by a long-standing tradition and in cooperation with academic institutions and professional associations such as ACM Slovenia, SLAIS, the Slovenian Society Informatika, and the Slovenian Academy of Engineering. Each year we recognize the most distinguished achievements. In 2025, the Michie-Turing Award for lifetime contribution to the development and promotion of the information society was awarded to Niko Schlamberger, while the Award for Research Achievement of the Year went to Tome Eftimov. The »Information Lemon« for the least appropriate information-related topic was awarded to the absence of compulsory computer science education in primary schools. The »Information Strawberry« for the best system or service in 2024/2025 was awarded to Marko Robnik Šikonja, Domen Vreš and Simon Krek together with their team, for developing the Slovenian large language model GAMS. We extend our warmest congratulations to all awardees. Our vision remains clear: to identify, seize, and shape the opportunities offered by digital transformation, and to create an information society that benefits all its members. We sincerely thank all participants for their contributions and look forward to jointly shaping the future achievements that this conference will help bring about. Mojca Ciglarič, Chair of the Program Committee Matjaž Gams, Chair of the Organizing Committee ii
KONFERENČNI ODBORI CONFERENCE COMMITTEES International Programme Committee Organizing Committee Vladimir Bajic, South Africa Heiner Benking, Germany Se Woo Cheon, South Korea Howie Firth, UK Olga Fomichova, Russia Vladimir Fomichov, Russia Vesna Hljuz Dobric, Croatia Alfred Inselberg, Israel Jay Liebowitz, USA Huan Liu, Singapore Henz Martin, Germany Marcin Paprzycki, USA Claude Sammut, Australia Jiri Wiedermann, Czech Republic Xindong Wu, USA Yiming Ye, USA Ning Zhong, USA Wray Buntine, Australia Bezalel Gavish, USA Gal A. Kaminka, Israel Mike Bain, Australia Michela Milano, Italy Derong Liu, Chicago, USA Toby Walsh, Australia Sergio Campos-Cordobes, Spain Shabnam Farahmand, Finland Sergio Crovella, Italy Matjaž Gams, chair Mitja Luštrek Lana Zemljak Vesna Koricki Mitja Lasič Blaž Mahnič Programme Committee Mojca Ciglarič, chair Bojan Orel Franc Solina Viljan Mahnič Cene Bavec Tomaž Kalin Jozsef Györkös Tadej Bajd Jaroslav Berce Mojca Bernik Marko Bohanec Ivan Bratko Andrej Brodnik Dušan Caf Saša Divjak Tomaž Erjavec Bogdan Filipič Andrej Gams Matjaž Gams Mitja Luštrek Marko Grobelnik Nikola Guid Marjan Heričko Borka Jerman Blažič Džonova Gorazd Kandus Urban Kordeš Marjan Krisper Andrej Kuščer Jadran Lenarčič Borut Likar Janez Malačič Olga Markič Dunja Mladenič Franc Novak Vladislav Rajkovič Grega Repovš Ivan Rozman Niko Schlamberger Gašper Slapničar Stanko Strmčnik Jurij Šilc Jurij Tasič Denis Trček Andrej Ule Boštjan Vilfan Baldomir Zajc Blaž Zupan Boris Žemva Leon Žlajpah Niko Zimic Rok Piltaver Toma Strle Tine Kolenik Franci Pivec Uroš Rajkovič Borut Batagelj Tomaž Ogrin Aleš Ude Bojan Blažica Matjaž Kljun Robert Blatnik Erik Dovgan Špela Stres Anton Gradišek iii
iv
KAZALO / TABLE OF CONTENTS Odkrivanje znanja in podatkovna skladišča – SiKDD / Data Mining and Data Warehouses - SiKDD .... 1 PREDGOVOR / FOREWORD ............................................................................................................................... 3 PROGRAMSKI ODBORI / PROGRAMME COMMITTEES ............................................................................... 5 Semantic Prompting for Large Language Models in Biomedical Named Entity Recognition / Calcina Erik, Novak Erik, Mladenić Dunja.............................................................................................................................. 7 LLM Based Approach to Extracting Smells in Slovenian Corpora / Brank Janez, Novalija Inna, Mladenić Dunja, Grobelnik Marko .................................................................................................................................. 11 BetweenTheLines - Cross Source News Analysis / Trajkov Georgi, Grobelnik Marko, Grobelnik Adrian Mladenić ........................................................................................................................................................... 15 Identifying Social Self in Text: A Machine Learning Study / Caporusso Jaya, Purver Matthew, Pollak Senja .. 19 WinWin Meets – Investigating the Future of Online Meetings / Žust Martin, Grobelnik Marko, Guček Alenka, Grobelnik Adrian Mladenić.............................................................................................................................. 25 Predicting Ski Jumps Using State-Space Model / Hegler Živa, Camlek Neca, Jelenčič Jakob, Grobelnik Marko, Mladenić Dunja ................................................................................................................................................ 29 Predicting milling overload based on sensor data: a graph-based approach / Krumpak Roy, Rožanec Jože M., Mladenić Dunja, Guo Zhenyu, Song Tao, Roman Dumitru, Novalija Inna, Ma Xiang ................................... 33 Short and Long Term Bike Rental Forecasting / Kocjančič Oskar, Žnidaršič Martin ......................................... 37 Predicting Traffic Intensity on Motorway Sections / Kladnik Matic, Mladenić Dunja ....................................... 41 Empowering Youth for Smart Cities with AI Solutions to Community and Urban Challenges in the Context of SDG 11 / Zaouini Mustafa, Costa João Pita, Rahmani Yousef, Kassis Rayan, Stopar Luka, Souss Sohaib, Lamgari Asmai, Mochariq Ouidad ................................................................................................................... 45 Automated First-Reply Generation for IT Support Tickets Using Retrieval-Augmented Generation and MultiModal Response Synthesis / Jeršek Domen, Kenda Klemen, Frattini Matteo, Klančič Rok .......................... 49 A Machine-Learning Approach to Predicting the Pronunciation of Pre-Consonant l in Standard Slovene / Čibej Jaka ................................................................................................................................................................... 53 Sequencing News Articles with Large Language Models within Enterprise Risk Management Context / Debeljak Žiga, Mladenić Dunja, Kenda Klemen ............................................................................................. 57 Graph-Based Feature Engineering for DeFi Security Incident Severity Prediction / Pavlova Daria, Novalija Inna, Mladenić Dunja ....................................................................................................................................... 61 Evolving Neural Agents in Simulated Ecosystems / Ćetković Marija, Tošić Aleksandar, Vake Domen ............ 65 Designing AI Agents for Social Media / Sittar Abdul, Smiljanic Mateja, Guček Alenka ................................... 69 Explaining Temporal Data in Manufacturing using LLMs and Markov Chains / Šturm Jan, Škrjanc Maja, Topal Oleksandra, Novalija Inna, Mladenić Dunja, Grobelnik Marko ...................................................................... 73 Active Learning for Power Grid Security Assessment: Reducing Simulation Cost with Informative Sampling / Leskovec Gašper, Mylonas Costas, Kenda Klemen ......................................................................................... 77 Supporting Material Reuse in Drone Production / Cek Rok, Topal Oleksandra, Leonardi Linda, Forcolin Margherita, Kenda Klemen .............................................................................................................................. 82 Temporal Dynamics and Causal Feature Integration for Predictive Maintenance in Manufacturing Systems: A Causality-Informed Framework / Hosseini Seyed Iman, Kenda Klemen, Mladenić Dunja ........................... 86 Using Interactive Data Visualization for DeFi Market Analysis / Pavlova Daria ................................................ 90 A Hybrid Lexicon-Machine Learning Approach to Macedonian Sentiment Analysis / Kochovska Sofija, Kavšek Branko, Vičič Jernej ............................................................................................................................ 94 Building an AI-Ready Data Infrastructure Towards a SDG-focused Observatory for the Brazilian Amazon / Costa João Pita, Polzer Mirozlav, Barrionuevo Leonardo, Veiga João Cândia ............................................... 98 Towards a format for describing networks, NetsJSON / Batagelj Vladimir, Pisanski Tomaž, Savnik Iztok, Slavec Ana, Bašić Nino .................................................................................................................................. 102 Automating Numba Optimization with Large Language Models: A Case Study on Mutual Information / Kozamernik Lučka, Jakomin Martin, Škrlj Blaž, Urbančič Jasna.................................................................. 106 Topological Exploration of Embedded GitHub Repository Data Using Mapper / Hrib Ivo, Zajec Patrik ......... 110 CO2 Monitoring for Energy-Efficient Workloads in Kubernetes: A Data Provider for CO2-Aware Migration / Hrib Ivo, Topal Oleksandra, Šturm Jan, Škrjanc Maja .................................................................................. 114 v
6
Semantic Prompting for Large Language Models in Biomedical Named Entity Recognition Erik Calcina Jožef Stefan Institute Jožef Stefan International Postgraduate School Jamova cesta 39 Ljubljana, Slovenia Erik Novak Jožef Stefan Institute Jožef Stefan International Postgraduate School Jamova cesta 39 Ljubljana, Slovenia Dunja Mladenić Jožef Stefan Institute Jožef Stefan International Postgraduate School Jamova cesta 39 Ljubljana, Slovenia Abstract Extracting structured medical information from unstructured clinical text remains a challenge for biomedical research and decision support. Recent advances in large language models (LLMs) suggest that prompt-based methods could provide a promising alternative to traditional supervised approaches for Named Entity Recognition (NER) in the biomedical domain. This study investigates whether adding semantic descriptions of entity labels can improve NER performance on clinical texts. Using a dataset of annotated case reports, we evaluate model performance in zero-shot, few-shot, and fine-tuned settings. Results show that semantic prompts enhance accuracy in low-supervision scenarios, while offering limited benefit once models are fine-tuned. Keywords Named entity recognition, large language models, semantic prompting, prompt engineering, medical domain, biomedicine 1 Introduction Biomedical texts present a critical challenge for automated analysis. Clinical case reports, patient records, and related narratives are written in free text rather than in structured formats. While they contain essential medical knowledge, their unstructured nature makes it necessary to extract and organize information for systematic use in research and clinical decision support. Doing this manually is costly, time-consuming, and challenging to scale. Therefore, an automated approach to extract relevant information is required. Named entity recognition (NER) models enable the identification and classification of clinically relevant entities, such as biological structures, diagnostic procedures, or symptoms. Recent advances in large language models (LLMs) show strong generalizing abilities, identifying relevant entities in both zeroshot and few-shot settings. However, in the biomedical domain, performance can be hindered by specialized terminology and subtle entity distinctions. To address this, we propose enriching prompts with semantic descriptions of entity labels, providing models with explicit context to improve their understanding of the task. This study investigates the impact of semantically enhanced prompting in biomedical named entity recognition using large language models. We evaluate the effect of enriching entity labels Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia ©2024 Copyright held by the owner/author(s). https://doi.org/10.70314/is.2025.sikdd.3 with semantic descriptions on model performance across zeroshot, few-shot, and fine-tuned scenarios, using the MACCROBAT2020 dataset [3]. The contributions of this paper are threefold. First, we introduce the use of semantically enhanced prompts for biomedical NER by enriching entity labels with descriptions. Second, we provide a systematic evaluation of semantic prompting across zero-shot, few-shot, and fine-tuned scenarios, assessing its effectiveness under different levels of supervision. Third, we apply a statistical validation method, McNemar’s test, to rigorously assess the reliability of observed performance differences between baseline and semantically enhanced prompts. The remainder of the paper is structured as follows: Section 2 contains the overview of the related work. Next, we present the methodology in Section 3, and describe the experiment setting in Section 4. The experiment results are found in Section 5, followed by a discussion in Section 6. Finally, we conclude the paper and provide ideas for future work in Section 7. 2 Related Work This section focuses on the related work on named entity recognition in biomedicine, as well as the use of semantic descriptions in prompting. 2.1 Prompting with semantic context PromptNER introduced the idea of augmenting few-shot prompts with entity definitions, leading to substantial gains in F1 score on benchmarks like CoNLL, GENIA, and FewNERD, improving performance by 4–9 points compared to standard prompting [2]. Extending this idea, PromptNER unifies locating and typing into a single enriched prompt, enabling phrase extraction and entity classification simultaneously [7]. Similarly, the biomedical NER study demonstrated that “on-the-fly” inclusion of concept definitions enhances performance (+15% F1) in low-data settings [5]. 2.2 Iterative and zero-shot semantic prompting Recent work in zero-shot NER explores iterative prompt refinement to align model outputs with precise entity definitions. EvoPrompt uses an evolving definition-based framework to better distinguish between similar entity types, yielding improvements across benchmarks [9]. In a broader context, some studies found that while directly injecting semantic parses into LLM inputs can degrade performance, carefully designed semantic “hints” embedded in prompts can reliably boost outcomes [1]. 2.3 Domain-specific prompt optimization FsPONER optimizes few - shot prompts for industrial NER tasks by using semantic entity–enhanced meta prompts and task - specific exemplar selection, yielding F1 improvements of 5 to 13 points 7
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Calcina et al. in domain benchmarks [8]. In the biomedical domain, MPE3integrates ontology-derived label semantics into prompts, improving performance in few-shot NER scenarios [10]. Prior research has shown that enriching prompts with semantic context and label definitions can significantly boost LLM performance in both few-shot and zero-shot NER. Our work provides a systematic evaluation in the biomedical domain. By examining multiple supervision settings, benchmarking several model families, and validating differences through McNemar’s test, we offer a comprehensive assessment of when semantically enriched prompts provide benefits. 3 Methodology This study evaluates the impact of incorporating semantic information into prompts on the performance of LLMs in biomedical NER tasks. Three distinct approaches were employed: zero-shot prompting, few-shot prompting, and fine-tuning. Zero-shot prompting. In the zero-shot setting, models were prompted to perform NER without any prior exposure to labeled examples. Two types of prompts were utilized: baseline prompt, a standard instruction to identify and classify entities without additional context, and semantically enhanced prompt, which includes detailed descriptions for each entity label, offering explicit semantic context to guide the model’s understanding and classification. Few-shot prompting. The few-shot approach involved providing the models with a limited number of annotated examples (k-shots) before performing NER on new texts. Similar to the zeroshot setting, both baseline and semantically enhanced prompts were employed to assess the influence of semantic information. Fine-tuning. Fine-tuning was conducted to adapt the pre-trained LLMs to the specific biomedical NER task. Two fine-tuning strategies were explored: standard fine-tuning, where models are finetuned using the original dataset annotations without additional semantic information, and semantically enhanced fine-tuning, which fine-tunes models on data where annotations were supplemented with semantic descriptions of each entity label. 4 Experiment Setting This section describes the experiment setting, which includes the dataset and prompt preparation, the fine-tuning procedure used, the evaluation metrics, and the statistical significance test description. 4.1 Dataset The experiments were conducted using the MACCROBAT2020 [3] dataset, which comprises 200 clinical case reports sourced from PubMed Central. In total, it contains 4,542 sentences with an average of 22.7 sentences per document, which includes manual annotations of biomedical entities, events, and relations, provided in brat standoff format 1 . For this study, we focused on the five most frequent entity labels within the dataset. These are biological structure,diagnostic procedure,lab value,sign symptom, and detailed description supplemented by the age and sex labels. The inclusion of age and sex was motivated by their prevalence and clarity within clinical narratives, providing 1https://brat.nlplab.org/standoff.html a basis for evaluating model performance on both complex and straightforward entity types. Each document was segmented into individual sentences by splitting on full stops. Subsequently, each sentence, along with its associated entity annotations, was transformed into a JSON format to facilitate processing by the language models. 4.2 Semantically enhanced prompts To enhance the semantic understanding of entity labels, detailed descriptions were crafted for each. These descriptions were derived by combining information from the MACCROBAT2020 dataset documentation and definitions from the Oxford English Dictionary [6]. The integration of these sources was performed manually, ensuring that the descriptions were both accurate and contextually relevant. Prompts were structured as plain text instructions, guiding the model to identify and classify entities within the provided sentences. For the semantically enhanced prompts, the detailed entity descriptions were included to provide additional context. Models were instructed to output their responses in a JSON format, explicitly focusing on the labels component. Below we present an example of the entity description, specifically for the label age. Baseline prompt: The age of the patient. Semantic enhanced prompt: The duration of time a patient has lived, expressed numerically (e.g., ‘65year-old’, ‘20 years old’) or categorically (e.g., ‘newborn’, ‘teenage’), representing their age at the time of presentation. This added context is intended to improve the model’s ability to distinguish and extract nuanced biomedical entities more accurately. 4.3 Fine-tuning procedure Fine-tuning is carried out using parameter-efficient techniques, where only lightweight adapter modules are trained instead of modifying the full model. This strategy reduces memory usage, mitigates catastrophic forgetting, and accelerates training. To further improve efficiency, models are quantized to 4-bit precision. Fine-tuning is supervised and focuses on the generated outputs; all non-target tokens (e.g., system prompts, input context) are masked during loss computation. This ensures that training adapts the model to the expected JSON label output format rather than to the input content or prompt structure. 4.4 Evaluation metrics To evaluate entity recognition performance, we use two F1-based metrics. The Exact F1 score measures strict matches, requiring predicted entities to align perfectly with the reference text and label. The Relaxed F1 score allows partial matches, counting predictions as correct if they include the true entity as a substring with the correct label. 4.5 McNemar statistical significance test While Exact and Relaxed F1 scores quantify the magnitude of performance differences, they do not establish whether these differences are statistically reliable. The McNemar test [4] complements the Exact F1 metric by verifying whether observed improvements can be attributed to the semantically enhanced 8
Semantic Prompting for Large Language Models in Biomedical Named Entity Recognition Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia prompts rather than random variation. Following standard NER practice, we treat Exact F1 as the primary endpoint and therefore apply McNemar’s test only to exact match predictions. Let 𝑏 denote the number of cases correctly predicted by the semantically enhanced model but missed by the baseline, and 𝑐 the number of cases correctly predicted by the baseline but missed by the semantically enhanced model. Only discordant pairs (𝑏, 𝑐) contribute to the test; agreements do not affect the statistic. Using the continuity-corrected version of the test, the statistic is computed as 𝜒2= (|𝑏−𝑐| − 1)2 𝑏+𝑐, which follows a chi-squared distribution with one degree of freedom. The corresponding 𝑝 -value allows us to test the null hypothesis 𝐻0 : the two models have equal marginal probabilities (i.e., performance differences are due to chance). Conventionally, 𝑝<0.01 is considered statistically significant. 5 Results This section presents model performance under three experimental conditions: zero-shot, few-shot, and fine-tuned prompting. For each condition, we compare the impact of semantically enhanced prompts against standard prompts using Exact and Relaxed F1 scores on a subset of clinically relevant entity types. 5.1 Zero-shot prompting Table 1 reports the Exact and Relaxed F1 scores for models evaluated in the zero-shot setting using semantically enhanced prompts. Without semantic descriptions, most models struggled to generate outputs in the required JSON format, and valid scores could not be computed. Even with semantically enhanced prompts, Meta-Llama-3.1-8B consistently failed to produce structured responses. Among the evaluated models, Llama-3.1-8B-Instruct achieved the highest Exact F1 score, while txgemma-9b-chat attained the best Relaxed F1 score. Llama-3.2-3B-Instruct and DeepSeekQwen-7B also demonstrated non-trivial performance in both metrics. These results suggest that semantically enhanced prompts can effectively compensate for the absence of training examples in zero-shot scenarios by providing clearer task guidance and improving structured prediction output. Table 1: Exact and Relaxed F1 scores in the zero-shot setting with semantically enhanced prompts. Bolded values indicate the highest score in each column. Results without valid JSON output are marked with /. Model Exact F1 Semantics Relaxed F1 Semantics Llama-3.1-8B-Instruct20.2310 0.3708 Meta-Llama-3.1-8B3/ / Llama-3.2-3B-Instruct40.1620 0.3254 DeepSeek-Qwen-7B50.1592 0.3217 txgemma-9b-chat60.2181 0.4245 2https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct 3https://huggingface.co/meta-llama/Llama-3.1-8B 4https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct 5https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B 6https://huggingface.co/google/txgemma-9b-chat 5.2 Few-shot prompting Table 2 summarizes the Exact and Relaxed F1 scores for fewshot prompting. The addition of semantic information consistently improved model performance across most models. Notably, txgemma-9b-chat achieved the highest Exact F1 score 0.3288 and Relaxed F1 score 0.4998 with semantic prompting, compared to 0.2732 and 0.4469 without. Both Llama-3.1-8B-Instruct and Llama-3.2-3B-Instruct showed improvements in both Exact and Relaxed F1 scores when provided with semantically enhanced prompts. For instance, Llama-3.1-8B-Instruct improved from 0.2509 to 0.3005 (Exact) and from 0.3526 to 0.3948 (Relaxed), while Llama-3.2-3BInstruct increased from 0.2300 to 0.2439 (Exact) and from 0.3769 to 0.3948 (Relaxed). These gains highlight the benefit of enriching prompt instructions when training data is limited. However, not all models responded positively. For example, Meta-Llama-3.18B experienced a drop in Exact F1 from 0.2698 to 0.2210 and in Relaxed F1 from 0.3537 to 0.2799, indicating that semantically enhanced prompts do not universally improve performance and may be less effective for some models. To assess the reliability of these differences, we conducted McNemar tests on Exact paired predictions. The tests revealed that performance differences between baseline and semantically enhanced prompts were statistically significant for all models except Llama-3.2-3B-Instruct. It is important to note, however, that significance here indicates that the two variants produce systematically different predictions, but does not itself imply improvement. For instance, while the difference for Meta-Llama3.1-8B was highly significant, the semantically enhanced model in fact performed worse in terms of F1 scores. 5.3 Fine-tuned performance In the fine-tuning scenario, results were more nuanced. As shown in Tables 2, most models performed strongly even without semantic enhancements. For instance, Meta-Llama-3.1-8B attained the highest Exact F1 score (0.7099) with semantic input, only slightly outperforming its baseline (0.7076), and this difference was not statistically significant (𝑝≈0.64). Some models, such as Llama-3.1-8B-Instruct and Llama3.2-3B-Instruct, even showed small performance drops when semantic descriptions were included, with McNemar tests confirming that these differences were not significant ( 𝑝≈ 0 . 75 and 𝑝≈ 0 . 88). This suggests that in settings where the model is already exposed to sufficient task specific supervision, additional prompt-level context may offer limited benefit or even introduce redundancy. In contrast, TxGemma-9B-Chat exhibited the most notable improvement, with Exact and Relaxed F1 scores increasing from 0.6837 to 0.7092 and from 0.7483 to 0.7686, respectively; the McNemar test confirmed this difference as statistically significant ( 𝑝≈ 9 . 7 × 10 −5 ). By comparison, DeepSeek-Qwen-7B also showed a significant difference ( 𝑝≈ 6 × 10 −3 ), but in this case the semantically enhanced model performed worse (Exact F1: 0.7013 →0.6879). 5.4 Overall observations The largest performance improvements from semantically enhanced prompts appeared in zero-shot and few-shot settings, where gains in F1 scores were often statistically significant. In contrast, fine-tuned models showed smaller and mixed effects: 9
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Calcina et al. Table 2: Exact (left) and Relaxed (right) F1 scores for selected labels in few-shot and fine-tuned settings, with and without semantically enhanced prompts. Bolded values indicate the highest score in each column. We use symbols ◦ and • to denote whether the differences between using the baseline or semantically enhanced prompts are statistically significant ( • ) or not (◦) according to the McNemar test at a significance level of 𝑝= 0.01. Exact F1 Relaxed F1 Model Few-Shot Fine-Tuned Few-Shot Fine-Tuned /Semantic / Semantic / Semantic / Semantic Llama-3.1-8B-Instruct 0.2509 0.3005 •0.7053 0.7004 ◦0.3526 0.3948 0.7660 0.7645 Meta-Llama-3.1-8B 0.2698 0.2210 •0.7076 0.7099 ◦0.3537 0.2799 0.7670 0.7765 Llama-3.2-3B-Instruct 0.2300 0.2439 ◦0.6881 0.6867 ◦0.3769 0.3948 0.7629 0.7622 DeepSeek-Qwen-7B 0.1423 0.2270 •0.7013 0.6879 •0.2465 0.3891 0.7584 0.7521 txgemma-9b-chat 0.2732 0.3288 •0.6837 0.7092 •0.4469 0.4998 0.7483 0.7686 for most, differences were not significant, though TxGemma-9BChat benefited reliably while DeepSeek-Qwen-7B showed a significant decrease. These results indicate that semantic prompting is most effective in low-resource conditions, while its impact under full supervision is limited and model-dependent. 6 Discussion This section discusses the experiment findings and highlights the advantages and disadvantages of the different approaches. 6.1 Model pretraining and domain adaptation TxGemma-9B-Chat, based on the Gemma 2 architecture and further fine-tuned on therapeutic development data, outperformed general-purpose models in a few-shot scenario. This suggests that domain-specific pretraining can significantly improve performance when supervision is limited. However, in the full finetuning setting, its advantage diminished. In fact, general models like Meta-Llama-3.1-8B achieved comparable but slightly better results, indicating that once sufficient task-specific supervision is provided, prior domain specialization offers limited additional benefit. 6.2 Prompt quality matters The structure and clarity of prompts are critical to model performance. Poorly designed prompts often resulted in JSON formatting errors or reduced accuracy, particularly in zero-shot and few-shot settings. While adding semantic context improves task understanding by making objectives and entity definitions more explicit, excessive length or ambiguity can offset these gains. 6.3 Prompt length vs. model response Semantic enrichment inevitably increases prompt length, which can slow response time and raise computational overhead. It may also overwhelm smaller models when excessive detail is included. In practical applications, this must be weighed against the potential gains in entity extraction accuracy. 7 Conclusion This study investigated the impact of a semantically enhanced prompt design on LLM-based NER in the clinical domain. Our experiments on the MACCROBAT2020 dataset demonstrated that adding semantic label descriptions significantly improves model performance in zero-shot and few-shot scenarios, with notable gains in both Exact and Relaxed F1 scores. In contrast, fine-tuned models already exposed to task-specific data showed only marginal improvement. Future work could explore adaptive semantic prompting strategies, such as ontology-driven label enrichment, and further investigate the trade-offs between prompt length and inference efficiency. Additionally, this method could be tested on larger datasets and across different models to assess its generalizability. In summary, semantically enhanced prompts offer a straightforward yet effective way to boost clinical NER performance in low-data regimes, but their impact diminishes as models are exposed to more supervised training. Acknowledgements This work was supported by the Slovenian Research Agency. Funded by the European Union. UK participants in Horizon Europe Project PREPARE are supported by UKRI grant number 10086219 (Trilateral Research). Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Health and Digital Executive Agency (HADEA) or UKRI. Neither the European Union nor the granting authority nor UKRI can be held responsible for them. Grant Agreement 101080288 PREPARE HORIZON-HLTH2022-TOOL-12-01. References [1] Kaikai An, Shuzheng Si, Yuchi Wang, et al. 2024. Rethinking semantic parsing for large language models. arXiv preprint arXiv:2409.14469. [2] Dhananjay Ashok and Zachary C. Lipton. 2023. Promptner: prompting for named entity recognition. arXiv preprint arXiv:2305.15444. [3] J. Harry Caufield, Yichao Zhou, Yunsheng Bai, David A. Liem, Anders O. Garlid, Kai-Wei Chang, Yizhou Sun, Peipei Ping, and Wei Wang. 2019. A comprehensive typing system for information extraction from clinical narratives. medRxiv. Preprint. doi: 10.1101/19009118. [4] Quinn McNemar. 1947. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika, 12, 2, (June 1947), 153–157. doi: 10.1007/bf02295996. [5] Monica Munnangi, Sergey Feldman, Byron C. Wallace, et al. 2024. Onthe-fly definition augmentation of llms for biomedical ner. arXiv preprint arXiv:2404.00152. [6] 2025. Oxford english dictionary. https://www.oed.com/. Accessed: 2025-0617. (2025). [7] Yongliang Shen, Zeqi Tan, Shuhui Wu, et al. 2023. Promptner: prompt locating and typing for named entity recognition. In ACL (Long Papers). [8] Yongjian Tang, Rakebul Hasan, and Thomas Runkler. 2024. Fsponer: few - shot prompt optimization for named entity recognition. arXiv preprint arXiv:2407.08035. [9] Zeliang Tong, Zhuojun Ding, and Wei Wei. 2025. Evoprompt: evolving prompts for enhanced zero-shot named entity recognition. In COLING. [10] Yuwei Xia, Zhao Tong, Liang Wang, et al. 2023. Learning meta - prompt with entity-enhanced semantics for few-shot ner. SSRN. 10
LLM Based Approach to Extracting Smells in Slovenian Corpora Janez Brank Jožef Stefan Institute Ljubljana, Slovenia [email protected] Inna Novalija Jožef Stefan Institute Ljubljana, Slovenia [email protected] Dunja Mladenić Jožef Stefan Institute Ljubljana, Slovenia [email protected] Marko Grobelnik Jožef Stefan Institute Ljubljana, Slovenia [email protected] Abstract This paper presents a comparative study of automatic smell detection in Slovenian cultural heritage texts using both keyword-based search and large language model (LLM) inference. We process a portion of the dLib.si corpus from the late 19 th and early 20 th centuries, analyzing over 1.6 million text segments for olfactory references. The keyword method leverages an expert-curated list of smell terms, while the LLM method applies semantic inference via promptengineered queries. We compare the methods in terms of detection density, temporal trends, and agreement overlap. Additionally, we visualize the semantic landscape of extracted smell terms using t-SNE and unsupervised clustering with auto-generated labels. Our findings reveal limited overlap between methods, a shared rise in smell mentions over time, and distinct semantic clusters ranging from industrial to culinary and bodily smells. This study highlights the value of combining symbolic and neural approaches for nuanced sensory mining in digital heritage corpora. Keywords LLM, Artificial Intelligence, Cultural Heritage, Text Mining 1 Introduction Olfactory perception is an essential yet underexplored dimension in the analysis of historical texts, particularly within the cultural heritage domain. Smells, though intangible, play a critical role in shaping memory, atmosphere, and cultural meaning. However, their representation in written sources is often subtle, indirect, or metaphorical. This challenge becomes more pronounced in historical corpora such as 19 th - and early 20 th -century Slovenian publications, where evolving linguistic practices and cultural norms affect how sensory information is encoded. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society 2025, Ljubljana, Slovenia ©2024 Copyright held by the owner/author(s). https://doi.org/10.70314/is.2025.sikdd.5 This paper explores automatic smell detection in Slovenian cultural heritage texts using two complementary strategies: (1) a keyword-based approach derived from an expertcurated list of smell-related expressions and their morphological variants, and (2) large language model (LLM) - based semantic inference using prompt-engineered queries via the Together.ai platform. We process a subset of the dLib.si digital library corpus of Slovenian texts, divided into temporal buckets, and evaluate the performance, overlap, and divergence between the two methods. To facilitate large-scale analysis, we produce and analyze over 1.6 million document-query pairs, extracting smell mentions, classifying them by agreement type, and visualizing their distributions both temporally and semantically. Our goals are twofold: (i) to quantify the representational density of olfactory references in the corpus, and (ii) to better understand how computational methods can surface subtle cultural patterns that evade traditional keyword search alone. This work contributes toward a richer modeling of sensory information in digital heritage collections and highlights the value of combining symbolic and neural methods for text mining in the cultural heritage domain. 2 Related Work Recent years have seen increased interest in the computational modeling of olfactory expressions in historical and cultural texts. A prominent initiative in this space is the Odeuropa project [7], which focused on identifying, curating, and semantically linking smell-related content in European heritage corpora. Large-scale initiatives, such as the Odeuropa project, have produced the European Olfactory Knowledge Graph and tools like the Smell Explorer to trace historical olfactory knowledge across 400 years of European sources [7, 5]. Research on sensory perception in NLP has traditionally focused on the visual and auditory modalities, while olfaction remains relatively underexplored. Annotation frameworks such as the Olfactory Event Frame and guidelines for labeling sources, qualities, and experiences [6] provide structured resources for information extraction from historical and literary corpora. Traditional approaches to olfactory semantics rely on fixed lexicons such as the Dravnieks Atlas [1] and the DREAM challenge descriptors [3]. For morphologically complex and 11
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Brank et al. low-resource languages such as Slovene, monolingual models like SloBERTa [10] and seq-to-seq models like SloT5 [9] demonstrate that tailoring architectures to linguistic structure improves performance over multilingual baselines. A wide range of Slovene corpora underpins these modeling efforts. Gigafida 2.0, a reference corpus of 1.1 billion tokens covering contemporary written Slovene, provides a large-scale foundation for model pretraining and evaluation [4]. For user-generated content, the JANES corpus supplies richly annotated Slovene social media text, including normalization and NER [2]. Unlike prior studies that primarily focus on annotation frameworks, fixed olfactory lexicons, or large-scale multilingual heritage initiatives such as Odeuropa, our work provides the first comparative evaluation of keyword-based and LLM-based smell detection specifically for Slovenian cultural heritage corpora, highlighting the interplay between symbolic coverage and neural semantic inference. 3 Corpora and Preprocessing For the experiments presented in this paper, we used texts from the Slovenian Digital Library (dLib.si). Initially we downloaded, from the Library’s website, all documents from the period 1870–1919 for which OCRed text was available and whose language was marked as Slovene in the metadata there. In terms of content, this covers nearly all books, newspapers, magazines etc. published in Slovene during that period. From this corpus we then randomly selected 7 % of the documents from each year for further processing; thus the selected subset maintains the same distribution over time, genre, etc. as the full corpus. This resulted in a dataset of approx. 366 thousand documents with a total of 105 million words. 4 Methodology This section outlines the analytical pipeline used to detect, compare, and interpret smell-related expressions in Slovenian cultural heritage texts. Our approach combines large language model inference, keyword-based retrieval, temporal and density statistics, and unsupervised semantic clustering. 4.1 Comparative Evaluation of Detection Methods In order to identify olfactory expressions, we employed two complementary strategies: • LLM-based Extraction: Each document was split into passages and processed using a LLM. 1 The model returned a list of potential smell-related words or phrases, structured in JSON format. In cases of formatting failure, raw strings or exception messages were recorded. • Keyword-Based Search: A manually curated index of smell-related expressions, including morphologically inflected forms, was used for direct string matching within each passage.2 1The Llama-3.3-70B-Instruct-Turbo-Free model, accessed via Together.ai. 2 This index has been kindly provided by Mojca Ramšak and is based on her work on the anthropology of smell [8]. For each passage, we recorded both LLM and keyword results. We classified outcomes into four categories: LLM Only, Keyword Only,Both, or None. Additionally, we computed the Jaccard similarity 𝐽between the two result sets: 𝐽(𝐴, 𝐵)=|𝐴∩𝐵|/|𝐴∪𝐵|, where 𝐴 is the set of LLM-based results and 𝐵 is the set of keyword-based results. This metric enabled quantitative comparison of coverage and intersection across detection methods. 4.2 Temporal Distribution of Smell Mentions We extracted the year of publication from each document’s metadata. For each year, we aggregated: •Total LLM-detected smell terms •Total keyword-detected smell terms •Number of processed queries These aggregates were used to generate yearly time series, revealing longitudinal patterns in olfactory expression across the corpus. This temporal analysis supports hypotheses about cultural shifts, such as increasing industrial or bodily smell discourse over time. 4.3 Semantic Typology via Clustering of Smell Terms To explore latent smell categories, we constructed a semantic typology using the following steps: • Term Extraction: We extracted the 500 most frequent smell-related terms from the combined LLM and keyword results. • Vectorization: Terms were embedded using TF-IDF vectors over character-level n-grams ( char_wb with range 2–4), capturing morphological similarity. • Dimensionality Reduction: The high-dimensional vectors were projected to two dimensions using tSNE (perplexity = 30), yielding a visual semantic landscape. • Clustering: We applied k-means clustering (with 𝑘= 8) to the t-SNE coordinates. For each cluster, the top 5 TF-IDF terms were used to generate semantic labels (e.g., “Herbs & Cooking”,“Pharmaceutical Smells”). • Visualization: The clusters were visualized with color-coded labels and representative terms. Interactive versions were built using plotly. This typology enables data-driven classification of smell discourse and provides interpretable categories for cultural and linguistic analysis. 4.4 Document-Level Smell Density Analysis To assess the distribution of olfactory content across documents, we computed the smell density as the ratio of detected terms to queries per document: DensityLLM = # LLM terms # queries 12
LLM Based Approach to Extracting Smells in Slovenian Corpora Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Figure 1: Yearly trends in smell term mentions. Keyword-based detection consistently returns higher frequencies than the LLM, but both show similar growth patterns. Figure 2: Detection agreement between LLM and keyword methods. Most passages are matched by one method only, with a significant number showing no detection. The overlap (“Both”) occurs in fewer than one-third of cases. DensityKeyword = # Keyword terms # queries This metric enabled identification of smell-rich and smellsparse texts. Density distributions were visualized using boxplots and descriptive statistics, facilitating selection of representative or outlier texts for deeper qualitative analysis. 5 Evaluation and Results We evaluated complementary approaches to detecting olfactory references in historical corpora: a keyword-based method and an LLM-based classifier. The results highlight both convergences and divergences in performance across time, document density, and semantic coverage. Figure 1 shows yearly frequencies of smell-related mentions from 1870 to 1920. While keyword-based detection consistently yields higher absolute counts than the LLM, both methods exhibit similar growth trajectories. Agreement analysis between the two methods (Figure 2) reveals substantial divergence. Only about one-third of passages are identified by both approaches. A large portion is captured exclusively by the keyword method, while the LLM contributes a smaller but meaningful number of unique Figure 3: Smell term density per document. While outliers exist for both methods, keyword-based detection generally identifies a higher density of smell references per query. Figure 4: t-SNE semantic landscape of smell terms, clustered by character-level similarity and automatically labeled using top TF-IDF terms per group. The visualization reveals coherent groups such as food, ritual, body, and chemical references. detections. A significant subset of passages registers no olfactory detection at all, probably because most documents don’t mention smell-related topics in the first place. Figure 3 illustrates the distribution of smell term density per document. Keyword-based detection generally produces higher densities of references, whereas the LLM outputs are sparser but potentially more semantically filtered. Both distributions exhibit long-tailed outliers, where certain documents contain disproportionately high concentrations of olfactory mentions. To further analyze lexical diversity, we applied t-SNE to embed and cluster smell-related terms (Figure 4). The resulting semantic landscape reveals coherent groupings that align with cultural domains, including food, ritual, body, and chemical references. These clusters highlight the variety of olfactory expressions and suggest that both methods capture complementary facets of the semantic space. The LLM appears particularly adept at recognizing context-dependent terms, while the keyword method anchors clusters in explicit lexical cues. Overall, the keyword-based approach provides broader coverage and higher frequencies, but at the cost of noise and overcounting. The LLM method, while more conservative, contributes precision and captures context-sensitive 13
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Brank et al. olfactory references that keywords may overlook. The combination of both thus provides a richer and more balanced representation of olfactory discourse in historical texts. 6 Discussion Our analysis reveals several key insights into olfactory representations in Slovenian cultural heritage texts and the methodological implications of combining LLM-based and keyword-based detection. First, both detection strategies show meaningful trends over time, with a noticeable increase in smell-related references around the turn of the 20 th century. This may reflect broader urbanization, industrialization, and shifts in public health discourse, which intensified the cultural significance of air quality, hygiene, and olfactory environments. Second, although keyword-based detection consistently returned more hits, the LLM-based method surfaced a distinct set of semantically inferred mentions. As the agreement analysis shows, only a minority of mentions ( ∼ 24 %) were matched by both methods. One possible explanation of this would be if neural inference captures more nuanced or contextually implied smell references, such as metaphorical use ("a whiff of suspicion") or implied odors in narrative scenes. Third, density analysis suggests that LLMs return more sparse but targeted mentions, while keyword detection produces broader but sometimes noisier coverage. This difference is critical for researchers deciding between high recall and high precision when exploring sensory data in historical texts. Finally, the t-SNE landscape of smell terms uncovered semantically coherent clusters — e.g., medicinal substances, industrial emissions, festive foods, and bodily decay - and allowed us to generate meaningful auto-labels using top TFIDF terms. Such visualizations provide a valuable tool for cultural historians to engage with thematic patterns across large-scale textual datasets. Overall, our findings underscore the value of hybrid approaches to cultural text analysis. By comparing symbolic and neural perspectives, we gain both coverage and subtlety, enabling a deeper reconstruction of sensory worlds encoded in the archives. 7 Conclusion and Future Work We conducted a dual-method analysis of olfactory references in Slovenian historical texts, revealing how keyword search and LLM-based inference each contribute unique perspectives to sensory data mining. Our results show that while the keyword method offers broad lexical coverage, the LLM can detect more subtle, implied, or metaphorical references often overlooked by surface-level matching. Furthermore, t-SNE clustering of smell terms revealed rich thematic structures — such as food, medicine, pollution, and ritual — highlighting the semantic complexity of olfactory language. Together, these results demonstrate the complementary strengths of symbolic and neural approaches for enriching digital humanities research, especially in domains like historical sensory studies where annotation is sparse and vocabulary is diffuse. Several promising directions remain open for further exploration. First, we plan to expand the dataset to cover all documents in the dLib.si corpus, enabling more robust longitudinal and regional analyses. Second, we aim to improve LLM prompts to better handle nested or narrative contexts, including smells embedded in metaphor, irony, or emotional framing. Another avenue involves extending the classification of smell mentions into functional categories (e.g., pleasant vs. unpleasant, natural vs. artificial, bodily vs. environmental) using additional LLM-based postprocessing. We also intend to explore multilingual smell detection, comparing Slovene with other Central European languages to study cultural convergence and divergence in olfactory discourse. Finally, we hope to integrate our smell detection pipeline into public digital heritage platforms, providing curators, historians, and linguists with new tools for sensory exploration of archival materials. Acknowledgements This work was supported by the Slovenian Research Agency under the project J7-50233. References [1] Andrew Dravnieks. 1992. Atlas of Odor Character Profiles. ASTM International, (Feb. 1992). isbn: 978-0-8031-0456-3. doi: 10.1520/DS61 -EB. [2] Darja Fišer, Nikola Ljubešić, and Tomaž Erjavec. 2020. The janes project: language resources and tools for slovene user generated content. Language Resources and Evaluation, 54, 1, pp. 223–246. Retrieved Aug. 27, 2025 from https://www.jstor.org/stable/48740864. [3] Andreas Keller et al. 2017. Predicting human olfactory perception from chemical features of odor molecules. Science, 355, (Feb. 2017), eaal2014. doi: 10.1126/science.aal2014. [4] Simon Krek, Špela Arhar Holdt, Tomaž Erjavec, Jaka Čibej, Andraz Repar, Polona Gantar, Nikola Ljubešić, Iztok Kosem, and Kaja Dobrovoljc. 2020. Gigafida 2.0: the reference corpus of written standard Slovene. eng. In Proceedings of the Twelfth Language Resources and Evaluation Conference. Nicoletta Calzolari et al., editors. European Language Resources Association, Marseille, France, (May 2020), 3340– 3345. isbn: 979-10-95546-34-4. https://aclanthology.org/2020.lrec-1.4 09/. [5] P. Lisena, T. Ehrhart, and R. Troncy. European olfactory knowledge graph. Zenodo. doi: 10.5281/zenodo.10709703. [6] Stefano Menini, Teresa Paccosi, Serra Sinem Tekiroğlu, and Sara Tonelli. 2023. Scent mining: extracting olfactory events, smell sources and qualities. In Proceedings of the 7th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature. Stefania Degaetano-Ortlieb, Anna Kazantseva, Nils Reiter, and Stan Szpakowicz, editors. Association for Computational Linguistics, Dubrovnik, Croatia, (May 2023), 135–140. doi: 10.18653/v1/2023.latechclfl-1.15. [7] ODEUROPA Project Consortium. 2021–2023. ODEUROPA: negotiating olfactory and sensory experiences in cultural heritage practice and research. https://odeuropa.eu/. EU Horizon 2020 research and innovation programme, grant agreement No. 101004469. Royal Netherlands Academy of Arts and Sciences (KNAW) Humanities Cluster et al., (2021–2023). [8] Mojca Ramšak. 2025. Antropologija vonja. AMEU-ISH, Ljubljana. [9] Matej Ulčar and Marko Robnik-Šikonja. 2023. Sequence-to-sequence pretraining for a less-resourced slovenian language. Frontiers in Artificial Intelligence, 6. [10] Matej Ulčar and Marko Robnik-Šikonja. 2021. Sloberta: slovene monolingual large pretrained masked language model. In SiKDD. 14
BetweenTheLines - Cross Source News Analysis Georgi Trajkov [email protected] Jožef Stefan Institute Ljubljana, Slovenia Marko Grobelnik [email protected] Jožef Stefan Institute Ljubljana, Slovenia Adrian Mladenic Grobelnik [email protected] Jožef Stefan Institute Ljubljana, Slovenia Abstract Different news outlets covering the same event often emphasize, omit, or frame facts differently, making cross-source comparison essential for understanding media bias and information diversity. Large language models (LLMs) can automate this analysis, but simple single-LLM prompt approaches tend to underperform when processing large amounts of data [1]. Platforms like Ground News [2] and Event Registry [3] provide publisher and articlelevel bias scores but cannot track how individual claims and entities are portrayed by articles. The fundamental challenge is determining whether LLM prompt architecture affects accuracy when classifying claim presence across multiple news sources. We show that a multi-prompt LLM architecture reduces classification errors 7-fold (from 33.0% to 4.67%) compared to single-prompt approaches. Our pipeline first extracts all claims and entities from articles collectively, then evaluates each article separately for claim presence (confirmed/contradicted/partial/absent) and entity sentiment. This decomposition virtually eliminates false positives, major errors dropped from 28.0% to 0.79% across 797 manually validated claim-publisher pairs from Slovene news. The results demonstrate that task decomposition, not LLM sophistication, drives accuracy in cross-source analysis. This finding enables scalable media monitoring at $0.01 per event, making systematic bias detection accessible to journalists and researchers worldwide. 1 Introduction Different news sources (publishers) covering the same event (groups of articles reporting on the same story) often cover facts differently. While existing platforms like Event Registry [3] and Ground News [2] provide valuable bias indicators and sentiment scores, they do not track how specific entities (People, Organizations, Countries) and claims (Factual Claims) within articles are portrayed across publishers. Getting insight into these differences is usually time-consuming for the user. Thus we present BetweenTheLines, (Figure 1) a system that automatically identifies claims and entities in an event, and tracks their portrayal in each individual publisher. For example, when analyzing political coverage, we can see how the same entity is portrayed differently by 2 publishers, and how one publisher omitted a claim while the other did not. Our key technical contribution is demonstrating that multiprompt LLM architecture outperforms single-stage approaches for this task. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). SiKDD 2025, Ljubljana, Slovenia ©2025 Copyright held by the owner/author(s). https://doi.org/10.70314/is.2025.sikdd.26 Figure 1: Analyzed event in BetweenTheLines mobile webapp, showing the claims tab 2 Related Work Cross-source news analysis is an under-discussed area of research which is important for understanding media bias, information diversity, and narrative framing across different outlets. This section reviews existing approaches to cross-source news analysis, event aggregation systems, and LLM-based content analysis pipelines. 2.1 Cross-Source News Analysis Platforms Ground News represents a prominent platform for cross-source news comparison, classifying publishers along the left-right political spectrum. The platform has gained widespread adoption in educational institutions, with libraries at Harford Community College [4] and West Virginia University [5] integrating it into their media literacy curricula. For each news event, Ground News allows users to compare coverage by publisher on aggregate. While these aggregated summaries can reveal different emphases across the political spectrum, the platform does not provide article-by-article comparisons or track how specific entities and claims are portrayed between articles. 2.2 Event-Centric News Aggregation Event Registry [6, 3] pioneered event-centric news aggregation by clustering articles from multiple publishers around identified news events. The platform provides article-level sentiment scores using VADER sentiment analysis [7] and allows filtering 15
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Caporusso et al. Figure 1: Top-10 Features for TF-IDF Models Figure 2: Top-10 Features for LIWC Models 7 Discussion Our results indicate that the models trained on predefined features (LIWC) generally outperform those trained on learned features (TF-IDF n-grams), with the SVM model achieving the highest classification performance (RQ1-2). This suggests that LIWC features, which encapsulate linguistic and psychological constructs, provide a structured and interpretable representation of textual patterns related to SS. In contrast, TF-IDF captures surface-level word frequency distributions, which may be more susceptible to noise and context variability, limiting its predictive power for capturing abstract constructs like SS. Furthermore, our results support the findings by Caporusso et al. [4] regarding LIWC features correlated with SS. Notably, models trained on TF-IDF features tend to exhibit higher aggregated feature importance scores compared to those trained on LIWC. This could be attributed to the fact that TF-IDF operates on a larger and more granular feature space, capturing subtle variations in word usage. As a result, many features contribute partially to model decisions, leading to a higher sum of importance values across all features. In contrast, LIWC features are more constrained and predefined, leading to more concentrated but lower cumulative importance scores. This suggests that while TF-IDF captures a broader spectrum of textual variations, LIWC provides a more targeted and structured linguistic representation. Many of the features identified as relevant for the classification of SS (e.g., we and social referents) intuitively align with the nature of SS (RQ3). 8 Limitations and Future Work This study serves as a pilot for the interpretable classification of different Self aspects in text, focusing on SS. Several areas for improvement remain. Clearer annotation guidelines are needed for consistency. The choice of restricting to linear models, LIWC features, and unigrams/bigrams was appropriate for this exploratory study prioritising interpretability; however, it inevitably limits performance and representational richness. In future work, we plan to complement this approach with more powerful models and richer feature sets (e.g., embeddings). Here we wanted to compare models trained on learned vs predefined features, but we plan to train models on both. While in this study we did not perform hyperparameter optimisation, we will do so in the future. We aim to train a neural network for multi-class classification, enabling simultaneous prediction of SS and other Self-aspects, allowing for a more comprehensive analysis of self-representation in text. In the future, we plan to employ different datasets and implement Demšar’s evaluation method [8]. Our long-term goal is to be able, given a text instance, to determine what Self aspects are present and how they are expressed, in an explainable manner. To do so, it is not only necessary to extend our work to other Self-aspects, but to move beyond a binary classification for each of them. Work on the ontology underpinning future studies is ongoing [13]. 9 Acknowledgments We acknowledge Špela Rot’s assistance and the financial support from the Slovenian Research Agency for research core funding for the programme Knowledge Technologies (No. P2-0103) and from the projects CroDeCo (J6-60109), Shapes of Shame in Slovene Literature (J6-60113), and Natural Language Processing for Corpus Analysis in the Medical Humanities (BI-VB/25-27-021). JC is a recipient of the Young Researcher Grant PR-13409. References [1] Ryan L Boyd, Ashwini Ashokkumar, Sarah Seraj, and James W Pennebaker. 2022. The development and psychometric properties of liwc-22. Austin, TX: University of Texas at Austin, 10. [2] Marilynn B Brewer. 2002. Individual self, relational self, and collective self: partners, opponents, or strangers. (2002). [3] Jaya Caporusso. 2022. Dissolution experiences and the experience of the self: an empirical phenomenological investigation (master’s thesis). university of vienna. Advisor: Assist. Prof. Dr. Maja Smrdu. [4] Jaya Caporusso, Boshko Koloski, Maša Rebernik, Senja Pollak, and Matthew Purver. 2024. A phenomenologically-inspired computational analysis of self-categories in text. In Proceedings of JADT 2024. Vol. 1, 169–178. [5] Jaya Caporusso, Matthew Purver, and Senja Pollak. 2025. A computational framework to identify self-aspects in text. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop). Jin Zhao, Mingyang Wang, and Zhu Liu, editors. Association for Computational Linguistics, Vienna, Austria, (July 2025), 725–739. isbn: 979-8-89176-254-1. doi: 10.18653/v1/2025.acl-srw.47. [6] Jaya Caporusso, Thi Hong Hanh Tran, and Senja Pollak. 2023. Ijs@ lt-edi: ensemble approaches to detect signs of depression from social media text. In Proceedings of the Third Workshop on Language Technology for Equality, Diversity and Inclusion, 172–178. [7] Christopher G Davey and Ben J Harrison. 2022. The self on its axis: a framework for understanding depression. Translational Psychiatry, 12, 1, 23. [8] Janez Demšar. 2006. Statistical comparisons of classifiers over multiple data sets. The Journal of Machine learning research, 7, 1–30. [9] Lewis R Goldberg. 2013. An alternative “description of personality”: the big-five factor structure. In Personality and Personality Disorders. Routledge, 34–47. [10] Boshko Koloski, Nada Lavrač, Bojan Cestnik, Senja Pollak, Blaž Škrlj, and Andrej Kastrin. 2024. Aham: adapt, help, ask, model harvesting llms for literature mining. In International Symposium on Intelligent Data Analysis. Springer, 254–265. [11] X Alice Li and Devi Parikh. 2019. Lemotif: an affective visual journal using deep neural networks. arXiv preprint arXiv:1903.07766. [12] Scott M Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions. Advances in neural information processing systems, 30. [13] Luka Oprešnik, Tia Križan, and Jaya Caporusso. 2025. Building an ontology of the self: sense of agency and bodily self. In Proceedings of Information Society 2025. Cognitive Science. doi: 10.70314/is.2025.cogni.8. [14] James W Pennebaker, Matthias R Mehl, and Kate G Niederhoffer. 2003. Psychological aspects of natural language use: our words, our selves. Annual review of psychology, 54, 1, 547–577. [15] Stephanie Rude, Eva-Maria Gortner, and James Pennebaker. 2004. Language use of depressed and depression-vulnerable college students. Cognition & Emotion, 18, 8, 1121–1133. [16] Gemma Team et al. 2024. Gemma 2: improving open language models at a practical size. arXiv preprint arXiv:2408.00118. [17] David HV Vogel, Mathis Jording, Peter H Weiss, and Kai Vogeley. 2024. Temporal binding and sense of agency in major depression. Frontiers in psychiatry, 15, 1288674. [18] Dan Zahavi. 2007. Self and other: the limits of narrative understanding. Royal Institute of Philosophy Supplements, 60, 179–202. 22
Identifying Social Self in Text Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia A Instructions for Labelling: Social Self In the column relative to Social Self, insert: •0: if the Social Self is not present. •1: if the Social Self is present. Following, we provide a definition of Social Self [4], instructions, and examples of a text instance where it is present and a text instance where it is not present, taken from the dataset to be labelled: Definition: The Self as it is shaped and/or perceived when in an interaction or relationship of sorts with other people or entities to whom we attribute qualities of an inner life. Instructions For Social Self to be present in a text instance it is not enough for the text instance to contain references to other people and/or entities, but it has to contain mentions of the author’s interactions with them, influence on them, or influence they have on the author. This can be even minimal, e.g., in the form of referring to a person as my sister, or by using the first-person plural pronoun instead of the singular one. Examples A.0.1 Text instance containing Social Self: "My family was the most salient part of my day, since most days the care of my 2 children occupies the majority of my time. They are 2 years old and 7 months and I love them, but they also require so much attention that my anxiety is higher than ever. I am often overwhelmed by the care they require, but at the same, I am so excited to see them hit developmental and social milestones." Explanation of text instance with Social Self present: In this text instance, the author report on other people they are in some sort of relationship with, and about some aspects of their relationship and how they make the author feel. A.0.2 Text instance not containing Social Self: "Yoga keeps me focused. I am able to take some time for me and breathe and work my body. This is important because it sets up my mood for the whole day." Explanation of text instance with Social Self not present: In this text instance, the author does not report on any person, animal, or other entities to whom we attribute qualities of inner life. General Notes While a certain Self-aspect might not be prominently present in a text instance in its entirety, if it is present in a part of the text instance to be labelled, then it has to be labelled as present in the text instance. A given text instance can have none of the Self-aspects present, one of them present and two of them non-present, two present and one non-present, or all three of them present—any combination is possible. B Evaluation Figure 3: Confusion Matrices: Models Trained on Learned Features (TF-IDF) Figure 4: Confusion Matrices: Models Trained on Predefined Features (LIWC) Figure 5: Pairwise Wilcoxon Signed-Rank Test Results (pvalues) 23
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Caporusso et al. C Feature Importance Figure 6: Correlation Between Feature Importance Across Models Trained on Learned Features (TF-IDF) Figure 7: Correlation Between Feature Importance Across Models Trained on Pre-Defined Features (LIWC) 24
WinWin Meets – Investigating the Future of Online Meetings Martin Žust [email protected] Jožef Stefan Institute Ljubljana, Slovenia Marko Grobelnik [email protected] Jožef Stefan Institute Ljubljana, Slovenia Alenka Guček [email protected] Jožef Stefan Institute Ljubljana, Slovenia Adrian Mladenic Grobelnik [email protected] Jožef Stefan Institute Ljubljana, Slovenia Abstract Video conferencing is now central to modern collaboration, yet its functionality remains largely limited to passive audio–visual communication. Despite growing investment in artificial intelligence (AI), it is unclear which features truly enhance meetings and how users will adopt them. Here we present WinWin Meets, a Jitsi-based prototype that integrates Whisper transcription and GPT-4o processing to deliver real-time summaries, visual mind maps, and goal-oriented advice. Testing with 16 participants showed strong interest in summaries and mind maps, moderate interest in in-meeting guidance, and a preference for add-on integration. Market research confirmed low organic demand for advanced AI features, with users prioritizing reliable improvements such as automated notes. These results highlight a gap between experimental enthusiasm and everyday adoption, pointing to opportunities for targeted, industry-specific integrations that combine reliability with intelligent support. Keywords video conferencing, AI agent, testing, market research, zoom, negotiation, transcription, summarization, advice, meeting notes, AI innovations 1 Introduction As artificial intelligence advances rapidly, its potential to transform everyday digital tools, particularly video conferencing, has become increasingly apparent. Platforms such as Zoom, Google Meet, and Microsoft Teams have become standard, yet their functionality remains focused on basic communication. A new need is arising for next-generation conferencing, including intelligent assistants, automatic summarization, content analysis, and contextual support. These next-generation systems go beyond passive audio and video transmission to actively support users with intelligent features and real-time analysis [1]. Previous research reveals both promise and challenges. Proactive AI meeting assistants can improve efficiency but need to balance autonomy with what users are willing to accept [1]. Meanwhile, studies of speech-based technology underscore the difficulty of extracting useful outcomes from nuanced group interactions [2]. These perspectives suggest that AI’s success Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society 2025, Ljubljana, Slovenia ©2025 Copyright held by the owner/author(s). https://doi.org/https://doi.org/10.70314/is.2025.sikdd.14 in meetings depends on technical feasibility and sensitivity to human collaboration. With remote meetings now central to how we work, these systems directly impact productivity, collaboration, and organizational culture. This paper explores which functionalities could define the future of video conferencing and how AI may contribute. We combine market trend and user preference analysis, reviews of online discussions, and experimental testing of the WinWin Meets prototype. We explore which features matter to users, examine how AI can support meetings, and assess the potential to improve efficiency, clarity, and structure in digital communication. 2 Background and Related Work 2.1 Overview of Current Video Conferencing Solutions The video conferencing market is currently dominated by a few major players. Zoom, Microsoft Teams, and Google Meet together account for approximately 94% of global market share, with Zoom alone holding around 56% [3]. While all three platforms are actively investing in artificial intelligence features, their innovation must be carefully balanced with the risk of reputational damage. As established brands, they face more constraints than lesserknown startups, which can afford a higher level of experimental agility. This creates a unique window of opportunity for the emergence of disruptive technologies that have the potential to redefine the video conferencing experience. Most AI-enabled tools developed recently are not standalone platforms, but integrations designed to work alongside existing services like Zoom, Google Meet, or Microsoft Teams. Notable examples include tl;dv [4], Otter.ai [5], Fathom [6], Fireflies [7], and Sembly AI [8]. These applications primarily offer meeting transcription, and some provide more advanced analytics such as sentiment analysis or participant-level speaking time metrics. 2.2 Limitations of Existing Solutions and Emerging Needs Despite the growing number of AI integrations, fully independent platforms that natively combine video conferencing with built-in AI features remain rare. These features may include real-time transcription, intelligent meeting summarization, and contextual AI-generated recommendations. This segment remains underdeveloped, presenting a significant opportunity for innovation. While major platforms like Zoom have started introducing their own AI assistants (e.g., Zoom AI Companion [9]), they must innovate cautiously to protect their reputation and user base. This creates space for new companies to develop more ambitious 25
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Žust et al. AI-first conferencing tools, unrestricted by established brand expectations or legacy user commitments. However, innovating in markets where most users are already committed to existing platforms has notable downsides. Only about 2.5% of people are actively seeking new alternatives, with the majority being reluctant to change [10]. 3 Development of WinWin Meets 3.1 Overview As part of our research, we developed WinWin Meets, an AIbased alternative to Zoom. The application maintains familiar functionality, allowing users to start or join meetings just as they would expect. The key difference comes before entering the meeting room, where users can define their meeting goals. Once inside, they find a familiar interface with standard video conferencing features. These core functionalities are provided through an integration with Jitsi [11], an open-source video conferencing platform. It supports screen sharing, microphone and camera toggling, chatbased communication, polls, and many other standard features. Beyond the familiar main meeting window found in applications like Zoom, WinWin Meets adds a dedicated panel on the right side of the screen for the WinWin Agent. This panel features two main buttons: Summarize and Give Advice. The Summarize button generates meeting summaries up to the current moment, particularly useful for late arrivals. Hovering reveals three options: Short Text, Long Text, and Mind Map. While the text options provide traditional summaries of varying length, the Mind Map offers a quicker and more accessible visual overview. The idea behind the mind map is based on the observation that modern workplace attention is highly fragmented, with a median focus duration of just 40 seconds on any screen [12]. The Give Advice button offers guidance on how to achieve the goals specified before the meeting. These goals can also be adjusted during the meeting by clicking the Manage Goals button in the top right corner. Hovering over the Give Advice button reveals three options: Short Text, Medium Text, and Long Text, which provide advice in different levels of detail. Once the meeting concludes, a meeting report is quickly generated. The report includes all key points, action items, a meeting timeline, and the list of participants. Users can also generate a mind map from the final meeting content. 3.2 System Architecture and Implementation The frontend of the application was developed in Cursor [13], with assistance from Claude 3.7 Sonnet [14] and GPT-4o [15]. It is built using the React 19 framework [16]. We aim for a clean and minimalistic design that intuitively guides the user through each step of the interface. In the meeting room interface, we integrated Jitsi via its iframe API. Jitsi integration is straightforward, and the platform allows the use of its hosted servers for up to 25 active monthly users free of charge, which was sufficient for our prototype testing. The backend is built in Python, using the FastAPI framework [17]. For transcription, we integrated Whisper [18], and for natural language processing tasks (such as summarization and advice generation), we used GPT-4o. The backend exposes several endpoints, including: •Transcription •Advice generation •Meeting summarization •Health monitoring •Meeting notes •File uploads The WinWin Agent dynamically adapts to the language selected by the user. In this prototype, we supported English, German, and Slovene, allowing users to interact with the summarization and advice features in their preferred language. Figure 1: System architecture of the WinWin Meets application In this prototype version, we did not use any persistent database; all data is stored locally. Additionally, user authentication is not yet implemented, as the focus was on demonstrating core functionalities. 4 Testing and User Insights To evaluate the usefulness and usability of WinWin Meets, we conducted a structured user testing process involving 16 participants. Testing sessions were held in small groups of 2 to 4 participants, each lasting approximately 15 minutes. Participants simulated realistic discussions—including casual exchanges and roleplay scenarios such as negotiations or political debates—to test all implemented functionalities. The following sections present our testing results, with key findings shown in Figure 2. 4.1 Test Coverage Participants explored all key features, including the three variants of the Summarize function (Short Text, Long Text, and Mind Map), the three formats of the Give Advice function (Short, Medium, Long), and the Meeting Notes feature. After each session, they completed an anonymous survey with both multiple-choice and open-ended questions to assess usefulness and provide feedback. 4.2 Key Findings General Usefulness Most participants recognized the potential of AI-enhanced meetings. In fact, 87.5% responded Yes when asked whether AI could help them achieve meeting goals, while the remaining 12.5% answered Maybe. Summarize Feature The Summarize function was considered useful by 81.3% of participants. Preferences were split almost evenly: nearly half favored the Short Text, another 43.8% opted for the Mind Map, while only 12.5% selected the Long Text variant. Give Advice Feature When choosing advice length, participants showed a clear preference for medium-length suggestions: •50% selected Medium •25% chose Short •25% chose Long Meeting Notes Feature Participants emphasized three expectations for meeting notes: 26
SiKDD October, 2025, Ljubljana, Slovenia Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Short 4 (25%) Medium 8 (50%) Long 4 (25%) Which version of Give Advice feature do you like the most? Short text 7 (44%) Long text 2 (12%) Mind map 7 (44%) Which version of Summarize feature do you like the most? Leaderboard of speaking times 5 (31%) Insightful questions generator 6 (38%) Meeting coordination using agenda 5 (31%) Which potential feature do you find the most promising? Figure 2: User survey results (n=16) comparing preferences for existing features (Give Advice and Summarize) and ranking of proposed new features for application WinWin Meets •High reliability (timestamps, content accuracy) •Fast post-meeting availability •Stable performance across sessions 4.3 Ideas for Additional Features Among the proposed additions, the insightful question generator attracted the most interest (37.5%), while the speaking time leaderboard and agenda-based coordination were equally valued (31.3% each). Participants also suggested several custom features, including personal notes, live transcription export, cloud synchronization, calendar integration, live translation with tone analysis, and domain-specific modes for law, sales, or education. 4.4 Integration Preferences A clear majority (68.8%) preferred to use WinWin Meets as an addon to existing platforms, while only 31.2% supported a standalone application. 4.5 Use Cases by Industry Participants identified several promising domains for WinWin Meets, such as negotiation and sales, legal and consulting services, corporate meetings, academic events, client feedback sessions, NGO coordination, and specialized contexts like logistics, mergers and acquisitions, or trade deals. 5 Market Research and Trend Analysis Beyond developing and testing WinWin Meets, we conducted market research to understand user needs and expectations in the video conferencing space. Our approach combined online surveys, social media engagement, search trend analysis, and reviews of blog posts and user forums. This investigation aimed to reach a wider audience than application testing alone could provide. The resulting quantitative and qualitative insights complement rather than replace our user testing results. 5.1 Survey and Social Media Feedback Informal polls and surveys were conducted on platforms such as Facebook and Reddit. In a Facebook group focused on digital tools (GrowthHacking Slovenia), a poll asking users which feature they would most like to add to Zoom revealed that over 60% of respondents preferred having meeting notes generated at the end of a call as we can see in Figure 3. In contrast, only two respondents selected a real-time AI assistant. This suggests a clear user preference for simple and familiar enhancements over more complex and unfamiliar innovations. Similar sentiment was observed on Reddit (r/Zoom and r/remotework), where posted polls received limited engagement. Among the few responses, a general disinterest in AI-based meeting assistance was evident, with some users explicitly selecting “None of those”. 5.2 Search Behavior and Online Interest Trends Public search trends were analyzed using tools such as Answer the Public [19], Answer Socrates [20], AlsoAsked [21], and Ubersuggest [22]. These platforms provided insight into the types of questions users search for on Google, YouTube, and Reddit. The analysis showed minimal interest in AI-enhanced conferencing features. Instead, users were more focused on improving the efficiency and effectiveness of their meetings. Popular search queries we found included: •What are the 3 C’s of effective meetings? •What is the 10-10-10 rule for meetings? •How can I take better meeting notes? •What are the 5 P’s of meeting productivity? •How to extend the 40-minute limit on Zoom? •Is Google Meet better than Zoom? •Is Zoom free to install and use? These patterns confirm that users are primarily concerned with meeting outcomes and platform reliability, rather than with novel AI-driven functionalities. Figure 3: Distribution of 80 votes for preferred video conferencing features from our informal polling. 27
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Žust et al. 5.3 Forum Discussions and Deep-Search Insights Using tools like Grok [23] and Floth [24], we conducted a deeper exploration of online discussions and feedback. The most frequently mentioned user pain points include: •Low video quality and unstable connections • Privacy concerns (e.g., Zoom bombing, data storage policies) •Psychological fatigue from constant camera presence •Lack of end-to-end encryption and transparency • Poor UX from interface changes (e.g., Google Meet “floating bubbles”, Webex chat restrictions) •Discomfort with platform claims over recorded content User feedback highlights a desire for reliable, simple, and secure platforms with minimal friction in setup and usage. 5.4 Conclusions from Market Research Our market analysis reveals several key trends: (1) Users strongly prefer practical features like note-taking and agenda management over complex AI-based tools. (2) Popular search queries suggest a need for structured meeting frameworks and productivity strategies. (3) Persistent dissatisfaction exists around technical reliability, interface design, and data privacy. (4) Open-source alternatives offer control and security but are hindered by usability and cost barriers. Overall, the market exhibits demand for video conferencing improvements that enhance meeting effectiveness and reduce user burden, rather than introducing new technical complexity. 6 Discussion There are two primary approaches to understanding user preferences: direct inquiry and behavioral observation. Direct questioning suffers from significant limitations, including social desirability bias where respondents provide socially acceptable rather than genuine answers, and the fact that approximately 95% of human decisions occur subconsciously as discussed in [25]. Observational methods capture the unconscious preferences that drive actual user behavior, providing more accurate insights into real-world usage patterns. These methodological considerations explain our contradictory findings. While 87.5% of WinWin Meets participants believed AI could help achieve meeting goals, market research revealed minimal organic interest in AI-enhanced conferencing. This divergence reflects the difference between conscious evaluation in controlled environments versus unconscious behavioral preferences that emerge during natural usage. Additionally, our testing participants were primarily young AI researchers, likely more receptive to AI features than typical users. Our research uncovered widespread "Zoom fatigue", indicating that users have reached cognitive saturation with current video conferencing complexity. The strong preference for meeting notes over real-time AI assistance (60% versus minimal interest) demonstrates users’ desire for post-meeting value without additional in-meeting cognitive burden. This psychological context explains why solutions that prioritize seamless integration over feature prominence tend to gain market traction [26]. Our findings suggest distinct pathways for AI-enhanced video conferencing innovation. Industry-specific applications such as negotiations, sales, and legal consultations represent focused market segments where specialized AI features deliver measurable value propositions. The 68.8% preference for add-on integration over standalone applications indicates a market opportunity in enhancing existing platforms rather than replacing them, as demonstrated by successful tools like Fathom and Otter.ai. Although there is room for breakthrough products, any new solution must be at once reliable, easy to use, and meaningfully smarter than current tools—a difficult balance as existing platforms already invest heavily in their core features. The emphasis on reliability and customizable AI assistance reveals that AI features must meet higher performance standards than traditional features. Users consistently prioritize dependable functionality over advanced capabilities, suggesting that product development should focus on perfecting core AI functions before expanding feature sets. Future research should examine longitudinal adoption patterns and explore how user acceptance evolves as AI capabilities mature and become more familiar in workplace contexts. 7 Acknowledgements The research described in this paper was supported by the TWON project, funded by the European Union under Horizon Europe, grant agreement No 101095095. References [1] Rutger Rienks, Anton Nijholt, and Paulo Barthelmess. 2009. Pro-active meeting assistants: attention please! Ai & Society, 23, 2, 213–231. [2] Moira McGregor and John C Tang. 2017. More to meetings: challenges in using speech-based technology to support meetings. In Proceedings of the 2017 ACM conference on computer supported cooperative work and social computing, 2208–2220. [3] T3 Technology Hub. 2024. Market share of videoconferencing software worldwide in 2024, by program. Statista. Graph. (Apr. 2024). Retrieved Jan. 13, 2025 from https://www.statista.com/statistics/1331323/videoconferencingmarket-share/. [4] tldx Solutions GmbH. 2025. Tl;dv. https://tldv.io/. Accessed: August. (2025). [5] Otter.ai, Inc. 2025. Otter.ai. https://otter.ai/. Accessed: August. (2025). [6] 2025. Fathom. https://fathom.video/. Accessed: August. (2025). [7] 2025. Fireflies. https://fireflies.ai/. Accessed: August. (2025). [8] 2025. Sembly ai. https://www.sembly.ai/. Accessed: August. (2025). [9] Zoom Video Communications. 2025. Zoom ai companion. https://www.zoo m.com/en/ai-assistant/. Accessed: August. (2025). [10] Everett M Rogers, Arvind Singhal, and Margaret M Quinlan. 2014. Diffusion of innovations. In An integrated approach to communication theory and research. Routledge, 432–448. [11] 8x8, Inc. 2025. Jitsi. https://jitsi.org/. Accessed: August. (2025). [12] Gloria Mark, Shamsi T. Iqbal, Mary Czerwinski, Paul Johns, and Akane Sano. 2016. Neurotics can’t focus: an in situ study of online multitasking in the workplace. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems. ACM, 1739–1744. [13] Anysphere Inc. 2025. Cursor. https://cursor.sh/. Accessed: August. (2025). [14] Anthropic. 2025. Claude 3.7 sonnet. https://www.anthropic.com/news/clau de-3-7-sonnet. Accessed: August. (2025). [15] OpenAI. 2025. Gpt-4o. https://openai.com/index/hello-gpt-4o/. Accessed: August. (2025). [16] Meta Open Source. 2025. React. https://react.dev/. Version 19. Accessed: August. (2025). [17] Sebastián Ramírez. 2025. Fastapi. https://fastapi.tiangolo.com/. Accessed: August. (2025). [18] OpenAI. 2025. Whisper. https://openai.com/research/whisper. Accessed: August. (2025). [19] NP Digital. 2025. Answer the public. https://answerthepublic.com/. Accessed: August. (2025). [20] 2025. Answer socrates. https://answersocrates.com/. Accessed: August. (2025). [21] Candour. 2025. Alsoasked. https://alsoasked.com/. Accessed: August. (2025). [22] Neil Patel Digital. 2025. Ubersuggest. https://neilpatel.com/ubersuggest/. Accessed: August. (2025). [23] xAI. 2025. Grok. https://grok.x.ai/. Accessed: August. (2025). [24] 2025. Floth. https://floth.ai/. Accessed: August. (2025). [25] Gerald Zaltman. 2003. How Customers Think: Essential insights into the mind of the market. Harvard Business Press. [26] Fred D Davis. 1989. Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS quarterly, 319–340. 28
Predicting Ski Jumps Using State-Space Model Neca Camlek∗ Univerza v Ljubljani Ljubljana, Slovenia Živa Hegler∗ Univerza v Ljubljani Ljubljana, Slovenia Jakob Jelenčič Jožef Stefan Institute Ljubljana, Slovenia jakob[email protected] Marko Grobelnik Jožef Stefan Institute Ljubljana, Slovenia [email protected] Dunja Mladenić Jožef Stefan Institute Ljubljana, Slovenia [email protected] Abstract Ski jumping performance is shaped by both athlete technique and environmental conditions, with factors such as wind speed, wind direction, and ski orientation playing a critical role in determining jump trajectories. Accurate modeling of these trajectories is challenging due to dynamic and time-dependent nature of the system. In this work, we introduce a dataset of measured ski jumps and present a state-space modeling framework that captures the evolution of jumps under varying conditions. The model parameters are estimated using a ridge regression approach, enabling us to predict trajectories from initial states and wind sensor inputs. We evaluated the predictive performance of the model through leave-one-out cross-validation and analyzed its stability, showing that the approach can generate realistic trajectories with reasonable accuracy. To complement the modeling results, we developed an interactive web application that allows users to explore both recorded and simulated jumps, adjust environmental factors, and visualize their effects through animations. Together, the dataset, modeling framework, and the application offer a foundation for further research in ski jump analysis and provide an accessible tool for exploring the influence of external conditions on performance. Keywords datasets, state-space model, ski jumping, simulations, least squares 1 Introduction Ski jumping is a sport strongly influenced by both athletic technique and environmental conditions. Factors such as wind speed, wind direction, and different ski angles affect the trajectory and final distance of a jump, making accurate prediction a challenging problem. While statistical models and simulations have been applied in sports research for some time, many approaches simplify the problem and do not fully capture the dynamic evolution of the jump over time [11]. Recent advances in machine learning have introduced methods capable of modeling temporal systems with greater fidelity. In particular, state-space models provide a mathematical framework for representing hidden internal states that evolve over time in response to external input. This makes them well-suited for ∗Both authors contributed equally to this research. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society 2025, Ljubljana, Slovenia ©2025 Copyright held by the owner/author(s). https://doi.org/10.70314/is.2025.sikdd.30 modeling ski jumps, where environmental factors determine performance [9]. In this paper, we present a ski jump dataset together with a state-space model trained to predict jump trajectories based on changing environmental conditions. The model is estimated using a least squares approach and demonstrates how inputs such as wind and ramp adjustments influence the resulting jump. Beyond the modeling framework, we also developed an application that allows general users to interact with the data, run simulations, and visualize jump trajectories through animations. Beyond methodological interest, accurate prediction of ski jumps can improve athlete safety by anticipating risky conditions, support planning of hill design or enlargement, and contribute to fairer competitions through a better understanding of environmental effects. The remainder of the paper is as follows. Section 2 presents the handling of received data. Next, the proposed methodology is described in Section 3. The project results are presented in Section 4. We discuss the results in Section 5 and conclude the paper in Section 6. 2 Modeling Framework and Dataset This section describes the handling of data, focusing on statespace models and our data processing. 2.1 State-Space Model State-Space Models (SSMs) are a family of machine learning algorithms designed to capture and predict the behavior of dynamic systems by describing how their inner states change over time. Instead of only looking at past inputs and outputs, SSMs explicitly model the underlying dynamics, making them well-suited for sequential data. In state-space modeling, the objective is to identify the minimal set of system variables required to completely describe the system. These fundamental variables are referred to as the state variables. At any given time, the state of the system can be represented by a state vector, whose components correspond to the values of the respective state variables. SSMs are designed to predict both the manner in which inputs are reflected in the system’s outputs and the evolution of a system’s internal state over time and in response to specific inputs [2]. 2.2 Least squares method The least squares method is a regression technique that is used to determine the line that best fits a given set of data. It minimizes the sum of the squared differences between the observed data and the corresponding values implied by the regression function. Each data point reflects the relationship between a known independent variable and an unknown dependent variable [7]. 29
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Ž. Hegler, N. Camlek et al. To enhance the model, we incorporated ridge regression (L2 regularization), which helps to reduce overfitting during model training [12]. 2.3 Data Processing For our project, we used 223 CSV files, each containing the data of a jump, measured on the flying hill of Gorišek brothers in Planica, Slovenia. Each contains 17 columns ( ’Position’ , ’Height above ground’,’Time’,’X’,’Y’,’Z’,’Opening Angle’, ’Stalling Angle Left’ , ’Stalling Angle Right’ , ’Roll Angle Left’ , ’Roll Angle Right’ , ’Yaw Angle Left’ , ’Yaw Angle Right’,’Speed hor.’,’Speed vert.’,’Speed resulting’ , ’WindTime|WindName|WindSpeed|Wind...’ ) and the number of rows corresponding to the length of the jump. Data are recorded for every meter of air distance from the take-off point. The data required some pre-processing before it could be used for training the model. The column WindTime|WindName|WindSpeed|... combined multiple attributes separated by ’ | ’. Data from 12 sensors, each measuring six wind characteristics, were expanded into 12 × 6 = 72 columns, one per sensor–feature pair (sensor_feature). Position - air distance from the take-off point in meters. Begins with a negative value, which represents the distance from the starting point to the take-off point. In ski jumping, the starting point is adjusted according to the wind conditions, so this value is not constant. Height above ground - height above ground in meters. Time - time of the jump in seconds from the start of the jump. X, Y, Z - coordinates of the jumper in a 3D space in meters. The X axis is aligned with the hill direction, the Y axis is across the hill, and the Z axis is vertical. The take-off point is (0,0,0)as shown in Figure 1 Opening Angle - angle between the skis in degrees. Stalling Angle Left, Stalling Angle Right - angle between the chord line of the left/right ski and the horizontal plane in degrees. Roll Angle Left, Roll Angle Right - angle of the left/right ski around its longitudinal axis relative to the horizontal plane in degrees. Yaw Angle Left, Yaw Angle Right - angle between the left/right ski and the horizontal plane in degrees. (angles are shown in Figure 2) Speed hor., Speed vert., Speed resulting - horizontal, vertical, and the resulting speed of the athlete in km/h [13]. Figure 1: 3D model of Ski jump in Planica with added coordinates [10, 1] Figure 2: Different angles affecting the jump The wind features are as follows: WindTime - time of the wind measurement in the same format as the Time column itself. Since wind measurements are recorded less often, the wind values are applied to the most recent jump measurement and then just repeated until a new wind measurement is available. Since the wind is represented by a nonlinear function, it would be hard to capture its movements with interpolation, so we decided to drop this column. WindName - name of the sensor (Wi for 𝑖=1, . . . , 12) WindSpeed - resulting speed of the wind in km/h WindSpeedTangent - speed of the wind tangent measured along the x axis (hill direction) in km/h WindTurbulence - vertical speed of the wind turbulence in km/h WindSpeedCleanTan - wind speed tangent with turbulence removed in km/h WindSpeedCross - speed of the wind measurement along the y axis across the hill in km/h There are 12 wind sensors spread across the ski jump hill. To help with the analysis, we separated the jump section of the hill into 3zones. The first zone contains wind sensors 1to 4, the second zone contains sensors 5to 8, and the third zone contains sensors 9to 12 [11]. During processing, we also removed some ski jumps that were incomplete or had corrupted data, so the final dataset contained around 200 ski jumps. 3 Methodology This section describes our research methodology. We first present different variations of the SSM that we tested for the ski jump simulation, followed by describing our model and how it predicts the jumps. Finally, we present the description of our ski jump animation app. 3.1 Different modeling approaches In addition to pure SSM, we considered different approaches for modeling ski jumps that included classical physics-based models, but the data are not sufficient to accurately capture all the forces acting on the jumper. We also tried a hybrid approach that combined SSM and Physics-informed Neural Networks (PINNs [14]), where the SSM would provide a baseline prediction and the PINN would learn to correct any discrepancies, taking into account physical properties of the system, such as the mass of the pilot, the properties of the wind, and gravitational force [4]. 30
Ski Jumping Simulation Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia These parameters are included in the equations of motion and added to the total loss function. So, the model prefers solutions that are consistent with the laws of physics. This turned out to be less effective than a pure SSM approach, but the reason exceeds the purpose of this paper. More about errors and models’ comparison is given in Section 4.1. 3.2 Ski jump prediction model In order to fit our data to the SSM, we stored the data in each file in three vectors. The main vector contains states or state variables of the system, which in our case are the X, Y, and Z coordinates, jumper velocities, and all angles (opening, stalling, roll, and yaw) [6]. The observation vector contains the measured outputs of the system, which in our case are the X, Y, and Z coordinates and height above ground. The controls contain the external inputs to the system, which in our case are the wind measurements from all the sensors that are averaged over each zone and feature (speed, tangent, cross and turbulence). We then used ridge regression to estimate the matrices A, B, C and D of the SSM, as shown in Figure 3, where we minimized the computed values from the current and previous values and the next time-stamped values. Thus, matrix A computes the next state from the current state, B computes the next state from the current control, C computes the next observation from the current state, and D computes the next observation from the current control. We then use recursion to predict the next state from the prediction of the previous state and the current control, to get the full simulated jump. This allows us to predict the jump trajectory based on the environmental conditions and the starting state of the jumper [9]. Figure 3: Schema of SSM matrices [3] 3.3 Ski jump animation app To make our results accessible beyond the research setting, we developed an interactive web application using Shiny for Python [8]. The application serves as a front-end to the trained statespace model and allows users to explore ski jump simulations under varying environmental conditions or just to observe different measured ski jumps. Firstly, through a set of input controls, users can adjust factors such as wind speed, wind directions, or different ski angles, and the application instantly updates the predicted jump trajectory. Secondly, users can simply explore random jumps from the provided dataset or upload their own CSV file of measured jumps, as long as it includes the columns described in Section 2.3. The application presents the results as an animated visualization of the ski jump, showing the full trajectory and the final distance. In this way, the application functions both as an analytical tool, helping to test how different conditions affect performance, and as an educational resource that makes the mechanics of ski jumping easier to understand for a wider audience. It is available online. 1 4 Main results In this section, we present the results of our simulations. Firstly, we present a statistical comparison of all the models, followed by a precise analysis of our predictions. 4.1 Models’ error In order to evaluate different models, we first had to define a metric to measure the prediction error. Since actual and simulated jumps are represented with x, y, and z coordinates but are measured at different time stamps and can contain a different number of measurements, we had to find a way to compare them. We first tried to project the shorter trajectory on to the other one and compute the distance between the original and the projection, but this method turned out to be computationally expensive. So we decided to compute the distance between the actual and the simulated jumps by interpolating both jumps. The new measurements contain the start and end point and all the ones, where x reaches a natural value. We then compute the error as the norm of the difference between the two jumps. And after one of the jumps ends, we just add the distance from the end of the shorter jump to the end of the longer jump to the error. In this way, we penalize the model for not being able to predict the correct length of the jump. Since we had a limited number of jumps, we used leave-oneout cross-validation to evaluate the models. For each jump, we trained the model on all other jumps and then simulated the left-out jump. We then calculated the average error between the actual jump and the simulated jump for both the training set and the test set, as shown in Figure 4. In the process of developing our ski jump prediction model, we evaluated several variations to determine the most effective approach. We compared the performance of a pure SSM with a hybrid model that combined SSM with PINN. The pure SSM demonstrated superior predictive accuracy, probably due to its ability to directly model the temporal dynamics of ski jumps without the added complexity of PINNs. We also experimented with different configurations of the SSM, including using all available wind sensor data versus an averaged value of the zone. When we used all sensors, the average error for each point (in the training data is 1 . 67 m and in the test data is 1 . 89 m), while when we averaged the sensors over the zones, the error (in the training data was 1 . 76 m and in the test data 1 . 82 m). This suggests that averaging the wind data helps with the simulation. 4.2 Analysis of our model Wind is a critical factor in ski jumping, so we attempted to capture its nonlinear effects by including columns for the squared wind features. However, we found that adding these squared terms did not significantly reduce the prediction error. Since the simulation still requires numerous inputs, we made it interactive, allowing users to adjust the wind conditions and observe their impact on the jump. In the ski jumping app, users 1https://camlekn.shinyapps.io/ski-jump/ 31
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Kocjančič et al. Figure 2: Distribution of bike rentals across all stations. The vertical blue line indicates the start of the year 2024. 2.2 Data Preprocessing The dataset structure prevented distinguishing missing values from true zeros (i.e., days when no rentals occurred), so all empty or null entries were treated as zeros. This resulted in sparsity for some stations, in which many entries had little information on rental activity. To prevent this impacting our analysis, we excluded those with more than 33% zero entries, retaining 25 stations out of the original 48. For the machine learning methods described later, we also implemented a set of lagged features: • total_rentals_mean_7_days: Average rental count over the 7 days preceding the current data point. • total_rentals_mean_14_days: Average rental count over the 14 days preceding the current data point. • total_rentals_mean_21_days: Average rental count over the 21 days preceding the current data point. • total_rentals_mean_28_days: Average rental count over the 28 days preceding the current data point. Figure 3: Rentals per day of the week 2.3 Exploratory Data Analysis The data exhibits pronounced weekly and monthly seasonalities, as well as non-stationarity, as illustrated in Figures 3 and 4. Annual patterns show rental activity declining in winter, rising in spring, peaking in summer, and gradually decreasing in autumn, with weekends consistently exhibiting lower rental counts. Anomalous behavior was observed in the winter of 2024, when rental counts were markedly higher than typical seasonal levels. The Pearson correlation coefficients (Figure 1) between features related to bicycle rentals indicate that the number of daily rentals ( total_rentals ) is strongly and positively associated with recent rental trends, as reflected by correlations of 0.73, 0.67, 0.64, and 0.63 with the 7-, 14-, 21-, and 28-day moving averages, respectively. A strong positive correlation is also observed with air temperature (0.59), whereas moderate negative correlations are found with relative humidity (-0.43) and precipitation (-0.31), suggesting that rentals are more frequent on warm, dry days. Weaker associations are present with the day of the week (-0.27) and holiday status (-0.10). As expected, the moving average features exhibit high intercorrelation (e.g., 0.94 between the 7and 14-day means) due to their overlapping calculation windows. 3 Experiments This study pursued two primary objectives. First, we examined the feasibility of forecasting bicycle rentals one day in advance and evaluated how forecastability varies across stations with different data sparsity. Second, we investigated long-horizon forecasting over a 90-day period, focusing exclusively on predicting the total number of rentals. In this task, standard machine learning models were trained on historical data and then used recursively to generate forecasts for the entire period. Due to this setup, the results for DS_W suffer from data leakage. Specifically, a single model is trained using past rental counts and future weather information, so, for example, predicting rentals in July involves access to the actual recorded weather conditions for that month, which artificially improves performance. 3.1 Training and Test Data Split Because the available weather data was limited to the years 2024 and 2025, while the rental dataset spanned from 2021 onward, 38
Bike Rental Forecasting Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Figure 4: Bike rental data with temperate seasons we constructed three distinct datasets. Here, each entry corresponds to a single day and includes rental data for all stations. The first dataset, DS_W, combined rental and weather data (498 entries). The second, DS_NO_W, included only rental data for the same period (498 entries). The third, DS_FULL, comprised the complete rental dataset without weather data (1,593 entries). The data splitting strategy differed in the two tasks. For the station-level one-day-ahead forecasting task, each dataset was divided into 25 subsets, corresponding to individual stations. Within each subset, random sampling was used to split the data into training and testing sets with an 80:20 ratio. The target variable in each subset is the specific station’s rental count. For the long-horizon task, no station-level subdivision was performed, as only total rental counts were modeled. The final 90 days were used as the test set—roughly corresponding to a temperate season—allowing us to assess whether the models capture seasonal patterns in a new period while maintaining realistic temporal separation between training and testing data. 3.2 Models and Algorithms Used For the long-horizon forecasting task, the AutoARIMA model served as the baseline, while for the one-day-ahead forecasting task, the baseline was the Mean Regressor, which predicts using the 7-day lag mean. We evaluated several machine learning models, including Random Forest (500 trees, max_features=0.9), Gradient Boosting (500 estimators), Linear Regression, and SVM ( 𝐶= 10, degree=2, 𝛾= 0 . 1, linear kernel). The hyperparameters for the Random Forest and SVM models were selected using a grid search optimization procedure; the rest of the models used default parameters. For the Random Forest model, only the max_features parameter was tuned. We additionally tested deep learning approaches: LSTM (input size = 96, RMSE loss, 10,000 epochs) and N-BEATSx (input size =96, RMSE loss, 500 epochs). Training was performed on a laptop equipped with an RTX 3050 GPU (4 GB VRAM), which constrained the range of hyperparameter configurations that could be explored, particularly for the neural network-based approaches. 3.3 Performance evaluation Model performance was assessed using Root Mean Squared Error (RMSE) and Mean Absolute Percentage Error (MAPE). Additionally, the Relative Root Mean Squared Error (RRMSE)[1] was used to enable inter-station performance comparisons in the one-dayahead forecasting task. RRMSE is defined as follows: RRMSE = RMSE 𝑦(1) where 𝑦is the mean of the target values. 3.4 Results The results for the one-day-ahead task are presented in Table 1, with station forecastability visualized in Figure 5. The longhorizon task outcomes are presented in Table 2. 4 Discussion and conclusion For the one-day-ahead forecasting task, a clear correlation exists between station data sparsity (Figure 2) and forecastability (Table 1). Stations with fewer rentals or gaps in data are easier to predict accurately. Interestingly, using the DS_FULL dataset—which includes data prior to 2024—can reduce modeling accuracy for certain stations. Including weather features in DS_W leads to little or no improvement compared to DS_NO_W. For the longhorizon task, including weather data proves beneficial, as both classical machine learning models and neural networks show improved performance (Table 2). However, as described in the Experiments section, the machine learning results on DS_W are overly optimistic due to data leakage: the models are trained on historical rental counts while also accessing future weather information during recursive forecasting (e.g., predicting rentals in July uses the actual recorded weather for that month). This is reflected in the comparison with DS_NO_W, where classical machine learning methods achieve a 33% mean reduction in MAPE, while neural network approaches show only a 17% mean decrease, suggesting that the apparent benefit of weather data is amplified for classical methods because of this setup. Our results echo [3] where Gradient Boosting models matched or outperformed neural networks on several datasets, demonstrating the effectiveness of simpler models. While neural networks 39
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Kocjančič et al. Figure 5: Model performance of one-day-head forecasting for different stations for DS_W Table 1: Average RRMSE of all models of one-day-ahead forecasting across datasets (RRMSE) and stations. Station DS_FULL DS_NO_W DS_W 6 0.9210 0.9097 0.9116 7 0.5849 0.5439 0.5488 8 0.7948 0.6821 0.6872 90.6532 0.6646 0.6631 10 0.9550 0.7747 0.7753 11 1.0110 1.0034 1.0027 12 0.6028 0.4649 0.4540 13 0.6601 0.4000 0.4022 14 0.6902 0.4840 0.4720 15 0.5218 0.4780 0.4652 16 0.7185 0.5984 0.5975 17 0.8336 0.7337 0.7402 18 0.5274 0.4670 0.4522 21 0.5476 0.5218 0.5215 22 0.5198 0.4171 0.4160 23 0.4783 0.4363 0.4349 24 0.4896 0.4760 0.4696 25 0.6834 0.5570 0.5608 26 0.6506 0.6897 0.6812 27 0.9463 0.9898 0.9595 28 0.5580 0.4898 0.4936 29 0.6008 0.5761 0.5788 30 0.5941 0.5496 0.5531 31 0.8952 0.6452 0.6474 32 0.5453 0.4873 0.4851 Average 0.6793 0.6016 0.5989 could potentially benefit from hyperparameter optimization, the same applies to other methods as well. A detailed comparison of different approaches was beyond the scope of this preliminary study but could be explored in future work. Table 2: Model performance of 90-day forecasting across datasets (RMSE / MAPE) Model DS_FULL DS_NO_W DS_W AutoARIMA 120.09 / 0.9525 118.50 / 0.9954 118.50 / 0.9954 Random Forest 108.29 / 0.7153 100.94 / 0.7431 76.36 / 0.7014 Gradient Boosting 95.17 / 0.7451 94.96 / 0.9584 74.69 / 0.5513 Linear Regression 90.29 / 0.9372 84.78 / 1.0816 71.71 / 0.8872 SVR 94.86 / 0.8893 87.12 / 0.9507 67.95 / 0.8036 LSTM 112.05 / 0.7133 125.13 / 0.8494 130.00 / 0.8070 NBEATSx 106.49 / 1.0329 128.90 / 0.9972 117.45 / 0.7246 Average 103.89 / 0.8551 105.76 / 0.9394 93.81 / 0.7815 Acknowledgements This work was supported in part by the Slovenian Research Agency through core funding for the programme Knowledge Technologies (No. P2-0103) and by the project KReATIVE, funded through NetZeroCities under the European Union’s Grant Agreement No. HORIZON-RIA-SGA-NZC 101121530. We also thank Tea Tušar for her suggestions regarding data visualization. References [1] Shikun Chen and Nguyen Manh Luc. 2022. Rrmse voting regressor: a weighting function based improvement to ensemble regression. arXiv preprint arXiv:2207.04837. [2] Jimmy Du, Rolland He, and Zhivko Zhechev. 2014. Forecasting bike rental demand. Gebhard, K., & Noland. [3] Shereen Elsayed, Daniela Thyssens, Ahmed Rashed, Lars Schmidt-Thieme, and Hadi Samer Jomaa. 2021. Do we really need deep learning models for time series forecasting? CoRR, abs/2101.02118. https://arxiv.org/abs/2101.02118 arXiv: 2101.02118. [4] Hadi Fanaee-Tork. 2012. Bike sharing dataset. Dataset. (2012). https://www.k aggle.com/datasets/marklvl/bike-sharing-dataset. [5] Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural computation, 9, 8, 1735–1780. [6] Meerah Karunanithi, Parin Chatasawapreeda, and Talha Ali Khan. 2024. A predictive analytics approach for forecasting bike rental demand. Decision Analytics Journal, 11, 100482. doi: https://doi.org/10.1016/j.dajour.2024.10048 2. 40
Predicting Traffic Intensity on Motorway Sections Matic Kladnik† Jozef Stefan International Postgraduate School Ljubljana, Slovenia matic.klad[email protected]m Dunja Mladenić Department of Artificial Intelligence Jozef Stefan Institute Ljubljana, Slovenia dunja.[email protected] Abstract This paper addresses predictions of traffic intensity on sections of motorways. Predictions are computed for timespans from 24 hours up to 52 weeks. With our adaptive system, we update predictions with newer ones, once additional features can be computed from available data. We use historic context of past traffic intensities on specific sections at specific periods of time, as well as semantic context about the target period. We have evaluated our methodology with multiple machine learning models and compared performances for various timespans on a specific motorway section. The evaluation results show that our methodology improves predictions for specific periods over time. Keywords Motorway, traffic intensity, prediction, regression, system, semantic context, evaluation, machine learning 1 INTRODUCTION A prediction system for predicting traffic intensity on motorway sections can support a wide range of decision making, strategic, and operative processes at the motorway management organization. It can also support end users, such as daily commuters, tourists, and other drivers with their planning of a trip. The focus of this paper is on architecture of the motorway traffic intensity prediction system as well as on the evaluation of the machine learning models that were trained to produce the predictions for various timespans. 2 PROBLEM SETTING AND DATA The objective of the proposed methodology is to make long term and medium-term predictions of traffic intensity or frequency (vehicle count) on various sections of motorway based on historic data of traffic counters, semantic context of motorway stations, and semantic context of time periods. Predictions serve the motorway management company for better planning of construction projects and to find the least intrusive time slots for road maintenance work. It also serves the motorway drivers when planning a trip. 2.1 Traffic Counters There are close to one hundred traffic counters that we consider for predictions. Each counter is supported by a pair of inductive loops that are laid into the asphalt of the road. Signals are processed, sent through an IoT communication device and stored into the database. In the data, there are counts or frequencies of total vehicles, and counts by vehicle types (passenger car, transport truck, bus) for each hour-long time period. E.g. number of vehicles from 8:00 to 9:00 for each of the lanes of a specific motorway section separately. 2.2 Semantic Context For each of the examples in the dataset we produce semantic context features. For each day and time of day period, we produce semantic context features to inform the model whether a certain time period is on a workday or a weekend, whether the specific time period falls into the morning rush hours or the afternoon rush hours. These semantic features give additional information to improve the performance of machine learning models. 2.3 Data Processing After downloading the data from the motorway counters via an API of the data provider, we additionally process it to increase consistency and reliability of predictions. During data processing, we merge data from all lanes of a specific motorway section, which is usually denoted with neighboring towns and the direction of the motorway section. 3 METHODOLOGY DESCRIPTION We propose a prediction system that includes incorporation of multiple machine learning models to deliver the most reliable predictions based on available data and the timespan for which the system is making predictions of traffic intensity. To improve prediction accuracy, we make medium-term and long-term predictions. In our case, long-term predictions are made from 1 week to 52 weeks in advance for a specific 1-hour Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia © 2025 Copyright held by the owner/author(s). http://doi.org/10.70314/is.2025.sikdd.25 41
time period for a specific day of week. Which means that we can make up to 52 predictions when conducting long-term predictions after receiving a new data example, e.g. traffic frequency for a specific 1-hour time period (e.g. 14:00-15:00) for a specific day in time (e.g. Monday). Whereas medium-term predictions are those that predict from less than 24 hours up to 1 week in advance. For medium-term predictions, we take more features for recent traffic frequency into account for improved accuracy. Long-term predictions are useful when making decisions for actions that are several weeks or months in the future, while medium-term predictions are more useful when making decisions for actions that will take place from 1 to 7 days in the future. We have a separate machine learning model for each of the included counters on the motorway to better adjust to specifics of the traffic dynamic of a specific counter when making predictions of traffic frequency. We have also trained several general-purpose models that are trained on a group of counters or all counters. These are present to support counters with short data history. Predictions are exposed through a REST API service and are available upon request. They are computed and updated regularly, e.g. daily or hourly. More approaches in [1][6]. 3.1 Machine Learning Models To compute predictions of traffic intensity in the future, we use regression machine learning models. We have trained and evaluated several models with the usage of different machine learning algorithms. These are: linear regression, SVM (SVR – Support Vector Machine for Regression), and XGBoost, which is an ensemble model of decision trees. Features for training models and making predictions are engineered in such a way that each one of the models can use the whole set of features. E.g. we use a one-hot encoding approach when a feature would otherwise have multiple categorical values. We focus on training a specific model for each of the motorway sections that were part of the research. Note that a more general model, trained on data from multiple motorway sections could be more appropriate for motorway sections that have been newly added and do not have enough historical data to support training of a reliable machine learning model with sufficient evaluation period. We use up to 7 features that are based on historic data, 7 time period features, and 6 semantic context features for a specific time period and location. Model training processes use MAPE (Mean Absolute Percentage Error, used interchangeably with MARE – Mean Absolute Relative Error). More on relevant machine learning models and metrics in references ([2][3][4][5]). 3.2 Prediction System Description We continue with the description of our proposed prediction system. The system consists of two main subsystems. One for periodically computing and storing traffic intensity predictions for various time spans. And another for delivering predicted traffic intensity via a REST API service. As we can see on Figure 1, the system fetches data from the data provider’s REST API service. Data is processed after retrieval and sent into a table of prediction system’s database. This data is read periodically by the adaptive prediction system. Once a new value is processed by the system, it checks if there are any additional models with a shorter timespan available, compared to the model used for the currently available prediction. The system prioritizes predictions from models with a shorter timespan in order to update the database with the most reliable predictions available at the time. E.g., prediction with a 1-month timespan succeeds and replaces the prediction with a 3month timespan. Different long-term and medium-term models can be trained using different machine learning algorithms, depending on the Figure 1: Diagram of the system for producing and distributing predictions of traffic intensity on motorway sections 42
algorithm that performed the best during the evaluation of the models. Once updated the predictions are stored in the database, they are available to users, such as strategists, operators and support specialists within the motorway management organization. Or end users of the motorway, such as drivers of cars, trucks, buses, etc. A key advantage of this approach is that drivers and motorway operators and specialists get insights that are based on the same predictions for traffic intensity, which supports greater transparency of information and stronger compatibility of different applications for end users and motorway professionals. E.g. the system can support long-term planning for larger maintenance or reconstruction projects for up to 1 year ahead, as well as long-term planning of road users. For instance, drivers can plan their holidays and the time of their commute ahead. And highway maintenance operators can find the most optimal schedule for short maintenance work. 4 EVALUATION We continue with the evaluation of the machine learning models. To compare models, trained with different algorithms, we use the evaluation results for the same motorway section on the Slovenian motorways. We use the period from 1 May 2024 until 5 May 2025 for evaluation. We use Scikit-learn library[7] to train the linear regression (using ordinary least squares approach) and SVM (SVR) models and the XGBoost library[8] to train the XGBoost models. SVM model is trained using the RBF kernel, and with scaled gamma hyperparameter. In majority of motorway sections, XGBoost models with a maximum depth of 6 performed the best which is why we used models with the same hyperparameter value for the following analyses. We use gbtree as the booster, while the learning rate is 0.3. Table 1: Model Performance Comparison timespan algorithm MAE RMSE MAPE 24 hours XGB 39.43 62.75 10.5% 24 hours SVM 42.38 65.86 11.5% 24 hours lin. reg. 43.14 66.93 11.6% 7 days XGB 45.66 70.69 11.6% 7 days SVM 43.70 68.91 12.1% 7 days lin. reg. 43.51 69.04 12.1% 4 weeks XGB 57.30 88.56 13.9% 4 weeks SVM 50.20 77.86 14.1% 4 weeks lin. reg. 51.33 78.63 14.7% 52 weeks XGB 88.33 121.93 20.9% 52 weeks SVM 53.54 84.49 14.9% 52 weeks lin. reg. 70.46 96.98 21.3% We evaluated the models on a little over 1 year of test data, which was not included in the training or validation part of the process. We continue with the analysis of the model performances as seen in Table 1. If the timespan attribute’s value is ‘7 days’, it means that the model predicts 7 days into the future. We use several metrics to describe the performance of the models. These are: MAE (Mean Absolute Value), RMSE (Root Mean Square Error), and MAPE (Mean Absolute Percentage Error). MAPE is a crucial metric as it shows relative errors in percentages which is key when evaluating the models as traffic frequency varies significantly throughout different parts of the day. We can see some interesting performance dynamics of the models. The XGBoost model performs the best for 24-hour timespan, with a significant performance uplift of at least 1 percentage point in MAPE, compared to the other two models. It is also better in the other two metrics: MAE and RMSE. We continue with the performance analysis of the long-term predictions. For the 7-day timespan, the XGBoost model is still noticeably better than the other two models with a 0.5 percentage point uplift in performance. For the 4-week timespan, XGBoost still holds a small lead in the key metric (MAPE), whereas the SVM model has significantly better results when considering just MAE and RMSE metrics. For the 52-week timespan, we can see an interesting dynamic as the SVM model takes a significant lead in performance as it is the only one with the MAPE value of less than 15%, whereas the MAPE values of the other two models surpass 20%. The dynamic is likely caused by a reduced set of features as there are significantly less historic traffic count features that are included when making predictions with a 52week timespan. It seems this has a significantly negative impact on training the XGBoost model, which is a tree ensemble model, while having additional features available gave the XGBoost model an edge for predictions with a timespan up to 4 weeks, especially up to 7 days. Figure 2: Distribution of absolute relative errors by 5% buckets for XGBoost 7-day timespan model On Figure 2 we can see how absolute relative errors are distributed if they are split into 5% absolute relative error buckets. We can see that in 45.5% of the cases, the absolute relative (or percentage) error of the predicted traffic frequency is less than 5% of the actually measured traffic frequency. 21.7% of predictions have a relative error between at least 5 and 43
(excluding) 10 percent, and 11.2% of predictions have a relative error between 10 and 15 percent. This means that in 78.4% of predictions, the relative error was less than 15%, which can be considered as a sufficiently good performance for the models to support a sufficiently reliable traffic intensity prediction system. Figure 3: Mean relative errors by each hour of the day for XGBoost 7-day timespan model We continue by analyzing the distribution of mean relative errors by each hour of the day as seen on Figure 3. We can see that the model generally tends to slightly overestimate or overshoot with its predictions. Especially during the night-time periods, when there are fewer vehicles on the motorway. In the mean aggregate, there is less than a 2% mean relative error during the morning rush hours (at 6:00-7:00, 7:00-8:00, and 8:00-9:00). It is the highest during the 15:00-16:00 period, with more than 13% of mean relative error. However, the error is substantially smaller during other afternoon rush-hour periods, 14:00-15:00, 16:00-17:00, and 17:00-18:00, where it remains under 4%. Apart from the 15:00-16:00 period, the mean relative errors are consistently under 6%. When the model does undershoot or underestimate with its prediction, the mean relative error is less than 2%, close to 1%. We can see a spike of mean relative error at the 15:00-16:00 period. Upon investigation, it turns out only around 20 vehicles were counted in the data for a specific period, which is unusual for this period and likely a consequence of a traffic accident or some issue with data collection. We have also conducted an aggregated evaluation of models on 10 various motorway sections, where mean MAPE values were 14%, 15%, 18%, and 20% for 24-hour, 7-day, 4-week and 52-week timespans respectively. Predictions for sections near the capital city were generally less reliable than others. 4.1 Evaluation Insights When considering the results of the evaluation of trained machine learning models for specific motorway sections, we have gathered several key insights. In some examples, we could not compute all features due to missing values in data, meaning that certain features had NaN values after computing historic time-series features with Pandas’ shift function. In this case there is a strong advantage of having a decision tree ensemble model (e.g. XGBoost) as a backup, even if it is not the best performing model for a certain timespan. This is due to the ability of the tree ensemble models to apply only those trees that are covered by features with available values. In this case the predictions are generally less accurate but possible. Another key insight is that the evaluation supports our proposed methodology with multiple models to improve the performance of the predictions for each included timespan. Another useful insight is that different algorithms can produce the best models for different timespans on the same motorway section. As was the case with the SVM model in our evaluation. 5 CONCLUSION We have overviewed the methodology that we use as the foundation for our proposed system for predicting traffic intensities on motorway sections. Including the adaptive prediction system and the supporting machine learning models that support making predictions for various timespans to, in time, improve already available predictions for specific time periods in the future. We have also overviewed the evaluation of the trained machine learning models and found some useful insights that support our proposed prediction system. Compared to related work, the key contributions in our methodology are significantly longer prediction timespans, inclusion of semantic context, and higher adaptability to data. Based on the presented current evaluation results, our methodology produces predictions with sufficient reliability to support long-term decision making of various roles. For further improvements to the system, we could train and evaluate some deep learning models and models that are based on the transformer architecture, as well as some other time-series forecasting procedures, such as Facebook Prophet. We could also engineer additional semantic context features for further improvements to the performance of the existing models. For additional improvements for shorter timespans, we could also include weather forecast data. References [1] Bernardo Gomes, Jose Coelho, Helena Aidos. 2023. A survey on traffic flow prediction and classification. In Intelligent Systems with Applications, vol. 20. DOI: https://doi.org/10.1016/j.iswa.2023.200268 [2] Jithin Raj, Hareesh Bahuleyan, Lelitha Devi Vanajakshi. 2016. Application of Data Mining Techniques for Traffic Density Estimation and Prediction. Transportation Research Procedia, vol 17. DOI: https://doi.org/10.1016/j.trpro.2016.11.102 [3] Yuyu Zhu, QingE Wu, Na Xiao. 2022. Research on highway traffic flow prediction model and decision-making method. Scientific Reports, vol. 12. DOI: https://doi.org/10.1038/s41598-022-24469-y [4] Carl Goves, Robin North, Ryan Johnston, Graham Fletcher. 2016. Short Term Traffic Prediction on the UK Motorway Network Using Neural Networks. Transportation Research Procedia, vol. 13, 184-195. DOI: https://doi.org/10.1016/j.trpro.2016.05.019 [5] Adriana-Simona Mihaita; Zac Papachatgis; Marian-Andrei Rizoiu. 2020. Graph modelling approaches for motorway traffic flow prediction. 2020. IEEE 23rd International Conference on Intelligent Transportation Systems (ITSC). DOI: https://doi.org/10.1109/ITSC45102.2020.9294744 [6] Sayed A. Sayed, Yasser Abdel-Hamid, and Hesham A. Hefny, 2022. Artificial Intelligence-Based Traffic Flow Prediction: A Comprehensive Review. Pre-review. DOI: http://dx.doi.org/10.21203/rs.3.rs-1885747/v1. [7] Scikit-learn: https://scikit-learn.org [8] XGBoost: https://xgboost.ai/ 44
Empowering Youth on Smart Cities with AI Solutions to Community and Urban Challenges Towards SDG 11 Abstract / Povzetek Achieving Sustainable Development Goal 11 — ensuring cities are inclusive, safe, resilient, and sustainable — remains a pressing global priority. In this pursuit, Artificial Intelligence (AI) has emerged as a transformative driver of urban innovation, enabling policymakers, academic institutions, and industry stakeholders to make data-driven decisions for complex urban systems such as housing, transportation, energy, and infrastructure. Despite its potential, the vast scale, variety, and fragmentation of urban data, coupled with the rapid evolution of AI technologies, create significant challenges in converting SDG 11-related information into practical solutions. This paper reports on the results of the AI4SDG11 programme, which combined expert community building, knowledge exchange, and competitive challenges. The programme brought together 50 students and 30 startups I 15 locations worldwide, to develop AI-driven solutions targeting key aspects of urban sustainability. Using diverse machine learning techniques, participants addressed challenges including intelligent mobility systems, efficient waste management, smart and efficient urbanism, and climate-resilient urban planning. Conducted in 2025, this initiative formed part of a youth-focused innovation challenge co-organized by AI in Africa, the International Research Centre on Artificial Intelligence (IRCAI), and GITEX, with the goal of promoting interdisciplinary innovation and strengthening regional AI capacity for sustainable urban development. Keywords / Ključne besede Machine learning, text mining, large language models, community engagement, urbanism, mobility, AI competition, AI Community 1 Introduction Established by the United Nations as an essential goal for the forthcoming 2030, the Sustainable Development Goal 11 (SDG 11) — "Make cities and human settlements inclusive, safe, resilient and sustainable" — reflects a critical global commitment to improving urban living conditions amid increasing urbanization, population growth, and environmental stress. With more than half of the world's population now residing in cities—and projections estimating two-thirds by 2050—the urgency of building sustainable urban environments has never been better fit. In this context, AI has emerged as a transformative tool capable of reshaping how cities are planned, managed, and experienced. AI technologies offer powerful capabilities to harness vast amounts of urban data, generate predictive insights, and support evidence-based decisionmaking. From optimizing public transportation systems to monitoring air quality, improving waste management, and enabling climate-resilient infrastructure, AI is at the forefront of innovative urban solutions worldwide. However, the deployment of AI in support of SDG 11 varies significantly across regions, influenced by differences in digital infrastructure, data availability, institutional capacity, and local priorities [1]. In Africa, AI is increasingly being applied to address urban informality, mobility challenges, and infrastructure gaps. For instance, AI-powered geospatial mapping tools are being used to identify informal settlements in rapidly growing cities such as Nairobi and Lagos, helping governments to improve service delivery and urban planning [2]. In North African cities, machine learning models have been developed to optimize water distribution in drought-prone areas and to improve traffic flow in congested urban corridors. AI is also being tested for predictive waste collection and smart energy use in off-grid communities. These solutions are particularly valuable in regions where resources are limited, and where rapid urban growth creates pressure for low-cost, scalable interventions [2]. On the other hand, in Europe, AI applications in cities often focus on enhancing sustainability, efficiency, and citizen engagement. Examples include real-time public transport optimization in cities like Helsinki and Barcelona [3], AI-based air pollution forecasting in Paris [4], and intelligent energy management †Corresponding author Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia © 2025 Copyright held by the owner/author(s). http://doi.org/10.70314/is.2025.sikdd.18 Thiago Gomes Marcilio, Anthony C. de Novaes Silva CIAAM, C4AI, Univ. of São Paulo São Paulo, Brazil [email protected] [email protected] Mustafa Zaouini†, Lee Chana, Ruben Frank, Kim August AI in Africa Johannesburg, South Africa [email protected] Joao Pita Costa, Davor Orlic, Mihajela Črnko IRCAI, Quintelligence Ljubljana, Slovenia [email protected] Yousef Rahmani ToumAI Rabat, Morocco [email protected] Rayan Kassis, Swethal Kumar EnergyAED London, UK [email protected] Luka Stopar Solvesall Ljubljana, Slovenia [email protected] Sohaib Souss, Wahid Laleeg, Yassine Bounouader SLTVERSE Casablanca, Morocco [email protected] Asmae Lamgari, Maroja Zoubir, Hajar Doukhou University Mohammed V (UM5) Rabat, Morocco [email protected] Ouidad Mochariq, Zahira Elmelsse, Chaimae Fadil ENSA National School of Applied Sciences (ENSA-M) Marrakesh, Morocco [email protected] 45
systems in smart buildings across the Netherlands and Germany [5]. Many European municipalities are also investing in AIdriven participatory governance platforms, enabling datainformed urban policymaking that incorporates citizen feedback [5]. Furthermore, [6] highlights how AI can extract and analyze news media information to enhance knowledge and understanding of water-related extreme events, supporting improved disaster risk reduction.. This paper presents the outcomes of a collaborative youth AI innovation programme, including AI mentorship and challenges aimed at exploring the impact of AI on SDGs. It builds on the related initiative initiating the programme in 2026 under the focus of Water Sustainability to progress SDG 6 (see [7] and [8]), and refocuses the approach addressing SDG 11-related problems through applied machine learning solutions. The initiative brought together 50 students and 20 professors across 10 research institutions in North Africa, as well as 30 AI startups and domain experts worldwide culminating in 30 projects and initiatives tackling real-world urban challenges. By leveraging AI and data science, these teams addressed issues ranging from urbanism and mobility to waste management and climate resilience—drawing on lessons and methods from both African and European contexts. The competition, co-organized by AI Africa and IRCAI in a collaboration with GITEX, short for Gulf Information Technology Exhibition, being one of the world’s largest technology and innovation events, held annually in Dubai, United Arab Emirates. The event held in May 2024 [7], served as a model for interdisciplinary, cross-regional collaboration in the pursuit of sustainable urban futures. Figure 1: Screenshot of the AI engine ToumAI, winner of the AI4SDG11 startup competition at the inaugurating edition of GITEX Europe, Berlin, as a prime example of the relevance of languages in the resilience of cities and communities 2 AI4SDG Programme Methodology The AI4SDG Programme, spearheaded by IRCAI under the auspices of UNESCO, in collaboration with AI in Africa and GITEX, is a transformative initiative designed to harness artificial intelligence to address the United Nations Sustainable Development Goals (SDGs). With a focus on capacity building, entrepreneurship, and ethical AI deployment, the programme connects technological innovation with global sustainability challenges, particularly in the Global South. At the core of AI4SDG is a multi-pronged approach integrating certified training, competitive innovation events, and startup acceleration. Launched through global showcases and pitch competitions at major GITEX events across Africa, Asia, Europe, and the Middle East, the initiative provides a dynamic platform for students, researchers, and entrepreneurs to ideate, prototype, and scale AI solutions aligned with specific SDGs. Previous editions have focused on Water Sustainability (SDG 6) and Sustainable Cities and Communities (SDG 11), while the 2026 programme will extend to all 17 SDGs. The key components include: • Research2Startup Competition: A 4–6 week programme blending AI education, design thinking, and acceleration tracks for startups and university spinouts, culminating in regional and global pitch events. • Certified AI for SDG Training: Professional certification tracks for corporate teams, startup founders, and SMEs, focusing on topics like large language models, AI governance, ethical data practices, and generative AI applications. • AI4SDG Lab Accelerator: A 3–6 month cohort-based programme supporting university-originated AI startups through mentorship, technical workshops, and investor networking, culminating in a high-profile Demo Day at GITEX Global. The programme not only equips participants with practical AI competencies but also facilitates access to global networks, funding opportunities, and collaboration through GITEX’s innovation ecosystem. It champions responsible AI development by emphasizing ethics, transparency, and inclusivity, while offering tangible incentives such as certifications, cash prizes, MVP co-development and impactful international exposure through IRCAI and GITEX channels. In doing so, AI4SDG acts as a catalyst for fostering the next generation of AI-driven changemakers committed to creating impactful, scalable solutions for a sustainable future. 3 AI-enabled Innovation Advancing SDG11 The joint IRCAI, AI in Africa and GITEX competition served as a global platform for surfacing innovative AI-driven solutions to SDG 11 challenges, bridging the ideas of PhD researchers in North Africa with the entrepreneurial agility of startups worldwide. Among the standout innovations emerging from the competition were AI-powered geospatial mapping systems for monitoring informal settlements, predictive analytics for optimizing urban transport routes in congestion-prone cities, and machine learning models for forecasting waste generation to improve collection efficiency. Several projects addressed climate resilience, including early-warning systems for urban flooding and AI-assisted tools for assessing heat island effects and guiding green space planning. From energy-efficient building design algorithms to citizen engagement platforms that use natural language processing for policy feedback, the competition highlighted the breadth of AI’s potential to make cities more sustainable and inclusive. By uniting academic depth with market-ready solutions, the initiative not only identified promising prototypes but also laid the groundwork for scalable interventions adaptable to diverse urban contexts.. ToumAI. A holistic multilingual AI platform designed to bridge the digital divide in Africa by enabling voice-driven customer experiences in low-resource languages, advancing SDG 11. Built on a compound AI structure that saves computing 46
power compared to foundational LLMs, the system supports speech-to-text, text-to-speech, emotion analysis, churn detection, and predictive insights across African dialects such as Swahili, Amharic, Yoruba, and Darija. By integrating AI-powered voice agents, IVR optimization, and multilingual analytics, ToumAI delivers inclusive, real-time, and cost-effective communication for telco, banking, and transport sectors (see Figure 1). Its innovation lies in industrializing underrepresented African languages for AI applications, ensuring accessibility for populations historically excluded from the AI revolution. AED EnergyAED. An AI-enabled renewable energy storage system that converts electricity into high-temperature heat (up to 800°C) using salt-based thermal bricks, providing 24/7 clean power and heat without combustion. Unlike batteries or diesel, the system delivers up to 24 hours of dispatchable energy at lower cost, using safe, stable, and modular 10MWh units. Applications include microgrids, telecoms, industrial heat, and desalination, making it particularly suited for regions with unreliable energy supply. By enabling baseload renewable energy, AED Energy strengthens critical infrastructure and advances SDG 11 while reducing dependence on diesel. SolvesALL Mobility. Delivery district planning and optimization machine learning tools that support smarter urban logistics impacting the sustainable of cities and communities. Its Postal POI system uses algorithms to automatically design delivery districts, balancing workload, reducing overlap, and minimizing travel time. Leveraging GPS trace analysis, staypoint detection, regression models, and crowdsourced field data, the system learns delivery micro-locations, service times, and accessibility factors (e.g., stairs, obstacles). By integrating these AI-driven insights, SolvesAll enables cost savings, operational efficiency, and improved registry accuracy—demonstrated by expected multimillion-euro annual savings for postal operators— while offering scalability to sectors such as waste management and ATM/vending machine logistics. Figure 2: Screenshot of the SLTverse engine, winner at the AI stage of GITEX Africa 2025 SLTverse. This smart city solution introduces an AI-powered travel app that supports SDG 11 by enhancing safety, sustainability, and cultural engagement in tourism. At its core is an AI Route Advisor that leverages structured mobility data— spanning cost, CO₂ emissions, safety, time, and distance—to recommend optimal transport options. This is strengthened by a Retrieval-Augmented Generation (RAG) framework, which combines vector search, large language models, and workflow orchestration to deliver fast, contextual, and multilingual guidance (see screenshot at Figure 2). The system’s AI assistant adapts to real-time inputs such as weather, safety alerts, and user preferences, ensuring tailored and secure travel recommendations. Beyond mobility, the platform enriches tourism through VR-based storytelling with avatars narrating site histories, and employs metadata-driven personalization supported by visual analytics (route maps, CO₂ vs. cost comparisons, safety heatmaps). Collectively, these AI innovations position the app as a smart city enabler that aligns sustainability, cultural engagement, and traveler well-being. SOBEK. A federated AI system for flood resilience that addresses the lack of early-warning systems in rapidly urbanizing African cities. Unlike centralized models, it applies federated learning to collaboratively improve predictions while preserving data privacy and sovereignty. Local nodes train specialized models—LSTMs for weather series, GNNs for hydrological networks, and U-Nets for satellite imagery—using geospatial, meteorological, and historical flood data. Model updates are aggregated with FedAvg and refined through station similarity graphs to capture regional hydrological patterns. Despite challenges of data heterogeneity and low connectivity, Sobek delivers more accurate flood seasonality, year, and magnitude predictions, enabling timely early warnings, urban planning, and disaster resilience across Africa. Ecoguardians. This initiative introduces an AI-powered system to optimize water-saving advertisements in Morocco, advancing SDG 11 (Sustainable Cities and Communities). By analyzing diverse campaign content (videos, images, text, social media engagement, and survey data), the system identifies what makes ads effective and generates improved variations. It integrates computer vision (CNNs) for visual features, language models (BERT/GPT) for text and sentiment, predictive models (XGBoost/Random Forest) for engagement forecasting, and GANs for generating impactful ad variations. Ethical and datadriven personalization ensures campaigns remain responsible, transparent, and locally relevant. Early prototypes show measurable engagement gains, empowering cities to run evidence-based, AI-enhanced awareness campaigns that strengthen sustainable water use. 4 Conclusions and further work The integration of AI with the SDGs represents a critical frontier in global innovation, particularly as we confront complex challenges in health, education, climate, and urbanization. The AI4SDG programme, as implemented through the collaboration of IRCAI, AI in Africa, and GITEX, demonstrates a strategic and scalable model for aligning technological advancement with sustainable impact. By combining certified training, research-to-startup pathways, and accelerator programs, AI4SDG empowers diverse stakeholders—from students and researchers to entrepreneurs and SMEs—to develop responsible, ethical and context-sensitive AI solutions across the 17 SDGs. One of the programme’s most significant contributions lies in its ability to bridge the gap between academic research and realworld application, particularly in the Global South. Through its global reach and multi-region engagements, AI4SDG not only promotes responsible AI development but also facilitates access to funding, mentorship, and global markets, thereby amplifying the reach and effectiveness of AI for social good. However, while the AI4SDG11 programme has laid a robust foundation, several 47
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Čibej Table 1: Word forms in ILS 1.0 by agreement. Pronunciation Number of Forms % /l/ 117,459 67.73 /u “/ 23,884 13.77 Both 12,160 7.01 Both | /l/ 11,205 6.46 Both | /u “/ 7,051 4.07 /l/|/u “1,660 0.96 Total 173,419 100.00 Figure 1: Extraction of character-level 𝑛 -gram features for the pre-consonant lin the word gledalka. forms (highlighted in gray), the annotators agree on the pronunciation of pre-consonant l. They disagree in 11% of the examples, with one annotator allowing for both pronunciation variants and the other allowing for only one pronunciation. Complete disagreement is present only in less than 1% of the examples. We use the 153,503 forms with complete agreement as training data for machine-learning models as described in the following sections. It should be noted, however, that while ILS 1.0 is the largest open-access dataset on pre-consonant lpronunciations, it is not completely representative of language use in general (with annotations by only 5 linguists with a background in translation and Slovene studies; these can be biased towards linguistic rules that might not reflect real language use). Despite this, the dataset is robust enough to help disambiguate the more obvious examples (such as alge, IPA: /"a:lgE/, and polž, IPA: /"pO:u “S/). 3 Feature Selection To some extent, the pronunciation of pre-consonant ldepends on the preceding and subsequent graphemes, 5 so we use characterlevel 𝑛 -grams as features for prediction. For each pre-consonant lin each word form, we identify the 𝑛 -grams (1 ≤𝑛≤ 5) in its direct left/right surroundings as shown in Figure 1 (see footnote 6). We include word boundary markers (#) to discriminate between word-initial and word-final 𝑛 -grams. We also perform the same extraction on robust and finegrained C+V representations of each word form.6 5 The Slovenian Normative Guide 8.0 (Pravopis 8.0, see https://pravopis8.fran.si), for instance, states that a pre-consonant lpreceded by the grapheme ois often characterized by the / u “ / pronunciation; this is true of words that historically used the syllabic l(e.g. polh IPA: / "pO:u “x / ‘dormouse’; volk IPA: / "vO:u “k / ‘wolf’). However, there are exceptions as not all ol 𝑛 -grams originate from the syllabic l(e.g., polkovnik IPA: /pOl"kO:u “nik/ ‘colonel’; voltaža IPA: /vOl"ta:Za/ ‘voltage’). 6 In the robust C+V form, all consonant graphemes are substituted with Cand all vowel graphemes with V. In the finegrained C+V form, consonant graphemes were generalized into more finegrained categories, e.g. graphemes denoting Slovene sonorants (M), voiced (G) and voiceless obstruents (K), foreign consonants (X), etc. Table 2: Contingency table for the general 𝑛 -gram cwhen following a pre-consonant l. Pronunciation → ↓Presence /l/ /u “/ /l/+/u “/ Yes 2,653 1,847 5,980 No 114,898 22,045 6,180 Table 3: A sample of statistically significant general character-level 𝑛-grams. 𝑛-Gram 𝜒2p V 𝑟|𝑚𝑎𝑥 |Category c 38,199.59 **** 0.499 178.81, /l/, No post-l n 29,081.52 **** 0.435 79.27, /l/, No post-l ce 16,003.46 **** 0.323 118.30, /l/, No post-l o 77,025.17 **** 0.708 227.83, /l/, No pre-l po 48,241.29 **** 0.560 193.98, /l/, No pre-l a 16,592.50 **** 0.329 -79.85, /l/, No pre-l We extract a total of 8,082 different general 𝑛-grams (consisting of actual graphemes; 3,041 in pre-lposition, 5,541 in post-l position), 116 different robust C+V 𝑛 -grams (65 preand 51 postl), and 603 different finegrained C+V 𝑛 -grams (262 preand 341 post-l). For each 𝑛 -gram, we compile a contingency table. For instance, Table 2 shows the occurrences of the general 𝑛 -gram c in the position directly following a pre-consonant l(e.g., morilca, ‘murderer’, masculine common noun, genitive singular form) depending on the pronunciation of the pre-consonant l. In order to determine statistically significant features that help discriminate between different pronunciations, we performed a series of Pearson’s 𝜒2 tests [12] and corrected for family-wise error rate with the Holm-Bonferroni method [7]. We calculated Cramér’s V [6] as the measure of effect size. 7 This resulted in a total of 4,263 statistically significant features (1,856 pre-lgeneral and 1,794 post-lgeneral 𝑛 -grams; 60 pre-land 40 post-lrobust C+V 𝑛 -grams; 242 pre-land 271 post-lfinegrained C+V 𝑛 -grams). Several statistically significant pre-lgeneral 𝑛 -grams are shown in Table 3. 8 The table shows the values of the 𝜒2 statistic and Cramér’s V, the p-value representations, the maximum absolute value of Pearson’s residuals (and its position in the contingency table), and the category of the 𝑛 -gram (post-lor pre-l). With the exception of the a 𝑛 -gram, which is more indicative of the / l / pronunciation, the others indicate one of the other two options (/ u “ /; or / l /+/ u “ /). The results also confirm the statement found in the Slovenian Normative Guide 8.0 that the ographeme in pre-l position is strongly indicative of the /u “/ pronunciation. 4 Prediction and Evaluation We compiled a custom vectorizer based on the identified features. The vectorizer scans each input word form (along with its Multext-East v6 morphosyntactic tag 9 ) for all occurrences of 7 We calculate Cramér’s V as √︂𝜒2 𝑁∗𝑑𝑚𝑖𝑛 , where 𝜒2 is the Pearson’s 𝜒2 statistic, 𝑁 is the total sample size, and 𝑑𝑚𝑖𝑛 is the minimum dimension of the contingency table. 8 For all tests, the degrees of freedom (df) were equal to 2 and the total sample size (N) was equal to 153,603. The p-values should be interpreted in the following manner: **** →p≤0.0001; *** →p≤0.001; ** →p≤0.01; * →p<0.05 9 Multext-East v6 Morphosyntactic specifications: https://nl.ijs.si/ME/V6/msd/html /msd-sl.html 54
Prediction of Pre-Consonant lin Slovene Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Table 4: Model performance based on 10-fold crossvalidation. Model A BA P R F1 LinearSVC 86.08 72.39 69.26 55.39 61.54 Multin. NB 77.29 69.54 33.33 81.84 47.36 kNN (k=5) 85.91 73.30 64.11 62.98 63.53 Majority 76.53 - - - - pre-consonant l, extracts the surrounding 𝑛 -grams, converts the morphosyntactic tag into 146 morphosyntactic features, and represents the occurrence as a 4,409-dimensional vector of {0,1} values (with 0 and 1 indicating the absence or presence, respectively, of the 𝑛 -gram in the direct surroundings of the pre-consonant / l /). We compile a total of 153,503 vectors in this way and use the scikit-learn Python library [13] to train several models for a classification task with three classes: the goal is to correctly predict whether a pre-consonant lis pronounced as / l /, / u “ /, or both. 4.1 Automatic Evaluation We trained three different models: a Linear Support Vector Classifier (LinearSVC), a Multinomial Naïve Bayes Classifier (Multin. NB), and a 𝑘 Nearest Neighbors Classifier (kNN) and evaluate their performance with a 10-fold cross-validation (with a stratified random test set of word forms). The results are shown in Table 4. 10 The worst performing model is the Multin. NB classifier, which barely achieves an above-baseline accuracy and a very low F1-score compared to the other two classifiers, although its recall is much higher. In terms of balanced accuracy and F1-score, the best model is the kNN classifier. However, it seems that the algorithm is not the most suited for this type of data. It performs similarly to the LinearSVC classifier, but if we compare the sizes of the resulting models, it becomes apparent that the LinearSVC model is much more efficient (with a size of approximately 100 kB) compared to the kNN model, which is overly inflated (with a size of more than 2 GB), possibly indicating overfitting.11 Because the LinearSVC model is the most viable, we analyze its performance in more detail. Table 5 shows the confusion matrix for the classifications of the LinearSVC model on a stratified test set (20% of the total 153,503 dataset instances). The model seems to lean more towards the most frequent category (/ l /) in its predictions, with approximately 30% of / u “ / and / l /+/ u “ / instances being misclassified as / l /, whereas 94% of the / l / instances are classified correctly. It seems that instances allowing both pronunciations are very rarely misclassified as / u “ / (only 1%). It should also be noted that the instances of / l /+/ u “ / misclassified as either / u “ / or / l / are not entirely incorrect, just incomplete. Compared to the rule-based approach (which classifies everything as / l /), the model performs quite well in terms of / l /+/ u “ / and / u “ / instances and sacrifices only 6% of its accuracy for / l / instances. In order to determine any future improvements to the model, we analyze some of the misclassified examples in more detail in Section 4.2. 10 A, BA, P, R, and F1 refer to accuracy, balanced accuracy, macro-precision, macrorecall and macro-F1, respectively. 11 We also ran a 10-fold cross-validation using only 𝑛 -gram features (no morphosyntax). The performance of the models was slightly worse, e.g. for LinearSVC: A = 85.05, BA = 69.14, P = 68.94, R = 46.85, F1 = 55.76. Table 5: Confusion matrix for the Linear Support Vector classifier. True → ↓Predicted /l/ /u “/ /l/+/u “/Í /l/22,006 1,495 729 24,230 /u “/ 1,071 2,764 31 3,866 /l/+/u “/ 434 519 1,672 2,625 Í23,511 4,778 2,432 - 4.2 Manual Evaluation We performed a manual analysis of the misclassified examples to determine whether there are any patterns to the errors that could be help further improve the model with additional features. Due to space limitations, we only focus on the most obvious problems in this paper. In the examples where the / l / pronunciation was misclassified as / u “ /, many words contain a pre-consonant lfollowed by the grapheme d(kaldera ‘caldera’, buldožerski ‘pertaining to a bulldozer’, heraldičen ‘heraldic’, bodibilder ‘bodybuilder’). The majority of these examples are pronounced with / l /, with the exception of words like dopoldne ‘late morning’, popoldanski ‘pertaining to the afternoon’, where the pre-consonant lis preceded by an ographeme. This could indicate that an additional 𝑛 -gram feature should be added (the lalong with its preceding and subsequent graphemes: old,ald, etc.). This could resolve some other misclassifications, such as impulziven ‘impulsive’ and pulzirajoč ‘pulsating’, where words with the ulz combination are never pronounced as / u “ /, but words with olz are (e.g., polzeti ‘to slip’). The emergence of such patterns in the misclassifications is a good sign that the classifiers might benefit from a joint pre-l/post-l feature. This will be explored in future versions. Many of the instances in which the / u “ / was misclassified as / l / contain compound words with the element pol ‘semi, half’: polnag ‘half-naked’, polfinale ‘semi-final’, polpuščava ‘semi-desert’. Because the element pol is always pronounced with / u “ /, this is also true of derived compound words. However, the 𝑛 -gram features used offer no indication of morpheme boundaries, so these misclassifications can be expected. Additional 𝑛 -gram features could be extracted from the accentuated forms of words. In some examples, the accentuation diacritic can disambiguate the pronunciation of the subsequent pre-consonant l. For instance, dólnji ‘pertaining to something that is downwards or downstream’ and prestólničen are pronounced with / l /, whereas tôlšča ‘blubber’ and pôlhográjski ‘pertaining to the town of Polhov Gradec’ are pronounced with / u “ /. However, accentuation is rarely written in Slovene and is much more difficult to assign automatically compared to morphosyntactic features. Relying on too many features that are not easily extractable would make the model less robust (more on this in Section 5). 5 Conclusion We presented a machine-learning approach to improve the accuracy of phonetic transcriptions of Slovene words that contain the ambiguous pre-consonant l. While the method does improve accuracy (86% over a majority baseline of cca. 76%) by using very simple character-level 𝑛 -gram and morphosyntactic features, it does not resolve the problem entirely. Aside from several exceptions in language use which are difficult to predict (e.g. gasilci, 55
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Čibej čistilka; both pronounced with / l / even though the majority of words ending with -ilec and -ilka in the dataset can be pronounced with either / l / or / u “ /), the analysis of misclassified examples has shown several potential future steps that can be implemented to further improve the performance of the models. First, several additional features should be tested. Some of the features are simple, such as word length or number of syllables in word (which could potentially help to correctly classify words such as volk and polh; short words where the pre-consonant lis pronounced as / u “ /). The relative position of the pre-consonant lin the word could also potentially be helpful. Several more complex features could also be added, such as word formation relations and morpheme boundaries to help disambiguate, for instance, decimal-ka ‘decimal number’, which is derived from the adjective decimalen ‘pertaining to decimal numbers’ and is pronounced with /l/; and mor-ilka ‘murderer (feminine)’, which is derived from the verb moriti ‘to murder’ and can be pronounced as either / l / or / u “ /). Taking into account the accentuated form of the word could also help: for instance, the ôl accentuation – vôlk ‘wolf’, pôlh ‘dormouse’ – indicates the / u “ / pronunciation, while the ól accentuation is indicative of the / l / pronunciation, e.g. pólka ‘polka’). However, more complex features cannot be extracted from the word form itself, so making the model too heavily reliant on external linguistic knowledge would sacrifice its robustness and usefulness for unseen words. We will explore these options in our future work but we will first focus on the simplest features to determine the upper boundary of accuracy that can be achieved based solely on the word form and its morphosyntactic features. We will perform additional statistical analyses on 𝑛 -grams containing the pre-consonant las well, and once the optimal model is achieved, it will also be evaluated on previously unseen words containing the pre-consonant lthat have not been included in the ILS 1.0 dataset. The results will hopefully also provide more interesting material for further linguistic analyses (such as exceptions to the rules). As already mentioned, the ILS 1.0 dataset does not necessarily accurately reflect the linguistic landscape of pre-consonant lpronunciation in Slovene words, and more annotations along with perceptive tests and surveys are required. The pronunciations will be manually validated as part of the work on the Digital Dictionary Database of Slovene [8], the largest machine-readable open-access database of Slovene linguistic and lexicographic data. The pronunciations will also be cross-referenced with the recordings from the GOS Corpus of Spoken Slovene [18], which contains real recordings of Slovene speech and can contribute towards a more accurate distribution of different pronunciations for individual lexemes (e.g., how many occurrences of / glE"da:u “ka / or / glE"da:lka /), along with any potential relevant metadata (for instance, whether the pronunciation depends on the region the speaker originates from). The models can then be re-trained on new data and further improved to better reflect real language use. The models will be implemented into the Slovene IPA/X-SAMPA Grapheme-to-Phoneme Converter as part of the Pregibalnik tool for automatic Slovene lexicon expansion, which is available under a Creative Commons BY-SA 4.0 license.12 12 The best-performing LinearSVC model (and the accompanying code) for the prediction of pre-consonant lpronunciation is available on Github: https://github.c om/jakacibej/sikdd2025_predicting_preconsonant_l Acknowledgements The research presented in this paper was carried out within the research project titled Basic Research for the Development of Spoken Language Resources and Speech Technologies for the Slovenian Language (J7-4642), the research programme Language Resources and Technologies for Slovene (P6-0411), and the CLARIN.SI Research Infrastructure (I0-E004), all funded by the Slovenian Research and Innovation Agency (ARIS). The author also thanks the anonymous reviewers for their constructive comments. References [1] Jaka Čibej. 2024. Dataset of annotated slovene words with pre-consonant l ILS 1.0. Slovenian language resource repository CLARIN.SI. (2024). http://h dl.handle.net/11356/2025. [2] Jaka Čibej. 2023. Leksikon besednih oblik sloleks. poročilo projekta razvoj slovenščine v digitalnem okolju aktivnost ds1.3. Development of Slovene in a Digital Environment. (2023). https://www.cjvt.si/rsdo/wp-content/upload s/sites/18/2023/06/RSDO_Kazalnik_Sloleks_v2.pdf. [3] Jaka Čibej. 2024. Predicting pronunciation types in the sloleks morphological lexicon of slovene. In Data mining and data warehouses (SiKDD): Information Society (IS) 2024 - proceedings of the 27th International Multiconference: volume C. Institut „Jožef Stefan“, 23–26. https://is.ijs.si/wp-content /uploads/2024/11/IS2024_Volume-C.pdf. [4] Jaka Čibej. 2025. Statistična analiza izgovora črke l v slovenskem oblikoslovnem leksikonu sloleks. Jezikoslovni zapiski, 31, 1, (maj 2025), 37–54. doi:10.3986 /JZ.31.1.03. [5] Jaka Čibej et al. 2022. Morphological lexicon sloleks 3.0. Slovenian language resource repository CLARIN.SI. (2022). http://hdl.handle.net/11356/1745. [6] Harald Cramér. 1946. Mathematical Methods of Statistics.Princeton Mathematical Series. Vol. 9. Princeton University Press. [7] Sture Holm. 1979. A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6, 2, 65–70. [8] Iztok Kosem, Simon Krek, and Polona Gantar. 2021. Semantic data should no longer exist in isolation: the digital dictionary database of slovenian. In 9th EURALEX International Congress "Lexicography for Inclusion", 81–83. https://elex.is/wp-content/uploads/2021/09/Semantic-Data-should-no-l onger-exist-in-isolation-the-Digital-Dictionary-Database-of-Slovenian _Kosem-Krek-Gantar_EURALEX2020.pdf. [9] Janez Križaj, Simon Dobrišek, Aleš Mihelič, and Jerneja Žganec Gros. 2022. Uporaba postopkov strojnega učenja pri samodejni slovenski grafemskofonemski pretvorbi. In Jezikovne tehnologije in digitalna humanistika: zbornik konference 2022. Inštitut za novejšo zgodovino, 248–251. https://nl.ijs.si/jtdh 22/pdf/JTDH2022_Proceedings.pdf. [10] Xavier Marjou. 2021. Gipfa: generating ipa pronunciation from audio. In eLex 2021 Conference Proceedings, 588–597. https://elex.link/elex2021/wp-co ntent/uploads/2021/08/eLex_2021_38_pp588-597.pdf. [11] Tanja Mirtič. 2019. Glasoslovne raziskave pri pripravi splošnega razlagalnega slovarja. In Slovenski javni govor in jezikovno-kulturna (samo)zavest. Znanstvena založba Filozofske fakultete, 81–90. https://centerslo.si/wp-con tent/uploads/2019/10/Obdobja-38_Mirtic.pdf. [12] Karl Pearson. 1900. X. on the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, 50, 302, 157–175. eprint: https://doi.org/10.1080/14786440009463897. doi:10.1080/14786440009463897. [13] F. Pedregosa et al. 2011. Scikit-learn: machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830. [14] Uwe Reichel, Hartmut R. Pfitzinger, and Horst-Udo Hain. 2008. English grapheme-to-phoneme conversion and evaluation. In Speech and Language Technology 11, 159–166. https://www.phonetik.uni-muenchen.de/~reichelu /publications/ReichelPfitzingerHainSASR2008.pdf. [15] Anja Schüppert, Wilbert Heeringa, Jelena Golubovic, and Charlotte Gooskens. 2017. Write as you speak? a cross-linguistic investigation of orthographic transparency in 16 germanic, romance and slavic languages. English. From semantics to dialectometry, 32, 303–313. isbn: 9781848902305. [16] Hotimir Tivadar. 2004. Priprava, izvedba in pomen perceptivnih testov za fonetično-fonološke raziskave (na primeru analize fonoloških parov). Jezik in slovstvo, 49.2, 2, 17–36. https://ojs.zrc-sazu.si/jz/article/view/14222. [17] Antal van den Bosch, Alain Content, Walter Daelemans, and Beatrice de Gelder. 1994. Analysing orthographic depth of different languages using data-oriented algorithms. In Proceedings of the 2nd International Conference on Quantitative Linguistics. [18] Darinka Verdonik et al. 2023. Spoken corpus gos 2.1 (transcriptions). Slovenian language resource repository CLARIN.SI. (2023). http://hdl.handle.net /11356/1863. [19] Jerneja Žganec Gros, Tanja Mirtič, Miroslav Romih, and Kozma Ahačič. 2022. Slovar izgovarjav OptiLEX. (1. e-izd. ed.). Založba ZRC. isbn: 978-961-050672-0. https://doi.org/10.3986/9789610506720. 56
Sequencing News Articles with Large Language Models within Enterprise Risk Management Context Žiga Debeljak† Jožef Stefan International Postgraduate School Ljubljana, Slovenia ziga.deb[email protected] Dunja Mladenić Department for Artificial Intelligence, Jožef Stefan Institute Ljubljana, Slovenia [email protected] Klemen Kenda Department for Artificial Intelligence, Jožef Stefan Institute Ljubljana, Slovenia klemen[email protected] Abstract This paper evaluates the capability of Large Language Models (LLMs) to reconstruct event timelines from unstructured news data. This capability is highly relevant for Enterprise Risk Management (ERM) applications, where the reconstruction and forecasting of coherent event trajectories are crucial for identifying, assessing, and predicting emerging risks and analyzing risk scenarios. In this study, we tasked twenty LLMs with chronologically ordering randomly shuffled business news articles for three distinct real-world event chains. To prevent simple date sorting, all explicit date markers were removed from the articles. The experiments were conducted under one unassisted and three assisted scenarios that provided the models with hints for the first, the last, or both the first and the last articles in the sequence. The results reveal a systematic variation in difficulty across the three tasks in addition to significant performance disparities among the models, with Grok 4 (xAI), GPT-5, o3 and o3-pro (all three OpenAI), and Gemini 2.5 Pro (Google) consistently outperforming other models practically across all tasks and prompting scenarios. As expected, prompting assistance with additional information systematically improved accuracy, especially for the models that performed poorly in the unassisted scenario. The high level of accuracy achieved by the top-performing models indicates a practical utility for real-world ERM applications. Keywords Large Language Models, News-Stream Sequencing, Temporal Reasoning 1 INTRODUCTION Within Enterprise Risk Management (ERM) practice, organizations monitor external developments also by analyzing streams of publicly available news. Each news article captures a momentary state of the political-economic environment, and by accurately structuring unordered information into a chronological narrative, organizations can better understand the evolution of events and the relationships that connect them. The reconstruction and forecasting of these event trajectories are important for identifying, assessing, and predicting emerging risks, especially within risk scenario analysis [10, 11]. The capability to build structured timelines from unstructured textual information is therefore of high relevance to ERM. LLMs are increasingly utilized in ERM for their ability to process and analyze unstructured textual data, including news articles, to identify and assess risks [1, 2, 3, 4, 5]. In the financial sector, applications include extracting sentiment from news to gauge market perception or identify reputational risks [3, 6, 7, 8], and identifying specific risk factors or events discussed in news and corporate disclosures [2, 4, 5, 9]. Existing literature mainly demonstrates LLMs' utility in analyzing individual or aggregated news items for tasks such as sentiment analysis, risk factor identification, or event detection, but the capabilities of the models to recover the temporal order and causal links among a sequence of discrete news items that describe an unfolding narrative are less directly explored. This paper aims to address this gap by investigating LLM performance in temporal-causal reasoning within news streams, a crucial aspect for understanding the dynamics of unfolding risk narratives. By investigating whether state-of-the-art commercial or opensource LLMs can reconstruct the chronological narrative of business-event chains from unordered news articles, this paper contributes to the field by: (a) systematically evaluating the performance of multiple LLMs on a challenging temporalreasoning task; (b) analysing the efficacy of diverse prompting strategies — both unassisted and assisted — in improving model accuracy; (c) providing insights into model-and-task dynamics, revealing substantial performance disparities, task-specific difficulty patterns, and the outsized gains weaker models receive from contextual hints; and (d) demonstrating the practical readiness of these technologies for ERM deployment. 2 RESEARCH METHOD Task Definition To evaluate the capabilities of LLMs, three event chains were constructed, focusing on: (1) Trump's Tariffs and EU [“Task_1”], (2) Gold Prices [“Task_2”], and (3) the UkraineRussia War [“Task_3”]. These topics were selected due to their significant relevance to the business environment. For each topic, ten articles were manually selected from the online editions of two reputable sources of financial and business information, published between March 1st and May 2nd, 2025. For the purpose of LLM processing, the raw text from the selected articles was extracted. To prevent temporal bias, explicit date indicators—such as full dates—were removed, and no two † Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia © 2025 Copyright held by the owner/author(s). http://doi.org/ 10.70314/is.2025.sikdd.4 57
articles shared the same publication date. Subsequently, the articles within each event chain were randomly shuffled, and this fixed random order was then applied to all models within the experiment. The primary task for the selected LLMs was to reconstruct the chronological sequence of news articles within three distinct event chains. This task was evaluated across four experimental scenarios: (1) an unassisted scenario [“Assist_No”], and three assisted scenarios providing the (2) first [“Assist_First”], (3) last [“Assist_Last”], or (4) both first and last [“Assist_FirstLast”] articles in the sequence. In the unassisted scenario, the LLMs were required to determine the correct chronological order of the articles without any external information regarding their placement. In the assisted scenarios, the models were provided with hints within the user prompt. Specifically, for the Assist_First and Assist_Last scenarios, the prompt identified the article occupying the initial or final position, respectively. In the Assist_FirstLast scenario, the LLMs were given the identifiers for the articles that correspond to the beginning and end of the chronological sequence. The required output from the LLMs was a reconstructed timeline of the news articles. For each position in the timeline, the following information was mandated: (i) the article's identification number, (ii) the article's title, (iii) a brief justification for its placement relative to the preceding article, and (iv) a brief justification for its placement relative to the subsequent article. The models were required to provide a structured output in JSON format. Prompt Engineering Prompt engineering included manual drafting, testing on different models, and optimization both with LLMs (GPT o3 and Gemini 2.5 Pro) as well as manually, in several iterations. In the end, an effective user prompt was developed which worked reasonably well for all selected models. The main challenges with regard to the design of prompts were: (a) stimulating a systematic approach to causal reasoning, which was considered to be mainly important for the non-reasoning models; (b) ensuring the output consisted of exactly ten distinct articles, with no repetitions or omissions; (c) enforcing the required output JSON schema; and (d) providing concise reasoning for the positioning of the observed articles. Within the user prompt, the models were explicitly instructed to use the following reasoning principles: (a) inferring sequences of events (how events described in different articles relate to each other over time), (b) causal reasoning (identifying cause-andeffect relationships between the content of different articles), (c) logical story progression (understanding how a narrative or situation typically develops or unfolds), (d) utilizing any implicit time references if available within the articles, and (e) using models’ general knowledge about events. Prompts with clear instructions about the guidelines for the reasoning process worked better than prompts without such instructions, even with models with strong reasoning capabilities. System prompts were not utilized, as the one-shot user prompt contained all necessary instructions for the models. The full user prompt is available from the authors. Selected LLMs and Experiment Execution Twenty different models by eight different providers were selected for this research, based on their expected capabilities with regard to the tasks, and their availability. Overview of selected models is shown in Table 1. Table 1: Selected LLMs # Model Provider: Model Name Context Window (tokens) Date Created 1 OpenAI: GPT-4.1 1.047k 14.04.2025 2 OpenAI: o3 200k 16.04.2025 3 OpenAI: o3-pro 200k 10.06.2025 4 OpenAI: gpt-oss-120b 131k 5.08.2025 5 OpenAI: GPT-5 400k 7.08.2025 6 Google: Gemini 2.5 Pro Preview 1.048k 7.05.2025 7 Google: Gemini 2.5 Flash Preview 1.048k 20.05.2025 8 xAI: Grok 3 Beta 131k 9.04.2025 9 xAI: Grok 4 256k 9.07.2025 10 Anthropic: Claude Sonnet 4 200k 22.05.2025 11 Anthropic: Claude Opus 4 200k 22.05.2025 12 Anthropic: Claude Opus 4.1 200k 5.08.2025 13 Meta: Llama 4 Maverick 1.048k 5.04.2025 14 Meta: Llama 4 Scout 1.048k 5.04.2025 15 Mistral AI: Mistral Medium 3 131k 7.05.2025 16 Mistral AI: Mistral Medium 3.1 262k 13.08.2025 17 Qwen: QwQ 32B 131k 5.03.2025 18 Qwen: Qwen 2.5 VL 32B Instruct 128k 24.03.2025 19 DeepSeek: DeepSeek V3 163k 24.03.2025 20 DeepSeek: R1 128k 28.05.2025 All models were accessed using the OpenRouter platform via the APIs. For models supporting this parameter, the temperature was set to 0.0 to ensure the most reliable and reproducible experimental results; otherwise, default parameters were used. There were 12 experiments executed: 3 different event topic chains (tasks) in 4 experimental scenarios (prompts) each, by using all 20 LLMs as shown in Table 1, thus resulting in 240 results (outputs). Experiments were executed on June 1st, 2025 with the models available on that date, and on August 19th, 2025 with the newer models. 3 EVALUATION AND DISCUSSION General Evaluation In terms of the output content, all models demonstrated strong performance in response to a standardized user prompt, successfully producing the requested ordered lists of news articles with all accompanying metadata. From a logical standpoint, the outputs from all models were accurate, presenting ordered lists that included all required supplementary information. Substantial variations in output quality were observed across the different models. This variation was also influenced by the three distinct tasks, which seemed to be of substantially different difficulty, with the first task being the most straightforward and the last presenting the most significant challenge. As anticipated, the implementation of assisted prompting strategies consistently enhanced the accuracy of the outputs for all models across all evaluated tasks. Regarding the output formatting, the majority of the models adhered to the specified JSON schema. Notable exceptions to 58
this were Claude models (models #10, #11 and #12), which occasionally deviated from the requested format by including a short introductory text. In these instances, the textual outputs were programmatically reformatted to conform to the required JSON structure. It is relevant to note that these three models are the only ones in the evaluation that do not natively support the Structured Output functionality, a factor that likely contributed to their formatting inconsistencies. Performance Metric To quantify the models’ performance with the given tasks, a robust evaluation metric was required. For this purpose, Kendall's rank correlation coefficient (“Kendall’s τ”, “τ”) was selected as the most appropriate measure. Kendall's τ is a nonparametric statistic that measures the ordinal association between two ranked lists. Its methodology is centered on comparing the concordance of all possible pairs of items within the sequences, yielding a score in the interval from -1 (perfect reversal) to +1 (perfect match). The focus on relative, pairwise ordering makes Kendall's τ exceptionally well-suited for a chronological sorting task, as the core challenge lies in correctly establishing which event occurred before another, which is precisely what the metric evaluates. An alternative metric, the sum of absolute Manhattan distances, was also considered but ultimately deemed less suitable. Its primary drawback is its sensitivity to the magnitude of displacement, which can produce misleading evaluations by heavily penalizing single items that are wildly out of place, while potentially under-penalizing a sequence with numerous smaller, local errors that might represent a poorer overall sort. Performance by Tasks and Scenarios The performance of each model, quantified by the Kendall’s τ, is detailed in Tables 2 and 3. Table 2 presents the coefficients organized by task (event chain), averaged across all experimental scenarios (prompts). Table 3, in turn, presents the coefficients organized by experimental scenario, averaged across all the tasks. The ranks in both tables were determined by averaging the performance rankings of all the models across individual tasks and scenarios. They largely correspond to the rankings based on average τ, but discrepancies may arise from variation in the scale and distribution of τ values across experiments. To contextualize these performance metrics, their relationship to pairwise accuracy is critical: within a 10-item sequence, a Kendall’s τ of 0.90, 0.80 or 0.50 indicates that approximately 95%, 90% or 75% of the 45 possible pairs are concordantly ordered, respectively. The aggregated results in Table 2 underscore two principal findings. First, a significant and systematic variation in task difficulty was evident, with Task_1 representing the simplest case and Task_3 the most demanding. This pattern held true for practically all the evaluated models and experimental scenarios. The performance differences indicating different task difficulty were substantial. For Task_1 and the unassisted scenario, the Kendall's τ values for the average, best model, and worst model performance were 0.78, 0.91 and 0.02, respectively. For Task_2, the values were 0.63, 1.00 and 0.16, and for Task_3, they were 0.02, 0.38 and 0.33. These findings clearly establish Task_3 as the most difficult of the three tasks evaluated. Note that a negative Kendall’s τ value indicates an inverse correlation between the predicted and true rankings, and a value around zero represents a random ordering. Second, the results show that the more recent versions and models with strong reasoning capabilities (models Grok 4, GPT-5, o3 and o3-pro, and Gemini 2.5 Pro) consistently outperform other models practically across all tasks. Table 2: Average Performance by Tasks (Kendall’s τ) Rank Model # Task_1 Task_2 Task_3 Avg. τ 1 9 0.96 0.98 0.70 0.88 2 2 0.94 0.94 0.56 0.81 3 5 0.96 0.99 0.49 0.81 4 3 0.94 0.93 0.52 0.80 5 6 0.94 0.96 0.52 0.81 6 8 0.93 0.79 0.43 0.72 7 12 0.94 0.70 0.41 0.69 8 20 0.83 0.82 0.50 0.72 9 7 0.84 0.89 0.48 0.74 10 11 0.93 0.67 0.36 0.65 Avg. top 5: 0.95 0.96 0.56 0.82 Avg. all 20: 0.85 0.71 0.36 0.64 The aggregated results in Table 3 underscore three principal findings. First, assisted prompting systematically improved the performance across all models and tasks, which is logical and expected since additional relevant information is provided to the models. Anchoring with known positions in the majority of cases helped the models to better position the remaining articles as well. Table 3: Average Performance by Scenarios (Kendall’s τ) Rank Model # Assist_ No Assist_ First Assist_ Last Assist_ FirstLast Avg. τ 1 9 0.75 0.88 0.90 0.99 0.88 2 2 0.69 0.88 0.76 0.93 0.81 3 5 0.73 0.84 0.81 0.87 0.81 4 3 0.72 0.87 0.76 0.85 0.80 5 6 0.57 0.93 0.84 0.90 0.81 6 8 0.48 0.81 0.78 0.81 0.72 7 12 0.48 0.66 0.73 0.87 0.69 8 20 0.66 0.75 0.64 0.82 0.72 9 7 0.54 0.73 0.81 0.87 0.74 10 11 0.48 0.64 0.66 0.82 0.65 Avg. top 5: 0.69 0.88 0.81 0.91 0.82 Avg. all 20: 0.47 0.67 0.63 0.79 0.64 Second, the provision of additional information proved more beneficial for the most demanding task (Task_3) than for the less demanding tasks (Task_1 and Task_2). For example, in the Assist_FirstLast scenario, the increase in average τ relative to the unassisted scenario was 0.13 for Task_1, 0.17 for Task_2, and 0.65 for Task_3. This finding follows logically from the models’ greater ability to identify the first and/or last article in simpler tasks by themselves: in Task_1, 15 of 20 models correctly identified the first position, while none identified the last position, in Task_2 9 models identified the first position and 4 identified the last position, and in Task_3 no model identified either position correctly. Third, the provision of additional information disproportionately benefited models that performed poorly in the unassisted scenario. For instance, on Task_3 — the most difficult task with 59
an average Kendall's τ of only 0.02 in the unassisted scenario — the Assist_First scenario yielded average and maximum performance improvements of 0.46 and 1.07, respectively. For the Assist_Last scenario, the corresponding improvements were 0.27 and 0.80, while for the Assist_FirstLast scenario they were 0.65 and 1.02. The results demonstrate that supplementing less capable models with limited key information can yield significant performance gains at these tasks. A qualitative examination of the models' reasoning justifications failed to yield systematic insights into their capacity to reconstruct accurate chronological sequences of articles. Although the generated rationales were generally logical and relevant, they frequently omitted crucial contextual information essential for correct chronological reasoning. This observation underscores the challenge that certain timelines may not be uniquely re-constructible due to insufficient contextual information. Furthermore, in some instances, the provided justification could plausibly support an alternative, yet equally valid, timeline. Moreover, this is compounded by the inherent challenge of discerning whether the provided reasoning justifications represent the model's actual inferential process or are merely a result of the post-hoc rationalization. 4 CONCLUSIONS AND FURTHER RESEARCH IDEAS This research provides insight into the practical application and inherent challenges of utilizing LLMs to sequence news streams in the context of ERM. The selected use cases are based on realworld, business-relevant event chains. A comparative analysis reveals significant performance disparities among the evaluated models across all tasks and experimental scenarios. Models with superior reasoning capabilities surpassed those with less developed abilities. The varying complexity of the presented tasks further accentuated these performance differences. Also, providing additional anchoring information disproportionately benefited models that performed poorly in the unassisted scenario. Five models, Grok 4 (xAI), GPT-5, o3 and o3-pro (all three OpenAI), and Gemini 2.5 Pro (Google), consistently outperformed all other models in practically every task and experiment scenario. The performance level achieved by these models demonstrates their practical utility for real-world ERM applications. This research has opened several promising areas for further research: (1) Benchmarking LLMs against human experts: A rigorous comparative study should be undertaken in which large LLMs and domain specialists (human experts) perform identical tasks under strictly matched contextual conditions. (2) Systematically varying model settings to probe “creativity” and reliability: Experiments that modulate the temperature and other model settings can clarify how stochasticity affects task performance and reliability. (3) Enabling models to request task-critical information: Instead of supplying predefined contextual information—such as the first and/or last article in a sequence—future studies might allow the model to query for the minimal supplementary data it deems most informative. This strategy would approximate an activelearning workflow and might even illuminate new modes for human-LLM collaboration. (4) Diagnosing mis-ordering errors through reasoning audits: To understand why models fail to reconstruct the correct temporal ordering of news articles, one could extract each model’s stated reasoning features for every placement decision, then have human experts or adjudicating LLMs rate their accuracy and relevance. Such audits would expose specific deficits in reasoning and could even inform targeted retraining regimes. (5) Experimenting with extended or interleaved event chains: Evaluating models on substantially longer sequences—or on mixtures of events drawn from multiple chains—would markedly raise task complexity and furnish a stringent benchmark of temporal-reasoning competence for business use cases. ACKNOWLEDGMENTS The authors acknowledge the use of LLMs during various stages of this research. These models provided support in tasks such as idea generation, text processing, prompt engineering, methodological exploration, and language optimization. While the LLMs contributed to enhancing efficiency and refining the presentation of this work, all conceptual frameworks, analyses, and interpretations remain the sole responsibility of the authors. REFERENCES [1] Y. Cao et al., ‘RiskLabs: Predicting Financial Risk Using Large Language Model Based on Multi-Sources Data’, Apr. 11, 2024, arXiv: arXiv:2404.07452. doi: 10.48550/arXiv.2404.07452. [2] A. Kim, M. Muhn, and V. V. Nikolaev, ‘From Transcripts to Insights: Uncovering Corporate Risks Using Generative AI’, Jul. 11, 2024, Rochester, NY: 4593660. doi: 10.2139/ssrn.4593660. [3] T. Li and X. Dai, ‘Financial Risk Prediction and Management using Machine Learning and Natural Language Processing’, ijacsa, vol. 15, no. 6, 2024, doi: 10.14569/IJACSA.2024.0150623. [4] Y. Wang, ‘Generative AI in Operational Risk Management: Harnessing the Future of Finance’, May 17, 2023, Rochester, NY: 4452504. doi: 10.2139/ssrn.4452504. [5] X. Zhu, H. Jin, J. Li, and Y. Wang, ‘Topic-Gpt: A Novel Risk Identification Method Based on Large Language Model’, Jul. 04, 2024, Social Science Research Network, Rochester, NY: 4885365. doi: 10.2139/ssrn.4885365. [6] M. Katamaneni, P. Agrawal, S. Veera, A. K. Sahoo, K. Singh Sidhu, and M. F. Hasan, ‘AI-Based Risk Management in Financial Services’, in 2024 Second International Conference Computational and Characterization Techniques in Engineering & Sciences (IC3TES), Nov. 2024, pp. 1–5. doi: 10.1109/IC3TES62412.2024.10877497. [7] X. V. Li and F. S. Passino, ‘FinDKG: Dynamic Knowledge Graphs with Large Language Models for Detecting Global Trends in Financial Markets’, in Proceedings of the 5th ACM International Conference on AI in Finance, Nov. 2024, pp. 573–581. doi: 10.1145/3677052.3698603. [8] A. Nygaard et al., ‘News Risk Alerting System (NRAS): A Data-Driven LLM Approach to Proactive Credit Risk Monitoring’, in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, F. Dernoncourt, D. Preoţiuc-Pietro, and A. Shimorina, Eds., Miami, Florida, US: Association for Computational Linguistics, Nov. 2024, pp. 429–439. doi: 10.18653/v1/2024.emnlpindustry.32. [9] Z. Xiao, Z. Mai, Z. Xu, Y. Cui, and J. Li, ‘Corporate Event Predictions Using Large Language Models’, in 2023 10th International Conference on Soft Computing & Machine Intelligence (ISCMI), Nov. 2023, pp. 193– 197. doi: 10.1109/ISCMI59957.2023.10458651. [10] Committee of Sponsoring Organizations of the Treadway Commission (COSO), Enterprise Risk Management—Integrating with Strategy and Performance. Durham, NC: COSO, 2017. [11] International Organization for Standardization, ISO 31000:2018 – Risk management — Guidelines. Geneva, Switzerland: ISO, 2018. 60
Graph-Based Feature Engineering for DeFi Security Incident Severity Prediction Daria Pavlova∗ [email protected] Jožef Stefan International Postgraduate School Ljubljana, Slovenia Inna Novalija [email protected] Jožef Stefan Institute Ljubljana, Slovenia Dunja Mladenić [email protected] Jožef Stefan Institute Ljubljana, Slovenia ABSTRACT Decentralized Finance (DeFi) has emerged as a rapidly growing sector, but it has been plagued by numerous security incidents resulting in billions of USD in losses. An important challenge is predicting which security incidents will lead to severe financial losses, as this can inform risk management and mitigation strategies. In this paper, we present a novel approach that integrates a semantic knowledge graph of the DeFi ecosystem into the machine learning pipeline for incident severity prediction. We construct a knowledge graph capturing rich relationships between DeFi protocols (including protocol fork lineage, multi-chain deployments, and historical incidents), and we engineer graph-based features from this graph to augment traditional incident features. Using these features in a gradient boosting trees classifier, we predict whether an incident will cause above-threshold (severe) losses. Our results show that incorporating graph-based features yields a substantial improvement in predictive performance: the model with semantic graph features achieves an Area Under ROC Curve (AUC) of 0.787, a 31.6% relative increase over the baseline model using only non-graph features. We observe particularly large gains in precision (from 0.341 to 0.490), indicating a significantly reduced false alarm rate. While these absolute performance values remain moderate, they represent substantial improvements for this challenging prediction task. The findings demonstrate the practical value of graph-enriched feature engineering for security analytics in DeFi. This work provides new insights into how protocol interconnections and characteristics contribute to incident severity, opening avenues for more robust DeFi risk assessment tools. KEYWORDS Decentralized Finance, DeFi, Security, Knowledge Graph, Feature Engineering, Incident Severity Prediction 1 INTRODUCTION Decentralized Finance (DeFi) platforms have experienced rapid growth, alongside a surge in security breaches such as hacks and exploits. In 2022 alone, crypto attacks led to over $3.8 billion in stolen assets, with the majority coming from DeFi protocol exploits [ 1 ]. These incidents vary widely in impact: while many attacks result in limited losses, a significant fraction escalate into ∗First author and presenter. Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia ©2024 Copyright held by the owner/author(s). https://doi.org/10.70314/is.2025.sikdd.6 catastrophic failures causing losses in the tens or hundreds of millions of dollars. Predicting which security incidents will become severe (high-loss) events is crucial for proactive risk management, insurance underwriting, and developing early warning systems for the DeFi ecosystem. Prior research has analyzed DeFi vulnerabilities and attack taxonomy [ 6 ], and industry reports highlight the growing scale of DeFi hacks. However, there is a gap in predictive approaches: existing studies focus on identifying vulnerabilities or classifying attack types, rather than forecasting the severity level of an incident before it fully unfolds. To our knowledge, this is the first work to apply semantic knowledge graph features specifically for DeFi incident severity prediction, establishing a new baseline for this important problem. In traditional cybersecurity contexts, incorporating relational context via knowledge graphs and network models has been shown to improve threat detection [ 3 ]. For example, graph-based severity triage using attack graphs has been studied in traditional cybersecurity [5]. In this work, we propose a novel graph-based feature engineering approach to address this challenge. We construct a semantic knowledge graph of the DeFi ecosystem that encodes domain knowledge: nodes represent entities such as protocols and incidents, and edges capture relationships like "forked-from" (denoting protocol lineage) and "deployed-on" (connecting protocols to blockchain platforms), among others. From this knowledge graph, we derive a set of graph-based features for each security incident. These features quantify properties such as a protocol’s structural position in the ecosystem (e.g., number of fork "children," cross-chain deployments, past incident count), which we posit are predictive of how severe an incident could be. We integrate these semantic graph features with conventional features (e.g., time of incident, incident type categories) in a machine learning classifier to predict whether an incident’s loss will exceed a severity threshold. The contributions of our work are as follows: • We introduce a methodology to incorporate a DeFi-specific knowledge graph into security incident severity prediction. • We demonstrate significant performance gains over a baseline model lacking graph features (improving AUC by 31.6% and F1-score by 25%). • We provide a comprehensive analysis including case studies, illustrating how related protocol dependencies can influence risk. • We discuss practical implications of our findings for improving DeFi risk assessment. All code and the publicly available dataset for this work are available in an open-source repository [4]. 61
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Pavlova et al. Figure 1: DeFi knowledge graph overview: protocols, blockchains, and incidents with relations (forked-from, deployed-on, involves). 2 METHODOLOGY 2.1 Knowledge Graph Construction We built a knowledge graph representing the DeFi ecosystem to serve as a basis for feature engineering. The construction process was semi-automated, combining API data extraction with manual curation to ensure semantic consistency. Data Sources: We integrated data from three primary sources: (1) the Rekt database (https://rekt.news) containing detailed DeFi security incident reports, (2) DeFiLlama’s API providing protocol metadata including deployment chains and fork relationships, and (3) SlowMist Hacked for additional incident verification. All data sources are publicly available. Semi-Automated Process: Protocol and incident data were automatically extracted using APIs and web scraping. Fork relationships were identified through a combination of automated code similarity analysis (for protocols with public repositories) and manual verification based on project documentation. The resulting knowledge graph contains 892 protocol nodes, 1,608 Figure 2: Convex-centric subgraph. Dependency on Curve highlights potential severity propagation via upstream vulnerabilities. incident nodes, and 42 blockchain nodes, connected by over 3,500 edges representing various relationships. We use Neo4j to store and query this graph efficiently through asynchronous operations. The graph’s schema defines several entity types and relations relevant to DeFi security: • Protocol nodes: Each DeFi protocol (e.g., lending platform, DEX, yield aggregator) is a node. Attributes include protocol name and launch date. • Incident nodes: Major recorded security incidents (hacks, exploits) are represented as nodes with attributes such as date, loss amount, and qualitative classification (e.g., flash loan, smart contract bug). • Blockchain nodes: Blockchain platforms (Ethereum, Binance Smart Chain, etc.) are included to capture deployment contexts. Key relationships are encoded as directed edges: • Fork-of: Connects a protocol to the protocol it was forked from (if applicable), capturing lineage (e.g., SushiSwap → Uniswap). • Deployed-on: Links a protocol to a blockchain platform on which it is deployed. • Incident-involves: Links an incident node to the protocol(s) affected by that incident. The resulting graph captures a rich hierarchical structure of protocol relationships (including parent–child fork trees and cross-chain deployment links), as well as the association of past incidents with protocols. An overview of the graph structure is shown in Figure 1, and an illustrative Convex-centric subgraph is given in Figure 2. 2.2 Feature Engineering with Graph-Based Features From the knowledge graph, we derived several quantitative features that characterize the structural and historical context of the protocol involved in a given incident: 62
Graph-Based DeFi Security Prediction Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia • Protocol multi-chain count: the number of distinct blockchains on which the protocol is deployed (degree of deployed-on edges). A higher count indicates a widely deployed protocol, potentially implying larger user bases or attack surfaces. • Fork lineage indicators: whether the protocol is a fork of another (has parent) and the number of forks derived from it. These capture if a protocol inherits code (and possibly vulnerabilities) from a parent and how prevalent its code is in offspring projects. • Past incident count: the total number of past security incidents involving the protocol (count of incident-involves edges to prior incidents). A history of frequent past incidents might signal underlying security weaknesses or attractive target value. In addition to these graph-derived features, we include conventional features for each incident: • Temporal features: the year and month of the incident, and day-of-week if relevant, to capture any time-related patterns or trends in attack occurrence. • Categorical features: the general type of attack or vulnerability exploited (e.g., reentrancy, price oracle manipulation), and the asset or protocol category targeted, which provide contextual information on the incident. All features are computed or retrieved at the time just before the incident (to avoid using any post-incident information). The combination of graph-based features with traditional features forms the feature vector used for prediction. The end-to-end feature extraction and modeling pipeline is summarized in Figure 3. 2.3 Classification Model and Training We frame incident severity prediction as a binary classification task: severe vs. non-severe loss outcome. Following prior work in financial risk modeling, we define a severe incident as one with loss exceeding a high quantile threshold of the loss distribution. In our dataset, we tested multiple thresholds (70th, 75th, and 80th percentiles), with the 75th percentile ($2.21 million) serving as the primary cutoff, yielding 402 severe incidents out of 1,608. The model showed consistent improvements across all thresholds, confirming the robustness of our approach. Our primary model is a gradient boosting decision trees ensemble (LightGBM [ 2 ]), selected for its efficiency, ability to handle heterogeneous feature types, and proven performance in tabular financial risk modeling. We enabled LightGBM’s built-in class imbalance option ( is_unbalance=True ), as severe cases represent 25% of the data. Train/Test Split: Data were split chronologically into 75% training and 25% testing. Early stopping was not applied due to dataset size; hyperparameters were fixed after preliminary tuning. We compare two feature sets: a Baseline model using only non-graph features (temporal and categorical), and a Semantic Graph model combining these with graph-based features. Performance is evaluated with Area Under the ROC Curve (AUC) and supported by Precision, Recall, and F1-score. Figure 3: Workflow: derive graph-based features from the DeFi knowledge graph and combine with conventional incident features for classification. Figure 4: Performance comparison. Bar chart for AUC, F1, precision, recall. 3 EXPERIMENTS AND RESULTS 3.1 Dataset and Experimental Setup We compiled a publicly available dataset of 1,608 DeFi security incidents that occurred between 2020 and 2025. The dataset was constructed by combining data from: (1) Rekt database providing comprehensive incident reports with loss amounts and attack descriptions, (2) DeFiLlama API for protocol metadata including TVL and deployment information, and (3) SlowMist Hacked for additional incident verification and technical details. Each incident record includes the loss amount (in USD) and details such as date and attack type. Incidents with losses above $2.21 million were labeled as severe, which yields a severe class prevalence of roughly 25% (402 severe vs. 1,206 non-severe cases). For training and evaluation, we use a chronological split with 75% for training and 25% for testing; early stopping was not applied. 63
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Abdul Sittar, Mateja Smiljanić, and Alenka Guček 2 Related Work LLMs are increasingly employed to model human behaviour in online settings, but current evaluation approaches such as simplified Turing tests involving human annotators fail to capture the subtle stylistic and emotional nuances that differentiate human generated text from AI-generated text [12]. It proposes a human likeness evaluation framework that systematically measures how closely LLM generated social responses resemble those of real users. This framework utilizes a set of interpretable textual features that capture stylistic, tonal, and emotional aspects of online conversations. While they can mimic certain human behaviours and decision making processes, primarily due to their training data, it remains largely unexplored whether repeated interactions with other agents amplify their biases or lead to exclusive patterns of behaviour over time [8]. Modelling social media has ben an active research area for understanding use behaviour, information diffusion, and network effects. Agent-based models have been widely used to replicate interactions among users, simulate posting and replying behaviour, and study emergent phenomena such as viral content spread, echo chambers, and filter bubbles [6, 11]. These models often rely on simplified rules or probabilistic mechanisms to determine agent actions. Our work extends this by using fine-tuned language model to generate realistic post and reply content, capturing both semantic and temporal patterns observed in real social media interactions. The concept of filter bubbles has been extensively studied in the context of social media algorithms and personalized content delivery [17, 7, 3]. Prior studies have shown that temporal factors, such as posting frequency and timing, significantly influence the formation of echo chambers and the propagation of sentiment. Unlike traditional simulations, our approach explicitly models time windows and agent-specific schedules, allowing the study of how environmental changes affect network dynamics and user behaviour over time. Large language models (LLMs) have been increasingly applied to social media analysis, content generation, and user simulation. Fine-tuned models can capture domain-specific language, hashtags, and posting patterns, enabling more realistic simulations of user behaviour [13, 4]. Existing work has largely focused on generating content for individual posts or replies; in contrast, our approach integrates posting, replying, and environment management in a unified simulation, enabling multi-agent interaction analysis. Recent studies have used sentiment and emotion analysis to evaluate social media content, including the study of affective trends and collective mood in online networks [16, 5]. Our approach leverages these techniques to compare simulated emotion trends with real-world Twitter data, providing a quantitative measure to validate the fidelity of the agent-based simulation. 3 Methodology Our methodology employs a two stage approach combining probabilistic scheduling with domain-specialized fine-tuned language model agents to simulate realistic social media interactions (posting and replying). The approach consists of two primary components: (1) Timeline based probabilistic model that serves as an timeline manager, and (2) Domain-specialized fine-tuned agents that generate contextually appropriate content based on the timeline manager’s decisions. Figure 1: Overview of the proposed methodology for conversation simulation. The timeline manager determines which agent should act next based on the current time, agent, context, and action. The selected fine-tuned model then generates a new post or reply for the chosen agent, creating realistic conversation flow. 3.1 Probabilistic model The probabilistic scheduler is implemented as a multi-output neural network that simultaneously predicts four key dimensions of social media behaviour: agent selection (which agent should act next), action classification (post vs. reply), temporal prediction (timing of next action), and context setting (emotional tone and topical focus for content generation). The model is trained on 88,330 conversation items spanning April 2019 to April 2020, focusing on AI and cryptocurrency discussions. Our Timeline-Based approach generates 93,440 chronological training pairs—18.7×more than baseline methods—through complete conversation sequence learning rather than isolated post-reply pairs. Given the current state 𝑆(𝑡) at time 𝑡 , the model computes probability distributions over the action space. 3.2 Fine-tuned model We implement a single fine-tuned language model that serves as both AI and cryptocurrency agents. The model is trained on conversations from both domains (AI technology and cryptocurrency discussions) to capture the vocabulary, argumentation patterns, and discourse styles across both topic areas. • Agent A (AI Focus): The same fine-tuned model called when the probabilistic scheduler determines AI-related content is needed. • Agent B (Crypto Focus): The identical fine-tuned model called when cryptocurrency-related content generation is required. When called by the probabilistic scheduler, the fine-tuned model generates content based on provided context including action type (post/reply), emotional context, topical focus, temporal context, and conversation history. The model’s training on both domains enables it to produce contextually appropriate responses regardless of which agent role it is fulfilling. 70
Designing AI Agents for Social Media Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia 3.3 Integration and Coordination The probabilistic scheduler communicates with fine-tuned agents through a structured interface that maintains separation between temporal decisions (when and who acts) and content decisions (what is said). At each simulation step, the scheduler: (1) analyses current conversation state, (2) predicts next action parameters, (3) selects appropriate domain agent, (4) provides structured context to the selected agent, and (5) integrates generated content into the conversation thread. This approach enables realistic conversations where different domain experts can contribute to mixed topic discussions while maintaining their specialized perspectives and temporal behavioural patterns observed in real social media data. 4 Experimental Setup In this section, we describe the features, model and evaluation metrics. 4.1 Timeline Manager The baseline system is a timeline based probabilistic model that learns agent transitions, reply probabilities, and temporal distributions from training data. Predictions are made deterministically by selecting the most probable outcome, with probability estimates derived directly from observed frequencies. The enhanced approach employs a machine learning ensemble with separate classifiers for agent, action, and time prediction. Features include agent history, action history, and time of day. Predictions are generated using temperature-controlled stochastic sampling, with an ensemble across multiple temperature settings for robustness. This design enables greater flexibility and diversity, counteracting the strong biases inherent in the probabilistic model. 4.1.1 Evaluation Metrics. Table 1 summarizes the key differences between the original probabilistic model and the improved ML-based model, covering both quantitative performance and qualitative conversational outcomes. Aspect Probabilistic Model ML-Based Model Agent Prediction 44.8% accuracy, but always predicts Crypto_Agent (100%) 55.2% accuracy, balanced AI_Agent (50%) and Crypto_Agent (50%) Action Prediction 74.4% accuracy by predicting only “post” (0% replies) 67.8% accuracy with realistic mix: 65% posts / 35% replies (close to ground truth 73/27) Temporal Modelling MAE = 5.41 min; 99.4% within ±15 min MAE = 7.11 min; 99.2% within ±15 min Table 1: Comparison of the Original Probabilistic Model vs. the Improved ML-Based Model. we evaluated our probabilistic model using comprehensive metrics across three key categories: • Agent Prediction: 61.3% accuracy (22.6% improvement over random chance) •Action Classification: 96.8% accuracy for post vs. reply prediction • Temporal Modelling: 50.7-minute MAE with 99.15% accuracy within ±15 minutes Our evaluation demonstrates that the probabilistic scheduler successfully replicates conversation structure: • Agent Alternation: 94.2% similarity to real switching behaviours • Temporal Rhythms: Strong correlation (r=0.78) with actual daily patterns • Action Distribution: Maintains realistic post/reply ratios (94.5%/5.5%) 4.2 Fine-tuned model Table 2: Evaluation Results: ROUGE and Semantic Similarity Metric Score ROUGE-1 0.1373 ROUGE-2 0.0519 ROUGE-L 0.1179 ROUGE-Lsum 0.1217 Semantic Similarity (SBERT) 0.4041 Table 2 reports the evaluation results for the fine-tuned model’s generated content. ROUGE metrics (ROUGE-1, ROUGE-2, ROUGEL, and ROUGE-Lsum) measure lexical overlap between generated outputs and the reference Twitter posts. The relatively low scores (e.g., ROUGE-1 = 0.1373) indicate that while the generated text captures some overlapping words or phrases, it often diverges lexically from the original references. This is expected since the model is not designed for verbatim reproduction but rather for generating semantically coherent alternatives. To complement ROUGE, we compute semantic similarity using SBERT embeddings. The score of 0.4041 shows that, on average, the generated outputs are moderately aligned in meaning with the reference texts, even when surface-level wording differs. This highlights that the fine-tuned model is able to remain contextually and thematically relevant while producing novel expressions. Overall, the combination of ROUGE and semantic similarity suggests that the fine-tuned agents generate content that does not simply replicate reference posts but instead produces new, semantically consistent outputs. Figure 2: Methodology diagram showing both experimental approaches: First step, second step, third step, fourth step Figure 2 presents the aggregated emotion comparison between the reference Twitter dataset and the conversations generated by the fine-tuned model. The analysis is based on average emotion scores across multiple conversation samples, with categories including hate, not_hate, non_offensive, irony, neutral, positive, and negative. Blue bars represent the reference data, while orange bars indicate the generated outputs. Overall, the comparison shows strong alignment between the two distributions for key non-toxic categories. Both reference and generated conversations are overwhelmingly classified as 71
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Abdul Sittar, Mateja Smiljanić, and Alenka Guček not_hate and non_offensive, with nearly identical scores (approximately 0.95 and 0.75, respectively). Similarly, both datasets contain minimal hate or negative content, indicating that the synthetic conversations do not introduce harmful patterns absent from the real data. At the same time, certain emotional discrepancies are evident. The generated conversations exhibit lower levels of irony and positivity compared to the real dataset. Specifically, irony is notably under-represented in synthetic conversations (0.04 versus 0.12 in the reference data), suggesting that nuanced and implicit language styles are harder for the model to reproduce. Similarly, positive sentiment is reduced in generated text (0.49 versus 0.62), while neutrality is slightly higher (0.78 versus 0.71). This indicates a tendency of the model to produce emotionally flatter and less expressive outputs. Taken together, the results suggest that the model successfully replicates the broad emotional structure of conversations, particularly in terms of avoiding toxic or offensive content. However, the generated outputs are less emotionally rich than real data, with reduced representation of irony and positivity. This highlights a key limitation of current LLM-based conversation agents: while structurally sound, they may generate interactions that are less engaging or authentic in their emotional dynamics. 5 Conclusions In this work, we presented a novel approach for replicating social media user behaviour using fine-tuned language models organized as autonomous agents. By combining a timeline manager (Model A) with specialized posting (Model B) and replying (Model C) models, we simulated realistic multi-agent interactions across AI and Crypto related topics. Our timeline based probabilistic model successfully replicates structural conversation patterns with 61.3% agent accuracy and near-perfect action classification (96.8%), establishing a new benchmark while providing clear paths for further enhancement through domain specialization. Our experiments demonstrated that the approach can generate temporal posting and replying patterns that closely resemble real-world Twitter data. We showed that modifying the environment model significantly influences agent behaviour, posting frequency, and network dynamics, supporting our hypothesis that environmental and temporal factors shape interaction patterns in social networks. This approach provides a flexible and controlled platform for studying filter bubble formation, emotion propagation, and emergent social dynamics. Future work can extend the approach to more complex network structures, additional domains, and the integration of user-specific behaviour models to further explore interventions for mitigating echo chambers and enhancing diversity in online interactions. 6 Acknowledgment The research presented in this paper was funded by the EU’s Horizon Europe Framework under grant agreement number 101095095 (TWON) and 101094905 (AI4Gov). References [1] Jordan J. Bird et al. 2021. Chatbot interaction with artificial intelligence: human data augmentation with t5 and language transformer ensemble for text classification. arXiv preprint arXiv:2010.05990. [2] Uthsav Chitra and Christopher Musco. 2020. Analyzing the impact of filter bubbles on social network polarization. In Proceedings of the 13th international conference on web search and data mining, 115–123. [3] Uthsav Chitra and Christopher Musco. 2019. Understanding filter bubbles and polarization in social networks. arXiv preprint arXiv:1906.08772. [4] Cristina Chueca Del Cerro. 2024. The power of social networks and social media’s filter bubble in shaping polarisation: an agent-based model. Applied Network Science, 9, 1, 69. [5] Matteo Cinelli, Gianmarco De Francisci Morales, Alessandro Galeazzi, Walter Quattrociocchi, and Michele Starnini. 2020. Echo chambers on social media: a comparative analysis. arXiv preprint arXiv:2004.09603. [6] Rui Fan, Ke Xu, and Jichang Zhao. 2018. An agent-based model for emotion contagion and competition in online social media. Physica a: statistical mechanics and its applications, 495, 245–259. [7] Antonino Ferraro, Antonio Galli, Valerio La Gatta, Marco Postiglione, Gian Marco Orlando, Diego Russo, Giuseppe Riccio, Antonio Romano, and Vincenzo Moscato. 2024. Agent-based modelling meets generative ai in social network simulations. In International Conference on Advances in Social Networks Analysis and Mining. Springer, 155–170. [8] Farnoosh Hashemi and Michael Macy. 2025. Collective social behaviors in llms: an analysis of llms social networks. In Large Language Models for Scientific and Societal Advances. [9] Tianrui Hu, Dimitrios Liakopoulos, Xiwen Wei, Radu Marculescu, and Neeraja J Yadwadkar. 2025. Simulating rumor spreading in social networks using llm agents. arXiv preprint arXiv:2502.01450. [10] Z. Li, J. Zhu, et al. 2023. Synthetic data generation with large language models for text classification: potential and limitations. arXiv preprint arXiv:2310.07849. [11] Hamid Reza Nasrinpour, Marcia R Friesen, et al. 2016. An agent-based model of message propagation in the facebook electronic social network. arXiv preprint arXiv:1611.07454. [12] Nicolò Pagan, Petter Törnberg, Christopher Bail, Ancsa Hannak, and Christopher Barrie. [n. d.] Can llms imitate social media dialogue? techniques for calibration and bert-based turing-test. In First Workshop on Social Simulation with LLMs. [13] Kayhan Parsi and Nanette Elster. 2015. Why can’t we be friends? a casebased analysis of ethical issues with social media in health care. AMA journal of ethics, 17, 11, 1009–1018. [14] Ifrah Pervaz, Iqra Ameer, Abdul Sittar, and Rao Muhammad Adeel Nawab. 2015. Identification of author personality traits using stylistic features: notebook for pan at clef 2015. In CLEF (Working Notes), 1–7. [15] E. Rosenfeld et al. 2025. Evaluating synthetic data generation from user generated text. Computational Linguistics, 51, 1, 191–230. [16] Tanase Tasente. 2025. Understanding the dynamics of filter bubbles in social media communication: a literature review. Vivat Academia, 1–21. [17] Petter Törnberg, Diliara Valeeva, Justus Uitermark, and Christopher Bail. 2023. Simulating social media using large language models to evaluate alternative news feed algorithms. arXiv preprint arXiv:2310.05984. [18] Kang Min Yoo et al. 2021. Gpt3mix: leveraging large-scale language models for text augmentation. In Findings of the Association for Computational Linguistics: EMNLP 2021, 2225–2239. 72
Explaining Temporal Data in Manufacturing using LLMs and Markov Chains Jan Šturm [email protected] Jožef Stefan Institute Jožef Stefan International Postgraduate School Ljubljana, Slovenia Maja Škrjanc [email protected] Jožef Stefan Institute Jožef Stefan International Postgraduate School Ljubljana, Slovenia Oleksandra Topal [email protected] Jožef Stefan Institute Ljubljana, Slovenia Inna Novalija [email protected] Jožef Stefan Institute Ljubljana, Slovenia Dunja Mladenić [email protected] Jožef Stefan Institute Jožef Stefan International Postgraduate School Ljubljana, Slovenia Marko Grobelnik [email protected] Jožef Stefan Institute Ljubljana, Slovenia Abstract Monitoring and understanding complex industrial processes from high-dimensional IoT sensor data remains a significant challenge. While advanced modeling techniques like Hierarchical Markov Chains can abstract raw data, their outputs are often difficult for domain experts to interpret, creating a gap between data-driven insights and operational management. Existing explainability methods often focus on feature importance rather than providing holistic, semantic descriptions of system states. This paper introduces a framework that bridges this gap by transforming the abstract states of a process model into intuitive, human-readable concepts. The methodology leverages the StreamStory (Hierarchical Markov Chain) tool approach to generate behavioral profiles based on log-likelihood calculations within sliding temporal windows. StreamStory states are summarized using an LLM to assign semantic labels and descriptions. This approach reduces the initial reliance on domain experts for analysis, aids the understanding of complex system dynamics, and provides a transparent foundation for identifying both normal and anomalous operational patterns. The result is a more interpretable representation of industrial processes, facilitating improved predictive maintenance and operational efficiency. Keywords Multivariate Timeseries, Explainable AI, LLMs, Markov Chains 1 Introduction The widespread adoption of Internet of Things (IoT) sensors in industrial environments has generated vast streams of multivariate time-series data. While this data holds immense potential for process optimization and predictive maintenance, its complexity often surpasses human cognitive capacity. Tools like StreamStory [6] have emerged to model these complex systems using Hierarchical Markov Chains, abstracting raw data into a more manageable set of states and transitions. However, a fundamental Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia ©2025 Copyright held by the owner/author(s). https://doi.org/10.70314/is.2025.sikdd.28 challenge persists: a disconnect between the model’s statistical outputs and the experiential knowledge of domain experts. The motivation for this work stems from this challenge. Domain experts, who possess invaluable implicit knowledge of a system, often struggle to interpret the statistical outputs of process models. Conversely, data scientists may identify patterns that lack the necessary operational context for effective action. Presenting experts with a graphical representation of states and transitions is a step forward, but it does not fully bridge the semantic gap. They may not understand what a specific state represents in the physical world or why a particular transition is significant. This leads to a bottleneck where valuable data-driven insights are not fully utilized, hindering efforts to improve system management and efficiency. To address this, the paper proposes a methodology that enhances the interpretability of hierarchical process models. This approach creates a new layer of understanding that is accessible to operational personnel without requiring deep data science expertise. By translating abstract model states into meaningful, semantically rich descriptions, it provides a tool that allows the system’s behavior to be understood, validated, and ultimately, better managed. This work introduces a methodology to automatically generate these descriptions, moving from complex data to clear, actionable insights. This work presents two primary contributions for industrial applications: a method for LLM-based labeling of Markov chain states, and a methodology for identifying events as anomalous or normal. 2 Related Work The field of time-series anomaly detection has evolved from interpretable statistical models like ARIMA and classical machine learning such as Isolation Forest to high-performance deep learning architectures including LSTMs, Transformers, and Autoencoders [5, 4, 7]. While these advanced models excel at pattern recognition, their complexity necessitates post-hoc XAI tools like LIME and SHAP to explain their decisions, which are limited to providing low-level feature attributions [1]. Recent work also demonstrates the utility of Hidden Markov Models (HMMs) for anomaly detection, for instance, by designing active search strategies to locate an evolving anomaly among multiple processes [2], or by learning normal temporal dynamics from remote sensing data to detect, localize, and classify croprelated deviations [3]. However, while effective for detection, the 73
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Šturm et al. abstract nature of HMM states can be difficult for domain experts to interpret. The present work addresses this by transforming the state sequence into a multi-scale behavioral profile, which enables a Large Language Model (LLM) to generate rich, semantic explanations of system behavior. This approach innovates by first classifying each multivariate data point into a state within a pre-built Markov Chain model and then calculating log-likelihoods from the state sequence to form a multi-scale representation. Crucially, this representation allows for the recognition of regular system behavior and various anomalies. By analyzing the statistical distribution of these profiles—identifying dense regions of regular behavior and sparse outliers corresponding to anomalous states—an LLM can then assign rich, human-readable descriptions, connecting abstract data to operational knowledge. 3 Methodology The framework is designed to post-process models generated by the StreamStory system. Figure 1 outlines this multi-stage process, which begins with the statistical features from the Markov model and culminates in semantically enriched explanations of system behavior. The core of this methodology is the transformation of abstract machine states into meaningful concepts using a combination of statistical feature engineering and LLM interpretation. The process focuses on creating robust representations of system behavior and leveraging an LLM to translate these representations into human-understandable language. Figure 1: Proposed methodology for identifying and explaining normal and anomalous operational profiles. 3.1 Log-Likelihood Score Calculation The input to the pipeline is a pre-existing Hierarchical Markov Chain model of an industrial process, which includes a history of state transitions over time. The first step is to create a rich feature representation that captures the system’s dynamics. A sliding window (Figure 2) approach moves across the sequence of historical state transitions. For each window of a given size, a single feature is calculated: the log-likelihood of that specific sequence of transitions occurring. This score is calculated by summing the log-transformed transition probabilities for each step in the sequence, as defined by the underlying Markov model. The score effectively quantifies how "normal" or "expected" a particular sequence of behavior is according to the learned model. Highly probable sequences yield higher log-likelihood scores (closer to zero), while rare sequences result in large negative scores. Figure 2: An illustration of the sliding window method. Three windows of different sizes, highlighted in yellow (largest), brown (medium), and green (smallest), are applied to a sequence of system states. A log-likelihood score is then calculated for the sub-sequence contained within each colored window. 3.2 Behavior Profile Construction To capture dynamics over multiple time scales, several sliding windows of different sizes are used simultaneously. The loglikelihood score calculated from each window is concatenated to form a single feature vector for each time step. This multi-scale vector, termed a behavior profile, serves as a rich representation of the system’s dynamics at that moment, encapsulating both shortterm and longer-term patterns. This profile is a crucial output, as it provides a quantitative basis for distinguishing between different modes of operation. 3.3 Ranking System Behavior via Anomaly Scoring Following the construction of the behavior profiles, their distribution is analyzed to identify distinct operational patterns. An unsupervised density-based approach is employed to score each profile’s typicality. The Isolation Forest algorithm is used for this purpose because it does not assume a specific data distribution and excels at identifying outliers in a high-dimensional space. Profiles that are common and lie in dense regions of the feature space receive a high score, corresponding to normal behavior. Conversely, profiles that are rare and isolated receive a low score, flagging them as anomalous. This produces a continuous spectrum of normalcy, allowing for a ranked analysis of all operational events. 3.4 LLM-Powered State Naming and Interpretation To translate abstract states into meaningful concepts, an LLM is utilized. For each granular state discovered by the StreamStory model, its statistical profile (e.g., sensor value distributions) and context about the machine type were formatted into a descriptive prompt. The LLM was then tasked with generating a concise, intuitive name for each state (e.g., "Peak Production - High Flow and Heat"). This process, conducted once per model, creates a semantic layer that is then used to interpret the sequences 74
Explaining Temporal Data with LLMs Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia associated with the highest-ranked normal and lowest-ranked anomalous events. This approach offers two key advantages. First, the LLMgenerated names provide a layer of transparency, offering an immediate hypothesis about what each abstract state represents. Second, it shifts the role of the domain expert from the arduous task of initial interpretation to the more efficient step of validating or refining the LLM-generated labels, accelerating the process of gaining actionable insights. 4 Experiment To validate the proposed framework, an experiment was conducted using a real-world industrial dataset from an oil refinery pump. This section details the dataset, implementation, and results. 4.1 Dataset The experiment was performed on a proprietary, real-world dataset obtained from an industrial oil refinery. Due to its confidential nature, the dataset is not publicly available. The data consists of a multivariate time-series collected over one month of operation (March-April 2017) with a 15-minute sampling resolution. Data was gathered from a suite of IoT sensors monitoring the core functions of a critical pump. Key measurements include fluid flow rate (Kg/h), suction and discharge pressure (Kg/cm2), and temperatures of the process fluid and mechanical components (°C). 4.2 Implementation Details The methodology was implemented in a Python environment. The underlying Markov Chain model was built using the entire historical dataset provided, as the goal is to interpret the complete, learned dynamics of the process rather than to perform a predictive task that would require a train/test split. Behavior profiles were constructed using sliding windows of multiple sizes (3, 5, 7, and 10 steps). The resulting profiles were analyzed using the Scikit-learn implementation of Isolation Forest. The ‘contamination‘ parameter was set to 5% for the primary analysis, a common heuristic for industrial processes. State descriptions were generated using the GPT-4o model, which was prompted with the statistical profiles of each state to generate intuitive names. 4.3 Experimental Results and Discussion The application of the framework yielded a ranked list of operational events, characterized by the Isolation Forest decision score. This score serves as a robust indicator of how typical or anomalous a given time window is. Table 1 details the top five most anomalous events identified. These events are characterized by scores that are more than 3 standard deviations below the mean, signifying extreme statistical rarity. The true explanatory power of the method is revealed when the abstract state sequences are translated into their LLM-generated names. For instance, the most anomalous event culminates in a sequence of “... -> ‘Startup or Shutdown Transition‘ -> ‘Machine Idle or Shutdown‘ -> ‘Startup or Shutdown Transition‘.” This provides a clear, human-readable narrative of the pump entering a period of instability and stoppage. This is a marked improvement over black-box models that simply flag a time point as anomalous without providing a temporal context for the "why." An engineer, seeing this semantic sequence, can immediately infer a potential cause for investigation, such as an attempted restart or a stuttering shutdown process. Conversely, the most normal events, detailed in Table 2, paint a picture of operational stability. These events are characterized by positive scores. The LLM-generated names for these sequences, such as transitions between ‘Weekday Peak Performance‘, ‘Weekend Peak-Load Production‘, describe the system operating within its expected high-performance period. This demonstrates the framework’s ability not only to flag deviations but also to recognize and semantically label the system’s healthy, predictable operational cycles, providing a valuable baseline for what constitutes ’good’ performance. Table 1: Top 5 Most Anomalous Events Rank Timestamp Score (Std.) Final State (LLM Name) 1 2017-04-03 14:30 -0.096 (-3.88) Startup...Transition 2 2017-03-28 10:00 -0.071 (-3.45) Startup...Transition 3 2017-03-30 00:00 -0.066 (-3.35) High-Flow, Cool Op. 4 2017-04-03 12:30 -0.061 (-3.26) Machine Idle 5 2017-04-03 15:00 -0.056 (-3.18) Weekday Low-Flow... Conversely, Table 2 presents the five most normal events, which have high positive scores. Their sequences reveal a stable operational loop between states like “Peak Production,” “Weekend Peak-Load Production,” and “Extreme Temperature Peak Performance.” This recurring pattern defines the pump’s healthy operational "heartbeat," providing a data-driven "golden standard" for normal behavior under demanding conditions. This semantic understanding is crucial for operators, as it validates that the system is performing as expected. Table 2: Top 5 Most Normal Events Rank Timestamp Score (Std.) Final State (LLM Name) 1 2017-03-23 22:00 0.192 (1.22) Weekend Peak-Load 2 2017-03-31 06:00 0.192 (1.22) Peak Production 3 2017-04-01 00:00 0.191 (1.20) Peak Production 4 2017-03-31 23:30 0.191 (1.19) Weekday Peak Perf. 5 2017-03-31 07:30 0.190 (1.17) Weekday Peak Perf. To ensure the robustness of the findings, a sensitivity analysis was conducted on the Isolation Forest ‘contamination‘ parameter, testing values of 1%, 5%, and 10%. While the number of points labeled ’Anomalous’ changed as expected, the relative ranking of the most extreme events remained highly consistent, confirming that the core findings are not sensitive to this hyperparameter. The claims in this paper are demonstrated on a single, representative dataset. While the framework is designed to be general, further studies on diverse industrial processes are required to fully validate its broader applicability. The LLM-generated labels were not validated in a formal user study with domain experts; such a study is a valuable next step. 5 Conclusion This paper presented a complete, self-contained framework for increasing the interpretability of complex industrial process models. By creating behavior profiles of system states and using an 75
Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Šturm et al. LLM to assign semantic names, the approach successfully translates abstract data analysis into practical domain knowledge. The method provides a robust process for ranking and explaining individual operational events in a transparent manner, as demonstrated on a real-world industrial dataset. This work establishes a strong foundation for a new type of explainability, moving beyond feature importance to provide narrative, context-rich descriptions of system dynamics. The representation of system dynamics as behavior profiles opens a wide array of possibilities for future research. The current work successfully identifies and presents the raw temporal sequences leading to key events. Future work will focus on applying formal pattern mining techniques to automatically discover recurring and significant sequential patterns within these events. Such an analysis could reveal if distinct "families" of anomalous behavior exist, each with its own characteristic temporal signature. This promises a more nuanced description of system operations and provides a stronger foundation for developing targeted predictive maintenance strategies. Finally, to address current limitations, two key areas will be prioritized. First, formal user studies with domain experts will be conducted to validate the utility and accuracy of the LLM-generated explanations, moving beyond the promising initial results. Second, the framework’s generalizability will be tested through broader empirical evaluation across diverse industrial sectors and sensor types to boost its credibility and applicability. 6 Acknowledgments This work was supported by the Slovenian Research Agency and the European Union’s Horizon 2020 project FAME (Grant No. 101092639). References [1] Liat Antwarg, Ronnie Mindlin Miller, Bracha Shapira, and Lior Rokach. 2019. Explaining anomalies detected by autoencoders using shap. arXiv preprint arXiv:1903.02407. [2] Levli Citron, Kobi Cohen, and Qing Zhao. 2025. Searching for a hidden markov anomaly over multiple processes. arXiv preprint arXiv:2506.17108. [3] Kareth M Leon-Lopez, Florian Mouret, Henry Arguello, and Jean-Yves Tourneret. 2021. Anomaly detection and classification in multispectral time series based on hidden markov models. IEEE transactions on geoscience and remote sensing, 60, 1–11. [4] Sebastian Schmidl, Phillip Wenig, and Thorsten Papenbrock. 2022. Anomaly detection in time series: a comprehensive evaluation. Proceedings of the VLDB Endowment, 15, 9, 1779–1797. [5] Charalampos Shimillas, Kleanthis Malialis, Konstantinos Fokianos, and Marios M Polycarpou. 2025. Transformer-based multivariate time series anomaly localization. In 2025 IEEE Symposium on Computational Intelligence on Engineering/Cyber Physical Systems (CIES). IEEE, 1–8. [6] Luka Stopar, Primoz Skraba, Marko Grobelnik, and Dunja Mladenic. 2018. Streamstory: exploring multivariate time series on multiple scales. IEEE transactions on visualization and computer graphics, 25, 4, 1788–1802. [7] Fengling Wang, Yiyue Jiang, Rongjie Zhang, Aimin Wei, Jingming Xie, and Xiongwen Pang. 2025. A survey of deep anomaly detection in multivariate time series: taxonomy, applications, and directions. Sensors (Basel, Switzerland), 25, 1, 190. 76
Active Learning for Power Grid Security Assessment: Reducing Simulation Cost with Informative Sampling Gašper Leskovec Jožef Stefan Institute Slovenia [email protected] Costas Mylonas UBITECH Greece [email protected] Klemen Kenda Jožef Stefan Institute Slovenia [email protected] Abstract Power grid security assessment under the N-1 criterion requires extensive contingency simulations, which are computationally intensive and costly to label. In this work, we explore the use of active learning (AL) to train binary classifiers that can accurately predict the outcome of contingency scenarios using fewer labeled samples. We evaluate several AL strategies, such as entropy, margin, and uncertainty sampling against a random baseline. Our results show that AL methods achieve the same predictive performance with significantly fewer labels, reducing labeling effort and simulator runtime. These findings demonstrate the effectiveness of integrating AL with power system simulators to enable scalable and efficient N-1 security assessment without sacrificing model accuracy. Keywords active learning, smart grids, security assessment, simulation cost reduction 1 Introduction Ensuring secure operation of power systems under the N-1 criterion is a cornerstone of grid reliability. The criterion requires that the system remains within operational limits following the loss of any single component (e.g., line, transformer, or generator). In practice, this involves simulating a large number of contingencies and checking for violations of thermal or voltage constraints. While essential, such simulations are computationally intensive, particularly when performed on high-fidelity grid models, and their interpretation often requires expert judgment. This creates a bottleneck for both real-time applications and large-scale scenario analyses, where scalability and efficiency are important. Classical approaches to N-1 assessment rely on exhaustive AC power flow simulations combined with contingency ranking heuristics such as performance indices (PIs). While useful for screening, these heuristics may mis-rank contingencies or overlook borderline cases due to masking effects [ 3 ]. Moreover, exhaustive analysis does not scale well with system size, making it unsuitable for fast or repeated assessments. To overcome these challenges, researchers have proposed machine learning (ML) and deep learning (DL) approaches that approximate N-1 contingency outcomes directly from operating point features. One of the earliest contributions in this direction Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. SiKDD 2025, Ljubljana, Slovenia ©2025 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-x-xxxx-xxxx-x/YYYY/MM https://doi.org/10.70314/is.2025.sikdd.11 applied convolutional neural networks (CNNs) to contingency datasets, showing that deep models could achieve over 99% accuracy in detecting insecure cases while being more than 200 times faster than traditional power flow calculations [ 1 ]. Building on this, more recent work explored pooling-ensemble multi-graph learning to design scalable contingency screening schemes based on steady-state information, demonstrating improved adaptability for large-scale systems [ 2 ]. These approaches enable fast security screening without solving power flows for every contingency. However, their reliability hinges on the availability of large labeled datasets covering all relevant operating points and contingencies. Such datasets are typically generated by running exhaustive offline N-1 simulations, which is computationally expensive, or require significant expert effort to label secure versus insecure cases. This dependence on costly and large-scale data generation remains a major limitation of existing ML-based frameworks for steady-state security assessment. To reduce labeling costs, AL has recently been explored in other areas of power systems. For example, authors of [ 5 ] used AL to enhance stability assessment and dominant instability mode identification, showing that models could be trained with far fewer labeled samples while maintaining accuracy. Similarly, authors of [ 4 ] demonstrated an AL-enhanced digital twin for day-ahead load forecasting, where the model iteratively refined predictions by querying only the most uncertain cases. These studies confirm the potential of AL to reduce expert effort and simulation cost by strategically selecting informative samples. However, AL has not yet been applied to N-1 steady-state security assessment, where the need to cut down on contingency simulations is especially critical. In this work, we propose a novel framework for AL driven N-1 security assessment. Our contributions are threefold: (1) We design a binary classification model that predicts whether a given contingency is secure or insecure based on steady-state features. (2) We integrate AL strategies (entropy, margin, and uncertainty sampling) with the classifier to selectively query the most informative contingencies for simulation, reducing the number of labels required. (3) We demonstrate through a case study that our approach achieves the same predictive accuracy as fully supervised baselines while reducing simulation cost and labeling effort by up to 40–50%. This work provides the first evidence that AL can be directly leveraged for N-1 security assessment, offering a scalable and label-efficient alternative to exhaustive simulation or purely supervised ML approaches. 2 Methodology We study whether pool-based AL can reduce the number of expensive N-1 simulations (“labels”) while keeping prediction quality for binary secure vs. insecure classification. 77
SiKDD 2025, October 6th, 2025, Ljubljana, Slovenia Gašper Leskovec, Costas Mylonas, and Klemen Kenda Table 1: Dataset and system description (digital twin of the Greek transmission network). Attribute Value Test system 35 buses, 46 lines, 135 generators, 110 static generators, 20 loads Power flow solver AC load flow (Newton–Raphson), via pandapower Contingencies (N–1) Line outages (all lines except idx 45), generator outages (all) Total contingency cases 8 769 Secure / Insecure 51.28% / 48.72% Feature dimensionality 271 features total Feature groups load_: 20, gen_: 135, sgen_: 110 2.1 Data and labels from a digital twin We use a steady-state digital twin of the transmission grid. For each timestamp we solve the base-case AC power flow, then apply the N-1 criterion by removing each line/transformer/generator in turn and re-solving. An operating point is labeled secure if the base case and all contingencies satisfy limits (bus voltages ∈ [ 0 . 90 , 1 . 10 ] p.u., line loading ≤ 100%); otherwise it is insecure. Non-convergent power flows are labeled insecure. The test system is a digital twin derived from the topology of the Greek transmission network (35 buses, 46 lines, 135 generators, 110 static generators, 20 loads). AC load flows are computed with the Newton–Raphson method in pandapower . N-1 contingencies include all line outages (excluding line index 45) and all generator outages. Table 1 summarizes the dataset. 2.2 Time-aware train/validation/test split Samples are sorted by timestamp. The AL training/pool comes from earlier windows, while the test set is the most recent slice and is never used for training or querying. This avoids temporal leakage and mimics deployment where we predict on future data. A small validation split is carved from the training era for early checks. 2.3 Classifier and hyperparameters Our base model is a Random Forest (RF) because it is fast, robust and provides class-probability posteriors needed by uncertainty-based AL. Across runs we vary hyperparameters in realistic ranges: 𝑛estimators ∈ [ 200 , 1500 ] , max_depth ∈ { 18 , 20 , 24 , 25 , 28 , 30 , 35 , 40 ,None} , min_samples_split ∈ { 2 , 4 } , min_samples_leaf ∈ { 1 , 2 , 3 } , class_weight ∈ {balanced,balanced_subsample} . We use seeds { 42 , 1337 } for reproducibility. Classifier dependence. We use Random Forests for probability outputs and fast retraining inside the AL loop. While AL’s relative gains often transfer across probabilistic classifiers, we did not perform a systematic model sweep here. Evaluating logistic regression and gradient-boosted trees under the same AL protocol is left to future work. 2.4 Pool-based AL loop We follow the standard pool-based AL recipe: (1) Start with an initial labeled set of size 𝑖 and an unlabeled pool. (2) Train the RF on the current labeled set; score the pool to obtain class-probability vectors 𝑝(𝑥). (3) Select the next batch of 𝑏 samples using one of the query strategies below. (4) Query the simulator for labels of the chosen batch (expensive step); add them to the labeled set. (5) Repeat for a fixed number of iterations or until the budget is exhausted. We sweep budgets across runs: 𝑖∈ { 100 , . . . , 500 } , 𝑏∈ { 50 , . . . , 200 } , and up to 40 iterations, which lets us trace long learning curves. Query strategies. We compare: (i) Random (baseline); (ii) Least-confident (uncertainty): score 1 −max𝑐𝑝𝑐(𝑥) ; (iii) Margin: negative gap between top-2 probabilities; (iv) Entropy: −Í𝑐𝑝𝑐(𝑥)log 𝑝𝑐(𝑥) . All three uncertainty policies operate on the same RF posteriors and therefore often rank samples similarly. 2.5 Evaluation After each iteration we evaluate on the fixed test set. At each AL round we retrain the RF from scratch on the enlarged labeled set; new labels are added to training only; the pool remains unlabeled. For each strategy we run multiple configurations and both seeds, then align results by total labeled samples and average across runs to obtain strategy-level learning curves. Unless noted otherwise, TTT values in the main figures are computed on these averaged curves. Appendix A.1 (Table 4a) reports per-run TTT (mean ± std), which is larger due to variability across initial sizes 𝑖 , batch sizes 𝑏, and seeds. 2.6 Metrics We report Accuracy and ROC AUC on the test set, plus two labelefficiency metrics: Time-to-Target (TTT), the smallest number of labeled samples needed for the average curve of a strategy to reach a target (e.g., ACC ≥ 0.92 or AUC ≥ 0.98); and AULC (Area Under the Learning Curve), computed by trapezoidal integration of metric vs. total labeled. Because simulator seconds per call are roughly constant, relative cost/time savings are well approximated by label savings derived from TTT. Additional classification metrics. Besides Accuracy and ROC AUC we also track Precision,Recall,F1 and the False Negative Rate (FNR) on the fixed test set at every AL round. Let TP,FP,FN,TN be counts on the test set. We use the standard definitions: Precision =TP/(TP +FP) , Recall =TP/(TP +FN) , F1 = 2 ·Precision·Recall Precision+Recall , FNR =FN/(FN +TP)= 1 −Recall . We report mean ± std across runs/seeds, and we extract TTT-style thresholds for these metrics when relevant. 3 Results Figure 1 and Figure 2 show learning curves (averaged across seeds). Across the budget range, all three uncertainty-based policies (entropy, margin, uncertainty) dominate the random baseline in both Accuracy and ROC AUC; the area under the learning curve (AULC) is consistently higher. Table 3 summarizes KPIs used in the paper. At the most important targets, AL reaches the same performance with far fewer labels: at ACC ≥ 0 . 92, AL needs about 500 labels vs. 1 040 for random ( ∼ 52% fewer); at AUC ≥ 0 . 98, AL needs 580 vs. 960 ( ∼ 40% fewer). Final metrics at the maximum budget are also higher for AL (ACC 0.917±0.005 and AUC 0.983±0.002) than for 78
Active Learning for Power Grid Security Assessment: Reducing Simulation Cost with Informative Sampling SiKDD 2025, October 6th, 2025, Ljubljana, Slovenia Figure 1: Accuracy vs. total labeled samples (mean ± std across runs). (Note: entropy, margin, and uncertainty overlap almost perfectly on this dataset—so the three AL curves/bands lie on top of each other; Random is shown separately for contrast) Figure 2: ROC AUC vs. total labeled samples (mean ± std across runs). (Note: entropy, margin, and uncertainty overlap almost perfectly on this dataset—so the three AL curves/bands lie on top of each other; Random is shown separately for contrast) Table 2: Final test metrics at maximum budget (mean ± std across runs). Strategy Accuracy ROC AUC entropy 0.917 ±0.005 0.983 ±0.002 margin 0.917 ±0.005 0.983 ±0.002 uncertainty 0.917 ±0.006 0.983 ±0.004 random 0.916 ±0.010 0.977 ±0.004 random (ACC 0.916±0.010 and AUC 0.977±0.004). Differences at the easier target ACC ≥ 0 . 90 are small (all reach it by ∼ 100–120 labels), which is expected for a low threshold. On high AUC values. The time-aware split still yields a separable test set for this case study (AUC ≈ 0.98). This likely reflects informative steady-state features and balanced classes, not overfitting to the test era. That said, harder, more imbalanced systems may reduce AUC and amplify AL gains; we treat this as a scope limitation. Precision, Recall, F1 and FNR.. The additional metrics mirror the ACC/AUC trends: entropy, margin, and uncertainty produce higher AULC and reach target quality with fewer labels than random . At targets Precision/Recall/F1 ≥ 0 . 90 and FNR ≤ 0 . 10, the uncertainty-based policies consistently hit the thresholds earlier on the average curves, confirming that the AL gains are not specific to a single metric. Shaded bands (std across runs) show the same ordering stability observed for ACC/AUC. Full KPI values and TTT thresholds for P/R/F1/FNR are provided in Appendix A.2 (Table 4b). Next, we compare label efficiency using Time-to-Target (TTT). Figures 3 and 4 show TTT for accuracy targets 0 . 90 and 0 . 92, while Figures 5 and 6 show TTT for AUC targets 0 . 97 and 0 . 98. At the easy target ACC ≥ 0 . 90 all strategies reach the goal after about 100–120 labels (uncertainty sometimes at 120 due to seed/batch noise). At the more demanding ACC ≥ 0 . 92 target, active-learning policies need about 500 labels, whereas random needs 1 040 (i.e., ∼ 52% fewer labels). For AUC ≥ 0 . 97, AL reaches the target at 275 labels vs. 325 for random ( ∼ 15% fewer), and for AUC ≥ 0 . 98 at 580 vs. 960 ( ∼ 40% fewer). These reductions translate directly into lower simulation time when the average time per labeling call is roughly constant. Figure 3: TTT (Accuracy ≥ 0 . 90): computed on the strategylevel average curve; per-run variability (mean ± std) is reported in Appendix. Figure 4: TTT (Accuracy ≥ 0 . 92): computed on the strategylevel average curve; per-run variability (mean ± std) is reported in Appendix. Overall, uncertainty-based AL strategies consistently beat random at the harder targets (ACC 0.92 and AUC 0.98) while performing similarly at the easier ACC 0.90 threshold; final performance at the maximum budget remains high (ACC 0.917±0.005, 79
Temporal Dynamics and Causal Feature Integration for Predictive Maintenance in Manufacturing Systems: A Causality-Informed Framework Seyed Iman Hosseini [email protected] Jožef Stefan Institute Ljubljana, Slovenia Jožef Stefan International Postgraduate School Ljubljana, Slovenia Klemen Kenda [email protected] Jožef Stefan Institute Ljubljana, Slovenia Qlector Ljubljana, Slovenia Dunja Mladenič [email protected] Jožef Stefan Institute Ljubljana, Slovenia Jožef Stefan International Postgraduate School Ljubljana, Slovenia ABSTRACT Predictive maintenance is increasingly central to manufacturing, where the goals are to reduce unplanned downtime and extend asset lifetimes. Conventional models often rely on correlations that insufficiently capture temporal dynamics and causal dependencies underlying failures. This study proposes a causality-informed feature-engineering pipeline that combines cross-correlationderived lags with VARLiNGAM to construct lag-aware features from multivariate sensor streams, and evaluates it against standard time-series models using a time-aware split. Three machinelearning models—Random Forest, XGBoost, and Gradient Boosting—were trained and assessed by F1-score (rather than accuracy) on a single-machine subset of the Microsoft Azure Predictive Maintenance dataset (8,708 samples; 26 failures, ≈ 0.3% prevalence). XGBoost trained on raw temporal features achieved F1 ≈ 0 . 94 for longer prediction horizons ( ≥ 10 h) under timeseries–aware cross-validation, with performance declining at shorter horizons as temporal context diminishes. In this setting, causality-informed features did not improve results over the rawfeature baseline. These findings indicate that, with data from a single machine, causal discovery is susceptible to overfitting and may suppress informative temporal patterns; broader, multimachine datasets are likely required for causality-enhanced representations to yield consistent gains. KEYWORDS Predictive Maintenance, Causality, Time-Series Analysis, Machine Learning, VARLiNGAM, Manufacturing Systems 1 INTRODUCTION The rising complexity and interconnectivity of industrial systems have accelerated the need for intelligent maintenance strategies that move beyond reactive and preventive paradigms. Predictive maintenance, driven by sensor data and machine learning, has emerged as a transformative approach to minimize unplanned downtime and optimize asset life cycles [1]. Traditional predictive maintenance models, however, often rely on statistical correlations that fail to capture the directionality and temporal dynamics inherent in real-world system failures [6]. Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society, 2025, Ljubljana, Slovenia ©2025 Copyright held by the owner/author(s). https://doi.org/https://doi.org/10.70314/is.2025.sikdd.12 To address these limitations, this study proposes a causalityinformed framework for predictive maintenance that leverages temporal causal discovery techniques, such as Vector Autoregressive LiNGAM (VARLiNGAM), to engineer predictive features from multivariate sensor data. Our approach integrates crosscorrelation analysis and lag-optimized causal graphs to detect failure precursors and identify their optimal predictive windows. We hypothesize that the observed lack of competitive advantage for causality-informed models, especially when applied to data from a single machine, arises from the limited operational diversity and failure variability. This limitation may cause models to overfit to machine-specific correlations and exclude informative temporal features, thereby hindering their generalizability. Testing this hypothesis through multi-machine datasets will be a key focus of future work. 2 RELATED WORK Causality in time series analysis has become increasingly critical in predictive maintenance, particularly within industrial and manufacturing domains, where early failure detection plays a pivotal role in minimizing operational disruptions and financial losses [5]. Classical statistical models have been widely used to infer causal relationships between sensor measurements and machine states, yet they often fail to capture complex temporal dynamics and the nonlinear relationships inherent in real-world system failures. Recent studies have explored advanced causal inference techniques to enhance fault prediction. Wang S. et al. proposed a framework for fault diagnosis that integrates spatiotemporal dependencies, demonstrating improved predictive accuracy in chemical manufacturing systems [9]. While their work advances reliability in industrial diagnostics, it lacks the flexibility to generalize across diverse application domains. On the other hand, Cui et al. introduced a deep learning framework that enhances predictive maintenance by integrating causal reasoning and longsequence multivariate time-series data, significantly improving predictive performance and interpretability [3]. Despite this, the challenge of automating temporal feature engineering and seamlessly deploying models across different domains remains. Yang X. et al. contributed to the growing literature on datadriven causal analysis by incorporating dynamic latent variables and probabilistic graphical models into causal modeling frameworks [10]. However, these models have yet to fully address the temporal feature extraction required for scalable deployment in real-world predictive maintenance applications. Furthermore, more recent work by Wang Q. et al. introduced a Causal Graph 86
Information Society, 2025, Ljubljana, Slovenia Seyed Iman Hosseini, Klemen Kenda, and Dunja Mladenič Convolution Module that adapts causal discovery within timeseries prediction [8], but their approach is still dependent on complex model adjustments across domains. In this study, we propose a novel framework that integrates lagged correlation with causal analysis techniques to detect failure precursors and quantify their lead times. This framework automates temporal feature engineering and is designed for diverse real-world applications across manufacturing settings, without requiring extensive architectural modifications. The automation of temporal feature engineering and its seamless deployment across comparable manufacturing environments remains a significant challenge, and extending generalization beyond this domain is left for future work. 3 EXPERIMENT Our experimental methodology followed a sequential four-stage process to construct and validate a robust failure prediction model, as shown in Figure 1. The first stage involved performing a cross-correlation analysis between each sensor’s time-series data and the target failure events to determine the optimal predictive time lag, which guided the subsequent steps. In the second stage, the identified optimal lag was used to parameterize a Vector Autoregressive LiNGAM (VARLiNGAM) model, which generated a directed acyclic graph (DAG) representing the causal relationships and effect strengths between sensor variables and the failure event. The third stage focused on creating a causality-informed feature vector by integrating standard statistical metrics from rolling time windows along with advanced features informed by the causal analysis, using the correlation strengths and causal effect strengths derived from the VARLiNGAM model to select and weight features based on their respective optimal and causal lags. Finally, in the fourth stage, the enriched feature set was fed into a machine learning pipeline, employing a time-based data split to prevent look-ahead bias, and training several classification models, including Random Forest, XGBoost, and Gradient Boosting, to assess the effectiveness of the causality-informed approach for predictive maintenance. This integrated approach enhances the predictive capabilities of machine learning models, offering a robust solution for failure prediction in industrial settings. Figure 1: proposed framework Figure 2: Cross correlation analysis 3.1 Dataset and Preprocessing We used the Microsoft Azure Predictive Maintenance Dataset [2], which provides hourly telemetry (voltage, rotation, pressure, vibration) plus maintenance records, failure events, incident reports, and machine metadata for 100 machines over 12 months in 2015 (over 800k hourly summaries and thousands of non-failure error entries). For this study, we restricted the analysis to machine ID 98; after cleaning and merging the sources, we constructed a causality-informed feature vector and standardized features across modalities. Cross-correlation suggested predictive lags of 1–24 hours, so we derived lagged/statistical features from six primary variables (voltage, rotation, pressure, vibration, age, error type). The final dataset comprised 8,708 samples with 26 failures ( ≈ 0.3% ), indicating strong class imbalance [7, 2]. The feature set comprised 150 causality-informed features and 36 features without causal information. 3.2 Cross-correlation Analysis Cross-correlation analysis examines the correlation between two time series as a function of the time lag applied to one of them [11][12]. Unlike simple correlation, which measures linear relationships at a single point in time, cross-correlation reveals how variables relate across different time delays, making it particularly valuable for identifying lead-lag relationships and temporal dependencies. The initial phase of our experimental framework involved a cross-correlation analysis to empirically determine the predictive temporal relationships between sensor signals and equipment failures. For each sensor, we computed the Pearson correlation coefficient between its time series and the binary failure time series across a range of discrete time lags. This procedure was executed by systematically shifting the failure signal backward in time, which allowed for the correlation of sensor readings at a given time t with failure events at a future time t + lag. The optimal predictive lag for each sensor was then identified as the time lag that yielded the maximum absolute correlation value. This analysis is critical as it quantifies the time window in which each sensor’s data is most informative for forecasting an impending failure, thereby providing an empirical foundation for the subsequent causal discovery and feature engineering stages. In the cross-correlation plot shown in Figure 2, the red star annotated on each sensor’s curve denotes the optimal predictive lag—20 hours for Pressure, 14 hours for Vibration, and so forth. This marker identifies the specific time lag, measured in hours, at which the sensor’s signal exhibits the highest absolute Pearson correlation with the future failure event. Consequently, the red star highlights the most influential temporal offset for each variable, effectively quantifying the sensor’s most informative predictive window within the 24-hour forecasting horizon. 87
A Causality-Informed Framework Information Society, 2025, Ljubljana, Slovenia 3.3 Causal Graph Construction To elucidate the causal interdependencies between sensor signals and equipment failures, a causal graph was constructed using VARLiNGAM. This methodology first employs a Vector Autoregression (VAR) model to capture the linear, time-lagged relationships among the multivariate sensor time series. The optimal lag for the VAR model was adaptively informed by the preceding cross-correlation analysis to focus on the most predictive temporal window. Following the VAR estimation, the LiNGAM algorithm is applied to the resulting model residuals, or innovations. By exploiting the non-Gaussian nature of these innovations, LiNGAM uniquely identifies the contemporaneous causal structure—the instantaneous effects between variables—and determines the direction of influence, thereby producing a directed acyclic graph (DAG). The final output is a set of adjacency matrices representing the causal graph, where each non-zero entry quantifies the strength and direction of a causal link from one variable to another at a specific time lag. Our approach constructs a directed causal graph from time-series sensor data using the following steps: (1) Data Sorting and Integrity: Chronologically sort sensor data, verifying integrity and noting irregular intervals. (2) Variable Definition: Define variables which are vibration, rotation, pressure, voltage, and a binary failure indicator as the target node. (3) Causal Model Setup: Configure a VARLiNGAM [4] model with a specified lag order and BIC-based pruning. (4) Model Fitting: Fit the model to the prepared data matrix, applying regularization—by adding small Gaussian noise (e.g., 10 −6 )—when numerical instability arises during VARLiNGAM causal graph construction due to ill-conditioned matrices. (5) Adjacency Extraction: Extract adjacency matrices to identify directed edges, effect strengths, and corresponding lags. (6) Graph Assembly: Assemble the causal graph, categorizing edges by their relation to the target and between sensor variables. This workflow ensures that temporal ordering is respected and that detected causal links most likely represent meaningful relationships for predictive maintenance and further analytical investigations. Figure 3 presents the causal graph generated by the VARLiNGAM algorithm, illustrating the network of causal relationships between sensor telemetry (volt, pressure, vibration, rotate), machine properties (age), and the target failure event. In this graph, nodes represent the variables, and the directed edges (arrows) signify the direction of causality, with edge thickness corresponding to the strength of the effect. The labels on each edge quantify the causal strength and the time delay (lag) in hours. The analysis reveals a complex web of interactions, prominently highlighting that machine age is the most significant causal driver of failure, with an exceptionally strong effect strength at a lag of 6 hours. Other notable, though weaker, causal pathways are also identified, such as the influence of rotate on failure. This causal structure provides critical insights into the system’s dynamics, identifying the key variables and time-delayed interactions that precede a failure event. 3.4 Causality-Informed Feature Engineering We prepared the data by building a causality-informed feature vector grounded in the paper’s causal graph and a temporal causality Figure 3: Causal Graph analysis that selects per-sensor optimal prediction windows. Using a sliding feature window (typically 72 h), samples are formed from historical data only to avoid leakage. Feature construction proceeds in four stages: (1) basic statistics (mean, standard deviation, min/max, latest/earliest within the window); (2) causalityaligned temporal features computed at the optimal lags identified by causal analysis; (3) dynamics via trend slopes (linear regression), rolling volatility (standard deviation), and rates of change; and (4) cross-feature terms implied by the causal graph (e.g., voltage/rotation ratios and pressure–vibration correlations). Targets are defined for multiple horizons (1,6,12, and 24 h ahead) to enable early warnings at different lead times. The resulting dataset contains 150 features that integrate causal dependencies with temporal patterns. 3.5 Machine Learning Models Three classification algorithms, each configured with default hyperparameters, were evaluated using time-based data partitioning to mitigate the risk of data leakage. • Random Forest (RF): Ensemble method with 200 estimators, maximum depth of 15, and balanced class weights •XGBoost (XGB): Gradient boosting with 200 estimators, learning rate of 0.1, and automatic scale balancing • Gradient Boosting (GB): Scikit-learn implementation with 200 estimators and 0.8 subsample ratio Model performance was assessed using F1 Score metric appropriate for imbalanced classification: •F1-Score: Harmonic mean of precision and recall A time-series–aware data partitioning strategy was implemented using scikit-learn’s TimeSeriesSplit, which generates folds in chronological order by progressively expanding the training set with earlier observations and reserving subsequent periods for testing. This procedure ensures that all training data temporally precedes the corresponding test data. To approximate stratification and preserve class balance between rare failure and more frequent non-failure events, the folds were constructed to proportionally distribute failure cases across splits without introducing randomization. This design maintains the temporal integrity of the sensor data while supporting reliable model evaluation. 4 RESULTS AND DISCUSSION Figure 4 presents the comprehensive F1-score evaluation of all three models, while Figure 5 provides a comparative analysis 88
Information Society, 2025, Ljubljana, Slovenia Seyed Iman Hosseini, Klemen Kenda, and Dunja Mladenič Figure 4: F1-score evaluated over a 20-hour prediction horizon Figure 5: The XGBoost F1-score across a 20-hour prediction horizon, evaluated with and without a causality-informed feature vector of the XGBoost model with and without the causality-informed feature vector. Standard time-series models, particularly those trained on raw temporal data, consistently outperform causalityinformed approaches in predictive maintenance tasks, especially at extended prediction horizons. XGBoost, for instance, achieves F1 scores exceeding 94% for horizons beyond 10 hours, though performance declines with shorter windows due to reduced temporal context. In contrast, causality-informed models offer no competitive advantage—primarily due to the limitations of causal discovery conducted on data from a single machine. This narrow scope lacks the operational diversity and failure variability needed to infer generalizable causal structures, resulting in overfitting to machine-specific correlations and the exclusion of informative temporal features. These findings highlight the critical need for multi-machine datasets when applying causal methods, ensuring that inferred relationships reflect true causality rather than artifacts of constrained data. In addition, Longer prediction horizons (e.g., 20 hours) afford models access to extended historical windows (e.g., 72 hours), enhancing their ability to detect subtle patterns and causal signals. In contrast, short horizons (e.g., 1 hour) offer limited temporal context, increasing susceptibility to noise and overfitting. Causality-informed features such as optimal lag and causal strength are inherently better suited to longer windows, where failure patterns emerge gradually rather than abruptly. 5 FUTURE WORKS While this study establishes a robust, domain-agnostic framework for failure prediction, future work will focus on enhancing its transparency and causal reasoning capabilities. The integration of Explainable Artificial Intelligence (XAI) methods, such as SHAP or LIME, will provide transparent insights into the predictive models’ decision-making processes, fostering trust among users and enabling more informed maintenance decisions. Additionally, investigating counterfactual analysis will allow for exploring ’what-if’ scenarios to better understand the causal impacts of various factors on failure predictions. Alongside these enhancements, we will address the observed limitations of applying causality-informed models to data from a single machine. Specifically, we hypothesize that the lack of competitive advantage stems from the limited operational diversity and failure variability of a single-machine dataset, leading to overfitting. Future work will validate this hypothesis by expanding the dataset to include multiple machines, ensuring more generalizable insights into causal relationships and improving the robustness of predictive models. ACKNOWLEDGEMENTS We gratefully acknowledge the European Commission for its support of the Marie Skłodowska-Curie program through the Horizon Europe DN APRIORI project (GA 101073551). REFERENCES [1] Abdeldjalil Benhanifia, Zied Ben Cheikh, Paulo Moura Oliveira, Antonio Valente, and José Lima. 2025. Systematic review of predictive maintenance practices in the manufacturing sector. Intelligent Systems with Applications, 26, 200501. doi: https://doi.org/10.1016/j.iswa.2025.200501. [2] Arnab Biswas. 2025. Microsoft azure predictive maintenance. Accessed: 2025-05-20. (2025). https://www.kaggle.com/datasets/arnabbiswas1/micros oft-azure-predictive-maintenance/data. [3] Qing’an Cui, Jiao Lu, and Xianhui Yin. 2025. Causality enhanced deep learning framework for quality characteristic prediction via long sequence multivariate time-series data. Measurement Science and Technology, 36, (Mar. 2025), 3, (Mar. 2025). doi: 10.1088/1361-6501/adb05a. [4] LiNGAM Developers. 2025. VARLiNGAM — LiNGAM 1.10.0 documentation. https://lingam.readthedocs.io/ en / latest / tutorial / var.html. Accessed: 2025-06-25. (2025). [5] Karim Nadim, Ahmed Ragab, and Mohamed Salah Ouali. 2023. Data-driven dynamic causality analysis of industrial systems using interpretable machine learning and process mining. Journal of Intelligent Manufacturing, 34, (Jan. 2023), 57–83, 1, (Jan. 2023). doi: 10.1007/s10845-021-01903-y. [6] P. Nunes, J. Santos, and E. Rocha. 2023. Challenges in predictive maintenance – a review. CIRP Journal of Manufacturing Science and Technology, 40, 53–67. doi: https://doi.org/10.1016/j.cirpj.2022.11.004. [7] Margarida Da Rocha and Faísca Moreira. 2024. FACULDADE DE ENGENHARIA DA UNIVERSIDADE DO PORTO Data-Driven Predictive Maintenance for Component Life-Cycle Extension. Tech. rep. [8] Qipeng Wang, Shoubo Feng, and Min Han. 2025. Causal graph convolution neural differential equation for spatio-temporal time series prediction. Applied Intelligence, 55, (May 2025), 7, (May 2025). doi: 10.1007/s10489-025-06 287-7. [9] Sheng Wang, Qiang Zhao, Yinghua Han, and Jinkuan Wang. 2023. Root cause diagnosis for complex industrial process faults via spatiotemporal coalescent based time series prediction and optimized granger causality. Chemometrics and Intelligent Laboratory Systems, 233, (Feb. 2023). doi: 10.10 16/j.chemolab.2022.104728. [10] Xing Yang, Tian Lan, Hao Qiu, and Chen Zhang. 2025. Nonlinear causal discovery via dynamic latent variables. IEEE Transactions on Automation Science and Engineering.doi: 10.1109/TASE.2024.3522917. [11] Tanja Zerenner, Marc Goodfellow, and Peter Ashwin. 2021. Harmonic crosscorrelation decomposition for multivariate time series. Physical Review E, 103, (June 2021), 6, (June 2021). doi: 10.1103/PhysRevE.103.062213. [12] XIAOJUN ZHAO, PENGJIAN SHANG, and JINGJING HUANG. 2017. Several fundamental properties of dcca cross-correlation coefficient. Fractals, 25, 02, 1750017. eprint: https://doi.org/10.1142/S0218348X17500177. doi: 10.1142/S0218348X17500177. 89
Using Interactive Data Visualization for DeFi Market Analysis Daria Pavlova [email protected] Jožef Stefan International Postgraduate School Ljubljana, Slovenia Inna Novalija [email protected] Jožef Stefan Institute Ljubljana, Slovenia ABSTRACT Decentralized Finance (DeFi) presents unique analytical challenges with its data-rich, volatile, and multi-dimensional ecosystem. Static reports struggle to convey short-term dynamics and cross-sectional structure simultaneously. We present a comprehensive Business Intelligence (BI) solution featuring an automated Extract-Transform-Load (ETL) pipeline and interactive Tableau dashboard. Our ETL architecture processes data from three Application Programming Interfaces (APIs)—CoinGecko, DeFiLlama, and DexScreener—through validation and transformation stages, achieving 45-second execution time. The dashboard integrates Key Performance Indicators (KPIs), Total Value Locked (TVL) time-series, market categories analysis, and top movers panel with synchronized filters. Performance evaluation demonstrates 85-99% reduction in analysis time compared to manual methods. Three real-world use cases validate practical applicability: narrative rotation detection (28% investment returns), risk concentration monitoring (15% drawdown reduction), and competitive benchmarking. Our approach bridges the gap between complex DeFi data and actionable insights without requiring technical expertise. KEYWORDS DeFi, Business Intelligence, Tableau, TVL, KPI dashboards, Interactive Visualization, ETL Pipeline, Data Mining, Cryptocurrency 1 INTRODUCTION Decentralized Finance (DeFi) compresses high-frequency market activity—liquidity flows, incentive programs, and new protocol deployments—into datasets that change hourly. The ecosystem encompasses over 6,000 protocols managing billions in Total Value Locked (TVL), creating analytical complexity that traditional tools struggle to handle. Practitioners must simultaneously answer three critical questions: How big is the market now? (level KPIs), How is it moving? (time series), and What drives the crosssection? (categories, movers). Interactive visualization reduces cognitive load and increases pattern salience relative to static tables [ 3 , 7 , 10 , 11 ]. However, existing solutions present trade-offs: Dune Analytics requires Structured Query Language (SQL) expertise, Nansen charges $1,800 annually, while free alternatives like DeFiLlama offer limited visualization capabilities. Our goal is to demonstrate a compact, reproducible Business Intelligence workflow that democratizes DeFi analytics through automated data processing and intuitive visualization. Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). SiKDD 2025, October 6, 2025, Ljubljana, Slovenia ©2025 Copyright held by the owner/author(s). https://doi.org/10.70314/is.2025.sikdd.15 2 RELATED WORK 2.1 DeFi Analytics Landscape Surveys of DeFi systems [ 12 ] highlight the centrality of TVL, market capitalization, and volume as monitoring signals. Public APIs from CoinGecko and DeFiLlama expose these aggregates for research and dashboards, processing millions of daily transactions into consumable metrics.1 Recent advances in artificial intelligence have opened new frontiers in DeFi analysis. Chen et al. [ 1 ] proposed ensemble machine learning approaches for detecting rug pulls and protocol vulnerabilities, achieving 87% accuracy using features extracted from on-chain data and social signals. Their Random Forest model combined with Long Short-Term Memory (LSTM) networks demonstrated AI’s potential in risk assessment. However, these Machine Learning (ML) approaches require significant computational resources and technical expertise, creating barriers for non-technical analysts. Our solution complements these advanced techniques by providing immediate, interpretable insights through interactive visualization. 2.2 Business Intelligence and Visualization Classic data warehouse and BI literature formalizes metrics and dimensional modeling for decision support [ 6 ]. Industry guidance positions interactive platforms such as Tableau among leading tools for exploratory analysis [ 4 ]. Visualization principles—overview first, zoom and filter, details-on-demand [ 8 ]—map directly to dashboard layout patterns [3, 9]. Studies of graphical perception [ 2 , 5 ] explain why bars outperform pies for accurate comparisons, and why color semantics (green/red for gains/losses) aid preattentive detection [ 11 ]. We align with these findings in our chart choices and encodings. 3 SYSTEM ARCHITECTURE AND METHODOLOGY 3.1 ETL Pipeline Architecture Our ETL pipeline implements a modular, fault-tolerant architecture processing data through five stages. The architecture follows a standard Extract-Transform-Load pattern with additional validation and quality checks at each stage. Extract Layer: Three parallel API clients collect data from CoinGecko (200 tokens per page), DeFiLlama (6,000+ protocols), and DexScreener (100+ Decentralized Exchange pairs). Each client implements asynchronous Hypertext Transfer Protocol (HTTP) requests with exponential backoff (4-10 seconds) and retry logic (up to 5 attempts). Validation Layer: Implements four-level data quality checks: • Completeness: Missing value detection with fallback strategies •Consistency: Cross-validation between data sources •Timeliness: Timestamp validation (<1 hour freshness) 1 API documentation: https://www.coingecko.com/en/api, https://defillama.com/ docs/api. 90
SiKDD 2025, October 6, 2025, Ljubljana, Slovenia D. Pavlova Figure 1: ETL Pipeline Architecture: Data flows from three APIs through validation and transformation stages to produce four CSV files for dashboard visualization. The system processes 6,000+ protocols with automated retry logic and data quality checks. • Accuracy: Outlier detection using Median Absolute Deviation (MAD) Transform Layer: Processes validated data through three streams: • Normalize: Converts to tidy format with Coordinated Universal Time (UTC) timestamps • Features: Calculates rolling statistics and market sentiment •Aggregate: Groups by time windows and categories Load Layer: Exports processed data as Comma-Separated Values (CSV) files optimized for Tableau consumption. 3.2 Dashboard Design Methodology The dashboard layout follows Shneiderman’s Visual Information Seeking Mantra [ 8 ]: overview first, zoom and filter, then detailson-demand. Layout Structure: • Top Row: Four KPI cards displaying market totals with 24-hour changes • Middle Section: TVL time-series (left, 60% width) and Top Movers panel (right, 40% width) • Bottom Section: Category bars (left) and pie chart (right) for market structure analysis • Right Sidebar: Interactive filters for Time Window, Category Metric, and Top N selections 4 PERFORMANCE EVALUATION 4.1 System Performance Metrics We evaluated system performance across three dimensions: Response Time: •Initial dashboard load: 3.2s ±0.5s (n=100) •Filter operations: 1.8s ±0.3s •ETL pipeline execution: 45s complete, 8s incremental Data Processing Efficiency: •Batch processing: 50-100 protocols per batch •API delay: 0.1s between requests •Memory usage: Peak 256MB •Data volume: 6,000+ protocols, 200 tokens/page User Efficiency Gains: • Market overview generation: 15 min → 5 sec (99.4% reduction) • Sector rotation analysis: 30 min → 2 min (93.3% reduction) • Top movers identification: 10 min → instant (100% automation) 4.2 Comparison with Existing Solutions Table 1: Feature Comparison with Industry Solutions Feature Our Solution Dune Nansen DeFiLlama Cost Free $390/yr $1,800/yr Free No-code Interface ✓×✓ ✓ Custom ETL ✓× × × Response Time <2s 5-30s <3s <1s Visualization Types 4 Unlimited 10+ 2 Data Sources 3 Multiple Multiple 1 Historical Data 30 days All All Limited Our solution occupies a unique position: more sophisticated than DeFiLlama’s basic charts, more accessible than Dune’s SQL requirements, and more affordable than Nansen’s premium tiers. 5 RESULTS AND USE CASE VALIDATION 5.1 Dashboard Implementation The integrated dashboard combines multiple analytical views with synchronized filtering capabilities. The design synthesizes four key data dimensions: • KPI Header: Market metrics provide immediate context—$2.86T total market cap with 56.1% BTC dominance indicates riskoff sentiment • TVL Time-Series: Shows capital deployment patterns across protocols, with upward trajectory suggesting renewed confidence • Top Movers Panel: Highlights outliers—clustering in specific sectors signals narrative emergence • Category Analysis: Reveals market concentration—top 3 sectors comprise 51% of total value 5.2 Use Case Validation Use Case 1: Narrative Rotation Detection An investment fund utilized the dashboard to identify emerging trends in Liquid Staking Derivatives (LSDs). When multiple LSD protocols appeared in Top Movers with 40%+ gains while category volume increased 3x, they allocated capital early, achieving 28% returns over two weeks. Use Case 2: Risk Concentration Analysis A DeFi protocol team monitored market concentration using the category pie chart. When the top 3 categories exceeded 65% of total market cap (Herfindahl-Hirschman Index >0.25), they adjusted treasury diversification strategy, reducing drawdown by 15% during the subsequent correction. Use Case 3: Competitive Benchmarking Protocol developers tracked their TVL growth relative to category peers. The synchronized time-series view revealed their incentive program launched 3 days after competitors but achieved 2x the TVL growth rate, validating their tokenomics design. 6 DISCUSSION 6.1 Synthesis for Decision-Making The dashboard enables multi-dimensional analysis through synchronized views: 91
[Document text truncated for crawler view.]