scieee AI-readable full text Open interactive document viewer

Middle-European Conference on Applied Theoretical Computer Science (MATCOS) 2025

Jožef Stefan Institute

Full text

6. – 10. oktober 2025 l 6 – 10 October 2025 Koper, Slovenia IS 2025 Srednjeevropska konferenca o uporabnem teoretičnem računalništvu in informatiki (MATCOS) Middle-European Conference on Applied Theoretical Computer Science (MATCOS) Uredniki l Editors: Andrej Brodnik, Gábor Galambos, Rok Požar INFORMACIJSKA DRUZBA INFORMATION SOCIETY Zbornik 28. mednarodne multikonference Zvezek N Proceedings of the 28th International Multiconference Volume N ˇ Zbornik 28. mednarodne multikonference INFORMACIJSKA DRUŽBA – IS 2025 Zvezek N Proceedings of the 28th International Multiconference INFORMATION SOCIETY – IS 2025 Volume N Srednjeevropska konferenca o uporabnem teoretičnem računalništvu in informatiki (MATCOS) Middle-European Conference on Applied Theoretical Computer Science (MATCOS) Uredniki / Editors Andrej Brodnik, Gábor Galambos, Rok Požar http://is.ijs.si 9.–10. oktober 2025 / 9–10 October 2025 Koper, Slovenia Uredniki: Andrej Brodnik Univerza na Primorskem, Koper in Univerza v Ljubljani, Ljubljana Gábor Galambos Univerza v Szegedu, Szeged Rok Požar Univerza na Primorskem, Koper Založnik: Institut »Jožef Stefan«, Ljubljana Priprava zbornika: Mitja Lasič, Vesna Lasič, Lana Zemljak Oblikovanje naslovnice: Vesna Lasič, uporabljena slika iz Pixabay Dostop do e-publikacije: http://library.ijs.si/Stacks/Proceedings/InformationSociety Ljubljana, oktober 2025 Informacijska družba ISSN 2630-371X DOI: https://doi.org/10.70314/is.2025.matcos Kataložni zapis o publikaciji (CIP) pripravili v Narodni in univerzitetni knjižnici v Ljubljani COBISS.SI-ID 255599107 ISBN 978-961-264-332-4 (PDF) PREDGOVOR MULTIKONFERENCI INFORMACIJSKA DRUŽBA 2025 28. mednarodna multikonferenca Informacijska družba se odvija v času izjemne rasti umetne inteligence, njenih aplikacij in vplivov na človeštvo. Vsako leto vstopamo v novo dobo, v kateri generativna umetna inteligenca ter drugi inovativni pristopi oblikujejo poti k superinteligenci in singularnosti, ki bosta krojili prihodnost človeške civilizacije. Naša konferenca je tako hkrati tradicionalna znanstvena in akademsko odprta, pa tudi inkubator novih, pogumnih idej in pogledov. Letošnja konferenca poleg umetne inteligence vključuje tudi razprave o perečih temah današnjega časa: ohranjanje okolja, demografski izzivi, zdravstvo in preobrazba družbenih struktur. Razvoj UI ponuja rešitve za številne sodobne izzive, kar poudarja pomen sodelovanja med raziskovalci, strokovnjaki in odločevalci pri oblikovanju trajnostnih strategij. Zavedamo se, da živimo v obdobju velikih sprememb, kjer je ključno, da z inovativnimi pristopi in poglobljenim znanjem ustvarimo informacijsko družbo, ki bo varna, vključujoča in trajnostna. V okviru multikonference smo letos združili dvanajst vsebinsko raznolikih srečanj, ki odražajo širino in globino informacijskih ved: od umetne inteligence v zdravstvu, demografskih in družinskih analiz, digitalne preobrazbe zdravstvene nege ter digitalne vključenosti v informacijski družbi, do raziskav na področju kognitivne znanosti, zdrave dolgoživosti ter vzgoje in izobraževanja v informacijski družbi. Pridružujejo se konference o legendah računalništva in informatike, prenosu tehnologij, mitih in resnicah o varovanju okolja, odkrivanju znanja in podatkovnih skladiščih ter seveda Slovenska konferenca o umetni inteligenci. Poleg referatov bodo okrogle mize in delavnice omogočile poglobljeno izmenjavo mnenj, ki bo pomembno prispevala k oblikovanju prihodnje informacijske družbe. »Legende računalništva in informatike« predstavljajo domači »Hall of Fame« za izjemne posameznike s tega področja. Še naprej bomo spodbujali raziskovanje in razvoj, odličnost in sodelovanje; razširjeni referati bodo objavljeni v reviji Informatica, s podporo dolgoletne tradicije in v sodelovanju z akademskimi institucijami ter strokovnimi združenji, kot so ACM Slovenija, SLAIS, Slovensko društvo Informatika in Inženirska akademija Slovenije. Vsako leto izberemo najbolj izstopajoče dosežke. Letos je nagrado Michie-Turing za izjemen življenjski prispevek k razvoju in promociji informacijske družbe prejel Niko Schlamberger, priznanje za raziskovalni dosežek leta pa Tome Eftimov. »Informacijsko limono« za najmanj primerno informacijsko tematiko je prejela odsotnost obveznega pouka računalništva v osnovnih šolah. »Informacijsko jagodo« za najboljši sistem ali storitev v letih 2024/2025 pa so prejeli Marko Robnik Šikonja, Domen Vreš in Simon Krek s skupino za slovenski veliki jezikovni model GAMS. Iskrene čestitke vsem nagrajencem! Naša vizija ostaja jasna: prepoznati, izkoristiti in oblikovati priložnosti, ki jih prinaša digitalna preobrazba, ter ustvariti informacijsko družbo, ki koristi vsem njenim članom. Vsem sodelujočim se zahvaljujemo za njihov prispevek — veseli nas, da bomo skupaj oblikovali prihodnje dosežke, ki jih bo soustvarjala ta konferenca. Mojca Ciglarič, predsednica programskega odbora Matjaž Gams, predsednik organizacijskega odbora i FOREWORD TO THE MULTICONFERENCE INFORMATION SOCIETY 2025 The 28th International Multiconference on the Information Society takes place at a time of remarkable growth in artificial intelligence, its applications, and its impact on humanity. Each year we enter a new era in which generative AI and other innovative approaches shape the path toward superintelligence and singularity — phenomena that will shape the future of human civilization. The conference is both a traditional scientific forum and an academically open incubator for new, bold ideas and perspectives. In addition to artificial intelligence, this year’s conference addresses other pressing issues of our time: environmental preservation, demographic challenges, healthcare, and the transformation of social structures. The rapid development of AI offers potential solutions to many of today’s challenges and highlights the importance of collaboration among researchers, experts, and policymakers in designing sustainable strategies. We are acutely aware that we live in an era of profound change, where innovative approaches and deep knowledge are essential to creating an information society that is safe, inclusive, and sustainable. This year’s multiconference brings together twelve thematically diverse meetings reflecting the breadth and depth of the information sciences: from artificial intelligence in healthcare, demographic and family studies, and the digital transformation of nursing and digital inclusion, to research in cognitive science, healthy longevity, and education in the information society. Additional conferences include Legends of Computing and Informatics, Technology Transfer, Myths and Truths of Environmental Protection, Knowledge Discovery and Data Warehouses, and, of course, the Slovenian Conference on Artificial Intelligence. Alongside scientific papers, round tables and workshops will provide opportunities for in-depth exchanges of views, making an important contribution to shaping the future information society. Legends of Computing and Informatics serves as a national »Hall of Fame« honoring outstanding individuals in the field. We will continue to promote research and development, excellence, and collaboration. Extended papers will be published in the journal Informatica, supported by a long-standing tradition and in cooperation with academic institutions and professional associations such as ACM Slovenia, SLAIS, the Slovenian Society Informatika, and the Slovenian Academy of Engineering. Each year we recognize the most distinguished achievements. In 2025, the Michie-Turing Award for lifetime contribution to the development and promotion of the information society was awarded to Niko Schlamberger, while the Award for Research Achievement of the Year went to Tome Eftimov. The »Information Lemon« for the least appropriate information-related topic was awarded to the absence of compulsory computer science education in primary schools. The »Information Strawberry« for the best system or service in 2024/2025 was awarded to Marko Robnik Šikonja, Domen Vreš and Simon Krek together with their team, for developing the Slovenian large language model GAMS. We extend our warmest congratulations to all awardees. Our vision remains clear: to identify, seize, and shape the opportunities offered by digital transformation, and to create an information society that benefits all its members. We sincerely thank all participants for their contributions and look forward to jointly shaping the future achievements that this conference will help bring about. Mojca Ciglarič, Chair of the Program Committee Matjaž Gams, Chair of the Organizing Committee iiii KONFERENČNI ODBORI CONFERENCE COMMITTEES International Programme Committee Organizing Committee Vladimir Bajic, South Africa Heiner Benking, Germany Se Woo Cheon, South Korea Howie Firth, UK Olga Fomichova, Russia Vladimir Fomichov, Russia Vesna Hljuz Dobric, Croatia Alfred Inselberg, Israel Jay Liebowitz, USA Huan Liu, Singapore Henz Martin, Germany Marcin Paprzycki, USA Claude Sammut, Australia Jiri Wiedermann, Czech Republic Xindong Wu, USA Yiming Ye, USA Ning Zhong, USA Wray Buntine, Australia Bezalel Gavish, USA Gal A. Kaminka, Israel Mike Bain, Australia Michela Milano, Italy Derong Liu, Chicago, USA Toby Walsh, Australia Sergio Campos-Cordobes, Spain Shabnam Farahmand, Finland Sergio Crovella, Italy Matjaž Gams, chair Mitja Luštrek Lana Zemljak Vesna Koricki Mitja Lasič Blaž Mahnič Programme Committee Mojca Ciglarič, chair Bojan Orel Franc Solina Viljan Mahnič Cene Bavec Tomaž Kalin Jozsef Györkös Tadej Bajd Jaroslav Berce Mojca Bernik Marko Bohanec Ivan Bratko Andrej Brodnik Dušan Caf Saša Divjak Tomaž Erjavec Bogdan Filipič Andrej Gams Matjaž Gams Mitja Luštrek Marko Grobelnik Nikola Guid Marjan Heričko Borka Jerman Blažič Džonova Gorazd Kandus Urban Kordeš Marjan Krisper Andrej Kuščer Jadran Lenarčič Borut Likar Janez Malačič Olga Markič Dunja Mladenič Franc Novak Vladislav Rajkovič Grega Repovš Ivan Rozman Niko Schlamberger Gašper Slapničar Stanko Strmčnik Jurij Šilc Jurij Tasič Denis Trček Andrej Ule Boštjan Vilfan Baldomir Zajc Blaž Zupan Boris Žemva Leon Žlajpah Niko Zimic Rok Piltaver Toma Strle Tine Kolenik Franci Pivec Uroš Rajkovič Borut Batagelj Tomaž Ogrin Aleš Ude Bojan Blažica Matjaž Kljun Robert Blatnik Erik Dovgan Špela Stres Anton Gradišek iii iv KAZALO / TABLE OF CONTENTS Srednjeevropska konferenca o uporabnem teoretičnem računalništvu in informatiki (MATCOS) / Middle-European Conference on Applied Theoretical Computer Science (MATCOS) ....................... 1 PREDGOVOR / FOREWORD ............................................................................................................................... 3 PROGRAMSKI ODBORI / PROGRAMME COMMITTEES ............................................................................... 5 Merging Operations in the Open-Shop Scheduling Problem / Baldouski Daniil, Dávid Balázs, Krész Miklós .... 7 Flexibility vs. efficiency: a study in sawmill scheduling / Kebelei Csaba, Hegyháti Máté ................................. 11 Implementation of a Vehicle and Driver Scheduling Model: a Case Study / Árgilán Viktor, Békési József, Galambos Gábor, Papp Imre ............................................................................................................................ 15 ALGatorGraph: A Java Library for Graph Generation and Manipulation within the ALGator System / Hren Boštjan, Dobravec Tomaž ................................................................................................................................ 19 Engineering CSFLOC: A Subsumption-Driven Clause-Counting SAT Solver / Kusper Gábor ........................ 23 Non-redundant Systems of Independence Atoms in Relational Databases / Alland Lucas, Sali Attila, Wu Nicole .......................................................................................................................................................................... 27 Scrambler Automaton Block Cipher for IoT Devices / Dömösi Pál, Horváth Géza ............................................ 30 Automata for context-free trace languages and permutation languages / Nagy Benedek .................................... 35 Towards a Category-Theoretic Informatics Model of PSPP Linkages in Biomaterials / Tahalea Sylvert Prian, Krész Miklós .................................................................................................................................................... 39 Cost-Sensitive Overview of Model Ensembling for Machine-Generated Text Detection / Kiss Mihály, Berend Gábor ................................................................................................................................................................ 44 Hybrid Reinforcement Learning Enhanced Genetic Algorithm for the Capacitated Vehicle Routing Problem with Split Deliveries and Heterogeneous Fleet / Dabbous Ahmed, Bóta András ........................................... 48 Empiric results on the achievable performance gains by the inclusion of instance-specific information in a scheduling optimizer / Hegyháti Máté ............................................................................................................ 52 Bi-Level Routing and Scheduling / Quilliot Alain, Toussaint Hélène ................................................................. 56 Adiscrete event simulation model for analyzing the wood waste reverse supply chain / Kovačević Nikola, Tavzes Črtomir, Dávid Balázs.......................................................................................................................... 60 Robust (re)Design of Material Flow in Circular Networks – a Scientific Approach / Szaller Ádám, Dávid Balázs, Egri Péter, Krész Miklós, Váncza József ............................................................................................. 64 Harvest plan generation in precision agriculture / Horvat Štefan, Strnad Damjan, Mongus Domen, Brumen Matej ................................................................................................................................................................ 68 Reducing #SAT to k-clique enumeration / Szabó Sándor, Zaválnij Bogdán ....................................................... 72 Modularity aware graph clustering for exploratory tasks with a case study of the biomass supply chain / Tahalea Sylvert Prian, Kawa Arkadiusz, Dávid Balázs ................................................................................... 75 ASynthetic Multi-View Tracking and 3D Pose Dataset for Automated Airport Visual Surveillance / Mansour Ahmed, Beleznai Csaba, Oberweger Fabio F., Widhalm Verena, Kirillova Nadezda, Possegger Horst ......... 80 Sphere Target-Based Point Cloud Registration in a Railway Safety Application / Podgorelec David, Lukač Luka, Pečnik Sašo, Repnik Blaž, Žalik Borut .................................................................................................. 84 Indeks avtorjev / Author index ................................................................................................................... 89 v 6 Merging Operations in the Open-Shop Scheduling Problem Daniil Baldouski [email protected] University of Primorska, IAM Koper, Slovenia Balázs Dávid [email protected] InnoRenew CoE Izola, Slovenia University of Primorska, IAM and FAMNIT Koper, Slovenia Miklós Krész [email protected] InnoRenew CoE Izola, Slovenia University of Primorska, IAM and FAMNIT Koper, Slovenia University of Szeged Szeged, Hungary ABSTRACT In this work we consider the open-shop scheduling problem with operation batching (OSSP-OB), which extends classical openshop scheduling by allowing operations of the same category to be processed simultaneously on a machine. To address this problem, we develop an exact procedure for constructing optimal batches with the corresponding machine scheduling by developing a mixed-integer linear programming (MILP) model. The effectiveness of the method is evaluated on a dedicated benchmark instance set. KEYWORDS job scheduling, batch scheduling, open-shop scheduling, mixedinteger linear programming 1 INTRODUCTION We consider a scheduling problem that combines ideas from open-shop and batch scheduling. Each job consists of several operations, each associated with a category and executable on one of the suitable machines for a specified processing time. Operations within a job can be processed in arbitrary order, so the problem environment belongs to the class of open-shop scheduling (OSSP) [3, 4, 1]. In addition, we allow multiple operations of the same category to be executed simultaneously on a machine. This grouping procedure, which we call merging, is related to the notion of batching in batch scheduling [5, 2, 7], but differs as it applies to operations rather than jobs. The resulting formulation, referred to as the open-shop scheduling problem with operation batching (OSSP-OB), extends the classical OSSP by incorporating merging of operations and captures a wider range of scheduling scenarios. The OSSP can then be viewed as a special case of OSSP-OB in which a unique category is provided for every operation, thus leading OSSP-OB to be an NP-hard problem in the general case. 2 PROBLEM DESCRIPTION The environment of the open-shop scheduling problem with the merging of operations (operations batching) is given by the sets of jobs, operations, and machines, with various parameters for each of the sets. Formally, OSSP-OB environment consists of: Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia ©2025 Copyright held by the owner/author(s). • The time horizon of the problem is represented by the discretized set of consecutive time units 𝑇. • A finite set 𝐽 of jobs. For each job 𝑗∈𝐽 there is a time window [𝑎𝑗,𝑏𝑗] within the time horizon of the problem for which the operations of this job should be analyzed and approved. • A finite set 𝑂 of operations associated with jobs. For each job 𝑗∈𝐽 , a set 𝑂𝑗⊆𝑂 of operations associated with the job. • A finite set 𝐶 of operation categories associated with operations. Each operation 𝑜∈𝑂 has a category 𝐶(𝑜) and there are no two operations of the same category that belong to the same job. The time it takes to perform the operation is the same within the operation category and is then defined for each category 𝑐∈𝐶 as ℓ(𝑐) (the length of the category). The set of all operations of the same category is denoted by 𝑂(𝑐)for each category 𝑐∈𝐶. • A finite set 𝑀 of machines. For each machine 𝑚∈𝑀 , a set 𝐶𝑚⊆𝐶 of operation categories that can be carried out on the machine. • Operations can be performed in batches if they belong to the same category. This will be referred to as merging of operations. Each merged operation 𝑠∈𝑆 is a set of operations, that is, 𝑠⊆𝑂 such that 𝑠 contains operations of the same category. The set of all merged operations can be understood as a set 𝑆⊆ P(𝑂) that contains all the possible valid merged operations. Additional parameters can include, but are not limited to: • capacity parameters for the number of operations for each machine, •capacity of the number of merged jobs for each category, • time window for which the operations of the job should be analyzed, •internal deadline of the job, •periods of activity/inactivity, • time to merge operations (e.g. constant, flexible, proportional to the execution time, proportional to the number of merged operations, and more) • precedence relations between the operations within the same job, •precedence relations between jobs. In this work, we consider two additional parameters that are of high importance in the use-case of laboratory experiments: internal deadlines and activity periods. The internal deadline ( 𝑑𝑗 ) can be defined as 𝑘 time units before 𝑏𝑗 , which is the desired internal target for operations to be analyzed and approved, which then gives some time ( 𝑘 time units) for corrections in the case of exceptions (failures). A good 7 Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Baldouski et al. example of a merging parameter would be a limit on the number of operations that can be merged ( 𝑚(𝐶) ) for each category 𝑐∈𝐶 . In the context of QA laboratory experiments, personnel scheduling plays an important role, since most tasks (such as machine preparation, machine loading, validation of the results, and more) are performed manually, each requiring a specific employee qualification to handle it. In this work, we consider personnel scheduling in a simplified format, assuming that all manual tasks are performed instantaneously and with infinite personnel capacity. Let us call periods of inactivity all the gaps in the time horizon (e.g., weekends, national holidays, other special events) during which machines are unavailable and none of the tasks can be performed. Let us then consider an activity period to be any maximal interval of time units that does not contain periods of inactivity. An example of the problem (Fig. 1) with a feasible solution (Fig. 2) is demonstrated. In this example, we are given 3jobs ( 𝐽={𝐽 1 , 𝐽 2 , 𝐽 3 } , sorted by the earliest allowed start time), each of them containing 3 − 4operations. The time windows of the jobs are given ( [𝐴(𝑗), 𝐵(𝑗)] for each job 𝑗∈𝐽 ), as well as the lengths of the operations (categories). The colors of the operations represent their categories and there are 6categories in total 𝐶={𝐶 1 , . . . ,𝐶 6 } . For example, 𝑂(𝐽 2 )={𝑂 1 ,𝑂 2 ,𝑂 3 } such that 𝐶(𝑂3)=𝐶1. There are 3machines ( 𝑀={𝑀 1 , 𝑀 2 , 𝑀 3 } ), each of them corresponds to a set of categories, of which operations are allowed to be processed. In this example, 𝐶𝑀1={𝐶 1 , . . . ,𝐶 4 } , while 𝐶𝑀2=𝐶𝑀3={𝐶 4 , . . . ,𝐶 6 } . The example does not dive deep into all the possible parameters of the problem, assuming that merging is allowed in case operations belong to the same category, while satisfying the time constraints. A feasible solution (Fig. 2) is given in the form of performed merges, (merged) operations-machines assignment, with the (merged) operations timeline. There are 4potential merges for categories 𝐶 1 ,𝐶 2 ,𝐶 4and 𝐶 5, representing different types of merges that can occur. For example, merging of category 𝐶 1is not allowed due to time limitation ( ℓ(𝐶 1 )>𝐵(𝐽 1 ) −𝐴(𝐽 2 ) ). Because of this, the schedule of machine 𝑀 1becomes tight, making the merging of operations of category 𝐶 2necessary. Merges of categories 𝐶 4and 𝐶 5are possible but not necessary and would be performed (or not) based on the specifics of the objective function. In general, the goal of this work is to find the most efficient scheduling of operations to machines given the merging constraints, with the objectives to consider including tardiness, number of batches, flowtime, makespan, and overall costs. 3 SOLUTION METHODOLOGY In this work, we formulate the OSSP-OB as a mixed-integer linear programming (MILP) model that integrates both the merging of operations and the scheduling of merged operations. 3.1 Problem data For each 𝑠∈𝑆 , 𝑗∈𝐽 , 𝑜∈𝑂 , 𝑚∈𝑀 , 𝑐∈𝐶 , and 𝜏∈𝑇 , we establish the following parameters. (1) Assignment parameters: •𝑎𝑂𝐽 (𝑜) - the index of the job associated with the operation 𝑜. •𝑎𝑂𝐶(𝑜)- the index of the category of the operation 𝑜. •𝑎𝑀𝐶(𝑚, 𝑐)= 1if and only if operations of category 𝑐 can be carried out on the machine 𝑚. (2) Processing time parameters: •𝑡𝐶(𝑐) - time required to perform the operations of category 𝑐. •𝑡 𝐽𝑎(𝑗) - the start of the time interval for which the operations of job 𝑗should be analyzed. •𝑡 𝐽𝑏(𝑗) - the end of the time interval for which the operations of job 𝑗should be analyzed. •𝑡 𝐽𝑑(𝑗) - the internal deadline of the time interval for which the operations of job 𝑗should be analyzed. •𝐴set of activity periods, 𝐴⊆𝑇. (3) Auxiliary parameters. A set of derived parameters follows, simplifying the connection between the set of merged operations and: •Categories: –𝑎𝑆𝐶(𝑠,𝑐)= 1if 𝑐=𝑎𝑂𝐶(𝑜) for any 𝑜∈𝑠 ,0otherwise, for a merged operation 𝑠. – Let 𝑐𝑠∈𝐶 be the category of a merged operation 𝑠 , that is 𝑎𝑆𝐶(𝑠, 𝑐𝑠)=1. •Jobs: –𝑎𝑆𝐽 (𝑠, 𝑗)= 1if 𝑗∈ {𝑎𝑂𝐽 (𝑜) : 𝑜∈𝑠} and 0otherwise, for a merged operation 𝑠 . That is, if the merged operation 𝑠 and the job 𝑗 are associated with each other. •Machines: –𝑎𝑆𝑀(𝑠,𝑚)= 1if there is at least one category 𝑐∈𝐶 such that 𝑎𝑆𝐶(𝑠,𝑐)= 1and 𝑎𝑀𝐶(𝑚, 𝑐)= 1,0otherwise, for merged operation 𝑠and machine 𝑚. – Let set 𝑀𝑠 be the set of machines compatible with the merged operation 𝑠, that is, for every 𝑠∈𝑆: 𝑀𝑠={𝑚|𝑚∈𝑀, 𝑎𝑆𝑀 (𝑠,𝑚)=1}. Additionally, we introduce a big M parameter based on the time horizon 𝑇: M=2· |𝑇| + 1. 3.2 Variables For each 𝑠∈𝑆 , 𝑗∈𝐽 , 𝑜∈𝑂 , 𝑚∈𝑀 , 𝑐∈𝐶 , and 𝜏∈𝑇 , we establish the following variables. (1) Assignment variables: •𝑥𝑠,𝑚,𝜏 = 1if and only if merged operation 𝑠 is active on machine 𝑚in time unit 𝜏. •𝑦𝑠,𝜏 = 1if and only if the merged operation 𝑠 starts in the time unit 𝜏. •𝑧𝑠,𝑚 = 1if and only if the merged operation 𝑠 and machine 𝑚 are assigned to each other during the time horizon 𝑇. •𝑟𝑠= 1if and only if the merged operation 𝑠 is active during the time horizon 𝑇. (2) Auxiliary variables. A set of helpful variables to simplify the model formulation, grouping the merged operations parameters together with the decision variable 𝑟𝑠 that represents the activity status of a merged operation: •𝑟𝑐 𝑠,𝑐 =𝑟𝑠·𝑎𝑆𝐶(𝑠, 𝑐) for merged operation 𝑠 and category 𝑐. •𝑟𝑗 𝑠,𝑗 =𝑟𝑠·𝑎𝑆𝐽 (𝑠, 𝑗)for merged operation 𝑠and job 𝑗. •𝑟𝑚 𝑠,𝑚 =𝑟𝑠·𝑎𝑆𝑀(𝑠,𝑚) for merged operation 𝑠 and machine 𝑚. •𝑟𝑐 𝑗 𝑠,𝑐,𝑗 =𝑟𝑐 𝑠,𝑐 ·𝑟𝑗 𝑠,𝑗 for merged operation 𝑠 , category 𝑐 and job 𝑗. •𝑘𝑠,𝑚,𝑗,𝜏 =𝑟𝑗 𝑠,𝑗 ·𝑥𝑠,𝑚,𝜏 for merged operation 𝑠 , machine 𝑚, category 𝑐and time unit 𝜏. 8 Merging Operations in the Open-Shop Scheduling Problem Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Figure 1: Example of OSSP-OB with 3jobs and 3−4operations per job. Figure 2: Example of an OSSP-OB solution, demonstrating operations to machines scheduling. 3.3 Constraints (1) Auxiliary variables definitions. For each 𝑠∈𝑆 , 𝑚∈𝑀 , 𝑐∈𝐶,𝑗∈𝐽,𝜏∈𝑇: •𝑟𝑐 𝑠,𝑐 =𝑟𝑠·𝑎𝑆𝐶(𝑠, 𝑐). •𝑟𝑗 𝑠,𝑗 =𝑟𝑠·𝑎𝑆𝐽 (𝑠, 𝑗). •𝑟𝑚 𝑠,𝑚 =𝑟𝑠·𝑎𝑆𝑀(𝑠,𝑚). •Linearization of 𝑟𝑐 𝑗 : 𝑟𝑐 𝑗 𝑠,𝑐,𝑗 ≤𝑟𝑐 𝑠,𝑐 (1) 𝑟𝑐 𝑗 𝑠,𝑐,𝑗 ≤𝑟𝑗 𝑠,𝑗 (2) 𝑟𝑐 𝑠,𝑐 +𝑟𝑗 𝑠,𝑗 −1≤𝑟𝑐 𝑗 𝑠,𝑐,𝑗 (3) •Linearization of 𝑘: 𝑘𝑠,𝑚,𝑗,𝜏 ≤𝑟𝑗 𝑠,𝑗 (4) 𝑘𝑠,𝑚,𝑗,𝜏 ≤𝑥𝑠,𝑚,𝜏 (5) 𝑟𝑗 𝑠,𝑗 +𝑥𝑠,𝑚,𝜏 −1≤𝑘𝑠,𝑚,𝑗,𝜏 (6) (2) At any time unit, merged operations and machines have at most one unique schedule. For each time unit 𝜏∈𝑇: •For each merged operation 𝑠∈𝑆: ∑︁ 𝑚∈𝑀 𝑥𝑠,𝑚,𝜏 ≤1(7) •For each machine 𝑚∈𝑀: ∑︁ 𝑠∈𝑆 𝑥𝑠,𝑚,𝜏 ≤1(8) (3) For each merged operation 𝑠∈𝑆 to be scheduled ( 𝑟𝑠= 1): •There is exactly one starting time unit: ∑︁ 𝜏∈𝑇 𝑦𝑠,𝜏 =𝑟𝑠(9) •Exactly one machine is assigned: ∑︁ 𝑚∈𝑀 𝑧𝑠,𝑚 =𝑟𝑠(10) (4) Merged operations can not be scheduled on machines that do not share the corresponding category. For each merged operation 𝑠∈𝑆: •For each machine 𝑚∈𝑀\𝑀𝑠: ∑︁ 𝜏∈𝑇 𝑥𝑠,𝑚,𝜏 =0(11) 𝑧𝑠,𝑚 =0(12) •For each machine 𝑚∈𝑀𝑠: ∑︁ 𝜏∈𝑇 𝑥𝑠,𝑚,𝜏 =𝑟𝑠·𝑡𝐶(𝑐𝑠)(13) (5) Merged operations should start and end within the corresponding time interval. For each merged operation 𝑠∈𝑆 , job 𝑗∈𝐽and category 𝑐∈𝐶: ∑︁ 𝜏∈𝑇,𝜏<𝑡 𝐽 𝑎(𝑗) 𝑦𝑠,𝜏 ≤ (1−𝑟𝑗 𝑠,𝑗 )M (14) ∑︁ 𝜏∈𝑇,𝜏>𝑡 𝐽 𝑏(𝑗)−𝑡𝐶 (𝑐) 𝑦𝑠,𝜏 ≤ (1−𝑟𝑐 𝑗 𝑠,𝑐,𝑗 )M (15) (6) Merged operations can be performed only on the dedicated machines. For each merged operation 𝑠∈𝑆 , machine 𝑚∈𝑀and category 𝑐∈𝐶: 𝑧𝑠,𝑚 ≤𝑟𝑚 𝑠,𝑚 (16) (7) Each merged operation 𝑠∈𝑆 occupies 𝑡𝐶(𝑐) consecutive time units for a unique 𝑐∈𝐶 such that 𝑟𝑐 𝑠,𝑐 = 1. For each 9 Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Baldouski et al. merged operation 𝑠∈𝑆 , machine 𝑚∈𝑀 , category 𝑐∈𝐶 , time unit 𝜏∈𝑇and 𝜏′∈ {0,1, . . . , 𝑡𝐶 (𝑐) − 1}: 𝑥𝑠,𝑚,𝜏+𝜏′≥𝑦𝑠,𝜏 − (1−𝑧𝑠,𝑚)M − (1−𝑟𝑐 𝑠,𝑐 )M,(17) where 𝜏+𝜏′denotes the time 𝜏′units after time 𝜏. (8) Merged operations can not be active earlier than they start. For each merged operation 𝑠∈𝑆 , for each 𝜏∈𝑇 and 𝜏′∈𝑇 such that 𝜏′<𝜏: 𝑥𝑠,𝑚,𝜏′≤𝑦𝑠,𝜏 (18) (9) Machines are not assigned merged operations during the periods of inactivity. For each time unit 𝜏∈𝑇\𝐴: ∑︁ 𝑠∈𝑆,𝑚∈𝑀 𝑥𝑠,𝑚,𝜏 =0(19) (10) For each operation 𝑜∈𝑂 only one merged operation 𝑠∈𝑆 such that 𝑜∈𝑠 can be active during the time horizon. For each 𝑜∈𝑂:∑︁ 𝑠∈𝑆,𝑜 ∈𝑠 𝑟𝑠=1(20) 3.4 Objective function Minimize the overall tardiness of merged operations with respect to internal deadline of the corresponding jobs: Minimize ∑︁ 𝑠∈𝑆©  «∑︁ 𝑚∈𝑀 ∑︁ 𝑗∈𝐽 ∑︁ 𝜏∈𝑇,𝜏 ≥𝑡 𝐽 𝑑(𝑗) 𝑘𝑠,𝑚,𝑗,𝜏 ª ® ¬(21) 4 RESULTS To evaluate our approach, we generated a new benchmark set for the OSSP-OB following the ideas of E.Taillard[6]. Instances are defined by the number of jobs, machines, and categories. We considered jobs of sizes 3,5,7, and 10, and for each case the number of machines and categories was set to either 2or 3. All other parameters, such as processing times, time horizon, and number of operations per job, were kept constant. The construction procedure is flexible and can be extended to introduce randomness or additional parameters if needed. Each valid merge corresponds to a subset of operations of the same category, meaning that the set of possible merges is a subset of the power set P(𝑂) . The size of this set is strongly influenced by the number of available category capacities: smaller values lead to larger collections of potential merges. This effect directly impacts solver memory usage and thus scalability. Computational experiments were carried out using the Gurobi 11 . 0 . 0solver on an AMD Ryzen 7 5800 𝐻 3 . 2GHz CPU with 16 GB RAM, with a time limit of 60 minutes per instance. Table 1 demonstrates the maximum runtime observed for each instance configuration. All instances of this set were solved optimally. Overall, the model is able to solve instances up to 10 jobs in a reasonable time, but the growth of the merged operation space remains the main computational bottleneck. For example, increasing the number of jobs to 15 leads to out-of-memory errors. 5 CONCLUSION In this work, we introduced a complete MILP formulation for OSSP-OB. Our experiments show that the model can solve small instances to optimality within a reasonable time, but quickly becomes impractical as the number of jobs and possible merges increases. Rapid growth of the set of potential merged operations leads to excessive memory usage and limits scalability. N jobs N machines N categories Avg. Runtime (s) 3 2 2 1.11 3 2 3 0.56 3 3 2 1.50 3 3 3 0.67 5 2 2 6.90 5 2 3 2.55 5 3 2 6.36 5 3 3 3.09 7 2 2 16.81 7 2 3 7.54 7 3 2 20.33 7 3 3 10.70 10 2 2 124.92 10 2 3 40.62 10 3 2 156.67 10 3 3 47.94 Table 1: Average runtimes (in seconds) for OSSP-OB instances. Future research will therefore focus on approximation methods. A potential approach to consider is separating the OSSP-OB environment into two subproblems, merging of operations and scheduling of the merged operations, combining exact methods for the two parts separately. Another approach is to design dedicated (meta-)heuristics specific to the merging structure. Such approaches are expected to extend the range of solvable instances and provide practical solutions for larger problem sizes. ACKNOWLEDGEMENTS The research was supported by the BioLOG project: Balázs Dávid and Miklós Krész are grateful for the support of National Center of Science (NCN) through grant DEC-2020/39/I/HS4/03533, the Slovenian Research and Innovation Agency (ARIS) through grant N1-0223 and the Austrian Science Fund (FWF) through grant I 5443-N. This work was partially supported by the Slovenian Research Agency, research program P1-0404 and by the research program CogniCom (0013103) at the University of Primorska. Balázs Dávid is grateful for the support of ARIS through grant J1-50000, and gratefully acknowledges the Slovenian Research and Innovation Agency (ARIS) and the Ministry of the Economy, Tourism and Sport (MGTŠ) for the grant V4-2512. REFERENCES [1] Danyu Bai, Zhi-Hai Zhang, and Qiang Zhang. 2016. Flexible open shop scheduling problem to minimize makespan. Computers & operations research, 67, 207–215. [2] John W Fowler and Lars Mönch. 2022. A survey of scheduling with parallel batch (p-batch) processing. European journal of operational research, 298, 1, 1–24. [3] Wieslaw Kubiak. 2022. Book of Open Shop Scheduling. Springer. [4] B Naderi, SMT Fatemi Ghomi, Majid Aminnayeri, and Mostafa Zandieh. 2011. Scheduling open shops with parallel machines to minimize total completion time. Journal of Computational and Applied Mathematics, 235, 5, 1275–1287. [5] Chris N Potts and Mikhail Y Kovalyov. 2000. Scheduling with batching: a review. European journal of operational research, 120, 2, 228–249. [6] Eric Taillard. 1993. Benchmarks for basic scheduling problems. European journal of operational research, 64, 2, 278–285. [7] Tanya Y Tang and J Christopher Beck. 2020. CP and hybrid models for twostage batching and scheduling. In Integration of Constraint Programming, Artificial Intelligence, and Operations Research: 17th International Conference, CPAIOR 2020, Vienna, Austria, September 21–24, 2020, Proceedings 17. Springer, 431–446. 10 Flexibility vs. efficiency: a study in sawmill scheduling Csaba Kebelei Eötvös Loránd University Faculty of Informatics Budapest, Hungary [email protected] Máté Hegyháti University of Sopron Sopron, Hungary University of Pannonia Veszprém, Hungary Abstract Sawmills are an important stage in the primary wood industry, producing timber and other side products from harvested logs. Increasing efficiency in a sawmill is not only an economic desire, but it also propagates to the competitiveness of using more renewable resources in dependent industries such as construction, packaging, furniture. Although the road from a tree log to a shipped product involves many stages, the key processing step is the sawing performed by high-value sawing machines. As the operation of these machines plays a key role in overall efficiency, several research papers have addressed their scheduling. Naturally, allowing more freedom in the production plan may result in better financial results in exchange for increased computational needs of the optimizer. However, more complex schedules may also pose an additional burden in their real-life execution. This paper presents several variants of a model formerly proposed by the authors and investigates this trade-off relationship empirically. Keywords sawmill, scheduling, MILP 1 Introduction and literature The increasing utilization of renewable resources is a major directive of many nations and companies around the world. Wood is a natural resource that not only serves as a great carbon sequestration tool, but it is a suitable raw material for many industries, including construction and interior design, paper, packaging, and the boat industry. After logging, sawmills are generally the first stage in timber processing regardless of the end-use industry. Their main task is to turn unprocessed logs from the forest into well-cut lumber. While this is a multi-stage process, the main step is the sawing itself that can be done by two of the main machine types: band saws and frame saws. The efficient operation of these machines have a significant impact on the whole plant, thus, several works have addressed this issue. Efficiency is tackled on both the operational and the planning level by designing cutting patterns and scheduling, respectively [ 6 ]. A tree log may be cut in different patterns yielding different quantities of different products, as shown by in Figure 1. Designing patterns based on long-term demands for products, statistics about defects, etc. is a well-researched topic in the literature. Moreover, scanning tools available on modern sawing machines can provide real-time information about logs to make such decisions adaptive [ 5 ]. On the planning level, selecting the cutting patterns and the corresponding log quantities for each shift while considering orders, available logs, storage capacity, etc. is a MATCOS-25, Koper, Slovenia 2025. scheduling problem. The base problem was introduced by Zanjani et al. [ 7 ] and Maturana et al. [ 9 ] for the stochastic and deterministic cases, respectively. Subsequent works extended the approach in various directions, e.g., addressing uncertainties [ 2 , 11 ], integrating cutting pattern selections [3, 10], or minimizing waste [4]. Figure 1: Cutting patterns for test cases in Section 5 In this paper, a previously proposed model by the authors [ 8 ] is further investigated, that addressed issues prevalent for small-scale sawmills: workforce availability and operation differences between sawing technologies. Compared to band saws, the changeover time for switching cutting patterns on a frame saw is non-negligible, thus the same formulation is not applicable in short-term scheduling. The proposed model introduced a significantly different formulation for the behavior of the frame saw, that only allows one change per 11 MATCOS-25, October 09–10, 2025, Koper, Slovenia Kebelei and Hegyháti shift following industry practice. In later Sections this base model will be referred to as C–EK , where EK (expert knowledge) indicates at most one change in a shift, that may be timed continuously ( C ) at any time. The aim of this work is to investigate, whether allowing more cutting pattern changes on a frame saw would yield significantly better results due to a broader solution space, and/or does it increase the computational costs comparably. While allowing such flexibility in a mathematical model may be simple, one has to keep in mind, that the execution of overly complicated schedules in an non-automated small-scale plant is not realistic, thus a set of models will also be introduced, where cutting pattern changes may only happen at predefined discrete ( D ) time points, as detailed in Section 4. Another decision freedom, whose effect will be investigated in Section 5 is the possible reallocation of the workforce from one machine to another during a shift, as discussed in Section 3. 2 Problem definition The objective is the same as in [ 8 ]: minimizing the under-production cost in a small-scale sawmill equipped with one frame saw and one band saw. The sawmill may operate in multiple shifts, however, from the modeling point of view, only the total number of shifts is relevant. Thus, without the loss of generality, it is assumed, that each day has a single shift (of length 𝐻 ), and the two terms are used interchangeably. For each day ( 𝑑∈ D ), the number of specialists ( 𝐻𝑅𝑆𝑃 𝑑 ) and additional workers ( 𝐻𝑅𝐴𝑊 𝑑 ) are given, along with the requirements to operate each machine ( 𝑅𝑅𝑆𝑃,𝐹 , 𝑅𝑅𝑆𝑃,𝐵 , 𝑅𝑅𝐴𝑊 ,𝐹 , 𝑅𝑅𝐴𝑊 ,𝐵 ). Resource related data is also given, i.e., the quantity ( 𝐼𝑙 ) and volume ( 𝑉𝑙 ) of logs of different sizes ( 𝑙∈ L ) all available at the start of the planning horizon. For each log size several cutting patterns ( P𝑙⊆ P ) may be available with known yields ( 𝑌𝑡,𝑝 ) for lumber products ( 𝑡∈ T ). For each shift the production planner may decide which machines to operate and how many logs with which cutting patterns are processed. By allowing different flexibilities, several different problem definitions will be addressed by their corresponding models: C–EK Original problem from [ 8 ]: the frame saw may change cutting pattern once per shift, but a saw is either operating through the whole shift or not. C′–EK Same as C–EK , but the frame saw may also be turned on/off, and the workforce reallocated to the band saw instead of changing the pattern, once per shift. D[𝐾] The shift is subdivided into 𝑘 segments of equal length. Pattern changes and workforce reallocations are allowed only at the end of these intervals. D[𝐾]–EK Same as D[𝐾] , but the frame saw is allowed to change cutting patterns at most once a day. If 𝑂𝑃𝑇 (𝑀) denotes the optimal (minimal) solution for the problem/model 𝑀, the following inequalities will hold naturally: •𝑂𝑃𝑇 (C′–EK) ≤ 𝑂𝑃𝑇 (C–EK) •𝑂𝑃𝑇 (C′–EK) ≤ 𝑂𝑃𝑇 (D[𝑋]–EK) •𝑂𝑃𝑇 (D1)=𝑂𝑃𝑇 (D1–EK) •𝑂𝑃𝑇 (D[𝑋]) ≤ 𝑂𝑃𝑇 (D[𝑋]–EK) •𝑂𝑃𝑇 (D[𝑌]) ≤ 𝑂𝑃𝑇 (D[𝑋])if 𝑋|𝑌 •𝑂𝑃𝑇 (D[𝑌]–EK) ≤ 𝑂𝑃𝑇 (D[𝑋]–EK)if 𝑋|𝑌 3 Model with workforce reallocation: C’–EK The event, when worker reallocation is also possible in addition to pattern changes on the frame saw, will be referred to as the changeover. As this model is based on C–EK , constraints (1)–(5), (7)–(8), and (10) from [ 8 ] are copied verbatim, expressing the objective, storage balances, daily production quantities, inventory limits, frame saw production bounds before the changeover, and frame saw activation bounds. Binary variables 𝑤𝐹 𝑑 and 𝑤𝐵 𝑑 are split to 𝑤𝐹,− 𝑑 , 𝑤𝐵,− 𝑑 and 𝑤𝐹,+ 𝑑 , 𝑤𝐵,+ 𝑑 , indicating whether the machines are active before or after the changeover. Duplicated versions of constraints (15)–(16) with these variables are also copied to express human resource needs. Superscripts + and − will indicate such division similarly in other variables, and ±will be used to indicate both cases. The band saw time capacity constraint is reformulated as: ∑︁ 𝑝∈ P 𝑆𝑇𝐵 𝑝·𝑞𝐵 𝑑,𝑝 ≤𝜏𝐵,− 𝑑+𝜏𝐵,+ 𝑑∀𝑑∈ D (1) Where variables 𝜏𝐵,± 𝑑∈ [ 0 , 𝐻] represent the available band saw cutting time, linked to changeover timing by the following constraints: 𝜏𝐵,± 𝑑≤𝐻·𝑤𝐵,± 𝑑∀𝑑∈ D,± ∈ {−,+} (2) 𝜏𝐵,− 𝑑≤𝑡𝑑∀𝑑∈ D (3) 𝜏𝐵,+ 𝑑≤𝐻−𝑡𝑑∀𝑑∈ D (4) Constraint (2) activates 𝜏𝐵,± 𝑑 only when the band saw is active in the corresponding part of the shift, while constraints (3) and (4) bound them by 𝑡𝑑and 𝐻−𝑡𝑑, respectively. For the frame saw, the following constraint limits the production after the changeover: ∑︁ 𝑝∈ P 𝑆𝑇 𝐹 𝑝·𝑞𝐹,+ 𝑑,𝑝 ≤(𝐻−𝑡𝑑)−𝐶𝑇 𝐹·𝑧𝐹,𝐶𝑇 𝑑+𝐻·1−𝑤𝐹,+ 𝑑 ∀𝑑∈ D (5) Where 𝑧𝐹,𝐶𝑇 𝑑∈ { 0 , 1 } indicates whether changeover time must be deducted after the changeover. Constraint (6) sets 𝑧𝐹,𝐶𝑇 𝑑 to one, if the frame saw is active through the whole shift and the pattern is changed. 𝑧𝐹,𝐶𝑇 𝑑≥𝑤𝐹,− 𝑑+𝑤𝐹,+ 𝑑+𝑠𝐹,+ 𝑑,𝑝 −𝑠𝐹,− 𝑑,𝑝 −2∀𝑑∈ D, 𝑝 ∈ P (6) Constraint (7) ensures that only an active frame saw can have an active cutting pattern, which is preserved by constraints (8) and (9) for the next day if the saw remains active. ∑︁ 𝑝∈ P 𝑠𝐹,± 𝑑,𝑝 =𝑤𝐹,± 𝑑∀𝑑∈ D,± ∈ {−,+} (7) 𝑠𝐹,+ 𝑑,𝑝 ≥𝑠𝐹,− 𝑑+1,𝑝 −1·2−𝑤𝐹,+ 𝑑−𝑤𝐹,− 𝑑+1∀𝑑∈ D, 𝑝 ∈ P (8) 𝑠𝐹,+ 𝑑,𝑝 ≤𝑠𝐹,− 𝑑+1,𝑝 +1·2−𝑤𝐹,+ 𝑑−𝑤𝐹,− 𝑑+1∀𝑑∈ D, 𝑝 ∈ P (9) 12 Flexibility vs. efficiency: a study in sawmill scheduling MATCOS-25, October 09–10, 2025, Koper, Slovenia 4 Discrete time-slot models: D[𝐾],D[𝐾]–EK Similarly to the previous modification of the original model, some constraints from [ 8 ] remain unmodified, expressing the objective, storage balances: (1)–(3). Moreover, some variables are now defined for each segment of a shift, indexed by the set K={1,2, . . . , 𝐾}. For example, the binary variables 𝑤𝐹 𝑑,𝑘, 𝑤𝐵 𝑑,𝑘 indicate machine usage, which are used in segment-wise copies of (15)–(16) of [ 8 ] to express human-resource requirements. 𝑤𝐵 𝑑,𝑘 is also used in the band saw time capacity constraint, where 𝐻𝐾=𝐻/𝐾denotes the length of a segment: ∑︁ 𝑝∈ P 𝑆𝑇𝐵 𝑝·𝑞𝐵 𝑑,𝑝 ≤∑︁ 𝑘∈K 𝐻𝐾·𝑤𝐵 𝑑,𝑘 ∀𝑑∈ D (10) Another such variable is 𝑞𝐹 𝑑,𝑘,𝑝 ∈Z≥0 indicating the frame saw production, used in constraints (11) and (12) to calculate the daily production quantity and limit the inventory: 𝑦𝑑,𝑡 =∑︁ 𝑝∈ P 𝑌𝑡,𝑝 ·𝑉𝑙𝑝· 𝑞𝐵 𝑑,𝑝 +∑︁ 𝑘∈K 𝑞𝐹 𝑑,𝑘,𝑝 !∀𝑑∈ D, 𝑡 ∈ T (11) ∑︁ 𝑑∈ D ∑︁ 𝑝∈ P𝑙 𝑞𝐵 𝑑,𝑝 +∑︁ 𝑘∈K 𝑞𝐹 𝑑,𝑘,𝑝 !≤𝐼𝑙∀𝑙∈ L (12) Active consecutive segments on the frame saw should be merged together if the pattern is unchanged. These conditions are expressed by binary variable 𝑠𝐹 𝑑,𝑘,𝑝 which is 1 iff pattern 𝑝 is active in segment 𝑑,𝑘 and the one preceding it. Another binary indicator variable, 𝑠𝑐,𝐹 𝑑,𝑘,𝑝 is 1 iff 𝑝 is active in segment 𝑘 but not in the preceding one. This logic is enforced by the following constraints: ∑︁ 𝑝∈ P 𝑠𝑐,𝐹 𝑑,𝑘,𝑝 +𝑠𝐹 𝑑,𝑘,𝑝 =𝑤𝐹 𝑑,𝑘 ∀𝑑∈ D, 𝑘 ∈ K (13) 𝑠𝐹 𝑑,𝑘,𝑝 ≤𝑠𝐹 𝑑,𝑘−1,𝑝 +𝑠𝑐,𝐹 𝑑,𝑘−1,𝑝 +1−𝑤𝐹 𝑑,𝑘−1 ∀𝑑∈ D, 𝑘 ∈ K \ {1}, 𝑝 ∈ P (14) 𝑠𝐹 𝑑,1,𝑝 ≤𝑠𝐹 𝑑−1,𝐾,𝑝 +𝑠𝑐,𝐹 𝑑−1,𝐾,𝑝 +1−𝑤𝐹 𝑑−1,𝐾  ∀𝑑∈ D \ {1}, 𝑝 ∈ P (15) Constraint (13) ensures that when the frame saw is operated in segment 𝑘 , exactly one pattern is selected (either a changeover to 𝑝 or continuation with 𝑝 ), and when it is idle, neither indicator is active. Constraint (14) allows pattern continuation only if 𝑝 was active in the previous segment, and the frame saw is active. The same condition is set across days is by Constraint (15). The merging of active segments with the same pattern is modeled by constraints (16) , (17) , and (18) , introducing a spare time variable 𝜎𝑑,𝑘,𝑝 ∈ [ 0 , 𝐻𝐾] that records unused time in segment 𝑘 for pattern 𝑝and can be carried to the next segment. 𝜎𝑑,𝑘,𝑝 ≤𝑆𝑇𝐹 𝑝·𝑠𝐹 𝑑,𝑘+1,𝑝 ∀𝑑∈ D, 𝑘 ∈ K \ {𝐾}, 𝑝 ∈ P (16) 𝑆𝑇 𝐹 𝑝·𝑞𝐹 𝑑,1,𝑝 ≤𝐻𝐾−𝐶𝑇 𝐹·𝑠𝑐,𝐹 𝑑,1,𝑝 +𝐻𝐾·𝑠𝐹 𝑑,1,𝑝 −𝜎𝑑,1,𝑝 ∀𝑑∈ D, 𝑝 ∈ P (17) 𝑆𝑇 𝐹 𝑝·𝑞𝐹 𝑑,𝑘,𝑝 ≤𝐻𝐾−𝐶𝑇 𝐹·𝑠𝑐,𝐹 𝑑,𝑘,𝑝 +𝐻𝐾·𝑠𝐹 𝑑,𝑘,𝑝 +𝜎𝑑,𝑘−1,𝑝 −𝜎𝑑,𝑘,𝑝 ∀𝑑∈ D, 𝑘 ∈ K \ {1}, 𝑝 ∈ P (18) Constraint (16) limits the carryover from segment 𝑘 to one log’s sawing time 𝑆𝑇 𝐹 𝑝 and activates it only if the next segment continues with the same pattern, while the production time bounds follow from Eqs. (17) and (18) , accounting for changeover time 𝐶𝑇 𝐹 and the inflow/outflow of spare time. To obtain D[𝐾]–EK from D[𝐾] , Constraint (19) needs to be added, which enforces at most one cutting pattern change by the frame saw per day: ∑︁ 𝑘∈K ∑︁ 𝑝∈ P 𝑠𝑐,𝐹 𝑑,𝑘,𝑝 ≤1∀𝑑∈ D (19) 5 Empirical results To evaluate the new approaches, the randomly generated test set from our previous study [ 8 ] is used, where cutting pattern yields were precomputed with Pitago Optimizers [ 1 ] and technical parameters reflect industry practice and recommendations by experts. Figure 1 shows the five patterns defined for three log diameter classes in all of the test cases. For practical reasons in execution, the values 1, 2, 4, and 8 were considered for 𝐾 , and the EK variant for 𝐾= 1is omitted, as it is equivalent to the non-EK case. All problems in the dataset were solved with all of the models by Gurobi Optimizer 12.0.1 on a computer with an Apple M2 CPU and 16 GB of RAM available, with a time limit of 1500 seconds. Table 1 shows the aggregated results from the 50 cases with 4-week planning horizon. The left block shows objective value distributions (m 3 ) as box plots. The median relative deviation from the C–EK reference is also reported in the column labeled ˜ ΔC–EK (%). On the right, CPU times (s) are shown as box plots on a base-10 logarithmic scale. The results match the theory: •Adding on/off reallocation in C′–EK increases quality • Discretizing the changeover timing degrades it when only one cutting pattern change is allowed per shift ( EK models), while smaller slot lengths mitigate the loss. • Imposing the “at most one change” rule on the frame saw (–EK) has a negative impact on solution quality. It can be observed, however, that the differences in objective values across models are small: ˜ ΔC–EK ranges from − 0 . 1072% for C′–EK (best) to + 0 . 8536% for D1 (worst), and difference between D[𝐾]and D[𝐾]–EK is less than 0.05 percentage points. A key takeaway is that even D8 does not outperform C–EK . For these test cases, continuously timing a single change is more valuable than permitting multiple changes at fixed time-slot boundaries. Even so, the loss from discretization is small: as 𝑘 increases, the median gap shrinks from + 0 . 85% ( D1 ) to about + 0 . 0058% for D8–EK or +0.0024% for D8. Regarding runtimes, increasing 𝐾 for the discrete models increases CPU time, as expected, yet the 𝐾= 8variants still run faster 13 MATCOS-25, October 09–10, 2025, Koper, Slovenia Kebelei and Hegyháti Table 1: Results on randomly generated instances with a 4-week planning horizon across model variants. Objective value (m3) CPU time (s) Model 950 1025 1100 1175 1250 ˜ ΔC—EK (%) 10−210−0.625 100.75 102.125 103.5Avg. C–EK +0.0000 151.36 C′–EK -0.1072 125.98 D1 +0.8536 0.24 D2–EK +0.2908 1.37 D2 +0.2440 1.52 D4–EK +0.1065 5.00 D4 +0.0760 4.21 D8–EK +0.0058 67.76 D8 +0.0024 82.82 on average than the continuous baselines. Notably, D8 is within ≈ 0 . 11 percentage points of C′–EK in objective value yet is faster on average. The –EK restriction leaves quality essentially unchanged and tends to reduce runtime. Time-limit hits were infrequent (Table 2). Table 2: Time-limit hits (counts). Model 3-week 4-week 5-week 6-week C–EK 3254 C′–EK 0123 D8–EK 0110 D8 0111 These counts suggest that the discrete-slot models remain computationally stable across longer planning horizons. Overall, the discrete time-slot-based models offer a tunable tradeoff: coarse slotting is extremely fast but less accurate, whereas 𝐾= 8 achieves near-continuous quality at a substantially lower average runtime than the continuous baselines. 6 Concluding remarks We revisited a MILP scheduling model for small-scale sawmills and examined three design choices: (i) allowing workforce reallocation, (ii) restricting changeovers to discrete within-shift time-slot boundaries, and (iii) allowing multiple changes per shift. Computational tests on a large synthetic instance set show that objective differences across variants are modest. Allowing on/off operation with reallocation provides a consistent, albeit small, improvement in solution quality, whereas discretizing the changeover into time slots introduces a controllable loss that diminishes as the slot granularity is refined. Overall, the discrete formulations offer a tunable tradeoff between accuracy and speed. In our setting, even coarse slots perform well, while a moderate refinement (e.g., hourly segments) achieves near-continuous quality at substantially lower average runtimes. In practice, multiple within-shift changes bring marginal gains and complicate execution, whereas –EK yields simpler plans with similar quality and often lower runtime. For the investigated test cases, even completely forbidding changes within shifts resulted in less than 1 percent quality reduction. These findings support the discrete-slot approach as a practical planning tool for small, low-automation sawmills. Future work will expand the set of test instances, include an industrial case study to validate and calibrate the models on operational data, and, where appropriate, incorporate uncertainty in key inputs. References [1] Pitago Optimizers. Retrieved August 24, 2025 from https://pitago.eu/ [2] Pamela P Alvarez and Jorge R Vera. 2014. Application of robust optimization to the sawmill planning problem. Annals of Operations Research 219 (2014), 457–475. doi:10.1007/s10479-011-1002-4 [3] Diego Broz, Nicolás Vanzetti, Gabriela Corsano, and Jorge M. Montagna. 2019. Goal programming application for the decision support in the daily production planning of sawmills. Forest Policy and Economics 102 (2019), 29–40. doi:10.1016/ j.forpol.2019.02.004 [4] Giacomo Da Col, Philipp Fleiss, Alice Tarzariol, Erich C. Teppan, and Elena Wiegelmann. 2025. A Declarative Approach to Tackle Sawmill Production Scheduling with Answer Set Programming. In Computational Science and Computational Intelligence, Hamid R. Arabnia, Leonidas Deligiannidis, Farid Ghareh Mohammadi, Soheyla Amirian, and Farzan Shenavarmasouleh (Eds.). Springer Nature Switzerland, Cham, 313–323. doi:10.1007/978-3-031-90341-0_23 [5] Kamran Forghani, Mats Carlsson, Pierre Flener, Magnus Fredriksson, Justin Pearson, and Di Yuan. 2024. Maximizing value yield in wood industry through flexible sawing and product grading based on wane and log shape. Computers and Electronics in Agriculture 216 (2024), 108513. doi:10.1016/j.compag.2023.108513 [6] Seyed Mohsen Hosseini and Angelika Peer. 2022. Wood Products Manufacturing Optimization: A Survey. IEEE Access 10 (2022), 121653–121683. doi:10.1109/ ACCESS.2022.3223053 [7] Masoumeh Kazemi Zanjani, Daoud Ait-Kadi, and Mustapha Nourelfath. 2010. Robust production planning in a manufacturing environment with random yield: A case in sawmill production planning. European Journal of Operational Research 201, 3 (2010), 882–891. doi:10.1016/j.ejor.2009.03.041 [8] Csaba Kebelei and Mate Hegyhati. 2024. Sawmill scheduling: an applicationoriented model. Acta Technica Jaurinensis 17, 3 (2024), 104–110. doi:10.14513/ actatechjaur.00743 [9] Sergio Maturana, Enzo Pizani, and Jorge Vera. 2010. Scheduling production for a sawmill: A comparison of a mathematical model versus a heuristic. Computers & Industrial Engineering 59, 4 (2010), 667–674. doi:10.1016/j.cie.2010.07.016 [10] Nicolás Vanzetti, Diego Broz, Gabriela Corsano, and Jorge M. Montagna. 2018. An optimization approach for multiperiod production planning in a sawmill. Forest Policy and Economics 97 (2018), 1–8. doi:10.1016/j.forpol.2018.09.001 [11] Mauricio Varas, Sergio Maturana, Rodrigo Pascual, Ignacio Vargas, and Jorge Vera. 2014. Scheduling production for a sawmill: A robust optimization approach. International Journal of Production Economics 150 (2014), 37–51. doi:10.1016/j. ijpe.2013.11.028 14 Implementation of a Vehicle and Driver Scheduling Model: a Case Study Viktor Árgilán University of Szeged, Juhász Gyula Faculty of Education, Department of Applied Informatics Szeged, Hungary viktor.sandor[email protected] József Békési University of Szeged, Institute of Informatics, Department of Foundations of Computer Science Szeged, Hungary [email protected]szeged.hu Gábor Galambos University of Szeged, Juhász Gyula Faculty of Education, Department of Applied Informatics Szeged, Hungary GalambosGab[email protected] Imre Papp University of Szeged, Juhász Gyula Faculty of Education, Department of Applied Informatics Szeged, Hungary Abstract During combined vehicle and driver scheduling, we have to plan the daily work of vehicles and their drivers so that they perform a set of tasks at the lowest possible cost. The tasks are defined by time intervals, and the vehicles are located in different depots. In recent decades, several mathematical models have been defined, with which we can create regular schedules and search for optimal solutions based on different objective functions. However, in practical problems, there are many requirements that cannot be handled easily and usually have a significant computational demand. The results of our research and development project are presented through a case study based on real experiments carried out at the Budapest Transport Corporation. Keywords Optimization, Vehicle and driver scheduling problem, Public transport 1 Introduction Operating costs are usually a significant item in the budget of public transport providers. The main components of these costs are vehicle fleet acquisition costs, fuel and maintenance costs, and driver wages. With the help of various decision support systems, comprehensive solutions for both the vehicle and the driver have been developed in recent decades to solve the optimization task. In public transport, vehicle and driver scheduling can be very complex. In theory, we generally look for a global optimum that minimizes both vehicle-related costs and driver scheduling costs. These two types of costs affect each other, so it’s usually best to handle the tasks together [2]. If we want to find the optimal solution, combined vehicle and driver scheduling mathematical optimization models can be used. There are several such methods in the literature. The vehicle scheduling problem is usually formulated as a multicommodity network flow problem ([ 4 ], [ 10 ], [ 12 ]). The optimal schedule can be calculated as the solution of an integer programming problem. Other models are also known, for example the problem can be formulated as a set partitioning problem (see e.g. [14], [7]). MATCOS-25, Koper, Slovenia 2025. https://doi.org/10.1145/nnnnnnn.nnnnnnn The Generate and Select (GaS) method is the most wellknown technology for the driver scheduling part. In the initial phase, a substantial number of standard shifts are generated. In the subsequent selection phase, a subset of these regular shifts is selected to minimize cost and optimize coverage of trips. It is noteworthy that the execution of both phases requires substantial computational resources. The extent of this computational demand is contingent upon the number of trips and the intricacy of the operational guidelines. The selection phase can be modeled as a set covering or set partitioning problem. In 2005, Huisman et al. [ 9 ] extended the former combined models and algorithms of the single-depot case [ 5 , 6 ] to the multi-depot version. This was the first general mathematical formulation of the combined multi-depot problem. Later many authors investigated this version of the problem (see, for example, [8], [13], [15], [16]). In practice, however, it turns out that there are many company-specific details and constraints that cannot be uniformly addressed by general systems, but which are important for the transport companies. As an example, if a transport company also uses alternative fuel vehicles in its fleet, their scheduling must take into account the number of kilometers per refueling, known as the radius, which can be much lower than the mileage of a conventional fuel vehicle. Such cases have been examined for example in [ 1 ] and [11]. In the paper [ 3 ] Békési and Nagy presented how the methods used in the above mentioned papers were adapted to develop a decision support system for the Budapest Transport Corporation. The aim of this project was to automatically calculate optimal or approximately optimal vehicle and driver schedules for a given list of trips based on the master data and the company specific requirements and parameters in compliance with labor regulations. This paper overviews how the complete integration was implemented and what kind of specific developments were necessary to take into account all the practical requirements of the company. 2 The Automatic Solving Process We summarize the system’s key characteristics, requirements, input data, and settings based on [ 3 ]. All the data required for a computation is kept in packages so that it 15 B. Hren and T. Dobravec Figure 2: Comparison of the complexities of the Edmonds-Karp and Dinic algorithms on (a) DIRECTED_SCALE_FREE graphs (upper chart) and (b) LAYERED graphs (lower chart). DGraph < Integer , Integer > graph = input . graph ; int sour ce = input . source ; int sink = input . sink ; int maxFlow = edmondsKarp ( graph , source , sink ); Output output = new Output ( maxFlow ); return output; } } Testsets. In the project, several test sets were implemented. The first test set contains directed scale-free graphs generated by a Type1 test set generator. The test cases were created using the following for-loop–based description: $for {i ,1 ,30 ,1}: Type1 ::0: $ {100* i -1}: $ {100* i }: $ {500* i} This generates 30 Type1 test cases with parameters 𝑛= 100 ∗𝑖 , 𝑚= 500 ∗𝑖 , source = 0, and sink = 100 𝑖− 1, for 𝑖= 1 , 2 , . . . , 30. Running the algorithms on this test set yields the results shown in Figure 2 (upper chart). For this configuration, the complexity of both algorithms increases with the number of nodes. On average, the Edmonds–Karp algorithm performs about 33% faster than Dinic’s algorithm on this test set. Further tests revealed that the advantage of the Edmonds–Karp algorithm on directed scale-free graphs is greater when edges are evenly distibuted. As the disparity between node degrees increases, this advantage diminishes, and in some cases Dinic’s algorithm even outperforms Edmonds–Karp algorithm. Another test set used in this project consists of layered graphs generated by a Type4 test set generator. The test cases were defined using the following for-loop–based description: $for {i ,1 ,30 ,1}: Type4 ::0: $ {( i *3 -2)*( i +5)+2 -1}: ${ i *3}: $ {i +5} This generates layered graphs with 3 ∗𝑖 layers and 𝑖+ 5nodes per layer, for 𝑖= 1 , 2 , . . . , 30. Running both algorithms on this test set shows that, for these layered graphs, Dinic’s algorithm is faster than Edmonds–Karp algorithm (see Figure 2, bottom chart). The average speedup in this scenario is 2.9x. 5 Conclusions The ALGatorGraph Java library has been developed as a supporting tool for ALGator projects that address graph-related problems. Its main benefits are twofold: it offers a unified interface for implementing graph data structures, and it provides the capability to generate a wide variety of graphs—both through parameter-based generation and by reading graphs from the library. Through the MaxFlow project we have demonstrated how straightforward it is to implement graph-related problems in ALGator using ALGatorGraph. The results obtained in the experiments highlight some characteristics of the two implemented algorithms, but their specific values are not the main point. The real significance lies in how easily such results can be obtained by combining ALGatorGraph’s ability to handle graph data structures with ALGator’s functionality for conducting and analyzing experiments. References [1] S. Arora and B. Barak. 2009. Computational Complexity: A Modern Approach. Cambridge University Press. [2] Tomaž Dobravec. 2025. ALGator: Online Implementation of the System for Automated Algorithm Execution and Evaluation. (2025). https://algator.fri.unilj.si [3] Catherine C. McGeoch. 2012. A Guide to Experimental Algorithmics. Cambridge University Press. [4] Dimitrios Michail, Joris Kinable, Barak Naveh, and John V. Sichi. 2020. JGraphT–A Java Library for Graph Data Structures and Algorithms. ACM Trans. Math. Softw. 46, 2, Article 16 (May 2020), 29 pages. [5] Matthias Müller-Hannemann. 2010. Algorithm Engineering, Bridging the Gap Between Algorithm Theory and Practice. Springer-Verlag Berlin and Heidelberg GmbH & Co. K. [6] Ruiheng Zhang. 2024. Network Flow of Graph Theory and Its Application. In Proceedings of the 1st International Conference on Engineering Management, Information Technology and Intelligence - EMITI. INSTICC, SciTePress, 692–696. https://doi.org/10.5220/0012968800004508 22 Engineering CSFLOC: A Subsumption-Driven Clause-Counting SAT Solver Gábor Kusper [email protected] Eszterházy Károly Catholic University Eger, Hungary Abstract Clause counting, a technique from the #SAT domain, offers a complementary angle to conflict-driven clause learning (CDCL). We revisit CSFLOC (Counting Subsumed Full-Length Ordered Clauses) and present an engineered implementation that turns its theoretical "last-1-bit" observation into practical speedups. The solver represents full-length clauses with a binary counter and performs sign-aware bucket scans keyed by the last-literal index, allowing safe jumps that skip a block of consecutive fulllength clauses subsumed by input clauses. We map the known pseudocode of CSFLOC line-by-line to Java, detailing the data structures that make the inner loop fast: POS/NEG buckets per index, small effected/learned caches, and carry-like counter updates that implement jumps of size 2 𝑛−𝑗 without big integers. Preordering (variable renaming) strategies—composable flags B,C,H,I,R,S,W —shape the distribution of last-literal indices and, via an island/strait view of the counter, bias the solver toward longer trailing 1-blocks and shorter 0-gaps. A concise empirical snapshot shows where CSFLOC is competitive (over-constrained, dense-structure instances) and where CDCL remains preferable. Source code is available at http://fmv.ektf.hu/tools.html. Keywords CSFLOC, SAT solving, clause counting 1 Introduction The propositional satisfiability problem (SAT) asks whether a propositional CNF formula has an assignment of truth values to its variables that makes the formula true. SAT is one of the mostresearched NP-complete [ 7 ] problems in computer science, with applications ranging from theoretical computer science and artificial intelligence to hardware design and formal verification [ 4 ]. Modern SAT solvers build on conflict-driven clause learning (CDCL), an extension of the classical Davis Putnam Logemann Loveland (DPLL) procedure [ 9 ] that adds conflict analysis, clause learning, and non-chronological backtracking (backjumping) [ 4 ]. An interesting question is how many models a SAT instance has; this is the #SAT problem [ 11 , 15 , 3 ], i.e., counting the satisfying assignments of a CNF formula [10]. CSFLOC (Counting Subsumed Full-Length Ordered Clauses) is a clause-counting approach that enumerates full-length clauses via a binary counter and uses subsumption to jump over blocks. It was introduced in [ 14 , 13 ]. Although CSFLOC is a general SAT solver, it is inspired by classical #SAT techniques such as full-length clause counting. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). MATCOS-25, 9–10 October 2024, Koper,Slovenia ©2025 Copyright held by the owner/author(s). This paper briefly recalls CSFLOC and focuses on an optimized Java implementation. 1.1 Related work Modern SAT solving is dominated by conflict-driven clause learning (CDCL) with powerful preprocessing and restart heuristics; see the updated Handbook of Satisfiability for a broad survey of techniques and applications [ 4 ]. Alongside CDCL, a classical but distinct line is clause counting and inclusion–exclusion–based reasoning, which reasons about sets of full-length (maximal) clauses and their subsumption structure. Early work by Iwama established counting-style satisfiability tests with average-case guarantees [ 11 ]. Lozinskii developed exact propositional model counting (the so-called #SAT problem) in [ 15 ], while Birnbaum–Lozinskii connected counting tightly to Davis–Putnam–style branching [ 2 ]. A complementary strand is Andrei’s “inverting resolution,” which frames satisfiability counting via inverse propositional resolution and normalization [ 1 ]. These works already emphasize that ordering, last-literal positions, and subsumption can unlock substantial contiguous “jumps” in the enumeration space. The inclusion–exclusion principle has been repeatedly explored for model counting and even SAT itself. Bennett and Sankaranarayanan proposed an inclusion–exclusion counter with subsumption pruning for 𝑘 -SAT [ 3 ], while Zaleski implemented a SAT solver in Maple using inclusion–exclusion and Bonferroni inequalities [ 16 ]. More recently, inclusion–exclusion has been combined with dynamic programming over small treewidth for projected model counting (PMC), yielding practically competitive PMC/#SAT solvers [ 8 ]. These results corroborate the general message behind CSFLOC: structural regularities (e.g., small treewidth, strong subsumption) can be exploited to skip large contiguous blocks in the search space. CSFLOC itself belongs to the full-length clause–counting family. It is the successor of the Optimized CCC algorithm [ 14 ]. It uses a counter to count subsumed full-length clauses. By studying Optimized CCC we observed that its full-length clause counter can be increased on its last 1 bit in the best case [ 13 ]. CSFLOC is based on this observation. It also uses a data structure in which the clauses are ordered by the index of their last literal. These two improvements result in a faster algorithm which can compete with a state-of-the-art SAT solver on problems with lots of clauses, like Black-and-White 2-SAT problems [ 6 ] and weakly nondecisive SAT problems [5]. 1.2 Revisiting the CSFLOC algorithm CSFLOC was introduced in [ 13 ]. In this subsection, we recall its pseudocode and the main theoretical results. The following properties underpin Algorithm 1; we state them without proof, since these theorems are already proven in [13]. Lemma 1 (Observation 2.). In the inner loop of CSFLOC, if 𝑘= IndexOfLastPositiveLiteral( 𝐶 ) and 𝑗 is the last-literal index of a 23 MATCOS-25, 9–10 October 2024, Koper,Slovenia Gábor Kusper Algorithm 1 CSFLOC(𝑆) Require: 𝑆 is a non-empty list of ordered clauses with variable index function 𝐼. Ensure: If 𝑆 is satisfiable it returns a solution of 𝑆 , otherwise returns the empty set. 1: 𝑛:=number of variables in 𝑆; 2: 𝑆[𝑖] : ={𝐶|𝐶∈𝑆∧ IndexOfLastLiteral( 𝐶 ) =𝑖} , where 𝑖= 1 ..𝑛 ; 3: 𝑐𝑜𝑢𝑛𝑡 :=0; 4: while 𝑐𝑜𝑢𝑛𝑡 <2𝑛do 5: 𝑖𝑛𝑐𝑟𝑒𝑚𝑒𝑛𝑡 :=0; 6: 𝐶:=FullLengthClauseRepresentationOf(𝑐𝑜𝑢𝑛𝑡); 7: for 𝑗 : = IndexOfLastPositiveLiteral( 𝐶 ); 𝑗<=𝑛 ; 𝑗 : =𝑗+ 1 do 8: if ∃𝐷∈𝑆[𝑗]such that 𝐷subsumes 𝐶then 9: 𝑖𝑛𝑐𝑟𝑒𝑚𝑒𝑛𝑡 :=2𝑛−𝑗; 10: 𝑗:=𝑛+1; 11: end if 12: end for 13: if 𝑖𝑛𝑐𝑟𝑒𝑚𝑒𝑛𝑡 =0then 14: return ¬𝐶; 15: else 16: 𝑐𝑜𝑢𝑛𝑡 :=𝑐𝑜𝑢𝑛𝑡 +𝑖𝑛𝑐𝑟𝑒𝑚𝑒𝑛𝑡; 17: end if 18: end while 19: return {}; subsuming clause 𝐷∈𝑆[𝑗] , then 𝑗≥𝑘 , hence the largest possible jump is 2𝑛−𝑘. Theorem 1 (Soundness and Completeness of CSFLOC). Algorithm 1 is sound and complete. The 2. Observation Lemma, i.e., Lemma 1 states that in case of CSFLOC the largest possible jump is 2 𝑛−𝑘 , and Theorem 1 states that this trick is valid, i.e., CSFLOC is a general SAT solver algorithm. See the proof of these theoretical results in [ 13 ]. In this paper we discuss the practical implementation of CSFLOC. 2 Pseudocode-to-Java Mapping The Java implementation referenced in this paper is available at http://fmv.ektf.hu/files/CSFLOC19.java . This section is highly technical. We suggest consulting the Java source code to understand this section. 2.1 Entry Points and Call Flow The Java implementation uses the following functions and flow of calls: •main calls •DIMACSReader (parse DIMACS CNF), which calls •HighLevelReader (variable ordering, bucket build), which calls •CSFLOCSolver.CSFLOC_v7() (outer loop), which calls •usingBestClause_v4 (inner scan), which calls •increaseCounter_v5 (jump) •the outer loop returns eventually a model or UNSAT; • if a model is found, then main calls SimpleCheckSolution which validates the model. 2.2 Element-by-Element Mapping We map each step of Algorithm 1 with concrete classes and methods. Buckets 𝑆[1..𝑛](lines 1–2): built by HighLevelReader using two Clause arrays •clauseListOrderedByLastVarIndexPos[i] and •clauseListOrderedByLastVarIndexNeg[i] and 𝑖∈ {1..𝑛}, storing clauses with IndexOfLastLiteral(𝐶)=𝑖 whose last literal is positive / negative. Counter init (line 3): count is represented by the Boolean vector counter[1..n] (1 = positive, 0 = negative). Implementation uses a practical initial value: the best black clause; it sets the first 1-bit and calls addEffectedClause to prime caches. Full-length clause 𝐶(lines 4–6): Not materialized: the current counter encodes 𝐶 ; subsumption is tested directly via methods like Clause.subsumedBy(counter). Setting 𝑗(line 7): Computed as the index of the last 1-bit in counter ; passed as index to usingBestClause_v4( index, counter). Scan 𝑗..𝑛 for a subsumer (line 7): One of the main methods usingBestClause_v4(i) performs a sign-aware traversal. It tries to find a subsumer clause by testing •effected/oldEffected at level 𝑖, •learnedClausesPos at level 𝑖 •clauseListOrderedByLastVarIndexPos at level 𝑖, •clauseListOrderedByLastVarIndexNeg from level 𝑖+ 1in a loop. This realizes the theoretical 𝑗..𝑛 loop once the sign pattern of 𝐶is taken into account. Subsumption test (line 8): Bit-level subsumption check is done by D.subsumedBy(counter). Increment by 2𝑛−𝑗(line 9): is done by the other main function increaseCounter_v5 which executes the jump with carry-like bit flips on counter , equivalent to adding 2 𝑛−𝑗 to the integer counter; hooks update caches and may insert learned clauses via addEffectedClause. Return ¬𝐶(line 14): Complement the counter and emit the model (optionally map the model back to the original variable indices if a renaming was applied). Termination (lines 15–19): The outer loop in the method CSFLOC_v7() repeats until a model is returned or the jumps exhaust [0,2𝑛), in which case UNSAT is reported. 2.3 Supporting Data Structures and Caches • Sign-aware buckets: In the pseudocode 𝑆[𝑗] denotes the set of clauses whose last-literal index is 𝑗 . In the implementation we maintain two arrays per index, 𝑆+[𝑗] and 𝑆−[𝑗] , for clauses whose last literal is positive or negative, respectively; these are the Java arrays –clauseListOrderedByLastVarIndexPos[j] and –clauseListOrderedByLastVarIndexNeg[j]. The inner scan first probes 𝑆+[𝑗] and then 𝑆−[𝑘] for 𝑘>𝑗 , matching the positive/negative hit order. • Effected / oldEffected: per-index, per-sign small working sets prioritized before base buckets. •Learned clauses: bounded per-index pools –learnedClausesPos, –learnedClausesNeg filled after successful subsumption hits; probed before base buckets. 24 Engineering CSFLOC MATCOS-25, 9–10 October 2024, Koper,Slovenia • Best black clause fields: bestBlackClause is used to set the first index before the main loop to gain a small speed-up. • Renaming metadata: permutation 𝑡𝑟𝑎𝑛𝑠𝑙𝑎𝑡𝑒 maintained by HighLevelReader (see section 4); used for bucket indices and for mapping models back. 3 Counter Dynamics, Hits, Islands, and Straits In this section, we introduce the intuitive notions of positive hit,negative hit,islands and straits to help explain how CSFLOC works. Hit types. Let 𝐶 be the current full-length clause encoded by counter, and let 𝑘=IndexOfLastPositiveLiteral(𝐶) denote the index of the last 1-bit in 𝐶 . For 𝑗∈ {𝑘, . . . , 𝑛} we say there is a positive hit at level 𝑗 if there exists a clause 𝐷∈𝑆[𝑗] whose last literal is positive and 𝐷⊆𝐶 . For 𝑗>𝑘 we say there is a negative hit at level 𝑗 if there exists a clause 𝐷∈𝑆[𝑗] whose last literal is negative and 𝐷⊆𝐶 . Because positions greater than 𝑘 are 0s in 𝐶 , a positive last literal cannot subsume 𝐶 at those levels, hence only negative hits are possible above 𝑘 . In both cases, CSFLOC jumps by 2𝑛−𝑗(see Lemma 1). CSFLOC enumerates full-length clauses through a counter which is a Boolean array (1=positive, 0=negative) and leverages subsumption to jump. In the best case, the next iteration in the main loop increases the index of the last 1-bit by exactly 1, i.e., the trailing 1-block grows by one. In the general case, adding 1 in base-2 may also flip higher 1-bits to 0 (carry), so a single step can translate into a large jump in the integer view of count. Islands and straits. We interpret counter as a sea of zeros with occasional ones. A contiguous block of 1s is an island; a contiguous block of 0s between islands is a strait. Let the rightmost island (closest to index 𝑛 ) be the last island. Growing the last island by one (i.e., increasing the last 1-bit index by one) is the locally optimal step: it preserves previously gained structure and maximizes the chance of a larger safe jump because carries only affect positions to the left of the last island. When is a one-step growth possible? A one-step growth occurs if there exists a clause 𝐷 whose last-literal index equals 𝑗= IndexOfLastPositiveLiteral(𝐶) and 𝐷⊆𝐶 (a positive hit at level 𝑗 ). Otherwise the best we can hope for is that the next 1-bit sits immediately to the left of 𝑗 (i.e., a negative hit at level 𝑗+ 1), so the last island still grows by one. If the nearest 1-bit is farther left, the intervening 0s form a wider strait, which negatively affects the performance of CSFLOC. Implications for ordering and renaming. Because the inner scan starts at 𝑗=IndexOfLastPositiveLiteral(𝐶) and then checks for negative hits at 𝑘>𝑗 , we profit from (i) large islands (long trailing 1-blocks) and (ii) small straits between successive islands. This aligns with runtime data: fewer, longer islands typically enable larger 2 𝑛−𝑗 jumps, whereas many short islands correlate with smaller increments and more scans. Our preordering strategies aim to bias the distribution of last-literal indices accordingly; see the next section, especially the I (Island), S (Strait), and C (Clustering) variable-renaming strategies. Why no dynamic renaming? One might attempt to dynamically adjust the order so that the next 1-bit always follows the last one. However, the implementation reaches clauses by their last literal (a variant of watched-literal indexing). Swapping two variables (say, 8 and 25) can change the last-literal index in many clauses (e.g., {− 1 , 8 , 23 } becomes {− 1 , 25 , 23 } , moving the last literal from 23 to 25). Maintaining bucket memberships and all derived caches at runtime would impose a prohibitive burden, so our solver performs renaming only once before the main loop (see section 4). In conclusion, CSFLOC runs faster when the last island is long and the straits between islands are short. Proving these observations in full generality is challenging; for random 3-SAT, however, parts of them can be established rigorously. A complete theoretical treatment remains an open line of research. The next section provides more information on variable renaming strategies. 4 Variable Renaming Strategies CSFLOC scans buckets from 𝑗=IndexOfLastPositiveLiteral(𝐶) upwards. Hence the variable order determines the distribution of last-literal indices and which buckets are hit early. Renaming aims to (i) concentrate strong constraints early, (ii) lengthen useful trailing 1-bit islands in the counter (larger jumps), and (iii) improve cache locality. Interface. The solver accepts a composite string variableRenamingStrategy consisting of letters from {B,C,H,I,R,S,W} (uppercase or lowercase). Strategies are applied in the order of appearance; lowercase variants indicate a milder weighting/priority. Typical presets in the implementation include: "IWCR" , "HWCR" , "BHWCR" for random 3-SAT, and "B" for pigeonhole families. List of strategies and method names. B : renameBlackClauses. Renames variables so that black clauses (all-negative) receive smaller last-literal indices. Intuition: black clauses seed the formation of 1-bit islands in the counter; a long trailing island implies a highprobability large jump. Bringing black clauses forward reduces warm-up and favors long final islands or at least smaller straits. W : renameWhiteClauses. Symmetric to B. It orders white clauses (all-positive). It reduces fragmentation caused by early all-positive buckets and helps separate positive hits from later negative scans. Used together with Bto shape polarity structure. S : renameStraitClauses. A clause is strait if it contains exactly one negative literal (all others are positive). Such clauses tend to start a new island. This strategy assigns a low index to the unique negative variable (and, if still unassigned, to the remaining variables of these clauses), thereby shrinking the 0-gaps (straits) between islands and lowering last-literal positions. I : renameIslandClauses. Symmetric to S(which targets clauses with exactly one negative literal), Itargets clauses with exactly one positive literal (all others negative), i.e., definite Horn clauses. This strategy moves the unique positive variable to the largest index in the clause (its last literal), but still assigns it a small global index, increasing the chance of positive hits and, in the island–strait view, tending to cancel the last island. H : renameDefiniteHornClauses. Uses the same syntactic filter as I(clauses with exactly one positive literal), but is conservative: it processes a clause only if none of its literals has been renamed by any earlier strategy. Otherwise, it is the same as I. Because Htriggers only on untouched clauses, it composes well with polarity-oriented strategies such as W(white) and B(black); the combined effect is 25 MATCOS-25, 9–10 October 2024, Koper,Slovenia Gábor Kusper Table 1: Where CSFLOC excels vs struggles. Instance type Best solver Intuition Random SAT (uf50/uuf50) Glucose weak structure Pigeonhole (small 𝑛) Glucose black and white WSN BW 2-SAT CSFLOC over-constrained WnD UNSAT CSFLOC many clauses SM, dense >10% CSFLOC over-constrained SM, sparse <10% Glucose few subsumers to lower last-positive positions and stabilize long trailing islands. R : simplerVariableRenaming. It orders variables by their occurrence counts; high-frequency variables result in earlier hits. C : clusterVariables. Clustering orders variables so that variables that frequently co-occur in clauses are placed close to each other. The cluster size is controlled by the clustering factor (an integer ≥ 2). Implementation-wise, we compute variable-pair frequencies, greedily form and merge clusters up to the factor, and then assign low consecutive indices cluster by cluster (applying the resulting translate to the formula). This shortens gaps between variables that tend to appear together, compresses lastliteral positions, and improves the hit rate in the positive/negative hit scans. The option is enabled by including C in the renaming string and can be combined effectively with R,H, and S(and their compositions). Following Jebelean’s original clustering idea [ 12 ], our experiments confirm it is the most consistently helpful preordering among the options we evaluated. Heuristic notes. Promoting black clauses (B) early tends to create fewer but longer trailing islands, enabling larger jumps; many short islands are typically unfavorable, though we do not state formal lemmas due to space. Islandand Horn-aware orders (I,H) often dominate on over-constrained inputs; randomized tie-perturbation inside clusters mitigates adversarial cases. 5 Empirical Snapshot This section may contain informal terms due to lack of space, but these have been clarified in [13]. The tests were done on iMac macOS Sierra (CPU: 2,5GHz Intel Core i5, Memory: 4GB 1333MHz DDR3). We follow the measurement settings of our previous report [ 13 ]: the Java implementation of CSFLOC is compared against off-theshelf CDCL baselines (Glucose 3.0 1 ) on standard suites (SATLIB uf/uuf and pigeonhole), generator-based families (WnD UNSAT), a Black-and-White 2-SAT model of wireless sensor networks (WSN), and graph-induced SAT encodings (SM/BM/SBB). Timeouts and machine details are kept fixed across solvers; instances are run in a single thread. Where CSFLOC shines (and where it does not). CSFLOC is not a random-SAT solver; its strengths appear on over-constrained or structurally dense inputs where subsumption enables large jumps. Effect of preordering (clustering and renaming). Preordering is crucial for CSFLOC because the inner loop scans by last-literal 1http://www.labri.fr/perso/lsimon/glucose/ Table 2: Qualitative impact of clustering on SATLIB. Suite no clustering cluster-2 cluster-3 uf50 baseline ↓run time ↓↓ run time uf75 often timeout ✓many solved ✓✓ most solved uf100 timeout ✓some solved ✓✓ many solved index and sign. A simple clustering of variables into small contiguous groups (section 4) consistently improves run time on SATLIB: uf50 becomes notably faster; uf75/uf100, which tend to time out without clustering, become solvable under 2–3 variable clusters. Frequencyand Horn-biased orders (R/H) further reduce the last-positive indices encountered, increasing positive hits. Graph-induced encodings. On SM/BM/SBB encodings derived from directed graphs, we observe a density threshold: for edge densities above roughly 10%, CSFLOC tends to outperform Glucose; below that, CDCL remains preferable. The effect aligns with the island/strait picture: higher density increases the chance that a positive hit occurs or that a nearby negative bucket contains subsumers, extending the last island and triggering larger jumps. 6 Conclusion and Future Work We engineered a faithful and fast implementation of CSFLOC, clarifying how the last-1-bit observation can be realized efficiently via bucketed, sign-aware data structures, clause learning and variable preordering. Future work includes parallel traversal of counter ranges, richer learning schemes, and broader benchmarks. References [1] S.Andrei, Counting for Satisfiability by Inverting Resolution, Artificial Intelligence Review,Volume 22, Issue 4,339–366, 2004. [2] E. Birnbaum, E. L. Lozinskii, The Good Old Davis-Putnam Procedure Helps Counting Models, Journal Of Artificial Intelligence Research, Volume 10, pages 457–477, DOI: https://doi.org/10.1613/jair.601, 1999. [3] H. Bennett and S. Sankaranarayanan, Model Counting Using the InclusionExclusion Principle, Theory and Applications of Satisfiability Testing - SAT 2011 Lecture Notes in Computer Science, Volume 6695, 362–363, 2011. [4] A. Biere, M. Heule, H. van Maaren, T. Walsh, Handbook of Satisfiability, 2nd. edition, IOS Press, Amsterdam, 2021. [5] Cs. Biro and G. Kusper, How to generate weakly nondecisive SAT instances, Proceedings of 11th International IEEE Symposium on Intelligent Systems and Informatics (SISY), 265–269, Subotica, 2013. [6] Cs. Biro and G. Kusper, Equivalence of Strongly Connected Graphs and Blackand-White 2-SAT Problems, Miskolc Mathematical Notes, accepted manuscript, MMN-2140. [7] S. A. Cook, The Complexity of Theorem-Proving Procedures, Proc. of STOC’71, 151–158, 1971. [8] J. K. Fichte, et al., Solving Projected Model Counting by Utilizing Treewidth and its Limits, Artificial Intelligence 314:103810, 2023. [9] M. Davis, G. Logemann, D. Loveland, A Machine Program for Theorem Proving, Communications of the ACM, Volume 5, 394–397, 1962. [10] Carla P. Gomes, Ashish Sabharwal, and Bart Selman, Model Counting, Chapter 20 of Handbook of Satisfiability, IOS Press, Amsterdam, 2009. [11] K. Iwama, CNF-satisfiability test by counting and polynomial average time, SIAM Journal on Computing, Volume 18, Issue 2, 385–391, 1989. [12] T. Jebelean and G. Kusper, Multi–Domain Logic and its Applications to SAT, (invited talk), SYNASC’08, DOI: 10.1109/SYNASC.2008.93, IEEE Computer Society Press, ISBN 978-0-7695-3523-4, 3–8, 2008. [13] G. Kusper, Cs. Biró, Gy. B. Iszály,SAT solving by CSFLOC, the next generation of full-length clause counting algorithms, Proceedings of IEEE International Conference on Future IoT Technologies 2018, DOI: 10.1109/FIOT.2018.8325589, 2018. [14] G. Kusper, Cs. Biró,Solving SAT by an Iterative Version of the Inclusion-Exclusion Principle, Proceedings of SYNASC 2015, DOI: 10.1109/SYNASC.2015.38, pp. 189–190, 2015. [15] E. L. Lozinskii, Counting propositional models, Information Processing Letters, Volume 41, 327–332, 1992. [16] A. Zaleski, Solving Satisfiability using Inclusion-Exclusion, arXiv preprint, DOI: arXiv:1712.06587, 2017. 26 Non-redundant Systems of Independence Atoms in Relational Databases Lucas Alland∗ Swarthmore College Swarthmore, PA, United States Attila Sali HUN-REN Alfréd Rényi Institute Budapest, Hungary Nicole Wu∗ Harvey Mudd College Claremont, CA, United States Abstract Let 𝑅={𝐴1, 𝐴2, . . . , 𝐴𝑛} be a relational database schema and 𝑟 be a relation over 𝑅 . For 𝑋,𝑌 ⊂𝑅 with 𝑋∩𝑌=∅ , 𝑟 is said to satisfy independence atom 𝑋⊥𝑌 , if the projection 𝑟(𝑋𝑌) of 𝑅 to 𝑋∪𝑌 is the Cartesian product of the projections 𝑟(𝑋) and 𝑟(𝑌) , i.e 𝑟(𝑋𝑌 )=𝑟(𝑋) × 𝑟(𝑌) . Implication of independence atoms is defined naturally and a collection of independence atoms is nonredundant, if none of its members is impled by the remaining ones. In the present paper the maximum possible size of nonredundant system of independence atoms is investigated. Keywords relational databases, independence atoms, non-redundant system 1 Introduction In this paper, we investigate an efficient subclass of embedded multivalued data dependencies which are called – in accordance with [1] – independence atoms. Definition 1.1. A relation 𝑟 satisfies the independence atom 𝑋⊥𝑌 between two disjoint sets 𝑋 and 𝑌 of attributes, if for all tuples 𝑡1, 𝑡2∈𝑟 there is some tuple 𝑡∈𝑟 which matches the values of 𝑡1 on all attributes in 𝑋 and matches the values of 𝑡2 on all attributes in 𝑌. In other words, in relations that satisfy 𝑋⊥𝑌 , the occurrence of 𝑋-values is independent of the occurrence of 𝑌-values. If Σ is a collection of independence atoms and 𝜎 is an independence atom, then Σ implies 𝜎 , in notation Σ|=𝜎 , if any database relation 𝑟 that satisfies every atom in Σ also satisfies 𝜎 . The implication problem was axiomatized by Kontinen , Link and Väänänen [2]. The following four rules are sound and complete system for the implication of independence atoms. (1) 𝑋⊥ ∅ (trivial independence, T). (2) 𝑋⊥𝑌⇒𝑌⊥𝑋(symmetry, S). (3) 𝑋⊥𝑌∪𝑍⇒𝑋⊥𝑌(decomposition, D). (4) 𝑋⊥𝑌∧𝑋∪𝑌⊥𝑍⇒𝑋⊥𝑌∪𝑍(exchange, E). That is, Σ|=𝜎 iff 𝜎 has a finite derivation from Σ using the rules above. A set of independence atoms, Σ is non-redundant if, for each atom 𝜎∈Σ , Σ\ {𝜎}⊭𝜎 —no atom can be derived from the other atoms. The following interesting problem was asked by Sebastian Link [3]. Let 𝑓(𝑛) denote the largest size of a non-redundant ∗ This research was done under the auspicies of Research Opportunties course at Budapest Semesters in Mathematics Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). MATCOS-25, October 9th and 10th, 2025., Ljubljana, Slovenia ©2025 Copyright held by the owner/author(s). collection of independent atoms over a schema of 𝑛 attributes. Determine or give good bounds for 𝑓(𝑛). For the sake of convenience, we identify the schema of 𝑛 attributes {𝐴1, 𝐴2, . . . , 𝐴𝑛} with the set [𝑛] of the first 𝑛 positive integers. It was proven in [2] that every system Σ of independence atoms has an Armstrong database, that is a database that satisfies an independence atom 𝜎 iff Σ|=𝜎 . This implies that every nonredundant system of independence atoms actually occurs as a system satisfied by a particular database, so for us it is enough to consider implications at the schema level. 2 Lower Bounds One may check easily that 𝑓( 1 )= 0and 𝑓( 2 )= 1. An immediate lower bound for 𝑓(𝑛) follows from the observation that none of the derivation rules introduces new attributes. That is, 𝑋⊥𝑌 cannot be derived from a set Σ={𝑋𝑖⊥𝑌𝑖:𝑖∈𝐼} if 𝑋𝑖∪𝑌𝑖⊉ 𝑋∪𝑌for all 𝑖∈𝐼. Proposition 2.1. For 𝑛≥3we have 𝑓(𝑛) ≥ 𝑛 ⌈𝑛 2⌉. Proof. Let {𝑋𝑖∪𝑌𝑖:𝑖= 1 , 2 , . . . , 𝑛 ⌈𝑛 2⌉} be list of nontrivial partitions of the ⌈𝑛 2⌉ -element subsets of [𝑛] . The system Σ= {𝑋𝑖⊥𝑌𝑖:𝑖= 1 , 2 , . . . , 𝑛 ⌈𝑛 2⌉} is clearly non-redundant, as 𝑋𝑖∪ 𝑌𝑖⊈𝑋𝑗∪𝑌𝑗for 𝑖≠𝑗.□ It is interesting to observe that this simple lower bound is sharp for small 𝑛. Proposition 2.2. 𝑓(3)=3and 𝑓(4)=6. The proof of the first statement is easy, the latter one needs detailed case-by-case analysis. We have improved on the simple lower bound above by using a recursive construction. Theorem 2.3. For 𝑛≥ 2we have 𝑓(𝑛) ≥ 𝑓(𝑛− 1 ) + 𝑛−1 ⌈𝑛−1 2⌉ implying 𝑓(𝑛) ≥ 𝑓(1) + 𝑛 ∑︁ 𝑖=2𝑖−1 ⌈𝑖−1 2⌉= 𝑛−1 ∑︁ 𝑖=1𝑖 ⌈𝑖 2⌉. Proof. The proof is based on the following observation of [2]. Let Σ be a collection of independence atoms over [𝑛] and 𝑋⊥𝑌 be an independence atom. Let Σ′=Σ[𝑋𝑌]={𝑉∩ (𝑋∪𝑌) ⊥ 𝑊∩ (𝑋∪𝑌):𝑉⊥𝑊∈Σ} . If there is no non-trivial atom 𝑈⊥𝑉∈Σ′where 𝑈∪𝑉=𝑋∪𝑌, then Σ|=𝑋⊥𝑌. Now assume that Σ𝑛−1 is a non-redundant system of independence atoms of size 𝑓(𝑛− 1 ) over [𝑛− 1 ] . Define Σ𝑛= Σ𝑛−1∪{𝑛} ⊥ 𝑋:𝑋⊂ [𝑛−1]with |𝑋|=⌈𝑛−1 2⌉ . We claim that Σ𝑛 is non-redundant. Indeed, none of the atoms of the form {𝑛} ⊥ 𝑋 can be derived from Σ𝑛\ {{𝑛} ⊥ 𝑋} as no other atom contains all attributes of {𝑛} ⊥ 𝑋 . On the other hand, the atoms added to Σ𝑛−1 cannot be used in a derivation for an atom in Σ𝑛−1 since their intersection with [𝑛− 1 ] is a trivial atom of the form ∅ ⊥ 𝑋.□ 27 MATCOS-25, October 9th and 10th, 2025., Ljubljana, Slovenia Lucas Alland, Attila Sali, and Nicole Wu 3 Upper bounds A useful concept in obtaining upper bounds is the concept of atom shapes. Definition 3.1. An atom 𝑋⊥𝑌 is defined as having the shape (|𝑋|,|𝑌|) . Notice by symmetry ( S ) that an (𝑎,𝑏) shape is equivalent to an (𝑏, 𝑎)shape. Non-trivial shapes on five attributes, with examples. (1) (1,1):1⊥2 (2) (1,2):1⊥2 3 (3) (2,2):1 2 ⊥3 4 (4) (1,3):1⊥123 (5) (1,4):1⊥2345 (6) (2,3):1 2 ⊥345 Proposition 3.2. A maximum sized non-redundant system of independence atoms over 5 attributes cannot contain an atom of shape (1,4). Proof. Let 𝛼= 1 ⊥ 2 3 4 5 and Σ be a non-redundant set on 5attributes, containing 𝛼 . We see that 𝛼⊨ 1 ⊥𝑌 for all 𝑌⊆ {2,3,4,5}. We partition Σ\ {𝛼}by the sets 𝑆and 𝑇where: 𝑆={1∪𝑋⊥𝑌∈Σ|𝑋≠∅} 𝑇={𝑈⊥𝑉∈Σ|𝑈∪𝑉⊆ {2,3,4,5}}. Now define 𝑆′={𝑋⊥𝑌|1∪𝑋⊥𝑌∈𝑆}. We will show Σ′=𝑆′∪𝑇 is non-redundant. Take any 𝜎′∈Σ′ and consider two cases: 𝜎′=𝑈⊥𝑉∈𝑇 Every atom in 𝑆′ is derivable from its corresponding atom in 𝑆 . So if 𝜎′∈𝑇 were derivable from Σ′\ {𝜎′} then the same sequence of implications would show Σ\ {𝜎′}⊨𝜎′. 𝜎′=𝑋⊥𝑌∈𝑆′Let 𝜎=1∪𝑋⊥𝑌. Now, we observe: 𝛼⊨1⊥𝑋∪𝑌by D 1⊥𝑋∪𝑌∧𝑋⊥𝑌⊨𝜎by E. So if Σ′\ {𝜎′}⊨𝜎′ then with the above two implications we have Σ\ {𝜎}⊨𝜎. So we conclude Σ′ is non-redundant. Since 𝑓( 4 )= 6, and Σ′ is a set on four attributes, |Σ′| ≤ 6. Then, |Σ|=|{𝛼}| + |𝑆|+|𝑇|=1+ |𝑆′|+|𝑇|=1+ |Σ′| ≤ 7. □ We note that the previous argument generalizes to all 𝑛 from the given 𝑛= 5case. We can in general conclude that if Σ is a non-redundant set of maximum size on 𝑛 attributes that contain a(1,𝑛 −1)shape atom: |Σ| ≤ 1+𝑓(𝑛−1). Since our lower bound for 𝑓(𝑛) increases faster than 1when 𝑛≥ 3, the maximum-order non-redundant attribute set will never contain a (1, 𝑛 −1)shape atom. Proposition 3.3. A maximum sized non-redundant set of independence atoms on 5 attributes cannot contain an atom of shape (2,3). Proof. Let Σ be non-redundant, suppose we have 𝐾⊆Σ such that 𝐾⊨𝛼and then define 𝐿={1⊥𝑌∈Σ\𝐾} 𝑆={1∪𝑋⊥𝑌∈Σ\𝐾|𝑋≠∅} 𝑇={𝑈⊥𝑉∈Σ\𝐾|𝑈∪𝑉⊆ {2,3,4,5}}. Then, we may partition Σby Σ=𝐾∪𝐿∪𝑆∪𝑇 . Note, that |𝐿| ≤ 1by D . Now, recalling the proof that there are no (1,4)atoms, we can conclude |𝑆∪𝑇| ≤ 6. Indeed, since 𝛼 is derived from 𝐾 and 𝑆∪𝑇⊆Σ\𝐾 , we may use 𝛼in the derivations from 𝑆∪𝑇. This implies that |Σ|≤|𝐾| + 7. Suppose 𝛽= 1 2 ⊥ 345 ∈Σ , and Σ is non-redundant. Note that 𝛽 implies 𝛼 by exchange ( E ) if 𝛾= 1 ⊥ 2is derivable from Σ\ {𝛽} . Consider two cases: (1) If Σ⊭𝛾 then the only allowed atoms in Σ\ {𝛽} are from the sets: 𝑃={𝑈⊥𝑉|𝑈∪𝑉⊆ {3,4,5},𝑈 ≠∅,𝑉 ≠∅} 𝑄={2∪𝑈⊥𝑉|𝑈∪𝑉⊆ {3,4,5},𝑈 ≠∅,𝑉 ≠∅} 𝑅={1∪𝑈⊥𝑉|𝑈∪𝑉⊆ {3,4,5},𝑈 ≠∅,𝑉 ≠∅} 𝑊={1 2 ∪𝑈⊥𝑉|𝑈∪𝑉⊆ {3,4,5},𝑈 ≠∅,𝑉 ≠∅}. In particular, 𝐿=∅ , since 𝑌⊆ { 3 , 4 , 5 } would hold, thus 𝛽|=1⊥𝑌by D. Observe that for 𝑈∪𝑉⊆ { 3 , 4 , 5 },𝑈 ≠∅,𝑉 ≠∅ {𝑈⊥ 𝑉, 𝛽} |= 1 2 ∪𝑈⊥𝑉 by D,E . So for a given pair 𝑈∪𝑉⊆ { 3 , 4 , 5 },𝑈 ≠∅,𝑉 ≠∅ at most one atom from the sets 𝑃,𝑄, 𝑅,𝑊 can be in Σ . That is, if 𝑄′={𝑈⊥𝑉: 2 ∪𝑈⊥ 𝑉∈𝑄} , 𝑅′={𝑈⊥𝑉: 1 ∪𝑈⊥𝑉∈𝑅} and 𝑊′={𝑈⊥ 𝑉: 1 2 ∪𝑈⊥𝑉∈𝑊} , then these sets are pairwise disjoint and also 𝑃∪𝑄′∪𝑅′∪𝑊′ is a non-redundant system of independence atoms over the attribute set { 3 , 4 , 5 } . Thus, |𝑃∪𝑄∪𝑅∪𝑊|=|𝑃∪𝑄′∪𝑅′∪𝑊′| ≤ 3implying |Σ| ≤ 4<12. (2) If Σ⊨𝛾then we have an atom from the set 𝑆={1∪𝑋⊥2∪𝑌|𝑋∪𝑌⊆ {3,4,5}} With 𝛽 , one such atom will derive 𝛼 . Then, |𝐾| ≤ 2which forces |Σ| ≤ 9<12. □ Proposition 3.4. 𝑓(5) ≤ 30 Proof. We use the above two facts to find an upper bound on 𝑓( 5 ) . Let |Σ|= 31 = 5 𝑓( 4 ) + 1on 5attributes. Then, for Σ to be non-redundant: (1) Each attribute from [ 5 ] may miss at most 𝑓( 4 ) atoms, since those atoms must form a non-redundant system on the remaining four attributes. Thus, each attribute must appear in at least |Σ| − 𝑓(4)=31 −6=25 times; (2) Each atom can contain at most 4attributes. Comparing the possible number of appearances and the necessary number of appearances over all attributes would imply the following inequality: 125 =5(24 +1) ≤ 4|Σ|=124 28 Non-redundant Systems of Independence Atoms in Relational Databases MATCOS-25, October 9th and 10th, 2025., Ljubljana, Slovenia which clearly does not hold. So for any non-redundant atom-set Σ,|Σ| ≤ 30. Thus, 𝑓(5) ≤ 30.□ We can give a general upper bound. Proposition 3.5. 𝑓(𝑛)<(3−𝜀)𝑛 Proof. Consider the set of all independence atoms to be a partially ordered set according to the rule 𝑍⊥𝑊≼𝑋⊥𝑌 if 𝑋⊥𝑌⊨𝑍⊥𝑊 by decomposition ( D ). This is a graded poset by the rank function 𝜌(𝑋⊥𝑌)=|𝑋∪𝑌|. For example, the elements of level 1are the trivial atoms with the empty set on one side and one attribute on the other. A rank 𝑘 element of this poset is covered by 2 (𝑛−𝑘) elements of rank 𝑘+ 1, since we have two choices to which side of the atom the new attribute is added. On the other hand, a rank 𝑘+ 1 element covers 𝑘+ 1rank 𝑘 elements. This implies that this graded poset has the Sperner property — no antichain is larger than the largest rank level. Since every rank-level is itself an antichain, the largest rank-level will be a maximum-sized antichain. Thus, since any non-redundant set must be the subset of some antichain, 𝑓(𝑛)is bounded by the size of the largest rank-level: 𝑓(𝑛) ≤ max{𝐿(𝑘) | 1≤𝑘≤𝑛} where 𝐿(𝑘) denotes the size of rank-level 𝑘 . We consider how 𝐿(𝑘)relates to 𝐿(𝑘+1). The covering numbers determined above give us the following relation: 2· (𝑛−𝑘) · 𝐿(𝑘)=(𝑘+1) · 𝐿(𝑘+1). Using the above iteratively and 𝐿(1)=|{𝑘⊥ ∅,∅ ⊥ 𝑘|𝑘∈ [𝑛]}| =2𝑛 we obtain 𝐿(𝑘)=𝑛·2(𝑛−1) 1+1·. . . ·2(𝑛−𝑘+1) 𝑘−1+1=2𝑘𝑛 𝑘. Then, the maximum of 𝐿(𝑘) is reached (non-uniquely) when 𝑘=⌊2𝑛 3⌋. Hence, 𝑓(𝑛) ≤ 𝐿(⌊ 2𝑛 3⌋) =2⌊2𝑛 3⌋𝑛 ⌊2𝑛 3⌋<(3−𝜀)𝑛 where the final bound can be obtained by applying some asymptotics on the binomial coefficient. □ 4 Conclusions We have studied non-redundant systems of independence atoms of relational database schemata. We defined 𝑓(𝑛) to be the largest possible size of such a system over a schema of 𝑛 attributes. 𝑓(𝑛) was determined for small 𝑛 and we gave general lower and upper bounds. We conjecture that Conjecture 4.1. 𝑓(5)=12. Also, we ask whether the general lower bound construction is optimal. We believe so, but we do not dare to put it as a conjecture. References [1] Erich Grädel and Jouko Väänänen. 2013. Dependence and independence. Studia Logica, 101, 399–410. https://www.jstor.org/stable/23488329. [2] Juha Kontinen, Sebastian Link, and Jouko Väänänen. 2013. Independence in database relations. In Logic, Language, Information, and Computation. Leonid Libkin, Ulrich Kohlenbach, and Ruy de Queiroz, editors. Springer Berlin Heidelberg, Berlin, Heidelberg, 179–193. isbn: 978-3-642-39992-3. [3] Sebastian Link. 2024. Personal Communication. (2024). 29 Scrambler Automaton Block Cipher for IoT Devices Pál Dömösi University of Debrecen 4028 Debrecen, Kassai Road 26., Hungary [email protected] Géza Horváth University of Debrecen 4028 Debrecen, Kassai Road 26., Hungary [email protected] ABSTRACT In this article, we introduce a new block cipher based on finite automata. Its structure is simple, using few and inexpensive operations, making it particularly suitable for lightweight cryptographic applications. KEYWORDS finite automata, Internet of Things, lightweight cryptography, block cypher 1 INTRODUCTION Communication has undergone significant changes in the 21st century. While communication initially took place between people, and then in many cases between computers by the end of the 20th century, in the 21st century countless small smart devices communicate with each other and their environment via the Internet. These devices and this changed environment are collectively referred to as the Internet of Things, or IoT for short. These changes have inevitably forced changes in secret communication and encryption. Since encrypted communication must be implemented using inexpensive and simple tools, procedures that use few simple operations, require little memory and storage space, and still provide fast and secure communication have come to the fore. These procedures are collectively referred to as lightweight cryptography. A significant step towards lightweight cryptography based on automata theory was the stream cipher introduced by Pál Dömösi and Géza Horváth in 2017 [ 1 ]. This stream cipher was subjected to thorough testing, which confirmed that the system is resistant to side-channel attacks, a type of attack that plays a significant role in attacks against IoT devices [ 2 ]. The authors of paper [ 1 ] described a scrambler method in patent [ 5 ], demonstrating its use as a pseudo-random number generator in the articles [ 3 ] and [ 4 ]. This article describes the use of the scrambler method described in patent [ 5 ] as a block cipher. A well-known common weakness of symmetric encryption is that a suitable method (such as asymmetric encryption) is required to exchange (synchronize) the secret key without revealing confidential information. Therefore, the cipher presented in this paper does not provide a solution to this difficulty. 2 BASIC CONCEPTS In this lecture we consider a novel type of block cipher based on abstract finite automata. This cipher consists of three parts. One of them is a counter which sends its current state as input to a so-called scrambler automaton, which also belongs to the cipher. The second part is a feeder, which passes the plaintext Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society 2024, 7–11 October 2024, Ljubljana, Slovenia ©2024 Copyright held by the owner/author(s). to the cipher block by block for encryption. The third part is a scrambler automaton which changes from the state received from the feeder to another state in response to an input signal received from the counter. The new state will be the next output block of the cipher, i.e. the next ciphertext block. Decryption is essentially done in the same way, but the role of the scrambler automaton is taken over by a so-called inverse scrambler automaton. It also changes from the state, the next ciphertext block, received from the feeder to another state in response to an input signal received from the counter. This response is nothing but the next block of the plaintext (i.e. the decrypted secret text block). The counter starts its operation with the same secret initial state as during encryption. After then sends its current state as input to the inverse scrambler automaton of the cipher. The term "counter" is used to refer to a method and apparatus that is presumed to operate according to a discrete time scale such that at the start of its operation it is in a fixed state 𝑠0,and in every subsequent time instant 𝑡 its state is an element 𝑠𝑡 of the nonempty, finite state set 𝑆, where 𝑠𝑡∈𝑆 denotes the state of the counter at the time instant immediately preceding the time instant 𝑡, while 𝑓 : 𝑆→𝑆 is a function adapted to map the nonempty, finite set 𝑆 to itself in a bijective manner. The triplet S=(𝑆, 𝑠0, 𝑓 ) will herein after also be referred to as the base structure of the counter, set 𝑆will be called the state set of the counter, state 𝑠0∈𝑆 the core of the counter, and function 𝑓:𝑆→𝑆the state transition function of the counter. In the following, it is assumed that the state transition functions f applied in this application are very simple, preferably 𝑆={ 0 , 1 , . . . , 2 128 − 1 }, and for each 𝑘∈𝑆, 𝑓 (𝑘)=𝑘+ 1 , if 𝑘+1<2128 −1,and 𝑓(𝑘)=0,if 𝑘+1=2128 −1.) 3 TRANSPOSITION-CONTROLLED AUTOMATON For every 𝑚= 1 , 2 , . . . , define the permutations 𝑃1, 𝑃2, . . . , 𝑃𝑚 over the set {1, . . . , 2𝑚}in the following way. Let 𝑛 be a fixed positive integer power of 2 , and let us define the following permutations that are specified as a product of transpositions (for example, for a permutation 𝑃 the transposition ( 9 , 13 ) " which thus denotes such a pair " means that 𝑃( 9 )= 13 and 𝑃( 13 )= 9 ). For specifying these permutations, let us consider the following algorithm for the vector (1, . . . , 𝑛): If 𝑛=2,then let 𝑃1=(1,2),and we are ready. Else, let us consider the vectors ( 1 , . . . , 𝑛/ 2 ) and (𝑛/ 2 + 1 , . . . , 𝑛), and let us generate the permutation 𝑃1 such that 𝑃1 is a product of such transpositions where for each 𝑘∈ { 1 , . . . , 𝑛/ 2 } the first component of the 𝑘 -th factor of this transposition-product is the 𝑘 -th component of the vector ( 1 , . . . , 𝑛/ 2 ), while the second component thereof is the 𝑘 -th component of the vector (𝑛/ 2 + 1 , . . . , 𝑛) . Taken to an expression: 𝑃1=( 1 ,𝑛/ 2 + 1 )( 2 ,𝑛/ 2 + 2 ). . . , (𝑛/ 2 ,𝑛). If 𝑛=4,then let 𝑃2=(1,2)(3,4),and we are ready. Else, let us carry out the above process separately for the vector ( 1 , . . . , 𝑛/ 2 ) and for the vector (𝑛/ 2 + 1 , . . . , 𝑛). The product of the two permutations ( 1 ,𝑛/ 4 + 1 )( 2 ,𝑛/ 4 + 2 ) · · · (𝑛/ 4 , 2 𝑛/ 4 ) and ( 2 𝑛/ 4 + 1 , 3 𝑛/ 4 + 1 )( 2 𝑛/ 4 + 2 , 3 𝑛/ 4 + 2 ) · · · ( 3 𝑛/ 4 , 4 𝑛/ 4 ) thus 30 Information Society 2024, 7–11 October 2024, Ljubljana, Slovenia Pál Dömösi and Géza Horváth obtained will be the permutation 𝑃2=( 1 ,𝑛/ 4 + 1 )( 2 ,𝑛/ 4 + 2 ) · · · (𝑛/ 4 , 2 𝑛/ 4 ) ( 2 𝑛/ 4 + 1 , 3 𝑛/ 4 + 1 )( 2 𝑛/ 4 + 2 , 3 𝑛/ 4 + 2 ) · · · ( 3 𝑛/ 4 , 4𝑛/4). Else, like with the above, let us carry out the process separately for the vectors ( 1 , . . . , 𝑛/ 4 ),(𝑛/ 4 + 1 , . . . , 2 𝑛/ 4 ),( 2 𝑛/ 4 + 1 , . . . , 3 𝑛/ 4 ),( 3 𝑛/ 4 + 1 , . . . , 4 𝑛/ 4 ). The product of the four permutations ( 1 ,𝑛/ 8 + 1 )( 2 ,𝑛/ 8 + 2 ) · · · (𝑛/ 8 , 2 𝑛/ 8 ),( 2 𝑛/ 8 + 1 , 3 𝑛/ 8 + 1 )( 2 𝑛/ 8 + 2 , 3 𝑛/ 8 + 2 ) · · · ( 3 𝑛/ 8 , 4 𝑛/ 8 ),( 4 𝑛/ 8 + 1 , 5 𝑛/ 8 + 1 )( 4 𝑛/ 8 + 2 , 5 𝑛/ 8 + 2 ) · · · ( 5 𝑛/ 8 , 6 𝑛/ 8 ),( 6 𝑛/ 8 + 1 , 7 𝑛/ 8 + 1 )( 6 𝑛/ 8 + 2 , 7 𝑛/ 8 + 2 ) · · · ( 7 𝑛/ 8 , 8 𝑛/ 8 ) thereby obtained will be the permutation 𝑃3= ( 1 ,𝑛/ 8 + 1 )( 2 ,𝑛/ 8 + 2 ) · · · (𝑛/ 8 , 2 𝑛/ 8 )( 2 𝑛/ 8 + 1 , 3 𝑛/ 8 + 1 ) · · · ( 5 𝑛/ 8 , 6𝑛/8)(6𝑛/8+1,7𝑛/8+1)(6𝑛/8+2,7𝑛/8+2) · · · (7𝑛/8,8𝑛/8). If 𝑛=16,then let 𝑃4=(1,2)(3,4). . . , (15,16). Else, like with the above, let us carry out the process separately for the vectors ( 1 ,· · · ,𝑛/ 8 ),(𝑛/ 8 + 1 ,· · · , 2 𝑛/ 8 ),( 2 𝑛/ 8 + 1 ,· · · , 3 𝑛/ 8 ),( 3 𝑛/ 8 + 1 ,· · · , 4 𝑛/ 8 ),( 4 𝑛/ 8 + 1 ,· · · , 5 𝑛/ 8 ),( 5 𝑛/ 8 + 1,· · · ,6𝑛/8),(6𝑛/8+1,· · · ,7𝑛/8),(7𝑛/8+1,· · · ,8𝑛/8). The product of the eight permutations ( 1 ,𝑛/ 16 + 1 )( 2 ,𝑛/ 16 + 2 ) · · · (𝑛/ 16 , 2 𝑛/ 16 ),( 2 𝑛/ 16 + 2 , 3 𝑛/ 16 + 2 ) · · · ( 3 𝑛/ 16 , 4 𝑛/ 16 ), ( 4 𝑛/ 16 + 1 , 5 𝑛/ 16 + 1 ) ( 4 𝑛/ 16 + 2 , 5 𝑛/ 16 + 2 ) · · · ( 5 𝑛/ 16 , 6 𝑛/ 16 ), ( 6 𝑛/ 16 + 1 , 7 𝑛/ 16 + 1 ) ( 6 𝑛/ 16 + 2 , 7 𝑛/ 16 + 2 ) · · · ( 7 𝑛/ 16 , 8 𝑛/ 16 ), ( 8 𝑛/ 16 + 1 , 9 𝑛/ 16 + 1 )( 8 𝑛/ 16 + 2 , 9 𝑛/ 16 + 2 ) · · · ( 9 𝑛/ 16 , 10 𝑛/ 16 ), ( 12 𝑛/ 16 + 1 , 13 𝑛/ 16 + 1 )( 12 𝑛/ 16 + 2 , 13 𝑛/ 16 + 2 ) · · · ( 13 𝑛/ 16 , 14 𝑛/ 16 ),( 14 𝑛/ 16 + 1 , 15 𝑛/ 16 + 1 )( 14 𝑛/ 16 + 2 , 15 𝑛/ 16 + 2 ) · · · ( 15 𝑛/ 16 , 16 𝑛/ 16 ) thus obtained will be the permutation 𝑃4= ( 1 ,𝑛/ 16 + 1 )( 2 ,𝑛/ 16 + 2 ) · · · (𝑛/ 16 , 2 𝑛/ 16 )( 2 𝑛/ 16 + 1 , 3 𝑛/ 16 + 1 ) ( 2 𝑛/ 16 + 2 , 3 𝑛/ 16 + 2 ) · · · ( 3 𝑛/ 16 , 4 𝑛/ 16 )( 4 𝑛/ 16 + 1 , 5 𝑛/ 16 + 1 ) ( 4 𝑛/ 16 + 2 , 5 𝑛/ 16 + 2 ) · · · ( 5 𝑛/ 16 , 6 𝑛/ 16 )( 6 𝑛/ 16 + 1 , 7 𝑛/ 16 + 1 ) ( 6 𝑛/ 16 + 2 , 7 𝑛/ 16 + 2 ) · · · ( 7 𝑛/ 16 , 8 𝑛/ 16 )( 8 𝑛/ 16 + 1 , 9 𝑛/ 16 + 1 ) ( 8 𝑛/ 16 + 2 , 9 𝑛/ 16 + 2 ) · · · ( 9 𝑛/ 16 , 10 𝑛/ 16 ) ( 10 𝑛/ 16 + 1 , 11 𝑛/ 16 + 1 )( 10 𝑛/ 16 + 2 , 11 𝑛/ 16 + 2 ) · · · ( 11 𝑛/ 16 , 12 𝑛/ 16 ) ( 12 𝑛/ 16 + 1 , 13 𝑛/ 16 + 1 )( 12 𝑛/ 16 + 2 , 13 𝑛/ 16 + 2 ) · · · ( 13 𝑛/ 16 , 14 𝑛/ 16 ) ( 14 𝑛/ 16 + 1,15𝑛/16 +1)(14𝑛/16 +2,15𝑛/16 +2) · · · (15𝑛/16,16𝑛/16). If 𝑛= 32 , the let 𝑃5=( 1 , 2 )( 3 , 4 ) · · · ( 31 , 32 ) and we are ready. Else, by continuing the process in an analogous manner, for every 𝑛> 32 , where 𝑛 is a power of two, we get 𝑃1=( 1 ,𝑛/ 2 + 1 )( 2 ,𝑛/ 2 − 12 ) · · · (𝑛/ 2 ,𝑛), 𝑃2=( 1 ,𝑛/ 4 + 1 )( 2 ,𝑛/ 4 + 2 ) · · · (𝑛/ 4 , 2 𝑛/ 4 )( 2 𝑛/ 4 + 1 , 3 𝑛/ 4 + 1 )( 2 𝑛/ 4 + 2 , 3 𝑛/ 4 + 2 ) · · · ( 3 𝑛/ 4 , 4 𝑛/ 4 ), 𝑃3=( 1 ,𝑛/ 8 + 1 )( 2 ,𝑛/ 8 + 2 ) · · · (𝑛/ 8 , 2 𝑛/ 8 )( 2 𝑛/ 8 + 1 , 3 𝑛/ 8 + 1 ( 2 𝑛/ 8 + 2 , 3 𝑛/ 8 + 2 )) · · · ( 3 𝑛/ 8 , 4 𝑛/ 8 )( 4 𝑛/ 8 + 1 , 5 𝑛/ 8 + 1 )( 4 𝑛/ 8 + 2 , 5 𝑛/ 8 + 2 ) · · · ( 5 𝑛/ 8 , 6 𝑛/ 8 ) ( 6 𝑛/ 8 + 1 , 7 𝑛/ 8 + 1 )( 6 𝑛/ 8 + 2 , 7 𝑛/ 8 + 2 ) · · · ( 7 𝑛/ 8 , 8 𝑛/ 8 ), 𝑃4=( 1 ,𝑛/ 16 + 1 )( 2 ,𝑛/ 16 + 2 ) · · · (𝑛/ 16 , 2 𝑛/ 16 )( 2 𝑛/ 16 + 1 , 3 𝑛/ 16 + 1 ) ( 2 𝑛/ 16 + 2 , 3 𝑛/ 16 + 2 ) · · · ( 3 𝑛/ 16 , 4 𝑛/ 16 )( 4 𝑛/ 16 + 1 , 5 𝑛/ 16 + 1 ) ( 4 𝑛/ 16 + 2 , 5 𝑛/ 16 + 2 ) · · · ( 5 𝑛/ 16 , 6 𝑛/ 16 )( 6 𝑛/ 16 + 1 , 7 𝑛/ 16 + 1 ) ( 6 𝑛/ 16 + 2 , 7 𝑛/ 16 + 2 ) · · · ( 7 𝑛/ 16 , 8 𝑛/ 16 ) ( 8 𝑛/ 16 + 1,9𝑛/16 +1)(8𝑛/16 +2,9𝑛/16 +2) · · · (9𝑛/16,10𝑛/16)(10𝑛/16 + 1 , 11 𝑛/ 16 + 1 )( 10 𝑛/ 16 + 2 , 11 𝑛/ 16 + 2 ) · · · ( 11 𝑛/ 16 , 12 𝑛/ 16 )( 12 𝑛/ 16 + 1 , 13 𝑛/ 16 + 1 )( 12 𝑛/ 16 + 2 , 13 𝑛/ 16 + 2 ) · · · ( 13 𝑛/ 16 , 14 𝑛/ 16 )( 14 𝑛/ 16 + 1 , 15 𝑛/ 16 + 1 )( 14 𝑛/ 16 + 2 , 15 𝑛/ 16 + 2 ) · · · ( 15 𝑛/ 16 , 16 𝑛/ 16 ),· · · 𝑃𝑙𝑜𝑔2𝑛−1=( 1 , 3 )( 2 , 4 )( 5 , 7 )( 6 , 8 ) · · · (𝑛− 3 ,𝑛 − 1 )(𝑛− 2,𝑛), 𝑃𝑙𝑜𝑔2𝑛=(1,2)(3,4) · · · (𝑛−1,𝑛). For example, if 𝑛=16,then 𝑃1=(1,9)(2,10)(3,11)(4,12)(5,13)(6,14)(7,15)(8,16), 𝑃2=(1,5)(2,6)(3,7)(4,8)(9,13)(10,14)(11,15)(12,16), 𝑃3=(1,3)(2,4)(5,7)(6,8)(9,11)(10,12)(13,15)(14,16), 𝑃4=(1,2)(3,4)(5,6)(7,8)(9,10)(11,12)(13,14)(15,16). Let us define the automata B𝑖=(𝐴𝑛, 𝑋𝑛, 𝛿𝐵,𝑖) for each 𝑖∈ { 1 , . . . , 𝑙𝑜𝑔2𝑛} such that for any (the definition formula below specifies what was referred to above as the calculation "turns": based on the index 𝑗 a "with comma" or a "without comma" character component is used in the formula; the formula distinguishes the permutations according to indices 𝑎1, . . . , 𝑎𝑛∈𝐴, 𝑥1, . . . , 𝑥𝑛∈ 𝑋, 𝛿𝐵,𝑖 ((𝑎1, . . . , 𝑎𝑛),(𝑥1, . . . , 𝑥𝑛)) =(𝑎”1, . . . , 𝑎”𝑛),where 𝑎”1=𝛿(𝑎1, 𝛿 (𝑎𝑃𝑖(1), 𝑥𝑃𝑖(1) ) ), . . . 𝑎 ” 𝑗=𝛿(𝑎𝑗, 𝛿 (𝑎𝑃𝑖(𝑗), 𝑥𝑃𝑖(𝑗))), if 𝑗<𝑃𝑖(𝑗), and 𝑎 ” 𝑗=𝛿(𝑎𝑗, 𝛿(𝑎”𝑃𝑖(𝑗), 𝑥𝑃𝑖(𝑗))),if 𝑗≥𝑃𝑖(𝑗)(𝑗∈1, . . . , 𝑛), . . . 𝑎”𝑛=𝛿(𝑎𝑛, 𝛿 (𝑎”𝑃𝑖(𝑛), 𝑥𝑃𝑖(𝑛))). For example, if 𝑖= 1and 𝑃1=( 1 , 9 )( 2 , 10 )( 3 , 11 )( 4 , 12 )( 5 , 13 ) (6,14)(7,15)(8,16),then 𝑎”1=𝛿(𝑎1, 𝛿 (𝑎9, 𝑥9)),because 𝑃1(1)=9and 1<𝑃1(1)(=9), 𝑎 ” 2=𝛿(𝑎2, 𝛿 (𝑎10, 𝑥10)), because 𝑃1( 2 )= 10 and 2 <𝑃1( 2 )(= 10 ), . . . 𝑎 ” 8=𝛿(𝑎8, 𝛿 (𝑎16, 𝑥16)), because 𝑃1( 8 )= 16 and 8 <𝑃1( 8 )(= 16 ), 𝑎 ” 9=𝛿(𝑎9, 𝛿 (𝑎 ” 1, 𝑥1)), because 𝑃1( 9 )= 16 and 9 >𝑃1( 9 )(= 16 ), 𝑎 ” 10 =𝛿(𝑎10, 𝛿 (𝑎 ” 2, 𝑥2)), because 𝑃1( 10 )= 2and 10 >𝑃1( 10 )= 2, . . . 𝑎 ” 16 =𝛿(𝑎16, 𝛿 (𝑎 ” 8, 𝑥8)), because 𝑃1( 16 )= 8and 16 >𝑃1( 16 )(= 8). For the example included below, let us define the above such that the automaton B=(𝐴𝑛, 𝑋𝑛𝑙𝑜𝑔2𝑛, 𝛿𝐵) is defined such that for any 𝑎1, . . . , 𝑎𝑛∈𝐴, (𝑥1, . . . , 𝑥𝑛𝑙𝑜𝑔2𝑛) ∈ 𝑋𝑛𝑙𝑜𝑔2𝑛, the transition 𝛿𝐵((𝑎1, . . . , 𝑎𝑛),(𝑥1, . . . , 𝑥𝑛𝑙𝑜𝑔2𝑛)) is generated by first generating the state vector that can be obtained by applying the transition function 𝛿𝐵,1 for the vector (𝑎1, . . . , 𝑎𝑛) as a state, and for the vector (𝑥1, . . . , 𝑥𝑛) as an input signal. Taken to an expression: first the transitions 𝛿𝐵,1((𝑎1, . . . , 𝑎𝑛),(𝑥1, . . . , 𝑥𝑛)) are generated. Thereafter, the transition function 𝛿𝐵,2 is applied for the result of this transition and the vector (𝑥𝑛+1, . . . , 𝑥2𝑛) as an input signal. Taken to an expression: the transition 𝛿𝐵,3(𝛿𝐵,2(𝛿𝐵,1((𝑎1, . . . , 𝑎𝑛), (𝑥1, . . . , 𝑥𝑛)),(𝑥𝑛+1, . . . , 𝑥2𝑛)),(𝑥2𝑛+1, . . . , 𝑥3𝑛)) is generated (therefore, in the permutation approach the steps of this process have to be implemented applying the above-described permutations). This process is carried on in 𝑙𝑜𝑔2𝑛 steps, wherein, in the last step, the transition function 𝛿𝐵,𝑙𝑜𝑔2𝑛 is applied for the state vector obtained, and for the vector (𝑥𝑛(𝑙𝑜𝑔2𝑛−1), . . . , 𝑥𝑏𝑙𝑜𝑔2𝑛) as an input signal. Taken to an expression: the transition 𝛿𝐵,𝑙𝑜𝑔2𝑛(𝛿𝐵,𝑙𝑜𝑔2𝑛−1(· · · 𝛿𝐵,2(𝛿𝐵,1((𝑎1, . . . , 𝑎𝑛),(𝑥1, . . . , 𝑥𝑛)), (𝑥𝑛+1, . . . , 𝑥2𝑛)) . . . , 𝑥𝑛(𝑙𝑜𝑔𝑛−1)+1, . . . , 𝑥𝑛𝑙𝑜𝑔2𝑛)) is generated. The automaton B defined in such a manner is a scrambler automaton having 𝑙𝑜𝑔2𝑛 components that is determined by the automaton A, where the automaton A is the base automaton of the automaton B. By our definition, B is an automaton contolled by transpositions of its state components. So, for short, B will be called transposition-controlled automaton. It is also assumed that the transition matrix of the base automaton forms a Latin square. 4 THE NOVEL BLOCK CIPHER The scrambler unit B of the cipher is called transposition-controlled automaton, using which, receiving a character string 𝑤 having a length of power of two 𝑛 yields a string 𝑔(𝑤) (where 𝑔 : { 0 , 1 }𝑛→ { 0 , 1 }𝑛 is a function that bijectively maps the set of all possible strings of length 𝑛onto itself). The function g will hereinafter be referred to as a scrambler function. A matrix of which each row and each column is a permutation of the elements of 𝐻 is a Latin square over the set 𝐻. An automaton without output is an algebraic structure consisting of two non-empty sets, namely, the state set and the 31 Now, our example is continued. Example 1 (cont.) The language Lpis accepted by the PDATSBP M= ({a, b, c},{S, A, B, C}, δ, {(A, C),(A, B), (B, C),(C, B)}, S), where CBA ∈δ(λ, S),CBSA ∈δ(λ, S), λ∈δ(a, A),λ∈δ(b, B),λ∈δ(c, C). As a corollary of Theorems 1 and 2, we have the following important result. Corollary 1. There is a PDATSBP for every contextfree trace language that accepts it. 5. CONCLUSIONS Permutation grammars and the newly investigated variant of pushdown automata are capable to model parallel processes where the base language is context-free and also some linguistically important structures. It is known that the word problem for permutation grammars is NP-complete in general [21]. It is a future task to analyse what is the case for context-free trace languages or at least for those of them which may appear in modeling and in applications. Relation of these classes to languages accepted by pushdown automata with translucent letters [23, 24] should also be studied in the future. About the normal form for permutation grammars, it is an open problem whether the chain rules A→Bcan be eliminated. 6. REFERENCES [1] S. Crvenkovi´c, I. Dolinka and Z. ´ Esik. On equations for union-free regular languages. Inform. and Comput. 164/1 pp. 152–172, 2001. [2] J. Dassow and Gh. P˘aun. Regulated rewriting in formal language theory. Springer-Verlag, Berlin, 1989. [3] V. Diekert and G. Rozenberg (eds.). The book of traces. World Scientific, River Edge, NJ, USA, 1995. [4] K. Fogarasi and B. Nagy. A nondeterministic parser for Perm2grammars. Abstract volume of the 10th Joint Conference on Mathematics and Computer Science (MaCS), Cluj-Napoca, Romania, 2014. [5] M. A. Harrison. Introduction to formal language theory, Addison-Wesley, Reading, AM, USA, 1978. [6] T. Herendi and B. Nagy. Parallel Approach of Algorithms. Typotex, Budapest, 2014. [7] R. Janicki, J. Kleijn, M. Koutny and L. Mikulski. Paradigms of Concurrency – Observations, Behaviours, and Systems – a Petri Net View. Studies in Computational Intelligence 1020, Springer, 2022. [8] G. Madejski. Infinite hierarchy of permutation languages. Fundam. Inform. 130/3 pp. 263–274, 2014. [9] E. M¨ akinen. On permutative grammars generating context-free languages. BIT 25/4 pp. 604–610, 1985. [10] A. Mateescu, G. Rozenberg and A. Salomaa. Shuffle on trajectories: syntactic constraints. Theoretical Computer Science 197 pp. 1–56, 1998. [11] A. W. Mazurkiewicz. Trace Theory. In: Petri Nets: Central Models and Their Properties, Advances in Petri Nets, Part II, LNCS, 255, pp. 279–324, 1986. [12] B. Nagy. Union-free regular languages and 1-cycle-free-path-automata. Publ. Math. Debrecen 68 pp. 183–197, 2006. [13] B. Nagy. Languages generated by context-free and type AB →BA rules. In: Proc. CINTI 2007: 8th Int. Symp. of Hung. Research. Comput. Intelligence Inf., Budapest, Hungary, pp. 563–572, 2007. [14] B. Nagy. Languages generated by context-free and type AB →BA rules. J Autom. Lang. Combin. 14 pp. 175–186, 2009. [15] B. Nagy. Permutation languages in formal linguistics. In: Proc. IWANN 2009, Part I, LNCS 5517 pp. 504—511, 2009. [16] B. Nagy. On a hierarchy of permutation languages. In: Automata, Formal Languages and Algebraic Systems, pp. 163–178, World Scientific, Singapore, 2010. [17] B. Nagy. Derivation Trees for Context-Sensitive Grammars. In: Automata, Formal Languages and Algebraic Systems, pp. 179–199, World Scientific, Singapore, 2010. [18] B. Nagy. 50→30Sensing Watson-Crick Finite Automata. In: G. Fung, ed.: Sequence and Genome Analysis II - Methods and Applications, iConcept Press, pp. 39—56, 2010. [19] B. Nagy. Linguistic power of permutation languages by regular help. In: Bio-inspired models for natural and formal languages. pp. 135–152, Cambridge Scholars, 2011. [20] B. Nagy. On the NP-completeness of the word problem for permutation grammars. In: Abstract volume of the Int. Workshop in honor of Masami Ito’s 77th birthday and P´al D¨ om¨ osi’s 75th birthday: DLT’s Satellite Workshop in Kyoto, Japan, September 2018. [21] B. Nagy. On the Membership Problem of Permutation Grammars – A Direct Proof of NP-Completeness. Int. J. Found. Comput. Sci. 31/4 pp. 515–525, 2020. [22] B. Nagy and A. A. Akkeles. Trajectories and Traces on Non-traditional Regular Tessellations of the Plane. In: Combinatorial Image Analysis, 18th Int. Workshop, IWCIA 2017, LNCS 10256, pp. 16–29, 2017. [23] B. Nagy and F. Otto. An Automata-Theoretical Characterization of Context-Free Trace Languages. In: SOFSEM 2011: 37th Conference on Current Trends in Theory and Practice of Computer Science, LNCS 6543, pp. 406–417 2011. [24] B. Nagy and F. Otto. CD-systems of stateless deterministic R(1)-automata governed by an external pushdown store. RAIRO Theor. Informatics Appl. 45 pp. 413–448, (2011). [25] B. Nagy and S. Parchami. On deterministic sensing 50→30Watson–Crick finite automata: a full hierarchy in 2detLIN. Acta Informatica 58 pp. 153–175, 2021. [26] S. Parchami and B. Nagy. Deterministic Sensing 50→30Watson-Crick Automata Without Sensing Parameter. In: Unconventional Computation and Natural Computation, UCNC 2018, LNCS 10867, Springer, pp. 173–187, 2018. [27] M. Penttonen. One-sided and two-sided context in formal grammars. Inf. Control 25 pp. 371–392, 1974. [28] R. Schott and J. C. Spehner. Two optimal parallel algorithms on the commutation class of a word Theoretical Computer Science 324 pp. 107–131, 2004. [29] G. Rozenberg and A. Salomaa (eds.). Handbook of formal languages. Springer-Verlag, Berlin, 1997. 38 Towards a Category-Theoretic Informatics Model of PSPP Linkages in Biomaterials Sylvert Prian Tahalea [email protected]szeged.hu University of Szeged, Hungary Universitas Pembangunan Nasional Veteran Yogyakarta, Indonesia Miklós Krész [email protected] InnoRenew CoE, UP IAM and UP FAMNIT, University of Primorska, Slovenia University of Szeged, Hungary Abstract This paper introduces a category-theoretic framework for modeling the causal and structural relationships in the ProcessingStructure-Property-Performance (PSPP) paradigm, with a case study of bioluminescent bacterial interactions on engineered surfaces. The composition and interdependence of material behavior in mathematical form are captured by defining the stages of PSPP linkages into objects and their transformations as morphisms. This model is applied to a case study involving bacterial cultivation on three different engineered surfaces. The universal input configurations and multi-objective performance targets are identified using limit and colimit, respectively. Additionally, the functors and natural transformations applied compare the bacterial behaviors across material categories, illustrating the structural relationship between bacteria. While the current model assumes deterministic and discrete transitions, it opens pathways for future development using enriched, fibred, and computational categorical tools. This work demonstrates the potential of category theory as a unifying language for scientific modeling in biomaterials research. Keywords category theory, PSPP, biomaterials, natural transformations, functor 1 Introduction The property-structure-processing-performance (PSPP) linkages are a core paradigm in material science to conceptualise the interdependence between material properties, processing methods, structural features, and resulting performance [12, 8, 1]. The PSPP linkages come from process-structure-property-performance relationships, which build up like a chain in a bottom-up (simulationdriven approach) or top-down (process design approach) [7]. Process refers to the methods and techniques used to transform the raw materials into finished products with desired shapes and properties; structure refers to the arrangements of atoms, molecules, or phases within the materials; properties are the measurable responses of a material to the stimuli and can be classified into various categories; while performance is the materials’ ability to function effectively in a specific application, considering its properties and conditions [7, 6, 10, 23]. Despite its importance, PSPP linkages are described in an informal and domain-specific manner, which hinders systematic integration across disciplines and limits its computational implementation [15, 22]. This lack Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). MATCOS-25, 9–10 October 2025, Koper, Slovenia ©2024 Copyright held by the owner/author(s). of formalism becomes a barrier when attempting to analyse the PSPP relationships at scale, particularly in data-driven or interdisciplinary contexts [25, 11]. Recent advances in material informatics have highlighted the importance of formal representation for enabling computational reasoning, interoperability, and data integration [2, 24, 14, 7]. In parallel, category theory has emerged as a mathematical language for abstraction and composition, providing a rigorous way to describe the structures and relationships between them [19]. Originally developed in mathematics, category theory has found its way to be implemented in computer science, biology, data science, where it serves to represent the complex interactions [18, 4]. The category theory implementation in informatics offers a pathway to move beyond the classical representation of materials knowledge, such as graphs and ontologies, to ensure the consistency, scalability, and interoperability in computational models of materials science. It has been recognised as a unifying language for science, offering tools such as objects, morphisms, limits, and colimits to formally capture relationships between complex systems [18, 4]. To address this gap, a categorical model is proposed in this paper to formalise the PSPP interconnections in a structured and compositional way. Our contribution lies in defining this formal representation, demonstrating how it can unify disparate descriptions into a coherent framework. The scope of this work is focused on establishing the formal model rather than developing a full computational implementation. To illustrate its potential, a small example that highlights how categorical structure captures PSPP linkages is provided. The remainder of the paper is organised as follows. Section 2 explains the methodology, section 3 presents the case study example, section 4 concludes the contribution, limitations, and future research. 2 Methodology This section explains the category, subcategories, functors, natural transformations, limits, and colimits of the PSPP linkages. 2.1 PSPP Linkage PSPP illustrates the chain of relationships that connect processing parameters to the resulting structure (microstructure), which governs properties and thus impacts component performance—a critical multiscale, multiphysics mapping essential for the modeling and design of materials [7]. Processing techniques, resultant structures, material qualities, and performance outcomes are all PSPP stages that are frequently examined separately or depicted informally using diagrams that lack formal semantics and mathematical rigor [7, 26, 3]. The PSPP framework helps to formalize the causal pathway from how a material is processed to its performance outcome, e.g. bacteria-surface interaction for bioluminescence, which is later used as the study case. The bacteria-surface 39 MATCOS-25, 9–10 October 2025, Koper, Slovenia Tahalea & Krész interaction involves a complex and multilayered relationship between the processing method and the metabolic output. However, current approaches for modeling these interactions suffer from several limitations, such as fragmented representation, lack of a unified framework, and lack of composability, which can be addressed with a category-theoretical model. The PSPP linkage presented in Figure 1 shows how a specific processing method (e.g., oxygen plasma treatment) alters the surface structure, which in turn affects bacterial interaction properties (such as adhesion or viability), ultimately leading to a measurable performance outcome like bioluminescence intensity. Figure 1: PSPP linkages in bacteria-surface interaction. 2.2 PSPP Category The category theoretic model of PSPP is defined as follows. 2.2.1 Objects. Let C be a category. The set of the objects in C is defined as: Ob(C) ={𝑃, 𝑆, 𝑃𝑟, 𝑃𝑒} where: 𝑃 represents the manufacturing process or processing in short form, 𝑆 represents materials structure, 𝑃𝑟 represents material properties, and 𝑃𝑒 represents material performance. Thus, Ob(C) encodes the fundamental entities of the PSPP framework as objects within the categorical structure. 2.2.2 Morphisms. Define the set of morphisms (arrows) in the category Cas follows: 𝐻𝑜𝑚( C) (𝑋, 𝑌) Where 𝑋,𝑌 ∈𝑃, 𝑆, 𝑃𝑟, 𝑃𝑒 and each morphism represents a relationship or transformation from one object to another. The morphisms in the category Care defined as follows: 𝑓:𝑃→𝑆(Processing affect Structure) 𝑔:𝑆→𝑃𝑟 (Structure determine Properties) ℎ:𝑃𝑟 →𝑃𝑒 (Properties dictate Performance) 𝑖:𝑃𝑒 →𝑃(Performance feedback to Processing) 𝑗:𝑃→𝑆′(Processing results Modified Structure) 𝑘:𝑆′→𝑃𝑟 (Modified Structure determine Properties) 𝑙:𝑃𝑟 →𝑃𝑟′(Properties transform to Modified Properties) 𝑚:𝑃𝑟′→𝑃𝑒 (Modified Properties dictate Performance) 2.2.3 Composition. The composition of morphisms in the category Cfollows the associativity and identity properties: (ℎ◦𝑔) ◦ 𝑓=ℎ◦ (𝑔◦𝑓) and for each object 𝑋∈Ob(C) , there exists an identity morphism id𝑋:𝑋→𝑋such that: id𝑌◦𝑓=𝑓and 𝑔◦id𝑋=𝑔 In example for the bacterial surface interaction case, P = “oxygen plasma processing”, S = “hydrophilic glass surface”, Pr = “moderate bacterial adhesion, low EPS density”, Pe = “5 𝜇 W/cm 2 luminescence output”. Define the morphisms in category CPSPP as: 𝑓:𝑃→𝑆, ,𝑔 :𝑆→𝑃𝑟, , and ℎ:𝑃𝑟 →𝑃𝑒 Then the composite morphisms are: 𝑔◦𝑓:𝑃→𝑃𝑟 and ℎ◦𝑔:𝑆→𝑃𝑒 This composite morphism represents the entire PSPP causal chain from processing to performance. For this case, it maps: “plasma treatment” −→ “5 𝜇W/cm2luminescence” via structure and property stages. The identity morphism at each stage (e.g., id𝑆 : 𝑆→𝑆 ) satisfies: 𝑔◦id𝑆=𝑔, id𝑃𝑟◦𝑔=𝑔 confirming the category’s identity and associativity laws. 2.2.4 Diagram Representation. A typical morphism of the PSPP category is only represented by processing, structure, properties, and performance. However, the processing can result in modified structure and properties, which are shown in Figure 2. Figure 2: PSPP diagram representation including modified structure and properties. 2.3 Subcategories of PSPP Model The category model of PSPP linkage has several subcategories by nature (as presented in Figure 2), which are meaningful subsets of the objects and morphisms that still satisfy the definition of a category. 2.3.1 Canonical Path Subcategory. The main causal chain C𝑐𝑜𝑟𝑒 without branching. •Objects. Ob(C𝑐𝑜𝑟𝑒 )=𝑃, 𝑆, 𝑃𝑟, 𝑃𝑒 •Morphisms.𝑓:𝑃→𝑆,𝑔:𝑆→𝑃𝑟,ℎ:𝑃𝑟 →𝑃𝑒 This paper will be focusing only on this canonical path or the main linkages. 2.3.2 Modified Path Subcategory. This modified subcategory C𝑚𝑜𝑑 is an expansion of the canonical path that includes only the branching route. •Objects. Ob(C𝑐𝑜𝑟𝑒 )=𝑃, 𝑆, 𝑆′, 𝑃𝑟, 𝑃𝑟′, 𝑃𝑒 • Morphisms. 𝑓 : 𝑃→𝑆 , 𝑗 : 𝑃→𝑆′ , 𝑔 : 𝑆→𝑃𝑟 , 𝑘:𝑆→𝑃𝑟′,𝑙:𝑃𝑟 →𝑃𝑟′,𝑚:𝑃𝑟′→𝑃𝑒 2.3.3 Feedback Loop Subcategory. This loop subcategory C𝑙𝑜𝑜𝑝 focuses on feedback interactions within the model. •Objects. Ob(C𝑐𝑜𝑟𝑒 )=𝑃, 𝑆, 𝑃𝑟, 𝑃𝑒 •Morphisms.𝑓 ,𝑔,ℎ, 𝑖 :𝑃𝑒 →𝑃 40 Category-Theoretic Modeling of PSPP Linkages MATCOS-25, 9–10 October 2025, Koper, Slovenia 2.3.4 Application Specific Subcategory. This subcategory can be expanded based on the application focus field. The application on bacteria-surface interaction will be used as the case study in this article, illustrated by following examples. •Biofilm subcategory C𝑏𝑖𝑜𝑓 𝑖𝑙𝑚, with objects such as: 𝑃, 𝑆, 𝑃𝑟𝑎𝑑ℎ𝑒𝑠𝑖𝑜𝑛, 𝑃𝑒𝑏𝑖𝑜 𝑓 𝑖𝑙𝑚 • Bioluminescence subcategory C𝑏𝑖𝑜𝑙𝑢𝑚 , with objects such as: 𝑃, 𝑆, 𝑃𝑟𝑙𝑖𝑔ℎ𝑡, 𝑃𝑒𝑙𝑢𝑚 2.4 Functor and Natural Transformation To define a functor and natural transformation of this model, first, we need to define another category which can be processed with this functor and natural transformation logic. This subsection was prepared in advance to preserve the relation and transformation. 2.4.1 Functor. A functor 𝐹 : C1→ C2 between categories C1 and C2is defined by: (1) Object mapping: for each object 𝑋∈Ob(C1) , there is object 𝐹(𝑋) ∈ Ob(C2) (2) Morphism mapping: for each morphism 𝑓 : 𝑋→𝑌∈ HomC1(𝑋,𝑌 ) , there is exist a morphism 𝐹(𝑓) : 𝐹(𝑋) → 𝐹(𝑌) ∈ HomC2(𝐹(𝑋), 𝐹 (𝑌)) such that: •Composition is preserved: 𝐹(𝑓◦𝑔)=𝐹(𝑓) ◦ 𝐹(𝑔),∀𝑓 ,𝑔 ∈Mor(C1) •Identities are preserved: 𝐹(id𝑋)=id𝐹(𝑋),∀𝑋∈Mor(C1) Functor of PSPP category model example in our use case: Considering two distinct PSPP categories: • C𝑔𝑙𝑎𝑠𝑠 : a category of PSPP stages for Pseudomonas flourescens treated on glass surface. • C𝑤𝑜𝑜𝑑 : a category of PSPP stages for Pseudomonas flourescens on functionalized wood. Then we can define Functor 𝐹:C𝑔𝑙𝑎𝑠𝑠 → C𝑤𝑜𝑜𝑑 : •Objects –𝑃𝑔𝑙𝑎𝑠𝑠 →𝑃𝑤𝑜𝑜𝑑 : oxygen plasma cleaning → enzymbased-pretreatement –𝑆𝑔𝑙𝑎𝑠𝑠 →𝑆𝑤𝑜𝑜𝑑 : smooth hydropphilic silica → porous lignovellulosic surface –𝑃𝑟𝑔𝑙𝑎𝑠𝑠 →𝑃𝑟𝑤𝑜𝑜𝑑 : low biofilm density → moderate biofilm density –𝑃𝑒𝑔𝑙𝑎𝑠𝑠 →𝑃𝑒𝑤𝑜𝑜𝑑 : weak bioluminescence → stable bioluminescence overtime •Morphisms –𝑓𝑔𝑙𝑎𝑠𝑠 →𝑓𝑤𝑜𝑜𝑑 : plasma alters hydrophilicity → enzymatic treatement alters lignin exposure –𝑔𝑔𝑙𝑎𝑠𝑠 →𝑔𝑤𝑜𝑜𝑑 : smooth surface limit adhesion → roughness enhances microbial anchoring –ℎ𝑔𝑙𝑎𝑠𝑠 →ℎ𝑤𝑜𝑜𝑑 : low metabolic activity → higher growth and igh emission –𝑖𝑔𝑙𝑎𝑠𝑠 →𝑖𝑤𝑜𝑜𝑑 : bioluminescence feedback informas wood resin optimization Having established the functor, a later section presents the use case of the PSPP model. 2.4.2 Natural Transformation. Let 𝐹,𝐺 : C1→ C2 be functors. A natural transformation 𝜂 : 𝐹⇒𝐺 is a family of morphisms 𝜂𝑋 : 𝐹(𝑋) → 𝐺(𝑋),∀𝑋∈Ob(C1), such that for every morphism 𝑓:𝑋→𝑌in C1,𝜂𝑌◦𝐹(𝑓)=𝐺(𝑓) ◦ 𝜂𝑋. Natural transformation of PSPP category model of our use case: Suppose there are 𝐹,𝐺 : C𝑐𝑜𝑛𝑐𝑟𝑒𝑡𝑒 → C𝑏𝑎𝑐𝑡 two functors: •𝐹: using Pseudomonas flourescence •𝐺: using Serratia marcescens The natural transformation 𝜂 : 𝐹⇒𝐺 , for each PSPP stage (object) transformation, comes from one species to another. 2.5 Limit and Colimit 2.5.1 Cone. Let J be a category, C a category, and let 𝐷 : J → C be a functor (called a diagram in C ). A cone over 𝐷 consists of: (1) An object 𝑁∈Ob(C), called the apex of the cone, (2) A family of morphism in C , {𝜙𝑗 : 𝑁→𝐷(𝑗)}𝑗∈Ob( J) , such that for every morphism 𝑓 : 𝑗→𝑘 in J , the following commutative condition holds: 𝐷(𝑓) ◦ 𝜙𝑗=𝜙𝑘 In other words, a cone points into every object in the functor 𝐷from the apex using the morphisms. Cone of PSPP model of our use case: Let there be three different processing routes for bioluminescent bacterial coating: 𝐹1 : Processing 1, 𝐹2 : Processing 2, and 𝐹3 : Processing 3; which produce a full chain 𝑃𝑖 𝑓𝑖 −→ 𝑆𝑖 𝑔𝑖 −−→ 𝑃𝑟𝑖 ℎ𝑖 −−→ 𝑃𝑒𝑖 and the functor 𝐹 of this output can be expressed as: 𝐹:J → C, where 𝐹(𝑗)=𝑃𝑒𝑗. 2.5.2 Limit. Let 𝐹 : J → C be a functor. A limit of 𝐹 is an object 𝐿∈Ob(C) together with morphisms 𝜋𝑗 : 𝐿→𝐹(𝑗),∀𝑗∈ Ob(J),such that for every morphism 𝑓:𝑗→𝑗′in J, 𝐹(𝑓) ◦ 𝜋𝑗=𝜋𝑗′ Moreover, (𝐿, 𝜋𝑗) is universal: for any other cone {𝜓𝑗 : 𝑁→ 𝐹(𝑗)}𝑗∈Ob( J) , there exists a unique morphism 𝑢 : 𝑁→𝐿 such that 𝜓𝑗=𝜋𝑗◦𝑢, ∀𝑗∈Ob(J). Limit of PSPP Model of our use case: The output of the PSPP linkage lies at the performance stage; the subdiagram from processing to performance can be used to find the optimal solution through the PSPP model. Assume that there are three different processing routes: (1) 𝐹 1 : 𝑃1→𝑆1→𝑃𝑟1→𝑃𝑒1 , (2) 𝐹 2 : 𝑃2→𝑆2→𝑃𝑟2→𝑃𝑒2 , and (3) 𝐹 1 : 𝑃3→𝑆3→𝑃𝑟3→𝑃𝑒3 . The limit object 𝐿 is universal, thus there exist projections to each case: 𝜋𝑖:𝐿→𝑃𝑒𝑖for 𝑖=1,2,3 2.5.3 Colimit. Let 𝐹 : J → C be a functor. A colimit of 𝐹 is an object 𝐶∈Ob(C) together with morphisms 𝜄𝑗 : 𝐹(𝑗) → 𝐶, ∀𝑗∈Ob(J), such that for every morphism 𝑓 : 𝑗→𝑗′ in J , 𝜄𝑗′◦𝐹(𝑓)=𝜄𝑗 . Moreover, (𝐶, 𝜄𝑗) is universal: for any other cocone {𝜓𝑗 : 𝐹(𝑗) → 𝑁}𝑗∈Ob( J) , there exists a unique morphism 𝑣 : 𝐶→𝑁such that 𝜓𝑗=𝑣◦𝜄𝑗,∀𝑗∈Ob(J). Colimit of PSPP Model of our use case: In the study of the PSPP linkage, suppose that there are three different properties, as follows. (1) 𝐹1 : 𝑃𝑟𝐴→𝑃𝑒𝐴 , (2) 𝐹2 : 𝑃𝑟𝐵→𝑃𝑒𝐵 , (3) 𝐹3 : 𝑃𝑟𝐶→𝑃𝑒𝐶 ; one of the objectives is to merge or combine all the performances to produce a unified performance outcome 𝜄𝑖 : 𝑃𝑒𝑖→𝐶 and any projection results factors uniquely through 𝐶. 3 Case Study: Optimising Surface Design for Bioluminescent Bacteria The case study of this model is to determine the best surface configuration to support bioluminescence output from Pseudomonas flourescens. There are three different surface preparation routes 41 MATCOS-25, 9–10 October 2025, Koper, Slovenia Tahalea & Krész [16, 13, 17, 9, 20, 5, 21] using the PSPP framework and interpreting the model categorically, such as glass, wood, and concrete. 3.1 Experimental Pipelines Each processing route is a distinct PSSP chain. (1) Oxygen plasma on glass: 𝑃1 : Oxygen plasma cleaning; 𝑆1 : Enhanced surface hydrophilicity and cleanliness; 𝑃𝑟1 : Reduced adhesion of bacteria; 𝑃𝑒1 : Produce a low bioluminescence signal caused by the low density of biofilm. (2) Enzyme treatment on wood: 𝑃2 : Enzymatic pretreatment of the surface; 𝑆2 : Surface may degrade and reduce the cohesion; 𝑃𝑟2 : Low biofilm viability; 𝑃𝑒2 : Low signal output due to diminished bacterial presence. (3) Sol-gel coated concrete: 𝑃3 : Sol-gel encapsulation; 𝑆3 : Potentially surrounded by silica gel, contains methanol; 𝑃𝑟3 : Severely limits bacteria’s metabolic activity; 𝑃𝑒3 : Light output reduced drastically. The PSPP linkage on this study case is presented in Figure 1. 3.2 Categorical diagram Given that each route forms a PSPP morphism chain: 𝑃𝑖 𝑓𝑖 −→ 𝑆𝑖 𝑔𝑖 −−→ 𝑃𝑟𝑖 ℎ𝑖 −−→ 𝑃𝑒𝑖, which is presented in Figure 3. Figure 3: The routes of PSPP morphism chain 3.3 Limit Construction A universal configuration 𝐿 (e.g., a hybrid substrate combining features of all three processes) is defined as a cone in Figure 4. Universal configuration 𝐿 represents an ideal design that projects Figure 4: 𝐿 is the universal input that relates to all observed performances to each observed performance, supporting general optimization. The categorical properties are as follows. •𝐿→𝑃𝑒𝑖form morphims • Commutativity must hold: any internal morphism between 𝑃𝑒𝑖must be maintained if they exist •𝐿 is the limit and universal: every cone factors uniquely through 𝐿 3.4 Colimit Construction Conversely, a colimit object 𝐶 which integrates all outcomes into a general-purpose performance target should be defined as presented in Figure 5. 𝐶 serves as a composite performance Figure 5: A cocone with apex 𝐶 absorbing all performances output specification, capturing multi-objective behaviors (e.g., balancing intensity and longevity) with the following categorical properties. •Morphisms 𝜄𝑖:𝑃𝑒𝑖→𝐶 •𝐶 is the colimit of every projection from 𝑃𝑒𝑖 factors uniquely through 𝐶 3.5 Natural transformation Suppose there is another experiment with the same three surfaces but different bacteria, i.e. Serratia marcescens. The functors and the natural transformation are presented in Figure 6 and can be defined as follows. • Functors: 𝐹 : PSPP route using P. flourescens and 𝐺 : PSPP route using S. marcescens • Natural transformation: 𝜂 : 𝐹⇒𝐺 , where 𝜂𝑃 : translate the process (e.g plasma treatment) parameters, 𝜂𝑆 : map structure adaptation, and 𝜂𝑃𝑟,𝜂𝑃𝑒 : transform properties and performances across bacterial strains. Figure 6: Natural transformation from P. flourescence to S. marcescens 3.6 Discussion The categorical analysis of PSPP linkages in bioluminescent bacterial systems provides a structured lens to evaluate and compare multiple experimental design routes. In this case study, three distinct processing strategies—oxygen plasma treatment on glass, enzyme pretreatment on wood, and sol-gel silica encapsulation on concrete—were formalised as morphism chains within the category CPSPP . Each route defines a unique instance of a functor mapping the canonical PSPP stages (Processing → Structure → Property → Performance). The resulting categorical diagram allowed us to analyse their behaviours both individually and relationally. The introduction of a limit object 𝐿 , structured as a cone over the performance outcomes 𝑃𝑒1 , 𝑃𝑒2 , and 𝑃𝑒3 , offered a formal 42 Category-Theoretic Modeling of PSPP Linkages MATCOS-25, 9–10 October 2025, Koper, Slovenia approach to identifying an optimal or universal input configuration capable of reproducing all observed performance behaviors. Practically, 𝐿 can be interpreted as a generalized hybrid surface or processing strategy that maintains compatibility with the range of observed bioluminescence behaviors. Conversely, the dual construction of a colimit object 𝐶 , defined via morphisms 𝜄𝑖 : 𝑃𝑒𝑖→𝐶 , enabled the synthesis of a composite performance profile, which may serve as a design target in multi-objective optimization tasks. This dual analysis captures generalisation (limit) and integration (colimit), offering a flexible modelling approach. The model supports the comparative approach through natural transformation between bacteria through functors. For example, transitions between Pseudomonas fluorescens and Serratia marcescens can be formalised as structure-preserving morphism families, highlighting how microbial adaptation maps across processing contexts. This indicates category theory’s ability to enable reasoning and potential computational automation in material informatics. 4 Conclusion With a brief case study of bioluminescent bacterial interactions on artificial surfaces, this paper offers a category-theoretic framework for modelling PSPP linkage in biomaterials. In order to allow compositional thinking, a formal categorical structure CPSPP is designed to represent PSPP stages as objects and causal relationships as morphisms. The model captures both variability and unification across various experimental settings by using functors, natural transformations, and universal constructions like limits and colimits. A case study of Pseudomonas fluorescens across three surface treatments shows the effectiveness of this approach in determining optimal input configurations (by limits) and composite performance targets (via colimits). The application of natural transformations furthers biological comparison across strains. The main contributions of this work are: (1) a formal mathematical structure for PSPP links based on category theory and (2) the integration of experimental design with universal structures like cones and cocones. However, there are some limits. The current model assumes discrete and predictable transitions between PSPP phases, does not account for uncertainty, dynamics, or probabilistic behavior, and lacks an implemented computational instance functor to real data. Future work will focus on enriching the categorical model with quantitative structure (e.g., cost, confidence, time) via enriched categories; defining pullbacks and pushouts for constraints. Acknowledgements The research was supported by the BioLOG project: the second author is grateful for the support of the National Centre of Science (NCN) through grant DEC-2020/39/I/HS4/03533, the Slovenian Research and Innovation Agency (ARIS) through grant N1-0223, and the Austrian Science Fund (FWF) through grant I 5443-N. This work is supported by the ARIS research program P1-0404 and by the research program CogniCom (0013103) at the University of Primorska. References [1] Ankit Agrawal and Alok Choudhary. 2016. Perspective: materials informatics and big data: realization of the “fourth paradigm” of science in materials science. Apl Materials, 4, 5. doi: 10.1063/1.4946894. [2] Toshihiro Ashino. 2010. Materials ontology: an infrastructure for exchanging materials information and knowledge. Data Science Journal, 9, 54–61. doi: 10.2481/dsj.008-041. [3] Igor Baskin and Yair Ein-Eli. 2022. Electrochemoinformatics as an emerging scientific field for designing materials and electrochemical energy storage and conversion devices—an application in battery science and technology. Advanced Energy Materials, 12, 48, 2202380. doi: 10.1002/aenm.202202380. [4] Brendan Fong and David I Spivak. 2018. Seven sketches in compositionality: an invitation to applied category theory. arXiv preprint arXiv:1803.05316. doi: 10.48550/arXiv.1803.05316. [5] Wissam Ghach, M. Etienne, V. Urbanová, Frédéric P. A. Jorand, and A. Walcarius. 2014. Sol–gel based ‘artificial’ biofilm from pseudomonas fluorescens using bovine heart cytochrome c as electron mediator. Electrochemistry Communications, 38, 71–74. doi: 10.1016/J.ELECOM.2013.11.001. [6] D. Gu, Xinyu Shi, R. Poprawe, D. Bourell, R. Setchi, and Jihong Zhu. 2021. Material-structure-performance integrated laser-metal additive manufacturing. Science, 372. doi: 10.1126/science.abg1487. [7] Seyed Mahdi Hashemi, Soroush Parvizi, Haniyeh Baghbanijavid, Alvin TL Tan, Mohammadreza Nematollahi, Ali Ramazani, Nicholas X Fang, and Mohammad Elahinia. 2022. Computational modelling of process–structure– property–performance relationships in metal additive manufacturing: a review. International Materials Reviews, 67, 1, 1–46. doi: 10.1080/09506608.2 020.1868889. [8] Surya R Kalidindi and Marc De Graef. 2015. Materials data science: current status and future outlook. Annual Review of Materials Research, 45, 1, 171– 193. doi: 10.1146/annurev-matsci-070214-020844. [9] A. Kotlobay et al. 2018. Genetically encodable bioluminescent system from fungi. Proceedings of the National Academy of Sciences of the United States of America, 115, 12728–12732. doi: 10.1073/pnas.1803615115. [10] Wenqi Li and Jian Shi. 2023. Lignin-derived carbon material for electrochemical energy storage applications: insight into the process-structureproperties-performance correlations. Frontiers in Bioengineering and Biotechnology, 11. doi: 10.3389/fbioe.2023.1121027. [11] Zhanzhao Li, Te Pei, Weichao Ying, Wil V Srubar III, Rui Zhang, Jinyoung Yoon, Hailong Ye, Ismaila Dabo, and Aleksandra Radlińska. 2024. Can domain knowledge benefit machine learning for concrete property prediction? Journal of the American Ceramic Society, 107, 3, 1582–1602. doi: 10.1111/jac e.19549. [12] Gregory B Olson. 1997. Computational design of hierarchically structured materials. Science, 277, 5330, 1237–1242. doi: 10.1126/science.277.5330.1237. [13] Rahman Rahmanpour and T. Bugg. 2015. Characterisation of dyp-type peroxidases from pseudomonas fluorescens pf-5: oxidation of mn(ii) and polymeric lignin by dyp1b. Archives of biochemistry and biophysics, 574, 93–8. doi: 10.1016/j.abb.2014.12.022. [14] Balashanmuga Priyan Rajamohan et al. 2025. Materials data science ontology (mds-onto): unifying domain knowledge in materials and applied data science. Scientific Data, 12, 1, 628. doi: 10.1038/s41597-025-04938-5. [15] Krishna Rajan. 2005. Materials informatics. Materials Today, 8, 10, 38–45. doi: 10.1016/S1369-7021(05)71123-8. [16] Z. Remeš, O. Babčenko, Vítězslav Jarý, and K. Beranová. 2024. Enhanced photoluminescence of plasma-treated recycled glass particles. Nanomaterials, 14. doi: 10.3390/nano14131091. [17] Francis W. M. R. Schwarze et al. 2024. Taming the production of bioluminescent wood using the white rot fungus desarmillaria tabescens. Advanced Science, 11. doi: 10.1002/advs.202403215. [18] David I Spivak. 2014. Category theory for the sciences. MIT press. [19] David I Spivak, Tristan Giesa, Elizabeth Wood, and Markus J Buehler. 2011. Category theoretic analysis of hierarchical protein materials and social networks. PloS one, 6, 9, e23911. doi: 10.1371/journal.pone.0023911. [20] Aisha J Syed and James C. Anderson. 2021. Applications of bioluminescence in biotechnology and beyond. Chemical Society reviews.doi: 10.1039/d0cs01 492c. [21] J. Trögl, G. Kuncová, and P. Kurán. 2010. Bioluminescence of pseudomonas fluorescens hk44 in the course of encapsulation into silica gel. effect of methanol. Folia Microbiologica, 55, 569–575. doi: 10.1007/s12223-010-00919. [22] Kewei Wang et al. 2024. Deep learning based inverse modeling for materials design: from microstructure and property to processing. In 2024 international conference on machine learning and applications (ICMLA). IEEE, 236–241. doi: 10.1109/ICMLA61862.2024.00038. [23] Yayun Wang, S. Naleway, and Bin Wang. 2020. Biological and bioinspired materials: structure leading to functional and mechanical performance. Bioactive Materials, 5, 745–757. doi: 10.1016/j.bioactmat.2020.06.003. [24] Tian Xie and Jeffrey C Grossman. 2018. Crystal graph convolutional neural networks for an accurate and interpretable prediction of material properties. Physical review letters, 120, 14, 145301. doi: 10.1103/PhysRevLett.120.145301. [25] Han Zhang, Runsheng Li, Junjiang Liu, Kaiyun Wang, Qian Weijian, Lei Shi, Liming Lei, Weifeng He, and Shengchuan Wu. 2024. State-of-art review on the process-structure-properties-performance linkage in wire arc additive manufacturing. Virtual and Physical Prototyping, 19, 1, e2390495. doi: 10.10 80/17452759.2024.2390495. [26] Olga Zinovieva, Varvara Romanova, Ekaterina Dymnich, Aleksandr Zinoviev, and Ruslan Balokhonov. 2023. A review of computational approaches to the microstructure-informed mechanical modelling of metals produced by powder bed fusion additive manufacturing. Materials, 16, 19, 6459. doi: 10.3390/ma16196459. 43 Cost-Sensitive Overview of Model Ensembling for Machine-Generated Text Detection Abstract In this paper, we compare different ways of combining encoderbased neural language models for machine-generated text detection from a cost-centric perspective. We tested five ensembling approaches: soft voting with separately fine-tuned models, training ensembles on disjoint training data subsets (for reducing training costs), snapshot ensembles, models with multiple classification heads, and merging models by averaging their weights. We evaluated each method based on how accurate they are (using macro F1 scores) and how much computing costs they require during training and prediction. Our findings show that while typical ensemble methods can boost accuracy, they come with high resource usage. In comparison, model merging achieves the highest accuracy (macro F1 of 0.832) without increasing the inference time. Although model merging requires higher training costs, this is often less of a concern where inference time costs outweigh the costs of model construction. Keywords ensemble methods, ai text detection, encoder models 1 Introduction In recent years, improvements in generative models have made machine-generated text much more natural and realistic. Large volumes of AI-produced content are now being created and have even found their way into academic and scientific publications (Liang et al . , 2024b). For example, previous studies indicate that as much as 6.5% to 16.9% of peer review submissions to AI conferences may be heavily influenced or rewritten by large language models, highlighting a growing dependence on LLMs in scholarly environments (Liang et al . , 2024a). This trend is further reinforced by the widespread availability of powerful models, which can now be accessed by users with minimal technical expertise. Because of this, there is a growing need for automated tools that can distinguish between human-written and machine-generated text. This need has been underscored by several recent shared tasks and competitions focused on this problem (Sarvazyan et al . , 2023; Wang et al . , 2024; Chamezopoulos et al . , 2024; Wang et al . , 2025). Even with notable progress in the field, reliably separating machine-generated and human-written content remains difficult. Most state-of-the-art approaches for detecting AI-generated text rely on deep neural architectures, particularly transformer-based models fine-tuned on large annotated datasets (Sarvazyan et al . , 2024; Marchitan et al . , 2024). Recent research results have shown that the combination of multiple models can further improve the accuracy and robustness of machine-generated text detection (Gu and Meng, 2024; Kiss and Berend, 2025). However, these gains come at the expense of increased computational demands, Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society 2024, 7–11 October 2024, Ljubljana, Slovenia ©2024 Copyright held by the owner/author(s). both during training and inference, which can limit the practical adoption of such methods in large-scale, real-time applications. To address these concerns, this paper provides a cost-centric overview of different strategies adapted for machine-generated text detection. For our experiments, we relied on a recent publicly available dataset: Task 1, Subtask A of the Workshop on Detecting AI-Generated Content at COLING 2025 (Wang et al . , 2025), where the primary goal is to classify English text as human-written or AI-generated. We compare a range of approaches, including softvoting ensembles, ensembles on disjoint training data subsets, snapshot ensembles, multiple classification heads model, and weight space model merging. By evaluating each method not only on predictive accuracy, but also on their training and inference costs, we aim to identify solutions that offer a practical balance between effectiveness and efficiency. Our results show that, while traditional ensembling improves performance, it also results in significant computational overhead both during training and inference time. In contrast, model merging, where multiple models’ weights are combined, achieves higher accuracy without increasing inference costs compared to a single model. Although training becomes more expensive, this is a one-time expense, making model merging an efficient and practical solution. Our source code is available at http://retracted. for.anonimity. 2 Related Work Recent advancements in detecting machine-generated text have highlighted the efficacy of ensemble learning methods, particularly those that employ soft voting strategies (Gu and Meng, 2024; Kiss and Berend, 2025). In the SemEval-2024 Task 8 (Wang et al . , 2024), Gu and Meng (2024) introduced a class-balanced soft voting system that fine-tuned transformer-based models (Gu and Meng, 2024). Their approach effectively addressed data imbalance and achieved state-of-the-art performance in multi-class classification tasks involving various text generators. Due to the high computational cost of standard ensembles, which require training multiple independent models, several approaches have been proposed to reduce resource usage. One such method is the snapshot ensemble technique (Huang et al . , 2017), which captures multiple model “snapshots” at different points during a single training run. Although this approach decreases training costs by avoiding the need to train separate models, it usually does not reach standard ensemble performance. In the domain of multi-label and large-scale classification, Liang et al . (2025) proposed the Multi-Head Encoding (MHE) approach to tackle extreme label classification challenges(Liang et al . , 2025). Their innovation involves substituting a conventional single classifier with multiple specialized classification heads, with each head responsible for a distinct subset of the overall label space. Their experiments across several extreme classification benchmarks underscore the potential of multiple classification head architectures in improving both efficiency and generalization, which aligns with our motivation for using models with multiple classification heads in the context of AIgenerated text detection. 44 Information Society 2024, 7–11 October 2024, Ljubljana, Slovenia Trovato et al. 3 Methodology 3.1 Base models Dataset. The shared task dataset (Wang et al . , 2025) we used provides more than 610,000 English texts for training and 261,000 for validation. Each entry is labeled either machine-generated by one of 40 different LLMs or human-written, yielding a 41-class classification task. Single Model. For our ensemble experiments, we fine-tuned three separate DeBERTa-base models (He et al . , 2021) using a 41class setup. By adopting this multi-class approach, the models are better at capturing the diversity and differences among texts generated by different models, rather than treating all non-human written text as a single, heterogeneous category. Each DeBERTabase model was trained independently using the AdamW optimizer (learning rate 2e-5, weight decay 0.01), with a batch size of 16, a 10% warmup schedule, and early stopping based on validation performance. In our work, we intentionally chose not to rely on generative large language models (LLMs) as the core detection mechanism. This decision was motivated by practical considerations of cost efficiency and scalability. LLM-based solutions often have high computational costs and can be difficult to deploy at scale. As a result, our approach prioritizes methods that are both accurate and cost-effective. 3.2 Vanilla Ensemble Traditional ensemble methods, also known as output-space ensembling, where several independently trained models contribute to the final prediction, are widely used to improve both accuracy and robustness in machine learning (Ganaie et al . , 2022). It also works well for machine-generated text detection (Abburi et al . , 2023; Gu and Meng, 2024; Kiss and Berend, 2025). In our experiments, we used a standard soft voting approach, where each model was fine-tuned separately, and their predicted class probabilities were averaged to determine the final output. This typically led to better classification results than any single model. However, this improvement comes with significant resource costs as all models must be trained and evaluated individually. For our study, we fine-tuned three models and combined their predictions using soft voting, then compared the ensemble’s macro F1 score with the average score of the single models. 3.3 Disjoint Training Data (DTD) Ensembles As we wanted to keep the total fine-tuning cost of the ensemble close to that of a single model, while still benefiting from the advantages of ensembling, we employed an ensemble method, where we trained the model on disjoint training data subsets. Knowing the large size of our available training data (approx. 611,000 examples), we divided the dataset into three disjoint subsets of equal size. Each subset was independently used to fine-tune a separate model, resulting in three models trained on disjoint portions of the dataset. The smaller training subsets shorten the individual training durations, making the approach more efficient overall. After training, the three models were combined into an ensemble using soft voting. This setup uses the diversity of individually trained models while offering a cost-effective and time-efficient training alternative to standard ensemble strategies. 3.4 Snapshot Ensemble For the snapshot ensemble approach (Huang et al . , 2017), we generated several “snapshots” from a single model taken at different points throughout its fine-tuning. Unlike traditional ensemble methods, which require training multiple independent models from scratch, snapshot ensembling takes advantage of a cyclical learning rate schedule to encourage the model to converge to diverse regions of the parameter space within a single training run. This makes the method much more resource efficient, as it avoids the repeated cost of training from random initializations. In our experiments, we used a cyclical learning rate schedule, as described in the original method. This approach periodically increases and decreases the learning rate during training. At the end of each cycle, we saved the model as a snapshot (see Figure 1). We decided to save six snapshots to achieve sufficient diversity. Figure 1: Learning rate schedule for snapshot ensemble To construct the ensemble, we selected the three most diverse snapshots (based on validation performance) and combined their predictions via soft voting. This approach allowed us to obtain multiple strong models within a single fine-tuning process. However, since inference requires running each selected snapshot, the overall inference cost remains similar to that of standard ensembles. Although the improvement in predictive performance over a single model was modest, the snapshot ensemble presents a practical alternative when training resources are limited. Up to this point, the examined solutions mainly focused on reducing training costs, without lowering inference demands. From now on, we explore methods that aim to improve inference efficiency as well. 3.5 Model Merging Model merging, also known as weight space ensembling or model soups (Wortsman et al . , 2022), is a technique explored in machine learning as an alternative to traditional output space ensembling. Although assembling multiple independently trained models by averaging their predictions can improve performance, it comes at the cost of increased inference computation, as the output of several models must be calculated. Model merging addresses this by averaging the weights of several models to form a single, merged model. This merged model can match or exceed the performance of output space ensembles in many cases, but crucially, it incurs no additional inference costs compared to a single model. The effectiveness of weight space merging is supported by two fundamental concepts: mode connectivity (Frankle et al . , 2020) and sparsity. Mode connectivity suggests that independently trained neural networks, especially those derived from the same base model, can reside in regions of the loss landscape connected by paths of relatively constant performance. In this work, we applied a model merging approach using three distinct models. 45 Cost-Sensitive Overview of Model Ensembling for Machine-Generated Text Detection Information Society 2024, 7–11 October 2024, Ljubljana, Slovenia After merging, we experimented with two configurations. First, we attached the original three classification heads from the source models to the merged encoder, resulting in three separate models sharing the same backbone but differing in their classification heads. Second, we also fine-tuned new classification heads on top of the merged encoder, which model variant we refer to as Model Merging (FT). This setup kept the efficiency of a single-model inference pipeline while allowing the new head to adapt to the merged weights. 3.6 Models with Multiple Classification Heads In this architecture, multiple classification heads are attached to a shared backbone model (Chang et al . , 2023; Lee et al . , 2015). Each head is independently fine-tuned, allowing them to specialize and capture different aspects of the data distribution. During inference, the outputs of these heads are aggregated, typically by voting to produce the final prediction. This approach provides some of the benefits of ensembling, while maintaining a compact model structure and reducing deployment complexity. In our experiments, we implemented variants with 3 and 5 classification heads (referred to as MCH models). Each head independently produced a classification output, and the final label was determined by soft voting over them. The main advantage of the multiple classification headed (MCH) architecture is its low inference cost, equivalent to a single model since all heads share the same encoder. The only additional overhead comes from running the multiple classification heads, which is negligible compared to vanilla multi-model ensembles. 4 Results We compared several ensembling approaches, focusing on both their Macro-F1 scores and their training and inference costs. The vanilla ensemble and the fine-tuned model merging approach are averaged over three independent runs, the rest are based on a single evaluation. Costs Experiment Macro F1 Training Inference 3 Classifier Heads 0.805 Low Low Single model 0.806 Low Low Snapshot ensemble 0.810 Low High DTD Ensemble 0.814 Low High 5 Classifier Heads 0.817 Low Low Vanilla Ensemble 0.826 High High Model Merging (FT) 0.827 Highest Low Model Merging 0.832 High Low Table 1: Summary of results on test set with costs. Our experimental results are summarized in Table 1, which compares the predictive performance (macro F1 score) and computational costs (training and inference) of each ensemble method. The Vanilla Ensemble achieved an improved macro F1 score of 0.826, which was a substantial increase over the single model. However, this performance gain incurred significantly higher costs for both training and inference. The DTD ensemble approach resulted in a slightly lower macro F1 compared to the vanilla ensemble. However, it significantly reduced training costs, since each individual model was trained on only one-third of the training data, although inference costs remained high due to querying multiple models simultaneously. The Snapshot Ensemble offered moderate computational efficiency, yet resulted in a small improvement of macro F1 over the single model, while still retaining high inference costs due to combining predictions from multiple snapshots. MCH models demonstrated varying performance depending on the number of heads. The 3-head variant was matching single-model costs but did not outperform it. The 5-head variant improved F1 to 0.817 while maintaining low inference costs and moderate training expense. Model merging is the most promising approach. Directly merging pre-trained model weights showed the best overall performance, maintaining low inference cost equivalent to a single model. Further fine-tuning after merging increased training costs, but unlike ensembles, the merged model requires only a single model at inference. This greatly reduces inference time and resource usage. Considering both computational efficiency and classification performance, the model merging strategy without additional fine-tuning provided the best balance overall. 5 Conclusion We explored multiple ensemble strategies for detecting machinegenerated text, highlighting a cost-sensitive approach that balances computational efficiency with predictive performance. Our evaluation demonstrated that although ensemble methods consistently improve predictive accuracy, they typically require significantly higher training or inference resources, making them less practical for real-world scenarios where resources are limited. In particular, model merging was the most effective option among all the evaluated methods. Not only did it achieve the highest macro F1 score (0.832), but it did so without inflating inference costs, maintaining an efficiency comparable to running just a single model. This approach effectively captures the strengths of multiple models while avoiding the computational overhead commonly associated with traditional ensembling methods. Overall, our results show that the use of ensembles does not always have to be very costly. By carefully choosing methods such as model merging, it is possible to get most of the performance benefits of ensembling without using too many resources for predictions, especially during inference time. This can make these methods a good fit for building fast and scalable systems to detect AI-generated text. References Harika Abburi, Kalyani Roy, Michael Suesserman, Nirmala Pudota, Balaji Veeramani, Edward Bowen, and Sanmitra Bhattacharya. 2023. A Simple yet Efficient Ensemble Approach for AI-generated Text Detection. arXiv:2311.03084 [cs.CL] https://arxiv.org/abs/2311.03084 Savvas Chamezopoulos, Drahomira Herrmannova, Anita De Waard, Drahomira Herrmannova, Domenic Rosati, and Yury Kashnitsky. 2024. Overview of the DagPap24 Shared Task on Detecting Automatically Generated Scientific Paper. In Proceedings of the Fourth Workshop on Scholarly Document Processing (SDP 2024), Tirthankar Ghosal, Amanpreet Singh, Anita Waard, Philipp Mayr, Aakanksha Naik, Orion Weller, Yoonjoo Lee, Shannon Shen, and Yanxia Qin (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 7–11. https://aclanthology.org/2024.sdp-1.2 Haw-Shiuan Chang, Ruei-Yao Sun, Kathryn Ricci, and Andrew McCallum. 2023. Multi-CLS BERT: An Efficient Alternative to Traditional Ensembling. In Proceedings of the 61st Annual 46 Information Society 2024, 7–11 October 2024, Ljubljana, Slovenia Trovato et al. Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, Toronto, Canada, 821–854. https://doi.org/10.18653/v1/ 2023.acl-long.48 Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin. 2020. Linear mode connectivity and the lottery ticket hypothesis. In Proceedings of the 37th International Conference on Machine Learning (ICML’20). JMLR.org, Article 305, 11 pages. M.A. Ganaie, Minghui Hu, A.K. Malik, M. Tanveer, and P.N. Suganthan. 2022. Ensemble deep learning: A review. Engineering Applications of Artificial Intelligence 115 (Oct. 2022), 105151. https://doi.org/10.1016/j.engappai.2022.105151 Renhua Gu and Xiangfeng Meng. 2024. AISPACE at SemEval2024 task 8: A Class-balanced Soft-voting System for Detecting Multi-generator Machine-generated Text. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval2024), Atul Kr. Ojha, A. Seza Doğruöz, Harish Tayyar Madabushi, Giovanni Da San Martino, Sara Rosenthal, and Aiala Rosá (Eds.). Association for Computational Linguistics, Mexico City, Mexico, 1476–1481. https://doi.org/10.18653/v1/2024. semeval-1.212 Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. DeBERTa: Decoding-enhanced BERT with Disentangled Attention. In International Conference on Learning Representations. https://openreview.net/forum?id=XPZIaotutsD Gao Huang, Yixuan Li, Geoff Pleiss, Zhuang Liu, John E. Hopcroft, and Kilian Q. Weinberger. 2017. Snapshot Ensembles: Train 1, Get M for Free. In International Conference on Learning Representations. https://openreview.net/forum?id=BJYwwY9ll Mihaly Kiss and Gábor Berend. 2025. SzegedAI at GenAI Detection Task 1: Beyond Binary - Soft-Voting Multi-Class Classification for Binary Machine-Generated Text Detection Across Diverse Language Models. In Proceedings of the 1stWorkshop on GenAI Content Detection (GenAIDetect), Firoj Alam, Preslav Nakov, Nizar Habash, Iryna Gurevych, Shammur Chowdhury, Artem Shelmanov, Yuxia Wang, Ekaterina Artemova, Mucahid Kutlu, and George Mikros (Eds.). International Conference on Computational Linguistics, Abu Dhabi, UAE, 166–172. https://aclanthology.org/2025.genaidetect-1.15/ Stefan Lee, Senthil Purushwalkam, Michael Cogswell, David J. Crandall, and Dhruv Batra. 2015. Why M Heads are Better than One: Training a Diverse Ensemble of Deep Networks. CoRR abs/1511.06314 (2015). arXiv:1511.06314 http://arxiv.org/abs/ 1511.06314 Daojun Liang, Haixia Zhang, Dongfeng Yuan, and Minggao Zhang. 2025. Multi-Head Encoding for Extreme Label Classification. IEEE Transactions on Pattern Analysis and Machine Intelligence 47, 3 (March 2025), 2199–2211. https://doi.org/10. 1109/tpami.2024.3522298 Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, Daniel Mcfarland, and James Y. Zou. 2024a. Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews. In Proceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235), Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (Eds.). PMLR, 29575–29620. https: //proceedings.mlr.press/v235/liang24b.html Weixin Liang, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sheng Liu, Siyu He, Zhi Huang, Diyi Yang, Christopher Potts, Christopher D Manning, and James Y. Zou. 2024b. Mapping the Increasing Use of LLMs in Scientific Papers. In First Conference on Language Modeling. https://openreview.net/forum?id=YX7QnhxESU Teodor-george Marchitan, Claudiu Creanga, and Liviu P. Dinu. 2024. Team Unibuc - NLP at SemEval-2024 Task 8: Transformer and Hybrid Deep Learning Based Models for MachineGenerated Text Detection. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024), Atul Kr. Ojha, A. Seza Doğruöz, Harish Tayyar Madabushi, Giovanni Da San Martino, Sara Rosenthal, and Aiala Rosá (Eds.). Association for Computational Linguistics, Mexico City, Mexico, 403–411. https://doi.org/10.18653/v1/2024.semeval-1.63 Areg Mikael Sarvazyan, José Ángel González, and Marc Francosalvador. 2024. Genaios at SemEval-2024 Task 8: Detecting Machine-Generated Text by Mixing Language Model Probabilistic Features. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024), Atul Kr. Ojha, A. Seza Doğruöz, Harish Tayyar Madabushi, Giovanni Da San Martino, Sara Rosenthal, and Aiala Rosá (Eds.). Association for Computational Linguistics, Mexico City, Mexico, 101–107. https://doi.org/10.18653/v1/2024.semeval-1.17 Areg Mikael Sarvazyan, José Ángel González, Marc FrancoSalvador, Francisco Rangel, Berta Chulvi, and Paolo Rosso. 2023. Overview of AuTexTification at IberLEF 2023: Detection and Attribution of Machine-Generated Text in Multiple Domains. Proces. del Leng. Natural 71 (2023), 275– 288. http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/ article/view/6559 Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Osama Mohammed Afzal, Tarek Mahmoud, Giovanni Puccetti, and Thomas Arnold. 2024. SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval2024), Atul Kr. Ojha, A. Seza Doğruöz, Harish Tayyar Madabushi, Giovanni Da San Martino, Sara Rosenthal, and Aiala Rosá (Eds.). Association for Computational Linguistics, Mexico City, Mexico, 2057–2079. https://doi.org/10.18653/v1/2024. semeval-1.279 Yuxia Wang, Artem Shelmanov, Jonibek Mansurov, Akim Tsvigun, Vladislav Mikhailov, Rui Xing, Zhuohan Xie, Jiahui Geng, Giovanni Puccetti, Ekaterina Artemova, Jinyan Su, Minh Ngoc Ta, Mervat Abassy, Kareem Elozeiri, Saad El Dine Ahmed, Maiya Goloburda, Tarek Mahmoud, Raj Vardhan Tomar, Alexander Aziz, Nurkhan Laiyk, Osama Mohammed Afzal, Ryuto Koike, Masahiro Kaneko, Alham Fikri Aji, Nizar Habash, Iryna Gurevych, and Preslav Nakov. 2025. GenAI Content Detection Task 1: English and Multilingual Machinegenerated Text Detection: AI vs. Human. In Proceedings of the 1st Workshop on GenAI Content Detection (GenAIDetect). International Conference on Computational Linguistics, Abu Dhabi, UAE. Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S. Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt. 2022. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. arXiv:2203.05482 [cs.LG] https://arxiv.org/abs/2203.05482 47 Performance of an instance-specific scheduling optimizer MATCOS 2025, 9–10 October 2025, Koper, Slovenia i f not case . is_scheduled ( task ) : EXPLORE( case , unit , curr , task ) All implementations in Section 4 rely on this same exploration logic of the search space. Further algorithmic acceleration techniques of the S-graph framework are not utilized in any of the implementations. This adjusted Sgraph algorithm may not be the most efficient job-shop scheduler, however, it still retains the same important characteristics, as the general solvers for the wider problem class with precedential recipes, mixed intermadiate storage policies, changeover times, etc. 4 Generic and instance-specific implementations A shared feature of all of the S-graph implementations is to maintain a longest path matrix with a cached maximum instead of the schedule-graph itself. This removes the need for dynamic longest path calculations and cycle checks. Maintaining such a matrix when inserting an arc takes O(|V|2) time in the worst case, however, this technique was proven to be advantageous over the years. The recipe is always stored and accessible as constant global data. Altogether, there are 5 implementations that will be compared in the next section. G1 is the generic implementation, where a partial schedule is stored in such a structure: struct partial_schedule { vector<vector<uint>> longest_paths ; uint longest_path_max ; vector<vector<uint>> unscheduled_tasks ; }; In G2, the list of unscheduled tasks for each unit is not stored by a vector in this struct, but as a managed global boolean array, which is possible due to the DFSlike strategy. Comparing these two solvers will show, how much time is needed for managing several relatively short dynamic arrays. The implementations I1 and I2 are the problem scpecific equivalents, where fix-sized array s are employed instead of vector s, and the subroutine responsible for returning the root partial_schedule is constexpr with the return of a list-initialized structure instead of opening the input file, and reading the data dynamically. I3 is a special variant of I2, where the branching subroutine is instantiated for all possible unit , prev , curr triplets in compile time, thus, some of the conditionals and loops can be removed in exchange for longer compilation time and a larger binary. It is important to note, that all implementations traverse exactly the same DFS-like B&B tree with the same number of nodes, copies, etc. There are no algorithmic differences between them in this sense, and the C++ code for exploring a node is the same in G1 and I1, and similarly the same in the remaining three. 5 Empiric results Test cases All of the test cases were generated randomly and solved by all of the aforementioned implementations on the same machine. A single C++ source file, however, can result very different binaries based on the compiler used, and the applied configurations. All binaries used were compiled with g++ (GCC version 14.2). After testing various compiler configurations on a smaller set of examples, two binaries were generated for each implementation: one with optimization level 1, and another with optimization level 3. Initial tests showed, that -O2 and -Ofast did not show significant difference compared tot -O3 , and naturally, no optimization always provided worse results. Altogether 80 instances were tested with all of the 5 implementations and 2 compiling options. The size of the instances ranged from 3 to 8 products with 2-6 stages carried out by 5-8 units. Overheads Instance-specific solvers have an overhead for generating and compiling the source files. This is especially heavy for I3. The largest such overhead was less than 6 seconds for I3 with -O3 , for a problem, whose solution was measured in hours. For the same problem the overhad in case of I2 with -O1 was less then 1 second. On the other hand, constructing the root node is faster for the instance-specific binaries, but that only milliseconds and microseconds in the worst case for the generic and instance-specific solvers, respectively. As a conclusion, these overheads did not play a significant role, except for small problems, where the optimal solution can be found in a few seconds either way. Optimization levels A somewhat unexpected result between the -O1 and -O3 optimization option for the 5 implementations is shown on a box-plot in Figure 2. Figure 2: CPU time ratio of binaries with -O3 and -O1 for the 5 implementations The figure shows the distribution of CPU time with -O3 divided by that of with -O1 of all the test cases for all implementations. While several outliers are cropped from the diagram, it is clear to see, that in case of the general solvers, more optimization had minor benefits, but the instance-specific solvers performed significantly better with 54 MATCOS 2025, 9–10 October 2025, Koper, Slovenia Mate Hegyhati less optimization. Thus, in later comparisons results from the general solvers always refer to the -O3 build, and to the -O1 build for instance-specific solvers. Instantiated functions The overall performance of I2 and I3 were very close for all of the larger test cases. For the 5 largest ones, I2 outperformed I3 by 1,1,0,9,3 percent. I3 is removed from further discussion, as function instantiation did not bring measurable results for this problem class, only slight performance decrease on top of the compilation overhead. Unscheduled task set representation Figure 3 shows the effect of storing the set of unscheduled tasks differently in case of the generic and instance-specific solvers. Figure 3: Effect of different unscheduled task set management The Figure shows an XY plot, where the X and Y coordinates are CPU time ratios G2/G1, and I2/I1, respectively. For most of the instances, both coordinates are less than 1, the exceptions belong to smaller instances, i.e., maintaining a single boolean array was more beneficial, as expected. It can also be observed, that in general, this had a larger effect on instance-specific solvers, whose partial problem representation would not use vectors for other purpose. Largest instances Table 1 shows the overall results for several examples, including the largest cases and some medium-sized ones. The second column shows the number of tasks, which is the maximal depth of the B&B tree, thus the maximal number of function calls on the stack. Memory usage grows cubically with this in these DFS-like implementations. As is shown in the table, CPU becomes an issue much faster than RAM. Moreover, considering the 4 bytes for an unsigned int, the I2 solver takes 256 kB to store all weights in the matrices in the call stack when a complete schedule is examined. This number is definitely higher for G2, but it Table 1: Overall results for some examples case # tasks G2-O3 I2-O1 Reduction 66 40 11180 s 3246 s 71% 24 36 7392 s 2107 s 71% 45 30 62 s 14 s 77% 0 38 40 s 9 s 77% 1 39 19 s 4 s 79% 17 33 12 s 3 s 72% 69 37 8 s 2 s 78% is reasonable to assume, that syscalls for more heap space did not play a significant role for the larger examples. The last column of the table shows the reduction in CPU time by using insta-specific solvers, which are all in the 70-80% range. 6 Concluding remarks While computer programs are meant to be generic for many reasons, performance focused applications may benefit from generating binaries that have instance-specific information encoded. In this work, this question was put to the test on an S-graph-based custom scheduler for job-shop problems, and results showed that a 70% reduction in computational needs is achievable mostly by fixing container sizes in compile time. This initial investigation motivates more exhaustive experiments on more complex real-life case-studies. References [1] Klaus-Jürgen Bathe. 2002. Finite Element Procedures. PrenticeHall of India. isbn: 8120310756. http://www.worldcat.org/isbn /8120310756. [2] Yasemin Eryoldas and Alptekin Durmuşoğlu. 2021. A literature survey on instance specific algorithm configuration methods. Proceedings of the International Conference on Industrial Engineering and Operations Management. https://api.seman ticscholar.org/CorpusID:240158129. [3] M. Hegyhati and Zs. Tuza. 2023. Why implementation matters for graph algorithms? Presented at OPAL, Veszprém, Hungary. (June 2023). [4] Scott Meyers. 2005. Effective C++: 55 Specific Ways to Improve Your Programs and Designs. (3rd ed.). AddisonWesley Professional. [5] Zsolt Ádám Nemes. 2015. Heuristic Accelerations of the Sgraph Solver. Bachelor’s thesis. University of Pannonia. [6] E. Sanmartí, Luis Puigjaner, Tibor Holczinger, and Ferenc Friedler. 2002. Combinatorial framework for effective scheduling of multipurpose batch plants. AIChE Journal, 48, 2557– 2570, 11. doi:10.1002/aic.690481115. [7] G. R. Stibbs. 1958. Railway timetabling and early computers. Proceedings of the IEE. [8] Bjarne Stroustrup. 2013. The C++ Programming Language. (4th ed.). Addison–Wesley, Boston, MA. isbn: 978-0321958327. [9] Seth J. Teller and Carlo H. Séquin. 1991. Visibility preprocessing for interactive walkthroughs. SIGGRAPH Comput. Graph., 25, 4, (July 1991), 61–70. doi:10.1145/127719.122725. 55 Bi-Level Routing and Scheduling Alain Quilliot and H´ el` ene Toussaint LIMOS Lab. UCA, CNRS and EMSE Clermont-Ferrand, France Email: [email protected] Abstract—Decentralized renewable energy platforms promote self-consumption, a single agent becoming able to simultaneously act as a producer and a consumer of power. We deal here with the design of routing strategies for electric vehicles relying on a local photovoltaic (PV) platform. We propose an exact Mixed-Integer Linear Programming (MILP) setting solved through Branch and Cut, along with a heuristic approach involving an approximation of the behavior of the PV platform. I. INTRODUCTION Renewable energy sources are promoting the emergence of the self-consumption paradigm [6], according to which consumers become energy producers. This paradigm raises its own issues, operational ones (synchronization of production and consumption) and tactical (pricing). We consider here a Vehicle Routing under Energy Production Constraints VR EPC problem which consists in routing electric vehicles (EV) while synchronizing their activity with the time dependent production and purchase of power. Numerous routing models (see Green Vehicle Routing, Pollution Routing models [2], [7], [9],[3] ) have been designed, involving refueling transactions subject to time windows or shared access constraints, as well as CO2 emissions (see Kuo [4], Lajunen [5]). Nevertheless, few studies simultaneously dealt with both energy production and consumption, which requires integrating heterogeneous routing and scheduling processes (see [1]). Related models are complex ones, involving highly heterogeneous variables and relying on the existence of a single central decider, which may not fit most real life contexts. In any case, designing efficient algorithms remains a challenge. Because of the collaborative features of our problem, we adopt here the point of view of the EV and propose: •An exact MILP model that considers the vehicle as leader and includes Recharge Decomposition constraints requiring the application of a separation time-polynomial procedure. We solve it through branch-and-cut. •Heuristic algorithms, that approximate the behavior of the power producer. II. THE VR EPC PROBLEM We consider a photovoltaic (PV) platform, referred to as PVP, along with an electric vehicle (EV) in charge of visiting a set of customers. These two players interact through recharge events when the vehicles return to PVP to recharge. We suppose that the system is deterministic . Its main components are: •A Photo-Voltaic platform PVP The time horizon of PVP is divided into Nperiods i= 0, . . . , N −1, all with a same duration p. Thus period istarts at time p·iand ends at time p·(i+ 1). During this period, PVP is expected to produce Ri power units. It may also buy an additional amount yiof power, that cannot exceed a threshold CCh. The cost of yimay be written Φi(y), where Φiis a piecewise linear increasing convex function, convexity meaning that marginal power purchase prices are usually increasing. PVP is provided with a macro-battery, with storage capacity CP V P ≥CCh and initial load HP V P 0. It must meet vehicle’s demand, ending with a charge at least equal to HP V P 0under minimal purchase cost PCost. •An Electric Vehicle EV and a Set of Customers J EV is initially located at a Depot. It must visit, within the time horizon [0, N.p], a set J={1, . . . , M}of customers according to a TSP (Traveling Salesman) route Γ, before coming back to Depot. Moving from a customer jto a customer krequires Ej,k energy units and Tj,k time units. Vectors Tand Edefine distances on the set defined by Depot,PVP and J,Depot and PVP being considered as customers identified as Depot = 0 at the beginning of the process, Depot =M+ 1 at the end and PVP =−1.EV is provided with a battery, with storage capacity CEV and initial load HEV 0. It must end with a charge at least equal to HEV 0while minimizing the time EV Time when it is back to Depot. •Recharge Events EV must periodically move to PVP in order to recharge. Arecharge event requires a single period iand takes place between two customers j, k consecutive according to Γ:EV moves from jto PVP before time p·i, receives an amount m≤Inf(CCh, CEV )of power, and leaves PVP at time p·(i+ 1). This recharge event, denoted by ω= (i, j, k, m), induces a cost Ψiwhich depends on iand is independent on m(infrastructure resource cost). Purchasing power is forbidden during the period i. Recharge event ωmay force EV to wait at PVP until the beginning of period iin case it arrives before time p·i. If we denote by τjthe time when EV arrives at j, by 56 VEV jits energy load and by VP V P ithe energy load of the PV-plant at the beginning of i, then: –p·i≥τj+Tj,−1;τk=p·(i+ 1) + T−1,k; –VP V P i+1 =VP i−m≥0; –The load of EV at the end of iis equal to VEV j− Ej,−1+m≤CEV ; –The load of EV at the beginning of iis equal to VEV j−Ej,−1≥0. Given time versus money coefficient α, we set VR EPC as a bi-level mono-objective problem, with EV a leader. VR EPC: Vehicle Routing under Energy Production Constraints:{Compute the route Γfollowed by EV together with the recharge events linking PVP and EV in such a way that: •All customers are visited once within the time horizon. •Vehicle storage capacity constraints and energy requirements are satisfied. •Extended cost PCost+α.EV Time is minimized, under the PVP constraints: –ymeets the demand related the recharge events; –It meets PVP storage and charge capacities as well as the constraint about the final charge of PVP; –PCost =P i (Ψi·δi+ Φi(yi)).} III. A MILP MODEL SOLVED THROUGH BRANCH&CUT Since we adopt here the point of view of the vehicle, our master variables are 2 {0,1}-valued vectors Z= (Zj,k, k = j∈ {−1,0, . . . , M + 1})and X= (Xj,k, k =j∈ {0, . . . , M +1}), describing the routes followed by the vehicle with and without the recharge detours respectively. Those main EV variables are completed by time and charge EV variables, by PVP variables and by Recharge Event variables. •EV variables: –{0,1}-valued Zj,k, k =j∈ {−1,0, . . . , M + 1}: Zj,k = 1 iff EV moves from jto k: (-1 = PVP). –{0,1}-valued Xj,k, k =j∈ {0, . . . , M +1}:Xj,k = 1iff EV moves either from jto kor from jto PVP, and next from PVP to k. –Non negative LEV j, j = 0, . . . , M:LMP jis the power transferred to EV just after jif Zj,−1= 1; –Non negative VEV j, j = 0, . . . , M + 1:VEV jis the power stored by EV when it arrives at j; –Non negative τj, j = 0, . . . , M + 1:τjis the time when EV arrives at j; –Non negative τ∗ j, j = 0, . . . , M + 1:τ∗ jis the time when EV starts recharging after jif Zj,−1= 1. •PVP variables: –Non negative yi, i = 0, . . . , N −1:yimeans the power bought by PVP during period i; –{0,1}-valued δi, i = 0, . . . , N −1:δi= 1 means that some recharge event takes place at i; –Non negative VP V P i, i = 0, . . . , N −1, N:VP V P iis the charge of PVP at the beginning of i; –Non negative LP V P i, i = 0, . . . , N −1:LP V P iis the power transferred at iif δi= 1. •Recharge Event variables: –{0,1}-valued Ui,j, i = 0, . . . , N −1, j = 0, . . . , M: Ui,j = 1 means that some recharge event involving joccurs at period i; –Non negative mi,j, i = 0, . . . , N −1, j = 0, . . . , M: mi,j means related amount of power. A. Structural EV Constraints Zand Xdescribe the full route followed by EV. They must clearly meet the standard vehicle routing constraints: •ZM+1,0= 1;∀j:Zj,j = 0; (VR1) •∀j= 0, . . . , M + 1:Pk=−1,...,M+1 Zj,k =1= Pk=−1,...,M+1 Zk,j; (VR2) •Pj=0,...,M+1 Zj,−1=Pj=0,...,M+1 Z−1,j ≥1; (VR3) •∀j, k ∈ {0, . . . , M + 1}:Xj,k ≥Zj,k ; (VR4) •∀j∈ {0, . . . , M + 1}:Pk=0,1,...,M+1 Xj,k =1= Pk=0,1,...,M+1 Xk,j; (VR5) (VR1, ..., VR5) do not ensure that Z, X defines a route, consistent with the power requirements. We reinforce them by noticing that if EV spends Wenergy inside or at the border of some customer subset Awhich does contain PVP, then it must move at least ⌈W CEV ⌉times towards PVP in order to recharge. Let us set: •For any such a subset Aof {−1,0, . . . , M, M + 1}: –Cl(A) = {(j, k)s.t at least jor kis in A}; –δ(A) = {(j, k), s.t j /∈Aand k∈A}. •For any (j, k): –Πj,k =E0if (j, k) = (M+ 1,0) and Πj,k =CEV else. –Π∗ j,k =CEV −E0if (j, k)=(M+ 1,0) and Π∗ j,k =CEV else. Then we derive the Recharge Decomposition constraints: •For any A⊆ {0, . . . , M + 1},P(j,k)∈δ(A)Πj,k ·Zj,k ≥ P(j,k)∈Cl(A)Ej,k ·Zj,k (VR6) •For any A⊆ {0, . . . , M +1},P(j,k)∈δ(J−A)Π∗ j,k ·Zj,k ≥ P(j,k)∈Cl(A)Ej,k ·Zj,k (VR6-Bis) Lemma 1:The VRP constraints (VR1, ..., VR6-Bis) hold if and only if the arcs (j, k)such that Zj,k = 1 define a collection γof sub-tours γs, s = 0, . . . , S such that: •γ0starts from Depot = 0, ends into PVP =−1, and spends less than HEV 0power. (SUB1) •γSstarts from PVP =−1, ends into Depot = 0, and spends less than CEV –HEV 0power. (SUB2) •For any s= 1, . . . , S−1,γsstarts from PVP =−1, ends into −1and does not require more than CEV power. (SUB3) •Every customer j= 1, . . . , M is visited once. (SUB4) 57 Remark 1: Constraints (VR6, VR6-Bis) do not order the sub-routes γs, s = 0, . . . , S. Separating the Recharge Decomposition Constraints: Given 2 possibly non integral vectors (Z, X), separating the constraints (VR6, VR6-Bis) means checking that all those constraints are satisfied by (Z, X)and, in case they are not, computing a contradicting subset A⊆ {0, . . . , M+1}. Theorem 1:The Recharge Decomposition constraints can be separated in polynomial time, via a min cost flow algorithm. Principle of the Proof: We construct an auxiliary multi-graph G= (X, A)with X={Source =−1,0, . . . , M+1, M +2 = Sink}and Ais defined as the set of all simple-arc (j, k), j = k∈ {−1,0, . . . , M + 1}, augmented, for every j, with copyarcs (j, M + 1)k, k =j∈ {−1,0, . . . , M + 1}, connecting j to Sink = (M+ 2) and provided with label k. Every copyarc a= (j, M + 1)kis provided with a weight waequal to Ej,k ·Zj,k. Every simple-arc a= (j, k)is provided with a weight waequal to Πj,k ·Zj,k −Ej,k ·Zj,k. Then computing Athat contradicts VR6 (VR6-Bis) means computing some cut Bwhich separates Source =−1from Sink =M+ 2 in G and is such that Pa s.t. origin(a)∈B′,destination(a)/∈B′wadoes not exceed some threshold. This can be done in polynomial time through a Max-Flow algorithm. End-Proof. B. The VR EPC MILP MILP formulation We get it while distinguishing 3 main groups of constraints: •The PVP Constraints: They involve the variables related to the purchase of power and express the evolution along the periods of the load VP V P iof the PVP battery. •The EV Constraints: They contain VR1, ..., VR6-Bis, together with constraints related to the time and power values when the vehicle visit the customers or to PVP. •The Recharge Event Constraints: They link together the PVP periods and the time horizon [0, p ·N]of EV. Those constraints come as follows: VR EPC Constraints and Objective Function: •Objective: Minimize P i (Ψi.δi+ Φi(yi) + α.τM+1 •PVP Constraints –∀i= 1, . . . , N −1:yi≤CCh ·(1 −δi); (PC1) –∀i= 0, . . . , N:VP V P i≤CP V P ; (PC2) –VP V P 0=HP V P 0;VP V P N≥HP V P 0; (PC3) –∀i= 1, . . . , N:VP V P i=VP V P i−1+ yi−LP V P i; (PC4) •EV Constraints –(VR1, ..., VR6-Bis) involved in Lemma 1; –VEV 0=HEV 0;VEV M+1 ≥HEV 0; (VR7) –∀j= 0, . . . , M + 1:Ej,−1≤VEV j≤CEV ; (VR8) –∀j, k = 0, . . . , M + 1:Xj,k → (VEV k+Ej,k +(Zj,k −1)·(Ej,−1+E−1,k −Ej,k)) ≤ (VEV j+LEV j); (VR9) TABLE I BEHAVIOR OF VR EPC MILP Id (N, M, p)LBG LP UBG LP 3 (160, 10, 1) 51,6 171 3-2 (160, 10, 2) 51,7 174 5 (240, 20, 1) 43,8 148 6 (240, 20, 1) 59,5 176 5-2 (240, 20, 2) 43,9 148 6-2 (240, 20, 2) 59,1 201 8 (320, 30, 1) 58,0 151 9 (320, 30, 1) 51,7 117 8-2 (320, 30, 2) 58,0 153 9-2 (320, 30, 2) 51,2 127 –∀j= 0, . . . , M:Zj,−1→(VEV j≥Ej,−1); (VR10) –∀j= 0, . . . , M:Zj,−1→(VEV j+LEV j≤ Ej,−1+CEV ); (VR10-Bis) –τ0= 0;τM+1 ≤p·N; (VR11) –∀j, k = 0, . . . , M + 1:Zj,k → (τj+Tj,k ≤τk); (VR12) –∀j, k = 0, . . . , M + 1:(Xj,k −Zj,k = 1) → (τ∗ j+p+T−1,k ≤τk); (VR13) –∀j= 0, . . . , M + 1:(Zj,−1)→ (τj+Tj,−1≤ −τ∗ j); (VR13-Bis) •Synchronization Constraints –∀j= 0, . . . , M:Pi=0,...,N−1Ui,j =Zj,−1; (SY1) –∀j= 0, . . . , M: Pi=0,...,N−1mi,j =LEV j; (SY1-Bis) –∀i= 0, . . . , N −1:Pj=0,...,M Ui,j =δi; (SY2) –∀i= 0, . . . , N −1: Pj=0,...,M mi,j =LP V P i; (SY2-Bis) –∀j= 0, . . . , M:Pi=0,...,N−1p·i·Ui,j =τ∗ j; (SY3) –∀i= 0, . . . , N −1, j = 0, . . . , M: mi,j ≤Inf(CP V P , CEV )·Ui,j ; (SY4) Theorem 2:Above VR EPC MILP MILP model solves VR EPC in an exact way. We handle VR EPC MILP through Branch and Cut. C. Numerical Experiments Purpose: Evaluating the VR EPC MILP MILP model. Technical Context: The experiments are performed on an AMD EPYC 7H12 64-Core processor, under Gnu/linux Ubuntu 20.04.2CPLEX 12.10 is used in single-thread mode. Instances: The main parameters of every instance are the customer number M(from 5 to 30), the period number N (from 40 to 320), and the period length p∈ {1,2,4}. We derive instances with p= 2 (Id −2) from instances with p= 1 (Id) by merging the periods, p·Nremaining the same. Results: For every instance, Table II displays the lower bound LBG LP and the upper bound UBG UB computed through Branch and Cut by CPLEX library in 2 CPU hours. 58 IV. HEURISTIC HANDLING OF VR EPC We drive the leader route Γ(the vector Xof the MILP setting) while relying on the correlation which exists between the optimal value of VR EPC(Γ) and some key features of Γ. Let us W(Γ) be the optimal value VR EPC(Γ) and LT(Γ), LE(Γ) be the lengths of Γin the sense of Tand Erespectively. Standard TSP 2 Opt and Reloc operators act on any route Γ through 2 parameters j1, j2in J+{Depot}: •2 Opt(Γ, j1, j2)replaces the moves from j1(j2) to its successor ¯ j1(¯ j2) by moves from j1(¯ j1) to j2(¯ j2); •Reloc(Γ, j1, j2)relocates j1between j2and ¯ j2. A. Approximating W(Γ) through Path Search Once Γis fixed, a full solution is determined by the sequence of recharge events linking EV and PVP, augmented with the amounts of power bought between 2 consecutive recharge transactions. Thus, we design a PDYN EPC algorithm that searches for a path in a transition graph: •A state in the sense is a 4-uple S= (j, i, V V EV j, V P V P i), where jis a customer, ia period, VEV jis the power load of EV when it leaves jand VP V P ithe power load of PVP at the end of i, with the implicit meaning that a recharge transaction involving iand jhas just been performed. •A decision is a 4-uple (j1, i1, m1,¯y), where (j1, i1, m1) means the next recharge event and ¯ymeans the power bought by PVP during the periods i+ 1, . . . , i1−1. Related transition cost is p·(i1−i)augmented with the purchase cost of ¯y, which corresponds to the optimal value of some convex optimization program and is computed and stored as part of a pre-process. •Initial state S0is a 4-uple (Depot, −1, HEV 0, HP V P 0). A final state is any 4-uple (Depot, N −1, V EV j≥ HEV 0, V P V P j≥HP V P 0). We speed the algorithm by fixing an upper bound NDec on the number of feasible decisions and introduce filtering devices. B. Two Simple Heuristic Algorithms Experiments show that the evolutions of W(Γ),LT(Γ), LE(Γ) and the number NRT during a decent loop performed on W(Γ) via the 2 Opt and Reloc operators are strongly correlated. So we partially short-cut W(Γ) and rely on LT(Γ) and LE(Γ) in order to drive Γtowards good solutions. The GRASP Algorithm GRASP VR EPC:It considers a replication parameter Qand for any q= 1, . . . , Q: 1) It randomly generates parameters ωT, ωE≥0together with an initial Γ(q); 2) It applies 2 Opt and Reloc until Γ(q)becomes a local optimum with respect to ωT·LT(Γ)+ ωE·LE(Γ); 3) It computes W Aux(Γ(q)) viaPDYN EPC and updates the best current route Γ(qBest). TABLE II RESULTS FOR GRASP VR EPC Id GRASP 150 Desc +LS T Desc T GRASP 3 173 165 122,7 145.6 3-2 176 171 35,6 47.6 5 156 168 356,2 327.2 6 175 176 565,7 845.1 5-2 158 173 109,8 134.1 6-2 181 181 203,6 227.2 8 121 121 1998,6 542.5 9 119 122 2351,0 578.9 8-1 132 134 773,0 246.7 9-1 123 130 629,1 189.3 The Pseudo-Descent Algorithm Descent VR EPC: We allow 2 Opt and Reloc to slightly deteriorate LT(Γ) and LE(Γ). Given 2 parameters δT>0, δE>0and 2 routes Γ1,Γ2, we say that Γ2deteriorates Γ1by no more than (δT, δE)if LT(Γ2)−LT(Γ1)≤δ1or LE(Γ2)−LE(Γ1)≤δ2. Then the Descent VR EPC(δT, δE) algorithm works as follows: 1) Initialize Γby applying GRASP SVR EP(1); 2) At any iteration do a) Generate all the 2 Opt and Reloc parameters (j1, j2)such that resulting route deteriorates Γby no more than (δT, δE); b) For any (j1, j2)selected this way and any resulting route Γ1compute W Aux(Γ1)through PDYN EPC; If improving Γis possible, then do it according to a Best Descent strategy else stop. C. Numerical Experiments Table III involves the same instances as Table II. It displays: •The value GRASP 150 obtained by GRASP VR EPC(150) and related CPU time T GRASP. •The value Desc+LS by pipelining: GRASP VR EPC(1) (1 replication) →Descent VR EPC(4, 4), together with CPU time T Desc. Comments:GRASP VR EPC with 150 replications most often reaches quasi-optimality. The pipe-line GRASP VR EPC(1) →Descent VR EPC(4, 4) improves resulting value in 2 among the 10 instances. REFERENCES [1] Drexl M.: Synchronization in vehicle routing–a survey; Transportation Science 46, pp. 297-316, (2012) [2] Erdelic T., Caric T., Lalla-Ruiz E.: A Survey on the Electric Vehicle Routing Problem.Journal of Advanced Transportation (2019). [3] Koc¸ C., Jabali O., Mendoza J., Laporte G.. The electric vehicle routing problem with shared charging stations. ITOR, 26, (2018). [4] Kuo Y.. Simulated annealing to minimize fuel consumption for the timedependent VRP. Comp. Indust. Eng., 59 :157–165, (2010). [5] Lajunen A.: Energy consumption and cost analysis of electric city buses. Transportation Research Part C, 38 :1–15, (2014). [6] Luthander R., Widen J., Nilsson D., Palm J.: Photovoltaic selfconsumption in buildings. Applied Energy 142, p 80-94, (2015). [7] Macrina G., Di Puglia L., Guerriero F.: The Green-Vehicle Routing Problem: A Survey. Model. and Optim. in Logistics p 1-26 (2020). [8] Verma A.: Electric vehicle routing with recharging stations. Euro Journal of Transport. Logistics 7, p 415-451, (2018). 59 A discrete event simulation model for analyzing the wood waste reverse supply chain Nikola Kovačević [email protected] UP FAMNIT, University of Primorska Koper, Slovenia Črtomir Tavzes crtomir[email protected] InnoRenew CoE, UP IAM, University of Primorska Koper, Slovenia Balázs Dávid [email protected] InnoRenew CoE, UP IAM and UP FAMNIT, University of Primorska Koper, Slovenia Abstract This paper presents a discrete event simulation model for analyzing reverse supply chain operations of wood waste materials. The model enables comprehensive evaluation of storage management, processing capabilities and resource allocation in reverse logistics networks, tracking the flow and transformation of resources from generation to its final destination either as recovered and reused products of landfilled waste. The proposed model is capable of managing different inventory policies and stock strategies. The efficiency of the proposed model is shown on a network generated based on statistical data of waste wood in Slovenia. Keywords reverse supply chain, discrete event simulation, inventory management, wood waste 1 Introduction Wood waste represents a growing environmental and economic challenge worldwide, driven by expanding construction, manufacturing, and demolition activities. If not managed properly, it constitutes both a lost resource and a significant burden, contributing to greenhouse gas emissions, pollution, and rising landfill costs. Recent years have witnessed increased wood waste generation due to expanding construction sectors, including substantial increases in renovations and refurbishments for structural and energy improvements of building stock, alongside growing demand for wood-based packaging [14]. The environmental and economic consequences of wood waste mismanagement are profound. Conventional disposal methods, such as landfilling and incineration, release harmful greenhouse gases like methane and carbon dioxide, while slow decomposition under anaerobic conditions creates long-term environmental hazards [11, 10]. Beyond ecological damage, improper disposal poses public health risks due to associated pollution. Conversely, wood waste remains a vastly underutilized resource, with substantial volumes that could be repurposed into valuable raw materials or energy sources if managed effectively [2, 9]. To fully realize these opportunities, effective reverse supply chains are essential. Unlike traditional forward supply chains, which move products from raw materials to consumers, reverse supply chains focus on recovering value from used products through collection, reprocessing, and redistribution [5, 12]. Wood waste presents unique logistical and environmental challenges, including uncertain return volumes, variable product conditions, and complex processing requirements, but also significant potential for energy recovery and impact reduction. MATCOS 2025, 9–10 October 2025, Koper, Slovenia 2024. Reverse supply chains for wood waste operate within the framework of cascading use principles, where reclaimed wood undergoes successive down-cycling through high-value applications— from structural reuse to particleboard feedstock—before final energy recovery, thereby maximizing carbon storage while reducing virgin material demand [3]. Wood waste reverse supply chains involve diverse stakeholders spanning sawmills, panel producers, demolition firms, recyclers, transport providers, and municipal waste collectors. These actors coordinate complex logistical processes including collection, transportation, quality control, sorting, and processing operations. However, this complexity introduces substantial challenges: seasonal and spatial dispersion of biomass sources, wide variability in material form and contamination levels, high processing costs, and insufficient information systems for post-consumer wood recovery [7]. Discrete Event Simulation (DES) has emerged as a particularly effective tool for modeling these complex reverse logistics systems. Unlike analytical methods that often require simplifying assumptions, DES can capture the stochastic nature of return flows, processing variability, and resource constraints that characterize real-world reverse logistics operations [4]. Discrete-event simulation has proven effective for studying supply chain dynamics and transport logistics, particularly in complex systems where uncertainty and stochastic elements play critical roles [8]. By representing a supply chain as a sequence of discrete events, DES enables detailed representation of operational policies and aids in evaluating alternative strategies and system responses to uncertainty. Recent advances demonstrate DES capability to analyze complex multi-level resilience relationships that traditional analytical methods struggle to capture [6]. Current research predominantly employs deterministic optimization models that fail to capture the inherent stochasticity of wood waste systems. While existing research has advanced understanding of individual system components, there remains a gap in DES applications to wood waste reverse supply chains. The bidirectional flows, uncertain material quality, and variable availability patterns characteristic of these systems require simulation approaches capable of representing stochastic interactions and multi-stakeholder coordination mechanisms. This paper presents a DES model designed for analyzing reverse supply chain operations in wood waste management contexts. The model enables evaluation of processing capabilities, resource allocation strategies, and operational policies through a dual-axis inventory management framework that combines push/pull strategies with different stock management approaches. Using Slovenia as a case study, we demonstrate how the simulation framework supports evidence-based decision-making in reverse logistics network design and management. 60 MATCOS 2025, 9–10 October 2025, Koper, Slovenia Kovačević et al. Product A Demand Treatment (convert waste into product A) order inventory Collector (waste types P,Q,R) inventory transport excess manage storage Generator (waste type P) inventory manage storage needs waste P,Q,R Generator (waste type Q) inventory manage storage Generator (waste type R) inventory manage storage manage storage waste type P product A Treatment Inventory generate generate generate transport excess Handle excess (buy inventory or landfill) transport excess transportorder order transport Figure 1: Overview of the model. 2 Model Development A discrete-event simulation model was developed for the representation of the wood waste reverse supply chain. The proposed framework models waste flows from generation through collection to final processing across connected regions, with each region containing generators, collectors, and treatment facilities. The system tracks three primary entity categories that interact through event-driven communication protocols, responding to state changes and resource availability. Figure 1 provides an overview of this system, presenting all model entities, their connections and the possible decisions available to them. Table 1: Policy Configuration Framework Configuration Description PUSH ON DEMAND Continuous forecast-driven production PUSH REORDER 50 Production triggered at 50% capacity PUSH REORDER 90 Production triggered at 90% capacity PULL ON DEMAND Demand-triggered operation only PULL REORDER 50 Demand-driven with 50% threshold PULL REORDER 90 Demand-driven with 90% threshold 2.1 Entity Modeling Waste Generators represent real-world actors that produce wood waste, serving as material sources introducing specified quantities and types over time. Generator behavior incorporates variability through random seasonal fluctuations reflecting market and operational uncertainties. Generators produce waste proactively regardless of system policy. Collector Companies serve as commercial intermediaries operating as midstream buffers managing bidirectional material flows. Collected material is stored in finite-capacity collection centers that decouple generation events from treatment scheduling. The transport system operates at two levels: local collection from generators to collection centers with capacity and distance constraints, and long-haul transport from collection centers to treatment facilities using priority-based scheduling. Treatment Operators represent industrial facilities that serve as final waste customers and value-adding stages. These entities consume waste as raw material, performing transformation according to predefined recipes for products. 2.2 Decision strategies All entities operate under a dual-axis operational paradigm that collectively dictates production and procurement logic throughout the reverse supply chain. The Inventory Policy axis defines strategic philosophy: PUSH policies operate using forecast-driven approaches, managing operations based on internal state projections, while PULL policies implement lean, demand-driven approaches aligned with Just-in-Time principles. The Stock Strategy axis defines tactical execution rules: ON DEMAND represents continuous operation under PUSH or operation only upon signal receipt under PULL, while buffer-based strategies (REORDER 50 and REORDER 90) trigger actions when inventory levels cross 50% or 90% capacity thresholds respectively. This framework generates six distinct operational models (2×3) that enable systematic empirical comparison. Table 1 summarizes the six operational configurations evaluated in this study. Each configuration represents a unique combination of inventory policy (PUSH vs PULL) and stock strategy (ON DEMAND, REORDER 50, REORDER 90), enabling analysis of different supply chain management approaches. 61 A discrete event simulation model for analyzing the wood waste reverse supply chain MATCOS 2025, 9–10 October 2025, Koper, Slovenia The framework incorporates realistic capacity constraints through an overflow management system. When entities exceed storage capacity, they choose between capacity expansion or landfill disposal using cost-minimization logic with dynamic pricing that escalates with repeated use, reflecting real-world constraints and discouraging reactive management. 2.3 Stochastic Elements The model incorporates several stochastic components to capture real-world uncertainty. Entity failures use uniform recovery durations with failure probabilities varying by scenario severity. Waste generation includes daily variability factors from clipped normal distributions, while treatment conversion efficiency is normally distributed around base values. The model ensures reproducibility through hierarchical random number generation with deterministic seeding for consistent behavior across runs. 3 Results The simulation study analyzes six operational models across a 365-day simulation period representing one operational year. Each configuration underwent 100 Monte Carlo replications with unique random seeds to capture stochastic variability. System performance is evaluated through five complementary metrics that capture operational effectiveness across the complete material flow trajectory. Service level quantifies system reliability by measuring the percentage of generated waste successfully converted into final products, representing the fundamental waste diversion capability. Operational efficiency provides hierarchical assessment through three interconnected ratios: collection efficiency (waste collected relative to generated), processing efficiency (waste converted relative to collected), and overall system efficiency (final products relative to initial generation), thereby identifying bottlenecks and resource utilization patterns throughout the supply chain. Storage utilization examines capacity management by tracking current storage usage against maximum capacity across collection centers and processing facilities, revealing both potential constraints and the effectiveness of demand-supply synchronization under different inventory policies. Environmental impact quantifies the emissions in kgCO 2 e equivalent generated through transportation, processing operations, and landfill disposal activities. Finally, landfill overflow measures the volume of waste diverted to disposal due to system capacity limitations or processing constraints, representing both environmental burden and foregone resource recovery opportunities. Collectively, these metrics provide comprehensive assessment of reverse supply chain performance across operational, environmental, and resource recovery dimensions. The simulation framework was implemented with SimPy as the core simulation engine [13]. 3.1 Setup The analytical framework employed Slovenia as the case study, with each of the 12 statistical regions of the country represented by a single instance of each system entity (generator, collector and treatment). The system incorporated 14 waste codes and 3 product types. These types were selected to be engineered wood products: Particle Board (accepting five waste types), OSB (four waste types), and MDF (four waste types). Production prioritization follows ABC analysis based on total biogenic carbon impact, with OSB demonstrating the highest carbon storage potential (−1213.60 kgCO2e/m3). While not all waste codes contribute to product creation, unused codes were retained to maintain representation of Slovenia’s wood-related waste streams, and resources of these waste codes still contributed to decisions about inventory management and landfilling. Initial generation rates derived from empirical data by the Slovenian Environment Agency (ARSO) datasets [1] and underwent dynamic adjustment throughout the 365-unit simulation period. 3.2 Inventory Management Performance Storage utilization analysis reveals distinct patterns across entity types and strategies. PUSH strategies maintain consistently high storage utilization with frequent saturation events, particularly in the simulation’s latter half, reflecting proactive inventory prepositioning regardless of downstream demand. PULL strategies demonstrate more dynamic, demand-responsive behavior with lower average utilization and greater temporal variability. At the generation level, PUSH policies show progressive storage accumulation leading to widespread saturation across most regions, with particularly high utilization (80-100%) sustained throughout the simulation period. PULL strategies maintain more balanced levels with moderate utilization (20-60%) and better demand synchronization, avoiding the excessive accumulation patterns observed in PUSH configurations. Collection storage exhibits similar patterns, with PUSH policies creating sustained high utilization across all regions and stock strategies, while PULL policies show significantly lower average utilization with more efficient turnover. The temporal analysis reveals PUSH strategies experience early saturation that persists throughout the simulation, whereas PULL strategies maintain responsive storage levels that adapt to demand fluctuations. 3.3 Cost-Environmental Impact Analysis The cost-environmental impact analysis, based on single simulation runs for each configuration, reveals significant trade-offs across supply chain configurations. PULL REORDER 90 emerges as the best-performing strategy for environmental performance, achieving the lowest environmental impact (approximately 800k kgCO 2 e) at competitive costs (around 30 million €). PULL ON DEMAND demonstrates comparable performance with slightly higher emissions (approximately 970k kgCO 2 e) at lower cost levels. PULL REORDER 50 occupies an intermediate position within the PULL solution space, demonstrating suboptimal efficiency relative to alternative configurations due to higher costs despite moderate environmental performance. PUSH strategies generally show higher environmental impacts, with PUSH REORDER 90 representing the least favorable option, generating approximately 2mkgCO 2 e emissions at costs around 58 million €. Figure 2 illustrates the total costs and environmental impacts of each strategy, demonstrating PULL policy as the dominant policy across both cost and environmental dimensions. 3.4 Monte Carlo Analysis Results The statistical analysis based on 100 Monte Carlo replications reveals different performance patterns when accounting for stochastic variability. All strategies achieved high service levels 62 MATCOS 2025, 9–10 October 2025, Koper, Slovenia Kovačević et al. on demand reorder 90 reorder 50 on demand rder 90 reorder 50 30M 35M 40M 45M 50M 55M 60M 0.8M 1M 1.2M 1.4M 1.6M 1.8M 2M push - on demand push - reorder 90 push - reorder 50 pull - on demand pull - reorder 90 pull - reorder 50 Total Cost (€) Total Environmental Impact (kg CO₂e) Figure 2: Cost versus environmental impact analysis across supply chain configurations. above 97%, with PULL REORDER 50 reaching the highest average (99.9%), followed closely by PULL REORDER 90 (99.7%). PUSH strategies showed slightly lower but still strong service levels, ranging from 97.8% (PUSH ON DEMAND) to 98.6% (PUSH REORDER 90). Environmental impact analysis reveals PULL REORDER 50 as the most environmentally friendly option (278k kgCO 2 e), followed by PULL REORDER 90 (323k kgCO 2 e). PUSH REORDER 90 generated the highest emissions (629k kgCO 2 e), more than double that of the best-performing configuration. Landfill overflow analysis demonstrates superior waste management performance of PULL strategies. PULL REORDER 50 achieved zero landfill overflow, while PULL ON DEMAND generated minimal overflow (16.1m 3 ). In contrast, PUSH REORDER 90 produced substantial landfill overflow (2,014m 3 ), indicating significant waste management inefficiencies. Efficiency comparison reveals nuanced performance patterns across configurations. Collection efficiency ranges from 39.3% (PULL REORDER 50) to 58.1% (PUSH REORDER 90), while processing efficiency shows PULL strategies generally outperforming PUSH approaches, with PULL REORDER 50 achieving the highest processing efficiency (38.2%). Overall system efficiency demonstrates PUSH REORDER 90 as the top performer (19.0%), followed by PULL REORDER 90 (16.6%) and PULL REORDER 50 (14.9%). 4 Conclusion & Future Work In this paper, we developed a DES framework for analyzing waste and material flow and transformation in reverse supply chains. The proposed model was tested on a simplified network of wood waste flows built on statistical data by ARSO. The model is capable of evaluating different operational strategies through a dual-axis framework that combines inventory policies (PUSH vs PULL) with stock management approaches (ON DEMAND, REORDER 50, REORDER 90), enabling systematic comparison of supply chain configurations under realistic stochastic conditions. The simulation results demonstrate that PULL-based inventory strategies offer superior environmental performance while maintaining competitive service levels. The analysis reveals a fundamental trade-off between aggressive inventory accumulation and environmental sustainability, with PULL strategies achieving better demand synchronization and reduced waste overflow compared to forecast-driven PUSH approaches. These findings suggest that waste management practitioners should prioritize demand-driven operational models over forecast-based strategies. PULL REORDER configurations emerge as particularly effective, balancing high service levels with minimal environmental impact through responsive inventory management that avoids excessive accumulation and associated disposal costs. The framework provides valuable insights for reverse logistics network design and demonstrates the importance of aligning inventory policies with sustainability objectives in waste management systems. Future work should extend the model to incorporate multi-product interactions, dynamic pricing mechanisms, and regional policy variations to enhance applicability across diverse waste management contexts. Acknowledgements The research was supported by the BioLOG project: the authors are grateful for the support of National Center of Science (NCN) through grant DEC-2020/39/I/HS4/03533, the Slovenian Research and Innovation Agency (ARIS) through grant N1-0223 and the Austrian Science Fund (FWF) through grant I 5443-N. Balázs Dávid is also grateful for the support of the Slovenian Research and Innovation Agency (ARIS) through grant J1-50000. Balázs Dávid and Črtomir Tavzes gratefully acknowledge the Slovenian Research and Innovation Agency (ARIS) and the Ministry of the Economy, Tourism and Sport (MGTŠ) for the grant V4-2512. References [1] Agencija Republike Slovenije za okolje (ARSO). 2025. Poročila in publikacije o odpadkih. (2025). Retrieved Aug. 20, 2025 from https://www.arso.gov.si/v arstvo%20okolja/odpadki/poro%C4%8Dila%20in%20publikacije/. [2] Peter Akhator, Albert Obanor, and Anthony Ugege. 2017. Nigerian Wood Waste: A Potential Resource for Economic Development. Journal of Applied Sciences and Environmental Management, 21, 2, 246. doi: 10.4314/jasem.v21i 2.4. [3] Michael Burnard, Črtomir Tavzes, Aleksandar Tošić, Andrej Brodnik, and Andreja Kutnar. 2015. The role of reverse logistics in recycling of wood products. In Environmental Implications of Recycling and Recycled Products. Subramanian Senthilkannan Muthu, editor. Springer Singapore, 1–30. doi: 10.1007/978-981-287-643-0_1. [4] George S. Fishman. 2001. Discrete-Event Simulation. Springer New York, New York, NY. isbn: 978-1-4419-2892-4 978-1-4757-3552-9. doi: 10.1007/978 -1-4757-3552-9. [5] S.M. Gupta, ed. 2013. Reverse Supply Chains: Issues and Analysis. (1st ed.). CRC Press. doi: 10.1201/b13749. [6] Dmitry Ivanov. 2025. Comparative analysis of product and network supply chain resilience. International Transactions in Operational Research, itor.13612. doi: 10.1111/itor.13612. [7] Arkadiusz Kawa. 2023. REVERSE SUPPLY CHAIN OF RESIDUAL WOOD BIOMASS. Logforum. [8] C. Kögler and P. Rauch. 2018. Discrete event simulation of multimodal and unimodal transportation in the wood supply chain: a literature review. Silva Fennica, 52. doi: 10.14214/sf.9984. [9] Dorin Maier. 2023. A Review of the Environmental Benefits of Using Wood Waste and Magnesium Oxychloride Cement as a Composite Building Material. Materials, 16, 5, 1944. doi: 10.3390/ma16051944. [10] J. Owoyemi, H. Zakariya, and I. Elegbede. 2016. Sustainable wood waste management in nigeria. Environmental & Socio-Economic Studies, 4, 3, 1–9. doi: 10.1515/environ-2016-0012. [11] Kemal Parlak, Nural Yilgor, and Atakan Öngen. 2024. Hydrogen-rich syngas production from wood waste and wood waste pellet via gasification in updraft circulating fixed bed reactor. Research Square, (May 2024). Preprint. doi: 10.21203/rs.3.rs-4343729/v1. [12] Carol Prahinski and Canan Kocabasoglu. 2006. Empirical research opportunities in reverse supply chains. Omega, 34, 6, 519–532. doi: 10.1016/j.omega .2005.01.003. [13] Team SimPy. 2025. Simpy. Accessed: 2025-05-15. (2025). https://simpy.readt hedocs.io/en/latest/. [14] Statistical Office of the Republic of Slovenia. 2024. Več komunalnih odpadkov predvsem zaradi povečanja količine kosovnih odpadkov. Stat.si. Accessed: 2025-05-15. (2024). https://www.stat.si/StatWeb/News/Index/12770. 63 Harvest plan generation in precision agriculture Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia 2022-06-01 2022-06-08 2022-06-15 2022-06-22 2022-07-01 2022-07-08 2022-07-15 Date 0.0 0.1 0.2 0.3 0.4 0.5 Value Linear regression Data Linear fit y=0.02 x 2022-07-16 2022-02 2022-03 2022-04 2022-05 2022-06 2022-07 Date 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Value Polynomial regression - Degree: 2 Data Poly fit y=0.02 x 2022-01-16 x 2022-07-09 Figure 2: An example of fitting a linear model (top) and a quadratic polynomial model (bottom) for a field with wheat. Black dots mark the prediction which can be found on the intersection between the maturity threshold 𝑦𝑓= 0.02 and the time axis. GeoPackage 𝐹 , populate the fields with new metainformation and return the result to the user. The algorithms used in the process are described in 1. 3 Results A collection of eighteen fields located in northwestern Slovenia were used to evaluate the process. There were total of ten samples with barley and eight samples with wheat. Sometimes the actual harvests happened before or after the optimal time for harvesting because of bad weather or disruptions in resources (farming equipment breaks, organisational problems due to labor shortage). The algorithm was intended to be used near the start of the actual harvest season, which is usually determined by numerous exogenous factors. Our method assumes perfect conditions so it does not take those limitations into account. The vegetation index, the historic window length and maturity threshold 𝑦𝑓 were all chosen in conjunction with agricultural experts. We should point out that the historic window can be increased, but then the preprocessing step takes more time. When doing extrapolation, both linear and quadratic polynomial regression diverge far from the expected results or in the case of quadratic regression there can be no real solutions. To avoid this, we clamp the harvest dates to the period between the 20th of June and 31st of July, which are reasonable bounds for barley and wheat. Table 1: Mean difference and std. deviation between predicted and actual harvest times in days for linear regression. Crop Mean Std. deviation Accuracy Barley 15.10 12.51 5/10 Wheat 23.10 20.79 2/8 Table 2: Mean difference and std. deviation between predicted and actual harvest times in days for quadratic polynomial regression. Crop Mean Std. deviation Accuracy Barley 5.70 15.71 6/10 Wheat 5.75 11.12 6/8 0 2 4 6 8 Sample index 0 5 10 15 20 25 30 Deviation Deviation from baseline (0) - Barley 0 2 4 6 8 Sample index 10 0 10 20 30 Deviation Figure 3: Differences (in days) between predicted and actual harvest dates on barley fields for linear (top) and quadratic polynomial (bottom) model. The comparison of harvest dates predicted with linear and polynomial regression compared to actual harvest dates can be seen in Tables 1 and 2. The polynomial regression has a lower average error than the linear regression for both crops, so its predictions were closer to the actual harvest dates in both cases. Figures 3 and 4 show the residuals, i.e. the differences in days between actual and predicted harvest days, for barley and wheat fields, respectively. The opinion of agricultural experts is that all forecasts, that are within a ten day window from the actual 70 Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Horvat et al. harvesting date can be considered acceptable, which is represented with the accuracy column. The polynomial regression outperforms linear regression, as it scored an additional correct sample with barley and four additional samples with wheat. We can see that some predictions are off by multiple weeks. One of the reasons can be attributed to the small quantity of real samples. 4 Conclusion The described procedure is a simple and scalable solution for prediction of harvest dates and simple harvest plan generation. It utilizes Sentinel-2 imagery to obtain a timeseries before the nearing harvests and predicts harvest dates based on timeseries patterns. We found that polynomial regression did a better job at making those predictions and that the generated harvest plan can be organised as a simulation of how the harvest could be performed. The current method does not take into account real world limitations, such as weather events, which could prove vital if included in the harvest plan generation. Additional data would also be of great importance because with a large enough dataset a supervised machine learning model could be build that could predict the harvest dates with greater accuracy. 0 1 2 3 4 5 6 7 Sample index 0 10 20 30 40 50 60 Deviation Deviation from baseline (0) - Wheat 0 1 2 3 4 5 6 7 Sample index 5 0 5 10 15 20 Deviation Figure 4: Differences (in days) between predicted and actual harvest dates on wheat fields for linear (top) and quadratic polynomial (bottom) model. 5 Acknowledgments This research has received funding from the HORIZON Europe programme through project PrAEctiCe under grant agreement no. 101084248 and Research Programme P2-0041. I would also like to thank ITC for providing the actual harvest dataset. 6 References [1] Arturo G. Cauba, Roshanak Darvishzadeh, Michael Schlund, Andrew Nelson, and Alice Laborte. 2025. Estimation of transplanting and harvest dates of rice crops in the philippines using sentinel-1 data. Remote Sensing Applications: Society and Environment, 37, 101435. doi: https://doi.org/10.1016/j.rsase.202 4.101435. [2] European Commission. 1962. Retrieved August 21, 2025 from https://agricu lture.ec.europa.eu/common-agricultural-policy_en. [3] ESA. 2014. Sentinel-1. Retrieved August 21, 2025 from https://dataspace.cop ernicus.eu/data-collections/sentinel-data/sentinel-1. [4] ESA. 2015. Sentinel-2. Retrieved August 21, 2025 from https://www.esa.int /Applications/Observing_the_Earth/Copernicus/Sentinel-2. [5] Robin Gebbers and Viacheslav I. Adamchuk. 2010. Precision agriculture and food security. Science, 327, 5967, 828–831. eprint: https://www.science.org /doi/pdf/10.1126/science.1183899. doi: 10.1126/science.1183899. [6] Anatoly A. Gitelson, Yoram J. Kaufman, and Mark N. Merzlyak. 1996. Use of a green channel in remote sensing of global vegetation from eos-modis. Remote Sensing of Environment, 58, 3, 289–298. doi: https://doi.org/10.1016 /S0034-4257(96)00072-7. [7] Garima Gupta. 2025. Applications of ai in precision agriculture. Discover Agriculture.doi: 10.1007/s44279-025-00220-9. [8] Rob Hyndman and G. Athanasopoulos. 2021. Forecasting: Principles and Practice. English. (3rd ed.). OTexts, Australia. [9] Olena Kavats, Dmitriy Khramov, Kateryna Sergieieva, and Volodymyr Vasyliev. 2019. Monitoring harvesting by time series of sentinel-1 sar data. Remote Sensing, 11, 21. doi: 10.3390/rs11212496. [10] James Y. Kim. 2024. Open-source software for satellite-based crop health monitoring. Journal of Biosystems Engineering.doi: 10.1007/s42853-024-002 42-z. [11] Emogine Mamabolo, Makgabo Johanna Mashala, Ephias Mugari, Tlou Elizabeth Mogale, Norman Mathebula, Kabisheng Mabitsela, and Kwabena Kingsley Ayisi. 2025. Application of precision agriculture technologies for crop protection and soil health. Smart Agricultural Technology, 12, 101270. doi: https://doi.org/10.1016/j.atech.2025.101270. [12] Gordan Mimić, {Amit Kumar} Mishra, Miljana Marković, Branislav Živaljević, Dejan Pavlović, and Oskar Marko. 2025. Machine learning-based harvest date detection and prediction using sar data for the vojvodina region (serbia). English. Sensors, 25, 7, (Apr. 2025). Publisher Copyright: ©2025 by the authors. doi: 10.3390/s25072239. [13] David J. Mulla. 2013. Twenty five years of remote sensing in precision agriculture: key advances and remaining knowledge gaps. Biosystems Engineering, 114, 4, 358–371. Special Issue: Sensing Technologies for Sustainable Agriculture. doi: https://doi.org/10.1016/j.biosystemseng.2012.08.009. [14] Ali Nasrallah, Nicolas Baghdadi, Mohammad El Hajj, Talal Darwish, Hatem Belhouchette, Ghaleb Faour, Salem Darwich, and Mario Mhawej. 2019. Sentinel-1 data for winter wheat phenology monitoring and mapping. Remote Sensing, 11, 19. doi: 10.3390/rs11192228. [15] Dorijan Radočaj, Mladen Jurišić, and Mateo Gašparović. 2022. The role of remote sensing data and methods in a modern approach to fertilization in precision agriculture. Remote Sensing, 14, 3. doi: 10.3390/rs14030778. [16] Jan Verbesselt, Rob Hyndman, Achim Zeileis, and Darius Culvenor. 2010. Phenological change detection while accounting for abrupt and gradual trends in satellite image time series. Remote Sensing of Environment, 114, 12, 2970–2980. doi: https://doi.org/10.1016/j.rse.2010.08.003. [17] George Worrall, Jasmeet Judge, Kenneth Boote, and Anand Rangarajan. 2023. In-season crop phenology using remote sensing and model-guided machine learning. Agronomy Journal, 115, 3, 1214–1236. eprint: https://acse ss.onlinelibrary.wiley.com/doi/pdf/10.1002/agj2.21230. doi: https://doi.org /10.1002/agj2.21230. 71 Reducing #SAT to 𝑘-clique enumeration S´andor Szab´o sszab[email protected] University of P´ecs Pecs, Hungary Bogd´an Zav´alnij [email protected] HUN-REN Alfred Renyi Institute of Mathematics Budapest, Hungary Abstract Complexity theory states that some complex problem can be reduced to another problem. The reduction from SAT to 𝑘 -CLIQUE was already introduced by Karp. Here we propose an alternative graph reformulation, that can be efficiently used for solving the #SAT problem. Keywords model counting, combinatorial optimization, k-clique, SAT, reformulation 1 Introduction That is actually the base of complexity theory that some complex problem can be reduced to another problem. The reduction from SAT to 𝑘 -CLIQUE was already introduced by Karp [ 3 ]. Thus most textbooks that include chapter on NP-completeness include reduction from 3-SAT to 𝑘 - CLIQUE, as for example [ 2 , pp. 1087–1089]. A small textbook like example pictured in Figure 1. Here the SAT example consists of three clauses, each consisting of three literals. Namely Φ = 𝐶1∧𝐶2∧𝐶3, 𝐶1=𝑥1∨𝑥2∨ ¬𝑥3, 𝐶2=¬𝑥1∨ ¬𝑥2∨𝑥4, 𝐶3=𝑥2∨ ¬𝑥4∨𝑥3. In short, we construct a graph. The number of the nodes is equal to number of the literals in the SAT problem. Each node associated with a certain literal is not adjacent to any node adjacent to a literal in the same clause; and two nodes associated with literals 𝑙1, 𝑙2 from different clauses are not adjacent if 𝑙1∧𝑙2 is a contradiction. All other pairs are adjacent. The constructed graph is 𝑘 -partite where 𝑘 is the number of clauses, and each 𝑘 -clique is giving us a solution of the proposed SAT problem. There are obviously several ways to accomplish such reformulation. We will introduce one, that may possibly occurred elsewhere before. There is a small problem with the textbook way, and that causes two different drawbacks. The problem is basically that two different 𝑘 -cliques can encode the same solution of the SAT problem. The reason behind another reformulation is twofold. First, for the SAT reformulation use a symmetry breaking and produce a graph with less edges and less 𝑘 -cliques. Second, the textbook reformulation is hardly usable for #SAT but the Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society 2024, 7–11 October 2024, Ljubljana, Slovenia © 2024 Copyright held by the owner/author(s). Figure 1: Φ = ( 𝑥1∨𝑥2∨ ¬𝑥3 ) ∧ ( ¬𝑥1∨ ¬𝑥2∨𝑥4 ) ∧ ( 𝑥2∨ ¬𝑥4∨𝑥3) x1 x2 ¬x3 ¬x1 ¬x2 x4 x2 ¬x4 x3 proposed method is. This short paper is focusing only on the #SAT problem. 2 Proposed reformulation We will describe a reduction of the SAT problem to the 𝑘 -clique problem which differs from the commonly encountered reduction. The new 𝑘 -clique reformulation of the SAT problem has the added benefit that all the possible solutions of the SAT problem and all the possible solutions of the 𝑘 -clique problem are in a well defined connection. Consequently, the 𝑘 -clique problem can be used to list all possible solutions of the SAT problem. Let 𝐶1,...,𝐶𝑘 be clauses over the propositional variables 𝑥1,...,𝑥𝑛 . For our considerations we need a more detailed description of the clauses. Namely, for each 𝑖 , 1 ≤𝑖≤𝑘 set 𝐶𝑖=𝑙𝑖,1∨𝑙𝑖,2∨ · · · ∨ 𝑙𝑖,𝑠(𝑖), where 𝑙𝑖,1, 𝑙𝑖,2,...,𝑙𝑖,𝑠(𝑖) are literals of the variables 𝑥1,...,𝑥𝑛 . As a first step we assign the expression 𝐿𝑖,𝑗 =𝑙𝑖,𝑗 ∧(¬𝑙𝑖,1∧ · · · ∧ ¬𝑙𝑖,𝑗−1) to the literal 𝑙𝑖,𝑗 for each 𝑖,𝑗, 1 ≤𝑖≤𝑘, 1 ≤𝑗≤𝑠(𝑖). For our previous small textbook like example the described expressions will be the following: 𝐿1,1=𝑥1, 𝐿1,2=𝑥2∧ ¬𝑥1, 𝐿1,3=¬𝑥3∧ ¬𝑥1∧ ¬𝑥2. 𝐿2,1=¬𝑥1, 𝐿2,2=¬𝑥2∧𝑥1, 𝐿2,3=𝑥4∧𝑥1∧𝑥2. 𝐿3,1=𝑥2, 𝐿3,2=¬𝑥4∧ ¬𝑥2, 𝐿3,3=𝑥3∧ ¬𝑥2∧𝑥4. Altogether nine 𝐿𝑖𝑗 formulas are associated with the SAT instance. We would like to point out that the 𝐿𝑖𝑗 formulas are not clauses but it will not cause any problem in the course of the clique reformulation of the problem. So the equivalent formula for the original expression is: Φ = 𝐾1∧𝐾2∧𝐾3, 𝐾1=𝐿1,1∨𝐿1,2∨𝐿1,3, 𝐾2=𝐿2,1∨𝐿2,2∨𝐿2,3, 𝐾3=𝐿3,1∨𝐿3,2∨𝐿3,3. 72 Information Society 2024, 7–11 October 2024, Ljubljana, Slovenia Szab´o et al. The reader can easily verify that the two formulas are in fact equivalent. We skip the proof, as we do use this fact in our reasoning. As a second step we define a finite simple graph 𝐺 . The nodes of the graph 𝐺 are the conjunctions 𝐿𝑖,𝑗 for each 𝑖 , 𝑗 , 1 ≤𝑖≤𝑘 , 1 ≤𝑗≤𝑠 ( 𝑖 ). Two distinct nodes 𝐿𝑢,𝑣 , 𝐿𝑝,𝑞 of the graph 𝐺 are not adjacent in 𝐺 if the formula 𝐿𝑢,𝑣 ∧𝐿𝑝,𝑞 is a contradiction (always false). See as a small example the new graph for the same textbook problem we described previously in Figure 2. Figure 2: Φ = (𝑥1∨(𝑥2∧ ¬𝑥1)∨(¬𝑥3∧ ¬𝑥1∧ ¬𝑥2)) ∧(¬𝑥1∨(¬𝑥2∧𝑥1)∨(𝑥4∧𝑥1∧𝑥2)) ∧(𝑥2∨(¬𝑥4∧ ¬𝑥2)∨(𝑥3∧ ¬𝑥2∧𝑥4)) x1 x2 ¬x3 ¬x1 ¬x2 x4 x2 ¬x4 x3 ∧¬x1 ∧¬x1 ∧¬x2 ∧x1 ∧x1∧x2 ∧¬x2 ∧¬x2∧x4 Lemma 2.1. The set of nodes {𝐿𝑖,1,...,𝐿𝑖,𝑠(𝑖)} is an independent set in the graph 𝐺 for each 𝑖 ,1 ≤𝑖≤𝑘 . This means that the nodes of the graph 𝐺 can be well-colored using 𝑘 colors, the nodes 𝐿𝑖,1,...,𝐿𝑖,𝑠(𝑖) all can receive color 𝑖. Proof. In order to verify our observation let us consider two distinct nodes 𝐿𝑖,𝑝 , 𝐿𝑖,𝑞 of the graph 𝐺 such that 𝑝 < 𝑞 . The formula 𝐿𝑖,𝑝 contains the literal 𝑙𝑖,𝑝 and the formula 𝐿𝑖,𝑞 contains the literal ¬𝑙𝑖,𝑝 . Consequently, 𝐿𝑖,𝑝 ∧𝐿𝑖,𝑞 is a contradiction. Therefore, the nodes 𝐿𝑖,𝑝 , 𝐿𝑖,𝑞 of the graph 𝐺are not adjacent in 𝐺. Lemma 2.2. If the clauses 𝐶1,...,𝐶𝑘 can be satisfied simultaneously, then the graph 𝐺has a 𝑘-clique. Proof. Let us assume that there is an assignment of the truth values of the propositional variables 𝑥1,...,𝑥𝑛 that makes each of the clauses 𝐶1,...,𝐶𝑘 true. In particular for each 𝑖 , 1 ≤𝑖≤𝑘 there is a literal 𝑙𝑖,𝑢(𝑖) such that the assignment of the truth values makes 𝑙𝑖,𝑢(𝑖) true and makes 𝑙𝑖,1,...,𝑙𝑖,𝑢(𝑖)−1 false. We simply pick the first among the literals 𝑙𝑖,1, 𝑙𝑖,2,...,𝑙𝑖,𝑠(𝑖) which is made true by the truth assignment of the variables. In the 𝑢 ( 𝑖 ) = 1 particular case the set of indices { 1 ,...,𝑢 ( 𝑖 ) − 1 } is empty. Obviously, this assignment of the truth values of the variables makes the conjunction 𝐿𝑖,𝑢(𝑖)=𝑙𝑖,𝑗 ∧(¬𝑙𝑖,1∧ · · · ∧ ¬𝑙𝑖,𝑢(𝑖)−1) true for each 𝑖 , 1 ≤𝑖≤𝑘 . From this follows that 𝐿𝑝,𝑢(𝑝)∧ 𝐿𝑞,𝑢(𝑞) is true for each 𝑝 , 𝑞 , 1 ≤𝑝 < 𝑞 ≤𝑘 . In other words, the formula 𝐿𝑝,𝑢(𝑝)∧𝐿𝑞,𝑢(𝑞) cannot be a contradiction. As a consequence the nodes 𝐿1,𝑢(1), 𝐿2,𝑢(2),...,𝐿𝑘,𝑢(𝑘) of the graph 𝐺are the nodes of a 𝑘-clique in 𝐺.□ Lemma 2.3. If there is a 𝑘 -clique in the graph 𝐺 , then there is an assignment of the truth values of the variables 𝑥1,...,𝑥𝑛 that makes the clauses 𝐶1,...,𝐶𝑘 true simultaneously. Proof. Let us suppose that the graph 𝐺 contains a 𝑘 -clique ∆. We know that the nodes of the graph 𝐺 can be wel-colored using 𝑘 colors. The nodes of the 𝑘 -clique ∆ must receive 𝑘 pair-wise distinct colors. It means that the nodes of the 𝑘 -clique ∆ receive the colors 1 , 2 ,...,𝑘 . There are integers 𝑢(1), 𝑢(2),...,𝑢(𝑘) such that 𝐿1,𝑢(1), 𝐿2,𝑢(2),...,𝐿𝑘,𝑢(𝑘)(1) are the nodes of the 𝑘-clique ∆. Note that the expression 𝐿𝑖,𝑢(𝑖) is true when each of the literals 𝑙𝑖,𝑢(𝑖),¬𝑙𝑖,1,¬𝑙𝑖,2,...,¬𝑙𝑖,𝑢(𝑖)−1 is true. Using this information we can set the values of the corresponding propositional variables 𝑥1,...,𝑥𝑛 to be true or false. Note, that it is not possible to assign true and false values at the same time to a given variable, since the nodes (1) pair-wise adjacent in the graph 𝐺 . If a literal in the expression 𝐿𝑢,𝑣 forces us to set the value of the variable 𝑥𝑖 to be true and a literal in the expression 𝐿𝑥,𝑦 forces us to assign the variable 𝑥𝑖 the false value, then the expression 𝐿𝑢,𝑣 ∧𝐿𝑥,𝑦 contains 𝑥𝑖∧ ¬𝑥𝑖 and consequently 𝐿𝑢,𝑣 ∧𝐿𝑥,𝑦 is a contradiction. This violates the fact that the nodes 𝐿𝑢,𝑣,𝐿𝑥,𝑦 of the graph 𝐺are adjacent in the 𝑘-clique ∆. Of course, it well may happen that for some variables among 𝑥1,...,𝑥𝑛 we are not forced to assign any of the true or false values. In this case the value of such variable is not restricted and we are free to choose between the truth or false values. In other words a 𝑘 -clique in the graph 𝐺 may give rise to more than one truth assignments of the truth values of the variables that satisfy the clauses 𝐶1,...,𝐶𝑘 simultaneously. Namely there will be 2 𝑧 of such assignments if the number of unsigned variables was 𝑧 . □ Remember, that in the graph 𝐺 two distinct nodes 𝐿𝑢,𝑣 , 𝐿𝑥,𝑦 are not adjacent if the formula 𝐿𝑢,𝑣 ∧𝐿𝑥,𝑦 is a contradiction. Let 𝛾 be an assignment of truth values to the propositional variables 𝑥1,...,𝑥𝑛 . Using 𝛾 we color certain edges of the graph 𝐺 . The edge connecting the nodes 𝐿𝑢,𝑣 , 𝐿𝑥,𝑦 receives red color if the truth value assignment 𝛾 makes the expression 𝐿𝑢,𝑣 ∧𝐿𝑥,𝑦 true. If each edge of a 𝑘 -clique ∆ in the graph 𝐺 is red, then we call the 𝑘 -clique ∆ a red 𝑘 -clique in the graph 𝐺 . We call the assignment 𝛾 a satisfying assignment if 𝛾 makes each of the clauses 𝐶1,...,𝐶𝑘 true simultaneously. Lemma 2.2 can can be restated in the following way. Lemma 2.4. If 𝛾 is a satisfying assignment of the truth values of the variables 𝑥1,...,𝑥𝑛 , then there is exactly one red 𝑘-clique in the graph 𝐺. Proof. Lemma 2.2 proved that there is at least one such clique, as it is constructed one. Assume on the contrary that there are two non-identical red 𝑘 -cliques ∆ 1 , ∆ 2 in the graph 𝐺 . Since the red 𝑘 -cliques ∆ 1 , ∆ 2 are not identical, there is a color class, say the 𝑖 -th color class, such that the nodes 𝐿𝑖,𝑝 , 𝐿𝑖,𝑞 of ∆ 1 , ∆ 2 in the 𝑖-th color class are not identical. As 𝐿𝑖,𝑝 is an end point of a red edge in the red 𝑘 - clique ∆ 1 , the assignment 𝛾 makes the expression 𝐿𝑖,𝑝 73 Reducing #SAT to 𝑘-clique enumeration Information Society 2024, 7–11 October 2024, Ljubljana, Slovenia true. Similarly, 𝛾 makes the expression 𝐿𝑖,𝑞 true. But from Lemma 2.1 we know that 𝐿𝑖,𝑝 ∧𝐿𝑖,𝑞 is a contradiction, so 𝐿𝑖,𝑝 , 𝐿𝑖,𝑞 cannot be true at the same time, so ∆ 1 , ∆ 2 must be identical. □ Theorem 2.5. The proposed reformulation along with the 𝑘 -clique enumerating method, where for each clique one adds up numbers 2 𝑧 , where 𝑧 is the number of unset variables, solves the #SAT problem. Proof. Concludes from Lemmas 2.2,2.3 and 2.4. 3 Discussion One can notice, that if the ordering of the literals are fixed in the same way in each clause then the proposed method is very similar to the DPLL method adjusted for #SAT in [ 1 ]. The dedicated reader will realize that in the proposed method rearranging the literals among each other in a clause will lead to essentially different graphs in the clique reformulation. These graphs not only have different number of edges, but the number of 𝑘 -cliques can also differ. At the first glance this may look as a disadvantage. However, there is another way to look at this phenomenon. The sizes of the search spaces vary with permuting the literals in the original SAT instance. Thus it opens up an avenue reducing the size of the search tree by tactically choosing among the possible rearrangements of the literals. It gives a certain flexibility, as one can use different methods (nowadays for example artificial intelligence) to find out which ordering and thus reformulation would give the best approach in terms of solution speed. Also, one can apply preconditioning methods for the resulting graph based on graph theoretical considerations, like proposed in [ 5 ]. It may be much easier to enumerate all 𝑘 -cliques in the resulting graph, this is a future goal of our research. We also did a tiny computational experiment. We took two simple examples from [ 4 ], namely from Section 4, with parameters 𝑡 = 2 , 3, where the SAT formula has 8 𝑡 + 22 clauses and each of these clauses has length three, and the formula is unsatisfiable, that is the number of solution is zero. So our first 3-SAT example has 38 clauses, and the second example has 46. The textbook reformulation of the first problem gave a graph with 114 nodes and 6 218 edges. The clique search algorithm we used lead to a search tree of size 15 414 191. After an edge dominance preconditioning from [ 5 ] we got a simplified graph of 5 967 edges, and the clique search lead to a search tree of size 28 474. The proposed reformulation of the first problem gave a graph with 114 nodes and 5 953 edges. The clique search algorithm we used lead to a search tree of size 92 296. After an edge dominance preconditioning the simplified graph had 5 619 edges, and the clique search lead to a search tree of size 18 202. The textbook reformulation of the second problem gave a graph with 138 nodes and 9 174 edges. The clique search algorithm lead to a search tree of size 2 497 099 319. After an edge dominance preconditioning we got a simplified graph of 8 730 edges, and the clique search lead to a search tree of size 142 399. The proposed reformulation of the second problem gave a graph with 138 nodes and 8 805 edges. The clique search algorithm we used lead to a search tree of size 922 996. After an edge dominance preconditioning the simplified graph had 8 233 edges, and the clique search lead to a search tree of size 72 810. It is clear from these examples, that both preconditioning and the proposed reformulation may greatly aid the calculation in some cases. We close this section with some clarifying notes. The preconditioning rules used in connection with 𝑘 -clique problems can be sorted into two broader classes. Namely, 𝑘 - clique preserving and 𝑘 -clique loosing rules. The 𝑘 -clique preserving rules reduces the given graph 𝐺 to a new graph 𝐺′ in such a way that each 𝑘 -clique in 𝐺 is also a 𝑘 -clique in the reduced graph 𝐺′ . On the other hand, the 𝑘 -clique loosing rules reduces the given graph 𝐺 to a new graph 𝐺′′ such that some 𝑘 -clique in 𝐺 may not be a 𝑘 -clique in the reduced graph 𝐺′′ but at least one of the 𝑘 -cliques of the original graph 𝐺 still will be a 𝑘 -clique in the reduced graph 𝐺′′ . Therefore, when our purpose is to enumerate all the possible 𝑘 -cliques of 𝐺 , then we cannot simply apply 𝑘 -clique loosing preconditioning rules. However, when our purpose is to verify that the given graph 𝐺 does not have any 𝑘 -clique, then we may apply 𝑘 -clique loosing preconditioning rules. The edge doninance preconditioning rule is a 𝑘 -clique loosing preconditiong rule. In the same time it was known that the examples from [ 4 ] are unsatisfiable. Therefore using 𝑘-clique loosing rules are fully justified in this case. On the other hand, surprisingly, even when there are several 𝑘 -cliques present and the goal is to count them, one can use clique loosing preconditioning like dominance in the following way. In such a case, one need to keep a list of these transformations, and during the enumeration each solution needs to be check against that list. This is justified by the fact that such transformations always delete a node (edge) while pointing to another one as a possible substitution. That is, if node 𝑢 dominates node 𝑣 and we deleted node 𝑣 , then during the solution listing process we need to check each solution if it contains (the remaining) node 𝑢 , and if substituting back node 𝑣 instead of node 𝑢 also gives a valid solution. This approach however is not developed yet, and certainly needs a more elaborate program for listing solutions. References [1] Elazar Birnbaum and Eliezer L. Lozinskii. The good old davisputnam procedure helps counting models. Journal of Artificial Intelligence Research, 10(1):457–477, 1999. [2] Thomas H. Cormen, Charles E. Leiserson, Ronald L. Rivest, and Clifford Stein. Introduction to Algorithms. The MIT Press, third edition, 2009. [3] Richard M. Karp. Reducibility among combinatorial problems. In R. E. Miller, J. W. Thatcher, and J.D. Bohlinger, editors, Complexity of Computer Computations, pages 85–103. New York: Plenum, 1972. [4] Ming Ouyang. How good are branching rules in dpll? Discrete Applied Mathematics, 89(1):281–286, 1998. [5] S´andor Szab´o and Bogd´an Zav´alnij. Clique search in graphs of special class and job shop scheduling. Mathematics, 10(5), 2022. 74 Modularity aware graph clustering for exploratory tasks with a case study of the biomass supply chain Sylvert Prian Tahalea [email protected]szeged.hu University of Szeged, Hungary Universitas Pembangunan Nasional Veteran Yogyakarta, Indonesia Arkadiusz Kawa Poznań School of Logistics, Department of Logistics, Poznań Poznań University of Economics and Business, Department of Business Relationships and International Marketing, Poznań Poland Balázs Dávid [email protected] InnoRenew CoE, UP IAM and UP FAMNIT, University of Primorska, Slovenia Abstract This paper proposes a novel modularity-aware graph clustering algorithm that combines label propagation principle and global modularity. The algorithm consists of two phases: (1) the first phase is to quickly form clusters using LPA and approximate modularity to evaluate their quality, and (2) the second phase focuses on refining the cluster structures using global modulation. The proposed algorithm is later evaluated using several metrics with several datasets and applied to a real-world use case. The results reveal a clear trade-off between internal and external cluster quality, which is useful for exploratory tasks. The modularity-aware graph clustering algorithm performed well in the experimental and the real-world cases, producing reliable clusters for each case. Keywords graph clustering, modularity, LPA, exploratory data analysis 1 Introduction Graph clustering has arisen as a critical issue in network research, data mining, and machine learning due to its potential to reveal hidden structures within complex systems. Real-world phenomena such as social interactions, biological processes, communication networks, citation graphs, and transportation systems can be naturally modeled as graphs, with nodes representing things and edges representing their relationships. Identifying cohesive groups of nodes, also known as clusters or communities, allows researchers to get a better understanding of the modular organisation of networks and simplify large-scale data for analysis. Over the past decades, various algorithmic paradigms have arisen, ranging from modularity optimisation, spectral approach, and random walks to label propagation, statistical inference, and most recently, utilising graph neural networks. Each technique strikes a balance between accuracy, scalability, interpretability, and applicability to overlapping or hierarchical communities. Despite these achievements, there are some problems remaining, such as the low modularity for the fast algorithms, efficient methods for large-scale graphs, adaptive strategies for temporal graphs, and the lack of universally acknowledged ground truth. The Label propagation algorithm (LPA) is considered to be one of the fastest graph clustering techniques. However, it has low Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). Information Society 2024, 7–11 October 2024, Ljubljana, Slovenia ©2024 Copyright held by the owner/author(s). modularity [20, 17]. Modularity measurement objectively measures the quality of the communities, where the value near one indicates well-structured communities, and a value near zero indicates a bad community structure. There are several LPA extensions to increase the modularity utilizing its fast runtime, namely LPAm [1], LPAm+ [13], LPA-MNI [12], and FLPA [18]. This paper proposes a novel modularity-aware graph clustering algorithm utilising LPA, local modularity, and global modularity. This algorithm consists of two phases: fast label propagation and global modularity refinement. In the first phase, utilising LPA for the fast propagation and local modularity to regain its initial structure. The second phase focuses on increasing the modularity using global modularity refinement. Finally, the result will provide higher-quality partitions compared to the original LPA. The algorithm is later evaluated using different metrics with several datasets. Moreover, in order to invetsigate the quality of the methodology in real-world scenario, a use case of biomass supply chains is studied. The remainder of the paper is organised as follows. Section 2 provides the methodology and evaluation techniques. Section 3 provides the main results and highlights the main findings. Section 4 provides the use case on a real-world scenario. Section 5 concludes the contribution, limitations, and future research. 2 Methodology 2.1 Modularity Modularity is one of the most influential principles in graph clustering, which is evaluated by the strength of a network’s partition by comparing the density of edges inside communities [15]. High modularity value indicates that clusters have significantly more internal connections than external ones. Modularity optimisation has become a benchmark for measuring cluster quality, but the direct optimisation of modularity is challenging due to its NP-hard nature. The modularity measurement [15] is as follows. 𝑄= 1 2𝑚∑︁[𝐴𝑖 𝑗 −𝑘𝑖𝑘𝑗 2𝑚]𝛿(𝑐𝑖, 𝑐𝑗),(1) where 𝐴𝑖 𝑗 represents the actual connection between nodes 𝑖 and 𝑗 , the term 𝑘𝑖𝑘𝑗 2𝑚 is the expected number of edges between nodes 𝑖and 𝑗, and 𝑐𝑖is the community assignment. To avoid recomputing modularity for every possible partition, Blondel et al. [2] proposed a local measurement that evaluates the modularity change if a node were moved from its original cluster to the neighboring cluster and also provides the approximate modularity. The approximate modularity is as follows. Δ𝑄(𝑣,𝑙) ≈ 𝑘𝑣,𝑖𝑛 𝑚−𝑘𝑣·sum_tot(𝑙) 2𝑚2(2) 75 Information Society 2024, 7–11 October 2024, Ljubljana, Slovenia Tahalea et al. Algorithm 1 Fast propagation phase Require: Graph 𝐺=(𝑉, 𝐸) Ensure: Node labels 𝐿(𝑣)for all 𝑣∈𝑉 1: Assign each node 𝑣∈𝑉a unique label 𝐿(𝑣) ← 𝑣 2: repeat 3: for all nodes 𝑣∈𝑉in random order do 4: Compute label frequencies among neighbors of 𝑣 5: Select top-𝑘frequent labels (optional) 6: for all candidate labels 𝑙do 7: Estimate local modularity gain Δ𝑄(𝑣, 𝑙 ) 8: end for 9: Select label 𝑙∗with maximum gain (if gain >threshold) 10: Update 𝐿(𝑣) ← 𝑙∗ 11: end for 12: until no label changes or maximum iterations reached where 𝑘𝑣,𝑖𝑛 is the number of edges from node 𝑣 carrying label 𝑙 ; 𝑘𝑣 is the degree of node 𝑣 ; sum_tot( 𝑙 ) is the total degree of nodes with label (𝑙); and 𝑚is the number of edges in the network. 2.2 Label Propagation Algorithm The label propagation algorithm (LPA) was introduced as one of the simplest yet most scalable methods for graph clustering [16]. The algorithm begins by assigning a unique label to each node, then iteratively updates the node’s label based on the most frequent label among its neighbors. The process is performed in random sequential order to avoid bias, and a label is chosen at random when there are multiple labels that have the same frequency. The iterations continue until every node holds a label that is the majority among the neighbors, leading to convergence in the network. The clusters are formed by grouping the nodes that share the same label. LPA has near-linear time complexity, ability to uncover clusters without prior knowledge of their number, and is scalable for large networks. 2.3 Modularity-Aware Graph Clustering The modularity-aware graph clustering utilizes LPA’s ability to quickly cluster the nodes and modularity refinement to produce better cluster structures. This proposed method is divided into two phases: (1) The first phase is to quickly form clusters using LPA and approximate modularity to evaluate their quality, and (2) the second phase focuses on refining the cluster structures using global modularity measurement. 2.3.1 Phase 1: Fast Propagation Phase. The objective of the first phase is to obtain the local approximation of the modularity gain by rapidly propagating through the network. This step helps to form a coarse graph for further refinement. The initial label for each node is its own node ID; then the algorithm iterates over all the nodes in a randomized order and updates the labels using the local approximation of modularity gain as in Equation (2). This step also leverages the Top-K label filtering to reduce the chance of unstable changes while a minimum gain threshold is applied to prevent weak label changing. The iteration process will continue until there is no label change or the maximum number of iterations is reached. The fast propagation algorithm is presented in Algorithm 1. 2.3.2 Phase 2: Modularity Refinement. The objective of the second phase is to refine the approximation performed in the first phase. Whenever the label is spread across the network, maximal modularity is gained from the overall structure. Therefore, the second phase performs the refinement for each label using global modularity, as in Equation (1), for the entire network. The modularity refinement is presented in Algorithm 2. Algorithm 2 Global Modularity Refinement 1: Input: Label map 𝐿, Graph 𝐺, stability threshold 𝑅 2: Initialize stability counter 𝑆(𝑣) ← 0for all 3: repeat 4: for all nodes 𝑣∈𝑉not marked stable do 5: Compute label frequencies among neighbors 6: Select Top-𝑘candidate labels (optional) 7: for all labels 𝑙do 8: Temporarily assign 𝐿′(𝑣) ← 𝑙 9: Compute new global modularity 𝑄′ 10: end for 11: Select label 𝑙∗that maximizes 𝑄′ 12: if 𝑄′>𝑄then 13: Update 𝐿(𝑣) ← 𝑙∗,𝑆(𝑣) ← 0 14: else 15: 𝑆(𝑣) ← 𝑆(𝑣) + 1 16: if 𝑆(𝑣) ≥ 𝑅then 17: Mark 𝑣as stable 18: end if 19: end if 20: end for 21: until no label changes or improvement <𝜖 2.4 Evaluation The proposed algorithm is evaluated using several common metrics that are widely accepted for graph clustering, such as the modularity measurement, normalized mutual information (NMI), Adjusted Rand Index (ARI), and F1-score. The modularity measurement measures the internal structures of the cluster using Equation (1), while NMI measures the similarity between the result cluster and the ground truth of the dataset as follows [4]. 𝑁𝑀𝐼 (𝐴, 𝐵)= 2·𝐼(𝐴, 𝐵) 𝐻(𝐴) + 𝐻(𝐵)(3) where 𝐼(𝐴, 𝐵) is the mutual information between clusters 𝐴 and 𝐵; and 𝐻(𝐴)and 𝐻(𝐵)are the entropies of the clusters. The Adjusted Rand Index (ARI) is used to measure similarity between generated clusters and ground truth [6], while F1-score evaluates the precision and recall for each cluster [14, 3]. The ARI is measured using Equation (4) and F1-score is measured using Equation (5). ARI =Í𝑖 𝑗 𝑛𝑖 𝑗 2−hÍ𝑖𝑎𝑖 2Í𝑗𝑏𝑗 2i/𝑛 2 1 2hÍ𝑖𝑎𝑖 2+Í𝑗𝑏𝑗 2i−hÍ𝑖𝑎𝑖 2Í𝑗𝑏𝑗 2i/𝑛 2 (4) where 𝑛 is the number of nodes; 𝑛𝑖,𝑗 is the number of nodes in both the predicted cluster 𝑖 and the ground truth cluster 𝑗 ; 𝑎𝑖 is the number of nodes in the predicted cluster 𝑖 ; 𝑏𝑗 is the number of nodes in the ground truth cluster 𝑗 ; and 𝑥 2 is a binomial coefficient to count the number of pairs between clusters. 𝐹1= 2·precision ·recall precision +recall = 2·𝑇𝑃 2·𝑇𝑃 +𝐹𝑃 +𝐹 𝑁 (5) where 𝑇𝑃 are correctly predicted positive cases; 𝐹𝑃 are incorrectly predicted as positive; and 𝐹𝑁 are missed positive cases, which are predicted as negative. 3 Experiments and results . 3.1 Experimental Datasets and results There are four real-world datasets used in this research, such as Zachary’s Karate Club Network [5], Football [5], Polbooks [8], and Email [19, 10, 11] from SNAP datasets. The synthetic datasets are also used in this research, utilizing the LFR framework [9]. 76 Modularity aware graph clustering Information Society 2024, 7–11 October 2024, Ljubljana, Slovenia The summary of the datasets is presented in Table 1, where 𝑐 is the number of clusters, 𝑘 is the average degree, 𝜇 is the mixing parameter. The measurement is compared with several algorithms such as LPA, LPAm+, and Louvain. This is meant to compare the proposed algorithm with its predecessors and one of the top graph clustering algorithms. Table 1: Comparative metadata of benchmark datasets Dataset Nodes Edges 𝑐 𝑘 𝜇 Karate 34 78 2 - - Football 115 613 3 - - Polbooks 105 441 12 - - Email 1005 25571 42 - - LFR1 1000 10455 44 20.91 0.10 LFR2 2000 15495 32 15.49 0.15 LFR3 5000 93743 56 37.50 0.30 LFR4 10000 113290 61 22.66 0.10 LFR5 10000 120047 75 24.01 0.20 LFR6 10000 190440 84 38.09 0.25 3.1.1 Modularity and NMI. The modularity evaluation is meant to evaluate the structure of the clusters. The larger value, closer to 1, indicates a densely connected structure within the results, but it does not indicate the accuracy of the community detection. The results presented in Table 2 show that the modularity aware label propagation algorithm has the best performance across the datasets, from the small one to large datasets. The modularity aware label propagation performed at the same level as Louvain in the LFR4-LFR6 with the highest modularity score, showing that it can produce high-quality clusters. The normalized mutual information measures the similarity between clusters generated by the algorithms and the ground truth. Based on the results presented in Table 2, the LPA is the best performer, reaching the highest score for almost all the datasets. The modularity aware label propagation algorithm is the second best for NMI evaluation, with notably one highest score for the Email network dataset, having a huge gap with LPA, which is the best overall NMI performer. 3.1.2 ARI and F1-score. The Adjusted Rand Index (ARI) quantifies the similarity between the generated clusters and the ground truth, considering all pairs of samples and evaluating whether they are assigned to the same cluster. The results presented in Table 3 show that only LPA is the best performer. The modularity aware label propagation only has the highest score once, for the Email network case, while becoming the second best for most other cases. The F1-score serves as an evaluation metric for assessing the quality of cluster assignments. The results presented in Table 3 show that LPA produced high-quality clusters, reaching the highest score for most of the real-world and synthetic datasets. The modularity-aware label propagation algorithm performs as the second best and joins the LPA as the best performer for the three largest datasets in this research. 3.2 Discussion The results highlight a clear distinction between internal and external clustering quality. The modularity-aware label propagation consistently maximizes the modularity, as presented in Table Table 2: Comparison of algorithms in terms of modularity and NMI across datasets Dataset Modularity LPA LPAm+ MA-LPA Louvain Karate 0.309 0.418 0.445 0.427 Polbook 0.481 0.493 0.528 0.519 Football 0.583 0.557 0.604 0.596 Email 0.089 0.418 0.432 0.431 LFR1 0.807 0.809 0.810 0.808 LFR2 0.479 0.514 0.515 0.511 LFR3 0.569 0.735 0.742 0.731 LFR4 0.829 0.824 0.830 0.830 LFR5 0.839 0.840 0.841 0.841 LFR6 0.611 0.609 0.613 0.613 NMI LPA LPAm+ MA-LPA Louvain Karate 1.000 0.565 0.607 0.483 Polbooks 1.000 0.611 0.759 0.697 Football 1.000 0.919 0.937 0.939 Email 0.180 0.593 0.663 0.575 LFR1 0.932 1.000 0.964 0.993 LFR2 1.000 0.831 0.854 0.823 LFR3 0.301 0.602 0.615 0.598 LFR4 1.000 0.986 0.989 0.964 LFR5 1.000 0.979 0.994 0.995 LFR6 1.000 0.961 0.975 0.963 Table 3: Comparison of algorithms in terms of ARI and F1-score across datasets Dataset ARI LPA LPAm+ MA-LPA Louvain Karate 1.000 0.398 0.597 0.483 Polbooks 1.000 0.509 0.710 0.579 Football 1.000 0.833 0.842 0.892 Email 0.011 0.332 0.440 0.315 LFR1 0.778 1.000 0.905 0.976 LFR2 1.000 0.400 0.411 0.387 LFR3 0.177 0.471 0.487 0.462 LFR4 1.000 0.911 0.966 0.959 LFR5 1.000 0.955 0.980 0.979 LFR6 1.000 0.918 0.899 0.903 F1-score LPA LPAm+ MA-LPA Louvain Karate 1.000 0.807 0.956 0.850 Polbooks 1.000 0.688 0.897 0.665 Football 1.000 0.869 0.986 0.745 Email 0.078 0.128 0.261 0.130 LFR1 0.546 1.000 1.000 0.710 LFR2 1.000 0.907 0.998 0.855 LFR3 0.185 0.491 0.512 0.498 LFR4 1.000 0.661 1.000 0.877 LFR5 1.000 0.759 1.000 0.0.993 LFR6 1.000 0.796 1.000 0.837 2, indicating densely connected clusters in the graph topology. However, the standard LPA attains the strongest argument with ground truth across most datasets, suggesting that it works better 77 Information Society 2024, 7–11 October 2024, Ljubljana, Slovenia Tahalea et al. for datasets with known labels. The ARI and F1-score results in Table 3 show the consistent trade-offs of modularity-aware label propagation, which is optimizing the global structural objective not always gives the best intended partition against known labels. This means that the algorithm choice should reflect the evaluation goal. When the task is exploratory and the ground truth is not available, maximizing internal structure is attractive for discovering cohesive clusters; thus, modularity-aware label propagation algorithms are better. When the aim is to recover known categories or enable downstream tasks, LPA is safer given its superior NMI, ARI, and F1-score. 4 Case study 4.0.1 Dataset. This case study draws on a quantitative survey conducted in late 2022 within the BioLog project, aimed at developing a model of the reverse supply chain of residual wood biomass. The survey targeted enterprises across the wood sector—sawmills, furniture and joinery producers, recyclers, and logistics providers—using a stratified random sampling method based on PKD (The Polish Classification of Activities) industry codes and a project-specific list. Data were collected through 300 Computer Assisted Telephone Interviews (CATI) with company representatives knowledgeable about wood residue origins and management. The questionnaire gathered company profile data (value chain position, employment size, revenue) and explored types of wood residues used (forestry/agricultural, postproduction, post-consumer) as well as their applications in energy generation, energy carriers, and wood-based materials. It also investigated logistics and technological processes, cooperation networks, requirements, and constraints in biomass management, and adaptations in transport, storage, and unitization practices. The dataset was anonymised, and respondents were assured of confidentiality, informed consent, and voluntary participation. 4.0.2 Results. The dataset underwent a series of preprocessing steps before being analyzed. The answers to the questionnaire were transformed into categorical and numerical data. The Interquartile Range (IQR) performed to detect the outlier data and resulted in reducing the number of questions to only include relevant ones. The clean data was used to model a graph with the nodes representing the firms and the edge weight between any two companies using position-aware Jaccard similarity [7]: essentially measuring the fraction of survey questions on which the two firms gave identical answers. There were several versions of graph models such as unweighted graph, plain weighted graph, and weighted graph with thresholds (0.1, 0.125, and 0.25) The graph indicates that most businesses are connected through shared practices. Out of 300 firms, 272 (about 91%) form one large connected group, while the other 28 are isolated or in very small clusters, representing outliers. This core network isn’t very dense, but it does have clusters: companies that work with the same partner are commonly connected to each other, creating tiny groupings with shared qualities. The average distance between companies is modest (approximately 3 steps), and the longest distance is tiny (6–7 steps), which gives the network a "smallworld" shape. In practice, this means that any two firms, even from different parts of the sector, can be connected through only a few intermediaries. The proposed algorithm applied to this data, alongside with Louvain, LPA, and LPAm+ to identify the clusters. The graph clustering methods uncovered a non-trivial clustering structure in the data, partitioning the firms into clusters without any apriori categorization. However, the results demonstrate that explicitly optimizing for modularity yields a superior segregation of the network. Although no explicit categories (such as firm type or size) were given to the clustering algorithms, the communities they detected appear to reflect meaningful behavioral and operational patterns among the companies. In other words, firms ended up clustered together because they answered many survey questions in similar ways – a purely data-driven outcome that likely corresponds to real-world commonalities. For example, one cluster derived from the graph predominantly consists of sawmills and wood processors that generate large volumes of residues and use them for bioenergy, while another cluster groups manufacturers (e.g. furniture or flooring producers) that have different residue uses and logistics practices. Indeed, the analysis suggests that companies naturally form a few distinct sub-communities: even without predefining any segments, those with analogous supply chain roles, residue utilization strategies, or challenges tend to congregate in the same community. This insight is valuable for policymakers and industry stakeholders – it implies the sector can be segmented into groups with shared characteristics, which may each benefit from tailored strategies (for instance, a cluster of firms focused on energy production might face similar regulatory and technological issues). In constructing the similarity measure, all survey features were treated with equal weight. This methodological choice (i.e. not assigning higher importance to any particular question or topic) means the clustering was driven by overall similarity across many attributes. A side effect is that very common attributes (such as broadly adopted practices) contribute to linking many firms, potentially overshadowing rarer but distinctive features. In the present case, however, even unweighted features yielded intelligible clusters, indicating that the dominant patterns in the data were strong enough to shape the communities. Future work could experiment with feature weighting (for instance, giving more emphasis to specific key questions or rare responses) to see if even clearer or more nuanced groupings emerge. Nonetheless, the current graph clustering results already highlight consistent patterns: companies with similar operational profiles gravitated into the same clusters. In summary, the modularity-aware graph clustering of the BioLog survey data uncovers a modular structure in the wood biomass supply chain, revealing that despite the lack of explicit grouping criteria, firms naturally aggregate into network communities that mirror their shared behaviors and challenges. The performance of the algorithms is presented in Table 4. 5 Conclusion In conclusion, the results reveal a clear trade-off between internal and external cluster quality; the modularity-aware label propagation consistently attains the highest modularity, which is useful for exploratory tasks. The standard LPA achieved the best agreement with ground truth on NMI, ARI, and F1-score across most real and synthetic datasets, showing it’s better for optimizing downstream tasks. Overall, these findings argue against a single best method and support choosing the algorithm according to the evaluation objective and data regime. Moreover, a real-world case study of biomass supply chains demonstrates that our methodology provides results for strategic decision making. 78 Modularity aware graph clustering Information Society 2024, 7–11 October 2024, Ljubljana, Slovenia Table 4: Algorithms performance in the case study Unweighted Graph LPA LPAm+ MA-LPA Louvain No. of Cluster(s) 1 1 31 Modularity 0 0 0.259 0 Weighted Graph LPA LPAm+ MA-LPA Louvain No. of Cluster(s) 1 1 86 Modularity 0 0 0.145 0.147 Weighted Graph (0.1) LPA LPAm+ MA-LPA Louvain No. of Cluster(s) 1 1 10 8 Modularity 0 0 0.155 0.139 Weighted Graph (0.125) LPA LPAm+ MA-LPA Louvain No. of Cluster(s) 1 1 10 10 Modularity 1 1 0.153 0.146 Weighted Graph (0.25) LPA LPAm+ MA-LPA Louvain No. of Cluster(s) 1 1 4 10 Modularity 0 0 0.266 0.238 Finally, there are several methodological considerations that need to be improved. Future work could explore several improvements, such as (1) multi-objective formulation that balances modularity with information-theoretic alignment; (2) considering variability over multiple runs and using consensus clustering; and (3) examining performance under different mixing parameters and larger datasets. Acknowledgements The research was supported by the BioLOG project: the authors are grateful for the support of the National Center of Science (NCN) through grant DEC-2020/39/I/HS4/03533, the Slovenian Research and Innovation Agency (ARIS) through grant N1-0223, and the Austrian Science Fund (FWF) through grant I 5443-N. Balázs Dávid is also grateful for the support of the Slovenian Research and Innovation Agency (ARIS) through grant J1-50000, and gratefully acknowledges the Slovenian Research and Innovation Agency (ARIS) and the Ministry of the Economy, Tourism and Sport (MGTŠ) for the grant V4-2512. References [1] Michael J Barber and John W Clark. 2009. Detecting network communities by propagating labels under constraints. Physical Review E—Statistical, Nonlinear, and Soft Matter Physics, 80, 2, 026129. doi: 10.1103/PhysRevE.80.0261 29. [2] Vincent D Blondel, Jean-Loup Guillaume, Renaud Lambiotte, and Etienne Lefebvre. 2008. Fast unfolding of communities in large networks. Journal of statistical mechanics: theory and experiment, 2008, 10, P10008. doi: 10.1088/1 742-5468/2008/10/P10008. [3] Peter Christen, David J Hand, and Nishadi Kirielle. 2023. A review of the fmeasure: its history, properties, criticism, and alternatives. ACM Computing Surveys, 56, 3, 1–24. [4] Leon Danon, Albert Diaz-Guilera, Jordi Duch, and Alex Arenas. 2005. Comparing community structure identification. Journal of statistical mechanics: Theory and experiment, 2005, 09, P09008. doi: 10.1088/1742-5468/2005/09 /P09008. [5] Linton C Freeman. 1977. A set of measures of centrality based on betweenness. Sociometry, 35–41. doi: 10.2307/3033543. [6] Lawrence Hubert and Phipps Arabie. 1985. Comparing partitions. Journal of classification, 2, 1, 193–218. doi: 10.1007/BF01908075. [7] Paul Jaccard. 1901. Étude comparative de la distribution florale dans une portion des alpes et des jura. Bull Soc Vaudoise Sci Nat, 37, 547–579. [8] Valdis Krebs. 2004. Books about us politics. unpublished, http://www. orgnet. com. [9] Andrea Lancichinetti, Santo Fortunato, and Filippo Radicchi. 2008. Benchmark graphs for testing community detection algorithms. Physical Review E—Statistical, Nonlinear, and Soft Matter Physics, 78, 4, 046110. doi: 10.1103 /PhysRevE.78.046110. [10] Jure Leskovec, Jon Kleinberg, and Christos Faloutsos. 2007. Graph evolution: densification and shrinking diameters. ACM transactions on Knowledge Discovery from Data (TKDD), 1, 1, 2–es. doi: 10.1145/1217299.1217301. [11] Jure Leskovec and Andrej Krevl. 2014. SNAP Datasets: Stanford large network dataset collection. http://snap.stanford.edu/data. (2014). [12] Huan Li, Ruisheng Zhang, Zhili Zhao, and Xin Liu. 2021. Lpa-mni: an improved label propagation algorithm based on modularity and node importance for community detection. Entropy, 23, 5, 497. doi: 10.3390/e23050497. [13] Xin Liu and Tsuyoshi Murata. 2010. Advanced modularity-specialized label propagation algorithm for detecting communities in networks. Physica A: Statistical Mechanics and its Applications, 389, 7, 1493–1500. doi: 10.1016/j.p hysa.2009.12.019. [14] Christopher D Manning. 2008. Introduction to information retrieval. Syngress Publishing. [15] Mark EJ Newman and Michelle Girvan. 2004. Finding and evaluating community structure in networks. Physical review E, 69, 2, 026113. doi: 10.1103 /PhysRevE.69.026113. [16] Usha Nandini Raghavan, Réka Albert, and Soundar Kumara. 2007. Near linear time algorithm to detect community structures in large-scale networks. Physical Review E—Statistical, Nonlinear, and Soft Matter Physics, 76, 3, 036106. doi: 10.1103/PhysRevE.76.036106. [17] Seema Rani and Monica Mehrotra. 2017. Hybrid influential centrality based label propagation algorithm for community detection. In 2017 International Conference on Computing, Communication and Automation (ICCCA). IEEE, 11–16. doi: 10.1109/CCAA.2017.8229801. [18] Vincent A Traag and Lovro Šubelj. 2023. Large network community detection by fast label propagation. Scientific Reports, 13, 1, 2701. doi: 10.1038/s4 1598-023-29610-z. [19] Hao Yin, Austin R Benson, Jure Leskovec, and David F Gleich. 2017. Local higher-order graph clustering. In Proceedings of the 23rd ACM SIGKDD international conference on knowledge discovery and data mining, 555–564. doi: 10.1145/3097983.3098069. [20] Aiping Zhang, Guang Ren, Yejin Lin, Baozhu Jia, Hui Cao, Jundong Zhang, and Shubin Zhang. 2014. Detecting community structures in networks by label propagation with prediction of percolation transition. The Scientific World Journal, 2014, 1, 148686. doi: 10.1155/2014/148686. 79 Sphere Target-Based Point Cloud Registration Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia yet determined in red. These three colours are used consistently in Figures 2b–c. An important observation at this stage is that the centre 𝐶 of the sphere must represent the vertex of a right circular cone above 𝑐123 , i.e. it lies on 𝑙123 . In the second step, we determine the orthogonal distance 𝑓 from the fourth point 𝑃4 of the quartet to Σ123 , the distance 𝑒=|𝑃4𝐶123| , and angle 𝛽 (Figure 2b). Finally, we use 𝑒 , 𝛽 , and the cosine theorem in triangles Δ𝑃4𝐶123𝐶 and Δ𝑃1𝐶𝐶123 to determine the distance 𝑑=|𝐶𝐶123| and the sphere target radius 𝑟 . Through 𝑑 , we then also determine the sphere centre 𝐶(Figure 2c). Figure 2: Determination of the sphere centre 𝐶 and radius 𝑟 from points 𝑃1 , 𝑃2 , 𝑃3 , and 𝑃4 on the sphere, when 𝑃4 and 𝐶 are on the same side of Σ123 : a) determining Σ123 , b) describing the relation between 𝑃4 and Σ123 , c) final solution. Algorithm 1 explains this procedure in a compact form. Note that Figure 2 only shows the case where 𝑃4 and 𝐶 are on the same side of Σ123 . The situation with 𝑃4 in the spherical cap is handled by assigning the absolute value to 𝑑 , while the exception with both 𝑃4and 𝐶in Σ123 is caught with the first if-clause. 2.4 Registration The determination of the registration matrix 𝑀 is adopted from [6]. There, it was used to map the point cloud from the 𝐶𝑆𝑆 coordinate system of the LiDAR scanner to the robot’s 𝐶𝑆𝑅 via the intermediate 𝐶𝑆𝐼 . Here, 𝑆 (source) denotes the first LiDAR, 𝑅 (reference, registered) the target coordinate system of the second LiDAR, and 𝐼 the intermediate coordinate system. 𝑀 is determined by three translations and three rotations along/around the coordinate axes. The role of registration is thus to determine the parameters of these six elementary transformations with the best possible accuracy. A triplet of landmark points 𝐴 , 𝐵 , and 𝐶 Algorithm 1 Centre and radius of the sphere from 4 points function SphereThroughFourPoints(𝑃1,𝑃2,𝑃3,𝑃4) 𝐶123 ←centre of circumreference 𝑐123 of Δ𝑃1𝑃2𝑃3 𝑟123 ← |𝐶123𝑃1|⊲Radius of 𝑐123 n123 ←(𝑃1−𝐶123)×(𝑃2−𝐶123) |(𝑃1−𝐶123)×(𝑃2−𝐶123)| ⊲ Norm. vect. Δ𝑃1𝑃2𝑃3 if |𝐶123𝑃4|=𝑟123 then return (𝐶123,𝑟123) end if 𝑇4←orthogonal projection of 𝑃4on the plane of Δ𝑃1𝑃2𝑃3 𝑒← |𝐶123𝑃4| 𝑑← 𝑒2−𝑟1232 2|𝑃4𝑇4|  𝑟←√︁𝑑2+𝑟1232 if (n123 · (𝑃4−𝐶123)>0)⊕(𝑒>𝑟123)then 𝐶←𝐶123 +𝑑n123 else 𝐶←𝐶123 −𝑑n123 end if return (𝐶, 𝑟) end function (see Equation 3 and Figure 3) is used to establish 𝐶𝑆𝐼 with the origin 𝑂 and orthogonal unit coordinate vectors 𝑈 , 𝑉 and 𝑊 . 𝑀 is then computed as a composition of two transformations – 𝑀𝑆2𝐼 from 𝐶𝑆𝑆 to 𝐶𝑆𝐼 , and 𝑀𝐼2𝑅 from the latter to 𝐶𝑆𝑅 . Note that the system of formulas (3) and the interpretation from Figure 3 must be employed separately for {𝐴𝑆, 𝐵𝑆,𝐶𝑆, 𝑂𝑆,𝑈𝑆,𝑉𝑆,𝑊𝑆} expressed in 𝐶𝑆𝑆, and {𝐴𝑅, 𝐵𝑅,𝐶𝑅, 𝑂𝑅,𝑈𝑅,𝑉𝑅,𝑊𝑅}expressed in 𝐶𝑆𝑅. Figure 3: Construction of the intermediate coordinate system from non-collinear 𝐴,𝐵and 𝐶(adapted from [6]). 𝑂=𝐴, 𝑈= 𝐵−𝑂 |𝐵−𝑂|, 𝑊= 𝑈× (𝐶−𝑂) |𝑈× (𝐶−𝑂)| , 𝑉=𝑊×𝑈 . (3) The transformation 𝑀𝐼2𝑅 from 𝐶𝑆𝐼 to 𝐶𝑆𝑅 is given in homogeneous coordinates as the composition of a 3D rotation 𝑅𝑜𝑡𝑅 and translation 𝑇𝑟𝑎𝑛𝑅(𝑂𝑅) , as shown in (4). The matrix 𝑀𝐼2𝑆 from 𝐶𝑆𝐼 to 𝐶𝑆𝑆 can be generated in the same manner, but the inverse 𝑀𝑆2𝐼 , as shown in (5), is actually needed. 𝑀 is then obtained as the composition 𝑀=𝑀𝐼2𝑅𝑀𝑆2𝐼. 𝑀𝐼2𝑅=𝑇𝑟𝑎𝑛𝑅(𝑂𝑅) · 𝑅𝑜𝑡𝑅=          𝑈𝑅.𝑥 𝑉𝑅.𝑥 𝑊𝑅.𝑥 𝑂𝑅.𝑥 𝑈𝑅.𝑦 𝑉𝑅.𝑦 𝑊𝑅.𝑦 𝑂𝑅.𝑦 𝑈𝑅.𝑧 𝑉𝑅.𝑧 𝑊𝑅.𝑧 𝑂𝑅.𝑧 0 0 0 1          (4) 86 Information Society 2025, 6–10 October 2025, Ljubljana, Slovenia Podgorelec et al. 𝑀𝑆2𝐼=𝑀−1 𝐼2𝑆=𝑅𝑜𝑡−1 𝑆·𝑇𝑟𝑎𝑛−1 𝑆(𝑂𝑆)=𝑅𝑜𝑡𝑇 𝑆·𝑇𝑟𝑎𝑛𝑆(−𝑂𝑆) =          𝑈𝑆.𝑥 𝑈𝑆.𝑦 𝑈𝑆.𝑧 0 𝑉𝑆.𝑥 𝑉𝑆.𝑦 𝑉𝑆.𝑧 0 𝑊𝑆.𝑥 𝑊𝑆.𝑦 𝑊𝑆.𝑧 0 0 0 0 1          ·          100−𝑂𝑆.𝑥 010−𝑂𝑆.𝑦 001−𝑂𝑆.𝑧 0 0 0 1          (5) 3 Results 3.1 Test Setup In tests conducted on the lawn of the primary school in Kamnica (Figure 1), we simulated conditions at an actual railway level crossing, which would be monitored by two LiDARs at the diagonal ends of the crossing. The optical properties, installation height, and viewing angle of LiDARs were also adapted to the expected conditions. We used two Ouster OS1 LiDARs with a vertical resolution of 128 channels, a horizontal resolution of 2048 samples, and a field of view of 42 . 4 ◦ vertically and 360 ◦ horizontally. They were 40 m apart, at a height of 3.2 m, and inclined to cover the ground from 3 m onwards. Approximately halfway between them, we placed six styrofoam sphere targets with radii of 25 cm and at different heights, making sure that no three sphere centres were collinear. We manually measured their circumferences in several directions and concluded that the deviations of radii were below 2 mm. Given the optical properties of LiDARs, the nearest neighbour points captured at a distance of 20 m are approximately 6 cm apart in the horizontal direction and 11 cm in the vertical direction. In our railway safety application, we use voxels with sides of 12 cm, which are sufficient for reliable detection of standardized minimum obstacles of 100 × 50 × 50 cm. The spheres were mounted on slender stands, which we "erased" using the procedure described in Subsection 2.2. 3.2 Accuracy Analysis Each LiDAR acquired between 10 and 50 points on each sphere. We considered 10 frames for each LiDAR, i.e., 10 pairs of point clouds. For each pair, we tested all possible quartets of points on each sphere and selected the best result. After the Landmark identification step, we first checked whether both calculated radii were in the range [24.8 cm, 25.2 cm], and then selected the pair with the smallest centre offset. The centre offset is the deviation between the calculated centres of the same sphere, one originally from 𝐶𝑆𝑅 and the other one transformed (registered) from 𝐶𝑆𝑆 . The Excel spreadsheets of results contain a total of almost 200 MB of data. Table 1 shows the results of the best alignment. Of the two calculated radii, we write down the worse one, i.e. the one that deviates more from the expected 25 cm. Both metrics gave encouraging results with the radius error below 0.3 mm and the centre offset below 3.5 cm, which is half better than LiDAR resolution applied 20 m from the lens, and below 30 % of the used voxel size. Note that the worst among best offsets in individual frames was 6.1 cm (probably due noise). Besides, there were always quartets of points found that gave completely unusable results, as the calculated radii ranged from 5 cm to 2 m. 4 Conclusion We presented a variation of the registration of two LiDAR point clouds based on sphere targets, which uses our own geometric construction method to determine the centres and radii of the spheres. Tests have shown that the method achieves encouraging Table 1: Sphere identifier, calculated radius [cm], and the distance between both calculated centres [cm] Sphere ID Calculated radius Centre offset 1 25.02 1.15 2 25.02 1.47 3 25.03 0.34 4 25.04 3.13 5 24.99 1.73 6 25.02 2.30 results for use in a railway safety application for level crossing monitoring. The registration error is below 30 % of the voxel side length used there. Unlike the traditional landmark identification approach based on solving the system of linear equations, the proposed approach offers good geometric interpretability and error explainability, which has potential for the development of heuristics that would prune the solution space and, eventually, enable us to use the time saved to conduct a more detailed investigation in the vicinity of the current optima. Current experience suggests that the best quartets are those where all normals are directed as closely as possible toward the LiDAR, while at the same time the points are not too close together and are as non-coplanar as possible (defining a tetrahedron with a larger volume). Acknowledgements The research was funded by the Slovene Research and Innovation Agency under Research Project J2-4458 and Research Programme P2-0041. The authors are grateful to Fokus Tech d.o.o. from Celje and OŠ Kamnica for providing equipment and testing facilities. References [1] Burcin Becerik-Gerber, Farrokh Jazizadeh, Geoffrey Kavulya, and Gulben Calis. 2011. Assessment of target types and layouts in 3d laser scanning for registration accuracy. Autom. Constr., 20, 5, 649–658. doi:10.1016/j.autcon.2 010.12.008. [2] Ben Bellekens, Vincent Spruyt, Rafael Berkvens, Rudi Penne, and Maarten Weyn. 2015. A benchmark survey of rigid 3d point cloud registration algorithms. Int. J. Adv. Intell. Syst, 8, 5, 118–127. http://72.52.166.99/articles/ints ys_v8_n12_2015_10.pdf. [3] Menthy Denayer, Joris De Winter, Evandro Bernardes, Bram Vanderborght, and Tom Verstraten. 2024. Comparison of point cloud registration techniques on scanned physical objects. Sensors, 24, 7, 2142. doi:10.3390/s24072142. [4] Mengjin Lyu, Jie Yang, Zhiquan Qi, Ruijie Xu, and Jiabin Liu. 2024. Rigid pairwise 3d point cloud registration: a survey. Pattern Recognit., 110408. doi:10.1016/j.patcog.2024.110408. [5] Jaroslav Marek and Pavel Chmelař. 2023. Survey of point cloud registration methods and new statistical approach. Mathematics, 11, 16, 3564. doi:10.339 0/math11163564. [6] David Podgorelec, Suzana Uran, Andrej Nerat, Božidar Bratina, Sašo Pečnik, Marjan Dimec, Franc Žaberl, Borut Žalik, and Riko Šafarič. 2023. Lidar-based maintenance of a safe distance between a human and a robot arm. Sensors, 23, 9, 4305. doi:10.3390/s23094305. [7] Trung-Thien Tran, Van-Toan Cao, and Denis Laurendeau. 2016. Esphere: extracting spheres from unorganized point clouds: how to extract multiple spheres accurately and simultaneously. Vis. Comput., 32, 1205–1222. doi:10 .1007/s00371-015-1157-0. [8] Tilen Urbančič, Žiga Roškar, Mojca Kosmatin Fras, and Dejan Grigillo. 2019. New target for accurate terrestrial laser scanning and unmanned aerial vehicle point cloud registration. Sensors, 19, 14, 3179. doi:10.3390/s19143179. [9] Dongfang Xie, Wei Zhu, Fengxiang Rong, Xu Xia, and Huiliang Shang. 2021. Registration of point clouds: a survey. In 2021 Int. Conf. on Networking Systems of AI (INSAI). IEEE, 136–142. doi:10.1109/INSAI54028.2021.00034. [10] Ningli Xu, Rongjun Qin, and Shuang Song. 2023. Point cloud registration for lidar and photogrammetric data: a critical synthesis and performance analysis on classic and deep learning algorithms. ISPRS J. Photogramm. Remote Sens., 8, 100032. doi:10.1016/j.ophoto.2023.100032. [11] Dongho Yun, Sunghan Kim, Heeyoung Heo, and Kwang Hee Ko. 2015. Automated registration of multi-view point clouds using sphere targets. Adv. Eng. Inform., 29, 4, 930–939. doi:0.1016/j.aei.2015.09.008. 87 88 Indeks avtorjev / Author index Alland Lucas ................................................................................................................................................................................ 27 Árgilán Viktor .............................................................................................................................................................................. 15 Baldouski Daniil ............................................................................................................................................................................. 7 Békési József ................................................................................................................................................................................ 15 Beleznai Csaba ............................................................................................................................................................................. 80 Berend Gábor ............................................................................................................................................................................... 44 Bóta András .................................................................................................................................................................................. 48 Brumen Matej ............................................................................................................................................................................... 68 Dabbous Ahmed ........................................................................................................................................................................... 48 Dávid Balázs ................................................................................................................................................................ 7, 60, 64, 75 Dobravec Tomaž .......................................................................................................................................................................... 19 Dömösi Pál ................................................................................................................................................................................... 30 Egri Péter ...................................................................................................................................................................................... 64 Galambos Gábor ........................................................................................................................................................................... 15 Hegyháti Máté ........................................................................................................................................................................ 11, 52 Horvat Štefan ............................................................................................................................................................................... 68 Horváth Géza ............................................................................................................................................................................... 30 Hren Boštjan ................................................................................................................................................................................. 19 Kawa Arkadiusz ........................................................................................................................................................................... 75 Kebelei Csaba ............................................................................................................................................................................... 11 Kirillova Nadezda ........................................................................................................................................................................ 80 Kiss Mihály .................................................................................................................................................................................. 44 Kovačević Nikola ......................................................................................................................................................................... 60 Krész Miklós ...................................................................................................................................................................... 7, 39, 64 Kusper Gábor ............................................................................................................................................................................... 23 Lukač Luka ................................................................................................................................................................................... 84 Mansour Ahmed ........................................................................................................................................................................... 80 Mongus Domen ............................................................................................................................................................................ 68 Nagy Benedek .............................................................................................................................................................................. 35 Oberweger Fabio F. ...................................................................................................................................................................... 80 Papp Imre ..................................................................................................................................................................................... 15 Pečnik Sašo .................................................................................................................................................................................. 84 Podgorelec David ......................................................................................................................................................................... 84 Possegger Horst ............................................................................................................................................................................ 80 Quilliot Alain ............................................................................................................................................................................... 56 Repnik Blaž .................................................................................................................................................................................. 84 Sali Attila ..................................................................................................................................................................................... 27 Strnad Damjan .............................................................................................................................................................................. 68 Szabó Sándor ................................................................................................................................................................................ 72 Szaller Ádám ................................................................................................................................................................................ 64 Tahalea Sylvert Prian ............................................................................................................................................................. 39, 75 Tavzes Črtomir ............................................................................................................................................................................. 60 Toussaint Hélène .......................................................................................................................................................................... 56 Váncza József ............................................................................................................................................................................... 64 Widhalm Verena .......................................................................................................................................................................... 80 Wu Nicole .................................................................................................................................................................................... 27 Žalik Borut ................................................................................................................................................................................... 84 Zaválnij Bogdán ........................................................................................................................................................................... 72 89 Srednjeevropska konferenca o uporabnem teoretičnem računalništvu in informatiki (MATCOS) Middle-European Conference on Applied Theoretical Computer Science (MATCOS) Uredniki l Editors: Andrej Brodnik Gábor Galambos Rok Požar