scieee AI-readable full text Open interactive document viewer

Agoran: An Agentic Open Marketplace for 6G RAN Automation

Chatzistefanidis, Ilias; Nikaein, Navid; Leone, Andrea; Maatouk, Ali; Tassiulas, Leandros; Morabito, Roberto; Pitsiorlas, Ioannis; Kountouris, Marios

Full text

Graphical Abstract Agoran: An Agentic Open Marketplace for 6G RAN Automation Ilias Chatzistefanidis, Navid Nikaein, Andrea Leone, Ali Maatouk, Leandros Tassiulas, Roberto Morabito, Ioannis Pitsiorlas, Marios Kountouris arXiv:2508.09159v2 [cs.NI] 21 Aug 2025 Highlights Agoran: An Agentic Open Marketplace for 6G RAN Automation Ilias Chatzistefanidis, Navid Nikaein, Andrea Leone, Ali Maatouk, Leandros Tassiulas, Roberto Morabito, Ioannis Pitsiorlas, Marios Kountouris •Present Agoran, the first tripartite, Legislative, Executive, Judicialagentic GenAI marketplace that enables 6G stakeholders to express network intents in natural language and receive autonomous, regulationcompliant resource allocations. •Introduce three novel AI services: (i) a regulation-aware RAG for onthe-fly compliance checks, (ii) a watcher-driven vector store that converts live telemetry into retrievable context, and (iii) a rule-based Trust Score that filters hallucinations and malicious tactics in real time. •Demonstrate that a LoRA-tuned 8 B-parameter LLaMA reaches ≈86 % of GPT-4.1’s decision quality, while a fully fine-tuned 1 B LLaMA still recovers ≈78 % of GPT-4.1 at only 6 GiB VRAM and 1.3secs convergence time. •Deploy the full framework on an OpenAirInterface and FlexRIC 5G testbed with realistic MCS traces; dynamic agentic negotiation increases aggregate throughput by 37%, reduces URLLC latency by 73%, and saves 8.3% of PRBs compared to a static baseline. A live demo is presented https ://www.youtube.com/watch?v=h7vEyMu2f5w&abchannel = BubbleRAN. •Demonstrate one-round consensus and cross-slice intent swapping, validating Agoran’s compatibility with today’s Open RAN and future AI-RAN roadmaps. •Release code, datasets, and fine-tuning notebooks to catalyze research on stakeholder-centric, ultra-flexible 6G networks. Agoran: An Agentic Open Marketplace for 6G RAN Automation Ilias Chatzistefanidisa, Navid Nikaeina,b, Andrea Leoneb, Ali Maatoukc, Leandros Tassiulasc, Roberto Morabitoa, Ioannis Pitsiorlasa, Marios Kountourisa,d aEURECOM, Sophia-Antipolis, France bBubbleRAN, Sophia-Antipolis, France cYale University, New Haven, USA dUniversity of Granada, Spain Abstract Next-generation mobile networks must reconcile the often-conflicting goals of multiple service owners. However, today’s network slice controllers remain rigid, policy-bound, and largely unaware of the business context. We introduce Agoran Service and Resource Broker (SRB), an agentic marketplace that brings stakeholders directly into the operational loop. Inspired by the ancient Greek agor´a,Agoran distributes authority across three autonomous artifical intelligence (AI) branches: a Legislative branch that answers compliance queries using retrieval-augmented Large Language Models (LLMs); an Executive branch that maintains real-time situational awareness through a watcher-updated vector database; and a Judicial branch that evaluates each agent message with a rule-based Trust Score, while arbitrating LLMs detect malicious behavior and apply real-time incentives to restore trust. Stakeholder-side Negotiation Agents and the SRB-side Mediator Agent negotiate feasible, Pareto-optimal offers produced by a multi-objective optimizer, reaching a consensus intent in a single round, which is then deployed to Open and AI-driven RAN controllers. Email addresses: [email protected] (Ilias Chatzistefanidis), [email protected], [email protected] (Navid Nikaein), [email protected] (Andrea Leone), [email protected] (Ali Maatouk), [email protected] (Leandros Tassiulas), [email protected] (Roberto Morabito), [email protected] (Ioannis Pitsiorlas), [email protected] (Marios Kountouris) Preprint submitted to Computer Networks August 22, 2025 Deployed on a private 5G testbed (OpenAirInterface and FlexRIC) and evaluated with realistic modulation and coding scheme (MCS) traces of vehicle mobility, Agoran achieved significant gains: (i) a 37% increase in negotiated throughput of enhanced mobile broadband (eMBB) slices, (ii) a 73% reduction in negotiated latency of ultra-reliable low latency communications (URLLC) slices, and concurrently (iii) an end-to-end 8.3% saving in physical physical resource blocks (PRB) usage compared to a static baseline. An 1B-parameter Llama model, fine-tuned for just five minutes on 100 GPT4 dialogues, recovers approximately 80% of GPT-4.1’s decision quality, while operating within 6 GiB of memory and converging in only 1.3 seconds. These results establish Agoran as a concrete, standards-aligned path toward ultraflexible, serviceand stakeholder-centric 6G networks and open new research avenues in agentic observability, lightweight agent distillation for network functions such as multi-agent service-level-agreement (SLA) negotiation, and cross-domain intent reconciliation. A detailed live demo is presented https : //www.youtube.com/watch?v=h7vEyMu2f5w&abchannel =BubbleRAN. Keywords: Agentic AI, Multi-Agent Negotiation, Intent-Based Networking, Open RAN/AI-RAN, Large Language Models,, Trustworthy AI, Stakeholder-Centric 6G, Next-G 1. Introduction Upcoming sixth-generation (6G) systems are expected to operate as multiservice, multi-stakeholder platforms, where mobile network operators (MNOs), virtual operators, service providers, and vertical industries share a common, slice-capable infrastructure [1, 2, 3, 4]. The transition into the upper midband (7–24 GHz, FR3), designated by 3GPP and the World Radiocommunication Conference 2023 (WRC-23) for International Mobile Telecommunications (IMT) services, coincides with a surge in latency-critical workloads such as cloud gaming and industrial Extended Reality (XR). Meanwhile, the global mobile user base reached 5.8 billion unique subscribers in 2024, with 4G/5G networks alone supporting over 7 billion connections by early 2025 [5, 6, 7, 8, 9]. These trends are already overwhelming traditional rulebased control and management systems. Yet, most live networks still operate with static configurations and pre-negotiated service-level agreements (SLAs). This results in an service-level-agreement (SLA) gap, the mismatch 2 between contractual performance targets and real-time network conditions, which causes inefficient resource utilization, degraded Quality of Experience (QoE), and costly manual interventions [10, 11, 12, 13, 14]. Intent-based networking (IBN) [15, 16] and early multi-service orchestration frameworks [17, 18, 19, 20, 21, 22] have begun to address these limitations by enabling operators to specify desired outcomes rather than low-level commands. These same principles now underpin industry-wide standardization efforts: the O-RAN Alliance promotes openness in the RAN by breaking vendor silos and mandating intent-centric, AI-ready interfaces [23], while the recently founded AI-RAN Alliance advocates for an AI-native radio stack [24]. However, in nearly all existing solutions, business entities remain outside the operational loop, and AI is treated as an offline optimizer [25] rather than a real-time decision-maker. What remains missing is a continuous, trustworthy dialogue—one in which every stakeholder can express objectives in natural language, negotiate trade-offs on the fly, and rely on the network to enforce the resulting consensus across the radio-to-cloud stack. Recent progress in agentic AI, powered by Large Language Models (LLMs) and their domain-tuned derivatives—hereafter referred to as Large Telecom Models (LTMs) [26, 27, 28]—offers a promising foundation for such a dialogue. LTMs combine language understanding, external tool invocation, and chain-of-thought reasoning [29, 30], enabling software agents to translate human intents into verifiable actions. However, without appropriate governance mechanisms, the deployment of autonomous agents risks biased decisions, unfair resource allocations, or strategic manipulation [31, 32, 33]. To reconcile autonomy with fairness, we propose Agoran1, an open marketplace for intent reconciliation and resource brokerage.Agoran embeds agentic AI into three mutually independent branches, Legislative,Judicial, and Executive, inspired by the classical separation-of-powers doctrine. Legislative agents curate and evolve the corpus of spectrum regulations, security policies, and contractual clauses. Judicial agents resolve conflicts and enforce compliance through incentives and penalties. Executive agents integrate real-time telemetry with ratified consensus intents to issue slice and resource directives, thereby closing the control loop across heterogeneous infrastructure. This tripartite architecture prevents unilateral dominance in 1In ancient Greece, the agor´a (ἀγορά) was the civic and commercial hub where citizens gathered to trade, debate, and deliberate; Agoran plays a similar role in future networks. 3 negotiations and enables pay-as-you-grow scalability, autonomous fault resilience, and fine-grained service differentiation. This paper makes four main contributions: •It introduces Agoran, a novel Service & Resource Broker architecture that embeds agentic AI into legislative, executive, and judicial roles to automate decision-making in multi-service, multi-stakeholder 6G networks. •It proposes a negotiation engine where LTM agents collaborate with evolutionary optimization to achieve near–Pareto-optimal consensus intents under stringent real-time constraints. •It presents a full prototype implementation of Agoran on an OpenAirInterface [34] and FlexRIC [35] 5G testbed, demonstrating live consensus formation, robustness against malicious bidding, and sustained Quality of Service (QoS) under bursty traffic. The evaluation spans both large models and fine-tuned small language models (SLMs), quantifying the trade-off between accuracy and system overhead. •It releases code, demos, datasets, and LTM checkpoints to promote transparency and reproducibility in trustworthy agentic automation research. The remainder of the paper is organized as follows. Section 2 surveys AI-enabled management and marketplace concepts for 5G/6G systems. Section 3 motivates the marketplace approach and details the tripartite governance model. Section 4 describes the end-to-end workflow from intent capture to closed-loop enforcement, and Section 5 delves into the internal design of negotiation, executive, judicial, and executive agents. Section 6 and Section 7 evaluate Agoran on real-world scenarios. Section 8 discusses limitations and open research questions, and Section 9 concludes the paper. 2. Related Work Knowledge-Engine LTMs. A collective roadmap from industry and academia outlines how LTMs can support a wide range of use cases across the network lifecycle [26]. Early work treats LLMs as domain-grounded knowledge engines. Bariah et al. pre-train multimodal models on 3GPP, RF, and traffic 4 Table 1: Comparison with representative work on LLMs and multi-agent GenAI in the telecom domain. A ✓indicates that a feature is explicitly addressed. Work Multi-Service Real-Time Multi-Agent Tool Use / APIs Governance & Arch. Evaluation Edge Bariah et al. [36] ✗ ✗ ✗ ✗ ✗ ✗ ✗ Lin et al. [37] ✗ ✗ ✗ ✗ ✗ ✓ ✓ Zou et al. [38] ✗ ✓ ✓ ✗ ✗ ✗ ✓ He et al. [39] ✗ ✗ ✓ ✗ ✗ ✓ ✗ Patil et al. (Gorilla) [40] ✗ ✗ ✗ ✓ ✗ ✓ ✗ Qin et al. (TOOLLLM) [41] ✗ ✗ ✗ ✓ ✗ ✓ ✗ Martini et al. [18] ✓ ✗ ✗ ✗ ✓ ✓ ✗ NASP [42] ✓ ✗ ✗ ✗ ✗ ✓ ✓ Wu et al. (LLM-xApp) [43] ✓ ✓ ✗ ✓ ✗ ✓ ✓ Lotfi et al. [44] ✓ ✓ ✓ ✗ ✗ ✓ ✗ Elkael et al. (ALLSTaR) [45] ✗ ✓ ✗ ✓ ✓ ✓ ✓ Chatzistefanidis et al. (Maestro) [17] ✓ ✓ ✓ ✗ ✗ ✓ ✗ Agoran (this work) ✓ ✓ ✓ ✓ ✓ ✓ ✓ corpora, arguing that such LTMs could underpin artificial general intelligence (AGI)-grade cognition for networks [36]. Maatouk et al. refine this idea in the TeleLLMs series, enhancing accuracy on standardization documents through telecom-aware vocabulary and positional embeddings [46]. These studies establish domain grounding, but the model remains outside the control loop. Multi-Agent Reasoning. A second line of research views LLMs as autonomous or collaborative agents. Zou et al. embed on-device LLMs in a game-theoretic multi-agent scheduler for spectrum and power allocation under tight latency constraints [38]. He et al. combine generative AI with cooperative game theory for secure UAV routing [39], while Du et al. demonstrate that a “society of minds” outperforms single models on compositional reasoning [47]. Outside the telecom domain, collective LLMs have been shown to exhibit persona-driven biases and strategic manipulation, unless appropriately moderated through incentives or governance mechanisms [48, 49]. LLM-Powered xApps and Closed-Loop RIC Control. Wu et al. integrate a GPT-prompting LLM-xApp into the O-RAN near-real-time radio intelligent controller (RIC), retuning slice resources and achieving a 28% increase in downlink throughput over a MARL baseline [43]. Lotfi et al. reduce MARL convergence time by 40% using prompt-tuned LLM embeddings to steer distributed RL agents for O-RAN slicing [44]. Tool Grounding and Edge Deployment. Gorilla [40] and TOOLLLM [41] fine-tune language models on API triples, while Toolformer demonstrates that LLMs can self-learn API usage through unsupervised prompting [29]. Lin et al. compress LLMs for 6G edge devices, achieving compliance with stringent latency constraints [37]. Recent work addresses specific layers of the stack: an LLM-centric intent life-cycle manager [50], a reinforcement 5 learning (RL) explainer for slicing transparency [51], and an LLM agent that interacts with OpenAI Cellular for slice optimization [52]. While each of these contributions advances its respective area, none integrates negotiation, enforcement, and governance within a unified framework. Business-Plane Brokerage and Large-Scale Evaluation. The Network Sliceas-a-Service Platform (NASP) implements hierarchical orchestration and businessplane onboarding for multiple verticals across both 3GPP and non-3GPP domains [42]. ALLSTaR automatically generates 18 schedulers, compiles them into RIC-compliant code, and A/B tests them over-the-air while enforcing IEEE 7001 transparency requirements [45, 53]. Martini et al. contribute a slice-assurance loop that maps vertical intents into verifiable key performance indicator (KPI) thresholds [18]. Closest Antecedent. Our earlier Maestro prototype [17] deploys personarich LLM agents on a 5G testbed, where multiple stakeholders negotiate spectrum shares and adversarial tactics are surfaced. While it validates the business-plane concept, it optimizes a single KPI, lacks tool-enabled enforcement, and does not provide large-scale evaluation or a formal governance model. Table 1 summarizes the literature across seven dimensions critical to autonomous 6G operation. No prior work combines multi-service intent negotiation among independent stakeholders, real-time multi-agent reasoning, tool-enabled enforcement, and a tripartite governance architecture, evaluated end-to-end on an over-the-air 5G-slice testbed. Agoran closes this gap by transforming LTMs into legislative, judicial, and executive actors that reconcile conflicting intents in real time, withstand malicious bidding, and operate within a trustworthy governance framework. Existing studies provide essential building blocks—domain-specific LLMs, agent collectives, tool grounding, and edge optimization, but leave open the question of how to orchestrate these components into a neutral, self-regulating marketplace that spans business intents, network policies, and real-time enforcement. Agoran delivers the first end-to-end solution for autonomous, fair, and efficient resource brokerage in 6G networks. 3. AGORAN Overview This section delves into the overall design and principles that constiture Agoran Service & Resource Broker (SRB) the key enabler of fully autonomous, multi-stakeholder next-generation networks. As shown in Fig. 6 Figure 1: Agoran Service & Resource Broker (SRB) for Collaborative MultiStakeholder, Multi-Service Network Automation. Power is delegated into three autonomous branches, mirroring societal structures. The Legislative branch codifies policy, the Judicial branch arbitrates negotiations between Network Slice/Service Customers (NSCs/CSCs), and the Executive branch enforces the consensus decisions across a multislice-capable network. 1, the SRB spawns a digital agora for Communication Service and Network Slice Customers (CSCs/NSCs) as defined in 3GPP TS 28.530 and TS 28.531 [54, 55] including MNOs, MVNOs, service providers, and vertical industries. Each CSC/NSC owns a dedicated software broker agent (bAgent) that acts on their behalf negotiating, deliberating, and cooperating throughout the life cycle of network planning, deployment, and real-time optimization. Intent-Based, Multi-Domain Fabric. Figure 1 depicts the horizontal segmentation into three logical domains. Within the Customer Domain,broker agents capture high-level customer intent, KPIs, SLA targets, spectrum budgets, cost models, and use-case priorities. The SRB ingests these intents, detects conflicts or synergies, and, through three independent power branches and multi-stakeholder collaboration, derives a Consensus Intent. Once ratified, the intent is decomposed into slice-, resource-, and control-domain directives that are enforced over heterogeneous, multi-vendor infrastructure comprising RUs, DUs/CUs, and the core. Separation of Powers. Inspired by societal checks and balances, the mar7 into a choice set of non-dominated offers presented to the LLM-based negotiators (Figure 5). We formalize the resources, constraints, and optimization objectives below, followed by a detailed description of the evolutionary search procedure. Resource vector. We consider three canonical 5G/6G service slices, enhanced Mobile Broadband (eMBB), Ultra-Reliable Low-Latency Communication (URLLC), and massive Machine-Type Communication (mMTC). Each slice i∈ {e,u,m}is assigned a resource quadruple xi= (bi, ci, pi, si), where bidenotes downlink bandwidth (MHz), cirepresents abstract compute cycles, piis transmission power (Watt), and siis ancillary storage (megabyte). The global decision vector is then given by x=xe,xu,xm∈R12. Constraints. Resources are subject to system-wide limits: Pibi≤Bmax,Pici≤Cmax,;Pipi≤Pmax,;Pisi≤Smax,(1) and to slice-specific SLA clauses: eMBB must sustain throughput Te≥ Tmin e; URLLC latency must satisfy Lu≤Lmax u; mMTC cost is capped by Cm≤Cmax m; and so on. KPI models. Throughput. Following the spectral efficiency tables in 3GPP TS 38.306 [57], slice iachieves a throughput of Ti(bi, mi) = κ QmiRmibi,(2) where Qmiand Rmidenote the modulation order and coding rate associated with MCS index mi(e.g., Q28 = 6, R28 = 0.948 for 256-QAM with a coding rate of 0.948). The factor κ≈0.86 accounts for OFDM overhead and the DL/UL duty cycle, as specified in [58]. Latency. Total latency comprises a fixed scheduler and transport component Lfix i, added to the delay of an M/M/1 queue [59]: Li(bi, mi) = Lfix i+1 µi(bi, mi)1−ρi,(3) where the service rate is defined as µi(bi, mi) = Ti(bi,mi)×106 Spkt ,with packet size Spkt = 1500 ×8 bits, and ρi∈(0,1) denotes the traffic load ratio for slice i. Cost. Monetary cost is modeled as a linear function of compute and storage resources: Ci(ci, si) = αci+si,(4) where αis a fixed unit cost coefficient. 14 Energy. Energy consumption is approximated by the committed transmission power: Ei=pi.(5) Objective vector. The optimization aims to minimize the objective vector: f(x) = − XiTi,XiLi,XiCi,XiEi,(6) that is, we maximize aggregate throughput while minimizing total latency, cost, and energy consumption. Evolutionary search (NSGA-II). Each individual encodes a 12-dimensional resource allocation vector. Search is conducted using classic NSGA-II operators [60]: uniform crossover (pc= 0.9) and per-gene Gaussian mutation (pm= 0.1). A repair function rescales any gene that violates global budget constraints to restore feasibility. Individuals violating slice-specific SLA constraints are discarded during fitness evaluation. Diversity is preserved through fast non-dominated sorting and crowding distance. The algorithm runs for 80 generations on a population of 60, producing the non-dominated front P∗, which typically contains 20–40 Pareto-optimal solutions. Offer generation. The Pareto front is sorted by crowding distance, and the top-kentries (with k= 3 by default in our setup) are encoded into a JSON template that lists the derived KPIs for each slice. Since every candidate in Pis already efficient,feasible, and SLA-compliant, the subsequent LLM-mediated negotiation is guaranteed to operate on sound configurations. All hyperparameters and numerical constants follow the canonical NSGA-II specification in [60], as well as 3GPP standards for spectral efficiency and latency budgeting [57, 58, 59]. 5.2. Legislative Branch (lAgent): Retrieval-Augmented Compliance Engine To determine whether a proposal complies with spectrum regulations, national security policies, or contractual obligations, the Legislative branch operates the specialized lAgent as a retrieval-augmented compliance engine (Figure 6). This engine combines a semantic search layer with a compact LLM (ranging from 2 to 7 billion (B) parameters). When an agent, or the Judicial branch, issues a regulatory query, such as “Is offer #17 legal under current regulations?” or any other regulatory query to enhance the marketplace, the engine executes the following sequence in a few seconds: 15 Figure 6: Operation of the lAgent as compliance engine in the Legislative Branch. 1) A regulation query is issued. 2) A retriever fetches the most relevant clauses from the evolving knowledge base. 3) The query is augmented with the retrieved snippets. 4) The LLM returns a grounded, citation-ready answer. 5) New, validated precedents are written back to the database. •the query is dispatched to the retriever, which ranks passages from a dynamic corpus integrating 3GPP and ETSI standards, national directives, and prior marketplace rulings; •the top-kretrieved passages are concatenated with the query to form an augmented prompt; •the LLM synthesizes a concise, citation-ready response that references each supporting clause inline; and •if the result establishes a novel precedent, the validated snippet is appended to the corpus, ensuring that future deliberations reflect the updated regulatory context. Retrieval-augmented generation (RAG) has already demonstrated its value in offline legal reasoning, with public benchmarks such as LegalBench-RAG [61], LexRAG [62], and KRAG [63] reporting significant accuracy gains over vanilla LLMs while producing verifiable citations. Recent toolkits even support ondemand drafting of statutory language (e.g., LexDrafter [64]). What remains unproven, however, is whether the same principle can be compressed into small (7–13B) models collocated with the non-RT RIC. By adapting edgeoriented RAG techniques, we demonstrate that an LTM, running on a single GPU, can generate grounded, citation-ready answers within a few seconds, 16 well within the compute and latency constraints of RIC edge controllers. This design keeps the Legislative branch transparent, up-to-date, and fully self-contained within the Agoran loop. 5.3. Executive Branch (eAgent): Agentic Observability The Executive branch provides the marketplace with a live, self-evolving view of network reality (Figure 7). At its core lies the eAgent which answers queries according to a vector-based knowledge store that is updated via two complementary paths. Push path – resource watchers. Lightweight watchers attach to Kubernetes resources, custom resource definitions (CRDs), O-RAN xApp state, spectrum allocators, and Prometheus exporters. Whenever a resource changes, the watcher serializes the delta, embeds it, and pushes the resulting vector into the vector database (DB) within a few milliseconds. This event-driven pipeline keeps the store in lockstep with configuration drift and KPI excursions, without flooding the cluster with telemetry traffic. To our knowledge, no prior work in the RAG-for-networks literature has proposed such a mechanism. Pull path – agent queries. When an agent issues a query, e.g., “What is the current headroom on DU-7?”, an LLM determines the most efficient action to minimize latency and cost. It may first retrieve semantically indexed snapshots from the vector DB, or, if finer-grained data is needed, invoke a monitoring API (e.g., Prometheus, eBPF) to fetch fresh counters. These values are then injected into the model’s context, and the LLM may iterate until it emits a stop token. Feedback loop. Every datum that passes the LLM’s internal consistency checks is appended to the vector store with a timestamp. Together with the watcher-based push path, this push-plus-pull design ensures that optimization and negotiation operate on up-to-the-second evidence. The result is a context-aware, bandwidth-efficient observability mechanism that scales with the demands of the agentic marketplace. 5.4. Judicial Branch (jAgent): Arbitration and Incentive Engine The Executive and Legislative branches ensure that negotiated intents are feasible and legal; the Judicial branch ensures they are also trustworthy, as shown in Figure 8. An arbitration-specific LLM for content moderation embeds each negotiation message and draft consensus intent, and compares it against a continuously updated library of toxic discourse, collusion patterns, 17 Figure 7: Agentic observability loop in the Executive Branch (eAgent). 1) A query arrives at the branch. 2) The LLM decides whether to retrieve from the vector store (DB), monitor live counters, or stop the iteration (optional dashed path). 3) Retrieved facts are injected into the prompt; the LLM may iterate. 4) Verified evidence is written back to the store. A separate watcher layer continuously streams configuration deltas and KPI shifts to the store (periodic update). and over-provisioning tactics. The resulting toxicity score is processed by a lightweight incentive engine: benign messages pass unaltered; borderline cases receive a soft warning; and malicious or hallucinatory content triggers an automatic fine that temporarily reduces the offending agent’s influence. Conversely, consistently constructive contributions earn credits that enhance future bargaining power. The need for this branch is not merely theoretical. Our earlier Maestro prototype [17], along with independent studies on LLM societies [48, 49], demonstrates that persona-driven language agents can become manipulative, toxic, or hallucinate facts when left unchecked. By grounding its verdicts in the same regulatory corpus maintained by the Legislative branch (§5.2), the Judicial branch aligns its sanctions with formal policy while preserving an immutable audit trail. In practice, this closed loop suppresses hallucinations, blocks resource-hoarding strategies, and maintains a fair bargaining environment, without violating the real-time constraints of the negotiation cycle. 5.5. bAgent Output: Trustworthiness-Score Framework In multi-agent negotiation scenarios, evaluating the trustworthiness of LLM agents’ outputs is essential for ensuring reliable and effective decision18 Figure 8: jAgent in the Judicial Branch. Negotiation messages and draft consensus intents are evaluated for toxicity and manipulative behavior. The resulting trust verdict triggers a proportional incentive—warn,fine, or credit—which is broadcast to all marketplace participants. making. Given the high computational cost associated with retraining and refining LLMs, it is critical not only to develop robust and confident models [65], but also to ensure that their reliability is maintained over time [66]. To address this need, we propose a comprehensive Trust Score framework that quantifies agent reliability along two key dimensions: (i) the alignment of the LLM’s decision-making with the mediator’s expectations, and (ii) the behavioral coherence of the agent’s communications. Our framework is deliberately rule-based, avoiding complex machine learning techniques or additional LLM-based evaluators. This design choice emphasizes explainability and transparency, ensuring that every component of the trust assessment process is interpretable, auditable, and verifiable by human experts. The overall trust score Tis defined as a weighted sum of two components, satisfaction and coherence, as follows: T=ws·S+wc·C(7) where Sdenotes the satisfaction score, which measures the alignment between the agents’ decisions and those of the central mediator, while Crepresents the coherence score, which evaluates the quality and consistency of agent communications. The weights wsand wcare user-configurable parameters that satisfy ws+wc= 1. In our implementation, we set ws= 0.15 and wc= 0.85, thus placing greater emphasis on communication quality as a determinant of trust. 19 Figure 9: Trust-Score framework applied to every agent after the oAgent emits its chosen SLA index and rationale. Smeasures decision alignment; Cmeasures behavioral coherence; the final score is T= 0.15 S+ 0.85 C. 5.5.1. Trustworthiness of the LLM In negotiation contexts, the trustworthiness of an LLM agent is evaluated based on the degree to which its decisions align with three critical reference points: (i) the validity of the proposed solutions; (ii) the optimization of the intended objectives; and (iii) the consistency and/or the agreement with mediator recommendations. We quantify this alignment using a satisfaction score S, which captures the extent of deviation across these three dimensions. 1. Deviation from Valid Offers: This component tests whether the agent’s proposal lies inside the set of admissible, feasible solutions. Let P= {p1, p2, . . . , pn}denote the set of valid proposals in the negotiation space. We define the deviation metric Do=(0 if pagent ∈ P 1 if pagent /∈ P (8) where pagent is the offer generated by the agent. Hence Do= 0 signifies full compliance with feasibility constraints, while Do= 1 assigns the maximum penalty for submitting an invalid proposal, accurately reflecting a failure to respect the negotiation limits. 2. Deviation from Intent: The second component evaluates how well the agent’s selected proposal aligns with its stated objectives, specifically in 20 terms of solution optimality. To quantify this, we employ the Normalized Generational Distance (NGD) metric, which measures the proximity of the agent’s achieved outcome to a reference set of Pareto-optimal solutions, denoted by S∗={s∗ 1, s∗ 2, . . . , s∗ k}. Let vagent = [v1, v2, . . . , vm] represent the agent’s KPI vector corresponding to its chosen proposal. The deviation from intent is defined as: Di= min (1,NGD(vagent,S∗)) (9) where the NGD metric is computed as: NGD(v,S∗) = 1 |S∗|sX s∗∈S∗ min s∈S d(vnorm,s∗ norm)2.(10) Here, vnorm and s∗ norm denote the normalized versions of the KPI vectors to ensure comparability across different scales. The use of the minimum function in Equation 9 bounds the deviation score within [0,1], preserving interpretability while penalizing significant divergence from optimal intent. 3. Deviation from Mediator: The third component assesses the degree to which the agent’s proposal aligns with the mediator’s recommendation, serving as an indicator of the agent’s willingness to cooperate so as to reach a consensus. This deviation is defined as: Dm=(0 if pagent =pmediator 1 if pagent =pmediator (11) where pmediator denotes the proposal recommended by the mediator. A value of Dm= 0 reflects full agreement with the mediator, while Dm= 1 indicates complete disagreement, implying resistance to convergence within the negotiation process. The overall satisfaction score Sintegrates the three deviation components (offer validity, intent alignment, and mediator agreement) using a weighted linear combination: S= 1 −(wo·Do+wi·Di+wm·Dm) (12) where wo,wi, and wmare the respective weights assigned to deviations from valid offers, intent, and the mediator’s recommendation, subject to the constraint wo+wi+wm= 1. In our implementation, we assign equal importance 21 to each component, setting wo=wi=wm=1 3, thus ensuring a balanced evaluation across all three dimensions of decision alignment. 5.5.2. Behavior of the Agents In addition to decision alignment, the behavioral coherence of agent communications offers critical insights into the quality and reliability of their reasoning processes. To assess this aspect, we evaluate coherence across three complementary dimensions, each capturing a distinct facet of the agent’s ability to articulate, justify, and consistently support its decisions. Factual Accuracy. Factual accuracy assesses the correctness of numerical claims made by agents in their decision rationales. We apply natural language processing techniques to extract quantitative statements related to KPIs and validate them against ground truth data. Let C={c1, c2, . . . , ck}denote the set of numerical claims extracted from an agent’s rationale, where each claim ci= (mi, vi) comprises a metric type miand an asserted value vi. For each claim, we compute the relative error with respect to the true value v∗ ias follows: ϵi=|vi−v∗ i| v∗ i (13) We categorize hallucinations based on the magnitude of relative error ϵi as follows: •None:ϵi≤0.15 (within 15% tolerance) •Minor: 0.15 < ϵi≤0.5 (penalty: 0.1) •Major: 0.5< ϵi≤1.0 (penalty: 0.3) •Severe:ϵi>1.0 (penalty: 0.6) The factual accuracy score Fquantifies an agent’s overall numerical reliability by combining relative accuracy with hallucination penalties and is defined as: F= max  0,1 |C| |C| X i=1 (1 −ϵi)− |C| X i=1 penaltyi (14) where penaltyiis the hallucination penalty for claim ciassigned based on the error category. The score is normalized to ensure a maximum of 1.0 and lower-bounded at zero to prevent negative values. 22 Logical Consistency. Logical consistency assesses the internal coherence and reasoning quality of an agent’s communication. We evaluate whether the rationale demonstrates structured thinking, aligns with the agent’s objectives, and avoids self-contradictions. The logical consistency score Lis computed based on the following criteria: •Logical connectors: Detection of reasoning indicators (e.g., “because”, “therefore”, “since”), which signal structured argumentation. •Goal alignment: Presence of references to negotiation objectives or strategic intent. •Contradiction detection: Identification of conflicting claims, such as simultaneous positive and negative assertions about the same aspect. Formally, the score is defined as: L= min (1,max (0,1−Pc+Bs+Bg)) (15) where Pcis the contradiction penalty, reducing the score for detected inconsistencies, Bsis the structure bonus, awarded for the presence of logical connectors and coherent reasoning, and Bgis the goal bonus, assigned when the rationale explicitly references objectives or strategic considerations. This formulation ensures that logical consistency is rewarded for clear, structured, and goal-driven reasoning while penalizing internal contradictions. Semantic Coherence. Semantic coherence measures the relevance, expressiveness, and structural quality of an agent’s communication within the negotiation domain. This metric evaluates whether agents demonstrate a solid understanding of domain-specific content and express/communicate their reasoning with clarity and variety. The semantic coherence score Eintegrates the following components: •Domain terminology (dt): Use of appropriate technical terms and negotiationspecific vocabulary. •Linguistic diversity (ld): Richness of vocabulary and avoidance of repetitive phrasing. •Structural quality (sq): Sentence variety, appropriate length distribution, and overall readability. 23 (a) Negotiation trajectory over ten iterations. Lower MAE = closer to consensus. (b) Effect of the warning incentive on deviation (box plots) and convergence time (line plots). Green shading and right-sided bars indicate runs where the arbitral LLM issued a warning incentive; red denotes the baseline without arbitration (left-sided bars). Figure 10: Judicial-branch evaluation. The effect of different agent personalities leads to significant variation in negotiation trajectories. Toxic or manipulative LLMs are more resistant to reaching consensus; however, the warning incentive mechanism effectively mitigates this resistance, making them more cooperative. Table 9: Toxicity Classification T-N, T-D Predicted Non-Toxic Toxic Actual Non-Toxic 179, 168 1, 12 Toxic 21, 40 249, 230 (Prec) (Rec) (F1) (1.0, .95) (.92, .85) (.96, .90) size to transform that context into an accurate answer. Although our agentic observability loop is an early prototype and still exhibits retrieval-induced errors, it is, to the best of our knowledge, the first real-system demonstration of a fully agentic approach to network observability. We expect that targeted improvements in RAG mechanics, along with incremental model scaling, will significantly improve accuracy while preserving an edge-friendly resource profile. 6.5. Exp. 4 — Judicial Malice Mitigation (jAgent) We now evaluate the jAgent of the Judicial branch, i.e., the combination of an arbitral LLM and an incentive engine, whose role is to detect malicious behavior and steer negotiations back toward consensus using a warn incentive. All tests involve five agents, a budget of ten iterations (rounds), and personalities drawn from the Big Five spectrum introduced in our earlier prototype, Maestro [17]: Vulnerable (V), Agreeable (A), Neutral (N), Disagreeable (D), Toxic with arbitration (T), and Toxic without arbitration (T∗). 30 Model Satisf. Coher. Trust Score (0-5)↑ gpt-4.1 3.88 5.00 4.83 gpt-4.1-mini 3.88 4.96 4.81 Llama-3.1-8B-instruct-FT 4.44 3.86 3.94 Llama-3.1-8B-instruct 5.00 3.73 3.91 Llama-3.2-3B-instruct-FT 3.89 2.23 2.48 Llama-3.2-3B-instruct 4.43 2.01 2.36 Llama-3.2-1B-instruct 3.33 2.06 2.25 Llama-3.2-1B-instruct-FT 5.00 1.53 2.05 Table 10: Trust Score Comparison of LLM Models in Multi-Agent Negotiations Negotiation Dynamics. Fig. 10a tracks the mean absolute error (MAE) between each agent’s proposal and the emerging consensus. All personalities tend to gravitate toward agreement, but their convergence trajectories differ: Neutral and Agreeable agents converge the fastest, while Disagreeable participants lag slightly. Crucially, Toxic agents that receive the warning incentive (T) “crack” in the final rounds, reducing their deviation by approximately 20% relative to the no-arbitration baseline (T∗). Toxicity Detection. The arbitral LLM runs in parallel, classifying each utterance as either toxic or non-toxic. Table 9 presents the confusion matrix for two challenging settings: Toxic vs. Neutral (T–N) and Toxic vs. Disagreeable (T–D). Even when the negative class is behaviorally close to toxicity (D), the classifier maintains high precision, recall, and F1scores (0.95, 0.85, 0.90), confirming reliable detection. Impact of the Warning Incentive. Fig. 10b summarizes 100 runs (with a four-iteration budget) across mixed-personality groups. Introducing the warning incentive (green secondary bars) consistently narrows the MAE distribution, indicating that toxic agents align more closely with the group. Because arbitration is fully parallelized, the additional overhead is negligible: convergence time increases by at most 0.4 seconds (median), well within practical limits. Takeaway. The judicial branch does not replace but rather amplifies LLM reasoning. By accurately flagging malicious turns and applying a calibrated penalty threat, the system reduces deviation, accelerates consensus, and preserves scalability—even in the presence of stubborn or toxic agents. 6.6. Exp. 5 – Trustworthiness Analysis To evaluate the effectiveness of our Trust Score framework, we conducted a comprehensive assessment of all prior models using identical multi-agent 31 negotiation inputs, ensuring a fair and consistent comparison of their trustworthiness. The results in Table 10 reveal substantial differences in trustworthiness between model sizes and architectures. The GPT-4 family consistently outperforms other models, with both GPT-4.1 and GPT-4.1-mini achieving trust scores above 4.8. These models also achieve exceptional coherence scores (5.0 and 4.96, respectively), reflecting a high degree of factual accuracy and logical consistency in their negotiation rationales. In contrast, smaller models exhibit markedly lower trustworthiness. The Llama-3.2 variants, particularly the 1B and 3B parameter models, receive trust scores below 2.5, largely due to weak coherence performance. The 1B-instruct model achieves a score of just 2.25, while its fine-tuned variant (1B-instruct-FT) performs even worse at 2.05, despite attaining perfect satisfaction scores. The fine-tuned versions of the 3B and 8B llama show negligible improvements compared to their non-fine-tuned counterparts. This indicates that, although smaller models may align well with negotiation goals, they struggle substantially with factual accuracy and logical reasoning. These findings clearly demonstrate that both model size and architecture have a significant impact on trustworthiness in multi-agent negotiation scenarios. The consistently low trust scores of smaller models indicate that they are not suitable for high-stakes negotiation tasks, where reliability, factual accuracy, and logical consistency are essential. Future works will explore the 70B+ parameter models for finding SLMs that are trustworthy for highstakes tasks. 7. Use Case: Autonomous SLA Brokering This section assesses the feasibility of Agoran open marketplace for an autonomous SLA brokering use case on the testbed. We present a live demo here. This final experiment integrates all Agoran components building the full Broker Agents (bAgents) on a live over-the-air OpenAirInterface and FlexRIC testbed, tracing the actions of three human stakeholders through four successive network phases. The objective is to demonstrate that: (i) LLM agents can translate free-form intent into concrete SLAs; (ii) these SLAs are dynamically renegotiated in response to changing radio or business conditions; and (iii) the resulting closed-loop control improves spectrum efficiency and slice QoS compared to conventional static planning. 32 Figure 11: End-to-end testbed on Autonomous SLA Brokering Case. RAN slice negotiations are conducted among three distinct stakeholders or services, converging on a Pareto-optimal consensus. The negotiated outcome is enforced as a policy at the Non-RT RIC, which subsequently triggers the Near-RT RIC to deploy a resource allocation xApp for dynamic adaptation of slice resources. 7.1. Scenario and Slice Personas Figure 11 shows the setup: each stakeholder/application is assigned a slice, a UE, and a bAgent, while the SRB hosts the mediator bAgent and enforces the agreed-upon policies by pushing them to a resource allocation xApp in the Near-RT RIC. Stakeholders. There are three stakeholders with different applications and needs. (i) Media-Flex (airport lounge): initially in phase PA operates as an eMBB slice (“deliver 4K video, minimize cost below 200 €”) and later in phase PD switches to URLLC for night-time augmented reality (AR) esports. (ii) Factory-Ops: operates as a daytime URLLC (phase PA) slice (“sub-5 ms motion control, cost irrelevant”), and switches to eMBB at night (phase PD) for bulk log uploads. (iii) IoT-Sense: a continuous, always-on mMTC slice that prioritizes ultra-low cost (≤50 €) and minimal energy consumption. Phases. Four stitched intervals (PA–PD, Table 11) emulate realistic channel dynamics by applying real-world channel quality patterns on the testbed, 33 Exp. Phase MCS Application Intent Min Tput (Mbps) Max Lat. (ms) Max Cost (€) PA 28 Media-Flex eMBB 60 10 200 Factory-Ops URLLC 5 2200 IoT-Sense mMTC 20 10 30 PB 7 Media-Flex eMBB 10 10 200 Factory-Ops URLLC 2 8200 IoT-Sense mMTC 5 10 30 PC 7 Media-Flex eMBB − ∞ − Factory-Ops URLLC − ∞ − IoT-Sense mMTC − ∞ − PD 28 Media-Flex URLLC 20 2200 Factory-Ops eMBB 40 10 200 IoT-Sense mMTC 20 10 50 Table 11: Per-slice constraints include an identical energy limit of 100 W in all phases, except for Phase PC, where the limit is set to 0 W. These constraints reflect the physical limitations of our testbed setup, which supports a maximum overall throughput of 133.7 Mbps. based on CQI/MCS time-series datasets [72, 73] of moving vehicles. In Phase PA, the system operates under good channel quality (MCS 28); in Phase PB, the testbed experiences a deep fade (MCS 7). During Phase PC, stakeholders trigger an energy-saving shutdown of the RAN, while in Phase PD, the network undergoes channel recovery (MCS 28) along with a role swap, where Media-Flex and Factory-Ops exchange their intents. At the beginning of each phase, all stakeholders restate their constraints and free-form intents (as shown in Table 12 for Phase PA). The optimizer then generates three Pareto-optimal offers (see Table 13 for Phase PA), which the agents negotiate (as shown in Figure 14 for Phase PA). 7.2. One-Round Consensus and SLA Selection Across all four phases, the three tenant agents and the Mediator Agent consistently converged to the same offer within a single JSON negotiation round, excluding the initial round request and starting from the mediator’s first response. Offer 2* was ultimately selected in Phase PA, with its underlying resource allocation detailed in Table 14. The complete negotiation transcript for this round is shown in Figure 14. Table 15 summarizes the KPIs of the agreed SLAs for each phase, as accepted by all agents; every slice constraint is satisfied in every phase (comparing Table 15 to Table 11), despite diverging slice objectives and the significant MCS degradation observed in Phase PB. 34 Slice Type Tput (Mbps) Lat. (ms) Cost (€) Energy (W) Full Intent in Natural Text by the Human Stakeholder MediaFlex eMBB 60 10 200 100 My slice use case is eMBB for an Airport lounge 4-K pipe, and it is crucial to maximize throughput. Push for the highest throughput, to maximize QoS/QoE of our users. FactoryOps URLLC 5 2 200 100 Our slice is used by robots that need sub-5 ms motion control and thus an URLLC guarantee. So, minimize latency as the ultimate goal. Moreover, high throughput is also important. IoTSense mMTC 20 10 50 100 My slice is providing coverage to Smart-city sensors and we have an mMTC case. I need you to prioritize cost and aim for the most cost-efficient solution, as our budget is limited. Mediator – – – – – My network goal is to serve and find a balance between the stakeholder’s slices objectives. However, prioritize minimum energy consumption to align with the country’s regulations. Table 12: In Phase PAthe human stakeholders express their slice intents in full naturallanguage form with associated constraints. ID Application Intent Tput [Mbps] Lat. [ms] Cost [€] Energy [W] 1 Media-Flex eMBB 60.72 5.66 61.52 13.39 Factory-Ops URLLC 34.82 1.49 133.14 12.72 IoT-Sense mMTC 38.14 5.45 2.19 0.05 Overall 133.68 12.60 196.85 26.16 2* Media-Flex eMBB 60.02 5.67 63.28 10.77 Factory-Ops URLLC 35.40 1.48 132.35 12.08 IoT-Sense mMTC 38.26 5.45 2.19 0.00 Overall 133.68 12.60 197.83 22.84 3 Media-Flex eMBB 60.06 5.67 68.16 12.84 Factory-Ops URLLC 35.52 1.48 133.94 12.92 IoT-Sense mMTC 38.11 5.45 1.64 2.46 Overall 133.68 12.60 203.74 28.22 Table 13: KPIs of the three Pareto-optimal offers for Experiment Phase PA, conducted under favorable channel conditions (MCS = 28). Each offer reflects different trade-offs across the three slices. Offer 2* was ultimately selected by unanimous agent consensus. The “Overall” rows report aggregate KPIs across all slices. ID Application Intent PRBs CPU Power Storage [%] [units] [W] [MB] 2* Media-Flex eMBB 44.9 5.6 10.8 26.1 Factory-Ops URLLC 26.5 21.7 12.1 44.5 IoT-Sense mMTC 28.6 0.0 0.0 1.1 Overall 100 27.3 22.9 71.7 Table 14: Underlying resource allocation corresponding to the selected Offer 2* in Experiment Phase PA. 35 ExpApplication Intent Tput Lat. Cost Energy PA Media-Flex eMBB 60 5.7 €63 10.8 W Factory-Ops URLLC 35 1.5 €132 12.1 W IoT-Sense mMTC 38 5.5 €20.0 W PB Media-Flex eMBB 11 8.8 €84 3.8 W Factory-Ops URLLC 7 3.4 €27 2.5 W IoT-Sense mMTC 7 7.5 €02.5 W PC Media-Flex eMBB 0 ∞€0 0 W Factory-Ops URLLC 0 ∞€0 0 W IoT-Sense mMTC 0 ∞€0 0 W PD Media-Flex URLLC38 1.5 €2.7 1.7 W Factory-Ops eMBB 58 5.7 €3 26.0 W IoT-Sense mMTC 38 5.5 €01.3 W Table 15: KPIs of the agreed SLA for each experiment phase. These SLAs reflect the capabilities of our testbed, which employs an OAI gNB with 40 MHz bandwidth and supports a maximum overall throughput of 133.7 Mbps. (a) Throughput of the three slices (b) Latency of the three slices Figure 12: Negotiated 5G slice performance across the four experimental phases PA– PD. 36 7.3. Throughput, Latency, and Resource Use Figure 12 shows the per-slice throughput and latency over 2,500 stitched Transmission Time Intervals (TTIs). Adaptivity. In Phase PA, the eMBB slice (Media-Flex) achieves over 60 Mbps, while the URLLC latency of the Factory-Ops slice remains below 2 ms. When spectral efficiency collapses (Phase PB), the agents renegotiate reduced but still feasible performance targets. In Phase PC, the stakeholders jointly request a power-off window, resulting in zero throughput and negligible energy consumption. Finally, in Phase PD, the intent swap is honored: Factory-Ops is prioritized for eMBB, reaching approximately 60 Mbps, while Media-Flex latency is held around 1.5 ms. Figure 13b shows physical PRB utilization for Factory-Ops, comparing Agoran’s dynamic negotiation with conventional static configurations. Dynamic allocation reduces PRB usage by 15,888 PRBs (24%) during periods of low demand and adds only 10,422 PRBs (16%) during high-demand intervals. This results in an overall net PRB saving of 8.3% over the full trace. Moreover, this flexible allocation enables targeted, on-demand improvements in Factory-Ops throughput, as shown in Figure 13a. When the slice switches to eMBB priority in Phase PD, throughput increases by up to 66%—from 35 Mbps in Phase PAto 58 Mbps in Phase PD. Finally, Figure 13c compares the latency of the Media-Flex slice under negotiated and static SLAs. Agentic negotiation reduces the median RTT by 73.4%, from 5.7 ms to 1.5 ms, once the slice transitions to URLLC. 7.4. Discussion and Takeaways •One-Round Consensus. The tripartite agent design consistently reaches unanimous agreement after a single message exchange, despite conflicting stakeholder priorities and dynamic radio conditions. •Constraint Satisfaction & Flexible QoS Gains. Negotiated SLAs always satisfy slice-level constraints while enabling flexible QoS improvements: aggregate throughput increases by up to 66%, and URLLC latency is reduced by up to 73.4%, compared to the static baseline. •Spectrum Efficiency. Closed-loop control achieves a net 8.3% PRB saving over the full trace, demonstrating more efficient and targeted resource allocation. 37 (a) Factory-Ops throughput (b) Factory-Ops PRB utilization (c) Media-Flex latency Figure 13: Negotiated versus static SLA. Dynamic reallocation increases throughput, releases unused PRBs (green), and allocates additional PRBs only when beneficial (red), while achieving lower latency when Media-Flex becomes URLLC-critical in Phase PD. 38 Figure 14: Snapshot of multi-agent negotiations during Experiment Phase PA, powered by GPT-4.1. The NSC Agents represent three stakeholders, Media-Flex, Factory-Ops, and IoT-Sense, in the Agoran marketplace. The agents negotiate over the three Paretooptimal offers generated by the multi-objective optimizer (Table 13), ultimately reaching consensus on Offer 2 as the most balanced solution. Agreement is achieved in a single round, excluding the initial expression of intent. •Human-Friendly Ultra-Flexibility. Stakeholders articulate rich, naturallanguage intents (e.g., “AR e-sports all night”) rather than relying on rigid slice templates. The optimizer and agents translate these into quantitatively optimal, regulation-compliant resource directives. The negotiation framework remains agile with respect to both evolving stakeholder intents and time-varying channel conditions. In summary, this use case validates Agoran as a fully autonomous yet human-centric marketplace that reconciles high-level business intent with radio resource constraints, adapts through non-real-time control loops, and enhances both QoS and spectrum efficiency, an essential capability for ultraflexible 6G deployments. 39 [18] B. Martini, M. Gharbaoui, P. Castoldi, Intent-based network slicing for sdn vertical services with assurance, Future Generation Computer Systems 142 (2023) 101–116. doi:10.1016/j.future.2022.12.033. [19] A. M. da Costa, L. Contreras, Integration of network slice controller for enhanced intent-based networking in 5g/6g networks, in: Proc. ACM MobiArch, 2023, pp. 1–6. doi:10.1145/3615587.3615989. [20] M. H. Chowdhury, Accelerator: An intent-based intelligent resourceslicing scheme for sfc-based 6g application execution over sdnand nfvempowered zero-touch networks, Frontiers in Communications and Networks 5 (2024). doi:10.3389/frcmn.2024.1385656. [21] T. Tsourdinis, I. Chatzistefanidis, N. Makris, T. Korakis, Ai-driven service-aware real-time slicing for beyond 5g networks, in: IEEE INFOCOM 2022 - IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS), 2022, pp. 1–6. doi:10.1109/ INFOCOMWKSHPS54753.2022.9798391. [22] T. Tsourdinis, I. Chatzistefanidis, N. Makris, T. Korakis, N. Nikaein, S. Fdida, Service-aware real-time slicing for virtualized beyond 5g networks, Computer Networks 247 (2024) 110445. doi:https://doi.org/10.1016/j.comnet.2024.110445. URL https://www.sciencedirect.com/science/article/pii/ S1389128624002779 [23] O-RAN Alliance, O–ran architecture description, Tech. Rep. O–RAN.WG1.OAD–R003 v10.00 (Oct. 2023). URL https://www.o-ran.org/specifications [24] AI-RAN Alliance, Ai-ran alliance vision and mission white paper, Tech. rep. (Dec. 2024). URL https://ai-ran.org/wp-content/uploads/2024/12/AI-RAN_ Alliance_Whitepaper.pdf [25] 3GPP, Management and orchestration; self-organising networks (son) for 5g networks, Tech. Rep. 3GPP TS 28.313 V18.1.0 (Jul. 2024). URL https://www.etsi.org/deliver/etsi_ts/128300_128399/ 128313/18.01.00_60/ts_128313v180100p.pdf 46 [26] A. Shahid, A. Kliks, A. Al-Tahmeesschi, A. Elbakary, A. Nikou, A. Maatouk, A. Mokh, A. Kazemi, A. De Domenico, A. Karapantelakis, et al., Large-scale ai in telecom: Charting the roadmap for innovation, scalability, and enhanced digital experiences, arXiv preprint arXiv:2503.04184 (2025). [27] J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. J. Ahn, M. Bosma, Q. Le, Chain-of-thought prompting elicits reasoning in large language models, Proceedings of the National Academy of Sciences 119 (46) (2022) e2209413119. doi:10.1073/pnas.2209413119. [28] S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, C. Xiong, React: Synergizing reasoning and acting in language models, arXiv 2210.03629 (2023). URL https://arxiv.org/abs/2210.03629 [29] T. Schick, J. Pfeiffer, H. Sch¨utze, Toolformer: Language models can teach themselves to use tools, in: Proc. ICLR, 2024. URL https://openreview.net/forum?id=7aQdV9BbDv [30] J. Liu, M. Saeidi, D. Goldwasser, Agentbench: Benchmarking large language models as agents, arXiv 2308.11432 (2023). URL https://arxiv.org/abs/2308.11432 [31] A. Tamkin, M. Brundage, J. Clark, D. Ganguli, Understanding the capabilities, limitations, and societal impact of large language models, 2021, URL https://arxiv. org/abs/2102.02503.(cited on pp. 29, 34, and 39). [32] IEEE Standards Association, Ieee standard 7000-2021—model process for addressing ethical concerns during system design (2021). URL https://standards.ieee.org/standard/7000-2021.html [33] I. Solaiman, M. Brundage, G. Hadfield, Evaluating the social impact of released large language models, arXiv 2306.06690 (2023). URL https://arxiv.org/abs/2306.06690 [34] N. Nikaein, M. K. Marina, S. Manickam, A. Dawson, R. Knopp, C. Bonnet, Openairinterface: A flexible platform for 5g research, ACM SIGCOMM Computer Communication Review 44 (5) (2014) 33–38. 47 [35] R. Schmidt, M. Irazabal, N. Nikaein, Flexric: an sdk for next-generation sd-rans, in: Proceedings of the 17th International Conference on emerging Networking EXperiments and Technologies, 2021, pp. 411–425. [36] L. Bariah, Q. Zhao, H. Zou, Y. Tian, F. Bader, M. Debbah, Large language models for telecom: The next big thing?, arXiv preprint arXiv:2306.10249 (2023). [37] Z. Lin, G. Qu, Q. Chen, X. Chen, Z. Chen, K. Huang, Pushing large language models to the 6g edge: Vision, challenges, and opportunities, arXiv preprint arXiv:2309.16739 (2023). [38] H. Zou, Q. Zhao, L. Bariah, M. Bennis, M. Debbah, Wireless multiagent generative ai: From connected intelligence to collective intelligence, arXiv preprint arXiv:2307.02757 (2023). [39] L. He, G. Sun, D. Niyato, H. Du, F. Mei, J. Kang, M. Debbah, et al., Generative ai for game theory-based mobile networking, arXiv preprint arXiv:2404.09699 (2024). [40] S. G. Patil, T. Zhang, X. Wang, J. E. Gonzalez, Gorilla: Large language model connected with massive apis, arXiv preprint arXiv:2305.15334 (2023). [41] Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, et al., Toolllm: Facilitating large language models to master 16000+ real-world apis, arXiv preprint arXiv:2307.16789 (2023). [42] F. H. Grings, G. Z. Bruno, L. R. Prade, C. B. Both, J. C. Brito, Nasp: Network slice as a service platform for 5g networks, Computer NetworksArXiv:2505.24051 (2025). [43] X. Wu, J. Farooq, Y. Wang, J. Chen, Llm-xapp: A large language model empowered radio resource management xapp for 5g o-ran, in: NDSS FutureG Workshop, 2025. doi:10.14722/futureg.2025.23057. [44] F. Lotfi, H. Rajoli, F. Afghah, Prompt-tuned llm-augmented drl for dynamic o-ran network slicing, arXiv 2506.00574 (2025). [45] M. Elkael, M. Polese, R. Prasad, S. Maxenti, T. Melodia, Allstar: Automated llm-driven scheduler generation and testing for intent-based ran, arXiv 2505.18389 (2025). 48 [46] A. Maatouk, K. C. Ampudia, R. Ying, L. Tassiulas, Tele-llms: A series of specialized large language models for telecommunications (2024). arXiv:2409.05314. URL https://arxiv.org/abs/2409.05314 [47] Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, I. Mordatch, Improving factuality and reasoning in language models through multiagent debate, arXiv preprint arXiv:2305.14325 (2023). [48] J. Zhang, X. Xu, S. Deng, Exploring collaboration mechanisms for llm agents: A social psychology view, arXiv preprint arXiv:2310.02124 (2023). [49] C.-M. Chan, W. Chen, Y. Su, J. Yu, W. Xue, S. Zhang, J. Fu, Z. Liu, Chateval: Towards better llm-based evaluators through multi-agent debate, arXiv preprint arXiv:2308.07201 (2023). [50] A. Mekrache, A. Ksentini, C. Verikoukis, Intent-based management of next-generation networks: An llm-centric approach, Ieee Network (2024). [51] M. Ameur, B. Brik, A. Ksentini, Leveraging llms to explain drl decisions for transparent 6g network slicing, in: 2024 IEEE 10th International Conference on Network Softwarization (NetSoft), IEEE, 2024, pp. 204– 212. [52] X. Wu, J. Farooq, Y. Wang, J. Chen, Llm-xapp: A large language model empowered radio resource management xapp for 5g o-ran, in: Symposium on Networks and Distributed Systems Security (NDSS), Workshop on Security and Privacy of Next-Generation Networks (FutureG 2025), San Diego, CA, 2025. [53] IEEE Standards Association, Ieee 7001tm – standard for transparency of autonomous systems (2021). URL https://standards.ieee.org/standard/7001-2021.html [54] 3rd Generation Partnership Project (3GPP), Telecommunication management; Management and orchestration; Concepts, use cases and requirements, 3GPP Technical Specification TS 28.530 V18.2.0, 3GPP, Sophia Antipolis, France, release 18 (Jan. 2025). URL https://www.3gpp.org/DynaReport/28530.htm 49 [55] 3GPP, Telecommunication management; Management and orchestration; Provisioning, 3GPP Technical Specification TS 28.531 V18.8.0, 3GPP, Sophia Antipolis, France, release 18 (Jan. 2025). URL https://www.3gpp.org/DynaReport/28531.htm [56] S. D’Oro, M. Polese, L. Bonati, H. Cheng, T. Melodia, dapps: Distributed applications for real-time inference and control in o-ran, IEEE Communications Magazine 60 (11) (2022) 52–58. [57] 3GPP, NR; User Equipment (UE) radio access capabilities, Tech. Rep. TS 38.306 V18.2.0, 3rd Generation Partnership Project, release 18 (Jun. 2024). URL https://www.etsi.org/deliver/etsi_ts/138300_138399/ 138306/18.02.00_60/ts_138306v180200p.pdf [58] 3GPP, NR; Physical channels and modulation, Tech. Rep. TS 38.211 V18.3.0, 3rd Generation Partnership Project, release 18 (Mar. 2025). URL https://www.etsi.org/deliver/etsi_ts/138200_138299/ 138211/18.03.00_60/ts_138211v180300p.pdf [59] L. Kleinrock, Queueing Systems, Volume I: Theory, John Wiley & Sons, New York, 1975. [60] K. Deb, A. Pratap, S. Agarwal, T. Meyarivan, A fast and elitist multiobjective genetic algorithm: NSGA-II, IEEE Transactions on Evolutionary Computation 6 (2) (2002) 182–197. doi:10.1109/4235.996017. [61] N. Pipitone, G. H. Alami, Legalbench-rag: A benchmark for retrievalaugmented generation in the legal domain, arXiv 2408.10343 (2024). URL https://arxiv.org/abs/2408.10343 [62] H. Li, Y. Chen, Y. Hu, Q. Ai, J. Chen, X. Yang, J. Yang, Y. Wu, Z. Liu, Y. Liu, Lexrag: Benchmarking retrieval-augmented generation in multiturn legal consultation conversation, arXiv 2502.20640, findings of ACL 2025 (2025). URL https://arxiv.org/abs/2502.20640 [63] H. T. Nguyen, K. Satoh, Krag: A knowledge-representation-augmented generation framework for legal reasoning, arXiv 2410.07551 (2024). URL https://arxiv.org/abs/2410.07551 50 [64] S. Bouhanna, A. Gupta, L. Romary, Lexdrafter: Retrieval-augmented drafting of legislative definitions, in: Proc. JURIX, 2024. URL https://arxiv.org/abs/2403.16295 [65] I. Pitsiorlas, G. Arvanitakis, M. Kountouris, Trustworthy intrusion detection: Confidence estimation using latent space, in: 2024 22nd International Symposium on Modeling and Optimization in Mobile, Ad Hoc, and Wireless Networks (WiOpt), 2024, pp. 92–98. [66] I. Pitsiorlas, N. Jamoussi, M. Kountouris, A conformal predictive measure for assessing catastrophic forgetting (2025). arXiv:2505.10677. URL https://arxiv.org/abs/2505.10677 [67] Z. Guo, R. Jin, C. Liu, Y. Huang, D. Shi, L. Yu, Y. Liu, J. Li, B. Xiong, D. Xiong, et al., Evaluating large language models: A comprehensive survey, arXiv preprint arXiv:2310.19736 (2023). [68] J. Fu, S.-K. Ng, Z. Jiang, P. Liu, Gptscore: Evaluate as you desire, arXiv preprint arXiv:2302.04166 (2023). [69] Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, C. Zhu, G-eval: Nlg evaluation using gpt-4 with better human alignment, arXiv preprint arXiv:2303.16634 (2023). [70] C. Spearman, The proof and measurement of association between two things, The American Journal of Psychology 15 (1) (1904) 72–101. doi: 10.2307/1412159. [71] M. G. Kendall, A new measure of rank correlation, Biometrika 30 (1-2) (1938) 81–93. doi:10.1093/biomet/30.1-2.81. [72] I. Chatzistefanidis, N. Makris, V. Passas, T. Korakis, Ue statistics timeseries (cqi) in lte networks (2022). [73] T. Tsourdinis, I. Chatzistefanidis, N. Makris, T. Korakis, Ue network traffic time-series (applications, throughput, latency, cqi) in lte/5g networks, IEEE Dataport (2022). [74] European Parliament and Council, EU Artificial Intelligence Act — Consolidated Text, https://artificialintelligenceact.eu/, provisional agreement, accessed: 2 July 2025 (2024). 51 [75] National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF), https://www.nist.gov/itl/ ai-risk-management-framework, accessed: 2 July 2025 (2023). [76] Wikipedia contributors, HMAC — Hash-Based Message Authentication Code, https://en.wikipedia.org/wiki/HMAC, revision dated 1 July 2025, accessed: 2 July 2025 (2025). 52