scieee AI-readable full text Open interactive document viewer

Transformer-Empowered Actor-Critic Reinforcement Learning for Sequence-Aware Service Function Chain Partitioning

Hsu, Cyril Shih-Huan; Dalgkitsis, Anestis; papagianni, chrysa; Grosso, Paola

Abstract

In the forthcoming era of 6G networks, characterized by unprecedented data rates, ultra-low latency, and extensive connectivity, effective management of Virtualized Network Functions (VNFs) is essential. VNFs are software-based counterparts of traditional hardware devices that facilitate flexible and scalable service provisioning. Service Function Chains (SFCs), structured as ordered sequences of VNFs, are pivotal in orchestrating complex network services. Nevertheless, partitioning SFCs across multi-domain network infrastructures presents substantial challenges due to stringent latency constraints and limited resource availability. Conventional optimization-based methods typically exhibit low scalability, whereas existing data-driven approaches often fail to adequately balance computational efficiency with the capability to effectively account for dependencies inherent in SFCs. To overcome these limitations, we introduce a Transformer-empowered actor-critic framework specifically designed for sequence-aware SFC partitioning. By utilizing the self-attention mechanism, our approach effectively models complex inter-dependencies among VNFs, facilitating coordinated and parallelized decision-making processes. Additionally, we enhance training stability and convergence using -LoPe exploration strategy as well as Asymptotic Return Normalization. Comprehensive simulation results demonstrate that the proposed methodology outperforms existing state-of-the-art solutions in terms of long-term acceptance rates, resource utilization efficiency, and scalability, while achieving rapid inference. This study not only advances intelligent network orchestration by delivering a scalable and robust solution for SFC partitioning within emerging 6G environments, but also bridging recent advancements in Large Language Models (LLMs) with the optimization of next-generation networks.

Full text

1 Transformer-Empowered Actor-Critic Reinforcement Learning for Sequence-Aware Service Function Chain Partitioning Cyril Shih-Huan Hsu, Anestis Dalgkitsis, Paola Grosso, Chrysa Papagianni Abstract—In the forthcoming era of 6G networks, characterized by unprecedented data rates, ultra-low latency, and ubiquitous connectivity, effective management of Virtualized Network Functions (VNFs) is essential. VNFs are software-based counterparts of traditional hardware devices that facilitate flexible and scalable service provisioning. Service Function Chains (SFCs), structured as ordered sequences of VNFs, are pivotal in delivering complex network services. Nevertheless, splitting an SFC into multiple segments that are deployed across different network domains or infrastructure locations presents substantial challenges due to the potential heterogeneity of domain characteristic—such as terrestrial and satellite networks—along with quality of service (QoS) constraints and limited visibility of network state. Conventional optimization-based methods typically exhibit low scalability, whereas existing data-driven approaches often fail to adequately balance computational efficiency with the capability to effectively account for VNF inter-dependencies, inherent in SFCs. To overcome these limitations, we introduce a Transformer-empowered actor-critic framework specifically designed for sequence-aware SFC partitioning. By utilizing the selfattention mechanism, our approach effectively models complex inter-dependencies between VNFs, facilitating coordinated and parallelized decision-making processes. Furthermore, we improve training stability and convergence using the ϵ-LoPe exploration strategy as well as the Asymptotic Return Normalization. Comprehensive simulation results demonstrate that the proposed methodology outperforms existing state-of-the-art solutions in terms of long-term acceptance rates, resource utilization efficiency, and scalability while achieving fast inference. This study not only advances intelligent network orchestration by delivering a scalable and robust solution for SFC partitioning in 6G, but also bridges recent advancements in Large Language Models (LLMs) with the optimization of next-generation networks. Index Terms—service function chain partitioning, service optimization, network function virtualization, quality of service, zerotouch service orchestration, transformers, deep reinforcement learning I. INTRODUCTION THE rapid emergence of beyond 5G and 6G networks has made network slice and logical network abstraction central challenges in the optimization of network resources and services. Network services are forced to compete for a limited resource pool over the shared infrastructure to ensure robust Quality of Service (QoS) and meet stringent Service Level Agreements (SLA). New services not only demand higher cross-domain performance guarantees, but also operate in highly shared, multi-tenant environments, imposing significant University of Amsterdam, LAB42, Science Park 900, 1098 XH Amsterdam, The Netherlands (e-mail: [email protected], [email protected], [email protected], [email protected]). pressure on both pre-deployment and post-deployment optimization processes. This is where service partitioning becomes a crucial strategy. By distributing service functions across domains while considering key constraints, such as limited visibility of the network state (due to privacy regulations or domain-specific data governance policies) and cost-driven overbooking, network operators can preserve performance and meet SLA commitments. The 3rd Generation Partnership Project (3GPP) introduced End-to-End (E2E) logical network segmentation to enable perservice customization and management, powered by a range of technologies [1]. At the core, Software-Defined Networking (SDN) allows logically centralized, programmatic control of data flows. ETSI has been standardizing Network Functions Virtualization (NFV) that decouples Network Functions (NF) from dedicated hardware and enables scaling services as network demands evolve [2]. Service Function Chaining (SFC) adds another layer of dynamism, allowing for the dynamic linking of multiple NFs to create flexible, policy-driven services. Fig. 1 illustrates the process of SFC partitioning across multiple network domains. Fig. 1: Illustration of SFC partitioning across multiple network domains. The left side presents a logical SFC composed of five interconnected VNFs. The right side shows the outcome of partitioning, where the VNFs are assigned to three separate domains: A, B, and C. Solid lines denote intra-domain connections, while dashed lines represent inter-domain links. Finally, orchestration platforms integrate these technologies by automating network service lifecycle management, abstracting the complexities associated with creating, deploying, and reconfiguring network services and slices. This automation significantly reduces both Capital Expenditure (CAPEX) and Operational Expenditure (OPEX) by enabling efficient resource utilization, simplifying infrastructure management, and automating lifecycle operations, thereby lowering infrastructure investment and ongoing maintenance costs for telecommunication operators. However, the integration of heteroge- 2 neous domains, dynamic service requirements, and real-time decision-making introduces unprecedented orchestration complexity. Global service optimization algorithms have emerged to manage large-scale infrastructures involving multiple stakeholders and Infrastructure Service Providers (ISPs), raising privacy concerns among providers. Service partitioning alleviates these issues by logically dividing an SFC into segments, facilitating separation, scaling, and partial orchestration without requiring full visibility into each provider’s infrastructure. Despite its advantages, Service Function Chain Partitioning (SFCP) poses a complex and challenging optimization problem. The difficulty lies in jointly determining the optimal partitioning strategy and the placement of each segment across multiple network domains with heterogeneous and dynamic characteristics, including variations in resource capacity, latency, mobility, and reliability. These domains—spanning terrestrial, aerial, satellite, and other network infrastructures—exhibit diverse behaviors that must be carefully considered to satisfy stringent QoS constraints. This involves balancing resource availability, QoS requirements, inter-domain routing overhead, and data locality considerations, often under incomplete or uncertain network state information. The combinatorial nature of possible partitioning and placement decisions, coupled with dynamic network conditions and policy constraints, makes SFCP a highly non-trivial problem, requiring scalable, intelligent, and privacy-preserving solutions. In recent years, there has been a notable adoption rate of Artificial Intelligence (AI) and Machine Learning (ML) across a broad range of applications to optimize complex tasks. Studies have shown that these approaches can not only outperform humans, but also rival classical optimization techniques, all while reducing energy consumption, cost, and time [3], [4], [5]. Reinforcement learning (RL), in particular, has gained attention for its remarkable performance in decision-making and real-time control. Recently, the rise of Transformerbased architectures revolutionized the learning sector, offering unprecedented capabilities in processing and understanding language and large-scale datasets [6]. As these promising technologies continue to advance, they are increasingly being integrated into the networking domain, benefiting stakeholders across the spectrum, from end users to ISPs, particularly in the era beyond 5G, where efficiency and adaptability are critical. To this end, we investigate the problem of efficient SFC partitioning and propose Sequence-aware Differentiable ActorCritic RL (SDAC), a novel method that integrates Transformer layers within a differentiable actor-critic RL framework. We evaluated the effectiveness of SDAC through an extensive and diverse simulation-based environment and compared it with multiple state-of-the-art methods to showcase its ability to maintain a higher long-term acceptance rate. The contribution of this work is threefold: •We introduce a novel DRL framework, which leverages Transformer encoders within an actor-critic architecture to enable sequence-aware partitioning of SFCs. By leveraging the self-attention mechanism, our approach models inter-dependencies between VNFs, facilitating coordinated and parallel decision-making. Additionally, we introduce the ϵ-LoPe exploration strategy and Asymptotic Return Normalization to enhance training stability and convergence. These innovations address the limitations of existing sequential and parallel VNF partitioning methods, significantly improving the efficiency and effectiveness of SFCP in multi-domain network orchestration. •We introduce an innovative critic network, inspired by sentence encoders in Large Language Models (LLMs), that evaluates the entire sequence of VNF decisions holistically, providing a context-sensitive evaluation of the collective partitioning strategy. Alongside a relaxed representation for discrete actions, the policy can be updated in a fully differentiable manner, avoiding the stochastic sampling processes and credit assignment problem. Its adaptability to varying numbers of agents makes it a scalable solution for Multi-Agent Reinforcement Learning (MARL) problems, beyond the SFCP problem. •We provide a comprehensive evaluation of SDAC through extensive simulations, demonstrating its superior performance compared to state-of-the-art methods, including DRL-based and (meta-)heuristic-based approaches. SDAC achieves higher long-term acceptance rates, improved resource utilization, and greater scalability while maintaining efficient inference times. These results validate the effectiveness of the proposed method in managing network resources under varying conditions. The remainder of this paper is organized as follows. Section II offers a discussion of related work. Section III gives a comprehensive introduction to the problem formulation and system model. Section IV details the proposed method. Section V describes the experimental setup. Section VI presents an extensive performance evaluation study. Finally, the conclusion and future directions are given in Section VII. II. RELATED WORK The transition of networking from hardware-oriented network architectures to flexible, software-driven ones has made service partitioning, along with related concepts such as SFC, Forwarding Graphs (FGs), and VNG-FG embedding, critical to system performance and reliability [2]. The increasing complexity of multi-domain networks [1], pushed recent literature efforts to integrate data-driven approaches into service partitioning strategies. A. Optimization-based Approaches Initially, researchers leaned on classical optimization-based algorithmic and heuristic approaches for SFC optimization, based on multiple criteria. Some notable examples related to our work include the following: In [7], Quang et al., formulate the multi-domain VNF-FG embedding problem as an Integer Linear Programming (ILP) and utilize a heuristic algorithm for dynamic re-allocation. The authors in [8] study optimal network service decomposition and propose two algorithms, an ILP-based to minimize the cost of mapping, and a heuristic algorithm to solve the scalability issues of the ILP formulation, jointly studying the optimization of decomposition and embedding. The authors in [9] propose a Mixed Integer Programming (MIP) framework for optimal 3 virtual resource allocation in networked clouds, addressing the Virtual Network Embedding (VNE) problem with QoS constraints. Similarly, Dietrich et al. in [10] study the multiprovider Network Service Embedding (NSE), by decomposing the problem into SFC graph partitioning and mapping with ILP. The authors propose a centralized NSE orchestrator to address traffic scaling. Their evaluation emphasizes the optimization of the SFC graph partitioning and uncovers significant savings in service cost and resource consumption. The authors in [11] propose a hierarchical framework for virtual resource mapping in networked clouds, using Iterated Local Search (ILS) for cost-effective request partitioning among cloud providers and a Networked Cloud Mapping (NCM) approach for intra-domain resource embedding. In [12], the authors propose a game theory-inspired graph partitioning framework, and prove mathematically that a Nash equilibrium exists for satisfying server affinity, collocation, and latency constraints of SFCs. The formulated partitioning game is based on the various components of an SFC that have to be split in a set of partitions and executed as VMs or containers in specific servers. Li et al. [13] address dynamic SFC mapping and scheduling for Internet of Vehicles (IoV) services, proposing two Tabu Search-based algorithms to solve a mixed-integer linear programming (MILP) problem. Their approach optimizes service provider profit by enabling VNF live migration, reinstantiation, and rescheduling, considering migration costs and delays across heterogeneous nodes. Although most of these approaches guarantee convergence to a near-optimal solution, they suffer from scalability limitations due to the exponential growth of possible actions. Specifically, each SFC requires handling MNpotential actions, where Mrepresents the number of available deployment targets and Ndenotes the number of VNFs within the SFC. This combinatorial explosion, coupled with the lack of adaptability and long convergence times, substantially limits the practicality in real-world scenarios. B. Data-driven Approaches To overcome the limitations of classical optimization-based methods and embrace scalability, researchers pivoted their focus on AI data-driven solutions due to their remarkable performance that challenged the status quo in networking. The authors of [14] introduce an availability-aware SFC placement algorithm that leverages Proximal Policy Optimization (PPO). They evaluate the proposed algorithm by comparing it against two greedy algorithms in a variety of simulated scenarios using the Brazilian National Teaching and Research Network backbone topology. In stark contrast to the status quo, newer solutions have emerged with multi-agent implementations. Specifically, in [15], Pentelas et al. argue that most SFCP works develop centralized implementation hinders scalability. They propose a cooperative multi-agent RL scheme for decentralized SFCP. Experimental results show that their scheme outperforms centralized Double Deep Q-Learning (DDQN) and further analyze the behavior and information exchange of the agents. In a different note, Heo et al. in [16] propose a Graph Neural Network (GNN) based SFC model, which employs an encoder-decoder architecture. The encoder extracts a representation of the network topology, while the decoder estimates the probabilities of neighboring nodes processing a VNF. To support their case, the authors argue that related works with high-level intelligent models such as Deep Neural Networks (DNNs) are not efficient in utilizing the topology information of networks and cannot be applied to dynamically changing topologies, as the models expect a fixed topology. Their solution outperforms the DNN-based baseline model in the conducted experiments. Despite significant advancements, these works predominantly focus on either sequential or parallel VNF placement, each with inherent limitations. Sequential placement of VNFs, while offering strong coordination, is time consuming and constrained by the interdependence of prior decisions, which can impact overall efficiency [17], [18], [15]. Parallel VNF placement offers faster deployment, but suffers from weak coordination among VNFs, as individual placement decisions are made without full awareness of the broader context. C. Transformers and their Applications Recently Transformers have emerged as a groundbreaking architecture that disrupted Deep Learning (DL) research by demonstrating outstanding results. The use of attention mechanisms enables the effective modeling of complex, longrange dependencies within data sequences. The authors in [6] propose the groundbreaking DNN architecture of Transformers and demonstrate the remarkable results in natural language translation. Based on that, several studies [19], [20], [21] have showcased the successes of the architecture in various fields such as computer vision, natural language processing, audio processing and robotics. In the field of next-generation communication networks, authors in [22] discuss how the strong representation capabilities of Transformers can be used in wireless network design. Particularly, they explore various solutions for massive MultipleInput Multiple-Output (MIMO) beamforming and semantic communication problems. A recent work applies Transformers to automatic modulation classification [23], showing improved accuracy at low Signal-to-Noise Ratios (SNRs) and fewer model parameters compared to existing methods. A temporal Transformer was proposed for real-time QoS prediction in heterogeneous Internet of Things (IoT) applications [24], leveraging attention mechanisms for both shortand longterm sequence dependencies. Evaluated on diverse IoT traffic datasets, the model demonstrated robust and accurate QoS metric predictions. The authors in [25] propose a beamforming approach for 6G mmWave networks, combining a lightweight Transformer-CNN hybrid with multimodal data to ensure consistent QoS in dynamic environments. The method outperformed state-of-the-art baselines in beam prediction accuracy, received power, and efficiency across diverse scenarios. While researchers are actively investigating the use of Transformers in beyond 5G and 6G networking scenarios, substantial gaps remain in this area, particularly regarding their effective integration into practical, real-time network optimization and orchestration tasks. 4 In summary, although data-driven methods improve scalability and long-term performance over optimization-based ones, parallel decision-making often overlooks inter-VNF dependencies, while sequential decision-making captures them at the cost of computational efficiency. To address these gaps, we propose a Transformer-based actor-critic framework that models SFCs as token sequences, enabling parallel, context-aware, and scalable decision-making for efficient SFC partitioning. III. PROBLEM FORMULATION To effectively address the SFCP problem, we begin with a high-level overview, followed by a comprehensive description of the system model. This model serves as the foundation for capturing the interplay between the substrate network, SFC requests, and the constraints governing resource allocation. By formalizing these elements, we enable the optimization of long-term system performance while adhering to resource capacity and latency constraints. A. Problem Description An SFC is defined as an ordered sequence of VNFs that enforces a flow processing policy on network traffic. For example, a typical SFC might consist of VNFs such as VPN Gateway →Firewall →Encryption/Decryption →Packet Inspection, ensuring secure and efficient flow management for network applications. Each VNF in the SFC requires specific computing resources (e.g., CPU) to process the network traffic, while the connections between consecutive VNFs, referred to as virtual links (vLinks), consume network bandwidth for traffic forwarding. These VNFs and vLinks are hosted on a multi-datacenter system, where Data Centers (DCs) offer computing resources via their servers, and the physical links within and between DCs provide the necessary bandwidth. In network orchestration, SFC partitioning can occur in single-domain or multi-domain settings. In single-domain settings, a centralized controller manages the entire substrate network and makes decisions based on fine-grain information from all Point of Presences (PoPs) or DCs. In contrast, multidomain settings often involve decentralized control, where different domains (e.g., geographically distributed networks) operate independently, and only limited information can be shared between domains. In this work, we focus on the multidomain setting, where the centralized Network Functions Virtualization Orchestrator (NFVO) has access only to aggregated resource statistics (e.g., total available CPU and bandwidth) of each PoP. However, the limited granularity of information exchanged between the PoPs and the NFVO introduces significant uncertainty, as the NFVO cannot observe the exact state of the resource distribution within each PoP [15]. In the context of SFC management, partitioning typically refers to the process of logically dividing an SFC into smaller sub-chains or groups of VNFs without immediate consideration of the substrate network, whereas embedding refers to the subsequent mapping of VNFs or sub-chains onto physical resources. However, in the literature, the term partitioning is sometimes used broadly to describe both the partitioning and embedding phases together, particularly when the focus is on mapping VNFs across different domains or infrastructures. In this work, for clarity, we refer to partitioning as the mapping of VNFs at the inter-DC level, while embedding refers to the fine-grained mapping of VNFs within a single DC (intra-DC). Specifically, VNFs are mapped to PoPs, such as DCs, while vLinks are mapped to physical paths consisting of multiple physical links that connect the assigned PoPs. The SFCP problem becomes challenging due to the following factors: •Resource Constraints. Each PoP and physical link has limited CPU and bandwidth resources. Mapping VNFs and vLinks must respect these capacity limits. •Latency Constraints. An SFC must meet its E2E latency requirement, ensuring that network flows traverse the chain of VNFs within a specified SLA. •Limited Visibility. The system considers the settings where PoPs provide only aggregated descriptive statistics (e.g., total available CPU and bandwidth). This limitation imposes uncertainty when making decisions at the NFVO. •Resource Fragmentation. Even if aggregated resource capacity at a DC is sufficient to host a VNF, the VNF may still fail to be accommodated due to fragmented resources across the servers within the DC. •Long-Term Acceptance Rate. A mapping strategy that focuses excessively on optimizing on the current request may lead to inefficient resource allocation and fragmentation, making it difficult to accommodate subsequent requests. For example, aggressively placing VNFs on lowlatency data centers without considering future resource availability might leave scattered, unusable resources in data centers. Over time, this short-sighted strategy could reduce the overall acceptance rate. To summarize, the objective of SFCP is to find a strategy that places the VNFs and their corresponding vLinks onto the substrate network such that (i) The resource constraints are satisfied; (ii) The E2E latency requirement is met; and (iii) The long-term acceptance rate is maximized. Fig. 2: Architecture and flow of the NFVO with the proposed decision-making module (SDAC) for SFC partitioning. In this work, we consider a scenario where the embedding of partial SFCs within each DC is handled locally by its corresponding Virtual Infrastructure Manager (VIM). The focus of this work lies solely on the partitioning of SFCs at the 5 NFVO level. The integration of the proposed decision-making component (introduced in Section IV-B) within the NFVO for SFC partitioning is illustrated in Fig. 2. The process involves the interaction among the SFC requests, the NFVO, and the VIMs. The steps of this process are detailed as follows: Step 0: Reception of the SFC Request. The process starts with the NFVO receiving an SFC request through the Northbound Interface (NBI). The SFC request typically comprises a sequence of VNFs, denoted as VNF0,VNF1,...,VNFn, which must be deployed between a source (src) and a destination (dst) target DC. Upon receipt, the NFVO processes the request and extracts corresponding features, which may encompass the resource demands of each VNF, latency constraints, bandwidth requirements, and other QoS parameters. Step 1: Retrieval of the State of NFVIs. In the following, the NFVO retrieves the current state of the Network Function Virtualization Infrastructures (NFVIs) through the Southbound Interface (SBI). The SBI enables communication between the NFVO and the underlying VIMs, denoted as VIM0,VIM1,...,VIMm, each of which manages a corresponding DC (DC0,DC1,...,DCm). The NFVO gathers monitoting data regarding the available resources across these VIMs, including computational capacity and networking resources of the DCs. Step 2: Decision Making using the Joint State. The NFVO consolidates the processed SFC request (from Step 0) with the current state of the underlying infrastructure (from Step 1) and forwards this joint information to the decision-making module (introduced in Section IV-B). Step 3: Delivery of Results for Deployment. Upon receiving the assignments from the decision-making module, the NFVO distributes these results to the respective VIMs through the SBI. The results specify the allocation of VNFs within the SFC across the DCs. For instance, VNF0might be assigned to DC1, while VNF1and VNF2could be co-located to DC3, depending on resource availability and objectives. Each VIM subsequently embeds the assigned VNFs within its corresponding DC, ensuring that the SFC is instantiated and operational in accordance with the partition assignments. B. System Model Substrate Network. The substrate network, denoted as Gs= (Vs, Es), is modeled as a graph where Vsis the set of Data Centers (DCs), each u∈Vswith a CPU capacity of Cu(in units of available CPU resources) and Esis the set of physical links between data centers, each l∈Eswith a bandwidth capacity of Bland latency Ll. The inter-DC latency is modeled as the sum of latencies along the shortest path between two DCs, computed using Dijkstra’s algorithm. We consider intra-DC latency as negligible, thus only inter-DC latency contributes to the overall E2E latency. SFC Requests. SFC requests are received sequentially over time, with one request arriving during each time slot t. For simplicity, the notation tis omitted in this section, as we focus on the context within a single time slot. Each request is modeled as a 6-tuple (Gv, tarr, t∆, LSLA,src,dst). Gv= (Vv, Ev)represents the SFC as a directed graph, with TABLE I: Notation Table Notation Description GsSubstrate network graph with DCs as nodes and interDC links as edges. VsSet of data centers. EsSet of physical links between data centers. BlBandwidth capacity of the physical link l∈Es. LlLatency of the physical link l∈Es. ¯ CuAggregated remaining CPU capacity of DC u∈Vs. ¯ BuAggregated remaining bandwidth capacity of DC u∈ Vs. GvSFC request graph with VNFs as nodes and virtual links as edges. VvSet of VNFs in the SFC. EvSet of virtual links between VNFs in the SFC. DvCPU demand of VNF v∈Vv. DeBandwidth demand of the virtual link e∈Ev. LSLA End-to-end latency SLA for an SFC request. xv,u Binary variable: 1 if VNF vis assigned to DC u, 0 otherwise. ye,l Binary variable: 1 if the virtual link eis mapped to the physical link l, 0 otherwise. δBinary variable: 1 if the request is successfully admitted, 0 otherwise. Vvthe set of VNFs, each v∈Vvwith a CPU demand of Dv and Evthe set of virtual links between VNFs, each e∈Ev with a bandwidth demand of De.Tarr is the arrival time, t∆ is the service lifetime and LSLA is the E2E latency constraint for the request. The src and dst are two auxiliary VNFs with zero resource requirements (Dsrc =Ddst = 0) that represent the endpoints of the SFC. Accordingly, two additional edges are introduced to the service graph: (i) the edge from src to the first VNF v1∈Vvand (ii) the edge from the last VNF v|Vv|∈Vvto dst. The objective is to assign each v∈Vvto a DC u∈Vs, ensuring that the SFC’s E2E latency satisfies LSLA, while respecting the resource limitations of the substrate network. Partitioning Decision. The partitioning decision is made at the NFVO level, which assigns each VNF v∈Vvto a specific DC u∈Vs: xv,u =(1if VNF vis assigned to DC u, 0otherwise. (1) At the time of making this decision, the NFVO only has aggregated information of the available resources in each DC: •¯ Cu: the sum of remaining CPUs in DC u. •¯ Bu: the sum of available bandwidth for links associated with DC u. The actual embedding of VNFs inside each DC (e.g., mapping VNFs to individual computing nodes within the DC) is handled by the corresponding VIM of each DC. In contrast, the NFVO is responsible for inter-DC partitioning decisions, determining how VNFs are distributed across different DCs, rather than handling placements within a single DC. Additionally, the mapping of virtual links in Evis decided implicitly at the same time when the VNF-to-DC mapping is determined. The shortest paths between assigned DCs are computed using Dijkstra’s algorithm, and virtual links are mapped to the corresponding substrate links in Es, without the need to explicitly decide on link mappings. 6 Resource Constraints. The following defines the key constraints on CPU and bandwidth capacities, alongside the issue of resource fragmentation. •CPU Capacity. The CPU demand of VNFs assigned to a DC must not exceed its aggregated capacity: X v∈Vv Dv·xv,u ≤¯ Cu,∀u∈Vs.(2) •Bandwidth Capacity. There are two types of bandwidth constraints to consider: – Intra-DC Bandwidth Capacity Constraint. The bandwidth demand for virtual links mapped within a DC must not exceed the aggregated bandwidth capacity of that DC: X e∈Ev De·xv1,u ·xv2,u ≤¯ Bu,∀u∈Vs,(3) where ¯ Buis the aggregated bandwidth capacity available within DC u, and xv1,u and xv2,u indicate whether two consecutive VNFs (v1, v2)associated with the virtual link eare assigned to DC u. – Inter-DC Bandwidth Capacity Constraint. The bandwidth demand for virtual links mapped between DCs must not exceed the remaining bandwidth capacity of the physical links between them: X e∈Ev De·ye,l ≤Bl,∀l∈Es,(4) where ye,l indicates whether the virtual link eis mapped to the physical link l. •Resource Fragmentation. Despite the aggregated resource demands described in formula (2) and (3) are satisfied for a given request, the request could still be rejected due to resource fragmentation, where: –Available resources are scattered across multiple computing nodes within a DC, making it impossible to allocate sufficient contiguous resources to meet the request’s demands. –Similarly, bandwidth fragmentation on physical links within a DC can hinder the successful mapping of virtual links even if the total available bandwidth is sufficient. Thus, meeting the aggregated resource demands is a necessary but not sufficient condition for accepting an SFC request. E2E Latency Constraint. The E2E latency Le2e is computed by traversing the VNFs v1, v2, . . . , vkin the SFC sequentially and summing the inter-DC latencies between the DCs they are mapped to: Le2e =X e∈Ev Ll·ye,l,∀l∈Es,(5) where ye,l = 1 if the virtual link eis mapped to the substrate link l. The E2E latency must satisfy the SLA: Le2e ≤LSLA.(6) Objective. The goal is to maximize the long-term acceptance rate of SFC requests, where acceptance occurs only if all resource and latency requirements are satisfied. This can be formulated as: max {x(t)}T t=1 1 T T X t=1 δt(x(t)), where δt(x(t)) = (1,if the request at time tis accepted, 0,otherwise. (7) Here, x(t)denotes the assignment of decision variables at time tdefined in (1). The overall constraints below are the necessary conditions for achieving δt(x(t)) = 1: X u∈Vs xv,u(t)=1,∀v∈Vv(t), X v∈Vv(t) Dv(t)·xv,u(t)≤¯ Cu(t),∀u∈Vs, X e∈Ev(t) De(t)·xv1,u(t)·xv2,u(t) ≤¯ Bu(t),∀u∈Vs, X e∈Ev(t) De(t)·ye,l(t)≤Bl(t),∀l∈Es, X e∈Ev(t) Ll·ye,l(t)≤LSLA(t). (8) IV. PROPOSED METHOD Building on the problem formulation, we now present our proposed approach to address the SFCP problem. Given the NP-hard nature of SFCP [26], we reformulate the problem as a Markov Decision Process (MDP), enabling us to leverage RL techniques for a scalable and efficient solution. This section is organized into two parts: (i) the MDP formulation, where we define the states, actions, transitions, and rewards for the problem, and (ii) the proposed DRL solution, where we detail our DRL approach. A. MDP Formulation To solve the SFCP problem using reinforcement learning, we first model it as an MDP. An MDP is defined as a 4-tuple (S,A,T,R), consisting of: States (S). The state captures the characteristics of an individual VNF within an SFC request and the substrate network state, represented as a vector: s= [sreq, ssub],(9) where: •sreq = [Dv, De, Tarr, t∆, TSLA, src, dst]represents the request state, with the individual elements defined in Section III-B. •ssub =¯ C, ¯ Brepresents the substrate state, where: –¯ Ci∈¯ C1,¯ C2,..., ¯ Cmis the aggregated remaining CPU resources at data center i. –¯ Bi∈¯ B1,¯ B2,..., ¯ Bmis the aggregated remaining bandwidth at data center i. 7 –mis the total number of data centers. It is important to note that while the state sfor each VNF within the same SFC shares the same ssub, the sreq component will differ based on the resource requirements of the VNF. Actions (A). The action specifies the target DC for mapping each VNF in the SFC request. Formally, let Vv= {v1, v2, . . . , vn}denote the set of VNFs in the SFC request, and Vs={u1, u2, . . . , um}represent the set of available DCs in the substrate network. An action ai∈ A for VNF vi∈Vv is represented as a one-hot vector of size m, where: ai= [a1 i, a2 i, . . . , am i]∈ {0,1}m, m X j=1 aj i= 1.(10) Here, arg max(ai) = jexplicitly identifies the index jof the DC ujto which the VNF viis assigned. Transitions (T). The transitions describe how the state evolves as a result of an action. After taking an action ain state s: •The substrate state ssub is updated to reflect the resources consumed by the embedded VNFs and their virtual links. If the embedding fails, the substrate state remains unchanged. Additionally, resources consumed by expired services are released back into the substrate. •The request state sreq transitions to the next incoming SFC request, with its resource demands, arrival time, service lifetime, and SLA. The transition function T(s, a)ensures that the new state s′ reflects both the updated substrate state and the attributes of the next SFC request. The detailed steps of this process are provided in Section V-A. Rewards (R). The reward evaluates the outcome of the partitioning decision: r(s, a) = (1,if the request is accepted, 0,otherwise.(11) A reward of 1is assigned if the SFC request is accepted, meaning all VNFs are successfully embedded, and the E2E latency satisfies the SLA. A reward of 0is assigned if the request is rejected, either due to a failure in embedding any VNF or a violation of the SLA. Note that this immediate reward reflects the acceptance indicator δin (7). Objective. The goal is to maximize the expected cumulative discounted reward over time, which is formulated based on the standard RL framework. Formally, let r(st, at)denote the reward received at time step t, and γ∈[0,1] be the discount factor, the objective is to learn a policy πthat maximizes the expected cumulative discounted reward: Eτ∼π"T X t=0 γtr(st, at)#,(12) where Tis the time horizon, and the expectation is taken over the trajectories τinduced by the policy π. This formulation reformulates the problem in III-B as an MDP, enabling its solution via DRL algorithms. B. The Proposed Framework: SDAC We propose a DRL framework, named Sequence-aware Differentiable Actor-Critic RL, for efficient SFC partitioning. Our approach incorporates Transformer encoders in both the actor and critic networks, leveraging their ability to model inter-dependencies among VNFs and make informed, coordinated decisions, as shown in Fig. 3. Below, we detail the components of the framework. Algorithm 1: Learning Framework of SDAC Initialize actor network πθand critic network Qϕ Initialize target networks πθ′←πθand Qϕ′←Qϕ Initialize replay buffer D Initialize perturbation factor ϵ for episode = 1 to Ndo Reset environment and observe the initial state s0 for t= 1 to total number of SFC requests do // Parallel Action Selection at←ϵ-LoPe(st, πθ, ϵ) // Environment Interaction Run action at, get reward rtand next state st+1 // Transition Storing Store transition (st, at, rt, st+1)in D // Critic Update Sample Mmini-batch transitions (s, a, r, s′) from Dand compute target Q-values: y←r+γQϕ′(s′, πθ′(s′)) Update critic by minimizing the loss: Lcritic =1 MX(Qϕ(s, a)−y)2 // Actor Update Update actor by minimizing the loss: Lactor =−1 MXQϕ(s, πθ(s)) // Target Network Update Update target networks: ϕ′←τϕ + (1 −τ)ϕ′ θ′←τθ + (1 −τ)θ′ end end return πθ Actor-Critic Framework. The actor-critic architecture is a widely used framework in DRL that involves: •Actor Network. The actor generates the actions (partitioning decisions for VNFs) based on the current state of the system. It learns a policy πθ(s)that maps states sto actions a. •Critic Network. The critic evaluates the quality of the actions proposed by the actor by estimating the Q-value, Qϕ(s, a), which represents the expected cumulative reward for taking action ain state s. 8 Fig. 3: Overview of the SDAC framework with Transformer-based architecture for sequence-aware SFC partitioning. During training, the actor and critic are updated iteratively, with the critic providing feedback to guide the actor toward optimal policies. This architecture enables stable and efficient learning, as the critic guides the actor’s updates while the actor improves the policy. Transformer Layers. The Transformer architecture [6], originally proposed for NLP tasks, is known for its ability to model relationships between input tokens using self-attention. A Transformer encoder layer comprises two main components: a multi-head self-attention mechanism and a MultiLayer Perceptron (MLP). Furthermore, there are two common Transformer encoder variants: Post-Norm, where layer normalization follows the residual blocks, and Pre-Norm, where Layer Normalization precedes the residual blocks. Post-Norm can result in unstable training due to the compounded effects of residual connections and Layer Normalization placement, whereas Pre-Norm provides more stable weight updates [27]. In the SFCP problem, the VNFs within an SFC are highly inter-dependent because the partitioning decision for one VNF can significantly influence the feasibility and performance of partitioning for other VNFs. By interpreting each VNF as a “token” and the entire SFC as a sequence, the Transformerbased actor and critic can effectively accounts for these interdependencies. Concretely, the prediction process resembles a sequence-to-sequence translation, where the VNF sequence is “translated” into a corresponding sequence of target DC indices. This formulation allows the model to efficiently map VNFs to their target DCs while exploiting the contextual relationships within the SFC. •Self-Attention for Inter-Dependency Modeling. The self-attention mechanism, a core component of Transformers, allows the model to capture inter-dependencies between input elements by dynamically computing attention weights. For a sequence of nVNF tokens denoted as: S={s1, s2, . . . , sn} ∈ Rn×d,(13) where sis defined in (9) and dis its dimension. Selfattention computes a weighted combination of all tokens for each token in the sequence, resulting in S′. Specifically, self-attention involves three key steps: Projection of Input. Project the input embeddings into query Q, key K, and value Vmatrices: Q=S·WQ, K =S·WK, V =S·WV,(14) where WQ, WK, WV∈Rd×dkare learned projection matrices, and dkis the dimension of the query vector Qand key vector K. It is a hyperparameter of the Transformer model, typically smaller than the input dimension d. Computation of Attention Scores. Compute attention scores Eusing the dot product of queries and keys, scaled by the dimension dk: E=softmax QK⊤ √dk,(15) where E∈Rn×ncontains attention weights, representing the importance of each token with respect to others. Weighted Sum of Values. Compute the output as a weighted sum of values: S′=E·V. (16) This mechanism allows each VNF token to attend to all other tokens, capturing their inter-dependencies. This means that the partitioning decision for one VNF incorporates information on the resource demands and constraints of other VNFs, enabling coordinated decisions. •Actor Network with Transformers. The actor uses Transformer encoder layers to process the sequence of VNFs in the SFC. The input to the actor is the set of tokens of the entire SFC S, as defined in (13). This matrix is passed through the Transformer encoder layers. The actor outputs a probability distribution over the target DCs for each VNF, representing the policy πθ(s). Specifically, the policy can be expressed as: πθ(S) = [πθ(s1), πθ(s2), . . . , πθ(sn)],(17) where πθ(si) = aiis a one-hot vector that encodes the selected action over the available DCs for the i-th VNF. •Critic Network with Transformers. The critic Qϕ(s, a) also adopts Transformer encoder layers to process the sequence of VNFs and their corresponding partitioning 9 decisions. Instead of evaluating each decision individually, the critic evaluates the entire sequence of stateactions for the SFC as a whole. In particular, the input to the critic encodes the entire SFC, represented by the state matrix Sdefined in (13), concatenated with the partitioning decisions A, where A= [a1, a2, . . . , an] represents the assignment of each VNF to a specific DC. Formally, the input to the critic is defined as: H= [h1, h2, . . . , hn],(18) where hi= [si;ai],siis the feature of the i-th VNF, aiis the one-hot representation of the partitioning decision for i-th VNF, and [·;·]denotes concatenation. The output of the Transformer is then pooled into a single fixed-length vector using element-wise mean pooling. This approach borrows the idea from Universal Sentence Encoder (USE) architecture [28] commonly used in NLP tasks, where variable-length sequences are mapped into fixed-length representations. The pooled representation is then passed through a linear layer to compute the Q-value, providing a holistic evaluation of the decisions for the entire SFC. •Target Networks To stabilize training, two target networks are employed for both the actor πθ′(s)and critic Qϕ′(s, a). A target network is a separate network used to compute the target values for updates to the main network. This separation prevents the network from chasing its own moving predictions, which can lead to instability or divergence in training. The target network’s parameters are updated less frequently or with a smoothing mechanism (e.g., Polyak averaging) to improve convergence: ϕ′←τϕ + (1 −τ)ϕ′, θ′←τθ + (1 −τ)θ′,(19) where 0< τ ≪1is the target update rate. While several studies have proposed sophisticated interdecision communication mechanisms to improve the coherence of VNF placements, these approaches often introduce significant complexity, complicating both optimization and implementation. In contrast, the use of Transformers offers an elegant and well-validated solution for achieving coordination without unnecessary complexity. Training Process and Objectives. The training process adopts the standard actor-critic framework, which uses separate loss functions for the actor and critic networks. The complete training steps are given in Alg. 1 and 2. •Critic Loss. The critic is trained to minimize the Mean Squared Error (MSE) between the predicted Q-value Qϕ(s, a)and the target Q-value y. The target Q-value is computed using the Bellman equation: y=r+γQϕ′(s′, πθ′(s′)),(20) where ris the immediate reward, γis the discount factor, and Qϕ′and πθ′are the target critic and actor networks, respectively. The critic loss is defined as: Lcritic =Eh(Qϕ(s, a)−y)2i.(21) •Actor Loss. The actor is trained to maximize the Q-value of the actions it generates, as evaluated by the critic. This is equivalent to minimizing the negative expected Q-value: Lactor =−E[Qϕ(s, πθ(s))] .(22) By leveraging the critic’s evaluation, the loss function guides the actor to update its policy in the direction that increases the Q-value, ultimately improving decisionmaking performance. Continuous Action Representation. In the proposed framework, the decision to assign each VNF to a specific target DC is inherently a discrete action, which typically requires stochastic sampling, which is non-differentiable. To enable E2E training using gradient-based optimization, we employ a softmax layer on the actor’s output to represent actions. The softmax function serves as a smooth and differentiable approximation of the argmax operation [29], generating a relaxed one-hot vector that encodes the selected action. This approach effectively transforms the discrete decisionmaking problem into a continuous one, avoiding the need for stochastic sampling processes typically required for discrete action spaces in DRL, allowing gradients to propagate directly through the softmax layer during backpropagation. This continuous approximation bridges the gap between continuous and discrete actions, enabling fully differentiable hybrid policy optimization. Particularly, in formula (18), the input Hto the critic, formed by a concatenation of continuous states and relaxed discrete actions, enables seamless gradient flow through both the actor and critic networks. In addition, because the joint action vector is optimized in a single differentiable step, it inherently avoids the credit assignment problem that commonly arises in MARL settings. In general, this technique not only simplifies the training process, but also potentially accelerates convergence by enabling E2E differentiability through continuous representations of discrete actions. Epsilon-Logit Perturbation (ϵ-LoPe). Since discrete actions are encoded as continuous representations, the traditional ϵ-greedy exploration strategy becomes inapplicable. To this end, we introduce ϵ-LoPe, a noise injection strategy that facilitates exploration. Instead of selecting a random action with probability ϵ,ϵ-LoPe perturbs the logits before applying softmax, imposing the exploration at the representation level rather than in the final actions. Specifically, at each decision step, a Gaussian noise η∼ N(0,1) is scaled by the decaying perturbation factor ϵ∈[0,1] and added to the logits zt: ˜zt=zt+ϵ·η, (23) where ztrepresents the pre-softmax logits generated by the actor network. The perturbed logits ˜ztare then passed through a softmax function to produce the final action: at=softmax(˜zt).(24) The perturbation factor ϵdecreases over episodes, encouraging high exploration in early training while gradually moving toward deterministic policy as training progresses. To regulate the impact of noise injection, we apply z-score normalization to the logits ztacross features for each sample before adding the noise. This normalization ensures that the scale of logits 16 •Average Resource Utilization. The average resource utilization metric provides an overall measure of how effectively the computational resources across all DCs are used during peak hours. For a set of DCs VS, the average resource utilization is defined as: 1 |VS|X u∈VS Resources Used at DC u Total Resources at DC u.(36) A higher value indicates better overall utilization of available resources, while a lower value suggests potential under-utilization. •Perplexity of Load Distribution. The perplexity of resource utilization quantifies the distribution of resource usage across all DCs during peak hours. It reflects the degree of balance in resource allocation, where higher perplexity implies a more uniform distribution and lower perplexity indicates uneven utilization. For a set of DCs VS, the perplexity is computed as: exp −X u∈VS pulog(pu)!,(37) where puis the proportion of total utilized resources at DC u, given by: pu=Resources Used at DC u Pv∈VSResources Used at DC v.(38) A higher perplexity value suggests that resource usage is evenly distributed across DCs, reducing the risk of overloading specific DCs, while a lower value indicates concentration of resource usage in a few DCs. Fig. 9: Average CPU utilization over peak hours. Fig. 9 shows the average resource utilization over time, which highlights the superior acceptance rate achieved by SDAC, with seqDDQN and paraDDQN following closely behind. This trend is consistent with the behavior observed in Fig. 5. As shown in the figure, SDAC consistently achieves the highest CPU utilization, peaking above 0.91. This indicates that SDAC makes the most effective use of available computational resources, which contributes directly to its high acceptance rates. Its ability to efficiently pack services onto infrastructure under heavy load reflects strong decisionmaking and orchestration capabilities. seqDDQN and paraDDQN also exhibit high CPU utilization, though slightly lower than SDAC. seqDDQN maintains a slight advantage over paraDDQN across the time steps, consistent with its more coordinated decision-making framework. RAILS shows moderate utilization, gradually improving toward the end of the interval, but still trailing behind the DRL-based approaches. On the other end of the spectrum, ILS and GP demonstrate the lowest CPU utilization. GP, in particular, maintains a consistently conservative resource usage pattern, suggesting a more cautious allocation strategy. While this may help avoid overloading the system, it results in significant underutilization of available resources and ultimately limits acceptance performance. Overall, the CPU utilization trends during peak hours strongly correlate with the algorithms’ effectiveness, confirming SDAC’s superior ability to adapt under pressure. Fig. 10: Perplexity of load distribution over peak hours. Fig. 10 illustrates the perplexity of load distribution over time, highlighting the balance of resource allocation across all DCs. The DRL-based methods maintain relatively high perplexity, indicating their ability to distribute resource usage evenly. In contrast, GP, ILS, and RAILS exhibit consistently lower perplexity, suggesting a more conservative approach that concentrates resource usage in fewer DCs. In particular, SDAC sustains high perplexity throughout the timeline and remains unaffected by noticeable drops, demonstrating its robustness in optimizing resource allocation over time. E. Asymptotic Return Normalization To assess the impact of different normalization techniques on training stability and performance, we evaluate SDAC under four strategies: the proposed ARN, without normalization (w/o Norm), gradient normalization (grad Norm), and z-score normalization (z-score Norm), respectively. Fig. 11 demonstrates the impact of different return normalization strategies on SDAC’s learning performance. The proposed ARN consistently leads to faster convergence and achieves the highest final rewards. While z-score normalization and training without normalization also reach competitive performance, their curves exhibit noticeable fluctuations even after convergence, indicating less stable policy updates. This instability is possibly due to increasingly large Q-values as training progresses, which can result in larger gradients and unstable updates. Gradient normalization performs less effectively overall, showing slower learning and lower final rewards. These results highlight the advantage of ARN in promoting both efficient learning and post-convergence stability. 17 Fig. 11: Effect of various normalization techniques on SDAC. VII. CONCLUSION & FUTURE WORK In this work, we proposed a novel Transformer-empowered Actor-Critic Reinforcement Learning framework for sequenceaware Service Function Chain (SFC) partitioning in multidomain network orchestration. By integrating Transformer encoders within an Actor-Critic architecture, our method leverages the self-attention mechanism to model inter-dependencies among Virtual Network Functions (VNFs), enabling coordinated and parallel decision-making. The introduction of the ϵ-LoPe exploration strategy and Asymptotic Return Normalization (ARN) further enhances training stability and convergence, addressing the limitations of existing sequential and parallel VNF partitioning methods. Additionally, the innovative critic network, inspired by sentence encoders in Large Language Models (LLMs), enables holistic evaluation of VNF assignments, capturing contextual relationships within SFCs. Comprehensive evaluations through extensive simulations demonstrate that our method outperforms state-of-the-art methods, including both Deep Reinforcement Learning (DRL)-based and (meta-)heuristic-based approaches. The proposed method achieves higher long-term acceptance rates, improved resource utilization and scalability while maintaining efficient inference times. Furthermore, the proposed ARN technique ensures training stability, outperforming other normalization strategies in terms of convergence speed and final performance. These results validate the effectiveness of our method in optimizing SFC partitioning under the complexities of multi-domain networks, making it a robust solution for beyond 5G and 6G network environments. The promising results of our method pave the way for several future research directions. First, validating and extending the proposed method using real-world network data is a critical step toward practical deployment. To address potential limitations in data availability, advanced generative AI techniques could be employed to synthesize realistic network traffic, thus enriching the training and evaluation process. Second, exploring the integration of our method with emerging technologies, such as digital twins or in-band telemetry, could enable more precise resource allocation and proactive network management by providing real-time insights into network states. Third, investigating a decentralized scheme for the proposed method to enhance scalability and robustness in multi-domain settings. Specifically, designing a communication framework that allows limited information exchange tailored for attention mechanisms could facilitate distributed decision-making across multiple locations. This decentralized scheme is expected to reduce latency, enhance fault tolerance, and effectively handle imperfect information inherent in multidomain network environments. ACKNOWLEDGMENT This research was partially funded by the HORIZON SNS JU DESIRE6G project (grant no. 101096466) and the Dutch 6G flagship project “Future Network Services”. REFERENCES [1] G. Sun, Y. Li, D. Liao, and V. Chang, “Service function chain orchestration across multiple domains: A full mesh aggregation approach,” IEEE Transactions on Network and Service Management, vol. 15, no. 3, pp. 1175–1191, 2018. [2] D. Bhamare, R. Jain, M. Samaka, and A. Erbad, “A survey on service function chaining,” Journal of Network and Computer Applications, vol. 75, pp. 138–155, 2016. [3] Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, pp. 436–444, 2015. [4] V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al., “Human-level control through deep reinforcement learning,” nature, vol. 518, no. 7540, pp. 529–533, 2015. [5] D. Silver et al., “Mastering the game of Go without human knowledge,” Nature, vol. 550, pp. 354–359, 10 2017. [6] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 5998–6008. [7] P. T. A. Quang, A. Bradai, K. D. Singh, G. Picard, and R. Riggio, “Single and multi-domain adaptive allocation algorithms for VNF forwarding graph embedding,” IEEE Transactions on Network and Service Management, vol. 16, no. 1, pp. 98–112, 2019. [8] S. Sahhaf, W. Tavernier, M. Rost, S. Schmid, D. Colle, M. Pickavet, and P. Demeester, “Network service chaining with optimized network function embedding supporting service decompositions,” Computer Networks, vol. 93, pp. 492–505, 2015, cloud Networking and Communications II. [9] C. Papagianni, A. Leivadeas, S. Papavassiliou, V. Maglaris, C. CervelloPastor, and A. Monje, “On the optimal allocation of virtual resources in cloud computing networks,” IEEE Transactions on Computers, vol. 62, no. 6, pp. 1060–1071, 2013. [10] D. Dietrich, A. Abujoda, A. Rizk, and P. Papadimitriou, “Multi-provider service chain embedding with Nestor,” IEEE Transactions on Network and Service Management, vol. 14, no. 1, pp. 91–105, 2017. [11] A. Leivadeas, C. Papagianni, and S. Papavassiliou, “Efficient resource mapping framework over networked clouds via iterated local searchbased request partitioning,” IEEE Transactions on Parallel and Distributed Systems, vol. 24, no. 6, pp. 1077–1086, 2013. [12] A. Leivadeas, G. Kesidis, M. Falkner, and I. Lambadaris, “A graph partitioning game theoretical approach for the VNF service chaining problem,” IEEE Transactions on Network and Service Management, vol. 14, no. 4, pp. 890–903, 2017. [13] J. Li, W. Shi, H. Wu, S. Zhang, and X. Shen, “Cost-aware dynamic SFC mapping and scheduling in SDN/NFV-enabled space–air–groundintegrated networks for internet of vehicles,” IEEE Internet of Things Journal, vol. 9, no. 8, pp. 5824–5838, 2022. [14] G. L. Santos, T. Lynn, J. Kelner, and P. T. Endo, “Availability-aware and energy-aware dynamic SFC placement using reinforcement learning,” The Journal of Supercomputing, vol. 77, no. 11, pp. 12 711–12 740, Nov 2021. [15] A. Pentelas, D. De Vleeschauwer, C.-Y. Chang, K. De Schepper, and P. Papadimitriou, “Deep multi-agent reinforcement learning with minimal cross-agent communication for SFC partitioning,” IEEE Access, vol. 11, pp. 40 384–40 398, 2023. [16] D. Heo, S. Lange, H.-G. Kim, and H. Choi, “Graph neural network based service function chaining for automatic network control,” in 2020 21st Asia-Pacific Network Operations and Management Symposium (APNOMS), 2020, pp. 7–12. 18 [17] A. Dalgkitsis, L. A. Garrido, P.-V. Mekikis, K. Ramantas, L. Alonso, and C. Verikoukis, “SCHEMA: Service chain elastic management with distributed reinforcement learning,” in 2021 IEEE Global Communications Conference (GLOBECOM), 2021, pp. 01–06. [18] A. Dalgkitsis, L. A. Garrido, F. Rezazadeh, H. Chergui, K. Ramantas, J. S. Vardakas, and C. Verikoukis, “SCHE2MA: Scalable, energyaware, multidomain orchestration for Beyond-5G URLLC services,” IEEE Transactions on Intelligent Transportation Systems, pp. 1–11, 2022. [19] S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,” ACM computing surveys (CSUR), vol. 54, no. 10s, pp. 1–41, 2022. [20] A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al., “Rt-1: Robotics transformer for real-world control at scale,” arXiv preprint arXiv:2212.06817, 2022. [21] S. Islam, H. Elmekki, A. Elsebai, J. Bentahar, N. Drawel, G. Rjoub, and W. Pedrycz, “A comprehensive survey on applications of transformers for deep learning tasks,” Expert Systems with Applications, vol. 241, p. 122666, 2024. [22] Y. Wang, Z. Gao, D. Zheng, S. Chen, D. G¨ und¨ uz, and H. V. Poor, “Transformer-empowered 6G intelligent networks: From massive mimo processing to semantic communication,” IEEE Wireless Communications, vol. 30, no. 6, pp. 127–135, 2022. [23] J. Cai, F. Gan, X. Cao, and W. Liu, “Signal modulation classification based on the transformer network,” IEEE Transactions on Cognitive Communications and Networking, vol. 8, no. 3, pp. 1348–1357, 2022. [24] A. Hameed, J. Violos, A. Leivadeas, N. Santi, R. Gr¨ unblatt, and N. Mitton, “Toward qos prediction based on temporal transformers for iot applications,” IEEE Transactions on Network and Service Management, vol. 19, no. 4, pp. 4010–4027, 2022. [25] A. D. Raha, K. Kim, A. Adhikary, M. Gain, Z. Han, and C. S. Hong, “Advancing ultra-reliable 6G: Transformer and semantic localization empowered robust beamforming in millimeter-wave communications,” arXiv preprint arXiv:2406.02000, 2024. [26] M. Rost and S. Schmid, “On the hardness and inapproximability of virtual network embeddings,” IEEE/ACM Transactions on Networking, vol. 28, no. 2, pp. 791–803, 2020. [27] R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T. Liu, “On layer normalization in the transformer architecture,” in International conference on machine learning. PMLR, 2020, pp. 10 524–10 533. [28] D. Cer, Y. Yang, S.-y. Kong, N. Hua, N. Limtiaco, R. S. John, N. Constant, M. Guajardo-Cespedes, S. Yuan, C. Tar, et al., “Universal sentence encoder,” arXiv preprint arXiv:1803.11175, 2018. [29] E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with gumbel-softmax,” in International Conference on Learning Representations, 2017. [30] V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, “Playing Atari with deep reinforcement learning,” arXiv preprint arXiv:1312.5602, 2013. [31] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al., “Language models are few-shot learners,” Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020. [32] A. P. Steiner, A. Kolesnikov, X. Zhai, R. Wightman, J. Uszkoreit, and L. Beyer, “How to train your ViT? Data, augmentation, and regularization in vision transformers,” Transactions on Machine Learning Research, 2022. [33] C. S.-H. Hsu, C. Papagianni, and P. Grosso, “RAILS: Risk-aware iterated local search for joint SLA decomposition and service provider management in multi-domain networks,” in 2025 IEEE 26th International Conference on High Performance Switching and Routing (HPSR), 2025, to appear. [34] H. R. Lourenc¸o, O. C. Martin, and T. St¨ utzle, Iterated Local Search. Boston, MA: Springer US, 2003, pp. 320–353. [35] L. Matignon, G. J. Laurent, and N. Le Fort-Piat, “Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems,” The Knowledge Engineering Review, vol. 27, no. 1, pp. 1–31, 2012. [36] D. Hendrycks and K. Gimpel, “Bridging nonlinearities and stochastic regularizers with gaussian error linear units,” 2017. [37] K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14. Springer, 2016, pp. 630–645. [38] I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations, 2019. Cyril Shih-Huan Hsu is a Ph.D. candidate at the Informatics Institute, University of Amsterdam (UvA), The Netherlands, where he has been pursuing his degree since 2021. He earned his B.Sc. and M.Sc. degrees from National Taiwan University (NTU) in 2013 and 2015, respectively. From 2016 to 2021, he worked as a machine learning researcher at several international AI startups. He is currently a member of the Multiscale Networked Systems (MNS) Group. His recent research centers on leveraging AI and machine learning for network resource management. Anestis Dalgkitsis is a Postdoctoral Researcher at the University of Amsterdam’s Multi-Scale Networked Systems group. With a PhD in Signal Theory and Telecommunications from Universitat Polit` ecnica de Catalunya, he specializes in decentralized network orchestration, AI-driven automation, and programmable data planes. Anestis has a strong background in 5G/6G research, multi-agent systems, and high-performance network experimentation, with his work earning recognition in international competitions and conferences. His career includes roles in prestigious EU projects and collaborations with global industry and academic partners. Paola Grosso (Member, IEEE and ACM) is a Full Professor at the University of Amsterdam with a chair on ”Multiscale networks”. She is part of the Multiscale Networked Systems research group (mns-research.nl) and director of the Informatics Institute (ivi.uva.nl). Her research focuses on the creation of sustainable and secure computing infrastructures, which rely on the provisioning and design of programmable networks. Current interests cover developments of control planes for quantum networks, in-band telemetry and use of FPGA for enhanced network monitoring and adoption of digital twins to support enhanced network operations. She has an extensive list of publications on these topics (https://www.uva.nl/profiel/g/r/p.grosso/p.grosso.html) and currently contributes to several national and international projects in this area. Chrysa Papagianni is an Associate Professor at the Informatics Institute of the University of Amsterdam. She is part of the Multiscale Networked Systems group that focuses its research on network programmability and data-centric automation. Her research interests lie primarily in the area of programmable networks with an emphasis on network optimization and the use of machine learning in networking. She has worked on multiple research projects funded by the European Commission, the Dutch research council and the European Space Agency, such as 5Growth, CATRIN, etc. Chrysa is currently coordinating the SNS JU DESIRE6G project on 6G system architecture. She also leads the AI-assisted networking work package in the 6G flagship project for the Netherlands on Future Network Services.