IEEE TRANSACTIONS ON NETWORK AND SERVICE MANAGEMENT 1 The TES framework: Joint Statistical Modeling and Machine Learning for Network KPI Forecasting Leonardo Lo Schiavo, Genoveva Garcia, Marco Gramaglia, Marco Fiore, Albert Banchs, Xavier Costa-Perez Abstract—The vision of intelligent networks capable of automatically configuring crucial parameters for tasks such as resource provisioning, anomaly detection or load balancing largely hinges upon efficient AI-based algorithms. Time series forecasting is a fundamental building block for network-oriented AI and current trends lean towards the systematic adoption of models based on deep learning approaches. In this paper, we pave the way for a different strategy for the design of predictors for mobile network environments, and we propose the Thresholded Exponential Smoothing (TES) framework, a hybrid Statistical Modeling and Deep Learning tool that allows for improving the performance of network Key Performance Indicator (KPI) forecasting. We adapt our framework to two state-of-the-art deep learning tools for time series forecasting, based on Recurrent Neural Networks and Transformer architectures. We experiment with TES by showcasing its superior support for three practical network management use cases, i.e., (i) anticipatory allocation of network resources, (ii) mobile traffic anomaly prediction, and (iii) mobile traffic load balancing. Our results, derived from traffic measurements collected in operational mobile networks, demonstrate that the TES framework can yield substantial performance gains over current state-of-the-art predictors in the applications considered. Index Terms—Forecasting, prediction, mobile traffic, network KPI, network management, neural networks, statistical modeling. ©2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. I. INTRODUCTION EVERY new generation of mobile networks invariably raises the bar for the performance, reliability, and security of cellular communication systems. Adhering to such a trend, 6G systems are expected to support diverse classes of services and do so with near-zero latency, apparent infinite capacity, and 100% availability, making de-facto the communication infrastructure fully transparent to applications [1] and turning 6G networks into general-purpose platforms providing smart connectivity to a plethora of very heterogeneous terminals. While today’s mobile communication infrastructures are already extremely tangled architectures that entail significant challenges in terms of equipment management, traffic engineering, and capacity allocation [2], 6G systems will introduce several layers of substantial additional complexity [3], [4]. Indeed, meeting the ambitious 6G performance targets will require instant orchestration of physical resources and Virtual L. Lo Schiavo, G. Garcia, M. Gramaglia, and A. Banchs are with University Carlos III de Madrid. Corresponding email: [email protected], genovev[email protected], [email protected], [email protected]. M. Fiore and A. Banchs are with IMDEA Networks Institute. Corresponding email: [email protected], [email protected]g. X. Costa-Perez is with NEC Laboratories Europe GmbH, I2CAT, and ICREA. Corresponding e-mail:
[email protected]. Network Functions (VNFs) across different network domains, in concert with user demands and multi-tenancy requirements that rapidly shift in time. Machine Learning (ML) and Artificial Intelligence (AI) are largely regarded as fundamental enablers to realize such a vision. Integrating AI/ML solutions, supported by a native network architecture [5], will pave the road towards the efficient support of various use cases that dramatically enhance the performance of next-generation systems. Data-driven models have been repeatedly shown to offer enhanced quality for key network management tasks such as anomaly detection [6], traffic classification [7], resource orchestration [8], radio access operation [9], and energy saving [10], just to name a few. In many of those tasks, anticipatory decision-making is a very desirable –if not mandatory– feature, making prediction an essential building block to AI/ML-driven network management [11]. In this context, a plethora of works have proposed ever more accurate forecasting models [12], [13] and recent works have also shown how predictors can be tailored to the downstream network management task by steering their output [8], [14]. In this work, we focus on forecasting network Key Performance Indicators (KPIs), such as traffic demands or user throughput, as one of the cornerstones of future zero-touch network management [15]. While traffic prediction was carried out via statistical models until the first decade of the century [16], [17], AI/ML solutions have nowadays taken a clear lead and dominate the literature [18]–[21]. We depart from the common practice to propose pure AI/ML-based models and explore hybrid approaches where traditional statistical modeling is combined with deep learning [22]. This hybrid strategy is known to yield resilience to noisy data and wide excursions in time series of financial data or weather fluctuations [23], and we show that it can also help achieve higher prediction accuracy in the presence of real-world mobile traffic data, thus benefiting network management tasks that build upon such forecasts. By pioneering the adoption of hybrid predictors using statistical modeling and ML in the context of anticipatory network management, our work yields the following main contributions. •We introduce for the first time a hybrid model combining exponential smoothing (ES) with different deep learning models based on Recurrent Neural Network (RNN) and Transformer architectures, and demonstrate how it improves quality of the downstream anticipatory network management tasks, with improvements in the 4%-26%, 20%-40% and 6%-16% range depending on the scenario.
IEEE TRANSACTIONS ON NETWORK AND SERVICE MANAGEMENT 2 •We update the operation of the state-of-the-art ES-RNN architecture to cope with unique features of mobile traffic dynamics; the result is an original Thresholded ES-RNN (TES-RNN) model, i.e., a general-purpose network traffic forecasting technique that can be tailored to perform predictions for different network management functions. •We apply the same methodology to the Transformer neural network architecture, understanding the benefits and trade-offs of the different approaches. •We apply the proposed models to three practical zerotouch management use cases, i.e., (i) capacity allocation, (ii) anomaly detection, and (iii) load balancing, for which we train the models with appropriate loss functions. •We evaluate the performance of our solutions against recent works in the literature, demonstrating in all use cases above its superior performance with respect to stateof-the-art Deep Neural Network (DNN) architectures. This provides a substantial update with respect to our previous work in [24], as we discuss the usage of Transformer models in this context (to our knowledge, the first attempt to integrate statistical modeling with this architecture), improve the threshold selection of the TES framework through an automatic and generalizable Reinforcement Learning (RL) algorithm, and add the load balancing use case to further showcase the adaptability of our solution. The paper is structured as follows: we detail the context of the zero-touch network management and related work in time series forecasting for networks in Section II. We discuss the application of hybrid strategies that combine ML and statistical modeling for forecasting in Section III and present our proposed TES framework in Section IV. Finally, we analyze a set of relevant use cases in Section V and their performance evaluation in Section VI, before concluding in Section VII. II. RELATED WORK Relying on a precise time-series forecasting algorithm is a fundamental building block for many autonomous network management and operation solutions [25]. Indeed, the quality of the prediction plays a role in the overall performance of the autonomous network management algorithm: with a more precise forecast, the decision taken can guarantee a better outcome. In the following, we revise the state-of-the-art solutions for autonomous and zero-touch network management, with a focus on anticipatory networking. Finally, we explore the works in the field of joint statistical modeling and ML, which is the solution we adopt in this paper to improve the forecasting quality of pure DNN models. Network Intelligence for zero-touch management. Handling the escalating complexity of Beyond 5G (B5G) networks with traditional human-in-the-loop approaches will not be possible anymore. Instead, it is expected that current management models will be replaced by zero-touch network and service management technologies, which fully automate the network operation and are presently being standardized [26]. As a result of this transition, the success of B5G will vastly depend on the quality of the Network Intelligence (NI) that will run at schedulers, controllers, and orchestrators across network domains, de-facto managing the zero-touch infrastructure. Following a popular trend in many research and engineering domains, AI models relying on DNN architectures are regarded as a promising approach for the design of NI solutions. Indeed, AI models have proven remarkably effective at solving complex network operation tasks, and they thrive on the large amount of control and traffic data available within network architectures [27]. Forecasting for anticipatory networking. Many NI solutions build upon anticipatory networking principles and aim at proactively optimizing network configurations with respect to upcoming traffic conditions rather than to the current state [28]. The prominence of anticipatory NI makes predicting future network states a fundamental task for the effective operation of B5G systems. Forecasting is in fact a manifold problem in networking environments, where different applications require accurate future projections of diverse metrics, including computational resources [29], capacity requirements [14], or sheer traffic volumes [12], possibly separated by mobile service [13]. Similarly to what happens for other aspects of NI design, DNN models have lately been established as the prevailing approach to developing the predictors that will support proactive decisions by NI solutions. In the past few years, a fairly large body of works has explored varied DNN architectures, which target diverse forecasting objectives, and are typically proven to yield improved accuracy over legacy statistical models. Joint statistical modeling and DNN. While current state-ofthe-art predictors in the networking domain invariably rely on deep learning, very recent results from the ML community suggest that hybrid engines that integrate statistical modeling and DNN can, in fact, substantially outperform pure DNN approaches in time series forecasting tasks. The very first model of this kind comfortably won the renowned M4 Competition, a challenge for data scientists to develop ever more accurate time series predictors [23]. It did so by beating a variety of statistical and ML benchmarks, as well as 48 competitor solutions, across 100,000 experiments. The aforementioned engine combines a classical ES statistical model with a RNN architecture, hence it is named ESRNN [22]. It is a true hybrid predictor since the parameters of the ES model are optimized concurrently with the RNN weights using unified gradient descent. Thanks to this joint training, the ES-RNN model represents a leap forward with respect to previous attempts at mixing different statistical and/or ML methods: unlike simple combination [30] or ensemble [31] strategies used to date, this technique takes full advantage of the strengths of statistical and ML methods, while mitigating their respective limitations. III. HYBRID NETWORK KPI PREDICTION WITH MACHINE LEARNING AND STATISTICAL MODELING The hybrid prediction approach proposed in this paper builds upon the innovative design principles first introduced by the recent ES-RNN engine [22], which is presented in Section III-A. The considered ES-RNN predictor has limitations
IEEE TRANSACTIONS ON NETWORK AND SERVICE MANAGEMENT 3 when confronted with real-world mobile traffic dynamics, as discussed in Section III-B. Our proposed hybrid methodology enhances the structure proposed in [22] and solves such issues by enhancing the original engine with an automatically learned threshold parameter, as detailed in Section IV-B. A. ES-RNN and joint SGD optimization ES-RNN is a truly hybrid forecasting model for time series that mixes statistical modeling, i.e., ES, and ML, i.e., RNN. We consider the GPU implementation of ES-RNN [32] as the basis for our study: this variant presents a first pre-processing layer for adaptive and local normalization of input time series using ES formulas, followed by a neural network architecture that processes the normalized data and provides forecasts over a customizable time horizon. The original ES-RNN may adopt a variety of ES expressions, depending on the temporal features of the target data. In networking settings, 24-hour circadian rhythms are known to dominate the fluctuations of mobile data traffic [33], hence we opted for a Holt linear non-seasonal ES formula [34], which is the recommended expression for time series with daily periodicity [22]. At each time step t, the non-seasonal ES updates a normalization coefficient lt(called level) as lt=ωyt+ (1 −ω)lt−1,(1) where ω∈[0,1] is the exponential smoothing parameter, and ytrepresents the value of the input time series at time step t. The level ltis used for data normalization. At a given time step t, all values in the input window [t-tI,t] of size Iand in the output interval [t+1, t+tO] of size Oare divided by lt. During training, the normalized input window is fed to the RNN, whose (normalized) forecast is compared with the normalized output window using a loss function. In testing, or when running the model in production systems, de-normalization is performed by multiplying the normalized values forecasted in the prediction horizon Oby the level lt. The major novelty of the ES-RNN model is that the smoothing parameter ωis treated as a system variable that is learned together with the weights of the subsequent RNN architecture. In other words, the stochastic gradient descent (SGD) process, normally used to fit the RNN weights, backpropagates in this case before the neural network input layer, and into the preceding ES model, where it updates ω. In this way, a single SGD allows for jointly optimizing the parameters of the statistical model and the neural network, adapting them all to the characteristics of the target time series. The SGD optimization of ωoperated by ES-RNN results in a level ltthat is dynamically adapted to the input data. In turn, this enables a so-called local and adaptive normalization, which (i) ensures that all portions of the time series are equally important to the ensuing neural network training process, and (ii) suitably smooths the ML input so that the neural network can concentrate on predicting actual trends, without overfitting on spurious patterns [22]. Thus, this normalization helps forecast time series with severe fluctuations, like those observed in mobile networks. This is not the case with traditional global normalization of all values to the same [0,1] interval, which 0246810 Time [hours] 0 2 4 6 8 10 Normalized Load Real (a) 0246810 Time [hours] 0 20 40 60 80 Normalized Load Real ES-RNN TES-RNN (b) Fig. 1: Example of problematic prediction of real-world mobile traffic by ES-RNN. (a) Instagram demand at one BS. (b) The same demand is compared to the prediction generated by ESRNN trained with a Mean Squared Error (MSE) loss function, using an input window of size I= 6 and an output window of size O= 1. does not yield input smoothing and makes it hard for the RNN to learn to predict small values. B. Limitations of hybrid predictors with network KPIs The ES-RNN model is intended to operate on a time series with strictly positive values of comparable magnitude. However, this assumption is often violated in the mobile networking context, where KPIs observed at the radio access and edge network elements are highly irregular and bursty, with continued inactivity periods that lead to a possibly significant presence of zero or near-zero values and severe underutilization of the network. This consideration holds for both voice [35] and data [33] traffic, especially when predictions target demands generated by individual users or at single base stations. These characteristics of mobile traffic dynamics determine levels ltcomputed with (1) that are at times equal to zero, or close to that value. In the case of zero-level values, ES normalization is simply not possible, as it would involve a division by zero. In the case of values close to zero, value discontinuities between the input and output windows yield normalized outputs that are not numerically comparable with (and in fact much higher than) the values predicted by the neural network; the loss function returns then inflated costs that hinder the quality of the learning process. Figure 1 illustrates the latter problem in a practical scenario. Plot (a) portrays the real-world demand generated by Instagram at one base station for several hours: the inconsistent nature of the traffic, with a long period of very low or no activity, is evident. Plot (b) shows how, when a traffic peak occurs after such a sequence of low-traffic time steps, the network starts predicting amplified values largely above the real traffic demand. IV. THE TES FRAMEWORK Motivated by the analysis performed in Section III, we propose a hybrid methodology for the forecasting of time series for network management purposes. The methodology is composed of (i) a module that introduces the statistical
IEEE TRANSACTIONS ON NETWORK AND SERVICE MANAGEMENT 4 Normalized past traffic 𝜆 (t),…, 𝜆 (t - 𝑡𝐼 ) Traffic forecast θ (t + 1) ,…, θ (t + 𝑡𝑂 ) Holt linear non-seasonal formula 𝑙𝑡= max{τ, ω𝑦𝑡+ 1 − ω 𝑙𝑡−1} Local and adaptive normalization Past traffic Λ (t),…, Λ (t - 𝑡𝐼 ) 𝜆 (t) =Λ (t) 𝑙𝑡 … 𝜆 (t − 𝑡𝐼 ) =Λ (t − 𝑡𝐼 ) 𝑙𝑡 Thresholded Exponential Smoothing Forecasting Deep Learning Model Loss function Automatic Threshold τ setting Hybrid ES and model training via single backpropagation RL 𝜏 threshold hyperparameter configuration Custom expression based on target task θ (t) – 𝑦𝑡 F DRNN Layer 1 2 LSTM Cells with dilations (1,3), and state size 50 DRNN Layer 2 2 LSTM Cells with dilations (6,12), and state size 50 Non-linear (tanh) Layer, state size 50x50 … O RNN Transformer or RNN Transformer O Positional Encoding Encoder Layer Transformer Encoder State: μ, σ of timeseries Λ Reward: inverse of loss function Actor Critic TD error Entropy Action: τ Decoder Layer Fig. 2: The architecture. The traffic Λobserved over the past Itime steps is input to the TES component for a local and adaptive normalization of non-negligible traffic above τ. The resulting traffic λis fed to the model of choice, which outputs a forecast θof traffic within a horizon of Otime steps. During training, the loss computed from θis used to learn the threshold hyperparameter τof the TES block in an AutoML style via a Reinforcement Learning algorithm. modeling part through ES, and (ii) a dynamic thresholding (T) module that counters the problem discussed in Section III-B, building hence our TES framework. The overall framework is depicted in Figure 2. A. Neural network predictors One of the goals of our work is to demonstrate the wide applicability of the TES approach, independently of the model used for the actual forecasting of the time series. We consider hence two models for time forecasting: the Recurrent Neural Network already used in [36] and a Transformer architecture [37], whose capabilities in time series forecasting have been recently studied. 1) RNN: RNNs are a class of neural networks designed for processing sequential data, making them particularly wellsuited for time series forecasting. Unlike traditional feedforward neural networks, RNNs possess an internal state that allows them to retain information from previous inputs, enabling the modeling of temporal dependencies. This ability to maintain a memory of past observations makes RNNs effective in capturing patterns and trends over time, a very good feature for accurate time series prediction. 2) Transformer: Initially introduced for natural language processing tasks, the transformer architecture [38], [39] has been used for time series forecasting [40]–[42] due to its ability to handle long-range dependencies and parallelize computation. Unlike RNNs, Transformers do not rely on sequential data processing; instead, they employ self-attention mechanisms to weigh the importance of different time steps, capturing complex patterns and relationships within the data. This architecture consists of encoder and decoder layers, which enable the model to focus on relevant parts of the input sequence dynamically. The parallel processing capability significantly accelerates training and inference times, making Transformers suitable for large-scale time series datasets. We corroborated this fact in a set of training experiments that we executed for the experimental evaluation detailed in Section VI, where we compare the time elapsed for training and the retained loss on the same training data, for the RNN and Transformer architectures. Pure transformer-based solutions could be trained generally one order of magnitude faster (less than 1s vs. tens of seconds) than RNN solutions. The advantages in training time that the Transformer architecture has are beneficial for the automatic thresholding feature that we propose in this work. See Table IV for more details. B. Improved statistical modeling As introduced in Section III-B, deep learning models suffer from noisy data with zeros. For these reasons, we introduce
IEEE TRANSACTIONS ON NETWORK AND SERVICE MANAGEMENT 5 TABLE I: Training time over 20 epochs and normalized MSE of forecasting models predicting Facebook time series. Model Training time (min) Normalized MSE RNN 2.85 0.023 Transformer 0.17 0.018 Autoformer 2.05 0.008 Informer 2.83 0.007 N-Beats 1.61 0.126 D-Linear 1.07 0.140 a thresholded version of the two models discussed below, following the framework in Figure 2. 1) TES-RNN: To address the shortcomings of the original ES-RNN, we introduce the Thresholded ES-RNN (TES-RNN) model. Our solution employs a threshold τto bound the minimum value of lt, which is then updated at each time step tas lt= max{τ, ωyt+ (1 −ω)lt−1}.(2) The enhancement in (2) is simple yet effective in solving the issues observed for ES-RNN. A representative example is provided in Figure 1b: TES-RNN does not suffer from inflated predictions and correctly anticipates the growing traffic. 2) ES-Transformer and TES-Transformer: To improve the performance of the original Transformer model, we adopt equation 1 to introduce the ES-Transformer model with an unbounded adaptive normalization of the inputs. However, similar shortcomings observed for ES-RNN are also observed for ES-Transformer. Therefore, we introduce a Thresholded ES-Transformer (TES-Transformer) model to address those limitations. TES-Transformer adopts a conditional two-stage normalization scheme: in the first stage, a normalization coefficient ltis computed as in equation 1 to normalize the values in the input window [t-tI,t]. Then, a threshold τis used to scale the maximum value of the normalized input window to get a second-stage normalization coefficient lmax tas lmax t=τ·max [t−tI,t]λ(t).(3) In the second conditional stage, if lmax tis bigger than a guard value δ, then lmax tis used to further normalize the input window to avoid the training artifacts of the original Transformer and ES-Transformer models. C. Effect of ES and TES normalization The effect of the normalization discussed in III-A, attained by applying equations 1, 2 and/or 3 depending on the model, can be observed in Figure 3, which shows training inputs over two consecutive days. While the global normalization scales all the input values to the same [0, 1] interval, the other normalizations yield better input smoothing and ease the prediction task for small values, i.e., low overnight traffic values. The benefits of the latter normalization on the forecasting performance will be discussed in detail in Section VI. Mon Tue Days 0 0.25 0.5 0.75 1 1.25 1.5 1.75 Normalized Input Load Global Normalization ES Normalization (1) TES-RNN Normalization (2) TES-Transformer Normalization (1,3) Fig. 3: Input values with different kinds of normalization. TABLE II: Symmetric Mean Absolute Percentage Error (SMAPE) between stationary and non-stationary training. Service Stationary training Non-Stationary training Facebook 37.44% 36.42% Instagram 67.27% 66.89% Snapchat 69.73% 69.99% D. Comparative analysis and impact of stationarity To select the best internal design for the TES framework, we conducted a preliminary comparison of recent forecasting models based on Transformer and linear architectures, including Informer [37], Autoformer [43], pure Transformer [38], N-BEATS (Neural Basis Expansion Analysis for Time Series) [44], and D-Linear [45]. As shown in Table I, although Informer and Autoformer yielded a slightly lower normalized MSE, the Transformer offered a much shorter training time, a critical parameter given the complexity introduced by the TES framework for, e.g., finding the best τ. Importantly, TES compensates for the modest accuracy gap of the pure Transformer, improving the final performance as will be shown later in Section VI. Another important aspect we take into account while designing the internal model of the TES framework is the impact of the stationarity of the input time series, which has been discussed in the literature [46]–[48] as an important metric for assessing the complexity of forecasting problems. Following these works in the literature, we evaluated the stationarity of the traffic time series using Augmented DickeyFuller (ADF) tests and statistical moments (mean, variance, skewness, and kurtosis). We found two different patterns: daily data showed non-stationarity, with varying statistics and ADF p-values above 0.05, while weekly aggregation revealed stationary behavior, as illustrated in Figure 4. To assess the robustness of our model, we trained the pure Transformer model over temporally ordered weekly data (stationary) and shuffled individual days (non-stationary). The performance analysis, summarized in Table II, yields comparable accuracy in both setups, showing that these models generalize well across time shifts even for the non-stationary case, showing how the model copes well with complex problems. The results of Table I and Table II are obtained using the time series of popular services that will be introduced in Section VI, where we train our models in larger stationary scenarios. The results we presented in this section show that the TES
IEEE TRANSACTIONS ON NETWORK AND SERVICE MANAGEMENT 6 -20 0 20 40 60 80 0.00 0.02 0.04 0.06 0.08 0.10 Density - Mean: 6.02 - Var: 23.34 - Skew: 0.76 - Kurt: 0.48 - ADF (p-value): 2.73e-01 Day 1 -20 0 20 40 60 80 - Mean: 10.09 - Var: 78.81 - Skew: 1.34 - Kurt: 2.96 - ADF (p-value): 3.60e-01 Day 2 -20 0 20 40 60 80 - Mean: 8.50 - Var: 56.68 - Skew: 1.69 - Kurt: 6.69 - ADF (p-value): 2.60e-06 Week 1 -20 0 20 40 60 80 - Mean: 7.95 - Var: 47.05 - Skew: 1.75 - Kurt: 9.35 - ADF (p-value): 1.23e-05 Week 2 Theoretical distribution Actual distribution Traffic Fig. 4: Distribution of Instagram time series over 2 individual days (leftmost plots) and 2 consecutive weeks (rightmost plots). framework is general in nature, without any noticeable bias on the kind of input data (e.g., stationary or not), and can be applied to forecasting models beyond Transformers. E. Parametrization of the hyperparameter τ The result in Figure 1b is not obvious to achieve. In particular, the threshold τis challenging to configure, as it introduces an interesting trade-off. Generally, a threshold closer to the traffic peak ensures higher robustness to the problem of time series discontinuities highlighted above. However, it also triggers a global normalization to level τmore often, raising the issue of model insensitivity to low values below the threshold that the local and adaptive normalization aims at solving. Conversely, thresholds closer to the smallest possible level tend to preserve the desirable properties of the fine-tuned ES normalization but incur more often the issues related to discontinuous data. These problems, which are independent of the actual deep learning model used for forecasting, require an automatic algorithm for the correct setting of the τ. There is no one-size-fits-all solution to the trade-off above, and the best value of τdepends on the nature of the traffic time series that is relevant to the target networking functionality. Therefore, τalso needs to be adjusted to the settings of the considered task. Notably, τis a hyperparameter for the TES models, as it steers the overall system behavior. To ensure a smooth operation, it is highly desirable that the setting of τ does not require human intervention, but is fully automated and generalizable. The setup at hand calls for an Automated Machine Learning (or AutoML) approach, since our goal is to automate the design of complex neural network models [49]. For this task, while in an earlier version of the framework [36] we used a Golden-Section search algorithm based on convex loss functions [50], we now propose a more generalizable approach based on a RL algorithm that automatically selects the best τvalue. We resort to a soft actor-critic deep RL algorithm to maximize an arbitrary reward function while exploring as randomly as possible the space of possible τ (action) values at training time through an entropy component. The critic neural network estimates the effect of selecting a given τvalue for a state s, which captures the nature of the traffic time series and is represented by its mean µand standard deviation σ. Such an effect is estimated using an instantaneous reward function, which is the additive inverse of the prediction loss obtained by forecasting the time series with state susing the selected τ. Leveraging the estimates Environment Actor Critic TD Error Entropy Action 𝝉 State s 𝝁 , σ timeseries Agent Reward Inverse of prediction loss Fig. 5: Architecture of the soft actor-critic algorithm. of the critic for a given state s, the actor outputs the best τ from a discrete action space with 20 possible values in the range [0.05, 1] with step 0.05. The architecture of the soft actor-critic algorithm is depicted in Figure 5. V. APPLICATION USE CASES As discussed in Section I, forecasting future network KPIs is a cornerstone task for many networking problems, including admission control [51], capacity allocation [14], handovers [52] and power management [53], among others, which involve different operation time scales [8]. The TES architecture described in Section IV-B is general-purpose and agnostic to the specific networking application: it can be trained to support different NI instances, e.g., by combining it with a suitable loss function that allows optimizing the model for a given prediction task. Next, we present two practical anticipatory networking use cases where the TES framework can be used as the forecasting model, which also sets the ground for the experimental evaluation conducted in Section VI. A. Use case I: capacity allocation for network slicing A first use case of interest for forecasting in anticipatory networking is that of capacity allocation, i.e., reserving the resources needed to meet the upcoming demand for a given service. This functionality is especially relevant in network slicing settings, where (sets of) services run in different slices,
IEEE TRANSACTIONS ON NETWORK AND SERVICE MANAGEMENT 7 and the operator needs to dedicate sufficient resources to each slice, in agreement with the load generated by the corresponding service(s) [51]. The anticipatory NI in charge of capacity allocation to slices must rely on so-called capacity forecasting, i.e., predicting the minimum capacity sufficient to accommodate the future slice traffic. We highlight that capacity forecasting is a fairly unique problem, where sheer accuracy is not the most relevant metric. Instead, the prediction must stay above the actual load with a very high probability, because underestimation determines the allocation of insufficient capacity to slices, hence service disruption on the user side. Underprovisioning also triggers violations of the Service Level Agreement (SLA) between the slice tenant and the network operator, which thus incurs substantial economic penalties. Clearly, this must be avoided without allocating exceedingly large amounts of unnecessary resources, which also has a cost for the operator. The problem of capacity forecasting has recently received attention, with the proposal of dedicated predictors [8], [14]. These models rely on a loss function that drives the learning process to capture the actual cost of incurring SLA violations against that of overprovisioning the slice capacity. Specifically, the function handles negative and positive errors differently, to reflect the different costs they entail in the context of virtualized communication networks, as follows. •A constant penalty βis associated with each negative error, which causes an SLA violation during the predicted time interval. βcan be customized to the desired behavior: for instance, higher values may be used when reliability is paramount (e.g., for slices serving ultrareliable low-latency communications or URLLC), and lower penalties can be applied for slices with more relaxed requirements. •A monotonically increasing cost is attributed to positive errors, which imply the allocation of excess resources. Therefore, the cost is proportional to the amount of (unnecessarily) provisioned capacity. Typically, the expenditure is assumed to grow linearly with the overprovisioned capacity, with a fixed rate γof cost per surplus capacity. The configuration of the two costs can be, in fact, controlled by a single parameter α=β/γ, which represents the amount of overprovisioned capacity that the operator is willing to deploy to avoid committing an SLA violation. Formally, for a given prediction error x, the loss function that abides by the specifications above is expressed as L(x) = α−ϵ·xif x≤0 α−1 ϵxif 0< x ≤ϵα x−ϵα if x > ϵα, (4) where steep slopes (implemented with a small positive ϵ) ensure differentiability over the whole xdomain [14]. The parameter αserves as a knob to steer the operational point of the system towards higher expenses in deployed resources but reduced chances of SLA violations, or vice-versa. As a result, the loss function in (4) can be parametrized to the specifications of different network infrastructure locations (e.g., reflecting the higher cost of deploying resources at the Time Load Decision time Probability of future anomaly Past load Next-step reference interval Probabilistic next-step forecast Load 0 0.2 0.4 0.6 0.8 1 CDF ΔAΔA 1−PA Next load level Reference Anomalous (real) level Regular (real) level Fig. 6: Anomaly forecasting use case. (left) Representation of the anomaly forecasting problem. (right) Explanation of the operation of anomaly detection. network edge than at the core), resource types (e.g., capturing the fact that radio resources are sensibly more expensive than CPU resources), and SLA strategies (e.g., expressing the higher fees for violations affecting slices of critical services). B. Use case II: anomaly detection in mobile service traffic The second use case we study is an anticipatory anomaly detection framework, where the NI must trigger an alarm when an abnormal future traffic load is expected for a specific mobile service. The anomaly detection problem is summarized in Figure 6a. The predictor module is in charge of producing aprobability distribution of the traffic demand that the target service will generate in the next time slot. Such a probabilistic prediction is compared against a reference interval that encompasses the expected range of normal traffic values in the following time slot. Then, if the probability of the anticipated traffic being outside the reference interval is beyond a threshold, an alarm is raised. This allows the associated NI to perform some preventive actions, such as those detailed in the 3GPP TS 23.288 [54] technical specification under the “Abnormal behavior” analytics, which capture anomalies such as unexpected large rate flows generated by terminals. We consider a simple yet practical implementation1of the approach above that is commonly adopted in many fields, also outside networking [55]. First, it is worth noting that the output of the forecasting algorithm shall not be a scalar but a probability distribution of the future traffic load. This type of output is implicit in certain types of models like Bayesian Neural Networks, which are, however, computationally expensive and not suited for resource-constrained network environments. In order to generate a probabilistic forecast with a generic neural network, we resort to recent findings in uncertainty modeling [56]: specifically, by activating dropout layers in the predictor during the inference phase and performing a Monte Carlo test, the neural network returns a set of values that have been shown to closely approximate the probabilistic result of a deep Gaussian process implemented with a Bayesian network. 1Our goal is not to propose a novel anomaly detection algorithm, but to compare the effectiveness of different forecasting models in supporting such a task. Therefore, we are not interested in developing a complex algorithm for anomaly detection, and using a baseline solution is sufficient for our purpose.
IEEE TRANSACTIONS ON NETWORK AND SERVICE MANAGEMENT 8 The anomaly detection algorithm then operates on the empirical Probability Density Function (PDF) fpof the predicted traffic values for each decision interval, as illustrated in Figure 6b. First, the upper and lower limits that mark the boundaries between regular and anomalous values are computed as xl,h =R±∆A, where Ris a configurable reference value. From these two values, the probability of a future anomaly is empirically calculated as PA= 1 −(Fp(xh)−Fp(xl)), where Fp(x) = Pk<x fp(k)is the Cumulative Distribution Function (CDF) of the anticipated traffic. Finally, an alarm is triggered if PA> τA. The parameters R,∆A, and τA control the sensitivity of the algorithm. In our experiments, we set the reference values for the estimated load in the next prediction step Ras the average of the last three load values, ∆A= 0.9·R, and τA= 0.9. In other words, we trigger an alert when the model forecasts future traffic with a 90% probability to fall outside a range ±90% of the reference value. The correctness of the anticipatory anomaly detection can be determined by checking whether the actual traffic falls into the xl,h interval or not, and computing precision and recall scores. Clearly, a higher accuracy in the probabilistic traffic forecast, denoted by a lower variance around a value closer to the true one, yields better performance: a MSE loss function is thus a sensible choice for this use case. C. Use case III: anticipatory load balancing Load balancing aims at equally splitting the load among different entities and is another networking functionality for which forecasting is critical. Indeed, anticipatory load balancing can be applied at different levels, including the following •Load balancing at edge clouds. In edge cloud facilities, e.g., for Cloud Radio Access Network (C-RAN), it is desirable to associate base stations with data centers in such a way that the future traffic load channeled to each data center is balanced, and no data center will suffer congestion and reduced performance. •Load balancing for network slicing. Under slicing models, network entities must run dedicated, customized VNFs to serve the traffic associated with each slice. Operators then have to map slices to network nodes, ensuring that the upcoming VNF load is evenly shared across the latter, to optimize the global system performance. •Load balancing at base stations. In the presence of increasingly dense network deployments, users are offered an increased choice of candidate base stations for association. Forecasting allows operators to make informed decisions on new user associations and equalize the charge across base stations to ensure service continuity. Independent of the problem variant, anticipatory load balancing requires a prediction of future traffic that is as accurate as possible. In this case, whether the prediction falls above or below the actual load is irrelevant: any deviation of the prediction from the actual demand causes an imbalance in the resulting load that only depends on the error magnitude, hence a negative error is not more harmful than a positive one. For the purpose of evaluation, we focus on the third problem above, i.e., load balancing at base stations. This type of task is run in modern networks, for instance, by the Policy Control Function (PCF) through the User Equipment (UE) Route Selection Policy (RSP) [57]. The PCF assigns an incoming UE to a Protocol Data Unit (PDU) session or network slice, once the UE is activated. The availability of an accurate prediction of future traffic at each base station allows the NI deployed at the PCF to drive the assignment in a way that the ensuing load is leveled across the radio access infrastructure. As mentioned above, this type of load balancing requires a traditional mobile traffic forecasting model to operate in an anticipatory fashion. We thus rely on a conventional loss function that weights equally negative and positive errors, i.e., the MSE. Based on the traffic load predicted with this loss function, a load balancer performs the corresponding mapping to equalize the (expected) load at the different entities. In the case of load balancing at base stations, forecasting is naturally performed on traffic loads at the base station level; then, the load balancing NI engine running at the PCF manages each association request by assigning the soliciting UE to the base station with the lowest forecasted load. VI. PERFORMANCE EVALUATION We assess the performance of the proposed TES framework in the two use cases set out in Section V, hinging on real-world mobile traffic measurement data collected in an operational network. Specifically, we consider mobile data traffic time series recorded at more than 400 4G/LTE base stations that provide coverage to millions of subscribers in a metropolitan area.2The data was collected in the production infrastructure of a major operator during 11 continuous weeks, by passive probes tapping at interfaces of the Gateway GPRS Support Node (GGSN) and Packet Data Network Gateway (PDNG) in configurable intervals ranging from 5 to 15 minutes. The measurement probes leverage Deep Packet Inspection (DPI) to extract protocol information from packets in the GPRS Tunneling Protocol user plane (GTP-U). Such information is then fed to proprietary classifiers developed by the operator to determine the service associated with each session. As a result, the time series we use in our evaluation describes the traffic generated by individual popular services. All time series have the finest temporal granularity allowed by our dataset, which is 5 minutes. This temporal granularity is compatible with the requirements of the three target use cases, since (i) the reconfiguration periodicity of slice resources allowed by modern Virtual Infrastructure Managers (VIM) is in the order of minutes [58], and (ii) anticipating anomalies or (iii) balancing load by several minutes is largely sufficient to plan and enact countermeasures. Therefore, a prediction of the traffic in the next 5-minute time step (i.e., a point forecast of the time series) is aligned with the use cases, and we consider an output window size O= 1 in all our experiments. Also, we use an input window size of I= 6 time steps to feed the model 2Due to confidentiality reasons, we cannot disclose the identity of the operator, the target geographical region, or the absolute volumes of traffic captured in the data. We thus either normalize the traffic values or report them without the scaling factor that would reveal their order of magnitude.
IEEE TRANSACTIONS ON NETWORK AND SERVICE MANAGEMENT 9 TABLE III: Summary of experimental settings Parameter Value Temporal granularity 5 minutes Input window size (I) 6 time steps (30 minutes) Output window size (O) 1 time step (5 minutes) Data split (weeks) Training: 8; Validation: 2; Test: 1 (corresponding to a history of 30 minutes), and we employ 8, 2, and 1 different weeks of traffic for training, validation, and testing, respectively. The guard value for TES-Transformer second-stage normalization is δ= 0.001. All the experimental settings are summarized in Table III: training, validation, and test sets include both on-peak and off-peak behavior, following a night-day, weekday-weekend pattern [59]. As a final remark, we highlight that our study observes high privacy and ethical standards: (i) the network operator conducted the data collection abiding by applicable regulations at national and international levels; (ii) the competent national privacy agency and the data protection officer of the operator authorized the data processing; and, (iii) the time series we accessed for the purpose of this work solely describe traffic aggregated at individual base stations over large sets of users, and do not contain personal subscriber information. A. Forecasting for capacity allocation We set the capacity allocation use case presented in Section V-A in a network core Cloud scenario, where a data center runs VNFs for the traffic generated in the whole target region by three traffic-intensive mobile applications, i.e., Facebook, Instagram, and Snapchat. Each such service is assigned a dedicated network slice, and the NI responsible for capacity allocation at the data center must reserve in advance enough resources to accommodate the future demand of single slices. To address this problem, we train the models used in TES with the appropriate loss function in (4) and compare our hybrid solution against the following four relevant benchmarks: •INFOCOM19 [14] is the predictor designed by the study that first introduced the problem of capacity forecasting and proposed the loss function in (4). It relies on a DNN architecture fed with a 3D tensor of the spatiotemporal mobile data traffic and uses convolutional layers to capture geographical correlations in the demands. This is the state-of-the-art forecasting model for capacity allocation. •ES-RNN [22] is the GPU implementation of the original ES-RNN approach presented in Section III-A. For the sake of fairness, ES-RNN is trained with the loss in (4). •RNN uses the same RNN architecture of ES-RNN, but relies on a global normalization for the input data, thus without any of the optimizations proposed in this work and in [22]. This benchmark is useful for understanding how statistical modeling favors prediction accuracy. We also train this benchmark with the loss function in (4). •Transformer uses the model first introduced in [60], which is also the basis of our ES-Transformer and TESTransformer models. For this baseline model, we use the loss function in (4) as well. Facebook Instagram Snapchat Services 0 10 20 30 40 50 60 70 80 Normalized Cost [%] INFOCOM 19 25.42 % INFOCOM 19 36.15 % INFOCOM 19 60.36 % RNN 30.97 % RNN 42.50 % RNN 71.43 % Trans 28.51 % Trans 51.70 % Trans 72.15 % ES-RNN 30.54 % ES-RNN 51.39 % ES-RNN 54.67 % ES-Trans 20.01 % ES-Trans 35.71 % ES-Trans 36.30 % TES-RNN 17.79 % TES-RNN 32.90 % TES-RNN 27.99 % TES-Trans18.13 % TES-Trans 34.02 % TES-Trans 35.61 % SoA overprovisioning TES overprovisioning SLA violations Transformer models Fig. 7: Additional capacity allocation cost caused by stateof-the-art models (INFOCOM19, RNN, Transformer, and ESRNN, in blue), and the proposed in this paper (ES-Transformer and TES models, in red) prediction errors. Results refer to three slices assigned to specific services at a network core data center, with parameter α= 3. 1) Overall capacity forecasting performance: We start by comparing the total costs incurred by the operator when supporting capacity allocation with the different forecasting models, in Figure 7. In order to make these values interpretable, all costs are normalized to the (unavoidable) cost of the minimum resources needed to accommodate the exact demand for each service. In other words, costs are expressed as the percent excess over a baseline given by an oracle that makes a perfect prediction. In each case, the figure also tells apart the fraction of the cost resulting from the two sources of penalty, i.e., resource overprovisioning and SLA violations. We group results into state-of-the-art approaches (blue bars in Figure 7) and our three proposals: ES-Transformer, TES-RNN, and TES-Transformer. The key observation is that our approaches consistently outperform the benchmarks, with gains over the secondbest solution that range between 4% and 26%. The different TES approaches yield distinctive results. In general, the ES approach, with or without the τselection, improves the performance of the state-of-the-art algorithms. For instance, TESRNN steadily guarantees very low SLA violation probabilities, which is a desirable feature for the operator. And, it does so by causing an overprovisioning that is lower than or comparable to that produced by the other predictors. These are very encouraging findings, as one of the benchmarks is the stateof-the-art model designed for capacity forecasting. Interestingly, ES-RNN yields an allocation of unnecessary resources close to that of TES but incurs much more frequent SLA violations. Transformer, in both the ES and TES configurations, improves the performance of the pure Transformer counterpart for all the services, although TES-Transformer incurs a higher SLA violation rate due to its aggressive forecasting. INFOCOM19, RNN, and Transformer, when compared to TES, induce a substantially higher overprovisioning that often helps limit SLA violations, at the expense of an overall higher expenditure.
IEEE TRANSACTIONS ON NETWORK AND SERVICE MANAGEMENT 16 Albert Banchs (Senior Member, IEEE) received the M.Sc. and Ph.D. degrees from the Polytechnic University of Catalonia (UPC-BarcelonaTech) in 1997 and 2002, respectively. He is currently a Full Professor with the University Carlos III of Madrid (UC3M), with double affiliation as the Director of the IMDEA Networks Institute. Before joining UC3M, he was with ICSI Berkeley in 1997, Telefonica I+D in 1998, and NEC Europe Ltd., from 1998 to 2003. He was an Academic Guest with ETHZ in 2012, a Visiting Professor with EPFL in 2015, 2013, and 2018, and a Fulbright Scholar with The University of Texas at Austin in 2019. He is the author over 150 publications in international conferences and journals and is the co-inventor of several patents. Xavier Costa is ICREA Research Professor, Scientific Director at the i2cat Research Center and Head of 6G R&D at NEC Laboratories Europe. His team generates research results that are regularly published at top scientific venues, produces innovations that have received several awards for successful technology transfers, and participates in major European Commission R&D collaborative projects. He has held multiple leadership positions both in industry and research organizations, such as Deputy General Manager, Chief Researcher, Technology Board member, and Scientific Advisory Board Member. As a standards delegate, he contributed to multiple standardization bodies (e.g., IEEE 802.11, 802.16, WiFi Alliance, 3GPP, ...) and was recognized in several standards as a top contributor. He has served on the Organizing Committees of several conferences (including ACM MOBICOM, IEEE INFOCOM, WCNC, and Greencom), published papers of high impact, and holds about 100 granted patents. He has served as Editor at IEEE Transactions on Mobile Computing (TMC), IEEE Transactions on Communications (TCOM), and Elsevier Computer Communications journals (COMCOM). Xavier received both his M.Sc. and Ph.D. degrees in Telecommunications from the Polytechnic University of Catalonia (UPC) in Barcelona and was the recipient of a national award for his Ph.D. thesis.