scieee AI-readable full text Open interactive document viewer

Faro: a scalable and reliable outage detection algorithm for IoT Mobile Virtual Network Aggregators

Milani, Mattia; Vomhoff, Viktoria; Bega, Dario; Gramaglia, Marco; Geissler, Stefan; Carsai, Ksaba; Serrano, Pablo

Abstract

Internet of Things (IoT) Mobile Network Aggregators (MNAs) are an increasingly successful paradigm for providing ubiquitous mobile Internet connectivity for Smart Devices. They leverage the 4G/5G roaming architecture offered by Mobile Network Operators (MNOs) worldwide to offer a single gateway to the Internet. In this context, promptly detecting network outages is fundamental: Mobile Network Aggregators (MNAs) do not have control over the visited network infrastructure, and they offer connectivity to a huge number of possibly unmanaged devices. However, this activity is currently a painstaking process carried out manually by network engineers, who review trends in control plane signalling messages that are forwarded in the Mobile Network Aggregators (MNAs) managed core. In this paper, we present Faro, a self-supervised solution for the automatic detection of network outages. Faro leverages a contrastive learning framework that is particularly suited for this scenario, characterised by very skewed input data, few samples, and the use of confidence scores to automatically detect changes in the input data. We evaluate Faro using a real-world dataset comprising more than 5 million control plane messages across multiple countries and over 6 months, demonstrating its ability to accurately diagnose all outage scenarios in less than 15 minutes.

Full text

Faro: a scalable and reliable outage detection algorithm for IoT Mobile Virtual Network Aggregators MATTIA MILANI, Universidad Carlos III de Madrid, Spain and Nokia S&T, Germany VIKTORIA VOMHOFF, Chair of Communication Networks, University of Würzburg, Germany, Germany DARIO BEGA, Nokia Bell Labs, Germany MARCO GRAMAGLIA, Universidad Carlos III de Madrid, Spain STEFAN GEISSLER, Chair of Communication Networks, University of Würzburg, Germany, Germany CSABA KARSAI, EMnify GmbH, Germany PABLO SERRANO, Universidad Carlos III de Madrid, Spain Internet of Things (IoT) Mobile Network Aggregators (MNAs) are an increasingly successful paradigm for providing ubiquitous mobile Internet connectivity for Smart Devices. They leverage the 4G/5G roaming architecture offered by Mobile Network Operators (MNOs) worldwide to offer a single gateway to the Internet. In this context, promptly detecting network outages is fundamental: Mobile Network Aggregators (MNAs) do not have control over the visited network infrastructure, and they offer connectivity to a huge number of possibly unmanaged devices. However, this activity is currently a painstaking process carried out manually by network engineers, who review trends in control plane signalling messages that are forwarded in the Mobile Network Aggregators (MNAs) managed core. In this paper, we present Faro, a self-supervised solution for the automatic detection of network outages. Faro leverages a contrastive learning framework that is particularly suited for this scenario, characterised by very skewed input data, few samples, and the use of confidence scores to automatically detect changes in the input data. We evaluate Faro using a real-world dataset comprising more than 5 million control plane messages across multiple countries and over 6 months, demonstrating its ability to accurately diagnose all outage scenarios in less than 15 minutes. CCS Concepts: •Computing methodologies → Anomaly detection;Artificial intelligence;•Networks → Mobile networks;Network monitoring. Additional Key Words and Phrases: Similarity learning, Siamese networks ACM Reference Format: Mattia Milani, Viktoria Vomhoff, Dario Bega, Marco Gramaglia, Stefan Geißler, Csaba Karsai, and Pablo Serrano. 2025. Faro: a scalable and reliable outage detection algorithm for IoT Mobile Virtual Network Aggregators. Proc. ACM Netw. 3, CoNEXT4, Article 35 (December 2025), 23 pages. https://doi.org/10.1145/3768982 1 Introduction Mobile Network Aggregators (MNAs) such as Airalo [ 1 ], HolaFly [ 3 ] or Emnify [ 2 ] are an increasingly popular trend in the mobile network ecosystem [ 5 , 43 ]. They differ from Mobile Virtual Network Operators (MVNOs) as they offer customers mobile connectivity globally by using multiple Mobile Authors’ Contact Information: Mattia Milani, [email protected], Universidad Carlos III de Madrid, Madrid, Spain and Nokia S&T, München, Bayern, Germany; Viktoria Vomhoff, [email protected], Chair of Communication Networks, University of Würzburg, Germany, Würzburg, Germany; Dario Bega, [email protected], Nokia Bell Labs, Stuttgart, Germany; Marco Gramaglia, mg[email protected], Universidad Carlos III de Madrid, Madrid, Spain; Stefan Geißler, [email protected], Chair of Communication Networks, University of Würzburg, Germany, Würzburg, Germany; Csaba Karsai, [email protected], EMnify GmbH, Würzburg, Germany; Pablo Serrano, [email protected], Universidad Carlos III de Madrid, Madrid, Spain. This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. ©2025 Copyright held by the owner/author(s). ACM 2834-5509/2025/12-ART35 https://doi.org/10.1145/3768982 Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. 35:2 Mattia Milani et al. MNA Internet PGW HSS Data Monitoring Human analyst MNO MNO MNO Customer 1 Open ticket 2Investigate 3 React & Restore Fig. 1. Overview of the relationship between MNA and customers describing a possible flow in case of undetected outages. Network Operators (MNOs) instead of just one, and connecting through the one that best suits the customer’s needs at any given moment. This very flexible paradigm is especially well suited for the Internet of Things (IoT), as manufacturers can have a commercial agreement with just one MNA and enjoy seamless connectivity worldwide by visiting the networks of several MNOs. Enforcing high Quality of Experience (QoE) standards for customers is hence paramount, as sudden outages in one visited network can cause the disconnection of a very high number of unmanaged smart devices that depend on Internet connectivity. Indeed, this is a challenging task to be performed from the MNAs’ perspective: they do not have access to the MNO infrastructure, since only the core is operated by the MNA, and the only available information comes from control plane signalling traffic, usually authentication and registration messages. One of the most frequent, yet difficult to discover, root causes of service degradation are short-term outages occurring in the visited networks. These aspects, combined with the typically very broad geographical extent covered by IoT devices, make the currently adopted manual verification performed by the MNAs’ network engineers a painstaking process, as it entails the analysis of hundreds of thousands of data samples at the same time. This calls for semi-supervised solutions to address this problem, as currently state-of-the-art manual procedures employed by MNAs limit detection efforts to a small subset of operators or even customers operating on a set of specific MNOs, creating a cascading set of inefficiencies: engineers are forced to perform a tedious, error-prone, mostly manual, and time-consuming task, and achieve very limited coverage of the problem surface. Figure 1 depicts the scenario where an MNO that is not actively monitored by the MNA may experience an outage, prompting the customer to raise a ticket. Only then, the MNA network engineers investigate the problem and apply appropriate actions to mitigate it. However, this lack of a proactive approach can cause delays in the reaction from the MNA, ultimately reducing customer satisfaction. This creates a strong incentive to identify new diagnostic tools that can operate self-supervised, improving detection coverage and reducing manual effort. In this paper, we propose Faro, an unsupervised Deep Learning (DL) solution for the identification of service outages for worldwide IoTs MNAs, with the following characteristics: 1 Resiliency to Data Skewness. The working assumption is that the network, only in exceptional conditions, is suffering outages as shown in [20]. Thus, the vast majority of the data analysed by IoT MNAs does not exhibit unwanted behaviour, leading to a severe skewness in the training data. We overcome this problem by employing contrastive learning techniques on top of an Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. Faro: a scalable and reliable outage detection algorithm for IoT Mobile Virtual Network Aggregators 35:3 Siamese Neural Networks (SNN), separating faulty behaviour from normal network operation through a triplet loss. By training the SNN, the model learns to differentiate the characteristics that indicate a malfunction from those that are within a normal operational range. 2 Limited data labelling. Connectivity data arriving at the IoT MNAs are overwhelming (with an average of 55 million control plane messages per day) and including outages only for a minimal period. This makes data labelling for training a supervised model hard: hence, Faro training has been designed to be as efficient as possible, leveraging only a very limited amount of labelled data, and self-labelling unlabelled samples leveraging the Curriculum-Learning (CL) [8, 18] technique. This makes Faro a self-supervised model that can be used in the wild. 3 Confidence Score. Having a model that self-labels and autonomously assesses samples, to detect whether they are coming from an MNO that is suffering an outage, requires a control mechanism to constantly assess the quality of the taken decision. For this reason, we integrate in the Faro framework a confidence score that can be leveraged by network engineers (or automatically through a threshold-based algorithm, as we propose) for model lifecycle management, i.e., to trigger a retraining or sandboxing of the decision taken by the model. Faro addresses these requirements by providing a scalable and label-efficient solution for anomaly detection in real-world deployments. This allows for broader coverage, reduces the operational burden on engineers, and enables faster, more consistent detection. By aligning learning objectives with the structure and constraints of real-world data, Faro contributes both a scalable anomaly detection framework and a practically applicable solution that advances the state of the art in self-supervised monitoring of mobile infrastructures. We validate Faro over a large-scale dataset that collects information from an IoT MNA operating worldwide, emulating the real-world operation of an MNA connecting several MNOs worldwide. Faro detects 93.3 % of the outages in less than 10 minutes from their beginning during a supervised evaluation. Extending our test to a self-supervised scenario over the data collected from 119 MNOs Faro can identify outages for which we only had coarse information (i.e., the day when they happened). The paper is structured as follows: Section 2 lays the grounds of the problem, discussing current industrial methodologies, the state-of-the-art solutions for anomaly detection algorithms in mobile networks, and the data acquisition setup used for this work. Then, in Section 3, we discuss the design of our solution, detailing all the elements that fulfil the requirements discussed above. We show the performance of our approach in Section 4, before concluding in Section 5. 2 Context Detecting outages is a complex task that exhibits heterogeneous levels of maturity: while in the literature there are Artificial Intelligence (AI) based solutions for the problem, their complexity or their supervised nature often makes them not suitable for their application in a production system. In this section, we review the relevant solutions in the state-of-the-art and present the details of the data used to design Faro. 2.1 Limitations of current approaches In commercial MNAs and broader telecommunication infrastructure, anomaly detection and outage identification are typically performed using rule-based monitoring, threshold alarms, and manual inspection of aggregated Key Performance Indicators (KPIs). These systems rely heavily on domain expertise and historical baselines, often lacking adaptability to evolving traffic patterns or unseen failure modes. Alarms are triggered when metric thresholds are breached, but these thresholds are often statically defined, leading to high rates of false positives or undetected anomalies in skewed or low-volume regions of the network. Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. 35:4 Mattia Milani et al. Furthermore, the detection of service outages often depends on indirect indicators such as sudden drops in traffic, surge in attach failures, or changes in core signalling patterns (e.g., PDP context creation/drop ratios, Update Location events). These signals are aggregated over temporal bins (e.g., per-minute summaries), which limits fine-grained analysis and makes manual triage laborious and reactive rather than proactive. To bridge this gap, there is a growing need for intelligent, data-driven solutions that can operate in a fully automated manner while still supporting expert oversight. Our work is motivated by this practical context and proposes a system that leverages similarity learning and self-labelling to continuously adapt to evolving patterns in real-world data while maintaining trust through human-in-the-loop validation. 2.2 Anomaly detection for network data Anomaly detection in complex systems such as telecommunications and IoT infrastructures faces significant challenges due to the high volume and velocity of data, limited ground truth availability, and the presence of data skewness and noise. Traditional supervised learning proves largely impractical due to the prohibitive cost and delays of manual event labelling. For example, in telecommunications, prior work shows signalling message statistics can reveal network incidents [ 39 ], highlighting the inherent value of control plane data even without machine learning. Although such domain-specific insights are crucial, they are often insufficient alone, driving the field’s increasing interest in unsupervised, semi-supervised, and self-supervised learning paradigms to overcome these limitations. Recent surveys provide a broad overview of this evolving landscape. General-purpose anomaly detection methods for IoT and time-series data span a spectrum, from statistical models to advanced deep learning approaches and hybrid systems, all designed for robustness and scalability in noisy, real-world environments [11, 16, 45]. In MNO environments, DL techniques are crucial to analyse vast streams of network performance data, call detail records, and signalling messages. Architectures like autoencoders and LSTM-based predictors are frequently used to capture sequential dependencies in network traffic and detect deviations from the learnt dynamics that could indicate outages or security breaches [4, 29, 30]. Generative adversarial models [ 46 ] also find application in learning complex normal network behaviours. However, these reconstructionor prediction-based methods often struggle with the dynamic and highly imbalanced nature of real-world network anomalies, frequently requiring extensive tuning or being sensitive to novel, unseen failure patterns. Anomaly detection in IoT, by contrast, often deals with a wider heterogeneity of devices, resource constraints, and diverse data types from sensors to industrial control systems. Here, deep learning approaches, including various encoder-decoder architectures [ 7 , 37 ], are adapted to handle varying data rates and device-specific anomalies [ 38 ]. The emphasis often shifts to lightweight models or edgebased processing because of device limitations. Across both MNO and IoT contexts, ensemble methods remain a common strategy to improve robustness amid noise and limited supervision [4, 30, 38]. Given the inherent scarcity of labels across both telecommunications and IoT, self-supervised and semi-supervised learning have emerged as crucial paradigms. Contrastive learning has proven particularly effective for representation learning without explicit labels, producing semantically meaningful embeddings that strongly support downstream tasks like anomaly detection. Foundational work on pairwise and triplet-based similarity learning [ 15 , 19 , 40 ] paved the way for more recent advancements such as SimCLR [ 13 ] and MoCo [ 21 ]. Surveys categorise these methods and explore their diverse applications in both temporal (e.g., MNO traffic patterns) and structured (e.g., IoT device sensor readings) domains [23, 24]. Building on this, SNNs, which learn embeddings by processing input pairs through shared-weight encoders, have been widely adopted in few-shot and anomaly detection tasks. These architectures are particularly suited to scenarios with scarce labels and class imbalance. For example, few-shot Siamese Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. Faro: a scalable and reliable outage detection algorithm for IoT Mobile Virtual Network Aggregators 35:5 models have been proposed for anomaly detection in industrial cyber-physical systems [ 47 ], and intelligent IoT systems [42], showing strong generalisation from minimal supervision. The FaceNet model [ 34 ], while designed for face recognition, remains a canonical reference for contrastive and triplet-based metric learning. Furthermore, semi-supervised learning strategies offer another way to reduce the dependence on explicit labels. Approaches like consistency regularisation and pseudo-labelling frameworks, such as FixMatch [ 35 ] and DASH [ 41 ], leverage model confidence to iteratively expand labelled sets. These approaches are especially useful in scenarios where operators can only provide a small number of labelled examples, typical of many real-world anomaly detection pipelines. To operationalise machine learning in settings with limited ground truth and dynamic behaviour, Human-In-The-Loop (HITL) learning is gaining increasing attraction. Active learning [ 33 , 44 ] and few-shot learning techniques [6] have been proposed for intrusion detection, IoT security, and other applications, allowing models to query experts strategically or generalise from minimal data. These methods contribute to improving trust and explainability in deployed systems. Our work critically builds upon these established foundations by combining contrastive representation learning with a novel confidence-based self-labelling framework and the option for human feedback. This unified approach directly addresses three core challenges inherent in real-world anomaly detection systems: label scarcity, data skewness, and operational interpretability. Ultimately, Faro aims to narrow the persistent gap between theoretical advancements in machine learning and the practical constraints of deploying such systems in dynamic, high-volume environments. 2.3 Real-World dataset As global IoT connectivity scales, managed connectivity providers often acting as MNAs play a central role in enabling worldwide device deployment without operating their own Radio Access Network (RAN). Instead, they rely on roaming agreements and global mobile core infrastructure. Consequently, all devices in such networks are continuously roaming. These devices first connect to a local (visited) network and then access their home network via the IPX backbone, using a home-routed roaming approach. In this setup, not only user data but also signalling traffic (required for authentication, session establishment, and mobility management) is routed through a complex, multi-stakeholder infrastructure. Fig. 2 shows this roaming architecture for a 4G setup. It illustrates the flow of control-plane signalling and user-plane traffic across the visited network, IPX carrier, and home network. The Mobility Management Entity (MME) in the visited network handles mobility and session control, while the Home Subscriber Server (HSS) in the home network manages user identities and subscription information. These components communicate via the Diameter protocol (shown in red), using signalling messages such as Update Location for mobility management. Simultaneously, user-plane session management and data tunnelling are handled via the GTP-C protocol (shown in blue), with messages such as Create PDP Context and Delete PDP Context exchanged between the Serving Gateway (SGW) and Packet Gateway (PGW). 2.3.1 Dataset collection. The dataset (later identified by D ) used in this study was collected from an MNA that fits this architecture. It includes aggregated features derived from GTP-C signalling messages (specifically, Create PDP Context and Delete PDP Context, captured at the PGW) marked in blue in Fig. 2. It also includes mobility-related signalling messages such as Update Location and Update GPRS Location, indicated in red in Fig. 2. The data is aggregated per minute per MNO, and different statistical features (reported in Table 1 in the Appendix) are computed for each time bin. Each row in the dataset represents a one-minute interval for a specific operator, forming a twodimensional time series indexed by timestamp and operator ID. Each entry includes 14 numerical Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. 35:6 Mattia Milani et al. Roaming IPX Carrier Diameter Signaling Visited Network SGW MME Home Network PGW HSS End Device Internet GTP-C Diameter Fig. 2. End-to-end roaming architecture showing signalling (red: Diameter, blue: GTP-C) between visited and home network components in a 4G IoT deployment. features that summarise core network KPIs. These include counts of successful and failed attach or PDP context requests, uplink data attempts, and alerts triggered by unsupported or blocklist operator selections. In the pre-processing step described in Section 3.1, the dataset is extended to 40 features. This dataset allows the validation of Faro in a realistic scenario with data samples measured directly in an MNA core network and meaningful signalling traffic KPIs. Although the data used to train and test Faro has been collected in a live 4G industrial environment, Faro has been designed to operate seamlessly over 5G and beyond mobile networks. In fact, the metrics reported in Table 1 can also be mapped to 5G measurements: for example, the number of received PDP context creation requests can be easily translated into the number of received PDU session establishment requests. 3 Design of Faro In this Section, we discuss the fundamental building blocks used by Faro, and how they build towards a carrier-grade solution for detecting outages in a MNA. The overall framework is depicted in Fig. 3. Data is collected from the core network managed by an MNA, eventually separated into labelled subsets belonging to normal operation and outages. The samples coming from these distributions are used in a contrastive learning solution, which leverages an SNN to tell apart in an embedding space samples belonging to different classes (red arrows in Fig. 3) and moving closer samples coming from the same classes (blue arrows in Fig. 3). Through a distance layer, new samples are classified according to their category, and this information is used to further enrich the training dataset with newly labelled data, not available in the first iteration. Finally, a confidence score is computed for each classification decision. Samples with low confidence will raise an alarm that can be handled by a human or, as we propose in this world, automatically trigger a re-training. 3.1 Dataset description Two main sets compose the dataset D collected as described in Section 2.3. The first is a labelled subset, consisting of 6 months of data from 6 MNOs operating in six different countries, from June to November 2022 . Each time bin is manually annotated through collaboration with domain experts, indicating if it belongs to normal operations or to a network outage. This subset includes approximately 1.9 million samples, with outages representing a minority class due to their relative rarity. The second subset is unlabelled and includes data from a broader set of operators ( 119 ) over a period of one month (September 2024 ), totalling around 5.1 million samples for 53 different countries on six continents. The only information provided for this portion of the dataset, identified by a human expert, is that 3 known operators experienced at least an outage on a specific date. Neither the time interval for such an outage nor information about other potential outages on any date is available. 3.1.1 Pre-processing. As mentioned in Section 2.3.1, we expand D from the initial set of 14 numerical features into 40 features by computing additional statistical metrics. This list, detailed in Table 1 in Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. Faro: a scalable and reliable outage detection algorithm for IoT Mobile Virtual Network Aggregators 35:7 Separation Data Contrastive learning 𝑙0. . . 𝑙𝑛 𝑑layer Emb. space 𝑑(𝑠+,𝑠𝑎) 𝑑(𝑠+,𝑠𝑎) 𝑓𝑒 𝑓 𝑠+ 𝑠𝑎 𝑠− 𝑒+ 𝑒𝑎 𝑒− e 𝑁 Curriculum Learning ( e 𝑁, e 𝑂,𝐷𝑑) e 𝑂 e 𝐷 𝐷𝑡expansion Confidence Score Human analyst Classifier 𝑝 Labels Fig. 3. Faro infrastructure visual representation. the Appendix, has been constructed after a careful analysis of the dataset characteristics. We then use this extended version to derive the training, validation, and testing sets for our model. The objective is to generate samples from raw time-series data, discretised into 1 min time bins, a value we use to balance the noisiness of the data with the agility in identifying an outage. Each sample includes a look-back window of 120 min , allowing the model to capture the temporal dynamics of a possible outage event, estimating outage probabilities using statistics derived from raw KPIs trend changes within the selected time window. The duration of labelled outages in the dataset spans from 150 min to 67 h : for this reason, we use a 120 min look-back window to find a balance between covering a sufficiently long part of the outage event while avoiding overreacting to transient variations in KPI trends. Following this procedure, we obtain for each MNO a curated time series describing the evolution over the last 120 min . In the subset of the dataset where labels are available, they are assigned to the post-processed data samples: a sample is labelled as an outage if at least one minute within the window shows an outage behaviour. Otherwise, it is marked as normal operation (i.e., no outage occurs within the 120 min interval). As the sliding window moves forward, overlapping samples during an ongoing outage will increasingly contain anomalous data; still, the label remains binary. This labelling strategy is motivated by the fact that outages are temporally contiguous. Thus, a data sample containing a 1 min outage indicates either the start of an outage (if the 1 min slot is at the end of the sample) or the end of an outage (if it is at the beginning). This pre-processing step enhances the ability of Faro to anticipate outage detection compared to state-of-the-art approaches. The full set of statistics is reported in Table 1 in the Appendix, alongside their description. In total, each sample contains 40 features: some of them are aggregated using a roll function, which computes multiple statistical indicators over the time windows, others use a lag function that uses the exact value of the metrics at the beginning of the window. 3.2 Similarity learning SNNs[ 14 , 34 , 36 ] are a type of Deep Neural Network (DNN) specialised in identifying and aggregating samples that exhibit similar features to each other and telling them apart from the different ones, falling into the category of contrastive learning techniques. SNNs owe their name mainly to the fact that their architecture uses two or more identical subnetworks that share the same weights and parameters. In this way, multiple inputs are processed at the same time by the network. SNNs map Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. 35:8 Mattia Milani et al. input samples onto an embedding space (in Faro, represented in a 16 -dimensional space) that can be used to subsequently cluster and classify them, according to their mutual distance and labels. In Faro, this effect is achieved by training the network with a triplet loss function[ 17 , 34 , 40 ].This function uses a triplet (𝑠+,𝑠𝑎, 𝑠−) , where 𝑠𝑎 is the anchor sample considered for training and 𝑠+ , 𝑠− are the positive and negative references. Hence, in Faro, if 𝑠𝑎 label is normal behaviour, then 𝑠+ and 𝑠− are samples drawn from the normal and misbehaving sets, and vice-versa. Once the triplet samples are mapped in the embedding space, Euclidean distances 𝑑(𝑠+,𝑠𝑎) and 𝑑(𝑠−,𝑠𝑎) (also abbreviated in 𝑑+𝑎 and 𝑑−𝑎 ) are computed to quantify the sample similarity, and finally, the distances are minimised using Eq. (1): 𝑙=𝑚𝑎𝑥 (0,||𝑑(𝑠+,𝑠𝑎)||2− ||𝑑(𝑠−,𝑠𝑎)||2+𝜇)(1) where 𝜇 is a parameter used to enforce a minimum distance by which the negative reference must be farther from the anchor than the positive one. During one training backpropagation step, the optimisation of Eq. (1) achieves the goal of grouping 𝑠+ and 𝑠𝑎 while telling 𝑠𝑎 and 𝑠− apart. An illustration of a SNN is depicted in Fig. 3, where 𝑓 identifies the entire neural network, which is divided into 𝑓𝑒 that transforms samples into embeddings, and the 𝑑 layer to compute distances. Then, from the two distances, we compute the label according to their ratio in the embedding space, as follows: 𝑝=1−𝑑(𝑠−,𝑠𝑎) 𝑑(𝑠+,𝑠𝑎) + 𝑑(𝑠−, 𝑠𝑎)(2) Where 𝑝 represents the probability that 𝑎 belongs to the same class of 𝑠− , which translates into an outage probability when 𝑠−comes from that set of samples. 3.3 Making Faro self-supervised As discussed in the literature [ 34 ], the selection of the training samples is key for the performance of SNNs. Ideally, samples that belong to different categories and exhibit heterogeneous features provide the best results for the model. We call them easy samples. However, the training dataset obtained after pre-processing contains data samples, e.g., the ones with few time bins labelled as outage, that cannot be clearly marked as outage or normal operation. We call them difficult samples. Furthermore, in the scenario targeted by Faro, data is skewed: it is expected that visited networks are mostly exhibiting normal operation, while exceptionally presenting characteristics that indicate an outage. This skewness (in the training dataset, only 4.95 % of the samples are classified as outage, but there we focus on MNOs that at some point in time had an outage, so the real fraction in the full dataset is much lower) limits the extent of the usable data for training, as using a too unbalanced training set certainly leads to poor performance when the model is used in inference. These limitations, i.e., skewness and different levels of difficulty in the data samples, hinder the abstraction capability of the DNN. Indeed, training cannot be performed mixing difficult and easy samples from the start, because the model will not be able to correctly distinguish them. In Faro, we solve these problems first by employing a contrastive learning technique, and second, with a CL [ 18 , 28 ] technique to incrementally introduce more difficult and heterogeneous samples into the training dataset labelled in a self-supervised fashion. The basic idea behind CL is to start only from a small set of easy samples belonging to each of the two categories to build stronger generalisation foundations in the SNN. Samples that do not belong to them are then iteratively classified as follows: after the model is trained with the initial easy samples, it is used in inference to classify unknown difficult samples. These are re-injected into the training dataset with the label assigned by the model, the model is Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. Faro: a scalable and reliable outage detection algorithm for IoT Mobile Virtual Network Aggregators 35:9 Algorithm 1 CL-Clustering function 1: procedure CL-Clustering(𝑓𝑒, e 𝑁, e 𝑂,𝐷𝑑,𝐷𝑡) 2: f 𝑁𝑒, f 𝑂𝑒←𝑓𝑒( e 𝑁, e 𝑂)⊲Reference embeddings generation 3: 𝑒𝑁, 𝑒𝑂←𝐶𝑙𝑢𝑠𝑡𝑒𝑟𝑖𝑛𝑔( f 𝑁𝑒, f 𝑂𝑒)⊲Clustering and extract centroids 4: 𝐷𝑑 𝑒←𝑓𝑒(𝐷𝑑)⊲Transform 𝐷𝑑into the embedding space 5: 𝑑+𝑎,𝑑−𝑎←𝐺𝑒𝑡𝐷𝑖𝑠𝑡𝑎𝑛𝑐𝑒𝑠(𝑒𝑁, 𝑒𝑂, 𝐷𝑑 𝑒)⊲Get cluster centroids distances 6: 𝑙𝑑←𝐺𝑒𝑡𝐿𝑎𝑏𝑒𝑙 (𝑑+𝑎,𝑑−𝑎, 𝜇𝐶𝐿 )⊲Distances to labels 7: 𝐷𝑡←𝐴𝑝𝑝𝑒𝑛𝑑(𝐷𝑑,𝑙𝑑)⊲Add 𝐷𝑑samples to 𝐷𝑡 8: return 𝐷𝑡⊲Return updated set 9: end procedure re-trained, and the set of unlabelled difficult samples is re-evaluated. The controlled injection of new samples acts as a specialisation phase, where the SNN refines the weights to enhance its generalisation capabilities. This sequence is repeated until the SNN reaches a plateau, meaning that the model cannot further improve the quality of its categorisation. 3.3.1 Faro initial training. As described above, to implement CL, the labelled subset of the initial dataset D is split into three categories: two easy, containing data samples belonging to normal operation and outage; called e 𝑁 and e 𝑂 , respectively. In particular, samples with at least 10 minutes of data labelled as outage within the 120 minute interval belong to e 𝑂 . We set this as our SNN can effectively dissect them with very high accuracy with the above threshold. The third category contains the difficult cases (called e 𝐷 ), which are the samples harder to categorise given a shorter outage time interval. Samples that exhibit outage data for a time interval between 1 and 9 minutes over the last 120 minutes belong to this category. At this point, training, validation, and test datasets are generated (respectively identified by 𝐷𝑡 , 𝐷𝑣 , 𝐷𝑡𝑒𝑠𝑡 ): 𝐷𝑡 initially contains only data samples from e 𝑁 and e 𝑂 , while 𝐷𝑣and 𝐷𝑡𝑒𝑠𝑡 also use e 𝐷. As mentioned in Section 3.2, the SNN infrastructure requires an input triplet (𝑠+,𝑠𝑎, 𝑠−) . To respect the real-world skewness of the data (with sporadic outages over the whole dataset), we also applied a balancing function while generating 𝐷𝑡 , so that it contains 99 % samples from e 𝑁 as 𝑠𝑎 and 1 % from e 𝑂 . Training triplets 𝑠+ and 𝑠− are selected accordingly, while samples in 𝐷𝑣 and 𝐷𝑡𝑒𝑠𝑡 are equally balanced. The SNN is initially trained over 𝐷𝑡 while samples from e 𝐷 that were not included in 𝐷𝑣 and 𝐷𝑡𝑒𝑠𝑡 , are placed into 𝐷𝑑 set for the proposed CL algorithm (named CL-Clustering ) as described next. This is to keep 𝐷𝑣and 𝐷𝑡𝑒𝑠𝑡 sets consistent for performance evaluation. 3.3.2 In-Training Clustering Exploitation. To maximise the effectiveness of the CL approach when used in a production system, we propose the CL-Clustering algorithm (detailed in Algorithm 1), which leverages inference time decisions to autonomously label samples that were unknown (or not used because difficult) during training time. The triplet selection process described in Section 3.2 plays a critical role during the SNN training [ 34 ]. Applying the state-of-the-art CL technique implies randomly selecting at every iteration, 𝑠+ and 𝑠− reference samples from the respective categories. This results in several disadvantages, as a wrong selection may induce a counterproductive effect on the network. In contrast, CL-Clustering selects reference samples based on the embedding space generated by 𝑓𝑒 after the initial training, harvesting cluster information and using it to optimise the triplet generation, leading to multiple benefits over state-of-the-art CL techniques[ 28 , 34 ]: ( 𝑖 ) more deterministic and stable comparison samples resulting in more accurate and precise decisions; ( 𝑖𝑖 ) smaller memory footprint, since only the centroids of the clusters are kept in memory and refined over time, instead of all the data samples Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. 35:16 Mattia Milani et al. -5 0 Values 11:48 11:52 11:00 12:00 13:00 14:00 15:00 OP 30 - 5th Sept. 0 50 100 p [0-100] 11:00 12:00 13:00 14:00 15:00 OP 31 - 5th Sept. F1 F2 F3 p Outage Fig. 9. Outage probability 𝑝for two co-geolocated MNOs. an external event, such as a power outage, affected both operators at the same time. An alternative explanation is that one or both MNOs shared underlying infrastructure, and therefore an infrastructurelevel outage impacts them both. These results further support the notion that Faro can operate effectively in a real production system with minimal manual intervention. 4.3 Faro explainability While the Faro unsupervised features make it attractive for large-scale utilisation, they also pose challenges to understanding why the system made a certain decision. Human-in-the-loop approaches can be coupled with Faro to validate decisions in real-time, but further information can be extracted from analysing the samples marked as outages. We hence select the embeddings 𝑒 generated by Faro for 91 090 samples classified as outages and cluster them using the HDBSCAN [ 10 ] algorithm, optimising the Silhouette score [ 32 ] to obtain the best number of clusters from the set. We processed them using two feature explainability algorithms, namely LIME [ 31 ] and SHAP [ 26 ], to obtain the relative weight of the input features in each cluster. Thanks to this analysis, we can tell apart the two main kinds of outages: • User plane traffic anomalies: outages are identified thanks to spurious spikes and abnormal behaviour in the amount of exchanged traffic (in both uplink and downlink). We found 15 430 samples in this class, belonging to 107 operators. • Control plane traffic anomalies: in this case, outages exhibit frequent User Equipment (UE) (de-)registrations and re-connections, impacting the control plane of the network. We found 75 660 samples in this class, belonging to 119 operators. This side information can be further leveraged by MNAs to specifically target interactions with their final customers. Details on the most important features for each category are provided in Table 2 in the Appendix. 4.4 Model lifecycle management In this section, we evaluate Faro lifecycle management capabilities leveraging 𝜎(ℎ,𝑘) metric over the unlabelled portion of D , focusing on data from 3 MNOs ( 11 , 30 , and 714 ). We chose them because they exhibit different behaviours: 11 exhibits long abnormal patterns with low 𝑝 , while the MNO 30 presents a burst behaviour. Lastly, MNO 714 acts as a counter-test where Faro retraining should not affect a very good behaviour. This diversification allows us to test the effectiveness of the approach devised under different conditions. Fig. 10 depicts in the upper plot the outage probability 𝑝 evaluated by Faro, while at the bottom, the confidence metric 𝜎(ℎ,𝑘) is shown. MNO 11 has periods in which the outage probability 𝑝 is almost 0 and the confidence score is almost 1 meaning that no outages have Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. Faro: a scalable and reliable outage detection algorithm for IoT Mobile Virtual Network Aggregators 35:17 25 50 75 100 p [0-100] OP 11 OP 30 p ( h , k ) p ( u )Normal Outage OP 714 01 05 09 Days [1st - 11th] 0 0.25 0.5 0.75 1 ( h , k )[0-1] 01 05 09 Days [1st - 11th] 01 05 09 Days [1st - 11th] Fig. 10. Retraining labels used for the operator retraining phase, data from the 1St of September to to the 11Th 25 50 75 100 p [0-100] Retrain No Retrain 11 1818 24 30 0 0.25 0.5 0.75 1 ( h , k )[0-1] 11 1818 24 30 Op 11 - Sept. 2024 Retrain No Retrain 11 1818 24 30 11 1818 24 30 Op 30 - Sept. 2024 Retrain No Retrain 11 1818 24 30 11 1818 24 30 Op 714 - Sept. 2024 Fig. 11. Performance comparison between the retraining cycle and the static solution. Data from the 11 Th of September to the end of the month been detected, while there are others where 𝑝 is close to 50 % but at the same time the confidence score is very low, meaning that Faro identified patterns for which is not certain. MNO 30 is characterised by a different behaviour: in the first part of September, the outage probability 𝑝is almost 0along with a confidence score almost equal to 1. Only around the 10th and 11th of September, it exhibits higher values of 𝑝along with a lower confidence score. In the second part of September, the operator is characterised by unstable values of 𝑝 followed by low values of confidence score, meaning that Faro model was more uncertain about the taken decisions. Finally, MNO 714 is characterised by normal operation for almost all the months except for a few time intervals in which the outage probability 𝑝rises together with a low confidence score. Following the insights provided by the confidence score 𝜎(ℎ,𝑘) , data samples with low values are collected and used for retraining Faro. For this experiment, only data samples up to the 11 th of September are considered for retraining, allowing the model to be tested on unseen data. In particular, Fig. 10 highlights the subset of data used for retraining through colour patches, which indicate the labels assigned to the samples by Faro. The sets of samples have been selected by applying the threshold-based algorithm defined in Section 3.4.1, emulating a production environment where, when the 𝜎(ℎ,𝑘) score falls below the defined values, it triggers a model retraining. We leverage 𝑓 to automatically label these samples before using them to retrain the model. The results of the retraining are presented in Fig. 11, where the two models (original and re-trained) are compared over the second part of the data samples (from 11th of September). For all operators, Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. 35:18 Mattia Milani et al. the benefits of promptly identifying changes in data patterns and triggering countermeasures are evident. The re-trained model exhibits higher confidence scores, enabling more certain decisions that help avoid false alarms, or even more critically, false negatives that might overlook outages. Moreover, it is important to note that cases such as MNO 714 are not affected by the retraining of Faro: the model was already able to classify its data samples with high confidence and continues to do so after retraining. This confirms that by selecting for retraining only those data samples for which Faro reports low confidence, we preserve the model’s correct operation while improving its performance where needed. Still, in the worst case scenario (i.e., when a re-training procedure does not improve the confidence on subsequent evaluations), the load imposed by this procedure is affordable. In our dataset, the median interval (recorded for each MNO) between two successive periods of low confidence (that would trigger a re-training) is 1297 min (almost 22 hours), with the default settings used in Faro, and can be further reduced (at the expense of precision) by decreasing the threshold value to, e.g., 𝜎(𝛾𝑢)=0.175 increases the median to 1838 min . Hence, in the worst possible scenario, operators need to retrain Faro roughly once per day. 4.5 Faro resource footprint We trained and executed Faro on a COTS server equipped with an NVIDIA GeForce GTX 1070 GPU with 8 GB of internal memory, a setup that was enough to train the model used in Faro in 250 minutes on average. Faro also has a limited memory footprint as the model only uses 438.81KB with 112 336 parameters and can be executed in inference even on a CPU. In our GPU setup, an inference pass on a set of 130 000 elements takes on average 4.45 seconds. Moreover, the proposed clustering approach also improves the inference time compared to a standard triplet loss approach, where positive and negative references are randomly selected. By using the cluster centroids, inference times can be lowered to 2.25 seconds. 5 Conclusion In this paper, we proposed Faro a self-supervised solution for the timely detection of outages for an MNA scenario. Leveraging on contrastive learning techniques, Faro has been designed to promptly identify outage events: when labels are available, Faro can correctly identify 93.3 % of the outages within 10 min and 100 % in 15 min , much less than the time usually required for a ticketing system. We also evaluated Faro in the wild using an unlabelled dataset where only coarse indications for outage events were provided: all of them could be identified, while there is strong evidence that other detected events were also happening in reality. Finally, Faro also provides a confidence score associated with its operation that can be used by human operators to validate the solution while in production and to automatically trigger the re-training of the model, further improving the adaptability of the solution. Acknowledgements This work has been supported by the ORIGAMI project, funded by the Smart Networks and Services Joint Undertaking (SNS JU) under the European Union’s Horizon Europe research and innovation programme with Grant Agreement No. 101139270. It has also been supported by the Spanish Ministry of Digital Transformation through the UNICO I+D 5G project 6G-RIEMANN, by the Spanish Ministry of Science, Innovation and Universities through the project 6GINSPIRE (PID2022-137329OB-C42), and by the Region of Madrid through the project TUCAN6-CM (TEC-2024/COM-460). Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. Faro: a scalable and reliable outage detection algorithm for IoT Mobile Virtual Network Aggregators 35:19 References [1] 2025. Airalo. https://www.airalo.com/. Accessed: 2025-10-01. [2] 2025. Emnify. https://www.emnify.com/. Accessed: 2025-10-01. [3] 2025. HolaFly. https://esim.holafly.com/. Accessed: 2025-10-01. [4] Adel Abusitta, Glaucio H.S. de Carvalho, Omar Abdel Wahab, Talal Halabi, Benjamin C.M. Fung, and Saja Al Mamoori. 2023. Deep learning-enabled anomaly detection for IoT systems. Internet of Things 21 (2023), 100656. doi:10.1016/j.iot.2022.100656 [5] Sergi Alcalá-Marín, Aravindh Raman, Weili Wu, Andra Lutu, Marcelo Bagnulo, Ozgu Alay, and Fabián Bustamante. 2022. Global mobile network aggregators: taxonomy, roaming performance and optimization. In Proceedings of the 20th Annual International Conference on Mobile Systems, Applications and Services (Portland, Oregon) (MobiSys ’22). Association for Computing Machinery, New York, NY, USA, 183–195. doi:10.1145/3498361.3538942 [6] Theyab Althiyabi, Iftikhar Ahmad, and Madini O. Alassafi. 2024. Enhancing IoT Security: A Few-Shot Learning Approach for Intrusion Detection. Mathematics 12, 7 (2024). doi:10.3390/math12071055 [7] Julien Audibert, Pietro Michiardi, Frédéric Guyard, Sébastien Marti, and Maria A. Zuluaga. 2020. USAD: UnSupervised Anomaly Detection on Multivariate Time Series. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (Virtual Event, CA, USA) (KDD ’20). Association for Computing Machinery, New York, NY, USA, 3395–3404. doi:10.1145/3394486.3403392 [8] Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009. Curriculum learning. In Proceedings of the 26th Annual International Conference on Machine Learning (Montreal, Quebec, Canada) (ICML ’09). Association for Computing Machinery, New York, NY, USA, 41–48. doi:10.1145/1553374.1553380 [9] Andrea Bommert, Xudong Sun, Bernd Bischl, Jörg Rahnenführer, and Michel Lang. 2020. Benchmark for filter methods for feature selection in high-dimensional classification data. Computational Statistics & Data Analysis 143 (2020), 106839. [10] Ricardo J. G. B. Campello, Davoud Moulavi, and Joerg Sander. 2013. Density-Based Clustering Based on Hierarchical Density Estimates. In Advances in Knowledge Discovery and Data Mining, Jian Pei, Vincent S. Tseng, Longbing Cao, Hiroshi Motoda, and Guandong Xu (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 160–172. [11] Ayan Chatterjee and Bestoun S. Ahmed. 2022. IoT anomaly detection methods and applications: A survey. Internet of Things 19 (2022), 100568. doi:10.1016/j.iot.2022.100568 [12] Tianqi Chen and Carlos Guestrin. 2016. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (San Francisco, California, USA) (KDD ’16). Association for Computing Machinery, New York, NY, USA, 785–794. doi:10.1145/2939672.2939785 [13] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A Simple Framework for Contrastive Learning of Visual Representations. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119), Hal Daumé III and Aarti Singh (Eds.). PmLR, PMLR, Virtual Event, 1597–1607. https://proceedings.mlr.press/v119/chen20j.html [14] Davide Chicco. 2021. Siamese Neural Networks: An Overview. (2021), 73–94. doi:10.1007/978-1-0716-0826-5_3 [15] S. Chopra, R. Hadsell, and Y. LeCun. 2005. Learning a similarity metric discriminatively, with application to face verification. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), Vol. 1. IEEE, IEEE, San Diego, CA, USA, 539–546 vol. 1. doi:10.1109/CVPR.2005.202 [16] Andrew A. Cook, Göksel Mısırlı, and Zhong Fan. 2020. Anomaly Detection for IoT Time-Series Data: A Survey. IEEE Internet of Things Journal 7, 7 (2020), 6481–6494. doi:10.1109/JIOT.2019.2958185 [17] Xingping Dong and Jianbing Shen. 2018. Triplet loss in siamese network for object tracking. In Proceedings of the European conference on computer vision (ECCV). Munich, Germany, 459–474. [18] Antonio-Javier Gallego, Jorge Calvo-Zaragoza, and Robert B Fisher. 2020. Incremental unsupervised domain-adversarial training of neural networks. IEEE Transactions on Neural Networks and Learning Systems 32, 11 (2020), 4864–4878. [19] Weifeng Ge, Weilin Huang, Dengke Dong, and Matthew R. Scott. 2018. Deep Metric Learning with Hierarchical Triplet Loss. In Proceedings of the European conference on computer vision (ECCV). Munich, Germany, 269–285. [20] Stefan Geissler, Florian Wamser, Wolfgang Bauer, Michael Krolikowski, Steffen Gebert, and Tobias Hoßfeld. 2021. Signaling Traffic in Internet-of-Things Mobile Networks. In 2021 IFIP/IEEE International Symposium on Integrated Network Management (IM). 452–458. [21] Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Momentum Contrast for Unsupervised Visual Representation Learning. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Seattle, WA, USA, 9726–9735. doi:10.1109/CVPR42600.2020.00975 [22] Tin Kam Ho. 1995. Random decision forests. In Proceedings of 3rd international conference on document analysis and recognition, Vol. 1. IEEE, 278–282. [23] Haigen Hu, Xiaoyuan Wang, Yan Zhang, Qi Chen, and Qiu Guan. 2024. A comprehensive survey on contrastive learning. Neurocomputing 610 (2024), 128645. doi:10.1016/j.neucom.2024.128645 Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. 35:20 Mattia Milani et al. [24] Ashish Jaiswal, Ashwin Ramesh Babu, Mohammad Zaki Zadeh, Debapriya Banerjee, and Fillia Makedon. 2021. A Survey on Contrastive Self-Supervised Learning. Technologies 9, 1 (2021), 2. doi:10.3390/technologies9010002 [25] Keith Kendig. 2000. Is a 2000-year-old formula still keeping some secrets? The American Mathematical Monthly 107, 5 (2000), 402–415. [26] Scott M. Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17). Curran Associates Inc., Red Hook, NY, USA, 4768–4777. [27] Leland McInnes, John Healy, and James Melville. 2020. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. (2020). arXiv:1802.03426 [stat.ML] https://arxiv.org/abs/1802.03426 [28] Mattia Milani, Dario Bega, Marco Gramaglia, Pablo Serrano, and Christian Mannweiler. 2024. ATELIER: Service Tailored and Limited-Trust Network Analytics Using Cooperative Learning. IEEE Open Journal of the Communications Society 5 (2024), 3315–3330. doi:10.1109/OJCOMS.2024.3401746 [29] Mohsin Munir, Shoaib Ahmed Siddiqui, Andreas Dengel, and Sheraz Ahmed. 2019. DeepAnT: A Deep Learning Approach for Unsupervised Anomaly Detection in Time Series. IEEE Access 7 (2019), 1991–2005. doi:10.1109/ ACCESS.2018.2886457 [30] Hussain Nizam, Samra Zafar, Zefeng Lv, Fan Wang, and Xiaopeng Hu. 2022. Real-Time Deep Anomaly Detection Framework for Multivariate Time-Series Data in Industrial IoT. IEEE Sensors Journal 22, 23 (2022), 22836–22849. doi:10.1109/JSEN.2022.3211874 [31] Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. "Why Should I Trust You?": Explaining the Predictions of Any Classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (San Francisco, California, USA) (KDD ’16). Association for Computing Machinery, New York, NY, USA, 1135–1144. doi:10.1145/2939672.2939778 [32] Peter J. Rousseeuw. 1987. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. J. Comput. Appl. Math. 20 (1987), 53–65. doi:10.1016/0377-0427(87)90125-7 [33] Stefania Russo, Moritz Lürig, Wenjin Hao, Blake Matthews, and Kris Villez. 2020. Active learning for anomaly detection in environmental data. Environmental Modelling & Software 134 (2020), 104869. doi:10.1016/j.envsoft.2020.104869 [34] Florian Schroff, Dmitry Kalenichenko, and James Philbin. 2015. FaceNet: A unified embedding for face recognition and clustering. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Boston, MA, USA, 815–823. doi:10.1109/CVPR.2015.7298682 [35] Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin A Raffel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li. 2020. FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence. 33 (2020), 596–608. https://proceedings.neurips.cc/paper_files/paper/2020/file/06964dce9addb1c5cb5d6e3d9838f733Paper.pdf [36] Yaniv Taigman, Ming Yang, Marc’Aurelio Ranzato, and Lior Wolf. 2014. DeepFace: Closing the Gap to Human-Level Performance in Face Verification. In 2014 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, Columbus, OH, USA, 1701–1708. doi:10.1109/CVPR.2014.220 [37] Markus Thill, Wolfgang Konen, Hao Wang, and Thomas Bäck. 2021. Temporal convolutional autoencoder for unsupervised anomaly detection in time series. Applied Soft Computing 112 (2021), 107751. doi:10.1016/j.asoc.2021.107751 [38] Enkhtur Tsogbaatar, Monowar H. Bhuyan, Yuzo Taenaka, Doudou Fall, Khishigjargal Gonchigsumlaa, Erik Elmroth, and Youki Kadobayashi. 2021. DeL-IoT: A deep ensemble learning approach to uncover anomalies in IoT. Internet of Things 14 (2021), 100391. doi:10.1016/j.iot.2021.100391 [39] Viktoria Vomhoff, Stefan Geissler, Frank Loh, Wolfgang Bauer, and Tobias Hossfeld. 2022. Characterizing Mobile Signaling Anomalies in the Internet-of-Things. In NOMS 2022-2022 IEEE/IFIP Network Operations and Management Symposium. 1–6. doi:10.1109/NOMS54207.2022.9789902 [40] Kilian Q. Weinberger and Lawrence K. Saul. 2009. Distance Metric Learning for Large Margin Nearest Neighbor Classification. J. Mach. Learn. Res. 10, 2 (June 2009), 207–244. [41] Yi Xu, Lei Shang, Jinxing Ye, Qi Qian, Yu-Feng Li, Baigui Sun, Hao Li, and Rong Jin. 2021. Dash: Semi-Supervised Learning with Dynamic Thresholding. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139), Marina Meila and Tong Zhang (Eds.). PMLR, Virtual event, 11525–11536. https://proceedings.mlr.press/v139/xu21e.html [42] Li Yang, Ying Li, Jin Wang, and Neal N. Xiong. 2021. FSLM: An Intelligent Few-Shot Learning Model Based on Siamese Networks for IoT Technology. IEEE Internet of Things Journal 8, 12 (2021), 9717–9729. doi:10.1109/JIOT.2020.3022427 [43] Zengwen Yuan, Qianru Li, Yuanjie Li, Songwu Lu, Chunyi Peng, and George Varghese. 2018. Resolving Policy Conflicts in Multi-Carrier Cellular Access. In Proceedings of the 24th Annual International Conference on Mobile Computing and Networking (New Delhi, India) (MobiCom ’18). Association for Computing Machinery, New York, NY, USA, 147–162. doi:10.1145/3241539.3241558 Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. Faro: a scalable and reliable outage detection algorithm for IoT Mobile Virtual Network Aggregators 35:21 [44] Mohammed Zakariah and Abdulaziz S. Almazyad. 2023. Anomaly Detection for IOT Systems Using Active Learning. Applied Sciences 13, 21 (2023). doi:10.3390/app132112029 [45] Zahra Zamanzadeh Darban, Geoffrey I. Webb, Shirui Pan, Charu Aggarwal, and Mahsa Salehi. 2024. Deep Learning for Time Series Anomaly Detection: A Survey. ACM Comput. Surv. 57, 1, Article 15 (Oct. 2024), 42 pages. doi:10.1145/3691338 [46] Chuxu Zhang, Dongjin Song, Yuncong Chen, Xinyang Feng, Cristian Lumezanu, Wei Cheng, Jingchao Ni, Bo Zong, Haifeng Chen, and Nitesh V Chawla. 2019. A deep neural network for unsupervised anomaly detection and diagnosis in multivariate time series data. In Proceedings of the AAAI conference on artificial intelligence, Vol. 33. Hilton Hawaiian Village, Honolulu, Hawaii, USA, 1409–1416. [47] Xiaokang Zhou, Wei Liang, Shohei Shimizu, Jianhua Ma, and Qun Jin. 2021. Siamese Neural Network Based Few-Shot Learning for Anomaly Detection in Industrial Cyber-Physical Systems. IEEE Transactions on Industrial Informatics 17, 8 (2021), 5790–5798. doi:10.1109/TII.2020.3047675 A Data Details All the input features used by Faro contrastive learning solutions are reported in Table 1. The features have been ordered depending on the importance computed using the Analysis of Variance F-Score [ 9 ] from the most important to the lowest. This technique performs an analysis of variance of the features isolated by the given class. A high F-Score indicates a significant difference between the given classes on the same feature; therefore, such features are more useful to discern between the classes. Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. 35:22 Mattia Milani et al. Table 1. Features used by Faro. ID Feature Name F-Score Type Description 1medcross_n_cr_pdp 23 019.62 roll # times n_cr_pdp crossed the median computed over the window 2medcross_n_cr_pdp_success 23 019.62 roll # times n_cr_pdp_success crossed the median computed over the window 3medcross_n_ul_success 8989.14 roll # times n_ul_success crossed the median computed over the window 4entropy_n_ul_alert_operator_not_supported 7825.72 roll Shannon entropy of n_ul_alert_operator_not_supported computed over the window 5spec_entr_n_ul_gprs 7819.26 roll Spectral entropy of n_ul_gprs computed over the window 6medcross_n_ul_gprs_success 7663.68 roll # times n_ul_gprs_success crossed the median computed over the window 7entropy_n_ul_gprs_alert_operator_not_supported 7456.31 roll Shannon entropy of n_ul_gprs_alert_operator_not_supported computed over the window 8quan95_n_ul 5630.24 roll 0.95 quantile of n_ul computed over the window 9spec_entr_n_ul_gprs_success 5412.73 roll Spectral entropy of n_ul_gprs_success computed over the window 10 quan95_n_ul_gprs 4626.28 roll 0.95 quantile of n_ul_gprs computed over the window 11 tot_connline_n_ul_gprs 4329.22 roll Slope of the line connecting first and last point of n_ul_gprs in the window 12 medcross_n_ul 4252.45 roll # times n_ul crossed the median computed over the window 13 quan99_n_ul_gprs 3987.54 roll 0.99 quantile of n_ul_gprs computed over the window 14 quan99_n_ul_gprs_success 3731.38 roll 0.99 quantile of n_ul_gprs_success computed over the window 15 tot_connline_n_ul 3575.90 roll Slope of the line connecting first and last point of n_ul in the window 16 tot_energy_n_ul 3454.03 roll Total energy of n_ul computed over the window 17 max_n_ul_gprs 3314.20 roll Max value of n_ul_gprs computed over the window 18 quan_99_n_ul 3303.45 roll 0.99 quantile of n_ul computed over the window 19 std_n_ul_gprs 3293.11 roll Standard deviation computed on n_ul_gprs over the window 20 quan95_n_ul_success 3217.41 roll 0.95 quantile of n_ul_success computed over the window 21 tot_connline_n_ul_gprs_success 3164.77 roll Slope of the line connecting first and last point of n_ul_gprs_success in the window 22 max_n_ul_gprs_success 3142.80 roll Max value of n_ul_gprs_success computed over the window 23 spec_entr_n_ul 3063.38 roll Spectral entropy of n_ul computed over the window 24 meddiffmin_n_ul 2951.64 roll Distance between median and minimum of n_ul over the window 25 mean_n_ul 2841.47 roll Mean computed on n_ul over the window 26 meddiffmax_n_ul_gprs 2571.70 roll Distance between median and maximum of n_ul_gprs over the window 27 120_mean_n_ul 2506.06 lag Value of mean_n_ul 120 min in the past 28 n_ul 2443.47 Count Current value of n_ul 29 medquan90_n_ul 2361.70 roll Distance of median and 0.9quantile of n_ul over the window 30 max_n_ul 2273.69 roll Max value of n_ul computed over the window 31 std_n_ul 2273.39 roll Standard deviation computed on n_ul over the window 32 median_n_ul 2273.04 roll Median computed over n_ul during the window 33 quan99_n_ul_success 1936.46 roll 0.99 quantile of n_ul_success computed over the window 34 spec_entr_n_ul_success 1690.64 roll Spectral entropy of n_ul_success computed over the window 35 std_n_ul_success 1613.96 roll Standard deviation computed on n_ul_success over the window 36 meddiffmax_n_ul 1562.47 roll Distance between median and maximum of n_ul over the window 37 max_n_ul_success 1478.65 roll Max value of n_ul_success computed over the window 38 meddiffmax_n_ul_success 1304.18 roll Distance between median and maximum of n_ul_success over the window 39 medcross_n_ul_gprs 1127.77 roll # times n_ul_success crossed the median computed over the window 40 day 10.78 Number Day of the week as number Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025. Faro: a scalable and reliable outage detection algorithm for IoT Mobile Virtual Network Aggregators 35:23 Table 2. Main features characterising user plane and control plane related outages. Outage Category Most Relevant Features User plane traffic anomalies – Uplink variability: n_ul – PDP stability: medcross_n_cr_pdp_success,std_n_ul_gprs – Quantiles of UL usage: medquan90_n_ul,quan95_n_ul_gprs – Temporal dynamics: mean_n_ul Control plane traffic anomalies – Alert entropy: entropy_n_ul_alert_operator_not_supported – GPRS alert entropy: entropy_n_ul_gprs_alert_operator_not_supported – Spectral entropy of GPRS success: spec_entr_n_ul_gprs_success – Extreme quantiles: quan95/99_n_ul_gprs,quan99_n_ul – Maximal successful location updates: max_n_ul_success B Faro explainability The methodology discussed in Section 4.3 allowed us to dissect the features that characterise two main kinds of outages in our dataset: the ones related to Control Plane metrics and the ones related to User Plane metrics. Table 2 shows the most important features for each category. Received 5 June 2025; revised 6 October 2025; accepted 8 October 2025 Proc. ACM Netw., Vol. 3, No. CoNEXT4, Article 35. Publication date: December 2025.