Full text
1 Trust Score Prediction for IoT Device Onboarding Using Transfer and Few-Shot Learning in Consumer Electronics Ilias Politis1, Michail Bampatsikos2, Apostolis Zarras2and Christos Xenakis2 1Industrial Systems Institute, ATHENA Research Center, Patras, Greece 2Systems Security Laboratory, University of Piraeus, Greece Abstract—The rapid proliferation of Internet of Things (IoT) devices in consumer electronics has made efficient trust score prediction essential for secure device onboarding. This paper presents a hybrid Trust Management framework that integrates few-shot and transfer learning with a statistical Markov chain foundation to address data scarcity and adaptability challenges in dynamic IoT environments. The few-shot learning phase enables rapid adaptation from minimal data (as few as 5–20 onboarding samples), while transfer learning ensures robust cross-domain generalizability (e.g., from consumer to industrial IoT). Comprehensive evaluation demonstrates that, with 20 onboarding samples, the proposed approach achieves fine-tuned Mean Squared Error (MSE) as low as 4.45 (XGBoost) and 5.07 (Random Forest), R2scores exceeding 0.96, and average prediction error (MAE) below 1.55. Batched distributed ledger operations reduce total onboarding latency to under 750 ms for five devices, with system throughput averaging 137.7 devices per second across models. These results show the framework delivers accurate, low-latency, and scalable trust score predictions suitable for real-time onboarding in evolving IoT environments. Index Terms—IoT, Trust Score Prediction, Transfer Learning, Few-Shot Learning, Device Onboarding, Real-Time Prediction I. INTRODUCTION The rapid proliferation of Internet of Things (IoT) devices in consumer electronics has introduced significant challenges in ensuring secure and reliable device interactions, particularly during onboarding in dynamic, data-constrained environments. A robust Trust Management (TM) layer is indispensable for rapidly gauging a newcomer’s reliability, shielding the network from data-poisoning and zero-day threats, and achieving these objectives within the strict data and resource constraints typical of consumer-grade IoT. However, existing TM approaches present notable limitations. Statistical models, such as our prior Markov chain-based Multi Attribute Decision Making (MADM) framework [1], are robust and resourceefficient but exhibit limited adaptability to rapidly evolving threats and new attack vectors. Conversely, purely Machine Learning (ML)-based methods [2] often require large training datasets, limiting their practicality in real-time onboarding scenarios—such as in smart homes, healthcare wearables, and industrial IoT—where only minimal data is available for new devices. This motivates the need for a hybrid approach that preserves the efficiency and interpretability of statistical models while incorporating the adaptability of ML techniques under extreme data scarcity. The central problem addressed in this paper is therefore: how to reliably predict the trustworthiness of a new IoT device during its onboarding process when only a handful of labelled samples (5–20) are available, while meeting strict latency budgets (sub-500 ms) and ensuring resilience against adversarial manipulation. Solving this problem requires not only accurate trust score predictions under limited data but also system-level scalability to handle large device populations in real time. To address these challenges, this paper proposes a novel hybrid TM framework that integrates: (i) Few Shot Learning (FSL) for rapid adaptation from minimal onboarding data; (ii) transfer learning for cross-domain generalisation (e.g., consumer-to-industrial IoT); and (iii) a foundational Markov chain-based statistical model for interpretable and deterministic trust estimation. The novelty of our approach lies in this explicit integration: unlike prior TM solutions, we combine the statistical robustness of a Markov/MADM trust function with the adaptability of ML predictors and enhance it further with a mathematically justified blending strategy between base and fine-tuned models. This makes the framework resilient to poisoning of scarce onboarding samples while ensuring adaptability to device-specific behaviours. Our experimental evaluation, using Random Forest, k-NN, and XGBoost as trust predictors, demonstrates that the proposed framework achieves fine-tuned Mean Squared Error (MSE) as low as 4.45 (XGBoost) and 5.07 (Random Forest), R2scores exceeding 0.96, and Mean Absolute Error (MAE) below 1.55 with 20 onboarding samples. The approach achieves system throughput up to 175 devices per second (kNN) and consistently averages above 137 devices/s across models. Onboarding latency is reduced to under 750 ms for five devices using batched distributed ledger (DLT) operations. To further improve prediction accuracy and adaptability under extreme data scarcity, we employ a mathematically justified blending of the base and fine-tuned models, with optimal weights (α) empirically determined via grid search to minimise validation error. This blended prediction strategy ensures both robustness and rapid domain adaptation, as validated by detailed ablation and sensitivity analyses. To our knowledge, this is the first TM framework to explicitly integrate few-shot learning, transfer learning, and a Markov chain-based trust estimation model for secure IoT device onboarding. The contributions of this paper can be summarised as follows: •Problem formulation: we formally define the challenge of trust score prediction under scarce onboarding data and
2 strict latency constraints, a gap not directly addressed in existing TM research. •Novel hybrid framework: we introduce the first TM framework that fuses statistical Markov-based modelling, transfer learning, and few-shot adaptation, augmented with a principled blending mechanism. •Scalable, low-latency evaluation: we demonstrate experimentally that the framework achieves high accuracy, subsecond onboarding, and throughput exceeding 137 devices/s, validating its applicability to real-world consumer electronics. The remainder of the paper is organised as follows. Section II reviews related work on TM in IoT. Section III introduces the algorithmic design and logic. Section IV presents performance evaluation using real-world use cases. Section V discusses comparative advantages and security benefits. Section VI concludes the paper and outlines future research directions. II. LITERATURE REVIEW This literature review surveys key TM approaches in IoT networks, focusing on methodologies used to compute device trust scores based on factors such as Quality of Service (QoS), delays, security, and feedback. By examining existing solutions, we underscore how the proposed hybrid framework advances the state of the art. Special emphasis is placed on our prior work leveraging a Markov chain-based method [1], a critical foundational element in this research. A. Machine Learning Approaches in Trust Management ML techniques have been extensively employed in IoT trust management. For instance, Shayesteh et al. [3] apply Bayesian learning and Dempster-Shafer Theory to derive trust scores from entity and data reliability. Alghofaili and Rassam [4] present a Multi-Criteria Decision-Making approach integrated with a Deep Long Short-Term Memory (LSTM) model to evaluate trust through packet loss and throughput metrics, weighted by Shannon’s entropy. Researchers have also proposed a Federated Learning-based TM framework for Industrial IoT, where trust is calculated via a weighted sum [5]. Furthermore, Wang et al. [6] employ unsupervised learning and clustering to differentiate trustworthy vehicles in IoV, using Direct Trust and Indirect Trust. Meanwhile, Ma et al. [7] leverage LSTM to predict device behaviour by incorporating QoS, a measure of network performance, delays, security, and feedback, and time-dependent features. Despite providing valuable insights, these ML-based methods often depend on large datasets, limiting their effectiveness in data-scarce IoT settings. In addition, they tend to focus on a narrow range of metrics, reducing their broader applicability. B. Non-Machine Learning Approaches in Trust Management Non-ML TM methods rely on statistical or structural techniques. Alam et al. [2] propose a TM method that calculates trust scores using QoS and cooperation metrics. Latif [8] introduces a context-dependent approach for social IoT, integrating device capabilities and satisfaction. Bampatsikos et al. [1] present a two-dimensional Markov chain model, combined with MADM and a piecewise function, to predict trust evolution based on cyber risk, packet loss, and device utilisation. Bampatsikos et al. [9] also explore this approach with Hyperledger Fabric. This offers a strong statistical foundation but struggles with dynamic threats due to its static nature. Other works, such as Liu et al. [10] (i.e., Hidden Markov Model for VANETs) and Bai et al. [11] (i.e., game theory in supply chains), provide stability yet lack adaptability to new or evolving scenarios. C. Beyond the State of the Art The reviewed literature highlights limitations in both MLand non-ML-based TM approaches. ML methods require large training datasets, potentially diminishing efficiency [7], whereas non-ML approaches tend to be less adaptable to emerging threats [1]. Our prior Markov-based TM method provides a robust baseline by employing statistical modelling of key parameters such as cyber risk and QoS, but its predefined rules impede responsiveness in dynamic IoT environments [1]. The proposed hybrid framework addresses these issues by integrating transfer learning and FSL. Transfer learning leverages pre-trained models on a synthetic Markov chain dataset, ensuring efficiency and enabling cross-domain adaptability (e.g., from consumer to industrial IoT). Meanwhile, FSL enhances this framework by requiring minimal data (only five samples), facilitating rapid trust assessment (e.g., XGBoost inference at 0.9425 ms) with an average prediction error of 2.29. Unlike resource-intensive ML methods, the solution avoids large dataset dependencies, and unlike purely statistical models, it remains adaptable to evolving threats. Moreover, the statistical ground truth mitigates poisoning attacks by building upon the Markov foundation while resolving its inherent limitations. Consequently, the proposed framework significantly advances consumer IoT security and scalability. III. ALGORITHM DESIGN AND LOGIC The proposed framework for trust score prediction in IoT device onboarding leverages transfer learning and few-shot learning techniques to address the pervasive challenge of sparse data in consumer electronics environments. Building on a previous work which employed a 2D Markov chain model to simulate trust dynamics [1], this study introduces a novel hybrid ML and statistical approach for predicting trust scores in new IoT devices. By enabling secure, real-time onboarding in resource-constrained consumer IoT contexts, such as wearable health devices and autonomous vehicle systems, this method ensures the rapid trust assessment necessary for operational safety and security. This section elucidates the conceptual design logic of the proposed algorithm and provides a mathematically grounded description of its implementation, including feature engineering, model training, and prediction blending. A. Design Logic The proposed trust score prediction algorithm is driven by three primary challenges in consumer IoT device onboarding:
3 (i)the scarcity of labeled data for new devices, (ii)the need for rapid trust assessment to facilitate real-time decisionmaking, and (iii)the requirement for resilience against security threats (e.g., data poisoning, zero-day attacks). To address these issues, the algorithm employs a two-phase structure that combines transfer learning (to exploit prior knowledge) with FSL (to adapt efficiently to limited onboarding data). 1) Few-Shot Transfer Learning: Concept and Rationale: Few-shot transfer learning synthesises the benefits of transfer learning and FSL to achieve effective and flexible trust score prediction. Transfer learning relies on a pre-trained model, initially trained on a large, heterogeneous dataset. We utilise synthetic data from a 2D Markov chain model [1], providing a robust foundation for trust prediction. Few-shot learning then fine-tunes this model using a small number of onboarding samples (e.g., 5 samples) to adapt to the specific characteristics of new IoT devices. This hybrid approach ensures that the model can generalise from substantial prior knowledge while adapting rapidly to sparse, real-time data, a critical requirement for consumer electronics applications. Several factors motivated this approach. First, consumer IoT onboarding faces severe data limitations from resource constraints, privacy regulations, or the novelty of emerging devices (e.g., next-generation health wearables or autonomous vehicle modules). Few-shot learning addresses this issue by requiring only a small number of samples, thereby mitigating data collection overhead. Second, real-time operation in consumer IoT environments demands swift trust evaluation. Transfer learning accelerates this process by providing a robust starting point by reducing training time, while few-shot finetuning ensures rapid adaptation without requiring comprehensive retraining. Third, the approach reinforces security and robustness by leveraging a large, diverse training corpus. This pre-training promotes resilience against noise and attacks, while few-shot fine-tuning tailors the model to specific device attributes, mitigating overfitting risks on small datasets. This methodology yields multiple advantages. It offers efficiency by reducing data and computational demands; a pivotal factor for resource-constrained IoT devices. It provides adaptability by enabling the model to accommodate new devices using minimal samples; an important capability for dynamic consumer IoT ecosystems. It also enhances accuracy and stability by employing a blended prediction strategy, achieving high performance (e.g., a fine-tuned Random Forest (RF) yields R2= 0.8891). Finally, it enhances security via reliance on statistical ground truth and blending predictions, reducing data poisoning in few-shot samples, a common security concern in IoT contexts. B. Algorithm Description The implementation of the algorithm formalises this design logic in mathematical terms. The framework proceeds in three stages: transfer learning using a Markov-based statistical ground truth, few-shot fine-tuning with limited onboarding data, and blended prediction for robust decision-making. Figure 1 illustrates the workflow, while Algorithm 1 provides pseudocode. The mathematical representation of each stage is presented below. Load Prior Data (Dprior) Feature Engineering (Generate X, Calc y, Normalize) Transfer Learning (Train Mbase, Evaluate) Few-Shot Fine-Tuning (Collect Donboard, Fine-Tune Mfine) Predict for New Device (Preprocess Xnew, Blend) End Offline transfer learning On-device few-shot fine-tuning Fig. 1. End-to-end workflow of the proposed hybrid trust-score engine. The upper band (grey) represents offline transfer learning; the lower band (blue) depicts on-device few-shot fine-tuning. Their convergence shows how each inference blends long-term knowledge with real-time adaptation. The final stage includes latency evaluation (Eq. 10), where onboarding delay is decomposed into identity, communication, ledger, and inference components to ensure compliance with the 500 ms real-time target. 1) Statistical Ground Truth and Feature Representation: Each IoT device iis described by a feature vector xi={Ci, Ri, Si, Pi, RSi, RepSi, OSi, P kgi},(1) where Ci, Ridenote CPU and RAM utilisation, Siis the security level, Pithe packet loss ratio, RSi, RepSithe risk and reputation states, and (OSi, Pkgi)the inverted ages of the operating system and installed packages. The problem is to learn a regression function f:R8→[0,100], f(xi)≈Ti, where Tiis the trust score that quantifies the device’s reliability during onboarding. a) Ground Truth Trust Score: The statistical trust score is generated via the piecewise function: T= max(0,min(100,100 −20R−10C+ +4S+ 2OSi+ 1Pi−10P)),if RS ≥4 min(100,80 −20R−10C+ +4S+ 2OSi+ 1Pi−10P),if RS ≤2 max(20,min(100,70 −20R−10C+ +4S+ 2OSi+ 1Pi−10P)),otherwise (2) Equation 2 balances negative contributions from utilisation and packet loss with positive contributions from security and freshness. The coefficients (e.g., 20, 10, 4) are empirically derived from domain knowledge and prior analysis [1].
4 Algorithm 1 IoT Device Onboarding Trust Score Prediction 1: Input: Pre-generated Markov chain data (Dprior), onboarding samples (Donboard), device features (Xnew) 2: Output: Predicted trust score (Tfinal) for a new device 3: Load Prior Data: Load Dprior = {RepState, RiskState, TrustScores}from CSV files. 4: Feature Engineering: 5: a. Generate synthetic features X= {C, R, S, P, RS, RepS, OSi, Pi}. 6: b. Calculate trust scores yusing Eq. 2. 7: c. Normalize Xusing MinMaxScaler to obtain Xscaled. 8: Transfer Learning (Base Training): 9: a. Split Xscaled, y into training (80%) and test (20%) sets. 10: b. Train base models Mbase ={RF, k-NN, XGBoost}by minimising Eq. 3. 11: c. Evaluate Mbase on test set, computing metrics (MSE, R2, MAE, MAPE, RMSE, inference time). 12: Few-Shot Fine-Tuning: 13: a. Collect Donboard ={Xonboard, yonboard}with Ksamples. 14: b. Impute missing features in Xonboard with training set means, normalize to Xonboard,scaled. 15: c. Fine-tune models Mfine = {RFfine, k-NNfine, XGBoostfine}by minimising Eq. 5. 16: d. Blend predictions with Eq. 6: Tfinal =α·Mbase(Xtest) + (1 −α)·Mfine(Xtest). 17: e. Evaluate blended predictions on test set, computing metrics. 18: Prediction for New Device: 19: a. Preprocess Xnew: impute missing features, normalize to Xnew,scaled. 20: b. Predict Tbase =Mbase(Xnew,scaled),Tfine = Mfine(Xnew,scaled). 21: c. Compute Tfinal =α·Tbase + (1 −α)·Tfine per Eq. 6. 22: d. Return Tfinal. 23: System Latency Evaluation: 24: Compute end-to-end onboarding latency Laccording to Eq. 10, decomposing identity (LSSI), messaging (LMQT T ), ledger (LDLT ), and inference (Linf ) components. Importantly, Eq. 2 also encodes domain knowledge by penalising resource saturation (Ci,Ri) and packet loss (Pi) with relatively high weights, since these directly degrade device reliability. Positive contributions are drawn from security level (Si) and software freshness (OSi,Pkgi), which improve resilience and reduce vulnerability. The bounding functions (min,max) ensure that scores remain within [0,100], preventing extreme fluctuations due to any single feature. This design reflects practical IoT considerations (i.e., no device is entirely untrustworthy if its core behaviour is stable, and no device is fully trustworthy if critical resources are exhausted). b) Markov Dynamics of States: As it was detailed in [1], the risk (RS) and reputation (RepS) states evolve over time according to a two-dimensional Markov chain model [1]. Let p(t)denote the joint distribution of (RS, RepS)at time t. The state transitions follow: p(t+1) =p(t)P, where Pis the joint transition matrix. This captures temporal correlations in device behaviour, e.g., a device with declining reputation is more likely to remain in low-trust states, while a device with strong historical reputation tends to persist in higher-trust states. c) Feature Sampling: The remaining features are stochastically sampled to reflect heterogeneous device characteristics. CPU and RAM utilisation are drawn from truncated uniform distributions, Ci∼ U(0.2,0.8), Ri∼ U(0.1,0.9), security level is sampled from a categorical distribution biased towards medium security, Si∼Categorical({1,2,3,4,5}, pS), pS= (0.2,0.5,0.3), packet loss follows a Beta distribution skewed towards lowloss regimes, Pi∼Beta(α= 2, β = 8), and OS/package ages are sampled uniformly as inverted values to reflect freshness: OSi, Pkgi∼ U{1,2,3,4,5}. d) Synthetic Dataset.: The labelled dataset is then D={(xi, Ti)}N i=1, N = 1000, with Ticomputed from Eq. 2. This setup encodes domain knowledge such as resource saturation (Ci, Ri) and packet loss (Pi) receive heavy penalties, while higher security levels (Si) and updated software (OSi, Pkgi) improve trust. The resulting dataset provides both trusted and untrusted devices in realistic proportions (approximately 64% trusted, 36% untrusted), ensuring balanced training for transfer learning and robust evaluation. 2) Transfer Learning Phase:The base models are trained on Dby minimising mean squared error: fb= arg min f∈F 1 N N X i=1 (f(xi)−Ti)2,(3) where fb∈ {fRF, fkNN, fXGB}. RF provides robustness, kNN simplicity, and XGBoost high inference efficiency. These models encode transferable priors. The optimisation in Eq. 3 produces base models fbthat serve as priors: they learn a generic mapping from the feature space to trust scores across a wide range of devices. Training on synthetic data allows these models to capture general trust dynamics without relying on scarce real-world onboarding samples. Random Forest provides robustness against noise and non-linear feature interactions, k-NN offers sensitivity to local variations in feature space and XGBoost contributes high accuracy with millisecond-level inference time. Importantly, each model is trained and evaluated independently, not as an ensemble, to allow systematic comparison of strengths and weaknesses under sparse onboarding conditions. 3) Few-Shot Fine-Tuning Phase:For onboarding, only K samples are available per device: Dfew ={(xj, Tj)}K j=1, K ∈ {5,10,20}.(4) The fine-tuned model is trained with reduced complexity to avoid overfitting: ff= arg min f∈F 1 K K X j=1 (f(xj)−Tj)2.(5)
5 The dataset Dfew contains only Ksamples per device, with K∈ {5,10,20}chosen to reflect realistic onboarding conditions where only a handful of interactions are available before the device must be trusted or rejected. The optimisation in Eq. 5 adapts the base model to these limited samples. To reduce variance under such small K, the fine-tuned models are deliberately restricted in complexity (e.g., fewer trees in RF, shallower depth in XGBoost). Unlike typical N-way K-shot classification settings, our formulation is regression hence, each sample provides a continuous trust score label and K indicates the number of available regression pairs. This ensures the model adapts without overfitting to a handful of points. 4) Blended Prediction:To mitigate the risk of overfitting in few-shot scenarios while ensuring adaptability, the framework employs a blended prediction strategy that combines the base and fine-tuned models. The final trust score for a device is given by: Tfinal(x) = αfb(x) + (1 −α)ff(x), α ∈[0,1],(6) where fbis the base predictor (trained on the large synthetic dataset) and ffis the fine-tuned predictor (adapted to K onboarding samples). Equation 6 formalises the trade-off between robustness and adaptability: the base model fbtypically yields low-variance but sometimes biased predictions for new devices, while the fine-tuned model ffadapts to device-specific behaviour but exhibits higher variance under small K. Blending balances these effects, implementing a bias–variance compromise. The mathematically optimal solution weight α\∗ can be obtained by minimising the validation mean squared error (MSE). Let eb=fb(x)−Tand ef=ff(x)−Tdenote the residuals of the base and fine-tuned models with respect to the ground truth T. The expected blended error is: L(α) = E(αeb+ (1 −α)ef)2.(7) Expanding and differentiating with respect to αgives: α\∗ =Sf−Sbf Sb+Sf−2Sbf ,(8) where Sb=E[e2 b],Sf=E[e2 f], and Sbf =E[ebef]. In the special case of uncorrelated errors (Sbf ≈0), this reduces to: α\∗ ≈Sf Sb+Sf ,(9) which intuitively assigns more weight to the model with smaller error. In practice, we compute α\∗ on a validation split of the fewshot onboarding data. For numerical stability, α\∗ is clipped to [0,1] and regularised when the denominator of Eq. 8 is small. When Kis very small (e.g., 5), we also evaluate αvia grid search to confirm consistency with the analytical solution. Empirically, α\∗ ≈0.8across most experiments, reflecting the dominance of the base model under sparse conditions while still leveraging the corrective signal from the finetuned model. This convex combination ensures stable trust estimation without sacrificing adaptability, yielding robust predictions for secure IoT device onboarding. 5) Latency Model:End-to-end onboarding latency is decomposed as: L=LSSI +LMQT T +LDLT +Linf where LSSI ≈50 ms, LMQT T ≈20 ms, LDLT ≈200 ms, and Linf <7ms. For Ksequential samples: Ltotal(K)≈K·(LSSI +LMQT T +LDLT ) + Linf (10) Batching optimisations reduce LDLT per sample to meet the 500 ms target. Eq. 10 decomposes end-to-end onboarding latency into four measurable components: identity generation (LSSI ), communication (LMQT T ), distributed ledger logging (LDLT ), and inference time (Linf ). Among these, inference latency is consistently below 7 ms and therefore negligible compared to the ∼200 ms per-transaction DLT cost. The decomposition makes the bottleneck explicit and motivates batching: aggregating Bonboarding events per ledger transaction reduces the amortised LDLT to approximately 200/B ms per device, a critical optimisation to meet the 500 ms real-time requirement. This system-level model thus validates the feasibility of the framework under deployment conditions. Algorithm 1 presents the pseudocode implementation of these stages, directly corresponding to the mathematical framework defined by Equations 2–10. IV. USE CASE AND PERFORMANCE EVALUATION This section presents a practical use case for the proposed trust score prediction algorithm. It details the onboarding scenario for IoT devices in a consumer electronics context. An experimental setup evaluates the implementation of the algorithm. The performance assessment, supported by quantitative metrics and visual aids, illustrates the algorithm’s effectiveness and sheds light on its potential for real-world applications, including smart health devices and autonomous vehicles. A. Use Case and Experimental Setup The focus of this use case is on the secure onboarding and registration of a health-related smart device (e.g., a fitness tracker) within an IoT network, where a rapid and reliable verification of device credentials is essential for safety and operational trust. The onboarding sequence, as illustrated in Figure 2, starts with the IoT device generating a Decentralized Identifier (DID) managed by an Self-Sovereign Identity (SSI) management system, adhering to the W3C DID standard [12]. The SSI client then provisions cryptographic keys, offering three roots of trust: Physical Unclonable Function (PUF), Trusted Execution Environment (TEE), and Hyperledger Aries [13] as a fallback, providing flexibility across diverse hardware platforms. Upon successful DID creation and publication of the corresponding DID Document, the SSI Manager acknowledges registration. The IoT device then requests issuance of a Verifiable Credential (VC) embedding its unique cryptographic footprint. This credential is signed and the device is paired with its owner, who authenticates the DID using their own keys. Subsequently, the IoT device transmits trust data and its signed
6 Fig. 2. Secure onboarding message flow. The diagram highlights (a) cryptographic binding of device identity via DID/SSI and (b) tamper-evident logging of the computed trust score to the DLT. DID to the Trust Management Server (TMS) using MQTT. The TMS computes an initial trust score, then records this score and DID to the Distributed Ledger Technology (DLT) network via a gRPC interface. The DLT chaincode securely logs the onboarding event, with acknowledgements propagated back to the TMS, ensuring the onboarding completes within the target latency budget. The full onboarding message flow, including cryptographic binding via SSI and tamper-evident logging to the DLT, is shown in Figure 2. To evaluate the proposed trust management framework, we simulate a testbed of 50 heterogeneous IoT devices (including wearables and vehicle sensors) using an Intel i7 CPU, 16 GB RAM, and Ubuntu 20.04. Synthetic datasets are generated using the 2D Markov chain-based statistical trust modelling approach from [1], where device trust evolution is a function of both reputation and risk states. Specifically, the dataset consists of 1000 synthetic device samples, each comprising features such as CPU utilisation (C), RAM utilisation (R), security level (S), packet loss (P), risk state (RS), reputation state (RepS), and inverted OS/package ages (OSi,Pi). Trust scores (T) are computed for each device according to the piecewise statistical function defined in Eq. 2. The use of a Markov chain and MADM-driven synthetic simulation [1] enables the generation of feature vectors and trust scores spanning a wide spectrum of device behaviours and trustworthiness. By explicitly sampling from a range of risk, reputation, and performance states, the simulation provides a dataset that is both sufficiently large and balanced for model training. Consequently, no additional oversampling, undersampling, or post-processing techniques are needed to mitigate data imbalance. When binarising the trust score (e.g., trusted if T≥70), the dataset includes approximately 64% trusted and 36% untrusted devices, ensuring fair representation for both classes. In this formulation, resource usage (R,C,P) negatively impacts trust, while security and device/software freshness (S, OSi,Pi) contribute positively. The risk state thresholds ensure that devices with high risk are penalised, while low-risk states yield higher base trust. For each experiment, 80% of the synthetic dataset is used for training the base models (Random Forest, k-NN, XGBoost), and 20% for evaluation. Realistic onboarding scenarios are emulated using K={5,10,20}few-shot samples per device. All feature values are sampled using Python 3.10, NumPy, and pandas, with truncated normal or beta distributions for CPU, RAM, and packet loss, and categorical assignments for security, risk, and reputation states based on empirically observed device behaviors [1]. The ground truth trust score for each device is generated using the statistical function derived from the Markov chain and MADM-based model. These values serve as the reference labels for supervised training and evaluation of all machine learning models considered in this study. Model predictions are assessed against these Markovbased trust scores to determine accuracy and error metrics throughout our experiments. The overall pipeline, including model training, fine-tuning, and performance evaluation, is implemented in Python 3.9 with scikit-learn 1.0.2, XGBoost 1.5.0, and Hyperledger Fabric 2.2. The trust scores generated serve as ground truth labels for all supervised learning experiments. In line with few-shot learning evaluation practices, we adopt a K-shot regression approach in our experimental design. Here, Kdenotes the number of labelled onboarding samples per device used for model fine-tuning and evaluation, with K={5,10,20}. Unlike traditional classification-based few-
7 TABLE I PERFORMANCE METRICS FOR BASE AND FINE-TUNED MODELS (FINE-TUNED: 5 ONBOARDING SAMPLES) Model MSE R² MAE Inference Time (ms) RF (Base) 0.9904 0.9917 0.6796 2.9747 k-NN (Base) 18.3057 0.8460 2.6887 0.3256 XGBoost (Base) 0.4436 0.9963 0.4782 0.4660 RF (Fine-Tuned) 13.1759 0.8891 3.6297 3.4938 k-NN (Fine-Tuned) 33.0026 0.7223 5.7440 3.4716 XGBoost (Fine-Tuned) 20.5551 0.8270 4.5347 0.9425 shot learning (N-way K-shot), our task is formulated as a regression problem for continuous trust score prediction, so the concept of “N-way” (number of classes) does not apply. Instead, we analyse the model’s ability to rapidly adapt and generalise from a small set of onboarding data points per device, closely mirroring real-world IoT onboarding scenarios where only limited feature-label pairs are available at runtime. B. Performance Analysis The performance of the trust score prediction algorithm is evaluated using quantitative metrics and visual aids, focusing on accuracy, stability, and inference time across the simulated IoT network. The analysis compares RF, k-Nearest Neighbors (k-NN), and XGBoost models in both base and fine-tuned configurations, leveraging the few-shot transfer learning approach. The primary metrics include MSE, R-squared (R2), MAE, Mean Absolute Percentage Error (MAPE), Root Mean Squared Error (RMSE), and inference time. These metrics are computed on the test set for base models and the blended predictions for fine-tuned models. Table I summarises the results for the base models and for the fine-tuned models using five onboarding samples, averaged over 10 runs to account for variability. Table I shows that XGBoost achieves the highest R2 (0.9963) and lowest MSE (0.4436) among base models, indicating excellent predictive accuracy on the pre-trained dataset. After fine-tuning with only five onboarding samples, all models experience an increase in MSE and a decrease in R2, consistent with the expected risk of overfitting in extreme few-shot regimes. For example, RF maintains a strong R2of 0.8891 but MSE increases to 13.18, while k-NN exhibits the greatest sensitivity to limited data, with R2dropping to 0.72 and MSE rising to 33.00. These results reflect the inherent trade-off in few-shot learning: minimal data can limit precision but enable rapid adaptation and personalisation. To further analyse the impact of onboarding sample size, we systematically varied the number of samples used for finetuning. Table II reports the performance of the fine-tuned models as the onboarding sample size increases from 5 to 10 and 20. As shown, increasing the number of onboarding samples substantially improves both MSE and R2for RF and XGBoost: for RF, MSE drops from 13.18 (5 samples) to 5.07 (20 samples), while R2increases from 0.89 to 0.96. Similarly, XGBoost reaches an MSE of 4.45 and R2of 0.96 with 20 samples. k-NN remains less robust in the few-shot regime, though it also benefits from additional data. Random Forest k-NN XGBoost −20 −15 −10 −5 0 5 10 15 20 Residuals (Prediction - Ground Truth) Base Model Residuals Random Forest k-NN XGBoost −20 −15 −10 −5 0 5 10 15 Residuals (Prediction - Ground Truth) Fine-Tuned Model Residuals Fig. 3. Residual error distribution for base vs. fine-tuned models (5 onboarding samples). Random Forest maintains stable error distribution with a tight IQR, while k-NN exhibits wider variability and more outliers, reflecting higher sensitivity to sparse data. XGBoost shows moderate dispersion, balancing accuracy and robustness. These findings demonstrate that, while overfitting can occur when very few samples are available for fine-tuning, the hybrid few-shot transfer learning approach becomes robust as more onboarding data is provided. For real-world deployments, we recommend using at least 10 onboarding samples to ensure reliable trust score calibration and generalisation. This trend is consistent with statistical learning theory, which states that increased sample size leads to reduced empirical risk and improved generalisation performance. Fig. 4 visually compares the R2and MSE values for each model in both base and fine-tuned (5-sample) configurations. The results confirm that XGBoost achieves the highest accuracy in the base scenario, while all models experience a reduction in R2and an increase in MSE when fine-tuned with only five onboarding samples. This illustrates the trade-off between adaptability and generalisation in extreme few-shot regimes, as also reflected in Table I. Figure 3 presents the box plots of residuals for both base and fine-tuned models (5 onboarding samples). For the base models, Random Forest exhibits a tight interquartile range (IQR) of approximately ±1.0, with most residuals close to zero and few outliers, indicating stable and accurate predictions. After fine-tuning, the IQR for Random Forest widens to roughly ±3.0, consistent with the observed increase in prediction variability due to a limited few-shot data. The k-NN model exhibits a broader base IQR (approximately ±4.5), with its fine-tuned IQR expanding to around ±6.0, accompanied by several noticeable outliers, reflecting its higher sensitivity to small sample sizes and resulting in greater prediction error variance. XGBoost maintains a narrow base IQR (≈ ±0.7), while its fine-tuned IQR increases to about ±3.5, illustrating that it remains robust in the base configuration but, like the other models, exhibits more dispersed residuals after fewshot adaptation. Overall, these distributions highlight Random Forest’s robustness and XGBoost’s high precision under base training. The predictive capability of the models with five onboarding samples per device is illustrated in Figures 5, 6, and 7. For Random Forest, the predicted trust scores for a representative set of onboarding devices closely track the ground truth, with predictions (85.89, 77.18, 86.96, 99.99, 99.31) aligning well with calculated true values (85.25, 77.27, 85.92, 100.00, 100.00), and an average absolute error of 1.74. The k-NN
8 TABLE II PERFORMANCE OF FINE-TUNED MODELS VS. NUMBER OF ONBOARDING SAMPLES Samples Model MSE R2MAE MAPE (%) Inference Time (ms) Std. Residuals 5 Random Forest 13.18 0.889 3.14 3.48 3.33 3.12 k-NN 33.00 0.722 4.74 5.47 3.47 5.53 XGBoost 20.56 0.827 3.79 4.13 0.96 3.80 10 Random Forest 6.29 0.947 2.07 2.45 3.49 2.50 k-NN 27.83 0.766 4.11 4.87 3.33 5.25 XGBoost 4.54 0.962 1.54 1.91 0.96 2.10 20 Random Forest 5.07 0.957 1.54 1.89 3.70 2.23 k-NN 26.70 0.775 3.49 4.36 0.64 5.07 XGBoost 4.45 0.963 1.49 1.85 0.88 2.06 Random Forest k-NN XGBoost Models 0.0 0.2 0.4 0.6 0.8 1.0 R² Score R² Comparison: Base vs Fine-Tuned Models Base Fine-Tuned Random Forest k-NN XGBoost Models 100 101 MSE MSE Comparison: Base vs Fine-Tuned Models Base Fine-Tuned Fig. 4. Comparison of R2and MSE for base and fine-tuned models (finetuned with 5 onboarding samples). 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 Sample Index 50 60 70 80 90 100 Trust Score Predictions vs Ground Truth (Random Forest) Ground Truth Predictions Fig. 5. Random-Forest predictions on five unseen onboarding samples. The error bars (±σ) remain within ±3trust-score units, illustrating the 2.29 average absolute error reported in Table I predictions (82.66, 85.66, 77.48, 94.88, 88.67) show larger deviations from the corresponding ground truth (81.37, 87.51, 74.82, 100.00, 71.37), with an average absolute error of 5.20, indicating instability under few-shot fine-tuning. For XGBoost, predicted values (99.72, 83.63, 99.82, 99.32, 84.01) also follow the ground truth (100.00, 85.36, 100.00, 100.00, 85.77) closely, yielding an average absolute error of 2.32. These results validate the algorithm’s adaptability in the fewshot setting, with Random Forest and XGBoost demonstrating strong generalisation even with minimal onboarding data, while k-NN remains more sensitive to the sample size. 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 Sample Index 50 60 70 80 90 100 Trust Score Predictions vs Ground Truth (k-NN) Ground Truth Predictions Fig. 6. k-NN predictions exhibit larger deviations and two clear outliers, consistent with its higher post-tuning MSE (33.0). This confirms that distance-based methods struggle with the high-variance, low-sample regime. 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 Sample Index 50 60 70 80 90 100 Trust Score Predictions vs Ground Truth (XGBoost) Ground Truth Predictions Fig. 7. XGBoost balances accuracy and smoothness. Although not as precise as RF on tiny samples, XGBoost avoids the k-NN outliers while maintaining less than 1ms inference, supporting its role as a lightweight fallback. C. Scalability and Real-Time Performance The trust management algorithm demonstrates strong scalability across a 50-device network. With the adoption of batch DLT onboarding, the TMS now processes requests at an average throughput of 137.7 devices per second across the Random Forest, k-NN, and XGBoost models, as measured on
9 TABLE III SCALABILITY AND PERFORMANCE METRICS ACROSS MODELS (BATCH DLT ONBOARDING) Model Throughput (devices/s) Total Latency for 5 Devices (ms) Random Forest 106.77 745.41 k-NN 198.42 708.66 XGBoost 107.91 722.44 Average 137.70 725.50 Random Forest k-NN XGBoost Models 0 25 50 75 100 125 150 175 200 Throughput (devices/s) TMS Throughput Across Models Average (137.70 devices/s) Model Throughput Fig. 8. Bar chart comparing TMS throughput across Random Forest, k-NN, and XGBoost models (batch DLT onboarding), with the average throughput (137.70 devices/s) indicated. the simulated testbed. Table III and Figure 8 highlight the model-wise variability in throughput, with k-NN achieving the highest (198.4 devices/s) and Random Forest and XGBoost providing consistent performance above 100 devices/s. The onboarding workflow involves simulated SSI operations (mean 53–63 ms per device), MQTT transmission (mean 20–35 ms per device), and now a single batch DLT transaction for each group of 5 devices (approximately 300ms total). This batching reduces the cumulative DLT latency from over 1000 ms (in the previous per-transaction approach) to well within the 500 ms real-time target set for industrial IoT onboarding. The total latency for onboarding 5 devices now averages under 750 ms for all models, with DLT contributing only a fraction of the total delay. This confirms that the proposed system meets real-time requirements in practical, scalable IIoT deployments. Figure 8 presents the throughput comparison, showing that k-NN is fastest in raw inference speed, but all models deliver real-time onboarding performance under the batch-optimised pipeline. The impact of device count on throughput is further visualised in Figure 9, confirming stable scaling characteristics for each model. V. DISCUSSIONS A. Proposed Framework Benefits The proposed TM framework leverages transfer learning and FSL to address IoT device onboarding challenges in consumer electronics, offering notable benefits in efficiency and adaptability. Transfer learning utilises prior knowledge from a large synthetic dataset (1000 devices) generated by a 2D Markov chain model [1], enabling base models (RF, k-NN, XGBoost) 10 15 20 25 30 35 40 45 50 Number of Devices 80 100 120 140 160 180 200 Throughput (devices/s) Throughput vs Device Count Random Forest k-NN XGBoost Fig. 9. TMS throughput as a function of onboarding device count for each model. to learn general trust patterns without extensive real-time data collection. This approach reduces computational overhead and training time, which is critical for resource-constrained IoT environments where data scarcity and privacy concerns are prevalent. For instance, pre-trained models achieve high accuracy (e.g., XGBoost base R2of 0.9963), ensuring rapid trust assessment within real-time constraints. Each sample in the synthetic dataset consists of eight features—CPU utilisation, RAM utilisation, security level, packet loss, risk state, reputation state, inverted OS age, and inverted package age—and is labeled by computing the trust score via the piecewise function in Eq. (2). FSL enhances adaptability by fine-tuning these models with only a few onboarding samples (e.g., 5–20), addressing data scarcity in dynamic IoT settings. This enables the framework to quickly adapt to new devices, such as smart health wearables, with fine-tuned RF R2reaching 0.9574 (20 samples) and prediction error (MAE) as low as 1.54. The blended prediction approach—where the optimal base/fine-tuned weights are empirically determined for each model—balances robustness and adaptability, minimising overfitting risks as evidenced by the validation error curve. Additionally, the framework’s low inference times (e.g., XGBoost at 0.9999 ms) and scalability (average throughput 126.9 devices/s, peaking at 174.6 devices/s for k-NN) make it suitable for large-scale, timesensitive applications. Transfer learning also enables crossdomain applicability, allowing the framework to adapt to contexts like industrial IoT or vehicular networks, enhancing its versatility. To quantitatively compare the proposed framework with existing TM approaches, Table IV summarises key performance metrics across our models and state-of-the-art methods. The table highlights the framework’s competitive inference times and base model accuracy, although fine-tuned MSE values remain higher, reflecting the inherent trade-off in rapid fewshot adaptation. The comparison reveals that our framework excels in inference time (e.g., under 1 ms for XGBoost) compared to LSTM and Bi-LSTM, which require orders of magnitude