scieee AI-readable full text Open interactive document viewer

Route Optimization in Smart Villages: A Graph Neural Network Approach

Durán-López, Alberto; Bolaños-Martinez, Daniel; Almahmoud, Zaid; Pravin, Chandresh; De, Suparna; Bermudez-Edo, Maria

Abstract

Abstract In the Internet of Things (IoT) domain, modeling the intricate relationships within spatiotemporal data from sensors is a significant challenge. Conventional methods, which often rely on tabular or time-series formats, frequently fail to capture the underlying complex dynamics. In contrast, graph-based representations excel at this task, but they have seldom been explored in tourism and remain unexplored in small villages, where IoT networks involve only a handful of sensors or nodes. We propose an explainable and optimized graph neural network (GNN) methodology to predict return visits in smart villages. We collected data from License Plate Recognition cameras across three villages, tracking 467 773 distinct vehicles over 2 years and 11 months. For each vehicle, we construct a unique graph that maps its specific route across all cameras. We conduct a comprehensive study of various GNN architectures, including spectral, spatial, recurrent, and attention-based models, with a special focus on the impact of incorporating edge features. We also apply explainability methods to identify the most influential graph components driving model predictions. To optimize performance, we integrate neural architecture search and hyperparameter optimization (NAS-HPO) using a multivariate objective that maximizes the F1-score while minimizing computational time. Analysis of the results reveals that models using edge attributes perform better; specifically, the edge-featured graph attention network (E-GAT) model. We apply the NAS-HPO approach to this E-GAT base model, and the improved version achieves the highest performance across all metrics, demonstrating a 5.71% improvement in F1-score and a 20.87% decrease in execution time.

Full text

IEEE INTERNET OF THINGS JOURNAL, VOL. 12, NO. 21, 1 NOVEMBER 2025 45235 Route Optimization in Smart Villages: A Graph Neural Network Approach Alberto Durán-López , Daniel Bolaños-Martinez , Zaid Almahmoud , Chandresh Pravin , Suparna De ,Member, IEEE, and Maria Bermudez-Edo Abstract—In the Internet of Things (IoT) domain, modeling the intricate relationships within spatiotemporal data from sensors is a significant challenge. Conventional methods, which often rely on tabular or time-series formats, frequently fail to capture the underlying complex dynamics. In contrast, graph-based representations excel at this task, but they have seldom been explored in tourism and remain unexplored in small villages, where IoT networks involve only a handful of sensors or nodes. We propose an explainable and optimized graph neural network (GNN) methodology to predict return visits in smart villages. We collected data from License Plate Recognition cameras across three villages, tracking 467773 distinct vehicles over 2 years and 11months.Foreach vehicle,weconstruct aunique graphthat maps its specific route across all cameras. We conduct a comprehensive study of various GNN architectures, including spectral, spatial, recurrent, and attention-based models, with a special focus on the impact of incorporating edge features. We also apply explainability methods to identify the most influential graph components driving model predictions. To optimize performance, we integrate neural architecture search and hyperparameter optimization (NAS-HPO) using a multivariate objective that maximizes the F1-score while minimizing computational time. Analysis of the results reveals that models using edge attributes perform better; specifically, the edge-featured graph attention network (E-GAT) model. We apply the NAS-HPO approach to this E-GAT base model, and the improved version achieves the highest performance across all metrics, demonstrating a 5.71% improvement in F1-score and a 20.87% decrease in execution time. Index Terms—Explainability, graph neural network (GNN), Internet of Things (IoT), NAS and hyperparameter optimization (NAS-HPO). I. INTRODUCTION THE EXPANSION of Internet of Things (IoT) devices and accompanying sensor technologies has triggered an Received 18 March 2025; revised 13 June 2025 and 15 July 2025; accepted 10 August 2025. Date of publication 15 August 2025; date of current version 24 October 2025. This work was supported in part by Grant C-SEJ-128UGR23 funded by Consejería de Universidad, Investigación e Innovación and by ERDF Andalusia Program 2021-2027; project PID2023-149185OBI00 funded by MICIU/AEI/10.13039/501100011033 and co-funded by ERDF/EU. Funding for open access charge: Universidad de Granada/CBUA. (Corresponding author: Alberto Durán-López.) Alberto Durán-López, Daniel Bolaños-Martinez, and Maria Bermudez-Edo are with the CITIC-UGR and IATUR Institutes, University of Granada, 18012 Granada, Spain (e-mail: [email protected]; [email protected]; [email protected]). Zaid Almahmoud is with the Artificial Intelligence and Media Lab, Northwestern University in Qatar, Ar-Rayyan, Qatar. (e-mail: [email protected]). Chandresh Pravin and Suparna De are with the School of Computer Science and Electronic Engineering, University of Surrey, GU2 7XH Guildford, U.K. (e-mail: [email protected]; [email protected]). Digital Object Identifier 10.1109/JIOT.2025.3599235 increase in the generation of daily data [1]. While traditionally this type of data has been organized in tabular or timeseries formats [2], many real-world systems, ranging from sensor and transportation systems to social networks, are inherently interconnected and exhibit relationships that are most effectively represented as graphs [3]. This graph-structured data poses new challenges for conventional machine learning (ML) approaches, which typically struggle to capture the detailed relational and structural information embedded within graphs. Graph neural networks (GNNs) [4] have emerged as a powerful approach specifically developed to address these challenges [5]. These models operate directly on graph data by aggregating and propagating features along the connections between nodes [6]. This process not only captures local interactions but also integrates global structural information, ensuring that the intricate relationships inherent in the data are effectively preserved and exploited [7]. When data is represented as a graph, it must be transformed into embeddings before being input into GNNs, a preprocessing step that significantly impacts downstream performance. Generating representative embeddings is necessary for robust outcomes in tasks such as graph classification, clustering, and node or edge prediction. These embeddings must not only preserve the structural and semantic information of the data but also enhance it by introducing additional representational dimensions that make the data more useful for models [8]. Traditional GNNs, such as graph convolutional networks (GCNs) [4], gated graph convolutional networks (GGCNs) [9], and Graph Attention Networks (GAT) [10],have primarily focused on capturing node features and overall graph structure while often overlooking the information encoded within edge features [11]. Recent advances have addressed this limitation by modifying the message passing process to explicitly incorporate edge attributes, as seen in models such as the relational GCN (R-GCN) [12] and various heterogeneous GNN architectures [13]. This integration of edge information not only enhances model performance but also works to improve explainability, offering valuable insights into how specific nodes and edges contribute to the decisions made by models [14]. Over recent years, new GNNs have been developed to address the intricate and domain-specific challenges posed by different types of graph data [3]. These architectures include spectral convolutional [15], spatial [16], recurrent [9], and attention-based models [10]. Additionally, spatio-temporal models [17] have been introduced to capture changes in c 2025 The Authors. This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License. For more information, see https://creativecommons.org/licenses/by-nc-nd/4.0/ 45236 IEEE INTERNET OF THINGS JOURNAL, VOL. 12, NO. 21, 1 NOVEMBER 2025 graph structures over time. However, performance variations between these models are often driven by the inherent characteristics of the data more than by the models themselves [6]. Moreover, certain GNN architectures function as closed boxes, making it difficult to understand how predictions are made. Consequently, identifying the nodes and edges that most significantly influence predictions is important for understanding model behavior and ensuring robustness in real-world applications [18]. Explanation techniques, such as GNNExplainer [19], shed light on the key features and relationships driving model performance, thus improving transparency and facilitating validation. Despite the advancements in new GNN architectures, a fundamental challenge persists: finding an effective model configuration for a specific graph dataset. Neural architecture search (NAS) [20] has emerged as a solution to this challenge by systematically exploring various configurations of layers and hyperparameters. This automated approach is in response to the observation that performance variations of GNNs can often be attributed more to the characteristics of the data than model selection [6]. By leveraging NAS, researchers can efficiently identify optimal architectures tailored to specific graph tasks [21]. This not only accelerates development but also enhances performance by discovering architectures suited to the intrinsic properties of each dataset [22]. In real-world traffic and transportation applications, road networks are commonly represented as graphs, where roads correspond to edges and intersections to nodes, thereby facilitating further analysis [23]. This graph-based approach has proven to be effective for various tasks, such as route planning and traffic prediction [3]. In the broader smart tourism domain, ML has been used to optimize tourist itineraries by modeling attractions and paths as a graph network [24],aswellas to predict user behavior on online travel platforms for hotel recommendation [25]. However, the smart village domain remains under-explored, particularly in applications of GNNs for mobility analysis, where the resulting graphs tend to be small, raising questions about the effectiveness of approaches on transportation graphs with a small number of nodes. To address this gap, we present a methodology that leverages graph-based ML with explainability on a real smart village traffic dataset. The dataset, collected via IoT cameras with automatic number plate recognition (ANPR), is used to construct a road network graph reflecting vehicle movements. Using this data, we predict repeat tourist visits, yielding societal benefits by reducing overcrowding [26], economic benefits through steady revenue [27], and planning benefits by guiding infrastructure and resource allocation [28]. By applying GNN models along with explainability techniques, we can understand which features most influence traffic patterns in the village, turning closed-box predictions into actionable insights. While NAS has been employed in other domains to automatically improve deep neural networks (including GNNs [29]), to our knowledge, it has never been applied to tourist mobility classification in a smart village context. This work is therefore novel in combining graph-based route modeling, IoT data, GNN explainability, and an exploration of NAS to optimize models for smart village traffic management, addressing an important gap in the current research. This study addresses these gaps by proposing an integrated framework that combines IoT-based graph construction, GNN explainability, and automated GNN architecture search for rural mobility analysis. Our specific objectives are: first, to evaluate the predictive performance of GNN models on small-scale transportation networks; second, to assess the explainability of GNN predictions for tourism; and third, to determine the efficiency improvements of using automated NAS. The main contributions of our approach are as follows. 1) We present a methodology and validate it on our case study dataset to forecast tourism in a rural smart village. Our approach transforms real-world smart village data into graph representations and employs explainability techniques on the GNN. By evaluating both node and edge feature importance, we gain clear insights into the decision-making process. 2) We implement a multivariate NAS and hyperparameter optimization (NAS-HPO) strategy to explore various layer configurations and hyperparameter settings, aiming to maximize the F1-score while reducing execution time. 3) We construct a real-world dataset from a road network using IoT ANPR sensor data, covering 2 years and 11 months of vehicle routes in a smart village. This dataset, comprising 467 773 distinct vehicles and preserving the sequential patterns of tourist routes, enables thorough evaluation of our approach. The remaining sections of this article are organized as follows. Section II provides a review of the existing literature on GNN models and state-of-the-art on smart village case studies. The proposed methodology is outlined in Section III. Section IV details the experiments conducted on our smart village case study dataset, with results analyzed in Section V. Finally, Section VII concludes this article. II. RELATED WORK In this section, we provide a comprehensive review of GNN literature, covering spectral, spatial, attention-based networks, and other related architectures. In addition to foundational and theoretical works, we also incorporate recent empirical studies on tourism and smart villages relevant to our work, while highlighting our contribution. A. Spectral Graph Neural Networks Spectral GNNs leverage graph signal processing techniques to perform convolutions in the spectral domain, effectively capturing structural information from graph-structured data. ChebNet [15] introduced an efficient formulation of spectral convolutions using Chebyshev polynomial approximations, enabling fast and localized graph filtering with linear computational complexity. Building on this foundation, Kipf and Welling [4] proposed GCN, a first-order approximation of spectral convolutions that enables scalable semi-supervised learning by efficiently aggregating node features using a normalized Laplacian. However, GCN does not fully exploit edge DURÁN-LÓPEZ et al.: ROUTE OPTIMIZATION IN SMART VILLAGES 45237 attributes, which can encode essential relational information. To address this, E-GCN [11] extends the GCN framework by integrating edge features through a doubly stochastic normalization strategy and multidimensional edge-aware transformations, leading to enhanced representation learning for tasks such as node and graph classification. B. Spatial Graph Neural Networks Spatial GNNs operate directly on a graph’s structure by defining convolutional operations in the spatial domain, typically aggregating information from local node neighborhoods. One of the most well-known spatial GNNs, GraphSAGE [16], provides an inductive learning framework that generates node embeddings by sampling and aggregating features from neighboring nodes. This approach enables generalization to unseen nodes and scales efficiently to large graphs. Building on GraphSAGE, E-GraphSAGE [1] extends GraphSAGE by incorporating edge features to modulate the message passing based on relationship information, demonstrating superior performance on benchmark datasets. Meanwhile, message passing neural networks (MPNNs) [30] unify various GNN models under a common framework and have been particularly effective in molecular property prediction tasks. In the context of point clouds, dynamic graph CNN (DGCNN) [31] introduces EdgeConv, a module that dynamically computes graph structures at each layer to capture both local and global shape properties, significantly improving classification and segmentation tasks. For knowledge graphs, R-GCNs [12] extend GCNs to handle multirelational data by incorporating relationspecific transformations, proving effective in link prediction and entity classification. Furthermore, heterogeneous GNNs (HetGNNs) [21] address the challenges of learning representations in heterogeneous graphs by leveraging a random walk sampling strategy and a two-stage aggregation mechanism, achieving strong performance in tasks like recommendation and node classification. C. Attention-Based Graph Neural Networks Attention-based GNNs enhance message passing by dynamically weighting node interactions, allowing models to differentiate the importance of neighboring nodes. Graph Attention Networks (GATs) [10] introduce masked self-attentional layers to assign different importance scores to neighboring nodes, improving upon traditional graph convolutional approaches. By leveraging multihead attention, GAT enables efficient learning on both transductive and inductive tasks while avoiding expensive matrix operations. The model achieves state-of-the-art results on various benchmarks. Expanding on this, edge-featured graph attention networks (E-GATs) [32],[33] incorporate edge attributes into the attention mechanism, addressing limitations in node-centric GNNs. E-GAT learns representations by integrating edge features into attention weight calculations and updating edge attributes alongside node embeddings. This approach enhances performance on edge-sensitive graph tasks, making E-GAT particularly effective for applications in domains like financial transaction networks. D. Graph Autoencoders and Recurrent Graph Neural Networks Graph autoencoders (GAEs) are unsupervised learning models designed to encode graph-structured data into latent representations and reconstruct the graph topology from them. One of the most notable variants, variational GAE (VGAE) [34], extends the variational autoencoder (VAE) framework to graphs by leveraging a GCN as the encoder and an inner-product decoder. By incorporating node features into the latent variable model, VGAE achieves competitive performance in unsupervised link prediction tasks on benchmark citation network datasets. Recurrent GNNs extend GNN architectures by incorporating recurrent mechanisms, allowing iterative updates of node embeddings through multiple message-passing steps. One such model, gated graph neural networks (GGNNs) [9], improves upon GNNs by integrating gated recurrent units (GRUs) to iteratively refine node representations based on their local neighborhoods. GGNNs have demonstrated strong performance in various structured learning tasks, including program verification and algorithm learning. Building upon GGNN, edge-gated GNNs (E-GGNNs) [35] enhance message passing by introducing edge-aware gating mechanisms. These mechanisms dynamically regulate information flow along edges, allowing the model to capture more nuanced relational structures. E-GGNN has shown effectiveness in specialized domains, such as few-shot learning and molecular property prediction, where modeling edge information is critical for improving classification and regression tasks. E. Spatio-Temporal Graph Neural Networks Spatio-Temporal GNNs extend conventional GNNs by modeling temporal dependencies in addition to spatial dependencies, making them effective for dynamic graph-based tasks such as traffic forecasting and human activity recognition. These models simultaneously leverage spatial and temporal information to capture evolving patterns in structured data. Spatio-temporal GCN (STGCN) [36] integrates spatial graph convolutions with temporal convolutional networks to effectively model dynamic traffic patterns over time. Instead of relying on standard recurrent or convolutional layers, STGCN processes traffic data as a graph structure, enabling more efficient and accurate forecasting. By modeling multiscale traffic networks, STGCN significantly improves predictive performance on real-world datasets. Building upon spatio-temporal modeling, dynamic spatiotemporal graph transformer network (DST-GTN) [17] introduces a transformer-based approach that dynamically adapts to changing spatial relationships over time. Unlike traditional methods that treat spatial structures as static, DSTGTN captures dynamic adjacency relationships and refines both global and local spatio-temporal representations using 45238 IEEE INTERNET OF THINGS JOURNAL, VOL. 12, NO. 21, 1 NOVEMBER 2025 TABLE I OVERVIEW OF GNNS adaptive filters. This enables the model to extract rich spatiotemporal features from time-series traffic data, achieving state-of-the-art results in various forecasting tasks. Recently, a Bayesian approximation [37] of a Multivariate Time-series GNN (Bayesian approximation of a multivariate time-series Graph Neural Network (B-MTGNN)) model [38] was proposed for forecasting multivariate time-series data represented as a graph. This approach models spatial and temporal dependencies and adaptively learns hidden relationships between different variables to enhance multivariate time-series prediction, particularly for complex datasets such as cybersecurity trends. Additionally, the model quantifies epistemic uncertainty using the Monte Carlo dropout method [39], improving reliability through uncertainty estimation. F. Hierarchical Pooling Graph Neural Networks Hierarchical Pooling GNNs address the limitations of traditional GNN architectures by enabling the learning of hierarchical representations of graphs. This capability is particularly beneficial for graph classification tasks, where a comprehensive understanding of the graph’s structure is important. DiffPool [40] introduces a differentiable graph pooling (gPool) module that generates hierarchical representations. By learning soft cluster assignments for nodes at each layer, DiffPool coarsens graphs and creates fixed-size representations suitable for classification. This method can be integrated with various GNN architectures in an end-to-end manner, improving performance on graph classification benchmarks by 5-10% compared to existing pooling methods, and achieving state-of-the-art results on multiple datasets. Building on the concept of hierarchical representation learning, Graph U-Nets [41] introduce novel gPool and unpooling (gUnpool) operations to address the unique challenges posed by graph data. The gPool layer selectively chooses nodes based on their scalar projection values onto a trainable vector, effectively coarsening the graph for downstream classification tasks by selecting the highest scoring nodes (TopKPool). The inverse gUnpool layer reconstructs the original graph structure from the pooled representation. The introduction of these operations enables the development of an encoder-decoder architecture tailored for graph data, demonstrating superior performance in node and graph classification tasks. Table I summarizes the related work on GNNs. G. Comparative Analysis of Recent Works Recent studies have increasingly applied ML and deep learning techniques to tourism forecasting and smart village analytics. These efforts span a variety of approaches, from classical time-series models to graph-based and attentionbased neural architectures. In this section, we review recent empirical work in two main areas relevant to our contribution: DURÁN-LÓPEZ et al.: ROUTE OPTIMIZATION IN SMART VILLAGES 45239 1) tourism forecasting and 2) smart village systems. We end this section by positioning our contribution within this evolving landscape. 1) Tourism Forecasting: Some researchers employed classical time-series models to forecast tourism demand [43]. Others integrated online behavior and Internet-search trends into prediction models with kernel-based and mixed-frequency learning approaches [44]. The study in [45] applied RobustSTL to decompose the series into a trend component and a residual component capturing both seasonal cycles and irregular fluctuations. It forecasted the trend with ARIMA, modeled the residual with an LSTM, and merged the two forecasts to produce the final prediction. Chen et al. [46] picked Baidu keywords via Boruta, added visitor counts, and used a CNN-BiLSTM framework to forecast monthly flows, outperforming both traditional and single-network models across three Chinese attractions. The study in [47] focuses on daily inbound tourist volume in South Korea using a MultiHead Attention CNN framework that integrates temporal attention to deliver more accurate predictions than other deeplearning approaches. Transformer-based models have been used in [48], where temporal fusion transformers are combined with Bayesian optimization to produce robust forecasts. Extensions of the transformer architecture appear in [49], addressing spatial–temporal dependencies and interpretability through modular fusion networks and masked attention mechanisms. Recent work adopted graph-based architectures: the study in [50] proposed a GNN exploiting spatiotemporal correlations with dual correlation matrices and attention mechanisms, achieving high accuracy across Chinese tourist cities. A graph-guided network modeling lag effects by dynamically encoding historical data into a bipartite graph appears in [51]. Li et al. [52] used a hybrid GCN–LSTM model to forecast monthly tourist arrivals across European countries and hourly visits to Beijing attractions, demonstrating that incorporating spatial effects through graph structures significantly reduces forecasting errors. The approach in [53] employs a multigraph convolutional network to capture diverse relationships among attractions via multiple graph representations, outperforming baseline methods in tourist flow forecasting. Hajisafi et al. [54] introduced Busyness GNN (BysGNN) to predict POI visits in the United States by constructing dynamic graphs from temporal, geographic, and contextual features, yielding significant performance gains over state-of-the-art models. 2) Smart Villages: Data from satisfaction questionnaires have been employed to train and optimize genetic algorithms to develop an accommodation-recommendation system for smart villages [55]. Similarly, Surya et al. [56] proposed an integrated tourism model for smart villages and validates it via partial least squares structural equation modeling (PLS-SEM). Only a few works exploit sensor data within smart villages. For instance, one study leverages ANPR data, with socio-economic and seasonal features, to predict overnight stays in smart villages [57]. Similarly, Bolaños-Martinez et al. [58] combined questionnaires and contextual datasets with ANPR feeds to forecast tourists’ intent to visit over the next 12 months. 3) Research Contribution: While tourism studies offer valuable contributions, they primarily focus on forecasting aggregated tourist volume or demand across cities or attractions, often using either time-series or structured input representations. Among the reviewed graph approaches, none evaluate multiple GNN models to determine which architecture works best for the task. They also do not use explainability to reveal how edge attributes contribute to predictions, nor do they apply NAS HPO tuning to jointly optimise model structure and training settings. Moreover, existing methods depend on aggregate flow statistics or static graphs built from cumulative data rather than on the individual time ordered paths of each traveler. This limitation is especially problematic in smart village settings where only a small number of sensor nodes are available and every single route matters. Existing smart village studies use contextual databases with limited accessibility (such as questionnaires), which restricts model training. In contrast, our work focuses on repeat visitation prediction in smart village environments by modeling real-world route data as graphs derived from IoT-based vehicle detection, without relying on additional information. We explore a diverse set of GNN architectures, including spectral, spatial, recurrent, and attention-based models, while also incorporating edge features. Furthermore, we employ NAS-HPO to jointly optimize performance and efficiency, addressing scalability and deployment in rural sensor-based systems. We additionally apply explainability methods to uncover graph structures driving predictions, highlighting a novel intersection between predictive modeling, efficient inference, and model interpretability in micro-level tourism prediction. III. PROPOSED METHODOLOGY Our methodology (see Fig. 1) begins with the creation of a graph object for each dataset instance. This involves identifying which nodes and edges (with their attributes) form the graphs and converting them into graph objects (Section III-A). Next, we apply different GNN models to evaluate their performance in classifying these graphs. This step involves using a wide range of GNNs to determine which models adapt best to the data, and whether edge features provide additional information or performance improvements (Section III-B). We do this because GNNs often perform differently across datasets, showing that comparisons are task-specific [59]. We then use explainability techniques to understand feature importance in our best-performing model (Section III-C). This allows us to understand which nodes or edges are most influential and to determine the key drivers of the model’s predictions. Finally, after identifying the GNN model that best fits the classification task and using it as a base, we employ NAS-HPO to identify the optimal combination of layers and hyperparameters for the model’s architecture (Section III-D). A. Graphs Preprocessing and Representation If the initial dataset is not given as a graph object, it must be preprocessed into a graph structure. This involves defining 45240 IEEE INTERNET OF THINGS JOURNAL, VOL. 12, NO. 21, 1 NOVEMBER 2025 Fig. 1. Pipeline and architecture of our framework. which features become the nodes, which relationships become the edges, and whether nodes and edges carry additional attributes. We define a graph as G=(V,E,Xv,Xe)(see [30],[60]), where: 1) Vis the set of nodes; 2) E⊆V×Vis the set of edges; 3) Xv∈R|V|×dis the node feature matrix, with each node v∈Vassociated with a feature vector xv∈Rd; and 4) Xe∈R|E|×pis the edge feature matrix, where each edge eij ∈E, connecting nodes viand vj, is associated with a feature vector xeij ∈Rp. This transforms the problem of processing tabular or sequential camera detection data into a graph network problem, which is needed to structure the data in a way that GNN models can directly process. In our work, the graph structure is fixed for each instance, meaning that the nodes, edges, and their corresponding features remain constant over time. B. Graph Neural Network GNNs aim to learn expressive representations of nodes by iteratively aggregating and transforming information from their local neighborhoods. In the traditional formulation, the focus is on leveraging the graph structure and node features alone [4],[10]. For each node v∈V,leth(l) vdenote its hidden representation at layer l, with the initial representation given by h(0) v=xv. The update mechanism at layer lis defined as m(l) v=AGG(l){φ(l)h(l−1) v,h(l−1) u:u∈N(v)} h(l) v=ψ(l)h(l−1) v,m(l) v(1) where: 1) N(v)denotes the set of neighbors of node v; 2) φ(l)is a message function that computes the contribution from neighbor uto node v; 3) AGG(l)is an aggregation function (e.g., sum, mean, or max) that combines the messages from all neighbors; and 4) ψ(l)is an update function that fuses the previous representation of vwith the aggregated message. After Llayers, each node’s final representation h(L) vencapsulates information from its L-hop neighborhood. For graph-level tasks, a readout function R(·)aggregates these node representations into a single graph embedding hG=R{h(L) v:v∈V}.(2) While the framework above leverages only node features and the underlying graph structure, modern GNNs extend the message-passing process to incorporate edge features as well. In these models, the message function is modified to include edge attributes xeuv from the graph G=(V,E,Xv,Xe), leading to the following update: m(l) v=AGG(l){φ(l)h(l−1) v,h(l−1) u,xeuv :u∈N(v)} h(l) v=ψ(l)h(l−1) v,m(l) v.(3) In (3), while the message function φ(l)has been adapted as detailed above to include edge features xeuv , the remaining components, specifically the aggregation function AGG(l),the update function ψ(l), and the node representations h(l−1) vand h(l−1) u, preserve their definitions and operational roles from the traditional formulation (1). In our study, we aim to evaluate whether the inclusion of edge features xeuv enables the network to capture additional relational and contextual information between nodes. To this end, we first employ traditional GNN models that rely solely on node features and the graph structure. Subsequently, we DURÁN-LÓPEZ et al.: ROUTE OPTIMIZATION IN SMART VILLAGES 45241 experiment with GNN architectures that incorporate edge features, as in (3). This comparative analysis will help us determine if integrating edge attributes improves the models’ ability to learn more detailed representations of complex graph-structured data. C. Explainability To interpret the predictions of our GNN models, we use an explainability method that shows the importance of both node features and edges. We use GNNExplainer [19] because it produces masks for nodes and edges, which helps us understand what parts of the graph drive the prediction. This approach gives us insight into how the network works and whether including edge features improves the learned representations. Given a graph G=(V,E,Xv,Xe)and a trained GNN model f, the explainer produces two importance masks, the mask for node features Mvand the mask for edges Me Mv∈[0,1]|V|×dand Me∈[0,1]|E| which indicate how relevant each node and edge feature is to the model’s prediction. Specifically, GNNExplainer introduces soft masks Xvand Xeto modulate the original inputs ˜ Xv=XvMvand ˜ Xe=XeMe where denotes element-wise multiplication. The perturbed graph is then ˜ G=(V,E,˜ Xv,˜ Xe). The explainer learns these masks by minimizing the following objective: min Mv,Me Lpredf(G), f˜ G+λ1Mv1+λ2Me1 subject to Mv∈[0,1]|V|×dand Me∈[0,1]|E|. 1) The prediction loss Lpred(f(G), f(˜ G)) measures the discrepancy between the original model output f(G)and the output on the perturbed graph f(˜ G). We use crossentropy loss because our task is a classification problem. 2) The L1norms Mv1and Me1encourage sparsity in the masks, making the explanations simpler and easier to interpret. 3) The hyperparameters λ1and λ2balance the tradeoff between fidelity and sparsity in the learned masks. By optimizing this objective, GNNExplainer identifies a subgraph and subset of node features that most strongly influence the model’s prediction, thus providing interpretable insights into how the GNN processes graph-structured data. This process not only offers explainability into the decisionmaking process of the GNN but also informs the choice of model architecture by quantifying the contribution of node and edge information to overall performance. D. Neural Architecture Search We perform NAS-HPO to design an effective GNN architecture [61]. Our approach searches over a candidate configuration space , which consists of various combinations of network layers and hyperparameter settings. For the architectural hyperparameters, we allow 1–3 node and edge projection layers, where a single layer applies a linear transformation and using 2–3 layers enables richer, nonlinear representations. We also consider 1–4 GNN layers (e.g., E-GAT), since each layer performs one message-passing step by attending to neighbor information (and including edge features when available). Stacking more than four layers can lead to oversmoothing, so it is good practice to add residual connections for stability [62]. We set fusion and message MLP depths to 1–3 layers [63], with an option of 0 message MLP layers to disable the message transformation entirely. To combine neighbor messages into a single vector before updating a node’s representation, we use mean, sum, and max aggregators [64]. Finally, for activation functions, which allow GNNs and their internal MLPs to learn complex mappings beyond affine combinations, we include ReLU, LeakyReLU, and GELU as widely used options [65]. For each configuration θ∈, we train a GNN model fθ with parameters Wby minimizing the training loss W∗(θ)=arg min WLtrain(W,θ). After training, we evaluate each configuration on a validation set by measuring the F1-score F1(θ, W∗(θ)) and recording the training time T(θ). Our goal is to achieve both high predictive performance and low computational cost; hence, we formulate a multivariate optimization problem by simultaneously minimizing the negative F1-score (so that a higher F1-score corresponds to a lower objective value) and the training time. This is expressed by finding the Pareto front θ∗=ParetoFront−F1θ,W∗(θ),T(θ):θ∈. To efficiently explore the configuration space, we employ the tree-structured Parzen estimator (TPE) as implemented in the TPASampler [66]. TPE partitions the space based on a threshold y∗and estimates two conditional densities l(θ)=pθ|f(θ)<y∗and g(θ)=pθ|f(θ)≥y∗. Here, l(θ) is the probability density function θgiven that its performance metric f(θ) is good; and g(θ) is the probability density function of configurations θgiven that their performance f(θ) is not good. The next candidate configuration θnew is then selected by maximizing the ratio θnew =argmax θ∈ l(θ) g(θ) which serves as a surrogate for the expected improvement, directing the search toward configurations likely to yield higher F1-score with lower training times. E. Model Training and Evaluation Each GNN has its own characteristics and advantages. By adapting and testing different GNNs, we can evaluate which algorithms perform best for our problem, helping us identify the most suitable approach for the data. For the training hyperparameters, we consider the following ranges and options: the hidden dimension, representing each node’s intermediate embedding size, varies between 32 and 256, which covers the spot for many graph datasets [67].We apply dropout with rates up to 0.5 during training to randomly 45242 IEEE INTERNET OF THINGS JOURNAL, VOL. 12, NO. 21, 1 NOVEMBER 2025 mask node activations and include batch normalization at each layer to stabilize hidden activations and reduce internal covariate shift, accelerating convergence, since this combination markedly improves graph classification performance [68]. We set the learning rate to range from 10−4to 10−2,a common default for GNN training [69]. Finally, we include undersampling as an on/off option to balance classes [70]. It is important to select various metrics to compare algorithm performance on test data, such as F1-score, recall, and precision [71]. Recall measures the proportion of actual positive cases that are correctly identified, highlighting the model’s ability to detect all relevant instances. Precision quantifies the proportion of predicted positive cases that are truly positive, which is important for evaluating how many false positives occur. The F1-score, as the harmonic mean of recall and precision, provides a balanced measure that accounts for both false positives and false negatives. Additionally, for visualization purposes, we also consider the false positive rate and plot ROC curves to further assess the models’ performance. IV. EXPERIMENTS:SMART VILLAGE CASE STUDY We evaluate our proposed methodology for forecasting repeated visitation in a smart village context. The study area is located in the Alpujarra region of southern Spain and comprises three villages, such as Pampaneira, Capileira, and Bubion. We represent tourist visitation routes as graphs and perform graph classification to predict whether a tourist will return to the area. A. Graphs Preprocessing and Representation We collected license plate data using ANPR devices as vehicles entered or exited the area. The IoT scenario, illustrated in Fig. 2, consisted of four Hikvision IP cameras that leverage deep-learning-based ANPR technology. Each camera is equipped with vehicle detection sensors, a 2MP resolution, varifocal lenses (2.8–12 mm), and infrared LEDs with a 50-m range, ensuring complete coverage of vehicular movements. The ANPR devices recorded license plate numbers along with timestamps and the movement direction for each vehicle. After data collection, we preprocessed the raw data (see Table II) by correcting incomplete routes by merging partial capture records based on matching license plates and plausible time gaps, ensuring that each vehicle’s passage through the area was fully reconstructed [72]. In our study, we use four cameras: PAM1 and PAM2 at the entrance and exit of Pampaneira, BUB at the entrance of Bubion, and CAP at the entrance of Capileira. We recorded data from February 2022 to December 2024; the first six months were used for the experiments, while the remaining 2 years and 5 months were used to label the dataset for validation by determining whether a tourist returned. This dataset covers 4 67773 different vehicles and includes entry and exit times, node appearances, travel times between nodes, and the capture direction. We have made this real-world dataset from our smart village area publicly available in [73]. TABLE II RAW DATASET Fig. 2. Study area with the ANPR position. Fig. 3. Study area as a graph. We preprocess the data to represent our study area as a graph (see Fig. 3). In this representation, each tourist route is represented as a graph G=(V,E,Xv,Xe). 1) Each of the four cameras becomes a node v ∈V= {PAM1, PAM2, BUB, CAP}and is associated with a feature vector xv∈Rd, obtained by concatenating a onehot encoding of the camera identifier with a directional attribute, indicating whether the vehicle is entering (1) or exiting (0) through that camera. In this way, Xv= {xv:v∈V}collects all node feature vectors. 2) To form edges (vi,vj)∈E, we connect cameras in the exact sequence that a vehicle is detected; for example, if a vehicle appears first at PAM2 and then at BUB, there is a directed edge from PAM2 to BUB. Each edge carries an attribute xeij ∈Rpequal to the travel time between those two detections, calculated as the difference between their timestamps. Xe={xeij :(vi,vj)∈E} collects all edge feature vectors. Because direction is recorded at each node, all edges are directed. Finally, linking all camera detections builds the graph of the vehicle’s route trip. This fixed graph representation preserves the sequential pattern of each route and serves as the input for our binary graph classification task, where the target variable Yindicates whether the tourist will revisit the area. A tourist is considered a nonrepeater if they do not return (only one visit is recorded), and a repeater if they make more than one visit in the future. DURÁN-LÓPEZ et al.: ROUTE OPTIMIZATION IN SMART VILLAGES 45243 TABLE III CONFIGURABLE PARAMETERS B. Graph Neural Network Using these graph objects, we perform the graph classification task to predict whether a tourist will revisit the area using solely the graph information. To address this task, we train different GNNs to evaluate their performance on this task, including spectral-based models such as GCN [4], spatialbased models like GraphSAGE [16], EdgeConv [31], and MPNN [30], attention-based models, such as GAT [10], and recurrent models like GGNN [9]. In addition, we incorporate their edge-enhanced variants E-GCN [11], E-GraphSAGE [1], E-GAT [32],[33], and E-GGNN [35],[42] that leverage edge feature information to enrich the representation learning process, allowing us to assess the impact of incorporating edge features on classification performance. C. Explainability We apply GNNExplainer to our best-performing model to uncover the importance of both node and edge features in driving its predictions. By generating soft masks for nodes and edges, the explainer identifies which features and connections are most influential, allowing us to determine whether incorporating edge features yields a significant benefit. D. Neural Architecture Search For NAS-HPO (see Table III), we optimize our best performing GNN architecture by tuning both the layer configurations and training hyperparameters, following the settings detailed in Sections III-D and III-E. We employ Optuna [74] using the TPESampler along with a Hyperband pruner, running 200 trials. In each trial, we train the model for 30 epochs with a batch size of 64, with early stopping triggering if the validation F1-score does not improve over 5 consecutive epochs. Based on the evaluations (see Table IV), the E-GAT model outperformed other GNN models. Motivated by these findings, Table III presents the configurable architecture layers and parameters that define the search space for NAS-HPO. The works [32],[33] present similar E-GAT architectures, both of which include node and edge projection layers, E-GAT layers, and fusion MLP layers; however, we extend the architecture by incorporating residual connections to improve gradient flow and facilitate training [75]. While these works use sum as the aggregation function, we also experiment with mean and max. Additionally, although they use LeakyReLU, we further explore GELU and ReLU as activation functions. Moreover, none of the papers mention applying dropout within the MLP or attention layers [76], nor Batch Normalization [77], which could be beneficial for model regularization and faster convergence. E. Model Training and Evaluation In our approach, we employ the AdamW optimizer [78] together with the BCEWithLogitsLoss function [79] for training our GNN models. To address class imbalance, we compute class weights and incorporate them into the loss function. We use a fixed batch size of 64, and train the models for up to 30 epochs, with early stopping triggered after 5 consecutive epochs without improvement. We tune the remaining training hyperparameters of each GNN over the ranges described in Section III-E (listed in Table III) and select the optimal settings using Optuna. We implement our experiments in Python 3.12.3. The dataset is split in a stratified manner into 70% for training, 15% for validation, and 15% for testing. The validation set is used to monitor the loss and determine the best threshold that maximizes the F1-score based on the precisionrecall curve. After training, we save the best-performing model and evaluate it on the test set using the classification report. V. RESULTS AND DISCUSSION Table IV summarizes the classification performance for our two classes, reporting the precision, recall, and F1-score for each class along with the weighted averages across all evaluated models. The experimental results indicate that traditional models based solely on node features and graph structure do not perform well. Models such as GCN, GraphSage, and GGNN perform poorly, with precision, recall, and F1-score