ST-SplitVFL: Spatio-Temporal Split Vertical Federated Learning
Abstract
This paper introduces ST-SplitVFL, a novel framework for Spatio-Temporal Split Vertical Federated Learning, which extends vertical FL with split learning and spatial graph constraints. In this setting, clients maintain their local data while selectively exchanging encoded intermediate representations with geographically or logically connected peers.
Full text
ST-SplitVFL: Spatio-Temporal Split Vertical Federated Learning Anita Graser1[0000−0001−5361−2885], Jose Antonio Lorencio-Abril1[0009−0005−9127−4844], Axel Weissenfeld1[0000−0002−7246−2744], and Anahid Wachsenegger1[0000−0002−8889−0735] AIT Austrian Institute of Technology, 1210 Vienna, Austria [email protected] Abstract. The increasing demand for accurate spatio-temporal forecasting in domains such as mobility and urban management is often constrained by data privacy regulations and institutional reluctance to share sensitive information. Federated Learning (FL) offers a promising paradigm by enabling collaborative model training without centralized data aggregation. This paper introduces ST-SplitVFL, a novel framework for Spatio-Temporal Split Vertical Federated Learning, which extends vertical FL with split learning and spatial graph constraints. In this setting, clients maintain their local data while selectively exchanging encoded intermediate representations with geographically or logically connected peers. We evaluate ST-SplitVFL on two datasets and demonstrate that ST-SplitVFL consistently outperforms locally trained models and achieves predictive performance close to centralized and fully connected FL baselines, while substantially reducing model size and communication overhead. By preserving data locality, minimizing network congestion, and supporting multi-target forecasting, ST-SplitVFL provides an efficient and privacy-preserving solution for spatio-temporal learning across distributed stakeholders. Keywords: GeoAI ·Mobility Data Science ·Spatio-temporal Forecasting 1 Introduction The rapid progress of Artificial Intelligence (AI) and its broad uptake across industry have motivated public authorities and companies to deploy data-driven services for citizens and end users. Forecasting models, based on historical time series are one popular options since they have the potential to enable better planning and preparedness for future conditions. However, AI system training typically requires access to large, high-quality datasets that many organizations do not possess [19]. Even when data exists, regulatory requirements and reluctance to disclose proprietary information often constrain data sharing. These factors hinder the implementation of AI solutions that require centralized data aggregation.
2 A. Graser et al. Federated Learning (FL) addresses these challenges by enabling collaborative model training without centralizing raw data [15]. Instead of uploading data to a server, clients compute local updates or intermediate representations and share only these with a coordinator. This decentralization mitigates privacy risks and reduces the attack surface associated with single points of failure. Two primary FL paradigms are Horizontal Federated Learning (HFL) and Vertical Federated Learning (VFL) [7] (also known as Sample-based FL or Crosssilo FL and Feature-based FL, respectively). HFL applies when parties hold the same feature space for disjoint sets of entities; VFL applies when parties hold complementary feature spaces for (partially) overlapping entities. VFL is especially relevant when different organizations maintain feature-specific data that should remain private – for example, clinical records at a hospital and demographic records at a municipality for the same patient. Such settings frequently arise in spatio-temporal forecasting, where multiple stakeholders collect geographically distributed data over time and wish to obtain joint insights without exposing raw data. GeoAI combines machine learning with geography and GIScience to support geospatial analysis and forecasting [6,23]. While FL and GeoAI are natural complements, their intersection remains underexplored; early studies nevertheless point to strong synergies [5,24]. Spatio-temporal data introduce additional complexity: they couple temporal dynamics with spatial dependence and raise specific privacy risks [20]. When such data are fragmented across organizations, effective learning without centralization is particularly challenging. This work introduces a Spatio-Temporal Split Vertical Federated Learning framework (ST-SplitVFL) tailored to forecasting of space–time series. The framework extends VFL by incorporating spatial relations among clients via a spatial graph, enabling joint use of temporal histories and spatial context while keeping raw data local. In our setting, clients hold distinct forecasting tasks and interrelated datasets. ST-SplitVFL leverages these interconnections to improve predictive performance without compromising data locality. The main contributions are: –We introduce ST-SplitVFL, a split VFL approach for space–time series forecasting that connects spatially related clients through a spatial graph. The graph defines a SplitNN in which each client performs local forecasting with LSTM-based modules while selectively exchanging intermediate embeddings with geographically or logically adjacent clients. –We conduct comprehensive experiments showing that ST-SplitVFL outperforms locally trained models and approaches the performance of centralized models, while reducing model complexity and network congestion. –We provide a theoretical argument that graph-constrained sharing yields smaller models whose size grows linearly in the number of clients, and empirical evidence that this reduction does not degrade performance. The remainder of this paper is organized as follows: Section 2 reviews related work on spatio-temporal forecasting and federated learning. Section 3 presents
ST-SplitVFL: Spatio-Temporal Split Vertical Federated Learning 3 the ST-SplitVFL framework. Section 5 describes the experimental setup and results, and discusses implications and applications. Section 6 concludes. 2 Related Work Both previous works on spatial time series forecasting as well as on spatiotemporal federated learning research are relevant for the development of STSplitVFL. A detailed survey of Spatio-temporal ML is presented in [23], with the authors reviewing how NNs have been made spatially-aware for different tasks. In the following subsection, we focus on those papers that address specifically spatial time series. 2.1 Spatial Time Series Forecasting When predicting time series at different spatial locations, making models spatiallyexplicit can improve forecast results. Table 1 provides an overview of existing neural network approaches for spatio-temporal forecasting and how they approach modeling the spatial and temporal dimensions. Convolutional NNs (CNNs) are a common approach to build spatial models by performing a convolution operation on the spatial dimension. This represents the assumption that the spatial relationships should be similar in different locations, i.e., similar patterns in different areas should have similar effects. For example, [28][8] process each location’s time series independently using temporal convolutions and time segmentation. Then the results are processed to extract spatial relationships through a spatial convolution. Asadi et al. [1] replace the temporal CNN with recurrent NNs (RNNs). CNNs are also common in approaches that model spatial data as images [14,26]. Spatial Modeling Temporal Modeling Reference CNN [28,8] CNN on temporal modeling output RNN [1] CNN on images that represent spatial data RNN [14,26] GNN on spatial graph RNN [18,10] GNN on spatio-temporal graph [2] Global attention on temporal modeling output Local attention mechanism [3,25] Table 1. Classification of the approaches for spatio-temporal forecasting. Graph NNs (GNNs) present a different way to model spatial correlations by encoding the spatial information using a spatial graph. GNNs can process
4 A. Graser et al. the spatial graph representing the spatial information, while temporal data is processed with RNNs [18,10]. Instead of separating the processing of spatial and temporal features, GNNs can also be applied to unified spatio-temporal graphs [2], extracting spatial and temporal relationships simultaneously. Transformers [21] can also be used to tackle spatio-temporal forecasting, by using the attention mechanism on the spatial and temporal dimensions [3,25]. As this summary shows, in most works, the spatio-temporal dependencies are analyzed separately as temporal dependencies and spatial dependencies, and combined at some point to effectively model the mixed nature of the data. This approach can reduce the complexity of the model but could miss possibly relevant interactions. On the other hand, the completely spatio-temporal approach enables to take into account different dependencies, both spatial and temporal and facilitates multi-granular analysis, however, the model can become quite involved. For federated learning, the two-step approach is more appropriate, since each FL client holds only local temporal data, and spatial dependencies can only be detected with access to multiple clients’ data. We detail how we deal with the spatio-temporal dependencies in Section 3, and continue now to review how FL has been made spatially-explicit in the existing literature. 2.2 Spatio-Temporal Federated Learning FL can be adapted to cope with spatial relationships in different ways. Some of these are very natural, e.g. each local model is adapted to the local specificities of the data, while others require an involved process of constructing a graph or an attention mechanism. Table 2 summarizes different approaches for spatiallyexplicit FL. An early spatially-explicit HFL setting was proposed in [4], enabling the global model to learn general patterns in the data while allowing local models for customization to the local regional characteristics. A similar approach is followed in [16]. In these two HFL approaches, the global model is an aggregation of the local models, and therefore the customization depends on the local updating of the local data. Alternatively, GNNs can be used to model the interdependencies [22]. VFL approaches process each data source with a different NN module, which is locally trained at each cell [27,11]. The model parameters are shared with other clients and locally aggregated. Therefore, this approach leverages spatial information by looking at local contextual spatial information (weather conditions and POIs, shared by SplitVFL) and by sharing model parameters taking into account the spatial component. In the next chapter, we develop our proposed solution, which follows the natural idea of the split learning setting of local models coping with local specificities, but adding an extra level of locality by sharing only embeddings according to the topology of a spatial graph representing the spatial relationships between clients. In addition, our proposed framework copes with a multi-target environment, where each client may have different prediction needs, while to the
ST-SplitVFL: Spatio-Temporal Split Vertical Federated Learning 5 Spatial Modeling FL Type Global Local Reference Local training [4,16] HFL Through GNNs [22] Mix Local models aggregation Attention mechanism between client embeddings [27] VFL Centralization of embeddings Split learning local models [11] Table 2. Classification of the approaches for spatially-explicit FL. best of our knowledge, there is no previous research tackling this problem in a spatially-aware federated setting. 3 Spatio-Temporal Split Vertical Federated Learning Building on insights from our previous work-in-progress article [13] and the approach proposed by Lie et al. [12], we develop a variant that we denote STSplitVFL. It enables collaborative model training across multiple clients by constructing a global model capable of generating predictions locally at each client site. The model incorporates information from other clients through secure aggregation mechanisms, ensuring that raw data remains confined to local databases and that privacy is preserved throughout the training process. For a set of clients C={C1, ..., Cn}, each client holds a local data set Xci, i = 1, ..., n and is placed in a location lci∈ L (which could possibly repeat). Clients are either active if they have a forecasting task or passive if they do not have a forecasting task (and thus only serve as data providers). Active clients hold two neural modules: 1) multi-head encoder Ei,j with i= 1, ..., n and j∈ N(Ci), where N(Ci)are the neighbors of client Ciin the spatial graph (defined in the following paragraph), including itself; 2) forecaster Fi, i = 1, ..., n, responsible for generating predictions ˆ Yi. In contrast, passive clients serving as covariates maintain only the encoder module and do not perform forecasting tasks themselves. This approach is exemplified in Fig. 1, which depicts two active clients and one passive client. The illustration highlights how data exchange is structured based on spatial relationships among the clients. In particular, only encoded representations are exchanged between clients—not the original raw data—which is a key aspect in ensuring data privacy throughout the process. Firstly, we build a spatial graph of the system, G= (C,E), where Eis the set of edges connecting those clients. The spatial graph encodes the spatial relationships between the clients, determined by their locations. E.g. clients Ciand Cj are connected in the graph, if the distance between lciand lcjis smaller than some pre-defined threshold. Secondly, the topology of the spatial graph determines the architecture of the SplitNN for the system. Each client adds one head to the encoder for each neighboring client, and the intermediate results will be sent only to these neighboring
6 A. Graser et al. Fig. 1. Illustration of the Split VFL approach, with active clients A and B, and passive client C. Client C is only connected to client B. Dotted arrows represent information transmitted through the network, and continuous arrows represent local data flow.
ST-SplitVFL: Spatio-Temporal Split Vertical Federated Learning 7 clients. That is, each client Ciproduces Zi,j =Ei,j(Xi), and receives Z′ i= [Zj,i : (j, i)∈ E]. Finally, the forecaster module of each client receives all intermediate results from neighboring clients, concatenated, and uses them and the local data to make forecasts ˆ Yi=Fi(Z′ i). To process the temporal data, both the encoder and the forecaster are LSTM networks, while the spatial component is taken into account by defining the topology of the network according to the spatial graph, G. We illustrate this ST-SplitVFL framework in Figure 1, where we show how the number of heads varies according to the connectivity of the graph, as well as the different roles of active and passive clients. To train the system, a two-phase approach is followed: 1. Encoders pre-training: each client trains an encoder with its local data in an encoder-decoder fashion. Once it is trained, it makes as many copies as neighboring clients it has. These encoders are updated during the second phase. 2. Federated training: each batch is encoded by each client, and sent to its neighbors. All incoming encoded data are concatenated and input to the forecaster, which makes predictions. The loss is computed and the forecaster is updated via backpropagation. The gradients are sent to the appropriate clients to fine-tune each head of the encoder to the associated forecasting task, following the inverse path that the data followed in the forward pass, as shown in Figure 2. Formally, if encoded representations from client Ci, head Ei,j,Zi,j =Ei,j(Xi), are used by client Cjto make forecasts, ˆy=Fj([..., Zi,j, ...]), then we compute the loss of client Cjas Lj(yj,ˆyj), update Fjaccording to ∂Lj ∂WFj , where WFjare the weights of Fj, and transmit ∂Lj ∂WFj,Zi,j to client Ci, so that this client can compute ∂Lj ∂WEi,j and update the local encoder Ei,j associated to client Cj. 4 Experimental Setup 4.1 Datasets Experiments were performed on two different datasets: the public Porto dataset from Kaggle (originally released as part of the ECML/PKDD 2015 Discovery Challenge [17]) and a proprietary dataset from Scheveningen (Netherlands) that was part of our project. The basic dataset statistics are summarized in Table 3. The Scheveningen dataset consists of six areas with crowd data and an additional data source providing weather forecasts. Each client has 22,202 records.
8 A. Graser et al. Fig. 2. Illustration of the backpropagation process in SplitVFL, focusing on the loss of client B, and how the different gradients flow backward, represented as red arrows. The number of connections of the spatial graph is either four or five (depending on the location of the area). The Porto dataset provides the opportunity to test on a larger geographic region. It encompasses taxi trajectory data collected within the city of Porto, Portugal. To align the dataset with our framework, we preprocess it by aggregating spatial and temporal information. Spatially, we partition the city using the H3 grid system [9], resulting in 78 distinct hexagonal areas and 8,760 records per client. Temporally, the data is segmented into hourly intervals. Within each area, we compute the hourly count of active taxis. For the predictive task, we use data from the preceding four days to forecast taxi activity over the next four hours. To model spatial relationships, we construct a graph by connecting each hexagon to its six nearest neighbors. Table 3. Characteristics of the datasets after pre-processing: number of clients, records per client, and input/output sizes. Use case Types of clients No. clients Recs./client Input size Output size Scheveningen Crowd areas 6 22202 168 ×3 168 Weather 1 168 ×6Porto Taxi counts areas 78 8760 96 ×3 4
ST-SplitVFL: Spatio-Temporal Split Vertical Federated Learning 9 4.2 Training Each of the data sets is processed with a window function that creates input and output sequences for training and validation, with an 80/20 temporal split (first 80% of the data used for training, and final 20% for validation). As described in Section 3, ST-SplitVFL clients have a forecasting model and multiple encoders. We used a two-layer LSTM encoder and a two-layer LSTM decoder, each with a hidden state size of 16. The forecaster is implemented as a bidirectional LSTM with a hidden state size of 64. The model is optimized using Adam optimizer, with a dropout rate of 0.1, an initial learning rate of 1 x 10−2, and a learning rate scheduler that reduces the rate by half every 20 epochs. The batch size is 64 for Scheveningen and 32 for the Porto dataset. We used the standard loss function MSE (mean square error) and trained for 200 epochs. To improve training stability, the encoders are first pre-trained on local data. We then perform joint training, where the forecaster models are trained, and the encoders are fine-tuned simultaneously. Our networks were trained from scratch using PyTorch. All experiments were conducted on a system equipped with an NVIDIA TITAN RTX GPU with 24GB RAM. The training was repeated over 11 independent runs in order to provide a sound, statistically reliable estimate of the model performance. 4.3 Baseline Models For comparison, we trained two baselines: local LSTM models and SplitVFL. For the first baseline, we trained models where each client independently using only its local data, without any communication or information exchange with other clients. Hence, clients do not have any encoders but only a forecaster with bidirectional LSTM as described in Sec. 4.2 and where the input size adopted accordingly. We refer to this setup as the Local LSTM model. For the second baseline, the SplitVFL model is based on the ST-SplitVFL model but without the spatial graph, which defines the connections between clients. The client connections of the SplitVFL models slightly differ depending on the dataset (Scheveningen or Porto) they are trained on. The clients of the SplitVFL model trained on the Scheveningen dataset maintain full interconnectivity with one another. Since there are only a few clients in the Scheveningen dataset, the difference between full interconnectivity and the spatial graph is small - seven connections in the case of full interconnectivity versus four or five connections using the spatial graph. In contrast, the Porto dataset has 78 clients. In this case, the clients of the SplitVFL model were randomly connected with 39 other clients. Note, that the number of connected clients needed to be restricted to 39 because of the limited GPU capacity for training the models. The trainings of all models were carried out in the same way as described in the previous section.