Full text
SIARGKAS ET AL.: TMR BASED ON LOW-RATE ACCELERATION AND LOCATION SIGNALS WITH AN ATTENTION-BASED MIL NETWORK 1 Transportation mode recognition based on low-rate acceleration and location signals with an attention-based multiple-instance learning network Christos Siargkas, Vasileios Papapanagiotou,Anastasios Delopoulos Abstract—Transportation mode recognition (TMR) is a critical component of human activity recognition (HAR) that focuses on understanding and identifying how people move within transportation systems. It is commonly based on leveraging inertial, location, or both types of signals, captured by modern smartphone devices. Each type has benefits (such as increased effectiveness) and drawbacks (such as increased battery consumption) depending on the transportation mode (TM). Combining the two types is challenging as they exhibit significant differences such as very different sampling rates. This paper focuses on the TMR task and proposes an approach for combining the two types of signals in an effective and robust classifier. Our network includes two sub-networks for processing acceleration and location signals separately, using different window sizes for each signal. The two sub-networks are designed to also embed the two types of signals into the same space so that we can then apply an attention-based multiple-instance learning classifier to recognize TM. We use very low sampling rates for both signal types to reduce battery consumption. We evaluate the proposed methodology on a publicly available dataset and compare against other well known algorithms. Index Terms—transportation mode detection, accelerometer, location, GPS, multiple instance learning, attention, hidden Markov model I. INTRODUCTION To discover solutions and uncover opportunities to improve the quality of life at scale in cities, it would be useful to acquire a better understanding of commuting patterns and establish new grounds of collaboration between the transportation sector and other fields of research like biomedical research [1]–[3], urban and transportation planning [4], public-policy making, environmental research [5], [6], carbon footprint analysis [7], safe driving, and journey planning. Transportation mode recognition (TMR) systems can play an integral part in this by automating the process and eliminating the need of manual data collection and annotation. Moreover, real-time recognition can Christos Siargkas and Anastasios Delopoulos are with the Multimedia Understanding Group, Dpt. of Electrical and Computer Engineering, Faculty of Engineering, Aristotle University of Thessaloniki, Greece. Vasileios Papapanagiotou is with the IMPACT research group, Dpt. of Biosciences and Nutrition, Karolinska Institutet, Stockholm, Sweden and the Multimedia Understanding Group, Dpt. of Electrical and Computer Engineering, Faculty of Engineering, Aristotle University of Thessaloniki, Greece. E-mail: [email protected], v[email protected], antelopo@ ece.auth.gr Manuscript received August 21, 2023; revised February 13, 2024. This research was funded by the European Union’s H2020 programmes under Grant Agreement No. 965231 (REBECCA - “REsearch on BrEast Cancer induced chronic conditions supported by Causal Analysis of multi-source data”). assist in identifying crucial moments and providing immediate assistance. Given the rich arsenal of sensors available in common smartphone devices (i.e., smartphones and smartwatches), data fusion has become an integral part of TMR research, by combining multiple sensors to observe the same event from different viewpoints. Acceleration signal can elaborately distinguish between physical activities which are characterized by volatile body movement, however, the ability to distinguish between motorized transportation modes (TMs) that are characterized by low body movement is lesser [8], [9]. In particular, when using only acceleration signal, classifiers can recognize activities such as standing still, walking, and running robustly, but exhibit significant ambiguities between car vs. bus, train vs. subway, and standing still vs. train/subway. Findings of [8] align with these limitations, highlighting the model’s constrained ability to differentiate between Subway and Train. Contrary, location signals can capture a “higher-level” overview of the user’s movement, including distance, speed, etc; however, relying location signals requires uninterrupted interaction with satellites which is not always possible (i.e., when the user moves underground, inside a building, or along urban canyons). The inherent heterogeneity between these two sensor modalities presents an opportunity for highly-effective fusion, yet it also introduces a challenge for achieving seamless combination. Furthermore, the vastly different sampling rates used for collecting the two sensor modalities introduce another layer of complexity in robustly combining them. In this work, we propose an intermediate-fusion deeplearning TMR model which combines two sensor modalities: acceleration and location. To address the challenges posed by variations in these two sensor modalities, including heterogeneity, distinct sampling rates, and sensor unavailability, we propose a novel model that learns independent spatio-temporal features for each modality while simultaneously mapping the heterogeneous sensor data to a shared low-dimensional embedding space. Leveraging this common embedding space, we efficiently integrate information from the diverse sensors using an attention-based multiple-instance-learning (MIL) classifier. Finally, we employ a Hidden Markov Model (HMM) for postprocessing. To optimize the efficacy of MIL, we use sequences of acceleration windows instead of relying on single windows. This approach offers several advantages, such as enhancing the resolution of the acceleration input and identifying the most critical regions within the acceleration data that pertain arXiv:2404.15323v1 [eess.SP] 5 Apr 2024
SIARGKAS ET AL.: TMR BASED ON LOW-RATE ACCELERATION AND LOCATION SIGNALS WITH AN ATTENTION-BASED MIL NETWORK 2 to the specific TM being analyzed. We evaluate on a publicly available dataset and compare with our own implementations of other state-of-the-art algorithms. Our main contributions are: •We propose a novel, energy-efficient, and lightweight system for TMR that uniquely combines low samplingrate acceleration and location signals collected by smartphone sensors, optimizing energy consumption while maintaining high effectiveness. •We present Fusion-MIL, an attention-based MIL framework that effectively combines acceleration and location signals, overcoming challenges such as sensor heterogeneity, distinct sampling rates and sensor unavailability. We use multiple acceleration windows instead of single one, enhancing the resolution of the acceleration input and identifying regions most relevant to the final prediction. This approach inherits the interpretability of the instance-attention based MIL. •We introduce additional processes such as data preprocessing, feature engineering, data augmentation, and pre-training to boost effectiveness. •Extensive experimental evaluation is performed on a publicly available dataset, with a focus on cross-subject and cross-placement variability. These experiments provide in-depth insights into various factors influencing TMR, particularly the impact of device placement. Results demonstrate our method’s capability to accurately distinguish between eight different TMs in complex scenarios, surpassing the state-of-the-art methods and various alternative single-modal and multi-modal algorithms. The rest of the paper is organized as follows: Section II presents an overview of recent approaches for TMR based on inertial signals, location signals, or their combination. Section III presents our proposed approach and the architecture of the artificial neural network (ANN). Section IV presents the dataset, types of experiments, and details for model training, and Section V presents the results and experimental evaluation. Finally, Section VI concludes the paper. II. RELATED WORK Inertial and location signals offer different advantages and challenges when used for transportation mode recognition (TMR). In this section we present recent approaches for TMR from literature, organized based on their use of sensor type, with an emphasis on the requirements (i.e., sampling rate, battery consumption). A. Location-based approaches One of the most commonly used modalities for TMR is the location signal (i.e., geographical coordinates). In [10], James et al. explore this domain using the publicly available GeoLife Dataset [11] and process location trajectories (sampling interval of 1to 5sec) based on discrete wavelet transform (DWT) and deep learning techniques. In a different approach, Dabiri el al. [12] adjust all the location trajectories to a fixed length and design a CNN-based supervised travel mode identification system. This system extracts 4features from location data, including speed, acceleration, jerk, and bearing and identifies 5types of TMs. It is important to note that these studies, as well as the vast majority of the TMR studies, use the same set of users for training and testing, limiting the generalization potential to user variations such as different gait styles, walking/running patterns, speed, and travel behaviours. In [13], location-based TMR using a random forest (RF) is combined with geographic information system (GIS) to leverage the underlying transportation network. Location is sampled only once every 15 seconds to avoid battery drain. The trained classification model is deployed to the general public, accommodating new individuals for evaluation. Authors report an average detection accuracy of 93.5% at classifying 6TMs: car, bus, train, walking, biking and stationary. The 2021 SHL Recognition Challenge [14] also focuses on user-independent evaluation by separating participants in train and test sets (instead of separating segments of participant data). The contestant teams were provided with four radio sensor modalities, including: GPS location, GPS reception, WiFi reception, and Cell reception. These sensors were asynchronously sampled with a sampling rate of roughly 1Hz. Sekiguchi et al. [15], using an LSTM model with location-derived features (velocity, acceleration, and angular acceleration) and XGBoost with statistical features for when location signal was not available achieved the 5th highest F1-score of 58.2%. In [16], Balabka and Shkliarenko extract 975 hand-picked features from the 4 sensors and employ an AdaNet model, provided by Google AutoML service, taking the first place with F1 score as high as 75.4%. B. Inertial based approaches Location signal is not always available (especially when it is derived from GPS) in shielded areas; additionally, samplingrates can quickly drain the smartphone battery. Thus, several studies propose only inertial-based only approaches [17]. Most of them only rely on the accelerometer due to its low energy requirements and early integration into smartphones. In [18], 3D acceleration signals are collected from 4participants holding their phones freely in any orientation of their preference. Acceleration signal is captured at 50 Hz. After filtering and computing the signal magnitude, a CNN model discriminates between 7TMs: stationary, walk, bicycle, car, bus, subway and train. This approach achieves an accuracy of 94.48% by splitting the data into 80% -20% training and test sets respectively. Nonetheless, like several other studies, this study employs the same group of users for both training and testing. Contrary, Rosario et al. [19] observe that when a classifier is deployed on data of an older participant after being trained on a younger participant, its effectiveness considerably declines. In [20], 3common body placements are considered and acceleration is recorded both at 60 Hz and 100 Hz. A hierarchical classifier is proposed to identify 6TMs. This approach is evaluated using a leave-one-subject-out (LOSO) crossvalidation (CV) and a leave-one-placement-out CV, achieving F1-scores of 81% and 83.3% respectively.
SIARGKAS ET AL.: TMR BASED ON LOW-RATE ACCELERATION AND LOCATION SIGNALS WITH AN ATTENTION-BASED MIL NETWORK 3 However, some vehicles such as automobiles and buses (or subways and trains) produce similar patterns, thus limiting the discriminative power of acceleration-only based systems [21]. To counter this, some studies propose combining multiple inertial sensor modalities. Fang et al. [22] uses 8,311 hours of accelerometer, magnetometer, and gyroscope sensor data recorded from 224 users with a sampling rate of 30 Hz [23]. A deep neural network (DNN) is employed to recognize 5TMs (still, walk, run, bike, and vehicle), achieving accuracy of 95%. Yet, a key challenge in TMR, especially when relying solely on inertial sensors, is the discrimination ability between similar motorized classes [24], [25]. In Zhao et al. [26], a Bi-LSTM model is trained on accelerometer and gyroscope data, sampled at 50 Hz, to detect 6TMs: still, walk, run, bike, bus, subway. The model is trained on data from 8users and is evaluated on 3separate users, achieving an overall accuracy of 92.8%. However, all data is collected from smartphones with a fixed orientation and placement, limiting generalization to other body positions. In fact, according to [27], [28], inertial data may differ significantly across different body positions. In [8], Tang et al. use a deep multimodal fusion network on accelerometer, magnetometer, and gyroscope data sampled at 20 Hz and segmented into 60-second windows. Their proposed method achieves a remarkable accuracy of 95.1% and an F1-score of 94.7% in detecting eight different transportation modes. The 2018 SHL recognition challenge [29] focuses on timeinvariant evaluation using 62 days (272 hours) as training data and the remaining 20 days (95 hours) as test data, while providing the participant teams with 7different motionsensor modalities. Ito et al. [30] apply a deep CNN to twodimensional spectrogram images extracted from the sensor sequences, taking the third place with F1-score 88.8%. In the position-independent 2019 SHL recognition challenge [31], Choi and Lee [32] achieve the second highest F1-score 75.9% by introducing a deep multi-modal fusion model named EmbraceNet that combines multiple inertial sensor modalities which are independently pre-processed by a CNN. C. Hybrid approaches Compared to inertial sensors, location sensors can offer a “higher” overview of the user’s movement, thus serving as a complementary data source when used alongside inertial sensors. Moreover, location sensors are mostly position invariant and can be effectively combined with inertial sensors to minimize the impact of the varying smartphone placement and orientation. Consequently, a combination of location and inertial sensor modalities can enhance the robustness of a TMR system and improve its discrimination abilities. In fact, according to Wang et al. [9], the combination of location and accelerometer outperforms the combination of three inertial sensor modalities (accelerometer, gyroscope, and magnetometer), while the combination of all four modalities only improves the recognition effectiveness slightly. Additionally, [9] also emphasize that the combination of location and acceleration signals remarkably improves the recognition accuracy for each class compared to relying solely on accelerometer data. Notably, the accuracy of distinguishing between Car and Bus, as well as between Train and Subway, shows significant improvement with the combined use of location and accelerometer data which is what we propose in this work. In [33], Reddy et al. introduces a hybrid approach using both location and acceleration data captured from 6different smartphones (worn in 6different body positions) by 16 participants. Each smartphone contains an accelerometer that can sample at 32 Hz and a built-in GPS receiver that can sample at 1Hz. The proposed classifier consists of a DT followed by a discrete hidden Markov model (HMM) and recognizes 5TMs: still, walk, run, bicycle, motorized. The classifier is evaluated with LOSO CV, achieving an average and a minimum accuracy of 93.6% and 88.2% respectively. However, this method can’t handle location-signal loss. In [34], location (GPS) trajectories are resampled evenly with a frequency of 1Hz and the accelerometer readings are resampled to a frequency of 50 Hz. The proposed approach copes with location signal loss by including positioning data obtained from the cell network and relying solely on accelerometer features when the trajectory cannot be reconstructed with sufficient accuracy. A randomized ensemble of classifiers with an HMM is proposed to differentiate among 8 different TMs: bus, car, bike, tram, train, subway, walk, and motorcycle. Two experiments are presented: a 4-class task (walk, car, bus, bike) where the classifier achieves a high recall accuracy of 91%, and a 5-class task (walk, car, bus, subway, train) where classification effectiveness is degraded (recall accuracy of 76%). In [35], Priscoli et al. use accelerometer (50 Hz), gyroscope (50 Hz), and location (1Hz) signals to recognize seven TMs: walk, car, motorbike, tram, still, bus, and subway. This is the only work, in our knowledge, that proposes a deep learning TMR system based on hybrid information from both location and inertial sensor modalities. Specifically, recursive neural networks (RNNs) with statistical features achieve the second best effectiveness (88%), while a CNN model with 3D acceleration, 3D gyroscope, and speed achieves 98.6% without any pre-processing. In [36], location, accelerometer, and heart rate data are collected from 126 individuals over a period of 7days. Random Forest (RF) algorithms are used to perform TMR on a minute-by-minute basis. This study focuses on the participantlevel split and the proposed method demonstrates varying prediction rates, achieving a prediction rate of 95% for biking, 65% for public transport, and an overall prediction rate of 90%. Notably, this study, like the vast majority of hybrid studies, use early fusion to combine these sensor modalities by either extracting statistical features or concatenating the raw signals. To date, most approaches use location and inertial sensor data collected at high sampling rates; in most studies, location sampling rate falls within the range of 1to 5seconds [37], [38], while the sampling rate of inertial data ranges between 20 and 30 Hz [39]. Such requirements tend to have a strong impact on modern smartphones energy consumption. Some studies have suggested that inertial sensors sampled at 20
SIARGKAS ET AL.: TMR BASED ON LOW-RATE ACCELERATION AND LOCATION SIGNALS WITH AN ATTENTION-BASED MIL NETWORK 4 Hz can sufficiently discriminate between various TMs [40] and a sampling rate of 10 Hz is adequate for distinguishing between different mobility patterns [41]. Notably, one study demonstrated that reducing the sampling rate from 100 Hz to 12.5Hz led to a threefold increase in the duration of data collection on a single battery charge [42]. Regarding location data, power consumption decreases with increased sampling time and it is estimated that employing a sampling time of 40 seconds can, on average, extend the mobile phone battery life by 25.4% compared to the standard sampling setting of 1Hz [43]. Thus, one challenge in developing such a hybrid approach is to fuse the different sensor modalities efficiently and maintain low sampling-rate requirements. To address these challenges, we use accelerometer signals at 10 Hz and location signals at 1/60 Hz (one sample per minute) and propose a hybrid attention-based MIL system that extracts instance-level embeddings through single-sensor–based feature encoders and fuses the instance-level features into joint representations through a deep MIL mechanism [44]. The network’s ability of being selective enables it to effectively fuse the acceleration and location modalities, adeptly addressing situations where location signal is not available by solely relying on accelerometer data. Morevoer, the network’s interpretability capabilities offer a rich spectrum of insights into the TMR task. The classification task includes eight TMs: still, walk, run, bike, car, bus, train, and subway, and the model is evaluated with LOSO CV, achieving accuracy and F1-Score of 90.8% and 92.6%, respectively. III. MATERIALS AND METHODS Our approach is based on an artificial neural network (ANN) that separately processes windows of acceleration and location signal and then combines them with MIL. Each signal type is originally pre-processed and then windows are extracted. We use two different sub-networks for extracting features from each signal type and then map them to a common embedding feature space. A third sub-network is then used that performs attention-based MIL. Finally, we post-process the predicted labels using an HMM. A. Acceleration pre-processing While it is possible to use high sampling rates to capture acceleration in modern devices, continuous capturing at such rates can have a detrimental effect on battery life. Thus, we opt to sample acceleration at a low rate of fa s= 10 Hz. We extract consecutive windows of Ta= 60 sec. This choice is based on the assumption that TM does not change significantly within a single minute. Experimenting with the dataset also shows similar results (see Section V-B1). The accelerometer provides 3D measurements: a[n] = [ax[n], ay[n], az[n]]T. To remove the effect of phone and sensor orientation, we compute the magnitude a[n] = ∥a[n]∥ [45]. We also compute the magnitude of jerk as: j[n] = ∥a[n]−a[n−1]∥ · fa s(1) because it is orientation-independent and reflects body-related accelerations [46]. Thus, we obtain a signal of 2channels: [a[n], j[n]]T. It is important to note that we do not remove the gravity component from the mangitude a[n], as it proves to be a source of valuable information (based on early-stage experimentations with the dataset). However, the effect of gravity is indirectly removed when computing jerk, j[n]. Thus, our method takes advantage of both versions of the signal. Finally, we transform each of the 2channels into spectrograms, using the short-time Fourier transform (STFT). Power spectrum is estimated in 10-second segments with 9sec overlap, yielding 51 segments in total. We rescale the power spectrum bands into 51 bands by grouping frequencies; each band has double bandwidth compared to the previous one, i.e.: fb[i+ 1] −fb[i] fb[i]−fb[i−1] = 2, i = 1,...,50 (2) where fb[i−1] and fb[i]are the lowest and highest frequency of the i-th band. The spectrograms are then concatenated as two channels, resulting in a single “image” of 51 ×51 ×2. To reduce the risk of overfitting, particularly in cases when there is a considerable overlap between training samples, and to enhance the generalization capabilities of our ANN, we incorporate data augmentation. Typically, accelerometer signals are augmented using signal processing transformations such as jittering, scaling, rotation, and permutation [47], [48]. While these methods have shown improvements in recognition rates, they tend to disregard the correlations between signals, limiting the potential to uncover cross-signal features. We adopt an alternative augmentation policy, proposed by Park [49]. Inspired by cut-out [50], consecutive time steps or frequency channels within the spectrogram are randomly masked, reducing reliance on specific regions or features while preserving the intrinsic cross-signal correlations. In particular, we randomly choose 0,1, or 2frequency bands of random width (0to 5rows), and also randomly choose 0,1, or 2time intervals of random length (0to 5columns),and mask them. B. Location pre-processing Most studies use location data captured at high sampling rates (e.g. 1Hz). Such rates have a notable impact on battery usage, even in modern smartphones. Thus, we opt for a reduced sampling rate of fl s= 1/60 Hz, which translates to one sample per minute, in order to simulate a more batteryfriendly data collection approach. Capturing location data depends on GPS signal availability, resulting in certain sections of the time-series data being frequently absent due to poor satellite reception. We identify two types of “gaps”: (a) brief signal losses, lasting less than 3 minutes, for which we employ linear interpolation to estimate the missing values, and (b) extended periods of signal loss, exceeding 3minutes, during which we apply a masking value the missing values that signifies the presence of missing data. This approach is necessary since linear interpolation would yield erroneous data in such cases.
SIARGKAS ET AL.: TMR BASED ON LOW-RATE ACCELERATION AND LOCATION SIGNALS WITH AN ATTENTION-BASED MIL NETWORK 5 The signal is then segmented to windows, similarly to the accelerometer signal. However, we use windows of 12 minutes. This value of 12 minutes enables us to capture behavior that is not observable in shorter durations (i.e., moving through heavy traffic in a motor vehicle can create patters that last longer than 1minute). It should be also noted that choosing smaller window sizes (such as 1minute as in the case of the acceleration signal) would yield windows with very few signal samples (e.e., only one for 1-minute windows). The big difference in window lengths between the two modalities as well as the orders of magnitude of difference between the sampling rates creates a challenge when trying to combine tehm in a single detector. We face this challenge by designing our architecture appropriately (please see Section III-C for more details). Using raw location coordinates as input data is not practical as it may lead to the development of a location-dependent classifier. To overcome this we transform the coordinates into location-invariant features. Specifically, for a window of location signal, l[n]=(lat[n],lon[n]), at timestamps t[n]for n= 0,...,11, we consider 2features, namely speed as: vl[n] = d(l[n], l[n−1]) t[n]−t[n−1] (3) and acceleration as: al[n] = vl[n]−vl[n−1] t[n]−t[i−1] (4) where d(•,•)is the Havershine distance between two pairs of coordinates. Because of how vland alare computed, the result is a 10 ×2matrix. Additionaly, we calculate 5more features for each window: mean and standard deviation of speed, mean and standard deviation of speed derivative, and “movability” [51]. Movability indicates the ratio of total displacement to the total distance covered within the examined period: m=d(l[11], l[0]) k=11 P k=1 d(l[k], l[k−1]) (5) C. Embedding signals into a common space To effectively integrate the information from the two sensor modalities into a cohesive joint representation, we must account for the variations between these modal inputs, including their distinct nature, representation, sampling rate, and window size. To tackle these challenges we adopt a two-pronged approach, employing two parallel modality-specific branches to embed the information originating from acceleration and location into a common, embedding space. To achieve this, we project the data of Sections III-A and III-B onto Rd(where d= 256 in our case), using two feature encoders: •fa:R51×51×2→Rd, the acceleration-feature encoder •fl:R10×2→Rd, the location-feature encoder It is important to note that we have designed the two modalityspecific networks so that their output is the same, i.e., both produce vectors in Rdwhere d= 256. Thus, windows of both signal types are embedded in the same feature space, enabling the application of MIL. 0 5 6 7 12 Time (min) 1 2Target Acceleration Location Fig. 1. Visual representation of the windows that are used for a single bag (for MIL). Given a target timestamp (dashed line), we include nl= 1 window of location data that is 12 minutes long and na= 3 successive windows of acceleration data that are each 1minute long. The next bag is obtained by shifting everything by 1minute. Thus, we create bags of n= 4 instances, including na= 3 acceleration windows and nl= 1 location windows, as shown in Figure 1. This, combined with MIL, enables us to fuse the two types of signals despite their very different nature. The acceleration-feature encoder, fa, is a convolution model. It consists of an initial batch-normalization layer, followed by 3convolutional blocks, and then 2fully-connected (FC) blocks. Each convolutional block consists of a 2Dconvolutional layer with 3×3kernels, batch normalization, Re.L.U. activation, and finally max-pooling (ratio 2:1per dimension and stride of 2). The 3convolutional blocks have 16,32, and 64 filters respectively. A FC block consists of a dropout layer, a FC layer, batch normalization, and Re.L.U. activation. The first block acts as a bottleneck, by decreasing the dimension to 128. The second block increases the dimension to d. The location-feature encoder, fl, combines a bi-directional LSTM of 128 cells with the set of hand-crafted features (Section III-B) with 3FC blocks. Prior to the Bi-LSTM layer, batch normalization is applied. The output of the Bi-LSTM layer, which includes both the first and last states of the Bi-LSTM, is concatenated with 5hand-crafted features that are computed during the feature extraction phase. Three fully-connected blocks follow; each block consists of a fully-connected layer, batch normalization, and then Re.L.U. activation. D. Multiple instance learning and classification Our main focus is leveraging MIL to effectively integrate information from the two types of modalities. We use bags of ninstances, including naacceleration instances and nl location instances, as shown in Figure 1. However, directly combining acceleration and location instances in their raw form is not a feasible solution due to their inherent differences, such as different representation and captured information. To address this challenge, we map them into a common embedding space, as discussed in Section III-C. The acceleration feature encoder fatransforms the acceleration instances Sa= {sa,1,· · · ,sa,na} ∈ Rna×51×51×2into a set of accelerationbased embeddings Ha={ha,1,· · · ,ha,na} ∈ Rna×d. Similarly, the location feature encoder transforms the location instances Sl={sl,1,· · · ,sl,nl} ∈ Rnl×10×2into a set of location-based embeddings Hl={hl,1,· · · ,hl,nl} ∈ Rnl×d. Subsequently, we introduce the MIL ANN f:RN×d→Rd, which aggregates the multi-modal instance-level embeddings H={ha,1,· · · ,ha,na,hl,1,· · · ,hl,nl} ∈ RN×dinto a fused encoding of the same dimension d= 256.
SIARGKAS ET AL.: TMR BASED ON LOW-RATE ACCELERATION AND LOCATION SIGNALS WITH AN ATTENTION-BASED MIL NETWORK 6 Linear Sigmoid Car Bus Train Bike Run Walk Subway Still Fig. 2. Proposed multi-modal TMR network. The input is a bag containing two types of instances: acceleration (na) and location (nl). Instances are processed in parallel by two modality-specific feature encoders faand flaccordingly, which embed them in the same d-dimensional space. The attention-based MIL ANN f, follows; it aggregates ninstances (of the Rdembedding space) into a fused attentive encoding of the same dimension d. Finally, to make transportation mode predictions, a small classification network maps the fused encoding zinto category-wise decision scores. To emphasize the expressiveness of each instance and prioritize the most useful ones, we propose a weighted average of the instance-level embeddings. Inspired by the work of Ilse [44] as well as of Papadopoulos [52], we design an attentionbased fusion network responsible for determining the instance weights anfor the embedded instances hn. This network can be formalized as follows: an=e(wTtanh (VhT n)⊙sigm (UhT n)) PN j=1 e(wTtanh (VhT j)⊙sigm (UhT j))(6) where w∈R256×1,V∈R256×d, and U∈R256×dare trainable parameters of the attention network, and ⊙is the Hadamard product operator. Accordingly, the contribution of the individual instances is subject to the attention weights {aa,1,· · · , aa,na, al,1,· · · , al,nl}. By definition, it always holds that Pna i=1 aa,i +Pnl i=1 al,i = 1, where the first term indicates the importance attributed to the acceleration modality, while the second term represents the importance of the location modality. Consequently, a modal-fused encoding zis derived by calculating the weighted sum of the instance-level embeddings as: z= N X n=1 an·hn= na X i=1 aa,i ·ha,i + nl X i=1 al,i ·hl,i (7) Finally, the classification ANN, c:Rd→Rm, maps the fused attentive encoding z∈Rdto the label space (where m= 8 in our case); It is a small network, consisting of one fully-connected block (FC layer, batch normalization, and Re.L.U. activation) followed by one FC layer. The problem at hand is formulated as a multi-class classification task, in the sense that there is always a single “right” mode to be recognized. However, given that certain transportation modes, like “train” and “walking”, are not mutually exclusive, the sigmoid activation is employed to provide the final class probability of the j-th TM pj(out of 8TMs). A complete overview of the entire architecture is shown in Figure 2. E. Post-processing While our ANN classifies each example (bag) independently, it is worth noting that there are correlations in transportation mode patterns. For instance, it is unlikely that we would immediately take the bus after riding the subway, without walking in between. To account for such correlations and improve the accuracy of the classification output, a common approach is to apply a smoothing operation, such as a rolling median filter. In our work we employ an HMM that effectively captures temporal dependencies inherent in transportation mode patterns to determine the most probable sequence of classifications and to mitigate classification noise. We apply the HMM in the following manner: •States are defined as the TMs, i.e., a total of 8states are defined •State-transition probabilities P(cm|cm−1)are estimated from the transition matrix derived from the combined training and validation sets •Emission probabilities P(cm|ˆcm)are estimated using the prediction probability estimates p •Start probabilities P(c1)are uniformly set to 1/8(since 8is the number of TM classes) •The HMM is applied to each recording session and the most likely sequence of predictions is computed IV. EXPERIMENTATION SETUP A. Dataset To evaluate our approach we employ the publicly available “SHL Preview Dataset”. This dataset is available as part of the complete Sussex-Huawei Locomotion Dataset [9], [53] and includes 59 hours of annotated recordings; it was captured using 4different smartphones (attached in 4different body positions) on 3participants, as they went about their normal lives over the course of 3days. This dataset contain 8TMs: still, walk, run, bike, car, bus, train, and subway, and preserves all sensor modalities except for audio. In our study, we only consider the accelerometer and location sensor modalities. Acceleration signals are downsampled from the original 100
SIARGKAS ET AL.: TMR BASED ON LOW-RATE ACCELERATION AND LOCATION SIGNALS WITH AN ATTENTION-BASED MIL NETWORK 7 Hz to 10 Hz and location signals are downsampled from 1Hz to 1/60 Hz. The processed acceleration signals are segmented into discrete frames, each with 600 samples (60 sec). At this sample rate, the dataset yields a total of 3,421 acceleration frames, resulting in a total data size of 3,421 ×600 for each sensor placement. The transportation mode labels are synchronized with the acceleration frames, yielding a total of 3,421 labels. Location signals are segmented into 12minute windows, each containing 12 samples with an 11minute overlap with neighboring windows. Location data is recorded asynchronously with average availability of location signal of 80%. In total, 2,715 location samples are obtained against the overall 3,421 acceleration frames and labels. In our scenario, we divide the dataset using LOSO; each user’s data is used in turn for testing and the data from the remaining users are used for model development. The development data (for each LOSO iteration) consist of data from 2participants. We split into 80% training set and 20% validation set using a stratified split framework [54] so that both sets maintain the same label and user distribution. However, random splitting may produce heavily dependent sets since pairs of highly correlated neighboring windows could be assigned to the training and validation set respectively [55]. To avoid this, we select long, continuous streams of data (instead of individual windows) as the unit for random splitting. It is important to note that, since both training and validation data are captured from the same subjects, effectiveness on the validation set may be positively biased compared to the expected effectiveness on the unseen test set. B. Model Training While our proposed ANN (Figure 2) can be trained end-toend, we opt to separately train the acceleration-feature encoder fa. Subsequently, we freeze the trained weights and integrate them in the complete ANN model and trainit, as described in Section V-C2. Both training steps are implemented in the same manner. A standard categorical cross-entropy (CCE) is used as the loss function to compute the error between the prediction probabilities and the true labels. To minimize the CCE loss and update the trainable weights, the Adam optimizer [56] is used with an initial learning rate of 10−4. We use a batch size of 32 and train for at most 80 epochs. To avoid overfitting we use early stopping to identify the ideal number of epochs for training the network. V. EVALUATION AND DISCUSSION We evaluate effectiveness on the basis of accuracy and macro-averaged F1-score. All subsequent experiments are conducted in a user-independent manner (using LOSO). We also repeat each experiment 15 times to account for random initialization of model weights and random train and validation set split. A. Influence of device placement To assess the effectiveness of our model and gain insights into the influence of device placement, we propose a comprehensive assessment from three different perspectives: TABLE I RESULTS OF THE PROPOSED APPROACH ON THE EIGHT-CLASS TASK. RESULTS ARE AVERAGES ACROSS 15 RUNS. No HMM HMM Accuracy F1-score Accuracy F1-score Per-placement Bag 88.0 84.9 91.5 87.3 Hand 82.0 81.3 87.7 85.6 Hips 85.5 80.3 88.7 82.0 Torso 84.9 79.4 88.4 83.3 Average 85.1 81.5 89.1 84.6 All-placements Bag 90.5 89.0 94.6 92.6 Hand 84.8 84.5 90.3 89.5 Hips 89.1 87.8 93.0 91.2 Torso 88.2 85.4 92.5 89.7 Average 88.1 86.7 92.6 90.8 Mixed-placement One 86.6 84.8 91.3 88.6 Multiple 89.3 87.9 93.7 92.3 •Per-placement: we exclusively train and test our model using accelerometer data obtained from a single smartphone position •All-placements: we train our model using data from all four distinct acceleration data streams, each corresponding to a different recording position. Subsequently, we evaluate the model’s effectiveness using accelerometer data obtained exclusively from a single position •Mixed-placement: we generate new virtual streams by combining sequential acceleration data randomly obtained from different positions. This experiment aims to simulate real-world scenarios where the position of the smartphone changes over time. We conduct two experiments of this type; one utilizing a single mixed stream for training and another utilizing multiple mixed streams (four in total) for training. We include both experiments for a more thorough evaluation It is important to note that we do not include any information regarding the placement of the acceleration sensor in any of the experiments mentioned above. Instead, the model remains unbiased and agnostic towards sensor placement, enabling independent learning and distinguishing of the sources of the acceleration data. Furthermore, these experiments are specifically focused on the accelerometer sensors. In contrast, location signals exhibit minimal to no variation across different recording positions. This suggests that the inclusion of location signal in the recognition system enhances its robustness to variations in position. Based on Table I there is a clear ranking of sensor placement: Bag >Hips >Torso >Hand. The data collected from the Hand placement appears to be noisier compared to the other three placements. This is likely due to the physical contact and interaction between the hand and the phone, as it introduces more variability and interference in the captured data. Additionally, the Hand placement seems to be more sensitive to individual user variations and the way each subject
SIARGKAS ET AL.: TMR BASED ON LOW-RATE ACCELERATION AND LOCATION SIGNALS WITH AN ATTENTION-BASED MIL NETWORK 8 uses their smartphone. On the contrary, placing the phone in a Bag yields higher inference accuracy since this placement involves less movement variability and is in close proximity to the body. Hips and Torso exhibit similar behavior. Table I also demonstrates a significant improvement in effectiveness for the generalized classifier when compared to the per-placement experiment. Evidently, in the second experiment, the model is trained across different placements, thus learning more general features, while reducing the risk of overfitting the data from a specific placement. One might expect the effectiveness in the mixed-placement experiment to be on par with the average effectiveness of the all-placements experiment, since both experiments utilize all the available data from every recording placement. However, our model can benefit from the mixed streams and the dynamic smartphone position; the attention-based MIL mechanism can effectively allocate its attention among a sequence of acceleration windows obtained from different positions and leverage the changing position in its advantage (see Section V-D). For instance, if a sequence contains two windows obtained from the Hand position and one obtained from the Bag position, the model will likely prioritize the window acquired from the Bag position. This increase in effectiveness is evident on Table I. B. Acceleration-specific study In all subsequent studies, we maintain the structure of the second experiment (as mentioned in V-A), and the effectiveness is reported as the average score across all testing positions. To gain insights into the acceleration-specific aspect of our proposed model, we examine the impact of three important components: the pre-processing pipeline, using MIL to exploit acceleration data, and the importance of data augmentation. 1) Pre-processing pipeline: Initially, we conduct tests using various sampling frequencies and the findings indicate that the highest accuracy is attained at a sampling rate of 10 Hz. We, then, explore the importance of the acceleration-derived signals, comparing all potential axes a[n] = [ax[n], ay[n], az[n]]T, along with the magnitude a[n] and the jerk norm j[n]. Individually, the magnitude outperforms all other signals, while the jerk exhibits the poorest effectiveness. However, when combined, the magnitude and jerk yield better results compared to using the magnitude alone, and slightly worst to utilizing all the signals together. Next, we compare the different parameters involved in generating the 2D spectrogram representations. We observe that employing the logarithm of the power yields superior results compared to using the raw power. Additionally, a logarithmic interpolation for the frequency axis proves to be an effective approach for reducing data size while retaining essential information. Finally, after experimenting with different short-time Fourier transform (STFT) window durations, we conclude that the highest effectiveness is achieved for a window with a duration of 10 seconds. 2) MIL for acceleration: In this study, we employ MIL to leverage sequences of acceleration windows, instead of relying solely on individual acceleration windows. To assess the importance of this method, we compare three variations of our model: 1234567 Total Duration (minutes) 72 74 76 78 80 82 84 86 Accuracy [%] Method Acc-CNN Acc-MIL Fusion-MIL Fig. 3. Accuracy comparison between Acc-CNN, Acc-MIL and Fusion-MIL versus the duration dof acceleration input data •Acc-CNN: A single d-minute acceleration window is fed into Acc-CNN. This network consists of the accelerationfeature encoder fa(see Section III-C) followed by a classification layer (FC layer and sigmoid activation) that maps the final embedding hainto class probabilities p. •Acc-MIL: The same d-minute window is now segmented into a sequence of dsuccessive one-minute windows, which are fed to Acc-MIL. Similar to Acc-CNN, AccMIL also includes the acceleration-feature encoder fabut additionally incorporates the MIL ANN f, followed by the same classification layer. •Fusion-MIL: A sequence of done-minute acceleration windows along with a single 12-minute location window, as depicted in Figure 1. The proposed Fusion-MIL model exploits this bag of multi-modal instances to perform TMR, as illustrated in Figure 2 In Figure 3, we present a effectiveness comparison between these three approaches as the duration dof the input increases. Evidently, removing MIL causes a clear effectiveness drop. 3) Data augmentation: As we progressively increase the number of acceleration instances, we inadvertently increase the overlap across different examples, increasing the risk of over-fitting. To mitigate this issue and improve generalization, we opt to augment the acceleration data before feeding it into the model (see section III-A). Figure V-B3 investigates the influence of data augmentation on the effectiveness of AccMIL as the number of acceleration instances increases. C. Multi-modal study We assess the effect of three significant designs, i.e. alternative approaches to Fusion-MIL, pre-training, and Markov smoothing. 1) Comparison with alternative approaches: We compare the classification effectiveness of our proposed Fusion-MIL model to the following single-modal and multi-modal approaches (results in Table II): •Loc-LSTM: A location-only ANN with a 12-minute location window as its input. This network consists of the location-feature encoder flfollowed by a classification layer (FC layer and sigmoid activation) that maps the final embedding hlinto class probabilities p.
SIARGKAS ET AL.: TMR BASED ON LOW-RATE ACCELERATION AND LOCATION SIGNALS WITH AN ATTENTION-BASED MIL NETWORK 9 1234567 Number of instances 72 74 76 78 80 82 84 86 Accuracy [%] Augmentation Not-Applied Applied Fig. 4. Effect of data augmentation on the accuracy of Acc-MIL versus the number dof one-minute acceleration instances TABLE II EFFECTIVENESS COMPARISON WITH ALTERNATIVE SINGLE-MODAL (SM) AND MULTI-MODAL (MM) METHODS Method Accuracy F1-score without MIL Loc-LSTM (SM)72.3 61.4 Acc-CNN (SM)81.2 80.7 Fusion-Concat (MM)84.5 81.3 with MIL Acc-MIL (SM)82.5 81.9 Fusion-Concat++ (MM)85.2 82.5 Fusion-MIL (MM)86.1 84.1 •Acc-CNN: A acceleration-only ANN with a 3-minute acceleration window as its input (see Section V-B2). •Fusion-Concat: A multi-modal network that accepts a 3minute acceleration window and a 12-minute location window as a paired input. It combines the accelerationfeature encoder faand the location-feature encoder fl, followed by a fusion layer that concatenates the 256dimensional modal embeddings haand hl. The modalfused 512-dimensional encoding z, is fed to the classification ANN cto derive the final classification probabilities. •Acc-MIL: A acceleration-only ANN that uses MIL to leverage a bag of na= 3 successive one-minute acceleration windows for classification (see V-B2). •Fusion-Concat++: We enhance Fusion-Concat, by substituting its original fasub-network with Acc-MIL. The Acc-MIL branch exploits a set of na= 3 successive oneminute acceleration windows, mapping them into a joint 256-dimensional embedding ha, while the flbranch maps a12-minute location window into a 256-dimensional embedding hl. The 2embeddings are fused through the concatenation layer and the resulting modal-fused encoding zis mapped to the label space. 2) Pre-training: Training the proposed ANN end-to-end reduces effectiveness, primarily due to the limited availability of paired acceleration-location data in comparison to the available acceleration-only data (the same location recording corresponds to 4different accelerometer recordings from 4 TABLE III EFFECT OF PRE-TRAINING Encoder pre-training Accuracy F1-score No pre-training 86.1 84.1 Location encoder fl87.0 84.8 Acceleration encoder fa88.1 86.7 Both encoders fa&fl88.8 87.4 TABLE IV OVERALL EFFECTIVENESS OF OUR PROPOSED METHOD. AVERAGES AND STANDARD DEVIATIONS ARE COMPUTED OVER THE 4PLACEMENTS AND FOR EACH PLACEMENT WE PERFORM 15 RUNS WITH RANDOM INITIAL MODEL WEIGHTS AND TRAIN/VALIDATION SPLIT. Accuracy F1-score Precision Recall no HMM 88.1±0.7 86.7±1.0 87.8±0.8 86.6±1.2 HMM 92.6±1.0 90.8±1.5 92.8±1.0 90.6±1.6 different smartphone positions). Thus, we opt to separately pre-train the uni-modal feature encoders and then use the trained weights in the full ANN model: •Acceleration encoder pre-training: To train fawe adopt the Acc-CNN framework (Section V-B2) with 1-minute acceleration windows. After training fawith all available training acceleration data we discard the classification layer and train the multi-modal ANN. We randomly use only one acceleration-location pair (out of 4) per instance. To prevent over-fitting, we freeze the trained weights of faand solely train the remaining layers. This approach uses all available data and reduces the risk of overfitting. •Location encoder pre-training: We pre-train fafollowing a similar process; we include this baseline for a more comprehensive evaluation. •Both encoders pre-training: We pre-train both acceleration and location feature encoders. Table III provides a comparison of incorporating pre-trained uni-modal encoders. The results clearly indicate that the acceleration pre-trained model surpasses the no pre-training model. By pre-training both encoders there is only a marginal effectiveness improvement compared to solely pre-training the acceleration encoder. 3) Smoothing with HMMs: To determine the most probable sequence of TMs we employ Markov smoothing, which reduces classification noise by exploiting the temporal correlation between neighbouring samples. The HMM uses the prediction probability estimates (p) as emission probabilities P(cm|ˆcm)while the transition probabilities P(cm|cm−1)are estimated using the combined training and validation sets. Table IV presents the effect of applying HMM smoothing and Figure 5 shows the ROC curve without the HMM. It can be observed that the area under the curve is highest for nonmotorized TMs (still, walk, run and bike) and slightly lower for motorized modes of transportation. D. Model interpretability The significance of individual instances is determined by their respective attention weights, i.e.: