Full text
Received 26 March 2025, accepted 13 April 2025. Date of publication 00 xxxx 0000, date of current version 00 xxxx 0000. Digital Object Identifier 10.1109/ACCESS.2025.3561238 MobilitApp: A Deep Learning-Based Tool for Transport Mode Detection to Support Sustainable Urban Mobility GERARD CARAVACA IBAÑEZ, LUIS J. DE LA CRUZ LLOPIS , (Member, IEEE), ADRIAN CATALÍN DIACONEASA, ALBERTO BAZÁN GUILLÉN , AND MÓNICA AGUILAR IGARTUA , (Member, IEEE) Department of Network Engineering, Universitat Politècnica de Catalunya (UPC), 08034 Barcelona, Spain Corresponding author: Mónica Aguilar Igartua ([email protected]) This work was supported in part by Spanish Government funded by MCIN/AEI/10.13039/501100011033 under Research Projects ’’DIstributed Smart Communications with Verifiable EneRgy-optimal Yields (DISCOVERY)’’ under Grant PID2023-148716OB-C32 and ’’Enhancing Communication Protocols with Machine Learning while Protecting Sensitive Data (COMPROMISE)’’ under Grant PID2020-113795RB-C31; partially funded by MCIN/AEI/10.13039/501100011033 and by the European Union (EU) NextGenerationEU/PRTR (Plan de Recuperación, Transformación y Resiliencia) under Research Project ’’Anonymization Technology for AI-Based Analytics of Mobility Data (MOBILYTICS)’’ under Grant TED2021-129782B-I00; in part by the Predoctoral Scholarship ‘‘Generación de Conocimiento - Projects Call 2022’’ under Grant PRE2021-099830; and in part by the Generalitat de Catalunya under AGAUR (Agència de Gestió d’Ajuts Universitaris i de Recerca) Grant 2021-SGR-01413. ABSTRACT The visible effects of climate change in urban areas are driving a paradigm shift in societal and political priorities. Urban planners, public transport providers, and traffic managers are increasingly focused on redesigning cities to promote sustainable mobility and create green spaces for pedestrians, cyclists, and scooter users. In alignment with these objectives, the European Climate Law mandates a minimum 55% reduction in greenhouse gas emissions by 2030 and climate neutrality by 2050. Achieving these targets requires robust tools to collect and analyze mobility data, enabling the evaluation of citizens’ travel habits and the planning of sustainable urban infrastructure. This study presents MobilitApp, a tool developed based on a deep learning (DL) model for real-time detection of transportation modes using smartphone sensor data. Our approach leverages a hierarchical model combining convolutional neural networks (CNNs) for feature extraction and long short-term memory (LSTM) layers for temporal processing, enhanced by skip connections. To ensure computational efficiency on mobile devices, the system integrates statistical techniques for early motion detection, minimizing reliance on DL models. The model was trained on a dataset of multimodal trips in Barcelona, achieving over 80% accuracy for most transport modes and a weighted average accuracy of 88%. These results highlight the effectiveness of our approach for accurately predicting users’ transport modes during their trips. The MobilitApp tool provides an intuitive platform for collecting and analyzing urban mobility data. By analyzing travel patterns, transport modes, and modeswitching behaviors, it delivers actionable insights to city planners, aiding in the enhancement of urban mobility, promotion of sustainable development, and transition to greener cities. INDEX TERMS Transport mode detection, activity recognition, deep learning, mobility sensors, sustainable urban mobility, MobilitApp. I. INTRODUCTION The visible impact of climate change in urban areas is driving a significant transformation in societal and political priorities. The associate editor coordinating the review of this manuscript and approving it for publication was Mouquan Shen . This shift has prompted urban planners, public transport providers, and traffic managers to urgently redesign cities with a focus on sustainable mobility and the creation of green spaces dedicated to pedestrians, cyclists, and scooter users. In response to the climate crisis, the European Parliament enacted the European Climate Law, which mandates a VOLUME 13, 2025 2025 The Authors. This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License. For more information, see https://creativecommons.org/licenses/by-nc-nd/4.0/ 1
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection minimum 55% reduction in greenhouse gas emissions by 2030 and establishes climate neutrality as a legally binding objective by 2050. Private vehicles, which contribute 15% of the EU’s CO2emissions, remain a critical target for these efforts. To achieve these goals, urban planners require robust tools to collect and analyze real-world mobility data. Such tools enable the evaluation of citizens’ travel habits and support the planning of essential infrastructure changes, such as identifying optimal locations for mobility hubs, converting streets into pedestrian zones, and enhancing public transport systems to foster more sustainable urban environments. To support urban mobility planners in achieving their goals, this paper provides an in-depth exploration of smartphone-based systems for transportation mode detection. The widespread adoption of smartphones has revolutionized data collection, enabling dynamic and real-time monitoring of transportation patterns. These devices, endowed with diverse sensors and computing capabilities, enable continuous data collection on individual movements and transportation choices. To support this goal, we developed MobilitApp [4], a tool designed to analyze citizens’ mobility patterns, including travel origins and destinations, modes of transportation used, and transfer points between different modes. Our tool accurately identifies 12 distinct transport modes and human activities with an impressive accuracy of approximately 88%. Designed for seamless integration, the code can function as a library within applications managed by public transport entities, fully compliant with stringent privacy regulations. This functionality empowers urban mobility planners with precise insights into citizens’ mobility patterns, supporting informed decisions to enhance urban mobility and foster sustainable transportation solutions. Our goal is to support municipal entities, such as the Autoritat del Transport Metropolità (ATM) of Barcelona [1], in enhancing urban mobility and advancing towards more sustainable cities. This study aims to fully leverage smartphone technology for transport mode detection, automatically identifying a person’s mode of transport. We address key aspects such as algorithm development, data analysis, preprocessing techniques, privacy considerations, and practical implementation. By transforming the understanding and management of transportation, we aim to contribute to the creation of more sustainable, healthier, and user-friendly urban environments, with a specific focus on the city of Barcelona as a case study. To achieve the desired outcome, we have engineered a sophisticated hierarchical deep learning model, i.e. a machine learning model developed through the systematic construction of deep neural networks organized in a hierarchical structure. This framework marries the strengths of convolutional neural networks (CNNs) [32] for high-level feature extraction, with the capabilities of long short-term memory (LSTM) [24] networks for processing sequential data. To enhance its efficiency and robustness, this integrated approach is further augmented with frequency analysis techniques and GPS location methods. This combination guarantees a comprehensive and efficient system for addressing the task at hand. The main key contributions of this work are listed as follows: •In this work, we introduce MobilitApp, a tool capable of predicting citizens’ activities and transportation modes during their trips across a metropolitan area. MobilitApp features an advanced hierarchical deep learning model designed with a strong emphasis on minimizing computational costs, which is particularly important as the tool is intended to run on smartphones with limited processing power. This predictive framework provides urban mobility planners and public transport providers with precise insights into citizens’ mobility patterns, facilitating informed decision-making to enhance urban mobility and promote sustainable transportation solutions. •We have developed a hierarchical deep learning model to predict citizens’ activities and transportation modes, supporting urban planners in fostering sustainable urban mobility. To train the model, we created a dataset using three smartphone sensors (accelerometer, magnetometer, and gyroscope), achieving an average accuracy of 88%. This model is used to characterize citizens’ multimodal trips, offering valuable insights into urban mobility patterns and aiding efforts to promote more sustainable transportation practices. •We have developed an Android application, named MobilitApp [4], able to perform these tasks: (i) realtime data acquisition of smartphone mobility sensors; (ii) data pre-processing; (iii) transport mode prediction using our hierarchical deep learning model; and (iv) user feedback showing the predicted mode over a map. The system is structured hierarchically consisting of: (i) a kinematic motion classifier to fast detect Walk and Stationary states that define consecutive segments of the multimodal trip (commuting points); (ii) the proposed hierarchical deep learning model; and (iii) a stop-detection algorithm that automatically identifies the end of a trip and halts sensor data collection on the smartphone, ensuring that no additional data is gathered once the trip has concluded. •We have tested the MobilitApp tool in a large-scale multimodal travel data collection campaign within the Universitat Politècnica de Catalunya (UPC) community in Barcelona. This campaign produced an extensive dataset of multimodal trips, consisting of consecutive segments predicted by the MobilitApp tool. The collected data shows promising potential for future analysis of the mobility habits of a community (e.g., a Campus, a city, a company). The rest of this work is structured as follows. Section II provides background and reviews related studies, establishing the context for our research. Section III-D outlines the methods used for data pre-processing and the dataset collection process. Then, Section IV discusses the 2VOLUME 13, 2025
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection model architecture and its performance outcomes. Section V presents a comprehensive overview of the Android application MobilitApp [4], which encapsulates our complete system. Afterwards, Section VI includes a performance evaluation of the MobilitApp proposal, showcasing results from a pilot campaign on multimodal trip collection. Finally, Section VII offers concluding remarks and suggests potential future research directions. II. RELATED WORK This study is inspired by a diverse range of methodologies, encompassing both classical and modern approaches. It integrates traditional machine learning algorithms and sensor frequency analysis with insights from recent advancements in deep learning models. A. TRADITIONAL PROPOSALS Initially, the task of transport mode classification focused on manual feature extraction and training of machine learning models using traditional algorithms such as support vector machine [22], k-nearest neighbour [14] or decision trees [35]. An example of that is the Hemminki et Al. [23] proposal. This proposal is based on a hierarchical feature-based system composed of three stages. The hierarchy begins with a proposal that initially distinguishes between pedestrian motion and other types of motion at a coarse level. If the kinematic motion classifier fails to detect significant physical movement, such as walking, the process advances to a stationary classifier. This classifier then assesses whether the user is either stationary or on some form of motorized transportation. When motorized transportation is identified, the classification process moves on to a motorized transportation classifier. This classifier is responsible for categorizing the ongoing activity into one of five modalities: bus, train, metro, tram, or car, and it employs adaptive boosting (AdaBoost) as a statistical classification meta-algorithm for enhanced accuracy. There are other studies that, in addition to using sensor data, use GPS data to improve classification. An example of this is the approach proposed by L. Randleff et al. [33]. The plan was to primarily rely on accelerometer data unless it was not clear what mode of transportation was being used. This way, they save energy and keep the algorithm simple by using more power-hungry sensors and data processing only when necessary. It is worth noting that the algorithm was tested on a relatively small dataset, requiring users to label their trips with both device orientation and transportation mode during the training phase. Nonetheless, the approach of leveraging additional sensors to complement accelerometer data when necessary is interesting. B. DEEP LEARNING METHODS In the context of machine learning and artificial intelligence, the importance of feature selection has been a longstanding challenge. Traditionally, the process of identifying and selecting relevant features from raw data has been a crucial step in building effective predictive models. However, modern deep learning techniques have transformed this field by eliminating the need for manual feature selection, enabling the direct use of raw data for tasks such as activity recognition and transportation mode prediction. One of the pioneering techniques in this regard is the use of CNNs [18]. CNNs have shown remarkable capabilities in understanding spatial relationships within data, making them highly effective in tasks like image recognition. When applied to activity recognition, CNNs can directly process raw sensor data, allowing for the automatic extraction of relevant features from the input. On the other hand, inspired by the transformative success of deep learning in activity recognition and transportation mode prediction, J.V. Jeyakumar et al. proposed a novel deep learning model known as the Deep Convolutional Bidirectional LSTM (DCBL) [26]. DCBL takes raw sensor data as input, combining the power of convolutional layers to capture spatial patterns and bidirectional LSTMs to model temporal dependencies. This approach enables precise predictions of transportation modes without the complexity of feature engineering or selection. C. HUMAN ACTIVITY RECOGNITION Regarding proposal specifically oriented to Human Activity Recognition (HAR), the work in [40] proposes a radar-based HAR approach using hyper-dimensional computing (HDC), specifically designed to identify six human activities: walking, walking while carrying a carton, walking with a stick, walking in a squat position, running, and jumping. The HDC system proves to be energy-efficient, robust, lightweight, and capable of fast learning with low latency, making it ideal for real-time, on-chip human activity recognition in future radar systems. The work [17] develops OptiMapper, a uninformed cross-subject transfer learning framework for activity recognition that extracts abstract knowledge across subjects and utilizes this knowledge for developing a personalized and accurate activity recognition model in new subjects. They focus on daily and sport human activity recognition (jumping, descending stairs, walking, running, cycling). The work [9] proposes an heuristic algorithm for a decision tree to automate a decision-making for HAR to accurately detect passengers getting into public transportation systems. They use a public WiFi-based activity recognition dataset to extract human activity features from Channel State Information power, which changes caused by nearby human activity. D. TRANSPORT MODE DETECTION The work [25] proposes a combined solution of a Long-ShortTerm Memory network and a Healing algorithm (used to correct misdetected modes based on majority voting) capable of recognizing 12 transportation modes (walking, running, climbing stairs, descending stairs, bicycle, motorcycle, car, subway, train, high-speed train, tram and metrobus). The VOLUME 13, 2025 3
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection authors in [28] propose a model able to recognize 6 transportation modes: car, bus, train, bicycle, walking and vertical activities (elevator, escalator, and stairs) from a large dataset with over 16,000 participants going about their daily activity in the city-state of Singapore. The dataset includes barometer, accelerometer, and Wi-Fi scanner data. They focus their work on identifying vertical mobility at city-scale and its potential to track and identify vertical transportation in a densely built urban environment. The work [34] introduces a transportation mode recognition through fusing multimodal data from wearable sensors: motion, sound and vision. Three independent deep neural network (DNN) classifiers work with the three types of sensors, respectively. Then two schemes fuse the classification results from the three mono-modal classifiers. Results show that while performance is reduced for each individual classifier, the benefits of fusion are retained with performance improved by 15%. While motion and sound sensors are available in current smartphones, vision should be obtained from body-worn cameras such as eye-wear computers (e.g. Google Glass, Spectacles by Snap). Inspired by these and other approaches, in this work we have designed a novel hierarchical deep learning model to recognize transportation modes and related human activities (still and walk). Our model uses convolutional neural networks (CNN) for high-level feature extraction and LSTM networks for processing sequential data, to develop a predictive model for eleven types of transport and human activities: Bike, bus, car, motorbike, run, stationary, subway, train, tram, walk, e-scooter. Additionally, we employ preprocessing techniques to standardize, augment, and refine the dataset, enhancing the training and evaluation of machine learning models. Building on this, we developed three algorithms capable of characterizing complete multimodal trips (comprising multiple segments) each associated with a specific transport mode, in real time: •Micro-state algorithm. It classifies motion within each 20-second sample set, as detailed in Algorithm 2. •Macro-state algorithm. It classifies motion across every five micro-states (equivalent to 100 sec.) by employing a majority-value scheme, as described in Algorithm 1. •Stop algorithm. It autonomously detects when a user concludes their trip, as it is outlined in Algorithm 3. III. MOBILITY DATASET OF MOBILITY SENSORS This section outlines the methodology used for the creation and refinement of the dataset central to this study. It begins with a detailed description of the sensors employed, followed by an explanation of the data collection process. Subsequently, the pre-processing techniques applied to the raw data are detailed. To further enhance the dataset, this section also describes several data augmentation techniques. Finally, the section addresses feature extraction, explaining how the preprocessed data was transformed into a structured feature set for use in the outlier detection process. A. DISCUSSION ABOUT AVAILABLE DATASETS REGARDING MOBILITY Before creating our own mobility dataset, we first evaluated the suitability of existing public datasets such as Sussex-Huawei Locomotion Dataset (SHL) [19],[38], Transport Mode Detection Dataset (TMD) [11] and Collecty Dataset [16]. However, we determined that none of the available datasets satisfied the specific requirements of this project. The SHL dataset [19] is a rich resource for developing activity recognition technologies, offering a wide array of sensor data (2812 h during 7 months) that captures a detailed view of user environments and actions (8 transport modes). It includes data collected from 15 smartphones, varied sensor inputs, integrates third-party data for benchmarking, and provides crucial information on device status and precise positioning. However, its complexity may deter simpler applications, and redundancies, hardware specificity, and limited participant diversity pose challenges to its broader applicability and the generalization of derived models. The TMD dataset [11] is a robust resource for activity recognition research, marked by its diverse participant base (13 users) and rigorous pre-processing, offering a solid foundation for developing and testing models across various smartphone models. However, its limitations, including limited data duration (31 h), and a small range of activity classes (5 transport modes), may constrain its effectiveness in capturing the full complexity and diversity of human moving activities. Finally, the Collecty dataset [16] presents a comprehensive overview of transportation habits, encompassing a wide array of transport modes (8 transport modes, 15 users), high amount of data gathered (242 h during 5 month) and providing detailed usage duration while ensuring participant anonymity. However, its geographical focus on Zagreb, Croatia, the inaccessibility of the data collection app (to gather data from other cities), and the lack of device-specific information may limit its broader applicability and introduce challenges in ensuring replication and sensor data consistency in research. In conclusion, there is a notable absence of a comprehensive public dataset suitable for effectively training machine learning models to predict transportation modes in urban environments. An ideal dataset would encompass data from a diverse range of urban residents and include a wide variety of transportation modes. B. SMARTPHONE SENSORS USED IN MOBILITAPP In our dataset, data was collected from three motion sensors embedded in the smartphone: accelerometer,gyroscope, and magnetometer. These three sensors have the advantage of being present in nearly all smartphones on the market, providing a comprehensive view of the user’s movement. For the gyroscope and magnetometer, data was collected directly from the raw sensor outputs. In the case of the gyroscope and the magnetometer, data has been collected directly from 4VOLUME 13, 2025
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection the pure sensors. In the case of the accelerometer, however, it was decided to use the linear acceleration value, since it provides information about acceleration along the x,y, and zaxes, excluding the effects of gravity, to avoid the masking effects of gravity on the precise readings of the accelerometer sensor. For the purpose of data collection and subsequent dataset compilation, an Android smartphone application was developed. Called MobilitApp [4], this application is able to collect data from the phone sensors during users’ trips. It has two main functionalities: (i) the first one is to collect sensor data from trips labeled by our volunteers; and (ii) the second one is to test the live transport detection model during multimodal travel. A visual explanation of this comprehensive process of data acquisition is provided in Fig. 1. In addition, a detailed user manual and video tutorial are available on the MobilitApp website [4]. After validating the functionality and accuracy of the MobilitApp system, we carried out multiple data collection campaigns to obtain sufficient sensor samples for all transportation modes included in the MobilitApp dataset. These data samples were essential for the effective training of our prediction model. Consequently, the resulting dataset incorporates samples from 50 smartphones, exceeding the user counts of previously analyzed databases. This larger sample size is expected to improve the model’s ability to generalize effectively to new users. C. PRE-PROCESSING STEPS The efficient processing and analysis of sensor data collected during various journeys present a significant challenge, particularly due to the inherent variability in journey duration and conditions. To address this challenge, the subsequent subsections examine a series of advanced techniques designed to standardize,augment, and refine the dataset, enabling more effective training and evaluation of machine learning models. 1) SLIDING WINDOW The data collected from smartphone sensors during a journey are stored in CSV-format files. However, these datasets often display variability in journey durations. To standardize and preprocess this data effectively, a technique known as the sliding window method is applied. Adjusting the window size and overlap factor is not trivial; these parameters significantly influence the granularity and quality of the analyzed data. Smaller windows might capture finer details but risk missing broader trends, while larger windows might smooth over important nuances. Similarly, higher overlap might provide more continuous data analysis but increase redundancy and computational load. Conversely, lower overlap might miss critical transitional information. For our use case, after considerable testing, we determined that the optimal values for window size and overlapping factor were 512 and 50%, respectively. 2) DATA AUGMENTATION Data augmentation is a critical process in the development of robust machine learning models, particularly when working with datasets that need to represent a complex and variable real-world phenomenon. The objective of data augmentation is to artificially increase the size and enhance the quality of the dataset by generating new synthetic examples derived from the original data. These modifications enable the model to learn from a more diverse set of examples, which is essential for enhancing its generalization capabilities. For this purpose, in this work two data augmentation functions have been designed: obfuscation and rotation. On the one hand, obfuscation techniques try to avoid recognition of phone-specific patterns. Essentially, the process involves adding noise to the phone’s sensor data. Both additive and multiplicative noise are introduced to enhance the dataset. Additive noise serves to impose a consistent noise floor across the signal, while multiplicative noise modifies the signal in a scale-dependent fashion. The application of white noise is particularly advantageous because of its consistent frequency spectrum, which effectively masks device-specific noise signatures while preserving the signal’s spectral characteristics. On the other hand, rotation techniques try to address the variability of phone orientations, which can significantly affect sensor readings. Thus, the dataset is augmented by simulating various potential phone orientations that may occur in real-world scenarios during a user’s journey. This is achieved by performing rotations of the three-dimensional sensor data along the X, Y, and Z axes. For this, random rotation angles from 0◦to 180◦are generated. Subsequently, a corresponding three-dimensional rotation matrix is constructed based on these angles. The sensor data for each trip is then transformed by this rotation matrix, effectively reorienting the data to reflect a different phone position. These approaches ensure that the model remains focused on identifying different transportation modes, avoiding distractions caused by minor, irrelevant details unrelated to the primary task. 3) FEATURE EXTRACTION FOR OUTLIER DETECTION Feature extraction for outlier detection involves identifying and selecting the most relevant features from a dataset to detect outliers that deviate significantly from the norm and can represent anomalies, errors, fraud or rare events. Feature extraction helps algorithms to identify more effectively these anomalies, improving analysis accuracy and efficiency. Traditional methods like Z-score are often constrained by assumptions of linearity and normal distribution. In contrast, the Mahalanobis distance (MD) [29] offers a more nuanced approach, especially in multivariate contexts. Unlike Euclidean distance, MD accounts for correlations between variables, making it particularly effective for identifying outliers in multidimensional datasets. This characteristic is especially advantageous for this dataset due to the inclusion of data from multiple sensor types. VOLUME 13, 2025 5
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection FIGURE 1. Application interface of the MobilitApp tool for collecting sensor data. First, the user initiates data collection from the mobility sensors, ensuring that each segment is labeled with the corresponding transport mode. At the end of the trip, the user stops the data collection process and has the option to send the data to our server. [4]. TABLE 1. Summary of the features extracted from the MobilitApp sensor dataset. To conduct this study, we utilized features extracted from the raw data, which were further processed using the principal component analysis (PCA) technique. With the reduced feature set, we calculated the Mahalanobis distance for each observation. The final step involved identifying outliers based on the calculated MDs. We used the Chi-Square Distribution to determine a threshold for what constitutes an outlier. A 95% confidence interval was chosen, meaning that observations with a MD greater than the corresponding Chi-Square value were considered outliers. This statistical approach provides an objective criterion for outlier detection. For the feature extraction process, we chose to extract three types of features: statistical features, frequency domain features, and time domain features. Table 1shows a description of the features extracted from the MobilitApp sensor dataset. D. RESULTING MOBILITAPP SENSOR DATASET After completing various data collection campaigns, we have generated the MobilitApp Sensor Dataset, which is publicly available [8]. As mentioned above, it consists of the X, Y and Z axes of three motion sensors: Accelerometer, Gyroscope and Magnetometer. The file is 10,500 lines long and 2.9 GB in size. Note that the data gathered in this dataset were collected during 15 months in different campaigns, from March 2023 to June 2024 from a total of 50 different smartphones, representing 298 hours of data collecting time. Volunteers labeled their unimodal trips within the Barcelona Metropolitan Area using nine different transportation modes and three human activities related to these modes: Bike, Bus, Car, Motorbike, Run, Stationary, Subway, Tram, Train, Walk, e-Bicycle, e-Scooter. Volunteers collected activity samples while carrying the smartphone in various positions, including in their hand, pocket, and bag. The distribution of hours by transport mode, without including yet any data augmentation, is depicted in Fig. 2. The dataset presents a significant challenge: achieving a more balanced distribution of hours across all transportation modes, as the data is unevenly distributed. This imbalance is primarily due to volunteers’ tendency to use the most common modes of transportation (bus, subway, car, walking). This issue is commonly observed in other similar datasets that were analyzed [11],[16],[19]. It should be noted that, in this study, the e-bicycle mode of transport will not be considered for training the prediction model, as we have not yet collected sufficient samples. However, we plan to repeat the training in the near future to include this sustainable mode of personal transport. Consequently, Fig. 3illustrates the sensor dataset after applying classical data augmentation techniques to artificially expand its size and enhance its 6VOLUME 13, 2025
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection TABLE 2. Comparison of key features of MobilitApp sensor dataset compared to similar datasets. FIGURE 2. MobilitApp sensor dataset (July 2024). Class distribution of the dataset collected using MobilitApp [4], after applying pre-processing steps, excluding data augmentation techniques. FIGURE 3. MobilitApp Sensor Dataset (July 2024), after applying data augmentation techniques, see Sec. III-C2. quality by generating new synthetic examples derived from the original data, see Sec. III-C2. Compared to the other referenced datasets summarized in Section III-A, our dataset presents a higher volume of hours recorded than TDM and similar than Collecty. However, it is notably smaller when contrasted with the SHL dataset. Nevertheless, it is pertinent to note that our dataset includes a broader range of transport modes than its counterparts and a higher variety of users. Table 2summarizes the key features of the datasets analyzed, as well as those of our own dataset, MobilitApp. IV. DESIGN OF A HIERARCHICAL PREDICTION MODEL FOR ACTIVITY AND TRANSPORTATION MODE RECOGNITION To develop the hierarchical prediction model, we focused on minimizing computational costs, as MobilitApp is designed to run on smartphones with limited processing power. As a result, the model was designed to be lightweight and easily optimized, ensuring efficient operation on mobile devices. We experimented with different architectures, from fully connected networks to recurrent architectures including convolutional networks. After all the experimentation, the architecture that best fits our use case is the one shown in Fig. 4. Our proposal takes inspiration from the efficacy of the mixed model proposed in [37], and additionally it also includes a hierarchical design. Regarding the hyperparameter optimization process, we conducted extensive experiments to fine-tune the functions and parameters underlying the hierarchical deep learning model. Below, we summarize the final design choices and key features of the model, which are also depicted in Table 3. The model is composed of an initial trunk of convolutional blocks. This is the most important part of the model for feature extraction. The achievement of these blocks should allow the model to extract the features from the data, which allows the rest of the prediction model to differentiate between the various modes of transport. In this block various parameters and configurations have been experimented. As a result, the best configuration extracted from this part of the study is the one using dilated convolutions with kernel sizes of 7, 5 and 3; and with the initial block’s structure composed of convolutional layers using ReLU activation (a nonlinear function that outputs the input directly if it is positive and zero otherwise, enabling efficient training by introducing sparsity) [39], with batch normalization (to normalize the inputs or hidden states within the network across a batch, stabilizing and accelerating training), spatial dropout (a regularization technique that randomly drops entire feature maps during training to reduce overfitting) of 5% [36], and max pooling (a downsampling technique commonly used in neural networks to reduce the spatial dimensions of an input volume by capturing the most prominent features while discarding less relevant information). Retaining the convolutional framework described above, this model incorporates skip connections that route the feature maps directly into dedicated LSTM layers with 128 neurons. Concretely, the outputs from the first, second, third, and fourth max-pooling stages are each fed into separate LSTM units. Such a design enables individual LSTMs to capture and analyze temporal patterns at various levels of feature abstraction, enriching the model’s capacity to discern complex temporal relationships within the data. Finally, there is a dense layer (also known as a fully connected layer) applying a L2 regularizer [13] with parameter 0.001, which processes the features learned by the LSTM; and also a dense output layer with a softmax activation VOLUME 13, 2025 7
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection FIGURE 4. Architecture of the proposed hierarchical prediction model designed in MobilitApp. function which outputs the matching probabilities for each transport mode. A. DESIGN OF HIERARCHICAL MODEL PARAMETERS AND PERFORMANCE EVALUATION This section presents a performance evaluation of our proposed hierarchical deep learning model for predicting activity and transportation mode, trained using the configuration summarized in Table 3. Notice that our hierarchical deep learning model has been explicitly designed to account for the class imbalance present in our dataset. For instance, to quantify the difference between the model’s predictions and the actual labels and optimize the model during training, we employed a Weighted Categorical Crossentropy with 0.1 smoothing as the loss function (see Table 3). This approach includes the following elements: •Categorical Crossentropy (CCE): A loss function commonly used in multi-class classification tasks, which measures the divergence between the predicted probability distribution and the true one-hot encoded labels. In a one-hot encoding scheme, the correct class is represented by a probability of 1, while all other classes are assigned a probability of 0. •Weighted CCE: Assigns different weights to each class, penalizing misclassification of underrepresented classes more heavily, thereby mitigating the effects of class imbalance. •Label smoothing (0.1 smoothing): Replaces strict one-hot labels with softer probability distributions, reducing overconfidence in predictions and improving generalization. For example, in a four-class scenario, a traditional one-hot encoding of [0, 1, 0, 0] would be adjusted to [0.025, 0.9, 0.025, 0.025] with a smoothing factor of α= 0.1. This prevents the model from becoming overly confident in its outputs and enhances generalization, particularly in imbalanced datasets. To further mitigate class imbalance, we also applied undersampling to majority classes such as car and public transport, thereby reducing their dominance and improving overall dataset balance. Moreover, additional data rebalancing techniques could be explored to enhance representation in the bicycle, e-bicycle, e-scooter, and motorcycle classes. Two promising methods for our dataset include: •(Synthetic Minority Over-sampling Technique) [12]: Generates synthetic samples by interpolating existing data points from minority classes, increasing their representation. •ADASYN (Adaptive Synthetic Sampling) [21]: Expands on SMOTE by adapting sample generation based on the local distribution of minority classes, focusing on areas with higher classification difficulty. We plan to investigate these rebalancing strategies as part of our future work in the coming months. We partitioned the sensor dataset into three subsets: (i) Training set used for the training of the model (80% of the data); (ii) Validation set used to give an unbiased evaluation of the model trained by tuning the hyperparameters of the model (10% of the data); and (iii) Test set used as an unbiased evaluation of the final model (10% of the data). We have ensured that the validation and test sets consist of data from users not included in the training set. This strategy aims to better emulate real-world scenarios, where the model must generalize to unseen users, thereby providing a more robust evaluation of its performance. The total training time for the model was 2 hours on the i9-13900F computer specified in Table 3. We can see the performance evaluation of the model both in the confusion matrix in Fig. 5and in the classification report in Table 4. A confusion matrix represents the real class (true label) in rows and the predicted class (predicted label) in columns. In the confusion matrix (see Fig. 5), we provide an example for the case of the e-Scooter: (i) True positive (TP) is the number of positives correctly identified (yellow square zone); (ii) true negative (TN) is the number of negatives correctly identified (brown square zone); (iii) false positive (FP) is the number of negatives incorrectly identified as positive (red rectangle zone); and (iv) false negative (FN) 8VOLUME 13, 2025
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection TABLE 3. Hyperparameter configuration of the hierarchical model. is the number of positives incorrectly identified as negatives (blue rectangle zone). Table 4shows the most common metrics for evaluating the performance of prediction models: (i) Precision, as the percentage that the model correctly predicts positive when making a decision, see Eq. (1); (ii) Recall, also known as sensitivity, as the percentage of positives correctly identified out of all the existing positives, see Eq. (1); and (iii) F1score, the harmonic mean of Precision and Recall as more suitable metric for evaluating performance in scenarios with unbalanced classes, see Eq. (2). We calculated the model’s average weighted Accuracy, see Eq. 3, achieving a value of 88%, as detailed in Table 4. This average accuracy value was weighted by the number of samples for each mode to fairly reflect the contribution of each transport mode in its computation. Precision =TP TP +FPRecall =TP TP +FN (1) F1−score =2·Precision ·Recall Precision +Recall (2) Accuracy =TP +TN TP +TN +FP +FN (3) It is clear from the results shown in Fig. 5that while most of the transports are successfully predicted with an accuracy above 80%, it becomes evident that modes with a sparse amount of training data (see Bike, Train and e-Scooter in Fig. 2) tend to be misidentified more frequently by the model. This observation highlights an intriguing point: Bike and e-Scooter are frequently misclassified as one another by the model: (i) True Bike is wrongly predicted as e-Scotter with 27% probability; while (ii) true e-Scotter is wrongly predicted as Bike with 13% probability. This is both interesting and reasonable, considering that these modes of transport share similarities in acceleration and handling characteristics. Nevertheless, as we collect additional Bike and e-Scooter sensor samples from our volunteers (see Fig. 2), the predictive FIGURE 5. Confusion matrix of the hierarchical MobilitApp prediction model. Values are expressed as fractions of one. For example, Car is correctly recognized 96% of the time, it is confused with Bus 1% of the time, and it is confused with Train 1% of the time. TABLE 4. Classification report on the test set of the hierarchical MobilitApp prediction model. The average weighted Accuracy is 88%. accuracy of these transport modes in our MobilitApp model will certainly improve. The case of the Motorbike is also interesting, as it shows a higher accuracy despite having a similar number of samples to the bike and e-scooter. This is due to its inherent rolling-over driving behavior, which is well captured by the gyroscope, making it easier to identify this transport mode. After reviewing these results, we are now ready to conduct a comparative analysis of the four selected prediction models, evaluated through a robust methodological framework. The approach adopted for this comparison involved group-fold cross-validation with five folds [6], executed across three different random seeds. This procedure culminated in a total of 15 runs per model, ensuring a comprehensive assessment of each model’s performance consistency and resilience to varying data splits. The models compared in this section are the following: VOLUME 13, 2025 9
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection TABLE 6. Difference in storage size and performance of the former classification model and the current model optimized using Tensorflow Lite. TABLE 7. Comparison of the battery consumption of the MobilitApp application with other well-known Android applications. VI. PERFORMANCE EVALUATION OF MOBILITAPP To evaluate the performance of the MobilitApp tool, we focused on the integration of the state-based system for determining Macro-states (see Section V-A) and the kinetic motion classifier for Micro-state determination (see Section V-B). Four case study examples were analyzed to illustrate transitions between states: First experiment: Transition WALK to STATIONARY. Fig. 14 shows accelerometer samples and the state classification. In the observed experiment, two Micro-states were inaccurately labeled as MOVING due to potentially poor location data accuracy. Notably, the Macro-states show resilience against these anomalies, underscoring their robustness in filtering out isolated inaccuracies in Micro-state labeling. Second experiment: Transition STATIONARY to BUS. Fig. 15 shows acceleration peaks in the early samples, corresponding to the user boarding and settling into the bus. At this stage, the bus remains stationary as the passengers take their seats. Variability in the regularity of the analyzed windows is noted, attributable to inconsistent GPS signal quality which can cause delayed or occasionally premature location updates. We observe the absence of erroneous states, which underscores the clean and accurate capture of user activity during the experiment, along with the correct activity classification predicted by our MobilitApp tool. Third experiment: Transition BUS to WALK. The results depicted in Fig. 16 present some irregular windows although no spurious Micro-states. As in the other experiments, we can conclude that the MobilitApp tool accurately captured the user activity and correctly classified both the activity and transportation mode during the experiment. Fourth experiment: Real-life multimodal trips in Barcelona. After testing the system with the previous three experiments, the next phase involves a real-world evaluation of multimodal trips. Testing was conducted using several smartphone models not previously included in the dataset to simulate the experience of new users. Tests identified occasional inaccuracies in the location on the map, especially when using the subway. However, these inaccuracies were largely mitigated by the model’s robust prediction capabilities for the subway mode of transport, ensuring reliable classification despite GPS errors. To address this issue, we modified MobilitApp to handle GPS inaccuracies independently of the flowchart depicted in Fig. 9. Specifically, when MobilitApp detects a significant GPS error, it immediately invokes the machine learning model to predict the user’s mode of transport, allowing the system to identify whether the user is traveling on the subway. This approach effectively resolves issues related to GPS inaccuracies commonly encountered in the subway. Figure 17 presents two examples of tests performed with two smartphones that were not used during training (Samsung Galaxy S23 and Xiaomi Redmi Note 12). Similarly, Fig. 18 shows additional tests conducted with two other smartphones (Xiaomi HyperOS and Samsung Galaxy A54) that were not used during training. Neither of these devices contributed data to the dataset used for training the MobilitApp prediction model. Through these and numerous similar tests with various smartphones of new users, we observed that MobilitApp accurately predicts the different transportation modes used in experiments involving multimodal city travel. From extensive testing, we concluded that the detection system for Stationary and Walk activities is robust and performs effectively across most scenarios. An area of improvement is identified in high-traffic occasions with frequent stops, where the system sometimes incorrectly registers the Stationary activity for semi-stationary vehicles in a traffic jam. The emphasis of our design on robustness and accuracy over speed in detection produces this behavior. Nevertheless, we easily addressed this issue on the server with a script that detects semi-stationary patterns from consecutive vehicle predictions and classifies them as traffic jam events. This approach ensures accurate identification of traffic congestion while maintaining robust transport mode classification. In conclusion, the systematic evaluation of the Micro and Macro-state systems, along with the kinetic motion classifier across various scenarios, highlights MobilitApp’s robustness as a tool for urban planners to analyze citizens’ mobility habits. The system demonstrates resilience, particularly in accurately differentiating between WALK, STATIONARY, and MOVING states. Nonetheless, challenges were observed in high-traffic scenarios, where brief vehicle misclassifications as STATIONARY were observed during certain seconds. However, the majority voting scheme (see Fig. 7) after aggregating 100 seconds of data (5 microstates of 20 seconds each), effectively resolved this issue. The real-world multimodal trip trials showcased strong system performance for the most frequently used transport modes, even though minor misclassifications were occasionally observed during the initial seconds of movement. These findings emphasize the application’s significant potential for deployment in real-world urban environments. 16 VOLUME 13, 2025
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection FIGURE 14. Testing with a WALK →STATIONARY transition in the accelerometer samples. The 20-second Micro-states are shown in the left image, while the 100-second Macro-states are displayed in the right image. Total experiment time: 200 sec. FIGURE 15. Testing with a STATIONARY →BUS transition in the accelerometer samples. The 20-second Micro-states are shown in the left image, while the 100-second Macro-states are displayed in the right image. Total experiment time: 200 sec. FIGURE 16. Testing with a BUS →WALK transition in the accelerometer samples. The 20-second Micro-states are shown in the left image, while the 100-second Macro-states are displayed in the right image. A. MOBILITAPP USE CASE: MASSIVE MULTIMODAL TRIP DATA COLLECTION CAMPAIGN ‘‘HOW DO YOU COME TO THE CAMPUS?’’ Our work focuses on studying urban mobility through the analysis of citizens’ mobility habits. To this end, the aim is for the MobilitApp tool to serve as a resource for researchers and public entities responsible for urban mobility planning and management. MobilitApp allows users to upload a multimodal trip summary to the server included in a light text CSV file (each below 300 KB). The MobilitApp Multimodal Trip VOLUME 13, 2025 17
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection FIGURE 17. MobilitApp tests carried out with two smartphones that were not used during training. Two multimodal trips are shown: Car+Walk in the image on the left; Walk+Bus+Walk+Tram+Walk in the image on the right. FIGURE 18. Additional MobilitApp testing was performed with two other smartphones that were not used during training. Two multimodal trips are shown: Walk+Bus+Walk in the image on the left; Walk+Bus+Walk+Tram+Walk in the image on the right. dataset consists of files that capture optional user data (such as age range and gender) and record detailed information about multimodal trips, including predicted transport modes for each trip segment and transition location points (latitude and longitude) between different modes. Given the sensitivity of this user data and the associated privacy risks, the dataset cannot be publicly released in its current form. However, as mentioned in Section VII, we are actively researching privacy-preserving techniques that could be implemented to address this issue. In Fig. 19, an example of a multimodal trip (2nd and 3rd rows) is shown, identified with the same random ID (in the blue rectangle). The trip belongs to a woman aged 45-59, who first walked and then used the subway for the second section. The transfer point (from walking to subway) is marked in the GPS location highlighted by a green rectangle, with the transition time indicated by the red rectangle. The data from these multimodal captures are invaluable for studying mobility flows in any city, as exemplified by Barcelona in our case study. Two entities showing particular interest in this project are UPC Sostenible [5] and the ATM [1]. UPC Sostenible is keen on analyzing popular transport combinations among users at each UPC campus (more information in the UPC citizen science portal [7]). Meanwhile, ATM focuses on improving public transport in the Barcelona metropolitan are. MobilitApp facilitates a more comprehensive and faster analysis of daily mobility patterns compared to traditional methodologies, e.g. surveys of passengers about origin and end of the journey. Thus, once the MobilitApp prediction model was validated, we decided to carry out a Massive Multimodal Travel Data Collection Campaign within the community of the Universitat Politècnica de Catalunya (UPC) to compile a dataset of multimodal trips, composed of consecutive segments predicted using the MobilitApp tool. Additionally, many of the campaign volunteers also collaborated by generating unimodal labeled trips to help enlarge the MobilitApp Sensor Dataset [8], see Fig. 2. In collaboration with UPC Sostenible [5] and the ATM, we launched a 1-week data collection initiative in the UPC with the pilot data collection campaign ‘‘How do you come to the Campus?’’2[30], within the framework of the open and citizen science initiatives of the UPC [7]. This campaign included the participation of volunteers from different groups (students, professors, researchers and administrative staff), who contributed to testing the performance of the MobilitApp tool. They generated more than 200 multimodal trips during their trips to and from the university, which we stored in the MobilitApp Multimodal Trip Dataset for later analysis of the mobility habits of the UPC community. As mentioned above, this dataset contains sensitive information that could potentially reveal user locations. To comply with data privacy regulations and protect user confidentiality, we cannot publish it in its current form. However, we are actively investigating trajectory anonymization techniques that may enable the secure sharing of the dataset in the future. In the meantime, we can address potential privacy concerns by publishing aggregated data instead of raw datasets, ensuring that 2Mobility data collection campaign ‘‘How do you come to the Campus?’’ 15th-19th April 2023. 18 VOLUME 13, 2025
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection FIGURE 19. Example of multimodal trips collected with MobilitApp. Gender and Age range are asked just for statistical purposes. FIGURE 20. Screenshot of the dashboard with multimodal trips collected with the application. The MobilitApp Power Bi dynamic tool is available online [3]. FIGURE 21. Outcome of the MobilitApp Power Bi dynamic tool [3] filtered according to the selection shown in Fig. 20 for women, (left image) and for men (right image). Examples of what could be analyzed with the mobility information collected with our MobilitApp tool. individual user information remains protected. To achieve this, we have developed a MobilitApp Power Bi dynamic tool [3] to analyse multimodal trips. This dashboard implemented in Power Bi helps us to analyze the multimodal trip data collected by the MobilitApp tool during the campaign. Fig. 20 shows a screenshot of the dashboard, which shows a selection of 21 out of the 207 mobility segments. VOLUME 13, 2025 19
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection FIGURE 22. Outcome of the MobilitApp Power Bi dynamic tool [3] filtered showing all aggregated trips to the Les Corts district (where the UPC is located). A multimodal trip consists of one or more mobility segments, so that each phase of the trip corresponds to a different mode of transport. In the figure we have filtered the trips to match the options marked in black (gender =woman; age range =18-29, 45-59; transport mode =Bus, Metro, Car, Walk, or Bicycle; School =all; profiles =professors, researchers, students). The filtered results, depicted in Fig. 21, reveal walking as the most preferred mode of transportation. Additionally, gender-based differences in transportation choices were observed: women tended to rely more frequently on cars, while men showed a greater inclination toward using buses and subways during their multimodal trips. In addition, the tool shows the filtered trips added to a map, see Fig. 22, in this case from all the Districts of Barcelona to the Les Corts District (where the UPC is located). These are just a few examples of what can be analyzed with the multimodal travel information collected with our MobilitApp tool. The MobilitApp Power Bi dynamic tool can be accessed in this link [3]. VII. CONCLUSION AND FUTURE WORK In this study, we introduced MobilitApp, a deep learningbased tool for real-time transportation mode detection using smartphone sensor data, designed with a strong emphasis on minimizing computational costs. This optimization is crucial, as the tool is intended to operate on smartphones with limited processing power. MobilitApp provides valuable insights for urban planners, enabling data-driven strategies to improve mobility and support the transition to more sustainable cities. Our hierarchical prediction model combines CNNs for feature extraction and LSTM layers for temporal processing, enhanced with skip connections to improve performance. To further optimize computational efficiency on mobile devices, a kinematic motion classifier is integrated for early motion detection, reducing dependency on deep learning models. This research represents a thorough investigation into transport mode and human activity recognition, using the city of Barcelona as a case study to explore the dynamics of urban mobility. It bridges traditional analytical methods and advanced deep learning techniques, offering a detailed perspective on the state-of-the-art solutions in this domain. Through rigorous examination, the study highlights the transformative potential of deep learning models in solving complex urban mobility challenges. Key findings from this work not only illustrate the intricate nature of transport mode detection but also lay out directions for future exploration, pushing the boundaries of urban mobility research. This research demonstrates that deep learning provides a promising framework for accurately detecting and distinguishing transport modes, particularly when utilizing temporal data collected from various sensors. The implementation of these techniques sheds light on the complexities of managing and pre-processing sensory data, emphasizing the need for robust data-handling models to ensure reliable and high-quality outcomes. Importantly, the differentiation of user groups between training and evaluation phases is critical to simulate real-world conditions and evaluate the model under circumstances similar to its intended deployment. Additionally, it has been shown that combining recurrent architectures such as LSTM with convolutional networks helps to achieve a robust transport mode classification model using accelerometer, magnetometer, and gyroscope data. After completing various data collection campaigns, we have generated the MobilitApp Sensor Dataset [8], which is 20 VOLUME 13, 2025
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection publicly available. Finally, the adoption of a more complex hierarchical architecture has proven effective in improving the generalization capabilities of the model, leading to better performance and more accurate results. The developed DL model successfully predicts most transport modes with an accuracy exceeding 80%, achieving a weighted average accuracy of 88%. These promising results underscore the utility of our approach for predicting users’ transport modes during their travels. This capability is particularly valuable for urban planners and public service providers, as the MobilitApp tool enables detailed analysis of citizens’ mobility habits, contributing to the advancement of more sustainable urban mobility solutions. A significant achievement of this study is the development of an Android application named MobilitApp [4] that combines deep learning with statistical methods, proving robustness to the transport mode recognition system. This tool, capable of generating multimodal trip summaries, represents a leap forward in understanding urban mobility patterns, offering valuable insights for urban planners. The MobilitApp tool enables the collection of mobility data, which is stored in two datasets: •The MobilitApp Sensor Dataset, see Section III-D and Fig. 2. This dataset consists solely of values corresponding to the X, Y, and Z axes of three motion sensors: Accelerometer, Gyroscope, and Magnetometer. It does not contain any personal information and is publicly available at [8].This dataset is used to train our DL-based prediction model to determine the transportation mode used by the citizen. •The MobilitApp Multimodal Trip Dataset, see Section VI-A and Fig. 19. This dataset includes personal and sensitive user information, such as age range, gender, and the GPS location of each origin-destination pair for each segment of the trip (one segment per predicted transportation mode). This dataset includes sensitive location data, so it cannot be published in its current form due to privacy regulations. We are exploring anonymization techniques for secure sharing in the future. Meanwhile, only aggregated data will be released to protect user confidentiality. This dataset is designed to assist mobility planners in understanding users’ mobility habits and identifying potential strategies to enhance urban mobility, fostering a transition toward more sustainable transportation solutions. In conclusion, this work contributes to the existing body of knowledge in the area of transport mode recognition using deep learning and paves the way for future studies that will enhance and expand these systems’ capabilities. The potential impact of this research extends beyond academia, offering practical solutions and insights that could shape the future of urban mobility in cities like Barcelona. Moreover, the methodology developed in this work aims to contribute to more sustainable urban mobility in Barcelona and can be adapted for other cities as well. The MobilitApp tool can assist urban planners and public service providers to better understand the mobility habits of citizens, assess the impact of mobility improvement initiatives, and enhance public transportation services. For future work, the team is encouraged to collect additional data for underrepresented transport modes, such as e-scooter, Motorbike, Bike, and e-Bike. In this regard, we plan to conduct additional data-gathering campaigns at several universities that have invited us to collaborate during the European Sustainable Mobility Week. The goal is to raise awareness among the university community about adopting sustainable transportation for their daily commutes to campus. Moreover, additional data rebalancing techniques, such as SMOTE [12] and ADASYN [21], will be explored to improve the representation of these underrepresented classes with a limited number of gathered samples. Additionally, privacy concerns and the system’s interpretability must be carefully addressed before making the MobilitApp Multimodal Trip Dataset publicly available. In this regard, we are exploring anonymization techniques, such as trajectory obfuscation and differential privacy, to ensure user privacy and protect sensitive information, which is a necessary condition before making thist dataset publicly available. To effectively mitigate privacy concerns when the trajectory dataset is public, among a variety of privacy models available in the literature one might consider applying a differentially private data publishing mechanism. Essentially, differential privacy (DP) ensures that the presence or absence of a single individual trajectory in the dataset does not significantly affect the overall statistics. That is, the type of attacks DP protects against refers to those in which the adversary aims to determine whether an individual contributed their data to the dataset. This is achieved by adding carefully calibrated random noise to the data or query responses [41]. In the context of deep learning architecture research, we aim to explore the integration of Transformers into our framework. Additionally, the innovative application of Generative Adversarial Networks (GANs) [20] for data augmentation presents new opportunities for enhancing model performance and expanding research possibilities. These strategies represent the ongoing development of deep learning in interpreting complicated transport data, as they seek to overcome existing limits and investigate novel ways. On the one hand, Transformers are a deep learning architecture initially developed for Natural Language Processing (NLP) tasks, but their versatility extends to fields like computer vision and audio. Models like GPT (Generative Pretrained Transformer) leverage Transformers for tasks such as translation, text classification, and language generation. Transformers excel at capturing global dependencies, scaling to large datasets, and adapting to applications beyond text. On the other hand, GANs are lately being used for data augmentation, particularly in scenarios with limited datasets. GANs consist of a generator and a discriminator that compete VOLUME 13, 2025 21
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection in a minimax game, enabling the creation of synthetic data that closely resembles the original dataset. This approach enhances model performance by increasing data diversity, addressing class imbalances, and improving generalization. GAN-based augmentation is particularly effective in applications like image synthesis, speech generation, and medical data augmentation. ACKNOWLEDGMENT The authors would like to thank Miquel Gotanegra Estañol and Jaume Planas i Planas for their discussions and contributions during their degree and master’s theses, respectively. Additionally, they extend their gratitude to the volunteers who helped them develop the MobilitApp tool by sharing their mobility data. Finally, they also would like to thank Albert Villarroya Saladié from UPC Sostenible [5] and Francesc Calvet from the ATM [1] for their collaboration in organizing the data collection campaign ‘‘How do you come to the Campus?’’. ACRONYMS ADASYN Adaptive Synthetic Sampling. ATM Autoritat del Transport Metropolità. CCN Convolutional Neural Networks. DCBL Deep Convolutional Bidirectional LSTM. DNN Deep Neural Network. DP Differential Privacy. EWMA Exponentially Weighted Moving Average. FFT Fast Fourier Transform. HAR Human Activity Recognition. HDC HyperDimensional Computing. LSTM Long Short-Term Memory. ML Machine Learning. MD Mahalanobis distance. PCA Principal Component Analysis. PSD Power Spectral Density. SHL Sussex-Huawei Locomotion Dataset. SMOTE Synthetic Minority Over-sampling Technique. TMD Transport Mode Detection Dataset. UPC Universitat Politècnica de Catalunya. REFERENCES [1] Prof. Mónica Aguilar Igartua (UPC) and Mr. Francesc Calvet (ATM). Prediction of the Transportation Mode Used By the Citizens in Their Multimodal Trips From the Analysis of Smartphone Sensors. Collaboration agreement ATM. Accessed: Apr. 17, 2025. [Online]. Available: https://futur.upc.edu/35021523 [2] MobilitApp Application That Predicts the Transportation Mode Used. Accessed: Apr. 17, 2025. [Online]. Available: https://play.google.com/ store/apps/details?id=com.mobi.mobilitapp&pcampaignid=web_share [3] MobilitApp Power Bi Dynamic Tool to Analyse Multimodal Trips. Accessed: Apr. 17, 2025. [Online]. Available: https://acortar.link/o0Eszw [4] MobilitApp: Tool to Help Analyze the Mobility Flow of Citizens in the Metropolitan Area of Barcelona. Accessed: Apr. 17, 2025. [Online]. Available: https://mobilitapp.upc.edu [5] Gabinet D’Innovació I Comunitat De La Universitat Politècnica De Catalunya (UPC). Accessed: Apr. 17, 2025. [Online]. Available: https://sostenible.upc.edu [6] (2021). GroupKFold. [Online]. Available: https://scikit-learn.org/stable/ modules/generated/sklearn.model_selection.GroupKFold.html [7] (2023). UPC Citizen Science Portal. [Online]. Available: https://en.cienciaciutadana.upc.edu [8] M. A. Igartua, G. C. Ibáñez, and M. G. Estañol, ‘‘MobilitApp—Mobility sensor dataset,’’ CORA Repositori de Dades de Recerca, Barcelona, Spain, Dec. 2024, doi: 10.34810/data1948. Accessed: Apr. 17, 2025. [9] R. Alizadeh, Y. Savaria, and C. Nerguizian, ‘‘Characterization and selection of WiFi channel state information features for human activity detection in a smart public transportation system,’’ IEEE Open J. Intell. Transp. Syst., vol. 5, pp. 55–69, 2024, doi: 10.1109/OJITS.2023.3336795. [10] Android Developers. (2023). Jetpack Compose. Accessed: Sep. 19, 2023. [Online]. Available: https://developer.android.com/jetpack/compose [11] C. Carpineti, V. Lomonaco, L. Bedogni, M. D. Felice, and L. Bononi, ‘‘Custom dual transportation mode detection by smartphone devices exploiting sensor diversity,’’ in Proc. IEEE Int. Conf. Pervasive Comput. Commun. Workshops (PerCom Workshops), Mar. 2018, pp. 367–372, doi: 10.1109/PERCOMW.2018.8480119. [12] N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, ‘‘SMOTE: Synthetic minority over-sampling technique,’’ J. Artif. Intell. Res., vol. 16, pp. 321–357, Jun. 2002, doi: 10.1613/jair.953. [13] C. Cortes, M. Mohri, and A. Rostamizadeh, ‘‘L2 regularization for learning kernels,’’ 2012, arXiv:1205.2653. [14] P. Cunningham and S. J. Delany, ‘‘K-nearest neighbour classifiers—A tutorial,’’ ACM Comput. Surv., vol. 54, no. 6, p. 128, Apr. 2007, doi: 10.1145/3459665. [15] R. David, J. Duke, A. Jain, V. J. Reddi, N. Jeffries, J. Li, N. Kreeger, I. Nappier, M. Natraj, S. Regev, R. Rhodes, T. Wang, and P. Warden, ‘‘TensorFlow lite micro: Embedded machine learning on TinyML systems,’’ 2020, arXiv:2010.08678. [16] M. Erdelić, T. Erdelić, and T. Carić, ‘‘Dataset for multimodal transport analytics of smartphone users–Collecty,’’ Data Brief, vol. 50, Oct. 2023, Art. no. 109481, doi: 10.1016/j.dib.2023.109481. [17] R. Fallahzadeh, Z. E. Ashari, P. Alinia, and H. Ghasemzadeh, ‘‘Personalized activity recognition using partially available target data,’’ IEEE Trans. Mobile Comput., vol. 22, no. 1, pp. 374–388, Jan. 2023, doi: 10.1109/TMC.2021.3071434. [18] H. Gjoreski, J. Bizjak, M. Gjoreski, and M. Gams, ‘‘Comparing deep and classical machine learning methods for human activity recognition using wrist accelerometer,’’ in Proc. IJCAI Workshop Deep Learn. Artif. Intell., 2016, pp. 1–7. [Online]. Available: https://api.semantic scholar.org/CorpusID:34272612 [19] H. Gjoreski, M. Ciliberto, L. Wang, F. J. Ordonez Morales, S. Mekki, S. Valentin, and D. Roggen, ‘‘The University of Sussex-Huawei locomotion and transportation dataset for multimodal analytics with mobile devices,’’ IEEE Access, vol. 6, pp. 42592–42604, 2018, doi: 10.1109/ACCESS.2018.2858933. [20] I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, ‘‘Generative adversarial networks,’’ in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2014, pp. 2672–2680. [Online]. Available: https://papers.nips.cc/ paper/2014/hash/5ca3e9b122f61f8f06494c97b1afccf3-Abstract.html [21] H. He, Y. Bai, E. A. Garcia, and S. Li, ‘‘ADASYN: Adaptive synthetic sampling approach for imbalanced learning,’’ in Proc. IEEE Int. Joint Conf. Neural Netw. (IEEE World Congr. Comput. Intell.), Jun. 2008, pp. 1322–1328, doi: 10.1109/IJCNN.2008.4633969. [22] M. A. Hearst, S. T. Dumais, E. Osuna, J. Platt, and B. Scholkopf, ‘‘Support vector machines,’’ IEEE Intell. Syst. Their Appl., vol. 13, no. 4, pp. 18–28, Jul. 1998, doi: 10.1109/5254.708428. [23] S. Hemminki, P. Nurmi, and S. Tarkoma, ‘‘Accelerometer-based transportation mode detection on smartphones,’’ in Proc. 11th ACM Conf. Embedded Networked Sensor Syst., Nov. 2013, pp. 1–14, doi: 10.1145/2517351.2517367. [24] S. Hochreiter and J. Schmidhuber, ‘‘Long short-term memory,’’ Neural Comput., vol. 9, no. 8, pp. 1735–1780, Nov. 1997, doi: 10.1162/neco.1997.9.8.1735. [25] J. Iskanderov and M. A. Guvensan, ‘‘Breaking the limits of transportation mode detection: Applying deep learning approach with knowledge-based features,’’ IEEE Sensors J., vol. 20, no. 21, pp. 12871–12884, Nov. 2020, doi: 10.1109/JSEN.2020.3001803. [26] J. V. Jeyakumar, E. S. Lee, Z. Xia, S. S. Sandha, N. Tausik, and M. Srivastava, ‘‘Deep convolutional bidirectional LSTM based transportation mode recognition,’’ in Proc. ACM Int. Joint Conf. Int. Symp. Pervasive Ubiquitous Comput. Wearable Comput., Oct. 2018, pp. 1606–1615, doi: 10.1145/3267305.3267529. 22 VOLUME 13, 2025
G. C. Ibañez et al.: MobilitApp: A DL-Based Tool for Transport Mode Detection [27] D. P. Kingma and J. Ba, ‘‘Adam: A method for stochastic optimization,’’ 2014, arXiv:1412.6980. [28] A. Manivannan, E. J. Willemse, B. T. Balamurali, W. C. B. Chin, Y. Zhou, B. Tunçer, A. Barrat, and R. Bouffanais, ‘‘A framework for the identification of human vertical displacement activity based on multisensor data,’’ IEEE Sensors J., vol. 22, no. 8, pp. 8011–8029, Apr. 2022, doi: 10.1109/JSEN.2022.3157806. [29] G. J. McLachlan, ‘‘Mahalanobis distance,’’ Resonance, vol. 4, no. 6, pp. 20–26, Jun. 1999, doi: 10.1007/bf02834632. [30] M. Aguilar, A. Villarroya, M. Gotanegra, G. Caravaca, and M. De Mingo, ‘‘Anàlisi de la campanya de recollida de dades de mobilitat a la comunitat UPC amb l’eina MobilitApp,’’ Item, Revista De Biblioteconomia I Documentació, vol. 2024, no. 76, pp. 1–22, Jul. 2024, doi: 10.60940/itemn76id430338. [31] N. Nachar, ‘‘The mann-whitney U: A test for assessing whether two independent samples come from the same distribution,’’ Tuts. Quant. Methods Psychol., vol. 4, no. 1, pp. 13–20, Mar. 2008, doi: 10.20982/tqmp.04.1.p013. [32] K. O’Shea and R. Nash, ‘‘An introduction to convolutional neural networks,’’ 2015, arXiv:1511.08458. [33] L. R. Randleff, J. B. Wanscher, and M. Zabic, ‘‘Distributed travel mode estimation,’’ in Proc. 19th ITS World Congr., Jan. 2012, pp. 1–6. [34] S. Richoz, L. Wang, P. Birch, and D. Roggen, ‘‘Transportation mode recognition fusing wearable motion, sound, and vision sensors,’’ IEEE Sensors J., vol. 20, no. 16, pp. 9314–9328, Aug. 2020, doi: 10.1109/JSEN.2020.2987306. [35] L. Rokach and O. Maimon, ‘‘Decision trees,’’ in Data Mining and Knowledge Discovery Handbook, vol. 6. Boston, MA, USA: Springer, 2005, pp. 165–192, doi: 10.1007/0-387-25465-X_9. [36] S. Lee and C. Lee, ‘‘Revisiting spatial dropout for regularizing convolutional neural networks,’’ Multimedia Tools Appl., vol. 79, nos. 45–46, pp. 34195–34207, Dec. 2020, doi: 10.1007/s11042-020-09054-7. [37] Q. Tang, K. Jahan, and M. Roth, ‘‘Deep CNN-BiLSTM model for transportation mode detection using smartphone accelerometer and magnetometer,’’ in Proc. IEEE Intell. Vehicles Symp. (IV), Jun. 2022, pp. 772–778, doi: 10.1109/IV51971.2022.9827275. [38] L. Wang, H. Gjoreski, M. Ciliberto, S. Mekki, S. Valentin, and D. Roggen, ‘‘Enabling reproducible research in sensor-based transportation mode recognition with the sussex-huawei dataset,’’ IEEE Access, vol. 7, pp. 10870–10891, 2019, doi: 10.1109/ACCESS.2019.2890793. [39] B. Xu, N. Wang, T. Chen, and M. Li, ‘‘Empirical evaluation of rectified activations in convolutional network,’’ 2015, arXiv:1505.00853. [40] Y. Yao, W. Liu, G. Zhang, and W. Hu, ‘‘Radar-based human activity recognition using hyperdimensional computing,’’ IEEE Trans. Microw. Theory Techn., vol. 70, no. 3, pp. 1605–1619, Mar. 2022, doi: 10.1109/TMTT.2021.3134992. [41] À. Miranda-Pascual, P. Guerra-Balboa, J. Parra-Arnau, J. Forné, and T. Strufe, ‘‘An overview of proposals towards the privacy-preserving publication of trajectory data,’’ Int. J. Inf. Secur., vol. 23, no. 6, pp. 3711–3747, Dec. 2024, doi: 10.1007/s10207-024-00894-0. GERARD CARAVACA IBAÑEZ received the M.S. degree in artificial intelligence engineering from the Universitat Politècnica de Catalunya (UPC), Barcelona, Spain, in 2024. His expertise extends across computer vision, natural language processing, mobility, and a spectrum of machine learning challenges. His research interest includes deep learning models for predicting mobility activities. LUIS J. DE LA CRUZ LLOPIS (Member, IEEE) received the Graduate and Ph.D. degrees in telecommunications engineering from the Universitat Politècnica de Catalunya (UPC), Barcelona, Spain, in 1994 and 1999, respectively. He is currently an Associate Professor with the Department of Network Engineering, UPC. He is also part of the Smart Services for Information Systems and Communication Networks (SISCOM) Research Group. His current research interests include the application of machine learning techniques in wireless multi-hop and C-V2X networks and also the development of IoT smart services. ADRIAN CATALÍN DIACONEASA received the degree in telecommunications engineering from the Universitat Politècnica de Catalunya (UPC), minor in telematic systems. He has experience with various software development tools obtained through participation in multidisciplinary projects. His research interest includes developing deep learning models to predict mobility patterns and activities. ALBERTO BAZÁN GUILLÉN received the Engineering degree in telecommunications and electronics and the M.Sc. degree in telematics from the Central University ‘‘Marta Abreu’’ of Las Villas (UCLV), Santa Clara, Cuba, in 2020 and 2022, respectively. He is currently pursuing the Ph.D. degree in telematics engineering with the SISCOM Research Group, Universitat Politècnica de Catalunya (UPC). His research interests include wireless networks, vehicular networks, electrical vehicles, urban mobility, the IoT, and federated learning models. MÓNICA AGUILAR IGARTUA (Member, IEEE) received the M.Sc. and Ph.D. degrees in telecommunication engineering from the Universitat Politècnica de Catalunya, Barcelona, Spain, in 1995 and 2000, respectively. She is the author of more than 30 journal articles, has chaired several conferences and is a member of the Editorial Board of the Ad Hoc Networks journal. Her research interests include design and performance evaluation of routing protocols for vehicular networks, electric vehicle, platooning, machine learning, smart city services, and sustainable urban mobility. She belongs to the SISCOM Research Group, focused on smart services for information systems and communication networks. VOLUME 13, 2025 23