scieee AI-readable full text Open interactive document viewer

Lane type classification using advanced neural architectures

Moreno Punzano, Alex

Abstract

In this thesis we have explored the feasibility of using 3D Convolution Neural Networks (C3D) for a lane type classification task for micromobility vehicles, such as e-scooters or bicycles, with the objective of improve both drivers and pedestrian security. To accomplish our objectives, different configurations of expanded 3D architectures (X3D) have been trained and tested. The results obtained suggest that efficient C3D network models can perform lane type classification tasks better than 2D image classification systems in similar conditions.

Full text

Lane type classification using advanced neural architectures Master Thesis submitted to the Faculty of the Escola T`ecnica d’Enginyeria de Telecomunicaci´o de Barcelona Universitat Polit`ecnica de Catalunya by Alex Moreno Punzano In partial fulfillment of the requirements for the master in Master’s degree in Advanced Telecommunication Technologies ENGINEERING Advisor: Josep Ramon Morros, Elisa Sayrol Barcelona, 23 of January of 2022 Contents List of Figures 3 List of Tables 3 1 Introduction 8 1.1 Statementofpurpose.............................. 8 1.2 Requirements and specifications . . . . . . . . . . . . . . . . . . . . . . . . 8 1.3 Methods and procedures . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 1.4 GanttDiagram ................................. 9 1.5 Deviations and incidences . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 2 State of the art 10 2.1 Micromobility safety using AI . . . . . . . . . . . . . . . . . . . . . . . . . 10 2.2 Video classification Techniques . . . . . . . . . . . . . . . . . . . . . . . . . 10 2.2.1 Recurrent Convolutional Neural Networks and Attention Models . . 11 2.2.2 3D Convolutional Networks . . . . . . . . . . . . . . . . . . . . . . 11 2.2.3 Expanded 3D architectures: X3D . . . . . . . . . . . . . . . . . . . 12 3 Methodology 14 3.1 Dataset ..................................... 14 3.2 C3DModels................................... 15 3.3 End-to-endframework ............................. 15 3.4 Transfer learning and fine-tuning . . . . . . . . . . . . . . . . . . . . . . . 15 3.5 DataPreparation................................ 16 3.6 Training..................................... 16 4 Results 18 4.1 Modelcomparison................................ 18 4.2 Multi-class classification with X3D . . . . . . . . . . . . . . . . . . . . . . 19 4.3 Binary classification with X3D . . . . . . . . . . . . . . . . . . . . . . . . . 19 4.4 Binary classification simulation test . . . . . . . . . . . . . . . . . . . . . . 20 5 Conclusions and future development: 23 References 24 Appendices 26 A Results 26 A.1 Multiclass lanetype classification: Confusion matrices . . . . . . . . . . . . 26 A.2 Multiclass lanetype classification: Metrics . . . . . . . . . . . . . . . . . . . 27 A.3 Binary lanetype classification: Confusion matrices . . . . . . . . . . . . . . 29 A.4 Binary lanetype classification: Metrics . . . . . . . . . . . . . . . . . . . . . 30 A.5 Simulation Test: Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31 A.6 Results: Original Database . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 2 List of Figures 1 Project’sGanttdiagram ............................ 9 2 Representation of a 3D convolution. . . . . . . . . . . . . . . . . . . . . . . 12 3X3D networks progressively expand a 2D network across the following axes: Temporal duration γt, frame rate γτ, spatial resolution γs, width γw, bottleneck width γb, and depth γd[16]. (Image Source: [16]) . . . . . . . . 13 4 Examples of images of each class . . . . . . . . . . . . . . . . . . . . . . . 14 Listings List of Tables 1 Transformation values depending on the model used. . . . . . . . . . . . . 16 2 Modelcharacteristics.............................. 18 3 Results obtained in multi-class classification experiment. . . . . . . . . . . 19 4 Results obtained in binary classification experiment. . . . . . . . . . . . . . 20 5 F-score and inference time in X3D models, fr refers to inference time . . . 21 6 F-score and inference time in X3D models (original database) . . . . . . . 22 7 Multi-class classification confusion matrices (X3D) . . . . . . . . . . . . . . 26 8 Multi-class confusion matrices (ResNet) . . . . . . . . . . . . . . . . . . . . 27 9 Multi-class classification metrics (X3D) . . . . . . . . . . . . . . . . . . . . 27 10 Multi-class classification metrics (ResNet) . . . . . . . . . . . . . . . . . . 28 11 Binary classification confusion matrices (X3D) . . . . . . . . . . . . . . . . 29 12 Binary classification confusion matrices (ResNet) . . . . . . . . . . . . . . 29 13 Binary classification metrics (X3D) . . . . . . . . . . . . . . . . . . . . . . 30 14 Binary classification metrics (ResNet) . . . . . . . . . . . . . . . . . . . . . 30 15 Simulation test: Results (X3D) . . . . . . . . . . . . . . . . . . . . . . . . 31 16 Simulation test: Results (ResNet) . . . . . . . . . . . . . . . . . . . . . . . 32 17 Results original database (X3D) . . . . . . . . . . . . . . . . . . . . . . . . 33 3 Revision history and approval record Revision Date Purpose 0 16/01/2022 Document creation 1 19/01/2022 Document revision DOCUMENT DISTRIBUTION LIST Name e-mail Alex Moreno Josep Ramon Morros Elisa Sayrol Written by: Reviewed and approved by: Date 16/01/2022 Date 22/01/2022 Name Alex Moreno Punzano Name Ramon Morros, Elisa Sayrol Position Project Author Position Project Supervisor 4 Abstract In this thesis we have explored the feasibility of using 3D Convolution Neural Networks (C3D) for a lane type classification task for micromobility vehicles, such as e-scooters or bicycles, with the objective of improve both drivers and pedestrian security. To accomplish our objectives, different configurations of expanded 3D architectures (X3D) have been trained and tested. The results obtained suggest that efficient C3D network models can perform lane type classification tasks better than 2D image classification models in similar conditions. 5 Resum En aquesta tesi hem explorat la viabilitat d’utilitzar Xarxes Neuronals Convolucionals 3D (C3D) per portar a terme un sistema de classificaci´o de tipus de via per a vehicles de micromobilitat, com scooters el`ectrics o bicicletes, amb l’objectiu de millorar la seguretat tant de conductors com de vianants. Per complir amb els nostres objectius, s’han entrenat i testejat diverses configuracions d’una xarxa 3D expandida (X3D). Els resultats obtinguts en els experiments realitzats suggereixen que els models xarxes C3D eficients poden desenvolupar tasques de classificaci´o de tipus de v´ıa millor que sistemes de classificaci´o d’imatges 2D en condicions similars. 6 Resumen En esta tesis hemos explorado la viabilidad de utilizar Redes Neuronales Convolucionales 3D (C3D) para llevar a cabo un sistema de clasificaci´on de tipo de v´ıa para veh´ıculos de micromovilidad, como scooters el´ectricos o bicicletas, con el objetivo de mejorar la seguridad de conductores y peatones. Para cumplir con nuestros objetivos, hemos entrenado y testeado varias configuraciones de una red 3D expandida (X3D). Los resultados obtenidos sugieren que los modelos de redes C3D eficientes pueden desempe˜nar tareas de clasificaci´on de tipos de v´ıa mejor que modelos de clasificaci´on de imagen 2D en condiciones similares. 7 1 Introduction 1.1 Statement of purpose In recent years the technologies related to video classification have advanced a lot. Thanks to achievements in computational speed and deep learning architectures, the branch of the Convolutional 3-Dimensional Neural Networks (C3D) has been in the spotlight for computer vision engineers to design video classification tasks. The most relevant progress in C3D related to this project are the new types of architectures that have been implemented in the last few years, which have grown in efficiency with respect to the original models, and make it more feasible to run these models in real-time or in small devices such as mobile phones. The objective of this project is to train and test different C3Ds models in the task of classifying different lane types for micromobility vehicles. Once we have done that, we will explore how well these systems can perform in this task, having in mind the efficiency of the systems implemented. In the end we want to implement a system capable of detecting the type of lane the user is circulating through a camera located in front of the vehicle, and if it is detected that the user is circulating by the sidewalk, a warning signal must be sent to the driver. This thesis has been promoted by the micromobility project of the Computer Vision department of the UPC. The main goal of this project is to make city rides safer for both pedestrians and drivers. Within this project, several thesis related to microvehicle safety [20] and lane classification [18] have already been carried out. Although these works have been of great help in developing some aspects of this thesis, this project is not the direct continuation of any of them. 1.2 Requirements and specifications To conduct this project it will be required to implement a system capable of training and testing the different C3D architectures. Then, a comparative analysis among all tested models will be performed, therefore it will be necessary to find metrics to fairly compare the models in terms of accuracy and efficiency. Finally, conclusions will have to be drawn and further development and improvements will be determined. A Github directory has been created in order to save the final implementation, in case someone wants to further develop this project [19]. 1.3 Methods and procedures In order to implement our models, Python will be selected as the base programming language. This is because Python is an open-source language and that most of the video classification packages, models and other utility resources can be found implemented in this language. About the video dataset used to train and test the models, we are going to use a collection of video data obtained by first hand in Barcelona. 8 1.4 Gantt Diagram Now Phases of the Project 2021 2022 September October November December January . Planning 100% complete Research 100% complete Prepare the environment End-to-End Implementation 100% complete End-to-End Implementation Training Phase 100% complete Training 100% complete Test Documentation 100% complete State-Of-Art 100% complete Thesis Figure 1: Gantt diagram of the project 1.5 Deviations and incidences Respect the initial project planning, there has been a small delay. The end-to-end framework implementation took a couple of extra weeks due to several causes. Therefore the training phase started around mid-December. 9 To fine-tune the network for our purposes we are going to change the output layer of the pretrained model, whose default size is [2048,400] (2048 input nodes and 400 output nodes, one output node per action class of the Kinetics-400 dataset) to an input of size [2048,5], with 5 output nodes, as we have 5 output classes. Then we train the system as it’s explained on the next section. 3.5 Data Preparation Before starting the training phase, the dataset is loaded into a dataloader. The dataloader will encapsulate each video with its correspondent label and will prepare the data for the neural network. To prepare the data, some spatio-temporal transformations of the input videos are required. This transformation will depend on the X3D configuration, as can be seen in the table below: Model Temporal duration (γt) Frame Rate (γτ) Spatial Resolution (112γs) M 16 5 252 S 13 6 182 XS 4 12 182 Table 1: Transformation values depending on the model used. The transformation process will consist of the following steps: •First the input video is normalized and we extract a specific number of frames (γt) at a given frame rate (γtau) that will depend on the model. •Then we scale the image determining that the shortest side will be the selected spatial resolution (182 pixels in XS and S models and 252 in M configurations). •Finally, we extract a square central crop of the images that form the video, of the same size as the shortest side of the image. This transformation of the video will be inserted into the network. Additionally, long videos (of length superior to 4 seconds) will be divided in 2-second clips as a previous step to this transformation. 3.6 Training First it’s important to clarify that for this project training will consist of fine-tuning different pretrained X3D configurations, therefore we are not going to follow the training process explained in section 2.2.3. Once all the input data is on the dataloader, the training phase starts. The training phase is based on the forward-backward optimization algorithm. First, we define some initial weights to each node that defines the network (in our case, we start with the pre-trained networks’s weights). The forward step simply consists of passing the input through the network and saving the output. 16 Once we have the result, we need to define an assistant function that helps us interpret how well our network is working. We commonly use a Loss function, which will compare the differences between the predictions of our network and the true labels. This is a classification problem, with 5 different classes. In this case cross-entropy log-likelihood is used as criterion, and therefore a LogSoftmax layer will be placed as output of the network. We also are going to work with binary classification later, in this second case binary entropy loss is going to be used as a loss function. Back-propagation procedure will use this loss function and an optimizer to update the network’s weights, with the objective of improving the results on the next forward step. During back-propagation we are going to compute the derivative of the final node output with respect to nodes inputs. This will help to indicate which parameters are responsible for the most error, so the optimizer can decide which weights need to be updated while loss function will be used as criteria of how well our network is doing. In our case we are using stochastic gradient descent (SGD) as our optimizer. Parameters related about how our optimizer is modifying the weights are: a learning rate of 0.1, a weight decay of 10−4) and a momentum of 0.9. Fine-tuning adds a new layer of complexity to the training phase as we can choose which layers update and which not by freezing them. The training process will proceed as explained above but frozen layer’s weights won’t be updated during the back-propagation step. In any case, in this particular approach we have not frozen any layer, so all weights will be updated. Once weights have been updated, the forward-backward cycle will be repeated for the next video, and so on. Once we have trained with all our available data, we start again. A cycle through all the dataset is called an epoch. We are going to perform 120 epochs per training, as the validation’s accuracy stabilizes around that number. Each cycle does not necessarily improve the results with respect to the previous epoch, so we are going to save the model with the lowest validation loss among all epochs. 17 4 Results This section is focused on analyzing the results of different experiments that involve the distinct X3D configurations we have chosen to train. Additionally, a 2D ResNet has been fine-tuned to compare with a 2D image classification network that does not use information from previous frames. This resnet model has been trained with the same train dataset that we are going to use for video classification. Resnet input images will be normalised and resized into 256 x 256 pixels before inserting them into the network. Then, we are going to train during 20 epochs, when validation accuracy is stabilized. 4.1 Model comparison As we are keen on finding the most efficient system, first we are going to start comparing the different systems used in terms of number of parameters, memory size and GFLOPs required to obtain one prediction (from a 2-second clip in the case of X3Ds and from 1 image in the case of the resnet), as can be seen in Table 2. Model # Params (M) Size (MB) Prdiction cost (GFLOPS) X3D-M 3,79 23 6,8 X3D-S 3,79 23 2,8 X3D-XS 3,79 23 1,4 ResNet-18 11,2 128 1,8 Table 2: Model characteristics The first thing we can notice is that all X3D configurations have the same number the parameters and memory size. This is due to the axes related with the network (see figure 3) are the same in these configurations, therefore all these models are using the same base architecture. However, inputs are different depending on the configuration (see table 1), and consequently the computational cost increases in the bigger models. About how the computational cost increases in function of the input, the number of images that we stack at the input (γt) linearly affects the computational cost of the network whereas the system is very sensitive to modifying the size of the images (γs), where the computational cost increases exponentially as images grow. In comparison with the ResNet-18, X3D models are notably smaller in terms of the number of parameters and the memory size. In terms of computational cost, the only model capable of making a prediction with a smaller computational cost is the XS. However, remember that ResNets predicts one single frame and the X3D models give the prediction of a block of frames, and therefore they are highly efficient overall. In comparison with other state-of-the-art C3D architectures, the difference in terms of number of parameters and computational cost is notable; architectures like ResNet(2+1)D or Slowfast-R50 have approximately 30 millions of trainable parameters each one [2] and are more expensive in terms of FLOPs, obtaining similar results. 18 4.2 Multi-class classification with X3D In the first experiment we are going to test how well X3D classifies the different lanetypes mentioned. We are going to give as input the full video, and we are going to make 1 prediction for the video. Given as input the full video is not ideal because one axis of X3D models is the number of frames processed (γt). Most of the clips on the database have a duration inferior to the minimum necessary to extract the block of information with the axis proposed (which is γtγτ). Therefore we are going to omit γtfor this and the following experiment. In the case of the ResNet, we have transformed the train, validation and test datasets into images, and 2 experiments have been made: the first one consists of classificate all images of the test database while on the second experiment we classify 1 random frame of each video in the database. In order to test the network’s performance, we are going to generate a confusion matrix from the predictions obtained and then we are going to compute the precision and the recall of each class against the others. Finally we are going to estimate the average F1-score of all classes. Additionally we are going to compute a weighted F-score considering the number of videos of each class. As we have 4 different models, all the confusion matrices and the metrics computed can be seen in detail in the appendices in A.1 and A.2. Below we can see an overview of the results obtained: Model Avg. F1-Score (weighted) Avg. F1-Score (non-weighted) X3D-M 0,947 0,881 X3D-S 0,904 0,851 X3D-XS 0.841 0.796 ResNet-18 0,835 0,791 ResNet-18 (1 frame) 0,824 0,784 Table 3: Results obtained in multi-class classification experiment. We can observe in Table ?? that all X3Ds have outperformed the ResNet-18 architecture. Among all X3Ds, the M configuration has obtained the best results. Nevertheless, no configuration has been able to exceed a f-score of 0,9, although results are not bad at all. If we look at the confusion matrices in detail, we can see that whereas the X3D networks succeed in classifying most of the lanes selected, it notably under-performs when try to classify the class ’crosswalk’. This could be a consequence of the fact that we have less samples of this class and respect the others. 4.3 Binary classification with X3D The next experiment will be the same as described above but now we are going to divide the data into 2 different labels: ’sidewalk’ and ’no sidewalk’. The ’sidewalk’ class will remain the same with respect the previous experiment whereas all the other classes 19 will be merged as the class ’no sidewalk’. Clarify that we are not going to map results obtained in the previous experiments, a new network will be fine-tuned for this particular configuration. With this new binary division, the number of clips of both classes are more balanced than the previous experiment, as we have a total of 1235 videos labeled as ’sidewalk’ and 979 videos tagged as ’other’ (or ’no-sidewalk’). Below there is an overview of the results obtained on the test dataset. Detailed confusion matrices and metrics could be seen in detail in appends A.3 and A.4. Model Avg. F1-Score (weighted) Avg. F1-Score (non-weighted) X3D-M 0,991 0,990 X3D-S 0,988 0,988 X3D-XS 0,986 0,986 ResNet-18 0,9422 0,9421 ResNet-18 (1 frame) 0,9399 0,9389 Table 4: Results obtained in binary classification experiment. In this case all X3Ds models have achieved similar results, outperforming the ResNet-18 again and being the X3D-M the best of all them. Overall this experiment has given very good performance results for the X3D configurations. 4.4 Binary classification simulation test This last section will consist of simulating a real usage of the network developed. Specifically, this experiment will consist of a system that classifies if a vehicle is driving by the sidewalk or not given an input video. In this case we are going to use a new test dataset with ’large’ videos (about 15 seconds each). The system will divide these in clips of 2 seconds duration, to replicate the training clips, and then will classify them using the networks trained in ??. With these conditions we are going to perform 2 different experiments: The first experiment consists of estimating the accuracy for each individual video. To do so, the following conditions have been defined: •if prediction is ’sidewalk’ and label is ’sidewalk’ at any moment of the video →True Positive •if prediction is ’sidewalk’ and label is ’no sidewalk’ for all inputs of the same video →False Positive •if prediction is ’no sidewalk’ and label is ’no sidewalk’ at any moment of the video →True Negative •if prediction is ’no sidewalk’ and label is ’sidewalk’ for all inputs of the same video →False Negative 20 The second experiment will consist of classifying all the clips of all the videos, similarly to what we’ve done in the previous section. This way we could see in detail if the network is performing as we expect, as we expect a higher number of True Positives and Negatives in experiment 1. Finally, we are going to compute inference time. To do so, we are going to define the time that elapses between when we take the first set of images and when the model returns a result. It’s important to mention here in which system we are going to run our model. We are going to run this test in GPU model Nvidia GTX 1080 Ti with 3584 cores. To fairly compare the X3D architectures with the ResNet-18 in this experiment, we are going to test ResNet in 2 second clips too. We are going to accumulate the results obtained by processing all the frames that form that clip and we are going to compute the average probability of all frames given that 2-second period. Therefore, the inference time will be the time that has passed until all the clips have been processed. Additionally we have made this test at 2 different frame rates: 1 and 10, to compare the differences in terms of F-score and inference time. These have been the results obtained, and can be seen in detail in appendices A.5 andA.6: Model F-score (Exp1) F-score (Exp1) Inference time (ms) X3D-M 0,971 0,79 9606 X3D-S 0,694 0,477 8180 X3D-XS 0,667 0,464 5192 ResNet-18 (fr = 1) 0,890 0,770 14760 ResNet-18 (fr = 10) 0,799 0,772 2386 Table 5: F-score and inference time in X3D models, fr refers to inference time These results are quite inconsistent with the ones of the previous experiments. The Fscores in experiment 2 are much lower from what we expect, of both X3D and ResNet models, and the inference time is much higher with respect to previous experiments. There does not seem to be a clear cause for this results. Besides that, the model that has been capable of obtaining the best results is the X3D-M, having the higher inference time at the same time. In fact, the inference time seems to increase as the model is bigger. In any case, the inference times obtained are so big that a real-time implementation does not seem feasible. If we repeat this experiment with the original data (the same test set used in 4.3, formed by short two-second clips), results obtained are much more similar to the previous section, and inference time is much more smaller: 21 Model F-Score (Exp1) F-Score (Exp1) Inference time (ms) X3D-M 0,932 0,907 703 X3D-S 0,977 0,969 653 X3D-XS 0,918 0,905 570 ResNet-18 - - 71 Table 6: F-score and inference time in X3D models (original database) However, comparing this table with table 4 it’s observable that there have been a slight drop of the F-Scores in M and XS configurations, whereas the X3D-S configuration has given a similar performance. For this case, the inference time increases as bigger is the model too. In comparison, with this same dataset, the ResNet-18 has achieved an inference time of 71ms per frame, a much smaller value, but only is processing one frame, whereas X3D models are processing 2-second blocks. In this case a real-time implementation seems more feasible. Given the inconsistencies between the results that can be seen in Tables 5 and 6, which contain the results of the simulation experiments using the ’long’ test dataset and ’short’ test dataset respectively, some hypothesis about what could have happen can be extracted. •There seems to be a problem with the formatting in long videos. Remind that long videos have been divided in 2-second clips as a previous step of performing the experiments. In theory, the initial conditions of both experiments are very similar. However, their performance has been very different. The main proof in favour of this hypothesis is that the behaviour with the ’short’ dataset is similar to the one seen in 4.3 whereas the ’long’ dataset experiment performance it’s notably different, in terms of results and inference time. Another point is that there has been an under-performance in both X3D and ResNet models respect train with the ’short’ database. •We can’t exclude that there may be problems related with the training or the networks. Maybe the system does not generalize well, even though this theory does not explain the increase on the inference time and does not explain the underperformance of both X3D and ResNet models. 22 5 Conclusions and future development: In this thesis, we have explored how well C3D architectures can perform in a lanetype classification task designed for micromobility vehicles. To carry this through, first there has been done a research on which architecture could fit better the purpose of this thesis. The selected architecture has been an efficient C3D network, the expanding convolutional 3D networks, or simply X3Ds. Then, this was necessary to implement a framework capable of loading and fine-tuning different pretrained X3D configurations. Finally, the different models have trained and tested to see how they perform. In the light of the results obtained in short videos we have demonstrated that C3Ds architectures that have been trained for action recognition tasks can be fine-tuned and can give good results in lane classification tasks. Plus, we have seen that X3Ds can outperform an state-of-the-art image classification architecture as it’s a ResNet-18 in similar conditions (having 3 times less parameters). However, results obtained with the ’long’ videos dataset, have been inconsistent even though with the original dataset the system has achieved good results. The main hypothesis is that this under-performance is caused for something related with the ’long’ videos formatting. Further development will be necessary to confirm that this has been the problem and solve it if it’s possible, as we are designing an ADAS and it is essential that the results obtained are consistent for any input video. Regarding further development, this thesis can be considered a first approach to X3D architectures, and there are several ways to develop. One of the most interesting approaches could be to train an X3D from scratch in order to optimize for the best results for our task. However, this does not seem feasible due to the small amount of data that is currently available. Nonetheless there are still a lot of experiments that may improve the results obtained. For example, it could be interesting to make a partial X3D training, and fine-tune the same base model trying different spatio-temporal input configurations (would be like training an X3D from scratch but only modifying the axis related with the input size). Another feasible experiment that could uplift results is to train a system again using the lowercentral crop instead the central one, as probably it’s the part of the picture that contains the most substantial information to perform this particular task. To sum up, given the sudden appearance of micromobility vehicles in cities, there is a need to make these as safe as possible for everyone. Adapting the vehicle’s velocity to the type of lane is a way to accomplish this, thus systems capable of performing lane classification tasks are required. Efficient C3Ds and in particular X3Ds have proven to be one possible option to carry through this task, outperforming image classification systems in similar conditions. Even though X3Ds have it’s own problems and some of the results obtained are lower than expected, the performance of this systems has been quite impressive if we have in mind how small and efficient they are. This has been a first approach to X3Ds and there seems to be room for improvement, with further development these models could give very promising results. 23 References [1] Helbiz partners with drover ai to bring artificial intelligence to scooter sharing. URL: https://finance.yahoo.com/news/ helbiz-partners-drover-ai-bring-123000583.html. [2] Model zoo model’s benchmark. URL: https://pytorchvideo.readthedocs.io/ en/latest/model_zoo.html. [3] Pytorchvideo. URL: https://pytorchvideo.org/. [4] Taha Anwar. Introduction to video classification and human activity recognition. 2021. URL: https://learnopencv.com/ introduction-to-video-classification-and-human-activity-recognition/. [5] Various authors. Advanced driver-assistance systems. URL: https://en. wikipedia.org/wiki/Advanced_driver-assistance_systems. [6] Carreira and Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. 2018. URL: https://export.arxiv.org/pdf/1705.07750. [7] IBM Cloud Education. What are convolutional neural networks? 2020. URL: https: //www.ibm.com/cloud/learn/convolutional-neural-networks. [8] IBM Cloud Education. What are recurrent convolutional neural networks? 2020. URL: https://www.ibm.com/cloud/learn/recurrent-neural-networks. [9] Amin Ullah et al. Action recognition in video sequences using deep bi-directional lstm with cnn features. 2018. URL: https://ieeexplore.ieee.org/abstract/ document/8121994. [10] Ashish Vaswani et al. Attention is all you need. 2019. URL: https://papers.nips. cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf. [11] Du Tran et al. Learning spatiotemporal features with 3d convolutional networks. 2015. URL: https://arxiv.org/pdf/1412.0767.pdf. [12] Du Tran et al. A closer look at spatiotemporal convolutions for action recognition. 2017. URL: https://openaccess.thecvf.com/content_cvpr_2018/papers/Tran_ A_Closer_Look_CVPR_2018_paper.pdf. [13] Haoqi Fan et Al. Multiscale vision transformers. 2021. URL: https://arxiv.org/ pdf/2104.11227v1.pdf. [14] Kopuklu et al. Resource efficient 3d convolutional neural networks. 2019. URL: https://arxiv.org/pdf/1904.02422v4.pdf. [15] Will Kay et al. The kinetics human action video dataset. URL: https://arxiv. org/abs/1705.06950. [16] Christoph Feichtenhofer. X3d: Expanding architectures for efficient video recognition. 2020. URL: hhttps://openaccess.thecvf.com/content_CVPR_2020/ 24 papers/Feichtenhofer_X3D_Expanding_Architectures_for_Efficient_Video_ Recognition_CVPR_2020_paper.pdf. [17] Gonzalo Garc´ıa. La tecnolog´ıa ayuda a evitar el mal uso de los patinetes el´ectricos castigando al infractor. URL: https://www.hibridosyelectricos.com/articulo/tecnologia/ tecnologia-ayuda-evitar-mal-uso-patinetes-electricos-castigando-infractor/ 20210721143640047144.html. [18] Julio Gonz´alez. Environment perception for micromobility applications. 2021. [19] Alex Moreno. Github repository: micromobility c3d. URL: https://github.com/ alexmopun/tfm__micromobility_3dcnn. [20] Esteve Valls. Micromobility safety applications using ai. 2021. [21] Various. Transfer learning. URL: https://en.wikipedia.org/wiki/Transfer_ learning. 25 EXPERIMENT 1 Ground Truth Res18 (fr=1) sidewalk other Res18 (1) Precision Recall Fscore Pred sidewalk 32 8 sidewalk 0,8000 0,8889 0,8421 other 4 91 other 0,9579 0,9192 0,9381 Time (ms) std avg. Fscore 8180 1369 0,8901 Ground Truth Res18 (fr=10) sidewalk other Res18 (10) Precision Recall Fscore Pred sidewalk 25 10 0,7143 0,6944 0,7042 other 11 89 other 0,8900 0,8990 0,8945 Time (ms) std avg. Fscore 2386 6479 0,7993 EXPERIMENT 2 Ground Truth Res18 (1) sidewalk other Res18 (1) Precision Recall Fscore Pred sidewalk 235 105 sidewalk 0,6912 0,6583 0,6743 other 122 736 other 0,8578 0,8751 0,8664 avg. Fscore 0,7704 Ground Truth Res18 (10) sidewalk other Res18 (10) Precision Recall Fscore Pred sidewalk 235 95 sidewalk 0,7121 0,6438 0,6763 other 130 746 other 0,8516 0,8870 0,8690 avg. Fscore 0,7726 Table 16: Simulation test: Results (ResNet) A.6 Results: Original Database 32 EXPERIMENT 1 (original database) X3D-M sidewalk other Recall Precision Fscore sidewalk 220 27 sidewalk 0,8907 0,9865 0,9362 other 3 193 other 0,9847 0,8773 0,9279 Time (ms) std avg. Fscore 703 230 0,9320 X3D-S sidewalk other Recall Precision Fscore sidewalk 240 7 sidewalk 0,9717 0,9877 0,9796 other 3 193 other 0,9847 0,9650 0,9747 Time (ms) std avg. Fscore 653 230 0,9772 X3D-XS sidewalk other Recall Precision Fscore sidewalk 214 33 sidewalk 0,8664 0,9862 0,9224 other 3 193 other 0,9847 0,8540 0,9147 Time (ms) std avg. Fscore 570 206 0,9186 EXPERIMENT 2 (original database) X3D-M sidewalk other Recall Precision Fscore sidewalk 230 40 sidewalk 0,8519 0,9871 0,9145 other 3 193 other 0,9847 0,8283 0,8998 avg. Fscore 0,9071 X3D-S sidewalk other Recall Precision Fscore sidewalk 259 11 sidewalk 0,9593 0,9885 0,9737 other 3 193 other 0,9847 0,9461 0,9650 avg. Fscore 0,9693 X3D-XS sidewalk other Recall Precision Fscore sidewalk 229 41 sidewalk 0,8481 0,9871 0,9124 other 3 193 other 0,9847 0,8248 0,8977 avg. Fscore 0,9050 Table 17: Results original database (X3D) 33