Full text
A novel urban mobility classification approach based on convolutional neural networks and mobility-to-image encoding Peppino Fazio a,b, ⇑ , Miralem Mehic b,c , Miroslav Voznak b a DSMN, Ca’ Foscari University of Venice, Via Torino 155, 30172 Mestre, VE, Italy b VSB – Technical University of Ostrava, 17. listopadu 2172/15, 70800 Ostrava, Czechia c Department of Telecommunications, Faculty of Electrical Engineering, University of Sarajevo, Zmaja od Bosne bb, 71000 Sarajevo, Bosnia and Herzegovina article info Article history: Received 21 November 2022 Revised 22 March 2023 Accepted 13 April 2023 Available online 27 April 2023 Keywords: Convolutional neural networks Data-2-image conversion Machine learning Mobility classification Pattern prediction abstract Over the last few decades, the classification and prediction of mobility trajectories in dynamic networks have become major research topics. Switching of mobility areas (hand-over) in modern cellular networks is frequent due to restricted coverage area and node speeds (urban, highway, etc.). Accurate management of hand-over events is highly desirable to improve the system’s quality of service. We have exploited the high accuracy of machine learning to classify user mobility from mobility traces which we encoded into images. The method delivers high performance in mobility classification/prediction (exceeding 95%) and avoids the need to study and implement a dedicated neural network structure. The technique requires the conversion of mobility traces into image structures and the subsequent application of a convolutional neural network. We propose a novel approach to classifying mobility that involves data-to-image encoding and machine learning for image classification. Numerous simulations were performed to demonstrate the benefits of the proposed technique and to illustrate the variance in the accuracy of the functions of many encoding/classification parameters. The work represents a first preliminary step towards a new mobility prediction approach. We demonstrate that it is possible to achieve a very high level of prediction accuracy with low computational complexity, exploiting the strength of neural networks in image recognition. Ó2023 The Author(s). Published by Elsevier B.V. on behalf of King Saud University. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). 1. Introduction The advent of 5G technology and related studies concerning its potential latency and bandwidth performance (Liu et al., 2020)in mobile networks have demonstrated that it is possible to attain a good level of Quality of Service (QoS), especially if predictive approaches are integrated into the system (Fazio et al., 2017; Fazio et al., 2023) after correctly modeling traffic flows (Kromer et al., 2020; Kromer et al., 2018; De Rango et al., 2005). Machine and deep learning applications (Martin et al., 2021; Singh et al., 2021; Xu et al., 2021) have also steeply proliferated in the last few years and added enormous value to the software which can exploit them. In the current work, we especially highlight the possibility of using well-known Convolutional Neural Networks (CNNs) (Khan et al., 2018) to classify mobility after transposing the mobility data into suitable mobility images (instead of realworld images). Our work does not propose a new neural network layer scheme or image classification method. Still, we demonstrate the strength of CNNs (Luo et al., 2018) in their accuracy and how they can be applied to classifying/predicting mobility. Neuralbased image classification algorithms are known to achieve an accuracy of 95–98 %whereas mobility classification or prediction schemes such as those described in Jin et al. (2001), Gaiduchenko and Gritsyk (2019),Fazio et al. (2017), Zhang et al. (2018) can obtain, to the best of our knowledge, a maximum accuracy of about 85–87 %, which is well below 90 %.Fig. 1 shows a generic reference scenario: mobile hosts are free to move in any geographical area, each one covered by a 5G micro/femto-cell. For example, for the black path, a mobility predictive approach can be integrated with 5G architecture, in order to reserve resources (bandwidth channels) in-advance, avoiding service disruptions for the mobile host (darker cells). The encoding can be made locally, on-the-fly, https://doi.org/10.1016/j.jksuci.2023.101561 1319-1578/Ó2023 The Author(s). Published by Elsevier B.V. on behalf of King Saud University. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). ⇑ Corresponding author. E-mail addresses: [email protected] (P. Fazio), [email protected] (M. Mehic), [email protected] (M. Voznak). Peer review under responsibility of King Saud University. Production and hosting by Elsevier Journal of King Saud University – Computer and Information Sciences 35 (2023) 101561 Contents lists available at ScienceDirect Journal of King Saud University – Computer and Information Sciences journal homepage: www.sciencedirect.com
or remotely, by a dedicated server, while the remote CNN receives the images and the related belonging cell, having the possibility to be trained and, then, validated. Starting from this observation, a new method for converting mobility-data into image-data is proposed; in particular, the main contributions of the current work can be summarised as: A proposed innovative method for encoding mobility data into images, taking into account the entropy metric (Frank and Frank, 2020), to assist the CNN in reaching a better validation accuracy; No need for the design and implementation of a new neural network: the most straightforward CNN structures are taken into account (we do not design a new layering structure, we demonstrate how it can perform in its simplest configuration); Evaluation of CNN performance in the function of encoding parameters such as sampling frequency, encoding algorithm, image size, etc.; Base station capability of recognizing the geographical area where a mobile host is moving without the aid of GPS or other positioning devices. Mobile nodes and base stations can apply Angle-of-Arrival (AoA), Proximity, Trilateration, or Time Difference of Arrival (TDoA) methods. Our encoding scheme functions with normalized mobility data so the trained CNN recognizes the mobility region without accurate GPS data. The remainder of the paper is structured as follows: Section 2 introduces some recent works which examine mobility classification and prediction; Section 3describes the main proposed algorithm, specifying sufficient detail regarding the encoding approach; Section 4describes in detail the main results obtained through the proposed approach; Section 5concludes the paper, remarking on the main advantages of the proposed algorithm. 2. Related work This section reviews the main literature contributions concerning mobility classification, prediction, and image recognition. We examined the existing literature and found that no studies have integrated these features and exploited the strength of imageclassification CNNs or suggested methods for encoding mobility traces in images. We thoroughly review the most significant works related to our proposal. In particular, vehicle trajectory is considered as a time series (coordinates, speeds, and vehicle parameters are time-dependent). The great performance of CNNs (Karim et al., 2019) is exploited for multivariate time series classification: the authors introduced multiple convolutional layers for features extraction, showing the classification accuracy for the proposed method in the function of the dataset noise and obtaining an accuracy that ranges from 30%and 100%. The authors of Asad et al. (2020) focused on travelers profiling in train stations to analyze how humans behave according to their age (two age classes are considered: 16–59 and 60-and-over) during the Covid-19 disease pandemic. In particular, the authors employed six different classifiers (Logistic Regression, Multi-layer Perceptron, Support Vector Machine, Random Forest, K-nearest Neighbour, and Decision Tree) to establish automation in intelligent decisions while learning from history and adapting to the testing environment. Given the described dataset (London Underground and Overground - LUO), the proposed idea can monitor potential contacts/proximity travelers and advise them to safeguard vulnerable age-group travelers. Simulation results reached an accuracy of about 82%and 86%for the two age classes. The importance of mobility prediction is stated by several literature contributions, such as the ones in Fazio et al. (2017), Zhang et al. (2018), where the authors survey the main contributions, also in terms of different methodologies and approaches, in the world of mobility prediction. It has been studied for decades, and there is a wide variety of predictive models (Kalman filters, Markov chains, Neural Networks, Auto Regressive, etc.). In the mentioned works, the authors compare the different approaches, showing the potentialities of each predictor. It is also underlined that the reached accuracy is generally below 90%. The work in Wang et al. (2021) proposes an innovative attentional Markov model, considering the long-term correlation with historical trajectories and context information and predicting future host positions. The authors considered an extensive dataset (more than 20000 users), obtaining an improvement of the classical machine learning approaches (such as the Hidden Markov Model - HMM, Recurrent Neural Network - RNN, Friendship, and Mobility - MF, etc.), with a faster execution time. For many years, machine learning and neural networks have been widely used in image recognition. The convolutional layer of CNNs can discover some hidden features of the input images while the subsequent layers produce the correct output. The work in Chen et al. (2019) extends the well-known LeCun’s CNN (Lecun et al., 1998), adding convolutional and pooling layers and proposing a Multi-Convolution NN (MCNN). The authors tested the new CNN on the Cats vs. Dogs (Dogs vs Cats dataset, 2023), Cifar-10 (Krizhevsky, 2023), and Fer2013 (Facial Expression Recognition dataset, 2013) datasets, showing that the proposed MCNN outperforms the classical one in terms of complex-texture features recognition. The authors of Tiwari et al. (2020) proposed the Visual Geometry Group 16 (VGG16) model to classify images into two additional categories instead of performing feature extraction or segmentation. VGG16 offers an accuracy of 99%, and images are further categorized into additional sub-categories. The paper (Xu et al., 2020) is related to the compressed-domain image classification: images are considered before encoding in the desired format (the reconstruction step is bypassed), and the training is made with a dynamic Measurement Rate (MR) by selecting the needed MR with the help of a sensing matrix. The authors tested their proposal on a massive set of datasets (e.g. Cifar-10 (Krizhevsky, 2023), and Coil-100 (Nene et al., 2023)), and the performance has been very satisfactory, also in terms of noise robustness. Taking into account the main results discussed above, we will illustrate, in the next section, our proposal, to design a new mobility classification algorithm based on CNNs and mobility-to-image encoding. 3. Mobility classification through CNNs imaging This section is dedicated to our main proposal. First of all, we will introduce the main issue. Then we will give an in-depth overview of mobility encoding and the type of used CNN. We are not proposing a new CNN structure but a new approach to mobility classification, which can exploit the high accuracy of CNNs in image classification. In addition, we are not considering any particular kind of node (a vehicle, a human, a mobile sensor, an Unmanned Aerial Vehicle, etc.), so the approach is completely general (simulation results will be specialized for a particular case). Fig. 2 illustrates the main steps of our proposed idea through Data-Flow Diagram structures. As shown in figure (a), mobility traces are generated for all the considered geographical areas based on real maps and real traffic behaviors; then, we propose two possible algorithms for converting the created data into matrices and, in the end, into images. In this way, each path of a mobile user is stored as a set of images, so the users moving into the same area create images belonging to the same class. Then (b), a classical CNN is trained based on the previously generated set of images; we carried out several simulations to avoid over-fitting and find the P. Fazio, M. Mehic and M. Voznak Journal of King Saud University – Computer and Information Sciences 35 (2023) 101561 2
best trade-off between the subset of the images for training and the subset of the images for validation. Once the CNN has been trained (c), each new mobile trace can be classified adequately, giving the CNN a short portion of the trace as input. All the implementation details are given in the following subsections and as simulation results. 3.1. The general model for mobility classification We are considering a general set of Mobility Areas (e.g., a square area on the Earth’s surface, a portion of the sky, a volume in the sea/ocean, etc.) MAs ¼ma 1 ;ma 2 ;...;ma n fg , with jMAsj¼n. Each ma i 2MAs, can be adjacent to another ma j 2MAs, with i–jand i;j¼1;...;n, or it can be located in a completely different part of the Earth/sky. Let us indicate with Vthe set of mobile nodes which are moving into MAs: we assume that V¼ v 1 ;...; v m fg and jVj¼m. So, our system has nmobility areas and mmobile nodes. Each v k 2V moves into only one ma i 2MAs, so we can use the notation v k;i to indicate the kth mobile node is moving in ma i 2MAs. Under these assumptions, we cannot have v k;i and v k;j at the same time. Without loss of generality, we assume that each mobile node moves in a 3D environment, so each ma i 2MAs is characterized by a center c i ¼cx i ;cy i ;cz i ðÞand an extension radius r i , that is ma i ¼c i ;r i ðÞ, with cx i ;cy i ;cz i 2R. So, each ma i 2MAs can be represented as a sphere (or a circle in 2D mobility). Our model is valid if a cube or a square is considered. Each node v k;i , by its mobility in ma i , creates a pattern P k ¼ v x t k ; v y t k ; v z t k , where t¼1;2;...is the discrete time index, with t l t l1 ðÞ¼T, the sampling period (the period at which mobility position is sampled and stored). We can refer to P k also by the notation P k ¼ v x 1 k ; v y 1 k ; v z 1 k ; v x 2 k ; v y 2 k ; v z 2 k ;...; v x t k ; v y t k ; v z t k ;...g, given that it is composed by a sequence of triplets (in the general case of 3D space). An ordered subset of P k ;p t 0 k , of length pis defined as p t 0 k #P k : Fig. 1. An example of 5G cellular coverage and users trajectories encoded into images, able to train a remote server and its CNN. P. Fazio, M. Mehic and M. Voznak Journal of King Saud University – Computer and Information Sciences 35 (2023) 101561 3
p t 0 k ¼x 1 ;y 1 ;z 1 ðÞ;x 2 ;y 2 ;z 2 ðÞ;...;x p ;y p ;z p if 9t 0 j... x 1 ;y 1 ;z 1 ðÞ ¼ v x t 0 k ; v y t 0 k ; v z t 0 k ;... x 2 ;y 2 ;z 2 ðÞ¼ v x t 0 þ1 k ; v y t 0 þ1 k ; v z t 0 þ1 k ;... x p ;y p ;z p ¼ v x t 0 þp k ; v y t 0 þp k ; v z t 0 þp k : ð1Þ So, under the assumptions above, if we refer to v k;i , it means that: ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi cx i v x t k 2 þcy i v y t k 2 þcz i v z t k 2 hi r6r i 8t2N:ð2Þ Eq. 2indicates that each point of the mobility pattern of v k is bounded to the extension area (or volume) of ma i . The goal of our entire article is: given an ordered subset p t 0 k (extracted from P k ), a CNN should classify it by returning the ma i in which node v k is moving. To reach this aim, first of all, we propose a mobility-to-image encoding scheme, and then, the CNNs are trained and validated by the obtained encoded datasets. 3.2. Mobility to matrix encoding: MME, an alternative to raw data representation Based on the definitions above, we will propose a new way to encode mobility data into lossless images (basically, a matrix is obtained, then stored as an image through a proper lossless codec). Let us start with a mobility pattern subset p t 0 k ¼ v x t k ; v y t k ; v z t k jt2t 0 ;t 0 þp ½ . As earlier defined, it is a set of triplets x;y;zðÞ(in the general case of a 3D mobility). We can evaluate the related speed s t 0 k of node v k during the time range t 0 ;t 0 þp½as: s t 0 k ¼ 1 T v x t 0 þlþ1 k v x t 0 þl k ; v y t 0 þlþ1 k v y t 0 þl k ;... v z t 0 þlþ1 k v z t 0 þl k no ; ð3Þ with l¼0;1;...;p1, and sx t 0 k ¼vx t 0 þlþ1 k vx t 0 þl k ; sy t 0 k ¼vy t 0 þlþ1 k vy t 0 þl k ;sz t 0 k ¼vz t 0 þlþ1 k vz t 0 þl k . The length of s t 0 k will be p1. We can also evaluate the acceleration/deceleration a t 0 k of node v k during the time range t 0 ;t 0 þp½as: a t 0 k ¼ 1 T sx t 0 þqþ1 k sx t 0 þq k ;sy t 0 þqþ1 k sy t 0 þq k ;...sz t 0 þqþ1 k sz t 0 þq k no ;ð4Þ with q¼0;1;...;p2, and ax t 0 k ¼sx t 0 þqþ1 k sx t 0 þq k ;ay t 0 k ¼ sy t 0 þqþ1 k sy t 0 þq k Þ;az t 0 k ¼sz t 0 þqþ1 k sz t 0 þq k . The length of a t 0 k will be p2. In the following, we will use the shorter notations for p t 0 k ;s t 0 k and a t 0 k . In the case of Raw Mobility Data (RMD) representation, the information is converted from a vector to a matrix format. Three matrices are obtained for p t 0 k ;s t 0 k and a t 0 k . Without loss of generality, we assume that, for RMD, pis a square value, so it is easy to define r¼ffiffiffi p p. For speed and acceleration vectors, the missing elements (1 and 2, respectively) can be padded during the matrix construction (we assume that, for example, the values are replaced by zeros). Matrix dimensions are rx3r, because rcolumns are dedicated to each coordinate. Fig. 2. The Data-Flow Diagrams of the three steps provided in our proposal. P. Fazio, M. Mehic and M. Voznak Journal of King Saud University – Computer and Information Sciences 35 (2023) 101561 4
Algorithm 1 represents just a data-structure transformation with a computational complexity hr 2 . It is applied to position, speed, and acceleration vectors, accepting as input the vector v ec t 0 k containing position, speed, or acceleration mobility information (specified in the input parameter type) related to mobile host v k , that is one of the terms in Eq. 5. Firstly, the algorithm evaluates the right matrix dimension r. Then it executes two nested For cycles to fill up the rows and columns of the output matrix M t 0 k , which has been initially set to the null matrix and normalized in the range 0;1½at the end. Inside the nested loops, the notation v ec t 0 k c;ind v ðÞ, with c= 1,2,3 refers to the x;yand zcomponents of the v ec t 0 k cðÞtriplet. Let us see how the raw data can be effectively encoded while putting it into a matrix structure. We can start to build three separate matrices from the three vectors. A spreading factor sf is defined, and each matrix will be square: setting r¼psfðÞ, then the dimensions will be rxr. All the elements of p t 0 k ;s t 0 k and a t 0 k are initially re-scaled (rs) between 0 and 1 (they can also contain negative values), so we can write: p t 0 k ¼rs p t 0 k ;s t 0 k ¼rs s t 0 k ;a t 0 k ¼rs a t 0 k :ð5Þ The rs operation divides all the elements of a vector by the maximum one if there are no negative values; otherwise, the absolute value of the lowest negative element is added to each value before selecting the maximum and normalizing the elements. Then the structure of the matrices is defined as follows. First of all, the rxr matrices Mp t 0 k ;Ms t 0 k and Ma t 0 k are created as empty matrices, that is, each element is equal to zero. P. Fazio, M. Mehic and M. Voznak Journal of King Saud University – Computer and Information Sciences 35 (2023) 101561 5
For all of them, the xcomponent is encoded as a sequence of columns from top to bottom, the ycomponent is encoded as columns from the bottom to the top, and the zcomponent is encoded as rows, from left to right. The idea is to associate a row or a column of ”ones” which is proportional to the content of p t 0 k ;s t 0 k ;a t 0 k . We underline that our proposal is only one of the possible ways to fill the matrices: However, by knowing which kind of results can be obtained, one can find different ways to create images. So, the Algorithm 2 specifies how the three matrices are filled up; we illustrate the algorithm for the generic matrix M. Both RMD and MME algorithms are executed for: Position p t 0 k :Mp t 0 k ¼MME p t 0 k ;sf;p; ’position 0 Þ;or Mp t 0 k ¼ RMD p t 0 k ; 0 position 0 ; Speed s t 0 k :Ms t 0 k ¼MME s t 0 k ;sf;p; ’speed 0 Þ;or Ms t 0 k ¼RMD s t 0 k ; ’speed 0 Þ; Acceleration a t 0 k :Ma t 0 k ¼MME a t 0 k ;sf;p; ’accel 0 Þ;or Ma t 0 k ¼RMD a t 0 k ; ’accel 0 Þ; Referring to Algorithm 2, it accepts as input the normalized vector v ec t 0 k containing position, speed or acceleration mobility information (specified in the input parameter type) related to mobile P. Fazio, M. Mehic and M. Voznak Journal of King Saud University – Computer and Information Sciences 35 (2023) 101561 6
host v k , that is one of the terms in Eq. 5; the length of the subset p and the spreading factor sf are also given as input. The values of sf indicate the number of columns assigned to each component (at least it should be set to 3). Firstly, the MME evaluates the matrix dimension ras defined earlier and creates the empty matrix M t 0 k . Then it computes the right dimension of the input vector: from the definitions of Eqs. (3) and (4), we know that for speed and acceleration vectors, the lengths are p1 and p2 respectively. At this point the algorithm executes the first cycle (hvariable), considering for the xcomponents the columns (index x )1,1+sf, 1+2 sf , and so on, for the y components the columns (index y ), 3, 3 + sf, 3+2 sf , and so on, while for the zcomponents the rows (index z )1,1+sf, 1+2 sf, and so on. At this point, each row/column is filled with several ones, which is proportional to the content of the elements of v ec t 0 k c;hðÞ. For the z component, during the third internal cycle, if the generic matrix element M t 0 k index z ;tðÞhas already been modified by the two previous cycles (for the xand ycomponents), no further actions are made. Recalling that the elements of v ec t 0 k are normalized from 0 to 1, then the term b v ec t 0 k c;hðÞrewill be bounded to the maximum size of M t 0 k . The symbol be indicates the rounding operation (the nearest integer is chosen). Computational complexity is acceptable because the initial operations have constant complexity (negligible), the primary cycle is executed OpðÞtimes. In contrast, the internal cycles are executed Or ðÞtimes, so we can write that the overall computational complexity is Op3rðÞ=Op3psfðÞ which is bounded by Op 3 and pdoes not depend on the length of the overall pattern. At this point, it is useful to make an example of how the matrices are created from real values with MME (RMD is more intuitive since it is only a conversion between vectors and matrices). Let pbe equal to 4, sf be equal to 4 and p t0 k be f12:735;87:675;10:22ðÞ;18:143;79:884;ð11:31Þ;25:750;71:943;ð 12:7Þ;30:564;63:345;11:8ðÞg, with jp t 0 k j¼p¼4. By applying Eqs. (3) and (4) then we will have s t 0 k = f5:408;7:791;1:09ðÞ;7:607;7:941;1:39ðÞ;4:814;8:598;0:9ðÞ;g, with js t 0 k j¼p1¼3 and a t 0 k =f2:199;0:15;0:3 ðÞ ; 2:793;ð0:657;2:29Þg, with ja t 0 k j¼p2¼2. At this point, by the input value of sf, we will have r=4 4¼16, so the three matrices will have dimensions of 16 x 16. The normalized vectors will be: p t0 k :f0:145;1;0:116ðÞ;0:207;0:911;0:129ðÞ;0:294;0:820;ð 0:145Þ;0:349;0:722;0:134ðÞg; s t 0 k =f0:864;0:04978;0:598ðÞÞ;1;0:040;0:616ðÞ;0:827;0;0:475ðÞg; a t 0 k =f1;0:529;0:619ðÞ;0;0:427;0:101ðÞg. The MME is applied to each normalized vector at this point, and the matrices illustrated in Fig. 3 are obtained. From Fig. 3, it can be seen how the xcomponents are added by the unitary elements from top to bottom (red color) in the columns indexed as in the MME (indexes 1, 5, 9, 13), the unitary elements from the bottom add the ycomponents to the top (green color) in the columns indexed as in the MME (3, 7, 11, 15) and the z Fig. 3. An example of the three matrices as output of MME. P. Fazio, M. Mehic and M. Voznak Journal of King Saud University – Computer and Information Sciences 35 (2023) 101561 7
components are added by the unitary elements from left to right (dark blue color) in the rows indexed as in the MME (1, 5, 9, 13). Orange values represent the presence of more components in the same matrix elements: e.g., the element Mp t 0 k 1;1ðÞis generated by both the xcomponent and zcomponent. The following (and last) step is to encode the matrices into an image: to this aim, we associate the Mp t 0 k matrix to the RED image channel, the Ms t 0 k matrix to the GREEN image channel and the Ma t 0 k matrix to the BLUE image channel, after multiplying each matrix by 2 8 -1 (we are considering 8-bits resolution per layer), independently from RMD or MME. We did not care about investigating the particular image format. Still, we focused on a lossless encoding, such as the Portable Network Graphics (PNG) (The Portable Network Graphics specification, 2023), to avoid the creation of artifacts (JPG format, for example, is not suitable for our approach). In particular, we chose the RGB 24bit format for PNG, so each element in each matrix is represented by 8 bits (256 levels): a detail of the three matrices of Fig. 3 in a visual representation is given in Fig. 4 (please, note that the images have been magnified, because the width of each row/column is 1 pixel). 3.3. The image entropy: A possible evaluation metric It is known from the literature that the entropy concept plays a vital role in image classification (Gowdra et al., 2020). The classical metric used to evaluate the goodness of a slicing recognition and classification is the Maximum Entropy (ME): the majority of classification/recognition machine learning approaches try to divide big images into smaller ones, finding the slices which are characterized by the highest ME (i.e., the maximum level of information, the most visually representative portions of the image), which can be used for training the CNN, whose convolution operations can discover and extract the hidden information. Entropy, independently from Hartley’s (Hartley, 1928) or Shannon’s (Shannon, 1953) definitions, represents the amount of information contained in the image to be classified and the degree of randomness of the pixel values. It is intended that the more information is included in the image, the higher will be both the entropy and, also, the structural complexity of the CNN: generally, more than one convolution layer is needed to extract the right features from data. In our experiments, we evaluate the color image entropy as follows (assuming that ris the number of image pixels and 8-bit values represent each color channel): EIm½¼ 1 3E c Im RED ðÞþE c Im GREEN ðÞþE c Im BLUE ðÞ½;ð6Þ that is the average of the entropy of each color channel E c defined as: E c Im ch ½ ¼X 2 8 1 pix ch ¼0 pIm ch ;pix ch ðÞ log 2 pIm ch ;pix ch ðÞ½ ;ð7Þ with: pIm ch ;pix ch ðÞ¼ count Im ch ;pix ch ðÞ rð8Þ with the function count In ch ;pix ch ðÞcounting the number of times a pixel in Im ch assumes the value pix ch . The focus of our work is not related to the proposal of a new CNN structure but only to show that good results can be achieved in mobility classification with a novel approach. For these reasons, we are proposing a new MME algorithm to maintain the CNN complexity as low as possible, so images will be characterized by relatively low entropy values. As stated in the previous sub-section, the proposed MME is just one of the possible algorithms of data-to-image encoding. In this work, we show, for the first time, that this approach is suitable for mobility classification. The easiest way to encode mobility into an image is to put the raw data directly into a matrix and then encode it as an image. But this last approach has several drawbacks: Very tiny images will be generated: every CNN, as shown in the numerical results section, needs a minimum image size to recognize the image and classify it; if we do not think about an efficient encoding scheme, a huge number of samples is needed. Consider the previous example: we had p= 4, which means we have only four triplets (4 values for each coordinate component). It also means that we obtain three values of speed and two values of acceleration: we have nine different data for each component, for a total of 27 matrix elements. With an encoding scheme (as MME), we have 3 r 2 =768 matrix elements, starting from 27 data points. So, a proper representation of the data is mandatory to build an adequate training set; From an entropy point of view: if we represent the raw data directly into an image, the entropy value will be high due to the enormous color changes with high granularity (especially if mobility is represented by lat and lon couples); the MME, instead, as demonstrated later, will reduce the entropy sensibly. For example, we show the difference between RMD and MME encoding in terms of entropy values distribution. Fig. 5 shows an example of the results obtained by encoding 2D mobility data by MME or just leaving the raw values into a single RGB matrix (RMD). RMD and MME images contain 8192 and 7744 pixels, respectively (comparable sizes). Still, the RMD image is created by exactly 4096 samples (mobility points, which are doubled for xand ycomponents), while MME needs only 484 mobility points (with a gain of about 88.1%). We used 2500 2D mobility pattern subsets processed with RMD and MME to see an example of how the entropy is distributed. The subset length phas been set to 22; Fig. 6 shows the massive difference in terms of entropy between RMD and MME encoded images: for MME, the mean entropy is 1.22 J/K, while for RMD it is 4.26 J/K. In the case of MME it is almost constant (due to the presence of high black spaces), while for RMD it has a substantial oscillatory trend. It is also interesting to see how the entropy is distributed: Fig. 7 shows the pdf(E) trend. It confirms what was observed in Fig. 6, a very low deviation around the mean value for MME, while more spread values for RMD. In the numerical results section, we will show what happens to Entropy based on the parameters related to image encoding. In the next section, instead, the CNN model is deeply discussed. Fig. 4. A visual representation of Mp t 0 k ;Ms t 0 k ;Ma t 0 k (top) and the related RGB layers (bottom). P. Fazio, M. Mehic and M. Voznak Journal of King Saud University – Computer and Information Sciences 35 (2023) 101561 8
3.4. The considered convolutional neural network Our work shows which performance can offer a CNN-based image classification model if trained on mobility data (previously encoded into images). So, our aim does not regard the proposal of a new CNN model but a new way of using it in a different research context. So, in this sub-section, we are giving just the main description of the used LeCun-like (Lecun et al., 1998) CNN model. Using its Machine Learning Toolbox, we used MATLAB (MATLAB, 2021) for our implementation. In particular, we structured the CNN, following LeCun’s proposal, as follows: Input layer: it is designed for accepting an image as input; it is created by defining the size of the input image and the number of its layers (3 in our case: red, green, and blu); by default, in MATLAB, the image input layer normalizes pixel values by substracting their mean value; we can specify pixel values to be normalized between 0 and 1; Convolution layer: it applies sliding convolutional filters to the 2D input. The input is convolved by moving the specified filters along the 2D input (in horizontal and vertical directions) and evaluating the dot product of the weights and the input; Batch Normalization layer: it normalizes a data batch across all observations. It is generally used to speed up the training of the CNN, reducing the sensitivity to the beginning uncertainty; Rectified Linear Unit (ReLU) layer: it is one of the possible activation layers and applies a threshold operation to its input (in simple words, negative values are set to 0); Fully Connected layer: it is a dense layer of neurons and multiplies the input by a weight matrix; Soft Max layer: it applies a softmax function to the input. Softmax is a generalization of the logistic function, which can compress a k-length and arbitrary content vector to another klength vector, with the sum of elements equal to 1; Classification layer: it can receive an input and return the classified output (in our case, the ma i 2MAs). Fig. 8 shows the complete layering of the considered CNNs in the MATLAB window, with the main used parameters. Results can be, of course, enhanced. Still, in this paper, we are interested neither in investigating the structure of the CNN nor in increasing its complexity (on the contrary, one of our goals is to maintain the CNN structure as simple as possible). For more details about the possible CNN optimization, please refer to Bishop (2006). Fig. 7. Probability Density Function for the entropy values of Fig. 7. Fig. 6. Entropy trend for RMD and MME images, with p= 22 and T= 1s. Fig. 5. True images obtained by using Raw Mobility Data (RMD), size 64 128, or the MME algorithm, size 88 88. Fig. 8. The layering structure of the simple CNN used in our work. P. Fazio, M. Mehic and M. Voznak Journal of King Saud University – Computer and Information Sciences 35 (2023) 101561 9