scieee AI-readable full text Open interactive document viewer

Machine learning and image processing

Martins, Cecília Eduarda Coelho Machado da Cruz

Abstract

Portuguese legislation states the compulsory reporting of the addition of amenities, such as swimming pools, to the Portuguese tax authority. The purpose is to update the property tax value, to be charged annually to the owner of each real estate. According to Technavio and Market- Watch, this decade will bring a global rise to the number of swimming pools due to certain factors such as: cost reduction, increasing health consciousness, and others. The need for inspections to ensure that all new constructions are communicated to the competent authorities is therefore rapidly increasing and new solutions are needed to address this problem. Typically, supervision is done by sending human resources to the field, involving huge time and resource consumption, and preventing the catalogue from updating at a rate close to the speed of construction. Automation is rapidly becoming an absolute requirement to improve task efficiency and affordability. Recently, Deep Learning algorithms have shown incredible performance results when used for object detection tasks. Based on the above, the objective of this thesis is to study the various existing object detection algorithms and implement a Deep Learning model capable of recognising swimming pools from satellite images. To achieve the best results for this specific task, the RetinaNet algorithm was chosen. To provide a smooth user experience with the developed model, a simple graphical user interface was also created.

Full text

Cecília Eduarda Coelho Machado da Cruz Martins Machine Learning and Image Processing julho de 2020 UMinho | 2020 Cecília Eduarda Coelho Machine Learning and Image Processing Universidade do Minho Escola de Ciências Cecília Eduarda Coelho Machado da Cruz Martins Machine Learning and Image Processing Dissertação de Mestrado em Matemática e Computação Trabalho efetuado sob a orientação do Professor Doutor Luís Jorge Lima Ferrás e da Professora Doutora Maria Fernanda Pires da Costa Universidade do Minho Escola de Ciências julho de 2020 Direitos de Autor e Condições de Utilização do Trabalho por Terceiros Este é um trabalho académico que pode ser utilizado por terceiros desde que respeitadas as regras e boas práticas internacionalmente aceites, no que concerne aos direitos de autor e direitos conexos. Assim, o presente trabalho pode ser utilizado nos termos previstos na licença abaixo indicada. Caso o utilizador necessite de permissão para poder fazer um uso do trabalho em condições não previstas no licenciamento indicado, deverá contactar o autor, através do RepositóriUM da Universidade do Minho. Atribuição-NãoComercial-CompartilhaIgual CC BY-NC-SA https://creativecommons.org/licenses/by-nc-sa/4.0/ ii Acknowledgements First of all, I would like to acknowledge my supervisors, Professor Doctor Luís and Professor Doctor Fernanda for all their availability, concern and kindness during this process and for the flexibility given, by supporting and approving the concepts I wanted to use to achieve the goal of this thesis, after a lengthy session of questions. I am also grateful for the help of Professor Doctor Ana Jacinta for teaming up with my supervisors and contributing in the same terms. Furthermore, I want to express my gratitude to Accenture for the experience and for suggesting to address the detection of swimming pools in satellite images based on the theme of this thesis. Also, I want to thank my two supervisors provided by Accenture, Cristiana and Guilherme. Although they were busy they never forgot to check on me and their advice and opinions can be found scattered through all this work. To Diana and Luís for accompanying me in this journey. Thank you for your help, patience and motivation. No day would pass without some good laughter. To Ozzy, the pseudo engineer and future physicist that recently discovered the wonderful world of Machine Learning. Could not forget Fernando and João for all the hardships we overcame together in this master’s degree and for being the best work group I have ever encountered. Finally, I want to express my greatest appreciation for my mother and for all the people that have always supported me in all my choices, if I could go back in time I wouldn’t change a thing. To my grandfather and to the child that dreamt about working with cutting edge technology. iii Statement of Integrity I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the University of Minho. v Resumo Machine Learning and Image Processing A legislação Portuguesa declara a obrigatoriedade da comunicação de novas construções, como piscinas, à Autoridade Tributária e Aduaneira. Esta comunição permite o ajustamento do Imposto Municipal sobre Imóveis a pagar anualmente pelo proprietário. De acordo com o Technavio e o MarketWatch, irá ocorrer um aumento significativo do número de piscinas devido a vários fatores como a redução do custo da construção, o aumento da consciência para a adoção de um estilo de vida saudável, entre outros. Isto leva à necessidade de um reforço na inspeção de forma a garantir que todas as novas construções foram devidamente comunicadas à autoridade competente. Atualmente, estas inspeções são realizadas com a distribuição de recursos humanos pelo terreno, o que tráz um elevado custo operacional e temporal, impedindo uma catalogação a uma taxa próxima da de construção. Hoje em dia, a automatação de tarefas está a tornar-se muito requisitada devido a permitir o aumento da eficiência e a redução de custos. Recentemente, os algoritmos de Deep Learning tem demonstrado resultados incriveis quando usados para deteção de objetos. O objetivo desta dissertação é o estudo dos vários algoritmos de deteção de objetos existentes e a implementação de um modelo de Deep Learning capaz de detetar piscinas em imagens satélite. De forma a obter os melhores resultados na tarefa em questão, o algoritmo RetinaNet foi usado. Além disso e com o intuito de melhorar a experiência na utilização do modelo desenvolvido, foi construída uma interface gráfica simples. Palavras-chave: Visão por Computador, Deep Learning, Deteção de Objetos, ResNet,RetinaNet 2.22 Detected regions of interest. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 2.23 The reshaped regions with a RoI pooling layer are passed into a fully connected network........................................... 25 2.24 Bounding boxes reshaped and classified. . . . . . . . . . . . . . . . . . . . . . . . . 26 2.25Imagegivenasinput[4].................................. 26 2.26 The input is given to convolutional layers that produce a feature map. . . . . . . . 27 2.27 RoI are detected from the feature maps. . . . . . . . . . . . . . . . . . . . . . . . . 27 2.28 To create the input for a Fully-connected Network, the regions are reshaped using aRoIpoolinglayer. ................................... 28 2.29 Bounding boxes reshaped and classified. . . . . . . . . . . . . . . . . . . . . . . . . 28 2.30Inputimage[4]....................................... 29 2.31 Input image divided into an S×Sgrid. ........................ 29 2.32 Input image with randomly drawn bounding boxes. . . . . . . . . . . . . . . . . . . 30 2.33 Filtered bounding boxes according to the threshold value chosen by the user. . . . 30 2.34 Graph that shows the loss functions values for each prediction. The cross entropy represented by the blue line and the other full-lines are the focal loss function with aγvaryingfrom0.5to5[5]. .............................. 32 2.35 Illustration of the bottom-up and top-down pyramid along with a lateral connection scheme[6].......................................... 33 2.36 RetinaNet network architecture composed of four main components: a) bottom-up pathway; b) top-down pathway; c) classification subnetwork; d) regression subnetwork;[5]. ......................................... 34 3.1 Four training images of the chosen dataset [4]. . . . . . . . . . . . . . . . . . . . . . 40 5.1 Testing set image annotation using LabelImg [7]. ................... 53 5.2 Two testing images, containing swimming pools, of the chosen dataset used for evaluating the model’s performance [4]. . . . . . . . . . . . . . . . . . . . . . . . . 55 5.3 Two testing images without swimming pools, of the chosen dataset, used for evaluating the model’s performance [4]. . . . . . . . . . . . . . . . . . . . . . . . . . . . 55 5.4 Model’s predictions of two testing images of the dataset. . . . . . . . . . . . . . . . 56 5.5 Model’s predictions of two testing images without swimming pools. . . . . . . . . . 56 5.6 Model’s predictions of two testing images of the dataset using ResNet101 as backbonearchitecture. .................................... 57 5.7 Model’s predictions of two testing images without swimming pools. . . . . . . . . . 57 5.8 Model’s predictions of two testing images of the dataset using ResNet152 as backbonearchitecture. .................................... 58 5.9 Model’s predictions of two testing images without swimming pools. . . . . . . . . . 59 xiv 5.10 Eight test images taken from google maps with different sizes and resolutions. The images are numbered in yellow on the top right corner. . . . . . . . . . . . . . . . . 61 5.11 Eight test images taken from google maps with different sizes and resolutions. The detected objects are enclosed by blue boxes. The images are numbered on the top rightcorner......................................... 62 5.12 Eight test images taken from google maps with different sizes and resolutions. The detected objects are enclosed by blue boxes. The images are numbered on the top rightcorner......................................... 63 5.13 Eight test images taken from google maps with different sizes and resolutions. The detected objects are enclosed by blue boxes. The images are numbered on the top rightcorner......................................... 65 6.1 Desktop application main window. . . . . . . . . . . . . . . . . . . . . . . . . . . . 69 6.2 “SavePath”buttonaction................................. 70 6.3 “Upload Image” button action. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70 6.4 “Upload Folder” button action. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71 6.5 “Run”buttonaction.................................... 71 6.6 “Show”buttonaction. .................................. 72 6.7 Error prompt shown where the saving path and/or at least one image are not provided. 72 6.8 An error message prompts if the previous (left arrow) or next (right arrow) image is clicked but there are no more images with bounding boxes to show. . . . . . . . 73 B.1 Region Proposal Networks architecture [8]. . . . . . . . . . . . . . . . . . . . . . . 81 C.1 Adding more layers to a neural network in which the best possible state was attained results on the following layers learning the identity function. . . . . . . . . . . . . . 83 C.2 Training error (left) and test error (right) on CIFAR-10 [4] with 20-layer and 56layer “plain” networks. The deeper network has higher training error, and thus test error[9]........................................... 84 C.3 A residual block where the layer’s input is added to the output [9]. . . . . . . . . . 84 C.4 Classical neural networks on the left, ResNets on the right. Dashed lines denote training error, and bold lines denote testing error (training on CIFAR-10 [4]) [9]. . 85 xv List of Tables 5.1 Theconfusionmatrix. .................................. 51 5.2 Confusion matrix computed with the results obtained by comparing predictions and ground truth for the model that uses ResNet50 as backbone architecture. . . . . . 56 5.3 Confusion matrix computed with the results of comparing predictions and ground truth for the model that uses ResNet101 as backbone architecture. . . . . . . . . 58 5.4 Confusion matrix computed with the results of comparing predictions and ground truth for the model that uses ResNet152 as backbone architecture. . . . . . . . . 59 5.5 Confusion matrix values and computed evaluation quantities for each of the trained models. .......................................... 59 xvii List of Abbreviations ANN (or NN) Artificial Neural Network CNN Convolutional Neural Network CPU Central Processing Unit CSV Comma-Separated Values CV Computer Vision DL Deep Learning DNN Deep Neural Network EUSA European Union of Swimming Pool and Associations FCL Fully-Connected Layer FCN Fully-Connected Network FPN Feature Pyramid Network GUI Graphical User Interface IMI Imposto Municipal sobre Imóveis IoU Intersection over Union MCC Matthews Correlation Coefficient ML Machine Learning RCNN Region-based Convolutional Neural Network ReLU Rectified Linear Unit RGB Red, Green, Blue RoI Region of Interest RPN Region Proposal Network SVM Support Vector Machine UK United Kingdom XML Extensive Markup Language xix Chapter 1 Introduction 1 1.1 State-of-the-Art Fifty years ago, Machine Learning was still science fiction. Today it’s an integral part of our lives, helping us do everything from finding photos to driving cars. Machine Learning has evolved along time, mainly due to philosophers, filmmakers, mathematicians, and computer scientists who fuelled the dream of learning machines. One can distinguish three milestones. 1642: Blaise Pascal was 19 when he made an arithmetic machine for his tax collector father. It could add, subtract, multiply, and divide. Three centuries later, Machine Learning is used to combat tax evasion in income taxes. •1679: German mathematician, philosopher, and occasional poet Gottfried Wilhelm Leibniz devised the system of binary code that laid the foundation for modern computing; •1847: Philosopher George Boole created a form of algebra in which all values can be reduced to true or false. Essential to modern computing, Boolean logic helps a CPU (Central Processing Unit) decide how to process new inputs; •1936: Inspired by how we follow specific processes to perform tasks, English logician and cryptanalyst Alan Turing theorised how a machine might decipher and execute a set of instructions. His published proof is considered the basis of computer science; 1943: A neurophysiologist and a mathematician co-wrote a paper on how human neurons might work. To illustrate the theory, they modelled a neural network with electrical circuits. In the 1950s, computer scientists would begin applying the idea to their work. •1952: Machine Learning pioneer Arthur Samuel created a program that helped an IBM computer get better at checkers the more it played. Machine Learning scientists often use board games because they are both understandable and complex. Arthur Samuel is considered by many as the father of Machine Learning; •1959: In computing, a neural network is a system modeled on the human nervous system. The first neural network applied to a real world problem, Stanford’s MADALINE used an adaptive filter to remove echoes over phone lines. It’s still in use today; •1985: Invented by Terry Sejnowski and Charles Rosenberg, this artificial neural network taught itself how to correctly pronounce 20,000 words in one week. Early outputs sounded like gibberish, but with training its speech became clearer; •1997: IBM’s Deep Blue beat chess grandmaster Garry Kasparov, it was the first time a computer had bested a human chess expert. Kasparov demanded a rematch, but IBM declined and immediately retired Deep Blue; •1999: Computers can’t cure cancer, but they can help us diagnose it. The CAD Prototype Intelligent Workstation, developed at the University of Chicago, reviewed 22,000 mammograms and detected cancer 52% more accurately than radiologists did; 3 In this chapter, a brief description on object detection and the most prominent Deep Learning techniques for object detection are presented. 2.1 Object Detection When looking at an image, the human brain has the ability to instantly locate and recognise different objects on that image (object classification). Nowadays, the extraction of information from a digital image or video is becoming a necessity due to the increasing number of applications that rely on it. For instance, face recognition [17], autonomous driving [18], and pedestrian detection [19]. In the field of Computer Vision (CV), this challenging problem is called object detection [20]. In object detection, the goal is to locate and classify objects in an input image by outputting bounding boxes for each object, and a class 1label for each box. That is, given an input image X, the output is expected to be in the form (X, Y )where Yis an array containing the class label and the box edge’s coordinates. An image is a visual representation of a real-life object. Each image is made from small boxes called pixels. Each pixel’s colour is represented by an RGB (Red, Green, Blue) number from a colour space where each colour’s intensity is in the range of 0to 255 [21]. This colour space is used because it mimics the way the human eye works (the human eye has cone cells located in the retina that are able to identify three colours (red, green and blue) and their combinations [22]). Computers store RGB images in the form of three matrices (or channels), one for each main colour, where each pixel is converted to a number between 0and 255 for each channel. When three colour matrices are blended, by adding the values of each pixel, the image is computed. An example is shown in figure 2.1. 1A class is a group formed by objects with common attributes or characteristics and usually it refers to the object’s name, “chair” for example. 11 Figure 2.1: An RGB image is composed of three colour channels (red, green and blue) that when blended together, the image is computed. Traditionally, the CV techniques used for object detection are based on feature descriptors. These are a part of the broader family of Machine Learning methods. These have the ability to identify/extract features for each object class, a small list of interesting points that describe an image’s content (points, edges or colour variations) and that allow for image classification 2. To accomplish this, several CV algorithms are used, namely edge detection, corner detection or threshold segmentation [23]. For example, during training for image classification, all the detected interesting points (the most important features) form a definition’s list of the object class. When classifying new images, if a significant number of points belong to the definition list, then the image is classified as belonging to that class (for more information on feature descriptors refer to [24]). If an image contains several objects (classes) to be identified, feature extraction becomes computationally expensive. This problem will be aggravated with the increasing number of features per class. To prevent from this, the user needs to go through a trial and error process to determine which features best describe the different classes [25]. With the use of Deep Learning methods, the computer receives several images, in which the classes present have been annotated, and the DL method 3is able to discover the most descriptive features of each class by itself, acting as a human in feature descriptors. This allows for better performance than the traditional methods with the disadvantage of the user not being able to know which features were given higher emphasis (it is used as a black box procedure). The difference between a ML and a DL workflow is represented in figure 2.2. 2Image classification is a computer vision process in which an image is classified based on its visual contents. 3A DL method is a computer program developed to simulate the outcome of a situation. 12 Figure 2.2: (a) Traditional computer vision methods vs. (b) Deep learning workflow [1]. 2.2 Artificial Neural Networks As already mentioned, Deep Learning [26] is a part of a broader family of Machine Learning methods based on Artificial Neural Networks (ANN or NN) with feature learning (also known as representation learning). ANN were inspired by information processing and distributed communication nodes in biological systems [27]. ANN have various differences from biological brains. Specifically, NN tend to be static and symbolic, while the biological brain of most living organisms is dynamic (plastic) and analog (encodes information as a continuum). ANN may contain many layers of neurons that are used to train models by using an extensive amount of labelled data. This can be described as mathematical models of the human brain with the purpose of processing nonlinear relations between inputs and outputs [27]. It should be remarked that Deep Learning architectures have produced results comparable to and in some cases surpassing human expert performance [27]. Artificial Neurons Artificial neurons are the most elementary unit of an ANN. In the human brain, when an electrical signal is transmitted due to an external stimuli, through the dendrites, a neuron receives signals as inputs and sends output signals to other neurons via synapses [28], this is pictured in figure 2.3. 13 Figure 2.3: Components of a biological neuron network [2]. Similarly, in an artificial neuron (often called a perceptron) the received signals (xi) interact in a multiplicative way with the dendrites based on certain weights (wi) (expressing the strength of the synapse). The weights are learned by the network and control the impact that one neuron has on another. This impact can be excitatory or inhibitory according to positive or negative weights respectively. Additionally, a bias, b, that stores a value is added to the sum of the weighted inputs. This bias is not influenced by the weights but it contributes to the final output result, allowing the user to control the behaviour of the layer 4[29]. A representation of an artificial neuron can be viewed in figure 2.4. Figure 2.4: Components of an artificial neuron, designated by perceptron [2]. If there are multiple inputs to the network (x0, x1, x2, ..., xn), each is multiplied by a weight (synapse) (w0, w1, w2, ..., wn). In order to generate a result, the products are summed and fed to an activation function [29]. An activation function introduces non-linear properties to a NN . It determines whether a neuron is activated or not if its value is above a certain threshold. Based on the final result, the next layer of neurons may be activated [30]. This leads the NN to learn complex connections between input and output, this process is often designated by training. An ANN can be built by stacking numerous perceptrons into layers, as seen in figure 2.5. Thus, 4A layer refers to a group of neurons that work together at a certain depth of a NN, figure 2.5 14 the most basic ANN has three components: an input layer that receives the provided data to the NN; a hidden layer that has the task of discovering relationships between the features in the input; and an output layer that produces the result of the given inputs [30]. Figure 2.5: Composition of an Artificial Neural Network with a single-hidden layer. Deep learning is introduced when there is more than one hidden layer, meaning that the difference between a single-hidden layer ANN and Deep Learning lies in the depth of the model [31]. Figure 2.6 shows an example of a Deep Neural Network. The activation function must be a non-linear function so that adding more hidden layers contributes to obtain a better model. If a linear function is used, the neural network will behave like a perceptron due to the fact of the sum of several linear functions being a linear function. This means that adding more hidden layers won’t produce a higher performance Deep Learning model and will result in a neural network lacking the ability to learn complex patterns from the data. Figure 2.6: Composition of a Deep Artificial Neural Network. 15 The most common ANN are structured in a way where each neuron is connected to every other neuron in the next layer, named Feed-Forward ANN. However, better results can be achieved by connecting neurons to other neurons in different patterns, meaning that every possible combination can be studied (for more information on the most used patterns and their applications refer to [32],[29]). When a Deep Neural Network (DNN) is used for classification, the last layer is given by a function which computes the probability of the input belonging to the classes in study (it is known by logistic regression). Two of the most used functions are the Softmax and the Sigmoid functions [33]. The Softmax function extends the logistic regression idea into a multi-class problem by normalising the input vector into a vector of values in which the total sum is 1. This allows the usage of as many classes as needed for the problem we are dealing with. Unlike the latter, the Sigmoid function outputs the predicted probabilities in the range of 0and 1being the total sum of these not necessarily 1. In a DNN it is necessary to define weights for each connection. During the DNN’s training, the aim is to optimise the weights that minimise a loss function that measures the error, i.e., the difference between the expected values and those obtained by the network. There is no universal loss function and there are various factors to consider when choosing one for a given problem [34]. In this thesis, two loss functions will be presented, namely the cross entropy and focal loss. 2.3 Convolutional Neural Networks Object detection can be achieved by ANN, but a computational problem arises due to the proportional increasing number of parameters of the ANN with the increasing number of the input features associated to the image (image’s size). For instance, the two-dimensional image must be converted into a one-dimensional vector which is the input format required by an ANN. For example, if the dataset has images of size 64 ×64 ×3(since the image is RGB there will be 3 colour channels), then there will be 12288 features, but if the size is 1000 ×1000 ×3then 3000000 features will be provided. This leads to an exponential increase of time, computational cost and to the loss of the spatial features (pixels arrangement) of an image [35]. An alternative that fixes these problems is the Convolutional Neural Networks (CNN). A Convolutional Neural Network (CNN) takes an image as input, defines a weight matrix (explained later in this subsection), designated by filter, and then the input is convolved to extract specific features without losing the information about its spatial arrangement. In contrast to the ANN, the convolved matrices are smaller in size and consequently the CNN will have a smaller number of parameters. Thus, leading to a dramatically reduction of the time and computational cost. Just like the ANN, a CNN’s number of parameters increases with the size of the input but at a slower rate since the input image is reduced to a feature matrix [35][31]. To set up a basic Convolutional Neural Network it is necessary to define three basic components: 1. Convolutional layer: Every image can be represented as a matrix of pixel values. The concept behind this layer is to define a weight matrix (filter) in order to compute the dot 16 product between the values of the filter and the image pixel matrix. This is accomplished by sliding the filter across the width and height of the input matrix. With this approach, it is possible to extract certain features from the images (for a detailed explanation on filters refer to [36]). The next example illustrates how this process works. Suppose the input is an RGB image of size 6×6×3. The weight matrix, a 3×3×1filter, for example, runs across each channel of the image matrix in such a way that all the pixels are covered at least once (stride of 1), resulting in a convolved output. Each 6×6×1input channel matrix is now converted into a 4×4×1matrix as represented in figure 2.7. Figure 2.7: A 6×6×3RGB image example is seen by the computer as a matrix. When a filter is applied to each colour channel, the output is a convoluted matrix. The result is a 4×4×1 matrix. (In the figure only one colour channel is shown.) Several filters can be used as a weight combination with the aim of extracting different information, such as edges, a particular colour (one of the RGB channels as represented in figure 2.1, for example) or to remove unwanted noise. These filters have been studied and documented in the literature, the way the weights are distributed in the matrix is already predefined according to user needs. However, it is possible for the user to create a new filter. The output of the convolutional layer is called the activation map (or feature map) whose depth is equivalent to the number of filters applied. These filters slide across the input image with a jump of 1or more pixels, the size of this jump is called stride. When applying several filters, these are applied to the input image, meaning that there will be the same number of feature maps as filters applied. It is important to mention that if it is desired to keep the output with the same size of the input image, the “zero” padding technique needs to be applied to the input image [37]. This technique appends a border of zeros around the image 17 and depending on the application it may have advantages [38]. 2. Pooling layer: Due to the size of some inputs, it can be useful to periodically introduce pooling layers between subsequent convolution layers in order to reduce the spatial size of an image. This is achieved by reducing the dimension of the feature map. One of the most used pooling layers, Maximum Pooling (or Max Pooling), computes the maximum value for each patch (defined by the stride and pooling size, represented by the different colours in figure 2.8) of the feature map [39]. Given a stride and pooling size of 3, after the max operation is applied to each depth dimension of the convolved output, the 6×6×1feature map becomes a2×2×1matrix, see figure 2.8. Figure 2.8: A Max Pooling layer applied to the input image in figure 2.7, results in the reduced feature map of 2×2×1. It was used a stride of 3and a Pooling size of 3×3×1. 3. Output layer: The convolution and pooling layers are only able to extract features and reduce the number of features from the original images. In order to generate the desired output, it is required the use of Fully-Connected Layers (FCL). These layers are identical to the traditional DNN (see figure 2.6) with the addition of an activation function in the output layer. This output layer performs the task of classification by using the features extracted by the convolutional and pooling layers. The term “Fully-Connected” means that every neuron in the previous layer is connected to every neuron in the following layer. It is important to point out that if the output arriving to the FCL has more than one-dimension (more than one filter was applied) and since the last layer only receives a single vector of numbers, the last pooling layer output must be flattened into a single vector. Then, the flattened feature map is given as input to the FCL which gives the classes probability as output. [28][31][40]. The architecture of a CNN is schematised in figure 2.9: 18 Figure 2.9: Components of a convolutional neural network [3]. The CNN algorithm steps will be illustrated considering the image detection of swimming pools in satellite images. It is assumed that the CNN was previously trained to detect swimming pools. The steps are the following: 1. Take an image as input (figure 2.10); Figure 2.10: Input image given to the convolutional neural network [4]. 2. Divide the image into various regions (figure 2.11); Figure 2.11: Division of the input image in several new independent images. 19 Figure 2.24: Bounding boxes reshaped and classified. To summarise, Fast RCNN uses a single model (CNN) for feature extraction, classification and bounding box fitting. This architecture, like RCNN, uses the selective search algorithm to find the RoI making this algorithm slow, so its usage is not advised for large real-life datasets. This problem was addressed in the literature, leading to the development of the Faster Region-based Convolutional Neural Network [8]. 2.6 Faster Region-based Convolutional Neural Network The Faster Region-based Convolutional Neural Network (Faster RCNN) is an adjusted version of Fast RCNN. The main difference between the two versions is that, instead of the selective search algorithm, Faster RCNN uses a Region Proposal Network (RPN) [8]. A RPN generates kanchor boxes of various random sizes and shapes by sliding a window over the feature maps outputted by the CNN. The RPN predicts, for each anchor box, the probability of that anchor being an object and the best-fit bounding box for that object (for more information on RPN see appendix B.1). For performing object detection, the Faster RCNN approach uses the following steps: 1. Receive an image as an input (2.25); Figure 2.25: Image given as input [4]. 26 2. Pass the input to the CNN and get the image’s feature map (figure 2.26); Figure 2.26: The input is given to convolutional layers that produce a feature map. 3. Apply the RPN on the feature map, which returns the object proposals and their corresponding objectness score that gives the probability of a box enclosing an object (figure 2.27); Figure 2.27: RoI are detected from the feature maps. 4. Reshape the proposals using a RoI pooling layer, so all feature map’s sizes match (figure 2.28) (for more information on how this layer works see [45]); 27 Figure 2.28: To create the input for a Fully-connected Network, the regions are reshaped using a RoI pooling layer. 5. Pass the resized proposals onto a Fully-connected Network with two layers. One softmax layer for classification and one linear regression to output the bounding boxes (figure 2.29); Figure 2.29: Bounding boxes reshaped and classified. Like every object detection algorithm discussed above, Faster RCNN identifies objects using regions. This implies that the network requires more than one pass through a single image in order to extract all the objects. Another downside of this algorithm is the handling of different systems working in sequence, making the performance of a system to depend on how the previous systems performed (error propagation). 28 2.7 You Only Look Once The You Only Look Once (YOLO) algorithm uses the entire input image and predicts the bounding box and each corresponding class probability using a neural network similar to a CNN. The advantage of this algorithm is its remarkable speed. The object detection using the YOLO algorithm is performed by considering the following steps: 1. Receive an image as input (figure 2.30); Figure 2.30: Input image [4]. 2. Divide the image into an S×Sgrid of cells (figure 2.31); Figure 2.31: Input image divided into an S×Sgrid. 3. For each cell, predict Nbounding boxes and the corresponding confidence scores. These scores report how confident the model is, i.e., it encloses an object and its prediction accuracy. Each bounding box is described by the vector (x, y, w, h, c)where: xand yare the coordinates of the centre relative to the grid cell; wand hare the width and height relative to the whole image, respectively; cis the confidence. The confidence is given by equation 2.1 [46] c=P(object)×IoU (2.1) 29 where P(object)is the probability of the bounding box enclosing an object and IoU is the Intersection over Union (see appendix A.1) between the predicted box and the ground truth. In parallel, the network computes all the classes probability conditioned by the existence of an object in the bounding box, C=P(classi|object). Notice that only one set of class probabilities will be computed per cell, despite the number of existing bounding boxes (figure 2.32). This step output is usually referred to as a feature map [46]. Figure 2.32: Input image with randomly drawn bounding boxes. 4. In order to get class-specific confidence scores for each box, the conditional class probabilities and the box confidence predictions are multiplied [46]. At the end, according to a class probability threshold value provided by the user, the boxes that have the highest class probabilities are filtered in order to have one box for each detected object (figure 2.33). Figure 2.33: Filtered bounding boxes according to the threshold value chosen by the user. 30 2.8 Single-Shot Detector The Single-Shot Detector (SSD) algorithm follows the same strategy as YOLO but instead of using a single feature map (for prediction of classes and bounding boxes) it uses several activation maps with different scales. Therefore, SSD may achieve higher precision due to the improved ability to detect different sized objects on an image. The SSD algorithm divides the image into different sized S×Sgrids of cells instead of just one. This allows the SSD to find smaller objects (smaller cells) or bigger objects (bigger cells) [47]. Despite the SSD precision, YOLO would be a better choice if speed is preferred over accuracy. 2.9 RetinaNet The RetinaNet algorithm is a single-stage detector that has two new improvements over YOLO and SSD: Focal Loss and Feature Pyramid Network (FPN). 2.9.1 Focal Loss Single-stage detectors, like YOLO and SSD, classify the whole image. In general, the background of an image represents a big part of the whole image, these detectors experience class imbalances since most of the bounding boxes do not contain any object. Unlike YOLO and SSD, two-stage detectors like RCNN, Fast-RCNN and Faster-RCNN, start by predicting a few object locations and then use a CNN to classify each of these objects. The RetinaNet algorithm was proposed in order to fix the class imbalances by slightly changing the loss function, as can be seen in figure 2.34. The loss function used by the other algorithms is commonly called cross-entropy and has the following expression, when computed for all the data (equation 2.2): CE(p, y) = −X t ytlog pt,(2.2) where tis the class index, ytthe label (1 if the object belongs to class t, 0 otherwise) and ptis the probability of the object belonging to class t. Consider the following example: in an image, there are 10000 bounding boxes and the neural network predicts, with high accuracy, 9990 are background. If the other 10 boxes contain objects and the network isn’t sure to which class they belong to, the loss of these few true objects will be much smaller than the background loss. Therefore, the network won’t focus on the minority in which it had difficulties due to the overpowering of the large number of easily classified boxes over the minority, which reflects in a low value of the loss function. The creators of RetinaNet modified the cross-entropy loss function so that the contribution of the more easily classified examples (such as the background) is reduced and a more focused learning on the few interesting cases is performed (figure 2.34). This new loss function is called Focal loss and is given by [5]: 31 FL(p, y) = −X t yt(1 −pt)γlog pt,(2.3) where γ∈[0,5] is a modulating factor that reduces the loss contribution from easily classified examples, as seen in figure 2.34. According to Tsung-yi Ling et al. [5], γ= 2 is the best choice. Notice that a well (or easy) classified example is characterised by having a predicted probability above 0.6. Figure 2.34: Graph that shows the loss functions values for each prediction. The cross entropy represented by the blue line and the other full-lines are the focal loss function with a γvarying from 0.5 to 5 [5]. The graph in figure 2.34 shows how the focal loss is able to give less importance to easy classified examples by using a modulating factor, γ. If a part of the image is easily classified then the (1−pt) factor will be close to zero, which induces a very low or no learning, giving more attention to the cases of interest. Note that equations 2.2 and 2.3 expressions are computed for all the dataset while in figure 2.34 the goal is to study the loss variation in a single example. 2.9.2 Feature Pyramid Network Detecting objects in different scales is a difficult task, specially for smaller objects. To overcome this problem, a possible solution is to give the network several copies of the same image at different scales (resembling a pyramid). Processing multiple scale images is computationally expensive as well as time consuming, and therefore, a Feature Pyramid Network (FPN) is proposed. The Feature Pyramid Network is a feature extractor that generates multiple multi-scale feature maps with focus on both accuracy and speed. The FPN comprises a bottom-up and a top-down pathway, represented in figure 2.35. The bottom-up pathway uses ResNet, and, for each layer, the spatial resolution is decreased by a factor of 1/2by doubling the stride used for the next convolution block (which may contain several layers) , allowing more high-level structures to be detected (there is an increase of semantic value). The output of each layer is a feature map that, by lateral connection, will be used for enriching the top-down pathway. In the top-down pathway, the spatial resolution is increased by a factor of 2(for each layer) using the nearest neighbour 32 upsampling 5. Although semantically strong, the reconstructed layers display imprecise object locations. To prevent this from happening, lateral connections are added between reconstructed layers and the corresponding feature maps by applying a 1×1convolution 6to the feature maps and add them to the reconstructed layers, element-wise. A scheme is shown in figure 2.35 [6]. Figure 2.35: Illustration of the bottom-up and top-down pyramid along with a lateral connection scheme [6]. 5The nearest neighbour upsampling increases the size of images by assuming the new pixels have the same intensity as the closest pixel. 6A1×1convolution collapses all the input pixel channels into one pixel, allowing for a reduction of the number of feature maps while retaining the salient features. This is useful to reduce the number of depth channels so as to reduce the computational cost by reducing the number of parameters that will be passed to the next phase of a neural network. For example, if the input has a size of 64 ×64 ×3and a 1×1convolution is applied (1×1×3 filter), the output will have the same size of the input but only one channel, 64 ×64 ×1. 33 2.9.3 Network Architecture In this subsection the main components of RetinaNet are summarised: •Bottom-up pathway: Backbone network called Feature Pyramid Network (in this project, built on top of a ResNet) which computes several convolutional feature maps at different scales of an entire image, regardless of the input image size; •Top-down pathway and lateral connections: Upsampling of the spatially coarser feature maps from higher pyramid levels. Same size top-down and bottom-up layers are associated by lateral connections; •Classification subnetwork: Fully-Connected Network (FCN) 7responsible for predicting the probability of an object to belong to an anchor box and to perform the classification of the given objects. This subnetwork is composed of four 3×3convolutional layers with 256 filters (number arbitrarily chosen by the authors for design simplicity and robustness [6]) followed by a Rectified Linear Unit (ReLU) activation function 8. Finally, one more 3×3convolutional layer with K×Afilters is applied with a Sigmoid activation function, outputting (W, H, K ×A)feature maps (where Wand Hare the width and height of the classification subnetwork input map, Kand Aare the number of classes and anchor boxes, respectively), figure 2.36 9. The activation function used in this stage can be either Sigmoid or Softmax, in this thesis the Sigmoid function was used; •Regression subnetwork: Responsible for executing bounding box regression with the purpose of approximating the box to the ground-truth. The architecture is identical to the classification subnetwork with the exception of the last convolutional layer for which K= 4, outputting feature maps with shape (W, H, 4×A), figure 2.3610; Figure 2.36 shows the architecture of the RetinaNet. Figure 2.36: RetinaNet network architecture composed of four main components: a) bottom-up pathway; b) top-down pathway; c) classification subnetwork; d) regression subnetwork; [5]. 7A fully-connected network is an artificial neural network which makes usage of FCL. 8The ReLU activation function is a linear function that outputs the input if it is positive and 0otherwise. 9For example, in a 4×4feature map, for each 16 grid cells, 16 different anchor boxes will be used by RetinaNet. Since each box will be looking for Kclasses and there are Aboxes per grid, the output map of this subnetwork will have K×Achannels. 10This subnetwork outputs 4coordinates that characterise each anchor box. Since there are Aboxes per grid, the output feature map of the regression subnetwork will have 4×Achannels 34 As shown in figure 2.36, after the pyramid is concluded, a 3×3convolution is applied to each layer’s map in order to generate the final feature maps with a reduced aliasing effect (caused by the upsampling). Then, a classification and a regression networks are applied, in parallel, to each FPN level. From these results classified bounding boxes from all levels, which may result in an object being predicted in more than one layer. To prevent this from happening, the boxes with IoU greater than 0.5are filtered by keeping the one with the highest confidence score. Based on the Deep Learning techniques presented in this chapter, and their advantages and disadvantages that affect the detection of swimming pools in satellite images, one concluded that the RetinaNet algorithm is the most adequate to solve this problem. 35 In order to accomplish the task of swimming pools detection, it is necessary to exclude the cars labels from the files since the goal of this work doesn’t include car detection. A Python script was implemented to filter the training images by considering the following approach. Open every training labels file and for each file do: •Count the number of pools (denoted by 2) and cars (denoted by 1) labels; •If the image contains only cars, remove from the training set the labels file and the corresponding image. Since there are training images with pools that also have cars, the information relative to cars must be deleted from the labels file. To remove it, a Python script was implemented using the following steps: •Open every training labels file; •if the label is 1, then delete all the information related to that label. After applying this Python script to the training dataset, the number of training images is 1993. 42 Chapter 4 Training RetinaNet 43 The deep learning algorithm chosen to achieve the detection of swimming pools in satellite images was RetinaNet. The main reasons to chose this algorithm was the fact that the satellite images for detection may have different sizes and a huge amount of background. Several implementations can be found in the literature, being free to use and modify. The implementation used in this thesis can be found on Github [53]. 4.1 Data Preprocessing of the Training Set - Pascal VOC to CSV In order to use the RetinaNet algorithm, the input data must have the correct format. Thus, to train a model, it must be given as input a training set of images and two CSV files, namely: one CSV (Comma-separated Values) class mapping format file with the classes corresponding labels and one CSV with the annotations. In this case, the mapping file with the existing classes has only one label since the only objects to be detected are swimming pools. In case there were more, the file should have one class per line. The following structure must be used: class_name_0 , id_0 class_name_1 , id_1 where the class id must start at 0. Since there is only one class, it must have the first available id value, 0. Thus, the CSV mapping file has the following structure: pool , 0 The CSV file with annotations must have one box annotation per row. Therefore, images in which there are multiple objects (multiple bounding boxes) must have one row for each bounding box. Thus, each row must have the following specific format: path/ to /image . jpg , x1 , y1 , x2 , y2 , class_name An example from the chosen dataset must resemble the following: 000000012. jpg ,14 9 . 5 3 ,196. 1 1 , 1 93.97 , 2 2 4 .00 , pool 000000012. jpg ,12 0 . 2 4 ,212. 7 7 , 1 58.87 , 2 2 4 .00 , pool 000000014. jpg ,21 1 . 7 1 ,156. 4 1 , 2 24.00 , 1 9 6 .16 , pool Since the training dataset is not in the required format by the RetinaNet algorithm, it is imperative to convert the data labels files in XML format to the CSV format described above. As mentioned before, the PASCAL VOC provides a standardised form to annotate image datasets for object detection. The main attributes of this type are the existence of one annotations file per 45 image in the XML format. A Python script was written to transform the dataset annotations files to the input required format (CSV format), as follows: Create a new CSV file, open the PASCAL VOC files one by one and for each file do: •find the information under the tag “name”; •find the information under the tag “xmin”; •find the information under the tag “xmax”; •find the information under the tag “ymin”; •find the information under the tag “ymax”; •write a new row in the CSV file with the name, xmin, ymin, xmax, ymax and class name info for each bounding box. 4.2 Training of RetinaNet Models After the preprocessing of the training dataset, the data is ready to be used to train the Deep Learning model given by the RetinaNet algorithm. This algorithm allows a certain level of training customisation by having available certain options: •Backbone architecture: The FPN can be built on top of several architectures. The default is ResNet50 but both ResNet101 and ResNet152 are also available. The choice between the various ResNet influences on the model’s results since there is a difference of depth between them, represented by the numbers 50,101 and 152, being ResNet50 the shallowest and ResNet152 the deepest network (for more information on ResNet see appendix C.1) [5]; •Weights: To initialise the weights of the DNN given by the RetinaNet algorithm, an available weights file can be used. This means the training is not done from scratch, speeding up the training process (the algorithm will start to converge earlier) [31]. By default, ImageNet weights are used (more information on ImageNet and weights in appendix D.1); •Epochs: One epoch is when an entire dataset is passed through the neural network only once. By default, RetinaNet sets this option to 50 [31]; •Batch size: The number of training examples in a single batch, by default, is 1. A dataset can be divided into batches (subsets of the training dataset) if it is not possible to pass the entire training dataset into the neural network at once due to its size [29]; •Steps: (or iterations) 10000 by default, steps is the number of batches needed to complete one epoch [29]; The Deep Learning models training was done using the default settings with the exception of the epochs, iterations and backbone architecture parameters. The first two parameters were 1and 46 10000 by default, respectively. Due to the high computational cost, the training takes a huge amount of time. Therefore, 2epochs and 500 iterations were used, which lowered the training time to less than a day , for all models. Since the training dataset, from Kaggle, has 1993 images and a batch size of 1and 2epochs were chosen, the dataset will be divided into 1993 batches, each with one image. This means that one epoch will involve 1993 batches in which the model weights will be updated after each batch. Three different backbone architectures were used in order to analyse which would give better results. Therefore, using the training dataset with 1993 images, three separate models were trained, each with a different architecture, namely: ResNet50, ResNet101 and ResNet152. 47 Chapter 5 Testing RetinaNet - Results and Discussion 49 5.1 Performance Metrics To analyse the performance of each RetinaNet model, a confusion matrix (square matrix of order 2) was computed. This matrix allows a better understanding of the types of errors a model is making. Based on the confusion matrix it is possible to determine the following measures: accuracy, precision, sensitivity, specificity and the Matthews Correlation Coefficient (MCC) of the model. Additionally, two error types, Type I and Type II, can also be computed to deepen the understanding of the model’s fouls. To compute a confusion matrix it is necessary to have a test dataset with the respective expected outcome values. The expected outcome values and predictions are compared to get the four values of the confusion matrix. These values are obtained by counting the number of results in each of the following categories [54] [55]: •True Positive (TP): The expected and the predicted values are positive. In this case, the image has a swimming pool in a certain location and the model accurately identified it. For each swimming pool, correctly detected, is added 1to the TP value (note that an image with 3correctly classified objects contributes with 3to the TP value); •True Negative (TN): The expected and the predicted values are negative. In this case, the image has no swimming pools and the model doesn’t classify any part of the image as one. For each image, correctly classified as not having pools, is added 1to the TN value; •False Positive (FP): The predicted value is positive but the expected is negative. In this case, the image has no swimming pools but the model identifies a part of the image as being one. For each pool detected but not being one, is added 1to the FP value; •False Negative (FN): The predicted value is negative but the expected is positive. That is, the image has swimming pools but the model isn’t able to detect any. For each image, wrongly classified as not having pools, is added 1to the FN value. The confusion matrix is organised as it follows in table 5.1. Predicted value|Ground Truth Positive Negative Positive TP FP Negative F N TN Table 5.1: The confusion matrix. Using the four values that form the confusion matrix, it is possible to compute five performance metrics relative to the model, defined by equations 5.1, 5.2, 5.3, 5.4 and 5.5, and two error types, given by equations 5.6 and 5.7. The accuracy (equation 5.1) is given by the division between the total of correctly classified examples and the total of predictions made. However, using only this measure to assess the model’s performance is not recommended since it assumes equal weight for both types of errors (FN and FP). For instance, if only 1% of the images encompass swimming pools and the model 51 The 2703 test images were given as input to the model which gave the predictions as outputs. The model’s bounding boxes were compared with the corresponding ground truth in order to evaluate the model’s performance using the confusion matrix, see table 5.3. Predicted |Ground Truth Positive Negative Positive 480 7 Negative 133 2285 Table 5.3: Confusion matrix computed with the results of comparing predictions and ground truth for the model that uses ResNet101 as backbone architecture. From table 5.3, this model was able to correctly classify 480 objects as swimming pools, from the test images, (represented by the TP value) and 2285 as not (TN value). 7objects were classified as pools while not being one (FP value) and 133 images were wrongly classified as not having a pool (FN value). With these values and using the equations 5.1, 5.2, 5.3, 5.4 and 5.5, and it was computed an accuracy of 0.9522, a precision of 0.9626, a sensitivity of 0.7830, a specificity of 0.9969 and a MCC of 0.8519. The type I and II errors, given by equations 5.6 and 5.7, were also computed being 0.0031 and 0.2170, respectively. 5.3.3 Model using ResNet152 Finally, the images in figure 5.2 were passed to the model that uses a backbone structure built on top of a ResNet152. This model was able to detect all the swimming pools correctly, as can be seen in figure 5.8. Figure 5.8: Model’s predictions of two testing images of the dataset using ResNet152 as backbone architecture. The model’s certainty of the objects enclosed by the bounding boxes being swimming pools has minimum 0.674, maximum 0.871 and mean 0.783 values being the model with the lowest scores, obtaining the lowest minimum, maximum and mean confidence values of all three trained models. The images with no swimming pools (figure 5.3) were also given, resulting in figure 5.9. 58 Figure 5.9: Model’s predictions of two testing images without swimming pools. Once again, all the images were correctly classified however this model got the lowest confidence scores. This fact could indicate the other models are more suitable for the task. Nevertheless, the cases studied don’t represent a significant number of tests as to make conclusions about performance. Therefore, the 2703 test images were given as input to the model which gave the predictions as outputs. The model’s bounding boxes were compared with the corresponding ground truth in order to evaluate the model’s performance using a confusion matrix, see table 5.4. Predicted |Ground Truth Positive Negative Positive 503 6 Negative 111 2241 Table 5.4: Confusion matrix computed with the results of comparing predictions and ground truth for the model that uses ResNet152 as backbone architecture. Table 5.4 shows that this model was able to correctly classify 503 objects as swimming pools, from the test images, (represented by the TP value) and 2241 as not (TN value). 6objects were classified as pools while not being one (FP value) and 111 images were wrongly classified as not having a pool (FN value). With these values and using the equations 5.1, 5.2, 5.3, 5.4 and 5.5, it was computed an accuracy of 0.9591, a precision of 0.9882, a sensitivity of 0.8192, a specificity of 0.9973 and a MCC of 0.8766 . The type I and II errors, given by equations 5.6 and 5.7, for this model are 0.0027 and 0.1808, respectively. 5.3.4 Conclusions The results obtained with the three models are shown in table 5.5. TP TN FP FN Accuracy Precision Sensitivity Specificity MCC Type I Type II ResNet50 471 2200 18 131 0.9472 0.9631 0.7824 0.9919 0.8380 0.0081 0.2176 ResNet101 480 2285 7 133 0.9522 0.9626 0.7830 0.9969 0.8519 0.0031 0.2170 ResNet152 503 2241 6 111 0.9591 0.9882 0.8192 0.9973 0.8766 0.0027 0.1808 Table 5.5: Confusion matrix values and computed evaluation quantities for each of the trained models. 59 Analysing the results in table 5.5, it can be seen that every model’s accuracy is high and remarkably close to each other, indicating that the classifiers are correct most of the times (approximately 0.95 of accuracy). All RetinaNet models have high precision (above 0.95) indicating that an image labelled as positive is truly positive, as verified by the small number of false positives reported in the table. The model using a ResNet152 backbone architecture is capable of better recognising the class in study (lowest number of false negatives). This is also corroborated by the highest sensitivity value obtained for this model. The three models have high specificity, being above 0.99 for all three models, which means the models predict negative examples well (images without swimming pools). The Matthew’s Correlation Coefficient (MCC) is higher for the model built on top of a ResNet152, demonstrating, out of the three models, the effectiveness in classifying both classes (presence or absence of pools in an image). Moreover, the ResNet152 has the smallest Type I and Type II errors suggesting that the predictions are correct most of the times. A closer look at the false positive and false negative classifications, exposes the type of difficulties faced by each model and the kind of objects that compromise the swimming pools detection: •Every model detects blue rectangles, that may be canopies, for instance, as swimming pools; •All models have difficulties in detecting empty pools; •Both ResNet50 and ResNet101 models identify blue basketball fields as pools; •The model that uses a ResNet50 as backbone architecture wrongly detects circular blue looking canopies; From this analysis one may conclude that the model using a ResNet152 backbone architecture outperforms the models using ResNet50 and ResNet101. Furthermore, all three models have similar computational cost. 5.4 Test set from Google maps As to further test the limitations of the trained models, a few Google Maps cropped images were used. This images were taken with various zoom percentages and from areas with different image quality. The eight images chosen for testing are shown in figure 5.10. 60 Figure 5.10: Eight test images taken from google maps with different sizes and resolutions. The images are numbered in yellow on the top right corner. The test images in figure 5.10 were specifically chosen for enclosing objects very similar to swimming pools and swimming pools without the usual appearance. It is listed below the characteristics of each image: •image 1 -433 ×340 pixels, a lake and a blue tag are displayed; •image 2 -442 ×258 pixels, encloses two swimming pools and a swimming pool look-alike glass structure on top of a building; •image 3 -298 ×241 pixels, a pool without one of the top characteristics, the blue water; •image 4 -100 ×118 pixels, the image is very small with high zoom; •image 5 -366 ×270 pixels, unusual shape swimming pool present; •image 6 -1351 ×1364 pixels, very high-quality 3-dimensional picture with 2pools and a blue basketball court; •image 7 -1099 ×1333 pixels, the same court as in image 6and a green water swimming pool; •image 8 -732 ×613 pixels, the basketball court seen in images 6and 7but with higher zoom. 61 5.4.1 Model using ResNet50 The images in figure 5.10 were given to the RetinaNet model built on top of a ResNet50. The detected swimming pools were enclosed by a blue bounding box as shown in figure 5.11. Figure 5.11: Eight test images taken from google maps with different sizes and resolutions. The detected objects are enclosed by blue boxes. The images are numbered on the top right corner. It is listed below the comments on the results obtained for each image: •image 1 - The model was able to detect no swimming pool, even though the image contains a lake and a blue tag; •image 2 - It correctly detected the two swimming pools and did not identify the glass structure as a swimming pool; •image 3 - The unclean swimming pool was detected by this model; •image 4 - Several detections were made for a single pool. The very low resolution and size of this image might explain the multiple bounding boxes due to the pixels being easily distinguished; •image 5 - Three pools were detected at the real location of a single pool. The error might be explained by the unusual shape; •image 6 - The model exceeded the expectations by correctly identifying the swimming pools and not mistaking the blue basketball court by one; 62 •image 7 - Similarly to what happened with image 3, the unclean pool was not identified. This may be the reason why the lake in image 1was not detected. Also, the previously not detected court (image 6) was, this time, classified as a pool. The difference between the pictures indicate that this model is sensible to zoom percentages; •image 8 - The test done on this image supports the statement presented in the previous item. The higher zoom contributed to deceive this model; 5.4.2 Model using ResNet101 The images in figure 5.10 were given to the RetinaNet model built on top of a ResNet101. The detected swimming pools were enclosed by a blue bounding box as represented in figure 5.12. Figure 5.12: Eight test images taken from google maps with different sizes and resolutions. The detected objects are enclosed by blue boxes. The images are numbered on the top right corner. It is listed below the comments on the results obtained for each image: •image 1 - The model was able to detect no swimming pool, even though the image contains a lake and a blue tag; 63 •image 2 - It was not capable of detecting any of the swimming pools, falling behind the ResNet50 model for this image; •image 3 - The unclean swimming pool wasn’t detected by this model; •image 4 - Several detections were made for a single pool. Again, the very low resolution and size of this image might explain the multiple bounding boxes (due to the pixels being easily distinguished); •image 5 - Similarly to the previous model, three pools were detected at the real location of a single pool; •image 6 - The model correctly identified the swimming pools and did not mistake the blue basketball court by one; •image 7 - Both the unclean pool and court were correctly classified. The only difference between this image and image 3is the quality. Image 7has higher quality and presents 3-dimensional textures; •image 8 - The bigger zoom didn’t influence the model’s detection; 5.4.3 Model using ResNet152 Finally, the images in figure 5.10 were given to the RetinaNet model built on top of a ResNet152. The detected swimming pools were enclosed by a blue bounding box as represented in figure 5.13. 64 Figure 5.13: Eight test images taken from google maps with different sizes and resolutions. The detected objects are enclosed by blue boxes. The images are numbered on the top right corner. It is listed below the comments on the results obtained for each image: •image 1 - The model was able to detect no swimming pool, even though the image contains a lake and a blue tag; •image 2 - The ResNet152 model was capable to accurately detect the objects in this image; •image 3 - The unclean pool was properly identified; •image 4 - Unlike the previous models, it wasn’t able to detect the swimming pool in this image; •image 5 - In contrast to the ResNet50 and ResNet101 models, the model using ResNet152 correctly detected a single pool in this image; •image 6 - It correctly identified the swimming pools and did not mistake the blue basketball court by a swimming pool; •image 7 - Both the unclean pool and court were correctly classified; •image 8 - The basketball court was not able to trick this model; 65 5.4.4 Conclusions The results obtained with the three models are now summarised. •The models using ResNet50 and ResNet101 are not able to correctly classify a swimming pool when the water has an unusual colour (green as in the tested examples); •The results indicate that the ResNet50 model is sensitive to the image’s zoom (wrong detection of objects when the zoom percentage is high); •Classification of exceptionally small images may be difficult for the model using ResNet152. In conclusion, the model that had the best performance was the one built on top of a ResNet152. It was able to correctly detect a bigger number of swimming pools, except for the smallest image. 66 Chapter 6 Graphical User Interface 67 Chapter 7 Conclusions and Future Work 75 7.1 Conclusions The main goal of this work was to develop a software application to identify swimming pools in satellite images using a Deep Learning approach. After an intensive study on the advantages and disadvantages of the available algorithms used for object detection, the RetinaNet algorithm was chosen. Three RetinaNet models were trained, each built on top of a differing backbone architecture, using a dataset from Kaggle. After the training, the models were tested using the test set composed of 2703 images with 620 swimming pools distributed by 524 images. The results were analysed using seven performance metrics, that are computed based on confusion matrices. From these results, one may conclude that the model using a ResNet152 backbone architecture outperforms the models using ResNet50 and ResNet101. The models were further tested using images obtained from Google Maps. The idea is to test the models under extreme conditions. These images contain objects similar to a swimming pool, such as a blue basketball court, and also swimming pools with unusual colours, several uncommon shapes and sizes. The results indicate that the ResNet50 model is sensitive to the image’s zoom, while the ResNet101 model isn’t consistent with the characteristics of the pools detected. Furthermore, the classification of really small images may be difficult for the model using ResNet152. According to the results obtained, the model with a backbone architecture built on top of a ResNet152 has the best performance when used for the detection of swimming pools in satellite images. Therefore, the ResNet152 model was adopted to integrate the interface. A Graphic User Interface was developed in order to make the process of detecting illegal swimming pools using satellite images user friendly. The interface allows to select single images or a folder of images for detecting swimming pools. The classification information is stored in a CSV file and saved in the chosen path. Additionally, blue bounding boxes are drawn around the detected objects if the user desires to have a visual representation of the classifications. 7.2 Future Work Several improvements can be considered in the future. The most challenging one involves training the model using more epochs and iterations (steps) in order to achieve a better performance. To do that, one must use a cluster or a supercomputer so that the training computational time becomes feasible. Regarding the newly developed GUI, new functionalities can also be added. Namely, one can create a python script to perform the automatic acquisition of the satellite images and give the geographic coordinates of the swimming pools as an output. Furthermore, other recent object detection algorithms can be analysed and compared with the RetinaNet performance, for this specific problem. Namely, the EfficientDet, the Sniper and the Recurrent YOLO. 77 Appendix A A.1 Intersection over Union (IoU) Intersection over Union (also known as the Jaccard index) is the most used metric to evaluate the accuracy of an object detection algorithm on a specific dataset. Essentially it calculates the similarity between the predicted box and the ground truth. In order to apply the IoU metric two variables are needed, the ground truth bounding boxes (hand labelled bounding boxes from the testing set) and the predicted bounding boxes (output of a neural network). With this two information, the IoU is computed as follows (equation A.1) [57]: IoU =|A∩B| |A∪B|(A.1) where Aand Bare the prediction and ground truth bounding boxes, respectively. In the numerator is computed the area of overlap between the predicted and ground truth bounding boxes, denoted by |A∩B|and in the denominator is the area of union (the area enclosed by both the predicted and ground truth bounding boxes), denoted by |A∪B|. 79 Appendix B B.1 Region Proposal Networks Region proposal networks (RPN) is an algorithm that generates region proposals for the locations of each object in an image by sliding a convolution layer (the authors chose a 3×3window) [8]. The RPN uses two different layers: a regression layer where a convolution layer strides through the image and predicts kanchor boxes (since the sliding window is 3×3, there will be 9boxes for each pixel) at each anchor point location (note that is called an anchor to the central point of the sliding window); a classification layer that predicts the probability of an object being contained by a box [8]. The architecture of RPN is represented in figure B.1. Figure B.1: Region Proposal Networks architecture [8]. As seen in figure B.1, the classification layer outputs 2scores for each k, giving two probabilities, one for object and another for not object. The regression layer outputs 4×kcoordinates that represent the bounding box spatial coordinates. 81 Appendix C C.1 ResNet ResNet is considered a powerful deep neural network having obtained first place in ILSVRC and COCO competitions in 2015 [9]. The ResNet architecture has several versions in which the only change is the number of layers, this number is indicated by a two or more digit number following the name ResNet. In this project, three variants were studied with 50,101 and 152 layers named ResNet50, ResNet101 and ResNet152, respectively [9]. In theory, adding layers to a neural network has two possible effects on the performance, it either increases or remains the same. Accordingly, more layers are never a disadvantage since once the neural network is in its best possible state (achieved 100% accuracy and the loss function is at the global minimum), the new layers added should learn the identity function, g(x) = x; to retain the best possible state (figure C.1). Figure C.1: Adding more layers to a neural network in which the best possible state was attained results on the following layers learning the identity function. Nonetheless, when this concept is tested, the conclusion may be different. Adding more layers to a neural network might lead to a decreasing accuracy in training and testing, as seen in the graphs in figure C.2 [9]. 83 [11] MarketWatch. Swimming pool construction market (3.8% cagr) 2018-2026: Global business growth, size and forecast. (online resouce - https://www.marketwatch.com/pressrelease/swimming-pool-construction-market-38-cagr-2018-2026-global-business-growth-sizeand-forecast-2019-11-05, last accessed on 2020/02/26), 2019. [12] European Union of Swimming Pool and Spa Associations. Sector volume and market characteristics. (online resouce - https://www.eusaswim.eu/, last accessed on 2020/02/25), 2009. [13] Assembleia da República. Lei no60/2007. Diário da República n.o170/2007, Série I, pages 6258–6309, 2007. [14] Ministério do Equipamento do Planeamento e da Administração do Território. Decreto-lei n.o555/99. Diário da República n.o291/1999, Série I-A, pages 8912–8942, 1999. [15] Decreto-Lei n.o287/2003. Código do imposto municipal sobre imóveis e o código do imposto municipal sobre as transmissões onerosas de imóveis. pages 7568–7647, 2003. [16] A. Cabrita. Imi 2019: Um guia com tudo o que precisa saber. (online resource - https://www.doutorfinancas.pt/impostos/, last accessed on 2020/02/26), 2019. [17] Z. Yang and R. Nevatia. A multi-scale cascade fully convolutional network face detector. Proceedings - International Conference on Pattern Recognition, pages 633–638, 2016. [18] X. Chen, H. Ma, J. Wan, B. Li, and T. Xia. Multi-view 3d object detection network for autonomous driving. pages 6526–6534, 07 2017. [19] C. Wojek, B. Schiele, and P. Perona. Pedestrian detection: An evaluation of the state of the art. IEEE transactions on pattern analysis and machine intelligence, 34:743–61, 07 2011. [20] P. Felzenszwalb, R. Girshick, D. McAllester, and D. Ramanan. Object detection with discriminatively trained part-based models. IEEE transactions on pattern analysis and machine intelligence, 32:1627–45, 09 2010. [21] H. Singh. Practical Machine Learning and Image Processing For Facial Recognition Using Python. Apress, Berkeley, CA, 2019. [22] S. J. Ryan, C. P. Wilkinson, S. R. Sadda, and P. Wiedermann. Retina. Elsevier, 6th edition edition, 2017. [23] E. Salahat and M. Qasaimeh. Recent advances in features extraction and description algorithms: A comprehensive survey. Proceedings of the IEEE International Conference on Industrial Technology, pages 1059–1063, 2017. [24] A. I. Awad and M. Hassaballah. Image Feature Detectors and Descriptors: Foundations and Applications. Springer International Publishing, 1st edition edition, 2016. [25] B. Ghojogh, M. N. Samad, S. A. Mashhadi, T. Kapoor, W. Ali, F. Karray, and M. Crowley. Feature selection and feature extraction in pattern analysis: A literature review. 2019. 90 [26] A. Buetti-dinh, V. Galli, S. Bellenberg, O. Ilie, M. Herold, S. Christel, M. Boretska, I. V. Pivkin, P. Wilmes, W. Sand, M. Vera, and M. Dopson. Deep neural networks outperform human expert’s capacity in characterizing bioleaching bacterial biofilm composition. Biotechnology Reports, 22:e00321, 03 2019. [27] Y. Lecun, Y. Bengio, and G. Hinton. Deep learning. Nature, 521:436–44, 05 2015. [28] J. Chapmann. Neural Networks: Introduction to Artificial Neurons , Backpropagation Algorithms and Multilayer Feedforward Networks (Advanced Data Analytcs) (Volume 2). CreateSpace Independent Publishing Platform, 2017. [29] M. A. Nielsen. Neural Networks and Deep Learning. Determination Press, 2015. [30] J. Feng and S. Lu. Performance analysis of various activation functions in artificial neural networks. Journal of Physics: Conference Series, 1237(2), 06 2019. [31] I. Goodfellow, Y. Bengio, and A. Courville. Deep learning. (online resource - http://www.deeplearningbook.org, last accessed on 2020/03/02), 2016. [32] D. Anderson and G. McNeill. Artificial neural networks technology. (online resource - http://andrei.clubcisco.ro/cursuri/, last accessed on 2020/03/20), 1992. [33] A. Zhang, Z. C. Lipton, M. Li, and A. J. Smola. Dive into deep learning. Journal of the American College of Radiology, (online resource - https://d2l.ai, last accessed on 2020/04/05), 2020. [34] W. Di, A. Bhardwaj, and J. Wei. Deep learning essentials: Your hands-on guide to the fundamentals of deep learning and neural network modeling. 2018. [35] A. Khan, A. Sohail, U. Zahoora, and A. S. Qureshi. A survey of the recent architectures of deep convolutional neural networks. Artificial Intelligence Review, pages 1–62, 2020. [36] L. Shapiro and G. Stockman. Computer Vision. Pearson, 1st edition edition, 2001. [37] L. Shapiro. Images and filters notes. CSE/EE 576: Computer Vision - University of Washington, (online resource - https://courses.cs.washington.edu/courses/cse576/, last accessed on 2020/04/21), 2018. [38] M. Hashemi. Enlarging smaller images before inputting into convolutional neural network: zero-padding vs. interpolation. Journal of Big Data, 6(1), 12 2019. [39] D. Scherer, A. Müller, and S. Behnke. Evaluation of pooling operations in convolutional architectures for object recognition. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 6354 LNCS(PART 3):92–101, 01 2010. [40] A. Amidi and S. Amidi. Convolutional neural networks cheatsheet. C2 230 - Deep Learning University of Standford, (online resource - https://stanford.edu/ shervine/teaching/cs-230/, last accessed on 2020/04/07), 2019. 91 [41] R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 580–587, 2014. [42] J. R. R. Uijlings, K. E. A. Sande, T. Gevers, and A. W. M. Smeulders. Selective search for object recognition. International Journal of Computer Vision, 104:154–171, 09 2013. [43] M. Awad and R. Khanna. Efficient learning machines: Theories, concepts, and applications for engineers and system designers. Efficient Learning Machines: Theories, Concepts, and Applications for Engineers and System Designers, 04 2015. [44] R. Girshick. Fast r-cnn. Proceedings of the IEEE International Conference on Computer Vision (ICVV), pages 1440–1448, 12 2015. [45] M. Sewak, M. R. Karim, and P. Pujari. Practical Convolutional Neural Networks: Implement advanced deep learning models using Python. Packt Publishing, 2018. [46] J. Redmon, S. Divvala, R. Girshick, and A. Farhadi. You only look once: Unified, real-time object detection. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2016-Decem:779–788, 2016. [47] W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C. Y. Fu, and A. C. Berg. Ssd: Single shot multibox detector. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 9905 LNCS:21–37, 2016. [48] G. Zaccone. Getting Started with Tensorflow. Packt Publishing, 2016. [49] A. Gulli and S. Pal. Deep Learning with Keras: Implementing deep learning models and neural networks with the power of Python. Packt Publishing, 2017. [50] Python Software Foundation. Python standard library. (online resouce - https://www.python.org/, last accessed on 2020/05/15), 2020. [51] L. Richardson. Beautiful soup documentation. (online resource - https://readthedocs.org/projects/beautiful-soup-4/, last accessed on 2020/05/16), 2019. [52] GitHub. Electron.js. (online resource - https://www.electronjs.org/, last accessed on 2020/05/05), 2013. [53] Fizyr. Keras implementation of retinanet object detection. GitHub, (online resource - https://github.com/fizyr/keras-retinanet, last accessed on 2020/04/24), 2017. [54] D. M. W. Powers. Evaluation: from precision, recall and f-measure to roc, informedness, markedness and correlation. 2011. [55] C. Sammut and G. Webb. Encyclopedia of Machine Learning. Springer US, 2 edition, 2017. [56] H Nord and Eirik Chambe-Eng. Qt. (online resouce - https://www.qt.io/, last accessed on 2020/05/14), 1995. 92 [57] N. Tsoi, J. Gwak, I. Reid, and S. Savarese. Generalized intersection over union: A metric and a loss for bounding box regression. 02 2019. [58] M. Zabir, N. Fazira, Z. Ibrahim, and N. Sabri. Evaluation of pre-trained convolutional neural network models for object recognition. International Journal of Engineering & Technology, 7(3.15):95–98, 2018. 93