Full text
Artificial intelligence on the edge: Real-time object detection embarked on a drone Master Thesis submitted to the Faculty of the Escola T`ecnica d’Enginyeria de Telecomunicaci´o de Barcelona Universitat Polit`ecnica de Catalunya by Sergi Mercad´e Laborda In partial fulfillment of the requirements for the master in Master in Advanced Telecommunications Technologies ENGINEERING Advisor: Anna Calveras Aug´e Advisor: Josep Escrig Escrig Barcelona, October 2021
Contents List of Figures 4 List of Tables 7 1 Introduction 14 1.1 Statementofpurpose.............................. 14 1.2 Requirements and specifications . . . . . . . . . . . . . . . . . . . . . . . . 14 1.3 Methods and procedures . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 1.4 Workplan.................................... 15 1.4.1 Work breakdown structure . . . . . . . . . . . . . . . . . . . . . . . 15 1.4.2 Workpackages ............................. 15 1.4.3 GanttDiagram ............................. 18 1.4.4 Deviations................................ 18 2 Concepts and state of the art of the technology used or applied in this thesis 19 2.1 Brief introduction to AI and ML . . . . . . . . . . . . . . . . . . . . . . . . 19 2.2 Artificial neural networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 2.3 Convolutional neural networks . . . . . . . . . . . . . . . . . . . . . . . . . 25 2.3.1 Object detection using convolutional neural networks . . . . . . . . 27 2.4 YOLOv3..................................... 29 2.4.1 Introduction............................... 29 2.4.2 Architecture............................... 30 2.4.3 YOLOoutput.............................. 31 2.4.4 Lossfunction .............................. 32 2.4.5 Output post-processing . . . . . . . . . . . . . . . . . . . . . . . . . 34 2.5 EfficientAItechniques ............................. 35 2.5.1 Quantization .............................. 35 2.5.2 Pruning ................................. 36 2.5.3 Knowledge distillation . . . . . . . . . . . . . . . . . . . . . . . . . 36 2.5.4 Knowledge Distillation for YOLO: Objectness scaled Distillation . . 37 3 Methodology and project development 39 3.1 Modelarchitecture ............................... 40 3.1.1 Creatingthelayers ........................... 40 3.1.2 Initmethod............................... 41 3.1.3 Forwardmethod ............................ 41 3.1.4 Train step and validation step . . . . . . . . . . . . . . . . . . . . . 42 3.2 Modeldistillation................................ 44 3.2.1 Custom objectness scaled knowledge distillation . . . . . . . . . . . 45 3.2.2 Custom IoU objectness scaled knowledge distillation . . . . . . . . . 45 3.3 Dataset ..................................... 46 3.3.1 Dataset preparation . . . . . . . . . . . . . . . . . . . . . . . . . . 47 3.3.2 Data augmentation . . . . . . . . . . . . . . . . . . . . . . . . . . . 48 2
3.4 Modeltraining ................................. 50 3.5 Metrics...................................... 52 3.5.1 Precision and Recall . . . . . . . . . . . . . . . . . . . . . . . . . . 52 3.5.2 Intersection over Union . . . . . . . . . . . . . . . . . . . . . . . . . 54 3.5.3 Mean average precision . . . . . . . . . . . . . . . . . . . . . . . . . 55 3.5.4 Performance metrics . . . . . . . . . . . . . . . . . . . . . . . . . . 57 4 Results 58 5 Budget 60 6 Conclusions and future development 61 References 62 Appendices 64 A Appendix I: Darknet configuration files 64 A.1 YOLOv3 configuration file . . . . . . . . . . . . . . . . . . . . . . . . . . . 64 A.2 tiny-YOLOv3 configuration file . . . . . . . . . . . . . . . . . . . . . . . . 79 B Appendix II: Layers implementation 83 B.1 Convolutionallayer............................... 83 B.2 Maxpoollayer.................................. 84 B.3 Upsamplelayer ................................. 85 B.4 Routelayer ................................... 86 B.5 Shortcutlayer.................................. 87 B.6 YOLOlayer................................... 88 C Appendix III: Data augmentation pipelines 89 C.1 Train data augmentation pipeline . . . . . . . . . . . . . . . . . . . . . . . 89 C.2 Test data augmentation pipeline . . . . . . . . . . . . . . . . . . . . . . . . 90 D Appendix IV: KD algorithms 91 D.1 Custom objectness scaled distillation . . . . . . . . . . . . . . . . . . . . . 91 D.2 Custom IoU objectness scaled distillation . . . . . . . . . . . . . . . . . . . 93 3
List of Figures 1 Project’s Gantt diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 2 Basic artificial neuron scheme. . . . . . . . . . . . . . . . . . . . . . . . . . 20 3 Output of neuron m, where fis an activation function. . . . . . . . . . . . 20 4 Common activation functions. . . . . . . . . . . . . . . . . . . . . . . . . . 21 5 Common activation function plots. . . . . . . . . . . . . . . . . . . . . . . 21 6 Multi-layer perceptron example. . . . . . . . . . . . . . . . . . . . . . . . . 21 7 Output of layer l, with k+ 1 neurons fully connected to n+ 1 neurons of thepreviouslayer. ............................... 22 8 MNIST inference example using an MLP. The input layer has 784 neurons but has been simplified for representation purposes. . . . . . . . . . . . . . 23 9 Loss function examples. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 10 2D convolution operation example . . . . . . . . . . . . . . . . . . . . . . . 26 11 2D max pooling operation example . . . . . . . . . . . . . . . . . . . . . . 26 12 R-CNN approach. Source:[7] . . . . . . . . . . . . . . . . . . . . . . . . . . 27 13 Fast R-CNN architecture. Source:[8] . . . . . . . . . . . . . . . . . . . . . . 28 14 Faster R-CNN pipeline. Source:[9] . . . . . . . . . . . . . . . . . . . . . . . 28 15 Performance comparison between YOLOv3 and other contemporary object detectors.Source:[12].............................. 29 16 Different grid-sizes and their associated anchor boxes(in blue). Source:[14] . 30 17 Darknet-53 structure. Source:[12] . . . . . . . . . . . . . . . . . . . . . . . 31 18 Outputtensorshape............................... 31 19 First output layer tensor shape. . . . . . . . . . . . . . . . . . . . . . . . . 31 20 YOLOv3 loss function where ˆoi, ˆpi,ˆ biare the predicted objectness, class probability and bounding box coordinates and ogt i,pgt i,bgt iare ground truth values....................................... 32 21 Objectnessloss.................................. 32 22 Classificationloss................................. 33 23 Regressionloss. ................................. 33 24 Total number of bounding boxes. . . . . . . . . . . . . . . . . . . . . . . . 34 25 Basic quantization example. . . . . . . . . . . . . . . . . . . . . . . . . . . 35 26 Finalloss. .................................... 38 27 Some data augmentation transformations available in Albumentations library.Source:[20]................................ 49 28 Precisionformula................................ 52 29 Recallformula.................................. 52 30 Precision and recall visual interpretation. Source: [31] . . . . . . . . . . . . 53 31 IoUformula................................... 54 32 IoU visual interpretation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 54 33 Precision-recall curves. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56 34 Comparison of outputs between the baseline model and the best model. Although the best model is better in most cases, sometimes it fails for images that the baseline model detects properly. . . . . . . . . . . . . . . . 59 4
List of Algorithms 1 NN training algorithm pseudocode . . . . . . . . . . . . . . . . . . . . . . 24 2 Non-max suppression algorithm . . . . . . . . . . . . . . . . . . . . . . . . 34 3 Mean Average precision pseudocode . . . . . . . . . . . . . . . . . . . . . . 55 5
Listings 1 Forwardfunction ................................ 41 2 YOLOv3 training_step function......................... 42 3 tiny-YOLOv3 training_step function...................... 44 4 yolov3.cfg .................................... 64 5 yolov3-tiny.cfg.................................. 79 6 Convolutionalfunction............................. 83 7 Maxpoollayer.................................. 84 8 Upsamplelayer ................................. 85 9 Routelayer ................................... 86 10 Shortcutlayer.................................. 87 11 YOLOlayer................................... 88 12 Data augmentation pipeline used in training. . . . . . . . . . . . . . . . . . 89 13 Data augmentation pipeline used in validation and test stages. . . . . . . . 90 14 objectness scaled distillation implementation. . . . . . . . . . . . . . . . . 91 15 IoU objectness scaled distillation implementation. . . . . . . . . . . . . . . 93 6
List of Tables 1 PASCAL VOC 2012 trainval set and 2007 test set number of images and objectsperclass. ................................ 47 2 Trainval dataset number of images and object distribution. The total number of images is not the sum of cat and dog images as some of them contain instances of both classes. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 3 Test dataset number of images and object distribution. The total number of images is not the sum of cat and dog images as some of them contain instances of both classes. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 4 Hyperparameters used in YOLOv3 trainings. . . . . . . . . . . . . . . . . . 51 5 Hyperparameters used in tiny-YOLOv3 trainings. . . . . . . . . . . . . . . 51 6 Parameters used in the experiments. . . . . . . . . . . . . . . . . . . . . . 51 7 AP calculation example. The detector outputs 10 appearances of the target class, while in the dataset we have only 5 appearances in total. . . . . . . . 56 8 Experimentalresults............................... 58 9 Estimated hardware budget. . . . . . . . . . . . . . . . . . . . . . . . . . . 60 10 Estimated salary budget. . . . . . . . . . . . . . . . . . . . . . . . . . . . 60 11 Finalestimatedbudget............................. 60 7
Revision history and approval record Revision Date Purpose 0 15/02/2021 Document creation 1 19/08/2021 Document revision DOCUMENT DISTRIBUTION LIST Name e-mail Sergi Mercad´e Laborda Anna Calveras Aug´e Josep Escrig Escrig Written by: Reviewed and approved by: Date 19/08/2021 Date 20/08/2021 Name Sergi Mercad´e Laborda Name Anna Calveras Aug´e Position Project Author Position Project Supervisor 8
Abstract The integration of artificial intelligence (AI) in a drone belongs to the field known as AI on the edge. AI on the edge applications have more restricted requirements regarding the memory, energy consumption and computational power available with respect to a conventional deployment. This causes the application of optimization techniques to achieve the requirements, that usually have a negative impact on the performance of the model. The aim of this project consists in researching and implementing techniques known as ”Knowledge Distillation” (KD) applied to models that follow YOLO architectures, that reduce the negative effect of the optimizations. In KD, one or more networks act as teachers and help a student network during its training stage to obtain better results. YOLOv3 and tiny-YOLOv3 networks have been successfully implemented from scratch and different KD techniques have been developed and their effects on the training studied. The preliminary results obtained validate these techniques to obtain better performance on optimized edge models. 9
•Dates: None. WP2: Dataset •WP ref: 2 •Major constituent: Research, SW •Description: The purpose of this work package is to identify and collect open datasets available online for the task of object detection in research. The annotations of the collected datasets must be transformed to be used in the following work packages. •Tasks: –Internal task T1: Dataset research –Internal task T2: Dataset preparation •Planned start date: Week 4. •Planned end date: Week 6. •Start event: End of WP1: Research. •End event: Internal task T2: Dataset preparation complete. •Deliverables: Ready to use datasets. •Dates: WP3: AI development, training and deployment •WP ref: 3 •Major constituent: Research, SW •Description: The purpose of this work package is to develop from the scratch teacher and student neural networks models following YOLO-like architectures. These networks must be trained with a dataset. The knowledge distillation algorithms must be implemented and validated in this WP. •Tasks: –Internal task T1: CNN development. –Internal task T2: CNN testing and validation. –Internal task T3: KD algorithms development. •Planned start date: Week 6. •Planned end date: Week 18. •Start event: End of WP2: Dataset. •End event: Internal task T3: KD algorithms development complete. •Deliverables: CNN object detection code. 16
•Dates: WP4: Validation •WP ref: 4 •Major constituent: SW •Description: The impact of the knowledge distillation techniques in the training of the student network must be assessed and validated experimentally. •Tasks: –Internal task T1: Experimental validation of KD algorithms. •Planned start date: Week 18. •Planned end date: Week 24. •Start event: End of WP3: WP3: AI development, training and deployment. •End event: End of the project. •Deliverables: Experimental results. •Dates: WP1: Final documentation •WP ref: 5 •Major constituent: Documentation •Description: The methodology and the information of the validation experiments must be retrieved and documented. •Tasks: –Internal task T1: Final documentation. •Planned start date: Week 21. •Planned end date: Week 24. •Start event: First results of WP 4: Validation. •End event: End of the project. •Deliverables: Project documentation. •Dates: 17
1.4.3 Gantt Diagram Figure 1: Gantt diagram of the project 1.4.4 Deviations This project has experienced strong deviations in several stages and WPs: •The modification of the scope of the thesis to be more research-oriented. Instead of performing a PoC deployment on the drone itself and documenting it, this thesis became a research of Knowledge Distillation techniques for edge deployments. •Deviations in the drone project that motivated this thesis have resulted in a reduction of my availability to dedicate hours to this project. •Implementing such a complex neural network is not a trivial task, and debugging networks like this is very difficult. Debugging problems in the implementation stage caused an important delay. •The iterative nature of a machine learning project makes them difficult to dimension. •Training a complex neural network takes a lot of time that is difficult to know in advance. It depends on the hardware used in the training stage, the hyperparameters... 18
2 Concepts and state of the art of the technology used or applied in this thesis 2.1 Brief introduction to AI and ML Artificial Intelligence (AI) and Machine Learning (ML) are present in our everyday lives. They have become a natural part of our daily workflow and are used in many fields to push forward the state of the art and open the door to applications that seemed impossible just a few years back. ML is a branch of AI that studies and develops techniques that allow a machine to learn how to perform a certain task from data. The first formal definition of ML was introduced in 1959 by Arthur Samuel in [1]. Tom M. Mitchell[2] proposed a definition that nowadays is considered the most straightforward: ”A computer program is said to learn from experience E with respect to some class of tasks T and performance measure P if its performance at tasks in T, as measured by P, improves with experience E”. There are different strategies for ML algorithms. The main approaches are supervised and unsupervised algorithms. Supervised learning algorithms are capable of learning from inputs and outputs pairs in order to produce outputs for previously unseen data. Unsupervised learning algorithms learn relations and patterns on the provided data. These learned features can be later used in supervised algorithms, to compress data or in anomaly detection among other applications. 2.2 Artificial neural networks Artificial neural networks (NN) are computational models structured in connected layers that can solve complex problems by learning relations from the input data. NN can be used in supervised and unsupervised learning. Although the first artificial neuron model was proposed in 1943 by Warren S. McCulloch and Walter Pitts[3], the high computational cost and memory requirements implied in training and using this approach limited its usage. At the end of the last century, AI and NN, in particular, became practical technologies. Nowadays, with the democratization of computing power and memory at a reasonable cost and the usage of cloud computing, we are living a fast-paced revolution in the AI field that has an impact on all levels of our society, economy and all scientific fields. The basic unit of NNs is called the artificial neuron and mimics the behavior of a biological neuron. A basic artificial neuron is composed of 4 different elements: •Inputs: The inputs are numerical values that can come from sensors, features or other neurons. •Weights: There is one weight for each input. They multiply the input value, to let the neuron decide which inputs are important and which are not. •Bias: The bias is used to modulate the output of the neuron independently of the inputs. 19
•Activation function: Used to produce the output of the neuron. Activation functions normally are non-linear functions to make the neuron able to learn non-linear relations between the inputs and the outputs. An artificial neuron has at least one input and one output. A simplified scheme and the mathematic expression of an artificial neuron can be seen in figure 2 and figure 3 respectively. Figure 2: Basic artificial neuron scheme. ym=f x0. . . xn∗ wm0 . . . wmn +bm Figure 3: Output of neuron m, where fis an activation function. Activation functions are usually non-linear functions: •A non-linear function allows the neuron to solve more complex problems. This also reduces the number of neurons needed in NN. When the activation function is non-linear, a two-layer neural network can be proven to be a universal function approximator. •If the activation function is linear, the gradient is constant and this can be a problem for the gradient descent discussed later in this section. •The composition of linear functions can be expressed as a unique linear function, so NN could be expressed in a single linear expression, losing the ability to stack layers and take advantage of complex architectures. Activation functions don’t have to be continuously differentiable but is desirable. Some common activation functions are: •Sigmoid function: Widely used for classification problems. •Hyperbolic tangent function (tanh): A scaled version of the sigmoid function but with a stronger gradient. •Rectified Linear Unit function (Relu): One of the most popular activation functions, considered the default approach. Is more computationally efficient than the other two. 20
y=1 1 + e−x (a) Sigmoid function. y=ex−e−x ex+e−x (b) Tanh function. y=max(0, x) (c) Relu function. Figure 4: Common activation functions. (a) Sigmoid function. (b) Tanh function. (c) Relu function. Figure 5: Common activation function plots. The weights and bias are the variable elements of an artificial neuron and give the neuron the ability to solve simple classification and linear regression problems by adjusting these elements. For more complex problems, we can connect different artificial neurons in arbitrary architectures called neural networks. The most basic NN is called multi-layer perceptron (MLP). In MLP, artificial neurons are grouped in layers and these layers are stacked sequentially as seen in figure 6. Figure 6: Multi-layer perceptron example. 21
The MLP is composed of fully connected layers, where each neuron has a connection with all the neurons of the previous layer. The first layer is called the input layer, that acts as a placeholder for the input data. The layers in between are called hidden layers. The last layer is called the output layer, where we obtain the result of the inference. Following the expression in figure 3 and considering that in an MLP, all layers are fully connected, we can express the output of a layer of the network, l, with k+ 1 neurons fully connected to a previous layer of n+ 1 neurons, as in figure 7. ol=f wl,0,0. . . wl,0,n . . ..... . . wl,k,0. . . wl,k,n ol−1,0 . . . ol−1,n + bl,0 . . . bl,k Figure 7: Output of layer l, with k+ 1 neurons fully connected to n+ 1 neurons of the previous layer. Generalizing this expression, we can see the NN as a big complex function composed of a chain of expressions like the one in figure 7. The weights and bias of each artificial neuron are parameters of the function that can be adjusted, and the number of layers and the number of neurons for each layer determine the number of these parameters. Although the MLP architecture is very simple, it is so powerful and flexible that a similar structure is capable to solve the hand-written digit recognition problem with good results, considered the ”Hello World” problem in the AI field. In the hand-written digit recognition problem, the system must recognize hand-written digits from 0 to 9 in images of 28x28 pixels. It is not a trivial task with traditional computer vision algorithms, but using NNs we can approach this problem in a very straightforward manner. The images can be seen as matrices of 28x28, where each value is the level of the normalized brightness of the corresponding pixel. These matrices can be flattened into vectors and introduced in the input layer of a neural network, that must contain 784 neurons in the input layer, one for each pixel. After running the inference, we obtain at the output layer the probability of the correspondent number to appear in the image, so the output layer must have 10 artificial neurons, one for each possible result. To work as expected for the task to be solved, the weights and biases of a NN must be adjusted. The original implementation of an artificial network did not include an algorithm of how to properly adjust these parameters for an arbitrary task. Although some predecessors appeared in the 1960s and similar approaches were implemented in the 1970s, it wasn’t until 1986 when Rumelhart, Hinton and Williams proposed the term backpropagation for use in arbitrary NNs in [4], starting a new era in AI. The backpropagation algorithm allows a NN to learn the best parameters to perform an arbitrary task, only by seeing input-output examples in a process called training. The training stage is what makes a NN an ML algorithm and allow us to take advantage of the plasticity of its structure. To apply the backpropagation algorithm first we have to define a loss function. The loss 22
Figure 8: MNIST inference example using an MLP. The input layer has 784 neurons but has been simplified for representation purposes. function assesses the performance of the NN on a single input by comparing the actual output ˆyiwith the expected output yi. Its expression variates depending on the task that the NN is performing, but the loss function has to be differentiable. The mean square error loss function(figure 9a) and the binary cross-entropy loss function(figure 9b) are examples of common loss functions used in regression and binary classification problems respectively. Thanks to the backpropagation algorithm, we can compute the gradients of the loss function with respect to the NN parameters, the weights and the biases of the network. We can use a process called gradient descent to optimize the network parameters, allowing us to adjust them for the target task and reducing the overall loss. The gradient descent does not guarantee to find a global minimum, only a local minimum. In the training stage, we find the local minimum of the cost function through the analysis of its gradient and recalculate the weights and bias based on the results obtained from the gradient descent on the examples. This process is performed iteratively until the local minimum is found and then the obtained values of weights and biases are locked to allow the model to run inferences. 23
L(y, ˆy) = 1 N N X i=1 (yi−ˆyi)2 (a) Mean Square Error loss function. L(y, ˆy) = −1 N N X i=1 yilog(ˆyi)+(1−yi)log(1−ˆyi) (b) Binary cross-entropy loss function. Figure 9: Loss function examples. Algorithm 1 NN training algorithm pseudocode 1: procedure NN train(N, batchsize, dataset, model, optimizer)1 2: for Nepochs do 3: Split dataset in random minibatches of size batchsize 4: for every minibatch do 5: for sample in minibatch do 6: Evaluate model(forward) 7: Compute loss 8: Accumulate gradient of loss(backward) 9: Update model parameters with accumulated gradient using optimizer 10: return model All the parameters of the network that are not learned during the training stage (number of layers, activation function, number of neurons for each layer...) and the parameters related to the training process(batch size, optimizer algorithm, epochs...) are called hyperparameters. These hyperparameters have to be set up by humans, normally using a trial-error approach in a set of experiments. 1Based on figure 7.12.C from Deep Learning with Pytorch[6] 24
2.3 Convolutional neural networks As we have seen in the previous section, neural networks can handle images as input in a very straightforward way. For grayscale images, we can use as input the level of the normalized brightness of the corresponding pixel. For colour images, we can use the normalized pixel value for each channel. In the MNIST example, for small images of 28x28 pixels, the input layer has to have 784 placeholders. This means that for the next layer, each neuron has to have 784 weights plus its own bias. For larger images, these parameters easily increase and even more if the image has three channels. This makes this approach not very scalable. Additionally, the MLP architecture doesn’t naturally capture the local patterns of the images. A better approach is needed to deal with images. Convolutional neural networks (CNNs) are NNs that use specific kind of layers called convolutional layers and pooling layers. These NN can be used in applications where the input data has some kind of spatial structure. The most common usage of these layers is in the computer vision field. The first CNN, LeNet[5], was introduced in 1998, and since there, CNN revolutionized the computer vision field. CNN can be used in multiple tasks in the computer vision domain, some examples: •Classification: CNN is used to classify the content of an image between different classes. •Object detection: CNN is used to detect, locate and classify objects in an image. •Image segmentation: CNN classifies pixels of an image between different classes. Using this approach, we can obtain the shapes of objects of different classes in an image. •Object detection and instance segmentation: CNN not only classifies and locates multiple objects in an image, but also identifies their shape. The problem this project approaches is object detection. Convolutional layers perform a convolutional operation in a tensor using a kernel or filter of learnable weights. In a convolution, the scalar product between a kernel and a patch of the target tensor is computed in all channels of the tensor. An example of this operation can be seen in figure 10. By applying this operation, we achieve features that are translational invariant and take into account local patterns. These convolutions can be performed in 1 dimension (useful when working with sequential data), 2 dimensions (useful when working on images) and 3 dimensions (useful when working on spatial data or videos). The explanation will focus on the 2-dimensional case, as this project deals with object detection in images. The output of a convolutional layer is always an activation map or feature map, a tensor of one dimension that contains the results of the convolution between the input tensor and the learnable kernel of the layer. A convolutional layer composed of several kernels is like a filterbank with filters that are learned during the training stage to optimize the output of the network. This allows the convolutional layer to extract relevant features for the target task. 25
To recover the real bounding box coordinates and shape, we have to perform some operations on the tensor values assigned to this end. To compute the centre X, Y values we have to apply a sigmoid function, which makes these values go between 0 and 1. To compute the width and the height of the bounding box we have to calculate the exponent of the tensor value for the width and the height and multiply it by the width and the height of the associated anchor. 2.4.4 Loss function As we have seen in previous sections, the loss functions are a key element during the NN training. In some way, it models the logic behind the expected output for a given input. Object detection is not a trivial task, as it involves different subtasks like object classification and localization, so the loss function is a bit complex having several losses, one for each task. The loss function of the YOLO evolved during the successive versions of the network. The first major change was introduced in YOLOv2 when the model became all convolutional and the anchor concept was introduced for the first time. The second major change was performed in YOLOv3 when the objectness score prediction was introduced and a featurepyramid like approach to predict boxes at multiple scales was applied to improve the detection of small objects. The loss function of the following versions are very similar to the one used in YOLOv3. In YOLOv3, for each anchor box in each cell and scale the network produces the following outputs: objectness score,bounding box coordinates and shape and the class prediction. For each kind of output we need to define a particular loss function: regression loss for the bounding box predictions, objectness loss and classification loss. A general view of the loss function can be seen in figure 20. LY OLOv3=fobj(ˆoi, ogt i) + fcl(ˆpi, pgt i) + fbb(ˆ bi, bgt i) Figure 20: YOLOv3 loss function where ˆoi, ˆpi,ˆ biare the predicted objectness, class probability and bounding box coordinates and ogt i,pgt i,bgt iare ground truth values. fobj(ˆoi, ogt i) = λobj S2 X i=0 B X j=0 1obj ij (ogt i−ˆoi)2+λnoobj S2 X i=0 B X j=0 1noobj ij (ogt i−ˆoi)2 Figure 21: Objectness loss. 32
fcl(ˆpi, pgt i) = S2 X i=0 1obj iX c∈classes (pgt i(c)−ˆpi(c))2 Figure 22: Classification loss. fbb(ˆ bi, bgt i) = λcoord S2 X i=0 B X j=0 1obj i[(xi−ˆxi)2+(yi−ˆyi)2]+λcoord S2 X i=0 B X j=0 1obj i[(√wi−pˆwi)2+(phi−qˆ hi)2] Figure 23: Regression loss. 33
2.4.5 Output post-processing As we have seen in 2.4.3, one inference of an image produces A∗G∗Gpossible bounding boxes for each YOLO layer, where Ais the number of anchors and Gis the grid size. This means for a default YOLO network we obtain a total of 10647 bounding boxes for an image, as seen in figure 24. total bbox = 13 ∗13 ∗3 + 26 ∗26 ∗3 + 52 ∗52 ∗3 = 10647 Figure 24: Total number of bounding boxes. A lot of these bounding boxes contain background detections and only a few contain relevant detections of the objects in the image. We have to perform post-processing techniques to discard background detections and try to take only one bounding box per object. The first step is to filter the bounding boxes using a minimum threshold for their objectiveness values. Background detections should have a low objectiveness value while actual object detections should be high. A typical value used as the minimum threshold is 0.2. The inference of YOLO networks can produce several duplicates of the same detected object. To avoid these duplicates, an algorithm to remove duplicates with low confidence called Non-Maximum Suppression (NMS) is applied. This algorithm uses the metric Intersection Over Union (IoU) discussed in detail in 3.5.2. This algorithm is not perfect and it can discard valid detections if these are very close or overlapping. Algorithm 2 Non-max suppression algorithm 1: procedure NMS(bboxes, threshold) 2: Create an empty list called bNMS. 3: Sort predictions of bboxes by confidence score. 4: for bbox in bboxes do 5: for other bbox in bboxes −{bbox}do 6: if class id from bbox == class id from other bbox then 7: Compute the IoU between bbox and other bbox. 8: if IoU > threshold then 9: Remove other bbox from bboxes 10: Add bbox to bNMS and remove it from bboxes 11: return bNMS 34
2.5 Efficient AI techniques There are several motivations behind the usage of these techniques: reduce inference speed, reduce memory and resource utilization, reduce power consumption, security and/or cost concerns. . . Although efficient techniques can be applied in different environments, they usually imply a cost in terms of algorithm performance. This is the main reason why these techniques are usually relegated to edge applications, where the devices are normally more constrained. Several kinds of strategies are followed in order to deploy NN in constrained devices: quantization, pruning and knowledge distillation among others. 2.5.1 Quantization Normally, the parameters of a NN are stored in float32 variables to perform operations with good accuracy. Parameter quantization is a technique that reduces the bits needed to represent these parameters and as a result, the memory footprint and the complexity of the involved operations are reduced. As this technique reduces the precision of the parameters and the operations, it can cause an accuracy drop in the NN. Figure 25: Basic quantization example. The initial approach of this technique applied the same quantization to all the NN parameters, but deep research of how quantization affects the NN performance have shown that customized quantization for different types of parameters and layers can reduce the memory footprint and complexity without producing an important accuracy drop. The NN edge devices have usually specialized hardware to compute integer operations, which further increase the speed boost of this technique. There are two main methods to apply quantization to a NN: •Post-training quantization: The NN is trained normally. Once the NN is trained, a quantization algorithm is applied. This approach can cause an accuracy drop. 35
•Quantization aware training: The network parameters are fake quantized during the training stage. The accuracy drop in this approach is lower than in post-training quantization. 2.5.2 Pruning It has been demonstrated that not all network capacity contributes to the final output. This means that there are some parameters of the NN that are not relevant to the output and are redundant. They can be removed from the network without having a major impact on accuracy. This process is called pruning. By eliminating these redundant parameters we reduce the size of the network and we can improve its inference speed but, but as quantization, it can cause a drop in terms of accuracy. We can prune the network in two different ways: •Weight pruning: Some weights of the network are discarded, eliminating some connections between neurons. This makes the network sparse and decreases the number of parameters without changing the architecture. Weight pruning is easier to perform without a major impact on the accuracy of the network, although to boost the inference speed, the resulting network has to be deployed in an environment ready for sparse computation. •Neuron pruning: In neuron pruning, some neurons are removed from the network. This reduces the size of the network while allowing dense computations. This method has more impact on the NN accuracy performance as we are removing more parameters. There are two pruning strategies: •Pruning after training: In this approach, scores are assigned to the network parameters after training and the ones with the lowest scores are removed from the network. This process does not have an advantage during the training stage, only in the inference stage. To reduce the impact on the accuracy, this process can be performed iteratively while training the model. After some neurons are removed from the network, a fine-tuning stage is performed before evaluating the network again. •Pruning before training: In this method, the network is pruned before the training stage, decreasing the network’s capacity from the beginning but reducing the penalization in the final test accuracy. 2.5.3 Knowledge distillation In knowledge distillation (KD), a large model or an ensemble of models, known as teachers, are used to train a smaller model, called student, in order to improve its performance. The first approach to this idea was proposed in 2006 in the paper Model Compression [15], where the knowledge of an ensemble of models was compressed into a single model. It wasn’t until 2015 when the concept ”Knowledge distillation” was proposed by Oriol Vinyals and his team at Google in [16] and gained popularity. 36
The advantage of using KD on AI on the edge deployments is straightforward, as it allows to compress the knowledge of models that are too expensive in computational power or storage requirements, in a smaller model available to deploy in real-time applications in constrained devices. There are several strategies when applying knowledge transfer in the training stage, widely described in [17]: 1. Response-Based knowledge: The most straightforward is the response-based knowledge distillation. In this strategy, the output of the last layer of the teacher network or ensemble is used during the training stage of the student. The student can try to mimic the output teacher model or it can use it alongside its own output. 2. Feature-Based knowledge: The values of the intermediate layers like feature maps can also be used as hints to help the student network in the training stage. The basic idea is to force the matching of the feature activations maps between the teacher and the student. 3. Relation-Based knowledge: This approach tries to exploit the relationship between different layers or data as a source of knowledge. There are also different distillation schemes. The most common ones are: 1. Offline distillation: In this distillation approach, the knowledge of a pre-trained teacher is transferred to a student model. This means that, first, we have to train the teacher model and then transfer the knowledge to the student network in a different training stage. This is the most common distillation approach as it is easy to implement. 2. Online distillation: In this distillation scheme, the teacher model and the student model are trained and updated simultaneously in only one training stage. This scheme is relatively new and it allows also deep mutual learning, a training paradigm where multiple NNs work in collaboration during the training stage and any network can be the student and the others can be the teachers. 2.5.4 Knowledge Distillation for YOLO: Objectness scaled Distillation At the time this thesis was started, there wasn’t a lot of distillation approaches to perform knowledge distillation on object detectors. One of the first proposed methods for knowledge distillation proposed for YOLO was introduced in 2018 in [18]. The authors called this method objectness scaled distillation, as it uses the objectness score of the detections to weight their share on the final distillation loss A basic approach to performing object distillation in object detectors is an offline distillation using response-based knowledge: this implies using the last convolutional layer of the teacher network for transferring the knowledge. The authors encountered that this approach wasn’t optimum for YOLO, as it predicts a dense set of candidates that include detections in the background. These background detections are filtered and ignored in the inference stage, as we have seen in section 2.4.5, 37
but would be transferred to the student network at the training stage, making the student network learn irrelevant and erroneous background bounding boxes. To avoid this, the authors propose the objectness scaled distillation approach. The teacher networks assigns a low objectness value to the background candidates, while it gives a high value to the candidates that appear more to relevant objects. This method only transfers the knowledge of the detections of the teacher network that have a large objectness value. This is achieved by multiplying the loss between the detection of the student and the detections of the teacher by the teacher’s objectiveness score. The computation of the losses does not change. This approach also adds a new hyperparameter, called λD, that modulates the contribution of the distillation algorithm to the final loss. Lfinal =LY OLO +λD∗Ldistillation = fobj(ˆoi, ogt i) + fcl(ˆpi, pgt i) + fbb(ˆ bi, bgt i) + λD∗(fobj(ˆoi, oT i) + oT i∗fcl(ˆpi, pT i) + oT i∗fbb(ˆ bi, bT i)) Figure 26: Final loss. As discussed in 2.4.5, several cells can predict the same instance of an object in the image. The NMS algorithm is applied after the inference of the network to filter these duplicates and take only the best prediction. But this filtering is applied after the inference, so when the knowledge of the teacher is transferred to the student network, the duplicates are transferred too, resulting in redundant information that can cause over-fitting in the student network. To overcome this issue, the authors propose Feature Map-NMS (FM-NMS). If there are several candidates of the same class in a grid of KxK cells in the output tensor, they consider it highly likely to be of the same object and only take the candidate with the higher objectness value. The Kvalue used is another hyperparameter added in this approach, in [18] they use a grid of 3x3. The usage of the objectness scaled knowledge distillation allows training the student network even when there is no labelled data, using only the knowledge of the teacher network. With this approach, the teacher part of the loss can be used if there are not ground-truth data available, or use the combination seen in figure 26 when the groundtruth is available. With this technique, the authors achieve a mAP2increasing of several points with respect to the student baseline in their experiments. 2The mAP metric is discussed in detail in section 3.5.3. 38
3 Methodology and project development The main objective of this thesis is to develop knowledge distillation algorithms in Pytorch Lightning and compare their performance against the vanilla approach. This requires: •A teacher and student network: YOLOv3 has been chosen as the teacher network while tiny-YOLOv3 will be the student network. These models have been developed from scratch in Pytorch Lightning, using a methodology described in 3.1. •Knowledge distillation approach: knowledge distillation algorithms are required in order to compress the knowledge from the teacher into the student network. These are described in section 3.2 with their correspondent implementations. •A dataset: A set of images with their correspondent annotations is required to train the models. The dataset used in the experiments, the annotations format and the data augmentation pipelines are discussed in 3.3. •Metrics: A set of metrics is required to empirically evaluate the performance between the models trained without KD and KD. The metrics used are explained in 3.5. The algorithms of this thesis have been developed in Python, using the Pytorch Lightning[19] research framework to develop the models, manage the training stage and track the experiments. Pytorch Lightning3has been chosen as it simplifies the code structure, it makes easier to automatically track the experiments and allows seamlessly training and deployment in different hardware among other advanced usages. Other well known Python libraries have been used, like Albumentations4[21] to apply online data augmentation during the training stage, Pandas5[22][23] to handle the dataset and Pillow6[24] to process the dataset’s images. The code used in this thesis is inspired, taken or modified from several sources: 1. PyTorch Computer Vision Cookbook[25]: This book explains without much detail how to implement a YOLOv3 network. Although the code is not very adaptable, it gave me a general example of how the YOLO network is structured in a Pytorch code. 2. Online blogs[26]: Although some of them are based on the example provided in PyTorch Computer Vision Cookbook, these blogs provide details on how to structure the forward pass and details on how to implement the layers. These ideas have been implemented with adaptations to make the code more modular. 3. YOLOv3 GitHub repository[27]: The main structure of the code of this thesis is based on the content of this open-source repository, although strong modifications, corrections and adaptations have been performed. This repository is accompanied 3https://www.pytorchlightning.ai/ 4https://albumentations.ai/ 5https://pandas.pydata.org/ 6https://python-pillow.org/ 39
by a very useful video explanation of how it is implemented and the logic behind the code. 3.1 Model architecture The first stage is to implement the YOLOv3 model and the tiny-YOLOv3 model. The Pytorch-Lightning framework defines a model as a special class called LightningModule . This allows structuring the code in a self-contained environment, making it easier to develop, test and read. A LightningModule has to have some methods defined to work: •init: where the model architecture is defined. •training_step: defines the behavior of a training step of the model for a given input. •configure_optimizers: defines the optimizers to be used during the training stage. The user can define other methods related to the validation and test stages. For each model, additional LightningModule methods have been overridden: •validation_step: defines the behavior for a given input in the validation step. •forward: method to define the behavior of the model in the inference time. Additional functions have been also implemented inside the LightningModule, like the loss function or auxiliary functions. The usage of these methods also allows leaving the boilerplate code to the Pytorch-Lightning framework in a similar way as the Keras framework does. For example, the Pytorch-Lightning framework takes care of the training and validation loops for us. 3.1.1 Creating the layers Before even starting the definition of the architecture of the YOLO models in the init method, the different layers that are the basic blocks that compose these architectures must be defined and implemented. The original authors have published a framework called Darknet[28] to train YOLO based networks and run inferences. This framework uses special plain text files called configuration files that have codified how the networks are structured, and some hyperparameters used in the training stage of the networks. For this project, the configuration file of YOLOv3 and the configuration of tiny-YOLOv3 are downloaded from the official repository. These files can be seen on appendix A. After analyzing these files, several types of layers are identified: •Convolutional layer. •Max pooling layer. •Upsample layer. Only used in tiny-YOLOv3. •Route layer. •Shortcut layer. 40
•YOLO layer. Acts as an output layer, it reshapes the tensor. Each layer type class is implemented in a separate python file named modules.py as independent PyTorch nn.Modules, where its initialization procedure and forward pass is defined. The basic code structure is taken from [27] and [26] although strong modifications have been performed in order to adapt the code to a more modular and self-contained approach. The complete implementations of the layers can be seen in appendix B. A special mention to the Upsample layer, which is only used in the tiny-YOLOv3 architecture. In this layer, a difference between the padding algorithms of Darknet and PyTorch create an incompatibility issue with the output dimensions that makes the concatenation operations fail. A patch to this is to apply manual padding, as seen in several forums. 3.1.2 Init method As we have seen before, the init method is where all the architecture of the model is defined. To make the code of this thesis reusable and modular, the first step is to make a code that is capable to read the information of the configuration files and build the network described in them. To implement this, a configuration file parser function and a constructor that take advantage of the modular approach are required. The parsing of the configuration file is implemented in a function called parse_cfg. This function is taken from [26], although a modification in the return statement is performed to return a list of dictionaries. To build the network structure, a function called _create_layers is implemented inside the LightningModule. This function takes the list of dictionaries created by parse_cfg and builds the network described by creating the layers with the parameters specified in each dictionary using the nn.Modules described in modules.py as a base. As a result, we obtain a module list that represents the network described in the target configuration file. Additional parameters like the anchors, other hyperparameters or the losses are initiated in this method as well. 3.1.3 Forward method The forward function must be a general function capable of being valid for different YOLO based networks. Taking advantage of the modular approach in the network construction and that each layer knows how to perform its own forward pass, the following forward pass function is implemented. It iterates for each layer and delegates to them how to perform its own forward operation. It also stores the outputs of each one of the layers that use outputs from other layers different from the previous one, like route layers or shortcut layers. 1def forward ( self , x): 2 3outputs = [] # for each scale 4layer_outputs = [] 5 6for layer in self . layers : 41
–Information about the truncation of the object, if any. –Information about the detection difficulty of the object. –Top left and bottom right corners coordinate in pixels of the bounding box. These annotations must be converted to the YOLO format, where the annotations are plain text files that contain in each line a bounding box annotation with the following format: •Class id of the object. •Top left corner X pixel coordinate. •Top left corner Y pixel coordinate. •Bottom right corner X pixel coordinate. •Bottom right corner Y pixel coordinate. •Bounding box width. •Bounding box height. All the values of the bounding box annotations are relative to the image dimensions. An script[30] that translates the annotations from the VOC format to the YOLO format has been used. This script has also been modified to being capable of filter classes, allowing us to create subsets of the main datasets with only a few classes. 3.3.2 Data augmentation Data augmentation is a technique commonly used in the computer vision field, used as a regularizer to help to avoid the overfitting of the algorithms and help these to better generalize to new images. It consists in applying one or more transformations to an image to change its appearance and change accordingly the annotations. By using this technique we can have sightly different images from just only one source image, increasing the number of images that the network sees during training. These transformations can be rotations, shifts, changes in the colour channels, crops, affine transformations... There are two types of data augmentation: •Online data augmentation: Applied during the training stage each time an image is going to pass through the network. After passing the network the augmented image and annotations are discarded. •Offline data augmentation: Applied only once before the training. The new generated images and annotations are added to the original dataset, increasing the number of samples. Online data augmentation is preferable as it increases the variety of the images because a new data augmentation is applied each time an image enters the network, allowing almost 48
an infinite number of augmented images while saving space in the computer by avoiding storing the augmented images. This project uses online data augmentation during the training stage implemented with the library Albumentations. Figure 27: Some data augmentation transformations available in Albumentations library. Source: [20] This library allows building pipelines with the possible transformations that an image might have. The parameters of these transformations and their probability can be set, and when an image enters the pipeline, the image goes for each step randomly activating the transformations with the probabilities set before. The pipelines used in the training dataset and the validation and test datasets can be seen in appendix C. These pipelines are almost the same as [27]. 49
3.4 Model training The training stage is one of the most critical aspects of machine learning algorithms deployments. The validation of the KD algorithms is performed through the experimental part of this thesis, which requires the following trainings: •A baseline YOLOv3 model to act as the teacher. •A baseline tiny-YOLOv3 model to act as a reference for the KD algorithms. •Student tiny-YOLOv3 models trained with the KD algorithms. Training a model for an object detection task is not straightforward, especially for large and complex models like YOLO-like networks. As seen in the previous section, the loss function to minimize is very complex, and training a YOLOv3 from scratch in a mid-size or large dataset can take several days or even weeks. As the objective of this thesis is to validate the impact of the KD algorithms in the training stage of the student networks, and to do so several experiments are required. To be able to perform the experiments in the thesis time window, a small proof-of-concept dataset has been used. A subset of PASCAL VOC dataset containing only the samples of classes ”cat” and ”dog” has been created. The trainval set is shown in table 2 and test set is shown in 3. Only the test set has been used as a dataset in the training and validation steps for the experiments and no data augmentation has been applied to make it easier for the models to learn the features of the images. Class Images Objects Cat 1128 1277 Dog 1340 1597 Total 2433 2974 Table 2: Trainval dataset number of images and object distribution. The total number of images is not the sum of cat and dog images as some of them contain instances of both classes. Class Images Objects Cat 332 370 Dog 432 529 Total 748 899 Table 3: Test dataset number of images and object distribution. The total number of images is not the sum of cat and dog images as some of them contain instances of both classes. Multiple experiments have been performed to find the optimum hyperparameters to obtain the best results in the teacher and student models. The final hyperparameters used as a base in YOLOv3 and tiny-YOLOv3 models can be seen in tables 4 and 5 respectively. 50
Hyperparameter Value Batch-size 64 Optimizer Adam Learning rate 0.0001 Weight decay 0.0005 Data augmentation None Table 4: Hyperparameters used in YOLOv3 trainings. Hyperparameter Value Batch-size 64 Optimizer Adam Learning rate 0.001 Weight decay 0.0005 Data augmentation None Table 5: Hyperparameters used in tiny-YOLOv3 trainings. For evaluating the impact of the knowledge distillation algorithms, several experiments were performed while sweeping the values for the minimum threshold used in each knowledge distillation algorithms. To compare the different experiments correctly, a seed was used to eliminate most of the randomness in the training stage. The YOLOv3 was trained during 500 epochs and the tiny-YOLOv3 models were trained for 50 epochs. The values used in the experiments can be seen in table 6 Parameters Values KD [None, objectness, iou] KD threshold [0, 0.25, 0.5, 0.75, 0.95] Table 6: Parameters used in the experiments. 51
3.5 Metrics 3.5.1 Precision and Recall Precision and recall metrics are well known in the ML field. In this thesis, they are used as an intermediate metric to compute the Mean Average Precision (mAP). These metrics are computed in a binary classification problem, where a classifier must detect samples of a given class, for example a classifier that detects sick people from healthy people. In this scenario we have four possible outcomes: •True positives (TP): A positive sample classified as positive. Sick people detected as sick. •True negatives (TN): A negative sample classified as negative. Healthy people detected as healthy. •False positives (FP): A negative sample classified as positive. Healthy people classified as sick. •False negatives (FN): A positive sample classified as negative. Sick people classified as healthy. Precision evaluates how many of the samples detected as positive are actually positive samples. Precision =T P TP +FP Figure 28: Precision formula Recall evaluates how many positive samples have been classified as positive by the classificator. Recall =TP TP +FN Figure 29: Recall formula The figure 307shows a visual interpretation of both metrics. 52
Figure 30: Precision and recall visual interpretation. Source: [31] 7By Walber - Own work, CC BY-SA 4.0, https://commons.wikimedia.org/w/index.php?curid=36926283 53
3.5.2 Intersection over Union Intersection over Union (IoU) is a metric that measures the overlap between two bounding boxes. It is computed by dividing the area of overlap between the area of the union of two bounding boxes. It is normally used to compare a bounding box of a prediction with a bounding box from the ground truth. IoU =Intersection area Union area Figure 31: IoU formula The figure 32 shows a visual interpretation of the IoU metric. Figure 32: IoU visual interpretation. This is an intermediate metric used in several other metrics like Mean Average Precision (mAP). 54
3.5.3 Mean average precision This is one of the main metrics used to evaluate the performance of detection algorithms. This metric measures the average performance in all the classes in a dataset and can be used to compare the performance of different models in a specific dataset. As the name states, this metric is the average among all the classes of a metric called Average Precision (AP). The algorithm to compute the AP for each class and then compute the mAP is the following: Algorithm 3 Mean Average precision pseudocode 1: procedure MAP(IoU threshold) 2: Run the object detector algorithm on all test images. 3: for each class do 4: TP = 0 5: FP = 0 6: Sort the detections using their confidence score in descending order. 7: for each detection do 8: Get all the GT bounding boxes of the corresponding image. 9: Compute the IoU between the detection and all the GT bounding boxes. 10: if the maximum IoU value is 0 then 11: FP = FP+1 .The detection does not overlap with any GT bounding box, it is a FP. 12: else if the GT bounding box with maximum IoU value has the same class id then 13: TP = TP +1 and discard the GT bounding box. 14: else 15: FP = FP + 1. .The class id of the GT is not equal to the detection id, it is a FP. 16: Compute precision and recall using TP and FP values and the total class appearances. 17: Plot a point on the Precision-Recall curve. 18: Class AP = area under the Precision-Recall curve. 19: mAP = average of AP for each class. 20: return mAP To compute the AP, a precision-recall curve is built from the results obtained from the classifier. This curve relates the precision and the recall for each detection of the classifier. A well-known example of AP calculation is a dataset with a target class with 5 appearances, where the detector detects 10 objects. In table 7, we can see the results of ordering the detections by confidence in descending order and applying the algorithm 3. 55
Rank Correct Precision Recall 1 True 1 0.2 2 True 1 0.4 3 False 0.67 0.4 4 False 0.5 0.4 5 False 0.4 0.4 6 True 0.5 0.6 7 True 0.57 0.8 8 False 0.5 0.8 9 False 0.44 0.8 10 True 0.5 1 Table 7: AP calculation example. The detector outputs 10 appearances of the target class, while in the dataset we have only 5 appearances in total. The resulting precision-recall curve can be seen in the figure 33a. A usual approach to facilitate the computations is to smooth the PR curve by assigning to each recall point the higher recall value on the right of the axis. The smoothed PR curve of this example can be seen in the figure 33b. (a) Precision-recall curve (b) Smoothed precision-recall curve. Figure 33: Precision-recall curves. In this thesis, the implementation of the mAP metric is taken from [27], where algorithm 3 is split into two different functions: get_evaluation_bboxes and mean_average_precision. In the function get_evaluation_bboxes, all the target images are passed to the network and their outputs are stored in a convenient format alongside the target tensors for the mAP computation. In the mean_average_precision function, the mAP is computed. To use the mean_average_precision function with the outputs of the networks created in this thesis, a new custom Dataset class and a function called custom_get_evaluation_bboxes have been created. 56
The new custom Dataset is called YOLO_custom_MAP_Dataset, and instead of returning a batch of images and their corresponding target tensors, it returns a batch of images and their corresponding bounding boxes in YOLO format. The custom_get_evaluation_bboxes is very similar to the original one, but it adapts the format of the GT bounding boxes retuned by the YOLO_custom_MAP_Dataset to the format expected by the function mean_average_precision . 3.5.4 Performance metrics Other metrics that must be taken into account when deploying a model, especially on the edge, are the inference speed and the size of the model. •The inference speed is the average time that a model needs to perform inference for a given input. •The size of the model is the size in bytes that the model needs in memory. 57
Appendices A Appendix I: Darknet configuration files A.1 YOLOv3 configuration file 1[net ] 2#Testing 3#batch =1 4#subdivisions=1 5#Training 6batch =64 7subdivisions=16 8width =608 9height =608 10 channels =3 11 momentum =0.9 12 decay =0.0005 13 angle =0 14 saturation = 1.5 15 exposure = 1.5 16 hue =.1 17 18 learning_rate =0.001 19 burn_in =1000 20 max_batches = 500200 21 policy = steps 22 steps=400000,450000 23 scales =.1 ,.1 24 25 [ convolutional ] 26 batch_normalize =1 27 filters =32 28 size =3 29 stride=1 30 pad =1 31 activation = leaky 32 33 #Downsample 34 35 [ convolutional ] 36 batch_normalize =1 37 filters =64 38 size =3 39 stride=2 40 pad =1 41 activation = leaky 42 43 [ convolutional ] 44 batch_normalize =1 45 filters =32 46 size =1 47 stride=1 64
48 pad =1 49 activation = leaky 50 51 [ convolutional ] 52 batch_normalize =1 53 filters =64 54 size =3 55 stride=1 56 pad =1 57 activation = leaky 58 59 [ shortcut ] 60 from =-3 61 activation = linear 62 63 #Downsample 64 65 [ convolutional ] 66 batch_normalize =1 67 filters =128 68 size =3 69 stride=2 70 pad =1 71 activation = leaky 72 73 [ convolutional ] 74 batch_normalize =1 75 filters =64 76 size =1 77 stride=1 78 pad =1 79 activation = leaky 80 81 [ convolutional ] 82 batch_normalize =1 83 filters =128 84 size =3 85 stride=1 86 pad =1 87 activation = leaky 88 89 [ shortcut ] 90 from =-3 91 activation = linear 92 93 [ convolutional ] 94 batch_normalize =1 95 filters =64 96 size =1 97 stride=1 98 pad =1 99 activation = leaky 100 101 [ convolutional ] 102 batch_normalize =1 65
103 filters =128 104 size =3 105 stride=1 106 pad =1 107 activation = leaky 108 109 [ shortcut ] 110 from =-3 111 activation = linear 112 113 #Downsample 114 115 [ convolutional ] 116 batch_normalize =1 117 filters =256 118 size =3 119 stride=2 120 pad =1 121 activation = leaky 122 123 [ convolutional ] 124 batch_normalize =1 125 filters =128 126 size =1 127 stride=1 128 pad =1 129 activation = leaky 130 131 [ convolutional ] 132 batch_normalize =1 133 filters =256 134 size =3 135 stride=1 136 pad =1 137 activation = leaky 138 139 [ shortcut ] 140 from =-3 141 activation = linear 142 143 [ convolutional ] 144 batch_normalize =1 145 filters =128 146 size =1 147 stride=1 148 pad =1 149 activation = leaky 150 151 [ convolutional ] 152 batch_normalize =1 153 filters =256 154 size =3 155 stride=1 156 pad =1 157 activation = leaky 66
158 159 [ shortcut ] 160 from =-3 161 activation = linear 162 163 [ convolutional ] 164 batch_normalize =1 165 filters =128 166 size =1 167 stride=1 168 pad =1 169 activation = leaky 170 171 [ convolutional ] 172 batch_normalize =1 173 filters =256 174 size =3 175 stride=1 176 pad =1 177 activation = leaky 178 179 [ shortcut ] 180 from =-3 181 activation = linear 182 183 [ convolutional ] 184 batch_normalize =1 185 filters =128 186 size =1 187 stride=1 188 pad =1 189 activation = leaky 190 191 [ convolutional ] 192 batch_normalize =1 193 filters =256 194 size =3 195 stride=1 196 pad =1 197 activation = leaky 198 199 [ shortcut ] 200 from =-3 201 activation = linear 202 203 204 [ convolutional ] 205 batch_normalize =1 206 filters =128 207 size =1 208 stride=1 209 pad =1 210 activation = leaky 211 212 [ convolutional ] 67
213 batch_normalize =1 214 filters =256 215 size =3 216 stride=1 217 pad =1 218 activation = leaky 219 220 [ shortcut ] 221 from =-3 222 activation = linear 223 224 [ convolutional ] 225 batch_normalize =1 226 filters =128 227 size =1 228 stride=1 229 pad =1 230 activation = leaky 231 232 [ convolutional ] 233 batch_normalize =1 234 filters =256 235 size =3 236 stride=1 237 pad =1 238 activation = leaky 239 240 [ shortcut ] 241 from =-3 242 activation = linear 243 244 [ convolutional ] 245 batch_normalize =1 246 filters =128 247 size =1 248 stride=1 249 pad =1 250 activation = leaky 251 252 [ convolutional ] 253 batch_normalize =1 254 filters =256 255 size =3 256 stride=1 257 pad =1 258 activation = leaky 259 260 [ shortcut ] 261 from =-3 262 activation = linear 263 264 [ convolutional ] 265 batch_normalize =1 266 filters =128 267 size =1 68
268 stride=1 269 pad =1 270 activation = leaky 271 272 [ convolutional ] 273 batch_normalize =1 274 filters =256 275 size =3 276 stride=1 277 pad =1 278 activation = leaky 279 280 [ shortcut ] 281 from =-3 282 activation = linear 283 284 #Downsample 285 286 [ convolutional ] 287 batch_normalize =1 288 filters =512 289 size =3 290 stride=2 291 pad =1 292 activation = leaky 293 294 [ convolutional ] 295 batch_normalize =1 296 filters =256 297 size =1 298 stride=1 299 pad =1 300 activation = leaky 301 302 [ convolutional ] 303 batch_normalize =1 304 filters =512 305 size =3 306 stride=1 307 pad =1 308 activation = leaky 309 310 [ shortcut ] 311 from =-3 312 activation = linear 313 314 315 [ convolutional ] 316 batch_normalize =1 317 filters =256 318 size =1 319 stride=1 320 pad =1 321 activation = leaky 322 69
323 [ convolutional ] 324 batch_normalize =1 325 filters =512 326 size =3 327 stride=1 328 pad =1 329 activation = leaky 330 331 [ shortcut ] 332 from =-3 333 activation = linear 334 335 336 [ convolutional ] 337 batch_normalize =1 338 filters =256 339 size =1 340 stride=1 341 pad =1 342 activation = leaky 343 344 [ convolutional ] 345 batch_normalize =1 346 filters =512 347 size =3 348 stride=1 349 pad =1 350 activation = leaky 351 352 [ shortcut ] 353 from =-3 354 activation = linear 355 356 357 [ convolutional ] 358 batch_normalize =1 359 filters =256 360 size =1 361 stride=1 362 pad =1 363 activation = leaky 364 365 [ convolutional ] 366 batch_normalize =1 367 filters =512 368 size =3 369 stride=1 370 pad =1 371 activation = leaky 372 373 [ shortcut ] 374 from =-3 375 activation = linear 376 377 [ convolutional ] 70
378 batch_normalize =1 379 filters =256 380 size =1 381 stride=1 382 pad =1 383 activation = leaky 384 385 [ convolutional ] 386 batch_normalize =1 387 filters =512 388 size =3 389 stride=1 390 pad =1 391 activation = leaky 392 393 [ shortcut ] 394 from =-3 395 activation = linear 396 397 398 [ convolutional ] 399 batch_normalize =1 400 filters =256 401 size =1 402 stride=1 403 pad =1 404 activation = leaky 405 406 [ convolutional ] 407 batch_normalize =1 408 filters =512 409 size =3 410 stride=1 411 pad =1 412 activation = leaky 413 414 [ shortcut ] 415 from =-3 416 activation = linear 417 418 419 [ convolutional ] 420 batch_normalize =1 421 filters =256 422 size =1 423 stride=1 424 pad =1 425 activation = leaky 426 427 [ convolutional ] 428 batch_normalize =1 429 filters =512 430 size =3 431 stride=1 432 pad =1 71
433 activation = leaky 434 435 [ shortcut ] 436 from =-3 437 activation = linear 438 439 [ convolutional ] 440 batch_normalize =1 441 filters =256 442 size =1 443 stride=1 444 pad =1 445 activation = leaky 446 447 [ convolutional ] 448 batch_normalize =1 449 filters =512 450 size =3 451 stride=1 452 pad =1 453 activation = leaky 454 455 [ shortcut ] 456 from =-3 457 activation = linear 458 459 #Downsample 460 461 [ convolutional ] 462 batch_normalize =1 463 filters =1024 464 size =3 465 stride=2 466 pad =1 467 activation = leaky 468 469 [ convolutional ] 470 batch_normalize =1 471 filters =512 472 size =1 473 stride=1 474 pad =1 475 activation = leaky 476 477 [ convolutional ] 478 batch_normalize =1 479 filters =1024 480 size =3 481 stride=1 482 pad =1 483 activation = leaky 484 485 [ shortcut ] 486 from =-3 487 activation = linear 72
488 489 [ convolutional ] 490 batch_normalize =1 491 filters =512 492 size =1 493 stride=1 494 pad =1 495 activation = leaky 496 497 [ convolutional ] 498 batch_normalize =1 499 filters =1024 500 size =3 501 stride=1 502 pad =1 503 activation = leaky 504 505 [ shortcut ] 506 from =-3 507 activation = linear 508 509 [ convolutional ] 510 batch_normalize =1 511 filters =512 512 size =1 513 stride=1 514 pad =1 515 activation = leaky 516 517 [ convolutional ] 518 batch_normalize =1 519 filters =1024 520 size =3 521 stride=1 522 pad =1 523 activation = leaky 524 525 [ shortcut ] 526 from =-3 527 activation = linear 528 529 [ convolutional ] 530 batch_normalize =1 531 filters =512 532 size =1 533 stride=1 534 pad =1 535 activation = leaky 536 537 [ convolutional ] 538 batch_normalize =1 539 filters =1024 540 size =3 541 stride=1 542 pad =1 73
54 pad =1 55 activation = leaky 56 57 [maxpool] 58 size =2 59 stride=2 60 61 [ convolutional ] 62 batch_normalize =1 63 filters =128 64 size =3 65 stride=1 66 pad =1 67 activation = leaky 68 69 [maxpool] 70 size =2 71 stride=2 72 73 [ convolutional ] 74 batch_normalize =1 75 filters =256 76 size =3 77 stride=1 78 pad =1 79 activation = leaky 80 81 [maxpool] 82 size =2 83 stride=2 84 85 [ convolutional ] 86 batch_normalize =1 87 filters =512 88 size =3 89 stride=1 90 pad =1 91 activation = leaky 92 93 [maxpool] 94 size =2 95 stride=1 96 97 [ convolutional ] 98 batch_normalize =1 99 filters =1024 100 size =3 101 stride=1 102 pad =1 103 activation = leaky 104 105 ########### 106 107 [ convolutional ] 108 batch_normalize =1 80
109 filters =256 110 size =1 111 stride=1 112 pad =1 113 activation = leaky 114 115 [ convolutional ] 116 batch_normalize =1 117 filters =512 118 size =3 119 stride=1 120 pad =1 121 activation = leaky 122 123 [ convolutional ] 124 size =1 125 stride=1 126 pad =1 127 filters =255 128 activation = linear 129 130 131 132 [yolo] 133 mask = 3,4,5 134 anchors = 10 ,14 , 23 ,27 , 37 ,58 , 81 ,82 , 135 ,169 , 344 ,319 135 classes =80 136 num =6 137 jitter=.3 138 ignore_thresh = .7 139 truth_thresh = 1 140 random=1 141 142 [ route ] 143 layers = -4 144 145 [ convolutional ] 146 batch_normalize =1 147 filters =128 148 size =1 149 stride=1 150 pad =1 151 activation = leaky 152 153 [ upsample ] 154 stride=2 155 156 [ route ] 157 layers = -1, 8 158 159 [ convolutional ] 160 batch_normalize =1 161 filters =256 162 size =3 163 stride=1 81
164 pad =1 165 activation = leaky 166 167 [ convolutional ] 168 size =1 169 stride=1 170 pad =1 171 filters =255 172 activation = linear 173 174 [yolo] 175 mask = 0,1,2 176 anchors = 10 ,14 , 23 ,27 , 37 ,58 , 81 ,82 , 135 ,169 , 344 ,319 177 classes =80 178 num =6 179 jitter=.3 180 ignore_thresh = .7 181 truth_thresh = 1 182 random=1 Listing 5: yolov3-tiny.cfg 82
B Appendix II: Layers implementation B.1 Convolutional layer 1class Convolutional_Layer (nn . Module ): 2 3def __init__ (self , in_channels :int, layer_info : dict): 4super () . __init__ () 5 6if ’batch_normalize ’ in layer_info : 7apply_bn = True 8else: 9apply_bn = False 10 11 if int( layer_info [ ’pad ’]) == 1: 12 pad = int(int( layer_info [’size ’]) /2) 13 else: 14 pad = 0 15 16 self . conv = nn. Conv2d ( in_channels = in_channels , 17 out_channels=int( layer_info [’filters’]) , 18 kernel_size=int( layer_info [’size ’]) , 19 stride=int ( layer_info [’stride’]) , 20 padding =pad , 21 bias= not apply_bn ) 22 23 if apply_bn : 24 self .bn = nn . BatchNorm2d ( num_features = int( layer_info [’ filters’])) 25 else: 26 self.bn = None 27 28 if layer_info [’ activation ’] == " leaky ": 29 self . activation = nn . LeakyReLU (0.1) 30 else: 31 self . activation = None 32 33 def forward ( self ,x): 34 35 if self .bn is not None and self . activation is not None: 36 return self . activation ( self .bn( self .conv (x))) 37 elif self.bn is None and self . activation is not None: 38 return self . activation ( self .conv (x)) 39 else: 40 return self.conv(x) Listing 6: Convolutional function 83
B.2 Maxpool layer 1class Maxpool_Layer (nn . Module ): 2 3def __init__ (self , layer_info : dict): 4super () . __init__ () 5 6self . kernel_size = int ( layer_info [’size ’]) 7self . stride = int (int( layer_info [ ’stride’])) 8# There is an incompatibility issue due to a tiny difference in padding algorithms between darknet and tensorflow / pytorch . 9# In particular , layer 11 in yolov3 -tiny is a 2x2 maxpool with stride =1 , operating on 13 x13x512 input . In darknet this produces 10 # an output that is 13 x13x512 . However in tensorflow / pytorch , it produces 12 x12x512 ( default padding in TF is valid). 11 # Consecutive layers mismatch and concat complains . To solve this we must add padding in this case . 12 if self . kernel_size == 2 and self . stride == 1: 13 self. padding = nn. ZeroPad2d ((0 , 1, 0, 1)) 14 15 self . maxpool = nn. MaxPool2d ( kernel_size = self .kernel_size , 16 stride=self.stride) 17 18 19 def forward ( self ,x): 20 21 if self . kernel_size == 2 and self . stride == 1: 22 x = self . padding (x) 23 24 return self . maxpool (x) Listing 7: Maxpool layer 84
B.3 Upsample layer 1class Upsample_Layer ( nn. Module ): 2 3def __init__ (self , layer_info : dict): 4super () . __init__ () 5 6# To upscale your image by factor x you normaly use a stride of x. Check this . 7self . upsample = nn. Upsample ( scale_factor = int ( layer_info [" stride"]) , mode = " bilinear ") 8 9def forward ( self , x): 10 11 return self . upsample (x) Listing 8: Upsample layer 85
B.4 Route layer 1class Route_Layer (nn . Module ): 2 3def __init__ (self , layer_info : dict): 4super () . __init__ () 5 6self . layers = [ int(x) for xin layer_info ["layers"]. split (",")] 7# As the list position starts at 0 we need to change positive values: 8# as the output of the first layer would be output [0] 9self . layers = [x -1 if x >0 else xfor xin self . layers ] 10 11 if len( self . layers ) not in [1 ,2]: 12 raise Exception (" Route_Layer : Layers value len not valid ") 13 14 def forward ( self , outputs ): 15 16 if len( self . layers ) == 1: 17 x = outputs [ self . layers [0]] 18 19 else:# len ( self . layers ) == 2 20 map1 = outputs [ self . layers [0]] 21 map2 = outputs [ self . layers [1]] 22 23 x = torch . cat (( map1 , map2 ) , 1) 24 25 return x Listing 9: Route layer 86
B.5 Shortcut layer 1class Shortcut_Layer ( nn. Module ): 2 3def __init__ (self , layer_info : dict): 4super () . __init__ () 5 6self . _from = int( layer_info [" from "]) 7if self . _from >= 0: 8raise Exception (" Shortcut_Layer : from value not valid ") 9# An activation fied is also present but is always linear 10 11 def forward ( self , outputs ): 12 13 # We have just to add the previous output with the shortcut 14 x = outputs [ self . _from ]+ outputs [ -1] 15 16 return x Listing 10: Shortcut layer 87
B.6 YOLO layer 1class YOLO_Layer ( nn. Module ): 2 3def __init__ (self , layer_info : dict, num_classes:int): 4super () . __init__ () 5self . num_classes = num_classes 6 7def forward ( self ,x): 8return x. reshape (x. shape [0] , 3, self . num_classes + 5, x. shape [2] , x. shape [3]) . permute (0 , 1, 3, 4, 2) Listing 11: YOLO layer 88
C Appendix III: Data augmentation pipelines C.1 Train data augmentation pipeline 1 2train_transforms = A. Compose ( 3[ 4A. LongestMaxSize ( max_size = int( IMAGE_SIZE * scale )) , 5A.PadIfNeeded( 6min_height =int( IMAGE_SIZE * scale ), 7min_width = int( IMAGE_SIZE * scale ) , 8border_mode = cv2. BORDER_CONSTANT , 9), 10 A. RandomCrop ( width = IMAGE_SIZE , height = IMAGE_SIZE ) , 11 A. ColorJitter ( brightness =0.6 , contrast =0.6 , saturation =0.6 , hue =0.6 , p =0.4) , 12 A. OneOf ( 13 [ 14 A.ShiftScaleRotate( 15 rotate_limit =20 , p =0.5 , border_mode = cv2 . BORDER_CONSTANT 16 ), 17 A. IAAAffine ( shear =15 , p =0.5 , mode =" constant "), 18 ], 19 p=0.5 , 20 ), 21 A.HorizontalFlip(p=0.5), 22 A. Blur (p =0.1) , 23 A. CLAHE (p =0.1) , 24 A. Posterize (p =0.1) , 25 A. ToGray (p =0.1) , 26 A.ChannelShuffle(p=0.05), 27 A. Normalize ( mean =[0 , 0, 0] , std =[1 , 1, 1], max_pixel_value =255 ,) , 28 ToTensorV2 () , 29 ], 30 bbox_params =A. BboxParams ( format=" yolo ", min_visibility=0.4, label_fields=[],), 31 ) Listing 12: Data augmentation pipeline used in training. 89