scieee AI-readable full text Open interactive document viewer

Unsupervised domain adaptation for bladder segmentation by U-nets in Cone Beam CT.

Cuxart García, Anna

Abstract

Introduction: Radiotherapy is a medical treatment used to control or kill cancerous cells in cancer patients. At the beginning of the treatment planning process, the patient takes a CT scan in order to plan his radiation dose and, sometimes, he can take a Cone Beam CT scan some days after, right before receiving his radiation, to adjust his couch position for the delivery. The main difference between CT and CBCT scans is that the first one has a higher quality and contrast, and the second one is taken directly at the isocenter. As the treatment planning takes several days, when the patient receives his radiation his organs might not be in the same position as they were at the beginning, so the healthy tissues around the tumor area can receive more radiation than what it was planned and get damaged. Our aim is to implement an automatic segmentation of the bladder in CBCT 3D images using deep learning, in order to get a clearer idea of the position of those organs. Materials and Methods: In order to implement the segmentation we performed unsu- pervised domain adaptation between CT (the source) and CBCT 3D images (the target), as we didn?t have a large labeled CBCT dataset but we did for CT, as their segmenta- tion is already part of the treatment planning process. We have used a subset dataset of 120 patients: 60 CT and 60 CBCT of the male pelvic region. We have implemented a deep learning network using Unet as a segmenter and a regular CNN as a domain discriminator (an adversarial network), which also includes a gradient reversal layer. We have used the Dice score coefficient and the Hausdorff distance in order to evaluate the performance of our network and compare it with some previous works developed in the same field. Results: We have performed three main experiments, for which we have obtained the following DSC and Hausdorff distance (in voxels): (i) lower boundary: 0.383±0.260 and 36.47, (ii) upper boundary: 0.717±0.177 and 27.44, (iii) unsupervised domain adapta- tion: 0.623±0.149 and 39.18. With this implementation we have closed the gap between training the network only on CT and only on CBCT by a 72%. Conclusions: Cone Beam CT image segmentation using unsupervised domain adap- tation proves to be an improving methodology in radiotherapy and presents different applications in other fields.

Full text

École polytechnique de Louvain ! ! Unsupervised domain adaptation for bladder segmentation by U-net in Cone Beam CT Author : Anna CUXART Supervisor : Benoît MACQ Readers : Benoît MACQ, Eliott BRION, John LEE Academic year 2019–2020 Master in Advanced Telecommunications Technologies (UPC) (EPL) Polytechnic University of Catalonia (UPC) POLYTECHNIC UNIVERSITY OF CATALONIA (UPC) UNIVERSIT´ E CATHOLIQUE DE LOUVAIN (UCL) Abstract by Anna Cuxart Garcia Introduction Radiotherapy is a medical treatment used to control or kill cancerous cells in cancer patients. At the beginning of the treatment planning process, the patient takes a CT scan in order to plan his radiation dose and, sometimes, he can take a Cone Beam CT scan some days after, right before receiving his radiation, to adjust his couch position for the delivery. The main difference between CT and CBCT scans is that the first one has a higher quality and contrast, and the second one is taken directly at the isocenter. As the treatment planning takes several days, when the patient receives his radiation his organs might not be in the same position as they were at the beginning, so the healthy tissues around the tumor area can receive more radiation than what it was planned and get damaged. Our aim is to implement an automatic segmentation of the bladder in CBCT 3D images using deep learning, in order to get a clearer idea of the position of those organs. Materials and Methods In order to implement the segmentation we performed unsupervised domain adaptation between CT (the source) and CBCT 3D images (the target), as we didn’t have a large labeled CBCT dataset but we did for CT, as their segmentation is already part of the treatment planning process. We have used a subset dataset of 120 patients: 60 CT and 60 CBCT of the male pelvic region. We have implemented a deep learning network using Unet as a segmenter and a regular CNN as a domain discriminator (an adversarial network), which also includes a gradient reversal layer. We have used the Dice score coefficient and the Hausdorff distance in order to evaluate the performance of our network and compare it with some previous works developed in the same field. Results We have performed three main experiments, for which we have obtained the following DSC and Hausdorff distance (in voxels): (i) lower boundary: 0.383±0.260 and 36.47, (ii) upper boundary: 0.717±0.177 and 27.44, (iii) unsupervised domain adaptation: 0.623±0.149 and 39.18. With this implementation we have closed the gap between training the network only on CT and only on CBCT by a 72%. Conclusions Cone Beam CT image segmentation using unsupervised domain adaptation proves to be an improving methodology in radiotherapy and presents different applications in other fields. Acknowledgements First of all, I would like to thank my project supervisor Benoˆıt Macq from the ICTEAM Research Group at Universit´e Catholique de Louvain for his trust and help, and the opportunity to work with him. Special thanks to his PhD Eliott Brion, for guiding me through all the project, providing all the information needed and always being accessible for help. Thanks also to Jean L´eger for helping me with their 2D neural network model, and to S´ebastien Lugan for his dedication to solve the technical issues in order to properly work in remote. I want to thank CHU-UCL-Namur (Dr. Vincent Remouchamps) and CHU-Charleroi (Dr. Nicolas Meert) for providing the data for this project, as well as Sara Teruel Rivas for her technical support in the data acquisition and annotation and to Dr. Genevieve Van Ooteghem for the verification of the contours. I would like to thank also Prof. Ferran Marqu´es from UPC for highly recommending me to come at UCL to develop my thesis with ICTEAM. Thanks also to Prof. Enric Monte for agreeing to be my supervisor and tutor at UPC. Lastly, thanks to my family and friends for encouraging me to stick to my goals and never give up, despite unforeseen situations. Contents Acknowledgements Abbreviations 1 Introduction 1 1.1 Radiotherapy .................................. 1 1.1.1 Treatment planning .......................... 2 1.2 State of the Art and previous work ...................... 5 1.3 Introduction to machine learning ....................... 7 1.4 Unsupervised Domain Adaptation ...................... 9 1.5 Objective .................................... 10 1.6 Contribution of this master thesis ....................... 11 1.7 Structure of this master thesis ......................... 12 1.8 Summary of Chapter 1 ............................. 12 2 Background 14 2.1 Machine learning ................................ 14 2.1.1 Neural network hyper-parameters ................... 15 2.1.2 Convolutional Neural Networks .................... 17 2.2 Previous research ................................ 18 2.2.1 Unsupervised domain adaptation with adversarial networks . . . . 18 2.2.2 Unsupervised domain adaptation with backpropagation ...... 19 3 Materials and methods 23 3.1 Dataset ..................................... 23 3.2 Environment .................................. 26 3.3 Hyperparameters ................................ 26 3.4 Our architecture ................................ 29 3.4.1 Unet ................................... 30 3.5 Metrics ..................................... 30 4 Experiments & Results 33 4.1 Lower Bound - train on CT .......................... 34 4.2 Upper Bound - train on CBCT ........................ 35 4.3 Main experiment - Unsupervised domain adaptation ............ 36 4.3.1 Losses .................................. 37 4.3.2 Lambda ................................. 38 Contents 4.3.3 Discriminator Accuracy ........................ 38 4.3.4 Source DSC - CT ............................ 39 4.3.5 Target DSC - CBCT .......................... 39 4.4 Unsupervised domain adaptation with lambda = 0 ............. 40 4.5 Results examples ................................ 41 5 Discussion 44 6 Conclusions 49 A Convolutional Neural Networks 52 B User Guide 57 Bibliography 58 Abbreviations UDA Unsupervised Domain Adaptation ANN Artificial Neural Network AI Artificial Intelligence CNN Convolutional Neural Network NN Neural Network CT Computed Tomography CBCT Cone Beam Computed Tomography TP True Positive FP False Positive FN False Negative TN True Negative SGD Stochastic Gradient Descent GRL Gradient Reversal Layer GT Ground Truth DSC Dice Score Coefficient SMBD Symmetric Mean Boundary Distance Chapter 1 Introduction 1.1 Radiotherapy Radiation Therapy, also called Radiotherapy, is a medical treatment used to control or kill cancerous cells. It uses high doses of intense radiation in order to kill cancer cells and reduce the patient’s tumor. The main goal of Radiotherapy is to damage the DNA of the cancer cells, in order to slow their growth and avoid further dividing. The given radiation reacts with the water of the cells, damaging their genetic material that controls cell growth. Mostly, all regular cells can recover their functional behaviour after the radiation treatment, but that’s not the case for cancer cells. Unfortunately, it is quite a slow process, as it takes several days and weeks of constant treatment before the DNA is damaged enough for the cancer cell to die, and the healthy cells need to recover from the radiation between the sessions. There are two kinds of radiation therapy, and they can be given to the patient depending on different aspects: the specific kind of cancer, where is it placed, how deep should the radiation go into the body and the health of the patient amongst others. There are the external beam and internal beam, and their main difference is the source of the radiation given to the patient. The first one uses an external machine outside the body, while the second one uses an implant that is placed inside the patient’s body trough surgery to provide a more focused and controlled radiation. 1 Introduction 2 External beam is the most common one, and it can be used to treat different kinds of cancers, as well as to relieve pain that this illness can cause when it spreads to other parts of the body [1] [2][3][4]. 1.1.1 Treatment planning As there are different kinds of radiation and different methodologies, an important part of radiotherapy is the treatment planning. As a previous step of the radiation delivery, a group of medical experts (the radiation oncologist and the radiation therapists) analyze the specific case of the target patient and plan his whole treatment: where is his affected area placed in his body, which kind of radiation should he receive, the daily radiation dose, in which position should he receive it, etc. Figure 1.1: In radiotherapy, during the treatment planning, the patient takes first a CT scan in order to plan his radiation dose. On the last step, right before he receives his radiation, he can take a second scan, the Cone Beam CT, in order to adjust his position in the machine for the radiation delivery. - Source: Xavier Geets, Assessment of the quality of a treatment plan (unpublished). In the case of external radiation, the physicist and the physician will firstly execute a simulation method in order to examine where and how the patient’s treatment should be applied. In the very first step of the simulation the patient will lay down, as still as possible, in a table of examination. In order to know where the radiation should be applied, the therapists scan the patient using a special X-Ray machine, performing the Computed Tomography (CT) scan (Planning CT-scan in 1.1). Introduction 3 The CT scan uses a combination of different computer-processed X-Ray images taken from different angles of the patient. The x-rays utilize radiation from a radioactive contrast injected into the body of the patient to create cross-sectional (tomographic) images. In the particular case of the CT, the emitter of x-rays rotates around the immobile patient and the detector, placed in diametrically opposite side, picks up the image of a body section (beam and detector move synchronously). This simulation and scanning process may take several hours, as the medical experts want to assure that they get a clear image of the affected area, so they can plan how to provide the radiation while damaging the least amount of healthy tissue. Figure 1.2: Slice of a 3D CT scan image of the male pelvic area (left) with his bladder mask segmented by hand (right). Pixel spacing (mm) = [1.2, 1.2, 1.5]. Once the medical experts have all the CT scan images from the patient, they will carefully delineate manually the affected regions by the cancer, contouring precisely the organs and the tumors (Delineation in 1.1). They will use them to plan the radiation dose: adjusting where should it be projected and choosing the adequate intensity, so the main tumor can receive the main radiation but the healthy tissues around are minimum affected (Dose planning in 1.1). Formerly, they will check the performance of the planning by doing a Q&A test using the radiation machine on a bag of water, simulating the human body (QA in 1.1). Thereupon, some days after, the patient will be ready to start his first radiation session. Unfortunately, the treatment can not be given with the same machine that did the Introduction 10 the domain. Its goal is to train a learning function fa(∆) that can perform successfully on both domains. Unsupervised domain adaptation (UDA) [8] is a subclass of DA that assumes the availability of a labeled source domain database S = {Xs, Ys}and an unlabeled target domain database T = (Xt). In this environment, the main goals are first learn a representation ha(x) that maps Xs and Xt to a feature space invariant to the differences between the two domains; and secondly, learn a function fah(x) such that fa(x) = fah(ha(x)) approximates f0 s(∆) and it is closer to f0 T(∆) than any function fs(∆) learned only using the source data (Xs, Ys). In this project, we are going to work with a source domain database composed by labeled CT images and a target domain database composed by unlabeled CBCT images. 1.5 Objective Our main goal of this project is to accomplish the automatic segmentation of CBCT radiotherapy images using deep learning. Why? Because they are used in some processes to treat cancer, in order to analyze where the radiation has to be applied. And we think that, by segmenting the images, the treatment applied to the patient could be executed more precisely, so the healthy tissues and organs around the tumor area would be less affected. How are we going to do it? We are going to build a system based on neural networks architecture, and we will feed it with different examples of already segmented CBCT images, providing him the original image and the mask or segmentation. From this data, the algorithm will learn the different features of the images and their regions, so afterwards he will be able to make decisions by its own and predict or propose segmentations for new images. The system will autonomously learn how to do the segmentations, specifically bladder segmentations. Introduction 11 The main constraint is that CBCT images have a low contrast and quality, therefore they are not easy to segment. And the second constraint is that we do not have a dataset of already segmented CBCT images to train a deep learning algorithm, and in order to obtain it we would steal a lot of time of doctors. What we sure have is a dataset of segmented CT images from previous treatments, which are similar to our target but from another domain. Hence, we propose to implement the automatic CBCT image segmentation NN by using unsupervised domain adaptation. UDA is a deep learning technique that will allow us to work with a CT dataset, with already segmented examples, and a CBCT dataset, without segmented examples in order to find a joint representation of the two datasets that doesn’t depend on the domain. This way, we will be able to train our neural network with both datasets, make it learn the features of both images that are domain independent and use them to teach the network to segment the CBCT images based on the already segmented CT examples. In the following sections, we will proceed to explain in more detail what is exactly machine learning and an accurate definition of unsupervised domain adaptation. 1.6 Contribution of this master thesis Our main proposal is to implement unsupervised domain adaptation for cone beam CT images (CBCT) using deep learning and, in comparison with the previous works, the main novelty that this master thesis presents is that it combines the following three features at the same time: (i) we are working with a 3D dataset of CBCT images, (ii) we are implementing an architecture that involves adversarial networks with the gradient reversal layer and (iii) we are using Unet as the segmenter of our network. (i) We are working with a 3D dataset Compared to the previous works such as [6] and [8], in this project we are using a 3D dataset of CT and cone beam CT images. With this, we aim to work closer to the radiotherapy treatment cases described in this chapter, shortening the gap between the experimental work and the clinic applications. Introduction 12 (ii) We are implementing an architecture that involves adversarial networks with the gradient reversal layer As described in [8], we are going to perform unsupervised domain adaptation by implementing an adversarial neural network. But, in order to implement the backpropagation, we are going to join a gradient reversal layer to it, as defined in [7]. (iii) We are using Unet as the segmenter of our network In contradistinction to [8], for this project we are going to use the CNN Unet as the segmenter network of our architecture. Unlike the completely fast forward network that they present, which only plays with downsampling pass, we are going to work with this much more complex network that includes skipping connections and relies on data augmentation. Moreover, while our use case is in radiotherapy, our work proposal could be used for any application in image segmentation beyond medical environments, such as face recognition, self-driving vehicles and satellite imagery amongst many others. 1.7 Structure of this master thesis The organization of this master thesis is the following: first, in Chapter 2, we present some background information, by defining what is machine learning and presenting the previous works in which this project is based. In Chapter 3, we introduce the dataset that we have used and we explain in more details the methods that we have employed for this work, such as the environment, the hyperparameters of our network, the architecture that we have designed and the metrics that we have used to evaluate our results. In Chapter 4, we define the experiments that we have performed and share the obtained results, which we discuss in Chapter 5. And finally, in Chapter 6, we terminate with our main conclusions and possible developments for a future work. 1.8 Summary of Chapter 1 Radiation therapy is a medical treatment used in cancer patients that gives a high dose of intense radiation to kill or control the cancer cells and reduce the patient’s tumor. Sometimes, during the treatment process, two different kinds of scans are taken, first the computed tomography (CT) and then the cone beam computed tomography (CBCT Introduction 13 (section 1.1.1)). As they are taken in different days, the organs of the patient around the tumor area can be displaced, and therefore receive a higher amount of radiation than what it was firstly planned. Previous works (section 2.2) show that automatic segmentation of CBCT images can improve this situation, by curately showing the new position of the organs. Therefore, in this project we aim to implement CBCT image segmentation of the bladder by using unsupervised domain adaptation with adversarial neural networks and Unet. Chapter 2 Background In this section we present some detailed definitions of what is machine learning and how it works, followed by an accurate description of the previous works that we have been following. 2.1 Machine learning Our segmentation methodology will be based on convolution neural networks (CNN), as in the previous projects and in line with the last trends in the scientific community [8]. In this section, we are going to show concepts related to Neural Networks and some new techniques that could aid in the improvement of results. The algorithmic structures of artificial neural networks (ANN) allow models that are composed of multiple layers of processing to learn data representations with multiple levels of abstraction. This set of layers are composed by neurons and perform a series of linear and non-linear transformations to the input data in order to generate an output close to the expected (the label). The main scenario, supervised learning, consists in obtaining the parameters of these transformations: the wi(weights) and the b(biases); with the aim that the output of the neuron (y) differs as little as possible to the expected output. For each artificial neuron we have ”m” input signals (x), as usually the x0input is assigned the value +1. 14 Background 15 Neural networks use the error, or loss, of their outputs in a process called Backpropagation, whereby these weights and biases are updated. y= m X i=0 Wixi+b(2.1) A group of neurons form a layer, and multiple layers form a network, as for example the multilayer perceptron 2.1. This basic NN is composed by a minimum of three layers: one Input Layer (the first, where we will feed our data), one Output Layer (the last) and at least one Hidden layer. All the neurons of each layer are connected with a certain weight to every node in the following layer, and the amount of Hidden layers will depend on the complexity of the task that we want it to learn, more difficulty more layers. Figure 2.1: A Neural Network is composed by, at least, three layers: the input, the output and the hidden layer between them. The multilayer perceptron in the smallest version with three layers [15]. 2.1.1 Neural network hyper-parameters Neural networks have a lot of parameters that need to be adjusted in order to obtain better results. In the following sections we will cover the most important ones. 2.1.1.1 Epochs The number of epochs [16] means the number of times that we will pass all the training data through the neural network during the training process. The number of epochs is Background 16 usually increased until the accuracy of the validation data starts to decrease, even when the accuracy of the training data continues to increase (as this is when we could detect a potential overfitting). 2.1.1.2 Loss One of the most important hyper-parameters is the loss function. We will use it to estimate the error by comparing and measuring the performance of our prediction with the expected result. The main goal is to keep is as small as possible, as we don’t want divergences between the estimated and the expected values. Therefore, during the process of training the model, the weights of the interconnections of the neurons will be adjusted for a higher accuracy, by reducing the loss towards its minimum. The choice of the best function of loss resides in understanding what type of error is or is not acceptable for the problem in question. 2.1.1.3 Activation Function Each neuron of the neural network has an activation function that defines the output of the neuron. It is used to introduce non-linearity in the modeling capabilities of the network. We can find several options for activation functions, but the most common are: Linear,Sigmoid,Tanh,ReLU and Softmax [15]. The choice of a specific activation function depends on several factors: (a) the kind of data that are being processed, (b) the kind of prediction that we seek to obtain, (c) which part of the network we are in, and (d) in which kind of layer we are going to apply it. 2.1.1.4 Optimizer During the training process of the neural network, the main goal is to automatically adjust the network parameters (weights and biases) in order to minimize the loss function. We want to optimize the network to obtain the best results. We can use different optimizers, and the most common ones are the following: Stochastic Gradient Descent,RMSprop,Adagrad,Adadelta,Adam,Adamax and Nadam [15]. Background 17 2.1.1.5 Learning Rate In order to adjust and update the parameters of the network (weights and biases), the learning rate controls how much these components change per epoch, being αin the gradient descent equation [17]: Wij ←Wij −αdError dWij (2.2) The learning rate decay (α) is used to decrease the learning rate as epochs go by to allow the learning to advance faster at the beginning than towards the end (”fine-tuning”). As progress is made, smaller and smaller adjustments are made to facilitate the convergence of the training process to the minimum of the loss function. 2.1.2 Convolutional Neural Networks As their name self describes, artificial neural networks are an artificial simulation of the animal’s biological neural network, with the main goal of being trained, learned and be capable to make decisions by mimicking the human brain distribution. We can program different computers as if they were a network of millions of brain cells, the so called artificial neurons, and train them to learn how to make automatic decisions [18] [19]. There are many different classes or types of ANN, depending on the task we want the computers to learn, but the most common ones are the following: •MLP: Multilayer perceptron are the vanilla neural networks, composed by at least 3 layers of artificial neurons (input layer, minimum 1 hidden layer, and one output layer). Their main characteristic is that each neuron of the network is fully connected to all the neurons of the previous and following layers. •RNN: Recurrent neural networks are basically designed to work with sequence prediction problems. Their main characteristic is that they have a recurrent connection on their hidden states, the artificial neurons of the hidden layers, in order to ensure that the sequential information is captured in the input data. •CNN: Convolutional Neural Networks. Background 18 Convolutional Neuronal Networks [20] [21] appeared at the late 90s, and their name comes from their use of Convolutional Layers A.1, which look for spatial relationships inside the input data, so they can learn to recognize patterns across the space. Some of their applications can be system recommendation or natural language processing, but nowadays they are mostly used for image and video classification and recognition, and they have had a huge impact in the area of computer vision. The main characteristic of CNNs is that they encode certain properties of the input data into the architecture, so they can learn to recognize specific components of an image, from lines and curves to more complex figures like faces, objects or animals. In this project, we will use CNNs to extract features from both CT and CBCT images in order to automatically segment the affected areas and organs of the patients. However this kind of features are quite complex, requiring that we first extract the simpler ones such as lines and curves in the first layers, and as we get to the deeper ones we will get to more complex features like circles and textures. as well as the architecture we have designed for this project and the datasets that we have used. 2.2 Previous research In the previous section we have seen more detailed definitions of machine learning and how neural networks work. Formerly, we are going to present some previous projects that have been done in the same research area. As mentioned in 1.2, this project is based in three previous works that have been developed in this field. Two of them [8] and [7] propose to implement an unsupervised domain adaptation with two interesting features: 2.2.1 Unsupervised domain adaptation with adversarial networks The first paper [8] describes an architecture to implement unsupervised domain adaptation to segment brain lesion images by using adversarial networks. Background 19 The authors propose to use two different networks to achieve the UDA segmentation: the segmenter and the domain discriminator (the adversarial network). The first one is trained using only the source dataset, and the second one is trained with both source and target datasets, in order to learn those features that are invariant to their domains. With them, after the training, the segmenter is finally able to segment the target images without any example of their original segmentations. 2.2.2 Unsupervised domain adaptation with backpropagation The second paper [7] describes an architecture composed by three connected networks: the feature extractor, the label predictor and the domain classifier. And their interesting extra feature is that between the domain discriminator and the feature extractor they place a gradient reversal layer, which allows the network to be trained using backpropagation. Lets dig a little bit deeper into their proposal to understand how it works and how we adapted it to our project. In the following figure we can have a closer look at the architecture of the model: Figure 2.2: The unsupervised domain adaptation by backpropagation in [7] presents the following neural network architecture: a feed-forward network composed by a feature extractor and a label predictor, which allows the segmentation task, and a domain classifier that allows the unsupervised domain adaptation task. This last network is connected to the feature extractor by a gradient reversal layer, in order to compute the backpropagation by multiplying the gradient by a negative constant. This bans the domain classifier to distinguish the domain of the input features, finding therefore the domain-invariant features. Materials and methods 26 3.2 Environment For this project we have used Ubuntu 1and Anaconda 2as the virtual environment manager, and Jupyter notebook 3as a development tool, and the whole project was written in Python 4. We have trained and tested our CNNs with nVidia cuDNN 5acceleration on a computer equipped with a GPU nVidia GeForce GTX 1080 Ti 611Gb, and our training time is 11 minutes. Regarding other packages that we used, TensorFlow 7and Keras 8are the most important ones due to their powerful handling of deep learning. You can find our whole code containing the development of our 4 experiments in the following repository on Github: https://github.com/annacuxart/unsupervised_domain_adaptation_unet_CBCT_segmentation 3.3 Hyperparameters Each neural network has its own characteristic elements, and their initialization, optimization and tuning are key to make it optimal and powerful. In the previous sections we have been talking about the Parameters of the network, the coefficients that the network optimizes using different algorithms and strategies, for example the weights Wand the biases b. These parameters are initialised by us, but during the whole training process they are chosen and updated by the algorithm itself, in order to reduce the loss and improve the network. Another characteristical element of a neural network is its Hyperparameters 2.1.1, and the main difference with the other ones is that we set them at the beginning of the training process and they rest untouched, they are not optimized by the network. In order to set them, we require a minimum knowledge of deep learning, as they determine the structure of our network and how is it going to be trained. 1https://ubuntu.com/ 2https://www.anaconda.com 3http://jupyter.org/ 4https://www.python.org 5https://developer.nvidia.com/cudnn 6https://www.nvidia.com/en-sg/geforce/products/10series/geforce-gtx-1080-ti/ 7https://www.tensorflow.org/ 8https://keras.io/ Materials and methods 27 For this project, we have selected the following Hyperparameters, following the implementations of [6]: •Learning rate: at first we setted the same learning rate for both networks, but after several training trials we realized that the segmenter and the domain discriminator had different learning speeds, so we finally opted for splitting them as it follows. –Segmenter learning rate: µS(p) = µ0 (1 + αp)β(3.1) where p= (current step) / nsteps,µ0= 1e-2, α= 10, β= 1 and ηsteps = 1e4. –Domain discriminator learning rate: µD(p) = ψ×µS(p) (3.2) where ψ= 1e-2. •Total Loss: as descrived in 2.14, we will compute our total loss by adding both segmenter’s and domain discriminator losses, taking into account that the second one includes the gradient descend layer 2.12 and 2.13. ˜ E(θs, θd) = X i=1..N,di=0 Ls(Gs(x;θs)) + X i=1..N Ld(Gd(Rλ(x;θd))) (3.3) In this equation, Lsand Ldstands for the segmenter and discriminator’s losses respectively. They are both computed with the cross entropy of the predictions on a training batch, the first one using a sigmoid as an activation function, and the second one using a softmax. The subindex istands for all the samples in the batch, and di= 0 means that the segmenter’s loss is only computed using the samples from the target domain, contrary to the discriminato’s loss, which uses data from both domains. x corresponds to the input data, which is first mapped by the segmenter Gswith its vector of parameters θs. At the same time, the input data goes trough the gradient reversal layer Rλand enters the domain discriminator Gd, with its parameters θd. Finally, both losses are computed and added up to obtain the total loss of the network. Materials and methods 28 •Batch Size: as we are working with 3D images, which translates to a lot of parameters, we decided to feed them into the network by small batches at a time. So we finally chose a batch size = 4. •Number of Feature Maps: feature maps get their name by their function of mapping of where a certain kind of feature is found in the image, a high activation means a certain feature was found. In our case, we have chosen 12. •Domain discriminator architecture: another hyperparameter that we can adjust is the architecture of the domain discriminator network. In our case, following the examples of the original papers [8], we decided to implement a Convolutional Neural Network, described in more detail in 3.3. •Lambda: as described in [8], this hyperparameter adjusts the strength with which the segmenter is adapting its features in order to counter the domain discriminator, and it also controls the strength with which the discriminator is adapting his features. For our project we have chosen to use the same lambda schedule one as the original Kamnitsas paper proposed. λ(p) =                        0if p < p1 λ=λmax ×pcurr −p1 p2−p1 if p1< p < p2 λmax if p2< p (3.4) Where p1= 0.2, p2= 0.7 and λmax = 3. Materials and methods 29 3.4 Our architecture In order to implement the unsupervised domain adaptation to segment CBCT images, we have taken the architecture of [6], which was based on 2.2.1 and 2.2.2, and we have adapted it to work with 3D images instead of 2D. We have implemented two connected Neural Networks: the segmenter and the domain discriminator. But we have done a modification that has not been seen yet. In our case, we have used the CNN Unet [22] as the model for the segmenter network, and a regular fast-forward CNN for the domain discriminator. Figure 3.3: Our architecture is composed by: Unet structure as a segmenter to predict the labels, and a regular CNN as a domain discriminator (the adversarial network) in order to implement the unsupervised domain adaptation, finding those features that are invariant to the domain. The domain discriminator is connected with a gradient reversal layer to the 6th level of the segmenter, in order to compute the UDA with backpropagation, as discussed in 2.2. As we can see in 3.3, we have build a combination of a modified Unet (see more details in 3.4.1) and a CNN. Our Unet has a depth of 4 layers instead of 5 in order to adjust it to the dimensions of our input data, as we will explain in the following section. The second network, the domain discriminator, is implemented with two convolutional layers with 3x3x3 filters and a ReLU, a MaxPooling layer of 2x2x2, another two convolutional layers and two final fully connected layers (for more details of each layer see Appendix A). In the original paper of [8] they proposed different structures, by connecting the domain discriminator at different levels of the expanding part of Unet. In our case we decided Materials and methods 30 to connect it at the 6th level as a start, but for future steps it would be interesting to compare it with the results obtained at different levels, even a combination of them. And, as descrived in 3.3, we will compute the total loss of our network as the loss of the segmenter plus the loss of the domain discriminator using backpropagation. Let’s now take a look at the definition of the original Unet itself. 3.4.1 Unet Unet [22] is a known CNN that relies on data augmentation A.4.1 to take more advantage of the available labeled data samples. This network is composed by convolutional layers with 16 3x3 filters with stride 2, non padded and activated by ReLU. After each convolution layer there is a 2x2 max-pooling layer that reduce the amount of features. Its name appeals to its ”U” shape, due to its contracting path, that captures the context in the data, and its expanding path, that allows precise localization, and it is commonly used for biomedical image segmentation, such as [23] and [24]. 3.5 Metrics Finally, in order to evaluate our results and compare them with previous related works, we are going to use the following metrics: dice score coefficient and Hausdorff distance. The first one is overlap based and the second one is distance based, so they provide complementary information about the quality of the segmentation results: •Dice Score Coefficient: When working with neural networks for image segmentation, the DICE score (also known as the ”similarity coefficient”) [25] is a commonly used metric. It quantifies how closely the network’s outputs match the training dataset comprising hand-annotated ground truth segmentation. In our case, it will let us compare the automatic image segmentation that the NN estimates with the real segmentation, the one done by hand by the Doctors. It will compute the size of the overlap of the two segmentations divided by the total size of the two affected areas (the ones that we are trying to segment). Therefore, Materials and methods 31 it will provide us the information of the networks accuracy, so we can know how is it working and we can make the adequate arrangements. DICEScore =2∗TP 2∗TP +FP +FN (3.5) TP stands for the number of True Positive pixels, FP is the number of False Positive pixels and FN is the number of False Negative pixels, as illustrated in 3.4. Figure 3.4: The classification chart [26] is commonly used in medicine, in order to perform binary classification tests. It shows the original positive and negative samples (relevant elements and right side respectively) and the prediction that it’s been made (selected elements). With this, we can know which samples have been correctly predicted as positive (true positive), which ones have been wrongly predicted as negative (false positive), which ones have been wrongly predicted as negative (false negative) and which ones have been correctly predicted as negative (true negative). •Hausdorff distance: Besides the Dice Score Coefficient metric, in order to evaluate our results we also computed the Hausdorff Distance [27] between the original segmentations and our predicted ones. The Hausdorff Distance is a mathematical way to compute, using the Euclidean distance, the maximum distance of a set to the nearest point in another set, as it follows: h(A, B) = max a∈A{min b∈B{d(a, b)} } (3.6) Materials and methods 32 In our case, we have computed this metric for each one of our experiments, selecting the patient which DSC was closer to the mean. As our data consists in 3D images, we have measured the shortest distance of any point in the original segmentation to any point of the predicted segmentation, in the three dimensions. Chapter 4 Experiments & Results As we have described in the previous sections, our main goal is to train our network with the source labeled dataset (CT images) and the target unlabeled one (CBCT images), in order to be able to automatically segment the target images with the segmenter, based on the information of the domain invariant features that the domain discriminator provides. In order to achieve that, we have divided this project into four experiments. 1. The Lower Bound 2. The Upper Bound 3. The Unsupervised Domain Adaptation 4. The Unsupervised Domain Adaptation with λ= 0 The first two experiments aim to get information about the worst and the best case scenario in this work, those being train the network only using the segmenter without the domain discriminator, and feeding it with only CT examples and only CBCT examples respectively. The third experiment is the main one, where we train the network using both segmenter and domain discriminator in order to implement the unsupervised domain adaptation process. In that case, we are going to train it with both datasets, even though not both parts of the network will have access to them. And the last one is a variation of the third, to understand how the hyperparameter lambda affects the network. That being said, let’s proceed to describe each one of them in more detail and discuss their results. 33 Experiments & Results 34 4.1 Lower Bound - train on CT With this first experiment we wanted to observe the lowest result that we could expect from our network working only with the segmenter, hence we fed it only with samples from the source CT dataset and we tested it with both CT and CBCT datasets. Figure 4.1: With this experiment, the lower bound, we wanted to have a worst case scenario for our project, by training the network on CT using only the segmenter. We can see how we obtained high dice scores on CT and low ones on CBCT (0.383 on the test set, even though in this experiment both CBCT train and test are seen as test datasets), as they are both from different domains and the segmenter itself is not capable to segment CBCT by only seen CT samples. At the left we can see how the segmentation loss decreases during the training process. Observing the results from 4.1, we can say that the Dice Score Coefficient of the CT dataset (second graph) presents good results, as the final DSC value gets past 0.8 and that’s a result that we have seen in previous works that were also based in CT and CBCT image segmentation. We can appreciate as well that the CT train dataset has a higher DSC value than the test CT dataset. That’s the behaviour we would expect, as the network has never seen the CT test dataset before, but as it is similar to the train one it is able to segment it successfully. Regarding the DSC results of the CBCT datasets (third graph), we can see that the segmenter doesn’t work as accurate as on CT. That’s because CBCT images are not in the same domain as the CT, and therefore, as the segmenter is trained only on CT, its performance on this second group decreases. Experiments & Results 35 4.2 Upper Bound - train on CBCT With this second experiment we wanted to obtain our upper bound, we wanted to know how the network would behave in the best case scenario again, using only the segmenter: training it with CBCT and testing with CBCT. Therefore, we trained it only with CBCT samples, and we tested it on both CT and CBCT datasets, to compare the results obtained with the first experiment. Figure 4.2: With this experiment, the upper bound, we wanted to have the best case scenario for our project, by training the network on CBCT using only the segmenter. We can see how we obtained high dice scores on CBCT (0.717), and surprisingly high ones on CT as well. At the left we can see how the segmentation loss decreases during the training process. By taking a look at the right graph of 4.2, at the DSC results on the CBCT dataset, we can appreciate that they have increased in respect to the lower bound. And the CT ones work fine as well, even though the performance decreases in 0.1. We can settle that the segmenter achieves better results when training on CBCT and testing on CT that not the opposite. One guess to explain this behaviour is that it might be easier to segment CT images because they have a higher quality, a better contrast. Hence the gap between the DSC of source and target datasets when we train with CBCT samples is not as large as when we train on CT. In this case we can see that the CT test results are higher than the CT train, and that’s as well the expected behaviour. In this experiment, as we are training only with CBCT, both CT train and test datasets are seen as test types, as the network hasn’t seen either of them before. Thus it is a random behaviour which one gets greatest values, in this situation it could have been the CT train as well. Experiments & Results 42 Figure 4.5: Comparison between the original bladder segmentations (in green) and the predicted segmentations (in red) of CBCT images. These examples correspond to our three first experiments of one patient with results closer to the mean: lower bound (second column), upper bound (third column) and unsupervised domain adaptation (last column). The first column corresponds to the original CBCT images. For each experiment (column) their respective DSC scores and hausdorff distance values (in voxels) are: 0.378 and 40.93, 0.892 and 33.17, 0.780 and 48.22. Each row corresponds to a different slice of the 3D image, from top to bottom: slice number 0, 4, 9 and 12 (in total each 3D image has 32 slides). We can observe how the second column represents our worst case scenario, the third one represents the best and the last one represents the UDA with a performance between them. As we can see, their DSC values follow the same pattern: the lower bound gets the lowest DSC with 0.378, the upper bound gets the highest with 0.892 and the UDA stays in between with 0.780. What surprises us here is the UDA hausdorff distance, as it should be between both bounds but instead goes higher than the upper bound. In the following figure 4.6 we can observe the variations of the Dice Scores obtained for the CBCT test dataset, as well as the hausdorff distances computed. Experiments & Results 43 If we take a closer look at the DSC values, from left to right, we can appreciate that the source results are the most spread, and that’s what we would expect from the Lower Bound experiment, as it is the worst case scenario of our project. Secondly, moving to the best case scenario, we can see that for the target experiment the Dice values are the most concentrated, as in this case we train on CBCT and test on CBCT. Finally, the distribution of the unsupervised domain adaptation experiment is settled between the source and the target experiments. Regarding the right boxplot, corresponding to the Hausdorff distance computed between the original segmentations and the predicted ones, we can see that the source distances are higher than the target, and that’s the result we would expect, as they are the worst and best case scenarii. What surprises us here is the results from the unsupervised domain adaptation distances, that are higher than the source experiment and have a higher deviation. We are not sure about the reason of this behaviour, but we have to mention that the distance is computed in voxels. So in order to adjust it to mm, we should take the distances from the axis X, Y and Z and multiply them by their pixel spacing, which are respectively 4.8, 4.8 and 6 mm. Figure 4.6: Boxplots of the DSC values distribution (right) and Hausdorff distance in voxels (left) for the 30 test CBCT patient. Chapter 5 Discussion In the previous section we have been sharing and analyzing the results obtained in our three experiments: the lower bound, the upper bound and the unsupervised domain adaptation. Through the following pages, we will discuss the results obtained in our work, and we will compare them with the papers there were based on and some other related projects. First, let’s compare our results with the original Kamnitsas paper on unsupervised domain adaptation with Adversarial Networks [8]: 44 Discussion 45 Figure 5.1: Original results from the Kamnitsas paper [8] about the domain discriminator’s accuracy and the segmenter’s dice values. Figure 5.2: Results from our unsupervised domain adaptation experiment, showing the discriminator’s accuracy and the dice value of the target CBCT dataset across 100 steps. We can observe that they have a similar shapre compared to Kamnitsas results, but ours are more inestable. First of all, we didn’t got to contrast our network’s behaviour placing the domain discriminator at different layers of the segmenter as they proposed (see different colors of the legend in 5.1), even though we have that in mind for the possible future steps of this project. In our case, we tested the network with the discriminator placed at the 6th level of Unet 3.3, so those are the results we are going to analyse. Discussion 46 It is important to remind that our main goal is to achieve a high DSC value when testing on CBCT dataset. It is not necessary for the segmenter to work perfectly and the discriminator to do its worst, we just want to find the best combination of the two networks in order to achieve that high DSC score 2.7 and 2.8. If we take a look at our domain discriminator, what we could expect from it is to not be too close to 1 nor too close to 0. In the first case it wouldn’t find the features invariant to the domain, and on the second one it would mean that the segmenter puts too much effort in fooling the discriminator. So we would expect the discriminator’s accuracy to be between 0.5 and 1, and that’s precisely the range that we’ve obtained in our results. Comparing it with the original one presented in Kamnitsas paper, we can set that they have a similar behaviour, yet their curve is smoother than ours. One reason to explain that might be that we have way less data than them, and also we are working with 3D images. Another reason might be that our domains are much more different that the ones of the original paper, as they are doing the adaptation between two MRI datasets. So for them it’s the same modality but the samples are taken from two different machines, meaning that the domains are much closer to another. In our case we have a bigger challenge, as CT and CBCT are totally different domains. Besides, we are working at a lower resolution, as we explained in 3.1, and also we are working with a smaller batch size for the same reason. So we think those factors can contribute to our instability. The last reason to explain the differences between our domain discriminator’s accuracy and theirs is that we are using Unet architecture for our segmenter and they are using a classical CNN without skip connection. In the paper they are using domain adaptation in a completely fast forward network, only playing with downsampling pass, while we are working with a much complex network that includes skipping connections. That’s a significant difference and that’s the main contribution of our work, as we are studying UDA in a possibly more challenging setting, which might bring instability and more variability amongst others to our code. In the following table we can find the comparison between our DSC and Hausdorff distance results, with other related projects on CT, CBCT and other radiotherapy segmentation methods. Discussion 47 Table 5.1: Dice Score Coefficient and distance comparison between our results and related projects. For our three experiments (lower bound, upper bound and UDA) we have computed the hausdorff distance (HD) in voxles (in order to set them into mm we should multiply each axis by its pixel spacing value). ?SMBD: Symmetric mean boundary distance. †Results reported on a test set containing both CBCT and CT scans. ‡The authors compute the root mean square boundary distance. ∗The authors compute the mean boundary distance. The methods presented in the table stand for the following. DL: deep learning, with their datasets composed by the number of CBCT images (nCBCT ) and the number of CT images (nCT ), DIR: Deformable image registration and RS: RayStation. If we compare our results obtained with the other works we can clearly see that there’s an important gap. However, we believe that they are promising still, and we could expect that at full resolution and with sufficient training data our method could be competitive with others, as we should take into account that our work has faced the following restrictions: firstly, we have trained our network with a smaller dataset, as ours is a subset of the original L´eger et al. dataset; also, as described in 3.1, we are working with images of a lower resolution, as we had to downsample the original ones by a factor of 4 because of our GPUs restrictions 3.2. If we take a look at the Hausdorff distance results, we can observe that the distance in the lower bound is higher than the upper bound, and this is the behaviour that we would expect, as they are respectively the worst and best case scenarios for our project. Discussion 48 What surprises us here is that the distance for the third and main experiment, the unsupervised domain adaptation, is higher than the lower bound, and that is unexpected as it should be between the lower and the upper bound. We are not sure why it behaves like this, but we know how could we improve it as a following steps of this project: we could compute the mean Hausdorff distance in all the layers of the 3D image, by implementing the Symmetric mean boundary distance (SMBD) [6]. SMBD =¯ D(M, P) + ¯ D(P, M) 2(5.1) SMBD computes the mean distance ¯ D(M, P) from the contour of the manual segmentation (M) to the predicted segmentation (P) 3D binary masks, then the mean distance from P to M and finally computes the mean of their sum. The distance D(M, P) stands for: D(M, P) = {minx∈Ωpks(x−y)k, y ∈ΩM}(5.2) where ΩM,ΩPare the boundaries extracted from M and P, and sTstans for the pixel spacing in mm (4.8, 4.8, 6). That is why we think it could bring more accurate information about the results, as it is more robust to the outliers than the Hausdorff distance. Also, another way to improve the Hausdorff distance results could be to look at the 95% percentile of all the distances computed, so we wouldn’t take into account the outliers as well. Chapter 6 Conclusions The main goal of this master thesis is to implement an automatic segmentation of the bladder in CBCT images using deep learning. CBCT images are used in radiotheraphy, a treatment used in patients with cancer, right before the previously planned radiation is delivered to the patient. We believe that by segmenting CBCT images, we can help to improve the dosification of the radiation, conducting its focus to the tumor area so the healthy tissues around it don’t receive a higher amount of what they are supposed to. Unfortunately, we didn’t have a labeled dataset of CBCT images, but what we did have was one of CT images, as this other kind of scan is taken previously at the beginning of the treatment planning and needs to be segmented for the process. So we performed unsupervised domain adaptation between the CT dataset (the source) and the CBCT (the target). In order to do so, we have implemented an architecture that combines the CNN Unet as a segmenter and a regular CNN without skip connections as an adversarial network with the gradient reversal layer, the domain discriminator. With this, we have succesfully acomplished our goal, as by executing UDA (training our network with a source CT dataset and testing it on the target CBCT dataset) we have improved our results compared with training the network only with the source dataset using only a segmenter. If we take a closer look at the last chapter, we can observe that there’s a 20% gap between our results and the mean of the previous projects. We have identified our weaknesses and there are three main reasons to explain that gap, as during this project we have faced 49 Conclusions 50 some remarkable challenges. First of all, we have been working with a small dataset, 120 3D images with a low resolution. Due to our GPU memory limitations 3.2, we had to split the original dataset and downsize it from 160x160x128 to 40x40x32 in order to fit the data into our network. Secondly, we have dealt with a source and target domains that are completely different. Some of the previous works that were also implementing unsupervised domain adaptation were working with two similar domains, as in [8], where they were doing an adaptation between two MRI datasets. And third, we have brought the innovation of using Unet as a segmenter network, which has never been implemented before for this purpose and it has brought some instability to the network. We trust that, by implementing the following steps, we can improve these first results: firstly, we would like to work with a bigger dataset (for example the original [6] from which we extracted our subset), and in order to do so we would need a GPU with more than 11GB of memory. Also, we would like to do more research about some of the hyperparameters like λmax, as we think they might bring more stability to the network. As we have mentioned, we could also implement the symmetric mean boundary distance instead of the Hausdorff distance, to get more accurate information about the predicted segmentations and reevaluate our results. Finally, we could also implement other data augmentation techniques in order to enlarge our dataset and improve our results. As a future work, we would like to follow the steps of [8] and adapt the domain discriminator at different levels of Unet, as in this project we opted to place it only at the 6th. Kamnitsas observed that his method was robust at a specific layer and for a specific value of λmax, and we would like to investigate if this applies to our case as well. Also, we would like to extend this project in order to apply the CBCT segmentation to other male pelvic organs, like the rectum and the prostate, as we already have the datasets of images and masks for them. Hence, even though we obtained lower results than the other previous works shown in Chapter 5, we still believe that using Unet for 3D Cone Beam CT automatic image segmentation can improve the main concern in Radiotherapy regarding the administration of the radiation. By segmenting CBCT images right after they are taken, we could have a clear vision of each organ’s position, and we could use it for: (i) do an online re-planning of the radiation’s distribution for that patient, refocusing the higher radiation to the tumor area so the healthy tissues are less affected or (ii) do a selection of the Conclusions 51 patient, in order to know if he can have an online re-planning or if he needs an offline one, meaning that he must restart the radiation planning process. Also, we believe that this segmentation method that performs unsupervised domain adaptation with an adversarial network and Unet can be useful for many other applications, even outside the medical field. For example, if we switch to the automobile sector, we could apply this technique to self driving cars. In order to be trained to avoid pedestrians and objects in the road they need a large and specific dataset. So, for instance, they could use this method for the training process in different seasons. They could take the summer dataset as a source domain and the winter dataset as a target domain, and apply this UDA method to train the automobile to segment and recognize pedestrians and objects in winter only by seen summer examples. Therefore, overall, we think this is an interesting and promising project that can still go further for new improvements, and that can be applied to many different areas to perform automatic segmentation with unsupervised domain adaptation, for those circumstances where we don’t have enough labeled data to train our network. Bibliography [1] Medical News Today. What to know about radiation therapy?, September 3, 2019. url:https://www.medicalnewstoday.com/articles/158513#types. [2] Mayo Clinic access date May 2020. Radiation therapy, March 28, 2018.url: https : / / www . mayoclinic . org / tests - procedures / radiation - therapy / about/pac-20385162. [3] 2019 National Cancer Institute at the National Institutes of Health January 8. Radiation Therapy to Treat Cancer.url:https://www.cancer.gov/aboutcancer/treatment/types/radiation-therapy. [4] American Cancer Society. Getting External Beam Radiation Therapy, December 27, 2019.url:https://www.cancer.org/treatment/treatments-and-side- effects/treatment-types/radiation/external-beam-radiation-therapy. html. [5] SimpleElastix. Non-rigid Registration, Mars 2020.url:https://simpleelastix. readthedocs.io/NonRigidRegistration.html. [6] Jean L´eger, Eliott Brion, Paul Desbordes, Christophe De Vleeschouwer, and John Aldo Lee. “Cross-domain data augmentation for deep-learning-based male pelvic organ segmentation in cone beam CT”. In: ”Applied Sciences” (2020). url:http: //hdl.handle.net/2078.1/226998. [7] Yaroslav Ganin and Victor Lempitsky. Unsupervised Domain Adaptation by Backpropagation. 2014. arXiv: 1409.7495 [stat.ML]. [8] Konstantinos Kamnitsas et al. “Unsupervised domain adaptation in brain lesion segmentation with adversarial networks”. In: IPMI. 2017. [9] Thomas M. Mitchell. Machine Learning. 1st ed. USA: McGraw-Hill, Inc., 1997. isbn: 0070428077. 58 Bibliography 59 [10] Jeremy Jordan. Machine learning overview., May 29,2017.url:https://www. jeremyjordan.me/machine-learning-overview/. [11] Randy Lao. A Beginner’s Guide to Machine Learning, July 3, 2018.url:https: // towardsdatascience. com/ abeginners - guide- to - machinelearning - 5d87d1b06111. [12] Machine Learning Mastery. Machine Learning Mastery, 2019.url:https : / / machinelearningmastery.com/start-here/#algorithms. [13] Daffodil Software. 9 Applications of Machine Learning from Day-to-Day Life, July 31, 2017.url:https : / / medium . com / app - affairs / 9 - applications - of - machine-learning-from-day-to-day-life-112a47a429d0. [14] Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep Learning.http:// www.deeplearningbook.org. MIT Press, 2016. [15] Jordi Torres. Deep Learning: Introducci´on pr´actica. UPC, 2018. url:https : // torres .ai / artificialintelligence - content /first - contact - deeplearning-practical-introduction-keras/. [16] Jason Brownlee. Difference Between a Batch and an Epoch in a Neural Network, July 20, 2018.url:https : / / machinelearningmastery . com / difference - between-a-batch-and-an-epoch/. [17] Bottou L. Large-Scale Machine Learning with Stochastic Gradient Descent. In: Lechevallier Y., Saporta G. (eds) Proceedings of COMPSTAT ’ 2010. Physica- Verlag HD, 2010. isbn: 9783790826036. [18] Bernard Marr. What Are Artificial Neural Networks - A Simple Explanation For Absolutely Anyone, September 24, 2018.url:https://www.forbes.com/sites/ bernardmarr / 2018 / 09 / 24 / what - are - artificial - neural - networks - a - simple-explanation-for-absolutely-anyone/#7ecd1ec51245. [19] Wikipedia. Artificial Neural Network, April 8, 2020.url:https://en.wikipedia. org/wiki/Artificial_neural_network. [20] Jason Brownlee. When to Use MLP, CNN, and RNN Neural Networks, July 23, 2018.url:https://machinelearningmastery.com/when-to-use-mlp-cnn- and-rnn-neural-networks/. Bibliography 60 [21] Uniqtech. Multilayer Perceptron (MLP) vs Convolutional Neural Network in Deep Learning, Dec 22, 2018.url:https://medium.com/data-science-bootcamp/ multilayer-perceptron-mlp-vs-convolutional-neural-network-in-deep- learning-c890f487a8f1. [22] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. “U-Net: Convolutional Networks for Biomedical Image Segmentation”. In: arxiv (2015). url:https : //arxiv.org/pdf/1505.04597.pdf. [23] Perelman School of Medicine at the University of Pennsylvania. MICCAI BraTS 2017: Scope — Section for Biomedical Image Analysis (SBIA), Retrieved 2018-12- 24. url:www.med.upenn.edu. [24] Retrieved 2018-12-24. Perelman School of Medicine at the University of Pennsylvania. SLIVER07 : Home.url:www.sliver07.orgu. [25] Wikipedia. Sørensen–Dice coefficient, April 2020.url:https://en.wikipedia. org/wiki/S%5C%C3%5C%B8rensen%5C%E2%5C%80%5C%93Dice_coefficien. [26] Wikipedia. F1 score, April 2020.url:https://en.wikipedia.org/wiki/F1_ score. [27] Normand Gr´egoire and Mikael Bouillot. Hausdorff Distance between convex polygons, May 2020.url:http://cgm.cs.mcgill.ca/~godfried/teaching/cgprojects/98/normand/main.html. [28] Jason Brownlee. A Gentle Introduction to Cross-Entropy for Machine Learning, October 21, 2019.url:https://machinelearningmastery.com/cross-entropy- for-machine-learning/. [29] Tim Hartley. When Parallelism Gets Tricky: Accelerating Floyd-Steinberg on the Mali GPU, November 25, 2014.url:https://community.arm.com/developer/ tools-software/graphics/b/blog/posts/when-parallelism-gets-tricky- accelerating-floyd-steinberg-on-the-mali-gpu. [30] DeepAI. Max Pooling, Mars 2020.url:https://deepai.org/machine-learning- glossary-and-terms/max-pooling. [31] ETSETB TelecomBCN Universitat Politecnica de Catalunya. Deep Learning for Artificial Intelligence, Autumn 2019.url:https://telecombcn-dl.github.io/ dlai-2019/. UNIVERSITÉ CATHOLIQUE DE LOUVAIN École polytechnique de Louvain Rue Archimède, 1 bte L6.11.01, 1348 Louvain-la-Neuve, Belgique | www.uclouvain.be/epl