scieee AI-readable full text Open interactive document viewer

Exploring adversarial attacks and defenses

Branco, Tiago Manuel Sampaio

Abstract

Deep Learning classifiers are capable of an outstanding performance. Yet, they are vulnera ble to adversarial attacks, i.e. it is possible to craft a slightly modified version of a correctly classified image that, although its contents are still clearly recognisable to a human being, the classifier outputs an incorrect classification. In this thesis we evaluate the effectiveness of adversarial attacks, namely their trans ferability to other models, and some proposed defenses. Transferability occurs when an adversarial sample is crafted with a model, and it succeeds in achieving a misclassification in another model. To make this study as comprehensive as possible, we explore several attack methods, namely: Fast Gradient Sign Method (FGSM), Deepfool, Jacobian Saliency Map Attack (JSMA), Carlini, Projected Gradient Descent (PGD) and Few Pixels. To evaluate the impact of the model’s architecture in the transferability rate we use sev eral common architectures: VGG16, three ResNet with different depths, and a small Con volution Neural Network. Two common datasets were used for evaluation: CIFAR-10 and German Traffic Sign Recognition Benchmark (GTSRB). Different attack methods use different approaches and parameters to craft adversarial samples. Hence, it is not trivial to control the degree of perturbation. To be able to achieve the same level of perturbation with every method we resorted to an image comparison metric: Structural Similarity Index Measure (SSIM). For each method we performed a search within its parameter space to find the parameters that on average attain a specific level of perturbation. To evaluate the impact of the level of perturbation on transferability rates, we evaluate two different values for the SSIM metric. Our results show that while it is possible to craft an adversarial sample in a particular model, the transferability rates vary considerably from method to method. Regarding defensive methods we explored Adversarial Training and Defensive Distilla tion. The results show that the ability to prevent an adversarial attack, or robustness, varies significantly depending on the conditions that the attack is performed and on the defensive methods used. Furthermore, there is a trade-off between robustness and accuracy, with defensive models having lower accuracy than non-defended models.

Full text

Universidade do Minho Escola de Engenharia Departamento de Informática Tiago Manuel Sampaio Branco Exploring Adversarial Attacks and Defenses February 2024 Universidade do Minho Escola de Engenharia Departamento de Informática Tiago Manuel Sampaio Branco Exploring Adversarial Attacks and Defenses Master dissertation Master Degree in Master´s Informatics Engineering Dissertation supervised by António José Borba Ramires Fernandes February 2024 COPYRIGHT AND TERMS OF USE FOR THIRD PARTY WORK This dissertation reports on academic work that can be used by third parties as long as the internationally accepted standards and good practices are respected concerning copyright and related rights. This work can thereafter be used under the terms established in the license below. Readers needing authorization conditions not provided for in the indicated licensing should contact the author through the RepositóriUM of the University of Minho. license granted to users of this work: CC BY https://creativecommons.org/licenses/by/4.0/ i ACKNOWLEDGEMENTS I would like to express my gratitude to my primary supervisor, António Ramires, who guided me throughout this project. I would also like to thank my friends and family who supported me and offered deep insight into the study. ii STATEMENT OF INTEGRITY I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the University of Minho. iii ABSTRACT Deep Learning classifiers are capable of an outstanding performance. Yet, they are vulnerable to adversarial attacks, i.e. it is possible to craft a slightly modified version of a correctly classified image that, although its contents are still clearly recognisable to a human being, the classifier outputs an incorrect classification. In this thesis we evaluate the effectiveness of adversarial attacks, namely their transferability to other models, and some proposed defenses. Transferability occurs when an adversarial sample is crafted with a model, and it succeeds in achieving a misclassification in another model. To make this study as comprehensive as possible, we explore several attack methods, namely: Fast Gradient Sign Method (FGSM), Deepfool, Jacobian Saliency Map Attack (JSMA), Carlini, Projected Gradient Descent (PGD) and Few Pixels. To evaluate the impact of the model’s architecture in the transferability rate we use several common architectures: VGG16, three ResNet with different depths, and a small Convolution Neural Network. Two common datasets were used for evaluation: CIFAR-10 and German Traffic Sign Recognition Benchmark (GTSRB). Different attack methods use different approaches and parameters to craft adversarial samples. Hence, it is not trivial to control the degree of perturbation. To be able to achieve the same level of perturbation with every method we resorted to an image comparison metric: Structural Similarity Index Measure (SSIM). For each method we performed a search within its parameter space to find the parameters that on average attain a specific level of perturbation. To evaluate the impact of the level of perturbation on transferability rates, we evaluate two different values for the SSIM metric. Our results show that while it is possible to craft an adversarial sample in a particular model, the transferability rates vary considerably from method to method. Regarding defensive methods we explored Adversarial Training and Defensive Distillation. The results show that the ability to prevent an adversarial attack, or robustness, varies significantly depending on the conditions that the attack is performed and on the defensive methods used. Furthermore, there is a trade-off between robustness and accuracy, with defensive models having lower accuracy than non-defended models. Keywords: Adversarial Attacks, FGSM, DeepFool, JSMA, PGD, Carlini, Adversarial Training, Defensive Distillation, SSIM, CIFAR-10, GTSRB iv RESUMO Modelos classificativos de Deep Learning são capazes de performances extraordinárias, superando mesmo a capacidade humana, no entanto são vulneráveis a ataques adversariais, por exemplo é possível criar uma versão modificada de uma imagem bem classificada, que apesar do ser facilmente reconhecível por um ser humano, o modelo classifica incorrectamente. Nesta tese avaliamos a transferabilidade dos ataques adversariais, e algumas propostas de defesa. A transferabilidade occorre quando um adversarial gerado por um modelo é capaz de fazer outro modelo o classificar incorrectamente. Para fazer este estudo o mais compreensivo possível, explorámos vários tipos de métodos de ataques, nomeadamente: FGSM, Deepfool, JSMA, Carlini, PGD and Few Pixels. Para avaliar o impacto da arquitectura do modelo na taxa de transferabilidade, usamos várias arquitecturas conhecidas: VGG16, 3 ResNet com diferentes números de camadas, e uma rede pequena de 3camadas convolucionais. Usamos 2 datasets conhecidos para esta avaliação: CIFAR-10 eGTSRB. Diferentes métodos de ataques usam abordagens e parâmetros diferentes para gerar exemplos adversariais. Por isso, não é trivial controlar o grau de perturbação gerada pelos exemplos adversariais. Para sermos capazes de alcançar o mesmo nível de perturbação com todos os métodos recorremos a uma métrica de comparação de imagem: SSIM. Para cada método fizemos uma procura no domínio dos parâmetros para encontrar os valores que em média geram um nível específico de perturbação. Para avaliar o impacto do nível de perturbação na taxa de transferabilidade, avaliámos 2 valores diferentes para a métrica SSIM. Os nossos resultados mostram que enquanto é possível gerar um exemplo adversarial num modelo particular, a taxa de transferabilidade varia consideravelmente de método para método. Em relação às defesas exploramos Adversarial Training e Defensive Distillation. Os resultados mostram que a capacidade de prevenir um ataque adversarial varia significativamente dependendo das condições em que o ataque é feito e nos próprios métodos. Palavras-Chave: Ataques Adversariais, FGSM, DeepFool, JSMA, PGD, Carlini, Treino Adversarial, Defesa via Destilação, SSIM, CIFAR-10, GTSRB v CONTENTS 1 introduction 1 1.1motivation and objectives 2 1.2document structure 2 2 state of the art 4 2.1Adversarials Attacks 4 2.1.1Targeted vs. non-targeted attacks 5 2.1.2From White to Black Box Attacks 5 2.1.3Fast Gradient Sign Method and some variations 6 2.1.4Deepfool 8 2.1.5Jacobian Saliency Map Attack 11 2.1.6Carlini Method 16 2.1.7Projected Gradient Descent (PGD) 18 2.1.8One Pixel Attack 19 2.2Defense methods against Adversarial Attacks 21 2.2.1Adversarial Training 21 2.2.2Defensive Distillation 22 3 attack evaluation 25 3.1Datasets 25 3.1.1CIFAR-10 25 3.1.2German Traffic Sign Recognition Benchmark 26 3.2Models and training procedure 27 3.2.1VGG16 27 3.2.2ResNets 28 3.2.3ConvNet 28 3.3Building the adversarial sample set 28 3.4CIFAR-10: Adversarial Attacks 31 3.4.1Procedure details and attack parameters 31 3.4.2FGSM 35 3.4.3Deepfool 40 3.4.4JSMA 42 3.4.5Carlini 45 3.4.6PGD 48 3.4.7One Pixel/Few Pixels 52 3.4.8Adversarial attack comparison 54 vi Contents vii 3.5GTSRB: Adversarial attacks 64 3.5.1Procedural details and attack parameters 65 3.5.2FGSM 66 3.5.3Deepfool 71 3.5.4Carlini 73 3.5.5PGD 75 3.5.6Adversarial attack comparison 80 3.6Attack Analysis 87 4 defense evaluation 89 4.1Adversarial Training 89 4.2Defensive Distillation 90 4.3Defenses experiments settings 91 4.4Adversarial Defenses experiments on CIFAR-10 92 4.4.1Non-targeted environment 93 4.4.2Targeted environment 95 4.5Adversarial Defenses experiments on GTSRB 97 4.5.1Non-targeted environment 98 4.5.2Targeted environment 100 4.6Defense Analysis 103 5 conclusion 107 5.1Prospect for future work 108 a attack evaluation graphs 113 a.1CIFAR-10 search parameters 113 a.1.1FGSM 113 a.1.2Deepfool 117 a.1.3Carlini 120 a.1.4PGD 121 a.1.5One pixel/Few pixels 125 a.2GTSRB search parameters 127 a.2.1FGSM 127 a.2.2Deepfool 132 a.2.3Carlini 133 a.2.4PGD 135 List of Figures xiv Figure 97 Success attack rate of targeted PGD L2norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e124 Figure 98 Success attack rate of non-targeted Few pixels attack on each model and their corresponding SSIM and total successful adversarial attacks on model generator per number of pixels perturbed 125 Figure 99 Success attack rate of targeted Few pixels attack on each model and their corresponding SSIM and total successful adversarial attacks on model generator per number of pixels perturbed 126 Figure 100 Success attack rate of non-targeted FGSM L∞norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e127 Figure 101 Success attack rate of non-targeted FGSM L1norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e128 Figure 102 Success attack rate of non-targeted FGSM L2norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e128 Figure 103 Success attack rate of targeted for closest class FGSM L∞norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e129 Figure 104 Success attack rate of targeted for closest class FGSM L1norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e129 Figure 105 Success attack rate of targeted for closest class FGSM L2norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e130 Figure 106 Success attack rate of targeted for farthest class FGSM L∞norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e130 Figure 107 Success attack rate of targeted for farthest class FGSM L1norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e131 List of Figures xv Figure 108 Success attack rate of targeted for farthest class FGSM L2norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e131 Figure 109 Success attack rate of Deepfool on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e132 Figure 110 Non-targeted Carlini L2norm success attack rate on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per k133 Figure 111 Targeted Carlini L2for closest class success attack rate on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per k134 Figure 112 Targeted Carlini L2for farthest class success attack rate on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per k134 Figure 113 Success attack rate of non-targeted PGD L∞norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e135 Figure 114 Success attack rate of non-targeted PGD L1norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e136 Figure 115 Success attack rate of non-targeted PGD L2norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e136 Figure 116 Success attack rate of targeted for closest class PGD L∞norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e137 Figure 117 Success attack rate of targeted for closest class PGD L1norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e137 Figure 118 Success attack rate of targeted for closest class PGD L2norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e138 List of Figures xvi Figure 119 Success attack rate of targeted for farthest class PGD L∞norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e138 Figure 120 Success attack rate of targeted for farthest class PGD L1norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e139 Figure 121 Success attack rate of targeted for farthest class PGD L2norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e139 LIST OF TABLES Table 1Average accuracy and standard deviation on test set for all models trained on CIFAR-10 dataset and their corresponding number of trainable parameters. 31 Table 2CIFAR-10 data augmentation 31 Table 3VGG16 model training parameters on CIFAR-10 33 Table 4ResNet models training parameters on CIFAR-10 33 Table 5ConvNet model training parameters on CIFAR-10 34 Table 6ConvNet ReduceLROnPlateau callback parameters 34 Table 7FGSM non-targeted average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm 38 Table 8targeted FGSM average, SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm 40 Table 9Deepfool average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks 41 Table 10 NT-JSMA average SSIM, theta, gamma, success attack rate, success attacks after conversion and total successful attacks 45 Table 11 Targeted JSMA average SSIM, theta, gamma, success attack rate, success attacks after conversion and total successful attacks 45 Table 12 Non-targeted Carlini L2norm average SSIM, k, c, success attack rate, success attacks after conversion and total successful attacks 47 Table 13 Targeted Carlini L2norm average SSIM, k, c, success attack rate, success attacks after conversion and total successful attacks 48 Table 14 Non-targeted PGD average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm 52 Table 15 targeted PGD average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm 52 Table 16 Non-targeted Few Pixels attack average SSIM, number of pixels perturbed and success attack rate for targeted and non-targeted setting 53 xvii List of Tables xviii Table 17 CIFAR-10 non-targeted transferability table. I% is the initial attacks successful on generator, C% the percentage of attacks successful on generator after conversion to valid image, and F% final percentage of successful attacks on generator that were used to attack the other networks, this is a pure white box attack. Column TA is the transferability average of an attack across all networks. In bold are the highest values per attack, in green and blue cells are the highest transferability values of all attacks per model architecture for the L and S setting, respectively 57 Table 18 SSIM statistics of crafting examples on model generator for each CIFAR-10 non-targeted adversarials attacks using the parameters found in Section 3.4.260 Table 19 CIFAR-10 targeted transferability table. I% is the initial attacks successful on generator, C% the percentage of attacks successful on generator after conversion, and F% final percentage of successful attacks on generator that were used to attack the other networks, this is a pure white box attack. Column TA is the transferability average of an attack across all networks. In bold are the highest values per attack, and in green and blue cells are the highest transferability values of all attacks per model architecture, for the L and S setting, respectively 64 Table 20 Average accuracy and standard deviation on test set for all models trained on GTSRB dataset and their corresponding number of trainable parameters. 65 Table 21 VGG16 model callback used for training on GTSRB dataset 65 Table 22 GTSRB data augmentation 66 Table 23 Closest and farthest class by SSIM metric for each class. 67 Table 24 FGSM non-targeted average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm 71 Table 25 FGSM targeted attack for closest class average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm 72 Table 26 FGSM targeted attack for farthest class average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm 72 Table 27 Deepfool average SSIM, epsilons and transferability success rate 73 List of Tables xix Table 28 Non-targeted Carlini L2norm average SSIM, k, c, success attack rate, success attacks after conversion and total successful attacks 75 Table 29 Carlini L2targeted for closest and farthest classes, average SSIM, k, c, success attack rate, success attacks after conversion and total successful attacks 75 Table 30 PGD non-targeted average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm 79 Table 31 PGD targeted attack for closest class average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm 79 Table 32 PGD targeted attack for farthest class average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm 79 Table 33 GTSRB non-targeted transferability table. I% is the initial attacks successful on generator, C% the percentage of attacks successful on generator after conversion to valid image, and F% final percentage of successful attacks on generator that were used to attack the other networks, this is a pure white box attack. Column TA is the transferability average of an attack across all networks. In bold are the highest values per attack and in green and blue cells are the highest transferability values of all L and S attacks respectably, per model architecture 82 Table 34 GTSRB targeted for closest class transferability table. I% is the initial attacks successful on generator, C% the percentage of attacks successful on generator after conversion to valid image, and F% final percentage of successful attacks on generator that were used to attack the other networks. Column TA is the transferability average of an attack across all networks. In bold are the highest values per attack and in green and blue cells are the highest transferability values of all L and S attacks respectably, per model architecture 84 List of Tables xx Table 35 GTSRB targeted for farthest class transferability table. I% is the initial attacks successful on generator, C% the percentage of attacks successful on generator after conversion to valid image, and F% final percentage of successful attacks on generator that were used to attack the other networks, this is a pure white box attack. Column TA is the transferability average of an attack across all networks. In bold are the highest values per attack and in green and blue cells are the highest transferability values of all L and S attacks respectably, per model architecture 87 Table 36 PGD parameters used in Adversarial Training 90 Table 37 Defensive models and normal ResNet-50 average accuracies on CIFAR10 test set and corresponding standard deviation 93 Table 38 CIFAR-10 non-targeted transferability table. I% is the initial percentage of attacks successful on generator, C% is the percentage of successful attacks on generator after conversion to valid image, and F% is the final percentage of successful attacks on generator that were used to attack the other networks, this is a pure white box attack. In bold are the highest values per attack and in green and blue cells are the highest transferability values of all L and S attacks respectably per, model architecture. Attacks with * are those with norms that attained best transferability rates in Section 3.495 Table 39 CIFAR-10 targeted transferability table. I% is the initial percentage of attacks successful on generator, C% the percentage of attacks successful on generator after conversion to valid image, and F% final percentage of successful attacks on generator that were used to attack the other networks, this is a pure white box attack. In bold are the highest values per attack, and in green and blue cells are the highest transferability values of all L and S attacks respectably per model architecture. Attacks with * are those with norms that attained best transferability rates on Section 3.497 Table 40 Defensive models and normal ResNet-50 average accuracies on test set and corresponding standard deviation for GTSRB 98 List of Tables xxi Table 41 GTSRB non-targeted transferability table. I% is the initial attacks successful on generator, C% the percentage of attacks successful on generator after conversion to valid image, and F% final percentage of successful attacks on generator that were used to attack the other networks, this is a pure white box attack. In bold are the highest values per attack and in green and blue cells are the highest transferability values of all L and S attacks respectably per, model architecture. Attacks with * are those with norms that attained best transferability rates on Section 3.5100 Table 42 GTSRB targeted to closest class transferability table. I% is the initial percentage of attacks successful on generator, C% the percentage of attacks successful on generator after conversion to valid image, and F% final percentage of successful attacks on generator that were used to attack the other networks, this is a pure white box attack. In bold are the highest values per attack, and in green and blue cells are the highest transferability values of all L and S attacks respectably per model architecture. Attacks with * are those with norms that attained best transferability rates on Section 3.5102 Table 43 GTSRB targeted to farthest class transferability table. I% is the initial percentage of attacks successful on generator, C% the percentage of attacks successful on generator after conversion to valid image, and F% final percentage of successful attacks on generator that were used to attack the other networks, this is a pure white box attack. In bold are the highest values per attack, and in green and blue cells are the highest transferability values of all L and S attacks respectably per model architecture. Attacks with * are those with norms that attained best transferability rates on Section 3.5104 LIST OF ABBREVIATIONS CNN Convolutional Neural Network. 1,4,26,27 DE Differential Evolution. 19,20,21 DNN Deep Neural Network. 4 FGSM Fast Gradient Sign Method. iv,v,viii,ix,x,xi,xii,xiii,xiv,xvii,xviii,6,7,8,11,18, 22,28,31,34,35,36,38,40,41,54,64,65,66,68,69,70,71,72,80,84,86,87,89,91, 94,95,96,97,101,106,113,114,115,126,127,128,129,130,131 GTSRB German Traffic Sign Recognition Benchmark. iv,v,viii,xi,xii,xviii,xix,xx,xxi,2, 18,25,26,27,64,65,72,81,84,86,87,89,90,91,92,97,98,99,100,101,103,104, 105,113 JSMA Jacobian Saliency Map Attack. iv,v,viii,ix,x,xiii,xvii,11,15,18,29,31,42,43,44, 45,54,56,58,63,65,117,118 PGD Projected Gradient Descent. iv,v,viii,ix,x,xi,xiii,xv,xvi,xvii,xix,xx,18,19,29,31, 48,49,50,51,53,54,58,61,63,64,65,75,76,77,79,80,82,86,87,89,90,91,94,96, 99,101,104,106,107,120,121,122,123,134,135,136,137,138,139 SSIM Structural Similarity Index Measure. iv,v,viii,ix,x,xi,xii,xiii,xiv,xv,xvi,xvii, xviii,xix,2,28,29,30,31,35,36,35,37,36,38,39,40,41,42,43,44,45,46,47,48, 49,50,51,50,51,52,53,54,55,58,61,64,65,66,68,69,70,71,72,73,74,73,74,75, 76,77,78,79,80,82,83,85,86,87,90,93,100,101,104,107,108,113,114,115,116, 117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134, 135,136,137,138,139 xxii 1 INTRODUCTION It seems like everywhere one looks sees the words deep learning and neural networks. The ubiquity of these terms in the technological world is a reality. In the last few years, deep learning has led to very good performance on a broad variety of problems, such as visual recognition, speech recognition, and natural language processing. On the other hand, such a success will naturally provoke research on how to fool these systems. These are called adversarial attacks, consisting, in summary, in finding small perturbations to images such that these adulterated images end up being misclassified. This can cause major problems in some domains, for instance autonomous driving. Consider, for instance, that a Stop sign is adulterated in such way that a deep learning system sees it as a Right-of-way at an intersection. These perturbations can be made so small that the original image and the adversarial attack are indistinguishable to the human eye Goodfellow et al. (2015). Even visible perturbations may have the appearance of noise, making it harder for humans to detect adversarial images. Several research studies have shown that Convolutional Neural Network (CNN)s can be led to misclassify these adversarial images with a variable degree of success Moosavi-Dezfooli et al. (2016)Su et al. (2019), although a human would have still correctly classify them with high confidence. An adversarial attack can occur in a broad range of scenarios Zhang et al. (2020). The easiest being when the attacker has full knowledge of the object system it is attacking. In such a scenario, the attacker has access to the model architecture, its weights, and the datasets that were used to train the model. Such a scenario is called a white box attack. On the other extreme, the attacker has no knowledge about the system. This implies that the attacker must select a model, which is potentially different from the object model, and train it with a potentially different dataset. This is a black box attack. Between these two extremes there is a wide range of scenarios where the attacker has partial knowledge about the system, for instance, the dataset used to train the object model can be known. Attacks can also aim at not only causing a misclassification, but lead the model to output a classification in a particular class. This is a targeted attack. These are harder to achieve than non-targeted attacks, but nonetheless possible. 1 2.1. Adversarials Attacks 8 ClipX,e{XN}performs per-pixel clipping of the image XN, to limit pixel alterations to [xi− e,xi+e]. The authors of Kurakin et al. (2017b) used α=1 considering the range [0, 255], and defined the number of iterations to be min(e+4, 1.25e). They chose this function heuristically since it proved sufficient, yet restricted enough to keep the computational cost manageable. Target Class FGSM FGSM as presented before is a non-targeted attack. This version of the FGSM aims at make a model misclassify the adversarial samples as some specific class. This is achieved by, instead of increasing the loss for the original class, decreasing the loss for a particular target class. To increase the probability of some target class, p(ytarget|X), we perform the gradient descent, as shown in Equation 6 X∗=X−e·sign(∇XJ(X,ytarget)) (6) A variation of this method, as suggested in Kurakin et al. (2017a), is the least likely class method, where it is chosen the least likely class from the model prediction as the target class, yLL =arg min y {p(y|X)}. , thus we get Equation 7. X∗=X−e·sign(∇XJ(X,yLL)) (7) The targeted version of FGSM can also be defined as an iterative approach as seen in Equation 8. X0=X,XN+1=ClipX,e{XN−α·sign(∇XJ(XN,y∗))}(8) where y∗can be either ytarget or yLL. 2.1.4Deepfool Deepfool (Moosavi-Dezfooli et al. (2016)) is a non-targeted method that aims at minimizing the amount of perturbation required to craft an adversarial image by using an iterative linearization approach. In Figure 2we can see a comparison between FGSM and Deepfool adversarials and their corresponding perturbation, Deepfool leads to a smaller perturbation Moosavi-Dezfooli et al. (2016), found to be estimated as 5 times smaller than FGSM. The main idea of this attack is to take steps in the direction of the closest decision boundary until it crosses to the other side thus making the network misclassify the image as another class. To achieve this the authors assume that the neural networks classes are sep- 2.1. Adversarials Attacks 9 Figure 2.: Example of adversarial perturbations. First row: the original image Xthat is classified as "whale". Second row: the image X+rclassified as "turtle" and the corresponding perturbation rcomputed by Deepfool. On third row: the image classified as "turtle" and corresponding perturbations computed using FGSM. Note that these perturbations are amplified for visualization purposes. Source:Moosavi-Dezfooli et al. (2016) arated by hyper-planes. This is a linear approximation and the iterative behaviour of the algorithm allows to reach the other side of the decision boundary. In Figure 3we can see this linear approximation between the classes boundaries and the real ones. Figure 3.: Decision boundaries linear approximation and real boundaries of a model. Source:Moosavi-Dezfooli et al. (2016) 2.1. Adversarials Attacks 10 To calculate the distance between the image the closest decision boundary we perform an orthogonal projection. In Figure 4we can see this projection on a binary classifier, where F=x:wTx+b=0 is the hyperplane, x0is the data point, and ∆(x0;f)is the robustness of function fat point x0which is the smallest amount of perturbation needed to arrive to the decision boundary. To change the classification of data point x0we must add a small step to make the data point cross the boundary. This distance is calculated using the orthogonal projection of x0onto Fas seen in Equation 9, where wcorresponds to the gradient of function fwith respect to the data point x0. Figure 4.: Distance between a sample and the boundary on a linear classifier. Source:MoosaviDezfooli et al. (2016) r∗(x0):=arg minkrk2 subject to sign(f(x0+r)) 6=sign(f(x0)) =−f(x0) kwk2 2 w. (9) At each iteration f, the binary classifier, is linearized around the current point xiand the minimal perturbation of linearized classifier is computed as seen in Equation 10. arg min ri krik2subject to f(xi) + ∇f(xi)Tri=0. (10) Where perturbation riat iteration iis calculated using Equation 9. This can be generalized for the multiclass non linear classifiers, where we first need to calculate the closest decision boundary, and then apply the perturbation on that direction 2.1. Adversarials Attacks 11 that corresponds to the minimum perturbation r∗(x0)which is the vector that projects x0 on the hyperplane indexed by ˆ l(x0). ˆ l(x0) = arg min k6=ˆ k(x0) |fk(x0)−fˆ k(x0)(x0)| k∇ fk(x0)− ∇fˆ k(x0)(x0)k2 . (11) To calculate the closest decision boundary we use Equation 11, where ˆ kcorresponds to the class with highest score, and to calculate the perturbation we use Equation 12, where the minimum perturbation r∗(x0)is the vector that projects x0on the hyperplane indexed by ˆ l(x0)found in Equation 11. ˆ r∗(x0) = |fl(x0)−fˆ k(x0)(x0)| k∇ fl(x0)− ∇fˆ k(x0)(x0)k2 2 ·(∇fl(x0)− ∇fˆ k(x0)(x0)) (12) The pseudo code of this algorithm can be seen in Algorithm 1. Note that this algorithm operates in a greedy way and is not guaranteed to converge to the optimal perturbation. Algorithm 1:Deepfool: multi-class case Input: Image x, classifier f. Output: Perturbation ˆ r. 1Initialize x0←x,i←0. 2while ˆ k(xi) = ˆ k(x0)do 3for k6=ˆ k(x0)do 4w0 k← ∇fk(xi)− ∇fˆ k(x0)(xi) 5f0 k←fk(xi)−fˆ k(x0)(xi) 6ˆ l←arg min k6=ˆ k(x0) |f0 k| kw0 kk2 7ri←|f0 ˆ l| kw0 ˆ lk2 2 w0 ˆ l 8xi+1←xi+ri 9i←i+1, 10 return ˆ r=∑iri. 2.1.5Jacobian Saliency Map Attack JSMA is a class of algorithms to craft adversarial samples. Below is described the original JSMA algorithm based on Papernot et al. (2015), also known as Targeted JSMA, and the non-targeted attack variant introduced in Wiyatno and Xu (2018), where it removes the requirement to specify a target class as required in the original JSMA. 2.1. Adversarials Attacks 12 Figure 5.: Targeted adversarial samples using JSMA for different digits from MNIST dataset. Source:Papernot et al. (2015) The main difference regarding the result of this method when compared to FGSM, is that in a single iteration FGSM can potentially affect all pixels, whereas JSMA will only affect up to two pixels. Although only a small number of pixels may get disturbed, these perturbations end up as being clearly visible as will be shown, for instance, in 3.4.4. Targeted JSMA This method starts by finding the contribution of each pixel of the input image to each class. A selection of the pixels that positively influences the output for the given class is then gathered, with some perturbation being applied to those selected pixels. Some examples of this targeted attack can be seen in Figure 5. To calculate the contribution of each pixel from the input image to each class of the network output we can calculate the Jacobian matrix to obtain the partial derivatives. The Jacobian matrix give us the information about how each input pixel, or feature, influences the output, where large positive values mean that increasing the respective pixels will yield a large increase in the output by the model. Conversely, components with large negative values correspond to pixels that yield large decreases in the output by the model when their value is increased. In the second stage of the process, we must determine the best pixel candidates to be perturbed to guide the model into misclassifying an input image in a chosen target class. 2.1. Adversarials Attacks 13 To achieve this, a saliency map is computed, where each pixel is given a score based on how much impact that pixel will have on the target class and in all other classes. This pixel score is defined as the product of the gradient of the target class and the sum of the gradients of all other classes. We exclude pixels whose derivative of target class is negative or whose sum of the gradients of all other classes is positive. This ensures that we exclude pixels which negative impact the target class and those who have an overall positive contribution to the all other classes. S+(X,t)[i] =    0 if ∂Ft(X) ∂Xi <0 or ∑j6=t ∂Fj(X) ∂Xi >0 ∂Ft(X) ∂Xi∑j6=t ∂Fj(X) ∂Xiotherwise (13) In this saliency map, as presented in Equation 13,iis an input feature, and F(X)refers to the last hidden layer logits output. The condition on the first line sets the score to zero for input features that have a negative target derivative, or an overall positive derivative on other classes. The derivative ∂Ft(X) ∂Ximust be positive in order for Ft(X)to increase when the feature Xiincreases and analogous, the sum of all the derivatives of the other classes, ∑j6=t ∂Fj(X) ∂Xi, must be negative, or stay constant, so that if we increase the value of the input Xiit will not influence positively the output in those classes. Also, that is why the absolute value is taken in the second line, because the value should be negative. Thus, with the second line we can warrant that the pixels with highest score are the ones that increase the target class, or decrease all other classes significantly or both. This saliency map considers increasing the value of the input features in order to achieve misclassification. The counter part to this saliency map is one that considers decreasing the input features value instead to attain the misclassification. It’s follows the same construction principle, the only difference being in the inequalities that are flipped and the location of the absolute value in the second line as seen in Equation 14. S−(X,t)[i] =    0 if ∂Ft(X) ∂Xi >0 or ∑j6=t ∂Fj(X) ∂Xi <0  ∂Ft(X) ∂Xi∑j6=t ∂Fj(X) ∂Xiotherwise (14) The saliency maps corresponding to increase and decrease of pixel values are expressed as S+and S−respectively. In the last step in the algorithm, after finding the input feature with the highest score in the positive saliency map, we perturbed the pixels to realize the adversary sample. In each iteration of the algorithm this step is repeated until the misclassification is achieved. In Papernot et al. (2015) the authors proposed a more robust heuristic that instead of searching for the best candidate pixel it searches the best pair candidate. This approach is preferable because selecting two pixels instead of one is less strict and very few pixels 2.1. Adversarials Attacks 14 would meet the criteria when using the initial heuristic. This new search heuristic can be seen in Equation 15. arg max p1,p2 ∑ i=p1,p2 ∂Ft(X) ∂Xi!× ∑ i=p1,p2 ∑ j6=t ∂Fj(X) ∂Xi (15) For example, if pixel p1 has a target derivative of 5 but a sum of other classes derivatives equal to 0.1, while p2 as a target derivative equal to −0.5 and a sum of the other classes derivatives −6. Alone neither pixel won’t match the search criteria in Equation 13, but combined the pair matches the saliency criteria defined in Equation 15. Using pairs improves the chance of obeying the selection criteria more often, although the method becomes computationally more expensive. The full attack algorithm can be seen in Algorithm 2, showing 3 different stop conditions: 1. the adversarial sample is classified by the network as being the target class t 2. the max number of iterations has been reached 3. the feature domain Γis empty, Γrefers to the number of input features valid for perturbation, all the pixels that do not have the maximum value in the color range (ex: 1 or 255) if we are using the incremental saliency map S+, or the pixels that do not have the minimum value in the range( 0 ) if we are using the decremental saliency map S−. The number of maximum iterations max_iter is defined by max_iter =h# input features·γ 2i, where γis the percentage of the total number of input features allowed to be perturbed. 2.1. Adversarials Attacks 15 Besides γ, the algorithm has another parameter, θ, that defines the amount of perturbation applied to each pixel. Algorithm 2:JSMA full algorithm Input: input X,Y∗is the target classifier output, Fis the classifier, γis the percentage of the total number of input features allowed to be perturbed, θis the amount of perturbation applied to pixels, nthe number of features of X Output: Adversarial sample X∗ 1X∗←X 2Γ={1...|X|} // search domain is all pixels 3max_iter =n·γ 2·100 4s=arg max j F(X∗)j// source class 5t=arg max j Y∗ j// target class 6while s6=t & iter <max_iter & Γ6=∅do 7Compute derivative ∇F(X∗) 8p1,p2=Saliency_Map(∇F(X∗),Γ,Y∗) 9Modify p1and p2in X∗by θ 10 Remove p1from Γif p1== 0 for S−or p1== 1 for S+ 11 Remove p2from Γif p2== 0 for S−or p2== 1 for S+ 12 s=arg max j F(X∗)j 13 iter + + 14 return X∗ Non-Targeted JSMA (NT-JSMA) In Wiyatno and Xu (2018) proposed a variant of JSMA where there is no need to specify a target class to perform the attack. This algorithm instead of increasing the probability of some target class, it decreases the model’s prediction confidence of the true class label. This is done by swapping the saliency maps employed in JSMA, so if we are increasing the value of the pixels we use the S−map, and if we are decreasing the values we use the S+ instead. This variant has a more relaxed success criterion in comparison with JSMA because here we only need to ensure that the sample is misclassified, regardless of the final class. 2.1. Adversarials Attacks 16 2.1.6Carlini Method Figure 6.: Carlini L2adversary applied to the MNIST dataset performing a targeted attack for every source/target pair. Source: Carlini and Wagner (2017) In Carlini and Wagner (2017) the authors propose a new attack methodology. The main idea of this attack is to implement all the constraints required to craft an adversarial sample as an objective function that can be solved by using a gradient descent optimizer such as Adam. To understand how this is done first we need to formulate the constraints necessary to create an adversarial attack. This can be seen in Equation 16, where Xis the input image, δis the perturbation, Da distance function, Cthe classifier model, nare the dimensions of the input image, in case of a RGB image n=3 and tis the target class. minimize: D(X,X+δ) such that: C(X+δ) = tconstraint 1 X+δ∈[0, 1]nconstraint 2 (16) In the first line we want to ensure that the adversarial sample it’s close to original sample by using some distance metric, typically the distance metrics are specified in terms of Lp norms such as L0,L2or L∞. The second line shows a constraint that ensures that the adversarial sample is indeed misclassified by the classifier, this constraint is highly nonlinear Carlini and Wagner (2017). And finally in the last line we have the constraint 2 where we want to ensure that the crafted adversarial is a valid image. 2.1. Adversarials Attacks 17 To solve this optimization problem the authors suggest to replace C(X+δ) = tfor a function f, where this objective function fmust satisfy C(X+δ) = tif and only if the output of this objective function on the perturbed image X+δis less than or equal to zero, i.e. f(X+δ)≤0. The authors evaluated different candidates for the function fin original paper and found that best suited to be Equation 17. f(X) = max(max i6=tZ(X)i−Z(X)t,−k)(17) The term maxi6=tZ(X)irepresents what is the highest logits value among non-targeted classes, Z(X)tis the logits of the target class. Thus maxi6=tZ(X)i−Z(X)tis the difference between ’what the classifier thinks the current image probably is’ and ’what we want to misclassify as’. The purpose of the parameter kis to control the strength of the adversarial sample, adjust the trade-off between the size of the perturbation and the success of the attack. Larger values for kmay allow for more significant perturbations but could increase the risk of detection, on the other hand smaller values may result in smaller perturbations but may also reduce the likelihood of successful adversarial attacks. The optimal value for "k" may depend on the particular dataset, model architecture, or other factors. Constraint 2 from Equation 16 with lower and upper bounds set to 0 and 1 respectively, cannot be optimized using gradient descent because it is not a differentiable function. To solve this it is necessary to convert this constraint to a differentiable function with the same codomain. To achieve this purpose the authors suggest the use of the tanh function. The range of tanh is [−1, 1], hence we must add +1 and multiply by 1 2to force to be within the codomain required, see Equation 18. Note that Wis the image converted in tanh space, this is done by converting the image to the range [−1, 1]and then performing the tanh inverse function: arctanh. constraint: X+δ=1 2(tanh(W) + 1) δ=1 2(tanh(W) + 1)−X, (18) Now putting everything in an equation we get Equation 19, where kk2is the L2norm and cis a constant that steers the trade off between the left and right terms of the equation (image similarity vs misclassification confidence), cmust be greater than zero, having a small value for cresults in the attack rarely succeeding and having a large value of cresults in attacking being less effective (large value of L2distance) but always succeeding. The value of ccan be found empirically by performing binary search, we start with a very small 2.2. Defense methods against Adversarial Attacks 24 Using high temperature systematically reduces the model sensitivity to small variations of it’s inputs when defensive distillation is performed at training time. At test time, the temperature is set to T=1 in order to make predictions on unseen inputs. The intuition is that this doesn’t affect the model’s sensitivity as weights learned during training will not be changed by this change in temperature, and decreasing the temperature only makes the class probability vector more discrete, without changing the relative ordering of classes. In a way, the smaller sensitivity imposed by using a high temperature is encoded in the weights during training and is thus still observed at test time. 3 ATTACK EVALUATION This chapter presents an evaluation on the transferability of an attack generated by a particular model when tested on a set of different models. By transferability it is meant if an adversarial sample crafted on a model is also misclassified on other models. When considering targeted attacks we only consider an attack as successful if the models classify the samples as the target class. For each attack method, a model is used to generate a set of adversarial samples that will be used to evaluate the transferability to the other models. The chapter starts by describing the experimental setup, including the training procedures and attacks performed. The datasets used for these tests are introduced in Section 3.1, followed by the architectures and model training procedure in Section 3.2. The procedure to build the adversarial datasets is presented in Section 3.3. Section 3.4and Section 3.5present the results considering different attack methods for each dataset. 3.1 datasets For this work two datasets were considered: CIFAR-10 and GTSRB. The first is a widely used non-trivial dataset, while the second is relevant in real world applications such as autonomous driving. They differ significantly in their intra-class and inter-class variability, with GTSRB having a smaller variability in both items. 3.1.1CIFAR-10 CIFAR-10 (Krizhevsky (2009)) is a labeled subset of the Tiny Images dataset. The Tiny Images dataset had 80 million 32 ×32 images and was originally created in 2006 in MIT, however, it was withdrawn in June 2020 (Torralba et al. (2020)). This dataset was chosen due to the fact that it is widely known and reported. 25 3.1. Datasets 26 The CIFAR-10 dataset consists in 60000 32 ×32 colour images in 10 different classes, with 6000 images per class. These were labelled by Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton. The test set has 10000 images. From the remaining 50000 images a split for training/validation was performed at 80%/20%. Images in this dataset consists of natural shapes such as animals and some more geometrical ones like vehicles. Some samples can be seen in Figure 9. Figure 9.: CIFAR-10 samples from https://www.cs.toronto.edu/ kriz/cifar.html 3.1.2German Traffic Sign Recognition Benchmark GTSRB (GTSRB (2010)) is provided by the Real-Time Computer Vision group from the Institut für Neuroinformatik, Ruhr-Universität Bochum. It was initially made available for a competition at IJCNN 2011. This dataset was selected due to its relevance for autonomous driving. This is an image classification dataset based in traffic sign images, originally divided in a training set with 39209 images, and a test set with 12630 images, totalling 51839 samples divided by 43 classes. As opposed to CIFAR-10 this is not a balanced set, with the number of images per class ranging from 210 to 2250 in the training set. Each image is in RGB with dimension ranging from 25 ×25 to 232 ×266 pixels. This dataset is composed of the sequence of frames taken from 1 second videos at 30 fps. This implies that the validation set 3.2. Models and training procedure 27 must be built using full sequences, instead of randomly picking images from the original training set. The validation set was built with approximately 20% of the sequences for each class. Some samples can be seen in Figure 10. Figure 10.: GTSRB samples from http://benchmark.ini.rub.de 3.2 models and training procedure This project requires several architectures to be used to evaluate the transferability of adversarial attacks under different conditions. Well known architectures were selected, such as VGG16, ResNet (with three variations), together with a shallower conventional CNN architecture that we called ConvNet. For each model, 5 training runs were performed on each dataset. All models use pre-processing to feed the network subtracting from the image the training set mean for each channel. Training was performed with data augmentation with Keras ImageDataGenerator. Specific details on the training procedure for each dataset are provided in the respective sections. 3.2.1VGG16 VGG16 is a deep conventional CNN with 13 convolutional layers and 3 fully connected layers. This architecture won the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) competition in 2014. Our implementation of this architecture is based on Geifman (2018). 3.3. Building the adversarial sample set 28 3.2.2ResNets ResNet architecture was made famous when it won the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) competition in 2015. This architecture introduces the concept of short cut connections allowing the definition of models with a large number of layers without the vanishing gradients problem. All ResNets used in this work are based on a narrow version described in He et al. (2015), following the formula 6n+2 for the number of layers. We chose three architectures varying the number of layers: ResNet-20, 20 layers with n=3, ResNet-50, 50 layers with n=8, and ResNet-110, 110 layers with n=18. Our implementation is based on Chollet. 3.2.3ConvNet This is a shallower architecture built for this project that consists in only 3convolutional layers followed by 2dense layers. a diagram of this model and their parameters can be seen in Figure 11. 3.3 building the adversarial sample set An adversarial sample is an image that has been perturbed so that it causes a misclassification on a particular model. The level of perturbation present in an adversarial sample deeply influences the outcome, and some metric is required to evaluate this perturbation. Furthermore, different methods have different parameters, therefore, some measure of image perturbation is required in order to fairly compare the methods. We have adopted SSIM Wang et al. (2004), a well known and studied metric used to measure image degradation. The implementation used is provided by scikit-image library, in which the SSIM index for a pair of images is a value between −1.0 and 1.0, where 1.0 is only attainable with identical images. A value of −1.0 indicates that there are no structural similarities between the two images. To generate adversarial samples we considered two indices of perturbation: 0.95 and 0.80. The first index is achievable for pairs of images where there is almost no perceptual difference between them. The second index already implies some perceptual degradation, yet the items in the perturbed image are still easily recognisable. In this work, the SSIM value is computed between the original image and the perturbed, adversarial sample. The goal is therefore to create two sets of adversarial samples, each having an average SSIM value close to those specified above. This is done for all the methods discussed in Chapter 2, namely: •FGSM (Section 2.1.3) 3.3. Building the adversarial sample set 29 Figure 11.: ConvNet architecture diagram • Deepfool (Section 2.1.4) •JSMA (Section 2.1.5) • Carlini (Section 2.1.6) •PGD (Section 2.1.7) • One pixel/ Few Pixels (Section 2.1.8) 3.3. Building the adversarial sample set 30 The adversarial sample sets are defined by the pair (SSIM, method). For each pair we had to perform a parameter search for the particular method to achieve a dataset with the specified average SSIM value. A trained ResNet-50 was used to craft all sets of adversarial samples. As mentioned before, five models were trained for each architecture. We selected the ResNet-50 with median accuracy to be the adversarial sample generator. These datasets will then be used to evaluate the transferability ratio to other models. The selection of ResNet-50 allows to test in similar architectures both shallower (ResNet-20) and deeper (ResNet-110), as well as significantly different architectures (VGG16 and ConvNet). When attacking the other ResNet-50 models, we are doing an almost white box attack: only the weights differ from the generator model. Considering the other trained models, the attack is closer to a black box attack since only the training set is known. To build the adversarial sample sets, Nimages from the respective test set are selected such that all trained models correctly classify these images. The number of selected images and the selection procedure varies by dataset and is detailed in Section 3.4and Section 3.5. In this project, all models expect a preprocessed input by subtracting the training dataset mean from the image pixels. This shifts the input values to be centered around zero, meaning that the model input values are in the range [−1, 1]. When crafting an adversarial samples these are the max boundaries allowed. When the misclassification is attained, the train mean is added back to the adversarial sample, and the values clipped to the range [0, 1]to generate a valid RGB image. This remapping from different domains can cause an adversarial sample to lose its adversarial property, i.e, although the adversarial sample with values in [−1, 1]causes a misclassification, the RGB image obtained after summing the training set mean and clipping can still be correctly classified, or in the case of a targeted attack, it may not classify the image in the desired class. As mentioned before, for each adversarial sample configuration (SSIM, method), Nadversarial samples per class are attempted. Nis an upper bound on the final number of adversarial samples in each set. A set may have fewer samples due to two reasons: the method to generate the adversarial sample may fail, i.e. the method fails to converge; or the preprocessing issue mentioned above caused the sample to lose its adversarial effect. In practice, to generate the adversarial sample set we are carrying white box attack, and keeping the samples that result in a successful attack. The main goals of this work is to assert the transferability of the studied adversarial attacks methods between models with different architectures, or shared architecture but different weight states. 3.4. CIFAR-10: Adversarial Attacks 31 3.4 cifar-10:adversarial attacks This section details all the experiments performed on models trained with CIFAR-10. It starts by describing the specific procedure details for this dataset in Section 3.4.1.Section 3.4.2to Section 3.4.7detail the initial parameter search for each of the attacks involved, considering both SSIM values, where the attack is performed in a single model for each architecture. Finally, Section 3.4.8presents the aggregate test results, using all models for each architecture. 3.4.1Procedure details and attack parameters Tests on the CIFAR-10 dataset use all the trained models on this dataset: VGG16, ConvNet, ResNet-20, ResNet-50, and ResNet-110. The architectures of these models is described in Section 3.2. Table 1shows the average accuracy obtained in the test set for each architecture. VGG16 ConvNet ResNet-20 ResNet-50 ResNet-110 CIFAR-10 92.80 ±0.17 86.89 ±0.33 90.71 ±0.09 91.52 ±0.13 91.39 ±0.25 # train params 14991946 2725360 273066 760266 1734666 Table 1.: Average accuracy and standard deviation on test set for all models trained on CIFAR-10 dataset and their corresponding number of trainable parameters. Dynamic data augmentation was performed with settings as presented in Table 2. ImageDataGenerator width shift range 0.1 height shift range 0.1 fill mode constant horizontal flip True Table 2.: CIFAR-10 data augmentation All the attack methods described from Section 2.1.3to Section 2.1.8are used with this dataset, namely: FGSM, Deepfool, JSMA, Carlini, PGD, and One pixel/Few Pixels. The implementation for FGSM, Deepfool, JSMA (non-targeted), Carlini, and PGD are from the Adversarial-Robustness-Toolkit (ART) library by IBM Nicolae et al. (2018). For the nontargeted version of JSMA we used an implementation based on Youlixx (2020), and for Few Pixels attack we used the implementation of Song (2019) based on Su et al. (2019). To build a dataset for each pair (SSIM, method), first it is required to find the parameters for each method that generate an average level of distortion close to the selected SSIM indices. This implies exploring the parameter space of each attack method. 3.4. CIFAR-10: Adversarial Attacks 32 As mentioned before, the adversarial samples are created from a ResNet-50 model. Initially, 500 samples correctly classified in all trained models are randomly selected from CIFAR-10 test set. Note that, as mentioned in Section 3.3, this number is an upper bound on the cardinality of the datasets generated. Regarding the targeted version, on CIFAR-10 the attack uses as the target class the true class +1(mod 10), for example if the true class is 8 which contains horses the new target will be class 9, containing ships, if the true class is the last, in this case class 9 which contains truck, the new target will be 0, the airplane class. Starting from Section 3.4.2to Section 3.4.7the parameter exploration for each method is presented. In addition, these sections also present an evaluation for the trained models for each architecture. As mentioned before, to evaluate the best performing norm, initially we use a single model of each architecture (the chosen model was the one with median test accuracy across all 5 instances of the same architecture, in the case of the ResNet-50 we picked the second model with best test accuracy out of 4) and calculate the average transferability in these models on the adversarial samples. The norm that attains the highest average between the models will be chosen to report the final results. Section 3.4.8presents a more comprehensive test with all models. The remaining of this subsection presents the training details for each architecture. VGG16 The VGG16 was trained using the same protocol used in Geifman (2018), the parameters can be seen in Table 3. This training protocol has two callbacks functions: • ModelCheckpoint - saves current epoch model if the validation accuracy is better than the best saved model so far; • LearningRateScheduler - changes the learning rate along the epochs. The learning rate schedule in derived from the formula in Equation 24, where · ·denotes the integer division. The graph of this schedule can be seen in Figure 12. new_lr =learning_rate ·0.5jepoch lr_drop k(24) ResNet based models All the ResNet base models follow the same training protocol, which was based on Chollet. The training parameters can be seen in Table 4. This training protocol has the same callbacks functions as VGG-16 (Table 3.4.1) with a different equation for the LearningRateScheduler, as seen in Equation 25. A graph of this schedule can be seen in Figure 13. 3.4. CIFAR-10: Adversarial Attacks 33 VGG16 train params epochs 250 batch size 128 optimizer SGD learning rate 0.1 momentum 0.9 learning rate decay 0.000001 learning rate drop 20 nesterov True Table 3.: VGG16 model training parameters on CIFAR-10 Figure 12.: VGG16 learning rate schedule for CIFAR-10 ResNet train params epochs 200 batch size 32 optimizer Adam learning rate 0.001 Table 4.: ResNet models training parameters on CIFAR-10 new_lr(epoch) =                      0.0005 ·0.001 if epoch >180 0.001 ·0.001 if 160 <epoch <180 0.01 ·0.001 if 120 <epoch <160 0.1 ·0.001 if 80 <epoch <120 0.001 if epoch <80 (25) 3.4. CIFAR-10: Adversarial Attacks 40 • A similar behaviour can be seen with VGG16, but in this case the decrease in relative robustness starts earlier. In Figure 20 we can see this behaviour for targeted FGSM L2norm. This behaviour can also be observed in the other norms on the target setting. These graphs are available in Section A.1.1. Table 8shows the results of success targeted attack rate per SSIM and epsilon for each norm. As in the non-targeted attacks, the norm that achieved best average transferability for SSIM =0.80 was L∞and for SSIM =0.95 was L2. FGSM L∞L1L2 SSIM (±0.005)0.95 0.80 0.95 0.80 0.95 0.80 epsilon 0.026 0.070 42.0 130.0 1.23.6 avg transf atk 13.39 31.11 16.00 30.11 16.21 30.00 succ atk on generator 25.40 14.60 23.40 19.00 23.60 18.40 succ atk aft conv 100.00 98.63 98.29 91.58 98.31 93.48 total atks aft conv 25.40 14.40 23.00 17.40 23.20 17.20 Table 8.: targeted FGSM average, SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm 3.4.3Deepfool This attack, as opposed to FGSM (Section 2.1.3), has no norms. Furthermore, this attack is purely a non-targeted attack. Hence, we only searched the parameters that gave us average SSIM of 0.80 and 0.95. Figure 21 shows an adversarial sample for a range of evalues. Figure 21.: Deepfool image samples per epsilon Figure 22 shows the success attack rate on each model per e, their corresponding average SSIM and total percentage of successful adversarial samples on model generator after converting to valid images. For SSIM index 0.95 the epsilon found was e=5 attaining an average success attack rate of 30.75%. For SSIM index 0.80 the epsilon found was e=18 attaining an average success attack rate of 67.98% 3.4. CIFAR-10: Adversarial Attacks 41 Figure 22.: Deepfool success attack rate on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e An interesting finding, which differs from what was observed with FGSM, is that the percentage of successful adversarial attacks on generator model seems to decrease with the varying e. This behaviour seems to indicate that the algorithm increasingly overshoots the boundaries as egets larger. Therefore choosing eis important not only due to the amount of degradation it creates in the original image, but also due to its impact on the success rate of adversarial generation. The robustness relative ordering for the architectures is similar to the results obtained in non-targeted FGSM. Table 9presents the information obtained for both SSIM values. Deepfool SSIM (±0.005)0.95 0.80 epsilon 5.0 18.0 avg transf atk 30.75 67.98 succ atk on generator 96.20 87.40 succ atk aft conv 99.79 98.63 total atks aft conv 96.00 86.20 Table 9.: Deepfool average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks 3.4. CIFAR-10: Adversarial Attacks 42 3.4.4JSMA This attack does not have norms defined, and has two parameters: •θ: the amount of perturbation added to the pixel; •γ: the maximum amount of pixels as number or percentage in the image allowed to be perturbed. Non-Targeted For the non-targeted JSMA (NT-JSMA) algorithm we used an implementation based on Youlixx (2020). In Figures 23 and 24 we can see adversarial samples for different thetas and gammas. Figure 23.: NT-JSMA samples for different gammas and constant theta=0.60 Figure 24.: NT-JSMA samples for different thetas and constant gamma=20% This attack is very expensive computational-wise so instead of generating 500 samples for each parameter we only used 250. For the parameter search we varied θfrom 0.20 to 1.0 in 0.2 steps and varied γfrom 20% to 80%. For this particular dataset and ranges, the γvalue did not have any influence on the adversarial generation, i.e., the adversarials are the same regardless of γ.Hence, we fixed γ=20%. This behaviour can be seen in Figure 23. Figure 25 shows the success attack rate on each model per θwhile setting γ=20%, their corresponding SSIM and total percentage of successful adversarial samples on model generator after converting to valid images. The behaviour observed in Figure 25 shows some relevant features: 3.4. CIFAR-10: Adversarial Attacks 43 Figure 25.: NT-JSMA success attack rate on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per θ and γ=20% • Increasing θappears to have little influence in decreasing the SSIM between the original image and its adversarial counterpart. Together with what was observed for γ this corroborates the authors statement that JSMA crafts adversarials perturbing only a small number of pixels, see Section 2.1.5; • The percentage of successful adversarial attacks after conversion on the generator decreases significantly as θincreases. As can be seen in Table 10, for the selected SSIM values, all attacks before conversion have 100% success rate. This implies that JSMA tends to create adversarial samples outside the boundaries of valid images, that once converted loose their adversarial status. Note that there is no contradiction with the previous point. As seen in Figure 24 the perturbation is indeed small considering the number of pixels that are affected, however, this perturbation is very significant for the modified pixels; 3.4. CIFAR-10: Adversarial Attacks 44 • as the perturbation increases, the shallower networks have higher transferability rates: ConvNet, ResNet-20, and VGG16 show the highest transferability rate when θ=1.0. ResNet-110 is the most robust architecture in this test. As mentioned before, prior to the conversion to a valid range, all attacks had 100% success rate. Hence, we evaluated the transferability on the unconverted adversarial samples. Results are presented in Figure 26. In this graph we can see that the deepest model, ResNet110, remains the most robust. On the other hand, ConvNet is now the second most robust architecture when θ=1.0. This shows that converting to a valid image affects different architectures in different scales. Figure 26.: NT-JSMA success attack rate on each model and their corresponding SSIM and total successful adversarial attacks on model generator without converting to a valid image per θand γ=20% For SSIM index 0.95 the parameters found were: γ=20% and θ=0.20 with an average transferability success attack rate of 7.25%; we did not achieve a SSIM=0.80 with this algorithm with neither combination of parameters searched because the threshold needed to misclassify all the 250 samples tested were small enough that their average perturbation did not decreased below 0.8729. Although the targeted SSIM was not achieved, we used the 3.4. CIFAR-10: Adversarial Attacks 45 combination γ=0.2 and θ=1.0 as the highest perturbation possible in these conditions. The parameters chosen for the small and large perturbation can be seen in Table 10. NTJSMA SSIM(±0.005)0.95 0.8729 gamma 0.2 0.2 theta 0.2 1.0 avg transf atk 7.25 39.13 succ atk on generator 100.00 100.00 succ atk aft conv 77.20 27.60 total atks aft conv 77.20 27.60 Table 10.: NT-JSMA average SSIM, theta, gamma, success attack rate, success attacks after conversion and total successful attacks Targeted Figure 27 shows the success attack rate on each model for each θwhile setting γ=20%, their corresponding SSIM and total percentage of successful adversarial samples on model generator after converting to valid images. As expected there is a significant increase in robustness when considering targeted attacks. For the targeted attack the parameters chosen can be seen in Table 11. Target JSMA SSIM(±0.005)0.95 0.8729 gamma 0.2 0.2 theta 0.13 0.78 avg transf atk 0.98 9.00 succ atk on generator 100.00 100.00 succ atk aft conv 73.20 24.00 total atks aft conv 73.20 24.00 Table 11.: Targeted JSMA average SSIM, theta, gamma, success attack rate, success attacks after conversion and total successful attacks 3.4.5Carlini Carlini exist in both targeted and non-targeted versions. As these attacks are expensive computation-wise, and both L∞and L0are derived from the L2attack, which was the main attack devised by the authors, we only perform tests using the L2norm. 3.4. CIFAR-10: Adversarial Attacks 46 Figure 27.: Targeted JSMA success attack rate on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per θ with γ=20% In these tests we searched over k, used 1 for the starting value of c, and the other variables used the default values (learning rate =0.01, binary search steps=10, max iterations =10, max halving =5 and max doubling =5). Non-Targeted In Figure 28 we can see adversarial samples for different kvalues. Figure 28.: Non-targeted Carlini L2adversarial sample for some confidence values k 3.4. CIFAR-10: Adversarial Attacks 47 Figure 29 shows the success attack rate on each model per confidence k, their corresponding average SSIM and total percentage of successful adversarial samples on model generator after converting to valid images. Figure 29.: Non-targeted Carlini L2norm success attack rate of on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per k For SSIM=0.95 the kfound was 57 attaining an average transferability attack success rate of 63.80% and for SSIM=0.80 the kis 78 achieving an average transferability attack success rate of 87.00%, see Table 12. Carlini L2 SSIM(±0.005)0.95 0.80 k 57 78 c 1 1 avg transf atk 63.80 87.00 succ atk on generator 100.00 100.00 succ atk aft conv 100.00 100.00 total atks aft conv 100.00 100.00 Table 12.: Non-targeted Carlini L2norm average SSIM, k, c, success attack rate, success attacks after conversion and total successful attacks 3.4. CIFAR-10: Adversarial Attacks 48 The results in Figure 29 show a similar behaviour for all the ResNet model group. The ConvNet is by a relevant margin the most robust model regarding this attack, with VGG16 having a robustness in between the other models. Targeted For the targeted attack we followed the same steps of non-targeted attack, for this experiment the results are: for SSIM=0.95 the kfound was 30 attaining an average transferability targeted attack success rate of 17.80%, and for SSIM=0.80 the kfound was 49 achieving an average transferability attack success rate of 52.21%, see Table 13. Target Carlini L2 SSIM(±0.005)0.95 0.8729 k 30 49 c 1 1 avg transf atk 17.80 52.21 succ atk on generator 100.00 95.00 succ atk aft conv 100.00 100.00 total atks aft conv 100.00 95.00 Table 13.: Targeted Carlini L2norm average SSIM, k, c, success attack rate, success attacks after conversion and total successful attacks Figure 91, in Section A.1.3, shows the success attack rate on each model per confidence k. The behaviour is similar to the non-targeted attack, although with a lower overall transferability rate as expected. 3.4.6PGD PGD has a single parameter e. In this test we allowed this algorithm to run at most 100 iterations, and for each etested we defined the step size per iteration as e 100. All three norms (L∞,L1,L2) have been tested in both targeted and non-targeted versions. Non-targeted L∞norm In Figure 30 we can see adversarial samples per e. Figure 31 shows the success attack rate on each model per e, their corresponding average SSIM and total percentage of successful adversarial samples on model generator after converting to valid images. For SSIM index 0.95 the epsilon found was e=0.08 attaining an average transferability success attack rate of 71.28%, and for SSIM index 0.80 the epsilon found was e=0.45 attaining an average transferability success attack rate of 88.60%. 3.4. CIFAR-10: Adversarial Attacks 49 Figure 30.: Non-targeted PGD L∞norm adversarial sample per e Figure 31.: Success attack rate of non-targeted PGD L∞norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e Non-targeted L1norm In Figure 32 we can see adversarial samples per epsilon. Figure 33 shows the success attack rate on each model per e, their corresponding average SSIM and total percentage of successful adversarial samples on model generator after converting to valid images. For SSIM index 0.95 the epsilon found was e=180 attaining an average transferability success attack rate of 73.79%, and for SSIM index 0.80 the epsilon found was e=840 attaining an average transferability success attack rate of 89.71%. 3.4. CIFAR-10: Adversarial Attacks 56 Figure 39.: Sample for non-targeted attack with One Pixel method within [0.925, 0.975]and [0.75, 0.85]SSIMs for S and L settings respectively Based in Figure 38 and Figure 39 it is clear that both JSMA and Few Pixels produce clearly noticeable artifacts in the image, in the form of highly salient coloured pixels. The disturbance in the remaining methods resembles the existence of noise, not being so easily discernible for a human if they are in fact the result of adversarial attacks. In Figure 40 we see the different attacks and the their transferability rate between models. From the 1000 samples generated for each attack setting we picked only the samples that were misclassified by the generator model (ResNet50 Generator) when converted to the range [0, 1]. These samples were then predicted by the other models to calculate the transferability rate. The percentage of misclassified samples per attack can be seen in Table 17. 3.4. CIFAR-10: Adversarial Attacks 57 Figure 40.: Non-targeted attacks and corresponding transferability transfer rate. nt atk VGG16 ConvNet R20 R50 R110 TA I% C% F% FGSM L2S28.87 15.51 52.99 52.55 51.61 40.31 54.00 99.81 53.90 FGSM L∞L70.51 37.79 82.89 82.31 83.69 71.44 83.20 99.04 82.40 PGD L1S64.07 16.35 94.54 95.83 95.14 73.18 95.90 100.00 95.90 PGD L1L96.58 54.27 99.43 99.62 99.59 89.90 97.80 100.00 97.80 Deepfool S 16.33 6.16 42.09 42.01 40.81 29.48 95.90 99.90 95.80 Deepfool L 62.93 42.27 77.21 76.00 77.07 67.09 88.20 99.09 87.40 Carlini L2S37.70 8.88 85.02 85.78 85.30 60.54 100.00 100.00 100.00 Carlini L2L86.65 43.50 98.12 98.17 98.60 85.01 99.90 100.00 99.90 JSMA S 3.64 4.08 9.95 10.19 8.10 7.19 100.00 77.00 77.00 JSMA L 38.72 44.51 39.57 37.66 32.77 38.65 100.00 23.50 23.50 Few pixels S 9.04 7.43 14.77 12.87 10.84 10.99 84.50 100.00 84.50 Few pixels L 30.81 20.89 42.99 39.52 38.11 34.46 97.10 100.00 97.10 Table 17.: CIFAR-10 non-targeted transferability table. I% is the initial attacks successful on generator, C% the percentage of attacks successful on generator after conversion to valid image, and F% final percentage of successful attacks on generator that were used to attack the other networks, this is a pure white box attack. Column TA is the transferability average of an attack across all networks. In bold are the highest values per attack, in green and blue cells are the highest transferability values of all attacks per model architecture for the L and S setting, respectively From Figure 40 and Table 17 we can observe the following regarding the model architectures: 3.4. CIFAR-10: Adversarial Attacks 58 • Overall, the most robust architecture is ConvNet, possibly due to being the shallowest model, followed by VGG16, with the ResNet family being the most susceptible to the transferability of adversarial samples. This hints that architectural similarities may be relevant when considering the transferability of adversarial attacks. • The only exception to the first point is JSMA L (large perturbation), where all architectures show a more similar transferability rate. However, note that the transferability of JSMA is small compared to most other methods; • Within the ResNet family it is interesting to note that, although the adversarial samples were crafted in a ResNet-50, these models are not the least robust in some settings; • It is worth noting that, while Few Pixels does not use any architectural knowledge, it follows the general transferability trend regarding architecture similarity; • The attacks with best transferability overall are PGD and Carlini. On the other hand JSMA and Few Pixels are the less capable methods regarding transferability. Considering the two worst methods regarding transferability JSMA and Few Pixels, it is noticeable that both adopted the same approach of affecting only a small number of pixels with a high perturbation. All other methods produce a more diffuse perturbation. This may hint that the latter approach is preferable regarding transferability. Note that, as mentioned before, the SSIM values attained using these parameters are an average over the datasets. This implies that some samples have indexes values above or below that target values. This effect can be seen in Figure 41 presenting some Carlini L2 samples crafted using the same parameters and their distortion level varies significantly. This variability is directly related to the uncertainty of the level of perturbation obtained with an attack considering a fixed set of parameters. The higher the variability the less certain we can be of the amount of degradation present in an image. In Table 18 we present more information regarding the obtained SSIMs of each attack method for the adversarial examples crafted on the generator. We can see that the attacks with highest SSIM variability among samples are Deepfool and Carlini L2for L setting. On the other hand PGD is the more consistent method regarding the level of perturbation generated for an adversarial sample. In order to evaluate the influence of the level of perturbation, we repeated the experiments picking only the adversarials that have SSIMs within [0.925, 0.975]and [0.75, 0.85] for the S setting and the L settings respectively. In Figure 42 we can see the results, the overall trends are the same as in Figure 40. The differences between these two experiments can be seen in Figure 43. The increase in transferability when using threshold selection is 3.4. CIFAR-10: Adversarial Attacks 59 Figure 41.: CIFAR-10 Carlini L2samples using the same parameters (k=78) below 10% for most methods, the only exception being JSMA L where we have a significant increase in transferability on VGG16 and ConvNet models, with the latter obtaining an increase over 20%. Nevertheless, this does not change that JSMA relative performance regarding transferability still remains weak. 3.4. CIFAR-10: Adversarial Attacks 60 non-targeted adv SSIMs min max avg std % within correct SSIMs FGSM L2S0.675 1.000 0.967 0.039 37.80 PGD L1S0.815 1.000 0.949 0.027 75.50 Deepfool S 0.706 1.000 0.950 0.046 41.60 Carlini L2S0.261 0.998 0.959 0.058 30.60 JSMA S 0.638 0.999 0.948 0.039 53.00 Few pixels S 0.774 0.997 0.944 0.028 70.40 FGSM L∞L0.287 0.946 0.794 0.091 41.20 PGD L1L0.414 1.000 0.807 0.078 51.10 Deepfool L 0.316 1.000 0.783 0.134 27.20 Carlini L2L0.014 1.000 0.804 0.177 18.20 JSMA L 0.447 0.995 0.877 0.076 24.00 Few pixels L 0.412 0.942 0.801 0.070 54.10 Table 18.: SSIM statistics of crafting examples on model generator for each CIFAR-10 non-targeted adversarials attacks using the parameters found in Section 3.4.2 Figure 42.: Non-targeted attacks and corresponding transferability rate with adversarials samples that are within the threshold interval 3.4. CIFAR-10: Adversarial Attacks 61 Figure 43.: Non-targeted attacks transferability rate difference between attacks within threshold and without restrictions Concluding we can state that PGD is not only the most effective method regarding transferability on CIFAR-10, but it is also the method that provides the most consistent degradation when creating the adversarial samples. Targeted transferability In this setting we used all attacks methods in targeted environment. This excludes Deepfool because it does not have target behavior. 3.4. CIFAR-10: Adversarial Attacks 62 Figure 44.: Sample for targeted attacks within [0.925, 0.975]and [0.75, 0.85]SSIMs for S and L settings respectively In Figure 44 we can see all the different targeted attacks considered, for each type of perturbation used: small (S) and large (L). All these samples were generated using the parameters and norms that attained best average misclassification accuracy, found in the 3.4. CIFAR-10: Adversarial Attacks 63 previous subsections. As in the non-targeted attacks, JSMA and Few Pixels methods also show clearly visible artifacts in the form of coloured pixels. For the remaining methods, the perturbations follow the same noisy pattern as in the non-targeted case. This is to be expected as the methods do not have a fundamental algorithm change between non-targeted and targeted versions. Figure 45.: Targeted attacks and corresponding transferability transfer rate. In Figure 45 we can see the different target attacks and their transferability rate between models, the main takeaways being: • Compared to the non-targeted attacks, the transferability is smaller as expected. Recall that we only consider an attack as successful when the model misclassifies the adversarial sample on the target class; • Once more, PGD gets the best overall transferability rate; • Regarding the architectures, ConvNet is the architecture that achieves the lowest transferability rate, with only once going above the 20% mark, the only exception being JSMA L where the ConvNet is the architecture with highest transferability rate by a small margin; • Again the ResNet family has the highest transferability rates, suggesting that a similarity in architecture can be a significant factor in attack transferability rate; 3.5. GTSRB: Adversarial attacks 64 •JSMA and few pixels attack remains the worst attacks regarding transferability. The percentage of samples classified as target class per attack can be seen in Table 19. target atks VGG16 ConvNet R20 R50 R110 TA I C F FGSM L2S13.24 9.51 21.07 23.11 20.62 17.51 23.00 97.83 22.50 FGSM L∞L 37.06 28.53 32.06 33.46 35.88 33.40 13.90 97.84 13.60 PGD L∞S17.26 2.44 57.58 59.65 59.42 39.27 100.00 100.00 100.00 PGD L∞L56.42 12.74 82.78 83.62 81.40 63.39 100.00 100.00 100.00 Carlini L2S4.14 0.54 24.30 27.40 24.82 16.24 100.00 100.00 100.00 Carlini L2L38.00 7.55 66.76 67.93 65.27 49.10 94.10 100.00 94.10 JSMA S 0.45 0.28 1.35 1.72 1.35 1.03 100.00 71.20 71.20 JSMA L 9.12 12.84 10.49 11.40 10.20 10.81 100.00 20.40 20.40 Few pixels S 3.06 3.01 5.53 4.85 4.42 4.17 41.20 100.00 41.20 Few pixels L 8.31 5.83 13.42 12.14 11.90 10.32 66.90 100.00 66.90 Table 19.: CIFAR-10 targeted transferability table. I% is the initial attacks successful on generator, C% the percentage of attacks successful on generator after conversion, and F% final percentage of successful attacks on generator that were used to attack the other networks, this is a pure white box attack. Column TA is the transferability average of an attack across all networks. In bold are the highest values per attack, and in green and blue cells are the highest transferability values of all attacks per model architecture, for the L and S setting, respectively FGSM obtains the best transferability rate for ConvNet in both Sand Lsettings, while being the least successful method in generating adversarial samples for the generator model. Apart from ConvNet result with FGSM, these results follow the trends found with the non-targeted attacks, with PGD dominating in all the other architectures, although with a lower transferability rate as expected. 3.5 gtsrb:adversarial attacks This section details all the experiments performed on models trained with GTSRB. It starts by describing the specific procedure details for this dataset in Section 3.5.1. From Section 3.5.2to Section 3.5.5the parameter search for each of the attacks involved are detailed. This search is performed using a single network for each architecture. Finally, Section 3.5.6 presents the final aggregate test results, using all models, for the two SSIM indices. 3.5. GTSRB: Adversarial attacks 65 3.5.1Procedural details and attack parameters The overall procedural details for GTSRB are similar to those presented in 3.4.1for the CIFAR-10 dataset. In this section we only present the details which are specific for the traffic sign dataset. For this dataset we used the following architectures: VGG16, ConvNet, ResNet-20 and ResNet-50. For each architecture we trained 5 models. Table 20 presents the average accuracies for each architecture. VGG16 ConvNet ResNet-20 ResNet-50 GTSRB 98.79 ±0.20 98.81 ±0.08 98.94 ±0.32 98.86 ±0.31 # train params 15008875 2736943 275211 762411 Table 20.: Average accuracy and standard deviation on test set for all models trained on GTSRB dataset and their corresponding number of trainable parameters. Each model has been trained using the same training parameters presented in Section 3.4.1 with one exception, the VGG16: instead of using a LearningRateScheduler callback it uses a ReduceLROnPlateau during training, the callback parameters can be seen in Table 21. VGG16 ReduceLROnPlateau callback monitor val loss factor 0.2 patience 5 min lr 0.0001 Table 21.: VGG16 model callback used for training on GTSRB dataset Considering that in the original training set each class contains multiple 30 images sequences depicting the same sign, we removed 20% of these 30 images sequences from each class to build a validation set. We applied the same preprocessing as done with CIFAR-10, i.e., we subtracted each image with the training set mean per channel. The parameters used for dynamic data augmentation to train the networks can be seen in Table 22. The attacks chosen for this test on GTSRB dataset are: FGSM, Deepfool, Carlini L∞and PGD; we excluded JSMA and Few Pixels attacks due to their poor performance on the CIFAR-10 experiment. To perform the attacks we randomly picked 430 images that were all correctly predicted by all the models to perform the attacks. We searched for the parameters using the same methodology used on CIFAR-10. For the target attacks we calculate the closest and farthest classes for each class using the SSIM metric. To determine these classes we calculate the average image for each class using the training dataset and then we calculate the SSIM against each other average image 3.5. GTSRB: Adversarial attacks 72 FGSM closest cls L∞L1L2 SSIM(+-0.005)0.95 0.80 0.95 0.80 0.95 0.80 epsilon 0.012 0.060 18 98 0.60 3.25 avg transf atks 12.90 10.19 23.80 20.94 22.29 21.00 succ atks on generator 14.65 26.51 19.53 28.60 19.30 29.53 succ atks aft conv 98.41 94.74 98.81 95.12 100.00 98.43 total atks aft conv 14.42 25.12 19.30 27.21 19.30 29.07 Table 25.: FGSM targeted attack for closest class average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm FGSM farthest cls L∞L1L2 SSIM(+-0.005)0.95 0.80 0.95 0.80 0.95 0.80 epsilon 0.013 0.031 17.0 47.0 0.55 1.45 avg transf atks 0.00 4.49 0.00 2.88 0.00 0.00 succ atks on generator 4.19 9.07 1.86 6.74 1.63 5.81 succ atks aft conv 100.00 100.00 100.00 89.66 100.00 92.00 total atks aft conv 4.19 9.07 1.86 6.05 1.63 5.35 Table 26.: FGSM targeted attack for farthest class average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm Figure 53.: Deepfool adversarial sample for some epsilon Figure 54 shows the success rate on each model per e, their corresponding average SSIM and total percentage of successful adversarial samples on model generator after converting to valid images. For SSIM=0.95 we found e=0.5 attaining an average success transferability rate of 13.14%. For SSIM =0.80 the value found was e=3.1 attaining an average success transferability rate of 51.77%, these results are shown in Table 27. The results in Figure 54 are similar to those obtained in CIFAR-10, see Section 3.4.3. Another interesting note is that there is no significant change in transferability between CIFAR-10 and GTSRB trained models, as opposed to what occurred when attacking with FGSM, see Section 3.5.2. 3.5. GTSRB: Adversarial attacks 73 Figure 54.: Success attack rate of Deepfool on each model per e Deepfool SSIM(±0.005) 0.95 0.80 epsilon 0.5 3.1 avg transf atk 13.14 51.77 succ atk on generator 100.00 98.60 succ atk aft conv 100.00 100.00 total atks aft conv 100.00 98.60 Table 27.: Deepfool average SSIM, epsilons and transferability success rate 3.5.4Carlini Similarly, to Section 3.4.5only L2was used in this test. The parameter to search is confidence k, the other values are kept as constant as in Section 3.4.5. Figure 55 shows adversarial sample for different confidence k. Figure 56 shows the success attack rate on each model per confidence k, their corresponding average SSIM and total percentage of successful adversarial samples on model generator after converting to valid images. For SSIM=0.95 the kfound was 35 attaining an average attack success rate of 32.36%, and for SSIM=0.80 the kis 57 achieving an average attack success rate of 64.34%, these results 3.5. GTSRB: Adversarial attacks 74 Figure 55.: Non-targeted Carlini L2adversarial sample for some confidence values k Figure 56.: Non-targeted Carlini L2norm success attack rate on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per k are shown in Table 28.Figure 56 shows the same trend seen in Figure 29 where ConvNet is the more robust model regarding adversarials transferability followed by VGG16. Targeted A similar parameters search was performed for the target attack closest and farthest classes. Table 29 shows the results: for the closest class target attack the kfound for the SSIM 0.95 and 0.80 were 7 and 41 with average transfer attack success rate of 6.20% and 51.47% respectably; for the farthest class the kfound for SSIM 0.95 and 0.80 were 3 and 15 attaining an average transfer attack success rate of 1.20% and 6.15% respectably. The graphics related to these searches are shown in Section A.2.3. These graphics shows the same trend 3.5. GTSRB: Adversarial attacks 75 Carlini L2 SSIM(±0.005)0.95 0.80 k 35 57 c 1 1 avg transf atk 32.36 64.34 succ atk on generator 100.00 100.00 succ atk aft conv 100.00 100.00 total atks aft conv 100.00 100.00 Table 28.: Non-targeted Carlini L2norm average SSIM, k, c, success attack rate, success attacks after conversion and total successful attacks where ConvNet is the more resistant architecture against transferability followed closely by VGG16. Carlini L2closest cls farthest cls SSIM (±0.005)0.95 0.80 0.95 0.80 k 7 41 3 15 c 1111 avg transf atks 6.20 51.47 1.20 6.15 succ atks on generator 100.00 26.36 100.00 94.57 succ atks aft conv 100.00 100.00 96.90 100.00 total atks aft conv 100.00 26.36 96.90 94.57 Table 29.: Carlini L2targeted for closest and farthest classes, average SSIM, k, c, success attack rate, success attacks after conversion and total successful attacks 3.5.5PGD As mentioned in Section 3.4.6PGD has a single parameter epsilon. In this test we allowed this algorithm to run for 100 iterations, and for each epsilon tested we defined the step size per iteration as epsilon 100 . All three norms (L∞,L1, and L2) have been tested in both targeted and non-targeted versions. Non-targeted L∞norm In Figure 57 we can see adversarial samples for selected epsilon values. Figure 58 shows the success attack rate on each model per e, their corresponding average SSIM and total percentage of successful adversarial samples on model generator after converting to valid images. For SSIM index 0.95 the epsilon found was e=0.03 attaining an average success attack rate of 23.96%, and for SSIM index 0.80 the epsilon found was e=0.13 attaining an average success attack rate of 61.34%. 3.5. GTSRB: Adversarial attacks 76 Figure 57.: Non-targeted PGD L∞norm adversarial sample for some epsilon Figure 58.: Success attack rate of non-targeted PGD L∞norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e Non-targeted L1norm In Figure 59 we can see adversarial samples for selected epsilon values. Figure 60 shows the success attack rate on each model per e, their corresponding average SSIM and total percentage of successful adversarial samples on model generator after converting to valid images. For SSIM index 0.95 the epsilon found was e=47 attaining an average success attack rate of 48.03%, and for SSIM index 0.80 the epsilon found was e=240 attaining an average success attack rate of 76.49%. 3.5. GTSRB: Adversarial attacks 77 Figure 59.: Non-targeted PGD L1norm adversarial sample for some epsilon Figure 60.: Success attack rate of non-targeted PGD L1norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e Non-targeted L2norm In Figure 61 we can see adversarial samples for selected epsilon values. Figure 62 shows the success attack rate on each model per e, their corresponding average SSIM and total percentage of successful adversarial samples on model generator after converting to valid images. For SSIM index 0.95 the epsilon found was e=1.5 attaining an average success attack rate of 51.77%, and for SSIM index 0.80 the epsilon found was e=8.0 attaining an average success attack rate of 78.39%. 3.5. GTSRB: Adversarial attacks 78 Figure 61.: Non-targeted PGD L2norm adversarial sample for some epsilon Figure 62.: Success attack rate of non-targeted PGD L2norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e In Table 30 are the results of success attack rate per SSIM for each norm. As we can see the norm that achieved highest average transferability for SSIM=0.80 and SSIM=0.95 was the L2norm, this norm will be used for the corresponding SSIM for the full scale attacks. Summary for non-targeted attacks All non-targeted attacks exhibit the same relative behavior as other non-targeted attacks referred before in this chapter where ResNet models family group have similar behavior and are more susceptible to transferability of adversarials while ConvNet remains the architecture more robust to these followed by VGG16. 3.5. GTSRB: Adversarial attacks 79 PGD L∞L1L2 SSIM (±0.005)0.95 0.80 0.95 0.80 0.95 0.80 epsilon 0.03 0.13 47 240 1.5 8.0 avg transf atk 23.96 61.34 48.03 76.49 51.77 78.39 succ atk on generator 90.93 100.00 74.19 86.05 49.30 63.49 succ atk aft conv 98.72 100.00 99.69 100.00 100.00 100.00 total atks aft conv 89.77 100.00 73.95 86.05 49.30 63.49 Table 30.: PGD non-targeted average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm Targeted attack A similar search for the parameters was performed for each norm considering the two types of target attacks: closest and farthest classes. Table 31 shows the results of successful target attack rate per SSIM and epsilon for each norm for the closest class. The norm that achieved highest average transferability for SSIM=0.80 was L1and SSIM=0.95 was L2. Similarly, Table 32 shows the results for the target attack on the farthest class, the norm that achieved highest average transferability for SSIM=0.80 and 0.95 was L1, these will be the norms and parameters used for the full scale test. The graphics related to theses targeted parameters search can be found in Section A.2.4. All these targeted attack follow the same trend as non-targeted setting. PGD closest cls L∞L1L2 SSIM(+-0.005)0.95 0.80 0.95 0.80 0.95 0.80 epsilon 0.04 0.23 180 2750 24 104 avg transf atks 11.04 26.57 21.05 46.05 24.77 45.64 succ atks on generator 97.44 100.00 100.00 100.00 100.00 100.00 succ atks aft conv 98.33 100.00 100.00 100.00 100.00 100.00 total atks aft conv 95.81 100.00 100.00 100.00 100.00 100.00 Table 31.: PGD targeted attack for closest class average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm PGD farthest cls L∞L1L2 SSIM(+-0.005)0.95 0.80 0.95 0.80 0.95 0.80 epsilon 0.04 0.24 80 2000 6 82 avg transf atks 0.39 2.67 1.14 9.09 0.93 8.95 succ atks on generator 90.23 100.00 97.67 99.77 100.00 100.00 succ atks aft conv 97.94 100.00 99.05 100.00 100.00 100.00 total atks aft conv 88.37 100.00 96.74 99.77 100.00 100.00 Table 32.: PGD targeted attack for farthest class average SSIM, epsilons, success attack rate, success attacks after conversion and total successful attacks for each norm 3.5. GTSRB: Adversarial attacks 80 3.5.6Adversarial attack comparison This section uses the results attained in the search for the best parameters values and norms which resulted in best transferability success rate. For each pair (attack method - SSIM) we generated 860 samples using these parameters. From the 860 (20 for each class) samples generated for each attack setting we picked only the samples that were misclassified by the generator model when converted to a valid image (in the range [0,1]). These samples were then used to attack the other models and calculate the transferability rate. Similarly, as in Section 3.4.8we use 5 different models for each architecture (VGG16, ConvNet, ResNet-20) except for the ResNet-50 where we will use 4 (the 5th model is the one used to generate the samples). The final result will be the average of these models per architecture. Non-targeted transferability Figure 63 and Figure 64 show a sample for all the different types of non-targeted attacks considering both perceptual changes. All these samples were generated using the parameters and norms that attained best average transferability found in the previous sections. Figure 63.: Sample for non-targeted attacks (FGSM and PGD) within [0.925, 0.975]and [0.75, 0.85] SSIMs for S and L settings respectively 3.5. GTSRB: Adversarial attacks 81 Figure 64.: Sample for non-targeted attacks (Deepfool and Carlini) within [0.925, 0.975]and [0.75, 0.85]SSIMs for S and L settings respectively In Figure 65 we see the different attacks and their transferability rate between models. The percentage of misclassified samples per attack is presented in Table 33 Figure 65.: Non-targeted attacks and corresponding transferability transfer rate. 3.6. Attack Analysis 88 • Models trained for GTSRB have a much higher accuracy than the CIFAR-10 models, which could help to prevent transferability. Note, however, that further study is required to validate any of these hypotheses. Regarding the model architectures the ResNet family were consistently the models with highest transferability rates by a significant margin. This hints that architecture similarity may have an impact in transferability. On the other hand the ConvNet was the architecture with lowest transferability rates in all the experiments by a significant margin, in some cases attaining 0%, this could be due to this architecture being the shallowest. The hypothesis that there is a relationship between shallowness and transferability rate, is consistent with the results, although, as is, it requires further study to check if there is a significant correlation. 4 DEFENSE EVALUATION In this chapter we will evaluate two types of defenses: Adversarial Training and Defensive Distillation. The test methodology is similar to the one used in Chapter 3. Both CIFAR10 and GTSRB datasets were used. The attack methods are the ones used in Section 3.5, namely: FGSM,PGD, Deepfool and Carlini. Initially, the settings for each of the defenses are defined for each defense (Section 4.1and Section 4.2). Afterwards, reports on the performance of the defenses are presented in Section 4.4for CIFAR-10 and Section 4.5for GTSRB. The chapter ends with concluding remarks based on the test results, presented in Section 4.6. 4.1 adversarial training One defense against adversarial attacks is Adversarial Training introduced in Goodfellow et al. (2015), see Section 2.2.1. Adversarial Training consists in training the model with adversarial samples added to the original the training set. This defense methodology aims to make the model insensitive to perturbations present in the adversarial samples. This allows the network to be more resistant to external adversarial attacks (attacks with different algorithms and/or magnitude of perturbation). Adversarial training can be realized in two ways: • Prior to training, generate adversarial samples for each sample of the training set, and during training mix these samples with the original training set; • Dynamically, in each training step, generate new adversarial samples from original training samples and mixed them with the original samples. The first version has a major drawback: the adversarial samples are constant from start to the end of training. On the other hand, in the second version adversarial samples are evolving alongside training. As the weights of the model change, the adversarial samples 89 4.2. Defensive Distillation 90 crafted by the model will change too. This is a more effective train method because the model is trained under different perturbations per training step, making the model more robust and insensitive to a higher range of perturbations. The downside of this version it is that is computational more expensive as it requires crafting new adversarial samples at each training step. For this experiment we opted to train the model using adversarial samples generated dynamically in each training step. Specifically for each training step there is a 20% probability of all images from the batch to be adversarial samples, otherwise all images from the batch are unchanged. We chose this split of 20% heuristically. We trained 5 models for each norm (L1,L2and L∞) with adversarial samples generated with non targeted PGD algorithm. Adversarial samples are generated with esampled for each batch randomly from a uniform distribution that lies between 0 and the corresponding norm evalue for SSIM =0.80 (this values were found in Section 3.4.6and Section 3.5.5). All parameters can be found in Table 36. PGD attack dataset CIFAR-10 GTSRB norm L∞L1L2L∞L1L2 eps [0, 0.45] [0, 840] [0, 23] [0, 0.13] [0, 240] [0, 8] eps per iter e/5 iterations 10 Table 36.: PGD parameters used in Adversarial Training We used the ResNet-50 architecture for the adversarial training. The training hyperparameters, data-augmentation and preprocessing are as described in Chapter 3. 4.2 defensive distillation Defensive Distillation (Section 2.2.2) is an alternative to Adversarial Training, being first proposed in Papernot et al. (2016). This method aims to smooth the decision boundary of the model so that a larger perturbation will be needed for an adversarial attack to succeed. Commonly, the aim of an adversarial attack is to find the smallest perturbation to make the model misclassify. By smoothing the decision boundary this minimal perturbation must increase to achieve the misclassification. This defense consists in initially training a model with hard labels (one hot encode categorical labels) using a softmax function with temperature Tas can be seen in Equation 26 F(X) = "ezi(X)/T ∑N−1 l=0ezl(X)/T#i∈0..N−1 (26) 4.3. Defenses experiments settings 91 For this experiment we used the ResNet-50 architecture, with temperature T=100 (Papernot et al. (2016), Carlini and Wagner (2017)). Five teacher networks were trained and the teacher with the median accuracy in the test set was picked to teach five students networks. This was done for CIFAR-10 and GTSRB datasets, the training procedures are identical to those described in Chapter 3. 4.3 defenses experiments settings For this experiment we trained 5 defensive distilled models and 5 adversarial trained models for each of the 3 norms (L∞,L1,L2) plus the 4 ResNet-50 already used in the Section 3.4.8 and Section 3.5.6. In total, for both CIFAR-10 and GTSRB tests, we used 24 models to be attacked, and used the same Resnet-50 generator used in Section 3.4.8and Section 3.5.6to craft the adversarial attacks. To analyze the effect of different norms on adversarial trained models and defensive distillation we attacked with all three norms (L1,L2and L∞) for FGSM and PGD plus all other attacks used in Section 3.5.6. Figure 71.: Histogram of CIFAR-10 samples used to craft adversarial attacks. 4.4. Adversarial Defenses experiments on CIFAR-10 92 Figure 72.: Histogram of GTSRB samples used to craft adversarial attacks. Following the same methodology in Section 3.4.8and Section 3.5.6we generated, across all classes, 1000 samples for CIFAR-10 and 860 samples for GTSRB datasets. Due to the large number of models with varying test accuracy it was not possible to get an equally balanced sample distribution per class well classified by all models as in Section 3.4.8and Section 3.5.6. Figures Figure 71 and Figure 72 show the sample class histogram for each dataset. 4.4 adversarial defenses experiments on cifar-10 Before we start the defense evaluation per se, we tested how the trained models fared in the test set of CIFAR-10. As can be seen in Table 37 the defended models present a significant decrease in accuracy in the test set. This is particularly critical for Adversarial Trained models, where the drop in accuracy is between 6% and 10%. Regardless of the level of success obtained with the defensive methods, it is clear that will come with a cost in accuracy in regular test data. In this section we will evaluate this level of defense. 4.4. Adversarial Defenses experiments on CIFAR-10 93 R50 R50 distil R50 adv L∞R50 adv L1R50 adv L2 CIFAR-10 91.52 ±0.13 88.07 ±0.18 81.77 ±4.72 84.19 ±3.45 84.93 ±2.37 Table 37.: Defensive models and normal ResNet-50 average accuracies on CIFAR-10 test set and corresponding standard deviation For this experiment we used the same methodology as in Section 3.4. We gathered 1000 samples correctly predicted by all the networks (including 5 ResNet-50 with Defensive Distillation and 5 ResNet-50 with Adversarial Training for each of the 3 norms). Hence, this is a different set of samples from the one used in Section 3.4. With these samples we used the model generator to craft the adversarial samples using the same parameters found in Section 3.4and only kept the adversarial samples that misclassified the generator to attack the other models. This experiment has 2setups: • non-targeted environment: attack the defensive models with non-targeted adversarial samples; • target environment: attack the defensive models with targeted adversarial samples aiming to make the models misclassify the samples as a specific class. 4.4.1Non-targeted environment In this setup we compare transferability of the non-targeted adversarial samples against the defensive models. In this comparison we used ResNet-50 models from Section 3.4plus 20 new defensive models (considering 2 defensive methods and three norm variations in adversarial training). As in Section 3.4, we picked only the samples that were misclassified by the generator model when converted to valid images ([0, 1]range). Table 38 presents the percentage of samples used for each combination of attacking method/norm/SSIM. Theses samples were then fed to all the defensively trained models in order to evaluate the transferability rate. Figure 73 shows the different attacks and their transferability rate between models. The percentage of misclassified samples per attack can be seen in Table 38. 4.4. Adversarial Defenses experiments on CIFAR-10 94 Figure 73.: CIFAR-10 non-targeted attacks and corresponding transferability transfer rate. Attacks with * are those with norms that attained best transferability rates on Section 3.4 From Figure 73 and Table 38 we can observe the following regarding the defensive methods: • Defensive distillation shows a higher transferability rate than the original ResNet-50 for a significant number of attacks. This is true considering the attacks crafted with FGSM and Deepfool where the performance of the distilled networks is significantly worse than all the other tested models; • Adversarial training is more robust in these tests, with a significant decrease in transferability rate, across all attacks, when considering the Ssetting. For the Lsetting the defense is not as effective, having a higher robustness gain when compared to the non-defensive model; • Although PGD was used to train the adversarial models, as an attacking option it remains the stronger attacking method. Considering the norms used in the adversarial training, there does not seem to be a significant difference regarding the norms. Also regarding the norms, but considering the attack methods, FGSM and PGD, there is also no noticeable difference apart from a slight advantage FGSM L∞L. 4.4. Adversarial Defenses experiments on CIFAR-10 95 Comparing the transferability rates between the ResNet-50 and the defensive models, it is clear that, although there is a reduction when using Adversarial Training, none of the methods provides a true defense in the context of this experiment. Nevertheless, apart from FGSM (in all settings) and Deepfool (L), all other attacking methods suffer substantially in their transferability rate when confronted with Adversarial Trained models. Naturally, this increase in robustness is more noticeable when considering the Ssetting. On the other hand, there is a significant trade-off regarding the accuracy obtained in legitimate samples as shown in Table 37. nt atks R50 R50 distil R50 adv l1R50 adv l2R50 adv l∞I% C% F% FGSM L1S36.85 46.53 27.53 26.03 28.48 48.3 99.59 48.1 FGSM L1L66.42 71.91 64.79 65.19 63.27 54.5 98.9 53.9 FGSM L2S * 38.57 49.25 28.89 26.82 28.48 38.7 100.0 38.7 FGSM L2L66.71 72.07 64.66 63.99 63.08 41.8 99.52 41.6 FGSM L∞S39.09 55.72 28.17 26.13 27.8 71.9 99.44 71.5 FGSM L∞L * 79.47 88.15 75.43 75.53 73.17 81.1 99.26 80.5 PGD L1S * 94.57 92.86 47.28 51.27 49.15 90.7 100.0 90.7 PGD L1L * 99.73 99.51 80.75 84.95 82.13 93.7 100.0 93.7 PGD L2S 94.28 93.02 47.14 51.27 48.47 76.5 100.0 76.5 PGD L2L 99.35 98.98 80.81 84.5 82.16 84.4 100.0 84.4 PGD L∞S 94.191.9 42.64 44.3 42.84 100.0 100.0 100.0 PGD L∞L 99.62 99.48 79.68 82.14 79.5 100.0 100.0 100.0 Deepfool S 38.56 43.91 28.53 26.31 28.94 95.8 99.9 95.7 Deepfool L 75.17 78.75 72.14 71.88 69.73 87.3 99.08 86.5 Carlini L2S 83.62 82.52 34.82 37.02 37.6 100.0 100.0 100.0 Carlini L2L 98.097.44 73.64 75.72 74.38 100.0 100.0 100.0 Table 38.: CIFAR-10 non-targeted transferability table. I% is the initial percentage of attacks successful on generator, C% is the percentage of successful attacks on generator after conversion to valid image, and F% is the final percentage of successful attacks on generator that were used to attack the other networks, this is a pure white box attack. In bold are the highest values per attack and in green and blue cells are the highest transferability values of all L and S attacks respectably per, model architecture. Attacks with * are those with norms that attained best transferability rates in Section 3.4 4.4.2Targeted environment In this setup we compare transferability of the targeted adversarial samples against the defensive models. We test the transferability of misclassification for the target class. All the attack samples generated used the same norms and parameters found in Section 3.4. 4.4. Adversarial Defenses experiments on CIFAR-10 96 Recall that we only consider a successful transfer if the sample gets classified with the target class, i.e. a sample being misclassified in any other class is not considered a successful transferred attack. Figure 74.: CIFAR-10 targeted attacks and corresponding transferability transfer rate. Attacks with * are those with norms that attained best transferability rates on Section 3.4 From Figure 74 and Table 39 it is clear that the trend observed in the non-targeted attacks can also be observed in the targeted attacks, namely: • Comparing to the non-targeted environment, the distilled models manage to achieve some success with all methods except FGSM. However, Adversarial Training is still the most robust method by a comfortable margin; • Considering Defensive Distillation, the attack with the highest transferability rate is PGD L∞for both in S and L. For Adversarial Training models the highest transferability rate is obtained with FGSM L∞L. All defensive methods performed poorly in all FGSM attacks specially on L settings, in some instances performed worse than the normal models. • Adversarial trained models performed very well against PGD and Carlini L2. Although the adversarial trained models where trained using different L norms it seems that didn’t help to improve their performance against the same L norm attacks, all ad- 4.5. Adversarial Defenses experiments on GTSRB 97 versarial trained models performed similarly, this behavior was also seen on the non target setting. Although Adversarial Training achieves very robust results, reducing very significantly the transferability rate (apart from attacks with FGSM), the trade-off in accuracy is still worth taking into consideration when considering deploying these defenses. Regarding the attacking methods, FGSM is the most effective in the targeted mode, as opposed to the non-targeted version, where it has the lowest transferability rate in almost every setting. targeted atks R50 R50 distil R50 adv l1R50 adv l2R50 adv l∞I% C% F% FGSM L1S 16.82 14.88 9.57 9.19 8.15 21.5 98.14 21.1 FGSM L1L25.36 28.57 26.29 25.57 22.29 15.3 91.5 14.0 FGSM L2S * 16.47 15.98 9.25 9.35 8.5 21.7 98.62 21.4 FGSM L2L25.18 28.92 25.61 27.05 23.17 15.0 92.67 13.9 FGSM L∞S 15.06 14.06 8.92 7.63 6.43 25.3 98.42 24.9 FGSM L∞L * 35.78 36.33 36.735.6 30.09 11.2 97.32 10.9 PGD L1S 34.88 29.04 6.74 8.38 6.78 100.0 100.0 100.0 PGD L1L 62.62 48.66 20.28 23.62 18.08 100.0 100.0 100.0 PGD L2S 32.42 27.18 6.58 8.2 6.64 100.0 100.0 100.0 PGD L2L 58.12 45.44 18.96 22.64 17.54 100.0 100.0 100.0 PGD L∞S * 54.748.28 9.18 10.14 8.44 100.0 100.0 100.0 PGD L∞L * 81.37 70.68 25.44 27.22 21.42 100.0 100.0 100.0 Carlini L2S 24.63 20.54 2.36 3.4 3.0 100.0 100.0 100.0 Carlini L2L 65.11 56.77 16.92 19.45 16.35 95.5 99.79 95.3 Table 39.: CIFAR-10 targeted transferability table. I% is the initial percentage of attacks successful on generator, C% the percentage of attacks successful on generator after conversion to valid image, and F% final percentage of successful attacks on generator that were used to attack the other networks, this is a pure white box attack. In bold are the highest values per attack, and in green and blue cells are the highest transferability values of all L and S attacks respectably per model architecture. Attacks with * are those with norms that attained best transferability rates on Section 3.4 4.5 adversarial defenses experiments on gtsrb Similarly, to what was done for CIFAR-10, we tested how the trained models fared in the test set of GTSRB. Although not as severe as in CIFAR-10, the level of defense will come with a cost in accuracy in regular test data, see Table 40. The relative ranking in accuracy remains the same as in CIFAR-10 (Table 37)), with Defensive Distillation models performing better on non-adversarial samples than Adversarial Trained models. 4.6. Defense Analysis 104 far target atks R50 R50 distil R50 adv l1R50 adv l2R50 adv l∞I% C% F% FGSM L1L 2.56 1.54 0.00 0.00 0.00 4.88 92.86 4.53 FGSM L∞L * 2.78 0.74 0.00 0.00 0.00 6.4 98.18 6.28 PGD L1S * 2.09 0.26 0.00 0.00 0.00 98.72 98.82 97.56 PGD L1L * 18.69 3.72 0.05 0.00 0.05 100.0 100.0 100.0 PGD L2S 1.60.23 0.00 0.00 0.00 100.0 100.0 100.0 PGD L2L 17.67 4.12 0.05 0.05 0.07 100.0 100.0 100.0 PGD L∞S 0.74 0.03 0.00 0.00 0.00 88.95 96.86 86.16 PGD L∞L 6.87 1.12 0.00 0.00 0.00 100.0 99.88 99.88 Carlini L2S 0.50.1 0.02 0.00 0.00 100.0 97.91 97.91 Carlini L2L 8.15 2.17 0.24 0.17 0.1 97.33 100.0 97.33 Table 43.: GTSRB targeted to farthest class transferability table. I% is the initial percentage of attacks successful on generator, C% the percentage of attacks successful on generator after conversion to valid image, and F% final percentage of successful attacks on generator that were used to attack the other networks, this is a pure white box attack. In bold are the highest values per attack, and in green and blue cells are the highest transferability values of all L and S attacks respectably per model architecture. Attacks with * are those with norms that attained best transferability rates on Section 3.5 of transferability in CIFAR-10. While in GTSRB this does not occur, it still allows for significantly higher transferability rates than Adversarial Training. Even though the adversarial samples were crafted with PGD, it is still the most successful attack method, which attests once again to its effectiveness for adversarial attacks. The target environment experiments show a significant decrease in transferability as expected. When considering closest and farthest classes in GTSRB it becomes clear that the similarity of the classes, measured with SSIM, has a strong impact in transferability. Regarding the level of perturbation (S and L) it is clear that the smallest level of perturbation leads to a higher robustness in both methods as expected. Comparing the performance amongst datasets is harder as there are too many factors at stake. Both the inter-class and intra-class variability are significantly different, and the accuracies obtained in the test set are also substantially dissimilar. In CIFAR-10 the accuracy obtained with all models is very low compared to the accuracy for the GTSRB test set. CIFAR-10 classes have a very high intra-class variability, whereas in GTSRB the variability within each class is significantly lower. Furthermore, inter-class variability is also lower in the traffic sign dataset. Both variations are related to the boundaries of the classes in the sample space, and the accuracies on the test set for each dataset also reflect this variability, with models trained on CIFAR-10 having a much lower accuracy than models trained in GTSRB. In order to evaluate the impact of the accuracy of the trained models, we have performed another test on GTSRB. This time, the models were early stopped when an accuracy similar 4.6. Defense Analysis 105 to that obtained with CIFAR-10 was attained. This results in relatively poorly trained models considering the accuracy previously obtained in GTSRB. The goal is to evaluate how the accuracy of the model influences its transferability and its defense capabilities. Figure 78 presents the transferability rates for non-targeted GTSRB where the adversarial samples are crafted with a high accuracy ResNet-50. Figure 78.: GTSRB non targeted attacks generated by a ResNet-50 well trained, and corresponding transferability rate on normal and defensive trained models with high and low accuracies. Models with L letter in front of the architecture name are low accuracy models that imitate CIFAR-10 test accuracy. Attacks with * are those with norms that attained best transferability rates on Section 3.5 In this chart we can observe that the transferability in lower accuracy models is consistently higher than on well trained models. Note that the set of adversarial samples is now different from the one used in Section 4.5.1, since all samples used to craft the adversarial attacks must be well classified by a larger set of models. 4.6. Defense Analysis 106 Figure 79.: GTSRB non targeted attacks generated by a ResNet-50 with low accuracy, and corresponding transferability rate on normal and defensive trained models with normal and low accuracies. Models with L letter in front of the architecture name are low accuracy models that imitate CIFAR-10 test accuracy. Attacks with * are those with norms that attained best transferability rates on Section 3.5 Figure 79 presents a similar experiment, but this time the adversarial samples were crafted in the low accuracy ResNet-50. As in the previous chart we can also observe that the transferability in lower accuracy models is consistently higher than on well trained models. These two experiments suggest that the higher the accuracy of the model the more robust it is. Taking into account the accuracy of the model used to craft the adversarial samples we observe distinct behaviours for different attacks. Comparing Figure 78 to Figure 79 we can observe that while FGSM and Carlini show overall an increase in transferability, for PGD an increase of robustness for the defensive methods can be observed. Hence, this suggests that the accuracy of the generator model used to craft the adversarial samples has an impact on the robustness of the defensive models. However, due to having different results for different attacking methods, the effect of such an impact is not universal. Note that the set of adversarial samples used in each situation is different, which also makes it harder to establish a direct comparison. 5 CONCLUSION Nowadays convolutional neural networks are present in all kinds of fields and markets, from facial recognition to unlocking mobile phones to autonomous driving cars. Although this technology is widespread, the discovery of adversarial examples came to prove that not only these systems are not entirely secure but the attacks that create misclassification can be undetectable to the human eye. This can lead to serious security breaches such as impersonating other people identities, to misclassification of traffic signs in autonomous driving cars. With the appearance of these problems it became a necessity to devise techniques that aimed to improve resistance against these adversarial attacks. We focused our tests on transferability, i.e., the creation of an adversarial sample in a model that causes misclassification on another model. To test the relative performance of the attacks we used SSIM to quantify the level of perturbation. Therefore, we were able to test all the attacks with the same level of perturbation. To evaluate transferability across a range of models we selected models from the same architecture family of the adversarial sample generator, as well as significantly different models. We also used two datasets with different features to evaluate how the dataset itself affects transferability. Our results show that some attacks have high levels of transferability on non-targeted attacks, with PGD being the clear winner in all tests performed. Considering targeted attacks, we see the transferability rates drop to much lower levels, as expected. For targeted attacks we also showed that using a sample closer to the target class as a base to generate the adversarial sample, based on SSIM, provides a higher transferability rate. Regarding the architecture similarity of the tested models, our results hint that this may impact transferability. Finally, considering the datasets, we noticed a direct correlation between the accuracy of the trained models and their robustness to transferability. However, since the composition of the datasets differs significantly we can’t conclude that accuracy is the only relevant 107 5.1. Prospect for future work 108 factor. Intra and inter-class variability can also play a role, and further testing is required to evaluate its significance. Regarding defensive methods, in our assessment we found that, on non-targeted environment, Adversarial Training improved significantly the robustness against transferability attacks on both datasets. Defensive Distillation on the other hand performed much worse on both targeted and non-targeted settings, and in some instances it even performed worse than normal trained models. As expected, the defensive methods worked far better in the targeted tests, with SSIM playing an important role. The defense models provided a higher robustness when the sample used to craft the adversarial sample was from a class that is further away from the target class. Concerning the datasets both inter-class and intra-class variability seem to affect the performance of defensive methods. The accuracy of the models also appears to influence the robustness of the methods. 5.1 prospect for future work While in this work we have done a significant number of experiments, there are a lot of open questions at the end. These questions emerged when we got the results of our tests and were able to analyse them. The present work can be seen as the beginning of a long journey into the adversarial world. Avenues for future work can be derived from the analysis performed both on the attacks and defenses. Concerning defenses, there are still questions pending regarding the transferability effectiveness of an attack: • To what extent is architecture similarity relevant? • Is shallowness pertinent? • Does intra and/or inter-class variability have an impact? Regarding defense robustness the following questions also require further work: • How much does the level of perturbation in Adversarial Training affect its robustness and accuracy? • How does intra and/or inter-class variability impact on the robustness of the defenses? • how does the accuracy of the models affect the transferability and defense robustness? 5.1. Prospect for future work 109 Finally, there are other defensive and attack methods, see Chakraborty et al. (2018), that were not explored in this thesis due to time and hardware constraints. It would be interesting to evaluate them in the context of the questions above. BIBLIOGRAPHY Nicholas Carlini and David Wagner. Towards evaluating the robustness of neural networks. arXiv:1608.04644v2 22 Mar 2017,2017. Anirban Chakraborty, Manaar Aalm, VishalL Dey, Anupam Chattopadhay, and Debdeep Mukhopadhay. Adversarial attacks and defences: A survey. arXiv:1810.00069v1[cs.LG] 28 Sep 2018,2018. Francois Chollet. Trains a resnet on the cifar10 dataset. URL https://keras.io/zh/ examples/cifar10_resnet/. Francois Chollet. Deep Learning with Python. Manning Publications, 2017. ISBN 978-1-61729443-3. Yonatan Geifman. cifar10vgg. https://github.com/geifmany/cifar-vgg/blob/master/ cifar10vgg.py,2018. Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep Learning. MIT Press, 2016. http://www.deeplearningbook.org. Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. CoRR, abs/1412.6572,2015. GTSRB. German traffic sign repository benchmark, 2010. URL https://benchmark.ini. rub.de/. Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. arXiv:1512.03385v1 10 Dec 2015,2015. Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv:1503.02531v1[stat.ML] 9Mar 2015,2015. Andrew Ilyas, Shibani Santurkar, D. Tsipras, L. Engstrom, B. Tran, and A. Madry. Adversarial examples are not bugs, they are features. In NeurIPS,2019. Alex Krizhevsky. Learning multiple layers of features from tiny images. Technical report, 2009. URL https://www.cs.toronto.edu/~kriz/learning-features-2009-TR.pdf. A. Kurakin, Ian J. Goodfellow, and S. Bengio. Adversarial machine learning at scale. ArXiv, abs/1611.01236,2017a. 110 Bibliography 111 A. Kurakin, Ian J. Goodfellow, and S. Bengio. Adversarial examples in the physical world. ArXiv, abs/1607.02533,2017b. Fei-Fei Li, Ranjay Krishna, and Danfei Xu. Cs231n: Convolutional neural networks for visual recognition. URL http://cs231n.stanford.edu/. Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. arXiv:1706.06083v4 19 Jun 2017,2017. S. Moosavi-Dezfooli, A. Fawzi, and P. Frossard. Deepfool: A simple and accurate method to fool deep neural networks. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2574–2582,2016. doi: 10.1109/CVPR.2016.282. Maria-Irina Nicolae, Mathieu Sinn, Minh Ngoc Tran, Beat Buesser, Ambrish Rawat, Martin Wistuba, Valentina Zantedeschi, Nathalie Baracaldo, Bryant Chen, Heiko Ludwig, Ian Molloy, and Ben Edwards. Adversarial robustness toolbox v1.2.0.CoRR,1807.01069,2018. URL https://arxiv.org/pdf/1807.01069. Nicolas Papernot, Patrick D. McDaniel, Somesh Jha, Matt Fredrikson, Z. Berkay Celik, and Ananthram Swami. The limitations of deep learningin adversarial settings. 1st IEEE European Symposium on Security & Privacy, IEEE 2016,2015. Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami. Distillation as a defense to adversarial perturbations against deep neural networks. arXiv:1511.04508v2[cs.CR] 14 Mar 2016,2016. Nicolas Papernot, Patrick McDaniel, Ian Goodfellow, Somesh Jha, Z. Berkay Celik, and Ananthram Swami. Practical black-box attacks against machine learning, 2017. Jaewoo Song. one-pixel-attack-keras. https://github.com/Hyperparticle/ one-pixel-attack-keras/blob/master/1_one-pixel-attack-cifar10.ipynb,2019. Jiawei Su, Danilo Vasconcellos Vargas, , and Kouichi Sakurai. One pixel attack for fooling deep neural networks. arXiv:1710.08864v7[cs.LG] 17 Oct 2019,2019. Christian Szegedy, W. Zaremba, Ilya Sutskever, Joan Bruna, D. Erhan, Ian J. Goodfellow, and R. Fergus. Intriguing properties of neural networks. CoRR, abs/1312.6199,2014. Antonio Torralba, Rob Fergus, and Bill Freeman. 80 million tiny images, 2020. URL http: //groups.csail.mit.edu/vision/TinyImages/. Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image quality assessment: From error visibility to structural similarity. IEEE TRANSACTIONS ON IMAGE PROCESSING, VOL. 13, NO. 4,2004. Bibliography 112 Rey Wiyatno and Anqi Xu. Maximal jacobian-based saliency map attack. arXiv:1808.07945v1 23 Aug 2018,2018. Youlixx. tf-jsma-nt.py. https://github.com/probabilistic-jsmas/probabilistic-jsmas/ blob/master/attacks/tf_jsma_nt.py,2020. Yinghua Zhang, Yangqiu Song, Jian Liang, Kun Bai, and Qiang Yang. Two sides of the same coin: White-box and black-box attacks. ArXiv,2008.11089v1,2020. URL https: //arxiv.org/abs/2008.11089. A ATTACK EVALUATION GRAPHS In this section we present the graphics that show the search parameters for all attacks and their norms for both CIFAR-10 and GTSRB, and their corresponding settings: targeted and non-targeted attacks. For the sake of completeness, even graphics included in the main text are repeated in here for convenience. a.1 cifar-10 search parameters a.1.1FGSM Figure 80.: Success attack rate of non-targeted FGSM L∞norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e 113 A.1. CIFAR-10 search parameters 120 a.1.3Carlini Figure 90.: Non-targeted Carlini L2norm success attack rate of on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per k Figure 91.: Targeted Carlini L2adversarial sample for some confidence k A.1. CIFAR-10 search parameters 121 a.1.4PGD Figure 92.: Success attack rate of non-targeted PGD L∞norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e A.1. CIFAR-10 search parameters 122 Figure 93.: Success attack rate of non-targeted PGD L1norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e Figure 94.: Success attack rate of non-targeted PGD L2norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e A.1. CIFAR-10 search parameters 123 Figure 95.: Success attack rate of targeted PGD L∞norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e Figure 96.: Success attack rate of targeted PGD L1norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e A.1. CIFAR-10 search parameters 124 Figure 97.: Success attack rate of targeted PGD L2norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e A.1. CIFAR-10 search parameters 125 a.1.5One pixel/Few pixels Figure 98.: Success attack rate of non-targeted Few pixels attack on each model and their corresponding SSIM and total successful adversarial attacks on model generator per number of pixels perturbed A.1. CIFAR-10 search parameters 126 Figure 99.: Success attack rate of targeted Few pixels attack on each model and their corresponding SSIM and total successful adversarial attacks on model generator per number of pixels perturbed A.2. GTSRB search parameters 127 a.2 gtsrb search parameters a.2.1FGSM Figure 100.: Success attack rate of non-targeted FGSM L∞norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e A.2. GTSRB search parameters 128 Figure 101.: Success attack rate of non-targeted FGSM L1norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e Figure 102.: Success attack rate of non-targeted FGSM L2norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e A.2. GTSRB search parameters 129 Figure 103.: Success attack rate of targeted for closest class FGSM L∞norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e Figure 104.: Success attack rate of targeted for closest class FGSM L1norm on each model and their corresponding SSIM and total successful adversarial attacks on model generator after converting to a valid image per e