scieee AI-readable full text Open interactive document viewer

CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP

Xing, Songlong; Zhao, Zhengyu; Sebe, Niculae

Full text

CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP Songlong Xing1Zhengyu Zhao2*Nicu Sebe1 1University of Trento, Italy 2Xi’an Jiaotong University, China {songlong.xing, niculae.sebe}@unitn.it [email protected] Abstract Despite its prevalent use in image-text matching tasks in a zero-shot manner, CLIP has been shown to be highly vulnerable to adversarial perturbations added onto images. Recent studies propose to finetune the vision encoder of CLIP with adversarial samples generated on the fly, and show improved robustness against adversarial attacks on a spectrum of downstream datasets, a property termed as zero-shot robustness. In this paper, we show that malicious perturbations that seek to maximise the classification loss lead to ‘falsely stable’ images, and propose to leverage the pre-trained vision encoder of CLIP to counterattack such adversarial images during inference to achieve robustness. Our paradigm is simple and training-free, providing the first method to defend CLIP from adversarial attacks at test time, which is orthogonal to existing methods aiming to boost zero-shot adversarial robustness of CLIP. We conduct experiments across 16 classification datasets, and demonstrate stable and consistent gains compared to test-time defence methods adapted from existing adversarial robustness studies that do not rely on external networks, without noticeably impairing performance on clean images. We also show that our paradigm can be employed on CLIP models that have been adversarially finetuned to further enhance their robustness at test time. Our code is available here. 1. Introduction With the increasing availability of image-text data and the advancement of self-supervised learning techniques [7,8,16], vision-language models (VLM) have continued to spark research interests in both academia and industry [21,28,39,40,42,56]. As a representative VLM, CLIP [39] has shown impressive abilities to match an image with its descriptive text in a zero-shot manner. However, recent studies have shown that adding small imperceptible perturbations to an image can cause CLIP to misclassify it *Corresponding author Figure 1. Test-time counterattacks harness the expressive power of CLIP to generate a counterattack to defend CLIP against adversaries without finetuning the vision encoder. [26,27,32,44,50,59,63], a common problem plaguing nearly all neural networks [2,6,25,29,33,36,48,57,58]. As foundational models are deployed in real-world applications, their safety and reliability have become a pressing concern. In this paper, we focus on the robustness of CLIP against adversarial perturbations. Unlike conventional models for which adversarial robustness has been extensively studied [2,6,11,33,47,48], CLIP is a pre-trained foundation model that has learned massive amounts of real-world knowledge, and should be dealt with carefully to minimise damage to its generalisation abilities. Adversarial robustness of CLIP has just started to garner research attention [26,32,44,50] in recent years. Existing efforts fall into two categories. The first type is based on adversarial training [3,58], which alternately generates adversarial images on one dataset and uses them to finetune the vision encoder of CLIP [32,50]. This type of methods, known as adversarial finetuning (AFT), dynamically mimics a min-max game between CLIP and the threat model in the finetuning phase, and deploys the finetuned model in a wide variety of downstream classification tasks without further training. This method shows transferable robustness to downstream datasets, a property termed as zero-shot robustness [32,50]. The other type of This CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore. 15172 methods resorts to prompt tuning [61,62], which inserts learnable text tokens in the embedding space, aligns the corresponding text prompts with adversarial images, and tunes the learnable tokens by propagating gradients to the text embeddings. This approach is known as adversarial prompt tuning (APT) [26,59]. Although these methods have shown improved robustness over the original CLIP, there are apparent limitations. Firstly, they require time-consuming training, especially adversarial finetuning which involves generating adversarial images on the fly. Secondly, the model overfits to the training data, which compromises generalisation on clean and adversarial images from other data distributions. In the case of adversarial finetuning, the adversarially finetuned CLIP outperforms the original CLIP on clean images from the dataset used for finetuning, indicating that the model has overfitted to the distribution of the dataset (See Tab. 1in Sec. 4). Thirdly, for adversarial finetuning methods [32,50], robustness to adversarial samples comes at the cost of a significant decline of classification accuracy on clean images. To address these limitations, we draw inspiration from existing adversarial robustness studies that robustify nonfoundational target models at test time [1,17,31,52], and propose a test-time paradigm that utilises the expressive power of CLIP to defend itself from adversarial attacks (Fig. 1). Following previous studies, we focus on adversaries that aim to maximise the classification loss of CLIP given a test image. We observe that such adversaries cause images to be more stable when a small random noise is added, compared to clean images. We term this behaviour of adversarial images ‘false stability’, which can be interpreted as the images being trapped in a toxic surrounding in the latent space by an adversary. Intuitively, the pre-trained vision encoder of CLIP is highly expressive, and can be leveraged to push the adversarial image away from its toxic embedding. Therefore, we propose to employ the vision encoder of CLIP to counteract the false stability of adversarial images, thereby achieving robustness to these attacks. Since no label information is available at test time, we formulate the test image as an anchor, and iteratively update the counterattack perturbation such that it maximises the L2distance with the anchor in the embedding space [44]. However, pushing the test image away from its embedding risks hurting performance on clean images. To address this, we propose τ-thresholded weighted counterattacks, which employ a threshold to prevent further counterattacking if the test image does not exhibit false stability, thus preserving performance on clean images. To our best knowledge, our paradigm is the first work to utilise the expressive power of CLIP to defend itself from adversarial attacks, and can be categorised as a test-time input purification method for VLMs. We conduct extensive experiments and analyses on 16 classification datasets, establishing our method as an effective test-time defence for CLIP. We summarise the main contributions of this paper as follows: • We propose the first method that harnesses the power of CLIP to defend itself from adversarial attacks at inference time without relying on any auxiliary networks. Our method is simple and training-free, and can be easily employed in other VLMs. • We propose a test-time counterattack method based on Projected Gradient Descent (PGD). We show that our method can defend CLIP with a small number of counterattack steps without significantly impacting performance on clean images. • We conduct experiments across 16 classification datasets, and demonstrate superior performance compared to testtime defences adapted from existing adversarial robustness literature. Our paradigm can be employed on adversarially finetuned CLIP to further enhance robustness performance at test time. 2. Related Work Adversarial Robustness. Since the early development of deep neural networks [18,24,49], it has been found that they are vulnerable to adversarial attacks. Specifically, a small adversarial perturbation bounded by a Lp-radius ball, usually imperceptible by humans, can cause the network to misclassify the sample entirely [6,48]. To address this vulnerability, adversarial training (AT) [29,41,58] alternately generates adversarial samples with the target model on the fly, and trains the network with these adversarial samples. This practice has shown significantly improved robustness to adversarial attacks, and has become a de facto standard in adversarial machine learning, despite the presence of limitations such as expensive training [45,51]. Other types of methods are also proposed, of which the most related to this work is test-time defence. This can be achieved by employing a generative model to purify the test image with an auxiliary generative model [34,43,55], or by adjusting the test image based on an objective [1,20,31,52]. However, subsequent studies show that test-time defence can be circumvented by adaptive attacks specially designed for the defence [12]. Amongst these methods, Hedge Defense (HD) [52] is the most closely related to our work. They attack the test image by maximising the cross entropy loss with respect to all classes, based on the finding that the loss surface is smoother around the ground-truth label. An important difference between HD and this work is that they apply the defence method on an adversarially trained model. In contrast, we focus on CLIP and show that foundation models like CLIP possess the inherent ability to defend themselves against attacks that seek to maximise the classification loss, by producing a counterattack perturbation that leads the adversarial image away from its original embedding in the latent space, with no need for an adversarially trained model. 15173 Figure 2. Pipeline to generate an adversarial perturbation δgiven an image xand its ground-truth label based on CLIP. Black and red arrows denote the forward and backward pass, respectively. VLMs and Their Adversarial Robustness. Recently, adversarial robustness of foundation models have garnered increasing research attention [46,60]. This paper focuses on enhancing the adversarial robustness of CLIP [39] since it is a representative foundation model that aligns images and text. Existing methods can be divided into two types: (1) Adversarial finetuning. Mao et al. [32] propose TeCoA, which finetunes the vision encoder of CLIP using adversarial samples generated on the fly on one dataset, and transfers the learned robustness to downstream classification datasets. Based on this pipeline, Wang et al. [50] further propose to employ the original CLIP to guide adversarial finetuning by imposing two regularisation terms, showing improved generalization on clean and adversarial images across downstream datasets. (2) Adversarial prompt tuning. This line of research is built on prompt tuning of CLIP [61,62], where model weights are kept frozen. Li et al. [26] show that textual prompts play an important role in the effectiveness of both adversarial attacks and robustness. They propose to insert learnable tokens and tune the tokens by aligning the textual prompt with adversarial images. Zhang et al. [59] propose a similar pipeline by training robust text tokens, assuming that the attacker has access to the model but not to the text prompts employed by the end user. In this work, we discard any training and show that CLIP possesses the ability to defend itself from adversarial attacks by counterattacking adversarial images. Our method is the first test-time defence method for CLIP. 3. Methodology In this section, we first provide preliminaries regarding CLIP and adversarial robustness for CLIP in an image classification context. Then we proceed to introduce our testtime counterattack paradigm. 3.1. Preliminaries and Setup Zero-shot classification of CLIP. CLIP [39] is a visionlanguage model that matches images with their descriptive text. It has been contrastively pre-trained on 400 million image-text pairs. Specifically, it aligns an image xwith its Figure 3. Our test-time counterattack paradigm. We craft a counterattack perturbation δttc to lead an adversarial image away from its original embedding at test time without finetuning. corresponding text tthrough the cosine similarity between their representations produced by a vision encoder fθ(·)and a text encoder gϕ(·), respectively, where θand ϕare model weights. At inference time, CLIP performs classification in a zero-shot manner. Given a set of classes defined in their textual names c1, . . . , cK, CLIP matches a test image xagainst the textual prompts corresponding to candidate class names wrapped by a template T(usually ‘a photo of [CLASS]’): si=fθ(x)Tgϕ(T(ci)) ∥fθ(x)∥ · ∥ gϕ(T(ci)) ∥(1) The probability of xbelonging to class ciis calculated as the normalized similarity pi=exp(si) Pjexp(sj), and the candidate class with the highest probability is the predicted class. Adversarial attacks for CLIP. CLIP is highly vulnerable to adversarial perturbations [32]. In a setting where the attacker has full knowledge of the model weights and gradients of CLIP, a small perturbation δbounded by a Lp-radius can be maliciously designed to cause CLIP to misclassify: δa= arg max δL(x+δ, tc), s.t. ∥δ∥p≤ϵa(2) where tcis the ground-truth label of xand Lis a loss function which is usually a cross-entropy loss, and ϵais the attack budget, which ensures that the manipulation is subtle and imperceptible to humans. Eq. (2) can be approximated by Projected Gradient Descent (PGD) [6]. The adversarial image is the addition of the image and the perturbation x′:= x+δa. This process is illustrated in Fig. 2. Adversarial finetuning strengthens the adversarial robustness of CLIP [32,50] by alternately generating adversarial images x′following Eq. (2) and using them to finetune fθ. TeCoA [32] performs finetuning by aligning x′with the ground-truth text T(tc)on one dataset. PMG-AFT [50] imposes two CLIP-guided regularization terms on top of TeCoA to improve the generalization on clean and adversarial images. After the finetuning phase, CLIP has learned robustness to adversarial attacks, which transfers to downstream classification datasets without further training [32]. 15174 3.2. Test-time Counterattacks Although adversarial finetuning has been shown to significantly improve CLIP’s adversarial robustness, limitations are apparent, such as cumbersome training involving the generation of adversarial samples and finetuning of the vision encoder weights θ. In this paper, we investigate the ability of CLIP to defend itself at test time, with no need for any training, providing the first test-time defence method for CLIP. Our proposed paradigm, termed as Test-time Counterattacks (TTC), is illustrated in Fig. 3. Following previous studies [26,32,50,59], we focus on attacks that aim to maximise the classification loss of CLIP. Intuitively, a pre-trained vision encoder fθis highly expressive in capturing the nuanced pixel pattern in an image. In this sense, we speculate that an adversarial image that successfully fools CLIP is trapped in a toxic surrounding induced by the adversary, and that the pre-trained fθis able to lead the image away from this surrounding by utilising its expressiveness. Recently, Schlarmann et al. [44] propose an unsupervised attack method for CLIP, where a perturbation is updated such that it maximises the L2distance between the image and its original embedding. Inspired by this label-free attack method, we employ the same loss function in our paradigm. Specifically, we employ the original image embedding fθ(x)as the anchor, and craft a test-time counterattack perturbation δttc such that the L2 distance between the embedding of the counterattacked image fθ(x+δttc)and the anchor fθ(x)is maximised: δttc = arg max δ∥fθ(x+δ)−fθ(x)∥, s.t. ∥δ∥p≤ϵttc (3) This counterattack can also be approximated by PGD [6]. Since this counterattack is performed by the end user at test time, the counterattack does not need to be imperceptible, hence a large user-defined counterattack budget ϵttc. However, we hope to maintain a consistent attack style with existing studies, and still keep the counterattack budget low, bounded by a Lp-radius. In the experiments (Sec. 4), we show that a budget at ϵttc = 4/255 is able to improve CLIP’s adversarial robustness significantly. Note that the vision encoder weights θare kept frozen throughout. Among existing methods for non-foundational models, the most closely related to ours is hedge defense (HD) [52]. An important difference is that they employ HD on adversarially-trained models, whilst we show that CLIP without adversarial finetuning can harness the expressiveness of its vision encoder to defend itself. 3.3. τ-thresholded Weighted Counterattacks Sec. 3.2 has discussed the idea of defending CLIP with its pre-trained vision encoder. An undesirable risk is that the counterattacks can hurt natural images as well. Based on the idea of TTC, we further propose τ-thresholded (a) CIFAR10 (b) ImageNet Figure 4. Ratio of L2drift due to a random noise. The value of τ is the average τacross 100 randomly selected samples. weighted counterattacks to counterattack adversaries effectively while reducing the impact on clean images. Wu et al. [52] show that adversarial images are more vulnerable to a small noise than clean images. In this study, we find that adversarial images are actually more robust to small random noises, and are only vulnerable to sufficiently large noises, based on our analysis of adversarial images obtained by iterative attack methods (PGD [6] in our case). Specifically, we define a stochastic variable τinduced by a random noise n∈ RC×W×H∼U(−ϵrandom, ϵrandom), conditioned on an image x∈ RC×W×H: τ=∥fθ(x+n)−fθ(x)∥ ∥fθ(x)∥(4) which can be interpreted as the ratio of the L2drift in the latent space when a random noise nis applied on an image. The values of τare reported in Fig. 4for ImageNet and CIFAR10. We report more results and analysis on τfor other datasets in Appendix (Sec. 7). As can be seen from Fig. 4, when a small random noise (ϵrandom = 1/255,4/255) is imposed, the ratio of L2drift in the latent space is unusually small, showing that they are trapped in a toxic surrounding and rendered ‘falsely stable’ by an adversary. Adversarial images only become vulnerable when the strength of random noise is increased, as evidenced by the disproportionately rising values of τ. We term this behaviour of adversarial images obtained by maximising CLIP’s classification loss ‘false stability’, and provide more theoretical analysis in Appendix (Sec. 7). Built upon the analysis above, we propose τ-thresholded weighted counterattacks based on PGD [6]. Specifically, we follow a standard pipeline of PGD iterations, with the attack objective being Eq. (3). At the zero-th iteration, a random perturbation without any update δ0 ttc is applied, where we compute the τvalue based on Eq. (4) as an indicator. If it is higher than a user-defined threshold τthres, meaning that it is not ‘falsely stable’, we halt the counterattack and return the random noise δ0 ttc. Otherwise, the counterattack is resumed. Note that the selection of τthres is dependent only on the τvalue of clean images and the strength of random 15175 Algorithm 1 τ-thresholded weighted counterattacks. Require: Test image x, pre-trained CLIP vision encoder fθ, counterattack budget ϵttc, stepsize α, number of steps N, user-defined parameters τthres and β. 1: procedure TEST-TIME COUNTERATTACKS 2: δ0 ttc ∼U(−ϵttc, ϵttc). 3: Compute τbased on Eq. (4) using δ0 ttc. 4: if τ≥τthres then 5: w0= 1 6: return δttc =δ0 ttc 7: else if τ < τthres then 8: w,δttc := {},{} 9: for i= 1,2, . . . , N do 10: δi ttc=Π(δi−1 ttc +α∇δ∥fθ(x+δi−1 ttc )−fθ(x)∥) 11: wi= exp(β·i)/PN j=0 exp(β·j)(Eq. (5)) 12: w←wi,δttc ←δi ttc 13: end for 14: return δttc =PN i=0 wi·δi ttc (Eq. (6)) 15: end if 16: end procedure noise, irrespective of types and strengths of attacks. Since employing only one δttc may be suboptimal, we weight the counterattack perturbation vectors across all steps: wj=exp(β·j) PN j=0 exp(β·j)(5) δttc = N X j=0 wjδj ttc (6) where β > 0is a hyperparameter controlling the ascending rate of weights, Nis the number of steps for performing the counterattack, and δj ttc is the counterattack perturbation obtained after jsteps. We summarise our τ-thresholded weighted counterattacks in Algorithm 1. 4. Experiments In this section, we conduct extensive experiments to verify the effectiveness of our test-time counterattack paradigm. 4.1. Experiment setup Datasets. Following previous work [32,50], we conduct our experiments on 16 datasets, which include general object recognition datasets CIFAR10 [23], CIFAR100 [23], STL10 [10], ImageNet [13], Caltech101 [14] and Caltech256 [15]; fine-grained recognition datasets OxfordPets [37], Flowers102 [35], Food101 [5], StanfordCars [22]; scene recognition datasets SUN397 [53], Country211 [39]; domain-specific datasets FGVCAircraft [30], EuroSAT [19], DTD [9], and PCAM [4]. Implementation Details. We use a counterattack budget of ϵttc = 4/255 and a threshold τthres = 0.2. We set the number of steps for counterattacks as N= 2, unless otherwise stated. βis set to 2.0. All attacks and counterattacks in experiments are bounded by a L∞radius. Note that the selection of τthres is dependent on ϵttc, and is determined based on the τbehaviour of clean images (Fig. 4). A higher τthres trades off more clean accuracy for robustness. We provide more details in Appendix (Sec. 9and Sec. 11). Baselines. Since there are no test-time defence methods for CLIP, we implement several test-time methods from existing adversarial robustness studies that do not rely on auxiliary networks. Specifically, we implement Anti-adversary [1] and Hedge Defense (HD) [52], which are the most closely related to our method. Anti-adversary [1] generates a perturbation to reinforce the confidence of the classifier given a test image. HD [52] performs a counterattack on the test image by increasing the cross-entropy w.r.t. all candidate classes, based on their finding that the loss function surface is smoother around the ground-truth class. We also adapt their method in our experiments with CLIP. For these two methods, we employ a test-time perturbation budget of 4/255, which is equal to the counterattack budget of our TTC. Following the original papers, the number of steps are 2 and 20 for Anti-adversary and HD, respectively. Considering that previous studies establish image transformations as a simple and effective defence method [17,38,54], we also implement test-time transformation ensembling (TTE) [38], which ensembles image transformation as a defence. We implement TTE with image flip, 4 crops, and image flip of all crops, totaling 9 augmentative views. As a simplest baseline, we also include random noise (RN) which adds a random perturbation noise with the same strength as our ϵttc, i.e., n∼U(−ϵttc, ϵttc). As a useful reference, we also implement adversarially finetuning methods TeCoA [32], PMG-AFT [50] and FARE [44] by finetuning the CLIP vision encoder with adversarial images on TinyImageNet, based on the objective functions proposed in their papers1. We also finetune CLIP with clean images (CLIP-FT) on TinyImageNet. In the phase of finetuning, we use a 2-step PGD attack, with the stepsize α= 1/255 and attack budget ϵ= 1/255, following [32,50]. The learning rate for finetuning is 5e−5. After the finetuning phase, the finetuned models are deployed on 16 downstream datasets. 4.2. TTC on Original CLIP Robustness under ϵa= 1/255.We first test the robustness of all methods under the attack budget of ϵa= 1/255. Following previous studies on CLIP’s adversarial robustness 1Unlike original implementations, we hold out 10% of the training set of TinyImageNet for evaluation in our implementation, without consulting downstream datasets. We also find preprocessing significantly affects the performance of finetuned models on CIFAR10, CIFAR100, and STL10. We follow the preprocessing pipeline recommended by CLIP (Tab. 1). 15176 (%) CLIP Adversarial Finetuning Test-time Defence ∆ CLIP-FT TeCoA PMG-AFT FARE RN TTE Anti-adv HD TTC (ours) TinyImageNet Rob. 0.19 2.19 48.64 46.12 25.47 0.28±0.02 19.52±4.21 4.46±0.23 3.11±0.05 20.64±0.17 +20.45 Acc. 57.64 77.06 70.86 66.85 73.63 51.83±0.16 56.74±0.22 52.55±0.06 51.37±0.15 51.84±0.17 -5.80 CIFAR10 Rob. 0.74 3.34 33.61 40.66 19.65 2.01±0.08 41.35±6.14 12.39±0.07 17.22±0.45 28.75±0.18 +28.01 Acc. 85.12 84.90 64.61 70.69 74.44 81.18±0.07 84.74±0.40 83.52±0.09 78.23±0.16 81.18±0.07 -3.94 CIFAR100 Rob. 0.26 0.90 18.95 22.52 11.40 0.67±0.05 20.06±4.03 5.73±0.04 3.86±0.10 14.31±0.25 +14.05 Acc. 57.14 59.51 35.96 40.32 46.67 56.34±0.20 58.61±0.25 53.95±0.15 52.86±0.16 56.34±0.20 -0.80 STL10 Rob. 11.0 12.73 70.08 73.08 59.06 16.23±0.08 78.48±3.83 37.42±0.40 39.02±0.30 76.70±0.23 +65.70 Acc. 96.40 94.49 87.40 88.56 91.72 95.85±0.04 96.26±0.04 95.45±0.08 89.50±0.07 95.85±0.04 -0.55 ImageNet Rob. 1.15 0.93 18.89 21.43 14.00 1.77±0.03 31.01±4.40 8.67±0.05 6.63±0.05 38.41±0.07 +37.26 Acc. 59.69 54.24 34.89 36.12 48.79 59.34±0.06 60.02±0.12 54.27±0.14 54.54±0.05 49.39±0.00 -10.30 Caltech101 Rob. 14.67 14.21 55.51 61.08 50.74 18.90±0.14 67.56±3.88 34.81±0.16 31.53±0.22 65.78±0.07 +51.11 Acc. 85.66 83.63 71.68 75.45 80.95 86.61±0.10 85.84±0.09 84.02±0.10 82.33±0.04 86.53±0.07 +0.87 Caltech256 Rob. 8.47 6.76 43.19 45.91 38.79 11.33±0.04 60.09±4.03 25.36±0.17 23.48±0.10 60.11±0.04 +51.64 Acc. 81.72 78.53 61.14 62.24 73.32 81.25±0.03 82.49±0.08 79.38±0.12 79.12±0.01 79.66±0.04 -2.06 OxfordPets Rob. 1.04 2.10 38.35 41.18 31.07 1.86±0.01 50.33±7.30 20.42±0.22 12.04±0.16 57.87±0.15 +56.83 Acc. 87.44 84.14 62.12 65.88 79.37 87.41±0.12 88.13±0.13 80.62±0.35 80.91±0.05 83.35±0.21 -4.09 Flowers102 Rob. 1.14 0.54 21.94 23.43 17.14 1.52±0.01 35.88±4.72 7.16±0.41 7.29±0.06 39.14±0.28 +38.00 Acc. 65.46 53.37 36.80 37.00 47.98 64.62±0.19 65.18±0.22 62.66±0.14 58.22±0.12 64.16±0.19 -1.30 FGVCAircraft Rob. 0.00 0.00 2.49 2.22 1.35 0.00±0.00 6.23±1.37 1.27±0.07 1.26±0.07 13.77±0.38 +13.77 Acc. 20.10 14.04 5.31 5.55 10.86 19.25±0.18 20.19±0.36 15.88±0.23 16.36±0.03 18.00±0.16 -2.10 StanfordCars Rob. 0.02 0.06 8.76 11.65 6.75 0.16±0.02 22.36±4.17 4.40±0.30 2.71±0.09 33.01±0.07 +32.99 Acc. 52.02 42.11 20.91 25.44 38.68 52.14±0.09 52.73±0.31 36.21±0.27 44.28±0.02 48.16±0.16 -3.86 SUN397 Rob. 1.14 0.94 19.39 22.58 14.91 1.72±0.01 30.79±4.43 8.05±0.04 6.40±0.06 41.52±0.04 +40.38 Acc. 58.50 55.73 36.69 37.98 52.42 59.69±0.06 59.12±0.08 56.00±0.04 53.17±0.02 55.13±0.06 -3.37 Country211 Rob. 0.04 0.03 1.78 2.12 0.85 0.06±0.00 3.05±0.89 0.67±0.05 0.47±0.02 7.09±0.04 +7.05 Acc. 15.25 12.07 4.75 4.64 9.26 14.80±0.02 14.66±0.16 11.58±0.12 11.72±0.07 13.08±0.05 -2.17 Food101 Rob. 0.70 0.42 13.90 18.57 11.65 1.20±0.01 43.94±6.97 13.12±0.16 8.03±0.11 57.84±0.15 +57.14 Acc. 83.88 64.86 29.98 36.61 55.31 83.44±0.04 83.96±0.02 75.81±0.22 80.30±0.05 82.18±0.02 -1.70 EuroSAT Rob. 0.03 0.04 11.96 12.60 10.67 0.15±0.01 6.91±2.13 2.15±0.04 4.57±0.09 12.19±0.24 +12.16 Acc. 42.59 27.64 16.58 18.53 21.88 53.24±0.09 44.38±1.60 36.78±0.18 39.08±0.06 53.24±0.09 +10.65 DTD Rob. 2.98 2.39 17.61 14.95 15.64 3.71±0.09 23.90±2.34 5.62±0.07 11.63±0.17 27.32±0.25 +24.34 Acc. 40.64 36.49 25.16 21.76 32.07 37.96±0.13 41.33±0.32 38.92±0.22 34.89±0.35 36.98±0.21 -3.66 PCAM Rob. 0.08 1.11 48.24 46.18 16.23 0.41±0.01 10.62±3.22 4.97±0.12 44.74±0.17 52.85±0.20 +52.77 Acc. 52.02 47.21 49.96 50.03 52.54 52.73±0.07 51.01±0.08 52.49±0.02 50.38±0.04 52.73±0.07 +0.71 Avg. Rob. 2.70 2.91 26.54 28.76 20.00 3.86±0.02 33.28±3.98 12.01±0.04 13.81±0.06 39.17±0.02 +36.47 Acc. 61.51 55.80 40.25 42.30 51.02 61.61±0.03 61.79±0.13 57.35±0.03 56.62±0.02 59.75±0.06 -1.76 Table 1. Classification accuracy (%) on both adversarial images (Rob.) under 10-step PGD attack at ϵa= 1/255 and clean images (Acc.) across 16 datasets. We include the results on TinyImageNet because it is used to finetune the model for CLIP-FT, TeCoA [32], PMG-AFT [50], and FARE [44]. Comparison is made among our paradigm and test-time defences adapted from existing adversarial studies, with finetuning-based models implemented as a reference. We report the mean and standard deviation for test-time methods over 3 runs. The last column reports the gains w.r.t. original CLIP without any finetuning or test-time operations. [32,50], we test all baselines under 10-step PGD attacks across 16 datasets, assuming that the attacker has full access to the weights and gradients of the deployed model, but not to the test-time operations made by the end user. We report the accuracy on both adversarial images and clean images in Tab. 1. It can be seen that all finetuning-based methods overfit to the dataset used for adversarial finetuning to varying extents, as evidenced by the higher accuracy of clean images than the original CLIP on TinyImageNet. The improved robustness on downstream datasets comes at a cost of a noticeable clean accuracy drop. Among test-time methods, both Anti-adversary and HD, which generate an additive perturbation based on an objective, lead to limited improvement of robust accuracy. Our TTC, which utilises the pre-trained vision encoder of CLIP to produce counterattacks, shows the best robust accuracy on most downstream datasets, usually with a large gain. We also retain the best clean accuracy compared to these two perturbation update methods. Adding random noise (RN) brings little robustness, even though the added noise is four times larger than the attack budget, i.e., ϵttc ≫ϵa. RN can be viewed as a special case of TTC with the number of Nbeing 0. By exploiting the pre-trained model fθto optimize the noise, TTC significantly improves robustness. TTE ensembles a number of image transformations, which improves CLIP’s robustness at test time to an average accuracy of 33.28%. 15177 (%) Rob. Acc. CLIP 0.09 61.51 CLIP-FT 0.96 55.80 TeCoA1[32] 6.51 40.25 TeCoA4[32] 10.03 35.57 PMG-AFT1[50] 7.03 42.30 PMG-AFT4[50] 10.70 37.58 FARE1[44] 1.50 51.02 FARE4[44] 3.67 46.17 RN 0.06±0.00 61.61±0.03 TTE [38] 7.79±3.23 61.79±0.13 Anti-adv [1] 0.53±0.00 57.32±0.03 HD [52] 1.19±0.01 56.62±0.02 TTC (ours) 20.63±0.05 55.99±0.06 ∆+20.54 -5.52 Table 2. Classification accuracy (%) on adversarial images (Rob.) under 10-step PGD at ϵa= 4/255 and clean images (Acc.) averaged on 16 datasets. Superscripts indicate the attack budget used in the finetuning phase. The last row reports the gains compared to the original CLIP. However, this gain is generally unstable across runs, as indicated by the high standard deviation of robust accuracy. Overall, our proposed TTC leads to consistent gains on robust accuracy (+36.47%) averaged on downstream datasets with a slight loss (-1.76%) on clean accuracy compared to the original CLIP, serving as a stable defence at inference time. We test the robustness under CW attacks [6] in Appendix (Sec. 8.1) for limited space. It can also be seen that the robustness gains come at a cost of accuracy reduction on clean images to varying extents across datasets. We provide more analysis in Appendix Sec. 9. Robustness under ϵa= 4/255.We further test the robustness of all methods under a high attack budget ϵa= 4/255. For this setting, we increase the number of steps Nto 5 for more effective counterattacks, while other hyperparameters are unchanged. We also implement finetuning-based methods with attack budget ϵ= 4/255 during finetuning, in this setting. We report the average accuracy across 16 datasets in Tab. 2and provide the full table in Appendix (Tab. 5). It can be seen that a high attack budget at ϵa= 4/255 degrades the accuracy of all models to a very low level. Anti-Adversary [1] and HD [52] provide little to no robustness under this setting. TTE defends the model to a limited extent, but still with low reliability as indicated by the high standard deviation. In comparison, our proposed TTC provides a stable robustness gain averaged on 16 datasets. 4.3. TTC on Adversarially Finetuned CLIP Since our method performs counterattacks using the victim model at test time, it can also be employed on adversarially finetuned models in a plug-in manner. In this section, we apply TTC to finetuning-based models, assuming that the attacker has full access to the deployed model, but not to the operations by the end user. Note that we still employ the original vision encoder fθof CLIP to compute τ (Eq. (4)), because the sensitivity of adversarial finetuned vision encoders is largely reduced. We report the results in Tab. 3. It can be seen that TTC can further boost adversarial robustness by exploiting the finetuned model to perform counterattacks at test time. Specifically, TTC achieves a robustness accuracy of 29.06% and 30.81% when employed on TeCoA and PMG-AFT, surpassing the original finetuned models by 2.52 and 2.05 points, respectively. A significant gain of 13.85 points is achieved when we employ TTC on top of FARE, an unsupervised adversarially finetuned model. Interestingly, we find that adversarial finetuning greatly reduces the sensitivity of CLIP to variations in the pixel space, thus hurting the expressive power of the pre-trained encoder. We provide more in-depth analyses of such loss in Appendix (Sec. 10). Since our counterattacks rely heavily on the expressiveness of the pre-trained vision encoder fθ, this also explains the smaller gains achieved on adversarially finetuned models, compared to the original CLIP. The larger increase of robust accuracy on FARE implies that adversarially finetuning CLIP in an unsupervised manner better retains the expressiveness of the model. 4.4. Ablation studies We experimentally find that the number of steps Nof our TTC greatly affects performance on both adversarial and clean images. A general rule of thumb is that an attack with a higher budget ϵawould require more steps of counterattacks. In this section, we investigate the effect of Nand keep the other hyperparameters unchanged. We provide analysis on other hyperparameters in Appendix (Sec. 11). Fig. 5provides the performance of CLIP employing TTC on 14 datasets as Nvaries. For smaller attacks at ϵa= 1/255, it takes fewer than three steps for CLIP to defend itself effectively on most datasets. Excessive counterattacks can impair the images, as evidenced by the decline after a certain number of steps. In comparison, a strong attack ϵa= 4/255 requires a larger number of counterattack steps to reach a reasonable accuracy, showing that they are more resilient to counterattacks by the user side. TTC does not impact accuracy on clean images significantly on most datasets, except for SUN397 (Fig. 5d), OxfordPets (Fig. 5j), StanfordCars (Fig. 5n) and ImageNet (Fig. 5l), where clean images are found sensitive to the increase of N. 5. Limitations Although we show TTC improves robustness of CLIP to adversary that maximises the classification loss, there are limitations as discussed below. Firstly, the robustness gain of applying TTC on TeCoA [32] and PMG-AFT [50] is less obvious. This is due to the reduced expressiveness of CLIP caused by adversarial finetuning. We argue that for large pre-trained models like CLIP, adversarial finetuning should 15178 (%) CIFAR10 CIFAR100 STL10 ImageNet Caltech101 Caltech256 OxfordPets Flower102 FGVCAircraft StanfordCars SUN397 Country211 Food101 EuroSAT DTD PCAM Avg. Rob. Avg. Acc. TeCoA 33.61 18.95 70.08 18.89 55.51 43.19 38.35 21.94 2.49 8.76 19.39 1.78 13.90 11.96 17.61 48.24 26.54 40.25 TeCoA+TTC 34.68 20.00 71.65 23.14 59.44 48.49 42.66 25.13 2.78 12.09 23.91 2.49 17.79 12.75 18.87 48.44 29.02 39.85 ∆1.07 ↑1.05 ↑1.57 ↑4.25 ↑3.93 ↑5.30 ↑4.31 ↑3.19 ↑0.29 ↑3.33 ↑4.52 ↑0.71 ↑3.89 ↑0.79 ↑1.26 ↑0.20 ↑2.48 ↑−0.40 ↓ PMG-AFT 40.66 22.52 73.08 21.43 61.08 45.91 41.18 23.43 2.22 11.65 22.58 2.12 18.57 12.60 14.95 46.18 28.76 42.30 PMG-AFT+TTC 42.17 24.09 73.55 24.36 63.72 50.37 43.96 25.94 2.51 14.97 25.70 2.57 22.33 13.94 15.98 46.52 30.79 41.89 ∆1.51 ↑1.57 ↑0.47 ↑2.93 ↑2.64 ↑4.46 ↑2.78 ↑2.51 ↑0.29 ↑3.32 ↑3.12 ↑0.45 ↑3.76 ↑1.34 ↑1.03 ↑0.34 ↑2.03 ↑−0.41 ↓ FARE 19.65 11.40 59.06 14.00 50.74 38.79 31.07 17.14 1.35 6.85 14.90 0.85 11.65 10.67 15.64 16.23 20.00 51.02 FARE+TTC 35.55 22.34 76.65 30.52 67.39 59.20 51.53 29.85 5.03 20.46 33.42 4.04 31.76 15.49 23.17 35.79 33.89 49.91 ∆15.90 ↑10.94 ↑17.59 ↑16.52 ↑16.65 ↑20.41 ↑20.46 ↑12.71 ↑3.68 ↑13.61 ↑18.52 ↑3.19 ↑20.11 ↑4.82 ↑7.53 ↑19.56 ↑13.89 ↑−1.11 ↓ Table 3. TTC employed on adversarially finetuned models at test time. We report the robust accuracy at ϵa= 1/255 and the robustness gain by employing TTC for each dataset. (a) CIFAR10 (b) DTD (c) STL10 (d) SUN397 (e) EuroSAT (f) Caltech101 (g) PCAM (h) CIFAR100 (i) Food101 (j) OxfordPets (k) Flower102 (l) ImageNet (m) Caltech256 (n) StanfordCars Figure 5. Effects of the number of steps Nfor counterattacks performed on CLIP. The green lines represent accuracy on clean images, and red and blue lines accuracy on adversarial images at ϵa= 1/255 and ϵa= 4/255, respectively. be employed sparingly, considering that a fundamental difference from adversarial studies on non-foundational models is that they have learned massive amounts of real-world knowledge. Secondly, although TTC does not involve training on adversarial images, it incurs more computation expenses at inference time. Additionally, the number of counterattack steps affects robustness performance. It can be difficult to tune for the most suitable N, if the attack strength ϵais not known a priori (Fig. 5). We recommend fewer steps (no more than three) if the attack is unknown to avoid excessive counterattacks and unproductive computational overhead. In the future, we intend to explore methods to adjust the number of steps based on the test image. Thirdly, according to adversarial robustness studies on conventional models, test-time defence can be circumvented by adaptive attacks [12]. We discuss in Appendix (Sec. 12) possible adaptive attacks to break our counterattacks assuming the worst scenario where the attacker has access to the weights of the deployed CLIP model and TTC performed by the end user. 6. Conclusion We show that CLIP can leverage its own pre-trained vision encoder to defend against adversary maliciously manipulated to maximise its loss by performing counterattacks at test time, without relying on any auxiliary networks. Based on the finding that adversarial images are ‘falsely stable’, we propose τ-thresholded counterattacks to guide the adversarial image away from its original embedding in the latent space. Experiments on 16 datasets show that TTC employed on CLIP achieves stable and promising accuracy on adversarial images. TTC is also shown to further enhance robustness of adversarially finetuned CLIP models. We also find that finetuning CLIP with adversarial images compromises its own expressiveness, and recommend cautious use of adversarial finetuning as the only approach to robustifying large pre-trained models. Our paradigm is the first test-time method to defend CLIP at inference time without any finetuning. We hope this study will encourage future research of robustifying approaches for CLIP alternative to adversarial finetuning. Acknowledgement This work was supported by the MUR PNRR project FAIR (PE00000013) funded by the NextGenerationEU and the EU Horizon projects ELIAS (No. 101120237) and AI4Trust (No. 101070190). 15179 References [1] Motasem Alfarra, Juan C P´ erez, Ali Thabet, Adel Bibi, Philip HS Torr, and Bernard Ghanem. Combating adversaries with anti-adversaries. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 5992–6000, 2022. 2,5,7 [2] Anish Athalye, Nicholas Carlini, and David Wagner. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In International conference on machine learning, pages 274–283. PMLR, 2018. 1 [3] Tao Bai, Jinqi Luo, Jun Zhao, Bihan Wen, and Qian Wang. Recent advances in adversarial training for adversarial robustness. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, pages 4312–4321. International Joint Conferences on Artificial Intelligence Organization, 2021. Survey Track. 1 [4] Babak Ehteshami Bejnordi, Mitko Veta, Paul Johannes Van Diest, Bram Van Ginneken, Nico Karssemeijer, Geert Litjens, Jeroen AWM Van Der Laak, Meyke Hermsen, Quirine F Manson, Maschenka Balkenhol, et al. Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer. Jama, 318(22):2199–2210, 2017. 5 [5] Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101–mining discriminative components with random forests. In Computer vision–ECCV 2014: 13th European conference, zurich, Switzerland, September 6-12, 2014, proceedings, part VI 13, pages 446–461. Springer, 2014. 5 [6] Nicholas Carlini and David Wagner. Towards evaluating the robustness of neural networks. In 2017 ieee symposium on security and privacy (sp), pages 39–57. Ieee, 2017. 1,2,3, 4,7 [7] Mathilde Caron, Hugo Touvron, Ishan Misra, Herv´ e J´ egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021. 1 [8] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020. 1 [9] Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3606–3613, 2014. 5 [10] Adam Coates, Andrew Ng, and Honglak Lee. An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pages 215–223. JMLR Workshop and Conference Proceedings, 2011. 5 [11] Francesco Croce and Matthias Hein. Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In International conference on machine learning, pages 2206–2216. PMLR, 2020. 1 [12] Francesco Croce, Sven Gowal, Thomas Brunner, Evan Shelhamer, Matthias Hein, and Taylan Cemgil. Evaluating the adversarial robustness of adaptive test-time defenses. In International Conference on Machine Learning, pages 4421– 4435. PMLR, 2022. 2,8 [13] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 5 [14] Li Fei-Fei, Robert Fergus, and Pietro Perona. One-shot learning of object categories. IEEE transactions on pattern analysis and machine intelligence, 28(4):594–611, 2006. 5 [15] Gregory Griffin, Alex Holub, Pietro Perona, et al. Caltech256 object category dataset. Technical report, Technical Report 7694, California Institute of Technology Pasadena, 2007. 5 [16] Jean-Bastien Grill, Florian Strub, Florent Altch´ e, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. Advances in neural information processing systems, 33:21271–21284, 2020. 1 [17] Chuan Guo, Mayank Rana, Moustapha Cisse, and Laurens van der Maaten. Countering adversarial images using input transformations. In International Conference on Learning Representations, 2018. 2,5 [18] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 2 [19] Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(7):2217–2226, 2019. 5 [20] Duhun Hwang, Eunjung Lee, and Wonjong Rhee. Aidpurifier: A light auxiliary network for boosting adversarial defense. Neurocomputing, 541:126251, 2023. 2 [21] Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning, pages 4904–4916. PMLR, 2021. 1 [22] Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In Proceedings of the IEEE international conference on computer vision workshops, pages 554–561, 2013. 5 [23] Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 5 [24] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012. 2 [25] Alexey Kurakin, Ian J Goodfellow, and Samy Bengio. Adversarial examples in the physical world. In Artificial intelligence safety and security, pages 99–112. Chapman and Hall/CRC, 2018. 1 [26] Lin Li, Haoyan Guan, Jianing Qiu, and Michael Spratling. One prompt word is enough to boost adversarial robustness 15180