scieee AI-readable full text Open interactive document viewer

Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning

Zhang, Wanchang,Dimiccoli, Mariella,Lim, B.H.

Abstract

Model explanations such as saliency maps can improve user trust in AI by highlighting important features for a prediction. However, these become distorted and misleading when explaining predictions of images that are subject to systematic error (bias). Furthermore, the distortions persist despite model fine-tuning on images biased by different factors (blur, color temperature, day/night). We present Debiased-CAM to recover explanation faithfulness across various bias types and levels by training a multi-input, multi-task model with auxiliary tasks for explanation and bias level predictions. In simulation studies, the approach not only enhanced prediction accuracy, but also generated highly faithful explanations about these predictions as if the images were unbiased. In user studies, debiased explanations improved user task performance, perceived truthfulness and perceived helpfulness. Debiased training can provide a versatile platform for robust performance and explanation faithfulness for a wide range of applications with data biases.

Full text

Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning Wencan Zhang National University of Singapore Singapore [email protected] Mariella Dimiccoli Institut de Robòtica i Informàtica Industrial, CSIC-UPC Barcelona, Spain [email protected] Brian Y. Lim National University of Singapore Singapore [email protected] ABSTRACT Model explanations such as saliency maps can improve user trust in AI by highlighting important features for a prediction. However, these become distorted and misleading when explaining predictions of images that are subject to systematic error (bias) by perturbations and corruptions. Furthermore, the distortions persist despite model fine-tuning on images biased by different factors (blur, color temperature, day/night). We present Debiased-CAM to recover explanation faithfulness across various bias types and levels by training a multi-input, multi-task model with auxiliary tasks for explanation and bias level predictions. In simulation studies, the approach not only enhanced prediction accuracy, but also generated highly faithful explanations about these predictions as if the images were unbiased. In user studies, debiased explanations improved user task performance, perceived truthfulness and perceived helpfulness. Debiased training can provide a versatile platform for robust performance and explanation faithfulness for a wide range of applications with data biases. CCS CONCEPTS •Human-centered computing →Empirical studies in HCI ; • Computing methodologies →Computer vision ;Semi-supervised learning settings;•Security and privacy →Privacy protections. KEYWORDS Explainable AI, Misleading explanations, Class activation map, Robust machine learning, Image perturbations, User studies ACM Reference Format: Wencan Zhang, Mariella Dimiccoli, and Brian Y. Lim. 2022. Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning. In CHI Conference on Human Factors in Computing Systems (CHI ’22), April 29-May 5, 2022, New Orleans, LA, USA. ACM, New York, NY, USA, 32 pages. https://doi.org/10.1145/3491102.3517522 1 INTRODUCTION Machine learning models are increasingly capable to achieve impressive performance in many prediction tasks, such as image recognition [ 49 ], medical image diagnosis [ 29 ], captioning [ 86 ] and dialog Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA ©2022 Copyright held by the owner/author(s). ACM ISBN 978-1-4503-9157-3/22/04. https://doi.org/10.1145/3491102.3517522 systems [ 16 ]. Despite their superior performance, deep learning models are complex and unintelligible; this limits user trust and understanding [ 60 , 63 , 87 ]. This has driven the development of myriad explainable artificial intelligence (XAI) and interpretable machine learning methods [ 33 , 38 , 87 , 95 ]. Saliency maps [ 78 , 82 , 102 ] can provide intuitive explanations of Convolutional Neural Networks (CNN) for image prediction tasks by indicating which pixels or neurons were used for model inference. Amongst these, class activation map (CAM) [ 102 ], Grad-CAM [ 78 ] and extensions [ 13 , 89 ] are particularly useful by identifying pixels relevant to specific class labels. Users can verify the prediction correctness by checking whether expected pixels are highlighted. Models would be considered more trustworthy if their CAMs matched what users believe as salient. Despite the fidelity of CAMs on clean images, real-world images are typically subjected to systematic error, such as image blurring, color-distortion or lighting changes, which can affect what CAMs highlight. We call this systematic error bias 1 since it is directional based on a contextual factor or confound, and contrast it with noise that is based on non-directional random error. Also note that we are not referring to societal bias or discrimination (e.g., racism, sexism) [ 22 ]. Blurring can be due to accidental motion [ 50 ] or defocus blur [ 85 ], or deliberate obfuscation for privacy protection [ 21 ]. Unlike [ 100 ] which found explanation harms privacy, we find that privacy can harm explanation. Images may also be biased with shifted color temperature [ 4 ] due to mis-set white balance, or biased with daylight changes (e.g., day to night, sunrise/sunset). These biases decrease model prediction performance [ 4 , 21 , 85 ] and we further show that they also lead to deviated or distorted CAM explanations that are less faithful to the original scenes. For different bias types (image blur, and color temperature shift, day/night lighting), we found that CAMs deviated more as image bias increased (Fig. 1 and Fig. 3: Biased-CAMs from RegularCNN for σ > 0). Although Biased-CAM represents what the CNN considers important in a biased image, it is misaligned with people’s expectations [ 68 ], misleads users to irrelevant targets, and impedes human verification and trust [ 26 ] of the model prediction. For example, when explaining the inference of the “Fish” label for an image prediction, Biased-CAMs select pixels of the man instead of the fish (Fig. 1). To align with user expectations, models should not only have the right predictions but also have the right reasons [ 73 ]; however, current approaches face challenges in achieving this goal, particularly for biased data. First, while fine-tuning the model on biased data 1 Note that this does not refer to social bias that is presently popularly studied in AI fairness and algorithmic bias. We are using the word as defined in engineering and physics regarding measurements. CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA Wencan Zhang, Mariella Dimiccoli, and Brian Y. Lim Blur Bias Level (𝜎) 0 8 16 24 32 Biased Image RegularCNN Confidence: 1.000 CAM PCC: 1.000 Confidence: 0.988 CAM PCC: 0.814 Confidence: 0.390 CAM PCC: –0.314 Confidence: 0.044 CAM PCC: –0.199 Confidence: 0.053 CAM PCC: 0.420 FineTunedCNN Confidence: 1.000 CAM PCC: 1.000 Confidence: 0.973 CAM PCC: 0.826 Confidence: 0.944 CAM PCC: 0.167 Confidence: 0.562 CAM PCC: –0.330 Confidence: 0.583 CAM PCC: 0.130 DebiasedCNN Confidence: 0.988 CAM PCC: 0.981 Confidence: 0.983 CAM PCC: 0.632 Confidence: 0.891 CAM PCC: –0.102 Confidence: 0.646 CAM PCC: –0.257 Confidence: 0.476 CAM PCC: 0.013 Unbiased-CAM at 𝜎= 0 Biased-CAM of RegularCNN at 𝜎 = 16 Biased-CAM of FineTunedCNN at 𝜎 = 16 Debiased-CAM of DebiasedCNN at 𝜎 = 16 a b 1 0 Saliency RegularCNN (top predicted class label) Confidence: 1.000 CAM PCC: 1.000 Confidence: 0.988 CAM PCC: 0.814 Confidence: 0.390 CAM PCC: –0.314 Confidence: 0.731 CAM PCC: 0.126 Predicted: Dog Confidence: 0.642 CAM PCC: 0.088 Predicted: Dog c Figure 1: Deviated and debiased CAM explanations for prediction label "Fish". a) Debiased-CAMs (from DebiasedCNN) were most faithful to the Unbiased-CAM (from RegularCNN at σ= 0 ) as blur bias increased. In contrast, Biased-CAMs from RegularCNN and FineTunedCNN became very deviated with a much lower CAM Pearson Correlation Coefficient (PCC). The wrong CAMs can mislead users to think the predictions were wrong even if they were correct. b) Debiased-CAM selected similar important pixels of the Fish as Unbiased-CAM, while Biased-CAMs selected irrelevant pixels of the person or background instead. c) CAM of the top predicted class label with only RegularCNN at σ=32 predicting the wrong prediction label “Dog”. can improve its performance [ 21 , 85 ], this does not necessarily produce explanations aligned with human’s understanding. Indeed, we found that explanations remain deviated and unfaithful (Fig. 1, Fine- TunedCNN Biased-CAMs). Conversely, retraining the model with attention transfer [ 47 , 57 ] only improves explanation faithfulness for clean images, but cannot handle biased images. Finally, evaluating the human interpretability of explanations requires deep inquiry into user perception, understanding and usage [ 1 , 5 , 25 ], but typical evaluations of XAI involve only data simulations [ 8 , 30 , 73 , 83 , 94 ] or simple surveys [10, 13, 77, 78, 103]. Inspired by how people can “see through the blur” to recognize blurred images due to prior experiences with unblurred but unrelated images, we propose a debiasing approach such that models are trained to faithfully explain the event despite biased sources. Using CNNs with Grad-CAM saliency map explanations [ 78 ], we developed DebiasedCNN that interprets biased images as if predicting on the unbiased form of images and produces explanations, Debiased-CAMs, that are more human-relatable and robust. The approach has a modular design: 1) it is self-supervised which does not require additional human annotation for training; 2) it produces explanations as a secondary prediction task, so that they are retraininable to be debiased; 3) it models the bias level as a tertiary task to support bias-aware predictions. The approach not only enhances prediction performance on biased data, but also produces highly faithful explanations about these predictions as if the data were unbiased (Fig. 1: DebiasedCNN CAMs). To evaluate the developed model, we conducted simulation and user studies to address the research questions on 1) how bias decreases explanation faithfulness and how well debiasing mitigates this, and 2) how sensitive people are to perceiving explanation deviations and how well debiasing improves perceived explanation truthfulness and helpfulness. For generality, the simulation studies spanned different image prediction tasks object recognition, activity recognition with egocentric cameras, image captioning, and scene understanding), bias types (blur, color shift, and night vision interpolation) and various datasets. Across all studies, we found that while increasing bias led to poorer prediction performance and worse explanation deviation, Debiased-CAM showed the best improvement in task performance as well as explanation faithfulness. Instead of trading off task performance for explanation faithfulness, Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA our debiasing training improved both. We further demonstrated the usability and usefulness of Debiased-CAMs in two controlled user studies. Quantitative statistical and qualitative thematic analyses validated that users can perceive the improved truthfulness and helpfulness of Debiased-CAMs on biased images. In summary, this paper made the following contributions: (1) Assessed the deviations in model explanations due to bias in data across different bias types and levels. (2) Proposed a technical approach to accurately predict and faithfully explain inferences under data bias. (3) Validated the improvements in perceived truthfulness and helpfulness of debiased explanations. 2 RELATED WORK We review explainable AI methods for image predictions, how images get biased, how misleading explanations harm user experience and performance, and methods to improve explanation faithfulness. 2.1 Explainable AI for visual CNN models Many explainable AI (XAI) techniques have been proposed to understand the predictions of CNNs. These include saliency maps [ 13 , 70 , 78 , 81 , 102 ], feature visualization [ 10 , 39 , 66 ], neuron activations [ 41 ] and concept variables [ 43 , 46 ]. Saliency maps are intuitive to interpret deep CNN models, where important pixels are highlighted to indicate their importance towards the model prediction. Computing the prediction gradient [ 81 , 83 , 93 ] can identify sensitive pixels. Another approach divides prediction outcome across features by Taylor series approximation [ 8 ] or Shapley values [ 61 ]. Specific to CNNs, coarser saliency maps can be generated by aggregating activation maps as a weighted sum across convolutional kernels [ 13 , 70 , 78 , 89 , 102 ]. For this work, we evaluated Grad-CAM [ 78 ] to test if users can perceive truthful, biased, or debiased explanations, and expect our findings to be generalizable. 2.2 Systematic error and corruptions in images Although many models are trained on clean curated images, realworld images are subject to systematic errors (biases), perturbations and corruptions. Contextual or incidental biases include blurring, color distortions, or lighting changes. Blurring may be due to accidental motion blur [ 50 ], defocus blur [ 85 ] or deliberate obfuscation for privacy protection [ 21 ]. Image color shift [ 4 ] may be due to misset white balance. These biases can degrade model performance [ 4 , 21 , 85 ], and limit their usefulness in real-world applications. Images of outdoor scenes regularly change by time of day and seasons due to sunlight or weather changes [ 51 ]. Images can also be corrupted due to data processing, such as JPEG compression artifacts, Gaussian noise, brightness or contrast levels [ 36 ]. Mitigation strategies to handle such data errors include model fine-tuning with images at known blurred levels [ 21 ], or data augmentation with images blurred at multiple levels [ 85 ]. However, these approaches only aimed to improve prediction performance and not explanation faithfulness. In this work, we found that explanations remain deviated and we propose methods to debias them. Such data errors are related to the problem of model robustness, where small changes to data should not cause large changes in model behavior. This is an active area of research [ 36 , 37 , 101 ], but methods typically focus on improving performance by increasing decision boundary smoothness. In this work, we aim to make explanations more robust. Recent work by Dombrowski et al. [ 24 ] improved explanation robustness by similarly increasing decision smoothness relative to explanations, but this assumes clean data, and learns average explanations under bias. Instead, we debias explanations away from deviations due to biased data. Also, other than focusing on explanation robustness or stability towards the impression of global trustworthiness, we focus on faithful explanations that are verifiable per instance. 2.3 Risk of misleading model explanations User studies of model explanations aim to show that explanations can improve user understanding and trust [ 45 , 59 , 65 , 90 , 98 ]. These tend to study scenarios of correct model predictions and ideal explanations, but models can make prediction errors or may not be confident in their decisions. Studies have explored how this may lead to distrust, mistrust and over-trust [ 58 , 69 , 92 ]. For such cases, explanations can be avoided when there is a high chance of model error. However, explanations can still be wrong despite the model predicting correctly. For example, explanations may highlight spurious pixels [ 97 ], be adversarially manipulated [ 23 , 31 ], or subject to input error [ 88 ]. These cases are harder to detect, pose a serious risk to decrease user trust [ 18 , 54 ], or mislead users [ 54 ]. Unlike works that explore how different explanation formats affect trust [ 91 ], we investigate how slight data variations affect user performance and trust. Since data bias and corruption are prevalent in the real-world, it is tantamount to identify the severity of the problem and mitigate it with more robust explanations [ 35 ]. In this work, we quantify the extent of explanation deviation due to data bias, and evaluated how sensitive users are to these deviations. 2.4 Attention transfer to correct explanations While explanation techniques are primarily designed to improve human understanding of model behavior, they can be used to guide model training. One approach is to use transfer learning to regularize attention from a better model to the model under training, such as with student-teacher networks [ 47 ]. Another approach indirectly trains attention by ablating salient pixels from input images and maximizing the classification loss between the ablated and original images [ 57 ]. However, these approaches only train on clean data and will reinforce biased explanations if trained on obfuscated or biased data. Unlike conventional self-supervised learning with data augmentation and contrastive learning to improve feature learning [ 14 ], we use the unbiased explanation as a surrogate "label" to train the debiased model to predict a more faithful explanation. 3 TECHNICAL APPROACH We first describe baseline RegularCNN and FineTunedCNN approaches to predict on unbiased and biased image data, then our proposed DebiasedCNN architectures to predict on biased image data with debiased explanations. 3.1 Regular and Fine-tuned Models A regularly trained CNN model (RegularCNN) can generate a truthful CAM e M M M (Unbiased-CAM) of an unbiased image x , but will CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA Wencan Zhang, Mariella Dimiccoli, and Brian Y. Lim c 𝒙𝐶𝑁𝑁!𝑦% 𝑴 ' 𝒚 ) 𝒙" Bias 𝔅(𝒙, 𝑏) 𝑴 . 𝒚 𝐿# 𝐿" 𝑏 0 𝑏 Unbiased-CAM Debiased-CAM 𝐿$ 𝐶𝑁𝑁% 𝐿$ 𝜔&𝜔' 𝜔( LSTM LSTM LSTM 𝜎 Caption 𝒚 ))<S> 𝜎 𝜔)* 𝜎 <E> … LSTM 𝜎 33𝜔)& 𝒆&𝒆' 𝒆( Caption Task FC Layer b Grad-CAM Pool GAP2D Pool GAP2D Pool GAP2D Pool Input 𝒙" Max Pool/2 Convfinal Max Pool/2 Convfinal GAP2D Pool Shared FC Block 𝜶+ × ReLU Convlast Conv Blocks Label 𝒚 ) Label Task FC Layer Bias Level FC Block Bias Level 𝑏 0 a 𝑨 Grad-CAM 𝑴 ' Label 𝒆+ GAP2D Pool GAP2D Pool 𝜕𝐲!𝜕𝑨" ⁄ where 𝐲!= 𝒚 $∘ 𝒆! 𝑨, Figure 2: Architecture of DebiasedCNN. a) DebiasedCNN is a multi-input, multi-task convolutional neural network with two inputs image xband label ecfor class c, and three tasks for primary prediction task b y, CAM explanation task b M M M, and bias level prediction task b b. b) DebiasedCNN can be trained for different primary tasks, such as image captioning. c) Meta-architecture with self-supervised learning to minimize the CAM loss LMbetween Unbiased-CAM e M M Mfrom RegularCNN (CN N0) predicting on unbiased image xand Debiased-CAM b M M Mfrom DebiasedCNN (CN Nd) predicting on biased image xbat bias level b. produce a deviated CAM ˇ M M M (Biased-CAM) for the image under bias xb , i.e., e M M M(x),b M M M(xb) , due to the model not training on any biased images and learning spurious correlations with blurred pixels. A fine-tuned model trained on biased images can improve the prediction performance on biased images, but will still generate a deviated CAM ˇ M M M (Fig. 1a and Fig. 3a-c: CAMs of FineTunedCNN), as it was only trained with the classification loss and not explanation loss. While these models can be explained with Grad-CAM, they are not retrainable to improve their CAM faithfulness. 3.2 DebiasedCNN Model with Debiased-CAM Explanations 3.2.1 Trainable CAM as secondary prediction task. We enable CAM retraining by redefining Grad-CAM as a prediction task. Grad-CAM [ 78 ] computes a saliency map explanation of an image prediction with regards to class c as the weighted sum of activation maps in the final convolutional layer of a CNN. Each activation map Ak indicates the activation Ak ij for each grid cell (i,j) of the k th convolution filter k∈K (set of all filters). The importance weight αc k for the k th activation map is calculated by back-propagating gradients from the output ˆ yto the convolution filter, i.e., αc k=1 HW H Õ i=1 W Õ j=1 ∂ˆ yc ∂Ak ij ≡GAPij ∂ˆ yc ∂Ak(1) where H and W are the height and width of activation maps, respectively; yc is a one-hot vector indicating only the probability of class c ; GAPij (·) is the global average pooling operation. The class activation map (CAM) is the weighted combination of activation maps, followed by a ReLU transform to only show positive activations for class c, i.e., Mc=ReLU(Õ k αc kAk) ≡ b M M M=ReLU αcAT(2) which we rewrite as a matrix multiplication of all K importance weights αc=nαc koK and the transpose of activation maps A along the kth axis, i.e., AT=nAk ij oK×H×W. Therefore, the CAM prediction task can be redefined as three non-trainable layers (computational graph) in the neural network (orange in Fig. 2a) to compute ∂yc ∂Ak , αc k , and b M M M . By reformulating Grad-CAM as a secondary prediction task, we can train the model with faithful CAM based on differentiable CAM loss by backpropagating through this task. This task takes ecas the second input to the CNN architecture to specify the target class label for the CAM. c is set as the ground truth class label at training time, and chosen by the user at run time. We call the aforementioned approach Multi- Task DebiasedCNN, and call the conventional use of Grad-CAM as Single-Task DebiasedCNN. For single-task DebiasedCNN, the loss is added as a simple sum to the primary classification task, rather than predicted with secondary task. This will limit its learning since weights are not updated with gradient descent. 3.2.2 Training CAM debiasing with Self-Supervised Learning. To debias CAMs b M M M of biased images xb toward truthful Unbiased-CAMs e M M M of clean images x , i.e., b M M M(xb)≈e M M M(x) , we train DebiasedCNN with self-supervised learning to transfer knowledge of corresponding unbiased images in RegularCNN into DebiasedCNN. We aim to minimize the difference between Unbiased-CAM e M M M and Debiased- CAM b M M M . The training involves the following steps (see Fig. 2c): 1) Given a dataset with clean images x∈X and labels y , apply a bias transformation (e.g., blur) to create biased variants of each image xb∈Xb . 2) Train a RegularCNN to predict label ˜ y on clean image x . We assume that its Grad-CAM explanations e M M M are correct and serve as a good oracle for Unbiased-CAMs. 3) Train a DebiasedCNN to predict label ˆ y on corresponding biased image xb , and explain with CAM b M M M. DebiasedCNN is trained with loss function: L=Ly(y,by)+ωMLMe M M M,b M M M(3) where Ly is the classification loss, LM is the CAM loss, and ωM is a hyperparameter. The training can be interpreted as attention transfer from an unbiased model to the new model. DebiasedCNN can be generalized to image prediction other tasks (e.g., image captioning: Fig. 3b), other bias types (e.g., color temperature, lighting: Fig. 3c,d), different base CNN models (e.g., VGG16, Inception v3, ResNet50, Xception), and for privacy-preserving machine learning. Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA 3.2.3 Bias-agnostic, Multi-bias predictions with tertiary task. Image biasing can happen sporadically at run time, so the image bias level b may be unknown at training time. Instead of training on specific bias levels [ 21 ] or fine-tuning with data augmentation on multiple bias levels [ 85 ], we added a tertiary prediction task — bias level regression — to DebiasedCNN to leverage supervised learning (Fig. 2a: salmon-colored layers). This enables DebiasedCNN to be bias-aware (can predict bias level) and bias-agnostic (predict under any bias level). With the bias level prediction task, the training loss function for multi-bias, multi-task DebiasedCNN is: L=Ly(y,by)+ωMLMe M M M,b M M M+ωbLbb,b b(4) where Lbis the bias prediction loss, and ωbis a hyperparameter. 3.2.4 Training loss terms. In all, there are three loss terms: primary task loss Ly as cross-entropy loss for standard classification tasks, and as the sum of negative log likelihood for each word [ 86 ] in image captioning tasks; bias level loss Lb as the mean squared error (MSE), common for regression tasks; CAM loss LM as the mean squared error (MSE), since CAM prediction can be considered a 2D regression task, and this is common for visual attention tasks [ 47 ]. 3.2.5 Summary of DebiasedCNN Model Variants. DebiasedCNN has a modular design: 1) single-task (st) or multi-task (mt) to improve model training; and 2) single-bias (sb) or multi-bias (mb) to support bias-aware and bias-agnostic predictions. We denote the four DebiasedCNN variants as (sb, st), (mb, st), (sb, mt), (mb, mt), and conducted ablation studies to compare between them. Supplementary Fig. 1 and Supplementary Table 2 show details of each variant. 4 SIMULATION STUDIES To evaluate how much CAMs deviate with biased images and how well DebiasedCNN recovers CAM Faithfulness, we conducted five simulation studies with varying datasets, prediction tasks (classification, captioning), bias types (blur, color temperature, day/night lighting), and bias levels. These studies inform which applications explanation biasing is problematic, and show that our debiased training can successfully mitigate these deviations. 4.1 Evaluation Metrics We evaluated prediction performance and CAM explanation faithfulness to compare model variants. For classification, we measured the area under the precision-recall curve (PR AUC) as it is robust against imbalanced data [ 76 ], and calculated the class-weighted macro average to aggregate across multiple classes. For image captioning, we calculated the BLEU-4 [ 67 ] score that measures how closely 4-grams in the predicted and actual captions matched. For bias level regression, we calculated accuracy with R2 . We define the correctness of CAM explanations by their similarity or faithfulness to the original Unbiased-CAMs from RegularCNN that infers on unbiased data. To better compare CAMs beyond simple residual differences (e.g., MAE, MSE), we calculated CAM Faithfulness as the Pearson’s Correlation Coefficient (PCC) [ 11 , 56 ] of pixel-wise saliency as it closely matches the human perception to favor compact locations and match the number of salient locations [ 56 ], and it fairly weights between false positive and false negatives [11]. 4.2 Results In general, CAMs deviate more from Unbiased-CAMs as bias levels increased, but DebiasedCNN reduces this deviation. Debiased retraining also improved model prediction performance, which suggests that DebiasedCNN indeed "sees through the bias". Fig. 4 shows our evaluation Task Performance and CAM Faithfulness in ablation studies across increasing bias levels for different prediction tasks and datasets (Supplementary Table 1). Fig. 3 shows some examples of deviated and debiased CAMs. Next, we describe the experiment method and results for each simulation study. 4.2.1 Simulation Study 1 (Blur Bias). We evaluated CAMs for blur biased images of the object recognition dataset ImageNette [ 40 ]. We scaled images to a standardized maximum size of 1000 × 1000 pixels and applied uniform Gaussian blur at various standard deviations σ . We found that Task Performance and CAM Faithfulness decreased with increasing blur level for all CNNs, but DebiasedCNN mitigated these decreases (Fig. 4a). This indicates that model training with additional CAM loss improved model performance rather than trading-off explainability for performance [ 74 ]. RegularCNN had the worst Task Performance and the lowest CAM Faithfulness for all blur levels ( σ> 8). In comparison, trained with differentiable CAM loss, DebiasedCNN (sb, mt) showed marked improvements to both metrics, up to 2.33x and 6.03x over FineTunedCNN’s improvements, respectively. Trained with non-differentiable CAM loss, DebiasedCNN (sb, st) improved both metrics to a lesser extent than DebiasedCNN (sb, mt), confirming that separating the CAM task from the classification task enabled better weights update . Trained with an additional bias-level task, multi-bias DebiasedCNN (mb, mt) achieved high Task Performance and CAM Faithfulness for all bias levels that is only marginally lower than single-bias DebiasedCNN (sb, mt), because of the former’s good regression performance for bias level prediction (Supplementary Fig. 4). 4.2.2 Simulation Study 2 (Blur Bias, Egocentric). We evaluated the impact of blur biasing with a more ecologically realistic task — wearable camera activity recognition (NTCIR-12 [ 34 ]). This task 2 represents a real-world use case where egocentric cameras may capture blurred images accidentally due to motion or defocus, or deliberately for privacy protection. We found the same trends as for the ImageNette classification task with some differences due to the increased task difficulty (Fig. 4b). In particular, the differences between RegularCNN and DebiasedCNN in Task Performance and CAM Faithfulness were amplified, indicating that debiasing is more useful for this application. Task Performance and CAM Faithfulness decreased steeply for RegularCNN with increasing blur bias, while DebiasedCNN significantly recovered both metrics, demonstrating marginal decreases with increasing bias. FineTunedCNN marginally increased CAM Faithfulness from RegularCNN ( < 44%), while DebiasedCNN achieved a much larger improvement by up to 229%. We verified these trends for different CNN backbones and found that more accurate models produced more faithful CAMs even for stronger blur (Supplementary Figs. 5 and 6). Hence, Debiased-CAM 2 Note that we mean that the use case could have blurred images, not that the NTCIR- 12 dataset has blurred images. Blurring or biasing ImageNet photos (which may include curated stock photos) is an unrealistic use case, but there are more ecologically legitimate reasons for egocentric photos to be biased. CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA Wencan Zhang, Mariella Dimiccoli, and Brian Y. Lim Classification on Tran s At tr Night Bias Ratio (!) 0 0.4 0.8 Captioning on COCO Blur (&) 016 32 Classification on NTCIR-12 Color Tempe rat ure (Δ+) −3600K 0K +3600K True Classification Label: Biking True Classification Label: Cleaning & Chores True Caption: “A man on a horse on a street near people walking.” a b c Prediction Task Classification on NTCIR-12 Bias Type Blur (&) Bias Level 016 32 Biased Image Regular CNN FineTuned CNN Debiased CNN Confidence: 1.000 CAM PCC: 1.000 Confidence: 1.000 CAM PCC: 0.977 Confidence: 0.027 CAM PCC: 0.113 Confidence: 1.000 CAM PCC: 1.000 Confidence: 1.000 CAM PCC: 0.989 Confidence: 0.036 CAM PCC: –0.253 Confidence: 0.997 CAM PCC: 0.987 Confidence: 0.999 CAM PCC: 0.977 Confidence: 0.992 CAM PCC: 0.960 BLEU-4: 0.079 CAM PCC: 1.000 BLEU-4: 0.079 CAM PCC: 0.193 BLEU-4: 0.029 CAM PCC: 0.016 BLEU-4: 0.079 CAM PCC: 1.000 BLEU-4: 0.079 CAM PCC: –0.148 BLEU-4: 0.024 CAM PCC: –0.134 BLEU-4: 0.029 CAM PCC: 0.660 BLEU-4: 0.399 CAM PCC: 0.741 BLEU-4: 0.029 CAM PCC: 0.479 Confidence: 0.830 CAM PCC: 0.843 Confidence: 0.996 CAM PCC: 1.000 Confidence: 0.993 CAM PCC: 0.977 Confidence: 0.835 CAM PCC: 0.713 Confidence: 0.966 CAM PCC: 1.000 Confidence: 0.960 CAM PCC: 0.868 Confidence: 0.998 CAM PCC: 0.904 Confidence: 1.000 CAM PCC: 0.922 Confidence: 0.999 CAM PCC: 0.912 1 0 Saliency d True Labels: Is Snowy, Not Sunny, Not Cloudy, Not Dawn/Dusk Confidence: 0.985 CAM PCC: 1.000 Confidence: 0.994 CAM PCC: 0.967 Confidence: 0.032 CAM PCC: 0.368 Confidence: 0.985 CAM PCC: 1.000 Confidence: 0.997 CAM PCC: 0.729 Confidence: 0.926 CAM PCC: 0.630 Confidence: 0.991 CAM PCC: 0.877 Confidence: 0.988 CAM PCC: 0.778 Confidence: 0.909 CAM PCC: 0.826 Figure 3: Deviated and debiased CAM explanations from models trained on different prediction tasks (a-d) with varying bias levels. In general, RegularCNN and FineTunedCNN had deviated CAMs that missed selecting important pixels, while DebiasedCNN had CAMs similar to Unbiased-CAMs. At no bias, all CAMs from RegularCNN and FineTunedCNN are unbiased. Bias%Ra'o 0.0 0.3 0.6 0.9 Performance%(PR%AUC) 0.4 0.6 0.8 Bias%Ra'o 0.0 0.3 0.6 0.9 CAM%Faithfulness%(PCC) 0.4 0.7 Color%Temperature%Bias -5400 (6600) 5400 Performance%(PR%AUC) 0.85 0.90 0.95 Blur Bias Level (σ) 0 10 20 30 Performance (BLEU-4) 0.10 0.20 0.30 Blur Bias Level (σ) 0 10 20 30 CAM Faithfulness (PCC) 0.1 0.2 Blur Bias Level (σ) 0 10 20 30 Performance (PR AUC) 0.2 0.6 1.0 Blur Bias Level (σ) 0 10 20 30 Performance (PR AUC) 0.6 0.8 1.0 Blur Bias Level (σ) 0 10 20 30 CAM Faithfulness (PCC) 0.4 0.7 1.0 Task = Classification Bias = Blur Data = ImageNette Task = Captioning Bias = Blur Data = COCO Color%Temperature%Bias -5400 (6600) 5400 CAM%Faithfulness%(PCC) 0.70 0.85 1.00 Task = Classification Bias = Color Temperature Data = NTCIR-12 Blur Bias Level (σ) 0 10 20 30 CAM Faithfulness (PCC) 0.4 0.7 1.0 Task = Classification Bias = Blur Data = NTCIR-12 mb = multi-bias sb = single-bias mt = multi-task st = single-task a c d beTask = Classification Bias = Day-to-Night Data = TransAttr Night Bias Ratio (𝜌) Blur Bias Level (σ) 0 10 20 30 Performance (BLEU-4) 0.10 0.20 0.30 Figure 4: Task Performance and CAM Faithfulness for different prediction tasks with increasing bias levels. a)-e) All models’ Task Performance and CAM Faithfulness decreased with increasing blur, while DebiasedCNN decreased the least. For DebiasedCNN variants, multi-task had the highest CAM Faithfulness and Task Performance that is higher than single-task. enables privacy-preserving wearable camera activity recognition with improved performance and faithful explanations. 4.2.3 Simulation Study 3 (Blur Bias Captioning). We evaluated the influence of blur on a different prediction task — image captioning (COCO [ 15 ]). We found similar trends in Task Performance and CAM Faithfulness as before, though all models performed poorly at all blur levels (Fig. 4c). Furthermore, CAM Faithfulness was low for all models, even for RegularCNN at a small blur bias ( σ= 1). This could be because captioning is much harder than classification, and CAM retraining is weakened by vanishing gradients due to the long LSTM recurrence. Yet, DebiasedCNN improved CAM Faithfulness for all blur levels by up to 224% from RegularCNN. Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA 4.2.4 Simulation Study 4 (Color Temperature Biased). We evaluated color temperature bias on wearable camera images in NTCIR-12. This represents another realistic problem for the wearable camera use case, where the white balance may be miscalibrated. We set the neutral color temperature t to 6600K (cloudy/overcast) and perturbed the color temperature bias by applying Charity’s color mapping function to map a temperature to RGB values [ 12 ]. Color temperature can be bidirectionally biased towards warmer (more orange, lower values) or cooler (more blue, higher values) temperatures from neutral 6600K. Furthermore, image pixel values deviate asymmetrically with larger deviations for orange than for blue biases. Consequently, we found that orange bias led to a larger decrease in Task Performance and CAM Faithfulness than blue bias (Fig. 4d). Notably, CAM deviation was smaller across all color temperature biases than for blur biases, as indicated by the smaller decrease in CAM Faithfulness (compare Fig. 4b, d); hence, Task Performance also did not decrease as much as blur bias. FineTuned- CNN had similar Task Performance but lower CAM Faithfulness than RegularCNN; this suggests that color-biased images were too similar to improve model training with classification fine-tuning, and yet this significantly degraded explanation quality. In contrast, DebiasedCNN improved Task Performance and CAM Faithfulness compared to RegularCNN. Furthermore, due to bidirectional bias, multi-bias training enabled DebiasedCNN (mb, mt) to have significantly higher Task Performance even for unbiased images ( ∆tb= 0). 4.2.5 Simulation Study 5 (Lighting Bias). We evaluated lighting bias for outdoor scenes for a multi-label scene attribute recognition task (transient attribute database, TransAttr [ 51 ]). Lighting in outdoor scenes regularly change across hours or seasons due to transient attributes, such as sunlight or weather changes. Hence, models trained on images captured in one lighting condition may predict and explain differently under other conditions. Specifically, for the multi-label prediction task of classifying whether a scene is Snowy, Sunny, Foggy, or Dawn/Dusk, we biased whether the scene was daytime or nighttime. We performed a pixel-wise interpolation with ratio ρ to simulate interstitial periods between day and night (details in Appendix B.1.2). We found similar trends in Task Performance and CAM Faithfulness as with previous blur-biased classification tasks. The image prediction training was biased towards day-time photos, and as photos became darker to represent dusk or night time, all models generated more deviated, but least so for DebiasedCNN. Given the regularity and frequency of outdoor scenes changes, this study demonstrates the prevalence of biasing in model predictions and explanations, and emphasizes the need for Debiased-CAMs. 5 USER STUDIES Having found that DebiasedCNN improves CAM faithfulness, we next evaluated how well Debiased-CAM improves human interpretability over Biased-CAM. We conducted user studies to evaluate their perceived truthfulness (User Study 1) and helpfulness (User Study 2) in an AI verification task for a hypothetical smart camera with privacy blur filters, label predictions and CAM explanations, i.e., the Simulation Study 1 prediction task. Both studies had a 3 × 3 factorial design with two independent variables — Blur Bias level (None σ= 0, Weak σ= 16, Strong σ= 32) and CAM type (Unbiased, Debiased, and Biased). Unbiased-CAM is the CAM from RegularCNN predicting on the unbiased image regardless of blur bias level; Debiased-CAM is the CAM from DebiasedCNN (mb, mt) and Biased-CAM is the CAM from RegularCNN predicting on the biased image at corresponding Blur Bias levels. At the None blur level, Biased-CAM is identical to Unbiased-CAM. The user studies were approved by our university Institutional Review Board. 5.1 User Study 1 (CAM Truthfulness) The first study evaluated the perceived truthfulness of Unbiased, Debiased, and Biased CAMs. 5.1.1 Experiment Procedure. Participants: 1) read the introduction and gave consent; 2) studied a tutorial about automatic image labeling, privacy blurring, heatmap explanations, and how to interpret the survey questions; 3) answered four screening questions to test their labeling of an unblurred and a weakly blurred image and their selection of important locations in an image and a CAM; 4) if screening was passed (all correct answers), answered background questions on technology savviness and image comprehension, performed the main study with 10 trials; and ended with demographic questions. See Supplementary Figs. 11-13 for questionnaire details. In the main study (Fig. 5a), each participant viewed 10 repeated image trials, where each trial was randomly assigned to one of the three Blur Bias levels (within-subjects). All participants viewed the same 10 images (selection criteria described in Appendix C.1) in random order. For each trial, the participant: viewed a labeled unblurred image, indicated the most important locations on the image regarding the label with a “grid selection” UI (q1); and in the next page, viewed the blurred image, viewed CAMs of all 3 types generated from that and arranged randomly side-by-side, rated how well each CAM represented the image label on a 10-star scale (q2), and wrote her rating rationale (q3). 5.1.2 Experiment Apparatus and Measures. We used a “grid selection” user interface (UI) to measure objective truthfulness (Fig. 5b) to mitigate poor estimation of perceptions [ 6 , 7 , 32 ]. It overlays a clickable grid on the image for selecting important cells regarding the label. For usability, we limited the grid to 5 × 5 cells that can be selected or unselected (binary values). In the surveys, we referred to CAMs as “heatmaps”, which is a more familiar term. To compare the participant’s grid selection (User-CAM) with the heatmap shown (CAM), we aggregated CAM by averaging the pixel saliency in each cell and calculated CAM Truthfulness Selection Similarity as the Pearson’s Correlation Coefficient (PCC) between User-CAM and CAM. We also measured the CAM Truthfulness Rating as a subjective, self-reported rating on a uni-polar 10-point star scale (1 to 10). We collected the rationale of ratings as open-ended text. We measured the task time (per trial) as Task Time Level as low ( < 33 percentile), high ( > 66), medium, to account for response thoughtfulness. We tracked the Image Label of each image, since some types are easier to recognize even if blurred. 5.1.3 Participants. We recruited 36 participants from Amazon Mechanical Turk (AMT) with high qualification ( ≥ 5000 completed HITs with >97% approval rate). 32 participants passed screening, and completed the survey in a median time of 15.9 minutes and were compensated US$2.00. They were 41.7% female and 23-69 years old (Median = 35). CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA Wencan Zhang, Mariella Dimiccoli, and Brian Y. Lim Per-Image Trial: Image of a Dog Blur = Weak Randomized q1. Subjective saliency grid-select user selected Self-report Truthfulness Questionnaire q2. Representativeness Rating (10-star) q3. Rating rationale (free text) Unbiased- CAM Debiased- CAM Biased- CAM Randomized order, shown side-by-side Blur = Strong Blur = None Blur level 10× a b Figure 5: User Study 1 main study procedure (a) and grid selection UI to measure CAM Truthfulness Selection (b). 5.1.4 Statistical Analysis and Quantitative Results. For all dependent variables, we fit a multivariate linear mixed effects model with Blur Bias Level, CAM Types, Image Label and Task Time Level as fixed effects, Blur Bias Level × CAM Type, Image Label × Blur Bias Level, Image Label × CAM Type, Task Time Level × Blur Bias Level and Task Time Level × CAM Type as fixed interaction effects, and Participant as a random effect. Supplementary Table 3 reports the model fit ( R2 ) and significance of ANOVA tests for each fixed effect. Due to the large number of comparisons in our analysis, we consider differences with p< . 001 as significant. This is sufficiently strict for a Bonferroni correction for 50 comparisons ( α = .05/50). Furthermore, all results reported were significant at p< . 0001, unless otherwise stated. We performed post-hoc contrast tests for specific differences described. All statistical analyses were performed using JMP (v14.1.0). Fig. 6 summarizes our results. Unbiased-CAM had the highest CAM Truthfulness Selection Similarity, while Biased-CAM the lowest Similarity that was only 21.3-43.7% of the truthfulness of Unbiased-CAM. Debiased-CAM had significantly higher CAM Truthfulness Selection Similarity than Biased-CAM at 69.4-79.0% of the truthfulness of Unbiased-CAM. Similarly, for blurred images, participants rated Unbiased-CAM as the most truthful (M = 7.83 out of 10, standard error = 0.12), followed by Debiased-CAM (M = 6.00 ± 0 . 21 to 7.21 ± 0 . 18), and Biased-CAM as the least truthful (M = 3.05 ± 0 . 21 to 4.98 ± 0 . 26). In summary, Debiased-CAM improved CAM truthfulness, despite stronger blur that reduced CAM truthfulness by highlighting wrong or unexpected regions, sizes, and shapes. 5.1.5 Thematic Analysis and Qualitative Findings. We analyzed the rationale of participant ratings to better understand how participants interpreted different CAMs as truthful or untruthful, and what visual features they perceived in images and CAMs. We performed a thematic analysis with open coding [ 64 ]. Two authors 0.0 0.2 0.4 0.6 0.8 q1.)CAM)Truthfulness) Selec9on)Similarity)(PCC) Unbiased Debiased Biased CAM)Type 2 4 6 8 q2.'CAM'Truthfulness'' Ra7ng Unbiased Debiased Biased CAM'Type 2 4 6 8 q2.'CAM'Truthfulness'' Ra7ng Unbiased Debiased Biased CAM'Type 2 4 6 8 q2.'CAM'Truthfulness'' Ra7ng Unbiased Debiased Biased CAM'Type ab Figure 6: User Study 1 results. CAM Truthfulness decreased with blur, but was improved with Debiased-CAM. Dotted lines indicate very significant p< . 0001 comparisons; solid lines indicate no significance at p> . 01 . Error bars indicate 90% confidence interval. independently coded the rationales and discussed the coding until themes converged. Next, we first describe rationales for different blur levels, then describe themes spanning all blur levels. Note that all CAM types were shown anonymously (labeled A, B, and C) with randomly orders; we quote them specifically by type for clarity. For None blur, as expected, most participants perceived CAMs as identical, e.g., “all 3 images are the same and mostly representative” (Participant P23, “Fish” image); though some participants could perceive the slight decrease in the CAM truthfulness of Debiased- CAM, e.g., for the “Church” image, P1 wrote that Unbiased-CAM and Biased-CAM “had the most focus on *all* the crosses on the roof of the church and therefore I thought they were the most representative. [Debiased-CAM] gives less importance to the leftmost cross on the roof and therefore was rated lower.” For Weak blur, participants felt Unbiased-CAM was very truthful, Debiased-CAM was slightly less truthful, and Biased-CAM was untruthful; e.g., P29 felt that Biased-CAM “doesn’t show anything but blackness, [other CAMs] are much better in the way the heatmap shows details.” For Strong blur, participants perceived Debiased-CAM as moderately truthful, but Biased-CAM as very untruthful, e.g., P18 felt that “[Biased-CAM] is totally off, nothing there is a garbage truck. [Unbiased-CAM] shows the best and biggest area, and [Debiased-CAM] is good too but I’m thinking not good enough as [Unbiased-CAM].” Across blur conditions, we found that participants interpreted whether a CAM was truthful based on several criteria — primary object, object parts, irrelevant object, coverage span, and shape. Participants checked whether the primary object in the label was highlighted (e.g., “That heatmap that focuses on the chainsaw itself is the most representative.” P20, Chain Saw), and also checked whether specific parts of the primary object were included in the highlights (e.g., “[Unbiased-CAM and Debiased-CAM] correctly identify the fish though [Unbiased-CAM] also gives importance to the fish’s rear fin.” P1, Fish, Weak blur). P15 noted differences between the CAMs for the “French Horn” image: “[Unbiased-CAM] places the emphasis over the unique body of the French horn, and it places more well-defined, yellow and green emphasis on the mouthpiece and the opening of the horn itself. [Biased-CAM] is too vertical to completely capture the whole horn, and [Debiased-CAM]’s red area is too small to capture the body of the horn, and does not capture the opening of the horn or the mouthpiece.” Participants rated a CAM as less truthful if it highlighted irrelevant objects, e.g., “[Debiased-CAM] is quite close to capturing the entire church. (But) [Unbiased-CAM] captures more Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA of the tree.” (P26, Church). Much discussion also focused on the coverage of salient pixels. Less truthful CAMs had coverages that were either too wide (e.g., “[Debiased and Biased CAMs] are inaccurate. They are too wide.” P22, Garbage Truck), covering the background or other objects to get “less representative when it misleads you into the background or surroundings of the focus. It needs to only emphasize the critical area.” (P23, Church); or too narrow, not covering enough of the key object such that it “is very small and does not highlight the important part of the image. It is too narrow.” (P30, Fish). Finally, participants appreciated CAMs that highlighted the correct shape of the primary object, e.g., “[Debiased-CAM] perfectly captures the shape of the ball and all of its quadrants. [Unbiased-CAM] is a little more oblong than the golf ball itself, so it’s not as perfect. [Biased- CAM] is almost a vertical red spot and does not really capture the shape of the golf ball at all.” (P15, Golf Ball). In summary, we found that Debiased-CAM and Unbiased-CAM were perceived as truthful, because they: 1) highlighted semantically relevant targets while avoiding irrelevant ones, so concept or objectaware CNN models are important [ 10 , 43 ]; 2) had salient regions that were neither too wide nor narrow for the image domain; and 3) had accurate shape and edge boundaries for salient regions, which can be obtained from gradient explanations [74]. 5.2 User Study 2 (CAM Helpfulness) The second study evaluated the perceived helpfulness of each CAM type to verify predictions of blur biased images. 5.2.1 Experiment Procedure. The procedure is the same as User Study 1, except for the main study section. User Study 1 focused on CAM Truthfulness to obtain the participant’s saliency annotation of the unblurred image before revealing CAMs. In User Study 2, showing the unblurred image first will invalidate the use case of verifying predictions on blurred images, since the participant would have foreknowledge of the image. Hence, participants needed to see the blurred image and model prediction first, answer perception questions, then see the image unblurred. In the main study (Fig. 7a), each participant viewed 7 repeated image trials, each randomly assigned to one of 9 conditions (3 Blur Bias levels × 3 CAM types) in a within-subjects experiment design. Participants viewed 7 randomly chosen images from the same 10 images of User Study 1, instead of all 10, so that they could not easily conclude the class label for the remaining images by eliminating previous classes. For each trial, the participant performed the common explainable AI task to verify the label prediction of the model. On the first page, the participant viewed a labeled image at the assigned Blur Bias level with corresponding CAM for the assigned CAM type, indicated her likelihood choice(s) for the image label with the “balls and bins” question [ 32 ] to elicit user labeling (Fig. 7b) (q1); rated how well each CAM represented the image label (q2); rated how helpful the CAM was for verifying the label (q3), and wrote the rationale for her rating (q4). On the next page, participants saw the image unblurred and answered questions q2-4 again as questions q5-7. See Supplementary Fig. 14 for questionnaire details. 5.2.2 Experiment Apparatus and Measures. For q1, we asked the participant to indicate likelihoods of 10 possible image labels with Per-Image Trial: Auto labeled as Dog Blur Level CAM type Blur = Weak, CAM = Debiased Label Verification (Preconceived) q1. Likely label %’s (balls & bins) q2. Representativeness (10-star) q3. CAM helpfulness (7-pt Likert) q4. Helpfulness rationale (free text) Label Verification (Consequent) q5. Representativeness (10-star) q6. CAM helpfulness (7-pt Likert) q7. Helpfulness rationale (free text) Scene Blur Image CAM Blur Image CAM Weak Blur, Unbiased-CAM Weak Blur, Biased-CAM CAM type CAM type Weak Strong None DebiasedUnbiased Biased Random 7×(of 10) Randomized a b Figure 7: User Study 2 main study procedure (a) and “balls and bins” UI to elicit user labeling (b). the “balls and bins” question [ 19 , 32 , 80 ] to elicit her probability distribution p={pc}T over label classes c∈C . This question is reliable in eliciting probabilities from lay users [ 32 , 80 ] and avoids priming participants with the actual label c0 , since it asks about all labels. We calculated the participant’s selected label ´ c as the class with the highest probability, i.e., ´ c=argmaxc(pc) , Labeling Confidence as the indicated likelihood for the actual label pc0 , and Label Correctness as [´ c=c0] , where [·] is the Iverson bracket notation. CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA Wencan Zhang, Mariella Dimiccoli, and Brian Y. Lim A.2 Model Variants 𝐶𝑁𝑁!𝑦$ 𝑴 & Biased-CAM a 𝐿"𝐶𝑁𝑁!𝑦$ 𝑴 & Unbiased-CAM b 𝐿" 𝐶𝑁𝑁 ! "#,"% 𝑴 ( Biased-CA𝐌𝒇 𝒔𝒃,𝒔𝒕 g 𝑦) 𝐿"𝐶𝑁𝑁 ! *#,"% 𝑴 ( Biased-CA𝐌𝒇 𝒎𝒃,𝒔𝒕 h 𝑦) 𝐿" e 𝑴 *#$ 𝒚 ,#$ 𝐿% Debiased-CA𝐌𝒔𝒃,𝒔𝒕 𝐿" 𝐶𝑁𝑁! "#,"% + 𝑴 * 𝒚 , 𝐿% Debiased-CA𝐌𝒔𝒃,𝒎𝒕 𝐿" f 𝐶𝑁𝑁! "#,&% 𝐿& 𝑏 . 𝑴 *#$ 𝒚 ,#$ 𝐿% Debiased-CA𝐌𝒎𝒃,𝒔𝒕 𝐿" 𝐶𝑁𝑁! &#,"% + c 𝐿& 𝑏 . 𝐿" 𝑴 * 𝒚 , 𝐿% Debiased-CA𝐌𝒎𝒃,𝒎𝒕 d 𝐶𝑁𝑁! &#,&% /𝒙 //𝒙& //𝒙& //𝒙& //𝒙& //𝒙& //𝒙& //𝒙& Supplementary Fig. 1. Architectures of self-supervised DebiasedCNN variants and of baseline CNN models and their CAM explanations from a biased “Dog” image blurred at σ= 24 . a) RegularCNN on biased image. b) RegularCNN on unbiased image. c) DebiasedCNN (mb, st) with single-task loss as a sum of classification and CAM losses for the classification task, trained on multi-bias images with auxiliary bias level prediction task. d) DebiasedCNN (sb, mt) with multi-task for CAM prediction trained with differentiable CAM loss, and trained on multi-bias images with auxiliary bias level prediction task. e) DebiasedCNN (sb, st) with single-task loss as a sum of classification and CAM losses for the classification task. f) DebiasedCNN (sb, mt) with multi-task for the CAM prediction and differentiable CAM loss. g) FineTunedCNN (sb,st) retrained on images biased at a single-bias level. h) FineTunedCNN (mb,st) retrained on images biased variously at multi-bias levels. Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA Model Variant Training Loss Function Training Set Bias Levels RegularCNN ! ( # ) =!! ( &,& (( # )) )=0 FineTunedCNN (sb, st) ! ( # ) =!! ( &,& (( # )) )∈(0,)"#$] FineTunedCNN (mb, st) ! ( # ) =!! ( &,& (( # )) )∈-%#&'.~.0 ([ 0,)"#$ ]) DebiasedCNN (sb, st) ! ( # ) =!! ( &,& (( # )) +3(!( ( 4 5 ,4 6) )∈(0,)"#$] DebiasedCNN (mb, st) 7 ( # ) = 8 !! ( &,& (( # )) +3(.!( ( 4 5 ,4 6) 3).!) ( ),) 9( # )) : )∈-%#&' DebiasedCNN (sb, mt) 7 ( # ) = 8 !! ( &,& (( # )) 3(!( ( 4 5 ,4 6( # )): )∈(0,)"#$] DebiasedCNN (mb, mt) 7 ( # ) = ; !! ( &,& (( # )) 3(!( ( 4 5 ,4 6( # )) 3)!) ( ),) 9( # )) < )∈-%#&' Supplementary Table 2. CNN model variants with single-task (st) or multi-task (mt) architectures trained on a specific (sb) or multiple (mb) bias levels. Each training set image x∈Xis preprocessed by a bias operator Bat a selected level b , i.e., xb=B(x,|b|>0),∀x∈X.Bdepends on the bias type (e.g., blur, color temperature, day-night lighting). For DebiasedCNN, mt refers to including a CAM task with differentiable CAM loss separate from the primary prediction task, while st refers to the primary prediction task with non-differentiable CAM loss. Models trained for single-bias (sb) used training set images biased at a single level b> 0 , while models trained for multi-bias levels (mb) used training datasets with data augmentation where each image is biased to a level that is randomly selected from a uniform probability distribution Br and ∼U([ 0 ,bmax ]). Multi-bias DebiasedCNN also adds a task for bias level prediction. Loss functions in vector form specify one loss function per task in a multi-task architecture. CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA Wencan Zhang, Mariella Dimiccoli, and Brian Y. Lim A.3 Debiasing spurious explanations of privacy-preserving AI 𝒙𝑦# 𝑴 % 𝑦& 𝒙! Bias 𝔅(𝒙,𝑏) 𝑦 Public 𝐿" 𝐿! 𝑏 , Label !𝐴𝑛𝑛𝑜𝑡𝑎𝑡𝑜𝑟 ! Private Debiased-CAM 𝐿#𝐶𝑁𝑁$ 𝐿# a b 𝐶𝑁𝑁% 𝑏 𝑴 / Unbiased-CAM Supplementary Fig. 2. Architecture of multi-task DebiasedCNN model with self-supervised learning from private training data for privacypreserving machine learning. a) RegularCNN (CN N0) was trained on a private dataset with unblurred image x x xto generate Unbiased-CAM e M M M. b) DebiasedCNN (CN Nd) was trained on the corresponding public (privacy-protected) biased form of the private dataset with blurred imagex x xb and self-supervised with Unbiased-CAM e M M Mto generate Debiased-CAM b M M M. During model training, CN Ndhas access to the bias level bof each image x x xb, Unbiased-CAM e M M M, and actual label y, but has no access to them during model inference. C N Ndnever has access to any unblurred image x x x. At inference time, DebiasedCNN can generate relevant and faithful Debiased-CAMs from privacy-protected blurred images. Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA B SIMULATION STUDIES APPENDIX B.1 Supplemental Method: Calculating Bias Levels We provide details to calculate different bias for color temperature and lighting biases. B.1.1 Color Temperature Bias. Color temperature refers to the temperature of an ideal blackbody radiator as if illuminating the scene. We biased color temperature as follows. Each pixel in an unbiased image has color (r,д,b)T , where R,G,B represent the red, green, and blue color values within range 0-255, respectively. Each pixel is biased from neutral temperature t0 by ∆tb at bias level b by multiplying a diagonal correction matrix with its color, i.e., (rb,дb,bb)T=diaд(255/Rb,255/Gb,255/Bb) (r,д,b)T,(5) where (Rb,Gb,Bb)T=fCT (T)=fCT (t0+∆tb) are scaling factors obtained from Charity’s color mapping function fCT to map a blackbody temperature to RGB values [ 12 ] (Supplementary Fig. 3). We set the neutral color temperature t0 to 6600K, which represents cloudy/overcast daylight. Color temperature biasing is asymmetric about zero bias, because people are more sensitive to perceiving changes in orange than blue colors (Kruithof Curve [ 17 ]); and due to the non-linear monotonic relationship between blackbody temperature and modal color frequency (Wien’s Displacement Law). This asymmetry explains why orange biasing led to stronger CAM deviation than blue biasing. Supplementary Fig. 3. Color mapping function to bias color temperature of images in Simulation Study 4. Changes in Red, Green, Blue values are larger for orange biases (lower color temperature) than blue biases (higher temperature). Neutral color temperature is set to represent shaded/overcast skylight at 6600K. B.1.2 Lighting Bias. Lighting bias occurs when the same scene is lit brightly or dimly. In nature, this occurs as sunlight changes hour-to-hour, or season-to-season. The Transient Attributes database [ 51 ] contains photos of scenes from the same camera position taken across different times of the day and year. Attribute changes include whether the scene is daytime or nighttime, snowy, foggy, dusk/dawn or not. We sought to generate images with different degrees of darkness, but the dataset only contained photos that were very bright or very dark. Therefore, we interpolated photos to generate scenes with intermediate darkness. For each scene, with a daytime image Iday(x,y) and nighttime image Iniдht (x,y), we performed the pixel-wise interpolation as, Ibiased (x,y)=(1−ρ) × Iday(x,y)+ρ×Iniдht (x,y),(6) where ρ is the night/day ratio. An unbiased image has ρ= 0indicating daytime, and the most biased image has ρ= 1indicating nighttime. CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA Wencan Zhang, Mariella Dimiccoli, and Brian Y. Lim B.2 Supplemental Results Night&Bias&Ra+o 0.0 0.3 0.6 0.9 Predicted&Bias&Ra+o 0.0 0.3 0.6 0.9 R²:&0.863 Task = Classification Bias = Blur Data = NTCIR-12 bTask = Classification Bias = Color Temperature Data = NTCIR-12 d Task = Captioning Bias = Blur Data = COCO c Task = Classification Bias = Blur Data = ImageNette a Blur Bias Level (σ)Blur Bias Level (σ)Blur Bias Level (σ)Actual ∆t Bias Task = Classification Bias = Day-to-Night Data = TransAttr e Night Bias Ratio (𝜌) Supplementary Fig. 4. Regression performance for DebiasedCNN (mb, mt) measured as R2 for the bias level prediction task for five simulation studies. Very high R2values indicate that models trained for Simulation Studies 1-3 and 5 could predict the respective bias levels well. Color temperature bias level prediction depended on whether bias was towards lower (more orange) or higher (more blue) temperatures. Since bluebiased images were less distinguishable, the model was less well-trained to predict the blue color temperature bias level; it was more able to predict orange bias at a reasonable accuracy. Prediction Task VGG16 Blur Bias Level (𝜎)016 32 Biased Image RegularCNN FineTunedCNN DebiasedCNN a b c ResNet50 0 16 32 Xception 016 32 Confidence: 0.947 CAM PCC: 1.000 Confidence: 0.001 CAM PCC: –0.079 Confidence: 0.000 CAM PCC: –0.063 Confidence: 0.947 CAM PCC: 1.000 Confidence: 0.098 CAM PCC: 0.530 Confidence: 0.007 CAM PCC: 0.000 Confidence: 0.999 CAM PCC: 0.940 Confidence: 0.471 CAM PCC: 0.883 Confidence: 0.022 CAM PCC: 0.814 BLEU-4: 1.000 CAM PCC: 1.000 BLEU-4: 0.070 CAM PCC: 0.276 BLEU-4: 0.005 CAM PCC: 0.264 BLEU-4: 1.000 CAM PCC: 1.000 BLEU-4: 0.966 CAM PCC: 0.588 BLEU-4: 0.063 CAM PCC: 0.459 BLEU-4: 1.000 CAM PCC: 0.959 BLEU-4: 0.990 CAM PCC: 0.810 BLEU-4: 0.762 CAM PCC: 0.738 Confidence: 0.999 CAM PCC: 1.000 Confidence: 0.483 CAM PCC: 0.855 Confidence: 0.026 CAM PCC: 0.144 Confidence: 0.999 CAM PCC: 1.000 Confidence: 0.725 CAM PCC: 0.903 Confidence: 0.192 CAM PCC: 0.519 Confidence: 0.999 CAM PCC: 0.991 Confidence: 0.977 CAM PCC: 0.942 Confidence: 0.863 CAM PCC: 0.964 1 0 Saliency Supplementary Fig. 5. Deviated and debiased CAM explanations from various CNN models at varying bias levels of blur biased image from NTCIR-12 labeled as “Biking”. a) VGG16, b) ResNet50, c) Xception. a)-c), Models arranged in increasing CAM Faithfulness (see Supplementary Fig. 6, second row). CAMs from more performant models were more representative of the image label with higher CAM Faithfulness (PCC). Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA Model = ResNet50Model = Inception v3Model = VGG16 a b c Model = Xception d Blur Bias Level (σ)Blur Bias Level (σ)Blur Bias Level (σ)Blur Bias Level (σ) Supplementary Fig. 6. Comparison of model Task Performance and CAM Faithfulness for image classification on NTCIR-12 trained with different CNN models. a) VGG16, b) Inception v3, c) ResNet50, d) Xception. a)-d) Results agreed with Fig. 4 that higher bias led to lower Task Performance and CAM Faithfulness, but debiasing improved both. CNN models are arranged in increasing CAM Faithfulness from left to right. All models were pretrained on ImageNet and fine-tune on NTCIR-12. We set the last two layers of VGG16, and last block of ResNet50 and Xception as retrainable. b)-d) Newer base CNN models than VGG16 significantly outperformed it for both Task Performance and CAM Faithfulness. These newer models had similar Task Performance across bias levels, though their CAM Faithfulness differed more notably. CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA Wencan Zhang, Mariella Dimiccoli, and Brian Y. Lim C USER STUDIES APPENDIX C.1 User Studies Image Selection and CAMs For both user studies, we chose 10 images to select one instance per class label for 10 classes of ImageNette. This balanced between selecting a variety of images for better external validity, and too much workload for participants due to too many trials. CAMs were generated from specific CNN models in Simulation Study 1. At each blur level, Unbiased-CAM and Biased-CAM were generated from RegularCNN, while Debiased-CAM was generated from DebiasedCNN (mb, mt). A key objective of the user studies was to validate the results of the simulation studies regarding CAM types and image blur bias levels, hence, we selected canonical images that: (1) Had RegularCNN and DebiasedCNN predict correct labels for unblurred images, since we were not investigating the use of CAMs to debug model errors. CNN predictions on blurred images may be wrong, but we showed the CAM of the correct label. (2) Were easy to recognize when unblurred, so that users can perceive whether a CAM is representative of a recognizable image. This was validated in our pilot study. (3) Were somewhat difficult but not impossible to recognize with Weak blur, so that participants can feasibly verify image labels with some help from CAMs. (4) Were very difficult to recognize with Strong blur, such that about half of pilot participants were unable to recognize the scene, to investigate the upper limits of CAM helpfulness. (5) Had Unbiased-CAMs that were representative of their labels, to evaluate perceptions with respect to truthful CAMs. Conversely, debiasing towards untruthful CAMs is futile. (6) Had Biased-CAMs for Strong blur that were perceptibly deviated and localized irrelevant objects or pixels; otherwise, no difference between Unbiased-CAM and Biased-CAM will lead to no perceived difference between Unbiased-CAM and Debiased-CAM too. (7) Had Debiased-CAMs that were an approximate interpolation between the Unbiased-CAM and Biased-CAM of each image, to represent the intermediate CAM Faithfulness of Debiased-CAM found in the simulation studies. These criteria were verified with participants in a pilot study and the selected images had CAM Faithfulness representative of Simulation Study 1 for Debiased-CAM, but with slightly lower CAM Faithfulness for Biased-CAM to represent worse case scenarios. CAMs were different based on CAM type and Blur Bias level. Unbiased-CAMs were the same for all Blur Bias levels, and Unbiased-CAM and Biased-CAM were the same for None blur level. For other conditions, CAMs were deviated and debiased based on CAM type and Blur Bias level. We chose not to test participants with images in NTCIR-12 due to quality and recognizability issues. Since images were automatically captured at regular time intervals, many images were transitional (e.g., pointing at ceiling while “Watching TV”), which made them unrepresentative of the label. Furthermore, in pilot testing, participants had great difficulty recognizing some scenes (e.g., “Cleaning and Chores”) in images with Strong blur, such that the tasks became too confusing to test. Nevertheless, our results can generalize to wearable camera images with Weak blur, for users who are familiar with or can remember their personal recent or likely activities. Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA Blur Bias Level None Weak Strong Biased Image Unbiased- CAM Debiased- CAM Biased- CAM Label: Cassette Player Blur Bias Level None Weak Strong Biased Image Unbiased- CAM Debiased- CAM Biased- CAM Label: Dog Blur Bias Level None Weak Strong Label: Chain Saw Blur Bias Level None Weak Strong Label: Church Blur Bias Level None Weak Strong Label: Fish Blur Bias Level None Weak Strong Label: French Horn a b c d e f Supplementary Fig. 7. Images and CAMs at various Blur Bias levels and CAM types that participants viewed in both User Studies. CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA Wencan Zhang, Mariella Dimiccoli, and Brian Y. Lim Blur Bias Level None Weak Strong Biased Image Unbiased- CAM Debiased- CAM Biased- CAM Label: Garbage Truck Blur Bias Level None Weak Strong Biased Image Unbiased- CAM Debiased- CAM Biased- CAM Label: Golf Ball Blur Bias Level None Weak Strong Label: Gas Pump Blur Bias Level None Weak Strong Label: Parachute g h i j Supplementary Fig. 8. (Continued) Images and CAMs at various Blur Bias levels and CAM types that participants viewed in both User Studies. CAM$Type Unbiased Debiased Biased CAM$Faithfulness$(PCC) 0.0 0.5 1.0 Bias%Level%(Blur) None Weak Strong Supplementary Fig. 9. CAM Faithfulness of selected 10 image instances used in user studies. Faithfulness decreased as Blur Bias increased, was the highest for Unbiased-CAM, the lowest for Biased-CAM, and improved by Debiased-CAM. Error bars indicate 90% confidence interval. Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA C.2 User Study 1 and 2 Questionnaires We illustrate key sections in the questionnaire for the CAM Truthfulness User Study 1 and CAM Helpfulness User Study 2. Both questionnaires were identical except for the main study section. Supplementary Fig. 10. Tutorial to introduce the scenario background of a smart camera with privacy blur and heatmap (CAM) explanation. It taught the participant to i) interpret the “balls and bins” question ([32]), ii) understand why images were blurred, and iii) interpret the CAM. CHI ’22, April 29-May 5, 2022, New Orleans, LA, USA Wencan Zhang, Mariella Dimiccoli, and Brian Y. Lim C.4 Supplementary Results 5 7 9 CAM'Truthfulness'Ra4ng Unbiased Debiased Biased CAM'type 5 6 7 8 9 CAM)Truthfulness)Ra6ng Unbiased Debiased Biased CAM)type 5 7 9 CAM'Truthfulness'Ra4ng None Weak Strong Bias'Level'(Blur) 0 1 2 CAM'Helpfulness'Ra2ng Unbiased Debiased Biased CAM'type 0 1 2 CAM'Helpfulness'Ra2ng Unbiased Debiased Biased CAM'type 0 1 2 CAM'Helpfulness'Ra2ng None Weak Strong Bias'Level'(Blur) Unblurred)Disclosure Preconceived Consequent Unblurred)Disclosure Preconceived Consequent Unblurred)Disclosure Preconceived Consequent Supplementary Fig. 15. Comparisons of perceived CAM Truthfulness and CAM Helpfulness before (preconceived) and after (consequent) disclosing the unblurred image. There was a significant difference across Unblurred Disclosure for CAM Truthfulness Rating (p = .0013) but not for CAM Helpfulness Rating. Comparing preconceptual to consequent ratings, Unbiased-CAMs were rated as less truthful (M = 7.7 vs. 8.3, p = .0004), Debiased-CAMs were rated marginally less truthful (p = .0212), Biased-CAMs were rated similarly untruthful, and overall, CAMs of Strongly blurred images were rated as less truthful (M = 5.6 vs. 6.3, p< . 0001 ). These results suggest that even with the least biased CAM (Unbiased-CAM), the unfamiliarity of unblurred scenes can hurt trust (truthfulness) in the CAM, though there was no change in perceived helpfulness before or after disclosing the unblurred image. CAM Truthfulness Ratings were measured along a 1-10 scale, and CAM Helpfulness Ratings along a 7-point Likert scale (–3 = Strongly Disagree, 0 = Neither, +3 = Strongly Agree). Error bars indicate 90% confidence interval. Dotted lines indicate extremely significant p< .0001 comparisons, and solid lines indicate no significance at p> .01.