scieee AI-readable full text Open interactive document viewer

Multi-layer Representation Learning for Robust OOD Image Classification

Ballas, Aristotelis; Diou, Christos

Abstract

Abstract - Convolutional Neural Networks have become the norm in image classification. Nevertheless, their difficulty to maintain high accuracy across datasets has become apparent in the past few years. In order to utilize such models in real-world scenarios and applications, they must be able to provide trustworthy predictions on unseen data. In this paper, we argue that extracting features from a CNN’s intermediate layers can assist in the model’s final prediction. Specifically, we adapt the Hypercolumns method to a ResNet-18 and find a significant increase in the model’s accuracy, when evaluating on the NICO dataset.

Full text

Multi-layer Representation Learning for Robust OOD Image Classification Aristotelis Ballas∗ Department of Informatics and Telematics Harokopio University Athens, Greece [email protected] Christos Diou∗ Department of Informatics and Telematics Harokopio University Athens, Greece [email protected] Abstract Convolutional Neural Networks have become the norm in image classification. Nevertheless, their difficulty to maintain high accuracy across datasets has become apparent in the past few years. In order to utilize such models in real-world scenarios and applications, they must be able to provide trustworthy predictions on unseen data. In this paper, we argue that extracting features from a CNN’s intermediate layers can assist in the model’s final prediction. Specifically, we adapt the Hypercolumns method to a ResNet-18 and find a significant increase in the model’s accuracy, when evaluating on the NICO dataset. Keywords: deep learning, domain generalization, out of distribution, image classification ACM Reference Format: Aristotelis Ballas and Christos Diou. 2022. Multi-layer Representation Learning for Robust OOD Image Classification. In 12th Hellenic Conference on Artificial Intelligence (SETN 2022), September 7–9, 2022, Corfu, Greece. ACM, New York, NY, USA, 4pages. https: //doi.org/10.1145/3549737.3549780 1 Introduction In the past few years, Deep Learning has established itself in academia and industry. In particular, Convolutional Neural Networks have dominated image classification [ 11 ], achieving near-human, if not superhuman [ 7 ], accuracy. Despite their outstanding results in IID (independent and identically distributed) datasets, most models today fail to generalize well on unseen or out of distribution (OOD) settings [ 16 ], since they tend to incorporate statistical correlations present ∗Both authors contributed equally to this research. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. SETN 2022, September 7–9, 2022, Corfu, Greece ©2022 Association for Computing Machinery. ACM ISBN 978-1-4503-9597-7/22/09...$15.00 https://doi.org/10.1145/3549737.3549780 Figure 1. Visualization of the proposed model. The backbone of the architecture is a vanilla ResNet-18. By extracting feature maps from intermediate outputs of convolutional layers in the backbone model, our method takes advantage of the CNN’s ability to detect edge and bar-like figures in the early layers and combines them with the semantic features extracted from layers further down the network. We argue that by incorporating outputs from different intermediate levels of the network, we enable the model to disentangle the invariant qualities of an object. We extract features from the layers and residual connections marked in light blue. in the training data [ 1 ]. In real-world scenarios, data hardly ever originate from the same distribution, creating a need for approaches that are able to distinguish between biases and trivial features and make decisions based on invariant factors. This has been a long standing issue in the Machine Learning community and, as a result, in 2011 the Domain Generalization (DG) problem [ 3 ] was formally introduced. In the DG setting, the data used to evaluate the trained model originate from a different distribution than the training data. By making predictions on unseen data distributions, we can appropriately assess the model’s ability to generalize. In this work, we propose a method to tackle DG by adapting Hypercolumns [ 5 ] to extract local attributes of an image arXiv:2207.13678v1 [cs.CV] 27 Jul 2022 SETN 2022, September 7–9, 2022, Corfu, Greece Ballas and Diou. in earlier layers and semantics in layers further down the architecture. We hypothesize that by incorporating features from across the network, a classifier can be trained to ignore spurious correlations in the dataset and make predictions based on invariant features. We evaluate our method by experimenting on NICO [ 8 ], a dataset specifically designed for OOD image classification and are able to achieve promising results. Furthermore, to confirm our assumptions and intuition, we provide visual evidence of our model’s ability to distinguish between spurious and invariant characteristics present in an image. 1.1 Domain Generalization In this section, we formally introduce the notations and definitions of DG. Let 𝑋 be an input (feature) space and 𝑌 an output (label) space. A domain is defined as a joint distribution P(X,Y) ∼ PXY on X × Y . In DG, the training and test distributions are OOD, in the sense that we are given S source (training) domains and T target (test) domains, where P𝑖 XY ≠P𝑗 XY , 1 ≤ i, j ≤ S, T. Given labeled source domains S, the goal is to learn a model F , trained on data from S, which can adequately generalize to an unseen domain T. 2 Related Work Domain Generalization is arguably one of the most challenging problems in Machine Learning. To this end, a plethora of approaches have been proposed in the past few years. The most closely related fields to DG are: • Domain Adaptation [ 20 ] and Transfer Learning [ 23 ] methods , which are perhaps the most common, focus on boosting their accuracy on unseen data by finetuning pretrained models on the target domain(s). • Meta-Learning [ 10 ], aims to learn-to-learn and select the best method for solving the issue at hand. • Continual Learning [ 14 ] algorithms are used to overcome the issue of catastrophic forgetting by remembering the knowledge acquired over time and domains. • Zero-Shot Learning [ 21 ], like DG deals with unseen distributions but in the label space. With regard to learning stable or invariant features across domains, besides the baseline CNN proposed in the original paper [ 8 ], several other methods have been suggested. To address the issues of complex, non-linear correlations between data in DG, the authors of [ 22 ] propose StableNet, a model which utilizes Random Fourier Features for sample weighting. In [ 12 ], the authors introduce the LIRR algorithm in the Semi-Supervised Domain Adaptation setting, for learning invariant representations and risks. Another approach is to use gradient-based semantic augmentation [ 2 ] to improve the generalizability of a model. Finally, the causal structure of the data can also be utilized while a model is trained, as shown in [18]. 3 Methodology 3.1 Hypercolumns Hypercolumns were first introduced in Neuroscience by Hubel and Wiesel [ 9 ], in order to describe a vertical set of V1 neurons that behave similarly to optical stimuli of the retina. The authors of [ 5 ] borrowed this term and applied its fundamental attributes to a Convolutional Neural Network, in an attempt to leverage the different levels of information passed on the network’s intermediate layers. Namely, a Hypercolumn at a certain location is a stacked vector of the layer outputs of the CNN’s units above said location. In order to classify pixels using Hypercolumns, it is assumed that bounding boxes of the points of interest have been provided from an object detection system. For each bounding box, a 50 ×50 heatmap (locations) is predicted, which is projected onto the initial image and then passed into a CNN. Selected intermediate outputs of the CNN are then concatenated into a vector and each location is classified via 1×1 convolutional and fully connected layers. In the same paper, the Efficient Hypercolumn method is also described, where the 1×1 convolutions are replaced by 𝑛×𝑛 convolutions and upsampling. In a subsequent work [ 6 ], the authors were able to significantly speed up their Hypercolumn pipeline by passing the whole image through the CNN and cropping the bounding box locations afterwards. Hypercolumns have been used for semantic segmentation [ 6 ], object detection [ 4 ], visual correspondence [ 15 ] and in some cases for abnormality detection in the Biomedical domain [ 19 ]. However, all above implementations construct their hypercolumns by concatenating the upsampled images into a hyper vector, possibly without taking full advantage of the already extracted features of the original image. Thus, we propose a novel implementation of the original Hypercolumn method and adapt it for robust image classification in the DG setting. 3.2 Adapting hypercolumns to robust image classification Due to their convolutional nature, earlier layers in a CNN are prominent in detecting edges and bars, but cannot distinguish between edges that belong to a vehicle or an animal, per se. The extracted information is then passed down the network and generalized in the final layer, where inference occurs. An object or class consists of features which remain invariant [ 1 ] across domains. For example, a horse still has legs and a mane, whether it is standing aside a person, lying in sand or galloping through snow. Therefore, by disentangling an input image into distinct features, a causal decision can be made based upon the features present in the image. Following the above example, the presence of ‘legs’ in an image lead us to believe that an animal is most likely depicted. The presence of a ‘mane’, ‘long tail’ and ‘oval-shaped’ hoove make us confident that the animal is a horse. We argue that Multi-layer Representation Learning for Robust OOD Image Classification SETN 2022, September 7–9, 2022, Corfu, Greece by taking advantage of the early and intermediate features of a CNN, we can ‘push’ a model to learn these invariant features (i.e edges and bars which correspond to parts of the depicted objects). For our model, we follow the original Hypercolumns implementation and select the outputs from intermediate conv layers, pass them through a 1×1 conv layer and then upsample the outputs via bilinear interpolation. However, before concatenating the upsampled images to a hyper vector, we pass them through a 56 ×56 MaxPool2D layer, with 26 ×26 strides, in an attempt to capture the features of the depicted class. After the pooling layer, the outputs are concatenated into a hyper vector and passed through a classification head, which consists of a fully connected layer followed by a softmax activation. Our model architecture is depicted in Figure 1. 4 Experiments 4.1 Datasets In our experiments we adopt the NICO [ 8 ] dataset. The NICO dataset was created for OOD image classification and is therefore a good starting point for the evaluation of the robustness of classification algorithms. NICO contains 2 super-classes of Animal and Vehicle. The Animal superclass contains 10 classes and the Vehicle contains 9. Each class contains 9 or 10 contexts, which try to simulate real world scenarios, such as ‘airplane aside mountain’ and ‘airplane on grass’. The total images in the dataset are 25.000. 4.2 Experimental Setup To evaluate our model we follow the leave-one-domain-out protocol as described in [ 13 ], where in our case a domain is a context of a class. In order to demonstrate the robustness of our model, during training we select to hold out 3, 5 and 7 contexts from each class. For the backbone of our model, we select a ResNet-18, pre-trained on ImageNet. The selected intermediate layers include all residual half block layers (i.e., those leading to reduction of the output size), as well as selected convolutional layers. All later layers of the network are included. We train the model with SGD for 30 epochs and with a batch size of 32 images. The learning rate is initially set at 0.001 and decays with a rate of of 0.1 at epoch 24. For a baseline, we used a vanilla pre-trained ResNet-18 with the same hyperparameters as above. Both models were implemented with PyTorch on one NVIDIA RTX A5000 GPU. 4.3 Results Table 1 depicts the averaged results on NICO after 3 runs. We can observe that our model outperforms the baseline by approximately 3.2% when 3 contexts are left out, 3.2% in the case where 5 contexts are left out and 2.4% when we leave out 7. To validate our initial assumptions, we also visualize our model’s prediction with saliency maps. To be more specific, Table 1. Top-1% Accuracy Results on the NICO Dataset when leaving out N contexts. If an image class has C contexts, we create a training split with the C−N contexts and evaluate our model on the remaining N. The presented results are averaged over 3 runs. Model N=3 N=5 N=7 ResNet-18 79.6 78.1 78.9 Our Model 82.8 81.3 83.0 Figure 2. Visualization of saliency maps, produced by the baseline vanilla ResNet-18 model and our method, from images in the NICO dataset. The brighter the pixel, the more it contributes to the model’s prediction. To a certain degree, our approach disregards the pixels corresponding to the contexts and spurious correlations (i.e water, snow, grass and road) in each image, and focuses on the depicted object. we adopted the Image-Specific Class Saliency method, as proposed in [ 17 ]. By computing and visualizing the gradient of the loss function for the predicted class, with respect to the input pixels, one is able to produce a map of the pixels affecting the model’s prediction. As shown in Fig. 2, the brightness in the saliency maps indicate the pixels which the model pays more ‘attention’ to. Due to spurious correlations present in the training data, the vanilla ResNet-18 SETN 2022, September 7–9, 2022, Corfu, Greece Ballas and Diou. tends to make assumptions based on unimportant features of the image (e.g. water, snow, grass and road - indicated by the bright pixels around the object in the saliency maps), while our method seems to infer based on features of the object itself and to some extent overlook the context features. 5 Conclusion In this paper we attempt to tackle the Domain Generalization problem by adopting the Hypercolumns method and adapting it for robust image classification. We argue that by utilizing the extracted features from a CNN’s intermediate layers, the model can be forced to focus on the invariant features in an image. This claim is supported by the results of our experiments on NICO, a dataset dominated by spurious correlations, where we demonstrate our model’s ability to perform well on unseen data. Through visual examples, we show that our method is capable of emphasizing on the causal characteristics of an object and not on the inconsequential features of the input image. As future work, we aim to advance our method’s extraction mechanism and conduct further experiments on additional datasets. Acknowledgments The work leading to these results has received funding from the European Union’s Horizon 2020 research and innovation programme under Grant Agreement No. 965231 project REBECCA (REsearch on BrEast Cancer induced chronic conditions supported by Causal Analysis of multi-source data). References [1] Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz. 2020. Invariant Risk Minimization. arXiv:1907.02893 [cs, stat] (2020). [2] Haoyue Bai, Rui Sun, Lanqing Hong, Fengwei Zhou, Nanyang Ye, Han-Jia Ye, S. H. Gary Chan, and Zhenguo Li. 2020. DecAug: Out-ofDistribution Generalization via Decomposed Feature Representation and Semantic Augmentation. arXiv:2012.09382 [cs.LG] [3] Gilles Blanchard, Gyemin Lee, and Clayton Scott. 2011. Generalizing from Several Related Classification Tasks to a New Unlabeled Sample. In Advances in Neural Information Processing Systems, J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, and K. Q. Weinberger (Eds.), Vol. 24. Curran Associates, Inc. https://proceedings.neurips.cc/paper/2011/ file/b571ecea16a9824023ee1af16897a582-Paper.pdf [4] Alexey Bochkovskiy, Chien-Yao Wang, and Hong-Yuan Mark Liao. 2020. YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv:2004.10934 [cs, eess] (2020). [5] Bharath Hariharan, Pablo Arbelaez, Ross Girshick, and Jitendra Malik. 2015. Hypercolumns for object segmentation and fine-grained localization. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Boston, MA, USA, 447–456. https: //doi.org/10.1109/CVPR.2015.7298642 [6] Bharath Hariharan, Pablo Arbeláez, Ross Girshick, and Jitendra Malik. 2017. Object Instance Segmentation and Fine-Grained Localization Using Hypercolumns. IEEE Transactions on Pattern Analysis and Machine Intelligence (2017). [7] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. CoRR (2015). [8] Yue He, Zheyan Shen, and Peng Cui. 2021. Towards non-iid image classification: A dataset and baselines. Pattern Recognition 110 (2021), 107383. [9] D. H. Hubel and T. N. Wiesel. 1962. Receptive fields, binocular interaction and functional architecture in the cat’s visual cortex. The Journal of Physiology 160, 1 (Jan. 1962), 106–154.2. https://www.ncbi.nlm.nih. gov/pmc/articles/PMC1359523/ [10] Mike Huisman, Jan N. van Rijn, and Aske Plaat. 2021. A survey of deep meta-learning. Artif Intell Rev 54, 6 (Aug. 2021), 4483–4541. https://doi.org/10.1007/s10462-021-10004-4 [11] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. ImageNet Classification with Deep Convolutional Neural Networks. In Advances in Neural Information Processing Systems, F. Pereira, C. J. C. Burges, L. Bottou, and K. Q. Weinberger (Eds.), Vol. 25. Curran Associates, Inc. https://proceedings.neurips.cc/paper/2012/file/ c399862d3b9d6b76c8436e924a68c45b-Paper.pdf [12] Bo Li, Yezhen Wang, Shanghang Zhang, Dongsheng Li, Kurt Keutzer, Trevor Darrell, and Han Zhao. 2021. Learning Invariant Representations and Risks for Semi-Supervised Domain Adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 1104–1113. [13] Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M. Hospedales. 2017. Deeper, Broader and Artier Domain Generalization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV). [14] Zheda Mai, Ruiwen Li, Jihwan Jeong, David Quispe, Hyunwoo Kim, and Scott Sanner. 2022. Online continual learning in image classification: An empirical survey. Neurocomputing 469 (Jan. 2022), 28–51. https://doi.org/10.1016/j.neucom.2021.10.021 [15] Juhong Min, Jongmin Lee, Jean Ponce, and Minsu Cho. 2019. Hyperpixel Flow: Semantic Correspondence With Multi-Layer Neural Features. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). [16] Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. 2019. Do ImageNet Classifiers Generalize to ImageNet?. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research), Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR. [17] Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014. Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps. arXiv:1312.6034 [cs.CV] [18] Xinwei Sun, Botong Wu, Xiangyu Zheng, Chang Liu, Wei Chen, Tao Qin, and Tie yan Liu. 2021. Latent Causal Invariant Model. arXiv:2011.02203 [cs.LG] [19] Mesut Toğaçar, Zafer Cömert, and Burhan Ergen. 2021. Enhancing of dataset using DeepDream, fuzzy color image enhancement and hypercolumn techniques to detection of the Alzheimer’s disease stages by deep learning model. Neural Computing and Applications (2021). [20] Mei Wang and Weihong Deng. 2018. Deep Visual Domain Adaptation: A Survey. Neurocomputing 312 (Oct. 2018), 135–153. https://doi.org/ 10.1016/j.neucom.2018.05.083 arXiv: 1802.03601. [21] Yongqin Xian, Bernt Schiele, and Zeynep Akata. 2017. Zero-Shot Learning - the Good, the Bad and the Ugly. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). [22] Xingxuan Zhang, Peng Cui, Renzhe Xu, Linjun Zhou, Yue He, and Zheyan Shen. 2021. Deep stable learning for out-of-distribution generalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5372–5382. [23] Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, and Qing He. 2021. A Comprehensive Survey on Transfer Learning. Proc. IEEE 109, 1 (Jan. 2021), 43–76. https://doi.org/10.1109/JPROC.2020.3004555