scieee AI-readable full text Open interactive document viewer

Benchmarking Domain Generalization Algorithms in Computational Pathology

Zamanitajeddin, Neda; Jahanifar, Mostafa; Xu, Kesi; Siraj, Fouzia; Rajpoot, Nasir

Abstract

Deep learning models have shown immense promise in computational pathology (CPath) tasks, but their performance often suffers when applied to unseen data due to domain shifts. Addressing this requires domain generalization (DG) algorithms. However, a systematic evaluation of DG algorithms in the CPath context is lacking. This study aims to benchmark the effectiveness of 30 DG algorithms on 3 CPath tasks of varying difficulty through 7,560 cross-validation runs. We evaluate these algorithms using a unified and robust platform, incorporating modality-specific techniques and recent advances like pretrained foundation models. Our extensive cross-validation experiments provide insights into the relative performance of various DG strategies. We observe that self-supervised learning and stain augmentation consistently outperform other methods, highlighting the potential of pretrained models and data augmentation. Furthermore, we introduce a new pan-cancer tumor detection dataset (HISTOPANTUM) as a benchmark for future research. This study offers valuable guidance to researchers in selecting appropriate DG approaches for CPath tasks.

Full text

Benchmarking Domain Generalization Algorithms in Computational Pathology Neda Zamanitajeddina,1, Mostafa Jahanifara,1, Kesi Xua, Fouzia Sirajc, Nasir Rajpoota,b,∗ aTissue Image Analytics centre, Department of Computer Science, University of Warwick, Coventry, UK, bHistofy Ltd, Coventry, UK, cICMR-National Institute of Pathology, New Delhi, India, Abstract Deep learning models have shown immense promise in computational pathology (CPath) tasks, but their performance often suffers when applied to unseen data due to domain shifts. Addressing this requires domain generalization (DG) algorithms. However, a systematic evaluation of DG algorithms in the CPath context is lacking. This study aims to benchmark the effectiveness of 30 DG algorithms on 3 CPath tasks of varying difficulty through 7,560 cross-validation runs. We evaluate these algorithms using a unified and robust platform, incorporating modality-specific techniques and recent advances like pretrained foundation models. Our extensive cross-validation experiments provide insights into the relative performance of various DG strategies. We observe that selfsupervised learning and stain augmentation consistently outperform other methods, highlighting the potential of pretrained models and data augmentation. Furthermore, we introduce a new pan-cancer tumor detection dataset (HISTOPANTUM) as a benchmark for future research. This study offers valuable guidance to researchers in selecting appropriate DG approaches for CPath tasks. Keywords: Domain Generalization, Computational Pathology, Domain Shift, Deep Learning, Benchmarking 1. Introduction Deep learning (DL) models have shown great potential in addressing fundamental problems in computational pathology (CPath) [1, 2] such as histology image classification [3, 4, 5], tissue segmentation [6, 7, 8], and nuclei detection [9, 10, 11]. Furthermore, leveraging DL, more advanced problems are being tackled in the field, such as gene expression prediction [12, 13, 14] and biomarker discovery [15, 16, 17]. Regardless of the versatility and accuracy of DL models on the training domain data, it has also been shown that testing on unseen data can degrade the performance metrics considerably [18] (see Fig. 1A) which is a phenomenon usually caused by domain shift (DS) [19]. Defining a data ‘domain’ as the joint distribution of feature (X) and label (Y) spaces, domain shift can be characterized by discrepancies in the joint distribution of features and labels across source (s) and target (t) domains, that is, Ps XY ,Pt XY . According to the Bayes’ theorem, the joint distribution is given by PXY =PX|YPY=PY|XPXwhich can be leveraged to describe DS manifested in various ways [18]: •Covariate Shift: The feature distributions differ between the source and target domains, i.e., Ps X,Pt X.Example: Tissue samples scanned using different scanners exhibit distinct colors and features. •Prior Shift: The label distributions vary between the source and target domains, i.e., Ps Y,Pt Y.Example: A ∗Corresponding authors: [email protected] 1Joint first authors, contributed equally. model trained on a dataset with a specific proportion of cancerous to non-cancerous samples is applied to a new dataset with a different ratio of these classes. •Posterior Shift: The conditional label distributions differ, i.e., Ps Y|X,Pt Y|X.Example: Subjective labeling in mitosis detection where different annotators assign different labels to the same data due to varying interpretations. •Class-Conditional Shift: The data characteristics for a specific class differ between the source and target domains, i.e., Ps X|Y=y,Pt X|Y=y.Example: Morphological traits of cancer cells in early-stage cancers differ from those in latestage cancers, leading to variations in the same class across datasets. A schematic presentation of different DS types is adopted from [18] and shown in Fig. 1B. An ideal way of, addressing domain shifts would involve training models across all conceivable data distributions. However, this approach is typically infeasible due to the limited availability of comprehensive, multi-domain data during the training phase. Consequently, there is an urgent need for algorithms specifically designed to improve domain generalization (DG). DG refers to the capability of a model trained on data from source domains Dsto perform well on unseen target domains Dtdespite distributional differences (Ps XY ,Pt XY ). It is important to note that, unlike domain adaptation techniques [20], DG algorithms enhance generalization to novel target domains without access to target domain data during training [21, 18]. There remains a critical gap in the utilization of existing DG algorithms within CPath. Many sophisticated DG algorithms Preprint submitted to Elsevier September 26, 2024 arXiv:2409.17063v1 [cs.CV] 25 Sep 2024 BC Metastasis Detection Pan-cancer Tumour Detection Pan-tumour Mitosis Detection (C) Three Tasks (D) 30 DG Algorithms Run 1 Run 2 Run 3 Run 4 (E) Robust Cross-validation Domains: Train Validation Test Sets: Domain 1 95% Accuracy 65% Accuracy Domain 2 Feature Distribution Domain 1 Domain 2 (A) Domain Shift Example Covariate Shift Prior Shift Posterior Shift Class-conditional Shift (B) Domain Shift Types Classes Sources Total Runs = ෍ 𝑑 ∈ datasets 9×30×𝑁domains 𝑑=7560. Figure 1: Domain shift in computational pathology can cause degradation in performance when testing on an unseen dataset (A). Different types of DS are illustrated in (B) with shapes as classes, colors as features, and each circle as a domain. In (B), covariate shift is presented by changing the color of objects in two different domains, prior shift happens when the distribution of classes differs between the two domains, and posterior shift is shown when the same objects are labeled differently by the observers (highlighted shapes), and in class-conditional shift, the color of only one class is changing between domains. Leveraging three different tasks (C) in this work, we benchmark the performance of 30 domain generalization algorithms (D) in a series of robust cross-validation experiments (E). have not been systematically explored in this field. Motivated by this gap, we present a rigorous evaluation of DG methods in the CPath context. This research aims to benchmark the effectiveness of 30 different DG algorithms (Fig. 1D) on three different CPath tasks of various difficulties (Fig. 1C) in a unified and robust platform through extensive cross-validation experiments (Fig. 1E). Our goal is to fairly compare existing DG algorithms to provide insights that could help researchers to select a better strategy for their DG needs. To this end, we build on the DomainBed platform [22] as our main repository of DG algorithms, by including new DG algorithms (CPath-specific algorithms such as stain normalization and stain augmentation, and more recent approaches such as a pretrained foundation model), more robust evaluation metric (F1 score), and two new multidomain histology image classification datasets. This work’s main contributions can be summarized as follows: •Presentation of a unified and robust framework for benchmarking DG algorithms in the area of CPath; •Benchmarking 30 DG algorithms on 3 different tasks in CPath using the proposed framework; •Comprehensive and robust cross-validation experiments, covering 7,560 training-validation runs; •Making recommendations for selecting effective DG strategies in CPath; and •Releasing a large-scale tumor patch dataset (that we term HISTOPANTUM) comprising 280K+images in 4 different cancer types and capturing three types of DG and making the benchmarking framework HistoDomainBed publicly available at: https:// github.com/mostafajahanifar/HistoDomainBed. In the remainder of the paper, datasets, algorithms, and crossvalidation procedure are described in Section 2, results are presented and discussed in sections 3 and 4, and finally, paper is concluded in Section 5. 2. Material and Methods 2.1. Datasets and tasks In this study, we examine three datasets — CAMELYON17 [23], MIDOG22 [24], and HISTOPANTUM — each chosen for its specific challenges and domain shifts to facilitate a comprehensive evaluation of DG algorithms applied to various classification tasks in CPath. 2.1.1. CAMELYON17 With the CAMELYON17 grand challenge on the detection of breast cancer metastases in sentinel lymph nodes, a dataset with the same name has been released which comprises Whole Slide Images (WSIs) of lymph node resections in breast cancer patients and their corresponding lesion-level annotations [23]. The provided annotations have been leveraged for the extraction of a patch-level dataset of metastatic and normal images of breast cancer. In particular, we used the CAMELYON17 dataset part of the WILDS toolbox [25] (which is designed to test machine-learning models against significant distribution shifts) to allow for the reproducibility of the results. 2 (A) CAMELYON17 Metastasis Normal (B) MIDOG22 Mitosis Mimicker (C) HISTOPANTUM Tumour Normal Figure 2: Tasks and datasets used in the benchmarking process: (A) Breast cancer metastasis detection leveraging Camleyon17 dataset [23], (B) Mitosis detection in MIDOG22 dataset [24], and (C) tumor detection in our proposed HISTOPANTUM dataset. For every dataset, an example from each domain and class is provided. All the tasks are designed as a binary classification task, where the name and population of positive and negative classes are shown in red and blue color bars, respectively. The hatched region in each bar represents the fraction of samples used to generate small datasets (see Section 2.1.4). CAMELYON17 is gathered from various medical centers in the Netherlands, including Radboud University Medical Center (RUMC), Canisius-Wilhelmina Hospital (CWZ), University Medical Center Utrecht (UMCU), Rijnstate Hospital in Arnhem (RST), and the Laboratory of Pathology East-Netherlands (LPON), each representing a unique domain. CAMELYON17 comprises 455,953 image patches, categorized into metastasis and non-metastasis (normal) classes, with each patch measuring 96 ×96 pixels at a resolution of 0.5 microns per pixel (mpp). A major challenge presented by CAMELYON17 is the covariate shift, primarily caused by variations in imaging equipment and procedures across different centers. These variations are evident in the noticeable differences in color and texture among images from various sources as shown in Fig. 2A. Despite covariate shift, the dataset is well balanced in label space, with an equal number of tumor and non-tumor patches in each domain, effectively eliminating any potential prior shift in label distribution. The objective nature of the classification task ensures that there is no posterior shift, focusing the analysis solely on addressing the implications of covariate shifts. 2.1.2. MIDOG22 The second task, mitosis detection, involves classifying mitotic figures versus mimickers (cells of other types that are very similar to mitotic figures in appearance), which is a binary classification task previously explored in the literature [18, 11, 26, 27]. For this purpose, we utilize the MIDOG22 dataset [24], which comprises five domains: Canine Lung Cancer, Human Breast Cancer, Canine Lymphoma, Canine Cutaneous Mast Cell Tumor, and Human Neuroendocrine Tumor. From the original MIDOG22 dataset and based on the annotations provided, we extract 20,552 image patches, each sized 128 ×128 pixels at a resolution of 0.25 mpp, and categorize them into mitosis and mimicked classes. MIDOG22 is particularly challenging for achieving DG due to the presence of all four types of DS: covariate shift is evident as the images come from different centers using various scanners, leading to variations in color schemes. There is a prior shift, as different labels are distributed variably across domains, clearly shown in the dataset Fig. 2B. The task also involves a posterior shift due to the highly subjective nature of mitosis labeling. Additionally, class-conditional shifts occur because different tumor types and species influence the appearance of non-mitotic regions in the images, significantly varying from one domain to another. The complexity of this dataset makes it a rigorous test bed for assessing the DG capabilities of different algorithms in CPath. 2.1.3. HISTOPANTUM The last task we address is pan-cancer tumor detection, leveraging the HISTOPANTUM dataset that we were releasing in this study. This dataset captures four different cancer types: Colorectal (CRC), Uterus (UCEC), Ovary (OV), and Stomach (STAD), collectively referred to as four domains. During data curation, we source 40 WSIs for each cancer type from its related study in The Cancer Genome Atlas Program (TCGA) [28]. For sampling, we make sure to include a variety of tumor subtypes (adenocarcinoma and mucinous carcinoma), genders, ethnicities, and centers in order to make the dataset as diverse as possible. Then, an experienced pathologist (FS) meticulously annotates tumor and non-tumor regions in the slides, which are used to extract tumor and non-tumor patches from the WSIs to form the HISTOPANTUM dataset. The HISTOPANTUM dataset includes 281,142 patches, each 512 ×512 pixels at a resolution of about 0.5 mpp which are subsequently resized to 224×224 pixels during training and evaluation. HISTOPANTUM patches are classified into two 3 classes: tumor and non-tumor (normal). This dataset presents three significant types of DS. Firstly, a covariate shift arises due to images being sourced from different centers using different slide preparation and scanning devices, introducing notable variations in color and stain schemes. Secondly, a prior shift is observed with a distinct distribution of classes across domains, as depicted in Fig. 2C. Lastly, the class-conditional shift is evident as different cancer types influence the appearance of tumor region in the images (morphology of tumor cells varies significantly from one tumor to another), while non-tumor regions remain relatively consistent across domains (for example, the morphology of stromal and inflammatory regions are very similar across different cancer types). Notably, this dataset does not suffer from the posterior shift, thanks to the objective nature of the labeling process, eliminating subjectivity in tumor detection. We are making the HISTOPANTUM publicly available as a benchmark for pan-cancer tumor detection 2. 2.1.4. Subsampled datasets In addition to the primary experiments, we conduct a series of tests to understand how different DG algorithms perform under a low-data budget scenario. This analysis is crucial for applications where data availability is limited. For this purpose, we create smaller versions of the original datasets, maintaining similar distributions but significantly smaller populations. The reduced (small) datasets are generated by randomly sampling the following percentages of each class in the original datasets (as hatched regions shown in the bar plots of Fig. 2): •sCAMELYON17: 1% of the original CAMELYON17 dataset (N=4,560). •sHISTOPANTUM: 3% of the original HISTOPANTUM dataset (N=8,434). •sMIDOG22 : 30% of the original MIDOG22 dataset (N=6,166). The sampling percentage in each dataset is set to considerably reduce dataset size while keeping enough samples for convergence of DG algorithms. The same experimental setup, model selection strategy, and algorithms used in the original experiments are applied to these smaller datasets. This approach allows us to directly compare the performance and generalization capability of the algorithms in both ‘large dataset’ and ‘small dataset’ scenarios. 2.2. Algorithms We utilize DomainBed [22] as our main benchmarking tool because it offers a robust and well-tested platform for fair and reproducible comparison of different DG algorithms. Furthermore, DomainBed is well maintained and contains the most number of state-of-the-art DG algorithms implemented in comparison to other DG benchmarking tools (such as DeepDG 2Data is being uploaded to a web server. Please check this link for updates: https://github.com/mostafajahanifar/HistoDomainBed Figure 3: Benchmarked algorithms categorized into different domain generalization methodologies, as introduced in [18] [29]). Using DomainBed, we can test different DG algorithms while also controlling algorithms and training hyperparameters. As well as the DomainBed implemented algorithms, we also investigate the two CPath-specific algorithms, namely stain normalization and stain augmentation, and a selfsupervised learning (SSL) based algorithm in the same structured experimental setup of DomainBed. Furthermore, we have also added the F1-score evaluation metric to the platform which originally included only the Accuracy metric. The code base for our updated DomainBed platform, called HistoDomainBed, is available at: https://github.com/mostafajahanifar/ HistoDomainBed. All the algorithms in our experiments use a standard ResNet50 [30] model for feature extraction due to its proven generalization capabilities and popularity. The preferred model selection strategy in our work is the “training-domain validation set”, recognized for its effectiveness in different scenarios and datasets as shown previously in [22]. More information on how we performed cross-validation experiments using DomainBed is given in Section 2.3. In the rest of this section, we introduce the DG algorithms investigated in this work. Explaining the methodology of each algorithm is outside of the scope of this work, although we have categorized these algorithms into 6 distinct categories of DG methods based on their working principles and the introduction of the categories in [18]. Domain alignment techniques bridge domain gaps by harmonizing feature representations, employing methods like stain normalization and generative models. Data augmentation enhances model generalization through image transformations and generative networks. Metalearning enables quick adaptation to new domains or tasks using techniques like MAML [31]. Tailored model design strate4 gies leverage the unique characteristics of histopathology images with specialized network architectures and loss functions. Pretraining strategies use self-supervised, unsupervised, and semi-supervised learning to enhance feature encoding and generalizability. Regularization strategies prevent overfitting and improve performance on unseen data by introducing constraints and penalties. Domain separation learns disentangled domainspecific and domain-agnostic features. A chart showing the investigated DG algorithms and their related category is presented in Fig. 3. For more information on these categories or related methods, please refer to [18] or the respective cited articles. 2.2.1. DomainBed algorithms DomainBed [22], developed by the Facebook Research group, is a comprehensive PyTorch suite. At the time of writing this manuscript, it encompassed support for 27 DG algorithms, 10 computer vision datasets, and one CPath dataset (Camelyon17 WILDS [25]). The toolkit includes algorithms such as Empirical Risk Minimization (ERM) [32], Interdomain Mixup (Mixup) [33], Group Distributionally Robust Optimization (GroupDRO) [34], Conditional Domain Adversarial Neural Network (CDANN) [35], Learning Explanations that are Hard to Vary (AND-Mask) [36], Deep CORAL (CORAL) [37], Selfsupervised Contrastive Regularization (SelfReg) [38], Marginal Transfer Learning (MATL) [39], and Adaptive Risk Minimization (ARM) [40]. It also supports Invariant Risk Minimization (IRM) [41], Domain Adversarial Neural Network (DANN) [42], Style Agnostic Networks (SagNet) [43], Learning Representations that Support Robust Transfer of Predictors (TRM) [44], Optimal Representations for Covariate Shift (CAD and CondCAD) [45], Representation Self-Challenging (RSC) [46], Maximum Mean Discrepancy (MMD) [47], Outof-Distribution generalization with Maximal Invariant Predictor (IGA) [48], Variance Risk Extrapolation (VREx) [49], and Invariance Principle Meets Information Bottleneck for Out-ofDistribution generalization (IB-ERM) [50]. Additional algorithms include Empirical Quantile Risk Minimization (EQRM) [51], Spectral Decoupling (SD) [52], Quantifying and Improving Transferability in Domain generalization (Transfer) [53], Smoothed-AND mask (SAND-mask) [54], Meta-Learning Domain generalization (MLDG) [55], and Invariant Causal Mechanisms through Distribution Matching (CausIRL with CORAL or MMD) [56]. 2.2.2. CPath-specific algorithms Ideally, the same tissue specimens, stained in different laboratories, should yield identical results, but this ideal is often unattainable. Stain variation can arise from differences in slide scanners, stain quality and concentration, and staining procedure [57]. Pathologists can easily disregard irrelevant features (such as stain variation) in a WSI that do not impact their diagnosis. However, deep learning models sometimes struggle with this task [58]. Stain Normalization. A preprocessing step that aligns the color distributions of histology images to a reference image, countering discrepancies from varied staining procedures and scanners. Stain normalization techniques range from linear scaling and histogram matching to advanced methods by Ruifrok [59], Macenko [60], and Vahadane [61]. These methods adjust the stain matrix of source images while maintaining their stain concentrations, ensuring color consistency without altering structural details. In this work, we compare Macenko [60] stain normalization algorithm implemented in TIAToolbox [62]. To this end, we normalize the stain of all images in all datasets offline, utilizing the same reference image, to maintain training efficiency. Stain Augmentation (StainAug). A method that involves decomposing RGB histology images into stain components, perturbing them, and recomposing the images to introduce variability. This technique has proven effective in numerous tasks and helps improve model generalizability [18, 58, 63]. In HistoDomainBed, stain augmentation was randomly done on the fly using TIAToolbox [62]. To this end, the Macenko method [60] was used to extract the stain matrix and components from the RGB image. It is important to note that both Macenko and StainAug algorithms are employed on top of the Empirical Risk Minimization (ERM) [32] approach. In other words, Macenko refers to the scenario where we use ERM on stain-normalized images and StainAug is the scenario where the stain augmentation technique is used during the training of the ERM model on original images. 2.2.3. Self-supervised learning (SSL) For our SSL experiments, we use a ResNet50 model pretrained on histology images to start each run instead of initializing the model with ImageNet weights. In particular, we choose a foundation model from the work of Kang et al. [64] which has a ResNet50 architecture and is pretrained on 19M image patches from different studies of TCGA using Barlow Twin self-supervised learning algorithm [65]. The utilized SSL algorithm also works on top of the ERM algorithm in our HistoDomainBed platform. 2.3. Cross-validation and model selection Gulrajani et al. [22] meticulously examined three distinct model selection scenarios, which are crucial for determining how the performance of a model on a validation set can guide the selection of the best training epoch for application on unseen test sets. The scenarios we explore are: 1. Training-domain validation set: This involves pooling a validation set from all training domains, a standard practice of training/validation split across domains. The bestperforming model of the validation set is then tested on hold-out test domains. 2. Leave-one-domain-out validation: In this scenario, one of the training domains is reserved for validation. The model that performs best on this holdout domain is then re-trained on all domains before being applied to the test domains. 5 3. Test-domain validation set (oracle): This scenario selects the model that maximizes accuracy on a validation set mirroring the test domain’s distribution, albeit with limited queries. Although this is theoretically not a valid model selection method due to its reliance on access to the test domain, in the original DomainBed [22] it was included for comparative analysis. Based on findings of the DomainBed study [22], the “training-domain validation set” model selection scenario consistently yielded the best results on different datasets and using different algorithms. Thus, we have chosen to adopt this model selection strategy for our experiments to allow us to minimize computational demands and focus our resources on evaluating the effectiveness of the DG algorithms themselves, rather than delving into various model selection techniques. In our model selection strategy, we allocate 20% of the training data for validation purposes. For robust cross-validation, one domain was systematically left out for testing, and this process was repeated across the number of domains available in each task (Run 1, 2, 3, ...). An illustration of the utilized cross-validation process is given in Fig. 1E, where there are 4 domains and therefore 4 runs, and in each run data from three domains are used for training and validation (80%-20%), and 1 domain is left out for testing. All the metrics reported in this work are based on results of experiments on unseen test domains. To ensure the reliability of our results, we varied the hyperparameters (such as batch size, learning rate, and algorithmspecific parameters) randomly three times for each experiment. Each set of hyperparameters was then used to conduct three independent runs, leading to a total of nine training runs for each domain and method combination. The working range of algorithms’ hyperparameters is selected based on convergence experiments done for each algorithm beforehand. In summary, this comprehensive cross-validation approach resulted in a total of 7,560 training-validation runs for both full and small datasets (Pd∈datasets 9×30 ×Nd domains), illustrating the extensive scale and rigorous nature of our experimental design. In all the runs, we utilize an ImageNet-pretrained ResNet50 model [30] (except for SSL algorithm which uses histologypretrained weights as a starting point), trained for approximately 30 epochs using an Adam optimizer. All the experiments are performed using 8 NVidia Tesla V100 GPUs on a DGX2 machine. 3. Results This section presents the performance metrics of different algorithms across various datasets. The metrics reported are binary F1 score and accuracy. We added the F1 score to HistoDomainBed to account for the scenarios where the data is significantly imbalanced, rendering accuracy a sub-optimal metric. Using Accuracy and F1 Score, we can comprehensively understand our models’ performance across different domains, ensuring that our evaluation is robust and reliable. 3.1. Results for full datasets We present the performance metrics of various algorithms across full-scale datasets in Table 1. Accuracy and F1 metrics are reported for each dataset (task) separately as well as for the average performance across all the tasks. The rows in Table 1 are ordered by the average F1 score across all tasks. In each column, cells are colored from red to green based on the performance values, red indicating worse and green indicating better performance. The majority of methods exhibited similar performance, with average F1 scores ranging from 81% to 85%, except for the top 2 algorithms. SSL [64] and StainAug [58] methods consistently outperform all other methods on average, both in terms of F1 score (87.7%, 86.5%) and accuracy (88.9%, 87.4%). This advantage is particularly pronounced in the MIDOG22 and HISTOPANTUM datasets. The third-ranking algorithm, ARM [40], has done relatively worse on the MIDOG22 dataset. Furthermore, the Macenko stain normalization algorithm ranked 6th, outperforming 24 other DG algorithms, but not as good as StainAug. All algorithms performed exceptionally well on the CAMELYON17 dataset (F1>90%), which can be attributed to the abundance of data to help the model generalize better and the relatively simpler nature of the problem. The high performance across algorithms indicates that the CAMELYON17 dataset poses fewer challenges in terms of DS. The DS in CAMELYON17 is primarily stain variation between different hospitals, making StainAug one of the best candidates to improve DG in this dataset (as Table 1 shows the highest F1 96.1% for StainAug). The tasks associated with the MIDOG22 and HISTOPANTUM datasets are more challenging, involving more significant domain shifts. Specifically, the performance metrics in the MIDOG22 task are generally lower compared to other tasks, reflecting the increased difficulty. Notably, the baseline algorithm, ERM[32], demonstrated strong performance (ranked 17th), comparable to other SOTA methods. This suggests that combining simple augmentations with the ERM approach is sufficient to train a robust classifier. On the other hand, SANDMask [54] and IGA [48] algorithms struggled to converge on the CAMELYON17 and all datasets, respectively, indicating potential issues in handling domain shifts or complexities in these tasks. 3.2. Results for Sub-sampled (small) datasets To investigate how DG algorithms perform under a low-data budget scenario, we subsample each dataset at different rates to create smaller datasets (as explained in Section 2.1.4) and repeat the experiments. The results for these experiments are reported in Table 2. Interestingly, SSL and StainAug algorithms are still among the top 3 performing algorithms with Transfer algorithm [53] place on the second rank and achieving F1 of 82.8% (almost on a par with StainAug, F1=82.7%). However, SSL considerably outperforms other algorithms by gaining the F1 score of 85.4%, showing an advantage in small-dataset scenarios as has been 6 Table 1: Benchmarking results for full-scale datasets. In each column, cells are colored from red to green representing worst to best performance. CAMELYON17 MIDOG22 HISTOPANTUM Average Algorithm ACC F1 ACC F1 ACC F1 ACC F1 SSL 95.4±0.2 95.2±0.2 79.9±0.2 76.1±0.6 91.2±0.6 91.7±0.5 88.9 87.7 StainAug 96.4±0.9 96.1±0.9 79.9±0.3 76.0±0.4 85.9±0.0 87.3±0.0 87.4 86.5 ARM 94.7±0.3 94.5±0.3 78.5±0.3 73.4±0.2 87.6±0.8 88.6±0.8 87.0 85.5 CausIRL CORAL 93.3±0.6 92.8±0.6 78.9±0.3 74.9±0.5 85.4±0.2 87.1±0.2 85.9 84.9 SelfReg 94.6±0.3 94.3±0.4 79.2±0.3 74.6±0.4 83.0±0.4 85.2±0.4 85.6 84.7 Macenko 93.3±0.3 92.9±0.2 77.8±0.3 73.9±0.1 84.4±0.0 86.2±0.0 85.2 84.3 Transfer 93.7±0.4 93.6±0.7 78.5±0.3 74.2±0.6 82.6±1.2 83.9±0.8 84.9 83.9 TRM 93.5±0.3 93.2±0.4 79.4±0.4 74.4±0.3 83.5±1.6 84.2±1.9 85.5 83.9 IB ERM 94.0±0.1 94.0±0.2 79.5±0.2 74.2±0.3 81.4±0.6 82.8±0.5 85.0 83.7 CondCAD 93.5±0.1 93.3±0.1 78.7±0.3 74.6±0.6 80.9±0.7 82.7±1.2 84.3 83.6 ANDMask 93.7±0.2 93.5±0.3 78.6±0.2 74.5±0.3 78.1±1.0 81.8±0.6 83.5 83.2 Mixup 94.8±0.2 94.5±0.2 79.4±0.3 74.7±0.3 78.5±0.8 80.4±0.0 84.2 83.2 EQRM 95.3±0.1 95.1±0.1 80.2±0.1 75.2±0.1 77.4±0.2 78.9±0.4 84.3 83.1 CausIRL MMD 94.4±0.1 94.2±0.1 78.7±0.6 72.9±1.9 79.0±2.3 81.8±1.8 84.0 83.0 CORAL 96.0±0.3 95.9±0.3 79.2±0.4 74.9±0.2 77.1±0.3 78.0±0.9 84.1 82.9 VREx 94.2±0.6 93.8±0.7 79.3±0.2 74.3±0.2 79.7±0.7 80.4±0.8 84.4 82.8 ERM 95.6±0.0 95.4±0.0 79.1±0.2 74.7±0.2 76.8±0.1 77.6±0.9 83.8 82.6 GroupDRO 94.6±0.6 94.4±0.8 78.8±0.1 74.5±0.4 77.8±0.3 78.6±0.1 83.7 82.5 CAD 93.7±0.2 93.4±0.3 77.6±0.2 73.5±0.3 77.7±3.1 80.2±2.8 83.0 82.4 MTL 93.8±0.0 93.4±0.0 78.9±0.4 73.9±0.7 78.1±1.7 79.9±2.3 83.6 82.4 MLDG 95.2±0.3 94.9±0.4 78.9±0.3 74.2±0.4 75.5±0.2 77.6±0.4 83.2 82.3 RSC 94.1±0.1 93.7±0.2 78.8±0.4 75.7±0.2 76.6±1.2 77.3±0.4 83.2 82.2 SD 95.5±0.3 95.3±0.4 78.5±0.6 73.9±0.3 78.1±0.9 77.3±0.9 84.0 82.2 IRM 94.9±0.4 94.7±0.5 78.1±0.5 72.9±0.5 77.6±0.4 78.9±0.3 83.5 82.1 CDANN 91.6±1.2 90.9±1.4 78.8±0.6 74.3±0.3 80.1±1.9 80.2±2.1 83.5 81.8 SagNet 93.6±0.0 93.2±0.0 78.7±0.2 74.8±0.3 76.2±1.0 77.0±0.9 82.8 81.6 DANN 91.8±1.8 91.5±1.8 79.2±0.3 74.2±0.4 78.1±0.5 77.0±0.8 83.0 80.9 MMD 94.2±0.4 94.3±0.2 75.3±1.8 69.0±2.7 77.8±0.1 78.8±0.1 82.4 80.7 IGA 55.5±0.9 67.7±0.3 49.4±3.5 60.9±0.0 52.8±2.5 67.1±0.9 52.6 65.2 SANDMask 44.6±4.9 37.1±13.7 79.4±0.6 74.7±0.4 77.9±0.3 79.0±0.9 67.3 63.6 shown before for other algorithms based on self-supervised learning [66]. The majority of SSL superiority is owed to the performance of sCAMELYON17 and sHISTOPANTUM datasets. However, on the hardest DG task using the sMIDOG22 dataset, Transfer algorithm [53] gains the highest F1 of 77.9%, considerably outperforming SSL. StainAug does not perform as high as SSL on sHISTOPANTUM, nevertheless, it keeps a good performance on sMIDOG22 and sCAMELYON. On the sHISTOPANTUM dataset, except for the SSL algorithm, most of the other algorithms perform on par. The baseline ERM algorithm archives impressive results in the small-scale dataset, outperforming all other algorithms on the sHISTOPANTUM dataset excluding SSL (F1=87.1%) and very good performance in the other small datasets. On the other hand, IGA [48] still struggles to converge on small datasets whereas the SANDMask [54] algorithm works relatively well on sCAMELYON17 dataset although it could not converge on large-scale CAMELYON17 dataset. 3.3. Domain-level performance We comprehensively evaluate the performance of various algorithms across different domains, the results of which are presented in Fig. 4 in the form of bar plots. In Fig. 4, each domain in every dataset is represented by a uniquely colored bar. The average performance of all algorithms over each domain is indicated by horizontal dashed lines in the same color as the domain. In Fig. 4, algorithms ”IGA” and ”SANDMask” are excluded due to poor performance and to better visualize the working performance range of other algorithms. CAMELYON17. Performance across centers is generally high, with average F1 scores around 93% to 96%. This is highlighted by the closely clustered bars and the horizontal dashed lines at the top of the graph (only a 3.5% difference between the best and worst domains). Results for CWZ and RUMC centers are consistently among the highest scores, presumably because slides from these two centers were scanned using the same scanner (at RUMC center) and there is a lower domain shift between these two datasets, hence a model trained on the data from one of these domains will perform reasonably good 7 Table 2: Benchmarking results for small-scale datasets. In each column, cells are colored from red to green representing worst to best performance. sCAMELYON17 sMIDOG22 sHISTOPANTUM Average Algorithm ACC F1 ACC F1 ACC F1 ACC F1 SSL 94.9±0.3 93.2±1.0 76.6±0.7 71.1±0.7 91.6±0.5 92.0±0.5 87.7 85.4 Transfer 90.1±0.8 88.7±0.7 77.9±0.3 73.7±0.6 85.1±0.2 86.0±0.1 84.3 82.8 StainAug 90.2±0.0 88.4±0.8 77.0±0.2 72.2±1.0 86.4±0.1 87.4±0.1 84.5 82.7 Macenko 92.0±0.5 90.3±0.7 75.7±0.7 70.5±0.7 85.6±0.2 86.9±0.1 84.5 82.6 VREx 90.2±0.7 88.4±1.4 75.7±0.2 71.8±0.3 86.6±1.1 87.3±1.0 84.2 82.5 SagNet 90.1±0.3 88.9±0.2 75.6±0.7 72.2±0.5 84.7±0.4 85.5±0.3 83.5 82.2 IB ERM 87.9±1.0 88.3±0.6 76.9±0.5 71.1±1.7 86.2±0.4 87.1±0.3 83.7 82.1 ERM 88.2±0.4 85.8±0.4 77.3±0.3 72.6±0.5 87.1±0.1 87.3±0.3 84.2 81.9 GroupDRO 88.1±1.6 87.7±1.0 77.0±0.2 71.1±0.1 85.8±0.2 86.0±0.6 83.7 81.6 ANDMask 90.9±0.6 89.2±0.6 75.8±0.1 71.5±1.0 82.6±0.2 83.7±0.0 83.1 81.5 TRM 87.9±0.7 86.0±0.7 77.6±0.2 71.1±0.6 86.8±0.3 87.2±0.1 84.1 81.4 CDANN 88.9±0.2 88.0±0.2 76.0±0.8 70.1±0.0 84.5±0.2 85.7±0.9 83.1 81.3 ARM 88.3±0.8 85.7±0.7 77.1±1.0 71.3±0.6 86.1±0.6 86.3±1.0 83.8 81.1 EQRM 87.3±0.1 85.0±0.2 77.1±0.2 71.7±0.2 85.0±1.1 86.2±0.7 83.1 81.0 SelfReg 87.2±1.7 84.4±2.8 76.8±0.2 71.6±0.1 86.1±0.4 87.1±0.5 83.4 81.0 MTL 89.3±1.1 87.5±1.6 77.6±0.1 71.4±1.1 83.5±1.1 83.9±0.5 83.5 80.9 MLDG 87.4±0.6 86.6±0.5 76.4±0.3 71.5±1.4 82.5±1.5 84.2±0.7 82.1 80.8 CausIRL CORAL 86.8±0.9 84.8±1.5 76.8±0.2 69.8±0.6 86.4±0.8 87.2±0.8 83.3 80.6 SANDMask 88.0±0.6 86.8±0.1 76.0±0.3 70.8±1.2 83.3±0.2 84.1±0.5 82.4 80.6 SD 86.0±1.1 82.9±1.5 76.8±0.6 72.8±0.0 85.4±0.5 86.2±0.3 82.7 80.6 CondCAD 87.0±0.0 84.8±0.1 77.5±0.2 70.1±0.7 85.7±1.6 86.6±1.4 83.4 80.5 Mixup 86.8±0.6 84.5±1.1 76.7±0.4 70.7±1.4 85.1±0.2 86.0±0.0 82.9 80.4 RSC 87.0±2.0 84.6±2.8 77.2±0.6 70.2±0.4 85.9±1.9 86.1±0.9 83.4 80.3 CAD 87.7±2.0 85.7±2.8 74.6±0.4 69.8±0.1 84.3±0.4 84.9±0.1 82.2 80.1 CORAL 85.3±0.5 83.2±1.8 77.5±0.6 69.9±0.1 86.0±0.1 85.9±0.3 82.9 79.7 IRM 88.1±0.4 86.8±1.2 76.0±0.2 66.0±1.7 85.3±1.9 85.8±1.9 83.1 79.5 DANN 84.9±1.7 79.7±1.3 74.6±0.9 67.7±1.8 83.2±0.6 84.4±0.4 80.9 77.3 CausIRL MMD 86.6±0.7 83.8±1.5 64.3±8.8 59.2±7.7 85.5±0.9 86.3±1.2 78.8 76.4 MMD 89.7±0.7 88.6±1.0 62.9±5.6 43.0±14.4 85.0±0.2 86.6±0.7 79.2 72.7 IGA 52.9±0.3 63.9±1.0 49.3±1.1 60.9±0.2 57.1±0.8 67.9±0.6 53.1 64.2 for the other domain too. Furthermore, from Fig. 4B it is evident that StainAug [58], SD [52], and ERM [32] are among the most consistent algorithms over different domains whereas the performance of DAN [42], CDANN [35], CausIRL MMD [56], and Macenko [60] changes considerably for different CAMELYON17 domains. MIDOG22. In this dataset, performance varies more significantly across the different domains (11% difference between highest and lowest average performance). In particular, based on the F1 score, the human endocrine cancer is the hardest domain (F1=66%) and the canine cutaneous mast cell tumor is the easiest (F1=77%) for out-of-domain mitosis detection. Interestingly, the patterns of different algorithms’ performances over different domains are similar, i.e., best to worst performing domains being canine cutaneous mast cell, human breast, canine lung, canine lymphoma, and human neuroendocrine. The worst performance on human neuroendocrine can be justified by three reasons: the tumor type is completely different from all other domains (class-conditional shift), slides in this domain were scanned with a Hamamatsu NanoZoomer XR scanner unlike other datasets that used Aperio or 3DHistech scanners (covariate shift), and the label distribution in this domain is significantly different from other domains (prior shift). On the other hand, the performance on the canine cutaneous mast cell domain is higher because in that case 2 other similar tumor types from the same species and scanner are utilized for model training, hence model seeing similar data can generalize better to this unseen domain. In terms of consistency over different domains, although the accuracy metric shows consistency for most algorithms (such as SSL in Fig. 4A), F1 score values tell another story where consistency over different domains drastically decreases for all algorithms. This is mostly due to an imbalanced class population across different domains in the MIDOG22 dataset (see Fig. 2B). HISTOPANTUM. Average accuracy over different domains in around the same range for CRC, OV, and STAD domains (around 81%) with a notable accuracy drop for the UCEC domain (77%). The F1 score on the UCEC domain is also the worst among all domains (77% compared to 80-85% for other domains). This can be accounted for by a prior shift in the 8 label distribution when comparing the UCEC domain with others. CRC domain gets the highest performance over because it shares a similar label distribution and tissue phenotypes appearance (especially with the STAD domain). In HISTOPANTUM, the SSL algorithm is the best performing and consistent algorithm across all domains, mostly because the utilized SSL algorithm has already seen TCGA slide during its pertaining procedure. Furthermore, StainAug, ARM [40], CausIRL CORAL [56] and Transfer [53] algorithms also show decent consistency and performance over different domains. 4. Discussion Benchmarking DG algorithms is essential for evaluating their performance across diverse datasets and scenarios to provide insights that can increase model robustness in real-world applications. In this study, we benchmarked various DG algorithms with different working principles, including SOTA algorithms from DomainBed collection [22], self-supervised learning [64], and pathology-specific techniques [60, 58], on datasets with different DS and size properties. Our unified and fair benchmarking process reported both accuracy and F1 scores for comprehensive evaluation through robust cross-validation experiments. At a glance, the average best-performing algorithms are SSL and StainAug. SSL methods excelled especially on the HISTOPANTUM dataset, which is the main reason SSL ranked first in the full-scale dataset (Table 1) and small-scale dataset (Table 2) scenarios. This is mostly because SSL pertaining was done on an extensive set of patches extracted from TCGA, the same source used to curate the HISTOPANTUM dataset. Although the same image patches and labels are not shared between HISTOPANTUM and the dataset used during pertaining of SSL, the SSL has indirect access to the test data in HISTOPANTUM and already has seen the possible variation of image data within HISTOPANTUM. Therefore, the evaluation of SSL on HISTOPANTUM is more of a “domain adaptation” exercise rather than “domain generalization” and comparing its performance (F1=91.7%) with the rest of the algorithms (such as ARM and StainAug with F1 scores of 88.6% and 87.3%, respectively) is not fair. Nevertheless, SSL shows excellent performance in CAMELYON17, MIDOG22, and the low-databudget scenario of the sCAMELYON17 dataset. However, that is not the case with the sMIDOG22 dataset where there are all sorts of DS and data shortage problems. CPath-specific algorithms–stain augmentation (StainAug) and stain normalization (Macenko)–outperform most complex DG algorithms in the literature while being simple and easy to implement. In particular, StainAug excelled in the CAMELYON17 and MIDOG22 datasets. StainAug helps the model to learn more stain-invariant feature representations from the image during the training by randomly tweaking the stain information on the fly. Considering stain variation as a confounding factor, StainAug can be thought of as a causal approach for DG in CPath as it helps to learn features that are irrelevant to stain variation, hence achieving outstanding results on the CAMELYON17 dataset where the main DS is covariate shift (or changes in stain appearance). Macenko stain normalization algorithm has also shined in the CAMELYON17 and HISTOPANTUM tasks where stain variation is the dominant DS. Specifically, in the small dataset scenario of sCAMELYON, Macenko outperforms all other algorithms (except SSL) and ranks second. However, we should note that stain normalization methods add another preprocessing step to every CPath pipeline that uses them and sometimes are not stable in practice [67, 18]. Therefore, we suggest using stain augmentation over stain normalization when training models on H&E images. To gain a deeper understanding of how algorithm performance varies between small and full dataset regimes, we plot their rankings based on F1 and Accuracy metrics over small and full datasets in Fig. 5. Each point on the plot represents an algorithm, with its x-axis position indicating its rank in the full dataset regime and its y-axis position indicating its rank in the small dataset regime. The color of each point reflects the performance difference between the two dataset sizes. Algorithms located along the diagonal line performed similarly in both the full and small dataset regimes. Examples of such algorithms based on F1 score in Fig. 5 include SSL [64], StainAug [58], Macenko [60], ANDMask [36], EQRM [51], and RSC [46] among others. Conversely, some algorithms showed improved performance rankings on smaller datasets, evident from their positions below the diagonal line. Notable examples include Transfer [53], VREx [49], and SagNet [43]. Ideally, algorithms that are versatile and perform consistently across different data conditions are desirable. Therefore, algorithms close to the diagonal line and lower left corner of plots in Fig. 5 (such as SSL and StainAug) are more suitable ones to investigate regardless of the size of the dataset at hand. As mentioned before, in CAMELYON17 and HISTOPANTUM, most algorithms perform well due to the large amount of data and simpler domain shifts. However, the level of performance for all algorithms drops on MIDOG22, which encompasses all kinds of DS. In the large-scale MIDOG22 dataset, SSL and StainAug showed the best performance in terms of F1 score, whereas in the small-scale sMIDOG22 dataset, SSL is not among the top-performing algorithms. Instead, Transfer [53] and SD [52] algorithms achieve high F1 scores along with StainAug. Alternatively, focusing on the accuracy metrics obtained for the MIDOG22 dataset, we can see that the EQRM algorithm [51] is very promising. Although we have found that SSL and StainAug are generally good options to consider for DG applications, there is no one “best” algorithm that fits all the situations. Depending on the dataset size, the types of DS in the dataset, and the difficulty of the task different DG algorithms can perform differently. Notably, the baseline ERM algorithm [32] (which simply uses data from different domains in the mini-batches with standard data augmentation during training) performed consistently better than most DG algorithms, indicating that simple methods can be effective if implemented properly. This is in line with the findings of other works that investigated various DG algorithms [68, 22]. This underscores the importance of careful experiment design and incorporating well-performing baseline algorithms when looking for an optimal DG algorithm in any application. 9