FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentation
Full text
FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentation Dong Zhao1, Jinlong Li2, Shuang Wang1B, Mengyao Wu, Qi Zang1B, Nicu Sebe2, Zhun Zhong3 1School of Artificial Intelligence, Xidian University, Shaanxi, China 2Department of Information Engineering and Computer Science, University of Trento, Italy 3School of Computer Science and Information Engineering, Hefei University of Technology, China Abstract Vision Foundation Models (VFMs) excel in generalization due to large-scale pretraining, but fine-tuning them for Domain Generalized Semantic Segmentation (DGSS) while maintaining this ability remains a challenge. Existing approaches either selectively fine-tune parameters or freeze the VFMs and update only the adapters, both of which may underutilize the VFMs’ full potential in DGSS tasks. We observe that domain-sensitive parameters in VFMs, arising from task and distribution differences, can hinder generalization. To address this, we propose FisherTune, a robust fine-tuning method guided by the Domain-Related Fisher Information Matrix (DR-FIM). DR-FIM measures parameter sensitivity across tasks and domains, enabling selective updates that preserve generalization and enhance DGSS adaptability. To stabilize DR-FIM estimation, FisherTune incorporates variational inference, treating parameters as Gaussian-distributed variables and leveraging pretrained priors. Extensive experiments show that FisherTune achieves superior cross-domain segmentation while maintaining generalization, outperforming both selectiveparameter and adapter-based methods. 1. Introduction Vision Foundation Models (VFMs), such as CLIP [46], DINOv2 [5], and EVA02 [15], have emerged as powerful tools in computer vision, achieving remarkable generalization across diverse downstream tasks, including crossdomain perception [35,53,56,71,76], few-shot [37,65,70, 79] and zero-shot perception [26,31,54,72]. Pre-trained on massive data fields, VFMs encapsulate rich visual repreThis work was supported in part by the National Natural Science Foundation of China No. 62271377, the Key Research and Development Program of Shannxi Program No. 2021ZDLGY0106 and No. 2022ZDLGY0112, the Key Scientific Technological Innovation Research Project by Ministry of Education, the MUR PNRR project FAIR (PE00000013) funded by the NextGenerationEU and the EU Horizon projects ELIAS (No. 101120237) and AI4Trust (No. 101070190). BCo-corresponding author. VFMs VFMs VFMs (b) Selective Tuning (a) Adapter Tuning : Domain-sensitive parameters (c) Fisher Tuning (Ours) : Extra adapters : Handmade parameters Source Data Source Data Source Data Figure 1. Comparison of principles for different VFM adjustment methods: (a) tuning by adapter insertion [60,66], (b) tuning by manually selected [57] or automatically selected parameters [55], (c) our method for tuning domain-sensitive parameters. sentations that can be transferred to numerous applications with minor adaptation [61,66]. Despite this, when it comes to Domain Generalized Semantic Segmentation (DGSS), where the goal is to segment unseen domain images without explicit access to their training data, effectively fine-tuning VFMs while preserving their strong generalization capabilities remains an open challenge. Existing methods for adapting VFMs to DGSS tasks typically involve fine-tuning via adapter layers to remap pretrained tokens, as shown in Fig. 1(a) [60,66]. While this approach reduces overfitting, it does not fully leverage the internal representations of the VFM, as the core content of the model remains unchanged. Furthermore, when the pretraining tasks of the VFM (e.g., MAE, DINOv2) significantly deviate from the DGSS task, the adaptation improvement is limited. An alternative solution is to fine-tune a subset of parameters related to the target DGSS task, which activates the representations of VFMs, as illustrated in Fig. 1 (b). However, we observe that traditional parameter selection methods, whether manually defined [57] or automatically chosen [36,62], fail to guarantee the generalization ability of the VFM, and in fact, perform even worse than simply adding adapters, as shown in Fig. 2. This CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore. 15043
Figure 2. Comparison of average performance across multiple VFMs in DGSS experiments on GTA →Cityscapes + BDD100K + Mapillary using different fine-tuning methods, including adapter-based Rein [60], manually selected parameterbased VQT [57], adaptively selected parameter-based ChildTune [62], and our FisherTune. Our experiments reveal that certain parameters of Vision Foundation Models (VFMs) are crucial for maintaining generalization, while others are key to adapting to new domains and tasks. Traditional selective fine-tuning methods focus solely on task-sensitive parameters, risking the disruption of the VFM’s generalization ability. Therefore, we propose identifying and fine-tuning these domain-sensitive parameters. To address this, we introduce FisherTune, a novel robust fine-tuning method based on the DomainRelated Fisher Information Matrix (DR-FIM). This method allows us to preserve the generalization capability of the pre-trained VFM while activating its adaptability for DGSS tasks. Specifically, we first introduce the DR-FIM metric, which measures domain sensitivity by evaluating the fluctuation of parameters across different domains. Unlike FIM metrics, DR-FIM accounts for domain shifts, extending FIM for cross-domain tasks. To mitigate potential degradation in DR-FIM estimation, FisherTune innovatively employs variational inference. By treating model parameters as random variables following a Gaussian distribution and incorporating prior information from the pre-trained VFM, FisherTune stabilizes the DR-FIM estimation process, ensuring accuracy and robustness in cross-domain tasks. This novel estimation method improves computational efficiency and enhances the scalability of FisherTune to large-scale VFM models. Through extensive experiments on multiple DGSS benchmarks, we demonstrate that FisherTune consistently outperforms both selective-parameter and adapterbased methods. In summary, our contributions are, • We propose FisherTune, a novel fine-tuning strategy that leverages Fisher Information Matrix to selectively fine-tune VFMs for DGSS, preserving the generalization capabilities of VFMs while improving domain adaptability. • We introduce Domain-Related FIM (DR-FIM), a novel metric matrix that quantifies the sensitivity of parameters to domain shifts. • We employ variational inference to treat model parameters as Gaussian-distributed, ensuring stable and accurate DR-FIM estimation. • We validate the effectiveness of FisherTune through extensive experiments, showing superior generalization compared to state-of-the-art methods. 2. Related Work Domain Generalized Semantic Segmentation (DGSS) focuses on enhancing a model’s ability to generalize to unseen domains by training on source domains [8,43,74]. Common strategies include domain-invariant representation learning methods and domain augmentation techniques. Domain-invariant representation learning approaches involve splitting learned features into domain-invariant [68, 69] and domain-specific components [32,58,59,75], or employing meta-learning to develop more robust models [11, 30,73]. Additionally, several methods have succeeded by learning feature normalization or whitening schemes [9, 39,44]. Domain augmentation techniques, on the other hand, improve segmentation results through style transfer at image-level [23,43,45,78] or feature-level [6,8,77] and the introduction of additional data [29,38,67]. Some recent work has shown that text-guided feature enhancement [12,13] through CLIP [47] or synthetic data of diffusion models can also benefit model generalization [3,40,76]. Parameter-Efficient Fine-Tuning(PEFT) [18] customizes pre-trained models by fine-tuning a subset of parameters [17], improving performance and generalization with lower computational cost. The dominance of ViT [10] in vision tasks has spurred the development of PEFT methods. Adaptor-based Prompt Tuning [34,55] has shown strong performance in vision transfer tasks by adding learnable prompts. For instance, Visual Prompt Tuning (VPT) [25] introduces learnable prompts for each Transformer layer’s input embeddings, AdaptFormer [7] adds a bottleneck fully connected layer parallel to the MLP block, and VQT [57] optimizes prompts through bypassing to effectively leverage intermediate features of VFMs. In semantic segmentation, several works have applied visual prompts for model transfer. [35] uses frequency and spatial prompts to transfer pre-trained ViTs to low-level segmentation tasks, while [64] applies mask prompts to aid continual adaptation. [71] uses weak supervision for LoRA adapter [21] to adapt VFMs across domains. Rein [60] adds LoRA adapters to transform tokens and activate VFMs for 15044
DGSS. [66] enhances cross-domain adaptation by adding Fourier transform prompts to intermediate tokens of VFMs. While most visual semantic segmentation methods rely on adapter-based prompt fine-tuning, our work pioneers selective fine-tuning methods for visual model adaptation. Closely related to our approach are selective fine-tuning methods in NLP, like ChildTune [62] and Fishermask [36], which use Fisher matrices to identify task-sensitive parameters. In contrast, our method introduces domain-sensitive parameters and a stable estimation method to identify parameters highly sensitive to both tasks and domains, making it particularly suited for DGSS tasks. 3. Methodology 3.1. Preliminaries Domain Generalized Semantic Segmentation (DGSS) aims to train models that can generalize across unseen domains. Formally, given a set of labeled source domains Ds={(xi, yi)}Ns i=1, where xirepresents the input image and yirepresents the corresponding pixel-wise label, the goal is to train a model fθparameterized by θthat performs well on unseen target domains Dt={xj}Nt j=1, where the labels for Dtare not available during training. The optimization objective for DGSS can be written as: min θ E(xi,yi)∼Ds[L(fθ(xi), yi)] ,(1) where Lis the segmentation loss (e.g., cross-entropy) that evaluates the difference between the predicted segmentation map fθ(xi)and the ground truth yi. The challenge lies in ensuring that the learned model fθgeneralizes well to unseen target domains Dt, which can be written as a generalization objective: min θ Exj∼DtL(fθ(xj), y∗ j),(2) where y∗ jrepresents the true (but unknown) labels for the target domain Dt. Since the labels y∗ jare not available, the optimization focuses on learning domain-invariant representations in θ, enabling strong performance across both seen and unseen domains. Vision Foundation Models (VFMs), such as CLIP [46], MAE [19], SAM [28], EVA02 [14], and DINOv2 [41], almostly use the Vision Transformer (ViT) architecture. ViT typically consist of Lstacked blocks, each containing two main submodules: multi-head attention (MHA) and a feed-forward network (FFN). Specifically, the attention score for each head is calculated as: MHA(X) = Concat(head1,...,headh)θo,headi= Softmax Xθqi(Xθki)T √dhXθvi,where θois the output projection matrix, and θqi,θki, and θvirepresent the query, key, and value projections for head i. The FFN consists of two linear layers with a ReLU activation: FFN(X) = Freeze Full 使用DinoV2-Large 在GTA→ (Cityscapes, BDD, Mapi,的DGSS实验 平均性能) 我们观察到,仅微调Vision Foundation Models (VFM)中的部分 关键参数相比于全微调或完全冻结 其他参数,能显著提高模型的泛化 性能。这一现象启发了我们提出一 个假设:VFM中的某些预训练参数与 特定任务和领域的适应性密切相关, 而其他参数则较为通用,能够更好 地适应不同的域和任务。 基于这一观察,我们提出识别VFM中 与任务和域适应性密切相关的“域敏 感参数”,并对这些参数进行精细微 调,从而在保持VFM预训练泛化能力 的前提下,提升其在Domain Generalized Semantic Segmentation (DGSS)任务中的适应性。 Figure 3. Observations of fine-tuning different VFM layers for DGSS experiments using DINOV2-large under GTA → Cityscapes + BDD100K + Mapillary. It shows that fine-tuning different layers has different effects on the generalization performance of the VFMs. B means blocks. ReLU(Xθffn1+θb1)θffn2+θb2.Both the MHA and FFN are followed by residual connections and layer normalization. Let θdenote the set of the parameters of those VFMs that we aim to fine-tune: θ= [θ(1) Q, θ(1) K, θ(1) V, θ(1) FFN, . . . , θ(L) Q, θ(L) K, θ(L) V, θ(L) FFN]⊤, (3) where θ(l) Q= [θqi, ..., θqh], and so are θ(l) Kand θ(l) K. Fisher Information Matrix (FIM) is a fundamental concept in statistical estimation theory [16], which measures the amount of information that an observable random variable carries about an unknown parameter [2]. In the context of neural networks, the FIM provides insights into the sensitivity of the loss function with respect to the model parameters [36], reflecting the curvature of the loss landscape [48]. Mathematically, the FIM Fθis defined as: Fθ=ExEy∼fθ(y|x)∇θL(fθ(x), y)· ∇θL(fθ(x), y)⊤, (4) where Fθ∈R|θ|×|θ|is the symmetrical matrix, ∇θL(·)denotes the gradient of the loss function to the parameters. Intuitively, the Fisher Information Matrix captures how much changing the parameters affects the model’s output, thus quantifying the “informativeness” of the parameters. 3.2. Motivation This study is motivated by an intriguing experimental observation. We group the θQ,θK,θV, and θFFN components from different blocks of DINOv2 separately and fine-tune all possible pairwise combinations of these groups, and analyzed their impact on model generalization, as shown in Fig. 3. We found that tuning specific layers led to different levels of generalization in VFMs, with some configurations 15045
even outperforming a fully tuned model. This suggests that certain parameters are critical for maintaining generalization, while others are key to adapting to new domains and tasks. Based on this, we hypothesize that VFMs contain domain-sensitive parameters suited for specific tasks and domains, while other parameters remain broadly generalizable. Consequently, we propose identifying and fine-tuning these domain-sensitive parameters, enhancing adaptability in DGSS for improved cross-domain performance. 3.3. FisherTune In this section, we present our FisherTune, fine-tuning ViT-based Vision Foundation Models (VFMs) guided by the Fisher Information Matrix (FIM) while preserving their generalization strengths. Our idea is to use FIM to find taskand domain-sensitive parameters in θand fine-tune these sensitive parameters to improve the generalization ability of the model on unseen domains while maintaining the pretrained knowledge of VFMs. 3.3.1 Domain-Related FIM Domain-Related FIM. In Eq. 4, FIM quantifies the importance of model parameters for the current task by measuring the sensitivity of parameters to model output. However, FIM can not sufficiently capture the behavior of parameters in cross-domain scenes, especially in DGSS tasks, where the sensitivity of different parameters to varying data distributions may differ significantly. For the DGSS task, we need a parameters estimation metric that can capture the variation of parameters across different domains. To address this issue, we propose to calculate the Fisher information change ∆Fθbetween different data domains (seen domain and simulated unseen domain) to measure the sensitivity difference of parameters across domains. Formally, given the seen single-source domain Ds={(x, y)}, the ∆Fθis calculated as: ∆Fθ=|Fθ(x, y)−Fθ(x′, y)| min(Fθi(x),Fθi(x′)) + ϵ,(5) where Fθ(x, y)and Fθ(x′, y)is the FIM for in the seen and simulated unseen domain. ϵ= 1 ×10−8is a small constant to prevent division by zero. The numerator, |Fθ(x, y)−Fθ(x′, y)|, computes the difference between the FIM across domains, reflecting the model’s varying sensitivity to parameter changes in different environments or data distributions. The denominator, min(Fθi(x),Fθi(x′)), normalizes this difference to ensure the relative nature of the metric. A higher ∆Fiindicates that the parameter is more sensitive to domain changes. To simulate an unseen domain sample x′from a seen domain x, we leverage the uncertainty modeling method inspired by [33]. Specifically, the unseen domain feature x′ 𝑥𝑥=𝑥𝑥𝑥 𝒙𝒙𝟏𝟏𝑥 FIM DR-FIM = FIM DR-FIM < FIM DR-FIM < FIM DR-FIM < Domain Shift: 𝒙𝒙𝟐𝟐𝑥𝒙𝒙𝟑𝟑𝑥 𝑥𝑥1𝑥𝑥𝑥2𝑥 𝑥𝑥3𝑥 Figure 4. Comparison of FIM and DR-FIM under different degrees of domain shift. The size of the circle indicates the value. It shows that DR-FIM is a generalization of FIM as it additionally considers the cross-domain sensitivity of parameters. is simulated by modifying the feature statistics (mean and variance) of the seen domain sample x. The perturbed mean is generated as, α(x) = µ(x) + ϵµΣµ(x), where µ(x)represents the mean of the feature, Σµ(x)is the uncertainty estimation for the mean, and ϵµ∼ N(0,1) is noise sampled from a standard normal distribution. Next, the perturbed variance is generated as, β(x) = σ(x) + ϵσΣσ(x), where σ(x)is the standard deviation of the feature, Σσ(x) is the uncertainty estimation for the standard deviation, and ϵσ∼ N(0,1). Using the perturbed mean α(x)and variance β(x), the unseen domain sample x′is generated with the following formula, x′=β(x)·x−µ(x) σ(x)+α(x).(6) Using ∆Fθ, we introduce a unified metric, DomainRelated FIM (DR-FIM), to account for both task-sensitive and domain-sensitive parameters as, DRFθ=Fθ(x, y) | {z } task-sensitive +e−(ϵµ+ϵσ)|Fθ(x, y)−Fθ(x′, y)| min(Fθi(x),Fθi(x′)) + ϵ | {z } domain-sensitive . (7) The DRFθis a linear combination of Fθand ∆Fθ, and the combination coefficients are determined by domain shift control factors ϵµand ϵσ. When ϵµand ϵσare large, the simulated domain shift is significant, and ∆Fθis scaled appropriately to balance with Fθ. The relationship between the simulated domain shift and the numerical values of DRFIM and FIM is shown in Fig. 4. 3.3.2 Stable Estimation of DR-FIM Although DRFθprovides an estimation of the domain sensitivity of VFMs parameters, the dimensions of the parameters list θare very high, which makes it impractical to directly calculate in computation and storage, i.e.,O(|θ|2). Therefore, it is necessary to approximate the FIM to reduce the computational complexity. Diagonal Approximation. Following [36], by assuming that the off-diagonal elements are negligible, the FIM can 15046
be efficiently approximated by using a diagonal approximation, i.e.,ˆ Fθ=diag(Fθ1,Fθ2, ..., Fθ|θ|),(8) where each individual Fθncan be calculated as, Fθn=1 N N X i=1 Eyi∼fθ(yi|xi)(∇θnL(fθ(x), yi))2,(9) In the above diagonal approximation, only the individual contribution of each parameter to the loss is considered, while the interaction terms between different parameters are ignored. This diagonal approximation effectively simplifies aO(|θ|2)matrix to a vector of length O(|θ|), greatly reducing the computational complexity. Variational Estimation of DR-FIM. Diagonalization provides an efficient approximation method for computing the FIM. However, due to the differences between the pretraining tasks of the VFM and the DGSS task, the estimated FIM parameters often exhibit high sensitivity, leading to inaccuracies (See Fig. 6). To address this issue, we introduce a variational inference approach [4], treating the finetuning model’s parameters θas random variables following a Gaussian distribution. This introduces an additional regularization term into the FIM estimation, helping to learn a smoother prior distribution. Consequently, variational inference stabilizes the gradient update process during FIM calculation, mitigating the instability caused by high gradient noise. Specifically, assuming that the posterior distribution of the model parameters follows a Gaussian distribution: q(θ) = N(ˆ θ,Λ−1), where ˆ θis the mean of the current parameters estimation, Λ−1is the covariance matrix of the parameters. To preserve the pre-trained knowledge of VFMs, we introduce the prior parameter distribution as a regularizer to prevent degradation during prediction: p(θ) = N(θpt, τ2I),(10) where θpt is the pre-trained parameters of VFMs, τ2is the variance controlling the flexibility of fine-tune parameters, and Iis the identity matrix. We utilize the variational free energy (also called the evidence lower bound, ELBO [20]) as the loss function for optimizing Λ, L(ˆ θ,Λ−1) = Eθ∼q(θ)[L(θ)] + γ KL(q(θ)∥p(θ)),(11) where γis the regularization coefficient, controlling the influence of the prior, and KL(q(θ)∥p(θ)) is the KullbackLeibler divergence between the posterior q(θ)and the prior p(θ). Connection with DR-FIM. To simplify the first term in Eq. (11), we perform a second-order Taylor expansion of the loss function L(θ)around the current parameters estimate θ=ˆ θ. Taking the expectation over the weight distribution q(θ),Eθ∼q(θ)[L(θ)] ≈ L(ˆ θ) + 1 2Tr ∇2 θL(ˆ θ)Λ−1.According to the definition of FIM and its connection with the Hessian matrix [16], the FIM can be approximated by the Hessian matrix near ˆ θ, ∇2 θL(ˆ θ)≈Fθ.Thus, Eθ∼q(θ)[L(θ)] ≈ L(ˆ θ) + 1 2Tr FθΛ−1.(12) The second term in Eq. (11), KL divergence between two Gaussian distributions is simplified by, KL(q(θ)∥p(θ)) = 1 2τ−2Tr(Λ−1) + τ−2∥ˆ θ−θpt∥2 −k+kln τ2+ ln det Λ. (13) Substituting the Eq. (12) and Eq. (13) back into Eq. (11), and taking the derivative of the loss function with respect to Λand we obtain (See Appendix A for detailed derivation), Fθ=γΛ−γτ−2I. (14) Then, the DR-FIM defined in Eq. (7) is updated as, DRFθ=γ Λx−τ−2I+e−(ϵµ+ϵσ)|Λx−Λx′| min(Λx,Λx′) + ϵ γ!. (15) It shows that the DR-FIM can be estimated from the covariance matrix Λ, with γand τas hyperparameters. Using Eq. (15) to estimate the DR-FIM has several advantages over using Eq. (7) and Eq. (9). 1) Stability in Estimation: Eq. (15) introduces a more stable estimation mechanism by incorporating prior knowledge from the pre-trained VFMs p(θ)in Eq. (10) and the posterior distribution q(θ). This approach helps prevent the degradation of FIM estimation caused by the task shift between VFM pre-training tasks and DGSS tasks, ensuring more robust performance across unseen domains. 2) Computational Efficiency: The covariance matrix Λcan be efficiently computed by directly minimizing the loss L(ˆ θ,Λ−1)in Eq. (11) using reparameterization trick [27] and stochastic gradient variational Bayes [1], reducing both computational and memory overhead compared to traditional FIM estimation in Eq. (9). 3.3.3 Training Scheduler of FisherTune We follow Rein [60] which adds a mask decoder to the backbone network of VFMs as a segmentation model for DGSS. Different from Rein, we do not modify the backbone structure or add additional adapters. During training, we first fix the backbone network of VFMs and use the original data to warm-up the decoder to adapt the whole segmentation model to the DGSS task. After that, we tune the VFMs and decoder by our FisherTune as follows. 15047
Algorithm 1 FisherTune Process 1: Input: source dataset D={(xi, yi)}N i=1; Hyperparameters: regularization coefficient γ, variance coefficient τ, warm-up iterations T1, FIM estimation iterations T2, number of tune iterations T3; pretrained VFM θVFM; segmentation decoder θdec. 2: Step 1: Warm-up decoder: 3: Train the decoder θdec on Dfor T1steps, froze θVFM. 4: Step 2: Sampling and DR-FIM Calculation: 5: for t= 1 to T2do 6: Sample batch (x, y)∼ D 7: Simulate unseen domain data x′via Eq. 6. 8: Optimize covariance matrix Λvia Eq. 11 9: Estimate DR-FIM using the optimized Λvia Eq. 15 10: Step 3: Parameter Fine-Tuning: 11: for t= 1 to T3do 12: Sample a batch (x, y)∼ D 13: Select parameters ˆ θVFM via Eq. 16 14: Update the selected ˆ θVFM and θdec via Eq. 1. 15: Output: Fine-tuned θVFM and θdec. In FisherTune, the selection of parameters for fine-tuning is guided by the DR-FIM (DRFθ), which quantifies the sensitivity of parameters to task and domain shifts. To optimize the fine-tuning process, we propose a dynamic training schedule that adjusts the number of trainable parameters based on their DR-FIM values. At the beginning of training, we fine-tune only the most sensitive δmin% of parameters, as ranked by DRFθi. As training progresses, we gradually increase the percentage of fine-tuned parameters, reaching δmax% by the end. This ensures that the model starts with a focused fine-tuning process, targeting only the most critical parameters, and progressively expands the fine-tuning scope as the model becomes more stable. Formally, at each training step t, the dynamic threshold DRFthresh(t)is updated as follows: DRFthresh(t) = δmin +(δmax −δmin)·exp −t T,(16) where Tis the total number of training steps. Parameters with DRFθvalues higher than the threshold DRFthresh(t) will be selected for training. The detailed fine-tuning process of our FisherTune is in Algorithm 1. 4. Experiments 4.1. Datasets & Setup See Appendix B. 4.2. Comparison with State-of-the-art Alternatives GTAV →C, B, M. Table 1demonstrates that our approach significantly outperforms other fine-tuning methods across multiple vision foundation models (VFMs). ComGTAV →Cityscapes (Citys) + BDD100K (BDD) + Mapillary (Map) VFM type Fine-tune Method Trainable Params Citys BDD Map Avg. CLIP [46] (ViT-Large) Full 304.20M 51.3 47.6 54.3 51.1 Freeze 0M 53.7 48.7 55.0 52.5 LoRA [22] 0.79M 54.0 49.8 55.1 53.0 VPT [25] 3.69M 54.0 51.8 57.5 54.4 Rein [60] 2.99M 57.1 54.7 60.5 57.4 VQT [57] 3.01M 54.3 51.2 56.7 55.3 ChildTune [63] 15.21M 57.9 53.4 58.2 56.5 Ours 15.21M 59.2 57.5 61.0 59.2 MAE [19] (Huge)) Full 304.20M 53.7 50.8 58.1 54.2 Freeze 0M 43.3 37.8 48.0 43.0 LoRA [22] 0.79M 44.6 38.4 52.5 45.2 VPT [25] 3.69M 52.7 50.2 57.6 53.5 Rein [60] 2.99M 55.0 49.3 58.6 54.3 VQT [57] 3.01M 53.3 50.3 57.7 53.8 ChildTune [63] 15.21M 55.4 50.6 58.1 54.7 Ours 15.21M 56.6 51.9 59.7 56.1 SAM [28] (Huge) Full 632.18M 57.6 51.7 61.5 56.9 Freeze 0M 57.0 47.1 58.4 54.2 LoRA [22] 0.79M 57.4 47.7 58.4 54.5 VPT [25] 3.69M 56.3 52.7 57.8 55.6 Rein [60] 2.99M 59.6 52.0 62.1 57.9 VQT [57] 3.01M 56.7 53.9 59.3 56.6 ChildTune [63] 15.21M 60.8 49.6 61.2 57.2 Ours 15.21M 60.9 54.4 63.9 59.7 EVA02 [15] (Large) Full 304.20M 62.1 56.2 64.6 60.9 LoRA [22] 0.79M 55.5 52.7 58.3 55.5 AdaptFormer [7] 3.17M 63.7 59.9 64.2 62.6 VPT [25] 3.69M 62.2 57.7 62.5 60.8 Rein [60] 2.99M 65.3 61.1 63.9 63.4 VQT [57] 3.01M 61.3 55.1 62.2 59.5 ChildTune [63] 15.21M 61.6 59.3 62.3 61.1 Ours 15.21M 65.8 61.5 66.0 64.4 DINOv2 [41] (ViT-Large) Full 304.20M 63.7 57.4 64.2 61.7 LoRA [22] 0.79M 65.2 58.3 64.6 62.7 AdaptFormer [7] 3.17M 64.9 59.0 64.2 62.7 VQT [25] 3.01M 64.6 59.0 65.7 63.1 Rein [60] 2.99M 66.4 60.4 66.1 64.3 ChildTune [63] 15.21M 65.6 59.3 65.3 63.4 Ours 15.21M 68.2 63.3 68.0 66.5 EVA02 VLTSeg [24] 304.2M 65.3 58.3 66.0 63.2 DINOV2 SDT [66] 6.94M 68.1 61.6 67.7 65.8 CLIP+SAM CLOUDS [42] 304.2M 60.2 57.4 67.0 61.5 EVA02 tqdm [42] 304.2M 68.9 59.2 70.1 66.1 EVA02 Ours 15.21M 65.8 61.5 66.0 64.4 DINOV2 Ours 15.21M 68.2 63.3 68.7 66.6 Table 1. Performance and Trainable Parameters Comparison with the proposed FisherTune across Multiple VFMs as Backbones under the GTAV →Cityscapes (Citys) + BDD100K (BDD) + Mapillary (Map) generalization setting. pared to adapter-based methods (e.g., LoRA and Rein), our approach achieves an average of 4.3% higher mIoU than Rein across five VFM models. Additionally, it surpasses the self-focused parametric fine-tuning method VQT by 3.1% on average. Notably, for models with a substantial gap between pre-training and downstream tasks, such as MAE and EVA02, adapter methods yielded modest improvements of 1.3% and 1.7% mIoU, respectively, whereas our approach achieved 4.6% and 6.6% improvements. Besides, we added comparisons with the state-of-the-art methods using VFMs, and our method remains competitive. The tqdm [42] and VLTSeg [24] method leverages features of the language model, while Rein-series methods and ours focus on visual models. These results highlight our method’s enhanced adaptability to downstream tasks and its signifi15048
Cityscapes →BDD100K Fine-tune Method Trainable Params road side. build. wall fence pole light sign vege terr. sky pers. rider car truck bus train moto. bicy. mIoU DINOv2 [5] (Large) Full 304.20M 89.0 44.5 89.6 51.1 46.4 49.2 60.0 38.9 89.1 47.5 91.7 75.8 48.2 91.7 52.5 82.9 81.0 30.4 49.9 63.7 Freeze 0M 92.1 55.2 90.2 57.2 48.5 49.5 56.7 47.7 89.3 47.8 91.1 74.2 46.7 92.2 62.6 77.5 47.7 29.6 47.2 63.3 REIN [60] 2.99M 92.4 59.1 90.7 58.3 53.7 51.8 58.2 46.4 89.8 49.4 90.8 73.9 43.3 92.3 64.3 81.6 70.9 40.4 54.0 66.4 VQT [57] 3.01M 88.3 49.9 85.9 50.7 47.9 44.3 55.6 39.2 86.1 42.8 87.5 71.3 45.4 89.4 53.5 82.6 74.9 46.1 57.4 63.1 ChildTune [62] 15.21M 92.1 56.1 91.0 58.8 46.9 52.0 58.6 47.2 90.8 47.9 93.3 72.0 47.1 93.0 63.9 76.2 47.9 28.8 48.3 63.8 Ours 15.21M 92.1 55.4 90.2 58.9 50.9 54.5 59.8 49.1 92.5 52.8 91.0 73.7 51.5 92.7 67.4 82.9 72.8 44.3 54.1 67.7 EVA02 [15] (Large) Full 304.20M 89.3 46.9 89.9 47.7 45.6 50.1 56.8 42.2 88.8 48.4 89.9 75.8 49.0 90.5 45.3 69.2 55.9 44.4 55.1 62.1 Freeze 0M 93.1 52.7 88.0 47.4 31.1 41.7 46.0 39.6 85.7 41.4 89.5 67.5 39.7 89.0 47.0 72.8 46.3 19.2 35.2 56.5 REIN [60] 2.99M 91.7 51.8 90.1 52.8 48.4 48.2 56.0 42.0 89.1 44.1 90.2 74.2 47.0 91.1 54.5 84.1 78.9 47.2 59.4 65.3 VQT [57] 3.01M 90.1 46.6 91.1 46.9 46.4 51.7 56.5 43.2 89.3 49.6 92.3 75.0 50.3 90.3 44.6 71.8 57.4 44.0 55.8 62.8 ChildTune [62] 15.21M 87.9 46.5 88.1 46.5 46.1 46.1 56.0 41.5 87.9 50.3 89.6 77.7 45.6 91.4 42.4 68.1 54.7 46.0 56.8 61.5 Ours 15.21M 92.6 49.9 95.9 51.1 53.0 50.8 59.8 45.7 92.9 54.6 94.0 83.5 52.2 93.9 45.1 69.4 57.1 47.2 62.4 65.8 Cityscapes →ACDC DINOv2 [5] (Large) Full 304.20M 92.8 75 87.4 55.7 54.1 55.6 71.2 69.6 82.4 56 92.2 66.8 45.6 89 79.7 87.9 87.5 51.4 62.7 71.7 Freeze 0M 86.0 68.1 80.2 52.4 47.8 48.2 65.5 65.3 80.0 54.7 86.2 65.0 44.9 86.4 73.3 80.5 86.9 50.1 60.9 67.5 REIN [60] 2.99M 94.6 78.3 92.0 61.9 55.0 64.8 73.8 72.7 88.4 67.4 95.4 77.1 60.2 92.6 84.1 86.9 92.5 67.6 68.6 77.6 VQT [57] 3.01M 93.3 76.4 89.2 55.0 53.9 53.9 72.0 67.3 83.4 55.3 95.1 67.7 47.0 90.5 81.6 86.3 88.2 50.1 61.9 72.0 ChildTune [62] 15.21M 92.9 72.8 84.7 56.6 54.1 56.8 70.9 67.7 82.3 55.7 93.6 65.9 45.3 89.6 77.6 87.8 87.0 52.5 62.2 71.4 Ours 15.21M 95.6 79.0 96.5 60.5 58.3 64.9 75.6 77.7 85.0 61.3 98.6 73.6 51.5 94.8 85.4 94.7 93.8 59.0 66.7 77.5 EVA02 [15] (Large) Full 304.20M 90.2 68.8 81.0 53.7 49.9 48.1 68.7 64.2 80.1 57.4 88.1 68.8 41.8 89.7 74.1 82.1 89.7 50.0 56.8 68.6 Freeze 0M 86.0 60.5 76.3 49.0 41.7 46.1 60.5 61.0 72.1 49.8 77.7 56.7 40.6 80.3 68.3 77.2 85.5 46.7 56.4 62.8 REIN [60] 2.99M 88.7 71.8 81.7 55.2 51.7 50.5 70.5 64.9 83.7 59.0 90.3 72.0 48.3 93.0 79.3 83.3 91.3 50.8 62.0 70.9 VQT [57] 3.01M 90.3 71.2 81.4 54.3 53.1 49.1 67.9 64.3 82.0 60.5 86.9 66.8 41.3 89.3 76.6 81.7 91.3 47.2 55.7 69.0 ChildTune [62] 15.21M 86.4 68.8 81.0 54.4 50.6 48.9 69.6 64.5 83.2 57.8 88.2 69.0 47.9 90.2 74.8 82.8 90.3 51.0 61.4 69.5 Ours 15.21M 90.5 75.2 83.6 58.8 54.6 52.2 73.1 66.6 85.7 60.5 90.2 70.7 51.5 92.3 82.6 88.2 91.9 54.0 62.4 72.9 Table 2. DGSS generalization performance for each category from the Cityscapes source domain to mixed-domain BDD100K and ACDC, with comparison methods including adaptor-based Rein [60] and selective parameter fine-tuning methods VQT [57] and ChildTune [62]. Cityscapes →Adverse Weather Fine-tune Method Trainable Params Foggy Zurich [49] Foggy Driving [49] Dark Zurich [50] Nighttime Driving [52] ACDC-Rain [51] ACDC-Snow [51] mIoU DINOv2 [5] (Large) Full 304.20M 50.4 55.3 62.7 47.7 75.2 76.8 61.3 Freeze 0M 50.3 43.7 54.3 40.8 66.1 71.7 54.5 REIN [60] 2.99M 55.5 58.2 64.3 50.3 78.2 79.5 64.3 VQT [57] 3.01M 54.1 57.1 61.9 47.4 76.1 75.3 62.0 ChildTune [62] 15.21M 55.2 56.9 64.5 50.7 77.7 78.3 63.9 Ours 15.21M 56.9 60.0 66.6 53.2 78.6 82.2 66.3 Table 3. DGSS performance comparison for Cityscapes as the source domain under diverse weather conditions. Cityscapes →BDD100K Cityscapes →ACDC EVA02 [5] (Large) Full 62.1 68.6 Freeze 56.5 62.8 Random 61.1 67.6 Random Q62.8 69.1 Random K61.9 68.1 Random V62.9 69.2 Fθ63.8 69.5 ∆Fθ63.1 71.3 DRFθ65.8 (+3.7) 72.9 (+5.3) DINOv2 [5] (Large) Full 63.7 71.7 Freeze 63.3 67.5 Random 62.7 71.0 Random Q63.2 72.0 Random K63.5 72.3 Random V63.2 72.9 Fθ63.8 71.4 ∆Fθ64.5 76.1 DRFθ67.7 (+4.0) 77.5 (+5.8) Table 4. Ablation study on generalization with 5% fine-tunable parameters in terms of mIoU. cant improvement in model generalization. Cityscapes →BDD100K, ACDC. In migrating from Cityscapes to BDD100K and ACDC, our method achieved strong results, with average mIoU scores of 67.7% and 77.5%. As shown in Table 2, our method’s average mIoU on BDD100K is 2.4% higher than REIN. Compared to VQT and ChildTune, our method improved mIoU by 4.4% and 2.6% on the respective datasets, addressing issues in parameter tuning, data adaptation, and migration strategy. These results highlight our method’s superior adaptability and generalization in complex scenes. Cityscapes →Adverse Weather. We evaluated various fine-tuning strategies for DINOv2 models across challenging weather conditions, as shown in Table 3. Our approach achieved an average of 2.0% higher mIoU than the adapterbased REIN method and 4.3% higher than the self-focused VQT approach. This improvement likely stems from the substantial difference between pre-training and downstream tasks. Besides, ChildTune showed limited performance gains, and our method surpassed ChildTune by an average of 2.4% mIoU, demonstrating superior adaptability and generalization under complex weather scenarios. 4.3. Ablation Studies Ablation of DR-FIM effectiveness As shown in Table 4, randomly selecting Q,K, and Vparameters for fine-tuning does not fully leverage the generalization ability of VFMs, leading to lower mIoU. Using FIM (Fθ) for parameter selection improves performance over random choice. Further gains are achieved with ∆Fθ, which better identifies domain-sensitive parameters—especially on ACDC, where severe weather differences pose greater challenges. Our proposed DR-FIM, combining Fand ∆F, delivers the best 15049
Method EVA02 EVA02+FP DINOV2 DINOV2+FP AdaptFormer 62.6 63.3 (+0.7) 62.7 63.7 (+1.0) VPT 60.8 61.8 (+1.0) 63.3 64.1 (+0.8) Rein 63.6 63.9 (+0.3) 64.3 65.0 (+0.7) Ours 64.4 64.5 (+0.1) 66.3 66.5 (+0.2) Table 5. Ablation study on Feature Perturbation (FP) using [33]. results, boosting mIoU by +3.7% and +5.3% on Cityscapes →BDD100K and Cityscapes →ACDC for EVA02 (Large), and by +4.0% and +5.8%, respectively. These results highlight the effectiveness of our method. Ablation of DR-FIM Estimation Fig. 5presents the ablation study on the proposed stable estimation method. The results show that while DR-FIM outperforms FIM in parameter evaluation, but the effectiveness of DR-FIM is limited by traditional estimation methods. The stable estimation method significantly enhances the accuracy of parameter evaluation for both FIM and DR-FIM. Notably, applying stable estimation to DR-FIM results in an average improvement of 2.6% mIoU, demonstrating superior overall generalization performance. Ablation of Feature Perturbation Since we adopt domain simulation augmentation from [32], which is generally considered effective for DG, we also apply it to existing VFM methods for a fair comparison. FisherTune uses feature perturbation (FP) solely for identifying domainsensitive parameters, not during fine-tuning. As shown in Table 5, FP yields a modest improvement (+1.0% mIoU) on GTA→Avg., yet our method still outperforms others. 4.4. Discussion Captured domain-sensitive parameters. Fig. 6illustrates the impact of different estimation methods on parameter sensitivity estimation. (a) shows that parameter sensitivity estimated from original FIM is generally high, making it difficult to identify the most valuable parameters. (b) demonstrates that incorporating ∆Fθredefines parameter sensitivity by comprehensively considering both task relevance and domain sensitivity. (c) presents the DR-FIM estimated using a robust way, which highlights important parameters more effectively, aiding in the selection of valuable parameters. Additionally, (c) reveals that important parameters tend to be concentrated in the Q,K,Vand FFM parameters of deeper blocks. Furthermore, the overall sensitivity of Qand Kis higher than that of V. Feature Visualization. Fig. 7compares the T-SNE visualizations of feature distributions between Rein [60] and FisherTune. FisherTune exhibits a more balanced feature distribution across multiple unseen domains, indicating reduced domain bias and improved generalization. The Ratio of Fine-tuned Parameters.See Appendix C. Segmentation Result Visualization.See Appendix D. Influence of Hyper-parameters.See Appendix E. Figure 5. Ablation study of estimation ways on Cityscapes →BDD100K (C2B), →ACDC (C2A), and GTAV → Cityscapes(G2C), →BDD100K(G2B) and →Mapillary(G2M). (b) DR-FIM without Stable Estimation (c) DR-FIM with Stable Estimation (a) FIM Figure 6. Diagram of parameter sensitivity estimated by FIM and our DR-FIM using DINOV2-large, trained on GTAV for DGSS experiments. The Q,K,V, and FFN parameters are arranged in ascending order according to their block indices. Night •Rainy•Snow . •Foggy Figure 7. Comparison of T-SNE feature visualizations: Rein [60] (left) and the proposed FisherTune (right). The model is trained on the Cityscapes →ACDC DGSS task. FisherTune shows a more balanced feature distribution across multiple unseen domains. 5. Conclusion We propose FisherTune, a fine-tuning method for Vision Foundation Models (VFMs) in DGSS. It introduces the Domain-Related Fisher Information Matrix (DR-FIM) to measure parameter sensitivity to domain shifts, using variational inference for stable estimation. FisherTune enhances domain adaptability while maintaining generalization. We hope it encourages further research on selective fine-tuning to better unlock the generalization potential of VFMs in DGSS and beyond. 15050
References [1] Alessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran, Subhransu Maji, Charless C Fowlkes, Stefano Soatto, and Pietro Perona. Task2vec: Task embedding for meta-learning. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6430–6439, 2019. 5 [2] Alessandro Achille, Giovanni Paolini, and Stefano Soatto. Where is the information in a deep neural network? arXiv preprint arXiv:1905.12213, 2019. 3 [3] Yasser Benigmim, Subhankar Roy, Slim Essid, Vicky Kalogeiton, and St´ ephane Lathuili` ere. Collaborating foundation models for domain generalized semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3108–3119, 2024. 2 [4] David M Blei, Alp Kucukelbir, and Jon D McAuliffe. Variational inference: A review for statisticians. Journal of the American statistical Association, 112(518):859–877, 2017. 5 [5] Mathilde Caron et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 1,7 [6] Prithvijit Chattopadhyay, Kartik Sarangmath, Vivek Vijaykumar, and Judy Hoffman. Pasta: Proportional amplitude spectrum training augmentation for syn-to-real domain generalization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 19288–19300, 2023. 2 [7] Shoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang, Yibing Song, Jue Wang, and Ping Luo. Adaptformer: Adapting vision transformers for scalable visual recognition. Advances in Neural Information Processing Systems, 35:16664–16678, 2022. 2,6 [8] Ziyuan Cheng, Ruinian Wan, Meng Li, Feiyu Wang, Chao Xu, and Xiaofei He. Domain generalization via styleefficient perturbation and clustering of intra-domain heterogeneous data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3938– 3947, 2022. 2 [9] Sungha Choi, Sanghun Jung, Huiwon Yun, Joanne T Kim, Seungryong Kim, and Jaegul Choo. Robustnet: Improving domain generalization in urban-scene segmentation via instance selective whitening. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11580–11590, 2021. 2 [10] Alexey Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 2 [11] Qiong Dou, Daniel Caro de Castro, Konstantinos Kamnitsas, and Ben Glocker. Domain generalization via model-agnostic learning of semantic features. In Advances in Neural Information Processing Systems, pages 6450–6461, 2019. 2 [12] Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick P´ erez, and Raoul De Charette. Poda: Prompt-driven zeroshot domain adaptation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 18623– 18633, 2023. 2 [13] Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick P´ erez, and Raoul de Charette. A simple recipe for languageguided domain generalized segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23428–23437, 2024. 2 [14] Hao Fang et al. Eva-02: A visual learner for more generalized visual representation learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3 [15] Hao Fang et al. Eva-clip: Improving vision-language models with masked modeling. arXiv preprint arXiv:2303.13495, 2023. 1,6,7 [16] Ronald A Fisher. On the mathematical foundations of theoretical statistics. Philosophical transactions of the Royal Society of London. Series A, containing papers of a mathematical or physical character, 222(594-604):309–368, 1922. 3,5 [17] Zhangwei Gao, Zhe Chen, Erfei Cui, Yiming Ren, Weiyun Wang, Jinguo Zhu, Hao Tian, Shenglong Ye, Junjun He, Xizhou Zhu, et al. Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance. Visual Intelligence, 2(1):1–17, 2024. 2 [18] Zeyu Han, Chao Gao, Jinyang Liu, Jeff Zhang, and Sai Qian Zhang. Parameter-efficient fine-tuning for large models: A comprehensive survey. arXiv preprint arXiv:2403.14608, 2024. 2 [19] Kaiming He et al. Masked autoencoders are scalable vision learners. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3,6 [20] Matthew D Hoffman, David M Blei, Chong Wang, and John Paisley. Stochastic variational inference. Journal of Machine Learning Research, 2013. 5 [21] Edward J Hu et al. Lora: Low-rank adaptation of large language models. International Conference on Learning Representations (ICLR), 2022. 2 [22] Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan AllenZhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021. 6 [23] Jiaxing Huang, Dayan Guan, Aoran Xiao, and Shijian Lu. Fsdr: Frequency space domain randomization for domain generalization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6891– 6902, 2021. 2 [24] Christoph H¨ ummer, Manuel Schwonberg, Liangwei Zhou, Hu Cao, Alois Knoll, and Hanno Gottschalk. Strong but simple: A baseline for domain generalized dense perception by clip-based transfer learning. In Proceedings of the Asian Conference on Computer Vision, pages 4223–4244, 2024. 6 [25] Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. Visual prompt tuning. In European Conference on Computer Vision, pages 709–727. Springer, 2022. 2,6 [26] Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan. Maple: Multi-modal prompt learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19113–19122, 2023. 1 15051