Full text
MULTI-STAGE PATCH DIFFUSION AND CONTRASTIVE-SDE Thesis submitted to National Institute of Technology Andhra Pradesh for the award of the degree of Bachelor of Technology by KOTYADA VENKATA NARENDRA 421202 PEDAMALLA NITHIN 421236 KOMMURU ABHISHEK REDDY 421172 Under the supervision of Dr. NAGESH BHATTU. S DEPARTMENT OF COMPUTER SCIENCE & ENGINEERING NATIONAL INSTITUTE OF TECHNOLOGY ANDHRA PRADESH TADEPALLIGUDEM-534101, INDIA APRIL 2025 ©2025 Venkata Narendra Kotyada. This is the author’s version of the undergraduate thesis.
DECLARATION This written submission is an accurate representation of our own thoughts and if there are any other ideas or words that have been borrowed from someone else, we have given them credit through proper citation. We assure you that we have not cheated in any way and all the information in this work is true as well as obtained legally. We acknowledge that failure to comply with these terms may result in severe consequences such as expulsion from college; moreover it could also lead to civil or criminal charges against us by those individuals whose intellectual property rights were violated due either not citing them correctly nor seeking permission where needed. PEDAMALLA NITHIN K VENKATA NARENDRA K ABHISHEK REDDY 421236 421202 421172
Acknowledgement It is a pleasant aspect that I have now the opportunity to express my gratitude for all of them. We owe our sincere gratitude to our project guide Dr.Nagesh Bhattu.S, Department of Computer Science, National Institute of Technology, Andhra Pradesh, who took keen interest and guided us all along, till the completion of our project work by providing all the necessary information and referred many websites . We avail ourselves of this proud privilege to express our gratitude to all the faculty of the department of Computer Science and Engineering at NIT Andhra Pradesh for emphasizing and providing us with all the necessary facilities throughout the work. We offer our sincere thanks to all our fellow mates and other persons who knowingly or unknowingly helped us to complete this project.
LIST OF FIGURES S.No FIGURE NAME PAGE NO. 1 Unified and Multi-stage Architecture 8 2 Architecture of contrastive model F 14 3 FID vs Iterations 18 4 Qualitative comparison of Contrastive SDE with several baselines on three I2I translation tasks 18 5 Comparison of faithfulness with initial time P 20 LIST OF TABLES S.No TABLE NAME PAGE NO. 1 Quantitative comparison of Contrastive SDE with several baselines on three I2I translation tasks. The results marked with * came from [25]. The default setting (λ= 500, initial time P= 0.5T), and †denotes the model results with the modified setting (λ= 25, initial time P= 0.6T) F 19 2 Results of baseline (Multi stage Diffusion) vs (Mutistage Patch diffusion) 19 3 Results of both score functions for different λfor Cat →Dog I2I task 20
Abstract In this work, we address two key problems in diffusion models: (1) Efficiency of diffusion models, and (2) unpaired image-to-image (I2I) translation. Score-based diffusion models have demonstrated state-of-the-art performance in generative tasks. Their ability to approximate complex data distributions through stochastic differential equations (SDEs) enables them to generate high-fidelity and diverse outputs, making them particularly effective for unconditional image generation and unpaired I2I translation. To enhance efficiency, we propose an integrated approach that combines Patch Diffusion with a Multi-Stage Framework to enhance both training efficiency and generative quality. Patch Diffusion introduces a patchwise training methodology that employs randomized patch sizes and conditional score functions with location-based coordinate channels, enabling faster convergence and improved performance on smaller datasets. The Multi-Stage Framework, on the other hand, optimizes parameter utilization across different diffusion timesteps by segmenting them into distinct intervals, employing tailored multi-decoder U-Net architectures, and addressing inefficiencies such as gradient dissimilarity. For unpaired I2I translation, we propose a time-based contrastive learning method, Contrastive SDE, where a model is trained using SimCLR by treating an image and its domain-invariant representation as a positive pair. This helps the model retain features consistent across domains while suppressing those specific to individual domains. The trained contrastive model is then used to guide the inference process of a pretrained SDE for image-to-image translation. Experimental results demonstrate a substantial reduction in training time, although the FID scores remain suboptimal across datasets such as CIFAR10 and CelebA, highlighting both the promise and current limitations of the proposed framework in democratizing diffusion model training for broader applications. We further conduct empirical comparisons of Contrastive SDE with several baselines across three standard unpaired I2I tasks, evaluated using four common metrics. Contrastive SDE achieves performance comparable to state-of-the-art methods on several metrics. Moreover, our model converges significantly faster and requires neither label supervision nor classifier training, making it a more efficient alternative for unpaired image-to-image translation.
TABLE OF CONTENTS Page No. Title i Project work Approval ii Declaration iii Certificate iv Acknowledgements v List of Figures vi List of Tables Abstract vii Table of Contents viii
NIT ANDHRA PRADESH Contents 7 1 Introduction 1 2 Literature Review 2 3 Preliminary 3 3.1 Denoising Diffusion Probabilistic Models . . . . . . . . . . . . . . . . . . 3 3.2 Score Based Diffusion Models . . . . . . . . . . . . . . . . . . . . . . . . 5 3.3 ContrastiveLearning ............................ 6 3.3.1 Unsupervised Contrastive Loss . . . . . . . . . . . . . . . . . . . . 6 3.3.2 Semi-Supervised Contrastive Learning . . . . . . . . . . . . . . . . 6 3.3.3 Supervised Contrastive Loss . . . . . . . . . . . . . . . . . . . . . 6 4 Multi-Stage-Patch Approach 7 4.1 Multi-StageFramework........................... 7 4.1.1 Multi-stage U-Net Architectures . . . . . . . . . . . . . . . . . . . 7 4.1.2 Optimal Denoiser based TimeStep Clustering . . . . . . . . . . . . . 8 4.2 PatchDiffusion............................... 9 4.2.1 Patch-Wise Score Matching . . . . . . . . . . . . . . . . . . . . . 9 4.2.2 Progressive and Stochastic Patch Size Scheduling . . . . . . . . . . . 10 4.2.3 Conditional Coordinates for Patch Location . . . . . . . . . . . . . . 10 5 Contrastive-SDE 11 5.1 Contrastive Model Training . . . . . . . . . . . . . . . . . . . . . . . . . 11 5.2 GuidingDiffusion ............................. 12 6 Datasets 13 7 Experiments 14 7.1 Implementation............................... 14 7.2 Evaluationmetrics ............................. 15 7.3 Results................................... 15 7.4 Analysis.................................. 15 7.5 Ablation.................................. 18 8 Conclusion 19 7 7
NIT ANDHRA PRADESH 1 Introduction Diffusion models have gained prominence for their generative capabilities across a wide range of applications which includes unconditional image generation, image inpainting, super-resolution, text-to-image, and video generation. Despite their impressive generative abilities, diffusion models face challenges with slow training and sampling, limiting their use in applications that require real-time generation. These challenges mainly arise from the need of iterative forward and reverse diffusion processes with a large model of many parameters to train and infer over multiple timesteps. Addressing this challenge, one of the approaches introduced by [27] A Multi-Stage Framework comprising two components: (1) a U-net architecture with multiple decoders, and (2)a novel algorithm for partitioning timesteps into separate stages through clustering. Complementing this, Patch Diffusion [26], a general patch-based training framework, substantially reduces training time while enhancing data efficiency. This method achieves faster training, enhanced data efficiency, and robust performance on smaller datasets. By integrating both approaches, we propose a unified training approach that combines patch-wise conditioning with multi-stage optimization. This unified framework offers a promising solution for accessible, efficient, and scalable diffusion model training. While training a diffusion model is generally more computationally intensive than inference, our focus is on a task that can be addressed using inference alone that is unpaired image-to-image translation. In this setting, there is no direct correspondence between samples from the source and target domains. To solve this problem, we rely solely on a pretrained diffusion model trained on the target domain. The objective is to guide the inference process of this model to translate an image from the source domain into the target domain, while retaining the essential characteristics of the source. These characteristics, referred to as domain-invariant features, include aspects such as color and texture. Meanwhile, features specific to the source domain—such as pose or identity—should be removed. Formally, a successful translation preserves domain-invariant features while discarding domain-specific ones. Early approaches to unpaired image-to-image translation relied on GAN-based methods [7, 16]. However, these models often face issues such as mode collapse, unstable training, and limited diversity in generated outputs. In recent years, score-based diffusion models (SBDMs) [17] have emerged as state-of-the-art for generating high-quality, diverse images. Existing diffusion models often use classifier guidance to ensure that generated images retain domain-invariant features—such as color and texture—while discarding domain-specific features like identity, pose, or shape. More recent strategies introduce energy-based guidance [23] to guide generation toward semantically meaningful outputs. We take a different perspective by framing this challenge as a contrastive learning problem—where the goal is to pull domain-invariant features closer in the representation space while pushing domain-specific features apart. Motivated by this perspective, we propose a timedependent contrastive model that learns to retain domain-invariant features and discard domainspecific ones at various diffusion time steps. This model is then used to guide the inference of a pretrained diffusion model, leading to translations that match the target domain while ignoring 1
NIT ANDHRA PRADESH irrelevant source-specific traits. To evaluate both the guidance strategy and training efficiency, we conduct experiments across two settings. First, we assess our contrastive guidance model for unpaired image-to-image translation on standard datasets such as AFHQ [13] and CelebA [9], covering tasks like Cat → Dog, Wild →Dog, and Male →Female. Our method performs competitively with state-ofthe-art approaches across multiple evaluation metrics, while offering a simpler and more targeted alternative to classifier-based guidance. Second, to validate the training efficiency of our unified framework, we evaluate the multi-stage patch diffusion approach on CIFAR-10 and CelebA for unconditional image generation. The results demonstrate improved training speed and data efficiency but, quality of generated image was comprimised. 2 Literature Review As generative models, diffusion-based approaches have shown remarkable effectiveness, excelling in tasks such as image synthesis, video generation, and text-to-image translation. Building on their foundational principles, numerous methods have been proposed to address the challenges of training efficiency and scalability. One direction focuses on optimizing the training process. Techniques such as DDIM [20], TDPM [24], and EDM [21]-Sampling have reduced inference times but have limited impact on overall training efficiency. Meanwhile, hierarchical approaches like Cascaded Diffusion stabilize training by adopting multi-resolution designs, though these methods often require significant computational resources. The Multi-Stage Framework [27] takes this further by segmenting timesteps into manageable intervals, leveraging shared encoder architectures with interval-specific decoders to reduce parameter redundancy and enhance gradient consistency. This method has demonstrated significant improvements in training efficiency and scalability across large-scale datasets. A complementary area of exploration addresses data efficiency and computational bottlenecks through patch-based strategies. The Patch Diffusion [26] framework introduces a patch-wise training methodology, incorporating randomized patch sizes and location-based conditioning. These innovations allow the model to learn cross-region dependencies while dramatically reducing computational costs. Patch Diffusion achieves remarkable results on smaller datasets, significantly lowering the barrier to entry for researchers and practitioners with limited resources. By merging these ideas, a unified methodology integrating Patch Diffusion into the MultiStage Framework provides an opportunity to leverage the strengths of both approaches. While the Multi-Stage Framework ensures efficient parameter utilization and resource allocation, Patch Diffusion adds a scalable patch-based conditioning mechanism that enhances both training speed and generative quality. Together, these advancements promise to push the boundaries of diffusion model accessibility, scalability, and efficiency, making them more practical for diverse applications. As diffusion models are computationally expensive to train, one of the tasks that does’nt require training a diffusion model from scratch is guiding a diffusion model inference for an image-to2
NIT ANDHRA PRADESH t2←arg min τEt∼[τ,1][S(µ∗ t, µ∗ 1)] ≥α. Algorithm 1 Optimal Denoiser-Guided Timestep Clustering •1: Input: Total samples K, optimal denoiser function µ∗ ϕ(x, t), thresholds α,η, dataset pdata 2: Initialize: S0=S1=∅ 3: Output: Timesteps t1,t2 4: for k= 1 to Kdo 5: Sample yk∼pdata, ϵk∼ N(0, I), tk∼[0,1] 6: S0 k← D(ϵtk, ϵ0, yk, ϵk) 7: S1 k← D(ϵtk, ϵ1, yk, ϵk) 8: S0← S0∪{(tk, S0 k)} 9: S1← S1∪{(tk, S1 k)} 10: end for 11: t1= arg maxτ P (tk,S0 k)∈S0 S0 k·1(tk≤τ) P (tk,S0 k)∈S0 1(tk≤τ)≥α 12: t2= arg minτ P (tk,S1 k)∈S1 S1 k·1(tk≥τ) P (tk,S1 k)∈S1 1(tk≥τ)≥η 4.2 Patch Diffusion 4.2.1 Patch-Wise Score Matching We build a denoiser, Dθ(x;σt). For every patch that is divided, we use score-matching just like the score function in SBDM. The goal of the proposed method is to minimize the expected L2 denoising error for samples drawn from the data distribution p(y), independently for each σt: Ey∼p(y)Eµ∼N(0,σ2 tI)∥Dθ(y+µ;σt)−x∥2 2(24) Thus, we define the score function as: sθ(y, σt) = Dθ(y;σt)−y σ2 t .(25) Rather than applying score matching to entire images, we propose learning the score function on patches of random sizes. For any y∼p(y), We begin by randomly cropping small patches yi,j,s, where (i, j)denotes the top-left corner pixel coordinates of the patch, and srepresents the patch size (e.g., s= 32). Denoising score matching is then performed on these patches, conditioned on their corresponding sizes and locations, and can be formulated as: Ey∼p(y),µ∼N(0,σ2 tI),(i,j,s)∼U∥Dθ(˜yi,j,s;σt, i, j, s)−yi,j,s∥2 2(26) where ˜yi,j,s =yi,j,s +µ, and Udenotes the uniform distribution over the corresponding value ranges (e.g., i∼[−1,1]). Consequently, our conditional score function sθ(y, σt, i, j, s)is 9
NIT ANDHRA PRADESH defined over each local patch, with the scores being learned by conditioning on the patch’s location and size. Training with Equation (26) significantly accelerates convergence, as it operates on small local patches. However, a key challenge arises from the fact that the score function sθ(y, σt, i, j, s) observes local regions and may fail to capture global, cross-patch dependencies. In other words, the scores predicted from neighboring patches must collectively form a coherent score field to enable consistent image reconstruction during sampling. To overcome this limitation, we introduce two strategies: 1. Random Patch Sizes: Patch sizes are sampled from both small and large patches. Larger patches can be interpreted as compositions of smaller ones, encouraging sθto learn how to integrate scores across local regions into a unified and coherent score map when trained with larger patches. 2. Incorporating Full-Size Images: To ensure that the reverse diffusion process reliably converges to the original data distribution, a small proportion of full-resolution images is included during training. 4.2.2 Progressive and Stochastic Patch Size Scheduling We introduce a patch-size scheduling strategy. Let lr epresent the proportion of training iterations that use full-resolution images as input, and let Rdenote the original resolution of image. The available patch size options are defined as: s∼ls:= Rwhen s=R, 3 5(1 −l)when s=R 2, 2 5(1 −l)when s=R 4. (27) We consider two approaches for patch-size scheduling: 1. Stochastic: In this method, for each mini-batch during training, we randomly sample the patch size s∼lswhere the probability mass function is defined in Equation (27). 2. Progressive: Here, we gradually increase the patch size throughout training. For the first 2 5(1 −l)portion of the training iterations, we fix the patch size to s=R 4. In the next 3 5(1 −l)iterations, we increase the patch size to s=R 2.Finally, full-size images are used for the remaining lfraction of the total training iterations Empirically, we observe setting l= 0.5with stochastic scheduling strikes a good balance between training speed and generation quality. Regarding the denoiser Dθhandle image patches of varying sizes. This is made possible by the fully convolutional nature of the U-Net architecture [5]. Since convolutional layers can adapt to different input resolutions by sliding their filters across the spatial dimensions, our patch-based training approach seamlessly integrates with any U-Net-based diffusion model. This architectural flexibility also contributes to faster and more efficient sampling. 4.2.3 Conditional Coordinates for Patch Location To simplify how patch locations are incorporated, we introduce a pixel-level coordinate system. Specifically, we normalize pixel coordinates to the range[−1,1] based on the original image 10
NIT ANDHRA PRADESH resolution, where the top-left corner is set to (−1,−1) and the bottom right corner as (1,1). For any image patch xi,j,s, we extract its pixel coordinates iand jand encode them as two additional channels. During training, each sample in a batch is independently and randomly cropped using a sampled patch size, and the corresponding coordinate channels are computed. These two coordinate channels are then concatenated with the original image channels to serve as input to the denoiser Dθ. When evaluating the loss specified in Equation (27), we exclude the reconstructed coordinate channels and focus solely on minimizing the loss for the image channel. The combination of the pixel-level coordinate system and random patch sizes can be interpreted as a form of data augmentation. For instance, given an image of resolution 128 ×128, and patch size s= 32, there are (128 −32 + 1)2= 9409 possible distinct patch locations. This diversity in patch selection effectively increases the variety of training samples. Therefore, we argue that training diffusion models on patch-wise data enhances their data efficiency. By using patch training approach for Multi stage architecture can significantly improve the training efficiency of the model. Experiments are conducted on both the approaches 5 Contrastive-SDE For unpaired image to image translation, we introduce a time-dependent contrastive learning method where a neural network Fis trained to focus on domain-invariant features while suppressing domain-specific ones using the contrastive NT-Xent loss. This network is later used to guide the inference of a pretrained SDE via a guidance function Q. 5.1 Contrastive Model Training The core idea is to treat the image and its domain-invariant feature as a positive pair, while considering the image and its domain-specific feature as a negative pair. However, the NT-Xent loss treats all samples except the positive as negatives, which introduces a challenge in our setup. To address this, we ensure that domain-specific features are implicitly ignored during training. A theoretical explanation for this behavior is provided in the appendix. We begin by training a U-Net-based model F, where each input is conditioned on the diffusion timestep through a learned time embedding. As shown in Figure 2, this architecture consists of a U-Net with residual and attention blocks operating at multiple scales. For each image xand its low-pass filtered version ¯x, the model extracts hidden representations hiand hj. A projection head then transforms these into contrastive features ziand zj, which are used to compute the NT-Xent loss as follows: L=1 2N N X k=1 [ℓ(2k−1,2k) + ℓ(2k, 2k−1)] (28) 11
NIT ANDHRA PRADESH Figure 2: Architecture of the contrastive model F ℓi,j =−log exp(sim(zi, zj)/τ) P2N k=1 1 [k=i]exp(sim(zi, zk)/τ)(29) Here, sim(u, v) = uTv |u||v|is the cosine similarity between vectors uand v, and τis a temperature parameter. The indicator function 1 [k=i]is 1 when k=iand 0 otherwise. This contrastive setup is simpler to train compared to building a classifier from scratch for guidance purpose. 5.2 Guiding Diffusion Contrastive-SDE models the conditional distribution p(y0|x0)by combining a pretrained SDE with a guidance function Q(y, x, t)derived from the contrastive model F. The reverse-time guided SDE is given by: dy=f(y, t)−g(t)2(sθ(y, t)−∇yQ(y, x0, t))dt+g(t),d ¯w, (30) where wis the standard reverse-time Wiener process and sθ(y, t)is the score function from the pretrained SDE. The guidance function Qis defined as: Q(y, x, t) = −λS(y, x, t)(31) The similarity function S(y, x, t)measures how similar the generated sample xtis to the source image y, based on hidden representations from the contrastive model F. Let ht=F(xt, t) and h0=F(y, t)be the features of shapes C×M×N, then: 12
NIT ANDHRA PRADESH S(y, xt, t) = 1 MN M X m=1 N X n=1 hmn t⊤hmn 0 |hmn t|2,|hmn 0|2 (32) Here, hmn represents the channel-wise feature at spatial location (m, n). Alternatively, we can define Susing negative squared L2distance: S(y, xt, t) = −|h0−ht|2 2(33) We compare these similarity metrics through ablation studies. Similar to Eq. (??), the update equation for the guided diffusion becomes: yt−1=yt−[f(yt, t)−g(t)2(sθ(yt, t)−∇yQ(yt, x0, t))] + g(t)z, z ∼ N(0, I). (34) Algo. 2 shows the inference of pretrained SDE guided by contrastive model. Algorithm 2 Contrastive-SDE for unpaired image-to-image translation Require: Source image x0, initial time P, denoising steps R, hyper-parameter λ, similarity function S(·,·), score function s(·,·) 1: y∼qP|0(y|x0)▷Start point 2: l←P R 3: for i=Rto 1do 4: t←il 5: x∼qt|0(x|x0)▷Sample perturbed source image 6: Q(y,x, t)← −λS(y,x, t) 7: y←y−[f(y, s)−g(t)2s(y, t)−∇yQ(y,x, t)]h 8: if i > 1then 9: z∼ N(0,I) 10: else 11: z←0 12: end if 13: y←y+g(t)√lz 14: end for 15: y0←y 16: return y0 6 Datasets •CelebA: CelebA is a large-scale dataset for facial recognition, containing over 200,000 celebrity images across 10,177 identities. It is commonly used for tasks like facial recognition, attribute prediction, and generative models. The original images are 178x218 pixels. CelebA provides 40 attribute labels, such as "Smiling", "Male", "Young", "Eyeglasses", etc., for each face, making it popular for both supervised and semi-supervised learning tasks. When resized to 32x32, much of the detailed facial features and facial attribute information will be lost due to the reduction in resolution. The dataset may still be useful for tasks 13
NIT ANDHRA PRADESH like testing models on low-resolution images. It is widely used for facial recognition and generative adversarial networks (GANs), though resizing it to lower resolutions reduces the available information. •CIFAR-10: CIFAR-10 is a well-known image dataset for object classification. It contains 60,000 32x32 color images in 10 different classes, with each class containing 6,000 images. The classes are: airplane, automobile, bird, cat, deer, dog, frog, horse, ship, and truck. The images are originally 32x32 pixels, and the dataset is often used for image classification tasks and for testing various neural network architectures. Resizing CIFAR-10 images is typically unnecessary since the original images are already 32x32. However, resizing them to larger or smaller sizes can be used to test the robustness of models to resolution changes. CIFAR-10 is commonly used for benchmarking models in image classification tasks, particularly with Convolutional Neural Networks (CNNs) •AFHQ: AFHQ (Animal Faces-High Quality) is a high-resolution dataset designed for image-to-image translation and generative modeling. It consists of 15,630 high-quality 512×512 color images across three domains: cats, dogs, and wildlife (e.g., foxes, tigers, lions). Each domain contains roughly 5,000 images, with diverse breeds, poses, and backgrounds. Unlike lower-resolution datasets, AFHQ’s large image size makes it suitable for testing high-fidelity synthesis and domain transfer tasks. The dataset is often used to evaluate generative adversarial networks (GANs), diffusion models, and other image translation methods. AFHQ is particularly popular for benchmarking unpaired translation methods (e.g., CycleGAN, StarGANv2) due to its clear domain separation and visual richness. 7 Experiments 7.1 Implementation For the efficiency strategy, we trained both the baseline model and our proposed model for approximately 5×105iterations. We used a batch size of 256 instead of the default 128 from the baseline, applying this setting consistently across both CelebA and CIFAR-10 datasets while keeping all other training parameters identical to the multi-stage diffusion approach. For unpaired image-to-image translation tasks, we trained the contrastive model for 5,000 iterations with a batch size of 32, learning rate of 3e-4, and weight decay of 0.05 on AFHQ and CelebA datasets. For all unpaired translation tasks, we maintained the default settings of λ= 500 and P= 0.5T, though we adjusted these to λ= 25 and P= 0.6T when optimizing specifically for FID improvement. All experiments are conducted on NVIDIA RTX A6000 GPU. 14
NIT ANDHRA PRADESH 7.2 Evaluation metrics For efficiency, we use metrics FID, Iterations and time taken for training the diffusion models. And for Unpaired I2I translation, we evaluate our results using two key criteria: realism and faithfulness. To measure realism, we compute the Fréchet Inception Distance (FID) [6] between the translated images and the target domain. For faithfulness, we quantify how well the output preserves the input’s structure using L2 distance, PSNR, and SSIM [3] between source and translated image pairs. 7.3 Results The quantitative results of baseline model from our experiments and results reported in [27], we achieve better FID scores (improving from 1.44 to 1.16) in fewer iterations (4×105vs 4.3×105) just by changing batch size from 128 to 256. This suggests that using iteration count alone as an efficiency metric may be misleading. Whereas our approach reduced training time by about 25%, the final FID results still fell short of expectations on both datasets. For the unpaired I2I we have achieved remarkable results using contrastive SDE. We evaluate Contrastive-SDE against several state-of-the-art (SOTA) baselines: ILVR, EGSDE (with results reproduced from publicly available code), and SDDM (results reported in [25]). The quantitative comparison in Table 1, demonstrates that Contrastive-SDE achieves near state-of-the-art performance in faithfulness metrics (L2, PSNR, SSIM), often closely matching EGSDE. For realism, as measured by FID, Contrastive-SDE does not reach the SOTA level but shows superior performance compared to ILVR in default setting. Furthermore, we demonstrate that FID can be improved by adjusting some hyperparameters, offering a flexible trade-off with other metrics. Qualitative results, illustrated in Fig. 4, showcases the quality of generated samples across all baselines. We also compare the training cost of our contrastive model with the classifier-based setup used in EGSDE on the Cat →Dog task. As shown in Table 3, EGSDE requires approximately 7 hours of training for 5K iterations, along with labeled data to train the domain-specific classifier. In contrast, our contrastive model converges within 2K iterations in just 2 hours with no labeled data. Additionally, the average training speed is 3.6 seconds per iteration for Contrastive-SDE, compared to 5.06 seconds per iteration for EGSDE. These results show that our method not only removes the need for external classifier supervision but also offers a simpler and more efficient training pipeline. 7.4 Analysis The Table 2 displays the time taken for both models during training for 4.0×105iterations, indicating a promising 25% reduction in training time The FID score after implementing patch diffusion for CelebA, and CIFAR10 datasets the FID 126, 240 which is not a good score for a generative model. The poor FID performance of our patch-based multi-stage diffusion approach may be attributed to working with low-resolution 32×32 images. When these small images are divided into 8×8 patches, most patches contain insufficient meaningful information for the model to 15
NIT ANDHRA PRADESH Figure 3: FID vs Iterations Figure 4: Qualtitative comparison of Contrastive SDE with several baselines on three I2I translation tasks 16
NIT ANDHRA PRADESH Method FID ↓L2 ↓PSNR ↑SSIM ↑ Cat →Dog ILVR [18] 74.82 ±0.88 57.04 ±0.16 17.78 ±0.02 0.360 ±0.001 EGSDE [23] 65.88 ±0.71 48.68 ±0.08 19.31 ±0.01 0.415 ±0.001 SDDM* [25] 62.29 ±0.63 – – 0.422 ±0.001 Contrastive-SDE 72.61 ±0.72 48.72 ±0.07 19.31 ±0.01 0.425 ±0.001 Contrastive-SDE†61.68 ±0.71 59.41 ±0.10 17.54 ±0.02 0.383 ±0.001 Wild →Dog ILVR [18] 74.85 ±0.90 63.51 ±0.10 16.84 ±0.02 0.304 ±0.001 EGSDE [23] 60.13 ±0.52 55.97 ±0.08 18.13 ±0.01 0.342 ±0.001 SDDM* [25] 57.38 ±0.53 – – 0.328 ±0.001 Contrastive-SDE 66.72 ±0.62 56.29 ±0.11 18.08 ±0.01 0.345 ±0.001 Contrastive-SDE†59.86 ±0.39 66.19 ±0.15 16.64 ±0.02 0.304 ±0.001 Male →Female ILVR [18] 46.11 ±0.44 52.10 ±0.09 18.60 ±0.01 0.511 ±0.001 EGSDE [23] 42.31 ±0.41 43.00 ±0.03 20.35 ±0.01 0.574 ±0.000 SDDM* [25] 44.37 ±0.23 – – 0.526 ±0.001 Contrastive-SDE 50.16 ±0.18 44.18 ±0.04 20.11 ±0.01 0.574 ±0.001 Contrastive-SDE†45.15 ±0.15 53.53 ±0.06 18.44 ±0.02 0.533 ±0.001 Table 1: Quantitative comparison of Contrastive SDE with several baselines on three I2I translation tasks. The results marked with * came from [25]. The default setting (λ= 500, initial time P= 0.5T), and †denotes the model results with the modified setting (λ= 25, initial time P= 0.6T) . learn effectively. This explains why we observed reduced training time but unsatisfactory FID results. For the unpaired translation tasks, the moderate FID performance might result from using low-pass filtered versions of images during domain-invariant feature extraction. These filters may not completely eliminate domain-specific information, allowing some to persist in the hidden representations. While more sophisticated feature extractors could potentially improve results, our current faithfulness metrics confirm that the contrastive model successfully preserves domain-invariant features during the diffusion process. Model Dataset p Training Time FID ↓ Multistage EDM CelebA 0.5 8 days 1.16 Multistage-Patch EDM CelebA 0.5 6 days 126 Multistage EDM CIFAR10 0.5 7 days 2.67 Multistage-Patch EDM CIFAR10 0.5 5 days 4 hrs 240 Table 2: Results of baseline (Multi stage Diffusion) vs (Mutistage Patch diffusion). 17
NIT ANDHRA PRADESH Method Training time ↓Iterations ↓sec/iteration ↓batch size EGSDE 7hr 5000 5.04 32 Contrastive SDE 2hr 2000 3.6 32 Table 3: Comparison of training cost for Cat →Dog task . Figure 5: Comparison of faithfulness with inital time P 7.5 Ablation We conduct ablation study on unpaired I2I task on parameters like initial time and score function. Choice of Initial time P:Fig 5 illustrates an inverse relationship between P and image faithfulness: as P increases, realism improves, but faithfulness declines. Choice of Score function S:Table 4 shows the effect of score function on metrics on different λ. Regarding score functions, the NS L2 similarity demonstrated greater sensitivity to λ. changes compared to cosine similarity, with the quantitative results clearly showing the expected trade-off between faithfulness metrics and realism metrics. Simlarity λFID ↓L2 ↓PSNR ↑SSIM ↑ Cosine 500 72.61 48.72 19.31 0.425 Cosine 150 72.96 49.01 19.24 0.423 NS L25e-05 72.86 48.73 19.30 0.424 NS L25e-03 77.80 46.42 19.75 0.428 Table 4: Results of both score functions for different λfor Cat →Dog I2I task . 18