Full text
MambaCT: Feature Enhancement-Based Low-Dose CT Image Denoising Using Vision Mamba and a Scaling Adapter Abdelkarim Cherhabil1, Lahc`ene Mitiche1, and Amel Baha Houda Adamou-Mitiche1 1Department of Electronics and Telecommunications, Ziane Achour University of Djelfa, Laboratoire de Mod´elisation, Simulation et Optimisation des Syst`emes Complexes R´eels Abstract With the rapid development of Mamba models, Vision Mamba is replacing convolutional neural networks (CNNs) and vision transformers (ViTs), emerging as the dominant trend in computer vision tasks. After achieving remarkable success in natural language processing, Mamba models have garnered increasing interest in the medical imaging community for their ability to understand global context. However, there has been limited research on medical image denoising based on the Vision Mamba architecture. In this paper, we propose MambaCT, a model that integrates the key features of UNet and Mamba with a Scaling Adapter in the visual state space (VSS) Block. Additionally, we propose skip connection spatial-channel processing attention (SCSPA) to enhance feature integration and robustness as a pathway in place of traditional skip connections. MambaCT outperforms previous state-of-the-art (SOTA) models across various architectures in both visual quality and quantitative performance, requiring only 0.83G MACs and achieving an SSIM of 0.9104 and RMSE of 9.3423. The model was evaluated on the AAPM-Mayo Clinic low-dose computed tomography (LDCT) Grand Challenge Dataset. Keywords: Low-dose CT, Vision Mamba, Medical Image Denoising, Adapter, State Space Models, Auto-encoder. 1 Introduction Computed tomography is a diagnostic imaging method that precisely aligns X-ray, gamma, ultrasound, and ion beams to create cross-sectional images of the human body [1]. Clinical, industrial, and other fields make extensive use of CT [1] [31]. It is particularly effective in reconstructing organ structures at various depths and angles [2] [3]. Several algorithms have been created to improve image quality in low-dose CT (LDCT) scans in order to address this issue. Using physical models and existing data, researchers employ iterative techniques in classical ways to reduce noise and artifacts. For instance, some image priors are expressed as sparse transforms utilizing compressive sensing (CS) to address issues in internal CT, low-dose, few-view, and finite-angle CT [4]. Examples include dictionary learning [5], lowrank [6], non-local means (NLM) [7][8][9], total variation (TV), and its variations [10][11][12][13], among other methods. Li et al. [49] reconstructed feature similarities in large neighborhood images using NLM. Aharon et al. used dictionary learning [50] to denoise LDCT images, drawing inspiration from sparse representation theory, resulting in considerable improvement in denoising quality while reconstructing abdominal images [51]. Block-matching 3D (BM3D) has been shown by Feruglio et al. to be efficient for a range of X-ray imaging applications [52]. However, this method’s inaccuracy in determining the noise distribution in the image domain prevents the optimal balance between noise reduction and structure preservation. Due to restrictions on data volume, the accuracy of these conventional approaches is typically still poor [14]. Since the advent of deep learning, CNNs have been the dominant method for denoising low-dose CT (LDCT) images. CNNs extract features through convolutional operations, where the kernel moves across the entire image, resulting in a relatively small parameter volume for the CNN model. This approach effectively captures significant local features, and the network’s receptive field is incrementally expanded through layer stacking. Numerous traditional deep learning algorithms have been applied to the field of low-dose CT (LDCT) image denoising for reconstructing high-quality images. These include convolutional neural networks (CNNs) [15] [16] [29] [30], encoder-decoder networks with residual connections [17][18][19], and generative adversarial networks (GANs) [20][21]. The work of Chen et al. can be considered groundbreaking, as they were among the first to utilize convolution, deconvolution, and shortcut 75
connections to design a prototype of a residual encoder-decoder convolutional neural network, known as RED-CNN [17]. To improve the quality of denoised photos, Yang et al. used a generative adversarial network with Wasserstein distance (WGAN) and a perceptual loss mechanism [20]. Compared to other CT denoising techniques, Fan et al. developed a quadratic neuron-based autoencoder that is more resilient and useful for model efficiency [?]. The retrieval of detailed structural details in the denoised images may be adversely affected by CNNs’ limits in capturing long-range contextual information within images, notwithstanding their intriguing results for LDCT [23]. The integration of Transformer models into image denoising has significantly improved performance, resulting in higher accuracy and reduced processing times [24][25]. This advancement has brought about a revolution in image processing. Recent studies reveal that Transformer modules can effectively replace traditional convolutions in deep neural networks. They work by processing sequences of image patches, leading to the development of Vision Transformers (ViTs). Dosovitskiy et al. first proposed the vision transformer (ViT) in the CV field by mapping an image into 16×16 sequence words [24]. Wang et al. propose an innovative approach called the Convolution-free Token2Token Dilated Vision Transformer [26]. Luthra et al. introduce a fresh approach named Eformer, which stands for Edge Enhancement-based Transformer. Eformer is a unique architectural framework that constructs an encoder-decoder network using transformer blocks [27]. Jian et al. propose SwinCT, which utilizes a feature enhancement module (FEM) inspired by the Swin Transformer architecture. The FEM in SwinCT is employed to capture and enrich the high-level features within medical images [28]. According to the above analysis, Mamba models offer significant advantages over both CNN and transformer models, including greater visual interpretability due to their intrinsic Visual State Space (VSS) blocks [32]. Beyond their effectiveness, Mamba models are appealing to physicians because their self-explanatory nature allows doctors to understand the model’s reasoning. ¨ Ozt¨urk et al. [33] pioneered the application of an innovative SSM architecture, named DenoMamba, to enhance LDCT image denoising without increasing model complexity. This novel approach employs an hourglass-shaped structure, featuring encoder-decoder stages built with custom-designed FuseSSM blocks. Li et al. [34] introduced a CACTSR, which integrates VMamba and Transformer technologies with Mixed Attention Blocks and Cross Attention Blocks to enhance feature utilization and facilitate cross-window information interaction. The above studies demonstrate the promising results of Mamba-based deep learning models in visual tasks. Motivated by the aforementioned study, we introduce MambaCT, a paradigm that combines the key components of Mamba and UNet by integrating a Scaling Adapter within the VSS Block. Our approach incorporates an SCSPA module pathway in place of traditional skip connections, resulting in images that exhibit superior performance both quantitatively and visually. According to experimental results, our model outperforms other state-of-the-art models, achieving the highest SSIM value and the lowest RMSE value. The main contributions of this paper are as follows: 1. We propose MambaCT, a model that utilizes a U-Net-based Mamba network architecture and incorporates a Scaling Adapter within the VSS Block. The Scaling Adapter enhances the restoration of detailed and structural information in denoised images. 2. We propose the Skip Connection Spatial-Channel Processing Attention (SCSPA) module pathway as an alternative to traditional skip connections. 3. Assess the model’s performance by comparing it with previous works using various metrics, including MACs, SSIM, and RMSE. 76
2 Methods The architecture of the proposed MambaCT, as shown in Fig. 1, draws inspiration from both U-Net [35] and VMamba [36]. Designed specifically for LDCT denoising, MambaCT comprises four key modules: 1) Patch Extraction, 2) VSS Blocks, 3) SCSPA Module, and 4) Resizing Modules. Figure 1: The overall structure of MambaCT (a).The VSS Block serves as the primary building block of MambaCT, with SS2D and the Adapter as its core operations (b). 2.1 Patch extraction To train deep learning models effectively, a large volume of samples is crucial, which can be particularly challenging in clinical imaging. In our study, we addressed this issue by using CT scans with overlapping slices. This method has proven to be both effective and successful, as it helps capture perceptual differences in local regions and greatly boosts the number of samples available [37][38][39]. 2.2 VSS Block The core of MambaCT is the VSS Block, which serves as the primary building block of the model. The VSS Block incorporates SS2D and the Adapter as its core operations and is derived from [36], as shown in Figure 1(b). Its structure begins with Layer Normalization, followed by a split into two branches. The first branch applies a linear layer and the activation function SiLU [40]. The second branch processes the input through a linear layer, depthwise separable convolution, activation, and the SS2D module. the SS2D module provides contextual information to image patches via a compressed hidden state along scanning paths (Figure 2(b)), reducing computational complexity from quadratic to linear compared to self-attention mechanisms (Figure 2(a)). After SS2D, the features undergo Layer Normalization and are combined with the first branch’s output through element-wise multiplication. A Scaling Adapter is then applied, allowing for the learnable adjustment of the adapted features’ contribution. Directly after the adapter, we scale the embedding by a scale factor s[41]. Finally, the result passes through a linear mixing layer and is combined with a residual connection to produce the block’s output. this architecture efficiently processes spatial information while maintaining linear complexity, making it well-suited for medical image denoising in LDCT. 77
Figure 2: Comparison of correlation establishment between image patches via (a) self-attention and (b) the proposed 2D-Selective-Scan (SS2D). Red boxes indicate the query image patch, with patch opacity representing the degree of information loss. 2.2.1 2D-Selective-Scan for Vision Data A scan expansion operation, an S6 block, and a scan merging operation are the three primary parts of the SS2D module. According to Figure 3, SS2D first unfolds input patches into sequences along four distinct traversal paths (i.e., scan expanding), processes each patch sequence using a separate S6 block in parallel, and then reshapes and merges the resultant sequences to form the output map (i.e., scan merging). By adopting complementary 1D traversal paths, SS2D enables each pixel in the image to effectively integrate information from all other pixels in different directions, facilitating the establishment of global receptive fields in the 2D space. Figure 3: The overall structure of the 2D Selective Scan (SS2D) process. 78
2.2.2 Scaling Adaptor The Adapter operates as a bottleneck model. The down-projection layer reduces the dimensionality of the input embedding using a basic MLP layer with parameters Wdown ∈Rd׈ d, and the up-projection layer restores the compressed embedding to its original dimensionality with an additional MLP layer with parameters Wup ∈Rˆ d×d, where ˆ dis the bottleneck middle dimension and satisfies ˆ d < d. Additionally, there is a ReLU layer [42] between these projection layers for non-linear properties. Residual connections remain a crucial aspect, helping in training deeper networks by preventing gradient vanishing problems. 2.3 Skip Connection Spatial-Channel Processing Attention In contrast to using a single attention mechanism, the combination of channel attention and spatial attention, especially in a sequential manner, significantly enhances the model’s ability to capture important feature information [43]. Inspired by [44], we propose a SCSPA mechanism that applies sequential channel-spatial attention to the skip connections. As illustrated in Fig. 4, the SCSPA module consists of two key components: one for spatial attention and one for channel attention. A channel reduction action (C →C/rate), a ReLU activation, and a channel expansion operation (C/rate →C) comprise the Channel Attention Submodule. The two 7x7 convolutions that make up the Spatial Attention Submodule are followed by batch normalization (BN) and ReLU activation in the first one, and batch normalization and a sigmoid activation in the second. After that, element-wise multiplication is used to merge the outputs of these two routes. The dimensions of the input and the final output are [B, H, W, C]. Figure 4: The overall structure of SCSPA. 2.4 Resizing module The Patch Merging components function as downsampling mechanisms, diminishing the spatial dimensions of feature maps while amplifying the channel count. This approach enables the network to capture hierarchical features across various scales. Conversely, the Patch Expanding modules in the decoder act as counterparts to the Patch Merging modules in the encoder. They reverse the downsampling process, progressively restoring spatial resolution while decreasing the number of channels. 79
3 Experiment In this section, we begin by listing the languages and tools utilized: Python, PyTorch, and CUDA. Dataset:The publicly accessible clinical dataset from the 2016 NIH-AAPM Mayo Clinic LDCT Grand Challenge was used to train and evaluate the model [45]. Ten anonymous individuals’ 2,378 low-dose (quarter) and 2,378 normal-dose (full) CT scans with 3.0-mm whole-layer slices are included in this dataset. We chose patient L506’s data, which consists of 211 slice images with numbers ranging from 000 to 210, for testing. The model was trained using the data from the remaining nine cases. Experiment setup: PyTorch 1.11.0 [46] and CUDA 12.4.0 were used in the experiments, which were conducted on an Ubuntu 22.04 LTS system with an Intel(R) Core(TM) i7-12700k CPU @ 2.70 GHz. The four NVIDIA RTX 3070 Ti 8G GPUs were used to train the model. Four blocks were chosen at random from each image’s available slices for training. For 4,000 epochs, the batch size was fixed at 16. The ADAM-W optimizer, which has a learning rate of 1.0×10−5, was used to reduce the mean squared error loss. After training, the model’s performance was assessed using the standard metrics in the field. 4 Discussion To evaluate denoising performance in LDCT images, we retrained all models using their officially available code. We propose MambaCT, a model that combines key features from UNet and Mamba, enhanced with a Scaling Adapter in the VSS Block. Additionally, our approach incorporates an SCSPA module pathway in place of traditional skip connections to improve feature integration. As shown in Table 1for the L506 dataset, MambaCT achieved the highest quantitative metrics, surpassing all other methods. Table 1: Quantitative comparison of different methods on L506 in terms of learnable parameters (#param.), MACs, SSIM, and RMSE. Bold values represent our method’s performance. Method #param. MACs SSIM↑RMSE↓ LDCT – – 0.8759 14.2416 RED-CNN [17] 1.85M 5.05G 0.8952 11.5926 WGAN-VGG [20] 34.07M 3.61G 0.9008 11.6370 MAP-NN [47] 3.49M 13.79G 0.8941 11.5848 AD-NET [48] 2.07M 9.49G 0.9041 9.7166 MambaCT 62.08M 0.83G 0.9104 9.3423 To provide a thorough evaluation of denoising performance, we use both qualitative and quantitative methods. The quantitative analysis focuses on two key metrics: SSIM, and RMSE. Additionally, model complexity is assessed based on the number of trainable parameters (#param.) and multiply-accumulate operations (MACs). Table 1presents the average SSIM, and RMSE across all slices of L506. Our MambaCT model achieves the highest SSIM of 0.9104, the lowest RMSE of 9.3423. Figure 5presents the results of various networks on L506 with Lesion No. 575, while Figure 6displays the regions of interest (ROIs) from the rectangular area highlighted in Figure 5. Visual analysis of these figures 5and 6demonstrates MambaCT’s superior capability in achieving three key objectives: noise and artifact removal, maintenance of high-level spatial smoothness, and preservation of target image details. While RED-CNN, built on convolutional networks, shows proficiency in noise and artifact elimination while retaining image details, it faces limitations in structural recovery. This constraint stems from its computational architecture, which prioritizes high-frequency information extraction, such as texture details. Furthermore, RED-CNN’s effectiveness is hampered by its finite receptive field size, impeding comprehensive global information capture. Detailed examination of the ROIs in Figure 6reveals varying performance across methods: 1. WGANVGG and MAP-NN introduce unwanted artifacts, manifesting as additional shadows and tissue-like structures. 2. RED-CNN and AD-NET yield improvements in image clarity and smoothness compared to WGAN-VGG and MAP-NN, though residual blotchy noise persists around lesion areas. 80
Figure 5: various networks’ denoised findings on L506 with Lesion No. 575. These include LDCT (a), RED-CNN (b), WGAN-VGG (c), MAP-NN (d), AD-NET (e), MambaCT (f), and NDCT (g). The window for display is [-160, 240] HU. Figure 6: Fig. 5shows the ROIs of the rectangle. These include LDCT (a), RED-CNN (b), WGAN-VGG (c), MAP-NN (d), AD-NET (e), MambaCT (f), and NDCT (g). Comparatively, MambaCT performs best on all metrics: it effectively suppresses noise and artifacts, maintains high-level spatial smoothness, and preserves structural information in images that have been restored. The quantitative measurements shown in Table 1, where MambaCT consistently performs better than other comparative models. Concerning model complexity, MAP-NN has the highest MACs at 13.79G due to its numerous repeated modules, While MambaCT has the highest number of parameters at 62.08M, it remarkably uses the least MACs at 0.83G, demonstrating its computational efficiency. while WGAN-VGG has the greatest number of trainable parameters at 34.07M due to its use of VGG as a feature extractor. This balance between high performance and low complexity underscores the efficiency of MambaCT compared to other state-of-the-art methods, such as MAP-NN and WGAN-VGG and AD-NET, which exhibit higher complexity but lower performance. 81
5 Conclusion In this work, we propose MambaCT, a model that integrates a Scaling Adapter within the VSS Block, combining the essential elements of Mamba and UNet. To further enhance feature integration, our method replaces conventional skip connections with an SCSPA module pathway, resulting in images that demonstrate superior performance both quantitatively and visually. Experimental results indicate that our model outperforms other cutting-edge models, achieving the lowest RMSE value and the highest SSIM value. References [1] T. M. Buzug. Computed tomography. In Springer Handbook of Medical Technology, pages 311–342. Springer Berlin Heidelberg, Berlin, Heidelberg, 2011. [2] H. Zhang, L. Zhang, Y. Sun, and J. Zhang. Low dose CT image statistical reconstruction algorithms based on discrete shearlet. Multimedia Tools and Applications, 76:15049–15064, 2017. [3] A. H. Behzadi, Z. Farooq, J. H. Newhouse, and M. R. Prince. MRI and CT contrast media extravasation: a systematic review. Medicine, 97(9):e0055, 2018. [4] D. L. Donoho. Compressed sensing. IEEE Transactions on Information Theory, 52(4):1289–1306, 2006. [5] Q. Xu, H. Yu, X. Mou, L. Zhang, J. Hsieh, and G. Wang. Low-dose X-ray CT reconstruction via dictionary learning. IEEE Transactions on Medical Imaging, 31(9):1682–1697, 2012. [6] J. F. Cai, X. Jia, H. Gao, S. B. Jiang, Z. Shen, and H. Zhao. Cine cone beam CT reconstruction using low-rank matrix factorization: algorithm and a proof-of-principle study. IEEE Transactions on Medical Imaging, 33(8):1581–1591, 2014. [7] Y. Chen, D. Gao, C. Nie, L. Luo, W. Chen, X. Yin, and Y. Lin. Bayesian statistical reconstruction for low-dose X-ray computed tomography using an adaptive-weighting nonlocal prior. Computerized Medical Imaging and Graphics, 33(7):495–500, 2009. [8] J. Ma, H. Zhang, Y. Gao, J. Huang, Z. Liang, Q. Feng, and W. Chen. Iterative image reconstruction for cerebral perfusion CT using a pre-contrast scan induced edge-preserving prior. Physics in Medicine and Biology, 57(22):7519, 2012. [9] Y. Zhang, Y. Xi, Q. Yang, W. Cong, J. Zhou, and G. Wang. Spectral CT reconstruction with image sparsity and spectral mean. IEEE Transactions on Computational Imaging, 2(4):510–523, 2016. [10] E. Y. Sidky and X. Pan. Image reconstruction in circular cone-beam computed tomography by constrained, total-variation minimization. Physics in Medicine and Biology, 53(17):4777, 2008. [11] Y. Zhang, W. Zhang, Y. Lei, and J. Zhou. Few-view image reconstruction with fractional-order total variation. Journal of the Optical Society of America A, 31(5):981–995, 2014. [12] Y. Zhang, Y. Wang, W. Zhang, F. Lin, Y. Pu, and J. Zhou. Statistical iterative reconstruction using adaptive fractional order regularization. Biomedical Optics Express, 7(3):1015–1029, 2016. [13] Y. Zhang, W. H. Zhang, H. Chen, M. L. Yang, T. Y. Li, and J. L. Zhou. Few-view image reconstruction combining total variation and a high-order norm. International Journal of Imaging Systems and Technology, 23(3):249–255, 2013. [14] P. Kaur, G. Singh, and P. Kaur. A review of denoising medical images using machine learning approaches. Current Medical Imaging Reviews, 14(5):675–685, 2018. [15] H. Chen, Y. Zhang, W. Zhang, P. Liao, K. Li, J. Zhou, and G. Wang. Low-dose CT via convolutional neural network. Biomedical Optics Express, 8(2):679–694, 2017. [16] C. Tan, M. Yang, Z. You, H. Chen, and Y. Zhang. A selective kernel-based cycle-consistent generative adversarial network for unpaired low-dose CT denoising. Precision Clinical Medicine, 5(2):pbac011, 2022. 82
[17] H. Chen, Y. Zhang, M. K. Kalra, F. Lin, Y. Chen, P. Liao, et al. Low-dose CT with a residual encoder-decoder convolutional neural network. IEEE Transactions on Medical Imaging, 36(12):2524–2535, 2017. [18] K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. [19] C. You, Q. Yang, H. Shan, L. Gjesteby, G. Li, S. Ju, et al. Structurally-sensitive multi-scale deep neural network for low-dose CT denoising. IEEE Access, 6:41839–41855, 2018. [20] Q. Yang, P. Yan, Y. Zhang, H. Yu, Y. Shi, X. Mou, et al. Low-dose CT image denoising using a generative adversarial network with Wasserstein distance and perceptual loss. IEEE Transactions on Medical Imaging, 37(6):1348–1357, 2018. [21] G. Wang and X. Hu. Low-dose CT denoising using a progressive Wasserstein generative adversarial network. Computers in Biology and Medicine, 135:104625, 2021. [22] F. Fan, H. Shan, M. K. Kalra, R. Singh, G. Qian, M. Getzin, et al. Quadratic autoencoder (Q-AE) for low-dose CT denoising. IEEE Transactions on Medical Imaging, 39(6):2035–2050, 2019. [23] C. Corti, M. Cobanaj, E. C. Dee, C. Criscitiello, S. M. Tolaney, L. A. Celi, and G. Curigliano. Artificial intelligence in cancer research and precision medicine: Applications, limitations and priorities to drive transformation in the delivery of equitable and unbiased care. Cancer Treatment Reviews, 112:102498, 2023. [24] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. [25] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, et al. Attention is all you need. Advances in Neural Information Processing Systems, 30, 2017. [26] D. Wang, F. Fan, Z. Wu, R. Liu, F. Wang, and H. Yu. CTformer: convolution-free Token2Token dilated vision transformer for low-dose CT denoising. Physics in Medicine and Biology, 68(6):065012, 2023. [27] A. Luthra, H. Sulakhe, T. Mittal, A. Iyer, and S. Yadav. Eformer: Edge enhancement based transformer for medical image denoising. arXiv preprint arXiv:2109.08044, 2021. [28] M. Jian, X. Yu, H. Zhang, and C. Yang. SwinCT: feature enhancement based low-dose CT images denoising with swin transformer. Multimedia Systems, 30(1):1, 2024. [29] L. Jia, X. He, A. Huang, B. Jia, and X. Wang. Highly efficient encoder-decoder network based on multi-scale edge enhancement and dilated convolution for LDCT image denoising. Signal, Image and Video Processing, 18(8):6081-6091, 2024. [30] H. Yan, C. Fang, and Z. Qiao. A multi-attention Uformer for low-dose CT image denoising. Signal, Image and Video Processing, 18(2):1429-1442, 2024. [31] R. S. Jebur, M. H. B. M. Zabil, D. A. Hammood, and L. K. Cheng. A comprehensive review of image denoising in deep learning. Multimedia Tools and Applications, 83(20):58181-58199, 2024. [32] J. Ruan and S. Xiang. Vm-unet: Vision mamba unet for medical image segmentation. arXiv preprint arXiv:2402.02491, 2024. [33] S¸. ¨ Ozt¨urk, O. C. Duran, and T. C¸ukur. DenoMamba: A fused state-space model for low-dose CT denoising. arXiv preprint arXiv:2409.13094, 2024. [34] Y. Li, M. Yang, T. Bian, and H. Wu. Enhancing low-dose CT images by 4x using CACTSR: a deep learning model. In International Conference on Cloud Computing, Performance Computing, and Deep Learning (CCPCDL 2024), volume 13281, pages 295-301. SPIE, September 2024. [35] O. Ronneberger, P. Fischer, and T. Brox. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III, volume 18, pages 234-241. Springer, 2015. 83