Full text
Large-Scale Empirical Image Enhancement Studies with Diverse HPC systems Lujun Zhai [email protected] , Suxia Cui [email protected] Prairie View A&M University, Prairie View, TX, USA Abstract—High-performance computing (HPC) is vital for advancing AI research in computer vision, where training on high-resolution datasets requires significant computational power. Using the NSF-funded Accelerating Computing for Emerging Sciences (ACES) testbed at Texas A&M University, we leveraged Graphcore IPUs, NVIDIA H100 GPUs, and A30 GPUs to conduct a large-scale empirical study of 170 model configurations spanning CNN-, GAN-, Transformer-, and Diffusion-based architectures. This enabled us to address an open question: Is underwater image enhancement (UIE) truly beneficial for underwater object detection? We trained five object detectors across 17 enhancement domains and two datasets. Our results show that most UIE methods degrade detection accuracy, while select diffusion-based approaches that preserve key features can mitigate this drop. HPC resources also allowed us to compare GPU and IPU performance. These findings guide the practical use of UIE in marine vision and highlight the importance of equitable HPC access for large-scale AI research. I. INTRODUCTION High-performance computing (HPC) has become an indispensable enabler for modern artificial intelligence (AI) research, particularly in domains such as computer vision, where the training and evaluation of models on large-scale, highresolution datasets require substantial computational power. At one of minority-serving institutions HBCUS, however, access to such large-scale computing resources has traditionally been limited, creating barriers for conducting computationally intensive research. To address these challenges, the National Science Foundation (NSF) has invested in initiatives such as the Accelerating Computing for Emerging Sciences (ACES) testbed, hosted by Texas A&M University’s High Performance Research Computing (HPRC) facility. ACES provides cuttingedge hardware, including NVIDIA H100 and A30 GPUs as well as Graphcore Intelligence Processing Units (IPUs), enabling researchers from a broad range of institutions to perform large-scale AI experimentation [1]. For Prairie View A&M University (PVAMU), an HBCU [2], this access has been transformative—enabling projects that would otherwise be infeasible within on-campus computing constraints. In this work, we leverage ACES to conduct a large-scale empirical study on the impact of underwater image enhancement (UIE) techniques on downstream object detection. While UIE is widely used to improve visual quality in marine imagery, its actual effect on object detection accuracy remains an open question. Addressing this gap requires systematic evaluation across diverse enhancement algorithms, object detection architectures, and datasets—a computationally intensive task well beyond the capacity of conventional workstations. By utilizing ACES’ heterogeneous HPC infrastructure, we evaluate 170 unique model configurations, spanning convolutional neural networks (CNNs), generative adversarial networks (GANs), transformer-based methods, and diffusionbased models. This work not only provides new insights into the role of UIE in underwater object detection but also serves as a case study on how equitable access to HPC resources can enable impactful AI research at institutions historically underrepresented in large-scale computing. II. METHODOLOGY A. Experimental Design We conducted a large-scale empirical study to evaluate the impact of UIE on object detection. The study covered: •Enhancement Methods: 16 state-of-the-art UIE approaches, including CNN-based [3]–[6], GANbased [7]–[12], Transformer-based [13]–[15], and Diffusion-based [16]–[18] models. •Object Detectors: Five representative one-stage and twostage architectures [19]–[23], chosen for strong generalpurpose performance. •Datasets: Two public underwater object detection datasets—RUOD [24] and URPC2020 [25]. Each dataset contained a raw image domain plus 16 enhanced variants, producing 17 domains per dataset. Every domain was used to train all five detectors, yielding 170 unique training configurations (5 detectors ×17 domains × 2 datasets). All models were trained from scratch to ensure consistent evaluation. B. HPC Infrastructure All experiments were performed on the NSF-funded ACES testbed at Texas A&M University’s HPRC facility, including: •Graphcore IPUs: Optimized for fine-grained parallelism in AI workloads [26], enabling efficient training of transformerand diffusion-based models. •NVIDIA H100 and A30 GPUs: High-memory, highthroughput accelerators suited for large-scale convolutional and generative models. C. Evaluation Metrics Detection performance was evaluated using mean Average Precision (mAP), while image enhancement quality was measured with UIQM [27]. Comparative analyses of raw versus
enhanced images quantified the impact of each UIE method on downstream object detection accuracy. We measured HPC performance by recording training throughput (images processed per second) across systems. All tests used a fixed batch size of 16, identical hyperparameters, and equal compute resources. III. RESULTS AND DISCUSSION We trained 170 models by applying 16 UIE methods to two datasets and evaluating them with five object detectors. Table 1 presents representative results from each enhancement category (CNN, GAN, Transformer, and Diffusion) to illustrate key trends in image quality (UIQM) and detection performance (mAP) for raw and enhanced images. Contrary to the common assumption that enhancement universally improves detection, our results show that: •Most enhancement methods degrade accuracy. On both RUOD and URPC2020 datasets, most of UIE reduced detect acc.(mAP) compared to the raw baseline. •Diffusion-based approaches are more robust. A subset of well-designed diffusion models preserved or slightly improved accuracy by enhancing contrast without distorting low-level features critical for object localization. These results suggest that UIE should be applied selectively in downstream vision tasks. TABLE I COMPARISON OF ENHANCED IMAGE QUALITY (UIQM) AND DETECTION ACCURACY (MAP) WITH YOLO-NAS [23] FOR DIFFERENT UIE METHODS. BEST RESULTS ARE IN BLUE,SECOND-BEST IN PINK. SHOWN ARE REPRESENTATIVE SAMPLES FROM 170 MODELS. UIEBD Dataset UIEBD Dataset METHODS Type UIQM ↑mAP50:95 ↑UIQM ↑mAP50:95 ↑ Raw - 1.59 63.46 1.87 49.62 UWCNN [28] CNN 3.34 58.18 2.78 47.14 UWGAN [9] GAN 1.94 58.42 2.30 45.47 Spectroformer [14] Trans. 2.29 61.41 2.68 49.13 WF-Diff [16] Diffusion 3.80 62.30 4.04 49.57 A. Role of HPC in Enabling Large-Scale Analysis The exhaustive evaluation of 170 training configurations required thousands of GPU and IPU compute-hours. The ACES testbed’s heterogeneous architecture enabled: •Parallel execution: Multiple training jobs ran concurrently across different accelerators, reducing total wallclock time. •Architecture-specific optimization: Large CNN models such as WaterNet achieved higher throughput on H100 and A30 GPUs, while Transformer-based models such as Spectroformer benefited significantly from IPUs due to their fine-grained parallelism and attention acceleration. •Memory-bound workload efficiency: IPUs executed many small iterative steps efficiently, leveraging their high on-chip memory bandwidth to reduce data movement and latency, especially for attention-heavy layers. These architectural advantages are reflected in Table II and Fig. 1, where CNNs show peak throughput on GPUs, while Transformers gain a relative advantage on IPUs due to their parallelism and memory architecture. TABLE II THROUGHPUT COMPARISIONS (IMAGES/SEC AT BATCH = 16). Model Type HPC Type Throughput UWCNN CNN GPU (A30) 6.30 UWCNN CNN GPU (H100) 74.94 UWCNN CNN IPU 71.26 Spectroformer Transformer GPU (A30) 2.72 Spectroformer Transformer GPU (H100) 29.64 Spectroformer Transformer IPU 35.91 Spectroformer (Transformer) WaterNet (CNN) Fig. 1. Throughput comparison across different HPC types, highlighting that CNN-based WaterNet has an advantage on GPUs, while the Transformer-based Spectroformer shows an advantage on IPUs. IV. CONCLUSION AND FUTURE WORK This study provides large-scale, systematic evaluation of UIE methods for object detection, conducted using advanced HPC infrastructure. Our results challenge the common belief that image enhancement always improves detection accuracy, showing that many enhancement techniques actually degrade performance. By leveraging the heterogeneous resources of the NSFfunded ACES testbed at Texas A&M University’s High Performance Research Computing (HPRC) facility, we executed 170 unique training configurations that would have been infeasible to run at a minority-serving institution without equitable HPC access. This work highlights how shared HPC infrastructure can bridge resource gaps, enabling rigorous, competitive AI research across a diverse range of institutions. In the future, we will explore how IPUs and GPUs perform differently on diffusion-based and GAN-based models, further advancing the effective use of HPC resources for the global research community. ACKNOWLEDGMENT This work utilized the High Performance Research Computing (HPRC) facility at Texas A&M University and NSFfunded ACCESS resources. The authors gratefully acknowledge the support of the HPRC team at Texas A&M University, as well as additional support from the NSF-funded ACCESS program. REFERENCES [1] “Nsf aces,” Texas A&M University, High Performance Research Computing (HPRC), 2023, web page. [Online]. Available: https: //hprc.tamu.edu/aces/ [2] “Full list of hbcus,” https://www.thehundred-seven.org/hbculist.html, accessed 2025, list sourced from The Hundred-Seven.
[3] C. Li, S. Anwar, and F. Porikli, “Underwater scene prior inspired deep underwater image and video enhancement,” Pattern Recognition, vol. 98, p. 107038, 2020. [4] C. Li, C. Guo, W. Ren, R. Cong, J. Hou, S. Kwong, and D. Tao, “An underwater image enhancement benchmark dataset and beyond,” IEEE Transactions on Image Processing, vol. 29, pp. 4376–4389, 2020. [5] Y. Wang, J. Guo, H. Gao, and H. Yue, “Uiec2-net: Cnn-based underwater image enhancement using two color spaces,” Signal Processing: Image Communication, vol. 96, p. 116250, 2021. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0923596521001004 [6] F. Huo, B. Li, and X. Zhu, “Efficient wavelet boost learning-based multi-stage progressive refinement network for underwater image enhancement,” in 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2021, pp. 1944–1952. [7] J. Li, K. A. Skinner, R. Eustice, and M. Johnson-Roberson, “Watergan: Unsupervised generative network to enable real-time color correction of monocular underwater images,” IEEE Robotics and Automation Letters (RA-L), 2017, accepted. [8] C. Fabbri, M. J. Islam, and J. Sattar, “Enhancing underwater imagery using generative adversarial networks,” in 2018 IEEE International Conference on Robotics and Automation (ICRA), 2018, pp. 7159–7165. [9] N. Wang, Y. Zhou, F. Han, H. Zhu, and Y. Zheng, “Uwgan: Underwater gan for real-world underwater color restoration and dehazing,” 2019. [10] C. Desai, B. S. S. Reddy, R. A. Tabib, U. Patil, and U. Mudenagudi, “Aquagan: Restoration of underwater images,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2022, pp. 295–303. [11] Z. Wang, L. Shen, M. Xu, M. Yu, K. Wang, and Y. Lin, “Domain adaptation for underwater image enhancement,” IEEE Transactions on Image Processing, vol. 32, pp. 1442–1457, 2023. [12] Z. Jiang, Z. Li, S. Yang, X. Fan, and R. Liu, “Target oriented perceptual adversarial fusion network for underwater image enhancement,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 10, pp. 6584–6598, 2022. [13] Y. Tang, T. Iwaguchi, H. Kawasaki, R. Sagawa, and R. Furukawa, “Autoenhancer: Transformer on u-net architecture search for underwater image enhancement,” in Proceedings of the Asian Conference on Computer Vision (ACCV), December 2022, pp. 1403–1420. [14] M. R. Khan, P. Mishra, N. Mehta, S. S. Phutke, S. K. Vipparthi, S. Nandi, and S. Murala, “Spectroformer: Multi-domain query cascaded transformer network for underwater image enhancement,” in 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2024, pp. 1443–1452. [15] B. Wang, H. Xu, G. Jiang, M. Yu, T. Ren, T. Luo, and Z. Zhu, “Uieconvformer: Underwater image enhancement based on convolution and feature fusion transformer,” IEEE Transactions on Emerging Topics in Computational Intelligence, vol. 8, no. 2, pp. 1952–1968, 2024. [16] C. Zhao, W. Cai, C. Dong, and C. Hu, “Wavelet-based fourier information interaction with frequency diffusion adjustment for underwater image restoration,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 8281–8291. [17] D. Du, E. Li, L. Si, W. Zhai, F. Xu, J. Niu, and F. Sun, “Uiedp: Boosting underwater image enhancement with diffusion prior,” Expert Systems with Applications, vol. 259, p. 125271, 2025. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0957417424021389 [18] Y. Tang, H. Kawasaki, and T. Iwaguchi, “Underwater image enhancement by transformer-based diffusion model with non-uniform sampling for skip strategy,” in Proceedings of the 31st ACM International Conference on Multimedia, ser. MM ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 5419–5427. [Online]. Available: https://doi.org/10.1145/3581783.3612378 [19] S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster r-cnn: Towards realtime object detection with region proposal networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 39, no. 6, pp. 1137–1149, 2017. [20] Z. Cai and N. Vasconcelos, “Cascade r-cnn: Delving into high quality object detection,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 43, no. 5, pp. 1483–1498, 2021. [21] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Doll´ ar, “Focal loss for dense object detection,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), vol. 42, no. 2, pp. 318–327, 2020. [22] W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” European Conference on Computer Vision (ECCV), pp. 21–37, 2016. [23] D. AI, “Yolo-nas: Faster and more accurate object detection with nas,” Deci AI Research, 2023. [Online]. Available: https: //github.com/Deci-AI/YOLO-NAS [24] “Rethinking general underwater object detection: Datasets, challenges, and solutions,” Neurocomputing, vol. 517, pp. 243–256, 2023. [25] C. Liu, H. Li, S. Wang, M. Zhu, D. Wang, X. Fan, and Z. Wang, “A dataset and benchmark of underwater object detection for robot picking,” in 2021 IEEE international conference on multimedia & expo workshops (ICMEW). IEEE, 2021, pp. 1–6. [26] A. Nasari, L. Zhai, Z. He, H. Le, S. Cui, D. Chakravorty, J. Tao, and H. Liu, “Porting ai/ml models to intelligence processing units (ipus),” in Practice and Experience in Advanced Research Computing 2023: Computing for the Common Good, ser. PEARC ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 231–236. [Online]. Available: https://doi.org/10.1145/3569951.3603632 [27] M. Yang and A. Sowmya, “An underwater color image quality evaluation metric,” IEEE Transactions on Image Processing, vol. 24, no. 12, pp. 6062–6071, 2015. [28] C. Li, S. Anwar, and F. Porikli, “Underwater scene prior inspired deep underwater image and video enhancement,” Pattern Recognition, vol. 98, p. 107038, 2020.