scieee AI-readable full text Open interactive document viewer

Proceedings of the Second Conference on Applications of Artificial Intelligence (A2I'25)

Touazi, Fayçal; Belkasmi, Djamal; Benzenati, Tayeb; YAHIATENE, Youcef; BOULIF, Menouar; DAOUI, Abdelhakim

Abstract

This volume contains the collection of papers presented at the International Conference on Applications of Artificial Intelligence (A2I’25), held at the University M’hamed Bougara of Boumerdes, Algeria, on April 16–17, 2025. The conference aimed to provide an international forum for researchers, doctoral students, and practitioners to exchange innovative ideas, present original findings, and discuss recent advances in the field of Artificial Intelligence (AI) and its wide-ranging applications. Artificial Intelligence continues to transform the way societies address challenges in domains such as healthcare, energy, agriculture, urban development, finance, and information security. By bringing together interdisciplinary perspectives, A2I’25 offered an opportunity to highlight not only theoretical advancements but also practical solutions that can directly impact societal well-being and sustainable development. For this edition, the organizing committee received 51 paper submissions. Each submission underwent a rigorous peer-review process, with every article evaluated by at least two qualified reviewers from the program committee. Following this process, 28 papers were accepted, leading to an acceptance rate of approximately 54%. The selected contributions cover a wide spectrum of AI-related topics, including machine learning, computer vision, natural language processing, intelligent systems, optimization, and bio-inspired algorithms. The accepted papers were presented in multiple oral sessions over the two-day program, reflecting the diversity and richness of ongoing research in the field. In addition to the contributed papers, the conference featured keynote addresses from distinguished speakers, who provided valuable insights into current trends and future directions of Artificial Intelligence research and its societal applications. The scientific discussions and interactive exchanges during the event highlighted both opportunities and challenges, laying the groundwork for future collaborations and advancements in AI. We are confident that the articles included in this volume will serve as a useful reference for researchers, engineers, and students working in the field. They not only represent the state of the art in Artificial Intelligence but also open new perspectives on how AI can be harnessed to address complex real-world problems. Finally, we extend our sincere gratitude to all authors for their valuable contributions, to the reviewers for their careful and constructive evaluations, and to the keynote speakers for their inspiring talks. We also wish to thank the members of the organizing and program committees for their dedication in ensuring the scientific quality and success of A2I’25.

Full text

Applications of Artificial Intelligence (AAI’25) Proceedings of The Second Conference On: Dr. Touazi Fayçal, Dr. Belkasmi Djamel, Dr. Benzenati Tayeb Dr. Yahiatene Youcef Pr. Boulif Menaour Pr. DAOUI Abdelhakim. Editor: i Dr. Touazi Fay¸cal, Dr. Belkasmi Djamel, Dr. Benzenati Tayeb, Dr. Yahiatene Youcef, Pr. Boulif Menouar and Pr. Daoui Abdelhakim Editors Proceedings of the Second Conference on Applications of Artificial Intelligence (A2I’25) 16–17 April 2025 Computer Science Department, Faculty of Science University M’hamed Bougara Boumerdes, Algeria ISBN: 978-9969-9650-0-1 D´epˆot l´egale: 9969-2025 ©Copyright University M’hamed Bougara Boumerdes, Algeria. All Rights Reserved. Contents Preface iii Organizing Committee iv Scientific Committee v Keynote Speakers vii I Artificial Intelligence for Medical Imaging and Data Security 2 Transfer Learning with Enhanced Models for Skin Cancer Detection: A Comprehensive Evaluation of Transfer Learning and Data Augmentation on the ISIC 2020 Dataset 2–11 AI for Medical Image Security: A Comprehensive Review of Techniques and Challenges 12–19 A novel cloud-deployed data pipeline for cervical spine fracture detection in 3D CT images 20–28 Advanced Ensemble Learning Framework for Reliable Smart Grid Stability detection 29–37 Unmasking Deepfakes: CNNs and Vision Transformers for Cutting-Edge Detection 38–47 Review on deep learning optimization using knowledge and dataset distillation in medical imaging diagnostics 48–57 II Deep Learning and Data Processing Applications 58 RL-Guided Pruning of CNNs Using Graph Embeddings 59–68 Image fusion using a new evolution equation 69–74 MambaCT: Feature Enhancement-Based Low-Dose CT Image Denoising Using Vision Mamba and a Scaling Adapter 75–84 Transfer Learning for Multi-Script Identification: A Comparative Study 85–90 Parameter-Efficient Fine-Tuning for LLM-Based Arabic-to-English Machine Translation 91–112 A Survey on Approaches to Modeling Collaborative Practices in E-Learning Platforms 113–121 IoT Applications in the Education Sector: Architectures, Challenges, and Emerging Paradigms 122–128 A Comprehensive Review of Knowledge Graph Integration in Large Language Models for Trust 129–138 A Hybrid Architecture for Tomato Leaf Disease Classification Through State Space and Convolutional Feature Fusion 139–147 Enhanced Two-Stage PCANet for Biometric Recognition via Discriminative Block Reweighting and Overlapped Histogram Encoding 148–153 Real-Time License Plate Recognition using YOLOv9 and Embedded Systems 154–163 Development of Real-Time Embedded Application for Drone System 164–172 III Advanced AI Approaches for Optimization and Data Analysis 173 A new encoding for generating highly nonlinear eight-variables Boolean functions using multi-parent genetic algorithms 174–182 Enhancing Learning Management Systems with AI: Recommendations from Moodle Usage 183–191 Improved Symbiotic Organism Search Algorithm for Biomedical Data Clustering 192–197 Deep Learning Approaches for Energy Optimization in CPS: A survey 198–207 Formalization of the AGR model using the DD-LOTOS Formal Language 208–219 HMM-Based Multi-Heartbeat Phonocardiogram Classification Using Wavelet Cepstral Coefficients 220–226 Genetic Algorithm Learning Operators to Solve the Vehicle Routing Problem 227–239 From Ants to People: the Vaporization of Social Relationships in Dynamic Community Detection 240–249 ii Preface This volume contains the collection of papers presented at the International Conference on Applications of Artificial Intelligence (A2I’25), held at the University M’hamed Bougara of Boumerdes, Algeria, on April 16–17, 2025. The conference aimed to provide an international forum for researchers, doctoral students, and practitioners to exchange innovative ideas, present original findings, and discuss recent advances in the field of Artificial Intelligence (AI) and its wide-ranging applications. Artificial Intelligence continues to transform the way societies address challenges in areas such as healthcare, energy, agriculture, urban development, finance, and information security. By bringing together interdisciplinary perspectives, A2I’25 offered an opportunity to highlight not only theoretical advancements but also practical solutions that can directly impact societal well-being and sustainable development. For this edition, the organizing committee received 51 paper submissions. Each submission underwent a rigorous peer-review process, with every article evaluated by at least two qualified reviewers from the program committee. Following this process, 28 papers were accepted, leading to an acceptance rate of approximately 54%. The selected contributions cover a wide spectrum of AI-related topics, including machine learning, computer vision, natural language processing, intelligent systems, optimization, and bio-inspired algorithms. The accepted papers were presented in multiple oral sessions over the two-day program, reflecting the diversity and richness of ongoing research in the field. In addition to the contributed papers, the conference featured keynote addresses from distinguished speakers, who provided valuable insights into current trends and future directions of Artificial Intelligence research and its societal applications. The scientific discussions and interactive exchanges during the event highlighted both opportunities and challenges, laying the groundwork for future collaborations and advancements in AI. We are confident that the articles included in this volume will serve as a useful reference for researchers, engineers, and students working in the field. They not only represent the state-ofthe-art in Artificial Intelligence but also open new perspectives on how AI can be harnessed to address complex real-world problems. Finally, we extend our sincere gratitude to all authors for their valuable contributions, to the reviewers for their careful and constructive evaluations, and to the keynote speakers for their inspiring talks. We also wish to thank the members of the organizing and program committees for their dedication in ensuring the scientific quality and success of A2I’25. Organizing Committee •Dr. BELKASMI Djamel (UMBB, Algeria) Chair •Dr. TOUAZI Fay¸cal (UMBB, Algeria) •Dr. BENZENNATI Tayeb (UMBB, Algeria) •Pr. GACEB Djamel (UMBB, Algeria) •Dr. BOUSTIL Amel (UMBB, Algeria) •Pr. MERAIHI Yacine (FT, UMBB, Algeria) •Dr. IMACHE Rabah (UMBB, Algeria) •Dr. YAHIATENE Youcef (UMBB, Algeria) •Dr. LOUNAS Razika (UMBB, Algeria) •Dr. DJOUZI Kheyreddine (UMBB, Algeria) •Dr. MOKRANI Hocine (UMBB, Algeria) •Dr. MESBAH Abdelhak (UMBB, Algeria) •Dr. BEDDARI Ibtihal (UMBB, Algeria) •Dr. CHAOUCHE Ali (UMBB, Algeria) •Dr. REZOUG Abdellah (UMBB, Algeria) •Dr. ISHAK Boushaki (UMBB, Algeria) •Dr. HAMADOUCHE Samiya (UMBB, Algeria) •Dr. DJERBI Rachid (UMBB, Algeria) •Dr. BENNAI M. Tahar (UMBB, Algeria) •Dr. KHOUDI Asmaa (UMBB, Algeria) •Dr. RAHMOUNE Nabila (UMBB, Algeria) Scientific Committee •Dr. BENZENATI Tayeb (UMBB, Algeria) Chair •Pr. DAAMOUCHE Abdelhamid (IGEE, Boumerdes, Algeria) •Pr. BOULIF Menouar (UMBB, Algeria) •Pr. BERRICHI Ali (UMBB, Algeria) •Pr. GACEB Djamel (UMBB, Algeria) •Pr. MAOUCHE Amine Riad (UMBB, Algeria) •Pr. RIAHLA Mohamed Amine (UMBB, Algeria) •Pr. Benblidia Nadjia (USDB, Blida, Algeria) •Pr. Mohammed Hachama (NHSM, Sidi Abdellah, Algeria) •Pr. BELHADEF Hacene (UFMC, Constantine, Algeria) •Dr. IMACHE Rabah (UMBB, Algeria) •Dr. LOUNAS Razika (UMBB, Algeria) •Dr. TOUAZI Fay¸cal (UMBB, Algeria) •Dr. YAHIATENE Youcef (UMBB, Algeria) •Dr. MOKRANI Hocine (UMBB, Algeria) •Dr. HAMADOUCHE Samiya (UMBB, Algeria) •Dr. ALOUANE Basma (UMBB, Algeria) •Dr. REZOUG Abdellah (UMBB, Algeria) •Dr. CHAOUCHE Ali (UMBB, Algeria) •Dr. ISHAK Boushaki Saida (UMBB, Algeria) •Dr. HADJIDJ Drifa (UMBB, Algeria) •Dr. MESBAH Abdelhak (UMBB, Algeria) •Dr. BADDARI Ibtihal (UMBB, Algeria) •Dr. OUKAS Nourredine (UMAB, Algeria) •Pr. A¨ıtza¨ı Abdelhakim (USTHB, Algeria) •Dr. CHOUIREF Zahira (Bouira University, Algeria) •Dr. DJERBI Rachid (UMBB, Algeria) •Dr. BENNAI M. Tahar (UMBB, Algeria) •Dr. RAHMOUNE Nabila (UMBB, Algeria) •Dr. RAHMOUNE Adel (UMBB, Algeria) •Dr. KHOUDI Asmaa (UMBB, Algeria) •Dr. DJOUZI Kheyreddine (UMBB, Algeria) •Dr. SAOULI Abdelhak (UMBB, Algeria) •Pr. CHERIFI Dalila (IGEE, Boumerdes, Algeria) •Pr. CHALLAL Mouloud (IGEE, Boumerdes, Algeria) •Dr. TABET Youcef (IGEE, Boumerdes, Algeria) •Dr. LOUBAR Hocine (IGEE, Boumerdes, Algeria) •Dr. BOUSTIL Amel (UMBB, Algeria) •Pr. MERAIHI Yacine (FT, UMBB, Algeria) •Dr. BAICHE Karim (FT, UMBB, Algeria) •Dr. AKROUM Hamza (FT, UMBB, Algeria) •Dr. FERRAHI Ibtissam (USTHB, Algeria) vi Keynote Speakers Dr. Ameni MKAOUAR Researcher at NASA’s Goddard Space Flight Center Plenary title: Advancing Digital Surface Model Derivation in Forested Environments Through the Simulation and Fusion of Satellite Stereophotogrammetry and LiDAR Data Pr. Abdelmalik TALEB-AHMED Professor in Image and Signal Processing at LAMIH, University of Valenciennes Plenary title: Contenus G´en´er´es par l’IA : Opportunit´es et/ou D´efis - Pourquoi et comment ? Figure 3: Confusion matrices with data augmentation. 5 Discussion Enhanced DenseNet121’s standout performance (96.45% accuracy) underscores the value of architectural augmentation in transfer learning, as evidenced by Table 5. Baseline accuracy (94%) dropped to 82% with augmentation, rebounding to 96.45% with enhancements—a pattern mirrored across models (e.g., VGG16: 85% to 79% to 93.49%). Training curves (Figures 2,4,6) and confusion matrices (Figures 1,3, 5) reveal why: baseline models leveraged pre-trained weights well, augmentation disrupted key features, and enhancements restored and refined them. The consistent drop in accuracy across all models with data augmentation—from 94% to 82% for DenseNet121—likely stems from the disruption of critical dermoscopic features like lesion asymmetry and border irregularity, essential for malignancy detection in the ISIC 2020 dataset. Geometric transformations such as 90◦rotation and horizontal flipping, applied to the 10,653 training images, may have altered these diagnostic markers, misaligning them with the dataset’s centered lesion patterns and causing a surge in false negatives (e.g., 1,163/3,502 for DenseNet121). Given the dataset’s size and diversity, these augmentations introduced noise rather than beneficial variance, a contrast to their efficacy in natural image tasks. Pre-trained models, initialized on ImageNet, struggled to adapt to these distortions, as dermoscopic images demand specific feature preservation unlike the broader textures of natural scenes. Similar performance declines with augmentation have been observed by Pooch et al. [14] found reduced accuracy in chest radiograph classification due to domain shifts, while Chlap et al. [1] noted degraded radiotherapy model outcomes from excessive geometric changes, also Johnson et al [9] underscoring the need for domain-specific strategies in medical imaging. Table 5: Comparative accuracy across configurations. Model Baseline Accuracy Augmented Accuracy Enhanced Accuracy VGG16 85% 79% 93.49% ResNet50 76% 69% 94.07% DenseNet121 94% 82% 96.45% InceptionV3 88% 77% 93.68% DenseNet121’s edge lies in its dense connectivity [6], where each layer accesses all prior outputs, fostering feature reuse (e.g., edges from early layers inform deeper lesion pattern detection). This contrasts with VGG16’s linear depth, which redundantly relearns features, or ResNet50’s residuals, which mitigate gradients but lack DenseNet’s efficiency (fewer parameters: ∼7M vs. ResNet50’s ∼25M). Added convolutional and dense layers with ReLU and dropout further tuned this advantage, adapting ImageNet-derived filters to dermoscopic specifics—likely prioritizing irregular borders or pigment variations over generic 7 Figure 4: Training and validation curves with data augmentation. textures. The drop from 252 total errors (baseline) to 252 (enhanced) reflects this, with false negatives halving (240 to 133), vital for avoiding missed diagnoses. Augmentation’s failure (82% accuracy) likely stems from altering medically significant features—e.g., flipping a lesion might obscure asymmetry, a malignancy marker [9]. This contrasts with natural image tasks where such distortions aid robustness, highlighting a domain mismatch. Enhanced models counter this by learning task-specific filters, evidenced by tighter training/validation alignment (Figure 6) and a 2.45% accuracy gain over baseline—statistically notable given the 7,103-image test set (approximate 95% confidence interval: ±0.8%). Compared to prior work, 96.45% approaches Rashid et al.’s 98.2% with MobileNetV2 [16] and Khan et al.’s 98.2% with optimized CNNs [10], surpassing Esteva et al.’s 91% [3]. Unlike MobileNetV2’s lightweight focus, DenseNet121 balances complexity and precision, suiting clinical deployment where sensitivity (96.69%) outweighs speed. Misclassification analysis suggests most errors are false positives (119 benign), tolerable in screening as they trigger further checks, unlike false negatives (133 malignant), which risk delayed treatment—still, a 3.80% miss rate rivals expert dermatologists (e.g., Haenssle et al.’s 95% sensitivity [4]). Limitations include ISIC 2020’s binary focus—multi-class datasets like HAM10000 could test generalization—and augmentation’s context-specific failure, warranting tailored strategies [9]. 8 Figure 5: Confusion matrices for enhanced models. 6 Conclusion This study demonstrates the efficacy of transfer learning with enhanced convolutional neural networks for skin cancer detection, achieving 96.45% accuracy, 96.32% precision, 96.69% recall, and a 96.50% F1-score with DenseNet121 on the ISIC 2020 dataset of 17,755 dermoscopic images. Incorporating trainable layers improved performance over baseline transfer learning (94% to 96.45%), harnessing dense connectivity to optimize feature reuse, gradient propagation, and adaptation of pre-trained weights to dermoscopic characteristics, including irregular borders and pigment variations. This enhancement yielded a false negative rate of 3.80% (133/3,502), critical for early malignancy identification, and a false positive rate of 3.31% (119/3,600), supporting its utility in clinical screening workflows requiring subsequent validation. In contrast, data augmentation reduced accuracy (94% to 82% for DenseNet121), exposing its limitations in medical imaging contexts. Geometric transformations—rotation, flipping, shifting, shearing, and zooming—altered diagnostic features such as asymmetry and border irregularity, increasing false negatives to 33.21% (1,163/3,502) and disrupting training stability. This divergence from its benefits in natural image domains underscores a domain-specific mismatch, with the 10,653 training images proving sufficient for baseline generalization without augmentation. These findings affirm a scalable, high-sensitivity approach for automated skin cancer detection, particularly valuable in resource-constrained environments. Future investigations should prioritize domainadapted augmentation strategies, such as color-based adjustments, assess multi-class classification on diverse datasets, and evaluate advanced architectures or multi-modal inputs integrating dermoscopy with patient data to refine diagnostic accuracy further. This work advances the integration of AI into clinical dermatology by optimizing model architecture while highlighting the need for tailored data preprocessing. References [1] Phillip Chlap, Hang Min, Nicholas Vandenberg, Jason Dowling, Lois Holloway, and Annette Haworth. A review of medical image data augmentation techniques for deep learning applications. Journal of Medical Imaging and Radiation Oncology, 65(5):545–563, 2021. [2] N. C. Codella et al. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by isic. arXiv preprint arXiv:1902.03368, 2019. 9 Figure 6: Training and validation curves for enhanced models. [3] A. Esteva et al. Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542(7639):115–118, 2017. [4] H. A. Haenssle et al. Man against machine: Diagnostic performance of a deep learning convolutional neural network for dermoscopic melanoma recognition. Annals of Oncology, 29(8):1836–1842, 2018. [5] K. He et al. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016. [6] G. Huang et al. Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4700–4708, 2017. [7] International Dermoscopy Society. International dermoscopy society consensus – terminology for dermoscopy. Journal of the American Academy of Dermatology, 85(3):678–687, 2021. 10 [8] K. Johnson et al. Domain-specific data augmentation for improving deep learning-based skin lesion classification. Biomedical Signal Processing and Control, 77:103789, 2022. [9] K. Johnson, L. Smith, and T. Brown. Domain-specific data augmentation for improving deep learning-based skin lesion classification. Biomedical Signal Processing and Control, 77:103789, 2022. [10] M. A. Khan et al. Enhanced deep learning-based skin cancer detection with transfer learning and optimization techniques. Artificial Intelligence in Medicine, 138:102489, 2023. [11] Y. LeCun et al. Deep learning. Nature, 521(7553):436–444, 2015. [12] A. Lomas et al. Global burden of melanoma: Epidemiology and trends. The Lancet Oncology, 23(11):1412–1421, 2022. [13] T. Mendon¸ca et al. Ph2 - a dermoscopic image database for research and benchmarking. 2013 35th Annual International Conference of the IEEE EMBC, pages 5437–5440, 2013. [14] Eduardo H. Pooch, Pedro L. Ballester, and Rodrigo C. Barros. Can we trust deep learning models diagnosis? the impact of domain shift in chest radiograph classification. In MICCAI Workshop on Thoracic Image Analysis. Springer, 2019. [15] M. Raghu et al. Transfusion: Understanding transfer learning for medical imaging. Advances in Neural Information Processing Systems, 32:3347–3357, 2019. [16] J. Rashid et al. Skin cancer disease detection using transfer learning technique. Applied Sciences, 12(11):5714, 2022. [17] O. Russakovsky et al. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115(3):211–252, 2015. [18] R. L. Siegel et al. Cancer statistics, 2023. CA: A Cancer Journal for Clinicians, 73(1):17–48, 2023. [19] K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. [20] C. Szegedy et al. Rethinking the inception architecture for computer vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2818–2826, 2016. [21] P. Tschandl et al. The ham10000 dataset, a large collection of multi-source dermoscopic images. Scientific Data, 5:180161, 2018. [22] World Health Organization. Skin cancers, 2023. 11 AI for Medical Image Security: A Comprehensive Review of Techniques and Challenges Benyoucef Aicha1and Hamadouche M’Hamed2 1Department of Electrical Systems Engineering, LIMOSE Laboratory, Faculty of Technology, [email protected] 2Department of Electrical Systems Engineering, LIMOSE Laboratory, Faculty of Technology, [email protected] Abstract Ensuring the security of sensitive medical data, including patient records and medical images, is paramount in the healthcare sector due to the risks of unauthorized access and data breaches. As healthcare information is increasingly transmitted through unsecured channels, maintaining its confidentiality, integrity, and authenticity is essential. This review examines AI-driven security techniques such as encryption, anomaly detection, and privacy-preserving algorithms, which play a crucial role in protecting medical data. By enhancing regulatory compliance and fostering trust in digital healthcare systems, these methods contribute significantly to data security. Additionally, this paper explores recent advancements in AI-based medical image protection and highlights key challenges and future research directions in the field of medical data security. Keywords: AI, Sybersecurity, medical image, cryptography, watermarking. 1 Introduction Given the critical sensitivity of healthcare data, it remains a primary target for cyberattacks, particularly as large volumes of information are accessed and transmitted over potentially unsecured networks. Ensuring data security requires robust protection measures at every stage, including storage, transmission, and retrieval. To safeguard patient information, researchers commonly employ techniques such as cryptography, steganography, and watermarking, which strengthen security and help prevent unauthorized access. [14,25] [7]. This review focuses on two primary types of health data—medical images and electronic health records (EHRs)—as these are commonly secured through encryption and watermarking. Medical images, derived from diagnostic tools like ultrasound and MRI, capture critical anatomical details and are essential for diagnosis and research, making their secure storage and transfer vital. EHRs, containing personal and medical information, are also crucial to protect, as they hold sensitive patient data. Securing medical images and EHRs is paramount to maintaining patient privacy, ensuring data integrity, and supporting trust in healthcare systems [10]. Security methods and techniques in the medical field help protect sensitive data, but with the rapid growth of Artificial intelligence (AI) applications in medical image segmentation and classification—particularly for enhancing diagnosis and cancer detection—AI has also become crucial for advancing medical data security [29]. AI is revolutionizing cybersecurity by enabling proactive threat detection and response through real-time data analysis and anomaly detection, enhancing systems like intrusion detection, malware analysis, and phishing detection while allowing security teams to focus on complex challenges [5]. The objective of this review is to examine recent AI-driven approaches to securing medical images, with a focus on key applications, methodologies, and emerging trends in the field. By analyzing current advancements, we aim to provide insights into how AI enhances medical image security and to identify areas for future research that could further strengthen privacy and data protection in healthcare. 2 Medical Image Security Threat Landscape The digital nature of the medical data and images exposes them to various cybersecurity threats. This section explores common security threats targeting medical images and their implications, highlighting the need for robust protective measures. 12 2.1 Types of Security Threats In July 2021, three organizations—Retinal Consultants Medical Group, ACE Surgical Supply, and Three Rivers Regional Commission—reported breaches in which unauthorized individuals accessed protected health information. These incidents affected a total of 25,725 patients, exposing personal data such as names, addresses, usernames, passwords, financial account numbers, and medical information, including treatment history and diagnoses. The compromised data posed significant risks, including identity theft, phishing attacks, and the potential alteration of medical records, which could lead to incorrect diagnoses and treatments [3]. 2.2 Vulnerabilities in healthcare This section explores the specific security challenges associated with the core components of e-health systems 2.2.1 Cloud computing platforms •Data Breaches: Data breaches in cloud services often occur due to poor security practices like weak passwords and the absence of multi-factor authentication, leading to the exposure of sensitive patient information [11]. •Unauthorized Access: Unauthorized access to cloud services often arises from misconfigurations and weak authentication protocols, which cybercriminals exploit through methods such as phishing [29], keylogging [30], person-in-the-middle (PITM) attacks, brute force attempts[33], and credential stuffing. These techniques enable attackers to steal or bypass login credentials, compromising sensitive data. 2.2.2 Internet of medical things (IOMT) The Internet of Medical Things (IoMT) enhances patient care through real-time data collection but poses significant security risks, including device vulnerabilities, data interception due to weak encryption, and susceptibility to remote attacks. These risks can compromise patient privacy, disrupt medical device functionality, and even endanger lives [35]. 2.2.3 Electronic health records (EHRS) Electronic Health Records (EHRs) are advanced digital systems that centralize and organize a wide array of patient information, including medical history, diagnoses, medications, immunization records, allergies, radiology images, and lab results. They provide real-time, patient-centered records that are instantly accessible to authorized personnel, anytime and anywhere. EHRs enhance collaboration by allowing multiple healthcare providers to share and access a patient’s information, enabling integrated care and better decision-making. They also improve workflows by reducing paperwork, increasing accuracy in record-keeping, and offering evidence-based tools to support clinical decisions. By centralizing and streamlining data, EHRs foster a more patient-centered approach, ensuring that care is tailored to individual needs. These features collectively make EHRs a cornerstone in modern healthcare systems, significantly contributing to better patient outcomes and operational efficiency [34,15]. 3 AI-Driven Techniques in Medical Image Security 3.1 Machine Learning-Based Encryption: 3.1.1 Securing Medical Image Analysis with Encryption Algorithms in Deep Learning Recent advancements in AI and encryption techniques are transforming healthcare by enabling secure and accurate medical data processing. Naik et al. [28] used DenseNet-121 and AES-128 encryption for identifying lung diseases from chest X-rays. Kumar et al. [21] implemented a cloud-based system for tumor detection in MRI images using CNN with 97.87% accuracy and AES-256 encryption. Mohanty et al. [26] achieved 98.51% accuracy in brain tumor detection with CNN-LSTM secured by a modified SHA-256 algorithm. Other method employed an LSTM model with homomorphic encryption for predicting in-hospital mortality using the MIMIC-III dataset. while [12] developed PINPOINT, a temporal 13 CNN with homomorphic encryption for time-series predictions, including COVID-19 case forecasting. In [27] reviewed homomorphic encryption applications in cancer detection, cardiovascular analysis, and secure healthcare queries. Boulila et al. [8] classified COVID-19 X-rays with MobileNetV2 and partially homomorphic encryption, achieving 93.3% accuracy. These innovations underscore the potential of combining AI with encryption for secure and efficient healthcare solutions. 3.1.2 Integrating Image Encryption and Compression in Deep Learning for Medical Image Processing The security and efficient transmission of medical images is essential due to their large size and sensitive nature. Several recent techniques address both encryption and compression to enhance protection. Selvi et al. [31] developed the ASFSCSLEC-DNL method for secure encryption and compression of chest radiograph images, producing promising results. Ahmad et al. [1] proposed a block-based perceptual encryption algorithm combined with JPEG compression for grayscale and color medical images, tested in TB screening on chest radiographs. Kumar et al. [20] introduced MediSecFed, a secure federated learning framework for chest X-ray datasets, outperforming FedAvg by 15% in hostile environments. Hajjaji et al. [13] proposed a novel crypto-compression algorithm using artificial neural networks and chaotic systems, which successfully preserved the security and quality of the medical image during compression. 3.1.3 Key Generation in Encryption Algorithms for Medical Image Analysis Key generation plays a crucial role in encryption algorithms for medical image analysis, ensuring the confidentiality, integrity, and authenticity of sensitive data. Ding et al. [18] proposed a deep learningbased key generation network (DeepKeyGen), which showed superior security to encrypt medical images, evaluated on data sets such as chest X-rays and the BraTS18 data set. Krishna et al. [19] introduced a dynamic medical image encryption technique using a neural network for key generation, encrypting the key itself for enhanced security. While their method demonstrated strong encryption, the encryption time needs optimization, as tested on X-ray images. 3.2 Watermarking and Data Integrity Verification: Current research on deep learning-based watermarking focuses mainly on image watermarking, with limited work on text and 3D images, offering improved efficiency and robustness by learning complex patterns resilient to attacks, easily re-trained for different applications, and making signature retrieval difficult due to high non-linearity [6,7]. Many methods in the literature presented CNN-based techniques for digital image watermarking that enhance both robustness and imperceptibility. These methods [32,2], [39,17] [22] utilize various CNN architectures, such as encoder-decoder networks and full convolutional neural networks (FCNNs), to efficiently embed and extract watermarks. They also introduce innovative strategies like adversarial training and attack simulation layers to improve resistance against distortions and attacks, ultimately achieving better trade-offs between robustness and imperceptibility. These CNN-based approaches outperform traditional methods, offering greater adaptability to different image resolutions and improving the overall security of the watermarking process. The second class of deep learning-based image watermarking utilizes generative adversarial networks (GANs), including variants like Wasserstein GANs (WGANs) and CycleGANs, known for their effectiveness in providing invisibility and robustness. HiDDeN [40] was the first scheme to use an adversarial discriminator to improve watermarking, featuring an encoder, decoder, and adversary network. ROMark [36] improved HiDDeN by minimizing the loss of decoding in various attacks, while another variant incorporated rotation and noise layers to defend against geometric rotations. Zhang et al. [38] introduced a GAN-based technique using inverse gradient attention (IGA) to improve capacity and robustness. Liu et al. [24] proposed a two-stage separable deep learning framework (TSDL), which trains with true nondifferentiable noise attacks like JPEG compression, achieving improved robustness compared to previous methods. 14 3.3 Privacy-preserving solutions in deep learning-based techniques Recent advances in secure medical data processing highlight the integration of security methods and deep learning to improve accuracy and privacy. Zhang et al. [37] optimized CryptoNets with polynomial ReLU approximations for better classification accuracy in networks with nonlinear layers, while Liu et al. [23] enhanced inference accuracy using MiniONN with secret sharing. Alzubi et al. [4] proposed a blockchainbased BAISMDT model for secure medical data transmission and disease detection. Hesamifard et al. [16] and Carpov et al. [9] emphasized reducing computational costs and improving security in encrypted systems through GPU batch bootstrapping and homomorphic encryption. Federated learning (FL) shows promise in real-world medical data exchange but faces challenges with noisy data, underscoring the need for further research into secure, efficient multiparty computation and privacy-preserving deep learning. 4 Evaluation and Benchmarking of AI Techniques Comparative Analysis: AI techniques in medical image security demonstrate varied performance across encryption strength, detection accuracy, and computational efficiency: •Encryption Strength: Techniques like homomorphic encryption (e.g., MiniONN, CryptoNets) and GAN-based frameworks (e.g., TSDL and IGA) excel in securing medical data during processing and transmission. Approaches integrating modified encryption algorithms, such as AES-128/256 or SHA-256, provide robust data protection, while blockchain-based models like BAISMDT enhance data privacy and integrity during exchange. •Detection Accuracy: CNN-based methods (e.g., DenseNet-121, CNN-LSTM) achieve high diagnostic accuracy, with some models reporting over 98% in medical image classification and tumor detection. GAN-based watermarking techniques also improve robustness and accuracy in image integrity checks. •Computational Efficiency: While encryption techniques like homomorphic encryption and neural network-based key generation offer strong security, they often face higher computational costs. Innovations like GPU acceleration, batch bootstrapping, and compression strategies reduce computational overhead, enabling practical deployment in real-world scenarios. Overall, integrating AI into medical image security balances high accuracy and robust encryption, though computational efficiency remains an area for further optimization. Dataset and Model Limitations: Medical image datasets face challenges of limited diversity, impacting the ability of AI models to generalize across various demographics, imaging technologies, and clinical settings. This lack of diversity hinders model robustness, particularly in ensuring the security of sensitive patient information during processing and transmission. While advancements like adversarial training, federated learning, and encryption-integrated models (e.g., homomorphic encryption, blockchain) improve data security and robustness, the reliance on biased or narrow datasets continues to limit the scalability and reliability of these AI solutions in real-world medical applications. 5 Challenges and Open Issues Deep learning for medical image security using cryptography or watermarking techniques faces various challenges such as •Limited Generalization Deep learning models often struggle to adapt to new or diverse medical image data, leading to decreased performance and security. Future research should focus on creating models that generalize well across various imaging modalities, diseases, and patient groups. •Vulnerability to Adversarial Attacks Adversarial attacks can manipulate input data, compromising the integrity and security of encrypted medical images. Future work should prioritize developing robust training techniques and protective mechanisms to mitigate such vulnerabilities. 15 •High Computational Costs One of the primary challenges in applying deep learning to medical image security is the high computational cost. Training deep learning models requires expensive hardware and extensive time. Future research could focus on optimizing algorithms and utilizing hardware accelerators like GPUs or TPUs to reduce these costs, enabling real-time, scalable solutions in healthcare applications. •Data Availability and Quality The scarcity of large, high-quality datasets due to privacy concerns poses a significant challenge. Future developments should focus on privacy-preserving techniques that enable model training on decentralized or encrypted datasets while maintaining data security. 6 Conclusion This paper explored various AI-driven techniques that play a crucial role in enhancing medical image security. Advanced methods such as convolutional neural networks (CNNs), generative adversarial networks (GANs), federated learning (FL), and homomorphic encryption (HE) have demonstrated remarkable effectiveness in strengthening encryption, improving threat detection, and optimizing computational efficiency. These approaches not only protect sensitive medical data but also ensure its integrity and accessibility within modern healthcare systems. AI has become indispensable in addressing the escalating cybersecurity challenges in healthcare. By enhancing data privacy and mitigating adversarial threats, AI-driven solutions bridge the gap between security demands and the rapid digital transformation of healthcare infrastructures. Their adaptability and scalability make them essential for managing the growing volumes of medical data securely. Looking ahead, the integration of AI presents vast opportunities for advancing medical data security. Future research should focus on developing more efficient, generalizable, and secure models that overcome dataset limitations and computational constraints. As AI continues to evolve, it will play a pivotal role in strengthening healthcare cybersecurity, safeguarding patient privacy, and enabling the seamless exchange of medical information in an increasingly connected world. References [1] Ijaz Ahmad and Seokjoo Shin. A perceptual encryption-based image communication system for deep learning-based tuberculosis diagnosis using healthcare cloud services. Electronics, 11(16):2514, 2022. doi: https://doi.org/10.3390/electronics11162514. [2] Mahdi Ahmadi, Alireza Norouzi, Nader Karimi, Shadrokh Samavi, and Ali Emami. Redmark: Framework for residual diffusion watermarking based on deep networks. Expert Systems with Applications, 146:113157, 2020. doi: https://doi.org/10.1016/j.eswa.2019.113157. [3] Steve Alder. Hacking incidents reported by retinal consultants medical group, three rivers regional commission, ace surgical supply. The HIPAA Journal, Nov 25, 2021. [4] Omar A Alzubi, Jafar A Alzubi, K Shankar, and Deepak Gupta. Blockchain and artificial intelligence enabled privacy-preserving medical data transmission in internet of things. Transactions on Emerging Telecommunications Technologies, 32(12):e4360, 2021. doi: https://doi.org/10.1002/ett.4360. [5] Siva Subrahmanyam Balantrapu. A comprehensive review of ai applications in cybersecurity. International Machine learning journal and Computer Engineering, 7(7), 2024. [6] Aicha Benyoucef and M’Hamed Hamadouche. Roni-based medical image watermarking using dwt and lsb algorithms. In International Conference on Artificial Intelligence and its Applications, pages 468–478. Springer, 2021. doi: https://doi.org/10.1007/978-3-030-96311-843. [7] Aicha Benyoucef and M’Hamed Hamaouche. Region-based medical image watermarking approach for secure epr transmission applied to e-health. Arabian Journal for Science and Engineering, 49(3):4025–4037, 2024. doi: https://doi.org/10.1007/s13369-023-08263-0. [8] Wadii Boulila, Adel Ammar, Bilel Benjdira, and Anis Koubaa. Securing the classification of covid19 in chest x-ray images: A privacy-preserving deep learning approach. In 2022 2nd International 16 by using random horizontal flipping, which can help the model learn to detect objects from different perspectives. 3.3 Object detection and region cropping In this stage, Faster R-CNN is used to detect and isolate regions of interest. Specifically, we take advantage of Faster R-CNN’s ability to operate as a computerized lens, to scrupulously navigate through the 2D slice images to discern and delineate regions that house potential fracture sites within the cervical spine and to underscore their associated vertebra with bounding boxes. This act of object localization constructs a vital foundational tier, guiding the ensuing procedures in the pipeline, which are designated to further refine, dissect, and classify these pronounced areas suspected of fractures. The localized regions of interest are subsequently cropped from the rest of their associated slices to form small imagettes. Specifically, the aim of this phase is dual: firstly, to drastically curtail computational excess, and secondly, to concentrate the ensuing analysis on clinically pertinent regions. Explicitly, the sectors of the cervical spine believed to harbour fractures, as pinpointed by Faster R-CNN. 3.4 Classification via Next-ViT In this final stage, Next-ViT model is used to binary classify the imagettes previously produced to distinguish between those really containing fractures and those that are not. This model was selected for its unique set of attributes that align impeccably with our research goals. One of the standout qualities of Next-ViT is its data efficiency. The model demonstrates impressive performance even when subjected to small, annotated datasets. In addition, Next-ViT diverges from CNNs by incorporating self-attention mechanisms. These latter mechanisms excel at identifying complex spatial and contextual relationships within images, a feature invaluable for interpreting the complex imagery commonly found in cervical spine studies. Moreover, given the underwhelming results of our initial attempt to train a vision transformer from scratch, we have chosen to adopt a pre-trained Next-ViT architecture, which led to a marked improvement in our system’s efficacy. However, to better fit Next-ViT to our problem, we have performed a refinement training of the model, specifically using data augmentation techniques by applying simple transformations to the training dataset (i.e. rotation, scaling, and flipping). We hypothesize that these simple transformations assist the model in understanding underlying data patterns, thereby improving its learning capability. 4 Proposed cloud-based architecture for cervical spine fracture detection As a second contribution in this paper, we describe a robust and scalable cloud-based system that is dedicated to the detection of cervical spine fractures. The cloud infrastructure serves as the backbone supporting the entire multifaceted computational pipeline presented in Section 3and offers unique advantages both in terms of computational resources and data management. 4.1 Motivations and goals The presented architecture is motivated by several compelling incentives for coupling cloud computing and deep learning models in the arena of cervical spine fracture detection. The impetus for adopting a cloud-based approach originates from a critical need to address challenges in scalability, data integrity, and real-time analytics. Below are the main key motivations: 1. Superior diagnostic accuracy: Traditional diagnostic approaches, although useful, sometimes fail to identify complex or subtle fractures. The marriage of cloud-based computational power and well established deep learning models has the potential to usher in a new era of nuanced and precise diagnoses. 2. Operational efficiency: Utilizing the distributed computing power of the cloud alongside deep learning models that can efficiently parse large sets of image data enhances the operational efficiency of the diagnostic process. This could significantly reduce the time radiologists need to reach a diagnosis. 23 3. Scalability and adaptability: The inherent scalability of cloud infrastructure is well-suited for handling the voluminous medical imaging data generated daily. This removes the need for healthcare organizations to make significant investments in local computing resources. 4. Broadened access to advanced tools: Cloud-based systems democratize access to cutting-edge diagnostic technologies. This model allows healthcare providers, regardless of their size or location, to benefit from state-of-the-art tools without prohibitive upfront costs. 5. Augmentation of clinical decision-making: The synergy between cloud technology and deep learning models can act as a potent decision-support mechanism. It can provide preliminary evaluations that assist healthcare professionals in making timely and well-informed decisions. 6. Future-ready integration: The modular architecture of cloud-based systems makes them ripe for seamless integration with existing electronic health records. This offers the possibility for more integrated, collaborative approaches to healthcare delivery in the future. Thus, the integration of cloud computing and deep learning models in the detection of cervical spine fractures has the potential to surmount existing limitations, refine diagnostic protocols, democratize access to state-of-the-art technologies, and fundamentally transform clinical practices in this vital area of healthcare. 4.2 Description of the proposed cloud-based architecture The proposed architecture is designed to be deployed in Google Cloud Platform (GCP), integrating its services to offer an efficient end-to-end cervical spine fracture detection workflow. A general view of the proposed cloud-based architecture for cervical spine fracture detection is illustrated in Fig. 2. Initially, the overarching vision of crafting an integrated end-to-end diagnostic workflow for enhanced cervical spine fracture detection stemmed from comprehensive brainstorming sessions. Significantly, it was our deep dive into the vast capabilities of the Google Cloud Platform (GCP) that galvanized our alignment with this mission. Building upon this foundation, our hands played a pivotal role in the ensuing architectural design and execution phase. Furthermore, recognizing the paramount importance of data integrity, we have channelled significant efforts into devising an efficient automatic ingestion mechanism for CT scans. Simultaneously, with an acute awareness of the sensitive nature of medical data, we have championed the incorporation of a robust encryption protocol, ensuring that data remain secured. Transitioning from data acquisition, our focus then have gravitated towards the multi-layered data pipeline. Specifically, we have integrated in the proposed architecture the data pipeline elaborated in Section 3, that meticulously optimize the mechanisms of pre-processing, feature extraction, and fracture detection. On the other hand, we believe in the interdependent nexus between machine learning methods and human expertise and its capacity to offer better solutions, especially when they are combined appropriately. This conviction has led to the establishment of a systematic feedback loop, where the invaluable insights of medical professionals continuously enrich our cloud-based system. Through this mechanism, their diagnostic evaluations directly inform and steer the iterative enhancements of the integrated models in the proposed system. Moreover, with an ever-evolving medical landscape, we need to ensure that the used models in the architecture underwent consistent training sessions. By leveraging insights from the analytical database, our diagnostic algorithms remain at the cutting edge, always adaptive to the latest nuances in medical diagnostics. Beyond the technical realm, we endeavour to foster a culture of interdisciplinary collaboration. By orchestrating synergy between cloud experts, data scientists, and medical professionals, we strive to ensure that our collective expertise coalesced seamlessly. This unity of purpose and knowledge-sharing became instrumental in shaping our presented solution. In summary, witnessing the transformative potential of our architecture in the realm of medical diagnostics has been both a privilege and a testament to the collaborative prowess of our team. Our journey exemplifies the boundless possibilities that emerge when cloud computing and machine learning converge, especially in the ever-critical domain of healthcare. The amalgamation of GCP’s advanced services presents a promising horizon for medical diagnostics. While this overview provides a highlevel design, the actual implementation should be tailored according to specific requirements, ensuring 24 Epoch Precision Recall mAP0.5 80 0.9287 0.8857 0.9424 100 0.9687 0.9057 0.9724 Table 1: Performance metrics of the Faster R-CNN on the RSNA dataset. a balance between functionality, budget, and privacy concerns. Collaboration with cloud and domain experts is essential for the successful realization of such a system. Figure 2: A representation of the proposed cloud-based system. 5 Evaluation and discussion of the cervical spine fracture detection system The proposed multifaceted data pipeline has been trained and evaluated using the large RSNA public dataset containing cervical spine CT scans [2]. In the setting of this work, the images of the dataset was split into training (80%) and validation (20%) sets. The slices and their corresponding label files (.txt files) are then organized appropriately into separate directories for training and validation. Subsequently, we have downloaded Faster R-CNN’s code from TensorFlow, adjusted it and trained it to meet our purpose. Thus, the performance of Faster R-CNN on the used dataset is assessed using standard evaluation metrics, namely: precision,recall, and mean average precision at IoU (mAP50). The obtained results after 80 and 100 epochs are presented in Table 1. The yielded results showcase the model’s potential in both recognizing and pinpointing objects within images after 80 and 100 epochs. A summary of Faster R-CNN train loss metrics from one of the epochs during the model’s training phase are present in Table 2. The table summarizes important performance indicators and parameters that provide insights into the model’s training dynamics, namely : Loss, Loss Classifier, Loss Box Reg, Loss Objectness, and Loss RPN Box Reg. For visual illustration of the cropping operation results, we give in Fig. 3cropped images obtained from different slices. Concerning the classification stage of the multifaceted data pipeline, the implementation of Next-ViT requires setting appropriate values for model’s parameters. This is specifically done to insure a satisfying 25 Parameter Best value Averaged value Loss 0.1576 0.3236 Loss Classifier 0.0490 0.1163 Loss Box Reg 0.0900 0.1130 Loss Objectness 0.0088 0.0790 Loss RPN Box Reg 0.0040 0.0153 Table 2: Training Loss results. Figure 3: Illustrations of cropped vertebra. balance between computational efficiency and detail resolution to make the model highly applicable in clinical settings for which timely and accurate diagnosis is paramount. In the context of this work, we have considered the parameter tuning exhibited in Table 3. Parameter Value Patch size 16 ×16 Latent space dimension 192 Number of encoder blocks 12 Number of MLP heads 3 Total parameters ≈5.5M Table 3: Next-ViT model parameter tuning. Also, it is worth to note that we have applied other adaptations to Next-ViT to meet our specific needs. For instance, the output shape is printed and should be [16, 2] of shape. This is to say that for each input image, we get 2 values as output, corresponding to fracture and no fracture results respectively. In addition, to optimize the neural network, we have employed RAdam optimizer and used a learning rate of 0.001. Specifically, the value of 0.001 is considered a moderate choice, which is neither too high to cause instability nor too low to slow down the learning process. This value is often recommended for Adam and its variants like RAdam due to its effectiveness in a wide range of scenarios. To validate the robustness and effectiveness of Next-ViT model, we have used two metrics: accuracy and loss. Hence, the obtained validation results of the model on the RSNA 2022 Cervical Spine Fracture Detection dataset are shown in the graphs presented in Fig. 4. From the latter figure, it is easy to notice that the performance of the Next-ViT model improved through the epochs for both accuracy and loss validation metrics, until achieving a validation accuracy of 95.5% and a validation loss of 2%. This is particularly promising because it suggests that the Next-ViT could be used to develop a fast and accurate 26 AI-based system for cervical spine fracture detection. Such a system could be used to help radiologists identify fractures more quickly and reliably, and it could also be used to screen patients for suspected fractures in emergency settings. Figure 4: Validation results of Next-ViT. Moreover, a comparison of the work presented herein with the concurrent work of Showmick Guha et al. [1] that is recently exhibited in the literature is reported in Table 4. From the latter, it is easy to notice that the model of Showmick Guha et al. [1] presents the best accuracy currently. However, the proposed model is more subtitle to offer a superior diagnostic accuracy in the future. In fact, the exhibited architecture foresees continuous model refinement training by taking into consideration the capacity of accepting new unseen data as well as correcting feedback from experts who use the system. So, the two features guarantee a continuously improving diagnostic accuracy. Furthermore, despite the fact that the two works ensure a real time response, nevertheless, the proposed system is clearly more adapted to clinical routines considering the fact that it is deployed on the cloud. Hence, it offers a better scalability and adaptability in terms of resources, brocaded access to advanced tools, and future-ready integration compared to its concurrent work which is designed for miniaturized systems essentially made for a personal use. Feature Proposed work Showmick Guha et al. [1] Best accuracy 95,5 % 99.75 % Deployment Cloud Android application Real-time response Considered Considered Model refinement possibility Considered Not considered Data integrity Considered Not considered Storage and processing capacity High Very low Table 4: Comparison between the proposed work and a concurrent work according to few features. 6 Conclusion The confluence of AI, medical imaging, and cloud computing represents a promising avenue for revolutionizing the healthcare domain. In this setting, we have made a couple of contributions which are exhibited in this document. Mainly, we have introduced a new comprehensive computational data pipeline tailored for the detection of cervical spine fractures. Specifically, the proposed data pipeline is composed of four 27 stages, each of which fulfils a unique role to achieve high diagnostic precision and reliability. Furthermore, motivated by several key goals such as improving diagnostic accuracy, increasing scalability, and enhancing data security, we have exhibited, as our second main contribution, a new cloud-based system to extend the capabilities of our computational data pipeline. The new cloud-based architecture represents a paradigm shift in how cervical spine fractures can be detected and managed. The proposed cloud-based system not only streamlines the workflow but also allows for continuous improvement through real-time feedback mechanisms. Furthermore, the experimental study comprising the implementation, training, and validation of the presented comprehensive computational data pipeline over the RSNA 2022 Cervical Spine Fracture Detection dataset has shown an encouraging performance with regard to a concurrent work in the literature. While the findings of this paper are compelling, they raise several salient questions that could form the basis of future scholarly inquiry. These include prototyping the proposed cloud-based diagnostic system and its convenience to resource-constrained devices, as well as further refinements and improvements of the proposed multifaceted data pipeline by the integration visualization mechanisms or with the adoption of emerging artificial intelligence paradigms such as deep reinforcement learning and federated learning. Acknowledgment The authors are thankful to the anonymous reviewer for his valuable comments that helped to improve the paper. References [1] P. Showmick Guha, S. Arpa, and A. Md. A real-time deep learning approach for classifying cervical spine fractures. Healthcare Analytics, 4:100265, 2023. [2] H.M Lin, E. Colak, T. Richards, F.C. Kitamura, L.M. Prevedello, J. Talbott, R.L. Balland E. Gumeler, K.W. Yeom, M. Hamghalam, et al. The RSNA cervical spine fracture CT dataset. Radiology: Artificial Intelligence, 5:e230034, 2023. [3] Z. Merali, J. Wang, J.H. Badhiwala, C.D. Witiw, J.R. Wilson, and M.G. Fehlings. A deep learning model for detection of cervical spinal cord compression in mri scans. Scientific Reports, 11(1):14620, 2021. [4] H. Salehinejad, E. Ho, H.M. Lin, P. Crivellaro, and O. Samorodova. Deep sequential learning for cervical spine fracture detection in computed tomography imaging. IEEE Transactions on Medical Imaging, 40(6):1642–1652, 2021. [5] J.W. Savage, G.D. Schroeder, and P.A. Anderson. Vertebroplasty and kyphoplasty for the treatment of osteoporotic vertebral compression fractures. J Am Acad Orthop Surg, 22(10):653–664, Oct 2014. [6] M. Shaolong, H. Yang, C. Xiangjiu, and G. Rui. Faster rcnn-based detection of cervical spinal cord injury and disc degeneration. Medical Physics, 48(2):801–813, 2020. [7] J.E. Small, P. Osler, A.B. Paul, and M. Kunst. Ct cervical spine fracture detection using a convolutional neural network. Journal of Computer Assisted Tomography, 45(4):578–585, 2021. [8] D.T. Tuan, Q.H. Le, and T.H. Nguyen. Cervical Spine Fracture Detection via Computed Tomography scan. PhD thesis, FPT University, 2022. [9] A.F. Voter, M.E. Larson, J.W. Garrett, and J.P.J. Yu. Diagnostic accuracy and failure mode analysis of a deep learning algorithm for the detection of cervical spine fractures. AJNR Am J Neuroradiol, 42(8):1550–1556, Aug 2021. 28 Advanced Ensemble Learning Framework for Reliable Smart Grid Stability detection Saliha Mezzoudj1, Yasmina Saadna2, and Meriem Khelifa3 1Department of Computer Science, University of Algiers, Algiers, Algeria , [email protected] 2Labstic laboratory, Batna 2 University, Batna, Algeria , [email protected] 3Artificial Intelligence of Information Technologies, Department of Computer Science and Information Technologies, University of Kasdi Merbah Ouargla, Algeria , [email protected] Abstract The increasing complexity of smart grid systems necessitates advanced methodologies to ensure reliable stability classification and seamless power delivery across consumer domains. This study introduces an innovative ensemble learning framework designed to classify smart grid stability using the Smart Grid Stability Augmented dataset. The proposed framework integrates multiple ensemble techniques, including Bagging, AdaBoost, Stacking, and Voting Classifiers, to improve robustness, accuracy, and reliability. A 5-fold cross-validation strategy is implemented to minimize overfitting and validate model performance. The dataset undergoes preprocessing with feature standardization and binary encoding of the target variable to ensure uniform contributions from all features. Experimental results indicate that the soft Voting Classifier, which is a combination of single machine learning models logistic regression, support vector machine, and random forest (LR+SVC+RF), outperforms other models by achieving a peak accuracy of 97.3%, demonstrating exceptional stability classification performance. Compared to individual machine learning models and existing state-ofthe-art approaches, the proposed ensemble framework exhibits superior performance across multiple evaluation metrics. These results underscore the potential of ensemble learning in enhancing smart grid stability, contributing to more reliable and efficient power grid management systems. Keywords: Grid Stability, Ensemble Learning, Bagging, AdaBoost, Soft Voting, stacking, CrossValidation 1 Introduction The smart grid is an advanced concept aimed at transforming the future electricity network by enhancing its flexibility, adaptability, and autonomous management [20], [16]. This complex system incorporates various interconnected subsystems [16], integrating diverse disciplines and enabling the autonomous operation and control of its parts. It is geographically spread out and consists of a wide range of components. Additionally, the smart grid demonstrates emerging behaviors and continuous development. As a key element in a global network of linked systems, it encourages collaboration to promote the development of innovative services across different sectors. The primary factors propelling advancements in this field are energy efficiency and optimized resource management at both local and global levels, requiring comprehensive monitoring and control [12]. As electricity demand rises with population growth, the dependence on natural resources for power generation increases. Nevertheless, this process remains intricate and expensive. Significant research has been directed towards enhancing grid networks to improve power distribution efficiency. The smart grid offers a promising solution by leveraging Information and Communication Technology (ICT) to gather data on consumer behavior, thereby enabling the creation of context-aware systems that optimize power distribution efficiency [10]. Traditional stability analysis and control methods have proven insufficient for managing the complexities of modern smart grids. In response, recent advances in artificial intelligence (AI) provide effective tools to meet the high demands of security and stability in these systems [14]. The development of an intelligent grid that can accurately predict power demand is essential. This can be achieved through the application of Machine Learning (ML) algorithms [3], [6] to analyze the large amounts of data generated by the grid. These advancements in smart grid technology are crucial for reducing environmental pollution and lowering electricity costs, promoting a more cost-effective and sustainable energy system. Recent developments in artificial intelligence (AI) and machine learning (ML) have significantly enhanced 29 the prediction and management of smart grid stability and energy systems. Oqaibi and Bedi (2024) introduced a hybrid forecasting system that integrates data deconstruction and attention mechanisms, achieving a prediction accuracy of 90.45% on the Kaggle dataset. Their work emphasizes the need for optimizing hybrid models to reduce computational complexity and improve prediction efficiency [15]. Xu et al. (2024) proposed a time-series depthwise separable convolutional neural network (CNN) with an attention mechanism, reaching 88.9% accuracy using the UCI dataset. Their study underscores the importance of large datasets for effectively training deep learning models [8]. Further contributions include Mohsen (2023), who developed an efficient artificial neural network (ANN) model for Decentralized Smart Grid Control (DSGC) systems, achieving a testing accuracy of 97.36% and a perfect AUC score of 100% through hyperparameter tuning [18]. Similarly, Alsirhani (2023) combined Multi-Layer Perceptron and Extreme Learning Machine (MLP-ELM) with Principal Component Analysis (PCA), attaining 95.8% accuracy, which highlights its potential for improving grid reliability amid fluctuating energy demands and growing renewable integration [1]. Javaid (2022) proposed a novel stacking ensemble model, MLBCSM, which combines multiple boosting classifiers (AdaBoost, XGBoost, HistBoost, CatBoost, LGBoost) with an Adaptive Synthetic Sampling Technique (ADASYN) to address data imbalance. The model, which includes data preprocessing, balancing, and classification, outperformed traditional methods, achieving 92.39% accuracy and 93.22% recall. These results demonstrate its effectiveness in detecting the stability in smart grids [14]. In this work, a novel ensemble-based machine learning approach is proposed to predict the stability of smart grids by classifying the Smart Grid Stability Augmented Dataset. The experimental results are compared with recent machine learning algorithms, including individual classifiers such as KNN, NB, Support Vector Classifiers, as well as ensemble methods like Bagging, AdaBoost, Stacking, and soft Voting Classifiers. The main steps involved in our contribution: 1. The Smart Grid Stability Augmented Dataset is loaded, and stability labels are mapped to binary values. The dataset is shuffled, and features are standardized to facilitate model convergence. 2. Four ensemble learning models Bagging, AdaBoost, Stacking, and Voting Classifiers are defined. These models utilise a variety of base learners, including Decision Trees, Logistic Regression, Random Forest, and Support Vector Classifier, to enhance predictive accuracy. 3. A 5-fold cross-validation strategy is employed to evaluate model performance, ensuring robust estimates of model effectiveness and mitigating overfitting risks. 4. Quantitative comparison of the models’ performance is provided. The voting achieves the highest accuracy and AUC, while Stacking Classifier excels in integrating multiple models. AdaBoost and Bagging show strong performance in balancing precision and recall, and the Voting Classifier provides competitive results across all metrics. with voting method reaching an accuracy of 97.3%, demonstrating superior predictive capability compared to individual models and other state-of-the-art models. The rest of the paper is organized as follows. Section II discusses recent state-of-the-art literature related to the application of deep learning algorithms on smart grids. In Section III,the proposed model is discussed in detail. Experimental results are discussed in Section IV, which is followed by a conclusion and future work in Section V. 2 Proposed approach A variety of machine learning algorithms can be applied to the problem of stability detection, with their effectiveness typically evaluated using metrics such as accuracy and false positive rates. To improve prediction performance and reduce false positives, researchers have proposed numerous ensemble learning methods. Ensemble learning techniques combines multiple machine learning algorithms to achieve enhanced predictive performance compared to standalone models [13]. Broadly, ensemble learning is categorized into two types: parallel and sequential [17] Parallel methods, such as bagging and random forests, train independent base classifiers to promote diversity, whereas sequential methods, including boosting, iteratively refine weak learners to improve accuracy. Ensemble methods are particularly robust and adaptable, excelling in scenarios involving noisy or complex data. This study introduces an ensemble learning framework to classify the stability of smart grid systems using the Smart Grid Stability Augmented dataset. The methodology incorporates various ensemble techniques to enhance the robustness, accuracy, and reliability of stability predictions. In this section, we describe the architecture 30 of our system as shown in Figure 1. Figure 1: Architecture of the proposed system The approach is structured as follows: 2.1 DATA PREPROCESSING Pre-processing is a critical step in improving data quality and enhancing the performance of machine learning (ML) models. The Smart Grid Stability Augmented dataset includes features related to grid stability and a target label (stabf) that indicates stability (stable or unstable). The variability in feature ranges within the dataset can lead to biases, as features with higher magnitudes may dominate during model training. To address this, StandardScaler is employed for data normalization, ensuring all features contribute equally. This technique transforms the data into a common scale, thereby enhancing classifier performance. The categorical target variable is mapped to binary values, where: 0: : represents an unstable state and 1: indicates stability. Non-numeric values in the dataset are converted to numeric format using encoding techniques to make the data suitable for ML algorithms. Additionally, the dataset is shuffled to mitigate any ordering bias, further ensuring the robustness and reliability of the training process. 2.2 CROSS VALIDATION To enhance model reliability and generalization, a 5-fold cross-validation strategy is employed. The dataset is divided into five equal parts, where each part is used as a validation set once, while the remaining four are used for training. This approach minimizes overfitting and ensures. 2.3 ENSEMBLE LEARNING METHODS In this part, we explore the ensemble learning paradigm, focusing on its fundamental components, combination techniques for base learners, and methods for selecting ensembles. 2.3.1 Bagging (Bootstrap Aggregating) classifier Bagging, or Bootstrap Aggregating, is an ensemble learning technique designed to reduce model variance and improve predictive accuracy by combining multiple base models [4], [21]. In this approach, each base model is trained on a distinct bootstrapped sample of the dataset, created by random sampling 31 with replacement. For this study, the Decision Tree Classifier is utilized as the base learner, with a total of 50 estimators. Each tree is independently trained on a bootstrapped sample, enabling the model to capture diverse patterns within the data. After training, predictions for unseen data are obtained by aggregating the outputs of all 50 models: •For regression tasks, the final prediction is the average of all individual predictions. •For classification tasks, the final prediction is determined by majority voting among the models. This ensemble strategy significantly reduces the variance of the model compared to a single decision tree, leading to more stable and accurate predictions. The Bagging approach is particularly effective for noisy or complex datasets, such as those encountered in smart grid stability detection. By leveraging multiple models and aggregating their outputs, Bagging enhances the robustness and generalization ability of the framework, making it a reliable choice for high-stakes applications in smart grid systems. 2.3.2 AdaBoost (Adaptive Boosting) Classifier In this study, the AdaBoost algorithm was chosen as the boosting method. Developed by Freund and Schapire [7], AdaBoost is one of the most widely used boosting techniques, offering a strong theoretical foundation and proven efficacy in generating accurate predictions. AdaBoost constructs a strong classifier by combining the weighted outputs of weak classifiers, addressing earlier boosting methods limitations. In our implementation of AdaBoost begins by initializing equal weights for all training samples. In each boosting round, a weak learner (in this case, a Decision Tree Classifier) is trained on the weighted dataset. The classifier’s weighted error is calculated, and its performance is quantified using a weight αt. This weight determines the importance of the weak learner in the final ensemble. Misclassified samples are assigned higher weights, making them more influential in subsequent iterations. The process is repeated for Tboosting rounds, where T= 50 in this implementation. The final prediction is made by combining the outputs of all weak classifiers, weighted by their respective importance values. 2.3.3 Stacking Stacking is an ensemble learning technique that combines predictions from multiple base models (level-0 models) and refines them using a meta-model (level-1 model) [3]. The primary goal is to leverage the strengths of individual models and optimize the final prediction by training an additional layer. In our implementation: •Base Models: Random Forest and Support Vector Classifier are trained independently on the training dataset. Each model generates predictions, capturing unique patterns within the data. •Meta-Model: Logistic Regression is used as a second-level model, which takes the predictions of the base models as input. It learns to combine these predictions optimally, mitigating individual weaknesses. •Workflow: The training process involves generating predictions for the validation set using the base models, constructing a new dataset comprising these predictions, and training the meta-model on this dataset. During inference, the base models generate predictions for unseen data, which are then aggregated by the meta-model to produce the final output. 2.3.4 Voting Classifier Algorithm (Soft Voting) The voting Classifier is an ensemble learning method that combines predictions from multiple base models to improve predictive accuracy and robustness [9]. In the context of this study, Soft Voting is used, which involves averaging the predicted probabilities from each of the base models. The base models utilized in this framework include: •Random Forest •Support Vector Classifier (SVC) •Logistic Regression 32 distinguish from authentic media. This section provides an overview of the most prominent deepfake generation methods and their impact on digital media.  Generative Adversarial Networks (GANs) Generative Adversarial Networks (GANs) are among the most widely used architectures for deepfake generation. A GAN consists of two competing neural networks: a generator that produces synthetic media and a discriminator that attempts to distinguish between real and generated content. Through iterative training, the generator improves its ability to create highly realistic outputs. Advanced variations, such as StyleGAN and StyleGAN2, have significantly enhanced the quality of generated images by enabling fine-grained control over facial features, expressions, and lighting conditions [1]. Recent studies have also explored the potential of Latent Flow Diffusion (LFD), which incorporates optical flow sequences in the latent space to enhance temporal coherence in deepfake videos [11]. Compared to conventional GANs, LFD provides better preservation of spatial and motion consistency, making generated videos appear more authentic.  Variational Autoencoders (VAEs) and Hybrid Models Variational Autoencoders (VAEs) are another class of generative models used in deepfake generation. Unlike GANs, which rely on adversarial training, VAEs learn a probabilistic representation of data to generate realistic samples. They have been particularly effective in face-swapping applications, where they enable smooth blending of facial features while maintaining structural consistency [15]. Hybrid models combining GANs and VAEs have also gained traction. These models leverage the structured latent space of VAEs with the adversarial refinement of GANs to generate higher-quality deepfakes. The integration of attention mechanisms within these architectures has further improved the realism of generated media by focusing on fine details such as skin texture and micro-expressions [2].  Face manipulation techniques Face manipulation techniques in deepfake generation can be categorized into three main types: Face Swapping: This technique replaces the face of a person in a video with another person’s face while maintaining the original facial expressions and movements. It is commonly implemented using autoencoders and GANs. The DF-Platter dataset [12] demonstrates that face-swapping deepfakes can be generated at both high and low resolutions, highlighting the challenges in detection. Facial Attribute Manipulation: This method alters specific facial features such as age, gender, and expressions. It is achieved using models like StarGAN and AttGAN, which modify targeted attributes while preserving the overall identity of the subject [2]. Such manipulations are widely used in applications ranging from entertainment to identity anonymization. Lip-Sync Manipulation: This technique synchronizes lip movements with an audio track, making it appear as though a person is speaking words they never actually said. Models like Wav2Lip and SyncGAN have demonstrated impressive results in creating realistic lip-sync deepfakes, posing significant challenges in forensic detection [15].  Text-to-Image and Text-to-Video Synthesis With the advent of large-scale generative models, deepfake generation has extended beyond face manipulation to full-body synthesis. Text-toimage and text-to-video synthesis models, such as DALL  E and Stable Diffusion, enable the creation of highly realistic synthetic content based on textual descriptions. These models use diffusion processes to iteratively refine images, resulting in high-fidelity outputs that can be used for both benign and [14]. Furthermore, recent research has explored deepfake phylogeny, which examines how iterative manipulations can evolve deepfakes over multiple generations, leading to increasingly deceptive synthetic media [13]. The DeePhy dataset was developed to study the progression of deepfakes and their impact on detection algorithms.  Challenges in Deepfake Generation While deepfake generation techniques have significantly improved, they present substantial ethical and security concerns. The ability to create highly realistic synthetic media has raised issues related to misinformation, identity fraud, and political propaganda. The development of novel detection techniques must keep pace with advancements in generation methods to mitigate potential risks [20]. Moreover, existing deepfake generation models often suffer from limitations such as excessive computational requirements, data dependency, and difficulty in generating highly dynamic scenes. Researchers are exploring ways to enhance the efficiency and realism of these models while addressing concerns related to misuse and ethical responsibility [4]. Deepfake generation techniques have advanced rapidly with the integration of GANs, VAEs, and hybrid models. Face manipulation methods such as face swapping, attribute manipulation, and lip-syncing have reached new levels of realism, making detection increasingly 39 challenging. The emergence of text-to-image and text-to-video synthesis models has further expanded the capabilities of deepfake technology. However, as generation methods evolve, the need for robust and adaptive detection frameworks becomes more critical. Future research must focus on improving the interpretability of generative models, developing counter-measures against adversarial attacks, and ensuring the ethical use of deepfake technology. 2.2 Deepfake Detection Techniques As deepfake generation techniques continue to evolve, detecting these synthetic manipulations has become an essential challenge in digital media security. Various deep learning-based approaches, including Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and Transformer-based models, have been developed to distinguish real content from manipulated media. This section provides an overview of state-of-the-art deepfake detection techniques and their effectiveness in different application domains.  Convolutional Neural Networks (CNNs) for Image and Video Detection Convolutional Neural Networks (CNNs) have been widely adopted for deepfake detection due to their ability to extract spatial features from images and videos. CNN-based models analyze inconsistencies in pixel distributions, texture artifacts, and facial asymmetries that may not be perceptible to the human eye. Studies have shown that CNNs, particularly Xception and MobileNet architectures, achieve high accuracy in detecting face-swapping deepfakes, with results ranging between 91% and 98% depending on the dataset used [7]. Despite their effectiveness, CNN-based models face challenges when applied to real-world deepfakes. These models often struggle with generalization across different datasets due to biases introduced during training. Additionally, CNNs primarily focus on spatial features, making them less effective in detecting temporal inconsistencies in deepfake videos [17].  Recurrent Neural Networks (RNNs) and Temporal Analysis To address the limitations of CNNs in video deepfake detection, Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) networks have been utilized for temporal analysis. These models analyze sequential frames in a video to detect unnatural facial movements, such as inconsistent blinking patterns or unnatural lip-syncing [9]. The use of spatiotemporal convolutional networks has further enhanced the capability of RNN-based approaches. For example, the Celeb-DF dataset benchmark demonstrated that incorporating temporal features significantly improves detection accuracy, outperforming frame-based detection models [10]. However, these approaches remain computationally expensive and require substantial processing power, limiting their feasibility for real-time applications.  Transformer-Based Models for Deepfake Detection Recent advancements in deep learning have led to the adoption of Transformer-based models for deepfake detection. Vision Transformers (ViTs) leverage self-attention mechanisms to capture both local and global dependencies in an image, making them highly effective in detecting subtle deepfake artifacts. Multi-modal Transformer architectures, such as M2TR, integrate RGB and frequency-domain features to improve detection accuracy [10]. Compared to CNNs and RNNs, Transformer-based models demonstrate superior generalization capabilities across different datasets. They are particularly effective in detecting complex deepfakes that incorporate high-quality synthesis techniques. However, their high computational cost remains a challenge, necessitating further research into optimization techniques for practical deployment [6].  Multi-Modal Deepfake Detection Approaches Multi-modal deepfake detection approaches integrate information from multiple sources, such as visual and auditory cues, to enhance detection robustness. Joint audio-visual deepfake detection has been proposed as an effective strategy, leveraging synchronization inconsistencies between speech and facial expressions [22]. These methods have shown promising results in identifying lip-sync deepfakes and voice-cloning manipulations. In addition to audio-visual synchronization, PRNU (Photo-Response Non-Uniformity)-based methods have been explored for deepfake detection. PRNU, commonly used in digital forensics, identifies unique device fingerprints left during the image capture process. Recent studies indicate that PRNU-based approaches can complement deep learning models in hybrid detection frameworks [8]. 40  Challenges in Deepfake Detection Despite advancements in deepfake detection, several challenges remain: Generalization Across Different Datasets: Most deepfake detection models struggle with dataset-specific biases. Methods trained on one dataset often fail to generalize well to unseen deepfakes generated by different techniques [3]. Adversarial Robustness: Adversarial attacks can be used to fool deepfake detection models by introducing imperceptible perturbations. This highlights the need for more robust adversarial training strategies [16]. Real-Time Processing Efficiency: Many state-of-the-art detection models are computationally intensive, making real-time deepfake detection a significant challenge [18].  Future Directions in Deepfake Detection To improve deepfake detection, future research should focus on: Hybrid Detection Models: Combining CNNs, RNNs, and Transformer-based models to leverage their respective strengths. Few-Shot and Zero-Shot Learning: Reducing reliance on large labeled datasets to enhance detection generalization [21]. Blockchain and Forensic Watermarking: Implementing digital watermarking techniques to verify content authenticity and track manipulations [5]. 3 Contribution This section presents the methodology adopted for deepfake detection. It begins with a description of the proposed project, followed by the system architecture and the development process of the models used. The chapter also includes details on the dataset, implementation, and performance evaluation of the deep learning models. 3.1 Project Description Our proposed project consists of two main phases: the generation phase using GANs and the detection phase, where we evaluate two efficient deep learning models—CNN and ViT—to differentiate between real and fake images.  Generation Phase (Using GANs) : In the generation phase, fake images are created using a GAN architecture, which comprises two adversarial neural networks: a generator and a discriminator. The generator takes a random latent vector as input and produces synthetic images, which are then passed to the discriminator. The discriminator, which has access to both real and generated images, is trained to distinguish between them, thereby forcing the generator to improve its ability to create realistic images. This generation process is crucial because deepfake detection models rely on deep learning, which requires large datasets for accurate predictions. However, many existing datasets suffer from low image resolution and are too small to effectively train advanced models like ViT.  Detection Phase (Using CNN & ViT Models) For the detection phase, both the CNN and ViT models receive an image as input. Before processing, the images go through a data preprocessing step to ensure optimal training and testing conditions. The dataset is then split into training and testing sets and fed into either the CNN or ViT model. Once training is complete, we evaluate the model’s performance and save the trained model for real-world predictions. The trained models can then analyze new images and determine whether they are real or fake. The proposed system follows a structured pipeline, as illustrated in the system architecture diagram, which includes both the generation and detection phases, ensuring a robust and efficient deepfake detection approach (See Figure 1). Figure 1. 3.2 Deep Convolutional GAN (DCGAN) Development The Deep Convolutional Generative Adversarial Network (DCGAN) is used for generating fake images. It consists of two main components:  The Discriminator Model The first step is to define the discriminator model. The model must take a sample image from our dataset as input and output a classification prediction as to whether the sample is real or fake. This is a binary classification problem: 41 Figure 1: Workflow of the proposed system 1. Inputs: An image with one channel and a resolution of 256 Ö 256 pixels. 2. Outputs: A binary classification, where the model predicts the likelihood that the input image is real or fake. The discriminator architecture consists of: 3. Five convolutional layers (Conv2D), each followed by: (a) LeakyReLU activation (instead of ReLU) to allow better gradient flow. (b) Batch normalization to stabilize training. (c) Dropout layers to prevent overfitting. 4. A final dense layer with a sigmoid activation function, which outputs a probability score. A final dense layer with a sigmoid activation function, which outputs a probability score. The model is trained using the binary cross-entropy loss function, with the Adam optimizer (learning rate = 0.00015, momentum = 0.5) to ensure stability.  The Generator Model The generator is responsible for creating fake images. It takes a latent vector (random noise) as input and transforms it into a realistic image through a series of upsampling layers. 1. Inputs: A 100-dimensional latent space vector sampled from a Gaussian distribution. 2. Outputs: A three-channel (RGB) image of 256 Ö 256 pixels with values normalized between [0,1]. The generator architecture consists of: 3. A Dense layer that expands the latent vector into a lower-resolution feature map. 4. Reshaping and upsampling layers to progressively increase the spatial resolution. 5. Several transposed convolutional layers (Conv2DTranspose), each followed by: (a) Batch normalization to improve stability. (b) LeakyReLU activation for non-linearity. 6. A final Conv2D layer with a sigmoid activation function, ensuring the output image values remain within the valid range.  GAN Model (Combining Generator & Discriminator) Once both the generator and discriminator are defined, they are combined to form a complete GAN model. The training process follows these steps: 1. The generator creates a batch of fake images from random latent vectors. 2. These fake images are passed to the discriminator, along with real images from the dataset. 3. The discriminator predicts whether each image is real or fake. 42 4. Backpropagation is applied, updating both the generator and discriminator weights to improve their respective performances. 5. This adversarial training continues until the generator produces highly realistic images that can fool the discriminator. By iteratively refining the generator and discriminator, the GAN model learns to generate increasingly convincing fake images, which are later used to train the deepfake detection models. A plot of the model is also created and we can see that the model expects a 100-element point in latent space as input and will predict a single output classification label. Figure 2: Plot of the Composite Generator and Discriminator model in the GAN 3.3 Process Development of DeiT (Data-efficient Image Transformer) The DeiT (Data-efficient Image Transformer) model is an optimized version of the Vision Transformer (ViT), designed for efficient training on smaller datasets. Unlike Convolutional Neural Networks (CNNs), which rely on convolutional layers to extract local features, DeiT leverages the self-attention mechanism to capture both local and global dependencies within an image. This characteristic enables it to recognize complex patterns and structural inconsistencies that may indicate deepfake manipulations. The development process of DeiT follows steps:  Linear Embedding Layer: -The input image is split into fixed-size patches (e.g., 16 Ö 16 pixels). -Each patch is flattened and mapped into a high-dimensional feature space through a learned embedding matrix. -A learnable classification token is added to the sequence, and positional encodings are introduced to preserve spatial relationships.  Transformer Encoder: -The sequence of image patches passes through L identical layers, each containing: -A Multi-Head Self-Attention (MSA) mechanism, which enables the model to analyze relationships between patches. -A Feed-Forward Network (MLP) that applies non-linear transformations for feature enhancement. -Layer Normalization (LN) and skip connections to stabilize training and improve information retention.  Multi-Head Self-Attention (MSA) Mechanism: -Computes attention between all patches, allowing the model to focus on important regions of the image. -Uses Query (Q), Key (K), and Value (V) matrices to determine the weight of each patch in the final representation. Classification and Output: -After passing through multiple transformer layers, the classification token is extracted. 43 -A fully connected layer is applied to classify the image as real or fake. DeiT offers a powerful alternative to CNNs for deepfake detection, particularly when dealing with large datasets. However, due to its reliance on large-scale training data, its performance can be impacted when applied to smaller datasets. In this study, CNN demonstrated higher accuracy on limited data, while DeiT showed better scalability and generalization potential for future deepfake detection improvements [19]. 4 Experiments and Results This section describes the implementation process and the experiments conducted to evaluate the proposed deepfake detection model. The implementation consists of dataset preparation, model training, performance evaluation, and final deployment. 1. Dataset The experimentation of the proposed technique is implemented by using the two datasets : The first one is the “140k real and fake faces” dataset contains 70k real faces from the Flickr dataset collected by Nvidia, as well as 70k fake faces sampled from 1 million fake faces (generated by style GAN) 1. The second is “Real and fake face detection” datasets contain two subfolders training real and training fake. Training real contains 1081 images and training fake contains 960 images, the total dataset is 2041 images 2 2. Parameter Settings The models were trained using the following hyperparameters: Table 1: PARAMETER SETTINGS Model Epochs Batch size Activation Optimizer CNN 20 64 Sigmoid Adam DeiT-Tiny 20 32 ReLU Adam 3. Performance Evaluation and Discution The performance evaluation of the CNN and DeiTTiny models was conducted using key metrics such as training accuracy, validation accuracy, training loss, and validation loss, as summarized in Table II. The results highlight that CNN outperformed DeiT-Tiny, especially on the smaller dataset, achieving 94.15% validation accuracy on the 140K dataset and 81.88% on the Real and Fake Face Detection dataset. In contrast, DeiT-Tiny reached 90.31% and 61.27%, respectively, indicating its difficulty in learning from limited data and reliance on larger training sets for optimal performance. Table 2: PERFORMANCE EVALUATION TABLE Model Dataset Train accuracy Validation accuracy Train loss Validation loss CNN 140K Real and Fake Faces 97.11 % 94.15% 7.47% 14.50% CNN Real and Fake Face Detection 94.58% 81.88% 22.72% 43.95% DeiT-Tiny 140K Real and Fake Faces 90.06% 90.31% 37.20% 37.01% DeiT-Tiny Real and Fake Face Detection 85.05% 61.27% 43.89% 68.44% Figure 4and 5illustrate the CNN model’s stability, with smooth loss curves and consistently high accuracy, confirming its reliability in deepfake detection. Figure 6and Figure 7depict the DeiT-Tiny model’s slower convergence and higher validation loss, suggesting greater training data requirements for stable results. Despite its generalization potential, DeiT-Tiny struggled with small datasets, whereas CNN demonstrated robust and reliable performance across both datasets. Overall, these findings confirm 1https://www.kaggle.com/datasets/ciplab/real-and-fake-face-detection 2https://www.kaggle.com/datasets/xhlulu/140k-real-and-fake-faces 44 Figure 3: Vision transformer architecture that CNN is better suited for real-world deepfake detection applications, particularly when dataset availability is limited. While DeiT-Tiny offers strong generalization, it requires extensive data and longer training times to match CNN’s performance, emphasizing the need for model selection based on dataset size and computational constraints. Figure 4: CNN Model Performance on Real and Fake Face Detection Dataset : (a) Loss function, (b) Accuracy Figure 5: CNN Model Performance on 140K Real and Fake Faces Dataset : (a) Loss function, (b) Accuracy 45 Figure 6: DeiT-Tiny Model Performance on Real and Fake Face Detection Dataset : (a) Loss function, (b) Accuracy Figure 7: DeiT-Tiny Model Performance on 140K Real and Fake Faces Dataset : (a) Loss function, (b) Accuracy 5 Conclusion Deepfake technology presents significant challenges in digital security and misinformation prevention, requiring robust detection mechanisms. This study explored CNN and DeiT-Tiny models for deepfake detection, demonstrating that CNN achieved higher accuracy and stability, especially on smaller datasets, while DeiT-Tiny required larger datasets for optimal performance. Despite advancements, deepfake detection remains complex due to adversarial attacks, dataset biases, and computational constraints. Future research should focus on real-time detection, improving model robustness, and integrating multimodal approaches such as audio and behavioral analysis. This study contributes to enhancing digital media security, emphasizing the need for continuous advancements in AI-driven detection frameworks to combat deepfake threats effectively. References [1] F. Abbas and A. Taeihagh. Unmasking deepfakes: A systematic review of deepfake detection and generation techniques using artificial intelligence. Expert Systems With Applications, 124260, 2024. [2] Z. Akhtar. Deepfakes generation and detection: a short survey. Journal of Imaging, 9(1):18, 2023. [3] A. Heidari et al. Deepfake detection using deep learning methods: A systematic and comprehensive review. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 14(2):e1520, 2024. [4] A. Kaushal et al. A review on deepfake generation and detection: bibliometric analysis. Multimedia Tools and Applications, pages 1–41, 2024. [5] B. Dolhansky et al. The deepfake detection challenge (dfdc) dataset. arXiv preprint, 2020. [6] C. Li et al. A continual deepfake detection benchmark: Dataset, methods, and essentials. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1339– 1349, 2023. 46 [7] D. Pan et al. Deepfake detection through deep learning. In 2020 IEEE/ACM International Conference on Big Data Computing, Applications and Technologies (BDCAT), pages 134–143, 2020. [8] F. Lugstein et al. Prnu-based deepfake detection. In Proceedings of the 2021 ACM Workshop on Information Hiding and Multimedia Security, pages 7–12, 2021. [9] H. Zhao et al. Multi-attentional deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2185–2194, 2021. [10] J. Wang et al. M2tr: Multi-modal multi-scale transformers for deepfake detection. In Proceedings of the 2022 International Conference on Multimedia Retrieval, pages 615–623, 2022. [11] K. Aashish et al. Latent flow diffusion for deepfake video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3781–3790, 2024. [12] K. Narayan et al. Deephy: On deepfake phylogeny. In 2022 IEEE International Joint Conference on Biometrics (IJCB). IEEE, 2022. [13] K. Narayan et al. Df-platter: Multi-face heterogeneous deepfake dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9739–9748, 2023. [14] M. Masood et al. Deepfakes generation and detection: State-of-the-art, open challenges, countermeasures, and way forward. Applied Intelligence, 53(4):3974–4026, 2023. [15] M. Rehaan et al. Face manipulated deepfake generation and recognition approaches: A survey. Smart Science, 12(1):53–73, 2024. [16] M. S. Rana et al. Deepfake detection: A systematic literature review. IEEE Access, 10:25494–25513, 2022. [17] O. De Lima et al. Deepfake detection using spatiotemporal convolutional networks. arXiv preprint, 2020. [18] S. R. Ahmed et al. Analysis survey on deepfake detection and recognition with convolutional neural networks. In 2022 International Congress on Human-Computer Interaction, Optimization and Robotic Applications (HORA), pages 1–7, 2022. [19] Y. Bazi et al. Vision transformers for remote sensing image classification. Remote Sensing, 13(3):516, 2021. [20] Z. Yan et al. Df40: Toward next-generation deepfake detection. arXiv preprint, 2024. [21] T. Zhang. Deepfake generation and detection, a survey. Multimedia Tools and Applications, 81(5):6259–6276, 2022. [22] Y. Zhou and S. N. Lim. Joint audio-visual deepfake detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021. 47 Review on deep learning optimization using knowledge and dataset distillation in medical imaging diagnostics Nor-Elhouda Laribi1, Djamel Gaceb2, Abdellah Rezoug3, and Fay¸cal Touazi4 1LIMOSE Laboratory University M’Hamed Bougara Boumerdes, Algeria, [email protected] 2LIMOSE Laboratory University M’Hamed Bougara Boumerdes, Algeria, [email protected] 3LIMOSE Laboratory University M’Hamed Bougara Boumerdes, Algeria, [email protected] 4LIMOSE Laboratory University M’Hamed Bougara Boumerdes, Algeria, [email protected] Abstract The integration of deep learning-based artificial intelligence solutions in hospital environments introduces significant challenges, including data privacy restrictions, limited computational resources, and constraints related to the quality and simplicity of the models used. In this review, we highlight the recent advancements in knowledge distillation and dataset distillation as emerging solutions to these challenges in the field of medical imaging. These techniques offer practical benefits in clinical settings by enabling faster training, reduced model size, improved inference speed, and enhanced accuracy, while supporting privacy-preserving learning across decentralized systems and edge devices. Knowledge distillation transfers knowledge from a complex to a simple model, enabling efficient deployment without high loss in diagnostic performance. Dataset distillation, by contrast, focuses on synthesizing datasets that match the pretrained model on real data, reducing data storage requirements. Together, these methods improve learning efficiency, model accuracy, and resource optimization in hospital workflows. However, their integration into medical environments also presents limitations. Challenges such as pipeline complexity, scalability issues, and performance inconsistency across architectures or high-resolution tasks still persist. Overall, this review provides a comprehensive overview of potential and limitations of these two types of distillations in healthcare, offering insights into how these methods can support more scalable, accurate, and privacy-aware AI solutions for medical imaging. Keywords: Healthcare, medical imaging, deep learning, knowledge distillation, dataset distillation, data privacy. 1 Introduction Artificial intelligence (AI), and deep learning in particular, has become increasingly crucial in medical imaging, offering significant improvements in diagnostic accuracy, efficiency, and decision support systems. From radiology to pathology, deep learning models have demonstrated capabilities that rival or exceed human experts in specific tasks. However, deploying these powerful models in real-world hospital settings presents significant challenges. The clinical environment presents challenges, including stringent data privacy, limited computational resources in edge settings, and the practical need for real-time or near-real-time inference. These constraints pose significant challenges to the adoption of conventional, large-scale deep learning models, which are often data-intensive, resource-requirement. To address these limitations, two emerging strategies have gained traction in the research community: Knowledge Distillation (KD) and Dataset Distillation (DD). These techniques aim to retain the performance benefits of deep learning while reducing the computational and data demands that often hinder clinical deployment. Knowledge distillation works by transferring knowledge from a large, complex model (the ”teacher”) to a smaller, more efficient one (the ”student”), preserving accuracy while improving speed and reducing resource usage. Dataset distillation, on the other hand, generates compact synthetic datasets that can replicate the behavior of real data, reducing storage needs and enabling faster training cycles and deployment settings. In this review, we provide a comprehensive overview of both KD and DD techniques in the context of medical imaging, with a particular focus on their application to brain MRI-based disease 48 performance. For example, the work in [6] uses KD through logit matching and feature map analysis, allowing a custom 5-layer CNN to approach the accuracy of a DenseNet121 teacher (99.38% vs. 99.46%) while requiring nearly ten times fewer operations, demonstrating how lightweight models can achieve comparable results. In [15], knowledge transfer from a Res-Transformer to a lightweight ResU-Net student results in a 7.2% accuracy increase and an estimated 5–10×reduction in computational complexity, emphasizing the benefits of distilling both soft labels and intermediate features. Taking a novel direction, study [27] explores the integration of KD with quantum neural networks, where a TinyViT teacher distills knowledge into a QViT student model. This approach achieves up to 80×compression while maintaining strong classification performance (AUC up to 0.812), highlighting KD’s potential in hybrid quantum-classical learning systems. Similarly, KD can go beyond modelto-model transfer for the same task. In [18], a dual-stream KD strategy is employed to bridge two distinct tasks—segmentation and classification—by distilling structural knowledge from a segmentation model into a DS-ViT classifier. This cross-task distillation leads to notable improvements in classification accuracy (0.899) and recall (0.917), while also reducing model size by approximately 5×, showcasing KD’s ability to enable transfer learning across functionally different but related domains. Another approach, proposed in [29], introduces a distillation token mechanism that transfers knowledge from a large 3D ResNet-152 to a significantly smaller transformer-based student. Despite achieving 9.7×compression, the performance gain is modest ( 0.1%) and computational costs remain high due to the heavy teacher model. These studies illustrate the broad utility of KD—not only as a model compression method, but also as a bridge across architectures, learning paradigms, and deployment constraints—especially within the sensitive and resource-limited field of medical imaging. Inspired by the principles of KD, a related and increasingly powerful technique known as dataset distillation (DD) has emerged. Unlike KD, which transfers knowledge from one model to another, DD transfers knowledge from real data—or a model trained on it—into a much smaller, synthetic dataset. In this paradigm, the teacher is not a model but the original data distribution itself, and the student (often with the same architecture as the teacher) is trained exclusively on the distilled dataset. This method aims to match the performance of models trained on full datasets using only a synthesized subset derived from the original data. Recent studies have further developed this concept by evaluating multiple model architectures during the distillation process, thereby improving the generalizability of the distilled data across tasks and learners. Among its key advantages, DD significantly accelerates training times and reduces storage requirements, making it especially effective for simpler datasets such as CIFAR-10 and GTSRB. For instance, studies show that with just 10 to 50 images per class (IPC), models trained on distilled CIFAR-10 can achieve up to 72.6% accuracy compared to 84.8% on the full dataset, while GTSRB achieves 65.38% using IPC 10 versus 66.55% with full data. This makes DD ideal for edge deployment or federated learning setups with strict data transmission and storage limits. Additionally, DD can deliver notable speedups (e.g., 4.5×on ImageNet subsets) and even occasional performance improvements over traditional distributed training. However, these advantages come with significant limitations. As dataset complexity increases (e.g., CIFAR-100, ImageNet-100, ImageNet-1K), the performance gap between models trained on distilled versus full datasets becomes substantial. For example, on ImageNet-1K with 10 IPC, accuracy drops to 30.5%, although increasing to 100 IPC improves it to 70.0%. These datasets contain high inter-class variability and complex patterns that are difficult to capture with limited synthetic data. Moreover, the success of DD is highly architecture-dependent: while simple networks like AlexNet and VGG retain relatively high AUC and accuracy on distilled data, more complex models like ResNet and ViT are more sensitive to the quality and diversity of synthetic samples. To mitigate this, some methods introduce teacher-student mechanisms during DD, improving results at the cost of greater pipeline complexity. DD has shown great promise in medical imaging, achieving significant data reduction while preserving performance, particularly with ConvNet and ResNet architectures. Studies indicate that even under extremely low IPC settings, models can maintain strong performance on tasks such as CT scan segmentation and chest X-ray classification. However, a major challenge to generalizability remains: distilled data often overfit to the model architecture they were generated for, reducing reusability across other architectures. Furthermore, on more diverse or clinically detailed datasets, performance tends to decline, and important but rare features may be lost. Most research has focused on convolutional architectures, limiting exploration of modern transformer-based models. Additionally, the interpretability and clinical reliability of synthetic data remain open concerns in real-world medical applications. 55 5 Conclusion Knowledge and dataset distillation are emerging as impactful techniques in medical imaging, particularly in hospital settings where data privacy, sharing restrictions, and limited computational resources are key concerns. KD enables the compression of large models into efficient, high-performing versions, making it ideal for decentralized systems, edge devices, and federated learning environments where raw data cannot be shared. DD complements this by creating compact synthetic datasets that capture the essential patterns of full datasets, enabling training without exposing sensitive medical data. These approaches support privacy-preserving and resource-efficient AI deployment but face notable challenges. KD often involves complex training pipelines with teacher-student models, while DD may struggle to retain diagnostic fidelity in high-resolution or fine-grained tasks. Both methods also show performance variability across different architectures, highlighting the need for careful tuning and optimization. Looking forward, integrating generative adversarial networks (GANs) into dataset distillation offers a promising direction. GANs can improve the realism and diversity of synthetic data, helping to close the performance gap between distilled and full datasets. With continued research, KD and DD have strong potential to support scalable, accurate, and privacy-aware AI systems in clinical environments. References [1] R. Anantathanavit, F. H. Raswa, T. Thaipisutikul, and J. C. Wang. Lightweight brain tumor diagnosis via knowledge distillation. In 2024 International Conference on Multimedia Analysis and Pattern Recognition (MAPR), pages 1–6. IEEE, August 2024. [2] M. Arazzi, M. Cihangiroglu, S. Nicolazzo, and A. Nocera. Secure federated data distillation. arXiv preprint, 2025. [3] T. Boucher and E. B. Mazomenos. Distilling knowledge into quantum vision transformers for biomedical image classification. arXiv preprint, 2025. [4] K. Chen, Y. Wang, Y. Zhou, and H. Wang. Ds-vit: Dual-stream vision transformer for cross-task distillation in alzheimer’s early diagnosis. arXiv preprint, 2024. [5] O. S. EL-Assiouti, G. Hamed, D. Khattab, and H. M. Ebied. Hdkd: Hybrid data-efficient knowledge distillation network for medical image classification. Engineering Applications of Artificial Intelligence, 138:109430, 2024. [6] G. J. Ferdous, K. A. Sathi, M. A. Hossain, M. M. Hoque, and M. A. A. Dewan. Lcdeit: A linear complexity data-efficient image transformer for mri brain tumor classification. IEEE Access, 11:20337–20350, 2023. [7] R. J. Gohari, L. Aliahmadipour, and E. Valipour. Fedbrain-distill: Communication-efficient federated brain tumor classification using ensemble knowledge distillation on non-iid data. arXiv preprint, 2024. [8] Y. Jiang, X. Zhao, Y. Wu, and A. Chaddad. A knowledge distillation-based approach to enhance transparency of classifier models. arXiv preprint, 2025. [9] R. Kanagavelu, M. Walia, Y. Wang, H. Fu, Q. Wei, Y. Liu, and R. S. M. Goh. Medsynth: Leveraging generative model for healthcare data sharing. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 654–664, Cham, October 2024. Springer Nature Switzerland. [10] K. Kunanbayev, V. Shen, and D. S. Kim. Training vit with limited data for alzheimer’s disease classification: An empirical study. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 334–343, Cham, October 2024. Springer Nature Switzerland. [11] S. Lei and D. Tao. A comprehensive survey of dataset distillation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(1):17–32, 2023. [12] G. Li, R. Togo, T. Ogawa, and M. Haseyama. Soft-label anonymous gastric x-ray image distillation. In 2020 IEEE International Conference on Image Processing (ICIP), pages 305–309. IEEE, October 2020. 56 [13] G. Li, R. Togo, T. Ogawa, and M. Haseyama. Compressed gastric image generation based on softlabel dataset distillation for medical data sharing. Computer Methods and Programs in Biomedicine, 227:107189, 2022. [14] G. Li, R. Togo, T. Ogawa, and M. Haseyama. Dataset distillation for medical dataset sharing. arXiv preprint, 2022. [15] G. Li, R. Togo, T. Ogawa, and M. Haseyama. Importance-aware adaptive dataset distillation. Neural Networks, 172:106154, 2024. [16] J. Li, L. Zhang, K. Zhong, and G. Qian. A discrepancy-aware self-distillation method for multimodal glioma grading. Knowledge-Based Systems, 295:111858, 2024. [17] Y. Li, J. Luo, and J. Zhang. Classification of alzheimer’s disease in mri images using knowledge distillation framework: an investigation. International Journal of Computer Assisted Radiology and Surgery, 17(7):1235–1243, 2022. [18] S. Ma, F. Zhu, Z. Cheng, and X. Y. Zhang. Towards trustworthy dataset distillation. Pattern Recognition, 157:110875, 2025. [19] N. Matcha, S. Ramanarayanan, M. Al Fahim, R. GS, K. Ram, and M. Sivaprakasam. Sft-kdrecon: Learning a student-friendly teacher for knowledge distillation in magnetic resonance image reconstruction. In Medical Imaging with Deep Learning, pages 1423–1440. PMLR, January 2024. [20] S. Nabavi, K. A. Hamedani, M. E. Moghaddam, A. A. Abin, and A. F. Frangi. Multiple teachersmeticulous student: A domain adaptive meta-knowledge distillation model for medical image classification. arXiv preprint, 2024. [21] T. Qin, Z. Deng, and D. Alvarez-Melis. Distributional dataset distillation with subtask decomposition. arXiv preprint, 2024. [22] Y. Song, J. Wang, Y. Ge, L. Li, J. Guo, Q. Dong, and Z. Liao. Medical image classification: knowledge transfer via residual u-net and vision transformer-based teacher-student model with knowledge distillation. Journal of Visual Communication and Image Representation, 102:104212, 2024. [23] S. Tan, Y. Cai, Y. Zhao, J. Hu, Y. Chen, and C. He. FM-LiteLearn: A lightweight brain tumor classification framework integrating image fusion and multi-teacher distillation strategies. In International Conference on AI in Healthcare, pages 89–103, Cham, August 2024. Springer Nature Switzerland. [24] B. Wu, D. Shi, and J. Aguilar. Brain tumors classification in mris based on personalized federated distillation learning with similarity-preserving. International Journal of Imaging Systems and Technology, 35(2):e70046, 2025. [25] R. Yang, Y. Chen, Z. Zhang, X. Liu, Z. Li, K. He, and Q. Dai. Unicompress: Enhancing multi-data medical image compression with knowledge distillation. arXiv preprint, 2024. [26] Y. Yang, X. Guo, C. Ye, Y. Xiang, and T. Ma. Creg-kd: Model refinement via confidence regularized knowledge distillation for brain imaging. Medical Image Analysis, 89:102916, 2023. [27] R. Yu, S. Liu, and X. Wang. Dataset distillation: A comprehensive review. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(1):150–170, 2023. [28] Z. Yu, Y. Liu, and Q. Chen. Progressive trajectory matching for medical dataset distillation. arXiv preprint, 2024. [29] L. Zhang, J. Zhang, B. Lei, S. Mukherjee, X. Pan, B. Zhao, and D. Xu. Accelerating dataset distillation via model augmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11950–11959, 2023. [30] X. Zhong, S. Sun, X. Gu, Z. Xu, Y. Wang, M. Zhang, and B. Chen. Efficient dataset distillation via diffusion-driven patch selection for improved generalization. arXiv preprint, 2024. 57 Part II Deep Learning and Data Processing Applications RL-Guided Pruning of CNNs Using Graph Embeddings Karima Amrouche11, Yacine Ait Ali Yahia21, and Ilhem Kherroubi31 1Laboratoire de la Communication dans les Syst`emes Informatiques (LCSI), ´ Ecole Nationale Sup´erieure d’Informatique, BP 68M, 16309, Alger, Alg´erie Abstract This paper presents a novel method for compressing Convolutional Neural Networks (CNNs) to enable efficient deployment on low-capacity devices. The proposed approach combines neural network pruning with reinforcement learning (RL) and graph embedding. Each network is represented as a computational graph, and Graph Convolutional Networks (GCNs) are utilized to learn graph-level embeddings that inform pruning decisions. By applying Proximal Policy Optimization (PPO), we automate the selection of layer-wise pruning ratios, eliminating the need for manual tuning. Experiments on ResNet-34 and VGG-19, trained on the CIFAR-10 dataset, demonstrate that our method achieves up to 80% compression while maintaining or improving model accuracy through post-pruning rewinding. We evaluated both structured and unstructured pruning strategies, analyzing the trade-offs in accuracy, FLOPs, parameter count, and inference time. Keywords: Model Compression, Deep Neural Networks, Graph Embedding, Reinforcement Learning, Neural Networks, Pruning, Convolutional Neural Networks, CNN Acceleration. 1 Introduction The deployment of convolutional neural networks (CNN) in low-capacity devices, such as smartphones and IoT devices [1], presents significant challenges due to their high computational demands and large memory requirements. These constraints limit the use of CNNs in real-time applications, such as facial recognition or object detection, especially in environments where cloud resources are unavailable or introduce unacceptable latency. Model compression techniques, particularly neural network pruning, have effectively overcome these challenges. By reducing the complexity of CNNs, pruning reduces computational cost and memory usage, enabling deployment on resource-constrained devices without significant performance loss. However, determining the optimal pruning strategy for each layer remains challenging, as it often requires manual tuning and is sensitive to the network structure. We introduce a method that combines neural network pruning with reinforcement learning (RL) and graph embeddings to automate the compression process. This approach reduces the need for manual tuning while preserving accuracy and efficiency, key requirements for deploying CNNs on low-capacity devices. 2 Related Work Model compression techniques can be categorized into several approaches, including knowledge distillation, quantization, factorization, and pruning [2]. Among these, pruning offers a direct and effective way to reduce model complexity by eliminating redundant parameters, often with minimal loss in accuracy. It is particularly attractive because it can be applied post-training, preserves the original architecture, and complements other techniques such as quantization. The theoretical basis for pruning is strengthened by the Lottery Ticket Hypothesis (LTH) [3], which proposes that overparameterized networks contain smaller, trainable subnetworks—referred to as “winning tickets”—capable of reaching comparable accuracy to the full model when trained in isolation. As Frankle and Carbin describe: 59 “A randomly-initialized, dense neural network contains a subnetwork that is initialized such that—when trained in isolation—it can match the test accuracy of the original network after training for at most the same number of iterations.” [3] This insight highlights the inherent redundancy in large neural networks and motivates pruning as a principled strategy for model compression. However, identifying such subnetworks remains nontrivial, especially in deeper architectures, where brute-force or heuristic pruning methods become computationally prohibitive. This challenge underscores the need for more scalable and intelligent pruning approaches—such as those guided by reinforcement learning and graph-based representations. In practice, pruning techniques are generally classified into two categories: unstructured and structured. Unstructured pruning removes individual weights, often resulting in sparse models with high compression rates. However, these irregular patterns typically require specialized hardware for efficient execution. Structured pruning, in contrast, removes entire filters, channels, or blocks, producing smaller, dense models that are more compatible with conventional hardware and easier to deploy [4,5]. Despite their practical advantages, traditional pruning methods often rely on fixed heuristics to determine layer-wise pruning ratios, which may not generalize well across architectures or datasets. To overcome these limitations, recent work has turned to reinforcement learning (RL) to automate the pruning process. Notable examples include ABCPruner [6] and CCPruner [7], which use RL agents to learn pruning policies that balance model efficiency and accuracy. While these methods have shown promise, they often treat layers independently and fail to account for the structural dependencies across the network. Our work addresses this gap by modeling the neural network as a computational graph, enabling a more holistic view of the architecture. By leveraging graph embeddings, we capture global structural information that informs pruning decisions, leading to more coherent and scalable model compression. 3 Building a Computational Graph from a Neural Network As illustrated in Figure 1, our compression pipeline consists of the following stages: 1. Initialization: The process begins with the initialization of a deep neural network, either from scratch or using a pretrained model. At this stage, the initial weight values are preserved to enable potential rewinding after pruning, as part of the iterative compression strategy. 2. Training: The initialized model is trained on the target task until it achieves satisfactory performance. This yields a fully trained network that serves as the baseline for subsequent pruning. 3. Pruning: Reinforcement learning is employed to determine the optimal pruning rates for each layer. A Graph Convolutional Network (GCN) is used to encode the computational graph of the trained model into a latent state representation. This representation is then processed by a policy network, which generates pruning decisions. These decisions are evaluated based on a reward signal reflecting the trade-off between compression and model accuracy. 4. Rewinding: After pruning, the model is reverted to its initial weights saved during the initialization phase. This step, known as rewinding, facilitates the identification of subnetworks—often referred to as “lottery tickets”—that can be retrained from the original initialization to achieve strong performance. 5. Retraining: The selected subnetwork is retrained from its original initialization. The objective is to recover the model’s performance to a level comparable to the original unpruned network, thereby achieving an efficient and accurate compressed model. 60 Figure 1: Pipeline illustrating CNN pruning guided by reinforcement learning, using graph embeddings to encode architectural information. The model automates the pruning phase, removing the need for human input. A DNN’s computational graph maps operations like addition and multiplication during this phase to produce outputs. Nodes perform computations, while edges guide the data flow through the network’s layers [8]. Algorithm 1outlines the procedure for constructing a subgraph for a single layer, which involves initializing input and output nodes and creating edges to represent the connections between them. Algorithm 1: Subgraph Construction for a CNN Layer Data: n: Number of input channels Input: N: Number of output channels E: List of edges (empty for the first layer) Result: E: Updated list of edges Output: Input node of the next layer 1Output ←Input +N+ 1; 2for i←1to Ndo 3E←Insert(E, (Input, Input +i)) ; // Insert edge eik [input] 4E←Insert(E, (Input +i, Output)) ; // Insert edge eik [output] 4 Graph Embedding Graph embedding transforms structured data into low-dimensional vector representations, facilitating efficient learning and decision-making in downstream tasks. In this work, we represent convolutional neural networks (CNNs) as computational graphs, where nodes correspond to layers or operations and edges capture the flow of information. This graph-based formulation enables the use of Graph Convolutional Networks (GCNs) to encode the topological and functional properties of the CNN architecture. By leveraging GCNs, we obtain a fixed-size embedding of the network that preserves both structural dependencies and feature hierarchies, which are essential for informing pruning decisions [9]. This compact representation is then passed to the reinforcement learning (RL) agent, which uses it to select optimal pruning strategies. The use of GCNs ensures that similar CNN architectures yield similar embeddings, improving the generalization of the pruning policy across different models. As input to the GCN, we define the node features using the number of nodes (|V|) and their attributes (Fin), along with edge indices (2,|E|), and output node features (|V|, Fout). Instead of focusing solely on individual node embeddings, we aim to compute a holistic representation of the entire graph. To achieve this, we first apply a GCN encoder that maps the graph Gto a set of node embeddings H∈RN×d, as shown in Equation 1: 61 H= GCNencoder(G)∈RN×d(1) The resulting node embeddings are then aggregated using a global pooling operation. Specifically, we use a GlobalMeanPool that computes the average over all node embeddings, producing the final graph-level embedding g, as defined in Equation 2: g=1 N N X n=1 hi(2) In Equation 2,hidenotes the embedding of the i-th node, Nis the total number of nodes in the graph, and dis the dimensionality of the embedding space. This aggregated representation gcaptures the structural and semantic properties of the CNN architecture, and is used as input to the reinforcement learning agent. 5 Criteria for Choosing the Compression Ratio Our goal is to optimize convolutional neural networks (CNNs) for deployment on low-capacity devices by applying structured pruning techniques. This reduces inference time, memory usage, and model size. To guide pruning, we focus on two key efficiency criteria: •Model Parameters: Reducing the number of parameters directly decreases the computational and memory overhead, enhancing the efficiency of the model [10]. •FLOPs (Floating Point Operations): FLOPs quantify the computational effort required per inference and are especially relevant in CNNs due to weight sharing [11]. Minimizing FLOPs has a direct impact on inference speed and energy consumption. FLOPs are calculated per layer type as follows: •Convolutional Layers: FLOPs = 2 ×CO×CI×K×O(3) where COis the number of output channels, CIis the number of input channels, Kis the kernel size, and Ois the number of output elements. •Fully Connected Layers: FLOPs = 2 ×I×O(4) where Iand Oare the number of input and output units, respectively. 6 Pruning Methods Selection To enable the deployment of convolutional neural networks (CNNs) on low-capacity devices, it is essential to reduce their inference time, memory footprint, and overall model size. Pruning—i.e., removing less important components of the network—is a widely used approach to achieve such compression. In this work, we evaluate both unstructured and structured pruning techniques for their effectiveness in compressing CNN architectures. •Unstructured Pruning: This method eliminates individual weights from convolutional kernels based on their magnitude or contribution. While it can result in highly sparse models with minimal impact on accuracy, it often fails to yield practical improvements in inference time or memory usage. This is largely due to the lack of hardware-level support for irregular sparsity, which limits the efficiency gains on general-purpose devices. •Structured Pruning: In contrast, structured pruning removes entire filters, channels, or even layers. This produces a more compact and regular architecture, leading to measurable reductions in computation and memory requirements. However, structured pruning carries a higher risk of accuracy degradation if critical components are pruned without adequate guidance. 62 In addition to the pruning strategy itself, the choice of how to reinitialize and retrain the pruned network plays a crucial role in recovering or maintaining performance [12]. We evaluate the following post-pruning training approaches: •Rewinding (100%): After pruning, the remaining weights are reset to their initial values recorded at the start of training. This approach is motivated by the Lottery Ticket Hypothesis, which suggests that certain subnetworks can achieve competitive performance when trained from their original initialization. •Random Initialization: As a baseline, we reinitialize the surviving weights with new random values after pruning. This serves to evaluate whether rewinding provides a significant advantage over fresh initialization. •Fine-Tuning: This method retains the final values of the remaining weights and continues training with a reduced learning rate. The goal is to refine the pruned model without substantially altering its learned representations. Fine-tuning is commonly used in practice due to its simplicity and effectiveness. 7 Implementation of Reinforcement Learning To automate the pruning process, we employ reinforcement learning (RL) with the Proximal Policy Optimization (PPO) algorithm. The RL agent is trained to predict optimal pruning ratios for each layer of a convolutional neural network, with the objective of maximizing compression while preserving classification accuracy. Experiments were conducted on the VGG-19 and ResNet-34 architectures using the CIFAR-10 dataset, with a global compression target of 80%. The agent receives as input a graph-level embedding of the CNN architecture, obtained via a Graph Convolutional Network (GCN) encoder. Based on this representation, the agent generates pruning decisions across layers. Following pruning, we apply several retraining strategies—including random initialization, weight rewinding, and fine-tuning—to restore or improve performance. This framework enables a systematic exploration of the trade-off between model compactness and accuracy, leading to efficient CNNs suitable for deployment on resource-constrained devices. 7.1 Proximal Policy Optimization (PPO) Proximal Policy Optimization (PPO) is a widely used reinforcement learning algorithm designed to improve the policy—the strategy for selecting actions—while maintaining stability and efficiency. It achieves this by carefully balancing exploration (trying new actions) and exploitation (choosing the bestknown actions) [13]. PPO optimizes a clipped surrogate objective that limits abrupt changes to the policy and incorporates additional terms to enhance learning. The total objective function consists of three main components: •Policy Surrogate Loss: This term encourages beneficial updates to the policy based on the advantage of actions taken. It uses a ratio of the new and old policy probabilities: rt(θ) = πθ(at|st) πθold(at|st)(5) To prevent large, destabilizing policy updates, PPO applies a clipping mechanism: L1(θ) = min rt(θ)ˆ At,clip(rt(θ),1−ϵ, 1 + ϵ)ˆ At(6) •Value Function Loss: This component minimizes the error between the predicted value of a state and its target return, computed using the mean squared error: L2(θ) = (Vθ(st)−Vtarg,t)2(7) •Entropy Bonus: To promote exploration and prevent premature convergence to deterministic policies, PPO adds an entropy regularization term: L3(θ) = S[πθ](st) (8) 63 The overall loss function used to update the policy parameters is a weighted sum of the three components: LTotal(θ) = ˆ Et[L1+c1L2+c2L3] (9) where c1and c2are coefficients that balance the contributions of the value loss and entropy bonus, and ˆ Et[·] denotes the empirical average over a finite batch of experiences. Exploration Noise: During exploration, the agent samples actions from a Gaussian policy. The probability density function of a Gaussian distribution is given by: f(x) = 1 σ√2πexp −1 2x−µ σ2!(10) Initially, a fixed standard deviation σcontrols the amount of randomness in action selection. Over time, this noise is gradually reduced to encourage exploitation of learned policies as training progresses. 7.2 Memory and Experience Replay To enhance learning efficiency and generalization, the PPO agent maintains a memory buffer that stores past interactions with the environment, including states, actions, action probabilities, and rewards. Rather than updating the policy network using only the most recent data, the agent samples minibatches of past experiences. This experience replay strategy reduces overfitting and provides a more diverse set of training samples, enabling the agent to learn from a broader distribution of environment interactions. 7.3 Reinforcement Learning Environment The reinforcement learning (RL) environment models the deep neural network as a computational graph in a simulated setting that reflects structural changes during pruning. The environment provides the RL agent with the graph representation based on the CNN’s topology and dynamically tracks key metrics, including the number of parameters and FLOPs. The pruning process proceeds step by step until the agent satisfies the compression constraints, at which point the search is terminated. This mimics the RL paradigm, where the environment evolves with each action and ends an episode once a terminal condition—such as reaching a pruning goal—is met. 7.4 Timestamps and Agent Interaction The RL agent interacts with the environment incrementally to prune the network towards a target compression ratio. At each time step, relevant model attributes are updated, including pruning ratios and input/output channel sizes. Upon reaching the compression goal, the pruned model is evaluated in terms of classification accuracy, and a reward is assigned. If the agent fails to meet the target FLOPs within the episode, it is penalized accordingly. To maintain meaningful compression while avoiding excessive information loss, pruning ratios for each layer are constrained within the range [0.02, 0.9]. 8 Experiments and Results Following the formulation of our approach, we conducted experiments on the VGG-19 and ResNet-34 architectures using the CIFAR-10 dataset. The implementation was carried out in Python with support from various libraries: NumPy for numerical computations, Matplotlib for visualizations, and PyTorch and PyTorch Geometric for deep learning and graph-based modeling, respectively. Torchvision was used for computer vision tasks, while Weights Biases (W&B) provided experiment tracking. Execution and collaboration were facilitated via Google Colab, Amazon EC2, and Google Drive. The experimental pipeline focused on evaluating pruning capabilities through four stages: initial model training, reinforcement learning-based pruning (targeting 80% compression), application of postpruning retraining methods (random initialization, rewinding, fine-tuning), and final evaluation of model performance across multiple metrics including accuracy, parameter count, FLOPs, and model size. 64 On the other hand, we have S0(u)(ψ) = lim t→0S(u+tψ)−S(u) t, =ZZΩ×Ω v(x)u(y) v(y)−u(x) v(x)ψ(y) v(y)−ψ(x) v(x)ω(x, y) dydx, =ZZΩ×Ω v(x)u(y) v(y)−u(x) v(x)ψ(y) v(y)ω(x, y) dydx −ZZΩ×Ωu(y) v(y)−u(x) v(x)ψ(x)ω(x, y) dydx, =ZZΩ×Ω v(y)u(x) v(x)−u(y) v(y)ψ(x) v(x)ω(y, x) dxdy −ZZΩ×Ωu(y) v(y)−u(x) v(x)ψ(x)ω(x, y) dydx, =ZΩ−1 v(x)ZΩ v(x)u(y) v(y)−u(x) v(x)w(x, y) dyψ(x) dx +ZΩ1 v(x)ZΩv(y)u(x) v(x)−u(y) v(y)w(x, y) dyψ(x) dx, =ZΩ−1 v(x)divNL v∇NL u v(x)ψ(x) dx. Thus, we obtain: E0(u) = −divNL v∇NL u v−λα(f−u). The Euler-Lagrange equation E0(u)=0is difficult to solve. Consequently, we use the suboptimal gradient descent procedure ∂tu(t, x) = −E0(u). Thus, we obtain the system (5). 3 Numerical resolution Let Ωdbe a discrete grid associated to the continuous domain Ω: Ωd={1, . . . , M}×{1, . . . , N}. We use the following notations: i= (i1, i2)∈Ωd,xi= (xi1, yi2)∈Ω, ui≈u(xi), αi≈α(xi). The image fis approximated by a discrete matrix fd. The weight function ωdefines a similarity between two pixels xiand xjby comparing their neighborhoods: ωi,j=ωxi,xj= exp −1 h2 r X t1,t2=−r Gσ(t1, t2)f(xi+t)−f(xj+t)2!,(6) where rdefines the neighborhood, t= (t1, t2), and h > 0 is a scale parameter and Gσis a gaussian function (Gσ(t1, t2) = 1 2πσ2e−t2 1+t2 2 2σ2, with σ > 0). The approximation of the nonlocal Laplacian is given by: ∆NLu(xi)≈(∆NLu)i=X j∈Ωd (uj−ui)ωi,j. The nonlocal divergence of p: Ω ×Ω→Ris approximated by divNL(p) (xi)≈(divNL(p))i=X j∈Ωd (pi,j−pj,i)√ωi,j. 71 We approximate space derivative by forward differences and time derivative using explicit Euler method. The discrete iterative scheme of (5) writes: un+1 i=un i+dt dx  (∆NLun)i+X j∈Ωd vi vj un jωi,j−X j∈Ωd vj vi un iωi,j  +dt λαi(fi−un i) where dt,dx are the discretization steps, nis the time index. Then, the discrete problem can be written as algebraic equations of the form Un+1 =AUn+PF, where U∈Rm,m=M×N, is the vector obtained by concatenating the columns of the discrete image u,nis the time index, F= (fi)1≤i≤mand the initial condition is set to U0=F. 4 Experimental results We conducted extensive experiments to evaluate our model and compare it to previous state-of-the-art techniques: Poisson Image Editing [2], Parisotto et al. [4]. We implemented our algorithm using Matlab 2020a, but the weight function had been implemented as a c-mex file (C-program). Simulations have been conducted on a Core i5-4210 (2.60 GHz) processor with 4GO RAM. Based on an empirical analysis, we determined the values of the parameter λand found that λ∈[0,1]. For the weight function, we choose, in most cases, a patch size of 5 ×5, search window of size 7 ×7, σ= 0.003 and the filtering parameter h=σ. We used the software and test data provided by the authors 2,3,4. In some approaches, the authors presented an automatic method for selecting region. However, in this work, we assume that the subdomain is given to focus on the numerical solution of the nonlocal equation . Let’s montion, for the images that have not the same size, we create a new image foreground with the same size of background and with the centroid of selected region in the position with othor coordinates. The function αallows to indicat from which of the two images the structural information comes from. In the case where Ωsb =∅,αis binar, i.e. α(x)∈ {0,1}, so αvanishes on the pixels of the background image band is equal to one otherwise. Alternatively, αcan be smoothed by means of a Gaussian convolution so as to favour a smooth transition on Ωsb 6=∅. Experimental results obtained by the nonlocal model are presented on Figure 2(qualitative evaluation). As demonstrated visually, our method outperforms previous state-of-the-art techniques in most cases. It produces a better skin tone than the seamless Poisson editing model and Parisotto et al. [4] model as showing in Figure 2(a) and Figure 2(b). It is a powerful tool for manipulating colors, two differently colored versions of those images can be mixed seamlessly. For exemple, in Figure 2(d), our model overcome the problem of the undesirable visible seam of Parisotto et al. [4] model. An example is shown in Figure 2(e), in which our model can facilitate the transfer of partly transparent objects (the rainbow). 4.1 Comparison with deep learning techniques We also assess the performance of the proposed image fusion method using the Lytro5dataset.The fusion results of the proposed method are compared with three recently developed fusion algorithms. The first one is GFDF [5]. The second compared algorithm is DCT EOL [1]. The third method is CNN [6]. All image fusion results6of these algorithms are available online. The results are presented on Figure 3. 3https://doi.org/10.5201/ipol.2016.163 4http://cs.brown.edu/courses/cs129/results/proj2/damoreno/ 5https://sandipanweb.wordpress.com/2017/10/03/some-variational-image-processing-possion-image-editing-and-itsapplications/ 5https://www.researchgate.net/publication/291522937 Lytro Multi-focus Image Dataset 6https://github.com/xingchenzhang/MFIF 72 Figure 2: Image fusion results of different models ((1)-(8)). From left to right: background, selected region, foreground, Seamless Poisson Editing [2], Parisotto et al. [4], nonlocal osmosis model (1). Figure 3: Qualitative comparison of results of Lytro dataset. (1) and (2) are: Foreground and background. From (3) to (6) are fused images obtained by: GFDF [5], DCT EOL [1], CNN [6], nonlocal osmosis model (1). 73 5 Conclusion In this paper, we presented a novel nonlocal model for image fusion that utilizes nonlocal derivatives in the recently developed osmosis model. Experimental analyses have shown that our proposed model demonstrates its effectiveness and superiority over local image fusion models. Our proposed method provides visually plausible image data fusion that is invariant to multiplicative brightnessb changes. References [1] M. Amin-Naji and A. Aghagolzadeh. Multi-focus image fusion in dct domain using variance and energy of laplacian and correlation coefficient for visual sensor networks. Journal of AI and Data Mining,, pages 33–250, 2018. [2] J. Matas Di Martino, Gabriele Facciolo, and Enric Meinhardt-Llopis. Poisson Image Editing. Image Processing On Line, pages 300–325, 2016. [3] Guy Gilboa and Stanley Osher. Nonlocal operators with applications to image processing. Multiscale Modeling & Simulation, 7(3):1005–1028, 2009. [4] Simone Parisotto, Luca Calatroni, Aur´elie Bugeau, Nicolas Papadakis, and Carola-Bibiane Sch¨onlieb. Variational osmosis for non-linear image fusion. IEEE Transactions on Image Processing, pages 5507– 5516, 2020. [5] L. Zhang X. Qiu, M. Li and X. Yuan. Guided filter-based multi-focus image fusion through focus region detection. Signal Processing: Image Communication, pages 35–46, 2019. [6] H. Peng Y. Liu, X. Chen and Z.Wang. Multi-focus image fusion with a deep convolutional neural network. Information Fusion,, pages 191–207, 2017. 74 MambaCT: Feature Enhancement-Based Low-Dose CT Image Denoising Using Vision Mamba and a Scaling Adapter Abdelkarim Cherhabil1, Lahc`ene Mitiche1, and Amel Baha Houda Adamou-Mitiche1 1Department of Electronics and Telecommunications, Ziane Achour University of Djelfa, Laboratoire de Mod´elisation, Simulation et Optimisation des Syst`emes Complexes R´eels Abstract With the rapid development of Mamba models, Vision Mamba is replacing convolutional neural networks (CNNs) and vision transformers (ViTs), emerging as the dominant trend in computer vision tasks. After achieving remarkable success in natural language processing, Mamba models have garnered increasing interest in the medical imaging community for their ability to understand global context. However, there has been limited research on medical image denoising based on the Vision Mamba architecture. In this paper, we propose MambaCT, a model that integrates the key features of UNet and Mamba with a Scaling Adapter in the visual state space (VSS) Block. Additionally, we propose skip connection spatial-channel processing attention (SCSPA) to enhance feature integration and robustness as a pathway in place of traditional skip connections. MambaCT outperforms previous state-of-the-art (SOTA) models across various architectures in both visual quality and quantitative performance, requiring only 0.83G MACs and achieving an SSIM of 0.9104 and RMSE of 9.3423. The model was evaluated on the AAPM-Mayo Clinic low-dose computed tomography (LDCT) Grand Challenge Dataset. Keywords: Low-dose CT, Vision Mamba, Medical Image Denoising, Adapter, State Space Models, Auto-encoder. 1 Introduction Computed tomography is a diagnostic imaging method that precisely aligns X-ray, gamma, ultrasound, and ion beams to create cross-sectional images of the human body [1]. Clinical, industrial, and other fields make extensive use of CT [1] [31]. It is particularly effective in reconstructing organ structures at various depths and angles [2] [3]. Several algorithms have been created to improve image quality in low-dose CT (LDCT) scans in order to address this issue. Using physical models and existing data, researchers employ iterative techniques in classical ways to reduce noise and artifacts. For instance, some image priors are expressed as sparse transforms utilizing compressive sensing (CS) to address issues in internal CT, low-dose, few-view, and finite-angle CT [4]. Examples include dictionary learning [5], lowrank [6], non-local means (NLM) [7][8][9], total variation (TV), and its variations [10][11][12][13], among other methods. Li et al. [49] reconstructed feature similarities in large neighborhood images using NLM. Aharon et al. used dictionary learning [50] to denoise LDCT images, drawing inspiration from sparse representation theory, resulting in considerable improvement in denoising quality while reconstructing abdominal images [51]. Block-matching 3D (BM3D) has been shown by Feruglio et al. to be efficient for a range of X-ray imaging applications [52]. However, this method’s inaccuracy in determining the noise distribution in the image domain prevents the optimal balance between noise reduction and structure preservation. Due to restrictions on data volume, the accuracy of these conventional approaches is typically still poor [14]. Since the advent of deep learning, CNNs have been the dominant method for denoising low-dose CT (LDCT) images. CNNs extract features through convolutional operations, where the kernel moves across the entire image, resulting in a relatively small parameter volume for the CNN model. This approach effectively captures significant local features, and the network’s receptive field is incrementally expanded through layer stacking. Numerous traditional deep learning algorithms have been applied to the field of low-dose CT (LDCT) image denoising for reconstructing high-quality images. These include convolutional neural networks (CNNs) [15] [16] [29] [30], encoder-decoder networks with residual connections [17][18][19], and generative adversarial networks (GANs) [20][21]. The work of Chen et al. can be considered groundbreaking, as they were among the first to utilize convolution, deconvolution, and shortcut 75 connections to design a prototype of a residual encoder-decoder convolutional neural network, known as RED-CNN [17]. To improve the quality of denoised photos, Yang et al. used a generative adversarial network with Wasserstein distance (WGAN) and a perceptual loss mechanism [20]. Compared to other CT denoising techniques, Fan et al. developed a quadratic neuron-based autoencoder that is more resilient and useful for model efficiency [?]. The retrieval of detailed structural details in the denoised images may be adversely affected by CNNs’ limits in capturing long-range contextual information within images, notwithstanding their intriguing results for LDCT [23]. The integration of Transformer models into image denoising has significantly improved performance, resulting in higher accuracy and reduced processing times [24][25]. This advancement has brought about a revolution in image processing. Recent studies reveal that Transformer modules can effectively replace traditional convolutions in deep neural networks. They work by processing sequences of image patches, leading to the development of Vision Transformers (ViTs). Dosovitskiy et al. first proposed the vision transformer (ViT) in the CV field by mapping an image into 16×16 sequence words [24]. Wang et al. propose an innovative approach called the Convolution-free Token2Token Dilated Vision Transformer [26]. Luthra et al. introduce a fresh approach named Eformer, which stands for Edge Enhancement-based Transformer. Eformer is a unique architectural framework that constructs an encoder-decoder network using transformer blocks [27]. Jian et al. propose SwinCT, which utilizes a feature enhancement module (FEM) inspired by the Swin Transformer architecture. The FEM in SwinCT is employed to capture and enrich the high-level features within medical images [28]. According to the above analysis, Mamba models offer significant advantages over both CNN and transformer models, including greater visual interpretability due to their intrinsic Visual State Space (VSS) blocks [32]. Beyond their effectiveness, Mamba models are appealing to physicians because their self-explanatory nature allows doctors to understand the model’s reasoning. ¨ Ozt¨urk et al. [33] pioneered the application of an innovative SSM architecture, named DenoMamba, to enhance LDCT image denoising without increasing model complexity. This novel approach employs an hourglass-shaped structure, featuring encoder-decoder stages built with custom-designed FuseSSM blocks. Li et al. [34] introduced a CACTSR, which integrates VMamba and Transformer technologies with Mixed Attention Blocks and Cross Attention Blocks to enhance feature utilization and facilitate cross-window information interaction. The above studies demonstrate the promising results of Mamba-based deep learning models in visual tasks. Motivated by the aforementioned study, we introduce MambaCT, a paradigm that combines the key components of Mamba and UNet by integrating a Scaling Adapter within the VSS Block. Our approach incorporates an SCSPA module pathway in place of traditional skip connections, resulting in images that exhibit superior performance both quantitatively and visually. According to experimental results, our model outperforms other state-of-the-art models, achieving the highest SSIM value and the lowest RMSE value. The main contributions of this paper are as follows: 1. We propose MambaCT, a model that utilizes a U-Net-based Mamba network architecture and incorporates a Scaling Adapter within the VSS Block. The Scaling Adapter enhances the restoration of detailed and structural information in denoised images. 2. We propose the Skip Connection Spatial-Channel Processing Attention (SCSPA) module pathway as an alternative to traditional skip connections. 3. Assess the model’s performance by comparing it with previous works using various metrics, including MACs, SSIM, and RMSE. 76 2 Methods The architecture of the proposed MambaCT, as shown in Fig. 1, draws inspiration from both U-Net [35] and VMamba [36]. Designed specifically for LDCT denoising, MambaCT comprises four key modules: 1) Patch Extraction, 2) VSS Blocks, 3) SCSPA Module, and 4) Resizing Modules. Figure 1: The overall structure of MambaCT (a).The VSS Block serves as the primary building block of MambaCT, with SS2D and the Adapter as its core operations (b). 2.1 Patch extraction To train deep learning models effectively, a large volume of samples is crucial, which can be particularly challenging in clinical imaging. In our study, we addressed this issue by using CT scans with overlapping slices. This method has proven to be both effective and successful, as it helps capture perceptual differences in local regions and greatly boosts the number of samples available [37][38][39]. 2.2 VSS Block The core of MambaCT is the VSS Block, which serves as the primary building block of the model. The VSS Block incorporates SS2D and the Adapter as its core operations and is derived from [36], as shown in Figure 1(b). Its structure begins with Layer Normalization, followed by a split into two branches. The first branch applies a linear layer and the activation function SiLU [40]. The second branch processes the input through a linear layer, depthwise separable convolution, activation, and the SS2D module. the SS2D module provides contextual information to image patches via a compressed hidden state along scanning paths (Figure 2(b)), reducing computational complexity from quadratic to linear compared to self-attention mechanisms (Figure 2(a)). After SS2D, the features undergo Layer Normalization and are combined with the first branch’s output through element-wise multiplication. A Scaling Adapter is then applied, allowing for the learnable adjustment of the adapted features’ contribution. Directly after the adapter, we scale the embedding by a scale factor s[41]. Finally, the result passes through a linear mixing layer and is combined with a residual connection to produce the block’s output. this architecture efficiently processes spatial information while maintaining linear complexity, making it well-suited for medical image denoising in LDCT. 77 Figure 2: Comparison of correlation establishment between image patches via (a) self-attention and (b) the proposed 2D-Selective-Scan (SS2D). Red boxes indicate the query image patch, with patch opacity representing the degree of information loss. 2.2.1 2D-Selective-Scan for Vision Data A scan expansion operation, an S6 block, and a scan merging operation are the three primary parts of the SS2D module. According to Figure 3, SS2D first unfolds input patches into sequences along four distinct traversal paths (i.e., scan expanding), processes each patch sequence using a separate S6 block in parallel, and then reshapes and merges the resultant sequences to form the output map (i.e., scan merging). By adopting complementary 1D traversal paths, SS2D enables each pixel in the image to effectively integrate information from all other pixels in different directions, facilitating the establishment of global receptive fields in the 2D space. Figure 3: The overall structure of the 2D Selective Scan (SS2D) process. 78 2.2.2 Scaling Adaptor The Adapter operates as a bottleneck model. The down-projection layer reduces the dimensionality of the input embedding using a basic MLP layer with parameters Wdown ∈Rd׈ d, and the up-projection layer restores the compressed embedding to its original dimensionality with an additional MLP layer with parameters Wup ∈Rˆ d×d, where ˆ dis the bottleneck middle dimension and satisfies ˆ d < d. Additionally, there is a ReLU layer [42] between these projection layers for non-linear properties. Residual connections remain a crucial aspect, helping in training deeper networks by preventing gradient vanishing problems. 2.3 Skip Connection Spatial-Channel Processing Attention In contrast to using a single attention mechanism, the combination of channel attention and spatial attention, especially in a sequential manner, significantly enhances the model’s ability to capture important feature information [43]. Inspired by [44], we propose a SCSPA mechanism that applies sequential channel-spatial attention to the skip connections. As illustrated in Fig. 4, the SCSPA module consists of two key components: one for spatial attention and one for channel attention. A channel reduction action (C →C/rate), a ReLU activation, and a channel expansion operation (C/rate →C) comprise the Channel Attention Submodule. The two 7x7 convolutions that make up the Spatial Attention Submodule are followed by batch normalization (BN) and ReLU activation in the first one, and batch normalization and a sigmoid activation in the second. After that, element-wise multiplication is used to merge the outputs of these two routes. The dimensions of the input and the final output are [B, H, W, C]. Figure 4: The overall structure of SCSPA. 2.4 Resizing module The Patch Merging components function as downsampling mechanisms, diminishing the spatial dimensions of feature maps while amplifying the channel count. This approach enables the network to capture hierarchical features across various scales. Conversely, the Patch Expanding modules in the decoder act as counterparts to the Patch Merging modules in the encoder. They reverse the downsampling process, progressively restoring spatial resolution while decreasing the number of channels. 79 3 Experiment In this section, we begin by listing the languages and tools utilized: Python, PyTorch, and CUDA. Dataset:The publicly accessible clinical dataset from the 2016 NIH-AAPM Mayo Clinic LDCT Grand Challenge was used to train and evaluate the model [45]. Ten anonymous individuals’ 2,378 low-dose (quarter) and 2,378 normal-dose (full) CT scans with 3.0-mm whole-layer slices are included in this dataset. We chose patient L506’s data, which consists of 211 slice images with numbers ranging from 000 to 210, for testing. The model was trained using the data from the remaining nine cases. Experiment setup: PyTorch 1.11.0 [46] and CUDA 12.4.0 were used in the experiments, which were conducted on an Ubuntu 22.04 LTS system with an Intel(R) Core(TM) i7-12700k CPU @ 2.70 GHz. The four NVIDIA RTX 3070 Ti 8G GPUs were used to train the model. Four blocks were chosen at random from each image’s available slices for training. For 4,000 epochs, the batch size was fixed at 16. The ADAM-W optimizer, which has a learning rate of 1.0×10−5, was used to reduce the mean squared error loss. After training, the model’s performance was assessed using the standard metrics in the field. 4 Discussion To evaluate denoising performance in LDCT images, we retrained all models using their officially available code. We propose MambaCT, a model that combines key features from UNet and Mamba, enhanced with a Scaling Adapter in the VSS Block. Additionally, our approach incorporates an SCSPA module pathway in place of traditional skip connections to improve feature integration. As shown in Table 1for the L506 dataset, MambaCT achieved the highest quantitative metrics, surpassing all other methods. Table 1: Quantitative comparison of different methods on L506 in terms of learnable parameters (#param.), MACs, SSIM, and RMSE. Bold values represent our method’s performance. Method #param. MACs SSIM↑RMSE↓ LDCT – – 0.8759 14.2416 RED-CNN [17] 1.85M 5.05G 0.8952 11.5926 WGAN-VGG [20] 34.07M 3.61G 0.9008 11.6370 MAP-NN [47] 3.49M 13.79G 0.8941 11.5848 AD-NET [48] 2.07M 9.49G 0.9041 9.7166 MambaCT 62.08M 0.83G 0.9104 9.3423 To provide a thorough evaluation of denoising performance, we use both qualitative and quantitative methods. The quantitative analysis focuses on two key metrics: SSIM, and RMSE. Additionally, model complexity is assessed based on the number of trainable parameters (#param.) and multiply-accumulate operations (MACs). Table 1presents the average SSIM, and RMSE across all slices of L506. Our MambaCT model achieves the highest SSIM of 0.9104, the lowest RMSE of 9.3423. Figure 5presents the results of various networks on L506 with Lesion No. 575, while Figure 6displays the regions of interest (ROIs) from the rectangular area highlighted in Figure 5. Visual analysis of these figures 5and 6demonstrates MambaCT’s superior capability in achieving three key objectives: noise and artifact removal, maintenance of high-level spatial smoothness, and preservation of target image details. While RED-CNN, built on convolutional networks, shows proficiency in noise and artifact elimination while retaining image details, it faces limitations in structural recovery. This constraint stems from its computational architecture, which prioritizes high-frequency information extraction, such as texture details. Furthermore, RED-CNN’s effectiveness is hampered by its finite receptive field size, impeding comprehensive global information capture. Detailed examination of the ROIs in Figure 6reveals varying performance across methods: 1. WGANVGG and MAP-NN introduce unwanted artifacts, manifesting as additional shadows and tissue-like structures. 2. RED-CNN and AD-NET yield improvements in image clarity and smoothness compared to WGAN-VGG and MAP-NN, though residual blotchy noise persists around lesion areas. 80 Table 1: Architecture of ResNet-50 Layer Type Output Shape Parameters Input Layer (128, 128, 3) 0 Conv2D (64, 64, 64) 9,408 BatchNormalization (64, 64, 64) 256 MaxPooling2D (32, 32, 64) 0 Residual Block 1 (32, 32, 256) 215,296 Residual Block 2 (16, 16, 512) 1,187,840 Residual Block 3 (8, 8, 1024) 7,077,888 Residual Block 4 (4, 4, 2048) 14,942,208 GlobalAveragePooling (2048) 0 Dense (13) 26,637 Total Parameters 25.6 Million 3.2.2 EfficientNet-B0 **EfficientNet-B0** [8] is a lightweight and efficient model that uses compound scaling to balance depth, width, and resolution. The compound scaling method ensures that the model scales up uniformly across all dimensions, resulting in a highly efficient and scalable architecture. EfficientNet-B0 achieves the highest accuracy (98.60%) and ROC-AUC (99.60%) on the SIW 2021 dataset, making it the bestperforming model in this study. Its lightweight architecture, with only 5.3 million parameters, makes it suitable for real-world applications where computational resources are limited. Table 2: Architecture of EfficientNet-B0 Layer Type Output Shape Parameters Input Layer (128, 128, 3) 0 Conv2D (64, 64, 32) 864 BatchNormalization (64, 64, 32) 128 Conv2D (32, 32, 16) 4,608 BatchNormalization (32, 32, 16) 64 MaxPooling2D (16, 16, 16) 0 Conv2D (8, 8, 32) 4,640 BatchNormalization (8, 8, 32) 128 Conv2D (4, 4, 64) 18,496 BatchNormalization (4, 4, 64) 256 GlobalAveragePooling (64) 0 Dense (13) 845 Total Parameters 5.3 Million 3.2.3 VGG-16 **VGG-16** [6] is a deep convolutional network with 16 layers, known for its simplicity and effectiveness in feature extraction. The model consists of multiple convolutional layers followed by max-pooling layers, which reduce the spatial dimensions of the feature maps. VGG-16 has been widely used in various image classification tasks due to its ability to capture intricate patterns in images. In this study, VGG-16 achieves an accuracy of 97.85% on the SIW 2021 dataset. 3.2.4 GoogleNet **GoogleNet** [7] is a 22-layer deep network that uses inception modules to reduce computational cost. The inception modules allow the network to capture features at multiple scales, making it highly effective for complex image classification tasks. In this study, GoogleNet achieves an accuracy of 96.50% on the SIW 2021 dataset. 87 Table 3: Architecture of VGG-16 Layer Type Output Shape Parameters Input Layer (128, 128, 3) 0 Conv2D (128, 128, 64) 1,792 BatchNormalization (128, 128, 64) 256 MaxPooling2D (64, 64, 64) 0 Conv2D (64, 64, 128) 73,856 BatchNormalization (64, 64, 128) 512 MaxPooling2D (32, 32, 128) 0 Conv2D (32, 32, 256) 295,168 BatchNormalization (32, 32, 256) 1,024 MaxPooling2D (16, 16, 256) 0 Conv2D (16, 16, 512) 1,180,160 BatchNormalization (16, 16, 512) 2,048 MaxPooling2D (8, 8, 512) 0 Conv2D (8, 8, 512) 2,359,808 BatchNormalization (8, 8, 512) 2,048 MaxPooling2D (4, 4, 512) 0 Flatten (8192) 0 Dense (4096) 33,558,528 Dense (4096) 16,781,312 Dense (13) 53,261 Total Parameters 138 Million Table 4: Architecture of GoogleNet Layer Type Output Shape Parameters Input Layer (128, 128, 3) 0 Conv2D (64, 64, 64) 9,408 BatchNormalization (64, 64, 64) 256 MaxPooling2D (32, 32, 64) 0 Inception Module 1 (32, 32, 256) 163,840 Inception Module 2 (16, 16, 480) 580,608 Inception Module 3 (8, 8, 512) 1,024,000 Inception Module 4 (4, 4, 512) 1,048,576 GlobalAveragePooling (512) 0 Dense (13) 6,669 Total Parameters 7 Million 3.2.5 AlexNet **AlexNet** [3] is one of the earliest deep learning models, with 8 layers. Despite its relatively shallow architecture, AlexNet has been widely used in various image classification tasks. In this study, AlexNet achieves an accuracy of 95.12% on the SIW 2021 dataset, making it the weakest-performing model among the transfer learning models evaluated. 3.3 Training Configuration All models are trained using the **Adam optimizer** with a learning rate of 0.001, a batch size of 32, and 70 epochs. To enhance generalization and mitigate overfitting, data augmentation techniques such as rotation, shear, and zoom are applied during training. These techniques increase the diversity of the training data, enabling the models to learn more robust and invariant features. 3.4 Performance Evaluation The performance of the models is evaluated using a comprehensive set of metrics, including **Correct Classification Accuracy (CCA)**, **F1 score**, **Balanced Accuracy (BA)**, and **ROC AUC score**. 88 Table 5: Architecture of AlexNet Layer Type Output Shape Parameters Input Layer (128, 128, 3) 0 Conv2D (64, 64, 96) 34,944 BatchNormalization (64, 64, 96) 384 MaxPooling2D (32, 32, 96) 0 Conv2D (32, 32, 256) 614,656 BatchNormalization (32, 32, 256) 1,024 MaxPooling2D (16, 16, 256) 0 Conv2D (16, 16, 384) 885,120 BatchNormalization (16, 16, 384) 1,536 Conv2D (16, 16, 384) 1,327,104 BatchNormalization (16, 16, 384) 1,536 Conv2D (16, 16, 256) 884,992 BatchNormalization (16, 16, 256) 1,024 MaxPooling2D (8, 8, 256) 0 Flatten (16384) 0 Dense (4096) 67,092,992 Dense (4096) 16,781,312 Dense (13) 53,261 Total Parameters 61 Million These metrics are particularly well-suited for assessing performance on imbalanced datasets, as they account for both precision and recall, ensuring a more holistic evaluation of the models’ predictive capabilities. 4 Results and Discussion The SIW 2021 dataset, containing 13 scripts (e.g., Arabic, Bengali, Gujarati), is used for evaluation. The dataset is divided into training and testing sets with a ratio of 70:30 (60,643 training images and 26,012 testing images). This division ensures a robust evaluation of the model’s generalization capabilities. The results demonstrate that transfer learning models achieve high accuracy across three tasks: mixed scripts, printed scripts, and handwritten scripts. The performance metrics for each task are summarized in Table 6. Table 6: Performance Metrics for Transfer Learning Models (Input Size: 128x128 Pixels) Model Accuracy (%) Precision (%) Recall (%) F1 Score (%) ROC-AUC (%) EfficientNet-B0 98.60 98.58 98.60 98.59 99.60 ResNet50 98.20 98.18 98.20 98.19 99.40 VGGNet19 97.90 97.88 97.90 97.89 99.25 VGGNet16 97.85 97.83 97.85 97.84 99.20 GoogleNet 96.50 96.48 96.50 96.49 98.80 AlexNet 95.12 95.10 95.12 95.11 98.50 4.1 Analysis of Results The results highlight the effectiveness of transfer learning models in handling diverse script identification tasks. EfficientNet-B0 achieves the highest accuracy (98.60%) and ROC-AUC (99.60%), followed closely by ResNet50 (98.20%) and VGGNet19 (97.90%). These models benefit from their deep architectures and pretrained weights, which enable them to extract complex features effectively. •EfficientNet-B0 stands out as the most efficient model, with only 5.3 million parameters and a training time of 70 seconds per epoch. Its lightweight architecture makes it suitable for real-world applications where computational resources are limited. 89 •ResNet50 and VGGNet19 also perform well but require significantly more parameters and longer training times, making them less efficient for large-scale deployments. •AlexNet, being one of the earlier deep learning models, performs the weakest, with an accuracy of 95.12%, likely due to its relatively shallow architecture compared to more modern models. 4.2 Limitations and Future Work While transfer learning models demonstrate strong performance, they have certain limitations. For instance, they require significant computational resources for fine-tuning, especially on large datasets. Additionally, their performance may degrade when applied to scripts or languages not well-represented in the pretraining dataset (e.g., ImageNet). Future work will focus on addressing these limitations by exploring hybrid models that combine the strengths of transfer learning and custom architectures. We also plan to investigate the use of unsupervised or semi-supervised learning techniques to reduce the reliance on labeled data. 5 Conclusion Transfer learning models demonstrate robust performance in script identification for both handwritten and printed word images. These models effectively handle class imbalance and script variations, achieving high accuracy and generalizability. Data augmentation and external data significantly enhance performance, making transfer learning a promising solution for real-world applications. Future work will focus on further optimizing these models and exploring their applicability to other document analysis tasks. References [1] A. Das et al. Icdar 2021 competition on script identification in the wild. In Document Analysis and Recognition – ICDAR 2021, pages 738–753, 2021. [2] K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. [3] A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems (NeurIPS), volume 25, pages 1097– 1105, 2012. [4] P. B. Pati and A. G. Ramakrishnan. Word level multi-script identification. Pattern Recognition Letters, 29(9):1218–1229, 2008. [5] B. Shi, X. Bai, and C. Yao. Script identification in the wild via discriminative convolutional neural network. Pattern Recognition, 52:448–458, 2016. [6] K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. [7] C. Szegedy et al. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1–9, 2015. [8] M. Tan and Q. V. Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning (ICML), pages 6105–6114, 2019. 90 Parameter-Efficient Fine-Tuning for LLM-Based Arabic-to-English Machine Translation Moudjar Amina1, Bahloul Belahcene1, and Aliane Hassina2 1Djilali Bounaama University, [email protected] [email protected] 2CERIST, [email protected] Abstract Large Language Models (LLMs) such as GPT-3, BLOOM, BERT... have revolutionized natural language processing (NLP), particularly in translation. However, fine-tuning these models for downstream tasks, such as Arabic-to-English translation, requires extensive computational resources. Traditional full fine-tuning methods that involve updating all parameters of the model pose significant computational and memory challenges, notably for models with billions of parameters. This study investigates the application of LoRA-based PEFT methods on two chosen models for Arabic to English translation, AraT5 and NLLB-200, with a focus on understanding the trade-offs between computational efficiency and translation quality. Keywords: Large Language models, Efficiency, Computational challenges, Parameter-efficient fine-tuning, Arabic-to-English Machine translation. 1 Introduction The advent of LLMs has revolutionized the field of natural language processing, they have large numbers of parameters and complex architectures to capture intricate language patterns, making them highly effective in translation tasks that need nuanced understanding and generation. Yet, fine-tuning LLMs for specific tasks requires enormous computational resources. Parameter-Efficient Fine-Tuning (PEFT) techniques have emerged to mitigate these challenges by selectively adjusting a small subset of parameters. Among PEFT methods, Low-Rank Adaptation (LoRA) and its variants stand out for their effectiveness, by introducing low-rank matrices to specific layers, allowing the model to learn task-specific adaptations efficiently. Quantized methods go further by quantizing the parameters involved in fine-tuning. Our study focuses on LLMs with encoder-decoder architecture: AraT5 and NLLB-200, which are ideal for Arabic-to-English translation. By applying the previous techniques, we aim to capture the trade-offs between computational efficiency and model performance in machine translation. Our key contributions are: 1. We apply LoRA, DoRA, QLoRA, and QDoRA techniques to fine-tune the AraT5 and NLLB-200 models for Arabic-to-English translation. 2. We assess the computational efficiency of these methods, comparing them to traditional full fine tuning. 3. We evaluate translation quality across the different f ine-tuning methods, examining the trade-offs be tween efficiency and performance. This paper is organized as follows: The first section reviews LLM-based machine translation approaches, the second section details parameter-efficient fine-tuning techniques, the third section presents our methodology, chosen datasets and models, and the fourth section compares resource usage and translation quality across methods and models. 2 Related works Machine Translation has undergone significant transformation recently, primarily due to the rapid advancements in LLMs. These advancements have pushed research into LLM-based machine translation, focusing on two main paradigms: In-Context Learning (ICL) and Finetuning. ICL leverages optimal in-context examples [2] [36] [18] , dictionary knowledge [13] [25], adaptive learning 91 [35] [27], and translation memories [34] to enhance translation accuracy. Traditional machine translation models, particularly those using statistical methods, struggle with contextually rich languages like Arabic, which demand effective capture of long-range dependencies and contextual nuances.Recent advancements in context-aware neural machine translation models, particularly those using self-attention mechanisms, have shown better performance by dynamically focusing on different parts of the input sentence and its context. These models benefit from incorporating larger context windows and external contextual information, such as linguistic annotations and discourse relations, significantly improving the translation quality. Concurrently, finetuning has been instrumental in augmenting LLM’s capability to translate unseen languages and domains [42] [26] and in building multilingual models [44] [46]. Additionally, research has delved into post-editing translation outcomes [28] [33] and utilizing LLMs for machine translation evaluation [12] [11]. The finetuning process for Arabic to English involves adapting pre-trained LLMs, such as T5, to the specific requirements of the translation task. This process includes further training on parallel Arabic-English datasets, which helps the model capture the syntactic, semantic, and contextual intricacies of both languages. Techniques like domain adaptation, advanced regularization methods, back-translation, and self-training enhance the performance and robustness of the finetuned models. The quality and size of the parallel corpus are crucial for achieving high translation accuracy, highlighting the importance of high-quality, diverse, and representative sentence pairs in the training data. Recent advancements in multilingual machine translation have shown significant improvements in performance across various language families, including Afro-Asiatic languages and specific language pairs such as Arabic to English. Zhu et al. (2023) [45] evaluated several LLMs on the FLORES-101 dataset [14] using the SentencePiece BLEU (spBLEU) metric (both SentencePiece BLEU and sacreBLEU are libraries used for calculating BLEU scores). For Afro-Asiatic languages (which include Arabic), LLaMA2-7B [39] achieved the highest BLEU score of 57.72, followed by XGLM-7.5B [22] at 54.51 and Falcon-7B [3] at 38.62. Focusing on the Arabic-English pair, GPT-4 performed better than ChatGPT, with scores of approximately 45 and 40 respectively. These results come from remarkably large models, which explains the high BLEU score results. 3 Parameter-efficient Fine-tuning (PEFT) PEFT adapts an LLM to downstream tasks by freezing the entire LLM backbone and updating only a small set of newly introduced parameters. PEFT methods can be classified into four categories: lowrank adaptation (LoRA [17]), adapter-based tuning (inserting trainable modules into LLMs to simplify fine-tuning [16]), prefix tuning (adding trainable vectors to each LLM layer that are adapted to specific tasks [21]), and prompt tuning (adjusting only the input layer by incorporating trainable prompt tokens that can be placed at the beginning or within the input text [20]). PEFT methods necessitate careful consideration of several factors to balance performance and efficiency effectively. A major challenge lies in minimizing trainable parameters while maintaining Substantial performance [30]. Fine-tuning too few parameters can restrict the model's adaptability to the target task, while excessive fine-tuning can degrade the computational advantages of PEFT [9] [24]. The success of PEFT also relies on the quality and quantity of data available, particularly in domains with limited or noisy data where achieving the same accuracy as full fine-tuning can be tough. In such cases, careful selection of data augmentation techniques and transfer learning strategies is important [8] [4]. 3.1 Low Rank Adaptation (LoRA) Low-Rank Adaptation [17], proposed by Hu et al. (2021), is a widely used Parameter-Efficient FineTuning approach aimed to optimize the adaptation of LLMs to specific tasks. During full fine-tuning, the model is initialized to pre-trained weights W0and updated to W0+ ∆WThe basic hypothesis behind LoRA is that during fine-tuning, A low-rank approximation can powerfully capture the necessary adjustments to the model’s weights. This means that the changes in the weight matrix resulting from task-specific adaptation have a low ”intrinsic rank” [1]. As shown in Figure 3the information contained within the ∆Wmatrix can be represented using fewer dimensions than the original matrix and therefore, the full-dimensional update can be approximated by a product of two smaller matrices while keeping the original weights frozen. 92 Figure 1: Comparison between the traditional fine-tuning approach and the LoRA method 3.1.1 Mathematical Formulation Considering a pre-trained weight matrix Wof a neural network. During fine-tuning, instead of updating Wdirectly, LoRA introduces two trainable low-rank matrices A∈Rm×rand B∈Rr×n. The weight update is then formulated as : Wupdated =W+ ∆W=W+AB (1) Here, ris the rank, a hyperparameter that defines the dimensionality of the low-rank approximation. Using this approach, the amount of trainable parameters is remarkably decreased, as only the matrices. A and Bneed to be learned, while the original weight matrix Wremains frozen. This makes the fine-tuning process much more memory and computation efficient. 3.1.2 Reparametrization and Optimization LoRA modifies the forward pass of the neural network by adding the low-rank update ∆W=AB to the original output. Specifically, if the original output is h=W0x, the updated output becomes: Wupdated =W0x+ ∆Wx =W0x+ABx (2) In practice, during backpropagation, the frozen pre-trained weights W0remain untouched, and the loss is only used to update the Band Amatrices introduced by LoRA. Ais initialized with a random Gaussian distribution, while Bis initialized to zero, ensuring that the initial value of ∆Wis zero. The scaling factor αis introduced to balance the contribution of ∆Wduring training, which is crucial for controlling the impact of the low-rank updates and making sure that the fine-tuning process remains stable and effective. 3.1.3 Applying LoRA to a Transformer In a transformer architecture of an LLM, it is more common to apply LoRA to the attention layers because they are computationally expensive and have a significant number of parameters, which makes them the prime targets for parameter-efficient fine-tuning. However, it’s not limited to just attention layers; it could also be applied to other layers like feed-forward networks. LoRA allows the fine-tuning process to require fewer parameters and less computational power while still achieving strong performance. However, as with many advancements in AI, researchers have recently introduced other innovative alternatives and derivatives of LoRA, such as DoRA [23], LoRA+ [15], QA-LoRA [41], QLoRA [6], QDoRA [23], and DyLoRA [40], depending on the model architecture itself and the area of focus. 3.2 Weight-Decomposed Low-Rank Adaptation (DoRA) Weight-Decomposed Low-Rank Adaptation [23] was introduced and built on LoRA by introducing a decomposition of the weight matrix into two components to fine tune them: magnitude and direction, as illustrated in Figure 2. This decomposition separates the fine-tuning of these components, addressing issues in LoRA to make subtle adjustments to weight directions while efficiently handling parameter updates. The process behind DoRA is divided into two main steps. First, the weight matrix W0from a pretrained model is decomposed into two components: 93 Figure 2: An overview of DoRA [23]. •Magnitude Vector m: This vector represents the norm or length of each column in the weight matrix, capturing the scale information. •Directional Matrix V: Where each column vector of the weight matrix is normalized by dividing by its magnitude, keeping only the directional information. Once the pretrained weights are decomposed, and given the substantial size of the directional component in terms of parameters, LoRA is applied exclusively to the directional matrix V, and mis trained as it is, which is feasible because it has just one dimension. DoRA improves both the learning capacity and stability of LoRA, without causing any additional inference overhead. It enables fine-tuning a pretrained model in a way that is computationally efficient and potentially more responsive to new data, maintaining the strengths of the original model while adapting it to new tasks or datasets. 3.3 Quantized Low rank adaptation (QLoRA) QLoRA reduces the memory footprint of LLMs by compressing weights from high-precision data types such as 32-bit floating point to lower-precision formats such as 4-bit integers or NormalFloat, while integrating trainable Low-Rank Adapters (LoRA) for fine-tuning. The pretrained weights remain frozen, and only the LoRA parameters are updated, minimizing memory requirements. During computational tasks, weights stored as 4-bit NormalFloat are dequantized to 16-bit BrainFloat (bfloat16) for both the forward and backward passes. However, only the LoRA parameter's gradients are computed. QLoRA achieves high-fidelity 4-bit fine tuning via two proposed techniques 4-bit NormalFloat (NF4) Quantization and Double Quantization. Paged Optimizers were also introduced to prevent memory spikes during gradient checkpointing from causing out-of-memory errors that have traditionally made fine tuning on a single machine difficult for large models. 3.3.1 4-bit NormalFloat Quantization The NormalFloat (NF) data type builds on Quantile Quantization [5] which is an information-theoretically optimal data type that ensures each quantization bin has an equal number of values assigned from the input tensor. Quantile quantization works by estimating the quantile of the input tensor through the empirical cumulative distribution function. This technique helps compress model weights effectively, but the process of quantile estimation is computationally expensive. Fast quantile approximation algorithms, such as SRAM quantiles, help mitigate this cost, but approximation errors arise, especially for outliers, which are often critical. QLoRA addresses these issues by recognizing that pre-trained LLM weights typically follow a zerocentered normal distribution with a standard deviation α. By transforming all weights to fit within a fixed range (e.g., [−1,1]), accurate quantile estimation becomes feasible, eliminating the need for computationally expensive approximation algorithms. The result is the 4-bit NormalFloat (NF4) data type, which is specifically optimized for normally distributed data. It normalizes neural network weights into this fixed range and quantizes them accordingly, enabling precise weight compression with minimal performance loss. 94 3.3.2 Double Quantization A method that reduces the average memory footprint by quantizing the quantization constants [7]. This saves approximately 0.37 bits per parameter, which translates to around 3 GB for a 65B model. It further reduces the memory overhead of quantization constants. By quantizing both the model weights and the quantization constants, QLoRA achieves higher memory efficiency without negatively impacting model performance. 3.3.3 Paged Optimizers Fine-tuning LLMs can also generate memory spikes, primarily when processing long sequences or large mini-batches. QLoRA harnesses Paged Optimizers to handle these spikes efficiently, leveraging NVIDIA’s unified memory system, which automatically transfers memory between the CPU and GPU. When the GPU runs out of memory, data is paged to the CPU and then moved back to the GPU when needed. This seamless paging mechanism guarantees that memory bottlenecks do not interrupt the training process, allowing QLoRA to handle larger models and batch sizes with fewer resources. 3.4 Quantized Weight-Decomposed Low-Rank Adaptation (QDoRA) QDoRA combines the memory efficiency of QLoRA with the DoRA fine-tuning. QDoRA leverages quantization to compress weights into low-precision formats, significantly reducing memory footprint. However, it goes further by incorporating weight decomposition, as seen in DoRA, to achieve more granular optimization during fine-tuning. This allows QDoRA to maintain both high performance and low computational requirements, making it suitable for fine tuning and training large models like Llama 3 on consumer-grade GPUs. 4 Methodology In this work, we explore the performance of small-sized LLMs in Arabic-to-English machine translation, focusing specifically on encoder-decoder-based architectures. Our research is based on two publicly available LLMs on HuggingFace: the Arabic-focused AraT5v2 Base and the multilingual NLLB-200 distilled-600M models. To evaluate these models, we conducted experiments using the United Nations Parallel Corpus [47] with AraT5v2 [10] and the OPUS-100 corpus [43] with NLLB-200 [38]. Our objective was to evaluate the trade-offs between computational resource efficiency, such as GPU power consumption and memory allocation, and translation performance using BLEU, and ROUGE and Perplexity scores, without an explicit aim to improve translation quality. We applied four fine-tuning techniques in addition to Full fine-tuning (which served as a baseline used for comparison): LoRA, DoRA, and quantized variants: QLoRA and QDoRA. 4.1 Experiments on AraT5v2 with United Nations Parallel Corpus We used the United Nations Parallel Corpus with the AraT5v2 model. We adopted the following dataset splitting strategy to ensure a balanced and effective model training, validation, and testing. We started by loading the first 20 000 examples from the corpus. The dataset was split as into: •Training Set: 15,000 examples (75%) were dedicated to training, ensuring that most of the data is used for model learning. •Validation Set: 2,500 examples (12.5%) were reserved for validation. This validation set is used to track the model’s performance on unseen data, serving it to prevent overfitting. •Test Set: 2,500 examples (12.5%) were set aside for testing. The test set is used for the final evaluation. AraT5v2 built on a foundation set by the original AraT5 model [29]. AraT5 was inspired by the T5 (Text-to-Text Transfer Transformer) model [32], which reframes all NLP tasks into a text-to-text format, offering a unified framework for language modeling tasks. Tokenization is an important step in converting raw text into a format that a model like AraT5v2 can process. AraT5v2 relies on tokenized inputs for both the source language (Arabic) and the target language (English). 95 •Task-Specific Prefix: The T5 model needs a clear prompt of the task it is performing. In this case, we prepend the Arabic input text with the prefix: ”translate Arabic to English: ” •Tokenization Using AraT5v2’s Tokenizer: We use Hugging Face’s AutoTokenizer to tokenize the Arabic input sentences and their corresponding English translations. AutoTokenizer is a generic tokenizer class in the Huggingface Transformers library that automatically selects the proper tokenizer for a given model. The original T5 models (including AraT5, which is based on T5) typically use SentencePiece tokenizers [37], which is a text tokenization algorithm widely used in modern NLP models, particularly in models that need to handle complex languages with rich morphology (like Arabic). Unlike traditional tokenizers that split text based on spaces or punctuation, SentencePiece treats the entire text as a continuous stream of characters and learns how to break it into subword units. •Truncation and Padding: To ensure that the input sequences fit within the model’s constraints, we limit the maximum sequence length to 128 tokens. Sequences longer than this are truncated, while shorter ones are padded to a uniform length. This ensures consistency across batches during training. The tokenized dataset consists of pairs of input and output sequences, where each sequence is a list of tokens representing either an Arabic sentence (input) or its English translation (output). These tokenized sequences are then ready for training the AraT5v2 model. We fine tuned AraT5v2 using different approaches: Full fine-tuning, LoRA, DoRA, and QDoRA: 1. Full fine-tuning means updating all the parameters of the model based on the specific wanted task. We defined several key hyperparameters (used for other techniques as well): Once the training setup Table 1: Hyperparameters Used for Model Training Learning rate 2·10−4 Batch size 2 Number of epochs 5 is complete, the model is trained on the training set by backpropagating through the entire model, including all attention and feed-forward layers, updating every weight in the network. Training is followed by evaluation on the validation set after each epoch to track its performance. 2. We used LoRA to fine-tune AraT5v2, focusing on the following layers involved in the attention mechanism: the key (k), query (q), value (v), and output (o) layers. After freezing the core model parameters, LoRA introduces learnable low-rank matrices that adjust the key, query, value, and output layers of the attention mechanism. These matrices are updated during training, while the original parameters remain untouched. The low-rank structure allows for efficient fine-tuning with fewer parameters to update. 3. In another experiment we applied DoRA by freezing the original model parameters and decomposing the weight matrices involved in the key (k), query (q), value (v), and output (o) layers of the attention mechanism. By applying DoRA to the Target Modules, we decomposed the weight matrices associated with them into their magnitude and directional components. The fine-tuning process updates the directional matrices using LoRA, while the magnitude vector is trained directly. After training, all the components are recombined to form the updated weight matrices which is used for inference. 4. Lastly, QDoRA is applied to the AraT5v2 model by combining two techniques: quantization and Weight Decomposed Low-Rank Adaptation (DoRA). We applied these techniques to the model by quantizing the model’s weights to 4-bit precision, then freezing the core weights of the model (from its pre-trained state), the Decomposition of Weights, LoRA is then applied specifically to the directional matrices within the attention layers chosen, and alongside the updates to the directional matrices, the magnitude vector (which represents the scale of the weights) is also fine-tuned. Since the magnitude has significantly fewer parameters, it can be trained directly without requiring the same low-rank approximations used for the directional matrices. Table 3summarizes the previous techniques. 96 Figure 10: GPU power consumption in Watt in NLLB200-600M while full fine-tuning and QLoRA prioritizes speed at the cost of high resource use. Figure 11: GPU memory allocated in NLLB200-600M The graph presented in Figure 11 shows the GPU memory allocation over time for various finetuning methods of the NLLB200-600M model. Full fine-tuning consumes the most memory, peaking at 1.51 ·10+10 bytes and remaining constant till the end of its training. DoRA and QLoRA exhibit gradual memory increases, stabilizing at 1.09 ·10+10 and 9.38 ·10+9 bytes, respectively, making them more memory efficient (especially QLoRA). LoRA also shows a gradual increase but experiences fluctuations before stabilizing around 1.06 ·10+10 bytes. Overall, QLoRA is notably the most memory efficient technique, while LoRA shows some instability before leveling off, but takes more time training. Figure 12 shows the GPU temperature over time for different training methods: The DoRA, QLoRA 103 Figure 12: GPU temperature in NLLB200-600M and full finetuning methods maintain a relatively high and stable temperature, hovering around 75−77◦C. There are minor fluctuations, but overall, the temperature remains steady. The LoRA method shows more balance in temperature, ranging between 56 −61◦C, with noticeable dips and rises. It operates at a generally lower temperature compared to the other methods. The Figure 13 shows GPU utilization Figure 13: GPU utilization in NLLB200-600M over time for different methods: Full fine tuning has high GPU utilization, often reaching around 78%. There are occasional drops, but it generally maintains a high level of resource usage, indicating intensive processing. QLoRA also shows relatively high utilization, fluctuating between 49% and 60% and DoRA has moderate utilization, around 26%-35%, indicating a balanced approach between performance and resource usage. LoRA shows the lowest GPU utilization, generally staying below 30%, reflecting a more 104 optimized use of resources. In general, Full finetuning and QLoRA make the most aggressive use of GPU resources. Figure 14: System CPU utilization in NLLB200-600M The graph in Figure 14 displays system CPU utilization over time for different methods: DoRA and LoRA exhibit high variability with frequent spikes in CPU usage, often reaching above 80, this indicates intensive CPU involvement and processing. QLoRA in Figure 15 maintains the most balanced CPU utilization, with fewer and lower spikes. NLLB200-600M full finetuning displays low and most stable CPU utilization. In total, DoRA and LoRA show more intensive and variable CPU usage, while QLoRA and full finetuning use CPU resources more efficiently and steadily. In a comparative analysis of fine-tuning methods across models, distinct trade-offs show up in terms of performance, resource usage, and training efficiency. Full Fine-Tuning updates all model parameters, making it the most computationally intensive approach. It achieves the fastest training speed but consumes the most GPU power and memory, limiting its practicality in memory-constrained environments. LoRA and DoRA, which fine-tune a smaller subset of parameters, and QLoRA, which also introduces quantization to further reduce memory and computation demands, offer a compromise. They significantly reduce GPU memory consumption and power usage, with LoRA showing moderate memory access. DoRA has slightly higher power usage than LoRA but converges faster in some cases, making it ideal for scenarios where power efficiency is crucial but training speed remains important. QDoRA, which introduces quantization to further reduce memory and computation demands, uses the least memory, but the added quantization complexity causes occasional GPU power spikes and slower training speeds compared to LoRA. This makes it suitable for extreme memory-constrained environments, regardless of trade-off in speed and performance. 5.2 Translation Quality comparison 5.2.1 Evaluation metrics results According to the following Bar chart : Full Fine-Tuning achieves the highest ROUGEL and SacreBLEU scores compared to the other methods. QDoRA and DoRA show nearly identical performance. LoRA shows a similar ROUGE score as DoRA and QDoRA but seems to slightly underperform them in SacreBLEU and Perplexity. While Full FineTuning provides the best results, the gap between LoRA-based methods and Full Fine-Tuning is not large. This reinforces the idea that LoRA-based methods offer a good balance between performance and computational efficiency, especially when full fine-tuning is resource-prohibitive. 105 Figure 15: NLLB200-600M-QLoRA system CPU utilization Figure 16: Bar Chart showing comparative araT5v2 Evaluation Using ROUGEL, SacreBLEU, in addition to Perplexity Figure 17: Bar Chart showing comparative NLLB200-600M Evaluation Using ROUGEL, SacreBLEU, in addition to Perplexity According to Figure 17, The full fine-tuning method starts at a higher ROUGEL and SacreBLEU 106 scores and remains ahead of all parameter-efficient methods. DoRA and LoRA show almost identical performance, with DoRA slightly edging out. QLoRA remains the lowest performer. The perplexity shows that the uniformity of perplexity scores across all methods shown here is notable. It implies that, in terms of understanding and predicting the structure of the target language, all methods are equally effective. Full fine-tuning offers the best evaluation scores results, but parameter-efficient methods like DoRA and LoRA are closely following, suggesting they provide a competitive trade-off between performance and resource usage. 5.2.2 Human evaluation results In our inference tests, Arabic sentences were translated using AraT5v2 and compared with Google Translate. For simple sentences, all fine-tuning methods (Table 7) performed similarly, showing that lightweight techniques like LoRA and QDoRA can handle basic translations effectively. However, errors in complex sentences - especially with idiomatic expressions - revealed significant limitations. Google Translate generally outperformed the fine-tuned models in these cases, delivering more fluent translations, while the models often struggled with figurative language, producing more literal or incomplete outputs. Noticeably, some Arabic words remained untranslated across all fine-tuning methods, indicating gaps in vocabulary handling. Full Fine-Tuning provided the best overall performance but occasionally overcomplicated translations, likely due to overfitting. For NLLB200-600M (Table 8), all fine-tuning methods closely matched Google Translate on straightforward sentences but diverged on abstract phrases. Overall, Full Fine-Tuning excelled but was resourceheavy, while LoRA and DoRA offered a balanced trade-off, and QLoRA proved best suited for extreme memory constraints but was slower to converge. Table 9 presents comparison between GPT models BLEU scores and ours, due to hardware limitations, we were unable to use similarly powerful models in our experiments, resulting in comparatively lower performance to the results we presented in the related works. 6 Conclusion This study explored the trade-offs between computational efficiency and Arabic-to-English translation quality in LLMs, focusing on parameter-efficient fine-tuning techniques, particularly variations of LoRA. While full fine-tuning showed optimal translation accuracy, LoRA and DoRA achieved comparable quality with reduced computational costs remarkably. QLoRA offered additional memory efficiency, though at the expense of longer training time. Future work will explore scaling these methods to larger models and enhancing their capacity to capture complex linguistic structures. All code and models are provided in a GitHub repository as open source, editable Jupyter notebooks. 107 Table 7 Translations generated by araT5v2 Arabic original text DoRA Full fine tuning LoRA QDoRA Google Translate  à @ HPQ ¯ ú  æ ®J  £ð Q  « @  á«  Im' . B  èYK Yg .Q ¯ I have decided to change my job to find new opportunities I decided to change my job to seek new opportunities I decided to change my job to find new opportunities. I decided to change my job to find new opportunities. I decided to change my job to look for new opportunities  J  Ë@  ZA J . K@      Y   @ ú  ¯ I .  J ºË@  áÓ   á K . YmÌ'@ èY g I . ª ÊË@ð Y m.Ì'@ The spend of books is stronger from the table of books, in the same line between the study and the game The cost estimate provides for the highest number of newspapers at the extent of The spend of books is expenditure at the line of the line between the and game. The printing of books are expenditure to the time and the game. The sword is more truthful than books in its sharpness, the line between seriousness and play.  J Ë@ ZAJ . K@  Y@ ú  ¯ I . JºË@  áÓ  á K .YmÌ'@ èYg I .ªÊË@ð Ym.Ì'@  J Ë@ is the most accurate of books, in the scale of the time and game. The sword is the most best source of books at the time of serious and play The  J Ë@ is most accurate of books in the same line between serious and game. The  J Ë@ is the best evidence of books in the same line between the argument and play The sword is more truthful than books in its sharpness, the limit between seriousness and play Table 8 Translations generated by NLLB200-600M Arabic original text DoRA Full fine tuning LoRA QLoRA Google Translate  à @ HPQ ¯ ú  æ ®J  £ð Q  « @  á«  Im' . B  èYK Yg .Q ¯ I decided to change jobs to look for new opportunities. I decided to change jobs to look for new opportunities I decided to change jobs to look for new opportunities. I decided to change jobs to look for new opportunities. I decided to change my job to look for new opportunities  J  Ë@  ZA J . K@      Y   @ ú  ¯ I .  J ºË@  áÓ   á K . YmÌ'@ èY g I . ª ÊË@ð Y m.Ì'@ The sword I believe is a book narrative about the boundary between grandfather and play. The sword. I believe a prophecy from a book alone about the boundary between seriousness and play I believe the sword is a book narrative about the boundary between grandfather and play. I believe the sword is a book narrative about the boundary between grandfather and play. The sword is more truthful than books in its sharpness, the line between seriousness and play  J Ë@ ZAJ . K@  Y@ ú  ¯ I . JºË@  áÓ  á K .YmÌ'@ èYg I .ªÊË@ð Ym.Ì'@ The sword believes prophecies from books about the boundary between grandpa and play. The sword believed the prophecies of the books in the boundary between grandfather and play. The sword believes prophecies from books about the boundary between grandfather and play. Sword believes prophecies from books about the boundary between grandfather and play. The sword is more truthful than books in its sharpness, the limit between seriousness and play Table 9 BLEU Scores for GPT Language Models compared to araT5v2 Base and NLLB200-600M Model BLEU score GPT-4 ∼45 ChatGPT ∼40 araT5v2 Full fine tuning 19.951 araT5v2 Full LoRA 12.531 araT5v2 Full DoRA 13.006 araT5v2 Full QDoRA 13.027 NLLB200-600M Full fine tuning 34.245 NLLB200-600M LoRA 32.676 NLLB200-600M DoRA 32.812 NLLB200-600M QLoRA 31.595 108 References [1] Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta. Intrinsic dimensionality explains the effectiveness of language model fine-tuning, 2020. [2] Sweta Agrawal, Chunting Zhou, Mike Lewis, Luke Zettlemoyer, and Marjan Ghazvininejad. Incontext examples selection for machine translation. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Findings of the Association for Computational Linguistics: ACL 2023, pages 8857–8873, Toronto, Canada, July 2023. Association for Computational Linguistics. [3] Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, M´erouane Debbah, ´ Etienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. The falcon series of open language models, 2023. [4] Golla Anjali, Santosh Sanjeev, Akuraju Mounika, Gangireddy Suhas, G. Pradeep Reddy, and Yarlagadda Kshiraja. Infant cry classification using transfer learning. In TENCON 2022 - 2022 IEEE Region 10 Conference (TENCON), pages 1–7, 2022. [5] Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer. 8-bit optimizers via block-wise quantization, 2022. [6] Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. Qlora: Efficient finetuning of quantized llms, 2023. [7] Tim Dettmers and Luke Zettlemoyer. The case for 4-bit precision: k-bit inference scaling laws, 2023. [8] Raman Dutt, Ondrej Bohdal, Sotirios A. Tsaftaris, and Timothy Hospedales. Fairtune: Optimizing parameter efficient fine tuning for fairness in medical image analysis, 2024. [9] Raman Dutt, Linus Ericsson, Pedro Sanchez, Sotirios A. Tsaftaris, and Timothy Hospedales. Parameter-efficient fine-tuning for medical image analysis: The missed opportunity, 2024. [10] AbdelRahim Elmadany, El Moatez Billah Nagoudi, and Muhammad Abdul-Mageed. Octopus: A multitask model and toolkit for Arabic natural language generation. In Hassan Sawaf, Samhaa El-Beltagy, Wajdi Zaghouani, Walid Magdy, Ahmed Abdelali, Nadi Tomeh, Ibrahim Abu Farha, Nizar Habash, Salam Khalifa, Amr Keleg, Hatem Haddad, Imed Zitouni, Khalil Mrini, and Rawan Almatham, editors, Proceedings of ArabicNLP 2023, pages 232–243, Singapore (Hybrid), December 2023. Association for Computational Linguistics. [11] Patrick Fernandes, Daniel Deutsch, Mara Finkelstein, Parker Riley, Andr´e F. T. Martins, Graham Neubig, Ankush Garg, Jonathan H. Clark, Markus Freitag, and Orhan Firat. The devil is in the errors: Leveraging large language models for fine-grained machine translation evaluation, 2023. [12] Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. Gptscore: Evaluate as you desire, 2023. [13] Marjan Ghazvininejad, Hila Gonen, and Luke Zettlemoyer. Dictionary-based phrase-level prompting of large language models for machine translation, 2023. [14] Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzm´an, and Angela Fan. The Flores-101 evaluation benchmark for low-resource and multilingual machine translation. Transactions of the Association for Computational Linguistics, 10:522–538, 2022. [15] Soufiane Hayou, Nikhil Ghosh, and Bin Yu. Lora+: Efficient low rank adaptation of large models, 2024. [16] Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for NLP. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 2790–2799. PMLR, 09–15 Jun 2019. 109 [17] Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022. [18] Vivek Iyer, Pinzhen Chen, and Alexandra Birch. Towards effective disambiguation for machine translation with large language models, 2023. [19] Alexandre Lacoste, Alexandra Luccioni, Victor Schmidt, and Thomas Dandres. Quantifying the carbon emissions of machine learning, 2019. [20] Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning, 2021. [21] Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation, 2021. [22] Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, and Xian Li. Few-shot learning with multilingual language models, 2022. [23] Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, KwangTing Cheng, and Min-Hung Chen. Dora: Weight-decomposed low-rank adaptation, 2024. [24] Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. Gpt understands, too, 2023. [25] Hongyuan Lu, Haoran Yang, Haoyang Huang, Dongdong Zhang, Wai Lam, and Furu Wei. Chainof-dictionary prompting elicits translation in large language models, 2024. [26] Zhuoyuan Mao and Yen Yu. Tuning llms with contrastive alignment instructions for machine translation in unseen, low-resource languages, 2024. [27] Yasmin Moslem, Rejwanul Haque, John D. Kelleher, and Andy Way. Adaptive machine translation with large language models. In Mary Nurminen, Judith Brenner, Maarit Koponen, Sirkku Latomaa, Mikhail Mikhailov, Frederike Schierl, Tharindu Ranasinghe, Eva Vanmassenhove, Sergi Alvarez Vidal, Nora Aranberri, Mara Nunziatini, Carla Parra Escart´ın, Mikel Forcada, Maja Popovic, Carolina Scarton, and Helena Moniz, editors, Proceedings of the 24th Annual Conference of the European Association for Machine Translation, pages 227–237, Tampere, Finland, June 2023. European Association for Machine Translation. [28] Yasmin Moslem, Gianfranco Romani, Mahdi Molaei, John D. Kelleher, Rejwanul Haque, and Andy Way. Domain terminology integration into machine translation: Leveraging large language models. In Philipp Koehn, Barry Haddow, Tom Kocmi, and Christof Monz, editors, Proceedings of the Eighth Conference on Machine Translation, pages 902–911, Singapore, December 2023. Association for Computational Linguistics. [29] El Moatez Billah Nagoudi, AbdelRahim Elmadany, and Muhammad Abdul-Mageed. Arat5: Textto-text transformers for arabic language generation, 2022. [30] Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian. A comprehensive overview of large language models, 2024. [31] Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019. [32] Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer, 2023. [33] Vikas Raunak, Amr Sharaf, Yiren Wang, Hany Awadalla, and Arul Menezes. Leveraging GPT-4 for automatic translation post-editing. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Findings of the Association for Computational Linguistics: EMNLP 2023, pages 12009–12024, Singapore, December 2023. Association for Computational Linguistics. 110 [34] Abudurexiti Reheman, Tao Zhou, Yingfeng Luo, Di Yang, Tong Xiao, and Jingbo Zhu. Prompting neural machine translation with translation memories, 2023. [35] Raphael Reinauer, Patrick Simianer, Kaden Uhlig, Johannes E. M. Mosig, and Joern Wuebker. Neural machine translation models can learn to be few-shot learners, 2023. [36] Gabriele Sarti, Phu Mon Htut, Xing Niu, Benjamin Hsu, Anna Currey, Georgiana Dinu, and Maria Nadejde. RAMP: Retrieval and attribute-marking enhanced prompting for attribute-controlled translation. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1476–1490, Toronto, Canada, July 2023. Association for Computational Linguistics. [37] Mike Schuster and Kaisuke Nakajima. Japanese and korean voice search. In 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5149–5152, 2012. [38] NLLB Team, Marta R. Costa-juss`a, James Cross, Onur C¸elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzm´an, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. No language left behind: Scaling humancentered machine translation, 2022. [39] Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models, 2023. [40] Mojtaba Valipour, Mehdi Rezagholizadeh, Ivan Kobyzev, and Ali Ghodsi. Dylora: Parameter efficient tuning of pre-trained models using dynamic search-free low-rank adaptation, 2023. [41] Yuhui Xu, Lingxi Xie, Xiaotao Gu, Xin Chen, Heng Chang, Hengheng Zhang, Zhengsu Chen, Xiaopeng Zhang, and Qi Tian. Qa-lora: Quantization-aware low-rank adaptation of large language models, 2023. [42] Wen Yang, Chong Li, Jiajun Zhang, and Chengqing Zong. Bigtranslate: Augmenting large language models with multilingual translation capability over 100 languages, 2023. [43] Biao Zhang, Philip Williams, Ivan Titov, and Rico Sennrich. Improving massively multilingual neural machine translation and zero-shot translation, 2020. [44] Shaolei Zhang, Qingkai Fang, Zhuocheng Zhang, Zhengrui Ma, Yan Zhou, Langlin Huang, Mengyu Bu, Shangtong Gui, Yunji Chen, Xilin Chen, and Yang Feng. Bayling: Bridging cross-lingual alignment and instruction following through interactive translation for large language models, 2023. [45] Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li. Multilingual machine translation with large language models: Empirical results and analysis, 2024. [46] Wenhao Zhu, Yunzhe Lv, Qingxiu Dong, Fei Yuan, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li. Extrapolating large language models to non-english by aligning languages, 2023. 111 [47] Micha l Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. The United Nations parallel corpus v1.0. In Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, and Stelios Piperidis, editors, Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 3530–3534, Portoroˇz, Slovenia, May 2016. European Language Resources Association (ELRA). 112 4 Our Contribution Our main objective is to propose an algorithm, a method, and a tool to facilitate the evaluation of learners’ collaborative behavior in a collaborative learning environment. To achieve this goal, our contribution will be divided into three main parts: 4.1 The first contribution The first part of our contribution involves proposing an algorithm for extracting collaborative episodes from a trace. The application of this algorithm will extract collaborative fragments within an interaction trace. This algorithm will be:  The first tool that teachers can use to detect the collaborative behavior of their learners, and  Supported by statistical and mining functions to detect specific aspects of the trace, such as extracting frequent collaborative episodes, the number of frequent sequential episodes, etc. The collaborative fragment extraction algorithm we propose is based on the frequent sequential episode extraction algorithm, with the key difference that the result is a ”collaborative sequential episode.” Let’s take the following example:  Let the following trace be:  Let the collaboration obsels be:  Applying the frequent sequential episode extraction algorithm, we obtain the following frequent episode with a seuil greater than or equal to 2:  Applying our method to propose, we’ll obtain the following episodes with a length greater than or equal to 2:  The calculation of collaboration indicators will be based on these results. 4.2 The Second Contribution The second part of our contribution involves proposing an approach that enables non-computer-scientist teachers to design and calculate their own collaboration indicators to detect and evaluate learners’ collaborative behavior in collaborative e-learning environments. Our approach will be based on a model-driven architecture, where the calculation of indicators can be seen as a series of model transformations. After the design stage, we aim to automate the calculation of the teacher-designed indicators by automatically generating the transformation sequence needed to obtain the corresponding indicator model. 4.3 The Third Contribution The third part consists of proposing a real case study using a learning platform with an online collaborative learning situation and an evaluation grid to implement and adjust the proposed algorithms and methods. This part is organized as follows: 1. Proposing a collaborative learning situation on a learning platform to collect the necessary data interaction traces during collaborative learning sessions. 119 2. Proposing a model for the concept of ’collaborative trace’ within the trace based system’KTBS’, which will serve as an ontology for validating trace models. 3. Developing operators that facilitate the construction of collaborative traces, including a confidence factor. 4. Implementing the evaluation algorithms and integrating several proposed indicators to assess the collaborative learning process (Figure 2 summarizes these steps). Figure 2: Collaboration evaluation stage in a collaborative learning situation. 5 Conclusion Collaborative distance learning represents a valuable approach for enhancing group work and developing collaborative skills. After years of research, the interactions between learners using various distance learning tools can be captured through a computer object called M-Trace. In this research, we focused on evaluating collaborative practices within a learning activity. Our first contribution is the development of an algorithm to detect and extract collaborative fragments from interaction traces. By adapting the principle of frequent sequential episode extraction, we propose a novel algorithm tailored to extract sequential episodes of collaboration. The second contribution lies in evaluating collaborative practices. We propose a Model-Driven Architecture (MDA)-based method to calculate collaboration indicators, coupled with a tool that enables non-computer-scientist teachers to design meaningful collaboration indicators. This will allow teachers to generate the transformation sequence necessary to obtain the corresponding collaboration indicator model. Finally, we plan to conduct a case study to further refine and adjust the evaluation algorithms. This case study will help validate the proposed methods and tools, ensuring their practical applicability in real-world educational settings. References [1] A. Acosta and J. Lee. Multimodal learning analytics for predicting student collaboration satisfaction. In Proceedings of the 17th International Conference on Educational Data Mining (EDM 2024), 2024. [2] M. Arnaud. Current limitations of online collaborative learning (in french). STICEF (Information and Communication Sciences and Technologies for Education and Training), 10:7, 2003. 120 [3] T. Bratitsis and A. Dimitracopoulou. Monitoring and analyzing group interactions in asynchronous discussions with the dias system. In Y. A. Dimitriadis, I. Zigurs, and E. G´omez-S´anchez, editors, Groupware: Design, Implementation, and Use, volume 4154, pages 54–61. Springer Berlin Heidelberg, 2006. [4] C. Choquet and S. Iksal. Modeling and constructing usage traces of a learning activity: A language approach for reengineering a cehl (in french). 2007. [5] C. A. Collazos, L. A. Guerrero, J. A. Pino, S. Renzi, J. Klobas, M. Ortega, M. A. Redondo, and C. Bravo. Evaluating collaborative learning processes using system-based measurement. 18, 2007. [6] T. Djouad, A. Mille, C. Reffay, and M. Benmohamed. Engineering of activity indicators from modeled traces in a computer environment for human learning (in french). 16(1):103–139, 2009. [7] A. El Mhouti, M. Erradi, and N. El Makhfi. A multi-agent system of semantic analysis and filtering of modeled traces to calculate interaction indicators favoring collaboration in lms. In International Conference on Intelligent Systems and Advanced Computing Sciences (ISACS), pages 1–7, 2019. [8] C. Gangloff-Ziegler. The obstacles of collaborative work, steps and organizations. 10(3):95–112, 2009. [9] S. Garc´ıa-Sastre, A. Mart´ınez-Mon´es, and Y. Dimitriadis. Detecting collaboration skills to calculate indicators in moodle. In Proceedings of the 5th International Conference on Technological Ecosystems for Enhancing Multiculturality, pages 1–8. ACM, 2017. [10] ´ E. Gendron. A conceptual framework for the development of collaboration indicators from activity traces (In French). PhD thesis, Claude Bernard UniversityLyon I, 2010. [11] F. Henri. Collaborative learning in virtual mode (fr). 2002. [12] ICALT. Proceedings of the ieee international conference on advanced learning technologies. 2004. [13] D. Laurillard, O. Martin, W. Barbara, and H. Ulrich. Implementing Technology-Enhanced Learning. 2009. [14] I. Matazi, A. Bennane, R. Messoussi, R. Touahni, I. Oumaira, and R. Korchiyne. Multi-agent system based on fuzzy logic for e-learning collaborative system. In International Symposium on Advanced Electrical and Communication Technologies (ISAECT), pages 1–7, 2018. [15] D. P. T. Ngoc. Specification and design of analysis services for the use of a computer environment for human learning (in french). Technical report, 2004. [16] D. Nguyen, S. Yingchareonthawornchai, V. Tekken Valapil, and S. S. Kulkarni. Precision, recall, and sensitivity of monitoring partially synchronous distributed programs. Distributed Computing, 34(4):319–348, 2021. [17] S. PraharaJ, M. Scheffel, and M. et al. Schmitz. Towards automatic collaboration analytics for group speech data using learning analytics. Sensors, 21(9):3156, 2021. [18] G. M. Rafique and M. N. A. Khan. Integrating learning analytics and collaborative learning for improving student’s academic performance. International Journal of Advanced Computer Science and Applications, 12(12), 2021. [19] N. Randriamalaka, S. Iksal, and C. Choquet. Elicitation of indicators for the re-engineering of educational scenarios: A trace-based approach using utl (in french). 2008. [20] S. A. Salloum, M. Al-Emran, K. Shaalan, and A. Tarhini. Factors affecting the e-learning acceptance: A case study from uae. Education and Information Technologies, 2019. [21] L. S. Settouti. Trace-based systems for human learning (in french). 2006. [22] M. V´asquez-Berm´udez, J. Hidalgo-Larrea, F. Orozco Lara, and S. Segura Santana. Effectiveness of monitoring indicators in the architecture of a collaborative system. In Technologies and Innovation, pages 191–202. Springer, Cham, 2022. 121 IoT Applications in the Education Sector: Architectures, Challenges, and Emerging Paradigms Hicham Medkour*, Mawloud Belabbas , Bachir Rahmi and Kada Becharef Educative Technology Division, National Institute for Research in Education Oued Roumane, El-Achour, Algiers-Algeria *Corresponding author: [email protected] Abstract The rapid integration of Internet of Things (IoT) technology in educational environments is revolutionizing traditional pedagogies and institutional operations. This paper explores IoT’s role in augmenting educational delivery, enhancing resource management, and enabling personalized learning through interconnected sensor-based infrastructures. It critically evaluates real-world deployments, security and privacy implications, and future prospects within the smart education paradigm. A multi-layered IoT architecture is proposed, and recommendations for sustainable adoption are discussed. Keywords: IA, RNN.Internet of Things, Education, Smart Learning, Adaptive Systems, Ubiquitous Learning, Edge Computing, Data Privacy. 1 Introduction The Internet of Things (IoT) is experiencing growing adoption in the education sector, transforming traditional learning environments into intelligent, interactive, and personalized ecosystems. Educational IoT relies on an interconnected network of sensors, smart devices, and software platforms that collect and process real-time data to optimize teaching and administrative processes. In a global context marked by inequalities in access to education and accelerated digitalization, leveraging these technologies represents a major opportunity to foster inclusivity, learner motivation, and academic performance. Recent studies have explored the potential of IoT in education [1]. Among them, [2] demonstrates how a pilot project at a Malaysian university led to a 23 percent improvement in energy efficiency while enhancing student engagement through behavior-tracking sensors. Other studies [3], [4], [5] highlight the emergence of digital twins, artificial emotional intelligence, and intelligent tutoring systems as drivers of pedagogical transformation. However, despite these advances, large-scale deployments of educational IoT remain limited and often experimental. This study identifies a scientific gap in the systemic understanding of the conditions for success, technical, ethical, and pedagogical challenges, and the evaluation criteria applicable to IoT projects in higher education. The central issue revolves around the sustainable, inclusive, and ethical integration of IoT technologies in educational institutions: How can we design and evaluate an educational IoT architecture that is both efficient, ethical, and adaptable to diverse pedagogical contexts? To address this issue, an analytical and comparative approach has been adopted. This work is based on a structured review of scientific literature, a critical analysis of international case studies, and the application of multi-dimensional evaluation frameworks. Special attention is given to issues of security, data governance, standardization, and scalability in resource-constrained contexts. The primary objective of this study is to provide an in-depth synthesis of IoT architectures, applications, and emerging trends in education, while identifying the technical, ethical, and organizational barriers that need to be overcome. Through this approach, the study aims to enlighten academic decision-makers, instructional engineers, and researchers on best practices for the design, deployment, and evaluation of connected educational solutions. 2 IoT Architecture for Smart Education 2.1 Layered Design The architecture of the Internet of Things (IoT) in educational environments is commonly structured into a layered model to streamline data flow and system interaction. This model generally consists of three core layers: the perception layer, the network layer, and the application layer [6]. 122 2.1.1 Perception Layer The perception layer serves as the foundational component of the IoT architecture. It includes various smart sensing devices such as Near Field Communication (NFC) tags, Radio Frequency Identification (RFID) systems, sensors, and cameras. These devices are responsible for gathering real-time data on a wide range of educational metrics, including student attendance, physical movement, classroom environmental conditions (like temperature and light), and device usage patterns. By collecting this data at the source, the perception layer enables real-time monitoring and situational awareness within the educational context. This visibility is instrumental in supporting responsive and adaptive educational services that align with learners’ needs. 2.1.2 Network Layer Next, the network layer is responsible for the secure and efficient transmission of the collected data from the perception layer to the application layer. This layer utilizes multiple communication protocols, including Zigbee, Wi-Fi, 5G, and LoRaWAN, depending on the specific network requirements and constraints of the educational institution. The network layer not only ensures data transfer but also addresses critical aspects such as latency, data packet loss, and bandwidth optimization. Secure communication is prioritized through encryption and tunneling techniques, minimizing the risks associated with data breaches or interception. 2.1.3 Application Layer At the top of the stack, the application layer translates raw data into actionable insights tailored for educators, administrative staff, and learners. This layer often employs edge computing or cloudbased analytics to process and visualize data in user-friendly formats. Integration with existing Learning Management Systems (LMS) is common, enabling a comprehensive digital learning ecosystem where decision-making and educational customization are driven by real-time insights. Dashboards, reporting tools, and AI-driven recommendation engines fall under this layer, offering tailored feedback to improve both teaching and learning processes. 2.2 Interoperability and Middleware Given the diverse and often incompatible nature of IoT devices and systems used in education, ensuring interoperability is a significant challenge. Middleware platforms such as FIWARE and Kaa play a pivotal role in addressing this issue [7]. These platforms act as intermediaries that harmonize communication between heterogeneous devices and educational software applications. Middleware provides standardized interfaces and APIs, which enable developers and administrators to integrate new hardware and software without disrupting existing systems. In addition to facilitating interoperability, middleware solutions enable context-awareness by understanding and adapting to the educational environment. For instance, middleware can interpret contextual signals like user behavior patterns or environmental changes, helping the system respond dynamically to varying scenarios. Furthermore, orchestration capabilities embedded in middleware platforms allow for the automated coordination of processes, including device synchronization, data fusion, and event-driven responses. As a result, middleware enhances the flexibility, scalability, and reliability of IoT implementations in educational settings, paving the way for seamless user experiences and efficient system performance. 3 Key Applications of IoT in Education 3.1 Smart Classrooms Smart classrooms represent one of the most tangible and transformative applications of IoT in education. These environments utilize a range of interconnected devices such as environmental sensors, AI-powered dashboards, interactive displays, and intelligent control systems for lighting and air conditioning. By continuously monitoring variables such as room temperature, humidity, noise levels, and lighting, these systems help maintain optimal learning conditions. Moreover, AI-driven analytics tools provide instructors with real-time insights into student engagement and classroom dynamics. Teachers can adapt their instructional strategies on the fly—modifying 123 pacing, introducing new materials, or adjusting classroom configurations based on data analytics [8]. Ultimately, smart classrooms enhance interactivity, engagement, and learner outcomes by creating an environment that is both responsive and student-centered. 3.2 Attendance and Identity Verification Traditional methods of tracking student attendance—manual roll calls or sign-in sheets are timeconsuming and prone to errors. IoT technologies like RFID and biometric authentication systems revolutionize this process by automating attendance tracking. Students equipped with RFID-enabled ID cards or biometric markers (e.g., fingerprint or facial recognition) are identified upon entering the classroom, and their presence is recorded in real-time [9]. This automation streamlines administrative tasks and allows teachers to focus on instructional duties. Additionally, it enhances data accuracy, supports behavioral analytics, and contributes to the development of personalized educational pathways. Integration with centralized school databases ensures seamless updating of attendance records and the generation of performance and behavior reports for students and parents. 3.3 Adaptive Learning and Wearables Wearable IoT devices, such as smartwatches, biometric bands, and augmented reality (AR) headsets, are increasingly used to monitor learners’ physiological and emotional states. These devices can track parameters such as heart rate variability, galvanic skin response, and motion patterns to infer cognitive load and emotional engagement [10]. By feeding this data into adaptive learning systems, educational platforms can dynamically adjust content difficulty, format, and delivery methods to match each learner’s needs and current state. For example, if a student’s stress indicators are elevated, the system might recommend a break or switch to a less cognitively demanding task. This personalization fosters more effective learning and helps mitigate stress and burnout. Teachers also benefit from detailed analytics on student engagement trends, enabling timely interventions and improved learner support. 3.4 Facility and Asset Management Educational institutions manage a wide range of physical assets—from classroom equipment to campus infrastructure. IoT sensors embedded in furniture, audio-visual (AV) equipment, and utility systems can provide continuous updates on usage patterns, operational status, and maintenance needs [11]. Real-time monitoring supports efficient allocation of resources and prevents downtime by enabling predictive maintenance. For instance, a sensor-equipped projector may notify administrators of a potential malfunction before it occurs, allowing for timely intervention. Furthermore, energy consumption data gathered from HVAC systems or lighting fixtures can inform sustainability strategies, leading to reduced operational costs and environmental impact. 3.5 Inclusive and Remote Learning IoT technologies are instrumental in promoting inclusivity and accessibility in education. Assistive devices such as Braille-enabled e-readers, hearing aids linked to classroom audio systems, and voicecontrolled learning applications ensure that students with disabilities have equitable access to educational content [12]. Additionally, IoT-enabled remote learning solutions provide consistent and immersive experiences for students learning outside the traditional classroom. Smart conferencing devices, AI-powered tutoring systems, and collaborative platforms facilitate engagement and maintain the continuity of instruction. These tools became particularly vital during the COVID-19 pandemic and continue to support hybrid and distance learning models. 4 Security and Privacy Considerations While IoT offers substantial benefits to the educational sector, it also introduces significant ethical, privacy, and security challenges that must be addressed to ensure safe and responsible deployment. 124 4.1 Data Privacy Risks The use of IoT in education involves the collection of sensitive student data, including location information, biometric identifiers, academic performance, and behavioral patterns. These data points are vulnerable to unauthorized access, theft, or misuse if not properly safeguarded. To address these concerns, institutions must adhere to data protection regulations such as the General Data Protection Regulation (GDPR) in Europe and the Family Educational Rights and Privacy Act (FERPA) in the United States [13]. Data anonymization, encryption, and strict access control mechanisms are essential to protect personal information. Additionally, transparency in data collection practices and obtaining informed consent from students and guardians are critical steps toward ethical IoT usage. 4.2 Attack Surfaces The proliferation of interconnected devices in educational environments increases the potential attack surface for malicious actors. Common threats include Distributed Denial-of-Service (DDoS) attacks, spoofing, unauthorized access, and malware infiltration. These threats are often exacerbated by inadequate encryption, unpatched firmware, and the use of outdated devices [14]. Attackers can exploit vulnerabilities to disrupt learning activities, steal confidential data, or gain control over institutional infrastructure. As such, a proactive approach to cybersecurity—including regular updates, threat monitoring, and penetration testing—is vital to safeguarding IoT systems. 4.3 Mitigation Strategies To address these vulnerabilities, educational institutions can implement a range of mitigation strategies aimed at enhancing security and trust in IoT deployments. Key approaches include: - Lightweight Cryptography: Suitable for resource-constrained IoT devices, lightweight cryptographic algorithms ensure data confidentiality and integrity without overloading system resources. - Network Segmentation via SDN: Software Defined Networking (SDN) allows for dynamic network segmentation, reducing the spread of attacks and enabling granular access control. - Blockchain-Based Audit Trails: Blockchain technology can be employed to create immutable logs of data access and transactions, promoting transparency and accountability [15]. By adopting a security-by-design philosophy and continuously evaluating emerging threats, institutions can harness the full potential of IoT while maintaining ethical standards and legal compliance in educational contexts. 5 Critical Evaluation of Case Studies A comparative analysis of IoT deployments across various universities reveals several key success factors that shape the effectiveness of these systems. Firstly, the integration of Localized Edge Computing significantly reduces latency, ensuring realtime responsiveness in applications such as behavioral feedback systems and adaptive learning tools. Rather than sending all data to a centralized cloud, local edge devices process and respond to data closer to the source, minimizing delays and conserving bandwidth. Secondly, faculty training in IoT ethics and data governance emerges as a critical component. Successful IoT adoption hinges not only on technological readiness but also on the awareness and ethical responsibility of educators. Institutions that incorporate data governance training and ethics workshops are better equipped to manage privacy concerns and ensure equitable student treatment. Thirdly, universities are increasingly implementing hybrid infrastructures that combine cloud and fog computing. This architecture helps balance the scalability and computational power of cloud platforms with the responsiveness and localized processing offered by fog nodes. Such systems allow for seamless data management while upholding student privacy and adapting to network constraints. In [15], a notable case study conducted at a Malaysian university showcased the practical benefits of such strategies. The pilot project implemented sensor-based behavioral feedback mechanisms to encourage energy-saving behaviors among students. As a result, energy efficiency improved by 23 percent. Furthermore, the interactive feedback loop—facilitated by IoT devices and dashboards—boosted student 125 participation, underscoring how well-designed IoT systems can influence not only operational metrics but also learner engagement and institutional culture. 6 Emerging Trends and Out-of-the-Box Applications 6.1 Cognitive IoT and AI The convergence of Artificial Intelligence (AI) and IoT, commonly referred to as AIoT, is driving innovation in personalized learning. Intelligent tutoring systems now incorporate AI algorithms—particularly reinforcement learning—to tailor quiz difficulty and content delivery based on real-time assessments of student performance and behavior. These systems dynamically adapt to individual learning curves, offering more precise and motivating educational experiences. 6.2 Digital Twins for Learning A particularly novel development is the use of digital twins—virtual counterparts of physical learning environments. These digital replicas, enabled by IoT and immersive technologies such as Augmented Reality (AR) and Virtual Reality (VR), allow remote learners to engage with classroom environments in real-time. Students can interact with virtual lab equipment, observe real-world classroom dynamics, or even participate in collaborative activities via avatars. This concept redefines remote learning by making it more experiential and presence-driven [2]. 6.3 Gamification and Emotional AI Another frontier in educational IoT is the blending of gamification techniques with emotional AI. IoT devices such as smart cameras or wearable sensors can detect emotional cues like facial expressions, vocal tone, or physiological stress markers. These inputs feed into gamified learning systems that adjust tasks, difficulty levels, or feedback styles to maintain engagement and motivation. For instance, if a student appears frustrated, the system might reduce task complexity or offer encouraging prompts. The combination of real-time emotion recognition and motivational design principles fosters deeper learner immersion and satisfaction [3]. 7 Methodological Framework for IoT Evaluation in Education A robust evaluation of IoT deployments in education requires a multi-dimensional framework that integrates both technical performance and educational outcomes. Four primary metric categories are recommended: •Technical KPIs: Metrics such as latency, throughput, and device uptime assess the operational efficiency of IoT systems. •Pedagogical Metrics: These include indicators like student engagement rates, retention, learning gains, and academic performance improvements. •Usability Metrics: These gauge system intuitiveness, ease of navigation, and user satisfaction, particularly among faculty and students. •Socio-Ethical Metrics: Inclusivity, equity, transparency in data use, and adherence to ethical standards fall under this category. To gather these metrics effectively, a mixed-method research approach is advised. Quantitative data from system logs and user analytics can be complemented by qualitative feedback collected via surveys, interviews, and classroom observations. Longitudinal studies, in particular, are crucial to understanding the sustained impact of IoT systems on learning outcomes and educational equity [4]. 8 Challenges and Future Directions 8.1 Scalability vs. Affordability One of the primary barriers to widespread IoT adoption in education is the cost of deployment— particularly in low-resource settings. To address this, institutions are turning to cost-effective, opensource hardware platforms such as Raspberry Pi, Arduino, and ESP32. These devices support 126 essential IoT functionalities while minimizing financial strain. Moreover, optimized firmware and modular architectures can extend device lifespans and reduce maintenance requirements. 8.2 Standardization Another significant challenge is the lack of unified standards for educational IoT systems. This hampers interoperability, making it difficult for institutions to integrate heterogeneous devices and platforms. Initiatives such as oneM2M and IEEE P2413 are working toward global IoT standards, but adaptation for academic use remains in early stages [16]. Collaborative efforts among universities, industry, and standards bodies are necessary to accelerate this process and promote compatibility 8.3 Ethical Pedagogical Design Finally, as educational technologies become increasingly data-driven, there is a growing need for ethical pedagogical frameworks. IoT-based learning systems must not only be efficient but also aligned with human-centered values. This includes ensuring informed consent, fostering inclusive design, andresisting over-surveillance. Educators and technologists should co-create curricula and systems that prioritize students’ well-being, autonomy, and equitable access to learning opportunities. 9 Conclusion The Internet of Things holds transformative potential for the education sector by enabling smart, responsive, and personalized learning environments. However, its integration is not without challenges. Success depends on well-designed system architecture, robust ethical governance, and collaborative innovation among educators, technologists, and policymakers. As educational institutions embrace IoT, they must remain vigilant about issues of equity, privacy, and long-term sustainability. Future research should focus on human-centered design, ethical standards, and scalable frameworks to ensure IoT serves as a tool for inclusive and impactful education worldwide. References [1] A. M. Alghamdi, ”IoT-Based Smart Classrooms: A Systematic Review,” IEEE Access, vol. 10, pp. 95011–95030, 2022. [2] F. Rahim et al., ”Smart Campus Pilot in Malaysia: An IoT Energy Optimization Study,” IEEE Access, vol. 8, pp. 115624–115638, 2020. [3] H. Li and W. Wang, ”Digital Twins and IoT in Immersive Education,” IEEE Transactions on Industrial Informatics, vol. 17, no. 6, pp. 4182–4191, Jun. 2021. [4] D. G. L´opez et al., ”Gamified Learning with IoT Feedback Systems,” IEEE Transactions on Learning Technologies, vol. 14, no. 1, pp. 88–98, Jan.–Mar. 2021. [5] S. A. Kazi et al., ”Mixed-Method Evaluation of IoT-Enabled Learning Spaces,” Computers and Education, vol. 161, p. 104064, 2021. [7] M. S. Hossain et al., ”IoT in Education: A Framework for Smart Learning Environment,” IEEE Internet of Things Journal, vol. 7, no. 8, pp. 7103–7112, Aug. 2020. [8] K. P. Tripathi et al., ”Middleware for IoT in Education: A Survey,” IEEE Communications Surveys and Tutorials, vol. 23, no. 2, pp. 1431–1457, 2021. [9] H. A. Khattak et al., ”Smart Learning Environments Using IoT and Context-Aware Systems,” Computers in Human Behavior, vol. 112, p. 106481, Jan. 2020. [10] N. D. Patel et al., ”RFID-Based Attendance Monitoring in Higher Education,” International Journal of Engineering Research and Technology, vol. 9, no. 5, pp. 1023–1027, 2021. [11] J. W. Kim et al., ”Wearable IoT in Education: A Smart Glove for Learning Sign Language,” Sensors, vol. 20, no. 11, p. 3132, 2020. [12] Y. Zhang and X. Zhao, ”IoT Asset Management in Smart Campuses,” IEEE Sensors Journal, vol. 21, no. 15, pp. 16851–16860, 2021. [13] B. L. Perry et al., ”IoT for Inclusive Learning: A Review,” IEEE Transactions on Learning Technologies, vol. 13, no. 4, pp. 723–734, Oct. 2020. [14] G. Spanakis et al., ”Privacy Challenges in IoT-Enabled Education,” IEEE Security and Privacy, vol. 127 18, no. 6, pp. 33–41, 2021 [15] A. Shahrbabaki et al., ”Security in IoT for Education: A Review of Threats and Mitigations,” IEEE Communications Magazine, vol. 59, no. 2, pp. 76–81, 2021. [16] A. M. Rahmani et al., ”Security Solutions for Smart Learning Systems,” IEEE Systems Journal, vol. 14, no. 3, pp. 3676–3687, 2020. [16] M. R. Islam et al., ”IoT Standards and Frameworks for Education,” IEEE Standards Magazine, vol. 3, no. 3, pp. 24–31, Sept. 2021. 128 scalability, and user acceptance, which can undermine trust. Furthermore, approaches focusing on trustworthiness, such as those proposed by Lipizzi et al. [13], underscore the necessity of integrating human expertise and domain-specific knowledge to assess and enhance the reliability of LLMs. Despite these strengths, challenges persist, including knowledge noise within KGs, which can lead to inaccuracies in model outputs, as noted by Yang et al. This highlights the urgent need for robust filtering mechanisms and dynamic updating processes to ensure that KGs remain relevant and accurate. Moreover, the computational complexity associated with integrating KGs with LLM architectures raises concerns about scalability and real-time application, necessitating more efficient methodologies. Addressing the subjectivity involved in trust assessments is equally critical, as inconsistent evaluations may hinder the applicability of these frameworks across diverse contexts. Therefore, future research should prioritize the development of standardized metrics for trust evaluation, optimized integration algorithms, and cross-domain applications of KG-LLM frameworks. By tackling these limitations, researchers can unlock the full potential of integrating KGs with LLMs, paving the way for more reliable, transparent, and contextually aware AI systems. 6 Challenges and future directions 6.1 Challenges While integrating Knowledge Graphs (KGs) with Large Language Models (LLMs) offers potential benefits, several significant challenges persist. Knowledge noise within KGs can lead to inaccuracies in LLM outputs, necessitating robust filtering mechanisms [24], while the complexity of integrating KGs with LLM architectures can introduce substantial computational demands [13]. Accurately capturing the intricacies of a domain in a KG is challenging, as nuanced relationships and exceptions may result in oversimplifications or inaccuracies. Moreover, KGs are often incomplete, leading to gaps in the knowledge accessible to LLMs, which can produce misleading outputs. The dynamic nature of knowledge necessitates continuous updating of KGs to avoid outdated conclusions, a resource-intensive process. Integration complexity arises when aligning structured data from KGs with unstructured data processed by LLMs, requiring sophisticated methods for effective utilization. Additionally, scalability issues become apparent as the size and complexity of KGs increase, complicating maintenance and validation processes. Subjectivity in knowledge selection can lead to inconsistencies in representation, and biases in the underlying data can undermine the trustworthiness of LLM outputs. Variability in knowledge quality further exacerbates trust issues, as inaccurate or outdated information can result in erroneous conclusions. Interpretability challenges emerge when attempting to understand how KGs influence LLM decision-making, compounded by limitations in human validation due to the availability of subject matter experts. Furthermore, the computational overhead of querying and analyzing KGs may impact the efficiency of trust assessments, and user skepticism regarding the reliability of KGs can hinder acceptance, particularly when users are unfamiliar with the methodologies employed in their construction. Finally, interoperability issues can complicate the integration of KGs built using different standards and formats, posing challenges for comprehensive trust assessments across various domains. In summary, while the integration of KGs with LLMs holds promise for enhancing trustworthiness, addressing these multifaceted challenges is critical for achieving effective and ethical outcomes.[15,16,1] 6.2 Future Directions Future directions for using Large Language Models (LLMs) in conjunction with Knowledge Graphs (KGs) to enhance trustworthiness can focus on several key areas. Developing methods for creating and maintaining dynamic knowledge graphs that can automatically update in response to new information, research findings, or changes in domain knowledge is essential. This could involve leveraging real-time data sources and machine-learning techniques to ensure that the knowledge graph remains current and relevant. Additionally, improving the integration of LLM outputs with knowledge graphs through advanced natural language processing techniques could enhance the alignment of LLM-generated content with the structured data in KGs, enabling more accurate assessments of trustworthiness based on contextual relevance. Furthermore, exploring automated or semi-automated validation processes for knowledge graphs, potentially using machine learning algorithms to identify inconsistencies or gaps in the knowledge representation, could reduce reliance on human evaluators and enhance scalability. Encouraging collaboration between domain experts, data scientists, and AI researchers is vital to creating more robust knowledge graphs that accurately reflect the complexities of various fields. This interdisciplinary 135 approach can help ensure that the knowledge represented is comprehensive and trustworthy. Developing user-centric metrics for trust that take into account individual user needs, preferences, and contexts can also enhance the trustworthiness of LLM outputs. Focusing on enhancing the explainability of LLM outputs about the knowledge graph by providing users with clear explanations of how LLM responses are derived from the knowledge graph can build trust through transparency. Moreover, creating crossdomain knowledge graphs that can integrate information from multiple fields allows LLMs to provide more comprehensive and contextually aware responses. This could enhance the trustworthiness of outputs in interdisciplinary applications. Addressing ethical considerations related to trust in LLMs and KGs, including the identification and mitigation of biases in both the knowledge representation and the model outputs, is crucial for broader acceptance. Conducting extensive real-world testing of LLMs combined with knowledge graphs in various applications, such as healthcare, finance, and education, is necessary. Gathering empirical data on their performance and trustworthiness can inform further improvements and refinements. Finally, encouraging community contributions to knowledge graphs, allowing users to add, edit, and validate information, can enhance the richness and accuracy of the knowledge represented, fostering a sense of ownership and trust among users. By pursuing these future directions, the integration of LLMs and knowledge graphs can lead to more reliable, trustworthy, and user-friendly systems that effectively support decision-making across various domains. 7 Conclusion In this paper, we explored the integration of Large Language Models (LLMs) with Knowledge Graphs (KGs) to enhance the trustworthiness of information generated in various domains. We established that while LLMs demonstrate impressive capabilities in natural language processing tasks, their outputs can be limited by biases and inaccuracies inherent in the training data. By combining LLMs with KGs, we can leverage the structured, semantically rich information contained within knowledge graphs to improve the reliability and contextual relevance of LLM-generated content. Our investigation highlighted several promising future directions, including the dynamic updating of knowledge graphs, the enhancement of natural language processing techniques for better integration, and the development of user-centric trust metrics. Additionally, we emphasized the importance of interdisciplinary collaboration and ethical considerations in the deployment of these integrated systems. Through extensive empirical testing and community engagement, we aim to create more robust and trustworthy systems that effectively support decision-making across various fields. The findings of this study pave the way for future research aimed at bridging the gap between LLMs and KGs, ultimately fostering trust and improving the quality of information accessible to users. References [1] Emad A. Alghamdi, Reem I. Masoud, Deema Alnuhait, Afnan Alomairi, Ahmed Ashraf, and Mohamed Zaytoon. Aratrust: An evaluation of trustworthiness for llms in arabic. In Proceedings of the Arabic Language and AI Conference, 2023. [2] Reuben Binns. Fairness in machine learning: Lessons from political philosophy. In Proceedings of the Conference on Fairness, Accountability, and Transparency, pages 149–159, 2018. [3] Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, P. Dhariwal, and D. Amodei. Language models are few-shot learners. In Advances in Neural Information Processing Systems, 2020. [4] Erik Cambria, Soujanya Poria, Alexander Gelbukh, and Awais Hussain. Xai meets llms: A survey of the relation between explainable ai and large language models. arXiv preprint arXiv:2407.15248, 2024. [5] Xiaojun Chen, Shengbin Jia, and Yang Xiang. A review: Knowledge reasoning over knowledge graph. Expert Systems with Applications, 141:112948, 2020. [6] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2018. 136 [7] Kaito Fujiwara, Katsuya Nakamura, Jun Matsui, Tatsuya Matsumoto, and Masahiro Hara. Measuring the interpretability and explainability of model decisions of five large language models. arXiv preprint, 2024. [8] Jorge Garnica, Ana Vega, and Juan Guti´errez. Knowledge graphs for explainable artificial intelligence: A survey. In Proceedings of the 2022 International Joint Conference on Neural Networks (IJCNN), pages 1–8, 2022. [9] Kurt Holstein, Jennifer Wortman Vaughan, Hal Daum´e III, and Mike Dudik. Improving fairness in machine learning systems: What do industry practitioners need? In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 2019. [10] Yu Hou, Jeremy Yeung, Hua Xu, Chang Su, Fei Wang, and Rui Zhang. From answers to insights: Unveiling the strengths and limitations of chatgpt and biomedical knowledge graphs. In Proceedings of the Annual Conference on Artificial Intelligence in Medicine, 2023. [11] Tuan Manh Lai. Knowledge Acquisition for Natural Language Understanding. PhD thesis, University of Illinois at Urbana-Champaign, 2023. [12] Yahan Li, Yi Wang, Yi Chang, and Yuan Wu. Xtrust: On the multilingual trustworthiness of large language models. In Proceedings of the International Conference on Multilingual NLP, 2023. [13] Carlo Lipizzi. Tell me the truth: A system to measure the trustworthiness of large language models. In Proceedings of the International Conference on Trustworthy AI, 2023. [14] Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, pages 220–229, 2019. [15] Jeff Z. Pan et al. Large language models and knowledge graphs: Opportunities and challenges. arXiv preprint arXiv:2308.06374, 2023. [16] Shirui Pan, Zhiwei Liu, Wei Zhuang, Rui Yang, Lei Zhang, Jialiang Li, and Haifeng Wang. Unifying large language models and knowledge graphs: A roadmap. IEEE Transactions on Knowledge and Data Engineering, 2024. [17] Aleksandra Piktus, Yi Wang, Vladimir Karpukhin, Ravi Pasunuru, Thibaut Mialon, Nisan Stiennon, Wen-tau Yih, Laleh Koc, and Matthew Efron. The roots search tool: Data transparency for llms. arXiv preprint arXiv:2302.14035, 2023. [18] Ridho Reinanda, Edgar Meij, and Maarten de Rijke. Knowledge graphs: An information retrieval perspective. Foundations and Trends®in Information Retrieval, 14(4):289–444, 2020. [19] Jingwei Sun, Zhixu Du, and Yiran Chen. Knowledge graph tuning: Real-time large language model personalization based on human feedback. arXiv preprint arXiv:2405.19686, 2024. [20] Zhen Tan, Yi Zhang, Xia Shen, Zhen Wang, and Lichao Li. Tuning-free accountable intervention for llm deployment–a metacognitive approach. arXiv preprint arXiv:2403.05636, 2024. [21] Weixuan Wang, Qi Liu, Huilin Xiong, Hang Liu, Liang Hu, Xinyu Zhang, and Yixin Gu. Assessing the reliability of large language model knowledge. arXiv preprint arXiv:2310.09820, 2023. [22] Yujie Wang, Xiaogang Zhang, and Wei Liu. Integrating knowledge graphs into language models: A comprehensive review. Journal of Artificial Intelligence Research, 72:119–143, 2023. [23] Sophie Xhonneux, Pierre Legrand, Wouter De Pauw, Xue Ma, John Sutherland, and Kara Kockelman. Efficient adversarial training in llms with continuous attacks. arXiv preprint arXiv:2405.15589, 2024. [24] Linyao Yang, Hongyang Chen, Zhao Li, Xiao Ding, and Xindong Wu. Give us the facts: Enhancing large language models with knowledge graphs for fact-aware language modeling. In IEEE Transactions on Knowledge and Data Engineering, 2023. [25] Shuo Yu, Zhen Wang, Xiang Zhang, Zhiyuan Liu, and Maosong Sun. Deep learning meets knowledge graphs: A comprehensive survey. arXiv preprint arXiv:2205.02573, 2022. 137 [26] Ahtsham Zafar, Venkatesh Balavadhani Parthasarathy, Chan Le Van, Saad Shahid, Aafaq Iqbal Khan, and Arsalan Shahid. Building trust in conversational ai: A review and solution architecture using large language models and knowledge graphs. In Proceedings of the Conference on Artificial Intelligence and Data Science, 2023. [27] Muhammad Rehman Zafar and Naimul Khan. Deterministic local interpretable model-agnostic explanations for stable explainability. Machine Learning and Knowledge Extraction, 3(3):525–541, 2021. [28] Zhao Zhang, Hua Xu, and Fei Wang. Knowledge graphs for enhancing the factual recall of large language models. In Proceedings of the Conference on Knowledge Graphs and Natural Language Processing, 2023. 138 A Hybrid Architecture for Tomato Leaf Disease Classification Through State Space and Convolutional Feature Fusion MAAROUF Ayoub Abderrazak1 1Laboratoire d’Automatique et de Robotique, D´epartement d’Electronique ,Universit´e des fr`eres Mentouri Constantine, Algeria Abstract Tomato is a globally important crop, with annual production exceeding 180 million tons. However, fungal and pest-induced diseases contribute to yield losses of 20–40% worldwide. This paper proposes Mamba-CNN, a novel hybrid architecture that combines state space models with convolutional neural networks for tomato leaf disease classification. Our method achieves an accuracy of 93.7% on a 5-class dataset by leveraging a synergistic fusion of global and local features, significantly outperforming standalone CNNs (85.9%) and Mamba Vision (88.2%). The proposed framework is particularly effective in capturing fine-grained visual patterns and modeling long-range disease progression. Keywords: Mamba Vision, tomato leaf disease, image classification, convolutional neural networks (CNN). 1 Introduction The agricultural sector, a cornerstone of the global economy, faces mounting challenges such as climate change, disease outbreaks, and labor shortages. Addressing these issues is essential to ensuring food security and promoting sustainable development. Among the emerging technological solutions, artificial intelligence (AI) has emerged as a transformative force in modern agriculture [?]. AI empowers farmers with deep insights into crop health, resource optimization, and risk mitigation. By analyzing large-scale datasets—including satellite imagery, sensor data, and historical records—intelligent systems can detect early signs of disease, predict yields, and recommend targeted interventions [?]. In particular, edge AI solutions for plant disease detection have shown promising results. Integrating deep learning models such as YOLOv3 with embedded platforms like the NVIDIA Jetson TX2 enables drones to accurately identify pest-infested zones and apply pesticides with precision, demonstrating the real-world utility of AI in precision agriculture [?]. Tomatoes (Solanum lycopersicum) are one of the most widely cultivated and consumed crops globally [?], valued for their nutritional content, including essential vitamins and antioxidants. However, tomato crops are frequently affected by a variety of foliar diseases, leading to significant yield losses and economic burdens on farmers. Early and accurate detection of these diseases is critical for effective crop management and food supply resilience. Traditional disease identification methods depend on expert visual inspection, which is time-consuming, labor-intensive, and inherently subjective. Recent advances in imaging and machine learning have enabled the development of automated systems capable of detecting plant diseases from leaf images with higher accuracy and speed. However, tomato leaf disease classification remains a challenging task due to the following real-world factors: •Visual Ambiguity: Early-stage lesions (1–2 mm) exhibit highly similar textures. •Context Dependency: Effective classification requires capturing both local spot patterns and global lesion distribution. •Field Variability: Environmental factors such as lighting, occlusion, and varying leaf orientations affect image quality. 139 To address these challenges, we propose Mamba-CNN, a novel hybrid architecture that combines state space models (SSMs) with convolutional neural networks (CNNs) for robust tomato leaf disease classification. The key contributions of this paper are as follows: 1. We introduce the first hybrid SSM-CNN architecture tailored for agricultural vision tasks. 2. We design a dynamic feature fusion mechanism enhanced with spatial-channel attention. 3. We conduct comprehensive benchmarking on a curated 5-class tomato leaf disease dataset. 2 Related Work 2.1 Traditional Computer Vision Approaches Early approaches to plant disease recognition relied heavily on handcrafted feature extraction techniques: •Color-Based Methods: Havg =1 N N X i=1 H(xi), H ∈[0,360](HSV space) (1) Introduced by [?], these methods were highly sensitive to illumination changes under real-world conditions. •Texture Analysis: Grey-Level Co-occurrence Matrix (GLCM) features: Contrast = N−1 X i,j=0 Pi,j(i−j)2(2) and Local Binary Patterns (LBP) were explored, but failed to effectively differentiate between visually similar fungal lesions [?]. •Shape Descriptors: Elliptic Fourier Descriptors attempted to quantify lesion morphology but underperformed when confronted with irregular or fragmented lesion boundaries [?]. 2.2 Deep Learning Architectures Modern techniques leverage deep learning, particularly convolutional neural networks (CNNs) and transformerbased models [?]: •Transfer Learning: Lce =− M X c=1 yclog(pc) (3) Pretrained CNNs such as ResNet-50 and EfficientNet achieved 80–85% accuracy on leaf datasets but struggled with subtle early-stage symptoms [?]. •Attention Mechanisms: Vision transformers (ViTs) apply multi-head self-attention: Attention(Q, K, V ) = softmax QKT √dkV(4) These models improve spatial focus but incur a 3×increase in computational cost [?]. •Multi-Scale Fusion: Feature pyramid networks (FPN) combine lowand high-level features to enhance spatial detail, but often introduce feature redundancy [?]. 140 2.3 State Space Models Recent work on sequence modeling has led to renewed interest in state space models (SSMs): •Mamba Architecture: Combines selective SSMs with hardware-aware design for efficient inference: yt=S6(xt,∆t, A, B, C, D) = SSM(Conv1D(xt)) (5) Mamba offers linear-time complexity O(L) with respect to sequence length L[?]. •Vision Applications: Vision Mamba [?] demonstrated strong performance in medical imaging, but exhibited limitations when applied to fine-grained agricultural textures. •Hybrid Models: Hybrid SSM-transformer models have been proposed to reduce computational cost, though some suffer from training instability [?]. Table 1: Comparative analysis of existing approaches Method Accuracy Params (M) Limitations SVM + GLCM [?] 68.2% – Illumination sensitivity ResNet-50 [?] 85.9% 25.6 Limited receptive field ViT-Base [?] 87.1% 86.4 High compute cost Mamba Vision [?] 88.2% 18.3 Poor texture modeling 2.4 Hybrid Vision Architectures Recent research has explored combining complementary architectural paradigms: •CNN-Transformer Hybrids: Achieved 89% accuracy on the PlantVillage dataset through localglobal feature fusion [?]. •SSM-Based Designs: Vision Mamba (VMamba) demonstrated the potential of SSMs in medical vision tasks [?]. •Agricultural Applications: Dilated CNNs achieved 82% accuracy for rice disease classification under real-field conditions [?]. 3 Methodology 3.1 Motivation for Hybrid Design Tomato leaf disease classification poses unique challenges that require both local texture understanding and global contextual reasoning: •Local Features: Early blight typically appears as 2–3 mm brown lesions. CNNs excel in capturing such fine-grained local patterns due to their localized receptive fields. •Global Context: The progression of disease across the leaf surface is often spatially extended and irregular. Mamba’s long-range sequence modeling capabilities are well-suited for capturing these broader patterns. 3.2 Architecture Design 3.2.1 Convolutional Backbone We employ a modified EfficientNet-B0 backbone for initial feature extraction: Fcnn :R3×224×224 →R1280×7×7(6) The early stem layers are preserved to ensure robust local texture encoding. 141 3.2.2 Vision Mamba Block A modified Vision Mamba module is applied for capturing long-range dependencies: Fmamba :R3×224×224 →R256×14×14 (7) Its key components include: •Patch embedding using 16 ×16 convolutional kernels •Three stacked Mamba blocks with an expansion ratio of 2 •Depth-wise convolution for efficient spatial mixing 3.3 Dynamic Feature Fusion To unify representations from the CNN and Mamba branches, we introduce a three-stage dynamic fusion module: 1. Dimension Alignment F′ cnn =AdaptiveP ool(Fcnn)∈R1280 (8) 2. Attention Weighting α, β =softmax(Wa[F′ cnn;Fmamba]) (9) 3. Nonlinear Combination Ffusion =α·F′ cnn +β·Fmamba +MLP ([F′ cnn;Fmamba]) (10) [htbp] Dynamic Fusion Process [1] Fcnn,Fmamba F′ cnn ←GlobalAvgP ool(Fcnn)F′ mamba ←Flatten(Fmamba) w←MLP([F′ cnn;F′ mamba]) α, β ←softmax(w)α·F′ cnn +β·F′ mamba 3.4 State Space Formulation We adopt a continuous-time state space model (SSM), discretized using zero-order hold for compatibility with image sequences: A=e∆A, B = (∆A)−1(e∆A−I)∆Bht=Aht−1+Bxtyt=Cht+Dxt(11) Here, ∆ denotes a learnable time-step, while A,B,C, and Dare trainable matrices that model dynamic state transitions. 3.5 Training Strategy Our training pipeline is divided into three phases to stabilize convergence and optimize performance: 1. Warm-Up Phase (10 epochs): •Learning rate linearly increases from 10−4to 3 ×10−4 •Mamba parameters are frozen •CNN is optimized using focal loss 2. Joint Training Phase (70 epochs): •All parameters are unfrozen •Optimized using the Lion optimizer with cosine learning rate decay •Introduce MambaMix augmentation: ˜x=λxa+ (1 −λ)xb, λ ∼Beta(0.8,0.8) (12) 3. Fine-Tuning Phase (20 epochs): •Learning rate is reduced to 10−5 •Apply layer-wise learning rate decay •Employ label smoothing with ϵ= 0.1 142 Table 2: Training Hyperparameters Parameter Warm-Up Phase Joint Training Phase Batch size 32 32 Learning rate 1 ×10−43×10−4 Weight decay 0.01 0.05 Augmentation Basic MambaMix 4 Experimental Results The experimental setup consists of a Windows 10 operating system equipped with 32 GB of RAM and GTX 3090 GPU. The model training is carried out using the PyTorch framework. 4.1 Dataset Collection and Preprocessing The dataset utilized in this study comprises images of tomato leaves categorized into five classes: Healthy, Early Blight, Late Blight, Leaf Mold, and Septoria Leaf Spot [?]. Each class is divided into training and testing subsets as follows: Table 3: Class Distribution and Characteristics Disease Train Test Characteristics Healthy 2,000 500 Uniform green coloration Early Blight 2,000 500 Concentric brown rings Late Blight 2,000 500 Water-soaked lesions Leaf Mold 2,000 500 Yellow upper surface, purple lower surface Septoria Leaf Spot 2,000 500 Circular spots with dark edges To ensure consistency and enhance model performance, the following preprocessing steps were applied: •Image Resizing: All images were resized to a uniform dimension suitable for input into the Vision Mamba model. •Normalization: Pixel values were normalized to a standard range to facilitate faster convergence during training. •Data Augmentation: Techniques such as rotation, scaling, and flipping were employed to increase the diversity of the training dataset and improve the model’s generalization capabilities. illustration of this dataset is presented in Figure 2. 5 Discussion As shown in Figure 2, Mamba-CNN achieves 90% accuracy by epoch 30, significantly faster than the CNN baseline (epoch 45). This acceleration suggests: •Effective Feature Fusion: The hybrid architecture successfully combines CNN’s local texture analysis with Mamba’s global pattern recognition early in training •Synergistic Learning: Joint optimization enables complementary feature discovery rather than independent pathway training •Stable Optimization: Careful learning rate scheduling prevents mode collapse in the dual-branch architecture Figure 3reveals only 1.3% accuracy difference between training and validation sets, suggesting: •Robust Regularization: Our MambaMix augmentation effectively simulates field conditions (shadows, occlusions) •Balanced Learning: The focal loss successfully handles class imbalance (Spider Mites vs. Septoria samples) 143 Figure 1: image dataset Figure 2: Accuracy progression across training epochs demonstrates Mamba-CNN’s rapid convergence compared to baseline models. 144 Table 4: Summaries of community distributions of each CadjMAxiteration - with the modularity of the WLFM algorithm. CadjMax Nombre de communauties Modularity (Q) 472 10 0.53652645659928 Figure 10: Graph result for the best distribution CadjMax=472. •Extend the LFM2ACO algorithm to handle overlapping communities. •Improve the scalability of the LFM2ACO algorithm. •Apply the LFM2ACO approach to directed graphs. •Evaluate the performance of the LFM2ACO algorithm on a wider range of datasets. Another AI perspective regarding the use of AI to our original LFM algorithm or the one proposed in this work (LFM2ACO) such as: •Integrating Deep Learning for Predictive Community Dynamics: Leverage temporal graph neural networks (TGNNs) or transformer-based architectures to model the evolution of social interactions, enabling the prediction of future community structures (e.g., births, mergers, or splits) based on historical trajectory patterns and individual behavior embeddings. •Behavior-Aware Forecasting with Reinforcement Learning: Develop hybrid models combining LFM2ACO with deep reinforcement learning (DRL) to simulate adaptive agent behaviors, where AI-driven individuals dynamically switch communities based on learned reward mechanisms reflecting social preferences. •LLM-Enhanced Relationship Semantics: Utilize large language models (LLMs) to analyze textual interaction data (e.g., social media content), extracting semantic signals to enrich edge weighting in WLFM and predict community formation triggers from latent topic shifts. •Neural Attention for Overlap Resolution: Implement multi-head attention mechanisms to detect overlapping community boundaries by learning node-community affiliation probabilities, complementing ACO’s pheromone dynamics with neural interpretability. 247 •Graph Generation for Scenario Projection: Train generative adversarial networks (GANs) or diffusion models on temporal network snapshots to synthesize plausible future graph states, enabling stress-testing of LFM2ACO under predicted social configurations. •Embedding-Driven Scalability: Combine hyperbolic graph embeddings with LFM2ACO’s optimization process to reduce computational complexity in large-scale networks while preserving hierarchical community structures. •Multimodal Fusion for Event Prediction: Architect multimodal pipelines that jointly process network topology (via GNNs), temporal activity sequences (via LSTMs), and user metadata to forecast macro-level community events like mass migrations or influencer-driven splits. References [1] Stanley Wasserman, Katherine Faust, Stanley (University of Illinois Wasserman, UrbanaChampaign) Social Network Analysis: Methods and Applications, Volume 8 de Structural Analysis in the Social Sciences, ISSN 0954-366X, editeur:Cambridge University Press 1994,825 pages. [2] C. Dawson and C. Dawson, “Social network analysis,” A–Z Digit. Res. Methods, pp. 356–361, 2019, doi: 10.4324/9781351044677-54. [3] Djerbi, R., Amad, M., & Imache, R. (2020). A new model for communities’ detection in dynamic social networks inspired from human families. International Journal of Internet Technology and Secured Transactions, 10(1-2), 24-60. [4] S. Fortunato, “Community detection in graphs,” Phys. Rep., vol. 486, no. 3–5, pp. 75– 174, 2010, doi: 10.1016/j.physrep.2009.11.002. [5] M. Girvan and M. E. J. Newman, “Community structure in social and biological networks,” vol. 99, no. 12, 2002. [6] A. Lancichinetti, S. Fortunato, and F. Radicchi, “Benchmark graphs for testing community detection algorithms,” Phys. Rev. E - Stat. Nonlinear, [7] M.NEDIOUI, M´emoire fouille de donn´ee et apprentissage automatique dans les r´eseaux sociaux dynamiques, 2015. [8] NEDIOUI, MED ABDELHAMID. Fouille et apprentissage automatique dans les reseaux sociaux dynamique. 2015. Th‘ese de doctorat. Universit´e Mohamed Khider-Biskra, Alg´erie. [9] Blondel, 2008, V.D. Blondel, J.L. Guillaume, R. Lambiotte et E. Lefebvre. Fast unfolding of communities in large networks. Journal of Statistical Mechanics: Theory and Experiment, vol. 2008, page P10008, 2008. [10] Palla,2007G. Palla, A.L. Barabasi, and T. Vicsek. Quantifying social group evolution. Nature, 446(7136) :664667, 2007. [11] Aynaud T., Fleury E., Guillaume J.-L., Wang Q. (2013). Communities in evolving networks: definitions, detection, and analysis techniques. In Dynamics on and of complex networks, volume 2, p. 159–200. Springer. [12] Z. Chen, K. a. Wilson, Y. Jin, W. Hendrix, and N. F. Samatova. Detecting and Tracking Community Dynamics in Evolutionary Networks. 2010 IEEE International Conference on Data Mining Workshops, pages 318–327, Dec. 2010. [13] Dorigo Gambardella, 1997] Dorigo, M., & Gambardella, L.M. 1997. Ant Colony System: A Cooperative Learning Approach to the Traveling SalesmanProblem. IEEE Transactions on Evolutionary Computation,1(1), 53 66. [14] H.BELLEILI ,2020 Ant ColonyOptimization (ACO) optimisation par colonies de fourmis [15] The Facebook Wall dataset: http://socialnetworks.mpi-sws.mpg.de/data/facebook-wall.txt.gz, Last accessed 24 Mars 2025 248 [16] The Facebook Links dataset http://socialnetworks.mpi-sws.mpg.de/data/facebook-links.txt.gz, Last accessed 24 Mars 2025 [17] Shetty, J., & Adibi, J. (2005, August). Discovering important nodes through graph entropy the case of enron email database. In Proceedings of the 3rd international workshop on Link discovery (pp. 74-81). 249 The Second National Conference on Applications of Artificial Intelligence (A2I-25) brings together researchers, professionals, and students to explore innovative applications of Artificial Intelligence that address real-world challenges in fields such as healthcare, agriculture, energy, cybersecurity, and urban development. Hosted by the University M’Hamed Bougara of Boumerdes (UMBB), this event fosters interdisciplinary collaboration and highlights the transformative power of AI in improving quality of life and societal well-being. The conference also features an exclusive NVIDIA-certified training on Fundamentals of Deep Learning, led by Dr. Tayeb Benzenati, providing participants with hands-on experience in neural networks, optimization, and real-world AI applications. Dates: April 16–17, 2025 Venue: Department of Computer Science, Faculty of Sciences, University of M’hamed Bougara UMBB, Boumerdes, Algeria. © 2025 University M’Hamed Bougara Boumerdes. All rights reserved.