Full text
Frontiers in Artificial Intelligence 01 frontiersin.org Implementing federated learning for privacy-preserving emotion detection in educational environments RommelGutiérrez 1, WilliamVillegas-Ch 1* and SergioLuján-Mora 2 1 Escuela de Ingeniería en Ciberseguridad, FICA, Universidad de Las Américas, Quito, Ecuador, 2 Departamento de Lenguajes y Sistemas Informáticos, Universidad de Alicante, Alicante, Spain Emotion detection has become an essential tool in educational settings, where understanding and responding to students’ emotions is crucial to improving their engagement, academic performance, and emotional well-being. However, traditional emotion detection systems, such as DeepFace, and hybrid transformer-based models face significant data privacy and scalability limitations. These models rely on transferring sensitive data to central servers, compromising student confidentiality and making deployment in large or diverse populations difficult. In this work, wepropose a federated learning-based model designed to detect emotions in educational settings, preserving data privacy by processing them locally on students’ devices (smartphones, tablets, and laptops). The model was integrated into the Moodle platform, allowing its evaluation in a conventional educational environment. Advanced anonymization and preprocessing techniques were implemented to ensure the security of emotional data and optimize its quality. The results demonstrate that the proposed model achieves a precision of 87%, a recall of 85%, and an F1-score of 86%, maintaining its performance under adverse conditions, such as low lighting and ambient noise. In addition, a 15% increase in academic participation and a 12% improvement in the average academic performance of students were observed, highlighting the system’s positive impact on educational dynamics. This innovative method combines privacy, scalability, and performance, positioning itself as a viable and sustainable solution for emotion detection in contemporary educational environments. KEYWORDS federated learning, emotion detection, data privacy, educational environments, artificial intelligence 1 Introduction Emotion detection has emerged as a critical area in developing intelligent systems, particularly in educational contexts, where emotions play a pivotal role in student learning and behavior (Mutawa and Hassouneh, 2024; Shmelova etal., 2024). Understanding and responding to student emotions can significantly improve the personalization of teaching strategies, optimize academic engagement and performance, and contribute to the overall emotional well-being of students (Elisondo etal., 2024). However, the implementation of emotion detection systems faces significant challenges related to data privacy, scalability, and integration into diverse educational settings (Wang A. etal., 2024). Centralized models, such as DeepFace by An etal. (2023) facial features have been widely used for emotion detection due to their high performance on metrics such as precision and OPEN ACCESS EDITED BY Rita Orji, Dalhousie University, Canada REVIEWED BY Oladapo Oyebode, Dalhousie University, Canada Grace Ataguba, Dalhousie University, Canada Dario Di Dario, University of Salerno, Italy *CORRESPONDENCE William Villegas-Ch [email protected] RECEIVED 13 June 2025 ACCEPTED 28 October 2025 PUBLISHED 07 November 2025 CITATION Gutiérrez R, Villegas-Ch W and Luján-Mora S (2025) Implementing federated learning for privacy-preserving emotion detection in educational environments. Front. Artif. Intell. 8:1644844. doi: 10.3389/frai.2025.1644844 COPYRIGHT © 2025 Gutiérrez, Villegas-Ch and Luján-Mora. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms. TYPE Original Research PUBLISHED 07 November 2025 DOI 10.3389/frai.2025.1644844
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 02 frontiersin.org F1-score. However, these systems require the transfer of sensitive data to external servers for training and inference, which poses significant privacy risks. In contrast, models based on Takase and Kiyono (2023) transformers have demonstrated their ability to integrate multiple modalities, such as text, audio, and images, offering a more robust approach. Despite their advantages, high computational costs and configuration complexity limit their implementation in educational settings. In education, privacy and accessibility are crucial factors. Transferring students’ emotional data to external servers compromises confidentiality and poses ethical and legal challenges in handling sensitive information (Lee etal., 2024). At the same time, accessibility refers to the ability of emotion detection systems to function effectively across diverse educational contexts, including institutions with limited infrastructure or students with varying levels of technological access. Systems that require high-performance computing or stable connectivity may exclude part of the student population, reinforcing educational inequality. Despite these limitations, few studies have addressed these issues by developing emotion detection systems specifically designed to beintegrated into learning platforms, such as Moodle, ensuring both data protection and adaptability to the technological realities of educational institutions (Labidi etal., 2021; Woodward etal., 2024). This work introduces a model based on federated learning designed for emotion detection in educational settings, which addresses these critical challenges. Federated learning allows training models to berun directly on users’ local devices, eliminating the need to transfer sensitive data to central servers (Sengupta etal., 2024). In addition to model development, this work emphasizes the practical integration of the federated emotion detection system into real educational environments, assessing its influence on student engagement and academic outcomes. This feature improves privacy and enables greater scalability by allowing the system to operate on large and heterogeneous student populations (Wang etal., 2024a). The proposed methodology includes a multi-stage approach, starting with the collection of emotional data through images, audio, and text generated during academic activities. The data was preprocessed using advanced anonymization and feature extraction techniques, such as random facial point mapping and prosody analysis in speech. Subsequently, the federated model was trained locally on devices such as smartphones, tablets, and laptops, using federated averaging algorithms to combine the model updates on a central server without compromising data privacy (Doriguzzi-Corin and Siracusa, 2024). The results of this approach show that the proposed model achieves competitive metrics in terms of precision 87%, recall 85%, and F1-score 86%, which positions it as a robust alternative to centralized systems such as DeepFace and commercial solutions such as Affectiva SDK (Hammann etal., 2022). Furthermore, robust tests performed under adverse conditions, such as variations in lighting and environmental noise, demonstrated that the model maintains consistent performance, outperforming centralized models in similar scenarios. For example, the model’s precision in low lighting conditions was 80%, compared to 75% for centralized models evaluated under the same conditions. Integrating the model into Moodle, a widely used learning management system, enabled us to evaluate its practical applicability in a conventional educational environment (Shchedrina etal., 2021). This process demonstrated the system’s ease of adoption and highlighted its positive impact on student behavior. The results indicate that positive emotions, such as motivation, detected by the system are associated with a 15% increase in academic engagement and a 12% improvement in students’ average academic performance. In contrast, although it is more challenging to detect negative emotions, such as stress and frustration, it provides valuable data to adjust educational strategies and provide targeted emotional support. Despite these advances, the model faces limitations inherent to the federated approach, such as dependence on heterogeneous devices and sensitivity to the quality of network connections during the model aggregation process. Although significant, these limitations do not compromise the system’s viability; instead, they highlight the need for future research to optimize its performance in environments with limited technological infrastructure. This study’s main contribution lies in combining privacy, scalability, and performance in an emotion detection system specifically designed for educational environments. It aims to determine the extent to which a federated learning model can accurately identify students’ emotional states, both explicit and nuanced, using data from fundamental academic interactions across multiple modalities. The work further explores how decentralized training affects model reliability under real-world constraints, including limited infrastructure and diverse emotional expression patterns. Unlike existing solutions, the proposed approach ensures the confidentiality of emotional data while providing an adaptable and practical tool for academic institutions. The remainder of this article is structured as follows: Section 2 presents a literature review on emotion detection systems, highlighting the current limitations in terms of privacy and scalability. Section 3 describes the materials and methods, including the data collection process, preprocessing techniques, and the design of the federated learning architecture. Section 4 presents the experimental results, including performance comparisons, robustness evaluations, and assessments of real-world impact. Section 5 discusses the findings about existing literature, addresses limitations, and outlines future research directions. Finally, Section 6 summarizes the main contributions and conclusions of the study. 2 Literature review Emotion detection has been the subject of numerous studies examining various approaches to identifying human emotions in diverse contexts. Among these approaches, centralized systems such as DeepFace (An etal., 2023) have demonstrated high performance in emotion classification based on facial features (Anand and Babu, 2024). DeepFace uses highly trained convolutional neural networks (CNNs) to process images on central servers (Zhang Y. etal., 2024), achieving precision levels of up to 90% in emotion detection tasks. However, this centralized model faces significant criticism due to the need to transfer sensitive personal data to external servers, compromising user privacy. This aspect limits its applicability in educational settings, where data protection is a priority. Another prominent approach is hybrid transformer-based models, such as those presented by Teng et al. (2024), which combines image, audio, and text processing to achieve more robust emotion detection. These systems can analyze multiple modalities,
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 03 frontiersin.org integrating contextual and temporal features into their architecture. Although transformers offer advantages in terms of flexibility and performance, their high computational cost and reliance on large volumes of data limit their deployment in resource-constrained educational settings. Moreover, like centralized systems, these models often require transferring data to external servers, posing similar risks to privacy. Commercial systems, such as the Affectiva SDK, have been specifically designed for practical applications in marketing and behavioral analysis (Kulke etal., 2020). This software utilizes advanced computer vision techniques to identify facial emotions in real-time and is optimized for commercial platforms. Affectiva stands out for its ease of use and competitive performance (Shwe Sin and Khin, 2022), with accuracies ranging from 85 to 88%. However, its closed approach and high licensing costs hinder its adoption in educational settings, where budgets are often limited, and custom configurations are essential to integrate into existing platforms such as Moodle. Federated learning emerges as an innovative solution to address the privacy and scalability limitations of centralized models. Huang etal. (2023) demonstrated the potential of federated architectures for emotion detection; however, their framework exhibited limitations in device heterogeneity, requiring uniform client capabilities and stable communication channels. These constraints limited scalability and reduced effectiveness in dynamic educational environments, where device resources and connectivity vary considerably. Moreover, their work did not include practical integration with learning platforms, which limited its pedagogical impact and real-time applicability within classroom systems. Compared to the reviewed models, the federated approach proposed in this work stands out for its ability to balance performance, privacy, and scalability. It explicitly addresses the technical challenges noted by Huang etal. (2023) by introducing adaptive preprocessing techniques that tolerate device variability, optimizing local training for constrained devices, and integrating directly with Moodle and other learning management platforms. This reduces deployment complexity and supports institutions with limited infrastructure (Mukta etal., 2024). Despite advances in federated learning, existing literature has yet to explore its comprehensive implementation in hybrid academic environments that combine real users, platform integration, and privacy-by-design principles. This work seeks to address this shortcoming by presenting an integrated and deployable system designed for emotion detection in educational settings. 3 Materials and methods 3.1 Description of the test environment The federated learning-based emotion detection system was implemented in a university educational environment, specifically in the Faculty of Technologies, which includes approximately 650 students. This environment is characterized by a hybrid education modality, meaning that students attend classes both in person and online. This hybrid modality presents an interesting challenge for implementing emotion detection technologies, as students interact with content and teachers in multiple ways—either in the physical classroom or through digital platforms—enabling the collection of emotional data in diverse contexts. The Faculty of Technologies offers training programs in disciplines related to computer science, electronic engineering, and communication networks. This academic profile makes the federated learning approach particularly suitable, as most students are familiar with using technological tools and are active users of smart devices, which facilitates the adoption of the proposed technology for emotion detection. A total of 150 students were selected to participate in the study, representing approximately 23% of the faculty’s total student population. This group was chosen randomly but representatively, ensuring the inclusion of students from different majors within the faculty and capturing a diverse sample of emotions. In addition, 20 teachers actively participated in the study, allowing for the monitoring of students’ emotional well-being throughout the course, both in faceto-face and online classes. The selected sample consisted of undergraduate students with an average age of 21.2 years (SD = 1.7), ranging from 18 to 25. Gender distribution was approximately 56% male and 44% female. Students came from three main academic programs: Computer Science, Electronic Engineering, and Communication Networks. All participants were enrolled in hybrid courses that combined in-person and virtual components, ensuring a wide range of interaction modalities with the system. This diversity supports the generalizability and robustness of the experimental findings. The implementation occurs in an online and hybrid education environment, providing an ideal opportunity for collecting emotional data in real-time and asynchronous interactions (Pirrone etal., 2021). Students interact with the system through various devices, either during online classes, remote exams, or discussion forums and activities within the Moodle platform, which served as the Learning Management System (LMS) in this pilot test. The emotional data collection process is performed through various smart devices, such as smartphones, tablets, and laptops, which are standard in the faculty and integrated into the students’ daily activities. These devices capture emotional data through facial expression analysis, emotion detection through tone of voice during oral interactions, and text analysis in written responses on LMS platforms, mainly in forum activities and assessment tasks. Each device acts as a node in the federated system, where the emotional data captured on each one is processed locally to preserve the privacy of the students (Ribeiro Junior and Kamienski, 2024). The students’ devices preprocess the emotional data through applications developed specifically for this test, extracting relevant features from facial images, vocal tone, and textual responses. The emotion detection model is trained locally on these devices, using the data collected in real-time, without such data leaving the device (Almalki etal., 2024). Students can participate in the system through the mobile app and on their desktop devices without requiring constant direct interaction with the system, thus allowing the federated learning model to adapt to variations in emotions throughout the educational day. Teachers can access the emotion reports generated without compromising students’ privacy and use these reports to adjust their pedagogical strategies in real time, especially regarding student workload and stress during classes and assessments. The system infrastructure is based on federated architecture, where student devices train the emotion detection model independently. Communication between the devices and the central server is limited to model updates only, ensuring that sensitive data is
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 04 frontiersin.org not shared at any time (Zhou etal., 2024)Model updating employs techniques such as federated averaging, which enables the central server to aggregate updates to local models without requiring the original data for each student to becentralized. Figure1 presents the architecture of the proposed system for emotion detection using federated learning. It demonstrates how local student devices, such as smartphones, laptops, and tablets, interact with the system. Each device performs local preprocessing of the emotional data, extracting relevant features from facial expressions, tone of voice, and textual interactions. Once the regional models are trained on the devices, the model updates are sent to the central server for aggregation through a federated averaging process. This process allows the central server to combine the model updates without centralizing the original student data, ensuring data privacy (Zhang H. etal., 2024). Furthermore, aggregated reports of student emotions are visualized through the Teacher Dashboard, providing teachers with valuable information about the class’s emotional well-being without requiring access to individual personal data. The connection to Moodle enables the capture of students’ academic opinions in realtime, while model updates continually improve as more emotional data is collected. 3.2 Data collection 3.2.1 Emotion capture method Three specific techniques are used for emotion detection: facial expression analysis, voice tone detection, and text analysis. These techniques are applied complementarily to ensure a complete and accurate assessment of students’ emotional states during educational interactions. Facial expression analysis is based on the premise that human emotions are reliably reflected in facial movements, which are detected and classified with high precision using computer vision techniques (Lyu, 2023). This process is carried out using a CNN-based model, which allows the identification of key facial features such as eye, mouth, and eyebrow movements. Through the front-facing cameras of the devices, the system captures the students’ facial expressions in real-time. The data obtained is processed locally on each device to extract the relevant emotional features, allowing the detection of emotions such as happiness, sadness, anger, surprise, contempt, and disgust, which correspond to the basic emotions identified by Ekman etal. (1998). The application of this model is carried out continuously during the student’s interactions with the academic environment, ensuring the accurate capture of emotions in various situations. However, it is acknowledged that some of these interactions may occur outside the core academic platform (e.g., Moodle), were external, non-educational factors could influence emotional variation. These factors lie beyond the teacher’s control and could introduce biases in interpreting students’ emotional states, a limitation also highlighted in recent ethical studies on emotion recognition in educational settings (Di Dario etal., 2024). Voice pitch detection is another crucial method in emotion detection. This process involves capturing and analyzing variations in the acoustic features of the voice, such as fundamental frequency, intensity, duration, and prosody (Jiang etal., 2023). These features indicate emotional variations in speech, as vocal pitch and rhythm change in response to the individual’s emotional state. Microphones in the devices pick up the student’s voice during oral interactions. Using audio signal processing algorithms, such as acoustic feature analysis and time-sequence modeling, variations in speech are analyzed to identify emotions, including stress, confusion, or satisfaction. A model based on Recurrent Neural Networks (RNN), specifically Long Short-Term Memory (LSTM), is used, which can identify emotional patterns throughout voice sequences, allowing accurate detection in dynamic situations (Chen etal., 2017). FIGURE1 Architecture of the federated learning-based emotion detection system.
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 05 frontiersin.org Text analysis examines the emotional content of students’ written responses on platforms such as Moodle, especially in forum interactions, assignments, and exams. The system identifies linguistic patterns related to emotions using natural language processing (NLP) techniques (Merity et al., 2018). Advanced models, such as Bidirectional Encoder Representations from Transformers (BERT), are applied, which can analyze the semantic context of words and phrases within texts. This analysis allows the classification of underlying emotions, such as anxiety, motivation, or confusion, from written interactions (Subakti etal., 2022). Text data is processed in real time, assessing students’ emotions based on their written expressions during academic activities. 3.2.2 Devices and sensors Emotional data is collected using various devices and sensors embedded in students’ devices, specifically smartphones, tablets, and laptops. Each device has specific technologies that accurately capture emotional data based on the interaction modality. The cameras capture facial expressions, enabling the real-time visual analysis of emotions. The cameras can identify and classify facial patterns related to basic emotions, such as happiness, sadness, anger, surprise, and others. On the other hand, the microphones built into the devices allow for capturing the acoustic characteristics necessary to analyze voice tone. These microphones are designed to capture sounds at an appropriate frequency, enabling the detection of variations in pitch, volume, and speech rate that indicate different emotional states. Through these microphones, the system can identify whether the student is experiencing emotions such as stress or satisfaction, which correlates with the tone and dynamics of their voice. The Moodle platform is used for collecting textual data. Students interact on the platform through forums, assignments, and exams, generating written responses that are then processed to assess the underlying emotions in their content. The system analyzes the words, phrases, and text structure using natural language processing models to identify emotional states related to the content of the responses, such as anxiety, motivation, or confusion. Table 1 summarizes the devices, sensors, and platforms employed, along with their respective functions in the emotional data collection process. All devices used for data collection were the personal property of the students. The emotion detection system was not pre-installed; instead, it was accessed entirely through the Moodle Learning platform, which provided a seamless interface for data capture and analysis. This integration ensured that no additional software needed to beinstalled on student devices, thereby minimizing intrusiveness and maintaining user autonomy. Additionally, the system’s architecture ensures that all data is processed locally on the device, aligning with the principles of privacy-by-design. 3.3 Data preprocessing In the filtering and anonymization process, specific techniques are applied to protect sensitive data, especially students’ facial and audio features (Hanisch etal., 2024). Facial expression data is processed to remove backgrounds and lighting variations irrelevant to emotion detection. This filtering is performed by a face segmentation algorithm using the OpenCV library, which detects the exact location of the face within the image and crops only the region of interest. A Gaussian smoothing filter is applied to the face region to reduce background noise and ensure that only relevant facial features are processed (Nandan etal., 2024). Data anonymization is performed by modifying the detected facial points so the individual cannot beidentified. In the case of facial analysis, key landmarks, such as the eyes, eyebrows, and mouth, are replaced by generic points that do not correspond to a specific identity. This technique uses facial mapping algorithms that randomly relocate facial features within a range of standard facial parameters (Wang, 2024). In addition, the data is not stored in its original form; instead, only the model updates are sent, implying that the facial images never leave the local devices and do not contain identifiable information. In encryption, the Advanced Encryption Standard (AES-256) encrypts the model parameters when they are sent from the local devices to the central server (Ajagbe etal., 2024; Mishra etal., 2024). This encryption ensures that even if the data is intercepted, it cannot bedecrypted without the proper key, protecting the students’ privacy during model communication. The feature extraction process for each data type (facial expressions, voice, and text) is carried out using specific algorithms designed for each modality. To ensure clarity and reproducibility, the emotional states targeted by each modality were explicitly defined and consistently applied throughout the training and evaluation phases. Each modality was associated with a distinct subset of emotional labels based on the nature of the data and the capabilities of the corresponding model. These labels were selected from well-established emotional taxonomies that have been adapted for educational settings. Table2 summarizes the exact emotions detected by facial expressions, voice signals, and textual content. 3.3.1 Facial expression detection In facial expression analysis, Dlib’s facial point detection algorithm identifies critical points on the face, such as the contours of the eyes, nose, and mouth. The mathematical process underlying this algorithm is based on supervised learning and nonlinear regression techniques. Once these points are detected, the Active Shape Model (ASM) is used to model the variability in the shape of the face (Alavi etal., 2024). The ASM can be described mathematically by an elastic deformation model that adjusts parameters to capture facial variation. The warping algorithm uses affine transformation matrices, where TABLE1 Devices, sensors, and platforms used for collecting emotional data. Device/sensor Function Associated technique Smartphone/tablet/laptop Capture of emotional data through a camera and a microphone Facial analysis, voice detection, text analysis Camera Capture of facial expressions for real-time analysis Facial expression analysis Microphone Capture of tone of voice, variations in frequency, and amplitude for emotion detection Voice tone analysis Moodle (LMS) A platform for collecting textual responses through educational interactions Text analysis
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 06 frontiersin.org each transformation T is represented by a parameter matrix Ɵ that adjusts the position of each facial point p(x, y) in the image according to the warping model of the Equation (1): ( ) ( ) ( ) θ = ′ ⋅ ,, p xy T p xy (1) T is the transformation matrix, which describes how facial points are adjusted according to emotional variations. In addition, the Euclidean distance between the detected facial points is calculated to measure the degree of change in facial expressions as shown in Equation (2): ( ) ( ) = − +− 22 21 21 d xx yy (2) This allows the quantification of deformation in the face related to emotions such as surprise or sadness. Speech signal analysis is based on acoustic features such as the fundamental frequency (F 0 ) extracted using the Fast Fourier Transform (FFT). Mathematically, the FFT decomposes a signal x(t) into its frequency components, representing the signal in the frequency domain as a sum of sinusoids, as shown in Equation (3): ( ) ( ) π ∞ − −∞ = ∫ 2j ft X f x t e dt (3) X(f) is the frequency domain representation of the signal, f is the frequency, and x(t) is the time domain signal. The fundamental frequency (F 0 ) is the lowest component of the audio signal and is related to the pitch of the voice. This parameter is extracted to measure emotional variations in the voice pitch, such as when anger or joy is detected. Additionally, prosody analysis is employed, which examines the intensity and rhythm of the voice. Mathematically, rhythm can bemeasured in terms of syllable duration and speech rate, and intensity is evaluated as the amplitude of the audio signal in each time window using energy measurement formulas, as shown in Equation (4): ( ) ( ) = = + ∑2 0 N n Et xt n (4) where E(t) is the energy in a time window, x(t + n) is the value of the audio signal at time t + n, and N is the number of samples within the time window. Prosody analysis is then used to feed the LSTM model, which applies backpropagation through time (BPTT) to update the neural network weights and model emotions based on speech’s pitch and temporal variability. LSTMs use activation functions such as sigmoid or tanh, which classify emotions based on the temporal content of the signal. 3.3.2 Text analysis Text analytics is based on transformer models, such as BERT, designed to capture the bidirectional context of words within a sentence (Kotwal etal., 2022). Mathematically, this model is a word embedding, which maps words to high-dimensional vectors in a vector space, using functions such as softmax to generate classification probabilities. In mathematical terms, the embedding process is described by a projection of each word w i into a d-dimensional feature space, as shown in Equation (5): ( ) = ii v fw (5) where vi is the feature vector of the word wi, and f is the projection function learned during training. Using self-attention, the BERT model captures contextual relationships between words, which computes the weighted relationship between words within a given context. Attention is mathematically defined in Equation (6): ( ) = ,, T k QK Attention Q K V softmax V d (6) where Q, K, and V are the query, key, and value matrices, respectively, and d k is the dimension of the keys. This attention mechanism enables the model to capture long-range dependencies within the text, allowing it to detect complex emotions such as frustration or motivation. Once the vector representations of the words are obtained, they are used to classify the emotions associated with the text through a deep neural network that adjusts the weights using the backpropagation algorithm and the softmax activation function to obtain the probability of each emotion, as shown in Equation (7): ( ) =∑ i j z iz j e P emotion e (7) where z i are the network outputs for each emotional class, and P(emotioni) is the probability that the emotion is present in the text. 3.3.3 Model training by modality and dataset description For emotion detection in educational contexts, three specialized models were developed, each adapted to a different modality: facial images, voice signals, and written text. These models were trained using public datasets widely validated in the literature, ensuring their availability and validity for emotion classification tasks. Furthermore, invasive collection processes or those dependent on sensitive information were avoided, aligning with the privacy principles defined in the overall system design. For emotion detection using facial expressions, the JAFFE dataset was utilized, which comprises 213 images with a resolution of 48 × 48 TABLE2 Emotional states detected by each modality. Modality Emotional states detected Facial expression Happiness, sadness, anger, surprise, disgust, contempt Voice (audio) Stress, confusion, satisfaction, boredom, engagement Textual content Motivation, anxiety, confusion, frustration, curiosity
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 07 frontiersin.org pixels, categorized into seven emotional states: happiness, sadness, anger, surprise, fear, disgust, and neutrality. These images were captured in uncontrolled scenarios, enabling the model to generalize more effectively in real-life conditions. For speech-based detection, the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) dataset was used. It consists of 1,440 audio clips recorded by professional actors expressing eight emotions (calm, happiness, sadness, anger, fear, disgust, surprise, and neutral). Finally, for the textual modality, the EmotionX dataset, which focuses on fundamental conversational interactions, was utilized. This corpus includes brief responses manually labeled with emotions such as happiness, sadness, anger, motivation, frustration, or surprise. All datasets were openly accessible, and no private data was included; additionally, no manual labeling was performed. Each model was designed to respond to the specific characteristics of its modality. The facial image model employed a VGG-13-based architecture, comprising two convolution blocks with 64 and 128 filters, respectively, followed by max pooling operations and dense layers of 512 and 128 units, before the softmax output layer. ReLU activation functions and a dropout value of 0.5 were used to prevent overfitting. For speech modality, the model was built on an LSTM network with 256 hidden units, followed by a dense layer with 64 neurons and a softmax output tuned to eight classes. The input consisted of sequences of MFCC coefficients extracted from 25-ms segments. Batch normalization, a dropout of 0.3, and the categorical cross-entropy loss function were employed. For text, the BERT model (uncased version, 110 million parameters) was implemented, to which a dense layer with 128 neurons and a softmax output of six classes was added. Fine-tuning was performed only on the last four layers of the transformer to preserve the pre-trained semantic capacity. The three models were trained using a familiar hyperparameter setting, with slight variations tailored to the computational needs of each modality. A batch size of 32 was used for images and speech, and 16 for text. The initial learning rate was 0.0001, with the Adam optimizer and a weight decay penalty of 1e-5. The maximum number of epochs was set to 50, with an early stopping mechanism activated if no improvement was observed in validation after 10 iterations. In all cases, the sets were divided into 70% for training, 15% for validation, and 15% for testing, following a consistent protocol across modalities. The training environment consisted of notebooks developed in Python 3.9 using PyTorch 2.0, HuggingFace Transformers, and the librosa library for acoustic feature extraction. The experiments were conducted on Google Colab Pro+ with access to a 16 GB Tesla T4 GPU and 52 GB of RAM, enabling efficient and reproducible training. It is essential to clarify that these models were not trained directly on student data, but rather pre-trained on the datasets above and subsequently deployed in a federated architecture. The federated process involved three to five local fine-tuning cycles per device, enabling the models to gradually specialize according to the emotional characteristics of the real-life educational environment, while preserving user privacy. To complement these public datasets, the models were not deployed in their pre-trained form only. Once integrated into the federated environment, each modality was fine-tuned locally using anonymized records derived from fundamental student interactions within the Moodle platform, including forum messages, voice participation, and facial expressions captured during hybrid sessions. This local fine-tuning process ensured that the models adapted to the specific linguistic, acoustic, and behavioral characteristics of the target educational population, while respecting privacy constraints. Importantly, no raw interaction data was centralized; only model updates were transmitted following the principles of federated learning. In this way, the training strategy combined the robustness of publicly validated datasets with the contextual specificity of real-world data, ensuring methodological consistency and ecological validity. To address the mismatch between the target emotional categories and the labels present in the pre-training corpora, additional open datasets and a harmonization strategy were incorporated. For facial modality, supplementary corpora such as AffectNet (Mollahosseini etal., 2019) were used to include classes not covered by FER2013 (Santoso and Kusuma, 2022), particularly contempt, while still relying on FER2013 as the baseline for basic facial emotions. In the audio modality, RAVDESS was expanded with resources like EMO-DB, IEMOCAP, and RECOLA (Joudeh etal., 2023; Khurana etal., 2024; Ong etal., 2024), which provide categories closely aligned with stress, confusion, and boredom. Prosodic dimensions from these corpora, mapped along valence–arousal axes, enabled the derivation of satisfaction and engagement-related cues. For text, datasets such as GoEmotions and education-specific corpora were integrated, ensuring coverage of states like motivation, anxiety, and curiosity through semantic mapping and weak supervision techniques (Demszky etal., 2020). This process followed a label-space harmonization approach in which semantically equivalent or proximate categories from different sources were merged into a unified taxonomy. Mapping was supported by distributional similarity measures and embedding-based alignment to maintain consistency across modalities. When labels were absent from the pre-training corpora but present in supplemental ones, transfer learning mechanisms were employed to transfer representations into the federated fine-tuning stage. Regarding FER2013, only the 32,298 publicly available images were used, as the remaining portion of the original corpus is restricted and inaccessible. The dataset is distributed into a predefined training split of 28,709 images and a test split of 3,589 images. To introduce a validation stage consistent with the 70/15/15 strategy applied across modalities, we further partitioned the training split by reallocating 15% of its samples (≈4,307 images) as a validation subset, while retaining 24,402 images for training. The original test split of 3,589 images was preserved without modification to serve as the final evaluation set. This procedure ensured methodological uniformity across modalities while maintaining compatibility with the canonical FER2013 evaluation protocol, thereby avoiding the use of non-public data and reinforcing the reproducibility of the experiments. Finally, specific high-level affective constructs, such as engagement, were not directly predicted by a single classifier but inferred through multimodal fusion. In these cases, the system combined audio-prosodic indicators, facial activation levels, and behavioral traces from LMS interactions to derive a composite state. This ensured that all emotional categories defined in the study were technically grounded, either through explicit dataset coverage, mapped proxies, or composite modeling strategies aligned with the federated architecture.
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 08 frontiersin.org 3.3.4 Definition of emotion classes by modality The definition and categorization of the emotional states targeted in this study were established to guarantee clarity, reproducibility, and technical rigor. Each modality—facial expressions, voice signals, and textual content—was associated with a set of emotional labels aligned with validated taxonomies in affective computing and educational psychology. This structured definition ensured that the classification tasks were consistent across modalities and that the evaluation of the federated system was traceable and comparable with existing models. For facial modality, the taxonomy proposed by Ekman and Rosenberg was adopted, as it provides a robust foundation for identifying emotions that are consistently expressed through facial movements (Ekman et al., 1998). Emotions such as happiness, sadness, anger, surprise, contempt, and disgust were selected because they present distinctive visual cues that can bequantified using convolutional neural networks. These categories have been repeatedly validated in emotion recognition studies, allowing for a reliable mapping between observable facial features and underlying affective states. In the voice modality, the classes of stress, confusion, and satisfaction were defined, given their strong correlation with variations in prosodic features such as pitch, intensity, and rhythm. These emotions are adequately represented in corpora like RAVDESS and are critical in educational contexts, where vocal modulation often reflects cognitive load and affective responses to learning tasks. By focusing on these specific states, the system captures meaningful indicators of students’ emotional dynamics in oral interactions (Bilotti etal., 2024). In addition to stress, confusion, and satisfaction, the model also incorporated two derived affective states—boredom and engagement. These states were not directly annotated in the base RAVDESS corpus. Still, they were obtained through the integration of IEMOCAP and RECOLA datasets, where prosodic patterns were mapped along the valence–arousal plane. Boredom was associated with low arousal and neutral-to-negative valence speech segments, while engagement corresponded to high arousal and positive valence prosodic patterns. These derived states were incorporated through label harmonization and validated during the federated fine-tuning phase, allowing the model to infer motivational intensity from voice cues. In the textual modality, emotions such as anxiety, motivation, and frustration were prioritized. These categories are highly relevant in written academic interactions, where students frequently express their affective states indirectly through language. Using transformer-based semantic embeddings, particularly BERT, the system was able to analyze the bidirectional context of written responses, capturing subtle variations in meaning that reflect students’ affective conditions (Sayeed etal., 2023). Beyond motivation, anxiety, and frustration, two additional affective states—curiosity and confusion—were integrated through semantic mapping using the GoEmotions corpus and education-specific text samples. Curiosity was identified through linguistic constructions reflecting positive exploratory intent (e.g., interrogative forms combined with positive sentiment). In contrast, confusion emerged as a composite category derived from frustration and uncertainty labels through weak supervision. These categories were retained during fine-tuning as they frequently occur in learning contexts, enabling more accurate modeling of cognitive-affective dynamics in student writing. The taxonomy, organized by modality, aligns with Table2, where basic emotions (e.g., happiness, sadness) coexist with derived and context-specific states (e.g., engagement, curiosity). This threefold definition of emotional classes provides a rigorous framework for the federated model, ensuring that each modality contributes in a complementary manner to the global detection process. The careful alignment of modalities with distinct emotional categories avoids overlaps, reduces ambiguity in classification, and reinforces the interpretability of the results obtained in world educational environments. 3.4 Development of the emotion detection model 3.4.1 Emotion detection models Different types of deep learning models are used to address emotion detection in students, tailored to the specific characteristics of each data modality: facial images, audio, and text. These models have been selected for their ability to learn complex, high-level representations of emotional data, and each one specializes in the type of data it is provided with. First, CNNs are employed for facial expression analysis, which can extract spatial features from facial images. CNNs are especially effective in computer vision tasks due to their ability to identify hierarchical patterns of information, ranging from simple features such as edges and textures to complex patterns, including emotions expressed on the face. The model is trained using high-resolution facial images, where the network learns to identify spatial relationships between key points on the face. CNNs operate by applying convolutional filters to images, where each filter Wk generates a feature map Ck as defined in Equation (8): = ∗ kk C WI (8) where I is the input image and * denotes the convolution operation. These feature maps are combined to extract emotions such as happiness, sadness, anger, or surprise. An RNN and an LSTM are used for voice tone analysis and are ideal for processing temporal data sequences such as audio signals. LSTMs are designed to capture long-term dependencies in audio sequences, which is crucial for identifying emotions that evolve in a conversation or speech (Hashmi and Yayilgan, 2024). LSTM parameters, such as input, output, and forget gates, allow the network to remember and forget information based on the temporal characteristics of the signal. Mathematically, the LSTM model is defined by the following Equations (9)– (13): ( ) σ − =⋅+ 1, t ft t f f WhX b (9) ( ) σ − =⋅+ 1, t it t i i Wh X b (10) ( ) − =⋅+ 1 ˆ tan , t ct t c C hW h X b (11) − = ∗ +∗ 1ˆ t t t tt C fC ic (12) ( ) = ∗tanh tt t ho C (13)
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 09 frontiersin.org where f t , i t , and o t are the forget, input, and output gates, x t is the temporal input (audio signal), h t is the hidden state, and C t is the cell state. A transformer-based model, specifically BERT, is used for text analysis, which is highly effective at processing text bidirectionally. The BERT model can understand the full context of a word within a sentence, as it examines both the preceding and following contexts of the word. This approach outperforms traditional one-way models and is particularly useful for understanding complex emotions in language. Text analysis in BERT is done through word embedding, where each word is transformed into a high-dimensional vector that captures its context. 3.4.2 Federated model The federated learning model implemented in this study enables emotion detection models to betrained in a decentralized manner, i.e., without centralizing sensitive data on a server. This approach is crucial for ensuring student privacy, as only model updates are shared, not the original data. Each student device uses data to train the emotion detection model during local training. This process is conducted locally, meaning that each student’s emotional data remains on their device. The model on each device is continuously tuned and improved as more emotional data is collected from the student’s interactions with educational content. The local model performs parameter updates using the gradient descent algorithm. Since the data is not centralized, training is carried out in parallel on each device without sharing information about the students’ data. The model parameters, which are weight vectors wi, are updated based on the local error calculated at each device, and the update follows the standard gradient rule, as expressed in Equation (14): η = − ⋅∇ i ii w ww L (14) where η is the learning rate, and ∇i wL is the gradient of the loss function L concerning the parameters wi. To integrate emotional information obtained from the three data modalities, facial images, voice recordings, and text inputs, the system employs a late fusion strategy. Each modality is processed independently on the student’s device using the respective specialized models: a CNN for facial expressions, an LSTM for voice tone, and a transformer (BERT) for textual data. Each model outputs a probability distribution over the predefined set of emotional classes. These three distributions are then combined using a weighted average, where the weights were empirically tuned during the development phase to optimize overall classification performance. The final emotional prediction corresponds to the class with the highest combined probability. This modular approach enables flexible processing even in scenarios where one or more modalities are temporarily unavailable (e.g., no audio input), ensuring the robustness and adaptability of the federated learning system. 3.4.3 Model aggregation Once the local model has been trained on each device, the model parameter updates are sent to the central server, which combines them using an aggregation process. In federate learning, this is done using the federated averaging algorithm. This method allows the central server to combine model updates without accessing the original learner data (Ren etal., 2024). Mathematically, federated averaging can be expressed as a weighted average of the local model updates, denoted as the local model updates ∆i w from each device i as shown in Equation (15): = ∆= ∆ ∑ 1 1N i i ww N (15) where N is the total number of devices participating in the training, the central server calculates the weighted average of the parameter updates ∆ i w and fine-tunes the global model, which is then distributed back to the devices to continue the training process. This local training and federated aggregation process enables continuous improvement of the emotion detection model without requiring centralization of data. It ensures that learners’ privacy is preserved while the model continues to learn collectively. The result is a more accurate and robust global model that can detect emotions in realtime, with sensitive data never shared outside local devices. 3.5 Evaluating model precision Several standard metrics are used in machine learning to evaluate the performance of the emotion detection model. These metrics are essential for understanding how the model identifies emotions, both in terms of precision and recall, and for gaining a comprehensive view of its performance. The metrics used in this study are as follows. Precision: Precision measures the proportion of correct predictions of a positive class (e.g., the “happy” emotion) among all projections of that class. Mathematically, it is expressed as: Precision is defined as shown in Equation (16): = + TP Precision TP FP (16) where: • TP (True Positives) are the correct predictions of the positive class. • FP (False Positives) are the incorrect predictions of the positive class. Recall measures the ability of the model to detect all positive instances (specific emotions) in the data. It is calculated as shown in Equation (17): = + TP Recall TP FN (17) where: • FN (False Negatives) are the positive instances that the model incorrectly classified as negative.
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 16 frontiersin.org Virtual environments demonstrate excellent precision and stability, suggesting that a wholly digital environment, without faceto-face interaction, enables the model to work with more consistent data. In comparison, hybrid environments, which combine in-person and virtual interactions, present more variability in results, likely due to social interactions and variations in nonverbal communication, which introduce additional noise into the emotion detection process. 4.6 Impact of the system on student behavior The emotion detection system implemented in the educational environment has a significant impact on students’ emotional and academic behavior. The results demonstrate how the detected emotions impact students’ academic engagement and performance in educational activities. In Figure7, three graphs clearly illustrate how the detected emotions impact various aspects of student behavior. Figure 7A shows the temporal evolution of students’ engagement and academic performance before and after receiving emotional feedback. A general improvement in both parameters is observed after feedback, especially in those students with positive emotions, such as motivation. However, the variability of the results suggests that negative emotions, such as stress and frustration, have an uneven impact on academic behavior, resulting in less consistent outcomes. Figure7B analyzes the relationship between the detected emotions and the levels of academic engagement. The results indicate that positive emotions, such as motivation, are associated with remarkably high levels of engagement. In contrast, emotions such as stress and FIGURE5 Evolution of model precision during the dynamic fitting process. FIGURE6 Model performance analysis in real-world conditions. (A) Device comparison. (B) Model precision for complex emotions. (C) Performance comparison in virtual vs. hybrid environments.
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 17 frontiersin.org frustration are related to lower engagement in academic activities. This behavior highlights the direct impact of emotions on students’ degree of engagement. Figure7C illustrates the impact of detected emotions on students’ academic performance. The results reflect that positive emotions are strongly associated with better grades, while negative emotions, such as stress and frustration, contribute to lower academic performance. This analysis underscores the importance of positive emotions in overall academic performance and highlights the need to mitigate the negative impact of emotions. The tabulated data offers a quantitative analysis that complements the results displayed in the charts. Table6 details the levels of academic engagement according to the detected emotions. Positive emotions, such as motivation, are associated with higher engagement in educational activities. In contrast, emotions such as stress and frustration are associated with significantly lower levels of participation, indicating that these emotions negatively impact the degree of student involvement in educational activities. To provide a more granular understanding of the relationship between emotional states and academic outcomes, Figure8 introduces two additional visualizations. These graphs complement the findings presented in Figure7 by disaggregating the data at the student level and exploring the distribution of academic performance across defined score brackets and emotional categories. Figure8A presents a scatter plot of academic performance for each student, grouped by the dominant emotion detected. This individualized analysis reveals that students who are consistently motivated achieve high scores, clustering between 80 and 100%. In contrast, students under emotional states such as stress, anxiety, or frustration display more dispersed outcomes, with frustration being most associated with scores below 70%. This visualization confirms the general trends observed in the aggregated results, while also highlighting outliers and inter-individual variability, which emphasizes the importance of emotional profiling for personalized interventions. Figure8B provides a histogram of performance distribution, where students are grouped into score brackets (e.g., 40–49, 50–59, …, 90–100) and categorized by emotional state. This analysis reveals a high concentration of motivated students in the top two brackets (80–89 and 90–100), while frustrated students are primarily found in the 50–69 range. The anxiety and stress groups present a broader distribution, reinforcing the notion of emotional heterogeneity in academic contexts. The histogram offers a frequency-based perspective, supporting the interpretation that positive emotions not only improve performance averages but also reduce variability in student outcomes. It is essential to clarify that the dataset used in Figure8 encompasses a broader academic population (N = 150) than the group involved in the system’s field validation (N = 58). While the 58-student subset was used to evaluate the system’s effectiveness in a real deployment, the extended analysis in Figure8 was designed to assess the variability and distribution of academic performance across different emotional profiles. This allows for a more robust statistical exploration of how distinct emotional states correlate with performance brackets, without conflicting with the empirical validation phase. Table 7 presents the results related to academic performance according to the emotions detected. Students who experience positive emotions tend to exhibit higher academic performance than those who face negative emotions, such as stress and frustration. Furthermore, grade variability, measured through standard deviation, is higher in students with negative emotions, suggesting more diverse responses in this group. This highlights the need for targeted interventions to support students experiencing these complex emotions and enhance their academic performance. To assess the practical effect of the emotion detection system on students’ engagement and academic performance, a comparison was conducted between participation records and academic outcomes registered before and after the system’s implementation. Specifically, interaction logs from Moodle and educational records from the FIGURE7 Impact of detected emotions on academic behavior. (A) Time evolution of participation and academic performance. (B) Relationship between emotions and levels of educational participation. (C) Relationship between emotions and academic performance. TABLE6 Levels of academic participation according to the emotions detected. Emotion Participation level (%) Average participation per student (%) Stress 60% 62% Anxiety 65% 63% Frustration 55% 58% Motivation 85% 82%
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 18 frontiersin.org institutional platform were analyzed over two equivalent academic periods, each covering a whole semester. The results show a 15% relative increase in average academic engagement, measured by the frequency of student interactions in forums, assignment submissions, and feedback requests. Likewise, the average academic performance, based on final course grades, improved by 12% after the integration of the emotion-aware feedback system. These findings suggest a positive shift in both behavioral and academic outcomes associated with the system’s deployment. Additionally, students’ responses to the adapted TPA questionnaire (Jian etal., 2000; Scharowski etal., 2025) revealed highly perceived usefulness (M = 4.2, SD = 0.6), reliability (M = 4.1, SD = 0.7), and privacy confidence (M = 4.4, SD = 0.5). These results suggest that the system’s federated design made a positive contribution to its acceptance and usability. The favorable perception of privacy protection likely enhanced students’ trust in the system, reinforcing their engagement and receptiveness to emotion-aware feedback during the academic period. While the improvements observed in student engagement and academic performance are strongly aligned with the implementation of the emotion-aware feedback system, it is essential to acknowledge that isolating the specific contribution of the federated learning component remains a methodologically complex task. However, the successful deployment of the privacy-preserving system in a real educational setting, without degrading performance or usability, reinforces the practical viability of our approach. These findings suggest that privacy-preserving, real-time emotional feedback can effectively support student engagement and learning outcomes, even in heterogeneous device environments. The evaluation of such integrated systems over more extended periods and in more diverse learning contexts will bekey to further confirming these benefits. Table8 summarizes the comparison. 4.7 Comparison with other emotion detection models Comparing the proposed model and other existing approaches to emotion detection is essential to highlight its advantages and areas for improvement. The performance analysis is based on precision, recall, and F1-score metrics, comparing our federated learning-based approach with centralized models such as DeepFace and hybrid transformer-based systems. Regarding privacy, we evaluate how centralized models rely on transferring sensitive data, whereas our proposal processes data locally. Furthermore, scalability is analyzed based on the system’s ability to manage large student populations without compromising performance, highlighting the flexibility of our solution to adapt to platforms such as Moodle. Table 9 summarizes the proposed model’s main features and results in comparison to well-known systems, including DeepFace, hybrid transformer-based models, and the commercial Affectiva SDK system. The table includes critical performance metrics, as well as aspects of privacy, scalability, and ease of integration into educational platforms. In terms of performance, the proposed model shows competitive results in precision, recall, and F1-score, approaching the values obtained by systems such as DeepFace. However, it outperforms centralized models by maintaining data privacy and avoiding the transfer of data to central servers. Furthermore, its federated approach enables greater scalability, allowing it to handle large student populations without significant performance degradation. Regarding integration, the proposed model stands out for its ability to integrate directly with the platform. One such feature is Moodle, which is not yet fully available in commercial systems such as the Affectiva SDK. This facilitates its adoption in educational FIGURE8 Relationship between emotional states and academic performance. (A) Individual academic performance grouped by predominant emotion. (B) Frequency of students by grade range according to detected emotion. TABLE7 Academic performance according to the emotions detected. Emotion Average rating (%) Average deviation (%) Stress 70% 7% Anxiety 75% 6% Frustration 65% 8% Motivation 85% 5%
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 19 frontiersin.org environments, reducing configuration costs and adapting to the specific needs of institutions. 5 Discussion The results obtained in this study show that the federated learningbased model for emotion detection presents significant advantages in terms of privacy, scalability, and performance, aligning with existing literature in several key aspects. Compared to DeepFace and hybrid transformer-based models, our approach achieves competitive metrics of precision (0.87), recall (0.85), and F1-score (0.86) while maintaining data privacy by processing it locally on users’ devices. This finding is consistent with those of Kairouz etal. (2021), who demonstrated that federated learning could preserve privacy without compromising performance. However, the observed variability in model performance under different conditions, such as changes in illumination and ambient noise, suggests that adapting the system to dynamic environments remains challenging, as pointed out by previous work on centralized models by Saha etal. (2024). The methodological process involved a design that combined emotional data collection using multiple modalities (images, audio, and text) with preprocessing to anonymize the data before local training. This approach addresses the need to protect users’ identities and optimizes the quality of the processed data (Ribeiro Junior and Kamienski, 2024). Implementing the federated model enabled the updating of parameters on local devices without transferring sensitive information, demonstrating its viability in educational environments with high privacy standards. However, the reliance on devices with heterogeneous capabilities introduced challenges in performance uniformity, particularly on platforms with limited resources. This problem is inherent to federated models and has been identified in the literature as an area requiring further optimization (Briguglio etal., 2024). In practical terms, the integration with Moodle facilitated real-world adoption and assessment of the model. The results obtained in the hybrid learning environment, encompassing both face-to-face and online classes, demonstrate that positive emotions, such as motivation, are associated with increased engagement and improved academic performance. Furthermore, the proposed approach addresses a critical need in emotion detection: the ability to operate on a scale without compromising students’ privacy. This represents a significant improvement over commercial systems such as Affectiva SDK, which, although efficient in terms of performance, do not offer the same data protection or customized integration with educational platforms (Kulke etal., 2020). This advance has direct implications for the design and deployment of scalable and ethically responsible educational technologies. Despite its contributions, the work presents limitations that must bediscussed to contextualize the findings appropriately. One of the main restrictions is the dependence on the quality of the devices the students use. Although federated learning is highly scalable, its performance can be affected by devices with limited processing capabilities, particularly in terms of latency and precision. This factor could bias the results in populations with unequal access to technology, posing equity challenges in implementing the system across different educational institutions (Mohapatra et al., 2024). Furthermore, although the model maintains high levels of privacy by processing data locally, variability in the quality of network connections could influence the effectiveness of model updates aggregated at the central server, especially in environments with inconsistent network infrastructure. Additionally, although the devices used—smartphones, tablets, and laptops—were heterogeneous and reflected typical student hardware, no stratified benchmarking was performed to assess model behavior across different device types. The system was designed to belightweight and platform-independent; however, variations in CPU, memory, or sensor resolution could have introduced minor discrepancies in inference time or prediction accuracy. Future work should include performance audits across device categories to better understand and optimize real-world deployments in diverse educational settings. Finally, an additional key limitation of this study is that user perceptions regarding the ability of the federated model to preserve their privacy were not evaluated. While the technical design ensures that sensitive data TABLE8 Academic indicators before and after system deployment. Indicator Before implementation After implementation Relative change Avg. engagement score (%) 67.5 77.6 +15.0% Avg. academic performance (%) 71.2 79.7 +12.0% Std. dev. of performance 7.4 6.1 — The engagement score was computed from normalized interaction metrics (forum activity, task completion, and time-on-task). Performance data were drawn from final course grades in both terms. TABLE9 Comparison of the proposed federated model with other emotion detection systems. Feature Proposed model (federated) DeepFace (centralized) Transformer-based hybrid Affectiva SDK (commercial) Precision 87% 90% 85% 88% Recall 85% 88% 84% 86% F1-score 86% 89% 84.5% 87% Privacy Local data protected Centralized data Partially localized data Centralized data Scalability Highly adaptable to mediumto large-sized populations Low, limited to centralized environments Moderate, computationally dependent Low, designed for small groups Ease of integration with LMS Direct integration with Moodle Requires advanced configuration Partial support Not specified
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 20 frontiersin.org remains on local devices, the actual trust and acceptance of such mechanisms by students and educators remain unexplored. Future work will incorporate user-centered studies to assess these perceptions, complementing technical validation with empirical evidence of usability and trustworthiness. Another significant limitation is the system’s sensitivity to adverse conditions, such as sudden changes in lighting or high ambient noise levels. Although the preprocessing methods improved the system’s robustness against these conditions, less common emotions, such as frustration or anxiety, presented higher error rates in their detection. This could bedue to insufficient examples of these emotions in the dataset used for initial training, a problem widely documented in the literature on emotion detection (Wang etal., 2024b). Addressing this limitation will require expanding the dataset to include more representative examples of these emotions and implementing transfer learning techniques to improve the model’s generalization. From a methodological perspective, the model assumes that emotions detected in educational activities are consistent with students’ emotional states. However, this assumption might beinvalid, as external factors unrelated to the academic environment may influence the emotional expressions detected. This aspect may introduce biases in interpreting the results, especially if used to assess students’ emotional well-being or personalize educational feedback. Mitigating this problem will require a more holistic approach that combines emotion detection with other contextual metrics, such as cognitive load or social interaction. The study demonstrates the viability and potential of federated emotion detection systems for real-world educational settings, while identifying key challenges that must beaddressed for widespread implementation. The findings provide a foundation for future research to improve the robustness, equity, and contextual awareness of emotion-aware learning technologies. The results show a 15% relative increase in average academic engagement, measured by the frequency of student interactions in forums, assignment submissions, and feedback requests. Likewise, the average academic performance, based on final course grades, improved by 12% after the integration of the emotion-aware feedback system. However, it is essential to note that these findings are based on descriptive analysis. No statistical significance tests, such as regression models or paired hypothesis tests, were applied to determine whether these differences are statistically significant or attributable solely to the system’s deployment. As such, the results suggest a potential positive shift in both behavioral and academic outcomes, but do not establish a causal relationship. Future work should include inferential statistical analysis to validate the observed improvements. The successful integration of the system into Moodle highlights its practical applicability and offers several implications for large-scale educational deployment. Unlike commercial systems such as Affectiva SDK (Kulke etal., 2020), which often require proprietary environments and lack educational customization, our open architecture facilitates direct alignment with existing learning platforms. Given its privacypreserving design and low computational requirements, the system can beadopted in institutions with varied infrastructure levels without significant technical constraints. However, effective implementation depends on institutional policies and the readiness of educators. Prior studies (Huang etal., 2023) there is a need to address teacher training in interpreting emotional analytics and to establish governance frameworks that regulate ethical use. Our findings emphasize that teacher capacity-building and policy alignment are necessary to translate emotional insights into meaningful pedagogical actions. Moreover, unlike transformer-based systems (Teng etal., 2024) our system supports decentralized scalability, which is particularly beneficial for applications that require extensive cloud resources, suggesting feasibility for national or multi-institutional deployments with minimal cost and strong alignment to educational values. 6 Conclusions and future work This study demonstrates that a federated learning-based approach for emotion detection is both effective and practical in educational environments. The model achieved high precision, recall, and F1-score values while preserving student data privacy and enabling scalability. Its integration into Moodle confirmed that the system can operate in real academic settings with minimal friction, offering real-time emotional feedback without compromising confidentiality. The implementation showed a measurable impact on student outcomes: students whose positive emotions were detected and responded to exhibited a 15% increase in academic engagement and a 12% improvement in performance. These results support the effectiveness of the approach and its value as a tool for emotionally adaptive learning. However, the system still faces limitations. Its performance depends on the heterogeneity of user devices, and detecting complex emotions like frustration or anxiety remains challenging due to their underrepresentation in the dataset. These constraints did not compromise the system’s viability, but instead highlighted areas for future improvement. Future work will focus on optimizing the model for low-resource devices and incorporating synthetic data and transfer learning techniques to enhance the diversity of emotions. Additionally, exploring further behavioral and cognitive indicators will help refine emotional inference and expand the system’s pedagogical impact. Data availability statement The data analyzed in this study is subject to the following licenses/ restrictions: the dataset used in this study contains sensitive emotional information collected from students within a university environment. Due to privacy considerations and institutional regulations, the dataset is not publicly available. Access is restricted to authorized researchers under data-sharing agreements that ensure compliance with ethical guidelines and privacy laws. Requests for data access may beconsidered on a caseby-case basis and require approval from the institutional ethics committee and the data controller. Requests to access these datasets should bedirected to william.vil[email protected]. Ethics statement The studies involving humans were approved by “Gamificación educativa potenciada por inteligencia artificial”. File number UA-2025-05-24. The studies were conducted in accordance with the local legislation and institutional requirements. Written informed consent for participation was not required from the
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 21 frontiersin.org participants or the participants’ legal guardians/next of kin because According to Ministerial Agreement No. 0005-2022 of the Ministry of Public Health of Ecuador, written informed consent is not required for studies that do not involve direct intervention on human beings, use of biological samples, participation of vulnerable populations, or access to confidential personal data. This study complied with all these conditions, as it was limited to the use of anonymized emotional data collected through non-invasive educational interactions. Author contributions RG: Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft. WV-C: Conceptualization, Formal analysis, Investigation, Methodology, Supervision, Validation, Visualization, Writing – review & editing. SL-M: Conceptualization, Supervision, Validation, Visualization, Writing– review & editing. Funding The author(s) declare that no financial support was received for the research and/or publication of this article. Conflict of interest The authors declare that the research was conducted in the absence of any commercial or financial relationships that could beconstrued as a potential conflict of interest. Generative AI statement The authors declare that no Gen AI was used in the creation of this manuscript. Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If youidentify any issues, please contact us. Publisher’s note All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may beevaluated in this article, or claim that may bemade by its manufacturer, is not guaranteed or endorsed by the publisher. References Ajagbe, S. A., Adeniji, O. D., Olayiwola, A. A., and Abiona, S. F. (2024). Advanced encryption standard (AES)-based text encryption for near field communication (NFC) using Huffman compression. SN Comput. Sci. 5:156. doi: 10.1007/s42979-023-02486-6 Alavi, H., Seifi, M., Rouhollahei, M., Rafati, M., and Arabfard, M. (2024). Development of local software for automatic measurement of geometric parameters in the proximal femur using a combination of a deep learning approach and an active shape model on X-ray images. J. Imaging Inform. Med. 37, 633–652. doi: 10.1007/s10278-023-00953-3 Almalki, J., Alshahrani, S. M., and Khan, N. A. (2024). A comprehensive secure system enabling healthcare 5.0 using federated learning, intrusion detection and blockchain. PeerJ Comput. Sci. 10:e1778. doi: 10.7717/peerj-cs.1778 An, Y., Lee, J., Bak, E. S., and Pan, S. (2023). Deep facial emotion recognition using local features based on facial landmarks for a security system. Comput. Mater. Contin. 76:1817–1832. doi: 10.32604/cmc.2023.039460 Anand, M., and Babu, S. (2024). Multi-class facial emotion expression identification using DL-based feature extraction with classification models. Int. J. Comput. Intell. Syst. 17:2–17. doi: 10.1007/s44196-024-00406-x Bilotti, U., Bisogni, C., De Marsico, M., and Tramonte, S. (2024). Multimodal emotion recognition via convolutional neural networks: comparison of different strategies on two multimodal datasets. Eng. Appl. Artif. Intell. 130:1–13. doi: 10.1016/j.engappai.2023.107708 Briguglio, W., Yousef, W. A., Traore, I., and Mamun, M. (2024). Federated supervised principal component analysis. IEEE Trans. Inf. Forensics Secur. 19, 646–660. doi: 10.1109/TIFS.2023.3326981 Chen, Q., Ling, Z., Jiang, H., Zhu, X., Wei, S., and Inkpen, D. (2017). Enhanced LSTM for natural language inference, in ACL 2017– 55th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (Long Papers). doi: 10.18653/v1/P17-1152 Demszky, D., Movshovitz-Attias, D., Ko, J., Cowen, A., Nemade, G., and Ravi, S. (2020). GoEmotions: a dataset of fine-grained emotions, in Proceedings of the Annual Meeting of the Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.372 Di Dario, D., Pentangelo, V., Colella, M. I., Palomba, F., and Gravino, C. (2024). Collecting and implementing ethical guidelines for emotion recognition in an educational metaverse, in Adjunct proceedings of the 32nd ACM conference on user modeling, adaptation and personalization, (New York, NY, USA: Association for Computing Machinery), 549–554. doi: 10.1145/3631700.3665232 Doriguzzi-Corin, R., and Siracusa, D. (2024). Flad: adaptive federated learning for DDoS attack detection. Comput. Secur. 137:103597. doi: 10.1016/j.cose.2023.103597 Ekman, P., Rosenberg, E. L., and Heller, M. (1998). What the face reveals. Basic and applied studies of spontaneous expression using the facial action coding system (FACS). Psychothérapies 18, 179–180. doi: 10.1177/00030651221107681 Elisondo, R. C., Suárez-Lantarón, B., and García-Perales, N. (2024). University students´ emotions in forced remote education. An exploratory study in Spain and Argentina. Multidiscip. J. Educ. Res. 14:59–78. doi: 10.17583/remie.10251 Hammann, T., Schwartze, M. M., Zentel, P., Schlomann, A., Even, C., Wahl, H. W., et al. (2022). The challenge of emotions—an experimental approach to assess the emotional competence of people with intellectual disabilities. Disabilities 2, 611–625. doi: 10.3390/disabilities2040044 Hamon, R., Junklewitz, H., Sanchez, I., Malgieri, G., and De Hert, P. (2022). Bridging the gap between AI and explainability in the GDPR: towards trustworthiness-by-design in automated decision-making. IEEE Comput. Intell. Mag. 17, 72–85. doi: 10.1109/MCI.2021.3129960 Hanisch, S., Todt, J., Patino, J., Evans, N., and Strufe, T. (2024). A false sense of privacy: towards a reliable evaluation methodology for the anonymization of biometric data. Proc. Priv. Enhancing Technol. 2024, 116–132. doi: 10.56553/popets-2024-0008 Hashmi, E., and Yayilgan, S. Y. (2024). Multi-class hate speech detection in the Norwegian language using FAST-RNN and multilingual fine-tuned transformers. Complex Intell. Syst. 10, 4535–4556. doi: 10.1007/s40747-024-01392-5 Hellmann, F., Mertes, S., Benouis, M., Hustinx, A., Hsieh, T.-C., Conati, C., et al. (2024). GANonymization: a GAN-based face anonymization framework for preserving emotional expressions. ACM Trans. Multimedia Comput. Commun. Appl. 21, 1–27. doi: 10.1145/3641107 Huang, F., Yang, N., Chen, H., Bao, W., and Yuan, D. (2023). Distributed online multilabel learning with privacy protection in internet of things. Appl. Sci. 13:2713. doi: 10.3390/app13042713 Jian, J.-Y., Bisantz, A. M., and Drury, C. G. (2000). Foundations for an empirically determined scale of trust in automated systems. Int. J. Cogn. Ergon. 4, 53–71. doi: 10.1207/S15327566IJCE0401_04 Jiang, W., Li, M., Shabaz, M., Sharma, A., and Haq, M. A. (2023). Generation of voice signal tone sandhi and melody based on convolutional neural network. ACM Trans. Asian Low-Resour. Lang. Inf. Process. 22, 1–13. doi: 10.1145/3545569
Gutiérrez et al. 10.3389/frai.2025.1644844 Frontiers in Artificial Intelligence 22 frontiersin.org Joudeh, I. O., Cretu, A. M., Bouchard, S., and Guimond, S. (2023). Prediction of continuous emotional measures through physiological and visual data. Sensors 23:5613. doi: 10.3390/s23125613 Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., et al. (2021). Advances and open problems in federated learning. Found. Trends® Mach. Learn. 14, 1–210. doi: 10.1561/2200000083 Kaiss, W., Mansouri, K., and Poirier, F. (2023). Pre-evaluation with a personalized feedback conversational agent integrated in Moodle. Int. J. Emerg. Technol. Learn. 18, 177–189. doi: 10.3991/ijet.v18i06.36783 Kerman, N. T., Banihashem, S. K., Karami, M., Er, E., van Ginkel, S., and Noroozi, O. (2024). Online peer feedback in higher education: a synthesis of the literature. Educ. Inf. Technol. (Dordr) 29, 763–813. doi: 10.1007/s10639-023-12273-8 Khurana, Y., Gupta, S., Sathyaraj, R., and Raja, S. P. (2024). Robinnet: a multimodal speech emotion recognition system with speaker recognition for social interactions. IEEE Trans. Comput. Soc. Syst. 11, 478–487. doi: 10.1109/TCSS.2022.3228649 Kotwal, K., Bhattacharjee, S., Abbet, P., Mostaani, Z., Wei, H., Wenkang, X., et al. (2022). Domain-specific adaptation of CNN for detecting face presentation attacks in NIR. IEEE Trans. Biom. Behav. Identity Sci. 4, 135–147. doi: 10.1109/TBIOM.2022.3143569 Kulke, L., Feyerabend, D., and Schacht, A. (2020). A comparison of the Affectiva iMotions facial expression analysis software with EMG for identifying facial expressions of emotion. Front. Psychol. 11:1–9. doi: 10.3389/fpsyg.2020.00329 Labidi, A., Ouamani, F., and Saoud, N. B. B. (2021). An ontology based text approach for culture aware emotion mining: a Moodle plugin, in 2021 International Conference of Women in Data Science at Taif University, WiDSTaif 2021. doi: 10.1109/WIDSTAIF52235.2021.9430232 Lee, J. P., Jang, H., Jang, Y., Song, H., Lee, S., Lee, P. S., et al. (2024). Encoding of multimodal emotional information via personalized skin-integrated wireless facial interface. Nat. Commun. 15:530. doi: 10.1038/s41467-023-44673-2 Lyu, S. (2023). Facial expression recognition based on Mini_Xception. Highl. Sci. Eng. Technol. 39, 1178–1187. doi: 10.54097/hset.v39i.6726 Merity, S., Keskar, N. S., and Socher, R. (2018) Regularizing and optimizing LSTM language models, in 6th International Conference on Learning Representations, ICLR 2018– Conference Track Proceedings Mishra, S. K., Kumar, N. S., Rao, B., Brahmendra, and Teja, L. (2024). Role of federated learning in edge computing: a survey. J. Auton. Intell. 7:1–22. doi: 10.32629/jai.v7i1.624 Mohapatra, R. K., Jolly, L., Lyngdoh, D. C., Mourya, G. K., Changaai Mangalote, I. A., Alam, S. I., et al. (2024). A comprehensive survey to study the utilities of image segmentation methods in clinical routine. Netw. Modeling Anal. Health Inform. Bioinform. 13:2–26. doi: 10.1007/s13721-023-00436-z Mollahosseini, A., Hasani, B., and Mahoor, M. H. (2019). AffectNet: a database for facial expression, valence, and arousal computing in the wild. IEEE Trans. Affect. Comput. 10, 18–31. doi: 10.1109/TAFFC.2017.2740923 Mukta, M. S. H., Ahmad, J., Zaman, A., and Islam, S. (2024). Attention and metaheuristic based general self-efficacy prediction model from multimodal social media dataset. IEEE Access 12, 36853–36873. doi: 10.1109/ACCESS.2024.3373558 Mutawa, A. M., and Hassouneh, A. (2024). Multimodal real-time patient emotion recognition system using facial expressions and brain EEG signals based on machine learning and log-sync methods. Biomed. Signal Process. Control. 91:2–6. doi: 10.1016/j.bspc.2023.105942 Nandan, D., Kanungo, J., and Mahajan, A. (2024). An error-efficient Gaussian filter for image processing by using the expanded operand decomposition logarithm multiplication. J. Ambient. Intell. Humaniz. Comput. 15:1045–1052. doi: 10.1007/s12652-018-0933-x Ong, K. L., Lee, C. P., Lim, H. S., Lim, K. M., and Alqahtani, A. (2024). MaxMViTMLP: multiaxis and multiscale vision transformers fusion network for speech emotion recognition. IEEE Access 12, 18237–18250. doi: 10.1109/ACCESS.2024.3360483 Pedrycz, W. (2023). Advancing federated learning with granular computing. Fuzzy Inf. Eng. 15, 1–13. doi: 10.26599/FIE.2023.9270001 Pirrone, C., Varrasi, S., Platania, G. A., and Castellano, S. (2021) Face-to-face and online learning: the role of technology in students’ metacognition, in CEUR Workshop Proceedings Rai, R., and Gupta, S. (2022). A novel learning management system based on microservice architecture using Moodle. Int. J. Creat. Res. Thoughts 10:d356–d363. Ren, H., Anicic, D., and Runkler, T. A. (2023). TinyReptile: TinyML with federated meta-learning, in Proceedings of the International Joint Conference on Neural Networks. doi: 10.1109/IJCNN54540.2023.10191845 Ren, H., Deng, J., Xie, X., Ma, X., and Wang, Y. (2024). Fedboosting: federated learning with gradient protected boosting for text recognition. Neurocomputing 569:1–12. doi: 10.1016/j.neucom.2023.127126 Ribeiro Junior, F. M., and Kamienski, C. A. (2024). Federated learning for performance behavior detection in a fog-IoT system. Internet Things 25:101078. doi: 10.1016/j.iot.2024.101078 Saha, P. K., Arya, D., and Sekimoto, Y. (2024). Federated learning–based global road damage detection. Comput. Aided Civ. Inf. Eng. 39, 2223–2238. doi: 10.1111/mice.13186 Santoso, B. E., and Kusuma, G. P. (2022). Facial emotion recognition on FER2013 using VGGSpinalNet. J. Theor. Appl. Inf. Technol. 100:2088–2102. Sayeed, M. S., Mohan, V., and Muthu, K. S. (2023). BERT: a review of applications in sentiment analysis. HighTech Innov. J. 4, 453–462. doi: 10.28991/HIJ-2023-04-02-015 Scharowski, N., Perrig, S. A. C., von Felten, N., Aeschbach, L. F., Opwis, K., Wintersberger, P., et al (2025). To trust or distrust AI: a questionnaire validation study, in Proceedings of the 2025 ACM conference on fairness, accountability, and transparency, (NewYork, NY, USA: Association for Computing Machinery), 361–374. doi: 10.1145/3715275.3732025 Sengupta, D., Khan, S. S., Das, S., and De, D. (2024). FedEL: federated education learning for generating correlations between course outcomes and program outcomes for internet of education things. Internet Things 25:101056. doi: 10.1016/j.iot.2023.101056 Shchedrina, E., Valiev, I., Sabirova, F., and Babaskin, D. (2021). Providing adaptivity in Moodle LMS courses. Int. J. Emerg. Technol. Learn. 16:95. doi: 10.3991/ijet.v16i02.18813 Shmelova, T., Smolanka, V., Sikirda, Y., and Sechko, O. (2024). Real-time monitoring and diagnostics of the person’s emotional state and decision-making in extreme situations for healthcare. Decis. Mak. Anal.. 11–32. doi: 10.55976/dma.22024121911-32 Shwe Sin, T., and Khin, O. (2022). Facial expressions classification on android smartphone for a user, in Lecture notes in networks and systems. doi: 10.1007/978-981-16-1781-2_21 Souali, K., Rahmaoui, O., and Ouzzif, M. (2023). MentorBot: a traceability-based recommendation Chatbot for Moodle, in Lecture notes in networks and systems. doi: 10.1007/978-3-031-26384-2_37 Subakti, A., Murfi, H., and Hariadi, N. (2022). The performance of BERT as data representation of text clustering. J. Big Data 9:15. doi: 10.1186/s40537-022-00564-9 Takase, S., and Kiyono, S. (2023). Lessons on parameter sharing across layers in transformers, in Proceedings of the annual meeting of the Association for Computational Linguistics. doi: 10.18653/v1/2023.sustainlp-1.5 Teng, S., Liu, J., Huang, Y., Chai, S., Tateyama, T., Huang, X., et al. (2024). An intraand inter-emotion transformer-based fusion model with homogeneous and diverse constraints using multi-emotional audiovisual features for depression detection. IEICE Trans. Inf. Syst. E107.D, 342–353. doi: 10.1587/transinf.2023HCP0006 Wang, X. (2024). Research on the digital preservation of intangible cultural heritage of folk dance art category. Appl. Math. Nonlinear Sci. 9:1–16. doi: 10.2478/amns.2023.2.01500 Wang, R., Chen, T., Wang, J., Lai, J., Li, J., Zhang, M., et al. (2024a). Integrated comprehensive analysis method for education quality with federated learning, in Communications in Computer and Information Science 691. doi: 10.1007/978-981-99-9492-2_37 Wang, A., Meng, Z., Zhao, B., and Zhang, F. (2024). Using social media data to research the impact of campus green spaces on students’ emotions: a case study of Nanjing campuses. Sustainability 16. doi: 10.3390/su16020691 Wang, R., Song, P., and Zheng, W. (2024b). Graph-diffusion-based domain-invariant representation learning for cross-domain facial expression recognition. IEEE Trans. Comput. Soc. Syst. 11, 4163–4174. doi: 10.1109/TCSS.2024.3355113 Wong, R. Y., Chong, A., and Cooper Aspegren, R. (2023). Privacy legislation as business risks: how GDPR and CCPA are represented in technology companies’ investment risk disclosures. Proc. ACM on Hum.-Comput. Interact. 7, 1–26. doi: 10.1145/3579515 Woodward, K., Kanjo, E., and Tsanas, A. (2024). Combining deep learning with signal-image encoding for multi-modal mental wellbeing classification. ACM Trans. Comput. Healthcare 5, 1–23. doi: 10.1145/3631618 Xu, B., Gu, X., Wang, J., Zhang, Y., Zhu, J., and Zhang, W. (2024). Deformation control method for active shape morphing lattice structure using topology optimization approach. Smart Mater. Struct. 33:025006. doi: 10.1088/1361-665X/ad1bab Zhang, Y., Li, Z., Han, X., Ding, S., Li, J., Wang, J., et al. (2024). Pseudo-data based self-supervised federated learning for classification of histopathological images. IEEE Trans. Med. Imaging 43, 902–915. doi: 10.1109/TMI.2023.3323540 Zhang, H., Wu, T., Cheng, S., and Liu, J. (2024). Cc-fedavg: computationally customized federated averaging. IEEE Internet Things J. 11, 4826–4841. doi: 10.1109/JIOT.2023.3300080 Zhou, J., Pal, S., Dong, C., and Wang, K. (2024). Enhancing quality of service through federated learning in edge-cloud architecture. Ad Hoc Netw. 156:103430. doi: 10.1016/j.adhoc.2024.103430