scieee AI-readable full text Open interactive document viewer

Effect of Action Units, Viewpoint and Immersion on Emotion Recognition Using Dynamic Virtual Faces

Miguel A, Vicente-Querol; Antonio, Fernández-Caballero; Pascual, González; Luz M, González-Gualda; Patricia, Fernández-Sotos; José P, Molina; Arturo S, García

Abstract

Facial affect recognition is a critical skill in human interactions that is often impaired in psychiatric disorders. To address this challenge, tests have been developed to measure and train this skill. Recently, virtual human (VH) and virtual reality (VR) technologies have emerged as novel tools for this purpose. This study investigates the unique contributions of different factors in the communication and perception of emotions conveyed by VHs. Specifically, it examines the effects of the use of action units (AUs) in virtual faces, the positioning of the VH (frontal or mid-profile), and the level of immersion in the VR environment (desktop screen versus immersive VR). Thirty-six healthy subjects participated in each condition. Dynamic virtual faces (DVFs), VHs with facial animations, were used to represent the six basic emotions and the neutral expression. The results highlight the important role of the accurate implementation of AUs in virtual faces for emotion recognition. Furthermore, it is observed that frontal views outperform mid-profile views in both test conditions, while immersive VR shows a slight improvement in emotion recognition. This study provides novel insights into the influence of these factors on emotion perception and advances the understanding and application of these technologies for effective facial emotion recognition training.

Full text

November 4, 2025 12:11 ImmersionFAR International Journal of Neural Systems, Vol. 33, No. 08 (2023) 1–18 © World Scientific Publishing Company EFFECT OF ACTION UNITS, VIEWPOINT AND IMMERSION ON EMOTION RECOGNITION USING DYNAMIC VIRTUAL FACES Miguel A. Vicente-Querol Instituto de Investigaci´on en Inform´atica, Universidad de Castilla-La Mancha, Albacete, 02071, Spain Antonio Fern´andez-Caballero and Pascual Gonz´alez Departamento de Sistemas Inform´aticos, Universidad de Castilla-La Mancha, Albacete, 02071, Spain Instituto de Investigaci´on en Inform´atica, Universidad de Castilla-La Mancha, Albacete, 02071, Spain Biomedical Research Networking Centre in Mental Health, Instituto de Salud Carlos III, Madrid, 28029, Spain Luz M. Gonz´alez-Gualda Servicio de Salud Mental, Complejo Hospitalario Universitario de Albacete, Albacete, 02004, Spain Patricia Fern´andez-Sotos Servicio de Salud Mental, Complejo Hospitalario Universitario de Albacete, Albacete, 02004, Spain Biomedical Research Networking Centre in Mental Health, Instituto de Salud Carlos III, Madrid, 28029, Spain Jos´e P. Molina and Arturo S. Garc´ıa Departamento de Sistemas Inform´aticos, Universidad de Castilla-La Mancha, Albacete, 02071, Spain Instituto de Investigaci´on en Inform´atica, Universidad de Castilla-La Mancha, Albacete, 02071, Spain E-mail: arturosimon.gar[email protected] Facial affect recognition is a critical skill in human interactions that is often impaired in psychiatric disorders. To address this challenge, tests have been developed to measure and train this skill. Recently, virtual human (VH) and virtual reality (VR) technologies have emerged as novel tools for this purpose. This study investigates the unique contributions of different factors in the communication and perception of emotions conveyed by VHs. Specifically, it examines the effects of the use of action units (AUs) in virtual faces, the positioning of the VH (frontal or mid-profile), and the level of immersion in the VR environment (desktop screen vs. immersive VR). Thirty-six healthy subjects participated in each condition. Dynamic virtual faces (DVFs), VHs with facial animations, were used to represent the six basic emotions and the neutral expression. The results highlight the important role of the accurate implementation of AUs in virtual faces for emotion recognition. Furthermore, it is observed that frontal views outperform mid-profile views in both test conditions, while immersive VR shows a slight improvement in emotion recognition. This study provides novel insights into the influence of these factors on emotion perception and advances the understanding and application of these technologies for effective facial emotion recognition training. Keywords: emotion recognition; virtual humans; immersion level; virtual reality. 1. Introduction Emotion recognition1is a fundamental aspect of social interaction, providing us with the ability to interpret the emotions and intentions of other individuals.2–5 In particular, facial affect recognition6, 7 allows us to discriminate between different types of emotional expressions displayed on faces. However, several neuropsychiatric conditions8can im1 November 4, 2025 12:11 ImmersionFAR 2M.A. Vicente-Querol et al. pair this ability, including schizophrenia,9depression,10–14 Parkinson’s disease,15 and autism spectrum disorder.16 It is noteworthy that in the field of emotion recognition, not every study is related to interpreting what emotion is being shown to participants. Applying machine learning or deep learning techniques17–20 and combining them with the participants’ physiological signals measured by different types of medical techniques such as EEG,21–24 we can gain a better and deeper understanding of how the brain works when recognizing emotions, so that the tools to express these emotions can be improved to adapt to different type of situations and persons, or even recognize emotions depending on the speaker’s speech25 or writing,26 which can lead to the development of cognitive therapies for different types of neuropsychiatric conditions, as it was stated previously. In emotion recognition research using participants, the six basic facial expressions (anger,fear, surprise,sadness,disgust,joy) plus the neutral expression, as identified by Ekman et al.,27 are a widely cited and commonly used benchmark. The authors also developed the Facial Action Coding System (FACS),28 which identifies specific muscle movements, called Action Units (AUs), that are associated with different emotions. Each emotion is a combination of several of these AUs. For example, in surprise the changes in facial appearance can be found in the eyebrows being raised, the jaw dropping, and the mouth remaining open, while in fear the eyebrows are also raised but tense, the eyes are also open, as is the mouth, but the lips are tensed, in contrast to surprise.Inanger the eyebrows are lowered and tense, and the mouth may remain closed or open, but the lips are tightened. All these changes in the face are translated into different AUs, for example surprise is broken down into the sum of AUs 1, 2, 5 and 26, fear shares the same AUs as surprise plus AU4 and AU20 and anger is the sum of AUs 4, 5, 7, 17, 23 and 25, which makes it clear that some emotions share AUs and facial similarity. In the literature, fear is usually the least recognized emotion and it is highly confounded with surprise.29–31 The reason why some participants fail to accurately identify emotions is still unclear. One of the most discussed and commented on reasons for confusion between emotions is that some of them share AUs. For example, previous studies32, 33 discussed the high confusion between fear and surprise because they share the most AUs (as it was presented earlier). Other studies34, 35 have addressed the role these AUs play in emotion recognition. Also, as shown in previous works,30, 31, 36 the presentation angle of the faces (frontal or in mid-profile view) affects the accuracy in emotion recognition, especially in the recognition of fear and increases its confusion with surprise. Virtual humans (VHs) have been used in previous studies as stimuli for facial affect recognition.37, 38 The use of VHs offers several advantages over traditional methods using paper-based stimuli,39 including the ability to change facial expressions according to different races, ages, and genders of virtual characters, add animation to expressions, incorporate different contexts and backgrounds, and enhance the dataset with images taken from different angles. However, previous studies have mainly used VHs in desktop virtual reality (VR) setups, which does not take full advantage of the increased immersion of VR technology. VR technology aims to overcome the limitations of traditional human-computer interaction by providing a more natural and humanlike interaction, free from the constraints of devices such as keyboards, mice, or flat screens.40 Numerous studies have compared the performance of various tasks between desktop and VR configurations41–44 to assess the maturity of VR technology. The improvement in task performance in immersive VR experiences is often associated with an increased sense of presence.45, 46 Presence in VR refers to the feeling of “being there” within the computer-generated virtual world, even if it is a virtual representation of a distant planet.47 Previous studies on emotion recognition using VHs can be divided into two main groups: those that compare sets of pictures or videos of real people with VHs in a non-immersive VR setup38, 48–51 and those that compare these sets with VHs in an immersive VR setup.52 Overall, they concluded that, at least with VR setups, emotion recognition results can be as good as those obtained using real people as stimuli. They found that the dynamism and personalization of VH is a great advantage. But there are some drawbacks. Disgust38, 48, 52 is the worst recognized emotion in VR setups (both immersive and non-immersive), concluding that this lack of success is due to the difficulty of reproducing in avatars the wrinkles that appear in the nose with AU9. November 4, 2025 12:11 ImmersionFAR Emotion Recognition using Dynamic Virtual Faces 3 To the best of our knowledge, only one study has compared the results of using a non-immersive VR against an immersive VR in emotion recognition.53 They concluded that the degree of immersion in virtual reality has no negative impact on the recognition task, reaching an accuracy of 71.33% of hits in both conditions, and that the accuracy achieved in face recognition implies that immersive virtual reality has potential for exploring the emotion recognition process. As in the previous studies, they also found the same problems with disgust, as it was the emotion with the worst hit rate of all emotions (44.44%). This may be related to the fact that there is no mention of wrinkles in the paper, and they are not visible in the screenshots provided. This omission obviously affects the accurate representation of emotions, especially disgust, and possibly other emotions as well. In addition, surprise is not included in this study, which limits its findings and potential generalization by failing to address the potential confusion between surprise and fear, which, as noted above, is common in the literature. Following this trend, our research group has conducted a series of studies on virtual humans and their use for emotion recognition. The design process and validation of our set of dynamic virtual faces (DVFs) has already been described.54 In another paper,55 a comprehensive analysis was performed comparing our DVFs with the validated Penn Emotion Recognition Test (ER-40),39 taking into account gender, age, and camera angle. The DVF set outperformed the ER-40, achieving 88.25% accuracy compared to 82.30%. After successfully validating this set of DVFs with healthy participants, a subsequent phase of experimentation involved conducting tests with people with neurological conditions that could affect their ability to recognize emotions. Thus, the same set of DVFs was used to assess patients diagnosed with schizophrenia56 and depression.57, 58 The next step was to move from a nonimmersive virtual reality environment (desktop) to an immersive virtual reality environment (using a head mounted display, HMD). Thus, a study was conducted to find the optimal interpersonal distance between humans and DVFs in an immersive virtual reality setup.59 Finally, or last step, was the comparison between immersive and non-immersive virtual reality.60 A preliminary study briefly outlined a comparison between these two conditions in terms of accuracy and the influence of camera angle on emotion recognition. The results showed a slight advantage for the VR condition. The current paper builds on previous research60 and introduces several novel aspects in the study of the communication and perception of emotions conveyed by VHs. These novel factors include the use of action units to reproduce emotions in VHs, the investigation of the effect of VH position (frontal or midprofile) on emotion recognition, and the investigation of the level of immersion in emotion recognition applications, comparing the experiences of 72 participants using both desktop computers and HMDs. To our knowledge, no previous studies have conducted a similar comparison between these two conditions for emotion recognition using HMDs, focusing specifically on the aforementioned factors. In contrast, previous work has mainly focused on recognizing emotions experienced by participants immersed in a VR environment,61 rather than investigating participants’ ability to recognize emotions presented by VHs. This paper is organized as follows. Section 2 outlines the methods used in the study. This includes details about the subjects involved, the tools and equipment used, the cues needed to represent each emotion, and the steps followed during the experiment. Section 3 presents the results of the experiment. In Section 4, the results are discussed and analyzed in the context of the research objectives. Finally, Section 5 provides a summary of the main conclusions. 2. Materials and Methods 2.1. Participants Two experiments were conducted to evaluate the use of a non-immersive (desktop) and an immersive virtual reality (VR) application. The participants in the study consisted of 72 healthy volunteers, aged 20 to 59 years, with no prior experience with the dynamic virtual faces used in the evaluation. 36 participants took part in the desktop experiment, while the remaining 36 participated in the VR experiment using an HMD. Individuals with a previous diagnosis of mental illness, personal medical history, or firstdegree family history of psychosis were excluded from the study. The study was conducted in accordance with the guidelines of the Declaration of Helsinki and November 4, 2025 12:11 ImmersionFAR 4M.A. Vicente-Querol et al. was approved by the Clinical Research Ethics Committee of the Complejo Hospitalario Universitario de Albacete (protocol code 2019/07/073, approved on September 24, 2019). Informed consent was obtained from all subjects included in the study. To ensure that the sample size provided sufficient statistical power, a sensitivity test was performed using G*Power software (version 3.1.9.7). Given the nonparametric tests used in the study, namely the Mann-Whitney U test and the Wilcoxon Signed Rank test, we calculated the sensitivity for t tests and determined the required effect size. With a significance level of α= 0.05, a power (1−β) of 0.80, and a sample size of 36 for both groups, we obtained a non-centrality parameter of δ= 2.842, a critical t value of 1.996, and an effect size of d= 0.686, indicating a medium effect size. 2.2. Experimental Setup The computer used in the study was a laptop with a 17.3” display, an Intel Core i7-9750H processor, 16GB of RAM, and an NVIDIA RTX2070 Super graphics card. Participants used a mouse to make selections in the desktop experiment and a gamepad controller to make selections in the immersive experiment. The desktop experiment was conducted on the laptop screen, while the immersive experiment used an HMD, specifically FOVE (https://fove-inc. com/), which was equipped with a WQHD OLED display of 2560 ×1440 pixels and a field of view of 100 degrees. 2.3. Stimuli As already mentioned, the DVFs used as stimuli for emotion recognition in both experiments have been previously validated in both healthy individuals55 and individuals diagnosed with schizophrenia.56 The DVFs include the six basic emotions (anger, disgust, fear, joy, sadness, and surprise) plus the neutral expression. The stimuli consist of two avatar races (2 Caucasian and 2 African) of approximately 30 years of age and two of old age (male and female for all cases), and two viewpoints for presenting the DVFs: frontal and mid-profile (45 degrees). Ekman and Friesen62 conducted a comprehensive study of the facial changes associated with the six basic emotions. As their work is widely accepted in the literature, the expressions of our DVFs were designed on the basis of this study and using the wellknown FACS system.28 The technological process and the workflow used for their creation have been previously described in detail,54 while the present article focuses on the design of emotions based on the research of Ekman and Friesen, who discovered that each emotion has its own unique cues. As a result of these studies, the emotions were recreated in the DVFs, along with the wrinkles. Figure 1 shows how the six basic emotions plus the neutral one are represented in the DVFs, and also how each emotion is broken down into the sum of several AUs. 2.4. Procedure The experimental procedure, as shown in Figure 2, consisted of several steps. Figure 2. Procedure followed in the tests. First, participants completed a pretest phase in which they filled out a sociodemographic form with information such as age and gender. They also completed the Spanish version of the Positive and Negative Affect Schedule (PANAS) form,63 a questionnaire that assesses positive and negative affect. Participants with scores below 25 (PA <25) or above 35 (NA >35) were excluded due to possible mood alter- November 4, 2025 12:11 ImmersionFAR Emotion Recognition using Dynamic Virtual Faces 5 Figure 1. The six basic emotions plus the neutral expression and their AUs. The first row shows neutral,disgust,joy, surprise and anger emotions. The second row shows fear and sadness emotions, the names of each AU and the breakdown of each emotion into AUs. ations. After filling out the forms, participants in the VR experiment had to go through a short training phase. This phase was designed to familiarize them with the gamepad controller and the environment. Five samples of DVFs displaying random emotions were shown during this training phase. The test phase consisted of 52 trials, including 8 trials for each basic emotion (4 for each intensity level) and 4 for the neutral expression. The pseudocode in Listing 1 defines the procedure for the test phase, in which the DFVs are presented to the participants. There was no time limit for them to choose an answer. 1Initialize variables 2number_of_trials <- 52 3number_of_emotions <- 7 4transition_time <- 0.5 seconds 5display_time <- 1.5 seconds 6 7for each trial in number_of_trials do 8Select a random emotion 9Display neutral expression 10 Transition smoothly to the selected 11 emotion over transition_time 12 Display the selected emotion for 13 display_time 14 Transition back to neutral expression 15 over transition_time 16 Wait for user response 17 Transition environment to light gray 18 background 19 end for The order in which the emotions were presented, the virtual characters that represented them, and the camera angles varied across participants. The order of presentation of the virtual characters was randomized, as was their angle (with half of the trials presented from a frontal view and half from a side view - left and right). The gender distribution of the virtual characters was 50% male and 50% female. The race (African and Caucasian) and age (young and old) of the avatars were also randomized, including slight variations in eye color, skin tone, and hair, while maintaining the specified distributions. The experimental procedure was identical for both the desktop and VR conditions, with the only difference being the method of interaction. In the desktop condition, participants selected the correct emotion by clicking on buttons displaying the corresponding emotion names on the screen. In the VR condition, participants used the gamepad sticks to select an emotion and confirmed their choice by pressing a button. 2.5. Data Analysis The data collected for the emotion recognition experiment, including information about the emotion presented, type of presentation (frontal or mid-profile), identification accuracy, user response, and response time, were stored in the rows of a .CSV file and later analyzed using Microsoft Excel and IBM SPSS v28. Descriptive statistics, including mean and percentages, were obtained. Because the distribution of the data was not normal, nonparametric methods were used to compare the results between the two ex- November 4, 2025 12:11 ImmersionFAR 6M.A. Vicente-Querol et al. perimental conditions. The Mann-Whitney test was used to compare results between the desktop and VR conditions, while the Wilcoxon Signed Rank test was used to compare results within each condition (frontal vs. mid-profile). Spearman’s rank correlation coefficients were used for the identification of associations between variables. A p-value less than 0.05 was considered statistically significant. 3. Results The results of the study will be presented as follows. Section 3.1 will provide a general overview of the results of the two experiments in terms of accuracy of recognition, viewing time, and level of intensity of each emotion, while Section 3.2 will focus on presenting the results for the front and mid-profile views in each condition in terms of accuracy of recognition and viewing time. 3.1. Accuracy of Recognition Figure 3 shows the confusion matrices obtained from both the desktop and immersive VR experiments. The main diagonal shows the percentage of success (hits) in identifying each emotion, and the rows show the percentage of failure (misses). The color background of each cell represents the magnitude of the value: dark black represents the lowest values, dark red represents the highest values, while lighter (black, orange, red) colors represent values in between. For example, the second row represents “surprise” as the correct answer, and it is crossed by six columns, each representing the other answers given by the participants. Thus, the value shown in the second row, second column represents how often the participants chose the correct answer shown in dark red color (93.06% in desktop and 94.72% in VR), while the rest of the values in the second row represent how often the participants did not choose the correct answer shown in dark black color (they mainly confused surprise with fear in both experiments, 6.25% in desktop and 4.44% in VR). The average hit rate is 90.03% for the desktop experiment, with fear (77.78%), disgust (85.42%), and sadness (84.72%) below that average. For the VR experiment, the average is 92.34%, again with fear (76.67%), disgust (91.11%), and sadness (88.33%) below that average. Fear is mainly confused with surprise in both cases, but more so in VR (14.93% vs. 21.39%). Disgust is also confused with anger in both cases, but the confusion is slightly higher on desktop than in VR (11.81% vs 8.06%). Despite the low percentages (no more than 5%), sadness is most often confused with other emotions, although the percentages are low (no more than 5%). Figure 3. Confusion matrices showing the percentages of the correct answers for both the desktop and VR experiments. Fear is divided into its two intensity levels, Fear1 and Fear2. Although there are apparent differences in the percentages mentioned in the previous paragraph, they do not show statistical significance regarding the average hit rate of the two experiments (MannWhitney U= 530.0, p = 0.183). However, there is a significant difference for anger (U= 496.5, p = 0.013), with a higher hit rate in the VR experiment. The average response times for correctly and incorrectly recognized emotions per trial and per emotion were recorded and presented in Table 1. The November 4, 2025 12:11 ImmersionFAR Emotion Recognition using Dynamic Virtual Faces 7 results show that the time to recognize emotions was longer for misses compared to hits, and the average response time was higher in the VR condition for each emotion. A correlation between misses and average reaction time was observed in the desktop environment (r= 0.444, p= 0.007), but no such correlation was found in the VR environment (r= 0.116, p= 0.502). Table 1. Desktop and VR average viewing times for hits and misses views in seconds per experiment and emotion. Hits Misses Emotion Desktop VR Desktop VR Neutral 2.62 3.76 2.97 14.84 Surprise 1.78 4.86 2.90 8.95 Fear 2.18 5.57 2.62 5.33 Anger 2.04 4.58 4.06 8.14 Disgust 1.96 4.44 2.85 7.92 Joy 1.74 4.52 2.72 4.97 Sadness 2.17 4.62 3.52 5.97 Each emotion was also presented at two levels of expressive intensity. The results indicate that recognition accuracy increased with increasing intensity levels in both experiments, or remained largely unchanged. However, significant variations were only observed for the emotion fear. As the intensity level increased from fear1 to fear2 (see Figure 3), the confusion with the emotion surprise increased in both desktop and VR conditions (from 9.72% to 20.14% in desktop and from 10% to 32.78% in VR). However, the accuracy of recognizing fear decreased in immersive VR (from 86.11% to 67.22%) and increased slightly in the desktop condition (from 77.08% to 78.47%). In VR, the hit rate for fear1 is significantly higher than that for fear2 (Wilcoxon Z=−3.487, p<0.001), while this difference is not observed in the desktop condition. Moreover, fear2 is significantly more confused with surprise than fear1 in both conditions (Z=−2.027, p= 0.0413 and Z=−3.841, p<0.001, for VR and desktop, respectively). Finally, the confusion between fear2 and surprise is significantly higher in VR than in the desktop condition (U= 479.5, p= 0.049), while no difference is observed for fear1 (U= 630.5, p= 0.791). In addition, the accuracy of detecting sadness decreased with increasing intensity in the desktop condition (from 88.89% to 80.56%) and not so much in the VR condition (from 88.9% to 87.8%), although both are statistically significant (Z=−1.897, p= 0.058 and Z=−0.159, p= 0.874, respectively). 3.2. Accuracy by Orientation of the View Figure 4 shows the results of the desktop experiment, but separated according to how the avatar is presented, front or mid-profile. Figure 4. Confusion matrix showing the percentages of the correct answers for the desktop experiment for front and mid-profile. The Wilcoxon Signed Rank Test revealed no statistically significant differences between the two viewing angles for the desktop condition (Z= −0.246, p = 0.806), while the difference was found to be significant for the VR condition (Z=−2.819, p = 0.005). Furthermore, the differences between the two conditions were examined separately for the front and side views. In the case of the frontal view, statistically significant differences were observed (MannWhitney U= 444.5, p = 0.021), with higher percentages obtained for VR. However, no differences were found for the side view (U= 607.5, p = 0.647). November 4, 2025 12:11 ImmersionFAR 8M.A. Vicente-Querol et al. Examining the individual emotions, the results are similar in both cases the main differences being that in fear the confusion with surprise increases in the mid-profile view (11.64% versus 18.31%), although the hit rate remains similar (79.45% versus 76.06%). This difference in confusion rates is statistically significant (Z=−2.070, p = 0.038). The hit rate for sadness decreases from 86.81% to 82.64%, mainly due to increased confusion with neutral and surprise (from 3.47% to over 6%). Conversely, for disgust, the hit rate increases (82.64% versus 88.49%) mainly due to a decrease in confusion with anger (from 15.97% to 7.64%). However, the difference in confusion rates is not statistically significant (Z=−1.812, p = 0.070). In Figure 5, the results of the VR experiment are split between the front and mid-profile views. The main differences between the two views are similar to those in the desktop experiment. Figure 5. Confusion matrix showing the percentages of the correct answers for the Virtual reality experiment for front and mid-profile. Fear is more likely to be confused with surprise in the mid-profile view than in the front view, resulting in an error rate of 29.51%, which significantly decreases the hit rate (84.18% versus 69.40%). Similar to the desktop condition, this difference is statistically significant (Z=−3.683, p < 0.001). The hit rate for sadness decreases from 90.91% in the front view to 85.87% in the mid-profile view. Contrary to the desktop experiment, the hit rate for disgust also decreases, from 95.51% to 86.81%, while the confusion with anger increases significantly from 3.37% to 12.64% (Z=−2.993, p = 0.003). In the mid-profile view, the hit rate slightly increases for surprise (from 94.15% to 95.24%) and joy (97.73% to 100%). The hit rate for anger remains approximately the same in both views, at around 99%. Table 2 shows the average time spent exploring each emotion based on the angle at which the DVF was presented, including both hits and misses. The results suggest that the time spent exploring each emotion was relatively similar, with minimal differences when the emotion was presented frontally or in mid-profile, for both experiments and for each emotion. Table 2. Average viewing time of emotions when correctly identified from the front and mid-profile views in seconds. Front view Mid-profile view Emotion Desktop VR Desktop VR Neutral 2.70 3.84 2.53 3.69 Surprise 1.60 4.83 1.95 4.90 Fear 2.01 5.45 2.36 5.71 Anger 2.10 4.46 1.99 4.70 Disgust 2.01 4.51 1.91 4.36 Joy 1.72 4.46 1.77 4.58 Sadness 2.18 4.66 2.15 4.57 4. Discussion The results showed some differences in the recognition rates per emotion and, more importantly, in the confusion rates. We suspect that this may be related to how the emotions are decomposed into AUs. Therefore, this Section first discusses how AUs can affect emotion recognition. Second, the differences in the results with respect to the presentation angles are discussed. Next, the results of the two experimental conditions, immersive and desktop VR, are compared. Finally, we discuss the brain processes that might have led to better recognition rates when using our DVFs compared to static stimuli. November 4, 2025 12:11 ImmersionFAR Emotion Recognition using Dynamic Virtual Faces 9 4.1. Effects of the Action Units One of the goals of this study was to see how AUs were developed in the set of DVFs, how they might affect emotion recognition, and if the errors produced when trying to decode them were similar to previous work using pictures of real people. Some prototypes of AUs have been identified to represent each emotion.28, 38, 48, 51, 64 These prototypes contain more or less AUs, but there are some that appear more often: 6+12 for joy, 1+4+15 in sadness, 1+2+5 in surprise, 9+15 in disgust, 4+5 in anger and 1+4+5 in fear. A recent paper35 analyzed the AUs used by actors to represent the six basic emotions plus neutral, excluding surprise. For fear, the authors found that the most commonly used AUs were AU1, 4, 5, and 25, which appeared in more than 80% of the models and were primarily located around the eyes. For joy, AU6 was present in over 75% of the models, while AU12 was present in all models. For anger, only AU4 was found in more than 90% of the models. For disgust, AU9 was present in 75% of the models, and for sadness, AU1 and AU4 were present in over 90% of the models. According to previous research,34 these AUs were considered characteristic of each emotion, although not all AUs were positively correlated with recognition. Specifically, AU6 was found to be positively associated with joy, while AU4 and 17 were positively correlated with sadness, AU4 and 5 were positively associated with anger, while AU1, 5, and 26 were associated with fear. The remaining AUs are either not associated or negatively associated. They did not include surprise and disgust in their study. The present study suggests that the AUs commonly portrayed by actors are also present in the set of DVFs, as shown in Figure 1, thus enhancing the realism of the expressed emotions. Looking at the effects of AUs on the confusion between emotions, as previously shown in Figure 3, fear is the worst recognizable emotion, and it is mainly confused with surprise. The confusion between these two emotions is higher when the emotion to be decoded is fear in both desktop and VR conditions (14.93% and 21.39%). Fear and surprise share four AUs and the difference is that fear has two more AUs, AU4 (Brow Lowerer) and AU20 (Lip Stretcher), the first located around the eyes and the second in the mouth area. According to a very recent paper,65 in fear the eyes are the most viewed area of interest (AOI) among other AOIs such as the mouth or the nose. This was also found in other experiments using photographs of real people.29, 66 However, surprise the most viewed AOIs were both the eyes and the mouth. It can be assumed that the AUs that contribute most to the recognition of fear are those in the eyes, i.e. AU1, AU2, AU4 and AU5. This does not mean that the AUs involving the mouth are not seen (AU20 and AU26), but that they seem to be less relevant for decoding fear than for surprise. Since these two emotions share three AUs in the eyes, one reason for the confusion between them could be the mislabeling of these AUs, resulting in a failure to decode these emotions. In addition, according to Roy-Charland et al.32 and Chamberland et al.,33 the confusion between these two emotions could be a lack of attention to the forehead region when the forehead shows the discriminative cues, supporting the perceptualattentional limitation hypothesis. This means that it is more difficult to distinguish emotions perceptually when the similarity between their facial expressions is greater. They also found that the qualitative value of these cues is also important. Ekman and Friesen62 made a deep analysis of the six basic emotions. In the analysis of fear they said that it can blend with other emotions, but the most common emotion to blend is surprise and that after surprise usually another emotion is shown, which depends on the type of stimulus the person is exposed to. They also defined startle, which is considered an extreme form of surprise and is closely related to fear. In our DVFs, two levels of intensity were created for each emotion, and participants confused the most intense form of fear emotion with surprise. As presented in Section 3.1, fear2 is more confused with surprise than fear1, with the confusion being higher in the VR condition for fear2 (20.1% vs. 32.8%). Figure 6 shows a comparison of fear and surprise in terms of intensity. The changes in the most extreme expressions are: the mouth and eyes are more separated, the eyebrows are higher, and in fear the lips and the whole face look more tense. Looking at the time spent by the participants in recognizing emotions, as previously shown in Table 1, it becomes apparent that when they were unable to decode an emotion, the time invested is greater, November 4, 2025 12:11 ImmersionFAR 16 M.A. Vicente-Querol et al. tion, International Journal of Neural Systems 32(03) (2022) p. 2250005. 25. J. De Lope and M. Gra˜na, A hybrid time-distributed deep neural architecture for speech emotion recognition, International Journal of Neural Systems 32(06) (2022) p. 2250024. 26. P. Hajek, A. Barushka and M. Munk, Neural networks with emotion associations, topic modeling and supervised term weighting for sentiment analysis, International journal of neural systems 31(10) (2021) p. 2150013. 27. P. Ekman and D. Cordaro, What is meant by calling emotions basic, Emotion Review 3(4) (2011) 364– 370. 28. P. Ekman and W. Friesen, Facial Action Coding System: A Technique for the Measurement of Facial Movement (Consulting Psychologists Press, Palo Alto, CA, 1978). 29. K. Guo, Holistic gaze strategy to categorize facial expression of varying intensities, PLoS ONE 7(8) (2012). 30. K. Guo and H. Shaw, Face in profile view reduces perceived facial expression intensity: An eye-tracking study, Acta Psychologica 155 (2015) 19–28. 31. P. Surcinelli, F. Andrei, O. Montebarocci and S. Grandi, Emotion recognition of facial expressions presented in profile, Psychological Reports 125 (2022) 2623–2635. 32. A. Roy-Charland, M. Perrona, O. Beaudrya and K. Eadya, Confusion of fear and surprise: A test of the perceptual-attentional limitation hypothesis with eye movement monitoring, Cognition and Emotion 28 (2014) 1214–1222. 33. J. Chamberland, A. Roy-Charland, M. Perron and J. Dickinson, Distinction between fear and surprise: an interpretation-independent test of the perceptual-attentional limitation hypothesis, Social Neuroscience 12 (2017) 751–768. 34. C. G. Kohler, T. Turner, N. M. Stolar, W. B. Bilker, C. M. Brensinger, R. E. Gur and R. C. Gur, Differences in facial expressions of four universal emotions, Psychiatry Research 128 (2004) 235–244. 35. F. Poncet, R. Soussignan, M. Jaffiol, B. Gaudelus, A. Leleu, C. Demily, N. Franck and J. Y. Baudouin, The spatial distribution of eye movements predicts the (false) recognition of emotional facial expressions, PLoS ONE 16(1 January) (2021) 1–24. 36. Y. Busin, K. Lukasova, M. K. Asthana and E. C. Macedo, Hemiface differences in visual exploration patterns when judging the authenticity of facial expressions, Frontiers in Psychology 8(2018) p. 2332. 37. A. Garc´ıa, P. Fern´andez-Sotos, A. Fern´andezCaballero, E. Navarro, J. Latorre, R. RodriguezJimenez and P. Gonz´alez, Acceptance and use of a multi-modal avatar-based tool for remediation of social cognition deficits, Journal of Ambient Intelligence and Humanized Computing 1(2020) p. 4513–4524. 38. J. Guti´errez-Maldonado, M. Rus-Calafell and J. Gonz´alez-Conde, Creation of a new set of dynamic virtual reality faces for the assessment and training of facial emotion recognition ability, Virtual Reality 18(1) (2014) 61–71. 39. C. G. Kohler, T. H. Turner, W. B. Bilker, C. M. Brensinger, S. J. Siegel, S. J. Kanes, R. E. Gur and R. C. Gur, Facial emotion recognition in schizophrenia: intensity effects and error pattern, American Journal of Psychiatry 160(10) (2003) 1768–1774. 40. G. C. Burdea and P. Coiffet, Virtual Reality Technology (John Wiley & Sons, 2003). 41. D. Roberts, R. Wolff, O. Otto and A. Steed, Constructing a gazebo: supporting teamwork in a tightly coupled, distributed task in virtual reality, Presence 12(6) (2003) 644–657. 42. I. Heldal, A. Steed and R. Schroeder, Evaluating collaboration in distributed virtual environments for a puzzle-solving task, HCI International, (Las Vegas, NV, USA, 2005). 43. A. S. Garc´ıa, D. Mart´ınez, J. P. Molina and P. Gonz´alez, Collaborative virtual environments: you can’t do it alone, can you?, International Conference on Virtual Reality, Springer, (Beijing, China, 2007), pp. 224–233. 44. S. Baceviciute, T. Terkildsen and G. Makransky, Remediating learning from non-immersive to immersive media: Using eeg to investigate the effects of environmental embeddedness on reading in virtual reality, Computers & Education 164 (2021) p. 104122. 45. M. Slater, V. Linakis, M. Usoh and R. Kooper, Immersion, presence and performance in virtual environments: An experiment with tri-dimensional chess, ACM Symposium on Virtual Reality Software and Technology, (ACM, Hong Kong, 1996), pp. 163–172. 46. J. A. Stevens, J. P. Kincaid et al., The relationship between presence and performance in virtual simulation training, Open Journal of Modelling and Simulation 3(02) (2015) p. 41. 47. A. S. Garc´ıa, D. J. Roberts, T. Fernando, C. Bar, R. Wolff, J. Dodiya, W. Engelke and A. Gerndt, A collaborative workspace architecture for strengthening collaboration among space scientists, 2015 IEEE Aerospace Conference, IEEE, (Big Sky, MT, USA, 2015), pp. 1–12. 48. M. Dyck, M. Winbeck, S. Leiberg, Y. Chen, R. C. Gur and K. Mathiak, Recognition profile of emotions in natural and virtual faces, PloS one 3(11) (2008) p. e3628. 49. E. G. Krumhuber, L. Tamarit, E. B. Roesch and K. R. Scherer, Facsgen 2.0 animation software: generating three-dimensional facs-valid facial expressions for emotion research., Emotion 12(2) (2012) p. 351. 50. C. C. Joyal, L. Jacob, M. H. Cigna, J. P. Guay and P. Renaud, Virtual faces expressing emotions: An initial concomitant and construct validity study, Frontiers in Human Neuroscience 8(SEP) (2014) 1– November 4, 2025 12:11 ImmersionFAR Emotion Recognition using Dynamic Virtual Faces 17 6. 51. R. Amini, C. Lisetti and G. Ruiz, Hapfacs 3.0: Facsbased facial expression generator for 3d speaking virtual characters, IEEE Transactions on Affective Computing 6(4) (2015) 348–360. 52. C. Geraets, S. K. Tuente, B. Lestestuiver, M. Van Beilen, S. Nijman, J. Marsman and W. Veling, Virtual reality facial emotion recognition in social environments: An eye-tracking study, Internet interventions 25 (2021) p. 100432. 53. C. Faita, F. Vanni, C. Tanca, E. Ruffaldi, M. Carrozzino and M. Bergamasco, Investigating the process of emotion recognition in immersive and nonimmersive virtual technological setups, Proceedings of the 22nd ACM Conference on Virtual Reality Software and Technology, (Christchurch , New Zealand, 2016), pp. 61–64. 54. A. S. Garc´ıa, P. Fern´andez-Sotos, M. A. VicenteQuerol, G. Lahera, R. Rodriguez-Jimenez and A. Fernandez-Caballero, Design of reliable virtual human facial expressions and validation by healthy people, Integrated Computer-Aided Engineering 27(3) (2020) 287–299. 55. P. Fern´andez-Sotos, A. S. Garc´ıa, M. A. VicenteQuerol, G. Lahera, R. Rodriguez-Jimenez and A. Fern´andez-Caballero, Validation of dynamic virtual faces for facial affect recognition, PLoS ONE 16(1 1) (2021) 1–15. 56. N. I. Muros, A. S. Garc´ıa, C. Forner, P. L´opez-Arcas, G. Lahera, R. Rodriguez-Jimenez, K. N. Nieto, J. M. Latorre, A. Fern´andez-Caballero and P. Fern´andezSotos, Facial affect recognition by patients with schizophrenia using human avatars, Journal of Clinical Medicine 10(9) (2021) p. 1904. 57. M. Monferrer, A. S. Garc´ıa, J. J. Ricarte, M. J. Montes, A. Fern´andez-Caballero and P. Fern´andezSotos, Facial emotion recognition in patients with depression compared to healthy controls when using human avatars, Scientific Reports 13(1) (2023) p. 6007. 58. M. Monferrer, A. S. Garc´ıa, J. J. Ricarte, M. J. Montes, P. Fern´andez-Sotos and A. Fern´andezCaballero, Facial affect recognition in depression using human avatars, Applied Sciences 13(3) (2023) p. 1609. 59. J. del ´ Aguila, L. M. Gonz´alez-Gualda, M. A. J´ativa, P. Fern´andez-Sotos, A. Fern´andez-Caballero and A. S. Garc´ıa, How interpersonal distance between avatar and human influences facial affect recognition in immersive virtual reality, Frontiers in Psychology 12 (2021) p. 675515. 60. M. A. Vicente-Querol, A. Fern´andez-Caballero, J. P. Molina, P. Gonz´alez, L. M. Gonz´alez-Gualda, P. Fern´andez-Sotos and A. S. Garc´ıa, Influence of the level of immersion in emotion recognition using virtual humans, International Work-Conference on the Interplay Between Natural and Artificial Computation, Springer, (Puerto de la Cruz, Tenerife, Spain, 2022), pp. 464–474. 61. J. Mar´ın-Morales, C. Llinares, J. Guixeres and M. Alca˜niz, Emotion recognition in immersive virtual reality: From statistics to affective computing, Sensors 20(18) (2020) p. 5163. 62. P. Ekman and W. Friesen, Unmasking the face 2003. 63. B. Sand´ın, P. Chorot, L. Lostao, T. E. Joiner, M. A. Santed and R. M. Valiente, Escalas PANAS de afecto positivo y negativo: Validacion factorial y convergencia transcultural, Psicothema 11(1) (1999) 37–51. 64. N. Hussain, H. Ujir, I. Hipiny and J.-L. Minoi, 3d facial action units recognition for emotional expression https://arxiv.org/abs/1712.00195, (2017), Preprint. 65. M. A. Vicente-Querol, A. Fern´andez-Caballero, J. P. Molina, L. M. Gonz´alez-Gualda, P. Fern´andez-Sotos and A. S. Garc´ıa, Facial affect recognition in immersive virtual reality: Where is the participant looking?, International Journal of Neural Systems 32(10) (2022) p. 2250029. 66. M. W. Schurgin, J. Nelson, S. Iida, H. Ohira, J. Y. Chiao and S. L. Franconeri, Eye movements during emotion recognition in faces, Journal of Vision 14(13) (2014). 67. M. G. Calvo, A. Guti´errez-Garc´ıa and M. D. L´ıbano, What makes a smiling face look happy? visual saliency, distinctiveness, and affect, Psychological Research 82 (2018) 296–309. 68. E. Krumhuber and A. Kappas, Moving smiles: The role of dynamic components for the perception of the genuineness of smiles, Journal of Nonverbal Behavior 29 (2005) 3–24. 69. H. Hill, P. G. Schyns and S. Akamatsu, Information and viewpoint dependence in face recognition, Cognition 62 (1997) 201–222. 70. S. Du and A. M. Martinez, The resolution of facial expressions of emotion, Journal of Vision 11 (2011) 1–13. 71. C. E. Waugh, E. Z. Shing and B. M. Avery, Temporal Dynamics of Emotional Processing in the Brain, Emotion Review 7(4) (2015) 323–329. 72. M. Arsalidou, D. Morris and M. J. Taylor, Converging evidence for the advantage of dynamic facial expressions, Brain Topography 24(2) (2011) 149–163. 73. S. A. Trautmann, T. Fehr and M. Herrmann, Emotions in motion: Dynamic compared to static facial expressions of disgust and happiness reveal more widespread emotion-specific activations, Brain Research 1284 (2009) 100–115. 74. N. Torro-Alves, I. A. d. O. Bezerra, R. G. e Claudino, M. R. Rodrigues, J. P. Machado-de Sousa, F. d. L. Os´orio and J. A. Crippa, Facial emotion recognition in social anxiety: The influence of dynamic information, Psychology and Neuroscience 9(1) (2016) 1–11. 75. E. G. Krumhuber, A. Kappas and A. S. Manstead, Effects of dynamic aspects of facial expressions: A review, Emotion Review 5(1) (2013) 41–46. 76. S. A. Grainger, J. D. Henry, L. H. Phillips, E. J. November 4, 2025 12:11 ImmersionFAR 18 M.A. Vicente-Querol et al. Vanman and R. Allen, Age Deficits in Facial Affect Recognition: The Influence of Dynamic Cues, Journals of Gerontology - Series B Psychological Sciences and Social Sciences 72(4) (2017) 622–632. 77. H. Hoffmann, H. C. Traue, K. Limbrecht-Ecklundt, S. Walter and H. Kessler, Static and Dynamic Presentation of Emotions in Different Facial Areas: Fear and Surprise Show Influences of Temporal and Spatial Properties, Psychology 04(08) (2013) 663–668. 78. E. Bould, N. Morris and B. Wink, Recognising subtle emotional expressions: The role of facial movements, Cognition and Emotion 22(8) (2008) 1569–1587. 79. Z. Ambadar, J. W. Schooler and J. F. Conn, Deciphering the enigmatic face the importance of facial dynamics in interpreting subtle facial expressions, Psychological Science 16(5) (2005) 403–410. 80. A. Leleu, M. Dzhelyova, B. Rossion, R. Brochard, K. Durand, B. Schaal and J. Y. Baudouin, Tuning functions for automatic detection of brief changes of facial expression in the human brain, NeuroImage 179 (2018) 235–251. 81. C. Biele and A. Grabowska, Sex differences in perception of emotion intensity in dynamic and static facial expressions, Experimental Brain Research 171(1) (2006) 1–6. 82. S. Uono, W. Sato and M. Toichi, Brief report: Representational momentum for dynamic facial expressions in pervasive developmental disorder, Journal of Autism and Developmental Disorders 40(3) (2010) 371–377. 83. W. Sato, T. Kochiyama, S. Yoshikawa, E. Naito and M. Matsumura, Enhanced neural activity in response to dynamic facial expressions of emotion: an fmri study, Cognitive Brain Research 20(1) (2004) 81–91.