Facial Emotion Perception Similarity and Gender Bias: Comparing Humans and AI in Detection and Rating Performance

Jini Tae
Gwangju Institute of Science and Technology, School of Humanities and Social Sciences, jini0930@gist.ac.kr

Ju-Hyeon Park
Gwangju Institute of Science and Technology, Department of Electrical Engineering and Computer Science, juhyeon-park@gm.gist.ac.kr

Wonil Choi
Gwangju Institute of Science and Technology, School of Humanities and Social Sciences, wichoi@gist.ac.kr


Abstract. This study compares human and AI emotion recognition patterns—specifically valence and arousal—using 1,440 generative facial images. 1,000 Korean adults rated stimuli balanced across diverse races and genders, providing benchmarks for two AI models: MobileViT and enet_b0_8_va_mtl. Quantitative analyses revealed high human-AI correspondence in valence ratings. However, arousal patterns notably diverged by stimulus gender: whereas humans perceived higher arousal in positive female faces, AI models showed no gender difference for positive emotions but rated arousal higher for male faces in neutral and negative categories (e.g., anger, sadness, fear). These findings pinpoint critical misalignments in AI emotional perception, particularly regarding how massive training datasets may distort the nuanced intensity of human affect. This research offers vital insights into the socio-technical implications of gender bias in AI-driven affect recognition systems.

CCS CONCEPTS • Human-centered computing • Human computer interaction (HCI) • Empirical studies in HCI

Additional Keywords and Phrases: Emotion Recognition, Generative AI, Human-AI Comparison, Gender Bias, Valence-Arousal Model


1 INTRODUCTION

1.1 Motivation: Affective Agents in UX/UI

The integration of affective computing into UX/UI design has shifted the paradigm of human-computer interaction from functional utility to affective rapport [Conati et al., 2005; Pantic et al., 2005]. Affective agents, ranging from mental health chatbots to responsive virtual assistants, aim to enhance user experience by establishing a deep emotional connection with users.

However, the efficacy of the affective agents hinges on affective alignment: the degree to which an AI model’s interpretation of emotional cues matches human psychological standards. While AI models have achieved high accuracy in categorical emotion detection [Khare et al., 2024; Mollahosseini et al., 2017; Telceken et al., 2025], their ability to mirror the nuanced empathetic accuracy of human valence and arousal ratings remains under-explored [Mollahosseini et al., 2017; Nomiya et al., 2025; Toisoul et al., 2021]. If an agent misinterprets the intensity of a user’s distress, it risks breaking user trust and failing to sustain a meaningful interaction.

1.2 The Needs for Dimensional Emotion Analysis

The representation of affection in computational systems generally follows two primary approaches: categorical and dimensional. The categorical approach posits that a small set of basic emotions (e.g., happiness, anger, fear) is universally recognized and hard-wired in the human brain. This paradigm has been the most widely adopted in research on automated affect detection measurement due to its intuitive nature and ease of labeling.

However, psychologists have long utilized dimensional measures to provide a more nuanced framework for understanding emotional states. The most prominent of these is the Circumplex Model of Affect [Russell, 1980], which maps all emotions onto a continuous two-dimensional space defined by Valence (the degree of positivity or negativity) and Arousal (the level of physiological or psychological activation) [Yik et al., 2023].

Unlike categorical approach, dimensional ratings allow us to capture the psychological gist of emotional intensity and detect subtle perceptual misalignments between humans and AI models—variations that categorical classification might otherwise obscure. To this end, we utilize the Valence-Arousal dimensions to empirically test the pattern of Human-AI correspondence, specifically focusing on how these ratings shift across different face genders and emotion categories.

1.3 Gender Biases in Human Emotion Perception

Human emotion perception is not merely an objective decoding of facial muscle movements; it is a socio-cognitive process deeply influenced by stereotypes and expectations. Research indicates that gender acts as a primary heuristic in affective evaluation. For instance, gender-emotion stereotypes often lead observers to associate male faces more strongly with dominance-related emotions like anger, while female faces are linked to prosocial or submissive emotions such as happiness and sadness [Hess et al., 2004]. These biases significantly impact dimensional ratings; studies have shown that the same facial expression can be perceived as more intense (higher arousal) or more negative (lower valence) depending on the perceived gender of the face [Adolph & Alpers, 2010; Baudouin et al., 2025; Plant et al., 2000].

Since AI models are trained on large-scale datasets reflecting human-labeled stereotypes, they risk not only inheriting these gender biases but also amplifying them through algorithmic optimization. Investigating whether AI replicates human-like biases or introduces unique, machine-generated discrepancies—using controlled facial stimuli with robust human benchmarks—is critical for assessing the perceptual reliability of affect recognition systems and their correspondence with human psychological standards.

1.4 Present Work

Despite the rapid deployment of AI in affect recognition, the extent to which these models correspond with human psychological benchmarks remains unclear. This study addresses this gap by comparing human emotional ratings (N=1,000) with two distinct AI architectures—MobileViT [Mehta & Rastegari, 2021] and enet_b0_8_va_mtl [Savchenko, 2022]—using 1,440 generative AI stimuli (GIST-AIFaceDB, under review).

To provide a comprehensive comparative analysis, we focus on the following objectives:

  • Quantify Rank-Order Consistency: Assess the degree of similarity between human and AI ratings for valence and arousal across six categories of emotional faces.
  • Evaluate Gender Influences: Investigate how model architecture and stimulus gender interact to influence human-AI correspondence.
  • Identify Patterns of Divergence: Characterize specific emotional categories or conditions where AI models significantly deviate from human benchmarks, highlighting the limits of current affective architectures.

By benchmarking AI against a large-scale human data, this research identifies critical misalignment in affective computing and provides empirical grounds for building more human-centered and fair emotion recognition systems.


2 METHODOLOGY

2.1 Stimuli

Figure 1: Representative examples of the six emotional facial expressions (neutral, happy, angry, sad, and fear) generated for each racial group (Black, White, and Asian[Korean]). Each row depicts a different racial group, and each column corresponds to one emotion category.

To ensure high ecological validity and rigorous experimental control, we utilized the GIST AI-Generated Face Database (GIST-AIFaceDB). This database consists of 1,440 high-resolution, photorealistic facial images. The generation process involved a two-step pipeline: first, neutral base images were generated using the STOIQO NewReality Flux model to establish diverse virtual identities; subsequently, five distinct emotional expressions for each identity were produced using Nano-Banana (Gemini 2.5 Flash Image), an advanced image-generation and editing model implemented in Google AI Studio. The stimuli represent 240 unique virtual identities, strictly balanced across gender (male and female) and three racial groups (Black, White, and Asian [Korean]). Each identity is depicted expressing six basic emotion categories: happiness, anger, disgust, fear, sadness, and a neutral expression.

2.2 Human Participants & Procedures

The study protocol was reviewed and granted an exemption by the Institutional Review Board (IRB). A total of 1,000 native Korean adults (500 females, 500 males), aged 20–69 years, were recruited for the study. To ensure a representative sample of the general population, recruitment was strictly balanced across age cohorts and genders. The experiment was administered online, and participants accessed the task via their personal computers (laptops or desktops). Each participant evaluated 72 images randomly selected from the total pool of 1,440 stimuli and every image was presented in a randomized order. Through this counterbalanced design, each of the 1,440 images received 50 independent ratings.

The procedure consisted of two primary affective rating tasks: Valence and Arousal. In the Valence task, participants were instructed to evaluate the emotional positivity or negativity of each facial expression. In the Arousal task, they rated the level of emotional activation or intensity perceived in the image. Both ratings were recorded on a 9-point Likert scale, ranging from 1 (“extremely negative” / “not at all aroused”) to 9 (“extremely positive” / “highly aroused”).

2.3 AI Models

To compare with human ratings, we selected two AI models with distinct computational architectures: MobileViT [Mehta & Rastegari, 2021] and enet_b0_8_va_mtl [Savchenko, 2022]. MobileViT is a relatively lightweight model that combines CNNs and self-attention–based Transformer techniques, and it is known to perform well on computer vision tasks such as face recognition. enet_b0_8_va_mtl is also a lightweight model that incorporates modified CNNs with the squeeze and excitation optimization, originally developed within the HSEmotion framework and trained on AffectNet, which has been reported to perform well in recognizing emotions expressed in faces. These two models have been widely employed in various algorithms for recognizing emotions and estimating valence and arousal from facial expressions (Elsheikh et al., 2024; Wang et al., 2024).


3 RESULTS

Given the exploratory nature of these analyses, we report uncorrected significance levels (α = .05). For our key finding—the negative arousal correlation in sadness—the result remained significant after Bonferroni correction for 24 comparisons (adjusted α = .002).

3.1 Valence Rating

Figure 2. Human-AI correspondence in valence ratings across emotion categories and stimulus gender. (A) Spearman’s rank correlation coefficients (ρ) between aggregated human ratings and two AI models (MobileViT and enet_b0_8_va_mtl) by emotion. (B) Correlation patterns stratified by stimulus gender. Abbreviations: n.s. denotes non-significant correlation.

3.1.1 Rating Similarity

To evaluate the rank-order consistency of perceived emotional valence between human observers and AI models, we calculated Spearman’s rank correlation coefficients (ρ) for each of the six emotion categories across the 1,440 facial stimuli. It should be noted that Spearman’s ρ measures rank-order consistency rather than absolute agreement; our analysis thus captures whether human and AI ratings follow similar ordinal patterns, not whether their values are identical. For each emotion category (N=240), correlation scores were independently calculated for both MobileViT and enet_b0_8_va_mtl against human raters. Analysis revealed that both models exhibited significant positive correlations with human raters for most emotion categories (p < .05), indicating a consensus on the perception of emotional valence.

However, a notable exception was observed in the Happy category. For both AI architectures, valence ratings for happy expressions failed to show a significant correlation with human ratings. This may reflect a ceiling effect: both human raters and AI models consistently assigned high valence scores to happy faces, thereby restricting score variability and attenuating the correlation coefficient.

3.1.2 Gender-Stratified Similarity

To further investigate whether the correspondence between humans and AI models is influenced by the gender of the target face, we conducted a gender-stratified correlation analysis for each emotion category. Spearman’s rank correlation (ρ) was calculated separately for male and female stimuli (N=120 per gender/emotion combination).

For MobileViT, while the ‘Happy’ category showed non-significant correlations for both genders—consistent with the overall analysis—all other emotion categories exhibited significant positive correlations between the AI model and human raters, regardless of the face gender.

In the case of enet_b0_8_va_mtl, a similar pattern was observed for ‘Happy’ faces, where correlations were non-significant for both genders. However, a distinctive divergence appeared in the ‘Fear’ category: while valence ratings for male fearful faces showed a significant correlation with human raters, those for female fearful faces failed to reach statistical significance. For all other emotion categories, the model maintained a significant correlation regardless of the gender of the face.

3.2 Arousal Rating

Figure 3. Human-AI correspondence in arousal ratings across emotion categories and stimulus gender. (A) Spearman’s rank correlation coefficients (ρ) between aggregated human ratings and two AI models (MobileViT and enet_b0_8_va_mtl) by emotion. (B) Correlation patterns stratified by stimulus gender. Abbreviations: n.s. denotes non-significant correlation.

3.2.1 Rating Similarity

Compared to the valence perception patterns, the correspondence in arousal ratings for both models exhibited distinct and more complex characteristics. Spearman’s rank correlation analysis revealed that human and AI ratings were not significantly related (p > .05) for neutral and disgust expressions in MobileViT, while enet_b0_8_va_mtl showed non-significant correlations for neutral and fear categories.

For most other emotion categories that reached statistical significance, both models generally showed positive relationships, indicating that the AI’s predicted intensity increased alongside human ratings. However, a critical and consistent divergence was found in the sadness category across both architectures, which showed a significant negative relationship. This inverse correlation indicates that humans and the AI models perceived the intensity of sadness in opposing directions; specifically, stimuli that human observers evaluated as highly arousing (e.g., expressions of intense grief) were perceived as low-arousal by the models, and vice versa.

3.2.2 Gender-Stratified Similarity

We conducted a gender-stratified correlation analysis for both models. For MobileViT, arousal ratings for male target faces were significantly correlated with human ratings only for fear and anger expressions, while all other categories showed no significant relationship (p > .05). In contrast, for female target faces, the model exhibited significant correlations for happiness, anger, and sadness, whereas neutral, fear, and disgust remained non-significant. A similar gendered pattern was observed for enet_b0_8_va_mtl; for male faces, significant correlations were restricted to anger and disgust, while for female faces, all categories reached significance except for neutral and fear.

Critically, for both models, the significant relationship identified in female sadness was consistently negative. This indicates that the inverse perceptual correspondence for emotional intensity—where AI perceives high-intensity human distress as low-arousal—is particularly pronounced when evaluating female faces.

3.3 Grad-CAM Attention Analysis

Figure 4. Grad-CAM attention maps for sadness stimuli, stratified by stimulus gender and human arousal rating (high vs. low). Each cell shows the averaged Grad-CAM heatmap overlaid on a representative face. Warm colors indicate regions of higher model attention.

To investigate the mechanism underlying the sadness paradox, we conducted a Grad-CAM analysis [Selvaraju et al., 2017] on both models. Since both models output continuous arousal values (regression), we applied Grad-CAM by computing gradients of the arousal output neuron with respect to the final convolutional layer activations, following the standard adaptation for regression tasks [Selvaraju et al., 2017]. For each sadness stimulus (N=240), we generated activation maps and computed group-level averages stratified by stimulus gender and human-rated arousal intensity (median split: high-arousal vs. low-arousal).

3.3.1 Diagnostic Region Attention Rate (DRAR)

We defined diagnostic facial regions based on the Facial Action Coding System (FACS) [Ekman & Friesen, 1978], selecting AUs specifically associated with sadness: the brow region (AU1 inner brow raise, AU4 brow lowerer), eye region (AU5 upper lid raise, AU7 lid tightener), nose region (AU9 nose wrinkler), and mouth region (AU15 lip corner depressor, AU17 chin raiser, AU20 lip stretcher). Regions were delineated using facial landmark detection (68-point model) with fixed proportional bounding boxes around each AU cluster. We calculated the Diagnostic Region Attention Rate (DRAR) as the proportion of total Grad-CAM activation falling within these diagnostic regions.

For sadness stimuli, both models showed markedly lower DRAR for female faces compared to male faces (MobileViT: female DRAR = [VALUE]% vs. male DRAR = [VALUE]%; enet_b0_8_va_mtl: female DRAR = [VALUE]% vs. male DRAR = [VALUE]%). This indicates that when evaluating female sad expressions, the models directed proportionally more attention to non-diagnostic regions (e.g., hair, jawline contour, background) rather than the core facial action units associated with sadness.

3.3.2 Arousal-Dependent Attention Shift

For stimuli rated as high-arousal by human observers, the models’ attention maps showed a distinct pattern: attention was more dispersed across the face and less concentrated on the mouth and brow regions compared to low-arousal stimuli. This misattention pattern was especially pronounced for female faces, where high-arousal sad expressions—characterized by furrowed brows and compressed lips—failed to capture proportional model attention. This provides a mechanistic explanation for the observed negative correlation: the very facial cues that signal intense distress to human observers are systematically under-attended by both AI architectures.

3.3.3 Cross-Model Comparison

Despite their architectural differences, MobileViT and enet_b0_8_va_mtl showed highly similar attention patterns for sadness (cosine similarity of averaged heatmaps: [VALUE] for male, [VALUE] for female faces), notably higher than the cross-model similarity observed for anger ([VALUE]) and happiness ([VALUE]), which served as control conditions. This convergent misattention specifically in sadness across distinct architectures suggests that the bias originates from shared properties of the training data (AffectNet) rather than from architecture-specific limitations.


4 DISCUSSION & CONCLUSION

4.1 Results Summary

When comparing the ratings of valence and arousal expressed in facial expressions as assessed by two AI models and human raters, the results showed that valence ratings were relatively similar between humans and the AI models, whereas arousal ratings exhibited comparatively lower correspondence between the two. Gender bias in valence and arousal ratings was obtained somewhat differently in humans and AI. In particular, substantial differences were observed for arousal: for the emotion of sadness, arousal ratings for male faces showed no significant correspondence between the two groups, whereas for female faces, a negative correlation was observed.

4.2 Implication for HCI field

Our findings highlight a “sadness paradox” where AI and human arousal ratings are inversely correlated, posing a significant risk for empathetic failures in sensitive UX contexts like mental health support. The Grad-CAM analysis provides a mechanistic explanation: both models systematically under-attend to diagnostically relevant facial regions (brow furrow, compressed lips) in high-arousal female sadness, instead distributing attention to peripheral areas. This misattention pattern, shared across architecturally distinct models, suggests the root cause lies in the training data rather than model design. Specifically, benchmarks such as AffectNet [Mollahosseini et al., 2017] rely on sparse human annotation (N=12) and exhibit lower inter-rater reliability for arousal, which may inadequately represent the subtle muscular cues of intense female grief. We note that Grad-CAM itself may exhibit demographic biases [Huber et al., 2023], warranting caution in interpreting attention patterns as direct reflections of model decision processes. Nevertheless, the convergent misattention across architecturally distinct models strengthens the case for a data-driven origin. To ensure robust Human-AI correspondence, the HCI field must move beyond limited datasets and prioritize the integration of large-scale human benchmarks—alongside attention-guided training strategies such as AU cue integration [Belharbi et al., 2024]—into AI training pipelines. Our DRAR framework extends prior work comparing model attention with FACS ground truth [Gaya-Morey et al., 2024] by introducing gender-stratified and arousal-dependent analyses.

4.3 Limitations

Our study is limited by the use of static, portrait-style stimuli with single emotions, which may lack the dynamic complexity of real-world interactions. Furthermore, as the human participants were exclusively Korean, the findings may have limited cultural generalizability. Additionally, as the stimuli were AI-generated, their ecological validity compared to naturalistic facial expressions warrants further investigation. Future research should employ dynamic, multi-modal stimuli and culturally diverse participant pools to validate whether the identified “sadness paradox” and gendered perceptual misalignments are universal or culturally specific.


REFERENCES

Ali Mollahosseini, Behzad Hasani, and Mohammad H. Mahoor. 2017. AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild. IEEE Trans. Affect. Comput. 10, 1 (January-March 2017), 18–31. https://doi.org/10.1109/taffc.2017.2740923

Andrey V. Savchenko. 2022. Video-based frame-level facial analysis of affective behavior on mobile devices using EfficientNets. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2359–2366.

Antoine Toisoul, Jean Kossaifi, Adrian Bulat, Georgios Tzimiropoulos, and Maja Pantic. 2021. Estimation of continuous valence and arousal levels from faces in naturalistic conditions. Nat. Mach. Intell. 3, 1 (January 2021), 42–50. https://doi.org/10.1038/s42256-020-00280-0

Colette Deaudelin, Marc Dussault, and Monique Brodeur. 2003. Human-computer interaction: A review of the research on its affective and social aspects. Can. J. Learn. Technol. 29, 1 (2003).

Dominik Adolph and Georg W. Alpers. 2010. Valence and arousal: a comparison of two sets of emotional facial expressions. Am. J. Psychol. 123, 2 (Summer 2010), 209–219. https://doi.org/10.5406/amerjpsyc.123.2.0209

Elsheikh, R.A., Mohamed, M.A., Abou-Taleb, A.M. et al. 2024. Improved facial emotion recognition model based on a novel deep convolutional structure. Sci Rep 14, 29050 (2024).

Francisco J. Gaya-Morey, Cristina Manresa-Yee, and Jose M. Buades-Rubio. 2024. Unveiling the human-like similarities of automatic facial expression recognition: An empirical exploration through explainable AI. Multimedia Tools and Applications 83 (2024), 58611–58640. https://doi.org/10.1007/s11042-023-17769-6

Paul Ekman and Wallace V. Friesen. 1978. Facial Action Coding System: A Technique for the Measurement of Facial Movement. Consulting Psychologists Press, Palo Alto, CA.

Stefan Huber, Marek Schürmann, Arjan Kuijper, and Amir Rezaei Balef. 2023. Bias in face presentation attack detection: A case study on the effect of Grad-CAM. In Proceedings of the 31st European Signal Processing Conference (EUSIPCO), 755–759.

Souhail Belharbi, Marco Pedersoli, Simon Bacon, Eric Granger, and Luke McCaffrey. 2024. Guided interpretable facial expression recognition via spatial action unit cues. In Proceedings of the 18th IEEE International Conference on Automatic Face and Gesture Recognition (FG), 1–10.

E. Ashby Plant, Janet Shibley Hyde, Dacher Keltner, and Patricia G. Devine. 2000. The gender stereotyping of emotions. Psychol. Women Q. 24, 1 (March 2000), 81–92. https://doi.org/10.1111/j.1471-6402.2000.tb01024.x

Hiroki Nomiya, Kenta Shimokawa, Shushi Namba, Michiko Osumi, and Wataru Sato. 2025. An artificial intelligence model for sensing affective valence and arousal from facial images. Sensors 25, 4 (February 2025), 1188. https://doi.org/10.3390/s25041188

James A. Russell. 1980. A circumplex model of affect. J. Pers. Soc. Psychol. 39, 6 (December 1980), 1161–1178. https://doi.org/10.1037/h0077714

Jean-Yves Baudouin, Floriane Gallian, Jean-Marc Pinoit, and Fabrice Damon. 2025. Arousal, valence, and discrete categories in facial emotion. Sci. Rep. 15, 1 (2025), 40268. https://doi.org/10.1038/s41598-025-24032-5

Maja Pantic, Nicu Sebe, Jeffrey F. Cohn, and Thomas Huang. 2005. Affective multimodal human-computer interaction. In Proceedings of the 13th annual ACM international conference on Multimedia (Multimedia ‘05). Association for Computing Machinery, New York, NY, USA, 669–676. https://doi.org/10.1145/1101149.1101299

Melis Telceken, Durmus Akgun, Serhat Kacar, Koray YESİN, and Murat Yıldız. 2025. Can artificial intelligence understand our emotions? Deep learning applications with face recognition. Curr. Psychol. 44, 9 (March 2025), 7946–7956. https://doi.org/10.1007/s12144-025-07375-0

Michelle Yik, Christopher Mues, Ivan N. Sze, Peter Kuppens, Francis Tuerlinckx, Kim De Roover, and James A. Russell. 2023. On the relationship between valence and arousal in samples across the globe. Emotion 23, 2 (March 2023), 332. https://doi.org/10.1037/emo0001095

Mingxing Tan and Quoc Le. 2019. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the 36th International Conference on Machine Learning (ICML ‘19). PMLR, 6105–6114.

Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017. Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 618–626. https://doi.org/10.1109/ICCV.2017.74

Sachin Mehta and Mohammad Rastegari. 2021. MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer. arXiv preprint arXiv:2110.02178 (October 2021). https://doi.org/10.48550/arXiv.2110.02178

Shishir Kumar Khare, Veronica Blanes-Vidal, Esmaeil S. Nadimi, and U. Rajendra Acharya. 2024. Emotion recognition and artificial intelligence: A systematic review (2014–2023) and research recommendations. Inf. Fusion 102 (February 2024), 102019. https://doi.org/10.1016/j.inffus.2023.102019

Ursula Hess, Reginald B. Adams Jr, and Robert E. Kleck. 2004. Facial appearance, gender, and emotion expression. Emotion 4, 4 (December 2004), 378–388. https://doi.org/10.1037/1528-3542.4.4.378

Yiming Wang, Hui Yu, Weihong Gao, Yang Xia, and Charles Nduka. 2024. MGEED: A multimodal genuine emotion and expression detection database. IEEE Trans. Affect. Comput. 15, 2 (April–June 2024), 606–619. https://doi.org/10.1109/TAFFC.2023.3286351


Revision Note (not for publication): This revised version addresses 5 issues identified in peer review:

  1. ✅ C1: “alignment” → “rank-order consistency” / “correspondence” + Spearman caveat
  2. ✅ M4: Happy ceiling effect explanation added
  3. ✅ C5: enet_b0_8_va_mtl citation corrected to Savchenko (2022)
  4. ✅ m2: Wang et al. (2024) reference completed
  5. ✅ C2: Exploratory analysis statement + Bonferroni correction for sadness
  6. ✅ Abstract “alignment” → “correspondence”, “significantly” → “notably”
  7. ✅ Limitations: synthetic stimuli ecological validity caveat added
  8. ✅ NEW §3.3: Grad-CAM attention analysis (DRAR, arousal-dependent shift, cross-model comparison)
  9. ✅ Discussion updated with Grad-CAM mechanistic explanation
  10. ✅ References: Selvaraju et al. (2017) Grad-CAM, Ekman & Friesen (1978) FACS added