Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models
Researchers identify emotion-sensitive neurons (ESNs) — sparse decoder neurons selectively associated with emotion categories — across three multimodal foundation models: Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B Visual ESNs are causally meaningful: deactivating them impairs facial emotion recognition, while steering their activations enhances recognition of the associated emotion relative to others Acoustic and visual ESNs show emotion-matched overlap and similar layer-wise distributio
Analysis
TL;DR
- Researchers identify emotion-sensitive neurons (ESNs) — sparse decoder neurons selectively associated with emotion categories — across three multimodal foundation models: Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B
- Visual ESNs are causally meaningful: deactivating them impairs facial emotion recognition, while steering their activations enhances recognition of the associated emotion relative to others
- Acoustic and visual ESNs show emotion-matched overlap and similar layer-wise distributions, indicating partial structural alignment of affective representations across speech and faces
- Cross-modal interventions reveal bidirectional causal transfer: ESNs identified from one modality produce emotion-specific effects when applied to the other
- This provides one of the first cross-modality activation-level analyses of affective functional units in MFMs, showing emotion recognition converges onto sparse decoder-level components localizable and manipulable without training
Why It Matters
This research directly addresses a fundamental question in multimodal AI: whether models process emotions through shared or modality-specific mechanisms, with implications for model interpretability and robustness. For AI practitioners, the finding that emotion-sensitive neurons can be identified and manipulated without retraining opens practical pathways for improving emotion recognition systems and debugging affective failures. The cross-modal transfer results also suggest that insights from one affective modality can inform improvements in another, enabling more efficient model development.
Technical Details
- Models studied: Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B — three leading multimodal foundation models with speech, vision, and language capabilities
- Methodology: Emotion-sensitive neurons (ESNs) were identified as sparse decoder neurons selectively associated with specific emotion categories, using speech emotion recognition and facial expression recognition as complementary probing tasks
- Causal interventions: Visual ESNs were tested through selective deactivation (impairing associated facial emotion recognition) and activation steering (enhancing recognition of the target emotion relative to others), establishing causal rather than merely correlational relationships
- Cross-modal analysis: Acoustic and visual ESNs were compared for emotion-matched overlap and layer-wise distribution patterns, revealing partial structural alignment; bidirectional cross-modal interventions demonstrated that ESNs from one modality produce emotion-specific effects when applied to the other
Industry Insight
- The ability to localize and manipulate emotion-sensitive neurons without retraining suggests that post-hoc interpretability and intervention techniques could become standard tools for auditing and improving affective AI systems in production
- Cross-modal causal transfer implies that investment in emotion recognition for one modality (e.g., facial expressions) may yield compounding returns when extended to others (e.g., speech), supporting integrated multimodal emotion pipelines over siloed single-modality approaches
- As regulatory and ethical scrutiny of affective AI increases, these findings provide a mechanistic basis for explainability — practitioners can point to specific decoder neurons and demonstrate how emotional classifications are produced, which could support compliance with emerging AI transparency requirements
Disclaimer: The above content is generated by AI and is for reference only.