Research Papers 论文研究 3d ago Updated 2d ago 更新于 2天前 47

Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models 跨越语音与面部的情感:多模态基础模型中的共享情感机制

Researchers identify emotion-sensitive neurons (ESNs) — sparse decoder neurons selectively associated with emotion categories — across three multimodal foundation models: Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B Visual ESNs are causally meaningful: deactivating them impairs facial emotion recognition, while steering their activations enhances recognition of the associated emotion relative to others Acoustic and visual ESNs show emotion-matched overlap and similar layer-wise distributio 研究在多模态基础模型(MFMs)中探索情绪敏感神经元(ESNs),发现语音和面部情绪识别存在共享的情感功能单元 视觉ESNs具有因果意义:去激活选择性损害相关面部情绪识别,引导激活选择性增强识别 声学和视觉ESNs显示情绪匹配的重叠和相似的层分布,表明跨模态情感表征存在部分结构对齐 跨模态干预揭示双向因果转移:从一个模态识别的ESNs在应用于另一模态时产生情绪特异性效应 这是MFMs中情感功能单元的首次跨模态激活级别分析之一,表明情绪识别汇聚到可定位和操作的稀疏解码器组件

62
Hot 热度
72
Quality 质量
68
Impact 影响力

Analysis 深度分析

TL;DR

  • Researchers identify emotion-sensitive neurons (ESNs) — sparse decoder neurons selectively associated with emotion categories — across three multimodal foundation models: Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B
  • Visual ESNs are causally meaningful: deactivating them impairs facial emotion recognition, while steering their activations enhances recognition of the associated emotion relative to others
  • Acoustic and visual ESNs show emotion-matched overlap and similar layer-wise distributions, indicating partial structural alignment of affective representations across speech and faces
  • Cross-modal interventions reveal bidirectional causal transfer: ESNs identified from one modality produce emotion-specific effects when applied to the other
  • This provides one of the first cross-modality activation-level analyses of affective functional units in MFMs, showing emotion recognition converges onto sparse decoder-level components localizable and manipulable without training

Why It Matters

This research directly addresses a fundamental question in multimodal AI: whether models process emotions through shared or modality-specific mechanisms, with implications for model interpretability and robustness. For AI practitioners, the finding that emotion-sensitive neurons can be identified and manipulated without retraining opens practical pathways for improving emotion recognition systems and debugging affective failures. The cross-modal transfer results also suggest that insights from one affective modality can inform improvements in another, enabling more efficient model development.

Technical Details

  • Models studied: Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B — three leading multimodal foundation models with speech, vision, and language capabilities
  • Methodology: Emotion-sensitive neurons (ESNs) were identified as sparse decoder neurons selectively associated with specific emotion categories, using speech emotion recognition and facial expression recognition as complementary probing tasks
  • Causal interventions: Visual ESNs were tested through selective deactivation (impairing associated facial emotion recognition) and activation steering (enhancing recognition of the target emotion relative to others), establishing causal rather than merely correlational relationships
  • Cross-modal analysis: Acoustic and visual ESNs were compared for emotion-matched overlap and layer-wise distribution patterns, revealing partial structural alignment; bidirectional cross-modal interventions demonstrated that ESNs from one modality produce emotion-specific effects when applied to the other

Industry Insight

  • The ability to localize and manipulate emotion-sensitive neurons without retraining suggests that post-hoc interpretability and intervention techniques could become standard tools for auditing and improving affective AI systems in production
  • Cross-modal causal transfer implies that investment in emotion recognition for one modality (e.g., facial expressions) may yield compounding returns when extended to others (e.g., speech), supporting integrated multimodal emotion pipelines over siloed single-modality approaches
  • As regulatory and ethical scrutiny of affective AI increases, these findings provide a mechanistic basis for explainability — practitioners can point to specific decoder neurons and demonstrate how emotional classifications are produced, which could support compliance with emerging AI transparency requirements

TL;DR

  • 研究在多模态基础模型(MFMs)中探索情绪敏感神经元(ESNs),发现语音和面部情绪识别存在共享的情感功能单元
  • 视觉ESNs具有因果意义:去激活选择性损害相关面部情绪识别,引导激活选择性增强识别
  • 声学和视觉ESNs显示情绪匹配的重叠和相似的层分布,表明跨模态情感表征存在部分结构对齐
  • 跨模态干预揭示双向因果转移:从一个模态识别的ESNs在应用于另一模态时产生情绪特异性效应
  • 这是MFMs中情感功能单元的首次跨模态激活级别分析之一,表明情绪识别汇聚到可定位和操作的稀疏解码器组件

为什么值得看

  • 为理解多模态基础模型如何处理跨模态情感信息提供了机制性洞察,填补了情感功能单元研究的空白
  • 展示了无需训练即可定位和操作情感相关神经元的可行性,为模型可解释性和可控性研究开辟新路径

技术解析

  • 研究在三个多模态基础模型中识别情绪敏感神经元:Gemma-4-12B-it、MiniCPM-o-4.5和Qwen2.5-Omni-7B,使用稀疏解码器神经元作为分析对象
  • 采用语音情绪识别和面部表情识别作为互补探针,分别识别声学ESNs和视觉ESNs
  • 通过因果干预验证视觉ESNs的功能意义:去激活实验显示选择性损害相关面部情绪识别,引导激活实验显示选择性增强特定情绪识别
  • 发现声学ESNs和视觉ESNs在情绪匹配上存在重叠,且层分布模式相似,表明跨模态情感表征的部分结构对齐
  • 跨模态干预实验揭示双向因果转移效应,证明ESNs可在不同模态间迁移并产生情绪特异性影响

行业启示

  • 多模态模型的情感处理能力可能基于共享的稀疏解码器组件,这为模型可解释性研究和情感计算应用提供了新的技术路径
  • 跨模态情感表征的结构对齐表明不同模态的情感信息可能在模型内部以相似方式编码,为跨模态情感识别系统的开发提供了理论依据
  • 无需重新训练即可定位和操作情感相关神经元的能力,为情感AI的可控性和安全性研究开辟了新方向

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Multimodal 多模态 Speech 语音 Research 科学研究 Evaluation 评测 Dataset 数据集