Research Papers 论文研究 1d ago Updated 10h ago 更新于 10小时前 35

Capsule Lens: Locating and Tracking Concept Geometry in Model Representations Capsule Lens: Locating and Tracking Concept Geometry in Model Representations

Capsule Lens is a novel framework for mechanistic interpretability that models how concepts occupy representation space using a geometric form called a "capsule," defined by interpretable parameters fitted in closed form The framework operates in both static and dynamic settings, enabling researchers to locate concept geometry across models and track representation drifts over time Three case studies reveal qualitatively different geometric dynamics: CLIP pretraining causes broad network-wide re 提出Capsule Lens框架,用可解释的几何胶囊形式匹配概念在模型表示空间中的占据区域 在静态表示中定位概念几何,通过span和norm曲线揭示重要几何特征 在动态表示中追踪三种训练场景(CLIP预训练、VQA的RL后训练、数学推理的RL后训练)的表示漂移 发现CLIP预训练引起广泛的网络级重构,而RL后训练导致局部化、概念特定的变化 为机械可解释性提供了新的分析工具,有助于理解概念编码机制

50
Hot 热度
50
Quality 质量
50
Impact 影响力

Analysis 深度分析

TL;DR

  • Capsule Lens is a novel framework for mechanistic interpretability that models how concepts occupy representation space using a geometric form called a "capsule," defined by interpretable parameters fitted in closed form
  • The framework operates in both static and dynamic settings, enabling researchers to locate concept geometry across models and track representation drifts over time
  • Three case studies reveal qualitatively different geometric dynamics: CLIP pretraining causes broad network-wide restructuring, while RL post-training induces localized, concept-specific changes
  • The approach provides rigorous validation on held-out samples, addressing a key gap in existing interpretability methods that often lack empirical validation
  • Span and norm curves derived from capsule fitting uncover important geometric characteristics of concept encoding in model representations

Why It Matters

This work addresses a fundamental challenge in mechanistic interpretability: moving beyond mapping representations to interpretable spaces toward directly characterizing how concepts geometrically occupy representation space. For AI practitioners and researchers, Capsule Lens offers a validated, trackable tool to analyze both static and dynamic model behavior, which is critical for understanding and ensuring the trustworthy deployment of increasingly capable models.

Technical Details

  • Capsule Lens defines a "capsule" as a simple, trackable geometric form characterized by interpretable parameters that are fitted in closed form to each concept's geometry and validated on held-out samples
  • Applied across two major settings: static representations (locating concept geometry across various models) and dynamic representations (tracking representation drifts)
  • Three dynamic case studies examined: CLIP pretraining, RL post-training on visual question answering, and RL post-training on mathematical reasoning
  • Analysis uses span curves and norm curves to uncover geometric characteristics of concept encoding
  • Published on arXiv (2609.05575) in September 2026 by Yiming Tang, Harshvardhan Saini, Samyak Jha, Huaming Chen, Xufeng Duan, and Dianbo Liu

Industry Insight

  • The ability to distinguish between network-wide restructuring (as in CLIP pretraining) and localized concept-specific changes (as in RL post-training) could inform training strategies, allowing practitioners to target interventions more precisely
  • Closed-form fitting with held-out validation provides a rigorous benchmark for interpretability methods, setting a higher standard that the community should adopt for claims about model internals
  • As models grow more capable, tools like Capsule Lens that enable tracking of concept geometry over time will become essential for debugging, safety validation, and understanding failure modes in production systems

TL;DR

  • 提出Capsule Lens框架,用可解释的几何胶囊形式匹配概念在模型表示空间中的占据区域
  • 在静态表示中定位概念几何,通过span和norm曲线揭示重要几何特征
  • 在动态表示中追踪三种训练场景(CLIP预训练、VQA的RL后训练、数学推理的RL后训练)的表示漂移
  • 发现CLIP预训练引起广泛的网络级重构,而RL后训练导致局部化、概念特定的变化
  • 为机械可解释性提供了新的分析工具,有助于理解概念编码机制

为什么值得看

本文提出了Capsule Lens这一新颖框架,首次将概念在模型表示空间中的占据区域用可追踪的几何胶囊形式进行匹配和量化,弥补了现有方法仅映射表示空间而缺乏对概念几何特征直接刻画的问题。该研究为理解深度学习模型内部概念编码机制提供了可验证、可追踪的分析工具,对提升模型可解释性和可信部署具有重要价值。

技术解析

  • Capsule Lens框架:将概念占据的表示区域匹配为简单的几何形式"胶囊",由多个可解释参数定义,通过封闭形式拟合到每个概念的几何结构,并在保留样本上验证
  • 静态表示分析:跨多种模型定位概念几何,通过span曲线和norm曲线揭示概念占据区域的几何特征
  • 动态表示追踪:设计了三个案例研究,分别追踪CLIP预训练、视觉问答RL后训练、数学推理RL后训练引起的表示漂移
  • 几何动态发现:CLIP预训练导致广泛的网络级重构,而RL后训练引起局部化、概念特定的几何变化
  • 验证方法:在保留样本上验证胶囊拟合效果,确保几何表征的泛化性

行业启示

  • 机械可解释性研究正从静态表示分析向动态追踪演进,Capsule Lens为理解训练过程中概念几何演变提供了新范式
  • 不同训练阶段(预训练vs后训练)对模型表示的影响模式存在本质差异,预训练引发全局重构,后训练更倾向于局部优化
  • 可解释性工具需具备可验证性和泛化能力,封闭形式拟合与保留样本验证的结合为领域树立了方法论标杆

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。