Capsule Lens: Locating and Tracking Concept Geometry in Model Representations
Capsule Lens is a novel framework for mechanistic interpretability that models how concepts occupy representation space using a geometric form called a "capsule," defined by interpretable parameters fitted in closed form The framework operates in both static and dynamic settings, enabling researchers to locate concept geometry across models and track representation drifts over time Three case studies reveal qualitatively different geometric dynamics: CLIP pretraining causes broad network-wide re
Analysis
TL;DR
- Capsule Lens is a novel framework for mechanistic interpretability that models how concepts occupy representation space using a geometric form called a "capsule," defined by interpretable parameters fitted in closed form
- The framework operates in both static and dynamic settings, enabling researchers to locate concept geometry across models and track representation drifts over time
- Three case studies reveal qualitatively different geometric dynamics: CLIP pretraining causes broad network-wide restructuring, while RL post-training induces localized, concept-specific changes
- The approach provides rigorous validation on held-out samples, addressing a key gap in existing interpretability methods that often lack empirical validation
- Span and norm curves derived from capsule fitting uncover important geometric characteristics of concept encoding in model representations
Why It Matters
This work addresses a fundamental challenge in mechanistic interpretability: moving beyond mapping representations to interpretable spaces toward directly characterizing how concepts geometrically occupy representation space. For AI practitioners and researchers, Capsule Lens offers a validated, trackable tool to analyze both static and dynamic model behavior, which is critical for understanding and ensuring the trustworthy deployment of increasingly capable models.
Technical Details
- Capsule Lens defines a "capsule" as a simple, trackable geometric form characterized by interpretable parameters that are fitted in closed form to each concept's geometry and validated on held-out samples
- Applied across two major settings: static representations (locating concept geometry across various models) and dynamic representations (tracking representation drifts)
- Three dynamic case studies examined: CLIP pretraining, RL post-training on visual question answering, and RL post-training on mathematical reasoning
- Analysis uses span curves and norm curves to uncover geometric characteristics of concept encoding
- Published on arXiv (2609.05575) in September 2026 by Yiming Tang, Harshvardhan Saini, Samyak Jha, Huaming Chen, Xufeng Duan, and Dianbo Liu
Industry Insight
- The ability to distinguish between network-wide restructuring (as in CLIP pretraining) and localized concept-specific changes (as in RL post-training) could inform training strategies, allowing practitioners to target interventions more precisely
- Closed-form fitting with held-out validation provides a rigorous benchmark for interpretability methods, setting a higher standard that the community should adopt for claims about model internals
- As models grow more capable, tools like Capsule Lens that enable tracking of concept geometry over time will become essential for debugging, safety validation, and understanding failure modes in production systems
Disclaimer: The above content is generated by AI and is for reference only.