Do Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma
The study tests the "semantic-concentration prediction" by measuring holonomy on active Sparse Autoencoder (SAE) feature planes in Gemma 2 2B. Holonomy is quantified via restricted-Jacobian transport of local frames around small loops in the residual stream, normalized by enclosed area. The preregistered hypothesis was falsified: active-feature planes exhibited significantly less holonomy than matched mixed-feature controls. The result indicates an operational reversal rather than a causal claim
Analysis
TL;DR
- The study tests the "semantic-concentration prediction" by measuring holonomy on active Sparse Autoencoder (SAE) feature planes in Gemma 2 2B.
- Holonomy is quantified via restricted-Jacobian transport of local frames around small loops in the residual stream, normalized by enclosed area.
- The preregistered hypothesis was falsified: active-feature planes exhibited significantly less holonomy than matched mixed-feature controls.
- The result indicates an operational reversal rather than a causal claim that meaning suppresses holonomy, leaving alternative geometric explanations open.
Why It Matters
This research highlights the importance of preregistration and rigorous falsification in mechanistic interpretability, demonstrating how empirical data can overturn theoretical assumptions about semantic concentration. It provides practitioners with concrete methods for measuring geometric properties like holonomy in transformer residual streams, offering new tools for analyzing feature interactions. The findings caution against assuming that high activation or semantic relevance directly correlates with specific geometric curvatures, urging deeper investigation into dictionary geometry and transport mechanisms.
Technical Details
- Model & Scope: Analysis focused on Gemma 2 2B, specifically measuring holonomy at the layer-12 to layer-13 residual-stream readout.
- Measurement Method: Holonomy was calculated by transporting a local frame around small loops using a restricted-Jacobian transport rule and normalizing the resulting rotation by the enclosed area.
- Experimental Design: The study was preregistered with frozen design, materiality thresholds, analysis plans, and verdict rules prior to inspection of measurements.
- Key Finding: Active-feature planes carried less holonomy than matched mixed-feature controls, with an adjusted log contrast of -0.29439 and a 95% confidence interval of [-0.43989, -0.14889].
- Alternative Explanations: The study identified potential confounds including activation-strength geometry, degree of feature engagement, dictionary geometry, matched-center displacement, activation-manifold proximity, and transport shear.
Industry Insight
Researchers should prioritize preregistration in interpretability studies to ensure robustness and avoid confirmation bias when testing complex geometric hypotheses. Practitioners investigating semantic concentration should consider that active features may not always exhibit expected geometric properties, necessitating broader diagnostic checks beyond simple activation magnitude. Future work should explore the identified live alternatives, such as transport distortion and dictionary geometry, to better understand the relationship between feature engagement and manifold curvature.
Disclaimer: The above content is generated by AI and is for reference only.