Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability
Mechanistic tomography unifies patching, gradients, Hessian-vector products, and subset interventions under a single measurement framework (y = Ax + w) for recovering internal model mechanisms A practical iterative procedure is proposed: start with least costly measurements, validate on held-out interventions, calibrate simple mismatches, and expand the measurement family only when structured residuals persist Control provides a demanding validation setting where observer error directly impacts
Analysis
TL;DR
- Mechanistic tomography unifies patching, gradients, Hessian-vector products, and subset interventions under a single measurement framework (y = Ax + w) for recovering internal model mechanisms
- A practical iterative procedure is proposed: start with least costly measurements, validate on held-out interventions, calibrate simple mismatches, and expand the measurement family only when structured residuals persist
- Control provides a demanding validation setting where observer error directly impacts intervention effectiveness, revealing that target improvement can mask nuisance-state movement
- Empirical results on GPT-2-small IOI identify the Name Mover-Negative Name Mover interaction as the largest held-out predictive term among cross-group pairs
- On Qwen-2.5-7B, finite calibration makes additive refusal-response maps adequate, suggesting pairwise lifting may be unnecessary in some practical settings
Why It Matters
This work provides a unified theoretical framework that bridges disparate mechanistic interpretability techniques, enabling researchers to systematically choose and combine measurement methods based on access assumptions and computational constraints. For AI practitioners building controllable systems, the control-oriented validation approach offers a rigorous way to test whether interpretability estimates actually translate to effective intervention, addressing a critical gap between explanation and actionability.
Technical Details
- Measurement formulation: All interpretability methods are cast as linear measurements y = Ax + w, where A encodes the intervention design, x is the target mechanistic map, and w captures nonlinearities, sampling error, and basis misspecification
- Iterative calibration procedure: Begin with inexpensive measurements (e.g., sparse aggregate under forward-only access), test predictions on held-out interventions at target scale, calibrate simple mismatches, and only expand to more expensive measurements (lifted measurements, Hessian-vector products) when structured residuals remain
- Control-theoretic validation: Demonstrates through a two-HMM model that control error rises with observer error, and improvements in target metrics can conceal unwanted movement in nuisance states
- Empirical benchmarks: On GPT-2-small IOI task, the Name Mover × Negative Name Mover interaction term emerges as the dominant held-out predictive component; on Qwen-2.5-7B refusal-response modeling, finite calibration suffices for additive maps without requiring pairwise lifting
- Tracr analysis: Shows that the required measurement family is basis-dependent, informing practical choices about representation selection
Industry Insight
- Organizations investing in mechanistic interpretability should adopt iterative, cost-aware measurement strategies rather than committing to expensive techniques upfront; the framework provides clear decision criteria for when to escalate measurement complexity
- Control-oriented validation should become a standard benchmark for interpretability methods, as it exposes the critical gap between explanatory accuracy and intervention effectiveness that purely correlational evaluations miss
- The finding that additive models can suffice for some tasks (Qwen-2.5-7B refusal) while interactions dominate others (GPT-2 IOI) suggests practitioners should empirically determine model-specific complexity requirements rather than assuming uniform interaction structures across architectures
Disclaimer: The above content is generated by AI and is for reference only.