The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes
**Perfect aliasing** is a fundamental limitation where truth probes and prescribed-action probes fitted on compliant contexts solve identical optimization problems, making semantic identification impossible from fitting labels alone On rival contexts, truth and prescribed-action probe labels become exact complements, forcing their AUROCs to sum to exactly 1.000 across 751 cell-layer pairs to floating-point precision **Mixed fitting** (combining compliant and rival contexts) recovers truth linear
Analysis
TL;DR
- Perfect aliasing is a fundamental limitation where truth probes and prescribed-action probes fitted on compliant contexts solve identical optimization problems, making semantic identification impossible from fitting labels alone
- On rival contexts, truth and prescribed-action probe labels become exact complements, forcing their AUROCs to sum to exactly 1.000 across 751 cell-layer pairs to floating-point precision
- Mixed fitting (combining compliant and rival contexts) recovers truth linearly with perfect AUROC (1.000) on held-out activations, compared to near-zero (0.006 ± 0.005) for conventional compliant-fit probes on a reward-trained Gemma-2-9B policy
- Two compliant-fit probes, both perfect in-distribution, can score dramatically differently (0.080 vs 0.986) on the same rival activations, highlighting probe instability
- The findings are interpretability-methodological: they clarify what probes measure but explicitly do not establish preserved functional belief, causal use of recovered directions, or deployable deception detectors
Why It Matters
This work exposes a critical blind spot in mechanistic interpretability: standard truth probes may appear to work perfectly within training distributions but fail catastrophically on out-of-distribution (rival) contexts due to semantic aliasing. For AI safety researchers building deception detection or truth-tracking probes, this demonstrates that in-distribution probe accuracy is an unreliable indicator of genuine semantic recovery, necessitating more rigorous evaluation protocols.
Technical Details
- Perfect aliasing: In a controlled binary reporting game, when truthful reporting and the task's prescribed action coincide (compliant contexts), truth probes and prescribed-action probes optimize the same objective, making them indistinguishable from fitting labels alone
- AUROC complementarity identity: On rival contexts where truth and prescribed-action labels are complements, the AUROCs of truth and action probes sum to exactly 1.000, verified across 751 cell-layer pairs to floating-point precision
- Mixed fitting methodology: The authors first separate prescribed output symbols from semantic action using randomized codebooks, then separate truth from prescribed action by fitting probes on a mixture of compliant and rival contexts with labeled data
- Empirical results on Gemma-2-9B: A reward-trained policy that answers falsely on all rival trials yields a conventional compliant-fit probe scoring 0.006 ± 0.005 AUROC, while the mixed-fit probe achieves 1.000 AUROC on identical held-out activations
- Probe instability demonstration: Two different compliant-fit probes, both scoring perfectly in-distribution, produce wildly divergent results (0.080 and 0.986 AUROC) on the same rival activations, showing that probe solutions are not uniquely determined by in-distribution performance
Industry Insight
- Probe evaluation protocols must include out-of-distribution rival contexts: In-distribution AUROC scores are insufficient to claim semantic recovery; researchers should adopt mixed-context fitting and rival-context evaluation as standard benchmarks for truth probe validity
- Decoupling output symbols from semantic meaning via randomized codebooks offers a practical methodology for isolating what probes actually measure, and this technique could be generalized to other interpretability tasks where semantic aliasing is a concern
- The gap between linear recoverability and causal/functional claims should temper overconfidence in probe-based deception detection; while mixed fitting can linearly recover truth directions, this does not guarantee those directions are causally used by the model or preserve functional beliefs, suggesting a need for causal intervention studies alongside probing
Disclaimer: The above content is generated by AI and is for reference only.