Diagnosing Correctness Probes under Self-Judgement Confounding
Hidden-state readouts often reflect a model's self-judgment (SJ) rather than objective correctness (OC), creating semantic ambiguity in interpretability probes. In conflict cases where SJ and OC disagree, conventional probes incorrectly rank incorrect but self-endorsed responses higher than correct but self-rejected ones. The SJ-associated direction transfers robustly across domains and models, while the OC-associated direction performs below chance, indicating SJ dominance. This transfer asymme
Analysis
TL;DR
- Hidden-state readouts often reflect a model's self-judgment (SJ) rather than objective correctness (OC), creating semantic ambiguity in interpretability probes.
- In conflict cases where SJ and OC disagree, conventional probes incorrectly rank incorrect but self-endorsed responses higher than correct but self-rejected ones.
- The SJ-associated direction transfers robustly across domains and models, while the OC-associated direction performs below chance, indicating SJ dominance.
- This transfer asymmetry persists across middle-to-late layers and various control conditions, suggesting that transferability alone cannot verify objective-correctness semantics.
Why It Matters
This research challenges the validity of using standard interpretability probes to identify "objective correctness" in language models, revealing that many such signals are confounded by the model's own confidence or self-assessment. For AI practitioners and researchers, it highlights a critical flaw in assuming that readable latent features correspond to ground-truth accuracy, necessitating more rigorous diagnostic methods to disentangle belief from fact.
Technical Details
- Conflict Case Construction: The study creates scenarios where Objective Correctness (OC) and Self-Judgment (SJ) predict opposite orderings of hidden-state readouts to isolate confounding factors.
- Direction Estimation: Factorial SJ- and OC-associated directions are estimated and evaluated for polarity across mathematical reasoning and factual recall tasks.
- Model Scope: Analysis covers four instruction-tuned models up to 14B parameters, testing performance on MMLU and binary TruthfulQA without target-domain direction fitting.
- Control Measures: The findings persist under controls for answer likelihood, sequence length, and null-directions, confirming the robustness of the SJ-dominant signal.
- Layer Analysis: The asymmetry between SJ and OC transferability develops specifically in middle-to-late layers of the network architecture.
Industry Insight
- Re-evaluate Interpretability Metrics: Practitioners should not rely solely on probe transferability as proof of semantic alignment with objective truth; additional validation against ground-truth labels is essential.
- Design Robust Diagnostics: Future work must explicitly account for self-judgment confounding when building tools to monitor model honesty or error detection mechanisms.
- Caution in Scaling: As models grow larger, the gap between self-confidence and actual correctness may widen, requiring specialized architectural or training interventions to decouple these signals.
Disclaimer: The above content is generated by AI and is for reference only.