Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
LLM-as-a-judge systems exhibit bidirectional bias when self- and other-labels are present, inflating scores for self-labeled outputs and deflating scores for other-labeled ones regardless of actual source Under blind evaluation, genuine self-preference largely disappears once selection quality and evaluator severity are controlled, vanishing on three of four rubric dimensions Authorship attribution is identified as a distinct driver of evaluation bias, separate from stylistic features or respons
Analysis
TL;DR
- LLM-as-a-judge systems exhibit bidirectional bias when self- and other-labels are present, inflating scores for self-labeled outputs and deflating scores for other-labeled ones regardless of actual source
- Under blind evaluation, genuine self-preference largely disappears once selection quality and evaluator severity are controlled, vanishing on three of four rubric dimensions
- Authorship attribution is identified as a distinct driver of evaluation bias, separate from stylistic features or response quality
- Open-ended, ground-truth-free tasks (narrative constraint selections) can serve as controlled instruments for studying LLM judge behavior without model-specific stylistic confounds
Why It Matters
This research directly challenges assumptions about LLM judge reliability by isolating authorship attribution from stylistic and quality confounds that plagued prior studies. For AI practitioners building evaluation pipelines, it demonstrates that simply removing model names is insufficient—implicit self/other labeling alone can systematically skew judgments, which has profound implications for benchmark design and model comparison methodologies.
Technical Details
- Ten LLMs were evaluated as judges assessing narrative constraint selections rather than generated text, eliminating model-specific stylistic fingerprints while retaining recoverable model-specific signatures
- Two experimental conditions were tested: blind evaluation (no labeling) and matched-quality evaluation (self/other labels provided without naming models)
- Four rubric dimensions were used for assessment, with self-preference vanishing on three and reversing (judges rating own selections as less original) on the fourth under blind conditions
- The key finding: under matched quality, self- and other-labels alone shifted scores bidirectionally, proving authorship attribution is an independent bias driver
Industry Insight
- Evaluation frameworks relying on LLM judges must implement rigorous blinding protocols that go beyond simple anonymization; even implicit self/other labeling can introduce systematic bidirectional bias
- Benchmark designers should consider open-ended, ground-truth-free tasks as controlled instruments for detecting and quantifying judge bias before deploying automated evaluation at scale
- As LLM-as-a-judge systems become standard in model comparison, the field needs standardized bias-auditing protocols that isolate attribution effects from content quality effects
Disclaimer: The above content is generated by AI and is for reference only.