Research Papers 论文研究 2d ago Updated 1d ago 更新于 1天前 48

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges 自我与他者标签在LLM裁判中引发双向偏差

LLM-as-a-judge systems exhibit bidirectional bias when self- and other-labels are present, inflating scores for self-labeled outputs and deflating scores for other-labeled ones regardless of actual source Under blind evaluation, genuine self-preference largely disappears once selection quality and evaluator severity are controlled, vanishing on three of four rubric dimensions Authorship attribution is identified as a distinct driver of evaluation bias, separate from stylistic features or respons 研究揭示LLM-as-a-judge系统中自我偏好偏差的独立驱动因素是作者归属而非文本风格 盲评条件下自我偏好基本消失,但在质量匹配时自我/其他标签会导致双向评分偏移 使用叙事约束选择作为评估对象,成功分离了风格特征与响应质量的影响 开放式无标准答案任务可作为控制实验工具研究LLM评估行为 十种LLM参与实验,覆盖四个评分维度的系统性偏差分析

65
Hot 热度
72
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • LLM-as-a-judge systems exhibit bidirectional bias when self- and other-labels are present, inflating scores for self-labeled outputs and deflating scores for other-labeled ones regardless of actual source
  • Under blind evaluation, genuine self-preference largely disappears once selection quality and evaluator severity are controlled, vanishing on three of four rubric dimensions
  • Authorship attribution is identified as a distinct driver of evaluation bias, separate from stylistic features or response quality
  • Open-ended, ground-truth-free tasks (narrative constraint selections) can serve as controlled instruments for studying LLM judge behavior without model-specific stylistic confounds

Why It Matters

This research directly challenges assumptions about LLM judge reliability by isolating authorship attribution from stylistic and quality confounds that plagued prior studies. For AI practitioners building evaluation pipelines, it demonstrates that simply removing model names is insufficient—implicit self/other labeling alone can systematically skew judgments, which has profound implications for benchmark design and model comparison methodologies.

Technical Details

  • Ten LLMs were evaluated as judges assessing narrative constraint selections rather than generated text, eliminating model-specific stylistic fingerprints while retaining recoverable model-specific signatures
  • Two experimental conditions were tested: blind evaluation (no labeling) and matched-quality evaluation (self/other labels provided without naming models)
  • Four rubric dimensions were used for assessment, with self-preference vanishing on three and reversing (judges rating own selections as less original) on the fourth under blind conditions
  • The key finding: under matched quality, self- and other-labels alone shifted scores bidirectionally, proving authorship attribution is an independent bias driver

Industry Insight

  • Evaluation frameworks relying on LLM judges must implement rigorous blinding protocols that go beyond simple anonymization; even implicit self/other labeling can introduce systematic bidirectional bias
  • Benchmark designers should consider open-ended, ground-truth-free tasks as controlled instruments for detecting and quantifying judge bias before deploying automated evaluation at scale
  • As LLM-as-a-judge systems become standard in model comparison, the field needs standardized bias-auditing protocols that isolate attribution effects from content quality effects

TL;DR

  • 研究揭示LLM-as-a-judge系统中自我偏好偏差的独立驱动因素是作者归属而非文本风格
  • 盲评条件下自我偏好基本消失,但在质量匹配时自我/其他标签会导致双向评分偏移
  • 使用叙事约束选择作为评估对象,成功分离了风格特征与响应质量的影响
  • 开放式无标准答案任务可作为控制实验工具研究LLM评估行为
  • 十种LLM参与实验,覆盖四个评分维度的系统性偏差分析

为什么值得看

这篇论文为LLM评估可靠性提供了关键实证证据,帮助AI从业者理解评估偏差的真正来源。研究结果为设计更公正的LLM-as-a-judge系统提供了方法论基础和实践指导。

技术解析

  • 实验设计采用叙事约束选择作为评估对象,这些选择无模型特定风格指纹但保留可恢复的模型特定签名,成功分离了风格与质量混淆因素
  • 盲评实验显示:控制选择质量和评估者严格度后,自我偏好基本消失,在四个评分维度中三个维度消失,第四个维度出现逆转(自我选择被评更低原创性)
  • 标签实验显示:在质量匹配条件下,仅自我/其他标签(不提及模型名称)就会导致双向偏差——自我标签提升分数,其他标签降低分数
  • 研究贡献:作者归属是评估偏差的独立驱动因素;开放式无标准答案任务可作为研究LLM评估行为的控制工具

行业启示

  • LLM-as-a-judge系统必须引入去偏机制,特别是在评估者知晓输出来源时,标签信息会独立于内容质量影响评分
  • 评估设计应优先采用盲评方式,避免作者归属信息干扰评分结果,确保评估公正性
  • 未来研究应关注如何分离风格特征与内容质量,开发更可靠的LLM评估方法论

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Evaluation 评测 Research 科学研究 Alignment 对齐