Research Papers 论文研究 3h ago Updated 46m ago 更新于 46分钟前 48

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks 每一个错误答案都算数:LLM多项选择题基准的选项级心理测量学

Introduces LLM-NRM, an option-aware psychometric framework that models the full distribution over answer choices rather than binary correct/incorrect scoring Demonstrates that incorrect responses carry systematic information about LLM behavior, with distractor identity contributing +101% additional Fisher Information beyond correctness alone Achieves Spearman correlation of 0.920 with external human-preference Elo leaderboard and 0.943 when using incorrect responses alone for ability estimation 提出LLM-NRM(LLM名义响应模型),一种选项感知的心理测量框架,建模LLM对所有答案选项的完整分布而非仅关注正确/错误二元结果 在189个LLM、31,554个题目、14个基准测试上验证,LLM-NRM预测LLM-题目交互的准确性优于传统二元IRT模型和名义响应基线 干扰项(distractor)身份贡献了超出正确性判断+101%的额外Fisher信息,错误回答本身即可恢复高保真能力估计(Spearman 0.943) 基于学习到的题目参数实现高效基准测试:仅41个精选题目即可保持完整题库排名(Kendall 0.85),实现770倍压缩 能力估计与外部人类偏好Elo排行榜达到Spear

62
Hot 热度
76
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Introduces LLM-NRM, an option-aware psychometric framework that models the full distribution over answer choices rather than binary correct/incorrect scoring
  • Demonstrates that incorrect responses carry systematic information about LLM behavior, with distractor identity contributing +101% additional Fisher Information beyond correctness alone
  • Achieves Spearman correlation of 0.920 with external human-preference Elo leaderboard and 0.943 when using incorrect responses alone for ability estimation
  • Enables efficient benchmarking: 41 selected items preserve full-bank ranking (Kendall's 0.85), achieving a 770x reduction in evaluation cost
  • Validated across 189 LLMs, 31,554 items from 14 benchmarks, outperforming both binary Item Response models and conventional nominal-response baselines

Why It Matters

This work fundamentally challenges the binary scoring paradigm that dominates LLM evaluation, showing that the pattern of wrong answers is as informative as correctness itself. For AI practitioners, this means benchmark results can be dramatically compressed without losing ranking fidelity, reducing evaluation costs by orders of magnitude. For researchers, it opens new avenues for diagnosing model behavior through psychometric analysis of option-level preferences.

Technical Details

  • LLM-NRM Framework: Extends the classical Nominal Response Model from psychometrics to LLM evaluation, jointly estimating model ability and option-level item characteristics while disentangling response calibration sharpness, positional preference, and difficulty-dependent fallback behavior
  • Evaluation Scale: Tested across 189 LLMs and 31,554 items spanning 14 diverse multiple-choice benchmarks
  • Information-Theoretic Analysis: Quantified distractor identity contribution as +101% additional Fisher Information per item beyond binary correctness, demonstrating that wrong-answer patterns are highly informative
  • Benchmark Compression: Learned item parameters enable subset selection, where only 41 items (from the full bank) preserve ranking with Kendall's correlation of 0.85, yielding 770x evaluation reduction
  • Superior Predictive Performance: LLM-NRM outperforms both binary Item Response Theory models and conventional nominal-response baselines on held-out LLM-item interaction prediction

Industry Insight

  • Benchmark design should evolve from binary scoring to option-aware psychometric frameworks that extract maximum signal from every response, potentially becoming the new standard for reliable LLM evaluation
  • The 770x compression ratio suggests that expensive full-bank evaluations may be replaceable by carefully selected item subsets, enabling faster iteration cycles and more frequent model assessment
  • The strong correlation (0.920) with human-preference Elo rankings validates that psychometric ability estimates align closely with perceived model quality, giving practitioners a theoretically grounded alternative to raw accuracy metrics

TL;DR

  • 提出LLM-NRM(LLM名义响应模型),一种选项感知的心理测量框架,建模LLM对所有答案选项的完整分布而非仅关注正确/错误二元结果
  • 在189个LLM、31,554个题目、14个基准测试上验证,LLM-NRM预测LLM-题目交互的准确性优于传统二元IRT模型和名义响应基线
  • 干扰项(distractor)身份贡献了超出正确性判断+101%的额外Fisher信息,错误回答本身即可恢复高保真能力估计(Spearman 0.943)
  • 基于学习到的题目参数实现高效基准测试:仅41个精选题目即可保持完整题库排名(Kendall 0.85),实现770倍压缩
  • 能力估计与外部人类偏好Elo排行榜达到Spearman 0.920的相关性,验证了选项级分析的有效性

为什么值得看

本文挑战了LLM评估中长期沿用的二元评分范式,证明错误答案中蕴含系统性且可量化的测量信息,为构建更精确、更高效的LLM能力评估体系提供了新的方法论基础。对AI从业者而言,这直接影响了基准测试设计、模型能力排序和评估效率优化。

技术解析

  • LLM-NRM模型架构:将经典心理测量学中的名义响应模型(Nominal Response Model)扩展至LLM评估场景,同时估计LLM能力参数和题目级参数(包括选项吸引力、位置偏差、难度依赖的回退行为),实现选项级别的细粒度建模。
  • 实验规模与验证:覆盖189个LLM、31,554个题目、14个主流基准测试,在保留测试集上验证预测准确性,并与二元IRT模型及传统名义响应基线进行对比。
  • 关键量化指标:干扰项身份贡献+101%额外Fisher信息;仅错误回答即可恢复Spearman 0.943的能力估计;与人类偏好Elo排行榜相关性达0.920;41题精简集保持Kendall 0.85排名一致性。
  • 效率优化应用:利用学习到的题目参数进行题目选择,实现从完整题库到精简集的770倍压缩,同时保持排名保真度,为低成本高频评估提供可行方案。

行业启示

  • 评估范式升级:MCQ基准测试应从二元评分转向选项级分析,错误答案的分布模式可作为模型能力诊断和校准的重要信号,推动评估指标从"对错"向"偏好结构"演进。
  • 高效基准设计:基于心理测量参数进行题目筛选,可在保持评估信度的前提下大幅缩减测试规模,降低大规模模型评估的计算成本和时间开销。
  • 模型诊断与对齐:选项级分析可揭示模型的位置偏差、校准锐度等系统性行为特征,为模型对齐、安全评估和人类偏好建模提供更细粒度的诊断工具。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Evaluation 评测 Benchmark 基准测试 Research 科学研究