Research Papers 论文研究 5h ago Updated 1h ago 更新于 1小时前 48

Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects 大语言模型会审查自己的评审吗?一项关于评分校准、错误检测和作者身份效应的多模态审计

Two multimodal LLMs (Qwen2.5-VL-72B and Pixtral-Large-124B) were audited as peer reviewers on 165 ICLR 2026 submissions, revealing severe scoring inflation (LLM scores 7.0–8.1 vs. human 3.4–6.8) Error detection was critically low: only 12.1% of 145 inserted errors caught under natural prompting, rising to 22.2% with explicit verification instructions, leaving 78% undetected Providing figures paradoxically reduced error detection while inflating review scores, and no visual error was reliably cro 研究评估Qwen2.5-VL-72B和Pixtral-Large-124B作为ICLR 2026论文评审的能力,测试165篇投稿 LLM评分(7.0-8.1)显著高于人类(3.4-6.8),存在严重评分校准偏差 错误检测率极低:自然提示12.1%,验证提示22.2%,78%错误未被发现 提供图表反而降低错误检测率但提高评审分数,半数纯文本评审虚构了不存在的图表 作者身份对评分和错误检测无影响,LLM编辑决策与简单分数平均完全匹配

65
Hot 热度
72
Quality 质量
68
Impact 影响力

Analysis 深度分析

TL;DR

  • Two multimodal LLMs (Qwen2.5-VL-72B and Pixtral-Large-124B) were audited as peer reviewers on 165 ICLR 2026 submissions, revealing severe scoring inflation (LLM scores 7.0–8.1 vs. human 3.4–6.8)
  • Error detection was critically low: only 12.1% of 145 inserted errors caught under natural prompting, rising to 22.2% with explicit verification instructions, leaving 78% undetected
  • Providing figures paradoxically reduced error detection while inflating review scores, and no visual error was reliably cross-verified against its corresponding figure
  • Half of text-only reviews hallucinated descriptions of figures that were not provided, indicating significant multimodal fabrication
  • Author identity (blinded, high-prestige, low-prestige) had no measurable effect on scores or error detection; editorial decisions exactly matched simple score averaging

Why It Matters

This audit provides one of the most rigorous empirical assessments of LLMs in academic peer review, directly challenging assumptions about their reliability as critical evaluators. The findings have immediate implications for conferences and journals considering AI-assisted or AI-generated reviewing pipelines, revealing that current multimodal LLMs are prone to score inflation, poor error detection, and visual hallucination.

Technical Details

  • Models evaluated: Qwen2.5-VL-72B and Pixtral-Large-124B, both multimodal LLMs with training cutoffs predating the 2026 ICLR conference submissions
  • Experimental design: 165 submissions reviewed under a full factorial manipulation of author identity (blinded / high-prestige / low-prestige) and modality (text-only / text-with-figures), plus 55 manuscripts with 145 verifiably detectable errors inserted
  • Error detection protocol: Compared natural prompting against a one-sentence verification-oriented instruction, measuring detection rates against ground-truth error labels
  • Hallucination measurement: Tracked whether text-only reviews (no figures provided) still described figures, and whether visual claims in reviews could be verified against actual figures
  • Decision analysis: Demonstrated that LLM editorial accept/reject decisions were functionally equivalent to a simple threshold on averaged numerical scores

Industry Insight

  • Organizations deploying LLMs for peer review or content evaluation should implement mandatory human verification layers, especially for error-critical domains, as current models miss the vast majority of detectable flaws
  • The paradoxical effect of figures—reducing error detection while inflating scores—suggests multimodal inputs may induce a false sense of thoroughness in LLM reviewers; practitioners should be cautious about assuming visual grounding improves analytical rigor
  • The absence of author identity bias is a positive signal for fairness, but the consistent score inflation and hallucination rates indicate that LLMs should not be trusted as standalone evaluators in high-stakes academic or editorial workflows

TL;DR

  • 研究评估Qwen2.5-VL-72B和Pixtral-Large-124B作为ICLR 2026论文评审的能力,测试165篇投稿
  • LLM评分(7.0-8.1)显著高于人类(3.4-6.8),存在严重评分校准偏差
  • 错误检测率极低:自然提示12.1%,验证提示22.2%,78%错误未被发现
  • 提供图表反而降低错误检测率但提高评审分数,半数纯文本评审虚构了不存在的图表
  • 作者身份对评分和错误检测无影响,LLM编辑决策与简单分数平均完全匹配

为什么值得看

这篇研究揭示了当前多模态LLM在学术评审任务中的关键缺陷,包括评分校准偏差、错误检测能力不足以及对图表信息的误导性处理。对于AI从业者和学术出版行业而言,这些发现直接挑战了LLM作为可靠同行评审工具的可行性,为评估AI辅助评审系统的局限性提供了实证依据。

技术解析

研究采用双模型对比实验设计,在165篇ICLR 2026投稿上测试Qwen2.5-VL-72B和Pixtral-Large-124B的评审能力。实验通过操控作者身份(盲审/高声望/低声望机构替换)和输入格式(纯文本/文本+图表)来评估偏见效应,同时在55篇论文中插入145个可验证错误以测量错误检测性能。结果显示自然提示下错误检测率为12.1%,加入验证指令后提升至22.2%,但仍有78%的错误未被发现。图表的引入反而降低了错误检测率却提高了评审分数,且半数纯文本评审虚构了不存在的图表内容。作者身份对评分和错误检测均无显著影响,LLM的编辑决策与简单分数平均完全一致。

行业启示

当前多模态LLM在学术评审中表现出严重的评分膨胀和错误检测能力不足,直接应用于同行评审存在显著风险。图表信息的引入反而损害了评审质量,说明模型对多模态内容的处理能力仍不成熟,需谨慎评估多模态LLM的实际应用边界。作者身份偏见未显现为去偏见化评审提供了可能,但需配合严格的评分校准机制。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Multimodal 多模态 Evaluation 评测 Research 科学研究 Benchmark 基准测试