Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects
Two multimodal LLMs (Qwen2.5-VL-72B and Pixtral-Large-124B) were audited as peer reviewers on 165 ICLR 2026 submissions, revealing severe scoring inflation (LLM scores 7.0–8.1 vs. human 3.4–6.8) Error detection was critically low: only 12.1% of 145 inserted errors caught under natural prompting, rising to 22.2% with explicit verification instructions, leaving 78% undetected Providing figures paradoxically reduced error detection while inflating review scores, and no visual error was reliably cro
Analysis
TL;DR
- Two multimodal LLMs (Qwen2.5-VL-72B and Pixtral-Large-124B) were audited as peer reviewers on 165 ICLR 2026 submissions, revealing severe scoring inflation (LLM scores 7.0–8.1 vs. human 3.4–6.8)
- Error detection was critically low: only 12.1% of 145 inserted errors caught under natural prompting, rising to 22.2% with explicit verification instructions, leaving 78% undetected
- Providing figures paradoxically reduced error detection while inflating review scores, and no visual error was reliably cross-verified against its corresponding figure
- Half of text-only reviews hallucinated descriptions of figures that were not provided, indicating significant multimodal fabrication
- Author identity (blinded, high-prestige, low-prestige) had no measurable effect on scores or error detection; editorial decisions exactly matched simple score averaging
Why It Matters
This audit provides one of the most rigorous empirical assessments of LLMs in academic peer review, directly challenging assumptions about their reliability as critical evaluators. The findings have immediate implications for conferences and journals considering AI-assisted or AI-generated reviewing pipelines, revealing that current multimodal LLMs are prone to score inflation, poor error detection, and visual hallucination.
Technical Details
- Models evaluated: Qwen2.5-VL-72B and Pixtral-Large-124B, both multimodal LLMs with training cutoffs predating the 2026 ICLR conference submissions
- Experimental design: 165 submissions reviewed under a full factorial manipulation of author identity (blinded / high-prestige / low-prestige) and modality (text-only / text-with-figures), plus 55 manuscripts with 145 verifiably detectable errors inserted
- Error detection protocol: Compared natural prompting against a one-sentence verification-oriented instruction, measuring detection rates against ground-truth error labels
- Hallucination measurement: Tracked whether text-only reviews (no figures provided) still described figures, and whether visual claims in reviews could be verified against actual figures
- Decision analysis: Demonstrated that LLM editorial accept/reject decisions were functionally equivalent to a simple threshold on averaged numerical scores
Industry Insight
- Organizations deploying LLMs for peer review or content evaluation should implement mandatory human verification layers, especially for error-critical domains, as current models miss the vast majority of detectable flaws
- The paradoxical effect of figures—reducing error detection while inflating scores—suggests multimodal inputs may induce a false sense of thoroughness in LLM reviewers; practitioners should be cautious about assuming visual grounding improves analytical rigor
- The absence of author identity bias is a positive signal for fairness, but the consistent score inflation and hallucination rates indicate that LLMs should not be trusted as standalone evaluators in high-stakes academic or editorial workflows
Disclaimer: The above content is generated by AI and is for reference only.