The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure
The paper challenges prior claims that LLMs exhibit a "Teleological Bias" in inferring completed events from progressive descriptions, arguing the benchmark itself is flawed Three conceptual mis-specifications are identified, with "Aspectual Reduction" being the most critical, affecting benchmark construction and conclusions Under strict NLI standards, 76% of benchmark instances do not explicitly rule out event culmination, and native-speaker annotations confirm significant ambiguity (38% of Gro
Analysis
TL;DR
- The paper challenges prior claims that LLMs exhibit a "Teleological Bias" in inferring completed events from progressive descriptions, arguing the benchmark itself is flawed
- Three conceptual mis-specifications are identified, with "Aspectual Reduction" being the most critical, affecting benchmark construction and conclusions
- Under strict NLI standards, 76% of benchmark instances do not explicitly rule out event culmination, and native-speaker annotations confirm significant ambiguity (38% of Group A and 29% of Group C examples)
- The authors introduce "Sufficiency Bias" as a new failure mode: models accept simple-past hypotheses without affirming culmination
- With proper prompts and evaluation, models like Qwen-7B, GPT-5.4, and Qwen-72B can achieve human-comparable performance on aspectual classification
Why It Matters
This paper directly challenges widely cited findings about LLM semantic reasoning failures, demonstrating that benchmark design flaws can produce misleading conclusions about model capabilities. For AI practitioners, it underscores the critical importance of rigorous evaluation methodology and native-speaker validation before drawing strong claims about model limitations. The findings have broad implications for how the community interprets NLI benchmarks and attributes reasoning failures to models versus evaluation artifacts.
Technical Details
- Conceptual Mis-specifications: The authors identify three flaws in the original benchmark, with "Aspectual Reduction" being central—it oversimplifies aspectual distinctions, affecting benchmark construction, analysis, experiments, and conclusions
- Native-Speaker Annotation: Human annotation revealed that 38% of Group A examples and 29% of Group C examples permit alternative interpretations, undermining the original benchmark's ground truth assumptions
- Lexically Matched Minimal Pairs: A new evaluation dataset was constructed to control for lexical variation and isolate aspectual reasoning from confounding factors
- Multi-step Reasoning Framework: Event-semantic NLI is reformulated as a multi-step reasoning problem, assessing both intermediate semantic decisions and final predictions to reveal hidden failure modes
- Identified Failure Modes: Beyond Sufficiency Bias, the authors find "Surface-form Attraction" (models gravitating toward surface-associated answers) and errors in compositional aspectual classification
- Experiments: Tested on Qwen-7B with suitable prompts, GPT-5.4, and Qwen-72B, showing context sensitivity in aspectual classification and human-comparable performance under proper conditions
Industry Insight
- Benchmark-driven claims about model limitations should be treated as provisional until independently validated with native-speaker annotation and controlled minimal-pair evaluations
- Prompting interventions may produce superficial label shifts without genuine semantic improvement—researchers should measure intermediate reasoning steps, not just final predictions
- The gap between benchmark failure and model failure highlights an opportunity: well-designed evaluations can reveal that models possess capabilities previously thought absent, suggesting current benchmarks may systematically underestimate LLM semantic competence
Disclaimer: The above content is generated by AI and is for reference only.