OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings
OpenAI claims GPT-5.6 Sol achieves 38.3% on ARC-AGI-3 using custom API settings ("Retained Reasoning" and "Compaction"), surpassing Claude Opus 5's 30.2%. The official ARC-AGI-3 benchmark uses a standardized test harness without provider-specific features, where GPT-5.6 Sol scored only 7.8%, highlighting the impact of technical setup on performance. ARC Prize co-founder François Chollet acknowledged that using different API settings creates a potential parity issue but accepted it as long as set
Analysis
TL;DR
- OpenAI claims GPT-5.6 Sol achieves 38.3% on ARC-AGI-3 using custom API settings ("Retained Reasoning" and "Compaction"), surpassing Claude Opus 5's 30.2%.
- The official ARC-AGI-3 benchmark uses a standardized test harness without provider-specific features, where GPT-5.6 Sol scored only 7.8%, highlighting the impact of technical setup on performance.
- ARC Prize co-founder François Chollet acknowledged that using different API settings creates a potential parity issue but accepted it as long as settings and costs are transparently reported.
- The debate centers on whether benchmarks should measure pure model capability or include the influence of surrounding infrastructure and API configurations.
Why It Matters
This article underscores a critical tension in AI evaluation: how much of a model’s performance is attributable to its architecture versus the technical environment in which it operates. For researchers and practitioners, it raises important questions about benchmark fairness, reproducibility, and the need for standardized testing protocols that account for real-world deployment conditions. As models become increasingly integrated with specialized APIs and context management tools, understanding these nuances becomes essential for accurate comparison and progress tracking.
Technical Details
- OpenAI used its proprietary Responses API with two advanced features: “Retained Reasoning,” which preserves the model’s chain-of-thought across multiple steps, and “Compaction,” which summarizes prior context rather than truncating it—both designed to enhance long-context reasoning.
- In contrast, the official ARC-AGI-3 benchmark employs a generic test harness that does not support such features, resulting in a significantly lower score (7.8%) for GPT-5.6 Sol when evaluated under standard conditions.
- Claude Opus 5 was tested using Anthropic’s own API, which inherently supports similar advanced reasoning mechanisms, giving it an advantage over earlier evaluations of GPT-5.6 Sol that lacked comparable tooling.
- The discrepancy between scores (38.3% vs. 7.8%) demonstrates that model performance on complex reasoning tasks can be dramatically influenced by the availability and integration of contextual memory and reasoning persistence mechanisms within the inference pipeline.
Industry Insight
The incident highlights the growing importance of evaluating AI systems not just in isolation but within their operational ecosystems. Providers must clearly document and disclose all non-standard configurations used during benchmarking to ensure fair comparisons. Future evaluation frameworks may need to incorporate modular components like retained reasoning and dynamic compaction as part of the baseline assessment, reflecting real-world usage patterns. Additionally, this case suggests that competitive advantage in AI may increasingly stem from innovations in inference infrastructure—not just model size or training data quality—pushing companies to invest heavily in optimized serving layers and context-aware architectures.
Disclaimer: The above content is generated by AI and is for reference only.