AI News AI资讯 4h ago Updated 1h ago 更新于 1小时前 49

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings OpenAI声称GPT-5.6 Sol在ARC-AGI-3上击败Opus 5,使用最新API和两个附加设置

OpenAI claims GPT-5.6 Sol achieves 38.3% on ARC-AGI-3 using custom API settings ("Retained Reasoning" and "Compaction"), surpassing Claude Opus 5's 30.2%. The official ARC-AGI-3 benchmark uses a standardized test harness without provider-specific features, where GPT-5.6 Sol scored only 7.8%, highlighting the impact of technical setup on performance. ARC Prize co-founder François Chollet acknowledged that using different API settings creates a potential parity issue but accepted it as long as set OpenAI声称GPT-5.6 Sol在ARC-AGI-3基准测试中得分38.3%,超越了Anthropic的Claude Opus 5(30.2%)。 OpenAI使用了其自定义的Responses API,启用了“Retained Reasoning”和“Compaction”两项设置,以提升模型表现。 ARC Prize联合创始人François Chollet指出,OpenAI使用的API设置并非通用,可能影响比较的公平性。 OpenAI认为基准测试不仅衡量模型本身,还涉及技术设置,而ARC-AGI-3旨在测试纯模型性能。 争议点在于ARC Prize是否使用了较旧的“OpenAI-s

75
Hot 热度
65
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • OpenAI claims GPT-5.6 Sol achieves 38.3% on ARC-AGI-3 using custom API settings ("Retained Reasoning" and "Compaction"), surpassing Claude Opus 5's 30.2%.
  • The official ARC-AGI-3 benchmark uses a standardized test harness without provider-specific features, where GPT-5.6 Sol scored only 7.8%, highlighting the impact of technical setup on performance.
  • ARC Prize co-founder François Chollet acknowledged that using different API settings creates a potential parity issue but accepted it as long as settings and costs are transparently reported.
  • The debate centers on whether benchmarks should measure pure model capability or include the influence of surrounding infrastructure and API configurations.

Why It Matters

This article underscores a critical tension in AI evaluation: how much of a model’s performance is attributable to its architecture versus the technical environment in which it operates. For researchers and practitioners, it raises important questions about benchmark fairness, reproducibility, and the need for standardized testing protocols that account for real-world deployment conditions. As models become increasingly integrated with specialized APIs and context management tools, understanding these nuances becomes essential for accurate comparison and progress tracking.

Technical Details

  • OpenAI used its proprietary Responses API with two advanced features: “Retained Reasoning,” which preserves the model’s chain-of-thought across multiple steps, and “Compaction,” which summarizes prior context rather than truncating it—both designed to enhance long-context reasoning.
  • In contrast, the official ARC-AGI-3 benchmark employs a generic test harness that does not support such features, resulting in a significantly lower score (7.8%) for GPT-5.6 Sol when evaluated under standard conditions.
  • Claude Opus 5 was tested using Anthropic’s own API, which inherently supports similar advanced reasoning mechanisms, giving it an advantage over earlier evaluations of GPT-5.6 Sol that lacked comparable tooling.
  • The discrepancy between scores (38.3% vs. 7.8%) demonstrates that model performance on complex reasoning tasks can be dramatically influenced by the availability and integration of contextual memory and reasoning persistence mechanisms within the inference pipeline.

Industry Insight

The incident highlights the growing importance of evaluating AI systems not just in isolation but within their operational ecosystems. Providers must clearly document and disclose all non-standard configurations used during benchmarking to ensure fair comparisons. Future evaluation frameworks may need to incorporate modular components like retained reasoning and dynamic compaction as part of the baseline assessment, reflecting real-world usage patterns. Additionally, this case suggests that competitive advantage in AI may increasingly stem from innovations in inference infrastructure—not just model size or training data quality—pushing companies to invest heavily in optimized serving layers and context-aware architectures.

TL;DR

  • OpenAI声称GPT-5.6 Sol在ARC-AGI-3基准测试中得分38.3%,超越了Anthropic的Claude Opus 5(30.2%)。
  • OpenAI使用了其自定义的Responses API,启用了“Retained Reasoning”和“Compaction”两项设置,以提升模型表现。
  • ARC Prize联合创始人François Chollet指出,OpenAI使用的API设置并非通用,可能影响比较的公平性。
  • OpenAI认为基准测试不仅衡量模型本身,还涉及技术设置,而ARC-AGI-3旨在测试纯模型性能。
  • 争议点在于ARC Prize是否使用了较旧的“OpenAI-style completions API”,导致对OpenAI不公平。

为什么值得看

这篇文章揭示了当前大模型基准测试中的关键争议,即不同技术设置对模型性能的影响。对于AI从业者而言,理解这些细节有助于更准确地评估模型能力,并推动行业建立更公平的评测标准。

技术解析

  • 模型与设置:OpenAI使用GPT-5.6 Sol,通过其Responses API运行,启用“Retained Reasoning”和“Compaction”两项设置,分别用于保留推理链和压缩上下文。
  • 基准测试对比:在官方测试环境中,GPT-5.6 Sol仅得7.8%,而使用自定义设置后得分提升至38.3%。相比之下,Claude Opus 5在官方设置下得分为30.2%。
  • 争议焦点:ARC Prize认为官方测试应使用标准化设置以确保公平性,而OpenAI则强调技术设置对模型表现的重要性。
  • API差异:争议的核心在于ARC Prize是否使用了较旧的“OpenAI-style completions API”,该API缺乏Claude API提供的某些功能,可能导致对OpenAI的不利影响。

行业启示

  • 评测标准的透明化:行业需要进一步明确基准测试的技术设置要求,确保所有参与者在相同条件下进行比较,避免不公平竞争。
  • 技术设置的优化潜力:模型性能的显著提升表明,技术设置(如推理链保留和上下文管理)对大模型的表现有重大影响,未来研究应更多关注此类优化。
  • 多方协作的重要性:基准测试的公正性依赖于开发者和评测机构的紧密合作,共同制定合理的测试框架和规则,以促进技术的健康发展。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

GPT GPT Evaluation 评测 Benchmark 基准测试 LLM 大模型