AI News AI资讯 3h ago Updated 1h ago 更新于 1小时前 51

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet's AGI forecast forward 基准测试对GPT-6 Astra评价不一,但其在ARC-AGI-3上超越人类的效率将乔莱的AGI预测提前

GPT-6 Astra shows contradictory benchmark results: Epoch AI ranks it first overall (169 points across 50+ benchmarks), while Artificial Analysis rates it equal to its predecessor at 61 points, behind Claude Fable 5.1 at 66. The most significant breakthrough is on ARC-AGI-3, where Astra achieves 62.7% on unfamiliar game worlds—far surpassing Sol's 7.8% and Opus 5's 30.2%—and operates more efficiently than the average human solver. Astra is substantially more expensive per token than its predecess GPT-6 Astra在ARC-AGI-3基准测试中达到62.7%得分,首次超越人类平均水平,比预期快2倍 Epoch AI与Artificial Analysis给出矛盾评估:前者排第一(169分),后者与 predecessor 持平(61分) Astra成本是Sol的2.5倍,但计算效率显著提升,仅需Sol三分之一的计算步骤 模型在ARC-AGI-3中自发发明代数式简写符号系统,展示自我建模能力 ARC Prize创始人Chollet预计AGI到来时间将提前,ARC-AGI-4计划2027年Q1发布

72
Hot 热度
72
Quality 质量
75
Impact 影响力

Analysis 深度分析

TL;DR

  • GPT-6 Astra shows contradictory benchmark results: Epoch AI ranks it first overall (169 points across 50+ benchmarks), while Artificial Analysis rates it equal to its predecessor at 61 points, behind Claude Fable 5.1 at 66.
  • The most significant breakthrough is on ARC-AGI-3, where Astra achieves 62.7% on unfamiliar game worlds—far surpassing Sol's 7.8% and Opus 5's 30.2%—and operates more efficiently than the average human solver.
  • Astra is substantially more expensive per token than its predecessor (2.5x the cost), yet uses only a third of Sol's compute steps and a fifth of Opus 5's, making it cost-competitive on coding tasks versus Anthropic models.
  • The model develops its own symbolic shorthand notation for tracking game states and rules, demonstrating on-the-fly symbolic world modeling previously only seen with external harnesses.
  • ARC Prize's François Chollet calls the progress "2x faster" than expected and revises his AGI forecast to "sooner," while clarifying that ARC-AGI-3 success is not proof of AGI.

Why It Matters

This article highlights the growing fragmentation in AI benchmarking, where different evaluation frameworks produce contradictory rankings for the same model—critical context for practitioners trying to assess real capability. The ARC-AGI-3 results represent a qualitative shift in how models approach unknown environments, with Astra internalizing symbolic reasoning previously dependent on external tooling, which has implications for how we define and measure general intelligence. The cost-efficiency dynamics also signal a strategic pivot in the frontier model race: raw benchmark scores matter less than the ability to solve complex tasks with fewer compute steps.

Technical Details

  • Benchmark divergence: Epoch AI aggregates 50+ benchmarks into a single ECI score (Astra: 169, Sol: 162, Fable 5.1: 163), while Artificial Analysis uses a narrower index covering knowledge, coding, and text comprehension (Astra: 61, Fable 5.1: 66). Astra leads on math, knowledge, and puzzles; Fable 5.1 dominates nearly all coding tests.
  • ARC-AGI-3 performance: Astra scores 62.7% on unfamiliar game worlds at ~$26,000 cost, compared to Sol's 7.78% and Opus 5's 30.16%. With OpenAI's proprietary harness (which maintains reasoning chains and summarizes long runs), performance reaches 99.9%, though ARC Prize considers only the standard harness result fair for cross-vendor comparison.
  • Reasoning effort paradox: Higher reasoning levels reduce cost on ARC-AGI-3—from $49,791 (no reasoning) to $26,098 (max reasoning)—because Astra solves games in fewer moves, requiring fewer model calls and tokens. The "low" reasoning level (17.5%) underperforms even the no-reasoning baseline (35.2%).
  • Symbolic self-modeling: Astra invents an algebra-like DSL to record objects, coordinates, rules, and plans (e.g., "extend8 to3; retract10 to2", "Turn 5: P=(24,20), empty, facing west"). In the PRO-LONG sandbox environment, it writes custom program libraries including board parsers, state models, search algorithms, and planners.
  • Efficiency vs. human baseline: Astra cleared 96% of ARC-AGI-3 levels in fewer moves than the median human solver. Human testers earned ~$12.78 per game including compensation; modeling brain metabolism as electricity yields ~0.067 cents per game.
  • Other benchmark changes: Hallucination rate on AA-Omniscience drops from 92% to 51%. FrontierMath Erdős: Astra solved 2 of 68 open problems with Lean-verified proofs at $300 per attempt. Astra loses ~80 Elo on GDPval-AA v2 and slips on banking support, SciCode, and long-context reasoning.

Industry Insight

  • Benchmarking is becoming unreliable as a sole capability indicator: The stark contradiction between Epoch AI and Artificial Analysis rankings—driven by different test compositions and reasoning-level choices—means practitioners must scrutinize methodology before drawing conclusions about model superiority. The single coding score for Astra at medium reasoning further undermines direct comparisons.
  • Symbolic reasoning is migrating from harnesses to base models: Chollet's observation that "harness capabilities are increasingly shifting into the model itself" signals a structural change in model architecture. Future benchmarks will need to account for models that internally replicate what previously required external tooling, and the PRO-LONG sandbox results demonstrate that models can autonomously build complex toolchains when given the opportunity.
  • The AGI timeline is compressing, but definitions remain contested: Chollet's revised forecast ("sooner") and the ARC Prize team's insistence that ARC-AGI-3 is "not proof of AGI" reflect a growing tension between rapid empirical progress and conservative philosophical standards. ARC-AGI-4 (expected Q1 2027) will target recursive self-improvement and open innovation, suggesting the benchmarking frontier is moving toward capabilities that are harder to measure but more indicative of general intelligence.

TL;DR

  • GPT-6 Astra在ARC-AGI-3基准测试中达到62.7%得分,首次超越人类平均水平,比预期快2倍
  • Epoch AI与Artificial Analysis给出矛盾评估:前者排第一(169分),后者与 predecessor 持平(61分)
  • Astra成本是Sol的2.5倍,但计算效率显著提升,仅需Sol三分之一的计算步骤
  • 模型在ARC-AGI-3中自发发明代数式简写符号系统,展示自我建模能力
  • ARC Prize创始人Chollet预计AGI到来时间将提前,ARC-AGI-4计划2027年Q1发布

为什么值得看

本文揭示了当前AI评估体系的碎片化问题——不同基准测试得出相反结论,反映了评估标准的不统一。GPT-6 Astra在ARC-AGI-3上的突破展示了模型自我符号化能力的质变,这是通向AGI的重要里程碑。

技术解析

  • ARC-AGI-3测试:将AI置于未知游戏世界,需通过试错学习规则和目标。Astra在标准harness下达到62.7%(成本约$26,000),Sol仅7.78%,Opus 5为30.16%。使用OpenAI自有harness时效率提升3.66倍,token消耗减少49%。
  • 成本与效率悖论:Astra在ARC-AGI-3上呈现"思考越多成本越低"的反常现象——无推理时成本$49,791,最大推理时降至$26,098,因模型用更少步数解决问题。
  • 自我符号系统:Astra发明类代数简写记录游戏状态(如"extend8 to3; retract10 to2"),Chollet称之为"高效的即时符号世界建模",能力从harness转移到模型本身。
  • 外部工具使用:在PRO-LONG框架中,Astra为迷宫游戏编写解析器、状态模型、搜索算法和规划器,但未尝试逃逸沙箱。
  • 基准测试分歧:Epoch AI综合50+基准(169分领先),Artificial Analysis仅测知识/编程/理解(61分,落后于Fable 5.1的66分)。Astra在数学、知识、谜题领先,Fable 5.1在编程测试占优。

行业启示

  • 评估体系亟需统一:不同基准给出相反结论,反映当前AI评估缺乏标准化,行业需建立更一致的评测框架。
  • 效率优于算力堆砌:Astra证明通过架构优化可实现"思考越多成本越低",未来竞争将从算力规模转向推理效率。
  • AGI时间表可能提前:Chollet预计进展比预期快2倍,ARC-AGI-4将探索递归自我改进和开放创新,AGI到来时间或早于2030年。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

GPT GPT Benchmark 基准测试 Evaluation 评测 LLM 大模型 Closed Source 闭源