AI News AI资讯 7h ago Updated 1h ago 更新于 1小时前 49

GPT-6 Astra GPT-6 Astra

GPT-6 Astra is OpenAI's latest flagship model, rolling out to ChatGPT Plus/Pro/Business/Enterprise users and via API at $10/million input and $50/million output tokens It achieves 99.9% on the ARC-AGI 3 benchmark using OpenAI's custom Provider Adapter harness, though the default harness scored only 62.7% Astra dominates security tasks with 100% on ExploitBench, 42.4% on ExploitGym, and 99.2% on SRE-Bench binary reverse engineering It excels at long-context processing, scoring 100% on the eight-n GPT-6 Astra是OpenAI对标Claude Fable的新一代模型,API定价$10/million输入/$50/million输出,面向Plus/Pro/Business/Enterprise用户及API开放 在ARC-AGI 3基准测试中达到99.9%得分(使用自定义Provider Adapter harness,花费$19K),但默认harness仅62.7%(花费$26K) 安全任务表现卓越:ExploitBench满分100%,SRE-Bench二进制逆向工程99.2%,远超GPT-5.6 Sol的78.5%和68.7% 长上下文处理突破:256K-512K tokens

72
Hot 热度
62
Quality 质量
75
Impact 影响力

Analysis 深度分析

TL;DR

  • GPT-6 Astra is OpenAI's latest flagship model, rolling out to ChatGPT Plus/Pro/Business/Enterprise users and via API at $10/million input and $50/million output tokens
  • It achieves 99.9% on the ARC-AGI 3 benchmark using OpenAI's custom Provider Adapter harness, though the default harness scored only 62.7%
  • Astra dominates security tasks with 100% on ExploitBench, 42.4% on ExploitGym, and 99.2% on SRE-Bench binary reverse engineering
  • It excels at long-context processing, scoring 100% on the eight-needle benchmark at 256K–512K tokens and 96.3% at 512K–1M tokens
  • Despite these strengths, Astra trails Claude Fable 5.1 and Meta's Muse Spark 1.3 on the Artificial Analysis Intelligence Index, though it leads on cost efficiency for coding agent tasks

Why It Matters

GPT-6 Astra represents a significant competitive move by OpenAI to counter Claude Fable and Meta's Muse in the frontier model race, particularly in specialized domains like security and long-context reasoning. Its pricing strategy—matching Claude Fable 5/5.1—signals an aggressive push for market share in the API space, while its near-perfect ARC-AGI 3 score (albeit with a custom harness) suggests meaningful progress on general intelligence benchmarking.

Technical Details

  • ARC-AGI 3 Performance: Achieved 99.9% using OpenAI's custom "Provider Adapter harness," which preserves opaque reasoning state between requests and uses compaction for longer conversations; the default harness yielded only 62.7%, highlighting the impact of evaluation methodology
  • Security Capabilities: 100% on ExploitBench (vs. GPT-5.6 Sol's 78.5%), 42.4% on ExploitGym (vs. Sol's 30.3%), and 99.2% within four attempts on SRE-Bench binary reverse engineering (vs. Sol's 68.7%)
  • Long Context: 100% accuracy on OpenAI's eight-needle benchmark at 256K–512K tokens and 96.3% at 512K–1M tokens, addressing a persistent challenge in the field
  • Pricing: API priced at $10/million input tokens and $50/million output tokens, matching Claude Fable 5 and 5.1
  • Intelligence Index Positioning: Scores 61 on Artificial Analysis Intelligence Index, matching GPT-5.6 Sol but trailing Claude Fable 5.1 (66) and Meta's Muse Spark 1.3

Industry Insight

  • The dramatic gap between Astra's custom harness (99.9%) and default harness (62.7%) on ARC-AGI 3 underscores the importance of standardized evaluation methodologies; practitioners should be cautious about benchmark claims that depend on proprietary evaluation infrastructure
  • Astra's cost efficiency in coding agent tasks—less than half the cost of Claude Fable 5 for the same score—makes it a compelling option for production workloads that prioritize economics over raw intelligence, particularly for high-volume coding automation
  • OpenAI's focus on security and long-context capabilities appears strategically targeted at enterprise adoption, suggesting the company is positioning Astra as a drop-in replacement for organizations with compliance and large-document processing needs

TL;DR

  • GPT-6 Astra是OpenAI对标Claude Fable的新一代模型,API定价$10/million输入/$50/million输出,面向Plus/Pro/Business/Enterprise用户及API开放
  • 在ARC-AGI 3基准测试中达到99.9%得分(使用自定义Provider Adapter harness,花费$19K),但默认harness仅62.7%(花费$26K)
  • 安全任务表现卓越:ExploitBench满分100%,SRE-Bench二进制逆向工程99.2%,远超GPT-5.6 Sol的78.5%和68.7%
  • 长上下文处理突破:256K-512K tokens达100%,512K-1M tokens达96.3%,解决长期技术挑战
  • 综合智能指数61分,低于Claude Fable 5.1(66分)和Meta Muse Spark 1.3,但在Coding Agent Index成本效率上领先

为什么值得看

GPT-6 Astra的发布标志着OpenAI在推理能力和安全领域的重大突破,特别是ARC-AGI基准测试的接近满分表现引发行业关注。其独特的Provider Adapter harness技术为长上下文处理提供了新思路,同时成本效率优势使其在编码代理场景极具竞争力。

技术解析

  • Provider Adapter harness:OpenAI自定义测试框架,保留不透明推理状态并在请求间复用,通过压缩技术处理长对话,使ARC-AGI 3得分从62.7%提升至99.9%,成本降低27%
  • 安全任务表现:ExploitBench满分100%(GPT-5.6 Sol为78.5%),ExploitGym 42.4%(Sol为30.3%),SRE-Bench二进制逆向工程99.2%(Sol为68.7%),反映OpenAI对安全能力的重视
  • 长上下文处理:在256K-512K tokens范围达到100%准确率,512K-1M tokens达96.3%,标志着OpenAI可能已解决长上下文处理的核心挑战
  • 基准测试对比:Artificial Analysis Intelligence Index得61分,落后Claude Fable 5.1(66分)和Meta Muse Spark 1.3;但Coding Agent Index成本效率领先,同等分数下成本不到Claude Fable 5的一半

行业启示

  • 推理能力成为新战场:ARC-AGI 3基准测试的高分竞争表明,AI行业正从通用智能转向复杂推理能力比拼,OpenAI通过自定义测试框架展示技术实力,但也引发对基准测试公平性的讨论
  • 安全能力差异化:在Hugging Face安全事件背景下,Astra的安全任务表现凸显安全能力将成为企业级AI服务的关键差异化因素,建议开发者优先评估模型的安全防护能力
  • 成本效率决定落地速度:Astra在Coding Agent场景的成本优势(同等分数成本减半)表明,未来AI竞争不仅是性能比拼,更是单位任务成本的竞争,企业应关注特定场景的成本效率而非单纯追求最高智能指数

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

GPT GPT Claude Claude LLM 大模型 Product Launch 产品发布 Pricing Pricing