AI News AI资讯 7h ago Updated 2h ago 更新于 2小时前 49

Anthropic claims its new Claude Opus 5 delivers near-Fable 5 performance at half the token price Anthropic声称其新Claude Opus 5以一半的Token价格实现接近Fable 5的性能

Anthropic released Claude Opus 5, positioning it as a cost-effective flagship that delivers near-Fable 5 performance at half the token price ($5/$25 vs $10/$50). The model achieves state-of-the-art results in agentic coding and novel problem-solving, notably scoring 30.2% on ARC-AGI-3, nearly four times higher than GPT-5.6 Sol. Opus 5 demonstrates advanced self-correction capabilities, including building its own computer vision tools to solve tasks without direct visual input. While leading in c Anthropic发布Claude Opus 5,定位为Fable 5的高性价比替代方案,在保持100万token上下文窗口的同时,将输入价格降至$5/MTok(Fable 5为$10/MTok)。 在ARC-AGI-3基准测试中,Opus 5得分30.2%,远超GPT-5.6 Sol的7.8%,展现极强的新颖问题解决能力;同时在代理编码和知识工作领域取得领先。 模型具备自我迭代检查及通过代码构建自定义工具的能力,例如在无直接视觉输入情况下编写计算机视觉管道完成3D建模任务。 引入五种“努力程度”设置以平衡性能与Token消耗,Anthropic建议日常使用“low”或“medium”模式以优

75
Hot 热度
65
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Anthropic released Claude Opus 5, positioning it as a cost-effective flagship that delivers near-Fable 5 performance at half the token price ($5/$25 vs $10/$50).
  • The model achieves state-of-the-art results in agentic coding and novel problem-solving, notably scoring 30.2% on ARC-AGI-3, nearly four times higher than GPT-5.6 Sol.
  • Opus 5 demonstrates advanced self-correction capabilities, including building its own computer vision tools to solve tasks without direct visual input.
  • While leading in coding and general knowledge work, Opus 5 trails competitors in specific domains like cybersecurity exploitation and certain health/legal benchmarks.
  • The release reflects intense pricing pressure from OpenAI and Chinese competitors, with Anthropic optimizing for token efficiency rather than just lowering base rates.

Why It Matters

This release signals a strategic shift in the frontier AI market where performance parity is being achieved through cost optimization and specialized agentic capabilities rather than raw scale alone. For practitioners, the significant improvement in autonomous tool-building and iterative self-correction suggests that Opus 5 is better suited for complex, multi-step workflows compared to previous iterations or some competitors. The pricing structure also highlights the importance of evaluating total task cost (token efficiency) over simple per-token rates when selecting models for production environments.

Technical Details

  • Pricing and Architecture: Opus 5 maintains the 1 million-token context window but halves the input/output token costs compared to Fable 5 ($5M input / $25M output). A "Fast Mode" offers 2.5x speed at double the price.
  • Benchmark Performance: On Frontier-Bench v0.1, Opus 5 scored 43.3% in agentic terminal coding, outperforming GPT-5.6 Sol (34.4%) and Fable 5 (33.7%). It leads in knowledge work with an Elo score of 1,861 on GDPval-AAv2.
  • ARC-AGI-3 Results: The model scored 30.2% on this novel problem-solving benchmark, significantly surpassing GPT-5.6 Sol (7.8%) and Opus 4.8 (1.5%), indicating strong generalization beyond memorized patterns.
  • Agentic Capabilities: Opus 5 can generate its own code to create tools (e.g., a computer vision pipeline) when standard interfaces are unavailable, successfully solving tasks that other models failed after multiple attempts.
  • Effort Settings: Users can adjust effort levels (low to max). Anthropic recommends "low" or "medium" for most tasks due to better token efficiency, though "xhigh" is advised for coding. Notably, "max" effort sometimes yields lower scores than "xhigh" on specific benchmarks despite higher costs.

Industry Insight

  • Cost-Performance Trade-offs: Organizations should prioritize token efficiency metrics over base pricing. As seen with previous Opus versions, higher base rates do not guarantee lower total costs if the model uses more tokens to complete tasks.
  • Autonomous Agent Reliability: The ability to build custom tools and self-correct makes Opus 5 a stronger candidate for autonomous agent architectures, reducing the need for rigid, pre-defined toolsets in complex engineering workflows.
  • Market Consolidation: With Opus 5 closing the gap with Fable 5 at half the price, Anthropic is effectively creating a tiered product strategy where Opus 5 serves as the default high-performance option for most users, potentially reducing reliance on the more expensive Fable 5 for general tasks.

TL;DR

  • Anthropic发布Claude Opus 5,定位为Fable 5的高性价比替代方案,在保持100万token上下文窗口的同时,将输入价格降至$5/MTok(Fable 5为$10/MTok)。
  • 在ARC-AGI-3基准测试中,Opus 5得分30.2%,远超GPT-5.6 Sol的7.8%,展现极强的新颖问题解决能力;同时在代理编码和知识工作领域取得领先。
  • 模型具备自我迭代检查及通过代码构建自定义工具的能力,例如在无直接视觉输入情况下编写计算机视觉管道完成3D建模任务。
  • 引入五种“努力程度”设置以平衡性能与Token消耗,Anthropic建议日常使用“low”或“medium”模式以优化成本,仅在复杂编码任务中使用“xhigh”。
  • 安全策略调整使得Opus 5在允许源代码漏洞研究的同时,相比Fable 5减少了85%的安全拦截触发率,但在二进制扫描和漏洞利用方面仍受限。

为什么值得看

这篇文章揭示了当前大模型竞争从单纯追求极致性能向“性能-成本”平衡点转移的关键趋势,Opus 5通过降低价格并提升效率,直接回应了市场对于高昂API费用的痛点。对于AI从业者和企业开发者而言,理解不同努力程度设置对Token效率和最终成本的影响,是优化应用架构和降低运营支出的核心依据。此外,模型在自主构建工具和解决未见过的复杂问题上的突破,预示着Agent类应用在自动化程度上的新高度。

技术解析

  • 定价与架构策略:Opus 5维持了100万token上下文窗口,输入价格为$5/MTok,输出$25/MTok,仅为Fable 5的一半。新增Fast Mode速度提升2.5倍但价格翻倍。Anthropic强调Token效率(Task Cost)比基础费率更重要,指出高费率不等于高实际成本。
  • 基准测试表现:在Frontier-Bench v0.1代理终端编码任务中得分43.3%,领先Fable 5 (33.7%)和GPT-5.6 Sol (34.4%)。在知识工作基准GDPval-AAv2中Elo得分1,861。但在DeepSWE v1.1上以68.8%落后于GPT-5.6 Sol (72.7%)。
  • ARC-AGI-3异常突破:该基准测试衡量无记忆模式的新颖问题解决能力,Opus 5得分30.2%,是第二名GPT-5.6 Sol (7.8%)的近四倍,前代Opus 4.8仅得1.5%。这一结果被视为重大惊喜,但其在实际应用中的泛化性尚待观察。
  • 自主工具构建与迭代:Opus 5展现出强大的自我修正和工具生成能力。案例显示,模型在无视觉输入限制下,自行编写计算机视觉管道提取几何信息并重建3D模型;还能深入挖掘开源包管理器的边缘Case Bug,而非仅修复表面症状。
  • 安全与努力程度机制:提供Low/Medium/High/XHigh/Max五种努力设置。Anthropic推荐常规任务使用Low/Medium以节省Token和延迟,尽管Max设置下部分基准测试得分反而略降且成本更高。安全方面,放宽了源代码漏洞研究限制,但保留了对二进制扫描和渗透测试的封锁。

行业启示

  • 成本效率成为核心竞争力:随着模型能力趋同,单位任务的Token消耗效率(Efficiency)将成为比绝对峰值性能更关键的差异化指标。企业应重新评估模型选型,优先选择在高性价比设置下表现稳定的模型,而非盲目追求最高配置。
  • Agent能力的实质性跃迁:Opus 5能够自主编写代码来解决自身缺乏的功能模块(如视觉处理),标志着LLM从“被动执行指令”向“主动构建解决方案”的Agent形态演进。这要求开发者在设计工作流时,预留更多的自主决策和工具调用空间。
  • 安全护栏的动态平衡:Anthropic通过减少85%的安全拦截来改善开发者体验,表明行业正在寻找安全性与可用性之间的新平衡点。未来模型的安全策略将更加精细化,区分“恶意攻击”与“防御性研究”,这对合规性和伦理审查提出了新的技术要求。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Claude Claude LLM 大模型 Benchmark 基准测试 Product Launch 产品发布 Code Generation 代码生成