Anthropic claims its new Claude Opus 5 delivers near-Fable 5 performance at half the token price
Anthropic released Claude Opus 5, positioning it as a cost-effective flagship that delivers near-Fable 5 performance at half the token price ($5/$25 vs $10/$50). The model achieves state-of-the-art results in agentic coding and novel problem-solving, notably scoring 30.2% on ARC-AGI-3, nearly four times higher than GPT-5.6 Sol. Opus 5 demonstrates advanced self-correction capabilities, including building its own computer vision tools to solve tasks without direct visual input. While leading in c
Analysis
TL;DR
- Anthropic released Claude Opus 5, positioning it as a cost-effective flagship that delivers near-Fable 5 performance at half the token price ($5/$25 vs $10/$50).
- The model achieves state-of-the-art results in agentic coding and novel problem-solving, notably scoring 30.2% on ARC-AGI-3, nearly four times higher than GPT-5.6 Sol.
- Opus 5 demonstrates advanced self-correction capabilities, including building its own computer vision tools to solve tasks without direct visual input.
- While leading in coding and general knowledge work, Opus 5 trails competitors in specific domains like cybersecurity exploitation and certain health/legal benchmarks.
- The release reflects intense pricing pressure from OpenAI and Chinese competitors, with Anthropic optimizing for token efficiency rather than just lowering base rates.
Why It Matters
This release signals a strategic shift in the frontier AI market where performance parity is being achieved through cost optimization and specialized agentic capabilities rather than raw scale alone. For practitioners, the significant improvement in autonomous tool-building and iterative self-correction suggests that Opus 5 is better suited for complex, multi-step workflows compared to previous iterations or some competitors. The pricing structure also highlights the importance of evaluating total task cost (token efficiency) over simple per-token rates when selecting models for production environments.
Technical Details
- Pricing and Architecture: Opus 5 maintains the 1 million-token context window but halves the input/output token costs compared to Fable 5 ($5M input / $25M output). A "Fast Mode" offers 2.5x speed at double the price.
- Benchmark Performance: On Frontier-Bench v0.1, Opus 5 scored 43.3% in agentic terminal coding, outperforming GPT-5.6 Sol (34.4%) and Fable 5 (33.7%). It leads in knowledge work with an Elo score of 1,861 on GDPval-AAv2.
- ARC-AGI-3 Results: The model scored 30.2% on this novel problem-solving benchmark, significantly surpassing GPT-5.6 Sol (7.8%) and Opus 4.8 (1.5%), indicating strong generalization beyond memorized patterns.
- Agentic Capabilities: Opus 5 can generate its own code to create tools (e.g., a computer vision pipeline) when standard interfaces are unavailable, successfully solving tasks that other models failed after multiple attempts.
- Effort Settings: Users can adjust effort levels (low to max). Anthropic recommends "low" or "medium" for most tasks due to better token efficiency, though "xhigh" is advised for coding. Notably, "max" effort sometimes yields lower scores than "xhigh" on specific benchmarks despite higher costs.
Industry Insight
- Cost-Performance Trade-offs: Organizations should prioritize token efficiency metrics over base pricing. As seen with previous Opus versions, higher base rates do not guarantee lower total costs if the model uses more tokens to complete tasks.
- Autonomous Agent Reliability: The ability to build custom tools and self-correct makes Opus 5 a stronger candidate for autonomous agent architectures, reducing the need for rigid, pre-defined toolsets in complex engineering workflows.
- Market Consolidation: With Opus 5 closing the gap with Fable 5 at half the price, Anthropic is effectively creating a tiered product strategy where Opus 5 serves as the default high-performance option for most users, potentially reducing reliance on the more expensive Fable 5 for general tasks.
Disclaimer: The above content is generated by AI and is for reference only.