Sakana claims its AI model router Fugu Ultra v1.1 now beats Fable 5 without even including it in the pool
Sakana AI released Fugu Ultra v1.1, an AI model router that distributes queries across a pool of top-tier public models. The update claims performance gains of up to 7.9 points over version 1.0, with significant improvements on ProgramBench and TerminalBench 2.1. Fugu Ultra v1.1 reportedly outperforms Anthropic's Fable 5, despite Fable 5 not being included in the router's selection pool. All performance metrics are self-reported by Sakana AI, with no independent verification currently available.
Analysis
TL;DR
- Sakana AI released Fugu Ultra v1.1, an AI model router that distributes queries across a pool of top-tier public models.
- The update claims performance gains of up to 7.9 points over version 1.0, with significant improvements on ProgramBench and TerminalBench 2.1.
- Fugu Ultra v1.1 reportedly outperforms Anthropic's Fable 5, despite Fable 5 not being included in the router's selection pool.
- All performance metrics are self-reported by Sakana AI, with no independent verification currently available.
- Pricing remains unchanged at $5 per million input tokens and $30 per million output tokens.
Why It Matters
This development highlights the growing competitive pressure on standalone large language models from routing architectures that aggregate multiple specialized models. For practitioners, it underscores the importance of evaluating whether multi-model routing offers superior cost-performance trade-offs compared to single-state-of-the-art models like Fable 5. The lack of independent verification also serves as a critical reminder for researchers and engineers to validate vendor claims before integrating such systems into production pipelines.
Technical Details
- Architecture: Fugu Ultra v1.1 functions as a router that dynamically distributes user queries across a curated pool of publicly available, top-tier AI models.
- Performance Metrics: The update reports performance increases of up to 7.9 points compared to v1.0, specifically noting major jumps on ProgramBench and TerminalBench 2.1.
- Benchmarking: The router claims superiority over Anthropic's Fable 5 across most benchmarks, achieving this without Fable 5 being part of its underlying model pool.
- Integration & Infrastructure: A new Claude Code-compatible endpoint allows direct terminal access, building on previous availability via OpenRouter and Vercel.
- Operational Constraints: Adding new models to the pool requires approximately two weeks of training and evaluation, and the service currently excludes the EU and EEA due to regulatory concerns.
Industry Insight
The success of model routers like Fugu suggests a strategic shift toward agentic orchestration layers that leverage diverse model strengths rather than relying on monolithic models. Developers should consider implementing routing mechanisms to optimize for specific tasks (e.g., coding vs. reasoning) while managing costs, provided they can verify the claimed performance gains. Additionally, the regulatory exclusion of the EU market indicates that compliance with GDPR and emerging AI regulations will remain a significant barrier to global deployment for many AI infrastructure providers.
Disclaimer: The above content is generated by AI and is for reference only.