Claude Fable 5.1 made me a really nice animated pelican
Anthropic released Claude Fable 5.1, claiming it sets a new standard for coding, knowledge work, and long-running problem-solving tasks Fable 5.1 achieves 52.6% on the new Terminal-Bench-Science 0.1 benchmark, a dramatic improvement over Fable 5 (24.7%), Opus 5 (29.0%), and GPT-5.6 Sol (22.4%) The "pelican SVG" benchmark reveals that Fable 5.1's reasoning effort levels (low/medium/high/xhigh/max) produce wildly different results, with low and medium settings apparently skipping reasoning entirel
Analysis
TL;DR
- Anthropic released Claude Fable 5.1, claiming it sets a new standard for coding, knowledge work, and long-running problem-solving tasks
- Fable 5.1 achieves 52.6% on the new Terminal-Bench-Science 0.1 benchmark, a dramatic improvement over Fable 5 (24.7%), Opus 5 (29.0%), and GPT-5.6 Sol (22.4%)
- The "pelican SVG" benchmark reveals that Fable 5.1's reasoning effort levels (low/medium/high/xhigh/max) produce wildly different results, with low and medium settings apparently skipping reasoning entirely for this prompt
- At max reasoning effort, the model produced its best pelican SVG yet using 65,927 output tokens, $3.30 cost, and nearly 14 minutes of generation time
- The author notes the pelican benchmark's declining correlation with general model capability, though it remains useful for within-family comparisons across reasoning effort levels
Why It Matters
This release highlights Anthropic's continued push into scientific reasoning capabilities, with Terminal-Bench-Science showing the most dramatic improvement among all benchmarks. The reasoning effort tiering system also reveals important practical insights: lower effort settings may silently skip reasoning on certain prompts, which has direct implications for cost-performance tradeoffs in production deployments.
Technical Details
- Fable 5.1 introduces five reasoning levels: low, medium, high, xhigh, and max, with no option to disable reasoning entirely
- Terminal-Bench-Science 0.1 is a new benchmark (announced August 27, 2026) where Fable 5.1 scored 52.6%, more than doubling its predecessor's 24.7%
- At low/medium effort, the model generated ~2,000 output tokens in ~24 seconds for the pelican prompt with no visible reasoning trace
- At xhigh effort, the model used 36,767 tokens, took 7m51s, and cost $1.83, producing detailed SVG planning reasoning
- At max effort, the model used 65,927 tokens, took 13m54s, and cost $3.30, producing the most detailed and visually coherent pelican SVG with features like a blue hat, fish basket, and anatomically correct leg positioning
Industry Insight
- The dramatic cost and latency scaling at higher reasoning levels ($0.01 at low vs. $3.30 at max) suggests practitioners should carefully calibrate effort settings per use case rather than defaulting to max
- The observation that low/medium reasoning may silently skip reasoning on creative/visual prompts indicates a need for better transparency in how reasoning effort is applied across different task types
- Anthropic's focus on scientific benchmarks signals that the competitive frontier is shifting toward specialized, agentic task performance rather than pure language understanding
Disclaimer: The above content is generated by AI and is for reference only.