AI News AI资讯 4h ago Updated 2h ago 更新于 2小时前 49

Claude Fable 5.1 made me a really nice animated pelican Claude Fable 5.1 让我制作了一只非常可爱的动画鹈鹕

Anthropic released Claude Fable 5.1, claiming it sets a new standard for coding, knowledge work, and long-running problem-solving tasks Fable 5.1 achieves 52.6% on the new Terminal-Bench-Science 0.1 benchmark, a dramatic improvement over Fable 5 (24.7%), Opus 5 (29.0%), and GPT-5.6 Sol (22.4%) The "pelican SVG" benchmark reveals that Fable 5.1's reasoning effort levels (low/medium/high/xhigh/max) produce wildly different results, with low and medium settings apparently skipping reasoning entirel Claude Fable 5.1在Terminal-Bench-Science 0.1基准测试中取得52.6%的突破性成绩,远超Fable 5的24.7%、Opus 5的29.0%及GPT-5.6 Sol的22.4% Fable 5.1引入五级推理强度设置(low/medium/high/xhigh/max),但low和medium级别对创意SVG任务似乎完全跳过推理过程 高强度推理显著提升输出质量:max级别生成65,927 tokens、耗时13分54秒、花费$3.30,产生详细的SVG设计决策痕迹 推理成本与质量呈非线性关系:从low到max,成本从10美分跃升至$3.30,但输出质量差

72
Hot 热度
68
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Anthropic released Claude Fable 5.1, claiming it sets a new standard for coding, knowledge work, and long-running problem-solving tasks
  • Fable 5.1 achieves 52.6% on the new Terminal-Bench-Science 0.1 benchmark, a dramatic improvement over Fable 5 (24.7%), Opus 5 (29.0%), and GPT-5.6 Sol (22.4%)
  • The "pelican SVG" benchmark reveals that Fable 5.1's reasoning effort levels (low/medium/high/xhigh/max) produce wildly different results, with low and medium settings apparently skipping reasoning entirely for this prompt
  • At max reasoning effort, the model produced its best pelican SVG yet using 65,927 output tokens, $3.30 cost, and nearly 14 minutes of generation time
  • The author notes the pelican benchmark's declining correlation with general model capability, though it remains useful for within-family comparisons across reasoning effort levels

Why It Matters

This release highlights Anthropic's continued push into scientific reasoning capabilities, with Terminal-Bench-Science showing the most dramatic improvement among all benchmarks. The reasoning effort tiering system also reveals important practical insights: lower effort settings may silently skip reasoning on certain prompts, which has direct implications for cost-performance tradeoffs in production deployments.

Technical Details

  • Fable 5.1 introduces five reasoning levels: low, medium, high, xhigh, and max, with no option to disable reasoning entirely
  • Terminal-Bench-Science 0.1 is a new benchmark (announced August 27, 2026) where Fable 5.1 scored 52.6%, more than doubling its predecessor's 24.7%
  • At low/medium effort, the model generated ~2,000 output tokens in ~24 seconds for the pelican prompt with no visible reasoning trace
  • At xhigh effort, the model used 36,767 tokens, took 7m51s, and cost $1.83, producing detailed SVG planning reasoning
  • At max effort, the model used 65,927 tokens, took 13m54s, and cost $3.30, producing the most detailed and visually coherent pelican SVG with features like a blue hat, fish basket, and anatomically correct leg positioning

Industry Insight

  • The dramatic cost and latency scaling at higher reasoning levels ($0.01 at low vs. $3.30 at max) suggests practitioners should carefully calibrate effort settings per use case rather than defaulting to max
  • The observation that low/medium reasoning may silently skip reasoning on creative/visual prompts indicates a need for better transparency in how reasoning effort is applied across different task types
  • Anthropic's focus on scientific benchmarks signals that the competitive frontier is shifting toward specialized, agentic task performance rather than pure language understanding

TL;DR

  • Claude Fable 5.1在Terminal-Bench-Science 0.1基准测试中取得52.6%的突破性成绩,远超Fable 5的24.7%、Opus 5的29.0%及GPT-5.6 Sol的22.4%
  • Fable 5.1引入五级推理强度设置(low/medium/high/xhigh/max),但low和medium级别对创意SVG任务似乎完全跳过推理过程
  • 高强度推理显著提升输出质量:max级别生成65,927 tokens、耗时13分54秒、花费$3.30,产生详细的SVG设计决策痕迹
  • 推理成本与质量呈非线性关系:从low到max,成本从10美分跃升至$3.30,但输出质量差异显著

为什么值得看

本文通过"生成骑自行鹈鹕SVG"这一具体任务,揭示了Claude Fable 5.1推理机制的实际运作方式,为开发者理解不同推理强度对输出质量的影响提供了实证参考。

技术解析

  • 基准测试突破:Fable 5.1在Terminal-Bench-Science 0.1基准测试中得分52.6%,较Fable 5的24.7%实现翻倍提升,Anthropic重点强调其在科学研究的进展
  • 推理强度机制:模型提供五级推理设置,但low和medium级别对创意类任务(如SVG生成)似乎跳过推理,直接输出结果;high级别仅有少量推理摘要
  • max级别推理深度:max级别产生65,927 tokens输出,推理痕迹包含详细的SVG设计决策,如坐标计算、视觉元素布局、颜色选择及细节优化
  • 成本-质量权衡:max级别耗时13分54秒、花费$3.30,相比low级别的23.8秒和10.017美分,成本增加约33倍,但输出质量显著提升

行业启示

  • 推理强度选择策略:开发者应根据任务类型(创意vs逻辑)和成本预算选择合适的推理级别,对于SVG生成等创意任务,low/medium级别可能无法发挥模型潜力
  • 科学推理能力成为新竞争焦点:Anthropic将Terminal-Bench-Science作为Fable 5.1的主要卖点,表明AI模型在科学研究领域的表现正成为差异化竞争的关键维度
  • 推理成本透明化趋势:模型提供商需提供清晰的推理成本-质量曲线,帮助用户在精度和效率之间做出明智决策

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Claude Claude Benchmark 基准测试 Product Launch 产品发布 LLM 大模型 Evaluation 评测