AI News AI资讯 1d ago Updated 1d ago 更新于 1天前 48

Anthropic's most capable model, codenamed "Model 2," is for internal use only Anthropic 最强大的模型(代号"Model 2")仅限内部使用

Anthropic is running an unreleased internal model codenamed "Model 2" that outperforms every publicly available version of Claude The model is classified in the "Mythos" tier and scores approximately 1.5 points above Claude Mythos 5 on Anthropic's internal AECI capability index Despite being stronger overall, Model 2 is weaker in certain areas and does not represent a dramatic capability leap comparable to the Opus 4.6 to Mythos transition The model is used extensively for coding, data generatio Anthropic内部运行代号"Model 2"的未发布模型,整体能力略超Claude Mythos 5 该模型属于Mythos类别,在内部能力指数AECI上比Mythos 5高出约1.5分 Model 2主要用于编码、数据生成和研发工程,部分通过持续运行的agent使用 Anthropic风险评估为"低",未发现新的对齐问题,暂无对外发布计划

72
Hot 热度
65
Quality 质量
68
Impact 影响力

Analysis 深度分析

TL;DR

  • Anthropic is running an unreleased internal model codenamed "Model 2" that outperforms every publicly available version of Claude
  • The model is classified in the "Mythos" tier and scores approximately 1.5 points above Claude Mythos 5 on Anthropic's internal AECI capability index
  • Despite being stronger overall, Model 2 is weaker in certain areas and does not represent a dramatic capability leap comparable to the Opus 4.6 to Mythos transition
  • The model is used extensively for coding, data generation, and research/engineering, often through continuously running agents
  • Anthropic rates the overall misalignment risk as "low" and has no plans to release Model 2 externally at this time

Why It Matters

Anthropic's internal deployment of a more capable unreleased model signals the accelerating pace of capability gains happening behind closed doors, even as public releases remain incremental. The fact that Claude already writes most of Anthropic's production code underscores how deeply AI integration has penetrated core engineering workflows, making internal model performance a direct competitive differentiator. For the broader industry, this highlights the growing gap between what companies can do internally and what they choose to release publicly.

Technical Details

  • Model Classification: "Model 2" is placed in the Mythos class, positioned slightly above Claude Mythos 5 on Anthropic's internal AECI (Anthropic Evaluation Capability Index), approximately 1.5 points higher
  • Capability Profile: The model shows modest overall improvement rather than a dramatic jump; the gain is notably smaller than the transition from Mythos Preview to Mythos 5, and it exhibits weaknesses in certain unspecified areas
  • Internal Use Cases: Heavily utilized for software coding, synthetic data generation, and research and engineering tasks, with deployment sometimes occurring through autonomous agents running continuously
  • Safety Review: Underwent internal review prior to deployment but received less rigorous testing than Mythos 5; no new or more concerning misalignments were identified, and overall misalignment risk was rated "low"
  • Release Status: No external release planned; the model remains strictly internal as of the August 2026 Risk Report

Industry Insight

  • The incremental nature of Model 2's improvements—despite being Anthropic's most capable model—suggests that the easiest gains may have already been captured, and future progress could require fundamentally different approaches rather than scaling alone
  • Anthropic's reliance on Claude for production code and the deployment of Model 2 through persistent agents reflects a broader industry trend where AI is moving from assistive tool to autonomous infrastructure, raising both productivity and safety questions
  • The deliberate decision to keep a clearly superior model internal signals that capability withholding is becoming a strategic norm among top AI labs, which could slow public benchmark progress while widening the gap between leading labs and everyone else

TL;DR

  • Anthropic内部运行代号"Model 2"的未发布模型,整体能力略超Claude Mythos 5
  • 该模型属于Mythos类别,在内部能力指数AECI上比Mythos 5高出约1.5分
  • Model 2主要用于编码、数据生成和研发工程,部分通过持续运行的agent使用
  • Anthropic风险评估为"低",未发现新的对齐问题,暂无对外发布计划

为什么值得看

本文揭示了Anthropic内部模型迭代的实际进展,为行业提供了关于大模型能力演进速度的真实参考。风险报告的披露方式也体现了AI公司在安全透明度方面的新趋势。

技术解析

  • Model 2被归类为Mythos类别,整体能力略强于Claude Mythos 5,但在某些特定领域表现较弱,未出现类似Opus 4.6到Mythos那样的跨越式提升
  • 内部能力评估采用AECI指数,Model 2得分比Mythos 5高约1.5分,增幅小于Mythos Preview到Mythos 5的跨度
  • 该模型已通过内部审查流程,但未经历与Mythos 5同等程度的测试验证
  • Anthropic生产系统中的代码主要由Claude编写,Model 2在编码、数据生成和研发工程中承担重要角色

行业启示

  • 大模型迭代正从"跨越式突破"转向"渐进式优化",能力边际提升成为新常态
  • 内部模型优先用于提升研发效率(如代码生成),反映了AI公司"用AI改进AI"的战略路径
  • 风险披露机制日益规范化,AI公司通过定期发布风险报告建立行业透明度标准

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Claude Claude LLM 大模型 Closed Source 闭源 Research 科学研究 Security 安全