AI Skills AI技能 3h ago Updated 1h ago 更新于 1小时前 52

Claude's Protein Design Hit Rate Was 26.8%. One Target Returned 0 for 90. Claude蛋白质设计命中率为26.8%,一个靶点90个设计全部失败

Anthropic's autonomous protein design campaign achieved a pooled hit rate of 26.8% (354 binders from 1,320 designs), but this masks enormous per-target variance ranging from 80% on TREM2 to 0% on maltose-binding protein. A beta-binomial model fitted to published per-target counts reveals that for a new unseen target, the probability of falling below the 10% industry floor is approximately 38%, with a median hit rate of only 18.6%. The pipeline's in-silico confidence scores failed to distinguish Anthropic Claude蛋白质设计活动总体命中率26.8%(354/1320)是16个目标的组合平均值,非单目标概率 单目标命中率差异极大:TREM2达80%,而MBP为0%(90个设计无一结合) Beta-binomial模型拟合显示新目标命中率中位数为18.6%,38%概率低于10%行业基准 模型置信度评分无法区分"好目标"与"难目标",对MBP和BBF-14的评分与成功目标相当 单目标35.1%命中率与多目标26.7%的差异无法分离"聚焦"与"计算预算"两个因素

72
Hot 热度
78
Quality 质量
75
Impact 影响力

Analysis 深度分析

TL;DR

  • Anthropic's autonomous protein design campaign achieved a pooled hit rate of 26.8% (354 binders from 1,320 designs), but this masks enormous per-target variance ranging from 80% on TREM2 to 0% on maltose-binding protein.
  • A beta-binomial model fitted to published per-target counts reveals that for a new unseen target, the probability of falling below the 10% industry floor is approximately 38%, with a median hit rate of only 18.6%.
  • The pipeline's in-silico confidence scores failed to distinguish between targets that would succeed and those that would completely fail, scoring MBP and BBF-14 designs similarly to designs against working targets.
  • Under the fitted model, achieving 95% confidence of at least one binder requires approximately 638 designs per target — a 64× increase over what the headline 26.8% rate would suggest.
  • The campaign's single-target runs at 35.1% hit rate conflated focused prompting with a 2.8× larger compute budget per target, making the "narrow scope is free" narrative incomplete.

Why It Matters

This analysis exposes a critical flaw in how agentic AI benchmark results are reported and interpreted: pooled pass rates across heterogeneous tasks obscure the variance that any single deployment will actually experience. For AI practitioners building autonomous systems for scientific discovery, the finding that confidence scores cannot discriminate between solvable and closed targets is a direct warning about over-reliance on in-silico filtering without wet-lab validation. The protein design case serves as a template for evaluating every agentic benchmark — SWE-bench, ARC-AGI, τ-bench — where a single aggregate score similarly hides the distribution of per-item difficulty.

Technical Details

  • Campaign architecture: Claude (Mythos Preview and Opus 4.8) orchestrated 16 open-source protein design tools across 24 distinct workflows without human intervention on design decisions. Backbones came from PXDesign (358 designs), RFdiffusion3 (267), Genie 3 (185), FreeBindCraft (135), BoltzGen (134), RFdiffusion (118), and Proteina-Complexa (100), with sequences from SolubleMPNN filtered through ESMFold2 and Protenix v2.
  • Validation: Independent wet-lab validation at Adaptyv Bio using surface plasmon resonance at five target concentrations in duplicate. Designs arrived anonymized. 95% of designs expressed; 354 of 1,320 bound.
  • Statistical methodology: A beta-binomial model was fitted by maximum likelihood using the Nelder-Mead optimizer from nine starting points in Python 3.11 with SciPy 1.15.3. The fitted parameters were α = 0.459, β = 1.169, implying a mean of 28.2%. A likelihood-ratio test comparing a single shared rate against per-target rates rejected the shared-rate model at p ≈ 1.8 × 10⁻⁶³.
  • Per-target results: TREM2 — 72/90 (80.0%); VEGF-A — 54/90 (60.0%); IL-7Rα — 49/90 (54.4%); RBX1 — 28/90 (31.1%); TNFα — 12/150 (8.0%); BBF-14 — 3/90 (3.3%); 15-PGDH — 1/30 (3.3%); MBP — 0/90 (0.0%).
  • Confidence score calibration: The co-folding model's confidence scores were discriminative within a target (top-ranked designs bound 49% of the time vs. 26.8% overall) but blind across targets, providing no warning for the MBP and BBF-14 failures.

Industry Insight

  • Demand per-item distributions, not just means: Any agentic AI benchmark reporting a single aggregate pass rate should be treated as incomplete. The variance across items is the missing half of the result. Practitioners should request or infer per-item difficulty distributions before making deployment decisions based on published scores.
  • Budget and scope are confounded in agent performance claims: The 35.1% single-target hit rate cannot be attributed to narrow scoping alone — it came with 2.8× the compute budget per target. Organizations should be skeptical of claims that focusing an agent's scope is a free improvement; resource allocation changes are likely contributing factors.
  • Uncalibrated self-assessment is a systemic risk for agentic discovery: When an agent's internal confidence scores cannot distinguish solvable from closed targets, the agent will confidently ship failures at scale. This pattern — high average performance, enormous variance, and uncalibrated self-assessment — is likely to appear across agentic AI benchmarks in 2026 and beyond, making independent validation a non-negotiable component of any deployment pipeline.

TL;DR

  • Anthropic Claude蛋白质设计活动总体命中率26.8%(354/1320)是16个目标的组合平均值,非单目标概率
  • 单目标命中率差异极大:TREM2达80%,而MBP为0%(90个设计无一结合)
  • Beta-binomial模型拟合显示新目标命中率中位数为18.6%,38%概率低于10%行业基准
  • 模型置信度评分无法区分"好目标"与"难目标",对MBP和BBF-14的评分与成功目标相当
  • 单目标35.1%命中率与多目标26.7%的差异无法分离"聚焦"与"计算预算"两个因素

为什么值得看

这篇文章揭示了AI代理基准测试中普遍存在的"聚合数字陷阱"——26.8%的 headline 数字掩盖了目标间巨大的性能方差,对依赖单一命中率做决策的从业者具有警示意义。其方法论(beta-binomial建模)可直接迁移至SWE-bench、ARC-AGI等其他AI代理基准评估,帮助行业建立更科学的性能预期。

技术解析

  • 实验设计:Claude(Mythos Preview和Opus 4.8)在16个蛋白质目标上各生成30个minibinder设计,使用PXDesign、RFdiffusion3、Genie 3等12种开源工具组合,经SolubleMPNN序列生成和ESMFold2/Protenix v2过滤,无人类干预设计决策
  • 统计建模:对8个已公布的目标命中率数据拟合beta-binomial模型(α=0.459, β=1.169),揭示命中率分布的长尾特征;α<1意味着"完全失败"概率以N^(-α)衰减而非指数衰减
  • 验证方法:Adaptyv Bio独立验证,表面等离子共振(SPR)在5个靶标浓度下双重复测量,95%设计成功表达
  • 置信度校准问题:co-folding模型评分在目标内具有判别力(Top-1设计49%结合率vs整体26.8%),但在目标间完全盲态,无法预警MBP等"不可设计"靶标

行业启示

  • 基准评估范式需变革:单一聚合命中率(如SWE-bench的62%)掩盖了任务难度分布,应强制报告逐项分布而非仅均值,否则用户会错误地将组合概率当作单任务概率
  • 计算预算与性能不可解耦:35.1% vs 26.7%的差异源于2.8倍计算预算(2500 vs 12500 H100-hours),"聚焦策略"的红利被高估,资源分配需纳入ROI评估
  • 代理系统的自我评估存在系统性盲区:置信度评分无法跨目标校准的问题在蛋白质设计、代码生成、科学发现等代理基准中普遍存在,需建立跨域校准机制或引入外部验证回路

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Claude Claude Research 科学研究 Agent Agent Healthcare AI 医疗AI Evaluation 评测