AI News AI资讯 2d ago Updated 2d ago 更新于 2天前 50

Anthropic says any lab can now let a language model agent run the whole protein design stack Anthropic称任何实验室现在都可以让语言模型代理运行整个蛋白质设计流程

Anthropic's Claude models (Mythos Preview and Opus 4.8) achieved a 26.8% hit rate in de novo protein binder design across 15 targets, significantly outperforming the typical industry range of 10–15% Claude operated as an autonomous agent orchestrating 24+ open-source tools (RFdiffusion3, ProteinMPNN, ESMFold2, etc.) without human intervention on individual design decisions, completing multi-target campaigns in 48 hours In direct comparison on the RBX1 target, Claude's best design bound 10× tight Anthropic的Claude模型在早期药物发现实验中成功设计了针对16种靶蛋白的minibinders,其中15种产生可用测量结果,14种成功结合,整体命中率26.8% Claude通过编排24种开源工具工作流自动完成靶点研究、对接位点选择、程序安装和候选物排序,无需人类干预设计决策 在RBX1靶点上,Claude最佳设计结合强度达3.9 nM,击败竞赛获胜设计(45 nM),约强10倍;TNFα靶点此前多家方法零命中,Claude产出12个结合物 第二个实验中,Opus 5在23和19分钟内完成NMR和LC-MS原始数据分析,并自行解码LC-MS专有格式,交叉验证2664个数据点完全复现

72
Hot 热度
75
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Anthropic's Claude models (Mythos Preview and Opus 4.8) achieved a 26.8% hit rate in de novo protein binder design across 15 targets, significantly outperforming the typical industry range of 10–15%
  • Claude operated as an autonomous agent orchestrating 24+ open-source tools (RFdiffusion3, ProteinMPNN, ESMFold2, etc.) without human intervention on individual design decisions, completing multi-target campaigns in 48 hours
  • In direct comparison on the RBX1 target, Claude's best design bound 10× tighter (3.9 nM) than the winning entry from an open design contest (45 nM)
  • Claude also demonstrated rapid raw analytical chemistry interpretation, decoding proprietary LC-MS and NMR file formats and delivering results in under 25 minutes with high accuracy
  • The authors acknowledge significant limitations: no human expert control group, no structural validation of designs, potential prompt knowledge leakage from contest targets in the reading list, and single-run experiments preventing separation of model effects from chance

Why It Matters

This represents a significant step toward autonomous AI-driven scientific discovery, demonstrating that general-purpose language models can orchestrate complex, multi-step laboratory workflows without domain-specific fine-tuning. For AI practitioners and computational biologists, it validates the agentic approach to scientific tool use and establishes a reproducible benchmark dataset now available on Hugging Face. The results also raise important questions about the future role of human experts in drug discovery pipelines and the accessibility of advanced protein design for under-resourced labs.

Technical Details

  • Agent Architecture: Claude (Mythos Preview and Opus 4.8) functioned as a general-purpose agent using a ~16,000-word protocol prompt (one-third scientific guidance, two-thirds scheduling/delegation/verification/budget discipline). The model autonomously installed open-source tools from public repositories, selected docking sites (epitopes) without explicit instruction, and orchestrated 24 different workflows across tools including PXDesign, RFdiffusion3, Genie 3, FreeBindCraft, BoltzGen, SolubleMPNN, ESMFold2, and Protenix v2
  • Compute and Budget: Multi-target campaigns operated under a $50,000 budget across 16 targets within 48 hours; single-target runs received $10,000 each. All compute ran through Modal cloud infrastructure with minimal human intervention (only non-technical instructions needed after infrastructure outages)
  • Experimental Validation: 1,320 designs were synthesized and tested by contract labs (Adaptyv Bio and Twist Bioscience). Binding was measured via KD values in nanomolar. 354 of 1,320 designs (26.8%) bound to targets; 49% of Claude's top-ranked designs bound. Cross-species binding to mouse counterparts was achieved for 130 of 233 tested binders
  • Analytical Chemistry Task: Opus 5 interpreted raw NMR and LC-MS files from proprietary instrument formats, decoding the LC-MS format autonomously and reproducing all 2,664 stored summary values exactly. The model also proposed follow-up experiments and self-corrected errors in real-time
  • Limitations and Exclusions: AlphaFold-3, Rosetta, and ESM3 were excluded due to licensing. No designs were structurally resolved experimentally. Four of six contest targets had known solutions in Claude's reading list. Each model-format-target combination ran exactly once, preventing statistical separation of model performance from randomness

Industry Insight

  • Democratization of Protein Design: Since all tools used are open-source and the prompts/datasets are publicly available, sophisticated de novo protein design campaigns that previously required specialized expertise and significant compute budgets may become accessible to any lab with cloud computing resources, potentially accelerating the pace of early-stage drug discovery across the industry
  • Agentic AI as a Force Multiplier: The success of a general-purpose language model orchestrating complex scientific toolchains without domain-specific training validates the agentic AI paradigm for scientific automation. Organizations should invest in developing robust protocol prompts, verification layers, and budget discipline mechanisms rather than solely focusing on model architecture improvements
  • Critical Need for Rigorous Benchmarking: The acknowledged limitations—single runs, potential knowledge leakage, lack of structural validation, and absence of human expert controls—highlight that these results, while promising, require independent replication and controlled comparisons before claiming definitive superiority. The scientific community should prioritize establishing standardized benchmarks and parallel human expert campaigns to properly calibrate AI capabilities in drug discovery

TL;DR

  • Anthropic的Claude模型在早期药物发现实验中成功设计了针对16种靶蛋白的minibinders,其中15种产生可用测量结果,14种成功结合,整体命中率26.8%
  • Claude通过编排24种开源工具工作流自动完成靶点研究、对接位点选择、程序安装和候选物排序,无需人类干预设计决策
  • 在RBX1靶点上,Claude最佳设计结合强度达3.9 nM,击败竞赛获胜设计(45 nM),约强10倍;TNFα靶点此前多家方法零命中,Claude产出12个结合物
  • 第二个实验中,Opus 5在23和19分钟内完成NMR和LC-MS原始数据分析,并自行解码LC-MS专有格式,交叉验证2664个数据点完全复现
  • 报告明确承认局限性:无人类专家对照实验、无结构解析、模型差异与随机因素无法分离,且部分竞赛靶点的设计方案已在Claude阅读列表中

为什么值得看

本文展示了通用语言模型在复杂科学工作流编排中的突破性能力,证明AI可以替代领域专家完成从靶点研究到候选物排序的全流程自动化决策。对药物发现行业而言,这意味着早期筛选成本可能大幅降低,开源工具链的可及性让中小实验室也能参与AI驱动的药物设计竞争。

技术解析

Claude并未使用自研蛋白质模型,而是通过约16,000词的协议提示词作为系统提示,编排PXDesign、RFdiffusion3、Genie 3、FreeBindCraft、BoltzGen、Proteina-Complexa等24种开源工具,结合SolubleMPNN生成氨基酸链,再用ESMFold2、Protenix v2进行折叠预测和置信度评分。AlphaFold-3、Rosetta、ESM3因许可原因被排除。计算预算为多靶点战役5万美元、单靶点1万美元,通过Modal云运行。

实验采用de novo设计策略,从计算机端重新设计minibinders而非从自然界筛选。多靶点模式下Mythos Preview和Opus 4.8分别达到26.7%和22.6%命中率;单靶点模式下Mythos Preview提升至35.1%,但计算预算为多靶点模式的2.8倍。49%的命中率来自Claude自评排名第一的设计。

第二个实验聚焦分析化学数据处理。Opus 5从原始NMR和LC-MS文件出发,在23分钟和19分钟内完成分析。对于LC-MS专有格式,模型自行解码并复现全部2664个数据点的汇总值;NMR分析中模型先报告4个缺失信号,经内部核查修正为2个,并提出了与实验室三天后独立执行的后续实验相同的建议。

行业启示

AI驱动的药物发现正从"辅助工具"向"自主工作流编排者"演进。通用语言模型无需领域微调即可整合24种专业工具、完成跨学科决策,这为其他复杂科学领域(如材料发现、化学合成路径设计)的自动化提供了可复制范式。

开源工具链+大模型提示工程的组合大幅降低了AI药物设计的门槛。Anthropic公开了提示词、设计数据和测量数据集,意味着任何实验室均可复现或改进该流程,可能加速整个领域的迭代速度。

实验结果的验证仍需谨慎解读。缺乏人类专家对照、无结构解析、模型与随机因素无法分离等问题表明,当前AI药物发现仍处于早期验证阶段。行业应建立更严格的基准测试和对照实验标准,避免过度解读初步成果。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Claude Claude Agent Agent Healthcare AI 医疗AI Research 科学研究 LLM 大模型