Anthropic says any lab can now let a language model agent run the whole protein design stack
Anthropic's Claude models (Mythos Preview and Opus 4.8) achieved a 26.8% hit rate in de novo protein binder design across 15 targets, significantly outperforming the typical industry range of 10–15% Claude operated as an autonomous agent orchestrating 24+ open-source tools (RFdiffusion3, ProteinMPNN, ESMFold2, etc.) without human intervention on individual design decisions, completing multi-target campaigns in 48 hours In direct comparison on the RBX1 target, Claude's best design bound 10× tight
Analysis
TL;DR
- Anthropic's Claude models (Mythos Preview and Opus 4.8) achieved a 26.8% hit rate in de novo protein binder design across 15 targets, significantly outperforming the typical industry range of 10–15%
- Claude operated as an autonomous agent orchestrating 24+ open-source tools (RFdiffusion3, ProteinMPNN, ESMFold2, etc.) without human intervention on individual design decisions, completing multi-target campaigns in 48 hours
- In direct comparison on the RBX1 target, Claude's best design bound 10× tighter (3.9 nM) than the winning entry from an open design contest (45 nM)
- Claude also demonstrated rapid raw analytical chemistry interpretation, decoding proprietary LC-MS and NMR file formats and delivering results in under 25 minutes with high accuracy
- The authors acknowledge significant limitations: no human expert control group, no structural validation of designs, potential prompt knowledge leakage from contest targets in the reading list, and single-run experiments preventing separation of model effects from chance
Why It Matters
This represents a significant step toward autonomous AI-driven scientific discovery, demonstrating that general-purpose language models can orchestrate complex, multi-step laboratory workflows without domain-specific fine-tuning. For AI practitioners and computational biologists, it validates the agentic approach to scientific tool use and establishes a reproducible benchmark dataset now available on Hugging Face. The results also raise important questions about the future role of human experts in drug discovery pipelines and the accessibility of advanced protein design for under-resourced labs.
Technical Details
- Agent Architecture: Claude (Mythos Preview and Opus 4.8) functioned as a general-purpose agent using a ~16,000-word protocol prompt (one-third scientific guidance, two-thirds scheduling/delegation/verification/budget discipline). The model autonomously installed open-source tools from public repositories, selected docking sites (epitopes) without explicit instruction, and orchestrated 24 different workflows across tools including PXDesign, RFdiffusion3, Genie 3, FreeBindCraft, BoltzGen, SolubleMPNN, ESMFold2, and Protenix v2
- Compute and Budget: Multi-target campaigns operated under a $50,000 budget across 16 targets within 48 hours; single-target runs received $10,000 each. All compute ran through Modal cloud infrastructure with minimal human intervention (only non-technical instructions needed after infrastructure outages)
- Experimental Validation: 1,320 designs were synthesized and tested by contract labs (Adaptyv Bio and Twist Bioscience). Binding was measured via KD values in nanomolar. 354 of 1,320 designs (26.8%) bound to targets; 49% of Claude's top-ranked designs bound. Cross-species binding to mouse counterparts was achieved for 130 of 233 tested binders
- Analytical Chemistry Task: Opus 5 interpreted raw NMR and LC-MS files from proprietary instrument formats, decoding the LC-MS format autonomously and reproducing all 2,664 stored summary values exactly. The model also proposed follow-up experiments and self-corrected errors in real-time
- Limitations and Exclusions: AlphaFold-3, Rosetta, and ESM3 were excluded due to licensing. No designs were structurally resolved experimentally. Four of six contest targets had known solutions in Claude's reading list. Each model-format-target combination ran exactly once, preventing statistical separation of model performance from randomness
Industry Insight
- Democratization of Protein Design: Since all tools used are open-source and the prompts/datasets are publicly available, sophisticated de novo protein design campaigns that previously required specialized expertise and significant compute budgets may become accessible to any lab with cloud computing resources, potentially accelerating the pace of early-stage drug discovery across the industry
- Agentic AI as a Force Multiplier: The success of a general-purpose language model orchestrating complex scientific toolchains without domain-specific training validates the agentic AI paradigm for scientific automation. Organizations should invest in developing robust protocol prompts, verification layers, and budget discipline mechanisms rather than solely focusing on model architecture improvements
- Critical Need for Rigorous Benchmarking: The acknowledged limitations—single runs, potential knowledge leakage, lack of structural validation, and absence of human expert controls—highlight that these results, while promising, require independent replication and controlled comparisons before claiming definitive superiority. The scientific community should prioritize establishing standardized benchmarks and parallel human expert campaigns to properly calibrate AI capabilities in drug discovery
Disclaimer: The above content is generated by AI and is for reference only.