AI News AI资讯 18h ago Updated 17h ago 更新于 17小时前 43

Why biological data matters more in AI drug discovery 为什么生物数据在AI药物发现中更重要

GSK is expanding its AI-drug-discovery partnership with Relation Therapeutics through an $110 million collaboration focused on generating large-scale functional single-cell datasets to train AI models for target identification Relation's "Lab-in-the-Loop" approach integrates wet-lab perturbation experiments with computational analysis, combining single-cell/spatial transcriptomics, sequencing, and machine learning to validate disease targets Recent research in Nature Methods demonstrates that si GSK与英国生物技术公司Relation Therapeutics达成高达1.1亿美元的合作,扩展AI辅助药物发现领域 Relation将生成大规模人类细胞对基因变化和药物干预反应的 datasets,用于训练AI模型识别潜在药物靶点 2025年《Nature Methods》研究显示单细胞基础模型在训练少量数据后即达到性能瓶颈,不存在类似大语言模型的数据缩放定律 制药公司正转向获取专业化、疾病特定的数据集,如Relation的Osteomics骨图谱和GSK与Ochre Bio的肝脏单细胞数据合作 高质量、非冗余数据集的构建与模型架构同样重要,单纯增加生物训练数据并不保证性能提升

62
Hot 热度
65
Quality 质量
58
Impact 影响力

Analysis 深度分析

TL;DR

  • GSK is expanding its AI-drug-discovery partnership with Relation Therapeutics through an $110 million collaboration focused on generating large-scale functional single-cell datasets to train AI models for target identification
  • Relation's "Lab-in-the-Loop" approach integrates wet-lab perturbation experiments with computational analysis, combining single-cell/spatial transcriptomics, sequencing, and machine learning to validate disease targets
  • Recent research in Nature Methods demonstrates that single-cell foundation models do not follow clear data-scaling laws like large language models, plateauing after training on only a fraction of available data
  • High-quality, non-redundant, disease-specific datasets are proving more valuable than sheer data volume, driving pharma companies to pursue proprietary specialized datasets rather than relying solely on public repositories
  • The broader industry trend shows AI-focused biopharma deals increasingly centering on specialized dataset providers, with larger upfront payments and greater participation from major pharmaceutical companies

Why It Matters

This collaboration and the accompanying research highlight a critical inflection point in AI-driven drug discovery: the field is moving beyond the assumption that bigger public datasets automatically produce better models. For AI practitioners and biopharma researchers, the findings underscore that dataset curation, quality control, and disease-specific specialization are now as important as model architecture. The results also challenge the direct transfer of LLM scaling strategies to biological foundation models, urging a more balanced approach to computational resources, model capacity, and data diversity.

Technical Details

  • Relation's Lab-in-the-Loop platform combines laboratory experimentation (tissue profiling, single-cell and spatial transcriptomics, sequencing, perturbation experiments) with machine learning for target identification, prioritization, validation, and experimental design, creating a closed loop between computation and wet-lab generation
  • MORGAN platform is Relation's AI model infrastructure trained on large-scale datasets measuring human cellular responses to genetic changes and drug interventions, designed to identify and validate potential drug targets
  • Osteomics is Relation's proprietary functional single-cell bone atlas, integrating patient-derived samples with single-cell and spatial omics, imaging, genomics, proteomics, and clinical phenotype data to study osteoporosis disease biology and therapeutic targets
  • Public repository challenges include batch effects, technical noise, dataset overlap causing data-leakage risks, and disproportionate influence of repeated cells across resources like CZ CELLxGENE (100+ million cells), Human Cell Atlas, and NCBI Gene Expression Omnibus
  • Nature Methods study trained 400 single-cell foundation models across 6,400 experiments using a 22.2 million-cell corpus, finding performance plateaus well before exhausting available data and no clear scaling laws analogous to LLMs
  • Genome Biology 2025 study evaluated Geneformer and scGPT on zero-shot tasks, finding they did not consistently outperform simpler approaches and cautioning against assuming larger pretrained models yield better biological representations

Industry Insight

  • Pharma companies are increasingly treating high-quality, disease-specific datasets as strategic assets rather than commodities, with deals like GSK-Ochre Bio ($37.5M for liver single-cell data) and AstraZeneca-Pathos AI-Tempus ($200M for oncology foundation models) signaling that proprietary data access is becoming a key competitive moat in AI drug discovery
  • Researchers and AI practitioners should prioritize dataset curation, composition balancing, and quality control over simply scaling data volume, as the Nature Methods and Genome Biology studies demonstrate diminishing returns from larger but redundant or noisy biological datasets
  • The failure of simple scaling laws in single-cell models suggests that the next wave of breakthroughs will come from integrating multi-modal data (spatial omics, proteomics, clinical phenotypes) and causal modeling approaches rather than from larger pretraining corpora alone, making partnerships with specialized data generators like Relation Therapeutics strategically valuable

TL;DR

  • GSK与英国生物技术公司Relation Therapeutics达成高达1.1亿美元的合作,扩展AI辅助药物发现领域
  • Relation将生成大规模人类细胞对基因变化和药物干预反应的 datasets,用于训练AI模型识别潜在药物靶点
  • 2025年《Nature Methods》研究显示单细胞基础模型在训练少量数据后即达到性能瓶颈,不存在类似大语言模型的数据缩放定律
  • 制药公司正转向获取专业化、疾病特定的数据集,如Relation的Osteomics骨图谱和GSK与Ochre Bio的肝脏单细胞数据合作
  • 高质量、非冗余数据集的构建与模型架构同样重要,单纯增加生物训练数据并不保证性能提升

为什么值得看

本文揭示了AI药物发现领域的重要趋势:从依赖公开数据集转向获取专业化、高质量疾病特定数据集。研究证明单细胞基础模型存在性能瓶颈,为制药公司数据战略提供了关键指导。

技术解析

  • Relation的"Lab-in-the-Loop"方法结合实验室实验与计算分析,包括组织分析、单细胞和空间转录组学、测序和目标验证,机器学习用于靶点识别、优先排序、验证和实验设计
  • 2025年《Nature Methods》研究使用2220万个细胞的语料库训练400个模型,评估6400次实验,发现单细胞基础模型在训练少量数据后达到性能平台期,不遵循数据缩放定律
  • 公开存储库如CZ CELLxGENE提供超过1亿个标准化细胞数据,但存在采样方法、测序协议、实验程序和数据处理管道的差异,以及技术噪声和批次效应问题
  • 数据集重叠可能导致训练和测试数据泄漏风险,高质量非冗余数据集的组装与模型架构同样重要
  • Relation的Osteomics项目结合单细胞和空间组学、成像、基因组学、蛋白质组学和临床表型数据,用于骨质疏松症研究

行业启示

  • 制药公司正从通用公开数据转向获取专业化、疾病特定的高质量数据集,这将成为AI药物发现竞争的关键差异化因素
  • 单纯增加训练数据规模不再能保证模型性能提升,需要平衡模型容量、数据集大小和计算资源
  • 2025年《Nature Biotechnology》分析显示AI生物制药合作呈现更大预付款、新治疗模式和大型生物技术公司更多参与的趋势

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Healthcare AI 医疗AI Dataset 数据集 Funding 融资 Research 科学研究