Why biological data matters more in AI drug discovery
GSK is expanding its AI-drug-discovery partnership with Relation Therapeutics through an $110 million collaboration focused on generating large-scale functional single-cell datasets to train AI models for target identification Relation's "Lab-in-the-Loop" approach integrates wet-lab perturbation experiments with computational analysis, combining single-cell/spatial transcriptomics, sequencing, and machine learning to validate disease targets Recent research in Nature Methods demonstrates that si
Analysis
TL;DR
- GSK is expanding its AI-drug-discovery partnership with Relation Therapeutics through an $110 million collaboration focused on generating large-scale functional single-cell datasets to train AI models for target identification
- Relation's "Lab-in-the-Loop" approach integrates wet-lab perturbation experiments with computational analysis, combining single-cell/spatial transcriptomics, sequencing, and machine learning to validate disease targets
- Recent research in Nature Methods demonstrates that single-cell foundation models do not follow clear data-scaling laws like large language models, plateauing after training on only a fraction of available data
- High-quality, non-redundant, disease-specific datasets are proving more valuable than sheer data volume, driving pharma companies to pursue proprietary specialized datasets rather than relying solely on public repositories
- The broader industry trend shows AI-focused biopharma deals increasingly centering on specialized dataset providers, with larger upfront payments and greater participation from major pharmaceutical companies
Why It Matters
This collaboration and the accompanying research highlight a critical inflection point in AI-driven drug discovery: the field is moving beyond the assumption that bigger public datasets automatically produce better models. For AI practitioners and biopharma researchers, the findings underscore that dataset curation, quality control, and disease-specific specialization are now as important as model architecture. The results also challenge the direct transfer of LLM scaling strategies to biological foundation models, urging a more balanced approach to computational resources, model capacity, and data diversity.
Technical Details
- Relation's Lab-in-the-Loop platform combines laboratory experimentation (tissue profiling, single-cell and spatial transcriptomics, sequencing, perturbation experiments) with machine learning for target identification, prioritization, validation, and experimental design, creating a closed loop between computation and wet-lab generation
- MORGAN platform is Relation's AI model infrastructure trained on large-scale datasets measuring human cellular responses to genetic changes and drug interventions, designed to identify and validate potential drug targets
- Osteomics is Relation's proprietary functional single-cell bone atlas, integrating patient-derived samples with single-cell and spatial omics, imaging, genomics, proteomics, and clinical phenotype data to study osteoporosis disease biology and therapeutic targets
- Public repository challenges include batch effects, technical noise, dataset overlap causing data-leakage risks, and disproportionate influence of repeated cells across resources like CZ CELLxGENE (100+ million cells), Human Cell Atlas, and NCBI Gene Expression Omnibus
- Nature Methods study trained 400 single-cell foundation models across 6,400 experiments using a 22.2 million-cell corpus, finding performance plateaus well before exhausting available data and no clear scaling laws analogous to LLMs
- Genome Biology 2025 study evaluated Geneformer and scGPT on zero-shot tasks, finding they did not consistently outperform simpler approaches and cautioning against assuming larger pretrained models yield better biological representations
Industry Insight
- Pharma companies are increasingly treating high-quality, disease-specific datasets as strategic assets rather than commodities, with deals like GSK-Ochre Bio ($37.5M for liver single-cell data) and AstraZeneca-Pathos AI-Tempus ($200M for oncology foundation models) signaling that proprietary data access is becoming a key competitive moat in AI drug discovery
- Researchers and AI practitioners should prioritize dataset curation, composition balancing, and quality control over simply scaling data volume, as the Nature Methods and Genome Biology studies demonstrate diminishing returns from larger but redundant or noisy biological datasets
- The failure of simple scaling laws in single-cell models suggests that the next wave of breakthroughs will come from integrating multi-modal data (spatial omics, proteomics, clinical phenotypes) and causal modeling approaches rather than from larger pretraining corpora alone, making partnerships with specialized data generators like Relation Therapeutics strategically valuable
Disclaimer: The above content is generated by AI and is for reference only.