Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut it
Former OpenAI researcher Andrew Ho launched a startup focused on high-quality training data, arguing that scaling alone won't achieve true generalization in AI models. Existing datasets lack economically valuable skills, especially in bioinformatics and routine lab work, where current models perform poorly (e.g., ~30% success rate in complex analyses). Research from Cambridge and Google Deepmind supports this view: LLMs are becoming more specialized rather than versatile, with stagnation or decl
Analysis
TL;DR
- Former OpenAI researcher Andrew Ho launched a startup focused on high-quality training data, arguing that scaling alone won't achieve true generalization in AI models.
- Existing datasets lack economically valuable skills, especially in bioinformatics and routine lab work, where current models perform poorly (e.g., ~30% success rate in complex analyses).
- Research from Cambridge and Google Deepmind supports this view: LLMs are becoming more specialized rather than versatile, with stagnation or decline in language quality and simple logic despite improvements in programming and math.
- Reinforcement learning works well in domains like code due to clear reward signals and complete data, but fails elsewhere where such data is absent—suggesting early generalization was an artifact of broad text corpus training, not true understanding.
- A structural limitation identified by Google Deepmind’s Tom Zahavy: LLMs excel at deduction/induction but fail at creative abduction (inventing causes without linguistic precedent), pointing toward action-controllable world models as a potential fix.
Why It Matters
This article challenges the prevailing assumption that continued model scale will yield increasingly capable and general AI systems. For practitioners and researchers, it underscores a critical bottleneck: data quality and domain specificity may now be more important than sheer model size. The shift toward specialized, context-rich datasets could redefine R&D priorities, investment strategies, and evaluation metrics across the industry.
Technical Details
- Andrew Ho’s startup targets two initial domains: bioinformatics datasets for complex scientific analysis and everyday lab work involving photo-based experiment evaluation (chemistry, materials science, healthcare follow-up planned).
- Current LLMs achieve only ~30% success rates in advanced bioinformatics tasks, indicating severe gaps in real-world applicability despite large-scale pretraining.
- Adam Hunt observes that reinforcement learning succeeds in coding/math because these areas have unambiguous rewards and full observability—unlike most economic or creative tasks where outcomes aren’t easily graded.
- Google Deepmind’s “LLMs can’t jump” paper identifies a core architectural weakness: inability to perform creative abduction—the generation of novel causal explanations lacking prior linguistic examples—as a barrier to true innovation.
- Proposed solution involves integrating LLMs into action-controllable world models that enable counterfactual experimentation, allowing systems to simulate and evaluate alternative paths beyond pattern matching.
Industry Insight
- Expect a surge in demand for curated, domain-specific training data sets, particularly in life sciences and experimental workflows, potentially creating new market segments worth over $100 billion as predicted by Ho.
- Venture capital and corporate AI investments should pivot from pure model scaling toward data infrastructure, annotation pipelines, and synthetic data generation tailored to niche applications where generalization currently fails.
- Companies relying solely on off-the-shelf LLMs for decision-making in specialized fields face significant risk; hybrid architectures combining symbolic reasoning, simulation environments, and targeted fine-tuning will likely become standard for reliable deployment.
Disclaimer: The above content is generated by AI and is for reference only.