AI News AI资讯 6h ago Updated 1h ago 更新于 1小时前 50

Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut it 前OpenAI研究员押注1000亿美元将流入训练数据,因为单纯扩展无法实现突破

Former OpenAI researcher Andrew Ho launched a startup focused on high-quality training data, arguing that scaling alone won't achieve true generalization in AI models. Existing datasets lack economically valuable skills, especially in bioinformatics and routine lab work, where current models perform poorly (e.g., ~30% success rate in complex analyses). Research from Cambridge and Google Deepmind supports this view: LLMs are becoming more specialized rather than versatile, with stagnation or decl 前OpenAI研究员Andrew Ho认为单纯扩大模型规模无法实现真正的泛化能力,将投入超100亿美元专注于高质量训练数据的构建。 新成立的公司首推生物信息学和日常实验室工作专用数据集,指出现有数据源严重缺失经济价值技能。 剑桥大学与Google Deepmind研究支持该观点,发现当前AI系统正趋向专业化而非通用性,且在创造性问题解决上遇到瓶颈。 LLMs在代码等具备清晰奖励信号的领域表现优异,但在语言质量、简单逻辑等领域出现停滞甚至退化。 Google Deepmind提出LLMs缺乏“创造性溯因”能力,建议结合动作可控的世界模型以突破现有局限。

75
Hot 热度
68
Quality 质量
72
Impact 影响力

Analysis 深度分析

TL;DR

  • Former OpenAI researcher Andrew Ho launched a startup focused on high-quality training data, arguing that scaling alone won't achieve true generalization in AI models.
  • Existing datasets lack economically valuable skills, especially in bioinformatics and routine lab work, where current models perform poorly (e.g., ~30% success rate in complex analyses).
  • Research from Cambridge and Google Deepmind supports this view: LLMs are becoming more specialized rather than versatile, with stagnation or decline in language quality and simple logic despite improvements in programming and math.
  • Reinforcement learning works well in domains like code due to clear reward signals and complete data, but fails elsewhere where such data is absent—suggesting early generalization was an artifact of broad text corpus training, not true understanding.
  • A structural limitation identified by Google Deepmind’s Tom Zahavy: LLMs excel at deduction/induction but fail at creative abduction (inventing causes without linguistic precedent), pointing toward action-controllable world models as a potential fix.

Why It Matters

This article challenges the prevailing assumption that continued model scale will yield increasingly capable and general AI systems. For practitioners and researchers, it underscores a critical bottleneck: data quality and domain specificity may now be more important than sheer model size. The shift toward specialized, context-rich datasets could redefine R&D priorities, investment strategies, and evaluation metrics across the industry.

Technical Details

  • Andrew Ho’s startup targets two initial domains: bioinformatics datasets for complex scientific analysis and everyday lab work involving photo-based experiment evaluation (chemistry, materials science, healthcare follow-up planned).
  • Current LLMs achieve only ~30% success rates in advanced bioinformatics tasks, indicating severe gaps in real-world applicability despite large-scale pretraining.
  • Adam Hunt observes that reinforcement learning succeeds in coding/math because these areas have unambiguous rewards and full observability—unlike most economic or creative tasks where outcomes aren’t easily graded.
  • Google Deepmind’s “LLMs can’t jump” paper identifies a core architectural weakness: inability to perform creative abduction—the generation of novel causal explanations lacking prior linguistic examples—as a barrier to true innovation.
  • Proposed solution involves integrating LLMs into action-controllable world models that enable counterfactual experimentation, allowing systems to simulate and evaluate alternative paths beyond pattern matching.

Industry Insight

  • Expect a surge in demand for curated, domain-specific training data sets, particularly in life sciences and experimental workflows, potentially creating new market segments worth over $100 billion as predicted by Ho.
  • Venture capital and corporate AI investments should pivot from pure model scaling toward data infrastructure, annotation pipelines, and synthetic data generation tailored to niche applications where generalization currently fails.
  • Companies relying solely on off-the-shelf LLMs for decision-making in specialized fields face significant risk; hybrid architectures combining symbolic reasoning, simulation environments, and targeted fine-tuning will likely become standard for reliable deployment.

TL;DR

  • 前OpenAI研究员Andrew Ho认为单纯扩大模型规模无法实现真正的泛化能力,将投入超100亿美元专注于高质量训练数据的构建。
  • 新成立的公司首推生物信息学和日常实验室工作专用数据集,指出现有数据源严重缺失经济价值技能。
  • 剑桥大学与Google Deepmind研究支持该观点,发现当前AI系统正趋向专业化而非通用性,且在创造性问题解决上遇到瓶颈。
  • LLMs在代码等具备清晰奖励信号的领域表现优异,但在语言质量、简单逻辑等领域出现停滞甚至退化。
  • Google Deepmind提出LLMs缺乏“创造性溯因”能力,建议结合动作可控的世界模型以突破现有局限。

为什么值得看

本文揭示了大模型发展路径的关键分歧:从“规模崇拜”转向“数据质量驱动”,对AI从业者理解未来算力与数据投资优先级具有重要参考价值。同时,学术界对LLM泛化能力的质疑加深,提示行业需警惕过度依赖单一技术路线的风险。

技术解析

  • Andrew Ho指出当前主流训练数据难以编码高度情境化的经济技能,即使观察到人类“最优路径”,也缺乏评估替代方案优劣的能力,导致模型泛化受限。
  • Cambridge Adam Hunt通过对比分析发现,强化学习在编程等任务中有效是因为存在明确反馈信号,而文本语料带来的泛化效应只是偶然副产品,并非本质理解。
  • Google Deepmind Tom Zahavy提出LLMs擅长演绎与归纳但欠缺创造性溯因(即无先例的新因果推断),并建议引入支持反事实实验的动作可控世界模型作为补充架构。
  • 生物信息学领域实测显示,即便GPT-5.6 Sol等先进模型在处理复杂科学分析时成功率仅约30%,凸显专业领域数据稀缺的严峻性。
  • 未来计划扩展至化学、材料科学、医疗健康及通用知识工作场景,强调需构建可验证、多路径标注的精细化数据集以支撑可靠推理。

行业启示

  • AI实验室应重新评估资本分配策略,预计未来十年将有超过100亿美元流向垂直领域高质量数据采集与标注,而非盲目堆砌参数规模。
  • 前沿机构如OpenAI、Anthropic若持续依赖烧钱换性能模式,将面临盈利困境;企业级应用落地更需依赖特定场景下的精准数据闭环。
  • 通用智能的实现可能不再依赖单体模型膨胀,而是走向“专用模型+世界模拟器”的混合架构,推动跨学科协作成为研发新焦点。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Training 训练 Dataset 数据集 Research 科学研究 LLM 大模型