AI News AI资讯 5h ago Updated 3h ago 更新于 3小时前 54

Simulation: the new Scaling Law — Joon Sung Park, Simile AI 模拟:新的扩展定律——Joon Sung Park,Simile AI

Simile AI raised a $2B Series B backed by GreenOaks, Index Ventures, Fei-Fei Li, and Andrej Karpathy, marking the "Second Summer of Simulation" following the 2023 Generative Agents paper The company builds behavioral foundation models ("social physics") that simulate human behavior at 85-99% accuracy compared to human focus groups, used by Fortune 100 clients like CVS Joon Sung Park's approach combines long-form interviews, observational/transaction data, randomized controlled trials, and post-t Simile AI完成20亿美元B轮融资,由GreenOaks和Index Ventures领投,Fei-Fei Li和Andrej Karpathy等知名投资人参投,标志着"模拟AI"从学术概念走向商业落地 创始人Joon Sung Park延续2023年Smallville/Generative Agents研究,提出"Simulation is the new Scaling Law",用行为基础模型替代传统LLM预测 技术核心:通过长访谈、观察数据、交易记录、随机对照试验构建因果机制模型,数字孪生达到85%行为准确率(接近人类自我预测的90%+) 已为CVS等财富100强客户运行数千万

82
Hot 热度
72
Quality 质量
75
Impact 影响力

Analysis 深度分析

TL;DR

  • Simile AI raised a $2B Series B backed by GreenOaks, Index Ventures, Fei-Fei Li, and Andrej Karpathy, marking the "Second Summer of Simulation" following the 2023 Generative Agents paper
  • The company builds behavioral foundation models ("social physics") that simulate human behavior at 85-99% accuracy compared to human focus groups, used by Fortune 100 clients like CVS
  • Joon Sung Park's approach combines long-form interviews, observational/transaction data, randomized controlled trials, and post-training on causal decision mechanisms rather than relying on prompting frontier LLMs
  • The technology aims to scale from simulating 1,000 individuals to all 8 billion people, enabling pre-deployment testing of products, policies, UBI, climate strategies, and democratic stability
  • Key insight: rational-optimized models fail to capture real human behavior; accurate simulation requires reproducing human biases and irrationality through weight-level training, not just prompting

Why It Matters

This represents a paradigm shift from using LLMs as text generators to using them as behavioral simulators—fundamentally changing how companies and governments can test decisions before acting. For AI practitioners, it signals that the next frontier isn't just bigger models but deeper models of human behavior, with significant commercial validation already underway.

Technical Details

  • Data pipeline: Combines long-form interviews, observational data, transaction records, and randomized controlled trials to build population-level and individual-level behavioral models
  • Post-training approach: Models are fine-tuned on the causal mechanisms behind human decisions rather than relying on in-context prompting of frontier LLMs; this weight-level training is essential for reproducing irrational behavior and biases
  • Evaluation methodology: Digital twins of 1,000 real people achieved 85% behavioral accuracy (how accurately people reproduce their own responses), with Fortune 100 simulations reaching 85-99% accuracy versus human focus groups
  • Scaling ambition: Moving from individual-level to population-level to society-scale multi-agent simulations, with the long-term goal of simulating all 8 billion people on Earth—potentially requiring data-center-scale compute
  • Architectural lineage: Evolved from the Smallville/Generative Agents project (memory architectures, Markdown-based state, emergent social behaviors) toward foundation models of human behavior

Industry Insight

  • The $2B valuation and Fortune 100 adoption signal that synthetic populations will disrupt market research, policy testing, and product development—replacing expensive human panels with scalable simulation
  • The distinction between prediction and simulation is strategic: companies should invest in understanding how to shape outcomes through simulation rather than merely forecasting them
  • The emphasis on post-training over prompting suggests a new competitive moat: behavioral models trained on proprietary RCT and transaction data will be harder to replicate than general-purpose LLM capabilities

TL;DR

  • Simile AI完成20亿美元B轮融资,由GreenOaks和Index Ventures领投,Fei-Fei Li和Andrej Karpathy等知名投资人参投,标志着"模拟AI"从学术概念走向商业落地
  • 创始人Joon Sung Park延续2023年Smallville/Generative Agents研究,提出"Simulation is the new Scaling Law",用行为基础模型替代传统LLM预测
  • 技术核心:通过长访谈、观察数据、交易记录、随机对照试验构建因果机制模型,数字孪生达到85%行为准确率(接近人类自我预测的90%+)
  • 已为CVS等财富100强客户运行数千万次模拟,准确率85-99%超越人类焦点小组,验证了合成人口替代传统市场调研的可行性
  • 长期愿景:模拟全球80亿人口,应用于气候变化、民主稳定性、UBI等社会级问题,AGI与模拟技术被视为先进文明的双生技术

为什么值得看

这篇文章揭示了AI发展范式的潜在转变——从"预测下一个token"到"模拟人类行为",为从业者提供了理解下一代AI基础设施的框架。Simile AI的融资规模和技术验证表明,行为模拟正成为AI落地商业场景的新基础设施,值得密切关注其技术路线和应用拓展。

技术解析

  • 数据与训练方法:采用多源数据融合策略,包括深度访谈、观察数据、交易记录,结合随机对照试验(RCT)进行后训练,重点学习人类决策的因果机制而非表面相关性
  • 模型架构:区分人口级模型(population-level)和个人级模型(individual-level),前者捕捉群体行为规律,后者构建高保真数字孪生,两者结合实现从微观到宏观的行为涌现
  • 评估体系:以85%行为准确率为里程碑(对比人类自我预测约90%+准确率),强调模拟应复现人类偏见和错误而非追求"理性最优",避免过度优化导致的失真
  • 规模化路径:提出模拟扩展定律(Scaling Laws for Simulation),预测未来模拟整个社会可能需要整个数据中心规模,涉及多智能体交互和涌现行为建模
  • 与传统LLM的区别:强调仅靠提示词工程无法捕捉真实人类行为,需要改变模型权重进行专门训练,web数据更多反映"人们说什么"而非"人们实际做什么"

行业启示

  • 市场研究范式革命:合成人口可替代昂贵的人类焦点小组,为消费品、政策制定提供低成本、高可扩展的测试环境,传统调研行业面临结构性颠覆
  • AI基础设施新赛道:行为基础模型(Behavioral Foundation Models)与语言/视觉基础模型形成互补,"社会物理学"可能成为下一波AI投资热点
  • 战略建议:企业应尽早建立内部模拟能力或与Simile类平台合作,在产品发布和政策实施前进行虚拟测试,降低试错成本;同时关注模拟技术在社会治理、公共政策领域的潜在应用

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Agent Agent Evaluation 评测 Funding 融资 LLM 大模型