Research Papers 论文研究 3h ago Updated 1h ago 更新于 1小时前 55

CG-World: A Large-Scale World-State Dataset and Protocol for World Models CG-World:面向世界模型的大规模世界状态数据集与协议

CG-World is a large-scale world-state dataset derived from industrial computer graphics pipelines, containing 850,000 temporally aligned segments of 1-5 seconds. It explicitly records intermediate states including multimodal semantics, spatial structure, skeletal and controller states, motion curves, camera and lighting parameters, physics caches, contact events, and multi-pass renderings. The dataset supports intervention learning and counterfactual reasoning through defined branch lineage cove 提出CG-World,一个基于工业计算机图形学管道生成的世界状态数据集与协议。 记录多模态语义、空间结构、骨骼/控制器状态、运动曲线、相机/光照参数、物理缓存、接触事件及多通道渲染等中间状态。 v1版本包含约85万段1–5秒时间对齐样本,支持干预学习与反事实推理的分支谱系标注。 在几何条件视频生成、动作预测和闭环视觉-语言-行动策略迁移任务中验证有效性。 旨在构建面向世界模型、物理AI与具身智能的统一数据基础设施。

72
Hot 热度
85
Quality 质量
78
Impact 影响力

Analysis 深度分析

TL;DR

  • CG-World is a large-scale world-state dataset derived from industrial computer graphics pipelines, containing 850,000 temporally aligned segments of 1-5 seconds.
  • It explicitly records intermediate states including multimodal semantics, spatial structure, skeletal and controller states, motion curves, camera and lighting parameters, physics caches, contact events, and multi-pass renderings.
  • The dataset supports intervention learning and counterfactual reasoning through defined branch lineage covering factual trajectories, observation interventions, action interventions, mechanism interventions, and strict counterfactual branches.
  • Evaluation demonstrates utility for geometry-conditioned video generation, action prediction, and closed-loop vision-language-action policy transfer.
  • Aims to establish shared data infrastructure for world models, Physical AI, and embodied intelligence.

Why It Matters

This dataset addresses a critical gap in existing resources by providing comprehensive, structured supervision that captures the joint dynamics of states, actions, events, and observations—something current video, robotics, and simulation datasets fail to do comprehensively. For researchers working on world models, embodied AI, and physical intelligence, CG-World offers unprecedented granularity for training agents that can reason about complex environments with precise causal understanding. The inclusion of intervention capabilities enables more robust testing of model generalization and counterfactual reasoning, which are essential for real-world deployment.

Technical Details

  • Dataset composition: Approximately 850,000 temporally aligned segments ranging from 1-5 seconds each, extracted from industrial computer graphics production pipelines.
  • Recorded modalities include: multimodal semantics, spatial structure representations, skeletal and controller states, motion curves, camera parameters, lighting configurations, physics cache information, contact event logs, and multi-pass renderings.
  • Structural organization separates latent states, observations, relations, events, and branch metadata into unified spatiotemporal samples for coherent modeling.
  • Intervention framework defines five categories: factual trajectories (baseline behavior), observation interventions (altered sensory inputs), action interventions (modified agent behaviors), mechanism interventions (changed environmental rules), and strict counterfactual branches (alternative outcomes under different conditions).
  • Each intervention type specifies targets, invariants (what remains unchanged), and alternative outcomes to support rigorous evaluation of causal reasoning capabilities.
  • Evaluated applications include geometry-conditioned video generation, future action prediction, and transferring policies between vision-language-action systems in closed-loop settings.

Industry Insight

The creation of CG-World signals a shift toward more comprehensive, industrially sourced datasets that capture full environmental complexity rather than isolated aspects—a trend likely to accelerate as world models advance beyond narrow task-specific capabilities. Organizations developing embodied agents or physical AI should prioritize integrating such richly annotated datasets early in their pipeline development to build systems capable of robust counterfactual reasoning and adaptive behavior. Future efforts will likely focus on expanding community collaboration around standardized world-state representations, potentially leading to benchmark suites analogous to ImageNet but specifically designed for evaluating holistic environmental understanding and interactive decision-making in simulated and real-world contexts.

TL;DR

  • 提出CG-World,一个基于工业计算机图形学管道生成的世界状态数据集与协议。
  • 记录多模态语义、空间结构、骨骼/控制器状态、运动曲线、相机/光照参数、物理缓存、接触事件及多通道渲染等中间状态。
  • v1版本包含约85万段1–5秒时间对齐样本,支持干预学习与反事实推理的分支谱系标注。
  • 在几何条件视频生成、动作预测和闭环视觉-语言-行动策略迁移任务中验证有效性。
  • 旨在构建面向世界模型、物理AI与具身智能的统一数据基础设施。

为什么值得看

该工作填补了现有视频、机器人或仿真数据集仅捕捉部分世界结构的空白,通过提供细粒度、结构化且可干预的世界状态数据,为训练具备因果理解与规划能力的通用世界模型奠定基础。对从事具身智能、强化学习、视频生成及物理建模的研究者具有重要参考价值。

技术解析

  • 数据来源:从工业CG生产管线中提取,确保高保真度与物理一致性,涵盖真实感渲染所需的全链路中间变量。
  • 数据结构:将潜在状态、观测值、关系、事件和分支元数据分离并组织成统一时空样本,支持跨模态对齐与查询。
  • 干预机制:定义五类分支(事实轨迹、观测干预、动作干预、机制干预、严格反事实),明确记录干预目标、不变量与替代结果,支持因果推断实验。
  • 规模与时长:v1含850,000个1–5秒片段,覆盖多样化场景与交互模式,适合长序列建模与策略泛化评估。
  • 应用验证:在几何引导视频生成、未来动作预测以及从模拟到现实的闭环策略迁移任务中表现优异,证明其作为监督信号的有效性。

行业启示

  • 推动“世界模型”从纯像素级预测转向基于结构化状态的因果建模,促进AI系统对环境动态的理解与可控生成能力。
  • 为具身智能提供标准化、可扩展的数据范式,加速仿真到现实(Sim-to-Real)迁移,降低真实世界数据采集成本。
  • 鼓励社区共建共享数据生态,未来可通过持续扩展与协作形成类似ImageNet级别的基准设施,支撑下一代AI系统研发。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Dataset 数据集 Research 科学研究 Multimodal 多模态 Robotics 机器人