AI News AI资讯 6h ago Updated 57m ago 更新于 57分钟前 49

Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs Flux 3 生成带有原生音频的长达 20 秒的视频,这是 Black Forest Labs 的首创

Black Forest Labs released Flux 3, a multimodal foundation model that simultaneously learns from images, video, and audio to generate videos up to 20 seconds long with native audio. The model utilizes the "Self-Flow" architecture, a unified learning approach that outperforms standard flow-matching methods in generation quality and physical world understanding. Early evaluations show Flux 3 preferred over competitors like Luma Ray 3.2 (93%), Runway Gen-4.5 (77%), and Grok Imagine Video (69%) in h Black Forest Labs 发布多模态基础模型 Flux 3,首次实现图像、视频和音频的同步学习与生成。 模型支持生成长达 20 秒且带有原生音频的视频,在早期对比测试中表现优于 Luma Ray 3.2 和 Runway Gen-4.5 等竞品。 采用 Self-Flow 架构,通过统一的多模态 Transformer 将不同模态映射到共享内部表示,提升物理世界理解力。 推出针对机器人应用的 Flux-mimic 视频动作模型,目前已在奥迪(Audi)进行生产任务测试。 BFL 计划分阶段发布能力,并将在未来以“Flux 3 Dev”名称开放权重版本。

75
Hot 热度
65
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Black Forest Labs released Flux 3, a multimodal foundation model that simultaneously learns from images, video, and audio to generate videos up to 20 seconds long with native audio.
  • The model utilizes the "Self-Flow" architecture, a unified learning approach that outperforms standard flow-matching methods in generation quality and physical world understanding.
  • Early evaluations show Flux 3 preferred over competitors like Luma Ray 3.2 (93%), Runway Gen-4.5 (77%), and Grok Imagine Video (69%) in human preference tests.
  • BFL is developing Flux-mimic for robotics applications, currently being tested at Audi, with plans to release open-weight access via "Flux 3 Dev."

Why It Matters

This release marks a significant shift toward integrated multimodal systems, demonstrating that joint training on visual and auditory data yields superior realism and temporal coherence compared to unimodal or loosely coupled approaches. For AI practitioners, it highlights the growing importance of "world models" that can perceive, predict, and act, bridging the gap between generative media and practical robotics applications.

Technical Details

  • Architecture: Built on "Self-Flow," a method teaching a single model to generate and understand content simultaneously using a multimodal transformer with dedicated encoders/decoders for images, video, audio, and actions.
  • Capabilities: Supports text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven chaining for longer sequences.
  • Performance: In 720p, 10-second clip tests, Flux 3 showed strong preference rates against major rivals, particularly excelling in human facial expressions and sound-event synchronization.
  • Robotics Integration: Includes an action prediction component developed with Mimic Robotics for industrial tasks, such as those tested at Audi.

Industry Insight

The integration of native audio and action prediction suggests that future video models will increasingly serve as foundational layers for embodied AI and robotics, rather than just creative tools. Companies should monitor the "Flux 3 Dev" open-weight release to assess how unified multimodal architectures impact downstream application development in simulation and autonomous systems.

TL;DR

  • Black Forest Labs 发布多模态基础模型 Flux 3,首次实现图像、视频和音频的同步学习与生成。
  • 模型支持生成长达 20 秒且带有原生音频的视频,在早期对比测试中表现优于 Luma Ray 3.2 和 Runway Gen-4.5 等竞品。
  • 采用 Self-Flow 架构,通过统一的多模态 Transformer 将不同模态映射到共享内部表示,提升物理世界理解力。
  • 推出针对机器人应用的 Flux-mimic 视频动作模型,目前已在奥迪(Audi)进行生产任务测试。
  • BFL 计划分阶段发布能力,并将在未来以“Flux 3 Dev”名称开放权重版本。

为什么值得看

Flux 3 标志着视频生成从单纯的视觉模拟向包含听觉和物理逻辑的多模态感知迈出了关键一步,为构建真正的“世界模型”提供了技术路径。对于 AI 从业者和行业而言,其原生音频生成能力和在机器人领域的落地尝试(如与奥迪合作),展示了生成式 AI 从内容创作向具身智能和现实世界交互扩展的巨大潜力。

技术解析

  • 多模态同步学习架构:Flux 3 基于名为 Self-Flow 的方法,使用多模态 Transformer 配合专用的编码器和解码器,将图像、视频、音频和动作数据转换为共享的内部表示,从而让不同模态相互补充信息,优于传统的流匹配方法。
  • 原生音频与长视频生成:模型能够生成长达 20 秒的视频片段,并内置原生音频,支持文本到视频、图像到视频、关键帧过渡及多语言对话,特别擅长处理人类面部表情和声音与物理事件的匹配。
  • 基准测试表现:在 720p 分辨率、10 秒片段的早期评估中,Flux 3 在 93% 的情况下优于 Luma Ray 3.2,77% 优于 Runway Gen-4.5,69% 优于 Grok Imagine Video;与 Kling v3 Pro、Seedance 2.0 和 Gemini Omni Flash 等顶尖模型的差距较小(胜率 52%-60%)。
  • 机器人动作预测组件:模型包含专门的动作预测组件,已与 Mimic Robotics 合作开发 Flux-mimic,用于机器人应用,并在奥迪的生产环境中进行测试,验证了其在物理世界行动预测方面的能力。
  • 发布路线图:Flux 3 Video 已上线,Flux 3 Image 即将进入早期访问阶段,动作预测功能初期仅面向合作伙伴,长期计划发布名为“Flux 3 Dev”的开源权重版本,并研发结合感知、动作和语言预测的下一代模型。

行业启示

  • 多模态融合成为视频生成新标准:单纯的视频生成已不足以维持竞争优势,集成音频、物理逻辑和跨模态理解的多模态模型将成为下一代视频生成技术的核心方向。
  • 生成式 AI 向具身智能延伸:Flux 3 在机器人领域的应用表明,视频生成模型正在演变为具备物理世界理解和行动预测能力的“世界模型”,这将加速 AI 在自动驾驶、智能制造等实体场景中的落地。
  • 开源与闭源策略的动态平衡:BFL 采取分阶段发布和最终开放权重的策略,既通过早期访问收集反馈确保安全性,又通过开源吸引开发者生态,这种模式可能成为大型基础模型商业化推广的参考范式。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Video Generation 视频生成 Multimodal 多模态 Product Launch 产品发布