Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs
Black Forest Labs released Flux 3, a multimodal foundation model that simultaneously learns from images, video, and audio to generate videos up to 20 seconds long with native audio. The model utilizes the "Self-Flow" architecture, a unified learning approach that outperforms standard flow-matching methods in generation quality and physical world understanding. Early evaluations show Flux 3 preferred over competitors like Luma Ray 3.2 (93%), Runway Gen-4.5 (77%), and Grok Imagine Video (69%) in h
Analysis
TL;DR
- Black Forest Labs released Flux 3, a multimodal foundation model that simultaneously learns from images, video, and audio to generate videos up to 20 seconds long with native audio.
- The model utilizes the "Self-Flow" architecture, a unified learning approach that outperforms standard flow-matching methods in generation quality and physical world understanding.
- Early evaluations show Flux 3 preferred over competitors like Luma Ray 3.2 (93%), Runway Gen-4.5 (77%), and Grok Imagine Video (69%) in human preference tests.
- BFL is developing Flux-mimic for robotics applications, currently being tested at Audi, with plans to release open-weight access via "Flux 3 Dev."
Why It Matters
This release marks a significant shift toward integrated multimodal systems, demonstrating that joint training on visual and auditory data yields superior realism and temporal coherence compared to unimodal or loosely coupled approaches. For AI practitioners, it highlights the growing importance of "world models" that can perceive, predict, and act, bridging the gap between generative media and practical robotics applications.
Technical Details
- Architecture: Built on "Self-Flow," a method teaching a single model to generate and understand content simultaneously using a multimodal transformer with dedicated encoders/decoders for images, video, audio, and actions.
- Capabilities: Supports text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven chaining for longer sequences.
- Performance: In 720p, 10-second clip tests, Flux 3 showed strong preference rates against major rivals, particularly excelling in human facial expressions and sound-event synchronization.
- Robotics Integration: Includes an action prediction component developed with Mimic Robotics for industrial tasks, such as those tested at Audi.
Industry Insight
The integration of native audio and action prediction suggests that future video models will increasingly serve as foundational layers for embodied AI and robotics, rather than just creative tools. Companies should monitor the "Flux 3 Dev" open-weight release to assess how unified multimodal architectures impact downstream application development in simulation and autonomous systems.
Disclaimer: The above content is generated by AI and is for reference only.