Introducing FLUX 3 Video
FLUX 3 is Black Forest Labs' first multimodal foundation model. Here is what its video generation can do and why the unified image-video-audio architecture matters.

FLUX 3 Video is the video generation capability inside FLUX 3, Black Forest Labs' new multimodal foundation model. Announced in July 2026 and currently available in early access, FLUX 3 is BFL's first model trained jointly on images, videos, and audio within a single architecture.
A world model, not just a video model
Most video models are optimized to generate pixels. FLUX 3 is built around a different idea: it learns one shared representation of the world from multiple sensory inputs. Images provide spatial structure and relationships at a single moment. Video adds time, motion, and physical laws. Audio adds causal sound relationships that vision alone cannot capture. Language ties these perceptions to instructions and abstractions.
By training on all modalities at once, the model uses their mutual constraints to learn how the world actually behaves. As BFL puts it, the sound has to match the impact, the motion has to obey mass, and the future has to follow from the past. This is what BFL calls real-world visual intelligence, and video is the hardest part of it.
What FLUX 3 Video can do
FLUX 3 can generate highly diverse videos with native audio up to 20 seconds in length from a single generation. Its core video capabilities include:
- Text-to-video — describe a scene and get a complete clip.
- Image-to-video — animate a starting frame, or use images as visual references.
- Video-to-video — carry central elements of a source clip into a new scene or context.
- Generative video-audio continuation — extend an input video and audio with matching sound.
- Keyframe-to-video — control transitions between defined moments.
- Agentic chaining — connect individual clips into longer, multi-shot sequences.
- Broad styles and aspect ratios — from candid camcorder footage to animation and cinematics.
- Multilingual dialogue and animated typography.
A notable detail is that audio is generated natively alongside the video. That means speech can be synchronized to lip movement and sound effects can match the physical events causing them, rather than being added later.
The architecture behind it
FLUX 3 builds on Self-Flow, BFL's method for aligning multimodal generation and understanding within the same underlying architecture. The team scaled up compute and data to train across video, images, and audio simultaneously.
The same backbone that generates content also powers physical AI. In the FLUX-mimic collaboration with mimic robotics, an early version of FLUX 3 drives a video-action model for robots that has been tested and deployed at Audi. BFL notes that video prediction accounts for over 95% of the model's total training compute, which underscores why realistic motion is the central challenge — and why getting it right unlocks both creative and physical applications.
What this means for builders
FLUX 3 is still in early access, so API pricing and general availability are rolling out gradually. If you are building video workflows, it is worth applying for access to test how unified audio-video generation, longer coherent clips, and stronger style control fit into your pipeline.
The key shift is that FLUX 3 is not a narrow video tool. It is a multimodal foundation model where video generation is one expression of a broader understanding of the physical world. For product teams, that means fewer separate models to orchestrate and more coherent outputs across image, video, and sound.