Model guide

What Is FLUX 3?

FLUX 3 is the multimodal foundation model from Black Forest Labs (BFL), announced on July 23, 2026. Unlike earlier FLUX models that focused on images, FLUX 3 learns from images, video, and audio inside one architecture — so it can generate clips up to 20 seconds long with dialogue, sound effects, and music created natively alongside the picture.

This guide covers what each FLUX 3 capability does, how the Self-Flow training approach works, how FLUX 3 compares with today's leading video models, and how the staged rollout unfolds — plus where to try Flux 3 video generation with native audio.

Official Black Forest Labs collage labeling FLUX 3 image, video, audio, and action modalities around cinematic sample frames

FLUX 3 at a glance

Developer
Black Forest Labs (BFL), the German frontier lab behind the FLUX family
Announced
July 23, 2026, with capabilities rolling out through phased early access
Modalities
Image, video, native audio, language, and action prediction in one model
Video length
Up to 20 seconds per generation, chainable into longer multi-shot sequences
Audio
Native audio generated with the frame: multilingual dialogue, effects, and music
Inputs
Text prompts, start images, visual references, keyframes, and existing video or audio
Availability
FLUX 3 Video and Action in early access now; Image follows in the coming weeks
Open weights
An open-weight multimodal backbone, FLUX 3 Dev, is planned for a later release

One model, four product lines

Every FLUX 3 capability is built from the same underlying multimodal flow-matching model, and each ships after its own early access phase.

FLUX 3 Video

Early access

Text to video, image to video, video to video, and keyframe-controlled transitions — with native audio generated in the same pass. Single generations run up to 20 seconds, and agentic chaining links clips into longer multi-shot stories with consistent characters and style.

FLUX 3 Image

Coming weeks

Image synthesis and editing across styles, compositions, aspect ratios, and resolutions. Early results show stronger complex-prompt handling, more accurate multilingual text rendering, and better editing performance than earlier FLUX generations.

FLUX 3 Action

Partner access

The same foundation applied to physical AI: predicting how actions reshape a scene. FLUX-mimic, built with mimic robotics, already drives dexterous robot manipulation tested in production environments at Audi.

FLUX 3 Dev

Planned

An open-weight multimodal backbone for content creation and action prediction, continuing BFL's tradition of pairing frontier capability with open access. Timing and license details are still to be announced.

How Self-Flow makes it work

FLUX 3 builds on Self-Flow, BFL's approach for aligning multimodal generation and understanding within a single architecture. A multimodal transformer converts images, video, and audio into one shared internal representation, then turns that representation back into output. Trained this way, the model learns how scenes hold together, how objects move, and which sounds belong to which physical events.

That is why the audio is native rather than dubbed on afterwards: picture and sound come out of the same generation, so footsteps land on frames, lips match dialogue, and ambience follows the camera. Video prediction accounts for over 95% of FLUX 3's training compute — which is also what lets the same backbone extend to action prediction for robotics without losing its creative abilities.

Official FLUX 3 architecture diagram showing text, image, video, audio, and action encoders feeding a shared multimodal transformer

How FLUX 3 compares

In BFL's announcement, viewers compared 10-second 720p text-to-video clips with native audio head-to-head. The numbers show how often they preferred FLUX 3 over each competing model.

vs Luma Ray 3.2

93%

vs Runway Gen-4.5

77%

vs Grok Imagine Video

69%

vs Kling v3 Pro

60%

vs Happy Horse v1

59%

vs Happy Horse 1.1

57%

vs Seedance 2.0

52%

vs Gemini Omni Flash

52%

Figures come from Black Forest Labs' internal evaluation of a development-stage FLUX 3 candidate. Sample size, evaluator count, and full methodology have not been published, so treat them as an early signal rather than a final benchmark; BFL says complete results will follow at general availability.

FLUX 3 FAQ

See FLUX 3 in motion

Reading about a video model only goes so far. Open the Flux 3 video generator, write a prompt or drop in an image, and get a cinematic clip with native audio — or browse the official demo reel first.