One Model for Video, Image, and Audio
FLUX 3 is trained on video, images, and audio at the same time, inside a single architecture built on Self-Flow. Black Forest Labs calls the approach a real world model: instead of treating each modality separately, it learns how objects hold together, how things move, and how events sound — because no single modality describes the world on its own.
Up to 20 Seconds of Video
FLUX 3 handles text-to-video, image-to-video, and video-to-video, generating clips of up to 20 seconds. That's long enough to carry a full scene rather than a single beat, and every mode runs through the same model instead of a separate pipeline per task.
Native Audio, Generated With the Video
Sound isn't bolted on after the fact. Because FLUX 3 learns audio alongside video during training, it generates both together — so what you hear is produced by the same model that decided what you see, with no separate audio pass or post-production step.
Multi-Style Images and Sharper Text
Alongside video, FLUX 3 generates and edits images across styles, with improved text rendering — the part most image models still get wrong. The same world understanding that drives its video work carries over to stills, from photographic to illustrated.
Overchat AI brings you the power of the world’s top AI models: ChatGPT, Claude, Gemini, Mistral, and more.

What can you create with FLUX 3? Get inspired with these ideas:
Short-Form Video
Generate TikToks, Reels, and Shorts with sound already in the clip — up to 20 seconds from a single prompt, no separate audio pass.
AI Films
Build short scenes with motion and atmosphere that hold together, where the audio comes from the same model that generated the picture.
Image Generation
Create stills across styles — photographic, illustrated, graphic — with the improved text rendering FLUX 3 brings to image work.
Image Editing
Edit existing images by describing the change, keeping the rest of the frame intact — the same multimodal model handles generation and editing.
Product Marketing
Turn a product shot into a moving clip with sound, or generate campaign stills and video from one description of the idea.
Image to Video
Start from a still and let FLUX 3 animate it, or feed it existing footage and reshape that instead — text-to-video, image-to-video, and video-to-video all run through one model.
Create with FLUX 3 in 3 simple steps
Describe What You Want
Write your prompt for a video or an image, or start from a still or a clip you already have.
FLUX 3 Generates It
The model produces your video with its audio, or your image, from the same multimodal network.
Download and use
Get your result ready to share, post, or integrate into your projects.
What is FLUX 3?
FLUX 3 is a multimodal model from Black Forest Labs that generates video, images, and audio from a single architecture. Rather than training a separate model per modality, it learns from video, images, and audio together — an approach Black Forest Labs calls a real world model, one that learns how objects hold together, how things move, and how events sound.
What makes FLUX 3 different from other AI video models?
Most video models generate picture and sound in separate steps. FLUX 3 is trained on video, images, and audio at once, so it produces video together with its audio from the same network. It also covers image generation and editing, and the same video representations are used for action prediction in robotics.
How long can FLUX 3 videos be?
FLUX 3 generates video of up to 20 seconds with native audio. It supports text-to-video, image-to-video, and video-to-video, so you can start from a written description, a still image, or existing footage.
Does FLUX 3 generate audio?
Yes. Audio generation is native and integrated with video rather than added as a post-production step. Because audio was part of training alongside video, the sound comes from the same model that generated the picture.
Who made FLUX 3?
FLUX 3 was built by Black Forest Labs, the team behind the FLUX family of image models. It is built on Self-Flow, their approach for aligning multimodal generation and understanding inside a single architecture.
What can FLUX 3 do beyond video and images?
FLUX 3 also supports action prediction. In a collaboration with mimic, a lightweight action decoder reads the model's video representations to drive robots, and Audi has tested the resulting system on soft-body manipulation tasks such as inserting electronic components and handling flexible materials.