FLUX 3 Reaches General Availability, Adding 'Draft Mode' to Its Unified Multimodal Model

The development of generative models for visual content has moved through distinct stages over the past several years. 

Early systems focused almost exclusively on producing still images from text descriptions. Later iterations improved resolution, prompt adherence, and editing control. More recent efforts have begun treating motion and sound as integral rather than secondary elements. 

This progression reflects a growing recognition that real world scenes involve spatial structure, temporal change, and accompanying audio all at once.

Black Forest Labs entered this landscape in 2024 with a family of image generation models. Those systems established a strong reputation for quality and flexibility in text to image and image to image tasks.

In late July 2026, the laboratory announced FLUX 3, describing it as a multimodal foundation model trained jointly on images, video, and audio within a single architecture. 

The announcement positioned video generation as the initial public component, with image generation, action prediction for robotics, and eventual open weight versions planned to follow. 

At that time the video capabilities entered a gated early access program available only to approved applicants. 

Samples and limited partner integrations appeared, but broad API access was not yet open.

On August 4 the company moved the generation features of FLUX 3 Video into general availability. 

The model can now be accessed through the Black Forest Labs API and selected partner platforms. It produces clips lasting up to twenty seconds at HD and Full HD resolutions. Native audio is generated together with the visual frames rather than added afterward. Supported modes include text to video generation from natural language prompts, image to video conversion that accepts one or more keyframes to guide composition and motion, and continuation of short existing video clips while preserving movement, camera behavior, and audio continuity. 

The system also handles multilingual speech with corresponding lip movements and can maintain coherence across multiple shots within a single generation.

A practical addition in the general release is 'Draft mode.' 

This option allows a user to produce a rapid lower cost preview of a given prompt. 

The preview runs at a reduced price point and returns a quick indication of how the model interprets the request. 

Once the draft captures the intended subjects, framing, and motion, the same parameters can be submitted again for a full quality render. The higher fidelity version retains the core elements established in the draft while applying the complete computational resources. 

Pricing for the draft step is substantially lower than the full generation rate, which makes iterative exploration more accessible before committing to final output.

FLUX 3 differs from the earlier FLUX.1 and FLUX.2 models in its fundamental training approach. 

Previous systems optimized primarily for image data. 

The new model allocates the majority of its training compute to video prediction while incorporating image and audio signals within the same network. This joint training is intended to create stronger internal representations of how objects move, how scenes evolve over time, and how sound relates to visual events. 

The company has indicated that the same backbone will later support more advanced image synthesis and editing as well as action prediction for physical systems, including early work already underway with robotics partners.

The August release therefore represents the transition of one component of the larger FLUX 3 system from restricted testing into open commercial availability. 

Documentation and API endpoints are now public, and the draft to full quality workflow is explicitly supported. 

Higher resolutions beyond Full HD, additional reference capabilities, and open weight versions remain scheduled for later stages of the rollout. 

In the meantime the general availability of FLUX 3 Video provides developers and creators with a unified multimodal model that generates synchronized video and audio from a range of input types while offering a cost efficient path for refining results through Draft mode.

Published