AIToday
Large Language ModelsVideo GenerationReplicate BlogPublished: Aug 10, 2026, 16:01 JST4 min read

Black Forest Labs unveils FLUX 3, a multimodal video model learning from images, audio, and video

Black Forest Labs unveils FLUX 3, a multimodal video model learning from images, audio, and video

Key takeaway

  • Black Forest Labs has released FLUX 3, a multimodal foundation model that learns from images, audio, and video simultaneously to generate video with embedded physical realism.

  • Unlike prior video models that treat modalities separately, FLUX 3 uses mutual constraints between sound, motion, and temporal sequences to produce outputs where audio matches impact and motion obeys physics.

  • The model supports text-to-video, image-to-video morphing, video continuation, and multi-scene editing, and is accessible through Cloudflare AI Gateway.

3 Key Points

  1. What happened

    Black Forest Labs released FLUX 3, a new multimodal foundation model that learns jointly from images, audio, and video to generate video outputs. The model uses a consolidated architecture where the different modalities inform each other—audio constrains motion, video adds temporality to spatial relationships, and language links perceptions to instructions.

  2. Why it matters

    FLUX 3 appears to encode physical laws more directly than prior models, since training on multiple modalities simultaneously creates mutual constraints (sound must match impact, motion must obey mass, future must follow from past). This means users can prompt with either simple or dense descriptions, and the model handles realistic physics without requiring overwrought technical detail.

  3. What to watch

    FLUX 3 supports four main capabilities: text-to-video generation, image-to-video transformation using start and end frames, video continuation from an existing clip (via a start_video parameter that preserves momentum and framing), and timestamped multi-scene cuts in a single prompt. It also handles stylized outputs (stop-motion, comic art, VHS-style footage) beyond hyper-realistic defaults. The model is available via Cloudflare AI Gateway without requiring a Black Forest Labs API key.

In Depth

Read the full story

Black Forest Labs, the independent research lab behind the widely adopted FLUX.1 image model, has released FLUX 3, a multimodal foundation model designed to generate video by learning jointly from images, audio, and video during training. The consolidated architecture is the key differentiator: rather than treating each modality as a separate input stream, FLUX 3 learns the relationships between them. As Black Forest Labs explains in their release, images capture spatial structures at a point in time, video restores the temporal dimension and reveals physical dynamics, audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect, and language links these perceptions to goals and instructions. Learning from one modality produces a good model of that projection; learning from all simultaneously reveals deeper truth—the sound must match the impact, the motion must obey mass, the future must follow from the past, and the modalities become evidence of a single underlying reality rather than separate channels.

The model excels at text-to-video generation, accepting both simple prompts and dense, cinematic descriptions. Users can describe visuals alone or pair them with audio details (e.g., "the roar of the engine and the crunch of sand under tires"), and FLUX 3 integrates both. The model also defaults to multiple scene cuts unless told otherwise, allowing cinematic pacing without manual stitching. For image-to-video, users provide a start frame and an end frame along with a description of the physical transformation—the model then generates a continuous shot showing the natural change between them, preserving camera angle and object identity while morphing conditions (e.g., a rusted junkyard car becoming a restored, roadworthy vehicle). A new parameter, start_video, allows video continuation: users supply an existing clip, and FLUX 3 picks up from the final frame while preserving momentum, framing logic, and audio. Timestamped prompts enable multiple scenes and camera angles in a single generation; users write [0-4s]: wide shot, [4-8s]: close-up, [8-12s]: tracking shot directly into the prompt, and the model cuts between them in one pass. FLUX 3 is not locked into photorealistic hyper-realism; it can generate stop-motion claymation (with visible fingerprints and seams), comic-inked art with chromatic aberration, and vintage VHS home-video footage complete with autofocus hunting, blown highlights, and interlaced noise—all prompted descriptively rather than via style tokens. The model is accessible via Cloudflare AI Gateway, which handles provider credentials and billing through Unified Billing, eliminating the need for a separate Black Forest Labs API key.

Context & Analysis

Black Forest Labs built FLUX 3 on the insight that learning from multiple modalities simultaneously creates stronger constraints on what the model can generate. The release blog quote highlights the core principle: each modality (images, video, audio, language) captures different aspects of reality, and training on all of them at once reveals their mutual dependencies. A sound without matching physical impact, motion that violates mass, or a future that doesn't follow from the past becomes implausible when the model has learned all these relationships together. This architectural choice appears to reduce the need for over-detailed prompting—users can supply a simple sentence or a dense one, and the model's learned physics fills in convincing motion and realism. The lab's first major success was FLUX.1, an open-source image model that competed with proprietary alternatives. FLUX 3 extends that ambition into video and multimodal generation, positioning Black Forest Labs as an independent research lab tackling capabilities (fast, high-quality, open-source) that larger labs have dominated.

FAQ

What are the main generation modes FLUX 3 offers?
FLUX 3 supports text-to-video (from plain or detailed text prompts), image-to-video (using a start frame and end frame to guide a transformation), video continuation (taking an existing clip and continuing from its final frame while preserving momentum and framing), and timestamped multi-scene cuts (specifying different shots with time codes like [0-4s], [4-8s] in a single prompt for automatic cutting).
How does FLUX 3's multimodal training affect what it generates?
Because FLUX 3 learns from images, audio, and video jointly, the modalities constrain each other: sound must match the impact it represents, motion must obey the mass and physics of objects, and future frames must logically follow from the past. This means the model encodes physical laws more directly, allowing simpler prompts to produce realistic results without requiring dense technical descriptions.
Can FLUX 3 generate styles other than photorealistic video?
Yes. The model can generate stop-motion claymation, comic-inked and hand-drawn styles, Into the Spider-Verse aesthetics, VHS home-video footage, and other stylized outputs. Users describe the visual style and recording conditions directly in the prompt.
Replicate BlogRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAI chip demand strains memory supply through 2027; Inversion eyes ASML with X-ray lithography

The AI news that matters, in one minute each morning.

Sign up free