
Black Forest Labs has released FLUX 3, a multimodal foundation model that learns from images, audio, and video simultaneously to generate video with embedded physical realism.
Unlike prior video models that treat modalities separately, FLUX 3 uses mutual constraints between sound, motion, and temporal sequences to produce outputs where audio matches impact and motion obeys physics.
The model supports text-to-video, image-to-video morphing, video continuation, and multi-scene editing, and is accessible through Cloudflare AI Gateway.
What happened
Black Forest Labs released FLUX 3, a new multimodal foundation model that learns jointly from images, audio, and video to generate video outputs. The model uses a consolidated architecture where the different modalities inform each other—audio constrains motion, video adds temporality to spatial relationships, and language links perceptions to instructions.
Why it matters
FLUX 3 appears to encode physical laws more directly than prior models, since training on multiple modalities simultaneously creates mutual constraints (sound must match impact, motion must obey mass, future must follow from past). This means users can prompt with either simple or dense descriptions, and the model handles realistic physics without requiring overwrought technical detail.
What to watch
FLUX 3 supports four main capabilities: text-to-video generation, image-to-video transformation using start and end frames, video continuation from an existing clip (via a start_video parameter that preserves momentum and framing), and timestamped multi-scene cuts in a single prompt. It also handles stylized outputs (stop-motion, comic art, VHS-style footage) beyond hyper-realistic defaults. The model is available via Cloudflare AI Gateway without requiring a Black Forest Labs API key.
Black Forest Labs, the independent research lab behind the widely adopted FLUX.1 image model, has released FLUX 3, a multimodal foundation model designed to generate video by learning jointly from images, audio, and video during training. The consolidated architecture is the key differentiator: rather than treating each modality as a separate input stream, FLUX 3 learns the relationships between them. As Black Forest Labs explains in their release, images capture spatial structures at a point in time, video restores the temporal dimension and reveals physical dynamics, audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect, and language links these perceptions to goals and instructions. Learning from one modality produces a good model of that projection; learning from all simultaneously reveals deeper truth—the sound must match the impact, the motion must obey mass, the future must follow from the past, and the modalities become evidence of a single underlying reality rather than separate channels.
The model excels at text-to-video generation, accepting both simple prompts and dense, cinematic descriptions. Users can describe visuals alone or pair them with audio details (e.g., "the roar of the engine and the crunch of sand under tires"), and FLUX 3 integrates both. The model also defaults to multiple scene cuts unless told otherwise, allowing cinematic pacing without manual stitching. For image-to-video, users provide a start frame and an end frame along with a description of the physical transformation—the model then generates a continuous shot showing the natural change between them, preserving camera angle and object identity while morphing conditions (e.g., a rusted junkyard car becoming a restored, roadworthy vehicle). A new parameter, start_video, allows video continuation: users supply an existing clip, and FLUX 3 picks up from the final frame while preserving momentum, framing logic, and audio. Timestamped prompts enable multiple scenes and camera angles in a single generation; users write [0-4s]: wide shot, [4-8s]: close-up, [8-12s]: tracking shot directly into the prompt, and the model cuts between them in one pass. FLUX 3 is not locked into photorealistic hyper-realism; it can generate stop-motion claymation (with visible fingerprints and seams), comic-inked art with chromatic aberration, and vintage VHS home-video footage complete with autofocus hunting, blown highlights, and interlaced noise—all prompted descriptively rather than via style tokens. The model is accessible via Cloudflare AI Gateway, which handles provider credentials and billing through Unified Billing, eliminating the need for a separate Black Forest Labs API key.
Black Forest Labs built FLUX 3 on the insight that learning from multiple modalities simultaneously creates stronger constraints on what the model can generate. The release blog quote highlights the core principle: each modality (images, video, audio, language) captures different aspects of reality, and training on all of them at once reveals their mutual dependencies. A sound without matching physical impact, motion that violates mass, or a future that doesn't follow from the past becomes implausible when the model has learned all these relationships together. This architectural choice appears to reduce the need for over-detailed prompting—users can supply a simple sentence or a dense one, and the model's learned physics fills in convincing motion and realism. The lab's first major success was FLUX.1, an open-source image model that competed with proprietary alternatives. FLUX 3 extends that ambition into video and multimodal generation, positioning Black Forest Labs as an independent research lab tackling capabilities (fast, high-quality, open-source) that larger labs have dominated.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Amazon and Google are intensifying competitive efforts against The Trade Desk (TTD), a major digital advertisi…

OpenAI introduced Premium Seats for ChatGPT Business, priced at $125 per user per month ($100 with annual bill…

Computer scientists at University of Tübingen, Max Planck Institute, MATS Research, and Snyk discovered a meth…

Anthropic pledged to embed machine-readable watermarks in Claude-generated text and digitally signed provenanc…

Anthropic has signed the EU AI Act Code of Practice and will embed invisible watermarks in Claude-generated te…

Anthropic has agreed to pay $9.1 billion over 20 years to Riot Platforms Inc., a Bitcoin miner turned data cen…

The AI news that matters, in one minute each morning.
Sign up free