AIToday
Large Language ModelsVideo GenerationReplicate BlogPublished: Aug 10, 2026, 16:01 JST

Black Forest Labs unveils FLUX 3, a multimodal video model learning from images, audio, and video

Black Forest Labs unveils FLUX 3, a multimodal video model learning from images, audio, and video

3 Key Points

  1. What happened

    Black Forest Labs released FLUX 3, a new multimodal foundation model that learns jointly from images, audio, and video to generate video outputs. The model uses a consolidated architecture where the different modalities inform each other—audio constrains motion, video adds temporality to spatial relationships, and language links perceptions to instructions.

  2. Why it matters

    FLUX 3 appears to encode physical laws more directly than prior models, since training on multiple modalities simultaneously creates mutual constraints (sound must match impact, motion must obey mass, future must follow from past). This means users can prompt with either simple or dense descriptions, and the model handles realistic physics without requiring overwrought technical detail.

  3. What to watch

    FLUX 3 supports four main capabilities: text-to-video generation, image-to-video transformation using start and end frames, video continuation from an existing clip (via a start_video parameter that preserves momentum and framing), and timestamped multi-scene cuts in a single prompt. It also handles stylized outputs (stop-motion, comic art, VHS-style footage) beyond hyper-realistic defaults. The model is available via Cloudflare AI Gateway without requiring a Black Forest Labs API key.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Black Forest Labs built FLUX 3 on the insight that learning from multiple modalities simultaneously creates stronger constraints on what the model can generate. The release blog quote highlights the core principle: each modality (images, video, audio, language) captures different aspects of reality, and training on all of them at once reveals their mutual dependencies. A sound without matching physical impact, motion that violates mass, or a future that doesn't follow from the past becomes implausible when the model has learned all these relationships together. This architectural choice appears to reduce the need for over-detailed prompting—users can supply a simple sentence or a dense one, and the model's learned physics fills in convincing motion and realism. The lab's first major success was FLUX.1, an open-source image model that competed with proprietary alternatives. FLUX 3 extends that ambition into video and multimodal generation, positioning Black Forest Labs as an independent research lab tackling capabilities (fast, high-quality, open-source) that larger labs have dominated.

FAQ
What are the main generation modes FLUX 3 offers?
FLUX 3 supports text-to-video (from plain or detailed text prompts), image-to-video (using a start frame and end frame to guide a transformation), video continuation (taking an existing clip and continuing from its final frame while preserving momentum and framing), and timestamped multi-scene cuts (specifying different shots with time codes like [0-4s], [4-8s] in a single prompt for automatic cutting).
How does FLUX 3's multimodal training affect what it generates?
Because FLUX 3 learns from images, audio, and video jointly, the modalities constrain each other: sound must match the impact it represents, motion must obey the mass and physics of objects, and future frames must logically follow from the past. This means the model encodes physical laws more directly, allowing simpler prompts to produce realistic results without requiring dense technical descriptions.
Can FLUX 3 generate styles other than photorealistic video?
Yes. The model can generate stop-motion claymation, comic-inked and hand-drawn styles, Into the Spider-Verse aesthetics, VHS home-video footage, and other stylized outputs. Users describe the visual style and recording conditions directly in the prompt.
Replicate BlogRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • OpenAI, Anthropic probe tens of thousands of AI incidentsTHE DECODER · 2h ago
  • Microsoft folds Word, Excel, PowerPoint into CopilotYahoo Finance AI · 6h ago
  • Google tests Flipkart checkout inside Gemini, AI ModeTechCrunch AI · 9h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAI chip demand strains memory supply through 2027; Inversion eyes ASML with X-ray lithography