AIToday
Image GenerationVideo GenerationRoboticsLatent SpacePublished: Jul 24, 2026, 16:00 JST

Black Forest Labs launches FLUX 3 multimodal model spanning video, audio, robotics

Black Forest Labs launches FLUX 3 multimodal model spanning video, audio, robotics

3 Key Points

  1. What happened

    Black Forest Labs announced FLUX 3, a unified multimodal model trained in a single architecture that handles image, video, audio generation, and action prediction. The company also unveiled FLUX-mimic, a video-action robotics model built on FLUX 3 and tested with Audi for real factory deployment on a single on-premises GPU.

  2. Why it matters

    FLUX 3 reproduces and extends capabilities previously demonstrated separately by other frontier labs (Gemini Omni, Grok Imagine, Seedance 2.0), with the team committing to open an open-weights developer version. The robotics application signals that video world modeling can directly transfer to robot control, potentially lowering barriers for labs building competitive open models without reliance on closed ecosystems.

  3. What to watch

    FLUX 3 Video is currently in early access. The model's ability to unify image, video, audio, and robotics control in one architecture—trained jointly rather than as separate modules—will shape whether multimodal systems become industry standard or remain fragmented by task.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

Black Forest Labs' FLUX 3 marks a significant shift in how frontier multimodal systems are structured. Rather than assembling image, video, and audio generators as separate modules—as competitors like Google (Gemini Omni), xAI (Grok Imagine), and others have done—FLUX 3 trains all modalities in a single unified architecture. This approach matters because unified training is claimed to enable emergent capabilities: the company argues the same backbone can be extended toward robotics, which FLUX-mimic demonstrates concretely. The robotics instantiation is not merely a proof-of-concept; testing with a real manufacturer (Audi) in production settings signals confidence that video world modeling transfers directly into robot control quality and sample efficiency.

The open-weights commitment also positions FLUX 3 within a broader competitive dynamic around model access. As The Stack v3 release (a 114 TB open code dataset) reinforces in the same news cycle, open datasets and open models materially raise the floor for labs that want to build competitive systems without reliance on closed ecosystems. By promising an open-weights developer version alongside the commercial early-access product, Black Forest Labs is signaling that reproduction and extension of frontier capabilities in open form is now expected, not exceptional.

FAQ
What can FLUX 3 actually do?
FLUX 3 can generate video from text or images, continue video from starting frames (animation), generate video from reference clips while changing context, create audio alongside video, generate keyframe-to-video transitions, produce multilingual dialogue, and support diverse visual styles from camcorder footage to cinematics and animation. It also natively generates audio with every output.
Is FLUX 3 available now?
FLUX 3 Video is in early access. Black Forest Labs announced an open-weights developer version is on the way, though no release date was provided.
How is FLUX-mimic being tested?
FLUX-mimic is a video-action model built on FLUX 3, trained on robot and wearable data. It runs on a single on-premises GPU and is already being tested with Audi in real factory settings for general-purpose dexterity tasks.

Get the latest Image Generation news every morning

For example, today's edition would include:

  • Suzuki's Osamu Suzuki sayings become karuta cards, AI drawsTop Companies AI · 6h ago
  • OpenAI buys Glass Imaging for over $300 millionTechCrunch AI · 7h ago
  • Hitachi unveils selective re-encoding that halves ViT image AI processing timeTop Companies AI · 1d ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleThe 'Sad Wives of AI' — a marriage crisis in tech