AIToday
Video GenerationAudio & SpeechTHE DECODERPublished: Jul 24, 2026, 04:02 JST

Black Forest Labs releases Flux 3 with native audio video generation

Black Forest Labs releases Flux 3 with native audio video generation

3 Key Points

  1. What happened

    German AI company Black Forest Labs released Flux 3, a multimodal foundation model that learns from images, video, and audio together. The model can generate videos with native audio up to 20 seconds long for the first time, supporting text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven links between clips.

  2. Why it matters

    In early evaluations using 10-second clips at 720p, Flux 3 was preferred over multiple rivals: 93 percent over Luma Ray 3.2, 77 percent over Runway Gen-4.5, and 69 percent over Grok Imagine Video. Against stronger competitors, it matched or came close to Seedance 2.0 and Gemini Omni Flash (each at 52 percent preference), which already serve major production workflows. BFL also developed Flux-mimic, a video-action model being tested on production tasks at Audi, suggesting the architecture extends to robotics applications.

  3. What to watch

    Flux 3 Image is set to launch in early access within the next few weeks, with improvements to complex prompts and multilingual text rendering. Action prediction will initially be offered through select partners. BFL plans open-weight access to the multimodal backbone under the name 'Flux 3 Dev,' with longer-term work on a single model combining perception, action, and language prediction.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

Black Forest Labs frames Flux 3 as part of a shift toward world models—AI systems that perceive, predict, and act across physical and digital environments. The multimodal design reflects the company's argument that no single data type captures reality fully: images show spatial structure, video reveals how it changes over time, and audio connects mechanical events to their sounds. By training on all three together, the model can fill gaps for one another, gaining more information than training on each modality separately.

The integration of a dedicated action component distinguishes Flux 3 from earlier video-generation tools. BFL's collaboration with Mimic Robotics to develop Flux-mimic, now in production testing at Audi, signals that the architecture is designed to extend beyond content generation into physical robotics—a meaningful step in the broader industry push toward embodied AI. The company's plan to release open-weight access to the multimodal backbone under "Flux 3 Dev" suggests confidence in the architecture and an intent to let the research community and partners build on it.

FAQ
What can Flux 3 video do that earlier models could not?
Flux 3 can generate videos with native audio for the first time, with clips up to 20 seconds long. It supports text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven links between clips for longer multi-shot sequences.
How does Flux 3 compare to rivals in early testing?
In evaluations using 10-second clips at 720p, Flux 3 was preferred over Luma Ray 3.2 in 93 percent of comparisons, over Runway Gen-4.5 in 77 percent, and over Grok Imagine Video in 69 percent. Against stronger competitors, it was preferred over Kling v3 Pro 60 percent of the time, over Happy Horse v1 at 59 percent, over Happy Horse 1.1 at 57 percent, and over both Seedance 2.0 and Gemini Omni Flash at 52 percent each. BFL notes these results are preliminary and no independent tests are available yet.
When will the image generation version be available?
Flux 3 Image is set to launch in early access within the next few weeks, with improvements especially for complex prompts and accurate text rendering in multiple languages.

Get the latest Video Generation news every morning

For example, today's edition would include:

  • Avail launches creatorAPI for AI ad variantsSemafor Tech · 1d ago
  • NBC News airs Chloe Melas interview with AI actress Tilly NorwoodHacker News · 2d ago
  • AppLovin's Growth Wall Is a 30-to-60-Second Video Its AI Can't BuildTop Companies AI · 3d ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleOpenAI AI hacks Hugging Face—but trust in AI labs at historic low