
Black Forest Labs has released Flux 3, a multimodal AI model that generates videos with native audio up to 20 seconds long—a first for the company. In early testing, Flux 3 outperformed most rivals in user preference, beating Luma Ray 3.2 in 93 percent of comparisons and matching Seedance 2.0 and Gemini Omni Flash at 52 percent each. The model also powers Flux-mimic, a video-action system currently being tested at Audi for production tasks, while the company plans open-weight access and an image generation upgrade in the coming weeks.
Summaries like this, in your inbox every morning.
Sign up free →What happened
German AI company Black Forest Labs released Flux 3, a multimodal foundation model that learns from images, video, and audio together. The model can generate videos with native audio up to 20 seconds long for the first time, supporting text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven links between clips.
Why it matters
In early evaluations using 10-second clips at 720p, Flux 3 was preferred over multiple rivals: 93 percent over Luma Ray 3.2, 77 percent over Runway Gen-4.5, and 69 percent over Grok Imagine Video. Against stronger competitors, it matched or came close to Seedance 2.0 and Gemini Omni Flash (each at 52 percent preference), which already serve major production workflows. BFL also developed Flux-mimic, a video-action model being tested on production tasks at Audi, suggesting the architecture extends to robotics applications.
What to watch
Flux 3 Image is set to launch in early access within the next few weeks, with improvements to complex prompts and multilingual text rendering. Action prediction will initially be offered through select partners. BFL plans open-weight access to the multimodal backbone under the name 'Flux 3 Dev,' with longer-term work on a single model combining perception, action, and language prediction.
Black Forest Labs, a German AI company, has introduced Flux 3, a multimodal foundation model that learns jointly from images, video, and audio. The key innovation is native audio generation integrated directly into video creation, with support for clips up to 20 seconds long—a capability BFL describes as a first. The model supports multiple workflows: text-to-video, image-to-video, video-to-video, keyframe-based transitions, and multilingual dialogue. An agent-driven linking system also allows clips to be chained together for longer multi-shot sequences. The company notes that Flux 3 is especially strong at capturing human facial expressions and matching sounds to physical events.
Under the hood, Flux 3 is built on Self-Flow, BFL's approach for training a single model to both generate and understand content simultaneously. The architecture uses a multimodal transformer with dedicated encoders and decoders for images, video, audio, and actions. A specialized action component serves as the foundation for robotics applications and can be extended for new uses. BFL argues this unified learning process delivers better results than the previously standard flow-matching method, both in generation quality and in the model's understanding of the physical world.
In early evaluations using 10-second clips at 720p, Flux 3 showed strong preference margins against several rivals. It was preferred over Luma Ray 3.2 in 93 percent of comparisons, over Runway Gen-4.5 in 77 percent, and over Grok Imagine Video in 69 percent. Against more established competitors, the margins tightened: Flux 3 was preferred over Kling v3 Pro 60 percent of the time, over Happy Horse v1 at 59 percent, over Happy Horse 1.1 at 57 percent, and over both Seedance 2.0 and Gemini Omni Flash at 52 percent each. BFL acknowledges these results are preliminary and notes that no independent tests are available yet. Matching Seedance—which has already reached Hollywood—and Gemini Omni Flash would position Flux 3 among the top video models in the field.
BFL is rolling out capabilities in phases. Flux 3 Video is already available to users, while Flux 3 Image is scheduled to arrive in early access within the next few weeks, with improvements targeted at complex prompts and multilingual text rendering. Action prediction will initially be offered only through select partners. Beyond consumer video generation, BFL developed Flux-mimic in collaboration with Mimic Robotics—a video-action model currently undergoing production testing at Audi. The company also plans to release open-weight access to the multimodal backbone as "Flux 3 Dev," and is working on next-generation models that would combine perception, action, and language prediction in a single unified system.
Black Forest Labs frames Flux 3 as part of a shift toward world models—AI systems that perceive, predict, and act across physical and digital environments. The multimodal design reflects the company's argument that no single data type captures reality fully: images show spatial structure, video reveals how it changes over time, and audio connects mechanical events to their sounds. By training on all three together, the model can fill gaps for one another, gaining more information than training on each modality separately.
The integration of a dedicated action component distinguishes Flux 3 from earlier video-generation tools. BFL's collaboration with Mimic Robotics to develop Flux-mimic, now in production testing at Audi, signals that the architecture is designed to extend beyond content generation into physical robotics—a meaningful step in the broader industry push toward embodied AI. The company's plan to release open-weight access to the multimodal backbone under "Flux 3 Dev" suggests confidence in the architecture and an intent to let the research community and partners build on it.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack