
What happened
German AI company Black Forest Labs released Flux 3, a multimodal foundation model that learns from images, video, and audio together. The model can generate videos with native audio up to 20 seconds long for the first time, supporting text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven links between clips.
Why it matters
In early evaluations using 10-second clips at 720p, Flux 3 was preferred over multiple rivals: 93 percent over Luma Ray 3.2, 77 percent over Runway Gen-4.5, and 69 percent over Grok Imagine Video. Against stronger competitors, it matched or came close to Seedance 2.0 and Gemini Omni Flash (each at 52 percent preference), which already serve major production workflows. BFL also developed Flux-mimic, a video-action model being tested on production tasks at Audi, suggesting the architecture extends to robotics applications.
What to watch
Flux 3 Image is set to launch in early access within the next few weeks, with improvements to complex prompts and multilingual text rendering. Action prediction will initially be offered through select partners. BFL plans open-weight access to the multimodal backbone under the name 'Flux 3 Dev,' with longer-term work on a single model combining perception, action, and language prediction.
Summaries like this, in your inbox every morning.
Black Forest Labs frames Flux 3 as part of a shift toward world models—AI systems that perceive, predict, and act across physical and digital environments. The multimodal design reflects the company's argument that no single data type captures reality fully: images show spatial structure, video reveals how it changes over time, and audio connects mechanical events to their sounds. By training on all three together, the model can fill gaps for one another, gaining more information than training on each modality separately.
The integration of a dedicated action component distinguishes Flux 3 from earlier video-generation tools. BFL's collaboration with Mimic Robotics to develop Flux-mimic, now in production testing at Audi, signals that the architecture is designed to extend beyond content generation into physical robotics—a meaningful step in the broader industry push toward embodied AI. The company's plan to release open-weight access to the multimodal backbone under "Flux 3 Dev" suggests confidence in the architecture and an intent to let the research community and partners build on it.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Treble raised $18 million in a Series A extension led by Paladin Capital Group, with KOMPAS VC, Frumtak Ventur…

Spotify and industry groups are moving to label, restrict or ban AI-generated tracks, after the New York Times…

Tech-media startup Avail launched creatorAPI, a tool letting brands make dozens of AI-edited variants of video…

Google Deepmind released Gemini 3.8 Live and 3.8 Live Extended Thinking via the Gemini API and Google AI Studi…

Google says its products now support more than 300 languages used by over 7 billion people, 86% of the global…

Nuance Labs, a Seattle startup led by ex-Apple researcher Fangchang Ma, closed a $50 million Series A led by L…