
Black Forest Labs released FLUX 3, a single unified model that generates images, videos, audio, and can predict actions for robot control—capabilities that frontier labs had previously built as separate systems. An open-weights developer version is planned, and the company has already partnered with Audi to test a robotics variant (FLUX-mimic) in real factory settings on consumer-grade hardware.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Black Forest Labs announced FLUX 3, a unified multimodal model trained in a single architecture that handles image, video, audio generation, and action prediction. The company also unveiled FLUX-mimic, a video-action robotics model built on FLUX 3 and tested with Audi for real factory deployment on a single on-premises GPU.
Why it matters
FLUX 3 reproduces and extends capabilities previously demonstrated separately by other frontier labs (Gemini Omni, Grok Imagine, Seedance 2.0), with the team committing to open an open-weights developer version. The robotics application signals that video world modeling can directly transfer to robot control, potentially lowering barriers for labs building competitive open models without reliance on closed ecosystems.
What to watch
FLUX 3 Video is currently in early access. The model's ability to unify image, video, audio, and robotics control in one architecture—trained jointly rather than as separate modules—will shape whether multimodal systems become industry standard or remain fragmented by task.
Black Forest Labs announced FLUX 3 on July 23, 2026, positioning it as a breakthrough in multimodal model design. Unlike most frontier approaches that combine separate image, video, and audio generators, FLUX 3 is trained as one unified architecture spanning image generation, video generation (text-to-video, image-to-video, video-to-video), audio generation, and action prediction. The model handles text-to-video generation, image-to-video generation either as animation or visual reference, video-to-video generation that carries key elements of a source video into new scenes, generative video-audio continuation, keyframe-to-video generation for controlled transitions, and multilingual dialogue. It supports a broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics, and includes strong typography generation and animated designs. Critically, every output includes native audio generation.
The unified training approach is grounded in Self-Flow research and explicitly extends beyond media generation into robotics. Black Forest Labs partnered with Mimic Robotics to develop FLUX-mimic, a video-action model built on the FLUX 3 backbone combined with Mimic's expertise in robot learning for dexterous manipulation. FLUX-mimic is deployable on a single on-premises GPU and is already undergoing testing in real factory settings with Audi, proving that the video world model learned by FLUX 3 can transfer directly into robot control. The team's central claim is that better video world modeling improves robot control quality and sample efficiency without task-specific specialization per manipulator.
FLUX 3 Video is currently available in early access. Black Forest Labs has announced that an open-weights developer version is on the way, though no specific release date was provided. The release prompted significant industry discussion about the competitive landscape: FLUX 3 independently reproduces and claims superiority over capabilities previously demonstrated by Seedance 2.0, Gemini Omni (Google), and Grok Imagine (xAI), while opening the path for broader reproducibility in open-weight form. The timing of the announcement—arriving alongside The Stack v3, a 114 TB open code dataset representing the largest publicly released code corpus—reinforces a theme that open datasets and models are becoming the infrastructure floor for competitive system development, reducing dependence on closed ecosystems.
Black Forest Labs' FLUX 3 marks a significant shift in how frontier multimodal systems are structured. Rather than assembling image, video, and audio generators as separate modules—as competitors like Google (Gemini Omni), xAI (Grok Imagine), and others have done—FLUX 3 trains all modalities in a single unified architecture. This approach matters because unified training is claimed to enable emergent capabilities: the company argues the same backbone can be extended toward robotics, which FLUX-mimic demonstrates concretely. The robotics instantiation is not merely a proof-of-concept; testing with a real manufacturer (Audi) in production settings signals confidence that video world modeling transfers directly into robot control quality and sample efficiency.
The open-weights commitment also positions FLUX 3 within a broader competitive dynamic around model access. As The Stack v3 release (a 114 TB open code dataset) reinforces in the same news cycle, open datasets and open models materially raise the floor for labs that want to build competitive systems without reliance on closed ecosystems. By promising an open-weights developer version alongside the commercial early-access product, Black Forest Labs is signaling that reproduction and extension of frontier capabilities in open form is now expected, not exceptional.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion



Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack