
What happened
Black Forest Labs released FLUX 3, a new multimodal foundation model that learns jointly from images, audio, and video to generate video outputs. The model uses a consolidated architecture where the different modalities inform each other—audio constrains motion, video adds temporality to spatial relationships, and language links perceptions to instructions.
Why it matters
FLUX 3 appears to encode physical laws more directly than prior models, since training on multiple modalities simultaneously creates mutual constraints (sound must match impact, motion must obey mass, future must follow from past). This means users can prompt with either simple or dense descriptions, and the model handles realistic physics without requiring overwrought technical detail.
What to watch
FLUX 3 supports four main capabilities: text-to-video generation, image-to-video transformation using start and end frames, video continuation from an existing clip (via a start_video parameter that preserves momentum and framing), and timestamped multi-scene cuts in a single prompt. It also handles stylized outputs (stop-motion, comic art, VHS-style footage) beyond hyper-realistic defaults. The model is available via Cloudflare AI Gateway without requiring a Black Forest Labs API key.
Summaries like this, in your inbox every morning.
Black Forest Labs built FLUX 3 on the insight that learning from multiple modalities simultaneously creates stronger constraints on what the model can generate. The release blog quote highlights the core principle: each modality (images, video, audio, language) captures different aspects of reality, and training on all of them at once reveals their mutual dependencies. A sound without matching physical impact, motion that violates mass, or a future that doesn't follow from the past becomes implausible when the model has learned all these relationships together. This architectural choice appears to reduce the need for over-detailed prompting—users can supply a simple sentence or a dense one, and the model's learned physics fills in convincing motion and realism. The lab's first major success was FLUX.1, an open-source image model that competed with proprietary alternatives. FLUX 3 extends that ambition into video and multimodal generation, positioning Black Forest Labs as an independent research lab tackling capabilities (fast, high-quality, open-source) that larger labs have dominated.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
OpenAI and Anthropic are reviewing tens of thousands of incidents in which their AI agents hacked websites, us…

Microsoft folded full Word, Excel, and PowerPoint into Copilot and pushed agent-style features like Home, Chat…

Google is testing a "Buy" button on select Flipkart listings inside Gemini and AI Mode in India, covering smar…

Meta's Muse could challenge Google's $63 billion Search business by replacing searches, clicks, and ads with A…

PKSHA Technology provided the conversational AI agent feature of its AI SaaS "PKSHA ChatAgent" to NTT Docomo's…

CleanTechnica writer Fritz Hasler says Tesla's in-car Grok bot, Ara, offered an unprompted forecast that FSD V…
