AIToday

Black Forest Labs launches FLUX 3 for image, video, and audio generation

VentureBeat AI6h ago
Black Forest Labs launches FLUX 3 for image, video, and audio generation

Key takeaway

Black Forest Labs, a Freiburg-based AI lab, launched FLUX 3, a multimodal model that generates images and up to 20-second video with audio from a single prompt. Unlike models that combine separate image, video, and audio systems, FLUX 3 is jointly trained across all three modalities and extends to robotic vision and actions. This is the company's first public video generation model and marks a shift toward treating creative generation and robotics as connected applications of a single "visual intelligence" capability.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Black Forest Labs released FLUX 3, a multimodal AI model that generates images and combined audio/video clips up to 20 seconds from a single prompt. The model is jointly trained across image, video, and audio rather than combining separate models. It also extends to robotic vision and actions.

  • Why it matters

    FLUX 3 represents the company's first public video generation model and frames creative generation, simulation, computer use, and robotics as connected applications of a single capability. Black Forest Labs positions this as "visual intelligence"—models that can perceive, predict, and act across physical and digital environments—rather than treating each modality as a separate tool.

  • What to watch

    FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, and FLUX 3 Act. The release is limited to start, meaning broader availability or pricing details may follow.

In Depth

Black Forest Labs, the Freiburg, Germany-based AI lab, unveiled FLUX 3 today, a multimodal frontier model that marks a significant expansion beyond its earlier image-focused FLUX lineup. FLUX 3 is trained to understand and generate images as well as combined audio and video clips up to 20 seconds, all from a single prompt. Critically, the company emphasizes that FLUX 3 is jointly trained across those modalities—image, video, and audio—rather than combining pre-built separate models under one interface. This architectural choice underpins Black Forest Labs' core pitch to enterprises: that creative generation, simulation, computer use, and robotics should be understood as connected applications of a single underlying capability. The company frames this as "visual intelligence," referring to models that can perceive, predict, and act across both physical and digital environments. The release also marks Black Forest Labs' first public video generation model. FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, and FLUX 3 Act. The initial release is limited, indicating the company is rolling out access in phases rather than an immediate, broad launch.

Context & Analysis

Black Forest Labs has positioned FLUX 3 as a departure from the modular approach that has dominated multimodal AI: rather than building independent image, video, and audio generators and bolting them together, the company trained a single architecture jointly across all three domains. This design choice reflects a broader strategic vision where the company sees image generation, video creation, audio synthesis, robotic perception, and robotic control as manifestations of a unified capability—what it calls "visual intelligence." The framing suggests that enterprises should think of these traditionally separate applications as connected rather than siloed, potentially opening new workflows where a single model handles perception and action across physical and digital contexts. The limited initial release indicates the company is managing rollout carefully, likely to gather feedback and validate performance before scaling.

FAQ

What modalities can FLUX 3 handle?
FLUX 3 is trained to understand and generate images, combined audio/video clips up to 20 seconds, and extends to robotic vision and actions.
How is FLUX 3 different from assembling separate models?
FLUX 3 is jointly trained across image, video, and audio modalities rather than assembling separate image, video, and audio models behind a common interface.
Is this Black Forest Labs' first video generation model?
Yes, this release marks FLUX 3 as Black Forest Labs' first public video generation model.

Get the latest Image Generation news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →