AIToday
Video GenerationAI Business & IndustryTHE DECODERPublished: Aug 6, 2026, 01:00 JST3 min read

Black Forest Labs releases FLUX 3 Video, tops text-to-video rankings

Black Forest Labs releases FLUX 3 Video, tops text-to-video rankings

Key takeaway

  • Black Forest Labs has released FLUX 3 Video for general availability, claiming it ranks first in text-to-video generation with an Elo score of 1,135, ahead of competitors like Seedance 2.0.

  • The model generates HD and Full HD clips up to 20 seconds with native audio, lip-synced dialogue in over 14 languages, and supports multiple input modes including text and images.

  • Pricing is usage-based, starting at $0.06 per second for draft-mode text-to-video.

3 Key Points

  1. What happened

    Black Forest Labs made FLUX 3 Video generally available through its API and select partners. The model generates HD and Full HD video clips up to 20 seconds long with native audio (dialogue, sound effects, ambient noise), supports text-to-video, image-to-video, keyframes, video continuation, and generates lip-synced dialogue in more than 14 languages.

  2. Why it matters

    According to Black Forest Labs' own Elo rankings, FLUX 3 scores 1,135 for text-to-video and 1,051 for image-to-video, placing it ahead of Gemini Omni Flash, Minimax H3, and Seedance 2.0. This positions the model as a leading option for video generation among businesses and creators evaluating AI video tools.

  3. What to watch

    Pricing is per-second-of-output. Draft mode costs $0.06 per second (text-to-video or image-to-video) or $0.12 (video-to-video); full HD at top quality costs $0.29 and $0.53 per second respectively, with audio included.

In Depth

Read the full story

Black Forest Labs announced the general availability of FLUX 3 Video, its latest model for video generation, accessible through the BFL API and select partners. The model can generate HD and Full HD video clips up to 20 seconds in length, each with native audio containing dialogue, sound effects, and ambient noise.

FLUX 3 supports multiple workflows: text-to-video conversion, image-to-video generation, keyframe-based control, video continuation, and the ability to include multiple scenes and camera angles within a single output clip. A distinctive feature is its capacity to render text and typography directly into scenes, follow complex user prompts, and apply world knowledge—enabling applications such as documentary production. The model also generates lip-synced dialogue in more than 14 languages, broadening its appeal for international content creation.

According to Black Forest Labs' internal Elo rankings—a comparative scoring system for model outputs—FLUX 3 Video achieves a score of 1,135 in text-to-video generation and 1,051 in image-to-video generation, rankings that place it ahead of Gemini Omni Flash, Minimax H3, and Seedance 2.0. Pricing operates on a per-second-of-output basis. In draft mode (HD quality only), text-to-video and image-to-video generation costs $0.06 per second, while video-to-video costs $0.12 per second. Full quality HD pricing is $0.17 per second for text-to-video or image-to-video, and $0.41 for video-to-video. Full HD quality costs $0.29 per second for text-to-video or image-to-video, and $0.53 for video-to-video. Audio is included at all pricing tiers. Additional video examples are available on the Black Forest Labs blog.

Context & Analysis

Black Forest Labs' release of FLUX 3 Video marks a significant step in the competitive landscape of AI video generation. The company claims its model ranks first in both text-to-video and image-to-video capabilities according to Elo rankings—a scoring system that compares model outputs—with scores of 1,135 and 1,051 respectively, positioning it ahead of established competitors. This ranking system appears to be internal to Black Forest Labs; the body does not identify an independent third-party evaluator.

The model's feature set reflects maturation in the field: support for multiple input modes (text, images, keyframes), lip-synced dialogue across 14+ languages, and the ability to incorporate complex prompts and world knowledge for specialized use cases like documentaries. The pricing structure—graduated by output quality and input type—suggests the company is targeting both price-sensitive (draft mode at $0.06/second) and quality-focused (full HD at $0.53/second for video-to-video) use cases.

FAQ

How long can FLUX 3 Video generate at once?
The model generates HD and Full HD clips up to 20 seconds long.
What input methods does FLUX 3 Video support?
It supports text-to-video, image-to-video, keyframes, video continuation, and multiple scenes and camera angles within a single clip.
How much does FLUX 3 Video cost?
Draft mode costs $0.06 per second for text-to-video or image-to-video, and $0.12 for video-to-video. Full HD at top quality costs $0.29 per second for text-to-video or image-to-video, and $0.53 for video-to-video, with audio included.
In how many languages does FLUX 3 generate lip-synced dialogue?
The model generates lip-synced dialogue in more than 14 languages.

Get the latest Video Generation news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleSnowflake CoCo migrates Spark code to Snowflake in one prompt, delivering up to 5.1x speed gains

The AI news that matters, in one minute each morning.

Sign up free