AIToday
Open-Source AITHE DECODERPublished: Aug 4, 2026, 01:01 JST4 min read

MiniMax H3 becomes first open AI video model to top ranking

MiniMax H3 becomes first open AI video model to top ranking

Key takeaway

  • MiniMax has released the open weights of its H3 video model, making it the first open-source model to rank at the top of an AI video benchmark — placing first in Video Editing according to Artificial Analysis.

  • The 33-billion-parameter model can process text, images, video, and audio together to generate short clips, though some components like the 2K resolution module remain closed, and commercial use is only permitted for companies making under $20 million in revenue.

3 Key Points

  1. What happened

    MiniMax released the weights of its H3 video model, a 33-billion-parameter model that processes text, images, video, and audio together. Artificial Analysis ranks H3 first in Video Editing, second in Text-to-Video, and third in Image-to-Video — marking the first time an open model has topped a video ranking.

  2. Why it matters

    The open release allows developers to fine-tune H3 on custom footage, characters, or visual styles locally, though the highest resolution available to users (768p in ComfyUI) falls short of the closed 2K module MiniMax retained. Commercial use is restricted to companies earning under $20 million in revenue, which limits adoption among larger enterprises.

  3. What to watch

    H3 generates four- to 15-second clips with stereo sound and accepts up to nine reference images, three video clips, and three audio clips in a single prompt. ByteDance released its closed Seedance 2.5 the same day, which generates 30-second clips with built-in audio.

In Depth

Read the full story

MiniMax, a Chinese AI company, has released the open weights of H3, its video generation model, achieving a historic first: ranking at the top of a video benchmark despite being fully open-source. Artificial Analysis, a model evaluation platform, ranks H3 first in the Video Editing category, second in Text-to-Video, and third in Image-to-Video — positions previously held only by closed proprietary systems.

H3 is built as a 33-billion-parameter multimodal model that processes text, images, video, and audio as inputs in a single unified interface. The model generates short video clips ranging from four to 15 seconds in length with stereo audio included. A single prompt can include up to nine reference images, three separate video clips, and three separate audio clips, giving users substantial flexibility in composition. The open weights enable developers to fine-tune the model on custom footage, specific characters, or particular visual styles — a capability unavailable with purely closed systems.

Yet the open release is partial. MiniMax has withheld two key components: the 2K resolution module, which generates higher-quality output, and H3-Context-IR, the internal system that translates user prompts and reference materials into a structured intermediate format that the model uses internally. When users run the open H3 locally using ComfyUI, a popular open-source interface for AI video tools, output is capped at 768p resolution. Users must manually prepare context using prompting guides that MiniMax has published, adding a workflow step compared to the closed version. On the licensing side, commercial use is permitted only for companies with revenues under $20 million, which excludes larger enterprises from deploying H3 for paid services.

The timing of the H3 release coincides with ByteDance's announcement of Seedance 2.5, a closed-source competitor. Seedance 2.5 generates longer 30-second video clips and includes built-in audio generation, differentiating it from the open H3 on length and audio handling. This parallel release illustrates the ongoing division in the video generation market between open models targeting developers and researchers, and closed systems optimized for commercial deployment.

Context & Analysis

MiniMax's H3 release marks a significant milestone in open-source AI video generation: for the first time, an openly available model ranks at the top of a professional benchmark. Artificial Analysis places H3 first in Video Editing, ahead of both proprietary systems and prior open models, signaling that the performance gap between closed and open video generation is narrowing. The model's multimodal architecture — accepting text, images, video, and audio as inputs — reflects the direction of the broader AI field toward unified input handling.

However, the open release comes with notable limitations that preserve MiniMax's commercial advantages. The 2K resolution module and the H3-Context-IR system, which structures prompts and reference material for the model, remain proprietary. Users running H3 locally are constrained to 768p output and must manually prepare context using published guides. The revenue cap on commercial use ($20 million) further restricts enterprise adoption. These constraints suggest MiniMax views the open release as a way to build developer mindshare and fine-tuning flexibility while protecting higher-margin closed offerings. On the same day, ByteDance's closed Seedance 2.5 — which generates longer 30-second clips — underscores that the competitive landscape remains split between open and proprietary systems, each optimizing for different trade-offs.

FAQ

What video length and resolution can H3 produce?
H3 generates four- to 15-second clips with stereo sound. When running locally in ComfyUI, the resolution tops out at 768p; a closed 2K resolution module is not included in the open release.
Who can use H3 commercially?
Commercial use is only permitted for companies making under $20 million in revenue.
What inputs can a single H3 prompt include?
A single prompt can include up to nine reference images, three video clips, and three audio clips.

Get the latest Open-Source AI news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleSam Altman built $3.3B fortune outside OpenAI via contrarian investing

The AI news that matters, in one minute each morning.

Sign up free