
MiniMax has released the open weights of its H3 video model, making it the first open-source model to rank at the top of an AI video benchmark — placing first in Video Editing according to Artificial Analysis.
The 33-billion-parameter model can process text, images, video, and audio together to generate short clips, though some components like the 2K resolution module remain closed, and commercial use is only permitted for companies making under $20 million in revenue.
What happened
MiniMax released the weights of its H3 video model, a 33-billion-parameter model that processes text, images, video, and audio together. Artificial Analysis ranks H3 first in Video Editing, second in Text-to-Video, and third in Image-to-Video — marking the first time an open model has topped a video ranking.
Why it matters
The open release allows developers to fine-tune H3 on custom footage, characters, or visual styles locally, though the highest resolution available to users (768p in ComfyUI) falls short of the closed 2K module MiniMax retained. Commercial use is restricted to companies earning under $20 million in revenue, which limits adoption among larger enterprises.
What to watch
H3 generates four- to 15-second clips with stereo sound and accepts up to nine reference images, three video clips, and three audio clips in a single prompt. ByteDance released its closed Seedance 2.5 the same day, which generates 30-second clips with built-in audio.
MiniMax, a Chinese AI company, has released the open weights of H3, its video generation model, achieving a historic first: ranking at the top of a video benchmark despite being fully open-source. Artificial Analysis, a model evaluation platform, ranks H3 first in the Video Editing category, second in Text-to-Video, and third in Image-to-Video — positions previously held only by closed proprietary systems.
H3 is built as a 33-billion-parameter multimodal model that processes text, images, video, and audio as inputs in a single unified interface. The model generates short video clips ranging from four to 15 seconds in length with stereo audio included. A single prompt can include up to nine reference images, three separate video clips, and three separate audio clips, giving users substantial flexibility in composition. The open weights enable developers to fine-tune the model on custom footage, specific characters, or particular visual styles — a capability unavailable with purely closed systems.
Yet the open release is partial. MiniMax has withheld two key components: the 2K resolution module, which generates higher-quality output, and H3-Context-IR, the internal system that translates user prompts and reference materials into a structured intermediate format that the model uses internally. When users run the open H3 locally using ComfyUI, a popular open-source interface for AI video tools, output is capped at 768p resolution. Users must manually prepare context using prompting guides that MiniMax has published, adding a workflow step compared to the closed version. On the licensing side, commercial use is permitted only for companies with revenues under $20 million, which excludes larger enterprises from deploying H3 for paid services.
The timing of the H3 release coincides with ByteDance's announcement of Seedance 2.5, a closed-source competitor. Seedance 2.5 generates longer 30-second video clips and includes built-in audio generation, differentiating it from the open H3 on length and audio handling. This parallel release illustrates the ongoing division in the video generation market between open models targeting developers and researchers, and closed systems optimized for commercial deployment.
MiniMax's H3 release marks a significant milestone in open-source AI video generation: for the first time, an openly available model ranks at the top of a professional benchmark. Artificial Analysis places H3 first in Video Editing, ahead of both proprietary systems and prior open models, signaling that the performance gap between closed and open video generation is narrowing. The model's multimodal architecture — accepting text, images, video, and audio as inputs — reflects the direction of the broader AI field toward unified input handling.
However, the open release comes with notable limitations that preserve MiniMax's commercial advantages. The 2K resolution module and the H3-Context-IR system, which structures prompts and reference material for the model, remain proprietary. Users running H3 locally are constrained to 768p output and must manually prepare context using published guides. The revenue cap on commercial use ($20 million) further restricts enterprise adoption. These constraints suggest MiniMax views the open release as a way to build developer mindshare and fine-tuning flexibility while protecting higher-margin closed offerings. On the same day, ByteDance's closed Seedance 2.5 — which generates longer 30-second clips — underscores that the competitive landscape remains split between open and proprietary systems, each optimizing for different trade-offs.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
STATION AI and Denso have launched RE-CORE, an open innovation support program aimed at manufacturers

Wells Fargo analysts, led by Michael Turrin, argue that falling inference costs from rapid open-model developm…

Gregory Kurtzer, founder of CentOS and Rocky Linux, has launched OpenWALDO (Open Weights, Artifacts, Licenses…

CloudSEK identified more than 2,500 organisations potentially exposed by a March 2026 LiteLLM incident, an ope…

After years of broadly deploying AI tools, tech leaders at companies like Samsara, Docusign, Yum Brands, and C…

At the Ai4 conference in Las Vegas, Nobel Prize winner Geoffrey Hinton, World Labs CEO Fei-Fei Li, and Courser…

The AI news that matters, in one minute each morning.
Sign up free