MiniMax announced H3 as an API product on July 31, 2026 and published the weights to Hugging Face three days later. It is a 33B-parameter system with native audio generation, and it is served today by SGLang, vLLM, diffusers and ComfyUI.
"Open weights" is doing a lot of work in that sentence. H3 is a three-module system and only the middle module shipped. Here is what you can actually run on your own hardware, what still calls MiniMax's servers, and the licence clause that decides whether any of this is available to you at all.
H3 is an omni-modal generative system: it takes text, images, video and audio as a single multimodal context and generates video with native stereo audio β the video and audio latents are predicted jointly by one transformer, not stitched together afterwards.
| Spec | Value |
|---|---|
| Duration | 4β15 seconds |
| Frame rate | 24 fps |
| Resolution | 768p default; 2K via H3-Regenerate-2K |
| Audio | 32 kHz stereo, generated with the video |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 and others |
| Dialogue languages | 11 with stable support, including English, Chinese, Japanese, Korean, Arabic and Spanish |
Inputs come in two mutually exclusive shapes, which is the detail most summaries drop:
You pick one. Frame conditioning and reference inputs cannot be combined in a single request β they are separate model variants, served on separate endpoints.
A request to the H3 API passes through three stages:
Modules 1 and 3 are not in the release. MiniMax is explicit about why module 1 is missing β it "relies on a multi-stage workflow and multiple hosted models and services" β and equally explicit about what that costs you: H3-Context-IR "is critical to the quality of the final output, so we strongly recommend incorporating it into your generation pipeline."
That recommendation is not decoration. The IR is not a rewritten prompt; it is a structured document with named subjects, a retention analysis saying which parts of each reference must survive, a shot-by-shot description with dialogue inline, and a soundscape section. One published example runs past 33,000 prompt tokens for a five-second clip. You can build your own from the two prompt-writing guides MiniMax ships, but you are then rebuilding the component they identify as the main determinant of output quality.
The same is true of resolution. The card's own "Full 2K Workflow" is a local SGLang deployment calling two MiniMax hosted APIs around it. Self-hosted and offline, H3 is a 768p model.
This is the part that changes whether the rest of the article applies to you. From the MiniMax H3 Community License Agreement:
"Applicable Territory" means worldwide, excluding the Excluded Territories. β¦ "Excluded Territories" means the European Union, the United Kingdom, the Republic of Korea and the United States of America.
The open-weights grant does not extend to those four. Three consequences worth being precise about:
The API is global. MiniMax draws the line at distribution, not geography: when they operate the serving infrastructure they can enforce their own safeguards, so the hosted model is available everywhere the weights are not.
H3-Omni-Transformer is a 33B dense, single-stream transformer. Two details make the practical number smaller than that:
MiniMax's reference SGLang configuration serves each variant across four GPUs with Ulysses sequence parallelism, and the two variants run as separate services on separate ports. SGLang, vLLM, diffusers and ComfyUI all have documented paths, and the repository ships both the original checkpoint and a diffusers conversion so you only download the format your framework needs.
Self-host if you are outside the excluded territories, you have four GPUs per variant to dedicate, and either 768p is genuinely enough or you are prepared to build your own context-processing stage. Fine-tuning is the strongest reason of all β full weights, and the licence permits derivatives within the territory.
Use the API if you are in the EU, UK, Korea or the US and have not applied for a licence; if you need 2K; if you want the official IR quality without rebuilding it; or if your volume does not justify standing GPUs. It is also the only version of H3 that is genuinely one model rather than one model plus an integration project.
H3 is in Imgo's video generator as MiniMax H3, via MiniMax's official v2 API β the full three-module path, so 2K and the official IR are both included, and the territory question does not arise. Two constraints our model registry enforces, both inherited from the model rather than added by us:
H3 has no seed, so it is not the model to reach for when you need a shot to render identically next quarter β Veo 3.1 is the only one of the three majors that gives you one. What H3 gives you instead is the widest input surface of any model we route to: nine reference images, three reference videos and three reference audio clips in a single request, with the soundtrack generated alongside the picture.
If you want to compare it against Seedance 2.5 and Veo 3.1 before committing a pipeline, the three are one dropdown apart and share a request shape and a credit balance.
No image
Seedance 2.5's headline numbers β thirty seconds, fifty references β are registry facts. What they buy you in practice, and where Veo 3.1 still takes the job, is what this review is for.
No image
H3's weights are on Hugging Face, but H3 is a three-module system and only H3-Base shipped. Here is what runs locally, why 2K and the context orchestrator stay hosted, what four GPUs get you, and the territory clause most coverage skipped.
No image
OpenAI switches off the Sora API on September 24, 2026 β the app already closed in April. Here is what actually ends, how Sora 2's capabilities map onto Seedance 2.5 and Veo 3.1, and a checklist for migrating this week.