SGLang
The SGLang cookbook documents multi-GPU serving for MiniMax-H3.
This reference covers the open base checkpoint for an omni-modal video generation system, with higher-resolution modules remaining hosted. Its community license excludes the European Union, the United Kingdom, the Republic of Korea, and the United States of America.
MiniMax H3 is the vendor reference for MiniMaxAI/MiniMax-H3, as of 2026-08-12. Omni Transformer: H3-Omni-Transformer is a 33B-parameter dense, single-stream Transformer, with approximately 13B parameters residing in AdaLN-related branches.; Output duration: 4-15 seconds, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | MiniMaxAI/MiniMax-H3 | SourceRetrieved 2026-08-12 |
| Identity | Vendor description | MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. | SourceRetrieved 2026-08-12 |
| Identity | Launch publication date | 2026-07-31 | SourceRetrieved 2026-08-12 |
| Identity | Open-source publication date | 2026-08-03 | SourceRetrieved 2026-08-12 |
| Identity | License-stated date | MiniMax H3 release date/License date: August 2, 2026. | SourceRetrieved 2026-08-12 |
| Architecture | Omni Transformer | H3-Omni-Transformer is a 33B-parameter dense, single-stream Transformer, with approximately 13B parameters residing in AdaLN-related branches. | SourceRetrieved 2026-08-12 |
| Architecture | AdaLN deployment | AdaLN-related branches are cacheable for inference-only deployment. | SourceRetrieved 2026-08-12 |
| Architecture | Position encoding | 3D MM-RoPE over (t, h, w) | SourceRetrieved 2026-08-12 |
| Architecture | Text and vision encoder | uses the full pretrained weights of Qwen3-VL-32B; layer-50 hidden states | SourceRetrieved 2026-08-12 |
| Architecture | Video VAE | f16t4d24; temporally causal | SourceRetrieved 2026-08-12 |
| Architecture | Audio VAE | 32 kHz stereo to 40 Hz latents per channel | SourceRetrieved 2026-08-12 |
| Architecture | Joint prediction | joint audio and video prediction in one transformer | SourceRetrieved 2026-08-12 |
| Architecture | Initial attention path | The initial open-source release provides inference with full attention only. | SourceRetrieved 2026-08-12 |
| Architecture | Released checkpoints | The released checkpoints are CFG-distilled Omni Transformer model weights. | SourceRetrieved 2026-08-12 |
| Context | Output duration | 4-15 seconds | SourceRetrieved 2026-08-12 |
| Context | Output frame rate | 24 FPS | SourceRetrieved 2026-08-12 |
| Context | Output audio | 32 kHz stereo | SourceRetrieved 2026-08-12 |
| Context | Output resolution | Shorter side 768 pixels by default; 2K generation can be achieved with H3-Regenerate-2K. | SourceRetrieved 2026-08-12 |
| Context | Aspect ratios | including but not limited to 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 | SourceRetrieved 2026-08-12 |
| Context | Languages | Stable support for 11 languages | SourceRetrieved 2026-08-12 |
| Modalities | FL2VA checkpoint | FL2VA supports t2va and fl2va with text plus optional first and last frames; output is video and audio. | SourceRetrieved 2026-08-12 |
| Modalities | Ref2VA checkpoint | Ref2VA accepts text with up to 9 reference images, up to 3 videos, up to 3 audio clips, and a maximum of 12 files; output is video and audio. | SourceRetrieved 2026-08-12 |
| License | License identity | MiniMax H3 COMMUNITY LICENSE AGREEMENT | SourceRetrieved 2026-08-12 |
| License | Excluded territories | 'Excluded Territories' means the European Union, the United Kingdom, the Republic of Korea and the United States of America. | SourceRetrieved 2026-08-12 |
| License | Territorial output restriction | You may not use, reproduce, modify, distribute, or display the MiniMax H3 Works or any of their Outputs or results outside the Applicable Territory. | SourceRetrieved 2026-08-12 |
| License | Revenue clause | Separate authorization is required if your commercial products and services generate more than 20 million US dollars (or equivalent in other currencies) in yearly revenue. | SourceRetrieved 2026-08-12 |
| License | Output ownership | MiniMax claims no rights over the Outputs you generate. | SourceRetrieved 2026-08-12 |
| License | Encoder license | the encoder of MiniMax H3 uses Qwen3-VL-32B, which is licensed under Apache 2.0 License | SourceRetrieved 2026-08-12 |
| License | Excluded-territory application | https://platform.minimax.io/h3-license | SourceRetrieved 2026-08-12 |
| Pricing | Vendor price comparison | At 2K, H3's per-second price is less than a third of mainstream models | SourceRetrieved 2026-08-12 |
| Availability | Repository access | Ungated repository. | SourceRetrieved 2026-08-12 |
| Availability | Repository formats | Original FL2VA/ and Ref2VA/ formats and diffusers format are provided side by side in one repository. | SourceRetrieved 2026-08-12 |
| Availability | Open base module | H3-Base: Generates audio and video based on the H3-Context-IR output, producing results at 768p resolution. | SourceRetrieved 2026-08-12 |
| Availability | Context module | H3-Context-IR is not included in this open-source release. | SourceRetrieved 2026-08-12 |
| Availability | Higher-resolution module | H3-Regenerate-2K is not yet open-sourced. | SourceRetrieved 2026-08-12 |
| Availability | Diffusers integration | Modular Diffusers blocks integration was merged on August 5, 2026. | SourceRetrieved 2026-08-12 |
| Availability | ComfyUI support | Native support requires version 0.30.0 or later. | SourceRetrieved 2026-08-12 |
| Availability | Task checkpoint weights | MiniMaxAI/MiniMax-H3 | SourceRetrieved 2026-08-12 |
"Today, we're officially launching MiniMax H3, a general-purpose omni-modal generation model. H3 can jointly understand multimodal contexts spanning text, images, video, and audio. It generates video with native stereo audio at up to 2K resolution and 15 seconds in length."
SourceRetrieved 2026-08-12
"Today, we are officially open-sourcing MiniMax H3, our next-generation general-purpose video model."
SourceRetrieved 2026-08-12
These results are vendor-claimed, not independently measured by RunInfra.
No vendor-claimed benchmark results for this exact build were present in the cited material as of 2026-08-12.
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of 2026-08-12, release identity: MiniMaxAI/MiniMax-H3.
As of 2026-08-12, repository access: Ungated repository.
As of 2026-08-12, repository access: Ungated repository.
© 2026 RunInfra. All rights reserved.