RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading related term
Loading related term
Loading related term
Loading related term
Home/Catalog/MiniMax H3
Model reference

MiniMax H3

This reference covers the open base checkpoint for an omni-modal video generation system, with higher-resolution modules remaining hosted. Its community license excludes the European Union, the United Kingdom, the Republic of Korea, and the United States of America.

MiniMax H3 is the vendor reference for MiniMaxAI/MiniMax-H3, as of 2026-08-12. Omni Transformer: H3-Omni-Transformer is a 33B-parameter dense, single-stream Transformer, with approximately 13B parameters residing in AdaLN-related branches.; Output duration: 4-15 seconds, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.

Cited specifications

GroupFactSource-cited display valueSource and date
IdentityRelease identityMiniMaxAI/MiniMax-H3
SourceRetrieved 2026-08-12
IdentityVendor descriptionMiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds.
SourceRetrieved 2026-08-12
IdentityLaunch publication date2026-07-31
SourceRetrieved 2026-08-12
IdentityOpen-source publication date2026-08-03
SourceRetrieved 2026-08-12
IdentityLicense-stated dateMiniMax H3 release date/License date: August 2, 2026.
SourceRetrieved 2026-08-12
ArchitectureOmni TransformerH3-Omni-Transformer is a 33B-parameter dense, single-stream Transformer, with approximately 13B parameters residing in AdaLN-related branches.
SourceRetrieved 2026-08-12
ArchitectureAdaLN deploymentAdaLN-related branches are cacheable for inference-only deployment.
SourceRetrieved 2026-08-12
ArchitecturePosition encoding3D MM-RoPE over (t, h, w)
SourceRetrieved 2026-08-12
ArchitectureText and vision encoderuses the full pretrained weights of Qwen3-VL-32B; layer-50 hidden states
SourceRetrieved 2026-08-12
ArchitectureVideo VAEf16t4d24; temporally causal
SourceRetrieved 2026-08-12
ArchitectureAudio VAE32 kHz stereo to 40 Hz latents per channel
SourceRetrieved 2026-08-12
ArchitectureJoint predictionjoint audio and video prediction in one transformer
SourceRetrieved 2026-08-12
ArchitectureInitial attention pathThe initial open-source release provides inference with full attention only.
SourceRetrieved 2026-08-12
ArchitectureReleased checkpointsThe released checkpoints are CFG-distilled Omni Transformer model weights.
SourceRetrieved 2026-08-12
ContextOutput duration4-15 seconds
SourceRetrieved 2026-08-12
ContextOutput frame rate24 FPS
SourceRetrieved 2026-08-12
ContextOutput audio32 kHz stereo
SourceRetrieved 2026-08-12
ContextOutput resolutionShorter side 768 pixels by default; 2K generation can be achieved with H3-Regenerate-2K.
SourceRetrieved 2026-08-12
ContextAspect ratiosincluding but not limited to 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16
SourceRetrieved 2026-08-12
ContextLanguagesStable support for 11 languages
SourceRetrieved 2026-08-12
ModalitiesFL2VA checkpointFL2VA supports t2va and fl2va with text plus optional first and last frames; output is video and audio.
SourceRetrieved 2026-08-12
ModalitiesRef2VA checkpointRef2VA accepts text with up to 9 reference images, up to 3 videos, up to 3 audio clips, and a maximum of 12 files; output is video and audio.
SourceRetrieved 2026-08-12
LicenseLicense identityMiniMax H3 COMMUNITY LICENSE AGREEMENT
SourceRetrieved 2026-08-12
LicenseExcluded territories'Excluded Territories' means the European Union, the United Kingdom, the Republic of Korea and the United States of America.
SourceRetrieved 2026-08-12
LicenseTerritorial output restrictionYou may not use, reproduce, modify, distribute, or display the MiniMax H3 Works or any of their Outputs or results outside the Applicable Territory.
SourceRetrieved 2026-08-12
LicenseRevenue clauseSeparate authorization is required if your commercial products and services generate more than 20 million US dollars (or equivalent in other currencies) in yearly revenue.
SourceRetrieved 2026-08-12
LicenseOutput ownershipMiniMax claims no rights over the Outputs you generate.
SourceRetrieved 2026-08-12
LicenseEncoder licensethe encoder of MiniMax H3 uses Qwen3-VL-32B, which is licensed under Apache 2.0 License
SourceRetrieved 2026-08-12
LicenseExcluded-territory applicationhttps://platform.minimax.io/h3-license
SourceRetrieved 2026-08-12
PricingVendor price comparisonAt 2K, H3's per-second price is less than a third of mainstream models
SourceRetrieved 2026-08-12
AvailabilityRepository accessUngated repository.
SourceRetrieved 2026-08-12
AvailabilityRepository formatsOriginal FL2VA/ and Ref2VA/ formats and diffusers format are provided side by side in one repository.
SourceRetrieved 2026-08-12
AvailabilityOpen base moduleH3-Base: Generates audio and video based on the H3-Context-IR output, producing results at 768p resolution.
SourceRetrieved 2026-08-12
AvailabilityContext moduleH3-Context-IR is not included in this open-source release.
SourceRetrieved 2026-08-12
AvailabilityHigher-resolution moduleH3-Regenerate-2K is not yet open-sourced.
SourceRetrieved 2026-08-12
AvailabilityDiffusers integrationModular Diffusers blocks integration was merged on August 5, 2026.
SourceRetrieved 2026-08-12
AvailabilityComfyUI supportNative support requires version 0.30.0 or later.
SourceRetrieved 2026-08-12
AvailabilityTask checkpoint weightsMiniMaxAI/MiniMax-H3
SourceRetrieved 2026-08-12

What the vendor says is new

"Today, we're officially launching MiniMax H3, a general-purpose omni-modal generation model. H3 can jointly understand multimodal contexts spanning text, images, video, and audio. It generates video with native stereo audio at up to 2K resolution and 15 seconds in length."

SourceRetrieved 2026-08-12

"Today, we are officially open-sourcing MiniMax H3, our next-generation general-purpose video model."

SourceRetrieved 2026-08-12

Vendor-claimed benchmarks

These results are vendor-claimed, not independently measured by RunInfra.

No vendor-claimed benchmark results for this exact build were present in the cited material as of 2026-08-12.

Serving support

Listed rows have a cited upstream support signal. They are not RunInfra measurements.

SGLang

The SGLang cookbook documents multi-GPU serving for MiniMax-H3.

SourceRetrieved 2026-08-12

vLLM

The official vLLM recipes site publishes a MiniMax-H3 recipe linked from the model card.

SourceRetrieved 2026-08-12

Serving concepts

  • Model loading->
  • Tensor parallelism->
  • HBM->
  • Cold start->

What RunInfra measures when we measure it

RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.

  • Measurement methodology->
  • Measured package catalog->
  • Published benchmarks->

Questions about this model reference

What is MiniMax H3?

As of 2026-08-12, release identity: MiniMaxAI/MiniMax-H3.

Where are the MiniMax H3 weights?

As of 2026-08-12, repository access: Ungated repository.

Can I run MiniMax H3 myself?

As of 2026-08-12, repository access: Ungated repository.

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy