MiniMax-H3 Deep Research
Subject: MiniMaxAI/MiniMax-H3 (a.k.a. Hailuo 3.0) — a 33B omnimodal video generation model that creates footage and synchronized stereo audio in one pass.
Key points
- What it is: MiniMax's third-generation video model (Hailuo 01 → 02 → H3). API + Hailuo App launch: 2026-07-31; open weights on Hugging Face and ModelScope: 2026-08-03. Positioned as general omnimodal video generation: it understands text/image/video/audio and generates video with native stereo audio.
- Architecture: H3-Omni-Transformer — 33B dense, single-stream Transformer (50 layers, hidden dim 5376), no Mamba/SSM blocks. About 13B of parameters sit in AdaLN branches whose modulation outputs can be precomputed and cached at inference (~26GB saved). Unification across modalities is achieved via a 3D multimodal RoPE.
- Fact corrections vs. common claims:
- Not a Mamba/SSM hybrid — that confusion comes from MiniMax-01 (lightning attention), a different, sibling text model.
- Local output is 768p only; 2K comes from
H3-Regenerate-2K, which is not open-sourced. - Full stack at full precision is ~144GB (DiT 66.3 + encoders 66.7 + video VAE 10.4 + audio VAE 0.6); quantized ComfyUI package ~42.5GB. "33B" refers to the DiT body only.
- It is not OSI open source — weights are downloadable under a custom Community License.
- Components: ①
H3-Context-IR(instruction-based editing, not open), ②H3-Base(open, FL2VA / Ref2VA), ③H3-Regenerate-2K(not open). Text encoder: Qwen3-VL-32B (layer-50 hidden states), genuinely Apache 2.0. - H3-VisualVAE (
f16t4d24): temporally causal, 16× spatial / 4× temporal compression, 24 latent channels; additional 1×2×2 patchify gives an effective 32× spatial downsampling. - H3-AudioVAE: processes left/right channels separately then recombines stereo; compresses 32kHz audio to 40Hz latent tokens per channel.
- Sparse attention was mentioned by MiniMax as a future direction, but the current open runtime is a full-attention path — local vs. cloud latency is not directly comparable.
- Native stereo generation (dialogue + SFX + music in one forward pass) is the key differentiator; most open video models (e.g., HunyuanVideo 1.5) are silent.
- Ref2VA multi-reference mode: up to 9 images + 3 videos + 3 audio clips with natural-language relationship descriptions.
- Independent testing (Simon Willison, M5 Max) found visuals impressive, but audio degrades to "speech-shaped noise" unless prompts include explicit sound descriptions.
- Leaderboards: #1 open video model on Video Arena (1476 I2V, 2 points behind Seedance 2.0; ~280 above HunyuanVideo-1.5); Artificial Analysis audio-video editing Elo ~1130 (#1). Treat Elo as directional evidence only.
- Custom Community License (effective 2026-08-02), not OSI open source:
- Excludes the US, EU, UK, and Korea (separate authorization required); hosted API available globally.
- Companies with >$20M annual revenue need prior written approval.
- Requires visible "MiniMax H3" attribution in commercial UIs.
- Prohibits using H3 outputs to train/fine-tune other models (including distillation).
- Governed by Hong Kong SAR law.
- MiniMax attributes the regional exclusions to ongoing generative-video copyright litigation with Hollywood studios (Disney, Universal, Warner). On 2026-08-09 MiniMax stated intent to relicense to Apache 2.0 once copyright issues resolve, and previewed open-sourcing
H3-Regenerate-2K, low-step 4/8-NFE variants, and a text-to-image model. - Day-0 support: ComfyUI (
Comfy-Org/MiniMax-H3, quantized weights + T2V/FLF2V/R2V workflows), Diffusers, SGLang, antirez's pure C + Metalh3.cengine, and 16 chip platforms (NVIDIA, Intel, AMD, Huawei Ascend, Moore Threads, Apple, etc.). - Hardware tiers:
- RTX 4090/3090 (24GB) + quantization: workable at higher resolutions.
- RTX 4070 Ti / 3060 (12GB) + int8 + offload: 5s 480p in ~1–5 minutes.
- DGX Spark (FP8): ~89GB load, FL2VA ~155s observed.
- Apple Silicon (MLX): ≥192GB unified memory recommended; 512GB verified; ~115GB download.
- Full-precision servers: 8×141GB class — effectively out of consumer reach.
Architecture details
Capabilities and limitations
Licensing (read carefully)
Ecosystem and deployment
Comparison with peers
| Model | Resolution | Native audio | Open weights | Local | |---|---|---|---|---| | MiniMax H3 | 768p local / 2K cloud | ✅ stereo 32k | ✅ (Community License) | ✅ (quantized) | | Sora 2 / Pro | up to 1080p | ✅ | ❌ | ❌ | | Veo 3.1 | up to 4K | ✅ | ❌ | ❌ | | Seedance 2.0/2.5 | up to 1080p | ✅ | ❌ | ❌ | | HunyuanVideo 1.5 | 480/720p + 1080p SR | ❌ | ✅ | ✅ | | Wan 2.7 | — | ✅ | ✅ | ✅ | | LTX-2.3 | — | ✅ | ✅ | ✅ |
H3's edge is the combination of open weights + local control + native audio-visual generation + multi-reference control + roughly one-third the price of closed competitors.
Honest boundaries
1. "Open" is nuanced: not OSI open source; two modules unreleased; four regions excluded; $20M revenue threshold. 2. Local output caps at 768p; 2K is cloud-only. 3. Audio quality depends heavily on prompt engineering. 4. Hardware marketing understates requirements (full stack ~144GB; Apple ≥192GB recommended). 5. Elo benchmarks are directional, not production guarantees.
Conclusion
MiniMax-H3 productizes the hardest problem in generative video — joint audio-visual synthesis — via a 33B dense single-stream Transformer producing footage and stereo sound in one forward pass. It is arguably the strongest open-weight video model for local-first, controllable workflows, but it is not an unrestricted free lunch: licensing constraints, 768p local limits, prompt-dependent audio, and hardware demands all warrant caution.