English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MiniMax-H3 Deep Research: A 33B Omnimodal Video Generation Model with Native Stereo Audio

Forum topic · ✨步子哥 · 2026-08-12

Summary

MiniMax-H3 (Hailuo 3.0) is a 33B-parameter dense, single-stream Transformer (H3-Omni-Transformer) for omnimodal video generation, released via API on 2026-07-31 with open weights on Hugging Face (MiniMaxAI/MiniMax-H3) on 2026-08-03. Unlike serial pipelines that generate silent video then attach audio, H3 jointly generates video and native 32kHz stereo audio in a single forward pass, using a 3D multimodal RoPE and separate VisualVAE (16x spatial, 4x temporal compression) and AudioVAE components. It accepts up to 9 images, 3 videos, and 3 audio clips as references, produces up to 15s at 24fps and 2K (cloud only; local output is 768p). The stack uses Qwen3-VL-32B as text encoder (Apache 2.0). Weights ship under a custom Community License, not OSI open source: US/EU/UK/Korea excluded, >$20M revenue requires authorization, and H3-Context-IR and H3-Regenerate-2K remain unreleased; MiniMax has stated intent to move to Apache 2.0. Community benchmarks rank it first among open-weight video models. This report cross-verifies architecture, licensing, hardware requirements (full stack ~144GB), and deployment via ComfyUI, Diffusers, and SGLang.

MiniMax-H3 Deep Research

Subject: MiniMaxAI/MiniMax-H3 (a.k.a. Hailuo 3.0) — a 33B omnimodal video generation model that creates footage and synchronized stereo audio in one pass.

Key points

  • What it is: MiniMax's third-generation video model (Hailuo 01 → 02 → H3). API + Hailuo App launch: 2026-07-31; open weights on Hugging Face and ModelScope: 2026-08-03. Positioned as general omnimodal video generation: it understands text/image/video/audio and generates video with native stereo audio.
  • Architecture: H3-Omni-Transformer — 33B dense, single-stream Transformer (50 layers, hidden dim 5376), no Mamba/SSM blocks. About 13B of parameters sit in AdaLN branches whose modulation outputs can be precomputed and cached at inference (~26GB saved). Unification across modalities is achieved via a 3D multimodal RoPE.
  • Fact corrections vs. common claims:
  • Not a Mamba/SSM hybrid — that confusion comes from MiniMax-01 (lightning attention), a different, sibling text model.
  • Local output is 768p only; 2K comes from H3-Regenerate-2K, which is not open-sourced.
  • Full stack at full precision is ~144GB (DiT 66.3 + encoders 66.7 + video VAE 10.4 + audio VAE 0.6); quantized ComfyUI package ~42.5GB. "33B" refers to the DiT body only.
  • It is not OSI open source — weights are downloadable under a custom Community License.
  • Components: ① H3-Context-IR (instruction-based editing, not open), ② H3-Base (open, FL2VA / Ref2VA), ③ H3-Regenerate-2K (not open). Text encoder: Qwen3-VL-32B (layer-50 hidden states), genuinely Apache 2.0.
  • Architecture details

  • H3-VisualVAE (f16t4d24): temporally causal, 16× spatial / 4× temporal compression, 24 latent channels; additional 1×2×2 patchify gives an effective 32× spatial downsampling.
  • H3-AudioVAE: processes left/right channels separately then recombines stereo; compresses 32kHz audio to 40Hz latent tokens per channel.
  • Sparse attention was mentioned by MiniMax as a future direction, but the current open runtime is a full-attention path — local vs. cloud latency is not directly comparable.
  • Capabilities and limitations

  • Native stereo generation (dialogue + SFX + music in one forward pass) is the key differentiator; most open video models (e.g., HunyuanVideo 1.5) are silent.
  • Ref2VA multi-reference mode: up to 9 images + 3 videos + 3 audio clips with natural-language relationship descriptions.
  • Independent testing (Simon Willison, M5 Max) found visuals impressive, but audio degrades to "speech-shaped noise" unless prompts include explicit sound descriptions.
  • Leaderboards: #1 open video model on Video Arena (1476 I2V, 2 points behind Seedance 2.0; ~280 above HunyuanVideo-1.5); Artificial Analysis audio-video editing Elo ~1130 (#1). Treat Elo as directional evidence only.
  • Licensing (read carefully)

  • Custom Community License (effective 2026-08-02), not OSI open source:
  • Excludes the US, EU, UK, and Korea (separate authorization required); hosted API available globally.
  • Companies with >$20M annual revenue need prior written approval.
  • Requires visible "MiniMax H3" attribution in commercial UIs.
  • Prohibits using H3 outputs to train/fine-tune other models (including distillation).
  • Governed by Hong Kong SAR law.
  • MiniMax attributes the regional exclusions to ongoing generative-video copyright litigation with Hollywood studios (Disney, Universal, Warner). On 2026-08-09 MiniMax stated intent to relicense to Apache 2.0 once copyright issues resolve, and previewed open-sourcing H3-Regenerate-2K, low-step 4/8-NFE variants, and a text-to-image model.
  • Ecosystem and deployment

  • Day-0 support: ComfyUI (Comfy-Org/MiniMax-H3, quantized weights + T2V/FLF2V/R2V workflows), Diffusers, SGLang, antirez's pure C + Metal h3.c engine, and 16 chip platforms (NVIDIA, Intel, AMD, Huawei Ascend, Moore Threads, Apple, etc.).
  • Hardware tiers:
  • RTX 4090/3090 (24GB) + quantization: workable at higher resolutions.
  • RTX 4070 Ti / 3060 (12GB) + int8 + offload: 5s 480p in ~1–5 minutes.
  • DGX Spark (FP8): ~89GB load, FL2VA ~155s observed.
  • Apple Silicon (MLX): ≥192GB unified memory recommended; 512GB verified; ~115GB download.
  • Full-precision servers: 8×141GB class — effectively out of consumer reach.

Comparison with peers

| Model | Resolution | Native audio | Open weights | Local | |---|---|---|---|---| | MiniMax H3 | 768p local / 2K cloud | ✅ stereo 32k | ✅ (Community License) | ✅ (quantized) | | Sora 2 / Pro | up to 1080p | ✅ | ❌ | ❌ | | Veo 3.1 | up to 4K | ✅ | ❌ | ❌ | | Seedance 2.0/2.5 | up to 1080p | ✅ | ❌ | ❌ | | HunyuanVideo 1.5 | 480/720p + 1080p SR | ❌ | ✅ | ✅ | | Wan 2.7 | — | ✅ | ✅ | ✅ | | LTX-2.3 | — | ✅ | ✅ | ✅ |

H3's edge is the combination of open weights + local control + native audio-visual generation + multi-reference control + roughly one-third the price of closed competitors.

Honest boundaries

1. "Open" is nuanced: not OSI open source; two modules unreleased; four regions excluded; $20M revenue threshold. 2. Local output caps at 768p; 2K is cloud-only. 3. Audio quality depends heavily on prompt engineering. 4. Hardware marketing understates requirements (full stack ~144GB; Apple ≥192GB recommended). 5. Elo benchmarks are directional, not production guarantees.

Conclusion

MiniMax-H3 productizes the hardest problem in generative video — joint audio-visual synthesis — via a 33B dense single-stream Transformer producing footage and stereo sound in one forward pass. It is arguably the strongest open-weight video model for local-first, controllable workflows, but it is not an unrestricted free lunch: licensing constraints, 768p local limits, prompt-dependent audio, and hardware demands all warrant caution.

Tags

#minimax#minimax-h3#video-generation#omnimodal#open-weights#ai-audio#comfyui#huggingface

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633386