MiniMax Open-Sources H3: An Omni-Modal Video Model with Native 2K Stereo Audio
赛博禅心 · wechat · 2026-08-03
MiniMax has officially open-sourced MiniMax-H3, a new universal video model. It is an omni-modal generation system designed to understand complex inputs consisting of text, images, video, and audio, capable of generating videos up to 2K resolution and 15 seconds long with native stereo sound.
Core Architecture & Capabilities
- H3-Base: The core is a 33B parameter single-stream Transformer that jointly predicts video and audio latents. The text encoder utilizes the pre-trained weights of Qwen3-VL-32B.
- H3-Context-IR: A preprocessing system designed to deeply parse and distill complex multi-modal instructions into a structured representation to enhance generation quality.
- H3-Regenerate-2K: Instead of a traditional super-resolution module, it feeds the 768p output along with the original context back into the model to regenerate a 2K high-resolution video.
Deployment & Availability
The model supports text-to-video, first/last frame generation, and multi-modal reference generation (Ref2VA). Currently, the H3-Base weights are fully open and compatible with mainstream frameworks like SGLang, vLLM, diffusers, and ComfyUI for local deployment. The Context-IR and Regenerate-2K modules are provided via API to reproduce the official end-to-end workflow.
More from Models
- Run 2.78T Parameter Kimi K3 on a Single CPU in 8.24GB RAM — Saboo_Shubham_ · 2026-08-03
- MiniMax-H3 Open Weights Restrict Access in US, UK, EU, and South Korea — tokenbender · 2026-08-03
- Biotech Pros Urge Using DeepSeek and Qwen for Better Chinese Technical Info Retrieval — MWCvitkovic · 2026-08-03
- tinygrad Teases Local Deployment Product, Hints at Upcoming Qwen3.6-27B — max_paperclips · 2026-08-03
- US No Longer Safe for Open Weights? MiniMax Shift Sparks Concerns — cocktailpeanut · 2026-08-03
- Qwen3-Max Rumored to Open Source: 2.4T Parameters, Sonnet-Class Performance at Low Cost — bindureddy · 2026-08-03