MiniMax Open-Sources H3: Universal Video Model with Native Stereo Audio and 2K Resolution
MiniMax 稀宇科技 · wechat · 2026-08-03
MiniMax has officially open-sourced MiniMax-H3, its new universal video generation model. H3 is an all-modal generation system capable of understanding complex multimodal contexts (text, image, video, audio) and generating videos up to 2K resolution, 15 seconds long, with native stereo audio.
Key Architecture & Features:
- Multimodal Inputs: Supports text, first/last frames, and up to 12 mixed reference files (images/videos/audio) for complex tasks like identity preservation, lip-sync editing, and voice referencing.
- System Pipeline: Consists of H3-Context-IR (multimodal instruction preprocessing), H3-Base (768p base generation), and H3-Regenerate-2K (in-context 2K regeneration).
- Architecture: H3-Base is a 33B parameter single-stream Omni-Transformer utilizing Qwen3-VL-32B as its text and vision encoder. It natively supports sparse attention to reduce long-sequence computational overhead.
Deployment & Availability:
The H3-Base weights are fully open-source and compatible with popular inference frameworks like SGLang, vLLM, diffusers, and ComfyUI. While the Context-IR and 2K regeneration modules remain closed-source, MiniMax provides APIs and comprehensive workflows to help developers reproduce the official end-to-end 2K generation results.
More from Models
- Run 2.78T Parameter Kimi K3 on a Single CPU in 8.24GB RAM — Saboo_Shubham_ · 2026-08-03
- MiniMax-H3 Open Weights Restrict Access in US, UK, EU, and South Korea — tokenbender · 2026-08-03
- Biotech Pros Urge Using DeepSeek and Qwen for Better Chinese Technical Info Retrieval — MWCvitkovic · 2026-08-03
- tinygrad Teases Local Deployment Product, Hints at Upcoming Qwen3.6-27B — max_paperclips · 2026-08-03
- US No Longer Safe for Open Weights? MiniMax Shift Sparks Concerns — cocktailpeanut · 2026-08-03
- Qwen3-Max Rumored to Open Source: 2.4T Parameters, Sonnet-Class Performance at Low Cost — bindureddy · 2026-08-03