MiniMax Open-Sources H3: Universal Video Model with Native Stereo Audio and 2K Resolution

MiniMax 稀宇科技 · wechat · 2026-08-03

MiniMax has officially open-sourced MiniMax-H3, its new universal video generation model. H3 is an all-modal generation system capable of understanding complex multimodal contexts (text, image, video, audio) and generating videos up to 2K resolution, 15 seconds long, with native stereo audio.

Key Architecture & Features:

Deployment & Availability:

The H3-Base weights are fully open-source and compatible with popular inference frameworks like SGLang, vLLM, diffusers, and ComfyUI. While the Context-IR and 2K regeneration modules remain closed-source, MiniMax provides APIs and comprehensive workflows to help developers reproduce the official end-to-end 2K generation results.

Related event: MiniMax Releases Open-Source Omni-modal Model H3 with Native 2K Audio-Video Generation(14 posts)→

Original post →

More from Models

Models channel →