MiniMax Open-Sources H3: An Omni-Modal Video Model with Native 2K Stereo Audio

赛博禅心 · wechat · 2026-08-03

MiniMax has officially open-sourced MiniMax-H3, a new universal video model. It is an omni-modal generation system designed to understand complex inputs consisting of text, images, video, and audio, capable of generating videos up to 2K resolution and 15 seconds long with native stereo sound.

Core Architecture & Capabilities

Deployment & Availability

The model supports text-to-video, first/last frame generation, and multi-modal reference generation (Ref2VA). Currently, the H3-Base weights are fully open and compatible with mainstream frameworks like SGLang, vLLM, diffusers, and ComfyUI for local deployment. The Context-IR and Regenerate-2K modules are provided via API to reproduce the official end-to-end workflow.

Related event: MiniMax Releases Open-Source Omni-modal Model H3 with Native 2K Audio-Video Generation(14 posts)→

Original post →

More from Models

Models channel →