MiniMax Open-Sources H3 Video Model: Omnimodal Inputs, Native 2K Stereo Audio

智东西 · wechat · 2026-08-03

MiniMax has officially open-sourced MiniMax H3, its new generation universal video model. H3 is an omnimodal generation system capable of understanding a multimodal context of text, images, video, and audio to generate videos up to 15 seconds long with up to 2K resolution and native stereo sound. Currently, H3 ranks first on the ArtificialAnalysis sound video editing leaderboard with an Elo score of 1130.

The H3 system consists of three main modules: H3-Context-IR for multimodal instruction parsing and orchestration, H3-Base for actual audio-video generation, and H3-Regenerate-2K for regenerating 2K resolution using the original context. The model weights for H3-Base are fully open-source and support deployment via mainstream frameworks like vLLM, SGLang, and ComfyUI, while the other two advanced processing modules require official API calls. This release is accompanied by adaptation support from 16 domestic and international chip and platform ecosystem partners, including Huawei Ascend, AMD, and Intel.

Related event: MiniMax Releases Open-Source Omni-Modal Model H3 for Native Audio-Video Generation(33 posts)→

Original post →

More from Multimodal

Multimodal channel →