MiniMax Open-Sources H3 Video Model: Omnimodal Inputs, Native 2K Stereo Audio
智东西 · wechat · 2026-08-03
MiniMax has officially open-sourced MiniMax H3, its new generation universal video model. H3 is an omnimodal generation system capable of understanding a multimodal context of text, images, video, and audio to generate videos up to 15 seconds long with up to 2K resolution and native stereo sound. Currently, H3 ranks first on the ArtificialAnalysis sound video editing leaderboard with an Elo score of 1130.
The H3 system consists of three main modules: H3-Context-IR for multimodal instruction parsing and orchestration, H3-Base for actual audio-video generation, and H3-Regenerate-2K for regenerating 2K resolution using the original context. The model weights for H3-Base are fully open-source and support deployment via mainstream frameworks like vLLM, SGLang, and ComfyUI, while the other two advanced processing modules require official API calls. This release is accompanied by adaptation support from 16 domestic and international chip and platform ecosystem partners, including Huawei Ascend, AMD, and Intel.
More from Multimodal
- Maestro: A Local AI Studio for Generating Full Music Videos from a Single Prompt — cocktailpeanut · 2026-08-04
- Open-Source Tool Maestro 1.5 Released, Rebuilds Video Recast & Repaint — cocktailpeanut · 2026-08-04
- GPT 5.6 Sol One-Shot Image Generation Aesthetics Tested — kms_dev · 2026-08-04
- MiniMax H3 video model test: Accurately infers unspecified character details — the_bollo · 2026-08-03
- Dreamina Seedance 2.5 Goes Global: Comparative Test of Video Models — azed_ai · 2026-08-03
- Creator Uses Hailuo AI to Generate 'The Odyssey' Themed Short Film — DavidmComfort · 2026-08-03