MiniMax-H3 Released: Omni-modal System Generates 2K Video with Native Stereo Audio

VoidAsuka · x · 2026-08-03

MiniMax has publicly released MiniMax-H3, a general-purpose omni-modal generative system, on Hugging Face. The model offers unified understanding of multimodal contexts—including text, images, video, and audio—and can generate videos up to 2K resolution and 15 seconds in length with native stereo sound.

Designed with task generalization in mind, H3 possesses broad multimodal comprehension and generation capabilities straight out of the pre-training stage, enabling it to follow complex instructions effectively. It supports various input modes like first-and-last-frame control and handles 11 languages, including Chinese and English.

Related event: MiniMax Releases Open-Source Omni-modal Model H3 with Native 2K Audio-Video Generation(14 posts)→

Original post →

More from Multimodal

Multimodal channel →