MiniMax H3 Omni-modal Generative System Hits Hugging Face with Native 2K Audio-Video

jiqizhixin · x · 2026-08-03

MiniMax has officially released MiniMax H3, a general-purpose omni-modal generative system, on Hugging Face. The system offers unified understanding of multimodal contexts (text, images, video, audio) and can generate videos with native stereo audio at up to 2K resolution and 15 seconds in duration.

H3 supports various input/output specifications, including text-to-video, image-to-video, and video-to-video, alongside multiple aspect ratios like 21:9 and 9:16. Additionally, the model provides stable support for 11 dialogue languages, including Chinese, English, Japanese, and Korean.

Related event: MiniMax Releases Open-Source Omni-modal Model H3 with Native 2K Audio-Video Generation(14 posts)→

Original post →

More from Models

Models channel →