MiniMax-H3 Released: Omni-modal System Generates 2K Video with Native Stereo Audio
VoidAsuka · x · 2026-08-03
MiniMax has publicly released MiniMax-H3, a general-purpose omni-modal generative system, on Hugging Face. The model offers unified understanding of multimodal contexts—including text, images, video, and audio—and can generate videos up to 2K resolution and 15 seconds in length with native stereo sound.
Designed with task generalization in mind, H3 possesses broad multimodal comprehension and generation capabilities straight out of the pre-training stage, enabling it to follow complex instructions effectively. It supports various input modes like first-and-last-frame control and handles 11 languages, including Chinese and English.
More from Multimodal
- Developer Creates 'Song of the Sirens' Concept Art Using Grok — bennash · 2026-08-03
- xAI Updates Grok Imagine 1.5, Wowing Users with Image Generation Upgrades — minchoi · 2026-08-03
- Seedance Video Model Excels at Crisp Vector-Style Graphics — bennash · 2026-08-03
- Four Years of AI Image Generation Evolution: Same Prompt, Worlds Apart — DreamFly_13 · 2026-08-03
- Motion Skill Turns Claude into a Video Team: Generate Launch Videos from a URL — Scobleizer · 2026-08-03
- MiniMax Launches H3: Native 2K Video with Stereo Audio Generation — Mobile-Pumpkin7944 · 2026-08-03