MiniMax Releases Open-Source Omni-modal Model H3 with Native 2K Stereo Video

MiniMax has officially released the omni-modal generation system MiniMax-H3. The 33B parameter model features open weights on Hugging Face and is adapted for 🧨 Diffusers. It unifies multimodal context across text, image, video, and audio, supporting diverse inputs like text-to-video, image-to-video, and reference-image-to-video. Users can even combine video, image, text, and audio as reference materials in a single generation. For outputs, H3 natively generates end-to-end videos up to 2K resolution and 15 seconds long with native stereo sound, eliminating the need for post-dubbing.

Confirmed

Why it matters

2026-08-02 ~ 2026-08-03 · 13 related posts

Full story(14 episodes)→

Primary sources

7 near-duplicate retellings: multimodalart · VoidAsuka · MiniMax 稀宇科技 · 赛博禅心 · jiqizhixin · MiniMaxAI · aigclink