MiniMax Launches H3 Omni-Modal Model: 2K Video with Native Stereo Audio
multimodalart · x · 2026-08-03
MiniMax has released the H3 omni-modal generative system on Hugging Face. The model supports unified understanding of text, images, video, and audio, capable of generating videos with native stereo sound.
Key Features:
- High-Spec Output: Generates videos up to 15 seconds long, with resolutions up to 2K at 24 FPS, and 32kHz stereo audio.
- Multimodal Inputs: Supports text-to-video, image-to-video, reference-to-video, and first-last-frame generation.
- Multilingual: Stable support for instructions in 11 languages including Chinese, English, and Japanese.
- Consumer Hardware Ready: Despite its 33B parameters, the model is optimized for 🧨 Diffusers and ComfyUI, making it runnable on consumer GPUs.
More from Models
- MiniMax Officially Releases H3 Video Generation Model with Audio Support — MiniMaxAI · 2026-08-03
- Users Call for 200B Parameter Open-Source Release from Qwen — adrianscottcom · 2026-08-03
- Code Arena Leaderboard: Claude Opus 5 Max Tops Web Dev Eval — arena · 2026-08-03
- Qwen3.8-Max Pricing Revealed: Nearly 3x Cheaper Than Kimi K3 — cedric_chee · 2026-08-03
- Anthropic's Opus Criticized for Being Sycophantic and Missing Instructions — yacineMTB · 2026-08-03
- Qwen3.8-27B Open Weights Coming, Runs Locally on 17GB RAM — danielhanchen · 2026-08-03