MiniMax open-sources H3 video model: 33B params, tops HuggingFace, rivals closed-source

大模型之路 · wechat · 2026-08-13

MiniMax open-sourced its video generation model H3 on Aug 3, topping HuggingFace trending within 3 days and attracting 100+ enterprises on day one. H3 uses a 33B-parameter dense single-stream transformer unifying text, image, video, and audio. It generates up to 15s, 2K, 24fps video with native 32kHz stereo, supporting 6 aspect ratios and 11 languages. Two checkpoints: FL2VA for text-to-video and reference-based generation, Ref2VA for multi-modal reference (up to 9 images, 3 videos, 3 audios). It compresses hundreds of K tokens into 4K description via ContextualOmniRepresentation. H3 ranks first among open-source models on DesignArena for multi-image-to-video, image-to-video, and video editing; on ArtificialAnalysis, video editing Elo 1127 (global #1), text-to-video with audio 1237 (second only to Gemini Omni Flash). Price 0.8 RMB/sec, one-third of closed-source flagships. However, H3-Regenerate-2K module is not open-sourced, and license excludes some regions. MiniMax stock rose 78% in a week post-release.

Original post →

More from Models

Models channel →