MiniMax Launches Open-Source Video Model H3: Direct 2K Output, Tops Editing Arena
量子位 · wechat · 2026-07-31
MiniMax has officially released its next-generation video model, MiniMax H3, marking it as their first open-source video model. Breaking the traditional paradigm where models only generate raw clips, H3 integrates editing logic, typography, transitions, background music, and VFX end-to-end, supporting full-modal input to directly output 2K resolution, publish-ready videos.
In hands-on testing, H3 demonstrated strong intent comprehension and multimodal synergy. It effortlessly generates hand-drawn effects and dynamic titles while significantly improving text rendering in videos—a known AI weak point—ensuring stability and lighting consistency across frames. It also supports audio input for emotion-matched dialogue generation.
Pricing is highly competitive, with 2K generation costs at less than 1/3 of mainstream models per second. Technically, the team built the H3-OmniTransformer architecture, utilizing a custom Caption model to achieve deep multimodal emergence. Its open-source nature provides enterprises with a top-tier model for private deployment and custom fine-tuning.
Related event: MiniMax Launches Omni-Modal Model H3 with 2K Video and Open Weights(14 posts)→
More from Models
- DeepSeek V4-Flash Cracks Complex Russian Joke That Trips Up Other LLMs — teortaxesTex · 2026-07-31
- Model Selection is Becoming Org Design: Structuring AI Workflows — every · 2026-07-31
- Testing Inkling Small: A Vision-Equipped Model That Can Build Flappy Bird — LiTianleli · 2026-07-31
- DeepSeek V4-Flash Hits Index Score of 50 at a Cost of Just $0.20 — xeophon · 2026-07-31
- DeepSeek V4-Flash Scores 50 on Artificial Analysis Index, 1 Point Below GLM-5.2 — MagicZhang · 2026-07-31
- V4 Flash Scores 82.7 on Terminal Bench at Just $0.28/M Tokens — jiayuan_jy · 2026-07-31