MiniMax Launches H3 Omni-modal Model: Native Stereo Audio and 2K Video
MiniMax 稀宇科技 · wechat · 2026-07-31
MiniMax has officially released H3, an omni-modal generative model. It unifies the understanding and generation of text, images, video, and audio, supporting native stereo sound and up to 15 seconds of 2K resolution video generation.
Key Features & Technology:
- Multimodal Contextual Understanding: Users can describe relationships between reference materials (video, image, audio) and the target video via natural language for complex cross-modal editing.
- Architecture & Cost Efficiency: Introduces the new H3-VAE and H3-OmniTransformer architectures, significantly boosting compression rates and training throughput. The cost per second at 2K resolution is less than 1/3 of mainstream models.
- In-context Regeneration: Abandons traditional super-resolution modules, directly using the base model to regenerate high-res 2K output leveraging original multimodal context.
Open Source & Commercialization:
Designed for commercial scenarios like advertising, e-commerce, and UI/UX, H3 offers industry-leading cost-efficiency. MiniMax also announced plans to open-source the model weights in the coming days to boost the open-source community and facilitate domestic chip adaptation.
Related event: MiniMax Launches Omni-Modal Model H3 with Native 2K Stereo Video(5 posts)→
More from Models
- Gary Marcus Rounds Up Seven Shambolic AI Moments: OpenAI's 80% Price Cut and Gov Map Fail — GaryMarcus · 2026-07-31
- OpenAI Exec Teases Delivering 'Intelligence Too Cheap to Meter' This Week — daniel_mac8 · 2026-07-31
- MiniMax Launches H3 Video Model: Native 2K Stereo Audio, Open-Weights Soon — Hoodfu · 2026-07-31
- GPT-5.6 Luna Tested After 80% Price Drop: Blazing Fast Full-Stack Code Generation — BorisMPower · 2026-07-31
- Glass AI Tops Medical Benchmark, Beating OpenAI and Google — GlassHealthHQ · 2026-07-31
- Specific Prompts Trigger Bizarre Claude Opus Behavior, Sparking Prediction Market — rgblong · 2026-07-31