MiniMax Releases Open-Source H3: An Omni-Modal Model with Native Audio-Video Generation
aigclink · x · 2026-08-03
MiniMax has officially released and open-sourced the weights for H3, a universal omni-modal generation system. It can uniformly understand text, images, video, and audio, directly generating video content with native stereo sound.
Core Capabilities & Advantages:
- Unified Multimodality: Supports up to 15s of video generation at 2K resolution with native dual-channel audio. Features image-to-video, video-to-video (V2V) motion transfer, and audio-to-audio reference editing.
- Extreme Cost-Efficiency: The official cost per second at 2K resolution is less than 1/3 of mainstream models, and 768P is priced at half of what competitors charge for 720P.
- Technical Innovations: Utilizes H3-VAE tokenizer (4x sequence length efficiency), H3-Omni Transformer unified architecture, and heterogeneous compute optimization (30% training throughput boost). Training data consists entirely of real natural data.
Limitations: The H3-Regenerate-2K module and sparse attention implementation are not open-sourced initially, and detail precision in some scenarios still needs improvement.
More from Models
- Run 2.78T Parameter Kimi K3 on a Single CPU in 8.24GB RAM — Saboo_Shubham_ · 2026-08-03
- MiniMax-H3 Open Weights Restrict Access in US, UK, EU, and South Korea — tokenbender · 2026-08-03
- Biotech Pros Urge Using DeepSeek and Qwen for Better Chinese Technical Info Retrieval — MWCvitkovic · 2026-08-03
- tinygrad Teases Local Deployment Product, Hints at Upcoming Qwen3.6-27B — max_paperclips · 2026-08-03
- US No Longer Safe for Open Weights? MiniMax Shift Sparks Concerns — cocktailpeanut · 2026-08-03
- Qwen3-Max Rumored to Open Source: 2.4T Parameters, Sonnet-Class Performance at Low Cost — bindureddy · 2026-08-03