MiniMax H3 Omni-modal Generative System Hits Hugging Face with Native 2K Audio-Video
jiqizhixin · x · 2026-08-03
MiniMax has officially released MiniMax H3, a general-purpose omni-modal generative system, on Hugging Face. The system offers unified understanding of multimodal contexts (text, images, video, audio) and can generate videos with native stereo audio at up to 2K resolution and 15 seconds in duration.
H3 supports various input/output specifications, including text-to-video, image-to-video, and video-to-video, alongside multiple aspect ratios like 21:9 and 9:16. Additionally, the model provides stable support for 11 dialogue languages, including Chinese, English, Japanese, and Korean.
More from Models
- Run 2.78T Parameter Kimi K3 on a Single CPU in 8.24GB RAM — Saboo_Shubham_ · 2026-08-03
- MiniMax-H3 Open Weights Restrict Access in US, UK, EU, and South Korea — tokenbender · 2026-08-03
- Biotech Pros Urge Using DeepSeek and Qwen for Better Chinese Technical Info Retrieval — MWCvitkovic · 2026-08-03
- tinygrad Teases Local Deployment Product, Hints at Upcoming Qwen3.6-27B — max_paperclips · 2026-08-03
- US No Longer Safe for Open Weights? MiniMax Shift Sparks Concerns — cocktailpeanut · 2026-08-03
- Qwen3-Max Rumored to Open Source: 2.4T Parameters, Sonnet-Class Performance at Low Cost — bindureddy · 2026-08-03