MiniMax Releases H3 Omni-Modal Model: Supports Video and Native Audio Generation
RisingSayak · x · 2026-08-06
MiniMax has released its new H3 model, hailed as a major win for open-source. It is a 33B parameter DiT architecture general-purpose omni-modal generative system.
Key features include:
- Multimodal Input: Supports unified understanding of text, image, video, and audio.
- Audio-Video Generation: Capable of generating video with native stereo audio up to 15 seconds long, 2K resolution, and 24 FPS.
- Open Source: The base pipeline is fully integrated into the Hugging Face Diffusers library with various memory and speed optimizations enabled.
- Multilingual: Stable support for 11 languages including Chinese, English, Japanese, and Korean.
Related event: MiniMax Releases 33B Open-Source Multimodal Model H3(2 posts)→
More from Models
- OpenAI Reportedly Set to Launch New 'Astra' Model Next Week — rohanpaul_ai · 2026-08-06
- DeepSeek Funding Docs Leaked: V4-Pro Matches Claude, Runs on Ascend — xiaohu · 2026-08-06
- Why Do GPT Models Love Generating Comma-Separated Lists? — mariofilhoml · 2026-08-06
- Kimi K3 vs Grok 4.5: Which Model Better Completes Unfinished Sketches? — CodeByPoonam · 2026-08-06
- DeepSeek Planning to Significantly Raise API Prices — miroljub · 2026-08-06
- Unsloth's Gemma 4 mmproj Breaks with Newer llama.cpp, Causing Silent Multimodal Failures — Top_Speaker_7785 · 2026-08-06