MiniMax Releases H3: 33B Open-Source DiT for Image, Video, and Audio
RisingSayak · x · 2026-08-06
MiniMax has released its latest open-source model, H3, which is being celebrated as a major win for the open-source community. It is a 33B parameter DiT architecture with powerful multimodal generation capabilities.
Key Features:
- Omni-generation: Capable of generating images, videos, and audio.
- Multimodal Input: Can take multimodal inputs (including audio) for generating video and audio.
- Long Video: Supports generating videos up to 15 seconds long.
The base pipeline is now fully supported in the Hugging Face Diffusers library with various memory and speed optimizations enabled.
Related event: MiniMax Releases 33B Open-Source Multimodal Model H3(2 posts)→
More from Models
- OpenAI Reportedly Set to Launch New 'Astra' Model Next Week — rohanpaul_ai · 2026-08-06
- DeepSeek Funding Docs Leaked: V4-Pro Matches Claude, Runs on Ascend — xiaohu · 2026-08-06
- Why Do GPT Models Love Generating Comma-Separated Lists? — mariofilhoml · 2026-08-06
- Kimi K3 vs Grok 4.5: Which Model Better Completes Unfinished Sketches? — CodeByPoonam · 2026-08-06
- DeepSeek Planning to Significantly Raise API Prices — miroljub · 2026-08-06
- Unsloth's Gemma 4 mmproj Breaks with Newer llama.cpp, Causing Silent Multimodal Failures — Top_Speaker_7785 · 2026-08-06