MiniMax Releases Open-Weights Omni-Modal H3 Model with Native Stereo Audio and 2K Video
mishig25 · x · 2026-08-04
MiniMax has released MiniMax-H3, an open-weights omni-modal generative system, now available on Hugging Face. Designed with task-generalization in mind, the system possesses broad multimodal context understanding and generation capabilities straight out of the pre-training stage.
Key Features & Specs:
- Unified Multimodal Understanding: Supports unified comprehension of text, images, video, and audio.
- High-Quality Audio-Video Generation: Generates video up to 15 seconds long with 2K resolution at 24 FPS, featuring native 32kHz stereo audio generation.
- Versatile Input Modes: Offers multiple variants like First-and-Last-Frame (FL2VA), supporting comprehensive workflows including image-to-video, video-to-video, and audio-to-video.
- Multilingual: Provides stable support for 11 languages including Chinese, English, Japanese, and Korean.
Related event: MiniMax Releases Open-Weight Multimodal Model H3(6 posts)→
More from Models
- DiffusionGemma Tech Report: Parallel Decoding Breaks LLM Inference Speed Limits — bodonoghue85 · 2026-08-04
- ICML Winner FutureSim: GPT-5.6 Executes Over 15K Tool Calls in One Run — scaling01 · 2026-08-04
- Google's DiffusionGemma: Parallel Decoding Breaks LLM Text Generation Speed Limits — bodonoghue85 · 2026-08-04
- Testing GPT-5.6 and Gemini Robotics Models: Impressive but Need Real Deployment Data — m_wulfmeier · 2026-08-04
- Far Behind GPT? Devs Complain Gemini Live Lacks Emotional Nuance — TheBuzzer4625kHz · 2026-08-04
- Wired: Amid US Tech Turmoil, Open-Weight Pioneer Mistral Hits Its Stride — Wired AI · 2026-08-04