MiniMax H3 video model debuts on Together AI with 2K output and native stereo audio
togethercompute · x · 2026-09-16
MiniMax H3, the third generation of its video line, is now live on Together AI with open weights under the Community License. The omni-modal model accepts up to 9 images, 3 videos, and 3 audio clips as unified context, with reference/editing relationships described in natural language — replacing separate text-to-video, first-last-frame, and editing expert models. It generates 4–15s clips at up to 2K via in-context regeneration, and jointly models native 32kHz stereo audio (voice, SFX, music) with stable dialogue in 11 languages. MiniMax says early testing shows it ready for commercial content creation.
Related event: MiniMax H3 Omni-Modal Video Model Lands on Together AI(2 posts)→
More from Multimodal
- GMI launches MCP server exposing 150+ multimodal models to Claude, ChatGPT and Cursor — _jaydeepkarale · 2026-09-16
- ComfyUI stable release ships YuE2, the local music generation model — lazyspock · 2026-09-16
- Audio8 open-sources on-device ASR/TTS models down to 0.1B, including iPhone offline transcription — FinanceYF5 · 2026-09-16
- GPT Image 2.5 versus seven other image models on the same prompt — ZootAllures9111 · 2026-09-16
- FP8+AOTI-optimized Wan2.2 video model space trends on Hugging Face — zerogpu-aoti · 2026-09-16
- MiniMax Unveils H3 IP Edition with Officially Licensed Japanese IP — MiniMax_AI · 2026-09-16