MiniMax H3 video model debuts on Together AI with 2K output and native stereo audio

togethercompute · x · 2026-09-16

MiniMax H3, the third generation of its video line, is now live on Together AI with open weights under the Community License. The omni-modal model accepts up to 9 images, 3 videos, and 3 audio clips as unified context, with reference/editing relationships described in natural language — replacing separate text-to-video, first-last-frame, and editing expert models. It generates 4–15s clips at up to 2K via in-context regeneration, and jointly models native 32kHz stereo audio (voice, SFX, music) with stable dialogue in 11 languages. MiniMax says early testing shows it ready for commercial content creation.

Related event: MiniMax H3 Omni-Modal Video Model Lands on Together AI(2 posts)→

Original post →

More from Multimodal

Multimodal channel →