MiniMax H3 lands on Together AI: 33B omni-modal model generates 2K clips with native stereo audio

togethercompute · x · 2026-09-16

MiniMax's H3 video generation model is now live on Together AI. The 33B omni-modal model generates 4–15 second clips at up to 2K resolution with native stereo audio.

It accepts text, images, video, and audio as context, enabling multimodal reference-driven generation.

Related event: MiniMax H3 Omni-Modal Video Model Lands on Together AI(2 posts)→

Original post →

More from Multimodal

Multimodal channel →