MiniMax H3 generates singing and dancing videos from an image and audio

aziib · reddit · 2026-09-08

A Reddit user demos MiniMax H3: given just a reference image and an audio clip, the model generates a video of the character singing and dancing along — a showcase of its motion consistency and audio-visual sync.

Related event: MiniMax H3 tested: generating singing and dancing videos from one image and audio(2 posts)→

Original post →

More from Multimodal

Multimodal channel →