MiniMax H3 generates singing and dancing videos from an image and audio
aziib · reddit · 2026-09-08
A Reddit user demos MiniMax H3: given just a reference image and an audio clip, the model generates a video of the character singing and dancing along — a showcase of its motion consistency and audio-visual sync.
More from Multimodal
- Deemos Launches Hyper3D MCP to Generate 3D Models Directly Inside Codex — sidahuj · 2026-09-08
- ID-V2V weights land on Hugging Face: Wan 2.1 + VACE, two checkpoints including normal/depth control — minchoi · 2026-09-08
- Netflix releases ID-V2V, an AI that restyles video while keeping faces and lip sync — minchoi · 2026-09-08
- Generating one-actor-face videos: two start frames, two renders, then merge — cocktailpeanut · 2026-09-08
- Minimax H3 Workflow: Image + Audio to Singing-Dancing Video in ~30 Minutes — aziib · 2026-09-08
- Extending MiniMax H3 videos seamlessly via a ComfyUI latent-inheritance node — blackdatafilms · 2026-09-08