Minimax H3 Workflow: Image + Audio to Singing-Dancing Video in ~30 Minutes

aziib · reddit · 2026-09-08

A full reproducible workflow for generating singing-and-dancing videos from an image reference and audio with Minimax H3: the better-human-motion LoRA on HuggingFace, a fast-minimax-h3 workflow on Civitai, and a fused-turbo-int8 quantized checkpoint. The character comes from a VRoid 3D model and the singing voice was trained with RVC from ElevenLabs. Settings: 6 steps at 768p, 30 minutes for a 15-second video.

Related event: MiniMax H3 tested: generating singing and dancing videos from one image and audio(2 posts)→

Original post →

More from Multimodal

Multimodal channel →