Minimax H3 Workflow: Image + Audio to Singing-Dancing Video in ~30 Minutes
aziib · reddit · 2026-09-08
A full reproducible workflow for generating singing-and-dancing videos from an image reference and audio with Minimax H3: the better-human-motion LoRA on HuggingFace, a fast-minimax-h3 workflow on Civitai, and a fused-turbo-int8 quantized checkpoint. The character comes from a VRoid 3D model and the singing voice was trained with RVC from ElevenLabs. Settings: 6 steps at 768p, 30 minutes for a 15-second video.
More from Multimodal
- A video pipeline separating planning, Seedance 2.5 generation, and conversational CapCut editing — HeyZoyaKhan · 2026-09-09
- Spatial-first AI video: map the scene with GPT-6 Astra before rendering — HeyZoyaKhan · 2026-09-09
- GPT-6 Astra codes geometry, Blender + Clay plugin renders, Dreamina finalizes footage — thetripathi58 · 2026-09-09
- GPT-6 Astra + Blender + Dreamina Seedance 2.5 forms a working 3D-to-video pipeline — thetripathi58 · 2026-09-09
- Claude skips the training set: HTML/CSS rendered in a headless browser makes his site images — shashib · 2026-09-09
- The missing step: Clay Renderer plugin bridges Blender exports into Dreamina renders — socialwithaayan · 2026-09-09