Testing MiniMax H3 music + lip sync locally: 25 steps minimum, 12 min per 15s clip on a 5090
dassiyu · reddit · 2026-08-21
A Reddit user shared hands-on experience running a MiniMax H3 music generation + lip sync workflow locally:
- A decent 15s clip needs at least 25 steps; using LoRA is not recommended;
- On an RTX 5090 in ComfyUI Kitchen, 15s at 0.9 takes about 12 minutes;
- H3's music generation is solid — noticeably better than the author's results with LTX;
- The main pain point is seamlessly stitching 15s segments; some claim 20s works, but the author's GPU nearly dies at that length.
More from Multimodal
- H3 Model Demonstrates Impressive Identity Preservation from Single Images — linoy_tsaban · 2026-08-22
- MiniMax H3 in Practice: Replaces CGI Workflows for Instant VFX — aziz4ai · 2026-08-22
- Hyperreal AI music video "Saisho Ha Kakusei" made in Cascade Studio — tess-tipple · 2026-08-21
- Creator builds a 30s retro anime OP with Krea 2 + Minimax H3, shares full toolchain and pitfalls — Portable_Solar_ZA · 2026-08-21
- HyCreator Agent Harness: Zero-Intervention 10-Minute Video Generation — bdsqlsz · 2026-08-21
- One user benchmarked 80 Krea 2 checkpoints against the same 5 prompts, results in a public sheet — diffusion_throwaway · 2026-08-21