Full-Song Lip-Synced MV with MiniMax H3 at 3 Steps: Method and Benchmarks

Disastrous_Page_301 · reddit · 2026-09-16

A Redditor detailed a full pipeline for a 2:50 lip-synced K-pop style MV using MiniMax H3 Extender: slicing audio at vocal pauses into 17 clips, isolating vocals with htdemucs, and using the fullycopy retention syntax (without it, lip movements degrade to gibberish). With the TaoMate 3-step turbo LoRA, a 9.4s clip renders in 4–5 min on an RTX 3090 (vs 12.5 min at 8 steps), achieving 0.944 audio-envelope correlation and zero seam loss. An automated four-jury QC loop (audio alignment, Vision LLM lipsync scoring, camera stability, defect voting) auto-rerolls failing clips. Limitations: suppressed camera motion at 3 steps and occasional phoneme drift.

Original post →

More from coding & agent

coding & agent channel →