Seamless 1-Shot Lip-Sync with Minimax H3 via Per-Token Noise Masking

stonyleinchen · reddit · 2026-08-15

Release of a ComfyUI workflow update enabling seamless lip-sync using Per-Token Noise Masking on audio and video latents. This method outperforms guidance/reference-based approaches by enabling strong convergence from step 0 and running faster without latent expansion. It pins the music track to the latent space, protecting it from denoising to create strong conditioning, allowing perfect lip-sync even with the FL model. The repo includes custom nodes and example workflows.

Original post →

More from Multimodal

Multimodal channel →