Seamless 1-Shot Lip-Sync with Minimax H3 via Per-Token Noise Masking
stonyleinchen · reddit · 2026-08-15
Release of a ComfyUI workflow update enabling seamless lip-sync using Per-Token Noise Masking on audio and video latents. This method outperforms guidance/reference-based approaches by enabling strong convergence from step 0 and running faster without latent expansion. It pins the music track to the latent space, protecting it from denoising to create strong conditioning, allowing perfect lip-sync even with the FL model. The repo includes custom nodes and example workflows.
More from Multimodal
- Archviz creator wants to commission a ComfyUI workflow to AI-enhance 3D animation frames — PerfectoAllstar · 2026-08-15
- Converted 2D anime to 3D SBS video despite traditional depth sensor limitations — bdsqlsz · 2026-08-15
- Midjourney Prompting: Storytelling Through Object Still Life — tisch_eins · 2026-08-15
- Creator announces shift from 2D to exclusive dynamic 3D filming — TheTuringPost · 2026-08-15
- Test shows Qwen3.8-27b water surface rendering beats Gemini Flash — pbaylies · 2026-08-15
- Krea 2 Tiled Upscale Workflow for Detailed D&D Maps — 05032-MendicantBias · 2026-08-15