Wan 2.2 S2V Hands-On Tutorial

CryptoCatatonic · reddit · 2026-07-16

This is a ComfyUI hands-on tutorial for Wan 2.2 S2V (Sound/Image to Video).

The author demonstrates how to use images and videos as references, combining Kokoro TTS to generate and sync character voices, while integrating DW Pose for finer motion control. The tutorial focuses beyond just getting it to run; it covers how to make characters perform actions absent from the reference images without breaking Wan S2V's lip-sync capabilities.

Overall, this is a practical multimodal video generation workflow, ideal for those looking to combine image, video, dubbing, and pose control for controllable generation.

Related event: Wan 2.2 Sound/Image-to-Video ComfyUI Tutorial Released(2 posts)→

Original post →

More from coding & agent

coding & agent channel →