Wan 2.2 S2V Hands-On Tutorial
CryptoCatatonic · reddit · 2026-07-16
This is a ComfyUI hands-on tutorial for Wan 2.2 S2V (Sound/Image to Video).
The author demonstrates how to use images and videos as references, combining Kokoro TTS to generate and sync character voices, while integrating DW Pose for finer motion control. The tutorial focuses beyond just getting it to run; it covers how to make characters perform actions absent from the reference images without breaking Wan S2V's lip-sync capabilities.
Overall, this is a practical multimodal video generation workflow, ideal for those looking to combine image, video, dubbing, and pose control for controllable generation.
Related event: Wan 2.2 Sound/Image-to-Video ComfyUI Tutorial Released(2 posts)→
More from coding & agent
- AI agents are starting to strain code hosting platforms — craigsdennis · 2026-07-21
- Async OPD distillation doubles throughput while matching synchronous math accuracy — _lewtun · 2026-07-21
- Omnigent 0.6.0 adds Claude Code imports, Slack approvals and desktop apps — matei_zaharia · 2026-07-21
- Google appears to have quietly shipped Gemini 3.6 Flash, with lower pricing and better agentic scores — xiaohu · 2026-07-21
- Open-source CLI audits AI tools, MCP configs, and agent skills on local machines — Initial-Copy332 · 2026-07-21
- Coding agents feel less stressful when the 5-hour limits are temporarily removed — iamrobotbear · 2026-07-21