ComfyUI Automated Video Face Swap Tutorial: Combining Florence 2 and SAM2
petewoodbridge · x · 2026-07-30
A VFX artist shared a fully automated video face swap workflow built in ComfyUI. Rather than a simple pixel replacement, the pipeline achieves high fidelity through multi-model collaboration.
Technical Pipeline:
- Detection & Segmentation: Uses Florence 2 for object detection and bounding box generation, combined with SAM2 for precise segmentation masking. The tutorial also covers crucial Grow Mask tuning techniques.
- Motion & Identity Transfer: Captures original performance motion via pose and face detection. It then uses a reference image combined with QWEN auto-prompting to build semantic understanding and transfer the replacement identity.
- Video Generation: Finally uses WAN Video for sampling, ensuring that lighting, motion, and timing are perfectly preserved across all frames.
More from Multimodal
- ByteDance's Seedance Empowers China's AI Drama: Production Costs Slashed to $3.6k — pstAsiatech · 2026-07-30
- MJ v8.1 Meets Uisato: AI Music Video Workflow Breakdown — Chuka444 · 2026-07-30
- Hyper3D Launches Bang to Parts: One-Click Generative Decomposition for 3D Models — JaynitMakwana · 2026-07-30
- NVIDIA's NHT Outperforms ZipNeRF in 3D Reconstruction, Code Released — ZGojcic · 2026-07-30
- Baseten Merges Kimi Vision Encoder into GLM 5.2 for Multimodal Release — Practical-Collar3063 · 2026-07-30
- Cohere Shares Framework for Balancing Background and Foreground in AI Video — Cohere_Labs · 2026-07-30