Netflix’s ID-V2V preserves identity while restyling videos from one source clip
netflix · hf · 2026-07-28
Netflix researchers present ID-V2V, a video-to-video framework for identity-preserving video restylization.
- Goal: propagate scene, lighting, and style edits from an edited keyframe across a source video while preserving facial likeness, expressions, eye gaze, and lip sync.
- The main challenge is the lack of paired training data, so the method reconstructs training pairs from a single video.
- Their key idea is to decouple identity preservation from edit-driven synthesis: preserve faces via a relighting-style formulation, and use edited keyframes plus depth sequences to guide temporally coherent generation.
- Controls include relit facial regions and facial normal maps for likeness/performance, plus edited keyframes and depth for propagation.
- The authors say the system outperforms prior work on facial likeness and fine-grained performance in both single- and multi-subject scenes, and they release code on GitHub.
More from Multimodal
- dRAE scales visual tokenization to 131,072 codes without codebook collapse — burny_tech · 2026-07-28
- New paper maps compute-optimal scaling laws for native multimodal pre-training — burny_tech · 2026-07-28
- New ComfyUI node converts audio into MIDI for music workflows — MuziqueComfyUI · 2026-07-28
- Boogu-Image-0.1 says 2.08 billion images and $400,000 were enough to reach open-source SOTA — 机器之心 · 2026-07-28
- Exploring ComfyUI Basics: Why Separate Checkpoint and KSampler in Workflows? — DavidThi303 · 2026-07-28
- Running Z Image Turbo on RTX 4060: How to Break Through the Quality Ceiling? — Dangerous_Ring_435 · 2026-07-28