Identity-Locked Talking Video from Single Image and Audio

DaLyon92x · reddit · 2026-07-10

The post showcases a workflow for generating identity-locked talking videos using a "single photo + personal audio recording," emphasizing it avoids traditional face-swapping or driving videos. The author explains that voice recordings can be frozen into the audio latent, allowing joint audio-video denoising to generate lip movements and sound consistent with the recording. Workflows, examples, and repositories for both CUDA and Apple Silicon are provided.

Original post →

More from Embodied

Embodied channel →