Identity-Locked Talking Video from Single Image and Audio
DaLyon92x · reddit · 2026-07-10
The post showcases a workflow for generating identity-locked talking videos using a "single photo + personal audio recording," emphasizing it avoids traditional face-swapping or driving videos. The author explains that voice recordings can be frozen into the audio latent, allowing joint audio-video denoising to generate lip movements and sound consistent with the recording. Workflows, examples, and repositories for both CUDA and Apple Silicon are provided.
More from Embodied
- Hands-on robotics workshop on Saturday may be the last in-person session before August — StewartalsopIII · 2026-07-22
- NVIDIA pitches World Foundation Models as a way to scale physical AI data generation — MonaJalal_ · 2026-07-22
- RoboMME Podcast Preview: Benchmarking Memory for Robotic Policies — chris_j_paxton · 2026-07-21
- Gritt says an 8-person crew now installs 3,000 to 4,000 solar panels a day — HaktanSuren · 2026-07-21
- A helium-powered flying robot whale aims to be a quiet companion pet — chris_j_paxton · 2026-07-21
- Orlando robotaxi ride goes unsupervised in a Model Y — aelluswamy · 2026-07-21