A 5-second talking-head loop in 6GB VRAM: local graph, hosted lip-sync API

Extreme-Shock1930 · reddit · 2026-10-01

Pushing back against 40-node talking-head workflows, the author built the smallest possible graph: load video, load audio, plus a sync custom node set (video in, audio in, api key, generate, output, save) that turns a 5-second clip and a voice recording into a synced video.

Loading and preview run on a 3060 with barely any VRAM movement; the actual lip-sync is a hosted API call, so local GPU specs don't matter — you pay per second of output instead of waiting on VRAM.

Open questions: where to cache results so a graph re-run doesn't re-bill the same segment, and whether to keep the API key in the graph JSON or in env. He asks how others split local vs hosted stages.

Original post →

More from coding & agent

coding & agent channel →