LibTV Adds Depth-Video Extraction: Copy Motion and Camera Work, Not Faces

卡尔的AI沃茨 · wechat · 2026-09-01

When using reference videos for AI video generation, the original faces, clothing, and background often leak into the output. This post describes a fix: convert the source video to a depth video first, then feed that to the model — motion, camera moves, positioning, and spatial relations are preserved while identity and style interference drops sharply. LibTV now has one-click depth extraction: drop the source into the canvas, convert to grayscale depth, and pipe it into downstream generation nodes.

The author tested chase/fight scenes, classic orbiting camera work, and Monument Valley–style spatial transformations, with strong motion replication and face stability. Caveats: depth reference works best on live-action with clear spatial layers and large motion; flat anime or static shots benefit little. The post also covers LibTV's new talking-head smart editing: auto-removing filler, rough-cut confirmation, auto review, and natural-language tweaks, all in one canvas.

Original post →

More from coding & agent

coding & agent channel →