One repeated sentence keeps a character’s face and voice consistent across shots

Minute_Eye_6270 · reddit · 2026-07-23

The author says they combined JoyAI-Echo’s cross-shot character memory with LTX-2.3’s voice so that one repeated sentence can keep both a face and a voice consistent across every shot.

What’s included

The post is essentially a showcase of a video-generation pipeline that preserves identity across scenes while experimenting with different quantization formats.

Original post →

More from Multimodal

Multimodal channel →