One repeated sentence keeps a character’s face and voice consistent across shots

Minute_Eye_6270 · reddit · 2026-07-23

The author says they combined JoyAI-Echo’s cross-shot character memory with LTX-2.3’s voice so that one repeated sentence can keep both a face and a voice consistent across every shot.

What’s included

The post is essentially a showcase of a video-generation pipeline that preserves identity across scenes while experimenting with different quantization formats.

Related event: New Workflow Achieves Multi-Shot Audio-Visual Consistency for Characters(2 posts)→

Original post →

More from Multimodal

Multimodal channel →