One repeated sentence keeps a character’s face and voice consistent across shots
Minute_Eye_6270 · reddit · 2026-07-23
The author says they combined JoyAI-Echo’s cross-shot character memory with LTX-2.3’s voice so that one repeated sentence can keep both a face and a voice consistent across every shot.
What’s included
- The workflow used
- Weight formats: bf16, fp8, Q8, Q5, INT8
- A free demo Space
The post is essentially a showcase of a video-generation pipeline that preserves identity across scenes while experimenting with different quantization formats.
More from Multimodal
- Qwen-image 3.0 is now available on Runware — aziz4ai · 2026-07-23
- Directing Character Emotions: Grok Imagine Video Generation Tested — chaitu · 2026-07-23
- HeyGen’s HyperFrames adds a storyboard-first workflow for AI video generation — HeyGen · 2026-07-23
- Creator Combines Midjourney and GPT-4o for Impressive AI Art — aziz4ai · 2026-07-23
- LTX 2.3 Workflow: Direct Video Generation from Storyboards — Sad_Coach_1433 · 2026-07-23
- Seedance 2.0 prompt share stitches GPT Image 2, Midjourney and Suno together — azed_ai · 2026-07-23