How to keep LTX 2.3 ad videos readable when text starts degrading after 7 seconds

Opposite_Working5185 · reddit · 2026-07-22

A Reddit user asks how to get LTX 2.3 to render social-media ad videos with accurate on-screen text. They say text starts degrading around the 7-second mark, and that making the text larger and placing it on a flat horizontal plane helps, but not enough for smaller copy.

They also describe a second pain point: animating a person to point to specific parts of a computer or product screen. In practice, the model mostly produces vague gestures rather than precise pointing.

The user is looking for a better local workflow for this kind of electronic-product ad work, and notes that generating the static image first in ChatGPT has worked surprisingly well for them.

Original post →

More from Multimodal

Multimodal channel →