LTX-2.3 Workflow: Audio Drives Facial Expressions and Body Language

egeberkina · x · 2026-07-29

A creator shared a hands-on workflow using the LTX-2.3 Pro model for audio-to-video generation. By uploading a start frame, an audio track, and a prompt, the model generates a video where the audio does more than just sync lips. It actively influences the timing, facial expressions, body movements, and overall pacing of the shot.

Related event: LTX-2.3 Tested: Generates Highly Synced Video from Single Image and Audio(2 posts)→

Original post →

More from Multimodal

Multimodal channel →