MiniMax H3 in Action: Building an Audio-Driven Lip Sync Video Workflow

CountFloyd_ · reddit · 2026-08-05

A developer discovered that LTX nodes originally for audio-to-video also work with the MiniMax H3 model, prompting them to quickly hack together an audio-driven lip-sync video workflow. The setup uses a reference image and supplied audio, leveraging the faster native I2V workflow instead of R2V. It also employs an extra model to extract voice from the audio stream for optimal lip synchronization. The workflow was shared via Pastebin.

Original post →

More from Multimodal

Multimodal channel →