LTX-2.3 face-and-voice LoRA training can work on 12GB VRAM with heavy tradeoffs

__alpha_____ · reddit · 2026-07-21

A Reddit user reports that training an LTX-2.3 face-plus-voice LoRA on an RTX 3060 with 12GB of VRAM is possible, but only with careful setup.

The post describes a long debugging process and recommends a clean dataset of short 2–5 second face clips with clear dialogue and no background music. To avoid out-of-memory errors and excessive training time, the author suggests offloading, unloading the text encoder, 4-bit quantization, rank 4, cache latents, auto frame count, and audio training. After several days of retries, they got a working 1,200-step rank-4 LoRA, though the result is still rough and mainly proves the workflow can be made to work.

Original post →

More from Multimodal

Multimodal channel →