TurboT2VA distills 19B text-to-video-audio model to 4 steps for 54.67x speedup

青稞AI · wechat · 2026-09-25

Tsinghua's TurboT2VA accelerates LTX-2, a 19B open-source joint text-to-video-audio model. Training distills the 40-step teacher into a 4-step student via a three-stage curriculum: discrete consistency distillation warmup, continuous consistency distillation refinement, then joint sCM+DMD optimization—establishing trajectory structure first before distribution matching, with modality-specific timestep shifts and loss normalization. Inference stacks W8A8 quantization, SageSLA sparse attention, text-padding compression and fused kernels.

At 1024×1792 and 121 frames on a single NVIDIA H20, generator latency drops from 318.74s to 5.83s, a 54.67x speedup, while audio-visual alignment metrics improve. Paper and code (thu-ml/TurboDiffusion) are public.

Related event: TurboT2VA Speeds Up LTX-2 Audio-Video Generation 54.67x via Distillation(2 posts)→

Original post →

More from Multimodal

Multimodal channel →