TurboT2VA distills 19B text-to-video-audio model to 4 steps for 54.67x speedup
青稞AI · wechat · 2026-09-25
Tsinghua's TurboT2VA accelerates LTX-2, a 19B open-source joint text-to-video-audio model. Training distills the 40-step teacher into a 4-step student via a three-stage curriculum: discrete consistency distillation warmup, continuous consistency distillation refinement, then joint sCM+DMD optimization—establishing trajectory structure first before distribution matching, with modality-specific timestep shifts and loss normalization. Inference stacks W8A8 quantization, SageSLA sparse attention, text-padding compression and fused kernels.
At 1024×1792 and 121 frames on a single NVIDIA H20, generator latency drops from 318.74s to 5.83s, a 54.67x speedup, while audio-visual alignment metrics improve. Paper and code (thu-ml/TurboDiffusion) are public.
Related event: TurboT2VA Speeds Up LTX-2 Audio-Video Generation 54.67x via Distillation(2 posts)→
More from Multimodal
- User generates a music video from old material with Opus 5.5 — repligate · 2026-09-25
- Same prompt, Opus 5.5 one-shot video generation put to a public retest with different tools — drrickio · 2026-09-25
- One prompt: Claude agent wired to Runway MCP delivers a Netflix-style superintelligence doc — CurieuxExplorer · 2026-09-25
- Seedance 2.5 + GPT Image 2.5 One-Shot the Most Iconic Sci-Fi Rivalry — CurieuxExplorer · 2026-09-25
- Opus 5.5 writes songs and music videos entirely from code in a playable demo — pbaylies · 2026-09-25
- Study: Reasoning hurts 15.7% of multimodal embeddings; training-free SURE router fixes it — _reachsumit · 2026-09-25