TurboT2VA distills LTX-2 to 4 steps, cutting audio-video generation time by 54.67x
量子位 · wechat · 2026-09-24
- Three-stage curriculum distillation: dCM warmup → sCM refinement → sCM+DMD joint optimization. Discrete consistency distillation first stabilizes a 19B student's joint denoising structure, then continuous consistency preserves the teacher's trajectories, and distribution matching (DMD) improves perceptual quality. Ablations show staged training beats direct joint training on JavisScore.
- Modality-specific handling: separate diffusion timestep shifts for audio and video, per-modality loss normalization, paired latents through a joint Transformer so each modality can attend to the other.
- Quantization stack: W8A8 INT8 for linear layers, SageSLA attention with INT8 Q/K and FP8 V on H20, operator fusion and text-padding removal.
- Results: at 512×768, generator time drops from 50.52s to 2.51s (20x) with JavisScore improving 0.165→0.196; at 1024×1792 on one H20, 318.74s → 5.83s (54.67x). Paper: arxiv.org/abs/2608.24674, code: thu-ml/TurboDiffusion.
Related event: TurboT2VA Speeds Up LTX-2 Audio-Video Generation 54.67x via Distillation(2 posts)→
More from Multimodal
- Reddit User Explores AI Art With Only Steps, CFG and Denoise Tweaks, No LoRAs — Extreme_Nice · 2026-09-25
- Pose Blueprint: A Browser-Based 3D Pose Editor for ComfyUI and ControlNet — OkConfusion6667 · 2026-09-25
- World Labs Previews Chisel in Atlas Beta: Sketch a World and Bring It to Life — theworldlabs · 2026-09-25
- A sub-$20 LoRA makes Qwen-Image 2.1 rotate transparent objects with a prompt — ben_burtenshaw · 2026-09-25
- Lingbot World v2 runs at 60 FPS, hinting world models could reshape game dev — bingxu_ · 2026-09-25
- Image-video feedback loop yields bizarre organic shape-shifting dance moves — pixlpa · 2026-09-25