LightX2V Enables 14B Video Model on Single RTX 5090 with 720p Real-time Generation
新智元 · wechat · 2026-08-07
The LightX2V team introduced an extreme acceleration framework for the Wan2.2-A14B video generation model, drastically reducing computational load and VRAM peak usage via step distillation, NVFP4 quantization-aware training, and dynamic sparse attention.
- Core Optimizations: Uses PhasedDMD to compress denoising steps from 40 to 4. Combines NVFP4 quantization and dynamic sparse attention to lower per-step compute overhead.
- Single GPU Performance: The 14B model runs efficiently on a single RTX 5090, generating a 5-second 720p video in just 22.5 seconds, achieving over 100x speedup.
- Multi-GPU Real-time Generation: On an 8-GPU setup, T2V and I2V for 5-second 720p videos take only 3.8s and 4.5s respectively, both shorter than playback time, achieving up to 705.8x acceleration.
More from Infra
- ByteDance Rumored to Pre-Train 10T Parameter Model; Distillation Predicted for Serving — zephyr_z9 · 2026-08-07
- AI Inference is Memory-Constrained: A Shift Could Bullish Memory Chips — toptickcrypto · 2026-08-07
- Inference Compute Jumps to Two-Thirds of AI Workloads, Reshaping Profit Pools — msharmas · 2026-08-07
- Analysis: LLM Inference Value Shifts from Engines to Data Centers and GPU Capacity — zhyncs42 · 2026-08-07
- Nvidia Seeks China 6G Base Station Suppliers; Microsoft Expands India Cloud — 创业邦 · 2026-08-07
- Winbond Expands to Become Top SLC NAND Supplier by 2027 Amid AI Demand — zephyr_z9 · 2026-08-07