QuantWM: Training-Free 2-Bit KV Cache Quantization for Video World Models, 6.2x Compression
NationalUniversityofSingapore · hf · 2026-10-06
Existing 2-bit KV cache quantization looks lossless on VBench but causes severe temporal flickering in video world models. The authors show Key quantization has smaller reconstruction error than Value yet larger output degradation, because Key perturbations shift attention logits and the spatiotemporal tokens Queries select.
QuantWM is a training-free framework with two techniques:
- Quantization-sensitivity-aware clustering (QSAC) picks INT2-friendly Key centroids using historical Query sensitivity and residual ranges
- Principal-subspace attention compensation (PSAC) restores residual Key errors via low-rank projections along the dominant Query subspace
Across LingBot-World-v2, HY-World 1.5, Matrix-Game-2, Longcat-Video, and Causal-Forcing it markedly improves visual quality and temporal consistency, with up to 6.20x KV cache compression at low overhead.
More from Research
- Subsampling and extrapolation keep the Mandelbrot area estimate unbiased near the boundary — geoffreyirving · 2026-10-06
- New estimate sits 6.5e-9 below Hsing Lo's 2025 value; reproduction suggests the gap is a fluctuation — geoffreyirving · 2026-10-06
- Claude-assisted CUDA compute pins Mandelbrot set area to 1.506591883653, 60x tighter than 2012 record — geoffreyirving · 2026-10-06
- Group-Evolving Agents: a new paradigm where the unit of agent self-improvement is a group — xwang_lk · 2026-10-06
- SLIM paper at COLM: design principles for long-horizon agentic search systems — xiye_nlp · 2026-10-06
- The Nobel optogenetics drama: forgotten inventor Zhuo-Hua Pan had the stronger claim — _onionesque · 2026-10-06