QuantWM: Training-Free 2-Bit KV Cache Quantization for Video World Models, 6.2x Compression

NationalUniversityofSingapore · hf · 2026-10-06

Existing 2-bit KV cache quantization looks lossless on VBench but causes severe temporal flickering in video world models. The authors show Key quantization has smaller reconstruction error than Value yet larger output degradation, because Key perturbations shift attention logits and the spatiotemporal tokens Queries select.

QuantWM is a training-free framework with two techniques:

Across LingBot-World-v2, HY-World 1.5, Matrix-Game-2, Longcat-Video, and Causal-Forcing it markedly improves visual quality and temporal consistency, with up to 6.20x KV cache compression at low overhead.

Original post →

More from Research

Research channel →