DeCoPrune Prunes 85% of Video Diffusion KV Cache Training-Free, 4x Faster Continuation
mmlab-ntu · hf · 2026-10-08
NTU MMLab introduces DeCoPrune, a training-free KV-cache pruning method for autoregressive video diffusion models.
- Problem: KV cache grows unboundedly with generated history; existing fixed-window or similarity-based token selection can't directly measure whether a new chunk adds information beyond retained context.
- Method: Frames cache compression as a denoising-consistency problem — empirically, denoising difficulty proxies a token's long-term retention value. Tokens with larger discrepancy between intermediate clean predictions and final denoised values carry less predictable visual evidence and are kept in long-term cache.
- Benchmark: Releases CMBench with 58 1-minute context episodes and 116 Reappear/Revisit continuation tasks testing recall of previously observed objects/scenes.
- Results: On LingBot World v2, prunes over 85% of historical KV tokens while preserving near-FullKV long-range recall and accelerating continuation generation by 4x+, substantially beating compression baselines.
More from Research
- Chrome silently ships a 425KB neural echo cancellation model; reverse-engineer finds 5.4dB gain — chadwallacehart · 2026-10-08
- Snorkel scales open benchmark grants 10x to $30M, adds red team for benchmarks — ajratner · 2026-10-08
- AI Theorem Prover for Hardware Formal Verification Presented at REBASE 2026 Industry Track — satnam6502 · 2026-10-08
- 1985 practitioner sides with Schmidhuber: backprop history isn't the US narrative — SchmidhuberAI · 2026-10-08
- NVIDIA's PivotOPD teaches agents to prevent and recover from pivotal early mistakes — NVIDIAAI · 2026-10-08
- Scott Aaronson's 'The Mathocalypse': math education and research in the age of capable AI — TMWNN · 2026-10-08