Does DeepSeek V4.1-Flash's SWA Bounded Replay sacrifice recall to save KV cache memory?

Top-Handle-5728 · reddit · 2026-09-11

A Reddit post questions DeepSeek-V4.1-Flash's SWA Bounded Replay: the model discards sliding-window KV cache and approximately reconstructs it from only the last Nwindow tokens on replay, saving substantial memory.

The author argues the original SWA states carry a much larger effective receptive field due to depth, so reconstructed K/V vectors lack the long-range context that shaped the originals. They ask whether cache hits or restarted sessions show recall degradation in early interactions, and whether the parallel global KV augmentation compensates enough. They seek anyone who has benchmarked this or observed practical recall loss after eviction.

Original post →

More from Models

Models channel →