Does DeepSeek V4.1-Flash's SWA Bounded Replay sacrifice recall to save KV cache memory?
Top-Handle-5728 · reddit · 2026-09-11
A Reddit post questions DeepSeek-V4.1-Flash's SWA Bounded Replay: the model discards sliding-window KV cache and approximately reconstructs it from only the last Nwindow tokens on replay, saving substantial memory.
The author argues the original SWA states carry a much larger effective receptive field due to depth, so reconstructed K/V vectors lack the long-range context that shaped the originals. They ask whether cache hits or restarted sessions show recall degradation in early interactions, and whether the parallel global KV augmentation compensates enough. They seek anyone who has benchmarked this or observed practical recall loss after eviction.
More from Models
- GPT-6 Astra drives a robot arm on first try via physical ICL; Ken Goldberg touts Agentic Robotics — zhaoran_wang · 2026-09-11
- OpenAI's internal model claims a Navier-Stokes millennium prize proof, says analyst — QuintinPope5 · 2026-09-11
- DeepSeek 4.1 flash reportedly uses large ngram embeddings, echoing Qwen4 architecture — ccerrato147 · 2026-09-11
- ValsAI launches RSI Index, first third-party benchmark measuring how close AI is to self-improvement — JenniferHli · 2026-09-11
- Assistant Benchmark goes live: 61 assistants scored across 15 real-use dimensions — Scobleizer · 2026-09-11
- OpenAI pauses new $200 ChatGPT Pro signups as GPT-6 Astra demand overwhelms capacity — 机器之心 · 2026-09-11