DeepSeek V4.1 Flash architecture reset cuts KV cache to a quarter

DeepSeek has released V4.1 Flash, which—despite the model being roughly twice the size of its predecessor—compresses the working KV cache to 1/4 and persistent cache storage to 1/8, with measured speeds approaching 420 tokens/s. Multiple independent commentators call it "a masterpiece of AI engineering and research," and one architecture analyst even argues it deserves to be called "DeepSeek-V5 Flash." The model achieves an architectural reset by actively deleting its own older designs, sharply cutting cache-holding costs for long agent sessions, though the community remains puzzled by the mechanics of the 190B+ engram lookup table.

Confirmed

Not yet confirmed

Why it matters

2026-09-15 ~ 2026-09-17 · 5 related posts

Primary sources