DeepSeek-V4.1 Flash deep dive: pushing KV cache compression to the limit at 420 tok/s

teortaxesTex · x · 2026-09-17

A detailed architecture analysis of the DeepSeek-V4.1 Flash technical report, arguing it is substantial enough to be called "DeepSeek-V5 Flash".

Context: the model runs at nearly 420 tokens/s in practice, and DeepSeek took all V4 Pro models offline after release — signaling an architecture-level generational change, not a post-training iteration.

Key optimizations:

Related event: DeepSeek V4.1 Flash architecture reset cuts KV cache to a quarter(5 posts)→

Original post →

More from Infra

Infra channel →