DeepSeek's new tech report: 890 bytes/token KV, inference-first architecture dissected

nrehiew_ · x · 2026-09-11

Researcher nrehiew dissects DeepSeek's latest tech report, calling it cleaner than v4's HSA+CSA combo. Key points: the architecture is clearly inference-first (RL blurs the training/inference line); KV size is a startling 890 bytes/token at benchmark performance; the reasoning-performance plot is oddly non-linear (possible quality penalties), while the agent swarm plot is cleaner; agent team mode is explicitly RL-trained with rewards combining task performance and collaboration. He doubts OpenAI/Anthropic would pursue such architectural frankenstein designs given their custom inference chips.

Related event: DeepSeek V4.1 Tech Report Deep Dive: RL Infrastructure, Sandbox Design and Inference Stack(8 posts)→

Original post →

More from Models

Models channel →