DeepSeek's new tech report: 890 bytes/token KV, inference-first architecture dissected
nrehiew_ · x · 2026-09-11
Researcher nrehiew dissects DeepSeek's latest tech report, calling it cleaner than v4's HSA+CSA combo. Key points: the architecture is clearly inference-first (RL blurs the training/inference line); KV size is a startling 890 bytes/token at benchmark performance; the reasoning-performance plot is oddly non-linear (possible quality penalties), while the agent swarm plot is cleaner; agent team mode is explicitly RL-trained with rewards combining task performance and collaboration. He doubts OpenAI/Anthropic would pursue such architectural frankenstein designs given their custom inference chips.
More from Models
- Novel Reasoning Effort Control Scheme Analyzed: Graded GRPO Training — stochasticchasm · 2026-09-11
- Forcing models to always max effort is like humans evolving on Adderall, researcher argues — voooooogel · 2026-09-11
- Pushing models to always show 'maximum effort' drags along its corollaries, dev argues — voooooogel · 2026-09-11
- Claude is the distillation target of choice because agentic RL seed data is scarce — teortaxesTex · 2026-09-11
- Why Chinese labs distill from Anthropic: Claude's agent data is the scarce training signal — teortaxesTex · 2026-09-11
- ApprenticeBench: closed model scores 72% vs open Kimi K3 at 18% on real jobs — ysu_nlp · 2026-09-11