DeepSeek's New Model: 4x Smaller KV Cache Than DSV4-Flash and More Stable Training
stochasticchasm · x · 2026-09-11
Discussing DeepSeek's new model tech report, stochasticchasm highlights: strikingly high benchmark scores, a KV cache 4x smaller than DSV4-Flash, a break from the neat v<N>-equals-new-pretrain release convention, and an architecture that looks like a simplification of V4 — with claims it trained far more stably.
Related event: DeepSeek's New Model Report Highlights 4x Smaller KV Cache(3 posts)→
More from Infra
- 'Legacy Infrastructure' Is Suddenly the Future: Why Enterprise AI Is Moving Back On-Prem — DavidLinthicum · 2026-09-11
- ThunderKittens lands on NVIDIA Vera Rubin, pushing NVFP4 GEMMs past 22 PFLOPS — togethercompute · 2026-09-11
- Open-Source Go Gateway Stops Runaway Agent Loops and Attributes LLM Spend by Run — SnooCauliflowers2631 · 2026-09-11
- Can 2x RTX 3090 Plus 512GB DDR5 Reach 15 t/s on Large Local LLMs? Redditor Asks — levoniust · 2026-09-11
- Rebuilding a homelab with remote agents: tools and checks matter more than the LLM — HankYeomans · 2026-09-11
- LayerLens: Open-Source Profiler Breaks Down LLM Inference Timing by Token and Layer — Dry_Mixture130 · 2026-09-11