Why DeepSeek's Cache Rates Are So Hard to Replicate, Explained
thdxr · x · 2026-09-19
Responding to mitsuhiko's question about why DeepSeek achieves better aggregate cache rates than others, thdxr argues the cache rate reflects overall engineering excellence: DeepSeek does "1000 small things" across datacenter buildout and the software stack, all heavily optimized for its specific models and products. Even if you manage to rent 1000 GPUs, a generic setup will never match it. He notes OpenAI and Anthropic are also excellent, but was referring to models others can host.
More from Infra
- The rig built to run Emacs and doomscroll X is now worth more than its owner's car — tetsuoai · 2026-09-19
- Apple M4 sustains 10 instructions per cycle, beating most rivals; M5 speedup explained — lemire · 2026-09-19
- Apple M6 bumps cores to 12 with two super cores; CPUs keep improving fast — lemire · 2026-09-19
- Apple M-series chips gained ~50% Geekbench 6 performance over three years — lemire · 2026-09-19
- Inside OpenAI's inference routing: why the proportional controller had to go — AI Engineer · 2026-09-19
- "Normal people can't afford 2x DGX Spark": local AI hardware cost debate — FlolightC · 2026-09-19