Why DeepSeek's Cache Rates Are So Hard to Replicate, Explained

thdxr · x · 2026-09-19

Responding to mitsuhiko's question about why DeepSeek achieves better aggregate cache rates than others, thdxr argues the cache rate reflects overall engineering excellence: DeepSeek does "1000 small things" across datacenter buildout and the software stack, all heavily optimized for its specific models and products. Even if you manage to rent 1000 GPUs, a generic setup will never match it. He notes OpenAI and Anthropic are also excellent, but was referring to models others can host.

Original post →

More from Infra

Infra channel →