DeepSeek pushes single-token KV cache under 1KB—at brutal infra cost
kalomaze · x · 2026-09-11
Researcher kalomaze notes DeepSeek continues its efficiency trend, somehow pushing single-token KV cache storage costs into the sub-kilobyte regime—which he calls the most agonizing, gut-wrenching, misery-inducing possible training and serving infrastructure regime.
More from Infra
- NVIDIA and Red Hat team up with vLLM on a bounty for real builds with small open models — NVIDIAAI · 2026-09-11
- Dev questions whether OpenAI's prompt_cache_key design wastes massive compute — YouJiacheng · 2026-09-11
- SF Compute signs $245M in take-or-pay contracts for NVIDIA Blackwell B300 capacity — mattshumer_ · 2026-09-11
- SpaceX CFO: vertical integration is core, Starship paves way for orbital compute — elonmusk · 2026-09-11
- Bezos: Power Supply Chain Bottleneck Forces AI Labs to Slow Development Pace — beffjezos · 2026-09-11
- vLLM v0.29.0 cuts Blackwell E2E latency 33.6%, with 6.6-7.6x kernel speedups for Kimi-K3 — vllm_project · 2026-09-11