The hidden cost of agents is KV cache: DeepSeek compresses to ~890 bytes per token
altryne · x · 2026-09-12
Chris Alexiuk points out that the quiet cost lever for agents is the KV cache — cache hits vs misses make up most of the bill. DeepSeek's compression path gets it down to 890 bytes per token, keeping more context hot without burning GPU, which is key to affordable long-context agent inference.
More from coding & agent
- DeepSeek V4-Pro-0813 and GLM-5.3 Launch on Nebius, Targeting Coding Agents — Arindam_1729 · 2026-09-12
- Every MCP tool-poisoning detector is blind on the first tools/list — here's a baseline-free fix — blitzcrieg11 · 2026-09-12
- Brex CEO argues PM playbooks are dead, demos running his $5B company via OpenClaw — petergyang · 2026-09-12
- Three runs, two significant losses: your cheap-model switch may just have been luck — Ok-Challenge-7810 · 2026-09-12
- bough visualizes your Claude Code and Codex session history locally, open source — VoidEqualZero · 2026-09-12
- GPT-6 Astra Powers an Open-Source AI Software Factory That Ships Code 24/7 — Cole Medin · 2026-09-12