Token compression can miss the real bill: transcript resend and cache writes dominate
KitchenAmoeba4438 · reddit · 2026-07-25
Context compression often targets the wrong cost layer
The post argues that token compression usually saves the latest message, while the real bill comes from repeatedly resending the entire transcript in tool loops.
- In a long agent loop, turn 40 pays again for turns 1–39, so costs scale roughly with the square of turn count.
- Cache reads are cheaper than base input, but a high hit rate on a large prefix can still be expensive because you keep paying to reread old context.
- Cache writes are even more costly, and because Anthropic-style caching is prefix-based, editing anything inside the cached prefix can force a rewrite of the tail.
- The practical advice: order context by mutation rate, avoid editing inside a cached prefix, and summarize at a boundary where you intentionally rebuild the prefix.
The author also notes that token counters can be misleading if you are not tracking cache creation vs. cache reads, and suggests measuring cost per successful task in paired runs instead of trusting token savings alone.
More from coding & agent
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- 105 hidden bugs, 2 repos: DeepSeek V4.1 Flash fixes 24 at $1.80 vs Opus 5's 27 at $51.33 — ChartsJournalX · 2026-09-11
- Investment Analyst Asks How to Build a Claude-Based Diligence Agent Stack — Careless_Tie2286 · 2026-09-11
- Treating agents like 50 First Dates: a 3-layer context system so every conversation doesn't start from zero — evielync · 2026-09-11
- Running the Firefox MCP on Android via Termux, ngrok, and mcp-proxy — Nervous-Strain7544 · 2026-09-11