How a Self-Appending Summarizer Quadrupled Token Costs in a 1.4M-Conversation Agent
Good_Education4713 · reddit · 2026-10-08
An engineer dissects a cost anomaly in a 9-turn support agent handling 1.4M conversations/month: input tokens grew from 5,800 (turn 2) to 38,400 (turn 9), pushing p95 cost from $0.19 to $0.83 with no quality gain. Root cause: the summarizer appended the previous summary into each new one instead of replacing it, replaying its own history as context. Span-level token attribution revealed the bug; replacing the summary each turn flattened the curve without quality loss. Open question: reliably bounding evolving conversation state without losing detail.
More from coding & agent
- 3 AI agent-generated optimizations merged upstream into SGLang-Omni — ChengleiSi · 2026-10-08
- Addy Osmani on coding agents: separate adversarial-review models add little over same-session prompts — addyosmani · 2026-10-08
- DuckDB shows a database with no data: a few hundred KB catalog over your data lake — RexDouglass · 2026-10-08
- Dev open-sources smoke effect from Replit-built Drift game with agent prompt workflow — techartist_ · 2026-10-08
- Engineer Pushes Back on AI Code Analogy: Topology-Optimized Parts Look Cool but Are Unusable — rms80 · 2026-10-08
- Massive powers AI trade agents with x402 micropayments at a cent per request — kleffew94 · 2026-10-08