Our AI bill hit $11,400 with no attribution: a cautionary tale of unbounded retries and prompt bloat
vigilAPI · reddit · 2026-10-10
A team running 15 LLM-backed features saw inference spend grow 8x in a quarter to $11,400/month and couldn't attribute it — all calls hid inside SDKs scattered across the codebase. Post-mortem found three culprits: an uncapped retry helper that turned one call into forty for six weeks, a summarization prompt that bloated from 1.2k to 9k tokens (7x cost), and a customer's injection-like prompts they paid for. Fix: route all provider traffic through one proxy for per-feature/per-key cost attribution, and actually cap your retries.
More from coding & agent
- Three plugins to squeeze more token savings out of the Pi coding agent — solyarisoftware · 2026-10-10
- AI-Generated Proposals Are Making Reviewers Work Harder Than the Authors — yangyi · 2026-10-10
- "Superintelligence at $5/Task, But AI Mods Still Suck" — Devs Complain — Aryvyo · 2026-10-10
- TokenMaster launches: a prettier ccusage-style AI token usage tracker — cneuralnetwork · 2026-10-10
- Matt Pocock: Turn Repeated Tasks into CLI + Skills Agents Own to Cut Tokens and Hallucinations — mattpocockuk · 2026-10-10
- Claude Opus 5.5 Designs a Machine That Builds Machines, Producing a Working Relay in 27 Layers — dosco · 2026-10-10