Our AI bill hit $11,400 with no attribution: a cautionary tale of unbounded retries and prompt bloat

vigilAPI · reddit · 2026-10-10

A team running 15 LLM-backed features saw inference spend grow 8x in a quarter to $11,400/month and couldn't attribute it — all calls hid inside SDKs scattered across the codebase. Post-mortem found three culprits: an uncapped retry helper that turned one call into forty for six weeks, a summarization prompt that bloated from 1.2k to 9k tokens (7x cost), and a customer's injection-like prompts they paid for. Fix: route all provider traffic through one proxy for per-feature/per-key cost attribution, and actually cap your retries.

Original post →

More from coding & agent

coding & agent channel →