Agent Stuck in Retry Loop Burns $4,700 Overnight on a $400/mo Customer
Sufficient_Cause_43 · reddit · 2026-10-06
A lean 3-person B2B tool runs agents on its own infra for 3 customers, paying model bills upfront. Between 1:12am and 6:50am, a customer's agent hit a tool that kept erroring and decided to retry with slightly different prompts — 31,000 times, burning $4,700 in tokens on a $400/mo plan.
Postmortem: they had global rate limits and total-spend alerts, but no per-customer cap, since usage only rolled up into month-end invoices nobody checked. The customer wasn't at fault — agents do exactly this when nothing tells them to stop.
Fixes worth copying:
- Check a per-customer budget before every model call
- Soft alert at 1x plan, hard stop at 2x
- Add a max retry count per tool call
More from coding & agent
- A2A Protocol ships official CLI to discover, message and manage agents from the terminal — rseroter · 2026-10-06
- 462GB DeepSeek model runs on two desk-side DGX Sparks with experts squeezed to 2.77 bits — Teknium · 2026-10-06
- Dev pitches 'Business in a Box': ship a whole AI agent company inside one Claude Code account — RileyRalmuto · 2026-10-06
- Armin Ronacher: a Pi release took just 4 hours thanks to GitHub Actions — mitsuhiko · 2026-10-06
- Building a Proactive Personal Agent With Pi Durable: Checkpoints, Memory, Own Linux Desktop — omarsar0 · 2026-10-06
- ThePrimeagen's Desktop QA Agent: Cut QEMU Runs From 15 Minutes to 2, Still Too Slow — cyrus_zei · 2026-10-06