Dev builds tiny logger to find what eats OpenAI budget — token counts pointed to the wrong feature
Atm1n9 · reddit · 2026-10-11
The author wanted to know which feature in their app was burning their OpenAI budget, so they built a small logger: after each call, log the usage object's token counts with a feature tag and price them manually.
Pricing yourself is trickier than expected: cached reads bill differently from normal input, and cache writes can cost more; reasoning tokens bill as output though invisible; providers disagree on whether cached tokens sit inside the prompt count (the author double-counted for a while); audio tokens can cost several times the text rate on the same model.
Others got hit harder: one Google AI dev forum user saw €290 in search charges attached to €10 of actual LLM usage, bypassing their spend cap; a DEV post described a timeout-plus-auto-retry on long-context calls quietly doubling costs while every dashboard stayed green; one startup reported a $113k single-month AI bill.
Counterintuitive result: the code-gen feature did eat 70% of the bill, but by token count the summarizer was far bigger (129M vs 30M) yet cost only $14 total — token counts alone would have blamed the wrong feature. The logger later grew into a product, LLMtrack.
More from coding & agent
- Anatomy: a Claude Code skill that turns any idea into an interactive machine drawing — _AustinCalvert_ · 2026-10-11
- Fine-tuned SD1.5 generates 16×16 Minecraft item sprites you can drop into resource packs — yuuki202800 · 2026-10-11
- 16x Microsoft MVP demos Jev, a PowerShell AI workflow for ranking and routing requests — dfinke · 2026-10-11
- rec-rs: a clean-room Rust rewrite of OBS for macOS using only Apple frameworks — Rasmic · 2026-10-11
- Provenance gateway turns AI agent tool calls into audit-grade, tamper-evident records — hutsonlabs · 2026-10-11
- Evals are replays: store tasks in R2, sandbox in Docker, and 90% of the work is measuring — danshipper · 2026-10-11