Production AI budgets include retries, routing, caching and observability—not just token prices
arx-go · reddit · 2026-07-22
The post argues that production AI costs are much bigger than a model’s per-token price.
It breaks the budget into several parts:
- LLM/API usage
- Retries and failures
- Routing across multiple models
- Caching, or the lack of it
- Embeddings and vector databases
- Guardrails and moderation
- Monitoring and observability
- Infrastructure and orchestration
The main point is that AI spending should be treated as a system-level budget, not a single line item. The author links to a short visual article explaining the idea and asks how other startups budget for AI workloads.
More from Infra
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11