6 teams, one opaque LLM bill: a postmortem on tagging, proxies and gateway budget caps
Fit_Program7076 · reddit · 2026-09-17
When finance asked which teams were driving the company's LLM bill, nobody could answer: six teams hit providers directly with their own API keys, and invoices arrived as one line item per provider. The author walks through three fixes: manual per-team reporting (broke down after six weeks), a lightweight proxy adding team headers (visibility but only after the money is spent), and gateway-level enforcement with per-team/per-project caps (e.g. via orqai), which blocks overspend before the request goes out and gives finance real-time spend. The open question they're wrestling with: handling burst usage when a team legitimately needs to exceed its cap. Note the post carries a promotional slant toward one vendor.
More from Infra
- GlobalFoundries and Marvell expand Vermont SiGe capacity for AI optical networking — zephyr_z9 · 2026-09-17
- $26,100 desktop AI datacenter: dual RTX PRO 6000 Blackwell workstation goes open source — dee_hw · 2026-09-17
- Huawei Rumored Scale-Up Node with 4,096 Accelerators Could Pack 384TB of HBM — zephyr_z9 · 2026-09-17
- Brad Gerstner: 43GW AI compute forecast too aggressive, sees ~25GW with half for OpenAI/Anthropic — firstadopter · 2026-09-17
- Cohere's CUDA Megakernel Serving Hits 292 tok/s at Batch 1 on a 30B Model — dl_weekly · 2026-09-17
- Stanford's Ousterhout: AI workloads are outgrowing TCP — AI Engineer · 2026-09-17