How to Attribute LLM Inference Costs by Team
Extreme_Tangelo8336 · reddit · 2026-07-18
The author notes that as LLMs scale from isolated features to internal tools, agents, customer service workflows, and evaluations, inference cost attribution becomes tricky:
- Vendor consoles typically show only token usage, while finance needs to know exactly which team or project incurred the cost.
- Infra teams see raw usage, and finance sees a single total bill, but middle-tier attribution capabilities remain immature.
- The author suggests combining application-layer tagging with internal reporting, asking the community how they currently manage this formally versus treating it as shared infrastructure overhead.
More from Infra
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11
- Your p99 latency benchmark may be lying: a deep dive into coordinated omission — Franc0Fernand0 · 2026-09-11
- Running MiniMax H3 on 12GB VRAM: quantization, Turbo LoRAs and attention backends compared — Possible_Mood676 · 2026-09-11
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11