Inference Auctions: bidding for scarce LLM serving capacity without breaking KV-cache gains
nhaghtal · x · 2026-10-10
The author argues that when compute is scarce in LLM inference serving, scheduling still relies on a handful of coarse priority tiers that don't let users express latency sensitivity—real-time chat, a coding agent awaiting its next step, and overnight research queries have very different needs. The proposal is to use auctions so users can bid to express preferences over scarce inference resources. But inference isn't a normal good: modern serving leans heavily on KV-cache reuse and optimized scheduling, and naively serving highest bidders first destroys those latency gains. Their Inference Auctions provably maximize welfare without giving up these system optimizations or hurting latency.
More from Infra
- llama.cpp only hits 8-10 t/s on a 4bit 27B while bitsandbytes + transformers manages 27 t/s — cephaloform · 2026-10-10
- Two H100 price indices show just 0.17 weekly correlation, clouding compute futures hedge — BenBajarin · 2026-10-10
- Chrome's new echo canceller halves voice agent word error rate, stops agents answering their own greeting — chadwallacehart · 2026-10-10
- Price war math: does cheaper AI tokens boost or drain compute investment? — dnlkwk · 2026-10-10
- New Paper 'Inference Auctions' Brings Market Mechanisms to LLM Inference Serving — nhaghtal · 2026-10-10
- SGLang-Diffusion serving framework for diffusion models to be unveiled at PyTorch Conference 2026 — PyTorch · 2026-10-10