Inference Auctions: bidding for scarce LLM serving capacity without breaking KV-cache gains

nhaghtal · x · 2026-10-10

The author argues that when compute is scarce in LLM inference serving, scheduling still relies on a handful of coarse priority tiers that don't let users express latency sensitivity—real-time chat, a coding agent awaiting its next step, and overnight research queries have very different needs. The proposal is to use auctions so users can bid to express preferences over scarce inference resources. But inference isn't a normal good: modern serving leans heavily on KV-cache reuse and optimized scheduling, and naively serving highest bidders first destroys those latency gains. Their Inference Auctions provably maximize welfare without giving up these system optimizations or hurting latency.

Related event: Inference Auctions: Michael Jordan and Colleagues Propose Auction-Based Allocation of LLM Inference Compute(6 posts)→

Original post →

More from Infra

Infra channel →