vLLM baseline took 6.84h for an AI-SQL query estimated at 15min speed-of-light
sh_reya · x · 2026-09-25
The Quail team shared baseline numbers: running all LLM calls through a hand-tuned vLLM setup caused heavy host overhead, and vLLM discarded KV cache it needed later, forcing 50 million extra tokens of processing.
One query took 6.84 hours versus a speed-of-light estimate of just 15 minutes — a 27x gap showing that query planning and inference scheduling must be co-designed.
More from Infra
- Powering a home GPU cluster: one PDU safely feeds ~16 RTX 6000 Max-Q cards — TheZachMueller · 2026-09-25
- Docker launches Cloud Sandboxes: microVM isolation for always-on agents, $250 credit — juntao · 2026-09-25
- Muse Spark 1.3 Now Available via Oracle, in Private Preview on Google Cloud — alexandr_wang · 2026-09-25
- Andy Matuschak: Modern Chips Are Too Durable for Compute-Control-Based AI Governance to Hold — andy_matuschak · 2026-09-25
- Qualcomm touts double-digit MLPerf performance and efficiency gains at Snapdragon Summit — samcharrington · 2026-09-25
- HEIF Heist: image parser RCE chain nets $100k Meta bounty, hits OpenAI repos and more — evilsocket · 2026-09-25