vLLM baseline took 6.84h for an AI-SQL query estimated at 15min speed-of-light

sh_reya · x · 2026-09-25

The Quail team shared baseline numbers: running all LLM calls through a hand-tuned vLLM setup caused heavy host overhead, and vLLM discarded KV cache it needed later, forcing 50 million extra tokens of processing.

One query took 6.84 hours versus a speed-of-light estimate of just 15 minutes — a 27x gap showing that query planning and inference scheduling must be co-designed.

Related event: UC Berkeley Open-Sources Quail, an AI-SQL Engine Hitting 1B+ Tokens/min on a Single H100(8 posts)→

Original post →

More from Infra

Infra channel →