Google engineers: LLM benchmark harnesses silently drop requests — 200 QPS in, 38 out
AI Engineer · youtube · 2026-09-20
Google engineers Ashok Chandrasekar and Jason Kramberger show why many published LLM inference benchmarks don't reproduce: Python GIL-bound harnesses asked for 200 QPS quietly delivered 38; busy clients inflated measured latency by up to 58 seconds; a "20% throughput gain" came from temperature 0; and two harnesses tokenized the same dataset differently. Their fix is Inference Perf, a CNCF multiprocess load generator with client-side telemetry vs server metrics, verified at 5,000 QPS, plus a declarative workload catalog and the Prism UI from llm-d.
More from Infra
- SGLang x Datawhale Add New Chapters to Open-Source Inference Engine Course — ying11231 · 2026-09-20
- Baseten's Kiely: Speculative Decoding Is the Fastest-Moving Front in Inference — AI Engineer · 2026-09-20
- CoreWeave inference chief: 80-90% of agentic input is repeat, so cache shapes the whole stack — AI Engineer · 2026-09-20
- Leaving DigitalOcean, one site at a time — carnevalem · 2026-09-20
- The rig built to run Emacs and doomscroll X is now worth more than its owner's car — tetsuoai · 2026-09-19
- Apple M4 sustains 10 instructions per cycle, beating most rivals; M5 speedup explained — lemire · 2026-09-19