Community pushes llm-inference-bench as standard for local LLM inference speed measurement
TheZachMueller · x · 2026-10-04
Community members are calling for a unified standard to measure local LLM prefill/decode throughput, arguing that self-reported numbers aren't comparable and should be questioned — including their own.
- The proposed de facto tool is Local Inference Lab's open-source llm-inference-bench, which benchmarks a full matrix of concurrency levels (1/2/4/8...) and context lengths (0K–128K) with a Rich TUI dashboard
- It covers three benchmark layers — prefill, sustained decode, and optional Burst/E2E decode — supports SGLang and vLLM with auto-detection, and works with any OpenAI-compatible API (OpenRouter, Together AI, etc.)
- It's easily reproducible across hardware setups, making cross-machine comparisons feasible
Zach Mueller says he'll propose a standardization effort "by Tuesday" to align terminology with industry practice and capture the full range of values needed to understand local inference speeds.
More from Infra
- Running a 100B+ Qwen3 model locally on 64GB RAM: good vibes, short 192K context — lxfater · 2026-10-04
- TPU cost per million tokens beats NVIDIA Blackwell, giving Google an edge — cgarciae88 · 2026-10-04
- Strata calibrate nearly tripled decode speed: 256K context on a 16GB GPU — MoonsvnLyn · 2026-10-04
- Dev shares concurrency sweep method: TTFT, ITL and tok/s on 4xB200 for local models — TheZachMueller · 2026-10-04
- NVIDIA AIPerf docs go live: a package for performance-testing AI models — TheZachMueller · 2026-10-04
- One prompt freed 39.7 GB: Claude Code + ccmd MCP safely cleans dev caches — julsimon · 2026-10-04