Dev shares concurrency sweep method: TTFT, ITL and tok/s on 4xB200 for local models
TheZachMueller · x · 2026-10-04
TheZachMueller shared a work-in-progress benchmarking approach for local model inference: run a concurrency sweep and report TTFT, ITL, and tok/s to see how a model behaves under load on your specific hardware. He posted example results on 4xB200 GPUs and says more local model reports are coming.
More from Infra
- Running a 100B+ Qwen3 model locally on 64GB RAM: good vibes, short 192K context — lxfater · 2026-10-04
- TPU cost per million tokens beats NVIDIA Blackwell, giving Google an edge — cgarciae88 · 2026-10-04
- Strata calibrate nearly tripled decode speed: 256K context on a 16GB GPU — MoonsvnLyn · 2026-10-04
- NVIDIA AIPerf docs go live: a package for performance-testing AI models — TheZachMueller · 2026-10-04
- One prompt freed 39.7 GB: Claude Code + ccmd MCP safely cleans dev caches — julsimon · 2026-10-04
- Community pushes llm-inference-bench as standard for local LLM inference speed measurement — TheZachMueller · 2026-10-04