Dev shares concurrency sweep method: TTFT, ITL and tok/s on 4xB200 for local models

TheZachMueller · x · 2026-10-04

TheZachMueller shared a work-in-progress benchmarking approach for local model inference: run a concurrency sweep and report TTFT, ITL, and tok/s to see how a model behaves under load on your specific hardware. He posted example results on 4xB200 GPUs and says more local model reports are coming.

Original post →

More from Infra

Infra channel →