A year of LLM inference tuning: vLLM vs SGLang, 2000-rpm load tests, and the real bottleneck

abhijithneil · x · 2026-09-03

The author recounts a year spent speeding up LLM inference: ablations on single-GPU boxes and multi-node fleets, benchmarking vLLM against SGLang, and load tests that ramped to 2000 rpm before falling over at minute six. His biggest lesson: the real bottleneck was lacking a mental model of the machine, making every flag feel like a coin flip. The thread will expand into concrete inference optimization methodology.

Related event: A Year of LLM Inference Tuning: vLLM vs SGLang(2 posts)→

Original post →

More from Infra

Infra channel →