SGLang Hits 500+ tok/s on NVIDIA B300
mchiang0610 · x · 2026-07-15
This repost focuses on optimizing the serving stack for SGLang + GLM-5.2 + NVIDIA B300, highlighting how inference and serving stacks can run cutting-edge open-source models faster and more reliably, rather than just discussing the model release itself.
Key details from the quote include:
- On real-world, multi-turn agentic coding workloads, SGLang achieved 500+ tok/s/user on 8xB300, bs=1.
- Within two weeks since day-0, single-user interactivity improved by 18% to 34%.
- Peak throughput in high-concurrency scenarios increased by 6% to 11%.
- The new TopK-V2 kernel is 2.33x faster at 80K ISL and up to 10.17x faster at 1M ISL, while interactivity remains largely stable for long contexts.
The post also mentions that GLM-5.2 incorporates architectural designs like IndexShare and a stronger MTP head, emphasizing that a fully open-source stack is faster, more stable, and more reliable for agentic coding.
Related event: SGLang v0.5.15 tunes GLM-5.2 serving to 500+ tok/s on 8×B300(6 posts)→
More from Infra
- Burning through two ChatGPT resets a day, user coins the "Huang-Altman Law" — yihui_indie · 2026-09-11
- OpenRouter agents now out-consume humans as AI usage arrives in three waves — AccBalanced · 2026-09-11
- Nvidia Is Now Core to Every Major Robotaxi Stack at Commercial Scale — pdamodaran · 2026-09-11
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11