SGLang Pushes GLM-5.2 to 500+ tok/s
ying11231 · x · 2026-07-15
This deep dive explores how to use SGLang on 8×B300 to drive GLM-5.2 NVFP4 Agentic Workload to 500+ tok/s/user (bs=1).
Key results highlighted in the text include:
- In real-world, multi-turn agentic coding workloads, single-user interactivity improved by 18% to 34% compared to day-0.
- Under high concurrency, peak throughput increased by 6% to 11%.
- The new TopK-V2 kernel is 2.33x faster at 80K ISL, reaching up to 10.17x at 1M ISL.
- Interactivity is effectively maintained up to 1M tokens.
The author notes that some of these gains stem from the model architecture itself: GLM-5.2 applies IndexShare within its DSA layers and introduces a stronger MTP head that reuses IndexShare and KVShare. The remaining improvements are attributed to serving optimizations.
Related event: SGLang v0.5.15 tunes GLM-5.2 serving to 500+ tok/s on 8×B300(6 posts)→
More from Infra
- OpenRouter agents now out-consume humans as AI usage arrives in three waves — AccBalanced · 2026-09-11
- Nvidia Is Now Core to Every Major Robotaxi Stack at Commercial Scale — pdamodaran · 2026-09-11
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- RunningHub open-sources H3Lightning, speeding up MiniMax H3 video generation 12x — 智东西 · 2026-09-11