SGLang Pushes GLM-5.2 to 500+ tok/s

ying11231 · x · 2026-07-15

This deep dive explores how to use SGLang on 8×B300 to drive GLM-5.2 NVFP4 Agentic Workload to 500+ tok/s/user (bs=1).

Key results highlighted in the text include:

The author notes that some of these gains stem from the model architecture itself: GLM-5.2 applies IndexShare within its DSA layers and introduces a stronger MTP head that reuses IndexShare and KVShare. The remaining improvements are attributed to serving optimizations.

Related event: SGLang v0.5.15 tunes GLM-5.2 serving to 500+ tok/s on 8×B300(6 posts)→

Original post →

More from Infra

Infra channel →