SGLang Optimizes GLM-5.2: 8xB300 Hits 500+ tok/s for Single User
BanghuaZ · x · 2026-07-15
The SGLang team shared details on their inference optimization for the GLM-5.2 NVFP4 model, achieving a generation speed of 500+ tok/s for a single user (bs=1) on 8xB300 hardware.
- Performance: Within two weeks of release, single-user interactive speed increased by 18% to 34%, and peak throughput under high concurrency rose by 6% to 11%. Thanks to the new TopK-V2 kernel, interactive latency remains nearly flat for ultra-long contexts ranging from 80K to 1M tokens.
- Optimization Strategies: Focused on two main areas. First, reducing system overhead via zero-bubble scheduling, removing synchronization operations, and kernel fusion. Second, rewriting core operators, including TopK-V2 and CuTe DSL-based GEMM.
- Architectural Advantages: GLM-5.2 applies IndexShare technology in its DSA layers and is equipped with a stronger MTP head that reuses IndexShare and KVShare, further boosting inference performance.
Related event: SGLang v0.5.15 tunes GLM-5.2 serving to 500+ tok/s on 8×B300(6 posts)→
More from Infra
- Burning through two ChatGPT resets a day, user coins the "Huang-Altman Law" — yihui_indie · 2026-09-11
- OpenRouter agents now out-consume humans as AI usage arrives in three waves — AccBalanced · 2026-09-11
- Nvidia Is Now Core to Every Major Robotaxi Stack at Commercial Scale — pdamodaran · 2026-09-11
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11