SGLang v0.5.15 tunes GLM-5.2 serving to 500+ tok/s on 8×B300
SGLang released v0.5.15 focused on production-grade inference serving, with the headline case serving GLM-5.2 NVFP4 agentic coding workloads on 8×B300 at 500+ tok/s/user (bs=1). The release drew attention because the conversation has shifted from "can the model run" to "how to run frontier open models faster and more stably," treating single-user interactivity and high-concurrency throughput as a joint goal.
Key details
Across the forwards, the consensus is that v0.5.15 is a server-side optimization release tuning production serving for GLM-5.2 NVFP4, hitting 500+ tok/s/user (bs=1) on 8×B300 and also showing 4×B300 results, though no specific figure was given for the latter. The work targets real multi-turn interactive tasks rather than a single benchmark.
Disclosed results
@BanghuaZ relayed that single-user interactivity improved 18% within two weeks of release.
Background and significance
The focus has moved toward system-level optimization for real production inference, making v0.5.15 meaningful beyond the version bump itself.
2026-07-15 ~ 2026-07-15 · 6 related posts
- SGLang Optimizes GLM-5.2: 8xB300 Hits 500+ tok/s for Single User — BanghuaZ · 2026-07-15
- SGLang Pushes GLM-5.2 to 500+ tok/s — ying11231 · 2026-07-15
- SGLang v0.5.15 Released — ying11231 · 2026-07-15
3 near-duplicate retellings: BanghuaZ · mchiang0610 · TheZachMueller