SGLang v0.5.15 tunes GLM-5.2 serving to 500+ tok/s on 8×B300

SGLang released v0.5.15 focused on production-grade inference serving, with the headline case serving GLM-5.2 NVFP4 agentic coding workloads on 8×B300 at 500+ tok/s/user (bs=1). The release drew attention because the conversation has shifted from "can the model run" to "how to run frontier open models faster and more stably," treating single-user interactivity and high-concurrency throughput as a joint goal.

Key details

Across the forwards, the consensus is that v0.5.15 is a server-side optimization release tuning production serving for GLM-5.2 NVFP4, hitting 500+ tok/s/user (bs=1) on 8×B300 and also showing 4×B300 results, though no specific figure was given for the latter. The work targets real multi-turn interactive tasks rather than a single benchmark.

Disclosed results

@BanghuaZ relayed that single-user interactivity improved 18% within two weeks of release.

Background and significance

The focus has moved toward system-level optimization for real production inference, making v0.5.15 meaningful beyond the version bump itself.

2026-07-15 ~ 2026-07-15 · 6 related posts

3 near-duplicate retellings: BanghuaZ · mchiang0610 · TheZachMueller