SGLang Hits 500+ tok/s on NVIDIA B300

mchiang0610 · x · 2026-07-15

This repost focuses on optimizing the serving stack for SGLang + GLM-5.2 + NVIDIA B300, highlighting how inference and serving stacks can run cutting-edge open-source models faster and more reliably, rather than just discussing the model release itself.

Key details from the quote include:

The post also mentions that GLM-5.2 incorporates architectural designs like IndexShare and a stronger MTP head, emphasizing that a fully open-source stack is faster, more stable, and more reliable for agentic coding.

Related event: SGLang v0.5.15 tunes GLM-5.2 serving to 500+ tok/s on 8×B300(6 posts)→

Original post →

More from Infra

Infra channel →