SGLang Optimizes GLM 4.5 Inference to 535 token/s

BanghuaZ · x · 2026-07-15

The SGLang team continues to optimize the inference performance of the GLM 4.5 model, with the latest version achieving a throughput of 535 token/s. This progress was made possible through close collaboration with NVIDIA and Zhipu AI (Z.ai).

Original post →

More from Infra

Infra channel →