SGLang Optimizes GLM 4.5 Inference to 535 token/s
BanghuaZ · x · 2026-07-15
The SGLang team continues to optimize the inference performance of the GLM 4.5 model, with the latest version achieving a throughput of 535 token/s. This progress was made possible through close collaboration with NVIDIA and Zhipu AI (Z.ai).
More from Infra
- SkyPilot exits stealth with $20M seed round and an AI compute platform for fragmented clouds — jfiance · 2026-07-22
- Nothing phone mockup turns a film joke into a modular design meme — ZeYanjie · 2026-07-22
- Actual Computer says its inference stack is tuned for Nvidia’s consumer Blackwell lineup — markjeffrey · 2026-07-22
- Ben Bajarin says CPU demand is still being badly underestimated — BenBajarin · 2026-07-22
- An energy model says the U.S. could run short of natural gas starting in 2028 — churchkey · 2026-07-22
- Devin adds e2b sandboxes for remote agent execution — badphilosopher · 2026-07-22