GLM-5-32B Hits Over 200 tok/s Ultra-Fast Inference

Xianbao_QIAN · x · 2026-07-13

The author shared an ultra-fast inference experience using an open-source model, emphasizing that the triumph of open source lies not just in weight control, but in the robust infrastructure built by global experts. In testing, the GLM-5-32B FP8 base model achieved over 200 tok/s per stream during parallel multi-sub-agent decoding. This speed allows agents to operate faster than human thought and proactively deliver results before users even make a request. The test relied on an agent framework developed by the @Zai_org team, with an inference backend using vLLM and LiteLLM. It utilized default parameters with built-in MTP enabled, without any complex special optimizations.

Related event: Open-Source Inference Stack Achieves Over 200 tok/s(2 posts)→

Original post →

More from coding & agent

coding & agent channel →