GLM-5-32B Hits Over 200 tok/s Ultra-Fast Inference
Xianbao_QIAN · x · 2026-07-13
The author shared an ultra-fast inference experience using an open-source model, emphasizing that the triumph of open source lies not just in weight control, but in the robust infrastructure built by global experts. In testing, the GLM-5-32B FP8 base model achieved over 200 tok/s per stream during parallel multi-sub-agent decoding. This speed allows agents to operate faster than human thought and proactively deliver results before users even make a request. The test relied on an agent framework developed by the @Zai_org team, with an inference backend using vLLM and LiteLLM. It utilized default parameters with built-in MTP enabled, without any complex special optimizations.
Related event: Open-Source Inference Stack Achieves Over 200 tok/s(2 posts)→
More from coding & agent
- Why vector databases slow AI agents down after constant writes — PrajwalTomar_ · 2026-07-21
- A 13-minute GitHub Copilot video digs into prompt caching — lee_stott · 2026-07-21
- Codex turns out 123 screensavers in one playful batch — intellectronica · 2026-07-21
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- CHAP defines approvals, handoffs, and audit logs for human-agent workflows — DeliveryTechnical199 · 2026-07-21
- The author says Codex reached 20x and is now debugging spec decoding on a hybrid parallel setup — TheZachMueller · 2026-07-21