GLM-5.2 API tuning reaches 280 tokens per second before Kimi K3 launch
baseten · x · 2026-07-27
The team says it wrapped up the performance work for GLM-5.2 right before the Kimi K3 release.
An attached article claims they built the fastest API for GLM-5.2, with:
- peak throughput of 280 tokens/sec,
- average throughput around 100 tokens/sec.
The post suggests the main story is not the model itself, but the engineering needed to serve it at high speed.
Related event: Baseten Doubles GLM-5.2 API Performance with Low-Latency Fast Version(2 posts)→
More from Infra
- Triton backend pushes Falcon3-10B to 97.5 tok/s on an RTX 5070 — OCV_Researcher · 2026-07-27
- Sparrow switches its Standard mode to Ministral 3 14B for local document extraction — andrejusb · 2026-07-27
- Nvidia supplier Wistron opens $700 million Texas plant for GB300 and Vera Rubin systems — Beth_Kindig · 2026-07-27
- BeeLlama.cpp v0.4.1 adds KV-cache precision tails and new quantization modes — Anbeeld · 2026-07-27
- Running 100 million tokens through GLM 5.2 NVFP4 locally costs about $1 — _akhaliq · 2026-07-27
- Cloudflare’s AI-training block can also stop Googlebot after September 15 — daluoseo · 2026-07-27