Baseten Pushes GLM-5.2 to 280 tok/s
Baseten said it more than doubled GLM-5.2 API performance, reaching 280 tokens per second at peak and about 100 on average. It also released a low-latency Fast version aimed at coding and agent workloads, positioning it as the fastest GLM-5.2 API at the time.
2026-07-26 ~ 2026-07-27 · 3 related posts
- Episode 1: Baseten Launches GLM-5.2 Fast API with 2-3x TPS Boost(2026-07-24, 6 posts)
- Episode 2: Baseten Pushes GLM-5.2 to 280 tok/s(2026-07-26, 3 posts)
- Baseten doubles GLM-5.2 API performance and launches a lower-latency fast tier — philipkiely · 2026-07-26
- GLM-5.2 API tuning reaches 280 tokens per second before Kimi K3 launch — baseten · 2026-07-27
- Baseten says its GLM-5.2 API hits 280 tokens/s peak and adds vision support — iamrobotbear · 2026-07-27