FULL STORY
Baseten Launches GLM-5.2 Fast API with Major Speedup
Inference platform Baseten launched a new Fast API tier for GLM-5.2, optimizing the service to over twice its initial throughput with peaks reaching 280 tokens/s to meet high-demand real-time needs.
2026-07-24 ~ 2026-07-27 · 2 episodes · 9 posts
Episode 1 · Baseten Launches GLM-5.2 Fast API with 2-3x TPS Boost (2026-07-24, 6 posts)
On July 24, inference platform Baseten officially launched the GLM-5.2 Fast Model API, a new tier designed for highly demanding real-time scenarios. The official statement notes that the Fast tier achieves 2-3x higher TPS compared to the standard GLM-5.2 Model API, maintaining a stable, ultra-fast experience even under high concurrency. The API is now live and available for users to trial.
Confirmed
Baseten confirmed that the GLM-5.2 Fast tier is built for the most rigorous real-time applications, delivering 2 to 3 times higher TPS than the standard GLM-5.2 MAPI. Additionally, developers have already started integrating this API into practical workflows such as `opencode`.
Why it matters
The introduction of the Fast tier represents a performance upgrade at the model serving and inference infrastructure level. For developers and enterprises handling massive real-time interactive tasks, this API can significantly reduce response latency and improve the overall concurrent processing capabilities of their systems.
- Baseten launches GLM-5.2 Fast API with 2-3x higher TPS for real-time use cases — baseten · 2026-07-24
- Baseten launches GLM-5.2 Fast with 2-3x higher TPS for real-time workloads — baseten · 2026-07-24
- Baseten launches a GLM-5.2 Fast API tier with 2-3x higher throughput — baseten · 2026-07-24
- Baseten launches GLM-5.2 Fast API with 2–3x higher TPS — baseten · 2026-07-24
- Baseten Ships GLM-5.2 Fast API for Demanding Real-Time Use Cases — baseten · 2026-07-24
- Baseten says GLM-5.2 Fast delivers 2–3x more throughput for real-time use — baseten · 2026-07-24
Episode 2 · Baseten Pushes GLM-5.2 to 280 tok/s (2026-07-26, 3 posts)
Baseten said it more than doubled GLM-5.2 API performance, reaching 280 tokens per second at peak and about 100 on average. It also released a low-latency Fast version aimed at coding and agent workloads, positioning it as the fastest GLM-5.2 API at the time.
- Baseten doubles GLM-5.2 API performance and launches a lower-latency fast tier — philipkiely · 2026-07-26
- GLM-5.2 API tuning reaches 280 tokens per second before Kimi K3 launch — baseten · 2026-07-27
- Baseten says its GLM-5.2 API hits 280 tokens/s peak and adds vision support — iamrobotbear · 2026-07-27