FULL STORY

Baseten Launches GLM-5.2 Fast API with Major Speedup

Inference platform Baseten launched a new Fast API tier for GLM-5.2, optimizing the service to over twice its initial throughput with peaks reaching 280 tokens/s to meet high-demand real-time needs.

2026-07-24 ~ 2026-07-27 · 2 episodes · 9 posts

Episode 1 · Baseten Launches GLM-5.2 Fast API with 2-3x TPS Boost (2026-07-24, 6 posts)

On July 24, inference platform Baseten officially launched the GLM-5.2 Fast Model API, a new tier designed for highly demanding real-time scenarios. The official statement notes that the Fast tier achieves 2-3x higher TPS compared to the standard GLM-5.2 Model API, maintaining a stable, ultra-fast experience even under high concurrency. The API is now live and available for users to trial.

Confirmed

Baseten confirmed that the GLM-5.2 Fast tier is built for the most rigorous real-time applications, delivering 2 to 3 times higher TPS than the standard GLM-5.2 MAPI. Additionally, developers have already started integrating this API into practical workflows such as `opencode`.

Why it matters

The introduction of the Fast tier represents a performance upgrade at the model serving and inference infrastructure level. For developers and enterprises handling massive real-time interactive tasks, this API can significantly reduce response latency and improve the overall concurrent processing capabilities of their systems.

Episode 2 · Baseten Pushes GLM-5.2 to 280 tok/s (2026-07-26, 3 posts)

Baseten said it more than doubled GLM-5.2 API performance, reaching 280 tokens per second at peak and about 100 on average. It also released a low-latency Fast version aimed at coding and agent workloads, positioning it as the fastest GLM-5.2 API at the time.