Baseten Pushes GLM-5.2 to 280 tok/s

Baseten said it more than doubled GLM-5.2 API performance, reaching 280 tokens per second at peak and about 100 on average. It also released a low-latency Fast version aimed at coding and agent workloads, positioning it as the fastest GLM-5.2 API at the time.

2026-07-26 ~ 2026-07-27 · 3 related posts

Full story(2 episodes)→