Baseten doubles GLM-5.2 API performance and launches a lower-latency fast tier

philipkiely · x · 2026-07-26

Baseten says its GLM-5.2 serving stack now delivers more than 2× the launch-day performance, with peak speeds up to 280 tokens/s initially and a new benchmark showing much higher throughput and lower latency. The company also introduced a separate GLM-5.2-Fast API tuned for coding and agent workloads.

The optimization work included:

Both APIs run on NVIDIA B200 GPUs. The fast API trades throughput for latency and charges 50% higher input/output token prices. The post also notes that real-world performance depends heavily on traffic patterns and sequence lengths, and that more speculative-decoding improvements are planned.

Original post →

More from Infra

Infra channel →