GLM-5.2 API tuning reaches 280 tokens per second before Kimi K3 launch

baseten · x · 2026-07-27

The team says it wrapped up the performance work for GLM-5.2 right before the Kimi K3 release.

An attached article claims they built the fastest API for GLM-5.2, with:

The post suggests the main story is not the model itself, but the engineering needed to serve it at high speed.

Related event: Baseten Doubles GLM-5.2 API Performance with Low-Latency Fast Version(2 posts)→

Original post →

More from Infra

Infra channel →