Baseten says its GLM-5.2 API hits 280 tokens/s peak and adds vision support

iamrobotbear · x · 2026-07-27

Baseten says it built the fastest API for GLM-5.2, with peak throughput of 280 tokens/s and average speeds around 100 tokens/s.

The post also says the API adds vision support, and the author reports that vision testing “works perfectly.”

Related event: Baseten Pushes GLM-5.2 to 280 tok/s(3 posts)→

Original post →

More from Infra

Infra channel →