Baseten Claims Fastest Inference for GLM-5.3-Flash at 122+ TPS

baseten · x · 2026-08-28

Baseten claims to offer the fastest inference for GLM-5.3-Flash on Artificial Analysis, OpenRouter, and Hugging Face, achieving over 122 TPS. The service is US-only with ZDR enabled by default, with further optimizations planned.

Original post →

More from Infra

Infra channel →