Baseten doubles GLM-5.2 API performance and launches a lower-latency fast tier
philipkiely · x · 2026-07-26
Baseten says its GLM-5.2 serving stack now delivers more than 2× the launch-day performance, with peak speeds up to 280 tokens/s initially and a new benchmark showing much higher throughput and lower latency. The company also introduced a separate GLM-5.2-Fast API tuned for coding and agent workloads.
The optimization work included:
- scheduler tuning and bug fixes across the stack
- updated NVFP4 weights
- a revised speculative decoding profile
- latency-focused configuration changes for the fast API
- smaller max batch sizes and a shift away from Attention Data Parallelism toward tensor and expert parallelism
Both APIs run on NVIDIA B200 GPUs. The fast API trades throughput for latency and charges 50% higher input/output token prices. The post also notes that real-world performance depends heavily on traffic patterns and sequence lengths, and that more speculative-decoding improvements are planned.
Related event: Baseten Pushes GLM-5.2 to 280 tok/s(3 posts)→
More from Infra
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11