Baseten doubles GLM-5.2 API performance and launches a lower-latency fast tier
philipkiely · x · 2026-07-26
Baseten says its GLM-5.2 serving stack now delivers more than 2× the launch-day performance, with peak speeds up to 280 tokens/s initially and a new benchmark showing much higher throughput and lower latency. The company also introduced a separate GLM-5.2-Fast API tuned for coding and agent workloads.
The optimization work included:
- scheduler tuning and bug fixes across the stack
- updated NVFP4 weights
- a revised speculative decoding profile
- latency-focused configuration changes for the fast API
- smaller max batch sizes and a shift away from Attention Data Parallelism toward tensor and expert parallelism
Both APIs run on NVIDIA B200 GPUs. The fast API trades throughput for latency and charges 50% higher input/output token prices. The post also notes that real-world performance depends heavily on traffic patterns and sequence lengths, and that more speculative-decoding improvements are planned.
More from Infra
- SK Hynix says HBM is not the final answer to AI’s memory-wall bottleneck — rwang07 · 2026-07-26
- vLLM community backs an open ecosystem as Inferact joins the open letter — vllm_project · 2026-07-26
- Reddit user weighs $3,000 M1 Ultra 128GB against M2 Ultra 64GB for local models — jaqueh · 2026-07-26
- MCP success is not correctness: builders debate who verifies side effects after tool calls — marcin_michalak · 2026-07-26
- Can vLLM and Ray push hot experts onto an RTX 5090 while Sparks hold a 300B MoE? — 2thleZ · 2026-07-26
- AI distillation is colliding with IP law as China’s chip supply hits 41% of demand — Exponential View (Azeem Azhar) · 2026-07-26