Baseten says its GLM-5.2 API hits 280 tokens/s peak and adds vision support
iamrobotbear · x · 2026-07-27
Baseten says it built the fastest API for GLM-5.2, with peak throughput of 280 tokens/s and average speeds around 100 tokens/s.
The post also says the API adds vision support, and the author reports that vision testing “works perfectly.”
Related event: Baseten Pushes GLM-5.2 to 280 tok/s(3 posts)→
More from Infra
- Chutes says it trained a 20B model for under $10 an hour using rented GPUs across two continents — markjeffrey · 2026-07-27
- User reports 44 tok/s ingestion and 8 tok/s generation for GLM 5.2 on a $900 rig — naunen · 2026-07-27
- llama.cpp merges support for Minimax M3 with MSA — Time_Reaper · 2026-07-27
- Windows 11 ComfyUI user gets Sage Attention and FlashAttention 2 running on RTX 5090 FE — Left_of_Laniakea · 2026-07-27
- Tech Giants Back Open Source, But Is AI Infrastructure the Real Bottleneck? — myllmnews · 2026-07-27
- Open-source GGUF VRAM calculator estimates context memory before you download — Dry_Wing_ · 2026-07-27