GLM-5.3 Open Weights: SGLang Reports Over 500 tok/s on Agentic Workloads
BanghuaZ · x · 2026-08-28
Zai has released GLM-5.3 as an open-weight model, positioning it as their most capable model for agentic coding and cyber defense. SGLang has announced day-0 support with performance benchmarks on NVIDIA Blackwell/Hopper and AMD MI300 series hardware.
Key Performance Metrics:
- On real-world multi-turn agentic workloads (BS=1, TP8):
- 537.6 tok/s/user on NVFP4
- 413 tok/s/user on FP8
The model inherits optimizations from GLM-5.2 and is production-ready on NVIDIA Blackwell, Hopper, and AMD MI300X/325X/355X.
More from Infra
- Microsoft reportedly reassures employees on AI data center energy and emissions impact — Polymarket · 2026-08-29
- Google Cloud SQL Introduces Performance Assessments Preview — rseroter · 2026-08-29
- Decathlon Switches to Chronos-2 for Large-Scale Demand Forecasting on AWS — AWS ML Blog · 2026-08-29
- Salesforce Achieves Multi-AZ HA with SageMaker Inference Components — AWS ML Blog · 2026-08-29
- Cloud Rethink: Enterprises pull back from blind migration to public cloud — DavidLinthicum · 2026-08-29
- OpenAI Python SDK migrates to HTTPX2, drops httpx dependency — mitsuhiko · 2026-08-29