Quantized GLM-5.3 fits on 8x RTX Pro 6000, peaks at 486 decode tok/s
TheZachMueller · x · 2026-09-30
NVFP4 quantization of GLM-5.3 runs on a single node of 8x RTX Pro 6000 GPUs: the nvidia/GLM-5.3-NVFP4 build delivers 54 decode tok/s at 251k context (476 peak), while the incoai build with DFlash2 reaches 93 decode tok/s at 123k context (486 peak). Zach Mueller adds that inference saturates NVLink far less than training does.
More from Infra
- Vultr signs $1.2 billion deal with HPE, securing first AMD Helios order — wkmyrhang · 2026-09-30
- Cloudflare AI Gateway adds User Insights to spot teams overusing overly capable models — michellechen · 2026-09-30
- Cloudflare Birthday Week: 6x faster containers, AI Gateway auto-routing, Workers error monitoring — ritakozlov · 2026-09-30
- Tianqi Chen's team open-sources a book on compiler-driven agentic GPU kernel optimization for MLSys — sh_reya · 2026-09-30
- computesdk cuts cold-start latency 6x: median 4.05s to 648ms via co-located containers and VM restore — dinasaur_404 · 2026-09-30
- Golem Builds a Log Incident Pipeline With Jev Classifier and Durable Agents — blaizedsouza · 2026-09-30