Quantized GLM-5.3 fits on 8x RTX Pro 6000, peaks at 486 decode tok/s

TheZachMueller · x · 2026-09-30

NVFP4 quantization of GLM-5.3 runs on a single node of 8x RTX Pro 6000 GPUs: the nvidia/GLM-5.3-NVFP4 build delivers 54 decode tok/s at 251k context (476 peak), while the incoai build with DFlash2 reaches 93 decode tok/s at 123k context (486 peak). Zach Mueller adds that inference saturates NVLink far less than training does.

Original post →

More from Infra

Infra channel →