Self-hosted GLM model reportedly hits 175 tok/s on dual RTX 6000 Pro

A user reports running a suspected GLM model (labeled GLM-5.3-EX3, unverified) on two RTX 6000 Pro GPUs, achieving 175 tok/s generation with 98% draft acceptance rate using speculative decoding.

2026-10-02 ~ 2026-10-02 · 2 related posts