Self-hosted GLM model reportedly hits 175 tok/s on dual RTX 6000 Pro
A user reports running a suspected GLM model (labeled GLM-5.3-EX3, unverified) on two RTX 6000 Pro GPUs, achieving 175 tok/s generation with 98% draft acceptance rate using speculative decoding.
2026-10-02 ~ 2026-10-02 · 2 related posts
- Local GLM setup reportedly hits 175 tok/s with 98% draft acceptance on dual RTX 6000 Pros — HankYeomans · 2026-10-02
- Local GLM-5.3-EX3 on dual RTX 6k Pros: 98% draft acceptance at 175 tok/s — HankYeomans · 2026-10-02