Local GLM setup reportedly hits 175 tok/s with 98% draft acceptance on dual RTX 6000 Pros

HankYeomans · x · 2026-10-02

User HankYeomans claims a self-hosted GLM model (written as GLM-5.3-EX3, name unverified) runs at 175 tok/s on two RTX 6000 Pros with a 98% speculative-decoding draft acceptance rate, calling the numbers ridiculous himself. Unverified claim.

Related event: Self-hosted GLM model reportedly hits 175 tok/s on dual RTX 6000 Pro(2 posts)→

Original post →

More from Infra

Infra channel →