Local GLM-5.3-EX3 on dual RTX 6k Pros: 98% draft acceptance at 175 tok/s

HankYeomans · x · 2026-10-02

A user reports local deployment numbers: GLM-5.3-EX3 running on two RTX 6k Pro cards hits a 98% draft acceptance rate with 175 tok/s generation and roughly 4100 tok/s prompt throughput—figures the author himself finds almost too good to be true.

Related event: Self-hosted GLM model reportedly hits 175 tok/s on dual RTX 6000 Pro(2 posts)→

Original post →

More from Infra

Infra channel →