Local GLM setup reportedly hits 175 tok/s with 98% draft acceptance on dual RTX 6000 Pros
HankYeomans · x · 2026-10-02
User HankYeomans claims a self-hosted GLM model (written as GLM-5.3-EX3, name unverified) runs at 175 tok/s on two RTX 6000 Pros with a 98% speculative-decoding draft acceptance rate, calling the numbers ridiculous himself. Unverified claim.
Related event: Self-hosted GLM model reportedly hits 175 tok/s on dual RTX 6000 Pro(2 posts)→
More from Infra
- Modal launches BYOC 2.0 to keep cloud coding agents' sensitive data inside your account — charles_irl · 2026-10-02
- GPT-6 Astra Ultrafast launches on NVIDIA Blackwell with up to 8x faster tokens — nvidia · 2026-10-02
- Google's Project Suncatcher prototype satellite is now in orbit — Gaiden206 · 2026-10-02
- OpenAI's GPT-6 Astra Ultrafast Delivers Up to 8x Faster Token Generation on NVIDIA Blackwell — NVIDIA Blog · 2026-10-02
- Ben Lorica: The AI data problem moved downstream, from finding data to making it usable — bigdata · 2026-10-02
- California man arrested for allegedly smuggling over $300M in AI servers to China — Polymarket · 2026-10-02