Local GLM-5.3-EX3 on dual RTX 6k Pros: 98% draft acceptance at 175 tok/s
HankYeomans · x · 2026-10-02
A user reports local deployment numbers: GLM-5.3-EX3 running on two RTX 6k Pro cards hits a 98% draft acceptance rate with 175 tok/s generation and roughly 4100 tok/s prompt throughput—figures the author himself finds almost too good to be true.
Related event: Self-hosted GLM model reportedly hits 175 tok/s on dual RTX 6000 Pro(2 posts)→
More from Infra
- MachGen pushes MiniMax H3 past its 15s cap with 30-second continuous video — MiniMax_AI · 2026-10-02
- Redditor builds fully local LLM-powered radio site on two DGX Sparks and a 5090 — jwhh91 · 2026-10-02
- VC quip: many neoclouds are closer to 95% than five nines of reliability — saranormous · 2026-10-02
- Report: lenders demand up to 25% collateral from Nvidia as GPU-backed loans wobble — GaryMarcus · 2026-10-02
- Microsoft Backs Snowflake-Led Effort to Standardize Business Metrics for AI — xiaosun86 · 2026-10-02
- GLM-5.3-Flash NVFP4 benchmarks show no per-user speedup beyond 8 concurrent requests — TheZachMueller · 2026-10-02