GLM 5.3 Flash Benchmark: Hits 881 tok/s on Dual DGX
teortaxesTex · x · 2026-08-27
A user shared benchmark results for GLM 5.3 Flash, achieving 881 tok/s at C=64 and 232 tok/s at C=1 on an untuned 2x DGX Station setup (TP=2). The model supports long context, realistically serving 4-8 concurrent users on a single machine, and includes vision capabilities. Comments highlight the remaining optimization potential in kernels and prediction heads.
More from Infra
- Concerns Arise Over HuggingFace's Hardware Neutrality After NVIDIA Acquisition — QuixiAI · 2026-08-27
- Nvidia controls who succeeds in AI via allocation, shifting allies to enemies — matt_slotnick · 2026-08-27
- Nvidia posts blowout earnings; analyst calls it historic inflection point — PTrubey · 2026-08-27
- Qwen Training Details: Muon Usage and TP Load Balancing — nrehiew_ · 2026-08-27
- Worried about HuggingFace? A Guide to Legally Back Up AI Models via Torrent — SeyAssociation38 · 2026-08-27
- Two vLLM recipes for Blackwell: NVFP4 KV cache buys 262K context and more streams — SeanHighness · 2026-08-27