GLM 5.3 Flash Benchmark: Hits 881 tok/s on Dual DGX

teortaxesTex · x · 2026-08-27

A user shared benchmark results for GLM 5.3 Flash, achieving 881 tok/s at C=64 and 232 tok/s at C=1 on an untuned 2x DGX Station setup (TP=2). The model supports long context, realistically serving 4-8 concurrent users on a single machine, and includes vision capabilities. Comments highlight the remaining optimization potential in kernels and prediction heads.

Original post →

More from Infra

Infra channel →