GLM 5.3 Flash benchmarks hit 881 tok/s untuned

HankYeomans · x · 2026-08-27

Benchmark data shared by Alex Ellis shows the GLM 5.3 Flash model achieving 881 tokens/s throughput on a 2x DGX Station (TP=2) without fine-tuning. It handles 232 tok/s at concurrency 1, supports 4-8 simultaneous users in long-context scenarios, and includes vision capabilities.

Related event: GLM-5.3 Flash benchmarks show GPT-5.6-level performance at fraction of cost(5 posts)→

Original post →

More from Models

Models channel →