GLM-5.3-Flash Runs at 160t/s with 1.4M Context on RTX6000 Pro

AutonomousHangOver · reddit · 2026-08-27

A developer adapted GLM-5.3-Flash to run on sm120 architecture (4 x RTX6000 Pro). Performance metrics show support for 1.4 million context (approx. 5.45 sessions of 262k tokens), with a prefill speed of 3.7k t/s and a token generation speed of 160-230 t/s (MTP enabled).

The implementation uses vLLM inside a Docker container, with the source code available on GitHub.

Original post →

More from Infra

Infra channel →