GLM-5.3-Flash hits #1 on Artificial Analysis with 294 tok/s

Arindam_1729 · x · 2026-09-01

Zai's open-source model GLM-5.3-Flash (aka Ox Alpha) has claimed the top spot on the Artificial Analysis leaderboard. Running on Nebius, the model achieves a record-breaking inference speed of 294 tokens per second.

Key features:

Related event: GLM-5.3-Flash Tops Inference Speed Charts on Nebius(2 posts)→

Original post →

More from Infra

Infra channel →