GLM-5.3-Flash hits #1 on Artificial Analysis with 294 tok/s
Arindam_1729 · x · 2026-09-01
Zai's open-source model GLM-5.3-Flash (aka Ox Alpha) has claimed the top spot on the Artificial Analysis leaderboard. Running on Nebius, the model achieves a record-breaking inference speed of 294 tokens per second.
Key features:
- Blazing Speed: 294 output tokens/s, ideal for latency-sensitive apps.
- Massive Context: 1M token context window.
- Privacy First: Zero data retention policy.
- Multimodal: Capable of processing multiple modalities.
Related event: GLM-5.3-Flash Tops Inference Speed Charts on Nebius(2 posts)→
More from Infra
- MTP released for Qwen3.8-Flash-Next GGUF, promising big local TPS gains — vini542reddit · 2026-09-01
- Google Cloud Monitoring MCP Connector Released — modelcontextprotocol · 2026-09-01
- Distributed.systems发布可审计的Agent基础设施 — arthurcolle · 2026-09-01
- Does enabling ChatGPT Memory or history reference increase token usage? — ssunki · 2026-09-01
- Engineer fixes ROCm inference crash on MI350X, uncovers 9 bugs in deep dive — AnushElangovan · 2026-09-01
- 40nm Neural-Dynamics Chip Uses Conductance Drift for 2.12ms Iteration Latency — maier_ak · 2026-09-01