GLM-5.3 hands-on: 750B post-training lands top-10 on AA board, matches Kimi K3 in coding

卡尔的AI沃茨 · wechat · 2026-08-20

Zhipu launched GLM-5.3, targeting complex coding, long-horizon agent tasks, and defensive cybersecurity. It scores 60 on the AA board (top-10; only six models score 60+), and at 750B is the smallest among them — built on the same base as GLM-5.2 with pure post-training gains, which the author reads as evidence that scaling isn't just about parameters. Pricing stays flat and open-weights are promised for next week.

The author ran six real cases from Goodcase.ai — traffic simulation, text-to-3D games, a 2D rhythm runner, 3D Angry Birds, and a music workstation — benchmarking against DeepSeek v4 pro, ByteDance Seed-Evolving, Kimi K3, and Grok 4.6. Highlights: roughly 2x faster generation than DeepSeek; polished UI details (night headlights, brake lights); and on a real macOS menubar bug in the Ice app, it traced the root cause through private APIs and upstream GitHub issues to a code-level fix, outperforming Grok 4.6's "just switch tools" approach. On safety, Zhipu's team surfaced 2,404 vulnerabilities pre-launch (1,088 medium-to-high severity, some dating back 40 years).

Related event: Zhipu GLM-5.3 Hits 60 on Artificial Analysis Index, Tying Kimi K3(9 posts)→

Original post →

More from coding & agent

coding & agent channel →