GLM-5.3 launches: beats Opus 4.8 on coding benchmarks, Terminal Bench jumps from 4.6 to 28.3

ConfessionDiariesPH · reddit · 2026-08-14

GLM-5.3 is out, clearing Opus 4.8 on several coding benchmarks and roughly matching DeepSeek V4 Pro. The standout is Terminal Bench 3.0 jumping from 4.6 on 5.2 to 28.3 on 5.3, a long-horizon benchmark where such a leap is significant. Notably, 5.3 uses the same base model as 5.2; all gains come from post-training without new pretraining. Fable 5 and GPT-5.6 Sol still lead, but the gap has narrowed. The author notes improved tool calling and security, with the model finding real bugs in old open-source software and following proper disclosure.

Related event: Zhipu Releases GLM 5.3 with Enhanced Coding and Cyber Skills(33 posts)→

Original post →

More from Models

Models channel →