GLM-5.3 launches: beats Opus 4.8 on coding benchmarks, Terminal Bench jumps from 4.6 to 28.3
ConfessionDiariesPH · reddit · 2026-08-14
GLM-5.3 is out, clearing Opus 4.8 on several coding benchmarks and roughly matching DeepSeek V4 Pro. The standout is Terminal Bench 3.0 jumping from 4.6 on 5.2 to 28.3 on 5.3, a long-horizon benchmark where such a leap is significant. Notably, 5.3 uses the same base model as 5.2; all gains come from post-training without new pretraining. Fable 5 and GPT-5.6 Sol still lead, but the gap has narrowed. The author notes improved tool calling and security, with the model finding real bugs in old open-source software and following proper disclosure.
Related event: Zhipu Releases GLM 5.3 with Enhanced Coding and Cyber Skills(33 posts)→
More from Models
- Qwen3.8-Max open weights now on Nebius Token Factory Day 0 — Arindam_1729 · 2026-08-14
- Qwen 3.8-Max Now on Modal: 2.4T Params, 1M Context, Custom DFlash Speculator — AAAzzam · 2026-08-14
- Qwen3.8 27B vs Qwen3.6 27B vs Gemma 4 31B: Comparing 24GB GPU Options — MaySaki2 · 2026-08-14
- Qwen3.8-27B open-sourced with SGLang Day-0 support, 206 tok/s on RTX 5090 — ying11231 · 2026-08-14
- Qwen3.8-Max launches on DigitalOcean Serverless with 1M context — Alibaba_Qwen · 2026-08-14
- Qwen3.8-27B Becomes 4th Most Liked Model on Hugging Face of All Time — multimodalart · 2026-08-14