Terminal-Bench 4.0: GLM-5.3 Surpasses GPT-5.6 in New Ranking
eyishazyer · x · 2026-08-29
The new Terminal-Bench 4.0 leaderboard delivered surprising results, with scores dropping across the board due to increased difficulty, moving away from the clustering seen in previous versions.
- Claude Opus 5: Tops the list with 51.8% accuracy, barely solving half the tasks, a significant drop from near 90% on the old version.
- GLM-5.3: Ranked third, surprisingly ahead of OpenAI's flagship agent model GPT-5.6 Sol.
This shift indicates that harder benchmarks provide better differentiation of actual model capabilities.
More from Models
- MiniMax H3 Max Criticized for Extreme Censorship on Fal — MrUtterNonsense · 2026-08-29
- User Predicts Qwen-4-27B Will Be a Game Changer — Steus_au · 2026-08-29
- AlayaWorld Tops World Model Rankings as First Open-Weights Leader — Obvious_Set5239 · 2026-08-29
- Minimax H3 limits: Identity leaking and quality degradation — Kooky-Mode3047 · 2026-08-29
- Qwen3.8 27B runs at 50 tok/s with 100k context on 16GB GPU — qaf23 · 2026-08-29
- Uncensored GLM-5.3-Flash weights released, refusal rates drop to 11% — lipeng0820 · 2026-08-29