GLM misses only “Computer locked” in nearly perfect ProgramBench cmatrix run
jyangballin · x · 2026-07-23
GLM nearly reproduces the cmatrix task flawlessly
A quoted post says GPT 5.5 was the first model to fully solve a ProgramBench task, getting 505/506 tests on cmatrix, just one short of perfect.
The same quote adds that GLM rebuilt cmatrix almost perfectly and only missed the terminal text “Computer locked.”—a tiny detail that became the whole point of the comparison.
More from Models
- Bug Hunt Bench ranks frontier coding models on 105 planted real-repo bugs — PawelHuryn · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- 105 hidden bugs, 2 repos: DeepSeek V4.1 Flash fixes 24 at $1.80 vs Opus 5's 27 at $51.33 — ChartsJournalX · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11