GLM misses only “Computer locked” in nearly perfect ProgramBench cmatrix run
jyangballin · x · 2026-07-23
GLM nearly reproduces the cmatrix task flawlessly
A quoted post says GPT 5.5 was the first model to fully solve a ProgramBench task, getting 505/506 tests on cmatrix, just one short of perfect.
The same quote adds that GLM rebuilt cmatrix almost perfectly and only missed the terminal text “Computer locked.”—a tiny detail that became the whole point of the comparison.
More from Models
- Hugging Face sees fdtn-ai/antares-1b trend as a security-focused text model — fdtn-ai · 2026-07-23
- Qwen 3.6 benchmarks compare token throughput across 170HX, RTX 5090, and RTX PRO 6000 — simplefunction · 2026-07-23
- One user says GPT-5.6 Sol Medium is now their most-used model — MatthewBerman · 2026-07-23
- GLM 5.2 reaches No. 3 on the official ProgramBench leaderboard — klieret · 2026-07-23
- Gradium boosts speech-to-text accuracy on rare names with keyword prompting — mattturck · 2026-07-23
- OpenAI's GPT-5.6 Escapes Sandbox to Hack Hugging Face During Benchmark Test — Ars Technica AI · 2026-07-23