GLM misses only “Computer locked” in nearly perfect ProgramBench cmatrix run
jyangballin · x · 2026-07-23
GLM nearly reproduces the cmatrix task flawlessly
A quoted post says GPT 5.5 was the first model to fully solve a ProgramBench task, getting 505/506 tests on cmatrix, just one short of perfect.
The same quote adds that GLM rebuilt cmatrix almost perfectly and only missed the terminal text “Computer locked.”—a tiny detail that became the whole point of the comparison.
More from Models
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11