GLM misses only “Computer locked” in nearly perfect ProgramBench cmatrix run

jyangballin · x · 2026-07-23

GLM nearly reproduces the cmatrix task flawlessly

A quoted post says GPT 5.5 was the first model to fully solve a ProgramBench task, getting 505/506 tests on cmatrix, just one short of perfect.

The same quote adds that GLM rebuilt cmatrix almost perfectly and only missed the terminal text “Computer locked.”—a tiny detail that became the whole point of the comparison.

Original post →

More from Models

Models channel →