GLM 5.2 ranks third on ProgramBench, but the thread warns averages can mislead

jyangballin · x · 2026-07-23

GLM 5.2 has landed 3rd place on the ProgramBench leaderboard as the first open-weight model evaluated in the latest update. The discussion around the result also highlights a broader evaluation caveat: reporting only average pass rates can hide cases where a model is still failing a meaningful fraction of instances.

The thread argues that looking at the number of almost-resolved and fully resolved tasks is more informative than a single aggregate score, because a few stubborn failures can reveal major shortcomings even when the average looks strong.

Related event: GLM-5.2 Ranks Third on ProgramBench Leaderboard(3 posts)→

Original post →

More from Models

Models channel →