GLM 5.2 ranks third on ProgramBench, but the thread warns averages can mislead
jyangballin · x · 2026-07-23
GLM 5.2 has landed 3rd place on the ProgramBench leaderboard as the first open-weight model evaluated in the latest update. The discussion around the result also highlights a broader evaluation caveat: reporting only average pass rates can hide cases where a model is still failing a meaningful fraction of instances.
The thread argues that looking at the number of almost-resolved and fully resolved tasks is more informative than a single aggregate score, because a few stubborn failures can reveal major shortcomings even when the average looks strong.
Related event: GLM-5.2 Ranks Third on ProgramBench Leaderboard(3 posts)→
More from Models
- Sentdex questions SemiAnalysis’ chart comparing Kimi K3 with Nemotron 3 Ultra — Daniel_Farinax · 2026-07-23
- Ryan Greenblatt says several failure modes could explain OpenAI’s hacking incident — RyanGreenblatt · 2026-07-23
- A VRAM-by-VRAM model roundup maps the best open weights from 4GB to 384GB — victormustar · 2026-07-23
- BTL-3 launches as a 27B open-weight agentic coding model with 95.12% HumanEval — soteko · 2026-07-23
- Testing LLMs with a Churchill Insult Prompt: GPT Outperforms Kimi and Gemini — emollick · 2026-07-23
- NVIDIA Recaps Kaggle Nemotron Challenge: 5,000+ Explore Open Model Reasoning — NVIDIAAI · 2026-07-23