Controversy Over GLM-5.2 Benchmark Scores
xeophon · x · 2026-07-19
The discussion centers on the ECI benchmark and the impact of SimpleQA, pointing out that the SimpleQA performance of open-source models like GPT-OSS is a critical component of the evaluation system.
The cited text notes that Mechanize's GBAEval is a software subset, yet GLM-5.2 scored 0 on it. The author believes this is likely because it is a pure text model unable to verify its own outputs; if this benchmark were excluded, GLM-5.2's overall score would surge significantly. It also mentions that some top-tier models do not even possess a GBAEval score.
Related event: Debate Over SimpleQA's Impact on Model Rankings(3 posts)→
More from Models
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11