Controversy Over GLM-5.2 Benchmark Scores

xeophon · x · 2026-07-19

The discussion centers on the **ECI benchmark** and the **impact of SimpleQA**, pointing out that the SimpleQA performance of open-source models like GPT-OSS is a critical component of the evaluation system. The cited text notes that Mechanize's **GBAEval** is a software subset, yet **GLM-5.2 scored 0** on it. The author believes this is likely because it is a **pure text model** unable to verify its own outputs; if this benchmark were excluded, **GLM-5.2's overall score would surge significantly**. It also mentions that some top-tier models do not even possess a GBAEval score.

Related event: Debate Over SimpleQA's Impact on Model Rankings(3 posts)→

Original post →

More from Models

Models channel →