GLM 5.3 Scores 75.4% on WeirdML, Behind Kimi-K3

teortaxesTex · x · 2026-09-01

On the WeirdML benchmark, Zhipu GLM 5.3 (max) scored 75.4%, up from 70.1% for GLM 5.2.

Comparisons:

Analysis: While GLM improved, its relative standing might be overestimated. The test shows minor explore-vs-score issues, impacting the final score by less than 1%.

Original post →

More from Models

Models channel →