GLM-5.2 Benchmark Scores Spike Suspiciously

teortaxesTex · x · 2026-07-18

An X user pointed out Y-axis data anomalies in CAISI's benchmark leaderboard. They observed that while GLM-4 previously scored around 800, GLM-5.2's score recently jumped to roughly 1200, with other models also seeing a general score inflation. This suggests the organizers may have secretly altered the benchmark mix. The cited tweet claims GLM-5.2's cybersecurity capabilities are on par with Opus 4.6, and its overall level matches GPT 5.2.

Original post →

More from Models

Models channel →