GLM-5.2 Benchmark Scores Spike Suspiciously
teortaxesTex · x · 2026-07-18
An X user pointed out Y-axis data anomalies in CAISI's benchmark leaderboard. They observed that while GLM-4 previously scored around 800, GLM-5.2's score recently jumped to roughly 1200, with other models also seeing a general score inflation. This suggests the organizers may have secretly altered the benchmark mix. The cited tweet claims GLM-5.2's cybersecurity capabilities are on par with Opus 4.6, and its overall level matches GPT 5.2.
More from Models
- Claude 20x users report sharply tighter limits and faster quota burn — MarcJSchmidt · 2026-07-21
- Cola launches July, the latest model jokingly billed as “second only to Fable” — oran_ge · 2026-07-21
- Kimi K3 looks stronger and about 5× cheaper on a frontend dashboard task — OwariDa · 2026-07-21
- Last Week in AI recap: Anthropic’s $65B round, IPO filing, and Microsoft’s MAI push — Last Week in AI · 2026-07-21
- A user says Claude 4.6 felt worse yesterday and asks whether model quality can drift over time — Rahios · 2026-07-21
- Kimi K3 hits 89.4% peak on software tasks while Fable 5 is slightly steadier — FinanceYF5 · 2026-07-21