GLM 5.3 Scores 75.4% on WeirdML, Behind Kimi-K3
teortaxesTex · x · 2026-09-01
On the WeirdML benchmark, Zhipu GLM 5.3 (max) scored 75.4%, up from 70.1% for GLM 5.2.
Comparisons:
- Kimi-K3: 82.6%
- Claude Opus/Fable: 92%
Analysis: While GLM improved, its relative standing might be overestimated. The test shows minor explore-vs-score issues, impacting the final score by less than 1%.
More from Models
- Qwen 3.8 27b oneshots a Super Mario clone in single attempt — zannix · 2026-09-01
- Company offers unlimited Claude, yet most employees stick to default Sonnet 5 — LegitimateLength1916 · 2026-09-01
- Local LLM SVG Generation Test: Qwen 3.8 Leads — pmigdal · 2026-09-01
- Mystery Gemma Model Spotted on Arena Leaderboard — Hot_Example_4456 · 2026-09-01
- antirez shows DeepSeek v4 Flash vision running fast locally on an M5 Max; Metal/CUDA/ROCm support nearly ready — antirez · 2026-09-01
- OPSA boosts AIME24 by 35 points using self-entropy without teacher distillation — heghbalz · 2026-09-01