Reddit user calls GLM leaderboard results a lie: "worse than every model I use"
Thin_Pollution8843 · reddit · 2026-09-23
- A Reddit user who uses many open and closed models professionally and as a hobby says he doesn't believe z.ai benchmarks: regardless of harness or prompt engineering, GLM models consistently underperform what their leaderboard numbers suggest.
- He concedes Qwen3.8-27B and DeepSeek v4/4.1 flash may be somewhat benchmaxxed but roughly deliver; GLM-5.2/5.3 ranking on leaderboards he calls a joke, describing them as "lobotomized" — always worse than Qwen3.8-flash, any DeepSeek model, and even a mediocre gpt5.6-Luna.
- He admits he has no explanation but finds it suspicious that only these models fail to live up to their numbers. Unverified personal experience, not an independent evaluation.
More from Models
- Xiaomi open-sources 1.02T-parameter MiMo-V2.6-Pro under MIT with training code — emmanuelvivier · 2026-09-23
- Bug Hunt Bench: GPT-6 Sol (max) matches GPT-5.6 medium but trails Opus 5.5 — PawelHuryn · 2026-09-23
- Claude Opus 5.5 adds time budget: let the model decide how long to work on a task — JeremyNguyenPhD · 2026-09-23
- Buried in the Opus 5.5 system card: METR used an undisclosed 'additional source of information' — burny_tech · 2026-09-23
- Dev switches lineup: pi for headless harness, Claude's comeback, cheaper better Codex — intellectronica · 2026-09-23
- Why is dormant Aider still LLMs' top recommendation for open-source coding harnesses? — marlene_zw · 2026-09-23