Heavy user says AI benchmarks are now 'harmful to trust', not just misleading
aziham · reddit · 2026-09-08
After weeks of heavy use of Qwen 3.8 Max and comparisons with Gemini, GLM-5.3-Flash, and Muse Spark 1.3, a Reddit user argues none come close to Qwen 3.8 Max — except GLM-5.3, which was clearly better on cybersecurity tasks. His broader point: benchmarks diverge so much from real-world experience that they are no longer just misleading but actively harmful to trust, and shouldn't be used to rank models.
More from Models
- AI's Astra designs an original Magic: The Gathering deck and beats a bot on Arena — emollick · 2026-09-08
- Model slips on a science knowledge test, author calls it acceptable — felpix_ · 2026-09-08
- Native video inference beats sampled frames, but OpenRouter providers don't support it yet — spillai · 2026-09-08
- Users game Tibo's token reset: spend 80% fast, stagger reset cycles to dodge surprise resets — tinyfool · 2026-09-08
- Economists in the top 10% of AI use see no step change from GPT 5.6 to 6 — aniketapanjwani · 2026-09-08
- Mathematician: Astra solved two of my unpublished theorems in ~60 hours each, $20k in API — basedjensen · 2026-09-08