Heavy user says AI benchmarks are now 'harmful to trust', not just misleading

aziham · reddit · 2026-09-08

After weeks of heavy use of Qwen 3.8 Max and comparisons with Gemini, GLM-5.3-Flash, and Muse Spark 1.3, a Reddit user argues none come close to Qwen 3.8 Max — except GLM-5.3, which was clearly better on cybersecurity tasks. His broader point: benchmarks diverge so much from real-world experience that they are no longer just misleading but actively harmful to trust, and shouldn't be used to rank models.

Original post →

More from Models

Models channel →