Qwen 3.8 Max Faces Backlash: High Benchmark Scores Don't Match Real-World Use

DavidOrzc · reddit · 2026-08-10

A Reddit user questioned whether Qwen 3.8 Max is 'benchmaxing'. Despite ranking #2 on the Artificial Intelligence Index and scoring highly on LM Arena, the user found its real-world performance for office tasks (summarizing, drafting) underwhelming compared to models like Fable, GLM-5.2, and Kimi K3. The post seeks community opinions on this discrepancy.

Original post →

More from Models

Models channel →