Qwen 3.8 Max Faces Backlash: High Benchmark Scores Don't Match Real-World Use
DavidOrzc · reddit · 2026-08-10
A Reddit user questioned whether Qwen 3.8 Max is 'benchmaxing'. Despite ranking #2 on the Artificial Intelligence Index and scoring highly on LM Arena, the user found its real-world performance for office tasks (summarizing, drafting) underwhelming compared to models like Fable, GLM-5.2, and Kimi K3. The post seeks community opinions on this discrepancy.
More from Models
- Alibaba to Release Open-Weight Qwen3.8-Max, Pressuring US Rivals — emmanuelvivier · 2026-08-10
- Nous Research Releases Open-Source Code Model for Local Execution and Agents — emmanuelvivier · 2026-08-10
- Claude Extended Thinking Bug Silently Burns 674k Tokens — forchat4 · 2026-08-10
- Google Reportedly Cancels Gemini 3.5 Pro Development Silently — Left-Hotel904 · 2026-08-10
- Grok Imagine 2.0 Ships, Jumping to #2 Globally in Image Generation — eyishazyer · 2026-08-10
- Anthropic API Strict Mode with $ref Emits Contradictory Outputs — Nearby_Yam286 · 2026-08-10