Is Qwen 3.8 Max Really 56 Points or Just Benchmaxxed?
ideaofsoul · reddit · 2026-08-06
A user questioned the validity of Qwen 3.8 Max's benchmark scores, suggesting they might be artificially inflated. Despite reports of the model scoring 56 points, the poster found it lacking in actual intelligence when using the official app for research tasks.
This sparked a discussion about whether new models are gaming benchmarks. The author argues that the industry relies too heavily on metrics that fail to reflect a model's true capability in everyday, complex scenarios.
More from Models
- Hark Launches Handoff Model, Claims to Outperform ChatGPT 5.4 and Opus 4.8 — peterjliu · 2026-08-06
- Developers Shift to Kimi K3 Amid Anthropic's Limits and Restrictions — casper_hansen_ · 2026-08-06
- DeepSeek Flash v4 Back Online with Significant Speed Improvements — bindureddy · 2026-08-06
- ChatGPT-5.6-luna Cuts Routine Text Processing Costs by 3-5x vs Competitors — zakkohane · 2026-08-06
- Alibaba Releases Qwen3.8-Max with Native Multimodal Reasoning — FellMentKE · 2026-08-06
- Alibaba Releases Qwen3.8-Max: 2.4T Parameters and 1M Context — FellMentKE · 2026-08-06