Redditor calls benchmark scores meaningless: 1.6T-param DeepSeek only 1 point ahead of 27B Qwen
Nerfariox · reddit · 2026-08-18
A Reddit user argues benchmark scores have become meaningless, pointing to Artificial Analysis' Intelligence Index: DeepSeek V4 Pro 0813 (1600B-A49B, 1.6T params) is only 1 point ahead of Qwen3.8 27B despite 60x the parameters. He also notes people cite the index to claim Qwen3.8 27B beats Gemma 4 31B, while his own day-to-day tests on puzzles and C++/Java refactoring consistently favor Gemma 4 31B.
More from Models
- OpenAI President: Model Capabilities to Increase Significantly Along Roadmap — rohanpaul_ai · 2026-08-18
- GPT Provides Analytical Solution and Constructive Proof — YouJiacheng · 2026-08-18
- Qwen3.8 27B outperforms larger models on local MacBooks — appenz · 2026-08-18
- Open source project claims massive performance boost for Deepseek model via new framework — EAccelerate_42 · 2026-08-18
- Qwen3.8-27B Open Weights Beat Claude Opus on SWE-bench Pro, 262K Context, $0.40/M Input — markjeffrey · 2026-08-18
- Paper reveals massive activations in hybrid linear attention LLMs — rohanpaul_ai · 2026-08-18