Redditor calls benchmark scores meaningless: 1.6T-param DeepSeek only 1 point ahead of 27B Qwen

Nerfariox · reddit · 2026-08-18

A Reddit user argues benchmark scores have become meaningless, pointing to Artificial Analysis' Intelligence Index: DeepSeek V4 Pro 0813 (1600B-A49B, 1.6T params) is only 1 point ahead of Qwen3.8 27B despite 60x the parameters. He also notes people cite the index to claim Qwen3.8 27B beats Gemma 4 31B, while his own day-to-day tests on puzzles and C++/Java refactoring consistently favor Gemma 4 31B.

Original post →

More from Models

Models channel →