Gemma 4 31B vs Qwen3.8 27B: benchmark leaderboards flatly contradict each other

uncle_leon · reddit · 2026-08-26

A Reddit user picking a model for a hobby project found two leaderboards giving opposite verdicts: Artificial Analysis rates Qwen3.8 27B far ahead, while Arena ranks Gemma 4 31B almost 20 places higher and crushing Qwen in many categories. Community sentiment favors Qwen as more tenacious at reasoning, at the cost of overthinking simple tasks. The thread probes how different evaluation methodologies (speed-weighted intelligence vs blind user preference) produce conflicting rankings.

Original post →

More from Models

Models channel →