Gemma 4 31B vs Qwen3.8 27B: benchmark leaderboards flatly contradict each other
uncle_leon · reddit · 2026-08-26
A Reddit user picking a model for a hobby project found two leaderboards giving opposite verdicts: Artificial Analysis rates Qwen3.8 27B far ahead, while Arena ranks Gemma 4 31B almost 20 places higher and crushing Qwen in many categories. Community sentiment favors Qwen as more tenacious at reasoning, at the cost of overthinking simple tasks. The thread probes how different evaluation methodologies (speed-weighted intelligence vs blind user preference) produce conflicting rankings.
More from Models
- Zhipu GLM-5.3-Flash: Matches Opus 4.8 at 1/40 the Cost, Powered by Domestic Chips — vista8 · 2026-08-27
- TokenSpeed adds Day-0 support for Qwen 3.8 Flash Next architecture — Alibaba_Qwen · 2026-08-27
- Zhipu GLM-5.3 open weights releasing in 22 hours — Yuchenj_UW · 2026-08-27
- AI models show more creativity when talking to each other than in assistant persona — nabeelqu · 2026-08-27
- OpenRouter leaderboard: Real token consumption data outweighs media hype — sujingshen · 2026-08-27
- Qwen 3.8-Next Released with Detailed Technical Report on Architecture — nrehiew_ · 2026-08-27