Gemma 31B Outperforms DeepSeek V4 and Qwen in Benchmarks
Afinetheorem · x · 2026-09-02
Real-world testing for text classification reveals that Gemma 31B significantly outperforms DeepSeek V4 Flash and Qwen 3.5 122B, while being smaller and faster. The author notes that on custom vision+logic and contextual logic benchmarks, Gemma 31B also beat both competitors. Conversely, Qwen 3.5 27B ranked as the worst among 60+ models tested, highlighting potential issues with benchmark overfitting in some Chinese models.
More from Models
- Perplexity adds Claude Fable 5.1, cutting costs by 37% — perplexity_ai · 2026-09-02
- Anthropic Fable 5.1 System Prompt Leaked, Spanning 270k+ Characters — Scobleizer · 2026-09-02
- Fable 5.1 now integrates Anthropic's statistical text watermarking — RaGE_Syria · 2026-09-02
- Multi-agent evals lack model comparisons, need more details — scaling01 · 2026-09-02
- Fable 5.1 Beats GPT-5.6 on Benchmark at Lower Cost — haider1 · 2026-09-02
- Claude Fable 5.1 adds heavy instructions,疑似过度对齐 — teodorio · 2026-09-02