Gemma 31B Outperforms DeepSeek V4 and Qwen in Benchmarks

Afinetheorem · x · 2026-09-02

Real-world testing for text classification reveals that Gemma 31B significantly outperforms DeepSeek V4 Flash and Qwen 3.5 122B, while being smaller and faster. The author notes that on custom vision+logic and contextual logic benchmarks, Gemma 31B also beat both competitors. Conversely, Qwen 3.5 27B ranked as the worst among 60+ models tested, highlighting potential issues with benchmark overfitting in some Chinese models.

Original post →

More from Models

Models channel →