Ollama benchmarks: DeepSeek V3 Flash leads, Qwen wins quality but 30x slower

ollama · x · 2026-08-18

Ollama reports that DeepSeek V3 Flash offers the best average performance. For local deployment, the optimized Qwen 3.8 is recommended. Benchmarks on 9 complex tasks show Qwen edges Flash in quality when reasoning is on, but scores worst when off. The tradeoff is significant: Qwen is 30x slower and 4.5x pricier due to heavier computation.

Original post →

More from Infra

Infra channel →