Ollama benchmarks: DeepSeek V3 Flash leads, Qwen wins quality but 30x slower
ollama · x · 2026-08-18
Ollama reports that DeepSeek V3 Flash offers the best average performance. For local deployment, the optimized Qwen 3.8 is recommended. Benchmarks on 9 complex tasks show Qwen edges Flash in quality when reasoning is on, but scores worst when off. The tradeoff is significant: Qwen is 30x slower and 4.5x pricier due to heavier computation.
More from Infra
- Grid Bottlenecks Stall AI: Interconnection Queues Surge to 45 Months — PeterDiamandis · 2026-08-18
- Cooling solutions for multi-3090 setup for local inference in 2026 — Sevealin_ · 2026-08-18
- Nvidia secures 35%-40% of global HBM supply for next year — JOBhakdi · 2026-08-18
- pagedMark: Invisible SynthID watermark removal optimized for Apple Silicon — d0ofz · 2026-08-18
- Ex-SpaceX engineers build AI robotic factory for steel parts — Ars Technica AI · 2026-08-18
- Benchmarking Qwen3.8-27B on 4x RTX 3090: Topology Matters — Mr_Moonsilver · 2026-08-18