Local Model Search Benchmarks: Qwen 3.8 27B and Parallel Turbo Win

AccBalanced · x · 2026-09-01

Scott Sanchez released Agentic Search benchmarks for local models to address the lack of great web search options. Tests ran on NVIDIA DGX Spark, evaluating DeepSeek V4 Flash, GLM 5.3 Flash, Qwen 3.8 27B, and Qwen 3.8 Flash Next against 10 search providers (Brave, Exa, Serper, Tavily, etc.). With 2,000 graded answers and real costs incurred, Qwen 3.8 27B + Parallel Turbo emerged as the winner: 50/50 accuracy, 5.7s avg latency, and $1.26 per 1k questions. The runner-up was DeepSeek V4 Flash + Serper. Parallel Turbo led across all models in accuracy (197/200), speed (7.1s), and cost ($1.26).

Original post →

More from Infra

Infra channel →