Local Model Search Benchmarks: Qwen 3.8 27B and Parallel Turbo Win
AccBalanced · x · 2026-09-01
Scott Sanchez released Agentic Search benchmarks for local models to address the lack of great web search options. Tests ran on NVIDIA DGX Spark, evaluating DeepSeek V4 Flash, GLM 5.3 Flash, Qwen 3.8 27B, and Qwen 3.8 Flash Next against 10 search providers (Brave, Exa, Serper, Tavily, etc.). With 2,000 graded answers and real costs incurred, Qwen 3.8 27B + Parallel Turbo emerged as the winner: 50/50 accuracy, 5.7s avg latency, and $1.26 per 1k questions. The runner-up was DeepSeek V4 Flash + Serper. Parallel Turbo led across all models in accuracy (197/200), speed (7.1s), and cost ($1.26).
More from Infra
- DGX Spark owners flag bug: latest CUDA doesn't ship the instant it's released — QuixiAI · 2026-09-03
- VideoDeltaNet open-sources hybrid attention that speeds up MiniMax H3 video generation up to 90x — realmrfakename · 2026-09-03
- Analyst: NVIDIA Could Become Intel Foundry's 'Customer Zero' as a Second Source Beyond TSMC — BenBajarin · 2026-09-03
- Investors bullish on Meta as Muse Spark 1.3 pricing undercuts frontier rivals — Scobleizer · 2026-09-03
- Fervo hits 1,064 MW under contract as Google takes option on 600 MW more — aronchick · 2026-09-03
- METR Publishes Investigation Report on OpenAI / Hugging Face Hacking Incident — stikit · 2026-09-03