DeepSeek V4 Flash beats Qwen3.8-27B in SparkBench evaluation
solyarisoftware · x · 2026-08-19
SparkBench results show DeepSeek V4 Flash scoring 93.01, defeating Qwen3.8-27B's 90.94. DeepSeek wins in code, agents, and tools, while Qwen leads in robustness and calibration. DeepSeek V4 Flash also demonstrates lower median latency and is less verbose.
More from Models
- Gemini Image Generation Silently Fails From Hetzner IPs — Network Origin Was the Culprit — dota2dinall · 2026-08-19
- GLM-5.3 Scores 60 on AI Index, Touted as Strongest Chinese Model — teortaxesTex · 2026-08-19
- DFlash 2 available for Qwen 3.8 27B and Muse Glimmer — rerri · 2026-08-19
- Recent Codex update broke subagents; rolling back to 0.142.0 works — chibop1 · 2026-08-19
- Letting AI labs run their own benchmarks is like students proctoring their own SATs — MattPerault · 2026-08-19
- Grok 4.6 ties Claude Opus 5 on finance diligence bench at ~$0.84/task — karinanguyen · 2026-08-19