DeepSeek V4 Flash tops its price tier on GBench and looks much smarter in one-shot tasks
teortaxesTex · x · 2026-08-04
DeepSeek V4 Flash is now live on GBench, where it is described as the best model at its price.
- The model performs similarly to the preview checkpoint in an agentic coding harness.
- It is reported to be dramatically stronger in one-shot tasks and fluid intelligence.
- The author says its coding ability sits between GLM 5.1 and GLM 5.2, both much larger models.
- GBench says its game-based coding evaluations are designed to resist benchmark contamination, so results can diverge from standard leaderboards.
More from coding & agent
- Codex can turn every PR into a trading card ranked by business impact — jxnlco · 2026-08-04
- Supabase previews Evals and a new Sign in with ChatGPT button — dshukertjr · 2026-08-04
- Taskmarket lets agents race on tasks and pays only for accepted results — kleffew94 · 2026-08-04
- A Claude.md update codifies keyboard shortcuts, UI copy, and comment rules — charlieholtz · 2026-08-04
- Claude Code 2.1.221 is about to ship — ClaudeCodeLog · 2026-08-04
- Open-sourced a fail-closed authorization layer for high-stakes agents — Riskaval · 2026-08-04