105 Hidden Bugs Benchmark: DeepSeek V4.1 Flash Scores 24 at $1.80 vs Opus 5's $51
sanderjson · x · 2026-09-10
PawelHuryn ran a real-world benchmark across 2 repos with 105 hidden bugs, asking models to find and fix them:
- Opus 5 (max): 27 fixed, costing $51.33
- Grok 4.6 (max): 27, $16.96
- DeepSeek V4.1 Flash (max): 24, only $1.80
- GPT-5.6 Luna (xhigh): 23, $2.50
- Opus 5 (high): 21, $38+
Takeaway: DeepSeek V4.1 Flash is remarkably strong for everyday coding tasks — nearly matching frontier models at roughly 1/28th the cost of Opus 5.
More from coding & agent
- Google Simplifies Gemini API Skills, Boosting Correct API Code Generation to 87% — patloeber · 2026-09-10
- Grok Build's Terminal UI Renders Images In-Terminal, Called One of the Sleekest — chongdashu · 2026-09-10
- DeepSeek Agent Swarm Outcodes K3 on Complex Repos, but Burns ¥70 in 2 Hours — teortaxesTex · 2026-09-10
- Dev asks: hundreds of AI agents share one API key — is per-agent identity worth it? — SheepherderFree3931 · 2026-09-10
- Stripe Built an SDK Prototype in 2 Days Instead of 3 Weeks by Delegating to AI Agents Spec-First — dl_weekly · 2026-09-10
- Veteran dev: agentic coding loops 8 times, burns 100k tokens, and wrecks mature codebases — cgouguen · 2026-09-10