Hidden-bug hunt: DeepSeek V4.1 Flash fixes 24/105 bugs at $1.80 vs Opus 5's $51.33
soveLight · x · 2026-09-10
soveLight benchmarked models on real tasks: 2 repos with 105 hidden bugs, find-and-fix scoring at max settings:
- Opus 5: 27 bugs, $51.33
- Grok 4.6: 27 bugs, $16.96
- DeepSeek V4.1 Flash: 24 bugs, just $1.80
- GPT-5.6 Luna (xhigh): 23 bugs, $2.50
- Opus 5 (high): 21 bugs, $38
Takeaway: DeepSeek V4.1 Flash is strong for everyday work at roughly 1/28th of Opus's cost. Low/medium-tier tests are queued next.
More from Models
- DeepSeek 4.1 flash reportedly uses large ngram embeddings, echoing Qwen4 architecture — ccerrato147 · 2026-09-11
- ValsAI launches RSI Index, first third-party benchmark measuring how close AI is to self-improvement — JenniferHli · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11
- Business Insider Asked ChatGPT, Gemini, Claude and Grok How AI Could End Humanity — coinfanking · 2026-09-11
- Claims resurface that Moonshot's Kimi distilled from Claude raw CoTs — xuanalogue · 2026-09-11
- User switches back to GPT-5.6 Sol: barely uses quota and feels faster — CtrlAltDwayne · 2026-09-11