105 hidden bugs, 2 repos: DeepSeek V4.1 Flash fixes 24 at $1.80 vs Opus 5's 27 at $51.33
ChartsJournalX · x · 2026-09-11
PawelHuryn ran a hands-on test using two real code repos containing 105 hidden bugs, asking top models to find and fix as many as they could.
- Opus 5 (max): 27 fixed, $51.33
- Grok 4.6 (max): 27 fixed, $16.96
- DeepSeek V4.1 Flash (max): 24 fixed, $1.80
- GPT-5.6 Luna (xhigh): 23 fixed, $2.50
- Opus 5 (high): 21 fixed, $38
Takeaway: DeepSeek V4.1 Flash lands near flagship-level performance on real tasks and is a very strong everyday model, while costing only 3.5% of Opus 5 (max) for the same workload.
More from coding & agent
- MathModelAgent gains traction: auto-solves math modeling and writes a submission-ready paper — jihe520 · 2026-09-11
- alphaXiv open-sources OpenResearch to run parallel research agents with any model — alphaXiv · 2026-09-11
- DeskcommCRM: open-source AI sales CRM with native agents and WhatsApp hits 1k stars — melgarafael · 2026-09-11
- hyperresearch: agent-driven knowledge base that turns web research into a searchable wiki — jordan-gibbs · 2026-09-11
- Forter's 13 lessons from its agent sprint: skip custom RAG, lean on mature enterprise search — bibryam · 2026-09-11
- Two real 'company brains' opened up live: Gorgias' in-house Cortex vs Slite — femke_plantinga · 2026-09-11