105 real bugs benchmarked: Opus 5.5 fixes 36, DeepSeek fixes 19 for just $0.31
PawelHuryn · x · 2026-10-08
Pawel Huryn ran 105 replanted bugs (missed by frontier models in early 2026) across two real repos to benchmark coding models (Bug Hunt Bench):
- Opus 5.5 (xhigh): 36 fixed, $34.98
- Sonnet 5.5 (xhigh): 36, $19.85
- Haiku 5.5 (max): 21.5, $5.83
- DeepSeek V4.1 Flash (high): 19, only $0.31
- GPT-6 Luna (max): 18.3, $0.52 (but slow)
- Gemini 3.8 Flash (high): 18, $11.03
- Haiku 5.5 (high): 16, $4.80
- Haiku 4.5 (default): 1.5, $2.22 (no effort support)
Key takeaway: between 4.5 and 5.5, Haiku went from useless to capable of real engineering tasks; DeepSeek and GPT-6 Luna stand out dramatically on cost-efficiency.
Related event: Pawel Huryn's Bug Hunt Benchmark Puts New Coding Models to the Test(5 posts)→
More from Models
- Dev finds 6.1 Sol surprisingly good at designing native iOS apps — Dimillian · 2026-10-08
- Claude Code lead resets usage limits after community vote: users pick capacity over new features — CurieuxExplorer · 2026-10-08
- Prompting Opus 5.5 with a "grumpy senior engineer reviewer" kept it benchmarking for 2 days — remilouf · 2026-10-08
- Claude Haiku 5.5 beats GPT-6 Luna at matching price, but burns ~3x the tokens — Latent Space · 2026-10-08
- Gemini 3.8 Flash: 3 images and 5 chat messages already eat 8% of the quota — Horror-Airport-7606 · 2026-10-08
- Haiku 5.5 benchmarks 2-3x faster than GPT-6 Luna and performs better — PawelHuryn · 2026-10-08