105 real bugs benchmarked: Opus 5.5 fixes 36, DeepSeek fixes 19 for just $0.31

PawelHuryn · x · 2026-10-08

Pawel Huryn ran 105 replanted bugs (missed by frontier models in early 2026) across two real repos to benchmark coding models (Bug Hunt Bench):

Key takeaway: between 4.5 and 5.5, Haiku went from useless to capable of real engineering tasks; DeepSeek and GPT-6 Luna stand out dramatically on cost-efficiency.

Related event: Pawel Huryn's Bug Hunt Benchmark Puts New Coding Models to the Test(5 posts)→

Original post →

More from Models

Models channel →