Bug Hunt Bench Tests Frontier Models on 105 Real Bugs
Pawel Huryn introduced Bug Hunt Bench, testing frontier coding models on 105 real-world bugs across two repositories. Claude Sonnet 5.5 topped the leaderboard with a score of 55.5, well ahead of other flagship models.
2026-09-29 ~ 2026-09-29 · 2 related posts
- 105 real bugs benchmarked: Sonnet 5.5 max scores 55.5, beating GPT-6 Astra at 45 — PawelHuryn · 2026-09-29
- Bug Hunt Bench: New Blind-Graded Benchmark Tests Frontier Models on 105 Real Bugs — PawelHuryn · 2026-09-29