Bug Hunt Bench adds GPT-6 Astra scores: 27/105 low, 34/105 medium

PawelHuryn · x · 2026-09-10

Pawel Huryn's Bug Hunt Bench — a blind-graded leaderboard testing frontier coding models against real planted bugs (105 total, one prompt per repo) — added GPT-6 Astra runs on Sep 4-5: 27/105 at low and 34/105 at medium reasoning tiers, viewable via filters. The board, run by The Product Compass newsletter, also offers score-vs-cost (log axis, 200x spread) and score-vs-time views.

Related event: Bug Hunt Bench Tests 105 Real Bugs, DeepSeek-V4.1-Flash Stands Out on Value(2 posts)→

Original post →

More from Models

Models channel →