Bug Hunt Bench adds GPT-6 Astra scores: 27/105 low, 34/105 medium
PawelHuryn · x · 2026-09-10
Pawel Huryn's Bug Hunt Bench — a blind-graded leaderboard testing frontier coding models against real planted bugs (105 total, one prompt per repo) — added GPT-6 Astra runs on Sep 4-5: 27/105 at low and 34/105 at medium reasoning tiers, viewable via filters. The board, run by The Product Compass newsletter, also offers score-vs-cost (log axis, 200x spread) and score-vs-time views.
Related event: Bug Hunt Bench Tests 105 Real Bugs, DeepSeek-V4.1-Flash Stands Out on Value(2 posts)→
More from Models
- Polymarket Odds Put Next DeepSeek Pro Release by End of November at 48% — Polymarket · 2026-09-10
- DeepSeek Reportedly Launches V4.1 Flash at a Fraction of a Cent per Million Tokens — Polymarket · 2026-09-10
- New Paper: On-Policy Reverse Distillation Lets Stronger Students Surpass Weak Teachers — algo_diver · 2026-09-10
- willcb declares transformers done, says future belongs to 'freaky-looking sorta-transformers' — willcb · 2026-09-10
- Hy4 preview tested: playable 3D survival game from a single prompt in WorkBuddy — mhdfaran · 2026-09-10
- Gemini 2.5 Pro's search grounding may inflate its benchmark scores vs. rivals — Afinetheorem · 2026-09-10