Real-bug benchmark shows frontier models 'bug-maxing': real 2026 flaws still missed
PawelHuryn · x · 2026-09-05
- PawelHuryn maintains a real-bug benchmark built from genuine flaws frontier models missed in early 2026; answer keys stay private to keep it working.
- Notably, unplanted bugs are excluded: OpenAI models report many real but irrelevant issues, which the author calls "bug-maxing".
- Data (including .) and live stats are public, more effort levels are coming, and the harness will be released so others can test models themselves.
Related event: Bug Hunt Bench: GPT-6 Astra Fixes 48 of 105 Real Bugs to Top Leaderboard(11 posts)→
More from Models
- GPT-6 Astra builds a musical redstone reactor called HELIOS in Minecraft — Angaisb_ · 2026-09-05
- Gemini 3.8 Flash fixes 3.7's "rushing to conclusions" behavior, user reports — DynamicWebPaige · 2026-09-05
- Astra on a Plus plan: two tests torched two 5-hour limits in half an hour — flowersslop · 2026-09-05
- voxelbench 'speechless' over Astra result, points to Fable-5 comparison — legit_api · 2026-09-05
- Author's custom PvZ benchmark is now saturated since GPT 3.5 era — ___Patrice___ · 2026-09-05
- ChatGPT Astra rollout reaches Plus users, first reports from Oceania — OrangeRobots · 2026-09-05