Same Bug Benchmark: GPT-6 Astra Medium Fixes 34/105, Low Scores 27
PawelHuryn · x · 2026-09-10
Benchmark author PawelHuryn added missing data: GPT-6 Astra low previously scored 27/105 and medium 34/105 on the 105-hidden-bugs test, with results published Sep 4-5 and viewable via filters on the leaderboard. Against the rest of the series (Opus 5 max 27 at $51.33, DeepSeek V4.1 Flash 24 at $1.80), GPT-6 Astra medium is currently the top bug-fixer on this benchmark.
More from coding & agent
- Google Simplifies Gemini API Skills, Boosting Correct API Code Generation to 87% — patloeber · 2026-09-10
- Grok Build's Terminal UI Renders Images In-Terminal, Called One of the Sleekest — chongdashu · 2026-09-10
- DeepSeek Agent Swarm Outcodes K3 on Complex Repos, but Burns ¥70 in 2 Hours — teortaxesTex · 2026-09-10
- Dev asks: hundreds of AI agents share one API key — is per-agent identity worth it? — SheepherderFree3931 · 2026-09-10
- Stripe Built an SDK Prototype in 2 Days Instead of 3 Weeks by Delegating to AI Agents Spec-First — dl_weekly · 2026-09-10
- Veteran dev: agentic coding loops 8 times, burns 100k tokens, and wrecks mature codebases — cgouguen · 2026-09-10