Same Bug Benchmark: GPT-6 Astra Medium Fixes 34/105, Low Scores 27

PawelHuryn · x · 2026-09-10

Benchmark author PawelHuryn added missing data: GPT-6 Astra low previously scored 27/105 and medium 34/105 on the 105-hidden-bugs test, with results published Sep 4-5 and viewable via filters on the leaderboard. Against the rest of the series (Opus 5 max 27 at $51.33, DeepSeek V4.1 Flash 24 at $1.80), GPT-6 Astra medium is currently the top bug-fixer on this benchmark.

Original post →

More from coding & agent

coding & agent channel →