105 real bugs benchmarked: GPT Astra fixes 48, beats Fable 5.1 at half the cost

PawelHuryn · x · 2026-09-05

PawelHuryn tested GPT-6 Astra on 2 real repos with 105 bugs at max effort. Results: Astra 48/105, Claude Fable 5.1 43/105, GPT-5.6 Sol 42/105, Gemini 3.8 Flash 20/105. Astra was also 2x faster than GPT-5.6 Sol, and 2x faster and cheaper than both GPT-5.6 Sol and Fable 5.1. Full time and cost breakdown in the thread.

Related event: GPT-6 Astra Tops 105-Real-Bug Benchmark with 48 Fixes(10 posts)→

Original post →

More from coding & agent

coding & agent channel →