105-bug real-repo test: GPT-6 Astra fixes 48/105, beats Fable 5.1 and Gemini 3.8 Flash

PawelHuryn · x · 2026-09-05

PawelHuryn ran a find-and-fix test of 105 bugs across two real repos at max effort: GPT-6 Astra fixed 48/105, ahead of Fable 5.1 (43) and GPT-5.6 Sol (42), while Gemini 3.8 Flash managed only 20/105—and Astra was notably efficient on time and cost. He also complains about access: OpenRouter's effort-level support is unreliable, and Meta's Muse Spark 1.3 API can't be paid for from Europe (4 Visa/MasterCard attempts failed).

Original post →

More from Models

Models channel →