105-bug real-repo test: GPT-6 Astra fixes 48/105, beats Fable 5.1 and Gemini 3.8 Flash
PawelHuryn · x · 2026-09-05
PawelHuryn ran a find-and-fix test of 105 bugs across two real repos at max effort: GPT-6 Astra fixed 48/105, ahead of Fable 5.1 (43) and GPT-5.6 Sol (42), while Gemini 3.8 Flash managed only 20/105—and Astra was notably efficient on time and cost. He also complains about access: OpenRouter's effort-level support is unreliable, and Meta's Muse Spark 1.3 API can't be paid for from Europe (4 Visa/MasterCard attempts failed).
More from Models
- Leaker claims GPT-6 Astra beats Claude Fable 5.1 at metaprompting — whoiskatrin · 2026-09-05
- OpenAI docs briefly list "GPT-6 Astra", hinting at next-gen model — EAccelerate_42 · 2026-09-05
- Claude Pro 20x user reports usage down 15% despite heavy multi-hour sessions — Angaisb_ · 2026-09-05
- Mystery model Astra tops VoxelBench with 2600+ Elo, 300+ point lead over GPT-5.5 — legit_api · 2026-09-05
- GPT-6 Astra rolls out to all Plus users, and APPSO benchmarks it against Fable 5.1 — APPSO · 2026-09-05
- MiniMax H3 is underrated at voices and acting, far beyond image animation — LudovicCreator · 2026-09-05