105 real bugs benchmarked: GPT Astra fixes 48, beats Fable 5.1 at half the cost
PawelHuryn · x · 2026-09-05
PawelHuryn tested GPT-6 Astra on 2 real repos with 105 bugs at max effort. Results: Astra 48/105, Claude Fable 5.1 43/105, GPT-5.6 Sol 42/105, Gemini 3.8 Flash 20/105. Astra was also 2x faster than GPT-5.6 Sol, and 2x faster and cheaper than both GPT-5.6 Sol and Fable 5.1. Full time and cost breakdown in the thread.
Related event: GPT-6 Astra Tops 105-Real-Bug Benchmark with 48 Fixes(10 posts)→
More from coding & agent
- Dev builds entire project with Astra: code, demo video, and post all agent-made — daniel_mac8 · 2026-09-05
- astra-advisor: open-source tool lets GPT-6 Astra delegate to smaller models and cut API costs — daniel_mac8 · 2026-09-05
- astra-advisor hits GitHub: a Codex plugin for dynamic subagent routing and verification — daniel_mac8 · 2026-09-05
- MIT's SwarmWorld paper: agent swarms win when discoveries accumulate, not when agents get smarter — rohanpaul_ai · 2026-09-05
- Memory poisoning on a delay: one bad fact in agent memory seeds every future decision — sierracatalina · 2026-09-05
- GPT-6 Astra tops Terminal Bench 4.0 at half the cost of #2 — charliermarsh · 2026-09-05