Bug Hunt Bench: GPT-6 Astra has unique blind spots, prompting a new model-routing strategy
PawelHuryn · x · 2026-09-05
PawelHuryn's Bug Hunt Bench — blind-graded runs of frontier coding models against 105 real planted bugs in repos — shares new data and FAQ:
- Key finding: GPT-6 Astra has completely different blind spots from other models, leading to a new routing strategy — implementation with Astra, review with Fable or Sol.
- Methodology notes:
- Unplanted bugs don't count; OpenAI models report many real-but-irrelevant issues ("bug-maxing")
- All bugs were missed by frontier models in early 2026
- Answer keys stay private, which is why the benchmark works
- Costs span 200x (log axis); run times span under 7x
- The harness will be open-sourced so others can test with their own repos; data and JSON are public.
Related event: GPT-6 Astra Tops 105-Real-Bug Benchmark with 48 Fixes(10 posts)→
More from coding & agent
- Dev builds entire project with Astra: code, demo video, and post all agent-made — daniel_mac8 · 2026-09-05
- astra-advisor: open-source tool lets GPT-6 Astra delegate to smaller models and cut API costs — daniel_mac8 · 2026-09-05
- astra-advisor hits GitHub: a Codex plugin for dynamic subagent routing and verification — daniel_mac8 · 2026-09-05
- MIT's SwarmWorld paper: agent swarms win when discoveries accumulate, not when agents get smarter — rohanpaul_ai · 2026-09-05
- Memory poisoning on a delay: one bad fact in agent memory seeds every future decision — sierracatalina · 2026-09-05
- GPT-6 Astra tops Terminal Bench 4.0 at half the cost of #2 — charliermarsh · 2026-09-05