Hidden-bug eval across 105 issues: Fable 5.1 finds 43, none fixes all — cost per model compared
PawelHuryn · x · 2026-09-03
PawelHuryn ran an agent bug-hunting eval on the Antigravity CLI using 2 real repos with 105 hidden bugs:
- Fable 5.1 (max): 43/105 for $77.55
- GPT-5.6 (max): 42/105 for $69.61
- Grok 4.6 (xhigh): 27/105 for $16.96
- Opus 5 (max): 21/105 for $51.33
- Gemini 3.8 Flash (high): 20/105 for $9.78
Gemini 3.8 Flash is a clear jump from 3.6 Flash — not the top, but strong value. These aren't random bugs but hard issues agents still miss in early 2026; "unplanted" issues are excluded since models commonly report and fix real but irrelevant issues.
43 bugs remain unfixed by any model. Practical tip: verify your work with a different model family — same-family models share blind spots.
Related event: 105 Hidden Bugs Benchmark: No Model Finds Them All(2 posts)→
More from coding & agent
- Developer asks: is LangChain still worth it vs rolling your own agent harness? — curious_vii · 2026-09-03
- Inference Engineering Is Just a Recipe: vLLM/SGLang, Replicas, Cache-Aware Routing — GabGarrett · 2026-09-03
- Developer accidentally built an entire agent factory with Fable 5.1 — 0xkarasy · 2026-09-03
- Microsoft adds Fabric data agents to Foundry agents via Fabric IQ (preview) — adnan_hashmi · 2026-09-03
- Databricks pitches agent-native data infrastructure, Lakebase Postgres at VLDB 2026 — matei_zaharia · 2026-09-03
- doodlestein ships a comprehensive web app review skill after months of debugging — doodlestein · 2026-09-03