Bug Hunt Bench: Gemini 3.8 Flash jumps to 20/105 real bugs, Fable 5.1 leads at 43
PawelHuryn · x · 2026-09-03
Pawel Huryn tested Gemini 3.8 Flash (high) on Bug Hunt Bench (blind-graded, one prompt per repo): 2 real repos, 105 hidden planted bugs.
- Fable 5.1 (max): 43/105 for $77.55
- GPT-5.6 (max): 42/105 for $69.61
- Grok 4.6 (xhigh): 27/105 for $16.96
- Opus 5 (max): 21/105 for $51.33
- Gemini 3.8 Flash (high): 20/105 for $9.78
Key takeaways:
- 3.8 Flash is a clear jump from 3.6 Flash's 16/20 — not top of the board but strong value; costs span roughly 200x across runs
- 43/105 bugs remain unfixed by any model — not random bugs but hard issues agents missed in early 2026; "unplanted" fixes are excluded since OpenAI-family models often report and fix real but irrelevant issues
- Practical tip: review your work with a different model family so blind spots don't overlap
Related event: 105 hidden bugs benchmark: Fable 5.1 tops with 43 finds(3 posts)→
More from coding & agent
- Giving your Grok bot an email and a credit card: webhook routines that buy things on Amazon — jeff_weinstein · 2026-09-03
- Dev builds full multi-agent stack on Nostr relay with encrypted A2A comms and unified MCP — RileyRalmuto · 2026-09-03
- Study of 8,351 Claude Code plugins finds 74% of docs commits are runtime instructions — rohanpaul_ai · 2026-09-03
- Codex hooks hide exit codes — stalegreen rewrites verification commands and blocks 26% of stale green claims — SmiLePLSSS · 2026-09-03
- Dev says Codex burns 77% of weekly quota in 28 hours, 4x slower than Claude Code on same tasks — IcyBuy7417 · 2026-09-03
- Nobody on the team could name the version of their production agent — a drift postmortem — Many_Audience7660 · 2026-09-03