Bug Hunt Bench: Fable 5.1 Low Beats Opus 5 Max at Lower Cost
PawelHuryn · x · 2026-09-02
Pawel Huryn points to Bug Hunt Bench results showing Fable 5.1 low is cheaper and better than Opus 5 (max). Bug Hunt Bench blind-grades frontier coding models on 105 planted bugs in real repos with one prompt per repo, ranking by score, cost (log-scaled, roughly 200x spread) and time (under 7x spread).
More from Models
- Gemini 3.8 Flash reportedly delivers ~Opus 5 performance at much lower cost — inductionheads · 2026-09-02
- Tester Claims Fable 5.1 Excels at Long-Horizon Tasks, Finance and Consulting 'Wiped Out' — felpix_ · 2026-09-02
- Gemini 3.8 Flash spotted on Gemini API pricing page, release looks imminent — rajsharm404 · 2026-09-02
- Google DeepMind publishes Gemini 3.8 Flash model card focused on software engineering and agentic workflows — cedric_chee · 2026-09-02
- 'Pelican on a bike' is saturated — what's the next vibe check for models? — reach_vb · 2026-09-02
- Reddit users call the newly spotted Gemini 3.8 Flash 'insane' — Last_Conclusion_8984 · 2026-09-02