Bug Hunt Bench details: Data and receipts for Fable 5.1 vs competitors

PawelHuryn · x · 2026-09-02

Bug Hunt Bench released detailed data and receipts verifying the performance of frontier coding models on real repositories with planted bugs.

Benchmarks Features:

The leaderboard ranks models by the number of planted bugs fixed (out of 105) and provides visualizations comparing cost (logarithmic axis) and time (linear axis), illustrating the Pareto frontier of performance, speed, and cost across different models.

Related event: Anthropic's Fable 5.1 Tops Bug Hunt Bench, Beating GPT-5.6(4 posts)→

Original post →

More from coding & agent

coding & agent channel →