Anthropic's Fable 5.1 Tops Bug Hunt Bench, Beating GPT-5.6

Anthropic's Fable 5.1 topped the Bug Hunt Bench by fixing 43 of 105 injected bugs across two real codebases, beating GPT-5.6 Sol (42) and Grok 4.6 (27), and became the first Pareto-optimal model—faster and cheaper while leading in performance.

2026-09-02 ~ 2026-09-02 · 4 related posts