FrontierSWE v2 opens 24.1-point gap: Claude Fable 5.1 scores 56.29% vs GPT-5.6's 32.2%

geoffwolfe · x · 2026-09-20

ReasonCoreAI's FrontierSWE v2 benchmark reveals a 24.1-point gap between models that look similar elsewhere: Claude Fable 5.1 scored 56.29% while GPT-5.6 scored 32.2%.

The benchmark uses 34 ultra-long-horizon tasks, each up to 20 hours, with a harness designed to keep agents working instead of submitting early. The authors argue long-horizon capability is a task-and-harness problem, not a one-shot coding score. ReasonCore also maintains FrontierSWE-style tasks across scientific computing, research, and engineering.

Original post →

More from Models

Models channel →