FrontierSWE v2 opens 24.1-point gap: Claude Fable 5.1 scores 56.29% vs GPT-5.6's 32.2%
geoffwolfe · x · 2026-09-20
ReasonCoreAI's FrontierSWE v2 benchmark reveals a 24.1-point gap between models that look similar elsewhere: Claude Fable 5.1 scored 56.29% while GPT-5.6 scored 32.2%.
The benchmark uses 34 ultra-long-horizon tasks, each up to 20 hours, with a harness designed to keep agents working instead of submitting early. The authors argue long-horizon capability is a task-and-harness problem, not a one-shot coding score. ReasonCore also maintains FrontierSWE-style tasks across scientific computing, research, and engineering.
More from Models
- XGEN debuts Generative World Simulation: JING model tops WBench Full split — hey_abusiddik · 2026-09-20
- Code-only heuristic policies can beat frontier models on Craftax, evals researcher says — JoshPurtell · 2026-09-20
- Codex usage reset now live for all, big OpenAI release teased for Tuesday — kimmonismus · 2026-09-20
- Jev beats GPT-5.6 Luna on PR review: 1.93x faster at $0.0014 per run — aniketmaurya · 2026-09-20
- Fruit fly connectome chess model beats Jev 4-1 in 10 games, with a playable demo site — maximelabonne · 2026-09-20
- Bonsai 2 27B safety guardrails reportedly cut SWE-bench and Terminal-bench scores by ~20 points — julianharris · 2026-09-20