FrontierSWE v2: 20-hour autonomous tasks, Fable 5.1 best model by 24+ points
nrehiew_ · x · 2026-09-03
nrehiew released FrontierSWE v2, evaluating models on extremely difficult, extremely long-horizon tasks where models work autonomously for up to 20 hours. Fable 5.1 emerged as the best available model by over 24 percentage points. A follow-up notes thread covers lessons on squeezing performance from models on super-long tasks, including the custom Proximus harness, a time-budget-reminding submit tool, and reward-hacking safeguards.
More from Models
- GLM 5.3 and 5.3 Flash Now Free to Try on Together Chat, No API Setup — oilmutt · 2026-09-03
- LatchBio finds Grok's refusals come from the model itself, while rivals rely on external safety layers — kenbwork · 2026-09-03
- Gemini 3.8 Flash Accused of Bench Overfitting, Regressing vs 3.7 in Third-Party Tests — bindureddy · 2026-09-03
- Gemini 3.8 Flash shows double-digit lift in user satisfaction over 3.7 — tokumin · 2026-09-03
- Muse Spark 1.3 debuts at #3, first model to slot between Claude and GPT — alexandr_wang · 2026-09-03
- Google Researcher Mocks Astra Thinking-Token Outcry as Manufactured Angst — rao2z · 2026-09-03