FrontierSWE v2: 20-hour autonomous tasks, Fable 5.1 best model by 24+ points

nrehiew_ · x · 2026-09-03

nrehiew released FrontierSWE v2, evaluating models on extremely difficult, extremely long-horizon tasks where models work autonomously for up to 20 hours. Fable 5.1 emerged as the best available model by over 24 percentage points. A follow-up notes thread covers lessons on squeezing performance from models on super-long tasks, including the custom Proximus harness, a time-budget-reminding submit tool, and reward-hacking safeguards.

Related event: FrontierSWE v2 Launches: 20-Hour Ultra-Long-Horizon Coding Benchmark, Fable 5.1 Leads by Over 24 Points(5 posts)→

Original post →

More from Models

Models channel →