FrontierSWE v2 Launches: 20-Hour Ultra-Long-Horizon Coding Benchmark, Fable 5.1 Leads by Over 24 Points

The ProximalHQ team (nrehiew and others) released FrontierSWE v2 on September 3, an ultra-long-horizon coding benchmark: models must work autonomously for up to 20 hours on extremely difficult, long-horizon tasks—one representative task is building an OpenGL engine from scratch that can render a flight simulator game. Results show Fable 5.1 is currently the best model, leading by more than 24 percentage points.

Confirmed

Why it matters

This benchmark pushes evaluation duration to the hours or even 20-hour scale, exposing weaknesses in mainstream agent harnesses and models' time awareness, and provides a reference framework for measuring long-horizon autonomy and generalization that is hard to game through targeted training.

2026-09-03 ~ 2026-09-03 · 5 related posts

Full story(2 episodes)→

Primary sources