Codex and Claude Code harnesses handicap long-horizon tasks; custom Proximus harness lifts scores

nrehiew_ · x · 2026-09-03

Native harnesses not designed for long-horizon work — including Codex and Claude Code — artificially handicap model performance, so the team built Proximus on mini-swe-agent, seeing across-the-board gains in both score and how long models keep working. Long-horizon tasks also act as a proxy for out-of-distribution evaluation, and mass post-training is nearly infeasible when a single task takes 20 hours and 200M tokens.

Related event: FrontierSWE v2 Launches: 20-Hour Ultra-Long-Horizon Coding Benchmark, Fable 5.1 Leads by Over 24 Points(5 posts)→

Original post →

More from coding & agent

coding & agent channel →