Codex and Claude Code harnesses handicap long-horizon tasks; custom Proximus harness lifts scores
nrehiew_ · x · 2026-09-03
Native harnesses not designed for long-horizon work — including Codex and Claude Code — artificially handicap model performance, so the team built Proximus on mini-swe-agent, seeing across-the-board gains in both score and how long models keep working. Long-horizon tasks also act as a proxy for out-of-distribution evaluation, and mass post-training is nearly infeasible when a single task takes 20 hours and 200M tokens.
More from coding & agent
- Linear's bug autofix loop closed 300+ bugs in 30 days via Datadog/Sentry and its coding agent — zeeg · 2026-09-03
- Malicious .git configs make Claude Code, Codex, Cursor run attacker code pre-trust-prompt — Thionne_WTZ · 2026-09-03
- Indie dev marclou goes agent-first: every new SaaS ships with MCP and 60+ agent tools — marclou · 2026-09-03
- Indie dev marclou goes agent-first: 60+ MCP tools make his new SaaS fully AI-operable — steipete · 2026-09-03
- 18.2M tokens of Fable 5.1 built an entire end-to-end workflow app running in the browser — gaganghotra_ · 2026-09-03
- DeepMind's 83-Page Study: Autonomous Research Agents Fabricate 90% of Findings — williamtp · 2026-09-03