Same model, 5x cost gap: harness choice matters more than success rate
JeremyCMorgan · x · 2026-09-25
Berkeley and Arena ran 7 models through Claude Code, Codex CLI and Pi. Success rates moved only a couple of points, but cost moved up to 5x on the same tasks. If you pay per token, the harness is a line item — check sample-size limits before switching.
More from coding & agent
- Developers hunt for the gnarliest 'unmergeable' AI-generated code slop screenshots — pvncher · 2026-09-25
- Wake: open-source desktop app unifies and full-text searches all local coding-agent sessions — tom_doerr · 2026-09-25
- curf: a 250KB open-source C++ browser built for AI agents to drive — jasonkneen · 2026-09-25
- Reddit survey hunts real stories of runaway AI agents: infinite retries and burned credits — masterai01 · 2026-09-25
- Auto-Research Arena: 6,300 runs, agents rediscover MQA/MLA, memory layers and more — qixing_huang · 2026-09-25
- The 2026 AI divide: digitized project context plus agentic tools like Claude Code — Afinetheorem · 2026-09-25