Study of 7 models across Claude Code, Codex, Pi: harness barely affects success but swings cost
DavideCrapis · x · 2026-09-17
Melissa Pan's team evaluated 7 models across the Claude Code, Codex, and Pi harnesses and found three surprising results: (1) harness choice has little effect on task success rate but can significantly affect cost; (2) a simple harness can be competitive; (3) the model's native harness isn't always the best. The study addresses a long-standing gap — millions use coding agents, yet the impact of harness choice remained unclear.
More from coding & agent
- Computer-use agents breeze through GitHub token provisioning, sparking calls to open up CLI access — aronchick · 2026-09-17
- The Superdark Factory: by 2029, code review becomes a formality as agents write 35T tokens a month — bratton · 2026-09-17
- He runs his entire company on OpenClaw: an early look at agent-native operations — heyneighbor · 2026-09-17
- Qwen Code Desktop v0.24.0 ships Linux bwrap sandbox, DingTalk support, cross-session messaging — github-actions[bot] · 2026-09-17
- Nvidia's Agora: 13 LLM Agents Run 12 Days Unsupervised via Git-Based Shared Memory — nvidia · 2026-09-17
- Open-source AgenC ships its biggest release: 658 commits in 28 days, plus a desktop app — tetsuoai · 2026-09-17