Evaluating 7 models across Claude Code, Codex, and Pi: harness choice drives cost, not success rate
CShorten30 · x · 2026-09-17
Matei Zaharia shares Melissa Pan's research evaluating 7 models across three agent harnesses (Claude Code, Codex, Pi). Three surprising findings: harness choice has little effect on task success rate but can significantly affect cost; a simple harness can be competitive; and the native harness isn't always the best. With millions using coding agents, the impact of harness choice had remained unclear.
More from coding & agent
- Diorama gives OpenAI Codex coding agents a visual office you can watch work in real time — davidfromkansas · 2026-09-17
- Code-first, UI on top: building bespoke brand design tools with AI — floguo · 2026-09-17
- Study of 7 models across Claude Code, Codex, Pi: harness barely affects success but swings cost — DavideCrapis · 2026-09-17
- AI trading bot built with Jev is down 85%, owner shrugs it off — generativist · 2026-09-17
- Redditor's 3-Day SoL-Pi Test: Memory Objects Save ~12k Tokens Per Tool Run — Garblyx · 2026-09-17
- Reviewing AI code through Steve Jobs' lens: unseen internals deserve beauty too — sergeykarayev · 2026-09-17