Evaluating 7 models across Claude Code, Codex, and Pi: harness choice drives cost, not success rate

CShorten30 · x · 2026-09-17

Matei Zaharia shares Melissa Pan's research evaluating 7 models across three agent harnesses (Claude Code, Codex, Pi). Three surprising findings: harness choice has little effect on task success rate but can significantly affect cost; a simple harness can be competitive; and the native harness isn't always the best. With millions using coding agents, the impact of harness choice had remained unclear.

Related event: Testing 7 Models Across 3 Agent Harnesses: Harness Choice Drives Cost, Not Success Rate(5 posts)→

Original post →

More from coding & agent

coding & agent channel →