7 models across Claude Code, Codex, Pi: harness choice drives cost, not success
matei_zaharia · x · 2026-09-17
Melissa Pan's team evaluated 7 models on Claude Code, Codex, and Pi coding-agent harnesses, with three surprising findings: (1) harness choice has little effect on task success rate but significantly affects cost; (2) a simple harness can be competitive; (3) the native harness isn't always the best. Millions use coding agents, yet harness impact was previously unclear — full details in the thread.
More from coding & agent
- Vercel's 27.5KB WebGPU lexer model sparks wave: 5 browser-side tiny models in one week — iamrobotbear · 2026-09-17
- Mitchell Hashimoto's 'whiteboard defense' sets the bar for responsible AI coding — josh_wills · 2026-09-17
- OpenClaw creator to join Cloudflare Connect 2026 panel on agentic platforms — steipete · 2026-09-17
- Your LLM doesn't understand MCP — keeping the tool-call boundary clear makes agents easier to debug — gethackteam · 2026-09-17
- YC-backed Extend launches Parse Router to route each page to the right parsing engine — ycombinator · 2026-09-17
- He had Codex make a phone call to redeem a $500 gift card — and it worked — brandon_galang · 2026-09-17