Benchmark Finds Coding Harness Choice Drives Cost, Not Accuracy
A benchmark running 7 models across Claude Code, Codex, and Pi found harness choice barely affects success rates but greatly affects cost, with Claude Code costing up to twice as much as Pi for only about 1.1 points more accuracy on SWE-bench Lite.
2026-09-17 ~ 2026-09-17 · 3 related posts
- 7 models across Claude Code, Codex, Pi: harness choice drives cost, not success — matei_zaharia · 2026-09-17
- Testing 7 models across Claude Code, Codex and Pi: native harness isn't always best — CShorten30 · 2026-09-17
- SWE-bench tests reveal a hidden 'harness tax': Claude Code costs 2x Pi for 1.1-point gain — dbreunig · 2026-09-17