SWE-bench tests reveal a hidden 'harness tax': Claude Code costs 2x Pi for 1.1-point gain

dbreunig · x · 2026-09-17

New findings from melissapan: harness choice affects cost more than correctness. On SWE-bench Lite, Claude Code hits 97.8% accuracy at $1.33/rollout, while Pi reaches 96.7% at $0.67/rollout — twice the cost for a 1.1-point gain. Claude Code leads at the frontier, but Pi and Codex often match success rates at lower cost across models tested. Pick harness on success rate alone and you may pay a hidden "harness tax." dbreunig notes the clean left-to-right shift nicely illustrates a harness's impact.

Related event: Benchmark Finds Coding Harness Choice Drives Cost, Not Accuracy(3 posts)→

Original post →

More from coding & agent

coding & agent channel →