SWE-bench tests reveal a hidden 'harness tax': Claude Code costs 2x Pi for 1.1-point gain
dbreunig · x · 2026-09-17
New findings from melissapan: harness choice affects cost more than correctness. On SWE-bench Lite, Claude Code hits 97.8% accuracy at $1.33/rollout, while Pi reaches 96.7% at $0.67/rollout — twice the cost for a 1.1-point gain. Claude Code leads at the frontier, but Pi and Codex often match success rates at lower cost across models tested. Pick harness on success rate alone and you may pay a hidden "harness tax." dbreunig notes the clean left-to-right shift nicely illustrates a harness's impact.
Related event: Benchmark Finds Coding Harness Choice Drives Cost, Not Accuracy(3 posts)→
More from coding & agent
- Paper2Agent in Nature converts research papers into MCP-based AI agents automatically — ValerioCapraro · 2026-09-17
- Paper2Agent turns research papers into interactive AI agents, Nature paper shows — ValerioCapraro · 2026-09-17
- Enactra releases BuildingBench, a leaderboard scoring coding agents on 3D building reconstruction — Lianhuiq · 2026-09-17
- User burns 4 Codex banked resets in 30 minutes to reset 5-hour limits back-to-back — flowersslop · 2026-09-17
- Turso launches Accident Protection: restore agent-deleted databases for 5 days free — glcst · 2026-09-17
- Stand out in AI+DS in 2026: build a decision-making machine with RAG + agents — mdancho84 · 2026-09-17