FULL STORY
HarnessTax: Agent Frameworks Drive Cost, Not Success
A UC Berkeley-backed evaluation, HarnessTax, tested 7 models across 3 coding agent harnesses. The study went viral on X and Hacker News, finding frameworks barely affect success rates but can multiply costs fivefold.
2026-09-17 ~ 2026-09-22 · 3 episodes · 12 posts
Episode 1 · 7 Models × 3 Coding Harnesses: Harness Choice Drives Cost, Not Success (2026-09-17, 7 posts)
Melissa Pan's team benchmarked 7 models across Claude Code, Codex, and Pi coding agent harnesses. The counterintuitive finding: harness choice barely affects task success rates but significantly affects costs, and a model's native harness isn't necessarily optimal. The results were widely shared by Matei Zaharia, Hamel Husain, and others.
Confirmed
- The evaluation covered 7 models × 3 harnesses (Claude Code, Codex, Pi); harness choice had little impact on success rates but large impact on cost
- Per dbreunig citing melissapan's data, on SWE-bench Lite Claude Code achieved 97.8% success at $1.33 per rollout, while Pi hit 96.7% at roughly half the cost — just a 1.1pp gap
- Quantitative analysis shared by Hamel Husain: running Fable 5, Pi and Claude Code averaged nearly identical turns (15.4 vs 15.3), yet cost differed by about 2x
Why it matters
- Direct relevance for team tooling decisions: harness cost differences reach 2x while success differences are on the order of 1 percentage point
- Suggests evaluations and procurement decisions shouldn't default to the model vendor's native harness; simpler harnesses can materially cut agent operating costs
- Attention from DavideCrapis, CShorten30, blaizedsouza, mateizaharia, and Hamel Husain shows the finding resonated across the community
- 7 models across Claude Code, Codex, Pi: harness choice drives cost, not success — matei_zaharia · 2026-09-17
- Testing 7 models across Claude Code, Codex and Pi: native harness isn't always best — CShorten30 · 2026-09-17
- SWE-bench tests reveal a hidden 'harness tax': Claude Code costs 2x Pi for 1.1-point gain — dbreunig · 2026-09-17
- Evaluating 7 models across Claude Code, Codex, and Pi: harness choice drives cost, not success rate — CShorten30 · 2026-09-17
- Study of 7 models across Claude Code, Codex, Pi: harness barely affects success but swings cost — DavideCrapis · 2026-09-17
- The harness tax: Claude Code costs 2x Pi at the same 15.3 turns, with 10x initial context on SWE-bench Lite — HamelHusain · 2026-09-17
- 7 Models, 3 Coding Harnesses: Native Harness Isn't Always Best and Costs Vary Widely — blaizedsouza · 2026-09-17
Episode 2 · Benchmark of 21 model-harness combos: agent framework changes cost, not success rate (2026-09-19, 2 posts)
A viral study benchmarked 7 models across 3 coding harnesses (21 combos), finding that the choice of agent framework barely affects success rates but dramatically changes cost.
- 21 model-harness pairs tested: framework choice barely moves success rate but swings cost — iScienceLuvr · 2026-09-19
- Evaluating 7 Models Across Claude Code, Codex, and Pi: Harness Choice Drives Cost, Not Success — CShorten30 · 2026-09-20
Episode 3 · Harness Choice Barely Affects Success Rate but Can Quintuple Coding Agent Costs (2026-09-21, 3 posts)
UC Berkeley's HarnessTax study tested 21 model-harness combinations across 60 tasks and found that while success rates remain similar, costs can differ up to 5x depending on the harness. The findings suggest harness selection is a key lever for cutting coding agent costs.
- Arena: coding-agent harnesses show up to 5x cost differences at similar success rates — thione · 2026-09-21
- Same model bills 5x more in a different harness: 21 combos tested across 60 tasks — CShorten30 · 2026-09-22
- UC Berkeley's HarnessTax study: agent harness barely changes success rate but costs vary 5x — solyarisoftware · 2026-09-22