The harness tax: Claude Code costs 2x Pi at the same 15.3 turns, with 10x initial context on SWE-bench Lite
HamelHusain · x · 2026-09-17
Melissa Pan (RT'd by Hamel Husain) quantifies the "harness tax": on SWE-bench Lite with Fable 5, Pi and Claude Code average nearly identical turns per attempt (15.4 vs 15.3), yet Claude Code costs about twice as much for only a 1.1-percentage-point success gain. The gap is visible at the first model call — Claude Code's mean initial context is over 10x Pi's, from longer instructions and larger tool schemas. Her takeaway: as models improve, less scaffolding is needed, and harness design should prioritize cost efficiency and reliability.
More from coding & agent
- Common ML pitfall: trusting external benchmarks over product-grounded evals — yunta_tsai · 2026-09-17
- PostHog: Agent-Opened PRs Jumped From 20% to 70%, So What Do Engineers Do? — rseroter · 2026-09-17
- Non-developer builds GreekSoup, an open-source AI equity research desk, mostly by prompting Claude — Practical-Rise-1188 · 2026-09-17
- Google ships full agent lifecycle stack: context layers, self-heal, one-command deploy — blaizedsouza · 2026-09-17
- tldraw split its mascot favicon into three and serves one at random per page load — max__drake · 2026-09-17
- Stop writing LLM instructions in prose: structure criteria as JSON instead — iamrobotbear · 2026-09-17