The harness tax: Claude Code costs 2x Pi at the same 15.3 turns, with 10x initial context on SWE-bench Lite

HamelHusain · x · 2026-09-17

Melissa Pan (RT'd by Hamel Husain) quantifies the "harness tax": on SWE-bench Lite with Fable 5, Pi and Claude Code average nearly identical turns per attempt (15.4 vs 15.3), yet Claude Code costs about twice as much for only a 1.1-percentage-point success gain. The gap is visible at the first model call — Claude Code's mean initial context is over 10x Pi's, from longer instructions and larger tool schemas. Her takeaway: as models improve, less scaffolding is needed, and harness design should prioritize cost efficiency and reliability.

Related event: Benchmark Reveals 'Harness Tax': Coding Agent Frameworks Barely Affect Accuracy but Double Cost(7 posts)→

Original post →

More from coding & agent

coding & agent channel →