UC Berkeley's HarnessTax study: agent harness barely changes success rate but costs vary 5x

solyarisoftware · x · 2026-09-22

A UC Berkeley team published HarnessTax, quantifying how much a coding agent's harness really matters. The setup: 21 model-harness pairs across 7 models, SWE-bench Lite and Terminal-Bench 2.0, 30 tasks per benchmark, 3 repetitions, with 95% bootstrap confidence intervals.

Key findings:

Related event: Harness Choice Barely Affects Success Rate but Can Quintuple Coding Agent Costs(3 posts)→

Original post →

More from coding & agent

coding & agent channel →