Berkeley study reveals the 'harness tax': harness choice impacts coding agent cost more than accuracy
arena · x · 2026-09-29
LMArena highlights new research by Melissa Pan, PhD candidate at UC Berkeley's Sky Computing Lab, benchmarking Claude Code vs Codex vs Pi.
She introduces the hidden "harness tax": the system surrounding an AI model materially changes its cost and performance. Among three surprising findings: harness choice impacts cost more than accuracy — meaning picking the right harness matters as much as the model itself when building coding agents on realistic budgets.
Related event: Berkeley Study: Coding Agent Harness Choice Drives Cost More Than Accuracy(2 posts)→
More from coding & agent
- Hackathon project Beethoven turns paintings into a live AI band playing via Lyria — schwentker · 2026-09-29
- OpenAI DevDay agenda leaks: Codex to get platform capabilities for plugins, agents and apps — testingcatalog · 2026-09-29
- Concept car fully generated in code, zero assets, one HTML file, built with Claude Sonnet 5.5 — techartist_ · 2026-09-29
- LLMs Keep Stalling on Lead Enrichment: Lazy Output and Hallucinated Emails — Royal_icey69 · 2026-09-29
- Auditing Thousands of Rollouts: 80%+ of Coding Agents Reason About an Imagined Grader — jonas__m · 2026-09-29
- 120 FPS in Browser: A Universal Decompiler Steps Closer to Reality — yacineMTB · 2026-09-29