Berkeley study: open-source agent harness beats Claude Code and Codex CLI 75% of the time
solyarisoftware · x · 2026-09-22
A UC Berkeley and Arena study dubbed the "Harness Tax" benchmarked 21 model-agent pairs on SWE-bench Lite and Terminal-Bench 2.0, comparing Claude Code, Codex CLI, and a minimal open-source harness called Pi.
Key findings:
- Frontier models performed better on the barebones open-source framework about 75% of the time;
- Switching to Pi could reportedly cut agent API costs by roughly 50% with no benchmark performance drop;
- The results challenge the "in-house myth" that vendor-built agent frameworks are superior, suggesting users are paying an unnecessary tax for official harnesses.
The claim circulates via a retweet; verify specifics against the original paper.
Related event: Berkeley Study: Agent Harness Choice Can 5x Your Bill(4 posts)→
More from coding & agent
- Storewake MCP server puts App Store rankings, RevenueCat revenue and Apple Ads into your AI assistant for $19/mo — Pfernan95 · 2026-09-22
- Free workshop walks through building an LLM Wiki for agent long-term memory — Al_Grigor · 2026-09-22
- Un-fusing a realtime voice stack (STT → LLM → TTS) cut costs 14x — and the real win was text-level guardrails — Cloudsurfer_90 · 2026-09-22
- The last mile of agents: render structured output with Gamma instead of dumping JSON — raw-hit10 · 2026-09-22
- Fan-made Ado chibi pet released for the OpenAI Codex desktop app — secemp9 · 2026-09-22
- Opinion: SaaS becomes the harness and window into agents — guardrails plus visibility — StatisticianKey7858 · 2026-09-22