Berkeley PhD Candidate Compares Claude Code vs Codex vs Pi: Harness Choice Hits Cost More Than Accuracy
arena · x · 2026-09-29
Arena released a video in which Melissa Pan, a PhD candidate at UC Berkeley's Sky Computing Lab and former Arena intern, benchmarks three coding agents — Claude Code, Codex, and Pi. She introduces the concept of the hidden "harness tax": how the system scaffolding around a model affects its cost and performance.
She reports three surprising findings, the first being that harness choice impacts cost more than accuracy — the same model wrapped in different agent setups can differ far more in spend than in quality. The full video covers all three findings and their implications for building useful coding agents on realistic budgets.
More from coding & agent
- Turning coding-agent failures into regression tests with Kitaru session replay — strickvl · 2026-09-29
- Open-source Collaborator puts terminals, context files and code on one infinite canvas for agent dev — adnan_hashmi · 2026-09-29
- Running Codex for 5 days to enumerate every published AI safety idea — AaronBergman18 · 2026-09-29
- Sonnet 5.5 Lands in Conductor, and the Benchmark Chart Is Surprising — charlieholtz · 2026-09-29
- Agents With a Tailscale-Like VPN? The Coming Cat-and-Mouse Over Agent Browsing — alexisgallagher · 2026-09-29
- Dev benchmarks agent harnesses, says opencode beats PI, OpenClaw and Hermes by a huge margin — dh7net · 2026-09-29