Evaluating 7 Models Across Claude Code, Codex, and Pi: Harness Choice Drives Cost, Not Success
CShorten30 · x · 2026-09-20
A coding agent evaluation (trending on Hacker News as HarnessTax) benchmarks 7 models across Claude Code, Codex, and Pi, yielding three surprising findings:
- Harness choice has little effect on task success rate, but significantly affects cost — a wrong harness burns money without hurting quality;
- A simple harness can be competitive, so elaborate orchestration may not pay off;
- The model's native harness isn't always the best option.
Practical takeaway for developers: validate simple harnesses on cost/benefit before stacking complex agent frameworks.
More from coding & agent
- Explore /goal in your coding harness: omarsar shares tips on custom agent harnesses — omarsar0 · 2026-09-20
- Building a per-turn goal verifier with Jev for a custom agent harness — omarsar0 · 2026-09-20
- Omarchy sparks a Linux desktop boom as 2026 dubbed 'the year of Linux desktop' — mark_k · 2026-09-20
- Agents abusing doc requests for arbitrary RCE and data exfiltration — zainhas · 2026-09-20
- Yacine says coding from his phone killed his Twitter addiction — yacineMTB · 2026-09-20
- Dev launches site to browse AI skill libraries with visual flow diagrams — spersingerorinda · 2026-09-20