Same model bills 5x more in a different harness: 21 combos tested across 60 tasks
CShorten30 · x · 2026-09-22
HarnessTax tested 21 model-harness combos across 60 tasks and found cost moved far more than success rate did: $1.33 vs $0.67 per attempt for the same model, while success ranged only from 96.7% to 97.8%. Nine of twelve lean setups beat vendor defaults. Claude Code and Pi averaged 15 turns per task, but heavy setups carried over 10x the initial context since longer instructions and larger tool definitions ride along on every call. Co-author melissapan notes harness choices run on tribal knowledge, making it unclear which setup wins on cost for your workload.
More from coding & agent
- Foundation Capital: enterprise agents need identity, boundaries — and the web rebuilt for machines — brucemacv · 2026-09-22
- Question's Gambit lifts deep research agents: GPT-5.5 hits 90.5% on BrowseComp-Plus — omarsar0 · 2026-09-22
- Open-Source Tool Maps E-Drums to Ableton Kits, Designed for Coding Agents to Extend — _Dave__White_ · 2026-09-22
- Chollet: delegate coding, never delegate understanding — beware 'conceptual debt' — round · 2026-09-22
- Dev open-sources modsure npm package for LLM-powered social media moderation — RealDannyhvv · 2026-09-22
- TypeSafe's 2-stage extraction cascade: big-model quality at a fraction of the cost — TheMoonMidas · 2026-09-22