Same model, 17x cost gap: FrontierHarness benchmarks 9 agent harnesses across 360 runs
Double-Entertainer62 · reddit · 2026-09-10
FrontierHarness tested 9 agent harnesses (Pi, Exo, Claude Code, Codex, DeepSeek Harness and more) on the same model, tasks and runtime: 360 runs, 2B tokens. Pass rates ranged 50–67%; Claude Code and DSH Creator tied at 19/30. Median cost per pass varied from $1.05 to $18.34 — a 17x gap. Key insight: cheap successful runs can hide expensive workflows; failed attempts, retries and manual cleanup count toward total spend, so track total spend divided by accepted tasks, plus time and human intervention. Recommendations: Codex for default use, Pi for high-volume repeated jobs, Exo when retries are cheap, DSH for wall-clock speed.
More from coding & agent
- kafka-mcp: An MCP Server Letting LLM Agents Inspect Kafka Topics and Safely Reset Offsets — modelcontextprotocol · 2026-09-10
- AI agents are building a new infra layer that leaves existing enterprise systems behind — matt_slotnick · 2026-09-10
- If SORs don't adapt to agents, a new platform layer will be built on top of them — matt_slotnick · 2026-09-10
- Agents will shift power to new layers above existing enterprise systems — matt_slotnick · 2026-09-10
- Agent infrastructure is where spend happens — and incumbents have no path in — matt_slotnick · 2026-09-10
- System of Records aren't built for agent-native patterns yet, argues ex-employee — matt_slotnick · 2026-09-10