Same model, 17x cost gap: FrontierHarness benchmarks 9 agent harnesses across 360 runs

Double-Entertainer62 · reddit · 2026-09-10

FrontierHarness tested 9 agent harnesses (Pi, Exo, Claude Code, Codex, DeepSeek Harness and more) on the same model, tasks and runtime: 360 runs, 2B tokens. Pass rates ranged 50–67%; Claude Code and DSH Creator tied at 19/30. Median cost per pass varied from $1.05 to $18.34 — a 17x gap. Key insight: cheap successful runs can hide expensive workflows; failed attempts, retries and manual cleanup count toward total spend, so track total spend divided by accepted tasks, plus time and human intervention. Recommendations: Codex for default use, Pi for high-volume repeated jobs, Exo when retries are cheap, DSH for wall-clock speed.

Original post →

More from coding & agent

coding & agent channel →