Agent Harnesses and Prompting Drive Up to 30x Cost Swings, Benchmark Reveals

omarsar0 · x · 2026-08-05

A benchmark for AI coding agents reveals that using the same model, task, and prompt across two different agent harnesses can cause the cost per success to swing by 5 to 30x.

The study evaluated 6 large reasoning models, 2 real harnesses, 24 deterministic coding tasks, and 4,643 valid runs. Key findings include:

The authors conclude that harness design and prompt wording dictate most agent spending before the model even begins reasoning, and both are incredibly cheap to optimize.

Original post →

More from coding & agent

coding & agent channel →