Agent harness study: prompts and tool schemas drive 3x cost gap on identical models
dair_ai · x · 2026-10-08
A paper shared by dair-ai systematically compares agent harnesses with the model held fixed.
- Key finding: swapping the harness (Claude Code, mini-SWE-agent, OpenCode) changes results about as much as rerunning the experiment. On a 45-task hard set both flipped 13% of tasks; Claude Code and mini-SWE-agent scored within 5 points across 447 SWE-bench Verified tasks.
- Cost source: system prompts and tool schemas are resent at every agent step—the more steps, the more you pay.
- Practical takeaway: keep your harness prompt and tools small—this cut costs up to 3x without lowering accuracy.
Related event: Swapping agent harnesses reshapes results: a systematic study(7 posts)→
More from coding & agent
- Jira shifts from planned to captured work with real-time AI session tracking — davidhoang · 2026-10-08
- Datalab rebuilds chart understanding: 93% fewer wrong values, 3x faster at same $3/1k pages — VikParuchuri · 2026-10-08
- Tencent's WorkForge scales verifiable training environments for long-horizon work agents — teortaxesTex · 2026-10-08
- Factory AI agent now assignable as a teammate inside Atlassian Jira — matanSF · 2026-10-08
- Jira shifts from planned work to captured work with real-time AI session tracking — davidhoang · 2026-10-08
- River open recipe: post-trained open models beat GPT-6 Astra Pro and Opus 5.5 on text-to-SQL for under 1% of cost — kylekosic · 2026-10-08