New study: more compute doesn't mean better accuracy across agent harnesses

xiye_nlp · x · 2026-10-01

A research team led by xiyenlp's students published a paper and website systematically analyzing cost vs. performance across agent harnesses for SWE-bench-style evaluation. Key findings: costs vary widely across harnesses for similar accuracy, more compute does not always mean better accuracy, and mini-swe stands out by achieving the best performance at a cost close to direct inference. For agent engineers, choosing the right harness is itself a major cost lever.

Related event: Same model, different agent harnesses: SWE-bench gap from 61% to 75%(2 posts)→

Original post →

More from coding & agent

coding & agent channel →