Agent harnesses show vast cost spreads: mini-swe hits best accuracy near direct-inference cost

xiye_nlp · x · 2026-10-01

A new evaluation across long-context agent harnesses shows more compute does NOT always mean better accuracy — similar performance comes at vastly different costs.

Key points:

The results come from the author's LongHarness benchmark evaluation.

Related event: LongHarness benchmark reveals 10x efficiency gaps across agent harnesses(2 posts)→

Original post →

More from coding & agent

coding & agent channel →