LongHarness benchmark reveals 10x efficiency gaps across agent harnesses
The new LongHarness benchmark evaluates long-context agent harnesses on both accuracy and efficiency, finding over 10x gaps between harnesses on the same model, with mini-swe achieving the best results at near-direct-inference cost.
2026-10-01 ~ 2026-10-01 · 2 related posts
- LongHarness benchmark: same LM, >10x efficiency gaps across long-context agent harnesses — xiye_nlp · 2026-10-01
- Agent harnesses show vast cost spreads: mini-swe hits best accuracy near direct-inference cost — xiye_nlp · 2026-10-01