LongHarness benchmark reveals 10x efficiency gaps across agent harnesses

The new LongHarness benchmark evaluates long-context agent harnesses on both accuracy and efficiency, finding over 10x gaps between harnesses on the same model, with mini-swe achieving the best results at near-direct-inference cost.

2026-10-01 ~ 2026-10-01 · 2 related posts