LongHarness benchmark: same LM, >10x efficiency gaps across long-context agent harnesses

xiye_nlp · x · 2026-10-01

The author introduces LongHarness, a challenging benchmark for evaluating long-context harnesses like RLMs and coding agents, measuring both accuracy and efficiency.

Key findings:

The benchmark provides a quantitative tool for studying how harness design affects long-context task performance and cost.

Related event: LongHarness benchmark reveals 10x efficiency gaps across agent harnesses(2 posts)→

Original post →

More from coding & agent

coding & agent channel →