LongHarness benchmark: same LM, >10x efficiency gaps across long-context agent harnesses
xiye_nlp · x · 2026-10-01
The author introduces LongHarness, a challenging benchmark for evaluating long-context harnesses like RLMs and coding agents, measuring both accuracy and efficiency.
Key findings:
- Clear accuracy gaps across harnesses
- >10x efficiency differences even with the same underlying LM and similar accuracy
- Corollary: more compute does not always mean better results — orchestration matters
The benchmark provides a quantitative tool for studying how harness design affects long-context task performance and cost.
Related event: LongHarness benchmark reveals 10x efficiency gaps across agent harnesses(2 posts)→
More from coding & agent
- MiniMax Teases MiniMax Code: A Fresh Start, Not Just for Code — iamaliveix · 2026-10-01
- Open-source 'headcount' gives your AI agent 172 specialists across 16 departments, free — alex_verem · 2026-10-01
- AI coding agents write insecure Supabase RLS policies — dev's CLI scan finds 12 high-severity issues — Real_KingZeotic · 2026-10-01
- AI Agent Polled Every 15 Minutes and Snagged a Fully-Booked 6-Seat Tokyo Restaurant in 13 Hours — armand_ruiz · 2026-10-01
- Higgsfield launches ChatGPT extension bringing computer use into Codex — SimplyAnnisa · 2026-10-01
- LibraryDesignBench tests whether AI agents can design and effectively use code libraries — a1zhang · 2026-10-01