New study: more compute doesn't mean better accuracy across agent harnesses
xiye_nlp · x · 2026-10-01
A research team led by xiyenlp's students published a paper and website systematically analyzing cost vs. performance across agent harnesses for SWE-bench-style evaluation. Key findings: costs vary widely across harnesses for similar accuracy, more compute does not always mean better accuracy, and mini-swe stands out by achieving the best performance at a cost close to direct inference. For agent engineers, choosing the right harness is itself a major cost lever.
Related event: Same model, different agent harnesses: SWE-bench gap from 61% to 75%(2 posts)→
More from coding & agent
- MiniMax Teases MiniMax Code: A Fresh Start, Not Just for Code — iamaliveix · 2026-10-01
- Open-source 'headcount' gives your AI agent 172 specialists across 16 departments, free — alex_verem · 2026-10-01
- AI coding agents write insecure Supabase RLS policies — dev's CLI scan finds 12 high-severity issues — Real_KingZeotic · 2026-10-01
- AI Agent Polled Every 15 Minutes and Snagged a Fully-Booked 6-Seat Tokyo Restaurant in 13 Hours — armand_ruiz · 2026-10-01
- Higgsfield launches ChatGPT extension bringing computer use into Codex — SimplyAnnisa · 2026-10-01
- LibraryDesignBench tests whether AI agents can design and effectively use code libraries — a1zhang · 2026-10-01