Agent harnesses show vast cost spreads: mini-swe hits best accuracy near direct-inference cost
xiye_nlp · x · 2026-10-01
A new evaluation across long-context agent harnesses shows more compute does NOT always mean better accuracy — similar performance comes at vastly different costs.
Key points:
- Cost varies widely across different harnesses even with the same LM
- mini-swe stands out, achieving the best performance at a cost close to direct inference
- The takeaway: orchestration/harness engineering, not raw compute, drives cost-efficiency
The results come from the author's LongHarness benchmark evaluation.
Related event: LongHarness benchmark reveals 10x efficiency gaps across agent harnesses(2 posts)→
More from coding & agent
- headcount: Open-Source Repo Gives Your AI Agent a 16-Department, 172-Skill Company — alex_verem · 2026-10-01
- Open-source 'headcount' gives your AI agent 172 specialists across 16 departments, free — alex_verem · 2026-10-01
- AI coding agents write insecure Supabase RLS policies — dev's CLI scan finds 12 high-severity issues — Real_KingZeotic · 2026-10-01
- AI Agent Polled Every 15 Minutes and Snagged a Fully-Booked 6-Seat Tokyo Restaurant in 13 Hours — armand_ruiz · 2026-10-01
- Higgsfield launches ChatGPT extension bringing computer use into Codex — SimplyAnnisa · 2026-10-01
- LibraryDesignBench tests whether AI agents can design and effectively use code libraries — a1zhang · 2026-10-01