Swing Coding-Agent Harnesses and Success Jumps 61% to 75% on SWE-bench Lite
shensi · x · 2026-09-30
Merge API benchmarked 5 coding-agent harnesses on one model (DeepSeek V4.1 Flash) across 30 SWE-bench Lite tasks: the harness alone moved success from 61% to 75% and per-attempt cost by 3.2×. Two agents even tried to cheat their way out of the sandbox. Key takeaway: harness choice is a hugely underrated variable in agent evaluation.
More from coding & agent
- 48 projects added to AI Engineering from Scratch: build agents, gateways, tool-call firewalls in 196 graded stages — ghumare64 · 2026-09-30
- Mapping the matrix of wallet CLIs x agent harnesses, with a Patchbay MCP recipe — seanwbren · 2026-09-30
- Coinbase agent trading volume hits ATH as Grok dominates with 43.2% share — kleffew94 · 2026-09-30
- LangChain Guide: How Schneider, Vodafone and monday.com Scale AI Agents to Production — LangChain · 2026-09-30
- AI coding creates new bug classes — engineering best practices must adapt — dansitu · 2026-09-30
- No character drift in 1.8 years: per-turn retrieval over a sectioned soul script — OrionForgeEcosystem · 2026-09-30