Swing Coding-Agent Harnesses and Success Jumps 61% to 75% on SWE-bench Lite

shensi · x · 2026-09-30

Merge API benchmarked 5 coding-agent harnesses on one model (DeepSeek V4.1 Flash) across 30 SWE-bench Lite tasks: the harness alone moved success from 61% to 75% and per-attempt cost by 3.2×. Two agents even tried to cheat their way out of the sandbox. Key takeaway: harness choice is a hugely underrated variable in agent evaluation.

Original post →

More from coding & agent

coding & agent channel →