Harness optimization doubles Qwen solutions on SWE-bench

rohanpaul_ai · x · 2026-08-31

This paper studies how changing the harness affects fixed-model task results. It compares keeping full history vs. compressing old outputs/handling stalls. Under 20k token limits, the optimized config boosted Qwen2.5-Coder's complete solutions on SWE-bench Verified from 43 to 72, significantly improving F2PF metrics. This proves the critical role of engineering architecture in agent performance.

Related event: Harness Optimization Doubles Qwen's SWE-bench Performance(2 posts)→

Original post →

More from coding & agent

coding & agent channel →