SWE-bench Test: Swapping Agent Harnesses Boosts Coding Scores More Than New Models

_lewtun · x · 2026-08-07

A developer tested 10 mainstream coding agent harnesses (e.g., Codex, Claude Code) on SWE-bench Pro across different models. The findings reveal that the choice of harness significantly impacts final scores, often more than upgrading the model itself.

This suggests the AI coding field may be over-indexing on tuning model weights while underestimating the engineering potential of the surrounding agent harness.

Related event: SWE-bench Tests Show Switching Agent Frameworks Beats Swapping Models(2 posts)→

Original post →

More from coding & agent

coding & agent channel →