SWE-bench Test: Swapping Agent Harnesses Boosts Coding Scores More Than New Models
_lewtun · x · 2026-08-07
A developer tested 10 mainstream coding agent harnesses (e.g., Codex, Claude Code) on SWE-bench Pro across different models. The findings reveal that the choice of harness significantly impacts final scores, often more than upgrading the model itself.
- Massive score variance: Swapping harnesses boosted pass@1 from 23% to 52% on GLM-5.2, and from 15% to 36% on Gemma 4 26B-A4B.
- Rankings don't transfer: Harness performance is highly model-dependent. Vendor-specific harnesses (Codex, Claude Code) rank high with large models but drop sharply with smaller ones, while model-agnostic harnesses (crush, opencode) climb. The rank correlation between the two models' leaderboards is just -0.05.
This suggests the AI coding field may be over-indexing on tuning model weights while underestimating the engineering potential of the surrounding agent harness.
Related event: SWE-bench Tests Show Switching Agent Frameworks Beats Swapping Models(2 posts)→
More from coding & agent
- Vercel Launches Passport: Enterprise Identity Authentication for AI Agents — brandon_galang · 2026-08-09
- Mold: An Open-Source Rust-Based Lightweight Image Generator — brinkjames · 2026-08-09
- Stopping Free Trial Abuse: Using ML to Flag Duplicate Accounts — cneuralnetwork · 2026-08-09
- Amazon's $1.8M Agent Blunder: Tech Giants Scramble to Cap Exploding AI Costs — 量子位 · 2026-08-09
- Open-source Agent Skill transforms story text into hand-drawn vertical animations — aigclink · 2026-08-09
- ComfyUI Node for MiniMax H3 Latent Upscaler Released — Herooftermina1998 · 2026-08-09