Paper: Harness-Induced Variance Is 7.8x Model Variance in LLM Agent Benchmarks
rohanpaul_ai · x · 2026-09-01
A new paper, "Stop Comparing LLM Agents Without Disclosing the Harness," argues that for long-horizon agents the harness can matter more than the model, so benchmark scores shouldn't be compared without disclosing or controlling the harness.
In a controlled 3-model × 3-harness study on 100 SWE-bench Verified tasks:
- Average harness-induced variance was 7.80x model-induced variance.
- Holding the model fixed, moving from minimal to full harness changed pass@1 by 8.5–13.0 points; holding the harness fixed, switching models changed it by only 2.5–5.0 points.
- Rankings were unstable too: 6 of 9 model-pair/harness-pair comparisons reversed under another harness.
Paper: arxiv.org/abs/2605.23950
More from coding & agent
- Hamel Husain: everyone is building this AI agent — don't — hugobowne · 2026-09-03
- Dev accidentally burns $100 in an instant running ultracode AI coding mode — zsakib_ · 2026-09-03
- Hidden-bug eval across 105 issues: Fable 5.1 finds 43, none fixes all — cost per model compared — PawelHuryn · 2026-09-03
- Indie author builds an agent-native distribution layer for his novel with A2A endpoints — patternflow · 2026-09-03
- Bezalel gives AI agents memory, email, money and a cloud desktop behind one MCP URL — Rasmic · 2026-09-03
- jjk-explain turns any concept into a Jujutsu Kaisen-style explainer video with one Claude Code command — teortaxesTex · 2026-09-03