Paper: Harness-Induced Variance Is 7.8x Model Variance in LLM Agent Benchmarks

rohanpaul_ai · x · 2026-09-01

A new paper, "Stop Comparing LLM Agents Without Disclosing the Harness," argues that for long-horizon agents the harness can matter more than the model, so benchmark scores shouldn't be compared without disclosing or controlling the harness.

In a controlled 3-model × 3-harness study on 100 SWE-bench Verified tasks:

Paper: arxiv.org/abs/2605.23950

Original post →

More from coding & agent

coding & agent channel →