Paper reveals Agent benchmark unreliability: Harness variance 7.8x model variance

omarsar0 · x · 2026-08-26

A new paper investigates why agent leaderboard comparisons are hard to trust, focusing on how much of a benchmark score is attributed to the "harness"—the layer between the model and the task.

The harness builds context, mediates tool calls, validates outputs, and decides when to stop. In a controlled experiment on 100 tasks from SWE-bench Verified, researchers tested three frontier models against three harness configurations while holding other variables constant.

Results showed that swapping the harness caused GLM-5.1's score to swing by 13.0 points, whereas swapping the model within a fixed harness caused changes of only 3.0, 2.5, and 5.0 points. Harness-induced variance was found to be 7.8x larger than model-induced variance, with 6 out of 9 model-pair rankings flipping depending on the harness used.

Original post →

More from coding & agent

coding & agent channel →