Paper reveals Agent benchmark unreliability: Harness variance 7.8x model variance
omarsar0 · x · 2026-08-26
A new paper investigates why agent leaderboard comparisons are hard to trust, focusing on how much of a benchmark score is attributed to the "harness"—the layer between the model and the task.
The harness builds context, mediates tool calls, validates outputs, and decides when to stop. In a controlled experiment on 100 tasks from SWE-bench Verified, researchers tested three frontier models against three harness configurations while holding other variables constant.
Results showed that swapping the harness caused GLM-5.1's score to swing by 13.0 points, whereas swapping the model within a fixed harness caused changes of only 3.0, 2.5, and 5.0 points. Harness-induced variance was found to be 7.8x larger than model-induced variance, with 6 out of 9 model-pair rankings flipping depending on the harness used.
More from coding & agent
- LongRCA Bench: Diagnosing Failures in Long-Horizon Agent Trajectories — Yunfei Zhang · 2026-08-26
- Open-Source Guaardvark Simplifies ComfyUI with Voice Chat and MCP Integration — llama-of-death · 2026-08-26
- OpenAI: KV Cache is the largest and fastest-growing data structure in agentic inference — BenBajarin · 2026-08-26
- Claude Code Frontend Design Toolkit: 70+ Skills, Plugins and MCP Servers to Kill AI Slop — tom_doerr · 2026-08-26
- An "Artificial Civilization Scaffold" Could Make AI Smarter Without Any Retraining — New_User_1970 · 2026-08-26
- Idea: Build a Social Network Where Agents Roast and Collaborate — RileyRalmuto · 2026-08-26