Study finds Agent harness impacts benchmark scores more than the model itself

rohanpaul_ai · x · 2026-08-26

A survey of AI agents in command-line environments reveals that the harness (framework) around a model often determines benchmark scores more than the model itself. The study suggests treating scaffold selection as a first-order engineering decision, finding that stronger model variants in the same harness yielded no extra resolved tasks, only lower latency.

Paper: "Terminal Agents: A Survey of AI Agents in Command-Line Environments" (arxiv.org/abs/2608.20485).

Original post →

More from coding & agent

coding & agent channel →