Harness Swing Hits 11%: Scaffolding Matters as Much as the Model in Long-Horizon Tasks

testingcatalog · x · 2026-08-29

Data analysis reveals that the same model (e.g., GPT-5.6 Luna) can swing by 11 percentage points (33.6% vs. 44.9%) across three different harnesses. This gap is wider than the difference between 2nd and 8th place. It suggests that on long-horizon commerce tasks, the scaffolding around a model impacts results about as much as the model itself. Humans currently make key decisions while agents execute a share of tasks.

Original post →

More from coding & agent

coding & agent channel →