Harness Swing Hits 11%: Scaffolding Matters as Much as the Model in Long-Horizon Tasks
testingcatalog · x · 2026-08-29
Data analysis reveals that the same model (e.g., GPT-5.6 Luna) can swing by 11 percentage points (33.6% vs. 44.9%) across three different harnesses. This gap is wider than the difference between 2nd and 8th place. It suggests that on long-horizon commerce tasks, the scaffolding around a model impacts results about as much as the model itself. Humans currently make key decisions while agents execute a share of tasks.
More from coding & agent
- LangChain Adds MCP Support in Open Source, Built on FastMCP — LangChain · 2026-08-29
- Dev built his own provider-agnostic artifact hosting after Claude's sharing limits — miihr_ · 2026-08-29
- Open-source 'universal pipe' connects cloud, local and self-hosted LLMs for free — conifer_v11 · 2026-08-29
- Better Models Need Good Design: 4 Levers for Coding Agents — rajistics · 2026-08-29
- LangChain Academy Hosting Live Workshop on Building Deep Agents — LangChain · 2026-08-29
- Building PromptTrail to Undo Single-Prompt Changes in Lovable Without Git — Mueller96 · 2026-08-29