Same models, new harness: scores jump 23% to 62%, says Rajiv Shah

rajistics · x · 2026-10-06

Rajiv Shah argues the biggest agent performance gains now come from harness engineering, not models: identical models scored 23% vs 62% under different harnesses.

His hands-on ODSC workshop tests six common claims, including "harnesses don't matter," "more tools make a better agent," "more instructions and memory improve performance," "always use the strongest model," "just give the agent an objective," and "you need a multi-agent system."

Related event: Harness Engineering, Not Models, Is the Agent Bottleneck: Study(2 posts)→

Original post →

More from coding & agent

coding & agent channel →