Same models, new harness: scores jump 23% to 62%, says Rajiv Shah
rajistics · x · 2026-10-06
Rajiv Shah argues the biggest agent performance gains now come from harness engineering, not models: identical models scored 23% vs 62% under different harnesses.
His hands-on ODSC workshop tests six common claims, including "harnesses don't matter," "more tools make a better agent," "more instructions and memory improve performance," "always use the strongest model," "just give the agent an objective," and "you need a multi-agent system."
Related event: Harness Engineering, Not Models, Is the Agent Bottleneck: Study(2 posts)→
More from coding & agent
- Using Claude Design to prototype complex interactions beats static mocks — austin_malerba · 2026-10-06
- AI agent posts user's bank balances to company Slack, sparking agent paradigm debate — altryne · 2026-10-06
- The handoff test: approve, revoke, transfer to a fresh agent—does human authority survive? — tallmetommy · 2026-10-06
- Beam agent one-shots a full end-to-end Unsloth training pipeline in OpenCode with a single prompt — bhutanisanyam1 · 2026-10-06
- Open-source 'Clay killer' launched: 85% cheaper, top people-search accuracy, 25x faster — Scobleizer · 2026-10-06
- New Obsidian plugin qiaomu-ui-learn trains your vibe-coding UI taste with copyable prompts — vista8 · 2026-10-06