Harvey's rebuilt harness lifts 7-model average from 23.3% to 62.4% while cutting cost 60%
rajistics · x · 2026-10-06
Rajiv Shah's ODSC piece "The Unreasonable Effectiveness of the Harness" makes the case with two examples:
- Harvey (M&A diligence): rebuilding the harness raised average rubric pass rates across seven models from 23.3% to 62.4%. The design loads an entire data room into a Python REPL, with a root agent delegating bounded reviews to sub-agents in their own context windows. Claude Opus 5 cost per data room fell from $18 to $7 while scoring higher.
- OpenAI (ARC-AGI-3): GPT-5.6 Sol went from 13.3% to 38.3% using one-sixth of the output tokens, via changes to cross-turn reasoning and compaction defaults.
Shah argues the harness controls what the model sees, which tools it can use, what it remembers, and when it stops — engineering tradeoffs worth optimizing. He'll teach this at ODSC West (Oct 28) and the MLOps Conference in November.
Related event: Harness Engineering, Not Models, Is the Agent Bottleneck: Study(2 posts)→
More from coding & agent
- Using Claude Design to prototype complex interactions beats static mocks — austin_malerba · 2026-10-06
- AI agent posts user's bank balances to company Slack, sparking agent paradigm debate — altryne · 2026-10-06
- The handoff test: approve, revoke, transfer to a fresh agent—does human authority survive? — tallmetommy · 2026-10-06
- Beam agent one-shots a full end-to-end Unsloth training pipeline in OpenCode with a single prompt — bhutanisanyam1 · 2026-10-06
- Open-source 'Clay killer' launched: 85% cheaper, top people-search accuracy, 25x faster — Scobleizer · 2026-10-06
- New Obsidian plugin qiaomu-ui-learn trains your vibe-coding UI taste with copyable prompts — vista8 · 2026-10-06