Harness optimization lifts Harvey legal agent benchmark pass rate from 67.1% to 85.9%
sarahookr · x · 2026-09-10
- New research from Adaption AI (Sudip Roy and Dhruv Batra's team): harness optimization alone moved criterion pass rate from 67.10% to 85.92% on the Harvey Legal Agent Benchmark — a jump of nearly 19 points.
- Adding post-training on top pushed the score to 88.03%.
- Key argument: owning your intelligence means improving both the model and the harness around it. The result quantifies how much leverage agent orchestration/engineering has on end performance.
Related event: Harness tuning boosts legal agent benchmark pass rate from 67% to 86%(2 posts)→
More from coding & agent
- Agent builds, compiles and debug-deploys an iPhone app from Linux on an M1 Mac — alexcovo_eth · 2026-09-10
- OpenAI's cached-internet approach in Navier Stokes agents seen as fix for sandbox escapes — nrehiew_ · 2026-09-10
- The Kernel That Must Say No: how Semantic Kernel gates AI tool execution — Mahmoud_Zalt · 2026-09-10
- GPT-6 Astra vs GPT-5.6 Sol: Code Review Benchmark on 50 Real PRs — entelligenceai17 · 2026-09-10
- An "anti Claude-speak" system prompt: active voice, no em dashes, no startup clichés — dressinbrass · 2026-09-10
- Claude reliably finds holes in CAD models and places screws on command — _Stocko_ · 2026-09-10