LangChain's Chase: Agent Gains Come from Harness, Not Model — Paper Shows 14pt Lift
hwchase17 · x · 2026-09-10
hwchase17 endorses an EMNLP 2026 paper arguing agent improvement is harness improvement, not model improvement — the highest-value interventions sit at the tool boundary: what context you pass, when tools are provided, how you recover from failure, and what gets measured afterward.
The paper formalizes harness optimization around a fixed model as budgeted selection, with edits as guarded intercepts at the tool boundary. Its PRISM method clusters failures and routes repairs to prompt, middleware, or joint edit surfaces, selecting candidates by both gate pass-rate and reliability (RelLift95).
On BFCL multi-round, tau2-Retail, and tau2-Telecom, PRISM delivers mean held-out lifts of 14.2, 14.9, and 10.1 percentage points with positive reliable lift across all three; ablations credit failure-surface routing and edit-pattern constraints.
More from coding & agent
- Autoresearch Loop with Tinker Reproduces Self-Distillation Papers at Predictable Cost — SRSchmidgall · 2026-09-10
- Together AI breaks down the open-source stack for agentic coding, layer by layer — togethercompute · 2026-09-10
- Researcher Lets Codex Run the Experiments, Publishes Recurrent Model Length-Extrapolation Paper — qixing_huang · 2026-09-10
- Gradle's eval harness pits Fable 5.1 vs Astra 6 on one bug—Astra 6 faster and cheaper — Party_Till_I_Die · 2026-09-10
- Yes, your AI assistant can leak one customer's data to another — here's the fix — ericelliott_ · 2026-09-10
- Perplexity Computer adds desktop/mobile previews as web-app usage climbs — inductionheads · 2026-09-10